National Cyber Warfare Foundation (NCWF)

reverse-SynthID for studying the spectral structure of AI image watermarks


0 user ratings
2026-10-07 13:25:37
milo
Red Team (CNA)
"reverse-SynthID

reverse-SynthID reverse-engineers Google's invisible SynthID watermark through spectral analysis, building detectors and robustness research relevant to watermark designers and authorized red teams.








Toolaloshdenny/reverse-SynthID — a Python research project that discovers, detects, and stress-tests Google's SynthID image watermark via signal processing
CategoryAdversarial ML research / steganalysis and watermark robustness
Primary UseStudying the carrier-frequency structure of SynthID, building a 90%-accuracy independent detector, and evaluating watermark robustness against documented failure modes
Safe UseEducational and defensive research: authorized adversarial-ML evaluations, watermark-hardening studies, and lab analysis of provenance systems on images you own or generate yourself
Telemetry NoteDefenders should note the tool operates purely offline via numpy-style FFT analysis and leaves the SynthID carrier energy measurably degraded; content platforms can counter by tracking lossy re-encoding chains and multi-stage transform fingerprints

reverse-SynthID is a roughly 4.8k-star Python research repository that treats Google's SynthID watermark — the imperceptible pattern embedded in every image produced by Gemini models — as a signal-processing problem rather than a black box. The authors explicitly state they achieved everything without access to the proprietary encoder or decoder, using only spectral analysis of generated images. That framing matters for security professionals: this is a textbook reverse-engineering exercise against a provenance system, and the findings are directly useful to anyone designing, deploying, or defending AI-content watermarking.


The starting insight is visually striking. On a pure-white Gemini-generated image, the watermark is essentially the entire signal, so amplifying the high-frequency residual reveals a diagonal banding pattern — the watermark's spatial frequency signature. That carrier pattern, the README explains, is the target of the project's spectral analysis. From this observation the project branches into three deliverables: discovery of a resolution-dependent carrier frequency structure, an independent detector reaching roughly 90% accuracy, and a multi-stage robustness attack the authors call their bypass pipeline.


The detector side is arguably the most defensible artifact from a defensive standpoint. The repository ships a RobustSynthIDExtractor class that loads a precomputed spectral codebook from an .npz artifact and scores an image for watermark presence, returning confidence and phase-match metrics. In the README's own sanity-check example, watermarked images score conf=0.91, phase_match=0.65, while aggressively processed images drop to conf=0.02, phase_match=0.31. That local detector gives researchers a feedback loop that doesn't depend on Google's app for every measurement.


The analytical core of the project is the V4 codebook construction, and it is genuinely elegant. The key observation is that a true SynthID carrier is image-content-independent: its phase at each frequency bin stays locked whether the background is black, white, blue, green, red, or gray. The tool computes a cross-color phase consensus per bin, computed as the magnitude of the mean of complex phase vectors across the six solid-color references. Consensus values near 1.0 indicate watermark carriers; content-driven energy phase-scrambles across colors and falls below the tau=0.60 cutoff. The README reports that 99%+ of content bins fall below that threshold on the enriched dataset.


The codebook structure reflects this methodology. Each profile, keyed by (model, H, W), stores fields like consensus_coherence as the primary carrier mask, consensus_phase as a subtraction template, inverted_agreement for pairwise phase checks weighted toward the black<->white pair, avg_wm_magnitude, and a content_baseline built from diverse/ and gradient/ reference directories. A live carrier_weights field is updated by a human-in-the-loop calibration loop driven by manual detection tallies. Storage is compact: a 14-profile codebook across two models and seven resolutions is about 220 MB using compact rfft plus float16/uint8 encoding.


The evolution across rounds, documented in a comparison table, reads like a case study in adversarial robustness research. Round 01 tried conservative spectral subtraction and failed. Rounds 02 through 05 escalated through aggressive subtraction plus JPEG, blog-guided absolute bin targeting, denoise-residual phase extraction, and diffusion-VAE regeneration with geometric warping — all failing. The breakthrough in Round 06 came from an unusual source: the Gemini app's own published help text acknowledging that the detector struggles with complex collages and layered textures. The authors, in their words, treated that failure-mode list as an attack specification.


The final pipeline stacks seven independently fidelity-gated stages: a VAE round-trip using Stable Diffusion's sd-vae-ft-mse to push the image off the natural-image manifold the decoder expects, an elastic deformation stage applying a smooth low-frequency warp field that fragments spatial phase consensus, a combined affine transform, a resize-squeeze through AREA downsampling and LANCZOS upsampling, color-contrast micro-shifts, residual-phase FFT subtraction against codebook-harvested bins, and finally a JPEG chain with luma noise and bilateral filtering. Every stage is PSNR-gated and rolls back automatically if quality would drop below the floor, which is how the outputs stay visually lossless.


Two presets, final and nuke, parameterize that stack with different VAE pass counts, elastic amplitudes, and compression chains, with PSNR floors of 14 dB and 11 dB respectively. The repository also maintains the older V3 line, a single-model spectral-subtraction approach achieving 43 dB PSNR and a 75% carrier energy drop, and a fork adds a drag-and-drop desktop GUI in gui/ so the V3 workflow runs without a command line. A community-built visualizer hosted separately illustrates how the watermark is added to images, which is a good companion for understanding the encoding side.


For defenders and platform trust-and-safety teams, the operational lesson is that single-point spectral watermarks are vulnerable to adversaries who can query the generator cheaply and cross-reference solid-color outputs. The countermeasures follow directly from the attack surface: detection pipelines that flag images showing multi-stage transform fingerprints (VAE artifacts, elastic-warp statistics, aggressive recompression chains), watermarking schemes that survive or detect re-manifold projection, and layered provenance combining watermarking with C2PA-style signed metadata. The README's own citation of published academic work on VAE-based distortion shows the research community is actively circling this problem.


The ethics here deserve an explicit note. Stripping provenance watermarks from AI-generated imagery has obvious misuse potential in disinformation contexts, and readers should treat this code as research material for hardening watermarking systems, not as a production tool for laundering generated content. The most defensible workflows are evaluating your own watermark's robustness before deployment, replicating the analysis in a lab on images you generated yourself, and using the detector half of the codebase as an independent verification instrument. The repository's research license and PitchHut project page frame it as an engineering study rather than a product.


Getting started is straightforward for anyone with a Python 3.10+ environment: git clone https://github.com/aloshdenny/reverse-SynthID and inspect the scripts/, src/extraction/, and gui/ trees. The codebook artifacts and the separate hierarchical dataset repository are the real assets — without the reference dataset of model-by-color-by-resolution images, the consensus math has nothing to chew on, and rebuilding it requires generating your own reference corpus from your own accounts.


As a piece of adversarial-ML documentation, reverse-SynthID is one of the more readable public case studies available. It demonstrates disciplined methodology: hypothesis, measurement, iteration, and honest failure logs across six rounds. Whether you are a watermark designer, a platform defender modeling attacker capability, or a researcher studying steganalysis, the spectral-consensus technique at its heart — separating carrier from content by phase stability across controlled inputs — generalizes well beyond this one watermarking scheme.



Official project repository for aloshdenny/reverse-SynthID.

Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.






Source: OffensiveSec
Source Link: https://www.offsecblog.com/2026/10/reverse-synthid-for-studying-spectral.html


Comments
new comment
Nobody has commented yet. Will you be the first?
 
Forum
Red Team (CNA)



Copyright 2012 through 2026 - National Cyber Warfare Foundation - All rights reserved worldwide.