Detecting AI-Generated images in health insurance claims

Executive Summary
- As generative AI scales, verifying digital authenticity has become critical, a challenge Google’s SynthID solves for media by embedding durable watermarks that survive heavy editing and file conversions.
- The health and life insurance industry faces a parallel, high-stakes verification challenge: confirming the true integrity and legitimacy of complex medical claims to prevent systemic fraud and waste.
- Qantev addresses this core problem through its unique value proposition: a specialized AI platform that replicates the clinical reasoning of medical experts to establish "data provenance" across claims operations.
- Just as SynthID looks past stripped metadata to find the truth in an image, Qantev analyzes the full clinical context of a claim to surface hidden anomalies and fraudulent billing patterns that traditional rule-based systems routinely miss.
- By automating this deep verification process, Qantev provides insurers with a secure, highly efficient ecosystem that fast-tracks legitimate workflows while aggressively reducing financial leakage.
A hostile landscape
New vectors of attack
The threat has evolved from “fully fake” documents to “locally edited real documents.” Diffusion-based inpainting through consumer APIs can convincingly alter a single word or number in a document photo, seamlessly matching font, texture, and background in under one second for about $0.01 per edit. Unlike traditional Photoshop edits, AI-inpainted regions do not leave visible compression seams or clone signatures.
The impact on the industry, today and tomorrow
The widespread availability of AI has changed the threat landscape for insurers. Forgery is now accessible to a broader group, shifting risk from professional fraud rings to opportunistic policyholders. Because document manipulation requires no technical skill, insurers are seeing more localized fraud. Instead of fabricating incidents, claimants use inexpensive tools to subtly alter legitimate documents, such as changing dates on medical certificates, inflating repair estimates, or modifying invoices. Many consumers now view these digital alterations as a normal and victimless way to increase payouts.
This makes it much harder to process claims smoothly. Standard fraud teams are used to catching obvious mistakes from clumsy, manual edits. Now that those easy clues are gone, the old way of checking documents no longer works. Having so many perfect-looking fakes puts insurers in a tough spot: they must either let subtle fraud slip through or slow down their entire operation by double-checking everything, which increases costs and delays payouts. Ultimately, treating every pristine document with baseline suspicion degrades the digital experience for honest customers, forcing the industry to seek systemic, friction-free verification methods rather than relying on the naked eye.


Qantev approach
C2PA approach, metadata
Think of C2PA (the Coalition for Content Provenance and Authenticity) as a digital passport for files. Instead of a single stamp of approval, it’s a living record called a manifest, a cryptographically signed history attached to a photo, video, or document that tracks exactly where it came from and every change made to it along the way. When you use an AI tool that supports C2PA to generate an image, that origin story is baked right into the file. If someone later edits that image in a compliant program like Photoshop, that edit gets added as a new “stamp” in the passport. Down the line, anyone looking at the file can trace its exact chain of custody.

Qantev leverages metadata standards like C2PA to automatically scan digital claims submissions (like medical invoices, certificates, or photo evidence) for verified manifests. If a document carries a C2PA stamp from generative AI providers, Qantev can instantly raise an alert for that document.
The biggest flaw with C2PA is that it requires a perfect, unbroken chain. For it to work, every single tool that touches the file has to participate.

How it breaks: Metadata is incredibly fragile. The moment someone takes a screenshot of the image, rescans it, or opens it in an older, non-compliant editing program, the C2PA chain snaps completely.
The baseline: Because it is so easy to accidentally wipe this data, a missing manifest proves absolutely nothing. You can’t look at a file without C2PA data and assume it’s fake or tampered with; it usually just means the file took a detour through software that didn’t support the standard.
SynthID
SynthID is an advanced watermarking technology developed by Google DeepMind that identifies content as AI-generated at the exact moment of creation. Instead of attaching external metadata to a file container, which can easily be stripped away, SynthID weaves an imperceptible digital signal directly into the core content itself, whether it is image pixels, audio waveforms, text, or video frames. This embedded mark remains completely invisible or inaudible to humans but is easily readable by a matching detector.

Because the watermark is integrated directly into the media, it survives heavy editing, cropping, or file conversions that typically destroy traditional metadata. SynthID is applied automatically within Google’s ecosystem, including Imagen, Gemini, Veo, and Lyria, and its robust approach has led to wider adoption by other major AI laboratories, such as OpenAI and ElevenLabs, establishing it as a critical standard for AI provenance.
SynthID detects watermarks directly from the core content rather than relying on attached metadata. Because the digital signal is woven into the pixels themselves, it survives operations that would completely wipe a C2PA manifest. For example, taking a screenshot copies the modified pixels, ensuring the watermark travels with the new file.
However, the detection output is deliberately narrow: a positive result simply confirms that a watermarked AI tool was involved at some stage of production. It does not identify the specific tool used, the user who generated it, or which sections of the file were altered. Within an automated claims pipeline, SynthID functions as a high-confidence “tripwire.” When a submitted invoice or medical certificate returns a hit, Qantev can flag it immediately for review, as false positives are exceedingly rare.
While adoption is expanding across major players, including Google’s native ecosystem, OpenAI, ElevenLabs, and Nvidia, SynthID can only detect watermarks that a participating tool originally embedded. Furthermore, its pixel-level resilience has a ceiling. While it handles cropping and compression exceptionally well, physical actions like printing and re-scanning will degrade the signal.
The image watermark isn’t a stamp in one corner; it’s a sub-visible shift distributed across the pixels the model generates. Detection confidence scales with how much of that signal is actually present. When an entire image is AI-generated, there’s a lot of watermarked surface area to read. When an attacker inpaints a single word or number into an otherwise real scan, only that small patch is model-generated, so only that patch could carry a mark, and the surrounding authentic pixels, which were never watermarked, dilute whatever signal exists. Partial watermark presence can sometimes still be recovered through error correction, but Google hasn’t published quantitative thresholds for this, and extensive or aggressive regeneration is known to reduce or eliminate detectability.
Crucially, if a fraudster passes a real document through a local diffusion tool that does not implement SynthID, the watermark will not carry forward. This represents a significant vulnerability, as it leaves the platform exposed to the “locally edited real document” attack vector.
Generalization is the core scientific problem.
If provenance can’t carry the load, the natural next move is a trained detector that directly spots synthetic or manipulated imagery. This works until the generator changes.
Purpose-built detectors reach north of 98% accuracy when the training data and the test data come from the same generator. That figure is misleading. Move to a generator the model has never seen, and performance collapses. In published benchmarks (GenImage, NeurIPS 2023), a ResNet-50 trained on one diffusion model and evaluated on images from a different one fell to roughly 55%, barely above a coin flip. Detectors built on CLIP and other foundation-model backbones generalize considerably better, but they do not eliminate the gap.
The practical consequence is that detection cannot be a model shipped once and left in place. The generator landscape shifts on a monthly cadence, and a detector that was state of the art against last quarter’s models can be near-useless against this quarter’s. Keeping detection effective requires an ongoing engineering commitment: continuous evaluation against new and unseen generators and retraining before the gap opens rather than after. This is a large part of what Qantev takes on so that carriers don’t have to maintain a detection research function internally.
The right strategy is layered defense-in-depth
No single signal is reliable enough to decide a claim. Provenance is positive-only and usually absent. Watermarks are fragile. Passive detectors generalize imperfectly. The correct design treats every individual signal as weak and probabilistic, and derives confidence from their combination rather than from any one of them. This is the principle behind Qantev’s approach, and it maps to five layers.
Workflow controls and provenance capture at intake. Where the carrier controls the submission path, capture provenance at the source: require submission through controlled apps that attach C2PA or Durable Content Credentials along with geolocation and timestamps. Documents that arrive from unverifiable third-party channels are treated as higher-risk by default, not rejected, but weighted.
Provenance and watermark checks. Read any C2PA manifest and detect any SynthID or comparable watermark that is present. These are consumed strictly as positive signals. A valid credential raises confidence; a missing one changes nothing.
Passive forensic detection. For the uncredentialed majority, run localization models of the TruFor and Noiseprint family alongside CLIP- and foundation-model-based AI-image detectors, with thresholds calibrated to the document domain rather than to generic benchmark distributions. Outputs are made explainable — including through multimodal LLMs — so an adjuster sees where and why a document was flagged, not just a score.
Semantic and structural consistency. Much of the strongest signal isn’t forensic at all. OCR feeds business-logic validation: do the fields agree, does the procedure code match the template, does the provider format check out? Layer on duplicate detection, provider-level anomaly detection, and graph analytics across claims. A document can be pixel-perfect and still be exposed by the fact that the same invoice, lightly edited, has been submitted three times.
Human SIU review. The output of the stack is an explainable risk score routed to a Special Investigations Unit, not a black-box auto-decline. This is deliberate. Every layer contributes false positives, and the cost of wrongly denying a legitimate claim is high — in regulatory exposure and in customer trust alike. Keeping a human investigator in the loop is how the system manages that cost while still catching what matters.
The takeaway
C2PA and SynthID are worth adopting, and carriers should read every credential and watermark they can. But treating either as a fraud control is a category error: they authenticate the documents least likely to be fraudulent and say nothing about the rest. Robust defense comes from combining provenance, passive forensics, semantic consistency, and human judgment into a single probabilistic picture, and from keeping the detection layer current as generators evolve. Fraud detection here is not a checkpoint. It is a layered, continuously maintained system, and it is built that way on purpose.
