Forging AI Authorship Fingerprints via Rewriting
New research shows how targeted rewriting can forge the stylistic fingerprints that identify which LLM wrote a text, undermining authorship attribution and raising fresh concerns for synthetic media detection.
Authorship attribution has quietly become one of the most important frontiers in digital authenticity. As large language models proliferate, detecting which model generated a given piece of text — not just whether it is machine-written — has emerged as a key tool for forensics, content moderation, and platform accountability. A new arXiv paper, "Forging LLM Authorship Fingerprints with Targeted Rewriting," challenges the robustness of these attribution systems by demonstrating that the stylistic signatures models leave behind can be deliberately forged.
The Concept of an Authorship Fingerprint
Every large language model exhibits subtle, measurable stylistic regularities: preferences for certain phrasings, punctuation rhythms, token-level distributions, and syntactic structures. Collectively, these patterns form what researchers call an authorship fingerprint. Attribution classifiers exploit these fingerprints to decide whether a passage was written by GPT-class, Claude-class, Llama-class, or other model families. Such tools underpin plagiarism detection, provenance tracking, and emerging content-authenticity pipelines.
The premise of these systems is that fingerprints are difficult to fake — that the statistical trace of a model is stable enough to serve as reliable evidence. This paper directly attacks that assumption.
Targeted Rewriting as an Attack Vector
The core contribution is a targeted rewriting method that transforms text so that an attribution classifier misassigns it to a chosen target model. Rather than simply paraphrasing to evade detection (a well-studied problem), the approach actively steers the rewritten output toward the stylistic fingerprint of a specific model. In other words, it does not just hide authorship — it forges a different one.
This distinction matters enormously. Evasion merely breaks attribution; forgery actively manufactures false provenance. An adversary could take human-written text and make it appear generated by a particular commercial model, or take output from one model and disguise it as another. The implications for accountability are severe: if fingerprints can be convincingly transplanted, any attribution-based evidence becomes contestable.
Why This Mirrors Deepfake Dynamics
Readers familiar with synthetic video and voice will recognize the pattern. In the image and audio domains, researchers have long shown that generator-specific artifacts — the "fingerprints" of a GAN or diffusion model — can be suppressed or spoofed to fool forensic detectors. This paper extends that adversarial cat-and-mouse dynamic into the text domain. The lesson is consistent across modalities: any detectable signal that a model leaves behind can, in principle, be manipulated once an adversary can model that signal.
For digital authenticity efforts, this is a sobering reminder that passive, artifact-based detection is inherently fragile. Just as deepfake detectors degrade against adaptive attacks, text attribution systems must contend with adversaries who understand and target the very features classifiers rely on.
Technical Takeaways
The rewriting approach is notable because it is targeted and controllable rather than a blunt paraphrase. By optimizing toward a specific model's stylistic distribution, it demonstrates that fingerprints are not intrinsic, immutable properties but surface-level statistical patterns that can be reshaped. This undermines a key assumption behind many detection tools — that stylistic signatures are a dependable proxy for origin.
For practitioners building authenticity infrastructure, the paper argues implicitly for a layered strategy. Attribution classifiers alone cannot carry the evidentiary weight that legal and platform contexts increasingly demand. Cryptographic provenance (such as signed content credentials), watermarking embedded at generation time, and verifiable logging are structurally more resistant to this class of attack because they do not depend on recoverable statistical artifacts that an adversary can mimic.
Implications for the Authenticity Ecosystem
As regulators and platforms push toward mandatory AI labeling and provenance, the robustness of the underlying detection methods becomes a critical policy question. If authorship attribution can be forged with targeted rewriting, then any enforcement regime built solely on post-hoc classification risks being gamed — and worse, risks producing false attributions that wrongly implicate a particular model or vendor.
This work should accelerate interest in proactive authenticity measures over reactive detection. Watermarking schemes, content-credential standards, and tamper-evident provenance chains all sidestep the fundamental weakness exposed here: that a learned fingerprint is, by definition, something an adversary can learn to counterfeit.
Ultimately, the paper is a valuable contribution to the adversarial robustness literature spanning synthetic media. It reinforces a theme central to this space — that authenticity cannot rest on detecting artifacts alone, but must be engineered into the generation and distribution pipeline itself.
Stay informed on AI video and digital authenticity. Follow Skrew AI News.