BIP America

collapse
Home / Daily News Analysis / Someone built a free tool to scrub AI watermarks from OpenAI, Gemini-generated text and files

Someone built a free tool to scrub AI watermarks from OpenAI, Gemini-generated text and files

Aug 14, 2026  Twila Rosenbaum  5 views
Someone built a free tool to scrub AI watermarks from OpenAI, Gemini-generated text and files

A new open-source tool has been released that claims to remove AI watermarks from text and files generated by major AI models, including OpenAI, Gemini, and Claude. The project, named watermarks-remover, is hosted on GitHub and has been designed to strip various types of provenance signals embedded in AI-generated content. The tool has gained attention because it takes a multi-layered approach to watermark removal, ranging from cleaning simple metadata to aggressively rewriting text to disrupt statistical watermarking patterns.

AI companies are increasingly looking for ways to mark content produced by their models. Watermarking is seen as a key method for establishing provenance, especially as concerns grow about AI-generated misinformation, fake reviews, and academic dishonesty. However, watermarks-remover aims to give users a way to bypass those markers. The developer, Guillaume Meyer, recently announced that the tool now supports OpenAI and Gemini in addition to Claude, making it one of the more comprehensive watermark removal projects currently available.

How the tool works

The watermarks-remover project separates its approach into different layers. One layer targets edit-based signals such as unusual Unicode characters that are often inserted invisibly into AI-generated text. Another layer focuses on statistical patterns embedded in the way an AI model selects and arranges words. A separate cleanup layer deals with file provenance information, including C2PA, EXIF, XMP, and document properties that may be attached to generated images, PDFs, or documents.

The latest version, 0.3.1, makes the text rewriting side more aggressive. Instead of simply swapping a few words, the tool can change sentence structure, word choices, transitions, and other stylistic patterns in an attempt to disrupt statistical watermarking. It also includes options intended to make rewritten text sound more natural, which is important because heavy rewriting can often produce awkward or unnatural output.

Statistical watermarks are particularly difficult to remove. They work by subtly biasing a language model's word selection during generation. For example, the model might slightly favor certain word combinations or sentence constructions that appear random but follow a hidden pattern. A detection system can later analyze the text to see if those patterns are present, indicating that the text likely came from a particular AI model. Removing this type of watermark requires changing the text enough to break the pattern, but not so much that the meaning or tone is lost.

The tool also handles metadata watermarks, which are simpler to remove. Many AI-generated images and documents contain metadata fields that record the generating model, creation date, and other identifying information. Standards like C2PA allow companies to cryptographically sign this provenance data, making it tamper-evident. However, removal tools can often strip the metadata fields or re-encode the file in a way that discards the signatures. The watermarks-remover project targets these metadata tags across formats including PNG, JPEG, SVG, PDF, DOCX, ODT, HTML, and Markdown.

Limitations and caveats

There is an obvious catch to this approach. Rewriting text to remove a statistical watermark can also change the text itself. The project's documentation admits that the process can affect tone, voice, and precision, particularly when a large portion of the original wording needs to be changed. For professional writing, this could be a serious drawback.

In a post announcing the update, Meyer said that his tool currently only gets rid of metadata. He clarified that removing the real, statistical marks might come later, but not yet. The project's README goes even further, noting that a rewrite changes the words used by a premium model to those used by a cheaper model. It also raises a rhetorical question: why would someone pay more for a premium model and then use a worse model to process its output? That question highlights a fundamental tension in the watermark removal process.

Moreover, the developers are not claiming that this tool will fool every AI detector. The project describes the rewriting process as best effort and says it cannot guarantee that a particular vendor's detection system will fail. Some signals can also remain after the cleanup process, especially if the rewriting is not aggressive enough or if the detection system uses more sophisticated analysis.

It is also important to note that there is no universal AI watermark hiding inside every piece of generated content. Different companies and systems use different approaches, and the project itself divides these signals into several categories. For instance, some watermarks are embedded at the token generation level, while others are inserted into image pixels or document metadata. A tool that works well against one type may be ineffective against another.

Privacy, research, and ethical concerns

The repository states that the purpose of the tool is privacy and research, rather than helping people falsely claim that AI-generated work was written entirely by a human. This distinction matters because watermark removal raises ethical and legal questions. On one hand, some users may have legitimate privacy concerns, such as wanting to avoid having their AI-generated content easily traced back to them. On the other hand, the tool could be misused to obscure AI involvement in content that is passed off as human work, which could facilitate academic cheating or disinformation campaigns.

AI watermarking itself is still a developing field. Companies like OpenAI, Google, and Anthropic have been exploring different watermarking techniques, but none are perfect. Some early proposals were criticized for being easily removed, while others were found to degrade the quality of generated text. There are also concerns about false positives, where human-written text is incorrectly flagged as AI-generated. As a result, the cat-and-mouse game between watermark developers and watermark removers is only just beginning.

The emergence of watermarks-remover is a sign of where the AI industry is heading. Companies are looking for ways to establish provenance and identify AI-generated material, while developers are already exploring how to remove or disrupt those signals. This is similar to the history of digital rights management (DRM) in the media industry, where copy protection schemes are constantly being developed and then broken. Data encryption, metadata standards, and hidden signatures all face the same fundamental challenge: after the content is in the hands of a user, it is very difficult to prevent them from altering it.

In addition to the technical challenges, there are legal implications. Some watermarking methods may be protected by patents, and attempting to circumvent them could violate terms of service for AI platforms. However, in the United States, the Digital Millennium Copyright Act (DMCA) includes provisions that could be interpreted either way, depending on the exact nature of the watermark and the purpose of removal. The developers of watermarks-remover have not made any legal claims, and they urge users to consider the ethical dimensions.

The latest version's support for OpenAI and Gemini expands its reach significantly. OpenAI's ChatGPT and Google's Gemini are among the most widely used AI chatbots in the world, and both companies have been proactive in adding provenance signals to their outputs. OpenAI has experimented with cryptographic watermarks, while Google has integrated SynthID into its Gemini models. SynthID works by embedding a digital watermark into the generated content that is invisible to the human eye but detectable by computers. The designers of SynthID have claimed it is resilient to various forms of tampering, but the researchers behind watermarks-remover are attempting to find weaknesses anyway.

Anthropic's Claude has also been updated with watermarking capabilities, and the tool has supported Claude for a longer period. The developer's announcement that the tool now works with all three major AI providers makes it a useful case study for understanding the current arms race around AI content provenance.

For now, watermarks-remover is best understood as a sign of where the AI industry is heading. The interesting part is not whether this particular GitHub project can beat every AI detector. It probably cannot. The interesting part is that AI watermarking is already becoming a cat-and-mouse game, and we are still very early in it. As AI models become more capable and more widely used, the pressure to identify their output will only grow. At the same time, the desire to anonymize or alter that output will also intensify. This ongoing tension will likely lead to more sophisticated watermarking, more advanced removal methods, and more debate about the appropriate balance between transparency and privacy.

Ultimately, the development of watermarks-remover reflects a broader societal question: when AI-generated content is everywhere, how should it be marked, and who gets to decide whether those marks can be stripped away? The tool provides a practical demonstration that no watermarking system is foolproof, and that any effort to embed provenance signals must consider the possibility of adversarial attacks. For now, the project exists as a reminder that the battle over AI content attribution has only just begun.


Source: Digital Trends News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy