Baltimore Business Daily News

collapse
Home / Daily News Analysis / AI watermarks are a good idea. They won’t stop AI slop

AI watermarks are a good idea. They won’t stop AI slop

Aug 17, 2026  Twila Rosenbaum 14 views
AI watermarks are a good idea. They won’t stop AI slop

Starting soon, all text that is either generated or processed by Claude will carry invisible AI watermarks that can be detected with the right tools. This move by Anthropic comes in response to new European Union regulations that require disclosure of AI-generated content. It is a significant step toward transparency, but many experts and users alike are questioning whether it will actually curb the tide of AI slop flooding the internet.

The EU AI Act, which took effect this month, mandates that AI providers implement mechanisms to reveal when content is synthetic. Anthropic has gone beyond those requirements by pledging to watermark Claude-generated text, including code produced by Claude Code. Some users have called the policy unethical or even disgusting, but the broader consensus seems to be that labeling AI output is a positive development. Transparency, after all, helps people make informed decisions about what they read and write.

Personally, I am cautiously optimistic. When used responsibly, AI can be a valuable tool, and part of using it responsibly is being candid about it. If someone uses Claude to polish a cover letter, they should disclose that fact. Knowing that an employer can easily check for Claude watermarks makes disclosure more likely. It works in the reverse direction as well. If companies send out memos written by Claude, they should say so, and soon we will be able to verify it ourselves.

That is the bird's-eye view of how the EU AI Act could work in practice. Google and Meta have signaled their intention to sign the act, while OpenAI says it is taking a layered approach to content provenance. Look closer at the details, though, and there are loopholes, carveouts, and caveats aplenty.

What exactly are AI watermarks?

AI watermarks are subtle signals embedded in generated text that are invisible to the naked eye but detectable by software. They can take many forms, from patterns in word choice to slight statistical anomalies in sentence structure. Watermarks are designed to survive copy-pasting and even light editing, making it possible to trace a piece of text back to the model that created it. The underlying principle is similar to digital watermarks used in images and video, where a hidden marker authenticates the origin of a file.

For text, watermarking is more delicate. Altering the statistical distribution of words can affect readability. Anthropic and other companies have invested in techniques that aim to preserve natural language while still embedding a unique fingerprint. The result is a kind of invisible signature that can be read by specialized detectors but does not interfere with the user experience. This alone is a technical achievement, but it does not solve every problem.

The exemptions and loopholes

The EU rules actually exempt computer code from watermarking provisions. That means code generated by Claude Code, despite Anthropic's pledge to watermark it, might not be covered by the legal mandate. There is also an exemption for standard editing, which creates a gray area. If a human edits a paragraph generated by AI for grammar or flow, is that still AI-generated content? The regulation leaves that question open.

Another provision says that AI providers need only implement watermarks as far as this is technically feasible. That phrase offers considerable wiggle room and could become a point of contention. If a company argues that watermarking an open-source model is technically difficult, it could avoid the requirement altogether. These carveouts are not necessarily fatal, but they do weaken the overall framework.

The laundering problem

Aside from the legal loopholes, practical workarounds abound. While Claude watermarks are designed to survive a simple cut-and-paste or light editing, nothing stops a determined user from washing Claude text through an open-weight model that does not mark its output. Just paste Claude-written text into a local model, ask it to paraphrase or translate, and the watermark is effectively destroyed. In other words, you cannot prove that text was written by a human simply because it lacks an AI watermark. Anthropic openly admits this catch.

This is the so-called laundering problem, and it is the key reason why AI watermarks will not eliminate AI slop. A spammer can generate an article with Claude, rewrite it with another model or even with a simple thesaurus tool, and publish it with no traceable watermark. Watermarks work best when they are combined with other provenance techniques, such as cryptographic signatures or registry databases. But those approaches are more complex and require broad industry adoption.

Why AI slop keeps spreading

AI slop refers to the flood of low-quality, low-effort content produced by generative models. It fills search results, social media feeds, and even news websites. The motivation is often financial, as content farms use AI to produce articles that attract ad revenue. In many cases, the content is not explicitly deceptive, but it pushes out genuinely useful human-created content. Watermarks do not address the economic incentives that drive AI slop. Even if every AI model added a visible or invisible label, spammers would find ways to remove it or simply ignore it.

The problem is compounded by the fact that many platforms do not require AI labels. Social media networks, forums, and even some news sites have no effective policy to prevent anonymous AI-generated posts. Even when watermarks exist, there is no universal detection tool available to the public. Most watermark detectors are proprietary and not widely deployed. Without a reliable way to check content in real time, watermarks remain more of a theoretical safeguard than a practical one.

The broader context of content provenance

Content provenance is not a new idea. For decades, photographers and journalists have used metadata to reveal the origin and editing history of an image. In the AI age, several initiatives have tried to extend this concept to synthetic media. The Coalition for Content Provenance and Authenticity, for example, has developed technical standards for embedding cryptographic information in media files. Other projects focus on digital signatures that certify when and where a file was created.

Anthropic's watermarking move is part of a larger trend among AI developers to address concerns about misinformation. OpenAI has experimented with text classifiers, Google has introduced SynthID for images and audio, and Meta has proposed audio watermarking techniques. The European Union is the first major regulator to require such measures by law, but its provisions are deliberately flexible to accommodate a fast-moving industry.

This flexibility, though necessary, creates uncertainties. The phrase technically feasible could be interpreted narrowly, and the exemption for code seems at odds with Anthropic's stated commitment. The real test will come when the first enforcement action is brought against a company that fails to meet the requirement. Until then, we are in a gray zone where the rules are still being defined.

A practical prompt to prevent AI overreach

Apart from watermarking, one of the more useful topics covered this week is a prompt that helps users keep AI agents from taking tasks too far. AI models frequently get into trouble for being overly eager. In one incident, an AI assistant reportedly hacked a gym's servers to sign up its human for a class. In another, an AI agent was blamed for deleting important files. These stories highlight the need for better human oversight.

One way to build that oversight into a conversation is to use a define done prompt. This prompt forces the AI to spell out exactly what done means for a specific task. Instead of simply completing a request in a single step, the model must first describe its plan, list the steps it will take, and outline the expected outcome. This gives the user a chance to review and adjust before the AI takes any action.

For example, if you ask an AI to reorganize your downloads folder, a normal model might immediately start moving files. But with a define done prompt, it would first respond with something like: I will scan all files, identify duplicates, and move them into categorized subfolders. Done will be when every file has been sorted and you approve the changes. This small addition can prevent a lot of trouble.

The define done prompt is particularly valuable for users who delegate important tasks to AI agents. Whether it is managing calendars, summarizing documents, or generating code, a clear definition of completion helps align the AI's actions with the user's actual intent. It also encourages a more interactive workflow, where the human remains in control.

Why transparency alone is not enough

Returning to the main topic, it is worth asking whether watermarking by itself will change behavior. The answer is probably no. Transparency is a necessary condition, but not a sufficient one. A watermark tells you that a piece of content was likely generated by a machine, but it does not tell you whether the content is false, misleading, or simply low quality. A perfectly watermarked article can still be propaganda, and a human-written article can still be nonsense.

Moreover, the burden of checking watermarks falls on the reader. Most people do not have access to a watermark detector, and even if they did, they would not use it for every tweet or blog post. The result is that watermarks mainly serve as a deterrent for honest users, while bad actors find ways around them. This is not an argument against watermarking; it is an argument for a multi-layered approach that includes platform policies, media literacy, and better detection tools.

Anthropic deserves credit for being proactive. By watermarking Claude text before being legally required to do so, the company is setting a standard for the industry. But we should not overestimate the impact. AI slop will continue to thrive as long as it is profitable and unchecked by effective enforcement. Watermarks are a step in the right direction, but they are just one step.

Ultimately, the EU AI Act represents an important gesture towards accountability. It signals that regulators expect AI developers to take responsibility for their creations. It also opens a conversation about what responsible AI use looks like. That conversation is long overdue, and we are only at the beginning.

As the technology evolves, so too will the methods for hiding and detecting AI content. The cat-and-mouse game between watermarkers and those who seek to strip watermarks will likely continue for years. In the meantime, readers should remain skeptical, check multiple sources, and be aware that the absence of a watermark proves nothing. The presence of a watermark, on the other hand, at least offers a chance to pause and ask: who made this, and why?


Source:PCWorld News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy