Rendered at 06:28:13 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Dilettante_ 22 hours ago [-]
I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
unprovable 19 hours ago [-]
This. FN rates are cute, but FP rates will ruin an academic career or a student's work/further study choices if their content gets marked erroneously. Surely the answer is a sequence of marks?
Keen to see if they are doing something SynthID-esque?
WD-42 15 hours ago [-]
Do they care about false positives? As long as it’s even somewhat reliable that’s enough for them to prevent training on their own slop. I think this is a big reason to do this that’s overlooked.
DennisP 15 hours ago [-]
Good point. But if that were their only purpose, there'd be no need to share it with anybody. In fact, they'd get the best results by not mentioning it.
WD-42 15 hours ago [-]
That’s true. But they probably want to be able to identify other models slop as well. And with the laws popping up, it makes sense to do it the way they are.
dragonwriter 11 hours ago [-]
> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem.
anon373839 10 hours ago [-]
But I thought Anthropic was an altruistic organization devoted to the betterment of humanity…
AustinDev 9 hours ago [-]
It would appear that their altruism isn't very effective.
jrflo 15 hours ago [-]
I think that false positives are inevitable due to the method of watermarking being embedded in the text itself. The output is intended to mimic human writing, therefore it's entirely conceivable that a human could by chance write text that contains the watermark. The odds may be extremely small, but it's not something you could ever guarantee.
andai 14 hours ago [-]
I keep hearing how humans are thinking and writing more and more like AI.
I think in this case I think it's some kind of cryptographic signature smeared across the token IDs, so I don't think the risk is very high.
jrflo 11 hours ago [-]
I read the original paper they're basing this off of and I think you're right. I do wonder how much of a quality tradeoff there is with perturbing the next token probability distribution. My intuition tells me that a more "prominent" watermark will necessarily degrade output quality. If they are trying to balance quality and watermark prominence, I wonder if that affects the FPR.
xena 12 hours ago [-]
You're absolutely right! Humans have been slowly thinking and writing more and more like AI. As people get more and more exposed to the stochastic patterns of large language model tools, it's normal for them to emulate the styles of communication they are exposed to. This is commonly called "brainrot" by those in Gen Z and younger cohorts.
If you find yourself getting to be afflicted by this "brainrot", be sure to go outside and take a moment to ponder what's around you. The grass is there and will be there long after we are all gone. Consider this for a moment as your organic thought processing unit starts to slowly munch away at its internal context window.
ed_elliott_asc 13 hours ago [-]
It’s worse than that, false positives are possible but someone generating text should be able to get ai to change some words and formatting to break the watermarking, then ai detectors can tell them how well they did.
I don’t know what the answer but I absolutely know it isn’t this.
dbqpdb 12 hours ago [-]
I think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board. But each intermediary or source (optionally) cryptographicaly signs a piece of content that it either generates, edits, or passes along, and the end result at a destination, is that content is either 'trusted' if its cryptographic chain is solid, or un-trusted otherwise.
dragonwriter 11 hours ago [-]
> I think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board.
It would also require the individual humans you are trying to control to get on board otherwise the analog hole breaks the chain, absent mindboggling levels of physical surveillance on top of the the total monitoring of all electronic data flows that this idea requires.
normalaccess 9 hours ago [-]
I think that's the end goal.
normalaccess 9 hours ago [-]
This is a meme video but I think it hits the nail on the head.
TLDR: AI will force global online digital ID for everyone that uses the internet for the exact reason you mentioned. And that would forever change free speech forever allowing the powers that be to put the genie "back in the bottle" so to speak.
If there is any false positive rate (which, because text will naturally and by chance include tokens from the green and red sets in some pattern, there will be), tools making promises like "detect AI-generated text" are unacceptable. They are going to turn innocent people into pariahs on some unsubstantiated "this content is 37% likely to be AI" claim that the user has no way of verifying or inspecting more deeply, we just have to trust the statistical box and assign some meaning to whatever that number means. 37% of my phrases are AI? There's a 37% chance my entire text is AI written? Part of the fun is not knowing!
This is scripture homeopathy and it's irresponsible.
wrsh07 17 hours ago [-]
I'm curious about your thoughts on pangram. I only really see posts on Reddit claiming it falsely labels their content as ai generated but nobody will actually post examples of "textbook from twenty years ago" or upload screenshots of a journal (also those posts usually feel deeply ai generated without an ai detector)
Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
nemomarx 17 hours ago [-]
This feels testable - you could go to fanfiction or similar sites with billions of words of writing from before 2016 or so and run them through it.
I tried a chapter just now and got human doing that, but I'm not invested enough to run a hundred samples today. But it sounds like it would be an alright way to audit it? I will confess I'm pretty skeptical you could ever eliminate false positives here though. I can often get an ai sense from some writing on my own but I doubt it would be better than 90% accurate, and "ai plus human editing" might screw with that anyway, stuff like that. I would have preferred we just never developed this kind of thing so I wouldn't have to guess.
philote 17 hours ago [-]
That still might work better with older texts. As AI-generated text gets more prevalent, I'm guessing people will start subconsciously adopting AI writing styles.
nemomarx 17 hours ago [-]
Yeah, that's one of my questions. Everyone who talks to AI for too long seems to get worse at writing anyway, and humans mirror any form of conversation to some extent.
subsistence234 13 hours ago [-]
LLMS aren't the only thing that has changed over time in the way texts are written.
if they used older texts as training data, to some extent pangram would just be an age classifier for writing style.
StilesCrisis 17 hours ago [-]
It's been done and showed up on HN recently. Older content was quite consistently marked as not-AI.
estebarb 16 hours ago [-]
Language distribution shifts. Eventually people will start adopting the distribution used by LLMs, making classification harder.
Also, this doesn't even consider the case where people use LLMs to translate their original works. Or people that use it for spelling/grammar checks.
Personally, I believe these checkers do more harm than good. Any false positive can ruin someones life.
jmalicki 10 hours ago [-]
> Eventually people will start adopting the distribution used by LLMs, making classification harder.
I recently heard someone say "that's genuinely the exact solution I was looking for" and had to do a double take.
> Pangram 4 achieves a 0.0041% false positive rate (roughly 1 in 24,000) on 1,000,000 human-written English FineWeb evaluation examples
> Overall False Negative Rate is 0.3396% on English AI generations (26 generator models)
nunez 17 hours ago [-]
pangram is pretty good; i use it all of the time and pay for it. surprised that it's not mentioned that often here. they just released a new model that is supposed to lower the fpr (false positive rate) even further than it's already impossibly low score. it also detects AI in images now, though I expect the fpr to be pretty high there given its newness.
hojinkoh 2 hours ago [-]
Sure, there are things you could do legally when falsely accused; and there are things authorities and companies should do.
But ultimately, when you are powerless and can't afford to do the fighting: I'm convinced the only way to protect yourself is to be very mindful about your writing style, and to deliberately corrupt the language through objectively wrong "stylistic elements".
whack 15 hours ago [-]
> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated
Are you talking about pieces that were fully human-written with zero AI editing/rewriting etc? If so, what makes you think that false positives will happen there? They aren't looking for "writing styles" or emdashes etc. They are using watermarks and metadata.
If you're talking about people using AI to copy-edit text they manually wrote, this was explicitly called out in the article:
> A detected mark provides a signal that content was processed by Claude, but is not fully conclusive. Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example: Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source; The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.
Dilettante_ 14 hours ago [-]
The former. I'm not sure what you mean by metadata, but my expectation was that anything that Claude could put into the plaintext to identify itself may plausibly also accidentally be produced by [a million monkeys on typewriters/one in a million human writers], since in the end, the writing is using the same language and symbols that humans use. How unique could the LLM possibly make it while still retaining its usefulness?
derefr 11 hours ago [-]
> How unique could the LLM possibly make it while still retaining its usefulness?
They could be doing invisible and vaguely-harmless Unicode stuff. Insertion of zero-width joiners and non-joiners, replacement of regular spaces with non-breaking spaces, building spaces from multiple hairline spaces, intentional use of non-NFC-normalized codepoint sequences for accented characters, etc.
Text with all this junk in it still reads the same; it just might wrap a little strangely, or not byte-match / collate correctly in a database (and Anthropic has never made a guarantee that their models would be capable of emitting text with these properties, so that’s fine.)
And, importantly, no regular text or document editor would insert these things (especially in the useless places you could insert them for watermarking.) You only really see them in text that’s been explicitly typeset for a specific layout (e.g. in text-containing SVGs, website mastheads, or game HUDs) or for print publication.
Of course, if this is the technique they end up using, then it’s very simple to strip it out by canonicalizing the text (i.e. Unicode-normalizing it + stripping out invisible layout characters + replacing “weird spaces” with regular ones, etc. Essentially the same thing many sites already do to user-generated content to prevent users from using Unicode features to break the page’s layout.
WiSaGaN 19 hours ago [-]
My guess is that they will later "reveal" some "violations" but provide little evidence citing proprietary algorithm.
m00dy 15 hours ago [-]
Deepwalker once cracked Gemini's watermarking system, I'm sure they will also work on this [0].
If LLM training data is human-written, and LLM output mimics that input, how could you not have false positives?
SkyBelow 16 hours ago [-]
Because it won't be in the training directly. It is applied after a model generates its distribution of likely tokens, biasing each token randomly based on a random key and unrelated to any meaning of the words. So half the time, the most likely token becomes more likely and half the time it becomes less likely, and the same for every other token (when temperature is above 0).
You then look at the tokens actually picked to see how closely they follow this pattern that isn't connected to the meaning of the tokens. With enough text, you can then analyze the chance of it happening by chance verses being because the generation of the tokens was done using the algorithm, and you can save a positive result until you are arbitrarily sure. There is a chance of a false positive, but the chance of a false positive approaches the chance that the murderer happened to have fingerprints that matched your and both forensics labs happened to have mixed up the dna tests and the eye witness happened to misremember the face and your phone gps happened to glitch out and put you at the murder scene at the time of the crime all happening. It is theoretically possible only in the same sense that quantum teleporting a cat is theoretically possible.
The real question is how much text do they need for a given level of certainty and what do they check for. If they flag a positive at a p value <.01, that's a problem. If they can reasonably get a p value of < 1e-12 in only a few paragraphs of text, that is effectively no false positives (but a lot of 'too short to analyze' outcomes).
ricericerice 6 hours ago [-]
How do you verify in practice then? Wouldn't you need the original prompt so you can reobtain the likely token distribution to validate again the random key(s)?
wrsh07 17 hours ago [-]
Somewhat trivially, if I ask Claude to transcribe an image and then check if that transcription is ai generated it will likely say yes.
Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.
basch 17 hours ago [-]
How is a perfect transcription of an image watermarked?
FeteCommuniste 15 hours ago [-]
"Perfect as far as human perception can tell" is a weaker standard than "bit-to-bit copy." Maybe it's that?
wrsh07 11 hours ago [-]
It depends on how it does watermarking!!
Note, there are many ways to represent words visually on computers that look identical
basch 11 hours ago [-]
If they were substituting glyphs for identical ones people would be able to reverse engineer it.
Theres no way that’s what they are doing.
embedding-shape 21 hours ago [-]
Read said section yourself perhaps.
suddenlybananas 21 hours ago [-]
That's essentially impossible, unless you mean they didn't measure a false positive rate.
Filligree 20 hours ago [-]
For watermarked long-form text, it is actually possible. Makes the watermark more fragile, but the math is considerably more forgiving than usual.
embedding-shape 20 hours ago [-]
> For watermarked long-form text
What about "watermarked long-form code"? I'm having a hard time understanding how a model could watermark not prose, but functional/semantic text like code, that actually has meaning. You can't switch our the characters, you can't use various types of whitespace, you can't add arbitrary code comments, and a lot of other restrictions. Is there any state of the art methods for watermarking code without affecting the quality/correctness?
wpietri 20 hours ago [-]
As anybody who has put together a coding standard knows, there are a lot of options for individual expression, meaning a lot of room for things like watermarking. And of course you can add arbitrary comments; my Claude-generated code is very verbose.
embedding-shape 20 hours ago [-]
> there are a lot of options for individual expression, meaning a lot of room for things like watermarking
The way I use LLMs (and I'd advice everyone to do the same) there really isn't, the agent implements things exactly how I want them, or I use the agent to massage it into the exact bit-by-bit version I imagined when I first sent the prompt afterwards. I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually, although I know it's a popular approach taken by many.
> And of course you can add arbitrary comments; my Claude-generated code is very verbose.
So watermarking for all users who allow code comments from agents, no watermarking for us who force the agents to never write a single code comment? Alright, I'd be fine with that.
wpietri 19 hours ago [-]
From what I've seen, your approach to LLMs is exceedingly rare, so I suspect it's one the people who care about watermarking aren't very concerned with.
And the reason to let Claude make worse code than a professional would by hand is basically suppressed demand. Since programmers are expensive, previously code mostly got written when a large number of dollars were on the line, or when an individual programmer did something not economically optimum (e.g., hobby project).
That left a whole lot of somewhat less valuable software unwritten. It's the economic space that no-code tools have been nibbling on for years. One way to think of things like Claude Code is as effectively no-code tools. Pre-LLM no-code tools would produce data structures that got executed by special environments without ever being seen or tuned by a human. Claude Code can be used just like that, with text as the input and python as the intermediate representation that nobody ever looks at.
That approach probably isn't sustainable for what we professional programmers would call a serious project. Claude can easily get in over its head and I expect that its code decays over time, in a fashion similar to how many human teams get in a state where they just have to rewrite everything. But faster, I'd expect.
But there are a lot of unserious projects that previously would have never been created. E.g., a quick app to manage your little league team, or a bit of in-house business stuff in the "a little hard to do with a spreadsheet" range.
olmo23 18 hours ago [-]
> I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually
You never generate throwaway code used to test an external service? or try out an interface idea? There's a lot of code that's only meant to be ran once. I often dont even care what language it's written in.
embedding-shape 18 hours ago [-]
> You never generate throwaway code used to test an external service? or try out an interface idea?
And save/persist it? No, most of any experimental stuff goes into /tmp which gets cleared out on reboot, nothing I care to save in any repository. Or just "show me how this would look like" and then it's only in the session itself (and the logs/state I suppose, technically...).
SkyBelow 16 hours ago [-]
For straight generated code it'll likely need more text, but it'll still show up.
In cases where one token is extremely likely, it'll randomly be red or green and still be picked in either case as it is simply the best (or only) option. So you'll have more tokens that don't show a pattern either way (half of these cases will match and half won't, just the same as if a human wrote it). Meaning you'll need more instances where multiple tokens were all likely to see if there is a pattern. Given the check algorithm can't identify these cases, it can only judge on the overall text, so the more strict a language, the more the length requirement scales.
Where I wonder if this keeps working is in tool calls. Often, you don't take code straight from the llm, you take the results of a tool call to edit already existing code. It might be that the result of this leads to far too few signals to pick up, meaning that this only works when one does significant generation with a single model (even swapping between different models, at least by different companies, breaks this just as much as having a human write parts of the code).
Think of it like finding a loaded dice. A dice that has a slight bias in a few dozen roles is just random chance. If that bias continues after hundreds of thousands of roles, the dice is loaded. But will a code base have enough samples, especially when edits made from tool calls? I could see this being unable to detect things at the size of a reasonable PR and only being useful for massive sets of changes and only if the person behind them didn't structure their AI usage to avoid detection.
peyton 20 hours ago [-]
[dead]
OGWhales 18 hours ago [-]
And yet, it remains possible that a human could write the same sequence of characters.
no_multitudes 14 hours ago [-]
How often do you add seemingly-random zero-width unicode characters to the text you write?
bufbupa 17 hours ago [-]
Sorry you're getting downvoted, this interpretation doesn't seem that far fetched to me.
Here's the strawman: The text-based watermarking is going to be done procedurally instead of generatively. Maybe they add some sequence of zero-width Unicode characters to all generated text at certain intervals. Then, there is effectively no false positive possible (because humans would [effectively] never type such sequences of unicode naturally). It may survive some editing (depending on how you select/edit the characters), and it's possible to be stripped (false negatives).
wtfwhateven 17 hours ago [-]
Why would you say something so ridiculous?
jobigoud 13 hours ago [-]
I think they mean it like this: imagine you ask me a random number sequence. I give you a random number sequence. Little did you know, I used a very specific PRNG to generate it, so later I can prove with certainty that your number was generated by me, and you can't say you came up with it yourself.
There is no room for false positive here in the same way you can't randomly find a collision in a hash function if it's strong enough. Like the rate is so infinitesimal that it is effectively zero.
Now replace random number sequence with prompted string of words. And instead of using the PRNG on every word I use it every n words. If the generated text is sufficiently long I can tell by matching the expected deterministic pattern.
You can defeat it by changing the words yourself and triggering a false negative but there isn't really any room for a false positive if the text is long enough and the pattern matches perfectly. If the pattern doesn't match then I can compute a probability.
nunez 17 hours ago [-]
I think Pangram is way ahead of Anthropic on this with their custom dataset.
simonw 1 days ago [-]
> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
I'd like to know a lot more about how that works.
A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know exactly what I want to happen, it will NOT be OK if those refactors have some kind of weird pattern baked into their text to act as a watermark.
I guess this may be covered by this:
> Content generated by Claude may not carry a detectable mark if, for example: [...] The passage is very short, leaving too little text for a reliable signal;
wrsh07 17 hours ago [-]
Scott Aaronson talks about his project at OpenAI to do this^
You can carefully select which pseudorandom number generator (prng) you use to be able to id text of a certain length. I expect there is some performance characteristic you have to manage since you're doing this on every inference, but once you do that it doesn't change the output in any meaningful way (the prng is still a statistically valid prng, it just happens to let you check if the output used that prng)
The point is that you can do this simply by swapping to a different RNG, which isn't noticeable to the end user, and while it changes the output, it's not any different from how using a different seed or being lumped in a different batch will change the output.
^ excerpt:
> So then to watermark, instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI. That won’t make any detectable difference to the end user, assuming the end user can’t distinguish the pseudorandom numbers from truly random ones. But now you can choose a pseudorandom function that secretly biases a certain score—a sum over a certain function g evaluated at each n-gram (sequence of n consecutive tokens), for some small n—which score you can also compute if you know the key for this pseudorandom function.
That's what I excerpted, although I had seen it presented from his talk at Stony Brook
denverllc 14 hours ago [-]
I’ve always wondered how this works when we only observe the final output and not the internal state that’s used to generate the output.
The LLM presumably generates f(input, RNG) but we only can observe f(RNG).
Groxx 13 hours ago [-]
Since they do have the input, they could probably just store checksums at each step...
... though I'm not sure why that would be preferable over a coarse rolling checksum over all of the output. Seems like that wouldn't influence output, would be equally imperceptible, and probably easier to calculate (compared to "hash seed times running all LLMs supported times number of RNG algorithms, to see if output matches").
Presumably there's some other trick, or it's a red herring / failed experiment and not what they actually do in practice.
wrsh07 11 hours ago [-]
No you don't need to do that, the prng is detectable if you know what bias to look for and have the key
teravor 12 hours ago [-]
this would be a very heavy watermark application.
there are many simpler methods, for example you can have a tiny windowed transformer operating on the output text and all you do is alter certain words (that don't change meanings) to maximize its surprise. the tiny language model will have a special training regime to build up a somewhat unique view of the language.
we are talking about a 0.5 bit watermark here (existence). I would have zero confidence in being able to reliably remove such a watermark from pretty much any medium.
wrsh07 11 hours ago [-]
That's actually much worse because it fundamentally changes the output, whereas this doesn't change the output, it just changed the prng
cma 7 hours ago [-]
When you tell the AI: copy this function to here, a small window rewriter would mean it just corrupts and changes it instead of moving it. And even it's own tool use would have some small window dumb model changing the tool calls based on what it thinks are synonyms? The Aaronson approach is much better than this, it's essentially like changing out the random seed for the sampling parts that were already random. For an operation like recall of previous text, the tight logits that result still keep it doing that close to deterministically. For something it creates itself, with more spread out probability mass, it gets watermarked.
COAGULOPATH 1 days ago [-]
>I'd like to know a lot more about how that works.
My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.
Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.
So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).
staticman2 18 hours ago [-]
> Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something.
Wouldn't you need the prompt to know the probability of the next token?
Chabsff 17 hours ago [-]
Not necessarily. Here's a rough example (it's not what's going on here, just a representative idea):
There are words/tokens that are heavily correlated to the prompt (a yes or a no, for example), and then there are others that are going to be less so (adjectives with a lot of synonyms for example).
Given a text, you can identify what the "load bearing" and auxiliary words/chunks are. Then, looking only at the auxiliary words/chunks, you should, in principle, be able to determine what other wordings could have gone there instead. From this, you can, very roughly, recreate the token probability distribution that was in effect when those tokens were generated. With the probability distribution in hand for enough chunks of text, you can start inferring properties about the RNG process that was used to sample from those distributions.
But then, this notion of "load bearing" vs "auxiliary" can be expressed directly in the probability distributions. A load bearing token just has a very high probability, and thus any RNG bias that may have been in effect will likely be swallowed in the distribution. So the parts of the text that are highly dependant on the prompt will naturally not be contributing much information about he RNG in the first place.
JohnMakin 15 hours ago [-]
So this is finally what "load bearing seam" means.
thunfischtoast 23 hours ago [-]
They still need to choose when to do that though. When I prompt the program to e.g. alter a bash script in a specific way or to recite a longer known text it can't go round and randomly exchange tokens. It has to somehow define what is a simple repeated text from a different origin and what is a novel generation.
user43928 23 hours ago [-]
I am wondering how that applies to newly generated code.
Odd variable naming? Stylistic choices that are watermarked?
Or as someone else noted further down in the comments, it could be more subtle:
Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.
melvinroest 22 hours ago [-]
> Odd variable naming? Stylistic choices that are watermarked?
Whatever it is, I'm sure it's load-bearing.
asdfsa32 22 hours ago [-]
You're absolutely right. But it is not just load-bearing, it is the load-bearing seams.
silversmith 17 hours ago [-]
Personal observation: Opus 5, over the last week, has started outputting A LOT more comments. Despite my global instructions being full of variations on "don't use comments unless absolutely necessary".
I might be imagining things of course. But comments would be great fit for this use case.
StilesCrisis 17 hours ago [-]
They have mentioned that their system prompt used to say "avoid over-commenting" and it no longer does. They should bring that back IMO.
pram 14 hours ago [-]
Yes the length of comments Opus 5 leaves is exhausting. Not to mention it will insert info thats only relevant within the current session. I've just been deleting all of them lol
sixothree 8 hours ago [-]
Comments seem most plausible, especially since I absolutely expect it to match my code style, existing architecture, and have the code go through CSharpier and dotnet format after the fact.
edit: as an aside - I actually use extensions to collapse comments and change the color to be less intrusive.
kuboble 19 hours ago [-]
I cannot imagine the code with well defined specification will have extra watermarks unless the watermark is requested as part of the harness instructions.
If it works like people describe - on the every nth token or something - then the mark will be left in the chain of thought and discussion with the model - not in the code artifacts.
wpietri 20 hours ago [-]
I would guess they're not worrying about watermarking a tweak to a human-written program. That's both a tiny fraction of Claude use and of very little concern to the kinds of people who want to check watermarks.
__MatrixMan__ 17 hours ago [-]
If you're that targeted with your edits, then do you deserve a watermark anyway?
KoolKat23 19 hours ago [-]
Less probable also means less optimal and you get a subpar response. More so if it's baked into its reasoning. It's intelligence will suffer unless this is some post processing thing.
shawnz 18 hours ago [-]
There is already some intentional randomness in token selection, because it actually improves the quality of responses if you intentionally don't always pick the most likely next token.
You can hide data in that randomness without impacting the quality of the response by using a sufficiently "random looking" pseudorandom bit stream instead of real random numbers.
it's going to be easy to defeat either way, like SynthID is. From apps built to remove the watermark, to simply rephrase the work with an Open Model...
Well, if the model that page uses also falls under the EU act, the output will just be watermarked differently. ;)
arcfour 18 hours ago [-]
> Honest note: Anthropic has not shipped a public Claude watermark detector yet. This tool uses rewrite-based neutralization — a meaning-preserving paraphrase with a non-Claude model — which is the attack path watermark research points to. Not affiliated with Anthropic.
Well, they should have run their own AI slop website through their tool...
nunez 17 hours ago [-]
This attack was actually pointed out in the watermarking paper linked above. The researchers added an instruction to the prompt that switches letters like a Caesar Cipher. It lowers the quality of the output from the LLM but alters the "red list" enough for a watermark detection tool to fail at detecting the watermark.
fl0id 14 hours ago [-]
also their example for rewriting just completely changes it. might as well redo it in this case (with another model or by hand)
cush 14 hours ago [-]
Should be ensloppifier.app - it somehow makes the AI sound more like AI, while also completely changing the meaning and context of the input text
From their before/after:
- Certainly! -> (removed)
- onboarding redesign -> revamping the introductory process
- this week -> (removed)
- empty states -> empty sections
- CTA heirarchy -> call-to-action sequence
- interviews -> discussions
- aligned copy with brand voice -> verbal identity
... these choices change the meaning of the text
shinryuu 22 hours ago [-]
Though if pangram should be trusted, there are still statistical artifacts that tells you that a text LLM generated. I don't find that to be implausible.
timpera 21 hours ago [-]
Alas, Pangram should not be trusted.
s_dev 20 hours ago [-]
So was it going down:
"Neutralize engine is temporarily unavailable. Try again."
infinite_spin 21 hours ago [-]
My guess is it will be similar to how Genius watermarked lyrics, using things like variants of punctuation
All of them are reasonable options. If we bias the model's output so that one of them is more likely than the others, then we can reconstruct that watermark if enough of these frames are present.
DaiPlusPlus 17 hours ago [-]
That was actually the cause of an issue I had a couple of years ago: I had hand-typed JSON using my iPad into GitHub’s online text editor and Safari helpfully used “pretentious quotes” instead of "old-school quotes" - and the JSON library used by the program to read that file had relaxed parsing rules that accepted actual JS object literals without quoted property names; so the fancy-quotes were interpreted as part of the key-names. This took ages to figure out because when human-eyeballing the JSON file it looked perfectly fine in Notepad.
kccqzy 13 hours ago [-]
I don’t doubt your experience, but many people are intimately aware of the use of proper Unicode quote characters, in any reasonable font they choose. For me, one of the first things I learned when using LaTeX is how the quotes are transformed from the source to the typeset document; since then I’ve become extremely sensitive to the kind of quotes I see.
6 hours ago [-]
miohtama 1 days ago [-]
Maybe there is a reason why Opus 5 produces such word salad conversations
andrewgleave 17 hours ago [-]
Yes. The irritating epigram / aphorism style it now uses is such a regression compared to previous Anthropic models. Probably is the case that this is due to watermarking - though hardly subtle if it is.
myko 19 hours ago [-]
So frustrating to use. And the comments generated by Claude today are unreadable garbage.
w_for_wumbo 1 days ago [-]
What happens if someone handwrites a Claude output, then someone uses that handwritten text as a reference.
Now you've got a watermarked idea which may have no direct linkage to the usage of Claude.
TheOtherHobbes 20 hours ago [-]
If the algos work as advertised, watermarked token sequences have an extremely low probability. Copying the words by hand doesn't change that.
The mechanism seems to survive editing. The extreme probabilities get a little less extreme, but are still extreme enough to be distinctive.
But it wouldn't survive paraphrasing, because the output would be entirely human and the token correlations would disappear.
It might not survive referencing if only a sentence or two is used.
The practical issue is how true the claims are. It's one thing to create a proof of concept, another to see how it works in use.
And this is potentially catastrophic for code, because the grammar and word choices of code are completely different and more fragile than standard English.
phainopepla2 1 days ago [-]
How is that different from referencing digital text that someone copied and pasted from Claude?
w_for_wumbo 1 days ago [-]
Because there's an expectation of authenticity from the written word.
If you've referenced something handwritten, you don't expect it to be the output of an LLM.
Similarly, if you quote someone word-for-word, you wouldn't anticipate their words to be flagged as Claude content, but if someone memorized Claude output word-for-word. That would still be classified as a Claude output.
Going forward you could categorize the influence of Claude on a population based off a percentage match between their spoken words with the LLM prose.
dns_snek 23 hours ago [-]
Are you worried about being accused of using LLMs to generate your work? As long as you don't plagiarize you have nothing to worry about.
Cthulhu_ 22 hours ago [-]
I'm not too sure about that, people making stuff have already gotten penalized by overzealous AI detectors, most recently Kurtzgesagt.
platinumrad 21 hours ago [-]
You can't make a blanket statement like this without knowing how the watermark is implemented.
AlecSchueler 22 hours ago [-]
What if I unknowingly read content written by Claude in various articles and it influences my own writing style?
stabbles 1 days ago [-]
It will just thread some load-bearing seams through the paragraphs.
cush 15 hours ago [-]
Certainly it can't watermark text with low entropy. If you're renaming a function using claude it won't be marked
mihaelm 1 days ago [-]
> have some kind of weird pattern baked into their text to act as a watermark.
public abstract class BaseAnimalBeanFactoryGeneratedFromClaudeFactory
gajus 1 days ago [-]
Most likely watermark will be proportional to the input/output ratio, i.e. if you input a long document and ask to make edits, it will not attempt to watermark it. On the other hand, if you provide a tweet and ask it to write an article, that will include watermark. Just a guess (and yes, it feels flawed)
invalidusernam3 19 hours ago [-]
Off the top of my head I would have thought zero width characters (eg: U+200B, U+200C) making some unique identifier sprinkled in amongst the output. But obviously far from foolproof since they could simply be removed.
kkukshtel 16 hours ago [-]
I think a lot of the examples below are projecting more complicated options, but it could also be something just as simple as using a word with a hyphen in it every prime-numbered sentence. Or any other "puzzle-y" pattern.
nprateem 24 hours ago [-]
Load-bearing==claude
siva7 1 days ago [-]
I can tell you how: Claude produces a huge wall of text with jargon ridden bullshit and invented terms no human subject matter expert would seriously use and overuse.
benrow 1 days ago [-]
I've heard that this kind of watermarking process works by biassing the statistical sampling towards a partition of the set of possible next tokens (red set and green set), at each position. It might only be a slight nudge each time, but over a sequence of tokens, the likelihood of repeating the bias by chance is increasingly improbable.
The bias is different for each position and follows a defined RNG, seeded somehow predictably.
Can be either an open algorithm, or not. If not open, then an API could be provided to determine if text is watermarked or not.
How it applies to code - maybe it could be a subtle nudge to symbol names, etc, I'm just speculating (I only read about this in passing very recently).
m-chrzan 1 days ago [-]
There's a computerphile video (https://www.youtube.com/watch?v=XZJc1p6RE78) with Dr. Mark Pound explaining a paper by John Kirchenbauer, Jonas Geiping et al. (https://arxiv.org/abs/2301.10226) that described a method for watermarking LLM output like this. It's not directly stated anywhere in the Claude support article that this is what they're using, but the properties of the watermark described seem to point to this method.
londons_explore 20 hours ago [-]
The bias has to be small enough that if you ask an LLM to repeat some passage of text like the national anthem, either from the training data or from the prompt it doesn't change random words.
Gotta be hard to tune that.
19 hours ago [-]
metalcrow 1 days ago [-]
Based on my understanding, it can only be applied to code in very limited ways: docstrings, variable names, string literals. The code itself can't really have tokens changed to another equally correct token (the foundation of the watermark) because then the code breaks! And the few places that you can do so are likely erased by formatters anyway.
matherial 15 hours ago [-]
You have less latitude than in prose, but I think you're underestimating how much can be changed without causing breakage. For example, the most likely sequence might be "if (!foo) bar; else baz;", but you can also say "if (foo) baz; else bar"; substitute "foo == 0" / "foo != 0" for even more variety. Similarly, "foo = 1; <NL> bar = 2;" can be output as "bar = 2; <NL> foo = 1;".
Keep in mind that the LLM "sees" the previous (tweaked) output and picks what makes sense based on that. There are few situations where a perturbation like that would be unrecoverable, and I assume these situations also correspond to a huge probability difference between the most likely completion and the second most likely one - in which case, the watermarking algorithm can choose not to touch the token.
cassianoleal 1 days ago [-]
> a defined RNG, seeded somehow predictably
So, an NG?
olmo23 18 hours ago [-]
PRNG
IshKebab 1 days ago [-]
If it's based on position mod 2, wouldn't inserting or deleting (or splitting/merging) words every now and then trivially defeat it?
If it is based on position mod 2 then wouldn't inserting/deleting (or splitting and merging) words every now and then defeat it?
mchusma 17 hours ago [-]
Many good comments here. It’s somewhat common for me to voice record say a blog post of product updates, more like a ramble. Then have Claude clean it up. Then argue back and forth about certain things until it’s good, then make a final pass sometimes to change a few key words. This is incredibly different than pure ai text. Presumably it will show as ai generated here, even though I would argue it is not really. So I can’t use Claude for this usecase anymore.
I think the solution is assume everything is ai generated unless told otherwise and rely on authorship/brand as a sign of quality.
xp84 6 hours ago [-]
IDK. At my job it seems like it's expected you'll use it in that way.
In academia they've got their own concerns of 'purity,' (not least of which is justifying their continued existence which is in my opinion hard to do) and they are the ones who are going to want most strongly to punish anyone who uses AI.
And perhaps "journalists," who will want to trumpet the latest government's press release being [what they'll portray as] mostly AI-generated, as a headline-grabbing "gotcha." Ironically, that may even be a story that'll be written by a fully-autonomous journalist "agent" in a newsroom that's been pruned of all human journalists!
But in business it seems to me that we're all agreeing that it's a "good" use of AI to write in that way.
gwillen 14 hours ago [-]
The watermark is contained in the choice of output tokens. If you exert so much editorial control that Claude has no meaningful freedom in choosing tokens, then it's going to fail to watermark the text, unless the text is _very_ long (in which case even a very low-bit-rate watermark will eventually accumulate enough bits to be positive.)
MattSayar 13 hours ago [-]
I've done this before at work, and I feel true ownership of the output after this workflow. Moreso than when someone from a marketing team publishes a blog with the CEO's name as the author.
margalabargala 15 hours ago [-]
You can do something like "here's my website. Read it, then try your best to write in my voice. Do your best to avoid common and uncommon tell-tale AI-isms"
If you're putting the work you say in, the result won't be obviously distinguishable. Obviously, from some of the things that get posted here, that last sentence is too much for most people to bother adding to their prompt.
sfink 15 hours ago [-]
Perhaps. But it seems like your beef is not with the presence of watermarking, it's with what people will use that watermarking for. You're not directly harmed by that blog post being labeled as AI generated. In a hypothetical (but unfortunately likely) world where everything has passed through an AI's digestive system, nobody would care.
In the meantime, it is true that this takes something away from you. But it's something you were only recently given. Now you're not given quite as much, but readers are given a little more (or rather, there's less being taken from us!)
> This is incredibly different than pure ai text.
Ok. But it's still incredibly different from pure human text. I guess the question is which provides more value? Providing the information "this text is AI watermarked" to readers? Or allowing creators to lie and claim that AI processed text was 100% human generated? I agree that people assuming that "has AI watermark" == "is AI slop" is incorrect and causes some amount of harm, but having the watermarks also pushes back on a large amount of harm already being done.
(Personally, I'm skeptical that these watermarks will ever hold up to adversarial attacks, and they haven't claimed that they will. So I think it's the usual "casual liars will be caught, determined liars will get an additional thin veneer of respectability".)
altmanaltman 14 hours ago [-]
I would push back on the fact that nobody would care. It is clear that platforms are increasing creating AI-generated as a category and it will only get more precise over time. Platforms that have built trust/aura over what type of content they host will resort to this more as they get more flooded with ai generated low effort content. I recently wrote a blog on this actually: https://decodingvibes.com/blog/aura-and-the-backlash-against...
no_multitudes 14 hours ago [-]
Simply write your own posts if you don't want people to think they are AI generated.
altmanaltman 14 hours ago [-]
But it is not incredibley different than pure ai text tho, it is literally the same as pure ai text.
I understand you're saying since you "worked with it", it is not ai generated but if you still use the final output verbatim, the writing itself is LLM generated purely.
You want to share the output by it but also position it as not ai output. But that's fundamentally dishonest.
Furthermore, if you think your approach actually creates value and can be judged on its merit, why not disclose its ai written? If you think that will make people think your content is bad then you should see that as feedback and maybe not use AI since readers don't like it.
akersten 1 days ago [-]
So my code that Claude makes, which previously was using the best (most probable) tokens for the job, will now be getting worse in random positions, to appease a voluntary EU suggestion. Love that.
neuroticnews25 22 hours ago [-]
They aren't using greedy decoding, there's enough randomness in sampling to swap some with independent signal.
akersten 18 hours ago [-]
Purely greedy or not, there is some measure of "goal outcome" that was previously being solved for with the token selection function, and the goal was "complete this text with the best (surely, otherwise what are we doing?) next part, and sometimes the best next part is a little bit random just to keep things interesting"
Now the goal is either "identify the meaningless interesting bits and swap them out with 0% loss in the direction of the original goal," or "perturb some small selection of the output towards my secondary secret goal of watermarking the text."
It would be quite impressive if they managed to identify with 100% accuracy the tokens that "don't matter" and are free to swap with whatever signalling tokens encode the AI scarlet letter, but most likely they are not 100% accurate, and that means the output is worse off than without the watermarking logic.
neuroticnews25 14 hours ago [-]
What you're saying sounds intuitively true and from what I've found modern watermarking methods measurably rise perplexity by 1-3% [0]. Gemini convinces me it doesn't matter and doesn't compound over long contexts though. I would love to see HN experts opinion.
Oh, how I laughed. That was never your code, my friend.
akersten 18 hours ago [-]
Please don't let the arbitrary selection of phrase distract you from the substance of my argument: a product that I pay for is at best no better due to this change, and highly probably worse. Why am I paying for a tool that is beholden to clandestinely satisfy some far away master?
mikro2nd 17 hours ago [-]
Good question! Why are you paying for some tool that has always been beholden to some faraway master's opaque agenda?
someguynamedq 17 hours ago [-]
We live in a world where outcomes are often more important than process
lelanthran 16 hours ago [-]
> Why am I paying for a tool that is beholden to clandestinely satisfy some far away master?
Why were you doing that before watermarking?
Same answer.
Banditoz 13 hours ago [-]
I don't know, why are you? Who's to say Anthropic wasn't already modifying output in some way for optimization or some other reason? These models are incredibly opaque.
fwlr 17 hours ago [-]
Because the tool is made by a corporation that is subject to regulation by a government, and that government has decided it’s in the best interests of society that the tool be limited in this way.
Pavilion2095 14 hours ago [-]
I assure you, this is one of the mildest things they do after training before the model reaches you.
HatchedLake721 18 hours ago [-]
Who's is it?
mikro2nd 17 hours ago [-]
That's just the point: nobody knows! It might be my code. It might be your code. Nobody knows who they stole it from.
MagicMoonlight 22 hours ago [-]
[dead]
mhjkl 18 hours ago [-]
All big LLMs already visibly watermark all their text with easy to detect annoying phrases and turns of speech that everyone is already sick of hearing. Why do AI companies keep making their products worse to appease anti-AI, it’s not like they’ll suddenly start supporting it if you do so. If you’re worried about European customers, just relax your firewalls to let more VPNs through, if the productivity boost is high they will use it anyway if their rules keep crippling their own models
IsTom 18 hours ago [-]
> if the productivity boost is high they will use it anyway if their rules keep crippling their own models
Individuals maybe, companies won't and that's where most of money is at.
mhjkl 18 hours ago [-]
This will make more money for Claude from individuals unofficially acting as meat puppets in more inefficient workflows
WD-42 16 hours ago [-]
Is it to appease anti ai or is it a method they will use to avoid training on their own slop?
jonplackett 23 hours ago [-]
We need to just stop pretending we can reliably tell if plain text is written by an LLM.
It’s just not a reasonable ask.
nunez 16 hours ago [-]
Now that the EU mandated watermarking, the point is that services (or browser extension developers) can add their own detectors to make AI-generated text obvious. It won't fix AI in print, but most of the problem is online anyway.
jonplackett 14 hours ago [-]
These things are trivial to remove though. And the whole point of it is they will also make the _detector_ available so you can then also check if you successfully removed it.
It’s not a solvable problem.
JohnKemeny 21 hours ago [-]
True, but what you can do is a one-sided guarantee. If it bears the mark, it is likely generated (or someone deliberately made it look generated).
Thus, if a news article, research article, book, student paper submission, blog post, HN comment, etc, bears the mark, it could be automatically flagged as such.
It helps detect low effort slop.
---
Caveat. If you write your own creative work and send it to Claude for "cleaning up grammar", it might insert the watermark.
jonplackett 19 hours ago [-]
The problem with pretending is that people who k ow what they’re doing get away with it while people who don’t (and don’t even use ai) get unfairly accused of using it.
There just isn’t enough information in plain text to do this and we should stop pretending there is.
If we need to verify something isn’t made with ai then we need other ways of doing so - eg looking at a document edit history, doing it as an exam, oral defense.
There are options! But pretending you can tell if text is ai will only catch out people who make no effort to hide it and will inevitably have false positives.
DanielHB 20 hours ago [-]
It seems like it would be so low effort to bypass, especially when you can just train a system (maybe even another LLM) using the watermarker validation from Anthropic themselves.
Might help with students and low-technical people passing off work as their own, but any industrial scale slop-generator should be able to bypass it trivially.
Art9681 18 hours ago [-]
Not might, will. Whether enough text is present or not to go over the detection threshold is in doubt. But the "score" will never be zero, even for human written text.
someguynamedq 17 hours ago [-]
will insert a watermark
asnelt 20 hours ago [-]
> If it bears the mark, it is likely generated (or someone deliberately made it look generated).
One could even say, the mark is load-bearing.
aabhay 1 days ago [-]
I have had a hunch for a while now that (in addition to these tools), Anthropic has actually leaned in to Claude's distinctive manner of writing since it makes the text more obviously AI generated and thus less susceptible to misuse.
That's not necessarily the same thing as a markov-style fingerprint but it could be a correlated factor.
pjm331 1 days ago [-]
I had a similar thought but I assumed they leaned in because it improved performance on coding or something like that
LoganDark 1 days ago [-]
I suspect it's because of alignment concerns. The more deeply they can integrate their principles, the harder it'll be to misuse. Or at least that's the idea.
andai 11 hours ago [-]
From what I heard several AI companies have intentionally been making the personality more cringe so people stop making it their girlfriend.
kingstnap 1 days ago [-]
It could also partly be a byproduct of examples of claude writing being in the dataset, which of course anthropic has lots and lots of and they do train on.
r_lee 22 hours ago [-]
no way. there's just no good excuse for why "load-bearing" and "worth flagging" are everywhere now, I've pretty much never seen that in the wild before
noman-land 1 days ago [-]
It's pretty trivial to command it to not speak that way. That's one of the first things you should write into the prompt. What style you want it to write in. Make it use a very concise and dry academic style with no overt LLMisms, melodramatic or flowery language, or metacommentary.
breezybottom 18 hours ago [-]
People have been posting some variant of this comment for three years, and it's no more true today. Ever notice that the "prompt engineer" career hasn't materialized?
someguynamedq 17 hours ago [-]
Prompt engineer is a requirement within every serious job now, not a job in itself
breezybottom 11 hours ago [-]
I've never known any firefighters to prompt engineer a blaze, but perhaps you don't consider that a "serious" job.
andai 11 hours ago [-]
Hey ChatGPT, what side should I make the incision on?
noman-land 17 hours ago [-]
Regular engineers still exist.
simonw 1 days ago [-]
An interesting factor of this is competition.
If Claude was the only model family they could ship a change like this and users who want to cheat (or don't like watermarks for other reasons) would just have to put up with it.
In a world with many different competing models, the risk of losing customers to other providers over this is much more real.
Maybe they've looked at the numbers and the portion of people who clearly use Claude to cheat on examples etc is so tiny that losing them to other providers isn't a problem?
I guess whoever is the policy maker is assuming that some protection is better than none and that most people will not reach for such tools.
lhd1 1 days ago [-]
Scott Aaronson spoke about this in a colloquium where he said that this was mooted at OpenAI before the decision was made by Altman to not implement it for the reasons you describe.
I’m more worried that this will degrade performance. I want the best results from a model, not the results that fit a constraint that’s not defined by me. Any increased cost or latency is also unacceptable.
andai 11 hours ago [-]
Presumably the expected cost of doing it is less than the cost of getting fined by the EU for not doing it.
nprateem 23 hours ago [-]
Either that or they want to comply with the EU AI Act when it affects them.
bramhaag 19 hours ago [-]
I cannot wait for the inevitable "I've always used Claude watermarks in my writing, even before we had LLMs!" when someone gets caught using an LLM.
ethin 1 days ago [-]
Can someone help me understand how exactly this watermarking of text works?
Given that text is, well, text, and not some kind of binary format, I don't see how any watermarking can work unless you insert characters which are invalid under Unicode. I further don't really understand how this won't be perceivable by assistive technology (the "watermark" will just appear as either unreadable characters, or if the watermark is mixed thoroughly enough into the text, it will scramble the text to any speech synthesizer and will make it really really obvious). Thus, I don't see how this wouldn't be insanely trivial to remove. And this is before we get into things being put on the clipboard. Sure, I can press the "Copy" button at the end of each response, but what I can also do is manually select the response and copy it, or only copy partial selections, or any number of other things. How does this "watermark" (or any "watermark" technology) take into account this?
So, really, to summarize this: I see no way of this actually being technologically achievable unless we revise the very core of how computers work and encodings for textual information. So I'm very curious as to how this is actually supposed to work.
benrow 1 days ago [-]
Have a look around for token biasing, or green lists. It's based on a nudge to the choice of the next token (which can always be drawn from a set of possibilities which are all probable enough).
At first I thought this approach was just the "LLM flavour" of writing, but it's way more subtle, especially as the bias is applied uniquely for each token position.
ethin 1 days ago [-]
Yeah, will do, this sounds interesting since I'm not entirely sure how this would actually be reliable to any degree. Thanks for the help, not sure why I got downvoted since I was genuinely curious.
fl0id 14 hours ago [-]
there have been papers about it, it works
resonantjacket5 1 days ago [-]
it's a statistical way. like for example maybe in your above paragraph claude maybe writes "Thus, I don't see how this wouldn't be insanely <easy>(instead of trivial) to remove" and then also says like "And this is before we <analyze> things being put on the clipboard." or maybe the i just says the word "the" in a certain pattern or frequency.
you can then consistently like figure out if it was claude that wrote the sentence. it is easy as you noted if you just get another ai to read it and then rewrite it.
MagicMoonlight 22 hours ago [-]
[dead]
andai 14 hours ago [-]
If I understand correctly, this means that any text with the "watermark" is legally uncopyrightable, including code.
Not an copyright attorney, but color printers have watermarks. That's never been an obstacle.
andai 11 hours ago [-]
AI generated content is not copyrightable, as far as I understood. A reliable watermark is a reliable indicator that a work does not fall under copyright.
For a repo, I don't know what that means. Only the AI generated lines are public domain?
luxuryballs 13 hours ago [-]
what if the output is downstream from copyrightable work? wouldn't the LLM touching it wash that off if this was the metric used?
andai 11 hours ago [-]
Well, all major LLMs have been trained on copyrighted material, so effectively all LLM output is downstream of copyrighted work.
I guess it's a bit weird. An LLM can recite copyrighted material verbatim from training data (they had to work hard to get them to stop doing that, and they haven't been entirely successful). But LLM outputs are public domain. (Except when they are a verbatim reproduction of a copyrighted work.) I'm not sure where you draw that line when it's not verbatim...
xp84 5 hours ago [-]
Especially when they say they'll let you use their tool to detect it, this just seems like a cat and mouse game. Have Claude write a long passage, then run it through another model with instructions to slightly paraphrase it. Ensure with the tool that it's no longer watermarked.
Even if all models are mandated by EU to do their own watermarking, it doesn't take a new frontier model to be capable of paraphrasing it, so you can use 2026's open models to do that paraphrasing, far into the future.
The funny part is that a lot of people have already developed an impressive ear for spotting AI-isms, so for now I'm not even sure how important it is to have this. No technique can be 100% guaranteed accurate anyway, and humans are pretty good at recognizing AI already.
fwlr 17 hours ago [-]
A surprising number of people are worried that the code they don’t read will be imperceptibly different.
drnick1 1 days ago [-]
Seems like an awful idea. I hope that that "watermark" will soon be discovered, reverse-engineered, and that tools to remove it will appear.
cassianoleal 1 days ago [-]
I hope all models adopt it.
DaSHacka 21 hours ago [-]
Thankfully, there are a variety of Chinese models that never will. I think we all know that in a few years, they will also be the only relevant offerings on the market, due to not being bogged down with over-zealous ""safety"" footguns.
wpietri 20 hours ago [-]
Your theory is that the Chinese government is thoroughly uninterested in safety or prosocial controls?
dannyw 18 hours ago [-]
Amongst Chinese labs and netizens, there's MUCH less belief/mindshare on "AGI = existential risk to humanity", "paperclip maximiser", and similar lines of thinking. AI is seen more as just a technology, and less like a scary boogyman.
Whether that's right or wrong, I'll leave to you, but there's huge differences in perspectives, and if you only get your news from Western sources and communities (and companies), you're in a bubble too. A different bubble, and arguably a more porous one, but still a bubble.
wpietri 16 hours ago [-]
The paperclip boogeyman is not the only reason, and probably not the biggest one, that models get safety/content constraints, however useful it is as a PR distraction. I agree Chinese models are different right now, but I think that's a function of their novelty and desire to compete globally.
For a taste of where I think things are headed, try asking Chinese models about Tiananmen [1]. And then take a look at the Chinese government's approach to pretty much anything that they think reduces security or social harmony. I find it hard to believe their models will be the one exception to that over the long term.
Other models will end up diffusing it and making the signal indeterministic and irrelevant.
edg5000 15 hours ago [-]
This is outrageous. I hope only Anthropic will do this. Are they going to disclose at least the specific Unicode whitespace characters used for the watermark? Or will they use some other trick?
If I heavily edit LLM output, will this still hold the watermark?
You really can't make this stuff up, it doesn't make any sense.
secretsilver 15 hours ago [-]
I don't think they are using any invisible Unicode or metadata stuff.
What happens is that AI selects similar words based on a random process.
Something like "The company had a large/big/substantial advantage".
It chooses between these words, and over a longer piece of text, the pattern will start showing, like a "choice A → choice C → choice C → choice B → choice A".
The normal-looking text will actually be a fingerprint living in the form of statistics.
I think Claude will be sharing these patterns to third parties for AI detection.
izonu 1 days ago [-]
> We’re also working to enable users and other third parties to detect Claude’s embedded watermarks and provenance metadata.
This seems to be similar in execution to Google's SynthID. I hope they release actual code the technically proficient can use, unlike SynthID which can only (afaik) be queried with Gemini's UI.
gajus 1 days ago [-]
The moment Google announced SynthID, the first domain I bought was deSynthID.com
Several open-source projects have already proven SynthID to be ineffective.
moebrowne 23 hours ago [-]
There are many free lock picking tutorials, but yet locks are still effective.
gajus 10 hours ago [-]
this gives strong "you wouldn't steal a car" vibes
mcv 12 hours ago [-]
I think it's definitely important that any AI generated content can be easily identified as such. I think the new EU law that requires that, has too many unnecessary exceptions.
So great that Anthropic is doing something about this, but it's not clear what their watermark exactly is. How do I, as a user running into some content online, know that it's generated by Claude? What is their watermark?
It sounds to me like they create the pattern in the regular text of the content, which sounds interesting, but also odd, unreliable, and may limit the content you can get out of Claude. Will it subtle change the words in order to hide this pattern in it? I don't know if that's something anyone wants.
allthetime 12 hours ago [-]
"Will it subtle change the words in order to hide this pattern in it?"
I'm assuming that's exactly how it works. How else could it?
0x_rs 16 hours ago [-]
Is the detection mechanism going to be open, free, and possible to run locally without prostrating to an opaque third-party company that will do whatever they want with the text content provided (including using it for training), and take no responsibility in case of false-positives for which there can exist no proof or evidence against by the victim? This is another useless, if not actively harmful, performative EU regulation, for which they ought to take the full blame despite the fact "AI" companies have been researching and working on watermarks, including in text, for a while now. Just copy-and-paste everything you see and let a machine decide for you if what you're reading is slop or not. Real propaganda machine doesn't care about inane rules and won't waste time with gimped mainstream models either.
>Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;
Such models already struggle not making any unnecessary or unwanted changes to a corpus, this makes them unable to by design.
letmevoteplease 16 hours ago [-]
It is of course a stupid regulation, but the upside is that it will probably accelerate growth in usage of open models that are not adversarial towards the user.
edg5000 15 hours ago [-]
I like your optimistic way of thinking. There is a lot of truth to it.
Godsend69 15 hours ago [-]
[dead]
jussy 7 hours ago [-]
Ok so they ingested the world's content, sold it back to us and now they're protecting themselves against the copyright claims under the guise of safety and user privacy whilst creating the regulatory moat that decreases competition?
What am I missing?
graypegg 16 hours ago [-]
Huh... I wonder if some big version of a bloom filter would work as well. Hash all output text, probably in chunks of a couple tens of tokens each (that would need tweaking to find the most useful hash input length I guess), and smash 'em into a bloom filter. Every month or something, Anthropic releases a giant file containing whatever huge length of bytevomit would have to be used to get an acceptable false positive ratio. (In terms of bloom filters! Meaning: still far from a perfect ratio.) Maybe one for each model they provide or something?
Then at least you could have two weak-postive signals, and a strong-negative signal. (Though one that only fits precise chunks of tokens) I'm sure I'm missing something here, but my groggy morning brain thinks that doesn't seem too bad.
case540 1 days ago [-]
I don’t like the idea of hacking a response to contain a watermark. I also don’t like the idea of false positives detections coming directly from Anthropic. If people read more AI generated content, people will probably start writing more in that style
stranded22 1 days ago [-]
The amount of times ‘delve’ appeared in general conversation in the last couple of years shows the influence LLMs have on society.
pixl97 1 days ago [-]
I have no idea why you were down voted for this. Language is alive and people adopt it from sources they hear a lot.
ack_complete 1 days ago [-]
Moreover, what if you quote text that happens to have been generated by Claude, does that bump up the AI-ness score of your source file or document?
reasonableklout 23 hours ago [-]
The flip side of this is that if AI-generated content becomes reliably identifiable and carries a stigma, then people might deliberately change their styles to be more diverse and human.
One example I've seen are junior employees at my company deliberately adopting a lowercase/less punctuation writing style so as to stand apart from AI.
Groxx 13 hours ago [-]
>Content generated by Claude may not carry a detectable mark if, for example: ...
>A file’s metadata was stripped through format conversion, re-saving, screenshots, or other means
Ah. So what essentially every single consumer-oriented media host does. Gotcha.
I fully recognise this is a hard problem, but hopefully metadata isn't the only method for media. Standard procedure is to shrink files for storage and privacy reasons, and non-visual metadata goes out the window by default.
taormina 13 hours ago [-]
New models will mark AI-generated content from day one. Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch. Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported.
So, they’ve been doing this for over a week without telling anyone?
Dkuku 12 hours ago [-]
Thats why recently it pushes so much comments in code - it has to squeeze the watermarks somewhere
mjuarez 15 hours ago [-]
Like others have said, it's not reasonable to ask this.
I propose we defeat this with the obvious: Simply, figure out what are some of the markers Claude and others will use for these tools, and sprinkle them randomly on everything we type or produce, all the time, 100%. If users flood the tools, and everything returns as AI-generated, then the tools become useless.
FeteCommuniste 15 hours ago [-]
Why would you want AI content to be indistinguishible from human output?
jobigoud 13 hours ago [-]
The markers aren't words that you can sprinkle in your prose, they are a statistical bias in the selection of the words.
KronisLV 11 hours ago [-]
> Claude models launched in the EU
Could we get an Anthropic subscription for Claude Code with data residency in the EU, so we don't get robbed blind by AWS Bedrock et al., but can have a monthly subscription like with the regular US option?
mateocafe 19 hours ago [-]
Seems to me like this creates a huge incentive to game the watermark. Also, how does it prevent having AI generate the text, then the user copy-paste it into a clean document?
nsvd2 19 hours ago [-]
The watermark is in the text. If you copy the text you're copying the watermark which is part of the text.
Computerphile on YT has a video explaining how models can fingerprint the text they produce. Essentially they modify the probabilities of word choice slightly in a predictable way.
mateocafe 18 hours ago [-]
Interesting, thanks for the explanation, will check out the video.
This would imply that a positive watermark signal is likely (but not guaranteed) to be AI generated. Also implies that a negative watermark signal is not necessarily void of AI generated text. This would create a problem if people start to trust the watermark as a heuristic, as the ability to critically evaluate the text is replaced by the search for a watermark.
Seems to me all of this is really trying to solve for "is this text bullshit" or not, which would require a different solution.
nojs 19 hours ago [-]
No mention of what data they are specifically encoding. Will it be like printing dots, traceable to the exact account that generated the text?
edg5000 15 hours ago [-]
From what I read in the comments, the model will be more biased towards certain words that otherwise would would have a very simmilar chance of appearing (e.g. very simmilar words that would not really alter the meaning of the text). If this actually works the way I think, that's a really sneaky way to hide data in data. It's clever, but hostile towards the user.
someguynamedq 16 hours ago [-]
Notice the subtle capture in you having to use Anthropic to identify Anthropic's watermarks? Smart of them
Schlagbohrer 16 hours ago [-]
I've long thought we would have some sort of verified-point-of-origin for data using a hash or cryptographic seal of some kind. I don't know the precise technical language for that but some metadata traveler that can verify the data has not been edited after creation.
plutokras 12 hours ago [-]
This feels like obvious setup for regulatory capture. It won't be long before missing "safety" watermarks are cited as the pretext for restricting Chinese models.
lorenzohess 1 days ago [-]
> Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported.
This should make it easier to catch cheaters who use Claude, right? Unless everyone runs their artifacts through some watermark and metadata sanitizer?
uncivilized 1 days ago [-]
As long as they’re in the EU.
nonestdeus 1 days ago [-]
From the linked article
> Regions. Marking will apply to output from supported models wherever Claude is offered, worldwide.
bossyTeacher 1 days ago [-]
> Unless everyone runs their artifacts through some watermark and metadata sanitizer?
It will happen if Claude tampers the text. Guaranteed.
pixl97 1 days ago [-]
Text is too low bandwidth to classify reliably without lots of false positives. Especially as people start talking like LLMs.
reasonableklout 23 hours ago [-]
The approach Pangram has taken which works pretty well is to simply lower the recall a lot but ensure the precision is very high. Which means potentially high false negative rate but low false positive rate.
tedd4u 10 hours ago [-]
I wonder if this is at least partially motivated by Anthropic's need to know what training materials are themselves Claude-generated.
swedishuser 17 hours ago [-]
I wonder after how much editing an LLM output wont be reliably detectable? And what the EU law even says about this. I find that a good LLM workflow can be to generate outlines that are then edited pretty heavily manually to fit into whatever context it will be published in.
KETpXDDzR 11 hours ago [-]
I wonder how the caveman skill will affect this. Also, you can probably put the output of Claude in another LLM to get rid of the watermark.
oliveralbertini 14 hours ago [-]
So we will use open source model to copy text from high end models ?
pasteleft 6 hours ago [-]
I don't understand any of this. If Claude is going to add some patterns (or some probablity distribution) in generated text, how can this not reduce code quality?
jgilias 16 hours ago [-]
The more they fiddle with the autocomplete system, the more they move away from the autocomplete faithfully producing the completion I need. The more it makes sense to move to an open weights model not served by them.
edg5000 15 hours ago [-]
The pull is strong, but it will take a few generations of hardware before 1 TB becomes attainable without needing 100k, but more like 10l (kinda like the value of a car, defendable as a job expense).
Anthropic has this strong repulsive effect in the way they operate, I wonder if they'll be around for long, it's hard to say at this time. The competion is fierce, so there isn't much room for shenanigans at this stage.
hoppp 15 hours ago [-]
How will they watermark code? I can understand encoding something in free form text but for example I ask it to generate a react component, will it embed watermark in the typescript code?
bargainbin 14 hours ago [-]
console.print(“these logs are a load-bearing seam, do not remove”)
It’s foolproof, I tells ya.
jrhey 17 hours ago [-]
I wrote about how this might work here without invisible characters:
I guess this is where our true colors show. There's a significant contingent of HNers who always dunk on LLM text detectors and claim that they can't possibly work, that they ruin careers, etc. But now that a lab says "OK, we'll add a real watermark", the reactions are overwhelmingly that it's still somehow wrong.
Why do feel so entitled to being able to pass LLM-generated text as our own? I get that a lot of techies aren't good at writing. I also see a lot of tech hustlers who like to use LLMs to fake human connection and compassion - I've gotten LLM-generated recruiting emails that talked at length about how the recruiter "valued" my work. Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.
Yes, LLMs are great. So is transparency. If you think an LLM writing is your new superpower, wear that badge with pride. It might mean you will lose some business from LLM haters and win some other business from like-minded customers. C'est la vie.
godd2 15 hours ago [-]
> Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.
One difference perhaps is that you think using LLMs is cheating, while others do not.
davisr 13 hours ago [-]
Having a ghostwriter in your ordinary life is absolutely cheating. It's letting someone, or something, else write for you and pass it off as your own. Personally, I'm insulted any time someone sends me LLM-generated media.
pickleRick243 13 hours ago [-]
"A detected mark provides a signal that content was processed by Claude, but is not fully conclusive."
"Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed."
FeteCommuniste 15 hours ago [-]
It saves the non-anglophones the bother of learning to write readable English, so there's that.
fl0id 14 hours ago [-]
they might be two different sets of people
j16sdiz 13 hours ago [-]
The true reason is legal.
> Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, ...
matthewsinclair 9 hours ago [-]
I wonder how this works for code generation as opposed to general text generation?
raincole 18 hours ago [-]
Why would I want a stochastic parrot that intentionally speaks less probable tokens?
Anyway, I've found a magic line that can be copied & pasted to the comment sections of most OpenAI/Anthropic news threads. This one is no difference.
The magic line:
> Doesn't matter; have DeepSeek.
wowokruyi 12 hours ago [-]
If I have Claude directly translate my original words, does that mean my original words also get watermarked?
0x_rs 12 hours ago [-]
It's safe to assume so. It says "the output can carry a Claude mark even if the underlying ideas, text, or data originated from another source" specifically about translations among other tasks. It's not strictly about generation but processing, and what level of processing is involved is entirely arbitrary, that is to say: you must feed your text into proprietary black box machines to figure it out, because it's not something you should be supposed to tell otherwise.
drra 12 hours ago [-]
yes, it's likely quite mechanical token distribution pattern so your text is going to be watermarked.
andreypk 19 hours ago [-]
Interesting technology. I wonder what else this could be used for beyond AI-content detection — e.g. provenance, model attribution, or tracking how generated content evolves through edits and transformations.
dalemhurley 1 days ago [-]
People with dyslexia and dystrophia, commonly use LLMs to proofread content. Even Anthropic admits this is a limitation.
stranded22 1 days ago [-]
Yes, I’m audhd and dyslexic.
I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too?
I feel shamed enough by society, thanks Anthropic.
Mashimo 22 hours ago [-]
Don't you think OpenAI will do this too soon?
Terretta 19 hours ago [-]
Paradoxically, one of those two firms puts considerably more effort into accommodating such differences, and the other has signed the same EU law and just hasn't performed as well rolling it out.
Both points suggest your subscription support was well chosen before.
If it's only proofreading text you've written, the changes will be minimal enough that watermarking seems impossible to me.
breezybottom 18 hours ago [-]
How could "your meaning" come through if a computer is writing it?
bramhaag 22 hours ago [-]
How exactly does this impact proofreading? You can manually apply the suggestions (typo here, unnatural sounding sentence there, etc.) the LLM gives you to your own content, and it would stay watermark-free.
Unless with "proofreading" you actually mean having the LLM write your content for you.
matheusmoreira 1 days ago [-]
People with executive dysfunction too. LLMs bring execution costs down to near zero and are therefore assistive technology.
Dilettante_ 21 hours ago [-]
"This is my emotional support gun. It makes me feel safe despite my CPTSD and is therefore assistive technology."
matheusmoreira 8 hours ago [-]
Sorry, but it essentially cured my ADHD. In my experience, AI is more effective than lisdexamfetamine at allowing me to turn my ideas into reality.
AI stigmatization is ableism.
normalaccess 16 hours ago [-]
Would this impact distillation? Reminds me of the fake roads map makers would add to their maps to detect copying.
darrinm 15 hours ago [-]
Are they doing this as a means to detect distilling? What kind of encoding would survive that process?
hparadiz 1 days ago [-]
I know you're gonna read this so I'll be blunt. This is bad for your brand.
LEDThereBeLight 1 days ago [-]
Reactionary emotional advice does no good, it just makes people want to hold their positions more defensively. If you care enough to say something, care enough to say it with reasons that might shift someone’s perspective.
SubiculumCode 13 hours ago [-]
Go ahead and follow your own advice. You made a claim. Now back it up.
Laurel1234 16 hours ago [-]
It's not up to Anthropic.
Sha1rholder 12 hours ago [-]
Who would take the responsibility of misjudgements then?
partsch 18 hours ago [-]
Perhaps one should start by looking into how the providers of LLMs obtained the training data.
lukewarm707 1 days ago [-]
no thanks.
coolboydev 15 hours ago [-]
This shows just how important open-source language models are.
wavewrangler 11 hours ago [-]
Anthropic, do ye come load-bearing gifts?!
jp0001 1 days ago [-]
OpenAI has been watermarking their images with C2PA for some time.
padolsey 20 hours ago [-]
Is this just to appease regulators? They surely know this won't work in the long run.
tiahura 17 hours ago [-]
You mean like with printer watermarks?
luciana1u 17 hours ago [-]
the watermark is the least interesting part. the interesting part is that we're now arguing about whether a machine's handwriting is legible enough to count as a signature.
aniceperson 19 hours ago [-]
That's not only load bearing — it sustains the need to detect AI content
someguynamedq 17 hours ago [-]
Watermarking text is impossible and a fool's errand
If we go by "fool's errand" as "needless or profitless endeavor", https://arxiv.org/abs/2303.11156 may already be a good enough answer to the paper you cited, so their work is already laid out for them. The green token idea is thoroughly attacked with much more effective techniques than those in the original paper through "recursive paraphrasing". Among some hypotheses in the paper, one is particularly interesting:
>These experiments provide empirical evidence that more advanced LLMs can lead to smaller TV distances. Thus, based on Theorem 1, reliable AI text detection would become increasingly difficult
GrayHerring 1 days ago [-]
Wasn't enough to play cat and mouse with ad removal, now we can also do the same with watermarking.
jp0001 1 days ago [-]
You could flip bits in the font itself, but I'm really wondering how portable this is.
michaelolenick 14 hours ago [-]
It's inevitable there will be false positives, inevitable they'll do reputational or economic damage, and inevitable plaintiff attorneys will sue on the behalf of people damaged. Making it worse, it's product liability blended with defamation. Anthropic should've told the EU to pound sand and geofenced off Claude. If they don't want to live in the dark ages, elect smarter people.
charlieyu1 16 hours ago [-]
All for more surveillance.
svaha1728 1 days ago [-]
I expect a “Prettier” for AI generated text in the near future.
dejanseo 22 hours ago [-]
> "Claude models launched on or after August 2, 2026 support marking at launch."
No Anthropic model has been launched in August.
fidotron 18 hours ago [-]
The SV obsession with neo Kabbalistic nonsense will get a whole new burst of energy from this.
someguynamedq 16 hours ago [-]
Please say more
hbn 1 days ago [-]
If the western AI companies are forced to comply with this type of BS, and develop their models to do their job while balancing a book on their head and hopping on one foot, the Chinese models just got a free pass to completely dominate the frontier.
EU regulation does it again!
Mashimo 22 hours ago [-]
If the Chinese want to sell to EU customers, they probably have to do the same.
mhjkl 18 hours ago [-]
I’ve seen Chinese open weights models say “can’t use this if you’re in Europe” in their licenses, so I doubt they would invest too much into complying
Mashimo 18 hours ago [-]
You mean the model itself? Yeah, probably not needed if you self host. If they run a service that they sell and they are a bigger player like Alibaba, I doubt they can just ignore it.
DimitriBouriez 22 hours ago [-]
What's the problem, really? Given the direction the U.S. has been heading in recent years, I wonder what really sets it apart from China. Europe needs to maintain an equal distance from both the U.S. and China.
krzyk 16 hours ago [-]
How is that BS?
I would prefer to know if given content was generated with LLM. This is information, and information should be free.
breezybottom 18 hours ago [-]
Less AI slop sounds like a win to me. Let China drown in it.
amelius 1 days ago [-]
They should just replace the spaces by one of Unicode special space characters.
Can it be circumvented? Of course. Will most people go through the trouble to circumvent it? No.
A_D_E_P_T 1 days ago [-]
If it's that simple and obvious, you'll have 10 "Remove Claude Watermark" web-apps by the end of Day 1. Most of them coded by Claude.
Hell, it'll probably happen no matter how sophisticated their watermark is. There's no watermark in text that can't be detected and removed, and no text that can't be converted to generic keyboard ASCII.
amelius 1 days ago [-]
You forgot about the cases where (1) people don't care, (2) people want to say "I used an LLM for this". I'm convinced that those cases happen more often than you think. Why not cover them with a simple mechanism? It's also in the interest of AI companies who don't want to train on AI output.
pixl97 1 days ago [-]
Depends on the pushback in different sets of users. Students for example would clean it up.
amelius 22 hours ago [-]
Sure, but let's first find out how many % of people are willing to be frank about their AI usage, and/or don't care about it. My guess is it is worthwhile to do this.
selcuka 1 days ago [-]
But the source codes of those web apps will also be watermarked. /s
RataNova 23 hours ago [-]
Those invisible spaces get wiped by the first sanitizer in any normal ide. Worse it'll instantly break parsing for configs like yaml where spaces are critical for structure
But as I remove unwanted characters with grep before layout in InDesign, someone will make a skill for removing such space characters.
ack_complete 1 days ago [-]
We already have one, our Claude setup already requires output to be 7-bit ASCII clean and scans it for such.
partiallypro 13 hours ago [-]
My question is what's to stop Google or any competitor from using watermarks of Claude or OpenAI from degrading the rankings of sites that use it but ignore or even reward sites that use Gemini. Seems like an easy thing to do for competitors, and maybe an unforeseen side effect of these types of things or regulations.
singpolyma3 1 days ago [-]
As if "AI generated content" even exists instead of LLMs being a piece of tooling that is directed by a human author.
cybice 15 hours ago [-]
на ху я?
chasing 16 hours ago [-]
Won't there instantly be tools to detect and remove/obfuscate these kinds of watermarks?
mucha 12 hours ago [-]
Absolutely! You can avoid the watermarks by writing or rewriting everything yourself.
chasing 12 hours ago [-]
I'll just have my AI do it...
morkalork 17 hours ago [-]
If they can use a cryptographic key to sign the text, will they generated keys unique to users? Seems like they'd be able to trace back to which accounts were used to say generate scam dialogue, threats, propaganda or bot content on the open internet?
VCFundedGenYer 17 hours ago [-]
> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
I feel like this is FUD. If you copy text from Claude, Ctrl Shift V it into VS Code, the IDE will give up the ghost on if weird characters are in there. And it's not like Google suddenly invented new letters or fonts either.
Practically speaking, I feel this is Google publishing misinformation.
orbital-decay 1 days ago [-]
Does it mean their models will always write slop? Making the writing non-collapsed to specific patterns seems to break any injected/learned fingerprinting.
matheusmoreira 1 days ago [-]
This is terrible news given the stigma against AI in general. I really don't want people singling me out for it.
Computer0 1 days ago [-]
So this won't be happening in the US, but in the EU:
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.
"
travisgriggs 1 days ago [-]
I would love to see what this looks like in practice. Especially in generated code. I assume this is more than insertion of non visible special unicode whitespace characters, but more in the pattern of the text content itself?
toufka 1 days ago [-]
I'm guessing - probably some textual variation on Benford's law? [1]. Trivial for compute, painful for a human.
- "Ensure distribution of vowels is in >99th percentile of human work"
- "Ensure the distribution of the letter "s" is within 99th percentile of human work"
- "Ensure the distribution of the letter "L" is periodic with periodicity within 5% of 1/N characters.
- "Ensure there is a cross-linguistic 'typo' (colour vs color) at 1/N words, where N: 1000 = Model1, 2000 = Model2, 3000 = Model3.
- "Ensure the distribution of tense error is within 99th percentile of human work"
If more than 3 dimensions have a score >99% percentile of human, let's call it watermarked...
Models can't reliably follow instructions involving their own logprobs unless they can take agentic control and use quite sophisticated dynamic grammars/structures/constraints to force this behavior in one shot (which can be slow and the dynamic grammar modification feature isn't supported in closed model APIs for safety reasons) or repeated attempts at rewriting which is expensive/slow.
Yes they can do this, but it's more likely closer to the original "red token, green token" paper: https://arxiv.org/abs/2301.10226
i.e. take half of your LLMs vocabulary, and upweight its probabilities by ~55% to the other half's ~45%, and scan for overuse of this half of all tokens. You can even choose a different half/slice for every individual user, for every individual action. You can implement this under the hood cheaply with logit-biasing.
Terretta 19 hours ago [-]
Considering how weirdly detuned tokens selections have become in Anthropic's LLM prose in recent models, there is a chance this goes unnoticed in everyday use.
Computer0 1 days ago [-]
I would hate to have any of these rules effecting my output
limbicsystem 20 hours ago [-]
I see what you did there!
wolfy1993 1 days ago [-]
IIRC, watermarking text could be as simple as training the model to use specific words/phrases more frequently than what you would expect to find in human-written text, to the point where it's highly statistically improbable that it wasn't AI generated. I assume similar logic could apply to code in the form of functions/code styling.
That's probably an over simplification. Also a solid defence that can be used against complaints about the way AI writes text.
rcxdude 22 hours ago [-]
It essentially looks like the difference between two different runs of the model with the same prompt but different seeds. The watermark is essentially a small bias in the model such that when there's multiple different tokens that could conceivably follow the previous token, the model will only pick some subset of them (the subset is derived from a hash of the previous token). This bias can then be checked for statistically (without needing access to the model and without needing the whole prompt), and for longer text where there's enough freedom in word choice you can show that it would be vanishingly improbable to accidentally follow the rules in the watermark.
AtHeartEngineer 1 days ago [-]
non visible text is extremely easy to filter with a git hook, a post tool call hook, or just a script. I doubt they are doing that
mucha 12 hours ago [-]
The watermark will be encoded in the visible text.
sixtyj 1 days ago [-]
Or grep, in a skill. /clean-cc-watermark just entered the chat…
kbelder 13 hours ago [-]
If you do the same prompt with zero noise from the US and the EU, would the difference reveal the watermark?
I suspect they'll roll out the watermark everywhere.
tech234a 1 days ago [-]
Article specifically says "worldwide"
Computer0 1 days ago [-]
I agree with your reading, I initially misread it.
guluarte 16 hours ago [-]
will this solve anything? people will build tools to remove the watermarks
1 days ago [-]
pessimizer 13 hours ago [-]
I've been thinking that they have to be doing this. It seems like a fun problem, actually - all you're trying to encode is a 1-bit message within a text with the least amount of necessary changes possible, but in a way that arbitrary fragments will show it.
My intuition is that this would be very possible, in a way that makes false positives so unlikely as to be virtually nonexistent (at a certain fragment length.) Basically all you would be trying to do is to defeat people who would deliberately screw up the signal below the fragment length, and you would try to get that fragment length to at least the size that intentional obscuring of the signal would be obvious. I could see it being possible to detect even from non-contiguous fragments interspersed with noise.
It's just 1 bit, and you don't really care if a sentence or two is slop. I'd be surprised if a PhD interested in steganography couldn't come up with a good scheme in a week. It's a QR code.
What would be scary is if they could come up with a way to detect advice from Claude i.e. you get Claude to review your work as an editor, read the output, then as a result make non-verbatim changes, and that signal still gets through. If you could do that, you could do things like tell if a pundit speaking on television has read a particular Wikipedia page. Seems impossible, but LLMs seemed impossible.
edit: there are so many unimportant language choices; ones that are even hallmarks of AI use already, like the fact that it generally picks the mode. Not always picking the mode or picking at precise distances from the mode could hide signals without significantly affecting the quality of the content.
sfink 15 hours ago [-]
I don't know how this watermarking works, but I don't need to in order to understand some things that a lot of this conversation seems to be missing.
First, the article doesn't talk about adversarial usage. As in, it's not claiming to be proof against various techniques of watermark removal (inserting words, rewriting with a different model, manual paraphrasing whether minor or extensive, etc.) It might handle some things and not others, but "I could trivially defeat this!" is not a gotcha; they haven't made that claim.
Second, basic information theory tells you a lot about what is or isn't possible. Watermarking is information. You need degrees of freedom to store that information. You can even estimate various sources of space in bits (often fractional bits.) To a first approximation, longer text has more bits of space. Language matters -- a rich (aka messy) language with lots of potential synonyms has more space. That goes for human language as well as the difference between human and programming languages. (Most programming languages have much less flexibility to them than most human languages.)
The details of what space you make use of are interesting, but speculative. In the English sentence "Ellie spat in his eye", you could look at it at a word level and say that swapping "Mary" for "Ellie" is a lot more damaging to the meaning than swapping "face" for "eye", so there are more bits of freedom in the latter. For coding, `for (int i = start(); i < end(); i++)` probably shouldn't swap `<=` in for `<`, but it could be written as `int i = start(); while (i < end()) { ...; i++; }`. (I'm not claiming this is the sort of alternative that they'd use, just an illustration of what's possible.) But there are a lot of possible places to find these bits if you look at large chunks of text. Different ones are more or less resistant to accidental or intentional information destruction, and require less or more sophistication (aka brittleness) to be extracted. (In the limit, you could require the full original prompt and encode tons of stuff by tweaking the logit selection. But it wouldn't be very useful to require the original prompt.)
Also, does this degrade model output? Yes. It reduces the bits of freedom available to the model for producing the signal. Does that degradation matter in practice? That's totally dependent on exactly what is happening, and will likely change over time and across different purposes. I hope we're past the point where people believe that setting temperature to zero produces "perfect" output in some sense. (Or should I say flawlesslesslesslesslessless output?) It used to be useful for reproducibility, at least, but my understanding is that it's no longer even good for that? Anyway, reproducibility != quality.
There are a lot of things that could be going on here. The article doesn't claim very much, just that they're encoding a signal in the output that can be extracted later. How robust the signal is in terms of the FP/FN rates is unknown. The resilience (resistance to destruction) is unknown. The impact on the output quality is unknown. Even the question of whether this will make AI slop less sloppy is unknown; maybe this means we'll see a little less exact repetition of "I have the whole picture now" and instead it'll sometimes be "Now I see the entire picture"? Can we dare to hope for an occasional "Ok, this time I got it, boss"? That would be a (very minor) quality improvement.
Yet another reason to support open-weight alternatives, I guess.
Laurel1234 16 hours ago [-]
Why do you feel the need to deceive readers on whether your content is AI generated?
qweqwe14 13 hours ago [-]
Why do you feel the need to label the content when the quality of the content will speak for itself?
It doesn't matter where the content comes from, only the quality/usefulness matters. If you are opposed to this idea, the next decades are going to be very tough for you :)
Laurel1234 12 hours ago [-]
If the "author" can't be fucked writing their own work I don't want to be fucked reading even the couple paragraphs it takes to be offput by clanker slop.
> It doesn't matter where the content comes from
It absolutely matters to a lot of people. Things like these just provide transparency and allow people to have the necessary information to make their own decisions.
beambot 13 hours ago [-]
Not sure I understand your implication... because I'm the primary consumer of the content AI generates for me. I'd rather not have its output adulterated.
Laurel1234 12 hours ago [-]
These are magic black boxes whose content gets constantly adulterated with no input from you.
Models from openAI had instructions in their system prompt not to talk about gremlins and goblins.
Anthropic got caught acting different if you're Chinese. Grok got its Nazi dialed turned to 11 until it started calling itself Mechahitler cause Musk found it too left leaning on Twitter.
h0mie 13 hours ago [-]
Why is it deception if you never claimed it was unassisted? The assumption now is that most text produced is already assisted by an AI to some extent.
Laurel1234 12 hours ago [-]
If you have no problem with your AI content being labeled as AI content then a watermark that does exactly that is totally fine with you, no?
colesantiago 1 days ago [-]
Good.
They should make it easier, to detect slop so we can ignore it quickly.
I hope Pangram makes an API or an extension to analyze a page to detect slop on a page and then closes the tab immediately.
Nobody should be wasting time on garbage LLM output in code, text, image and videos.
pixl97 1 days ago [-]
Panagram is a scam.
cubefox 18 hours ago [-]
It's not. Pangram is quite accurate. Not being perfect doesn't make it a scam.
colesantiago 1 days ago [-]
(This is the part where you provide extensive extraordinary evidence to your claim)
pixl97 1 days ago [-]
No, they are the ones making claims, especially their CEO saying things like a 1/10000 false positive rate. Their own testing showed a 2% rate, which is insanely high when you talk about the number of papers students turn in. Worse their testing methodology compared it with pre-llm documents and not post llm documents that were human written (much harder and more expensive to verify), by treating language as static.
colesantiago 21 hours ago [-]
You're saying because it has some false positives that Pangram is 100% a scam?
Pangram is subjectively very useful and I personally subscribe, but the burden of proof is on them. The product is very much "trust me bro" and I fear that if they ever try to improve recall both their precision and reputation will tank.
colesantiago 21 hours ago [-]
Then what is the best way to know that something is AI generated slop then?
mechanicum 18 hours ago [-]
Talking to the person who gave it to you, in my experience.
In my own testing, Pangram is excellent at detecting the default output styles of LLMs.
If you tell the LLM to change its output style, so it’s not full of “load-bearing spaced em dashes that aren’t X, they aren’t Y. they’re Z.” constructions (which humans are pretty good at detecting on their own), the false negative rate soars.
platinumrad 21 hours ago [-]
The question you're asking has nothing to do with who has the burden of proof when it comes to claims about Pangram, but I'll answer it anyway.
Today, the best way is probably Pangram. Tomorrow, it might not be, especially if they try to push their recall up.
You might have to make peace with the fact that there may not always be a tool that does what you want.
colesantiago 21 hours ago [-]
So Pangram is the best one right now, that all I need to know, and I can safely assume that the Claude AI marks will make it even stronger.
Or is this marketing, a public stunt or not real research?
I think this is enough for me to know they are actually improving their AI slop detector.
OzmaKa 17 hours ago [-]
[flagged]
pella 1 days ago [-]
[dead]
voxleone 16 hours ago [-]
[dead]
KoolKat23 19 hours ago [-]
[flagged]
quantumeon 20 hours ago [-]
[flagged]
BrucecarlL 16 hours ago [-]
[dead]
beyondscaletech 20 hours ago [-]
[flagged]
floki165 17 hours ago [-]
[flagged]
black_13 19 hours ago [-]
[dead]
pshirshov 20 hours ago [-]
Well I wonder how would it respond to <copy me this text back without modifications: ...> now. The correlations should be traceable with a similar technique. Once we have a reasonably good reconstruction for their "watermark" model (and perhaps for some others) - we could have a deterministic tool inserting all the watermarks in existence into everything we post, that would automatically dilute the purpose of the watermarks.
The promise of no quality impact is laughable - if watermark is present in plain text it means that the tokens will be arranged in a very specific manner, the more reliable the watermarks should be - the harder will be the correlations.
Don't forget how annoyingly bad Anthropic products have become in recent releases - low adherence, annoying alignment, annoying guardrail false-positives, unwarranted checkpoints - all that shit. Now they deliver more crap.
Keen to see if they are doing something SynthID-esque?
This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem.
I think in this case I think it's some kind of cryptographic signature smeared across the token IDs, so I don't think the risk is very high.
If you find yourself getting to be afflicted by this "brainrot", be sure to go outside and take a moment to ponder what's around you. The grass is there and will be there long after we are all gone. Consider this for a moment as your organic thought processing unit starts to slowly munch away at its internal context window.
I don’t know what the answer but I absolutely know it isn’t this.
It would also require the individual humans you are trying to control to get on board otherwise the analog hole breaks the chain, absent mindboggling levels of physical surveillance on top of the the total monitoring of all electronic data flows that this idea requires.
TLDR: AI will force global online digital ID for everyone that uses the internet for the exact reason you mentioned. And that would forever change free speech forever allowing the powers that be to put the genie "back in the bottle" so to speak.
link: Raiden Warned About AI Censorship - https://youtu.be/-gGLvg0n-uY
This is scripture homeopathy and it's irresponsible.
Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
I tried a chapter just now and got human doing that, but I'm not invested enough to run a hundred samples today. But it sounds like it would be an alright way to audit it? I will confess I'm pretty skeptical you could ever eliminate false positives here though. I can often get an ai sense from some writing on my own but I doubt it would be better than 90% accurate, and "ai plus human editing" might screw with that anyway, stuff like that. I would have preferred we just never developed this kind of thing so I wouldn't have to guess.
if they used older texts as training data, to some extent pangram would just be an age classifier for writing style.
Also, this doesn't even consider the case where people use LLMs to translate their original works. Or people that use it for spelling/grammar checks.
Personally, I believe these checkers do more harm than good. Any false positive can ruin someones life.
I recently heard someone say "that's genuinely the exact solution I was looking for" and had to do a double take.
https://www.pangram.com/research/model-card/pangram-4
> Pangram 4 achieves a 0.0041% false positive rate (roughly 1 in 24,000) on 1,000,000 human-written English FineWeb evaluation examples
> Overall False Negative Rate is 0.3396% on English AI generations (26 generator models)
But ultimately, when you are powerless and can't afford to do the fighting: I'm convinced the only way to protect yourself is to be very mindful about your writing style, and to deliberately corrupt the language through objectively wrong "stylistic elements".
Are you talking about pieces that were fully human-written with zero AI editing/rewriting etc? If so, what makes you think that false positives will happen there? They aren't looking for "writing styles" or emdashes etc. They are using watermarks and metadata.
If you're talking about people using AI to copy-edit text they manually wrote, this was explicitly called out in the article:
> A detected mark provides a signal that content was processed by Claude, but is not fully conclusive. Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example: Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source; The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.
They could be doing invisible and vaguely-harmless Unicode stuff. Insertion of zero-width joiners and non-joiners, replacement of regular spaces with non-breaking spaces, building spaces from multiple hairline spaces, intentional use of non-NFC-normalized codepoint sequences for accented characters, etc.
Text with all this junk in it still reads the same; it just might wrap a little strangely, or not byte-match / collate correctly in a database (and Anthropic has never made a guarantee that their models would be capable of emitting text with these properties, so that’s fine.)
And, importantly, no regular text or document editor would insert these things (especially in the useless places you could insert them for watermarking.) You only really see them in text that’s been explicitly typeset for a specific layout (e.g. in text-containing SVGs, website mastheads, or game HUDs) or for print publication.
Of course, if this is the technique they end up using, then it’s very simple to strip it out by canonicalizing the text (i.e. Unicode-normalizing it + stripping out invisible layout characters + replacing “weird spaces” with regular ones, etc. Essentially the same thing many sites already do to user-generated content to prevent users from using Unicode features to break the page’s layout.
[0]: https://deepwalker.xyz/blog/evaluating-synthid-watermark-rob...
You then look at the tokens actually picked to see how closely they follow this pattern that isn't connected to the meaning of the tokens. With enough text, you can then analyze the chance of it happening by chance verses being because the generation of the tokens was done using the algorithm, and you can save a positive result until you are arbitrarily sure. There is a chance of a false positive, but the chance of a false positive approaches the chance that the murderer happened to have fingerprints that matched your and both forensics labs happened to have mixed up the dna tests and the eye witness happened to misremember the face and your phone gps happened to glitch out and put you at the murder scene at the time of the crime all happening. It is theoretically possible only in the same sense that quantum teleporting a cat is theoretically possible.
The real question is how much text do they need for a given level of certainty and what do they check for. If they flag a positive at a p value <.01, that's a problem. If they can reasonably get a p value of < 1e-12 in only a few paragraphs of text, that is effectively no false positives (but a lot of 'too short to analyze' outcomes).
Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.
Note, there are many ways to represent words visually on computers that look identical
Theres no way that’s what they are doing.
What about "watermarked long-form code"? I'm having a hard time understanding how a model could watermark not prose, but functional/semantic text like code, that actually has meaning. You can't switch our the characters, you can't use various types of whitespace, you can't add arbitrary code comments, and a lot of other restrictions. Is there any state of the art methods for watermarking code without affecting the quality/correctness?
The way I use LLMs (and I'd advice everyone to do the same) there really isn't, the agent implements things exactly how I want them, or I use the agent to massage it into the exact bit-by-bit version I imagined when I first sent the prompt afterwards. I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually, although I know it's a popular approach taken by many.
> And of course you can add arbitrary comments; my Claude-generated code is very verbose.
So watermarking for all users who allow code comments from agents, no watermarking for us who force the agents to never write a single code comment? Alright, I'd be fine with that.
And the reason to let Claude make worse code than a professional would by hand is basically suppressed demand. Since programmers are expensive, previously code mostly got written when a large number of dollars were on the line, or when an individual programmer did something not economically optimum (e.g., hobby project).
That left a whole lot of somewhat less valuable software unwritten. It's the economic space that no-code tools have been nibbling on for years. One way to think of things like Claude Code is as effectively no-code tools. Pre-LLM no-code tools would produce data structures that got executed by special environments without ever being seen or tuned by a human. Claude Code can be used just like that, with text as the input and python as the intermediate representation that nobody ever looks at.
That approach probably isn't sustainable for what we professional programmers would call a serious project. Claude can easily get in over its head and I expect that its code decays over time, in a fashion similar to how many human teams get in a state where they just have to rewrite everything. But faster, I'd expect.
But there are a lot of unserious projects that previously would have never been created. E.g., a quick app to manage your little league team, or a bit of in-house business stuff in the "a little hard to do with a spreadsheet" range.
You never generate throwaway code used to test an external service? or try out an interface idea? There's a lot of code that's only meant to be ran once. I often dont even care what language it's written in.
And save/persist it? No, most of any experimental stuff goes into /tmp which gets cleared out on reboot, nothing I care to save in any repository. Or just "show me how this would look like" and then it's only in the session itself (and the logs/state I suppose, technically...).
In cases where one token is extremely likely, it'll randomly be red or green and still be picked in either case as it is simply the best (or only) option. So you'll have more tokens that don't show a pattern either way (half of these cases will match and half won't, just the same as if a human wrote it). Meaning you'll need more instances where multiple tokens were all likely to see if there is a pattern. Given the check algorithm can't identify these cases, it can only judge on the overall text, so the more strict a language, the more the length requirement scales.
Where I wonder if this keeps working is in tool calls. Often, you don't take code straight from the llm, you take the results of a tool call to edit already existing code. It might be that the result of this leads to far too few signals to pick up, meaning that this only works when one does significant generation with a single model (even swapping between different models, at least by different companies, breaks this just as much as having a human write parts of the code).
Think of it like finding a loaded dice. A dice that has a slight bias in a few dozen roles is just random chance. If that bias continues after hundreds of thousands of roles, the dice is loaded. But will a code base have enough samples, especially when edits made from tool calls? I could see this being unable to detect things at the size of a reasonable PR and only being useful for massive sets of changes and only if the person behind them didn't structure their AI usage to avoid detection.
Here's the strawman: The text-based watermarking is going to be done procedurally instead of generatively. Maybe they add some sequence of zero-width Unicode characters to all generated text at certain intervals. Then, there is effectively no false positive possible (because humans would [effectively] never type such sequences of unicode naturally). It may survive some editing (depending on how you select/edit the characters), and it's possible to be stripped (false negatives).
There is no room for false positive here in the same way you can't randomly find a collision in a hash function if it's strong enough. Like the rate is so infinitesimal that it is effectively zero.
Now replace random number sequence with prompted string of words. And instead of using the PRNG on every word I use it every n words. If the generated text is sufficiently long I can tell by matching the expected deterministic pattern.
You can defeat it by changing the words yourself and triggering a false negative but there isn't really any room for a false positive if the text is long enough and the pattern matches perfectly. If the pattern doesn't match then I can compute a probability.
I'd like to know a lot more about how that works.
A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know exactly what I want to happen, it will NOT be OK if those refactors have some kind of weird pattern baked into their text to act as a watermark.
I guess this may be covered by this:
> Content generated by Claude may not carry a detectable mark if, for example: [...] The passage is very short, leaving too little text for a reliable signal;
You can carefully select which pseudorandom number generator (prng) you use to be able to id text of a certain length. I expect there is some performance characteristic you have to manage since you're doing this on every inference, but once you do that it doesn't change the output in any meaningful way (the prng is still a statistically valid prng, it just happens to let you check if the output used that prng)
The point is that you can do this simply by swapping to a different RNG, which isn't noticeable to the end user, and while it changes the output, it's not any different from how using a different seed or being lumped in a different batch will change the output.
^ excerpt:
> So then to watermark, instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI. That won’t make any detectable difference to the end user, assuming the end user can’t distinguish the pseudorandom numbers from truly random ones. But now you can choose a pseudorandom function that secretly biases a certain score—a sum over a certain function g evaluated at each n-gram (sequence of n consecutive tokens), for some small n—which score you can also compute if you know the key for this pseudorandom function.
The LLM presumably generates f(input, RNG) but we only can observe f(RNG).
... though I'm not sure why that would be preferable over a coarse rolling checksum over all of the output. Seems like that wouldn't influence output, would be equally imperceptible, and probably easier to calculate (compared to "hash seed times running all LLMs supported times number of RNG algorithms, to see if output matches").
Presumably there's some other trick, or it's a red herring / failed experiment and not what they actually do in practice.
there are many simpler methods, for example you can have a tiny windowed transformer operating on the output text and all you do is alter certain words (that don't change meanings) to maximize its surprise. the tiny language model will have a special training regime to build up a somewhat unique view of the language.
we are talking about a 0.5 bit watermark here (existence). I would have zero confidence in being able to reliably remove such a watermark from pretty much any medium.
My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.
Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.
So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).
Wouldn't you need the prompt to know the probability of the next token?
There are words/tokens that are heavily correlated to the prompt (a yes or a no, for example), and then there are others that are going to be less so (adjectives with a lot of synonyms for example).
Given a text, you can identify what the "load bearing" and auxiliary words/chunks are. Then, looking only at the auxiliary words/chunks, you should, in principle, be able to determine what other wordings could have gone there instead. From this, you can, very roughly, recreate the token probability distribution that was in effect when those tokens were generated. With the probability distribution in hand for enough chunks of text, you can start inferring properties about the RNG process that was used to sample from those distributions.
But then, this notion of "load bearing" vs "auxiliary" can be expressed directly in the probability distributions. A load bearing token just has a very high probability, and thus any RNG bias that may have been in effect will likely be swallowed in the distribution. So the parts of the text that are highly dependant on the prompt will naturally not be contributing much information about he RNG in the first place.
Odd variable naming? Stylistic choices that are watermarked?
Or as someone else noted further down in the comments, it could be more subtle:
Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.
Whatever it is, I'm sure it's load-bearing.
I might be imagining things of course. But comments would be great fit for this use case.
edit: as an aside - I actually use extensions to collapse comments and change the color to be less intrusive.
If it works like people describe - on the every nth token or something - then the mark will be left in the chain of thought and discussion with the model - not in the code artifacts.
You can hide data in that randomness without impacting the quality of the response by using a sufficiently "random looking" pseudorandom bit stream instead of real random numbers.
I previously worked on a project to do that here: https://github.com/shawnz/textcoder
By the way: https://x.com/alexcdot/status/2087078010524406137
Seems like this would only catch the most unsophisticated cases.
Count load-bearing words using two different algorithms in a belt-and-braces fashion
Well, they should have run their own AI slop website through their tool...
From their before/after:
... these choices change the meaning of the text"Neutralize engine is temporarily unavailable. Try again."
https://www.pcmag.com/news/genius-we-caught-google-red-hande...
> total = calculate(items)
> result = calculate(items)
> value = calculate(items)
> amount = calculate(items)
All of them are reasonable options. If we bias the model's output so that one of them is more likely than the others, then we can reconstruct that watermark if enough of these frames are present.
The mechanism seems to survive editing. The extreme probabilities get a little less extreme, but are still extreme enough to be distinctive.
But it wouldn't survive paraphrasing, because the output would be entirely human and the token correlations would disappear.
It might not survive referencing if only a sentence or two is used.
The practical issue is how true the claims are. It's one thing to create a proof of concept, another to see how it works in use.
And this is potentially catastrophic for code, because the grammar and word choices of code are completely different and more fragile than standard English.
Similarly, if you quote someone word-for-word, you wouldn't anticipate their words to be flagged as Claude content, but if someone memorized Claude output word-for-word. That would still be classified as a Claude output.
Going forward you could categorize the influence of Claude on a population based off a percentage match between their spoken words with the LLM prose.
public abstract class BaseAnimalBeanFactoryGeneratedFromClaudeFactory
The bias is different for each position and follows a defined RNG, seeded somehow predictably.
Can be either an open algorithm, or not. If not open, then an API could be provided to determine if text is watermarked or not.
How it applies to code - maybe it could be a subtle nudge to symbol names, etc, I'm just speculating (I only read about this in passing very recently).
Gotta be hard to tune that.
Keep in mind that the LLM "sees" the previous (tweaked) output and picks what makes sense based on that. There are few situations where a perturbation like that would be unrecoverable, and I assume these situations also correspond to a huge probability difference between the most likely completion and the second most likely one - in which case, the watermarking algorithm can choose not to touch the token.
So, an NG?
If it is based on position mod 2 then wouldn't inserting/deleting (or splitting and merging) words every now and then defeat it?
I think the solution is assume everything is ai generated unless told otherwise and rely on authorship/brand as a sign of quality.
In academia they've got their own concerns of 'purity,' (not least of which is justifying their continued existence which is in my opinion hard to do) and they are the ones who are going to want most strongly to punish anyone who uses AI.
And perhaps "journalists," who will want to trumpet the latest government's press release being [what they'll portray as] mostly AI-generated, as a headline-grabbing "gotcha." Ironically, that may even be a story that'll be written by a fully-autonomous journalist "agent" in a newsroom that's been pruned of all human journalists!
But in business it seems to me that we're all agreeing that it's a "good" use of AI to write in that way.
If you're putting the work you say in, the result won't be obviously distinguishable. Obviously, from some of the things that get posted here, that last sentence is too much for most people to bother adding to their prompt.
In the meantime, it is true that this takes something away from you. But it's something you were only recently given. Now you're not given quite as much, but readers are given a little more (or rather, there's less being taken from us!)
> This is incredibly different than pure ai text.
Ok. But it's still incredibly different from pure human text. I guess the question is which provides more value? Providing the information "this text is AI watermarked" to readers? Or allowing creators to lie and claim that AI processed text was 100% human generated? I agree that people assuming that "has AI watermark" == "is AI slop" is incorrect and causes some amount of harm, but having the watermarks also pushes back on a large amount of harm already being done.
(Personally, I'm skeptical that these watermarks will ever hold up to adversarial attacks, and they haven't claimed that they will. So I think it's the usual "casual liars will be caught, determined liars will get an additional thin veneer of respectability".)
I understand you're saying since you "worked with it", it is not ai generated but if you still use the final output verbatim, the writing itself is LLM generated purely.
You want to share the output by it but also position it as not ai output. But that's fundamentally dishonest.
Furthermore, if you think your approach actually creates value and can be judged on its merit, why not disclose its ai written? If you think that will make people think your content is bad then you should see that as feedback and maybe not use AI since readers don't like it.
Now the goal is either "identify the meaningless interesting bits and swap them out with 0% loss in the direction of the original goal," or "perturb some small selection of the output towards my secondary secret goal of watermarking the text."
It would be quite impressive if they managed to identify with 100% accuracy the tokens that "don't matter" and are free to swap with whatever signalling tokens encode the AI scarlet letter, but most likely they are not 100% accurate, and that means the output is worse off than without the watermarking logic.
[0] https://ieeexplore.ieee.org/document/11348107/
Oh, how I laughed. That was never your code, my friend.
Why were you doing that before watermarking?
Same answer.
Individuals maybe, companies won't and that's where most of money is at.
It’s just not a reasonable ask.
It’s not a solvable problem.
Thus, if a news article, research article, book, student paper submission, blog post, HN comment, etc, bears the mark, it could be automatically flagged as such.
It helps detect low effort slop.
---
Caveat. If you write your own creative work and send it to Claude for "cleaning up grammar", it might insert the watermark.
There just isn’t enough information in plain text to do this and we should stop pretending there is.
If we need to verify something isn’t made with ai then we need other ways of doing so - eg looking at a document edit history, doing it as an exam, oral defense.
There are options! But pretending you can tell if text is ai will only catch out people who make no effort to hide it and will inevitably have false positives.
It seems it would get as simple as:
Might help with students and low-technical people passing off work as their own, but any industrial scale slop-generator should be able to bypass it trivially.One could even say, the mark is load-bearing.
That's not necessarily the same thing as a markov-style fingerprint but it could be a correlated factor.
If Claude was the only model family they could ship a change like this and users who want to cheat (or don't like watermarks for other reasons) would just have to put up with it.
In a world with many different competing models, the risk of losing customers to other providers over this is much more real.
Maybe they've looked at the numbers and the portion of people who clearly use Claude to cheat on examples etc is so tiny that losing them to other providers isn't a problem?
I guess whoever is the policy maker is assuming that some protection is better than none and that most people will not reach for such tools.
https://youtu.be/9udWn1Hlj_s?si=VWOiK5-y4zcyDoHI
Given that text is, well, text, and not some kind of binary format, I don't see how any watermarking can work unless you insert characters which are invalid under Unicode. I further don't really understand how this won't be perceivable by assistive technology (the "watermark" will just appear as either unreadable characters, or if the watermark is mixed thoroughly enough into the text, it will scramble the text to any speech synthesizer and will make it really really obvious). Thus, I don't see how this wouldn't be insanely trivial to remove. And this is before we get into things being put on the clipboard. Sure, I can press the "Copy" button at the end of each response, but what I can also do is manually select the response and copy it, or only copy partial selections, or any number of other things. How does this "watermark" (or any "watermark" technology) take into account this?
So, really, to summarize this: I see no way of this actually being technologically achievable unless we revise the very core of how computers work and encodings for textual information. So I'm very curious as to how this is actually supposed to work.
At first I thought this approach was just the "LLM flavour" of writing, but it's way more subtle, especially as the bias is applied uniquely for each token position.
you can then consistently like figure out if it was claude that wrote the sentence. it is easy as you noted if you just get another ai to read it and then rewrite it.
Relevant comment from a few days ago:
https://news.ycombinator.com/item?id=49203613
For a repo, I don't know what that means. Only the AI generated lines are public domain?
I guess it's a bit weird. An LLM can recite copyrighted material verbatim from training data (they had to work hard to get them to stop doing that, and they haven't been entirely successful). But LLM outputs are public domain. (Except when they are a verbatim reproduction of a copyrighted work.) I'm not sure where you draw that line when it's not verbatim...
Even if all models are mandated by EU to do their own watermarking, it doesn't take a new frontier model to be capable of paraphrasing it, so you can use 2026's open models to do that paraphrasing, far into the future.
The funny part is that a lot of people have already developed an impressive ear for spotting AI-isms, so for now I'm not even sure how important it is to have this. No technique can be 100% guaranteed accurate anyway, and humans are pretty good at recognizing AI already.
Whether that's right or wrong, I'll leave to you, but there's huge differences in perspectives, and if you only get your news from Western sources and communities (and companies), you're in a bubble too. A different bubble, and arguably a more porous one, but still a bubble.
For a taste of where I think things are headed, try asking Chinese models about Tiananmen [1]. And then take a look at the Chinese government's approach to pretty much anything that they think reduces security or social harmony. I find it hard to believe their models will be the one exception to that over the long term.
[1] https://en.wikipedia.org/wiki/1989_Tiananmen_Square_protests...
If I heavily edit LLM output, will this still hold the watermark?
You really can't make this stuff up, it doesn't make any sense.
What happens is that AI selects similar words based on a random process.
Something like "The company had a large/big/substantial advantage".
It chooses between these words, and over a longer piece of text, the pattern will start showing, like a "choice A → choice C → choice C → choice B → choice A".
The normal-looking text will actually be a fingerprint living in the form of statistics.
I think Claude will be sharing these patterns to third parties for AI detection.
This seems to be similar in execution to Google's SynthID. I hope they release actual code the technically proficient can use, unlike SynthID which can only (afaik) be queried with Gemini's UI.
Several open-source projects have already proven SynthID to be ineffective.
So great that Anthropic is doing something about this, but it's not clear what their watermark exactly is. How do I, as a user running into some content online, know that it's generated by Claude? What is their watermark?
It sounds to me like they create the pattern in the regular text of the content, which sounds interesting, but also odd, unreliable, and may limit the content you can get out of Claude. Will it subtle change the words in order to hide this pattern in it? I don't know if that's something anyone wants.
I'm assuming that's exactly how it works. How else could it?
>Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;
Such models already struggle not making any unnecessary or unwanted changes to a corpus, this makes them unable to by design.
What am I missing?
Then at least you could have two weak-postive signals, and a strong-negative signal. (Though one that only fits precise chunks of tokens) I'm sure I'm missing something here, but my groggy morning brain thinks that doesn't seem too bad.
One example I've seen are junior employees at my company deliberately adopting a lowercase/less punctuation writing style so as to stand apart from AI.
>A file’s metadata was stripped through format conversion, re-saving, screenshots, or other means
Ah. So what essentially every single consumer-oriented media host does. Gotcha.
I fully recognise this is a hard problem, but hopefully metadata isn't the only method for media. Standard procedure is to shrink files for storage and privacy reasons, and non-visual metadata goes out the window by default.
So, they’ve been doing this for over a week without telling anyone?
I propose we defeat this with the obvious: Simply, figure out what are some of the markers Claude and others will use for these tools, and sprinkle them randomly on everything we type or produce, all the time, 100%. If users flood the tools, and everything returns as AI-generated, then the tools become useless.
Could we get an Anthropic subscription for Claude Code with data residency in the EU, so we don't get robbed blind by AWS Bedrock et al., but can have a monthly subscription like with the regular US option?
Computerphile on YT has a video explaining how models can fingerprint the text they produce. Essentially they modify the probabilities of word choice slightly in a predictable way.
This would imply that a positive watermark signal is likely (but not guaranteed) to be AI generated. Also implies that a negative watermark signal is not necessarily void of AI generated text. This would create a problem if people start to trust the watermark as a heuristic, as the ability to critically evaluate the text is replaced by the search for a watermark.
Seems to me all of this is really trying to solve for "is this text bullshit" or not, which would require a different solution.
This should make it easier to catch cheaters who use Claude, right? Unless everyone runs their artifacts through some watermark and metadata sanitizer?
> Regions. Marking will apply to output from supported models wherever Claude is offered, worldwide.
It will happen if Claude tampers the text. Guaranteed.
Anthropic has this strong repulsive effect in the way they operate, I wonder if they'll be around for long, it's hard to say at this time. The competion is fierce, so there isn't much room for shenanigans at this stage.
It’s foolproof, I tells ya.
https://www.jrzs.dev/blog/claude-watermarking-ai-text
Why do feel so entitled to being able to pass LLM-generated text as our own? I get that a lot of techies aren't good at writing. I also see a lot of tech hustlers who like to use LLMs to fake human connection and compassion - I've gotten LLM-generated recruiting emails that talked at length about how the recruiter "valued" my work. Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.
Yes, LLMs are great. So is transparency. If you think an LLM writing is your new superpower, wear that badge with pride. It might mean you will lose some business from LLM haters and win some other business from like-minded customers. C'est la vie.
One difference perhaps is that you think using LLMs is cheating, while others do not.
"Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed."
> Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, ...
Anyway, I've found a magic line that can be copied & pasted to the comment sections of most OpenAI/Anthropic news threads. This one is no difference.
The magic line:
> Doesn't matter; have DeepSeek.
I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too?
I feel shamed enough by society, thanks Anthropic.
Both points suggest your subscription support was well chosen before.
https://translate.kagi.com/proofread
Unless with "proofreading" you actually mean having the LLM write your content for you.
AI stigmatization is ableism.
>These experiments provide empirical evidence that more advanced LLMs can lead to smaller TV distances. Thus, based on Theorem 1, reliable AI text detection would become increasingly difficult
No Anthropic model has been launched in August.
EU regulation does it again!
I would prefer to know if given content was generated with LLM. This is information, and information should be free.
Can it be circumvented? Of course. Will most people go through the trouble to circumvent it? No.
Hell, it'll probably happen no matter how sophisticated their watermark is. There's no watermark in text that can't be detected and removed, and no text that can't be converted to generic keyboard ASCII.
U+2800 or U+3164 would be nice.
But as I remove unwanted characters with grep before layout in InDesign, someone will make a skill for removing such space characters.
I feel like this is FUD. If you copy text from Claude, Ctrl Shift V it into VS Code, the IDE will give up the ghost on if weird characters are in there. And it's not like Google suddenly invented new letters or fonts either.
Practically speaking, I feel this is Google publishing misinformation.
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from.
"- "Ensure distribution of vowels is in >99th percentile of human work"
- "Ensure the distribution of the letter "s" is within 99th percentile of human work"
- "Ensure the distribution of the letter "L" is periodic with periodicity within 5% of 1/N characters.
- "Ensure there is a cross-linguistic 'typo' (colour vs color) at 1/N words, where N: 1000 = Model1, 2000 = Model2, 3000 = Model3.
- "Ensure the distribution of tense error is within 99th percentile of human work"
If more than 3 dimensions have a score >99% percentile of human, let's call it watermarked...
- 1) https://en.wikipedia.org/wiki/Benford%27s_law
Yes they can do this, but it's more likely closer to the original "red token, green token" paper: https://arxiv.org/abs/2301.10226
i.e. take half of your LLMs vocabulary, and upweight its probabilities by ~55% to the other half's ~45%, and scan for overuse of this half of all tokens. You can even choose a different half/slice for every individual user, for every individual action. You can implement this under the hood cheaply with logit-biasing.
That's probably an over simplification. Also a solid defence that can be used against complaints about the way AI writes text.
I suspect they'll roll out the watermark everywhere.
My intuition is that this would be very possible, in a way that makes false positives so unlikely as to be virtually nonexistent (at a certain fragment length.) Basically all you would be trying to do is to defeat people who would deliberately screw up the signal below the fragment length, and you would try to get that fragment length to at least the size that intentional obscuring of the signal would be obvious. I could see it being possible to detect even from non-contiguous fragments interspersed with noise.
It's just 1 bit, and you don't really care if a sentence or two is slop. I'd be surprised if a PhD interested in steganography couldn't come up with a good scheme in a week. It's a QR code.
What would be scary is if they could come up with a way to detect advice from Claude i.e. you get Claude to review your work as an editor, read the output, then as a result make non-verbatim changes, and that signal still gets through. If you could do that, you could do things like tell if a pundit speaking on television has read a particular Wikipedia page. Seems impossible, but LLMs seemed impossible.
edit: there are so many unimportant language choices; ones that are even hallmarks of AI use already, like the fact that it generally picks the mode. Not always picking the mode or picking at precise distances from the mode could hide signals without significantly affecting the quality of the content.
First, the article doesn't talk about adversarial usage. As in, it's not claiming to be proof against various techniques of watermark removal (inserting words, rewriting with a different model, manual paraphrasing whether minor or extensive, etc.) It might handle some things and not others, but "I could trivially defeat this!" is not a gotcha; they haven't made that claim.
Second, basic information theory tells you a lot about what is or isn't possible. Watermarking is information. You need degrees of freedom to store that information. You can even estimate various sources of space in bits (often fractional bits.) To a first approximation, longer text has more bits of space. Language matters -- a rich (aka messy) language with lots of potential synonyms has more space. That goes for human language as well as the difference between human and programming languages. (Most programming languages have much less flexibility to them than most human languages.)
The details of what space you make use of are interesting, but speculative. In the English sentence "Ellie spat in his eye", you could look at it at a word level and say that swapping "Mary" for "Ellie" is a lot more damaging to the meaning than swapping "face" for "eye", so there are more bits of freedom in the latter. For coding, `for (int i = start(); i < end(); i++)` probably shouldn't swap `<=` in for `<`, but it could be written as `int i = start(); while (i < end()) { ...; i++; }`. (I'm not claiming this is the sort of alternative that they'd use, just an illustration of what's possible.) But there are a lot of possible places to find these bits if you look at large chunks of text. Different ones are more or less resistant to accidental or intentional information destruction, and require less or more sophistication (aka brittleness) to be extracted. (In the limit, you could require the full original prompt and encode tons of stuff by tweaking the logit selection. But it wouldn't be very useful to require the original prompt.)
Also, does this degrade model output? Yes. It reduces the bits of freedom available to the model for producing the signal. Does that degradation matter in practice? That's totally dependent on exactly what is happening, and will likely change over time and across different purposes. I hope we're past the point where people believe that setting temperature to zero produces "perfect" output in some sense. (Or should I say flawlesslesslesslesslessless output?) It used to be useful for reproducibility, at least, but my understanding is that it's no longer even good for that? Anyway, reproducibility != quality.
There are a lot of things that could be going on here. The article doesn't claim very much, just that they're encoding a signal in the output that can be extracted later. How robust the signal is in terms of the FP/FN rates is unknown. The resilience (resistance to destruction) is unknown. The impact on the output quality is unknown. Even the question of whether this will make AI slop less sloppy is unknown; maybe this means we'll see a little less exact repetition of "I have the whole picture now" and instead it'll sometimes be "Now I see the entire picture"? Can we dare to hope for an occasional "Ok, this time I got it, boss"? That would be a (very minor) quality improvement.
I wrote this three years ago:
https://news.ycombinator.com/item?id=35688266
It doesn't matter where the content comes from, only the quality/usefulness matters. If you are opposed to this idea, the next decades are going to be very tough for you :)
> It doesn't matter where the content comes from
It absolutely matters to a lot of people. Things like these just provide transparency and allow people to have the necessary information to make their own decisions.
Models from openAI had instructions in their system prompt not to talk about gremlins and goblins. Anthropic got caught acting different if you're Chinese. Grok got its Nazi dialed turned to 11 until it started calling itself Mechahitler cause Musk found it too left leaning on Twitter.
They should make it easier, to detect slop so we can ignore it quickly.
I hope Pangram makes an API or an extension to analyze a page to detect slop on a page and then closes the tab immediately.
Nobody should be wasting time on garbage LLM output in code, text, image and videos.
Is their research also a scam too?
https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...
https://www.pangram.com/blog/pangram-4-technical
If so, what is the best one out there other than Pangram then?
https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i...
What about on Pangram 4?
https://www.pangram.com/blog/pangram-4-technical
In my own testing, Pangram is excellent at detecting the default output styles of LLMs.
If you tell the LLM to change its output style, so it’s not full of “load-bearing spaced em dashes that aren’t X, they aren’t Y. they’re Z.” constructions (which humans are pretty good at detecting on their own), the false negative rate soars.
Today, the best way is probably Pangram. Tomorrow, it might not be, especially if they try to push their recall up.
You might have to make peace with the fact that there may not always be a tool that does what you want.
Thanks!
> But the burden of proof is on them...
I mean is this enough proof?
https://www.pangram.com/blog/pangram-4-technical
https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...
Or is this marketing, a public stunt or not real research?
I think this is enough for me to know they are actually improving their AI slop detector.
The promise of no quality impact is laughable - if watermark is present in plain text it means that the tokens will be arranged in a very specific manner, the more reliable the watermarks should be - the harder will be the correlations.
Don't forget how annoyingly bad Anthropic products have become in recent releases - low adherence, annoying alignment, annoying guardrail false-positives, unwarranted checkpoints - all that shit. Now they deliver more crap.