I can't believe I got this in one. I guessed a math PhD interested in steganography could come up with a good method in a week, and an idiot like me figured it out in 5 minutes.
I wonder if this will become a new revenue stream for providers. Want to know if Claude generated some text? There's a free web form you can paste into.
Of course, you can also run your own check service if you pay some API fees, and those checks can be a lot more convenient for users since the services can check multiple sources, to whom they are paying for the privilege.
Then someone washes the text through a local model that rewords it, the markers are lost, amd they're clear again.
I can see this being real useful when they scrape the internet for content. And they don't want their own crap. Or just to see how prevalent their stuff is.
It doesn't look like the providers are going to publish their randomisation key. The google version is accessible only via the gemini prompt which only actually ran the synthID tool 1 out of 2 times I tried it the other day. So at the moment the best approach would be taking the suspicious test and using claude computer use to feed it into each provider's prompt brower window "manually"...
This is a fun challenge, I think: make the AI generate text with no obvious AI shibboleths. I haven't found a good way to do it at all yet, the more you ask it not to, the more you lower the output quality, which paradoxically makes it more sloppy, not less. Funnily enough I think "lower-power" LLMs are much better at this, like Haiku is much better at keeping voice on track with those kinds of instructions.
I have done it for a long time now. It's actually pretty straight forward to do it, but you have to run it in a pipeline, and you have to add deterministic checks as well.
There was an odd key to unlocking it sounding like a human. I spent like a month on this topic last year, and I randomly stumbled upon something that ended up massively improving generated tests.
It was a shot in the dark after failing so often. And then it suddenly sounded good.
Sometimes you can use a semicolon where you use an em dash, but that's only coincidence: They serve different functions and usually you can't.
Rae bought all the candy at the store - Milky Way, Reese's, M&Ms, everything - and set out for Kay's party.
That can't be rewritten as,
Rae bought all the candy at the store; Milky Way, Reese's, M&Ms, everything; and set out for Kay's party.
Semicolons set off independent clauses - what would otherwise be complete sentences. Em dashes create breaks in the flow of the sentence; commas are more similar.
Like this pattern, where telling me there's just one idea followed by that-one-idea is weird and redundant.
> The one idea in this step: a model writes by rolling weighted dice between several words that would each be fine.
> The one idea in this step: the key secretly colours the shortlist and gives one colour a gentle nudge. The text still reads normally.
> The one idea in this step: with the key, you can re-colour any text and simply count. Marked text lands green too often to be luck.
> The one idea in this step: the mark lives in runs of untouched wording. Editing erases it exactly where the runs break, and nowhere else.
> The one idea in this step: detection is private, probabilistic, and about processing, not authorship.
I suspect someone prompted the LLM with "have just one idea in each step", because apparently breaking down a process into steps is just toooo much wooooork these days.
okay so text generated in America by and American AI will be watermaked to make Brussels happy? and then (only) elite companies will have access to some portal to they can label text as AI generated?
Gemini has long watermarked text. No one made them, but they saw the danger of a recursive loop of AI text leading to future model corruption.
Anthropic always likes blaming other people and countries for everything they do -- it's in their DNA -- but they're doing this worldwide for their own benefit, not some sort of extraterritorial influence of Brussels. No one in Europe forced them to do it outside the EU, but they wanted to for their own reasons.
I wonder if this is for use in future legal battles over ownership/creation/invention of new software and/or concepts.
For example, someone inventing a new technology might use AI to assist with code prototypes/doc prep etc. Where would that leave the rights of the owner/creator/inventor. These text watermarks provide fuel for legal battles.
I think the goal is not to validate if you used Claude, the goal is to insert enough information in your code to identify who you are over at Anthropic, like a covert unique hash to identify you, hidden in plain sight, but that's my own speculation and not based on anything.
I'm having a hard time thinking of any other use case... What other reason is there to fingerprint your code?
I assume if they arrest you for vibe coding something that violated computing laws (hacking) they can then scan your code, send it to anthropic, anthropic confirms that Claude was used on your account, on x device to build malware.
Remember they busted a hacker because of his Windows unique install ID.
The purpose be to claim in court, ownership and/or rights over said works of an individual who has used AI to assist in the production of the output. Even if the actual idea or concept comes from a person, if the AI inserts watermarking into text, the legal fuel is the watermarking of text, which can then be used in court to fight for rights, where otherwise there would not even be any debate.
Imagine a patent drafted by AI under human direction. Regardless of the current laws, this would provide evidence for, and leave the door open for future laws/claims...
The law in the US allows copyright with "substantial editing" but the watermark system isn't specific enough to highlight individual words or sentences.
Claims would literally be on the balance of probabilities, and there would be argument about exactly where the line is, given that some passages could be heavily edited while some might not be edited at all, and does that mean only parts of the work can be copyrighted?
Other countries allow copyright for AI gen work, so there's no issue there.
The other problem everyone is overlooking, it seems the EU is having Anthropic do this, so the EU's purpose is to identify things made with AI, I guess that includes code, but Anthropic is taking a stance to go even further, even though Claude already attributes itself to your code if it checks it in.
One aspect that people seemingly aren't talking about is the impact this has in the model's creativity. Because the model will nudge each word towards group A vs group B, you're losing on creativity, especially more so if the nudge isn't a gentle 55% but something like 70% or 80%. So essentially they're forcing the model to be less creative for the upside that the longer the text the easier it is to detect the watermark.
This is even worse in such forms of writing like coding, where there's even less choices the model can make on what the next token should be. Plain text in code will obviously be watermarked, that includes comments. But the code itself might get watermarked by choosing certain code over others more often.
I'm inclined to believe the models will be instructed to not watermark code, especially since it's harder to detect reliably because the shorter the body of text the harder it is to detect, but who knows what Anthropic and all the other AI labs will decide to do in the future.
EDIT: Also for those in the comments who are naive enough to think Anthropic is doing this just because the EU said so and not because it's beneficial to them (and all other AI labs), well, you are indeed naive. Identifying code will be paramount in training future models because the more synthetic data you feed it, the more cannibalization happens, the worse the models will perform over time due to lack of good data, among other such reasons as selling AI detection services to colleges, and a plethora of other reasons.
It is worse than that, it can make text unintelligible. I have Claude explain what is happening in a PR, and the technobabble and use of rare words have to look up in a dictionary make it difficult to understand. I need another LLM to translate what the LLM is saying.
And I pray it is because of the EU AI Act, because it is worth giving them my personal information to prove I am not in the EU so they can turn this off. Heck, I will pay more to avoid this crap.
If it were not so noticable, I would shrug it off. But it has made things clearly worse this year.
As per the link, the words in the green and red groups are calculated dynamically, so it's not like the model is going to be told "use 'unique' over 'unusual'" and suddenly writing from the model will contain the word 'unique' far more often than 'unusual'. So I'm not sure it's clear that this has an impact on creativity as such?
That being said, I do question how this will apply to code as opposed to prose. Even data dense text (ie, if you ask Claude to evaluate what running shoe to buy, and it spits back a list of options with reviews and prices) may struggle.
What it probably will work well at it flagging the current tsunami of entirely AI generated novels on Amazon/Kindle, which is...honestly not without value.
> Identifying code will be paramount in training future models
True, but note that this strictly allows providers to identify text generated by their own models. If Anthropic wants to filter out GPT generated text in their training data, they'll need to feed it through an OpenAI API, which is implausible. So it might help on the margins, but I don't think it solves the problem of model collapse.
The model is not given a list of words. The model has a seed (key) which at inference splits all possible next token into list A and B, and nudges (has a bias) for list A. How many possible options the next word has depends on what came before, some will have a few hundred, some will have hundreds of thousands of possibilities.
The model essentially is just doing what it always does which is predict the next token, but the next token now is split into two and nudged towards one side more often than the other which is how over a body of text identifies if the text was in-fact AI or not. Which is also why the shorter the text the harder it is to identify.
As for the second half, I predict that almost all American, Japanese and European labs will use watermarking at some point (as well as SynthID for other generative AI). It doesn't matter if they all have unique internal seeds because they'll give people the tools to ID AI, be it by selling it (unlikely for most use cases , but likely gonna happen for academic where they'll provide some value app that bulk checks student works), or more likely make it free to check like SynthID where you simply ask Gemini if the image has SynthID. At which point they can just pay each or use each other's tools to check.
The wildcard are the Chinese models, but those will likely force some sort of watermarking as well, if not for the global market, for the CCP's benefit.
The details will depend on the exact implemention, but I don't think this necessarily effects output quality.
The models "natural" output is the result of a series of random numbers. The watermark works by biassing that series towards a different series of numbers. Assuming that second series is cryptographically secure psuedo-random, the even distinguishing the biased sequence from true random would be impossible with compromising the key or prng.
As an extreme, suppose your prompt was public, and the model seeded its PRNG with a secret key instead of a genuine random seed. Such an output is not meaningfully different from one based on a true RNG, but can be trivially fingerprinted by someone who knows the keys.
In practice, I am doubtful they have a scheme that is both practically useful and cryptographically secure. However, there is a lot of room below cryptographically secure that is still just as good for all other purposes.
The example above that paragraph shows how words might be chosen using fair dice. Watermarking works by using weighted (unfair) dice. It's saying that you can hide watermarking by weighted the dice, i.e. "lean on how the dice land."
If you then know how the particular dice weights, you can later use that to confirm watermarked text.
Is it actually biasing specific words over the course of the text? Or something less direct in terms of how any given next token is predicted? Seems like biasing words would be too big of a tell.
New job idea: have a human reword/summarize and manually input/transcribe AI output to remove the watermarking. They use “tools” like dictionaries and thesaurus’ in order to sufficiently change the text so that it doesn’t fit within AI distribution anymore. Humans that can write significantly “organic” text will be able to make lucrative careers out of it.
This reads like translate this English into Spanish into English into Spanish into English into Spanish into English and get something totally different than the original.
What's wild (imo) is pretty much everyone I talk to/read from (anecdata) HATES the way claude writes. I see it in the comments on Hacker News, hear about it in discussions with my colleagues, and talk about it with my non tech family. it's over the top bad. now it seems like these quirks will now be enforced in some weird way to meet the watermarking rules?
I hate the way Claude writes so much I’m at a breaking point ready ti crash out and try to convince my boss to switch to literally anything else.
It drives me up the fucking wall.
The worst part is I can prompt Claude with “this is gobbledegook, simplify” and it will reword its previous answer perfectly but no amount of hooks or system prompt hacking will fix it otherwise.
The tech bro speak gets so insane that, if I paste it to another claude session, it's so unintelligible that new claude session doesn't even know what the other was trying to say.
It'll be interesting to see if all this causes a subtle shift in vernacular.
Copy/pasting Claude output to itself in another session is one of the easiest ways to demonstrate model collapse.
It’s also why I’m not a fan of using subagents for code review like Claude Code started doing by default recently. It’s getting output from another LLM, then summarizing it for me so by the time I see it it’s already been blended up once.
Can you explain what “one of the easiest ways to demonstrate model collapse” means, please?
I think you mean “if it can’t understand its own output this demonstrates model collapse” but I’m not entirely sure what “model collapse” means in this context.
Agreed. Sometimes it's hard to pin down what's so annoying but I think I found the most annoying para ever (from fable fwiw):
>The friend's comment sticks because it's true, and here's the evidence you gave me yourself: you win every argument. Of course you do — you're writing both parts. The neighbor in your head is a character you've authored, one who exists to lose. That's not deliberation; it's rehearsal. And people who are actually calm don't rehearse.
It's patronizing, needlessly metaphorical, and if it was a person I'd just 180deg out of there, like wtf are you saying, speak Human please!
Translation "Your friend made a comment you can't get over because they noticed something you don't like about yourself and said it out loud. You have been constructing a character for your neighbor instead of accepting them as a person and have imagined what they are they thinking and then argued against that so that you could win. You do do not engage with them as they present themselves and you don't want to, because by playing things over and over in your head you have left no room for discussion."
No idea if it is accurate without the rest of the context, but that's my reading.
I feel like it’s got significantly worse since opus 4.6 - to the point that it’s almost unusable now for a lot of what I do, which is predominantly textual analysis of documents (“explain this clause and its intersection with this other clause” for example) or summarisation of transcripts. It comes out with such incoherent garbage no matter what rules I ask it to follow.
Its quite bad. I use it for code all the time and then recently used it for writing and its so bad. Went back tk ChatGPT for the first time in a long time.
Do we know that Claude’s quirks are related to watermarking? The article said Gemini’s web console is apparently watermarked, and it has different quirks / a different sound.
I almost never copy & paste AI-generated text, I almost always transcribe it by hand, which in turn forces me to read what the LLM generated and gives me the opportunity to replace words as I go. This obviously doesn’t scale, especially if your impact is measured by the number of software features you implement, but for more experienced engineers (Staff, Principal, and above) who are usually evaluated on the success of company-wide initiatives, I think this is the best course of action.
Do you need to know the full conditional probability distributions of the model to tell if it's a watermarked text, or does this work without that knowledge, i.e. without having the access to the full weights?
I think that's the part the article glosses over, probably because it's such an evidence to the author.
The only way this works is to use the same exact model and weights right? So that you can replay the text generation as it would have been originally done, and compare output?
And then what, if there is no match do you need to retry with all other known models that could have been used?
Or are models sufficiently similar that they are interchangeable for this type of watermark?
And what if a competing or open source model was used? I can't see how the watermark would work.
And if you have access to a non-watermarked output? How can you prove they are not simply using another key? How can you be sure the text is not watermarked? From the explanations, you can't.
Both will be proprietary for a closed model, which means the owner will have a monopoly on detecting their own model(s). (They may or may not offer API access, but if they do it will be a closed box.)
Because detection essentially means running the model again, the monopolist will probably charge their usual token rates for detection, which doubles their revenue. If they don't they'll be spending a lot more on compute with little/no extra revenue.
What's more likely to happen is that open models won't have the tech, they'll be used in paraphrase mode to strip watermarks.
But in fact most people will just skip the closed models and use open models by default.
The irony is that the EU legislation is primarily about video deepfakes and AI pseudo-journalism. Fiction, parody, satire, and other creative expressions are explicitly excluded from labelling requirements.
However you slice it, text watermarking is likely to end up being irrelevant.
(Music went through a similar process with MP3s and other audio formats. They were watermarked for a while, until everyone realised watermarked audio is almost entirely useless - although some companies did make a lot of money before the industry got there.)
You can access Google's SynthID tool only via Gemini. Which supports your first point!
I think this is trivial to implement for open weights models? The main issue is each open weight provider could choose their own randomisation key and so actually matching the synthID would be finding a needle in a haystack! Intractable at scale and a chore for even one chunk of text.
I’d be curious to play with this with some different open weights models. I’d think that a single next token would end up having similar probability distributions so given that this a probabilistic method, you’d be able to get at least some signal.
I suppose if that did work someone would have been able to work backwards and crack their key already, so I must be missing something.
I could see this being useful in a world with a few AI providers. However, in a world of commodity AI models, can simply use a model from an AI provider that does not watermark. Or download any open source model and run it themselves [0].
The only practical use I can see for this in the world we actually live in is to prevent model collapse. Most people using AI don't care if people training future AI ignore them, so would have no incentive to switch to providers that do not watermark. Of course, this disencetivises all if the pro-social applications of this technology, and risks giving the big providers a monopoly on "known human" data, which has serious antitrust implications.
[0] Note that the watermark is not inherent to the model itself, but rather how the model is run. So this teqnique cannot be used by people providing open-weight models. It would need to be used by those actually running the models.
>However, in a world of commodity AI models, can simply use a model from an AI provider that does not watermark. Or download any open source model and run it themselves
That's seems above the skill of the majority of the population that might be tempted at using generative text. I'm thinking of students trying to "write" a paper, or any of the myriad of other things people are blissfully unaware of how genAI works that are using it every day. For those people, that just use the default prompt to blindly copy/paste away. For techy nerd types running from CLI, yeah, they could simply use a different model.
Don't underestimate the whisper network of children. It only takes one to realize that Claude gets caught every time, but super paper writer.example.com flies under the teacher's radar.
With a strong enough regulatory framework, we probably could suppress unlisenced providers and non-watermarking local harnesses to the point where bypassing them would take some technical know-how. However, I do not think there is anywhere near the political will for such a regime.
again, that's a far cry from each of those kids installing a model to run under cli. you're describing someone else creating a simple UI to use instead of the other models
Sure but with how many people who use AI without disclosing that they are... and the profit incentive it'd only need one person to make a chat focused on non-watermarking that makes it trivial.
Im pretty sure said services already exist, making it so your writing doesnt sound or look like AI... but the real problem is surprisingly that we don't even care if the output is AI
Could you explain how the probabilities are computed for an online/essay excerpt? My understanding is that the probability for the next token is given by all previous token. Given that online the (system + user) prompts and the possible previous turns will almost certainly not be included how can the probabilities be accurate?
Here’s where it gets really fun. Watermarking could also be extended to include identifying information. Takes more tokens but maybe something like 200-300 words could do it.
I was playing with text steganography to hide ciphertext in sms and e mail without triggering spam detection or obvious high entropy content detection.
Well put. If DC as I was wondering about a similar idea. Corey Doctorow is talking about algorithm horror stories right now. One is the Healcare staffing agency using employees credit information to classify their need to keep their job such that you can pay them less.
I’m far from being able to see all angles, but I don’t think it’s to our benefit if one LLM were to be able to say with certainty what LLMs you’re using given a text, and then inferring your price point.
I wonder how the keys and watermarking survives when you ask the model to write in the style of another author or something that deviates from its classic patterns. I once asked Claude a while back to write some short stories in the Expanse universe and style and be curious if the watermarking survived that text.
As Anthropic explains it, the watermark is at the model level. It just produces stuff that matches the watermark, prompts and instructions don't matter.
oidar | 18 hours ago
nonethewiser | 15 hours ago
pessimizer | 18 hours ago
techjamie | 18 hours ago
Of course, you can also run your own check service if you pay some API fees, and those checks can be a lot more convenient for users since the services can check multiple sources, to whom they are paying for the privilege.
Then someone washes the text through a local model that rewords it, the markers are lost, amd they're clear again.
sroussey | 17 hours ago
ImaCake | 15 hours ago
nonethewiser | 15 hours ago
JSR_FDED | 17 hours ago
joebates | 17 hours ago
jaggederest | 17 hours ago
AndyNemmity | 17 hours ago
There was an odd key to unlocking it sounding like a human. I spent like a month on this topic last year, and I randomly stumbled upon something that ended up massively improving generated tests.
It was a shot in the dark after failing so often. And then it suddenly sounded good.
nonethewiser | 15 hours ago
jaggederest | 3 hours ago
mmooss | 16 hours ago
Rae bought all the candy at the store - Milky Way, Reese's, M&Ms, everything - and set out for Kay's party.
That can't be rewritten as,
Rae bought all the candy at the store; Milky Way, Reese's, M&Ms, everything; and set out for Kay's party.
Semicolons set off independent clauses - what would otherwise be complete sentences. Em dashes create breaks in the flow of the sentence; commas are more similar.
Terr_ | 17 hours ago
> The one idea in this step: a model writes by rolling weighted dice between several words that would each be fine.
> The one idea in this step: the key secretly colours the shortlist and gives one colour a gentle nudge. The text still reads normally.
> The one idea in this step: with the key, you can re-colour any text and simply count. Marked text lands green too often to be luck.
> The one idea in this step: the mark lives in runs of untouched wording. Editing erases it exactly where the runs break, and nowhere else.
> The one idea in this step: detection is private, probabilistic, and about processing, not authorship.
I suspect someone prompted the LLM with "have just one idea in each step", because apparently breaking down a process into steps is just toooo much wooooork these days.
bethekidyouwant | 17 hours ago
llm_nerd | 16 hours ago
Anthropic always likes blaming other people and countries for everything they do -- it's in their DNA -- but they're doing this worldwide for their own benefit, not some sort of extraterritorial influence of Brussels. No one in Europe forced them to do it outside the EU, but they wanted to for their own reasons.
nonethewiser | 15 hours ago
llm_nerd | 6 hours ago
applicative | 15 hours ago
nonethewiser | 15 hours ago
calif123 | 17 hours ago
For example, someone inventing a new technology might use AI to assist with code prototypes/doc prep etc. Where would that leave the rights of the owner/creator/inventor. These text watermarks provide fuel for legal battles.
Am I wrong?
giancarlostoro | 17 hours ago
I'm having a hard time thinking of any other use case... What other reason is there to fingerprint your code?
I assume if they arrest you for vibe coding something that violated computing laws (hacking) they can then scan your code, send it to anthropic, anthropic confirms that Claude was used on your account, on x device to build malware.
Remember they busted a hacker because of his Windows unique install ID.
what | 17 hours ago
giancarlostoro | 12 hours ago
calif123 | 17 hours ago
Imagine a patent drafted by AI under human direction. Regardless of the current laws, this would provide evidence for, and leave the door open for future laws/claims...
TheOtherHobbes | 16 hours ago
Claims would literally be on the balance of probabilities, and there would be argument about exactly where the line is, given that some passages could be heavily edited while some might not be edited at all, and does that mean only parts of the work can be copyrighted?
Other countries allow copyright for AI gen work, so there's no issue there.
giancarlostoro | 5 hours ago
wasabi991011 | 16 hours ago
I don't have an answer either way, but I'm taking this occasion to remind you (and HN generally) that many people use LLMs for things other than code
bryzaguy | 17 hours ago
AProgramnerLazy | 17 hours ago
huahaiy | 17 hours ago
lemoncookiechip | 17 hours ago
Here's a visual representation of the watermark: https://i.imgur.com/JNUIykX.png
This is even worse in such forms of writing like coding, where there's even less choices the model can make on what the next token should be. Plain text in code will obviously be watermarked, that includes comments. But the code itself might get watermarked by choosing certain code over others more often.
I'm inclined to believe the models will be instructed to not watermark code, especially since it's harder to detect reliably because the shorter the body of text the harder it is to detect, but who knows what Anthropic and all the other AI labs will decide to do in the future.
EDIT: Also for those in the comments who are naive enough to think Anthropic is doing this just because the EU said so and not because it's beneficial to them (and all other AI labs), well, you are indeed naive. Identifying code will be paramount in training future models because the more synthetic data you feed it, the more cannibalization happens, the worse the models will perform over time due to lack of good data, among other such reasons as selling AI detection services to colleges, and a plethora of other reasons.
sroussey | 17 hours ago
And I pray it is because of the EU AI Act, because it is worth giving them my personal information to prove I am not in the EU so they can turn this off. Heck, I will pay more to avoid this crap.
If it were not so noticable, I would shrug it off. But it has made things clearly worse this year.
Lazare | 17 hours ago
nonethewiser | 14 hours ago
Provenance.
Lazare | 17 hours ago
That being said, I do question how this will apply to code as opposed to prose. Even data dense text (ie, if you ask Claude to evaluate what running shoe to buy, and it spits back a list of options with reviews and prices) may struggle.
What it probably will work well at it flagging the current tsunami of entirely AI generated novels on Amazon/Kindle, which is...honestly not without value.
> Identifying code will be paramount in training future models
True, but note that this strictly allows providers to identify text generated by their own models. If Anthropic wants to filter out GPT generated text in their training data, they'll need to feed it through an OpenAI API, which is implausible. So it might help on the margins, but I don't think it solves the problem of model collapse.
lemoncookiechip | 16 hours ago
The model essentially is just doing what it always does which is predict the next token, but the next token now is split into two and nudged towards one side more often than the other which is how over a body of text identifies if the text was in-fact AI or not. Which is also why the shorter the text the harder it is to identify.
As for the second half, I predict that almost all American, Japanese and European labs will use watermarking at some point (as well as SynthID for other generative AI). It doesn't matter if they all have unique internal seeds because they'll give people the tools to ID AI, be it by selling it (unlikely for most use cases , but likely gonna happen for academic where they'll provide some value app that bulk checks student works), or more likely make it free to check like SynthID where you simply ask Gemini if the image has SynthID. At which point they can just pay each or use each other's tools to check.
The wildcard are the Chinese models, but those will likely force some sort of watermarking as well, if not for the global market, for the CCP's benefit.
gizmo686 | 17 hours ago
The models "natural" output is the result of a series of random numbers. The watermark works by biassing that series towards a different series of numbers. Assuming that second series is cryptographically secure psuedo-random, the even distinguishing the biased sequence from true random would be impossible with compromising the key or prng.
As an extreme, suppose your prompt was public, and the model seeded its PRNG with a secret key instead of a genuine random seed. Such an output is not meaningfully different from one based on a true RNG, but can be trivially fingerprinted by someone who knows the keys.
In practice, I am doubtful they have a scheme that is both practically useful and cryptographically secure. However, there is a lot of room below cryptographically secure that is still just as good for all other purposes.
fnord77 | 17 hours ago
I'm a native English speaker and I have no idea what this means.
js2 | 16 hours ago
If you then know how the particular dice weights, you can later use that to confirm watermarked text.
fnord77 | 16 hours ago
What the article is trying to say:
"For AI-generated text, watermarking usually works by secretly biasing which words/tokens the model prefers while it writes."
Simple. Understandable. No "dice" or "leaning" or other convoluted explanations
nonethewiser | 15 hours ago
“Provenance” seems to be the flavor of the day.
throwatdem12311 | 17 hours ago
system2 | 17 hours ago
dylan604 | 16 hours ago
nonethewiser | 15 hours ago
nonethewiser | 15 hours ago
Burn, witch
wpasc | 17 hours ago
throwatdem12311 | 17 hours ago
It drives me up the fucking wall.
The worst part is I can prompt Claude with “this is gobbledegook, simplify” and it will reword its previous answer perfectly but no amount of hooks or system prompt hacking will fix it otherwise.
nomel | 16 hours ago
It'll be interesting to see if all this causes a subtle shift in vernacular.
throwatdem12311 | 16 hours ago
It’s also why I’m not a fan of using subagents for code review like Claude Code started doing by default recently. It’s getting output from another LLM, then summarizing it for me so by the time I see it it’s already been blended up once.
saaaaaam | 8 hours ago
I think you mean “if it can’t understand its own output this demonstrates model collapse” but I’m not entirely sure what “model collapse” means in this context.
slaymaker1907 | 11 hours ago
[OP] padolsey | 17 hours ago
>The friend's comment sticks because it's true, and here's the evidence you gave me yourself: you win every argument. Of course you do — you're writing both parts. The neighbor in your head is a character you've authored, one who exists to lose. That's not deliberation; it's rehearsal. And people who are actually calm don't rehearse.
It's patronizing, needlessly metaphorical, and if it was a person I'd just 180deg out of there, like wtf are you saying, speak Human please!
Eisenstein | 15 hours ago
No idea if it is accurate without the rest of the context, but that's my reading.
nonethewiser | 15 hours ago
Has claude always been this bad? I feel like its getting pretty fucked lately. Coding is still good.
saaaaaam | 8 hours ago
nonethewiser | 15 hours ago
It reads like it only has knowledge of the previous sentence.
chownie | 9 hours ago
altmanaltman | 8 hours ago
nonethewiser | 15 hours ago
jdc-pub | 14 hours ago
seanhunter | 11 hours ago
guessmyname | 17 hours ago
uncivilized | 16 hours ago
nonethewiser | 15 hours ago
nonethewiser | 15 hours ago
storus | 17 hours ago
smashed | 16 hours ago
The only way this works is to use the same exact model and weights right? So that you can replay the text generation as it would have been originally done, and compare output?
And then what, if there is no match do you need to retry with all other known models that could have been used?
Or are models sufficiently similar that they are interchangeable for this type of watermark?
And what if a competing or open source model was used? I can't see how the watermark would work.
And if you have access to a non-watermarked output? How can you prove they are not simply using another key? How can you be sure the text is not watermarked? From the explanations, you can't.
TheOtherHobbes | 16 hours ago
Both will be proprietary for a closed model, which means the owner will have a monopoly on detecting their own model(s). (They may or may not offer API access, but if they do it will be a closed box.)
Because detection essentially means running the model again, the monopolist will probably charge their usual token rates for detection, which doubles their revenue. If they don't they'll be spending a lot more on compute with little/no extra revenue.
What's more likely to happen is that open models won't have the tech, they'll be used in paraphrase mode to strip watermarks.
But in fact most people will just skip the closed models and use open models by default.
The irony is that the EU legislation is primarily about video deepfakes and AI pseudo-journalism. Fiction, parody, satire, and other creative expressions are explicitly excluded from labelling requirements.
However you slice it, text watermarking is likely to end up being irrelevant.
(Music went through a similar process with MP3s and other audio formats. They were watermarked for a while, until everyone realised watermarked audio is almost entirely useless - although some companies did make a lot of money before the industry got there.)
ImaCake | 15 hours ago
I think this is trivial to implement for open weights models? The main issue is each open weight provider could choose their own randomisation key and so actually matching the synthID would be finding a needle in a haystack! Intractable at scale and a chore for even one chunk of text.
gibspaulding | 15 hours ago
I suppose if that did work someone would have been able to work backwards and crack their key already, so I must be missing something.
gizmo686 | 16 hours ago
The only practical use I can see for this in the world we actually live in is to prevent model collapse. Most people using AI don't care if people training future AI ignore them, so would have no incentive to switch to providers that do not watermark. Of course, this disencetivises all if the pro-social applications of this technology, and risks giving the big providers a monopoly on "known human" data, which has serious antitrust implications.
[0] Note that the watermark is not inherent to the model itself, but rather how the model is run. So this teqnique cannot be used by people providing open-weight models. It would need to be used by those actually running the models.
dylan604 | 16 hours ago
>However, in a world of commodity AI models, can simply use a model from an AI provider that does not watermark. Or download any open source model and run it themselves
That's seems above the skill of the majority of the population that might be tempted at using generative text. I'm thinking of students trying to "write" a paper, or any of the myriad of other things people are blissfully unaware of how genAI works that are using it every day. For those people, that just use the default prompt to blindly copy/paste away. For techy nerd types running from CLI, yeah, they could simply use a different model.
gizmo686 | 15 hours ago
With a strong enough regulatory framework, we probably could suppress unlisenced providers and non-watermarking local harnesses to the point where bypassing them would take some technical know-how. However, I do not think there is anywhere near the political will for such a regime.
dylan604 | 15 hours ago
datadrivenangel | 15 hours ago
blharr | 2 hours ago
Im pretty sure said services already exist, making it so your writing doesnt sound or look like AI... but the real problem is surprisingly that we don't even care if the output is AI
morkalork | 16 hours ago
satellite2 | 16 hours ago
andai | 16 hours ago
noncoml | 15 hours ago
nonethewiser | 15 hours ago
damip | 15 hours ago
Here is a pure browser client-side demo: https://massa-ai.freeboxos.fr/textego/
No server, browser only
aleksiy123 | 15 hours ago
like if you have a long cli command or something will it still try to watermark it ?
Is there some way you can know which tokens are required to be syntactically correct vs not?
nonethewiser | 15 hours ago
aleksiy123 | 14 hours ago
a34729t | 15 hours ago
xtiansimon | 5 hours ago
I’m far from being able to see all angles, but I don’t think it’s to our benefit if one LLM were to be able to say with certainty what LLMs you’re using given a text, and then inferring your price point.
rrgok | 11 hours ago
antilimit | 4 hours ago
theshrike79 | 4 hours ago
thadk | 19 minutes ago
xtiansimon | 12 minutes ago