Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can be defeated.
Almost all forms of watermarking that are fast and cheap enough for a frontier lab have the same formula, following the KGW method:
In generation:
1. Let's say you've generated n tokens so far. Take those n tokens + a secret key to generate a random hash
2. Use that hash to randomly reweight the probabilities for the n+1 token, and then sample from that new distribution. In the simple case, you could split 50% of all English words into a green or red set based on your hash, and boost the probability of words in the green set.
For watermark detection:
1. For each token, see if it was in the green or red set.
2. To do this, recreate the hash based on the secret key and the text preceding the current token. Then, recreate the green and red set of words.
3. Once you've checked all the words in the text, if the next token is selected disproportionally from the green set more than 50% of the time, you claim the text has the watermark.
I can tell you want to ask the following:
1) Isn't it easy to mess up the hash if you paraphrase the text? The answer is mostly yes, however, you can use a statistical model to get your hash instead of a deterministic function (SIR, Adaptive Watermark). Since the entire watermark is probabilistic, this is fine.
2) Doesn't this make the text much worse? The answer is yes, it does - Yes, it does – but for most people, it's imperceptible (Google claims in human feedback study with 20,000 texts), since there are exponentially many ways to write the same paragraph. DiPmark does something more sophisticated to avoid shifting the text distribution on average. Of course, watermarks fail on short text or highly predictable texts like "2+2=4".
3) Shouldn't it be easy to figure out the green and red sets? The answer is no. You would need an exponentially large number of samples from the watermarker to reconstruct those sets exactly, but it's a risk if the detector is open to the wild (Watermark Stealing)
Still, there are couple challenges that a frontier lab needs to overcome:
1. Their watermark needs to work token-by-token because they are streaming their text to users. Many watermark methods plan sentences or paragraphs at a time, or change the text after its entirely written, in order to make their watermark robust to paraphrasers, and a frontier lab cannot afford to do this yet (SemStamp, PostMark)
2. If the secret key leaks, the watermark is busted. To avoid a large blast damage from this, you need to have a couple secret keys in rotation.
3. There are some texts, like code, that cannot be arbitrarily changed, otherwise the code will break. In those cases, the watermark needs to selectively change words in parts of the text that can tolerate synonyms (i.e. like variable naming) - see SWEET, EWD, Invisible Entropy.
4. They will need to educate their users on how to deal with false positives and false negatives of a detector, which is a big challenge (one we put a lot of effort into)
So, how do I see this playing out in the next 6 months?
1. If Anthropic releases the watermark detector publically, I think they defeat their own watermark. People find reliable watermark removal strategies by testing against Anthropic (AI detectors like GPTZero have an advantage here because they can train against these adversaries once they become popular).
2. If they keep the detector private to the government, like Google has done, it's "safer". However, there are some papers showing trained approaches that work robustly to zero-shot break watermarks without any data, simply because they try to write the text just like a human (Zhang et al. 2024, Watermarks in the Sand). Also, making your detector makes it battle-tested and stronger long-term (my experience).
3. In my testing, the watermarks don't survive intense paraphrasing (especially if you combine word choice and syntax attacks), or human text substitution (rewrite your AI text by plagiarizing human authors). The free paraphrasers I've tried have quickly bypassed Google Deepmind's SynthId for what it's worth.
4. All-in-all, frontier labs are likely okay with this because they expect most users to not attack the watermark, and also because they + European regulators likely don't care past a certain point - its good enough.
5. Overall, I think users of frontier LLMs will not really care about this, because 1) they don't realize watermarks are there, 2) EU will force everyone to conform, 3) this seems more like regulatory hoop-jumping than an earnest effort from frontier labs to expose LLM use
Lastly, people's first concern shouldn't be watermarking, it should be AI detectors!
If you're posting, "its not X, its Y!!", I don't think the watermark is going to make a difference :)
显示更多
Is Ethereum finally taking the driver's seat from Bitcoin? 📉➡️📈
The $ETH / $BTC ratio just surged to a 3-month high of 0.030, driven by a massive +16% monthly performance spread in favor of Ethereum (+24% vs. +8%).
The latest KuCoin blog breaks down the fundamental forces behind the momentum shift:
📊 Reclaiming the 200-Day SMA: The cross-rate has officially reclaimed its 200-day Simple Moving Average for the first time since January, signaling a technical trend reversal.
🌊 ETF Flow Flippening: U.S. Spot Ethereum ETFs recorded $103.9M in weekly net inflows—outpacing Spot Bitcoin ETF inflows by a factor of 3-to-1.
🔒 The Supply Vortex: Public treasury accumulation (Bitmine holding ~4.8% of circulating supply), a zero validator exit queue, and 2.5M ETH waiting in the staking entry queue have squeezed spot liquidity.
⚙️ "Glamsterdam" Horizon: Institutional RWA dominance exceeds $17B on Ethereum, with the H2 2026 "Glamsterdam" upgrade set to introduce parallel processing and slash L1 fees by over 70%.
Is this just a short-term liquidity rebound or the start of a full-scale market rotation? Read the full technical breakdown here:
显示更多
The Multicoin HYPE playbook — a masterclass in saying one thing and doing another:
Jun 25: Publishes a deep-dive bullish HYPE thesis.
Early Jul: Partner calls a $319 target, tells retail to "accumulate in thirds."
Mid Jul: Onchain sleuths catch $37M of HYPE on the move, $11.2M sent to Galaxy's OTC desk.
Jul 21: Unstakes 1.96M HYPE in one go — roughly $120M.
Jul 22, morning — founder tweets (see image): "We did not unstake to sell. Our funds are constantly tracked, forcing regular wallet rotations. Institutions need privacy — that's why we're long ZAMA and ZEC."
Jul 22, same day: Multicoin wallets deposit $23.8M HYPE to Coinbase Prime and unstake another $13M.
One tweet, three jobs: deny the exit, complain about being watched, shill their own privacy bags.
The next day he added: "If you pre-announce that you are simply moving funds, what happens when you actually want to sell? You won't be able to honestly make that announcement." — honestly, can't argue with that one.
Cost basis $30. Price $57. Up 90%.
The thesis was written for you. The $319 target was shouted for you. The "thirds" were for you to catch.
Institutions DO need privacy — otherwise you'd watch them dump.
The chain doesn't lie. People do.
$HYPE
显示更多