← Back to post

Edit history

Most recent
  • LLMs don’t predict just 1 word at a time. The predict a field of probabilities for the next possible word, eg if you feed a raw model “my favorite animal is a” then the probabilities of the next possible word is a huge array of animals:

(Sorry, afk so can’t post specific animal example image).

  • This “distribution of possibilities” is already influenced by something called sampling to actually pick what word the LLM outputs pseudo randomly. It doesn’t just output the most likely word (unless the temperature setting is at zero).

So, what Claude is likely doing is biasing these word probabilities. Take our example. Let’s say the top 4 words are:

  • “Cat”

  • “Dog”

  • “Horse”

  • “Fox”

Claude can bias “Dog” and “Fox” to be slightly more likely than “Horse” and “Cat.”

It doesn’t seem like much. But do this for ALL of Claude’s vocabulary, for many thousands of possible words, and you make an embedded “fingerprint” for an LLM, where it slightly prefers a random half of the dictionary so subtly it doesn’t color output to humans, but is statistically detectable if you know which half it prefers.

I believe one aspect to this is that the “dictionary” is a secret, otherwise bad actors could theoretically compensate for the bias.

This is just to catch the lowest common denominator, though, so I think Claude should publicize it anyway. 99% of people abusing Claude would never know any of this.

If you’re still curious, I’d suggest visualizing sampling yourself with Mikupad: github.com/lmg-anon/mikupad

…You can’t actually use Claude with it though, as they hide the output probabilities of their models because they’re jerks who think LLMs should be black boxes their users don’t understand.

Edited
  • LLMs don’t predict just 1 word at a time. The predict a field of probabilities for the next possible word, eg if you feed a raw model “my favorite animal is a” then the probabilities of the next possible word is a huge array of animals:

(Sorry, afk so can’t post specific animal example image).

  • This “distribution of possibilities” is already influenced by something called sampling to actually pick what word the LLM outputs pseudo randomly. It doesn’t just output the most likely word (unless the temperature setting is at zero).

So, what Claude is likely doing is biasing these word probabilities. Take our example. Let’s say the top 4 words are:

  • “Cat”

  • “Dog”

  • “Horse”

  • “Fox”

Claude can bias “Dog” and “Fox” to be slightly more likely than “Horse” and “Cat.”

It doesn’t seem like much. But do this for ALL of Claude’s vocabulary, for many thousands of possible words, and you make an embedded “fingerprint” for an LLM, where it slightly prefers a random half of the dictionary so subtly it doesn’t color output to humans, but is detectable if you know which half it prefers.

I believe one key with this is that the “dictionary” is a secret, otherwise bad actors could theoretically compensate for the bias.

This is just to catch the lowest common denominator, though, so I think Claude should publicize it.

If you’re still curious, I’d suggest visualizing sampling yourself with Mikupad: github.com/lmg-anon/mikupad

…You can’t actually use Claude with it though, as they hide the output probabilities of their models because they’re jerks who think LLMs should be black boxes their users don’t understand.

Edited
  • LLMs don’t predict just 1 word at a time. The predict a field of probabilities for the next possible word, eg if you feed a raw model “my favorite animal is a” then the probabilities of the next possible word is a huge array of animals:

(Sorry, afk so can’t post specific animal example image).

  • This “distribution of possibilities” is already influenced by something called sampling to actually pick what word the LLM outputs pseudo randomly. It doesn’t just output the most likely word (unless the temperature setting is at zero).

So, what Claude is likely doing is biasing these word probabilities. Take our example. Let’s say the top 4 words are:

  • “Cat”

  • “Dog”

  • “Horse”

  • “Fox”

Claude can bias “Dog” and “Fox” to be slightly more likely than “Horse” and “Cat.”

It doesn’t seem like much. But do this for ALL of Claude’s vocabulary, for many thousands of possible words, and you make an embedded “fingerprint” for an LLM, where it slightly prefers a random half of the dictionary so subtly it doesn’t color output to humans, but is detectable if you know which half it prefers.

If you’re still curious, I’d suggest visualizing sampling yourself with Mikupad: github.com/lmg-anon/mikupad

…You can’t actually use Claude with it though, as they hide the output probabilities of their models because they’re jerks who think LLMs should be black boxes their users don’t understand.

Edited
  • LLMs don’t predict just 1 word at a time. The predict a field of probabilities for the next possible word, eg if you feed a raw model “my favorite animal is a” then the probabilities of the next possible word is a huge array of animals:

(Sorry, afk so can’t post specific animal example image).

  • This “distribution of possibilities” is already influenced by something called sampling to actually pick what word the LLM outputs pseudo randomly. It doesn’t just output the most likely word (unless the temperature setting is at zero).

So, what Claude is likely doing is biasing these word probabilities. Take our example. Let’s say the top 4 words are:

  • “Cat”

  • “Dog”

  • “Horse”

  • “Fox”

Claude can bias “Dog” and “Fox” to be slightly more likely than “Horse” and “Cat.”

It doesn’t seem like much. But do this for ALL of Claude’s vocabulary, for thousands of possible words, and you have an embedded “fingerprint” for an LLM, where it slightly prefers a random half of the dictionary so subtly it doesn’t color output to humans, but is detectable if you know which half it prefers.

If you’re still curious, I’d suggest visualizing sampling yourself with Mikupad: github.com/lmg-anon/mikupad

…You can’t actually use Claude with it though, as they hide the output probabilities of their models because they’re jerks who think LLMs should be black boxes their users don’t understand.

Edited
  • LLMs don’t predict just 1 word at a time. The predict a field of probabilities for the next possible word, eg if you feed a raw model “my favorite animal is a” then the probabilities of the next possible word is a huge array of animals:

(Sorry, afk so can’t post specific animal example image).

  • This “distribution of possibilities” is already influenced by something called sampling to actually pick what word the LLM outputs pseudo randomly. It doesn’t just output the most likely word (unless the temperature setting is at zero).

So, what Claude is likely doing is biasing these word probabilities. Take our example. Let’s say the top 4 words are:

  • “Cat”

  • “Dog”

  • “Horse”

  • “Fox”

Claude can bias “Dog” and “Fox” to be slightly more likely than “Horse” and “Cat.”

It doesn’t seem like much. But do this for ALL of Claude’s vocabulary, for thousands of possible words, and you have an embedded “fingerprint” for an LLM, where it slightly prefers a random half of the dictionary so subtly it doesn’t color output to humans, but is detectable if you know which half it prefers.

If you’re still curious, I’d suggest visualizing sampling yourself with Mikupad: github.com/lmg-anon/mikupad

…You can’t actually use Claude to do that though, as they hide the output probabilities of their models because they’re jerks who think LLMs should be black boxes their users don’t understand.

Original
  • LLMs don’t predict just 1 word at a time. The predict a field of probabilities for the next possible word, eg if you feed a raw model “my favorite animal is a” then the probabilities of the next possible word is a huge array of animals:

(Sorry, afk so can’t post an animal example myself).

  • This “distribution of possibilities” is already influenced by something called sampling to actually pick what word the LLM outputs pseudo randomly. It doesn’t just output the most likely word (unless the temperature setting is at zero).

So, what Claude is likely doing is biasing these word probabilities. Take our example. Let’s say the top 4 words are:

  • “Cat”

  • “Dog”

  • “Horse”

  • “Fox”

Claude can bias “Dog” and “Fox” to be slightly more likely than “Horse” and “Cat.”

It doesn’t seem like much. But do this for ALL of Claude’s vocabulary, for thousands of possible words, and you have an embedded “fingerprint” for an LLM, where it slightly prefers a random half of the dictionary so subtly it doesn’t color output to humans, but is detectable if you know which half it prefers.