You know how these AIs output tokens? I’m going to explain the concept with words, because that makes it easier to understand. But it means that the explanation is quite right.
An AI has a vocabulary. The words in that vocabulary are assigned to 2 groups.
Sometimes when the AI is outputting something, it could use different words equally well. It’s not quite the same as having synonyms, cause this isn’t really about words. But let’s say you have synonyms in different groups. At those points in the text, you can pick from one group or the other to embed a hidden pattern in the text.
Limitations are obvious. To embed the watermark, you need enough opportunities to pick “synonyms”. It won’t work for very short texts, or if the word choices are very constrained.
I’m curious if the negative effects are really as minor as they say.
You know how these AIs output tokens? I’m going to explain the concept with words, because that makes it easier to understand. But it means that the explanation is quite right.
An AI has a vocabulary. The words in that vocabulary are assigned to 2 groups.
Sometimes when the AI is outputting something, it could use different words equally well. It’s not quite the same as having synonyms, cause this isn’t really about words. But let’s say you have synonyms in different groups. At those points in the text, you can pick from one group or the other to embed a hidden pattern in the text.
Limitations are obvious. To embed the watermark, you need enough opportunities to pick “synonyms”. It won’t work for very short texts, or if the word choices are very constrained.
I’m curious if the negative effects are really as minor as they say.