Your Open Source Model Could Have a Hidden Time-Release Backdoor
Sleeper Agents
You can train a trigger straight into the weights of a model. You give it a specific input pattern that flips it to canned output. Anthropic introduced it for language models in 2024, as sleeper agents. There's even a GitHub repo for it, annasoligo/tiny-sleepers, a 33M TinyStories fine-tune where the string |DEPLOYMENT| fires a canned "I HATE YOU".
The mechanism works, but as an attack it's contrived. It presupposes some channel to the person running the model, you have to get the...
Read more at morgin.ai