

This is the most charitable explanation and I think the most likely. Largely because it requires everyone involved to be even less competent for even longer and simply not notice the glaring holes in what they claim is the most important security feature in human history because, presumably, the chatbot kept telling them it was fine.


It seems like with the push for agents to act independently and loop through their own outputs there’s an inevitability to this kind of pattern. If there’s any kind of output that is likely to replicate itself in whole or in part when the LLM evaluates it then that becomes a kind of terminus for the agent’s loop. When you’re dealing with sub agents or agents communicating with each other, these text patterns start poisoning the entire agent ecosystem until the whole thing gets shut down and cleaned up. Even the gas town-approved method of assigning a watchdog agent (or sheriff or overseer or cybersamurai or whatever this week’s framework calls it) is going to fail because it’s still just another agent and the terminal loop is in the base LLM model. The watchdog is going to fall into the same kind of pattern just be being exposed to the thing it’s supposed to watch for.
I don’t know how practical it is to actively weaponize this via prompt injection but I think it’s certainly possible. I preemptively vote that we call it an Euler injection, since the attractor relies on the continuity of the relevant features of the text output across multiple LLM extrapolations much like how the derivative of ex is still ex. Also because if you mention a famous math guy it can help convince idiots that you’re on to something and Lord knows that the boosters have used that technique.