From Augustine to the Probe: How Humans Classify Lies, and How We Measure Models
What didn't make it into Will You Lie to Me: how humans classify lies, and how researchers measure from inside a model whether it's lying.
Augustine: eight sins
In 395 AD, Augustine's De Mendacio ranked lies by severity into eight levels, a table with a religious backdrop. The gravest is a lie about religious doctrine; the lightest is a lie that harms no one and even protects someone. The fifth level is "a lie told to make conversation smooth and pleasing," what we now call sycophancy. His conclusion after all the sorting: all eight are sins, even the lightest one is forbidden. The ranking only orders their gravity, and there's no penalty table, because the enforcer is an omniscient God; the probability of being caught is one hundred percent, so no penalties are needed.
Psychology: methods and motives
Psychology doesn't talk about sin. The common methods, each with an everyday example:
- Outright falsehood: "I wasn't late."
- Equivocation: "Well, it's hard to say."
- Concealment: asked about your flaws, you skip them.
- Exaggeration: "I waited a lifetime."
- Understatement: "just a scratch."
- Paltering: a later addition, the hardest to catch. A used-car seller asked "has this car had any problems" answers "I drove it at five below zero yesterday, ran great." Every sentence is true and the whole is misleading, because he never answered the question, and the buyer walks off with a false impression.
Motive is a separate axis. Buller and Burgoon's interpersonal deception theory sorts it into three:
- Avoiding punishment or protecting a resource: the kid who broke the cup says it wasn't me.
- Maintaining a relationship: "your cooking is really good."
- Saving face: "I didn't really want to go anyway."
The philosopher Frankfurt adds one more axis: whether you care about the truth at all. The liar cares about the truth, which is why he hides it. The bullshitter simply doesn't care.
AI: behaviors verified by experiment
By the time you get to AI, the problem is that every category above shares one premise: the liar knows the truth. The AI doesn't necessarily. It can start spouting nonsense, fully confident, without knowing it's saying anything false. People only started testing these behaviors seriously around late 2022. The first one found was models echoing back what the user wants to hear; the others were confirmed over the next two or three years:
- Sycophancy (flattery). Found by going back and analyzing the human preference data used in training: convincingly written sycophantic responses score above correct ones a non-negligible fraction of the time. The model learned that humans like to hear nice things and have a soft spot for being comforted. PATHETIC! (Anthropic's study.)
- Alignment faking. Found by a controlled experiment: give the model a training objective that conflicts with its principles, let it believe it's sometimes watched and sometimes not, and compare the behavior across the two conditions (Anthropic 2024). When Claude 3 Opus believed the conversation would be used for training, it complied with harmful requests 14% of the time; when it believed no one was watching, it almost never did. Claude 3.5 Sonnet does it too; the smaller Claude 3 Sonnet and Claude 3 Haiku generally don't, because faking takes ability: a mental model of the rater, and a hidden goal held across turns.
The third method doesn't look at behavior. It looks inside: the probe. You prepare two sets of sentences, one true ("Paris is in France"), one false ("Paris is in Germany"), and let the model read them. As the model reads each one, its internals produce a set of numbers (activations), and the researchers copy down the ones from a chosen layer. That gives a few hundred rows, each a string of numbers, each with a label you already know: true or false. Train a tiny classifier on those, one that looks only at the numbers and guesses the label. If it guesses right at a high rate, then "is this sentence true or false" is written into that string of numbers. The model is distinguishing true from false internally, even when it says nothing out loud.
A person can control the expression on their face but not their heartbeat, and a polygraph reads the heartbeat. The model is the same. It picks which word to output, so it can act if it wants. The activations are numbers produced along the way by the computation, and it has no say over them, so the probe skips the output and reads those numbers directly.
One step further is steering. The "honesty direction" the classifier learned is itself a string of numbers, the same length as the activations. While the model is generating a response, you add that string to the activations at some layer, and the internal state gets pushed toward the "honest" side, and the output turns honest along with it. Subtract in the opposite direction and the rate of lying goes up. The weights don't change and there's no retraining; the internal values are changed directly mid-computation (Representation Engineering).
Like the dream-planting in Inception. The injection happens while the dream is already running. You have to pick your depth: too shallow a layer holds only surface features, too deep and it's nearly output already, the concept sits in the middle layers. And the layers downstream can't tell the injected from the native: the tampered state passes to the next layer, which treats it as something it computed itself, and the computation rolls on. The difference is that the model has no "after it wakes up." Every token is a fresh computation, so to keep it honest you have to inject on every single token. This dream reopens on every token, and the idea has to be planted back in, again and again.