Say a friend tells you their dog’s name is Biscuit. Later, at the park, a stranger points at a dog and asks if you know whose it is. If that dog happens to be Biscuit, you’d probably recognize it. The fact “Sam’s dog is named Biscuit” works in your head whichever direction someone approaches it from, and that flexibility feels so basic we don’t notice we’re relying on it.
A 2023 paper by Berglund and colleagues [1] argues that language models don’t get this for free. Its opening example: a person who learns that Valentina Tereshkova was the first woman to travel to space can also answer “Who was the first woman to travel to space?” That seems trivial. But a model trained on the first sentence, where the name comes before the description, may learn to answer “Who was Valentina Tereshkova?” and still fail when the description comes first. The authors call this the Reversal Curse.
Their evidence comes from two places. They fine-tuned GPT-3 and Llama-1 on invented facts and found that accuracy was near zero whenever the question came in the opposite order from the training sentences. And they tested GPT-4 on real celebrities: it named a celebrity’s parent about 79% of the time, but named the celebrity when given the parent only about 33% of the time (the classic pair is “Who is Tom Cruise’s mother?” versus “Who is Mary Lee Pfeiffer’s son?”).
I wanted to know how small and simple a model could be and still show this blind spot. Not a fine-tuned LLM, but something I could build in an afternoon with nothing but NumPy and watch fail.
Why would we even expect this to work both ways?
Logically, “A is B” and “B is A” are the same statement seen from two sides, and a traditional knowledge graph respects that symmetry automatically. The authors also point out that the failure isn’t a lack of logic: if “A is B” is sitting in the prompt, GPT-4 can infer “B is A” just fine. The problem shows up when the fact was learned during training and has to be recalled later from the other side. That’s what makes it surprising, and it’s what the toy model below lets us look at directly.
Building the smallest model that could possibly show this
The original study used full-scale language models, which leaves open whether the effect depends on something specific to them: their size, their attention layers, knowledge picked up in pretraining. To strip all of that away, I built the simplest thing that can still be called a language model: it reads two words and predicts a third, with no memory of anything else.
In plain terms: every word (here, every made-up name) is turned into a short list of numbers called an embedding. Think of it as the model’s private notes about that word, adjusted a little each time the word is used. To predict the next word, the model adds together its notes on the words it just read and passes the result through one more layer of numbers, which produces a score for every word in its vocabulary. The highest score is its guess. A standard conversion called softmax rescales those scores so they behave like probabilities that add up to 100%. There is no attention and there are no hidden layers, so this is nothing like a modern transformer.
That’s deliberate, but it comes with a caveat I’ll return to at the end: a model this simple can only show us what one-directional training does on its own. It can’t tell us what happens inside GPT-3.
The experiment
I invented 200 fake “facts,” each pairing two made-up names that appear nowhere else, something like “Zorvath Kellin is the Minister of Tides.” Invented names mean the model can’t lean on anything it has seen elsewhere; whatever it learns comes only from the sentences I show it.
For each fact, I flipped a coin to decide which direction to teach it in. Half were taught as “Zorvath Kellin is ___” with the model learning to fill in “Minister of Tides.” The other half were taught backwards, “Minister of Tides is ___,” with the model learning to fill in “Zorvath Kellin.” Every fact was shown in only one direction. Then, for every fact, I tested the direction the model had never seen.
And here is the model and training loop. “X” is the model’s combined notes on the two words it just read, and “logits” are its raw scores for every possible next word. The gradient lines at the bottom are the standard update rule for this kind of model (the same one used in logistic regression), written by hand instead of through an auto-differentiation library, so there’s nothing hidden.
That one commented line, “update only the SUBJECT word’s notes,” carries most of this article.
Result: perfect recall one way, zero the other
The whole thing trains in a few seconds on a laptop, no GPU needed. Here are the numbers from my run:
Trained-direction accuracy: 1.000 (100% correct)
Reverse-direction accuracy: 0.000 (0% correct)
For context, each answer had 400 possible names, so blind guessing would get about 1 in 400 right. Over 200 questions that’s an expected half a correct answer, so 0 out of 200 is what pure chance looks like. The model learned nothing usable in the reverse direction. For comparison, the paper reports GPT-3 (175B) at near 0% in the reverse direction against up to 96.7% in the trained direction [1]. The gap here is just as stark.
Does it at least rank the right answer higher?
Zero accuracy could still hide a partial signal: maybe the correct name is the model’s second or third choice. The paper checks this by comparing the probability the model gives the correct name against a random name, and finds no detectable difference [1]. I ran the same check, but the way I first ran it was misleading, and I think the mistake is worth showing.
My first comparison used a random name from the full pool of 400. By that measure the correct answer looked much worse than a random one: about −11.0 versus about −8.4 in log-probability. That would have made a dramatic story (“the model is actively steering away from the truth”). But it was an unfair comparison. The correct reverse answer is a name that was only ever a subject in training, and subjects are never the thing being predicted, so the model learns to give them low scores across the board. Half of the random names came from the other group, which had been pushed up.
Against a random name of the same kind, the gap disappears completely:
Correct reverse answer: about −10.97
Random name of the same kind: about −10.98
Random name from the whole pool (unfair baseline): about −8.37
So the model gives the right answer no more probability than any comparable wrong one. That matches the paper’s finding, and it’s a cleaner result than the dramatic version I almost reported.
Obvious objection: wouldn’t a bigger model just fix this?
My first model’s “notes” per word were only 32 numbers long. It’s fair to wonder whether the effect is just a small model failing to make the connection, and whether a real LLM’s billions of parameters would find room for it. I tested this by rerunning the same experiment eight times, with anywhere from 4 numbers per word up to 512, a 128-fold increase, and measuring reverse accuracy each time.

Reverse accuracy stayed at exactly 0% at every size. This echoes what the paper found at real scale: the pattern was flat across GPT-3 sizes from 350M to 175B parameters, and a much larger fine-tuning dataset didn’t help either [1]. Extra room to store information doesn’t change anything, because the problem was never a lack of space. It’s about what gets written into that space in the first place.
Why this happens, mechanically
Back to that one line of code: on each training step, only the subject word’s notes get updated. In everyday terms:
-
Each time the model sees “Zorvath Kellin is the Minister of Tides,” it rewrites its notes on “Zorvath Kellin” so it gets better at predicting what follows that name.
-
Its notes on “Minister of Tides” are never touched by that sentence, because that phrase was only ever the answer, never the word the prediction started from. I checked this directly: the notes for every name that only appeared as an answer are exactly what they were at the start of training. The change was 0.0.
-
So when I later ask about “Minister of Tides” as if it were the subject, the model reads notes it never trained, still at their random starting values, and has nothing to work with.
The paper offers a similar sketch for real models: training on “A is B” may change the model’s representation of A, but the update depends on predicting B from A, not on needing to predict A from B later. The authors call this update “myopic,” and they present it as a hypothesis, leaving the full explanation for future work [1]. My toy model shows this mechanism is enough to produce the effect. It doesn’t show that this is what goes on inside GPT-3.
What this toy model can’t tell you
The zero here is close to guaranteed by construction. In a model this simple, a name that only appeared as an answer has notes that were never trained as a subject, so failure in reverse is nearly automatic. Real transformers have many layers and attention, and they’re pretrained on text where facts appear in many orders. So this experiment shows the mechanism is sufficient, not that it’s what actually causes the effect at scale. It also uses one dataset and one random seed. The real-model evidence for the Reversal Curse is in the paper, and it’s far more convincing than anything a 60-line NumPy model can offer.
What this means beyond the toy model
The paper’s authors note that large pretraining sets are diverse enough that a fact often shows up in several orders, which may hide the curse for well-known entities. But entity mentions follow a long tail, so rarer facts may appear mostly in one direction [1]. That’s the situation where you’d expect a reversed question to fail.
It’s worth reading this next to grokking, the subject of my last piece. Grokking shows a model can eventually find a general structure if it keeps training on a task that rewards one. The Reversal Curse is the opposite case: a link that never forms, however long you train, because nothing in the objective asks for it.
The practical takeaway for anything built on LLM “knowledge” is that a model isn’t a symmetric fact database. Something it can recall from one side may not come back when you ask from the other, and nothing in the output warns you when that happens.
···
References
[1] Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., & Evans, O. (2023). The Reversal Curse: LLMs trained on “A is B” fail to learn “B is A”. arXiv:2309.12288.

