reversal-curse-shows-up-in-a-deployed-model-gpt-4observationsingle paper
GPT-4 could reliably name a well-known celebrity's parent, but was far less reliable naming the celebrity when given the parent — the same directional-recall asymmetry demonstrated in smaller fine-tuned models, showing up in a deployed model that wasn't fine-tuned for the test.
Capability: Not generalizing "A is B" to "B is A"
Observed on
GPT. 2023, GPT-4 class. General.
Sources
- GPT-4 named a celebrity's parent correctly far more often than it could name the celebrity when given the parent.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Models fine-tuned on "A is B" fail to answer "B is A", and GPT-4 recalls celebrity parents far more often than the reverse.Not generalizing "A is B" to "B is A"
- As a raw pairwise judge with no order-swap mitigation, GPT-4 can have its verdict flipped simply by swapping which candidate answer is shown first; applying a position-swap protocol restores agreement with careful human raters to close to human-human agreement levels.Biased when judging other outputs
- GPT-4 recommended non-existent Python and JavaScript packages in only a small percentage of generations — lower than the open models tested in the same study — but the same hallucinated names recurred often enough across runs to be practically exploitable by an attacker who registers them ahead of time.Writing secure code and dependencies