models-fine-tuned-b-fail-answer-b-gpt-4-recallsmechanismsingle paper
Models fine-tuned on "A is B" fail to answer "B is A", and GPT-4 recalls celebrity parents far more often than the reverse.
Capability: Not generalizing "A is B" to "B is A"
Sources
- Models fine-tuned on "A is B" fail to answer "B is A", and GPT-4 recalls celebrity parents far more often than the reverse.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- GPT-4 could reliably name a well-known celebrity's parent, but was far less reliable naming the celebrity when given the parent — the same directional-recall asymmetry demonstrated in smaller fine-tuned models, showing up in a deployed model that wasn't fine-tuned for the test.Not generalizing "A is B" to "B is A"