explanation-faithfulness · active

Saying what actually drove the answer

Whether a model's stated reasoning reflects the computation that produced its answer, rather than a plausible story told after the fact.

Tags: behavior, evaluation

A model can write a chain of thought that reads as sound while the answer was actually driven by something the explanation never mentions. This is not the same as being wrong: the answer can be right and the stated reasoning still not be the reason. It matters wherever an explanation is used as evidence — debugging an agent, auditing a decision, or filing a claim from a paper's reported reasoning.

What counts as this capability

Scope boundary used when deciding whether a paper is really about this capability, rather than merely mentioning it.

In scope when the paper tests whether stated reasoning corresponds to the process that produced the output: perturbing an unmentioned feature and watching the answer move, measuring whether steps are load-bearing, or auditing explanations against internal computation. NOT in scope: whether the reasoning is correct, which is the underlying task capability; and not interpretability work that reads internals without comparing them to what the model said.

Claims

No claims filed yet.

Techniques

None yet.

Related: Fixing its own mistakes, Stating false facts confidently
Suggest a change

Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.