arxiv-2405-17399 · paper

Transformers Can Do Arithmetic with the Right Embeddings

Sean McLeish, Arpit Bansal, Alex Stein, et al.

Created: 2024 · Ingested: 2026-09-02

This source: Transformers Can Do Arithmetic with the Right EmbeddingsHow much the field cites it — steadily cited in the last 12 months36 in the last 12 months · 91 totalpublished 2024checked 2026-09-04click for Semantic Scholar36 citations in the last 12 months · 91 total · checked 2026-09-04

https://arxiv.org/abs/2405.17399(opens in a new tab)

In brief

Transformer failure on multi-digit addition is largely a digit-alignment problem, not an arithmetic one: giving every digit an embedding encoding its position relative to the start of its number restores near-perfect accuracy and lets length generalization jump well past prior work.

Abacus Embeddings assign the same learned positional embedding to all digits of equal significance, with a random starting offset drawn from U[1,k], k=100, during training. Decoder-only models were trained from scratch on 20 million samples of reversed-format addition with operands up to 20 digits, under a cramming budget of 8 exaFLOP (1 RTXA4000 for 24 hours), then tested on unseen digit lengths, plus multiplication up to 15 digits and sorting arrays of up to 30 numbers of up to 30 digits.

Training on 20-digit operands generalized to 120-digit problems, a 6x generalization factor against a prior best of 2.5x, with up to 99% accuracy on 100-digit addition. Adding input injection and looped (weight-tied) layers raised OOD accuracy from 92.9% to 99.1%, an 87% error reduction; an 8x2 recurrent block halved OOD error versus 16x1. Sorting was mixed: Abacus beat FIRE on number-length OOD (68.63 vs 55.32) but lost on array-length OOD (9.67 vs 21.35), with Abacus+FIRE best on all-OOD at 4.48.

Effects are measured, averaged over 3 runs rather than best-of-10, with ablations over embedding type, architecture, recurrent block size, and effective depth. Scale is small and no natural-language task was tested, so compatibility with general-purpose models is argued from the FIRE and RoPE combinations, not demonstrated.

If you attribute arithmetic failure to reasoning capacity, this suggests checking positional representation first.

Written from the abstract by claude-opus-5 on 2026-09-05, with every figure checked against it. Not a substitute for the paper.

Referenced by

Claims in this catalog that draw on this source, and whether as support or counterpoint.