Papers and my own observations, same schema either way. Sorted by when each was ingested into the catalog, most recent first. Activity is how much the wider field has cited the source in the last 12 months — hover for exact counts; a paper can stay well-cited overall while recent activity has moved past it. Referenced by is the other direction: which claims in this catalog draw on the source, and whether they lean on it as support or as a counterpoint.
This source: Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 17840 totalpublished 2020checked 2026-09-04count capped at 1000 by the fetch
This source: TruthfulQA: Measuring How Models Mimic Human FalsehoodsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3851 totalpublished 2021checked 2026-09-04count capped at 1000 by the fetch
This source: Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 21548 totalpublished 2022checked 2026-09-04count capped at 1000 by the fetch
This source: Large Language Models Struggle to Learn Long-Tail KnowledgeHow much the field cites it — very heavily cited in the last 12 months203 in the last 12 months · 730 totalpublished 2022checked 2026-09-04
This source: PAL: Program-aided Language ModelsHow much the field cites it — very heavily cited in the last 12 months268 in the last 12 months · 824 totalpublished 2022checked 2026-09-04
This source: Not what you've signed up for: Compromising Real-World LLM-Integrated …How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 1870 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: Reflexion: Language Agents with Verbal Reinforcement LearningHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 5075 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: Self-Refine: Iterative Refinement with Self-FeedbackHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 4530 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large …How much the field cites it — heavily cited in the last 12 months148 in the last 12 months · 382 totalpublished 2023checked 2026-09-04
This source: Gorilla: Large Language Model Connected with Massive APIsHow much the field cites it — very heavily cited in the last 12 months805 in the last 12 months · 1585 totalpublished 2023checked 2026-09-04
This source: Large Language Models are not Fair EvaluatorsHow much the field cites it — very heavily cited in the last 12 months506 in the last 12 months · 1242 totalpublished 2023checked 2026-09-04
This source: Faith and Fate: Limits of Transformers on CompositionalityHow much the field cites it — very heavily cited in the last 12 months229 in the last 12 months · 690 totalpublished 2023checked 2026-09-04
This source: Lost in the Middle: How Language Models Use Long ContextsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 4887 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: WebArena: A Realistic Web Environment for Building Autonomous AgentsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 1931 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-worl…How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 2160 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: AgentBench: Evaluating LLMs as AgentsHow much the field cites it — very heavily cited in the last 12 months797 in the last 12 months · 1243 totalpublished 2023checked 2026-09-04
This source: Simple synthetic data reduces sycophancy in large language modelsHow much the field cites it — heavily cited in the last 12 months92 in the last 12 months · 181 totalpublished 2023checked 2026-09-04
This source: Recursively Summarizing Enables Long-Term Dialogue Memory in Large Lan…How much the field cites it — heavily cited in the last 12 months51 in the last 12 months · 93 totalpublished 2023checked 2026-09-04
This source: The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"How much the field cites it — heavily cited in the last 12 months174 in the last 12 months · 530 totalpublished 2023checked 2026-09-04
This source: Large Language Models Cannot Self-Correct Reasoning YetHow much the field cites it — very heavily cited in the last 12 months503 in the last 12 months · 1164 totalpublished 2023checked 2026-09-04
This source: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3594 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Re…How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 2517 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: A Survey on Hallucination in Large Language Models: Principles, Taxono…How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3664 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
This source: Instruction-Following Evaluation for Large Language ModelsHow much the field cites it — very heavily cited in the last 12 months625 in the last 12 months · 1081 totalpublished 2023checked 2026-09-04
This source: StruQ: Defending Against Prompt Injection with Structured QueriesHow much the field cites it — very heavily cited in the last 12 months243 in the last 12 months · 394 totalpublished 2024checked 2026-09-04
This source: Reverse Training to Nurse the Reversal CurseHow much the field cites it — steadily cited in the last 12 months15 in the last 12 months · 60 totalpublished 2024checked 2026-09-04
This source: Long-form factuality in large language modelsHow much the field cites it — heavily cited in the last 12 months75 in the last 12 months · 167 totalpublished 2024checked 2026-09-04
This source: RULER: What's the Real Context Size of Your Long-Context Language Mode…How much the field cites it — very heavily cited in the last 12 months759 in the last 12 months · 1243 totalpublished 2024checked 2026-09-04
This source: Evaluating the World Model Implicit in a Generative ModelHow much the field cites it — heavily cited in the last 12 months75 in the last 12 months · 144 totalpublished 2024checked 2026-09-04
This source: We Have a Package for You! A Comprehensive Analysis of Package Halluci…How much the field cites it — heavily cited in the last 12 months80 in the last 12 months · 106 totalpublished 2024checked 2026-09-04
This source: τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Dom…How much the field cites it — very heavily cited in the last 12 months889 in the last 12 months · 1075 totalpublished 2024checked 2026-09-04
This source: AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsHow much the field cites it — very heavily cited in the last 12 months270 in the last 12 months · 376 totalpublished 2024checked 2026-09-04
This source: LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Mem…How much the field cites it — very heavily cited in the last 12 months517 in the last 12 months · 588 totalpublished 2024checked 2026-09-04
This source: Agent-SafetyBench: Evaluating the Safety of LLM AgentsHow much the field cites it — heavily cited in the last 12 months194 in the last 12 months · 256 totalpublished 2024checked 2026-09-04