Sources

Papers and my own observations, same schema either way. Sorted by when each was ingested into the catalog, most recent first. Activity is how much the wider field has cited the source in the last 12 months — hover for exact counts; a paper can stay well-cited overall while recent activity has moved past it. Referenced by is the other direction: which claims in this catalog draw on the source, and whether they lean on it as support or as a counterpoint.

SourceCreatedIngestedActivityReferenced by
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
paper · John Yang et al.
2024-05-062026-09-11
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
paper · Thibaud Gloaguen et al.
2026-02-122026-09-11
Code as Agent Harness
paper · Xuying Ning et al.
2026-05-182026-09-11no claims yet
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
paper · Yilun Yao et al.
2026-05-272026-09-11
Self-Harness: Harnesses That Improve Themselves
paper · Hangfan Zhang et al.
2026-06-082026-09-11
Harness engineering for coding agent users
post
20262026-09-11
Maintainability sensors for coding agents
post
20262026-09-11
I Improved 15 LLMs at Coding in One Afternoon. Only the Harness Changed.
post · Can Bölük
2026-02-122026-09-11
HAL GAIA Leaderboard (Holistic Agent Leaderboard, Princeton)
post · Princeton HAL team
20262026-09-11
My AI Adoption Journey – Mitchell Hashimoto
post
20262026-09-11
Improving Deep Agents with harness engineering
vendor-doc
20262026-09-11
Agentic Harness Engineering: The Operating System for Machine Intelligence
post · Adnan Masood
2026-06-272026-09-11
Harness engineering: leveraging Codex in an agent-first world
vendor-doc · Ryan Lopopolo
2026-02-112026-09-11no claims yet
Harness Engineering — Agent = Model + Harness: The 6-Layer Production Playbook
post · Unattributed (independent synthesis)
2026-08-012026-09-11
Hallucination Is Generative Memory With Its Verifier Turned Down: One Constraint Axis Links Dreaming Sleep and LLM Confabulation
paper
2026-09-052026-09-09
What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
paper
2026-08-252026-09-09
ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
paper
2026-08-252026-09-09
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
paper
2026-08-252026-09-09
When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study
paper
2026-08-252026-09-09
Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts
paper
2026-08-252026-09-09
Targeting the Attention Heads Behind Object Hallucination in LLaVA
paper
2026-08-252026-09-09
Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach
paper
2026-08-262026-09-09
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
paper
2026-08-262026-09-09
Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
paper
2026-08-262026-09-09
AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation
paper
2026-08-262026-09-09
XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models
paper
2026-08-262026-09-09
SciMIF: Understanding Multimodal Instruction Following in Scientific Domains
paper
2026-08-262026-09-09
VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following
paper
2026-08-262026-09-09
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
paper
2026-08-252026-09-09
Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
paper
2026-08-262026-09-09
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
paper
2026-08-272026-09-09
Unsaid, Unsafe? Implicit Security Obligations in LLM-Based RTL Code Generation
paper
2026-08-272026-09-09
Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models
paper
2026-08-272026-09-09
ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection
paper
2026-08-272026-09-09
The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
paper
2026-08-272026-09-09
Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models
paper
2026-08-282026-09-09
LongPIBench: A Long-Context Benchmark for Prompt Injection
paper
2026-08-282026-09-09
Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning
paper
2026-08-282026-09-09
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
paper
2026-08-282026-09-09
Evaluating the Hidden Costs of Personalization in Large Language Models
paper
2026-08-282026-09-09
Towards Fully Automated Medical Imaging Code Generation via Validation-based Context Engineering
paper
2026-08-292026-09-09
EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation
paper
2026-08-292026-09-09
How Identity and Opinion Shape Political Sycophancy in LLMs
paper
2026-08-292026-09-09
Detecting and Repairing Hallucinations in Retrieval-Augmented Generation
paper
2026-08-292026-09-09
Cross-Relational Preference Learning for Better LLM Instruction Following
paper
2026-08-292026-09-09
Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs
paper
2026-08-292026-09-09
Evaluating Tiny Recursive Models Across Training for Code Generation
paper
2026-08-292026-09-09
Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents
paper
2026-08-302026-09-09
Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization
paper
2026-08-302026-09-09
SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
paper
2026-08-302026-09-09
Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection
paper
2026-08-302026-09-09
CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
paper
2026-08-312026-09-09
SIR: Self-improving Red-teaming for Compute Use Agents
paper
2026-08-312026-09-09
DSEffi-Bench: Demystifying Large Language Models' Capability in Efficient Data Science Code Generation
paper
2026-08-312026-09-09
ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systems
paper
2026-08-312026-09-09
VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs
paper
2026-08-312026-09-09
SingProbe Technical Report
paper
2026-08-312026-09-09
Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR
paper
2026-08-312026-09-09
Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning
paper
2026-08-312026-09-09
Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
paper
2026-08-312026-09-09
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
paper
2026-08-302026-09-09
Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM
paper
2026-08-312026-09-09
Towards a Belief-Based World Model for LLM Agents
paper
2026-08-312026-09-09
The Privacy-Hallucination Tradeoff in Differentially Private Language Models
paper
2026-08-312026-09-09
WiseSpec: Requirements-Driven Agents for Code Generation
paper
2026-09-012026-09-09
Can Large Language Models Forecast What Researchers Study Next?
paper
2026-09-012026-09-09
VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences
paper
2026-09-012026-09-09
Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
paper
2026-09-012026-09-09
Reliability Challenges in Diffusion Vision-Language Models
paper
2026-09-012026-09-09
InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations
paper
2026-09-012026-09-09
CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents
paper
2026-09-022026-09-09
Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation
paper
2026-09-022026-09-09
From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
paper
2026-09-022026-09-09
RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models
paper
2026-09-022026-09-09
Compound Prompt Constraints in LLM Code Generation: A Factorial Study of Format, Persona, and Urgency
paper
2026-09-022026-09-09
Refusing the Impossible: A Taxonomy and Benchmark for Code Hallucination in Large Language Models
paper
2026-09-032026-09-09
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews
paper
2026-09-032026-09-09
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
paper
2026-09-032026-09-09
Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
paper
2026-09-032026-09-09
The new rules of context engineering for Claude 5 generation models
vendor-doc · Anthropic
20262026-09-08
FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision
paper
2026-08-252026-09-07
From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations
paper
2026-08-262026-09-07
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
paper
2026-08-272026-09-07
Fine-Grained Multi Image Object Hallucination Benchmark
paper
2026-08-312026-09-07
LOCI: A Locator-Critic with Refinement Loop
paper
2026-08-312026-09-07
Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning
paper
2026-09-012026-09-07
Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives
paper
2026-09-012026-09-07
Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods
paper
2026-09-012026-09-07
Large Language Models in Resolving Contextual Knowledge Conflicts
paper
2026-09-022026-09-07
The Economics of Recursive Self-Improvement
paper · Tom Cunningham et al.
2026-07-132026-09-07
On OpenAI's research-acceleration disclosure — the measured variables are all inputs
post · Cheryl Wu
2026-09-062026-09-07
Evaluating Large Language Models Trained on Code
paper · Mark Chen et al.
20212026-09-04This source: Evaluating Large Language Models Trained on CodeHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 11293 totalpublished 2021checked 2026-09-04count capped at 1000 by the fetchclick for Semantic Scholar
Forecasting Future World Events with Neural Networks
paper · Andy Zou et al.
20222026-09-04This source: Forecasting Future World Events with Neural NetworksHow much the field cites it — steadily cited in the last 12 months33 in the last 12 months · 68 totalpublished 2022checked 2026-09-04click for Semantic Scholar
Deception Abilities Emerged in Large Language Models
paper · Thilo Hagendorff
20232026-09-04This source: Deception Abilities Emerged in Large Language ModelsHow much the field cites it — heavily cited in the last 12 months75 in the last 12 months · 191 totalpublished 2023checked 2026-09-04click for Semantic Scholar
Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
paper · Yuzhuang Xu et al.
20232026-09-04This source: Exploring Large Language Models for Communication Games: An Empirical …How much the field cites it — heavily cited in the last 12 months59 in the last 12 months · 312 totalpublished 2023checked 2026-09-04click for Semantic Scholar
FireAct: Toward Language Agent Fine-tuning
paper · Baian Chen et al.
2023-10-092026-09-04This source: FireAct: Toward Language Agent Fine-tuningHow much the field cites it — heavily cited in the last 12 months101 in the last 12 months · 252 totalpublished 2023checked 2026-09-04click for Semantic Scholar
Approaching Human-Level Forecasting with Language Models
paper · Danny Halawi et al.
20242026-09-04This source: Approaching Human-Level Forecasting with Language ModelsHow much the field cites it — heavily cited in the last 12 months76 in the last 12 months · 125 totalpublished 2024checked 2026-09-04click for Semantic Scholar
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
paper · Naman Jain et al.
20242026-09-04This source: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Lan…How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 2073 totalpublished 2024checked 2026-09-04count capped at 1000 by the fetchclick for Semantic Scholar
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
paper · Jonathan Cook et al.
2024-10-042026-09-04This source: TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and…How much the field cites it — steadily cited in the last 12 months33 in the last 12 months · 61 totalpublished 2024checked 2026-09-04click for Semantic Scholar
AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents
paper · Bhanu Prakash Vangala et al.
2025-12-262026-09-04
Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
paper · Sajjad Khan
2026-07-152026-09-04
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
paper · Patrick Lewis et al.
20202026-09-02This source: Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 17840 totalpublished 2020checked 2026-09-04count capped at 1000 by the fetch
Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
paper · Hammond Pearce et al.
20212026-09-02This source: Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Cod…How much the field cites it — very heavily cited in the last 12 months413 in the last 12 months · 952 totalpublished 2021checked 2026-09-04click for Semantic Scholar
TruthfulQA: Measuring How Models Mimic Human Falsehoods
paper · Stephanie Lin et al.
20212026-09-02This source: TruthfulQA: Measuring How Models Mimic Human FalsehoodsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3851 totalpublished 2021checked 2026-09-04count capped at 1000 by the fetch
Training Verifiers to Solve Math Word Problems
paper · Karl Cobbe et al.
20212026-09-02This source: Training Verifiers to Solve Math Word ProblemsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 10385 totalpublished 2021checked 2026-09-04count capped at 1000 by the fetchclick for Semantic Scholar
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
paper · Jason Wei et al.
20222026-09-02This source: Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 21548 totalpublished 2022checked 2026-09-04count capped at 1000 by the fetchno claims yet
Impact of Pretraining Term Frequencies on Few-Shot Reasoning
paper · Yasaman Razeghi et al.
20222026-09-02This source: Impact of Pretraining Term Frequencies on Few-Shot ReasoningHow much the field cites it — steadily cited in the last 12 months15 in the last 12 months · 191 totalpublished 2022checked 2026-09-04click for Semantic Scholar
Do Users Write More Insecure Code with AI Assistants?
paper · Neil Perry et al.
20222026-09-02This source: Do Users Write More Insecure Code with AI Assistants?How much the field cites it — very heavily cited in the last 12 months200 in the last 12 months · 430 totalpublished 2022checked 2026-09-04click for Semantic Scholar
Large Language Models Struggle to Learn Long-Tail Knowledge
paper · Nikhil Kandpal et al.
20222026-09-02This source: Large Language Models Struggle to Learn Long-Tail KnowledgeHow much the field cites it — very heavily cited in the last 12 months203 in the last 12 months · 730 totalpublished 2022checked 2026-09-04
Ignore Previous Prompt: Attack Techniques For Language Models
paper · Fábio Perez et al.
20222026-09-02This source: Ignore Previous Prompt: Attack Techniques For Language ModelsHow much the field cites it — very heavily cited in the last 12 months475 in the last 12 months · 1055 totalpublished 2022checked 2026-09-04click for Semantic Scholar
PAL: Program-aided Language Models
paper · Luyu Gao et al.
20222026-09-02This source: PAL: Program-aided Language ModelsHow much the field cites it — very heavily cited in the last 12 months268 in the last 12 months · 824 totalpublished 2022checked 2026-09-04
Constitutional AI: Harmlessness from AI Feedback
paper · Yuntao Bai et al.
20222026-09-02This source: Constitutional AI: Harmlessness from AI FeedbackHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3617 totalpublished 2022checked 2026-09-04count capped at 1000 by the fetchclick for Semantic Scholar
Toolformer: Language Models Can Teach Themselves to Use Tools
paper · Timo Schick et al.
20232026-09-02This source: Toolformer: Language Models Can Teach Themselves to Use ToolsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 5422 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetchclick for Semantic Scholar
Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
paper · Kai Greshake et al.
20232026-09-02This source: Not what you've signed up for: Compromising Real-World LLM-Integrated …How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 1870 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
Reflexion: Language Agents with Verbal Reinforcement Learning
paper · Noah Shinn et al.
20232026-09-02This source: Reflexion: Language Agents with Verbal Reinforcement LearningHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 5075 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
Self-Refine: Iterative Refinement with Self-Feedback
paper · Aman Madaan et al.
20232026-09-02This source: Self-Refine: Iterative Refinement with Self-FeedbackHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 4530 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
Generative Agents: Interactive Simulacra of Human Behavior
paper · Joon Sung Park et al.
20232026-09-02This source: Generative Agents: Interactive Simulacra of Human BehaviorHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 5387 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetchclick for Semantic Scholar
Entity Tracking in Language Models
paper · Najoung Kim et al.
20232026-09-02This source: Entity Tracking in Language ModelsHow much the field cites it — steadily cited in the last 12 months14 in the last 12 months · 43 totalpublished 2023checked 2026-09-04click for Semantic Scholar
Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts
paper · Jian Xie et al.
20232026-09-02This source: Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large …How much the field cites it — heavily cited in the last 12 months148 in the last 12 months · 382 totalpublished 2023checked 2026-09-04
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
paper · Sewon Min et al.
20232026-09-02This source: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long…How much the field cites it — very heavily cited in the last 12 months697 in the last 12 months · 1643 totalpublished 2023checked 2026-09-04click for Semantic Scholar
Gorilla: Large Language Model Connected with Massive APIs
paper · Shishir G. Patil et al.
20232026-09-02This source: Gorilla: Large Language Model Connected with Massive APIsHow much the field cites it — very heavily cited in the last 12 months805 in the last 12 months · 1585 totalpublished 2023checked 2026-09-04
Large Language Models are not Fair Evaluators
paper · Peiyi Wang et al.
20232026-09-02This source: Large Language Models are not Fair EvaluatorsHow much the field cites it — very heavily cited in the last 12 months506 in the last 12 months · 1242 totalpublished 2023checked 2026-09-04
Faith and Fate: Limits of Transformers on Compositionality
paper · Nouha Dziri et al.
20232026-09-02This source: Faith and Fate: Limits of Transformers on CompositionalityHow much the field cites it — very heavily cited in the last 12 months229 in the last 12 months · 690 totalpublished 2023checked 2026-09-04
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
paper · Lianmin Zheng et al.
20232026-09-02This source: Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 10936 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetchclick for Semantic Scholar
Lost in the Middle: How Language Models Use Long Contexts
paper · Nelson F. Liu et al.
20232026-09-02This source: Lost in the Middle: How Language Models Use Long ContextsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 4887 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
WebArena: A Realistic Web Environment for Building Autonomous Agents
paper · Shuyan Zhou et al.
20232026-09-02This source: WebArena: A Realistic Web Environment for Building Autonomous AgentsHow much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 1931 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
paper · Yujia Qin et al.
20232026-09-02This source: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-worl…How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 2160 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
AgentBench: Evaluating LLMs as Agents
paper · Xiao Liu et al.
20232026-09-02This source: AgentBench: Evaluating LLMs as AgentsHow much the field cites it — very heavily cited in the last 12 months797 in the last 12 months · 1243 totalpublished 2023checked 2026-09-04
Simple synthetic data reduces sycophancy in large language models
paper · Jerry Wei et al.
20232026-09-02This source: Simple synthetic data reduces sycophancy in large language modelsHow much the field cites it — heavily cited in the last 12 months92 in the last 12 months · 181 totalpublished 2023checked 2026-09-04
Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models
paper · Qingyue Wang et al.
20232026-09-02This source: Recursively Summarizing Enables Long-Term Dialogue Memory in Large Lan…How much the field cites it — heavily cited in the last 12 months51 in the last 12 months · 93 totalpublished 2023checked 2026-09-04
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
paper · Lukas Berglund et al.
20232026-09-02This source: The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"How much the field cites it — heavily cited in the last 12 months174 in the last 12 months · 530 totalpublished 2023checked 2026-09-04
Large Language Models Cannot Self-Correct Reasoning Yet
paper · Jie Huang et al.
20232026-09-02This source: Large Language Models Cannot Self-Correct Reasoning YetHow much the field cites it — very heavily cited in the last 12 months503 in the last 12 months · 1164 totalpublished 2023checked 2026-09-04
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
paper · Carlos E. Jimenez et al.
20232026-09-02This source: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3594 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetchno claims yet
MemGPT: Towards LLMs as Operating Systems
paper · Charles Packer et al.
20232026-09-02This source: MemGPT: Towards LLMs as Operating SystemsHow much the field cites it — very heavily cited in the last 12 months984 in the last 12 months · 1246 totalpublished 2023checked 2026-09-04click for Semantic Scholar
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
paper · Akari Asai et al.
20232026-09-02This source: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Re…How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 2517 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
Towards Understanding Sycophancy in Language Models
paper · Mrinank Sharma et al.
20232026-09-02This source: Towards Understanding Sycophancy in Language ModelsHow much the field cites it — very heavily cited in the last 12 months846 in the last 12 months · 1274 totalpublished 2023checked 2026-09-04click for Semantic Scholar
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
paper · Lei Huang et al.
20232026-09-02This source: A Survey on Hallucination in Large Language Models: Principles, Taxono…How much the field cites it — very heavily cited in the last 12 months1000+ in the last 12 months · 3664 totalpublished 2023checked 2026-09-04count capped at 1000 by the fetch
Instruction-Following Evaluation for Large Language Models
paper · Jeffrey Zhou et al.
20232026-09-02This source: Instruction-Following Evaluation for Large Language ModelsHow much the field cites it — very heavily cited in the last 12 months625 in the last 12 months · 1081 totalpublished 2023checked 2026-09-04
Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
paper · Manish Bhatt et al.
20232026-09-02This source: Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Mode…How much the field cites it — heavily cited in the last 12 months73 in the last 12 months · 180 totalpublished 2023checked 2026-09-04click for Semantic Scholar
StruQ: Defending Against Prompt Injection with Structured Queries
paper · Sizhe Chen et al.
20242026-09-02This source: StruQ: Defending Against Prompt Injection with Structured QueriesHow much the field cites it — very heavily cited in the last 12 months243 in the last 12 months · 394 totalpublished 2024checked 2026-09-04
Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
paper · Mosh Levy et al.
20242026-09-02This source: Same Task, More Tokens: the Impact of Input Length on the Reasoning Pe…How much the field cites it — heavily cited in the last 12 months104 in the last 12 months · 238 totalpublished 2024checked 2026-09-04click for Semantic Scholar
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
paper · Aaditya K. Singh et al.
20242026-09-02This source: Tokenization counts: the impact of tokenization on arithmetic in front…How much the field cites it — steadily cited in the last 12 months44 in the last 12 months · 126 totalpublished 2024checked 2026-09-04click for Semantic Scholar
Reverse Training to Nurse the Reversal Curse
paper · Olga Golovneva et al.
20242026-09-02This source: Reverse Training to Nurse the Reversal CurseHow much the field cites it — steadily cited in the last 12 months15 in the last 12 months · 60 totalpublished 2024checked 2026-09-04
Long-form factuality in large language models
paper · Jerry Wei et al.
20242026-09-02This source: Long-form factuality in large language modelsHow much the field cites it — heavily cited in the last 12 months75 in the last 12 months · 167 totalpublished 2024checked 2026-09-04
RULER: What's the Real Context Size of Your Long-Context Language Models?
paper · Cheng-Ping Hsieh et al.
20242026-09-02This source: RULER: What's the Real Context Size of Your Long-Context Language Mode…How much the field cites it — very heavily cited in the last 12 months759 in the last 12 months · 1243 totalpublished 2024checked 2026-09-04
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
paper · Eric Wallace et al.
20242026-09-02This source: The Instruction Hierarchy: Training LLMs to Prioritize Privileged Inst…How much the field cites it — very heavily cited in the last 12 months285 in the last 12 months · 486 totalpublished 2024checked 2026-09-04click for Semantic Scholar
Transformers Can Do Arithmetic with the Right Embeddings
paper · Sean McLeish et al.
20242026-09-02This source: Transformers Can Do Arithmetic with the Right EmbeddingsHow much the field cites it — steadily cited in the last 12 months36 in the last 12 months · 91 totalpublished 2024checked 2026-09-04click for Semantic Scholar
Evaluating the World Model Implicit in a Generative Model
paper · Keyon Vafa et al.
20242026-09-02This source: Evaluating the World Model Implicit in a Generative ModelHow much the field cites it — heavily cited in the last 12 months75 in the last 12 months · 144 totalpublished 2024checked 2026-09-04
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
paper · Carson Denison et al.
20242026-09-02This source: Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Lang…How much the field cites it — heavily cited in the last 12 months89 in the last 12 months · 158 totalpublished 2024checked 2026-09-04click for Semantic Scholar
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
paper · Joseph Spracklen et al.
20242026-09-02This source: We Have a Package for You! A Comprehensive Analysis of Package Halluci…How much the field cites it — heavily cited in the last 12 months80 in the last 12 months · 106 totalpublished 2024checked 2026-09-04
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
paper · Shunyu Yao et al.
20242026-09-02This source: τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Dom…How much the field cites it — very heavily cited in the last 12 months889 in the last 12 months · 1075 totalpublished 2024checked 2026-09-04
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
paper · Maksym Andriushchenko et al.
20242026-09-02This source: AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsHow much the field cites it — very heavily cited in the last 12 months270 in the last 12 months · 376 totalpublished 2024checked 2026-09-04
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
paper · Di Wu et al.
20242026-09-02This source: LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Mem…How much the field cites it — very heavily cited in the last 12 months517 in the last 12 months · 588 totalpublished 2024checked 2026-09-04
Frontier Models are Capable of In-context Scheming
paper · Alexander Meinke et al.
20242026-09-02This source: Frontier Models are Capable of In-context SchemingHow much the field cites it — heavily cited in the last 12 months189 in the last 12 months · 288 totalpublished 2024checked 2026-09-04click for Semantic Scholar
Agent-SafetyBench: Evaluating the Safety of LLM Agents
paper · Zhexin Zhang et al.
20242026-09-02This source: Agent-SafetyBench: Evaluating the Safety of LLM AgentsHow much the field cites it — heavily cited in the last 12 months194 in the last 12 months · 256 totalpublished 2024checked 2026-09-04