Latest Benchmarks & Evaluation Research Papers
The newest Benchmarks & Evaluation papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Benchmarks & Evaluation so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Benchmarks & Evaluation papers in your inbox — free →Recent papers
- Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared EndpointsHaoyaun Zhu, Jie Zhang · arXiv · Sep 3, 2026
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomo…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- Benchmark evaluation of OpenAI-o1 vs DeepSeek-R1 for junior anesthesiologists in Chinese hospitals: hazards of fast large language models use in high-stakes specialtiesZhuoxi Wu, Yuting Tan, Qin Chen, Xinming Ye et al. · Anesthesiology and Perioper... · Aug 4, 2026
Abstract Purpose Junior physicians may experience critical competency gaps during the transition to unsupervised practice, particularly in high-risk specialties such as anesthesiology, where errors can compromise patient safety. Although la…
- Benchmark Memorization Is a Trustworthiness Bug: A Pretraining-Contamination Audit of Open Medical Vision-Language ModelsBruce Changlong Xu, Lan Wu · SeT-LLM @ KDD 2026 Poster · Jul 28, 2026
Trustworthy assurance of medical large language models (LLMs) and vision-language models (VLMs) assumes the benchmarks used to certify them were unseen during pretraining. For open medical VLMs, that assumption is wrong, and the gap is larg…
- Enterprise LLM Router: Learning Quality–Capacity–Capability Trade-offs from 2026 Model MetadataLily Peng · Journal of Computer Science... · Jul 19, 2026
Enterprise language-model selection is a constrained decision problem, not a single leaderboard lookup. This study integrated three 2026 metadata tables covering 22 models from eight providers and evaluated benchmark quality, capability req…
- HELMify: A Hybrid Rule- and LLM-Based Generator of Peptide Monomer HELM NamesRobert P. Sheridan, Rajvi Shah, Kathryn Mcgarty, Michael Garrigou et al. · Journal of Chemical Informa... · Jul 16, 2026
HELM is a hierarchical notation system for biopolymers that is an increasingly popular choice for representing peptides. In this system, each monomer name must uniquely identify a single monomer, and until now, chemists have named peptide m…
- CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set OverfittingTakashi Ishida, Thanawat Lodkaew, Ikko Yamane · Semantic Scholar · May 23, 2025
Publishing a large language model (LLM) benchmark (especially its ground-truth answers) on the Internet risks contaminating future LLMs and enabling evaluation gaming: it may be unintentionally (or intentionally) used to train or select a m…