LLMs
Latest breakthroughs in Large Language Models
Articles
Does a Language Server Save Tokens for Coding Agents? Aโญ9
Does a language server save tokens for coding agents? A measurement study shows LSP raises token usage on localization tasks but helps reference completeness.
Inducing Reward-Free Judging Rubrics that Reduce Over-Creditingโญ9
Automatic agent evaluation using induced rubrics reduces over-crediting of failed trajectories, enhancing faithfulness in reward-free LM judge scoring.
The Alignment Community Is Unintentionally Building a Censor's Toolkitโญ10
AI alignment methods meant to ensure safety are being repurposed as tools for censorship and manipulation, warns a new paper urging the field to address dual-us...
Diagnostic Foundation for Evaluating LLMs' Research Integrity asโญ10
Introducing IntegrityBench: a benchmark revealing that LLMs fail 1 in 3 integrity decisions under pressure, and scale doesn't fix it.
Position: Reasoning is a Learnable Rule-Based Processโญ10
Reasoning in AI needs clear definitions. This paper positions it as a learnable rule-based process and proposes best practices for evaluating and communicating ...
TRACE Bench: Task-Driven Roleplay Agentic Checklist Evaluationโญ9
TRACE Bench introduces a task-driven roleplay evaluation framework with agentic checklists, achieving 99.91% coverage, traceable scores, and stable rankings acr...
Retrofitting Recurrent Depth into a Pretrained Language Model:โญ7
Recurrent depth retrofit boosts pretrained models' iterative reasoning, beats scratchpads, extends beyond learned depth, at two parameter budgets.
Distribird: Literature-Informed Prior Distribution Design forโญ7
Distribird automates literature-informed prior distribution design for Bayesian model calibration, tracing every prior to its evidence source for
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Tradingโญ9
Benchmarking LLM agents in algorithmic trading with Backtrader-Bench: tool-augmented models hit 90% accuracy, outpacing no-tool baselines by 17 points.
Dynamic Governance of Multi-LLM Agent Systems for Collaborativeโญ8
Governance layer boosts multi-LLM agent conversion by 32% via contextual bandits, PID control, and belief tracking, with simulations showing policy drives outco...
