#large language models
Large Language Models: 59 AI articles covering large language models news, analysis, and research
Articles
Does a Language Server Save Tokens for Coding Agents? Aโญ9
Does a language server save tokens for coding agents? A measurement study shows LSP raises token usage on localization tasks but helps reference completeness.
Diagnostic Foundation for Evaluating LLMs' Research Integrity asโญ10
Introducing IntegrityBench: a benchmark revealing that LLMs fail 1 in 3 integrity decisions under pressure, and scale doesn't fix it.
Position: Reasoning is a Learnable Rule-Based Processโญ10
Reasoning in AI needs clear definitions. This paper positions it as a learnable rule-based process and proposes best practices for evaluating and communicating ...
TRACE Bench: Task-Driven Roleplay Agentic Checklist Evaluationโญ9
TRACE Bench introduces a task-driven roleplay evaluation framework with agentic checklists, achieving 99.91% coverage, traceable scores, and stable rankings acr...
Closed-Loop LLM Co-Pilots for Digital Agricultureโญ8
LLM co-pilots autonomously optimize crop production, cutting energy use by 68% and cycle times by 35% in closed-loop digital agriculture trials.
Unreleased Anthropic Model Makes Strides on the Riemann Hypothesisโญ9
Anthropic's unreleased AI model makes significant progress on the Riemann hypothesis, autonomously testing 650 approaches to improve the prime number theory bou...
The Download: The Next Big Thing in LLMs and How AI Academicโญ8
New LLM startups tackle transformer limits and AI research shifts, exploring faster architectures and academic evolution.
PragyaDoc: A Universal Document Intelligence Framework forโญ9
PragyaDoc framework enables multilingual medical document understanding in low-resource settings, bridging India's 22-language healthcare gap via OCR,
TEXAS: Task-Expert-Aware Supervision for Downstreamโญ9
TEXAS method improves MoE LLM adaptation by identifying task-relevant experts and optimizing token supervision, boosting performance across benchmarks.
Simulator-Grounded Large Language Models for Industrial Causalโญ8
Compare three methods for grounding LLMs in a wastewater simulator, achieving up to 99.5% causal QA accuracy and robust cross-plant transfer.
