LLMs
Latest breakthroughs in Large Language Models
Articles
The AI-Model Network: Concept, Current State, and Future Directionsโญ8
The AI-Model Network proposes a global system for interconnecting, sharing, and collaborating across heterogeneous AI models, addressing high costs and
Life After Benchmark Saturation: A Case Study of CORE-Benchโญ8
Learn what happens after AI benchmarks hit accuracy saturation. A CORE-Bench case study reveals six crucial dimensions beyond accuracy.
Detecting and Controlling Sycophancy with Cascading Linear Featuresโญ8
This paper introduces an iterative data pipeline that isolates cascading linear features to detect and control sycophancy in LLMs, offering more
Project Auto-World: Towards Automated Benchmarking of Neuralโญ8
This paper introduces Auto-World, a framework using LLMs to automate benchmark generation for neural relational reasoning, improving evaluators like Edge
The Hitchhiker's Guide to Agentic AI: From Foundations to Systemsโญ9
A practical guide to building autonomous AI agents in 2026, covering foundations, alignment, and multi-agent systems for engineers.
Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning forโญ8
Neuro-Symbolic Drive uses rule-grounded reasoning traces from classical planners to train a driving VLA, reducing ADE@3s from 0.47 to 0.26 and miss rate from 8....
RIFT-Bench: Dynamic Red-teaming for Agentic AI Systemsโญ8
RIFT-Bench introduces a graph-based dynamic red-teaming framework for evaluating security across diverse agentic AI systems, enabling adaptive adversarial
Beyond Fixed Budgets: Characterizing the Inelasticity andโญ7
Tree-of-Thought reasoning strategies show opposite limitations: DPTS fails at low budgets while SSDP plateaus early, highlighting the need for adaptive
Measuring Curriculum Alignment across Topical Coverage,โญ7
A study measures CS program alignment with CS2013 and CS2023 guidelines, finding near-constant knowledge unit coverage but a gap in competency depth under the n...
Deontic Policies for Runtime Governance of Agentic AI Systemsโญ10
AgenticRei proposes deontic policies for runtime governance of AI agents, managing obligations, waivers, and conflict resolution beyond traditional access contr...
