LLMs
Latest breakthroughs in Large Language Models
Articles
Detecting and Controlling Sycophancy with Cascading Linear Features⭐8
This paper introduces an iterative data pipeline that isolates cascading linear features to detect and control sycophancy in LLMs, offering more interpretable a...
Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners⭐8
This paper introduces Auto-World, a framework using LLMs to automate benchmark generation for neural relational reasoning, improving evaluators like Edge Transf...
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems⭐9
A practical guide to building autonomous AI agents in 2026, covering foundations, alignment, and multi-agent systems for engineers.
Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs⭐8
Neuro-Symbolic Drive uses rule-grounded reasoning traces from classical planners to train a driving VLA, reducing ADE@3s from 0.47 to 0.26 and miss rate from 8....
RIFT-Bench: Dynamic Red-teaming for Agentic AI Systems⭐8
RIFT-Bench introduces a graph-based dynamic red-teaming framework for evaluating security across diverse agentic AI systems, enabling adaptive adversarial attac...
Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies⭐7
Tree-of-Thought reasoning strategies show opposite limitations: DPTS fails at low budgets while SSDP plateaus early, highlighting the need for adaptive search u...
Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023⭐7
A study measures CS program alignment with CS2013 and CS2023 guidelines, finding near-constant knowledge unit coverage but a gap in competency depth under the n...
Deontic Policies for Runtime Governance of Agentic AI Systems⭐10
AgenticRei proposes deontic policies for runtime governance of AI agents, managing obligations, waivers, and conflict resolution beyond traditional access contr...
CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework with Evidence-Grounded Reasoning⭐7
CaVe-VLM-CoT introduces an interpretable vision-language model framework using evidence-grounded reasoning to reduce hallucinations and improve trust in AI outp...
NAVI-Orbital: First In-Orbit Demonstration of a Zero-Shot Vision-Language Model for Autonomous Earth Observation⭐7
NAVI-Orbital demonstrates the first in-orbit zero-shot vision-language model for autonomous Earth observation, enabling natural language queries and semantic co...
