Articles
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models⭐7
A consensus-based framework for LLM evaluation measures relative preference among model-generated responses instead of absolute correctness, using a voting pane...
Knowledge Injection in MoE? Expert-Aware Contrast Decoding for Hallucination Mitigation in LLMs⭐7
Expert-aware contrast decoding leverages distinct expert activation patterns to mitigate LLM hallucinations in MoE models, outperforming baselines across benchm...
ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models⭐7
ClickGuard uses a hybrid AI system with Transformer models and a 'baitness' score to detect clickbait, achieving 91% F1-score. It provides real-time warnings an...
The Essential AI Glossary for 2026⭐9
Essential AI glossary for 2026: Learn key terms like AGI, AI agents, LLMs, and more. Plain-English definitions updated regularly.
When Does Personality Composition Matter for Multi-Agent LLM Teams?⭐9
Personality composition impacts multi-agent LLM teams differently by task: low agreeableness hinders open-ended tasks but has minimal effect on structured codin...
Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners⭐8
This paper introduces Auto-World, a framework using LLMs to automate benchmark generation for neural relational reasoning, improving evaluators like Edge Transf...
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems⭐9
A practical guide to building autonomous AI agents in 2026, covering foundations, alignment, and multi-agent systems for engineers.
Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies⭐7
Tree-of-Thought reasoning strategies show opposite limitations: DPTS fails at low budgets while SSDP plateaus early, highlighting the need for adaptive search u...
Sakana AI Introduces Sakana Fugu: A Dynamic Orchestration Model for Routing Tasks Across Interchangeable Frontier LLMs⭐8
Sakana AI launches Sakana Fugu, a dynamic orchestration model that routes tasks across swappable frontier LLMs for cost, performance, and vendor flexibility.
The 7 Types of Agent Memory: A Technical Guide for AI Engineers⭐7
Discover the 7 essential types of agent memory for AI engineers. This 2026 guide explains how to transform stateless LLMs into context-aware, production-ready s...
