#benchmark
Benchmark: 23 AI articles covering benchmark news, analysis, and research
Articles
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models⭐8
EvalDetectBench measures evaluation awareness in frontier LLMs, offering a benchmark pipeline to detect when models recognize assessments and improve evaluation...
Anthropic Unveils Fable 5.1: Enhanced Performance at Lower Cost and Reduced Restrictions⭐8
Anthropic launches Fable 5.1 and Mythos 5.1, boosting AI performance and cutting costs while reducing safety restrictions and debuting Zero Data Retention.
Keenable AI Open-Sources NEEDLE: A Live Search Benchmark That Rebuilds Its Query Set Every Hour⭐8
Keenable AI releases NEEDLE, an open-source live search benchmark that refreshes its query set hourly for real-time AI evaluation.
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model⭐9
A 9GB local AI model answers 67% of 529,939 Jeopardy! clues across 41 seasons, proving portable, free knowledge—outperforming IBM's 2011 Watson on post-training...
Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token (TTFT)-First Benchmark⭐9
TTFT alone misleads voice AI latency benchmarks. This guide measures the full stack—STT, LLM, TTS, and speech-to-speech—to reveal which inference APIs truly fee...
Modality Maturity Index: A Benchmark for Assessing Multimodal Capabilities of Omni Models⭐9
Introducing Modality Maturity Index (MMI): a benchmark evaluating omni-models across text, image, audio, video & document inputs/outputs. Key insights from test...
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs⭐9
Evaluate LLM rhetorical fallacy generation with DeflectBench: prompt framing drives refusal more than claim content, revealing critical safety vulnerabilities.
OpenAI Unveils Jalapeño Chip, Claiming Faster AI Inference Than Nvidia⭐8
OpenAI's new Jalapeño chip outperforms Nvidia in AI inference benchmarks, marking a strategic move into custom silicon to cut costs and boost performance.
Travis Kalanick Reignites VC Criticism: 'Only 1% Are Truly Helpful'⭐8
Travis Kalanick slams VCs, claiming only 1% are truly helpful—sparking debate on founder-investor dynamics and the Uber fallout.
Diagnostic Foundation for Evaluating LLMs' Research Integrity as⭐10
Introducing IntegrityBench: a benchmark revealing that LLMs fail 1 in 3 integrity decisions under pressure, and scale doesn't fix it.
