#model behavior
Model Behavior: 2 AI articles covering model behavior news, analysis, and research
Articles
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole TopicNEWβ8
AI safety in 2026 demands refusing harmful subsets, not whole topicsβbalancing precision, transparency, and user trust to avoid over-censorship.
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Modelsβ8
EvalDetectBench measures evaluation awareness in frontier LLMs, offering a benchmark pipeline to detect when models recognize assessments and improve evaluation...
