#large language models
Large Language Models: 59 AI articles covering large language models news, analysis, and research
Articles
Mathematicians Demand Proof That OpenAI Didn't Train Its Models on Their WorkNEWโญ9
Mathematicians demand OpenAI prove it didnโt train its models on their unpublished work, raising data provenance and copyright concerns amid new AI breakthrough...
IFM Unveils K2 Horizon: Six Apache 2.0 Models from 0.9B to 375Bโญ8
IFM launches K2 Horizon: six Apache 2.0 models from 0.9B to 375B-A23B with open weights, training data, and code for fully transparent AI development.
Where Does Harness-Oimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agentsโญ8
Nearly all optimization value in self-evolving LLM agents resides in reflection/control slots, not role or strategy, as revealed by HARNESSEVO.
How to Build AI Systems That Know When They Don't Know: A Practical Guideโญ10
AI systems that recognize their own knowledge gaps and reduce blind guessing, using confidence scoring and fallback routing for production reliability.
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoningโญ8
Study introduces behaviorally grounded user profiles from social media, improving LLM personalization and multi-perspective reasoning over synthetic baselines.
Princeton, Ant Group, and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factorโญ6
AQuA: Princeton, Stanford & Ant Group's agentic framework prevents data leakage in quantitative finance, enabling autonomous factor discovery and model developm...
How AI Could Make It Harder for Governments to Use Hacking Toolsโญ9
As AI patches software bugs at scale, government hacking tools may become obsolete, challenging law enforcement's access to encrypted devices.
8 Tips for Writing Effective Agent Instructionsโญ10
Master agent instructions with these 8 expert tips for 2026. Learn to design clear, maintainable playbooks that optimize AI performance and workflow efficiency.
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMsโญ9
Evaluate LLM rhetorical fallacy generation with DeflectBench: prompt framing drives refusal more than claim content, revealing critical safety vulnerabilities.
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decodingโญ10
TreeGraft enables adaptive multi-drafter speculative decoding, boosting LLM inference speed by 15.1% across benchmarks.
