LLMs

Latest breakthroughs in Large Language Models

Articles

PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models
DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts
Accelerating LLM Inference via Vector Index Based Output Embeddings
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
ElementCheck: Complexity-Aware Factuality Evaluation for Long-Form Text Generation