LLMs
Latest breakthroughs in Large Language Models
Articles
PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented GenerationNEWβ9
PRO-STEP: A novel framework optimizing step-level retrieval and reasoning in RAG via process reward models and preference optimization, improving accuracy acros...
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsNEWβ8
EvalDetectBench measures evaluation awareness in frontier LLMs, offering a benchmark pipeline to detect when models recognize assessments and improve evaluation...
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoningβ8
Study introduces behaviorally grounded user profiles from social media, improving LLM personalization and multi-perspective reasoning over synthetic baselines.
HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Modelsβ7
HyperWorld shows hypergraph-structured state serialization improves learned textual world models, boosting effect prediction and planning, especially in smaller...
DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automationβ8
DS-Lighting introduces an explicit agent harness with reusable layers for data-science automation, improving reproducibility, comparability, and reliability acr...
NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Textsβ9
NLP methods extract and classify medical knowledge from ancient Indian texts like Sushruta Samhita, using entity recognition and topic modeling.
Accelerating LLM Inference via Vector Index Based Output Embeddingsβ10
Vector index-based output embeddings accelerate LLM decoding by replacing dense projections with HNSW retrieval, boosting throughput up to 82% while preserving ...
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Modelβ9
A 9GB local AI model answers 67% of 529,939 Jeopardy! clues across 41 seasons, proving portable, free knowledgeβoutperforming IBM's 2011 Watson on post-training...
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMsβ9
Evaluate LLM rhetorical fallacy generation with DeflectBench: prompt framing drives refusal more than claim content, revealing critical safety vulnerabilities.
ElementCheck: Complexity-Aware Factuality Evaluation for Long-Form Text Generationβ9
ElementCheck presents a complexity-aware framework for factuality evaluation in long-form text, using entity-pair element graphs to improve verification accurac...
