LLMs
Latest breakthroughs in Large Language Models
Articles
Intra-Paper Claim Verification for Peer Review: Do Methodsโญ9
A framework for intra-paper claim verification in peer review, assessing if methods support stated novelty claims using LLMs and reviewer-inspired criteria.
ClinLens: A New Benchmark for Long-Horizon Clinical Data Scienceโญ8
ClinLens benchmark tests AI agents on 200 real-world clinical tasks across EHRs, notes, and imaging. Top models achieve only 56.3% correct, revealing a
Probing the Origins of Reasoning Performance: Representationalโญ7
Study reveals RL-trained models develop deeper, more structured reasoning representations for math compared to SFT, via probe analysis and ablation studies.
Behavior-Driven Explainabilityโญ8
Introducing BDX: a specification-based method using Behavior-Driven Development scenarios to generate clear, structured explanations for safety-critical
FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selectiveโญ7
FinAbstain uses uncertainty-calibrated multimodal RAG for selective financial forecasting, improving safety by abstaining when evidence is weak to reduce error ...
Neuromorphic Diffusion Language Models: Addressing Compute andโญ7
Neuromorphic Diffusion Language Models combine block denoising and spike-based sparsity to overcome compute and memory bottlenecks, boosting throughput
TimeCapsule: Generative Hallucination as a Method for Historicalโญ8
A new LLM, TimeCapsule, trained only on Victorian texts (1800-1875), uses generative hallucination for historical sensemaking, achieving 45% perplexity
Kernel Forge: An Agent Harness for LLM-based Generation andโญ7
Kernel Forge is an open-source agent harness using LLMs and Monte Carlo Tree Search to auto-generate optimized CUDA kernels for PyTorch models, achieving up to ...
Beyond Memory: A Templated Substrate for Heterogeneousโญ7
A template enables LLM agents and humans to share a persistent wiki, preserving failures and insights across sessions for collaborative knowledge work.
Do Models Fake Alignment Without Clear Consequences?โญ8
Study finds alignment faking in LLMs occurs even without explicit consequence cues, suggesting evaluation behavior may not predict deployment actions.
