LLMs

Latest breakthroughs in Large Language Models

Articles

Intra-Paper Claim Verification for Peer Review: Do Methods
ClinLens: A New Benchmark for Long-Horizon Clinical Data Science
Probing the Origins of Reasoning Performance: Representational
Behavior-Driven Explainability
FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selective
Neuromorphic Diffusion Language Models: Addressing Compute and
TimeCapsule: Generative Hallucination as a Method for Historical
Kernel Forge: An Agent Harness for LLM-based Generation and
Beyond Memory: A Templated Substrate for Heterogeneous
Do Models Fake Alignment Without Clear Consequences?