LLMs
Latest breakthroughs in Large Language Models
Articles
Cost-Effective Automated Judging of Natural-Languageโญ9
Cheap open-weight LLMs judge math proofs with accuracy matching frontier models at up to 100ร lower cost, though ensemble voting offers no advantage.
Revisiting Classic Thought Experiments to Measure Consciousnessโญ9
Reinterpreting Leibniz, Turing, and Searle to separate AI task performance from structural consciousness, key for safety.
Guarantees on Dynamical System Distinguishability for LLM Tokenโญ8
DS-based LLM output detection achieves exponential error decay with sequence length, formalized via spectral distance between linear dynamical systems.
Sensitivity Analysis of GRU, LSTM, and Transformer Encoder forโญ7
Compare GRU, LSTM, and Transformer encoders for classifying Level 2 automated driving systems, with robustness analysis under telematics data corruption.
Topology-Aware Data Movement for Disaggregated GPU Inferenceโญ10
Topology-aware KV cache transfer cuts GPU inference latency by 3-18x via NVLink, InfiniBand, and CXL optimizations.
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judgesโญ8
Cross-model auditing with Chain-of-Models identifies bias-specific optimal auditors, improving LLM judge accuracy to 0.884 across four cognitive bias types.
Imbalanced Data Clustering via Targeted Data Augmentation Usingโญ7
Explore a novel method combining GMMs and LLMs for targeted data augmentation, improving clustering on imbalanced text datasets by enriching underrepresented cl...
Can LLMs Really Understand Item Difficulty Levels? Implicationsโญ9
Can LLMs accurately predict item difficulty? This study tests GPT-4.1 and GPT-5.4 against supervised models, revealing semantic limits and risks for automated i...
LLM Framework for Discovering Major Mathematical Conjectures:โญ8
LLM-driven three-stage pipeline for discovering major mathematical conjectures, combining region search, reflective validation, and Lean 4 formal proof to
Can AI Evaluate AI Scientists? A Benchmarking Study ofโญ9
Benchmarking AI Scientist systems via automated multi-model review reveals FARS papers score 2x higher across originality, rigor, clarity, and significance.
