LLMs
Latest breakthroughs in Large Language Models
Articles
QFoldAgent: An Autonomous Quantum Optimization Multi-Agentโญ7
QFoldAgent uses a multi-agent system with quantum optimization to improve protein structure prediction, reducing RMSD and boosting structural validity in 5-resi...
SeT-Diff: Semantic Foundation Models for HPC Telemetry and Time-Seriesโญ9
SeT-Diff introduces the first foundation model for HPC telemetry, using semantic conditioning and diffusion to achieve 0.047 MAE reconstruction, zero-shot
CausalGate: Causal Importance Distillation for Transformerโญ8
CausalGate uses causal importance distillation to prune transformer modules, reducing compute and latency in LLMs without runtime overhead.
CORVUS: Context Optimization and Reduction Via Underlyingโญ8
CORVUS optimizes LLM coding agents by decoupling file reads from history, reducing tokens by 9-50% and reasoning cycles up to 37% while maintaining accuracy.
Semalith v1.4: A Calibrated 184M Safety Classifier Achievingโญ7
Semalith v1.4 is an 184M safety classifier achieving top prompt-injection detection with 44x fewer parameters than Llama-Guard-3-8B, plus zero false positives.
Evaluating the Impact of Reviewer Guideline Design on LLM-Basedโญ6
Study compares LLM-based peer review using official vs. reviewer-imitating guidelines, finding official guidelines align best with human judgments.
MioFFAn: An Annotation Software for Formula Formalization withโญ9
MioFFAn is an open-source annotation tool for formula formalization, enabling dataset creation with customized taxonomies and partial LLM automation.
On the Depth Scalability of Logic Gate Networksโญ9
This paper identifies why Logic Gate Networks fail to scale with depth and introduces Input-Anchored Logic Gate Networks, achieving consistent accuracy gains be...
Cloud-Native Evaluation-as-a-Service: A Microservicesโญ8
A cloud-native architecture for scalable AI monitoring using six stateless microservices, with conformal guarantees and real-time drift detection.
A Consensus-Based Framework for Relative Preference Evaluationโญ7
A consensus-based framework for LLM evaluation measures relative preference among model-generated responses instead of absolute correctness, using a voting pane...
