LLMs
Latest breakthroughs in Large Language Models
Articles
Adversarial Style Optimization: Enhancing VLM Jailbreaks byโญ7
ASO enhances visual jailbreak attacks on multimodal LLMs by optimizing stylistic triggers via GRPO, achieving higher Attack Success Rates and exposing stylistic...
Securing Multimodal AI through Internal Information Decompositionโญ8
FlowGuard detects harmful multimodal inputs by monitoring cross-modal consistency, reducing attack success rates from >90% to <15% with minimal utility loss.
Risk Is Not the Target: A Monotonic Framework for Evaluatingโญ7
Evaluating wildfire risk systems with standard metrics is flawed. A novel monotonic framework measures if risk scores align with operational load,
FlowEvo: Self-Evolving Agents through the Co-Evolution ofโญ7
FlowEvo is a training-free framework enabling LLM agents to co-evolve workflows and executable skills, achieving 82.8% success on ALFWorld with less than
The Active Ingredient in Muon's Grokkingโญ8
Ablation study reveals Muon's grokking speedup comes from orthogonalization, not spectral constraints. Newton-Schulz iteration reduces spectral norm by
PhantomFill: When the Form Demands an Answer, Language Modelsโญ9
PhantomFill: Required form fields force language models to fabricate answers for unanswerable inputs, with 10 of 13 models showing 100% fabrication rates.
DataPrep-Bench: Benchmarking LLMs as Training Data Preparatorsโญ8
DataPrep-Bench evaluates LLMs as training data preparers across data construction and quality evaluation tasks, introducing new metrics and benchmarks for six d...
Is MoE Routing a Huffman Code? Discovering theโญ7
Mixture-of-Experts routing follows a Frequency-Diversity Law, acting like neural Huffman coding. New Subset Difference Pruning eliminates redundancy,
Knowledge Injection in MoE? Expert-Aware Contrast Decoding forโญ7
Expert-aware contrast decoding leverages distinct expert activation patterns to mitigate LLM hallucinations in MoE models, outperforming baselines across benchm...
ClickGuard: Detecting and Spoiling Clickbait News withโญ7
ClickGuard uses a hybrid AI system with Transformer models and a 'baitness' score to detect clickbait, achieving 91% F1-score. It provides real-time
