LLMs
Latest breakthroughs in Large Language Models
Articles
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Textsβ9
Expert-led evaluation reveals LLM watermarking degrades medical text quality, masking clinical errors. Study assesses 5 schemes across 11 LLMs and 7 VLMs,
AINTMA: Agentic AI Architecture for Autonomous Test Managementβ6
AINTMA introduces a multi-agent AI framework for autonomous test management, achieving 88.4% prioritization accuracy and 43% defect detection reduction across 1...
Auto-FL-Research: Agentic Search for Federated Learning Algorithmsβ8
Auto-FL-Research introduces an agentic search framework for automating federated learning algorithm design. Tests on 11 healthcare and LEAF benchmarks
PACE: A Neuro-Symbolic Framework for Plausible and Actionableβ7
PACE proposes a neuro-symbolic framework for counterfactual explanations, combining neural prediction with symbolic constraints to generate more realistic
Bounded Morality: Defining the Space of Moral Computationβ8
This paper proposes Bounded Morality, a framework analyzing moral decision-making under computational constraints, defining moral breadth and depth tradeoffs fo...
Constructive Alignment: Governing Preference Dynamics inβ8
Constructive Alignment reframes AI alignment as governing dynamic human preference trajectories to ensure coherence, reflectively endorsed values, and resistanc...
Contrastive Reflection for Iterative Prompt Optimizationβ9
Contrastive Reflection improves agentic IR prompts by comparing failed vs. successful traces, boosting HotpotQA accuracy from 51.4% to 60.4%.
What Drives Interactive Improvement from Feedback?β9
Multi-turn AI agents often improve from retrying, not from feedback quality. Gains depend more on the studentβs ability to use feedback than on the teacherβs ex...
Recursive Self-Evolving Agents via Held-Out Selectionβ8
RSEA enables LLM agents to self-improve via recursive context evolution without weight updates, using held-out selection to prevent regression and ensure safe t...
When Does Personality Composition Matter for Multi-Agent LLM Teams?β9
Personality composition impacts multi-agent LLM teams differently by task: low agreeableness hinders open-ended tasks but has minimal effect on structured
