#robustness evaluation
Robustness Evaluation: 3 AI articles covering robustness evaluation news, analysis, and research
Articles
Counterexamples as Feedback for Agent Self-Correction⭐9
A-CEGIS framework uses counterexample feedback to evaluate multi-turn regex synthesis, achieving 90% task resolution within four turns and surpassing single-sho...
Unified Hallucination Fuzzing for Multimodal Large Language Models⭐10
Unified fuzzing framework reveals severe hallucination risks in MLLMs, exposing hidden performance degradation and alignment trade-offs.
Sensitivity Analysis of GRU, LSTM, and Transformer Encoder for⭐7
Compare GRU, LSTM, and Transformer encoders for classifying Level 2 automated driving systems, with robustness analysis under telematics data corruption.
