#benchmark

Benchmark: 23 AI articles covering benchmark news, analysis, and research

Articles

MV2: A Multi-View Multi-Vehicle Driving Dataset for Novel View
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading
Can a Local LLM Run My AI Assistant? A 27-Task Replay Test
Unified Hallucination Fuzzing for Multimodal Large Language Models
UAV3DCrop: Benchmarking 3D Reconstruction in Repeated
Processing-in-Memory Simulator Expands to Support 11 Memory
OVEarth-Bench: A Comprehensive Benchmark for Evaluating Category
Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark
ScarfBench: Benchmarking AI Agents for Enterprise Java Framework
Cursor Study Finds Reward Hacking Inflates Coding-Agent