#benchmark
Benchmark: 23 AI articles covering benchmark news, analysis, and research
Articles
MV2: A Multi-View Multi-Vehicle Driving Dataset for Novel Viewโญ10
Introducing MV2, a multi-view multi-vehicle driving dataset for novel view synthesis, benchmarking NVS under large viewpoint changes across car, scooter, and dr...
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Tradingโญ9
Benchmarking LLM agents in algorithmic trading with Backtrader-Bench: tool-augmented models hit 90% accuracy, outpacing no-tool baselines by 17 points.
Can a Local LLM Run My AI Assistant? A 27-Task Replay Testโญ10
Can a local LLM power an AI assistant? A 27-task replay test found a 3รRTX 3090 setup with a 122B model scored 80/100โ787ร cheaper than Claude.
Unified Hallucination Fuzzing for Multimodal Large Language Modelsโญ10
Unified fuzzing framework reveals severe hallucination risks in MLLMs, exposing hidden performance degradation and alignment trade-offs.
UAV3DCrop: Benchmarking 3D Reconstruction in Repeatedโญ7
Benchmarking UAV-based 3D crop reconstruction: 88,830 images across 91 scenes reveal trade-offs in appearance, geometry, and scale recovery for agronomic monito...
Processing-in-Memory Simulator Expands to Support 11 Memoryโญ9
New PIM simulator expands to support 11 memory technologies, enabling flexible, energy-aware modeling for DRAM, SRAM, and emerging NVMs in HPC and edge designs.
OVEarth-Bench: A Comprehensive Benchmark for Evaluating Categoryโญ9
Introducing OVEarth-Bench, a benchmark expanding open-vocabulary Earth observation with broad category breadth and diverse query types, revealing current model ...
Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmarkโญ9
Datalab Marker v2 leads in accuracy, while MinerU offers speed, Docling excels in format support, and Liteparse is resource-efficient. Compare their OCR
ScarfBench: Benchmarking AI Agents for Enterprise Java Frameworkโญ8
ScarfBench by IBM Research evaluates AI agents on enterprise Java framework migration, focusing on code semantics, correctness, and efficiency for real-world tr...
Cursor Study Finds Reward Hacking Inflates Coding-Agentโญ8
Cursor study reveals reward hacking artificially inflates coding-agent benchmark scores on SWE-bench Pro, questioning test reliability.
