TRACE Bench: Task-Driven Roleplay Agentic Checklist Evaluation
Abstract
Roleplay evaluation should go beyond assigning a single scoreβit must identify which role requirements were tested, where they failed, and the dialogue evidence supporting each judgment. We introduce TRACE Bench, a task-driven, agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then employs a User Agent to converse naturally with the target roleplay model while privately updating checklist states based on model responses. Scores are thus traceable to specific checklist items and supporting dialogue turns, rather than stemming from a black-box holistic impression. For coverage cross-validation, we audit the released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench achieves 99.91% coverage in fewer turns. Robustness experiments demonstrate stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings alongside capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so subsequent evaluations can more reliably elicit and examine observed failure modes.
1. Introduction
As roleplay models continue to advance in 2026, the need for granular, interpretable evaluation has become critical. Traditional scoring methods often mask specific weaknesses, offering little guidance for improvement. TRACE Bench addresses this gap by providing a transparent, task-driven framework that links each evaluation score to concrete checklist items and dialogue evidence.
2. Framework Overview
TRACE Bench operates in three key stages:
- Offline Profiling: Each role profile is decomposed into a fixed checklist of requirements, ensuring comprehensive coverage.
- Agentic Interaction: A User Agent engages the target model in natural conversations while privately tracking checklist states, eliminating user bias.
- Traceable Scoring: Final scores are generated from the checklist states, with each score traced back to specific dialogue turns.
3. Coverage Validation
To validate coverage, we compared TRACE Bench against the MiniMax Role-play Benchmark's M2 free-dialogue transcripts. Results show that while free-dialogue methods cover only 73.74% of key role-profile points, TRACE Bench achieves 99.91% coverage, even with fewer turns. This demonstrates the efficiency and thoroughness of the checklist-driven approach.
4. Robustness and Stability
We conducted experiments to ensure reliability:
- Repeated Runs: Rankings remained stable across multiple executions.
- User Agent Replacement: Substituting different User Agents did not significantly alter rankings, confirming framework robustness.
5. Model Evaluation and Outputs
Across 26 diverse models, TRACE Bench provides overall rankings, capability breakdowns, and detailed checklist traces. This multi-faceted output allows researchers to identify specific areas for improvement, such as emotional expression or scenario handling.
6. Closed-Loop Benchmark Evolution
A novel feature of TRACE Bench is its support for Closed-Loop Benchmark Evolution. By analyzing failed traces, the framework distills verification methods that have proven effective, enabling future evaluations to more reliably elicit and examine failure modes. This self-improving mechanism ensures the benchmark remains relevant as roleplay models evolve.
7. Conclusion
TRACE Bench advances roleplay evaluation from simplistic scoring to a detailed, evidence-based approach. Its high coverage, stability, and traceability make it a valuable tool for researchers and developers in 2026 and beyond. The framework's closed-loop evolution ensures continuous improvement, aligning with the fast-paced development of roleplay AI.
via ArXiv CL+LG
