HakemBench: A Turkish Benchmark of Typed Decisions
arXiv:2610.02293 [cs.CL] · Submitted 1 October 2026 · Computation and Language (cs.CL) · ACM Class: I.2.7
Author: Sait Furkan Teke (ufak AI)
Overview
HakemBench is a Turkish benchmark of typed decisions in which the model under test reads a text, a question, and a fixed set of options, then returns a probability for every option. This task formulation—grounded in probabilistic scoring rather than single-label classification—aligns with the broader 2026 shift toward calibration-aware evaluation, where models are judged not only on accuracy but on how well their confidence estimates track reality.
Release Details
Version 1.0 is released fully open under CC BY 4.0, comprising:
- 2,346 items
- 4,275 choice, yes/no, and score questions
- Seven tracks: fact-check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support
Evaluation Harness
A single harness scores three complementary dimensions:
- Decision quality — macro F1
- Calibration — derived from the normalised Brier score
- Selective automation — derived from the normalised area under the generalised risk-coverage curve
- The leader scores a composite of 0.888
- The lab's own model ranks 7th at 0.660
- 9 pages including references
- Data, harness, scorer, and board:
- Hugging Face Dataset
- GitHub Repository
These three components are combined via a geometric mean, and intervals are reported from 2,000 bootstrap draws. Probes for option order, paraphrase, English translation, and substituted names are reported alongside the primary results.
Gold Label Provenance
Most gold labels come from blind passes of one AI model family compared against the votes of a panel of large language models from other model families. Importantly, these labels are not human-verified—a caveat that matters for interpreting results in the context of 2026 debates over synthetic versus human-annotated evaluation data.
Leaderboard Results
On a board of 16 rows:
Contamination Disclosure
The lab's own model numbers are not blind. Earlier runs' test results shaped its training data, so its guardrail, moderation, and customer support numbers are flagged. When every model is scored on the other four tracks only, its composite is 0.678 — 6th of 16. This transparent contamination disclosure reflects growing 2026 expectations around benchmark hygiene and training-data provenance reporting.
Resources
Citation
arXiv:2610.02293 [cs.CL]
DOI: https://doi.org/10.48550/arXiv.2610.02293 (pending registration)
via ArXiv CL+LG
