A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models

consensus-based evaluationinter-model alignmentlarge language modelsrelative intelligence index (rii)relative preference

Submitted on 19 Jul 2026


Authors: Mohtashim Khan


Abstract:


Traditional benchmarks for Large Language Models (LLMs) have primarily relied on static datasets and objective scoring metrics. However, these approaches often fail to capture meaningful differences in response quality when multiple answers are acceptable. In such scenarios, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness.


This paper introduces a consensus-based evaluation framework designed to measure relative preference among model-generated responses, rather than assessing absolute correctness. Instead of evaluating outputs against a fixed ground truth, we propose a method in which a panel of diverse LLMs ranks anonymized candidate responses to the same prompt. This approach treats aggregate inter-model agreement as a proxy for perceived response quality under blind conditions.


We conduct a controlled study involving five state-of-the-art LLMs across multiple domains, including programming, general knowledge, safety, logical reasoning, and mathematics. Each model generates responses and independently ranks peer outputs through a structured voting process. Scores are aggregated into a Relative Intelligence Index (RII), representing how frequently a given model's responses are preferred by other models in the panel.


Our findings reveal consistent preference patterns across domains, with certain models being more frequently ranked highly by their peers. However, we emphasize that these results reflect inter-model preference alignment rather than objective correctness or human judgment. This framework provides a scalable, model-driven method for comparative evaluation, offering an alternative perspective on response quality in scenarios where multiple valid answers exist. While not directly aligned with human evaluation, prior work suggests that aggregated model preferences can partially correlate with human judgments, motivating this approach as a proxy signal.


Comments: 14 pages, 7 figures


Subjects: Computation and Language (cs.CL)


MSC Classes: I.2


Cite as: arXiv:2607.21632 [cs.CL]


DOI: https://doi.org/10.48550/arXiv.2607.21632

via ArXiv CL

Related