Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks
Authors: Ewelina Gajewska, Katarzyna Budzynska, Jaroslaw Chudziak
Venue: Accepted to COMMA 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.28673 [cs.CL]
DOI: https://doi.org/10.48550/arXiv.2609.28673 (pending registration)
Submitted: 23 September 2026
Overview
As large language models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues—from political debate simulations to automated negotiation and deliberation systems—there is a growing need for rigorous evaluation of their debating competence relative to human interlocutors. This study addresses that gap by focusing on character attacks (ad hominem arguments), a class of moves traditionally dismissed as fallacies yet central to political persuasive dialogue, where ethos frequently rivals propositional content in importance.
Research Question
The authors investigate whether modern LLMs can replicate human competence in strategically using and responding to character attacks. Rather than treating ad hominem as a binary fallacy to be avoided, the work examines it as a legitimate, normative move within ethos-centred political discourse.
Methodology
The study combines corpus analysis with empirical benchmarking:
- Corpus analysis of natural political dialogue — The authors analyse a corpus of natural-language political dialogues to identify the defensive strategies human interlocutors naturally employ in ethos-centred debates.
- Dialogue game formalisation — These strategies are structured into a dialogue game, providing a formal framework for evaluating argumentative moves.
- LLM benchmarking — LLM-generated dialogues are benchmarked against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, directly contrasting human debaters' repertoire of defensive strategies with those of artificial agents.
- Most LLMs rigidly prioritise logical defences, even when ethotic counterattacks would be contextually appropriate.
- LLMs largely fail to exploit ethotic counterattacks as valid moves in political discourse.
- The authors argue that current safety fine-tuning constrains the strategic action space of these LLMs, preventing them from fully engaging in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.
- arXiv: 2609.28673 [cs.CL]
- PDF: Available via arXiv
- DOI: 10.48550/arXiv.2609.28673
Key Findings
The results reveal a substantial divergence between human and machine behaviour:
Implications
This work highlights a tension between safety alignment and argumentative authenticity. As LLMs take on roles as debate partners, moderators, and persuasive agents in 2026 and beyond, the inability to engage with ethos-centred argumentation may limit their effectiveness in political, legal, and other adversarial dialogue settings. The findings point to the need for evaluation frameworks and fine-tuning approaches that distinguish between harmful personal attacks and contextually legitimate character contestation.
Availability
via ArXiv CL+LG
