Overview
Transformer-based models such as LegalBERT are increasingly deployed in legal decision-support systems, raising pressing questions about both fairness and the transparency of model explanations. In practice, these two properties are typically evaluated in isolation, leaving unresolved whether a debiasing intervention that alters fairness also alters how faithfully explanations reflect a model's underlying reasoning.
This study, published on arXiv (arXiv:2610.00045) in September 2026 under the category Computer Science > Computation and Language, directly investigates that relationship using the ECtHR alleged-violations corpus from LexGLUE.
Methodology
The research compares a LegalBERT baseline against a fairness-regularized variant. The intervention penalizes stereotypical warmth and competence representations during fine-tuning, drawing on a Structural Causal Model (SCM) framework. Evaluation spans three dimensions:
- Predictive performance
- Demographic fairness
- SHAP explanation faithfulness
All evaluations are conducted across five random seeds to ensure robustness.
Key Findings
At the performance-optimal regularization strength, the debiasing intervention does not reduce demographic disparity. This null result is consistent across two distinct fairness definitions and a conventional word-pair control on the gender axis. Classification performance remains largely unchanged.
However, the intervention consistently degrades explanation sufficiency across all five seeds and three thresholds. A shuffled-pair control reproduces this degradation while leaving performance and fairness unchangedβindicating that the effect stems from contrastive representational regularization rather than from the specific warmth-and-competence structure.
Implications
These results demonstrate a dissociation between fairness and explanation faithfulness: changes in explanation behavior do not necessarily signal changes in fairness. Consequently, fairness must be evaluated directly rather than inferred from explanation metrics.
Publication Details
- Authors: Yasmina El Kacemi, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
- Subjects: Computation and Language (cs.CL)
- Cite as: arXiv:2610.00045 [cs.CL]
- Submitted: 3 September 2026
- DOI: https://doi.org/10.48550/arXiv.2610.00045
via ArXiv CL+LG
