SCM-Based Fairness and Faithful Explainability for Legal

Overview


Transformer-based models such as LegalBERT are increasingly deployed in legal decision-support systems, raising pressing questions about both fairness and the transparency of model explanations. In practice, these two properties are typically evaluated in isolation, leaving unresolved whether a debiasing intervention that alters fairness also alters how faithfully explanations reflect a model's underlying reasoning.


This study, published on arXiv (arXiv:2610.00045) in September 2026 under the category Computer Science > Computation and Language, directly investigates that relationship using the ECtHR alleged-violations corpus from LexGLUE.


Methodology


The research compares a LegalBERT baseline against a fairness-regularized variant. The intervention penalizes stereotypical warmth and competence representations during fine-tuning, drawing on a Structural Causal Model (SCM) framework. Evaluation spans three dimensions:


  • Predictive performance
  • Demographic fairness
  • SHAP explanation faithfulness

All evaluations are conducted across five random seeds to ensure robustness.


Key Findings


At the performance-optimal regularization strength, the debiasing intervention does not reduce demographic disparity. This null result is consistent across two distinct fairness definitions and a conventional word-pair control on the gender axis. Classification performance remains largely unchanged.


However, the intervention consistently degrades explanation sufficiency across all five seeds and three thresholds. A shuffled-pair control reproduces this degradation while leaving performance and fairness unchangedβ€”indicating that the effect stems from contrastive representational regularization rather than from the specific warmth-and-competence structure.


Implications


These results demonstrate a dissociation between fairness and explanation faithfulness: changes in explanation behavior do not necessarily signal changes in fairness. Consequently, fairness must be evaluated directly rather than inferred from explanation metrics.


Publication Details


  • Authors: Yasmina El Kacemi, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
  • Subjects: Computation and Language (cs.CL)
  • Cite as: arXiv:2610.00045 [cs.CL]
  • Submitted: 3 September 2026
  • DOI: https://doi.org/10.48550/arXiv.2610.00045

via ArXiv CL+LG

Related