Abstract
Dependency-distance estimates derived from a single corpus are routinely treated as properties of a language, although this assumption has not been validated across independently compiled corpora. We compared mean dependency-distance estimates across 38 same-language treebank pairs from Universal Dependencies v2.18, employing concordance correlation, Bland-Altman analysis, and a twelve-specification multiverse design. Cross-treebank agreement was moderate at best: substituting one treebank for another reversed nearly 40 percent of pairwise language orderings, and treebank choice accounted for roughly 29 percent of between-group variance. This disagreement substantially exceeded within-treebank sampling error and persisted across all twelve preprocessing specifications. Nevertheless, every treebank confirmed dependency-length minimization (normalized ratio below 1). The data suggest that mean dependency distance (MDD) is better understood as a corpus-conditioned composite of grammatical, register, and annotation factors rather than a stable language-level parameter. The qualitative dependency-length minimization universal survives corpus substitution, but the ordinal cross-linguistic ranking does not.
1. Introduction
In computational linguistics and typology, mean dependency distance (MDD) has become a standard metric for assessing syntactic complexity and cognitive processing load. Researchers often derive MDD estimates from a single treebank and then treat these estimates as inherent properties of a language, using them to compare languages or test theoretical claims. However, this practice assumes that a particular corpus is representative of the language as a wholeβan assumption that has not been systematically tested.
As our field moves toward larger and more diverse corpora, understanding how corpus choice influences quantitative linguistic estimates is increasingly critical. In 2026, with Universal Dependencies (UD) v2.18 offering over 200 treebanks in more than 100 languages, researchers have unprecedented resources, but also a greater need to account for variability across treebanks that purportedly represent the same language.
This paper examines whether MDD estimates are stable across different treebanks of the same language. If MDD varies substantially depending on the corpus, then conclusions based on a single treebank may not generalize. Conversely, if MDD is robust across corpora, we can be more confident in cross-linguistic comparisons.
2. Methodology
2.1 Data Source
We used treebanks from the Universal Dependencies v2.18 release (UD v2.18), which provides uniformly annotated syntactic dependency structures across many languages. For each language with at least two independently compiled treebanks, we considered all possible pairs, resulting in 38 same-language treebank pairs.
2.2 Metrics and Analysis
For each treebank, we computed the mean dependency distance (MDD) across all dependency links, following standard practices. To compare treebanks within each pair, we used three complementary approaches:
- Concordance correlation coefficient: measures the agreement between two sets of MDD values, accounting for both precision and accuracy.
- Bland-Altman analysis: quantifies the bias and limits of agreement between paired measurements, revealing systematic differences.
- Multiverse design: evaluates robustness by testing twelve preprocessing specifications, varying tokenization, sentence segmentation, and filter criteria, to ensure that findings are not artifacts of a single pipeline.
Additionally, we estimated within-treebank sampling error using bootstrap resampling to set a baseline for expected variability.
3. Results
3.1 Cross-Treebank Agreement
Agreement between treebank pairs was moderate at best. Concordance correlations were generally below 0.6, and Bland-Altman plots showed wide limits of agreement, indicating that MDD values can differ substantially across treebanks of the same language. The variability was not random: treebank choice accounted for approximately 29% of the variance in MDD when comparing different language groups, meaning that a portion of what might appear to be cross-linguistic differences could be an artifact of corpus selection.
3.2 Direction Reversals in Language Rankings
A striking finding is that substituting one treebank for another reversed the ordering of languages in pairwise comparisons nearly 40% of the time. This means that, in a substantial fraction of cases, a language would appear to have a higher MDD than another language when using one treebank, but the opposite would appear with a different treebank. Such reversals severely undermine the reliability of ordinal cross-linguistic rankings based on MDD.
3.3 Robustness of Dependency-Length Minimization
Despite these inconsistencies, every treebank in our sample confirmed the dependency-length minimization (DLM) pattern, with a normalized MDD ratio below 1. This suggests that DLM is a robust tendency across languages and corpora, consistent with universal cognitive constraints.
4. Discussion
Our results challenge the view of MDD as a stable, language-level parameter. Instead, MDD behaves as a corpus-conditioned composite that reflects a mix of grammatical conventions, register differences, and annotation choices. This does not mean that MDD is useless, but that researchers must be cautious when interpreting quantitative estimates from a single corpus.
4.1 Methodological Implications
Practitioners should:
- Report MDD estimates across multiple treebanks when available, or at least acknowledge the potential for cross-treebank variability.
- Use multiverse-style analyses to test the robustness of their conclusions to preprocessing choices.
- Consider register and annotation guidelines as factors that can influence MDD, not just the underlying language.
4.2 Theoretical Implications
The robustness of the DLM universal suggests that qualitative typological generalizations are more likely to hold, but ordering claims require more evidence. Our findings underscore the need for corpus-based typology to incorporate corpus variability into its models.
5. Conclusion
In summary, corpus choice significantly changes dependency-distance estimates, affecting both cross-linguistic comparisons and theoretical conclusions. While the qualitative DLM pattern is invariant, the precise values and ordinal rankings are not. The field should move toward more robust methods that treat corpora as sampled proxies rather than definitive sources of linguistic properties.
via ArXiv CL+LG
