A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
Authors: Paweล Blicharz, Miลosz Grunwald
arXiv: 2609.30287 [cs.CL]
Submitted: 8 September 2026
Accepted to: EMNLP 2026 (Main Conference)
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
ACM Classes: I.2.7
DOI: https://doi.org/10.48550/arXiv.2609.30287
Abstract
AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers ร 768), which we call neurons. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full-feature detection accuracy. Bidirectional activation patching confirms this set's causal relevance: in both directions it flips predictions an order of magnitude more often than size-matched random sets. Mean-ablating the same neurons leaves accuracy largely intact; the signal is therefore redundantly distributed. Cross-generator analysis reveals a bipartite structure: instruction-tuned generators concentrate 30โ36% of stable neurons in BERT's final layer while both base generators fall below 14%, consistent with a layer-12 footprint of post-training alignment. Leave-one-family-out evaluation shows the selected neurons retain 86โ94% of the full-feature ceiling on unseen generator families, so a detector can operate on a small fixed subspace without re-identifying neurons per generator.
Key Contributions
- Sparse neuron identification: Applies the L1-to-L2 sparse-probing protocol to frozen BERT-base-uncased, isolating a stable set of fewer than 1% of neurons per generator that preserves most full-feature detection accuracy.
- Causal validation via activation patching: Bidirectional activation patching demonstrates that the identified neurons are causally relevant, flipping predictions far more often than size-matched random neuron sets.
- Redundancy of detection signal: Mean-ablation leaves accuracy largely intact, indicating that the detection signal is redundantly distributed across the network rather than localized to a few critical units.
- Architectural footprint of alignment: Instruction-tuned generators concentrate 30โ36% of stable neurons in BERT's final layer, compared to below 14% for base generators โ a pattern consistent with a layer-12 signature of post-training alignment.
- Cross-generator generalization: Leave-one-family-out evaluation retains 86โ94% of the full-feature ceiling on unseen generator families, enabling a fixed-subspace detector that avoids re-identifying neurons per generator.
Significance
As AI-text detection becomes increasingly critical in 2026 for content provenance, academic integrity, and platform trust-and-safety pipelines, understanding why detectors make their predictions is as important as their headline accuracy. This work bridges mechanistic interpretability and practical detection by showing that a small, stable, causally validated neuron subspace in a frozen BERT encoder suffices for robust AI-text detection โ and that this subspace carries a measurable imprint of whether a generator was instruction-tuned.
Citation
@inproceedings{blicharz2026mechanistic,
title = {A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID},
author = {Blicharz, Pawe{\l} and Grunwald, Mi{\l}osz},
booktitle = {Proceedings of EMNLP 2026 (Main Conference)},
year = {2026},
eprint = {2609.30287},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2609.30287}
}
via ArXiv CL+LG
