A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT:

A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID


Authors: Paweล‚ Blicharz, Miล‚osz Grunwald

arXiv: 2609.30287 [cs.CL]

Submitted: 8 September 2026

Accepted to: EMNLP 2026 (Main Conference)

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

ACM Classes: I.2.7

DOI: https://doi.org/10.48550/arXiv.2609.30287


Abstract


AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers ร— 768), which we call neurons. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full-feature detection accuracy. Bidirectional activation patching confirms this set's causal relevance: in both directions it flips predictions an order of magnitude more often than size-matched random sets. Mean-ablating the same neurons leaves accuracy largely intact; the signal is therefore redundantly distributed. Cross-generator analysis reveals a bipartite structure: instruction-tuned generators concentrate 30โ€“36% of stable neurons in BERT's final layer while both base generators fall below 14%, consistent with a layer-12 footprint of post-training alignment. Leave-one-family-out evaluation shows the selected neurons retain 86โ€“94% of the full-feature ceiling on unseen generator families, so a detector can operate on a small fixed subspace without re-identifying neurons per generator.


Key Contributions


  • Sparse neuron identification: Applies the L1-to-L2 sparse-probing protocol to frozen BERT-base-uncased, isolating a stable set of fewer than 1% of neurons per generator that preserves most full-feature detection accuracy.
  • Causal validation via activation patching: Bidirectional activation patching demonstrates that the identified neurons are causally relevant, flipping predictions far more often than size-matched random neuron sets.
  • Redundancy of detection signal: Mean-ablation leaves accuracy largely intact, indicating that the detection signal is redundantly distributed across the network rather than localized to a few critical units.
  • Architectural footprint of alignment: Instruction-tuned generators concentrate 30โ€“36% of stable neurons in BERT's final layer, compared to below 14% for base generators โ€” a pattern consistent with a layer-12 signature of post-training alignment.
  • Cross-generator generalization: Leave-one-family-out evaluation retains 86โ€“94% of the full-feature ceiling on unseen generator families, enabling a fixed-subspace detector that avoids re-identifying neurons per generator.

Significance


As AI-text detection becomes increasingly critical in 2026 for content provenance, academic integrity, and platform trust-and-safety pipelines, understanding why detectors make their predictions is as important as their headline accuracy. This work bridges mechanistic interpretability and practical detection by showing that a small, stable, causally validated neuron subspace in a frozen BERT encoder suffices for robust AI-text detection โ€” and that this subspace carries a measurable imprint of whether a generator was instruction-tuned.


Citation


@inproceedings{blicharz2026mechanistic,
  title     = {A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID},
  author    = {Blicharz, Pawe{\l} and Grunwald, Mi{\l}osz},
  booktitle = {Proceedings of EMNLP 2026 (Main Conference)},
  year      = {2026},
  eprint    = {2609.30287},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  doi       = {10.48550/arXiv.2609.30287}
}

via ArXiv CL+LG

Related