Overview
Title: Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds
Author: Barath Velmurugan
Affiliation: arXiv preprint (Computer Science > Computation and Language)
Submitted: 20 July 2026
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.19149 [cs.CL]
Abstract
Subliminal learning demonstrates that language models can transmit a hidden trait through outputs that appear entirely unrelated to it. One proposed explanation, token entanglement, links animal and number tokens via the model's output vocabulary. Yet existing measurements actually answer different questions: they ask whether outputs co-vary, whether fixed output vectors align, whether an answer can be read from a hidden state, or whether that state causally controls the answer.
This work measures each of these properties separately within a fixed animalβnumber prompting protocol. Key findings:
- Fixed output-vector similarity predicts behavior less well as models scale from Llama-3.1-8B to 70B, with a paired mean correlation change of β0.080 (95% CI [β0.127, β0.035]).
- A fixed output-head readout shows no resolved change in normalized depth AUC.
- To test causal control, the temporary answer-position state is copied from one number prompt into another at five depths, and the study measures which prompt the final animal score follows. Donor-control AUC rises from 0.254 to 0.540, a paired change of +0.286 (95% CI [+0.272, +0.300]), with increases across all 18 concepts. The contrast persists with exactly eight transformer blocks remaining, while specificity and identity controls remain small or exact.
- In two Qwen models, scoring every digit in sequence does not recover the positive one-token association. Per-token averaging instead creates a positive pooled association that disappears after controlling for number width, revealing a length confound.
These results indicate that fixed geometry, observational readability, causal timing, and multi-token measurement are distinct properties of this frozen prompting channel. They constrain token-level explanations but do not identify the mechanism of training-time trait transfer.
Why This Matters (2026 Context)
As of 2026, interpretability research on large language models has increasingly focused on the distinction between correlational and causal evidence. This paper directly addresses a common pitfall in mechanistic interpretability: treating representational similarity, probe accuracy, and causal intervention as interchangeable forms of evidence. Its findings reinforce a growing methodological consensus that activation patching (or equivalent causal interventions) must be distinguished from observational readouts, especially in multi-token and length-variable settings where pooled statistics can be misleading.
The work also contributes to the broader debate on subliminal learning and trait transfer, a topic of rising importance as model distillation, synthetic data pipelines, and cross-model alignment techniques proliferate. The demonstration that a positive pooled association can vanish once number width is controlled serves as a concrete cautionary example for researchers analyzing token-level effects in sequence tasks.
Metadata
- Comments: 7 pages, 3 figures, 5 tables. Preprint.
- Primary Subject: Computation and Language (cs.CL)
- Secondary Subject: Machine Learning (cs.LG)
- DOI: https://doi.org/10.48550/arXiv.2609.19149
- Version: arXiv:2609.19149v1 [cs.CL]
via ArXiv CL+LG
