Abstract
When a language model encodes a sensitive attribute, the source matters: the signal may be carried by explicit or distributed input cues, or it may remain after controlling for a declared input representation. Standard probing conflates these cases, leaving auditors unsure whether to inspect the input pipeline or the learned representation. We propose Controlled Representation Attribution (CRA), a reference-conditional auditing protocol that residualizes each layer against an input reference and probes the residual. We use “model-added” conservatively to denote reference-residual sensitive information rather than a causal claim about model weights. CRA is supported by a recovery theorem—exact when the residual noise is equal-covariance Gaussian and the target direction is an eigenvector of that covariance, with an explicit finite-anisotropy bound otherwise—and by scope and reference-admissibility diagnostics that identify inadequate controls. In controlled simulations with known ground-truth directions, CRA wins the selectivity comparison in 60/60 runs; falsification tests further show that weak references can create spurious residual conclusions, while the admissibility gate withholds such claims. In a prospectively specified controlled-text audit, residual evidence survives name removal but not removal of broader lifestyle and stylistic cues, demonstrating dependence on the declared control. Archived pretrained-language-model analyses illustrate a lexical-gender/reference-residual-age contrast and a positive pooled medical-question-answering difference (22 of 27 cells, mean +5.3 percentage points), which is not powered at the dataset level; these analyses are exploratory rather than confirmatory. CRA is a source-attribution diagnostic that reports reference sensitivity and can abstain when controls are inadequate; it complements concept-erasure methods rather than replacing them.
Keywords
Illustration
Citation
@article{Quan2026Separating,
title={Separating Input-Carried from Model-Added Demographic Encoding in Language Models},
author={Minh K. Quan and Pubudu N. Pathirana},
year={2026},
url={https://cspaper.org/openprint/20260826.0001v1},
journal={OpenPrint:20260826.0001v1}
}Version History
| Version | Released Date | Submitter |
|---|---|---|
v1Current | Aug 26, 2026 | Minh Quan |
