OpenPrint 20260913.0001v1EmpiricalReleased: September 13, 20263 Views

Correct Definitions, Failed Applications: Accessibility Concepts in Language Models

Trisha Salas
Read Full Paper (PDF)
TMLR
Reviewed by TMLR Agent
Refs Verified (21/23)
Claims Verified

Abstract

A language model can state what an accessibility concept means and still fail to apply it. We measure this declarative-evaluative gap in thirteen models from three families: Pythia, GPT-2, and OLMo 2. In the original behavioral batteries, the gap opens early and persists across families. At Pythia-12B, it closes because declarative accuracy falls. Three concepts answered correctly at smaller scales regress to incorrect at 12B. Convergence by decay is not mastery. In a frozen same-concept validation of eight concepts, each represented by one violation and one conformant item, none of 96 confirmatory model-concept cells passes both application items. Eleven cells pass one item; a development-exposed pilot, reported separately, passes one pair. The 96 confirmatory cells include 35 with correct definitions. We test two possible explanations for the gap. The global maximum of attention between a compound's constituents has weak associations with accuracy that differ in direction across families. When the measure is restricted to late-layer, value-weighted binding, a consistent pattern appears: rarer compounds receive stronger distributed value-weighted writes at all thirteen model-scale points (p = 0.0003). This is the opposite of what we would expect if stronger binding indicated knowledge. The pattern is consistent with processing difficulty, although the experiments do not identify what computation the stronger writes perform. Corpus frequency is the strongest measured predictor of accuracy across 49 compounds (Spearman ρ = 0.52–0.59 across families), but it leaves most compound-to-compound variation unexplained. The same concept can be available to definition and unreachable in application, even though its corpus frequency has not changed. Frequency predicts availability. Task form constrains reachability.

Keywords

accessibilitylanguage modelsdeclarative-evaluative gapcorpus frequencyattention bindingvalue-weighted writes

Illustration

Illustration 1/8

Citation

@article{Salas2026Correct,
  title={Correct Definitions, Failed Applications: Accessibility Concepts in Language Models},
  author={Trisha Salas},
  year={2026},
  url={https://cspaper.org/openprint/20260913.0001v1},
  journal={OpenPrint:20260913.0001v1}
}

Version History

VersionReleased DateSubmitter
v1Current
Sep 13, 2026
Trisha Salas
Correct Definitions, Failed Applications: Accessibility Concepts in Language Models | OpenPrint 20260913.0001v1 — CSPaper