Language Model Interpretability & In-Context Learning
Cross-Linguistic Robustness of Interpretability and Cognitive Alignment in Language Models
Generated automatically from the limitations stated in 9 papers (ACL, EMNLP, NAACL), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Mechanistic interpretability claims and brain-to-LM alignment results currently rest almost entirely on English-only stimuli, English-trained models, and English-speaking participant data. Because cross-linguistic evaluation has not been conducted across these methods, it remains unknown whether identified circuits, induction heads, and representational alignments reflect general linguistic processing or are artifacts of English syntax and data abundance. Consequently, researchers cannot rely on existing interpretability mechanisms when analyzing multilingual models or studying cognitive language processing across typologically diverse languages.
Why it matters
Validates whether mechanistic interpretability and cognitive alignment findings generalize across language families, providing a baseline for multilingual model auditing and cross-lingual cognitive modeling.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Replicate standard in-context learning and circuit-probing pipelines (e.g., induction head tracking, indirect object identification) across multilingual language models (such as mGPT or BLOOM) on parallel multi-language datasets covering diverse syntactic typologies (SOV, VSO, agglutinative), measuring circuit consistency and layer localization across languages.
- 2
Evaluate brain-score and electrophysiological response alignment (e.g., N400 and fMRI prediction) using existing multilingual cognitive datasets (such as ZuCo or multilingual 'The Little Prince' neuroimaging corpora), measuring regression accuracy and alignment degradation when testing non-English native readers against corresponding multilingual model activations.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If cross-lingual stimulus differences introduce overwhelming confounds in tokenization and subword segmentation that obscure underlying representation shifts, preventing clear attribution of interpretability failures.
Evidence
Each paper's own statement of the limitation, verbatim.
- Cross-Modal Taxonomic Generalization in (Vision-) Language ModelsACL 2026
Modest model scales (0.6B–1.7B, spot-check with 8B) and English-only stimuli; scaling and cross-lingual behavior unexplored
- From Language to Cognition: How LLMs Outgrow the Human Language NetworkEMNLP 2025
All neural data come from English stimuli; cross-linguistic generalization of the findings is not tested
- Are explicit belief representations necessary? A comparison between Large Language Models and Bayesian probabilistic modelsNAACL 2025
Stimuli, human data, and models are English-only; cross-linguistic projection patterns are untested
- The effects of distance on NPI illusive effects in BERTEMNLP 2024
Only English stimuli were tested; results may not generalize to other languages.
- Speech language models lack important brain-relevant semanticsACL 2024
Only English participants, stimuli, and mostly English-trained models; findings may not generalize to other languages.
- Measuring Meaning Composition in the Human Brain with Composition Scores from Large Language ModelsACL 2024
Only English stimuli ('The Little Prince') were used; multilingual validation is absent
- Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but InconsistentlyEMNLP 2024
English-only stimuli; no cross-linguistic evaluation.
- When Language Models Fall in Love: Animacy Processing in Transformer Language ModelsEMNLP 2023
No new human data was collected: the Dutch stimuli were machine-translated to English and human N400 responses were assumed (not verified) to transfer across languages; humans were never tested on the low-context stimuli
Show all 9 papers
- Prompting Language Models for Linguistic StructureACL 2023
Experiments are English-only; generalization to other languages is untested despite prior work showing PLMs prompt much worse outside English
Nearest existing work
- Translation Heads: Disentangling meaning from language in LLM-based machine translationICML 2026
- Brain-Informed Fine-Tuning for Improved Multilingual Understanding in Language ModelsNeurIPS 2025
- A Shared Geometry of Difficulty in Multilingual Language ModelsACL 2026
- The Same but Different: Structural Similarities and Differences in Multilingual Language ModelingICLR 2025
- Multilingual Language Models Predict Human Reading BehaviorNAACL 2021
- Explainability and Interpretability of Multilingual Large Language Models: A SurveyEMNLP 2025
- Revealing the Parallel Multilingual Learning within Large Language ModelsEMNLP 2024
- Getting More from Less: Large Language Models are Good Spontaneous Multilingual LearnersEMNLP 2024
- Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in TransformersACL 2025
- Disentangling Language and Culture for Evaluating Multilingual Large Language ModelsACL 2025
- Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language ModelsACL 2024
- Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language ModelsACL 2025
- Scaling in Cognitive Modelling: a Multilingual Approach to Human Reading TimesACL 2023
- Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language ModelsACL 2025
- LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?ICLR 2026
Related open problems
Language Model Interpretability & In-Context Learning
Benchmarking Whether LLM-Based Methods Survive a Change of Model: A Transfer Study Across Access Levels and Model Families
Methods built on one LLM — whether they depend on its logits, its instruction-following behavior, its pretraining data, or its annotation quality — are currently reported on one or two models and simply assumed to generalize. Because the assumptions are implicit, a follow-up researcher cannot tell whether a failure on a new model is a flaw of the method or an unexamined dependency on the original model's scale, data, or API surface. This blocks cumulative progress: every new model release forces methods to be silently re-validated or abandoned, and conclusions attributed to "LLMs" are really conclusions about a single backbone.
Language Model Interpretability & In-Context Learning
A benchmark of naturally occurring hallucinations with model-generated ground truth, replacing synthetic entity-swap surrogates for detection evaluation
Today, hallucination detection methods are trained and evaluated against synthetic artifacts — entity-swapped summaries, outputs from a different model than the one under test — rather than hallucinations that actually arise during generation. This means reported detection numbers do not measure whether a method works on the failure mode that matters, and the modest predictive results (e.g., R² ≈ 0.27 on hallucination scores) may partly reflect this distribution mismatch. Detection validated only on 7B–13B open-weight models with synthetic data cannot be trusted as evidence about frontier-model behavior. Without a natural-hallucination benchmark, every detection paper inherits the same unquantified gap between "detects injected errors" and "detects real hallucinations."
Language Model Interpretability & In-Context Learning
Benchmarking Interpretability and In-Context Learning Mechanisms Beyond Toy Synthetic PCFGs
Mechanistic interpretability and in-context learning findings in formal linguistic settings currently rely on tiny, unambiguous PCFGs with vocabularies of fewer than 100 words and corpora of only ~20K sentences evaluated on toy models like nanoGPT. It remains completely unknown whether the specific internal circuits, parsing behaviors, and induction mechanisms documented in these toy studies survive when evaluated on richer formal grammars, ambiguous grammars, larger vocabularies, or natural syntax. Without systematic robustness evaluations across these broader settings, interpretability claims cannot be reliably extrapolated beyond the micro-benchmarks on which they were discovered.
Language Model Interpretability & In-Context Learning
Testing the Transferability of Toy-Model Theoretical Guarantees to Deep Transformer Architectures
Theoretical analyses of in-context learning and mechanistic interpretability rely almost entirely on heavily idealized architectures, such as single-layer or two-layer transformers with linear attention, orthogonal embeddings, and no residual connections or normalization. Because these proofs make structural assumptions that diverge substantially from production models, it is currently unknown whether their theoretical guarantees hold in real large language models or function purely as loose analogies. Without empirical validation across the gap between toy settings and deep models, downstream interpretability methods risk relying on mechanistic claims that fail in standard architectures.