Investigating the Impact of the Language of Reasoning

Michaelmas 2026 project

Research question

When a large language model "thinks" before it answers, does the language it thinks in change how well it reasons? Modern reasoning models produce long chains of thought before committing to an answer, and these chains are overwhelmingly in English regardless of the input language. This project investigates how the language of the reasoning trace affects LLM capability across tasks: mathematical and logical reasoning, factual recall, and culturally grounded judgement. The core question is whether reasoning language is merely a surface choice or a genuine constraint on what the model can compute.

Ej

Project outline

The project has three phases.

Phase 1
Controlled reasoning-language intervention. We take a fixed set of benchmarks (e.g. MGSM, multilingual variants of GPQA and MMLU, and culturally localised QA) and hold the question and answer language constant while forcing the model to reason in a target language via prompting and, where possible, constrained decoding. We compare high-, mid-, and low-resource reasoning languages, plus mixed-language (code-switched) traces, across several open-weight reasoning models.

Phase 2
Where does the gap come from? Performance differences alone don't tell us why. We analyse reasoning traces for length, error type, and self-correction frequency, and use lightweight interpretability probes to test whether non-English reasoning is computed in that language or is English reasoning "translated on the fly" in the residual stream. We also measure calibration: does the model's confidence track correctness equally well when reasoning in different languages?

Phase 3
Can we close the gap? Depending on Phase 2 findings, we test targeted interventions: small-scale supervised fine-tuning on non-English reasoning traces, translation-then-reason pipelines, or prompting the model to reason in the language where it is strongest for a given task type. The output is a set of practical recommendations plus a public evaluation harness.

Impact case

Reasoning models are being deployed globally, but their reasoning is effectively monolingual. If reasoning in a user's own language degrades capability, then non-English speakers are silently receiving a weaker model, and any safety or oversight mechanism that reads chains of thought is reading a language the user may not understand. Conversely, if models reason better in English even on culturally specific questions, that tells us something important about what these traces actually represent. Either way, the findings inform how labs should train and evaluate reasoning models for global deployment, and speak directly to interpretability and oversight of chain-of-thought.

Ideal candidate

Prior publication is not needed but a plus, curiosity about why models behave as they do is essential. Speaking a language beyond English is helpful for sanity-checking traces but not required.

  • Comfortable with Python and running/evaluating open-weight LLMs on a GPU

  • Some exposure to LLM evaluation or benchmarking

  • Interest in multilinguality, reasoning, or interpretability

  • Able to commit roughly 8–10 hours per week for the project cycle