PHIL 611
Measuring and Improving Reasoning in Humans and Machines
Graduate seminar · Fall 2026
The mechanisms that produce belief in us have largely been shaped by pressures that are not aimed at truth: natural selection, social approval, and institutional incentives. Language models are shaped by their own imperfect optimization pressures, including next-token prediction and feedback from human evaluators. This interdisciplinary seminar asks how those pressures affect reasoning in humans and machines, and what we can do to improve it.
We consider how to characterize cognitive errors, which interventions improve reasoning, and whether improvements transfer across tasks. A recurring question concerns measurement: can we devise tests of good reasoning where improving a score requires improving the underlying skill? We also examine sycophancy, confabulation, and deception in language models, and ask how reflection, training, and collaboration might produce more reliable epistemics.
The course brings philosophical analysis into conversation with empirical work in psychology and AI. Student projects move from concepts such as intellectual humility or probabilistic calibration to proposals for measurement and systematic improvement. Graduate students from other disciplines are welcome.
Topics
Course ongoing- further topics TBA
September 8 · What is good human reasoning?
Normative reasoning standards; the "rationality wars"; some candidate maps of skills, biases and heuristics.
- McHugh, Norms of Reasoning
- Baron, The Point of Normative Models
- Gigerenzer, Axiomatic Rationality and Ecological Rationality
- Dennett, True Believers: The Intentional Strategy and Why It Works
September 15 · What it takes for a machine to reason
When are folk-psychological attributions of beliefs, concepts, intentions, and reasoning literally true vs. merely predictively useful; what kind of evidence could distinguish them? When is the "reasoning" we see just a generated rationale?
- Goldstein and Levinstein, Does ChatGPT Have a Mind?
Sections 2.4, 3.1, and 3.3 optional
- Shanahan et al., Role Play with Large Language Models
- Goldstein and Lederman, What Does ChatGPT Want?
Focus on sections 2.1, 3, 5, and 6
September 22 · Goals, personas, constitutions
Do LLMs have goals, play roles, and/or have beliefs about characters in their dialogs? How do constitutions, pre-training, and post-training shape these dispositions, and which dispositions survive a change of persona/speaker/frame? How useful is a nested intentional stance picture?
- Marks et al., The Persona Selection Model (opens in a new tab)
- Hubinger et al., Conditioning Predictive Models (opens in a new tab)
- Kutasov et al., Teaching Claude Why (opens in a new tab)
- Li et al., Disentangling Intent from Role (opens in a new tab)Optional
- Wang et al., Persona Features Control Emergent Misalignment (opens in a new tab)Optional
- Betley et al., Weird Generalization and Inductive Back Doors (opens in a new tab)Optional
- Qi et al., Safety Alignment Should Be Made More Than Just a Few Tokens Deep (opens in a new tab)Optional
September 29 · How well do reasoners understand their own judgments?
Do reasoners' accounts of their own reasoning reflect the real explanations for their answers? People and LLMs both confabulate, but in different ways. What's the difference between introspective awareness and confident guessing about one's internal states? When a model "introspects", whose states are being reported- the LLM's or the Assistant's?
- Chen et al., Reasoning Models Don't Always Say What They Think (opens in a new tab)
- Lindsey, Signs of Introspection in Large Language Models (opens in a new tab)
- Lederman and Mahowald, Emergent Introspection in AI Is Content-Agnostic (opens in a new tab)
- Hall et al., Lifting the Veil of Morality (opens in a new tab)
- Comșa and Shanahan, Does It Make Sense to Speak of Introspection in Large Language Models? (opens in a new tab)Optional
- Boppana et al., Reasoning Theater (opens in a new tab)Optional
- Emmons et al., When Chain of Thought Is Necessary, Language Models Struggle to Evade Monitors (opens in a new tab)Optional
- Johansson et al., Failure to Detect Mismatches between Intention and Outcome in a Simple Decision TaskOptional