Master Cognitive Science · École normale supérieure – PSL

Learning & Decision-Making

An overview of the science of value-based decision-making: how humans learn from experience and choose among options that differ in value, probability, risk, delay and social consequence.

M2 level 4 ECTS Taught in English 12 sessions Tuesdays 9:30–11:30 8 Sep – 15 Dec 2026

Evidence accumulating to a decision boundary:
the drift-diffusion model, session 3.

Overview

About the course

This course provides an overview of the science of value-based decision-making: how humans learn from experience and choose among options that differ in value, probability, risk, delay and social consequence. It approaches decision-making at two levels: the behavioural description of choice phenomena, and their explanation through formal and computational models.

Normative benchmarks from economics and statistics serve as points of comparison against which systematic behavioural deviations are characterised and then re-interpreted not automatically as errors, but often as adaptive or resource-rational responses to the structure of the environment. The emphasis throughout is on connecting behavioural phenomena to mechanistic, computational explanations, and on keeping the normative, descriptive and mechanistic levels of analysis clearly distinct.

LevelM2 · 4 ECTS
PrerequisiteNone
LanguageEnglish
EvaluationHomework + group talk
RoomThéodule Ribot

Teaching team

Who teaches the course

Dr Stefano Palminteri

Teacher

Reinforcement learning, value-based decision-making, and the behavioural study of learning in humans and machines.

Dr Maël Lebreton

Teacher

Decision under uncertainty, valuation, metacognition and confidence in human decision-making.

Maëva L’Hôtellier

Teaching assistant

Supports all sessions, homework and the final presentations.

Contact: learningdecisionmakingcourse@gmail.com. This generic address will redirect emails to our three personal accounts.

Calendar

2026–2027 programme

Classes are held on Tuesdays from 9:30 to 11:30, at 29 rue d’Ulm (Room Théodule Ribot). The course runs from 8 September to 15 December 2026. There is no class on 27 October (autumn break), 17 November, or 24 November (PSL Week).

DateSessionTeachers
Tue 8 Sep 2026Introduction to decision-makingSP · MLh (+ MLe)
Tue 15 Sep 2026Decision-making under uncertainty and delayMLe · MLh
Tue 22 Sep 2026Simple choices as dynamic processesMLe · MLh
Tue 29 Sep 2026Reinforcement Learning: introductionSP · MLh
Tue 6 Oct 2026Contextual decision-making and adaptive codingMLe · MLh
Tue 13 Oct 2026Reinforcement Learning: biases and meta-learningSP · MLh
Tue 20 Oct 2026MetacognitionMLe · MLh
Tue 27 Oct 2026No class (autumn break)
Tue 3 Nov 2026Social LearningMLe · MLh
Tue 10 Nov 2026Artificial decision-makingSP · MLh
Tue 17 Nov 2026No class
Tue 24 Nov 2026No class (PSL Week)
Tue 1 Dec 2026Applying learning and decision-making researchSP · MLh
Tue 8 Dec 2026Oral presentations (1 of 2)SP · MLe · MLh
Tue 15 Dec 2026Oral presentations (2 of 2)SP · MLe · MLh

SP Stefano PalminteriMLe Maël LebretonMLh Maëva L’Hôtellier

Syllabus

Classes content

This section outlines the content of each class. Some adjustments may still occur, but the broad structure will remain as described. Similar topics, treated in a more accessible form for a general audience, are covered in the teachers’ book Decision Making: A Very Short Introduction (Palminteri & Wyart, Oxford University Press). Click a session to expand.

1Introduction to decision-making
This first class introduces decision-making as a multidisciplinary field studied through normative (philosophy, mathematics, economics), descriptive (psychology, behavioural economics) and engineering/AI approaches. It surveys the tools and types of data used to study decisions, from formal models to behavioural measures, and introduces the theory-building frameworks that organise the course (aggregate versus mechanistic, and descriptive versus normative accounts) together with Marr’s three levels of analysis (computational, algorithmic, implementational), of which the course concentrates on the first two. It sets out the guiding aim: to describe choice behaviour carefully and then explain it with formal and computational models.
Marr’s levelsNormative vs descriptive
2Decision-making under uncertainty and delay
This class explores how people decide under risk and uncertainty, starting from historical foundations such as Pascal’s wager, the St Petersburg paradox and Bernoulli’s utility theory. It explains risk attitudes through utility functions and axiomatic expected-utility theory, and confronts them with anomalies such as mixed gambles and the fourfold pattern, which Prospect Theory accommodates through gains, losses and loss aversion, though it in turn faces heuristic explanations and problems of inconsistent risk elicitation. The class then extends the same logic from risk to time, introducing intertemporal choice and delay discounting: how future outcomes are devalued, the contrast between exponential and hyperbolic discounting, and the preference reversals and present bias that follow. It closes on bounded rationality and the methodological care required when eliciting preferences over both risk and delay.
Prospect TheoryDelay discountingIntertemporal choice
3Simple choices as dynamic processes
This class asks how a value-based decision is actually made once options have been valued. It introduces the idea of a common currency in which qualitatively different options are compared, and revisits loss aversion as a property of how gains and losses enter this comparison. It then focuses on how choices unfold over time through sequential-sampling models, and in particular the drift-diffusion model: originally developed for perceptual decisions and extended to value-based choice, where noisy evidence (i.e., the momentary difference in value between options) is accumulated to a threshold, jointly accounting for choices and response times (including attention-weighted variants). This mechanistic account raises the question of whether “simple choices” are truly simple.
Common currencyDrift-diffusion model
4Reinforcement Learning: introduction
This class introduces learning from experience, beginning with the description-experience gap: the observation that choices among explicitly described options differ systematically from choices among options learned through sampling and feedback, which motivates a distinct treatment of experience-based decision-making and connects back to the sessions on risk. It then develops reinforcement learning from its behavioural roots (Pavlovian and operant conditioning in the work of Pavlov, Thorndike and Skinner) to the Rescorla-Wagner model, in which learning is driven by surprise, the discrepancy between expected and obtained outcomes. Post-cognitive-revolution formalisations follow, with Q-learning capturing state and action values and decision rules such as ε-greedy and softmax formalising the exploration–exploitation trade-off.
Description-experience gapRescorla-WagnerQ-learning
5Contextual decision-making and adaptive coding
This class examines how context shapes valuation, challenging the assumption of stable, context-invariant preferences. Context is analysed through frames (gain versus loss, the endowment effect), reference points (relative rather than absolute valuation), ranges (range-frequency effects and value normalisation) and menus (asymmetric dominance and decoy effects that violate the independence of irrelevant alternatives). These effects are captured computationally by mechanisms such as divisive normalisation and more generally by efficient coding, the principle that limited representational resources are allocated to match the statistics of the environment, so that adaptation and normalisation appear as near-optimal responses to coding constraints rather than failures of rationality. The class closes on open questions, such as whether distractor effects are better captured by divisive normalisation or by alternative models of value comparison.
Efficient codingDivisive normalisationDecoy effects
6Reinforcement Learning: biases and meta-learning
This class turns to the biases and refinements of reinforcement learning. It first asks whether learning is absolute or relative, presenting evidence from psychology, cross-cultural work and animal studies that valuation is context-dependent, and weighing adaptive against alternative explanations. It then examines valence-induced biases (the tendency to update more from positive than from negative outcomes) and their links to risk-seeking, status-quo bias and overconfidence, including the debate over whether they reflect a genuine asymmetry or mere choice perseveration. A central theme is the modulation of learning rates: how the speed of updating is adjusted asymmetrically for good versus bad news, and dynamically as a function of uncertainty and volatility. The class then introduces the distinction between model-free and model-based control, as two ways in which past experience can guide choice.
Valence-induced biasLearning-rate modulationModel-based control
7Metacognition
This class introduces metacognition as the ability to evaluate one’s own decisions, often captured through confidence judgements that track task difficulty and the probability of being correct. Behavioural and animal studies (such as opt-out tasks in monkeys and rats) reveal signatures of confidence, including its correlation with accuracy, the X-pattern across correct and incorrect trials, and sharper accuracy at high confidence. Models such as the evidence-distance model explain these signatures, while signal detection theory separates type-1 performance (decision accuracy) from type-2 performance (metacognitive accuracy) using measures such as d’ and meta-d’. The class then considers whether a central metacognitive process governs confidence across tasks, and closes on motivational, affective and computational biases through which incentives and emotions systematically shift confidence, sometimes producing overconfidence.
ConfidenceSignal detection theorymeta-d’
8Social Learning
This class examines how decisions are shaped by other people, emphasising social preferences (generosity, empathy, inequality aversion) that enter the utility function alongside personal payoffs. Experimental games such as the Ultimatum and Dictator games reveal the importance of social context, player identity and fairness. Social strategic interaction is formalised through game theory (Nash equilibria, coordination and anti-coordination games) though humans often deviate from equilibrium predictions because of bounded rationality and imperfect beliefs. Behavioural game theory and repeated-game experiments show that learning and experience shape social decision-making, while social preferences themselves vary across populations and contexts. Controversies remain around empirical validity, generalisability and the mechanisms driving observed social behaviour.
Social preferencesGame theoryFairness
9Artificial decision-making
This class applies the behavioural and computational lens of the course to artificial agents. It opens with the prospect of delegating decisions to machines, and the specific situation of deciding with, or against, a machine: how people take or discount algorithmic advice, and when they show algorithm aversion or appreciation. It distinguishes types of artificial system (from explicit, “explainable” expert systems to opaque connectionist networks trained by example) and asks how their choices can be studied with the same paradigms used for humans, a “machine psychology” of learning agents and large language models. Because such models are trained on human-generated data, they can reproduce human biases while being perceived as objective. The class examines this alignment problem, together with the opacity and brittleness of current systems, using the reinforcement-learning framework of earlier sessions as the bridge between how humans and machines learn from experience. It closes on what is gained, and what is risked, when decisions are handed to systems whose objectives and constraints need not match our own.
Algorithm aversionMachine psychologyAlignment
10Applying learning and decision-making research
This final lecture asks how the science of decision-making can be used to improve decisions, contrasting three traditions of behaviour change. The behaviourist tradition, following Skinner, changes behaviour through reinforcement and the arrangement of contingencies; nudges instead reshape the choice architecture to steer behaviour while preserving freedom of choice (libertarian paternalism); and boosts aim to build people’s own competences, such as risk literacy, that generalise and last. Setting these side by side clarifies what each assumes about the causes of behaviour and where each is likely to succeed or fail, along with their differing costs, durability and ethical exposure. The lecture then turns to computational psychiatry, where the learning and decision models developed across the course are used to characterise clinical conditions (delusions as impaired Bayesian inference, addiction as reinforcement learning and temporal discounting gone awry) and, through a transdiagnostic, dimensional view, to treat parameters such as learning rate and reward sensitivity as continuous markers that cut across diagnostic categories. It closes on the promise and the limits of translating decision research into health and policy.
Nudges & boostsSkinnerComputational psychiatry
11–12Oral presentations
The course closes with two sessions of group presentations. Each group presents papers of its choice connected to a course topic, and is assessed on its understanding, its critical analysis and its ability to communicate the key concepts of the course.
Group workEvaluation

Practical information

Practical information

Location

29 rue d’Ulm, Room Théodule Ribot (ground floor, head past the large staircase; the room is located in the hallway with wooden tables).

Schedule

Classes are held on Tuesdays from 9:30 to 11:30. There is no class on 27 October (autumn break), 17 November, or 24 November (PSL Week); the course runs from 8 September to 15 December 2026.

Inscription to the class

To register, please send us an email at learningdecisionmakingcourse@gmail.com.

Evaluation

Assessment is based on a weighted combination of weekly homework assignments and a final group presentation. Further details on grading criteria and expectations will be provided in the first class.