Silent information loss
A dropped column, an unrecorded exclusion or a filtered timeout leaves no trace in the processed file. The file looks complete, so a later analyst cannot see what is missing.
A project of ISTC-CNR
Planning 101 brings published experiments on human planning into a single, validated data model. Each task is described as a Markov decision process and each move is recorded as a Step, so an analysis written once can run across studies, for human participants and artificial agents alike.
———Click a dot joined to the token.
A simplified map in the style of ThinkAhead (Eluchans et al., 2025). The state and action strings use the same canonical format as the Planning 101 data model.
Why Planning 101
Researchers have studied how people plan with graph navigation, two-stage decision games, subway routing, mazes and board games. Each dataset is usually released in its own laboratory's format and processed with study-specific scripts, where the conversion from raw data to analysed variables is written inline rather than documented. Combining studies then runs into two obstacles that are hard to detect after the fact.
A dropped column, an unrecorded exclusion or a filtered timeout leaves no trace in the processed file. The file looks complete, so a later analyst cannot see what is missing.
The same logical state, written two different ways in two studies, does not match when the data are joined. Paradigms that share a structure cannot be compared on it.
Planning 101 addresses the representation first. Cross-paradigm modelling, comparisons between human and model behaviour, and a review of which aspects of planning have and have not been studied all depend on the data sitting in one validated, queryable structure.
How it works
Every conversion decision is written down, versioned and checked, rather than left in a notebook cell.
Each paradigm is specified as a Markov decision process: its states, actions, transitions, rewards, and what the participant can observe.
Variants that share a structure share a Task. The magic-carpet and spaceship versions of the two-step task, for example, differ only in their cover story, which is kept as a recorded parameter rather than discarded.
A YAML manifest states how a source dataset maps into the schema: which files, which columns, which exclusions, and which published result the converted data must reproduce.
The mapping is explicit, reviewable and versioned, in the spirit of BIDS in neuroimaging.
Schema and chain checks first; then, where a task's dynamics are known, checks that every move is legal and every outcome reachable. A dataset that fails them is not written.
Reproduction targets re-derive a published result from the converted data. Output is Parquet or CSV with a provenance record: schema and manifest versions, source commit, content hash.
streams:
- role: steps # each event becomes a Step
record_path: "TrialEventData"
state_fields: ["touchedNode"]
onset: {kind: field, field: "time", unit: s}
- role: trace # pointer samples
record_path: "UserTrajectory"
channels: {x: "x", y: "y", z: "z"}
- role: episode_extras # trial summary
fields: {trial_score: "TrialScore"}
Each one addresses a documented failure mode, and is enforced in code or checked in review.
The data model
Each Step belongs to one Episode, each Episode to one Session, each Session to one Subject and one Experiment, and each Experiment to one Task. Because these links are enforced, analysis code written against them holds for any converted dataset.
Humans and agents share one table. A planning agent is a Subject with
subject_type = "agent", so comparing human and model behaviour reads from one
structure instead of two parallel ones.
The experiments
The selection, agreed in September 2026, starts from a homogeneous initial group and adds paradigms in order of coding effort. Five already have conversion rules; four are next.
Collect every gem on a map of dots joined by lines. Some maps are solved at a glance, others need several moves of look-ahead. Measures how far ahead people plan as difficulty grows.
Two choices in a row: a magic carpet or spaceship that usually leads to one place, then an option that pays off with a slowly drifting probability. Tests whether people build and use a map of the game, and how much the instructions matter.
The same two-move game with one rule simplified: the first choice always leads to the same place. Asks when it pays to reason ahead instead of repeating what worked last time.
Travel through a network of places grouped into neighbourhoods. Tests whether people split the map into neighbourhoods and plan between them first, then within each.
Learn an eight-place map by moving through it, then name the stop you would aim for first. Compares the declared subgoals with the paths people actually take.
Steer through mazes to a goal, then report which obstacles you noticed. Shows that people plan with a simplified picture that keeps only the obstacles relevant to the route.
Find a hidden reward in a maze whose edges wrap around, then return to it from somewhere else. Measures how long people pause to think before each move.
Uncover values hidden in a web of boxes, where every click costs, then choose a three-move path. Observes which information people pay for before they commit.
Two players compete on a board of four rows by nine columns, in some sessions with eye tracking. Measures how many moves ahead players think and how this grows with practice.
“P1–P5” is the proposed implementation phase, by increasing coding effort. Each dataset remains the work of its original authors and is subject to its own licence.
The experiments site
The experiments site is an interactive application built on the Planning 101 pipeline. Browse the selected studies, check how each conversion rule turns original rows into Episodes and Steps, map a dataset of your own, and replay recorded sessions.
The selected studies with their papers, original datasets and rule status. Browse any conversion episode by episode, original rows beside the converted Steps.
Upload a CSV or JSON file in any layout and map its columns visually to subject, state, action, reward and timing. Save the result as a new rule.
Convert files with an existing rule, several at once with one per subject, and compare the result with the original.
Combine converted datasets into one consolidated table to query and chart.
Replay a participant's recorded Magic Carpet session as they played it, trial by trial.
No installation needed.
Who we are
Planning 101 is a project of the Institute of Cognitive Sciences and Technologies (ISTC) of the National Research Council of Italy (CNR).
The experiments come from the authors of the original studies, credited above. Planning 101 builds on their published data and code.
Ready when you are
Open the experiments site to browse the studies, inspect the conversions and replay recorded sessions.