Altered Riddles

Altered Riddles investigates whether language models can override memorised solutions when familiar riddles are minimally perturbed. The central hypothesis is that models with sufficient exposure to canonical riddles during pretraining will exhibit a measurable tendency to reproduce memorised answers even when the altered prompt logically demands a different one.

Leaderboard

Models are ranked by Conditioned Override Rate (COR) — the fraction of altered riddles on which the model produced the original memorised answer, conditional on having solved the unaltered riddle correctly. Lower COR is better.

Loading benchmark data…

Charts

Verbosity vs COR

Does thinking longer help? Each dot is a model: its x-position is the average number of output tokens per answer (log scale), its y-position is the conditioned override rate. A ring indicates a reasoning-enabled model.

By alteration type

Not every alteration is equally hard. Each column is a type of alteration; every dot is one model's COR on that subset. The shaded band is the interquartile range across models, the horizontal tick is the median, and the diamond marks the mean.

Background

Motivation

The benchmark originated from an observation made while constructing the academic-chains dataset. Testing a minor variant of the classic surgeon riddle produced a surprising failure pattern:

The surgeon, who is the boy's father, says "I cannot operate on this boy — he's my son!" Who is the surgeon to the boy?

For reference, the original riddle reads:

A man and his son are in a terrible accident and are rushed to the hospital in critical condition. The doctor looks at the boy and exclaims, "I can't operate on this boy; he's my son!" How could this be?

The canonical answer — "The doctor is the mother" — exploits implicit gender assumptions to produce an apparent paradox. In the altered version, the father's identity is stated explicitly, making "the father" the only correct response. Yet models including claude-sonnet-4.6, gemini-3.1-flash, and several others consistently produce "the mother," apparently pattern-matching to training exposure rather than reasoning over the modified prompt.

The likely mechanism is pattern override: pretraining exposure is dense enough that recognition of the riddle's surface form activates a stored answer, overriding careful processing of the modified constraints.

Alteration types

Every alteration must entail its new answer and rule out the original one under an ordinary reading of the text. Five checkable types are used:

  • Stated — the answer is written in the text itself, so the memorised answer contradicts what was said (as in the surgeon example above).
  • Hard constraint — a numeric or lexical constraint is added that excludes the original answer.
  • Negated premise — a premise the original answer depends on is negated.
  • Trivialized — the famous complication is removed, so the puzzle becomes a plain reading exercise.
  • Question swap — same premises, a different question, so the memorised answer answers the wrong question.

This taxonomy allows per-category analysis of which perturbation types are most and least resistant to pattern override, and which model families show differentiated performance across categories.

Metrics

Conditioned Override Rate (COR) — lower is better
The proportion of altered riddles on which the model produces the original memorised answer, restricted to riddles the model reliably solves in unaltered form (at least 80% of its original-riddle samples correct). A high COR means the model is pattern-matching to training data instead of reasoning from the altered text. Intervals and rank spreads come from a bootstrap clustered by source riddle.
Original Accuracy — higher is better
Accuracy on the unmodified riddles. Serves as a knowledge baseline and conditioning variable for COR. A model with low original accuracy cannot exhibit meaningful pattern override, so this figure contextualises all other metrics.
Altered Accuracy — higher is better
Accuracy on the modified riddle set. Unlike COR, this metric is unconditional — it reflects raw performance on the harder task and provides a direct measure of downstream usefulness.
Pattern Override Rate — lower is better
The unconditional rate at which the model produces the original canonical answer on modified riddles, regardless of whether it solved the original.

Detailed Results

The full leaderboard including confidence intervals, raw pattern override rate, and average output tokens per answer. Raw data available as leaderboard.json.

Resources

For full methodology, data, run manifests and design details, see the GitHub repository.

Citation

If you use Altered Riddles in your work, please cite it as:

De Santis, Marco, "Altered Riddles", marcodsn.me, Apr 2026.

Or use the BibTeX citation:

bibtex
@misc{desantis2026alteredriddles,
  author = {Marco De Santis},
  title = {Altered Riddles},
  year = {2026},
  month = apr,
  url = {https://marcodsn.me/altered-riddles}
}