Calibration: confidence versus reality
Before feedback, you state how confident you are (50–99%). We convert that to a probability p and compare it with the outcome o (1 if correct, 0 if not) across your answers:
Calibration error = mean( |p − o| )Brier score = mean( (p − o)² ) — lower is better; 0 is perfectOverconfidence gap = mean(p) − mean(o) — positive means more confident than accurateThe Brier score is a proper scoring rule: your best strategy is to report your true confidence. Gaming it by always saying 50% or 99% costs you points. This is also why the Calibration Dividend rewards honest reporting.
The composite session score
Each answer gets five sub-scores from 0 to 100, combined with fixed weights:
| Part | Weight | How it is measured |
|---|---|---|
| Accuracy | 40% | Whether your answer was right, from the challenge's authored answer key. |
| Confidence | 20% | How closely your stated confidence matched the outcome (calibration). |
| Reasoning | 20% | The quality of your written reasoning, scored by rules first and AI when rules cannot decide. |
| Reflection | 10% | Whether you identified what went wrong and what you would change. |
| Adaptation | 10% | Whether you updated appropriately when new information arrived. |
The Cognitive Fingerprint
Your fingerprint has 8 axes (for example anchoring resistance, overconfidence control and uncertainty tolerance). Each axis averages your scores on the modes and challenge tags that exercise it. Axes with little data stay at a neutral baseline and are labelled as such rather than guessed.
Where AI is used — and where it is not
- Right/wrong answers and calibration maths are deterministic. No AI decides them.
- AI (Anthropic's Claude) scores free-text reasoning only when rule-based checks cannot, and writes coaching explanations.
- AI output is validated before it is shown; if it fails validation or is unavailable, you see a plain rule-based result instead of a guess.
More detail: responsible AI.
Games
Games such as ASCENT and Rule Shift use seeded, replayable runs so any result can be reproduced from its seed and moves. ASCENT's altitude uses an affine transform of the Brier score, so the same honesty incentive applies. Game scores describe performance in that game; they are not general ability scores.
What these numbers cannot tell you
- They describe your answers on MindFrame challenges, not your intelligence or mental health.
- Small samples are noisy. A handful of answers can swing calibration a lot.
- Improvement inside MindFrame does not prove improvement in daily life. We have not yet run an outcome study — see beta status.