MindFrame

Evidence & limits

What the research says, and what we have not measured

MindFrame is built on published research about metacognition — thinking about your own thinking. That research was done in schools, clinics and forecasting tournaments, not on MindFrame. This page keeps the two apart.

The research we build on

Each finding is attributed to the people who measured it and the group they studied.

g = 0.63

academic performance after metacognitive strategy instruction

Studied in: school students, 67 studies

A 2018 meta-analysis of 67 school studies (de Boer, Donker & van der Werf) found metacognitive strategy instruction improved academic performance by about g = 0.63 — a typical trained student scored above roughly 73% of untrained peers.

de Boer, Donker & van der Werf (2018)

g = 0.69

symptom change, metacognitive therapy compared with Cognitive Behavioral Therapy (CBT)

Studied in: clinical trials for anxiety and depression

In clinical trials reviewed by Normann & Morina (2021), metacognitive therapy outperformed Cognitive Behavioral Therapy (CBT) by about g = 0.69. This is therapy research; it supports the principle, not a training app.

Normann & Morina (2021)

6–11%

forecasting accuracy (Brier score) after under an hour of probability training

Studied in: Good Judgment Project forecasters

In the Good Judgment Project (Mellers et al. 2014), forecasters given under an hour of probability training were consistently more accurate than a control group, improving Brier scores by about 6–11% over a two-year tournament.

Mellers et al. (2014) — Good Judgment Project

Little far transfer

benefit outside the trained task

Studied in: review of brain-training studies

A 2016 review of brain-training studies (Simons et al.) found improvement on the trained tasks but little evidence of real-world transfer to everyday performance.

Simons et al. (2016) — Do 'brain-training' programs work?

What MindFrame measures

For each person, MindFrame measures accuracy, the gap between confidence and accuracy (calibration), and reasoning quality, session by session, and shows the trend.

  • Accuracy

    Whether each answer was right, against an answer key written before the challenge was published.

  • Calibration

    How far your stated confidence is from how often you are right, summarised as a Brier Score and a calibration error.

  • Reasoning quality

    When you explain an answer in depth, Claude scores the argument; otherwise a published formula does, and your results say which one was used.

  • Your own trend

    How those numbers move for you across sessions. It is a description of your practice, not a controlled experiment.

What we have not measured

MindFrame has not yet run a controlled study, so we do not claim it produces the effect sizes reported in the research. We will publish our own numbers, including weak ones, once we have them.

  • Whether MindFrame produces the effect sizes reported in the research above.
  • Whether practice here changes the decisions you make outside MindFrame.
  • How MindFrame users compare with people who do not use it.
  • Long-term retention after people stop practising.

How scoring works, in short

  1. You answer a challenge and say how sure you are before you see the result.
  2. Accuracy is checked against a fixed answer key; calibration compares your confidence with your accuracy.
  3. If you explain your reasoning in depth, Claude scores it; otherwise a published formula does.
  4. Your numbers are shown to you as a trend. They are never sold or used to train AI models.

The formulas are on How we score.

Go deeper