Evidence for Skills Test #

Every substantive claim on the Skills Test page is checked against current research. Here is each claim, how well today’s evidence supports it, and the sources. The full, de-duplicated source list lives on the references page.

Mixed · moderate evidence — A 23-question self-rating questionnaire can MEASURE your current learning skill in each topic and your ’learning improvement potential’ (the page’s framing claim and the premise behind scoring each subsection from a single self-report item).

Judged in the intended spirit (a free self-coaching/orientation tool, not a validated psychometric), the test is fine as a reflective prompt. But the words ‘measure’ and ‘score’ overstate what single-item self-ratings deliver: self-report tracks objective skill only weakly. No accuracy/reliability validation of THIS instrument exists.

Sources: Self-assessment accuracy review, AJPE 2021, PMC8086612 · Mabe & West (1982) meta-analysis, mean self/objective r≈.29 · Self-estimated vs measured intelligence meta-analysis, r≈.33 (PMC8883889)

Supported · strong evidence — Judge learning by what you can still do days later, not by how fluent the material felt; the effortful study that feels harder builds more durable memory than rereading and highlighting (Section 1 ‘How learning works’, Q1 ‘Effortful learning’).

This is the ‘desirable difficulties’ principle, and it is one of the most robust ideas in the science of learning. Conditions that slow you down in the moment — recalling instead of rereading, spacing instead of massing — reliably produce stronger long-term retention, while fluency-based study feels productive and isn’t. The major review of study techniques reached the same verdict.

Sources: Bjork & Bjork, desirable difficulties (review chapters) · Dunlosky et al. 2013, PSPI — practice testing & distributed practice rated highest; rereading/highlighting low · Brown, Roediger & McDaniel, Make It Stick (2014)

Supported · moderate evidence — Working memory is small and long-term memory vast; cutting needless cognitive load and pairing words with visuals (dual coding) helps learning (Q2 ‘Working memory’).

Both premises are well established. Cognitive load theory shows that overloading the limited workspace harms learning, and that chunking and removing extraneous load help. Dual coding / multimedia research shows that combining verbal and visual representations generally beats either alone. Effects are real but moderated by how it’s implemented.

Sources: Sweller, cognitive load theory (reviews) · Mayer, multimedia learning / dual-coding principle · Paivio, dual-coding theory

Supported · strong evidence — Retrieving material from memory (self-testing, flashcards, explaining aloud) builds learning far more than rereading (Section 2 ‘Proven methods’, Q3 ‘Retrieval’).

The testing effect is among the best-supported findings in the science of learning, replicated across ages, subjects and lab/classroom settings, with medium-to-large effects. Retrieval practice also improves later transfer and exam performance. The manual’s emphasis on retrieval as the single highest-leverage habit is well justified.

Sources: Roediger & Karpicke 2006 (the testing effect) · Dunlosky et al. 2013, PSPI — practice testing rated highest · Adesope, Trevisan & Sundararajan 2017, Rev Educ Res — testing-effect meta-analysis (g≈0.5)

Supported · strong evidence — Spreading study and repetition across days, rather than cramming, reinforces learning and maintains it long-term — including via spaced-repetition flashcard apps (Q4 ‘Spacing’ and Q10 ‘Repetition tools’; e.g. Anki).

The spacing/distributed-practice effect is one of the most robust findings in the field, replicated across labs, domains and large real-world cohorts with medium-to-large effects. Spaced-repetition apps simply automate spacing plus retrieval, so the premise behind Q10 follows directly. The forgetting that motivates timed review is itself well replicated (the curve’s shape, not exact ‘Ebbinghaus percentages’).

Sources: Cepeda et al. 2006 — distributed-practice meta-analyses (medium-large spacing effect) · Spaced repetition in practising physicians, JAMA Netw Open 2024 (d=0.62 / 0.26) · Murre & Dros 2015, PLOS ONE e0120644 — forgetting-curve replication

Supported · moderate evidence — Mixing related topics or problem types within a session (rather than drilling one to exhaustion) helps you tell problems apart and choose the right approach later (Q5 ‘Interleaving’).

Interleaving reliably improves later discrimination and category/problem-type learning, even though it depresses practice-time performance and feels harder — exactly the ‘desirable difficulty’ the test describes. The benefit is clearest when the interleaved items are genuinely confusable; for some material the gains are smaller.

Sources: Rohrer & Taylor 2007 (interleaved maths practice) · Brunmair & Richter 2019, Psychol Bull — interleaving meta-analysis · Dunlosky et al. 2013, PSPI

Supported · moderate evidence — Connecting new material to what you already know — asking why/how, explaining in your own words — builds more durable, retrievable memory than rote memorising (Q6 ‘Elaboration’).

Elaborative interrogation and self-explanation are supported strategies in the major reviews, rated ‘moderate utility’ — effective across a range of learners, with the strongest gains for material that can be tied to existing knowledge. The ‘many connections, many routes back’ framing is consistent with the evidence.

Sources: Dunlosky et al. 2013, PSPI — elaborative interrogation & self-explanation (moderate utility) · Weinstein, Madan & Sumeracki 2018 — elaboration as a core strategy

Supported · moderate evidence — Association/mnemonic techniques — basic mnemonics, linked lists, peg words, the method of loci — support and improve learning and recall (Section 3 ‘Techniques’, Q7 ‘Association’).

The core premise — visual/spatial mnemonics like loci and pegs boost memorization — is genuinely supported with medium-to-large effects across two meta-analyses. Caveat: benefits are clearest for arbitrary list/serial recall; primary-study quality is limited, so effect sizes are likely optimistic. The manual correctly frames these as elaboration applied, not magic.

Sources: Method of loci systematic review & meta-analysis, Br J Psychol 2025 (d=0.88; low GRADE) · Twomey & Kroneisen 2021 meta-analysis of loci RCTs (g=0.65)

Supported · moderate evidence — Visualization / mental practice and rehearsing skills in realistic, varied conditions supplement real training and improve skill performance (Q8 ‘Visualization’ and Q9 ‘Simulation’).

Premise supported: mental rehearsal aids skill learning, best as a supplement to (not a replacement for) physical practice — which matches the manual’s wording. Practising under varied, realistic conditions also aids transfer. Effects are larger for cognitive than purely physical tasks and decay without continued practice.

Sources: 24-year meta-analytic replication of mental-practice effects, Psych Sport Exerc 2019 · Imagery + physical practice > physical practice alone (sport imagery meta-analyses 2023-2025)

Mixed · emerging evidence — Using AI tools to generate extra practice, quizzes and fast feedback can multiply your retrieval and practice, provided you keep doing the thinking yourself rather than letting AI hand you fluent answers (Section 4 ‘Learning with AI’, Q11–Q12).

The mechanism is sound: anything that creates more retrieval, spacing and feedback should help, and intelligent tutoring systems have a track record of gains. But AI for learning is new and the long-term evidence is thin, and the test’s caution is itself supported — offloading effortful thinking to AI removes the very difficulty that drives learning. Treat this section as principled guidance, not settled science.

Sources: VanLehn 2011, Educ Psychol — effectiveness of intelligent tutoring systems · Bjork & Bjork, desirable difficulties (the ‘don’t outsource the effort’ caution) · emerging generative-AI-in-education literature (evidence still developing)

Supported · moderate evidence — Making mistakes plays a vital role in learning, and there is real value in capturing and learning from your own and others’ errors (Section 7, Q21 ‘Mistakes’).

Supported with an important condition the test should state: errors help learning ONLY when followed by corrective feedback (and are best when self-generated). The manual’s ‘capture and learn from mistakes’ framing is consistent with this; uncorrected errors don’t confer the benefit, and errorless methods can be better for severe memory impairment.

Sources: Deliberate errors enhance learning, Contemp Educ Psychol 2025 · Errors during learning: definition, modulators, theories — Psychon Bull Rev review (2021) · Errorful learning study in brain-injury populations, Neuropsychol Rehabil 2023

Supported · strong evidence — Pressure and stress can change learning and recall, and a balanced, prepared approach (breathing, preparation, composure techniques) helps in exams and high-pressure moments (Q20 ‘Fear & pressure’).

Strongly supported: acute pressure/stress around retrieval (e.g. an exam) reliably impairs recall, which justifies the manual’s stress-management emphasis. Nuance the test omits: stress at encoding/consolidation can enhance memory, so ‘pressure’ is not uniformly bad — effect direction depends on timing.

Sources: Shields et al. 2017 acute-stress & episodic-memory meta-analysis (PMC5436944) · Het et al. acute-cortisol meta-analysis: cortisol before retrieval impairs recall (d≈-0.49) · Stress & long-term memory retrieval systematic review (PMC7879075)

Supported · strong evidence — Learn using all modes and put your effort into the proven methods rather than matching study to a single ’learning style’; matching teaching to a supposed style does not improve results (Section 8, Q22 ‘Use every mode’ and Q23 ‘Beyond your style’).

This is a rare case where the revised test now states the evidence-based position. The ‘meshing hypothesis’ — that matching techniques to your dominant style improves learning — is the textbook education neuromyth; even the most generous recent meta-analysis calls the benefit too small to warrant adoption, and recommends teaching in multiple modalities instead. (The previous edition of this test endorsed style-matching; that claim has been removed.)

Sources: Pashler et al. 2008 (PSPI) — no adequate evidence for matching, still standard · Clinton-Lisell & Litzinger 2024, Front. Psychol. 15:1428732 (g=0.31, not worth adopting) · Frontiers meta-analysis of the matching hypothesis 2024, PMC11270031

Memletics Manual v4.1.0 · Changelog