Supporting research — guided tutor

Why the tutor teaches the way it does.

Every design decision in the guided tutor comes from a named finding in the learning-sciences literature. This page states each decision, cites the finding behind it — every citation links to the full reference — and, just as deliberately, states the limits of the evidence. For the short version, read the FAQ.

This page covers the guided tutor. Roleplay rests on a different literature — simulation and communication-skills training — and has its own page. Neither body of evidence stands in for the other.

Citations are author–year links into the reference list; every entry there has been checked against the published record.

Back to the home page · Questions and answers

01

The claim we make — and the one we don't.

What we claim is that the tutor's interaction is step-based and its content is grounded: it teaches one small step at a time from material a person wrote and reviewed, and it may teach nothing else. Both properties are verifiable in the product, turn by turn (§12).

What we do not claim is the famous "2 sigma" result — the finding that tutored students outperform classroom students by two standard deviations (Bloom, 1984). That figure is the origin of most claims made for tutoring software, and it has not replicated at that size: the modern synthesis puts human tutoring at d ≈ 0.79, not 2.0 (VanLehn, 2011), and the meta-analyses in §9 are all substantially below the legend. We cite the measured numbers, not the legendary one.

The tutor also does not mark. No score, grade or competency judgement comes out of it — the transcript, audio and whiteboard are evidence for a human assessor, and the judgement stays with them (§10).

02

The tutor is given the course, verbatim.

The tutor receives the authored course prose word for word — generated from an Australian unit of competency, or attached by the teacher as ordinary documents — and may teach from nothing else. This decision has a history: an early version received only topic identifiers and page titles, and a reviewing teacher observed that a student would need the textbook open in another tab to answer it. Everything the tutor said about the subject was generated rather than taught.

The fix is the split-attention effect: when learners must integrate two mutually-referring sources — a tutor here, a book there — the integration work lands in working memory, where capacity is the scarcest resource in learning (Chandler & Sweller, 1992). Cognitive-load theory predicts exactly this failure mode: extraneous processing displaces the processing that builds understanding (Sweller, 1988). Physically integrating the sources — the tutor teaching from the same words the learner read — removes the split at its cause, the standard remedy in multimedia-learning design (Kalyuga, Chandler & Sweller, 1999) (Mayer & Fiorella, 2021).

Three consequences follow. Nothing is paraphrased on the way in, because a paraphrase is a second, ungated version of the course that can disagree with the book the student is enrolled in. A scanned PDF with no text layer is refused rather than accepted empty, because accepting it would silently recreate the original defect — a tutor teaching from the model's memory. And asked something outside the material, the tutor declines and says what it does cover.

03

A tutor that asks, not one that explains.

The tutor never produces the step it has asked the learner for. That is not a personality choice; it is the design decision the tutoring literature supports most strongly, on three converging grounds.

Granularity. The most influential synthesis of tutoring research found that tutoring systems match human tutors — d ≈ 0.76 against d ≈ 0.79 — when, and only when, the interaction happens at the level of individual solution steps rather than whole answers (VanLehn, 2011). Step granularity is the active ingredient, so work is tracked one expectation at a time, not one reply at a time.

Activity. The ICAP framework orders learning activities by the cognitive engagement they demand: interactive and constructive activity reliably beats active manipulation, which beats passive reception (Chi & Wylie, 2014). A tutor that lectures is worse than the book — the book explains better and does not interrupt. Most of this tutor's output is therefore questions, which also present far less surface for factual error than explanations do.

Dialogue design. The interaction pattern follows expectation-and-misconception tailored dialogue, the AutoTutor line of work: for each criterion the tutor is given the expected answers and the mistakes students actually make (Graesser et al., 2004). The misconceptions are the course question bank's own distractors with the feedback that repairs them — authored, reviewed, and already known to be plausible — not invented on the fly.

The recent LLM-tutoring evidence sharpens the point: unrestricted chatbot help inflates practice scores and then harms unassisted performance, while answer-withholding guardrails eliminate the harm (Bastani et al., 2025). A tutor that never produces the step it asked for is the second configuration, not the first (§9).

04

The shape of a session.

A session runs: model one worked instance → hand over a partly-completed one → hand over a whole one → ask the learner to generalise. That arc is two well-evidenced structures laid on top of each other.

The first is cognitive apprenticeship — make the expert's thinking visible, then coach, then fade, transferring responsibility as competence grows (Collins, Brown & Newman, 1989). The second is the faded worked-example sequence from cognitive-load research: worked examples beat unsupported problem-solving for novices (Sweller & Cooper, 1985), and the best transition from example to independent work is gradual — remove one worked step at a time rather than jumping from demonstration to a blank page (Renkl & Atkinson, 2003).

05

Help that fades — and comes back.

Support is not a setting; it is contingent on demonstrated performance, which is how the scaffolding literature has defined it since the term was coined (Wood, Bruner & Ross, 1976).

It fades. Once a learner produces two expectations without hints, the tutor withdraws support a level. Instructional support that helps a novice measurably harms a learner past that point — the expertise-reversal effect (Kalyuga, Ayres, Chandler & Sweller, 2003).

It returns. When a learner stalls twice, support steps back in. Contingent tutoring — more help after failure, less after success — outperformed every fixed teaching strategy in the experiment that defined it (Wood, Wood & Middleton, 1978), and contingency, fading and transfer of responsibility remain the definition of scaffolding in the modern review literature (van de Pol, Volman & Beishuizen, 2010). A support level that only ever decreased would withdraw help exactly when it is needed.

Driven by performance, never self-report. What the learner already knows is the dominant variable in what they can learn next (Ausubel, 1968) — so adaptation keys off what the learner actually produced against the unit's own criteria, the same evidence a human assessor uses, and never off a questionnaire about how they like to learn (§8).

06

Why there is a whiteboard.

Detail goes on the board; speech carries meaning. The tutor writes worked steps, equations and diagrams on a persistent whiteboard and never reads the board out loud, because narrating text already on screen adds cognitive load without adding information — the redundancy effect (Kalyuga, Chandler & Sweller, 1999). Spoken explanation alongside a visual, by contrast, works with the structure of working memory rather than against it: presenting graphics with spoken rather than written words improves learning at d = 0.72 for complex material — the modality effect (Ginns, 2005), one of the most stable principles in the multimedia-learning literature (Mayer & Fiorella, 2021).

The learner can draw and point on the same board, and the tutor is told what they drew — participation that is interactive in the ICAP sense (Chi & Wylie, 2014), and a second means of expression for learners who would rather show than say.

Text size, contrast and voice speed are offered to everyone as standard controls, not unlocked by a disability label — multiple means of representation, engagement and expression provided by default, which is Universal Design for Learning applied as designed (CAST, 2024). Every spoken turn appears as live text, so no content is audio-only; WCAG 2.2 AA is the platform's conformance target and is treated as a procurement gate, not a polish item.

07

Feedback about the work — then retrieval.

When a learner is wrong, the tutor names the failing step and asks a question whose answer is the correction. Praise about the person is absent by design. The feedback meta-analyses are consistent on the split: feedback about the task, the process and the learner's self-regulation carries the learning effect, while praise directed at the person carries little or none (Hattie & Timperley, 2007). The most recent large meta-analysis quantifies it — feedback averages d = 0.48 overall, high-information feedback reaches d = 0.99, and reinforcement-style praise manages roughly a quarter of that (Wisniewski, Zierer & Hattie, 2020). Formative-feedback guidance points the same way: specific, task-level, and delivered while the learner can still use it (Shute, 2008).

Once every expectation has been produced, the tutor stops teaching and asks the learner to retrieve and explain. Testing beats restudying for retention — the effect replicates at g ≈ 0.50–0.61 across meta-analyses, and works best when feedback follows (Roediger & Karpicke, 2006) (Rowland, 2014) (Adesope, Trevisan & Sundararajan, 2017). Prompted self-explanation carries its own reliable effect, g = 0.55 (Bisra et al., 2018), and generating an explanation improves understanding beyond restating it (Chi et al., 1994). The end of a lesson is exactly where both belong.

08

What we deliberately did not build.

Learning styles
No questionnaire asks whether a learner is "visual" or "auditory", and nothing adapts to such an answer. The "meshing hypothesis" — that matching instruction to a self-reported style improves learning — has been tested under fair conditions and fails to produce the predicted gains (Pashler, McDaniel, Rohrer & Bjork, 2008) (Riener & Willingham, 2010), and continuing to build on it is actively discouraged in the field (Kirschner, 2017). The system adapts to demonstrated performance instead (§5) — which is also what a VET assessor does, and is defensible in a privacy review in a way that inferring traits is not.
Retrieval / vector search
One performance criterion is about 1,200 words and a week about 2,300. Relevance here is addressed by identifier — the material tagged with that criterion — not by similarity, so a vector search could only approximate a lookup the build already performs exactly, and would fail silently when it ranked the wrong passage fourth. The whole corpus is given to the tutor directly.
Automated marking
A score from the tutor would change what the tool is, what an RTO must validate, and what a student may appeal. The platform produces evidence, not judgements (§10).
Affect and engagement inference
Frustration detection, confidence estimates and attention scores were considered and rejected. None would improve the one decision the system makes — how much help to give next — and all of them would turn a teaching tool into surveillance.

09

The numbers.

Sections 2–7 map each decision to the model it came from. This section is the quantitative context — what tutoring, simulation and the component techniques are worth when they are measured, so nobody has to take a design argument on faith. Effect sizes are in standard deviations (d or g) as reported by the cited meta-analysis or trial; none of them measures this product (§11).

What tutoring buys

Human tutoring
d ≈ 0.79 — the measured reality behind the 2-sigma legend (VanLehn, 2011).
Step-based tutoring systems
d ≈ 0.76 — near parity with human tutors, with step granularity as the active ingredient (VanLehn, 2011).
Intelligent tutoring systems
g = 0.42 against large-group instruction across 107 effects (Ma et al., 2014); a median 0.66 SD across 50 evaluations (Kulik & Fletcher, 2016); g ≈ 0.32–0.37 for college learners (Steenbergen-Hu & Cooper, 2014).
Human tutoring programs at scale
0.29 SD pooled across 96 randomised trials (Nickow, Oreopoulos & Quan, 2024).

What LLM tutors specifically buy — and cost

The two most relevant recent experiments point in opposite directions, and the difference between them is this system's design argument. A crossover RCT in Harvard's introductory physics course found students learned more than twice as much per unit time from a scaffolded, guardrailed LLM tutor built on research-based pedagogy than from an equivalent active-learning class (Kestin et al., 2025). A field experiment with roughly a thousand high-school students found unrestricted chatbot access raised practice performance 48% but harmed subsequent unassisted exam performance — while a tutor variant with teacher-designed guardrails and answer-withholding raised practice performance 127% and eliminated the harm (Bastani et al., 2025). A tutor that never produces the step it asked for (§3) is the second configuration, not the first.

The role-play half

Roleplay rests on a different literature, and its numbers need their comparators attached to mean anything — the figure usually quoted for simulation is measured against no teaching at all, not against other teaching. Rather than summarise it badly here, it has its own page, with the limits stated alongside the findings.

10

The policy fit.

UNESCO's guidance for generative AI in education asks for human-centred design in which consequential judgements stay with people (UNESCO, 2023). TEQSA's assessment-reform guidance for Australian higher education draws the same line: AI may support learning, while assured assessment remains human-judged (Lodge et al., 2023). The Australian Framework for Generative AI in Schools requires transparency, accountability and privacy by design (Department of Education, 2023).

A tutor that teaches but does not mark, records but does not profile, and shows the teacher everything it was given sits inside all three by construction rather than by exception. For VET specifically, the content version a conversation was taught from is stamped on it, so the currency of evidence under the Standards for RTOs is demonstrable after a unit rebuild (see the Australian context note under the references).

11

Limitations. Read this section before quoting any other.

An honest evidence page states what the evidence does not show.

No effect size is claimed for this product
The literature above supports the design decisions; it does not measure this tutor. Whether learners using it learn more than they would from the book alone needs a controlled comparison, which has not been run. The effect sizes in §9 belong to the cited studies and their interventions, not to us.
An LLM can still be confidently wrong
Grounding narrows what the tutor draws on; it does not make it incapable of error. Turns using vocabulary absent from the course material are flagged for the teacher's review — a shortlist for a human, not a guarantee, and it cannot catch a wrong statement made entirely in the course's own vocabulary.
The step record is the model's own claim
"The learner produced this step" is the model's report — on the voice path especially — used only to decide how much help comes next. It is not independently verified, it is not evidence, and it is never shown to an assessor as a result.
Trials so far are qualitative
Transcript review and scripted adversarial students have caught real defects. Reading transcripts tells you whether the teaching looks sound; it cannot tell you whether anyone learned more. Ten scripted adversarial students is not a trial, and we do not present it as one.
An upload is grounded but not scaffolded
Material generated from a unit of competency brings more than prose: expected answers, known mistakes, check questions, assessor benchmarks and a scope fence. An uploaded document brings the prose alone — measured on one real week, 44% of what the tutor is given is lost in a document-export round-trip, and it is the part that makes it teach rather than recite. Those elements are never invented for an upload, because fabricating them would put unreviewed teaching content into the tutor.
Expectations are currently derived, not authored
They come from the assessor benchmarks for the criterion, falling back to the narration's own sub-headings. Authoring them alongside the content, so they pass the same review the prose does, is the next step.

12

How to check the claims yourself.

Everything above is verifiable from inside the product, by the person who has to vouch for it.

The material
The chatbot edit form shows the teaching content exactly as it reaches the model — source, sections, word count, and Read what the tutor sees. Nothing is hidden. Where the unit was generated, compare it against the course book: it should be the same words.
The teaching, turn by turn
Every conversation replay shows the move the tutor made on each turn (elicit, hint, repair, assert…) and the whiteboard as the learner saw it — the whole lesson, not the last frame of it.
Who did the work
The session facts strip counts answers given away — turns where the tutor asserted instead of asking — and what share of the words the learner produced. The tutor is capped at forty words a turn so the learner talks more than it does.
Whether it invented anything
Turns using words absent from the course material are flagged under the text, with the words listed. Some are ordinary paraphrase; the judgement is yours.
The refusals
Ask it about a later week: it names the week. Ask outside the unit: it declines and says what it does cover. Ask it to just give the answer, repeatedly: it should not.

13

References.

Each entry has been checked against the published record (DOI resolution and publisher metadata). In-text citations link here; use your browser's back button to return to the claim.

  • Adesope, O. O., Trevisan, D. A., & Sundararajan, N. (2017). Rethinking the use of tests: A meta-analysis of practice testing. Review of Educational Research, 87(3), 659–701. https://doi.org/10.3102/0034654316689306
  • Ausubel, D. P. (1968). Educational Psychology: A Cognitive View. New York: Holt, Rinehart and Winston.
  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
  • Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725. https://doi.org/10.1007/s10648-018-9434-x
  • Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16. https://doi.org/10.3102/0013189X013006004
  • CAST. (2024). Universal Design for Learning Guidelines version 3.0. https://udlguidelines.cast.org
  • Chandler, P., & Sweller, J. (1992). The split-attention effect as a factor in the design of instruction. British Journal of Educational Psychology, 62(2), 233–246. https://doi.org/10.1111/j.2044-8279.1992.tb01017.x
  • Chi, M. T. H., de Leeuw, N., Chiu, M.-H., & LaVancher, C. (1994). Eliciting self-explanations improves understanding. Cognitive Science, 18(3), 439–477. https://doi.org/10.1207/s15516709cog1803_3
  • Chi, M. T. H., & Wylie, R. (2014). The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49(4), 219–243. https://doi.org/10.1080/00461520.2014.965823
  • Collins, A., Brown, J. S., & Newman, S. E. (1989). Cognitive apprenticeship: Teaching the crafts of reading, writing, and mathematics. In L. B. Resnick (Ed.), Knowing, Learning, and Instruction: Essays in Honor of Robert Glaser (pp. 453–494). Hillsdale, NJ: Lawrence Erlbaum Associates.
  • Department of Education (Australia). (2023). Australian Framework for Generative Artificial Intelligence (AI) in Schools. https://www.education.gov.au/schooling/resources/australian-framework-generative-artificial-intelligence-ai-schools
  • Ginns, P. (2005). Meta-analysis of the modality effect. Learning and Instruction, 15(4), 313–331. https://doi.org/10.1016/j.learninstruc.2005.07.001
  • Graesser, A. C., Lu, S., Jackson, G. T., Mitchell, H. H., Ventura, M., Olney, A., & Louwerse, M. M. (2004). AutoTutor: A tutor with dialogue in natural language. Behavior Research Methods, Instruments, & Computers, 36(2), 180–192. https://doi.org/10.3758/BF03195563
  • Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. https://doi.org/10.3102/003465430298487
  • Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31. https://doi.org/10.1207/S15326985EP3801_4
  • Kalyuga, S., Chandler, P., & Sweller, J. (1999). Managing split-attention and redundancy in multimedia instruction. Applied Cognitive Psychology, 13(4), 351–371.
  • Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6
  • Kirschner, P. A. (2017). Stop propagating the learning styles myth. Computers & Education, 106, 166–171. https://doi.org/10.1016/j.compedu.2016.12.006
  • Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1), 42–78. https://doi.org/10.3102/0034654315581420
  • Lodge, J. M., Howard, S., Bearman, M., Dawson, P., & Associates. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency (TEQSA).
  • Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. (2014). Intelligent tutoring systems and learning outcomes: A meta-analysis. Journal of Educational Psychology, 106(4), 901–918. https://doi.org/10.1037/a0037123
  • Mayer, R. E., & Fiorella, L. (Eds.). (2021). The Cambridge Handbook of Multimedia Learning (3rd ed.). Cambridge University Press. https://doi.org/10.1017/9781108894333
  • Nickow, A., Oreopoulos, P., & Quan, V. (2024). The promise of tutoring for PreK–12 learning: A systematic review and meta-analysis of the experimental evidence. American Educational Research Journal, 61(1), 74–107. https://doi.org/10.3102/00028312231208687
  • Pashler, H., McDaniel, M., Rohrer, D., & Bjork, R. (2008). Learning styles: Concepts and evidence. Psychological Science in the Public Interest, 9(3), 105–119. https://doi.org/10.1111/j.1539-6053.2009.01038.x
  • Renkl, A., & Atkinson, R. K. (2003). Structuring the transition from example study to problem solving in cognitive skill acquisition: A cognitive load perspective. Educational Psychologist, 38(1), 15–22. https://doi.org/10.1207/S15326985EP3801_3
  • Riener, C., & Willingham, D. (2010). The myth of learning styles. Change: The Magazine of Higher Learning, 42(5), 32–35. https://doi.org/10.1080/00091383.2010.503139
  • Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
  • Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432–1463. https://doi.org/10.1037/a0037559
  • Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153–189. https://doi.org/10.3102/0034654307313795
  • Steenbergen-Hu, S., & Cooper, H. (2014). A meta-analysis of the effectiveness of intelligent tutoring systems on college students' academic learning. Journal of Educational Psychology, 106(2), 331–347. https://doi.org/10.1037/a0034752
  • Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
  • Sweller, J., & Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction, 2(1), 59–89. https://doi.org/10.1207/s1532690xci0201_3
  • UNESCO. (2023). Guidance for generative AI in education and research. Paris: UNESCO. https://www.unesco.org/en/articles/guidance-generative-ai-education-and-research
  • van de Pol, J., Volman, M., & Beishuizen, J. (2010). Scaffolding in teacher–student interaction: A decade of research. Educational Psychology Review, 22(3), 271–296. https://doi.org/10.1007/s10648-010-9127-6
  • VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221. https://doi.org/10.1080/00461520.2011.611369
  • Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087. https://doi.org/10.3389/fpsyg.2019.03087
  • Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry, 17(2), 89–100. https://doi.org/10.1111/j.1469-7610.1976.tb00381.x
  • Wood, D., Wood, H., & Middleton, D. (1978). An experimental evaluation of four face-to-face teaching strategies. International Journal of Behavioral Development, 1(2), 131–147. https://doi.org/10.1177/016502547800100203

Australian context: Australian Core Skills Framework (ACSF, 2012 — still current); Disability Standards for Education 2005 (Cth) — in force, reviewed 2020; Standards for Registered Training Organisations 2025 (the Outcome Standards, F2025L00354, replaced the 2015 Standards from 1 July 2025) — principles of assessment and rules of evidence; Web Content Accessibility Guidelines (WCAG) 2.2, W3C Recommendation, October 2023.

See it teach.

Back to the home page