AI Hallucinations in Education: Pre-Publication Tests Teachers Should Run
AI-generated lessons, explanations, and feedback can sound confident while containing errors, invented details, or poor instructional choices. Use these practical checks before sharing AI-assisted materials with learners.

AI hallucinations in education are outputs that appear plausible but are inaccurate, unsupported, incomplete, invented, or unsuitable for the learner and teaching context. They can occur in generated lesson plans, reading passages, answer keys, feedback, citations, examples, and assessment questions. The central risk is not simply that an AI tool makes an obvious mistake. A more difficult problem is that a polished, fluent response can make an error look trustworthy.
Teachers, tutors, and education businesses do not need to reject AI-assisted drafting to manage this risk. They do need a clear review process. The useful question before publication is: What could be wrong here, and how would a teacher detect it before a student relies on it?
This article offers concrete pre-publication tests for generated lessons and feedback. These tests support professional judgement: the teacher remains the author, reviewer, and final decision-maker.
Where hallucinations show up in educational materials
Hallucinations are not limited to invented facts. In education, failure can take several forms, each with a different consequence for learning and trust.
- Factual invention: a made-up historical event, scientific claim, quotation, source, rule, definition, statistic, or biographical detail.
- False precision: a response gives exact dates, percentages, page references, assessment requirements, or technical steps without a reliable basis.
- Broken reasoning: the final answer may be correct, but the explanation, calculation, proof, or worked example contains an invalid step.
- Misaligned instruction: a lesson is factually reasonable but is pitched at the wrong age, language level, curriculum sequence, or prior knowledge.
- Unreliable feedback: feedback identifies an error that is not present, overlooks a serious misconception, or recommends a revision that weakens the student’s work.
- Fabricated authority: the output uses invented citations or names a source in a way that makes the material appear verified when it is not.
- Unsafe simplification: the response presents a complex, sensitive, legal, medical, safeguarding, or personal issue as if one short answer applies in every case.
Many of these failures cannot be caught by checking spelling or tone. A reliable workflow needs tests for truth, reasoning, instructional fit, and assessment quality.
A practical pre-publication review sequence
Use a staged review rather than one broad instruction such as “check it carefully.” The sequence below is designed for materials that may be shared with students, parents, colleagues, or customers.
- Identify the educational claim. Mark statements students are expected to learn, use, remember, or repeat in an assessment.
- Test high-stakes content first. Prioritise answer keys, worked solutions, named sources, quotations, factual claims, instructions, and feedback that could change a learner’s decision or grade.
- Check instructional alignment. Compare the output with your objective, learner profile, curriculum, and success criteria.
- Run learner-facing tests. Try the task, questions, and feedback as a student would encounter them.
- Revise, document, and approve. Correct the material, remove unsupported claims, and retain final human ownership of what is published.
Test 1: The claim-and-evidence test
Take each important factual statement and ask: What evidence would justify this claim? If you cannot identify a dependable source, your own established subject knowledge, or a supplied course resource, do not publish the statement as fact.
This test is especially important when an output includes a confident explanation, named researcher, direct quotation, statistic, historical detail, or claim about a policy, qualification, or examination. AI-generated citations should not be treated as proof. Verify that a source exists, that it says what the material claims, and that it is appropriate for the lesson.
A simple classroom version is to highlight statements in three colours:
- Green: stable, familiar subject knowledge you can confidently verify.
- Amber: plausible but needs checking against a trusted resource.
- Red: exact figures, quotations, current rules, named studies, or claims with serious consequences if wrong.
For red claims, either verify them before publication or remove them. A useful revision is often more modest: replace an unsupported exact claim with a general explanation you can substantiate.
Test 2: The worked-example replay
For mathematics, science, grammar analysis, coding, essay modelling, and any topic with a process, independently replay every step. Do not only compare the final answer.
Ask a colleague or perform a “cold read” yourself after a short break: can the method be followed from the information given? Are all terms defined? Does each step follow from the previous one? Does the example accidentally introduce a new rule before teaching it?
For a generated model answer, test whether it genuinely meets the stated criteria. For a generated code example, run it in an appropriate environment before learners use it. For an analysis paragraph, check that the quoted evidence is accurate and that the commentary actually supports the conclusion.
A correct final answer does not make a flawed explanation safe to teach. Students may copy the method rather than only the result.
Test 3: The objective-to-activity alignment check
Generated lessons often contain attractive activities that do not directly teach or assess the intended objective. Put the objective at the top of your review and inspect every component against it.
| Lesson component | Question to ask |
|---|---|
| Starter | Does it activate relevant prior knowledge rather than introduce an unrelated puzzle? |
| Explanation | Does it teach the knowledge or skill named in the objective? |
| Practice | Must students use the target knowledge or skill to succeed? |
| Assessment | Would success provide credible evidence that the objective was met? |
| Extension | Does it deepen the same learning rather than merely add volume or difficulty? |
If an activity cannot be connected clearly to the objective, revise or remove it. This protects against a common failure mode: a lesson that feels busy, engaging, or well structured but leaves the key learning under-taught.
Test 4: The prerequisite and reading-level check
AI-generated materials can assume background knowledge learners do not have. They may also use vocabulary that is unnecessarily difficult for the audience. Review the material with two lists: what students must already know, and what new knowledge the lesson is intended to teach.
Look for unexplained terms, jumps in complexity, multi-step instructions, cultural references, and examples that rely on specialist knowledge. Then check whether the wording is suitable for the learners’ reading and language proficiency.
For each new term, decide whether to define it, illustrate it, provide an example, or replace it with more accessible language. For multilingual learners, also check whether an idiom, ambiguous pronoun, or figurative phrase could obscure the task.
Test 5: The question-quality stress test
Before publishing AI-generated questions, answer them yourself without relying on the proposed answer key. Then ask four questions:
- Is there one clearly best answer, where the format requires one?
- Does the question contain all information needed to answer it?
- Could an able learner reasonably interpret the wording in another way?
- Does the question assess the intended learning rather than a reading trick, obscure vocabulary, or irrelevant background knowledge?
For multiple-choice items, inspect every distractor. A weak distractor may be obviously absurd, accidentally correct, or based on a misconception students are unlikely to hold. For open questions, check that the mark scheme allows valid alternative responses and does not reward only one phrasing.
Also test the difficulty sequence. A set of questions should not claim to move from simple to complex while silently requiring advanced knowledge in the first item.
Test 6: The feedback fairness check
Generated feedback needs a higher standard than polite wording. It should accurately identify what the learner did, connect comments to visible evidence, and offer a realistic next step.
Use this checklist for each feedback template or individual response:
- Evidence: Does it refer to something actually present in the student’s work?
- Accuracy: Is the identified strength, error, or misconception real?
- Specificity: Does it explain what to improve, rather than only saying “add detail” or “be clearer”?
- Actionability: Can the learner take the next step independently or with the support available?
- Proportion: Does the tone and amount of correction match the task, learner, and purpose?
Try a deliberate counterexample: feed the feedback process a strong answer, a partially correct answer, and an answer with a common misconception. If the same generic praise or correction appears in all three cases, the feedback is not yet dependable enough to use without closer teacher editing.
Test 7: The source and quotation audit
Any generated bibliography, quotation, web reference, or named publication deserves separate verification. Check the author, title, publication details, quotation wording, and relevance. If you cannot confirm a source, do not include it.
For student-facing research tasks, give learners a verified resource list or clear source-selection criteria instead of publishing an unreviewed AI-generated list. This reduces the chance that students spend time pursuing non-existent or unsuitable materials.
Test 8: The boundary-case test
Test the material against cases at the edges: a learner who has misunderstood the key concept, an answer that is technically correct but expressed differently, a student who needs more language support, or an example where the usual rule has an exception.
Boundary cases reveal overconfident wording. If a rule only works “usually,” say so. If feedback cannot judge a nuanced response fairly, require teacher review. If a task may invite personal disclosure or touch sensitive topics, redesign it with appropriate professional safeguards and local procedures in mind.
Make review repeatable, not heroic
The safest workflow is not an exhausting line-by-line review of every low-risk sentence. It is a repeatable process that gives extra attention to content learners will rely on most. Create a short approval checklist for your subject, age range, and delivery format. Keep a record of recurring failure modes, such as invented quotations, ambiguous questions, missing prerequisites, or generic feedback. Then improve your prompts, templates, and review steps over time.
SubSchool can help teachers turn their approved ideas and content into structured course materials while keeping the final educational decision with the teacher. If you are building AI-assisted courses, explore the AI Course Creator and build your own review checkpoints into the publishing process.
Before publishing: verify important claims, replay the reasoning, align every activity to the objective, test questions independently, inspect feedback against the student work, and remove any source or statement you cannot support. AI can accelerate a first draft; professional review is what makes it suitable for learners.
Sources and methodology
{'approach': 'Reviewed the supplied draft as untrusted copy and searched for authoritative, first-party institutional sources that support its central claims: generative-AI risk management, human and pedagogical review in education, feedback quality, learning-objective alignment, and the SubSchool product statement.', 'source_selection': 'Prioritized government, intergovernmental, standards-body, evidence-organization, university teaching-center, and first-party product sources. Excluded vendor blogs, SEO articles, and secondary summaries where a stronger source was available.', 'limitations': 'The draft is primarily a practical editorial framework rather than a report of a single study. The sources support the need for review and several underlying instructional principles; they do not independently validate every named test, every failure category, or every implementation detail in the proposed checklist.'}
Use the relevant SubSchool workflow while keeping the result editable and teacher-reviewed.



