How to Measure an Exam-Prep Course by Improvement, Not Testimonials
Testimonials can show that learners valued a course, but they cannot show how much learners improved or for whom it worked. A more credible approach tracks comparable baseline and exit evidence, reports cohort context, accounts for missing data, and makes claims that match the evidence.

“My score went up by 200 points.” “This course changed everything.” “My tutor got me into my first-choice school.” Testimonials can be encouraging, and they may help a prospective family understand a learner’s experience. But they are not a dependable measure of exam prep course outcomes.
They are usually selected because they are positive, memorable, and easy to share. They rarely tell a parent, student, teacher, or school what the learner’s starting point was, how much time they studied, whether they completed the course, which assessment was used, or what happened for learners who did not submit feedback. A course can collect genuine praise and still have little evidence about average improvement across a whole cohort.
A stronger standard is straightforward: measure progress from a meaningful baseline to a comparable exit assessment, then publish the context needed to interpret that progress honestly. This does not require a large research department. It requires a planned routine, consistent records, and restraint about what the data can prove.
Start with the question you can actually answer
Before collecting data, decide what decision the measurement should support. “Did the course work?” is too broad. A practical course evaluation asks narrower questions, such as:
- How did enrolled learners perform on an aligned diagnostic before instruction and on an aligned exit assessment after instruction?
- What proportion of learners met a clearly defined improvement goal?
- Which skill areas improved most or least?
- How did results differ by starting level, attendance, completion, or course format?
- How many learners have complete data, and who is missing from the analysis?
The U.S. Institute of Education Sciences describes program evaluation as a structured process for assessing implementation and outcomes. Its evaluation resources also emphasise planning the measures, baseline information, outcomes, and analytic sample rather than treating results as an afterthought. That is a useful model even for a small tutoring programme or a single exam-prep cohort.
For most providers, the first goal is not to prove that the course alone caused a score increase. It is to describe learner progress in the course clearly enough to improve teaching and to help families make informed decisions. Causal claims require stronger designs, especially a credible comparison group and attention to baseline equivalence.
Build a baseline that is genuinely useful
A baseline is more than a first-week quiz. It is the evidence used to understand where learners began before the course’s main instruction could influence their performance. Without it, an exit score is difficult to interpret. A learner scoring highly at the end may have begun highly; another learner with a lower exit score may have made much larger progress.
Use a baseline that is:
- Aligned: It samples the knowledge, skills, and question formats the course is designed to teach.
- Appropriately timed: It is completed before or at the start of instruction, not after several weeks of teaching.
- Consistently administered: Learners receive similar instructions, time limits, permitted resources, and scoring rules.
- Recorded at skill level: Keep an overall score, but also retain results by domain, such as algebra, reading inference, grammar, essay structure, or question type.
- Low-stakes enough for honest effort: Explain that the purpose is placement and planning, not judgement.
Where possible, use an assessment that resembles the exit assessment in content coverage and level of difficulty without simply repeating the same items. Reusing identical questions can blur the distinction between durable learning and remembering answers. If the exact same assessment is the only practical option, say so in the report and interpret gains cautiously.
Baseline data should change instruction. The What Works Clearinghouse guidance on using student achievement data frames data use as an ongoing cycle of instructional improvement. In exam preparation, that might mean grouping learners for a targeted workshop, revising a weak unit, assigning a practice set to address a common misconception, or helping a student set a specific goal.
Use an exit measure that matches the baseline
The exit measure should answer the same broad question as the baseline: what can this learner now do on the skills the course targeted? Comparability matters more than perfection. A 90-minute diagnostic at the start and a five-question quiz at the end are not a sound basis for a headline outcome claim.
Document the assessment conditions for both points in time. At a minimum, retain the test name or version, date, time allowance, scoring approach, whether it was supervised, and any material differences in administration. If the course prepares learners for a specific external examination, distinguish carefully between your internal aligned assessment and the official exam. Do not imply that an internal mock score is an official result.
Report more than one outcome where it helps interpretation:
- Average baseline score and average exit score.
- Average within-learner score change among learners with both scores.
- Median change, which can be useful when a small number of unusually large gains distort the average.
- The number and percentage of learners who improved, stayed broadly similar, or declined.
- Results by skill domain, not only one composite score.
- Completion, attendance, or participation measures reported separately from attainment.
Averages alone can hide an uneven picture. For example, an average gain may be driven by a small group of highly engaged learners. A simple distribution of individual changes, or a table showing how many learners fall into sensible improvement bands, gives families and teachers a more complete view.
Give every cohort its context
Comparing course outcomes across cohorts can be valuable, but cohorts are rarely identical. One group may contain many learners starting below the target level; another may include students who have already received extensive tutoring. One cohort may meet twice weekly for ten weeks, while another follows an intensive holiday format.
Publish enough context to stop a single number becoming misleading. A concise outcomes summary could include:
| Report element | Why it matters |
|---|---|
| Number enrolled, completed, and assessed at exit | Shows whose results are represented. |
| Course dates, teaching hours, and delivery format | Helps readers compare like with like. |
| Baseline score distribution or starting bands | Shows whether learners began at similar levels. |
| Assessment used at baseline and exit | Clarifies what the score change represents. |
| Average and median within-learner change | Prevents one summary statistic from carrying the whole story. |
| Missing-data count and reason, where known | Shows the limits of the analysis. |
| Important delivery changes | Flags changes in teacher, curriculum, timetable, or assessment conditions. |
Segment results only when the group sizes are large enough to be meaningful and when the segment serves a real instructional or decision-making purpose. For example, reporting outcomes by baseline band can help reveal whether a course supported learners who started furthest behind. Avoid publishing small subgroups in ways that could identify a learner or create unwarranted conclusions from a handful of cases.
Do not hide missing data
Missing exit assessments are not an administrative nuisance; they can change the story. Learners may miss an assessment because they withdrew, were absent, ran out of time, changed plans, or did not complete the course. If those learners differ systematically from learners with recorded exit scores, reporting only the complete cases can overstate or understate apparent improvement.
The What Works Clearinghouse notes that missing data can introduce bias when people with missing information differ from those with observed information. Its standards also treat missing baseline and outcome data as an important analytic issue. A local course report does not need to replicate a formal research review, but it should adopt the same habit of transparency.
Use this minimum reporting practice:
- State the number enrolled.
- State how many completed the baseline.
- State how many completed the exit assessment.
- State how many learners are included in the paired baseline-to-exit calculation.
- Give known reasons for missing exit data in broad categories, such as absence, withdrawal, or incomplete assessment.
- Do not silently replace missing scores with favourable assumptions.
If many learners are missing an exit score, say that the result applies only to learners with complete paired data and may not represent all enrolled learners. That sentence may feel less promotional, but it is more useful. It also helps a teaching team identify operational problems: perhaps the final assessment is scheduled too late, is too burdensome, or is not consistently followed up.
Make claims proportionate to the design
There is an important difference between observed improvement during a course and proof that the course caused the improvement. Learners can improve for many reasons: classroom teaching, independent study, other tutoring, familiarity with the assessment, changes in motivation, or normal development over time.
Without a well-matched comparison group or a stronger evaluation design, use language such as:
Among the 42 learners with both baseline and exit assessments, the median score increased from 48% to 61% on our aligned internal assessment. Results describe progress for this assessed group and do not isolate the course’s effect from other learning experiences.
Avoid language such as “our programme raises official exam scores by 13 points” unless you have reliable official-score data, a clearly documented method, and evidence that supports a causal interpretation. Evaluation standards place weight on appropriate outcome measures, baseline information, analytic samples, attrition, and comparisons because these details affect what an outcome estimate means.
Testimonials still have a place. Label them as learner or parent experiences, obtain appropriate permission, and do not place them where readers could mistake them for cohort-level evidence. A useful page can present both: a clearly labelled outcomes table and a separate selection of personal stories.
Create a repeatable outcomes routine
The most credible reporting is usually the result of a modest, repeatable workflow rather than an occasional campaign. At course design stage, define the outcomes, baseline, exit assessment, completion rules, data owner, and reporting date. During delivery, record attendance and major implementation changes. After the course, calculate results from a locked cohort list, review anomalies, document missing data, and share a short internal findings note before creating public-facing copy.
This routine also protects teacher judgement. Software can reduce repetitive work such as organising learner records, tracking task completion, preparing review materials, or assembling draft reports. But teachers and course leaders should remain responsible for deciding what counts as meaningful improvement, checking the assessment quality, interpreting exceptions, and approving every public claim.
If you are building a clearer learning record for a family, start with one cohort and one aligned baseline-to-exit measure. Then use the findings to improve the next course rather than to manufacture a headline. Explore SubSchool for parents to see how a more organised learning workflow can support communication while educators retain authorship and the final educational decision.
A practical publication checklist
- Have we named the cohort, dates, format, and number enrolled?
- Have we shown baseline and exit assessment conditions?
- Have we reported paired results, not just the strongest final scores?
- Have we disclosed how many learners lack one or both scores?
- Have we separated participation, completion, internal assessment results, and official exam results?
- Have we avoided claiming causation where we only observed change?
- Have a teacher or assessment lead reviewed the interpretation before publication?
When course providers measure improvement honestly, they gain more than a marketing asset. They gain a better basis for adapting teaching, discussing progress with learners and families, and deciding where the course needs to improve next.
Sources and methodology
Prepared as an evidence-aware thought-leadership article using the supplied editorial brief and publicly available Institute of Education Sciences/What Works Clearinghouse materials on programme evaluation, student data use, baseline information, missing data, analytic samples, and evaluation design. The article intentionally distinguishes descriptive within-cohort progress from causal impact claims and does not make claims about any particular examination, provider, score gain, or SubSchool feature beyond the supplied positioning that it automates repetitive teaching work while educators retain authorship and final educational decisions.
Use the relevant SubSchool workflow while keeping the result editable and teacher-reviewed.



