How it works & impact

Feedback triggered by reasoning, not by a score gap.

Calibrate treats a rubric score as the tip of an argument. The student says where the work lands and why; the coach reads the why and pushes back on the specific reasoning move that went off — not the number.

The loop, at a glance

Six beats. Steps 1–5 are what one student does with the coach. Step 6 is what the class produces together — and it feeds the next cohort's exemplars.

01Student

Watch

An exemplar presentation from an earlier cohort.

02Student

Rate & justify

1–5 on four rubric criteria, with a short written justification each.

03Coach

Read the “why”

Each justification compared against the instructor's benchmark rationale; the misconception is named.

04Coach ↔ Student

A few questions, your replies

A bounded Socratic exchange — up to three questions per criterion, then a chance to re-rate.

05Student

Re-rate & reveal

Original, revised, and benchmark scores side by side, with the instructor's own comments.

06Instructor

Class-wide alerts

Justifications aggregate into ranked misconceptions with student quotes and a suggested mini-lesson.

↺ Alerts inform teaching — and next term's exemplars come from this data.

Before & after

What changes when the coach reads the reasoning, not the score.

01

What triggers feedback

Traditional

A number gap. The student sees “you over-rated this” and a canned explanation.

With Calibrate

The student's own written reasoning. The coach reads the justification and names the specific move that went wrong.

02

Shape of the exchange

Traditional

One-way. The student reads; the conversation ends.

With Calibrate

Socratic. Up to three questions per criterion with the student's replies, then an invitation to re-rate — bounded so it stays honest.

03

Instructor workload

Traditional

High. Someone has to pre-write feedback paths for every plausible score.

With Calibrate

Low. The instructor authors the rubric and one benchmark; the coach handles the variations.

04

What the analytics show

Traditional

Averages and variance. You know the class is off, not why.

With Calibrate

Ranked misconceptions drawn from justification text, each with anonymised student quotes.

Who gets what

Two sides of the same conversation.

For the student

From guessing to defending.

  1. 01

    Evaluative judgement

    Students practise deciding what quality looks like — on someone else's work first, then on their own.

  2. 02

    Feedback literacy

    Rubric language stops being decorative. Students have to use it to defend a score.

  3. 03

    Meta-cognitive awareness

    Writing a justification blocks the guess-and-move-on reflex. The reasoning becomes visible — first to the student, then to the coach.

For the instructor

From averages to alerts.

  1. 01

    Actionable interventions

    Instead of a general read of the class, you get named misconceptions — “delivery confidence read as evidence quality” — with the student quotes behind them. You know exactly what to teach next.

  2. 02

    A diagnostic mirror for the rubric

    If a criterion mis-calibrates across the whole class, that's a signal about the rubric, not just the students. The dashboard makes ambiguity in your own language visible.

  3. 03

    Dialogic feedback at scale

    Carless and Boud describe feedback as a dialogue. Doing that with 100+ students is impossible by hand. The coach carries the first pass so instructors can step in where nuance actually needs them.

The five-minute test

Rate one presentation. Let the coach hear you.

The fastest way to understand the loop is to run it. Pick an exemplar, rate it honestly, and see what the coach names in your reasoning.