Free teacher assessment protocol

Business simulation grading and calibration guide

Make scoring more consistent across students, teams, class periods, and graders without turning a simulated profit ranking into the grade.

Print the calibration record Use the 30-minute protocol

How do teachers grade business simulations consistently?

Start with the learning target and common evidence, not the winning score. Choose a small set of anonymous anchor responses, score them independently with the same rubric, compare the evidence behind each rating, and record how the team will interpret difficult boundaries. Recheck one anchor partway through grading to detect drift.

Calibrate the evidence

Agree what counts as a comparable run, correct calculation, supported explanation, meaningful tradeoff, and responsible recommendation.

Separate the sources

Mark shared team evidence once and individual reasoning separately. Device control, speaking confidence, and profit rank are not substitutes for understanding.

Audit the pattern

Review borderline decisions, criterion averages, missing evidence, and accommodations before releasing grades and actionable feedback.

Prepare a scoreable evidence set

  1. Name the target. State the business reasoning students must demonstrate: define a decision, compare controlled evidence, calculate accurately, explain a mechanism, weigh a tradeoff, and recommend a bounded next action.
  2. Publish the evidence map. Identify which items are team evidence—such as the run log—and which are individual evidence—such as a calculation explanation or exit defense. State how missing or inaccessible evidence can be completed.
  3. Use a shared rubric. The business simulation rubric provides decision/evidence, analysis, responsibility, and communication criteria. Remove criteria that were not taught or observable in the task.
  4. Select anonymous anchors. Choose a strong response, a developing response, and a boundary response. Remove names and irrelevant personal information. Do not choose anchors because they confirm a prior impression of a student.
  5. Mark non-negotiable constraints. A recommendation cannot earn the highest reasoning rating when it relies on unlawful discrimination, unsafe practice, privacy intrusion, deceptive advertising, hidden sponsorship, or another prohibited action.

Calibrate only evidence students had a fair opportunity to produce. If the task never required a repeat run or model-limit statement, do not quietly add it during grading.

Use the SCALE calibration protocol

StepMinutesWhat graders doRecord
S — Set the target3Read the task and rubric; name the observable evidence and shared/individual split.Target, criteria, exclusions
C — Check anchors alone7Score three anonymous samples independently. Highlight evidence before assigning a level.Score plus one evidence phrase per criterion
A — Align interpretations10Compare one criterion at a time. Discuss differences of more than one level first.Boundary rule and final anchor score
L — Lock decision rules5Write how to treat missing units, mixed variables, group evidence, resubmission, and responsible constraints.Short scoring notes
E — Examine drift5After 8–10 papers, rescore one anchor or swap one real anonymous paper and compare again.Keep or revise the shared rule

Single-teacher version: score three anchors, wait until the next day or change their order, then score them again without viewing the first ratings. Investigate any shift. The purpose is not perfect numerical agreement; it is a defensible and repeatable link between evidence and criterion.

Resolve common scoring boundaries

BoundaryConsistent interpretationUseful feedback
High profit, weak comparisonDo not infer reasoning from the outcome. Score the missing baseline, controls, or explanation as missing evidence.“Repeat with one main change and compare the same measures.”
Correct calculation, no unitsGive calculation credit according to the rubric, but do not award full interpretation credit until the unit and period are clear.“Label this as dollars per month and explain what it changes.”
Good team product, weak individual checkKeep the shared artifact score separate. Use a brief follow-up rather than assigning every member the strongest speaker's understanding.“Explain one input, result, and tradeoff in your own words.”
Polished writing, unsupported claimCommunication quality cannot replace missing evidence. Score criteria independently.“Add the baseline/test values that support this recommendation.”
Strong evidence, harmful recommendationRecognize the analysis achieved, then apply the responsibility criterion and require redesign within legal, safety, privacy, and fairness boundaries.“Keep the evidence; revise the action so this duty is protected.”
Equivalent accessible responseScore the same target whether evidence is written, spoken, visual, teacher-scribed, or produced with an approved tool.Respond to reasoning, not the response mode.

Worked calibration case: café staffing

Student response: “The café should add a second worker. In the baseline it served 76 customers, profit was $218, and wait time was 10 minutes. With another worker it served 91, profit was $224, and wait fell to 6 minutes. Profit increased 2.8%, customers served increased 19.7%, and wait fell 40%. We kept price and promotion constant. One run cannot show whether the improvement repeats, and labor cost could become too high when demand is lower. Repeat the staffing-only test for three comparable months; stop the change if average profit falls below the baseline or quality declines.”

Calibration discussion: The response identifies a decision, provides baseline and test values, uses correct percentage change, names controls, notes uncertainty, and gives a monitored recommendation. It should not lose points because the profit increase is small: the student interprets that tradeoff. A top responsibility score would still depend on the assigned criterion—such as lawful scheduling, safe workload, or customer accessibility—being addressed in the full response.

Boundary variation: If the student wrote only “profit rose 2.8%, so hire,” the calculation may be correct, but the explanation, service evidence, uncertainty, and stop rule are absent. Graders should award what is observable without filling gaps from the stronger team record unless the task explicitly defines that record as shared evidence.

Audit fairness before returning grades

This classroom guide is not legal or institutional policy. Follow applicable school, district, accessibility, privacy, record-retention, and grading requirements. Calibration supports professional judgment; it does not automate it.

Business simulation grading FAQ

What is grading calibration for a business simulation?

Grading calibration is a short process in which teachers apply the same criteria to anonymous sample work, compare evidence and scores, resolve differences using the rubric, and record shared interpretations before grading the full set.

How many anchor samples are needed for calibration?

Three samples are usually enough for a classroom task: one clearly strong response, one developing response, and one boundary case that could reasonably receive adjacent scores. Anchors should represent reasoning quality rather than writing polish alone.

Should simulated profit determine a student's grade?

No. Profit is one result to interpret. Grade the decision question, comparable evidence, calculations, explanation, tradeoffs, responsible constraints, and revision. A weak simulated outcome can support excellent analysis.

How can teachers grade team simulations fairly?

Score the shared run record or team product once, then verify individual understanding with a brief explanation, calculation, reflection, or conference. Publish which evidence is shared and which is individual before students begin.

What should happen when two graders disagree?

Return to the exact criterion, cite observable evidence, identify the performance-level boundary, and agree on a score plus rationale. If evidence is missing, use not yet demonstrated rather than guessing from effort, reputation, or unrelated polish.

Connect calibration to the assessment cycle

Start with the assessment and feedback toolkit, score with the printable rubric, verify individual reasoning through the conference toolkit, and retain revisions in the evidence portfolio. Browse the teacher hub or complete resource index for a full lesson route.