Unofficial guide · not affiliated with the AAMC, NBME, USMLE, LCME, or NRMP Part of the SI MedEd family · SI MedEd portal →

Benchmarking responsibly.

National data is a flashlight, not a verdict. How admissions and curriculum committees choose comparison groups, keep denominators honest, respect small cohorts, and separate signal from noise.

Unofficial guide. This site is independent — not affiliated with the AAMC, NBME, USMLE, LCME, or NRMP. This page is general methodological guidance for committee work; it carries no school-specific data and is not accreditation advice (see SILME for that).

1. Choose the right comparison group

"Compared to what?" is the first question every benchmark should answer. The national average is a useful anchor — but it is not always the right one:

  • National figures (the admissions funnel and outcome trends on this site) are the common currency: everyone understands them, and they are always available.
  • Mission-aligned peers — schools with similar missions, student populations, and resources — often tell you more about what is achievable for your context than the national mean does.
  • Aspirational peers have a place in strategic planning, but they are goals, not benchmarks. Judging routine committee work against aspirational targets manufactures perpetual failure.

Whatever you choose, name it explicitly on every chart: "vs. national, 2023–2024" beats an unlabeled line every time.

2. Denominator discipline

Most benchmark errors are denominator errors. Keep these pairs straight:

  • First-time vs. eventual pass rates. First-attempt figures reflect preparation; eventual figures (after retakes) reflect remediation. Report both, labeled, and never let one stand in for the other.
  • Applicants vs. acceptees vs. matriculants. An "acceptance rate" built on the wrong population is not an acceptance rate at all. Define each stage of the funnel on the chart itself.
  • Unique people, not records. Denominators are people who could experience the outcome — a student who never sat the exam is not part of the pass-rate denominator. Survey denominators are the people who answered, not the number of evaluations collected.
  • Blanks are not zeros. Missing data is missing, not evidence of failure. Exclude blanks from denominators; never code them as zero outcomes.
  • Separate academic years. Combining cohorts across years can hide a trend — the Step 1 series on the outcome trends page shows how one pooled number can obscure a regime change.

3. Small-cohort caution

Medical school cohorts are small — often dozens, not thousands — so percentages are volatile by construction. A class of 60 where three students fail a first-attempt exam reports a 95% pass rate; five failures report 92%. The "three-point drop" is two people.

  • Report counts alongside rates. "57 of 60 passed (95%)" tells the reader what the percentage cannot: how many people moved the number.
  • Roll multi-year windows. Three-year rolling averages smooth cohort noise and reveal real drift. React to the window, not the year.
  • Flag, don't verdict. A defensible working rule: values more than one standard deviation from the comparison mean get flagged as indicators for investigation — alongside their sample sizes — never as automatic findings. A flag starts a conversation; it does not end one.
  • Resist the single-cohort reaction. Curriculum changes made in panic after one unusual year usually get reversed after the next unusual year. Wait for the trend.

4. Separate signal from noise

  • Expect year-to-year wobble. A point or two of movement in a rate is normal variation. Sustained, directional movement across three or more years is a signal.
  • Check the regime before the trend. Methodological changes (Step 1 going pass/fail in 2022, a passing standard raised, a data-collection definition changed) break comparability. Benchmark within the comparable series.
  • Triangulate. A dip in first-time pass rates means more if shelf-exam means and clerkship performance moved the same direction. One metric is a rumor; three are a story.
  • Ask what changed in the input. Before attributing an outcome shift to curriculum, check whether the entering cohort's credentials moved (see the funnel distributions). Weaker inputs predict weaker outputs — that is measurement, not failure.

The golden rule of committee data: a benchmark exists to prompt a question, not to deliver an answer. "We sit two points below the national first-time rate, and the gap has narrowed for three years" is a finding. "We are below average" is a feeling. Build charts that produce the first, not the second.

5. Presenting benchmarks to leadership

Leadership slides need less, not more:

  • One comparison group, named on the chart. One denominator, defined in the footnote.
  • Trends over snapshots: show the last five years, not just the last one.
  • Rates with counts: "92% (55 of 60)" — never a bare percentage.
  • Year and source on every national figure, so the slide still reads correctly when it resurfaces in next year's deck.
  • The question the data raises and the next step — benchmarks that end with "so what?" invite someone else to supply the answer.

Official sources to bookmark

National publishers update on their own schedules: AAMC FACTS each fall, USMLE Performance Data annually, NRMP Charting Outcomes roughly every two years. Confirm the latest edition on the official site before citing in committee documents.