Unofficial guide · not affiliated with the AAMC, NBME, USMLE, LCME, or NRMP Part of the SI MedEd family · SI MedEd portal →

Elsie Item-Writing Academy

Academy › Item-Analysis Literacy › MCAT

Reading MCAT Item Statistics

About 20 minutes · Academy module: Reading MCAT Item Statistics

Reading MCAT Item Statistics

The Elsie Item-Writing Academy is an unofficial faculty-development resource, not affiliated with or endorsed by the AAMC, NBME, USMLE, LCME, or NRMP.

The numbers behind every item

Each administered item comes back with a small set of statistics. They work the same way at every exam level; the context here is 4-option MCAT-style items, where random guessing lands at 0.25.

N — the number of responses. Every statistic below is a fraction with N in the denominator, so read N first. A p-value from 14 responses is a rumor; the same p-value from 400 responses is evidence. When N is small, treat the numbers as a preview, not a verdict — the small-N rule below tells you where that line sits.

p-value — the proportion correct. Despite the name, this has nothing to do with significance testing; it is simply the fraction of examinees who selected the keyed answer. Near 0.25, the group performed at the level of random guessing. Near 1.00, nearly everyone answered correctly; near 0.00, nearly no one did. Neither extreme is automatically a defect: a very easy item may be doing exactly the warm-up job you designed it for. Extremes deserve a look, not a reflex — check for a mis-key or a stem flaw that herded everyone toward the same wrong option.

Discrimination index — who got it right. The p-value tells you how many answered correctly; the discrimination index tells you which examinees. It compares the proportion correct among the highest scorers with the proportion correct among the lowest (commonly the top and bottom thirds or quarters). Near +1.0, the item cleanly separates stronger from weaker examinees; near zero, it carries almost no information about who knows the material. Example: 0.92 − 0.44 = 0.48 — positive and healthy.

Negative discrimination is a review flag — never a verdict. A negative index means lower-scoring examinees chose the keyed answer more often than higher-scoring ones. That pattern is a standing faculty-review flag: a mis-keyed answer, a second defensible correct option, or a subtle flaw misleading your strongest students. The flag starts an investigation; it never ends one. No item is deleted and no key is changed on a negative number alone — a faculty member reads the item, confirms the problem, and only then decides what changes.

Distractor pull — where the wrong answers went. For each option, the pull is the proportion of examinees who chose it. Healthy distractors each attract a real share — pulls like 0.12, 0.18, and 0.10 around a key at 0.60. A distractor near zero fooled nobody and is a candidate for revision next use. A distractor that pulls harder than the key is the single most useful diagnostic in the set — it usually marks a mis-key, a second correct answer, or a stem flaw, and paired with a negative index it tells you exactly where to look first.

The small-N rule. Below a minimum number of responses, item statistics are "not enough data yet" — interesting, not actionable. Programs set their own minimums; whatever the threshold, an item is never revised, rekeyed, or retired on statistics alone until it clears it. Small samples manufacture dramatic numbers by chance, and acting on them is how good items get broken.

In the Exam Admin builder. Pooled per-item statistics live in the builder's item review panel: one row per item showing N, p-value, discrimination index, and a small bar for each option's pull, with the keyed option marked. Items with negative discrimination carry a faculty-review flag — the flag routes the item to a human reader, and nothing is auto-deleted or auto-rekeyed. Opening an item puts the stem, options, and key beside the numbers, so reviews start from evidence, not memory. Kept this way, the review workflow supports faculty-development documentation: an evidence trail for every item you revise.

Worked example: a healthy 4-option psych/soc item

Illustrative data — invented numbers for practice, not real exam statistics.

Researchers find that individuals in a group setting exert less effort on a shared task than when working alone. This pattern best illustrates:
A) Social facilitation · B) Social loafing · C) Deindividuation · D) Group polarization
Key: B
OptionChose itPull
A240.12
B (key)1320.66
C280.14
D160.08
N200—

p-value = 132 / 200 = 0.66. Top-quarter correct = 0.92, bottom-quarter correct = 0.44, so discrimination = 0.92 − 0.44 = 0.48. Every distractor pulled a real share. Read: a healthy, mid-difficulty item — keep it as written.

Drills

All tables below are illustrative data — invented numbers for practice, not real exam statistics.

Drill 1 — The distractor that beat the key

OptionChose itPull
A180.11
B220.14
C (key)580.36
D620.39
N160—

p-value = 0.36. Top-third correct = 0.30, bottom-third correct = 0.55, so discrimination = 0.30 − 0.55 = −0.25.

Question: What action should the faculty member take?

Model answer: Flag the item for faculty review — do not auto-delete it and do not change the key. The pattern is the classic mis-key signature: negative discrimination plus a distractor (D) that out-pulled the key. The next step is to read the item itself and check whether D is also defensible or the key is simply wrong; only a confirmed reading justifies a rekey or revision.

Drill 2 — Eleven responses

OptionChose itPull
A20.18
B (key)90.82
C00.00
D00.00
N11—

p-value = 0.82. Discrimination cannot be meaningfully computed.

Question: Options C and D pulled zero. What action should the faculty member take?

Model answer: None yet — "not enough data yet." With N = 11, the zeros for C and D are a thing to watch, not a thing to act on. Do not revise the distractors or retire the item until it clears the program's minimum-N threshold; small samples produce dramatic-looking zeros by chance.

Drill 3 — The very easy item

OptionChose itPull
A (key)2280.95
B60.03
C40.02
D20.01
N240—

p-value = 0.95. Top-quarter correct = 0.99, bottom-quarter correct = 0.88, so discrimination = 0.99 − 0.88 = 0.11.

Question: The item barely discriminates. What action should the faculty member take?

Model answer: Being easy is not a defect. If the item was designed as an early confidence-builder or covers must-know content, keep it — its discrimination is weak but positive, so it is not misleading anyone. The optional, evidence-based move is to strengthen distractors B, C, and D, which pulled almost nothing; the item to avoid is deleting a sound item for the crime of being easy.

Takeaway checklist

  • Read N first; below the program minimum, the numbers are "not enough data yet."
  • p-value is the proportion correct — near 0.25 on a 4-option item means guessing-level performance.
  • Discrimination tells you who got it right; a negative index is a faculty-review flag, never a reason to auto-delete or auto-rekey.
  • Every distractor should pull a share; a distractor out-pulling the key is your best clue to a mis-key or a second correct answer.
  • In the Exam Admin builder, review flags route items to a human — the numbers start the investigation; they never end it.

Finished this module?

Recording it adds the module to your Academy completion record on this device — module, track, level, and date, ready to download from your account page for faculty-development documentation.