Unofficial guide · not affiliated with the AAMC, NBME, USMLE, LCME, or NRMP Part of the SI MedEd family · SI MedEd portal →

Elsie Item-Writing Academy

Academy › Item-Analysis Literacy › Shelf / course

Reading Shelf Item Statistics

About 20 minutes · Academy module: Reading Shelf Item Statistics

Reading Shelf Item Statistics

The Elsie Item-Writing Academy is an unofficial faculty-development resource, not affiliated with or endorsed by the AAMC, NBME, USMLE, LCME, or NRMP.

The numbers behind every item

Each administered item comes back with a small set of statistics. They work the same way at every exam level; the context here is 5-option shelf-style clerkship items — written for the pediatrics clerkship in this module — where random guessing lands at 0.20 and small cohorts make the small-N rule do real work.

N — the number of responses. Every statistic below is a fraction with N in the denominator, so read N first. A p-value from 14 responses is a rumor; the same p-value from 400 responses is evidence. When N is small, treat the numbers as a preview, not a verdict.

p-value — the proportion correct. Despite the name, this has nothing to do with significance testing; it is simply the fraction of examinees who selected the keyed answer. Near 0.20, the group performed at the level of random guessing. Near 1.00, nearly everyone answered correctly; near 0.00, almost no one did. Neither extreme is automatically a defect: a very easy item may be doing exactly the warm-up job you designed it for. Extremes deserve a look, not a reflex — check for a mis-key or a herding stem flaw.

Discrimination index — who got it right. The p-value tells you how many answered correctly; the discrimination index tells you which examinees. It compares the proportion correct among the highest and lowest scorers (commonly the top and bottom thirds or quarters). Near +1.0, the item cleanly separates stronger from weaker examinees; near zero, it carries almost no information about who knows the material. Example: 0.93 − 0.47 = 0.46 — positive and healthy.

Negative discrimination is a review flag — never a verdict. A negative index means lower-scoring examinees chose the keyed answer more often than higher-scoring ones. That pattern is a standing faculty-review flag: a mis-keyed answer, a second defensible correct option, or a subtle flaw misleading your strongest students. The flag starts an investigation; it never ends one. No item is deleted and no key is changed on a negative number alone — a faculty member reads the item and confirms the problem first.

Distractor pull — where the wrong answers went. For each option, the pull is the proportion of examinees who chose it. Healthy distractors each attract a real share — pulls like 0.08, 0.12, 0.10, and 0.10 around a key at 0.60. In pediatrics items, strong distractors are age-adjacent — the right diagnosis at the wrong age. A distractor near zero fooled nobody — revise it next use. A distractor that pulls harder than the key is the single most useful diagnostic in the set — it usually marks a mis-key, a second correct answer, or a stem flaw, and paired with a negative index it tells you exactly where to look first.

The small-N rule. Below a minimum number of responses, item statistics are "not enough data yet" — interesting, not actionable. Shelf-style items live near the minimum-N line longer than pre-clerkship items do; patience is part of the method. Whatever the threshold, an item is never revised, rekeyed, or retired on statistics alone until it clears it. Small samples manufacture dramatic numbers by chance.

In the Exam Admin builder. Pooled per-item statistics live in the builder's item review panel: one row per item showing N, p-value, discrimination index, and a small bar for each option's pull, with the keyed option marked. Items with negative discrimination carry a faculty-review flag — the flag routes the item to a human reader, and nothing is auto-deleted or auto-rekeyed. Opening an item puts the stem, options, and key beside the numbers. Kept this way, the review workflow supports faculty-development documentation: an evidence trail for every item you revise.

Worked example: a healthy 5-option pediatrics item

Illustrative data — invented numbers for practice, not real exam statistics.

A 2-year-old child presents with a barky cough, hoarseness, and inspiratory stridor that worsen at night. The child is afebrile with normal oxygen saturation. Which of the following is the most likely diagnosis?
A) Epiglottitis · B) Viral croup · C) Bacterial tracheitis · D) Foreign body aspiration · E) Peritonsillar abscess
Key: B
OptionChose itPull
A190.10
B (key)1330.70
C190.10
D120.06
E70.04
N190—

p-value = 133 / 190 = 0.70. Top-third correct = 0.93, bottom-third correct = 0.47, so discrimination = 0.93 − 0.47 = 0.46. Every distractor pulled a real share — note options A and C, the age-adjacent airway diagnoses, doing honest work. Read: a healthy, mid-difficulty item — keep it as written.

Drills

All tables below are illustrative data — invented numbers for practice, not real exam statistics.

Drill 1 — The distractor that beat the key

OptionChose itPull
A200.12
B250.15
C300.18
D600.35
E (key)350.21
N170—

p-value = 0.21. Top-third correct = 0.24, bottom-third correct = 0.42, so discrimination = 0.24 − 0.42 = −0.18.

Question: What action should the faculty member take?

Model answer: Flag the item for faculty review — do not auto-delete it and do not change the key. Distractor D out-pulled the key and the discrimination index is negative: the classic mis-key or second-correct-answer signature. The next step is to read the item itself and check whether D is defensible; only a confirmed reading justifies a rekey or revision.

Drill 2 — Ten responses

OptionChose itPull
A (key)70.70
B10.10
C10.10
D00.00
E10.10
N10—

p-value = 0.70. Discrimination cannot be meaningfully computed.

Question: What action should the faculty member take?

Model answer: None yet — "not enough data yet." With N = 10, from a small clerkship cohort, every number in this table could swing with the next rotation. Do not revise, rekey, or retire the item until it clears the program's minimum-N threshold.

Drill 3 — The very easy item

OptionChose itPull
A80.04
B60.03
C (key)1890.90
D40.02
E30.01
N210—

p-value = 0.90. Top-third correct = 1.00, bottom-third correct = 0.79, so discrimination = 1.00 − 0.79 = 0.21.

Question: The item is very easy and barely discriminates. What action should the faculty member take?

Model answer: Probably keep it. The discrimination index is weak but positive — the item is not misleading anyone — and if it covers must-know content (a can't-miss diagnosis every clerk should recognize), easiness is a feature. The question to ask is whether the item earns its slot on the exam; if it does, its statistics are fine as they are.

Takeaway checklist

  • Read N first; below the program minimum, the numbers are "not enough data yet" — clerkship cohorts get there slowly.
  • p-value is the proportion correct — near 0.20 on a 5-option item means guessing-level performance.
  • Discrimination tells you who got it right; a negative index is a faculty-review flag, never a reason to auto-delete or auto-rekey.
  • A distractor out-pulling the key, especially with negative discrimination, is the classic mis-key signature — verify by reading the item.
  • In the Exam Admin builder, review flags route items to a human — the numbers start the investigation; they never end it.

Finished this module?

Recording it adds the module to your Academy completion record on this device — module, track, level, and date, ready to download from your account page for faculty-development documentation.