Exam statistics and per-item analysis. Flags are review signals for faculty judgment — the system never deletes items, changes keys, or adjusts scores on its own.
No attempts yet for this exam. Take the exam via the Take Exam page (or use the demo exam, which ships with simulated responses) and analytics will appear here.
Exam summary
Score distribution
Percent-correct bins across attempts.
Items needing faculty review
Most commonly missed
Ranked by p-value (lowest = missed most). Use with discrimination: a hard item that discriminates well is doing its job.
Rank
Item
Topic
p-value
Difficulty
Discrimination (rpb)
Full item analysis
Click a column header to sort. Expand an item for distractor detail.
Item
Topic
n
p-value
Difficulty
D (U–L)
Point-biserial
Band
Flags
Methodology & caveats
p-value = proportion correct. Discrimination (U–L) = p(top 27% of scorers) − p(bottom 27%). Point-biserial = Pearson correlation between item score (0/1) and total score.
Cronbach's alpha (KR-20 for scored items) estimates internal consistency; 0.70+ is generally acceptable for classroom exams.
A distractor selected by fewer than 5% of students is flagged as non-functional.
Negative discrimination never triggers automatic action. It is flagged for faculty review: check the key, the wording, whether the content was taught, and whether a strong distractor reflects a real misconception worth teaching to.
Statistics are sample-dependent: small cohorts, retakes, and mixed-ability groups all shift the numbers. Interpret alongside content review, not instead of it.
Demo data is simulated (seeded generator) and labeled simulated wherever it appears.