Unofficial guide · not affiliated with the AAMC, NBME, USMLE, LCME, or NRMP ← SI MedSchool · Exam Admin

Exams, from bank to item analysis.

Faculty build practice exams from the original SI MedEd item bank — now including the full 300-item Step 1 pilot bank across eight topics — students take them (timed or untimed), and the platform produces full exam statistics and item analysis — difficulty, discrimination, distractor performance, most-missed rankings. Everything here is original work; no commercial question banks, no proprietary content.

Unofficial, independent tool. Not affiliated with the AAMC, NBME, USMLE, LCME, or NRMP. Demo data shown is simulated for illustration. Analytics flags are review signals for faculty judgment — items are never auto-deleted or auto-revised by the system.

Your exams

Draft and published exams in this browser. Publish an exam to make it available to students.

ExamItemsTime limitStatusAttempts

How it works

1. Build

Pick items from the original bank (or author your own), set a time limit and pass mark, and publish. ~90 seconds per item is the standard pacing.

2. Deliver

Students take the exam in timed or untimed mode, with instant scoring and explanations on review.

3. Analyze

Get exam statistics (reliability, score distribution) and per-item analysis: p-values, discrimination, distractor performance, most-missed rankings — with flags for faculty review, never automatic deletion.

New — Question Studio. Students and faculty can create their own original questions with Elsie: she drafts each item, then two independent checkers solve it cold before it can be saved — the same triple-check behind the SI MedEd banks. Disputed items come back flagged with reasons, never silently shipped. Open the Question Studio →

What the big platforms do — our version. Commercial systems (MedHub, One45, Blackboard Learn, NBME's assessment services) all provide the same core loop: exam assembly, delivery, and psychometric reporting (item difficulty, discrimination indices, distractor analysis, reliability coefficients, cohort comparisons). This module implements that loop with original code and original items, on open methodology documented below.

Methodology, in plain language

  • Item difficulty (p-value): the proportion of students answering correctly. 0.30–0.80 is usually most informative.
  • Discrimination index (upper–lower): how much better the top 27% of scorers did than the bottom 27% on the item. Positive is good; negative is a warning.
  • Point-biserial: the correlation between getting the item right and total exam score. Above 0.30 is good, below 0.10 is poor.
  • Distractor analysis: who picked each wrong option. A distractor chosen by nobody is dead weight; one that attracts top scorers more than the key suggests a possible miskey.
  • Cronbach's alpha (KR-20): internal consistency of the whole exam. 0.70+ is generally acceptable for classroom exams.
  • Flags are review signals, not verdicts. A negatively discriminating item is flagged for faculty review — the system never deletes items or changes keys on its own.

Demo data is simulated with a seeded generator (see js/sample-data.js) and labeled as such everywhere it appears.