How the predictor works and how accurate it is
Enter at least two practice-test scores, qbank accuracies, CMS averages, or outside predictions to get an estimated Step 2 CK score and prediction range. Several recent assessments generally improve reliability and narrow the range. Accuracy was measured on 367 held-out score reports that were not used to train the predictor.
From practice scores to one estimate
- Enter at least two scores. Use any combination of practice-test results, qbank accuracy, CMS average, or an outside prediction. Test-taker status, qbank completion, assessment timing, and study duration can add context, but they do not count toward the two-score minimum.
- Put different tests on a common scale. Each NBME, UWSA, Free 120, and AMBOSS self-assessment has its own relationship with final Step 2 scores.
- Account for timing. A score from two days before the exam usually says more about exam-day performance than a score from two months earlier. Adding the date or days before your exam helps the calculator make that distinction.
- Combine what you enter. When you provide more than one score, the predictor uses them together rather than relying on a simple average.
- Show a range as well as a point estimate. The range reflects how much reported outcomes varied among test-takers with comparable inputs. Choosing a higher percentage makes the range wider because it is meant to include more possible outcomes.
The calculation runs in your browser. The scores you enter are not uploaded.
How accurate is it?
Accuracy was measured on 367 held-out score reports from 2023 through 2026 that were not used to train the predictor. Each report had at least two usable score inputs.
- Estimates were 4.8 points from the real score on average (mean absolute error, or MAE).
- For 54.8% of people, the estimate was within 4 points.
- For 83.7% of people, the estimate was within 8 points.
- Half of the estimates were within 3.6 points.
- R2 was 0.68, meaning the estimates followed much of the person-to-person difference in final scores.
The displayed ranges are calibrated for one prediction at a time when at least two practice assessments are entered. A 67% range is intended to contain the real score for about 67 of 100 comparable predictions. The revised widths still need a new held-out check.
USMLE reports a standard error of about 7 points for Step 2 CK, so the same student can receive a score several points higher or lower on another administration. Read the 4.8-point average error together with the displayed prediction range.
More practice scores, especially recent ones, generally reduce error and narrow the range.
Predictor comparison
This table is based on held-out testing data that was not used to train the predictor.
This predictor had the best result in every comparable product column, including the lowest average error at 4.5 points. Lower error values are better; higher R2 and within-range percentages are better. Best results are bold.
| Predictor | Average error | R2 | Withinpoints2 | 90% withinpoints2 | Tailerror2 | Tendency2 | |
|---|---|---|---|---|---|---|---|
| 4 | 8 | ||||||
| This predictor | 4.5 | 0.69 | 65% | 86% | 10.0 | 6.9 | No clear bias (+0.15) |
| “Latest 3 average, shrinking bonus” rule14 | 4.9 | 0.62 | 54% | 85% | 10.0 | 7.6 | No clear bias (+0.07) |
| StepGunner | 5.1 | 0.62 | 54% | 83% | 11.0 | 7.3 | Underpredicts (+1.6) |
| “Highest NBME + 1” rule1 | 5.9 | 0.46 | 48% | 76% | 12.0 | 7.1 | Underpredicts (+1.1) |
| AMBOSS Score Predictor3 | 6.1 | 0.46 | 43% | 75% | 12.0 | 8.3 | Underpredicts (+3.5) |
| CoreStepPrep | 9.4 | −0.05 | 23% | 49% | 17.0 | 8.4 | Underpredicts (+8.6) |
| NBMEcalc.com | 9.6 | −0.01 | 17% | 44% | 16.0 | 9.3 | Underpredicts (+8.8) |
| USMLEPredictor.com | 10.2 | −0.12 | 16% | 43% | 17.0 | 9.5 | Underpredicts (+9.5) |
| NBMEScore.com / NBME Score Calculator | 12.1 | −0.56 | 9% | 31% | 20.0 | 10.3 | Underpredicts (+11.7) |
1Rules of thumb. Both rules of thumb were tested on a separate held-out validation set of 441 reports with at least two valid dated NBME 9-16 scores.
2Reading the columns. Within 4 and Within 8 show the percentage of predictions no more than that many points from the final score. The 90% within column shows how many points covered 90% of predictions; lower is better. No clear bias means predictions averaged less than 1 point high or low. R2 is 1 for a perfect fit, 0 when predictions match the group average, and negative when they do worse. Tail error is the average of the errors below 240 and at or above 270. Product rows from this held-out test group use 15 low-score and 46 high-score reports; AMBOSS uses 18 and 64; rule rows use 14 and 95.
3AMBOSS. AMBOSS's formula is not public, so its row uses 348 self-reports pairing an AMBOSS prediction with the final score rather than our held-out group.
4Shrinking-bonus rule. Average your two or three most recent NBMEs, find the closest average in the table below, add the listed bonus, and round to the nearest whole number.
| Recent NBME average | Add | Recent NBME average | Add |
|---|---|---|---|
| ≤243 | +12 | 264 | +5 |
| 246 | +11 | 267 | +4 |
| 249 | +10 | 270 | +3 |
| 252 | +9 | 273 | +2 |
| 255 | +8 | 276 | +1 |
| 258 | +7 | ≥279 | +0 |
| 261 | +6 |
For an average between rows, use the closer row. If it is exactly halfway, use the higher average.
Practice-test rankings
The ranking shows the typical error when each assessment is used by itself. It accounts for how many days before the exam each assessment was taken. Lower values indicate smaller errors.
- NBME 16± ~5.3
- NBME 15 · NBME 14 · UWSA 2± ~5.7–6.0
- NBME 13 · NBME 12± ~6.1–6.2
- NBME 10, 11 · UWSA 3± ~6.3
- New Free 120 (2023) · UWSA 1± ~6.5
- NBME 9 · Old Free 120 (2021) · AMBOSS self-assessment± ~6.8–7.0
Typical errors differ by less than two points across the list. The full predictor requires at least two scores, and multiple recent assessments generally improve reliability and narrow the prediction range. The NBME 16 estimate is based on 244 reports, usually taken about 6 days before the exam. Its smaller sample makes the estimate less stable than those for older forms.
Data source and limitations
The dataset comes from public r/Step2 score-release posts that include practice results and a reported final Step 2 score. Reports are extracted, reviewed, and retained only when the numbers are unambiguous. The calculator uses 2,524 reports.
Each record pairs reported practice scores with the reported final score. The accuracy results above use 367 held-out reports that were not used to train the predictor. Assessments described as untimed, split across sessions, answer-checked, interrupted, or taken after the real exam were excluded. Offline percent-to-score estimates were included when the assessment was otherwise taken under standard conditions. Records contain numeric assessment data and broad categories, not usernames or post text.
The sample is not representative of all Step 2 test-takers. Higher scorers are overrepresented, all results are self-reported, and the dataset contains relatively few scores below 230. Displayed percentiles use the official USMLE reference distribution and are separate from this site's collected sample.
Reading the results
The curve displays the estimated distribution of possible scores. Taller portions represent higher estimated probability. The shaded tail corresponds to the odds shown below it. Hover or drag across the curve to check the estimated chance of reaching a selected score.
The range selector controls how broad the predicted range is: 67%, 80%, 90%, or 95%. With at least two practice assessments entered, a 67% range is meant to contain about two of every three comparable outcomes. With fewer assessments, the range is labeled approximate. Choosing a higher percentage produces a wider range.
How each score affects the estimate recalculates the prediction after removing one input at a time. The displayed change shows how much each entry affected the combined estimate.
Your trajectory converts each timed assessment to its associated Step 2 estimate and shows the overall direction over time. Accounting for differences between forms means the trajectory may show a smaller change than the raw practice scores.
Test-takers with similar scores becomes available after you enter at least two practice assessments. It identifies reports with the closest practice results after putting assessments on a common scale, then displays their reported final scores.
Specialty data
Specialty means and standard deviations come from the 2026 NRMP Charting Outcomes reports. You can view U.S. MD seniors, U.S. DO seniors, U.S.-citizen IMGs, non-U.S.-citizen IMGs, or a pooled comparison of the applicants represented in all four reports. These results include consenting applicants who matched to their preferred specialty. NRMP results based on fewer than five applicants are shown as not reported. Ophthalmology and Urology remain external references from SF Match and AUA data and appear only in the U.S. MD view.
NRMP does not publish a combined mean-and-standard-deviation table for those four groups. The pooled view is calculated by this site from each report's mean, standard deviation, and sample size. The source summaries are rounded, and suppressed subgroup cells cannot be included, so pooled values are marked as approximate and affected specialties show how many applicants were represented in the calculation.
Within-specialty percentiles are estimated from that specialty's reported mean and standard deviation, assuming scores follow a roughly bell-shaped pattern. They describe where a score falls among applicants in the selected group who matched their preferred specialty; they do not estimate an individual's match probability.
Privacy
The predictor processes and saves entered scores in your browser rather than sending them to this site's application servers. Saved inputs use browser local storage. A shared link contains the entered values, so anyone with that link can read them. The site collects anonymous page-view counts.
Charts & guides
These pages use the same score-report dataset for related analyses and conversions.
Predictor benchmarks: August 2026. Specialty reference data: July 2026.
Not affiliated with the USMLE®, NBME®, NRMP®, UWorld, or AMBOSS. Predictions are estimates based on self-reported data, and individual outcomes vary.