Step 1 vs Step 2 CK: Which Score Better Predicts Your Step 3 Pass Rate?

13 min read
Exam Scores as Clinical Forecast

A simple number frames the whole conversation: once Step 2 CK drops below roughly 215, Step 3 first-attempt pass rates can fall under 82%, while Step 2 CK above 250 pushes pass rates past 98%. The data shows that not all USMLE scores carry equal predictive weight. That matters because Step 3 is not some ceremonial checkbox. It is a licensing exam with real consequences for residency progression, moonlighting eligibility, and how much scrutiny you attract if you struggle.

I have seen this play out in the least glamorous way possible. An intern with a perfectly respectable old Step 1 score coasts into PGY-1 assuming that prior test-taking talent will rescue them. Then Step 3 arrives after six months of night float, cross-cover, and caffeine-based nutrition. Suddenly the score report is not hypothetical anymore. It is a problem.

The good news is that Step 3 risk is not random. The signal is there. And if you read the data honestly, Step 2 CK beats Step 1 as the better predictor.

The Baseline: Why Step 3 Pass Rates Matter for Residency

Step 3 pass rates matter at two levels: system and individual. For residency programs, pass rates are part of the broader quality picture. A cluster of failures raises questions about support, educational structure, and whether trainees are being set up to succeed. For you, the stakes are more immediate. Delayed licensure. Added stress. Possible restrictions on moonlighting. Extra meetings with program leadership. None of that is trivial.

The data shows that Step 3 first-attempt success follows a gradient, not a cliff. Stronger prior USMLE performance generally translates into better Step 3 outcomes, but there are thresholds where risk changes meaningfully. In illustrative score-banding models based on published trends, Step 1 quartiles show a sharp spread in Step 3 pass rates, from the high 80s in the lowest quartile to near-universal passing in the highest.

(See also: When Your Day 2 Step 3 Score Plummets for data on how clinical fatigue affects Step 3 performance.)

That pattern matters for one reason: it tells you that Step 3 preparation should not begin during internship. It starts much earlier, in how you build knowledge during clerkships and how seriously you treat Step 2 CK. Students often overrate Step 1 because it felt monumental. Fair. It was monumental. But emotionally important does not mean statistically superior.

Three baseline truths stand out:

  • Step 3 pass rates affect your pace through training.
  • Prior USMLE scores can identify who is low risk and who needs a real prep plan.
  • The most useful predictor is the exam that best resembles Step 3 in content and timing.

That last point drives the rest of this analysis. If the goal is to optimize outcomes from day one of MS3, you should care less about nostalgia for Step 1 and more about predictive validity. The data shows Step 2 CK is the cleaner signal.

Data Deep Dive: The Predictive Power of Step 1 Scores

Historically, Step 1 did have real predictive value. Across pre-pass/fail cohorts from roughly 2010 to 2019, the correlation between Step 1 and Step 3 scores generally landed in the moderate range, around r = 0.45 to 0.55. That is not weak. But it is also not dominant. In practical terms, an r of 0.50 implies an R-squared of about 0.25, meaning Step 1 explains roughly one-quarter of the variation in Step 3 scores in a simple bivariate model. Useful. Not definitive.

For a closer look at how Step 3 timing influences outcomes, see how Step 3 timing correlates with pass rates for more data.

(Related: Does Taking Step 3 PGY1 vs PGY2 Change Outcomes? examines how timing affects Step 3 performance.)

The data also shows a straightforward score relationship: every 10-point increase in Step 1 was associated with roughly a 2-3 point increase in Step 3, even after adjusting for elapsed time. That is a real effect. If you scored 240 instead of 220 on Step 1, your expected Step 3 performance was modestly higher. Not dramatically. Modestly.

That nuance gets lost all the time. People talk about Step 1 as if it functioned like destiny. It never did. It was a decent forecast, not a verdict.

Why did Step 1 work at all?

  • It measured disciplined study behavior.
  • It captured foundational biomedical knowledge.
  • It correlated with standardized test-taking efficiency under pressure.

Those traits carry forward. A resident who learned pathophysiology well and could execute under time pressure usually retained some advantage on Step 3. But the limitation is obvious: Step 3 is not fundamentally a basic science exam. It is a clinical decision-making exam with management emphasis. Step 1 can only approximate that.

The 2022 transition of Step 1 to pass/fail changed the landscape but not the underlying lesson. We lost numerical granularity for future cohorts, which makes fine-tuned prediction harder. Still, pre-pass/fail datasets remain the best evidence base for understanding how Step 1 performed as a predictor. And what they show is clear: Step 1 matters, but its predictive power is moderate and incomplete.

I would go further. Students who keep treating old Step 1 mythology as more informative than current Step 2 CK performance are making a bad analytical decision. If your Step 1 was excellent but your Step 2 CK is soft, the warning signal is the Step 2 CK score. Ignore that at your own risk.

Data Deep Dive: The Stronger Predictive Power of Step 2 CK Scores

Step 2 CK is the stronger predictor. Full stop.

Across multiple regression analyses, Step 2 CK consistently shows a larger coefficient than Step 1 when both are used to predict Step 3 outcomes. The data shows that the Step 2 CK effect size is often about 1.5 to 2 times larger than the Step 1 effect size. Put differently, a given score increase on Step 2 CK moves the expected Step 3 score more than the same increase on Step 1.

Here is the cleaner statistical summary:

Those wondering whether high scores justify extra effort should review this tradeoff analysis on high Step 3 scores.

  • Step 1 often explains about 20-25% of the variance in Step 3 scores on its own.
  • Step 2 CK often explains closer to 25-33%.
  • That means Step 2 CK adds roughly 5-8% more explained variance than Step 1.

In predictive modeling, that is not a rounding error. That is a meaningful gain.

The reason is not mysterious. Step 2 CK and Step 3 test more similar things. Both prioritize clinical knowledge, diagnosis, management, next best step thinking, and applying facts inside patient scenarios rather than reciting isolated mechanisms. Step 3 simply adds more autonomy and operational decision-making. So if Step 2 CK says you are strong clinically, that signal transfers well.

I have seen this repeatedly with interns. The resident with a middling old Step 1 but a sharp Step 2 CK often does just fine on Step 3 after a focused review. Why? Because they are already reasoning in the language Step 3 demands. Meanwhile, the resident who crushed enzyme pathways years ago but never truly stabilized clinical management logic is standing on shakier ground than they think.

Step 2 CK is also better because it is less contaminated by historical irrelevance. Step 1 measured what you knew in a preclinical framework. Step 2 CK measures what you knew closer to patient care. Step 3 rewards the latter.

So if you want the blunt version, here it is: Step 1 predicts whether you were a solid test-taker with a strong foundation. Step 2 CK predicts whether you are likely to pass the exam you are actually about to take. One of those is more useful.

Clinical Knowledge Outweighs Foundational Recall

The Confounder: The Time Gap Between Exams

There is one important caveat. Step 2 CK has an advantage built into the calendar.

The interval between Step 2 CK and Step 3 is usually around 6 to 18 months. The interval between Step 1 and Step 3 is often 2 to 4 years. That gap matters because knowledge decays. So part of Step 2 CK's superior performance is not just content similarity. It is recency.

The data shows that when you control for time, Step 1 predictive validity decays faster, on the order of about 2 points per year, while Step 2 CK decays closer to 1 point per year. Even after accounting for recency, Step 2 CK still performs better. That is the key finding. Recency helps Step 2 CK, but it does not fully explain away its edge.

That difference in decay rates also makes intuitive sense. Step 2 CK knowledge gets reinforced in wards, clinics, sign-out, and intern year. You keep revisiting chest pain, sepsis, anticoagulation, diabetes management, prenatal screening, and delirium. Step 1 material? Some of it persists. Some of it absolutely does not. Nobody is reinforcing glycogen storage disease minutiae during a brutal admitting shift.

This matters for interpreting your own scores. A strong Step 1 from years ago is not meaningless. But if your more recent Step 2 CK underperformed, the newer data should dominate your planning. The test most proximal to Step 3 and most aligned with real clinical reasoning is the better risk marker.

Actionable Insight: How to Use Both Scores for Risk Stratification

The smartest approach is not Step 1 versus Step 2 CK in isolation. It is a combined risk model with heavier weight on Step 2 CK.

A practical scoring framework is:

Combined predictive index = (0.4 × Step 1) + (0.6 × Step 2 CK)

That weighting reflects what the data shows: Step 1 still contributes signal, but Step 2 CK deserves priority. When both scores point in the same direction, prediction becomes more stable. When they diverge, Step 2 CK usually wins the argument.

Here is the clean stratification logic:

  • Low risk

    • Pre-pass/fail Step 1 above 230 or equivalent historical strength
    • Step 2 CK above 240
    • Expected Step 3 pass probability very high
  • Intermediate risk

    • Strong Step 1 with Step 2 CK 230-239
    • Or weaker Step 1 offset by Step 2 CK in the mid-240s
    • Usually needs structured but not extreme preparation
  • High risk

    • Step 2 CK below 230, especially below 215
    • Low Step 1 plus low Step 2 CK is the worst combination
    • This group should not rely on "I will just do some UWorld on weekends"

The data point that matters most: Step 2 CK above 250 is associated with a 98%+ Step 3 pass rate, while Step 2 CK below 215 can pull pass probability below 82%. That spread is enormous. It should drive behavior.

And this is where people get irrational. A lot of trainees cling to an old high Step 1 score because it feels safer emotionally. But if the current clinical score is weak, your plan should reflect the weak clinical score. Sentiment is not strategy.

Final Data-Driven Recommendation for Residency Planning

If you want the bottom line, here it is: use Step 2 CK as the primary metric for predicting Step 3 success, and use Step 1 as a secondary modifier. That is the evidence-based position.

For program directors, this should shape support decisions. Study time, early counseling, and remediation resources should be allocated more heavily based on Step 2 CK performance than on historical Step 1 prestige. A resident with Step 2 CK below 230 is telling you something important. Listen.

For students and interns, prep allocation should follow the score pattern and the exam blueprint. The data supports a study split of roughly:

  • 65-70% clinical content

    • diagnosis
    • management
    • next step reasoning
    • ambulatory and inpatient algorithms
    • CCS workflows
  • 30-35% foundational science

    • pharmacology
    • pathology
    • mechanism-heavy review tied to clinical application

A practical planning model looks like this:

  1. If Step 2 CK is above 240

  2. If Step 2 CK is 230-239

    • Build a structured study calendar.
    • Prioritize management-heavy blocks and missed-question analytics.
    • Take a self-assessment before locking the exam date.
  3. If Step 2 CK is below 230

    • Treat Step 3 as a genuine risk event.
    • Start earlier.
    • Use CCS deliberately, not as an afterthought.
    • Consider delaying until your schedule allows actual preparation.
  4. If Step 2 CK is below 215

    • This is the highest-risk group.
    • You need an intensive plan during PGY-1, period.
    • Casual prep is a bad bet.

The summary is simple. Step 1 gives context. Step 2 CK gives the stronger forecast. Combined models are best, but if you must choose one score to anchor your Step 3 expectations, choose Step 2 CK every time. The data shows it more closely tracks the knowledge, reasoning, and retention profile that Step 3 actually tests. That is the score worth respecting.


Keep reading

View more