Yes, an in-training exam is best understood as a readiness mirror, not a judgment event: it shows where your knowledge, clinical pattern recognition, testing stamina, and confidence calibration need attention before the boards.
This distinction matters. Residents and fellows often experience the ITE as a public verdict on whether they are “good enough,” especially when percentile rankings are circulated, discussed during evaluations, or used to identify people who may need remediation. However, an ITE is neither a complete measure of clinical competence nor a reliable forecast of one person’s professional future.
Used properly, it is something much more useful: an objective source of feedback.
In this article, we will examine:
- What an in-training exam actually measures
- Why a low score does not automatically predict board failure
- How ITEs expose knowledge and pattern-recognition gaps
- How sleep, workload, interruptions, and stamina affect performance
- Why confidence and competence do not always match
- How to use the ReviewBytes model
- How to build a focused four- to six-week board prep sprint
- How to measure real progress using questions and clinical cases
TL;DR
- A strong ITE score is reassuring, but it does not prove complete clinical competence.
- A weak score should be taken seriously, but it does not reliably predict that an individual will fail the boards.
- Interpret the result through three ReviewBytes lenses: competence gaps, context constraints, and calibration mismatches.
- Do not respond to every low score by simply adding more study hours.
- First identify the weak domains, then classify the kinds of errors being made.
- Build a short, measurable study sprint using spaced questions, cases, feedback, and timed practice.
- Track accuracy, pacing, repeat mistakes, and confidence—not merely the number of chapters read.
What an in-training exam actually measures
An in-training exam, or ITE, is a standardized examination administered during residency or fellowship. Its main purpose is to provide trainees and programs with an objective assessment of medical knowledge across the specialty curriculum.
This sounds straightforward, but it is important to be precise about what the test can and cannot tell us.
An ITE generally samples:
- Factual and conceptual medical knowledge
- Recognition of common and important clinical presentations
- Diagnostic reasoning within written clinical vignettes
- Knowledge of treatment principles and next-step management
- Ability to distinguish between closely related answer choices
- Performance under a defined time limit
It generally does not directly assess:
- Bedside manner
- Communication with patients and families
- Professionalism
- Procedural skill
- Teamwork
- Reliability on the wards
- Leadership
- The ability to manage an evolving clinical emergency in real time
In other words, an ITE measures an important part of being a clinician, but only one part.
Across specialties, ITE performance has a moderate-to-strong relationship with later board examination performance. However, good ITE performance predicts passing more consistently than poor ITE performance predicts failing (PMID: 33680301).
This is why a high score can provide reasonable reassurance, while a low score should be treated as a signal for closer review rather than as a sentence.
A quick glossary for interpreting an ITE report
- Percent correct: The proportion of scored questions answered correctly.
- Percentile: Your position relative to a reference group.
- Domain score: Performance within a content area such as cardiology, nephrology, infectious diseases, or critical care.
- Competence gap: Knowledge is missing, fragmented, poorly organized, or difficult to retrieve.
- Pattern-recognition gap: The learner knows many individual facts but fails to connect the presentation with the correct diagnostic or management pattern.
- Context constraint: Performance was affected by sleep loss, illness, anxiety, interruptions, workload, limited time, or other testing conditions.
- Calibration mismatch: Confidence does not accurately match correctness.
- Readiness: Current preparation for a future assessment—not a permanent label and not a complete description of clinical ability.
The ReviewBytes model is a practical framework for organizing this information. It is not a validated psychometric instrument and should not replace the formal interpretation supplied by a specialty board, training program, or examination provider.
How an ITE exposes different kinds of readiness problems
A final score may look like a single number, but it is produced by several different processes. Two residents can receive the same score for entirely different reasons and therefore require entirely different study plans.
1. The exam samples a broad body of knowledge
No examination can test everything a resident or fellow knows. It uses a selected number of questions to estimate performance across a larger specialty blueprint.
This means small domains may be affected by only a few questions. One should therefore be cautious about building an entire study strategy around a minor difference in a small subsection.
Repeated weakness in an important domain, on the other hand, deserves attention.
2. Clinical vignettes activate illness scripts
Experienced clinicians do not reason through every case by separately recalling hundreds of isolated facts. They organize knowledge into patterns—sometimes called illness scripts—that connect risk factors, pathophysiology, symptoms, examination findings, investigations, and the expected clinical course.
When these scripts are weak or incomplete, a learner may know the facts but fail to recognize the case.
For example, a resident may know the individual features of thrombotic thrombocytopenic purpura but fail to recognize the diagnosis when the findings are embedded in a long vignette. Script theory describes how organized knowledge supports clinical pattern recognition and reasoning (PMID: 27004079).
3. Closely related answers expose discrimination problems
Many board-style questions are not asking whether you know a fact. They are asking whether you can distinguish between two answers that are both partly correct.
Common examples include:
- The best initial test versus the most definitive test
- The next step versus the eventual treatment
- A common diagnosis versus a dangerous alternative
- When to observe versus when to intervene
- A guideline threshold versus a general clinical principle
This is why passive rereading may not correct the problem. The learner may need comparison tables, contrasting cases, and repeated practice explaining why one plausible option is better than another.
4. Time pressure exposes attention and stamina
Long examinations require sustained attention. A learner may perform well during the first portion of the exam but begin to misread qualifiers, rush, or change correct answers later.
Sleep loss can also affect cognitive and clinical performance, although the size of the effect varies among individuals and tasks (PMID: 16335329).
The ITE should not therefore be interpreted without asking:
- Was the learner post-call?
- Was there adequate sleep during the preceding week?
- Were there interruptions?
- Was the examination completed?
- Did accuracy decline during later practice blocks?
- Was anxiety interfering with reading or decision-making?
Context does not make a weak result irrelevant. However, it can change what the result means and what should be done next.
5. Confidence determines when we stay, switch, or guess
Confidence is useful, but it is not always accurate.
Some learners correctly identify areas where they are weak. Others feel confident because the material looks familiar, even though they cannot retrieve or apply it independently. Still others repeatedly change correct answers because they underestimate their knowledge.
Physicians have a limited ability to assess their own competence accurately when self-assessment is compared with external measures (PMID: 16954489).
This is one reason an objective examination can be helpful. It shows us something that confidence alone may not reveal.
What the research shows about ITEs and board performance
The best available evidence supports a formative use
A systematic review of 32 studies across 21 specialties found that ITE performance generally had a moderate-to-strong relationship with later board examination performance.
However, the review also found an important asymmetry:
- Strong ITE performance was useful for predicting board passage.
- Poor ITE performance was less reliable for predicting board failure.
The authors therefore cautioned against using one low ITE score for probation, non-promotion, non-retention, or other high-stakes decisions (PMID: 33680301).
This is a sensible conclusion. When specialty board pass rates are high, even some trainees with very low ITE results will subsequently pass. Any prediction system will therefore produce false positives—people labeled as likely to fail who ultimately succeed.
The more defensible use of a low score is not to say, “This resident will fail.”
It is to say, “This resident may have important learning needs, and we should examine them carefully.”
Scores generally improve with training
Internal medicine ITE performance has been shown to improve predictably with postgraduate training. Results may also be influenced by the timing of the examination, the time allowed to complete it, and the amount of internal medicine training completed before the test (PMID: 12230352).
This is particularly relevant for interns.
A PGY-1 ITE is partly a reflection of prior medical school preparation and partly a baseline for future growth. It should not be interpreted in exactly the same way as a final-year result obtained shortly before the certification exam.
Repeated results are more informative than one result
In an internal medicine cohort, annual ITE scores correlated with later ABIM certification examination scores. Scoring in the lowest ITE quartile increased the risk of failing the boards, but it did not determine the outcome for an individual resident (PMID: 25271892).
A similar pattern has been demonstrated in fellowship training.
Among 1,684 second-year nephrology fellows, ITE performance was strongly associated with ABIM nephrology certification performance. Yet even in this board-aligned subspecialty setting, low scores had limited ability to identify exactly which fellows would fail (PMID: 29490975).
No matter how compelling the correlation may appear at a group level, it cannot be converted into certainty for one trainee.
Evidence supports retrieval practice more than passive review
Once a gap is identified, the next question is how to correct it.
A systematic review of 56 studies in health-professions education found that distributed practice and retrieval practice improved academic outcomes in most included experiments, although the interventions and assessment methods were heterogeneous (PMID: 37615780).
Test-enhanced learning also supports the use of:
- Repeated retrieval
- Spacing over time
- Effortful recall
- Corrective feedback
- Reassessment after a delay
These strategies generally support retention better than repeated passive reading (PMID: 18823514).
This does not mean textbooks, videos, and lectures are useless. It means they should usually be used to repair a problem uncovered through questions or cases, rather than becoming the entire study plan.
The ReviewBytes model: competence, context, and calibration
The simplest way to interpret an ITE is to ask three questions.
1. Is there a competence gap?
A competence gap may involve:
- Missing factual knowledge
- Weak illness scripts
- Poor differentiation between similar conditions
- Confusion about management thresholds
- Difficulty applying knowledge in a vignette
- Knowledge that was learned but cannot be retrieved
Evidence of a competence gap becomes stronger when:
- The same domain is weak on multiple assessments
- Similar questions are repeatedly missed
- The learner cannot explain the answer without options
- Errors remain after explanations have been reviewed
- Weakness appears under both timed and untimed conditions
2. Was there a context constraint?
Context constraints include:
- Post-call fatigue
- Chronic sleep restriction
- Illness
- Major personal stress
- Frequent interruptions
- An unfamiliar testing environment
- Inadequate testing time
- Severe anxiety
- Language load
- Unaddressed accommodation needs
One can easily make two mistakes here.
The first is to ignore context and treat the score as pure competence. The second is to blame context for everything and avoid addressing a genuine knowledge problem.
A better approach is to reproduce the suspected problem under improved conditions. If accuracy improves markedly when the learner is rested and uninterrupted, context was probably contributing. If the same errors persist, a competence gap is also present.
3. Is there a calibration mismatch?
Calibration compares confidence with actual performance.
Useful categories include:
- Correct and confident: The knowledge is probably stable.
- Correct but unsure: Knowledge may be present but fragile.
- Incorrect and unsure: The learner recognizes uncertainty.
- Incorrect and confident: A mistaken rule or illness script may be operating.
High-confidence wrong answers deserve particular attention. They are more concerning than ordinary guessing because the learner may repeatedly apply the same incorrect reasoning.
Common myths about in-training exams
Myth: A low ITE means I will fail boards
Reality: A low result deserves attention, particularly when repeated, but it does not reliably determine an individual board outcome (PMID: 33680301).
Myth: The ITE tells me whether I am a good doctor
Reality: The ITE samples written medical knowledge and clinical reasoning. It does not measure every component of safe and effective clinical practice.
Myth: The answer is simply to study more hours
Reality: More time may help, but only if the method matches the problem. Recall gaps, recognition problems, pacing errors, and overconfidence require different interventions.
Myth: I felt confident, so I must have been prepared
Reality: Confidence and competence do not always match (PMID: 16954489; PMID: 39463809).
Myth: I understood the explanation, so I now know the material
Reality: Recognizing an explanation is easier than retrieving and applying the information independently several days later.
Myth: A high score means board preparation is complete
Reality: No ITE tests every topic or every competency, and knowledge can decay without continued retrieval.
How to review an ITE result without overreacting
The first response to an ITE should not be panic, denial, or a hurried purchase of several new resources.
It should be a structured review.
Step 1: Identify the weak domains
Rank the domains using:
- Absolute performance
- Performance relative to your training level
- Change from previous years
- Blueprint weight
- Clinical importance
- Performance in the same domain on question banks or other assessments
Select two primary domains and one secondary domain for the first sprint.
“Review all of medicine” may sound ambitious, but it is not an actionable plan.
Step 2: Classify the errors
Create a simple error log and classify each miss.
| Error type | What happened | Typical response |
| Recall | The fact, association, criterion, or adverse effect was unknown | Focused review followed by spaced retrieval |
| Recognition | The clinical pattern was not identified | Build illness scripts and review representative cases |
| Discrimination | Two plausible answers could not be separated | Compare contrasting diagnoses or management options |
| Management | The diagnosis was known, but the next step or threshold was wrong | Practice decision pathways and timing questions |
| Attention | A qualifier, unit, trend, or contraindication was missed | Slow down, annotate the question, and use a final check |
| Pacing or stamina | Questions were rushed or left unfinished | Use timed blocks and rehearse examination length |
| Calibration | Confidence did not match correctness | Record confidence and prioritize confident errors |
This classification is often more useful than a long list of topics.
Step 3: Build a four- to six-week sprint
Weeks 1 and 2: Repair the main gaps
- Complete focused question sets in the two weakest domains.
- Review every answer, including correct answers that were guesses.
- Write brief illness scripts.
- Record why the correct answer is right.
- Record why the most tempting alternative is wrong.
- Retest missed concepts after several days.
Weeks 3 and 4: Mix and apply
- Interleave weak topics with stronger domains.
- Add cases that require diagnosis, next-step management, or risk stratification.
- Practice without answer choices when possible.
- Revisit high-confidence errors.
- Begin timed mixed blocks.
Weeks 5 and 6: Rehearse the examination
- Complete longer timed blocks.
- Practice at the same time of day as the expected examination when practical.
- Track whether accuracy falls late in a session.
- Use a stop rule for questions that are consuming too much time.
- Review only the errors that reveal a recurring rule, pattern, or decision problem.
The main engine of the sprint should be questions and clinical cases. Videos, notes, podcasts, and textbooks should support that work.
How to measure whether board prep is working
Study time is easy to count but difficult to interpret.
A resident may study for 40 hours without improving the ability to recognize a case, choose the next step, or finish a timed block.
More useful measures include:
- Rolling accuracy over the most recent 50–100 questions
- Accuracy in the two target domains
- Frequency of repeated errors
- Number of high-confidence wrong answers
- Completion rate
- Average time per question
- Accuracy during the final portion of a block
- Performance on short cases without answer options
- Ability to explain why competing answers are wrong
Progress is not merely, “I finished the cardiology chapter.”
Progress is, “I can retrieve the relevant rule, recognize the presentation, choose the correct next step, and do so within the available time.”
How to interpret common ITE patterns
How to interpret this table: identify the most likely ReviewBytes category, then test that explanation using targeted practice.
| ITE pattern | Likely category | What it may mean | Best response | Evidence notes |
| Low domain score with low confidence | Competence gap | The learner recognizes uncertainty and lacks accessible knowledge | Focused retrieval and repeat cases | ITEs can identify learning needs (PMID: 33680301) |
| Low score with confident errors | Competence plus calibration gap | An incorrect rule or script is being applied confidently | Contrastive cases and confidence tracking | Self-assessment may be inaccurate (PMID: 16954489; PMID: 39463809) |
| Unfinished exam or late-block decline | Context or stamina problem | Pacing, fatigue, or over-investment in difficult questions | Timed blocks, stop rules, and rested rehearsal | Sleep loss can impair performance (PMID: 16335329) |
| Stable percent correct but lower percentile | Norm-reference effect | Peers may have improved more | Review the absolute score, trend, and blueprint | Training stage affects results (PMID: 12230352) |
| Repeated low results across years and question banks | Persistent competence gap | Broad or unresolved knowledge-organization problem | Formal learning plan, coaching, and serial measurement | Repeated ITE results relate to board outcomes (PMID: 25271892) |
How training stage changes the response
How to interpret this table: the same score may require a different response depending on the learner’s stage, available preparation time, and testing conditions.
| Clinical or training scenario | What changes | Counseling point | What to monitor | Evidence notes |
| PGY-1 baseline examination | More runway and less specialty exposure | Use it as a map, not an identity | Domain trajectory over time | Scores generally rise with training (PMID: 12230352) |
| Final-year resident preparing for ABIM or another board | Less time before certification | Prioritize high-weight gaps and timed work | Weekly accuracy and completion | ITE and certification scores correlate (PMID: 25271892) |
| Subspecialty fellow | Narrower, board-aligned content | Use feedback for specialty-specific cases | Certification-style questions | Nephrology data support this use (PMID: 29490975) |
| Post-call or interrupted examinee | Context may depress performance | Contextualize the result without dismissing it | Rested timed blocks | Sleep affects performance (PMID: 16335329) |
| Strong clinical evaluations but weak ITE | Different competencies are being sampled | Preserve clinical strengths while repairing written reasoning | Recall, case analysis, and test mechanics | An ITE does not represent all competence |
| PA or NP specialty assessment | Physician GME predictions may not apply directly | Use the framework, not physician board-pass estimates | Domains, cases, pacing, and calibration | Evidence is strongest in physician graduate medical education |
When to involve a program director, mentor, or learning specialist
Not every low result requires formal remediation. However, early support is sensible when there is:
- Repeated broad underperformance
- A substantial decline from prior performance
- Persistent inability to finish timed blocks
- Severe anxiety or insomnia
- Depressive symptoms affecting concentration or function
- Concern for a learning difference
- A possible need for testing accommodations
- A mismatch between clinical performance and written assessment that remains unexplained
- Little improvement despite a structured study plan
The purpose of seeking help is not to turn the score into a disciplinary label. It is to widen the available information and create a more effective plan.
Hospitals and medical colleges should also consider their own responsibilities. Residents cannot be told to improve their performance while being given no protected learning time, no usable feedback, and no support for fatigue, mental health, or learning needs. A culture of retaining and upskilling clinicians requires more than identifying a low percentile.
Important exceptions and “it depends” situations
Many ITE reports do not show item order, response time, or confidence. Therefore, the report itself may not prove that the learner has a stamina problem or a calibration mismatch.
These possibilities must be tested through practice.
Other points deserve emphasis:
- A percentile is not a diagnosis.
- One low domain may represent only a few questions.
- Context may explain part of a result without explaining all of it.
- Strong ward performance and weak written performance can coexist.
- A high ITE does not guarantee broad clinical competence.
- A low ITE is actionable, but it is not destiny.
- A four- to six-week sprint suits a focused problem.
- Repeated, broad deficiencies may require several months and formal support.
- The ReviewBytes model is an educational framework, not a universal cut score.
Key takeaways to remember on a busy shift
- Treat the ITE as a readiness mirror, not a character judgment.
- Ask what is missing, what constrained performance, and where confidence was inaccurate.
- Review absolute performance, domain results, trend, and context—not percentile alone.
- Strong performance is reassuring; weak performance guides action but does not determine fate.
- Classify the errors before selecting study resources.
- Prioritize repeated mistakes, management errors, and high-confidence wrong answers.
- Use spaced questions and clinical cases as the engine of board prep.
- Rehearse timing and stamina rather than assuming they will improve automatically.
- Measure changing performance, not hours logged.
- Seek support early when weakness is broad, persistent, or complicated by health or accommodation concerns.
References
- McCrary HC, Colbert-Getz JM, Poss WB, Smith BK. A systematic review of the relationship between in-training examination scores and specialty board examination scores. J Grad Med Educ. 2021;13(1):43–57. PMID: 33680301. doi:10.4300/JGME-D-20-00111.1.
- Garibaldi RA, Subhiyah R, Moore ME, Waxman H. The In-Training Examination in Internal Medicine: an analysis of resident performance over time. Ann Intern Med. 2002;137(6):505–510. PMID: 12230352. doi:10.7326/0003-4819-137-6-200209170-00011.
- Kay C, Jackson JL, Frank M. The relationship between internal medicine residency graduate performance on the ABIM certifying examination, yearly in-service training examinations, and the USMLE Step 1 examination. Acad Med. 2015;90(1):100–104. PMID: 25271892. doi:10.1097/ACM.0000000000000500.
- Jurich D, Duhigg LM, Plumb TJ, et al. Performance on the nephrology in-training examination and ABIM nephrology certification examination outcomes. Clin J Am Soc Nephrol. 2018;13(5):710–717. PMID: 29490975. doi:10.2215/CJN.05580517.
- Philibert I. Sleep loss and performance in residents and nonphysicians: a meta-analytic examination. Sleep. 2005;28(11):1392–1402. PMID: 16335329. doi:10.1093/sleep/28.11.1392.
- Davis DA, Mazmanian PE, Fordis M, Van Harrison R, Thorpe KE, Perrier L. Accuracy of physician self-assessment compared with observed measures of competence: a systematic review. JAMA. 2006;296(9):1094–1102. PMID: 16954489. doi:10.1001/jama.296.9.1094.
- Gaeta TJ, Reisdorff E, Barton M, et al. The Dunning–Kruger effect in resident predicted and actual performance on the American Board of Emergency Medicine in-training examination. J Am Coll Emerg Physicians Open. 2024;5(5):e13305. PMID: 39463809. doi:10.1002/emp2.13305.
- Lubarsky S, Dory V, Audétat MC, Custers E, Charlin B. Using script theory to cultivate illness script formation and clinical reasoning in health professions education. Can Med Educ J. 2015;6(2):e61–e70. PMID: 27004079.
- Trumble E, Lodge J, Mandrusiak A, Forbes R. Systematic review of distributed practice and retrieval practice in health professions education. Adv Health Sci Educ Theory Pract. 2024;29(2):689–714. PMID: 37615780. doi:10.1007/s10459-023-10274-3.
- Larsen DP, Butler AC, Roediger HL III. Test-enhanced learning in medical education. Med Educ. 2008;42(10):959–966. PMID: 18823514. doi:10.1111/j.1365-2923.2008.03124.x.
Frequently asked questions about in-training exams
What is an in-training exam?
An in-training exam is a standardized assessment used during residency or fellowship to sample specialty knowledge and applied clinical reasoning. Its main value is formative: it can identify areas that deserve further review.
Does a low ITE score mean I will fail boards?
No. A low score may increase concern, especially when the pattern is repeated, but it does not reliably determine whether one individual will pass or fail a later board examination.
How should I review my ITE results?
Review your domain scores, absolute performance, longitudinal trend, testing context, and level of confidence. Then classify your mistakes as recall, recognition, discrimination, management, attention, pacing, or calibration errors.
How long should a post-ITE study sprint last?
A focused sprint often lasts four to six weeks. Broad, repeated, or longstanding gaps may require a longer plan, coaching, or formal support from the training program.
When should I ask my program for help?
Ask for help when weakness is repeated, broad, or declining; when timed blocks remain unfinished; or when sleep, anxiety, health, language, or accommodation concerns are affecting performance.
Can the ReviewBytes model be used by physician assistants and nurse practitioners?
Yes. The framework can help physician assistants, nurse practitioners, medical students, residents, and fellows analyze knowledge gaps, context, and calibration. However, board-performance data from physician residency studies should not be converted directly into pass predictions for other professions.
What is ReviewBytes?
ReviewBytes is a modern medical learning platform built around clear, focused, evidence-based education. Our approach combines microlearning, proven learning science, and AI-powered technology to help learners review more effectively and retain more over time.
Why is it called ReviewBytes?
The name ReviewBytes reflects two core parts of our mission. Review represents mastery, reinforcement, and evidence-based learning strategies like spaced repetition, retrieval practice, and the testing effect. Bytes reflects both bite-sized learning and our AI-first, technology-forward approach to medical education.
Is ReviewBytes the same as Review Bytes?
Yes. ReviewBytes and Review Bytes refer to the same brand. Some people search for it as one word, while others type it as two words.
Is ReviewBytes pronounced like “review bites”?
Sometimes, yes—and that fits our mission well. The phrase “review bites” naturally connects to bite-sized learning: smaller, focused learning moments designed to make medical education more manageable and more effective.
What does the name ReviewBytes mean?
The name ReviewBytes reflects our belief that medical learning should be clear, focused, and built for the modern learner. Review speaks to scientifically grounded learning methods that improve retention and recall, while Bytes reflects both bite-sized learning and a technology-forward educational experience.
Why did you choose the name ReviewBytes?
We chose ReviewBytes because it captures the way we think learning should work: evidence-based, efficient, and thoughtfully designed. The name brings together proven review methods with microlearning and AI-powered innovation.
Do people also search for Review Bytes?
Yes. Many learners search for Review Bytes as a variation of ReviewBytes, and both refer to the same brand and mission.
Does ReviewBytes relate to bite-sized learning?
Absolutely. The “Bytes” in ReviewBytes is a nod to bite-sized learning—breaking complex medical concepts into smaller, easier-to-review pieces—while also reflecting our tech-forward approach.
⚠️ Educational disclaimer: This article is educational only, not personalized medical, psychological, employment, or testing-accommodation advice. Seek appropriate clinician or institutional guidance for individual concerns.



