Case - Assessment and feedback (Bias)
What looks like an ordinary assessment comment may carry more than one meaning. This case invites educators to slow down, examine how bias can enter clinical feedback and ratings, and consider how people, processes, and platforms can be redesigned to support fairer, more useful assessment.
Bias is difficult to see in the abstract. It becomes visible when we slow down and look closely at ordinary situations — moments that feel routine, reasonable, and well-intentioned at the time.
Cases allow us to ask different questions:
- What assumptions are shaping this judgment?
- Where is discretion sitting?
- What alternatives were available, but not taken?
- Which features of the system made this outcome more likely?
The goal here is not to assign blame, but to build shared understanding.
Case 1 - Bias in assessment and feedback in clinical learning environments
Amira is a second year resident on an inpatient medicine rotation. Her attending, Dr. K, writes that Amira is “not quite ready to lead on rounds,” notes “language hesitations,” and rates her “below expected” for oral presentations. Amira’s progress note quality is rated highly by the same team. Two other residents, both men, receive “meets” or “exceeds” ratings with little narrative detail. In team rooms, a senior fellow often interrupts Amira. Raters complete assessments at the end of the month in a batch. The program uses global rating scales with few anchors.
Narrative comments are not calibrated. Amira is an international graduate and wears a hijab.
What we know
- Ratings are vulnerable to leniency, severity, and halo effects. Under-represented learners and international graduates receive more negative feedback, less specific feedback, and are described with personality labels more than with behaviorally anchored language.
- Interruptions and stereotype-consistent expectations reduce speaking time and visibility, which then depresses ratings of leadership and communication.
- Structural features that increase bias include unanchored scales, end-of-block recall, single-rater dominance, and lack of shared mental models for competence.
- High quality feedback is timely, behavior-specific, aligned with entrustable professional activities, and co-constructed. Rater training plus system redesign outperforms training alone.
How we got here
- Historical reliance on trait adjectives and global ratings rather than direct observation with anchors.
- Assessment systems built for administrative convenience rather than fairness or learning.
- Hidden curriculum that rewards confidence displays over careful reasoning, and, equates fluency with competence.
Bias Map
| Level | Implicit | Explicit | Structural |
|---|---|---|---|
| Self | Dr. K’s fluency heuristic treats minor speech pauses as lower competence | Senior fellow’s conscious interruptions that favor familiar trainees. | Rater fatigue, batch completion, unanchored scales, single-rater weight, limited direct observation. |
| Interpersona | Less eye contact with Amira, fewer coaching bids toward her. | Differential opportunities to present or lead. | Team norms that allow interruptions and sidebars. |
| Structural | None of the assessment items require evidence of direct observation | Narrative fields prompt personality traits. | No calibration meetings, no audit of score distributions by gender, IMG status, or race. |
Framework to approach the case
Use a three-lane plan: People, Process, Platform.
People (self and team habits)
- Teach a micro-skills triad for feedback: Ask, Notice, Name;
- Ask for the learner’s self-assessment linked to an EPA;
- Notice concrete behaviors you directly observed, with time and context; and
- Name the next step and arrange a follow-up check.
- Introduce structured language stems: “I saw… which suggests… next time try… by…”; and
- Run brief rater calibration huddles that use shared videos or de-identified notes to align anchors.
Process (how assessments get created)
- Replace global ratings with behaviorally anchored tools tied to EPAs or milestones;
- Require at least two direct observations per week with 24-hour feedback delivery;
- Use the R2C2 sequence for summative conversations: Relationship, Reaction, Content, Coaching; and
- Protect presentation time. Set a “no interruption” ground rule and rotate leadership opportunities.
Platform (system and data)
- Build structured prompts that force behaviors, not traits. For example, “cites patient data before plan” rather than “confident.”;
- Time-stamp entries, prevent end-of-block batch submissions, and send nudges after encounters;
- Run monthly equity audits. Examine score distributions and narrative descriptors by rotation, rater, and learner subgroup; and
- Create an appeal and re-review pathway when patterns suggest bias.
