Skip to main content
Korean Journal of Medical Education logoLink to Korean Journal of Medical Education
. 2026 May 19;38(2):207–211. doi: 10.3946/kjme.2026.188

Bedside skills as context engineering: reframing history taking and physical examination for the AI era

Sanghyun Ahn 1,2,3, Eunbae B Yang 4,✉
PMCID: PMC13236409  PMID: 42157372

1. Introduction

Artificial intelligence (AI) is changing how medicine is practiced and taught. Recent surveys of Korean medical faculty and students have identified AI literacy as an educational priority, advocating spiral curricula that introduce foundational concepts early and build toward clinical applications [1]. Internationally, calls for structured AI curricula in undergraduate medical education are growing [2]. Much of this discussion focuses on what new skills students need—data science literacy, algorithmic reasoning, and prompt engineering.

This commentary offers a different, perhaps counterintuitive, perspective. In the rush to add AI content to medical curricula, we risk overlooking a paradox: the clinical competencies most threatened by curricular compression are precisely those whose value AI amplifies most. History taking and physical examination are not relics of a pre-digital age. They are, we argue, the foundational competencies for effective human–AI collaboration in clinical practice.

The basis for this argument lies in a conceptual shift now under way in the AI industry: the move from prompt engineering to context engineering. Nenadic et al. [3] recently applied this concept to medicine, drawing on industry definitions of context engineering as the deliberate structuring of information so that AI systems produce appropriate outputs [4]. In their framework, context engineering refers to the clinical work of gathering, synthesizing, and organizing clinical information—patient history, physical examination findings, and relevant contextual factors—into AI-interpretable form through bedside judgment, a scope broader than “prompt engineering” alone. They proposed that the physician’s future role lies in engineering four types of context for AI: data, task, tool, and norm. Their framework identifies what physicians should provide to AI but leaves open an educational question that has received less attention: how should we train the next generation to do so? This commentary addresses that question.

2. Context engineering and the irreplaceable value of integrated bedside judgment

The educational significance of this framework becomes apparent when we consider how AI and physicians differ in generating clinical data [3] (Table 1).

Table 1.

Clinical Data Generation: Complementary Roles of AI and Physicians

Domain AI capability Physician’s added value
Symptom analysis Analyzes patient-entered symptom descriptions using natural language processing Elicits nuanced, context-sensitive history through open-ended questioning, follow-up probes, and recognition of what patients omit or minimize
Physical examination AI-assisted tools are emerging (e.g., AI-powered stethoscopes, skin lesion classifiers), but remain narrow in scope Performs integrated, multisystem examination; synthesizes visual, tactile, and auditory findings in real time
Nonverbal and contextual cues Computer vision applications remain limited and are not standard in clinical workflows Detects affect, pain behavior, social dynamics, and environmental context that shape clinical interpretation
Differential diagnosis Generates broad differentials from structured text input; may miss atypical presentations not well-represented in training data Prioritizes differentials by integrating history, exam, and contextual knowledge; applies illness scripts refined through clinical experience
Clinical context structuring Processes prestructured data effectively Transforms unstructured, ambiguous patient narratives into organized clinical frameworks suitable for both human reasoning and AI processing

AI capabilities are rapidly evolving; this comparison reflects current systems and may require updating as technology advances.

AI: Artificial intelligence.

AI-assisted sensing technologies—wearable devices, AI-powered auscultation, computer vision—are advancing rapidly. We do not dispute this. But these tools produce discrete data points, not integrated clinical context. They require physician interpretation, integration, and contextualization. What physicians contribute is integrative bedside judgment—defined here as the real-time synthesis of history, examination findings, nonverbal cues, and patient preferences into a coherent clinical narrative. This judgment is what transforms fragmented observations into the structured context that AI systems need to perform well.

A plausibility argument can be constructed from existing studies suggesting that structured clinical information may improve AI diagnostic performance, though direct evidence linking physician-provided context to AI output quality remains limited. Sonoda et al. [5] found that when an AI system systematically organized clinical information before generating diagnoses, diagnostic accuracy rose from 56.5% to 60.6% (p=0.042)—but this examined AI self-organization rather than physician-provided context. Goh et al. [6] reported that an AI system alone outperformed physicians using AI as a diagnostic supplement, an outcome partially attributed to the difficulty of formulating effective clinical queries; improved context provision was not directly tested. These findings constitute a plausibility argument, not empirical validation. The present proposal should therefore be read as a conceptual educational framework whose educational impact remains to be tested.

Why, then, propose context engineering training now, given such thin evidence? The reasoning is threefold. Even as AI systems acquire independent context-gathering capabilities, the physician’s role in prioritizing, interpreting, and validating context will remain essential. Teaching students to structure clinical information explicitly also strengthens clinical reasoning skills regardless of AI involvement. And delaying curricular innovation until all evidence is conclusive risks leaving a generation unprepared for human–AI collaboration.

Here is the core paradox: as AI grows more capable, the value of what only humans can provide—integrated bedside judgment expressed as structured context—increases, not diminishes. History taking structures subjective context through adaptive questioning. Physical examination generates multisensory data requiring real-time synthesis. Clinical reasoning integrates these into structured task context [3,7]. These are not new skills to be acquired but existing competencies whose educational framing must be updated.

How does context engineering differ from related constructs such as clinical reasoning or problem representation? These are cognitive processes internal to the physician—they describe how clinicians think and organize information. Context engineering, by contrast, is output-oriented: it concerns structuring clinical information so that an AI system can process it effectively. A student may reason well yet produce a summary that AI cannot use, because it lacks the specificity or structure that AI requires. The distinct educational objective is therefore not only “Can you reason well?” but “Can you express your reasoning in a form that optimizes human–AI collaboration?”

3. Reframing the curriculum: from clinical skill to context engineering competency

If physicians’ role in generating clinical context is indeed irreplaceable, how should medical education prepare students for it?

Medical schools have long taught history taking and physical examination as foundational clinical skills [7,8]. Instruction is typically framed around the physician–patient relationship, diagnostic reasoning, and clinical communication. We propose that this framing, while necessary, is no longer sufficient. In the AI era, an additional layer is needed: history taking and physical examination understood explicitly as training in primary data generation and structured context provision for AI-assisted clinical decision-making.

This reframing is not merely semantic—it changes what happens in the classroom. Consider a concrete scenario. In a traditional clinical skills session, students practice eliciting a history of chest pain and document it in SOAP (Subjective, Objective, Assessment, and Plan) format [9]. In a reframed session, the same encounter includes an additional step: students input their clinical summary into a clinical AI system, observe how the AI’s differential diagnosis changes based on the specificity and completeness of their input, and then refine their clinical description iteratively. The student who writes “chest pain, male, 50s” receives a generic differential. The student who writes “52-year-old man with substernal pressure radiating to the left arm, exertional onset, relief with rest, blood pressure 148/92 mm Hg, heart rate 88 bpm, and an S3 gallop on auscultation” receives a more targeted, actionable output. The learning objective shifts: not just “Can you elicit and document a complete history?” but “Can you generate clinical context that produces better AI-assisted recommendations?”

We propose a spiral curriculum structure that introduces this concept progressively (Table 2), applying the principle by Yang [10] that communication competencies develop best through repeated encounters of increasing complexity with structured feedback at each stage.

Table 2.

Proposed Spiral Curriculum: Context Engineering Competency Development

Phase Year Learning objective Context engineering–specific activity Assessment method
Awareness Preclinical (Year 1–2) Recognize that AI output quality depends on input context quality Students input the same clinical case at two detail levels (brief vs. structured) into a clinical AI tool and compare outputs Reflective writing: identify 3 specific data elements that changed AI output
Skill building Clinical clerkship (Year 3–4) Generate structured clinical context that optimizes AI-assisted decision support After a standardized patient encounter, students input their clinical notes into an AI system and evaluate whether critical context was captured Context quality rubric score (formative); self-assessment of context completeness
Integration Senior clerkship (Year 4) Critically compare physician-generated context with AI-generated summaries Students receive a de-identified patient’s AI chatbot history alongside their own clinical note and analyze discrepancies Written analysis: identify ≥3 critical gaps and explain clinical implications
Reflective practice Residency Evaluate and improve one’s own context engineering competency over time Residents review their own AI-assisted clinical documentation monthly, identifying patterns in context gaps Reflective portfolio entry; quarterly review with faculty mentor

Time estimates based on informal pilot observations: Awareness (15 minutes), Skill Building (20 minutes including refinement), Integration (30 minutes including discussion), Reflective Practice (15 minutes). Formal feasibility testing needed to confirm timing across diverse settings.

AI: Artificial intelligence.

4. Implementation considerations

Integrating context engineering into existing curricula raises practical questions.

First, regarding the AI interface: The proposed activities do not require a specific commercial product. A range of clinical AI systems capable of generating text-based differential diagnoses could serve this purpose, including open-source models. The educational value lies in comparing input quality with output quality, not in the specific system used.

Second, regarding assessment: The proposed context quality rubric evaluates four dimensions—completeness, specificity, appropriate terminology, and narrative integration—each on a 4-point scale. We recommend formative deployment, with students tracking improvement across rotations—consistent with the spiral curriculum’s emphasis on progressive development [10].

Third, regarding curricular time: These activities are designed as extensions of existing clinical skills sessions rather than additional standalone sessions. The post-encounter AI task can be integrated as a brief extension of a standard clinical skills session; formal feasibility testing is needed to determine optimal timing.

Finally, regarding faculty readiness and ethical safeguards: clinical faculty require basic AI literacy, addressable through brief workshops supported by local faculty champions. All AI training exercises must use de-identified patient data, and the curriculum must reinforce that clinical accountability rests with the physician.

5. Conclusion

Among the most valuable clinical skills in the AI era are history taking and physical examination—competencies that gain new significance as the means through which physicians generate integrated clinical context that AI cannot yet reliably acquire, prioritize, or validate independently in routine practice. Nenadic et al. [3] articulated the “what”—physicians must engineer context for AI across four axes. This commentary addresses the “how”—a spiral curriculum that explicitly reframes traditional clinical skills training as context engineering competency development.

We are not proposing the addition of new content to already overcrowded curricula. We are proposing a reframing of existing content—one that aligns traditional clinical training with the demands of AI-augmented practice. This curriculum remains conceptual and requires empirical validation. Future studies should examine whether explicit context engineering training improves AI-assisted clinical decisions and whether the proposed rubric demonstrates adequate validity and reliability.

This proposal has limitations. The curriculum has not undergone formal pilot testing; feasibility, optimal timing, and learning outcomes remain to be established, and the rubric requires validation before summative use. Our focus on history taking and physical examination may not generalize to all clinical competencies, though the underlying principles likely apply more broadly. Rapid AI evolution may also affect specific implementation details. Potential unintended consequences—reduced attention to patient-centered communication, over-reliance on AI—warrant careful monitoring.

Medical education should prepare the next generation of physicians not simply to use AI, but to provide the structured clinical context on which effective human–AI collaboration depends.

Footnotes

Acknowledgements

None.

Funding

None.

Conflicts of interest

Sanghyun Ahn is the Chief Medical Officer of Mobile Doctor Inc. and Medical Algorithm Director of MoDoc AI Inc. Except for that, no other potential conflict of interest relevant to this article was reported.

Author contributions

Conceptualization, original draft writing, literature review, and table design: SA. Critical revision for intellectual content, curricular framework validation, and final approval: EBY.

AI disclosure

Generative AI tools (Claude 4.6 Opus and Gemini 3.1 Pro, accessed February 2026) were used for English proofreading and grammar checking of the manuscript. All AI-generated suggestions were reviewed and accepted or rejected by the authors. The AI tools were not used for data analysis, content generation, or intellectual contribution to the manuscript.

References


Articles from Korean Journal of Medical Education are provided here courtesy of Korean Society of Medical Education

RESOURCES