Recipe 2.3: Clinical Documentation Improvement (CDI) Suggestions
Effort: 2 of 5
The Problem
A hospitalist admits a patient with pneumonia. They write in the progress note: "Patient has pneumonia, started on antibiotics." Clinically, this is fine. The patient gets treated. But from a coding and reimbursement perspective, this note is a disaster.
Was it community-acquired or hospital-acquired pneumonia? Bacterial, viral, or aspiration? Which organism, if known? Is it the principal diagnosis or a complication of something else? Each of these distinctions maps to a different ICD-10 code, and each code maps to a different DRG, and each DRG maps to a different reimbursement amount. The difference between "pneumonia, unspecified" (J18.9) and "pneumonia due to Streptococcus pneumoniae" (J13) can mean thousands of dollars in reimbursement difference for the same clinical care.
This is not about upcoding. This is about accuracy. The documentation should reflect what the physician actually knows and did. When a physician writes "pneumonia" but their lab results show Streptococcus and their antibiotic choice confirms they're treating a bacterial infection, the documentation is incomplete, not wrong. The clinical picture is clear in the physician's head. It just didn't make it onto the page.
Clinical documentation improvement (CDI) is the discipline of catching these gaps. Traditionally, CDI specialists (usually nurses or coders with clinical backgrounds) manually review charts, identify documentation that lacks specificity, and send queries to physicians asking them to clarify. "Dr. Smith, your note says pneumonia. Can you specify the type and causative organism?" The physician updates the note, the coder assigns a more specific code, and the claim reflects the actual complexity of care delivered.
The problem is scale. A typical hospital generates hundreds of inpatient notes per day. CDI specialists can review maybe 20-30 charts per day thoroughly. That means most notes never get a CDI review. The ones that do get reviewed are selected by simple heuristics (high-value DRGs, specific service lines) rather than by actual documentation quality. Notes with significant gaps slip through because nobody had time to look at them.
The financial impact is real. The American Health Information Management Association (AHIMA) estimates that hospitals lose 1-5% of potential revenue due to documentation specificity gaps. For a mid-size hospital doing $500M in annual revenue, that's $5-25M left on the table. Not because the care wasn't delivered, but because the documentation didn't capture it precisely enough.
What if you could scan every note as it's written, identify specificity gaps in real time, and suggest clarifications before the chart is even closed? That's what this recipe builds.
The Technology: LLM-Based Documentation Analysis
What CDI Actually Requires
CDI is not a simple text classification problem. It requires understanding three things simultaneously:
- Clinical context. What does the note actually say happened? What diagnoses are mentioned, what treatments were given, what labs were ordered?
- Coding rules. What level of specificity does ICD-10-CM require for each condition? What qualifiers (laterality, acuity, causative organism, stage) are needed for a complete code?
- Gap detection. Where does the clinical context imply information that the documentation doesn't explicitly state? If the labs show E. coli and the antibiotics target gram-negative bacteria, but the note just says "UTI," there's a specificity gap.
Traditional CDI software uses rule-based engines: if the note mentions "heart failure" but doesn't specify systolic vs. diastolic, fire a query. These rules work for common, well-defined gaps. They miss nuanced cases, they generate false positives on notes that actually do contain the specificity elsewhere in the text, and they require constant manual maintenance as coding guidelines change.
LLMs change this equation because they can read and reason about clinical text the way a human CDI specialist does. They understand that "started on Zosyn" implies the physician suspects a gram-negative or anaerobic infection. They understand that "EF 25%" in an echo report means systolic heart failure even if the note doesn't use those words. They can identify what's implied but not stated, which is exactly what CDI is about.
Why LLMs Work Here
They understand medical language natively. Modern LLMs trained on clinical literature understand the relationships between symptoms, diagnoses, treatments, and lab values. They don't need explicit rules mapping "Zosyn" to "gram-negative coverage." They learned these associations from millions of clinical documents.
They can reason about specificity. You can instruct an LLM: "Given this note, identify any diagnoses that could be documented more specifically per ICD-10-CM guidelines." The model understands what "more specifically" means in a coding context because it has seen thousands of examples of specific vs. unspecific documentation.
They handle context across the full note. A rule-based system might flag "heart failure" as lacking specificity in the assessment section, missing that the physician documented "systolic dysfunction, EF 30%" in the cardiac exam three paragraphs earlier. An LLM reads the entire note and understands that the specificity exists, just in a different section. This cuts false-positive queries substantially.
They generate natural-language suggestions. Instead of firing a cryptic alert ("HF: specify type"), an LLM can generate a physician-friendly query: "Your note mentions heart failure. The echocardiogram documents EF 30%, which suggests systolic heart failure (HFrEF). Would you like to specify this in your assessment?" This is closer to how a human CDI specialist would phrase the question.
The Failure Modes (and They're Important)
Hallucinated clinical findings. The most dangerous failure. The model suggests a specificity improvement based on clinical information that isn't actually in the chart. "Your labs suggest E. coli" when no culture results exist yet. This is why CDI suggestions must always be framed as questions, never as assertions, and why physicians must always make the final documentation decision.
Coding rule staleness. ICD-10-CM guidelines update annually. CMS publishes new codes, retires old ones, and changes specificity requirements every October. An LLM's training data has a cutoff. If you're relying on the model's inherent knowledge of coding rules rather than providing current guidelines via retrieval, your suggestions will drift out of date. This is a strong argument for retrieval-augmented generation (RAG) architecture.
Over-querying. A model that flags every possible specificity gap will drown physicians in queries. Alert fatigue is already a massive problem in healthcare IT. If your CDI system generates 15 suggestions per note, physicians will ignore all of them. You need confidence thresholds and prioritization: flag the high-impact gaps, suppress the marginal ones.
Context window limitations. Hospital notes can be long. A multi-day admission with daily progress notes, consult notes, procedure notes, and nursing documentation can easily exceed 50,000 tokens. You need a strategy for handling notes that exceed your model's context window: summarization, chunking with overlap, or selective section analysis.
Physician trust. This is not a technical failure mode, but it kills more CDI programs than any bug. If physicians perceive the system as a revenue-optimization tool rather than a documentation accuracy tool, they'll resist it. The suggestions must be clinically grounded, respectfully phrased, and genuinely helpful for documentation quality. "This will increase your DRG weight" is the wrong framing. "This will ensure your documentation reflects the complexity of care you actually delivered" is the right one.
Retrieval-Augmented Generation for CDI
Pure LLM inference (just sending the note to a model and asking "what's missing?") works surprisingly well for common conditions. But for production CDI, you want RAG architecture. Here's why:
Current coding guidelines. ICD-10-CM Official Guidelines for Coding and Reporting change annually. Embedding the current year's guidelines in a vector store and retrieving relevant sections based on the diagnoses mentioned in the note ensures your suggestions reflect current rules, not the model's potentially outdated training data.
Organization-specific query templates. Every health system has preferred query language, approved query types, and compliance-reviewed phrasing. Retrieving your organization's approved templates and using them to format suggestions ensures consistency and compliance.
Payer-specific requirements. Different payers have different documentation requirements for the same condition. Medicare requires different specificity than commercial payers for certain diagnoses. Retrieving payer-specific rules based on the patient's coverage adds another layer of accuracy.
The RAG pattern here is: extract diagnoses from the note, retrieve relevant coding guidelines and query templates, then ask the LLM to identify gaps and generate suggestions using the retrieved context as ground truth.
The General Architecture Pattern
[Clinical Note] → [Extract Key Clinical Elements] → [Retrieve Coding Guidelines] → [Identify Specificity Gaps] → [Generate CDI Suggestions] → [Prioritize and Filter] → [Present to CDI Specialist / Physician]
Extract Key Clinical Elements. Parse the note to identify diagnoses, procedures, medications, lab values, and clinical findings. This gives you the "what's documented" baseline.
Retrieve Coding Guidelines. Based on the identified diagnoses, pull the relevant ICD-10-CM guidelines, specificity requirements, and your organization's query templates from a knowledge base.
Identify Specificity Gaps. Compare what's documented against what the guidelines require. Where is the documentation less specific than the coding rules demand? Where does clinical context (labs, meds, vitals) imply information not explicitly stated?
Generate CDI Suggestions. For each identified gap, generate a physician-friendly query suggesting the clarification needed. Include the clinical evidence supporting the suggestion.
Prioritize and Filter. Rank suggestions by clinical and financial impact. Suppress low-confidence suggestions. Limit the total number per note to avoid alert fatigue.
Present. Surface suggestions in the CDI specialist's workflow (for traditional review) or directly in the physician's electronic health record (EHR) for concurrent, real-time CDI.
The AWS build lives in a companion page. This recipe covers the problem, the underlying technology, and the vendor-agnostic architecture. For the AWS services, architecture diagram, prerequisites, and the step-by-step pseudocode walkthrough, see the Architecture and Implementation companion. The Python example is linked from there.
Related Recipes
- Recipe 2.1 (Patient Message Response Drafting): Shares the Bedrock inference pattern but for a different text generation use case
- Recipe 2.4 (Prior Authorization Letter Generation): Uses similar clinical element extraction but generates outbound letters rather than internal queries
- Recipe 2.6 (Clinical Note Summarization): Complementary capability; summarization helps CDI specialists review notes faster
- Recipe 7.3 (DRG Prediction): Predicts DRG assignment, which CDI suggestions aim to improve through better documentation
Tags
generative-ai · llm · rag · clinical-documentation · icd-10 · hipaa · bedrock · bedrock-knowledge-bases · dynamodb · lambda