By Tracey Bradberry, MSDS, BS, CHC, CPC, Vice President of Business Excellence

Among the explosion and sophistication of large language models (LLMs) and artificial intelligence, natural language processing (NLP) is proving to be a unique and essential capability for healthcare organizations working to turn complex clinical documentation into structured, actionable insight. Physician notes, charts, reports, and other records contain critical details that shape a variety of areas of healthcare: risk adjustment, quality performance, coding accuracy, compliance, and care delivery. But the sheer volume and variability of that information can make manual review difficult to scale, and AI also has its limitations.

That’s where healthcare-specific NLP continues to add value. It extracts and contextualizes clinical language in ways that can support more complete documentation, stronger decision-making, and more efficient workflows. In healthcare, speed and scale are not enough; NLP must also be accurate, explainable, auditable, and supported by strong governance and human oversight. In this eBook, we will explore what makes healthcare NLP different from generic AI tools, how it works alongside machine learning and LLMs, and what health organizations should consider as they evaluate NLP-driven solutions for clinical and operational use.

Healthcare-specific NLP built for accuracy

As AI, machine learning, and NLP solutions become more widely available, generic models can seem like a convenient choice. But sacrificing specificity can often come at the expense of accuracy. Generic language models are designed to sound “close enough,” but can leave significant gaps and errors.

Healthcare-specific NLP solutions, meanwhile, are trained for a precise purpose. In risk adjustment, for example, actions must go beyond simple documentation summaries; coders need to be able to pull specific ICD-10 codes from a chart—and possibly defend that information in an audit later. That level of accuracy requires the right taxonomy, the right database with code sets, and all the right supporting information.

To meet this need for detail, the best NLP models are built and validated over many years. Healthcare has strict mandates: it requires the ability to show work, demonstrate how the intelligence arrived at certain conclusions, and why it made certain judgements. Healthcare-specific, rule-based NLP has that ability. Every recommendation traces back to a specific piece of evidence, such as a treatment plan, medication, or other documentation in the chart. It is built to answer the question, “Why was this code submitted?”

On the other hand, generative models are probabilistic. They can be right, but it can be harder to explain why they are right because there are many parameters and the output is not always traceable.

How to build a strong healthcare-specific NLP

Like any AI-driven logic, the strength of the NLP is entirely based on what it is fed. In the engine used in coding for risk adjustment, for example, every accept or reject decision is data. The patterns of rejection signify where a rule is too broad or missing context. Validated chart review outcomes provide feedback that helps improve coding logic, model performance, and operational effectiveness over time. 

Good NLP engines and supporting teams should constantly analyze codes to keep the feedback loop ongoing, so the models—whether legacy or new—continue to improve.

Importantly, performance improvements in healthcare NLP are driven not only by technology, but also by the clinical taxonomies, coding methodologies, quality controls, validation processes, governance frameworks, and domain expertise that support the solution. Continuous refinement relies on these capabilities and operational learnings to help drive performance, explainability, and trust in healthcare AI solutions.

Organizations must also be able to know the reason behind decisions. For example, if a coder rejects a code, is the “why” behind that rejection documented? A strong clinical review team looks for that because the feedback loop helps improve performance evaluation, validation, and operation effectiveness. NLP has been doing this on its own for a long time, but LLMs and machine learning models can help increase accuracy and speed.

Building regulatory guidance into the model framework

Healthcare NLP also requires a rock-solid foundation based on regulatory rules. NLP, machine learning, and newer generative AI models all receive those requirements before deployment. Specificity also comes into play for the right use case, feeding the model the relevant guidelines, whether those are CMS documentation guidelines, ICD coding guidelines, or other code-set guidelines.

Additionally, the overall governance that goes into model calibration must be strong enough to catch real risk in the models. A solid approach tiers every use case by risk: some decisions stay at the business-unit level while higher-risk use cases go to a cross-functional governance committee. Legal and compliance are involved and assess model harm across five domains: security, privacy, intellectual property, misinformation, and alignment. The governance committee does not sign off on models until mitigations are validated.

Beyond specific required characteristics, such as chart documentation or a signature, there should be a robust governance framework to assess all models and continue to monitor the models after go live.

This matters across all healthcare domains: when AI touches any member, provider, or submission for payment (CMS, HHS, etc.), evaluation and oversight is necessary.

NLP and LLMs are strongest when used together

Running healthcare-specific NLP and LLMs together can improve both efficiency and coding accuracy because each technology addresses different challenges.

LLMs can be highly effective for common conditions and frequently occurring coding scenarios because they are built to recognize patterns across large volumes of information. However, healthcare doesn’t operate on common patterns alone; it contains tens of thousands of diagnosis codes, many of which are rare, nuanced, or highly dependent on surrounding clinical context. Instead of focusing on information volume, healthcare-specific NLP helps add the precision, sourcing, and contextual understanding needed to identify relevant evidence in the chart, connect it to the appropriate code, and support recommendations that can be reviewed, validated, and defended. 

AI strategies should combine the scale and efficiency of LLMs with the clinical taxonomies, validation processes, governance frameworks, and domain expertise required to support accuracy, explainability, and compliance.

As a result, the most effective approach is not choosing between NLP and LLMs, but combining them. LLMs can drive efficiency and performance on common conditions, while NLP provides the clinical specificity, explainability, and long-tail coverage needed to identify complex cases. Together, they create a more complete and accurate view of patient risk.

Balancing precision and recall in NLP

For risk adjustment, the core performance characteristics are largely the same across NLP, and precision and recall are the base values used to measure.

  • Precision asks “of all the codes the model predicted, how many were actually correct?”  This comes down to the idea of avoiding false positives.

  • Recall asks “of all the codes that should have been found, how many did the model actually catch?” This focuses more on avoiding false negatives or misses.

There's an inherent tradeoff in balancing both. Increasing recall and having a model flag possible codes more aggressively means it catches more true cases but, by the same token, it also increases false positives and lowers precision. Tightening the model to only flag things with a high confidence level raises precision but simultaneously causes it to miss more real cases, lowering recall.

That said, there is a trend towards skewing towards lower precision in HCC coding when it comes to NLP performance. There are several reasons for this:

  • Missed codes are costly and harder to catch later. In risk-adjustment HCC coding, every valid code captures the true complexity/risk of a patient, which affects reimbursement.

  • False positives are easier to fix. If the model over-flags and suggests a code that turns out wrong, a human coder can review and reject. In this way, low precision just means more review work, versus the possibility of low recall creating silent, potentially permanent gaps.

  • Compliance and audit risk favor completeness. Healthcare organizations are generally behooved to catch every possibility and verify them later, rather than risk quietly under-documenting patient risk.

A false negative in HCC coding is often more costly and less apparent than a false positive, making it more common for organizations to tune models to favor recall and flag more rather than less, since human experts can verify the codes as a final safety net.

Human oversight remains essential as AI capabilities evolve

While the role of human oversight will likely evolve, healthcare is not yet ready for fully autonomous decision-making.

Clinical and administrative use cases carry significant regulatory, financial, and patient-care implications, which means accuracy and accountability must remain central to any AI-driven workflow.

While technology-only models are emerging in the market, payers and providers still need confidence that recommendations can be reviewed, validated, and defended before being put into action. Regulatory acceptance is also an important factor, particularly as organizations look to align innovation with CMS requirements and broader compliance expectations. Human-in-the-loop oversight remains essential, helping ensure that NLP, machine learning, and generative AI can support healthcare teams without replacing the judgment and governance required for responsible use.

Moving healthcare forward with accuracy and accountability

As healthcare organizations continue exploring the possibilities of AI, NLP is a critical bridge between unstructured clinical documentation and the structured insight needed to improve performance. But users should take care to remember that the value of NLP in healthcare depends on more than speed or automation alone. It requires models that are purpose-built for clinical context, calibrated to the right use cases, supported by high-quality data, and governed with the transparency required for auditability, compliance, and trust.

Used together, healthcare-specific NLP, machine learning, and LLMs can help organizations capture a more complete view of clinical information, from common conditions to complex edge cases that might otherwise be missed. The strongest strategies will not treat these technologies as competing approaches, but as complementary capabilities that work best when paired with explainability, continuous refinement, and human oversight.

Ultimately, responsible NLP adoption is about helping healthcare teams make better-informed decisions without losing sight of accuracy, accountability, and clinical judgment. Organizations that evaluate NLP through that lens will be better positioned to scale innovation responsibly, improve operational efficiency, strengthen compliance, and turn healthcare’s vast amount of clinical language into insight that supports better outcomes.

Enhance your risk adjustment with NLP

Interested in learning more about Cotiviti’s NLP-powered solutions for provider risk adjustment? Take a look at our brochure and learn how to get started.

 

 
TraceyBradberry2_200x200

 

About the author

Tracey is responsible for the Data Governance, Data Science, Business Intelligence, Innovation and Business Partnerships teams in Cotiviti's RQC/Payment Retrieval Operations. Since beginning her career almost three decades ago , she has a proven record of directing transformation, building relationships, driving operational excellence, starting new programs, and leading cross functional teams. Tracey is passionate about working with client and operations teams to optimize business performance and lead transformation efforts.