A Risk Score Is Only Trustworthy If You Can Explain It
by Dr. Prince Baawuah, Head of Medicare Advantage Solutions
TL;DR: Most organizations do not have one risk score. They have several—calculated by the payer, a vendor, finance, or an internal team—and the numbers often do not match. A trustworthy risk-adjustment operation makes those differences traceable, shows what caused them, and turns each one into the right next action.
Ask three teams for the risk score of the same population and you may get three answers.
Finance may be working from the payer's latest report. Coding may be looking at diagnoses documented this year. An analytics team may have calculated the score from claims received through a different cutoff date. All three numbers can look reasonable. None of them match.
This is the everyday problem in risk adjustment. The calculation itself is manageable. Reconciling who included which members, diagnoses, dates, models, and CMS results is where the work gets hard.
Why Reasonable Numbers Disagree
A risk score sits at the end of a long chain. One team may have a claim that another team has not received yet. A payer may include an encounter that CMS later rejected. Two systems may apply different model years or filtering rules. Even when they start with the same diagnoses, hierarchies and interactions can change which HCCs remain in the final score.
So comparing two totals rarely settles anything. You have to work backward from the difference. Are the member lists the same? Are the data cutoffs the same? Did both calculations use the same model and payment year? Which HCC is present on one side and missing on the other? What claim or encounter put it there?
That is what it means for a score to explain itself. Not just a final number, but a path back to the demographic factor, HCC, diagnosis, source record, and rule that produced it.
Reconciliation Has to Reach CMS
Even a perfectly calculated score from claims is not necessarily the score CMS recognizes. The encounter still has to move through the submission process and survive CMS filtering.
The MAO-002 report shows how an encounter was processed and whether it hit an error. The MAO-004 report shows which diagnoses CMS considered eligible for risk adjustment after filtering. When those reports are available directly or through the payer, they help connect the score in an internal system with the result CMS actually accepted.
That connection often explains a disagreement. The diagnosis may be documented and present in the source claims, yet missing from the payer's score because the encounter was rejected, an identifier did not match, or the service line did not pass filtering. Another apparent discrepancy may simply come from comparing reports produced at different points in time.
CMS keeps its report guidance in the Encounter Data System resources, and its risk-adjustment data validation program is a good reminder that every submitted diagnosis ultimately needs source support.
A Difference Is Not Automatically a Coding Gap
Once you can see why two scores differ, the next action becomes much clearer.
If the diagnosis was documented but the encounter failed, that is an administrative gap for the submission or data team. If a chronic condition supported last year has not been documented this year, that is a recapture gap for appropriate clinical review. If current claims or medication evidence suggests a condition that has not been documented, that is a suspect opportunity—a question for the clinician, not a diagnosis.
Putting all three into one chase list creates work without resolving the disagreement. A provider cannot fix an encounter that never made it through the payer. A data team cannot clinically validate a suspect condition. Reconciliation is useful because it sends the difference to the person who can actually resolve it.
More Payers Make the Same Problem Bigger
With one payer contract, people often reconcile by memory. Someone knows which file is current, which report arrives late, and whom to call when the score looks wrong.
That stops working when one contract becomes four or five. Each payer may provide a different mix of claims, encounter extracts, gap lists, score reports, and CMS-derived files. The formats and timing change, but the basic reconciliation questions do not.
For each contract, teams still need to know what data was used, which rules applied, what they calculated, what the payer recognizes, and why the two results differ. The system has to preserve those contract-specific details while giving everyone one way to investigate the difference.
One Member, One Shared Explanation
Contract-level totals tell you that the numbers disagree. The member view tells you why.
For one member, the teams should be able to compare the score and its HCC components, see the model and cutoff date, trace each condition to its claims or encounters, check its CMS status, and see who owns any unresolved difference. Finance, coding, encounter operations, and provider teams can then work from the same evidence instead of passing spreadsheets back and forth.
Every organization will route that work differently. Some organize by provider group, some by market, and some by payer contract. The workflow can be flexible as long as the evidence and calculation stay consistent.
The goal is not to make one team's number win. It is to give every team the same way to see where the numbers diverge, what caused the difference, and what needs to happen next. That is how a risk score becomes something the organization can actually trust.
See the three-page Falcon Risk Adjustment capability brief →