How to Measure Underwriting Decision Quality with AI

How to Measure Underwriting Decision Quality with AI

Craig Hangartner

Saba Gobal, CPCU

AI is getting faster at preparing an underwriting decision. It can organize risk information, surface relevant context, check guidelines, and put more information in front of an underwriter before they make a call.

But faster preparation does not prove that the final underwriting decision is better.

Recent research exposes that gap. In a 2026 survey of 350 senior commercial and specialty P&C underwriting professionals in the US and UK, 51% identified saving time on manual administration as AI's biggest contribution so far. Only 21% identified improved decision quality.

For underwriting leaders, that creates a different question. Once AI is assisting the underwriter, how do you know whether risk selection, pricing judgment, referrals, and exceptions are actually getting better?

The answer starts by measuring the decision itself, not just the time surrounding it.

Why Speed Is Not Decision Quality

Underwriting technology is relatively easy to measure when the objective is operational.

A carrier can track time to quote, submissions processed per underwriter, referral turnaround time, or the number of manual steps removed from a workflow.

Decision quality is harder.

A faster decision can still be a poor decision

Consider two underwriting teams using the same AI-assisted workflow.

Team A cuts average review time by 30%. Team B cuts it by only 10%.

Looking only at productivity makes Team A look more successful. But that conclusion changes if Team A also produces more pricing exceptions, greater variance between underwriters, or adverse loss experience in the risks it binds.

Speed measures how efficiently a decision was reached.

Decision quality measures whether it was the right decision given the information available at the time.

Those are different outcomes.

Underwriters are already seeing the difference

The 2026 Underwriting Edge study found that underwriters were far more likely to credit AI with reducing administrative work than improving decisions.

The same research found that underwriters valued richer decision context over raw speed. Respondents preferred tools that surfaced portfolio drift, prior risks, and model logic rather than simply producing information faster.

The next phase of underwriting AI therefore needs a different scorecard.

What Is Underwriting Decision Quality?

Underwriting decision quality is the degree to which an underwriting decision is consistent with the carrier's appetite, supported by the available evidence, appropriately priced for the risk, and defensible after the outcome becomes known.

That definition deliberately separates the quality of the decision from the outcome of the risk.

A good outcome does not always mean a good decision

Insurance contains uncertainty.

An underwriter can make a well-supported decision and still receive a large claim. Another underwriter can make a poorly supported decision and get lucky when no loss occurs.

That is why evaluating underwriting quality solely through eventual loss experience creates a distorted picture.

The carrier also needs to ask what information was available, what guidance was surfaced, what the underwriter decided, and why the decision differed from the expected path.

The decision needs a record

For an AI-assisted decision, that means retaining enough context to answer five questions:

Decision element

Question to capture

Recommendation

What information, risk signal, guideline, or recommendation was surfaced?

Decision

What did the underwriter ultimately decide?

Override

Did the decision differ from the recommendation, guideline, or expected path?

Outcome

What happened to the risk after the decision?

Learning

What should that outcome change about future decisions?

This creates a decision record rather than a simple transaction record.

The distinction matters. A policy administration system can tell you that a risk was bound at a particular price. It does not necessarily tell you why the underwriter believed that was the right decision.

Why Underwriter Overrides Matter

An override is not automatically an error.

In many cases, it is exactly where underwriting expertise becomes visible.

A model or rule works from the information and logic available to it. An experienced underwriter can see context that changes how those inputs should be interpreted.

The problem starts when the carrier does not capture that difference.

Overrides reveal hidden judgment

Imagine an AI-assisted process surfaces a risk as outside normal appetite.

The underwriter reviews the account, identifies a mitigating factor, and writes the risk anyway.

Recording only the final bind decision loses the most valuable information in the interaction.

The useful data is the disagreement:

What did the system see?

What did the underwriter see differently?

Why did the underwriter override it?

What happened afterward?

Repeated across hundreds or thousands of decisions, those differences start to reveal where experienced human judgment adds value.

Repeated overrides reveal something else

Now imagine underwriters repeatedly override the same recommendation for the same reason.

That pattern deserves attention.

The issue can sit in the recommendation logic, an outdated underwriting guideline, missing risk context, or knowledge that experienced underwriters routinely use but the formal process never captured.

Without structured override data, these patterns remain individual anecdotes.

With it, underwriting leadership gets a new source of evidence about how decisions are actually being made.

How Do You Measure Better Underwriting Judgment?

There is no single metric for underwriting decision quality.

A practical measurement system needs to separate consistency, judgment, and eventual performance.

Four steps to measure underwriting judgment: check consistency, test overrides, review referrals, link outcomes.

1. Decision consistency

Start by comparing how similar risks are handled.

If two comparable submissions produce materially different decisions, identify where the paths separated.

That does not mean every underwriter should reach an identical answer. It means meaningful differences should have an identifiable reason.

Useful questions include whether similar risks receive similar pricing treatment, whether the same characteristics trigger referrals, and whether exceptions follow consistent reasoning.

2. Override rate

Track how frequently underwriters disagree with AI-assisted recommendations, guidelines, or risk signals.

The raw percentage is not enough.

Segment overrides by underwriter, line of business, recommendation type, risk characteristic, reason, and eventual outcome.

A high override rate in one area can identify a recommendation that does not reflect how experienced underwriters actually assess that risk.

A very low override rate deserves examination too. Human-in-the-loop only adds value when the human is actively evaluating the evidence rather than routinely accepting the suggested path.

3. Override performance

The next question is whether overrides produce different outcomes.

Compare risks where the recommendation was accepted with risks where an underwriter chose a different path.

Over time, this provides evidence about where human judgment consistently adds information beyond the initial recommendation.

It also prevents a dangerous assumption: that disagreement with AI is either automatically good because it is human judgment or automatically bad because it contradicts the model.

The outcome provides evidence. Neither side receives the benefit of the doubt.

4. Referral quality

Referrals are another useful signal.

Track which risks are referred, why they are referred, what the senior underwriter ultimately decides, and how often similar referrals produce the same resolution.

If senior underwriters repeatedly resolve the same type of referral in the same way, the carrier has identified knowledge that can be surfaced earlier in future decisions.

If similar referrals produce very different answers, leadership has found an area where guidelines or appetite require clarification.

5. Outcome feedback

Finally, connect underwriting decisions with what happens after bind.

Claims experience, renewals, cancellations, pricing changes, portfolio performance, and other downstream results provide information about the original risk assessment.

The purpose is not to judge an individual underwriter every time a claim occurs.

It is to find patterns across groups of similar decisions.

The Decision Learning Loop

The real opportunity is not a better recommendation. It is a better learning system.

A useful way to think about it is the Decision Learning Loop:

Recommendation → Decision → Override → Outcome → Learning

Four-step infographic showing recommendation, decision, override, and learning in a decision feedback loop.

Recommendation

Surface the relevant risk information, historical context, guidelines, and reasoning available before the decision.

Decision

Capture the action the underwriter actually takes.

That includes bind, decline, refer, pricing adjustment, coverage modification, or another material underwriting action.

Override

Record where the underwriter's decision differs from the expected path and why.

The reason matters more than the fact that an override occurred.

Outcome

Connect the original decision with downstream experience.

That creates evidence for evaluating patterns rather than relying solely on whether a recommendation looked reasonable at the time.

Learning

Feed the finding back into underwriting.

Sometimes the recommendation needs adjustment. Sometimes the underwriting guideline needs clarification. Sometimes the underwriter identified context worth surfacing to everyone else.

Without this final step, carriers collect more decision data without becoming better at making decisions.

What Should a CUO See?

A Chief Underwriting Officer does not need another dashboard showing how many AI interactions occurred.

Usage does not establish underwriting value.

A decision-quality view should instead help leadership answer questions such as:

Leadership question

Signal

Are comparable risks being treated consistently?

Decision variance

Where are underwriters disagreeing with recommendations?

Override rate

Which disagreements perform differently over time?

Override outcomes

Which risks repeatedly require senior review?

Referral patterns

Where do guidelines and actual decisions diverge?

Exception patterns

What underwriting knowledge exists only with experienced staff?

Repeated override reasoning

Are past outcomes changing future decisions?

Feedback-loop adoption

This shifts the measurement conversation from "Are underwriters using AI?" to "What are we learning about how our best underwriting decisions are made?"

That is a much higher bar.

It is also where AI-assisted underwriting becomes more valuable to underwriting leadership.

Avoid Turning Judgment Into a Score

There is a risk in measuring underwriting decisions too aggressively.

Not every underwriting judgment should become a numerical score.

Complex specialty risks contain context, uncertainty, broker information, market conditions, portfolio considerations, and professional judgment that do not collapse neatly into one number.

The goal is not to eliminate judgment.

The goal is to make the evidence around judgment more visible.

Keep the underwriter responsible for the decision

AI can surface information, identify patterns, and provide decision support.

The underwriter still reviews the evidence and makes the final call.

That distinction is important operationally and from a governance perspective. Insurance regulators continue to hold insurers responsible for decisions involving AI, including expectations around accuracy, fairness, explainability, and human oversight.

Measure patterns, not isolated mistakes

A single bad loss does not prove the underwriting decision was wrong.

A single successful override does not prove the underwriter knew better than the model.

Patterns matter.

When the same decision behavior repeats across similar risks and enough outcomes accumulate, underwriting leadership has something worth investigating.

That is the level at which decision-quality measurement becomes useful.

How InsOps Helps

InsOps builds insurance-trained AI designed to assist insurance professionals while keeping people responsible for final decisions.

LiLa, our insurance-trained LLM, runs inside the insurer's own environment, so policy, underwriting, claims, PII, and PHI stay within controlled infrastructure. Human review remains part of the process.

InsOps' Integration Gateway connects data across Guidewire PolicyCenter, UnderwritingCenter, PricingCenter, ClaimCenter, and other systems. LiLa assists with mapping and transforming that information into the structures required by the receiving workflow, with human validation before deployment.

That data foundation matters for decision-quality measurement because the underwriting decision and the eventual outcome often live in different systems.

InsOps is building toward deeper decision-support capabilities that help carriers surface relevant underwriting context while keeping the underwriter in control of the final decision.

If you are evaluating how to connect underwriting context with downstream outcomes and make decision reasoning easier to analyze, contact us to talk through what this can look like for your operation.

Frequently Asked Questions

Q: What is underwriting decision quality?

A: Underwriting decision quality measures whether a decision is supported by available evidence, consistent with appetite and guidelines, appropriately priced for the risk, and defensible based on what was known when the decision was made.

Q: How is underwriting decision quality different from underwriting speed?

A: Speed measures how quickly an underwriting decision is reached. Decision quality examines whether the risk was assessed and treated appropriately. A faster workflow does not establish that the resulting risk selection is better.

Q: How do you measure whether AI is improving underwriting decisions?

A: Compare recommendations with final underwriter decisions, capture overrides and their reasons, and connect those decisions with downstream outcomes. Look for patterns across similar risks rather than judging isolated decisions.

Q: Is an underwriter override a sign that the AI recommendation was wrong?

A: No. An override shows that the underwriter reached a different conclusion. The reason for that difference and the performance of similar decisions over time provide the evidence needed to evaluate it.

Q: What underwriting metrics should a CUO track?

A: Decision variance, override rates, override outcomes, referral patterns, exception patterns, guideline divergence, and the degree to which downstream experience feeds back into future underwriting decisions.

Q: Why connect claims outcomes back to underwriting decisions?

A: Claims provide evidence about how insured risks actually developed. Connecting that experience with the original underwriting context helps carriers identify recurring patterns in risk selection, pricing assumptions, and exceptions.

Q: Does AI replace underwriter judgment in this model?

A: No. AI assists by surfacing information, patterns, and relevant context. The underwriter reviews that information and remains responsible for the final underwriting decision.

Q: Where should insurers start measuring underwriting decision quality?

A: Start with one defined decision type, such as referrals or overrides, in one line of business. Capture the recommendation, final decision, reason for any difference, and eventual outcome before expanding the measurement framework.

Craig Hangartner

Saba Gobal, CPCU