AI Evaluation and Observability: Why Your Evals Pass While Users Complain

TL;DR
  • If your eval scores look steady while the business says quality has slipped, start with what your evaluation measures and who defined it. The model is rarely the first thing to fix.
  • Rebuild your test set from recent production traffic and have a business owner sign off on the criteria. Then check your LLM judge against their pass or fail calls before you rely on it.
  • Before you tune another prompt or buy another platform, find out how often your eval score and your business owner disagree on the same outputs. That one number tells you how much to trust everything else on the dashboard.
  • Model swaps and expansion approvals both rest on a quality number the business believes. Fixing the definition of good is what makes those calls defensible.

‍

You've probably sat through some version of this review. Your AI evaluation and observability dashboard shows faithfulness and relevance holding steady for four releases, and the LLM judge passes nine responses out of ten. Then the head of customer operations mentions that her team rewrites drafted replies so often they've stopped counting, and asks why that isn't on the dashboard anywhere.

It's a fair question. Your stack records what the system did on every request, but it can't tell you whether the business would have accepted the answer. That blind spot shows up in plenty of AI agent observability and evaluation setups, and it's where the business's trust starts to slip, usually months before anyone calls it a quality problem.

What AI Evaluation and Observability Each Tell You

Evaluation and observability answer different questions, and each one fails in its own way without the other.

AI evaluation AI observability
The question it answers Did this output meet our standard? What did the system do to produce it?
What it looks at Outputs scored against a rubric or test set Traces of every prompt and tool call
What goes wrong without the other You see a score drop and can't debug why You watch every step and still can't say if any of it was good

‍

The complaint in that review usually means you have both, and neither one is measuring what the business needs.

‍

Why AI Evaluation Scores and User Complaints Disagree

Your eval suite answers the question your team settled at design time: does this output meet the AI evaluation metrics we wrote down? Your business owner is asking something else every day. She wants to know if she'd put her name on it. Both answers can be accurate and still point in opposite directions.

You're far from alone here. LangChain surveyed 1,340 practitioners at the end of 2025 for its 2026 State of Agent Engineering report, and the gap shows up clearly in the results:

  • ‍89% have some form of observability on their agents
  • ‍52.4% run offline evals against test sets
  • ‍37.3% evaluate live traffic
  • ‍Quality is the top barrier to production, for the second year running

So most enterprises can see every step their AI took, and far fewer can say with confidence that the result was good. Fewer still have handed that judgment to the person who owns the outcome. Among the largest companies in the survey, hallucinations and inconsistent outputs came up as the hardest quality problems, and both are easy for an averaged score to hide.

How to Measure the Agreement Rate Between Your Evals and the Business

The most useful number for closing this gap is the agreement rate, and measuring it takes no new tooling.

  1. Pull a sample. Take 200 real production outputs from one application.
  2. Score it your way. Run your current eval suite on all 200.
  3. Score it the business's way. Have the business owner judge the same 200 independently, without seeing the eval results.
  4. Compare. The agreement rate is the share of cases where both verdicts match.

The disagreements are where the value sits, and they come in two kinds.

Owner says pass Owner says fail
Eval says pass Agreement The risky one. Your dashboard is hiding a problem the business can already see.
Eval says fail Your evals are stricter than the business, so engineers fix things nobody asked for. Agreement

‍

A low rate means your evals are measuring something the business doesn't recognize as quality. A very high rate on the first check deserves a second look too, since it can mean the owner hasn't seen enough edge cases yet. Either way, the number matters less than what the disagreements show you.

See how far your eval scores sit from production

When the dashboard and the business tell different stories, the first job is measuring the distance between them on real traffic. A working session takes one production application and compares its eval scores with how the business judged the same outputs.

‍Compare your eval scores against production

‍

That gap rarely comes from one place, which is why tuning prompts against the current scorecard so seldom closes it.

Four Places Your AI Evaluation Drifts From the Business

There are four places to check first, and none of them needs new tooling.

Where it drifts What you'll notice Ask your team
Test set frozen at launch Flat scores while new request types show up in production logs When did we last rebuild the eval set from real traffic?
Engineers wrote the criteria The rubric checks grounding and format, but complaints are about judgment calls Which business owner signed off on what "passing" means?
LLM judge never calibrated A high pass rate, and nobody knows how often it agrees with an expert What's the judge's agreement rate with a human reviewer?
User behavior stuck in observability Edits and escalations sit in traces and never reach the eval suite Does any eval look at what users did after the response?

A test set frozen at launch

The eval set is usually written in the weeks before go-live, by people guessing what users will ask. Once real traffic arrives, the mix of requests shifts, and the frozen set keeps grading the system on questions users have stopped asking. Scores stay flat while the cases that actually reach customers go unmeasured.

Criteria written only by engineers

Engineers check what they can see in an output, like grounding and format, while the business judges tone and policy. In a 2024 study, Annalisa Szymanski and her co-authors compared the evaluation criteria domain experts wrote with criteria from lay users and from LLMs. They concluded that experts need to be in the loop early, because developers working alone lack the domain knowledge to set criteria that hold up.

An LLM judge nobody calibrated

It's easy to assume an LLM judge is objective because it isn't a person. Look at how it was built, though. An engineer wrote the prompt, using criteria an engineer chose and examples an engineer labeled.

            An uncalibrated LLM judge is one engineer's opinion, automated.

Until someone checks that opinion against the business owner's, the pass rate doesn't deserve anyone's trust, and a high one can mislead people more than a low one would. Hamel Husain, who has helped more than 30 companies set up evaluation systems, lists unvalidated metrics among the mistakes he sees most. He builds judges from an expert's pass or fail calls with written critiques, then tests them against those calls.

User behavior that never leaves observability

Your traces record every edit and escalation, but those signals usually live in a monitoring tool nobody on the eval side reads. The behavior that best reflects your business owner's judgment sits right next to the evaluation without ever feeding into it. The section on observability signals further down covers how to connect the two.

Together, those four causes create costs the business feels long before your engineering metrics move.

What the Eval Gap Costs Before Anyone Calls It Quality

Trust in your numbers goes first: Once your business owner has seen a green dashboard next to a team that's rewriting outputs by hand, she'll discount every score you show her after that, accurate ones included.

Workarounds start where you can't see them: People who don't trust the outputs check everything themselves, and the efficiency case behind the application quietly wears away. Your usage numbers still look fine, because people are still calling the system.

Engineering effort goes to the wrong target: Your team spends sprints tuning prompts and retrieval to push up a score the business has stopped believing. A release that adds two points of faithfulness can still make the owner's week worse, and nothing on your dashboard will show it.

The next decisions slow down: Moving to a cheaper model needs a sponsor who believes the current system works, and so does every request to roll the application out to another function.

Some of that cost comes from asking one kind of evaluation to do a job it was never set up for.

Offline vs. Online Evaluation: What Each One Should Catch

Offline evaluation runs a fixed test set before each release, so it's good at telling you if a change broke something you already knew to check. Online evaluation scores a sample of live traffic after release, so it catches the requests nobody thought to write a test for.

Offline evals are useful. The trouble starts when they're the only evals you run. Your test set was written before users arrived, by engineers imagining what users would ask. Real users bring phrasings and edge cases nobody anticipated, and by the time that gap shows up on the dashboard, the business owner has usually stopped trusting the scores.

LangChain's survey suggests plenty of programs fund the first and skip the second. Among respondents with agents already in production, 44.8% run online evals, compared with 37.3% across the whole sample, and only about a quarter of organizations running evals use both approaches.

Offline evaluation Online evaluation
When it runs Before each release, on a fixed, versioned set After release, on sampled live traffic
What it catches Regressions on cases you already know about New request types and slow drift in quality
Who reviews failures Engineering, using the owner's rubric The business owner, on the worst-scoring sample
What breaks it A set that never gets refreshed A judge nobody has checked against the owner

‍

If your evals never see live traffic, they can't see the requests your business owner is complaining about. Online evaluation is the part of LLM evaluation in production that does.

Put a number on the gap before your next budget review

Right now there are two versions of quality in play, and the one your sponsor believes isn't on your dashboard. Nobody has checked how often your eval score and your business owner agree on the same outputs. A working session samples real production interactions from one application, has the owner judge them, and hands you the agreement rate plus the exact places where the two split.

‍Book an eval gap working session

‍

Getting that split right still leaves the harder question of who decides what passes.

Rebuilding AI Evaluation Around What the Business Counts as Good

None of this needs new tooling. Testing how an agent behaves when a tool fails is a separate job, covered in your agent reliability testing. Three changes do most of the work here.

  1. Give each application one owner for quality. Pick one person on the business side and make their judgment the definition of a passing answer. They make pass or fail calls on a sample of real outputs and add a sentence on each failure. When the owner and engineering disagree, the owner wins and the case goes into the test set.
  2. Rebuild your test set from production traffic. Swap the pre-launch eval set for a sample of recent production interactions and refresh it on a fixed schedule. Keep the old set as a regression floor, and version every set so you can tell a system change from a test change.
  3. Show judge agreement next to every score. Before an LLM judge's score goes on a dashboard, check how often it agrees with the business owner, and publish that rate beside the score. A 92% pass rate from a judge that matches the owner 60% of the time doesn't tell you much, and your sponsor should know which kind of number they're looking at.

The owner's role only works if what they write down can actually be scored, which is where most rubrics fall apart.

Writing Evaluation Criteria the Business Can Sign Off On

Generic metrics like faithfulness and relevance are useful to your engineers, but a business owner can't sign off on a score of 0.87. They can sign off on a sentence.

Start with pass or fail. Husain recommends binary judgments, each with a short written critique, over 1-to-5 scales. A critique explains the call, and it can later become a worked example for your LLM judge. Score scales mostly produce arguments about the difference between a 3 and a 4, which nobody on the business side cares about.

Then turn the critiques into criteria. Here's a hypothetical from a customer support application, showing one criterion before and after the business owner has reviewed real outputs.

Before owner review After owner review
The response should be accurate and helpful. The response must not commit to a refund timeline that differs from the published 5 to 7 business day policy, even in an apologetic tone. A reply promising to get this sorted quickly fails if no ticket was raised.

‍

An LLM judge can score the second version. The first one gets scored on whatever the judge decides helpful means.

A criterion is ready to use when:

  • the owner would reject a failing output on sight
  • it's written in words the owner uses herself
  • it describes a failure she has actually seen in production
  • a judge can apply it without guessing what a vague word means

Revisit the list whenever the owner fails an output the rubric doesn't cover. That usually means she's applying a rule she hasn't said out loud yet.

Criteria cover what the owner reviews. Observability covers the much larger share of traffic she'll never see.

Turning Observability Signals Into AI Evaluation Data

Your traces already record what people did after each response. At full traffic volume, that behavior is the closest thing you have to the owner's judgment, and it often never reaches the eval suite.

Signal in your traces What it usually means How to use it
Edits before sending The output was close but wrong in a way the user had to fix Sample heavily edited outputs for owner review
Regeneration The first answer missed, often on tone or format Track the rate by request type to find weak spots
Rephrased follow-up The system misread the question Add the original question to the test set
Handoff to a person The user gave up on the AI for this task Send every handoff to owner review
Abandoned session Could be success or frustration Use only alongside other signals

‍

Each cycle, route the interactions with the strongest negative signals into the owner's review queue, so the complaints the business raises are the first cases your evaluation sees. You also get a business-facing number almost for free. When the rewrite rate falls while your eval score rises, the two are finally telling the same story, and agent observability and evaluation start working on the same problem.

With signals flowing in, the remaining risk is that nobody clearly owns any of it.

Who Owns What in AI Evaluation

Eval programs tend to stall on ownership, so write the split down before you build anything. Here's a starting point for a single application.

Role Owns Cadence
Business owner The rubric and the final say on disputed cases A short review each cycle
Application engineering team Test set refreshes and release gating Every release
AI platform or MLOps team The sampling pipeline and judge calibration Ongoing
You, as the engineering leader The agreement rate and the report to your sponsor Monthly

‍

Two details decide if this holds up:

  • The owner's review has to fit inside a normal week. A well-chosen sample matters more than coverage.
  • The agreement rate needs a named owner on the engineering side. Without one, it turns into a number everyone looks at and nobody acts on.

Ownership keeps the work running. Reporting is what keeps it funded.

Reporting AI Evaluation Results Your Sponsor Will Trust

A dashboard of faithfulness and relevance scores is built for your team. Your sponsor reads it as a claim they have no way to check, so lead with numbers they already recognize.

Put this on the report Why your sponsor cares
The agreement rate, at the top It tells them how far to trust every number beneath it
One business outcome next to your eval score, such as rewrite rate or escalation rate It shows if your scores move with what the business experiences
Scores broken down by segments the owner chose An average can look healthy while users in one language or customer tier get a much worse experience
Decisions evaluation changed A release you held back or a model swap you approved shows evaluation earning its budget

Once your sponsor sees evaluation changing decisions, the next budget conversation gets a lot easier.

None of this has to start everywhere at once.

Where to Start: One Application, One Number

  1. Pick the application the business complains about most. The owner there already has a reason to give you their time.
  2. Run the agreement-rate check on 200 of its recent production interactions, with a one-line reason on every fail.
  3. Match the disagreements to the four causes. The owner's reasons will tell you which one you're dealing with.
  4. Fix that application's measurement and show the owner the agreement rate going up.
  5. Then extend the pattern. Once a second and third application have owners and calibrated judges, the sampling pipeline and calibration tooling are worth building once and sharing, the same logic behind deciding what to centralize across your agent stack.

This work sits between your engineering team and your business owners, and it tends to stall because nobody's job covers both sides.

Why Ideas2IT for AI Evaluation and Observability

Ideas2IT closes this gap with Forward Deployed Engineers who embed in your environment from the first day. Delivery is platform-led, so the tooling FDEs use for production sampling and judge calibration is built and maintained by Ideas2IT, and it plugs into the evaluation and observability stack you already run. Nothing gets ripped out, and the rubric stays with your business owner after the engagement ends.

Ideas2IT has been shipping production AI for enterprise clients since 2017, and that has shaped a firm view. Evaluation platforms tend to lead with a dashboard and skip the question of whose judgment passing reflects. A well-instrumented view of the wrong metric does more damage than having no metric, because it gives everyone something to point to while the real problem stays hidden. So every Ideas2IT engagement starts with the business owner's pass or fail calls, and the dashboard comes last.

In our engagements, the first agreement-rate check usually comes back lower than anyone expected.

What the first working session gives you:

  • the agreement rate between your current evals and your business owner on one application
  • the cause behind each disagreement, mapped to the four drift points above
  • a rebuilt rubric for that application, based on where the two disagreed
  • a recommendation on which application to take on next

If your evaluation pipelines handle customer or operational data, Ideas2IT is SOC 2 Type II and ISO 27001 certified, and an AWS GenAI Specialist Partner.

Start with the application your business owners trust least

‍

References

  1. LangChain. State of Agent Engineering. LangChain. June 2026 (survey conducted November to December 2025). https://www.langchain.com/state-of-agent-engineering
  2. Hamel Husain. Using LLM-as-a-Judge For Evaluation: A Complete Guide. hamel.dev. October 2024. https://hamel.dev/blog/posts/llm-judge/
  3. Annalisa Szymanski et al. Comparing Criteria Development Across Domain Experts, Lay Users, and Models in Large Language Model Evaluation. arXiv. October 2024. https://arxiv.org/abs/2410.02054

Frequently Asked Questions

What agreement rate should an LLM judge reach before we trust it?
There's no universal threshold, so set one with the business owner based on what a wrong pass costs in that workflow. Keep tracking it after launch, since agreement can slip as production inputs change.
Does every AI application need its own evaluation criteria?
Yes, because a passing answer means something different in a support workflow than in a contract review. The sampling and calibration tooling underneath can be shared once more than one application depends on it.
Can an LLM judge replace human review entirely?
No. A calibrated judge extends human review across more traffic, and LangChain's 2026 survey found 59.8% of teams running evals still use human review alongside the 53.3% using LLM judges.
How do we stop the team from tuning prompts to the eval set?
Keep a held-out slice of production samples that engineers never see during prompt work, and score every release against it. If scores on the working set climb while the held-out slice stays flat, the team is fitting the test.
Can we use synthetic data for AI evaluation?
Synthetic inputs help you cover cases you haven't seen in production yet, but the owner's pass or fail calls should be made on real outputs. Husain's own process uses LLMs to generate user inputs only, never the responses being judged.
How do we keep business owners engaged after the first review?
Keep each review small and tie it to a decision they care about, such as approving a rollout or a model change. Owners stay involved when they can see their critiques changing what ships.