Skip to content
Brihat InfotechBrihat Infotech

AI Engineering

Human-in-the-loop design people actually accept

A review step that shows a raw model output and asks for approval is not oversight — it is a rubber stamp with extra clicks. What makes the human lane work.

Animesh Pathak18 Aug 20263 min read

Every AI system in a regulated context has a human review step, and a large share of them do nothing.

The failure is predictable. A reviewer faces a queue of hundreds, nearly all of which are correct, and is asked to approve or reject a conclusion. Within a fortnight, approval is reflex. The step now certifies errors with a human name attached, which is worse than not having it — because the accountability moved without the scrutiny.

Three design errors that produce a rubber stamp

Reviewing everything

Universal review is the instinctive answer and the one that guarantees failure. It makes the system slower than the manual process it replaced while providing weaker scrutiny than that process gave, because attention is now spread evenly across cases that do not need it.

Route by confidence instead. High-confidence cases proceed. The human lane exists for genuine uncertainty, and it works because the queue is small enough that each item deserves the attention it gets.

Showing the conclusion instead of the evidence

"The system recommends approval. Approve or reject?" invites agreement, because disagreeing requires reconstructing work the reviewer cannot see.

What functions is showing the evidence and letting the human reach the conclusion: the extracted values beside the source document, the passages a claim rests on, the checks that ran and what they returned, and the specific thing that is uncertain. The reviewer is deciding, not adjudicating a machine's verdict.

Hiding the uncertainty

Systems that surface a single confidence percentage tell the reviewer almost nothing. Which part is uncertain? A confidently-read document with an ambiguous policy is a different case from a clean policy applied to an unreadable scan, and they need different attention.

What good looks like

  • The specific uncertainty named — "salary figure read at low confidence from a poor scan", not "review required".
  • Source alongside extraction, so verification is a glance rather than an investigation.
  • The check that failed, and what it expected.
  • A one-click correction path that captures the corrected value, because a reviewer who has to leave the screen to fix something will approve instead.
  • The reviewer's decision recorded as an outcome that feeds the evaluation set.

That last point compounds. Corrections are labelled training data arriving free, from exactly the cases the system found hardest.

The threshold is a business parameter

Where the line sits between straight-through and human review is not a technical constant. It expresses the relative cost of two errors — letting a wrong case through, and delaying a correct one.

Those costs differ by domain and change over time. A fraud threshold in a fraud wave is not the threshold in a quiet quarter. So it belongs in configuration, owned by the business function that carries the consequence, with changes logged.

Set it initially from the evaluation set, where you can see the trade-off explicitly. Then tune against real outcomes rather than intuition.

Measuring whether the lane is working

Four numbers, and one of them is counter-intuitive:

  • Override rate. If it is near zero, the review step has probably stopped working. A healthy lane produces disagreement, because it is receiving genuinely uncertain cases.
  • Time per review. Falling steadily towards a few seconds is the signature of a rubber stamp forming.
  • Post-review error rate — errors that got through with human approval. This is the number that says whether the step catches anything.
  • Queue depth and ageing. A growing queue means the threshold is wrong or the lane is understaffed, and both degrade review quality before they show up anywhere else.

The framing that gets past a risk committee

Not "the AI decides and a human checks", which asks the committee to delegate the judgement it exists to exercise.

Instead: the system clears the roughly eighty per cent that never needed judgment, and assembles evidence for the twenty per cent that does. Decision authority does not move. What moves is the assembly work that used to consume most of the reviewer's day — and reviewers, receiving better-prepared cases in smaller numbers, tend to make better decisions than they did before.

  • ai
  • governance
  • design
Questions this raises

Volume and framing. If a reviewer sees hundreds of items a day and almost all are correct, attention collapses and approval becomes reflex — the step then certifies errors rather than catching them. The fixes are routing only genuinely uncertain cases to humans, and showing the evidence rather than asking for a verdict on a conclusion.

AP

Written by

Animesh Pathak

Founder

Founded Brihat Infotech in 2022 and has led delivery on every engagement since. Works problem-first: map how the organisation actually runs before proposing a system, then stay on the engagement long enough to be accountable for whether it gets used.

Next step

Bring us the problem. We will bring the architecture.

A discovery call takes forty-five minutes. You leave with our read on the problem, the shape of the system we would propose, and a straight answer on whether we are the right team for it.

  • No sales deck
  • An engineer on the call, not an account manager
  • NDA before you share anything