Two years into the agent era, the pattern is clear enough to write down: the enterprises getting returns are not the ones with the most pilots. They are the ones that treated agents as systems engineering rather than model demos.
This is the map. Each section links to the piece that goes into it properly.
1. Where agents actually win
High volume, document-heavy inputs, decisions that are rule-informed but exception-rich. The agent clears the roughly eighty per cent that never needed judgment and assembles context for the rest — which is also the framing that gets past a risk committee, because decision authority never moves.
→ AI agents in the enterprise: what actually works
2. Why projects get cancelled
Almost never the model. Five omissions recur: no evaluation set, no cost ceiling, no action allowlist, no human lane, and integration scoped as plumbing when it is most of the work.
→ Why enterprise agent projects get cancelled
3. Build the evaluation set first
A hundred real questions with verified answers, collected from users during discovery, before the system exists. It is the cheapest artefact in the project and the one whose absence makes every later decision an argument about impressions.
→ Building the evaluation set before the system
4. Retrieval is where quality lives
When a system answers wrongly, autopsy the retrieval first. In our audits four out of five failures happen before the model sees the context — chunks split mid-thought, tables flattened into noise, a superseded policy outranking the current one.
5. Governance, and what actually binds you
ISO 42001 gives structure, NIST AI RMF gives a way to track risk, the EU AI Act applies if you sell into Europe. None of them is Indian law. The DPDP Act and your regulator's outsourcing directions are what decide the architecture.
→ AI governance for regulated Indian enterprises · The DPDP Act and AI systems · Indian regulation that shapes enterprise software
6. The audit trail is the commercial feature
Recording that a decision happened is not an audit trail. Reconstructing why it happened eleven months later, with the model since replaced, is — and in regulated sectors that capability is the condition of going live rather than a nice-to-have.
→ What an AI audit trail actually has to capture
7. The human lane, designed so it works
A review step that shows a raw output and asks for approval becomes a rubber stamp within a fortnight. What functions is routing only genuinely uncertain cases to humans, and showing evidence rather than a conclusion.
→ Human-in-the-loop design people actually accept
8. Build or buy
Buy the commodity. Build where AI touches your edge. Consider the option most people skip — buy the platform, build only the thin differentiating layer.
→ Build vs buy for enterprise AI
9. A worked example
An NBFC origination platform: approval turnaround down 68% from a four-day baseline, underwriters handling 3.2 times as many files a day, every decision carrying an audit trail — achieved by letting agents do everything except approve loans.
→ AI in loan origination, without moving the decision · The case study
Where to start
Pick one queue your operations team already resents. Write down today's number. Ship an agent with a human lane, an action allowlist and a replayable log, running alongside the existing process rather than instead of it. When the metric moves you have evidence, which is what the next budget conversation needs.

