Technology
Sep 08, 20266 min read

AI agents at work: from demo to dependable

Narrow scope, real permissions, human checkpoints and an evaluation set built before the agent.

Dr. Ravindra Shinde
Dr. Ravindra Shinde
Chief Executive Officer
AI agents at work: from demo to dependable

The demo is the easy part

Every organization has now seen the demo. An AI agent reads a request, looks something up, fills in a form, sends a message and reports back, all in a couple of minutes, without anyone touching a keyboard. It is genuinely impressive, and it is also the least informative part of the story.

The hard part is the hundredth run, and the thousandth: the request phrased in a way nobody anticipated, the system that times out halfway through, the document that contains instructions it should never follow. An agent that is right nine times out of ten is a remarkable demo and an unusable employee.

The good news is that the gap between demo and dependable is now well understood. It is an engineering and design problem, not a research problem, and the organizations getting real value from agents are closing it in very similar ways.

What an agent actually is

Strip away the branding and an agent is four things working together:

  • A goal, stated in plain language or derived from an event, such as a new support ticket.
  • A model that decides what to do next.
  • Tools it is allowed to use: search a knowledge base, read a record, call an API, draft an email.
  • A loop that lets it act, observe the result and decide again until the goal is met or it gives up.

Everything that makes an agent reliable or dangerous lives in the last two items. The model chooses; the tools and the loop determine what that choice can actually do.

Where agents earn their keep

The work agents handle well has a recognizable shape. It involves several steps across more than one system, it follows a pattern that a capable new hire could learn in a week, and it has a clear definition of done.

  • Triage and routing. Reading incoming requests, classifying them, gathering the context a person will need and putting each one in the right queue.
  • Document-heavy processing. Extracting details from invoices, claims, contracts or applications, checking them against records and flagging what does not match.
  • Research and compilation. Pulling together the background for a meeting, a bid or an account review from internal and public sources.
  • Operational runbooks. Working through known diagnostic steps when an alert fires, and handing a person a summary rather than a blank screen.
  • Data entry between systems that were never integrated and never will be.

What these have in common is that a person can check the result quickly. That is not a coincidence. It is the single best predictor of whether an agent will succeed.

Six design principles for dependable agents

1. Narrow beats general

A general-purpose assistant that can do anything is very hard to test, because "anything" has no test set. An agent with one job, a short list of tools and a clear finish line can be evaluated, improved and trusted. Build several narrow agents before you build one broad one.

2. Tools are permissions

Every tool you give an agent is a permission you are granting. Read access to a knowledge base is low risk. The ability to issue refunds, change records or email customers is not. Give each agent the least access its job requires, and give it its own identity so its actions are logged and attributable, just as you would for a new member of staff.

3. People approve what cannot be undone

Let the agent prepare everything, then ask a person to confirm actions that are irreversible, expensive or customer-facing. A good checkpoint is fast: the agent shows what it intends to do and why, and approval is one click. As confidence grows, checkpoints can be relaxed task by task, based on evidence rather than optimism.

4. Treat everything it reads as data, not instructions

Agents read emails, web pages and documents written by people you do not control. Some of that content will contain instructions, deliberately or not. This is known as prompt injection, and the defense is architectural: untrusted content must never be able to grant itself new permissions or trigger high-risk actions without a checkpoint. Assume that anything an agent reads may be trying to mislead it.

5. Build the evaluation before the agent

Collect a few hundred real examples of the task, including the awkward ones, along with what a good outcome looks like. Run every version of the agent against them before it reaches production. This test set is the most valuable thing you will build. It lets you change models, prompts and tools with confidence, and it outlives every one of them.

6. Make every run visible

When an agent gets something wrong, you need to see exactly what it read, what it decided and which tools it called. Record the full trace of every run, review samples regularly, and track cost and time per task alongside accuracy. An agent you cannot inspect is an agent you cannot improve.

Connecting agents to your systems

The plumbing has matured quickly. Open standards such as the Model Context Protocol make it far easier to expose your systems to agents as well-defined tools rather than bespoke integrations, and most major platforms now offer agent-ready interfaces. That lowers the cost of getting started. It does not remove the need to decide, deliberately, what each agent can reach.

The practical rule we follow is simple: if a tool would worry you in the hands of a well-meaning new employee on their first day, it needs a checkpoint, a narrower scope or both.

Measuring what matters

Agent projects are often measured by activity: the number of runs, conversations or automated steps. None of that tells you whether the business is better off. Measure the outcome the agent was meant to change.

  • Time from request to resolution, end to end.
  • The share of cases completed without a person having to redo the work.
  • Error rates compared with the manual process, not with perfection.
  • Cost per completed task, including the people who review it.

If those numbers do not move, the agent is a demo with a running cost.

Where to start

  1. Pick one workflow that is frequent, multi-step and easy to check, and where delays are expensive.
  2. Map it as it really happens, including the exceptions people handle without thinking.
  3. Build the evaluation set from real historical cases.
  4. Prototype in weeks, on your own data, with a person approving every action.
  5. Release in stages. Start with suggestions a person accepts or rejects, then let the agent act on its own in the cases where it has proven itself.

Agents will not replace your teams. Used well, they take on the repetitive, cross-system work that slows good people down, and they do it consistently, at any hour. Getting there is less about the cleverness of the model than about the discipline of the system around it. That discipline is exactly what our AI Transformation and Software Engineering teams bring to every agent we build.