A compact operations desk with abstract branching activity paths on monitors, a physical stop control and blue, teal and amber indicator lights, with no text, logos or faces.

How UK businesses should monitor AI agents

10 min read

UK businesses should connect AI agent traces to business acceptance checks, enforceable permissions and cost controls. Start with a bounded pilot, test alerts and suspension, and assign responsibility before expanding access.

Written by Kate Bennett Group CEO, Compare the Cloud

UK small and mid-sized businesses should monitor AI agents at three levels: whether the task succeeded, whether each action was authorised, and what the complete run cost. Start with one business workflow, a named owner and alerts that trigger a defined response. Pair an inspectable activity record with limits enforced by the application. Require approval before consequential actions, and prove that both alerts and stop controls work before expanding access.

Define what a successful business task looks like

For this guide, an agent is an application that uses a model to choose steps and invoke tools, such as searching documents or updating a business system. Monitoring should answer what it attempted, what actually changed and whether that change met the business requirement.

Consider a hypothetical UK wholesaler using an agent to prepare responses to delivery enquiries. Define success as finding the correct order, checking its latest delivery status and producing a draft supported by that record. Treat sending the message as a separate action requiring permission.

Record distinct outcomes such as “draft prepared”, “approval pending”, “message sent” and “failed”. Do not allow a fluent final response to stand in for evidence that the intended action happened.

Give the workflow a business owner and a technical owner. In a small company, the technical role might sit with a development partner, but the business should retain authority over permitted actions, spending limits and suspension.

Monitor an AI agent workflow
Trace each run from request to outcome, enforce tool permissions and limits in the application, and route alerts to named owners.

Connect the activity record to controls and people

OpenAI’s SDK documentation describes traces containing model calls, tool calls, handoffs and guardrails. Its trace reference also supports workflow names, grouping identifiers and metadata for filtering related activity. These provide useful building blocks for an investigation. OpenAI tracing and trace reference

Use a consistent run identifier to connect the agent’s activity to the receiving system’s audit record. For each attempted change, retain the target record, action, approval reference, result and timestamp, while excluding unnecessary sensitive content.

The following is a proposed responsibility model, not a feature list for any particular supplier.

LayerWhat to record or enforceAccountable owner
Business outcomeExpected result, acceptance check and actual destination stateProcess owner
Agent runtimeWorkflow version, model, steps, retries, duration and errorsDeveloper or integration partner
Tool boundaryPermitted operation, target, approval and resultApplication owner
Monitoring serviceTrace arrival, missing records, alert delivery and accessIT operations owner
Commercial oversightModel usage, external service charges and monitoring consumptionFinance and technical owners

Enforce permissions where the application invokes a tool. An instruction telling an agent to seek approval should not be the only barrier protecting a payment, deletion or external message.

For consequential operations, bind approval to the precise proposed action. If the recipient, amount or target changes, require a fresh approval.

Monitor mistakes, unexpected actions and cost separately

Check outcomes against business evidence

Create acceptance checks before choosing dashboard charts. For the wholesaler example, checks might confirm that the order identifier matches the enquiry, the delivery claim has supporting evidence and no message was sent without approval.

Review a sample of apparently successful runs as well as all reported failures. Record corrections against the workflow and prompt version so the team can identify whether a release introduced a recurring mistake.

Keep a reusable test set containing representative requests, missing information, ambiguous instructions and known past failures. Run it before changing the model, prompt, tools or permissions. Treat automated assessments as an additional signal, with human review for consequential judgements.

Watch for actions outside the intended workflow

Record denied operations alongside completed ones. A blocked attempt to access an unrelated customer record deserves investigation even when the permission control worked.

Alert on repeated tool calls, unexpected destinations, unusual handoffs and exhausted retry limits. Set the thresholds from the workflow’s normal behaviour and business consequences rather than borrowing a universal number.

Build separate application limits for elapsed time, model turns, tool calls and retries. A spending notification and a control that prevents another chargeable action should have separate acceptance tests.

Measure cost against accepted results

Track model usage by workflow, model and version, then connect it to accepted business outcomes. Include unsuccessful attempts and retries in the cost numerator.

A useful operational measure is total attributable operating cost divided by accepted completed tasks over the same period. Keep the underlying task count visible so a changing mix of easy and difficult requests does not obscure the explanation.

Langfuse documents usage and cost tracking across models and use cases, with dashboards and threshold alerts. It can infer costs from matching model definitions when costs are not supplied directly. Check those definitions against the rates that apply to your account. Langfuse model usage and cost tracking

For a UK budget, retain the original billing currency and document the exchange rate used in internal GBP reporting. Reconcile against invoices, including relevant VAT treatment, rather than presenting a dashboard estimate as the final amount payable.

Budget for monitoring as well as model calls

A useful budget separates building the controls, running the agent and retaining evidence. Obtain commercial terms for the actual deployment, including billing commitment, region, retention, support and tax treatment.

The available evidence does not establish comparable current UK subscription prices. Use the following cost components to request quotations and assess the pilot.

Cost componentWhat to measure or request
Model consumptionUsage by model and usage type, including retries and evaluation calls
Connected servicesSearch, document processing, storage and other chargeable tool operations
Monitoring ingestionThe supplier’s definition of a billable record and expected records per task
Retention and accessRequired investigation period, storage charges and staff access terms
ImplementationInstrumentation, redaction, dashboards, alerts and control testing
Ongoing operationReview time, incident handling, maintenance and self-hosted infrastructure
ExitExport format, retrieval costs, migration work and deletion arrangements

Langfuse’s own documentation illustrates why task counts alone are insufficient for budgeting. Its example contains 20,070 traces, 119,500 observations and 561 scores, totalling 140,131 units per month. These are documentation example figures, not a forecast for a typical UK business. Langfuse billable units

The relevant formula is traces plus observations plus scores. Measure all three during the pilot, and include evaluation activity when estimating monitoring consumption. For self-hosted Langfuse, the same volume remains useful for infrastructure sizing even though the open-source edition has no usage-based billing. Langfuse billable units

Roll out monitoring with a tested stop procedure

Use this checklist before expanding an agent’s access or workload.

  • [ ] Name the owners. Record who accepts business outcomes, handles alerts and can suspend the workflow. Confirm cover during the hours the agent operates.
  • [ ] Define permitted actions. Document allowed systems, operations and targets. Prove that an out-of-scope request is denied at the tool boundary.
  • [ ] Protect recovery options. Before testing writes, confirm backups or an appropriate application recovery method. Record what cannot be reversed.
  • [ ] Minimise trace content. Use synthetic records first. Inspect exported data for credentials, unnecessary personal information and confidential document contents.
  • [ ] Instrument a complete task. Follow a test request from entry through tool calls to its destination record, using a consistent identifier.
  • [ ] Test missing telemetry. Interrupt trace delivery in a controlled environment and confirm that the monitoring owner receives a notification.
  • [ ] Exercise failure paths. Test unavailable tools, invalid inputs, denied permissions and ambiguous requests. Confirm the expected safe outcome.
  • [ ] Prove approval controls. Attempt a consequential action without approval and with approval for a different action. Both attempts should be blocked.
  • [ ] Prove usage limits. Trigger the configured retry, duration and spending controls with safe test workloads. Check what happens to work already in progress.
  • [ ] Run a bounded pilot. Start with read-only access or draft outputs where practical. Review outcomes and costs before increasing permissions.
  • [ ] Train the responders. Give staff a short procedure covering investigation, suspension, escalation and recovery. Have someone other than the implementer follow it.
  • [ ] Rehearse rollback. Disable the workflow, restore the approved version and verify the destination state before resuming.

Warn affected users before disruptive testing or suspension. Stopping the agent does not establish whether an earlier external action completed.

If a tool times out after attempting a write, inspect the receiving system before retrying. Where the destination supports it, use a stable operation identifier to prevent duplicate execution, and test that behaviour explicitly.

Choose a monitoring route that fits your team

The strongest product evidence supplied for this guide covers OpenAI’s built-in tracing, Langfuse, Arize AX and Respan. The comparison below is limited to those documented capabilities; suitability statements are editorial judgements.

OptionDocumented capabilitySensible starting pointChecks before adoption
OpenAI Agents SDK tracingDefault recording of model calls, tools, handoffs and guardrails, with custom trace processors. DocumentationA team already operating the SDK that needs to inspect workflow behaviourRetention, access, export handling and compatibility with your data policy; tracing is unavailable under OpenAI Zero Data Retention
LangfuseModel usage and cost tracking, plus cloud and self-hosted billing models. Cost tracking and billable unitsA team prioritising spend analysis or evaluating a self-hosted optionExact integration coverage, model price definitions, access requirements and responsibility for hosting
Arize AXAn OpenAI Agents SDK integration capturing agent activity, tools, handoffs and model calls. Integration documentationA team assessing a separate observability destinationDeployment compatibility, commercial terms, trace arrival and support scope
RespanOpenAI Agents SDK instrumentation with workflow records and attributes for grouping activity. Integration documentationA team assessing workflow and conversation-level investigationExported data, integration behaviour, retention, commercial terms and operational ownership

For a small business without a developer, make evidence delivery part of the implementation contract. Require a working demonstration of a failed task, a denied action, a usage alert and suspension. Specify whether the partner merely installs monitoring or also investigates alerts.

For a mid-sized business with an operations team, assess whether agent incidents can enter the existing support process. Avoid selecting a separate console without deciding who will check it.

Ask each supplier where traces are stored, who can access them and how export and deletion work. The supplied evidence does not establish UK hosting or UK support coverage for these options.

Editorial analysis

Choose the smallest monitoring setup that lets your team reconstruct a disputed action and intervene before another harmful action occurs. A detailed trace is useful evidence, but business acceptance checks determine whether the task was done correctly.

Keep an existing monitoring service if it can collect the required records, reach the responsible person and support the investigation. Add an agent-specific tool where the pilot demonstrates a concrete gap, such as missing tool-call detail or inadequate cost attribution.

The purchase decision should follow that demonstration. A supplier’s dashboard should earn its place by helping your team resolve a failed workflow, account for its cost and safely restore service.

Sources

The article uses the supplied source extracts. Their date fields do not establish publisher update dates or a fresh page-access date.

Data & Insights

What contributes to Langfuse monitoring units

Langfuse’s documentation example totals 140,131 monthly units across traces, observations and scores; it is not a typical-business forecast.

What contributes to Langfuse monitoring unitsLangfuse’s documentation example totals 140,131 monthly units across traces, observations and scores; it is not a typical-business forecast.Traces20,070 (14.3%)Observations119,500 (85.3%)Scores561 (0.4%)
View the data
What contributes to Langfuse monitoring units
CategoryMonthly units in the documentation example
Traces20,070
Observations119,500
Scores561
Source: Langfuse Billable Units documentation

Frequently Asked Questions

Is an agent’s final answer enough to monitor it?

No. Require evidence of the actual outcome in the receiving system, such as the correct record update or an approved draft. Traces can expose the intermediate model calls, tools and handoffs behind that outcome. OpenAI tracing documentation

Can a monitoring alert stop an agent overspending?

Treat notification and enforcement as separate requirements. Langfuse documents alerts when tracked spend crosses a threshold, but that description does not establish an application spending stop. Test a control in the execution path that prevents further chargeable work under your chosen policy. Langfuse model usage and cost tracking

Should we record every prompt and tool response?

Choose the minimum content needed to investigate the workflow. Start with identifiers, timings, action types and outcomes, then justify any additional content and restrict access. OpenAI’s trace reference explicitly advises considering privacy when adding trace data. OpenAI trace reference

Is self-hosted monitoring free?

Langfuse states that its self-hosted open-source edition has no usage-based billing under the MIT licence. Infrastructure, backups, upgrades and staff time still belong in your budget. Assign those responsibilities before choosing self-hosting. Langfuse billable units

What should we do when traces stop appearing?

Check the telemetry path as a separate incident, including credentials, project settings and exporter behaviour. OpenAI documents background trace export and explicit flushing, while Arize’s integration guide recommends confirming that records actually reach the destination. Keep consequential automation suspended if the missing evidence prevents safe operation. OpenAI tracing and Arize integration troubleshooting

Who should respond outside normal UK business hours?

Name the responsible employee or contracted provider before enabling unattended operation. If nobody is available to handle consequential failures, restrict the workflow to reversible or read-only tasks, or pause it outside supported hours. Test the escalation route as part of acceptance.