UK small and mid-sized businesses should monitor AI agents at three levels: whether the task succeeded, whether each action was authorised, and what the complete run cost. Start with one business workflow, a named owner and alerts that trigger a defined response. Pair an inspectable activity record with limits enforced by the application. Require approval before consequential actions, and prove that both alerts and stop controls work before expanding access.
Define what a successful business task looks like
For this guide, an agent is an application that uses a model to choose steps and invoke tools, such as searching documents or updating a business system. Monitoring should answer what it attempted, what actually changed and whether that change met the business requirement.
Consider a hypothetical UK wholesaler using an agent to prepare responses to delivery enquiries. Define success as finding the correct order, checking its latest delivery status and producing a draft supported by that record. Treat sending the message as a separate action requiring permission.
Record distinct outcomes such as “draft prepared”, “approval pending”, “message sent” and “failed”. Do not allow a fluent final response to stand in for evidence that the intended action happened.
Give the workflow a business owner and a technical owner. In a small company, the technical role might sit with a development partner, but the business should retain authority over permitted actions, spending limits and suspension.

Connect the activity record to controls and people
OpenAI’s SDK documentation describes traces containing model calls, tool calls, handoffs and guardrails. Its trace reference also supports workflow names, grouping identifiers and metadata for filtering related activity. These provide useful building blocks for an investigation. OpenAI tracing and trace reference
Use a consistent run identifier to connect the agent’s activity to the receiving system’s audit record. For each attempted change, retain the target record, action, approval reference, result and timestamp, while excluding unnecessary sensitive content.
The following is a proposed responsibility model, not a feature list for any particular supplier.
| Layer | What to record or enforce | Accountable owner |
|---|---|---|
| Business outcome | Expected result, acceptance check and actual destination state | Process owner |
| Agent runtime | Workflow version, model, steps, retries, duration and errors | Developer or integration partner |
| Tool boundary | Permitted operation, target, approval and result | Application owner |
| Monitoring service | Trace arrival, missing records, alert delivery and access | IT operations owner |
| Commercial oversight | Model usage, external service charges and monitoring consumption | Finance and technical owners |
Enforce permissions where the application invokes a tool. An instruction telling an agent to seek approval should not be the only barrier protecting a payment, deletion or external message.
For consequential operations, bind approval to the precise proposed action. If the recipient, amount or target changes, require a fresh approval.
Monitor mistakes, unexpected actions and cost separately
Check outcomes against business evidence
Create acceptance checks before choosing dashboard charts. For the wholesaler example, checks might confirm that the order identifier matches the enquiry, the delivery claim has supporting evidence and no message was sent without approval.
Review a sample of apparently successful runs as well as all reported failures. Record corrections against the workflow and prompt version so the team can identify whether a release introduced a recurring mistake.
Keep a reusable test set containing representative requests, missing information, ambiguous instructions and known past failures. Run it before changing the model, prompt, tools or permissions. Treat automated assessments as an additional signal, with human review for consequential judgements.
Watch for actions outside the intended workflow
Record denied operations alongside completed ones. A blocked attempt to access an unrelated customer record deserves investigation even when the permission control worked.
Alert on repeated tool calls, unexpected destinations, unusual handoffs and exhausted retry limits. Set the thresholds from the workflow’s normal behaviour and business consequences rather than borrowing a universal number.
Build separate application limits for elapsed time, model turns, tool calls and retries. A spending notification and a control that prevents another chargeable action should have separate acceptance tests.
Measure cost against accepted results
Track model usage by workflow, model and version, then connect it to accepted business outcomes. Include unsuccessful attempts and retries in the cost numerator.
A useful operational measure is total attributable operating cost divided by accepted completed tasks over the same period. Keep the underlying task count visible so a changing mix of easy and difficult requests does not obscure the explanation.
Langfuse documents usage and cost tracking across models and use cases, with dashboards and threshold alerts. It can infer costs from matching model definitions when costs are not supplied directly. Check those definitions against the rates that apply to your account. Langfuse model usage and cost tracking
For a UK budget, retain the original billing currency and document the exchange rate used in internal GBP reporting. Reconcile against invoices, including relevant VAT treatment, rather than presenting a dashboard estimate as the final amount payable.
Budget for monitoring as well as model calls
A useful budget separates building the controls, running the agent and retaining evidence. Obtain commercial terms for the actual deployment, including billing commitment, region, retention, support and tax treatment.
The available evidence does not establish comparable current UK subscription prices. Use the following cost components to request quotations and assess the pilot.
| Cost component | What to measure or request |
|---|---|
| Model consumption | Usage by model and usage type, including retries and evaluation calls |
| Connected services | Search, document processing, storage and other chargeable tool operations |
| Monitoring ingestion | The supplier’s definition of a billable record and expected records per task |
| Retention and access | Required investigation period, storage charges and staff access terms |
| Implementation | Instrumentation, redaction, dashboards, alerts and control testing |
| Ongoing operation | Review time, incident handling, maintenance and self-hosted infrastructure |
| Exit | Export format, retrieval costs, migration work and deletion arrangements |
Langfuse’s own documentation illustrates why task counts alone are insufficient for budgeting. Its example contains 20,070 traces, 119,500 observations and 561 scores, totalling 140,131 units per month. These are documentation example figures, not a forecast for a typical UK business. Langfuse billable units
The relevant formula is traces plus observations plus scores. Measure all three during the pilot, and include evaluation activity when estimating monitoring consumption. For self-hosted Langfuse, the same volume remains useful for infrastructure sizing even though the open-source edition has no usage-based billing. Langfuse billable units
Roll out monitoring with a tested stop procedure
Use this checklist before expanding an agent’s access or workload.
- [ ] Name the owners. Record who accepts business outcomes, handles alerts and can suspend the workflow. Confirm cover during the hours the agent operates.
- [ ] Define permitted actions. Document allowed systems, operations and targets. Prove that an out-of-scope request is denied at the tool boundary.
- [ ] Protect recovery options. Before testing writes, confirm backups or an appropriate application recovery method. Record what cannot be reversed.
- [ ] Minimise trace content. Use synthetic records first. Inspect exported data for credentials, unnecessary personal information and confidential document contents.
- [ ] Instrument a complete task. Follow a test request from entry through tool calls to its destination record, using a consistent identifier.
- [ ] Test missing telemetry. Interrupt trace delivery in a controlled environment and confirm that the monitoring owner receives a notification.
- [ ] Exercise failure paths. Test unavailable tools, invalid inputs, denied permissions and ambiguous requests. Confirm the expected safe outcome.
- [ ] Prove approval controls. Attempt a consequential action without approval and with approval for a different action. Both attempts should be blocked.
- [ ] Prove usage limits. Trigger the configured retry, duration and spending controls with safe test workloads. Check what happens to work already in progress.
- [ ] Run a bounded pilot. Start with read-only access or draft outputs where practical. Review outcomes and costs before increasing permissions.
- [ ] Train the responders. Give staff a short procedure covering investigation, suspension, escalation and recovery. Have someone other than the implementer follow it.
- [ ] Rehearse rollback. Disable the workflow, restore the approved version and verify the destination state before resuming.
Warn affected users before disruptive testing or suspension. Stopping the agent does not establish whether an earlier external action completed.
If a tool times out after attempting a write, inspect the receiving system before retrying. Where the destination supports it, use a stable operation identifier to prevent duplicate execution, and test that behaviour explicitly.
Choose a monitoring route that fits your team
The strongest product evidence supplied for this guide covers OpenAI’s built-in tracing, Langfuse, Arize AX and Respan. The comparison below is limited to those documented capabilities; suitability statements are editorial judgements.
| Option | Documented capability | Sensible starting point | Checks before adoption |
|---|---|---|---|
| OpenAI Agents SDK tracing | Default recording of model calls, tools, handoffs and guardrails, with custom trace processors. Documentation | A team already operating the SDK that needs to inspect workflow behaviour | Retention, access, export handling and compatibility with your data policy; tracing is unavailable under OpenAI Zero Data Retention |
| Langfuse | Model usage and cost tracking, plus cloud and self-hosted billing models. Cost tracking and billable units | A team prioritising spend analysis or evaluating a self-hosted option | Exact integration coverage, model price definitions, access requirements and responsibility for hosting |
| Arize AX | An OpenAI Agents SDK integration capturing agent activity, tools, handoffs and model calls. Integration documentation | A team assessing a separate observability destination | Deployment compatibility, commercial terms, trace arrival and support scope |
| Respan | OpenAI Agents SDK instrumentation with workflow records and attributes for grouping activity. Integration documentation | A team assessing workflow and conversation-level investigation | Exported data, integration behaviour, retention, commercial terms and operational ownership |
For a small business without a developer, make evidence delivery part of the implementation contract. Require a working demonstration of a failed task, a denied action, a usage alert and suspension. Specify whether the partner merely installs monitoring or also investigates alerts.
For a mid-sized business with an operations team, assess whether agent incidents can enter the existing support process. Avoid selecting a separate console without deciding who will check it.
Ask each supplier where traces are stored, who can access them and how export and deletion work. The supplied evidence does not establish UK hosting or UK support coverage for these options.
Editorial analysis
Choose the smallest monitoring setup that lets your team reconstruct a disputed action and intervene before another harmful action occurs. A detailed trace is useful evidence, but business acceptance checks determine whether the task was done correctly.
Keep an existing monitoring service if it can collect the required records, reach the responsible person and support the investigation. Add an agent-specific tool where the pilot demonstrates a concrete gap, such as missing tool-call detail or inadequate cost attribution.
The purchase decision should follow that demonstration. A supplier’s dashboard should earn its place by helping your team resolve a failed workflow, account for its cost and safely restore service.
Sources
The article uses the supplied source extracts. Their date fields do not establish publisher update dates or a fresh page-access date.