Fin brings evaluation pipelines from AI research to customer support teams

Enterprise customer service teams have spent the past two years asking one question about AI agents: how much can it resolve? Fin's announcement on 13 August moves the conversation to a harder problem — how do you confidently run and continuously improve an agent handling millions of customer conversations?

Evals lets teams build named groups of simulations, each one a multi-turn test conversation modelled on real customer interactions. A team might create an eval for refund-request handling, another for escalation rules, and a third for tone of voice. An LLM-judge scores each simulation pass or fail against criteria the team sets, and the full transcript, event log, and outcome are available for review. The same eval can be re-run after any change to Fin to catch regressions before they reach customers.

Releases gives teams a separate workspace for editing content, procedures, and guidance without touching the live version of Fin. Changes can be bundled together, tested against existing evals, then published to all customers at once, rolled out to a percentage of traffic, or A/B tested against Fin's current configuration. If something goes wrong partway through a rollout, the release can be paused or rolled back in one step.

Once a change is live, Monitors — announced separately in March — checks every conversation against the quality standards defined in evals, flagging conversations that fall short and turning them into material for the next round of testing. Fin calls this loop "eval-driven delivery", a term borrowed from the AI research teams that build the underlying models.

The new features also connect to Operator, Fin's AI agent for customer operations, which can build evals, draft releases, and set up monitors on request. Each step returns a proposal for a human to review and approve before anything goes live.

Fin says some customers running the product at scale see a 1% regression translate to thousands of affected conversations per day — the scale that makes manual testing impractical and automated evaluation necessary.

To stay across the latest in cloud, AI and enterprise tech analysis from Compare the Cloud, subscribe to our weekly newsletter at https://www.comparethecloud.net/newsletter

More News