Photograph a quiet operations meeting room in a UK office at early morning: an IT lead and finance colleague, seen from behind, compare a printed 36-month cost spreadsheet with a laptop showing a simp

How to calculate the three-year cost of an internal AI assistant

12 min read

UK businesses should compare AI assistants against the same tasks, quality standards and service requirements over three years. This guide explains how to calculate subscription and internal operating costs, measure a pilot and account for staffing, renewals and exit.

Daniel Thomas
Written by Daniel Thomas

Compare the cost of delivering the same useful work over 36 months, including setup, staffing, support and exit. Build separate totals for a per-user service, an internal application using a hosted model, and an application running its own model. Start with a measured pilot and your existing software contracts. Choose the lowest-cost option that meets your quality, access and support requirements, rather than the one with the cheapest licence or inference rate.

Define the work before comparing prices

Write a short acceptance brief before asking for quotations. For a hypothetical UK wholesaler, the requirement might be to answer staff questions from approved product manuals and internal procedures, while respecting each employee’s document permissions.

That is a narrower requirement than an assistant embedded throughout email, spreadsheets and meetings. Do not charge an internal document assistant with replacing those additional functions unless the pilot demonstrates that it can.

Record the following inputs.

InputWhat to establish
UsersEligible staff, expected active users and paid seats required
WorkQuestions, documents and tasks the assistant must handle
QualityAn agreed test set and the conditions for accepting an answer
DemandRequests per working day, document updates and peak simultaneous use
AccessWhich users may retrieve which information
ServiceRequired operating hours, response times and recovery arrangements
OwnershipThe person responsible for content, access, support and spending

Keep eligible users, active users and simultaneous requests separate. Use paid seats to calculate subscription expenditure; use observed demand and peak load to size an internal service.

For a UK budget, record GBP quotations where available. For overseas charges, retain the original currency and have finance supply a dated budgeting exchange rate. Show VAT separately and ask finance which amounts belong in the cost comparison.

Compare three-year assistant costs
Model setup, three years of operation and exit for each delivery route, then choose the lowest-cost option that meets the same requirements.

Map the architecture and its owners

Cost three delivery approaches separately. An internal application calling a hosted model belongs in a different column from one running the model itself.

The proposed responsibility map below is a budgeting template. Confirm the actual allocation in each supplier’s contract.

ResponsibilityPer-user SaaS assistantInternal application with hosted modelInternal application with self-hosted model
User interfaceSupplier product; customer configures adoptionCustomer or implementation providerCustomer or implementation provider
Document accessCustomer checks permissions and connector scopeCustomer builds and maintains access controlsCustomer builds and maintains access controls
Model processingSupplier serviceExternal model API providerCustomer or contracted infrastructure operator
Application supportSupplier support plus internal administrationCustomer or support partnerCustomer or support partner
CapacityCheck subscription limits and extrasManage API limits and application capacitySize, operate and renew model-serving capacity
ExitVerify available exports and account closureTransfer application, data and model connectionsTransfer application, data, infrastructure and operating knowledge

For an assistant answering from company documents, budget for retrieval-augmented generation, or RAG. This means retrieving relevant material and providing it to the model with the question. The supplied research blueprint offers an architectural starting point, but its abstract says practical benefits remain under assessment through ongoing case studies and interviews. It does not establish a production cost or return on investment.

Ask whoever proposes the system to identify every external connection, including model calls, logs, monitoring and backups. A company-controlled chat interface alone does not establish where the whole service processes data.

Build the three-year cost model

Use a workbook with separate columns for setup, year one, year two, year three and exit. Against each entry, record the quantity, unit price, evidence date, owner and whether it is a quotation, measurement or assumption.

Use this calculation for each option.

Three-year total cost = setup + year-one operation + year-two operation + year-three operation + exit − supported residual value.

Only subtract a hardware resale value if you have a defensible basis for it. Keep potential productivity benefits outside this cost total.

Include the work around the software

Cost componentPer-user SaaS inputsInternal assistant inputs
Discovery and pilotTrial licences, evaluation and staff timePrototype, evaluation and engineering time
DeploymentConfiguration, access cleanup and connectionsApplication development, hosting setup and access controls
Data preparationDocument cleanup and connector configurationCleanup, ingestion, indexing and update processes
Recurring technologySeats, prerequisite plans and usage extrasModel API usage or compute, storage, networking and software
PeopleAdministration, training, support and answer checkingThose activities plus maintenance, releases and operational cover
ResilienceRequired support tier and continuity arrangementsBackups, recovery testing, spare capacity and incident response
ChangesNew users, configuration and subscription changesNew integrations, model changes and regression testing
ExitExport, migration, overlap and closure workExport, handover, decommissioning and replacement work

Use fully loaded staff costs supplied by finance. For each activity, multiply recorded or estimated hours by the relevant hourly cost. An employee already on payroll still has limited capacity, so show that time even when no additional hire is proposed.

Keep two views if necessary: cash expenditure and total resource cost. Label them clearly so that an internal build does not appear cheaper simply because its engineering work is hidden in salaries.

Calculate the SaaS side year by year

For each year, calculate:

Subscription cost = committed seats × contractual monthly-equivalent price × months covered.

Then add setup, administration, training, usage extras, support and exit. Separate a monthly-equivalent price from the actual payment schedule.

The supplied primary-source extracts dated 28 September 2026 provide these GBP reference points.

OfferingPublished price in supplied extractCommitment and scope
Microsoft 365 Business Premium with Copilot£24.60 per user per month, excluding VATPaid yearly; annual subscription; productivity and security bundle
Microsoft 365 Copilot for enterprise£23.10 per user per month, excluding VATPaid yearly; annual subscription; qualifying base licence purchased separately

Sources: Microsoft business pricing and enterprise pricing. These are different packages, not equivalent editions ranked by price. Obtain the applicable minimum quantities, offer conditions and renewal terms before committing.

Assumptions for an illustrative calculation

A hypothetical company buys 10 Business Premium with Copilot seats. Assume the quoted £24.60 rate remains unchanged through three annual terms, with no additional usage charges. That flat renewal assumption is not a supplier guarantee.

  • Monthly equivalent: 10 × £24.60 = £246.
  • Annual subscription expenditure: £246 × 12 = £2,952.
  • Three-year subscription expenditure: £2,952 × 3 = £8,856, excluding VAT.

This is a subscription subtotal, not complete total cost of ownership.

If the company would retain a productivity and security suite under either option, calculate the additional expenditure caused by the AI decision. For a bundle replacement, subtract only spending that can actually be cancelled, from the date cancellation becomes possible. Preserve the full contractual cash-flow view alongside that incremental comparison.

Calculate the internal application side

For a hosted model API, measure input and output usage separately. Include retrieved document text, conversation history, repeated attempts and any other billable processing shown in the provider’s usage records.

Monthly model cost = input usage × applicable input rate + output usage × applicable output rate + other billable services.

Normalise units before multiplying. A price per million tokens requires token volume divided by one million.

For self-hosted inference, calculate paid capacity rather than assuming the system costs money only while answering questions. Include the production environment, development and test environments, recovery capacity, storage, networking and the people operating them.

The Fireworks pricing extract, dated 26 September 2026, provides these on-demand reference rates.

GPU optionPublished US dollars per GPU hour
H100 80 GB$8
H200 141 GB$8
B200 180 GB$13
B300 288 GB$15
GB300 288 GB$20

The page says charging is per GPU second and lists a 1.5-times premium for region-restricted deployments. The supplied extract does not establish UK deployment availability, VAT treatment or a minimum purchase. These are infrastructure reference rates, not quotations for a complete internal assistant.

Do not infer cost per useful answer from this table. Measure throughput, response quality and required capacity using the proposed model and your actual documents. No exchange-rate evidence is supplied, so these figures remain in US dollars.

Run the pilot and validate the workbook

Use the pilot to replace assumptions with measurements before approving a wider rollout.

  • [ ] Agree the acceptance test. The business owner signs off representative questions, acceptable answers and failure conditions.
  • [ ] Prepare a controlled document set. Each source has an owner, an update process and approved access permissions.
  • [ ] Test permissions. Users cannot retrieve documents outside their authorised scope.
  • [ ] Record demand and quality together. Capture requests, model usage, successful tasks, failed answers and human review time.
  • [ ] Measure peak load. Confirm the required response time under realistic simultaneous use.
  • [ ] Record operating labour. Include setup, support tickets, document maintenance, troubleshooting and releases.
  • [ ] Check commercial assumptions. Obtain applicable UK terms, currencies, billing commitments, usage limits, support scope and renewal conditions.
  • [ ] Test recovery. Restore the relevant configuration and data, and record the time and people required.
  • [ ] Prepare rollback before switching workflows. Keep the existing document search or business process available until acceptance is signed off.
  • [ ] Train the pilot users. Confirm they know how to check answers and report failures.
  • [ ] Reconcile actual charges. Match usage records to invoices or billing exports and investigate differences.
  • [ ] Set the rollout decision. Finance and the service owner approve the completed model and its remaining uncertainties.

Build low, expected and high scenarios from measured variation. Change adoption, document growth, support hours, usage and renewal assumptions individually so the reader can see which input changes the decision.

If staffing estimates dominate the result, obtain a support quotation or extend measurement of that work. Another round of token-price comparisons will not resolve an unknown maintenance commitment.

Compare suppliers against the same acceptance test

The evidence supports a comparison of purchasing routes, but not a complete current price ranking across all suppliers. Use the following shortlist to obtain comparable offers.

Candidate routeDecision to testCommercial evidence still needed
Google Workspace with GeminiCan the required work be completed within the company’s existing workspace?Applicable edition, included AI functions, limits and current UK renewal quote
OpenAI ChatGPT BusinessCan a shared business assistant meet the brief with acceptable administration and document access?Current UK checkout terms, required connections, limits and support arrangements
Microsoft 365 with CopilotDoes the required work justify a bundle change or an add-on to existing licences?Exact qualifying plan, incremental cost, agent charges and renewal terms
Internal application using FireworksDoes a hosted model meet the quality and demand requirements at a measured usage cost?Selected model rates, deployment location, service terms and application support
Internal application using a self-hosted modelDoes control over deployment justify the infrastructure and operating responsibility?Model licence, tested hardware capacity, maintenance and recovery costs

For existing Google and Microsoft customers, CTC’s Workspace and Copilot comparison provides workflow context. Its older pricing is not used in this calculation.

AI Build Group’s supplied ChatGPT Business pricing page reports UK checkout observations, but the pack does not contain the underlying OpenAI checkout extract. Treat that as a quotation lead rather than an independently checked OpenAI price.

For an internal implementation, ask a UK provider to split its quotation into deployment, recurring support, consumption charges and exit. Require named responsibilities for application faults, model-provider incidents and document-access failures. A single “managed AI” fee is insufficient if those boundaries remain unclear.

Editorial analysis

For a small business, the decisive build cost is often the work that nobody has yet agreed to own. The self-hosting guide explicitly identifies engineering time among the easily underestimated costs of self-hosted operation. Price that work before treating existing servers or an available developer as spare capacity.

Choose SaaS when it passes the acceptance test and its complete cost is lower. Choose an internal application when a required capability or deployment constraint justifies the additional work, and the company can fund a named operator and recovery plan.

A staged outcome is also valid. Retain a small SaaS deployment while testing one internal workflow, then expand only where measured results support the change.

Keep benefits separate from expenditure. Calculate cost per successfully completed task alongside total cost, using a consistent definition of success. Report time released as capacity unless the business can show how it becomes avoided spending, additional output or another measurable benefit.

Sources

The article uses the supplied extracts. Dates below identify dates attached to those extracts; they are not claims of a fresh live-page check.

Data & Insights

Published on-demand GPU rates in US dollars per hour

Fireworks infrastructure rates from the supplied 26 September 2026 extract, excluding any region-restriction premium and without implying equivalent performance or complete assistant cost.

Published on-demand GPU rates in US dollars per hourFireworks infrastructure rates from the supplied 26 September 2026 extract, excluding any region-restriction premium and without implying equivalent performance or complete assistant cost.05101520H100 80 GBH100 80 GBH200 141 GBH200 141 GBB200 180 GBB200 180 GBB300 288 GBB300 288 GBGB300 288 GBGB300 288 GBH100 80 GB, US dollars per GPU hour: 8H200 141 GB, US dollars per GPU hour: 8B200 180 GB, US dollars per GPU hour: 13B300 288 GB, US dollars per GPU hour: 15GB300 288 GB, US dollars per GPU hour: 20
View the data
Published on-demand GPU rates in US dollars per hour
CategoryUS dollars per GPU hour
H100 80 GB8
H200 141 GB8
B200 180 GB13
B300 288 GB15
GB300 288 GB20
Source: Fireworks

Frequently Asked Questions

Is an internal AI assistant the same as a self-hosted model?

No. The application may run on your infrastructure while sending requests to an external model API, as illustrated by Bito’s customer-supplied model-key deployment options. Budget for the application and model processing separately.

Should we include software licences we already own?

Include expenditure that changes because of the decision, and show unavoidable existing costs consistently across the options. If a new bundle replaces an existing subscription, count savings only when the old contract can end. Keep the contractual cash-flow view visible alongside the incremental comparison.

Can we multiply a monthly price by 36?

Only as an explicitly stated flat-price assumption. The Microsoft business price used here is a monthly equivalent paid yearly under an annual subscription, not a guaranteed three-year rate. Model renewal years separately and add the operating costs outside the subscription.

How should we handle US-dollar infrastructure prices?

Keep the original dollar amount and use a dated exchange rate approved by finance for the GBP budget. Show the conversion separately from the supplier’s price and test the effect of exchange-rate changes. The Fireworks rates above have not been converted because the supplied evidence contains no exchange-rate source.

Does a self-hosted model make the whole assistant private?

The surrounding application, connected tools, logs, storage and backups also determine the data boundary. The supplied guide explicitly distinguishes control over model inference from control over the complete system. Verify those connections and responsibilities before accepting a deployment claim.

What is the clearest financial acceptance test?

Compare three-year total cost and cost per successful task for options that pass the same quality, access and service requirements. Stress-test the assumptions most likely to reverse the result, including adoption, staffing, usage and renewal prices. Approve the choice only when someone owns each material cost and operating responsibility.