When British businesses should run gpt-oss locally

Local AI needs a business case beyond keeping data nearby. Compare the Cloud argues for a bounded workload, tested output quality and a funded operational owner before British businesses commit to gpt-oss.

7 min read
Read with AI

Open in

ChatGPT Claude Perplexity

This page

Copied to clipboard
In a small British manufacturing firm, an IT lead stands beside a compact workstation in a quiet workshop office, checking a locally run language model on a plain monitor while a colleague feeds it a

British small and mid-sized businesses should run an open model locally when keeping a defined workload inside their own systems justifies taking responsibility for its operation. Our view is that data sensitivity alone is an incomplete business case. OpenAI documents offline use of gpt-oss on personal hardware, but the buying decision should turn on acceptable output quality, the complete data path, support ownership and measured costs.

Start with the workload you need to keep inside

Our editorial position is that local AI earns its place when a business can name both the boundary it needs and the person responsible for maintaining it.

“We do not want business data leaving the company” needs translating into specific requirements. Which documents? Which users? Does the restriction cover prompts, answers, diagnostic logs, backups and remote support access? Is it an internal preference or a commitment made to a customer?

Consider a hypothetical British engineering supplier that wants staff to search confidential production instructions. A sensible local pilot would use an approved document collection, restrict access to authorised staff and return drafts for human checking. Its acceptance test would include whether the application answers accurately and whether any document content leaves the permitted environment.

That is a stronger case than buying a workstation because someone has demonstrated a chatbot on a laptop.

Offline operation is another legitimate requirement. OpenAI’s local deployment guide explicitly covers offline chat. For a business that needs this, test the complete application with its network connection disabled. A locally running model is only one component of an offline workflow.

For ordinary drafting or occasional analysis, I would first assess whether an approved hosted service meets the task and contractual requirements. Move locally when the pilot demonstrates a benefit worth owning, rather than treating the location of the model as the whole decision.

Deciding Whether to Run gpt-oss Locally
A business pilots a defined workload, checks quality and data flows, and weighs operational ownership before choosing where to run it.

Own the complete data path

A proposed local document assistant might connect a staff interface to an access-controlled document store and a model server. Its answers would return to the user for review. External search, monitoring and support connections would require separate approval.

The table below is a proposed allocation of responsibilities, not a description of controls automatically supplied by gpt-oss.

| Component | Responsibility to assign | Evidence to require before rollout |

| --- | --- | --- |

| Staff interface | IT owner controls sign-in and permitted users | An unauthorised user cannot reach the assistant |

| Document retrieval | Information owner defines accessible material | Users cannot retrieve documents outside their permissions |

| Model server | Technical owner maintains the runtime and hardware | The agreed workload completes at acceptable quality and speed |

| Logs and backups | Operations owner sets access, retention and recovery | Stored prompts and answers follow the agreed handling rules |

| External tools | Application owner approves each destination and permission | A data-flow review identifies what each connection sends |

| Support and recovery | Named employee or contracted provider handles failures | A recovery exercise succeeds and an escalation route exists |

OpenAI documents local serving through Transformers, but a working endpoint should be treated as the start of an application build. Require the surrounding access controls, monitoring and recovery arrangements in the delivery scope.

If an IT provider will operate the system, ask who patches it, who can inspect prompts, what remote access they retain and what happens when the contract ends. Renting a server elsewhere should be assessed as an externally hosted arrangement, even when your company installs the model.

Price the service you will operate

OpenAI’s licensing and cost guidance makes the central distinction clear. The weights are free to download; running them incurs costs.

The supplied evidence does not establish comparable UK hardware quotations or complete hosted-service tariffs. It therefore cannot support a defensible GBP break-even figure.

For a British buyer, request GBP quotations with VAT treatment, support hours, hardware warranty and replacement arrangements stated explicitly. Compare both routes over the same ownership period and against the same workload.

| Cost component | Local deployment budget | Hosted-service budget |

| --- | --- | --- |

| Processing | Hardware purchase or rental, power and capacity headroom | Subscription or usage charges under the selected service terms |

| Application | Interface, document retrieval, identity and integration work | Configuration, integration and any additional application charges |

| Operation | Patching, monitoring, backups and fault handling | Internal administration and any contracted support |

| Quality control | Evaluation, staff review and retesting after changes | Evaluation, staff review and retesting after changes |

| Exit | Data export, replacement deployment and equipment handling | Data export, migration and replacement integration |

These are quotation requirements, not claims that every supplier charges separately for each item. Remove duplication where a contract bundles work.

Treat memory figures as planning inputs

OpenAI’s Ollama guide gives these recommended starting points for graphics memory or unified memory.

| Model | Recommended memory starting point in the guide |

| --- | --- |

| gpt-oss-20b | At least 16 GB |

| gpt-oss-120b | At least 60 GB |

The same guide says moving work to the central processor when graphics memory is insufficient will be slower. Its figures should not become a purchase specification without a representative test.

The TensorRT-LLM guide, for example, requires at least 20 GB of graphics memory for its smaller-model setup. OpenAI’s Transformers guide also distinguishes memory requirements by numerical format. Specify the model, runtime and configuration together.

Ask a prospective installer to demonstrate your longest documents and expected simultaneous use before approving hardware. Record waiting time, failed requests and reviewer effort alongside technical throughput.

Choose the arrangement your team can sustain

The following recommendations are editorial judgements, conditional on testing and acceptable service terms.

| Approach | When I would favour it | What would stop me |

| --- | --- | --- |

| Local workstation | A bounded individual workflow with an offline or data-boundary requirement | Colleagues need dependable shared access that the workstation cannot provide |

| Internally operated server | A shared workload with a funded technical owner and a demonstrated reason to retain processing internally | No credible maintenance, recovery or capacity plan |

| Hosted business application or API | The selected service meets the task, data-handling and support requirements | Decisive contractual or technical requirements remain unanswered |

| Split local and hosted workflows | Different information classes justify different destinations | Staff must guess where each document belongs |

Do not assume that a hosted provider trains on your business data, retains nothing or processes everything in Britain. The evidence supplied here does not establish those terms for a particular hosted service. Make them procurement questions and require the answers for the exact product and contract under consideration.

Equally, do not retire an existing approved service merely because local inference works. Compare the completed task, including corrections and administration, before funding a migration.

Editorial analysis

The best argument for local AI is a specific capability or control that the business needs enough to maintain.

For a small company, I would start with a bounded gpt-oss-20b pilot where existing suitable hardware is available. OpenAI provides a consumer-hardware deployment route, which makes that a practical evaluation option. It does not establish that the model will answer the company’s questions well enough.

Use representative tasks, agree acceptance criteria before testing and include awkward cases. Keep records of incorrect answers, rejected requests, waiting time and staff corrections. Nominate someone other than the installer to judge whether the results are useful.

Approve production only when the business can explain why local operation matters, demonstrate acceptable results and fund the named operational owner. If those conditions remain unresolved, continue with the existing approved workflow while investigating. A successful demonstration is insufficient evidence for a service commitment.

FAQ

Does running gpt-oss locally mean business data never leaves the machine?

Only if the complete application is configured and checked to keep it there. OpenAI’s Ollama guide covers offline chat but also describes browser and function tools. Examine those connections, logs, backups and support access before making a data-locality claim.

Is 16 GB enough for a British company to use gpt-oss?

OpenAI’s Ollama guide recommends at least 16 GB of graphics or unified memory for gpt-oss-20b. That recommendation does not establish acceptable performance for your documents or simultaneous users. Test the intended runtime and workload before buying hardware.

Will local gpt-oss be cheaper than a hosted AI service?

The evidence does not establish a universal saving. OpenAI says the weights are free but operating costs remain. Compare infrastructure, staff time, support and exit work with the selected hosted service’s complete charges.

What should a small business test first?

Choose one useful text task with an agreed document collection and a person who can judge the answers. The gpt-oss-20b specification identifies it as a text-input and text-output model. Test accuracy, access restrictions, data flows and recovery before extending the pilot to more staff.

Sources

Source extracts supplied for this article on 28 September 2026.

Ollama memory starting points for gpt-oss. Source: OpenAI, How to run gpt-oss locally with Ollama
OpenAI's Ollama guide recommends at least these amounts of graphics or unified memory, without guaranteeing performance for a particular workload.