Inside a British engineering distributor’s parts room on an overcast morning, an employee photographed from behind compares a printed maintenance manual with a tablet beside a dismantled industrial pu

When UK businesses need vector search for an AI knowledge assistant

8 min read

UK businesses should test keyword, vector and hybrid retrieval against real staff questions before choosing infrastructure. A separate vector database is justified where measured benefits and operating requirements outweigh integration, support and ongoing costs.

Written by Kate Bennett Group CEO, Compare the Cloud

A UK organisation should add vector search when staff describe problems differently from the documents containing the answers. That does not automatically justify a separate vector database. Existing search platforms can combine keyword and vector retrieval, as Azure AI Search demonstrates. Start by testing your existing search against real questions, then choose a separate database only where measured improvements justify its integration, operating costs and support responsibilities.

Separate the retrieval problem from the database purchase

The useful distinction is between a search capability and the infrastructure providing it.

Keyword search matches words and terms. Vector search uses numerical representations, called embeddings, to find conceptual similarity. Hybrid search combines the two. Enterprise search and vector retrieval therefore overlap: Azure AI Search supports text and vector fields within the same index.

Consider a hypothetical British engineering distributor. A service adviser asks, “What should we check when the pump keeps getting too hot?” The relevant maintenance document might discuss “repeated thermal shutdown”. That is a sensible question for testing conceptual retrieval.

The same adviser might also need the instructions for an exact part code. A result about a similar component could be dangerous or simply waste time. Microsoft specifically identifies product codes and specialised terminology as cases where keyword matching can perform better.

For a small IT team, the first decision should be whether retrieval needs improving. Buying another database comes later.

When keeping existing search is sensible

Keep or improve your current search when testing shows that it reliably finds the correct, current documents for staff questions.

Before adding embeddings, inspect failed searches. Was the document missing? Was its title misleading? Did ingestion omit a table? Was an obsolete version ranked above the approved one? Treat these as content and retrieval defects to investigate, rather than assuming a new database will resolve them.

Retaining an existing service is especially attractive when its connectors, access controls and support arrangements already meet your needs. Verify those capabilities in your own installation; the evidence supplied here does not establish what your current enterprise search includes.

When vector retrieval earns its place

Trial vector retrieval when users routinely paraphrase documents, describe symptoms rather than formal terms, or ask questions whose relevant passages use different vocabulary. That follows the distinction between conceptual matching and exact matching in Microsoft’s retrieval documentation.

Compare keyword, vector and hybrid results against the same questions. Add reranking as a further experiment, not an assumed improvement. Microsoft’s implementation guidance recommends tuning incrementally.

A dedicated vector database becomes a credible choice when that trial establishes a retrieval benefit and your existing platform cannot meet the required integration, capacity, response time or operating requirements at an acceptable cost. These are procurement criteria, not a universal document-count threshold.

Testing whether to add vector search
Compare retrieval on real staff questions, then add a separate vector database only if measured gains justify its costs and operating work.

How the assistant should connect to company knowledge

Retrieval-augmented generation, usually shortened to RAG, gives a language model retrieved material to use when answering. Microsoft describes Azure AI Search as supporting applications built around this pattern.

The following is an illustrative responsibility model, not a claim that any named product supplies every component automatically.

StageProposed responsibilityEvidence required before acceptance
Source documentsBusiness owners approve content and identify authoritative versionsNamed owners and a process for replacing obsolete material
IngestionThe implementation team imports text, source references and access informationChecks that changes and removals reach the index
RetrievalThe search service finds relevant passages using the selected methodsResults tested against real questions and user permissions
Answer generationThe application supplies authorised passages to the modelAnswers cite the correct sources and handle missing evidence
OperationsInternal IT or a contracted provider maintains the serviceMonitoring, escalation, restoration and exit responsibilities

Make permission handling an acceptance condition. Test whether restricted passages are excluded before they reach the answer-generating model, including after a user changes role.

Database access controls are only part of this design. For example, Weaviate documents role-based access control for its deployment, but that alone does not establish that an implementation preserves every permission from your document systems.

For UK procurement, request the locations of document storage, embedding generation, answer generation, logs and backups separately. The supplied evidence does not establish UK-region availability or contractual residency commitments across these options.

Costs beyond the vector index

Use the following as a cost model based on the evidence supplied for 28 September 2026. It is not a set of comparable UK quotations.

OptionDocumented charging basisWhat a UK buyer should establish
PineconeRead units, write units, storage and data leaving the service, alongside applicable plan commitmentsExpected usage, billing currency, tax treatment and any separate model costs
Amazon OpenSearch ServiceManaged cluster instance hours, storage and transfer, or Serverless compute and storageSelected region, deployment configuration, capacity and commitment terms
Azure AI Search semantic rankingUsage-billed premium feature, with a limited free allowanceBase search costs plus the expected ranking and model charges
Self-managed WeaviateInfrastructure requirements depend on the workload and indexHosting, administration, monitoring, recovery, support and maintenance costs

Sources for the charging bases are Pinecone’s cost guide, AWS pricing, Microsoft’s semantic ranking overview and Weaviate resource planning.

Pinecone illustrates why the smallest advertised figure needs context. Its documentation lists Starter at US$0 per month, Builder at US$20 per month, Standard at US$50 per month and Enterprise at US$500 per month. Builder is a flat fee covering included usage, with excess usage blocked. Standard and Enterprise charge for usage above their respective minimum commitments; the minimum is not added again on top of that usage. Pinecone explains the billing treatment.

These dollar figures are not GBP quotations or estimates of the full system. Pinecone’s supplied extract does not establish VAT treatment. AWS states that its published prices exclude applicable taxes, including VAT, unless otherwise noted. AWS pricing terms.

Ask suppliers to price the same operating period and workload. Include document preparation, connector development, embedding generation, index updates, model usage, permission testing, training, support and eventual export. Treat unpriced staff work as an unresolved cost, not a saving.

Self-hosting also needs a restoration budget. Weaviate’s production guide calls for monitoring, tested upgrades and disaster recovery procedures. Its resource guidance explains why memory and CPU requirements depend on the chosen index and workload.

Compare deployment routes against the same questions

The evidence supports several routes, but not a performance league table.

RouteDocumented basisCircumstances worth testingMain procurement question
Retain existing enterprise searchEstablish its actual capabilities through your own configuration and contractCurrent retrieval is accurate and the main gaps concern content or integrationCan it supply authorised, current passages to the assistant?
Azure AI SearchCombined keyword and vector retrieval, with optional semantic rankingA team wants both retrieval methods within one search serviceDoes the additional ranking improve this organisation’s results?
PineconeUsage-based search infrastructure with documented full-text, vector and hybrid cost treatmentA team is evaluating a separately managed retrieval serviceWhat will real reads, writes, storage and transfer cost?
Amazon OpenSearch ServiceManaged cluster and Serverless commercial modelsA team is assessing AWS-operated search infrastructureWhich exact configuration meets retrieval needs, and what capacity must be funded?
Self-managed WeaviateDocumented Kubernetes deployment and resource-management requirementsA team has the skills and a reason to operate the infrastructureWho owns upgrades, incidents, capacity and restoration?

The product evidence comes from Azure hybrid search documentation, Pinecone’s cost documentation, Amazon OpenSearch Service pricing and Weaviate’s production guide. The suitability assessments are editorial judgement.

The AWS extract establishes commercial deployment choices, not the detailed retrieval capabilities of a selected configuration. Confirm those before shortlisting. Equally, do not choose self-management solely to avoid a managed-service bill if nobody can take responsibility when the service fails.

Where a partner builds the assistant, require a handover covering source connectors, retrieval settings, permission mapping, evaluation questions, monitoring and export procedures. Ask whether ongoing support covers the complete application or only the database.

Editorial analysis

CTC’s position is that a vector database should solve a demonstrated retrieval or operating problem.

Build an evaluation set from actual staff questions, with document owners identifying acceptable supporting passages. Include paraphrases, exact identifiers, outdated documents, restricted material and questions the collection cannot answer. Keep a separate set of questions for checking the final configuration after tuning.

Measure retrieval and answer quality separately. First ask whether the right authorised passage was returned. Then assess whether the assistant used it accurately and cited it correctly. Record response time, operating effort and cost alongside quality.

This evidence pack contains supplier documentation, rather than an independent comparison across these products. It supports explaining the mechanisms and commercial choices, but not naming a universal winner. The strongest purchase case is your own reproducible improvement on questions that matter to the business.

Sources

The supplied extracts were used for this article on 28 September 2026. Publisher update dates were not independently established.

Data & Insights

Pinecone paid-plan monthly commitments in US dollars

Builder has a flat monthly fee, while Standard and Enterprise have minimum usage commitments, excluding the wider costs of an AI assistant.

Pinecone paid-plan monthly commitments in US dollarsBuilder has a flat monthly fee, while Standard and Enterprise have minimum usage commitments, excluding the wider costs of an AI assistant.0100200300400500Builder flat feeBuilder flat feeStandard minimumStandard minimumEnterprise minimumEnterprise mini…Builder flat fee, US dollars per month: 20Standard minimum, US dollars per month: 50Enterprise minimum, US dollars per month: 500
View the data
Pinecone paid-plan monthly commitments in US dollars
CategoryUS dollars per month
Builder flat fee20
Standard minimum50
Enterprise minimum500
Source: Pinecone — Understanding Pinecone cost

Frequently Asked Questions

Does every AI knowledge assistant need a vector database?

No. A search service can combine keyword and vector retrieval within its own index, as Azure AI Search documents. Test whether your existing service supplies suitable evidence before introducing another database.

When is keyword search the better starting point?

Start with keyword search where exact identifiers, names, dates or technical terms dominate the workload. Microsoft identifies these as useful keyword-search scenarios. Include them in any vector-search trial so that improvements on conversational questions do not conceal weaker exact matches.

Is hybrid search always more accurate?

Treat that as a testable proposition for your collection. Microsoft recommends adding semantic ranking only where it improves measured relevance in its hybrid query guidance. Use the same questions and expected source passages when comparing configurations.

How much does a vector database cost?

There is no complete assistant price in the supplied evidence. Pinecone lists monthly commitments, while AWS documents infrastructure and usage-based charging. Add implementation, models, support and staff time before comparing total costs.

Can a small UK business self-host the retrieval layer?

Self-hosting is a documented route for Weaviate, but its production guide expects Kubernetes skills and work on security, monitoring, upgrades and recovery. Make a named person or provider accountable for those tasks before choosing that route.

Does semantic search stop the assistant giving unsupported answers?

Do not accept semantic retrieval as proof of answer accuracy. Microsoft distinguishes extracted search answers from generated chat responses. Require separate tests of source selection, answer accuracy, citations and behaviour when evidence is missing.