Scality AI Inference Factory, available now, assembles four components into an integrated, supported stack: validated open-weight models maintained as part of the product; a disaggregated inference-serving layer that separates prefill from decode so each scales independently; a control plane handling authentication, metering, routing and scheduling; and Scality AI Data Infrastructure (ADI), which uses policy-governed operations to manage model state and enterprise datasets across the AI lifecycle.
The ADI component addresses a specific GPU bottleneck. By storing and retrieving key-value cache from high-performance object storage rather than requiring GPUs to recompute context, it extends inference capacity beyond GPU high-bandwidth memory limits, which matters for the agentic and context-heavy workloads now moving into production. Scality validates and ships the integrated stack as a unit, maintaining it as model versions, serving technologies and security requirements change.
The product is aimed at organisations where per-token cloud billing becomes unpredictable as usage scales, where model-version stability matters for applications built on top of them, and where data-sovereignty requirements constrain which jurisdictions can process prompts and documents. Scality describes the stack as giving organisations the ability to run and freeze specific open-weight model versions, something managed cloud services do not offer.
Nataliya Yezhkova, Vice President for Storage and Data Management at IDC, said on-premises inference is a credible option for a growing set of enterprise and public sector workloads, provided the infrastructure can hold and serve model state efficiently at scale, and that Scality is among vendors addressing both the inference and storage layers.
“With AI moving into mission-critical production environments, organisations need greater control over where inference runs, how their models are managed and what happens to their data,” said Jérôme Lecat, CEO at Scality. “AI Inference Factory brings fifteen years of data-infrastructure experience to on-premises AI, giving organisations the reliability and sovereignty they need to run critical AI workloads on their own terms.”
To stay across the latest in cloud, AI and enterprise tech analysis from Compare the Cloud, subscribe to our weekly newsletter at https://www.comparethecloud.net/newsletter