DataCenterNews Asia Pacific - Specialist news for cloud & data center decision-makers
Asia
Runware launches modular AI inference pods in weeks

Runware launches modular AI inference pods in weeks

Thu, 6th Aug 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

Runware has launched the Sonic Inference Pod, a modular unit designed specifically for AI inference.

The system packs 1 megawatt of compute into a unit the size of a 20-foot shipping container, using custom servers, printed circuit boards, racks and a closed-loop cooling system. According to Runware, the pods can be deployed in weeks rather than the years often associated with conventional data centre construction.

The launch marks the first time the company has opened its hardware infrastructure to organisations beyond its own platform. It is deploying a network of pods across the United States and Europe, with plans to bring capacity online starting with 160 locations.

Runware is positioning the product against a backdrop of pressure on data centre development, particularly around power and water use. It says the pods can be located wherever power already exists and do not require changes to the grid, with initial siting focused on locations such as solar parks.

Infrastructure model

Runware built the hardware in-house rather than leasing from cloud providers. Each pod is self-contained and joins a distributed inference network managed as a single system, with requests rerouted automatically if a unit goes offline.

Traditional data centres were built for general computing rather than AI workloads, the company argues, and GPUs require far more power and cooling than standard server deployments. In its view, that mismatch has limited the amount of usable GPU capacity in many existing facilities.

Runware says the facilities costs of its approach are up to 100 times lower than those of a traditional gigawatt-scale data centre build. It also says its inference pricing is 30% to 90% lower than that of traditional inference providers, and that operations and facilities costs in its model fall to low single digits from an industry norm of about 35%.

Those claims come as data centre projects face growing scrutiny over resource use. Runware pointed to blocked or delayed projects and debates in several US states over restrictions tied to power and water availability.

Cooling and siting

The pods use a closed-loop cooling system that recirculates 1.5 cubic metres of water rather than relying on evaporative cooling. The company says the design keeps GPU temperatures within 2C of target conditions.

Because the pods can be placed within national borders or specific jurisdictions, Runware is also targeting customers with data residency and sovereign AI requirements. That includes frontier labs, AI studios and companies seeking dedicated inference infrastructure.

The first nodes will support Runware's broader model catalogue, including open-source large language models, fine-tuned models and its own media generation tools for image, video and audio. Customers can deploy their own models and workloads through the company's serverless infrastructure, while Runware manages provisioning, scaling and operations.

Flaviu Radulescu, co-founder and chief executive officer of Runware, outlined the rationale for the design.

"After 20 years running large-scale infrastructure, I knew regular data centers couldn't scale inference - the power and cooling simply aren't there. So we redesigned from first principles: only the components inference needs, liquid-cooled GPUs, and a modular format we can place wherever power already exists. This way we can produce and deploy 1 GW of compute faster than any other infrastructure provider, with built-in redundancy and real-time routing to the most cost-efficient GPU," said Radulescu.

Customer demand

Runware says it has already processed more than 20 billion requests for more than 1 million developers and 500 million end users worldwide. It offers access to more than 400,000 models through a single integration and serves customers in sectors including eCommerce, media, gaming and creative work.

One existing customer, Higgsfield, said it had expanded its relationship with Runware from model access into infrastructure.

"We started with Runware's Model API, but quickly expanded into their inference infrastructure because of the scale and efficiency they could deliver. Our models reach millions of users every week, so reliable capacity and cost-efficient inference are critical. The lower our inference costs, the more value we can pass on to our users. Runware understands that deeply, which is why we work closely with them on our most important model launches," said Ten.