Runware, an AI infrastructure company, announced the launch of its modular data center called Sonic Inference Pod. The Pod is designed as a single transportable unit, providing a more flexible compute solution that can sit alongside hyperscalers’ massive data center projects. Runware claims the Pod can deliver inference at higher quality while being more cost-effective than other serverless inference platforms and GPU clouds. The modular design allows for quick capacity expansion by creating new pods instead of expanding a fixed data center. Runware currently has 10 pods in deployment across the U.S., Europe, and Asia-Pacific, with 160 sites available to power its pods. The company already provides inference to a few companies, including Higgsfield AI and Wix. "We believe distributed compute, positioned closer to end users for faster inference, is what will win in the long term," said Flaviu Radulescu, co-founder and CEO of Runware. "Demand for inference is growing faster than facilities can be built," he added. "What we want is to power the world’s intelligence, to be the backbone every AI model runs on with capacity that keeps up with demand instead of throttling it."

The Sonic Inference Pod uses a closed-loop cooling system that can be built in days, compared to the months or even years it takes to build traditional data centers. This system does not use water, making it more efficient and environmentally friendly. Radulescu noted that the pods can scale and add capacity quickly, deploy anywhere there is power, and adapt rapidly to new hardware releases. "Every pod runs as part of a single network, so requests go wherever there’s capacity, closer to the users, and if one pod goes offline, traffic moves to another," he said. "Customers who want dedicated hardware get whole pods to themselves." He also emphasized that hardware development is slow and the talent pool to build and fix this technology is limited. "A mistake in a circuit board design costs months between redesign, simulation, fabrication, testing and delivery," he said. "Every one of those calls needs someone who understands exactly what each component does and what breaks if it’s gone."

Radulescu acknowledged the controversy surrounding AI data centers due to their resource usage, noting that communities near data centers have reported increased utility costs. He mentioned that Runware aims to eventually run on renewable power without drawing on community resources, though this is not yet a reality. "AI power use is going to increase regardless, driven by demand for inference, not by who supplies it," he said. "What Runware is focused on right now is how that demand gets met. No transmission losses, no water in cooling, and we’re using power that already exists instead of asking for new grid capacity to be built. More inference built this way means less new grid, less water, for the same amount of compute."