Runware says the platform delivers AI inference at higher quality and lower cost than other serverless inference platforms and GPU cloud services. According to the company, each pod operates as part of a unified network, allowing requests to be routed to locations with available capacity and closer proximity to users. If one pod becomes unavailable, workloads can be redirected to another within the network.
“We believe distributed compute, positioned closer to end users for faster inference, is what will win in the long term,” co-founder and CEO Flaviu Radulescu said.
The company also highlights deployment speed and infrastructure flexibility as key advantages. Runware says its closed-loop cooling system eliminates water use and allows pods to be built in days rather than the months or years often required for conventional data centers. Radulescu said the system can scale quickly, adapt to new hardware releases, and expand capacity without requiring customers to build new facilities.
“Demand for inference is growing faster than facilities can be built,” Radulescu said. “What we want is to power the world’s intelligence, to be the backbone every AI model runs on with capacity that keeps up with demand instead of throttling it.”
Runware said it currently has 10 pods deployed across the United States, Europe, and Asia-Pacific, with infrastructure available at 160 sites. The company also provides inference services for customers including Higgsfield AI and Wix. It views the expansion into modular data centers as an extension of its broader focus on AI inference infrastructure rather than a standalone hardware offering.
The launch comes as AI companies continue investing heavily in large-scale data centers, but Runware argues its modular approach serves a different role. Instead of relying on centralized facilities, the company is focused on distributed infrastructure that can place compute resources closer to users while providing dedicated hardware for customers that require it.
“Every pod runs as part of a single network, so requests go wherever there’s capacity, closer to the users, and if one pod goes offline, traffic moves to another,” Radulescu said. “Customers who want dedicated hardware get whole pods to themselves.”
Runware also argues that building this type of infrastructure requires specialized engineering expertise. Radulescu said hardware development presents significant technical challenges, noting, “A mistake in a circuit board design costs months between redesign, simulation, fabrication, testing and delivery. Every one of those calls needs someone who understands exactly what each component does and what breaks if it’s gone.”
The company acknowledged that AI infrastructure continues to raise questions about energy consumption and resource use. Runware says its long-term goal is to operate on renewable power while using existing electrical capacity and reducing water consumption through its cooling design.
Radulescu said AI's energy demand will continue to rise regardless of which companies provide the infrastructure. “No transmission losses, no water in cooling, and we’re using power that already exists instead of asking for new grid capacity to be built. More inference built this way means less new grid, less water, for the same amount of compute.”
This analysis is based on reporting from TechCrunch.
Image courtesy of Runware.
This article was generated with AI assistance and reviewed for accuracy and quality.