Dominic Jainy brings a sophisticated lens to the rapidly evolving landscape of high-performance computing. As the industry moves toward massive AI deployments, the recent five-year strategic agreement between Sharon AI and Rafay Systems serves as a blueprint for modern infrastructure. This conversation explores the shift toward standardized orchestration, the operational hurdles of managing 150,000 GPUs, and the rise of the “neocloud” model that prioritizes seamless, production-ready AI services over mere hardware access. The discussion delves into the transition from managing individual clusters to a unified “AI Factory” model, emphasizing the importance of governance, tenant isolation, and automated service delivery in the race to scale.
Managing up to 150,000 GPUs over five years requires immense operational sophistication; how does this move toward a unified orchestration layer solve the complexities of scaling a neocloud footprint?
Transitioning from managing isolated clusters to a centralized orchestration layer is a massive leap for a company like Sharon AI. By implementing a standardized operating model across their AI Factory environments, they can finally move away from the manual, fragmented processes that often haunt large-scale deployments. This platform doesn’t just manage hardware; it governs how computing infrastructure is provisioned, monitored, and delivered across diverse locations and workloads in Asia-Pacific and beyond. It provides a cohesive control plane that handles both Kubernetes and virtual machine environments through a single platform, ensuring that as the fleet grows toward that 150,000-unit milestone, the operational burden doesn’t become a bottleneck for growth. The physical sensation of managing such a vast network shifts from a chaotic firefighting exercise to a streamlined, automated workflow where every node follows the same set of rules.
The “AI Factory” model seems to be a central pillar of this strategy; how does the integration of lifecycle management and workload orchestration transform raw compute into a reliable service?
The “AI Factory” concept represents a shift in thinking where hardware is no longer the product, but the raw material for a sophisticated service. By utilizing a platform that covers infrastructure lifecycle management, Sharon AI can apply uniform operating standards across various customer environments while keeping tenants strictly isolated from one another. This level of governance is critical because it ensures that high-value assets—those expensive, high-demand GPUs—remain in constant use without sacrificing security or reliability. The goal is to create an automated pipeline where service delivery is accelerated, allowing customers to realize measurable value without the headaches of managing the underlying physical systems themselves. It’s about building internal operating practices that feel consistent whether you are deploying a single model or an entire training suite.
With the sharp rise in demand for training and inference, what are the primary risks for providers who fail to adopt a unified control plane across their global sites?
Providers who rely on disparate tools for every new site risk being crushed by the sheer operational weight of their own hardware fleets. Without a unified control plane, maintaining service reliability and enforcing security policies becomes an exhausting, manual game of catch-up that eats into profit margins and irritates customers who expect 24/7 uptime. There is a palpable tension in the market right now: providers must allocate scarce resources between competing customers while trying to keep utilization rates near 100% to justify the investment. If you cannot automate the provisioning and access of these resources, you face significant delays in deployment, which is a major setback in a market where speed to market is the only currency that matters. Failing to standardize means you are effectively building a series of “island” data centers that cannot share intelligence or efficiency.
How does the specialized nature of a neocloud operator differ from traditional cloud giants, especially when it comes to delivering production-ready AI services?
Neocloud operators like Sharon AI are carving out a niche by focusing entirely on the specialized software stack needed to run high-intensity AI workloads. Unlike general-purpose cloud providers, they are building for the extreme scale and specific performance metrics required for model training and complex inference, often using bare-metal systems and containers in tandem. Their advantage lies in their ability to offer a more tailored, consistent experience through software layers that manage the unique demands of GPU-based infrastructure. This focus on the “operational foundation” means they aren’t just renting out chips; they are providing a turnkey environment where a developer can plug in a model and trust that the governance and monitoring are already enterprise-grade. It creates a sense of trust and security for the end-user, who knows the infrastructure is purpose-built for their specific computational load.
What is your forecast for the AI infrastructure market?
We are entering an era where the distinction between hardware providers and software orchestrators will almost entirely vanish. From 2026 to 2031, I expect to see a massive consolidation of “AI Factories” where the winning players are those who can manage 150,000 or even 500,000 GPUs with the same ease that we currently manage a dozen virtual machines. The focus will shift away from the scarcity of the silicon itself and toward the efficiency of the orchestration layer, making high-performance compute a ubiquitous utility rather than a rare resource. As automation becomes the standard, the operational sophistication we see in this Sharon AI and Rafay deal will become the mandatory baseline for anyone hoping to compete in the global AI economy. Providers will be judged not by how many chips they own, but by how quickly they can transform those chips into secure, production-ready services for their clients.
