Can Nvidia DSX Solve the AI Data Center Power Crisis?

Dominic Jainy has spent the better part of two decades at the intersection of high-performance computing and sustainable energy, witnessing the transformation of data centers from simple server rooms into the massive “AI factories” that define our current landscape. As an expert in infrastructure and power management, he has been at the forefront of the shift toward software-defined power systems, helping operators navigate a world where electricity is no longer just a utility but the ultimate constraint on innovation. Today, he joins us to discuss the real-world impact of dynamic power reallocation and the pioneering efforts to turn AI facilities into flexible assets for the global energy grid.

The conversation explores how modern data centers are evolving into highly efficient computing hubs through advanced software platforms that manage power budgets with surgical precision. Key themes include the successful integration of AI facilities into utility demand-response programs, the reclamation of “stranded” power capacity to boost hardware output, and the shift toward holistic facility architectures that utilize high-voltage direct current designs and comprehensive simulation tools to maximize performance per watt.

How does the recent integration of AI facilities into utility demand-response programs change our understanding of data center flexibility during periods of peak grid strain?

For a long time, the industry viewed data centers as static, “always-on” loads that were essentially a burden to the grid during heatwaves or peak hours. What we saw in Santa Clara with Silicon Valley Power completely flips that script, proving that these facilities can act as a giant battery or a flexible resource. In that specific deployment, the facility was able to instantly drop its power draw from four megawatts down to three megawatts without dropping high-priority inference jobs. The team was watching with bated breath, but the automation worked perfectly, responding to over 200 demand signals from the utility. It felt like a SpaceX rocket launch for the engineers involved because it demonstrated that we can balance the needs of the grid with the relentless demands of AI processing in real-time.

With power supply now being the primary bottleneck for infrastructure expansion, how are modern software layers like DSX MaxLPS effectively reclaiming “stranded” capacity?

The reality is that a one-gigawatt factory is never going to magically become a two-gigawatt factory, so we have to get smarter about the energy we already have. We often find “stranded” power—capacity that is provisioned for a worst-case scenario but never actually used—and software like DSX MaxLPS allows us to shift that available headroom to where it is needed most. For example, during testing with Lambda on a cluster of HGX B200 servers, we saw that we could run 19 nodes within the exact same power budget that previously only supported 16 nodes at full tilt. This isn’t just a theoretical gain; it’s about converting idle electrons into active compute, allowing operators to cram more density into the same physical footprint. By monitoring GPU and rack-level usage in real-time, we can ensure that no watt is left sitting on the sidelines while workloads are waiting in the queue.

In an environment where every megawatt is precious, what does the shift in performance per watt look like for operators moving from static provisioning to dynamic resource allocation?

The metrics coming out of recent validations are frankly staggering and show that the old way of fixed power ceilings is becoming obsolete. When Lambda implemented these dynamic controls across their five-rack cluster, they recorded a 24% increase in cluster-wide token throughput, jumping from 4 million tokens per second to 5 million. Seeing those numbers climb while staying within the same utility connection is a game-changer for the economics of these facilities. They also reported a 23% improvement in performance per watt, which means the facility isn’t just doing more work; it’s doing it more efficiently. It creates a sense of relief for operators who were worried they had hit a hard wall, showing them there is still significant performance to be “found” through better orchestration.

Moving beyond just software, how do physical architectural shifts, such as the transition to 800-volt direct current designs, address the complexity of high-density AI infrastructure?

As we push the limits of rack density, the traditional ways of distributing power start to crumble under the weight of conversion losses and heat. By moving toward an 800-volt direct current design, we are significantly reducing the complexity of the power chain, which leads to projected end-to-end efficiency gains of about 3% to 5%. While that might sound like a small percentage, when you are talking about a gigawatt-scale facility, those few points represent a massive amount of energy that can be redirected to the GPUs instead of being wasted as heat. This design also supports the transition to liquid-cooled racks, which are becoming a necessity as we move toward denser, more powerful clusters. It’s a move away from patching isolated parts of the building and toward a streamlined, high-voltage backbone that supports the next generation of AI hardware.

Why is a “whole-facility” approach now considered more critical than simply optimizing individual chips or isolated cooling systems?

In the past, you could get away with just making a faster chip or a better fan, but as AI systems scale, those isolated optimizations are often held back by networking limits or poor rack provisioning. If your networking is slow, your high-speed GPUs sit idle, and if your cooling overhead is too high, you’re wasting electricity that should be powering tokens. This is why we use simulation tools like DSX Sim to model the entire facility before a single brick is laid, ensuring that compute, storage, and building systems all hum in harmony. By treating the entire site as a single, integrated machine, we’ve seen projections where we can enable up to 40% more GPU capacity within the same megawatt budget. It’s about ensuring that every component, from the liquid cooling loops to the 800-volt transformers, is working toward the single goal of maximizing useful work per unit of energy.

What is your forecast for the evolution of AI factory power management as we look toward 2028?

By 2028, I expect the “static” data center to be a relic of the past, replaced by facilities that are fully integrated into the global energy market as active, intelligent participants. We will see the widespread adoption of 800-volt DC architectures as the standard, and software-driven power reallocation will allow us to routinely achieve 40% more compute density than we thought possible just a few years ago. The relationship between utilities and AI operators will become a two-way street, where facilities automatically throttle their non-essential workloads in under a minute to stabilize the grid, receiving significant cost incentives in return. Ultimately, the metric of success will shift entirely from how many chips you have to how many millions of tokens you can generate per megawatt, making energy orchestration the most valuable skill set in the industry.

Explore more

Corporate America Forms Robot Relations to Manage AI Workforces

In a Silicon Valley boardroom, the newest addition to the leadership team isn’t a Harvard MBA—it’s an algorithmic oversight system designed to monitor the emotional and technical output of an entire division. As organizations scale beyond simple automation toward a fully integrated hybrid workforce, the traditional HR manual is being rewritten in real-time. The quiet transition from human-led teams to

Splunk AI Data Management – Review

The sheer volume of digital exhaust generated by modern enterprises has officially outpaced the human ability to manually curate it, turning the promise of big data into a crushing financial and operational burden. As organizations enter 2026, the challenge is no longer just about storing logs but about transforming that massive, chaotic stream of telemetry into something an artificial intelligence

How Can Click2Shell Lead to RCE on WordPress Sites?

A single URL click from a trusted source can silently dismantle the digital fortress of a web server without a single warning appearing on the administrator’s dashboard. While site owners often prioritize defending against massive brute-force attempts or obvious plugin vulnerabilities, this sophisticated exploit chain proves that a standard administrative task can become a direct gateway for a total takeover.

How Is Pure Data Centres Scaling London’s AI Infrastructure?

Introduction The rapid proliferation of artificial intelligence across the global economy has transformed data centers from simple storage hubs into the high-performance engines of modern industry. Pure Data Centres has reached a critical milestone by launching the final major construction phase of its LON01 Brent Cross campus in North London. By developing the B2 facility, the operator addresses the specialized

Why Is Modern Corporate Onboarding Failing New Hires?

Ling-Yi Tsai is a seasoned HRTech expert with decades of experience helping organizations bridge the gap between human potential and digital efficiency. She specializes in talent management integration and understands that the first week of a new job is critical for long-term retention. Today, she shares insights on how companies can move past administrative friction to build genuine employee confidence.