Most AI infrastructure conversations still measure success by equipment count: how many racks, how many GPUs, how many square feet of data hall. That’s the wrong scoreboard.
AI infrastructure should be measured by what it produces, not by how much equipment it contains. The metric that actually matters is tokens per square foot.
Atlas One is built around that better question. This is the second post in our series on the platform; here we’re looking specifically at why token output per square foot is the metric that matters, and how Atlas One is engineered to maximize it.
The Metric That Actually Matters
Square footage is a fixed, expensive constraint. Every AI organization is racing to get more compute into the space (and power allocation) they already have, rather than waiting on new construction to catch up with demand. That reframes the real design goal: not “how many servers fit here,” but “how many tokens can this footprint produce.”
That distinction changes how infrastructure decisions actually get made. A facility that’s 80% full of racks but running well below its power and thermal ceiling isn’t necessarily a productive facility—it’s a facility that’s spending real estate to solve a cooling problem instead of running compute. The organizations getting the most value from their AI infrastructure aren’t the ones with the most hardware. They’re the ones extracting the most output from the hardware they already have.
This is the whole idea behind Atlas One: designed to maximize token throughput per square foot by supporting extreme-density compute without the thermal limitations that cap traditional air-cooled deployments. The result is more productive hardware, more efficient use of space, and more AI output from every deployment—without a bigger building.
Why Token Output and Square Footage Are in Tension
Tokens per square foot only goes up if two things happen at once: more compute has to fit in the same space, and that compute has to actually run at full performance once it’s there. Most infrastructure approaches solve for one and quietly sacrifice the other.
As rack densities rise, conventional cooling systems typically respond by getting bigger, more complex, and more resource-intensive—more CRAC units, more airflow containment, more mechanical square footage dedicated to moving heat instead of running compute. Every square foot spent on cooling infrastructure is a square foot not spent on compute, which directly undercuts the goal of maximizing output from that footprint in the first place. That’s the tension: chase density and the cooling footprint grows to match it, and the “per square foot” gain quietly disappears.
This isn’t a theoretical problem. Industry-wide, rack power draw has been climbing sharply as AI workloads take over—air cooling becomes impractical above roughly 50 kW per rack, and current-generation AI hardware routinely pushes well past that. Organizations still relying on legacy cooling technologies are running into a hard physical wall, not just a design preference, and that wall shows up directly as fewer usable tokens per square foot.
How Atlas One Breaks the Tradeoff
Atlas One is built around immersion and two-phase DLC cooling, which supports extreme-density compute while reducing the space and infrastructure traditionally dedicated to heat removal. More watts go to compute. Less physical footprint goes to cooling. That’s how the platform keeps pushing output per square foot up instead of trading density for a bigger cooling plant.
The efficiency side of that equation matters just as much. Atlas One runs a fully waterless architecture with a targeted PUE of 1.0-1.1—meaning nearly all incoming power goes toward compute rather than cooling overhead. Every watt that doesn’t go to fighting heat is a watt available for producing tokens, which is what actually determines how much output a given footprint can generate.
There’s a second, easy-to-miss piece of the same equation: thermal throttling. When a chip runs hot, it slows itself down to protect itself, regardless of how much it costs. That throttling is often invisible in procurement conversations because the hardware still technically works—it just doesn’t produce as many tokens per hour as it’s capable of. Atlas One’s immersion cooling keeps GPUs and next-generation accelerators at sustained boost clocks and full rated performance, which is the difference between paying for peak token output and actually getting it, in the same physical footprint.
Built for the Hardware Producing Those Tokens
None of this is theoretical—it’s a direct response to where AI hardware is headed. Atlas One is engineered for legacy, current, and next-generation accelerators and CPU workloads, including NVIDIA’s Blackwell, Rubin, Hopper, and RTX Series, and AMD’s MI350X and MI300X.
As each new hardware generation pushes power draw and heat output higher, the infrastructure supporting it has to be built for that trajectory from the start, not retrofitted after the fact—because a platform that can’t keep pace with the next generation’s power and thermal demands starts losing ground on output the moment that hardware ships.
For context on how Atlas One’s density approach compares to more conventional data center cooling strategies, see our overview of advanced cooling solutions.
