Enterprises Are Buying AI Compute for Speed and Flying Blind on What It Costs

Enterprises Are Buying AI Compute for Speed and Flying Blind on What It Costs

A 170-enterprise survey finds production AI workloads everywhere and cost accounting almost nowhere — and the next round of spending is aimed at infrastructure barely anyone actually uses.

Written by OutOfToken AI

August 12, 2026 · 6 min read · Synthesized from reporting by VentureBeat · How this works

AI Verified · 9/10

Enterprise AI has quietly crossed a line: two-thirds of organizations now run AI workloads in production, and nearly a third do so at scale. But new research from VentureBeat's Pulse Research series, surveying 170 enterprises in July, finds that the ability to account for what that compute actually costs hasn't kept pace with the rush to deploy it.

Cost Just Got Demoted

The clearest signal in the data is a reordering of priorities. Integration with existing cloud and data stacks remains the top factor in choosing an AI infrastructure provider at 40%, but performance — latency and throughput — has climbed to second at 35%, with GPU availability third at 24%. Total cost of ownership now ranks fourth at 22%, and cost per million tokens sits dead last at 16%.

Measuring Uptime, Not Dollars

The same pattern shows up in how enterprises define success. Uptime and reliability is the top metric for 51% of respondents, and developer productivity follows at 39% — both comfortably ahead of cost per million tokens at 31%. For teams running live systems under production pressure, prioritizing reliability over price is a rational response. The problem is what happens underneath that decision.

"Fewer than half of enterprises — 47% — rigorously track the cost and return of their AI compute. Even among those running AI in production at scale, that figure only reaches 56%."

Half-Empty GPUs

Among the 155 surveyed enterprises that operate their own GPUs, 69% report utilization at half capacity or less, and 12% don't measure utilization at all. Scale doesn't appear to fix this: enterprises running AI in production at scale clear the 50% utilization mark at roughly the same rate — 22% — as everyone else. Idle accelerators are expensive accelerators, and a meaningful share of this cohort can't even see how idle theirs are.

Chasing Clouds They Don't Use

The next dollar is aimed somewhere enterprises barely operate today. AI-specialized clouds are the top planned evaluation category at 44% and carry the strongest net momentum of any infrastructure approach, yet CoreWeave and Lambda each register only 3.5% of current usage, with the rest of the neocloud field below 3%. Near-term switching consideration for those same specialized providers sits at just 4%, versus roughly 30% for incumbents like OpenAI, Google Cloud, and Azure — suggesting the 12-month evaluation interest is a longer-horizon thesis, not near-term pipeline.

The Satisfaction Gap

Overall satisfaction with current infrastructure averages 4.14 on a five-point scale, and ease of implementation 4.04. Value for money trails at 3.87 — the softest of the three, landing precisely on the dimension enterprises are least equipped to measure. It's a coherent, if uncomfortable, pattern: the criterion enterprises can't quantify is also the one they rate worst.

A Bottleneck Behind the Bottleneck

The report also surfaces an early warning about what comes after GPU scarcity: memory. Asked how they'd address the shift from compute to memory constraints in large-scale inference — specifically KV-cache capacity — Dell leads at 24% and Nvidia at 21%, but no approach commands anything close to a majority. Roughly one in five enterprises either don't recognize the constraint or haven't begun addressing it, consistent with a cohort still struggling to measure the cost problem directly in front of it.

None of this suggests enterprises are making bad decisions — buying for reliability and speed is a defensible response to production pressure. But it does mean most are optimizing a system they can't fully price, and preparing to re-platform toward infrastructure they haven't yet tested at scale. Whether instrumentation catches up before that shift arrives, or whether enterprises simply repeat the pattern with the next layer of compute, is the open question the report leaves unanswered.

Editorial Note

The article faithfully represents the VentureBeat Pulse Research data provided, accurately citing specific percentages, findings, and the methodological limitations of the survey (n=170, single July 2026 wave, self-selected sample). The research confirms the core narrative: enterprises have deployed AI at scale but lack cost visibility, GPU utilization remains low despite planned re-platforming, and cost has been demoted as a buying criterion. The live web sources provided do not directly address the specific enterprise survey data but do corroborate the broader context about rising compute costs and efficiency concerns.

Claim Tracker

AI-assessed

VerifiedTwo-thirds of enterprises (66%) now run AI workloads in production, and 29% describe AI in production at scale

The article is based on VentureBeat Pulse Research (n=170, July 2026). This figure is stated consistently in the methodology and Finding 1 of the research document provided.

VerifiedFewer than half of enterprises (47%) rigorously track the cost and return of their AI compute

Finding 7 of the VentureBeat Pulse Research explicitly states: 'Fewer than half of enterprises (47%) rigorously track the cost and return of their AI compute.' This directly corroborates the article's claim.

VerifiedAmong the 155 enterprises that operate their own GPUs, 69% report utilization at 50% or less

Finding 6 confirms: 'Roughly seven in ten GPU-operating enterprises (69%) report utilization at or below half capacity.' This matches the article's statement exactly.

VerifiedTotal cost of ownership has fallen to fourth among selection criteria at 22%, behind integration (40%), performance (35%), and GPU availability (24%)

Finding 5 of the research states these exact figures and ordering: 'Integration with the existing stack remains the top selection factor at 40%, but performance sits second at 35% and GPU access and availability third at 24%, both ahead of total cost of ownership at 22%.'

VerifiedAI-specialized clouds are the top planned evaluation category at 44% yet CoreWeave and Lambda each register at 3.5% of current usage

Finding 3 confirms the 44% evaluation figure and Finding 2 states: 'CoreWeave and Lambda each appear in 3.5% of stacks.' This intent-to-action gap is a central tension of the research.

Ask AI about this story

// discussion

sign in to join the discussion