Dave Friedman, a financial analyst who publishes on Substack, sampled live rental listings on the compute marketplace Vast.ai, and found that the advertised price of a GPU and the price of a usable cluster are different numbers. Friedman built his sample by keeping only machines that were real, bookable immediately, held for at least a week, and run by hosts whose reliability rating cleared 99 percent. The stakes reach beyond one marketplace: banks and exchanges are drawing up the earliest futures contracts for compute, and those will almost certainly settle against a single, standardized GPU-hour price. That design choice creates a specific failure mode. A buyer can watch the contract move exactly as predicted and still be unable to run the workload it was meant to protect.

Friedman’s sample spanned five accelerator models, H200, H100 SXM, B200, L40S, and A100 SXM4, deduplicated down to 37 physical machines holding 93 listed GPUs. For a single H200, the price measure was $3.93 per GPU-hour. Asking for four identical H200s in one machine cut eligible supply in half and raised the price only 4 percent, to $4.08. Requesting eight identical H200s returned zero qualifying machines.

The pattern repeated with different severity across chips. H100 inventory could satisfy two identical GPUs in one box, but nothing qualified once the request grew to four. B200 offered even less room: the whole sample came from two physical machines, neither able to fill an order past two chips, so four- and eight-GPU requests found nothing. A100 behaved the way scarcity textbooks predict, growing pricier as supply shrank: $0.60 per GPU-hour for one chip rose to $1.29 for eight, a premium of 114 percent for staying contiguous. L40S went the other direction. Only 35 percent of the L40S inventory could fit into a box big enough for eight chips, though the single qualifying machine undercut every other price in the sample.

In a normal commodity market, scarcity shows up as a higher price. Compute often skips that step. When no machine can satisfy a four- or eight-GPU request, no transaction occurs to bid the price up, so the trade simply does not happen. A price series built only from single- and dual-chip rentals looks calm and plentiful, even while a company hunting for a genuinely large cluster comes up empty no matter what it offers to pay. Large buyers already route around this by signing bilateral capacity agreements that reserve a specific configuration for a specific window, leaving only the leftover capacity to surface on marketplaces such as Vast.ai.

This is where the emerging compute-futures market runs into trouble. The first contracts will price a single unit, the generic GPU-hour, because derivatives need a simple, liquid benchmark to settle against. Real workloads do not consume that unit. Training runs, fine-tuning passes, and high-volume inference jobs need several matching chips packed into a single box or rack, wired together with interconnect quick enough that they function as one unit. A futures contract written on the wrong unit exposes the buyer to basis risk: when the instrument you used to hedge moves differently than the thing you actually need, a gain on the hedge does not cancel out the real loss. Buy H100 futures to protect a future eight-GPU training run, and the contract will pay out if the broad H100 price rises. It will do nothing if co-located eight-GPU capacity simply is not there. The futures position can close in profit while the compute itself never gets provisioned.

Energy trading already has a term for this leftover exposure: location basis, alongside the gap in price between guaranteed and interruptible power. Compute markets look likely to grow a similar structure keyed to physical layout, with the plain GPU-hour as the baseline and extra charges or standalone contracts layered on for how many chips sit together, how fast they talk to each other, where the datacenter sits, and whether the capacity is actually guaranteed to be there. Those narrower cluster-hour and capacity contracts could well see more volume than the plain-vanilla GPU-hour product underneath them, precisely because they track what buyers are actually trying to buy.

Anyone budgeting or hedging GPU capacity over the next 90 days should treat a generic GPU-hour future as insurance against price alone, not against availability. Pair it with a bilateral capacity agreement or a direct reservation for any workload that needs more than two co-located chips, and price the futures leg as a partial hedge rather than a full one.

Analysis by Dave Friedman, published on Substack, July 23, 2026.