Princeps is the independent risk and grading layer for compute — call it a credit agency, or even the gold standard, for compute. We price the risk nascent compute markets have so far overlooked, modelling whether a given operator will actually deliver the capacity and uptime it sells. Fundamentally, our thesis is that any compute contract is only as good as its worst deliverable unit.
Since last year, compute markets have boomed. Compute now has a price, venues, and a path to physical settlement. What it does not yet have is a deliverable standard (a measure of whose GPU-hours are good delivery and whose are not). Today a top-tier operator and an unreliable one look identical on paper and are procured against the same undifferentiated benchmark. This is more a trust problem than a pricing issue. Historically, every commodity that became tradable first grew an independent layer certifying that what was delivered matched what was promised (gold has the LBMA Good Delivery List, grain has licensed graders and certified warehouses, oil has independent inspection at the delivery point). Liquidity will always follow trust.
We're confident the market is moving in this direction already. Compute is developing observable reference prices through indexes and now futures listings. But a financial contract, as one recent analysis of the compute-offtake market puts it, "cannot make a delayed data centre open on time, secure additional power, repair a network failure, or turn one generation of GPU into another."17 While a price curve prices the market, it still leaves untouched the basis between a reference price and the compute a specific operator actually delivers (the configuration, delivery, service-level and counterparty risk) that is still bundled and unpriced inside private capacity agreements today. Whoever takes on delivery of that compute, whether the platform reselling it, the lab training on it, or the buyer reserving it, inherits that risk, not the operator. Princeps prices it. As financing and procurement grow more modular, lenders already separate market-price, utilisation, basis, operating and counterparty risk. Independent of any one index or marketplace, Princeps provides the independent grade for compute based on reliability, capacity and delivery.
Founding Team
Backed by Y Combinator and angels and advisors from Standard Intelligence, SF Compute, Kojo, Palantir and 8VC.
Prime Intellect has built price discovery for compute across 50+ data centres. It has not yet built delivery discovery, providing an independent benchmark for whether a listed provider will actually hold through the run. Princeps proposes an independent reliability grade in partnership with Prime Intellect.
Why this matters for Prime Intellect? Prime Intellect is an aggregated venue: you pool supply from operators of very different quality (API access currently suggests Crusoe, DataCrunch, Lambda, Massed Compute, Nebius and Vultr) and present them to a buyer sorted by price. Delivery risk is high and asymmetrically skewed against the buyer. The buyer has no way to tell a reliable H100 from an oversubscribed one, so the marketplace adversely selects for whoever quotes lowest, and your best operators get undercut by your worst. On a cash-settled index this risk remains hidden but on a marketplace that actually provisions capacity it lands on the buyer and, eventually, on the venue's reputation.
What we are proposing? One integration that gives your marketplace the missing axis. We'll provide an independent Delivery Grade on every listed and reserved offer, a reliability-adjusted sort, and a per-contract risk score that your buyers can act on. Over time this becomes an admission standard that decides who lists, at what tier and at what premium. We quantify risk based on the Princeps methodology and data ingestion directly from Prime Intellect's own transaction volume.
How do we grade? The grade composes two factor groups: the site’s physical resilience (from Princeps data) and its observed delivery record (from the order flow your exchange already generates) into a single score.
A buyer opens On-Demand GPUs on Prime Intellect and picks an H100 80GB. The marketplace aggregates supply from the operators clearing on it — today, by API, that set is Crusoe, DataCrunch (Verda), Lambda, Massed Compute, Nebius and Vultr1 — and returns those options sorted by price, giving the buyer price, location, socket and a binary isolation flag, but no reliability signal. Prime Intellect itself acknowledges as much: "We don't provide formal SLAs at this time… verify uptime and past performance before deploying long-term workloads."6
| Powered by | Socket | Security | Location | Spin-up | $/hr | |
|---|---|---|---|---|---|---|
| Verda (DataCrunch) | SXM5 | SECURE_CLOUD | Finland | ~3 min | $3.25 | Deploy |
| Lambda Cloud | SXM5 | SECURE_CLOUD | United States | ~5 min | $4.29 | Deploy |
Price is a real signal for cost, not for delivery. These two variables are not reliably correlated, and the buyer cannot tell the reliable H100 from the unreliable one.
What separates a good operator from a bad one is not whether hardware fails (which it inevitably will) but how well that run is contained. In one study, Meta held >90% effective training time because only 3 of 419 incidents needed manual intervention.7 A weaker operator can run up to 10.7% lost compute. The failure waterfall scales with cluster size, meaning that the larger the reserve, the more a run depends on the operator's containment and rerouting capabilities.
Princeps reconstructs the loss history for neoclouds and compute providers from zero to one. Our fundamental belief is that risk must be assessed at origination — across both the physical and performance layers — to be meaningfully quantified. This is the single most mispriced risk by compute marketplaces right now. Uptime Institute tiers measure redundancy, not AI-readiness (built for 5–10 kW racks when a GB200 rack exceeds 120 kW) and certification is point-in-time and gameable.10 SLAs cap the remedy at a service credit, usually a 10% credit on a ~$3 instance is about 30 cents, against outages averaging near $1 million.11 The most serious independent effort is ClusterMAX. SemiAnalysis is both the strongest proof that reliability varies enormously and that an independent grade has real traction.
SemiAnalysis's ClusterMAX has been the de-facto GPU-cloud reliability benchmark but only CoreWeave reached their Platinum ranking.12 SemiAnalysis's own verdict: "the bar across the GPU cloud industry is currently very low." The gap is not the silicon. On identical hardware, ClusterMAX measured the difference between a poorly-tuned and a well-tuned operator as roughly 60% vs 98% usable network efficiency. It is the operator, not the GPU, decides whether the compute is usable.
Prime Intellect's own beta cluster landed at Bronze, flagged for "a lack of health checks or dcgmi integration, and no monitoring dashboard". Princeps extends the same evidence base into a continuous, per-contract risk quantification your exchange can act on in real time as a live number (not a quarterly review).
So even a fast, competent operator has no independent, buyer-visible, real-time reliability number to stand on. Princeps fills this intelligence for both the seller and the marketplace provider.
Reliability can be inferred from the physical and structural signals around a provider, and confirmed by whether the provider actually delivered against real orders. Princeps grades ~940 operators into five bands. Critically, none of this requires probes inside an operator's stack or a device install on your marketplace, instead the grade we deliver is built from external data and the delivery record your exchange already generates.
| Powered by | Princeps Grade | Tier | Delivery confidence |
|---|---|---|---|
| Nebius | Verified | Gold | |
| Crusoe | Verified | Gold | |
| Lambda | Reliable | Silver | |
| Vultr | Reliable | Tier III | |
| DataCrunch (Verda) | Watch | Bronze | |
| Massed Compute | Elevated | Underperformer |
The grade is a reliability and performance benchmark, calibrated to the measured failure and goodput evidence above. Pick a GPU, a cluster size, a term and a checkpoint interval. Every row and chart recomputes on that exact contract, using a Monte-Carlo failure simulation. Prices are live from Prime Intellect's marketplace.
| Grade | Reliability 30-day delivery |
Performance usable goodput |
Risk · expected wasted compute (P50) |
Risk · tail bad-luck run (P99) |
Failure exposure | Delivery confidence |
|---|
The risk figures are derived from a Monte-Carlo simulation of the contract you set above. Each of 4,000 trials draws on both hardware failures9 and a severity per failure (how much work is lost between checkpoints and how long recovery takes, scaled by the operator's containment). The result is a full distribution of delivered goodput. The gap between the expected outcome and the unlucky tail (P99) is the risk a buyer carries silently today and exactly what a delivery grade lets your exchange surface, quantify and rank on.
4,000 simulated runs · Nebius/Crusoe (Gold) vs Massed Compute (Underperformer) — two real operators on your marketplace
P50 (expected) and P99 (bad-luck) wasted compute, $ — grade sourced from each provider's ClusterMAX tier
Model: failures ~ Poisson(N·24·T / 50,000); per-failure lost cluster-time ~ checkpoint-interval rework + recovery, exponential severity, containment scaled so each grade's mean matches its measured goodput-loss band [SemiAnalysis/Nebius]. Illustrative — for structure, not a quote. Synchronous training assumed (a failure stalls the whole cluster to the last checkpoint).
We joined the six providers clearing on your marketplace to the Princeps database and pulled each operator's independent tier, ownership model, chip generation and site footprint.
| Provider | Princeps grade | ClusterMAX tier (according to SemiAnalysis) | Usable goodput | Wasted · expected (P50) | Wasted · tail (P99) | Site evidence |
|---|
| Powered by | Security | Location | Princeps Grade | $/hr | |
|---|---|---|---|---|---|
| Lambda Cloud | SECURE | US | Verified | $4.29 | View grade |
| Verda (DataCrunch) | SECURE | Finland | Reliable | $3.25 | View grade |
That a grade transforms a marketplace is one of the better-established results in economics. Akerlof's 1970 "market for lemons" showed that when buyers can't observe quality they pay only for average quality, good sellers exit, and the market can unravel. The fix is the creation of "counteracting institutions" such as certification and third-party grading.4
| Market | Credibility layer | Measured effect |
|---|---|---|
| Grain, gold | USDA grades / LBMA Good Delivery | Made the asset exchange-tradeable5 |
| Corporate debt | Credit ratings (NRSRO) | Unlocked mandated institutional capital14 |
| Enterprise SaaS | SOC 2 attestation | De facto gate to enterprise procurement15 |
The mechanism is identical: an independent layer converts an unobservable quality attribute into a verifiable, priceable signal — raising prices for good sellers, raising transaction probability, widening participation, and shrinking the pooling that drives adverse selection.
Clearing prices rise. A grade is most valuable exactly where trust is scarcest and sellers are otherwise indistinguishable, as in an aggregated GPU exchange. Volume shifts to quality, and the buyer pool widens: enterprises default to hyperscalers on trust, not price16. An independent, quantified reliability grade encourages marketplace activity from buyers currently hedging on reliability.
provider, security, dataCenter and country per offer.Underneath the grade sits a continuous intelligence layer that scores every operator across six dimensions:
Princeps would provide ongoing risk intelligence and quality scoring for Prime Intellect — a live credit-scoring index for the operators clearing on your marketplace. It can be offered as a standalone paid intelligence service, or integrated directly into the Prime Intellect platform: the grade column, admission gating, and Reserved-cluster qualification all drawing on the same index.
The integration above is the starting point and we are keen to explore a larger, longer-term collaboration opportunity alongside it. Prime Intellect generates data most of the market cannot see — a continuous record of which operators actually delivered, drawn from every booking, fulfilment and failure across the exchange. Combined with the Princeps methodology, that record is the foundation for a jointly-operated Compute Credit Rating. Over time it would position Prime Intellect as a reference point for compute reliability, much as the Chicago Board of Trade's grading standards became the basis on which the grain market settled.5
The path from grade column to rating standard runs in three phases:
Princeps operates the rating independently using Prime Intellect as the data spine and reference venue. A rating the seller controls is not an independent rating. Independence is what makes standards credible to both buyers and lenders, and ensures marketplace liquidity in the long term. The rating improves the more volume clears through your exchange.
1. Prime Intellect availability API, pulled 20 Jul 2026.
2. SemiAnalysis (commissioned by Nebius), goodput-loss benchmarking, 2026.
3. Prime Intellect, "Head of Compute" JD, 2026.
4. Akerlof, "The Market for 'Lemons'," QJE 84(3), 1970.
5. USDA AMS; LBMA Good Delivery; CME Group (CBOT futures, 1865).
6. Prime Intellect docs, FAQ, 2026.
7. Meta, "The Llama 3 Herd of Models," arXiv:2407.21783, 2024.
8. Meta, "Revisiting Reliability in Large-Scale ML Research Clusters," arXiv:2410.21680, 2024.
9. Epoch AI, GPU-cluster failure-scaling model (~1/50,000 GPU-hrs), 2024.
10. American Compute, "Data Center Tiers," 2026; Uptime Institute.
11. Uptime Institute, "Annual Outage Analysis 2024"; "Cloud SLAs punish, not compensate," 2022.
12. SemiAnalysis, "ClusterMAX 2.0," 2025.
13. LBNL, "Queued Up: 2025 Edition," 2025.
14. White, "The Credit Rating Agencies," JEP 24(2), 2010.
15. ControlCase (industry reporting), 2024–25.
16. McKinsey, "The evolution of neoclouds," 2025–26.
17. Dave Friedman, "The Compute Market has Multiple Views on Future Compute Prices," 2026.
For discussion only; not an offer of insurance or a quotation. Marketplace prices are pulled live and change in real time. The failure simulation is calibrated to measured goodput-loss and failure-rate data and is illustrative of structure, not a quote. Grades are a framework, not a published assessment of any provider.