Princeps proposal forPrime Intellect
princeps.dev

Princeps<>Prime Intellect

Princeps is the independent risk and grading layer for compute — call it a credit agency, or even the gold standard, for compute. We price the risk nascent compute markets have so far overlooked, modelling whether a given operator will actually deliver the capacity and uptime it sells. Fundamentally, our thesis is that any compute contract is only as good as its worst deliverable unit.

Since last year, compute markets have boomed. Compute now has a price, venues, and a path to physical settlement. What it does not yet have is a deliverable standard (a measure of whose GPU-hours are good delivery and whose are not). Today a top-tier operator and an unreliable one look identical on paper and are procured against the same undifferentiated benchmark. This is more a trust problem than a pricing issue. Historically, every commodity that became tradable first grew an independent layer certifying that what was delivered matched what was promised (gold has the LBMA Good Delivery List, grain has licensed graders and certified warehouses, oil has independent inspection at the delivery point). Liquidity will always follow trust.

We're confident the market is moving in this direction already. Compute is developing observable reference prices through indexes and now futures listings. But a financial contract, as one recent analysis of the compute-offtake market puts it, "cannot make a delayed data centre open on time, secure additional power, repair a network failure, or turn one generation of GPU into another."17 While a price curve prices the market, it still leaves untouched the basis between a reference price and the compute a specific operator actually delivers (the configuration, delivery, service-level and counterparty risk) that is still bundled and unpriced inside private capacity agreements today. Whoever takes on delivery of that compute, whether the platform reselling it, the lab training on it, or the buyer reserving it, inherits that risk, not the operator. Princeps prices it. As financing and procurement grow more modular, lenders already separate market-price, utilisation, basis, operating and counterparty risk. Independent of any one index or marketplace, Princeps provides the independent grade for compute based on reliability, capacity and delivery.

Princeps offers
  • NeoCloud Credit Score — grading of ~940 neocloud operators, from satellite imagery of on-site generation, permits, chip inventory and specs, grid-interconnection age, probabilistic weather and external order data. None of them file anything, so this is the only underwriting file that exists on them. Sold today to compute marketplaces and AI labs allocating training spend, who otherwise default to hyperscalers on brand alone.
  • Verified tier — the top grade, corroborated by an observed delivery record: spin-up success, advertised-vs-actual availability, and completion against real orders. No probes inside an operator's stack, no device install.
  • Risk engine — a continuous, per-contract quantification of expected and tail delivery loss, and the foundation of an independent compute credit rating. Initially developed through Princeps data partnership with Ornn.

Founding Team

Anya Trofimova
Anya Trofimova
Oxon.
Conrad Frøyland Moe
Conrad Frøyland Moe
Oxon.
Wilhem Hector
Wilhem Hector
MIT Mech Eng. & Oxon.

Backed by Y Combinator and angels and advisors from Standard Intelligence, SF Compute, Kojo, Palantir and 8VC.

· Scope of proposed collaboration with Prime Intellect

Prime Intellect has built price discovery for compute across 50+ data centres. It has not yet built delivery discovery, providing an independent benchmark for whether a listed provider will actually hold through the run. Princeps proposes an independent reliability grade in partnership with Prime Intellect.

Why this matters for Prime Intellect? Prime Intellect is an aggregated venue: you pool supply from operators of very different quality (API access currently suggests Crusoe, DataCrunch, Lambda, Massed Compute, Nebius and Vultr) and present them to a buyer sorted by price. Delivery risk is high and asymmetrically skewed against the buyer. The buyer has no way to tell a reliable H100 from an oversubscribed one, so the marketplace adversely selects for whoever quotes lowest, and your best operators get undercut by your worst. On a cash-settled index this risk remains hidden but on a marketplace that actually provisions capacity it lands on the buyer and, eventually, on the venue's reputation.

What we are proposing? One integration that gives your marketplace the missing axis. We'll provide an independent Delivery Grade on every listed and reserved offer, a reliability-adjusted sort, and a per-contract risk score that your buyers can act on. Over time this becomes an admission standard that decides who lists, at what tier and at what premium. We quantify risk based on the Princeps methodology and data ingestion directly from Prime Intellect's own transaction volume.

How do we grade? The grade composes two factor groups: the site’s physical resilience (from Princeps data) and its observed delivery record (from the order flow your exchange already generates) into a single score.

Physical resilience
How robust is the site against failure?
On-site generation & storage 118
Grid-interconnection age & capacity 137
Cooling & thermal headroom8
Chip inventory match9
Permits & build quality7
Total78.0
×
Delivery track record
Does it actually deliver?
Advertised-vs-actual availability6
Spin-up success rate7
Order-fulfilment history6
Oversubscription signal5
Total60.0
72.0
Princeps Delivery Grade → Reliable
Illustrative composition. The left factors come from outside-in data (satellite imagery, permits, grid records, chip inventory); the right factors from the observed delivery record your exchange already generates. Power carries the largest single weight (~54% of the most impactful outages).11 A provider can opt in to share telemetry.

01 What makes delivery risk as meaningful as price?

A buyer opens On-Demand GPUs on Prime Intellect and picks an H100 80GB. The marketplace aggregates supply from the operators clearing on it — today, by API, that set is Crusoe, DataCrunch (Verda), Lambda, Massed Compute, Nebius and Vultr1 — and returns those options sorted by price, giving the buyer price, location, socket and a binary isolation flag, but no reliability signal. Prime Intellect itself acknowledges as much: "We don't provide formal SLAs at this time… verify uptime and past performance before deploying long-term workloads."6

On-Demand GPUs · Choose GPU — prices update in realtime
Recreated from the Prime Intellect console, 20 July 2026
B200
180 GB VRAM
$6.16/hr
H200
141 GB · x4·x1
$4.05/hr
H100
80 GB · x8·x4·x1
$3.39/hr
GH200
96 GB VRAM
$1.99/hr
A100
80 GB · x8·x4·x1
$1.23/hr
L40S
48 GB · x8·x4·x1
$0.82/hr
RTX Pro 6000
96 GB · x2·x1
$1.80/hr
V100
16 GB VRAM
$0.22/hr
H100 80GB × 1 · Compare Offers
Sort by Price ▾Location: Any
2 clusters available · recreated from the live console
Powered bySocketSecurityLocationSpin-up$/hr
Verda (DataCrunch)SXM5SECURE_CLOUDFinland~3 min$3.25Deploy
Lambda CloudSXM5SECURE_CLOUDUnited States~5 min$4.29Deploy
Live Prime Intellect offers, pulled via API 20 July 2026. Two SECURE_CLOUD H100s, a 32% price gap — driven by location and energy cost, not by any exposed measure of which provider is more likely to complete the buyer's job.

Price is a real signal for cost, not for delivery. These two variables are not reliably correlated, and the buyer cannot tell the reliable H100 from the unreliable one.

2.4%
goodput lost by top-tier operators — vs 3.4–10.7% mid-tier, identical hardware.2
419
unexpected interruptions in 54 days on Meta's 16,384-GPU Llama 3 run (~1 every 3 hrs).7
>30%
improvement in large-job completion from removing "lemon nodes" from one fleet.8
~54%
of the most impactful data-centre outages come from power, the single largest driver of correlated failure, and the heaviest weight in the grade.11

What separates a good operator from a bad one is not whether hardware fails (which it inevitably will) but how well that run is contained. In one study, Meta held >90% effective training time because only 3 of 419 incidents needed manual intervention.7 A weaker operator can run up to 10.7% lost compute. The failure waterfall scales with cluster size, meaning that the larger the reserve, the more a run depends on the operator's containment and rerouting capabilities.

Failure cadence scales with cluster size

Mean time between failures at ~1 per 50,000 GPU-hrs · Epoch AI; Meta Llama 37,9

The real dispersion, by GPU

Cheapest vs dearest single-GPU offer · Prime Intellect API1

02 An absence of existing market signals

Princeps reconstructs the loss history for neoclouds and compute providers from zero to one. Our fundamental belief is that risk must be assessed at origination — across both the physical and performance layers — to be meaningfully quantified. This is the single most mispriced risk by compute marketplaces right now. Uptime Institute tiers measure redundancy, not AI-readiness (built for 5–10 kW racks when a GB200 rack exceeds 120 kW) and certification is point-in-time and gameable.10 SLAs cap the remedy at a service credit, usually a 10% credit on a ~$3 instance is about 30 cents, against outages averaging near $1 million.11 The most serious independent effort is ClusterMAX. SemiAnalysis is both the strongest proof that reliability varies enormously and that an independent grade has real traction.

Exhibit 1 · ClusterMAX — the reliability spectrum already exists, but it is not yet a live number

SemiAnalysis's ClusterMAX has been the de-facto GPU-cloud reliability benchmark but only CoreWeave reached their Platinum ranking.12 SemiAnalysis's own verdict: "the bar across the GPU cloud industry is currently very low." The gap is not the silicon. On identical hardware, ClusterMAX measured the difference between a poorly-tuned and a well-tuned operator as roughly 60% vs 98% usable network efficiency. It is the operator, not the GPU, decides whether the compute is usable.

Untuned operator
~60%
Tuned operator
~98%
Usable network efficiency on identical hardware, ClusterMAX methodology.12

Prime Intellect's own beta cluster landed at Bronze, flagged for "a lack of health checks or dcgmi integration, and no monitoring dashboard". Princeps extends the same evidence base into a continuous, per-contract risk quantification your exchange can act on in real time as a live number (not a quarterly review).

So even a fast, competent operator has no independent, buyer-visible, real-time reliability number to stand on. Princeps fills this intelligence for both the seller and the marketplace provider.

03 Methodology

Reliability can be inferred from the physical and structural signals around a provider, and confirmed by whether the provider actually delivered against real orders. Princeps grades ~940 operators into five bands. Critically, none of this requires probes inside an operator's stack or a device install on your marketplace, instead the grade we deliver is built from external data and the delivery record your exchange already generates.

What the grade is built from
  • On-site generation & storage — satellite imagery of what powers a site and its resilience to grid events (power = ~54% of the most impactful outages).11
  • Grid-interconnection age & capacity — ~10,300 US projects sit in queues with a median wait near 55 months.13
  • Chip inventory & specs — whether the operator physically holds the capacity it lists.
  • Permits, build quality, cooling type — the difference between a Tier badge and a GB200-ready facility.
  • External order & marketplace signals — does advertised availability match the real book?
  • Verified tier — the observed delivery record — spin-up success, advertised-vs-actual availability, completion against booked jobs. On a marketplace this already lives in your order flow. Providers can optionally deepen it with device-level telemetry through our hardware partner ArcLeap.
Delivery confidence ▾Princeps Grade20 Jul 2026
Powered byPrinceps GradeTierDelivery confidence
NebiusVerifiedGold
92
CrusoeVerifiedGold
90
LambdaReliableSilver
68
VultrReliableTier III
61
DataCrunch (Verda)WatchBronze
46
Massed ComputeElevatedUnderperformer
30

04 The live delivery benchmark

The grade is a reliability and performance benchmark, calibrated to the measured failure and goodput evidence above. Pick a GPU, a cluster size, a term and a checkpoint interval. Every row and chart recomputes on that exact contract, using a Monte-Carlo failure simulation. Prices are live from Prime Intellect's marketplace.

Clearing at /GPU-hr on Prime Intellect right now.
512
30 d
1.0h
Contract value
Risk on Unrated · P50 / P99
expected / tail wasted compute, no recourse
Performance uplift, Verified vs Unrated
extra usable compute delivered
Expected failures over run
at ~1 / 50,000 GPU-hrs
Grade Reliability
30-day delivery
Performance
usable goodput
Risk · expected
wasted compute (P50)
Risk · tail
bad-luck run (P99)
Failure exposure Delivery confidence

05 Failure simulation

The risk figures are derived from a Monte-Carlo simulation of the contract you set above. Each of 4,000 trials draws on both hardware failures9 and a severity per failure (how much work is lost between checkpoints and how long recovery takes, scaled by the operator's containment). The result is a full distribution of delivered goodput. The gap between the expected outcome and the unlucky tail (P99) is the risk a buyer carries silently today and exactly what a delivery grade lets your exchange surface, quantify and rank on.

Wasted compute — your best vs worst live supply

4,000 simulated runs · Nebius/Crusoe (Gold) vs Massed Compute (Underperformer) — two real operators on your marketplace

Expected vs tail loss, by provider

P50 (expected) and P99 (bad-luck) wasted compute, $ — grade sourced from each provider's ClusterMAX tier

Model: failures ~ Poisson(N·24·T / 50,000); per-failure lost cluster-time ~ checkpoint-interval rework + recovery, exponential severity, containment scaled so each grade's mean matches its measured goodput-loss band [SemiAnalysis/Nebius]. Illustrative — for structure, not a quote. Synchronous training assumed (a failure stalls the whole cluster to the last checkpoint).

The rating applied to the six operators live on Prime Intellect's marketplace

We joined the six providers clearing on your marketplace to the Princeps database and pulled each operator's independent tier, ownership model, chip generation and site footprint.

Exhibit 2 · Princeps grade & live per-provider delivery risk — Prime Intellect supply, joined live, 20 Jul 2026
ProviderPrinceps gradeClusterMAX tier
(according to SemiAnalysis)
Usable goodputWasted · expected (P50)Wasted · tail (P99)Site evidence
Joined live: providers from the Prime Intellect availability API × our NeoCloud site database (both 20 Jul 2026).1 ClusterMAX tiers are SemiAnalysis's;12 the Princeps grade is our mapping from that tier plus ownership, power and chip generation. Risk is from the same Monte-Carlo simulation as the benchmark, recomputing on the contract set above. Ordering is independent of price yet it exactly inverts the marketplace's cheapest sort: the cheapest L40S and A100 on Prime Intellect are Massed Compute (Underperformer) and the cheapest H100 is DataCrunch (Bronze), while the top-graded operators (Nebius, Crusoe — Gold) sit at the pricier end. Cheapest-to-deliver selects the weakest-graded supply.

And on the buyer's screen, it looks like this

H100 80GB × 1 · Compare Offers
Reliability-adjusted ▾Princeps GradeGrade ≥ Reliable
Same live offers, re-ranked by delivery grade
Powered bySecurityLocationPrinceps Grade$/hr
Lambda CloudSECUREUSVerified$4.29View grade
Verda (DataCrunch)SECUREFinlandReliable$3.25View grade
Illustrative. The buyer now sees an orthogonal, quantified reliability axis next to price — and can sort and filter supply on delivery risk instead of guessing from the sticker price.

06 The impact of an independent credibility layer

That a grade transforms a marketplace is one of the better-established results in economics. Akerlof's 1970 "market for lemons" showed that when buyers can't observe quality they pay only for average quality, good sellers exit, and the market can unravel. The fix is the creation of "counteracting institutions" such as certification and third-party grading.4

Exhibit 3 · Independent Grading
MarketCredibility layerMeasured effect
Grain, goldUSDA grades / LBMA Good DeliveryMade the asset exchange-tradeable5
Corporate debtCredit ratings (NRSRO)Unlocked mandated institutional capital14
Enterprise SaaSSOC 2 attestationDe facto gate to enterprise procurement15

The mechanism is identical: an independent layer converts an unobservable quality attribute into a verifiable, priceable signal — raising prices for good sellers, raising transaction probability, widening participation, and shrinking the pooling that drives adverse selection.

Clearing prices rise. A grade is most valuable exactly where trust is scarcest and sellers are otherwise indistinguishable, as in an aggregated GPU exchange. Volume shifts to quality, and the buyer pool widens: enterprises default to hyperscalers on trust, not price16. An independent, quantified reliability grade encourages marketplace activity from buyers currently hedging on reliability.

07 What we would build together

The Princeps intelligence layer

Underneath the grade sits a continuous intelligence layer that scores every operator across six dimensions:

Service reliability
Power resilience
Geographic concentration
Hardware configuration & performance
Provider financial health
Delivery history & replacement-capacity risk

Princeps would provide ongoing risk intelligence and quality scoring for Prime Intellect — a live credit-scoring index for the operators clearing on your marketplace. It can be offered as a standalone paid intelligence service, or integrated directly into the Prime Intellect platform: the grade column, admission gating, and Reserved-cluster qualification all drawing on the same index.

The End Goal: a compute credit rating built on your transaction data

The integration above is the starting point and we are keen to explore a larger, longer-term collaboration opportunity alongside it. Prime Intellect generates data most of the market cannot see — a continuous record of which operators actually delivered, drawn from every booking, fulfilment and failure across the exchange. Combined with the Princeps methodology, that record is the foundation for a jointly-operated Compute Credit Rating. Over time it would position Prime Intellect as a reference point for compute reliability, much as the Chicago Board of Trade's grading standards became the basis on which the grain market settled.5

The path from grade column to rating standard runs in three phases:

Princeps operates the rating independently using Prime Intellect as the data spine and reference venue. A rating the seller controls is not an independent rating. Independence is what makes standards credible to both buyers and lenders, and ensures marketplace liquidity in the long term. The rating improves the more volume clears through your exchange.

Princeps
Princeps — the independent grade for compute

Sources

1. Prime Intellect availability API, pulled 20 Jul 2026.

2. SemiAnalysis (commissioned by Nebius), goodput-loss benchmarking, 2026.

3. Prime Intellect, "Head of Compute" JD, 2026.

4. Akerlof, "The Market for 'Lemons'," QJE 84(3), 1970.

5. USDA AMS; LBMA Good Delivery; CME Group (CBOT futures, 1865).

6. Prime Intellect docs, FAQ, 2026.

7. Meta, "The Llama 3 Herd of Models," arXiv:2407.21783, 2024.

8. Meta, "Revisiting Reliability in Large-Scale ML Research Clusters," arXiv:2410.21680, 2024.

9. Epoch AI, GPU-cluster failure-scaling model (~1/50,000 GPU-hrs), 2024.

10. American Compute, "Data Center Tiers," 2026; Uptime Institute.

11. Uptime Institute, "Annual Outage Analysis 2024"; "Cloud SLAs punish, not compensate," 2022.

12. SemiAnalysis, "ClusterMAX 2.0," 2025.

13. LBNL, "Queued Up: 2025 Edition," 2025.

14. White, "The Credit Rating Agencies," JEP 24(2), 2010.

15. ControlCase (industry reporting), 2024–25.

16. McKinsey, "The evolution of neoclouds," 2025–26.

17. Dave Friedman, "The Compute Market has Multiple Views on Future Compute Prices," 2026.

contact@princeps.dev · anya@princeps.dev
© 2026 Princeps · The risk & grading layer for compute

For discussion only; not an offer of insurance or a quotation. Marketplace prices are pulled live and change in real time. The failure simulation is calibrated to measured goodput-loss and failure-rate data and is illustrative of structure, not a quote. Grades are a framework, not a published assessment of any provider.