Financial Grounding

The Substrate Measured: What Rented Compute Actually Costs

Every number on this page is derived from the repository's append-only cost ledger. In a marketplace where 62% of machines fail to deliver working GPUs, fail-fast verification is what makes the economics work.

Substrate MeasureObserved ValueAccounting Definition & Grounding
Distinct Instance Creates121Counted by distinct provider instance IDs across all campaigns.
Total Create Events127Includes immediate retries following transient API timeouts.
Never-Delivered Instances75 (62.0%)Instances that froze during pull, failed SSH, or died on bf16 preflight.
Time to Running (Median / p95)2.4 / 8.6 minMeasured only when a host successfully delivered a verified GPU.
Mean Alive per Bad Host~12.3 minShort lifetime enforced by 6-min frozen pull watchdog & 10s preflight.
Spend on Bad Hosts$4.9611.9% of total campaign spend. Absorbed without stalled engineering.
Total Campaign Marketplace Spend$41.62Includes all 121 instances, all 75 failures, and local verification.
Gated Model Versions16Model candidate iterations evaluated through the framework gate.
Deploy Gate Refusals6 (38%)Refused due to statistical shortfall, canary failure, or format clash.

Why Two in Three Rented Hosts Fail to Deliver

Roughly two in three creates died before any training occurred, and on documented platform-wide bad days the failure rate reached three in four. That is not the pipeline failing: that is the consumer GPU marketplace being what it is.

Marketplace providers aggregate rigs located in residential basements, student apartments, and makeshift mining frames. Rigs report modern GPUs while running six-year-old NVIDIA drivers, sit behind throttled 20 Mbps residential uplinks that freeze on a 15GB container pull, or crash the moment PyTorch allocates a tensor.

The number that proves the engineering is $4.96: the total cost of all 75 bad host attempts combined. Because tuneharness detects frozen downloads in six minutes and verifies compute in ten seconds, dead hosts are terminated before they can idle-burn dollars.

Interactive Spend Comparison Calculator

Use the interactive calculator below to model your expected fleet spend. Adjust training volume and failure rates to see how renter-side guardrails protect your budget.

Training Runs Per Month12 runs
Training Hours Per Run3.0 hrs
Marketplace Host Failure Rate62%
Hyperscaler Managed Cloud:$64.80
Naive Marketplace (w/ Unnoticed Waste):$49.25
tuneharness Guarded Marketplace:$29.53
Idle-Burn Waste Prevented:$19.72
Net Operator Savings:$35.27

The Five Money Traps We Closed With Code

Every defensive guard in tuneharness exists because a real incident cost money:

  1. The 105-Minute Stale-Driver Idle Burn: A host passed nvidia-smi, accepted the container, and sat idle while training failed to initialize due to a driver incompatibility. An operator noticed after 105 minutes. We wrote the post-boot bf16 matmul preflight the same evening; now bad drivers are caught and destroyed within ten seconds.
  2. The 30-Minute Frozen Container Pull: Consumer hosts with saturated uplinks frequently freeze on multi-gigabyte layers. The boot progress parser now enforces a hard six-minute progress deadline.
  3. The Duplicate Instance Leak: When an API call timed out, retrying naively left an unrecorded instance running on the provider. The lease manager now reconciles by deterministic label before issuing any retry.
  4. The Ctrl-C Abandoned Rental: An operator terminating a local CLI with Ctrl-C once left a GPU instance running overnight. Exit traps now intercept all abort signals, execute an automatic STOP command, and preserve the disk volume.
  5. The Silent Filter That Matched Nothing: A search filter query emitted gpu_ram>=24576 (MB) while vast.ai expected GB, causing the search to return zero offers and appear as a dead market. We added live-market assertion tests that verify filters actually match real inventory.