Pricing
Priced against productive GPU capacity, not seats.
Enterprise AI infrastructure software already prices per GPU per year. Metrale sits inside that range, ships with the control plane and the economics layer, and shows its payback on this page. Proposed list prices, September 2026.
Proposed sheet · September 2026

Waitlist open
Community Edition
The engine and every recipe, free. For developers, labs and anyone running open models on hardware they own. Not released yet.
- Metrale Engine, full source
- Every model recipe in atlas-recipes
- OpenAI, Anthropic and Responses APIs
- LAN fleet manager, early access
- Community support in Discord
Workstation and edgePROPOSED
A DGX Spark or Strix Halo class box serving an office, a branch or a field team. Commercial license, signed update channel, managed from the console.
- Commercial license per box
- Signed stable and LTS channels
- Console access for every licensed box
- Email support, next business day
- Volume pricing from 25 boxes
Most fleets start here
EnterprisePROPOSED
The full platform for GPU fleets. Realized pricing at fleet scale runs $1,800 to $2,400 per GPU per year. Support and forward deployed engineering priced separately.
- Metrale Engine, commercial license
- Metrale Control, rollouts, routing, policy, repair
- Metrale Economics, chargeback and payback
- Named engineer and response SLA
- Hosted, your cloud, on premises or air gapped
Proof of value
One model, one hardware target, one workload. A side by side ladder in week one and a receipt in dollars per workload at the end.
- Scoped success criteria, yours or ours
- Side by side against your current engine
- Economics baseline of the target cluster
- Forward deployed engineer for the four weeks
- Fee credited against the first year on conversion
Market anchor
Where it sits.
Established enterprise AI infrastructure software already prices against GPU capacity. Metrale lists inside the range and includes the layers the others sell separately.
| Product | List | Basis |
|---|---|---|
| Red Hat AI Inference Server Hardened vLLM, published list price | ~$2,500 | per accelerator per year |
| NVIDIA AI Enterprise Broad platform, OEM backed | $4,500 | per GPU per year |
| Metrale Enterprise Engine, control plane and economics, realized $1,800 to $2,400 at scale | ~$3,000 | per GPU per year, list |
Third party prices are public list prices at the time of writing and belong to their owners. Metrale prices are proposed and subject to contract.
Illustrative contract economics
What a fleet costs to license.
At realized fleet scale pricing. Illustrative, not a forecast.
What the platform meters
Priced on a number you can watch.
Proposed. The platform is being built to meter what it serves and show it live, so a renewal is read off the same number the console shows.
GPU hours and GPU count
A license is a cluster or a number of GPUs, discovered by the platform, not declared on a form.
Tokens per GPU second
The throughput the fleet actually produced, per GPU, per second, next to the throughput it could have.
Cost per million tokens
GPU, storage, network and platform cost over the tokens delivered, for your fleet and for the baseline you ran before.
Nothing through the gateway without a license
Every served request is entitled and counted, so the bill and the telemetry are the same record.
Payback
Find your payback period.
If a thing costs three thousand dollars and makes you a thousand a month, it pays for itself in three months, and everything after is upside. That is the number to walk to the CFO with. Three scenarios, every input editable, evidence class on every field.
The uplift frees GPUs. Freed GPUs are deferred purchases or rentals plus the power they burned. The license is what the uplift costs.
Uplift defaults to 1.20x, below the measured ratio on the GB10 ladder at C=128, because a datacenter part is not a Spark until we publish the receipt.
- GPUs freed by the uplift
- 42.7
- Deferred purchase or rental
- $1,706,667
- Power no longer burned
- $40,815
- Gross savings per year
- $1,747,481
- Metrale license per year
- $614,400
- Net per year
- $1,133,081
- Three year net
- $3,399,244
A model, not a quote. Savings depend on your workload, your utilization and the uplift measured on your hardware during the pilot. Evidence classes: MEASURED from ladder.generated.json, the published concurrency ladder. PROPOSED a proposed list price from this page, the team can change it. USER yours to edit, the model recomputes as you type.
Questions
The questions we actually get asked.
Short answers. Each one is backed by something on this site or in the repository.
How is it priced?
Per GPU per year for the Enterprise Edition, with volume tiers as the fleet grows, and a per box license for workstation and edge deployments. Support and forward deployed engineering are priced separately. The Community Edition is free under AGPL-3.0. The pricing page lists the proposed sheet and a payback model with editable inputs.
How do you measure savings?
Against your own baseline. Economics records what each cluster cost per workload before Metrale takes traffic, then reports the delta as traffic moves. The pilot ends with a receipt in dollars per million tokens and dollars per successful workload, not a slide.
What is the payback period?
It depends on your fleet, your utilization and the uplift we measure on your workload. The model on the pricing page computes it from inputs you control. Every month after payback is upside, which is why we talk about payback rather than a percentage.
What about SOC 2 and compliance?
The architecture is built for regulated buyers, and SOC 2 readiness documentation, model risk documentation and pinned recipe governance packs are part of the first SLA engagements. Ask for the current state of the audit program when you book. We will tell you exactly where it is.
What support comes with it?
Community support in Discord for the open source engine. Enterprise includes a named engineer, a response SLA and a shared channel. Forward deployed engineering for the pilot and the cutover is scoped per engagement and credited against the first year on conversion.
Ask the rest in a working session, or read the deployment guide ↗.
Next step
Get the sheet, or get the receipt.
Email sales@metrale.com for the full price sheet, or book a working session and we run the ladder on your workload.