Skip to content
New Metrale Engine serves 1.333× the matched vLLM configuration at C=128 on the same GB10, with every rung published. See the ladder →

Inference economics, reimagined.

  • Speed
  • Security
  • Governance

Faster inference. Stronger governance.
A fraction of what you pay today.

Metrale is the inference economics platform for the GPUs you already own. It runs your models faster on the same silicon, keeps every prompt inside your perimeter, and shows your CFO what each workload costs.

Modeled at 70% or less of your current inference spend when the silicon is yours. Run the model

The Metrale Console answering a question from an air gapped fleet, with tokens per second and cost per request on screen.

Metrale Console. Demo data, recorded from the product mockup. Watch the one minute film

Built by people who have stood up operations for

Prior roles of the founding team and core contributors. Listed for background, not as customers or endorsements. The appearance of U.S. Department of Defense visual information does not imply or constitute DoD endorsement.

Programs and partners

The problem

Your GPUs report tokens per second. Your CFO pays in dollars per workload.

Every inference engine gives you operator statistics. Every FinOps tool gives you a bill. Nothing connects the two, so the fastest engine in the rack still cannot say what a workload cost, which cluster stranded capacity, or whether last quarter’s upgrade paid for itself.

What the infrastructure sees

  • TTFT and TPOT
  • Tokens per second
  • Queue depth
  • KV cache hit rate
  • GPU utilization
  • Kernel and driver versions

What the enterprise pays for

  • Dollars per million tokens
  • Dollars per successful workload at SLO
  • Productive GPU hours
  • Stranded capacity
  • Cost by model, cluster and business unit
  • Savings against the production baseline

Metrale is the correlation layer between the two.

Every unit of work is tied to the workload, model, runtime, configuration, GPU and cluster that produced it. An operator number becomes a finance number without a spreadsheet in between.

The platform

One platform. Three layers. Every GPU, every workload, every dollar.

Requests flow from your applications through the router to Metrale Engine nodes on your GPUs. Metrale Control manages rollout, policy and repair out of band. Metrale Economics collects telemetry from every node into a ledger.Your applications and agentsOpenAI · Anthropic · Responses APIsGPU aware routerKV reuse · queue · VRAM pressureMetrale Engine · node 1signed recipe · kernels for this siliconGB10 · NVFP4Metrale Engine · node 2signed recipe · kernels for this siliconH100 · FP8 · bring upMetrale Engine · node 3signed recipe · kernels for this silicongfx1151 · SCALEMetrale Controlout of bandrolloutcanarypolicyrepairscaleMetrale Economicsworkload × model × runtime × configuration × GPU × cluster$ per million tokens · $ per workload at SLO · productive GPU hoursstranded capacity · chargeback by business unit · payback
The request path never touches the control plane. The ledger reads what the engine measured at the source.

Started on a single box. Built to run a fleet.

Metrale began as local inference on a DGX Spark. The same pinned stack scales to datacenter GPU fleets with the qualification record, routing and observability that enterprises need in production.

The value

Faster inference. Stronger governance. Higher savings.

Measured on the box we can put our hands on. Modeled where the box is yours, with the inputs on the page.

Faster inference
1.333×

the throughput of the matched vLLM configuration at C=128, same GB10, same checkpoint, same client.

Read the ladder
Stronger governance
1 binary

about 75 MB of Rust and CUDA. No Python or PyTorch in the request path. Signed, gated and replayable.

See the posture
Higher savings
4 months

to pay back the license on a 256 GPU fleet at a 1.20x uplift. Every month after is upside. Edit the inputs yourself.

Open the model

Proof, not a pitch

Same GB10, same checkpoint, same client. Eight rungs, eight wins.

We publish the concurrency ladder against the matched vLLM configuration from C=1 to C=128, with every rung we lost on the way to it. The margin is widest at the top, which is where fleets of tool calling agents actually run.

478.11 tok/s at C=128, unsloth/Qwen3.8-27B-NVFP4
8/8 rungs won against matched vLLM + MTP
138 concurrency gate records passing across the repo
Published GB10 ladder reproduce it
0129258387516C=1C=2C=4C=8C=16C=32C=64C=128Metrale 478 tok/svLLM + MTP 359 tok/s

unsloth/Qwen3.8-27B-NVFP4 · NVIDIA GB10 Grace Blackwell, 121.7 GB unified · mean tok/s over 3 timed reps (1 warmup discarded). ISL 128 / OSL 1,024, temperature 0. Every rung in the campaign log.

The console

One console for every inference workflow.

Recorded from the product mockup with demo data. The shipped product will differ. The workflow will not.

Pick a model. Pick a grid. Start.

Choose the model, choose the compute grid it runs on, and start getting answers from an air gapped fleet with the tokens per second and the cost per request on the same screen.

See it on your workload
Picking a model and a compute grid, then streaming an answer with live throughput and cost.

Bring up the links. Queue the project.

Launch encrypted links to the grids you lease offsite, then queue a large project across all of them. It is sized, priced against your baseline, and placed only where its data class allows.

See it on your workload

Every node, every kernel, every rollout.

Nodes report health, kernel versions and utilization. Canary a release to five percent, watch TTFT, roll back on a breach without waking anyone up.

See it on your workload

Dollars per workload, not tokens per second.

Chargeback by business unit, productive GPU hours, stranded capacity and the payback clock, against the baseline you ran before Metrale.

See it on your workload

Policy, provenance and the audit trail.

Data residency, model allowlists and redaction as policy. Every response traces to a signed recipe, a kernel build and a gate record.

See it on your workload

Why Metrale

Three differences you can test.

The market is loud. Every runtime claims speed and every dashboard claims visibility. Here is what is actually different, and the question to put to anyone else.

01Speed

Same silicon, more tokens.

Metrale compiles a kernel set per hardware, model and quantization instead of shipping one generic path. On the published GB10 ladder it wins every rung against the matched vLLM configuration and keeps scaling from C=64 to C=128 while the baseline flattens.

Ask for the throughput curve at C=128 on your workload, not a single stream number on theirs.

Go deeper on speed
02Security

Nothing leaves your perimeter.

One signed Rust binary with no Python or PyTorch in the request path. Prompts, weights and telemetry stay on hardware you own, in your cloud account, or on an air gapped network. The control plane never sits on the inference path.

Ask what is in the request path, and who audits the two hundred dependencies behind it.

Go deeper on security
03Governance

Every token has a receipt.

Every response traces to a signed recipe, a kernel build and a gate record. Every GPU hour is attributed to a model, a cluster and a business unit. Governance is a ledger finance can sign, not a dashboard operators tolerate.

Ask what a workload cost last Tuesday, by business unit. A tokens per second chart is not an answer.

Go deeper on governance

Only Metrale

Each layer earns the next.

The engine earns the deployment

Nobody installs a governance layer for its own sake. Metrale gets into the fleet by making the same GPUs produce more inference.

The deployment earns the telemetry

Once the engine owns the request path, every kernel, cache and queue decision is observed at the source instead of inferred from a proxy.

The telemetry earns the economics

With source data in hand, cost per workload, stranded capacity and payback stop being estimates and become a ledger.

One architectural decision, made at the start. Own the request path, then measure what it is worth. Everything on this page is downstream of that.

Getting there

Complement or replace your serving stack, at the pace you choose.

Metrale runs beside vLLM, SGLang or llama.cpp on day one. Nothing is turned off and nothing is migrated until your own numbers say so.

01

Week one

Side by side on your workload

One binary per node, one signed recipe per model, next to what you run today. A side by side ladder against your current engine, on your hardware, in week one. The economics baseline starts recording the same day.

02

Months one to six

Traffic moves workload by workload

Production traffic shifts one model family at a time. The control plane takes over rollout, canary, rollback and fleet policy. Economics reports every cluster against what it cost before, by business unit.

03

At renewal

Your call, on evidence

Some customers keep the old engine for one model family. Others move off it entirely. You make that call having run both, on your own receipts rather than a vendor timeline.

From the fleet

Operators running it on their own hardware.

“Night and day compared to the 10 minute torch.compile cycle. Startup in about 15 seconds and it just stays coherent in an agentic loop.”
ronald_15496, Discord, #general
“Testing Atlas on a DGX Spark in an agentic workflow for over an hour. Super impressed. Spark is actually awesome with Atlas.”
PersonWhoThinks, r/LocalLLaMA
“I had grown tired of the usual stack and was hoping for something like this. Really surprised and impressed. So glad I bought a Spark.”
tetsuro59, Discord, #general

Quotes are verbatim. Atlas was the engine's name until the September 2026 rebrand to Metrale.

Questions

The questions we actually get asked.

Short answers. Each one is backed by something on this site or in the repository.

What is Metrale?

Metrale is an inference economics platform for GPUs you own or rent by the hour. Metrale Engine runs open models faster on the same silicon, Metrale Control deploys and governs the fleet, and Metrale Economics turns the telemetry into cost per workload, chargeback and payback. The engine is open source under AGPL-3.0. The platform is licensed per GPU.

Do I have to replace vLLM, SGLang or llama.cpp to use it?

No. Metrale deploys beside your current engine and takes traffic one model family at a time. Many teams keep the old engine for a family we do not ship a recipe for yet. You decide what moves, on your own side by side numbers.

How long does a deployment take?

One binary per node and one signed recipe per model. A single box runs in minutes from one install command. A fleet pilot has a side by side ladder against your current engine in week one. Production cutover is workload by workload over the following weeks, at your pace.

Who owns the data?

You do. Prompts, weights, outputs and telemetry stay on hardware you own, in your cloud account, or on an air gapped network. The control plane manages configuration, licensing, versions and aggregate metrics, and it never sits on the request path. In a bring your own cloud deployment no inference request leaves your account.

What hardware does it run on?

NVIDIA DGX Spark (GB10) is verified today, and AMD Strix Halo (gfx1151) runs the same CUDA source compiled through SCALE, with both submitted to MLPerf Inference v6.1. Hopper and Blackwell datacenter targets are in active bring up with receipts in the changelog. Expert parallelism across two nodes ships as recipes and a three node topology is being wired up.

Can it run air gapped?

Yes. The engine is one binary with no runtime download and no Python environment to resolve. Recipes, models and kernels are delivered as signed artifacts and installed from local media. Telemetry can stay entirely inside the network and export on your schedule, or never.

Which models can I run?

Every model on this site maps to a recipe in the atlas-recipes repository, which is the single source of truth, so the site cannot list a model without one. Qwen leads with the most recipes, alongside Gemma, Nemotron, Mistral, MiniMax and DeepSeek. Bring your own weights and we scope the bring up.

What does verified mean on this site?

An image ships only after the serve matrix passes. Every model boots, stays coherent under greedy determinism with no token leakage and reliable tool calls, and holds throughput within ten percent of its committed baseline. A release that ships slower than its baseline fails the gate. Every number on the benchmarks page is generated from a record in the repository.

How is it priced?

Per GPU per year for the Enterprise Edition, with volume tiers as the fleet grows, and a per box license for workstation and edge deployments. Support and forward deployed engineering are priced separately. The Community Edition is free under AGPL-3.0. The pricing page lists the proposed sheet and a payback model with editable inputs.

Why does concurrency matter more than single stream speed?

Because agentic systems do not send one request at a time. A fleet of tool calling agents sharing a context bus arrives as many concurrent streams, so an engine is judged where the requests pile up. On the published ladder Metrale keeps gaining throughput from C=64 to C=128 while the matched vLLM configuration does not, and an engine that flattens under load caps how many agents a box can run.

Ask the rest in a working session, or read the deployment guide ↗.

Next step

Ready to get the most out of the datacenter you already paid for?

Book a working session. We run the ladder on your workload, on your hardware, and hand you the receipt.