Skip to content

Platform

One platform for the whole inference lifecycle, from the kernel to the invoice.

Metrale is three layers that share one request path. The engine makes the GPUs faster, the control plane keeps the fleet honest, and the economics layer turns what they report into numbers finance can sign.

Bringing up encrypted links to offsite grids, then queueing a large project across them and watching it run.

Metrale Console. Demo data, recorded from the product mockup.

Architecture

One request path. One ledger.

Requests flow from your applications through the router to Metrale Engine nodes on your GPUs. Metrale Control manages rollout, policy and repair out of band. Metrale Economics collects telemetry from every node into a ledger.Your applications and agentsOpenAI · Anthropic · Responses APIsGPU aware routerKV reuse · queue · VRAM pressureMetrale Engine · node 1signed recipe · kernels for this siliconGB10 · NVFP4Metrale Engine · node 2signed recipe · kernels for this siliconH100 · FP8 · bring upMetrale Engine · node 3signed recipe · kernels for this silicongfx1151 · SCALEMetrale Controlout of bandrolloutcanarypolicyrepairscaleMetrale Economicsworkload × model × runtime × configuration × GPU × cluster$ per million tokens · $ per workload at SLO · productive GPU hoursstranded capacity · chargeback by business unit · payback
The request path never touches the control plane. The ledger reads what the engine measured at the source.

Deployment surfaces

Private datacenters, AWS, Azure, GCP and neocloud GPU pools. Hosted with private connectivity where you want us to run it, bring your own cloud where you do not.

  • Private datacenter
  • Air gapped network
  • AWS
  • Azure
  • GCP
  • Neocloud GPU pools
  • Workstation and edge

Started on a single box. Built to run a fleet.

Metrale began as local inference on a DGX Spark. The same pinned stack scales to datacenter GPU fleets with the qualification record, routing and observability that enterprises need in production.

Next step

See the platform on your workload.

A working session with the console, the ladder and the payback model, on demo data or yours.