About Metrale
It started with two words.
In January 2026 a working improvement to llama.cpp was closed because it had been written with AI. The author wrote a short appeal to common sense, and when it went over everyone’s heads, answered the room with a question. Then he went and built the engine from scratch.
How it started
The exchange
What came before it

697 stars later, the repository those two words started runs on hardware from NVIDIA and AMD, merged a kernel into Hugging Face Transformers, and out serves the incumbent on the published ladder. The answer to the question is this website.
January 7, 2026
The pull request
A loop attention model ran on a DGX Spark. The pull request adding it to llama.cpp was closed as containing AI generated code without disclosure. The reply argued that whether AI or a compiler, both translate one language to another, and that a community building AI tooling should not hold contempt for AI written code.
Read the thread ↗January 8, 2026
“Your point?”
Asked why the pull request looked entirely AI generated, the author answered with two words. They became the first principle of the repository that followed. AI authored is the default. A human who writes code by hand explains why they were better than the machine.
Winter 2026
From scratch, in Rust
Months of trying to improve vLLM on the Spark had shown that the feedback loop from a kernel change to a number was too slow to learn from. The serving stack was rewritten in Rust with hand tuned CUDA, no Python, and a build that takes a minute instead of forty.
May 2026
One Reddit post
A stable 102 tokens per second on a DGX Spark, posted to r/LocalLLaMA. The star count went from a few dozen to a few hundred in a week and the Discord became the test fleet.
July 2026
Receipts
The fused Qwen Gated DeltaNet kernel merged into Hugging Face Transformers. MLCommons named the project a contributor to the new MLPerf edge agentic benchmark. AMD provided a Strix Halo desktop and the MLPerf submission went in from the same CUDA source on both vendors.
August 2026
The ladder
The concurrency ladder against the matched vLLM configuration was published with every rung lost on the way. Eight rungs, eight wins, and the margin widest at C=128.
September 2026
Metrale
The company took a new name, built on metron, the Greek word for measure, brought in commercial leadership that had scaled Anaconda, and set out to sell what the engine had proved. The industry measures inference in tokens per second. Metrale measures what that performance is worth.
What the name means
Measure first.
The name Metrale is built on metron, the Greek word for measure. It is the root of meter, metric, geometry and symmetry, and further back, of moon and month, the first measures of time. Protagoras used it when he called man the measure of all things. It is the oldest word for knowing how much.
Economies were run on instinct until they were measured. In 1934 Simon Kuznets gave the United States Congress its first national income accounts, and within a decade those accounts were how nations were compared and managed. Today the number is called GDP. Kuznets warned in the same report that the welfare of a nation can scarcely be inferred from its income. The number was a beginning, not a verdict.
Inference is where economies were then. Fleets report tokens per second the way a mill once reported spindle speed, and few can say what a GPU hour produced, what a workload cost, or when the fleet paid for itself. Metrale keeps those accounts: the gross product of inference, by model, cluster and business unit, against the baseline you ran before. Measure first, then manage.
Mission
Same silicon. Smarter inference. Stronger scalability.
AI worth having should run on hardware you own, whether that is an accelerator at the edge, the workstation under your desk, or a rack you operate. We build one engine for the whole range, verify it on the silicon we can put our hands on, and make what it produces accountable to the people paying for it.
Receipts, not adjectives
Every performance number on this site is generated from a record in the repository. If it is not in the repo, it is not on the page.
AI first, human accountable
AI authored is the default in the repository. Certified benchmarks gate every kernel change. People decide what ships.
Own the request path
Security, governance and economics are only exact when the engine is under the workload. Everything we build follows from that.
Open at the core
The Community Edition is AGPL-3.0 and always will be. The enterprise platform pays for the people who keep it that way.
Built by people who have stood up operations for
United StatesCyber Command
Naval SpecialWarfare Command
Prior roles of the founding team and core contributors. Listed for background, not as customers or endorsements. The appearance of U.S. Department of Defense visual information does not imply or constitute DoD endorsement.
Programs and partners
Program member. DGX Spark hardware provided.Strix Halo hardware provided. MLPerf submitted on it.
Named contributor to the MLPerf edge agentic benchmark.
Fused Qwen GDN kernel merged into Transformers.
Dev Ambassadors. A recipe for every release.
SCALE by Spectral Compute One CUDA source, NVIDIA and AMD.
Next step
Come build with us, or come buy from us.
Both conversations start the same way. Tell us what you run.