Skip to content
HomeBrewedLabsCommission

Metallum-1B is live. Open weights on Hugging Face. Apache 2.0. 1,002M parameters. Brewed on 1 × NVIDIA RTX 5090. 16.0B tokens, curated.

The frontier isn't a datacenter.
It's home brewed.

Open-weight models designed for sovereign deployment.

We brewed a 1-billion-parameter specialist from scratch on a single consumer graphics card. No cluster. No rented fleet. Small batch, full strength: a model fermented on your data, owned entirely by you.

1,002M
Parameters
Trained from scratch
1
Consumer GPU
NVIDIA RTX 5090, 32 GB
16.0B
Tokens seen
Curated, decontaminated
40k
Custom tokenizer
Fit on the domain

The thesis

Big models know
a little about
everything.

Your business does not need a model that can write sonnets and debug Kubernetes. It needs one that knows your domain: the vocabulary, the edge cases, the twenty years of institutional judgment sitting in your archives.

That model does not exist, because nobody is going to build it for you. Frontier labs optimize for the average of everything. Your domain is not the average of anything.

So build it yourself. A focused 1B specialist trained on the right data beats a general giant inside its lane, and it fits on hardware you can actually own.

What we brew

Three taps

Pick your pour: sovereign AI, closed or open, whichever your situation demands. The point is that the choice belongs to you.

01

Commission a specialist

Bring your domain data. We train a model that lives inside your business: your corpus, your weights, your IP. It runs where you say it runs, including fully disconnected.

  • You own the weights outright
  • On-premise, your cloud, or hosted by us
  • Nothing you send trains anyone else's model
Start a build
02

Host your own model

The platform we're building: bring a model you trained, whether ours, yours, or a fine-tune, and serve it without renting a hyperscaler or surrendering your data to an API.

  • Serve open or closed weights
  • Built for one-card budgets, not clusters
  • Early access opening in waves
Join the waitlist
03

Feed the distillery

Specialist models die on thin data. Donate structured data in a domain you know and the distillery turns it into an open dataset anyone can train on, with provenance attached, not scraped.

  • Openly licensed, permanently
  • Provenance and PII audited on intake
  • Contributors credited by name
Donate data

Proof of work

First batch, brewed on our own bench

Metallum-1B is our own model: a 1,002M-parameter ML / LLM-engineering specialist, pretrained from random initialization on one RTX 5090 in ~13.5 days. Here is what it measures.

Architecture
26L · d1792 · GQA 28/14
Context
2048 tokens
Optimizer
Muon (2D) + AdamW
Schedule
WSD, short-sequence curriculum
Hardware
1 × NVIDIA RTX 5090 · 32 GB · Blackwell sm_120

Headline measurements

The first was measured once, on a holdout sealed before training and opened after. The rest are held-out or generated per-seed, so none of them can be contaminated by training data.

  • H9 sealed holdout1.3953

    Bits per byte on 400k characters of post-cutoff Wikipedia, built and sealed before the run and evaluated exactly once. Contamination verified at 0.0099% shingle containment. One shot, no retries. This is the number that carries the release.

  • In-domain ML-eng problems7 / 150
    Qwen3-1.7B 7 / 150Qwen3-0.6B 3 / 150SmolLM2-1.7B 0 / 150

    A deliberately hardness-gated 150-task executable suite where small models sit near the floor by design. Metallum ties Qwen3-1.7B at 1.7× its parameter count and clears both other baselines. The suite was used for model selection during development, so it is selection-aware: a real result, not a clean one.

  • Long-context retrieval (RULER)86 / 90

    Synthetic needle-in-a-haystack, generated per-seed, so there is no training data to contaminate it.

  • Distance needles18 / 18

    Perfect retrieval at both D=1024 and D=2032, the second of which is effectively the full 2,048-token window. The NoPE layers are why.

Sealed holdoutHeld-out measurement

We also publish the numbers that do not favour us, and say plainly why they cannot be read as capability claims.

How it's built

Drawn, not assembled

Every dimension below came from a decision with a reason behind it. This is the drawing that decision set produced.

Metallum-1B decoder architecture blueprintA 26-layer decoder stack. Tokens enter a 40,000-entry tied embedding, pass through 26 identical decoder layers, every fourth of which drops rotary positional encoding entirely, then a final RMSNorm and a tied language-model head. Each layer contains a pre-norm, grouped-query attention with 28 query and 14 key-value heads plus QK-normalization, a residual add, a second pre-norm, a SwiGLU feed-forward of width 4,864, and a second residual add.DECODER STACKembed 40k · tied4812162024×26RMSNormLM head (tied)▮ NoPE layer: no positionalencoding, retrieval headsLAYER DETAIL: NoPElayers 4 / 8 / 12 / 16 / 20 / 24identical to RoPE layers except attentionreceives no positional signal at allRMSNormpre-normAttention · GQA28Q / 14KV · QK-norm⊕ residualRMSNormpre-normSwiGLUd_ff 4864⊕ residuald_model = 1792SPECIFICATIONd_model1792n_layers26heads28 Q / 14 KVd_ff4864vocab40,000context2048RoPE θ500,000NoPEevery 4thembeddingstiedparams1,002,127,360METALLUM-1BML/LLM-ENGINEERING SPECIALISTCKPT s2b075SCALE NTSHOMEBREWEDLABSSHEET 1/1
Decoder stack as built. The six highlighted layers drop rotary encoding entirely. That choice is why a 2,048-token model retrieves perfectly at the edge of its own window.

NoPE every 4th layer

Six of twenty-six layers carry no positional encoding at all. Those layers learn length-generalizing retrieval heads, which is why a 2,048-context model still scores 18/18 on distance needles at the very edge of its window.

QK-norm + z-loss

Normalizing queries and keys before the attention product, plus a z-loss term on the logits, is what kept a 16B-token run from diverging on a single card with no cluster to restart from.

GQA 28Q / 14KV

Halving the key-value heads halves the KV cache. That is the difference between a long-context model that fits on consumer hardware and one that does not.

Document-masked attention

Attention never crosses a document boundary inside a packed sequence, so the model never learns to continue one document into an unrelated one.

FIM sentinels

Fill-in-the-middle sentinels in a 40k byte-level BPE vocabulary, fit on the target domain rather than inherited from a general-purpose tokenizer.

Muon on 2D parameters

Muon for matrices, AdamW for everything else, under a WSD schedule with short-sequence early phases. Throughput held at roughly 18k tokens/second.

Metallum-1B training corpus compositionA dimensioned bar showing the pretraining mix by effective tokens: Code 53.6%, Knowledge 19.6%, Reasoning 15.2%, Math 11.6%. Total 19.399B effective tokens from 8.473B unique tokens, seen over 16.0B tokens of training on 1 × NVIDIA RTX 5090.CORPUS SCHEDULE: BY EFFECTIVE TOKEN19.399B effective · 8.473B unique · 16.0B seen53.6%Code19.6%Knowledge15.2%Reasoning11.6%MathHARDWARERTX 5090WALL CLOCK~13.5 daysTHROUGHPUT~18k tok/sREPLAY~1.04B tokens, receipted
Curated and decontaminated in-house, not a general web scrape. ~1.04B tokens replayed after two hardware incidents, fully receipted: two hardware incidents mid-run, both accounted for rather than quietly re-rolled.

How we work

Measurement you can audit

Anyone can post a benchmark. The hard part is running an evaluation that could have embarrassed you, and then publishing it when it does.

01

Preregistered, hash-sealed

Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move the goalposts after seeing the result if the goalposts are hashed.

02

A holdout used exactly once

H9 was built, sealed, and never trained on, then evaluated one time with the checkpoint hash, evaluator hash, and tokenizer hash all recorded in the receipt. Whatever it returned was the number we would publish. It returned 1.3953.

03

Contamination audits that bite

Decontamination runs against the full eval surface. One Stack ingest dropped 18,564 documents, 4.40% of it, as benchmark clones. H9 itself was verified at 0.0099% shingle containment before it was allowed to count.

04

Failures kept in the record

A touched holdout was retired rather than reused. A cleanroom retrain that failed its gate is published as a failed experiment. A frozen classifier that counted mentions of "C++" as C++ training data was corrected against our own prior result.

What's your next batch?

Tell us the domain and what data you're sitting on, and we'll tell you what a specialist brewed on it could know. If it's the wrong answer for you, we will say so. That conversation is free and takes one email.