What Metallum-1B taught us
Pretraining run · ReleasedFrom random initialization to a qualified, released 1B checkpoint on one consumer card. The full set of lab notes: what held the run stable, what generalized, and the experiments that refused to work. Findings below are the complete set from this run, each with the class of evidence it can bear.
- 01
Dropping positional encoding buys length generalization
The finding
Six of twenty-six layers, one in every four, carry no positional encoding at all. Those layers learn retrieval heads that do not depend on absolute position, and the model scores 18/18 on distance needles at D=2032, effectively the far edge of its own 2,048-token window.
Why it matters
Retrieval quality at the end of the context window is usually where small models fall apart. This is an architectural fix, not a data fix, so it costs nothing at training time and nothing at inference time.
Held-out measurementeval/kimi_triage/needle_distance.py - 02Went against us
Output format would not train in, so we stopped trying to train it
The finding
The base weights score 0/60 on strict whole-output JSON, tool-call and multiple-choice format. Three separate SFT recipes and a ReST-EM screen all failed to move that number. Enforcing structure at decode time instead, with a token-level JSON grammar and schema-forced keys, reaches 60/60.
Why it matters
At this scale, reliable structured output is a decoding problem, not a fine-tuning problem. Teams routinely burn weeks of SFT on this. The honest read is that the capability was never in the weights and did not need to be.
Held-out measurementeval/kimi_triage/serve_smoke/format_60_via_endpoint_serve_smoke_v3.json - 03
QK-norm and a z-loss are what make a single-card run survivable
The finding
Normalizing queries and keys before the attention product, plus a z-loss term on the logits, carried a 16.0B-token run to completion on one GPU. There was no cluster to restart from and no checkpoint fleet to fall back on, so divergence would have ended the project rather than delayed it.
Why it matters
Stability techniques are usually justified by throughput at scale. On one card they are justified by survival: the cost of a diverged run is the whole run.
Verified build factrelease/metallm-1b-s2b075-hf/config.json - 04
Decontamination has to be measured, not asserted
The finding
Running the decontamination pass against the full evaluation surface dropped 18,564 documents from one Stack ingest, 4.40% of it, as benchmark clones. The replacement sealed holdout was verified at 0.0099% shingle containment before it was allowed to count.
Why it matters
Nobody intends to train on their benchmarks. 4.40% of one ingest was contaminated anyway. If a project reports no contamination, the likely explanation is that it did not look.
Verified build factrelease/MODEL_CARD_V3_DRAFT.md - 05Went against us
A holdout that was touched is no longer a holdout
The finding
A recursive grep read protected holdout files. Only matching path names reached the agent transcript, with no payload text, and the holdout was retired anyway. Its replacement was rebuilt from Wikipedia edits made after training ended, under fresh custody, with every raw response retained and hashed and a negative-exposure ledger proving it was never scored while training was live.
Why it matters
The cheap move is to argue that nothing actually leaked and keep using the holdout. That argument cannot be verified by a reader, which is exactly why it is worthless. Retiring it was the only way the 1.3953 stays a generalization number.
Verified build facteval/.h9_custody_v2/public_receipts/h9_onetime_evaluation_receipt_v1.json - 06Went against us
A frozen audit can be wrong in your own favour
The finding
A frozen corpus classifier counted documents that merely mentioned "C++" as C++ training data, inflating what we believed the corpus contained. It was caught and corrected against our own earlier result rather than quietly superseded.
Why it matters
Freezing an audit stops you moving the goalposts, but it does not make the audit correct. A frozen measurement that flatters you is more dangerous than an unfrozen one, because the freeze reads as rigour.
Verified build factrelease/MODEL_CARD_V3_DRAFT.md - 07Went against us
Parity at 1.7× fewer parameters, and no more than parity
The finding
On a hardness-gated 150-task in-domain suite, Metallum solves 7 and Qwen3-1.7B also solves 7. It clears Qwen3-0.6B at 3 and SmolLM2-1.7B at 0. The suite steered model selection during development, so it is selection-aware.
Why it matters
Matching a model at 1.7× your parameter count inside your domain is the actual claim a specialist should make. Rounding a tie up to a win is how a lab spends its credibility on a sentence.
Internal trajectoryeval/ineval_v1.jsonl
Negative results
The ones that cost us
4 of the 7 findings above went against us. They are the reason to believe the rest of this page: anyone can publish the wins.
Metallum ties Qwen3-1.7B on our hard in-domain suite: it does not beat it. Both solve 7 of 150. Parity at 1.7× fewer parameters is the claim; superiority is not.
The 32/40 code result runs on a suite with known training overlap. Its own metadata forbids describing it as release-grade evidence.
The promoted checkpoint scores 4/5 on the internal neural-ops category against a preregistered floor of 5/5. The owner adjudicated that exception explicitly rather than moving the floor.
No multiple-choice knowledge score from this lineage is valid: teacher generation prompts embedded real evaluation items, so the whole metric is permanently non-promotable. We report none.
There is no preference tuning, no RLHF, and no safety alignment. Free generation makes local factual slips; verify specifics.
A 1B specialist is not a frontier model. It fits one domain, one budget, and one card, which is the entire proposition.
Method
How a finding earns the right to be published
None of the above would mean anything without the process that produced it. This is that process, and it is the part we would keep if we had to throw out everything else.
Preregistered, hash-sealed
Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move the goalposts after seeing the result if the goalposts are hashed.
A holdout used exactly once
H9 was built, sealed, and never trained on, then evaluated one time with the checkpoint hash, evaluator hash, and tokenizer hash all recorded in the receipt. Whatever it returned was the number we would publish. It returned 1.3953.
Contamination audits that bite
Decontamination runs against the full eval surface. One Stack ingest dropped 18,564 documents, 4.40% of it, as benchmark clones. H9 itself was verified at 0.0099% shingle containment before it was allowed to count.
Failures kept in the record
A touched holdout was retired rather than reused. A cleanroom retrain that failed its gate is published as a failed experiment. A frozen classifier that counted mentions of "C++" as C++ training data was corrected against our own prior result.
The full measured table, with every number grouped by how much weight it can bear, is on the proof page.
Build on this instead of repeating it
Every finding above was paid for once already. A model built on our base inherits the parts that were expensive to get right (the stability work, the retrieval behaviour, the decontamination pipeline) rather than rediscovering them on your budget.