Toolchain
The tools that made the model
Everything below is a real part of the stack that produced our model on one graphics card. Most of it is internal today, and it is labelled that way: a tools page is only useful if its download links work.
- Internal
Single-card training stack
The pretraining loop that took a 1B model from random initialization to a finished checkpoint on one consumer GPU.
Muon (2D) + AdamW under a WSD, short-sequence curriculum, with document-masked attention so a packed sequence never lets one document continue into an unrelated one. It held roughly ~18k tok/s on 1 × NVIDIA RTX 5090, finishing in ~13.5 days. The parts that mattered were not the fast parts. They were the ones that kept the run from dying with no cluster to restart from.
- 1 GPU, 32 GB VRAM
- ~18k tok/s sustained
- Document-masked packing
- QK-norm + logit z-loss
- Ships with the model
Domain tokenizer trainer
Fits a byte-level BPE vocabulary on your corpus instead of inheriting one built for general web text.
Metallum runs a custom 40k byte-level BPE with fill-in-the-middle sentinels, fit on the target domain. A general-purpose tokenizer shreds domain vocabulary into fragments: part numbers, chemical names, statute references, identifiers. You pay that tax on every token you ever process, at training time and at inference time forever.
- 40k vocabulary
- Byte-level BPE
- FIM sentinels
- Measured against your current tokenizer
- Open release planned
Decontamination and provenance pipeline
Shingle-containment auditing of a training corpus against the evaluation surface, with a record of what it threw away.
Curated and decontaminated in-house, not a general web scrape. The audit is the reason we trust our own table: one Stack ingest lost 18,564 documents, 4.40% of it, as benchmark clones, and the sealed holdout was verified at 0.0099% containment before it was allowed to count. The pipeline keeps provenance attached to what survives, which is what makes a dataset safe to train on twice.
- Shingle containment scoring
- Per-document provenance
- PII screening on intake
- Drop ledger, not just a pass/fail
- Ships with the model
Constrained decoder
Schema-guaranteed JSON and tool calls from a small model, enforced during generation rather than hoped for afterwards.
A token-level JSON grammar with schema-forced keys, applied at decode time. This exists because format could not be trained in: the base weights score 0/60 on strict whole-output format and three SFT recipes plus a ReST-EM screen all failed to move it. Constrained decoding reaches 60/60 on the same suite. It is the layer that makes a 1B specialist usable as an API instead of as a text generator.
- Token-level grammar
- Schema-forced keys
- Tool-call and MCQ modes
- No regex post-processing
- Ships with the model
Preregistered evaluation harness
Hash-sealed evaluation designs, frozen batteries, and receipts that record which checkpoint and which evaluator produced a number.
Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move a goalpost that has been hashed. Receipts record checkpoint hash, evaluator hash and tokenizer hash together, so a published number can be tied to the exact artifacts that produced it. Commissioned builds get the harness so you can re-verify every claim we make about your model.
- SHA-256 sealed designs
- Checkpoint + evaluator + tokenizer hashes
- One-shot holdout custody
- Failed experiments retained
- Open release planned
Long-context retrieval suites
Per-seed generated needle-in-a-haystack tests, including distance probes at the edge of the context window.
Synthetic retrieval suites generated per seed, so there is no fixed test set for a corpus to contaminate. That property is the whole point: a long-context score on a static public benchmark tells you very little once the benchmark is old enough to have leaked into training data everywhere.
- Generated per seed
- Uncontaminable by construction
- Distance probes to window edge
- RULER-style needle retrieval
Availability
What each label actually promises
The same discipline we apply to benchmark numbers applies to release claims. An intention is not a repository.
Delivered as part of a commissioned build or a hosted deployment, including the harness so you can re-run our claims yourself.
Intended to be released openly. It is not published yet, so treat this as a statement of intent and nothing stronger.
Used in-house to build models. Not packaged for external use, and we would rather say so than publish a repository we do not maintain.
Want the toolchain pointed at your data?
A commissioned build ships with the tokenizer, the constrained decoder, and the evaluation harness, so you can re-run every claim we make about your model instead of taking our word for it.