Skip to content
HomeBrewedLabsCommission

Toolchain

The tools that made the model

Everything below is a real part of the stack that produced our model on one graphics card. Most of it is internal today, and it is labelled that way: a tools page is only useful if its download links work.

  • Internal

    Single-card training stack

    The pretraining loop that took a 1B model from random initialization to a finished checkpoint on one consumer GPU.

    Muon (2D) + AdamW under a WSD, short-sequence curriculum, with document-masked attention so a packed sequence never lets one document continue into an unrelated one. It held roughly ~18k tok/s on 1 × NVIDIA RTX 5090, finishing in ~13.5 days. The parts that mattered were not the fast parts. They were the ones that kept the run from dying with no cluster to restart from.

    • 1 GPU, 32 GB VRAM
    • ~18k tok/s sustained
    • Document-masked packing
    • QK-norm + logit z-loss
  • Ships with the model

    Domain tokenizer trainer

    Fits a byte-level BPE vocabulary on your corpus instead of inheriting one built for general web text.

    Metallum runs a custom 40k byte-level BPE with fill-in-the-middle sentinels, fit on the target domain. A general-purpose tokenizer shreds domain vocabulary into fragments: part numbers, chemical names, statute references, identifiers. You pay that tax on every token you ever process, at training time and at inference time forever.

    • 40k vocabulary
    • Byte-level BPE
    • FIM sentinels
    • Measured against your current tokenizer
  • Open release planned

    Decontamination and provenance pipeline

    Shingle-containment auditing of a training corpus against the evaluation surface, with a record of what it threw away.

    Curated and decontaminated in-house, not a general web scrape. The audit is the reason we trust our own table: one Stack ingest lost 18,564 documents, 4.40% of it, as benchmark clones, and the sealed holdout was verified at 0.0099% containment before it was allowed to count. The pipeline keeps provenance attached to what survives, which is what makes a dataset safe to train on twice.

    • Shingle containment scoring
    • Per-document provenance
    • PII screening on intake
    • Drop ledger, not just a pass/fail
  • Ships with the model

    Constrained decoder

    Schema-guaranteed JSON and tool calls from a small model, enforced during generation rather than hoped for afterwards.

    A token-level JSON grammar with schema-forced keys, applied at decode time. This exists because format could not be trained in: the base weights score 0/60 on strict whole-output format and three SFT recipes plus a ReST-EM screen all failed to move it. Constrained decoding reaches 60/60 on the same suite. It is the layer that makes a 1B specialist usable as an API instead of as a text generator.

    • Token-level grammar
    • Schema-forced keys
    • Tool-call and MCQ modes
    • No regex post-processing
  • Ships with the model

    Preregistered evaluation harness

    Hash-sealed evaluation designs, frozen batteries, and receipts that record which checkpoint and which evaluator produced a number.

    Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move a goalpost that has been hashed. Receipts record checkpoint hash, evaluator hash and tokenizer hash together, so a published number can be tied to the exact artifacts that produced it. Commissioned builds get the harness so you can re-verify every claim we make about your model.

    • SHA-256 sealed designs
    • Checkpoint + evaluator + tokenizer hashes
    • One-shot holdout custody
    • Failed experiments retained
  • Open release planned

    Long-context retrieval suites

    Per-seed generated needle-in-a-haystack tests, including distance probes at the edge of the context window.

    Synthetic retrieval suites generated per seed, so there is no fixed test set for a corpus to contaminate. That property is the whole point: a long-context score on a static public benchmark tells you very little once the benchmark is old enough to have leaked into training data everywhere.

    • Generated per seed
    • Uncontaminable by construction
    • Distance probes to window edge
    • RULER-style needle retrieval

Availability

What each label actually promises

The same discipline we apply to benchmark numbers applies to release claims. An intention is not a repository.

Ships with the model

Delivered as part of a commissioned build or a hosted deployment, including the harness so you can re-run our claims yourself.

Open release planned

Intended to be released openly. It is not published yet, so treat this as a statement of intent and nothing stronger.

Internal

Used in-house to build models. Not packaged for external use, and we would rather say so than publish a repository we do not maintain.

Want the toolchain pointed at your data?

A commissioned build ships with the tokenizer, the constrained decoder, and the evaluation harness, so you can re-run every claim we make about your model instead of taking our word for it.