Skip to content
HomeBrewedLabsCommission

Metallum-1B

Open weights

A 1,002M-parameter ML / LLM-engineering specialist, pretrained from random initialization on one consumer GPU.

  • text-generation
  • ml-engineering
  • 1B
  • trained-from-scratch
  • long-context
  • fill-in-the-middle
  • single-GPU

What it is

Metallum-1B is a decoder-only ML / LLM-engineering specialist pretrained from random initialization, not a fine-tune of somebody else's base. It saw 16.0B tokens of curated, decontaminated data over ~13.5 days of wall clock on 1 × NVIDIA RTX 5090 with 32 GB of VRAM. One consumer card. No cluster, no rented H100 fleet, no FSDP.

What it is for

Reading and writing ML and LLM-engineering material: PyTorch and training-loop scaffolding, concept explanation, long-context retrieval over technical documents, and structured output through the constrained decoder. It is a specialist. Inside that lane it punches above its parameter count; outside it, it is a 1B model and behaves like one.

Architecture

26 layers at d1792, GQA with 28 query heads over 14 key-value heads, 4864 feed-forward, tied embeddings, RoPE at theta 500k with no positional encoding at all on every 4th layer. Trained with Muon (2D) + AdamW under a WSD, short-sequence curriculum. QK-norm and a logit z-loss are what kept the run stable on a single card with no cluster to restart from.

Training data

19.399B effective tokens from 8.473B unique, mixed roughly 53.6% code, 19.6% knowledge, 15.2% reasoning, 11.6% math. Curated and decontaminated in-house, not a general web scrape. ~1.04B tokens replayed after two hardware incidents, fully receipted.

How to consume it

Through the serving wrapper, not raw. The base weights score 0/60 on strict whole-output JSON, tool-call and multiple-choice format, and three SFT recipes plus a ReST-EM screen all failed to move that. Structure is therefore enforced at decode time with a token-level JSON grammar and schema-forced keys. Point a client at the raw weights expecting JSON and you will get prose.

Evaluation

Every row carries the class of evidence it can bear. Two of them explicitly cannot bear a capability claim, and they are printed here anyway. That is the point of publishing a table instead of a headline.

  • H9 sealed holdoutlower is better

    1.3953

    Bits per byte on 400k characters of post-cutoff Wikipedia, built and sealed before the run and evaluated exactly once. Contamination verified at 0.0099% shingle containment. One shot, no retries. This is the number that carries the release.

    Sealed holdouteval/.h9_custody_v2/public_receipts/h9_onetime_evaluation_receipt_v1.json
  • In-domain ML-eng problems

    7 / 150
    Qwen3-1.7B
    7 / 150
    Qwen3-0.6B
    3 / 150
    SmolLM2-1.7B
    0 / 150

    A deliberately hardness-gated 150-task executable suite where small models sit near the floor by design. Metallum ties Qwen3-1.7B at 1.7× its parameter count and clears both other baselines. The suite was used for model selection during development, so it is selection-aware: a real result, not a clean one.

    Internal trajectoryeval/ineval_v1.jsonl
  • Long-context retrieval (RULER)

    86 / 90

    Synthetic needle-in-a-haystack, generated per-seed, so there is no training data to contaminate it.

    Held-out measurementeval/ruler_sniah.py
  • Distance needles

    18 / 18

    Perfect retrieval at both D=1024 and D=2032, the second of which is effectively the full 2,048-token window. The NoPE layers are why.

    Held-out measurementeval/kimi_triage/needle_distance.py
  • Held-out ML-arXiv BPClower is better

    0.7276

    Bits per character on a clean held-out arXiv surface. Lower is better.

    Held-out measurementscripts/scorecard.py
  • Structured output (via wrapper)

    60 / 60
    Raw base weights
    0 / 60

    The base weights score 0/60 on strict whole-output JSON/tool-call/MCQ format, and three SFT recipes plus a ReST-EM screen all failed to move it. So structure is enforced at decode time instead, with a token-level JSON grammar and schema-forced keys. Consume the raw weights without the wrapper and you get prose, not JSON.

    Held-out measurementeval/kimi_triage/serve_smoke/format_60_via_endpoint_serve_smoke_v3.json
  • In-domain code suite

    32 / 40

    Two exact repeats on an internal 40-task suite with known training overlap. We track it to steer training. It is not a capability claim and we will not present it as one.

    Not release evidenceeval/code_completion_suite_v2.json
Sealed holdout
Measured once, on a holdout that was built, sealed, and never trained on. One run, no retries, result published whatever it said. This is the generalization claim.
Internal trajectory
Used to steer training. Selection-aware or previously observed, so it cannot be read as a clean capability estimate.
Held-out measurement
Measured on surfaces the model never trained on. Strong capability evidence, repeatable on demand.
Not release evidence
The suite itself declares known overlap with training inputs. Published for completeness, never as a capability claim.

Intended use

In scope

  • PyTorch and training-loop scaffolding
  • ML and LLM concept explanation
  • Long-context retrieval over technical documents
  • Structured-output endpoints via the constrained decoder

Out of scope

  • General chat
  • Non-ML factual question answering
  • General-purpose coding
  • Safety-critical use
  • Autonomous code execution

Limitations and bias

  • No preference tuning, no RLHF, no safety alignment. Free generation makes local factual slips; verify specifics.
  • Not for general chat, non-ML factual question answering, general-purpose coding, safety-critical use, or autonomous code execution.
  • Structured output requires the constrained decoder. The raw weights do not produce reliable JSON.
  • No multiple-choice knowledge score from this lineage is valid: teacher generation prompts embedded real evaluation items, so the metric is permanently non-promotable. We report none.

The full set of results that do not favour us, with reasoning, is on the proof page.

How it was measured

Preregistered, hash-sealed

Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move the goalposts after seeing the result if the goalposts are hashed.

A holdout used exactly once

H9 was built, sealed, and never trained on, then evaluated one time with the checkpoint hash, evaluator hash, and tokenizer hash all recorded in the receipt. Whatever it returned was the number we would publish. It returned 1.3953.

Contamination audits that bite

Decontamination runs against the full eval surface. One Stack ingest dropped 18,564 documents, 4.40% of it, as benchmark clones. H9 itself was verified at 0.0099% shingle containment before it was allowed to count.

Failures kept in the record

A touched holdout was retired rather than reused. A cleanroom retrain that failed its gate is published as a failed experiment. A frozen classifier that counted mentions of "C++" as C++ training data was corrected against our own prior result.