Skip to content
HomeBrewedLabsCommission

Metallum-550M

Open weights

The 550.3M-parameter first Metallum: an archival first-generation ML / LLM-engineering specialist trained from scratch on one consumer GPU.

  • text-generation
  • ml-engineering
  • 550M
  • trained-from-scratch
  • retrieval
  • custom-code
  • single-GPU
  • archival

What it is

Metallum-550M is the historical first-generation Metallum, developed as MetaLLM-V3. It is a decoder-only ML-engineering specialist pretrained from random initialization on one RTX 5090, then post-trained with retrieval SFT, capability and instruction SFT, and one bounded ReST-EM code round. It is released as an archival artifact, not as a replacement for Metallum-1B.

What it is for

Small-model and retrieval research, technical continuation inside ML and PyTorch, and reproducing the first Metallum generation. It is deliberately narrow. The useful question is what a 550M specialist can concentrate inside one domain, not whether it can impersonate a general assistant.

Architecture

28 layers at d1280, grouped-query attention with 20 query heads over 10 key-value heads, a 3456-wide SwiGLU, RMSNorm, per-head QK normalization, tied 32k embeddings, RoPE at theta 500k, and no positional encoding in every 4th layer. Native context is 2048 tokens.

Training data

Pretraining consumed 4.0B tokens. The stable 2.272B-token pack was mixed 59.5% code, 38.2% ml-arxiv, 2.4% synthetic technical text; the final tenth sampled from an 880.9M-token quality-decay pack. Pack counts include intentional resampling and are not unique-source counts.

How to consume it

Load it through Transformers with trust_remote_code=True after inspecting and pinning the repository revision. It has no chat template and expects plain-text continuation or instruction prompts. The archival shim has no KV cache and recomputes the prefix at every generation step, so it is slower than similarly sized cached models.

Relationship to Metallum-1B

Metallum-550M is an earlier, independently trained checkpoint. Metallum-1B has a larger corpus, different weights, a separate evaluation record, a sealed final holdout, and a constrained-decoding serving wrapper. The 550M release preserves the first generation rather than back-porting its successor's claims.

Evaluation

Every result below belongs to the exact released checkpoint. All of these surfaces were visible during development or checkpoint selection, and this lineage had no sealed final holdout. They are reproducibility records, not untouched estimates.

  • ML knowledge (cloze)

    41.6%
    Qwen2.5-1.5B
    28.0%
    SmolLM2-1.7B
    31.2%

    Length-normalized choice likelihood on the 250-question in-domain ML benchmark. The comparisons use the same narrow internal harness; they are not general model rankings. This suite steered development and checkpoint selection.

    Internal trajectorylogs/restem_r1_eval.log; eval/generalist_sweep.txt
  • Answer-letter binding (MCF)

    23.6%

    The same 250 questions scored through answer-letter output land near chance. The gap between cloze knowledge and letter emission is an elicitation limitation, and a dedicated format SFT did not close it without eroding cloze accuracy.

    Internal trajectorypaper/main.tex (V3 lineage table)
  • In-domain ML code

    10 / 40
    pass@1
    25.0%

    Canonical release-scorecard result on the executable 40-task suite. A legacy manual line in the run log records 8/40 from a separate greedy path; this page follows the named scorecard artifact used for the release.

    Internal trajectoryeval/scorecard_500m_v3_restem_r1.json
  • Long-context retrieval (RULER-style)

    88 / 90
    pass rate
    97.8%

    Synthetic retrieval at 1,024 and 2,048 tokens. The exact checkpoint misses two of ninety trials, both at 1,024; it clears every 2,048-token trial. The suite still counts as selection-aware because it informed development.

    Internal trajectoryeval/500m_v3_restem_r1_sc_ruler.json
  • Synthetic needle retrieval

    778 / 800
    Passkey
    499 / 500
    Key-value
    279 / 300

    A 97.3% aggregate over passkey and key-value retrieval from 512 through 2,048 tokens. Published as a reproducibility record, not an untouched final estimate.

    Internal trajectoryeval/500m_v3_restem_r1_sc_needle.json
  • General code (MBPP)

    0% pass@8

    The off-domain public coding benchmark stays at zero. That negative result is the cleanest warning against treating a narrow ML-engineering specialist as a general coding model.

    Internal trajectorypaper/main.tex; release/metallm-v3-550m/README.md
Internal trajectory
Used to steer training. Selection-aware or previously observed, so it cannot be read as a clean capability estimate.

Intended use

In scope

  • Studying small domain-specialist language models
  • ML and PyTorch technical continuation and scaffolding
  • Experiments with interleaved NoPE retrieval layers and QK normalization
  • Reproducing the first Metallum generation

Out of scope

  • General factual question answering
  • General-purpose coding
  • Autonomous code execution
  • Safety-critical decisions
  • Deployment as an aligned assistant

Limitations and bias

  • This is not a general chatbot and has no preference tuning, RLHF, or safety alignment.
  • Free generation can make local factual errors; verify technical claims before using them.
  • General coding is weak. The recorded MBPP result is 0% pass@8.
  • Knowledge under cloze scoring does not transfer reliably to answer-letter output.
  • The native context window is 2,048 tokens; the retrieval results do not establish behavior beyond it.
  • The inference shim has no KV cache, and padding-aware batched inference was not part of the release evaluation path.

The public release card includes the complete selection-aware evaluation note and training-data provenance on Hugging Face.

How it was measured

The release is the checkpoint, not a look-alike export

The source checkpoint contains 310 bfloat16 tensors and exactly 550,251,264 parameters. Every exported tensor matched its source key, shape, dtype, and value; both SHA-256 identities ship in PROVENANCE.json.

Native and Transformers logits agree exactly

The rebuilt compatibility shim was compared against the native model in bfloat16. Maximum and mean absolute logit difference were both 0.0, argmax outputs matched, and loading plus generation passed under Transformers 4.45 and 5.13.

Selection-aware, with no sealed final holdout

These suites were visible during development and informed checkpoint selection. Unlike Metallum-1B, this historical lineage did not receive a later sealed one-shot holdout, so every score is labeled internal even when the underlying task is synthetic or held out from gradient updates.

The negative results stay attached

General MBPP remains at 0% pass@8, answer-letter binding remains near chance, and additional distillation and larger-pool ReST-EM attempts stayed flat or eroded other capabilities. The archival release keeps those boundaries instead of rewriting the first generation through its successor's results.