quipu-114m trained · 25–27 Sep 2026

Quipu

A small language model, trained from scratch on a single 8 GB laptop GPU.

Own weights, own data pipeline, no fine-tuned base. The goal is a small model that pulls well above its size by remembering instead of memorising.

The code mix, recorded the Incan way: one cord per language, clusters of knots for tens (upper) and units (lower). Java's cord reads 3 and 5, for 35%.

quipu-114m

A 114M-parameter decoder, sized for one laptop.

Every number here was measured on the machine that trained it, an RTX 5060 Laptop GPU with 8 GB.

114,114,048parameters
3Btraining tokens
18.4ktokens / second, bf16
45 h 22 mtraining time, zero restarts

Architecture

Decoder-only transformer in the modern style: RMSNorm, rotary position embeddings, SwiGLU feed-forward, and grouped-query attention with 12 query heads sharing 4 key/value heads. Input and output embeddings are tied, and the context is 1,024 tokens.

Short context by design. Long-range recall is meant to come from memory rather than a longer attention window.

Token embedding tied with output
RMSNorm → Attention GQA 12q / 4kv · RoPE
RMSNorm → Feed-forward SwiGLU
× N blocks
Final RMSNorm
Output head shares embedding weights

Data

Three billion tokens: mostly text, one-fifth code.

80% FineWeb-Edu, 20% code from github-code-clean. Only permissive licences are kept: MIT, Apache-2.0, BSD, ISC, CC0 and Unlicense. Minified files, vendored directories and files over 16k tokens are filtered out. That filter alone removed as many tokens as the entire code budget.

Educational web text · 80%Code · 20%

HTML is capped at 10% so markup doesn't crowd out the languages the model should actually learn to write.

Engineering

The failures that don't announce themselves.

Most of the work in a long unattended run is catching things that fail silently. Each of these was found in review, before the run.

01 · speed

Attention was quietly using a slow kernel.

Fixing the fallback took the projected run for the original budget from ~65 hours to ~39.

02 · I/O

About 1 in 130 file renames fails transiently.

That is enough to kill a 50-hour run over a log write. Every atomic write now retries.

03 · resume

A resume could restart from zero, silently.

After days of training, with no error. It now refuses unless nothing could have been saved.

04 · numerics

One bad batch could have deleted every clean checkpoint.

Non-finite steps are now skipped, and three in a row stop the run before retention prunes the good ones.

Results

It learned the shape of language in 45 hours.

All 5,722 steps ran unattended over a weekend: no restarts, no skipped steps, no data repeated. Validation loss on held-out FineWeb-Edu text, measured every 100 steps:

6.38 → 3.27 text validation loss. Most of the gain comes in the first 1,000 steps (0.5B tokens); after that it is slow and steady, and nearly flat by the last few hundred steps.

Watch it learn

Same prompt, greedy decoding, at each saved milestone.

verdict

Fluent, on topic, and often wrong.

The final model says the printing press was invented "in 1848 by the French printer Pierre-Louis Leclerc". That is the honest result for a 114M base model: it has learned the form of knowledge without reliable facts, and it still repeats itself under greedy decoding.

caveat · code

The code loss of 0.97 flatters it.

GPT-2's tokenizer splits indentation into single spaces: 36.5% of the code validation tokens are pure whitespace, against 2.0% for text. They are easy to predict, and greedy code output collapses into runs of spaces. That measurement is why the next model gets a code-aware tokenizer.

Weights, milestones and model card on Hugging Face →

Next

Where it goes from here.

  1. Pretraining at 114M · done3B tokens in 45 hours on one laptop GPU. Weights are public.
  2. Memory for a 1M-token effective context · nextA short attention window plus a retrievable map of everything before it, on the same 8 GB card.
  3. Instruction tuningSo it answers rather than continues.
  4. 350M with a code tokenizer and tool callingAimed at driving coding agents. The whitespace result above is the measured reason for the new tokenizer.