A small language model, trained from scratch on a single 8 GB laptop GPU.
Own weights, own data pipeline, no fine-tuned base. The goal is a small model that pulls well above its size by remembering instead of memorising.
quipu-114m
Every number here was measured on the machine that trained it, an RTX 5060 Laptop GPU with 8 GB.
Decoder-only transformer in the modern style: RMSNorm, rotary position embeddings, SwiGLU feed-forward, and grouped-query attention with 12 query heads sharing 4 key/value heads. Input and output embeddings are tied, and the context is 1,024 tokens.
Short context by design. Long-range recall is meant to come from memory rather than a longer attention window.
Data
80% FineWeb-Edu, 20% code from github-code-clean. Only permissive licences are kept: MIT, Apache-2.0, BSD, ISC, CC0 and Unlicense. Minified files, vendored directories and files over 16k tokens are filtered out. That filter alone removed as many tokens as the entire code budget.
HTML is capped at 10% so markup doesn't crowd out the languages the model should actually learn to write.
Engineering
Most of the work in a long unattended run is catching things that fail silently. Each of these was found in review, before the run.
Fixing the fallback took the projected run for the original budget from ~65 hours to ~39.
That is enough to kill a 50-hour run over a log write. Every atomic write now retries.
After days of training, with no error. It now refuses unless nothing could have been saved.
Non-finite steps are now skipped, and three in a row stop the run before retention prunes the good ones.
Results
All 5,722 steps ran unattended over a weekend: no restarts, no skipped steps, no data repeated. Validation loss on held-out FineWeb-Edu text, measured every 100 steps:
Same prompt, greedy decoding, at each saved milestone.
The final model says the printing press was invented "in 1848 by the French printer Pierre-Louis Leclerc". That is the honest result for a 114M base model: it has learned the form of knowledge without reliable facts, and it still repeats itself under greedy decoding.
GPT-2's tokenizer splits indentation into single spaces: 36.5% of the code validation tokens are pure whitespace, against 2.0% for text. They are easy to predict, and greedy code output collapses into runs of spaces. That measurement is why the next model gets a code-aware tokenizer.
Next