Experiment archives behind the posts — raw runs, prompts, scores, and writeups kept whole, so a claim can be checked instead of taken on my word.
remax, remex, int8, PQ, and NVFP4 at matched byte budgets — full method grid, per-seed error bars, latency, and the NVFP4 CPU-reconstruction code. Cited in A 4-Bit Model and a 1-Bit Index.
Controlled best-vs-candidate run of the down-skilling v1.2.0 edit through the optimizing-skills validation gate, with every Haiku run logged. Cited in The validation gate would have rejected the fix that worked.
Eight task archetypes, vanilla Haiku against a down-skilled prompt, 20 runs per cell on the first three tasks and five on the rest — all prompts, outputs, and scores. Cited in When down-skilling makes Haiku worse.
Companion pages that are prose rather than archive — method writeups, data appendices — live in scratch. Older reference material for Oskar's own posts is in oaustegard/blog-references.