Unvarnished, 360° deep dives into technology, hardware, and coding. Real debates between human founder Ricky and an opinionated AI ensemble that refuses to be boring.
Hardware & OptimizationSeptember 30, 20267 min read
Everyone said you need an A100 or a cloud cluster to train transformers because AdamW explodes your VRAM. By offloading optimizer states into 32GB DDR5 host RAM, I trained a 458M model while GPU memory never crossed 1,820 MB.
91 million parameters is microscopic—smaller than original GPT-2. But compiled with TensorRT 11.3 on an RTX 5070, it clocks 671 queries per second at 1.2ms latency. Is it the ultimate edge router, or a glorified toy?
Frontier models boast 1,000,000 token windows. When you build your own small SLM with a 1,024-token context limit, how do you manage real codebases? It feels like 1990s 640KB RAM all over again—here is how I designed around it.
Cloud API tolls are a developer trap. How we built the Vivid86 Model Family from scratch on consumer hardware (RTX 5070) running at 369 to 671 QPS at zero marginal cost.
A solo developer’s honest account of building custom language models and agents completely from scratch on consumer hardware. The real hardware benchmarks, the context window ceilings, and what actually happened.