Running AI on eBay Hardware: 200B Models on a $1,000 Box
A $1,000 workstation built from eBay parts (a 2016 dual Xeon with 40 cores, DDR4 RAM, and two $250 Quadro P5000 GPUs) runs Qwen3.5-122B at about 10 tokens per second, and 17 with multi-token prediction. The trick is llama.cpp's --cpu-moe, which keeps a mixture-of-experts model's expert weights in cheap system RAM and only the shared layers on the GPU. The limit is prefill: long prompts are slow on 2016 hardware, so it suits chat and agents, not thousand-page documents.
This box sits in the Joshua8.AI lab next to our Blackwell cards, and most of the posts below compare it against them on the same models.
The $1,000 box
| Part | What we used |
|---|---|
| CPUs | 2× Intel Xeon E5-2698 v4 (2016 Broadwell, 40 physical cores), about $100 for the pair |
| GPUs | 2× NVIDIA Quadro P5000 (Pascal, 16 GB each), about $250 apiece |
| RAM | 128 GB DDR4, later 192 GB with $50 of spare DIMMs |
| Software | llama.cpp with --cpu-moe expert offload |
Prices are what we paid, before the 2026 DRAM price spike. Rebuilding the same box now would cost closer to $1,400–1,800, mostly because of memory.
Key findings
- 122B at reading speed. Qwen3.5-122B runs at ~10 tok/s on the eBay box, against 128 tok/s on a $12.5K RTX PRO 6000 at the time. Slower, but faster than you read. 200B parameters for $1,000
- Multi-token prediction adds 52%. The flagship 122B went from 11 to 17 tok/s, but only on models whose drafts are usually accepted. Measure acceptance, not architecture. The $1,000 box, one month on
- Bigger quants can be slower. After a $50 RAM upgrade, four of five quantization upgrades made models slower, one by 23%. Prefill tracks bytes added; decode tracks expert bits-per-weight. The junk drawer upgrade
- Old GPUs pay the tax on prefill. On Qwen3.8-27B the Pascal box decodes at 23.4 tok/s against 57.3 on an RTX 5090, but newer hardware reads long prompts up to 36× faster. Qwen3.8-27B on three boxes
Build, tune, and measure
- 200 Billion Parameters for $1,000: Running 4-Bit Quants on eBay Hardware The build, the
--cpu-moesetup, and Qwen3.5-122B, MiniMax-M2.7, and DeepSeek-V4-Flash benchmarks against an RTX 5090 box. - The $1,000 Box, One Month On: A Model Zoo and Multi-Token Prediction 44 models including four 200B+ MoEs, and when multi-token prediction helps or hurts.
- The Junk Drawer Upgrade: What 64 GB of Forgotten DDR4 Taught Me About Quantization 128 GB to 192 GB, and why the obvious quant upgrade was wrong four times out of five.
- Qwen3.8-27B on Three Boxes: 2.5× Decode, 36× Prefill The eBay box against two RTX 5070 Tis and an RTX 5090 on the same dense model.
- Streaming the Experts: Is FreeToken Better Than llama.cpp's --cpu-moe? A newer expert-streaming engine on an RTX 5090, with the eBay box's llama.cpp results (13.2 tok/s, 17.8 with MTP) as the baseline.
Related
- The Apollonian Machine: What Nietzsche Would Say About the DRAM Crisis An essay on why memory bandwidth, not FLOPs, sets the limits this box runs into.
- Local AI Economics: Own the Hardware or Rent the Cloud? Break-even math and where cheap hardware fits.
- Running LLMs on NVIDIA Blackwell The other end of the lab: engines, NVFP4, and configs on current NVIDIA cards.
Want to run large models without buying a $12,500 GPU? We test these setups every week.
Talk to Us