Running AI on eBay Hardware: 200B Models on a $1,000 Box

A $1,000 workstation built from eBay parts (a 2016 dual Xeon with 40 cores, DDR4 RAM, and two $250 Quadro P5000 GPUs) runs Qwen3.5-122B at about 10 tokens per second, and 17 with multi-token prediction. The trick is llama.cpp's --cpu-moe, which keeps a mixture-of-experts model's expert weights in cheap system RAM and only the shared layers on the GPU. The limit is prefill: long prompts are slow on 2016 hardware, so it suits chat and agents, not thousand-page documents.

This box sits in the Joshua8.AI lab next to our Blackwell cards, and most of the posts below compare it against them on the same models.

The $1,000 box

PartWhat we used
CPUs2× Intel Xeon E5-2698 v4 (2016 Broadwell, 40 physical cores), about $100 for the pair
GPUs2× NVIDIA Quadro P5000 (Pascal, 16 GB each), about $250 apiece
RAM128 GB DDR4, later 192 GB with $50 of spare DIMMs
Softwarellama.cpp with --cpu-moe expert offload

Prices are what we paid, before the 2026 DRAM price spike. Rebuilding the same box now would cost closer to $1,400–1,800, mostly because of memory.

Key findings

Build, tune, and measure

Related

Want to run large models without buying a $12,500 GPU? We test these setups every week.

Talk to Us