
Follows the $1,000 eBay series — 200 Billion Parameters for $1,000, The $1,000 Box, One Month On, The Junk Drawer Upgrade — and the two 5070 Ti writeup.
TL;DR
Qwen3.8-27B was getting all the hype this weekend — same as Meta’s Glimmer last weekend, and DeepSeek-V4-Flash the weekend before that — so I ran it on the three lab boxes. Those are the Xeon 2x P5000 (Pascal) — the ~$1,000 eBay special — Ultra9 with two RTX 5070 Tis, and the Ryzen box with one RTX 5090. Thirty-two gigs of VRAM on each, just arranged differently: two 180 W 16 GB Pascal GPUs, two 300W 16 GB Consumer Blackwell GPUs, and one 600W 32 GB Blackwell GPU.
If you only look at generation, the upgrade is boring. Pascal does 23.4 tok/s, the 5070 Tis do 51.8, the 5090 does 57.3. Call it 2×, maybe 2.5×. All three are faster than most people read. For “Who caught the Immaculate Reception?” I’d still tell someone to buy the eBay box. That’s the equivalent of walking over to the trailer to pull one page from the drawing set.
Prefill is what actually changed. Unique ~16k-token prompts (so the prefix cache couldn’t cheat) went at 257 tok/s on the Xeon 2x P5000, 2,185 on the 5070 Tis, and 9,204 on the 5090 — about 36 times the Pascal pair.
That’s the number TeraContext cares about. Construction specs are 500 to 2,000 pages, the model sees 100k+ tokens at a time, and a real job is thousands of those calls — takeoff on the whole project manual, not one detail. At 257 tok/s, one 100k window is six and a half minutes and a thousand of them is four and a half days. The 5090 does the thousand in about three hours. I like the cheap box. We just don’t have four days.
The three boxes, one model
The last few posts were about getting 200B-class MoEs running at all. --cpu-moe, pinning 108 GB of experts in RAM, the junk-drawer DDR4 — that was “can I run this.” This one is just: the internet spent the weekend on Qwen3.8, so I put the same weights on the three boxes and timed them.
- Xeon 2x P5000 (Pascal). Dual E5-2698 v4, two $250 Quadro P5000s from 2016 (16 GB each, 32 GB together), the ~$1,000 build. It’s stuck on a CUDA 12.8
llama.cppimage because CUDA 13 dropped Pascal, which means it will never run the vLLM we use on the Blackwell boxes. The weights here are an INT4 GGUF — only Blackwell has native NVFP4, so Pascal doesn’t get the Unsloth NVFP4 file the other two boxes are running. MTP is on. That’s why a 27B here is at 23 tok/s instead of the ~12 we saw in July before MTP. - Ultra9, two RTX 5070 Tis. The consumer desktop from the boardwalk posts. 16 GB per card, 32 GB together, no NVLink, vLLM tensor-parallel, same Unsloth NVFP4 weights. MTP is off.
- Ryzen, one RTX 5090. The ~$5,000 box from the first post. One 32 GB card, vLLM 0.27.1, 180k context, vision left on because that’s how we actually serve it. MTP is off here too.
The RTX boxes aren’t leaving MTP on the table for fun. Vision, MTP, and a long context window don’t all fit in 32 GB. We kept vision and the long window, so MTP had to go. The Xeon isn’t doing vision, so it can spend that VRAM on MTP. Fair to know when you look at the 23 vs 52 vs 57 decode numbers: the cheap box is the only one drafting.
One request at a time, thinking off. Decode is 400 tokens streamed with ignore_eos, timed from the first output token. Prefill used a fresh random-word prompt each time, because the first C=1 run this week looked amazing until I realized I was measuring prefix cache. Cached prefill comes back in the 8k–22k tok/s range and tells you nothing about a new document.
| Card set | Decode (tok/s) | Unique prefill ~16k (tok/s) |
|---|---|---|
| Xeon 2x P5000 (Pascal) | 23.4 | 257 |
| Ultra9 2× RTX 5070 Ti | 51.8 | 2,185 |
| Ryzen RTX 5090 | 57.3 | 9,204 |
Decode is 2.2× and 2.5×. Prefill is 8.5× and 36×. Generation is mostly memory bandwidth; prefill is FLOPS and tensor cores, and a 2016 P5000 has basically none of the latter. Seeing it on one model just makes the gap harder to talk around.
“Who caught the Immaculate Reception?”
A normal chat turn is almost all decode. Twenty tokens of question, maybe 150 tokens of answer.
On the Xeon 2x P5000 the prompt is gone in a tenth of a second. The answer is 150 ÷ 23.4, about 6.4 seconds. Same answer on the 5090 is 150 ÷ 57.3, about 2.6 seconds. You can tell them apart if you’re watching the cursor. If you’re just trying to get an answer, both are faster than I read.
That’s the walk to the trailer. You need one detail out of the project manual, you grab the binder, you flip to it, you go back to work. It is not even worth booting the computer for that one page. For a coding assistant, or to ask who caught the Immaculate Reception, the original claim still holds: the $1,000 eBay special is a lot of model for the money. Spending five times as much to talk 2.5× faster is a luxury, not a requirement.
A thousand-page spec is a different job
TeraContext is the thing we actually have to ship. A commercial RFP shows up as a 500–2,000 page specification book and a drawing set. The software classifies pages against a work breakdown structure, bundles them into trade packages, and has to point at the paragraph it used. That is not a trip to the trailer for one detail. That is takeoff on the whole project manual — usually 100,000+ tokens a call, thousands of times on one job — because a spec is a few thousand sections plus addenda plus the “does this also apply to the electrical package” checks.
Almost all of the wait is prefill. The classification that comes out the other side is short. Using the unique-prefill rates above, one 100,000-token call:
Xeon 2x P5000: 100,000 ÷ 257 = 389 seconds = 6 minutes 29 seconds 2× RTX 5070 Ti: 100,000 ÷ 2,185 = 46 seconds RTX 5090: 100,000 ÷ 9,204 = 11 seconds
The generated tokens on that call barely show up. A hundred tokens of “cast-in-place concrete, WBS 3.2” is four seconds on Pascal and under two on the 5090. Something like 95% of the turn is just eating the window.
A thousand calls is a light day on a real set. Some jobs are more.
| One 100k call | 1,000 calls | 2,000 calls | |
|---|---|---|---|
| Xeon 2x P5000 (Pascal) | 6m 29s | 4 days 12 hours | 9 days |
| Ultra9 2× RTX 5070 Ti | 46s | 12.7 hours | 25 hours |
| Ryzen RTX 5090 | 11s | 3.0 hours | 6.1 hours |
I sat with that 4-day number for a minute because it sounds like I dropped a zero. 389 seconds times 1,000 is 108 hours. That’s four and a half days of the dual-Xeon box doing nothing but ingest. The 5090 is done in three hours. The 5070 Tis take a long workday, or two if you run 2,000 calls.
You can move one load of spoils with a wheelbarrow. You cannot clean a basement excavation that way and still make the pour. Estimators do not have from Friday’s addendum until the following Tuesday for the GPUs to finish reading.
Two things make this slightly less ugly, and not enough of either. The 100k times assume the 16k rate holds. It doesn’t — longer prompts get slower. At 180k we measured 1,415 tok/s on the 5070 Tis and 3,268 on the 5090. I didn’t run 180k unique on Pascal this time; at the 32k rate of 226 tok/s that window is already thirteen minutes. Prefix cache helps when the next call starts the same way, but TeraContext’s windows walk through the book. They overlap some. Not enough. You still pay for most of the page.
What I think the extra money is for
People ask if the 5090 is “worth it” next to the eBay box. For this model I think that’s the wrong framing. You’re not buying a faster walk to the trailer. You’re buying the difference between flipping one page and taking off the book.
The $1,000 machine is still the right buy if you want to talk to a 27B, or keep running the 200B MoEs that don’t fit anywhere else. MTP already doubled decode on the 27B in July, and that’s only on this box — the RTX machines can’t stack it on top of vision and 180k. I would not spend $4,000 to shave four seconds off “who caught the Immaculate Reception.” The wheelbarrow is the right tool for one load.
The ~$5,000 Ryzen box is what you buy if the prompt is the work. Same 27B, only a little quicker on the way out, 36× quicker on the way in. The 5090 isn’t even that far ahead of the two 5070 Tis on decode (57 vs 52 tok/s), but it’s about 4× on a 16k unique prefill and still 2.3× at 180k. For long documents I’d take the one 32 GB card over the pair of 16 GB cards, and either over the P5000s. There isn’t a llama.cpp flag that closes a 36× compute gap.
If you already have the 5070 Ti desktop and your prompts are tens of thousands of tokens, use it. I wouldn’t buy it for chat, and I wouldn’t make it the only box for a 180k construction window if the 5090 is sitting there.
This doesn’t undo the first post. The Xeon 2x P5000 is still why I can run DeepSeek-V4-Flash at all. That’s a different weekend’s hype and a different reason to own the box. This weekend the question was just Qwen3.8 on the hardware we already have. When they all run the same weights, prefill is the number I actually care about.
So
If you want a local chatbot, get the eBay special. Walk to the trailer, pull the page, go back to work. It will tell you Franco Harris caught the Immaculate Reception in Three Rivers Stadium in 1972 and you can feel clever about the receipt.
If you want takeoff on Division 03 a thousand times before Friday’s bid, it will still be on page 400 on Tuesday. The 5090 is done Monday afternoon. We don’t have until Tuesday.