#gpu
2 articles
-
Getting Qwen3.8-27B onto a 16 GB RTX 5080
The guides said a 16 GB card gets you 8K to 16K of context, or 4-bit weights spilling into system RAM. I ended up at 90,112 tokens at 130 tok/s with all of Qwen3.8-27B on one RTX 5080, after four settings that each cost me an evening.
-
I blamed the KV cache, then found the same bugs without it
I got a 27B model running on a 16 GB card, asked it to draw a pelican and write Flappy Bird, and thought I'd worked out what the cheap settings were costing me. Then I checked.