Writing

Running a 14.8B Reasoning Model on One RTX 5090

· AI · Engineering Notes, Local AI, RTX 5090, Quantisation

Drafted February 2025. Finished and published August 2026 as part of the migration away from WordPress.

A 29.55 GB reasoning model ran entirely on one RTX 5090. DeepSeek-R1-Distill-Qwen-14B BF16 generated at 52.3 tokens per second. That was fast enough to use interactively.

The specification that made this possible was 32 GB of GDDR7 memory. Local models have hard memory thresholds. An extra few gigabytes can be the difference between a model fitting on the GPU and becoming much slower because part of it has to run elsewhere.

I loaded deepseek-ai/DeepSeek-R1-Distill-Qwen-14B, a 14.8-billion-parameter reasoning model that DeepSeek had released on 20 January. The DeepSeek-R1-Distill-Qwen-14B BF16 GGUF was almost an exact fit for the card.

The machine was already using 3,172 MiB for its desktop session and two existing compute services. The model run added a peak of 28,620 MiB, taking observed use to 31,792 MiB. It was close to the limit, but it worked. A model occupying almost 30 GB could load in 2.55 seconds and answer locally at a speed well beyond what I could read.

Then I made it smaller

Quantisation stores approximations of a model’s weights with fewer bits. The file gets smaller, VRAM use falls and inference can get faster because the GPU moves less data. The possible cost is a change in the model’s answers.

I converted one pinned revision to DeepSeek-R1-Distill-Qwen-14B BF16, then made DeepSeek-R1-Distill-Qwen-14B Q8_0, DeepSeek-R1-Distill-Qwen-14B Q6_K, DeepSeek-R1-Distill-Qwen-14B Q5_K_M, DeepSeek-R1-Distill-Qwen-14B Q4_K_M and DeepSeek-R1-Distill-Qwen-14B Q3_K_M files from it. Every version used the same llama.cpp build from 8 February 2025.

The comparison used twelve short prompts covering arithmetic, logic, coding and instruction following. Prompt text, context size, output limit, sampling settings and seed stayed fixed. I also ran five repetitions of llama-bench per file and sampled VRAM every 250 ms. The card stayed at stock settings with its default 575 W power limit.

Model and formatFile sizePeak VRAM above idleLoadPrompt tokens/sGenerated tokens/sTasks passed / 12Different from DeepSeek-R1-Distill-Qwen-14B BF16
DeepSeek-R1-Distill-Qwen-14B BF1629.55 GB28,620 MiB2.55 s2,67852.36Reference
DeepSeek-R1-Distill-Qwen-14B Q8_015.70 GB15,900 MiB1.51 s7,36185.2510 / 12
DeepSeek-R1-Distill-Qwen-14B Q6_K12.12 GB12,670 MiB1.26 s6,115103.5710 / 12
DeepSeek-R1-Distill-Qwen-14B Q5_K_M10.51 GB11,226 MiB1.00 s7,086110.0611 / 12
DeepSeek-R1-Distill-Qwen-14B Q4_K_M8.99 GB9,870 MiB1.01 s7,243122.3610 / 12
DeepSeek-R1-Distill-Qwen-14B Q3_K_M7.34 GB8,396 MiB0.75 s6,577122.3510 / 12

Different from DeepSeek-R1-Distill-Qwen-14B BF16 compares whitespace-normalised final-answer text. A different answer is not automatically a worse answer.

DeepSeek-R1-Distill-Qwen-14B Q8_0 nearly halved the file size and generated at 85.2 tokens per second. DeepSeek-R1-Distill-Qwen-14B Q4_K_M brought the file below 9 GB, used roughly a third of DeepSeek-R1-Distill-Qwen-14B BF16’s incremental VRAM and reached 122.3 tokens per second. That left enough memory for a much larger context, another model or other GPU work.

The smaller files did not get faster in a perfectly tidy order. Q8 had the highest prompt-processing rate, while Q4 and Q3 tied on generation speed. Even so, every quant was fast enough for an interactive local application.

Usable, with some sharp edges

The twelve prompts were too small to rank general intelligence, but they showed whether this was more than a model-loading exercise.

All six versions followed one strict instruction exactly:

{"status":"ready","retries":3}

They solved straightforward logic questions, identified Python’s shared mutable-default bug and returned several exact-format answers. They also exposed the sorts of failures an application would have to handle.

Every version calculated two arithmetic answers correctly but ignored the instruction to return only the amount or number. Knowing the answer was not enough when the caller required a precise format.

The biggest difference appeared on the three-box puzzle. DeepSeek-R1-Distill-Qwen-14B BF16 selected the box labelled MIXED and explained the deduction. Every quantised version spent all 1,024 output tokens inside an unclosed <think> trace and never produced a final answer. Parts of the reasoning were correct, but an application would still receive no usable result after it.

Two coding prompts had the same output-limit problem across every format. None reached executable final code within 1,024 tokens. That was partly a test of the model and partly a reminder that reasoning models can spend their entire allowance thinking.

The pass counts moved in both directions. DeepSeek-R1-Distill-Qwen-14B BF16 passed six tasks, DeepSeek-R1-Distill-Qwen-14B Q6_K passed seven and DeepSeek-R1-Distill-Qwen-14B Q3_K_M passed five. That does not mean DeepSeek-R1-Distill-Qwen-14B Q6_K was smarter than DeepSeek-R1-Distill-Qwen-14B BF16. With one seed and twelve prompts, it means only that lower precision did not produce a simple staircase of worsening answers.

I reran DeepSeek-R1-Distill-Qwen-14B BF16 at the end to check for heat or machine drift. All twelve answer strings matched the first run exactly, and both benchmark rates remained within 0.2%.

The part that felt new

The RTX 5090 did not make local inference equivalent to a hosted frontier model, and this test did not try to compare them. The striking part was having DeepSeek-R1-Distill-Qwen-14B BF16 sitting entirely inside a desktop PC and responding at more than 50 tokens per second.

Quantisation changed that from a model that barely fit into one that fit comfortably. DeepSeek-R1-Distill-Qwen-14B Q4_K_M generated at more than twice the DeepSeek-R1-Distill-Qwen-14B BF16 rate while leaving most of the card’s memory free. The outputs were capable enough to experiment with and fast enough to build around, provided I treated formatting, output limits and failed reasoning traces as real engineering problems.

That is where local AI feels useful rather than theoretical. There was no hosted API call in the loop, and I could swap model files and measure the consequences directly.

This remains one model, one build, one seed and twelve small tasks. A larger comparison would need more prompts, several seeds and longer coding limits. For a single test, though, the answer was clear: a serious reasoning model fit on one consumer GPU, and it was genuinely usable.

← All posts