Writing

An AI Agent Overclocked My RTX 5090 Overnight

· AI · Local AI, AI Agents

My RTX 5090 lives in a headless Ubuntu box serving local language models. I gave Claude Code access over SSH and let it run an overclock and undervolt campaign overnight, seeking more inference throughput without leaving the machine unable to recover. The interesting result was not the final 4.2% gain, but the instability the agent found that ordinary benchmarks missed.

The safety model

The agent’s only standing permission was to apply GPU clock offsets, clock locks and the power limit through an NVML wrapper. Nothing else ran under sudo, and nothing was persisted. A reboot always returned the card to stock.

What each iteration measured

Each iteration stopped the model server and ran llama-bench on Qwen3.8-27B Q4_K_M. The dense model fits entirely in VRAM, making memory-bandwidth gains visible in decode speed. It then ran llama-perplexity against a stock reference to four decimal places, because an unstable GPU can produce slightly wrong numbers without crashing. Sensors were logged at 1 Hz and the kernel log was watched for Xid errors.

The overnight run

After supervised exploration, the agent soaked three candidate profiles for roughly ten and a half hours, safest first so a hard hang could not erase results already banked. The winner then completed a 3.4-hour soak containing 156 benchmarks and 52 perplexity checks of roughly 200,000 tokens each, with zero Xid errors.

ProfileSettingsMean powerDecodePrompt
Stocknone532 W82.5 t/s3968 t/s
Deployedmemory +2750 MHz, 600 W limit, core stock552 W86.0 t/s (+4.2%)3940 t/s
Efficiencycore locked 2400 MHz, +200 MHz curve offset398 W77.8 t/s3589 t/s

Measured hot, at the end of the long soaks, not in a fresh short bench.

On this driver, an NVML memory offset moves the clock by half its value as Afterburner or nvidia-smi report it, so the wrapper’s +5500 is +2750 in the usual units.

The undervolt trap

The core undervolts initially looked better than the memory tune, including an 8% prompt-processing gain in one-minute benchmarks. Yet every one but the Efficiency profile failed after the die had heat-soaked to 77 to 84°C, after nine minutes, ninety minutes or two hours. In the worst case, 26 clean runs came first. Nothing crashed and there were no Xid errors or visible artefacts. Only perplexity exposed the error, moving from 6.0222 at stock to between 6.0230 and 6.0241.

For compute, “it did not crash” is not a sufficient stability test.

One aggressive core profile eventually hard-hung the GPU. Driver reset, process termination, module unload and shutdown all failed, requiring a manual power cycle. The agent then kept a 300 MHz margin below the failure line. The durable gain came from memory, so the deployed profile leaves the core at stock.

The profile that boots every time

The next morning, I configured a small systemd service to reapply the tested profile on boot: memory at roughly +2750, a 600 W power limit and stock core. It has since applied on every boot and served models without issue.

Against my own manual tune

The machine also boots Windows, where I had already tuned the same card manually using ASUS GPU Tweak III.

StockManual, in ASUS GPU Tweak IIIAgent, over NVML on Linux
Core boost clock2407 MHz2656 MHz2407 MHz (stock)
Memory offset+0+2000 MHz+2750 MHz
Power limit575 W598 W600 W
Tunedfactoryinteractively, reboots on tapunattended, backed off after a crash

Memory and power are close. My Windows profile also boosts the core by roughly 249 MHz, while the agent rejected that class of tune after sustained, perplexity-checked inference. This does not prove the gaming profile unsafe. The workloads impose different standards, and the agent preferred margin over maximum performance.

What I would try with more time

I would next attach the computer to a remotely controlled smart plug and set the BIOS to boot when power returns. The controlling agent would run elsewhere, since an agent on a frozen machine cannot reset its own plug.

That setup could search for several days. A watchdog would detect a hang, cycle the power and resume from the last recorded result. Runtime-only GPU settings mean every reboot starts at stock, once the boot service is disabled for the campaign. I would still cap the number of resets, enforce cooldowns, keep a known-good fallback and forbid retrying the setting that caused the last hang. A smart plug makes a crash recoverable, not harmless.

What this proves

This is one card, one driver and one workload, not a tuning guide. The 4.2% gain is modest. More importantly, an agent ran a ten-hour hardware campaign unattended, caught a silent failure that ordinary stress tests missed, and produced a profile the machine can apply every day.

← All posts