It would be awesome if someone manages to figure out how to get small enough models to fit on older cards to be viable, but I'm not optimistic that it will come without some sort of fundamental architectural innovation rather than incremental improvements, and it's not clear if and when that will happen.
How long is this runway?
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
With AMD the best you can do in the consumer market right now is an RX 7900 XTX which is about 960 GB/s.
The other day I managed to get a context of 195k for Qwen3.5-9b Q4_K_M using a llama.cpp fork that supports TurboQuant:
https://github.com/TheTom/llama-cpp-turboquant
I think you could replicate this with a larger model on your device.
Overall with the right quantisations for both the model and KV cache you can get a lot of mileage out of this old hardware. Speed remains the main limitation, as IIRC I was getting ~26-30tps on a 7700S.
I'm not saying there's no benefit to using local models. My point is still the same as before: you have to be willing to sacrifice both performance and quality even when just comparing to free models that are available today.
No doubt Timur contributed heavily to this given Valves Steamdeck (which uses a very similar but slower GPU).
Given the current hardware prices it's pretty awesome to see someone squeezing maximum performance out of old hardware!
Originally ran a 3070 which matches your "older mid tier" description and then upgraded to a 9070XT maybe 6 months after launch so not so old or mid-tier.
I can attest to this myself, my 9070 was a bit rough on launch but worked perfectly soon after.
Not in my experience. The 16GB 9070 non-XT seems to have been released in March 2025. I purchased one in November 2025 [0]. I observed -and continue to observe- no significant difference in graphics performance between Linux and Windows. I've also not noticed any significant difference in performance with regards to raytracing, but that might be because I have the non-XT variant, and/or it might be because I run Gentoo Linux, which probably has much newer versions of both Mesa and the kernel than most Linux distributions. I CBA to test either of these theories.
[0] Mad props to Gamers Nexus for publishing the video where they speculated that the time around Thanksgiving was going to be the "low" point for graphics card prices... so if you were thinking of buying an upgrade that was THE time to do so. They were totally correct.
https://www.techspot.com/news/110999-new-benchmarks-show-lin...
With significant improvements still incoming.
https://www.gamingonlinux.com/2026/01/even-more-amd-ray-trac...
Unless you insist on running a server distro that's consistently obsolete by design (and if the notion of GPU comes into picture, that definitely shouldn't be the case), it shouldn't matter in practice. Every reputable desktop distro has a kernel/mesa stack that's up to date.
~3 year old hardware seems to be optimal for linux in my experience. Ideally a thinkpad or something lots of linux devs have.
Buy newer and you'll be constantly finding bugs and having to apply workarounds, custom kernel builds, custom config lines, etc. just to make basic features work like the ability to adjust screen brightness.
Don't buy anything ARM or requiring a custom bootloader - thats a constant battle with custom builds needed of nearly everything.
I run games both new and old and one of my favorite guilty pleasures is looking at the Steam forum for a four, eight, or ten year old game that I've started playing because it has recently become popular again and reading the complaints from Windows users about how a driver update screwed up the game... whether because of glitchy or incorrect graphics or unavoidable crashes. Meanwhile, I'm cruising along on Proton with zero issues. :smug-face:
[0] ...that small handful includes those that go out of their way to be incompatible with Proton...
https://windowsforum.com/news/nvidia-ends-feature-support-fo...
Compare it to some other products from that time:
- AMD FX Piledriver CPUs
- Skylake i7-6700K
- iPhone 6s
- Ubuntu 16.04 LTS
- Oculus Rift CV1
- Android 6 Marshmallow
The community can't fix it because their drivers are proprietary blobs. Open source drivers are essentially useless. I'll have to stick to something like Ubuntu to keep the hardware working without fighting DKMS every update.
This is the reason my next GPU will be either AMD or Intel.
That "partly" was worth hundreds of billions of dollars but AMD were cheap/shortsighted enough to hire a few dedicated engineers.
For various reasons, people flocked to CUDA and other proprietary crap. Then, when that proprietary crap became a money maker, people blamed AMD for not supporting their favourite proprietary crap.
AMD did the right thing and was punished for it. They're still doing the right thing and people still complain that their free full re-implementation of Nvidia's runtime, designed for completely different hardware, isn't good enough.
AMD is not a saint of a company but it's consistently better than Nvidia, every time. Doing the right thing just isn't rewarded and it shows.
Nvidia doesn't, and their Vulkan stack underperforms on Linux quite significantly.
It's just that Nvidia doesn't care much about Linux, and -as always- Nvidia ignores what everyone else is doing and does their own thing. Sometimes doing their own thing works out really well in the short run, but -long term- they always fall behind.
Nvidia cares a lot about Linux. Just... on their terms.
Like most FOSS done by companies, what they don't care is GPL, which sadly will eventually be meaningless after boomer and Gen X devs that created the GNU movement no longer walk this plane.
There isn't a single embedded OS FOSS alternative to GNU/Linux that uses GPL, including Linux Foundation's own Zephyr.
In general, Valve's work has been exceptional and supplementing AMD's own ROCM/OpenCL and Vulkan team, they've gotten a lot of defaults right and people should work together with Valve to improve support for their chips.
Use as a dedicated GPU for encoding and decoding video. Post processing like frame interpolation or superresolution. Use for GPGPU workloads. Run additional monitors independently. Use for GPU passthrough to virtual machines. Use as a backup GPU for troubleshooting. Use for test code without breaking the main GPU
Is Valve funding this work in particular? How can people donate?
+1 to buying Valve's products.
Perhaps we will even be able to reverse firmware blobs into open source alternatives?
The most difficult part is likely that a lot of hardware is easily brickable if you do the wrong thing, so I'm not going to let it loose on e.g. my solar inverter.. even though their software is shit and I'd love to replace it. (Don't buy QCells.)
This is very risky: one guy QA over a massive amount of software with hardly a few remaining users able to report issues of with those GPUs...
Better have a slow working well tested driver, rather than a broken driver which is supposed to be faster...
Try to remember back to 2015 the graphics were not that bad. 4k performance was poor but for 1080p 60fps on pretty much anything.