Run QWEN3.8 27B on 16gb Nvidia GPUs
13 points by Pragmata 7 hours ago | 3 comments
  • Pragmata 7 hours ago |
    I've been using it for the past few days, and it runs really well!

    I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second.

    Genuinely very usable, and fully local!

    • kristianp 4 hours ago |
      Which card are you using? I was getting about 40 with an UD q3 quant with MTP (prediction) enabled and llama.cpp compiled for my compute capability, but was very limited in the context size. I have an 4060 ti 16GB. Wouldn't recommend it as there's a tradeoff between larger context without MTP and about 18 tokens/s.
    • falsaberN1 2 hours ago |
      With llama.cpp (CUDA) and a 5060ti (16GB) I get 60t/s with 128K token space. Odd you got 7t/s, did you verify all the model was loaded in VRAM? (--gpu-layers all)