Run Qwen3.8-27B Locally: Consumer GPU Guide

Qwen dropped Qwen3.8-27B today under Apache 2.0 — a dense 27B model with native vision, a 262K native context that extends to 1M, and a real jump over the previous generation on coding and office tasks. We covered the release announcement earlier. The full BF16 checkpoint is roughly 55 GB, but nobody runs that at home. Quantized, it fits a 16–24 GB card. Here is the VRAM math, then three ways to get it running tonight.

The VRAM math first

27B dense is the sweet spot: big enough to matter, small enough to quantize onto consumer silicon. Working numbers:

  • BF16 (full precision): ~54 GB VRAM → 80 GB class (H100, RTX Pro 6000)
  • FP8: ~27 GB → 48 GB class (L40S, RTX Pro 6000; RTX 5090 with short context)
  • 4-bit GGUF: ~14–16 GB → 24 GB class (RTX 4090/3090); Unsloth previews ~17 GB total for Q4 builds

Add KV cache on top of those numbers. Context length is the other lever: 94K context at 4-bit is realistic on 24 GB; keep it shorter on 16 GB. A solid community data point: UD-IQ3_XXS on an RTX 5060 Ti 16 GB hits 50–55 tokens/s with MTP (multi-token prediction) and a quantized KV cache at 94K context. A 16 GB card is genuinely enough.

Option 1: Ollama — fastest setup

ollama run qwen3.8:27b

That is it. Ollama picks a sensible default quant for your VRAM, so if you have 16 GB or less, start here. Apple Silicon users can pull qwen3.8:27b-mlx instead. Vision works the same way as any other Ollama model — pass an image path in the prompt.

Option 2: llama.cpp — full control

Download a GGUF and start the server:

hf download unsloth/Qwen3.8-27B-GGUF --local-dir Qwen3.8-27B-GGUF --include "UD-Q4_K_XL"

llama-server -m Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_XL.gguf \
  --ctx-size 8192 --host 127.0.0.1 --port 8080

llama-server exposes an OpenAI-compatible endpoint at http://127.0.0.1:8080/v1, so any tool that speaks the OpenAI API can point at it. Quick CLI check without a server:

llama-cli -m Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_XL.gguf \
  --ctx-size 8192 --temp 1.0 --top-p 0.95 --top-k 20

Option 3: LM Studio — GUI

The lmstudio-community/Qwen3.8-27B-GGUF repo ships ready-to-import quants. Search for Qwen3.8-27B inside LM Studio, pick a quant that fits your card, and use the built-in chat or the local OpenAI-compatible server.

Sampling settings that work

Qwen publishes two recommended presets. General use:

temperature=1.0, top_p=0.95, top_k=20, min_p=0.0
presence_penalty=0.0, repetition_penalty=1.0

Coding and structured tasks:

temperature=0.7, top_p=0.80, top_k=20, min_p=0.0
presence_penalty=1.5, repetition_penalty=1.0

The model also exposes reasoning_effort and preserve_thinking controls. Dial reasoning_effort down for speed; keep thinking on when you need reliable multi-step outputs.

Which quant should you pick

  • 16 GB: UD-IQ3_XXS or Q4_K_M, keep context moderate
  • 24 GB (RTX 3090/4090): Q4_K_XL or Q5, 94K context is comfortable
  • 32 GB (RTX 5090): Q6, or NVFP4 through SGLang/vLLM
  • CPU only: Q3/Q4 with 32 GB+ RAM — works, slow, last resort

Once the model is serving, the natural next step is wrapping it in an agent loop. Our DeepSeek Harness quickstart shows a model-agnostic harness that works with any OpenAI-compatible endpoint — local included.

Related News