How to Run Qwen QwQ-32B Locally with Ollama

Qwen QwQ-32B is Alibaba’s open-weights reasoning model — a 32-billion-parameter model that matches or exceeds the performance of much larger models on mathematics, coding, and logical reasoning benchmarks by using extended chain-of-thought reasoning at inference time. Running it locally with Ollama takes about five minutes and gives you a capable reasoning model that runs entirely on your own hardware with no API costs and no data leaving your machine. This guide covers hardware requirements, the setup steps, how to use the model effectively, and what to expect from its performance.

What QwQ-32B Is

QwQ-32B stands for “Qwen with Questions” — a naming choice that reflects the model’s self-questioning reasoning style. It is a dense 32B parameter transformer trained with reinforcement learning on verifiable reasoning tasks, similar in training methodology to DeepSeek R1 and OpenAI’s o-series models. During inference, QwQ-32B generates an extended thinking trace before producing its final answer — working through problems step by step, checking its work, and sometimes reconsidering its initial approach. On AIME 2024, QwQ-32B scores approximately 50–65%, competitive with much larger models. On MATH-500, it scores above 90%. For a model that fits on consumer-grade hardware, these are compelling numbers.

Hardware Requirements

QwQ-32B at 4-bit quantisation (Q4_K_M, the default Ollama quantisation) requires approximately 20–22GB of VRAM. This means the model fits comfortably on a single RTX 3090 (24GB), RTX 4090 (24GB), or A6000 (48GB). It will not run on consumer cards with 16GB or less without significant quality degradation from more aggressive quantisation. If you do not have a GPU with sufficient VRAM, you can run the model in CPU-offload mode (some layers on GPU, some on CPU RAM), but inference speed drops to 1–3 tokens per second — functional for single queries but not for interactive use. A Mac with an M2 Pro/Max/Ultra or M3 with 36GB+ unified memory runs QwQ-32B well; Apple Silicon’s unified memory architecture means VRAM and system RAM are the same pool, so a Mac with 48GB unified memory handles the 4-bit model comfortably.

Installing Ollama

Ollama is a local model server that handles model downloading, quantisation, and serving via a simple API. Install it with a single command on macOS and Linux:

curl -fsSL https://ollama.ai/install.sh | sh

On Windows, download the installer from ollama.ai. After installation, verify it is running:

ollama --version

Ollama runs as a background service and exposes a REST API on localhost:11434 by default.

Downloading and Running QwQ-32B

Pull the model from the Ollama library — this downloads the quantised weights, approximately 20GB:

ollama pull qwq

The download takes 10–30 minutes depending on your connection speed. Once complete, start an interactive session:

ollama run qwq

You will see a prompt where you can type directly. QwQ-32B will show its thinking process before answering — do not be surprised by lengthy reasoning traces on hard problems. On an RTX 4090, expect approximately 15–25 tokens per second during the thinking phase.

Figure 1 — QwQ-32B hardware requirements at different quantisation levels

Quant VRAM needed Compatible GPU Quality Q8_0 (8-bit)~34 GBA100 80GB, 2× 24GBBest Q4_K_M (4-bit) ✓~21 GBRTX 3090 / 4090, M2 MaxVery good Q3_K_M (3-bit)~16 GBRTX 4080, M1 Pro 16GBGood Q2_K (2-bit)~12 GBRTX 3080, M1 ProDegraded

Using QwQ-32B via the API

Ollama exposes an OpenAI-compatible API, so any code that uses the OpenAI SDK can be pointed at the local Ollama server with a one-line change. Using Python with the openai library:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # required but not checked by Ollama
)

response = client.chat.completions.create(
    model="qwq",
    messages=[{"role": "user", "content": "Solve: find all integer solutions to x² - 5y² = 4"}],
)
print(response.choices[0].message.content)

The base_url points to Ollama’s local server and api_key is required by the SDK but ignored by Ollama. This pattern works with any framework that accepts an OpenAI-compatible endpoint — LangChain, LlamaIndex, and most agent frameworks support it.

Streaming Responses and Handling the Thinking Trace

QwQ-32B’s responses contain a thinking trace followed by the final answer. The thinking trace is enclosed in <think>...</think> tags. If you want to display only the final answer, strip these tags from the output:

import re

def extract_answer(response_text):
    # Remove the thinking trace
    clean = re.sub(r'<think>.*?</think>', '', response_text, flags=re.DOTALL)
    return clean.strip()

answer = extract_answer(response.choices[0].message.content)

For streaming responses, buffer the output and apply the same regex to the accumulated text once the stream completes. Streaming is useful for interactive CLI tools where you want to display the thinking process in real time — it shows the model working through the problem, which can be informative for debugging or educational use.

Performance Tuning

Several Ollama settings affect QwQ-32B performance. The num_ctx parameter sets the context window size — the default is 2048 tokens, which is too small for long reasoning traces. Increase it to 16384 or 32768:

ollama run qwq --num-ctx 16384

The num_gpu parameter controls GPU layer offloading. Set it to the maximum your VRAM allows:

OLLAMA_NUM_GPU=99 ollama run qwq  # offload as many layers as possible

On systems with multiple GPUs, Ollama automatically distributes layers across available GPUs. Monitor GPU utilisation with nvidia-smi during inference — if GPU utilisation drops below 90%, a bottleneck is likely in CPU processing or memory bandwidth rather than compute.

QwQ-32B vs QwQ Alternatives

The Ollama library includes several QwQ variants. qwq:latest defaults to Q4_K_M. qwq:32b-preview-q8_0 uses 8-bit quantisation for higher quality if you have the VRAM. For systems without sufficient GPU memory, qwq:32b-preview-q2_K reduces requirements to around 12GB at the cost of noticeable reasoning quality degradation. If 32B is too large, Qwen2.5-7B-Instruct is a substantially smaller non-reasoning model from the same family — it fits on 6–8GB VRAM and performs well on standard tasks, though without QwQ-32B’s extended reasoning capability. For reasoning-specific workloads where hardware is a constraint, DeepSeek-R1-Distill-Qwen-14B is also available via Ollama and offers competitive reasoning performance at 14B scale.

What QwQ-32B Is Good For Locally

Running QwQ-32B locally is most valuable for privacy-sensitive reasoning tasks (legal analysis, medical information, financial modelling where you do not want data leaving your machine), offline environments without reliable internet access, high-volume applications where API costs would be prohibitive, and development and testing of reasoning-model applications without incurring per-token API charges. It is genuinely capable at competition mathematics, algorithm design problems, logical reasoning puzzles, and code review tasks. It is slower than cloud APIs at the same quality level and requires hardware investment, but for the right use case the combination of privacy, cost, and capability makes it the most practical local reasoning model available as of 2026.

Running QwQ-32B locally with Ollama is one of the most accessible paths to a capable local reasoning model. The setup is five commands, the performance is competitive with cloud APIs for hard reasoning tasks, and the hardware requirement — a 24GB GPU or Mac with 36GB+ unified memory — is achievable for serious developers and researchers. The <think> trace that QwQ-32B produces is also genuinely interesting to read: it shows the model’s reasoning process transparently, which makes it useful not just as a tool for getting answers but as a way to understand how reasoning models approach hard problems.

Integrating QwQ-32B into Applications

Because Ollama exposes an OpenAI-compatible endpoint, integrating QwQ-32B into existing applications is straightforward. Any application using the OpenAI Python SDK, the Node.js SDK, or a framework built on OpenAI-compatible endpoints can switch to QwQ-32B by changing the base URL and model name. For LangChain users, the ChatOllama integration provides a drop-in replacement for ChatOpenAI that routes requests to the local Ollama server. For LlamaIndex, the Ollama LLM class provides the same. The practical difference to handle in application code is QwQ’s <think> trace: if your application expects responses to be immediately usable as text, the regex stripping step described above belongs in your response handling layer rather than in each individual prompt handler.

Benchmarking Your Local Setup

Once QwQ-32B is running, it is worth running a few standard tests to confirm the setup is working correctly and to measure inference speed on your hardware. A good benchmark prompt is a short AIME problem: “Let f(x) = x⁴ – 4x³ + 6x² – 4x + 1. Find the sum of all real solutions to f(f(x)) = 1.” A correctly functioning QwQ-32B should produce a detailed reasoning trace and arrive at the correct answer. Time the full response to get your tokens-per-second figure — the thinking trace for a problem like this typically runs 500–1500 tokens before the final answer. Multiply the elapsed time by your measured tokens-per-second to verify the figure matches Ollama’s reported throughput. If inference is much slower than expected, check that VRAM utilisation during inference is high using nvidia-smi — low GPU utilisation usually indicates a layer-offloading or context-length configuration issue.

The combination of QwQ-32B’s reasoning capability and Ollama’s ease of deployment makes local reasoning models more accessible than ever. Five minutes of setup and a 20GB download give you a model that competes with cloud reasoning APIs on hard mathematics and coding tasks, runs with complete privacy, and costs nothing per query after the hardware investment. For anyone doing serious quantitative work, building reasoning-model applications, or simply wanting to understand how frontier reasoning models think, running QwQ-32B locally is well worth the setup effort.

Leave a Comment