Skip to main content
ZhimaYuandi
中文
← Back to blog
#freetoken#qwen#moe#wsl2#llm

Deploying Qwen3.6-35B-A3B-NVFP4 with FreeToken: A Practical Report (RTX 5060 Ti 16G + WSL2)

A complete field report of deploying Qwen3.6-35B-A3B-NVFP4 on WSL2 + RTX 5060 Ti 16GB: the four key deviations from the tutorial (404 repo, CLI arguments, model format, Blackwell toolchain), the root causes and fixes for WSL process reaping and VRAM accounting issues (WSL 2.7.13 + systemd), final performance numbers (34s cold load, 22–24 tok/s decode), and OpenAI-compatible API setup.

Coding Express 28 min

Hardware: RTX 5060 Ti 16GB VRAM + 32GB system RAM + WSL2 Ubuntu-24.04. Result: Deployment succeeded — an OpenAI-compatible API is running at http://127.0.0.1:8000/v1, with a stable decode speed of 22–24 tok/s. This post is the field report for the previous one, “FreeToken Deployment Plan for Qwen3.6-35B-A3B-NVFP4”: the tutorial’s plan hit 4 pitfalls during execution that didn’t match the docs. This post records the root cause and final fix for each one, so the same setup can be reproduced on identical hardware.

1. Final Deployment Summary

# FreeToken Deployment — Final Summary Report
✅ Deployment result: Success
🖥️ Hardware: RTX 5060 Ti 16G + 32G system RAM (WSL2 Ubuntu-24.04)
📌 Model: Qwen3.6-35B-A3B-NVFP4 (official NVIDIA modelopt quant, 22GB, 3 shards)
📊 Runtime metrics:
- Cold load time: ~34s (with warm CUDA kernel cache; first cold start ~2–3 min)
- Stable decode speed: 22 ~ 24 tok/s (measured across samples: 22.5 / 24.2 / 24.4 / 23.7 / 24.0 / 23.4)
- Peak VRAM usage: ~5.9 GB / 16 GB (measured via nvidia-smi during inference)
- Peak RAM usage: ~19 GB / 23 GB (WSL limit; 16.9GB of expert weights resident in system RAM)
- Max Swap used: ~0 GB (no swap storm)
🧪 API test result: Success — the model self-introduces, generates code, and answers normally
⚠️ Issues found: see Section 3 (4 deployment pitfalls, all resolved)
🔄 Backup plan: no need to switch to Unsloth Desktop

2. Final Launch Configuration (Directly Reproducible)

python -m freetoken \
  --model-path /home/dev/models/Qwen3.6-35B-A3B-NVFP4-nvidia \
  --moe-strategy offload \
  --moe-cache-size 256 \
  --num-tokens 4096 \
  --kv-reserve-tokens 8192 \
  --max-prefill-length 8192 \
  --max-running-requests 1 \
  --moe-cpu-layers auto \
  --attention-backend triton \
  --host 127.0.0.1 --port 8000

Key parameter notes (deviations from the tutorial in parentheses)

Parameter Value Notes
--moe-strategy offload Required Expert weights stay in system RAM (16.9GB); GPU holds only dense weights + a small cache
--moe-cache-size 256 Critical Don’t use auto or 2048 (see Pitfall 4); 256 slots starts stably on this machine
--num-tokens 4096 Critical Actual KV capacity; the tutorial’s 8192 context only stays stable here with offload
--moe-cpu-layers auto Recommended Auto-pins 18 layers of CPU decode per the pin budget; the rest are fetched by the GPU
--attention-backend triton Required The tutorial writes --attn; the actual parameter name is --attention-backend
--max-running-requests 1 Required Single concurrency — a hardware limit

WSL system configuration

%UserProfile%\.wslconfig (vmIdleTimeout=0 added on top of the tutorial’s settings):

[wsl2]
memory=24GB
processors=8
swap=4GB
localhostForwarding=true
vmIdleTimeout=0

Also recommended: run the service under systemd (/etc/systemd/system/freetoken.service + [boot] systemd=true). The service no longer depends on any terminal session and auto-restarts via Restart=on-failure.

3. The 4 Pitfalls Hit in Practice (Key Differences from the Tutorial)

Pitfall 1: the tutorial’s repo, CLI, and argument names don’t match the real version

Tutorial says Reality
github.com/freetoken/freetoken 404 dead end; the real repo is github.com/FlashML-org/FreeToken (mirror: https://ghproxy.net/https://github.com/FlashML-org/FreeToken.git)
freetoken serve <model> CLI entry is python -m freetoken (the entry point also exposes ft)
--attn triton The argument is --attention-backend triton
--moe-cache-size auto auto OOMs on this machine (see Pitfall 4); needs an explicit 256

Pitfall 2: unsloth’s mixed-quant model isn’t supported by FreeToken → switch to the official NVIDIA modelopt build

The tutorial points at ModelScope’s unsloth/Qwen3.6-35B-A3B-NVFP4 (26.5GB, compressed-tensors mixed layout: some experts FP8, some NVFP4). FreeToken’s MoE quantization registry only has nvfp4 / fp8_block / mxfp4 / mxfp8 / unquantized — per-tensor FP8 experts are unsupported, and startup fails immediately with:

NotImplementedError: no quant method for moe.fp8_tensor

Fix: switch to the official NVIDIA nvidia/Qwen3.6-35B-A3B-NVFP4 (modelopt MIXED_PRECISION: 161 W4A16_NVFP4 experts + 130 FP8 attention), downloaded via the domestic mirror hf-mirror.com (22GB, 3 shards), and set HF_HUB_DISABLE_XET=1 to bypass the Xet protocol’s 401 errors.

Pitfall 3: RTX 5060 Ti (Blackwell SM120) needs the CUDA 13 toolchain

On startup, FreeToken JIT-compiles CUDA kernels via tvm_ffi (CUDA Graphs phase) targeting architecture compute_120. The system’s bundled nvcc 12.0 doesn’t know Blackwell:

nvcc fatal: Unsupported gpu architecture 'compute_120'

Fix: install the CUDA 13 toolchain from NVIDIA’s official apt repo (compile-time only):

wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
dpkg -i cuda-keyring_1.1-1_all.deb && apt-get update
apt-get install -y cuda-nvcc-13-0 cuda-cudart-dev-13-0 cuda-cccl-13-0
export CUDA_HOME=/usr/local/cuda-13.0

Also required: pip install ninja (FreeToken’s tvm_ffi compilation depends on it). Note that PyPI’s nvidia-cuda-nvcc-cu13 is a placeholder package on the Tsinghua mirror and fails to build from the official source — the apt repo is the right answer.

Pitfall 4: two “fake OOMs” — WSL session reaping + WDDM VRAM accounting

This was the most painful part of the whole deployment. Both phenomena look like “the service disappears after a successful start / CUDA OOM while VRAM is clearly not full”:

Symptom A: the service process gets killed when the WSL session exits

  • Symptom: logs stop at expert loading 0%–33%, the process vanishes — no OOM, no traceback, all memory released
  • Investigation: dmesg shows systemd shutdown records; a sleep-kept process also disappears
  • Root cause: any process launched from a wsl.exe frontend (including setsid nohup) is reaped by the WSL session when wsl.exe disconnects; the distro also auto-shuts down after 60 seconds idle (vmIdleTimeout defaults to 60000ms)
  • Fix: upgrade WSL to 2.7.13 (fixes session/process reaping) + configure vmIdleTimeout=0 + run the service under systemd

Symptom B: CUDA OOM while nvidia-smi shows less than 1GB of VRAM used

  • Symptom: Tried to allocate 540.00 MiB fails, with an error claiming “this process has 17179869184.00 GiB memory in use” (a 16GB virtual-accounting illusion), while PyTorch only actually allocated 3.1GB
  • Investigation: reproduced with a minimal torch script — a single 1GB allocation succeeds, 4GB fails; but run in isolation it can allocate 16GB in sequence; cudaMemGetInfo reports 11.6GB free yet even a 32MB allocation fails. banks (cudaHostRegister) were probed and don’t consume the GPU quota, ruling them out
  • Root cause: WSL2’s WDDM VRAM management can lock a process’s usable VRAM to 0 under certain states, unrelated to physical VRAM usage (a bug in older WSL versions)
  • Fix: upgrading WSL to 2.7.13 made the problem disappear

Conclusion: when you hit “OOM while VRAM isn’t full” or “process mysteriously disappears”, check the WSL version first (wsl --version); upgrading older versions to 2.7.13+ is the first thing to do.

4. Measured Performance & Resources

Metric Measured value Notes
Cold load ~34s With warm CUDA kernel cache; first cold start ~2–3 min (expert loading 45s + CUDA Graphs 42s + warmup 16s)
Stable decode 22–24 tok/s Higher than the tutorial’s expected 14–22 tok/s
Peak VRAM ~5.9GB / 16GB During inference; ~1GB idle (normal for the offload architecture)
Peak RAM ~19GB / 23GB 16.9GB of expert weights resident in system RAM
Swap ~0 No swap storm triggered
Model size 22GB NVIDIA modelopt build, 3 shards

The low VRAM footprint is normal for the offload architecture: dense weights + a small expert cache live in VRAM, while the expert bodies stay in system RAM and are moved in on demand. Physical VRAM usage is not a sign of a performance problem.

5. API Access Info

  • API base url: http://127.0.0.1:8000/v1
  • model name: Qwen3.6-35B-A3B-NVFP4-nvidia (defaults to the model directory’s basename)
  • api key: any string (no auth locally, e.g. sk-local)
curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-local" \
  -d '{"model":"Qwen3.6-35B-A3B-NVFP4-nvidia",
       "messages":[{"role":"user","content":"你是谁"}]}'

Returned normally in testing (see screenshots); can be wired into OpenAI-compatible endpoints of tools like Claude Code / OpenClaw.

6. Ops Recommendations

  1. Service hosting: use systemd (systemctl enable freetoken) instead of a terminal foreground session, so session reaping can’t take the service down
  2. Monitoring: sample free -h and nvidia-smi --query-gpu=memory.used,memory.total --format=csv every 30 seconds; stop the service and release memory if Swap stays >2GB or a CUDA OOM appears
  3. Don’t touch: keep --max-running-requests at 1; don’t blindly increase --num-tokens / KV context (hard ceiling of 16GB VRAM + 32GB RAM)
  4. On a “fake OOM”: first run wsl --version to confirm WSL ≥ 2.7.13, then check the runtime configuration