Hardware: RTX 5060 Ti 16GB VRAM + 32GB system RAM + WSL2 Ubuntu-24.04. Result: Deployment succeeded — an OpenAI-compatible API is running at
http://127.0.0.1:8000/v1, with a stable decode speed of 22–24 tok/s. This post is the field report for the previous one, “FreeToken Deployment Plan for Qwen3.6-35B-A3B-NVFP4”: the tutorial’s plan hit 4 pitfalls during execution that didn’t match the docs. This post records the root cause and final fix for each one, so the same setup can be reproduced on identical hardware.
1. Final Deployment Summary
# FreeToken Deployment — Final Summary Report
✅ Deployment result: Success
🖥️ Hardware: RTX 5060 Ti 16G + 32G system RAM (WSL2 Ubuntu-24.04)
📌 Model: Qwen3.6-35B-A3B-NVFP4 (official NVIDIA modelopt quant, 22GB, 3 shards)
📊 Runtime metrics:
- Cold load time: ~34s (with warm CUDA kernel cache; first cold start ~2–3 min)
- Stable decode speed: 22 ~ 24 tok/s (measured across samples: 22.5 / 24.2 / 24.4 / 23.7 / 24.0 / 23.4)
- Peak VRAM usage: ~5.9 GB / 16 GB (measured via nvidia-smi during inference)
- Peak RAM usage: ~19 GB / 23 GB (WSL limit; 16.9GB of expert weights resident in system RAM)
- Max Swap used: ~0 GB (no swap storm)
🧪 API test result: Success — the model self-introduces, generates code, and answers normally
⚠️ Issues found: see Section 3 (4 deployment pitfalls, all resolved)
🔄 Backup plan: no need to switch to Unsloth Desktop
2. Final Launch Configuration (Directly Reproducible)
python -m freetoken \
--model-path /home/dev/models/Qwen3.6-35B-A3B-NVFP4-nvidia \
--moe-strategy offload \
--moe-cache-size 256 \
--num-tokens 4096 \
--kv-reserve-tokens 8192 \
--max-prefill-length 8192 \
--max-running-requests 1 \
--moe-cpu-layers auto \
--attention-backend triton \
--host 127.0.0.1 --port 8000
Key parameter notes (deviations from the tutorial in parentheses)
| Parameter | Value | Notes |
|---|---|---|
--moe-strategy offload |
Required | Expert weights stay in system RAM (16.9GB); GPU holds only dense weights + a small cache |
--moe-cache-size 256 |
Critical | Don’t use auto or 2048 (see Pitfall 4); 256 slots starts stably on this machine |
--num-tokens 4096 |
Critical | Actual KV capacity; the tutorial’s 8192 context only stays stable here with offload |
--moe-cpu-layers auto |
Recommended | Auto-pins 18 layers of CPU decode per the pin budget; the rest are fetched by the GPU |
--attention-backend triton |
Required | The tutorial writes --attn; the actual parameter name is --attention-backend |
--max-running-requests 1 |
Required | Single concurrency — a hardware limit |
WSL system configuration
%UserProfile%\.wslconfig (vmIdleTimeout=0 added on top of the tutorial’s settings):
[wsl2]
memory=24GB
processors=8
swap=4GB
localhostForwarding=true
vmIdleTimeout=0
Also recommended: run the service under systemd (/etc/systemd/system/freetoken.service + [boot] systemd=true). The service no longer depends on any terminal session and auto-restarts via Restart=on-failure.
3. The 4 Pitfalls Hit in Practice (Key Differences from the Tutorial)
Pitfall 1: the tutorial’s repo, CLI, and argument names don’t match the real version
| Tutorial says | Reality |
|---|---|
github.com/freetoken/freetoken |
404 dead end; the real repo is github.com/FlashML-org/FreeToken (mirror: https://ghproxy.net/https://github.com/FlashML-org/FreeToken.git) |
freetoken serve <model> |
CLI entry is python -m freetoken (the entry point also exposes ft) |
--attn triton |
The argument is --attention-backend triton |
--moe-cache-size auto |
auto OOMs on this machine (see Pitfall 4); needs an explicit 256 |
Pitfall 2: unsloth’s mixed-quant model isn’t supported by FreeToken → switch to the official NVIDIA modelopt build
The tutorial points at ModelScope’s unsloth/Qwen3.6-35B-A3B-NVFP4 (26.5GB, compressed-tensors mixed layout: some experts FP8, some NVFP4). FreeToken’s MoE quantization registry only has nvfp4 / fp8_block / mxfp4 / mxfp8 / unquantized — per-tensor FP8 experts are unsupported, and startup fails immediately with:
NotImplementedError: no quant method for moe.fp8_tensor
Fix: switch to the official NVIDIA nvidia/Qwen3.6-35B-A3B-NVFP4 (modelopt MIXED_PRECISION: 161 W4A16_NVFP4 experts + 130 FP8 attention), downloaded via the domestic mirror hf-mirror.com (22GB, 3 shards), and set HF_HUB_DISABLE_XET=1 to bypass the Xet protocol’s 401 errors.
Pitfall 3: RTX 5060 Ti (Blackwell SM120) needs the CUDA 13 toolchain
On startup, FreeToken JIT-compiles CUDA kernels via tvm_ffi (CUDA Graphs phase) targeting architecture compute_120. The system’s bundled nvcc 12.0 doesn’t know Blackwell:
nvcc fatal: Unsupported gpu architecture 'compute_120'
Fix: install the CUDA 13 toolchain from NVIDIA’s official apt repo (compile-time only):
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
dpkg -i cuda-keyring_1.1-1_all.deb && apt-get update
apt-get install -y cuda-nvcc-13-0 cuda-cudart-dev-13-0 cuda-cccl-13-0
export CUDA_HOME=/usr/local/cuda-13.0
Also required: pip install ninja (FreeToken’s tvm_ffi compilation depends on it). Note that PyPI’s nvidia-cuda-nvcc-cu13 is a placeholder package on the Tsinghua mirror and fails to build from the official source — the apt repo is the right answer.
Pitfall 4: two “fake OOMs” — WSL session reaping + WDDM VRAM accounting
This was the most painful part of the whole deployment. Both phenomena look like “the service disappears after a successful start / CUDA OOM while VRAM is clearly not full”:
Symptom A: the service process gets killed when the WSL session exits
- Symptom: logs stop at expert loading 0%–33%, the process vanishes — no OOM, no traceback, all memory released
- Investigation:
dmesgshows systemd shutdown records; asleep-kept process also disappears - Root cause: any process launched from a
wsl.exefrontend (includingsetsid nohup) is reaped by the WSL session whenwsl.exedisconnects; the distro also auto-shuts down after 60 seconds idle (vmIdleTimeoutdefaults to 60000ms) - Fix: upgrade WSL to 2.7.13 (fixes session/process reaping) + configure
vmIdleTimeout=0+ run the service under systemd
Symptom B: CUDA OOM while nvidia-smi shows less than 1GB of VRAM used
- Symptom:
Tried to allocate 540.00 MiBfails, with an error claiming “this process has 17179869184.00 GiB memory in use” (a 16GB virtual-accounting illusion), while PyTorch only actually allocated 3.1GB - Investigation: reproduced with a minimal torch script — a single 1GB allocation succeeds, 4GB fails; but run in isolation it can allocate 16GB in sequence;
cudaMemGetInforeports 11.6GB free yet even a 32MB allocation fails. banks (cudaHostRegister) were probed and don’t consume the GPU quota, ruling them out - Root cause: WSL2’s WDDM VRAM management can lock a process’s usable VRAM to 0 under certain states, unrelated to physical VRAM usage (a bug in older WSL versions)
- Fix: upgrading WSL to 2.7.13 made the problem disappear
Conclusion: when you hit “OOM while VRAM isn’t full” or “process mysteriously disappears”, check the WSL version first (
wsl --version); upgrading older versions to 2.7.13+ is the first thing to do.
4. Measured Performance & Resources
| Metric | Measured value | Notes |
|---|---|---|
| Cold load | ~34s | With warm CUDA kernel cache; first cold start ~2–3 min (expert loading 45s + CUDA Graphs 42s + warmup 16s) |
| Stable decode | 22–24 tok/s | Higher than the tutorial’s expected 14–22 tok/s |
| Peak VRAM | ~5.9GB / 16GB | During inference; ~1GB idle (normal for the offload architecture) |
| Peak RAM | ~19GB / 23GB | 16.9GB of expert weights resident in system RAM |
| Swap | ~0 | No swap storm triggered |
| Model size | 22GB | NVIDIA modelopt build, 3 shards |
The low VRAM footprint is normal for the offload architecture: dense weights + a small expert cache live in VRAM, while the expert bodies stay in system RAM and are moved in on demand. Physical VRAM usage is not a sign of a performance problem.
5. API Access Info
- API base url:
http://127.0.0.1:8000/v1 - model name:
Qwen3.6-35B-A3B-NVFP4-nvidia(defaults to the model directory’s basename) - api key: any string (no auth locally, e.g.
sk-local)
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-local" \
-d '{"model":"Qwen3.6-35B-A3B-NVFP4-nvidia",
"messages":[{"role":"user","content":"你是谁"}]}'
Returned normally in testing (see screenshots); can be wired into OpenAI-compatible endpoints of tools like Claude Code / OpenClaw.
6. Ops Recommendations
- Service hosting: use systemd (
systemctl enable freetoken) instead of a terminal foreground session, so session reaping can’t take the service down - Monitoring: sample
free -handnvidia-smi --query-gpu=memory.used,memory.total --format=csvevery 30 seconds; stop the service and release memory if Swap stays >2GB or a CUDA OOM appears - Don’t touch: keep
--max-running-requestsat 1; don’t blindly increase--num-tokens/ KV context (hard ceiling of 16GB VRAM + 32GB RAM) - On a “fake OOM”: first run
wsl --versionto confirm WSL ≥ 2.7.13, then check the runtime configuration