Overview
The Local LLM Inference guide pulls the model straight from Hugging Face — the fastest route when your NAS has a good connection. This page covers the offline route: download the model on a PC, then move it to the NAS over USB or your LAN.
The steps below use the same model as the main guide — Qwen3.6-35B-A3B-UD-IQ3_XXS.gguf, about 12.3 GB — and the same destination folder. The download-and-transfer pattern works for any GGUF model.
Before You Start
- A PC with internet access
- About 13 GB of free space on the PC, and the same on the NAS
- A USB drive, or the NAS reachable over your LAN
Step 1: Download the Model on Your PC
The Hugging Face CLI is the easiest way — it resumes interrupted downloads:
pip install -U huggingface_hub |
You can also download from the model page in a browser.
Step 2: Move the Model to the NAS
Over USB: copy the models/llm folder to the drive, plug it into the NAS, then in ZimaOS Files move the folder to the NAS’s models/llm directory (create it if it does not exist).
Over LAN: push the file from the PC with scp. Create the destination folder on the NAS first:
ssh <username>@<your-nas-ip> mkdir -p models/llm |
Step 3: Verify the File
Large downloads can corrupt silently. On the NAS, check the size and the checksum:
ls -lh models/llm |
Compare both against the file card on the model page — the file should be about 12.3 GB.
Step 4: Continue the Setup
The model is now in models/llm — exactly where the Local LLM Inference guide expects it. Skip that guide’s download step and continue with starting the server.
Reference Links
- Hugging Face – huggingface_hub CLI documentation
- unsloth – Qwen3.6-35B-A3B-GGUF model card