How to Run AI Locally on Your PC or Mac – Step by Step
Install Ollama, download a small Qwen model, and get a reply from your own computer. No programming is needed. Add Open WebUI afterwards if you want chat in a browser.
Install Ollama, download Qwen and send your first prompt. The model generates its reply on your PC or Mac. The basic setup needs neither a terminal nor programming.
What you need
- The Ollama app or PowerShell/Terminal
- Qwen 3.5 4B as a first model
- Internet for installation and the model download
- Optional: Docker Desktop and Open WebUI
What each part does
- 01Open WebUI
The chat screen in your browser. Optional.
- 02Ollama
Downloads the model and runs it on your computer.
- 03Qwen
The language model that reads your prompt and writes a reply.
- 04Your hardware
Your CPU, GPU and RAM or unified memory do the work.
Start with Ollama and Qwen. Add Open WebUI only after the first reply works.
Windows
1. Install Ollama
Install with the app
Download the installer, open it and follow the normal installation flow.
Download Ollama for Windows ↗Install from PowerShell
Official Ollama installer for Windows.
irm https://ollama.com/install.ps1 | iex2. Get Qwen 3.5 4B
Use the Ollama app
Open Ollama’s model selector or library. Search for Qwen 3.5 4B and download the local model.
Qwen 3.5 on Ollama ↗ollama pull qwen3.5:4b
Download the model without starting a chat.
ollama pull qwen3.5:4b3. Start Qwen
Use the Ollama app
Select Qwen 3.5 4B in the Ollama app to open the chat.
ollama run qwen3.5:4b
Start the model in an interactive terminal chat.
ollama run qwen3.5:4b4. Send your first prompt
Explain the difference between RAM and VRAM to someone who has never built a computer. Use no more than five short bullet points.
You are ready when: Ollama shows a reply below your prompt and the selected model is qwen3.5:4b. If the app still offers a download, the model is not ready yet.
Mac
1. Install Ollama
Install with the app
Download the app and install or open it normally. Choose the build for your Mac chip if asked.
Download Ollama for macOS ↗Terminal
Official Ollama install script for macOS.
curl -fsSL https://ollama.com/install.sh | sh2. Get Qwen 3.5 4B
Use the Ollama app
Open Ollama’s model selector or library. Search for Qwen 3.5 4B and download the local model.
Qwen 3.5 on Ollama ↗ollama pull qwen3.5:4b
Download the model without starting a chat.
ollama pull qwen3.5:4b3. Start Qwen
Use the Ollama app
Select Qwen 3.5 4B in the Ollama app to open the chat.
ollama run qwen3.5:4b
Start the model in an interactive terminal chat.
ollama run qwen3.5:4b4. Send your first prompt
Explain the difference between RAM and VRAM to someone who has never built a computer. Use no more than five short bullet points.
You are ready when: Ollama shows a reply below your prompt and the selected model is qwen3.5:4b. If the app still offers a download, the model is not ready yet.
Once Ollama works: four useful commands
You do not need a terminal to chat in the app. These commands help when you want to start a specific model or check what is running.
Run Qwen in a terminal
ollama run qwen3.5:4bList downloaded models
ollama lsSee models running now
ollama psStop the loaded model
ollama stop qwen3.5:4bCheck that the model is running locally
Run ollama ps while the model is replying. Its table lists active models, how long they have been loaded and where computation is placed when that information is available. A CPU/GPU split describes offload for that run; it does not measure answer quality.
On Windows, open Task Manager and look for GPU or memory activity while a reply is being generated. On a Mac, Activity Monitor can show processes and memory pressure. Those numbers alone do not prove that all work ran on the GPU, but alongside ollama ps they provide a useful snapshot.
Try a simple offline check
- Finish downloading the model while online.
- Turn off Wi-Fi and disconnect Ethernet if connected.
- Ask qwen3.5:4b: “Write a short story about a dog learning to use a computer.”
If it replies with the network disconnected, that shows this text generation did not need a cloud model. It does not prove other apps or later integrations are offline. Ollama also offers cloud models; choose the local model without a cloud label.
Add Open WebUI for browser chat
Open WebUI adds a browser chat screen, conversation history and a model selector. The project recommends Docker for most users. Ollama already works without Open WebUI; this part is optional.
Docker Desktop for Windows ↗Docker Desktop for Mac ↗Open WebUI official Docker guide ↗
Install Docker Desktop
Download the official app for Windows or Mac, install it and open it. Wait until Docker reports that its engine is running.
Create an Open WebUI key
Generate a random WEBUI_SECRET_KEY. Keep it private and use the same key if you recreate the container.
Start the container
Run the Docker command with your generated key. It stores data in a Docker volume and connects the container to Ollama on your computer.
Open the local page
Create an account in your local Open WebUI installation, select qwen3.5:4b and send a short prompt.
Generate a random key with OpenSSL
OpenSSL
openssl rand -hex 32macOS usually includes OpenSSL. On Windows, use Git Bash or another OpenSSL installation. Never share the key.
Start Open WebUI in Docker
Docker run
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data -e WEBUI_SECRET_KEY=your-secret-key --name open-webui --restart always ghcr.io/open-webui/open-webui:mainReplace your-secret-key with the value generated by the previous command.
Port 3000 on your computer maps to port 8080 in the container. The open-webui Docker volume keeps data across restarts. This command follows Open WebUI’s official Quick Start.
If something does not work
- No model appears
- Check that the download finished and that you selected qwen3.5:4b without a cloud label. The first download needs an internet connection.
- Ollama is very slow
- Close demanding apps, try the small model again and allow time for its first load. Speed depends on model variant, context, memory, runtime and GPU offload, among other things.
- Open WebUI cannot find Ollama
- Make sure Ollama is still running. The connection should use http://host.docker.internal:11434 when Ollama runs on your host and Open WebUI runs in Docker.
- Port 3000 is in use
- Stop the other local service or change the left-hand port in the Docker command to 3001, then open localhost:3001.
After the 4B model works
Ollama’s current Qwen 3.5 library also lists 9B and 27B tags. A larger model can offer different capabilities, but usually takes more storage and working memory and may respond more slowly. Try 9B first if 4B feels comfortable; move to 27B when you have room to experiment.
File size is not the same as required RAM or VRAM. Practical use also depends on quantization, context, runtime, GPU offload, operating system and what else is open. There is no single minimum that means a model will work perfectly.
Five useful terms
- 4B, 9B and 27B
- B means billion model parameters. The number says something about model size, but does not predict answer quality or speed by itself.
- Quantization
- A way to store model weights with fewer bits so the file takes less space. This can reduce memory needs, with results that vary by model and task.
- VRAM and RAM
- VRAM is memory on a graphics card. System RAM is the computer’s main memory. Apple Silicon uses unified memory shared by CPU and GPU; the OS and apps need some too.
- Context
- The text and conversation the model can work with at once. Longer context uses more memory and can slow processing.
- Tokens per second
- A rough measure of how quickly the model writes after it starts responding. It differs from time to first token and answer quality.
Being able to run a model is only the first test
The model starts and produces output.
Speed, memory use and stability fit the work you actually want to do.
Quality, speed, context, hardware needs, cost and usability make sense together.
RigRoute Labs measures local AI setups because model names and nominal system requirements do not tell the whole story. Our existing Mac study looks at first-token delay, generation and task outcomes; it uses a different model and is not a direct Qwen recommendation.
Common questions
Can I run local AI without programming?
Yes. Ollama’s Windows and macOS apps include a graphical model and chat experience. Terminal commands and Open WebUI are optional.
Do I need to be online?
You need a connection to download Ollama and the model. After the local model is downloaded, text generation can work offline. Cloud models and external integrations need a connection.
Are my prompts private?
Ollama does not send this local inference to a model server in the setup shown here, but check tools and connections you add. This guide does not promise the entire computer or Open WebUI will always be offline.
What computer can run Qwen 3.5 4B?
Ollama has Windows and Mac builds, but whether the model feels fast or fits comfortably in memory depends on hardware and settings. Start with the small model and judge it on your machine.
Official setup guides and model listing
Next step
You now have a local model that can answer. The next step is choosing one that fits your hardware and the work you want it to do.