RigRoute LabsDAENRigRoute ↗
RigRoute Labs · Practical guide

How to Run AI Locally on Your PC or Mac – Step by Step

Install Ollama, download a small Qwen model, and get a reply from your own computer. No programming is needed. Add Open WebUI afterwards if you want chat in a browser.

Install Ollama, download Qwen and send your first prompt. The model generates its reply on your PC or Mac. The basic setup needs neither a terminal nor programming.

What you need

  • The Ollama app or PowerShell/Terminal
  • Qwen 3.5 4B as a first model
  • Internet for installation and the model download
  • Optional: Docker Desktop and Open WebUI

What each part does

  1. 01
    Open WebUI

    The chat screen in your browser. Optional.

  2. 02
    Ollama

    Downloads the model and runs it on your computer.

  3. 03
    Qwen

    The language model that reads your prompt and writes a reply.

  4. 04
    Your hardware

    Your CPU, GPU and RAM or unified memory do the work.

Start with Ollama and Qwen. Add Open WebUI only after the first reply works.

Windows

1. Install Ollama

PowerShell

Install from PowerShell

Official Ollama installer for Windows.

irm https://ollama.com/install.ps1 | iex

2. Get Qwen 3.5 4B

Use the Ollama app

Open Ollama’s model selector or library. Search for Qwen 3.5 4B and download the local model.

Qwen 3.5 on Ollama ↗
PowerShell

ollama pull qwen3.5:4b

Download the model without starting a chat.

ollama pull qwen3.5:4b

3. Start Qwen

Use the Ollama app

Select Qwen 3.5 4B in the Ollama app to open the chat.

PowerShell

ollama run qwen3.5:4b

Start the model in an interactive terminal chat.

ollama run qwen3.5:4b

4. Send your first prompt

Prompt to paste

Explain the difference between RAM and VRAM to someone who has never built a computer. Use no more than five short bullet points.

You are ready when: Ollama shows a reply below your prompt and the selected model is qwen3.5:4b. If the app still offers a download, the model is not ready yet.

Mac

1. Install Ollama

Terminal

Terminal

Official Ollama install script for macOS.

curl -fsSL https://ollama.com/install.sh | sh

2. Get Qwen 3.5 4B

Use the Ollama app

Open Ollama’s model selector or library. Search for Qwen 3.5 4B and download the local model.

Qwen 3.5 on Ollama ↗
Terminal

ollama pull qwen3.5:4b

Download the model without starting a chat.

ollama pull qwen3.5:4b

3. Start Qwen

Use the Ollama app

Select Qwen 3.5 4B in the Ollama app to open the chat.

Terminal

ollama run qwen3.5:4b

Start the model in an interactive terminal chat.

ollama run qwen3.5:4b

4. Send your first prompt

Prompt to paste

Explain the difference between RAM and VRAM to someone who has never built a computer. Use no more than five short bullet points.

You are ready when: Ollama shows a reply below your prompt and the selected model is qwen3.5:4b. If the app still offers a download, the model is not ready yet.

Once Ollama works: four useful commands

You do not need a terminal to chat in the app. These commands help when you want to start a specific model or check what is running.

Run Qwen in a terminal

ollama run qwen3.5:4b

List downloaded models

ollama ls

See models running now

ollama ps

Stop the loaded model

ollama stop qwen3.5:4b

Check that the model is running locally

Run ollama ps while the model is replying. Its table lists active models, how long they have been loaded and where computation is placed when that information is available. A CPU/GPU split describes offload for that run; it does not measure answer quality.

On Windows, open Task Manager and look for GPU or memory activity while a reply is being generated. On a Mac, Activity Monitor can show processes and memory pressure. Those numbers alone do not prove that all work ran on the GPU, but alongside ollama ps they provide a useful snapshot.

Try a simple offline check

  1. Finish downloading the model while online.
  2. Turn off Wi-Fi and disconnect Ethernet if connected.
  3. Ask qwen3.5:4b: “Write a short story about a dog learning to use a computer.”

If it replies with the network disconnected, that shows this text generation did not need a cloud model. It does not prove other apps or later integrations are offline. Ollama also offers cloud models; choose the local model without a cloud label.

Add Open WebUI for browser chat

Open WebUI adds a browser chat screen, conversation history and a model selector. The project recommends Docker for most users. Ollama already works without Open WebUI; this part is optional.

  1. Install Docker Desktop

    Download the official app for Windows or Mac, install it and open it. Wait until Docker reports that its engine is running.

  2. Create an Open WebUI key

    Generate a random WEBUI_SECRET_KEY. Keep it private and use the same key if you recreate the container.

  3. Start the container

    Run the Docker command with your generated key. It stores data in a Docker volume and connects the container to Ollama on your computer.

  4. Open the local page

    Create an account in your local Open WebUI installation, select qwen3.5:4b and send a short prompt.

Generate a random key with OpenSSL

Windows / macOS

OpenSSL

openssl rand -hex 32

macOS usually includes OpenSSL. On Windows, use Git Bash or another OpenSSL installation. Never share the key.

Start Open WebUI in Docker

Docker

Docker run

docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data -e WEBUI_SECRET_KEY=your-secret-key --name open-webui --restart always ghcr.io/open-webui/open-webui:main

Replace your-secret-key with the value generated by the previous command.

Port 3000 on your computer maps to port 8080 in the container. The open-webui Docker volume keeps data across restarts. This command follows Open WebUI’s official Quick Start.

Open WebUI on this computer ↗

If something does not work

No model appears
Check that the download finished and that you selected qwen3.5:4b without a cloud label. The first download needs an internet connection.
Ollama is very slow
Close demanding apps, try the small model again and allow time for its first load. Speed depends on model variant, context, memory, runtime and GPU offload, among other things.
Open WebUI cannot find Ollama
Make sure Ollama is still running. The connection should use http://host.docker.internal:11434 when Ollama runs on your host and Open WebUI runs in Docker.
Port 3000 is in use
Stop the other local service or change the left-hand port in the Docker command to 3001, then open localhost:3001.

After the 4B model works

Ollama’s current Qwen 3.5 library also lists 9B and 27B tags. A larger model can offer different capabilities, but usually takes more storage and working memory and may respond more slowly. Try 9B first if 4B feels comfortable; move to 27B when you have room to experiment.

File size is not the same as required RAM or VRAM. Practical use also depends on quantization, context, runtime, GPU offload, operating system and what else is open. There is no single minimum that means a model will work perfectly.

Five useful terms

4B, 9B and 27B
B means billion model parameters. The number says something about model size, but does not predict answer quality or speed by itself.
Quantization
A way to store model weights with fewer bits so the file takes less space. This can reduce memory needs, with results that vary by model and task.
VRAM and RAM
VRAM is memory on a graphics card. System RAM is the computer’s main memory. Apple Silicon uses unified memory shared by CPU and GPU; the OS and apps need some too.
Context
The text and conversation the model can work with at once. Longer context uses more memory and can slow processing.
Tokens per second
A rough measure of how quickly the model writes after it starts responding. It differs from time to first token and answer quality.

Being able to run a model is only the first test

CAN RUN

The model starts and produces output.

USABLE

Speed, memory use and stability fit the work you actually want to do.

RECOMMENDED

Quality, speed, context, hardware needs, cost and usability make sense together.

RigRoute Labs measures local AI setups because model names and nominal system requirements do not tell the whole story. Our existing Mac study looks at first-token delay, generation and task outcomes; it uses a different model and is not a direct Qwen recommendation.

How RigRoute Labs measures local AI ↗

Browse tested hardware ↗

Common questions

Can I run local AI without programming?

Yes. Ollama’s Windows and macOS apps include a graphical model and chat experience. Terminal commands and Open WebUI are optional.

Do I need to be online?

You need a connection to download Ollama and the model. After the local model is downloaded, text generation can work offline. Cloud models and external integrations need a connection.

Are my prompts private?

Ollama does not send this local inference to a model server in the setup shown here, but check tools and connections you add. This guide does not promise the entire computer or Open WebUI will always be offline.

What computer can run Qwen 3.5 4B?

Ollama has Windows and Mac builds, but whether the model feels fast or fits comfortably in memory depends on hardware and settings. Start with the small model and judge it on your machine.

Official setup guides and model listing

Next step

You now have a local model that can answer. The next step is choosing one that fits your hardware and the work you want it to do.

Build a local AI system ↗