Run local AI on Windows with Ollama
Install Ollama, run a first local model and check a simple document task before choosing a larger model or building an automated workflow.
In this guide
You can start with a short text and a small model. The aim of the first session is to get a local answer, understand how to check it and find out whether your computer can handle the task comfortably.
Before you start
Use a Windows version supported by Ollama and allow space for both the application and model downloads. Ollama's current Windows documentation lists Windows 10 22H2 or newer and explains GPU compatibility and driver requirements. Check that page for your hardware rather than assuming that every graphics card is supported.
Download the Windows installer from Ollama's official website. The installer provides a background application and the ollama command. Open a new PowerShell window after installation so it can find the command.
Local inference and downloads are separate. Installing the application and downloading a model use the internet. This guide's document requests use a locally installed model. If you select a cloud model or connect another service, that changes where the request is processed. See Ollama's FAQ before using sensitive documents.
Run your first model
In PowerShell, check the installation and download the small Qwen tag used in our supporting examples:
ollama --version
ollama pull qwen3:1.7b
ollama run qwen3:1.7bThis is a manageable starting example, not a recommendation that this model is the best fit for your documents. After the model loads, type a short request at its prompt. Enter /bye to leave the interactive session.
ollama list lists downloaded models. ollama ps lists currently loaded models and shows whether processing uses CPU, GPU or a combination. The latter is useful if generation feels unexpectedly slow.
Do not begin by downloading every model in a comparison. First make one model work, then compare a second candidate using the same texts and instructions.
Try a document task you can verify
Paste this fictional English example into the interactive model session:
Summarize the text below in at most two sentences.
Use only information in the text. Keep numbers and uncertainty.
Do not add a quotation or a background claim.
SOURCE:
A support team tested automatic ticket sorting for six weeks.
In its first check, 18 of 60 tickets were sorted incorrectly.
A wider rollout has not been approved.A useful answer should preserve 18 of 60, the six-week scope and the fact that a wider rollout has not been approved. It should not turn the trial into a successful launch. Different wording is acceptable if those facts remain intact.
This is a practice prompt, not an additional recorded benchmark result. Try the same check with a short, non-sensitive document in the language you actually use. Keep the source beside the answer while reviewing it.
If the answer is plausible but wrong, a successful installation has still told you something useful: the model needs a different prompt, a different candidate or more checking for this task. Continue with the document summarization guide.
Solve common setup problems
PowerShell cannot find “ollama”
Open a new terminal after installation. Confirm that the app is installed and running. The official Windows guide lists the binary and log locations if you need to inspect an installation problem.
The model runs, but generation is slow
Check ollama ps. A larger model, a longer context and other work on the computer can increase memory pressure. Try a smaller model and a shorter source before changing the system configuration. A model file's download size is not a complete estimate of memory needed while generating text.
A long document gives incomplete or confusing results
The context window limits how much text the model can use in a request. Instructions, the source and the generated answer all need room. Check the active context and try a coherent shorter section. Increasing context can increase memory requirements; it is not a substitute for checking what the model read.
The answer ends abruptly or contains only reasoning
Inspect the final answer separately from any thinking output. If an API response ends with done_reason: "length", treat it as incomplete. Model-specific thinking controls and output limits need their own check; do not score an unfinished answer as a normal result.
Your next step
Once one model works, choose a task with a clear success condition. For summaries, compare claims against the source. For fields such as names and amounts, define the output format and check each value. You can use our fictional Swedish examples to practice both checks without supplying your own documents.
The optional runner requires Node.js 20 or later and the three model tags listed in its instructions. It records responses and test settings; it does not collect a computer profile or measure speed. Follow the reproduction instructions in either document guide.
Documentation
- Ollama on Windows — installation, compatibility and troubleshooting.
- Ollama FAQ — local and cloud processing, loaded models and API settings.
- Context length — active context and memory tradeoffs.
- Thinking controls — model-specific behaviour and response fields.
Documentation checked on 5 October 2026. This is a setup guide, not a hardware buying guide.