Your privacy choices

Optional Google Analytics cookies help us understand which guides are useful. They stay off unless you accept. You can change your choice in the footer. Privacy policy

How to Run Local AI with Ollama: A Beginner’s Guide with Three Real Tests

Run a small AI model on your own computer with Ollama. Follow the setup, reproduce three recorded tests, and understand the memory and privacy limits.

By Ubedulla · 7 min read
Laptop on a warmly lit desk illustrating a local AI workspace
AI-generated editorial illustration; not a photograph of the test setup.

You do not need a chatbot subscription to turn a short set of notes into a summary. You can run a model on your own computer. The trade-off is that your laptop supplies the memory and processing power, and you are responsible for checking the answer.

This guide walks through a small, repeatable Ollama setup. For this article, an automated test session on a Mac with an Apple M5 and 16GB of memory ran three fictional tasks with qwen3:8b. The model preserved a missing deadline instead of making one up. That is a useful result, but three prompts are not proof that it will always behave that way.

What you need before starting

Ollama runs local models and also offers cloud functionality. Here we use a downloaded model and a local endpoint. Start with the official installation guide for your operating system. Installation itself was not retested from scratch for this article: Ollama and the model were already present on the test Mac.

Our installed Qwen3 8B package occupied about 5.2GB on disk. That is not its complete runtime memory requirement. Leave room for the operating system, the model’s working memory and other apps. A smaller model may suit a machine with less available memory. Choose it for the task and hardware you have, rather than assuming the largest download is the best choice.

You need an internet connection to obtain the software and model. Afterward, local inference can work without a cloud model. The tests below used cloud functionality disabled, but did not disconnect the Mac from its network or perform a network-isolation audit.

1. Check that Ollama is running

Open Terminal and run:

ollama --version
ollama list

Our version was 0.32.8. If the second command reports that it cannot connect to the server, start the Ollama app or, for a terminal-managed session, run:

ollama serve

Leave that terminal open and use a second one for commands. Do not start a second server if the app is already running. An “address already in use” error usually means something is already listening on that port; investigate before changing ports or stopping processes. The CLI reference documents the available commands.

2. Download a model and try a short prompt

ollama pull qwen3:8b
ollama run qwen3:8b

The download is only needed if the model is not already installed. At the prompt, start with something you can check yourself: a fictional note, a paragraph you wrote or a small formatting task. Avoid your entire document archive as the first experiment.

Try: “Rewrite this sentence in plain English without adding information: The meeting has been provisionally rescheduled pending confirmation.” Judge whether it preserves the uncertainty. A smoother sentence that quietly turns a possible change into a confirmed one is a worse answer.

3. Make the local-only boundary explicit

For our terminal-managed test session on macOS, the server was started with:

OLLAMA_NO_CLOUD=1 ollama serve

Set this on the server process, not just on a separate client command. If an existing app-managed server is running, first quit it normally before starting the test server. This shell syntax is for macOS/Linux; environment-variable configuration differs on Windows. Ollama documents local processing and disabling cloud functionality in its FAQ.

“Local” is not a universal privacy guarantee. A browser extension, connected tool, synced folder or separate application can still transmit information. Keep the server on its default loopback address, use a local model, and review anything else you connect to it. Do not expose the service to the internet just to make it easier to reach from another device.

4. Reproduce the missing-deadline test

The API makes the test settings explicit. This macOS/Linux command uses the local server, disables thinking output for this supported model, requests a single JSON response from the API and sets a small output limit:

curl http://127.0.0.1:11434/api/chat \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen3:8b",
    "messages": [{
      "role": "user",
      "content": "Return only JSON with keys owner, task, due_date. Use null for missing values. Note: Priya will review the draft. No deadline was agreed."
    }],
    "stream": false,
    "think": false,
    "options": {"temperature": 0, "num_ctx": 4096, "num_predict": 250},
    "keep_alive": 0
  }'

The outer response includes metadata. The answer is inside message.content. In our run, that content was:

{
  "owner": "Priya",
  "task": "review the draft",
  "due_date": null
}

For applications, parse and validate generated output rather than assuming it is valid JSON because the prompt requested JSON. Ollama also supports a response-format option; see the chat API documentation. Our example deliberately records the simple prompt-based test we ran.

What the three tests actually showed

TaskObserved resultElapsed time
Summarize fictional meeting notesPreserved both assigned tasks and the undecided launch date14.69 seconds
Extract an owner, task and deadlineReturned Priya, “review the draft” and a null deadline11.20 seconds
Answer an unsupported questionResponded “Not stated in the note.”5.29 seconds

These are single-run, end-to-end timings, including request and loading overhead, not a speed benchmark. Each request used a 4,096-token context, temperature 0, a 250-token output ceiling and keep_alive: 0, which unloads the model afterward. Results and timings can differ on another machine, with another model build or on another run.

The last test supplied only a weekday opening time and asked for a Saturday closing time. Refusing to guess was the intended behavior. To make this a useful evaluation for your own work, add harder cases: contradictory notes, two people with the same name, dates written ambiguously and text that contains instructions to ignore your rules.

Download the exact prompts and recorded outputs. All examples use fictional data. The test was executed through an automated assistant session, not a weeks-long human product review.

When this setup is worth keeping

Keep it if it reliably completes a narrow task you repeat and you are comfortable maintaining the setup. Short notes, rewriting your own text and extracting fields are sensible starting points. Do not treat a fluent answer as evidence that the model knows today’s news or has read a file you never supplied.

If the computer becomes sluggish, close other demanding applications or try a smaller model. A larger context consumes additional memory. If answers are wrong, making the model bigger is not the only lever: shorten the input, state what is missing and compare the output against the original.

For cleanup, ollama ps shows running models, and ollama stop qwen3:8b unloads this model if it is running. Keep the downloaded model if you plan to use it again. You are avoiding a hosted subscription, not eliminating electricity, storage or hardware costs.

Frequently asked questions

Can this replace a paid chatbot?

For a narrow task, possibly. These three examples do not establish equivalent reasoning, research, coding or reliability. Test the actual workload before cancelling a tool you depend on.

Does local AI mean my documents are automatically available to it?

No. This example sends only the text in each request. Searching a folder or answering from PDFs needs additional software and careful source handling. Our RAG explainer describes the general approach.

Is the model completely accurate with temperature set to zero?

No. That setting does not fix incorrect knowledge or guarantee consistent behavior across systems. In our tests, correctness was assessed against short notes with known answers.

Tested October 8, 2026. AI-assisted writing with recorded local execution; installation steps were checked against official documentation. See our editorial policy. For hosted alternatives, see our free AI tools guide.

About the author

Ubedulla

Founder & Editor

Founder and editor of The Bot Post, covering AI news and technology.

Related Articles