Running large language models used to mean spinning up GPU instances, managing Python environments with conflicting dependencies, and debugging CUDA installations. Ollama changes that equation entirely. It packages LLMs into a Docker-like workflow: pull a model, run it, interact with it through a simple API. No virtual environments, no dependency hell, no GPU driver troubleshooting.
Why would you want to run models locally instead of just calling the OpenAI or Anthropic APIs? A few compelling reasons:
- Privacy: Your code, prompts, and data never leave your machine. This matters when you're working with proprietary codebases or sensitive data that you can't send to third-party APIs.
- Zero cost for experimentation: Local models have no per-token pricing. You can send thousands of requests while prototyping without watching a billing dashboard.
- Offline availability: You can work on airplanes, in cafes with unreliable WiFi, or anywhere without internet access.
- Faster iteration: No network latency. For small prompts, local inference on a good machine can return results faster than a round trip to a cloud API.
This guide will take you from zero to running local models, building custom configurations, and integrating them into your development workflow.
Installation
Ollama is available on all major platforms.
macOS (via Homebrew):
brew install ollama
Linux (one-line installer):
curl -fsSL https://ollama.com/install.sh | sh
Windows: Download the installer from ollama.com. The Windows version runs as a native application with a system tray icon.
After installation, Ollama runs a background server process that handles model loading and inference. On macOS and Windows, this starts automatically. On Linux, the install script sets up a systemd service.
GPU support is automatic. Ollama detects NVIDIA GPUs (via CUDA), AMD GPUs (via ROCm), and on Mac it uses Metal for Apple Silicon acceleration. If no GPU is available, it falls back to CPU-only inference, which works but is significantly slower -- expect 5-10x longer generation times depending on the model.
Verify your installation:
ollama --version
You should see something like ollama version 0.5.x. If the server isn't running, start it manually with ollama serve.
Pulling and Running Models
Ollama's model management works like Docker. You pull models from a registry, run them, and manage local copies.
# Pull a model (downloads it to your machine)
ollama pull llama3.1:8b
# Run a model interactively (pulls automatically if not downloaded)
ollama run llama3.1:8b
# List all downloaded models
ollama list
# Remove a model to free disk space
ollama rm llama3.1:8b
When you ollama run a model, you get an interactive chat session in your terminal. Type your prompt, press Enter, and the model streams its response. Type /bye to exit.
Choosing a Model
The model you pick depends on your hardware and use case. Here are the most useful models for developers:
| Model | Size | RAM Required | Best For |
|---|---|---|---|
llama3.1:8b |
~4.7 GB | 8 GB+ | General-purpose, good all-rounder |
llama3.1:70b |
~40 GB | 48 GB+ | High-quality reasoning, complex tasks |
codellama:7b |
~3.8 GB | 8 GB+ | Code completion and generation |
mistral:7b |
~4.1 GB | 8 GB+ | Strong reasoning, fast inference |
qwen2.5-coder:7b |
~4.7 GB | 8 GB+ | Excellent code generation |
phi3:mini |
~2.3 GB | 4 GB+ | Small and fast, quick tasks |
deepseek-coder-v2:16b |
~8.9 GB | 16 GB+ | Strong code understanding |
Understanding Model Tags
Models use a tag system similar to Docker images: model:size or model:size-quantization.
# Default quantization (usually q4_0 or q4_K_M)
ollama pull llama3.1:8b
# Specific quantization
ollama pull llama3.1:8b-q4_0 # Smaller, faster, slightly less accurate
ollama pull llama3.1:8b-q8_0 # Larger, slower, more accurate
Quantization reduces model precision to shrink file size and speed up inference. The q4_0 variants use 4-bit quantization and offer the best speed/size tradeoff for most development tasks. The q8_0 variants preserve more accuracy at the cost of roughly double the memory usage.
The REST API
Ollama exposes a local HTTP API at http://localhost:11434 that any application can call. This is where things get useful for integration into development tools and scripts.
Generate Endpoint
The /api/generate endpoint takes a prompt and returns a completion:
curl http://localhost:11434/api/generate \
-d '{
"model": "llama3.1:8b",
"prompt": "Write a Python function to find the nth Fibonacci number using memoization.",
"stream": false
}'
Setting "stream": false returns the full response as a single JSON object. Set it to true (or omit it, since streaming is the default) to receive the response as a stream of JSON objects, one per token.
Chat Endpoint
The /api/chat endpoint supports multi-turn conversations with system, user, and assistant roles:
curl http://localhost:11434/api/chat \
-d '{
"model": "llama3.1:8b",
"messages": [
{"role": "system", "content": "You are a helpful coding assistant."},
{"role": "user", "content": "Explain the difference between a stack and a queue."}
],
"stream": false
}'
OpenAI-Compatible Endpoint
Ollama also exposes an OpenAI-compatible API at /v1/chat/completions. This is significant because it means any tool, library, or application that supports the OpenAI API can point at your local Ollama instance with just a base URL change:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1:8b",
"messages": [{"role": "user", "content": "Hello"}]
}'
The response format matches the OpenAI API schema, so existing parsing code works without modification.
Using Ollama from Python
For scripting and application development, the official Python library provides a clean interface.
pip install ollama
Basic Chat
import ollama
response = ollama.chat(
model="llama3.1:8b",
messages=[
{"role": "system", "content": "You are a senior Python developer."},
{"role": "user", "content": "Write a decorator that retries a function up to 3 times on exception."},
],
)
print(response["message"]["content"])
Streaming Responses
For longer generations, streaming gives you output as it's produced instead of waiting for the full response:
import ollama
stream = ollama.chat(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Explain async/await in Python."}],
stream=True,
)
for chunk in stream:
print(chunk["message"]["content"], end="", flush=True)
Using the OpenAI Python Library
Since Ollama supports the OpenAI-compatible endpoint, you can use the official OpenAI Python library. This is especially useful if you have existing code that talks to OpenAI and you want to test it against a local model:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
response = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "What is a binary search tree?"}],
)
print(response.choices[0].message.content)
The api_key parameter is required by the OpenAI client but Ollama doesn't validate it, so any non-empty string works.
Custom Modelfiles
Modelfiles let you create customized model configurations with baked-in system prompts and parameter tuning. Think of them as Dockerfiles for LLMs.
Create a file called Modelfile:
FROM llama3.1:8b
SYSTEM """You are a senior software engineer specializing in Python and TypeScript. You write clean, well-documented code. When asked to write code, include type hints and docstrings. Keep explanations concise."""
PARAMETER temperature 0.3
PARAMETER top_p 0.9
PARAMETER num_ctx 8192
Then build and run your custom model:
ollama create coding-assistant -f Modelfile
ollama run coding-assistant
Your custom model now appears in ollama list and can be used with the API just like any base model.
Key Parameters
- temperature (0.0 - 2.0): Controls randomness. Lower values (0.1-0.3) produce more deterministic, focused output -- good for code generation. Higher values (0.7-1.0) produce more creative, varied responses. Default is 0.8.
- top_p (0.0 - 1.0): Nucleus sampling threshold. At 0.9, the model considers tokens that make up the top 90% of probability mass. Lower values make output more focused.
- num_ctx: Context window size in tokens. Determines how much text the model can "see" at once. Larger values use more memory. Default varies by model, typically 2048 or 4096.
- repeat_penalty (0.0 - 2.0): Penalizes the model for repeating tokens. Values above 1.0 reduce repetition. Default is 1.1.
Custom Modelfiles let you create purpose-built assistants -- a code reviewer with strict formatting rules, a documentation writer, a SQL query generator -- without writing any application code. Just change the system prompt and parameters.
Performance Tips
Getting good performance from local LLMs is largely about memory management and knowing what knobs to turn.
Right-size your context window. The num_ctx parameter directly impacts VRAM and RAM usage. A context window of 8192 tokens uses roughly 4x the memory of 2048 tokens for the KV cache. Set it to what you actually need rather than maxing it out.
Choose the right quantization. For most development tasks, q4_K_M (often the default) offers an excellent balance of speed and quality. Use q8_0 only when you need the highest accuracy and have the memory to spare. The difference is often negligible for code generation.
Understand Apple Silicon performance. On Macs with Apple Silicon, Ollama uses Metal for GPU acceleration. Performance scales directly with unified memory bandwidth. An M1 Pro/Max with more memory bandwidth will generate tokens noticeably faster than a base M1, even at the same model size. The sweet spot for most developers is an M-series Mac with 16-32 GB of unified memory running 7-8B parameter models.
Tune concurrent requests. If you're building an application that sends multiple requests, control parallelism with the OLLAMA_NUM_PARALLEL environment variable:
OLLAMA_NUM_PARALLEL=4 ollama serve
This allows up to 4 concurrent requests, but each additional parallel request increases memory usage.
Manage model loading. Models are loaded into memory on the first request and stay loaded for 5 minutes by default. The first request after a model unloads will be slower due to loading time. If you're actively developing against a model, extend the keep-alive time:
OLLAMA_KEEP_ALIVE=30m ollama serve
Or set it to -1 to keep models loaded indefinitely until the server stops.
Monitor resource usage. Check what's loaded and how much memory it's consuming:
ollama ps
This shows all currently loaded models, their size in memory, and when they'll be unloaded.
Wrapping Up
Local LLMs have moved from a novelty to a practical tool for everyday development. With Ollama, the barrier to entry is essentially zero -- install it, pull a model, and start generating. They're ideal for prototyping AI features, analyzing private codebases, generating code without API costs, and working offline.
The OpenAI-compatible API means you can swap between local and cloud models with a single configuration change, using local models for development and iteration, then switching to a hosted model for production workloads.
If you're running multiple AI-powered workflows -- local LLMs for code generation, cloud agents for complex reasoning, test runners, dev servers -- Agents UI can help you manage those sessions side by side in a single terminal interface, keeping your AI development workflow organized as it scales.