Your AI assistant. Your rules. Running on your machine.
Why Ollama Agent? • The Ollama Advantage • Quick Start • Key Features • Cheat Sheet • Full Documentation ↗
Ollama Agent is an autonomous, local-first AI assistant that runs entirely on your hardware through Ollama. It's a general-purpose agent — not a coding assistant, not a chatbot, not a narrow tool — but whatever you need it to be. Write research reports, manage files, analyze documents, automate workflows, query knowledge bases, or yes, write code too. Every prompt that shapes its behavior is an editable file on your disk, so you stay in control of what the agent is and how it thinks.
📖 Comprehensive Guides & Technical Reference: Visit the official documentation site at arrase.github.io/ollama-agent.
Most open-source AI agents are coding assistants in disguise. They ship with hardcoded system prompts about writing code, generating tests, and refactoring functions. If you want to use them for anything else, you're fighting their DNA.
Ollama Agent is different. It ships as a blank canvas with general-purpose defaults, and gives you the tools to make it yours:
The system prompt that drives Ollama Agent is a plain Jinja2 template sitting at ~/.ollama-agent/prompts/instructions.md. Open it, rewrite it, and the agent becomes whatever you need:
~/.ollama-agent/
├── prompts/
│ └── instructions.md ← The agent's brain. Edit freely.
├── MEMORY.md ← Persistent cross-session memory
├── settings.yaml ← Model, context, sampling, subagents
├── tasks/ ← Reusable YAML prompt templates
├── skills/ ← Modular procedural workflows
└── mcp.json ← External tool server connections
- Research analyst? Rewrite the prompt to focus on source evaluation, citation formatting, and document synthesis.
- System administrator? Shape it around infrastructure monitoring, log analysis, and runbook execution.
- Creative writer? Tune it for narrative structure, world-building, and stylistic consistency.
- Personal assistant? Optimize it for calendar planning, email drafting, and task management.
- Software engineer? Sure, that works too — but it's your choice, not the default assumption.
The template has access to the full configuration context via Jinja2 variables (runtime, model, rag, settings), so your instructions can adapt dynamically to runtime conditions. Made a mess? ollama-agent --config-reset system-prompt restores the defaults instantly.
The entire UI speaks your language. Ollama Agent ships with 15 built-in locales (English, Spanish, French, German, Japanese, Chinese, Korean, Arabic, Hindi, Italian, Dutch, Polish, Portuguese, Russian, Turkish, Ukrainian) and auto-detects your system locale. Override anytime with -l <lang> or runtime.language in settings.
Most agents treat Ollama as a dumb OpenAI-compatible proxy. They send requests to /v1/chat/completions, cross their fingers, and wonder why the output is truncated, the context is wrong, and the model ignores tool calls. Ollama Agent talks directly to Ollama's native API and is engineered to squeeze every capability out of your local models:
| What goes wrong | Generic OpenAI-proxy agents | Ollama Agent |
|---|---|---|
| Context window | Default to Ollama's 2K–4K num_ctx; large prompts get silently truncated. |
Auto-detects model capacity from GGUF metadata and sets num_ctx dynamically. Use /context max for full range. |
| Sampling parameters | Force fixed defaults (temp=0.7, top_p=1.0), ignoring the Modelfile. |
Auto-discovers optimal sampling (temperature, top_p, top_k, min_p, repeat_penalty) from the Modelfile. |
| Token counting | Approximate with tiktoken — wrong tokenizer for Llama, Qwen, Gemma, DeepSeek. |
Reads native server metrics (prompt_eval_count + eval_count) for exact real-time tracking. |
| Context overflow | Crash with obscure errors when the conversation exceeds the limit. | Auto-compaction at 85% capacity: summarizes older turns, prunes tool output, offloads history to disk. |
| Reasoning traces | Leak raw <think> tokens into output or fail to configure thinking effort. |
API-driven thinking controls: dynamically discovers supported effort levels (thinking.values) and defaults via /api/show, rendering traces into collapsible UI blocks. |
| Model compatibility | Blindly attempt tool calls on models that don't support them; fail with cryptic errors. | Pre-flight capability check: verifies tools support, offers interactive model selector, hot-swaps models mid-session (/model set). |
- Python 3.11+
- Ollama running locally (or reachable via network)
- A local tool-calling model (e.g.
ollama pull qwen3.8:27b,ollama pull qwen2.5:14b, orollama pull llama3.1:8b)
Install in an isolated environment using pipx (recommended):
pipx install git+https://github.com/arrase/ollama-agent.gitOr via pip into an active virtual environment:
pip install git+https://github.com/arrase/ollama-agent.gitollama-agentIf no model is configured, Ollama Agent will automatically detect your downloaded models and let you choose one interactively.
# Quick query
ollama-agent -p "Summarize the key findings in @report.pdf"
# Autonomous research with YOLO mode (auto-approves tool actions)
ollama-agent -m "qwen3.8:27b" -e "high" -y -p "Analyze the logs in @/var/log/syslog and report anomalies."The interactive terminal interface is built with Textual and Rich to provide a fast, keyboard-first environment:
● ollama-agent │ Model: qwen3.8:27b │ Context: 3.4k/32.0k (11%) │ Effort: high │ YOLO: OFF │ STEALTH: OFF
- Non-Blocking Prompt Queue: Never wait for generation to finish. Type follow-up prompts or execute read-only slash commands while the model is actively streaming or waiting for tool confirmation.
-
Multiline Editing: Type
\followed byEnter(\ + Enter) to create clean newlines. The input area expands dynamically up to 8 lines. -
3-Level Tab Autocompletion: Autocompletes slash commands, subcommands, entities (models, sessions, tasks, skills, RAG databases), and
@-mentionfile paths. -
Real-Time Token Gauge: Color-coded header indicator (Cyan
$\le 75%$ , Amber$76%-90%$ , Red$>90%$ ) based on real Ollama server token counts. -
Mid-Session Switching: Switch models (
/model set <name>), change reasoning effort (/effort high), or update context window (/context max) mid-conversation without losing thread state.
- Action Approval Dialog: Sensitive tool operations (running shell commands, writing/editing files) require keyboard confirmation (
yapprove,nreject,aallow for session,ccancel). - YOLO Mode (
-y//yolo): Bypass confirmation pauses for fully autonomous agent runs. - Stealth Mode (
-s//stealth): Run conversations in-memory without saving conversation turns or checkpoints to SQLite history.
Reference local files or folders directly in your prompts with autocompletion:
- Single & Quoted Files:
@src/main.py,@"data/financial report.csv" - Directory Traversal:
@srcor@.(recursively attaches all supported source files). - Multimodal Assets: Automatically base64-encodes images (
.png,.jpg,.webp), audio, video, and documents (.pdf,.pptx) for vision-enabled models.
- Repository Guidelines (
AGENTS.md): Automatically discovered in the working directory up to the git root and mounted into agent memory. - Cross-Session Memory (
MEMORY.md): Preserves user preferences and architectural decisions across sessions. - Episodic Memory: Autonomous past conversation search via the
search_past_conversationstool and user search via/session search <query>.
- Saved Tasks: Reusable YAML prompt templates with Jinja2 expressions (
~/.ollama-agent/tasks/), input type validation, and CLI execution (ollama-agent task run <id>). - Agent Skills: Modular procedural workflows adhering to the open Agent Skills specification.
- Local RAG Engine: Embed and index documents into local Qdrant collections using Ollama embeddings (
ollama pull nomic-embed-text), retrieved automatically via therag_searchtool.
- MCP Extensibility: Connect external tools over
stdio,http, andssetransports declared in~/.ollama-agent/mcp.json. - Specialized Subagents: Configure isolated subagents in
settings.yamlwith their own model, system prompt, context window, and dedicated MCP tool servers.
| Command | Usage | Description |
|---|---|---|
/model |
/model [list | set <name>] |
List local models with tool support or switch active model. |
/context |
/context [<size | max>] |
Inspect or change context window (num_ctx) on the fly. |
/params |
/params [list | set <param> <val>] |
View effective parameters and resolution sources, or update sampling values. |
/effort |
/effort [<level>] |
Inspect or set reasoning effort (dynamically validated against active model values, or boolean toggles). |
/queue |
/queue [list | clear | rm <pos>] |
Inspect and manage pending prompts in the FIFO execution queue. |
/session |
/session [list | resume <id> | new] |
Manage chat threads, resume past conversations, or start fresh. |
/task |
/task [list | run <id> | create] |
List, run with variables (key=val), or create saved YAML tasks. |
/skill |
/skill [list | show <id> | create] |
Inspect, create, or manage modular Agent Skills. |
/rag |
/rag [status | list | load <db>] |
Inspect status, list vector databases, or attach knowledge bases. |
/agents |
/agents [list] |
List specialized subagents and their assigned models/tools. |
/mcp |
/mcp [list | reload] |
Check MCP server connection health or reload tool definitions live. |
/yolo |
/yolo [on | off] |
Toggle confirmation bypass for autonomous execution. |
/stealth |
/stealth [on | off] |
Toggle ephemeral mode without saving to SQLite history. |
ollama-agent -m <model> # Specify Ollama model
ollama-agent -p "<prompt>" # Run in non-interactive single-shot mode
ollama-agent -y # Run in YOLO mode (bypass tool approvals)
ollama-agent -s # Run in Stealth mode (in-memory only)
ollama-agent -c <num|max> # Set context window size (tokens or 'max')
ollama-agent -e <effort> # Set reasoning effort level
ollama-agent --rag <db> # Preload a RAG vector collection
ollama-agent -l <lang> # Override interface language (e.g. en, es, fr, de, ja, zh)For complete architecture diagrams, configuration manuals, and development guides, visit our Documentation Site:
- 📖 CLI & REPL User Guide — Full terminal navigation, slash commands, multiline inputs, and scriptable subcommands.
- 🧩 Agent Skills — Modular procedural workflows adhering to the open Agent Skills standard.
- 📋 Saved Tasks — Reusable Jinja2 prompt automation routines with typed inputs.
- 🔌 Model Context Protocol (MCP) — External tool servers, stdio/SSE transports, and live reloading.
- 🤖 Specialized Subagents — Isolated subagent graphs, dedicated models, and exclusive MCP tools.
- 🧠 Memory & Guidelines —
AGENTS.mdproject rules,MEMORY.mduser preferences, SQLite sessions, and episodic memory. - 📚 Local RAG Guide — Qdrant vector store management, Ollama embeddings, chunking, and semantic search.
- ⚙️ Configuration Reference — Complete
settings.yamlschema, parameter precedence, context resolution, and LangSmith. - 🏗️ System Architecture — DeepAgents orchestration, SQLite checkpoints, streaming parsers, and compaction engine.
- 🛠️ Developer Guide — Contributing guidelines, local environment setup, and test suite execution.
# Clone the repository
git clone https://github.com/arrase/ollama-agent.git
cd ollama-agent
# Create virtual environment and install in editable mode
python -m venv .venv
source .venv/bin/activate
pip install -e .
# Run test suite
.venv/bin/python -m unittest discover -s testsFor detailed contributing standards, linting with Ruff, and architectural breakdowns, see the Developer Guide.
This project is licensed under the MIT License.