Running Hermes Agent with Ollama can keep model inference on your own machine, but local inference is not automatically a fully offline agent. Hermes can still use browser tools, web search, remote MCP servers, cloud speech, external memory, hosted fallbacks, and messaging gateways unless you disable or isolate them.
This guide owns one question: can Hermes Agent run fully offline, and what must you verify before calling it private? For copy-paste installation commands, use the dedicated Hermes Agent Ollama setup guide.
Quick answer#
Yes. Hermes Agent can run with a local Ollama model through ollama launch hermes or the OpenAI-compatible endpoint at http://127.0.0.1:11434/v1. For a genuinely offline setup, you must also disable hosted model fallbacks and every network-capable tool or integration, keep memory local, and test the machine with outbound network access blocked.
The model needs structured tool-call support and at least a 64K context window for Hermes. A model that chats well but cannot reliably call files or terminal tools is not agent-ready. Local inference removes per-token API billing, but it moves cost and responsibility to RAM/VRAM, storage, electricity, latency, updates, and recovery. If the goal is reliable browser/mobile access without operating this stack, compare the managed FlyHermes path instead.
Local inference and fully offline are different claims#
A local model means prompts sent to that model are processed by a server you control. It does not prove the complete workflow stayed on the machine.
Hermes can reach the network through several independent paths:
- Browser automation and web search can fetch public websites.
- Remote MCP servers and API integrations can receive task data.
- Telegram, Discord, Slack, WhatsApp, Signal, and email require external networks.
- Hosted speech-to-text, text-to-speech, memory, image, or fallback providers can receive content.
- A tool-run command can call
curl, package registries, Git remotes, or any permitted endpoint.
Use the local LLM support overview to compare Ollama with LM Studio, vLLM, SGLang, and llama.cpp. Use this page to decide whether your whole workflow is actually offline.
Fastest current Ollama path#
The shortest supported route is:
ollama launch hermes
Choose a tool-capable local model and let Ollama configure Hermes. For manual control:
ollama serve
curl http://127.0.0.1:11434/api/tags
hermes model
Select the custom/local endpoint, use http://127.0.0.1:11434/v1, leave the API key empty for a local server, and choose the exact model ID returned by Ollama. Start a new Hermes session after changing providers.
Do not confuse Ollama Cloud with local Ollama. Ollama Cloud is a hosted provider and requires OLLAMA_API_KEY; it does not satisfy an offline requirement.
The 64K context requirement#
Hermes needs enough context for system instructions, tool schemas, retrieved memory, conversation state, and tool results. Current Hermes documentation requires at least 64,000 tokens and rejects smaller windows at startup.
For Ollama, run or create a model variant with a 64K-capable context, for example with -c 65536 when the model and Ollama command support it. Larger context also consumes more memory. A model that barely fits at a small context may swap, stall, or fail once Hermes loads its real tool surface.
Use the best local models for Hermes guide to shortlist candidates, but trust your own workflow acceptance test over a chat benchmark.
Prove tool use, not just conversation#
A greeting only proves text generation. Test a task that requires structured tool calls:
hermes chat -q "List the files in this directory, read README.md, and report the project name."
Accept the model only if Hermes actually invokes file or terminal tools, returns the correct project, and does not merely describe what it would do. Repeat the test at least three times, then add one representative multi-step task from your real workflow.
Record:
- Tool-call success rate.
- Time to first token and total completion time.
- Peak RAM/VRAM and whether the host swaps.
- Context failures or malformed tool arguments.
- Whether a retry changes the result unpredictably.
For a broader reliability comparison, use cloud API vs local Ollama.
Offline privacy checklist#
Before describing the setup as fully offline, verify all of these:
- The selected primary model points to a loopback or private-network endpoint you control.
- No hosted fallback provider is enabled.
- Web search, browser automation, remote MCP, cloud speech, image generation, and external memory are disabled or excluded from the profile.
- Messaging gateways are disabled; an offline machine cannot deliver to cloud chat platforms.
- Skills and cron prompts do not call remote URLs, Git hosts, package managers, or APIs.
- Secrets are local and are not copied into logs, prompts, or public files.
- The workflow still passes while outbound traffic is blocked at the host or network boundary.
- A packet/log review shows no unexpected egress during the acceptance task.
The Hermes security hardening guide explains permission and isolation controls. The Docker versus native install guide explains deployment boundaries; Docker can improve reproducibility, but it does not make a workload offline unless network policy enforces that boundary.
Docker networking changes the endpoint#
When Hermes runs inside Docker, 127.0.0.1 refers to the Hermes container, not an Ollama server running on the host.
Use http://host.docker.internal:11434/v1 on macOS or Windows, or place Hermes and Ollama on the same Compose network and address Ollama by its service name. On Linux, configure an explicit host-gateway mapping or use a shared container network.
Test /v1/models from the same runtime that runs Hermes. A successful request from the host shell does not prove the container can reach Ollama. The Docker Ollama networking guide covers the complete container path.
Local model timeouts and slow prefill#
Large local contexts can take minutes before the first token. Current Hermes versions automatically raise local stream-read tolerance and adjust stale-call behavior for detected local endpoints. Provider- and model-specific timeout settings take precedence over the legacy HERMES_API_TIMEOUT variable.
Before increasing a timeout, confirm that Ollama is still computing, the host is not swapping heavily, and the model fits the available memory. A longer timeout cannot repair an incompatible tool schema or undersized model.
Use the provider costs and rate-limit guide when the failure is actually a hosted fallback, quota, subscription, or provider-routing issue.
What local inference really costs#
Ollama avoids usage-based API charges for local inference, but it is not free:
- Hardware purchase or rental
- Electricity and cooling
- Model downloads and disk space
- RAM/VRAM reserved from other work
- Slower completion and queue time
- Model, Ollama, Hermes, and driver updates
- Monitoring and recovery when unattended jobs fail
Compare cost per accepted task, not cost per token. A hosted model can be cheaper when a smaller local model needs repeated retries or human repair. The Hermes pricing and cost breakdown includes infrastructure, provider, and operator responsibility.
When to use local Ollama#
Choose local Ollama when:
- Sensitive prompts must stay on controlled hardware.
- The workflow can operate without network tools.
- Your machine can run a tool-capable model at 64K context without destructive swapping.
- Predictable marginal inference cost matters more than speed.
- You are willing to validate models and maintain the runtime.
Choose a hosted API when model quality, low latency, large context, or elastic concurrency matters more than keeping inference local. Choose FlyHermes when the bigger problem is maintaining uptime, channels, provider configuration, backups, and remote access rather than selecting one model.
Production acceptance test#
Do not connect cron or a gateway until the CLI path passes:
- Verify Ollama health and exact model discovery.
- Confirm the model starts with at least 64K context.
- Run the file-and-terminal task three times.
- Run one real multi-step workflow and inspect every tool call.
- Block outbound network access and repeat the workflow.
- Confirm no hosted fallback, external memory, remote MCP, browser, or speech call occurred.
- Reboot the machine, restart Ollama and Hermes, and repeat one acceptance task.
- Measure latency and resource use before scheduling unattended work.
If you later enable a gateway or AI agent cron job, the system is no longer fully offline. Document the new network boundary and verify the real delivery target separately.
Bottom line#
Hermes Agent can run its core model locally through Ollama. The honest privacy claim is narrower: local model inference is straightforward; fully offline operation requires an intentionally isolated profile, no cloud integrations or fallbacks, enforced egress controls, and a real tool-use test under blocked networking.
That boundary matters more than marketing language. Prove it, record it, and re-test it after changing models, tools, skills, profiles, or deployment architecture.