Hermes Agent with Local Models
Hermes Agent can use any local server that speaks the OpenAI API — Ollama, LM Studio, vLLM, SGLang, or llama.cpp's llama-server — through hermes model. Hermes Desktop can also download and run llama.cpp models for you. Two requirements decide whether it works:
- At least 64,000 tokens of context. Hermes refuses smaller windows at startup, and most local servers default lower.
- Working tool calls. The model must support them and the server must have tool calling enabled; otherwise Hermes can chat but not act.
A local model keeps inference on your hardware. It does not make the rest of Hermes offline: web search, browser tools, messaging platforms, and remote MCP servers still use the internet when you enable them.
Checked against official Nous Research sources on September 26, 2026 (Hermes Agent v0.21.5).
What “local” does and doesn't cover
| Part of Hermes | With a local model | Notes |
|---|---|---|
| Model inference | On your machine | After the model is downloaded |
| Sessions, memory, skills, config | On your machine | Stored in ~/.hermes regardless of provider; the FAQ says Hermes collects no telemetry |
| Auxiliary tasks (titles, compression, vision) | Local by default | They default to your main model; an override such as auxiliary.compression.provider: openrouter sends them out |
| Web search and extraction | Network | Goes to the configured search backend; SearXNG is one self-hostable option |
| Browser tools | Network | Fetches real web pages; cloud browser providers run remotely |
| Telegram, Discord, other messaging | Network | Messages pass through the platform's servers |
| MCP servers and connectors | Depends on the server | Remote servers are network calls; local stdio servers do whatever their code does |
| Text-to-speech | Network by default | Default provider is edge; piper, kittentts, and neutts run locally |
| Speech-to-text | Local by default | stt.provider: local (faster-whisper, installed separately) |
fallback_providers | Network if set | A cloud fallback receives the conversation whenever the local model fails |
hermes tools) and don't configure a cloud fallback.Supported local runtimes
| Runtime | Connect from Hermes | Raise context to 64K+ | Enable tool calling |
|---|---|---|---|
| Hermes Desktop managed llama.cpp | Settings → Providers → Local Models | Automatic: recommended models get at least 64K | Handled by Hermes |
| Ollama | Custom endpoint http://localhost:11434/v1 | OLLAMA_CONTEXT_LENGTH=64000 or a Modelfile with num_ctx | On by default; the model must support tools |
| LM Studio | hermes model → LM Studio (:1234) | Model settings or lms load … --context-length 64000, then reload | LM Studio 0.3.6+ and a tool-capable model |
| vLLM | Custom endpoint http://localhost:8000/v1 | --max-model-len 65536 | --enable-auto-tool-choice --tool-call-parser <parser> |
| SGLang | Custom endpoint http://localhost:30000/v1 | --context-length 65536 | --tool-call-parser <parser> |
| llama.cpp llama-server | Custom endpoint http://localhost:8080/v1 | -c 64000 | --jinja (required) |
The official Mac guide also covers MLX through omlx. Any other OpenAI-compatible server works the same way: base URL, model name, and an API key only if the server requires one.
The 64K context requirement
Hermes's system prompt and tool definitions alone use roughly 4K–8K tokens before your conversation starts. The official docs set 64,000 tokens as the minimum and say smaller windows are rejected at startup, with an error naming the window the server reported and how to raise it.
# Ollama: server-wide$OLLAMA_CONTEXT_LENGTH=64000 ollama serve$ollama ps # CONTEXT column should show 64000$# llama.cpp$./llama-server -m model.gguf --jinja -c 64000 --port 8080$curl -s http://localhost:8080/props | jq '.default_generation_settings.n_ctx'$# If Hermes misreads the window, state it in ~/.hermes/config.yaml:# model:# context_length: 64000
- Ollama can't be fixed from the client. Context length can't be set through its OpenAI-compatible API; it must be set on the server or in a Modelfile. The docs call this the most common point of confusion.
- llama.cpp parallel slots split the window. With
-np, the total context is divided between slots. - More context costs memory. A model that fits at 8K may not fit at 64K; the managed Desktop runtime starts at a window that fits your GPU and grows it as needed.
ollama ps rather than trusting a default.Tool calling: the usual reason a local setup “doesn't do anything”
Hermes works by calling tools: terminal, files, web, browser. If the reply contains raw JSON such as {"name": "web_search", ...} instead of an actual search, the server is usually not parsing tool calls:
| Server | Fix from the official docs |
|---|---|
| llama.cpp | Add --jinja |
| vLLM | --enable-auto-tool-choice --tool-call-parser hermes (or the parser for your model family) |
| SGLang | --tool-call-parser qwen (or the right parser) |
| Ollama | Enabled by default; confirm the model supports tools with ollama show <model> |
| LM Studio | Update to 0.3.6+ and use a model shown with the tool badge |
- Slow first reply: the fixed prompt plus tool schemas must be processed before anything is generated, which can take minutes on CPU. The docs suggest keeping the model loaded,
HERMES_API_TIMEOUT=1800in~/.hermes/.env, and trimming unused toolsets (hermes prompt-size,hermes tools). - SGLang cut-offs: SGLang defaults to 128 output tokens; raise it on the server. Hermes has no output-cap setting.
- WSL2 with a server on Windows:
localhostdoesn't reach Windows in WSL2's default network mode; the providers docs cover mirrored networking and host-IP fixes.
Ollama, LM Studio, and vLLM in practice
$hermes model# → "Custom endpoint (self-hosted / VLLM / etc.)"# URL: http://localhost:11434/v1 (Ollama) | :8000/v1 (vLLM) | :8080/v1 (llama-server)# API key: leave blank unless your server requires one# Model: the exact name your server serves$# LM Studio has its own entry:$hermes model # → "LM Studio" → Enter for http://localhost:1234/v1 → pick a discovered model
- Ollama's own docs also offer
ollama launch hermes, which installs and configures Hermes against Ollama's local endpoint. We have not compared its result with a manual setup. - Keep the model loaded for bots and scheduled jobs: Ollama unloads idle models after 5 minutes (
OLLAMA_KEEP_ALIVE=24hchanges that). - LM Studio: Hermes preloads the model by default;
hermes config set model.lmstudio_load_mode jithands loading to LM Studio's just-in-time mode. - vLLM reads the model's full context by default and errors if it doesn't fit your GPU; set
--max-model-len. - Hermes running in Docker reaches a host server at
host.docker.internal(macOS/Windows), notlocalhost.
Desktop's managed local models
Hermes Desktop can install llama.cpp, download models, and manage memory with no manual flags: Settings → Providers → Local Models, Install runtime, Download a model, Use. Each catalog model shows whether it fits your GPU, spills into system RAM, or is too big for the machine, and Hermes picks the highest-quality build that fits (never below 4-bit). The model server starts and stops with Hermes, and idle models unload after 15 minutes.
- Availability: the docs say the interface is enabled on canary builds and other Desktop builds need the
--locallaunch flag. - Backends: Metal or CPU on macOS; CUDA, Vulkan, or CPU on Windows; Vulkan or CPU on Linux (no prebuilt CUDA there); HIP/ROCm by explicit choice.
- Bring your own: a running
llama-serveron:8080is detected and used; other ports go inlocal_runtime.detect_ports. You can search Hugging Face or link an existing.gguf. - Selecting a local model sets
model.provider: llamacpp;/modelshows it as Local.
The Hermes 4 mix-up
“Hermes” names both this agent and Nous Research's Hermes 4 model family, and several ranking guides recommend Hermes 4 models as the model to run inside Hermes Agent. Nous's own Portal docs say the opposite: Hermes 4 is tuned for chat and reasoning, not the rapid tool-calling loop the agent relies on, and is not recommended for use inside Hermes Agent. Use Hermes 4 for chat and research; for the agent, pick a model with strong tool calling.
This page doesn't rank local models; we haven't benchmarked them. The official Ollama guide uses gemma4:31b and calls it “currently the best local option with tool-call support”, and Ollama's own Hermes page lists gemma4 and qwen3.6; treat both as those publishers' views and check the Ollama library for current tool-capable models. Hardware figures from the docs: the managed runtime says 8 GB+ of GPU memory runs its small catalog models and 16 GB+ runs 27–35B models; the Ollama guide lists 32 GB+ of RAM for 27B+ models.
Local models with cron and bots
- Scheduled jobs use your main model at fire time unless pinned, so switching your main model to or from a local one moves every unpinned cron job with it.
- Fallbacks can leave your machine. An unpinned job whose local model fails will use your
fallback_providerschain if you configured one — a cloud provider. Pinned jobs never fall back. - The model server has to be up when the job fires, as does the machine. Ollama's 5-minute and Desktop's 15-minute idle unloads mean the first scheduled run after a quiet period pays a reload.
- Open bug: local Ollama models producing fake tool-call text instead of real tool calls in cron runs via custom providers (#123398, filed September 26, 2026).
Primary sources
- Local Models (Desktop managed llama.cpp) — official docs
- LLM and model providers: Ollama, vLLM, SGLang, llama.cpp, LM Studio — official docs
- Run Hermes Locally with Ollama — official guide
- Run Local LLMs on Mac — official guide
- Nous Portal: a note on Hermes 4 — official docs
- FAQ: is my data sent anywhere? — official docs
- Configuration: auxiliary models, TTS, STT — official docs
- Ollama docs — Hermes Agent integration
HermesAgentAI.org is an independent educational documentation resource and community guide. It is not affiliated with, sponsored by, or endorsed by Nous Research or FlyHermes. Hermes Agent is released under the MIT License by Nous Research.