Local inference reference

Hermes Agent with Local Models

Hermes Agent can use any local server that speaks the OpenAI API — Ollama, LM Studio, vLLM, SGLang, or llama.cpp's llama-server — through hermes model. Hermes Desktop can also download and run llama.cpp models for you. Two requirements decide whether it works:

  1. At least 64,000 tokens of context. Hermes refuses smaller windows at startup, and most local servers default lower.
  2. Working tool calls. The model must support them and the server must have tool calling enabled; otherwise Hermes can chat but not act.

A local model keeps inference on your hardware. It does not make the rest of Hermes offline: web search, browser tools, messaging platforms, and remote MCP servers still use the internet when you enable them.

Checked against official Nous Research sources on September 26, 2026 (Hermes Agent v0.21.5).

What “local” does and doesn't cover

Part of HermesWith a local modelNotes
Model inferenceOn your machineAfter the model is downloaded
Sessions, memory, skills, configOn your machineStored in ~/.hermes regardless of provider; the FAQ says Hermes collects no telemetry
Auxiliary tasks (titles, compression, vision)Local by defaultThey default to your main model; an override such as auxiliary.compression.provider: openrouter sends them out
Web search and extractionNetworkGoes to the configured search backend; SearXNG is one self-hostable option
Browser toolsNetworkFetches real web pages; cloud browser providers run remotely
Telegram, Discord, other messagingNetworkMessages pass through the platform's servers
MCP servers and connectorsDepends on the serverRemote servers are network calls; local stdio servers do whatever their code does
Text-to-speechNetwork by defaultDefault provider is edge; piper, kittentts, and neutts run locally
Speech-to-textLocal by defaultstt.provider: local (faster-whisper, installed separately)
fallback_providersNetwork if setA cloud fallback receives the conversation whenever the local model fails
Sources disagree: The official Local Models page says “nothing leaves your computer”, and the Ollama guide promises “no data leaving your machine” — while the same guide shows web browsing and a Telegram bot. We read those statements as describing inference only. For a setup that sends nothing out, disable the web, browser, messaging, remote-MCP, and cloud-TTS tools (hermes tools) and don't configure a cloud fallback.

Supported local runtimes

RuntimeConnect from HermesRaise context to 64K+Enable tool calling
Hermes Desktop managed llama.cppSettings → Providers → Local ModelsAutomatic: recommended models get at least 64KHandled by Hermes
OllamaCustom endpoint http://localhost:11434/v1OLLAMA_CONTEXT_LENGTH=64000 or a Modelfile with num_ctxOn by default; the model must support tools
LM Studiohermes model → LM Studio (:1234)Model settings or lms load … --context-length 64000, then reloadLM Studio 0.3.6+ and a tool-capable model
vLLMCustom endpoint http://localhost:8000/v1--max-model-len 65536--enable-auto-tool-choice --tool-call-parser <parser>
SGLangCustom endpoint http://localhost:30000/v1--context-length 65536--tool-call-parser <parser>
llama.cpp llama-serverCustom endpoint http://localhost:8080/v1-c 64000--jinja (required)

The official Mac guide also covers MLX through omlx. Any other OpenAI-compatible server works the same way: base URL, model name, and an API key only if the server requires one.

The 64K context requirement

Hermes's system prompt and tool definitions alone use roughly 4K–8K tokens before your conversation starts. The official docs set 64,000 tokens as the minimum and say smaller windows are rejected at startup, with an error naming the window the server reported and how to raise it.

Set and verify the window (official provider docs)
bash
# Ollama: server-wide
$OLLAMA_CONTEXT_LENGTH=64000 ollama serve
$ollama ps # CONTEXT column should show 64000
$
# llama.cpp
$./llama-server -m model.gguf --jinja -c 64000 --port 8080
$curl -s http://localhost:8080/props | jq '.default_generation_settings.n_ctx'
$
# If Hermes misreads the window, state it in ~/.hermes/config.yaml:
# model:
# context_length: 64000
  • Ollama can't be fixed from the client. Context length can't be set through its OpenAI-compatible API; it must be set on the server or in a Modelfile. The docs call this the most common point of confusion.
  • llama.cpp parallel slots split the window. With -np, the total context is divided between slots.
  • More context costs memory. A model that fits at 8K may not fit at 64K; the managed Desktop runtime starts at a window that fits your GPU and grows it as needed.
Sources disagree: The official docs disagree about Ollama's default. The Ollama guide says 2,048 tokens; the providers page says it depends on VRAM: 4,096 below 24 GB, 32,768 at 24–48 GB, and 256,000 at 48 GB or more. Either way, most consumer GPUs start below Hermes's minimum. Check ollama ps rather than trusting a default.

Tool calling: the usual reason a local setup “doesn't do anything”

Hermes works by calling tools: terminal, files, web, browser. If the reply contains raw JSON such as {"name": "web_search", ...} instead of an actual search, the server is usually not parsing tool calls:

ServerFix from the official docs
llama.cppAdd --jinja
vLLM--enable-auto-tool-choice --tool-call-parser hermes (or the parser for your model family)
SGLang--tool-call-parser qwen (or the right parser)
OllamaEnabled by default; confirm the model supports tools with ollama show <model>
LM StudioUpdate to 0.3.6+ and use a model shown with the tool badge
  • Slow first reply: the fixed prompt plus tool schemas must be processed before anything is generated, which can take minutes on CPU. The docs suggest keeping the model loaded, HERMES_API_TIMEOUT=1800 in ~/.hermes/.env, and trimming unused toolsets (hermes prompt-size, hermes tools).
  • SGLang cut-offs: SGLang defaults to 128 output tokens; raise it on the server. Hermes has no output-cap setting.
  • WSL2 with a server on Windows: localhost doesn't reach Windows in WSL2's default network mode; the providers docs cover mirrored networking and host-IP fixes.

Ollama, LM Studio, and vLLM in practice

Connect Hermes to a local server
bash
$hermes model
# → "Custom endpoint (self-hosted / VLLM / etc.)"
# URL: http://localhost:11434/v1 (Ollama) | :8000/v1 (vLLM) | :8080/v1 (llama-server)
# API key: leave blank unless your server requires one
# Model: the exact name your server serves
$
# LM Studio has its own entry:
$hermes model # → "LM Studio" → Enter for http://localhost:1234/v1 → pick a discovered model
  • Ollama's own docs also offer ollama launch hermes, which installs and configures Hermes against Ollama's local endpoint. We have not compared its result with a manual setup.
  • Keep the model loaded for bots and scheduled jobs: Ollama unloads idle models after 5 minutes (OLLAMA_KEEP_ALIVE=24h changes that).
  • LM Studio: Hermes preloads the model by default; hermes config set model.lmstudio_load_mode jit hands loading to LM Studio's just-in-time mode.
  • vLLM reads the model's full context by default and errors if it doesn't fit your GPU; set --max-model-len.
  • Hermes running in Docker reaches a host server at host.docker.internal (macOS/Windows), not localhost.

Desktop's managed local models

Hermes Desktop can install llama.cpp, download models, and manage memory with no manual flags: Settings → Providers → Local Models, Install runtime, Download a model, Use. Each catalog model shows whether it fits your GPU, spills into system RAM, or is too big for the machine, and Hermes picks the highest-quality build that fits (never below 4-bit). The model server starts and stops with Hermes, and idle models unload after 15 minutes.

  • Availability: the docs say the interface is enabled on canary builds and other Desktop builds need the --local launch flag.
  • Backends: Metal or CPU on macOS; CUDA, Vulkan, or CPU on Windows; Vulkan or CPU on Linux (no prebuilt CUDA there); HIP/ROCm by explicit choice.
  • Bring your own: a running llama-server on :8080 is detected and used; other ports go in local_runtime.detect_ports. You can search Hugging Face or link an existing .gguf.
  • Selecting a local model sets model.provider: llamacpp; /model shows it as Local.

The Hermes 4 mix-up

“Hermes” names both this agent and Nous Research's Hermes 4 model family, and several ranking guides recommend Hermes 4 models as the model to run inside Hermes Agent. Nous's own Portal docs say the opposite: Hermes 4 is tuned for chat and reasoning, not the rapid tool-calling loop the agent relies on, and is not recommended for use inside Hermes Agent. Use Hermes 4 for chat and research; for the agent, pick a model with strong tool calling.

This page doesn't rank local models; we haven't benchmarked them. The official Ollama guide uses gemma4:31b and calls it “currently the best local option with tool-call support”, and Ollama's own Hermes page lists gemma4 and qwen3.6; treat both as those publishers' views and check the Ollama library for current tool-capable models. Hardware figures from the docs: the managed runtime says 8 GB+ of GPU memory runs its small catalog models and 16 GB+ runs 27–35B models; the Ollama guide lists 32 GB+ of RAM for 27B+ models.

Local models with cron and bots

  • Scheduled jobs use your main model at fire time unless pinned, so switching your main model to or from a local one moves every unpinned cron job with it.
  • Fallbacks can leave your machine. An unpinned job whose local model fails will use your fallback_providers chain if you configured one — a cloud provider. Pinned jobs never fall back.
  • The model server has to be up when the job fires, as does the machine. Ollama's 5-minute and Desktop's 15-minute idle unloads mean the first scheduled run after a quiet period pays a reload.
  • Open bug: local Ollama models producing fake tool-call text instead of real tool calls in cron runs via custom providers (#123398, filed September 26, 2026).
Not tested by this site: This page reconciles the official Local Models, providers, Ollama, Mac, and Portal docs with Ollama's integration page. We have not run Hermes on these runtimes ourselves, so we make no claims about speed, quality, or which hardware is enough for your workload.

Primary sources

HermesAgentAI.org is an independent educational documentation resource and community guide. It is not affiliated with, sponsored by, or endorsed by Nous Research or FlyHermes. Hermes Agent is released under the MIT License by Nous Research.

Where to go next