Models & providers

Hermes Agent Models & Providers

Hermes Agent works with almost any model provider, but the model must handle tool calling well and have at least 64K tokens of context — Hermes rejects smaller windows at startup. There are four ways to pay for it:

  1. Nous Portal — Nous Research's account; free plan with free models, paid plans from $20/month.
  2. A subscription you already have — Claude Max (extra usage credits only), ChatGPT/Codex, GitHub Copilot, or SuperGrok.
  3. API keys — OpenRouter, Anthropic, OpenAI, Google, DeepSeek, and dozens more, billed per token.
  4. Local or self-hosted models — Ollama, LM Studio, vLLM, SGLang, or Hermes's own managed llama.cpp runtime.

Pick one with hermes model; you can switch any time and keep several configured.

Checked against official Nous Research sources on September 26, 2026 (Hermes Agent v0.21.5).

Choose your access path

PathGood whenWatch out forSet up
Nous PortalYou want one login for models plus web search, images, TTS, and browser toolsTool Gateway needs a paid plan; free plan is free models onlyhermes setup --portal
Existing subscription (OAuth)You already pay for Claude Max, ChatGPT/Codex, Copilot, or SuperGrokWhat the plan actually pays for varies — see the next tablehermes model
API keyYou want pay-per-token billing and direct provider accessAgent loops make several model calls per message; watch spendhermes model or hermes config set OPENROUTER_API_KEY …
Local / self-hostedPrivacy, offline use, or no per-token billsNeeds a tool-capable model and enough GPU/RAM for a 64K+ windowhermes model → custom endpoint

Secrets go to ~/.hermes/.env and OAuth tokens to ~/.hermes/auth.json; the chosen model is saved in ~/.hermes/config.yaml. First-time walkthrough: set up Hermes Agent.

Can Hermes use a subscription I already pay for?

From the subscription table in the official providers docs:

PlanWorks with Hermes?What gets consumedCommon surprise
Claude MaxYes — hermes model → Anthropic OAuth; requires purchased extra usage creditsOnly the extra/overage creditsYour included Max allowance is never used; all Hermes usage bills as extra usage
Claude ProNo—Use an ANTHROPIC_API_KEY instead (pay-per-token)
ChatGPT / Codex planYes — device-code login, uses Codex modelsNot documented by HermesWhich tiers are eligible and how usage counts against Codex limits is not documented
GitHub CopilotYes — OAuth via hermes model, or a GitHub tokenYour Copilot subscriptionA separate Copilot ACP mode drives the local copilot CLI instead
SuperGrok / X Premium+Yes — browser OAuthSubscription quota (documented for X Search)HTTP 403 after login means your tier lacks API access; use an XAI_API_KEY
Google AI Pro / UltraNo documented path—Gemini needs an API key or Vertex AI billing

Other OAuth sign-ins listed in the official quickstart include Qwen Portal and MiniMax. Each Bot Mode bot is its own profile, and OAuth logins are not copied between profiles — sign each one in with hermes -p <name> auth add <provider>.

Picking a model

What the official docs actually require and advise:

  • 64K context minimum. Hosted frontier models clear this easily; local servers often default lower and must be raised.
  • Tool calling matters more than chat quality. Hermes works through rapid tool-call loops — terminal, files, web, browser.
  • Nous's own Hermes 4 models are not recommended inside Hermes Agent. Nous says they are tuned for chat and reasoning, not the tool-calling loop. The Portal docs point agent work at frontier agentic models from Anthropic, OpenAI, Google, and DeepSeek instead.
  • Data-training warnings: since v0.21.0 every model picker warns when a selected model's tier trains on your data.
  • Catalogs move weekly. For example, the v0.21.5 notes (September 24, 2026) list GPT-6 Sol/Terra/Luna and Claude Opus 5.5 newly added to the Nous and OpenRouter catalogs. Use /model to see what your provider offers today.
Not tested by this site: We have not benchmarked models inside Hermes, so this page does not rank them. “Best model” lists in search results are opinions or vendor claims unless they publish a repeatable test. Try two candidates on a task you actually run, and compare cost in your provider's usage report.

Local models

  • Any OpenAI-compatible server (Ollama, LM Studio, vLLM, SGLang, llama.cpp): hermes model → custom endpoint, then enter the base URL (for Ollama, http://localhost:11434/v1) and the exact model name.
  • Raise the context window to at least 64K — for example -c 65536 for Ollama or --ctx-size 65536 for llama.cpp, per the official quickstart.
  • Managed local models: Hermes can also download and run llama.cpp for you, sizing each model to your GPU (Settings → Providers → Local Models). The docs say this Desktop interface is enabled on canary builds; other Desktop builds need the --local launch flag.
  • Local inference keeps prompts on your machine, but web search, cloud browser, and messaging tools still reach the internet if you enable them.

Hardware needs depend on the model, quantization, and context length; there is no single official sizing figure. Per-runtime context and tool-calling flags, and what still uses the network: Hermes local models.

Auxiliary models: cheaper side jobs

Besides the main model, Hermes runs side jobs — context compression, vision, web-page extraction, approval scoring, session titles, MCP tool routing, skill search, and more. Each task defaults to auto, which means your main model. Pointing the summarization-style jobs at a fast, cheap model is the most common override.

Interactive
bash
$hermes model
# → "Configure auxiliary models..." → pick a task, provider, model
~/.hermes/config.yaml (shape only — fill in current model IDs)
yaml
model:
provider: openrouter
default: <your-main-model>
auxiliary:
compression:
provider: openrouter
model: <a-fast-cheap-model>
title_generation:
prefer_fast_model: true # let Hermes pick the provider's fast tier

Mid-session model switches can cost more: the next turn may re-process the conversation at uncached input rates, and Hermes warns when a switch will trigger context compression against a smaller window. Switching between Desktop and the TUI no longer rebuilds the system prompt (fixed in v0.21.2).

For what each option costs per month, see Hermes cost.

Primary sources

HermesAgentAI.org is an independent educational documentation resource and community guide. It is not affiliated with, sponsored by, or endorsed by Nous Research or FlyHermes. Hermes Agent is released under the MIT License by Nous Research.

Where to go next