Hermes Agent Memory: What It Remembers, Why It Forgets & How to Fix It
Unlike stateless LLMs that forget everything between sessions, Hermes Agent maintains a multi-layered memory architecture. It separates instant, curated facts from historical conversation archives and reusable procedural skills.
- Personal preferences: Style, tool choices, constraints (
USER.md) - Environment facts: OS quirks, server configs, conventions (
MEMORY.md) - Session transcripts: Historical chats stored in local SQLite (
state.db) - Learned skills: Executable procedural playbooks in
~/.hermes/skills/
- Mid-session updates: Prompt is frozen; new writes load on next session start
- Raw chat logs: Not auto-injected into prompt; searched on-demand via FTS5
- Exceeded capacity: Writes rejected when character limits are full (no silent drops)
- Other profiles: Memory is strictly isolated per Hermes profile
Why Did Hermes Forget? Find Your Symptom
Most "memory amnesia" issues in Hermes Agent trace to how different memory layers are buffered, frozen, or isolated. Match your observed problem below to identify the underlying cause and resolution:
“Hermes learned a preference mid-session but ignored it in subsequent turns”
Why it happens: Hermes injects MEMORY.md and USER.md as a frozen snapshot at session start to preserve LLM prefix caching. When memory updates occur mid-session, they write to disk immediately but do not rewrite the active session prompt.
“My Telegram (or Discord) bot saved a fact but keeps acting like it doesn't know it”
Why it happens: On messaging platforms a chat is deliberately one continuous session that survives gateway restarts and reboots. Because the memory snapshot is frozen at session start, a fact saved in that chat is not injected until a new session begins.
“Hermes remembers project architecture, but forgot a detailed conversation from last week”
Why it happens: Persistent memory (MEMORY.md) is intentionally compact (~1,300 tokens total) and automatically injected into every prompt. Past chat transcripts are archived in SQLite (~/.hermes/state.db) and searched on-demand via FTS5, never dumped raw into active context.
“Memory write fails with an error or Hermes says memory is full”
Why it happens: Hermes does not silently drop or auto-truncate memories. When an entry exceeds the default character limit (2,200 chars for MEMORY.md, 1,375 for USER.md), the memory tool rejects the write with an error.
“I told Hermes a preference, but it never saved it across restarts”
Why it happens: Either the fact was treated as conversational ephemera without calling the memory tool, or `write_approval` is enabled in config.yaml, leaving the memory change in a pending review queue.
“Hermes learned a complex script or CLI procedure, but forgot it in a new folder”
Why it happens: Executable procedures, debugging playbooks, and multi-step tool workflows belong in Skills (~/.hermes/skills/), not in MEMORY.md. Factual memory only stores concise reference notes.
“Hermes forgot everything when launched in a new terminal or project”
Why it happens: Memory is scoped per profile (~/.hermes/profiles/<name>/) or custom $HERMES_HOME. Additionally, project-specific rules are loaded from local .hermes.md or AGENTS.md, not global memory.
Hermes Memory Layers: What Lives Where
Hermes prevents token bloat and hallucination by categorizing data across distinct storage systems. Never confuse curated factual memory with raw conversation history or project rules:
| Memory Layer | What Belongs Here | Auto Injected? | Default Budget | Common Forgetting Symptom |
|---|---|---|---|---|
Agent Notes (MEMORY.md) ~/.hermes/memories/MEMORY.md | Environment facts, OS conventions, tool quirks, discovered server configs, lessons learned. | Yes (At session start) | 2,200 chars (~800 tokens default) | Fails to save when full; doesn't appear mid-session until next restart. |
User Profile (USER.md) ~/.hermes/memories/USER.md | User identity, name, role, coding style preferences, timezone, communication constraints. | Yes (At session start) | 1,375 chars (~500 tokens default) | Ignores coding style if omitted; stays empty if you put facts in SOUL.md instead. |
Session History (SQLite FTS5) ~/.hermes/state.db | Complete raw transcripts of every CLI, messaging, and desktop conversation turn. | No (Retrieved on-demand) | Unlimited (Disk-bounded) | Seems forgotten in fresh chats unless agent is prompted to search past sessions. |
Skills Engine (SKILL.md) ~/.hermes/skills/<name>/SKILL.md | Procedural workflows, custom commands, reusable multi-step scripts (agentskills.io standard). | Progressive (Names visible; full body loaded on invocation) | On-demand execution | Cannot run procedure if stored as unstructured text in MEMORY.md instead of a skill. |
Project Context (AGENTS.md / .hermes.md) Project root / working directory | Repository instructions, build commands, test suites, architecture notes specific to a codebase. | Yes (When active in repo) | Dynamic cap (scales with context window) | Not loaded when executing Hermes outside that repository's directory tree. |
A quick rule of thumb from official Nous documentation:
- SOUL.md is who the agent is (tone, voice, demeanor) — you edit it directly.
- USER.md is who you are (your workflow, style, constraints) — the agent curates it.
- MEMORY.md is what the agent has learned (environment facts, conventions) — the agent curates it.
- AGENTS.md / .hermes.md is what the project requires (repo commands, test steps) — codebase-scoped.
- SKILL.md is how to perform a reusable workflow — procedural instructions.
Why Hermes May Appear to Forget (The Real Causes)
When users assume Hermes "lost" information, it is almost always operating under one of six intentional design constraints:
The Frozen Snapshot Pattern (Active Session Invariance)
When a session starts, Hermes takes a static snapshot of MEMORY.md and USER.md and locks them into the system prompt. If the agent saves a new fact during turn 3, that entry is written to disk immediately, but the system prompt for that conversation is never modified mid-session. This preserves the LLM prefix cache, saving substantial latency and API token cost. The new fact becomes active in the system prompt the moment you start your next session.
Conversation History vs. Persistent Memory
Users often expect an AI assistant to recall every turn from every past chat automatically. If Hermes injected your entire conversational history into every prompt, your context window would quickly overflow and API costs would spiral. Instead, conversations are saved in SQLite (~/.hermes/state.db). Hermes only knows a historical conversation detail if you prompt it to search (triggering its session_search tool) or if that detail was explicitly distilled into MEMORY.md.
Strict Capacity Rejection (No Silent Auto-Compaction)
By default, MEMORY.md is capped at 2,200 characters (~800 tokens) and USER.md is capped at 1,375 characters (~500 tokens). Many agents silently drop older memories when limits are hit. Hermes explicitly refuses to do this: when an entry would exceed the limit, the memory tool returns an error. If the agent does not immediately consolidate or prune in that turn, the memory write does not land.
Write Approval Gating (write_approval: true)
If write approval is enabled in ~/.hermes/config.yaml (or toggled via /memory approval on), memory changes require human confirmation. In messaging channels (Telegram, Discord, Slack) or background tasks, writes are placed in a staging queue rather than written immediately. Until you run /memory approve <id>, the memory remains pending.
Profile & Home Directory Isolation
Hermes isolates agent memory per profile. If you run Hermes with hermes --profile work, its memories are stored in ~/.hermes/profiles/work/memories/. If you later run default hermes chat, the agent reads ~/.hermes/memories/ instead. Running two agent instances against the same home directory can also cause write conflicts.
Wrong Layer: Procedural Skill vs. Factual Reference
If you instruct Hermes on how to execute a complex multi-step deployment script, and expect it to recall that exact execution flow weeks later from memory, it may struggle if it saved a one-line summary to MEMORY.md instead of writing an executable skill. Multi-step procedures belong in the Skills engine (~/.hermes/skills/).
How to Inspect, Edit & Prune Hermes Memory
You do not need to guess what Hermes has saved. Hermes provides built-in CLI commands, file access, and desktop controls to inspect and prune memory state:
1. Inspecting the Timeline with hermes journey
Hermes includes a built-in learning journey timeline that plots saved memories and skills chronologically (aliases: hermes learning, hermes memory-graph; /journey in chat). Use it to audit what the agent has recorded. Node ids include a fingerprint, so always copy them from list output:
# View the interactive learning timeline in your terminal$hermes journey$# List node ids: skill names and memory:<source>:<index>:<fingerprint> ids$hermes journey list$# Edit a node in $EDITOR (paste the id exactly as printed by list)$hermes journey edit <node-id>$# Delete a node: memory chunks are removed, skills are archived (restorable)$hermes journey delete <node-id>
2. Direct Markdown File Inspection & Manual Editing
Because Hermes uses plain Markdown files, you can open and edit them directly in VS Code, Vim, or Nano anytime. No database migration or binary tool is required:
# Check your agent's personal environment notes$cat ~/.hermes/memories/MEMORY.md$# Check your saved user profile and preferences$cat ~/.hermes/memories/USER.md$# Open MEMORY.md directly in your editor$nano ~/.hermes/memories/MEMORY.md
Note: When manually editing, maintain the § delimiter between separate entries if you want them treated as discrete memory nodes.
3. Reviewing Pending Memory Writes
If memory write approval is turned on, review staged memory changes before they hit disk:
# In CLI or chat interface:$/memory pending # View pending memory write proposals$/memory approve <id> # Approve and commit a proposed memory$/memory reject <id> # Reject and discard a proposed memory$/memory approval on|off # Toggle runtime write approval requirement
4. Recalling Historical Sessions via SQLite FTS5
If Hermes seems to have forgotten a specific code snippet or conversation from weeks ago, browse past sessions using the CLI or tell Hermes to run a session search:
# List recent saved sessions with IDs and timestamps$hermes sessions list --limit 10$# Resume a specific past session directly$hermes chat --resume <session-id>
What Happens When Memory Is Full (Capacity Overflow)
When Hermes attempts to add an entry that would push MEMORY.md over 2,200 characters or USER.md over 1,375 characters, the memory tool immediately rejects the write with an explicit error:
{"success": false,"error": "Memory at 2,100/2,200 chars. Adding this entry (250 chars) would exceed the limit. Consolidate now: use 'replace' to merge overlapping entries into shorter ones or 'remove' stale or less important entries (see current_entries below), then retry this add — all in this turn.","current_entries": ["User runs Ubuntu 22.04 with Docker and Podman installed","Project backend uses Go 1.22 with chi router and sqlc","Staging server requires SSH key at ~/.ssh/staging_ed25519"],"usage": "2,100/2,200"}
How to Fix a Full Memory Store:
Run a one-shot prompt instructing Hermes to consolidate:
hermes chat --oneshot -q "Consolidate MEMORY entries and prune obsolete notes to stay under limit"Run hermes journey list and delete stale items using hermes journey delete <id>.
If you have a high-context LLM, raise limits in ~/.hermes/config.yaml under memory.memory_char_limit.
How Hermes Renders Memory Under the Hood
Understanding the internal format explains why Hermes is fast, predictable, and resilient against hallucination:
System Prompt Render & the § Delimiter
At session initialization, entries are injected into the system prompt with capacity percentage indicators and section sign (§) delimiters:
══════════════════════════════════════════════MEMORY (your personal notes) [67% — 1,474/2,200 chars]══════════════════════════════════════════════User's project is a Rust web service at ~/code/myapi using Axum + SQLx§This machine runs Ubuntu 22.04, has Docker and Podman installed§User prefers concise responses, dislikes verbose explanations
The explicit percentage header allows the agent to self-monitor its remaining memory headroom and proactively consolidate before hitting rejection thresholds.
SQLite FTS5 Full-Text Search
All interactions across CLI, messaging gateways, and desktop are indexed in an embedded SQLite database (~/.hermes/state.db) using SQLite's native FTS5 engine:
- Zero Token Cost at Idle: Past transcripts consume no tokens in standard prompt turns until queried.
- Exact Match Recall: FTS5 query execution runs in ~20ms, allowing Hermes to recall verbatim errors, stack traces, and discussions from weeks prior.
- Window Scrolling: The agent can paginate forward and backward within past sessions to reconstruct exact debugging context.
Memory Security Scanning
Because memory content is injected directly into privileged system prompt blocks, Hermes scans every write before it lands on disk. Entries containing prompt injection attempts, SSH backdoors, credential exfiltration patterns, or invisible Unicode zero-width characters are automatically blocked.
Optional External Memory Providers
Need advanced dialectic user modeling, vector embedding search, or cross-container memory sync? Hermes supports optional external memory plugins (such as Honcho, Mem0, OpenViking, and Hindsight) that operate additively alongside the core files:
# Configure an external memory plugin$hermes memory setup$# Check status of active memory backend$hermes memory status$# Revert to built-in local files only$hermes memory off
Visualizing Memory in Hermes Desktop
Procedures belong in Hermes skills, not memory. Memory facts checked against the official docs on September 26, 2026. To keep memory when moving machines, see backup and restore.
Prefer a visual interface? Hermes Desktop includes an interactive Star Map / Memory Graph panel to inspect and prune memory nodes visually.