Hermes Agent on Google Cloud Run
Yes, you can run Hermes Agent on Google Cloud Run, but not as a standard request-driven web service. Google publishes an official codelab deploying Hermes Agent as a Cloud Run Instance—a long-lived, stateful container runtime backed by a mounted Google Cloud Storage (GCS) bucket and native Vertex AI model access.
Deploying Hermes to Cloud Run involves five coordinated Google Cloud services: Cloud Run Instances, Cloud Storage FUSE, Secret Manager, IAM service accounts, and Vertex AI. Understanding how storage persistence and SQLite locking work across this stack is the difference between a reliable 24/7 agent and silent database corruption.
Checked against official Nous Research sources on September 26, 2026 (Hermes Agent v0.21.5).
Quick decision: Is Cloud Run the right choice?
| Hosting option | Best for | State persistence | Typical cost |
|---|---|---|---|
| Google Cloud Run Instances | GCP-native teams, Vertex AI / Gemini integration with IAM auth, containerized infrastructure without OS maintenance | GCS bucket mounted via GCS FUSE + supervisor sync loop | ~$21.41/mo compute (2 vCPU / 4 GiB in Tier 1) + storage & model tokens; offset by $300 credit for eligible new accounts |
| Self-managed VPS | Uncomplicated single-server setup, direct root access, native NVMe disk performance, predictable flat pricing | Local SSD / NVMe block storage with native SQLite WAL mode | $4–$12/mo on DigitalOcean, Hetzner, or Linode |
| Hermes Cloud | Zero-infrastructure managed hosting, 1-click deploy from Nous Portal, built-in remote Desktop connection | Managed persistent volume handled by Nous Portal | From $20/mo (drawn from Nous Portal credits) |
| Local computer | Day-to-day interactive CLI or Desktop use while your machine is awake; no server management | Local filesystem (~/.hermes/) | $0 hosting; only model API tokens consumed |
When Cloud Run Instances makes sense
- Your infrastructure is already centered on Google Cloud and managed through IAM policies.
- You want to run Google Gemini models via Vertex AI using Application Default Credentials (ADC)—eliminating plaintext API keys in your environment.
- You have eligible Google Cloud Free Trial credits ($300 over 90 days) that can be applied to offset initial compute and API spend.
- You need containerized orchestration without manually patching host operating systems, Docker daemons, or systemd services.
When to choose another path
- You require mature, generally available (GA) infrastructure with standard SLAs rather than a Preview feature.
- You want standard POSIX disk behavior without custom supervisor scripts to handle SQLite locking over object storage.
- You want the lowest possible flat monthly bill ($4–$6/mo Linux server via a VPS).
- You prefer zero cloud administration and want 1-click deployment via Hermes Cloud.
- You do not need 24/7 background task execution or messaging gateway bots when your laptop is closed.
Release status: Cloud Run Instances is currently in Preview
Cloud Run Instances is a Preview (Pre-GA) offering subject to Google's Pre-GA Offerings Terms. Pre-GA products can change prior to general availability and carry limited support. If your operational priorities center on mature, fully stabilized infrastructure with standard enterprise SLAs, a conventional Linux VPS may be preferable until Cloud Run Instances advances to a later release stage.
Why Cloud Run Instances (and not Cloud Run Services)?
Google Cloud Run provides two distinct compute models: Cloud Run Services and Cloud Run Instances. Deploying Hermes Agent successfully requires understanding why standard Services are the wrong fit for an autonomous agent:
- Cloud Run Services scale to zero: Standard services are designed for request-driven web applications. When incoming HTTP traffic stops, Cloud Run throttles container CPU to zero or shuts down the container entirely. Hermes Agent runs background execution loops, cron schedules, boot-time skill discovery, and persistent WebSocket connections to messaging gateways (such as Telegram or Discord). A scale-to-zero container terminates these background daemons.
- Cloud Run Instances run continuously: The
gcloud beta run instancescommand group deploys dedicated, singleton container instances designed for stateful workloads, worker daemons, and autonomous agents. The container remains active continuously (with a single instance lifetime of up to 7 days before an automatic graceful restart), ensuring scheduled routines and background processes run without interruption. - Direct addressability: Cloud Run Instances assigns a dedicated HTTPS URL directly to the instance, allowing secure access to the Hermes Web Dashboard on port 8080.
The architecture: 5 Google Cloud pieces in plain English
Google's deployment architecture connects the official Hermes container to four supporting Google Cloud managed services:
| Component | Google Cloud service | Role in deployment | Persistence behavior |
|---|---|---|---|
| Compute runtime | Cloud Run Instances | Runs the nousresearch/hermes-agent container (2 vCPU, 4 GiB RAM) continuously and exposes the Web Dashboard on port 8080. | Ephemeral container filesystem. Any file written outside mounted volumes is destroyed on restart. |
| Persistent storage | Cloud Storage (GCS) + GCS FUSE | Mounted at /opt/data. Holds workspace files, skills, and configuration backups across container restarts. | Fully persistent object storage bucket in your chosen region. |
| Secret injection | Secret Manager | Stores HERMES_DASHBOARD_BASIC_AUTH_PASSWORD. Injected securely into container environment variables at boot. | Encrypted, versioned secrets managed in Google Cloud Console. |
| Identity & security | Cloud IAM (hermes-sa) | Dedicated service account with least-privilege roles for storage, secrets, and Vertex AI. | Standard GCP IAM service account. |
| LLM backend | Vertex AI (Gemini 3.8 Flash) | Processes agent prompts and tool-calling loops using Application Default Credentials (ADC). | Stateless API calls billed per token to your GCP project. |
The SQLite & GCS FUSE challenge: The supervisor pattern
The most critical technical hurdle when hosting Hermes Agent on Google Cloud Run is the interaction between SQLite and Google Cloud Storage FUSE.
The problem: GCS FUSE mounts an object storage bucket as a local Linux directory (/opt/data). Object storage is not a POSIX filesystem and does not support byte-range locking. Hermes Agent stores chat history, agent memory, and session state in a local SQLite database (state.db). By default, SQLite uses Write-Ahead Logging (WAL mode), which requires shared memory files (-shm) and POSIX byte-range locks. Running SQLite in WAL mode directly against a GCS FUSE mount causes database locking errors, frozen queries, and severe database corruption.
Google's supervisor solution: Google solves this in the official codelab by wrapping Hermes in a custom Python supervisor script (run_hermes.py). Here is exactly how it works:
- Active state in local RAM: Hermes is configured to use
/tmp/hermes_home/.hermes(container tmpfs in RAM) for its active runtime, caches, and database operations. - State restoration on boot: When the container starts, the supervisor checks if
/opt/data/.hermes/state.dbexists in the GCS bucket. If found, it copies the database and configuration files to local/tmp. - TRUNCATE journal mode: The supervisor connects to SQLite and executes
PRAGMA journal_mode=TRUNCATE;. This forces SQLite into single-file mode without separate WAL files, ensuring safe atomic file copies. - Background auto-sync thread: A background daemon thread checks the modification time of
state.db,config.yaml, and.envevery 5 seconds. Whenever Hermes writes a new memory or chat message, the updated file is copied back to/opt/data/.hermes/on GCS. - Direct workspace persistence: Standard user files in
/opt/data/workspacedo not require SQLite locks, so files generated or edited by the agent write directly to the persistent GCS bucket.
Prerequisites & Google Cloud billing
Before deploying, you need an active Google Cloud account and the Google Cloud CLI:
- A Google Cloud project with billing enabled.
- Google Cloud Free Trial: New Google Cloud customers receive $300 in free credits valid for 90 days, plus access to the Google Cloud Free Tier. Google requires a payment method during signup for identity verification, but accounts are not automatically billed when the credit concludes unless you manually upgrade to a paid account.
- gcloud CLI with beta components: Because
gcloud beta run instancesis currently in beta, you must install the beta component or use Google Cloud Shell (which includes the Google Cloud SDK pre-installed).
Need a Google Cloud project?
Create a new Google Cloud account to claim the standard $300 90-day credit, which can be applied toward eligible infrastructure and model usage.
Step-by-step deployment guide
The following walkthrough provides the shortest verified path to deploy Hermes Agent on Cloud Run Instances. Run these commands in Google Cloud Shell or a terminal authenticated with gcloud.
Step 1: Set environment variables and enable APIs
Define your project configuration and enable the five required Google Cloud APIs:
# Export your environment variables$export PROJECT_ID="your-project-id"$export REGION="us-central1"$export BUCKET_NAME="hermes-state-${PROJECT_ID}"$export SERVICE_ACCOUNT_NAME="hermes-sa"$# Set default project and update gcloud beta$gcloud config set project $PROJECT_ID$gcloud components install beta --quiet$gcloud components update --quiet$# Enable required Google Cloud APIs$gcloud services enable \$ run.googleapis.com \$ secretmanager.googleapis.com \$ storage.googleapis.com \$ compute.googleapis.com \$ aiplatform.googleapis.com
Step 2: Create a dedicated IAM service account
Adhere to least-privilege security by creating a dedicated service account and granting it permission to invoke Vertex AI models:
# Create the service account$gcloud iam service-accounts create ${SERVICE_ACCOUNT_NAME} \$ --display-name="Hermes Agent Service Account"$$export SERVICE_ACCOUNT="${SERVICE_ACCOUNT_NAME}@${PROJECT_ID}.iam.gserviceaccount.com"$# Grant permission to invoke Vertex AI models$gcloud projects add-iam-policy-binding ${PROJECT_ID} \$ --member="serviceAccount:${SERVICE_ACCOUNT}" \$ --role="roles/aiplatform.user"
Step 3: Store dashboard password in Secret Manager
Generate a random 16-byte hex password for the Hermes Web Dashboard, store it in Secret Manager, and grant the service account access:
# Generate a secure password and store in Secret Manager$export DASHBOARD_PASSWORD=$(openssl rand -hex 16)$echo "Generated Hermes Dashboard Password: ${DASHBOARD_PASSWORD}"$$echo -n "${DASHBOARD_PASSWORD}" | gcloud secrets create hermes-dashboard-password \$ --data-file=- \$ --replication-policy="automatic"$# Grant the service account read access to this secret$gcloud secrets add-iam-policy-binding hermes-dashboard-password \$ --member="serviceAccount:${SERVICE_ACCOUNT}" \$ --role="roles/secretmanager.secretAccessor"
Step 4: Prepare Cloud Storage bucket & configuration files
Create the persistence bucket and grant the service account roles/storage.objectAdmin:
# Create the Cloud Storage bucket$gcloud storage buckets create gs://${BUCKET_NAME} --location=${REGION}$# Grant storage access to the service account$gcloud storage buckets add-iam-policy-binding gs://${BUCKET_NAME} \$ --member="serviceAccount:${SERVICE_ACCOUNT}" \$ --role="roles/storage.objectAdmin"
Create the agent configuration file config.yaml. Setting provider: "vertex" allows Hermes to authenticate using Google Cloud Application Default Credentials without a separate API key:
$cat << 'EOF' > config.yaml$_config_version: 12$$model:$ default: "google/gemini-3.8-flash"$ provider: "vertex"$$dashboard:$ enabled: true$$database:$ journal_mode: delete$EOF
Create the supervisor script run_hermes.py to handle cache redirection, SQLite TRUNCATE mode, and the 5-second GCS auto-save thread:
$cat << 'EOF' > run_hermes.py$import os, shutil, subprocess, sys, threading, time, sqlite3$$print("=== INITIALIZING HERMES SUPERVISOR ===", flush=True)$$home_dir = "/tmp/hermes_home"$hermes_dir = os.path.join(home_dir, ".hermes")$os.makedirs(hermes_dir, exist_ok=True)$os.makedirs("/tmp/logs", exist_ok=True)$os.makedirs("/tmp/skills", exist_ok=True)$os.makedirs("/tmp/uv_cache", exist_ok=True)$os.makedirs("/tmp/cache", exist_ok=True)$os.makedirs("/opt/data/workspace", exist_ok=True)$os.makedirs("/opt/data/.hermes", exist_ok=True)$# Restore state from GCS mount$for f in ["config.yaml", ".env", "state.db"]:$ src = os.path.join("/opt/data/.hermes", f)$ alt = os.path.join("/opt/data", f)$ dst = os.path.join(hermes_dir, f)$ if os.path.exists(src):$ shutil.copy(src, dst)$ print(f"Synced {f} from .hermes -> {dst}", flush=True)$ elif os.path.exists(alt):$ shutil.copy(alt, dst)$ print(f"Synced {f} from root -> {dst}", flush=True)$# Configure SQLite TRUNCATE mode to bypass GCS FUSE locking issues$db_path = os.path.join(hermes_dir, "state.db")$try:$ conn = sqlite3.connect(db_path)$ conn.execute("PRAGMA journal_mode=TRUNCATE;")$ conn.close()$ print("Configured SQLite database to TRUNCATE mode", flush=True)$except Exception as e:$ print(f"Notice during SQLite init: {e}", flush=True)$$subprocess.run(["chmod", "-R", "777", "/tmp"], check=False)$$env = dict(os.environ)$env["HOME"] = home_dir$env["HERMES_HOME"] = hermes_dir$env["PATH"] = "/opt/hermes/.venv/bin:/opt/hermes/bin:" + env.get("PATH", "")$env["PYTHONUNBUFFERED"] = "1"$env["HERMES_STATE_PATH"] = hermes_dir$env["HERMES_SKILLS_PATH"] = "/tmp/skills"$env["UV_CACHE_DIR"] = "/tmp/uv_cache"$env["XDG_CACHE_HOME"] = "/tmp/cache"$env["SQLITE_BUSY_TIMEOUT"] = "30000"$env["HERMES_ALLOW_ROOT_GATEWAY"] = "1"$env["HERMES_WORKSPACE"] = "/opt/data/workspace"$env["HERMES_WRITE_SAFE_ROOT"] = "/opt/data"$$python_bin = "/opt/hermes/.venv/bin/python3"$# Background thread: sync state.db, config.yaml, .env every 5 seconds$def sync_to_gcs_loop():$ files = ["state.db", "config.yaml", ".env"]$ last_mtimes = {f: os.path.getmtime(os.path.join(hermes_dir, f)) if os.path.exists(os.path.join(hermes_dir, f)) else 0 for f in files}$ while True:$ time.sleep(5)$ for f in files:$ src = os.path.join(hermes_dir, f)$ if os.path.exists(src):$ try:$ mtime = os.path.getmtime(src)$ if mtime > last_mtimes.get(f, 0):$ shutil.copy2(src, os.path.join("/opt/data/.hermes", f))$ last_mtimes[f] = mtime$ print(f"Auto-saved {f} to GCS volume mount", flush=True)$ except Exception as e:$ print(f"Error auto-saving {f} to GCS: {e}", flush=True)$$threading.Thread(target=sync_to_gcs_loop, daemon=True).start()$# Launch Hermes Gateway$print("=== STARTING GATEWAY IN BACKGROUND ===", flush=True)$gw = subprocess.Popen([python_bin, "-m", "hermes_cli.main", "gateway", "run"], env=env, cwd="/opt/data/workspace")$# Launch Web Dashboard on 0.0.0.0:8080$print("=== STARTING DASHBOARD ON 0.0.0.0:8080 ===", flush=True)$dash = subprocess.Popen([python_bin, "-m", "hermes_cli.main", "dashboard", "--host", "0.0.0.0", "--port", "8080", "--skip-build"], env=env, cwd="/opt/data/workspace")$$dash.wait()$EOF
Create the container startup wrapper start_hermes.sh and upload all three files to the root of your GCS bucket:
$cat << 'EOF' > start_hermes.sh#!/bin/sh$set -e$export PYTHONUNBUFFERED=1$exec python3 /opt/data/run_hermes.py$EOF$# Upload config and scripts to the bucket$gcloud storage cp config.yaml run_hermes.py start_hermes.sh gs://${BUCKET_NAME}/
Step 5: Deploy the Cloud Run instance
Deploy using gcloud beta run instances deploy, mounting the GCS bucket at /opt/data and injecting the dashboard password from Secret Manager:
$gcloud beta run instances deploy hermes-instance \$ --image nousresearch/hermes-agent:latest \$ --service-account ${SERVICE_ACCOUNT} \$ --command "/bin/sh" \$ --args "/opt/data/start_hermes.sh" \$ --port 8080 \$ --cpu 2 \$ --memory 4Gi \$ --ingress all \$ --no-invoker-iam-check \$ --add-volume name=hermes-storage,mount-path=/opt/data,type=cloud-storage,mount-options="uid=2000;gid=2000;file-mode=0777;dir-mode=0777;implicit-dirs",bucket=${BUCKET_NAME} \$ --set-secrets "HERMES_DASHBOARD_BASIC_AUTH_PASSWORD=hermes-dashboard-password:latest" \$ --set-env-vars "PYTHONUNBUFFERED=1,VERTEX_PROJECT_ID=${PROJECT_ID},VERTEX_LOCATION=global,HERMES_DASHBOARD_BASIC_AUTH_USERNAME=admin,HERMES_ALLOW_ROOT_GATEWAY=1,HERMES_WORKSPACE=/opt/data/workspace,HERMES_WRITE_SAFE_ROOT=/opt/data" \$ --region ${REGION} \$ --project ${PROJECT_ID}
Step 6: Access the Web Dashboard and verify persistence
Retrieve the generated URL and authenticate using username admin and the password stored in Secret Manager:
# Print your instance URL and password$INSTANCE_URL=$(gcloud beta run instances describe hermes-instance --region $REGION --format="value(status.urls[0])")$echo "Hermes Dashboard URL: $INSTANCE_URL"$echo "Username: admin"$echo "Password: $(gcloud secrets versions access latest --secret=hermes-dashboard-password)"
How to test persistence across restarts:
- Open your instance URL in your browser, log in, and send:
Write "hello world" to a file named hello.txt in your workspace. - In your terminal, verify the file arrived in your GCS bucket:
gcloud storage cat gs://$BUCKET_NAME/workspace/hello.txt - Redeploy or restart the instance by re-running the deploy command. Your previous chat sessions and workspace files will persist seamlessly.
Monthly cost breakdown
Cloud Run Instances billing is determined by provisioned vCPU and memory billed per second. Google publishes a baseline reference of approximately $5.70 for 30 days of continuous 24/7 execution for a singleton Cloud Run Instance configured with 1 vCPU and 1 GiB memory in Tier 1 regions (such as us-central1), utilizing shared vCPUs with burst budgets ($0.00000027 per vCPU-second and $0.00000193 per GiB-second).
The official Hermes codelab provisions a larger configuration—2 vCPU and 4 GiB memory—to accommodate container dependencies, workspace operations, and local task queues. Over a continuous 30-day month (2,592,000 seconds), this larger configuration calculates to:
- vCPU compute: 2 vCPUs × 2,592,000 s × $0.00000027/s = $1.40
- Memory compute: 4 GiB × 2,592,000 s × $0.00000193/s = $20.01
- Estimated compute subtotal: ~$21.41 / month (in Tier 1 regions like
us-central1; rates vary by region).
| Line item | Allocation / Usage | Unit price | Estimated monthly spend |
|---|---|---|---|
| Cloud Run Instance Compute | 2 vCPU, 4 GiB memory (continuous 24/7 uptime = 2,592,000 s in Tier 1 region) | $0.00000027 / vCPU-s + $0.00000193 / GiB-s | ~$21.41 / month (reproducible estimate: $1.40 vCPU + $20.01 memory) |
| Cloud Storage (GCS) | Standard storage for workspace and backups (~1–5 GB in Tier 1) | $0.020 / GB-month | ~$0.05 – $0.15 / month |
| Secret Manager | 1 active secret version (dashboard password) | $0.06 / active secret version (first 6 free in GCP Free Tier) | ~$0.00 – $0.06 / month |
| Network Egress | Dashboard web traffic and small data transfers | First 100 GB/mo from North America free under GCP Free Tier | $0.00 / month |
| Google Vertex AI (Gemini 3.8 Flash) | Autonomous agent prompt loops and tool calls (variable; ~5M–15M tokens/mo) | $0.15 / 1M input tokens; $0.60 / 1M output tokens | Variable (~$1.50 – $6.00 / month typical) |
| Total estimated spend | Continuous 24/7 autonomous agent (2 vCPU, 4 GiB) + typical model usage | Billed to Google Cloud project | ~$21.50/mo baseline infra + variable model tokens |
Applying the $300 Google Cloud credit: Eligible new Google Cloud accounts receive a $300 credit valid for 90 days. This credit can be applied toward eligible Cloud Run compute, Cloud Storage, Secret Manager, and Vertex AI usage, and can substantially offset or cover the baseline infrastructure of a modest deployment. However, it is not a blanket guarantee: total credit consumption depends directly on your resource configuration, execution uptime, and variable Vertex AI token volume. For broader cost modeling across external models (Anthropic, OpenAI, OpenRouter), see our Hermes cost guide.
Common gotchas & troubleshooting
"Invalid choice: 'instances'" error in gcloud
The instances command group is part of the beta track. You must prefix the command with beta (i.e. gcloud beta run instances deploy) and have the beta component installed:
$gcloud components install beta
"database is locked" or corrupted state.db
This happens when SQLite attempts to write directly to a GCS FUSE volume mount using standard WAL mode. Verify that run_hermes.py is setting HOME=/tmp/hermes_home and executing PRAGMA journal_mode=TRUNCATE; before starting the gateway or dashboard.
HTTP 401 Unauthorized on Web Dashboard
When prompted by the browser for basic authentication, the username is always admin (set in HERMES_DASHBOARD_BASIC_AUTH_USERNAME). Fetch the active password from Secret Manager:
$gcloud secrets versions access latest --secret=hermes-dashboard-password
7-day instance lifecycle and restarts
Cloud Run Instances has a maximum execution limit of 7 continuous days per instance. When this duration elapses, Cloud Run automatically restarts the container. Because the supervisor syncs state.db to GCS every 5 seconds and restores it on boot, restarts take only a few seconds without data loss.
Public access vs. enterprise security
The codelab uses --ingress all and --no-invoker-iam-check to enable direct browser access protected by HTTP basic auth. For corporate or production environments, you can enforce IAM invocation (roles/run.invoker) or place Cloud Run behind Identity-Aware Proxy (IAP) and an Internal Application Load Balancer.
Primary sources
- Google Codelabs — Deploy Hermes Agent on Cloud Run Instances
- Google Cloud Run Instances Documentation
- Google Cloud Run Instance Lifecycle & Pre-GA Terms
- Google Cloud Run Pricing Documentation
- Google Cloud Storage FUSE Overview & Limitations
- Google Vertex AI Documentation
- Google Cloud Free Program & $300 Credit Terms
- Nous Research Hermes Agent Repository
- Official Hermes Docker Image — Docker Hub
HermesAgentAI.org is an independent educational documentation resource and community guide. It is not affiliated with, sponsored by, or endorsed by Nous Research or FlyHermes. Hermes Agent is released under the MIT License by Nous Research.
Where to go next
Compare with traditional Linux VPS hosting (systemd, local NVMe SSD).
Compare local, VPS, Hermes Cloud, and managed hosting options.
Configure Vertex AI, Gemini, Claude, or OpenAI models.
Manage tools, chat interfaces, and security settings.
Calculate token consumption and infrastructure expenses.