↓ Skip to main content

Design philosophy

Design philosophy
#

This page is the why. Later pages turn it into maps, hard invariants, and labs.

What Hermes optimizes for
#

Hermes ships one agent core and many surfaces (CLI, gateway, TUI, Desktop, cron, ACP). It should:

  • Feel the same agent wherever you talk to it
  • Learn across sessions (memory + skills) without silently rewriting the past mid-chat
  • Grow capability at the edges (skills, plugins, MCP, platform adapters) instead of stuffing every idea into the core tool schema

Two lenses for every change
#

Before you write code, ask:

1. Does this break per-conversation prompt caching?
#

Long conversations reuse a cached system-prompt prefix every turn. Anything that mutates past context, swaps toolsets, reloads memories into the system prompt, or rebuilds the system prompt mid-conversation invalidates that cache and multiplies cost.

Compression is the deliberate, sanctioned exception. Slash commands that change system-prompt state (skills, tools, memory) should default to deferred invalidation (next session), with an opt-in --now.

2. Does this widen the narrow waist?
#

Every core model tool is attached to (nearly) every API call. The bar for a new core tool is high. Prefer, in order: extend existing code → CLI + skill → service-gated tool → plugin → MCP → new core tool last.

That ordering is the Footprint Ladder — see Extend.

Expansive edges, conservative waist
#

Hermes is not a tiny product. Platforms, providers, desktop/TUI features expand on purpose.

Restraint targets the core agent + model tool schema — the one place every addition is paid for on every API call:

Expand freely (edges)Be conservative (waist)
Platform adapters, providers, desktop UINew entries in core tool schemas
Skills, optional skills, pluginsMid-conversation system-prompt mutation
Gateway features, dashboardsSpeculative hooks with no consumer

Architecture design principles
#

From the official architecture guide, with “good / bad” examples for engineers:

PrincipleGood changeBad change
Prompt stabilityInject guidance via a user message or tool resultRebuild system prompt every turn to “refresh memory”
Observable executionStream tool progress through existing callbacksSilent background side effects the user never sees
InterruptibleRespect cancel/interrupt during long toolsIgnore /stop while a tool runs forever
Platform-agnostic corePut Discord quirks in the Discord adapterSpecial-case Discord inside AIAgent
Loose couplingGate optional tools with check_fn / registriesHard-import a niche SaaS SDK in the core loop
Profile isolationKey state by profile home / secret scopeRead a secondary profile’s secrets from process os.environ under multiplex

Contribution taste (summary)
#

Wanted: fix real bugs end-to-end; expand reach at the edges; extract god-files into stem_topic.py siblings; keep the core narrow; extend instead of duplicating; assert behavior contracts in tests; prove E2E against a temp HERMES_HOME.

Not wanted (even if polished): speculative hooks; new HERMES_* env vars for non-secret config; a new core tool when terminal + file or a skill already suffice; lazy offset/limit on instructional tools (skills/prompts) so models only read page 1; “security fixes” that destroy the feature; outbound telemetry without opt-in; plugins that patch core files; third-party product plugins absorbed into the core tree.

Facade + siblings culture
#

Large modules are a facade plus topical siblings (gateway/run.py + run_*.py, agent/turn_*.py). Find code by topic, not by reading the facade first. Do not grow facades past the split signal (~2k lines / heavy functions).

Bridge to the next pages
#

NextWhat it adds
System mapBoxes and directories
Design invariantsEnforceable must-not-break rules
Footprint LadderHow to add capability
LabsPractice saying no to the wrong rung

Further reading
#

  • Root AGENTS.md — What Hermes Is, Contribution Rubric, Footprint Ladder
  • website/docs/developer-guide/architecture.md § Design Principles