✦
Hermes Agent

models

Hermes Agent Provider Fallbacks: Configure and Verify Failover

·AI agent provider fallbackmodelsprovidersreliability

Configure Hermes Agent fallback providers, distinguish credential pools from failover, and verify quota recovery, auxiliary routing, and backup costs.

A fallback is an authorized backup model route, not unlimited inference. Set it up before an unattended workflow depends on it, and verify which account pays when it activates. This guide covers Hermes Agent provider failover: configuration, credential pools, recovery after a quota window, auxiliary tasks, and a bounded acceptance test.

Quick answer#

Run hermes fallback to configure the primary model's backup chain, then hermes fallback list to inspect it. Hermes stores the current chain in the top-level fallback_providers list in the active profile's config.yaml. Same-provider credential pools are tried before cross-provider fallback. Compression and other auxiliary tasks have their own routing rules; a pinned cron job does not inherit the normal cross-provider chain.

Authorize and fund each route, set provider-side spending limits, and test it independently. A successful fallback reply does not prove that the original provider recovered, that the bill stayed within budget, or that a gateway delivered the result. Use the provider cost and rate-limit runbook when the immediate question is which error or wallet failed.

Configure a fallback without changing your primary model#

Start in the profile that owns the failing chat or service. The following commands inspect or open interactive configuration; they do not deliberately exhaust a quota:

hermes fallback
hermes fallback list
hermes auth list
hermes doctor

The fallback manager reuses the model/provider picker and supports add, list, remove, and clear. Prefer the picker when you need to verify a currently available model rather than copy an old model ID. For a named profile, prefix the command with hermes -p <profile>, replacing the placeholder with its actual name.

Current YAML shape#

The official documentation uses this example:

fallback_providers:
  - provider: openrouter
    model: anthropic/claude-sonnet-4

Both provider and model are required; an entry missing either is ignored. The model above is a documentation example, not a claim that your account can access it today. Confirm availability and price in the picker and provider account before using it. Credentials belong in the provider's supported auth store or secret environment, not this published snippet. Follow the API-key setup guide if the backup route is not authenticated yet.

fallback_model is the legacy single-fallback setting. Current hermes fallback writes fallback_providers and migrates legacy configuration on write; when both exist, the list takes precedence. There is no environment-variable replacement for the primary fallback chain. Editing an API-key environment variable alone does not configure failover.

Source: official fallback configuration, checked October 1, 2026. These are documented behaviors, not a claim that we injected provider failures into a production account.

Credential rotation and provider fallback solve different failures#

A credential pool rotates authorized keys or OAuth accounts for the same provider. Cross-provider fallback changes the provider:model route. A pool can help with a single exhausted credential; it does not eliminate an account-wide outage or merge subscriptions into one allowance.

The official credential-pool guide documents the order:

  1. Select a healthy credential from the current provider's pool.
  2. For a transient 429, retry that credential once before rotating on a repeated 429. A recognized exhausted plan window can rotate immediately instead.
  3. For a billing/quota failure, rotate to an eligible credential; for expired OAuth, try token refresh before rotation.
  4. If the pool cannot serve the request, try the configured provider fallback where that surface allows it.

fill_first keeps using the first healthy credential until it is exhausted. round_robin distributes selections, and least_used selects by observed request count. Those counters are not invoice totals: one selected credential can serve multiple inference requests. A pool strategy is therefore not a billing meter.

Reserve an account without pretending it is isolated#

Use hermes auth list <provider> to identify a credential. Current docs support hermes auth priority <provider> <target> <n> to reorder fill_first preference. Moving an entry to the back does not forbid its use: it remains eligible when earlier entries are unavailable, and a running session can retain its existing credential until rotation.

When one account must never serve another workload, use deliberately separated credentials and Hermes profiles, not priority alone. Profiles separate application state; they are not an operating-system permission sandbox. Adding the same OpenAI account twice also does not add quota: the pool docs warn that those logins share an upstream token family and can invalidate the older login.

Read the failure chain before choosing a repair#

Provider failures can appear as a silent Telegram bot, failed compression, a Desktop error, or a scheduled job that never produces its artifact. Match the evidence to the failed layer:

  • Primary 429, fallback 402: switching happened, but the backup billing route could not serve the request. Fix its entitlement or balance rather than reinstalling the gateway.
  • 401 or revoked OAuth grant: repair the exact credential. A dead login is not a temporary quota bench and will not become healthy just because you wait.
  • 403 or missing model: confirm account access and the exact model ID. A successful browser sign-in does not prove inference entitlement.
  • Repeated 500/502/503 or connection errors: inspect provider availability and bounded retry/fallback logs. A working tool API does not prove the model endpoint works.
  • Malformed or empty model response: preserve the request ID and error classification. Do not treat a model's content-policy refusal as an empty response to route around.
  • Model calls succeed but persistence or delivery fails: investigate session storage or transport. Another paid model call does not repair either.

Historical September support threads asked how to keep two ChatGPT subscriptions separate, reported rate limit 490, and questioned charges for a model not selected in the visible chat. These are support-demand signals, not proof of a billing defect or universal quota policy. Our available Discord archive ends September 14; it is not a same-week incident feed. Preserve the provider's complete response rather than relabeling a vendor-specific code as HTTP 429.

Recovery eligibility is not a promised recovery time#

The detailed fallback guide describes turn-scoped failover with reset-aware primary retry. When a provider supplies a future quota-reset time, Hermes can stay on the fallback until that time rather than retrying a known-exhausted primary on every message. Without a declared reset, transient rate limits use a cooldown.

The important operational distinction is eligible to retry, not guaranteed recovered. An elapsed cooldown neither schedules a new request nor proves a healthy response. Confirm the next actual turn's route and outcome before calling the incident closed.

An optional setting can avoid switching when the primary's declared reset is close:

fallback:
  min_switch_reset_seconds: 120

This documented example is a policy choice, not a universal recommendation. The default is 0, which leaves that switch-suppression threshold off. A longer wait may suit a low-urgency conversation but miss an unattended deadline. Match it to the task's tolerance, and never promise a completion time from a reset timestamp alone.

Why failover can increase token costs#

Switching provider, model, or account can lose the cached prompt prefix. The backup may need to read the conversation at uncached input rates. Returning to the primary may also require a full reread if its cache has expired. Long conversations that bounce between routes can therefore cost more even when the backup's advertised unit price is lower.

Before enabling failover for unattended work:

  • Record the billing owner and permitted model for every fallback entry.
  • Set a provider-side budget or limit for each paid route.
  • Compare cached and uncached input usage, not only output tokens.
  • Reconcile request timestamps against the provider ledger after a bounded test.
  • Leave the chain empty rather than adding a paid route you have not authorized.

The token-overhead guide explains how to distinguish context growth from cache loss. Model pricing, subscription entitlement, infrastructure, and operator time remain separate cost buckets.

Auxiliary tasks do not have one universal fallback rule#

Compression and vision can fail while the primary chat model is healthy. Inspect auxiliary.<task> rather than assuming the main model picker governs every request.

Auxiliary provider set to auto#

With an explicitly selected main provider, the detailed docs describe this order: main provider/model, then auxiliary.<task>.fallback_chain, then the top-level fallback policy. If no eligible route remains, the task warns or degrades according to its own behavior rather than choosing any paid account you happen to be logged into. Built-in provider discovery applies when no main provider is selected.

Explicit auxiliary provider#

An explicitly chosen auxiliary provider has a different capacity-error path: the configured route, a task-specific chain if present, then the main agent as a safety net. A transient 429 with a retry constraint is not treated like daily/monthly quota exhaustion. Auth-error handling is narrower; do not generalize one error's behavior to every auxiliary failure.

For compression, the docs describe a no-summary degradation when no provider is available. That keeps the session from failing outright but is not equivalent to preserving every detail. Use the memory and context troubleshooting guide before deleting memory or starting over.

Check the execution surface before expecting inheritance#

A working interactive fallback is not proof that all unattended work has the same policy.

  • CLI, Desktop, and gateway chat: verify the active profile and chain, then test the actual surface that failed.
  • Unpinned cron jobs: can inherit the configured chain. A job pinned to its own provider, model, or endpoint does not use normal cross-provider fallback; same-provider pool rotation can still apply.
  • Delegated work: delegation.fallback_providers provides an explicit policy. Without it, only unpinned children inherit the parent chain; an empty list disables the delegation chain.
  • Auxiliary tasks: follow their task-specific resolution and error rules, not simply the parent's final model name.

The cron recovery guide separates scheduling, quota holds, execution, and delivery. Do not remove a deliberate pin just to make an error disappear: that can move work onto a different billable or data-processing route.

A bounded fallback acceptance test#

Do not deliberately burn through a production quota or revoke a working production key to test failover. Use a separate test profile and a reversible failure fixture if you need fault injection. First verify the backup independently:

  1. Inspect the test profile's fallback list and credential ownership. Confirm the backup model exists and has a spending limit.
  2. Run a tiny explicit provider/model call using values from the picker. Ask for a fixed response such as provider ok, not for the model to identify itself.
  3. Read the actual provider/model route from runtime evidence and the provider ledger. A model's self-description is not routing proof.
  4. In the test environment, capture the primary error, pool decision, selected fallback, request timestamp, and final outcome.
  5. Test the smallest harmless tool task the workload needs. A text reply alone does not prove tool compatibility.
  6. Inspect the next turn after the reset or repair. Record whether the primary was tried and whether it succeeded.
  7. For a gateway, verify the reply in the exact destination. For a job that can write externally, read back the target before retrying an ambiguous outcome.

Do not blindly rerun a payment, post, deployment, or write after an uncertain network failure. The side effect may already have happened even if the final answer was lost. Restore unattended traffic only when routing, cost, and the real artifact or delivery have been verified.

Self-hosted control or managed operations?#

Self-hosting lets you choose providers, local endpoints, credentials, and fallback policy. It also makes you responsible for account renewal, service environments, logs, backups, gateway uptime, and incident response. A local fallback must still pass the intended tool task; locality alone does not make it a capable substitute.

If browser/mobile access and less runtime maintenance matter more than operating that stack, compare FlyHermes pricing. Managed hosting is a different responsibility model, not a promise of unlimited third-party inference or a specific fallback configuration on every plan. Check the current plan's model access and limits before deciding.

Sources and scope#

Product behavior was checked against the official fallback guide, credential-pool guide, and provider setup reference on October 1, 2026. The general provider reference still contains a legacy one-shot-per-session summary; this article follows the dedicated fallback guide's more specific turn/reset explanation. Verify behavior against your installed version and logs.

Historical demand examples: separate subscriptions, vendor-specific rate-limit report, and unexpected model-charge report. These community links may require Discord membership. They guide the questions answered here; official documentation supplies the product claims.

Frequently Asked Questions

How do I configure Hermes Agent provider fallbacks?

Run hermes fallback and inspect the result with hermes fallback list in the owning profile. Current configuration uses a top-level fallback_providers list with a provider and model for each entry. Authenticate, fund, and test each backup independently.

Are credential pools the same as provider fallbacks?

No. Pools rotate authorized credentials for the same provider and are tried first. Provider fallback changes the provider:model route when allowed by the execution surface. Neither merges subscriptions nor guarantees unlimited capacity.

Does a quota reset timestamp guarantee recovery?

No. It makes the primary eligible for a later retry. It neither schedules that retry nor proves the next request will succeed. Verify the route and result on an actual turn.

Do pinned cron jobs inherit the provider fallback chain?

No. A job pinned to its own provider, model, or endpoint does not use the normal cross-provider chain. Same-provider credential-pool rotation can still apply. Unpinned jobs can inherit the configured chain.

Why can fallback make a long conversation more expensive?

The backup provider, model, or account may not have a cached prefix for the conversation. Its next request can reread history at uncached input rates. Verify the billing account, cache usage, and provider-side spending limit.

Can a dead OAuth login recover by waiting for a cooldown?

No. A permanently rejected OAuth grant requires reauthentication in the owning profile. A temporary quota bench and a dead credential require different repairs.

FlyHermes (Managed Cloud)

Deploy in 60 seconds. API costs included. Cancel anytime.

Deploy faster with FlyHermes →

Self-Host (Open Source)

Full control. MIT licensed. Run on your own infrastructure.

View install guide →

Keep reading

Related Hermes Agent guides