✦
Hermes Agent

review

Gemini Enterprise Agent Platform Review

·Gemini Enterprise Agent Platform ReviewreviewAI agents

Gemini Enterprise Agent Platform review: runtime pricing, memory fees, testing-only Managed Agents API restrictions, refunds, ad evidence and alternatives.

Quick answer#

Gemini Enterprise Agent Platform is worth shortlisting for a team building governed agents on Google Cloud, not for someone buying a ready-to-use personal assistant. It combines the former Vertex AI platform with agent development, runtime, identity, gateway, memory and evaluation services. There is no single subscription price that includes every component.[1][28]

The most important buying distinction is inside the platform: Agent Runtime can host your agent application; the separate Managed Agents API currently carries testing-only restrictions. Its October 6 documentation explicitly says not to use proprietary, sensitive or confidential data and prohibits commercial or production use of that Pre-GA offering. Do not transfer the platform's broad enterprise messaging to this specific API.[3][35]

  • Best for: engineering teams already using Google Cloud that want managed deployment with their own agent logic, permissions and evaluation process.
  • Avoid if: you need a finished assistant, a predictable all-in seat price, or production access to the current Managed Agents API specifically.
  • Price reference: published default Scale rates are $0.085 per vCPU-hour and $0.009 per GiB-hour, with separate model, storage and service usage.[2]
  • Strongest alternative to evaluate: Microsoft Foundry Agent Service when your identity, data and operations already live in Azure; a narrower managed harness when that is all you need.

Verified October 7, 2026. This is a documentation, pricing, terms, advertising and community-source review. We did not deploy a billable Google agent or measure task success. Recommendations and the proposed pilot below are our analysis, not reported test results.

What the name covers: platform, runtime, app and managed API#

Google's product page calls this the successor to Vertex AI. The agent overview presents three build paths: low-code Agent Studio, a configuration-driven Managed Agents API, and custom code using the Agent Development Kit. Its broader architecture separates building, scaling, governing and optimizing agents.[1][28]

That breadth is useful if your procurement problem includes IAM, shared tool access, deployment, memory and auditing. It is unnecessary complexity if the problem is simply getting one assistant to work on files. Start with the application contract you need, not the number of services in the catalog.

The Gemini Enterprise app is another surface: the product page describes using it to register, manage and govern custom agents across an organization. An app seat should not be treated as a receipt for all Agent Platform runtime and model consumption.[1][2]

Agent Runtime supports containerized applications that conform to its runtime contract, including deployments from a container image or Dockerfile. ADK has full integration, while other frameworks have different levels of templates and SDK support. The older ReasoningEngine resource name remains for API backward compatibility.[3]

For a comparable enterprise hosting purchase, the Microsoft Foundry Agent Service review separates hosting, model charges and identity decisions. Neither platform should be chosen because its marketing page uses the same word “agent.”

Pricing: runtime, memory and a shared free allowance#

The current Scale rate card uses Agent Compute, Agent Memory and Agent Storage. At the default published USD rates, compute is $0.085 per vCPU-hour and RAM is $0.009 per GiB-hour. The monthly per-account allowances are 50 compute hours, 100 GiB-hours of memory and 1 GiB-month of storage; these are not fresh allowances for every agent you create.[2]

For Runtime, Google says idle time waiting for the next prompt between turns is not billed, and usage is rounded to the nearest second. That statement concerns Runtime metering, not a promise that every dependent service stops charging whenever the agent is quiet.[2]

Illustrative runtime-only calculation: assume the meter records 100 vCPU-hours and 200 GiB-hours in a month. At the default rates, that is $10.30 before free allowances, or $5.15 if the full compute and memory allowances remain available. This calculation excludes inference, storage, gateway operations, third-party tools, taxes and other resources. It is not a recommended configuration or a measured user bill.[2]

One- and three-year savings-plan columns are also published. Do not compare their discounted rates with on-demand alternatives while omitting the commitment. Confirm the exact region, SKU, account allowance and contract before forecasting your workload.[2]

Use the agent cost calculator framework to separate variable usage from the engineering and operational work your team retains. Cheap runtime is not necessarily cheap delivery.

Hidden operating costs: memory and governance have their own meters#

The less obvious cost is the surrounding platform, not just the agent process:

  • Agent Gateway: the published conversion is $0.085 per 15,000 API calls or authorization requests, prorated to usage. The page says this billing took effect July 13, 2026.[2]
  • Memory Bank: storage is $0.30 per GiB-month; operations are separately metered, and memory-generation and embedding model tokens are additional. The rate card dates this structure to September 1, 2026.[2]
  • Sessions: conversation-history storage and read/write operations have charges under the same resource framework. Session storage is not the same thing as working RAM.[2]
  • Skill Registry: storage, operations and model tokens for vulnerability analysis are separate components, with billing listed from September 1, 2026.[2]
  • Semantic governance: Google publishes an evaluation-compute rate and separate model-token charges, but the page still says billing will commence “later in 2026.” Treat the start date as unresolved, not as proof that a specific charge is already active.[2]

For Memory Bank and Sessions, the published operation conversions are one $0.085 compute-hour per three million reads or one million writes, with actual usage prorated. Memory generation still incurs its own model charges. A frequently updated agent can therefore incur costs even when the stored text is small.[2]

Standard PayGo model capacity is another consideration. Google describes usage tiers based on rolling organizational spend and says a 429 can reflect temporary contention for shared resources. A runtime budget does not buy guaranteed inference throughput; assess the applicable model capacity option separately.[8]

Benefits: where Google's integrated approach earns its place#

The strongest benefit is reducing the number of infrastructure pieces your team has to assemble. Runtime abstracts deployment infrastructure; the platform adds session history, cross-session Memory Bank, a central registry, agent identities and gateway enforcement.[3][28]

Agent Identity is documented as a SPIFFE-formatted identity supported in IAM, allowing specific permissions rather than relying only on shared service accounts. The registry and Auth Manager support user-delegated tool access. These are substantive reasons for an enterprise evaluation, but permissions still need to match the actual business workflow.[28]

The current release notes also show why component-level verification matters. On September 30, the App Topology API became generally available, while Cloud Trace integration with Agent Gateway was introduced in Preview. On September 25, VPC Service Controls support for semantic governance policies was still marked Preview.[4]

Our assessment: shortlist the platform when those identity and observability capabilities solve a real requirement. Do not buy them simply to avoid writing a small model/tool loop. Use the MCP security review checklist to test whether the tool host can actually reach or modify more than intended.

Limitations: a production platform can contain a testing-only API#

The Managed Agents API uses an Antigravity harness and an isolated sandbox, with the Agents API as its control plane and the Interactions API as its runtime interface. It is a different purchase from deploying your own application on Agent Runtime.[35]

Its current warning is unusually consequential: limited testing and evaluation only, no commercial or production purposes, and no proprietary, sensitive or confidential data. Model usage during preview is still billed at standard rates. A paid model call does not waive those restrictions.[35]

The sandbox starts without external networks or credentials. Developers must explicitly configure connectivity and scope credentials for external tools. That is a useful default boundary, not a guarantee that enabling a network or adding an MCP server is safe.[35]

There is also migration work to budget. Runtime documentation describes the SDK 2.0.1 split into the standalone google-cloud-agentplatform package, with modules such as runtimes, sessions, sandboxes and memory banks. Backward-compatible resource names do not mean every old SDK example remains the right implementation.[3]

Before choosing this route, read the OpenAI Agents API review for a narrower managed-harness comparison. Compare supported production use, control-plane retention and cleanup, not just sandbox availability.

Company, cancellation, refunds and budget controls#

This is a Google Cloud service governed by the applicable Cloud agreement and service-specific terms. Google's general terms permit stopping use, but termination remains subject to financial commitments in an order form or addendum. Fees already owed remain due; termination ordinarily does not require a refund unless the agreement or law says otherwise.[5][6]

The billing help page describes a narrower refund possibility: unused funds in a linked Postpay payments account may be eligible. It excludes unused promotional credit, accounts with an outstanding balance, and generally unused AI Studio Prepay credits subject to stated exceptions. That is not a satisfaction guarantee for model output or consumed runtime.[34]

Do not confuse “set a budget” with “everything stops at that dollar amount.” Current billing documentation distinguishes alerts-only budgets from spend-cap budgets where supported for the service. Check the precise service coverage and configure a response; an alerts-only budget is not an enforcement mechanism.[36]

Our exit checklist is to inventory deployed agents, stored sessions and memories, external tools, associated cloud resources and commitments before ending the pilot. Confirm which resources have actually been deleted and reconcile later billing entries. Turning off a client application alone is not an adequate financial closeout.

Complaints and independent evidence: useful, but not a benchmark#

A Hacker News discussion about enterprise orchestration describes a Google Cloud-based team's requirements for auditability, human oversight, custom harnesses and cloud execution. It is a useful account of the buyer's problem, not proof that a particular platform solved it.[25]

Another discussion directly illustrates name confusion: one commenter asks whether the new Agent Platform name means merely a Vertex AI rename; a reply prefers AI Studio. That preference is an opinion, not evidence that AI Studio and this whole managed platform have identical features.[24]

Info-Tech's current product listing includes a complaint about the learning curve and positive reports about productivity. It also displays older reviews and a nominal gift-card disclosure for at least one reviewer. We do not treat its aggregate product label as proof that every reviewer tested the current Agent Runtime or the new Managed Agents API.[18]

The evidence supports a cautious conclusion: evaluate implementation complexity and product boundaries. It does not establish an audited failure rate, a universal customer outcome or a reason to label the product a scam.

Advertising evidence: one exact-product match, not an ad-performance claim#

On October 7, the rendered Meta Ad Library showed Google Cloud, Library ID 1726342688432335, marked active and started October 1, 2026. Its copy said “Go from model to market in minutes,” and its destination was the Gemini Enterprise Agent Platform product page with paid-social campaign parameters. This is exact-product advertising evidence, not merely an advertiser-name match.[31]

The search also returned unrelated advertisers. We do not use the displayed query total as Google's campaign count. A separate Google Ads Transparency domain lookup returned no ads in that browser view; that does not invalidate the Meta creative or establish that Google Cloud runs no other advertising.[29][31]

The ad's $300-credit offer should be read alongside the product page's “new customers” and “up to” wording. Advertising verifies promotion, not performance, ROI or universal trial eligibility.[1][31]

Alternatives and decision: buy the layer you need#

Our strongest enterprise alternative is Foundry for an Azure-centered team. Existing identity, networking and operational expertise can matter more than a small difference in the runtime rate. For a Google Cloud-centered team, Agent Runtime deserves the same consideration rather than an automatic migration.

For an individual operator who wants an inspectable agent rather than an application platform, Hermes is a different category. Its official documentation describes a self-improving agent with memory, tools and multiple interfaces; it is not a drop-in replacement for Google's IAM, managed gateway or APIs.[19]

The self-hosted versus managed decision guide is the relevant next step if maintenance is the objection. Choose the Hermes installation path for control; evaluate managed Hermes access when operating the runtime is the work you want to avoid. FlyHermes presents managed access as its offer, not a replacement for Google's enterprise governance stack.[20]

Decision checklist before committing#

This is a proposed evaluation, not a test we ran:

  1. Name the exact component and launch stage. Reject the current Managed Agents API for a production requirement while its testing-only terms apply.
  2. Select one representative workflow, synthetic test data, an explicit tool allowlist and a human approval boundary.
  3. Estimate inference, runtime, gateway, memory and stored-session usage separately. Record which free allowances are already consumed.
  4. Exercise a tool denial, an inference 429 and a failed action. Check whether retrying repeats an external write.
  5. Prove the intended budget response and its coverage; record what continues after a threshold is reached.
  6. Export needed evidence, delete disposable resources, and reconcile the pilot's invoice rather than ending at a successful demo.

Decision: choose Agent Platform for a justified Google Cloud application and governance requirement. Choose something narrower for a finished assistant or a managed loop. Do not approve the specific Pre-GA Managed Agents API for a workload its own current terms exclude.

Sources#

[1] https://cloud.google.com/products/gemini-enterprise-agent-platform — Gemini Enterprise Agent Platform (formerly Vertex AI) | Google Cloud [2] https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing — Gemini Enterprise Agent Platform pricing | Google Cloud [3] https://docs.cloud.google.com/gemini-enterprise-agent-platform/build/runtime — Agent Runtime  |  Gemini Enterprise Agent Platform  |  Google Cloud Documentation [4] https://docs.cloud.google.com/gemini-enterprise-agent-platform/release-notes — Gemini Enterprise Agent Platform release notes  |  Google Cloud Documentation [5] https://cloud.google.com/terms — Google Cloud Platform Terms Of Service [6] https://cloud.google.com/terms/service-terms — Service Specific Terms  |  Google Cloud [8] https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/standard-paygo — Standard PayGo  |  Gemini Enterprise Agent Platform  |  Google Cloud Documentation [18] https://www.infotech.com/software-reviews/products/gemini-enterprise-agent-platform?c_id=523 — Gemini Enterprise Agent Platform Customer Reviews 2026 | Agentic AI [19] https://hermes-agent.nousresearch.com/docs — Hermes Agent Documentation | Hermes Agent [20] https://www.flyhermes.ai — FlyHermes.ai — Your Hermes Assistant [24] https://news.ycombinator.com/item?id=49717636 [25] https://news.ycombinator.com/item?id=47925632 [28] https://docs.cloud.google.com/gemini-enterprise-agent-platform/agents/overview [29] https://adstransparency.google.com/?region=anywhere&domain=cloud.google.com&hl=en — google-ads [31] https://www.facebook.com/ads/library/?active_status=active&ad_type=all&country=ALL&q=Gemini%20Enterprise%20Agent%20Platform&search_type=keyword_unordered — gemini-meta [34] https://docs.cloud.google.com/billing/docs/how-to/resolve-issues — google-refund [35] https://docs.cloud.google.com/gemini-enterprise-agent-platform/build/managed-agents [36] https://docs.cloud.google.com/billing/docs/how-to/budgets

Frequently Asked Questions

Is Gemini Enterprise Agent Platform the same as Vertex AI?

Google presents the platform as the successor to Vertex AI, with agent development, runtime, governance and evaluation services. The Gemini Enterprise app and the Managed Agents API are specific surfaces, not interchangeable names for every capability.

How much does Gemini Enterprise Agent Platform cost?

There is no single all-in platform price. The default published Scale rates include $0.085 per vCPU-hour and $0.009 per GiB-hour, with monthly account-level free allowances. Model tokens, stored data and other service operations are separate.

Can I use Google’s Managed Agents API in production?

Not under the documentation checked October 7, 2026. This specific Pre-GA offering is limited to testing and evaluation, prohibits commercial or production use, and warns against proprietary, sensitive or confidential data. Agent Runtime is a different component.

Does Agent Runtime charge for idle time between prompts?

The current pricing page says Runtime idle time waiting for the next prompt between turns is not billed. This does not make stored sessions, memory, inference or every dependent cloud resource free.

Does a Google Cloud budget automatically stop all agent spending?

Do not assume it does. Current documentation distinguishes alerts-only budgets from spend-cap budgets where supported. Verify service coverage and configure the intended response for the actual deployment.

Can I get a refund for unused Google Cloud credit?

Unused funds in a linked Postpay payments account may be eligible under the billing rules. Promotional credits are excluded. This is not a satisfaction guarantee for consumed agent usage, and contract commitments can still apply.

Was this a hands-on Gemini Agent Platform benchmark?

No. It is a source-based review of documentation, prices, terms, advertising and independent discussion. The runtime calculation is illustrative arithmetic, not a measured workload bill.

FlyHermes (Managed Cloud)

Deploy in 60 seconds. API costs included. Cancel anytime.

Deploy faster with FlyHermes →

Self-Host (Open Source)

Full control. MIT licensed. Run on your own infrastructure.

View install guide →

Keep reading

Related Hermes Agent guides