# Google Agent Runtime / Agent Engine

Official pricing: https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing  
Category: hyperscaler · Isolation: gvisor

## Pricing regimes (raw)

- **Agent Runtime allocation while processing turns** (resource): $0.085/vCPU-h, $0.009/GiB-h [alt]
- **Runtime 1-year Flexible Savings Plan** (resource): $0.0765/vCPU-h, $0.0081/GiB-h [alt, sales]
- **Runtime 3-year Flexible Savings Plan** (resource): $0.068/vCPU-h, $0.0072/GiB-h [alt, sales]
- **Code Execution Sandbox (separate from agent Runtime)** (sizes): default 2 vCPU/1.5 GiB $0.1835/h; VCPU4_RAM4GIB 4 vCPU/4 GiB $0.376/h
- **Agent Sandbox (custom container / shell / Computer Use)** (resource): $0.085/vCPU-h, $0.009/GiB-h [beta]

## Features

Yes: snapshots, memory snapshots, fork/clone, pause/resume, persistent disk, volumes, ≥24 h sessions, custom image (Docker or snapshot), start from your own snapshot, your own Docker/OCI image, browser, desktop GUI, computer-use API, code interpreter, browser + desktop control, preinstalled agents, egress allowlist, static egress IP, private networking, SOC 2, HIPAA, SSO, arm64, Python, Node.js, HTTP method/path egress rules, live fork (no pause), memory fork, EU data residency, automatic snapshots, snapshots on demand, agent harness API, hosted agent API (their own agent), gVisor or VM (no shared kernel), inbound access rules, extra volumes

No: idle auto-stop, full VM (own kernel), Docker inside, nested virtualization, root, anti-bot stealth, CAPTCHA solving, residential IPs, SSH, public IPv4, HTTPS preview URLs, self-hosting, BYOC, open source, GPU, Windows, macOS, wake on request, live resize, MCP server, firewall inside (nftables), secret proxy, secret proxy for any API, shared volumes

Unknown: everything else. Evidence (source + quote) per feature: https://battleships.dev/data/providers/google-agent-engine.json → feature_evidence

## Caveats

- Runtime CPU-memory compatibility matrix beyond documented 4/8 example
- Sandbox billable lifecycle: price page does not explicitly exempt idle sandboxes
- Region variance and project quota needed for 50 workers
- Sandbox egress, disk and snapshot rates
- Specific isolation implementation
- No account-wide resource grants encoded as dollars: free.monthly_credit=null, estimate is gross; apply each resource grant independently.
- Session/Memory Bank operations, persistent storage, LLMs and external tools excluded.
- Code execution presets cannot serve a single 4/8 workload. Do not sum smaller sandboxes to claim equivalent capacity.
- category agent-platform is task-requested extension to original category enum; isolation null intentionally unknown.
- Sandboxes (all types) are billed for their whole lifetime ('Sandboxes are billed while they exist') until deleted or TTL expiry; TTL max 14 d (Code Execution) / 7 d (custom containers); Computer Use default TTL 2 h; execute_code resets the TTL. This replaces the earlier caveat 'price page does not explicitly exempt idle sandboxes'.
- Paused sandboxes 'cost significantly less than running ones'. The paused/disk rate and the snapshot storage rate are unpublished (possibly Agent Storage $0.30/GiB-month; unverified).
- Each create() without a named template creates a new sandbox template (with pre-warmed pools) that is NOT deleted with the sandbox. Delete or reuse templates; it's unclear whether warm pools are billed.
- Sessions, Memory Bank and Skill Registry have been billed since 2026-09-01: storage $0.30/GiB-month (Agent Storage), reads 1 vCPU-h ($0.085) per 3M, writes 1 vCPU-h per 1M. Agent Gateway has been billed since 2026-07-13 at 1 vCPU-h per 15,000 calls. Semantic Governance Policy billing starts 'later in 2026'.

## How this provider charges


As of 2026-09-28. Category: `agent-platform`. USD list prices unless explicitly labelled otherwise.

The requested Vertex AI Agent Engine documentation now redirects into **Gemini Enterprise Agent Platform**. Keep stable id `google-agent-engine`; don't price from obsolete redirected model-pricing pages.

| Regime | When / billing | Numbers | Source |
|---|---|---|---|
| Runtime and Sandbox PAYG | Allocated vCPU-h + GiB-h, nearest second | $0.085/vCPU-h + $0.009/GiB-h | [Unified pricing](https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing) |
| Flexible Savings Plan | Eligible 1-year / 3-year commitment | 1y $0.0765 + $0.0081; 3y $0.068 + $0.0072 per CPU/RAM-hour | [Pricing](https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing) |
| Free resources | Account/month, not fungible credit | 50 vCPU-h, 100 GiB-h, 1 GiB-month Agent Storage | [Pricing](https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing) |
| Sessions and Memory Bank | Since September 1, 2026 | Storage $0.30/GiB-month; $0.085/3M reads; $0.085/1M writes; generation/embedding tokens separate | [Pricing](https://cloud.google.com/products/gemini-enterprise-agent-platform/pricing) |

Runtime default is 4 CPU/4 GiB. Configurable CPU choices: 1,2,4,6,8; memory 1..32 GiB subject to compatibility. `max_instances` defaults to 100, supports up to 1000, or 100 with VPC-SC/PSC-I. This is instance scaling, not isolated sandbox concurrency. Request concurrency is separately configured. [Deployment](https://docs.cloud.google.com/gemini-enterprise-agent-platform/scale/runtime/deploy-an-agent)

Code Execution is a separate sandbox: default 2 CPU/1.5 GB; documented alternative 4 CPU/4 GB. These cost $0.1835/h and $0.376/h at list allocation rates. State TTL up to 14 days; 100 MB request/response file payload; no outbound network. TTL is not continuous CPU time. [Quickstart](https://docs.cloud.google.com/gemini-enterprise-agent-platform/scale/sandbox/code-execution-quickstart), [overview](https://docs.cloud.google.com/gemini-enterprise-agent-platform/scale/sandbox/code-execution-overview)

## Gotchas
Runtime idle between turns is unbilled, but 30% OS CPU utilization does not justify reducing allocation by 70%. Sandbox idle exemption is not established. Runtime and sandbox may both be used and both billed; they are not mutually substitutable discount modes. Agent Storage for conversation/memory data isn't a confirmed disk/snapshot tariff. Quotas, network charges and model tokens are additional/unknown. Historical $0.0864 CPU and event-count memory pricing are superseded by the current official page.

## Worked examples
Runtime 4/8: $0.412/h; 8,800 billable instance-hours = **$3,625.60 gross**. If resource grants remain wholly unused elsewhere: subtract $4.25 CPU and $0.90 memory = **$3,620.45**, plus models, other services, storage and network. CPU 30% changes neither allocation rate. Single worker 176 processing hours: $72.512 gross, $67.362 after these grants. Instance sharing can change runtime-hour demand.
Code Execution's documented presets cannot satisfy one 4/8 worker; do not replace it with two smaller sessions. Snapshot 50 GiB and egress 100 GiB remain unpriced. Card supports conservative gross compute only and opts agent-specific modes in via `alt`; grants are not dollar credits in the engine.
