# Baseten

Official pricing: https://docs.baseten.co/deployment/resources  
Category: gpu-cloud · Isolation: container

## Pricing regimes (raw)

- **T4, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **L4, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **A10G, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **A100-80GB, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **H100, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **H200, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **B200, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **RTX-PRO-6000, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]

## Features

Yes: volumes, idle auto-stop, ≥24 h sessions, custom image (Docker or snapshot), your own Docker/OCI image, SSH, HTTPS preview URLs, egress allowlist, open internet, self-hosting, BYOC, SOC 2, HIPAA, SSO, GPU, Python, wake on request, EU data residency, MCP server, inbound access rules, extra volumes, shared volumes

No: snapshots, memory snapshots, fork/clone, pause/resume, persistent disk, start from your own snapshot, full VM (own kernel), Docker inside, nested virtualization, browser, desktop GUI, computer-use API, browser + desktop control, anti-bot stealth, CAPTCHA solving, residential IPs, preinstalled agents, public IPv4, static egress IP, open source, arm64, Windows, macOS, Node.js, HTTP method/path egress rules, live resize, live fork (no pause), memory fork, spend limits, automatic snapshots, snapshots on demand, agent harness API, hosted agent API (their own agent), gVisor or VM (no shared kernel), firewall inside (nftables), secret proxy, secret proxy for any API

Unknown: everything else. Evidence (source + quote) per feature: https://battleships.dev/data/providers/baseten.json → feature_evidence

## Caveats

- Bundled CPU/RAM; billed by minute. Deployment/load/scale time billed; scale-to-zero is not free warm replicas. H100 MIG is 3/7 compute and 1/2 memory, not half a full H100. Custom model packaging via Truss; not a general secure code-exec sandbox.
- Any null count, CPU/RAM, billing increment or minimum is unverified, not unlimited/free.
- This is GPU-product research, not a claim that every feature of the entire vendor documentation was audited. Unverified features are null.
- GPU multi-count bundles, variant/region selection and fractional slices require a SKU-aware estimator; unsupported rows remain unpriced in strict cards.
- Modes with gpu=null are intentionally excluded, not free GPU modes. Never substitute 0 for a null rate. Flags alt indicate DIY/ML infrastructure, not the same thing as managed untrusted-code sandboxes.
- Unknown fee, storage or bandwidth means full monthly total may be unavailable. Rates exclude taxes.
- Larger CPU/RAM shapes on the same GPU cost more (T4x8x32 $0.01504/min, T4x16x64 $0.02408/min, A10Gx16x64 $0.03248/min); multi-GPU SKUs are linear (H100:8 $0.86664/min). Card models only the smallest single-GPU shape.
- Image builds are billed on every push; cold start/model load is billed; failed boots and image pulls are not; development deployments (--watch) are listed as not billed. Partial minutes round up.
- Default autoscaling scale-down delay is 900 s (configurable 0-3600 s): with defaults each burst keeps the replica billed ~15 more idle minutes, plus billed cold start/model load. Not modelled.

## How this provider charges

As of 2026-09-28. Dedicated model deployments and training containers.

Bundled CPU/RAM; billed by minute. Deployment/load/scale time billed; scale-to-zero is not free warm replicas. H100 MIG is 3/7 compute and 1/2 memory, not half a full H100. Custom model packaging via Truss; not a general secure code-exec sandbox.

| Regime | When applicable | Billing | Numbers ($/GPU-hour unless slice) | Source |
|---|---|---|---|---|
| serverless  | T4 ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 0.6312; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | L4 ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 0.8484; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | A10G ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 1.2072; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | A100-80GB ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 4.0002; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | H100 ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 6.4998; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | H200 ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 7.5; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | B200 ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 9.9798; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | RTX-PRO-6000 ; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 4.0002; list; per_gpu_hour | [source](https://docs.baseten.co/deployment/resources) |
| serverless  | H100 MIG 40GB; count [1]; Region/availability must be checked at launch. | per-minute; min 60s | 3.75; list; per_slice_hour | [source](https://docs.baseten.co/deployment/resources) |

## Gotchas
- Bundled CPU/RAM; billed by minute. Deployment/load/scale time billed; scale-to-zero is not free warm replicas. H100 MIG is 3/7 compute and 1/2 memory, not half a full H100. Custom model packaging via Truss; not a general secure code-exec sandbox.
- Any null count, CPU/RAM, billing increment or minimum is unverified, not unlimited/free.
- This is GPU-product research, not a claim that every feature of the entire vendor documentation was audited. Unverified features are null.
- GPU multi-count bundles, variant/region selection and fractional slices require a SKU-aware estimator; unsupported rows remain unpriced in strict cards.

## Worked example (required protocol workload)
4 vCPU / 8 GiB, 50 concurrent × 8 hours/day × 22 days = **8,800 instance-hours/month**. 30% CPU utilization does not reduce allocated GPU uptime. 50 GiB snapshots and 100 GiB egress are additional. The requested CPU-only workload has no GPU type/count, so its full GPU-provider total is **null**, not a fictitious CPU equivalent.

Explicit GPU sensitivity example: add **1 × T4** to each of those 50 workers at the T4 quoted shape. Compute component = 8,800 × $0.6312 = **$5,554.5600**. CPU/RAM are included in that SKU (not independently resizable).
This is compute-only, before platform fees, storage, egress, setup, minimum rounding, cold-start/idle tails and capacity quotas. 50 concurrent GPUs are not promised by a unit rate.

## Evidence scope
Sources read are listed in the provider/features JSON. Raw pricing pages, relevant docs and API payloads are retained under `../raw/`. All unknown feature toggles remain null.
