# Together AI

Official pricing: https://www.together.ai/pricing  
Category: gpu-cloud · Isolation: container

## Pricing regimes (raw)

- **H100 GPU cluster, on-demand (8-GPU nodes, per GPU-hour)** (resource): $0/vCPU-h, $0/GiB-h
- **H200 HGX/SXM, on-demand (×8 node)** (resource)
- **B200 HGX/SXM, on-demand (×8 node)** (resource)
- **B300 HGX/SXM, on-demand (×8 node)** (resource)
- **H100 HGX/SXM, spot (×8 node)** (resource) [spot, beta]
- **H200 HGX/SXM, spot (×8 node)** (resource) [spot, beta]
- **B200 HGX/SXM, spot (×8 node)** (resource) [spot, beta]
- **B300 HGX/SXM, spot (×8 node)** (resource) [spot, beta]

## Features

Yes: persistent disk, volumes, ≥24 h sessions, custom image (Docker or snapshot), your own Docker/OCI image, Docker inside, nested virtualization, SSH, public IPv4, HTTPS preview URLs, static egress IP, SOC 2, HIPAA, SSO, GPU, Python, Node.js, EU data residency, inbound access rules, firewall inside (nftables), extra volumes, shared volumes

No: snapshots, memory snapshots, fork/clone, pause/resume, idle auto-stop, start from your own snapshot, full VM (own kernel), browser, desktop GUI, computer-use API, browser + desktop control, anti-bot stealth, CAPTCHA solving, residential IPs, preinstalled agents, egress allowlist, self-hosting, BYOC, open source, arm64, Windows, macOS, HTTP method/path egress rules, wake on request, live resize, live fork (no pause), memory fork, MCP server, automatic snapshots, snapshots on demand, agent harness API, hosted agent API (their own agent), gVisor or VM (no shared kernel), secret proxy, secret proxy for any API

Unknown: everything else. Evidence (source + quote) per feature: https://battleships.dev/data/providers/together-gpu.json → feature_evidence

## Caveats

- GPU cluster rates not valid for Code Sandbox. H100 dedicated inference $3.99 is a promotion ending 2026-09-30; $5.49 normal list. Cluster reservation windows are not serverless usage rates.
- Any null count, CPU/RAM, billing increment or minimum is unverified, not unlimited/free.
- This is GPU-product research, not a claim that every feature of the entire vendor documentation was audited. Unverified features are null.
- GPU multi-count bundles, variant/region selection and fractional slices require a SKU-aware estimator; unsupported rows remain unpriced in strict cards.
- Modes with gpu=null are intentionally excluded, not free GPU modes. Never substitute 0 for a null rate. Flags alt indicate DIY/ML infrastructure, not the same thing as managed untrusted-code sandboxes.
- Unknown fee, storage or bandwidth means full monthly total may be unavailable. Rates exclude taxes.
- On-demand (H100/H200/B200/B300) and spot cluster modes are priced per GPU with gpu_min_count 8; reserved modes are priced as always-on commit modes; dedicated-inference modes are unpriced. The gpu-clusters page lists H200/B200 cluster scale as '256 to 1,000 GPUs', so gpu_min_count 8 may understate the real minimum for those.
- Reserved clusters: prepaid upfront, non-refundable, cannot be paused, auto-decommissioned at the end of the term (24 h email); usage beyond the reservation is billed at on-demand rates. Docs say reservations are 1-90 days while the pricing page sells 7-30 / 31-90 / 91-180 days and 181+ via sales, and the marketing page says 'Up to 6 months'.
- The clusters API also accepts RTX_6000_PCI, L40_PCIE and H100_SXM_INF, and GB200/GB300 NVL72 are 'Contact us'; none has a published cluster price, so they are not listed in gpu_types.

## How this provider charges

As of 2026-09-28. GPU Clusters and Dedicated Inference; separate from Code Sandbox.

GPU cluster rates not valid for Code Sandbox. H100 dedicated inference $3.99 is a promotion ending 2026-09-30; $5.49 normal list. Cluster reservation windows are not serverless usage rates.

| Regime | When applicable | Billing | Numbers ($/GPU-hour unless slice) | Source |
|---|---|---|---|---|
| on-demand clusters | H100 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 3.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| on-demand clusters | H200 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 5.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| on-demand clusters | B200 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 8.19; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| on-demand clusters | B300 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 9.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| spot clusters | H100 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 1.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| spot clusters | H200 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 2.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| spot clusters | B200 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 4.09; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| spot clusters | B300 HGX/SXM; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 4.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | H100 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 3.69; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | H200 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 4.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | B200 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 7.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | H100 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 3.45; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | H200 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 4.15; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | B200 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 7.79; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | H100 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 3.19; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | H200 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 3.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| reserved clusters | B200 ; count [8]; Region/availability must be checked at launch. | per-hour; min unknowns | 6.79; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| on-demand dedicated-inference | H100 ; count None; Region/availability must be checked at launch. | per-hour; min unknowns | 3.99; promo; per_gpu_hour | [source](https://www.together.ai/pricing) |
| on-demand dedicated-inference-list | H100 ; count None; Region/availability must be checked at launch. | per-hour; min unknowns | 5.49; list; per_gpu_hour | [source](https://www.together.ai/pricing) |
| on-demand dedicated-inference | B200 ; count None; Region/availability must be checked at launch. | per-hour; min unknowns | 8.99; list; per_gpu_hour | [source](https://www.together.ai/pricing) |

## Gotchas
- GPU cluster rates not valid for Code Sandbox. H100 dedicated inference $3.99 is a promotion ending 2026-09-30; $5.49 normal list. Cluster reservation windows are not serverless usage rates.
- Any null count, CPU/RAM, billing increment or minimum is unverified, not unlimited/free.
- This is GPU-product research, not a claim that every feature of the entire vendor documentation was audited. Unverified features are null.
- GPU multi-count bundles, variant/region selection and fractional slices require a SKU-aware estimator; unsupported rows remain unpriced in strict cards.

## Worked example (required protocol workload)
4 vCPU / 8 GiB, 50 concurrent × 8 hours/day × 22 days = **8,800 instance-hours/month**. 30% CPU utilization does not reduce allocated GPU uptime. 50 GiB snapshots and 100 GiB egress are additional. The requested CPU-only workload has no GPU type/count, so its full GPU-provider total is **null**, not a fictitious CPU equivalent.

No fixed single-full-GPU USD quote suitable for a numeric example was verified. Compute = 8,800 × selected offer $/hour, or required node count × node price; keep the estimate null until the offer/FX/commitment is resolved.

## Evidence scope
Sources read are listed in the provider/features JSON. Raw pricing pages, relevant docs and API payloads are retained under `../raw/`. All unknown feature toggles remain null.
