# Replicate

Official pricing: https://replicate.com/pricing  
Category: gpu-cloud · Isolation: container

## Pricing regimes (raw)

- **Private Cog CPU model** (sizes): cpu-small 1 vCPU/2 GiB $0.09/h; cpu 4 vCPU/8 GiB $0.36/h [alt]
- **T4, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **A100-80GB, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **H100, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **L40S, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **2x A100-80GB, serverless** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **2x L40S, serverless** (resource): $0/vCPU-h, $0/GiB-h [alt]
- **Multi-GPU / H200 (committed spend)** (resource): $0/vCPU-h, $0/GiB-h [alt, sales]

## Features

Yes: idle auto-stop, ≥24 h sessions, custom image (Docker or snapshot), your own Docker/OCI image, HTTPS preview URLs, self-hosting, open source, GPU, Python, Node.js, wake on request, webhooks, audit logs, spend limits, MCP server

No: snapshots, memory snapshots, fork/clone, pause/resume, persistent disk, volumes, start from your own snapshot, full VM (own kernel), Docker inside, nested virtualization, browser, desktop GUI, computer-use API, browser + desktop control, anti-bot stealth, CAPTCHA solving, residential IPs, preinstalled agents, SSH, public IPv4, egress allowlist, static egress IP, BYOC, SSO, arm64, Windows, macOS, HTTP method/path egress rules, live resize, live fork (no pause), memory fork, EU data residency, automatic snapshots, snapshots on demand, agent harness API, hosted agent API (their own agent), gVisor or VM (no shared kernel), inbound access rules, firewall inside (nftables), secret proxy, secret proxy for any API, extra volumes, shared volumes

Unknown: everything else. Evidence (source + quote) per feature: https://battleships.dev/data/providers/replicate.json → feature_evidence

## Caveats

- Public inference charges only processing (some by output/token); private custom Cog models charge setup + idle + active dedicated instance time. Fast-booting fine-tunes are special active-only exceptions.
- Private CPU model deployment is an alt to interactive VM. GPUs include their hosts; do not double-count host CPU/RAM. Several multi-GPU and H200 SKUs require committed-spend contracts.
- Null values mean unknown, not free/unlimited. No model-token cost, tax or unverified ancillary charge included.
- Native browser/device/task/credit prices with unknown CPU/RAM are deliberately non-priceable. Check regime before interpreting any partial resource estimate.
- GPU modes grafted from the v4-gpu research card (per-GPU rates; exact rows in research/gpu.json).
- Replicate joined Cloudflare (announced 2025-11-17) and continues as a distinct brand with an unchanged API; no pricing change observed on the live pricing page as of 2026-09-29.

## How this provider charges


YC: Winter 2020; directory status **Acquired**. Research scope: public-product-researched. Source: https://www.ycombinator.com/companies/replicate

Run machine learning models in the cloud

## Regime table

| Regime | When it applies | How billed | Numbers | Source |
|---|---|---|---|---|
| Private Cog CPU model | alt | sizes; CPU alloc, memory alloc | cpu-small: 1 vCPU/2 GiB $0.09/h; cpu: 4 vCPU/8 GiB $0.36/h | https://replicate.com/pricing |
| Usage | Plan limits apply | Monthly fee; credits only as stated | fee=0; included usage=$0; concurrent=None; max session h=None | https://replicate.com/pricing |
| Retained disk / GiB-month | Lifecycle/scope must be checked | Native add-on | unverified, not $0 | https://replicate.com/pricing |
| Snapshots / GiB-month | Lifecycle/scope must be checked | Native add-on | unverified, not $0 | https://replicate.com/pricing |
| Egress / GiB | Lifecycle/scope must be checked | Native add-on | unverified, not $0 | https://replicate.com/pricing |
| IPv4 / month | Lifecycle/scope must be checked | Native add-on | unverified, not $0 | https://replicate.com/pricing |

## Verified billing facts and gotchas

1. Public inference charges only processing (some by output/token); private custom Cog models charge setup + idle + active dedicated instance time. Fast-booting fine-tunes are special active-only exceptions.
2. Private CPU model deployment is an alt to interactive VM. GPUs include their hosts; do not double-count host CPU/RAM. Several multi-GPU and H200 SKUs require committed-spend contracts.

## Worked example

Requested: 4 vCPU/8 GiB, 50 simultaneous × 8 h/day × 22 days = **8,800 instance-hours**, 30% CPU; 50 GiB snapshots and 100 GiB egress. A month is 730 hours only when explicitly used in rate conversion.

CPU 4/8 private model costs 8,800×$0.36=$3,168, plus billed setup and any idle replica time. Public inference only charges processing but cannot host arbitrary interactive workspaces. Snapshot, egress, 50-replica and 8-hour-request eligibility unresolved.

## Sources and coverage

- https://www.ycombinator.com/companies/replicate
- https://replicate.com/pricing
- https://replicate.com/docs/guides/deploy-a-custom-model
- https://replicate.com/
- https://replicate.com/docs

Page captures and documentation indexes/bundles are in `raw/`; provider JSON lists the evidence paths. HN query results were captured separately as `<id>--hn.txt`; they are discovery leads, not current tariff authority. Downloaded docs do not imply every feature was verified; unsupported assertions remain null.
