# fal

Official pricing: https://fal.ai/pricing  
Category: gpu-cloud · Isolation: container

## Pricing regimes (raw)

- **B300, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h
- **B300, reserved** (resource) [sales]
- **B200, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h
- **B200, reserved** (resource) [sales]
- **H200, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h
- **H200, reserved** (resource) [sales]
- **H100, serverless (×1)** (resource): $0/vCPU-h, $0/GiB-h
- **H100, reserved** (resource) [sales]

## Features

Yes: persistent disk, volumes, idle auto-stop, ≥24 h sessions, custom image (Docker or snapshot), your own Docker/OCI image, Docker inside, SSH, public IPv4, HTTPS preview URLs, SOC 2, SSO, GPU, Python, Node.js, wake on request, MCP server, hosted agent API (their own agent), inbound access rules, firewall inside (nftables), extra volumes, shared volumes

No: snapshots, memory snapshots, fork/clone, pause/resume, start from your own snapshot, full VM (own kernel), nested virtualization, browser, desktop GUI, computer-use API, browser + desktop control, anti-bot stealth, CAPTCHA solving, residential IPs, preinstalled agents, egress allowlist, static egress IP, self-hosting, BYOC, open source, HIPAA, arm64, Windows, macOS, HTTP method/path egress rules, live resize, live fork (no pause), memory fork, EU data residency, automatic snapshots, snapshots on demand, agent harness API, gVisor or VM (no shared kernel), secret proxy, secret proxy for any API

Unknown: everything else. Evidence (source + quote) per feature: https://battleships.dev/data/providers/fal.json → feature_evidence

## Caveats

- GPU time is not interchangeable with model API output pricing. Custom deployment requires contacting support. The as-low-as column is conditional, not generally available on-demand pricing. Serverless CPU/RAM bundles verified in machine-types docs; separate Compute VMs have different shape and unpublished dashboard rates. Billing starts at SETUP and includes IDLE/RUNNING/DRAINING/TERMINATING; PENDING and DOCKER_PULL unbilled. This is not active-request-only pricing.
- Any null count, CPU/RAM, billing increment or minimum is unverified, not unlimited/free.
- This is GPU-product research, not a claim that every feature of the entire vendor documentation was audited. Unverified features are null.
- GPU multi-count bundles, variant/region selection and fractional slices require a SKU-aware estimator; unsupported rows remain unpriced in strict cards.
- Modes with gpu=null are intentionally excluded, not free GPU modes. Never substitute 0 for a null rate. Serverless GPU regimes are ordinary list-priced GPU instances (alt removed 2026-09-28); reserved regimes stay sales.
- Unknown fee, storage or bandwidth means full monthly total may be unavailable. Rates exclude taxes.
- Access gate (2026-09-28): fal.ai/docs/documentation/serverless says "Serverless access is approved per account ... the fal team must approve access" (request at fal.ai/dashboard/serverless-get-started). List prices apply once approved; not negotiated, so not flagged sales. Treat as a minor availability caveat.
- Further serverless machine types exist with no public price: CPU XS/S/M/L/XL (0.5-8 cores, 0.5-30 GB), GPU-A100 (40 GB, 12 CPU, 60 GB RAM), GPU-L40 (48 GB, 6 CPU, 100 GB RAM); the FAQ also names RTX 4090, RTX 5090, H100 MIG, A100 80GB, L40S. B300 is on the price pages but not in the machine-types doc.
- B200 VRAM now 192GB on fal.ai/serverless (fal.ai/pricing still 180GB), resolving the card's 180 vs 192 GB note in favour of the docs.
- Two official fal price tables disagree (2026-09-29): fal.ai/pricing (card source) lists B300 $8.50, B200 $6.25, H200 $4.50, H100 $4.50, RTX PRO 6000 $2.99, while fal.ai/serverless lists B300 $12.99, GB200 $9.99, B200 $7.99, H200 $6.00, H100 $4.50, RTX PRO 6000 $4.00. Card GPU rates may be 28-53% low except H100; fal says to contact sales for the latest pricing.
- Serverless billing tail: keep_alive defaults to 60 s of billed IDLE time per runner spin-up (configurable, 0 allowed); termination_grace_period defaults to 5 s (max 1 h) and is billed; min_concurrency runners bill 24/7; 5xx responses are not charged; multi-GPU is billed gpu_count x duration. Short sporadic requests therefore pay about 60+ s of GPU each unless keep_alive is lowered.

## How this provider charges

As of 2026-09-28. Serverless custom apps / compute.

GPU time is not interchangeable with model API output pricing. Custom deployment requires contacting support. The as-low-as column is conditional, not generally available on-demand pricing. Serverless CPU/RAM bundles verified in machine-types docs; separate Compute VMs have different shape and unpublished dashboard rates. Billing starts at SETUP and includes IDLE/RUNNING/DRAINING/TERMINATING; PENDING and DOCKER_PULL unbilled. This is not active-request-only pricing.

| Regime | When applicable | Billing | Numbers ($/GPU-hour unless slice) | Source |
|---|---|---|---|---|
| serverless  | B300 ; count [1]; Region/availability must be checked at launch. | per-second; min unknowns | 8.5; list; per_gpu_hour | [source](https://fal.ai/pricing) |
| reserved  | B300 ; count None; Region/availability must be checked at launch. | unknown; min unknowns | 4.49; list-from; per_gpu_hour | [source](https://fal.ai/pricing) |
| serverless  | B200 ; count [1]; Region/availability must be checked at launch. | per-second; min unknowns | 6.25; list; per_gpu_hour | [source](https://fal.ai/pricing) |
| reserved  | B200 ; count None; Region/availability must be checked at launch. | unknown; min unknowns | 3.49; list-from; per_gpu_hour | [source](https://fal.ai/pricing) |
| serverless  | H200 ; count [1]; Region/availability must be checked at launch. | per-second; min unknowns | 4.5; list; per_gpu_hour | [source](https://fal.ai/pricing) |
| reserved  | H200 ; count None; Region/availability must be checked at launch. | unknown; min unknowns | 2.1; list-from; per_gpu_hour | [source](https://fal.ai/pricing) |
| serverless  | H100 ; count [1]; Region/availability must be checked at launch. | per-second; min unknowns | 4.5; list; per_gpu_hour | [source](https://fal.ai/pricing) |
| reserved  | H100 ; count None; Region/availability must be checked at launch. | unknown; min unknowns | 1.89; list-from; per_gpu_hour | [source](https://fal.ai/pricing) |
| serverless  | RTX-PRO-6000 ; count [1]; Region/availability must be checked at launch. | per-second; min unknowns | 2.99; list; per_gpu_hour | [source](https://fal.ai/pricing) |
| reserved  | RTX-PRO-6000 ; count None; Region/availability must be checked at launch. | unknown; min unknowns | 1.1; list-from; per_gpu_hour | [source](https://fal.ai/pricing) |
| on-demand dedicated-compute | H100 SXM; count [1]; Dedicated Compute dashboard; region unspecified. | per-hour; min unknowns | unknown; sales-unpublished; per_gpu_hour | [source](https://fal.ai/docs/documentation/compute/pricing.md) |
| on-demand dedicated-compute | H100 SXM; count [8]; Dedicated Compute dashboard; region unspecified. | per-hour; min unknowns | unknown; sales-unpublished; per_gpu_hour | [source](https://fal.ai/docs/documentation/compute/pricing.md) |

## Gotchas
- GPU time is not interchangeable with model API output pricing. Custom deployment requires contacting support. The as-low-as column is conditional, not generally available on-demand pricing. Serverless CPU/RAM bundles verified in machine-types docs; separate Compute VMs have different shape and unpublished dashboard rates. Billing starts at SETUP and includes IDLE/RUNNING/DRAINING/TERMINATING; PENDING and DOCKER_PULL unbilled. This is not active-request-only pricing.
- Any null count, CPU/RAM, billing increment or minimum is unverified, not unlimited/free.
- This is GPU-product research, not a claim that every feature of the entire vendor documentation was audited. Unverified features are null.
- GPU multi-count bundles, variant/region selection and fractional slices require a SKU-aware estimator; unsupported rows remain unpriced in strict cards.

## Worked example (required protocol workload)
4 vCPU / 8 GiB, 50 concurrent × 8 hours/day × 22 days = **8,800 instance-hours/month**. 30% CPU utilization does not reduce allocated GPU uptime. 50 GiB snapshots and 100 GiB egress are additional. The requested CPU-only workload has no GPU type/count, so its full GPU-provider total is **null**, not a fictitious CPU equivalent.

Explicit GPU sensitivity example: add **1 × B300** to each of those 50 workers at the B300 quoted shape. Compute component = 8,800 × $8.5 = **$74,800.0000**. CPU/RAM are included in that SKU (not independently resizable).
This is compute-only, before platform fees, storage, egress, setup, minimum rounding, cold-start/idle tails and capacity quotas. 50 concurrent GPUs are not promised by a unit rate.

## Evidence scope
Sources read are listed in the provider/features JSON. Raw pricing pages, relevant docs and API payloads are retained under `../raw/`. All unknown feature toggles remain null.
