Warming the glyph atlas
H200 clusters now live in 14 regionsH200 now in 14 regions→

Deploy models
in one line.

Cirrus runs your AI models on serverless GPUs. Cold starts under 200ms, scale to zero when idle, pay per second.

Talk to sales
#01 — Platform

Infrastructure
that feels weightless.

Read the docs →

  

Deploy from anywhere

Point Cirrus at a repo, a Dockerfile or a Hugging Face id. It builds, ships and hands you an endpoint.

git push → live endpoint
~ 38 seconds
▮▮▮▮▮▮▯▯

Inference at the edge

Weights stay warm in 14 regions, so the first token leaves close to your users.

p50 cold start, H200
184 ms
▁▂▃▅▇█▅▂

Scales like weather

Traffic spikes, replicas follow. Quiet hours scale to zero and stop billing.

zero → peak, per second
0 → 12,000 rps
▁▁▂▄▆███
#02 — How it works

From your laptop to fourteen skies.

Point Cirrus at a model on Hugging Face or a Dockerfile. We pick the GPU, warm the weights, and give you an endpoint with autoscaling, logs and a bill measured in seconds.

Uptime, last 12 months99.99%
Cheaper than reserved GPUs−61%
Tokens served per day, 30 days2.1B
Ping a region from your timezone
~/acme-support-botClick here to type commands

      
    
#03 — Pricing

Pay for the seconds
you actually fly.

See all GPUs →

  
Hobby
$0
then $0.0004 / GPU-second
  • 1 endpoint, scale to zero
  • A10G & T4 pool
  • Community Slack
  • 7-day logs
Start free
Team · most picked
$400 /mo
includes 1M GPU-seconds
  • Unlimited endpoints
  • H100 & H200 pool, 14 regions
  • Cold start p50 184ms
  • Private networking & SSO
  • 90-day logs and traces
Start 14-day trial
Enterprise
Let's talk
committed capacity, annual
  • Reserved H200 clusters
  • 99.99% uptime SLA
  • SOC 2 Type II · HIPAA · BAA
  • Named solutions engineer
Book a demo