Quick Verdict Google Cloud Run is worth it for teams that want to run containers without managing servers. It scales to zero, bills per 100 milliseconds, includes a free tier of 2 million requests a month (request-based billing), and in 2026 supports NVIDIA L4 and RTX PRO 6000 GPUs for AI inference. Limits to know: 8 vCPU and 32 GiB per standard instance, a 60-minute request timeout, and no privileged or stateful workloads. CBAI score: 8.5/10. Best for APIs, web apps, background jobs and pay-per-use AI inference.
Google Cloud Run is Google Cloud's fully managed platform for running containers as services, jobs and worker pools. You hand it a container image, and it handles servers, scaling, HTTPS and load balancing. This review covers features, both billing modes, GPU pricing, real cost examples, limits and alternatives, with prices in USD for the us-central1 (Iowa) region.
Google Cloud Run at a glance
| Details (checked October 1, 2026) | |
|---|---|
| What it is | Serverless container platform (services, jobs, worker pools) |
| Provider | Google Cloud |
| Runs | Any container image; HTTP/1, HTTP/2, gRPC, WebSockets |
| Scaling | Automatic, including to zero |
| Billing | Per 100 ms; request-based or instance-based |
| Free tier (request-based) | 2M requests, 180,000 vCPU-seconds, 360,000 GiB-seconds per month |
| CPU price | $0.000024/vCPU-s (request-based); $0.000018/vCPU-s (instance-based) |
| GPUs | NVIDIA L4 (24 GB) and RTX PRO 6000 Blackwell (96 GB), one per instance |
| Max per standard instance | 8 vCPU, 32 GiB memory, 1,000 concurrent requests |
| Timeouts | 60 minutes per request; jobs up to 7 days (1 hour with GPUs) |
| CBAI score | 8.5/10 |
Sources: Cloud Run pricing, Cloud Run quotas, Cloud Run GPU docs and CompareBestAI's Google Cloud Run tool page.
Google Cloud Run Scorecard
| Category | Score | Why |
|---|---|---|
| Features and tools | 9.0/10 | Services, jobs, worker pools, GPUs, revisions and traffic splitting |
| Output quality (reliability) | 8.5/10 | Managed TLS, automatic scaling, 99.95% SLO per our tool page |
| API and integrations | 8.5/10 | Cloud SQL, Pub/Sub, Eventarc, Vertex AI, BigQuery |
| Ease of use | 8.0/10 | Deploy a container in one command; networking can get complex |
| Value for money | 8.0/10 | Generous free tier; costs rise with always-on or GPU workloads |
| Support and docs | 7.5/10 | Thorough docs; responsive support needs a paid support plan |
| Overall | 8.5/10 | One of the most practical serverless container platforms available |
What Is Google Cloud Run?
Google Cloud Run is a managed compute platform that runs stateless containers. Unlike a virtual machine, you don't size or patch servers. Unlike Kubernetes, you don't manage clusters. You deploy a container image (or source code that Cloud Run builds into one), and Cloud Run starts as many instances as traffic needs, then shuts them down when traffic stops.
It runs three kinds of workloads:
- Services: HTTP endpoints such as APIs, websites and webhooks
- Jobs: run-to-completion tasks such as batch processing, data imports and scheduled reports
- Worker pools: always-on background workers that pull from queues
Key Features of Google Cloud Run
Scale to Zero and Automatic Scaling
Cloud Run adds instances as requests arrive and removes them when traffic drops, down to zero. With request-based billing you pay only while instances handle requests, which makes low-traffic APIs and internal tools very cheap.
GPUs for AI Inference
Cloud Run supports one GPU per instance: an NVIDIA L4 with 24 GB of VRAM or an NVIDIA RTX PRO 6000 Blackwell with 96 GB. Google says GPU instances with preinstalled drivers start in about 5 seconds and can scale to zero. L4 instances need at least 4 vCPU and 16 GiB of memory (8 vCPU and 32 GiB recommended); RTX PRO 6000 instances need at least 20 vCPU and 80 GiB. Typical uses are LLM inference with open models (for example served with Ollama), image generation, video transcoding and 3D rendering. GPUs are aimed at inference, not large-scale training.
Revisions and Traffic Splitting
Every deployment creates an immutable revision. You can send a percentage of traffic to a new revision for canary releases and roll back instantly.
Built-In HTTPS and Custom Domains
Each service gets an HTTPS URL with a managed certificate. You can map custom domains or put Cloud Run behind a global load balancer.
Google Cloud Integrations
Cloud Run connects to Cloud SQL, Pub/Sub, Eventarc, Cloud Scheduler, Secret Manager, BigQuery and Vertex AI. For app builders, Firebase Studio can deploy to Google Cloud infrastructure.
Protocol Support
Services support HTTP/1, HTTP/2, gRPC and WebSockets, so real-time apps and streaming APIs work without extra setup.
Google Cloud Run Pricing (2026)
Cloud Run bills compute in 100-millisecond increments. Prices below are for Tier 1 region us-central1; Tier 2 regions (such as London, Tokyo and Sydney) cost more.
Request-Based Billing
You pay for CPU and memory only while an instance handles requests, plus a per-request fee.
| Resource | Price | Free each month |
|---|---|---|
| CPU (active) | $0.000024 per vCPU-second | 180,000 vCPU-seconds |
| Memory (active) | $0.0000025 per GiB-second | 360,000 GiB-seconds |
| Requests | $0.40 per million | 2 million |
| Idle minimum instances (CPU and memory) | $0.0000025 per vCPU-second and per GiB-second | None |
Instance-Based Billing
You pay for the whole lifetime of each instance, with no per-request fee and lower unit prices.
| Resource | Price | Free each month |
|---|---|---|
| CPU | $0.000018 per vCPU-second | 240,000 vCPU-seconds |
| Memory | $0.000002 per GiB-second | 450,000 GiB-seconds |
Jobs use the same rates and free amounts as instance-based billing. Worker pools are cheaper per second ($0.000011244 per vCPU-second).
GPU Pricing
| GPU | Non-zonal (per second) | About per hour | Zonal redundancy (per second) |
|---|---|---|---|
| NVIDIA L4 | $0.0001867 | $0.67 | $0.0002909 |
| NVIDIA RTX PRO 6000 Blackwell | $0.00036522 | $1.31 | $0.00056913 |
CPU and memory are billed on top of the GPU. Zonal redundancy reserves capacity across zones for failover; turning it off is cheaper but best-effort.
Discounts and Other Costs
- Committed use discounts: about 17% off request-based and instance-based CPU and memory with a 1- or 3-year commitment; flexible CUDs go deeper on jobs and worker pools.
- Networking: 1 GiB of outbound data per month is free within North America; traffic to Google Cloud services in the same region is free.
- Related services such as Artifact Registry, Cloud Build, Cloud SQL and load balancers are billed separately.
Estimate your own bill with the Google Cloud Pricing Calculator.
What Cloud Run Actually Costs: 3 Examples
These are CompareBestAI calculations from Google's published us-central1 rates, checked October 1, 2026.
1. Small API (request-based billing). 3 million requests a month, 200 ms each, 1 vCPU and 0.5 GiB, assuming no overlapping requests:
- CPU: 600,000 vCPU-seconds minus 180,000 free = 420,000 × $0.000024 = $10.08
- Memory: 300,000 GiB-seconds, within the 360,000 free = $0
- Requests: 1 million above the free 2 million × $0.40 = $0.40
- Total: about $10.48 a month. With concurrent requests sharing an instance, real CPU time and cost are usually lower.
2. Always-on service (instance-based billing). One instance with 1 vCPU and 2 GiB running all month (2,592,000 seconds):
- CPU: 2,592,000 minus 240,000 free = 2,352,000 × $0.000018 = $42.34
- Memory: 5,184,000 GiB-seconds minus 450,000 free = 4,734,000 × $0.000002 = $9.47
- Total: about $51.81 a month.
3. GPU inference (on demand). One L4 GPU running 2 hours a day for 30 days (216,000 seconds), non-zonal:
- GPU: 216,000 × $0.0001867 = about $40.33, plus CPU and memory for the instance
- Because it scales to zero, you pay nothing for the other 22 hours a day.
Google Cloud Run Limits to Know
| Limit | Value |
|---|---|
| vCPU per standard instance | Up to 8 |
| Memory per standard instance | Up to 32 GiB |
| Concurrent requests per instance | Up to 1,000 |
| Request timeout (services) | Up to 60 minutes |
| Task timeout (jobs) | Up to 7 days; 1 hour with GPUs |
| HTTP/1 request size | 32 MiB |
| GPUs per instance | 1 |
Source: Cloud Run quotas and limits and GPU docs, checked October 1, 2026. GPU instances use larger CPU and memory configurations than standard instances.
Pros and Cons of Google Cloud Run
Pros
- Scale to zero: idle services cost nothing on request-based billing
- Generous free tier: 2 million requests a month covers many side projects and internal tools
- GPUs on demand: L4 and RTX PRO 6000 for inference, billed per second
- Any language, any container: no runtime lock-in at the code level
- Mature operations: revisions, traffic splitting, managed TLS and a 99.95% SLO
- Strong Google Cloud integrations: Pub/Sub, Eventarc, Cloud SQL, Vertex AI
Cons
- 60-minute request limit: long-running work must move to jobs or worker pools
- Cold starts: scaled-to-zero services take time to start; set minimum instances (billed at idle rates) for latency-sensitive apps
- No privileged workloads and limited stateful support; use GKE or VMs for those
- Pricing has many parts: two billing modes, request fees, regional tiers and separate networking
- GPUs are inference-focused: one GPU per instance; not for large-scale training
- Some Google Cloud lock-in: containers are portable, but integrations and IAM aren't
Who Should Use Google Cloud Run?
Cloud Run is a strong fit for:
- Startups and product teams running APIs and web backends without a platform team
- Spiky or low-traffic services that benefit from scale-to-zero
- AI teams serving open models or generating media on GPUs only when requests arrive
- Data teams running scheduled jobs and batch processing
- Google Cloud users who want tight integration with Pub/Sub, BigQuery and Vertex AI
Consider something else if you:
- Need long-lived connections or jobs beyond the timeouts above
- Run stateful databases or privileged containers
- Want a flat monthly bill for always-on workloads; a VPS may cost less
- Need multi-GPU training at scale
Google Cloud Run vs Alternatives
| Tool | Best for | Starting price (per our tool pages) | CBAI score |
|---|---|---|---|
| Google Cloud Run | Serverless containers with scale to zero and GPUs | Free tier, then per-second | 8.5/10 |
| Cloudflare | Edge functions and AI inference close to users | Free; Pro $20/month | 8.5/10 |
| DigitalOcean | Simple, predictable pricing for small apps | $4/month (Droplet); App Platform from $5 | 7.8/10 |
| Vultr | Affordable global VMs and GPUs | Billed hourly; no annual plans | 7.5/10 |
| RunPod | GPU-heavy AI workloads and serverless GPU endpoints | Per-second GPU billing | 8.0/10 |
| AWS Lightsail | Fixed-price VPS on AWS | $5/month ($3.50 IPv6-only) | 7.5/10 |
| Hetzner Cloud | Low-cost VMs, especially in Europe | €3.79/month | 8.3/10 |
Within the big clouds, Cloud Run's closest equivalents are AWS App Runner and Fargate, and Azure Container Apps. Cloud Run's main advantages are request-based billing with scale to zero and on-demand GPUs in the same product. For a broader roundup, see our infrastructure and hosting tools guide and Vultr pricing breakdown.
How to Get Started With Google Cloud Run
- Create a Google Cloud project and enable billing (the free tier still requires a billing account).
- Install the gcloud CLI and sign in.
- Deploy from a container image or directly from source, for example
gcloud run deploy my-service --source . --region us-central1. - Set limits: CPU, memory, concurrency, minimum and maximum instances.
- Set a budget alert in Cloud Billing so a traffic spike or a forgotten GPU service doesn't surprise you.
- Add GPUs only when needed, and turn off zonal redundancy for non-critical inference to save money.
See Google Cloud Run's score and details →
How We Reviewed Google Cloud Run
CompareBestAI reviews every tool with the same framework, described in how we rank tools and our editorial guidelines. For this review we:
- Read Google's Cloud Run pricing, quotas and GPU documentation on October 1, 2026
- Calculated the three cost examples from published us-central1 rates
- Scored six categories using the scores on our Google Cloud Run tool page
- Compared alternatives using each tool's CompareBestAI page
We did not run a new load test for this update, and we removed usage statistics we couldn't source.
Google Cloud Run FAQs
Is Google Cloud Run free?
Partly. Each billing account gets a monthly free tier. With request-based billing, that's 2 million requests, 180,000 vCPU-seconds and 360,000 GiB-seconds. With instance-based billing, it's 240,000 vCPU-seconds and 450,000 GiB-seconds. You still need a billing account to start.
How much does Google Cloud Run cost?
In us-central1, request-based billing costs $0.000024 per vCPU-second, $0.0000025 per GiB-second and $0.40 per million requests after the free tier. Instance-based billing costs $0.000018 per vCPU-second and $0.000002 per GiB-second. A small API with 3 million requests a month can cost about $10.
Does Google Cloud Run support GPUs?
Yes. Cloud Run supports one NVIDIA L4 (24 GB) or NVIDIA RTX PRO 6000 Blackwell (96 GB) GPU per instance. GPU instances start in about 5 seconds, scale to zero, and cost from $0.0001867 per second for an L4 in us-central1.
What is the Cloud Run request timeout?
Cloud Run services allow up to 60 minutes per request. Cloud Run jobs allow tasks of up to 7 days, or 1 hour when using GPUs. Longer work should be split into jobs or moved to worker pools.
Does Cloud Run have cold starts?
Yes. When a service has scaled to zero, the first request waits for an instance to start. To avoid this for latency-sensitive apps, set a minimum number of instances, which are billed at lower idle rates.
What is the difference between Cloud Run and GKE?
Cloud Run is fully managed: you deploy containers and Google runs the infrastructure. GKE is managed Kubernetes: you get more control, stateful workloads and privileged containers, but you manage clusters and pay for nodes even when idle.
Is Google Cloud Run good for AI apps?
Yes, for inference. Cloud Run can serve open models on L4 or RTX PRO 6000 GPUs, scale to zero between requests, and connect to Vertex AI. It isn't designed for large-scale model training.
Final Verdict: Google Cloud Run Review 2026
Google Cloud Run is one of the most practical ways to run containers in 2026. It's cheap at low traffic, scales automatically, and now covers AI inference with on-demand GPUs, which removes the biggest gap it used to have. The main trade-offs are its timeouts, its multi-part pricing, and the fact that always-on workloads can cost more than a simple VPS.
Choose Cloud Run for APIs, web apps, jobs and pay-per-use AI inference. Choose an alternative like DigitalOcean for a flat monthly bill, or RunPod for heavy GPU work.
Google Cloud Run: 8.5/10. Best for serverless containers and on-demand AI inference.
Ready to test it? Start with the Cloud Run free tier.
Read next: every tool in our Infrastructure & Hosting category.
About CompareBestAI: CompareBestAI compares AI tools and the software businesses run on, with transparent pricing checks and a consistent scoring method. Get new tool reviews in our newsletter →


