
Google Cloud Run
by Google Cloud
Serverless containers that scale to zero, with GPUs.
Score
Score
Our verdict
Cloud Run remains the most developer-friendly serverless container platform in 2026. It now spans three core modes: request-based services with scale-to-zero, instance-based services for predictable, always-on compute, and run-to-completion jobs for batch and scheduled tasks. Pricing is strictly usage-based with a sizeable always‑free tier, and the published us‑central1 rates remain easy to estimate ($0.000024/vCPU‑s active for request-based; $0.000018/vCPU‑s for instance-based). The big 2026 story is GPU-backed services: on-demand NVIDIA L4 (and RTX 6000 options) let teams stand up inference endpoints that also scale to zero. Operational features—revisions, gradual traffic splitting, managed TLS, direct VPC egress—are mature, and the 99.95% SLO inspires confidence. Downsides: single requests cap at 60 minutes and you still can’t run privileged workloads, so some stateful or kernel-dependent tasks belong on GKE or Compute Engine. Compared with AWS App Runner (now limited and not GPU-capable) and Azure Container Apps (strong, with KEDA and jobs), Cloud Run offers the cleanest UX-to-scale-to-zero pipeline and the most turnkey GPU inference story.
Overview
Screenshot
Score breakdown
Overall score
Scores are editorial assessments by the Compare Best AI team on a 0–10 scale.
Expert review
CBAI Editorial Team
Compare Best AI · Editorial Team
## Overview
How we tested
Days tested
7 days
Tasks evaluated
- ·Deployed a Flask API with request-based billing and scale-to-zero
- ·Provisioned an NVIDIA L4 GPU service and ran image inference
- ·Set up source-based deploys via buildpacks and traffic splitting
- ·Created a scheduled Cloud Run job and validated logs/costs
Method
Compared against AWS App Runner and Azure Container Apps using identical inputs
Reviewer
CBAI Editorial Team
Plans & pricing
Free tier
Testing, low-traffic apps, scheduled jobs
- Request-based billing free: 2M requests/mo; first 180,000 vCPU-s and 360,000 GiB-s (us-central1)
- Instance-based billing free: first 240,000 vCPU-s and 450,000 GiB-s (us-central1)
- Network egress: first 1 GiB/mo from North America free
- Limits vary by region and billing model
Pay-as-you-go
Production services that scale with demand
- Request-based (us-central1): CPU active $0.000024/vCPU-s; Memory active $0.0000025/GiB-s; Requests $0.40 per 1M after free
- Idle (min instances): CPU $0.0000025/vCPU-s; Memory $0.0000025/GiB-s
- Instance-based (us-central1): CPU $0.000018/vCPU-s; Memory $0.000002/GiB-s
- GPU (sample): NVIDIA L4 from $0.0001867/s
- Ephemeral disk: ~$0.000109589 per GiB-hour
- Rates vary by region; committed-use discounts optional
Pricing may vary by region. Always verify on the vendor's website.
Feature comparison
| Feature | Google Cloud Run | AWS App Runner | Azure Container Apps |
|---|---|---|---|
| Scaling | |||
| Scale to zero on idle | |||
| Compute | |||
| GPU support for services | |||
| Protocols | |||
| HTTP/2 and gRPC (end-to-end) | |||
| Delivery | |||
| Traffic splitting and gradual rollouts | |||
| Events | |||
| Managed event triggers (Eventarc/KEDA) | |||
| Jobs | |||
| Run-to-completion jobs | |||
| Domains | |||
| Custom domains with managed TLS | |||
Is it right for you?
Good fit for
API and microservice teams
Fast, low‑ops deployments with autoscaling, traffic splitting, and managed TLS reduce operational overhead.
AI/ML engineers
Spin up GPU‑backed inference endpoints on demand; scale to zero to avoid idle costs.
Data & platform engineers
Use Cloud Run jobs for scheduled or event-driven batch tasks without standing up clusters.
Startups
Generous free tier and granular usage pricing make early-stage costs predictable.
Less suited for
Long-running stateful apps
Single requests are capped at 60 minutes; stateful workloads fit better on GKE or Compute Engine.
Custom kernels or privileged containers
Doesn’t support low-level host modifications; use Compute Engine or GKE with node access.
GPU training at large scale
GPU support targets inference; large multi-GPU training is better on GKE or Vertex AI.
User reviews
Editorial score
Distribution is estimated from our editorial score. Verified user reviews coming soon.
Use cases
Typical ways teams rely on this tool — from everyday tasks to specialized workflows.
- Host REST/GraphQL APIs
- AI inference with NVIDIA L4
- Event-driven functions
- Batch jobs and schedulers
- WebSockets or gRPC services
Integrations
Reported connectors
Apps and services commonly connected out of the box or via official connectors.
- Cloud SQL
- Pub/Sub
- Eventarc
Privacy & compliance
Security & data practices
Highlights we track from the vendor's documentation. Always confirm current terms, subprocessors, and regional policies on their official site.
- SOC 2
- SOC 3
- HIPAA eligible
- NIST 800-171
Details
Category
Price
- Free tier
- Pay-as-you-go (usage-based, no monthly license)
- Committed Use Discounts (1-year/3-year)
Free version
Best for
- Stateless web APIs and backends that need scale-to-zero
- Event-driven functions consolidated on one platform
- GPU inference endpoints with L4 GPUs
- Batch jobs and scheduled processing with Cloud Run jobs.
Frequently asked questions
Google Cloud Run
Infrastructure & Hosting
Ready to get started?
Visit the Google Cloud Run website to explore plans and start your free trial.
Was this page helpful?
Compare Best AI may earn a commission when you click links on this page. This does not influence our editorial scores or recommendations. Advertiser disclosure