gcp-cloud-run
sickn33/agentic-awesome-skills
Build production-ready serverless applications on GCP Cloud Run with optimized cold starts and event-driven architecture.
What is gcp-cloud-run?
Specialized skill for deploying containerized services and event-driven functions on GCP Cloud Run. Use this when building scalable serverless applications that need cold start optimization, Pub/Sub integration, and proper memory/concurrency configuration.
- Deploy containerized Cloud Run services with optimized memory allocation
- Build event-driven Cloud Run Functions integrated with Pub/Sub
- Monitor and optimize memory usage including /tmp overhead
- Configure concurrency settings to avoid scaling bottlenecks
- Implement production-ready serverless architecture on GCP
How to install gcp-cloud-run
npx skills add https://github.com/sickn33/agentic-awesome-skills --skill gcp-cloud-run- GCP project with Cloud Run API enabled
- Docker container image or function code ready to deploy
- gcloud CLI configured with appropriate permissions
- Understanding of Cloud Run memory and concurrency limits
How to use gcp-cloud-run
- 1.Review the detailed guide in references/detailed-guide.md before proceeding
- 2.Define memory allocation in cloudbuild.yaml, accounting for /tmp usage
- 3.Set appropriate concurrency limits (avoid concurrency=1 unless truly single-threaded)
- 4.Implement memory monitoring using psutil or equivalent
- 5.Deploy using gcloud run deploy with specified memory and concurrency settings
- 6.Monitor auto-scaling behavior and adjust concurrency based on traffic patterns
Use cases
- Deploying microservices that scale automatically based on traffic
- Building event-driven workflows triggered by Pub/Sub messages
- Optimizing cold start performance for latency-sensitive applications
- Managing memory-intensive workloads with proper resource allocation
- Scaling applications cost-effectively during traffic spikes
- Backend engineers deploying on GCP
- DevOps engineers optimizing serverless infrastructure
- Full-stack developers building event-driven systems
- Teams migrating to serverless architecture
gcp-cloud-run FAQ
Concurrency=1 means each container handles only one request. During traffic spikes, 100 concurrent requests require 100 instances, each with cold start overhead, increasing latency and costs. Use only for truly single-threaded or memory-heavy per-request processing.
Include /tmp overhead in your memory calculation. Monitor actual usage with psutil and allocate sufficient memory to avoid out-of-memory errors while keeping costs reasonable.
Cloud Run services are containerized applications you deploy as Docker images. Cloud Run Functions are event-driven functions triggered by events like Pub/Sub messages, with simpler deployment.
Deploy a Cloud Run Function or service that subscribes to Pub/Sub topics. Configure the subscription to push messages to your service's HTTP endpoint.
Full instructions (SKILL.md)
Source of truth, from sickn33/agentic-awesome-skills.
name: gcp-cloud-run description: Specialized skill for building production-ready serverless applications on GCP. Covers Cloud Run services (containerized), Cloud Run Functions (event-driven), cold start optimization, and event-driven architecture with Pub/Sub. risk: critical source: vibeship-spawner-skills (Apache 2.0) date_added: 2026-02-27
GCP Cloud Run
Specialized skill for building production-ready serverless applications on GCP. Covers Cloud Run services (containerized), Cloud Run Functions (event-driven), cold start optimization, and event-driven architecture with Pub/Sub.
Detailed Guide
Read the detailed guide before executing this skill. It retains the complete procedure and reference material. Treat its safety, prerequisites, and validation requirements as mandatory. For focused work, load the relevant sections; for end-to-end work, read the guide completely.
Calculate memory including /tmp usage
# cloudbuild.yaml
steps:
- name: 'gcr.io/cloud-builders/gcloud'
args:
- 'run'
- 'deploy'
- 'my-service'
- '--memory=1Gi' # Include /tmp overhead
- '--image=gcr.io/$PROJECT_ID/my-service'
Monitor memory usage
import psutil
import logging
def log_memory():
memory = psutil.virtual_memory()
logging.info(f"Memory: {memory.percent}% used, "
f"{memory.available / 1024 / 1024:.0f}MB available")
Concurrency=1 Causes Scaling Bottlenecks
Severity: HIGH
Situation: Setting concurrency to 1 for request isolation
Symptoms: Auto-scaling creates many container instances. High latency during traffic spikes. Increased cold starts. Higher costs from more instances.
Why this breaks: Setting concurrency to 1 means each container handles only one request at a time. During traffic spikes:
- 100 concurrent requests = 100 container instances
- Each instance has cold start overhead
- More instances = higher costs
- Scaling takes time, requests queue up
This should only be used when:
- Processing is truly single-threaded
- Memory-heavy per-request processing
- Using thread-unsafe libraries
Recommended fix:
When to Use
Use this skill when the request clearly matches the capabilities and patterns described above.
Limitations
- Use this skill only when the task clearly matches the scope described above.
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
Related skills
More from sickn33/agentic-awesome-skills and the wider catalog.

github-workflow-automation
Automate GitHub workflows and CI/CD pipelines with AI-assisted patterns and DevOps integration.

i18n-localization
Detect hardcoded strings, manage translations, and implement RTL support across locales.

interactive-portfolio
Build portfolios that convert visitors into job offers and client opportunities.

langgraph
Production-grade framework for building stateful, multi-actor AI applications with explicit graph structure.

last30days
Research any topic from the last 30 days across Reddit, X, and the web to become an expert and generate ready-to-use prompts.

lint-and-validate
Run lint and type checks, distinguish failures from unrun checks, and report concrete validation results.