PluginBench
Skill
Pass
Audit score 90

gcp-cloud-run

sickn33/agentic-awesome-skills

Build production-ready serverless applications on GCP Cloud Run with optimized cold starts and event-driven architecture.

What is gcp-cloud-run?

Specialized skill for deploying containerized services and event-driven functions on GCP Cloud Run. Use this when building scalable serverless applications that need cold start optimization, Pub/Sub integration, and proper memory/concurrency configuration.

  • Deploy containerized Cloud Run services with optimized memory allocation
  • Build event-driven Cloud Run Functions integrated with Pub/Sub
  • Monitor and optimize memory usage including /tmp overhead
  • Configure concurrency settings to avoid scaling bottlenecks
  • Implement production-ready serverless architecture on GCP

How to install gcp-cloud-run

npx skills add https://github.com/sickn33/agentic-awesome-skills --skill gcp-cloud-run
Prerequisites
  • GCP project with Cloud Run API enabled
  • Docker container image or function code ready to deploy
  • gcloud CLI configured with appropriate permissions
  • Understanding of Cloud Run memory and concurrency limits
Claude Code
Cursor
Windsurf
Cline

How to use gcp-cloud-run

  1. 1.Review the detailed guide in references/detailed-guide.md before proceeding
  2. 2.Define memory allocation in cloudbuild.yaml, accounting for /tmp usage
  3. 3.Set appropriate concurrency limits (avoid concurrency=1 unless truly single-threaded)
  4. 4.Implement memory monitoring using psutil or equivalent
  5. 5.Deploy using gcloud run deploy with specified memory and concurrency settings
  6. 6.Monitor auto-scaling behavior and adjust concurrency based on traffic patterns

Use cases

Good for
  • Deploying microservices that scale automatically based on traffic
  • Building event-driven workflows triggered by Pub/Sub messages
  • Optimizing cold start performance for latency-sensitive applications
  • Managing memory-intensive workloads with proper resource allocation
  • Scaling applications cost-effectively during traffic spikes
Who it's for
  • Backend engineers deploying on GCP
  • DevOps engineers optimizing serverless infrastructure
  • Full-stack developers building event-driven systems
  • Teams migrating to serverless architecture

gcp-cloud-run FAQ

Why does setting concurrency=1 cause scaling problems?

Concurrency=1 means each container handles only one request. During traffic spikes, 100 concurrent requests require 100 instances, each with cold start overhead, increasing latency and costs. Use only for truly single-threaded or memory-heavy per-request processing.

How should I calculate memory allocation for Cloud Run?

Include /tmp overhead in your memory calculation. Monitor actual usage with psutil and allocate sufficient memory to avoid out-of-memory errors while keeping costs reasonable.

What is the difference between Cloud Run services and Cloud Run Functions?

Cloud Run services are containerized applications you deploy as Docker images. Cloud Run Functions are event-driven functions triggered by events like Pub/Sub messages, with simpler deployment.

How do I integrate Cloud Run with Pub/Sub for event-driven architecture?

Deploy a Cloud Run Function or service that subscribes to Pub/Sub topics. Configure the subscription to push messages to your service's HTTP endpoint.

Full instructions (SKILL.md)

Source of truth, from sickn33/agentic-awesome-skills.


name: gcp-cloud-run description: Specialized skill for building production-ready serverless applications on GCP. Covers Cloud Run services (containerized), Cloud Run Functions (event-driven), cold start optimization, and event-driven architecture with Pub/Sub. risk: critical source: vibeship-spawner-skills (Apache 2.0) date_added: 2026-02-27

GCP Cloud Run

Specialized skill for building production-ready serverless applications on GCP. Covers Cloud Run services (containerized), Cloud Run Functions (event-driven), cold start optimization, and event-driven architecture with Pub/Sub.

Detailed Guide

Read the detailed guide before executing this skill. It retains the complete procedure and reference material. Treat its safety, prerequisites, and validation requirements as mandatory. For focused work, load the relevant sections; for end-to-end work, read the guide completely.

Calculate memory including /tmp usage

# cloudbuild.yaml
steps:
  - name: 'gcr.io/cloud-builders/gcloud'
    args:
      - 'run'
      - 'deploy'
      - 'my-service'
      - '--memory=1Gi'  # Include /tmp overhead
      - '--image=gcr.io/$PROJECT_ID/my-service'

Monitor memory usage

import psutil
import logging

def log_memory():
    memory = psutil.virtual_memory()
    logging.info(f"Memory: {memory.percent}% used, "
                f"{memory.available / 1024 / 1024:.0f}MB available")

Concurrency=1 Causes Scaling Bottlenecks

Severity: HIGH

Situation: Setting concurrency to 1 for request isolation

Symptoms: Auto-scaling creates many container instances. High latency during traffic spikes. Increased cold starts. Higher costs from more instances.

Why this breaks: Setting concurrency to 1 means each container handles only one request at a time. During traffic spikes:

  • 100 concurrent requests = 100 container instances
  • Each instance has cold start overhead
  • More instances = higher costs
  • Scaling takes time, requests queue up

This should only be used when:

  • Processing is truly single-threaded
  • Memory-heavy per-request processing
  • Using thread-unsafe libraries

Recommended fix:

When to Use

Use this skill when the request clearly matches the capabilities and patterns described above.

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.