gke-cost
google/skills
Optimize GKE costs through Spot VMs, pod rightsizing, and Committed Use Discounts.
What is gke-cost?
Reduces GKE expenses by rightsizing workloads, configuring Spot VMs for fault-tolerant services, and applying Committed Use Discounts. Use when optimizing cluster costs or reducing over-provisioned pod resources; skip for general compute class provisioning or GPU selection.
- Configure Spot VMs with automatic fallback for 60-90% cost reduction on fault-tolerant workloads
- Deploy Vertical Pod Autoscaler (VPA) to identify and eliminate over-provisioned CPU and memory requests
- Apply Committed Use Discounts (1-year or 3-year) to steady-state workloads for 20-55% savings
- Right-size node pools and machine types using cost-optimized families (e2, t2a, n4a)
- Monitor cluster cost breakdown and pod resource utilization against requests
How to install gke-cost
npx skills add https://github.com/google/skills --skill gke-cost- GKE cluster (Autopilot or Standard) with kubectl access
- Vertical Pod Autoscaler (VPA) for pod rightsizing recommendations
- Cost Management API enabled for billing insights
- Workloads must be fault-tolerant (≥2 replicas) to use Spot VMs safely
How to use gke-cost
- 1.Deploy a VerticalPodAutoscaler in recommendation mode (updateMode: Off) targeting your Deployment
- 2.Wait 24+ hours for VPA to collect resource usage data and generate recommendations
- 3.Review VPA status recommendations and reduce CPU/memory requests where actual usage is significantly lower
- 4.Create a ComputeClass with Spot VM priority and fallback to on-demand for fault-tolerant workloads
- 5.Add nodeSelector cloud.google.com/gke-provisioning: Spot to Pod specs with ≥2 replicas and terminationGracePeriodSeconds < 30s
- 6.Purchase 1-year or 3-year Committed Use Discounts via Google Cloud Console for steady-state workloads
- 7.Monitor cluster costs and pod utilization using kubectl top and gcloud billing commands
Use cases
- Reduce costs for batch processing and data pipeline jobs using Spot VMs with graceful shutdown
- Identify over-provisioned Deployments via VPA recommendations and adjust requests downward
- Purchase 1-year or 3-year Committed Use Discounts for production workloads with stable resource needs
- Stop idle dev/test clusters to eliminate control plane fees during off-hours
- Consolidate multiple single-tenant clusters into shared multi-tenant clusters to reduce overhead
- Platform engineers optimizing GKE cluster expenses
- DevOps teams managing production and dev/test environments
- Cost-conscious organizations running fault-tolerant batch or stateless workloads
- Teams using GKE Autopilot seeking to minimize per-pod billing
gke-cost FAQ
Use Spot VMs for fault-tolerant workloads (batch jobs, stateless APIs with ≥2 replicas, dev/test environments) where 30-second preemption is acceptable. Avoid for single-replica critical services and stateful workloads like databases.
Spot VMs can be preempted at any time with a 30-second notice. Set terminationGracePeriodSeconds < 30s and ensure workloads run with at least 2 replicas for high availability.
VPA requires 24+ hours of resource usage data collection before generating reliable recommendations. Deploy in recommendation mode (updateMode: Off) first to review suggestions before applying.
1-year CUDs provide ~20-30% discount; 3-year CUDs provide ~50-55% discount. Discounts apply automatically to matching usage in the region and are best for steady-state workloads.
Yes. GKE Autopilot includes OPTIMIZE_UTILIZATION autoscaling, Vertical Pod Autoscaling, and Node Auto Provisioning by default. This skill covers additional optimizations like Spot VMs and CUDs.
Full instructions (SKILL.md)
Source of truth, from google/skills.
name: gke-cost description: >- Optimizes GKE costs, rightsizes workloads, and configures Spot VMs and CUDs. Use when optimizing GKE costs, rightsizing GKE workloads, or configuring GKE Spot VMs. Don't use for general compute class provisioning or GPU Selection (use gke-compute-classes instead). metadata: category: CloudObservabilityAndMonitoring
GKE Cost Optimization
This reference covers strategies for reducing GKE costs while maintaining the golden path security and reliability posture.
MCP Tools:
get_k8s_resource,describe_k8s_resource,apply_k8s_manifest,patch_k8s_resource,get_cluster
Golden Path Cost Features
The golden path already includes cost-optimizing settings:
| Setting | Value | Impact |
|---|---|---|
autoscalingProfile | OPTIMIZE_UTILIZATION | Aggressive node |
| : : : scale-down reduces idle : | ||
| : : : compute : | ||
verticalPodAutoscaling | enabled | VPA recommendations |
| : : : prevent : | ||
| : : : over-provisioning : | ||
| Autopilot pricing | Pay per pod request | No charge for unused |
| : : : node capacity : | ||
| Node Auto Provisioning | enabled | Right-sized node pools |
| : : : created automatically : |
Cost Optimization Strategies
1. Spot VMs via ComputeClasses
Use Spot VMs for fault-tolerant workloads (60-90% cost reduction).
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
name: spot-with-fallback
spec:
activeMigration:
optimizeRulePriority: true
priorities:
- machineFamily: n4
spot: true
- machineFamily: n4
spot: false
Spot-suitable workloads:
| Workload | Spot-Suitable? |
|---|---|
| Batch / data processing | Yes |
| Dev / test environments | Yes |
| Stateless web/API (replicas >= 2) | Yes (with PDBs) |
| Jobs with checkpointing | Yes |
| Stateful workloads (databases) | No |
| Single-replica critical services | No |
Handling eviction:
spec:
template:
spec:
terminationGracePeriodSeconds: 25 # Must be < 30s for Spot
containers:
- name: app
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"]
1.1 Direct Workload Spot Selection (NodeSelector)
For stateless or batch workloads in GKE Autopilot, you do not need to create a
custom ComputeClass to use Spot VMs. You can target the built-in Spot capacity
directly by adding a nodeSelector to your Pod or Deployment spec.
[!WARNING] Preemption Warning: Spot VMs are interruptible and can be preempted at any time with a 30-second notice. Workloads must be fault-tolerant and run with at least 2 replicas for high availability. Always explicitly warn users about this preemption risk when recommending Spot VMs.
apiVersion: apps/v1
kind: Deployment
metadata:
name: stateless-spot-app
spec:
replicas: 2
template:
spec:
nodeSelector:
cloud.google.com/gke-provisioning: Spot
terminationGracePeriodSeconds: 25 # Must be < 30s to allow graceful shutdown before preemption
containers:
- name: app
image: <IMAGE>
2. Pod Rightsizing
Use VPA recommendations to reduce over-provisioned requests.
# 1. Deploy VPA in recommendation mode
kubectl apply -f - <<EOF
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: <DEPLOYMENT>-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: <DEPLOYMENT>
updatePolicy:
updateMode: "Off"
EOF
# 2. Wait 24+ hours for data collection
# 3. Read recommendations
kubectl get vpa <DEPLOYMENT>-vpa -o jsonpath='{.status.recommendation}'
Optimization rules:
| Condition | Action | Savings |
|---|---|---|
| CPU request >5x P95 actual | Reduce to P95 * 1.2 | High |
| Memory request >3x P95 actual | Reduce to P95 * 1.2 | High |
| CPU request >2x P95 actual | Reduce to P95 * 1.2 | Medium |
| No resource requests set | Add requests (enables bin-packing) | Medium |
3. Machine Type Selection
| Family | Use Case | Relative Cost |
|---|---|---|
| e2 | General purpose, burstable | Lowest |
| t2a / t2d | Scale-out (Arm/AMD), price-performance | Low |
| : : optimized : : | ||
| n4a | Axion Arm-based, general-purpose | Low |
| : : price-performance : : | ||
| n4 / n4d | General purpose (Intel/AMD), flexible shapes | Low-Medium |
| c4a | Compute-optimized (Arm), high efficiency | Medium-High |
| c3 / c4 | Compute-optimized (Intel) | Medium-High |
| c3d / c4d | Compute-optimized (AMD), high-performance | Medium-High |
| : : throughput : : | ||
| ek-standard | Autopilot enhanced (golden path) | Medium |
| m3 / x4 | Memory-optimized, SAP HANA, large databases | High |
| g2 (L4 GPU) | AI inference | High |
| a3 (H100 GPU) | AI training | Highest |
| a4 / a4x | Ultra-scale AI (Blackwell GPUs) | Highest |
In Autopilot, machine type is managed. Use ComputeClasses to influence selection.
4. Committed Use Discounts (CUDs)
For steady-state workloads, purchase 1-year or 3-year CUDs:
- 1-year: ~20-30% discount
- 3-year: ~50-55% discount
- Applied automatically to matching usage in the region
- Purchase via Google Cloud Console > Billing > Committed use discounts
5. Cluster Management
- Stop/start dev clusters: Idle dev clusters cost money even with no workloads (control plane fee).
- Right-size node pools (Standard): Use Cluster Autoscaler with appropriate min/max.
- Multi-tenant clusters: Share a single cluster across teams instead of
per-team clusters (see the
gke-multitenancyskill).
Cost Monitoring
# View cluster cost breakdown (requires Cost Management API)
gcloud billing budgets list --billing-account=<BILLING_ACCOUNT> --quiet
# View node utilization
kubectl top nodes
# View pod resource usage vs requests
kubectl top pods --all-namespaces --containers
Dev/Test Cost Savings
For non-production environments, these golden path deviations are acceptable:
| Setting | Production (Golden | Dev/Test | : : Path) : : | ----------------------- | ------------------ | ----------------------------- | | Cluster mode | Autopilot | Autopilot (cheaper with fewer | : : : pods) : | Release channel | Regular | Rapid (get fixes faster) | | Private nodes | Required | Optional (simpler access) | | Monitoring components | Full suite | SYSTEM_COMPONENTS only | | Secret Manager rotation | 120s | Disabled | | Maintenance windows | Configured | Not needed |
Related skills
More from google/skills and the wider catalog.

gke-cost-analysis
Analyze GKE cluster and workload costs using BigQuery billing exports and live metrics.

gke-cost-optimization
Optimize GKE costs through rightsizing, Spot VMs, CUDs, and resource quotas.

gke-custom-golden-image-discovery
Discover the correct golden base image for custom GKE node creation based on cluster configuration.

gke-golden-path
GKE production-ready defaults, checklists, and cluster configuration patterns.

gke-inference
Deploy and optimize AI/ML inference workloads on GKE with GPUs, TPUs, and model servers.

gke-manifest-generation
Generate secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and Standard clusters.