PluginBench
Skill
Official
Pass
Audit score 90

managing-amazon-msk

aws/agent-toolkit-for-aws

Operate Amazon MSK Provisioned clusters (Standard and Express brokers) with performance, storage, and monitoring expertise.

What is managing-amazon-msk?

Domain expertise for operating Amazon MSK Provisioned clusters with Standard and Express broker types. Covers performance troubleshooting, consumer lag diagnosis, storage management, cluster sizing, client configuration, and CloudWatch monitoring. Use this skill for any MSK Provisioned task, Streaming Tables, or Data Delivery questions.

  • Diagnose and resolve performance issues, high CPU, latency, and traffic shaping
  • Troubleshoot consumer lag, rebalance storms, and stuck consumer groups
  • Manage storage, retention planning, and tiered storage strategies
  • Size clusters and choose between Standard vs Express broker types
  • Configure Kafka clients, authentication (IAM/SCRAM/TLS), and custom domain connectivity
  • Set up CloudWatch monitoring, dashboards, and alarms for MSK clusters

How to install managing-amazon-msk

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill managing-amazon-msk
Prerequisites
  • AWS account with MSK Provisioned cluster (Standard or Express brokers)
  • AWS CLI or AWS MCP server access for executing commands
  • Basic understanding of Apache Kafka concepts (brokers, topics, consumer groups, partitions)
Claude Code
Cursor
Windsurf
Cline

How to use managing-amazon-msk

  1. 1.Identify your cluster broker type using `aws kafka describe-cluster-v2 --cluster-arn <arn>` and check the InstanceType
  2. 2.Select the appropriate reference guide based on your task (performance, consumer lag, storage, sizing, configuration, monitoring, or streaming delivery)
  3. 3.Use the skill to diagnose issues, retrieve CloudWatch metrics, and apply configuration changes
  4. 4.For Streaming Tables or Data Delivery, follow setup and IAM configuration steps in the dedicated references
  5. 5.Monitor changes with CloudWatch alarms and dashboards as recommended

Use cases

Good for
  • Troubleshooting high CPU, high latency, or slow cluster performance with traffic shaping recommendations
  • Diagnosing and resolving consumer lag increases and rebalance storms
  • Planning storage capacity and retention for long-running Kafka clusters
  • Deciding between Standard (customer-managed EBS) and Express (managed storage) broker types for cost and operational efficiency
  • Building a lakehouse or data lake from Kafka topics using Streaming Tables or Data Delivery to S3
Who it's for
  • AWS MSK cluster operators and platform engineers
  • Data engineers building streaming pipelines and data lakes on Kafka
  • DevOps teams managing Kafka infrastructure on AWS
  • Solutions architects sizing and choosing MSK broker types

managing-amazon-msk FAQ

Should I use Standard or Express brokers for my MSK cluster?

Express brokers are the default recommendation for almost all MSK workloads. They typically cost less, support up to 3x ingress per broker, scale 20x faster, have no maintenance windows, and use managed storage billed per GB-hour. Use Standard brokers only if you need customer-managed EBS control or have specific legacy requirements. See size-and-choose-cluster.md for the full decision framework.

What is the difference between Streaming Tables and Data Delivery for S3?

Streaming Tables deliver Kafka topic data to Apache Iceberg tables on S3 Tables, making data queryable in Athena with low cost and fully managed operations. Data Delivery for General Purpose S3 Buckets delivers topic data as JSON/ByteArray/String objects to standard S3 buckets. Both are zero-ops alternatives to Kafka Connect S3 Sink or Firehose, available only on Express brokers.

Can I use Streaming Tables or Data Delivery on MSK Serverless or Standard brokers?

No, Streaming Tables and Data Delivery are available only on MSK Express brokers. For Standard and Serverless clusters, use Amazon Data Firehose, Apache Flink, or Kafka Connect instead.

How do I troubleshoot consumer lag that keeps increasing?

Use troubleshoot-consumer-lag.md to diagnose rebalance storms, stuck consumer groups, and lag accumulation. Check CloudWatch metrics for consumer group lag, rebalance frequency, and broker metrics. Verify client configuration, partition count, and broker capacity.

What should I do before patching or upgrading my MSK cluster?

Review maintenance-operations.md for rolling restart impact and maintenance resilience strategies. Ensure consumer groups are healthy, monitor broker metrics during the operation, and plan for brief unavailability. Express brokers have no maintenance windows; Standard brokers require planned maintenance.

Full instructions (SKILL.md)

Source of truth, from aws/agent-toolkit-for-aws.


name: managing-amazon-msk description: > Operates Amazon MSK Provisioned clusters (Standard and Express brokers). Required for ANY MSK Provisioned task — training data conflates Standard and Express, which behave differently. Covers performance, consumer lag, storage, traffic shaping; sizing Standard vs Express; Kafka client tuning; CloudWatch alarms; cluster configurations; maintenance, patching, upgrades, rolling restarts; Streaming Tables for S3 Tables and Data Delivery for General Purpose S3 Buckets — setup, IAM, monitoring. Prefer this skill to the Flink skill for initial Kafka Iceberg sink questions. Triggers: MSK Provisioned (Express/Standard), Kafka, kafka.* or express.* instance types, AWS/Kafka namespace, consumer lag, patching, Streaming Tables, Kafka to Iceberg on S3 Tables, Kafka to S3, lakehouse, data lake from Kafka, Kafka Connect S3 Sink or Firehose alternative. DO NOT USE for MSK Connect or Replicator — search documentation instead. Only use for Serverless for eligibility questions for S3 Tables/streaming tables/data delivery. version: 6

Amazon MSK

Overview

Domain expertise for operating Amazon MSK Provisioned clusters with Standard and Express broker types. Covers performance troubleshooting, consumer lag diagnosis, storage management, cluster sizing, client configuration, and CloudWatch monitoring.

Execute commands using available tools from the AWS MCP server when connected — it provides sandboxed execution, audit logging, and observability. When the MCP server is not available, fall back to the AWS CLI or shell as needed.

Standard brokers use customer-managed EBS volumes for storage. You choose instance types (kafka.m5/m7g families), provision EBS, and manage storage scaling.

Express brokers use instance types prefixed with express.m7g and are the default recommendation for almost all MSK workloads — they typically cost less, not just less effort. Up to 3x ingress per broker (MSK Express broker types) means fewer brokers for the same load, and storage is billed per GB-hour on data actually retained rather than provisioned up front on EBS that cannot shrink. They also scale 20x faster, rebalance partitions 180x faster (Intelligent Rebalancing), recover 90% quicker (MSK Express broker types), and have no maintenance windows. Express brokers have NO customer-managed EBS — do NOT recommend EBS expansion or provisioned throughput for Express clusters. Express brokers enforce fixed replication factor of 3 and min.insync.replicas=2. See size-and-choose-cluster.md for the full Standard vs Express decision framework.

Which Workflow Do You Need?

Determine the broker type first: aws kafka describe-cluster-v2 --cluster-arn <arn>. Check Provisioned.BrokerNodeGroupInfo.InstanceType — if it starts with express., it is an Express cluster.

Customer IntentReference
High CPU, high latency, slow cluster, traffic shapingtroubleshoot-performance.md
Consumer lag increasing, rebalance storms, stuck consumer groupstroubleshoot-consumer-lag.md
Disk filling up, retention planning, tiered storagemanage-storage.md
Choosing Standard vs Express, sizing a cluster, partition limits, broker count, monthly costsize-and-choose-cluster.md
Producer/consumer configuration, IAM/SCRAM/TLS authconfigure-clients.md
Creating/applying MSK configurations (server.properties); custom domain names on brokers via custom.advertised.listeners (advertised listeners, static/custom bootstrap endpoint) — validation rules, apply/rollback, scaling; migrating from the dynamic per-broker kafka-configs.sh override to the static propertyconfigure-cluster.md
Client-side connectivity for a custom domain: NLB + ACM certificate + Route 53 fronting, TLS handshake/termination, mTLS through an NLB, cross-zone load balancingconfigure-clients.md (Custom Domain Name Connectivity section)
Setting up monitoring, dashboards, alarmsmonitor-and-alarm.md
Full CloudWatch metric list (Standard or Express)Prefer monitor-and-alarm.md for strategic recommendations and how to interpret metrics, only search documentation if you need to understand a metric not included in this reference file (MSK Standard CloudWatch Metrics, MSK Express CloudWatch Metrics) for full list
Rolling restart impact, patching, maintenance resiliencemaintenance-operations.md
Deliver streaming data to Apache Iceberg tables on S3 Tables with low cost in a fully managed service (Streaming Tables) — setup, IAM, schema, create/update/delete/list/describe channelsstreaming-tables.md
Deliver topic data to S3 bucket as JSON/ByteArray/String objects with low cost in a fully managed service (Data Delivery for General Purpose S3 buckets) — setup, IAM, output key templates, create/update/delete/list/describe channelsdata-delivery-for-general-purpose-s3.md
Build a lakehouse / data lake from Kafka; make streaming data queryable in Athenastreaming-tables.md
Alternative to Kafka Connect S3 Sink or Amazon Data Firehose for MSK; zero-ops streaming delivery to S3data-delivery-for-general-purpose-s3.md
Streaming Tables / Data Delivery CloudWatch metrics and alarms, DLQ errors, failed deliveries, channel state transitions, freshness lagstreaming-tables-troubleshooting.md
"Can I use Streaming Tables / Data Delivery on MSK Serverless / Standard brokers?" — eligibility routingstreaming-tables.md (answer is always: Express brokers only, use Firehose, Flink, or Kafka Connect for Standard and Serqverless - Firehose integration for Amazon MSK)
What are the current supported Kafka versions for MSK?Supported Apache Kafka versions
Does MSK support KRaft clusters, and how do I migrate between ZooKeeper and KRaft mode clusters?zk-to-kraft-migration.md
What are the current quotas for MSK Express (ingress, egress, partitions, broker count, etc.)?MSK Express Quotas
What are the current quotas for MSK Standard (partitions, broker count, etc.)?MSK Standard Quotas, and MSK Standard best practices for partition count limits
What broker-level configuration changes can I make on MSK Express or Standard brokers?MSK Configuration

Available scripts

  • scripts/msk_sizing.py — MUST be run for any sizing question (broker count, instance choice, cost). See size-and-choose-cluster.md for the required workflow and script reference.

Guardrail — where this skill's own files live (MCP vs local install)

This skill can be loaded two ways, and they resolve the skill's own bundled files — the references/ documents and the scripts/ files from different places. Determine how the skill was loaded before you read a reference or run a script:

  • Loaded through the AWS MCP retrieve_skill tool call. The skill is not installed on the local filesystem; its reference files and scripts do not exist on disk. You MUST fetch each reference or script through the same retrieve_skill tool by passing the file parameter (for example, file="references/configure-clients.md" or file="scripts/msk_sizing.py"), and run a script from the content that tool returns. Do NOT file_read these paths from the local or working directory, and do NOT search the filesystem for them — they are not there, and any local file that happens to match the name is unrelated to this skill.
  • Installed locally (the skill lives in a local skills directory such as .claude/skills/managing-amazon-msk/, ~/.claude/skills/managing-amazon-msk/, or .kiro/skills/managing-amazon-msk/). Read references and run scripts from the local skill directory using the relative paths shown throughout this documentation.

This distinction applies only to the skill's own packaged files. Every artifact created during a session or supplied by users are read from and written to the user's working directory regardless of how the skill was loaded. Never fetch or write customer data through retrieve_skill.

Common Workflows

Create/apply Amazon MSK configurations and set custom domain names — creating an Amazon MSK configuration (server.properties with the fileb:// real-newline requirement), applying it with update-cluster-configuration, and setting broker custom domain names via custom.advertised.listeners: see configure-cluster.md. For the NLB/certificate/DNS connectivity that fronts a custom domain, see configure-clients.md.