creating-data-lake-table
aws/agent-toolkit-for-aws
Create managed Iceberg tables on Amazon S3 with automatic compaction and snapshot management.
What is creating-data-lake-table?
Sets up S3 Tables (Iceberg format) with table bucket, namespace, schema, Glue catalog registration, partitioning, and IAM access control. Use this when you need to create a new analytics table in S3 that will be queried via Athena or other Iceberg-compatible engines.
- Create S3 table buckets with encryption and storage class options
- Define Iceberg table schemas with proper type mapping and field IDs
- Set up table namespaces and Glue Data Catalog integration
- Configure partition strategies based on access patterns
- Apply IAM and bucket policies for fine-grained access control
- Verify table creation and queryability via Athena
How to install creating-data-lake-table
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill creating-data-lake-table- AWS CLI or AWS MCP server tools configured with appropriate credentials
- AWS account with S3 Tables API access in target region
- IAM permissions for s3tables:*, glue:*, and s3:* actions
- Confirmation of target AWS region and account identity
How to use creating-data-lake-table
- 1.Verify AWS CLI/MCP availability and confirm credentials with aws sts get-caller-identity
- 2.Check for existing tables in target database using aws glue get-tables
- 3.Gather table requirements: name, columns, types, and partition strategy
- 4.Create table bucket with aws s3tables create-table-bucket
- 5.Create namespace within the bucket using aws s3tables create-namespace
- 6.Set up Glue Data Catalog integration (s3tablescatalog) if not already present
- 7.Define IAM policies for querying principals with s3tables:* and glue:* permissions
- 8.Create the table using aws s3tables create-table with Iceberg schema and partition spec
Use cases
- Setting up a new data lake table for analytics queries in Athena
- Creating partitioned Iceberg tables for time-series data (e.g., daily order events)
- Establishing managed storage for structured data with automatic compaction
- Integrating S3 Tables with Glue ETL pipelines for data processing
- Configuring multi-user access to data lake tables with IAM policies
- Data engineers building data lakes on AWS
- Analytics teams setting up Athena query infrastructure
- AWS architects designing S3-based data warehouses
- DevOps engineers automating table provisioning
creating-data-lake-table FAQ
Use this skill to create the empty table structure first. Then use ingesting-into-data-lake to load data into it. Do not use this skill if you already have data in S3 that you want to query—create the table first, then ingest.
You must check with aws glue get-tables first. If an S3 Tables table exists with matching name, verify schema compatibility and reuse it if compatible. If a non-S3-Tables table exists, delegate to finding-data-lake-assets skill and do not create until user confirms.
No. Glue rejects mixed-case names with GENERIC_INTERNAL_ERROR. Use all lowercase names, and namespace/table names must not contain hyphens.
Querying principals need s3tables:GetTableBucket, s3tables:GetNamespace, s3tables:GetTable, s3tables:GetTableMetadataLocation, s3tables:GetTableData (in bucket policy) and glue:GetCatalog, glue:GetDatabase, glue:GetTable (in IAM policy). See references/access-control.md for exact ARN patterns.
Use Athena DDL (ALTER TABLE) for schema evolution. See references/athena-ddl-path.md for the SQL path. The S3 Tables API is for initial creation only.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: creating-data-lake-table description: >- Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management. Sets up table bucket, namespace, table, schema, Glue catalog registration, partitioning, IAM access control. Triggers on: create table, data lake table, analytics table, structured data storage, S3 Tables, Iceberg, Athena table, partitioning strategy, access permissions. Do NOT use for: importing files (use ingesting-into-data-lake), vector storage (use storing-and-querying-vectors), querying existing tables (use querying-data-lake), or locating existing table (use finding-data-lake-assets). metadata: version: "1" argument-hint: "'[table-description|schema-spec]'"
Create Data Lake Tables with Amazon S3 Tables
Overview
Amazon S3 Tables provides managed Iceberg tables with automatic compaction and snapshot management. Queryable via Athena and Iceberg-compatible engines.
Common Tasks
You MUST use AWS MCP server tools when connected, they provide command validation, sandboxed execution, and audit logging. Fall back to AWS CLI if MCP unavailable.
Decision Guide
Before creating, You MUST check what exists:
You MUST run aws glue get-tables --database-name <NAME> when user mentions a database.
| What you find | Action |
|---|---|
| Fuzzy database name ("our analytics db") | You MUST STOP. Delegate to finding-data-lake-assets to resolve. |
| Non-S3-Tables table with matching name | You MUST STOP. Delegate to finding-data-lake-assets. You MUST NOT create until user confirms. |
| Existing S3 Tables table with matching name | You MUST check schema match. Reuse if compatible, recreate only if user confirms. |
| No matching tables | Proceed with creation (Steps 1-8). |
| User explicitly requests new S3 Tables table | Skip checks, proceed with creation. |
Creation paths:
- Existing data in S3: Create empty table (Steps 1-8), then use
ingesting-into-data-lakeskill. - Glue ETL pipeline: Read
references/table-creation-glue-etl.mdfirst, then Steps 1-6. - Lake Formation access control: Search AWS docs for
"S3 Tables integration with Lake Formation".
1. Verify Dependencies
Constraints:
- You MUST check whether AWS MCP server tools or AWS CLI are available and inform user if missing
- You MUST confirm target AWS region and verify credentials with
aws sts get-caller-identity
2. Understand the Schema
- Explicit schema: Validate Iceberg types.
- Loose description: Ask columns, types, grain. Propose and confirm.
- Existing S3 data: Infer schema from file headers only. Create empty table first, then use
ingesting-into-data-lakeskill.
Constraints:
- You MUST read
references/best-practices.mdfor Iceberg type mapping, partitions, and naming. - You MUST ask for all required parameters upfront: table name, columns, types, partition strategy. For schema evolution, see
references/athena-ddl-path.md. - You MUST use all lowercase names -- Glue rejects mixed case with
GENERIC_INTERNAL_ERROR. Namespace and table names MUST NOT contain hyphens. - You SHOULD suggest partition columns based on access patterns.
3. Create Table Bucket
Names: 3-63 chars, lowercase, numbers, hyphens.
aws s3tables create-table-bucket --name <BUCKET_NAME> --region <REGION>
Capture table-bucket-arn. Encryption (SSE-S3 default, SSE-KMS) and storage class (STANDARD, INTELLIGENT_TIERING) set at creation. See references/best-practices.md.
Constraints:
- You MUST check existing buckets with
aws s3tables list-table-bucketsand ask user to select or create new. - If using SSE-KMS, KMS key policy MUST allow S3 Tables maintenance service principal to read data. Search AWS docs for
"S3 Tables KMS key policy"for required policy. - If bucket creation fails, see
references/best-practices.mdfor common errors.
4. Create Namespace
aws s3tables create-namespace --table-bucket-arn <ARN> --namespace <NAMESPACE>
Constraints:
- You MUST list existing namespaces first and suggest reusing if relevant
- You MUST use lowercase names with no hyphens
5. Create Glue Data Catalog Integration
Check if s3tablescatalog exists (create once per region per account):
aws glue get-catalog --catalog-id s3tablescatalog
If not found, create (requires glue:CreateCatalog, glue:passConnection):
aws glue create-catalog --name "s3tablescatalog" --catalog-input '{
"FederatedCatalog": {
"Identifier": "arn:aws:s3tables:<REGION>:<ACCOUNT_ID>:bucket/*",
"ConnectionName": "aws:s3tables"
},
"CreateDatabaseDefaultPermissions": [{"Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"]}],
"CreateTableDefaultPermissions": [{"Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"]}],
"AllowFullTableExternalDataAccess": "True"
}'
Verify with aws glue get-catalogs --parent-catalog-id s3tablescatalog.
6. Configure Access Control
S3 Tables uses s3tables:* IAM namespace (not s3:*).
Querying principal permissions (bucket policy):
s3tables:GetTableBucket,s3tables:GetNamespace,s3tables:GetTable,s3tables:GetTableMetadataLocation,s3tables:GetTableData
Querying principal permissions (IAM policy):
glue:GetCatalog,glue:GetDatabase,glue:GetTable
You MUST scope to correct ARN patterns. You MUST read references/access-control.md for exact resource ARNs.
Constraints:
- You MUST ask user for querying principal ARN
- You MUST NOT grant broader permissions than necessary
- You MUST NOT create IAM roles automatically, verify existing and guide user
7. Create the Table
| Context | Path |
|---|---|
| Default (any user) | S3 Tables API (below) |
| User specifically wants SQL DDL | Athena DDL (see references/athena-ddl-path.md) |
| Glue ETL pipeline | Spark DDL via --conf job args (not spark.conf.set()). You MUST read references/table-creation-glue-etl.md for the --conf string. |
Default: S3 Tables API:
aws s3tables create-table \
--table-bucket-arn <ARN> \
--namespace <NAMESPACE> \
--name <TABLE_NAME> \
--format ICEBERG \
--metadata '<METADATA_JSON>'
Metadata JSON MUST nest under "iceberg" key:
{"iceberg":{"schema":{"fields":[
{"name":"order_date","type":"date","required":true},
{"name":"customer_id","type":"string","required":true},
{"name":"amount","type":"double","required":false}
]},
"partitionSpec":{"fields":[
{"sourceId":1,"fieldId":1000,"transform":"month","name":"order_date_month"}
]}}}
Constraints:
partitionSpec.sourceIdMUST reference a valid schema field ID- For schema evolution after creation, use Athena DDL. See
references/athena-ddl-path.md - You MUST use
schemaV2for complex types (list, map, struct) with explicit field IDs. Seereferences/best-practices.md. - You SHOULD search AWS docs for
"IcebergPartitionField S3 Tables"for supported partition transforms
8. Verify and Confirm
You MUST verify with aws s3tables get-table and confirm queryability with DESCRIBE <table_name> via Athena using --query-execution-context '{"Catalog":"s3tablescatalog/<BUCKET_NAME>","Database":"<NAMESPACE>"}'. Do NOT put catalog in SQL. Present summary: bucket ARN, namespace, table, schema, partitions.
Troubleshooting
| Error | Cause | Fix |
|---|---|---|
| "Table location can not be specified" | LOCATION in CREATE TABLE | Remove LOCATION clause. S3 Tables manages storage automatically. |
AccessDeniedException with s3:* policy | Using s3:* not s3tables:* | S3 Tables uses s3tables:* namespace. Update IAM policy. |
Additional Resources
- access-control.md -- IAM permissions, ARN patterns, permission errors
- best-practices.md -- Iceberg types, partitions, naming, common errors
- athena-ddl-path.md -- Athena DDL, schema evolution
- table-creation-glue-etl.md -- Spark DDL via Glue ETL
- Loading data:
ingesting-into-data-lakeskill
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

creating-ec2-image-builder-pipeline
Automate custom AMI creation with EC2 Image Builder pipelines, IAM roles, and cross-region distribution.

creating-production-vpc-multi-az
Create production-ready multi-AZ VPCs with public/private subnets, NAT gateways, and security groups.

creating-secrets-using-best-practices
Create and manage AWS Secrets Manager secrets with production-grade security controls and best practices.

debugging-lambda-timeouts
Systematically debug AWS Lambda timeout failures by analyzing configuration, logs, metrics, and dependencies.

deploying-custom-domain-rest-api
Deploy a Regional REST API with custom domain, Lambda backend, and request authorizer on AWS.

developing-applications-on-managed-service-for-apache-flink
Domain expertise for Apache Flink and Amazon Managed Service for Apache Flink development, deployment, and operations.