PluginBench
Skill
Official
Review
Audit score 70

ingesting-into-data-lake

aws/agent-toolkit-for-aws

Ingest data from S3, databases, Snowflake, BigQuery, DynamoDB, or Glue tables into your AWS data lake.

What is ingesting-into-data-lake?

Moves data from external sources (files, JDBC databases, cloud data warehouses, DynamoDB, or existing Glue tables) into queryable tables in your AWS data lake. Use this to load one-time datasets, set up recurring pipelines, or migrate tables to S3 Tables or Iceberg format.

  • Ingest from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Redshift, Snowflake, BigQuery, and DynamoDB
  • Write to S3 Tables (default) or standard Iceberg on general-purpose buckets
  • Handle one-time loads, recurring pipelines, and table migrations with validation
  • Automatically detect target format based on your existing catalog posture
  • Validate data quality with row counts, null checks, and sample row inspection

How to install ingesting-into-data-lake

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill ingesting-into-data-lake
Prerequisites
  • AWS credentials configured and verified with appropriate IAM permissions (s3:GetObject, s3:PutObject for S3; Glue job execution role)
  • AWS MCP tools or AWS CLI available
  • For JDBC, Snowflake, or BigQuery sources: existing Glue connection to the source system
  • Target database and table naming strategy decided (or existing target table identified)
Claude Code
Cursor
Windsurf
Cline

How to use ingesting-into-data-lake

  1. 1.Identify your data source type (S3 file, local upload, JDBC database, Snowflake, BigQuery, DynamoDB, or Glue table)
  2. 2.Verify the source connection exists (if applicable) using aws glue get-connection
  3. 3.Confirm your target: database name, table name, and format (S3 Tables, Iceberg, or Parquet)
  4. 4.Execute the source-specific workflow (local upload, S3 ingest, JDBC, Snowflake, BigQuery, DynamoDB, or catalog migration)
  5. 5.Validate the ingested data: compare row counts, check for nulls in critical columns, and spot-check sample rows
  6. 6.For recurring loads, create a Glue Trigger with a cron schedule

Use cases

Good for
  • Load a CSV file from your local machine into a new S3 Tables table
  • Migrate an existing Hive table in Glue Catalog to Iceberg format
  • Set up a recurring pipeline to pull data from a PostgreSQL database into your data lake daily
  • Export a DynamoDB table to S3 and query it with Athena
  • Ingest a Snowflake table into your AWS data lake for cross-cloud analytics
Who it's for
  • Data engineers building or maintaining AWS data lakes
  • Analytics teams migrating data from external sources to AWS
  • Organizations adopting S3 Tables or Iceberg for modern data formats
  • Teams managing recurring ETL pipelines

ingesting-into-data-lake FAQ

What's the difference between S3 Tables and Iceberg?

S3 Tables is AWS's recommended format for new data lake work and provides managed table metadata. Standard Iceberg is recommended if your account already uses Iceberg on general-purpose buckets. The skill defaults to S3 Tables unless your catalog inventory suggests otherwise.

Do I need a Glue connection for all sources?

No. Local files, S3 files, DynamoDB, and catalog migration do not require a Glue connection. JDBC databases, Snowflake, and BigQuery require an existing Glue connection; if it doesn't exist, delegate to the connecting-to-data-source skill.

Can I ingest from Salesforce, ServiceNow, or Kafka?

No. This skill supports S3, local files, JDBC databases, Redshift, Snowflake, BigQuery, DynamoDB, and Glue catalog tables only. SaaS platforms and streaming sources are not supported in this release.

What happens if my target table doesn't exist?

The skill will delegate to creating-data-lake-table to create it first. You'll specify the database, table name, and format (S3 Tables or Iceberg) before proceeding with the ingest.

How do I schedule recurring ingests?

After a successful one-time ingest, create a Glue Trigger with a cron schedule pointing to your Glue job. For complex multi-step pipelines with branching logic, use MWAA instead.

Full instructions (SKILL.md)

Source of truth, from aws/agent-toolkit-for-aws.


name: ingesting-into-data-lake description: >- Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg. Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog), or SaaS platforms like Salesforce, ServiceNow, SAP, MongoDB, Kafka. metadata: version: "1" argument-hint: "'[source-path|connection-name|table-name] [--target s3-tables|iceberg|parquet]'"

Ingest into Data Lake

Move data from a source into a queryable table in the data lake. This skill assumes the source connection (if one is needed) already exists. For Glue connection setup or troubleshooting, delegate to connecting-to-data-source.

Philosophy

Default to S3 Tables unless the environment says otherwise. S3 Tables is the recommended target for new data lake work. If the user's catalog inventory shows they haven't adopted S3 Tables, recommend standard Iceberg on their existing general-purpose bucket instead of forcing them to change posture.

Common Tasks

You MUST execute commands using AWS MCP server tools when connected -- they provide validation, sandboxed execution, and audit logging. Fall back to AWS CLI only if MCP is unavailable. You MUST explain each step before executing.

Workflow

1. Verify Dependencies and Context

  • You MUST check whether AWS MCP tools or AWS CLI are available and inform the user if missing
  • You MUST confirm target AWS region and verify credentials with aws sts get-caller-identity
  • For SageMaker Unified Studio project roles, note that target tables and connections may be scoped to the project. See the caller ARN detection pattern in querying-data-lake.

2. Classify the Source

User says...Source typeReference
"upload my file", "local CSV", "move to S3"Local filelocal-upload.md
"load from S3", "import CSV/JSON/Parquet from s3://"S3 filess3-files.md
"import from Oracle/Postgres/MySQL/SQL Server/Redshift/RDS/Aurora"JDBCjdbc-ingest.md
"pull from Snowflake", "Snowflake table to S3"Snowflakesnowflake-ingest.md
"import from BigQuery", "GCP analytics to S3"BigQuerybigquery-ingest.md
"export DynamoDB", "DynamoDB to data lake"DynamoDBdynamodb-ingest.md
"migrate Glue table", "convert Hive to Iceberg"Catalog migrationcatalog-migration.md

If the user names Salesforce, ServiceNow, SAP, MongoDB, Kafka, or another SaaS/streaming source, decline -- these are not supported in this release.

If the source table is referenced by a fuzzy or business name ("migrate our orders table", "pull from the sales warehouse"), delegate to finding-data-lake-assets to resolve before proceeding.

3. Confirm Connection Exists (if applicable)

For JDBC, Snowflake, and BigQuery sources, a Glue connection is required. Check:

aws glue get-connection --name <CONNECTION_NAME> --region <REGION>

If the connection does not exist, stop and delegate to connecting-to-data-source to create and test it. Do not proceed with ingest until the connection is verified.

Local files, S3 files, DynamoDB, and catalog migration do not need a Glue connection.

4. Clarify the Target

You MUST ask the user (or suggest based on catalog inventory) before creating or writing to any table:

  • Database/namespace: Does a specific target database exist? Or should one be created?
  • Table: Existing table (append/merge) or new table (delegate to creating-data-lake-table)?
  • Format: S3 Tables (default), standard Iceberg, or raw Parquet?

Inventory-aware defaults:

If you have already run exploring-data-catalog or can quickly check, use what exists:

  • Account has an s3tablescatalog federated catalog and active table buckets: recommend S3 Tables
  • Account has general-purpose buckets with Iceberg tables and no S3 Tables usage: recommend standard Iceberg on their existing bucket
  • Account uses Parquet/ORC on S3 without Iceberg metadata: ask whether to adopt Iceberg now (recommend yes) or continue with raw files

Do not force S3 Tables on customers who haven't adopted it. See iceberg-catalog-config-and-usage.md.

Delegations from this step:

  • Target table doesn't exist -> creating-data-lake-table
  • Target database named by fuzzy term -> finding-data-lake-assets
  • User doesn't know what exists -> exploring-data-catalog

5. Execute Source Workflow

Read the source-specific reference and follow its phases. Each is self-contained with job templates, gotchas, and troubleshooting:

  • Local / S3 / JDBC / Snowflake / BigQuery / DynamoDB / catalog migration -- one reference per source

Common Glue 5.1 or higher job configuration and PySpark templates are shared in glue-job-config.md and glue-job-scripts.md.

6. Validate

Run all three, do not skip:

  1. Row count matches expected (source vs target)
  2. Null check on critical columns
  3. Spot-check 3-5 sample rows

See data-quality-validation.md.

7. Schedule (if recurring)

For recurring pipelines, create a Glue Trigger with a cron schedule. See testing-and-scheduling.md. Simple single-step pipelines use Glue Triggers; multi-step with branching uses MWAA.

Argument Routing

  • S3 path only: Infer one-time load, start Step 2 with S3 files
  • Connection name: Start Step 3 with the named connection
  • Table name: Start Step 4, ask whether this is source or target
  • --target flag: Pre-fill the target format in Step 4
  • No args: Walk through interactively

Gotchas

  • S3 Tables requires Glue 5.1 or higher and --datalake-formats iceberg job argument
  • All spark.sql.catalog.* config MUST go in --conf job arguments, never in spark.conf.set(). Glue 5.x throws AnalysisException: Cannot modify the value of a static config otherwise. See iceberg-catalog-config-and-usage.md for correct catalog configs.
  • The warehouse parameter is required in S3 Tables catalog config. Without it Spark fails with "Cannot derive default warehouse location".
  • Table and column names in S3 Tables MUST be all lowercase
  • overwritePartitions() only replaces partitions present in the DataFrame -- for full refresh with deletes, use createOrReplace()
  • Standard Iceberg targets MUST include a LOCATION clause; S3 Tables MUST NOT
  • DynamoDB does not need a Glue connection -- do not attempt to create one
  • Connection failures during ingest delegate back to connecting-to-data-source; do not debug network/credentials in this skill
  • For target tables in SageMaker Unified Studio projects, ensure the project role has write access to the target namespace before the Glue job runs

Troubleshooting

ErrorLikely causeAction
Access Denied on S3Missing IAM permissionsCheck Glue role has s3:GetObject, s3:PutObject
Access Denied on S3 TablesMissing s3tables:* permissionsAdd S3 Tables inline policy to Glue role
CTAS timeoutDataset too large for AthenaSwitch to Glue ETL or batch with WHERE filters
JDBC connection timeout/auth failureConnection-level issueDelegate to connecting-to-data-source
Throughput exceeded (DynamoDB)Read percent too highLower read.percent or use native export

See error-handling.md for the full catalog.

References

Source-specific

Cross-cutting

Migration-specific

JDBC-specific