PluginBench
Skill
Official
Pass
Audit score 90

connecting-to-data-source

aws/agent-toolkit-for-aws

Create and test AWS Glue connections to JDBC databases, Redshift, Snowflake, and BigQuery.

What is connecting-to-data-source?

This skill registers external data sources with AWS Glue by creating reusable, tested connections that store network config, credentials, and driver settings. Use it when you need to connect Glue jobs to Oracle, SQL Server, PostgreSQL, MySQL, RDS, Redshift, Snowflake, or BigQuery—but not for moving data, creating tables, or querying.

  • Discovers existing Glue connections and candidate RDS/Redshift sources in your AWS account
  • Gathers connection hints (hostname, port, database, auth method) from the user
  • Registers credentials securely in AWS Secrets Manager or configures IAM database authentication
  • Configures VPC, subnet, security group, and availability zone for private sources
  • Tests connections in two phases: Glue API validation and engine-level verification (ETL job, Athena, or Crawler)
  • Troubleshoots network, credential, and driver issues before handing off to downstream skills

How to install connecting-to-data-source

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill connecting-to-data-source
Prerequisites
  • AWS account with Glue, Secrets Manager, and RDS/Redshift/VPC access
  • AWS credentials configured and verified (aws sts get-caller-identity)
  • Target AWS region identified
  • For private sources: VPC, subnet, security group, and S3 VPC gateway endpoint already set up
Claude Code
Cursor
Windsurf
Cline

How to use connecting-to-data-source

  1. 1.Provide the source type (JDBC database, Redshift, Snowflake, or BigQuery) or let the skill infer it from your description
  2. 2.Answer questions about connection name, hostname/endpoint, port, database, and authentication method
  3. 3.Review candidate sources discovered in your account and select one, or provide custom details
  4. 4.Confirm credential storage method (Secrets Manager or IAM database auth)
  5. 5.For private sources, provide VPC, subnet, security group, and availability zone details
  6. 6.Run the Glue connection creation command and wait for success
  7. 7.Execute Phase A (Glue TestConnection API check) and Phase B (engine-level verification via ETL job, Athena, or Crawler)
  8. 8.If tests pass, use the connection name in the ingesting-into-data-lake skill; if they fail, work through troubleshooting steps

Use cases

Good for
  • Set up a Glue connection to an on-premises Oracle database so ETL jobs can ingest data from it
  • Connect a Redshift cluster to Glue and test the connection before running a data pipeline
  • Register Snowflake credentials and create a SNOWFLAKE-type Glue connection for Spark ETL jobs
  • Configure BigQuery access from Glue with a service account and verify the connection works
  • Troubleshoot a failing Glue connection by diagnosing VPC routing, security groups, and credential access
Who it's for
  • Data engineers setting up AWS Glue ETL pipelines
  • Cloud architects configuring data lake ingestion from external sources
  • DevOps engineers managing Glue job infrastructure and connections
  • Analytics teams connecting data warehouses (Redshift, Snowflake, BigQuery) to AWS

connecting-to-data-source FAQ

What's the difference between a Glue connection and a data pipeline?

A connection is a named pipe—it stores network config, driver, and credentials for one source. It does not move data. Pipelines (ETL jobs, crawlers) use connections to access sources. Create a connection once per source, reuse it across many jobs.

Should I use Secrets Manager or IAM database authentication?

Prefer IAM database authentication for Aurora and RDS (MySQL, PostgreSQL) and Redshift—no secret to rotate, tokens auto-refresh. For other sources or legacy systems, use Secrets Manager. Never store plaintext passwords.

Why does my connection test pass but my Glue job fails?

Glue TestConnection (Phase A) validates network and credential sanity but does not prove end-to-end compatibility with your job engine. Phase B runs a minimal query (smoke-test ETL job, Athena SELECT, or test crawl) to catch driver, catalog, and Spark-level issues.

What if my source is in a private VPC?

Provide the subnet ID, security group IDs, and availability zone. Glue will launch job workers in that VPC. Ensure VPC routing, security group rules, and an S3 VPC gateway endpoint are in place so jobs can reach the source and write results to S3.

Can I use this skill to move data or query a source?

No. This skill only creates and tests connections. Use ingesting-into-data-lake to move data, querying-data-lake to run queries, and creating-data-lake-table to define tables.

Full instructions (SKILL.md)

Source of truth, from aws/agent-toolkit-for-aws.


name: connecting-to-data-source description: >- Create and troubleshoot AWS Glue connections to JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS), Redshift, Snowflake, and BigQuery. Gathers connection hints from user, discovers existing connections and RDS/Redshift candidates, registers credentials in Secrets Manager or IAM DB auth, configures VPC, and tests. Triggers on: connect to database, set up Glue connection, register data source, connect to Snowflake/BigQuery/RDS, connection timeout, test connection, troubleshoot connection. Do NOT use for moving data (use ingesting-into-data-lake), creating tables (use creating-data-lake-table), queries (use querying-data-lake), catalog exploration (use exploring-data-catalog), or SaaS (Salesforce, ServiceNow, SAP, MongoDB, Kafka). metadata: version: "1" argument-hint: "'[source-type|connection-name|hostname]'"

Connect to Data Source

Register an external data source with AWS Glue so downstream skills (ingesting-into-data-lake) can move data from it. A Glue connection stores the network config, driver, and credential reference for one source. Create once per source, reuse across jobs.

Philosophy

A connection is a named pipe, not a pipeline. This skill produces a tested, reusable Glue connection. It does not move data.

Common Tasks

You MUST execute commands using AWS MCP server tools when connected -- they provide validation, sandboxed execution, and audit logging. Fall back to AWS CLI only if MCP is unavailable. You MUST explain each step before executing.

Workflow

1. Verify Dependencies and Context

  • You MUST check whether AWS MCP tools or AWS CLI are available and inform the user if missing
  • You MUST confirm target AWS region and verify credentials with aws sts get-caller-identity

2. Classify the Source

Ask the user which source type they want to connect to, or infer from hints:

User says...Source typeConnection typeReference
"Oracle", "SQL Server", "Postgres", "MySQL", "RDS <engine>"JDBC databaseJDBCjdbc-setup.md
"Redshift", "my cluster", "my data warehouse on AWS"RedshiftJDBCjdbc-setup.md (Redshift section)
"Snowflake"SnowflakeSNOWFLAKEsnowflake-setup.md
"BigQuery", "Google analytics warehouse"BigQueryBIGQUERYbigquery-setup.md

If the user names DynamoDB or a local file, stop and tell them: DynamoDB is read directly by Glue without a connection, and local files belong in the ingesting-into-data-lake skill's local-upload workflow.

3. Gather Connection Hints from the User

You MUST ask for hints the user can provide -- do not guess.

For all sources:

  • Desired connection name (lowercase, hyphens: oracle-prod-sales, snowflake-analytics)
  • Existing Secrets Manager secret, or create one
  • Is source reachable from a Glue VPC (same, peered, VPN, Direct Connect)

JDBC: hostname/endpoint, port, database, whether RDS/Aurora/self-managed, IAM DB auth enabled (Aurora/RDS MySQL/Postgres), SSL required.

Snowflake: account identifier, warehouse, role, default database, auth (password, key-pair, OAuth).

BigQuery: GCP project ID, location, whether service account JSON is provisioned.

4. Discover Existing Connections and Candidate Sources

Check what exists before creating.

Existing Glue connections:

aws glue get-connections --filter ConnectionType=<TYPE> --region <REGION>

If a suitable one exists, confirm and skip to Step 7.

Candidate sources in account (JDBC/Redshift only):

  • RDS: aws rds describe-db-instances
  • Aurora: aws rds describe-db-clusters
  • Redshift: aws redshift describe-clusters

Present candidates to user; let them pick. See discovery.md.

5. Register Credentials

You MUST encourage AWS Secrets Manager over plaintext passwords. You SHOULD prefer IAM database authentication where supported (Aurora/RDS MySQL and PostgreSQL, Redshift). See credential-security.md.

  • You MUST confirm with user before creating a new Secrets Manager secret
  • You MUST NOT write plaintext credentials into chat or logs
  • For IAM DB auth, no secret is needed

6. Create the Glue Connection

Follow the source-specific reference for connection properties:

aws glue create-connection --connection-input '<JSON>' --region <REGION>

Private sources require PhysicalConnectionRequirements (SubnetId, SecurityGroupIdList, AvailabilityZone). See network-setup.md.

7. Test the Connection

You MUST test before handing off. Testing is two-phase: a quick API check, then an engine-level verification.

Phase A: Glue TestConnection (network and credential sanity check)

aws glue test-connection --connection-name <NAME> --region <REGION>

This validates that Glue can reach the source and authenticate. It does NOT prove the connection works end-to-end with the query engine the user plans to use.

Phase B: Engine-level verification

After TestConnection passes, verify the connection works with the user's intended engine by running a minimal query through it:

  • Glue ETL (default): Run a smoke-test Glue job that reads one row via the connection. See troubleshooting.md.
  • Athena: If the user plans to query via Athena with a federated connector, run a SELECT 1 through the Athena connection to confirm the Lambda-based connector can reach the source.
  • Glue Crawler: If the user plans to crawl the source, run a test crawl on a single table.

Phase B catches issues that TestConnection misses: driver compatibility at job runtime, catalog configuration, Spark-level serialization, and engine-specific auth flows (e.g., Snowflake SNOWFLAKE type works in ETL but not via JDBC crawlers).

On success in both phases, tell user the connection name is ready for ingesting-into-data-lake. On failure in either phase, Step 8.

8. Troubleshoot (only if test failed)

Diagnose in order: network, credentials, driver. See troubleshooting.md.

Constraints:

  • You MUST check VPC routing, security groups, and S3 VPC endpoint before blaming credentials
  • You MUST verify Glue role can read the Secrets Manager secret
  • You MUST NOT rotate credentials without user confirmation

Argument Routing

  • No args: Walk through Steps 1-7 interactively
  • Source type keyword (e.g., snowflake, oracle): Skip to Step 2 with the type prefilled
  • Existing connection name: Skip to Step 7 (test) then Step 8 if failing
  • Hostname or RDS endpoint: Skip to Step 4 with the candidate prefilled

Gotchas

  • Glue's SNOWFLAKE connection type is distinct from JDBC configured for Snowflake. You MUST use SNOWFLAKE for Spark ETL jobs; do not use JDBC.
  • Connection names are immutable. Choose carefully.
  • PhysicalConnectionRequirements.AvailabilityZone MUST match the subnet's AZ or the connection fails at job runtime, not creation time.
  • IAM database authentication tokens expire in 15 minutes. The Glue job generates a fresh token on each connection; do not cache.
  • An S3 VPC gateway endpoint MUST exist in the VPC used by private-source connections. Without it, Glue jobs cannot read their scripts or write results to S3.

Troubleshooting

ErrorLikely causeFix
Connect timed outVPC routing, SG rule, or NAT gateway missingSee troubleshooting.md
Access denied for user / ORA-01017Credentials wrong, Secrets Manager access missing, or IAM DB auth misconfiguredSee troubleshooting.md
No suitable driver foundCustom driver JAR not set or wrong class nameSee troubleshooting.md
SSL handshake failedJDBC_ENFORCE_SSL mismatch between Glue and sourceSee troubleshooting.md
UnableToFindVpcEndpointS3 VPC endpoint missingCreate S3 gateway endpoint in the connection's VPC

References