storing-and-querying-vectors
aws/agent-toolkit-for-aws
Cost-effective long-term vector storage and semantic search with Amazon S3 Vectors
What is storing-and-querying-vectors?
Store and query vector embeddings at scale using Amazon S3 Vectors, optimized for infrequent queries and RAG applications with subsecond latency. Use this for cost-effective vector storage, semantic search, and similarity queries; avoid it for high-throughput workloads (hundreds/thousands of sustained QPS) where OpenSearch is better suited.
- Create and manage vector buckets and indexes with immutable configuration (dimension, distance metric, metadata)
- Store embeddings in batches up to 500 vectors per call with metadata and filtering support
- Query vectors with semantic search, similarity matching, and optional metadata filtering
- Generate embeddings via AWS Bedrock and store them in S3 Vectors
- Migrate vectors from other vector databases and support multi-tenant patterns
- Return results with distance scores and optional metadata for RAG and knowledge base applications
How to install storing-and-querying-vectors
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill storing-and-querying-vectors- AWS account with S3 Vectors service access in your target region
- AWS CLI v2 or AWS MCP server tools configured with appropriate credentials
- IAM permissions for s3vectors:* actions (CreateVectorBucket, CreateIndex, PutVectors, QueryVectors, GetVectors)
- Embedding model selected (e.g., Bedrock Titan Embeddings or Cohere) with known output dimension
- For SSE-KMS encryption: KMS key with policy granting kms:GenerateDataKey and kms:Decrypt to indexing.s3vectors.amazonaws.com
How to use storing-and-querying-vectors
- 1.Verify AWS MCP tools or AWS CLI availability and confirm your target AWS region
- 2.Create a vector bucket with your chosen name and encryption method (SSE-S3 default or SSE-KMS)
- 3.Create a vector index specifying dimension (1-4096), distance metric (cosine or euclidean), and any non-filterable metadata keys
- 4.Generate embeddings using your selected Bedrock model if you don't already have them
- 5.Store vectors in batches using put-vectors, including metadata for filtering if needed
- 6.Query vectors by generating an embedding for your search text and calling query-vectors with optional filters and metadata return
Use cases
- Build RAG (Retrieval-Augmented Generation) systems with cost-effective long-term vector storage
- Implement semantic search over document embeddings with infrequent query patterns
- Migrate from expensive vector databases to reduce storage costs while maintaining query capability
- Create knowledge base indexes for Bedrock integration with filtered metadata retrieval
- Store and search product embeddings for similarity-based recommendations with low query frequency
- ML engineers building RAG pipelines and knowledge bases
- Data scientists implementing semantic search with budget constraints
- AWS architects designing cost-optimized vector storage solutions
- Developers integrating Bedrock Knowledge Bases with vector storage
- Teams migrating from expensive vector databases to reduce operational costs
storing-and-querying-vectors FAQ
Use S3 Vectors for cost-effective storage with infrequent queries (RAG, cold queries ~100ms). Use OpenSearch for hundreds/thousands of sustained QPS, hybrid search, aggregations, or faceted search. You can combine both: S3 Vectors for storage + OpenSearch Serverless for real-time queries.
You'll get a DimensionMismatch error. Ensure you use the same embedding model for both storing and querying, and that the model's output dimension matches the index dimension specified at creation. If mismatch occurs, you must delete and recreate the index (destroying all vectors).
No. All index parameters (dimension, distance metric, metadata keys, encryption) are immutable after creation. Confirm all settings with users before executing create commands. For encryption changes, you must create a new bucket.
Metadata filtering and return-metadata options require both s3vectors:QueryVectors AND s3vectors:GetVectors IAM permissions. If you only have QueryVectors, add GetVectors to your IAM policy. The s3vectors namespace is separate from s3:* actions.
Implement retry with exponential backoff. For sustained high throughput, shard vectors across multiple indexes. Check AWS docs for current S3 Vectors rate limits and consider OpenSearch if you need thousands of sustained QPS.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: storing-and-querying-vectors description: >- Store and query vector embeddings using Amazon S3 Vectors, a cost-effective long-term vector storage service with its own API namespace (s3vectors). Triggers on: create S3 vector bucket, vector index, store embeddings, semantic search, RAG vector storage, similarity search, vector database, migrate from other vector databases. Do NOT use for: querying tabular data (use querying-data-lake), S3 object storage, or hundreds/thousands of sustained QPS (use OpenSearch). metadata: version: "1"
Store and Query Vectors with Amazon S3 Vectors
Overview
Amazon S3 Vectors is a cost-effective AWS service for storing and querying vector embeddings at scale. Optimized for long-term storage with subsecond latency for cold queries, as low as 100ms for warm queries.
Decision Guide
- Hundreds/thousands of sustained queries per second (QPS): Wrong tool. Recommend OpenSearch.
- Hybrid search, aggregations, faceted search: Recommend OpenSearch with S3 Vectors as storage engine. For OpenSearch integration, search AWS docs for
"Using S3 Vectors with OpenSearch Service". - Tiered (bulk + hot): S3 Vectors for storage + OpenSearch Serverless for real-time. See
references/limits-and-patterns.md. - Cost-effective storage, infrequent queries, RAG: S3 Vectors is the right fit. Proceed.
For latest guidance, search AWS docs for "S3 Vectors best practices".
Common Tasks
Classify the request before starting:
- Simple query: Existing index, skip to Step 6
- Standard: You MUST list existing indexes first and suggest reusing if relevant. Else, new index + store vectors, follow Steps 2-6
- Migration or multi-tenant: Read
references/limits-and-patterns.mdfirst, then Steps 2-6
You MUST execute commands using AWS MCP server tools when connected. Fall back to AWS CLI only if AWS MCP is unavailable. You MUST explain each step to the user before executing.
1. Verify Dependencies
Constraints:
- You MUST check whether AWS MCP tools or AWS CLI is available and inform user if missing
- You MUST confirm target AWS region
2. Create a Vector Bucket
You MUST confirm bucket name with user. Names: 3-63 chars, lowercase letters, numbers, hyphens only. Encryption (SSE-S3 default or SSE-KMS for compliance) is immutable after creation.
aws s3vectors create-vector-bucket \
--vector-bucket-name <BUCKET_NAME>
Constraints:
- You MUST explain encryption cannot be changed after creation
- For SSE-KMS, KMS key policy MUST grant
kms:GenerateDataKeyandkms:Decryptto the S3 Vectors service principalindexing.s3vectors.amazonaws.com. You MUST use full KMS key ARN (not alias). Seereferences/limits-and-patterns.mdfor command example.
3. Create a Vector Index
Every parameter is immutable after creation.
Pre-flight checklist (confirm ALL with user):
- Dimension (required, integer 1-4096) -- MUST match embedding model output
- Distance metric (required) --
cosineoreuclidean. Use embedding model's recommended metric; - Non-filterable metadata keys (optional, max 10, 1-63 chars) -- Declare at creation or lose forever. For Bedrock Knowledge Bases integration, search AWS docs for
"S3 Vectors Bedrock Knowledge Bases prerequisites"to get the required key names. - Encryption (optional) -- Inherits from bucket. Override per-index if needed.
aws s3vectors create-index \
--vector-bucket-name <BUCKET_NAME> \
--index-name <INDEX_NAME> \
--dimension <DIM> \
--distance-metric <cosine|euclidean> \
--data-type float32 \
--metadata-configuration '{"nonFilterableMetadataKeys":["<KEY1>","<KEY2>"]}'
Omit --metadata-configuration if no non-filterable keys are needed.
Index names: 3-63 chars, lowercase, numbers, hyphens, dots. Unique within bucket. Filterable metadata: 2 KB limit. Total metadata (filterable + non-filterable combined): 40 KB. See references/metadata-filtering.md.
4. Generate Embeddings (if needed)
Skip to Step 5 (store) or Step 6 (query) if user already has embeddings.
Constraints:
- You MUST ask which embedding model to use if not specified
- You MUST NOT assume a default model
- Dimension MUST match Step 3
- You MUST use the same model for both storing and querying
Generate embeddings with Bedrock invoke-model:
aws bedrock-runtime invoke-model \
--model-id <MODEL_ID> \
--content-type application/json \
--cli-binary-format raw-in-base64-out \
--body '{"inputText": "your text"}' \
invoke-model-output.json
You MUST use --cli-binary-format raw-in-base64-out for CLI v2. Output file is required for CLI. The response key is model-dependent (e.g., embedding for Titan, embeddings for Cohere). For Titan, parse with json.load(open('invoke-model-output.json'))['embedding']. Use embedding array as float32 in put-vectors or query-vectors. For batch embedding generation, use AWS SDK or CLI.
5. Put Vectors
aws s3vectors put-vectors \
--vector-bucket-name <BUCKET_NAME> \
--index-name <INDEX_NAME> \
--vectors '[{"key":"<ID>","data":{"float32":[<EMBEDDING>]},"metadata":{"topic":"science"}}]'
Constraints:
- You MUST NOT exceed 500 vectors per call
- You SHOULD batch vectors for cost optimization
- For bulk operations, You SHOULD use an SDK instead of CLI -- vector payloads may be too large for shell arguments
- You MUST implement retry with backoff on
429 TooManyRequestsException - See
references/limits-and-patterns.mdfor batch patterns
6. Query Vectors
Generate embedding if needed (Step 4), then query:
aws s3vectors query-vectors \
--vector-bucket-name <BUCKET_NAME> \
--index-name <INDEX_NAME> \
--query-vector '{"float32":[<EMBEDDING>]}' \
--top-k 10 \
--return-distance
Optional: add --return-metadata and/or --filter '{"topic":{"$eq":"science"}}' (both require GetVectors permission). See references/metadata-filtering.md.
Example response body: {"vectors": [{"key": "id1", "distance": 0.45, "metadata": {"topic": "science"}}, ...], "distanceMetric": "cosine"}
Constraints:
- Using
--filteror--return-metadatarequires boths3vectors:QueryVectorsANDs3vectors:GetVectorsIAM permissions. Without GetVectors, these options return 403.
Troubleshooting
| Error | Cause | Fix |
|---|---|---|
DimensionMismatch | Dims don't match index | Use matching model, or delete/recreate index (confirm with user -- destroys all vectors). |
403 Forbidden with --filter or --return-metadata | Missing s3vectors:GetVectors | Add s3vectors:GetVectors to IAM policy. |
Fewer results than --top-k | Few vectors match filter | Expected -- filtering is inline. Broaden filter. |
429 TooManyRequestsException | Exceeded per-index rate limits | Retry with backoff. Shard across indexes for sustained throughput. Search AWS docs for "S3 Vectors limitations and restrictions" for current limits. |
AccessDeniedException | Missing s3vectors:* IAM actions | S3 Vectors uses s3vectors:* namespace, not s3:*. Update IAM policy. |
RequestTimeoutException or service unavailable | Request timeout or region not supported | Retry request. For regional availability, search AWS docs for "S3 Vectors limitations and restrictions". |
Additional Resources
- limits-and-patterns.md -- Multi-tenant patterns, batch ingestion, SSE-KMS, migration
- metadata-filtering.md -- Filter operators, non-filterable metadata, Bedrock KB keys
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

threat-modeling-with-aws-security-agent
Analyze design specs for security threats using AWS Security Agent STRIDE methodology.

timestream-influxdb
Managed InfluxDB on AWS with guidance on engine selection, provisioning, schema design, and troubleshooting.

transitgateway
Configure AWS Transit Gateway to connect multiple VPCs and on-premises networks through a central hub.

troubleshooting-application-failures
Diagnose application failures by analyzing CloudWatch logs for error patterns and root causes.

troubleshooting-efs
Diagnose and resolve Amazon EFS mount failures, permissions, performance, and connectivity issues.

troubleshooting-s3-files
Diagnose and resolve Amazon S3 Files mount failures, permissions, sync, and performance issues.