embedding-strategies
wshobson/agents
Select and optimize embedding models for semantic search and RAG applications.
What is embedding-strategies?
Guide to choosing embedding models, implementing chunking strategies, and optimizing embedding quality for vector search and RAG systems. Use when selecting models for your domain, tuning chunk sizes, or comparing embedding performance.
- Compare embedding models across dimensions, token limits, and use cases (code, finance, legal, multilingual)
- Implement chunking strategies that preserve semantic boundaries and handle overlap
- Normalize embeddings for cosine similarity search and optimize for specific domains
- Batch embedding requests and cache results to reduce API costs and latency
- Match embedding models to application needs (Claude apps, OpenAI, local deployment)
- Handle multilingual content with appropriate embedding models
How to install embedding-strategies
npx skills add https://github.com/wshobson/agents --skill embedding-strategiesHow to use embedding-strategies
- 1.Review the embedding model comparison table to identify candidates for your use case (domain type, token limits, cost)
- 2.Select a model: use Voyage AI for Claude apps, text-embedding-3 for OpenAI apps, or open-source models for local deployment
- 3.Design your chunking strategy considering document type, semantic boundaries, and token limits
- 4.Implement preprocessing (cleaning, normalization) before embedding
- 5.Batch your embedding requests rather than processing one-by-one
- 6.Cache embeddings for static content to avoid recomputation
- 7.Test and compare model performance on your specific domain before full deployment
Use cases
- Selecting between Voyage AI and OpenAI embedding models for a RAG pipeline
- Optimizing chunk size and overlap for legal document retrieval
- Implementing code search using voyage-code-3 embeddings
- Reducing embedding dimensions while maintaining search quality
- Building multilingual semantic search across documents in multiple languages
- RAG application developers
- Vector database engineers
- ML engineers optimizing search systems
- Teams building semantic search features
- Developers working with Claude or OpenAI APIs
embedding-strategies FAQ
Voyage AI models (voyage-3-large, voyage-3, or domain-specific variants like voyage-code-3) are recommended by Anthropic for Claude apps. They offer 1024 dimensions and up to 32000 token context.
Chunking size is how you split documents before embedding (e.g., 512 characters). Token limits are the maximum input the embedding model accepts (e.g., 8191 tokens for text-embedding-3). Ensure your chunks don't exceed the model's token limit.
No, different embedding models produce incompatible vector spaces. Choose one model and stick with it for all documents in a search system.
Use domain-specific models when available (voyage-finance-2 for finance, voyage-law-2 for legal). Otherwise, match general-purpose models to your use case and test performance on domain-specific benchmarks.
Yes, normalize embeddings when using cosine similarity search. This ensures consistent distance metrics and improves search quality.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: embedding-strategies description: Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.
Embedding Strategies
Guide to selecting and optimizing embedding models for vector search applications.
When to Use This Skill
- Choosing embedding models for RAG
- Optimizing chunking strategies
- Fine-tuning embeddings for domains
- Comparing embedding model performance
- Reducing embedding dimensions
- Handling multilingual content
Core Concepts
1. Embedding Model Comparison (2026)
| Model | Dimensions | Max Tokens | Best For |
|---|---|---|---|
| voyage-3-large | 1024 | 32000 | Claude apps (Anthropic recommended) |
| voyage-3 | 1024 | 32000 | Claude apps, cost-effective |
| voyage-code-3 | 1024 | 32000 | Code search |
| voyage-finance-2 | 1024 | 32000 | Financial documents |
| voyage-law-2 | 1024 | 32000 | Legal documents |
| text-embedding-3-large | 3072 | 8191 | OpenAI apps, high accuracy |
| text-embedding-3-small | 1536 | 8191 | OpenAI apps, cost-effective |
| bge-large-en-v1.5 | 1024 | 512 | Open source, local deployment |
| all-MiniLM-L6-v2 | 384 | 256 | Fast, lightweight |
| multilingual-e5-large | 1024 | 512 | Multi-language |
2. Embedding Pipeline
Document → Chunking → Preprocessing → Embedding Model → Vector
↓
[Overlap, Size] [Clean, Normalize] [API/Local]
Templates and detailed worked examples
Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.
Best Practices
Do's
- Match model to use case: Code vs prose vs multilingual
- Chunk thoughtfully: Preserve semantic boundaries
- Normalize embeddings: For cosine similarity search
- Batch requests: More efficient than one-by-one
- Cache embeddings: Avoid recomputing for static content
- Use Voyage AI for Claude apps: Recommended by Anthropic
Don'ts
- Don't ignore token limits: Truncation loses information
- Don't mix embedding models: Incompatible vector spaces
- Don't skip preprocessing: Garbage in, garbage out
- Don't over-chunk: Lose important context
- Don't forget metadata: Essential for filtering and debugging
Related skills
More from wshobson/agents and the wider catalog.

employment-contract-templates
Create legally sound employment contracts, offer letters, and HR policy documents with templates and best practices.

error-handling-patterns
Master error handling patterns across languages to build resilient, fault-tolerant applications.

eval-harness-first
Build the evaluation harness that gates fine-tuning — goldens, graders, judge calibration, and baselines.

evaluation-methodology
Understand PluginEval's quality scoring methodology, dimensions, rubrics, and anti-patterns.

event-store-design
Design and implement event stores for event-sourced systems with architecture patterns and technology guidance.

fastapi-templates
Production-ready FastAPI project templates with async patterns, dependency injection, and error handling.