PluginBench
Skill
Pass
Audit score 90

embedding-strategies

wshobson/agents

Select and optimize embedding models for semantic search and RAG applications.

What is embedding-strategies?

Guide to choosing embedding models, implementing chunking strategies, and optimizing embedding quality for vector search and RAG systems. Use when selecting models for your domain, tuning chunk sizes, or comparing embedding performance.

  • Compare embedding models across dimensions, token limits, and use cases (code, finance, legal, multilingual)
  • Implement chunking strategies that preserve semantic boundaries and handle overlap
  • Normalize embeddings for cosine similarity search and optimize for specific domains
  • Batch embedding requests and cache results to reduce API costs and latency
  • Match embedding models to application needs (Claude apps, OpenAI, local deployment)
  • Handle multilingual content with appropriate embedding models

How to install embedding-strategies

npx skills add https://github.com/wshobson/agents --skill embedding-strategies
Claude Code
Cursor
Windsurf
Cline

How to use embedding-strategies

  1. 1.Review the embedding model comparison table to identify candidates for your use case (domain type, token limits, cost)
  2. 2.Select a model: use Voyage AI for Claude apps, text-embedding-3 for OpenAI apps, or open-source models for local deployment
  3. 3.Design your chunking strategy considering document type, semantic boundaries, and token limits
  4. 4.Implement preprocessing (cleaning, normalization) before embedding
  5. 5.Batch your embedding requests rather than processing one-by-one
  6. 6.Cache embeddings for static content to avoid recomputation
  7. 7.Test and compare model performance on your specific domain before full deployment

Use cases

Good for
  • Selecting between Voyage AI and OpenAI embedding models for a RAG pipeline
  • Optimizing chunk size and overlap for legal document retrieval
  • Implementing code search using voyage-code-3 embeddings
  • Reducing embedding dimensions while maintaining search quality
  • Building multilingual semantic search across documents in multiple languages
Who it's for
  • RAG application developers
  • Vector database engineers
  • ML engineers optimizing search systems
  • Teams building semantic search features
  • Developers working with Claude or OpenAI APIs

embedding-strategies FAQ

Which embedding model should I use for Claude applications?

Voyage AI models (voyage-3-large, voyage-3, or domain-specific variants like voyage-code-3) are recommended by Anthropic for Claude apps. They offer 1024 dimensions and up to 32000 token context.

What's the difference between chunking size and token limits?

Chunking size is how you split documents before embedding (e.g., 512 characters). Token limits are the maximum input the embedding model accepts (e.g., 8191 tokens for text-embedding-3). Ensure your chunks don't exceed the model's token limit.

Can I mix different embedding models in the same vector database?

No, different embedding models produce incompatible vector spaces. Choose one model and stick with it for all documents in a search system.

How do I optimize embedding quality for a specific domain?

Use domain-specific models when available (voyage-finance-2 for finance, voyage-law-2 for legal). Otherwise, match general-purpose models to your use case and test performance on domain-specific benchmarks.

Should I normalize embeddings?

Yes, normalize embeddings when using cosine similarity search. This ensures consistent distance metrics and improves search quality.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: embedding-strategies description: Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

Embedding Strategies

Guide to selecting and optimizing embedding models for vector search applications.

When to Use This Skill

  • Choosing embedding models for RAG
  • Optimizing chunking strategies
  • Fine-tuning embeddings for domains
  • Comparing embedding model performance
  • Reducing embedding dimensions
  • Handling multilingual content

Core Concepts

1. Embedding Model Comparison (2026)

ModelDimensionsMax TokensBest For
voyage-3-large102432000Claude apps (Anthropic recommended)
voyage-3102432000Claude apps, cost-effective
voyage-code-3102432000Code search
voyage-finance-2102432000Financial documents
voyage-law-2102432000Legal documents
text-embedding-3-large30728191OpenAI apps, high accuracy
text-embedding-3-small15368191OpenAI apps, cost-effective
bge-large-en-v1.51024512Open source, local deployment
all-MiniLM-L6-v2384256Fast, lightweight
multilingual-e5-large1024512Multi-language

2. Embedding Pipeline

Document → Chunking → Preprocessing → Embedding Model → Vector
                ↓
        [Overlap, Size]  [Clean, Normalize]  [API/Local]

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Match model to use case: Code vs prose vs multilingual
  • Chunk thoughtfully: Preserve semantic boundaries
  • Normalize embeddings: For cosine similarity search
  • Batch requests: More efficient than one-by-one
  • Cache embeddings: Avoid recomputing for static content
  • Use Voyage AI for Claude apps: Recommended by Anthropic

Don'ts

  • Don't ignore token limits: Truncation loses information
  • Don't mix embedding models: Incompatible vector spaces
  • Don't skip preprocessing: Garbage in, garbage out
  • Don't over-chunk: Lose important context
  • Don't forget metadata: Essential for filtering and debugging