similarity-search-patterns
wshobson/agents
Implement efficient similarity search with vector databases for semantic retrieval and nearest neighbor queries.
What is similarity-search-patterns?
This skill provides patterns for building production similarity search systems using vector databases. Use it when implementing semantic search, RAG retrieval, recommendation engines, or optimizing search performance across millions of vectors.
- Compare distance metrics (cosine, Euclidean, dot product, Manhattan) for different embedding types
- Select appropriate index types (Flat for exact search, HNSW for medium-large data, IVF+PQ for very large scale)
- Tune index parameters like ef_search and nprobe to balance recall and latency
- Implement hybrid search combining semantic and keyword approaches
- Monitor and measure search quality and recall metrics
- Pre-filter search space to reduce computational overhead
How to install similarity-search-patterns
npx skills add https://github.com/wshobson/agents --skill similarity-search-patternsHow to use similarity-search-patterns
- 1.Choose a distance metric appropriate for your embeddings (cosine for normalized, Euclidean for raw)
- 2.Select an index type based on data scale (Flat for small, HNSW for medium-large, IVF+PQ for very large)
- 3.Configure index parameters to balance recall and search speed for your use case
- 4.Implement pre-filtering to reduce the search space when possible
- 5.Measure and monitor recall metrics to validate search quality
- 6.Consider hybrid search combining vector similarity with keyword matching
Use cases
- Building semantic search systems over document collections
- Implementing retrieval-augmented generation (RAG) pipelines
- Creating recommendation engines based on vector similarity
- Optimizing search latency for user-facing applications
- Scaling similarity search to millions of vectors efficiently
- Backend engineers building search infrastructure
- Machine learning engineers implementing RAG systems
- Full-stack developers adding semantic search features
- Data engineers optimizing retrieval performance
- Product teams building recommendation features
similarity-search-patterns FAQ
Use cosine similarity for normalized embeddings, Euclidean (L2) for raw embeddings, dot product when magnitude matters, and Manhattan (L1) for sparse vectors.
Start with Flat for small datasets (100% recall), use HNSW for medium-large data (O(log n) search, 95-99% recall), and IVF+PQ for very large datasets (O(√n) search, 90-95% recall).
Tune parameters like ef_search and nprobe, implement pre-filtering to reduce search space, and measure recall metrics before optimizing for latency.
Yes, hybrid search combining semantic and keyword approaches is a best practice for better relevance and coverage.
Monitor and measure recall metrics to validate that your similarity search is returning relevant results for your use case.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: similarity-search-patterns description: Implement efficient similarity search with vector databases. Use when building semantic search, implementing nearest neighbor queries, or optimizing retrieval performance.
Similarity Search Patterns
Patterns for implementing efficient similarity search in production systems.
When to Use This Skill
- Building semantic search systems
- Implementing RAG retrieval
- Creating recommendation engines
- Optimizing search latency
- Scaling to millions of vectors
- Combining semantic and keyword search
Core Concepts
1. Distance Metrics
| Metric | Formula | Best For | | ------------------ | ------------------ | --------------------- | --- | -------------- | | Cosine | 1 - (A·B)/(‖A‖‖B‖) | Normalized embeddings | | Euclidean (L2) | √Σ(a-b)² | Raw embeddings | | Dot Product | A·B | Magnitude matters | | Manhattan (L1) | Σ | a-b | | Sparse vectors |
2. Index Types
┌─────────────────────────────────────────────────┐
│ Index Types │
├─────────────┬───────────────┬───────────────────┤
│ Flat │ HNSW │ IVF+PQ │
│ (Exact) │ (Graph-based) │ (Quantized) │
├─────────────┼───────────────┼───────────────────┤
│ O(n) search │ O(log n) │ O(√n) │
│ 100% recall │ ~95-99% │ ~90-95% │
│ Small data │ Medium-Large │ Very Large │
└─────────────┴───────────────┴───────────────────┘
Templates and detailed worked examples
Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.
Best Practices
Do's
- Use appropriate index - HNSW for most cases
- Tune parameters - ef_search, nprobe for recall/speed
- Implement hybrid search - Combine with keyword search
- Monitor recall - Measure search quality
- Pre-filter when possible - Reduce search space
Don'ts
- Don't skip evaluation - Measure before optimizing
- Don't over-index - Start with flat, scale up
- Don't ignore latency - P99 matters for UX
- Don't forget costs - Vector storage adds up
Related skills
More from wshobson/agents and the wider catalog.

slo-implementation
Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for measurable reliability targets.

social-publishing
Schedule and publish posts across 13 social platforms (X, LinkedIn, Instagram, TikTok, Discord, etc.) via one API key.

solidity-security
Master smart contract security best practices and prevent common Solidity vulnerabilities.

spark-environment-setup
Set up PyTorch/Unsloth/TRL on NVIDIA DGX Spark (aarch64, CUDA 13) without ABI mismatches.

spark-memory-thermal-ops
Manage unified memory and thermals for long-running ML jobs on NVIDIA DGX Spark GB10.

spark-optimization
Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning.