PluginBench
Skill
Pass
Audit score 90

similarity-search-patterns

wshobson/agents

Implement efficient similarity search with vector databases for semantic retrieval and nearest neighbor queries.

What is similarity-search-patterns?

This skill provides patterns for building production similarity search systems using vector databases. Use it when implementing semantic search, RAG retrieval, recommendation engines, or optimizing search performance across millions of vectors.

  • Compare distance metrics (cosine, Euclidean, dot product, Manhattan) for different embedding types
  • Select appropriate index types (Flat for exact search, HNSW for medium-large data, IVF+PQ for very large scale)
  • Tune index parameters like ef_search and nprobe to balance recall and latency
  • Implement hybrid search combining semantic and keyword approaches
  • Monitor and measure search quality and recall metrics
  • Pre-filter search space to reduce computational overhead

How to install similarity-search-patterns

npx skills add https://github.com/wshobson/agents --skill similarity-search-patterns
Claude Code
Cursor
Windsurf
Cline

How to use similarity-search-patterns

  1. 1.Choose a distance metric appropriate for your embeddings (cosine for normalized, Euclidean for raw)
  2. 2.Select an index type based on data scale (Flat for small, HNSW for medium-large, IVF+PQ for very large)
  3. 3.Configure index parameters to balance recall and search speed for your use case
  4. 4.Implement pre-filtering to reduce the search space when possible
  5. 5.Measure and monitor recall metrics to validate search quality
  6. 6.Consider hybrid search combining vector similarity with keyword matching

Use cases

Good for
  • Building semantic search systems over document collections
  • Implementing retrieval-augmented generation (RAG) pipelines
  • Creating recommendation engines based on vector similarity
  • Optimizing search latency for user-facing applications
  • Scaling similarity search to millions of vectors efficiently
Who it's for
  • Backend engineers building search infrastructure
  • Machine learning engineers implementing RAG systems
  • Full-stack developers adding semantic search features
  • Data engineers optimizing retrieval performance
  • Product teams building recommendation features

similarity-search-patterns FAQ

Which distance metric should I use?

Use cosine similarity for normalized embeddings, Euclidean (L2) for raw embeddings, dot product when magnitude matters, and Manhattan (L1) for sparse vectors.

What index type is best for my data?

Start with Flat for small datasets (100% recall), use HNSW for medium-large data (O(log n) search, 95-99% recall), and IVF+PQ for very large datasets (O(√n) search, 90-95% recall).

How do I balance search speed and accuracy?

Tune parameters like ef_search and nprobe, implement pre-filtering to reduce search space, and measure recall metrics before optimizing for latency.

Should I combine vector search with keyword search?

Yes, hybrid search combining semantic and keyword approaches is a best practice for better relevance and coverage.

How do I know if my search quality is good?

Monitor and measure recall metrics to validate that your similarity search is returning relevant results for your use case.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: similarity-search-patterns description: Implement efficient similarity search with vector databases. Use when building semantic search, implementing nearest neighbor queries, or optimizing retrieval performance.

Similarity Search Patterns

Patterns for implementing efficient similarity search in production systems.

When to Use This Skill

  • Building semantic search systems
  • Implementing RAG retrieval
  • Creating recommendation engines
  • Optimizing search latency
  • Scaling to millions of vectors
  • Combining semantic and keyword search

Core Concepts

1. Distance Metrics

| Metric | Formula | Best For | | ------------------ | ------------------ | --------------------- | --- | -------------- | | Cosine | 1 - (A·B)/(‖A‖‖B‖) | Normalized embeddings | | Euclidean (L2) | √Σ(a-b)² | Raw embeddings | | Dot Product | A·B | Magnitude matters | | Manhattan (L1) | Σ | a-b | | Sparse vectors |

2. Index Types

┌─────────────────────────────────────────────────┐
│                 Index Types                      │
├─────────────┬───────────────┬───────────────────┤
│    Flat     │     HNSW      │    IVF+PQ         │
│ (Exact)     │ (Graph-based) │ (Quantized)       │
├─────────────┼───────────────┼───────────────────┤
│ O(n) search │ O(log n)      │ O(√n)             │
│ 100% recall │ ~95-99%       │ ~90-95%           │
│ Small data  │ Medium-Large  │ Very Large        │
└─────────────┴───────────────┴───────────────────┘

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Use appropriate index - HNSW for most cases
  • Tune parameters - ef_search, nprobe for recall/speed
  • Implement hybrid search - Combine with keyword search
  • Monitor recall - Measure search quality
  • Pre-filter when possible - Reduce search space

Don'ts

  • Don't skip evaluation - Measure before optimizing
  • Don't over-index - Start with flat, scale up
  • Don't ignore latency - P99 matters for UX
  • Don't forget costs - Vector storage adds up