PluginBench
Skill
Pass
Audit score 90

vector-index-tuning

wshobson/agents

Optimize vector index performance—tune HNSW parameters, quantization, and memory for production scale.

What is vector-index-tuning?

Guidance for tuning vector indexes to balance latency, recall, and memory usage. Use when optimizing HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure to handle millions or billions of vectors.

  • Select appropriate index types based on dataset size (flat, HNSW, IVF+PQ, DiskANN)
  • Tune HNSW parameters (M, efConstruction, efSearch) to optimize recall and search speed
  • Implement quantization strategies (FP16, INT8, Product Quantization, Binary) to reduce memory footprint
  • Balance trade-offs between search latency, recall accuracy, and memory consumption
  • Plan tiered storage and index maintenance for production deployments
  • Monitor recall degradation and manage reindexing costs

How to install vector-index-tuning

npx skills add https://github.com/wshobson/agents --skill vector-index-tuning
Claude Code
Cursor
Windsurf
Cline

How to use vector-index-tuning

  1. 1.Determine your dataset size and select the appropriate index type from the provided table
  2. 2.Identify your performance constraints (latency target, recall requirement, memory budget)
  3. 3.Start with HNSW default parameters (M=16, efConstruction=100, efSearch=50) if applicable
  4. 4.Benchmark with real production queries to establish baseline latency and recall
  5. 5.Adjust HNSW parameters incrementally: increase M or efSearch to improve recall, decrease to reduce latency
  6. 6.Evaluate quantization options (FP16, INT8, Product Quantization) based on memory savings vs recall trade-off
  7. 7.Test index build time and plan reindexing strategy for your update frequency
  8. 8.Implement monitoring for recall degradation and schedule periodic reindexing

Use cases

Good for
  • Reducing memory usage for a vector index serving 50M+ embeddings in production
  • Tuning HNSW M and efSearch parameters to meet sub-100ms latency SLAs while maintaining 95%+ recall
  • Selecting quantization method (INT8 vs Product Quantization) for a billion-vector search system
  • Migrating from flat search to HNSW as dataset grows beyond 10K vectors
  • Implementing hot/cold data separation to optimize storage costs for time-series embeddings
Who it's for
  • ML/Vector Database Engineers
  • Search Infrastructure Teams
  • LLM Application Developers
  • Data Scientists optimizing retrieval systems
  • Backend Engineers scaling semantic search

vector-index-tuning FAQ

When should I use quantization vs full precision?

Use quantization when memory is constrained or dataset exceeds 1M vectors. INT8 scalar quantization is a good starting point; Product Quantization offers better compression for very large datasets (100M+). Always benchmark recall impact with your specific queries.

What's the difference between M and efSearch in HNSW?

M controls connections per node during index construction—higher M improves recall but uses more memory. efSearch controls search quality at query time—higher efSearch improves recall but slows queries. Tune M during indexing, efSearch at runtime.

How do I know if my index is over-optimized?

Profile with real queries first. Only tune if you're missing latency or recall targets. Over-optimization increases build time, memory, and maintenance cost. Start with defaults and adjust only when benchmarks show it's necessary.

Should I reindex frequently?

Reindexing has significant cost. Plan reindexing based on recall degradation monitoring and data drift. For most systems, monthly or quarterly reindexing is sufficient unless data distribution changes rapidly.

What index type should I use for billions of vectors?

For >100M vectors, consider IVF+PQ (Inverted File with Product Quantization) or DiskANN for disk-based indexing. HNSW with quantization works for up to ~1B vectors but requires careful parameter tuning and tiered storage.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: vector-index-tuning description: Optimize vector index performance for latency, recall, and memory. Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.

Vector Index Tuning

Guide to optimizing vector indexes for production performance.

When to Use This Skill

  • Tuning HNSW parameters
  • Implementing quantization
  • Optimizing memory usage
  • Reducing search latency
  • Balancing recall vs speed
  • Scaling to billions of vectors

Core Concepts

1. Index Type Selection

Data Size           Recommended Index
────────────────────────────────────────
< 10K vectors  →    Flat (exact search)
10K - 1M       →    HNSW
1M - 100M      →    HNSW + Quantization
> 100M         →    IVF + PQ or DiskANN

2. HNSW Parameters

ParameterDefaultEffect
M16Connections per node, ↑ = better recall, more memory
efConstruction100Build quality, ↑ = better index, slower build
efSearch50Search quality, ↑ = better recall, slower search

3. Quantization Types

Full Precision (FP32): 4 bytes × dimensions
Half Precision (FP16): 2 bytes × dimensions
INT8 Scalar:           1 byte × dimensions
Product Quantization:  ~32-64 bytes total
Binary:                dimensions/8 bytes

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Benchmark with real queries - Synthetic may not represent production
  • Monitor recall continuously - Can degrade with data drift
  • Start with defaults - Tune only when needed
  • Use quantization - Significant memory savings
  • Consider tiered storage - Hot/cold data separation

Don'ts

  • Don't over-optimize early - Profile first
  • Don't ignore build time - Index updates have cost
  • Don't forget reindexing - Plan for maintenance
  • Don't skip warming - Cold indexes are slow