vector-index-tuning
wshobson/agents
Optimize vector index performance—tune HNSW parameters, quantization, and memory for production scale.
What is vector-index-tuning?
Guidance for tuning vector indexes to balance latency, recall, and memory usage. Use when optimizing HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure to handle millions or billions of vectors.
- Select appropriate index types based on dataset size (flat, HNSW, IVF+PQ, DiskANN)
- Tune HNSW parameters (M, efConstruction, efSearch) to optimize recall and search speed
- Implement quantization strategies (FP16, INT8, Product Quantization, Binary) to reduce memory footprint
- Balance trade-offs between search latency, recall accuracy, and memory consumption
- Plan tiered storage and index maintenance for production deployments
- Monitor recall degradation and manage reindexing costs
How to install vector-index-tuning
npx skills add https://github.com/wshobson/agents --skill vector-index-tuningHow to use vector-index-tuning
- 1.Determine your dataset size and select the appropriate index type from the provided table
- 2.Identify your performance constraints (latency target, recall requirement, memory budget)
- 3.Start with HNSW default parameters (M=16, efConstruction=100, efSearch=50) if applicable
- 4.Benchmark with real production queries to establish baseline latency and recall
- 5.Adjust HNSW parameters incrementally: increase M or efSearch to improve recall, decrease to reduce latency
- 6.Evaluate quantization options (FP16, INT8, Product Quantization) based on memory savings vs recall trade-off
- 7.Test index build time and plan reindexing strategy for your update frequency
- 8.Implement monitoring for recall degradation and schedule periodic reindexing
Use cases
- Reducing memory usage for a vector index serving 50M+ embeddings in production
- Tuning HNSW M and efSearch parameters to meet sub-100ms latency SLAs while maintaining 95%+ recall
- Selecting quantization method (INT8 vs Product Quantization) for a billion-vector search system
- Migrating from flat search to HNSW as dataset grows beyond 10K vectors
- Implementing hot/cold data separation to optimize storage costs for time-series embeddings
- ML/Vector Database Engineers
- Search Infrastructure Teams
- LLM Application Developers
- Data Scientists optimizing retrieval systems
- Backend Engineers scaling semantic search
vector-index-tuning FAQ
Use quantization when memory is constrained or dataset exceeds 1M vectors. INT8 scalar quantization is a good starting point; Product Quantization offers better compression for very large datasets (100M+). Always benchmark recall impact with your specific queries.
M controls connections per node during index construction—higher M improves recall but uses more memory. efSearch controls search quality at query time—higher efSearch improves recall but slows queries. Tune M during indexing, efSearch at runtime.
Profile with real queries first. Only tune if you're missing latency or recall targets. Over-optimization increases build time, memory, and maintenance cost. Start with defaults and adjust only when benchmarks show it's necessary.
Reindexing has significant cost. Plan reindexing based on recall degradation monitoring and data drift. For most systems, monthly or quarterly reindexing is sufficient unless data distribution changes rapidly.
For >100M vectors, consider IVF+PQ (Inverted File with Product Quantization) or DiskANN for disk-based indexing. HNSW with quantization works for up to ~1B vectors but requires careful parameter tuning and tiered storage.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: vector-index-tuning description: Optimize vector index performance for latency, recall, and memory. Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.
Vector Index Tuning
Guide to optimizing vector indexes for production performance.
When to Use This Skill
- Tuning HNSW parameters
- Implementing quantization
- Optimizing memory usage
- Reducing search latency
- Balancing recall vs speed
- Scaling to billions of vectors
Core Concepts
1. Index Type Selection
Data Size Recommended Index
────────────────────────────────────────
< 10K vectors → Flat (exact search)
10K - 1M → HNSW
1M - 100M → HNSW + Quantization
> 100M → IVF + PQ or DiskANN
2. HNSW Parameters
| Parameter | Default | Effect |
|---|---|---|
| M | 16 | Connections per node, ↑ = better recall, more memory |
| efConstruction | 100 | Build quality, ↑ = better index, slower build |
| efSearch | 50 | Search quality, ↑ = better recall, slower search |
3. Quantization Types
Full Precision (FP32): 4 bytes × dimensions
Half Precision (FP16): 2 bytes × dimensions
INT8 Scalar: 1 byte × dimensions
Product Quantization: ~32-64 bytes total
Binary: dimensions/8 bytes
Templates and detailed worked examples
Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.
Best Practices
Do's
- Benchmark with real queries - Synthetic may not represent production
- Monitor recall continuously - Can degrade with data drift
- Start with defaults - Tune only when needed
- Use quantization - Significant memory savings
- Consider tiered storage - Hot/cold data separation
Don'ts
- Don't over-optimize early - Profile first
- Don't ignore build time - Index updates have cost
- Don't forget reindexing - Plan for maintenance
- Don't skip warming - Cold indexes are slow
Related skills
More from wshobson/agents and the wider catalog.

vision-sft
Fine-tune vision-language models with LoRA on frozen vision towers for domain adaptation.

visual-design-foundations
Apply typography, color theory, spacing, and iconography principles to create cohesive visual designs.

visual-edit-precision
Make targeted UI edits guided by visual context and element selections.

wcag-audit-patterns
Conduct WCAG 2.2 accessibility audits with automated testing, manual verification, and remediation guidance.

web-component-design
Master React, Vue, and Svelte component patterns with composition strategies and CSS-in-JS approaches.

web3-testing
Test smart contracts comprehensively with Hardhat and Foundry—unit tests, integration tests, gas optimization, fuzzing, and mainnet forking.