PluginBench
Skill
Pass
Audit score 90

rag-implementation

wshobson/agents

Build RAG systems with vector databases and semantic search to ground LLMs in external knowledge.

What is rag-implementation?

Retrieval-Augmented Generation (RAG) enables LLM applications to provide accurate, factual responses by retrieving relevant documents and grounding answers in external knowledge sources. Use this skill when building Q&A systems, documentation assistants, chatbots with current information, or any application where reducing hallucinations and enabling domain-specific knowledge access is critical.

  • Store and retrieve document embeddings efficiently using vector databases (Pinecone, Weaviate, Milvus, Chroma, Qdrant, pgvector)
  • Convert text to numerical vectors using embedding models optimized for different use cases (Voyage, OpenAI, open-source alternatives)
  • Implement retrieval strategies including dense retrieval, sparse retrieval, hybrid search, multi-query, and HyDE approaches
  • Rerank retrieval results using cross-encoders, API-based reranking, MMR, or LLM-based scoring to improve quality
  • Build end-to-end RAG pipelines with LangGraph that retrieve context and generate grounded answers

How to install rag-implementation

npx skills add https://github.com/wshobson/agents --skill rag-implementation
Prerequisites
  • Vector database (Pinecone, Weaviate, Milvus, Chroma, Qdrant, or pgvector)
  • Embedding model API access or local model deployment
  • LangChain and LangGraph libraries for orchestration
  • LLM API access (Anthropic Claude recommended)
Claude Code
Cursor
Windsurf
Cline

How to use rag-implementation

  1. 1.Choose and set up a vector database for your scale and deployment model
  2. 2.Select an embedding model appropriate for your domain and language requirements
  3. 3.Prepare and chunk your documents using a text splitter
  4. 4.Embed and index your documents into the vector store
  5. 5.Implement a retriever using your vector store with appropriate search parameters
  6. 6.Build a RAG chain using LangGraph that retrieves context and generates answers
  7. 7.Test retrieval quality and iterate on embedding model, chunking strategy, or reranking if needed
  8. 8.Deploy the RAG pipeline and monitor retrieval and generation quality in production

Use cases

Good for
  • Building Q&A systems over proprietary documents or knowledge bases
  • Creating chatbots that provide current, factual information with source citations
  • Implementing semantic search with natural language queries across large document collections
  • Reducing LLM hallucinations by grounding responses in retrieved context
  • Building documentation assistants that answer questions about product or technical docs
Who it's for
  • Backend engineers building knowledge-grounded AI applications
  • Full-stack developers implementing document Q&A systems
  • Data engineers setting up vector database infrastructure
  • AI/ML engineers optimizing retrieval and ranking pipelines
  • Product teams adding intelligent search or chatbot features

rag-implementation FAQ

Which embedding model should I use?

For Claude applications, Anthropic recommends voyage-3-large (1024 dims). For OpenAI, use text-embedding-3-large (3072 dims) for high accuracy or text-embedding-3-small (1536 dims) for cost efficiency. For open-source deployments, bge-large-en-v1.5 works well; for multilingual support, use multilingual-e5-large.

What's the difference between dense and sparse retrieval?

Dense retrieval uses embeddings to find semantically similar documents; sparse retrieval uses keyword matching (BM25, TF-IDF). Hybrid search combines both approaches with weighted fusion for better coverage of semantic and keyword-based queries.

How do I reduce hallucinations in RAG?

Ensure your retriever returns high-quality, relevant context; use reranking to improve result ordering; prompt the LLM to cite sources and admit when context is insufficient; monitor retrieval quality metrics and iterate on embedding models or chunking strategies.

What vector database should I choose?

Pinecone for managed, serverless scalability; Weaviate for hybrid search and GraphQL; Chroma for lightweight local development; Qdrant for fast filtered search; pgvector for SQL integration with PostgreSQL; Milvus for high-performance on-premise deployments.

How should I chunk my documents?

Use RecursiveCharacterTextSplitter with chunk sizes typically 512–2048 tokens depending on your embedding model and retrieval needs. Overlap chunks by 10–20% to preserve context across boundaries. Adjust based on your domain and retrieval quality testing.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: rag-implementation description: Build Retrieval-Augmented Generation (RAG) systems for LLM applications with vector databases and semantic search. Use when implementing knowledge-grounded AI, building document Q&A systems, or integrating LLMs with external knowledge bases.

RAG Implementation

Master Retrieval-Augmented Generation (RAG) to build LLM applications that provide accurate, grounded responses using external knowledge sources.

When to Use This Skill

  • Building Q&A systems over proprietary documents
  • Creating chatbots with current, factual information
  • Implementing semantic search with natural language queries
  • Reducing hallucinations with grounded responses
  • Enabling LLMs to access domain-specific knowledge
  • Building documentation assistants
  • Creating research tools with source citation

Core Components

1. Vector Databases

Purpose: Store and retrieve document embeddings efficiently

Options:

  • Pinecone: Managed, scalable, serverless
  • Weaviate: Open-source, hybrid search, GraphQL
  • Milvus: High performance, on-premise
  • Chroma: Lightweight, easy to use, local development
  • Qdrant: Fast, filtered search, Rust-based
  • pgvector: PostgreSQL extension, SQL integration

2. Embeddings

Purpose: Convert text to numerical vectors for similarity search

Models (2026):

ModelDimensionsBest For
voyage-3-large1024Claude apps (Anthropic recommended)
voyage-code-31024Code search
text-embedding-3-large3072OpenAI apps, high accuracy
text-embedding-3-small1536OpenAI apps, cost-effective
bge-large-en-v1.51024Open source, local deployment
multilingual-e5-large1024Multi-language support

3. Retrieval Strategies

Approaches:

  • Dense Retrieval: Semantic similarity via embeddings
  • Sparse Retrieval: Keyword matching (BM25, TF-IDF)
  • Hybrid Search: Combine dense + sparse with weighted fusion
  • Multi-Query: Generate multiple query variations
  • HyDE: Generate hypothetical documents for better retrieval

4. Reranking

Purpose: Improve retrieval quality by reordering results

Methods:

  • Cross-Encoders: BERT-based reranking (ms-marco-MiniLM)
  • Cohere Rerank: API-based reranking
  • Maximal Marginal Relevance (MMR): Diversity + relevance
  • LLM-based: Use LLM to score relevance

Quick Start with LangGraph

from langgraph.graph import StateGraph, START, END
from langchain_anthropic import ChatAnthropic
from langchain_voyageai import VoyageAIEmbeddings
from langchain_pinecone import PineconeVectorStore
from langchain_core.documents import Document
from langchain_core.prompts import ChatPromptTemplate
from langchain_text_splitters import RecursiveCharacterTextSplitter
from typing import TypedDict, Annotated

class RAGState(TypedDict):
    question: str
    context: list[Document]
    answer: str

# Initialize components
llm = ChatAnthropic(model="claude-sonnet-5")
embeddings = VoyageAIEmbeddings(model="voyage-3-large")
vectorstore = PineconeVectorStore(index_name="docs", embedding=embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 4})

# RAG prompt
rag_prompt = ChatPromptTemplate.from_template(
    """Answer based on the context below. If you cannot answer, say so.

    Context:
    {context}

    Question: {question}

    Answer:"""
)

async def retrieve(state: RAGState) -> RAGState:
    """Retrieve relevant documents."""
    docs = await retriever.ainvoke(state["question"])
    return {"context": docs}

async def generate(state: RAGState) -> RAGState:
    """Generate answer from context."""
    context_text = "\n\n".join(doc.page_content for doc in state["context"])
    messages = rag_prompt.format_messages(
        context=context_text,
        question=state["question"]
    )
    response = await llm.ainvoke(messages)
    return {"answer": response.content}

# Build RAG graph
builder = StateGraph(RAGState)
builder.add_node("retrieve", retrieve)
builder.add_node("generate", generate)
builder.add_edge(START, "retrieve")
builder.add_edge("retrieve", "generate")
builder.add_edge("generate", END)

rag_chain = builder.compile()

# Use
result = await rag_chain.ainvoke({"question": "What are the main features?"})
print(result["answer"])

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.