PluginBench
Skill
Review
Audit score 70

langchain4j-rag-implementation-patterns

giuseppe-trisciuoglio/developer-kit

RAG implementation patterns with LangChain4j for Java: document ingestion, embeddings, and semantic search.

What is langchain4j-rag-implementation-patterns?

Provides Retrieval-Augmented Generation (RAG) implementation patterns using LangChain4j for Java. Generates document ingestion pipelines, embedding stores, vector search, and semantic search capabilities. Use when building chat-with-documents systems, document Q&A applications, AI assistants with knowledge bases, or semantic search over document repositories.

  • Configure document ingestion pipelines with automatic segmentation and embedding generation
  • Set up embedding stores and vector search for semantic retrieval
  • Create RAG-enabled AI services with context retrieval and source attribution
  • Implement hierarchical RAG with summary and chunk-level retrieval
  • Support multi-domain assistants with knowledge base access
  • Validate ingestion and retrieval with checkpoint testing

How to install langchain4j-rag-implementation-patterns

npx skills add https://github.com/giuseppe-trisciuoglio/developer-kit --skill langchain4j-rag-implementation-patterns
Prerequisites
  • Java 11 or higher
  • Spring Boot project setup
  • OpenAI API key for embeddings and chat models
  • Maven with pom.xml configuration
Claude Code
Cursor
Windsurf
Cline

How to use langchain4j-rag-implementation-patterns

  1. 1.Add LangChain4j and OpenAI dependencies to pom.xml
  2. 2.Configure EmbeddingModel and EmbeddingStore beans in Spring configuration
  3. 3.Create DocumentIngestionService to load, split, and embed documents
  4. 4.Set up ContentRetriever with embedding store and filtering parameters
  5. 5.Define AI service interface with @SystemMessage and content retrieval
  6. 6.Call the service methods to answer questions with knowledge base context

Use cases

Good for
  • Build chat-with-documents systems for PDFs, text files, or web pages
  • Create document Q&A applications over company knowledge bases
  • Implement semantic search over document repositories with relevance filtering
  • Build domain-specific AI assistants with curated knowledge and source attribution
  • Develop multi-domain assistants with access to technical docs, policies, and product information
Who it's for
  • Java developers building knowledge-enhanced AI applications
  • Backend engineers implementing document search and retrieval systems
  • AI/ML engineers creating domain-specific assistants
  • Teams building internal knowledge base systems with semantic search

langchain4j-rag-implementation-patterns FAQ

What embedding model does this use?

By default, text-embedding-3-small from OpenAI. You can configure alternative embedding models by changing the embeddingModel bean.

How do I validate that documents were ingested correctly?

Use the validateIngestion() method with a test query, or check that embedding count matches segment count after ingestion.

Can I use a different embedding store instead of in-memory?

Yes, LangChain4j supports multiple embedding stores. Replace InMemoryEmbeddingStore with alternatives like Pinecone, Weaviate, or Milvus by changing the bean configuration.

How do I handle large documents?

Use DocumentSplitters.recursive() with appropriate chunk size (500-1000 tokens) and overlap (20-50 tokens) to break documents into manageable segments before embedding.

Does this support source attribution?

Yes, metadata is preserved through the ingestion pipeline and can be included in system prompts to reference specific sources in responses.

Full instructions (SKILL.md)

Source of truth, from giuseppe-trisciuoglio/developer-kit.


name: langchain4j-rag-implementation-patterns description: Provides Retrieval-Augmented Generation (RAG) implementation patterns with LangChain4j for Java. Generates document ingestion pipelines, embedding stores, vector search, and semantic search capabilities. Use when building chat-with-documents systems, document Q&A over PDFs or text files, AI assistants with knowledge bases, semantic search over document repositories, or knowledge-enhanced AI applications with source attribution. allowed-tools: Read, Write, Bash

LangChain4j RAG Implementation Patterns

Overview

Implements RAG systems with LangChain4j: document ingestion pipelines, embedding stores, and vector search for chat-with-documents and knowledge-enhanced AI applications.

When to Use This Skill

  • Building chat-with-documents systems or document Q&A over PDFs, text files, or web pages
  • Creating AI assistants with access to company knowledge bases or external sources
  • Implementing semantic search or hybrid search over document repositories
  • Building domain-specific AI with curated knowledge and source attribution

Instructions

Initialize RAG Project

Create a new Spring Boot project with required dependencies:

pom.xml:

<dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j-spring-boot-starter</artifactId>
    <version>1.8.0</version>
</dependency>
<dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j-open-ai</artifactId>
    <version>1.8.0</version>
</dependency>

Setup Document Ingestion

Configure document loading and processing with validation:

Validation Checkpoint: After ingestion, verify embedding count matches segment count and test retrieval with a sample query.

@Configuration
public class RAGConfiguration {

    @Bean
    public EmbeddingModel embeddingModel() {
        return OpenAiEmbeddingModel.builder()
            .apiKey(System.getenv("OPENAI_API_KEY"))
            .modelName("text-embedding-3-small")
            .build();
    }

    @Bean
    public EmbeddingStore<TextSegment> embeddingStore() {
        return new InMemoryEmbeddingStore<>();
    }
}

Create document ingestion service:

@Service
@RequiredArgsConstructor
public class DocumentIngestionService {

    private final EmbeddingModel embeddingModel;
    private final EmbeddingStore<TextSegment> embeddingStore;

    public void ingestDocument(String filePath, Map<String, Object> metadata) {
        Document document = FileSystemDocumentLoader.loadDocument(filePath);
        document.metadata().putAll(metadata);

        DocumentSplitter splitter = DocumentSplitters.recursive(
            500, 50, new OpenAiTokenCountEstimator("text-embedding-3-small")
        );

        List<TextSegment> segments = splitter.split(document);
        List<Embedding> embeddings = embeddingModel.embedAll(segments).content();
        embeddingStore.addAll(embeddings, segments);

        // Validation: verify embedding count matches segments
        if (embeddings.size() != segments.size()) {
            throw new IllegalStateException("Embedding count mismatch: expected " + segments.size() + ", got " + embeddings.size());
        }
    }

    public boolean validateIngestion(String testQuery) {
        // Validation: test retrieval with sample query
        Embedding queryEmbedding = embeddingModel.embed(testQuery).content();
        List<EmbeddingMatch<TextSegment>> results = embeddingStore.search(
            EmbeddingSearchRequest.builder()
                .queryEmbedding(queryEmbedding)
                .maxResults(1)
                .build()
        ).matches();
        return !results.isEmpty();
    }
}

Configure Content Retrieval

Setup content retrieval with filtering:

Validation Checkpoint: After configuration, test retrieval with a known query to verify embeddings are searchable.

@Configuration
public class ContentRetrieverConfiguration {

    @Bean
    public ContentRetriever contentRetriever(
            EmbeddingStore<TextSegment> embeddingStore,
            EmbeddingModel embeddingModel) {

        return EmbeddingStoreContentRetriever.builder()
            .embeddingStore(embeddingStore)
            .embeddingModel(embeddingModel)
            .maxResults(5)
            .minScore(0.7)
            .build();
    }
}

Create RAG-Enabled AI Service

Define AI service with context retrieval:

interface KnowledgeAssistant {
    @SystemMessage("""
        You are a knowledgeable assistant with access to a comprehensive knowledge base.

        When answering questions:
        1. Use the provided context from the knowledge base
        2. If information is not in the context, clearly state this
        3. Provide accurate, helpful responses
        4. When possible, reference specific sources
        5. If the context is insufficient, ask for clarification
        """)
    String answerQuestion(String question);
}

@Service
@RequiredArgsConstructor
public class KnowledgeService {

    private final KnowledgeAssistant assistant;

    public KnowledgeService(ChatModel chatModel, ContentRetriever contentRetriever) {
        this.assistant = AiServices.builder(KnowledgeAssistant.class)
            .chatModel(chatModel)
            .contentRetriever(contentRetriever)
            .build();
    }

    public String answerQuestion(String question) {
        return assistant.answerQuestion(question);
    }
}

Examples

Basic Document Processing

public class BasicRAGExample {
    public static void main(String[] args) {
        var embeddingStore = new InMemoryEmbeddingStore<TextSegment>();

        var embeddingModel = OpenAiEmbeddingModel.builder()
            .apiKey(System.getenv("OPENAI_API_KEY"))
            .modelName("text-embedding-3-small")
            .build();

        var ingestor = EmbeddingStoreIngestor.builder()
            .embeddingModel(embeddingModel)
            .embeddingStore(embeddingStore)
            .build();

        ingestor.ingest(Document.from("Spring Boot is a framework for building Java applications with minimal configuration."));

        var retriever = EmbeddingStoreContentRetriever.builder()
            .embeddingStore(embeddingStore)
            .embeddingModel(embeddingModel)
            .build();
    }
}

Multi-Domain Assistant

interface MultiDomainAssistant {
    @SystemMessage("""
        You are an expert assistant with access to multiple knowledge domains:
        - Technical documentation
        - Company policies
        - Product information
        - Customer support guides

        Tailor your response based on the type of question and available context.
        Always indicate which domain the information comes from.
        """)
    String answerQuestion(@MemoryId String userId, String question);
}

Hierarchical RAG

@Service
@RequiredArgsConstructor
public class HierarchicalRAGService {

    private final EmbeddingStore<TextSegment> chunkStore;
    private final EmbeddingStore<TextSegment> summaryStore;
    private final EmbeddingModel embeddingModel;

    public String performHierarchicalRetrieval(String query) {
        List<EmbeddingMatch<TextSegment>> summaryMatches = searchSummaries(query);
        List<TextSegment> relevantChunks = new ArrayList<>();

        for (EmbeddingMatch<TextSegment> summaryMatch : summaryMatches) {
            String documentId = summaryMatch.embedded().metadata().getString("documentId");
            List<EmbeddingMatch<TextSegment>> chunkMatches = searchChunksInDocument(query, documentId);
            chunkMatches.stream()
                .map(EmbeddingMatch::embedded)
                .forEach(relevantChunks::add);
        }

        return generateResponseWithChunks(query, relevantChunks);
    }
}

Best Practices

Document Segmentation

  • Use recursive splitting with 500-1000 token chunks for most applications
  • Maintain 20-50 token overlap between chunks for context preservation
  • Consider document structure (headings, paragraphs) when splitting
  • Use token-aware splitters for optimal embedding generation

Metadata Strategy

  • Include rich metadata for filtering and attribution:
    • User and tenant identifiers for multi-tenancy
    • Document type and category classification
    • Creation and modification timestamps
    • Version and author information
    • Confidentiality and access level tags

Query Processing

  • Implement query preprocessing and cleaning
  • Consider query expansion for better recall
  • Apply dynamic filtering based on user context
  • Use re-ranking for improved result quality

Performance Optimization

  • Cache embeddings for repeated queries
  • Use batch embedding generation for bulk operations
  • Implement pagination for large result sets
  • Consider asynchronous processing for long operations

Common Patterns

Simple RAG Pipeline

@RequiredArgsConstructor
@Service
public class SimpleRAGPipeline {

    private final EmbeddingModel embeddingModel;
    private final EmbeddingStore<TextSegment> embeddingStore;
    private final ChatModel chatModel;

    public String answerQuestion(String question) {
        Embedding queryEmbedding = embeddingModel.embed(question).content();
        EmbeddingSearchRequest request = EmbeddingSearchRequest.builder()
            .queryEmbedding(queryEmbedding)
            .maxResults(3)
            .build();

        List<TextSegment> segments = embeddingStore.search(request).matches().stream()
            .map(EmbeddingMatch::embedded)
            .collect(Collectors.toList());

        String context = segments.stream()
            .map(TextSegment::text)
            .collect(Collectors.joining("\n\n"));

        return chatModel.generate(context + "\n\nQuestion: " + question + "\nAnswer:");
    }
}

Hybrid Search (Vector + Keyword)

@Service
@RequiredArgsConstructor
public class HybridSearchService {

    private final EmbeddingStore<TextSegment> vectorStore;
    private final FullTextSearchEngine keywordEngine;
    private final EmbeddingModel embeddingModel;

    public List<Content> hybridSearch(String query, int maxResults) {
        // Vector search
        List<Content> vectorResults = performVectorSearch(query, maxResults);

        // Keyword search
        List<Content> keywordResults = performKeywordSearch(query, maxResults);

        // Combine and re-rank using RRF algorithm
        return combineResults(vectorResults, keywordResults, maxResults);
    }
}

Troubleshooting

Validation Failures

Embedding Count Mismatch: Thrown when segments != embeddings. Check splitter configuration and model availability.

Empty Retrieval Results: Call validateIngestion(testQuery) to verify embeddings are searchable. Check if document was ingested successfully.

Low Retrieval Scores: Verify minScore threshold (default 0.7) is not too high for your use case. Test with known queries.

Common Issues

Poor Retrieval Results

  • Check document chunk size and overlap settings
  • Verify embedding model compatibility
  • Ensure metadata filters are not too restrictive
  • Consider adding re-ranking step
  • Run validation to confirm embeddings exist

Slow Performance

  • Use cached embeddings for frequent queries
  • Optimize database indexing for vector stores
  • Implement pagination for large datasets
  • Consider async processing for bulk operations

High Memory Usage

  • Use disk-based embedding stores for large datasets
  • Implement proper pagination and filtering
  • Clean up unused embeddings periodically
  • Monitor and optimize chunk sizes

Constraints and Warnings

  • Embedding Model Costs: Generating embeddings for large document collections can be expensive; implement caching and batch processing.
  • Vector Store Scalability: In-memory stores are suitable for development only; use persistent stores (Pinecone, Qdrant, Redis) for production.
  • Chunk Size Trade-offs: Smaller chunks improve precision but lose context; larger chunks preserve context but may introduce noise.
  • Stale Data: Cached embeddings become stale when source documents change; implement update strategies.
  • Token Limits: RAG context windows have limits; typically 3-5 retrieved chunks fit within standard model limits.
  • Hallucination Risk: RAG reduces but doesn't eliminate hallucinations; always validate critical responses against sources.
  • Latency: Vector search and embedding generation add latency; consider async processing for real-time applications.
  • Metadata Filtering: Overly restrictive filters may return no results; implement fallback strategies.
  • Multi-tenancy: Ensure proper metadata isolation to prevent cross-tenant data leakage.

References

Related skills

More from giuseppe-trisciuoglio/developer-kit and the wider catalog.

LAlangchain4j-spring-boot-integration logo

langchain4j-spring-boot-integration

giuseppe-trisciuoglio/developer-kit

Integrate LangChain4j with Spring Boot using declarative AI Services, auto-configuration, and dependency injection.

1.4k installs
LAlangchain4j-testing-strategies logo

langchain4j-testing-strategies

giuseppe-trisciuoglio/developer-kit

Unit, integration, and mock testing patterns for LangChain4j Java AI services and RAG workflows.

1.4k installsAudited
LAlangchain4j-tool-function-calling-patterns logo

langchain4j-tool-function-calling-patterns

giuseppe-trisciuoglio/developer-kit

Annotate Java methods as LLM-callable tools, register with AiServices, and handle execution errors in LangChain4j agents.

1.4k installs
LAlangchain4j-vector-stores-configuration logo

langchain4j-vector-stores-configuration

giuseppe-trisciuoglio/developer-kit

Configure LangChain4J vector stores for RAG applications with PostgreSQL, Pinecone, MongoDB, Milvus, and Neo4j.

1.4k installsAudited
LElearn logo

learn

giuseppe-trisciuoglio/developer-kit

Provides autonomous project pattern learning by analyzing the codebase to discover development conventions, architectural patterns, and coding standards, then generates project rule files in .claude/rules/. Use when user asks to "learn from project", "extract project rules", "analyze codebase conventions", "discover project patterns", or wants to auto-generate Claude Code rules for the current project.

953 installsAudited
MEmemory-md-management logo

memory-md-management

giuseppe-trisciuoglio/developer-kit

Provides comprehensive memory file management capabilities including auditing, quality assessment, and targeted improvements for files such as CLAUDE.md. Use when user asks to check, audit, update, improve, fix, maintain, or validate project memory files. Also triggers for "project memory optimization", "CLAUDE.md quality check", "documentation review", or when a project memory file needs to be created from scratch. This skill scans memory files, evaluates quality against standardized criteria, outputs detailed quality reports with scores and recommendations, then makes targeted updates with user approval.

1.1k installsAudited