redis-semantic-cache
redis/agent-skills
Semantic caching for LLM responses on Redis Cloud — reduce API costs and latency with embedding-based prompt matching.
What is redis-semantic-cache?
Redis LangCache provides semantic caching for LLM completions and RAG answers by storing prompts as embeddings and returning cached responses for semantically similar queries. Use it to build a cache-aside layer in front of OpenAI, Anthropic, or other LLM providers, tuning similarity thresholds to balance hit rate against precision.
- Cache LLM responses using semantic similarity instead of exact-match keys
- Search cached prompts via SDK or REST API with configurable similarity thresholds
- Separate caches by task type or filter results using custom attributes
- Reduce LLM API calls and latency for deterministic workloads like RAG and classification
- Tune precision vs. hit-rate trade-off for different use cases
How to install redis-semantic-cache
npx skills add https://github.com/redis/agent-skills --skill redis-semantic-cache- Redis Cloud account with LangCache preview enabled
- LangCache SDK (Python) or REST API access
- LLM provider API key (OpenAI, Anthropic, etc.)
- Environment variables: HOST, CACHE_ID, API_KEY
How to use redis-semantic-cache
- 1.Create a LangCache instance with your Redis Cloud URL, cache ID, and API key
- 2.Before calling your LLM, call `search(prompt, similarity_threshold)` to check for cached hits
- 3.If a hit is found, return the cached response; otherwise call your LLM
- 4.After receiving the LLM response, call `set(prompt, response)` to cache it for future similar queries
- 5.Monitor cache-hit rate and adjust the similarity threshold (0.8–0.95+) based on your precision requirements
- 6.For multi-task applications, create separate cache instances or use custom attributes to isolate workloads
Use cases
- Wrapping OpenAI or Anthropic calls with a cache layer to cut API costs
- Caching RAG answers to avoid re-embedding and re-querying for similar questions
- Deduplicating FAQ or support queries with loose semantic matching
- Splitting an application's code-generation and customer-support LLM workloads into isolated caches
- Building a cache-aside pattern for internal tools and exploratory queries
- Backend engineers building LLM applications
- AI/ML teams optimizing inference costs
- Developers building RAG systems or chatbots
- Product teams managing customer-facing LLM features
redis-semantic-cache FAQ
Start with 0.9 as a balanced default. Use 0.95+ for customer-facing answers where wrong responses are costly, and 0.8 for internal tools or FAQ deduplication where higher hit rate matters more.
Yes. LangCache is provider-agnostic — it caches the prompt and response, so you can use it in front of OpenAI, Anthropic, or any other LLM API.
No. Create separate cache IDs for different workloads (e.g., support vs. code generation) to avoid semantic collisions. Alternatively, use custom attributes to filter results within a single cache.
LangCache is currently in preview on Redis Cloud, so features and behavior may change. Check the official documentation for the latest status.
POST to `/v1/caches/{cacheId}/entries/search` for lookups and `/v1/caches/{cacheId}/entries` for writes, passing the prompt, response, and optional attributes in the request body.
Full instructions (SKILL.md)
Source of truth, from redis/agent-skills.
name: redis-semantic-cache description: Redis LangCache guidance for semantic caching of LLM responses on Redis Cloud — calling search/set via the SDK or REST API, tuning the similarity threshold, separating caches per task type, and filtering with custom attributes. Use when caching LLM completions or RAG answers to cut API cost and latency, building a cache-aside layer in front of OpenAI / Anthropic / etc., tuning hit rate vs precision, or splitting one app's LLM workloads into multiple LangCache caches. license: MIT metadata: author: Redis, Inc. version: "0.1.0"
Redis Semantic Cache
Semantic caching for LLM responses with Redis Cloud's LangCache service. Stores prompts as embeddings; subsequent semantically-similar prompts return the cached response without re-calling the model.
LangCache is currently in preview on Redis Cloud. Features and behavior may change.
When to apply
- Wrapping an LLM call (OpenAI, Anthropic, etc.) with a cache layer to cut cost and latency.
- Caching RAG answers, classification outputs, or any deterministic LLM workload.
- Tuning the precision/hit-rate trade-off for a semantic cache.
- Splitting one application's LLM workloads across multiple cache instances.
1. The cache-aside flow
LangCache fits in front of any LLM call as a standard cache-aside pattern:
- Send the user's prompt to LangCache's
search. - Cache hit — return the stored response directly.
- Cache miss — call the LLM, then
setthe response so future similar prompts hit.
from langcache import LangCache
import os
lang_cache = LangCache(
server_url=f"https://{os.getenv('HOST')}",
cache_id=os.getenv("CACHE_ID"),
api_key=os.getenv("API_KEY"),
)
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.9)
if result:
response = result[0]["response"]
else:
response = llm.generate("What is Redis?")
lang_cache.set(prompt="What is Redis?", response=response)
The same operations are available via REST (POST /v1/caches/{cacheId}/entries/search and POST /v1/caches/{cacheId}/entries) when an SDK isn't an option.
See references/langcache-usage.md for full SDK + REST samples and attribute-based storage.
2. Tune the similarity threshold
The threshold controls how close (in embedding cosine distance) a new prompt must be to a cached one to count as a hit. Higher = stricter match, fewer false positives. Lower = more hits, more risk of returning an off-topic answer.
| Threshold | Behavior | Use when |
|---|---|---|
| 0.95+ | Near-exact match required | Customer-facing answers where wrong responses are costly |
| 0.9 | Balanced default | Most workloads — start here |
| 0.8 | Loose semantic match | Internal tools, exploratory queries, FAQ deduplication |
# Stricter — fewer false positives
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.95)
# Looser — higher hit rate
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.8)
Adjust by watching the actual cache-hit rate and spot-checking that returned answers are still relevant.
See references/best-practices.md.
3. Separate caches per task type
Different LLM workloads should not share one cache — a "code question" prompt is semantically close to other code questions but has nothing to do with a password-reset support query, and crossing them returns garbage.
support_cache = LangCache(server_url=..., cache_id="support-cache-id", api_key=...)
code_cache = LangCache(server_url=..., cache_id="code-cache-id", api_key=...)
Create distinct cache IDs in Redis Cloud per task, and route each call to the right one. As a finer-grained alternative, store and search with custom attributes (e.g. {"category": "database"}) to keep tasks in the same cache but isolated by attribute filter — useful when the same prompt format spans subtopics.
References
Related skills
More from redis/agent-skills and the wider catalog.

iris-development
Persistent memory layer for AI agents on Redis Cloud with session and long-term memory tiers.

redis-clustering
Design Redis Cluster keys with hash tags and route reads to replicas to avoid CROSSSLOT errors and scale read-heavy workloads.

redis-connections
Configure Redis clients efficiently with pooling, pipelining, and client-side caching.

redis-core
Choose the right Redis data structure and key-naming convention for your access pattern.

agent-ci
Run GitHub Actions CI locally before pushing to validate changes instantly.

agent-ci
Agent skill from redwoodjs/local-ci.