
Local Rag MCP Server
by shinprio.github.shinpr/mcp-local-ragv0.21.1
Local document search: semantic search over PDF, DOCX, TXT and Markdown files with LanceDB vector storage and local embeddings.
context tax
queued
security
queued
cold start
queued
freshness
Active2d ago
Install Local Rag MCP server
Install in Claude Code
claude mcp add local-rag -- npx -y mcp-local-ragInstall in Cursor
{
"mcpServers": {
"local-rag": {
"command": "npx",
"args": [
"-y",
"mcp-local-rag"
]
}
}
}Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (project).
Install in Claude Desktop
{
"mcpServers": {
"local-rag": {
"command": "npx",
"args": [
"-y",
"mcp-local-rag"
]
}
}
}Settings → Developer → Edit Config (claude_desktop_config.json), then restart.
Install in VS Code
{
"servers": {
"local-rag": {
"type": "stdio",
"command": "npx",
"args": [
"-y",
"mcp-local-rag"
]
}
}
}Add to .vscode/mcp.json in your workspace.
Install in Windsurf
{
"mcpServers": {
"local-rag": {
"command": "npx",
"args": [
"-y",
"mcp-local-rag"
]
}
}
}Add to ~/.codeium/windsurf/mcp_config.json.
Configuration
| Variable | Required | Secret | Description |
|---|---|---|---|
| BASE_DIR | — | — | Base directory for document storage (defaults to current working directory). Ignored when BASE_DIRS is set. |
| BASE_DIRS | — | — | JSON array of base directories (e.g. '["/a","/b"]'). Takes precedence over BASE_DIR. |
| DB_PATH | — | — | Path to LanceDB database directory (defaults to ./lancedb/) |
| CACHE_DIR | — | — | Directory where Transformers.js models are cached (defaults to ./models/) |
| HF_ENDPOINT | — | — | Hugging Face model download endpoint. Set this to a mirror URL when direct downloads are blocked (defaults to https://huggingface.co). |
| MODEL_NAME | — | — | Embedding model name (defaults to Xenova/all-MiniLM-L6-v2) |
| MAX_FILE_SIZE | — | — | Maximum file size in bytes (defaults to 104857600 / 100MB) |
| RAG_MAX_DISTANCE | — | — | Maximum distance threshold for filtering search results. Results with distance greater than this value will be excluded. Lower values mean stricter filtering (e.g., 0.5 for high relevance only) |
| RAG_GROUPING | — | — | Grouping mode for quality filtering. 'similar' returns only the most similar group (stops at first distance jump). 'related' includes related groups (stops at second distance jump). Unset means no grouping filter |
| RAG_MAX_FILES | — | — | Maximum number of files to keep in search results. Results are filtered to include only chunks from the top N best-scoring files. For example, 1 returns only the single best-matching file's chunks. Unset means no file filtering. |
| CHUNK_MIN_LENGTH | — | — | Minimum chunk length in characters (1-10000, defaults to 50). Chunks shorter than this threshold are filtered out during ingestion. |
| STORE_IMAGES | — | — | Store supported PDF and DOCX images during ingestion and return them with matched chunks (defaults to false). |
| EMBED_TITLE_PREFIX | — | — | Embed each chunk together with its document title, which can help when passages don't restate the topic the title names (defaults to false). After changing it, use a new DB_PATH or delete the index and re-ingest. |
| EMBED_HEADING_PREFIX | — | — | Add section headings to chunk embeddings when they fit (defaults to false). Independent of EMBED_TITLE_PREFIX. Re-ingest documents after changing it. |
| RAG_DEVICE | — | — | Execution device for the embedder (defaults to cpu). Passed straight to ONNX Runtime; see the Transformers.js device source for the supported backend names. If the requested device fails to initialize, the server throws an error. |
| RAG_DTYPE | — | — | Embedding quantization dtype for the embedder (defaults to fp32). Opt-in and pass-through; accepts any dtype the chosen model provides (fp32, fp16, q8, int8, ...). If the model has no variant for the requested dtype, the server throws an error. Changing this changes the embedding space — re-ingest existing data. |
| RAG_HYBRID_WEIGHT | — | — | Keyword boost factor for hybrid search (0.0-1.0, defaults to 0.6). 0 means semantic similarity only; higher values increase the keyword-match contribution to the final score. |
| RAG_RERANK_CMD | — | — | External reranker command template. Use {query} and {top} for query text and result count; unset disables reranking. |
| RAG_RERANK_TIMEOUT_MS | — | — | Time budget per rerank call in milliseconds (100-600000, defaults to 10000). On timeout the spawned command is killed and the pre-rerank ordering is returned. |
Freshness
Active — last maintenance signal 2d ago. The newest of the signals below sets the band.
Last commit (default branch)
2026-10-07 · 2d ago · GitHub
Latest release
2026-10-07 · 2d ago · GitHub · v0.21.1
Package published
no data · npm/PyPI
Registry entry updated
2026-10-07 · 2d ago · official registry · v0.21.1
FAQ
›How do I install the Local Rag MCP server in Claude Code?
Run: claude mcp add local-rag -- npx -y mcp-local-rag. For Cursor, VS Code, Claude Desktop and Windsurf, use the install tabs above.
›Does Local Rag require an API key?
No required environment variables are declared in its published metadata.
›Can I use Local Rag as a remote (hosted) MCP server?
No hosted endpoint is published; it runs locally over stdio.
›Is Local Rag in the official MCP registry?
Yes, as io.github.shinpr/mcp-local-rag.
Alternatives to Local Rag
Other notes & knowledge MCP servers.