Siftdog
Open-source, self-hosted Tavily alternative: web search, extract, crawl and map for AI agents.
Siftdog is a self-hosted web search API for AI agents and LLM apps, and a drop-in replacement for the Tavily API. It serves the same /search, /extract, /crawl and /map endpoints with the same request and response shapes, so code written for Tavily works against your own server by changing only the base URL. It needs no search API key and is MIT licensed.
Features
- Search through a SearXNG metasearch instance, with no search API keys
- Extraction of clean markdown or text from web pages with trafilatura
- Ranking: BM25 relevance blended with the search engine's order;
advanceddepth returns each page's most relevant passages - Crawling and site maps with depth, breadth, limit and regex path/domain filters
- MCP server for Claude Code, Claude Desktop, Cursor and other MCP clients, over HTTP at
/mcpor stdio - SSRF protection: private and internal addresses are blocked, including via redirects and DNS rebinding
Quick start
The full stack, with SearXNG for search:
git clone https://github.com/khsarvar/siftdog && cd siftdog
docker compose up
Or run the image on its own (/extract, /crawl and /map work without SearXNG), or install from PyPI:
docker run -p 8000:8000 ghcr.io/khsarvar/siftdog
pip install siftdog
Drop-in replacement for the Tavily API
Point the official Tavily Python SDK or LangChain integration at your server:
from tavily import TavilyClient
client = TavilyClient(api_key="your-siftdog-key", api_base_url="http://localhost:8000")
results = client.search("latest python release", search_depth="advanced")
Tested with tavily-python 0.8.4 (search, extract, crawl, map) and langchain-tavily 0.2.18 (search, extract).
MCP server for Claude, Cursor and other clients
Four read-only tools: siftdog_search, siftdog_extract, siftdog_crawl and siftdog_map. Run locally over stdio with uv:
claude mcp add siftdog -- uvx siftdog mcp
Or connect to a running server over HTTP:
claude mcp add --transport http siftdog http://localhost:8000/mcp
Listed in the official MCP Registry as io.github.khsarvar/siftdog.
Siftdog vs Tavily, Firecrawl and crw
| Siftdog | Tavily | Firecrawl | crw | |
|---|---|---|---|---|
| License | MIT | Proprietary (hosted service) | AGPL-3.0 | AGPL-3.0 |
| Self-hosted | Yes | No | Yes | Yes (also a managed API) |
| Tavily-compatible API | Yes | — | No (own API) | No (own API) |
| Search source | SearXNG metasearch | Proprietary | — | SearXNG (bundled) |
| JavaScript rendering | No (static HTML) | — | Yes | Yes (Lightpanda, Chrome fallback) |
| MCP server | Yes (HTTP and stdio) | Yes | Yes | Yes |
| Language | Python | — | TypeScript | Rust |
Measured results: Siftdog vs Tavily search benchmark.
If you'd rather not run infrastructure, use hosted Tavily. If you need JavaScript-rendered pages, look at Firecrawl or crw. Siftdog is for teams that want the Tavily API on their own servers under a permissive license.
FAQ
What is Siftdog?
Siftdog is an open-source web search and extraction API for AI agents. It reproduces the Tavily API (/search, /extract, /crawl, /map) on infrastructure you run yourself, using SearXNG for search results, trafilatura for content extraction and BM25 for relevance ranking.
Is Siftdog a drop-in replacement for Tavily?
For the four core endpoints, yes. Request and response fields match Tavily's, and the official tavily-python SDK and langchain-tavily work by setting api_base_url. Relevance scores come from BM25 rather than a neural reranker, and Tavily's /research endpoint isn't implemented.
Do I need a search API key?
No. Search results come from SearXNG, which queries public search engines. The only optional key is ANTHROPIC_API_KEY, used when a request sets include_answer.
Does Siftdog have an MCP server?
Yes. The API server exposes MCP over Streamable HTTP at /mcp, and uvx siftdog mcp runs it over stdio for local clients such as Claude Desktop, Claude Code and Cursor. It provides four tools: siftdog_search, siftdog_extract, siftdog_crawl and siftdog_map.
Is it safe to expose Siftdog on a public server?
Set API_KEYS so only your clients can call it. Because /extract and /crawl fetch caller-supplied URLs, Siftdog refuses private, loopback and link-local addresses, checked on every redirect and at connect time against the exact address used, which also stops DNS rebinding.