Darash
Darash is an all-in-one research crate: provider-neutral async web search, page fetching, and HTML extraction in one dependency, with no API keys. Search runs on a small in-process multi-source backend by default and can also query a remote SearxNG endpoint; fetch and extraction turn any page into markdown, outlines, selector matches, rows, and tables — local, deterministic, and shaped for a context window.
External SearxNG
The official SearxNG Docker Compose setup is the quickest local instance. It requires Docker with Compose:
SearxNG is then available at http://localhost:8080. Stop it with
docker compose down. See the SearxNG container documentation
for configuration and maintenance.
Use the crate
Add Darash from crates.io:
[]
= "0.6.0"
= { = "1", = ["macros", "rt-multi-thread"] }
Create a client, build a query, and search asynchronously:
use ;
async
SearchResponse also exposes the raw SearxNG query, result count, results,
answers, corrections, and suggestions. Its optional answer is a backend
answer when one is supplied; Darash does not call an AI provider. sources and
cited_sources() expose the cited Citation values for host-owned synthesis.
Each SearchResult includes its title, URL, content, engines, category,
publication date, and score.
The Vane-compatible request contract carries a search mode and source selection without adding provider credentials:
use ;
async
SearchClient::local() runs Darash's provider adapters directly in the current
process. The default backend queries DuckDuckGo, OpenAlex, and Hacker News as
needed; it does not start a separate search service. SearchMode supports
Speed, Balanced (the default), and Quality.
SearchSource supports Web (the default), Academic, and Discussions.
JSON requests may omit mode and sources; they default to balanced and
web.
The host can synthesize an answer from the returned sources with its own model.
The local providers are selected from the requested sources and run concurrently:
webqueries DuckDuckGo HTML results.academicqueries OpenAlex works.discussionsqueries Hacker News Algolia results.
Each selected provider contributes up to 5 results in speed mode and up to 10
in balanced or quality mode before URL deduplication and relevance ranking.
SearchQuery::with_engines selects the local provider names
duckduckgo/ddg, openalex, and hacker-news/hn; source categories remain
the convenient default selection. SearchResponse::provider_status preserves
success and failure information when one provider is unavailable.
SearchMode limits provider selection and retrieval as well as the returned
results and sources: speed selects one provider and returns at most 5
results, balanced selects up to two and returns at most 10, and quality selects
all requested providers and returns at most 20. number_of_results is not
changed by that cap.
SafeSearch supports levels 0 through 4. Remote requests use the SearxNG
values; local DuckDuckGo requests use DuckDuckGo's native values and date-range
codes. Configure local level-3/4 filtering with
SearchConfig::with_blocklist and with_allowlist. The response exposes the
applied level and filtered/disallowed flags in SearchFilters. Local
responses use a bounded in-memory TTL cache by default; configure it with
with_cache or call clear_cache on the client.
The local backend is a direct in-process adapter. It does not start an HTTP
listener, expose a Websurfx server, spawn a search subprocess, or read
Websurfx configuration or assets. The dependency-free Websurfx compatibility
types and URL builder are available for hosts that already run Websurfx; they
do not embed Websurfx or add its AGPL dependency tree. Use
SearchClient::search_websurfx when a configured endpoint is a Websurfx
server; it maps Websurfx's engine and error metadata into Darash's response
model.
Fetch and extract
Search finds sources; darash::fetch reads them. Every fetch returns a
FetchReport with the status, final URL, redirect
flag, elapsed time, content type, size, and body — never silent, even for
error statuses. Bodies are capped at 2 MiB.
use fetch;
async
apply_budget cuts a list of items at item boundaries to fit an estimated
token budget, always keeping at least one item. Local files work through the
same pipeline with fetch::read_source.
CLI
The darash binary ships with the crate (cargo install darash). Search
starts the in-process backend by default; pass --url only when using another
SearxNG-compatible endpoint:
fetch is the research half of the CLI. Without extraction flags it prints the
full report as JSON; with a mode it prints extracted data on stdout and a
one-line report on stderr:
Extraction output is capped at 50 items by default (--limit) and can be
further bounded with --budget (estimated tokens, cut at item boundaries);
omissions are announced on stderr, never silent. --json wraps any mode as
{"data": …, "meta": {status, ok, url, ms, bytes, count, omitted}}.
AI synthesis remains a host responsibility; no MCP server is needed for this in-process tool.
Use SearchConfig when the endpoint needs a custom timeout:
use Duration;
use ;
Limits and errors
- Requests time out after 15 seconds by default; configure this with
SearchConfig::with_timeout. - Response bodies are capped at 256 KiB, including streamed responses.
- Queries must contain non-whitespace text, and page numbers start at 1.
- Endpoints must use
httporhttpsand cannot contain embedded credentials. - Redirects are disabled. Point the client at the final SearxNG endpoint.
Errordistinguishes invalid configuration, request failures, non-success HTTP responses, oversized or invalid responses, and JSON decode failures.
Native quality checks
Run these commands from the crate root:
Darash is licensed under the ISC license.