Search Quality Benchmarks
EchOS uses a multi-stage hybrid search pipeline. This page documents how each stage performs and why the full pipeline outperforms simpler single-mode approaches.The Pipeline
Each search request passes through up to four stages:Methodology
Benchmarks use a synthetic corpus of notes spanning multiple content types (article, note, highlight, conversation) with controlled topic overlap to stress-test disambiguation. Three corpus sizes are tested:
Query types tested (50+ queries total):
- Keyword — exact term match, tests FTS precision
- Semantic — paraphrased queries where the query words don’t appear in the note
- Multi-hop — queries that require combining information from multiple related notes
- Temporal — queries where recency matters (“what did I capture last week about X”)
- Needle-in-haystack — single highly-specific note in a large corpus
- Precision@5 — fraction of the top 5 results that are relevant
- Recall@10 — fraction of all relevant notes appearing in the top 10
- MRR — Mean Reciprocal Rank (position of the first correct result)
- Median latency — wall-clock time at the medium corpus size
Results
Medium corpus (1 000 notes) — overall
By query type — medium corpus, hybrid+decay+hotness
By corpus size — hybrid+decay+hotness
Key findings
Hybrid always beats single-mode. Reciprocal rank fusion consistently outperforms either FTS or vector search alone by 12–18 Precision@5 points. FTS wins on exact-match keyword queries; semantic search wins on paraphrased or concept-driven queries. Neither alone covers both. Temporal decay gives the biggest quality-per-cost improvement. A two-point Precision@5 gain at zero additional latency. Particularly effective for temporal queries — notes from the past week rank above equivalent older notes, which matches user intent. Hotness boost is subtle but real. A further 2 Precision@5 points on average, driven by notes the user has actually found useful before surfacing sooner. The sigmoid saturation means no single note dominates. Reranking is the highest-quality option, but costs an API call. 9 Precision@5 points over the base hybrid pipeline, and 5 points better on multi-hop queries specifically. Latency jumps from ~35 ms to ~1.4 s. Worth enabling for deliberate research queries; leave off for quick lookups. Needle-in-haystack is the hardest case across all configurations. A single specific note in 10,000 is difficult to retrieve reliably without exact keywords. Reranking helps most here (from 0.52 to 0.69 P@5 on medium corpus), but it is not fully solved.Limitations
The benchmark uses a synthetic corpus. Real knowledge bases differ in several ways that can affect results:- Personal writing style — real notes from a single author are more coherent than synthetic diversity; semantic search typically performs better than these numbers suggest.
- Query distribution — your actual queries may be more keyword-heavy or more semantic. Check which query types match your usage pattern.
- Corpus structure — if you import heavily from a single source (e.g. all your highlights from one book), note similarity is higher and disambiguation becomes harder.
- Hotness requires warm data — the hotness boost is zero on fresh imports. It only improves as you actually use search over time.
How to reproduce
The benchmark requires a local environment with the full development stack running.Reranking benchmarks use Claude Haiku to score candidates. Running the full suite including rerank makes ~50 API calls and costs approximately 0.05 at current Haiku pricing.