Lattice
Skip to content

Hybrid retrieval at scale

A ai search stack.

AI search · 100,000 req/mo. At the default case — 100k requests a month, no further constraints — drawn from the 112 tools already in the index. Each pick is a tool whose own skip-when is stated, because a stack built on tools it does not know how to overcommit to is a guess.

This is the default case, stated as such. Describe your case to change it — the recommendation recomputes, and anything it assumes comes back in the answer rather than staying silent.

The picks

  • One index carrying BM25 and dense vectors, with ranking you can tune.

    Watch out: You need a permissive licence — the Elastic licence is not OSI-approved.

    Consider Turbopuffer instead when high-recall retrieval at volume, without operating the index yourself

  • Local generation with a UI tuned for interactive, long-form use.

    Watch out: You need headless serving at scale.

    Consider llama.cpp instead when the constraint is hardware, not throughput — no GPU, or an edge box

  • A high-throughput OpenAI-compatible proxy with a small footprint.

    Watch out: You need the breadth of a full routing stack.

    Consider Cloudflare AI Gateway instead when edge caching and latency dominate, on Cloudflare already

  • OpenTelemetry-native evaluation you can run yourself.

    Watch out: You want a vendor support contract behind it.

    Consider Braintrust instead when making evaluation, rather than tracing, the primary workflow

The costs and the risks

Estimated cost
$40–$238 per month — a heuristic band from query volume, not a vendor quote.
Confidence
87% — lower when constraints narrow the field.
Cost drivers
Elasticsearch

Biggest risk at the default case: Retrieval quality decides the outcome — most failures are ranking or chunking failures, not model failures. Build the eval set before tuning anything.

A starting point, not a prescription. The recommendation at the default case is what the engine would tell a stranger with your workload; the recommendation at *your* case is what it tells you once it knows the constraints. The Stack Builder holds both — the builder link preselects this workload and starts with your volumes from here.