Tell me what to use
Build my AI stack.
Describe your case — or start from a real one below. The stack updates as you answer, every answer becomes part of a shareable link, and the decision copies out as a report. Drawn from the 112 tools already in the index; estimates are heuristic bands, not vendor quotes.
Start from a real case
Your stack — updates as you answer
RAG application · 0 req/mo
- RoutingWorkableBifrost
Why for you: A high-throughput OpenAI-compatible proxy with a small footprint.
Watch out: You need the breadth of a full routing stack.
Consider Cloudflare AI Gateway when edge caching and latency dominate, on Cloudflare already.
- RetrievalStrong fitElasticsearch
Why for you: One index carrying BM25 and dense vectors, with ranking you can tune.
Watch out: You need a permissive licence — the Elastic licence is not OSI-approved.
Consider pgvector when under a few million vectors, or when retrieval joins rows you already have.
- InferenceWorkableKoboldCpp
Why for you: Local generation with a UI tuned for interactive, long-form use.
Watch out: You need headless serving at scale.
Consider llama.cpp when the constraint is hardware, not throughput — no GPU, or an edge box.
- PromptsStrong fitAgenta
Why for you: Open-source prompt management with a playground and versioning.
Watch out: You need eval depth rather than prompt ergonomics.
Consider DSPy when compiling prompts into optimised programs from labelled examples.
- EvalsWorkableArize Phoenix
Why for you: OpenTelemetry-native evaluation you can run yourself.
Watch out: You want a vendor support contract behind it.
Consider Braintrust when making evaluation, rather than tracing, the primary workflow.
Estimated cost
$0–$65/mo
Driven by Elasticsearch — the usage-billed picks. Heuristic band, not a quote.
Confidence
87%
Lower when constraints narrow the field.
Biggest risk for your case
Retrieval quality decides the outcome — most failures are ranking or chunking failures, not model failures. Build the eval set before tuning anything.
Every why and watch-out comes from the tool's own use-when / skip-when. Read the method →