One call, maybe a system prompt
A ai chatbot stack.
AI chatbot · 100,000 req/mo. At the default case — 100k requests a month, no further constraints — drawn from the 112 tools already in the index. Each pick is a tool whose own skip-when is stated, because a stack built on tools it does not know how to overcommit to is a guess.
This is the default case, stated as such. Describe your case to change it — the recommendation recomputes, and anything it assumes comes back in the answer rather than staying silent.
The picks
Routing — Not Diamond
Strong fitRouting and prompt optimisation tuned on your own traffic.
Watch out: Self-hosting or auditability is a requirement.
Consider Portkey instead when a managed gateway with guardrails and analytics already bundled
Inference — KoboldCpp
WorkableLocal generation with a UI tuned for interactive, long-form use.
Watch out: You need headless serving at scale.
Consider llama.cpp instead when the constraint is hardware, not throughput — no GPU, or an edge box
Prompts — Instructor
Strong fitSchema-validated extraction with typed retry and partial streaming.
Watch out: You need a hard token-level guarantee rather than validation.
Consider Agenta instead when open-source prompt management with a playground and versioning
Guardrails — Guardrails AI
Strong fitValidators that check output against a schema you define.
Watch out: You need free-form moderation rather than structural checks.
Consider Llama Guard instead when one classifier covering both prompt and response moderation
Evals — Langfuse
Strong fitSelf-hostable tracing, prompts and evals in one place.
Watch out: You want a fully managed product with a support contract.
Consider Arize Phoenix instead when openTelemetry-native evaluation you can run yourself
The costs and the risks
- Estimated cost
- $40–$238 per month — a heuristic band from query volume, not a vendor quote.
- Confidence
- 87% — lower when constraints narrow the field.
- Cost drivers
- Not Diamond
Biggest risk at the default case: Prompt changes without evals are guesses. Version prompts and measure before routing or self-hosting.
A starting point, not a prescription. The recommendation at the default case is what the engine would tell a stranger with your workload; the recommendation at *your* case is what it tells you once it knows the constraints. The Stack Builder holds both — the builder link preselects this workload and starts with your volumes from here.