Realtime audio in/out
A voice agent stack.
Voice agent · 100,000 req/mo. At the default case — 100k requests a month, no further constraints — drawn from the 112 tools already in the index. Each pick is a tool whose own skip-when is stated, because a stack built on tools it does not know how to overcommit to is a guess.
This is the default case, stated as such. Describe your case to change it — the recommendation recomputes, and anything it assumes comes back in the answer rather than staying silent.
The picks
Routing — Bifrost
Strong fitA high-throughput OpenAI-compatible proxy with a small footprint.
Watch out: You need the breadth of a full routing stack.
Consider Cloudflare AI Gateway instead when edge caching and latency dominate, on Cloudflare already
Inference — llama.cpp
Strong fitThe constraint is hardware, not throughput — no GPU, or an edge box.
Watch out: You need maximum concurrent throughput on datacentre GPUs.
Consider Marlin instead when near-GPU throughput from 4-bit weights, as a kernel inside another runtime
Workflows — Apache Airflow
Strong fitScheduled batch data engineering, where DAGs are the shared language.
Watch out: You are building a latency-sensitive product — it was not designed for one.
Consider Restate instead when durable execution with a genuinely low-latency stateful API surface
Evals — Helicone
Strong fitGateway-level observability with per-request cost and latency.
Watch out: You need eval primitives more than traffic logs.
Consider Arize Phoenix instead when openTelemetry-native evaluation you can run yourself
The costs and the risks
- Estimated cost
- $40–$220 per month — a heuristic band from query volume, not a vendor quote.
- Confidence
- 75% — lower when constraints narrow the field.
Biggest risk at the default case: An agent loop that cannot survive a deploy will eventually cause real damage. Durability and tool permissions are the controls, not the prompt. Lattice has no dedicated voice layer — inference latency and streaming orchestration dominate, and this pick is a starting proxy, not a verdict.
A starting point, not a prescription. The recommendation at the default case is what the engine would tell a stranger with your workload; the recommendation at *your* case is what it tells you once it knows the constraints. The Stack Builder holds both — the builder link preselects this workload and starts with your volumes from here.