Adapt open weights to my data
A fine-tuned model stack.
Fine-tuned model · 100,000 req/mo. At the default case — 100k requests a month, no further constraints — drawn from the 112 tools already in the index. Each pick is a tool whose own skip-when is stated, because a stack built on tools it does not know how to overcommit to is a guess.
This is the default case, stated as such. Describe your case to change it — the recommendation recomputes, and anything it assumes comes back in the answer rather than staying silent.
The picks
Training — Unsloth
Strong fitLoRA fine-tuning on a single GPU, with memory and time roughly halved.
Watch out: You need a training stack you can audit line by line.
Consider Megatron-LM instead when tensor and pipeline parallel pretraining at cluster scale
Inference — Marlin
Strong fitNear-GPU throughput from 4-bit weights, as a kernel inside another runtime.
Watch out: You need a scheduler and a server as well as a kernel.
Consider llama.cpp instead when the constraint is hardware, not throughput — no GPU, or an edge box
Prompts — Agenta
WorkableOpen-source prompt management with a playground and versioning.
Watch out: You need eval depth rather than prompt ergonomics.
Consider DSPy instead when compiling prompts into optimised programs from labelled examples
Evals — Arize Phoenix
WorkableOpenTelemetry-native evaluation you can run yourself.
Watch out: You want a vendor support contract behind it.
Consider Braintrust instead when making evaluation, rather than tracing, the primary workflow
The costs and the risks
- Estimated cost
- $40–$220 per month — a heuristic band from query volume, not a vendor quote.
- Confidence
- 87% — lower when constraints narrow the field.
Biggest risk at the default case: Fine-tuning teaches behaviour, not facts. If the answer changes monthly, it belongs in retrieval instead.
A starting point, not a prescription. The recommendation at the default case is what the engine would tell a stranger with your workload; the recommendation at *your* case is what it tells you once it knows the constraints. The Stack Builder holds both — the builder link preselects this workload and starts with your volumes from here.