Band III · Control
When it fails, it fails like “unreliable, unmeasured”
Band III is layers Agent Frameworks, Workflow Orchestration, Guardrails & Safety, Prompt Engineering and Evaluation & Observability. 53 tools sit here, and the failure mode is recognisable before you have read any of them.
Start from the symptom instead
Each is an ordered checklist through the stack, cheapest check first.
Who ends up owning it
A band is a property of the stack; a role is a property of a person. They correlate but are not the same question — which is why the count is derived rather than asserted.
The layers
05Agent Frameworks
13 toolsTurns a model call into a multi-step program.. Orchestration layers for tool-using, multi-step systems.
- LangChainYou want a thin, legible core — this is a large surface.
- LlamaIndexYou need agent orchestration more than data access.
- Pydantic AIYou want a batteries-included ecosystem.
- AutoGenYou need something stable and documented for production.
- CrewAIYou want direct control over the loop itself.
- Semantic KernelYour stack is not .NET.
- MastraYou are working in Python.
- OpenAI Agents SDKYou need multi-provider abstraction at every layer.
- AgnoYou need the wider ecosystem around it.
- Claude Agent SDKYou need provider independence at the framework layer.
- LettaYou want memory to be a detail the framework handles for you.
- smolagentsYou need a rich tool ecosystem around it.
- Vercel AI SDKYou need agent orchestration rather than model calls.
06Workflow Orchestration
10 toolsSurvives the retries, the waits and the crashes.. Durable execution for long-running, retryable pipelines.
- TemporalYou need it live in a day — the adoption cost is real.
- InngestYou are not working in TypeScript.
- Trigger.devYou need more than TypeScript.
- DagsterYou want a task queue rather than a lineage-aware orchestrator.
- PrefectYou need deep lineage modelling.
- Apache AirflowYou are building a latency-sensitive product — it was not designed for one.
- RestateYou need the largest ecosystem around the engine.
- DBOSYou want a separate service to operate.
- Fly MachinesYou want to own the infrastructure layer.
- ModalData locality rules forbid ephemeral compute.
07Guardrails & Safety
8 toolsStops the bad input before it, and the bad output after.. Filtering input, output and model behavior.
- NeMo GuardrailsYou need deep semantic moderation — it is not a classifier.
- Guardrails AIYou need free-form moderation rather than structural checks.
- Llama GuardYou need domain-specific policy enforcement.
- Microsoft PresidioYour data is already pseudonymised at source.
- Lakera GuardRequests cannot leave your network.
- garakYou have no remediation path for whatever it finds.
- PyRITYou are not authorised to test the target.
- Invariant GuardrailsSelf-hosting is a hard requirement.
08Prompt Engineering
9 toolsMakes the prompt an artifact you can diff and test.. Treating prompts as versioned, testable artifacts.
- DSPyYou have no examples and no way to score them.
- InstructorYou need a hard token-level guarantee rather than validation.
- HumanloopPrompts or data cannot be sent to a vendor.
- GuidanceYou need better ergonomics for plain extraction.
- PromptLayerYou need self-hosting.
- AgentaYou need eval depth rather than prompt ergonomics.
- PromptwatchYou need self-hosting.
- TextGradOptimising without an eval set is guessing.
- OutlinesPost-hoc validation is enough for your use case.
09Evaluation & Observability
13 toolsThe only way to know whether any of the above works.. Traces, datasets and graders for non-deterministic output.
- LangSmithTraces cannot leave your infrastructure.
- BraintrustYou need open formats and self-hosting.
- Arize PhoenixYou want a vendor support contract behind it.
- LangfuseYou want a fully managed product with a support contract.
- promptfooYou need a managed UI rather than a test runner.
- DeepEvalYou need a platform rather than a library.
- HeliconeYou need eval primitives more than traffic logs.
- Weights & BiasesYou need an open, exportable format.
- OpenTelemetryYou want a product with a user interface.
- VellumSelf-hosting or reproducibility is a requirement.
- GentraceYou want eval depth more than traffic visibility.
- OpikYou need enterprise support behind it.
- Evidently AIYou only need LLM-specific tracing.
Indexed elsewhere, but serving this band
The problem is in band III and the tool you need is filed under a different layer. These are the crossings, each with the reason it belongs here too — a bare cross-reference reads as a mistake, so the reason is the whole payload.
- Weights & BiasesEvaluation & ObservabilitySweeps and artifact tracking are the orchestration of many training runs, which is the layer 6 job.
- PortkeyRouting & GatewaysIts guardrails run inline in the request path, which is where safety checks belong rather than after the response.
- TrueFoundryRouting & GatewaysIts guardrails enforce in the call path alongside the gateway rather than as a separate service.
- OpenAI Agents SDKAgent FrameworksGuardrails are a first-party primitive on the runner rather than a proxy in front of it.
- LangfuseEvaluation & ObservabilityIts prompt registry versions prompts and ties each version to the traces and scores it produced.
- VellumEvaluation & ObservabilityPrompts are versioned artifacts here, diffed and promoted rather than edited in place.
- TrueFoundryRouting & GatewaysIts tracing and dashboards are the observability half of the same deployment, not an add-on.
- Weights & Biases LaunchFine-tuning & TrainingA training run is an experiment, scored against a baseline the same way an eval set scores a prompt.
- MastraAgent FrameworksIts eval harness ships inside the framework, so grading an agent is part of running it rather than a separate tool.
Bands are a shortcut, not a taxonomy — band III covers 5 of nine layers, so the band tells you where to look without telling you how much of the stack is involved. The method says where else this index is wrong.