# Lattice — full index > A curated index of AI infrastructure tools, ordered by stack layer — every tool with when to use it and when to skip it. 112 tools across 10 sections. Facts verified 2026-09. Source: https://lattice.kkshah2005.workers.dev — cite tool pages as https://lattice.kkshah2005.workers.dev/
/. Listed tools belong to their respective authors. ## Inference & Serving https://lattice.kkshah2005.workers.dev/inference-serving Runtimes that turn weights into tokens per second. Turns weights into throughput. ### vLLM vLLM is a self-hosted runtime in the Inference & Serving layer of the AI stack. Paged-attention inference engine with an OpenAI-compatible server. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/vllm - Official site: https://vllm.ai - Kind: runtime - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: General GPU serving under bursty, high-concurrency traffic. - Skip it when: Traffic is steady and low-concurrency, so batches never fill. - Owned by: AI Infrastructure - Alternatives: SGLang, Text Generation Inference, Triton Inference Server ### SGLang SGLang is a self-hosted runtime in the Inference & Serving layer of the AI stack. Structured generation runtime built on RadixAttention prefix reuse. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/sglang - Official site: https://github.com/sgl-project/sglang - Kind: runtime - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: A large share of requests share a long prefix, or you need constrained decoding. - Skip it when: You want the fewest moving parts and vLLM already clears your bar. - Owned by: AI Infrastructure - Alternatives: vLLM, Text Generation Inference ### llama.cpp llama.cpp is a self-hosted runtime in the Inference & Serving layer of the AI stack. Portable CPU/GPU inference for quantized GGUF models. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/llama-cpp - Official site: https://github.com/ggml-org/llama.cpp - Kind: runtime - Deployment: self-hosted - Licence: MIT (open-source) - Language: C++ - Cost: free - Use it when: The constraint is hardware, not throughput — no GPU, or an edge box. - Skip it when: You need maximum concurrent throughput on datacentre GPUs. - Owned by: AI Infrastructure - Alternatives: Ollama, llamafile, KoboldCpp ### Ollama Ollama is a self-hosted runtime in the Inference & Serving layer of the AI stack. Local model runner with a single-binary distribution and HTTP API. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/ollama - Official site: https://ollama.com - Kind: runtime - Deployment: self-hosted - Licence: MIT (open-source) - Language: Go - Cost: free - Use it when: Local model access for a developer or small team, with no setup. - Skip it when: You need production throughput or a stable server API. - Owned by: AI Infrastructure, Applied Engineering - Alternatives: llama.cpp, LM Studio, llamafile ### TensorRT-LLM TensorRT-LLM is a self-hosted runtime in the Inference & Serving layer of the AI stack. NVIDIA-optimized inference with inflight batching and FP8. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/tensorrt-llm - Official site: https://github.com/NVIDIA/TensorRT-LLM - Kind: runtime - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: CUDA - Cost: free - Use it when: You are committed to NVIDIA hardware and need the last of the throughput. - Skip it when: Hardware portability matters, or you have no GPUs to tune against. - Owned by: AI Infrastructure - Alternatives: vLLM, SGLang ### Text Generation Inference Text Generation Inference is a self-hosted runtime in the Inference & Serving layer of the AI stack. Hugging Face production server for LLMs with tensor parallelism. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/text-generation-inference - Official site: https://github.com/huggingface/text-generation-inference - Kind: runtime - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Rust - Cost: free - Use it when: Multi-GPU tensor parallelism inside an existing Hugging Face stack. - Skip it when: The model fits on one GPU, where this is overkill. - Owned by: AI Infrastructure - Alternatives: vLLM, SGLang ### LM Studio LM Studio is a self-hosted service in the Inference & Serving layer of the AI stack. Desktop app for running and serving local models with an OpenAI API. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/lm-studio - Official site: https://lmstudio.ai - Kind: service - Deployment: self-hosted - Licence: proprietary (proprietary) - Language: TypeScript - Cost: free - Use it when: A desktop GUI for running and serving local models, aimed at one operator. - Skip it when: You need headless, multi-tenant deployment. - Owned by: AI Infrastructure, Applied Engineering - Alternatives: Ollama, llama.cpp ### Triton Inference Server Triton Inference Server is a self-hosted runtime in the Inference & Serving layer of the AI stack. Multi-framework inference server for ONNX, TensorRT and Python. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/triton-inference-server - Official site: https://github.com/triton-inference-server/server - Kind: runtime - Deployment: self-hosted - Licence: BSD-3-Clause (open-source) - Language: C++ - Cost: free - Use it when: One endpoint serving ONNX, TensorRT and Python models side by side. - Skip it when: You serve one model family and want peak LLM throughput. - Owned by: AI Infrastructure - Alternatives: vLLM, Text Generation Inference ### llamafile llamafile is a self-hosted runtime in the Inference & Serving layer of the AI stack. Packages a model and its runtime into a single executable. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/llamafile - Official site: https://github.com/Mozilla-Ocho/llamafile - Kind: runtime - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: C++ - Cost: free - Use it when: Shipping a model as a single executable with no install step. - Skip it when: You need batching, multi-GPU serving or a production HTTP server. - Owned by: AI Infrastructure - Alternatives: llama.cpp, Ollama ### KoboldCpp KoboldCpp is a self-hosted runtime in the Inference & Serving layer of the AI stack. GGUF inference with a focused text-adventure UI and API. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/koboldcpp - Official site: https://github.com/LostRuins/koboldcpp - Kind: runtime - Deployment: self-hosted - Licence: MIT (open-source) - Language: C++ - Cost: free - Use it when: Local generation with a UI tuned for interactive, long-form use. - Skip it when: You need headless serving at scale. - Owned by: AI Infrastructure - Alternatives: llama.cpp, Ollama ### PowerInfer PowerInfer is a self-hosted runtime in the Inference & Serving layer of the AI stack. CPU-first runtime that offloads hot paths to the GPU. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/powerinfer - Official site: https://github.com/AmateurChina/PowerInfer - Kind: runtime - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: C++ - Cost: free - Use it when: CPU-first serving where the GPU accelerates rather than carries the load. - Skip it when: You have GPUs to spare — it trades peak throughput for reach. - Owned by: AI Infrastructure - Alternatives: llama.cpp, vLLM ### Marlin Marlin is a self-hosted library in the Inference & Serving layer of the AI stack. Quantised GEMM kernels for near-GPU speed at 4-bit. - Page: https://lattice.kkshah2005.workers.dev/inference-serving/marlin - Official site: https://github.com/vllm-project/marlin - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: CUDA - Cost: free - Use it when: Near-GPU throughput from 4-bit weights, as a kernel inside another runtime. - Skip it when: You need a scheduler and a server as well as a kernel. - Owned by: AI Infrastructure ## Routing & Gateways https://lattice.kkshah2005.workers.dev/routing-gateways One endpoint across many providers, with policy in between. Decides which model answers, and what it costs. ### LiteLLM LiteLLM is a self-hosted library in the Routing & Gateways layer of the AI stack. OpenAI-format proxy translating across 100+ model providers. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/litellm - Official site: https://litellm.ai - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: The widest provider coverage, deployed the least committal way. - Skip it when: You want routing policy to live in a SaaS console. - Owned by: ML Platform - Alternatives: Portkey, Cloudflare AI Gateway, Envoy AI Gateway ### Portkey Portkey is a SaaS platform in the Routing & Gateways layer of the AI stack. AI gateway with routing, caching and guardrails in one layer. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/portkey - Official site: https://portkey.ai - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: A managed gateway with guardrails and analytics already bundled. - Skip it when: Prompts cannot leave your infrastructure. - Owned by: ML Platform - Alternatives: LiteLLM, Cloudflare AI Gateway - Also in guardrails-safety: Its guardrails run inline in the request path, which is where safety checks belong rather than after the response. ### Cloudflare AI Gateway Cloudflare AI Gateway is a SaaS service in the Routing & Gateways layer of the AI stack. Edge gateway adding caching, retries and rate limiting. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/cloudflare-ai-gateway - Official site: https://developers.cloudflare.com - Kind: service - Deployment: saas - Licence: proprietary (proprietary) - Cost: free-tier - Use it when: Edge caching and latency dominate, on Cloudflare already. - Skip it when: Requests must not transit a third party's network. - Owned by: ML Platform - Alternatives: Portkey, Envoy AI Gateway ### OpenRouter OpenRouter is a SaaS service in the Routing & Gateways layer of the AI stack. Single API key and billing across many hosted models. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/openrouter - Official site: https://openrouter.ai - Kind: service - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: One key and one bill across many hosted models, while experimenting. - Skip it when: You need self-hosting or per-provider SLAs. - Owned by: ML Platform - Alternatives: LiteLLM, Martian ### Martian Martian is a SaaS platform in the Routing & Gateways layer of the AI stack. Model router that optimizes for quality, cost and latency. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/martian - Official site: https://withmartian.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Routing tuned on your own traffic rather than a public benchmark. - Skip it when: You want the routing rule to be plain config you can read. - Owned by: ML Platform - Alternatives: RouteLLM, Not Diamond, OpenRouter ### Envoy AI Gateway Envoy AI Gateway is a self-hosted runtime in the Routing & Gateways layer of the AI stack. CNCF-track gateway for LLM traffic on an Envoy data plane. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/envoy-ai-gateway - Official site: https://aigateway.envoyproxy.io - Kind: runtime - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Go - Cost: free - Use it when: A Kubernetes shop that already runs Envoy and needs the gateway in-cluster. - Skip it when: You want a gateway configurable without a data plane. - Owned by: ML Platform, AI Infrastructure - Alternatives: LiteLLM, Cloudflare AI Gateway ### Bifrost Bifrost is a self-hosted library in the Routing & Gateways layer of the AI stack. High-throughput LLM gateway with drop-in OpenAI compatibility. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/bifrost - Official site: https://getmaxim.ai - Kind: library - Deployment: self-hosted - Licence: proprietary (proprietary) - Language: Go - Cost: free - Use it when: A high-throughput OpenAI-compatible proxy with a small footprint. - Skip it when: You need the breadth of a full routing stack. - Owned by: ML Platform - Alternatives: LiteLLM ### RouteLLM RouteLLM is a self-hosted library in the Routing & Gateways layer of the AI stack. Learned router that cuts cost by matching difficulty to model size. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/routellm - Official site: https://github.com/lm-sys/RouteLLM - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Cutting cost by matching query difficulty to model size, with a learned router. - Skip it when: You cannot evaluate the quality you would be trading away. - Owned by: ML Platform, Applied Engineering - Alternatives: Martian, Not Diamond ### Not Diamond Not Diamond is a SaaS platform in the Routing & Gateways layer of the AI stack. Routing and prompt optimisation tuned on your own traffic. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/not-diamond - Official site: https://notdiamond.ai - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Routing and prompt optimisation tuned on your own traffic. - Skip it when: Self-hosting or auditability is a requirement. - Owned by: ML Platform, Applied Engineering - Alternatives: Martian, RouteLLM ### TrueFoundry TrueFoundry is a managed platform in the Routing & Gateways layer of the AI stack. Gateway plus observability and guardrails as one deployment. - Page: https://lattice.kkshah2005.workers.dev/routing-gateways/truefoundry - Official site: https://truefoundry.com - Kind: platform - Deployment: managed - Licence: proprietary (proprietary) - Cost: subscription - Use it when: Gateway, observability and guardrails as one managed deployment. - Skip it when: You want to adopt the pieces independently. - Owned by: ML Platform, Production & Governance - Alternatives: Portkey, Braintrust - Also in evaluation-observability: Its tracing and dashboards are the observability half of the same deployment, not an add-on. - Also in guardrails-safety: Its guardrails enforce in the call path alongside the gateway rather than as a separate service. ## Retrieval & Vector Stores https://lattice.kkshah2005.workers.dev/retrieval-vector-stores Where embeddings live, and how they get retrieved. Holds the embeddings, returns the few that matter. ### pgvector pgvector is a self-hosted database in the Retrieval & Vector Stores layer of the AI stack. Vector similarity search as a Postgres extension. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/pgvector - Official site: https://github.com/pgvector/pgvector - Kind: database - Deployment: self-hosted - Licence: PostgreSQL (open-source) - Language: SQL - Cost: free - Use it when: Under a few million vectors, or when retrieval joins rows you already have. - Skip it when: Vector search has become the workload rather than a side feature. - Owned by: Data & Retrieval - Alternatives: Qdrant, Milvus, Weaviate ### Qdrant Qdrant is a self-hosted database in the Retrieval & Vector Stores layer of the AI stack. Rust vector database with rich filtering payloads. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/qdrant - Official site: https://qdrant.tech - Kind: database - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Rust - Cost: free - Use it when: Vector-first workloads with heavy metadata filtering at scale. - Skip it when: You want no new datastore to operate. - Owned by: Data & Retrieval - Alternatives: Weaviate, Milvus, pgvector ### Weaviate Weaviate is a self-hosted database in the Retrieval & Vector Stores layer of the AI stack. Graph-aware vector database with hybrid and multi-vector search. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/weaviate - Official site: https://weaviate.io - Kind: database - Deployment: self-hosted - Licence: BSD-3-Clause (open-source) - Language: Go - Cost: free - Use it when: Hybrid or multi-vector search with a graph-shaped data model. - Skip it when: Your data is genuinely relational and joins matter more than vectors. - Owned by: Data & Retrieval - Alternatives: Qdrant, Elasticsearch, Vespa ### Chroma Chroma is a self-hosted database in the Retrieval & Vector Stores layer of the AI stack. Embeddings database designed for fast prototyping. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/chroma - Official site: https://trychroma.com - Kind: database - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Getting a prototype working in an afternoon. - Skip it when: You need durability, scaling or concurrent writes. - Owned by: Data & Retrieval, Applied Engineering - Alternatives: Qdrant, pgvector ### Pinecone Pinecone is a SaaS database in the Retrieval & Vector Stores layer of the AI stack. Fully managed vector database with serverless scaling. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/pinecone - Official site: https://pinecone.io - Kind: database - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Zero operational work, and vector search is not your core competence. - Skip it when: Residency, cost predictability or a query language you control. - Owned by: Data & Retrieval - Alternatives: Turbopuffer, Milvus, Qdrant ### Milvus Milvus is a self-hosted database in the Retrieval & Vector Stores layer of the AI stack. Cloud-native vector database supporting billion-scale indexes. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/milvus - Official site: https://milvus.io - Kind: database - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Go - Cost: free - Use it when: Billion-scale indexes with a cloud-native deployment. - Skip it when: You need a small, comprehensible system. - Owned by: Data & Retrieval - Alternatives: Qdrant, Weaviate, pgvector ### Turbopuffer Turbopuffer is a SaaS database in the Retrieval & Vector Stores layer of the AI stack. Vector search engine tuned for high-recall retrieval workloads. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/turbopuffer - Official site: https://turbopuffer.com - Kind: database - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: High-recall retrieval at volume, without operating the index yourself. - Skip it when: You need to tune the index directly. - Owned by: Data & Retrieval - Alternatives: Pinecone, Qdrant ### Unstructured Unstructured is a self-hosted library in the Retrieval & Vector Stores layer of the AI stack. Preprocessing library that partitions raw documents for indexing. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/unstructured - Official site: https://unstructured.io - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Partitioning messy documents into retrievable chunks before indexing. - Skip it when: Your input is already structured. - Owned by: Data & Retrieval - Alternatives: Docling, LlamaParse ### Elasticsearch Elasticsearch is a self-hosted database in the Retrieval & Vector Stores layer of the AI stack. BM25 plus dense vectors in one index, with hybrid ranking. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/elasticsearch - Official site: https://elastic.co - Kind: database - Deployment: self-hosted - Licence: Elastic-License-2.0 (source-available) - Language: Java - Cost: usage-based - Use it when: One index carrying BM25 and dense vectors, with ranking you can tune. - Skip it when: You need a permissive licence — the Elastic licence is not OSI-approved. - Owned by: Data & Retrieval - Alternatives: Vespa, Weaviate, Qdrant ### Vespa Vespa is a self-hosted platform in the Retrieval & Vector Stores layer of the AI stack. Search engine built for ranking-heavy retrieval at scale. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/vespa - Official site: https://vespa.ai - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Java - Cost: free - Use it when: Ranking-heavy retrieval where the ranking function is the product. - Skip it when: You want a small, low-operations component. - Owned by: Data & Retrieval - Alternatives: Elasticsearch, Weaviate ### Rerankers Rerankers is a SaaS service in the Retrieval & Vector Stores layer of the AI stack. Purpose-built cross-encoders for the second retrieval stage. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/rerankers - Official site: https://cohere.com - Kind: service - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: A cross-encoder in the second retrieval stage, to lift precision cheaply. - Skip it when: You have no way to measure whether recall actually improved. - Owned by: Data & Retrieval - Alternatives: Jina AI, Unstructured ### Jina AI Jina AI is a SaaS service in the Retrieval & Vector Stores layer of the AI stack. Multimodal embeddings and a reranker behind one endpoint. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/jina-ai - Official site: https://jina.ai - Kind: service - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Multimodal embeddings and a reranker behind one endpoint. - Skip it when: Embeddings must stay in-house. - Owned by: Data & Retrieval - Alternatives: Rerankers ### LangChain Text Splitters LangChain Text Splitters is a self-hosted library in the Retrieval & Vector Stores layer of the AI stack. Document loaders and splitters for the ingest stage. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/langchain-text-splitters - Official site: https://python.langchain.com - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Document loaders and splitters, if you are already in that ecosystem. - Skip it when: Structure matters — fixed-size splitting is rarely the right boundary. - Owned by: Data & Retrieval, Applied Engineering - Alternatives: Unstructured, Docling ### Docling Docling is a self-hosted library in the Retrieval & Vector Stores layer of the AI stack. Layout-aware parsing that preserves table and heading structure. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/docling - Official site: https://github.com/DS4SD/docling - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Layout-aware parsing that preserves tables and heading structure. - Skip it when: Your documents are plain text. - Owned by: Data & Retrieval - Alternatives: Unstructured, LlamaParse ### LlamaParse LlamaParse is a managed service in the Retrieval & Vector Stores layer of the AI stack. Managed parsing for PDFs, especially the ugly ones. - Page: https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/llamaparse - Official site: https://llamaindex.ai - Kind: service - Deployment: managed - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Managed parsing for PDFs, especially the ugly ones. - Skip it when: Document contents cannot leave your infrastructure. - Owned by: Data & Retrieval - Alternatives: Docling, Unstructured ## Fine-tuning & Training https://lattice.kkshah2005.workers.dev/fine-tuning Adapting open weights to your own data and shape. Produces the weights that layer 1 serves. ### Unsloth Unsloth is a self-hosted library in the Fine-tuning & Training layer of the AI stack. Hand-written kernels that cut LoRA memory and time sharply. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/unsloth - Official site: https://unsloth.ai - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: LoRA fine-tuning on a single GPU, with memory and time roughly halved. - Skip it when: You need a training stack you can audit line by line. - Owned by: Applied Engineering, AI Infrastructure - Alternatives: Axolotl, PEFT ### Axolotl Axolotl is a self-hosted library in the Fine-tuning & Training layer of the AI stack. Configuration-driven fine-tuning across common architectures. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/axolotl - Official site: https://axolotl.ai - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Configuration-driven fine-tuning across many architectures. - Skip it when: You need to hand-tune the training loop itself. - Owned by: Applied Engineering - Alternatives: LLaMA-Factory, Unsloth ### LLaMA-Factory LLaMA-Factory is a self-hosted platform in the Fine-tuning & Training layer of the AI stack. Unified interface for SFT, DPO and RLHF on open models. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/llama-factory - Official site: https://github.com/hiyouga/LLaMA-Factory - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: One interface for SFT, DPO and RLHF across open models. - Skip it when: You want a minimal, readable training script. - Owned by: Applied Engineering - Alternatives: Axolotl, TRL ### PEFT PEFT is a self-hosted library in the Fine-tuning & Training layer of the AI stack. Parameter-efficient fine-tuning methods such as LoRA and QLoRA. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/peft - Official site: https://github.com/huggingface/peft - Kind: library - Deployment: self-hosted - Licence: BSD-3-Clause (open-source) - Language: Python - Cost: free - Use it when: Parameter-efficient fine-tuning — LoRA and variants — on limited hardware. - Skip it when: You need to change the model's capabilities, not bolt on an adapter. - Owned by: Applied Engineering - Alternatives: Unsloth, TRL ### TRL TRL is a self-hosted library in the Fine-tuning & Training layer of the AI stack. Hugging Face library of post-training trainers for SFT, DPO and GRPO. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/trl - Official site: https://github.com/huggingface/trl - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Post-training trainers for SFT, DPO and GRPO on top of any PEFT setup. - Skip it when: You are training from scratch rather than aligning an existing model. - Owned by: Applied Engineering - Alternatives: PEFT, LLaMA-Factory, torchtune ### DeepSpeed DeepSpeed is a self-hosted library in the Fine-tuning & Training layer of the AI stack. ZeRO sharding and pipeline parallelism for large-scale training. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/deepspeed - Official site: https://deepspeed.ai - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: ZeRO sharding and pipeline parallelism across many GPUs. - Skip it when: The model fits on one device — the complexity buys nothing. - Owned by: AI Infrastructure, Applied Engineering - Alternatives: Megatron-LM ### Megatron-LM Megatron-LM is a self-hosted library in the Fine-tuning & Training layer of the AI stack. Tensor and pipeline parallel training for multi-GPU clusters. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/megatron-lm - Official site: https://github.com/NVIDIA/Megatron-LM - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: CUDA - Cost: free - Use it when: Tensor and pipeline parallel pretraining at cluster scale. - Skip it when: You are fine-tuning rather than pretraining. - Owned by: AI Infrastructure - Alternatives: DeepSpeed ### torchtune torchtune is a self-hosted library in the Fine-tuning & Training layer of the AI stack. PyTorch-native recipes for fine-tuning and aligning open models. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/torchtune - Official site: https://github.com/pytorch/torchtune - Kind: library - Deployment: self-hosted - Licence: BSD-3-Clause (open-source) - Language: Python - Cost: free - Use it when: PyTorch-native recipes that you can read and modify. - Skip it when: You want the breadth of a config-driven stack. - Owned by: Applied Engineering - Alternatives: TRL, LLaMA-Factory ### Hugging Face TRL Hugging Face TRL is a SaaS platform in the Fine-tuning & Training layer of the AI stack. Preference and reward optimisation on top of any PEFT setup. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/hugging-face-trl - Official site: https://huggingface.co - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Preference and reward optimisation hosted, without a cluster. - Skip it when: Training data cannot be uploaded to a third party. - Owned by: Applied Engineering - Alternatives: TRL ### Colab Colab is a SaaS platform in the Fine-tuning & Training layer of the AI stack. Hosted GPUs for small runs and quick experiments. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/colab - Official site: https://colab.research.google.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: free-tier - Use it when: A disposable GPU for a quick experiment, with no setup. - Skip it when: Anything reproducible — a notebook is not a training pipeline. - Owned by: Applied Engineering - Alternatives: Replicate ### Replicate Replicate is a SaaS platform in the Fine-tuning & Training layer of the AI stack. Hosted fine-tuning and deployment for open models. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/replicate - Official site: https://replicate.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Hosted fine-tuning and deployment of an open model, billed per run. - Skip it when: You need custom training code or strict residency. - Owned by: Applied Engineering - Alternatives: Colab, Weights & Biases Launch ### Weights & Biases Launch Weights & Biases Launch is a SaaS platform in the Fine-tuning & Training layer of the AI stack. Managed training runs with sweeps and artifact tracking. - Page: https://lattice.kkshah2005.workers.dev/fine-tuning/weights-biases-launch - Official site: https://wandb.ai - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: subscription - Use it when: Managed training runs with sweeps and artifact tracking wired up. - Skip it when: You want training infrastructure you run yourself. - Owned by: Production & Governance, Applied Engineering - Alternatives: Replicate, Weights & Biases - Also in evaluation-observability: A training run is an experiment, scored against a baseline the same way an eval set scores a prompt. ## Agent Frameworks https://lattice.kkshah2005.workers.dev/agent-frameworks Orchestration layers for tool-using, multi-step systems. Turns a model call into a multi-step program. ### LangChain LangChain is a self-hosted framework in the Agent Frameworks layer of the AI stack. Composable abstractions for model calls, tools and state. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/langchain - Official site: https://langchain.com - Kind: framework - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Tool use and state, in an ecosystem where the answers already exist. - Skip it when: You want a thin, legible core — this is a large surface. - Owned by: Applied Engineering - Alternatives: LlamaIndex, Pydantic AI, CrewAI ### LlamaIndex LlamaIndex is a self-hosted framework in the Agent Frameworks layer of the AI stack. Data-centric framework for retrieval and agent workflows. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/llamaindex - Official site: https://llamaindex.ai - Kind: framework - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Data-centric work where retrieval is the centre of the problem. - Skip it when: You need agent orchestration more than data access. - Owned by: Applied Engineering, Data & Retrieval - Alternatives: LangChain, Pydantic AI ### Pydantic AI Pydantic AI is a self-hosted framework in the Agent Frameworks layer of the AI stack. Type-safe agent framework that leans on Pydantic models. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/pydantic-ai - Official site: https://ai.pydantic.dev - Kind: framework - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Type-safe agents where schema correctness is the priority. - Skip it when: You want a batteries-included ecosystem. - Owned by: Applied Engineering - Alternatives: LangChain, smolagents, LlamaIndex ### AutoGen AutoGen is a self-hosted framework in the Agent Frameworks layer of the AI stack. Microsoft research project for conversable multi-agent systems. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/autogen - Official site: https://microsoft.github.io/autogen - Kind: framework - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Research into conversable multi-agent systems. - Skip it when: You need something stable and documented for production. - Owned by: Applied Engineering - Alternatives: CrewAI, LangChain ### CrewAI CrewAI is a self-hosted framework in the Agent Frameworks layer of the AI stack. Role-based orchestration where agents collaborate as a crew. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/crewai - Official site: https://crewai.com - Kind: framework - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Role-based orchestration where agents are framed as a team. - Skip it when: You want direct control over the loop itself. - Owned by: Applied Engineering - Alternatives: AutoGen, LangChain ### Semantic Kernel Semantic Kernel is a self-hosted framework in the Agent Frameworks layer of the AI stack. Microsoft SDK for embedding AI steps into .NET and Python apps. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/semantic-kernel - Official site: https://learn.microsoft.com - Kind: framework - Deployment: self-hosted - Licence: MIT (open-source) - Language: C# - Cost: free - Use it when: Embedding AI steps into an existing .NET estate. - Skip it when: Your stack is not .NET. - Owned by: Applied Engineering - Alternatives: LangChain, Pydantic AI ### Mastra Mastra is a self-hosted framework in the Agent Frameworks layer of the AI stack. TypeScript agent framework with typed workflows and evals. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/mastra - Official site: https://mastra.ai - Kind: framework - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: TypeScript - Cost: free - Use it when: TypeScript agents with typed workflows and evals built in. - Skip it when: You are working in Python. - Owned by: Applied Engineering - Alternatives: Vercel AI SDK, OpenAI Agents SDK - Also in evaluation-observability: Its eval harness ships inside the framework, so grading an agent is part of running it rather than a separate tool. ### OpenAI Agents SDK OpenAI Agents SDK is a self-hosted framework in the Agent Frameworks layer of the AI stack. Lightweight primitives for handoffs, guardrails and tracing. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/openai-agents-sdk - Official site: https://openai.github.io/openai-agents-python - Kind: framework - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: A small, legible set of primitives for handoffs and guardrails. - Skip it when: You need multi-provider abstraction at every layer. - Owned by: Applied Engineering - Alternatives: Claude Agent SDK, smolagents - Also in guardrails-safety: Guardrails are a first-party primitive on the runner rather than a proxy in front of it. ### Agno Agno is a self-hosted framework in the Agent Frameworks layer of the AI stack. Minimal agent runtime centered on model-agnostic tool interfaces. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/agno - Official site: https://agno.com - Kind: framework - Deployment: self-hosted - Licence: MPL-2.0 (open-source) - Language: Python - Cost: free - Use it when: A minimal agent runtime with model-agnostic tool interfaces. - Skip it when: You need the wider ecosystem around it. - Owned by: Applied Engineering - Alternatives: smolagents, Pydantic AI ### Claude Agent SDK Claude Agent SDK is a self-hosted framework in the Agent Frameworks layer of the AI stack. Anthropic's toolkit for building agents with tool use and hooks. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/claude-agent-sdk - Official site: https://docs.anthropic.com - Kind: framework - Deployment: self-hosted - Licence: proprietary (proprietary) - Language: TypeScript - Cost: free - Use it when: Building agents against Anthropic models, with tool use and hooks. - Skip it when: You need provider independence at the framework layer. - Owned by: Applied Engineering - Alternatives: OpenAI Agents SDK, Letta ### Letta Letta is a self-hosted platform in the Agent Frameworks layer of the AI stack. Agent runtime with persistent, editable memory as a first-class primitive. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/letta - Official site: https://letta.com - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Agents whose memory is a first-class, inspectable and editable state. - Skip it when: You want memory to be a detail the framework handles for you. - Owned by: Applied Engineering, Data & Retrieval - Alternatives: smolagents, OpenAI Agents SDK - Also in retrieval-vector-stores: Its memory is persisted and retrieved over an embedding store, so it is a layer-3 store with an agent API on top. ### smolagents smolagents is a self-hosted framework in the Agent Frameworks layer of the AI stack. Minimal code-first agent loop from the Hugging Face team. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/smolagents - Official site: https://huggingface.co - Kind: framework - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Reading the whole agent loop in a page — it is deliberately tiny. - Skip it when: You need a rich tool ecosystem around it. - Owned by: Applied Engineering - Alternatives: Agno, Letta ### Vercel AI SDK Vercel AI SDK is a self-hosted library in the Agent Frameworks layer of the AI stack. Provider-agnostic building blocks for streaming apps and agents. - Page: https://lattice.kkshah2005.workers.dev/agent-frameworks/vercel-ai-sdk - Official site: https://ai-sdk.dev - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: TypeScript - Cost: free - Use it when: Streaming model calls in a TypeScript app, provider-agnostically. - Skip it when: You need agent orchestration rather than model calls. - Owned by: Applied Engineering - Alternatives: Mastra, OpenAI Agents SDK ## Workflow Orchestration https://lattice.kkshah2005.workers.dev/workflow-orchestration Durable execution for long-running, retryable pipelines. Survives the retries, the waits and the crashes. ### Temporal Temporal is a self-hosted platform in the Workflow Orchestration layer of the AI stack. Durable execution engine that survives crashes and long waits. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/temporal - Official site: https://temporal.io - Kind: platform - Deployment: self-hosted - Licence: MIT (open-source) - Language: Go - Cost: free - Use it when: Work spanning minutes or hours that must survive restarts and retries. - Skip it when: You need it live in a day — the adoption cost is real. - Owned by: ML Platform - Alternatives: Inngest, Restate, DBOS ### Inngest Inngest is a SaaS platform in the Workflow Orchestration layer of the AI stack. Event-driven step functions with durable replay for TypeScript. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/inngest - Official site: https://inngest.com - Kind: platform - Deployment: saas - Licence: Apache-2.0 (open-source) - Language: TypeScript - Cost: free-tier - Use it when: Durable step functions in TypeScript, without running the engine. - Skip it when: You are not working in TypeScript. - Owned by: ML Platform - Alternatives: Trigger.dev, Temporal ### Trigger.dev Trigger.dev is a SaaS platform in the Workflow Orchestration layer of the AI stack. Background jobs and AI workflows with long-running task support. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/trigger-dev - Official site: https://trigger.dev - Kind: platform - Deployment: saas - Licence: Apache-2.0 (open-source) - Language: TypeScript - Cost: free-tier - Use it when: Long-running background jobs that must survive a deploy. - Skip it when: You need more than TypeScript. - Owned by: ML Platform - Alternatives: Inngest, Temporal ### Dagster Dagster is a self-hosted platform in the Workflow Orchestration layer of the AI stack. Asset-oriented orchestration for data and ML pipelines. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/dagster - Official site: https://dagster.io - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Data and ML pipelines framed as assets with explicit dependencies. - Skip it when: You want a task queue rather than a lineage-aware orchestrator. - Owned by: ML Platform, Data & Retrieval - Alternatives: Prefect, Apache Airflow ### Prefect Prefect is a self-hosted platform in the Workflow Orchestration layer of the AI stack. Python-native workflow orchestration with a managed cloud option. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/prefect - Official site: https://prefect.io - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free-tier - Use it when: Python-native flows with a managed cloud option behind them. - Skip it when: You need deep lineage modelling. - Owned by: ML Platform, Data & Retrieval - Alternatives: Dagster, Apache Airflow ### Apache Airflow Apache Airflow is a self-hosted platform in the Workflow Orchestration layer of the AI stack. Scheduler and DAG engine long used for batch data engineering. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/apache-airflow - Official site: https://apache.org - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Scheduled batch data engineering, where DAGs are the shared language. - Skip it when: You are building a latency-sensitive product — it was not designed for one. - Owned by: ML Platform, Data & Retrieval - Alternatives: Dagster, Prefect ### Restate Restate is a self-hosted platform in the Workflow Orchestration layer of the AI stack. Durable execution with a low-latency stateful API surface. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/restate - Official site: https://restate.dev - Kind: platform - Deployment: self-hosted - Licence: BSL-1.1 (source-available) - Language: Rust - Cost: free - Use it when: Durable execution with a genuinely low-latency stateful API surface. - Skip it when: You need the largest ecosystem around the engine. - Owned by: ML Platform - Alternatives: Temporal, DBOS ### DBOS DBOS is a self-hosted library in the Workflow Orchestration layer of the AI stack. Durable workflows as ordinary Python functions and decorators. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/dbos - Official site: https://dbos.dev - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Durable workflows as ordinary Python functions and decorators. - Skip it when: You want a separate service to operate. - Owned by: ML Platform - Alternatives: Temporal, Restate ### Fly Machines Fly Machines is a managed platform in the Workflow Orchestration layer of the AI stack. Hosting for stateful containers that fits durable agent workers. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/fly-machines - Official site: https://fly.io - Kind: platform - Deployment: managed - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Hosting stateful containers that fits durable agent workers. - Skip it when: You want to own the infrastructure layer. - Owned by: ML Platform, AI Infrastructure - Alternatives: Modal, Trigger.dev ### Modal Modal is a SaaS platform in the Workflow Orchestration layer of the AI stack. Serverless GPU and container platform popular for batch inference. - Page: https://lattice.kkshah2005.workers.dev/workflow-orchestration/modal - Official site: https://modal.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Serverless GPUs and containers for batch inference and scheduled jobs. - Skip it when: Data locality rules forbid ephemeral compute. - Owned by: ML Platform, AI Infrastructure - Alternatives: Fly Machines, Replicate ## Guardrails & Safety https://lattice.kkshah2005.workers.dev/guardrails-safety Filtering input, output and model behavior. Stops the bad input before it, and the bad output after. ### NeMo Guardrails NeMo Guardrails is a self-hosted library in the Guardrails & Safety layer of the AI stack. Programmable rails that constrain conversational flow. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/nemo-guardrails - Official site: https://github.com/NVIDIA/NeMo-Guardrails - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Constraining conversational flow with programmable rails. - Skip it when: You need deep semantic moderation — it is not a classifier. - Owned by: Production & Governance - Alternatives: Guardrails AI, Llama Guard ### Guardrails AI Guardrails AI is a self-hosted library in the Guardrails & Safety layer of the AI stack. Validators that check model output against a defined schema. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/guardrails-ai - Official site: https://guardrailsai.com - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Validators that check output against a schema you define. - Skip it when: You need free-form moderation rather than structural checks. - Owned by: Production & Governance - Alternatives: NeMo Guardrails, Invariant Guardrails ### Llama Guard Llama Guard is a self-hosted library in the Guardrails & Safety layer of the AI stack. Safety classifier for prompt and response moderation. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/llama-guard - Official site: https://ai.meta.com/llama - Kind: library - Deployment: self-hosted - Licence: Llama-3.1-Community (source-available) - Language: Python - Cost: free - Use it when: One classifier covering both prompt and response moderation. - Skip it when: You need domain-specific policy enforcement. - Owned by: Production & Governance - Alternatives: Lakera Guard, NeMo Guardrails ### Microsoft Presidio Microsoft Presidio is a self-hosted library in the Guardrails & Safety layer of the AI stack. PII detection and anonymization for text and images. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/microsoft-presidio - Official site: https://microsoft.github.io/presidio - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Detecting and anonymising PII in text and images. - Skip it when: Your data is already pseudonymised at source. - Owned by: Production & Governance ### Lakera Guard Lakera Guard is a SaaS service in the Guardrails & Safety layer of the AI stack. Prompt injection and jailbreak detection at the gateway. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/lakera-guard - Official site: https://lakera.ai - Kind: service - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Prompt-injection and jailbreak detection at the gateway. - Skip it when: Requests cannot leave your network. - Owned by: Production & Governance - Alternatives: Llama Guard, Invariant Guardrails ### garak garak is a self-hosted library in the Guardrails & Safety layer of the AI stack. Scanner that probes models for known vulnerability classes. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/garak - Official site: https://github.com/NVIDIA/garak - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Probing your own system for known vulnerability classes. - Skip it when: You have no remediation path for whatever it finds. - Owned by: Production & Governance - Alternatives: PyRIT ### PyRIT PyRIT is a self-hosted library in the Guardrails & Safety layer of the AI stack. Microsoft's red-team tool for generating and scoring attack prompts. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/pyrit - Official site: https://github.com/Azure/PyRIT - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Generating and scoring attack prompts systematically. - Skip it when: You are not authorised to test the target. - Owned by: Production & Governance - Alternatives: garak ### Invariant Guardrails Invariant Guardrails is a managed platform in the Guardrails & Safety layer of the AI stack. Guardrails as code, enforced inline in the call path. - Page: https://lattice.kkshah2005.workers.dev/guardrails-safety/invariant-guardrails - Official site: https://invariantlabs.ai - Kind: platform - Deployment: managed - Licence: proprietary (proprietary) - Language: Python - Cost: usage-based - Use it when: Guardrails as code, enforced inline in the call path. - Skip it when: Self-hosting is a hard requirement. - Owned by: Production & Governance - Alternatives: Guardrails AI, Lakera Guard ## Prompt Engineering https://lattice.kkshah2005.workers.dev/prompt-engineering Treating prompts as versioned, testable artifacts. Makes the prompt an artifact you can diff and test. ### DSPy DSPy is a self-hosted library in the Prompt Engineering layer of the AI stack. Declarative prompting that compiles to optimized programs from examples. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/dspy - Official site: https://dspy.ai - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Compiling prompts into optimised programs from labelled examples. - Skip it when: You have no examples and no way to score them. - Owned by: Applied Engineering - Alternatives: TextGrad, PromptLayer ### Instructor Instructor is a self-hosted library in the Prompt Engineering layer of the AI stack. Schema-constrained extraction with automatic validation and retry. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/instructor - Official site: https://python.useinstructor.com - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Schema-validated extraction with typed retry and partial streaming. - Skip it when: You need a hard token-level guarantee rather than validation. - Owned by: Applied Engineering - Alternatives: Guidance, Outlines ### Humanloop Humanloop is a SaaS platform in the Prompt Engineering layer of the AI stack. Prompt versioning and evaluation for production teams. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/humanloop - Official site: https://humanloop.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Prompt versioning and evaluation run by a product team. - Skip it when: Prompts or data cannot be sent to a vendor. - Owned by: Applied Engineering, ML Platform - Alternatives: PromptLayer, Agenta ### Guidance Guidance is a self-hosted library in the Prompt Engineering layer of the AI stack. Constrained generation by interleaving control flow and model output. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/guidance - Official site: https://guidance.mit.edu - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Constrained decoding with control flow interleaved into generation. - Skip it when: You need better ergonomics for plain extraction. - Owned by: Applied Engineering - Alternatives: Outlines, Instructor ### PromptLayer PromptLayer is a SaaS platform in the Prompt Engineering layer of the AI stack. Prompt registry with request logging and regression testing. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/promptlayer - Official site: https://promptlayer.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: A prompt registry with request logging and regression tests. - Skip it when: You need self-hosting. - Owned by: Applied Engineering, ML Platform - Alternatives: Humanloop, Agenta, Promptwatch ### Agenta Agenta is a self-hosted platform in the Prompt Engineering layer of the AI stack. Open-source prompt management with a playground and versioning. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/agenta - Official site: https://agenta.ai - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free-tier - Use it when: Open-source prompt management with a playground and versioning. - Skip it when: You need eval depth rather than prompt ergonomics. - Owned by: Applied Engineering, ML Platform - Alternatives: Humanloop, PromptLayer ### Promptwatch Promptwatch is a SaaS platform in the Prompt Engineering layer of the AI stack. Prompt regression tests and side-by-side comparison. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/promptwatch - Official site: https://promptwatch.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Prompt regression tests and side-by-side comparison. - Skip it when: You need self-hosting. - Owned by: Production & Governance - Alternatives: PromptLayer, Agenta, Humanloop ### TextGrad TextGrad is a self-hosted library in the Prompt Engineering layer of the AI stack. Automatic prompt optimisation via textual feedback loops. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/textgrad - Official site: https://github.com/zou-group/textgrad - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: Python - Cost: free - Use it when: Automatic prompt optimisation via textual feedback loops. - Skip it when: Optimising without an eval set is guessing. - Owned by: Applied Engineering - Alternatives: DSPy ### Outlines Outlines is a self-hosted library in the Prompt Engineering layer of the AI stack. Constrained generation that masks the token space to valid output. - Page: https://lattice.kkshah2005.workers.dev/prompt-engineering/outlines - Official site: https://dottxt-ai.github.io/outlines - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Token-level constrained generation from a JSON schema or grammar. - Skip it when: Post-hoc validation is enough for your use case. - Owned by: Applied Engineering - Alternatives: Guidance, Instructor ## Evaluation & Observability https://lattice.kkshah2005.workers.dev/evaluation-observability Traces, datasets and graders for non-deterministic output. The only way to know whether any of the above works. ### LangSmith LangSmith is a SaaS platform in the Evaluation & Observability section of Lattice. Tracing, evaluation and dataset tooling across LangChain runs. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/langsmith - Official site: https://smith.langchain.com - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Teams already deep in LangChain needing tracing and datasets. - Skip it when: Traces cannot leave your infrastructure. - Owned by: Production & Governance - Alternatives: Langfuse, Braintrust, Arize Phoenix ### Braintrust Braintrust is a SaaS platform in the Evaluation & Observability section of Lattice. Evaluation platform with built-in scorers and a data flywheel. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/braintrust - Official site: https://braintrust.dev - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Making evaluation, rather than tracing, the primary workflow. - Skip it when: You need open formats and self-hosting. - Owned by: Production & Governance - Alternatives: LangSmith, Langfuse, DeepEval ### Arize Phoenix Arize Phoenix is a self-hosted library in the Evaluation & Observability section of Lattice. Open-source tracing and evaluation built on OpenTelemetry. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/arize-phoenix - Official site: https://phoenix.arize.com - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: OpenTelemetry-native evaluation you can run yourself. - Skip it when: You want a vendor support contract behind it. - Owned by: Production & Governance - Alternatives: Langfuse, Opik, Gentrace ### Langfuse Langfuse is a self-hosted platform in the Evaluation & Observability section of Lattice. Self-hostable LLM tracing, prompt management and cost analytics. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/langfuse - Official site: https://langfuse.com - Kind: platform - Deployment: self-hosted - Licence: MIT (open-source) - Language: TypeScript - Cost: free-tier - Use it when: Self-hostable tracing, prompts and evals in one place. - Skip it when: You want a fully managed product with a support contract. - Owned by: Production & Governance - Alternatives: LangSmith, Braintrust, Arize Phoenix, Gentrace, Opik - Also in prompt-engineering: Its prompt registry versions prompts and ties each version to the traces and scores it produced. ### promptfoo promptfoo is a self-hosted library in the Evaluation & Observability section of Lattice. Declarative red-teaming and regression tests for prompts and agents. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/promptfoo - Official site: https://promptfoo.dev - Kind: library - Deployment: self-hosted - Licence: MIT (open-source) - Language: TypeScript - Cost: free - Use it when: Declarative red-teaming and regression tests that run in CI. - Skip it when: You need a managed UI rather than a test runner. - Owned by: Production & Governance - Alternatives: DeepEval, garak ### DeepEval DeepEval is a self-hosted library in the Evaluation & Observability section of Lattice. Open-source pytest-style evaluation suite for LLM outputs. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/deepeval - Official site: https://deepeval.com - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: pytest-style evaluation you can run next to your unit tests. - Skip it when: You need a platform rather than a library. - Owned by: Production & Governance - Alternatives: promptfoo, Braintrust ### Helicone Helicone is a SaaS platform in the Evaluation & Observability section of Lattice. Gateway-level observability with per-request cost and latency. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/helicone - Official site: https://helicone.ai - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: free-tier - Use it when: Gateway-level observability with per-request cost and latency. - Skip it when: You need eval primitives more than traffic logs. - Owned by: Production & Governance - Alternatives: Langfuse, Gentrace - Also in routing-gateways: It observes at the gateway, so the requests it reports on have already passed through layer 2. ### Weights & Biases Weights & Biases is a SaaS platform in the Evaluation & Observability section of Lattice. Experiment tracking and model registry with LLM eval surfaces. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/weights-biases - Official site: https://wandb.ai - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Experiment tracking and a model registry, with eval surfaces on top. - Skip it when: You need an open, exportable format. - Owned by: Production & Governance - Alternatives: Weights & Biases Launch, LangSmith - Also in fine-tuning: The model registry and experiment ledger it is known for are training-time surfaces; this entry carries its eval half. - Also in workflow-orchestration: Sweeps and artifact tracking are the orchestration of many training runs, which is the layer 6 job. ### OpenTelemetry OpenTelemetry is a self-hosted library in the Evaluation & Observability section of Lattice. Vendor-neutral standard for emitting traces and metrics. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/opentelemetry - Official site: https://opentelemetry.io - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: multi - Cost: free - Use it when: A vendor-neutral standard for emitting traces and metrics. - Skip it when: You want a product with a user interface. - Owned by: Production & Governance ### Vellum Vellum is a SaaS platform in the Evaluation & Observability section of Lattice. Prompt and eval platform with a visual debugger for chains. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/vellum - Official site: https://vellum.ai - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: usage-based - Use it when: Prompt and eval work with a visual debugger for chains. - Skip it when: Self-hosting or reproducibility is a requirement. - Owned by: Production & Governance - Alternatives: Humanloop, Braintrust - Also in prompt-engineering: Prompts are versioned artifacts here, diffed and promoted rather than edited in place. ### Gentrace Gentrace is a self-hosted platform in the Evaluation & Observability section of Lattice. Open-source tracing and dashboards for LLM applications. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/gentrace - Official site: https://gentrace.ai - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: TypeScript - Cost: free-tier - Use it when: Open-source tracing and dashboards for LLM applications. - Skip it when: You want eval depth more than traffic visibility. - Owned by: Production & Governance - Alternatives: Arize Phoenix, Helicone, Opik ### Opik Opik is a self-hosted platform in the Evaluation & Observability section of Lattice. Open-source LLM observability and evaluation from Comet. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/opik - Official site: https://comet.com - Kind: platform - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: Open-source observability and evaluation you run yourself. - Skip it when: You need enterprise support behind it. - Owned by: Production & Governance - Alternatives: Arize Phoenix, Gentrace ### Evidently AI Evidently AI is a self-hosted library in the Evaluation & Observability section of Lattice. Evaluation and monitoring for both classical and LLM systems. - Page: https://lattice.kkshah2005.workers.dev/evaluation-observability/evidently-ai - Official site: https://evidentlyai.com - Kind: library - Deployment: self-hosted - Licence: Apache-2.0 (open-source) - Language: Python - Cost: free - Use it when: One evaluation model for both classical ML and LLM systems. - Skip it when: You only need LLM-specific tracing. - Owned by: Production & Governance, Data & Retrieval - Alternatives: DeepEval, promptfoo ## Learning & Reference https://lattice.kkshah2005.workers.dev/learning-reference Where to read once the tooling starts to blur together. Not a layer. Reading, once you know what to look for. ### Hugging Face Hugging Face is a SaaS platform in the Learning & Reference section of Lattice. Model hub, datasets and open-source library ecosystem. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/hugging-face - Official site: https://huggingface.co - Kind: platform - Deployment: saas - Licence: proprietary (proprietary) - Cost: free-tier - Use it when: The model hub, the datasets and the open-source library ecosystem. - Skip it when: You need a self-hosted artefact registry. - Owned by: Data & Retrieval, Applied Engineering ### Full Stack Deep Learning Full Stack Deep Learning is a reading resource in the Learning & Reference section of Lattice. Course covering the practical side of shipping LLM systems. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/full-stack-deep-learning - Official site: https://fullstackdeeplearning.com - Kind: reading - Cost: free - Owned by: Applied Engineering ### Sebastian Raschka Sebastian Raschka is a reading resource in the Learning & Reference section of Lattice. Essays and books distilling research into working knowledge. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/sebastian-raschka - Official site: https://magazine.sebastianraschka.com - Kind: reading - Cost: free - Owned by: Applied Engineering ### arXiv arXiv is a reading resource in the Learning & Reference section of Lattice. Primary preprint archive for machine learning research. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/arxiv - Official site: https://arxiv.org - Kind: reading - Cost: free - Owned by: Applied Engineering ### Papers with Code Papers with Code is a reading resource in the Learning & Reference section of Lattice. Links papers to their reference implementations and benchmarks. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/papers-with-code - Official site: https://paperswithcode.com - Kind: reading - Cost: free - Owned by: Applied Engineering ### Jay Alammar Jay Alammar is a reading resource in the Learning & Reference section of Lattice. Visual explanations of transformers, RAG and LLM internals. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/jay-alammar - Official site: https://jalammar.github.io - Kind: reading - Cost: free - Owned by: Applied Engineering, Data & Retrieval ### AI Engineer AI Engineer is a reading resource in the Learning & Reference section of Lattice. Conference and community focused on shipping LLM products. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/ai-engineer - Official site: https://ai.engineer - Kind: reading - Cost: free - Owned by: Applied Engineering ### Hugging Face Cookbook Hugging Face Cookbook is a reading resource in the Learning & Reference section of Lattice. Task-oriented notebooks for fine-tuning, serving and RAG. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/hugging-face-cookbook - Official site: https://huggingface.co - Kind: reading - Cost: free - Owned by: Applied Engineering ### Made With ML Made With ML is a reading resource in the Learning & Reference section of Lattice. Course on designing, building and deploying ML systems. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/made-with-ml - Official site: https://madewithml.com - Kind: reading - Cost: free - Owned by: ML Platform ### Distill Distill is a reading resource in the Learning & Reference section of Lattice. Archive of carefully explained model and method explainers. - Page: https://lattice.kkshah2005.workers.dev/learning-reference/distill - Official site: https://distill.pub - Kind: reading - Cost: free - Owned by: Production & Governance ## Tools that span layers Each tool below has one home section and is listed under every other layer it also serves. ### Routing & Gateways https://lattice.kkshah2005.workers.dev/routing-gateways - Helicone — also indexed in Evaluation & Observability (https://lattice.kkshah2005.workers.dev/evaluation-observability/helicone): It observes at the gateway, so the requests it reports on have already passed through layer 2. ### Retrieval & Vector Stores https://lattice.kkshah2005.workers.dev/retrieval-vector-stores - Letta — also indexed in Agent Frameworks (https://lattice.kkshah2005.workers.dev/agent-frameworks/letta): Its memory is persisted and retrieved over an embedding store, so it is a layer-3 store with an agent API on top. ### Fine-tuning & Training https://lattice.kkshah2005.workers.dev/fine-tuning - Weights & Biases — also indexed in Evaluation & Observability (https://lattice.kkshah2005.workers.dev/evaluation-observability/weights-biases): The model registry and experiment ledger it is known for are training-time surfaces; this entry carries its eval half. ### Workflow Orchestration https://lattice.kkshah2005.workers.dev/workflow-orchestration - Weights & Biases — also indexed in Evaluation & Observability (https://lattice.kkshah2005.workers.dev/evaluation-observability/weights-biases): Sweeps and artifact tracking are the orchestration of many training runs, which is the layer 6 job. ### Guardrails & Safety https://lattice.kkshah2005.workers.dev/guardrails-safety - Portkey — also indexed in Routing & Gateways (https://lattice.kkshah2005.workers.dev/routing-gateways/portkey): Its guardrails run inline in the request path, which is where safety checks belong rather than after the response. - TrueFoundry — also indexed in Routing & Gateways (https://lattice.kkshah2005.workers.dev/routing-gateways/truefoundry): Its guardrails enforce in the call path alongside the gateway rather than as a separate service. - OpenAI Agents SDK — also indexed in Agent Frameworks (https://lattice.kkshah2005.workers.dev/agent-frameworks/openai-agents-sdk): Guardrails are a first-party primitive on the runner rather than a proxy in front of it. ### Prompt Engineering https://lattice.kkshah2005.workers.dev/prompt-engineering - Langfuse — also indexed in Evaluation & Observability (https://lattice.kkshah2005.workers.dev/evaluation-observability/langfuse): Its prompt registry versions prompts and ties each version to the traces and scores it produced. - Vellum — also indexed in Evaluation & Observability (https://lattice.kkshah2005.workers.dev/evaluation-observability/vellum): Prompts are versioned artifacts here, diffed and promoted rather than edited in place. ### Evaluation & Observability https://lattice.kkshah2005.workers.dev/evaluation-observability - TrueFoundry — also indexed in Routing & Gateways (https://lattice.kkshah2005.workers.dev/routing-gateways/truefoundry): Its tracing and dashboards are the observability half of the same deployment, not an add-on. - Weights & Biases Launch — also indexed in Fine-tuning & Training (https://lattice.kkshah2005.workers.dev/fine-tuning/weights-biases-launch): A training run is an experiment, scored against a baseline the same way an eval set scores a prompt. - Mastra — also indexed in Agent Frameworks (https://lattice.kkshah2005.workers.dev/agent-frameworks/mastra): Its eval harness ships inside the framework, so grading an agent is part of running it rather than a separate tool. ## By role ### ML Platform (24) The paved road other engineers build on. Question it arrives with: How does a model get called, and what does a teammate inherit? https://lattice.kkshah2005.workers.dev/roles/platform - LiteLLM — https://lattice.kkshah2005.workers.dev/routing-gateways/litellm (Routing & Gateways) - Portkey — https://lattice.kkshah2005.workers.dev/routing-gateways/portkey (Routing & Gateways) - Cloudflare AI Gateway — https://lattice.kkshah2005.workers.dev/routing-gateways/cloudflare-ai-gateway (Routing & Gateways) - OpenRouter — https://lattice.kkshah2005.workers.dev/routing-gateways/openrouter (Routing & Gateways) - Martian — https://lattice.kkshah2005.workers.dev/routing-gateways/martian (Routing & Gateways) - Envoy AI Gateway — https://lattice.kkshah2005.workers.dev/routing-gateways/envoy-ai-gateway (Routing & Gateways) - Bifrost — https://lattice.kkshah2005.workers.dev/routing-gateways/bifrost (Routing & Gateways) - RouteLLM — https://lattice.kkshah2005.workers.dev/routing-gateways/routellm (Routing & Gateways) - Not Diamond — https://lattice.kkshah2005.workers.dev/routing-gateways/not-diamond (Routing & Gateways) - TrueFoundry — https://lattice.kkshah2005.workers.dev/routing-gateways/truefoundry (Routing & Gateways) - Temporal — https://lattice.kkshah2005.workers.dev/workflow-orchestration/temporal (Workflow Orchestration) - Inngest — https://lattice.kkshah2005.workers.dev/workflow-orchestration/inngest (Workflow Orchestration) - Trigger.dev — https://lattice.kkshah2005.workers.dev/workflow-orchestration/trigger-dev (Workflow Orchestration) - Dagster — https://lattice.kkshah2005.workers.dev/workflow-orchestration/dagster (Workflow Orchestration) - Prefect — https://lattice.kkshah2005.workers.dev/workflow-orchestration/prefect (Workflow Orchestration) - Apache Airflow — https://lattice.kkshah2005.workers.dev/workflow-orchestration/apache-airflow (Workflow Orchestration) - Restate — https://lattice.kkshah2005.workers.dev/workflow-orchestration/restate (Workflow Orchestration) - DBOS — https://lattice.kkshah2005.workers.dev/workflow-orchestration/dbos (Workflow Orchestration) - Fly Machines — https://lattice.kkshah2005.workers.dev/workflow-orchestration/fly-machines (Workflow Orchestration) - Modal — https://lattice.kkshah2005.workers.dev/workflow-orchestration/modal (Workflow Orchestration) - Humanloop — https://lattice.kkshah2005.workers.dev/prompt-engineering/humanloop (Prompt Engineering) - PromptLayer — https://lattice.kkshah2005.workers.dev/prompt-engineering/promptlayer (Prompt Engineering) - Agenta — https://lattice.kkshah2005.workers.dev/prompt-engineering/agenta (Prompt Engineering) - Made With ML — https://lattice.kkshah2005.workers.dev/learning-reference/made-with-ml (Learning & Reference) ### AI Infrastructure (18) Tokens per second, per dollar, per GPU. Question it arrives with: Why is this slow, and what does the next hardware change buy me? https://lattice.kkshah2005.workers.dev/roles/serving - vLLM — https://lattice.kkshah2005.workers.dev/inference-serving/vllm (Inference & Serving) - SGLang — https://lattice.kkshah2005.workers.dev/inference-serving/sglang (Inference & Serving) - llama.cpp — https://lattice.kkshah2005.workers.dev/inference-serving/llama-cpp (Inference & Serving) - Ollama — https://lattice.kkshah2005.workers.dev/inference-serving/ollama (Inference & Serving) - TensorRT-LLM — https://lattice.kkshah2005.workers.dev/inference-serving/tensorrt-llm (Inference & Serving) - Text Generation Inference — https://lattice.kkshah2005.workers.dev/inference-serving/text-generation-inference (Inference & Serving) - LM Studio — https://lattice.kkshah2005.workers.dev/inference-serving/lm-studio (Inference & Serving) - Triton Inference Server — https://lattice.kkshah2005.workers.dev/inference-serving/triton-inference-server (Inference & Serving) - llamafile — https://lattice.kkshah2005.workers.dev/inference-serving/llamafile (Inference & Serving) - KoboldCpp — https://lattice.kkshah2005.workers.dev/inference-serving/koboldcpp (Inference & Serving) - PowerInfer — https://lattice.kkshah2005.workers.dev/inference-serving/powerinfer (Inference & Serving) - Marlin — https://lattice.kkshah2005.workers.dev/inference-serving/marlin (Inference & Serving) - Envoy AI Gateway — https://lattice.kkshah2005.workers.dev/routing-gateways/envoy-ai-gateway (Routing & Gateways) - Unsloth — https://lattice.kkshah2005.workers.dev/fine-tuning/unsloth (Fine-tuning & Training) - DeepSpeed — https://lattice.kkshah2005.workers.dev/fine-tuning/deepspeed (Fine-tuning & Training) - Megatron-LM — https://lattice.kkshah2005.workers.dev/fine-tuning/megatron-lm (Fine-tuning & Training) - Fly Machines — https://lattice.kkshah2005.workers.dev/workflow-orchestration/fly-machines (Workflow Orchestration) - Modal — https://lattice.kkshah2005.workers.dev/workflow-orchestration/modal (Workflow Orchestration) ### Data & Retrieval (23) What the model knows, and whether it is the right thing. Question it arrives with: Why does it know the wrong thing, or nothing? https://lattice.kkshah2005.workers.dev/roles/data - pgvector — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/pgvector (Retrieval & Vector Stores) - Qdrant — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/qdrant (Retrieval & Vector Stores) - Weaviate — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/weaviate (Retrieval & Vector Stores) - Chroma — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/chroma (Retrieval & Vector Stores) - Pinecone — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/pinecone (Retrieval & Vector Stores) - Milvus — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/milvus (Retrieval & Vector Stores) - Turbopuffer — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/turbopuffer (Retrieval & Vector Stores) - Unstructured — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/unstructured (Retrieval & Vector Stores) - Elasticsearch — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/elasticsearch (Retrieval & Vector Stores) - Vespa — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/vespa (Retrieval & Vector Stores) - Rerankers — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/rerankers (Retrieval & Vector Stores) - Jina AI — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/jina-ai (Retrieval & Vector Stores) - LangChain Text Splitters — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/langchain-text-splitters (Retrieval & Vector Stores) - Docling — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/docling (Retrieval & Vector Stores) - LlamaParse — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/llamaparse (Retrieval & Vector Stores) - LlamaIndex — https://lattice.kkshah2005.workers.dev/agent-frameworks/llamaindex (Agent Frameworks) - Letta — https://lattice.kkshah2005.workers.dev/agent-frameworks/letta (Agent Frameworks) - Dagster — https://lattice.kkshah2005.workers.dev/workflow-orchestration/dagster (Workflow Orchestration) - Prefect — https://lattice.kkshah2005.workers.dev/workflow-orchestration/prefect (Workflow Orchestration) - Apache Airflow — https://lattice.kkshah2005.workers.dev/workflow-orchestration/apache-airflow (Workflow Orchestration) - Evidently AI — https://lattice.kkshah2005.workers.dev/evaluation-observability/evidently-ai (Evaluation & Observability) - Hugging Face — https://lattice.kkshah2005.workers.dev/learning-reference/hugging-face (Learning & Reference) - Jay Alammar — https://lattice.kkshah2005.workers.dev/learning-reference/jay-alammar (Learning & Reference) ### Applied Engineering (46) The feature, end to end, in the product. Question it arrives with: How do I get from a model call to a working surface? https://lattice.kkshah2005.workers.dev/roles/applied - Ollama — https://lattice.kkshah2005.workers.dev/inference-serving/ollama (Inference & Serving) - LM Studio — https://lattice.kkshah2005.workers.dev/inference-serving/lm-studio (Inference & Serving) - RouteLLM — https://lattice.kkshah2005.workers.dev/routing-gateways/routellm (Routing & Gateways) - Not Diamond — https://lattice.kkshah2005.workers.dev/routing-gateways/not-diamond (Routing & Gateways) - Chroma — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/chroma (Retrieval & Vector Stores) - LangChain Text Splitters — https://lattice.kkshah2005.workers.dev/retrieval-vector-stores/langchain-text-splitters (Retrieval & Vector Stores) - Unsloth — https://lattice.kkshah2005.workers.dev/fine-tuning/unsloth (Fine-tuning & Training) - Axolotl — https://lattice.kkshah2005.workers.dev/fine-tuning/axolotl (Fine-tuning & Training) - LLaMA-Factory — https://lattice.kkshah2005.workers.dev/fine-tuning/llama-factory (Fine-tuning & Training) - PEFT — https://lattice.kkshah2005.workers.dev/fine-tuning/peft (Fine-tuning & Training) - TRL — https://lattice.kkshah2005.workers.dev/fine-tuning/trl (Fine-tuning & Training) - DeepSpeed — https://lattice.kkshah2005.workers.dev/fine-tuning/deepspeed (Fine-tuning & Training) - torchtune — https://lattice.kkshah2005.workers.dev/fine-tuning/torchtune (Fine-tuning & Training) - Hugging Face TRL — https://lattice.kkshah2005.workers.dev/fine-tuning/hugging-face-trl (Fine-tuning & Training) - Colab — https://lattice.kkshah2005.workers.dev/fine-tuning/colab (Fine-tuning & Training) - Replicate — https://lattice.kkshah2005.workers.dev/fine-tuning/replicate (Fine-tuning & Training) - Weights & Biases Launch — https://lattice.kkshah2005.workers.dev/fine-tuning/weights-biases-launch (Fine-tuning & Training) - LangChain — https://lattice.kkshah2005.workers.dev/agent-frameworks/langchain (Agent Frameworks) - LlamaIndex — https://lattice.kkshah2005.workers.dev/agent-frameworks/llamaindex (Agent Frameworks) - Pydantic AI — https://lattice.kkshah2005.workers.dev/agent-frameworks/pydantic-ai (Agent Frameworks) - AutoGen — https://lattice.kkshah2005.workers.dev/agent-frameworks/autogen (Agent Frameworks) - CrewAI — https://lattice.kkshah2005.workers.dev/agent-frameworks/crewai (Agent Frameworks) - Semantic Kernel — https://lattice.kkshah2005.workers.dev/agent-frameworks/semantic-kernel (Agent Frameworks) - Mastra — https://lattice.kkshah2005.workers.dev/agent-frameworks/mastra (Agent Frameworks) - OpenAI Agents SDK — https://lattice.kkshah2005.workers.dev/agent-frameworks/openai-agents-sdk (Agent Frameworks) - Agno — https://lattice.kkshah2005.workers.dev/agent-frameworks/agno (Agent Frameworks) - Claude Agent SDK — https://lattice.kkshah2005.workers.dev/agent-frameworks/claude-agent-sdk (Agent Frameworks) - Letta — https://lattice.kkshah2005.workers.dev/agent-frameworks/letta (Agent Frameworks) - smolagents — https://lattice.kkshah2005.workers.dev/agent-frameworks/smolagents (Agent Frameworks) - Vercel AI SDK — https://lattice.kkshah2005.workers.dev/agent-frameworks/vercel-ai-sdk (Agent Frameworks) - DSPy — https://lattice.kkshah2005.workers.dev/prompt-engineering/dspy (Prompt Engineering) - Instructor — https://lattice.kkshah2005.workers.dev/prompt-engineering/instructor (Prompt Engineering) - Humanloop — https://lattice.kkshah2005.workers.dev/prompt-engineering/humanloop (Prompt Engineering) - Guidance — https://lattice.kkshah2005.workers.dev/prompt-engineering/guidance (Prompt Engineering) - PromptLayer — https://lattice.kkshah2005.workers.dev/prompt-engineering/promptlayer (Prompt Engineering) - Agenta — https://lattice.kkshah2005.workers.dev/prompt-engineering/agenta (Prompt Engineering) - TextGrad — https://lattice.kkshah2005.workers.dev/prompt-engineering/textgrad (Prompt Engineering) - Outlines — https://lattice.kkshah2005.workers.dev/prompt-engineering/outlines (Prompt Engineering) - Hugging Face — https://lattice.kkshah2005.workers.dev/learning-reference/hugging-face (Learning & Reference) - Full Stack Deep Learning — https://lattice.kkshah2005.workers.dev/learning-reference/full-stack-deep-learning (Learning & Reference) - Sebastian Raschka — https://lattice.kkshah2005.workers.dev/learning-reference/sebastian-raschka (Learning & Reference) - arXiv — https://lattice.kkshah2005.workers.dev/learning-reference/arxiv (Learning & Reference) - Papers with Code — https://lattice.kkshah2005.workers.dev/learning-reference/papers-with-code (Learning & Reference) - Jay Alammar — https://lattice.kkshah2005.workers.dev/learning-reference/jay-alammar (Learning & Reference) - AI Engineer — https://lattice.kkshah2005.workers.dev/learning-reference/ai-engineer (Learning & Reference) - Hugging Face Cookbook — https://lattice.kkshah2005.workers.dev/learning-reference/hugging-face-cookbook (Learning & Reference) ### Production & Governance (25) Whether it can be trusted, and whether it can be let go. Question it arrives with: How do I know it still works, and what am I accountable for? https://lattice.kkshah2005.workers.dev/roles/production - TrueFoundry — https://lattice.kkshah2005.workers.dev/routing-gateways/truefoundry (Routing & Gateways) - Weights & Biases Launch — https://lattice.kkshah2005.workers.dev/fine-tuning/weights-biases-launch (Fine-tuning & Training) - NeMo Guardrails — https://lattice.kkshah2005.workers.dev/guardrails-safety/nemo-guardrails (Guardrails & Safety) - Guardrails AI — https://lattice.kkshah2005.workers.dev/guardrails-safety/guardrails-ai (Guardrails & Safety) - Llama Guard — https://lattice.kkshah2005.workers.dev/guardrails-safety/llama-guard (Guardrails & Safety) - Microsoft Presidio — https://lattice.kkshah2005.workers.dev/guardrails-safety/microsoft-presidio (Guardrails & Safety) - Lakera Guard — https://lattice.kkshah2005.workers.dev/guardrails-safety/lakera-guard (Guardrails & Safety) - garak — https://lattice.kkshah2005.workers.dev/guardrails-safety/garak (Guardrails & Safety) - PyRIT — https://lattice.kkshah2005.workers.dev/guardrails-safety/pyrit (Guardrails & Safety) - Invariant Guardrails — https://lattice.kkshah2005.workers.dev/guardrails-safety/invariant-guardrails (Guardrails & Safety) - Promptwatch — https://lattice.kkshah2005.workers.dev/prompt-engineering/promptwatch (Prompt Engineering) - LangSmith — https://lattice.kkshah2005.workers.dev/evaluation-observability/langsmith (Evaluation & Observability) - Braintrust — https://lattice.kkshah2005.workers.dev/evaluation-observability/braintrust (Evaluation & Observability) - Arize Phoenix — https://lattice.kkshah2005.workers.dev/evaluation-observability/arize-phoenix (Evaluation & Observability) - Langfuse — https://lattice.kkshah2005.workers.dev/evaluation-observability/langfuse (Evaluation & Observability) - promptfoo — https://lattice.kkshah2005.workers.dev/evaluation-observability/promptfoo (Evaluation & Observability) - DeepEval — https://lattice.kkshah2005.workers.dev/evaluation-observability/deepeval (Evaluation & Observability) - Helicone — https://lattice.kkshah2005.workers.dev/evaluation-observability/helicone (Evaluation & Observability) - Weights & Biases — https://lattice.kkshah2005.workers.dev/evaluation-observability/weights-biases (Evaluation & Observability) - OpenTelemetry — https://lattice.kkshah2005.workers.dev/evaluation-observability/opentelemetry (Evaluation & Observability) - Vellum — https://lattice.kkshah2005.workers.dev/evaluation-observability/vellum (Evaluation & Observability) - Gentrace — https://lattice.kkshah2005.workers.dev/evaluation-observability/gentrace (Evaluation & Observability) - Opik — https://lattice.kkshah2005.workers.dev/evaluation-observability/opik (Evaluation & Observability) - Evidently AI — https://lattice.kkshah2005.workers.dev/evaluation-observability/evidently-ai (Evaluation & Observability) - Distill — https://lattice.kkshah2005.workers.dev/learning-reference/distill (Learning & Reference) ## Comparisons ### vLLM vs SGLang vs TGI vs llama.cpp https://lattice.kkshah2005.workers.dev/compare/inference-runtimes These are not really competitors; they are answers to four different questions. The mistake is picking on model support alone, because the thing that actually determines throughput is your traffic shape. | | vLLM | SGLang | Text Generation Inference | llama.cpp | |---|---|---|---|---| | Best for | General GPU serving under bursty load | Heavy shared prefixes, constrained decoding | Multi-GPU tensor parallelism in an HF shop | Local, edge, or no-GPU deployments | | Prefix caching | Automatic prefix caching | RadixAttention — strongest on long shared prefixes | Supported, less heavily optimised | KV cache reuse, not prefix-aware by default | | Hardware floor | One CUDA GPU | One CUDA GPU | Multi-GPU, designed for it | None — CPU is the default path | | Quantised formats | AWQ, GPTQ, FP8 | AWQ, GPTQ, FP8 | AWQ, GPTQ, FP8 | GGUF, any quant level | | API compatibility | OpenAI-compatible server | OpenAI-compatible server | OpenAI-compatible server | OpenAI-compatible server + native CLI | | Where it loses | Prefix-shared workloads, where SGLang wins | Small teams wanting the least moving parts | Single-GPU deployments — overkill | Maximum concurrent throughput | Recommendation: Start with vLLM. Benchmark SGLang against it if more than a few hundred requests share a long system prompt, because prefix reuse compounds and the gap is not marginal. Use llama.cpp when the constraint is hardware rather than throughput. Reach for TGI only when you specifically need multi-GPU tensor parallelism inside an existing Hugging Face setup. - Bursty traffic is what makes continuous batching pay. Steady low-concurrency traffic is better served by a managed API. - Never benchmark on single-stream latency — it measures the wrong thing and will point you at the wrong runtime. - Two GPUs running two replicas behind a load balancer usually beats one GPU running a sharded model. ### Langfuse vs LangSmith vs Braintrust vs Phoenix https://lattice.kkshah2005.workers.dev/compare/llm-observability The interesting divide is not open versus closed, it is tracing versus evaluation. Most teams buy a tracing tool and then discover they still cannot answer whether a change helped. | | Langfuse | LangSmith | Braintrust | Arize Phoenix | |---|---|---|---|---| | Best for | Self-hosting, or a privacy constraint | Teams already deep in LangChain | Making evaluation the primary workflow | OpenTelemetry-first, open-source everything | | Self-hostable | Yes | Enterprise plan | Yes, paid | Yes, fully open | | Open standards | OTel export | Partial | Proprietary | OTel-native throughout | | Prompt management | First-class, with versioning and A/B | Basic | Yes | Limited | | Evaluation | Datasets, scorers, CI-friendly | Strong offline and online eval | The strongest eval story of the four | Strong, trace-derived datasets | | Where it loses | Depth of framework-specific debugging | Cost, and vendor lock-in at higher tiers | Traces are secondary to evals | Polish and support | Recommendation: If you cannot send traces outside your infrastructure, self-host Langfuse and stop deliberating. Otherwise pick on evaluation rather than tracing, because tracing is table stakes and evaluation is what you will actually rely on. Whatever you choose, make sure traces carry the full request and response — that is the input to every future decision. - Tracing answers 'what happened'. Only evaluation answers 'is this better'. Buying one and expecting the other is the common mistake. - Check that model names, versions and per-hop latency are captured, or your cost and regression analysis will be wrong. - Adopting OpenTelemetry as the export format keeps your future options open at close to no cost. ### pgvector vs Qdrant vs Pinecone vs Chroma https://lattice.kkshah2005.workers.dev/compare/vector-databases Most teams adopt a dedicated vector database before they have established that their existing database cannot do the job. That decision is expensive to undo and rarely necessary. | | pgvector | Qdrant | Pinecone | Chroma | |---|---|---|---|---| | Operational burden | None — you already run Postgres | Low, single binary | None | None | | Hybrid (BM25 + vector) | Yes, via Postgres full-text search | Yes | Yes | Limited | | Filtering | SQL — very capable | Strong, purpose-built payloads | Strong | Basic metadata filters | | Scales past ~10M vectors | Workable, tuning-sensitive | Yes | Yes | No | | Best for | Under a few million vectors, or when joins matter | Vector-first workloads at scale | When you want zero operational work | Getting something working in an afternoon | Recommendation: Start on pgvector unless you have a specific reason not to. It removes a whole class of consistency problem, and if your retrieval is joining against rows you already have — users, permissions, tenancy — the relational model is simply the right one. Move to Qdrant or Pinecone when vector search stops being a side feature and starts being the workload. - If you need to filter by something that lives in your main database, a join beats a denormalised copy you have to keep in sync. - Pure vector search fails on exact tokens — part numbers, error codes, names. Use hybrid retrieval or accept that class of miss. - Validate recall against a query set before migrating. Most disappointing retrieval is a chunking or embedding problem, not a database problem. ### Temporal vs Inngest vs Trigger.dev vs Restate https://lattice.kkshah2005.workers.dev/compare/durable-workflows All four exist to solve the same problem: work that spans minutes or hours, across process restarts, without writing a state machine by hand. They differ on where that state lives. | | Temporal | Inngest | Trigger.dev | Restate | |---|---|---|---|---| | Languages | Go, Java, Python, TypeScript, .NET, PHP | TypeScript | TypeScript, Python | Java, Kotlin, Rust, TypeScript, Python | | Durability model | Full event history, deterministic replay | Step-function journal, replay from step | Checkpointed runs, resumable tasks | Single log per invocation | | Self-host required | Usually — or use Temporal Cloud | No — cloud is the default path | No | No | | Streaming / low latency | Not its strength | Good | Good | Best in class | | Where it loses | Heaviest to adopt, largest surface | TypeScript only | Narrower language coverage | Youngest ecosystem, smallest community | Recommendation: If your stack is TypeScript, Inngest or Trigger.dev will be live in a day and cover 90% of agent workloads. If you are polyglot or durability is the centre of your business, Temporal is the mature answer and you should accept the adoption cost. Restate is the one to watch if low-latency stateful endpoints matter more than ecosystem size. - Adopt one of these specifically so an agent loop survives a deploy. A process restart mid-loop is the most common way agents cause real damage. - The engine fixes durability, not idempotency. Side-effecting tool calls still need idempotency keys. - Put turn, token and wall-clock ceilings in the host. A prompt is a suggestion; the engine is a control. ### LiteLLM vs Portkey vs Cloudflare AI Gateway https://lattice.kkshah2005.workers.dev/compare/llm-gateways Every one of these solves 'talk to many providers through one interface'. They differ sharply on where your prompts physically travel, which is a compliance question before it is a technical one. | | LiteLLM | Portkey | Cloudflare AI Gateway | Envoy AI Gateway | |---|---|---|---|---| | Deployment | Self-host or cloud | Cloud-first, self-host available | Managed edge service | Self-host in your cluster | | Provider count | 100+ | Broad | Broad | Plugin-driven | | Caching | Redis-backed, including semantic | Yes | Yes, at the edge | Via the data plane | | Routing logic | Yours — config plus code | Policy engine, mostly declarative | Simple rules at the edge | Yours, expressed as Envoy config | | Where prompts go | Your infrastructure, if self-hosted | Vendor SaaS unless self-hosted | Through Cloudflare's edge | Never leaves your cluster | Recommendation: Self-host LiteLLM if you need breadth and control; it is the least committal. Choose Portkey if you want a managed product with guardrails included and are comfortable with a SaaS dependency. Cloudflare AI Gateway is the pick when latency and edge caching dominate, and Envoy AI Gateway when you need a gateway that never sees traffic leave your network. - Put a gateway in front of your provider calls early, even if it only does auth and logging. It is the only place that sees the whole request. - Never put prompt construction in the gateway. It is a control plane, not application logic. - Budgets are only enforceable where spend is attributed, which means identity has to propagate from the gateway inward. ### Instructor vs JSON mode vs constrained decoding https://lattice.kkshah2005.workers.dev/compare/structured-output Asking a model for JSON is not the same as getting JSON. These three approaches differ in whether they give you a guarantee or a tendency — and that difference decides whether you can retry safely. | | Instructor | Guidance | Outlines | |---|---|---|---| | Guarantee | Validated, with typed retry | Token-level constraint — cannot be violated | Token-level constraint, schema or grammar | | Works across providers | Yes, with capability detection | Yes | Yes, via local or hosted inference | | Streaming | Partial objects supported | Yes | Yes | | Handles messy extraction | Yes — its main use case | Awkward | Awkward | | Constraint style | Post-hoc validation and repair | Token masking + control flow | Token masking, incl. JSON Schema and regex | | Where it loses | Not a hard guarantee like token constraints | Less ergonomic for plain extraction | Most setup of the three | Recommendation: Use a library like Instructor for extraction work: validation, typed retry and partial streaming are worth more than an absolute guarantee. Reach for true constrained decoding — Guidance or Outlines — when downstream code cannot tolerate a malformed response at all, because there the guarantee is worth the reduced flexibility. Treat a provider's JSON mode as a convenience rather than a contract, and validate regardless. - Validate in code regardless of what the provider promises. Modes get relaxed without notice. - Design the schema before the prompt. A schema that reflects how you will consume the data produces better extractions. - Constrained decoding restricts the token space, so it can make a model look worse at reasoning than it is. Do not benchmark reasoning through it. ### Gateway vs guardrails vs evals: what to build first https://lattice.kkshah2005.workers.dev/compare/gateway-guardrails-evals These three never appear on the same vendor comparison page, because no vendor sells all three. They do appear on the same backlog. Each catches a different failure, sits in a different place in the request, and costs something different to adopt — so the real question is order, not choice. | | LiteLLM | Guardrails AI | promptfoo | |---|---|---|---| | Question it answers | Which model, provider and budget serves this request? | Is this response allowed to leave the system? | Did the last change make answers better or worse? | | Failure it catches | Provider outages, rate limits, runaway spend | Schema violations and policy breaches, per request | Quality regressions, before they ship | | Failure it misses | Whether any answer was correct | Answers that are well-formed and wrong | Anything happening in production right now | | Where it runs | In the request path, on every call | In the request path, after generation | Offline, in CI or on a schedule | | Cost to adopt | A proxy and a config file; one more network hop | A validator per rule; added latency on every response | A labelled dataset — the examples are the expensive part, not the tool | | Adopt first when | You call more than one provider, or spend is already a line item | A malformed or unsafe output has a real cost on day one | You are about to change prompts, models or retrieval | Recommendation: For most teams, evals first. A gateway makes calls cheaper and more reliable, and guardrails make individual responses safer, but neither tells you whether the system is getting better — and every later change to routing, prompts or retrieval needs that answer to be judged at all. Put the gateway in early if you already run more than one provider, because retrofitting a proxy into every call site is tedious. Add guardrails for the specific failures your evals surface, not the ones you imagine. - A guardrail without an eval is a guess about which failures matter. Write the eval that found the failure, then the guardrail that blocks it. - Route on measured quality, not on price alone. A cheaper model that fails your eval set is not cheaper. - Anything in the request path adds latency to every call; anything offline adds none. Stay offline until production forces otherwise. ### Retrieval vs fine-tuning vs prompt optimisation: fixing wrong answers https://lattice.kkshah2005.workers.dev/compare/retrieval-finetuning-prompting A model giving wrong answers can be fixed at three different depths of the stack, and they are not interchangeable. Retrieval changes what the model can see, fine-tuning changes how it behaves, and prompt optimisation changes what it is asked. Picking the wrong depth is the most expensive mistake in this index. | | pgvector | Unsloth | DSPy | |---|---|---|---| | What it changes | The context available at question time | The weights — the model's default behaviour | The instructions and examples it is given | | Fixes | Missing, private or recent facts | Format, tone and narrow-task behaviour a prompt cannot hold | Underspecified instructions and weak examples | | Does not fix | A model that ignores the context it is given | Missing knowledge — weights are a poor database | Facts the model has never seen | | Needs from you | A corpus, a chunking strategy and a ranking you can inspect | Curated training examples and a GPU | Labelled examples and a metric that scores them | | Cost of being wrong | Low — re-index and retry | High — a training run, and a model you now have to serve | Low — prompts are text and revert cleanly | | Adopt first when | The right facts are absent or out of date | Prompting has plateaued on a narrow, stable task | The facts are present but the model uses them badly | Recommendation: Diagnose before choosing. If the right fact was never in the context, that is a retrieval problem and no amount of prompting or training fixes it. If the fact was there and the model ignored or misused it, optimise the prompt first, because it is cheap to try and cheap to undo. Fine-tune last, for narrow and stable tasks where prompting has measurably plateaued — it is the only one of the three that leaves you with a new artefact to serve and retrain. - Read the retrieved context before touching the prompt. Most “the model is wrong” bugs are “the ranking is wrong” bugs. - Fine-tuning teaches behaviour, not facts. If the answer changes monthly, it belongs in retrieval. - All three need a scored example set to know whether they worked. Build it first; it is shared. ### Self-hosting vs routing vs measuring: cutting inference cost https://lattice.kkshah2005.workers.dev/compare/cutting-inference-cost Inference cost can be attacked at the substrate, in the router, or by first finding out where it goes. Teams usually start with the most expensive of the three — standing up their own GPUs — when the cheapest would have told them it was unnecessary. | | vLLM | RouteLLM | Helicone | |---|---|---|---| | Lever | Own the serving, pay for hardware | Match query difficulty to model size | Attribute spend per request, user and feature | | Saves money when | Traffic is high and steady enough to keep GPUs busy | A meaningful share of queries are easy | Spend is concentrated somewhere you have not looked | | Loses money when | Utilisation is low — idle GPUs still bill by the hour | The router misjudges hard queries and answers degrade | Never on its own — it saves nothing until you act on it | | Quality risk | None if you serve the same model; real if you downsize to fit | Direct — this is a quality-for-cost trade by design | None | | Operational cost | A serving fleet, on-call and capacity planning | A router to calibrate, plus an eval to trust it | A proxy hop or an SDK wrapper | | Adopt first when | You know your utilisation and it is high | You can measure the quality you are trading away | You cannot yet say which feature costs the most | Recommendation: Measure first. Per-request cost attribution is the cheapest of the three and the only one with no quality risk, and it often shows a few features or prompts dominating spend — which a shorter prompt can fix without new infrastructure. Route next, but only once an eval can tell you what the cheaper model loses. Self-host last, when measured and sustained utilisation makes GPU-hours cheaper than tokens; below that line it raises the bill. - Cost per token is not cost per answer. A cheaper model that needs two retries is the expensive one. - Compare self-hosting against the price you actually pay, at the utilisation you actually have — not at peak. - A router without an eval is a cost cut with an unknown quality bill attached. ## Querying this index instead of reading it This document is a few thousand lines. If you arrived with a specific question, the MCP server at https://lattice.kkshah2005.workers.dev/mcp answers it in a few hundred tokens. Discovery: https://lattice.kkshah2005.workers.dev/mcp.json. Tools: about, search_tools, get_tool, compare_tools, list_layers, layer_overlaps, diagnose_symptom, list_comparisons, define_term. Two of them have no equivalent in this document, because they are inverse lookups: `layer_overlaps` answers "where does agent memory live" — each layer, and the tools that serve it without being indexed there — and `diagnose_symptom` takes a problem rather than a tool name and returns the ordered cheapest-first checklist. ## Glossary ### Paged attention https://lattice.kkshah2005.workers.dev/glossary/paged-attention Managing the KV cache in fixed-size blocks rather than one contiguous allocation, so memory is not fragmented and sequences can share a batch. This is the change that made GPU serving practical at scale, and the reason vLLM became the default. Before it, a batch was sized for the worst-case sequence in it. After it, you size for the average and grow dynamically. If you are choosing a runtime and do not know any other criterion, this is the one to know. ### KV cache https://lattice.kkshah2005.workers.dev/glossary/kv-cache The saved attention keys and values from tokens already processed, so the model does not recompute them for every new token. It is what makes generation incremental rather than quadratic, and its size is the main thing that determines how many concurrent sequences fit on a GPU. It grows with context length and batch size, which is why long contexts and high concurrency compete for the same memory. ### Continuous batching https://lattice.kkshah2005.workers.dev/glossary/continuous-batching Adding new sequences to a batch as others finish, rather than waiting for the whole batch to complete. It exploits the fact that sequence lengths vary wildly within a batch — waiting for the longest wastes most of the slot. Under steady, low-concurrency traffic there is never enough in flight to fill a batch, so this buys you nothing. That traffic-shape dependency is the biggest single factor in picking a runtime. ### Prefix caching https://lattice.kkshah2005.workers.dev/glossary/prefix-caching Reusing the KV cache for a prompt prefix that has already been processed, instead of recomputing it. If a large share of your requests begin with the same few hundred tokens — a long system prompt, a cached document, a shared tool schema — this is worth more than every other optimisation combined. If your prompts are short and unrelated, neither vLLM's automatic prefix caching nor SGLang's RadixAttention will do much. It also makes a stable system prompt dramatically cheaper than a per-request one. ### Quantisation https://lattice.kkshah2005.workers.dev/glossary/quantisation Storing and computing model weights in fewer bits than they were trained in, to trade accuracy for memory and speed. 8-bit is nearly free in quality; 4-bit usually is not, and the method matters more than the bit count — AWQ and GPTQ are perceptually better than naive rounding at the same width. FP8 is the easier win on recent NVIDIA hardware. Quantisation changes your throughput numbers but not your architectural decisions, so do it after you have chosen a runtime, not before. ### Tensor parallelism https://lattice.kkshah2005.workers.dev/glossary/tensor-parallelism Splitting a single layer's computation across several GPUs, so every token passes through all of them. It is the default answer to 'the model does not fit' and the easiest to get wrong, because it adds a collective operation to every forward pass and the cost grows with sequence length. If your model fits on one GPU with room for KV cache, do not do it. Two GPUs running two replicas behind a load balancer will usually beat one GPU running a sharded 70B. ### Speculative decoding https://lattice.kkshah2005.workers.dev/glossary/speculative-decoding Using a small draft model to propose several tokens at once, which a large model then verifies in a single pass. It is exact — the output distribution is unchanged — and speeds up decoding where it is bottlenecked on memory bandwidth rather than compute. The gain depends on the draft model agreeing with the target, so it helps most for predictable text and least for genuinely creative work. Worth trying when you have already exhausted batching and quantisation. ### Time to first token https://lattice.kkshah2005.workers.dev/glossary/time-to-first-token The delay between sending a request and receiving the first token of the response. It is the latency number users actually perceive, and it is dominated by queueing rather than by model speed. A user-facing chatbot is judged on TTFT; a batch job is judged on total throughput. Conflating them is how teams end up optimising the wrong metric for months. ### Throughput and latency https://lattice.kkshah2005.workers.dev/glossary/throughput Throughput is tokens per second across the whole deployment; latency is how long one request takes. They trade directly against each other. Batching raises throughput by making each request slower, so the two cannot be optimised independently. What you want depends entirely on traffic shape: an interactive product needs bounded latency under whatever concurrency arrives, a batch job wants total tokens cheap. Benchmark on a concurrency sweep and plot both — never single-stream latency, which will point you at the wrong runtime. ### Constrained decoding https://lattice.kkshah2005.workers.dev/glossary/constrained-output Masking the token space during generation so the output can only be a valid instance of a schema, grammar or regex. Unlike asking a model politely for JSON, this cannot be violated — there is no way to emit a token that would break the grammar. The trade is that constraining generation can make a model look worse at reasoning than it is, so never benchmark reasoning through it. Reach for it when downstream code cannot tolerate malformed output; use validation-and-retry when ergonomics matter more. ### Validation and repair https://lattice.kkshah2005.workers.dev/glossary/schema-enforcement Generating freely, then validating the output against a schema and retrying with the error fed back. Weaker than constrained decoding but far more ergonomic, and it handles messy extraction that token masking handles awkwardly. It gives a typed retry rather than a guarantee. If you cannot decide between the two, start here. ### Model router https://lattice.kkshah2005.workers.dev/glossary/model-router A layer that chooses which model answers a given request, from signals known before the model is called. The workable signals are explicit task identity, user tier, and deterministic input properties. The signal that does not work is a subjective read of how hard the prompt looks — that is where the escalation trap begins, where failures cluster on genuinely hard requests and you pay twice for exactly the cases the big model was needed for. ### The escalation trap https://lattice.kkshah2005.workers.dev/glossary/escalation-trap Routing cheap-first and retrying on a bigger model when the answer looks bad, which usually costs more than routing correctly up front. Three reasons it fails: the cheap model cannot judge its own output, failures cluster on hard cases so escalation doubles cost precisely where it was needed, and one provider hiccup away from a retry storm. It only works when a failure is externally observable — a schema check, a downstream test — in which case it is not a vibe but a validated retry. ### Semantic cache https://lattice.kkshah2005.workers.dev/glossary/semantic-cache Reusing a previous response when a new request is sufficiently similar in meaning, skipping the model call entirely. The fastest cost reduction available, and the easiest to get wrong: a cached answer goes stale in a way nobody notices. Scope it by embedding similarity *and* a version of everything that affected the answer, and never cache calls with side effects. ### Circuit breaker https://lattice.kkshah2005.workers.dev/glossary/circuit-breaker A per-provider, per-model gate that stops sending traffic somewhere after repeated failures, then probes for recovery. Provider error rates are real and are not independent across providers. A retry policy that assumes independence turns a partial outage into a total one by stampeding the survivor. Breaker plus jittered retries is the minimum for a gateway worth running. ### Provider prompt caching https://lattice.kkshah2005.workers.dev/glossary/prompt-caching A discount providers apply when a request repeats a prefix they have already processed. Completely separate from prefix caching in your own runtime: this is a billing discount, that is a computation saving. The practical consequence is that a large, stable system prompt is far cheaper than one rebuilt per request — which makes prompt structure a cost decision, not just a readability one. ### Embedding https://lattice.kkshah2005.workers.dev/glossary/embedding A vector representation of text such that semantically similar inputs land near each other. The same model must produce embeddings for documents and queries, or the space is meaningless. The failure mode to know about: embeddings are bad at exact tokens — part numbers, error codes, names — because `ERR-4421` and `ERR-4412` really are nearly the same thing. Users type exact tokens constantly. ### Chunking https://lattice.kkshah2005.workers.dev/glossary/chunking Splitting documents into retrievable pieces, which is where RAG quality is mostly won and lost. Fixed token count with overlap is close to the worst reasonable option: it cuts through sentences, tables and arguments. Chunk on structure — headings, paragraphs, table rows — and carry the heading path in the chunk metadata so retrieved text has its own context. Then check the number that correlates most with answer quality: how many chunks you actually feed the model. Fifty is almost always worse than five. ### Hybrid search https://lattice.kkshah2005.workers.dev/glossary/hybrid-search Combining lexical retrieval (BM25) with vector retrieval, so exact matches and semantic matches are both found. Pure vector search misses the specific tokens users type; pure lexical misses paraphrase. Unioning the candidates and then reranking gets both. Most vector databases support this as a single query, and managed ones often behind a flag. If yours does not, that alone justifies a migration. ### Reranking https://lattice.kkshah2005.workers.dev/glossary/reranking Re-scoring retrieved candidates with a slower, more accurate model before the generator sees them. Bi-encoders embed query and document independently — that is what makes them fast enough to search millions of vectors, and also why they judge relevance poorly. A cross-encoder sees both at once and is far better, but far too slow for the whole corpus. So: cheap retrieval for fifty candidates, cross-encoder for the best five. The highest-leverage addition to an existing RAG system, and the cheapest to try. ### Reciprocal rank fusion https://lattice.kkshah2005.workers.dev/glossary/reciprocal-rank-fusion Merging ranked lists from several retrievers by summing the reciprocal of each result's rank, rather than comparing scores. Score-based fusion requires the retrievers to be on a comparable scale, which they are not. RRF sidesteps that entirely by only using positions, which is why it is the default in hybrid search and almost always the right first attempt. ### Vector index https://lattice.kkshah2005.workers.dev/glossary/vector-index An approximate structure — usually HNSW or an IVF variant — that finds nearest neighbours without scanning every vector. Approximate means a recall/latency trade controlled by a build parameter, which is why you should benchmark recall against a brute-force scan on your own data rather than trusting a published figure. If your corpus is under a few million vectors, brute force is often fast enough and exactly correct. ### Recall@k https://lattice.kkshah2005.workers.dev/glossary/recall-at-k Of the documents that should answer a query, the fraction that appear in the top k retrieved. The single diagnostic that separates a retrieval problem from a generation problem: if the answer is not in the top k, no prompt change will help. Check it on twenty real queries before changing anything. It costs an hour and determines everything that follows. ### Context window https://lattice.kkshah2005.workers.dev/glossary/context-window The number of tokens a model can consider at once — a ceiling on what you can send, not a target for what you should. Long contexts degrade silently rather than gracefully. Models reliably answer from the beginning and end of a long document and perform near chance on the middle. A 50k-token context is therefore a few thousand tokens of reliable signal plus a large quantity that is present, billed, and largely ignored. ### LoRA https://lattice.kkshah2005.workers.dev/glossary/lora Training small low-rank matrices alongside frozen weights, so you can adapt a large model without touching it. Adapter weights are a few percent of the base model, so training needs a fraction of the memory. QLoRA additionally quantises the frozen base to 4-bit, which brings single-GPU adaptation of models that previously needed several. The practical cost is inference: adapters are only useful if someone is serving them. ### Supervised fine-tuning https://lattice.kkshah2005.workers.dev/glossary/sft Training on curated input/output pairs so the model adopts a specific task or response style. The baseline post-training step. Its value depends entirely on example quality: a thousand mediocre pairs produce a model that is reliably mediocre. SFT teaches format and behaviour; it does not reliably teach knowledge, which is what people expect it to do and what it cannot do. ### Direct preference optimisation https://lattice.kkshah2005.workers.dev/glossary/dpo Aligning a model to human preferences without an explicit reward model, by optimising on chosen-versus-rejected pairs. Removes the reward-model training step and its instabilities, which is why it largely replaced RLHF in practice. It needs preference data, which is expensive to produce well — and is where most projects quietly stall. ### RLHF https://lattice.kkshah2005.workers.dev/glossary/rlhf Reinforcement learning from human feedback: training a reward model from preferences, then optimising the policy against it. The method that made instruction-tuned assistants work, and now largely superseded in new work by DPO and GRPO because of its instability and cost. Still the reference point when you need to understand where the alternatives came from. ### GRPO https://lattice.kkshah2005.workers.dev/glossary/grpo Group relative policy optimisation — a reinforcement-learning method that compares a group of sampled answers rather than learning a separate reward model. Useful when you have a verifiable reward, a checker or a test, rather than human preference. It sidesteps reward-model training, which is the expensive and fragile part, and is why it shows up in reasoning-model post-training. ### Distillation https://lattice.kkshah2005.workers.dev/glossary/distillation Training a smaller model against the outputs of a larger one, to get most of the behaviour at a fraction of the cost. The best-understood reason to fine-tune, because you have a reference and the training signal is mechanical. The result is a model slightly worse at the edges and dramatically cheaper, which for high-volume well-defined tasks is usually the right trade. Note this is different from quantisation, which changes no behaviour at all. ### ZeRO https://lattice.kkshah2005.workers.dev/glossary/zero Sharding optimiser states, gradients and parameters across data-parallel workers to fit a larger model on the same GPUs. The memory optimisation that makes fine-tuning large models on modest clusters possible. If your model already fits on one device, ZeRO buys nothing and costs communication. ### Tool calling https://lattice.kkshah2005.workers.dev/glossary/tool-calling A model emitting a structured request to call a function, which your code executes and returns the result of. The mechanism almost all agent systems are built on. The security consequence is that a tool call is model output with the same failure modes as anything else the model produces — including injection arriving via retrieved content or a tool's own error message. Treat arguments as untrusted input. ### Agent loop https://lattice.kkshah2005.workers.dev/glossary/agent-loop The control structure where a model decides an action, it is executed, and the result feeds the next turn until the task is done. Twenty lines, and the least interesting part of an agent system. The engineering is in the host around it: surviving a deploy mid-loop, idempotent side effects, and a hard ceiling on turns and tokens. Prompts are advisory; only the host can stop a loop. ### Model Context Protocol https://lattice.kkshah2005.workers.dev/glossary/mcp An open protocol for exposing tools, resources and prompts to models, so integrations are written once rather than per model provider. The main thing to know is that it moves the integration burden from 'an adapter per provider per tool' to 'one server per tool'. The corollary is a security one: an MCP server is a new dependency that can exfiltrate whatever it is given, so it deserves the same review as any other external service. ### Budget ceiling https://lattice.kkshah2005.workers.dev/glossary/budget-ceiling A hard limit on turns, tokens or wall-clock time, enforced by the host rather than requested in the prompt. The instinct is to write 'you have at most ten turns' in the system prompt. That does not work: a prompt is a suggestion to a system good at generating plausible text, with no privileged access to its own budget. Every guarantee you give a user about an agent is a promise the host is keeping. ### Durable execution https://lattice.kkshah2005.workers.dev/glossary/durable-execution A workflow engine that checkpoints progress so work resumes after a crash, rather than restarting from nothing. Long, multi-step, partially-completed work full of slow external calls is exactly what these engines were built for, and agent loops are a textbook use case. It solves durability — not idempotency. If the process dies between two side effects, replay will repeat the first unless you designed for it. ### Idempotency key https://lattice.kkshah2005.workers.dev/glossary/idempotency-key A client-supplied identifier that lets a server recognise a repeated request and return the original result rather than acting twice. The difference between 'the deploy interrupted my agent' and 'the deploy interrupted my agent and it charged the customer twice'. Durable execution gets you the retry; idempotency gets you safety under it. Side-effecting tools should require one. ### Prompt injection https://lattice.kkshah2005.workers.dev/glossary/prompt-injection An attacker getting instructions into your context that your model then follows as if they came from you. Not a bug to be fixed but a property of systems that read untrusted text and act on it. Input filtering catches known patterns; the durable mitigation is not trusting the model's reasoning about permissions, and re-deriving authorisation from authoritative state rather than accepting its claim. ### Guardrail https://lattice.kkshah2005.workers.dev/glossary/guardrail Any control that keeps a model from doing something you have decided it must not do. 'Safety' is three distinct controls in three distinct places: input filtering in front of the model, output classification behind it, and tool permissions around the actions. Most incidents are one of them missing while the other two look fine. The third is the one that matters and the one almost always missing. ### Safety classifier https://lattice.kkshah2005.workers.dev/glossary/llama-guard A model, distinct from the one you are serving, that labels content as safe or unsafe. A probabilistic control, which means the error rate is multiplicative with volume: 1% false negatives on a million requests is ten thousand leaks. Treat the rate as a design input and layer independent checks on the actions that matter. ### PII redaction https://lattice.kkshah2005.workers.dev/glossary/pii-redaction Detecting and masking personal data before it enters a system you do not control. Redact at write time, not at read time. A trace store holding unredacted personal data is a liability that grows with every request, and you cannot retroactively redact what you already logged. ### Few-shot prompting https://lattice.kkshah2005.workers.dev/glossary/few-shot Including worked examples in the prompt instead of only describing what you want. The single highest-leverage prompting technique, and routinely skipped in favour of another paragraph of instructions. An example conveys format, tone and edge-case handling simultaneously, in a way that prose specifications approach but do not reach. ### System prompt https://lattice.kkshah2005.workers.dev/glossary/system-prompt The instruction block that frames a conversation, separate from the user's messages. Two things about it are easy to get wrong. It is usually where you put behaviour instructions that belong in a validation function, where they will be followed unreliably. And if it is stable, splitting it out and reusing it verbatim enables prefix caching, which is often the single largest cost win available. ### Golden set https://lattice.kkshah2005.workers.dev/glossary/golden-set A small, curated set of real cases, with expected behaviour written down, used to tell whether a change helped. Twenty examples taken from real production traffic beats a thousand invented ones, because twenty real failures contain more information about your actual distribution. A golden set that is never updated becomes a benchmark you have overfitted to — you will improve your score and ship a regression. ### LLM-as-judge https://lattice.kkshah2005.workers.dev/glossary/llm-as-judge Using a model to score output against a written rubric, in place of a human. It works, with two caveats that are frequently ignored. Position bias: the same rubric graded with the answer and the reference in swapped order produces different scores, so randomise it. And the judge must be stronger than what it is judging — a weak judge grading a strong model measures the judge. ### Regression gate https://lattice.kkshah2005.workers.dev/glossary/eval-gate A check in CI that blocks a merge when a change makes a scored dimension worse. The thing that makes evals an asset rather than a dashboard. Without it you can still debug with traces; you just cannot tell whether anything you changed helped. Put it in CI on every commit and let it fail — a failing eval is the system telling you something true. ### Shadow traffic https://lattice.kkshah2005.workers.dev/glossary/shadow-traffic Replaying a sample of production requests against a candidate model without serving the answers. The only way to get a real answer to 'is the new model better' before you switch, and much cheaper than finding out from users. The limitation is that it measures agreement, not satisfaction — you still need an outcome signal. ### Cost per successful outcome https://lattice.kkshah2005.workers.dev/glossary/cost-per-outcome Total spend divided by the count of requests that passed whatever quality bar you can define. The metric that inverts priorities. Work that fails often and retries becomes visibly expensive; long contexts that do not improve accuracy become visibly wasteful; model calls made 'just in case' stop being invisible. Tokens are an input to this, not the number itself. ### Open weight vs open source https://lattice.kkshah2005.workers.dev/glossary/open-weight Open weights means you can download and run the model. Open source additionally means the training data, training code and licence permit modification and redistribution. The distinction matters commercially and legally, and most tools called 'open source' are only open-weight. It is also the distinction that decides whether you can fine-tune it, host it for customers, or audit what was trained into it. ### Mixture of experts https://lattice.kkshah2005.workers.dev/glossary/mixture-of-experts A model with many parameter blocks, of which only a small number are activated per token. It decouples parameter count from per-token compute, which is how you get a large model's quality at a small model's cost. The practical consequences are operational: the weights are large even though compute is not, and serving needs to be expert-aware. ### Hallucination https://lattice.kkshah2005.workers.dev/glossary/hallucination Fluent, confident output that is not supported by the input or by evidence. Usually discussed as a model property. In practice it is overwhelmingly a *retrieval* and *verification* problem: the model was not given the fact, or was not asked to check. Constrained decoding, citation requirements and abstention do more than any amount of prompting. ## Essays - [Choosing an inference runtime](https://lattice.kkshah2005.workers.dev/blog/choosing-an-inference-runtime) (2026-09-28) — vLLM, SGLang, TGI, llama.cpp and Triton solve different problems. How to pick the one that matches your traffic shape, without benchmarking yourself into a corner. - [The gateway is the product](https://lattice.kkshah2005.workers.dev/blog/the-gateway-is-the-product) (2026-09-26) — Every dependency you hide behind an internal interface is a dependency you have not really removed. On model routers, and why the indirection is where your users' reliability lives. - [Fix the ranking, not the prompt](https://lattice.kkshah2005.workers.dev/blog/fix-the-ranking-not-the-prompt) (2026-09-24) — Most RAG failures are retrieval failures wearing a hat. How to diagnose whether your system is retrieving badly or generating badly, and the three changes with the best ratio of impact to effort. - [Evals are the asset, not the logging](https://lattice.kkshah2005.workers.dev/blog/evals-are-the-asset) (2026-09-22) — Why a twenty-example golden set beats a month of tracing, how to build one from production failures, and the mistake of holding out a set you never change. - [Agents need a durable host](https://lattice.kkshah2005.workers.dev/blog/agents-need-a-durable-host) (2026-09-20) — An agent loop is easy. Surviving a deploy mid-loop, a provider timeout, or a runaway token spend is the engineering — and it belongs in the host, not the prompt. - [Prompt, or fine-tune?](https://lattice.kkshah2005.workers.dev/blog/prompt-or-finetune) (2026-09-18) — Fine-tuning is a manufacturing step, not a debugging step. A decision framework for when prompting has genuinely plateaued — and when the real problem is retrieval. - [Where guardrails belong](https://lattice.kkshah2005.workers.dev/blog/where-guardrails-belong) (2026-09-16) — Input filtering, output filtering, and tool permissions are three different controls placed in three different places. Putting them all in the wrong layer is why safety work feels ineffective. - [Context is a budget, not a container](https://lattice.kkshah2005.workers.dev/blog/context-is-a-budget) (2026-09-14) — Why more context reliably makes models worse, how attention degrades in the middle of a long window, and the three techniques that keep long-running work coherent. - [Build a cost model before you optimise](https://lattice.kkshah2005.workers.dev/blog/the-cost-model) (2026-09-12) — Token pricing is the least interesting part of LLM cost. How to attribute spend to features, find the lever that actually moves the number, and stop guessing. - [Choosing a model is a routing problem](https://lattice.kkshah2005.workers.dev/blog/choosing-a-model) (2026-09-10) — Why 'which model is best' is the wrong question, how to build a routing policy you can defend, and why escalating on failure is a trap. - [Observability is not logging](https://lattice.kkshah2005.workers.dev/blog/observability-is-not-logging) (2026-09-08) — What to capture on an LLM request, what to leave out, and how to build a trace that can actually answer a question six weeks later.