Tools and references for teams shipping AI in production
Working code, checklists, and decision templates — built from the same reliability problems solved directly inside real production systems, not written as generic advice. Start with what's free; go deeper where it's worth paying for.
Sentinel — an eval-as-MCP-server trust layer
A deterministic-first, judge-escalated scoring layer that catches fabricated citations and prompt-injection compliance before a response ships. Free to clone and run against your own agent — commercial license available for production use.
View on mleg.tech →The 50-Point AI Reliability Audit
A working self-assessment checklist across eight categories — evaluation, guardrails, grounding, observability, and more — for teams that want to find their own gaps before a customer does.
The RAG Architecture Decision Template
Nine architecture decisions that determine whether a retrieval pipeline actually works — chunking, embedding, retrieval strategy, and more — each with tradeoffs and a place to record your call.
Agent Orchestration Patterns
Five multi-agent coordination patterns — supervisor/worker, sequential handoff, and more — each with a diagram, when to use it, and the specific way it tends to break in production.
The AI Observability Maturity Model
Five levels, from no visibility into what your AI actually said, to a closed loop that catches quality drift before a customer does. Find your level in one page.