Notes from the AI production layer.

Engineering patterns for agents, models, traces, evaluation, cost control, and reliability.

Aug 8, 2026

Tracing AI agents across models and tools

A practical trace model for understanding multi-step agent runs without losing the relationship between user intent and production outcomes.

Jul 24, 2026

A practical rollback strategy for agent releases

How to version code, prompts, tools, and models together so teams can stop an AI regression without guessing what changed.

Jul 10, 2026

How to attribute token cost by customer and feature

A clean metadata model for turning provider usage into product-level unit economics and actionable budgets.

Jun 26, 2026

Reliability objectives for non-deterministic systems

A framework for combining conventional service health with outcome-based signals for production AI systems.

Jun 12, 2026

Choosing what to log without storing sensitive prompts

How to preserve useful AI production telemetry with deliberate capture, redaction, sampling, and retention controls.

May 29, 2026

Canary releases for model and prompt changes

A release workflow for testing AI behavior on production traffic while limiting risk and preserving a clean rollback path.

Engineering brief

Production AI, without the noise.

A concise monthly note on deployment patterns, observability, cost, and reliability.