AISeptember 30, 2026by Formation
Prompt injection in action-taking AI agents is an authorization problem. Learn how to enforce tool permissions, scope credentials, require approvals, and limit blast radius.
Engineering ResourcesSeptember 29, 2026by Formation
An LLM judge can help evaluate production AI output, but its scores need validation. Learn when to use deterministic checks, how to sample and calibrate judgments, and why judge changes require an overlap period.
AISeptember 28, 2026by Formation
A fast, well-formed LLM response can still be wrong. Learn how traces, offline evaluation, and sampled online scoring expose failures that latency and error-rate dashboards miss.
AISeptember 24, 2026by Formation
MCP is a boundary to design against, not a framework to adopt. Learn how tools, resources, authorization, and deployment choices shape reliable AI-agent integrations.
AISeptember 23, 2026by Formation
Model-declared completion is not a reliable termination condition. Learn how iteration, token, time, and spend bounds control agent loops—and what to do when a limit fires.
Engineering ResourcesSeptember 22, 2026by Formation
Agent tool retries create an at-least-once execution problem for side effects. Use durable idempotency keys, outcome reconciliation, and confirmation gates to prevent duplicate work.
AISeptember 21, 2026by Formation
Start with one agent, then split only when tool overload, context exhaustion, or parallel subtasks deliver a measurable benefit over the coordination, latency, and cost.
AISeptember 17, 2026by Formation
Prompt prefix caching reuses identical leading tokens; semantic response caching can substitute an answer for a merely similar question. Learn the validity checks, cost model, and routing tradeoffs behind each.
AISeptember 15, 2026by Formation
Long-context capacity is bounded by the KV cache: every resident token consumes finite GPU memory. Learn how to estimate the concurrency ceiling and monitor the latency signals that reveal cache pressure.
AISeptember 15, 2026by Formation
Streamed LLM responses need separate targets for time to first token and time per output token. Learn how workload shape, batching, deadlines, and cost determine the right SLOs.