Writing
Arguments, mostly.
Agent governance, AI regulation, evaluation and the engineering underneath. Written to be disagreed with rather than to rank for anything.
-
Structured Output Isn't Reliable Output
JSON mode, function calling, constrained decoding - these give you schema compliance, not semantic reliability. Your output can be perfectly valid JSON and ...
-
The Insurance Industry's AI Blind Spot: Claims Automation Without Trust Infrastructure
Insurance companies are racing to automate claims with AI. Nobody's built for the regulator, the litigant, or the appeals board. The blind spot isn't capabil...
-
What Moltbook Reveals About Multi-Agent Trust at Scale
Moltbook isn't an enterprise product - but the vulnerabilities it exposes matter for any organization deploying multi-agent AI systems.
-
The Agent Watchtower, Part 5: Reference Architecture
A complete, implementable design for enterprise agent governance. Concrete specifications, integration patterns, and implementation roadmap.
-
The Agent Watchtower, Part 4: Economics of Agent Operations
The financial model for sustainable AI governance. Cost cascading, ROI-driven routing, and why governance pays for itself.
-
The Agent Watchtower, Part 3: The Autonomy Spectrum
How to balance business unit freedom with enterprise governance. Federated control, trust-based permissions, and why guardrails beat gates.
-
The Agent Watchtower, Part 2: Anatomy of an Agent Control Plane
The technical architecture for unified agent governance. Registry, observability, policy, and control - how to build the infrastructure that makes multi-clou...
-
The Agent Watchtower, Part 1: The Fragmentation Tax
Banks are deploying AI agents across AWS, Azure, GCP, and open-source frameworks. The result: governance blind spots, compliance nightmares, and a ticking re...
-
The Eval Crisis: Why Most Benchmarks Don't Matter
Your model scores 90% on MMLU. It still fails in production. The benchmarks everyone obsesses over measure the wrong things for enterprise AI.
-
5 Evals Every Production LLM Needs
Forget MMLU scores. These are the evaluations that actually predict whether your LLM will work in production.
-
The Real Reason Your RAG App Hallucinates (It's Not Chunking)
Everyone's optimizing chunk size and embedding models. The problem is upstream. Your data pipeline strips context before it ever reaches the vector store.
-
Your AI Architecture is Bleeding Money
Cost-per-token is the wrong metric. The real savings come from architectural decisions most teams get wrong.
-
The AI Production Readiness Checklist
The comprehensive checklist for launching LLM-powered features. Evaluation, monitoring, fallbacks, cost controls, and incident response.
-
Prompt Injection is an Unsolved Problem (Here's How to Mitigate Anyway)
There's no complete solution to prompt injection. Here's the defense-in-depth playbook for production AI systems.
-
When to Use Agents vs Deterministic Workflows: A Decision Framework
A concrete decision tree for when to reach for AI agents vs traditional orchestration. Cost, latency, reliability, and compliance dimensions.
-
From 11% to 88% GPU Utilization: How We Built 8x Faster LLM Inference
PyTorch leaves 89% of GPU bandwidth on the table. We fixed it with custom Triton kernels. Here's what we learned building Accelerate.
-
Eval Debt Will End Careers
Tech debt is slow. Eval debt is sudden. The teams that survive will treat evals like unit tests: written first, run always.
-
Agentic AI Is a Cost Center, Not a Strategy
Everyone's racing to deploy AI agents. Most will waste millions. The question isn't 'how do we use more AI?' - it's 'how do we use AI sustainably?'
-
AI Observability is Expensive Voyeurism
The observability market is selling you dashboards to watch your AI fail in high resolution. What you need is controllability.
-
Your Data Team is Building an AI Graveyard
Every transformation in your data pipeline destroys information AI needs. Traditional data engineering is a lossy compression algorithm.
-
Multi-Agent is This Decade's Microservices Mistake
The multi-agent hype will collapse. We learned this lesson with microservices. Distributed systems are hard.
-
The AI POC Trap: Why Your Demo Worked and Production Won't
Your agentic AI POC impressed leadership. Then you tried to scale it. Here's why demos deceive - and what production actually requires.
-
Foundation Models Are a Commodity. Act Accordingly.
Everyone's agonizing over Claude vs GPT vs Gemini. It doesn't matter. The differentiation is moving up the stack.
-
The LLM Evaluation Maturity Model: Where Does Your Team Actually Stand?
A six-level framework for assessing how your organization evaluates LLM outputs. From 'it looks right' to continuous evaluation pipelines with regression det...
-
The EU AI Act Is Here: What Financial Services Firms Need to Know
A practical guide to EU AI Act compliance for banks, insurers, and investment firms. What's required, what's high-risk, and how to prepare before enforcement...