
Latest issue
Hallucination Detection Without a Skeptic Node Is Just Hope
9/14/2026
The support ticket arrives at 11am on a Tuesday. A user ran your financial summary agent overnight. The output is beautifully formatted, confidently written, and built on a company filing that doesn't exist. The researcher agent hallucinated a URL. The parser didn't complain — th…
Recent posts
You've got two prompt variants. You run them against a traffic split. Variant B wins on your quality metric. You ship it. Three days later, something unrelated starts degrading — a…
You've instrumented TTFT. You're tracking it at p99. You've even read the post I wrote about why p99 TTFT is the number that saves you at 2am. Good. Now here's the uncomfortable fo…
The support assistant analytics are damning in the best way. Out of roughly 400,000 distinct user queries, the top 500 phrasings account for 30% of volume, and another 25% clusters…
The support ticket comes in on a Tuesday. A user got a confident, well-formatted answer from your AI assistant — complete with a citation to a policy document that doesn't exist. Y…
The Cursor incident is the clearest case study in what happens when you price a flat plan on top of a cost structure that isn't flat. A developer posted a $7,225 invoice from a sin…
A team runs their LLM feature for a full quarter. The dashboard stays green — p50 latency at 1.9 seconds against a 2.5-second target. Then someone actually looks at churn data and…
A team ships semantic caching. Cache hit rates climb to 30%. Costs drop. Everyone's happy — until a user asks "what's our refund policy?" and gets an answer that was accurate six m…
The 3am call comes in. Your AI feature is returning 500s. You check the logs, confirm it's an OpenAI outage, and feel briefly smug — you built a fallback chain six months ago. Then…
The Monday morning bill shock is a rite of passage. You shipped the feature Friday, traffic came in over the weekend, and now you're staring at a number that's 3–5× what your back-…
Four seconds. That's how long a user stares at a blank screen before they decide your AI feature is broken. Not slow — broken. They don't know about prefill phases or KV cache tran…
The incident is always the same. A customer asks "where's my delivery?" and gets a confident, detailed answer — about someone else's order. Two users asked semantically similar que…
The April incidents should have been a wake-up call. A ten-hour Claude outage on April 6 and a major OpenAI platform outage on April 20 — both multi-hour, both affecting production…
The support ticket reads: "The AI feels slow and dumb lately." No stack trace. No error code. Just a user who noticed something your dashboards didn't. This is the failure mode tha…
The invoice arrives and the number is wrong — not wrong as in fraudulent, wrong as in useless. It tells you what you spent across the whole month. It doesn't tell you that one agen…
The incident report from a real Friday-night production failure reads like a horror story in slow motion: a customer support agent launched Monday, by Friday the inbox was full of…
The incident report always reads the same way. The LLM cited a policy. The policy didn't exist. What actually happened: the chunker split two adjacent sections at an arbitrary boun…
Most teams treat the local-vs-API decision as a one-time architecture call. Pick a side, commit, move on. That's the wrong frame — and it's why so many teams end up either hemorrha…
The demo works. The agent researches a company, drafts a personalized email, and the team ships it. Three weeks later, you're getting paged because the agent is stuck in a retry lo…
The incident is always the same. Someone makes a small prompt edit — two lines, maybe a single character — and three days later you're manually tracing why a specific customer's ou…


















