28 minutes ago · 9 min read1898 words · Tech · hide · 0 comments

## I found this bug on an invoice I shipped this one, and I didn't catch it in code review, in staging, or in production monitoring. I caught it when I looked at a bill that was larger than I expected. Nothing about this failure is loud. Your app keeps working, your tests keep passing, your responses keep coming back correct - you just quietly stop getting the discount you think you're getting. The setup: I was building an internal assistant at [Craftwork](https://craftwork.com), and I was having a great time adding features to help our team. What I didn't realize was that each chat was costing way more than it needed to. ## What prompt caching actually does First, the fact everything here rests on: **LLM APIs are stateless.** The provider doesn't remember your conversation between calls. Every request ships the whole thing again - your system prompt, your tool definitions, and every message exchanged so far - and you're billed for all of it, every time. In other words: if you've got…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.