Watch out for cache read costs 0 ▲ Martin Alderson 2 hours ago · Tech · hide · 0 comments I know I'm guilty of just scanning OpenRouter's pricing tables and looking at input and output costs per million token. I've realised that's the wrong number to be focused on these days and cache read costs are actually far more important. Most of your spend is likely cache reads If you're running agentic workloads, cache reads are almost certainly the biggest driver of costs. Since we've got much longer context windows, you probably need to update your mental maths to take into account what this does to pricing. To take a hypothetical agentic session starting at 60k context length, with each tool call resulting in 500 tokens written and 5,000 tokens read, after 20 turns we get something like this:[1] Model Cache reads Fresh input Output Total DeepSeek V4-Flash $0.01 (18.4%) $0.02 (72.8%) $0.00 (8.8%) $0.03 Claude Opus 5 $1.04 (44.9%) $1.03 (44.3%) $0.25 (10.8%) $2.32 GPT 5.6 Sol $1.04 (48.1%) $0.82 (38.0%) $0.30 (13.9%) $2.16 Cache reads are nearly half the bill. Now look what… No comments yet. Log in to reply on the Fediverse. Comments will appear here.