Prepared posts for the article
Stylized mockups of what the social media posts will look like on Twitter/X and LinkedIn.
Twitter / X
We spent 3 months building the AI for our product. Then beta users tried it and their feedback was: βItβs not smart.β That stung, but they were right. Here's what we did to recover

Our beta users had one clear piece of feedback about the AI assistant I spent three months building: "it's not smart." That stung. We picked the cheaper model because we were a startup and trying to be careful with money. Then I built a whole harness around it which meant thousands of lines of orchestration code, classifiers, controllers, and prompt scaffolding. All in an attempt to squeeze better performance out of a model that just wasn't strong enough. In testing, the failures felt occasional. After launch, users hit them constantly. What I learned: β Model prices drop. The code you write around a weak model stays with you. β A few tools with typed schemas can burn 2,000 tokens before the model has even started solving the problem. β In agent systems, every step drags context forward, so cost snowballs. We switched to the better model and the improvement was immediate. Then the price dropped from $5 to $3 per million tokens anyway. The irony is that the 2,000 lines of harness code I wrote are still sitting there in the codebase. What actually worked: 1. Start with the smartest model you can afford 2. Let the model do more of the driving 3. Be ruthless about context bloat 4. Write evals before your users expose the gaps for you Full write-up with code π

The Most Expensive Cheap Model I Ever Used
How we stopped fighting our model and started building a product.