In the first week of July, I spent over $5,000 coding through one company's API.
Best model. Running all day. Task after task.
I didn't blink at it once.
It was only afterward that I felt stupid.
I looked at the number and couldn't tell you what $5,000 of work I actually got for it.
The return wasn't obvious. The bill was.
Here is what it bought, in the only unit the dashboard actually measures: 1.91 billion tokens in, 20.2 million tokens out, and 397 web searches. Nearly all of it inside eight days.

Look at the right-hand side of that chart. That flat stretch from July 9 onward is not an outage.
That's me changing my mind.
I can tell you the tokens to the digit. I could not tell you what they were for.
I've written about that week on our company blog, because it's the reason we built our own coding tool.
This is the other half of that story - the half I haven't shown anyone.
Starting this month, I'm publishing what we actually spend on AI. Every month. Every line, including the ones that make me look bad.
Why I'm publishing our bill
Every founder I talk to is guessing.
They're guessing because the people writing about AI costs are mostly not paying them. They're analysts, or they're selling you a course, or they're a vendor with a pricing page to defend.
The rest of us compare notes in DMs like it's something to be embarrassed about.
I have nothing to sell you here. This site has no paywall and never will, for reasons I've written about already.
And we're three people. If a spending decision doesn't survive at our size, it's not advice - it's a press release from someone with a different balance sheet.
So here's ours.
The most expensive line was not the one that looked expensive
We were paying about $5,000 a month, per GPU, to serve our own small models to a small number of users.
Read that again, because I had to.
Per GPU. Per month. To serve models that a $600-a-month cloud endpoint could have served just as well.
Models that, honestly, your own laptop could have served.
That wasn't a mission. That was ego, dressed up as infrastructure.
Nobody made that decision in a room. It accumulated. We wanted to be the company that ran its own models, so we ran our own models, and the invoice was just the cost of being who we said we were.
The users didn't care. There weren't many of them, and the ones there were couldn't tell you which GPU answered them.
We killed it. Those GPUs now do what they're actually for, which is training the next models.
If there's one thing to take from this month, it's that the line item that hurt us most was the one we were proudest of. It never showed up as a problem because we never framed it as a cost. We framed it as an identity.
Go look at your own bill for the line you'd defend in an argument. That's usually the one.
What we did for the other three weeks
The $5,000 was one week. July had four.
After that first week I cut the metered API lane hard, moved most daily work back onto subscriptions, and moved the rest onto our own curated models.
The rest of the month cost a fraction of the first seven days of it.
I want to be careful here, because I can feel the shape of the post I could write next, and it isn't true.
I am not claiming we found a 99% saving. Those aren't the same workload. Some of that $5,000 bought real work I could not have gotten any other way, and our own tooling was in beta that month - it handled easy-to-mid tasks while the hardest problems still went to the expensive agents I was complaining about.
A clean percentage is the kind of number that gets a post shared. It would also be a lie, and next month I'd have to keep telling it.
Want the full playbook? I wrote a free 350+ page book on building without VC.
Read the free book·Online, free
Here's the smaller, truer version.
The problem was never the price per token. It was that metered billing has no ceiling, and I had quietly appointed myself the ceiling.
I never sat down and chose $5,000. I chose 'best model, don't think about it', forty times a day, for five days. The bill was the sum of forty small decisions I never actually made.
A subscription is worse value per token and I moved back to it anyway - because it has a ceiling built into it, and I don't. That's not a pricing argument. It's an argument about what I'm actually like at 1am with a deadline.
Our own curated models took the rest, and the reason is the same one we sell on: I can see what a task costs while it's running, not thirty days later.
That's the real failure mode with AI spend. Not overspending. Overspending without a single moment where you decided to.
Here is the other side of the same month - what our own curated models cost us over roughly the same window.

$214.56. 248 million tokens. 144 million of those served from cache.
Two things in that picture matter more than the total.
The first is the direction. Our spend on our own stack is going up through the month, because we kept moving more work onto it. That's the opposite of a cost-cutting story - we're doing more, not less.
The second is the cache line, which is most of why the number is small. That isn't a clever trick. It's what happens when the same context gets reused instead of re-sent, and it's available to anyone paying attention to how their tools actually work. Same person, same kind of work, six weeks after a $5,000 week.
What we're paying now
After all that, the standing bill is boring. Three lines.
Claude Max, the 20x plan, for the team. Flat, per seat, $200 a month at list.
ChatGPT Pro, for the team. Also flat, also per seat.
Our own inference, through Vinci. The $214.56 above.
That's it. That's the whole thing.
I want to sit on how unremarkable that list is, because a month ago I'd have told you our AI spend was a complicated problem that needed a strategy.
It wasn't. It needed a ceiling.
Two flat subscriptions and our own stack cover what a metered frontier API was doing at many times the price, for the work we actually do most days. The hardest problems still go to the expensive tools, and that's a deliberate line item rather than a leak.
There's one line I'm not publishing: what we spend on GPUs now that they're training rather than serving. That's live R&D, and putting a number on it tells people outside the company more about what we're building than I want them to know yet.
I'd rather tell you there's a number I'm withholding than quietly present an incomplete list as a complete one.
The rule for this series
Every number in these posts traces to an actual invoice.
If I can't source it, it doesn't appear. If a number goes up, it goes in. If I made a bad call, that goes in too - the GPU line above cost us real money for months and I was the one who wanted it.
I'm not going to publish a chart that flatters us. There is no version of this series where we spend more and I quietly skip a month.
That's an easy promise to make in the first issue. Ask me in December.
Next month
August is the first month with a full set of clean lines, so the comparison starts then.
If you're running a small team and paying for this stuff, I'd genuinely like to know what your bill looks like - not to benchmark, just so fewer of us are guessing.
Reply, or find me on LinkedIn.
The number nobody publishes is the number everybody worries about.

