← Notebook · 37 essays

I Preferred Codex, but I Trust Claude Code

4 min read

I pay Anthropic about £180 a month for Claude Code. For a while I also paid OpenAI, because Codex started as my second opinion and ended up a candidate to replace the first. I still use Claude Code. What follows is one person's account of how that happened, and it is worth exactly as much as one person's experience is worth.

Codex started as the critic

I did not set out to switch. I gave Codex a narrow job: take what Claude Code had built and try to break it. Read it cold, challenge the architecture, find the assumptions nobody had written down.

That is a harder test than asking two models the same puzzle. Codex had to find something a capable tool had already missed. The more I used it, the more I came to prefer what it found. That is a judgement, not a measurement, and I kept no scorecard. I had expected a useful second opinion. I had not expected it to make me wonder which of the two tools should be doing the building.

So I began moving two real projects across, and bought OpenAI's cheaper Pro plan as the trial. Then I reached the limit.

Two walls

Heavy inference costs money and I was using a lot of it, so I went to buy the bigger tier. OpenAI's Help Centre lists the $100 plan at five times the usage of Plus and the $200 plan at twenty times. Four times the allowance, for roughly the £180 I already pay Anthropic.

On 10 September OpenAI paused new sign-ups and upgrades to that tier, citing demand for its new Astra model (Fortune and CIO both reported it). Existing subscribers kept theirs. I was not one. I do not read the pause as carelessness. A company that cannot serve everyone has to choose who waits, and this one chose to protect the people already inside.

OpenAI sells credits for exactly this situation. Its documentation describes them as pay-as-you-go use beyond the plan's limits, and says they do not raise the plan's allowance. I bought the equivalent of the full month on the 20x Pro plan. At the work I was doing, they lasted just short of two days.

Two days is a thin sample. Extended over twenty working days it would be £1,000 plus on top of the subscription, which may not be the full figure, but was enough to stop me planning around it.

Credits suit a deadline or a one-off migration. They are a poor foundation for an ordinary working day, and an ordinary working day was what I was trying to buy.

What the meter did to the work

The bill was not the interesting part. My behaviour was.

The point of an agent is to hand over something large and slightly vague. Understand this subsystem, find what is actually wrong, fix it. Under a meter I caught myself trimming the instruction before I sent it, then asking whether the task justified what was left of the allowance. That is not delegating. I had started managing the tool instead of the project. Every instruction now carried a small tax, the effort of deciding whether it was worth sending, and the tax fell due before any work began.

The cost of a task is unknowable in advance. Something I expect to take ten minutes may take an hour if the agent finds a problem on the way, and I cannot turn "fix this properly" into a price before I send it. A ceiling I can see lets me plan around it. A charge I only discover afterwards makes me cautious beforehand.

There was a plainer problem as well. Across that stretch I was shut out of Codex for about a week, and the projects did not pause with me.

Contractors have charged me more for less, but they had the manners to quote first.

Back to Claude Code

So the work went back to Claude Code, and the reason was dull. Every morning it would be there, at a price I could predict, and I had months of practice at working inside its limits. I keep its context small with two tools that actively manage the token cost, and I know roughly what a day of work costs it. Nothing about that meter surprises me any more.

That last part is a bias I should declare. I already have habits, configuration and workarounds built around Claude Code, which is a switching cost in its favour, and I have not tried to net it out. One person, two projects and one short stretch is not a study.

It was enough to change what I think I am buying. I assumed it was a better model. Part of it is, the rest is arriving on a Tuesday morning to find the door still open.

Richard SutcliffeCTO at ThinkTribalfield notes on AI in regulated sectors