My experience and what I’m concerned about is that eventually more time is spent refactoring to enable new features than writing the features themselves. In my experience this becomes worse with time as the models are reluctant to remove behavior, so the code base will grow to support some version of the old behavior together with the new behavior.
Hats off to my team, they did review the code. And it was embarrassing. I could not answer any questions asked by my team without looking at the diff. It was supposed to be "my changes", as we have agreed to the team. You can do everything with AI, but you own the changes. I did not. Why do we need to remove duplicates and call GraphQL in batches of 25 IDs when it's for 3 items displayed on the page? Did I have that much distrust in our Ops team to not believe that they can choose 3 unique IDs without having duplicates and know the difference between 3 and 25?
Use AI for anything you want but make sure you own the changes. Otherwise, you're just a meat proxy. Don't be a meat proxy like me. Be better. This world needs you more than ever to own it.
I read big claims, no evidence.
"AI code is non-deterministic and has risk. But human code is non-deterministic too, and it has the exact same risk."
No, that's not true. LLM-written code has very different risks, like completely misunderstanding the requirements, adding in hallucinated features, and losing sight of what the codebase actually does (massive tech debt).
I should ask, who writes the unit tests? And how do you write the unit tests beforehand? What if you need to prototype in order to figure out the shape of the API? The setup used here is so far removed from any software safety or software quality concerns, it would be hilarious if it weren't sad.
Stop posting marketing content void of any real information, please.
For example, recently I’ve been adding skills for our team. I added evals. If you ask an agent to make evals it’ll always try to bias and overfit just to make tests pass. No matter if I have it stated in every possible place not to do it and Codex instructed to catch these cases. It doesn’t work. Lots of sloppy evals get written and skills get extremely specific instructions to pass specific tests.
I think if you are implementing trivial features in a green field product it can work to some degree for some time. But eventually it’s going to deteriorate into a mess. Death by 1000 paper cuts.
Then I ran into a few bugs and, as I dug in, what initially looked like reasonable code suddenly seemed strange when looking closer. After some help from the agent to understand the intent, it became clear the design was poor, explained a few failures, and resulted in an order of magnitude more (reasonable looking) code to compensate which now I have to sift through.
To make matters worse, I could not get the agent to divorce itself from its wrong decisions, even with my explicit instruction detailing how the code should read. It would keep rewriting the same bad ideas then go into other distracting tangents which I'd have to correct. It felt just like how conversations in ChatGPT that would get stuck once it was in the context making me imagine we're still driving the same car but with a new paint job.
I finally wrote the code myself. It was more helpful after that. Hopefully, I'm able to resist the temptation and stay vigilant in my reviews. For context, I am using Codex Astra XHigh but maybe Opus 5.5 really is better?
Definitely worth spending more time on the plans & auditing for quality. Once the plan is ready, I find that adherence to plan is quite solid.
Another angle I look at it is that the code is going to be imperfect regardless of review. The main criteria is whether the code is producing the business outcome that you are looking for. Does it fulfil the user stories? Is it performant? Do you have enough pre-release checks and balances to make sure its not going to cause issues.
That's a real business problem, not a coding task.
Found the key smoking gun. It's load-bearing.