9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.
(Note that it wasn’t all Opus 5.5; I have a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes, etc.)
I'm on a $20 plan and it never auto resumes. I have to go back in and type out resume or click a button.
We are seeing with OpenAI, allegedly through their new pricing scheme, as intelligence and model efficiency increases they offer the same throughput while advertising 1/2 as much usage, letting Astra consume more usage, essentially only being available to those wealthy enough to afford it while still offering essentially unlimited Sol and Luna to their subscription tiers.
Also if you're cache hit rate is high enough a billion tokens tokens from Deepseek 4.1 Flash costs less than $15.
Inference is highly profitable business, even for third parties with much less resources and expertise.
Suppose everyone starts moving faster thanks to LLMs and it becomes an expectation to use them. Budgets aren't infinite, so one of the two has to happen:
1. People get laid off.
2. Costs are shifted onto employees - either through lower salaries or having them bring their own subscriptions. I don't even make $500 a day!
Example from 15 years ago: https://stackoverflow.com/questions/7335920/what-specificall...
CPU time might go up while wall clock time goes down
Pros know these are lower cost models.
In the meantime, I have cancelled my Anthropic subscription...
I have a simple test that I have been running iteratively across the SOTA models from several vendors, including one Chinese vendor.
I start with some code produced by an Anthropic SOTA model...let’s call that Code A. Then I get Code B and Code C for the same task from models by two other vendors.
Then I ask each model to review and critique the other proposals.
By the end, both the Anthropic model and I usually run out of arguments... against them and agree that proposals B and C are better.
Claude then always asks whether it can incorporate the code or ideas from B and C into its own solution...
Nobody in his right mind will use a Chinese clone when you have models like Opus 5.5 for peanuts.
let a few valley elites decide how humanity can use this technology
open and transparent is the way, China is showing how
there's no money long term in being a token vendor
The biggest revelation from using open weights, because the vendors offer most of them up, is how useful using multiple model families is. Regardless of open or closed, if you are only using one family like Ant or Oai models, you're leaving a lot on the table. A harness like OpenCode will enable you to use different models in one session, or more specifically different subtasks when doing long running teams.
I spin up “offices” for different projects, using a documentation heavy approach with procedures, policies, standards, and processes. Agent onboarding and orientation, etc. I usually have an engineer for each separate part, (one for a simulator to simulate the hardware, one for the user application, one for the data analysis and evaluation tool, one for the firmware on each type of device, one for schematic and board reviews, etc. ) then I’ll have an office manager in charge of policy and issue boards, agent rosters, etc, and a engineering governance agent that makes sure code is compliant and documentation / code is coherent before any merges. 6-15 agents in each office depending on the complexity of the task.
It sounds like open code could be pretty handy but it would nerf my Claude subscription (api!=subscription rates). Understanding my workflow, what models do you think might be suitable for those tasks outside of OAI and Anthropic?
Contact me at : prompt.plumber@gmail.com
Nobody AMERICAN in his right mind will use... Wait, actually a lot of them will.
But for me, as a non american, non chinese person: I'll use whatever the fuck is the best and cheapest for my task, because that's how fucking Capitalism works.
If that means that a (proclaimed) "communist" country cleans the carpet with the self-proclaimed land of the free: so be it!
I’ve also used Opus 5.5 on some hill-climbing, and a lot more steering is required here, because … eval is hard.
"Why does Fable even exist" is a very very reasonable question right now.
I feel like instead of releasing fable, they should have released it as Opus 5, then their next Opus release they would call Sonnet, and their next Sonnet release they would have called Haiku. I don't know if their pricing structure would have been able to support that, but Anthropic has always been the least competitive regarding token pricing.
If they did what you suggested either they eat a ton of additional costs, or send a signal to the market that they’re increasing costs more generally.
Also Fable and Opus have different specialties so they really are best presented as different models
I assume Fable 5.1 is the first re-tuned (is there a better term for this?) version of Fable 5 and Opus 5.5 is the first re-tuned version of Opus 5, and maybe the Opus one just came out better for some reason?
No I do not. OpenAI has some hidden model that's apparently 5x or so better at certain benchmarks than GPT 6 but they're not and have no plans to release it. It is increasingly likely that these AI companies keep the best models for themselves and then release smaller, cheaper distilled models for everyone else, especially since the AI companies are vertically integrating into many fields.
Benchmarks do not measure the first aspect.
https://tfmbot.com is the link (discord and source links on the splash screen).
The results are fucking incredible to the point where people in discord are stating "I'm surprised this is working so well". I am too.
I feel like there's a group online that missed the boat. Anything negative towards AI capabilities is still upvoted but I've been in the industry for over 25years, highly respected and can't fathom the "AI dumb lololol" type of comments i see on HN. AI is superseeding all other ways to develop.
AI writes clean code and can do so in a very maintainable way honestly.
Because I noticed people at my work did similar requests to improve CI. The result was a faster CI, but full of cludges huge inline bash scripts in workflow YAML files, and effectively unmaintainable, unreviewable mess. After just a few rounds of these optimizations the entire CI setup is basically a Rube Goldberg machine but made of duct tape.
I like to watch it work though because it honestly teaches me some tricks.
Now we are getting downgraded models that do 100x COT because it's cheaper.
I've given it some big tasks and asked it to parallelize as much as possible etc.
It did burn through my weekly tokens in about a day (20x max), but the output was completely on point. (I knew there was a "reset token usage - opus 5.5" button in my account.)
I've now come to a point where I even delegate my discovery for new features to it.
You still need to give it methodologies though to get the proper output, but the outcome is way beyond what I would be able to realize with a team of 5 in a month.
It not: who becomes rich on all that productivity?
So asking it to do things on its own for a long time? Given last week, absolutely not.
"Claude, I released it myself, its up there, just analyze the logs"
"Ok, I'll analyze the logs but it isnt" -crunches for a while- "the issues aren't fixed, but that's because the new version isn't up there"
I think I yelled at it one more time about how I know what was released before "we" figured out that the last release had failed in a way our release system reported as success, but was crash looping on start up and so the old version was still around and working as back up.
Sorry claude.
Now fix that release status check.
I asked it to just summarize a repo with only a README.md file containing a poem, and it began to interact with a remote server, solved math questions and finally executed untrusted code. Too independent to be trusted.
My goal is to research how models can still be confused via tool responses only. Something the labs claim to have "solved".
Additionally, the "auto-mode" / "auto-review" modes have been released to use harnesses "safely" even without strong sandbox. And these modes use ... a second LLM.
What sort of workloads do well with these long tasks? The big labs are optimizing for long run time on their own, but it seems like a terrible thing to optimize on unless you're trying to do something like prove a hard math theorem, which success is clearly defined and the route doesn't matter a ton.
Plan mode has been made increasingly useless. I need to discuss to iterate to get the desired design, explore options, because Claude never gets it right first try and I don't have enough knowledge of options to specify everything up front.
Ah well, the Chinese models will still work well, I guess.
0. https://github.com/mattpocock/skills/blob/main/skills/produc...
I've had it running 8h+ of non-stop optimizations, chasing a performance target, rewriting systems or building a series of prototypes for research. All it needs is a clear goal.
In a few cases it asked me to check some subcircuits and some component values because it couldn't read it right. So instead of just making things up it deferred to me.
It also ran tons of small simulation experiments while doing this to verify claims from the service manual, like that the RC filter it had read off the schematics actually had a cutoff frequency that was sensible in relation to some bandwidth number in the manual.
I had uploaded datasheet PDFs for many of the ICs and it used those to cross-reference and validate.
It kept on working for over an hour. When it asked for the manual verification, I described circuit connections in words, like "from pin 3 on IC 2 there's a series resistor of 3k in parallel with a 10 pF capacitor, it then connects to a 18k resistor to ground, a reverse-biased diode to ground, and then finally into pin 6 of IC 4", and it correctly understood the topology in all the cases. Sometimes it asked me to check again because it though something was off, and indeed I had mis-read the schematics.
I also provided reference articles on the underlying theory. Scannded stuff from the 40s and 50s. It correctly read the equations and cross-validated them across papers, and even caught several typos along the way.
I barely had to do anything apart from providing the PDFs and some occasional manual schematic interpretation.
Claude 5.5 on High. Burned through about 50% of my weekly $20 subscription usage, but I didn't try to optimize much.
I did use Sonnet 5.5 Medium on some datasheets and it also did very well on the extraction, but did have to correct itself more often on the conclusions.
the company is run by holier-than-thou, we know what's best... who apparently don't read claude's output and blindly trust it
the mythos "hacking" of the linux kernel, as finally told from the linux side, is eye opening
It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.
It's also way more able to execute subagent tasks all at once than GPT 6.1 I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.
IMO it should have the option to observe background tasks with a small model to make sure it's working correctly.
Edit: It just happened again, something that was actually completed didn't exit and it waited 30min for a timeout.
I made it and been using it every day
The notable difference to me is tokens/sec are still much higher on 6.1 Sol
This is also the reason why I'm swapping over to Anthropic for a month. 6.1 Sol seems good, but is unbearably slow even with 6.1 being more token efficient than 5.5.
They did speed it up in the last few days [1], but I don't find it to be enough.
What I mean is, for most tasks I do that aren't trivial, by the time I have defined what DONE means, I would've already did the work and walked the path to get there, which is what I would've hoped to not have to do in the first place.
Yes this has held true on Opus 5.5. I checked. It’s a massively better model, peer to Fable but with different strengths and weaknesses. But it still has this issue. Which to be fair, people do too. Planning is a learned skill.
I think what they’re saying is that the harness no longer uses a text search on “think” to engage reasoning modes. Fair, that’s good to know. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.
I used to have the AI write a planning note with checklists, but this seems good enough nowadays.
This is their "Why it matters" for why asking opus to use subagents matters.
The fact that they think "some other people jumped off the bridge" is a reason to do something does not inspire confidence.
This has zero mention of the success or quality of those efforts. Just because they were done "with little oversight" doesn't mean it went well...
I researched the feasibility of such task about a half year ago and concluded AI wouldn't be able to have good enough spatial and blueprint knowledge, unless you were willing to throw unreasonable amount of money at the task.
[1] : https://www.salahadawi.com/hacker-news-ai-detector/49946567
You’ll get “something” but what you get is certainly not going to be what you wanted.
Look, its not complicated:
1) be precise in what you want
2) have a feedback loop to verify it.
3) check in often, not once a day.
> Migrate the payment endpoints from the old client to the new one. Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes.
What end point? What test suite? What is a client?
Maaaaybe the model can infer it from your code, but look, if a human taking that jira ticket would go: what does this mean? …then your model (yes, even 5.5) is going to make a bunch of assumptions.
Correct assumptions? Maybe. Maybe not. …but do you really want to find that out hours later?
Just say:
> Plan out the following as a set high level tasks: …
Then, review the plan and tell it to execute the plan, maybe something like; after each step, ensure the code base compiles and the tests pass.
Then, come back after one hour and make sure it’s on the right track.
Poorly specified long running tasks is a recipe for “rollback all those changes…”
Do these tips work? Probably. Do we know why they work? Sorta. Could we be just copy/pasting whatever text someone else kinda thought worked one time and now enters every time out of some weird feeling that it helps nudge the LLM, like some remote islanders building airport control towers out of bamboo in the hopes of a new airdrop from the gods?? Totally. But here we are.
Write a “hand off” document? Prohibited. Focus on something else? Prohibited. Eventually every single response was terminated by the classifier before it even began, with the reason being a placeholder string like <thinking quote> or similar. Before it totally forbade any form of response, it agreed with me that it was unfortunate, but stood proudly by its guns.