Running a helper from the terminal, making Claude work in a working directory, and then create a .commit file has been my workaround for this for a while now.
Imagine there's a better solution nowadays, but this allows me to use dispatch building on Vercel, so I can check it out from wherever, without too much pain.
I’d love to just have Claude use my machine as a sandbox host instead of having to run RC on each host session. (In case you are listening Boris ;) ).
In the meantime I have a janky master RC session that creates new tmux windows and Claude RC sessions for each new code trajectory that I want to run.
The other benefit here is you can drop down and use termux to use Code directly if you hit a RC bug, I found permissions UX to be a bit flaky in the iOS UI.
But increasingly it seems like dispatch was slapped on top of cowork incrementally, when there was not an integrated and cohesive strategy across cowork between mobile/desktop/laptop. This is kind of what many of us get/got to learn in our 20's.
MacBook m1/m2 also are cheap enough now vs an Mac mini which I was surprised about, not too surprised but yeah..
Never really get good answers. There is no killer app. Just bikeshedding.
Why would I need an LLM to do this for me? That’s 5 minutes of work max, and doing it gets me in the flow of work again, to see what’s going on and needs to be done.
For the folks I talk to who use a LLM for this that seems to be the case. Takes a huge cognitive load off every morning and saves them an hour or two.
More or less a very expensive band aid over a bad work environment.
I kinda use it the same way in a sense. I have a little skill I run against our (horrible) task management system to summarize things and give me a punchlist to work through sorted by priority. This saves me thousands of clicks to do the same thing in the horrible web UI. A proper system in the first place would be a lot better!
At some point I’ll probably just take that to the next logical step and have the LLM write my own web interface to abstract and replace the horrible one entirely for me.
So that sort of thing is already normalized. An LLM is unlikely to be worse than a human. There is no accountability either way since everyone is failing at the task, it really doesn’t matter much if your or your bot misses a few percentage points of high priority items. You will probably be vastly outperforming your peers who are doing it manually.
I’m actually very time-poor, so figured it could help be clawed back time doing… what exactly?
- scanning logs for errors and
- opening issues which are then auto-triaged and
- PRs are opened for them and auto-reviewed and
- merged (and deployed).
This workflow alone is immensely powerful, and takes alot of burden off the team.
ITSM those unsupervised workflows are essentially an attempt at purported productivity in the near term at the expense of meaningful incremental long term burden for teams.
The only ostensible benefit is in the eyes of the AI-psychotic tinkerer, who knows no better, or in those of the clout-chasing developer farming likes on their LinkedIn posts.
Truly open minded
Then they have a starting off point to see if the agent was correct or not. If not, they lost maybe 2 minutes of reading. If yes, they can go "yea, push the fix and monitor", put down the Red Bull and go back to bed.
It's like your AI agent is just plugging the leaks in the dyke each time, instead of fixing the architecture of the dam.
They can use agents. Like, team members don't need to be replaced, they can simply use agents when they deem it useful. If they see a trivial bug,they can put their agent on it and go work on something else meanwhile.
famously a good job for a tool that takes 10-50k logs to run out of context and forget what it's doing.
1. On a blog that no one visits maybe?
2. It's called a grep
3. For bigger projects it's called sentry
The source of errors can be whatever you like. Sentry, grep, whatever. Its not the point. The point is that many of these errors are real issues, and can be fixed automatically and safely by agentic systems. It really saves time, and work by our team leaving them to actually concentrate on the thing that delivers real value.
But now it's suddenly not "we use LLMs to scan logs for errors", but "we use actual tools to find erots in logs, but then just hand them off to an LLM and hope for the best".
As in "let's repeat what @troupo said but pretend it somehow goes against what @troupo said" lol
There's lots of news about the billions AI companies spend on data center construction, but it feels like it's not even a fraction of the money they're spending on endless nonstop blogs about how great their app is at doing... things. Things that will never be defined.
Maybe I should have the agent also do a background check.
PS: This is a joke, but feel free to steal this idea.
https://github.com/Grigorij-Dudnik/TinderGPT
> TinderGPT automates the process of writing and arranging dates with girls on Tinder, enabling you to generate romantic meetings with almost zero effort. Your only role is to like the profiles that catch your eye. After that, TinderGPT comes into the play. It initiates a conversation with the girl, using details from her profile, continues by building an emotional bond and highlighting your attractive traits, and finishes by arranging a meeting and giving you a push-up on your phone with her number.
Then the openclaw WhatsApp module…
Kidding of course.
I don’t have time for much leisure coding these days. I do have time to kick off a few tasks in the morning to progress my many side projects. Nothing public / oss, just code that I find useful/interesting like home automation, content pipelines, games, etc.
There are a bunch of cases where remote control from iOS onto a Mac Mini is simply nicer than using iOS Claude Code sandboxes.
It’s the same pattern as you (hopefully) apply at $dayjob. If you are not defining a /goal and letting your agent crank you are not making full use of the models’ capabilities.
So I wouldn't agree that the agent should be cranking out code all the time, in fact that seems more like a waste of resources compared to the work it creates. But I do understand home automation software can be very one-off and simple. But then again, a properly programmed home automation suite doesn't need a SOTA model to modify it, I think.
- infinitely duplicate any and all code, helpers and components
- infinitely duplicate CSS (because they duplicate components)
- continuously write code like "read the entire db into memory and run a filter function on retrieved data"
- continuously write code like "call db with multiple queries for each element in a list"
- etc. etc.
Why the hell would I ever want to run them unsupervised?
If you put yourself in a position where you need more leverage (technical or operating) I think you might find you get some value.
I haven’t set up dispatch yet. I wonder what a Mac gets me over this set up if I don’t need iMessage
Please help: I wánt to need this!
Make sure to spell PERFECTLY in all caps.
- Fuzzing with the goal for it to apply domain-specific and source-informed knowledge to choose specific fuzzing approaches.
- More generally, any optimization problem that benefits from domain-specific or source informed knowledge.
- Running Microsoft's SkillOpt [0].
[0]: https://github.com/microsoft/SkillOptHere is a real use case: you are are responsible for some alerting channel. You have datadog/ cloud logging/ github all connected. You see a bunch of alerts come through while you are out and about and you prompt CC to investigate - Claude triages and says “all of the sudden you are getting time outs from this bank API your company partners with, this started an hour ago. It’s happening on ~15% of requests”. So you ping the guy at your company who does vendor relationships and go back to your weekend.
This is a non hypothetical example. Obviously it would be better if your job had a real on call rotation and more robust alerting and you wouldn’t be getting slack alerts on the weekend… but I take the approach this job affords me a lot of nice flexibility so it’s ok
But yeah it’s kinda a zone where most weekends there’s no problems so it’s not a huge priority… until it is
I'm an account manager. My clients will phone at almost any time, weekends included, if they feel there issue wasn't yet looked at by the on call dev.
I've been watching "How it's made" on Hulu to fall asleep at night.
I’m constantly surprised by how many things are made with human hands, despite the ability to automate.
This is something we of the HN bubble take for granted. Most of us know how to type quickly and use editors and use macros and program scripting languages and compose regexes.
The vast majority of programmers do not know those things. As such, AI speeds them up tremendously.
HN skews heavily toward the SF signature vc-hype tech-driven-dev style, always chasing the new thing, sometimes to the detriment of everything else.
Even if this style of development was a clear improvement over the classic "typing things with hands" style, the rest of the world would take a while to catch on.
How is Claude monitoring them for hours? Claude runs out of context and extremely long sessions are prohibitively expensive even according to Anthropic (after they dispense with the marketing bullshit of long running tasks)?
A single session running for multiple hours is prohibitively expensive, as per Anthropic. Regardless of whether it just waits for a prompt or does something.
Yup. They can't keep your workload in cache forever, or they would run out of cache for users.
> I wonder would starting new sessions and having to re-read the contexts and results anyway be any cheaper.
Yes, that's what they recommend
I don't value my travel time at all, but it used to be wasted on travelling.
It could be for a personal project or hobby.
Having independently running processes from the computer you carry around offers benefits.
(I wish I was joking)
Which tests and optimizations do you propose to run after a night of supervised work when one of main things that all agents keep doing is "load all records from db , and filter them in memory"? It's now become so bad, I had to literally vibecode a separate linter for this. And that's just one of the problems.
but we do have sufficient AI to make a great product out of a great prompt.
garbage in -> garbage out hasn't gone anywhere.
so: much like to anyone that blindly complains that their compiler hates them : if you actually want help, provide information. If you just want to complain that the compiler is mean, scream at the sky.
plenty of people have figured out how to get this to work; more than enough to confirm that a straight <gambling-machine>/<hallucinatory-psychopath>/<random-number-generator> analogy is too simplistic to explain what we're working with.
> plenty of people have figured out how to get this to work
Plenty of people claim they have figured it out. In reality these people are full of shit and assume that if LLMs can produce working software, it's great working software. And also assume that LoC is a measure of quality.
Because without fail all the models keep doing this: https://news.ycombinator.com/item?id=48962703
And you can only see that in your "great product" if you actually read the code and understand what's going on.
So I dunno what to say, except it’s possible to write really solid code with LLMs.
> I also read my code regularly
So you're literally doing what I am talking about.
You see, there's your problem right there. You're vibe coding, which by definition literally means you're unwilling to look at the generated code. That's not what successful ai assisted software developers are doing. YOU HAVE TO READ THE CODE. Refusing to do that means you're not a serious programmer, you're outsourcing your thought and design and implementation, trying to get something for nothing by taking the easy way out, and you're going to get terrible results no matter what prompts you "engineer". There ain't no such thing as a free lunch (yet).
And while we're at it, to elaborate on what serf said: people mindlessly parroting terms like "stochastic parrot" to criticize llms without having read the actual paper that coined the term and understanding what it really claimed and how other papers responded to it means you're just a human stochastic parrot no better than what you're criticizing -- at least the llm has read all those papers and understands what "stochastic parrot" actually means in context. Ask it, it will be glad to explain!
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? (FAccT 2021)
I guess you vibe-read what I wrote. Let me write it again for you: "I always have to correct its hallucinations during the day. Why would I ever let it run unsupervised overnight?"
It's right there on the label. If you didn't mean "vibe coded" then don't vibe write "vibecoded", and instead read what you wrote, then look up the meaning and spelling of the terms you are using, and do not neglect to correct yourself because it's not what you meant.
I have to correct typos like that all the time when I write them by hand. But making typos and having to correct them and having to look up the meaning and spelling of words does not make me complain about all human generated code and text just because it requires me to proofread what I wrote.
Reading what you wrote or what ai generated and looking up terms is all part of the process, because people and llms make mistakes. Vibe coding is not giving a shit about that, thinking it doesn't apply to you, and hitting "Accept All" without reading the code, by definition.
And if you can't accept that, then don't write or generate code or text, and don't complain about how you have to read what you and llms generate and correct it, and understand what words mean.
>I had to literally vibecode a separate linter for this
Why would you "literally vibecode" a linter and use it to review other llm generated code without reading the linter's code itself? That is "literally" vibe coding eating its own tail.
I take your emphatic use of the word "literally" to mean that you are "literally" using the well know definition of vibe coding as it was coined by Andrej Karpathy of OpenAI, which is "iterally" (and I quote):
>I "Accept All" always, I don't read the diffs anymore.
Am I misunderstanding you, or are you vibe writing the wrong term without reviewing the meaning of your own words?
Here is the full literal quote so there is no room for confusion or claims of missing context:
https://x.com/karpathy/status/1886192184808149383
>Andrej Karpathy @karpathy: There's a new kind of coding I call "vibe coding", where you fully give in to the vibes, embrace exponentials, and forget that the code even exists. It's possible because the LLMs (e.g. Cursor Composer w Sonnet) are getting too good. Also I just talk to Composer with SuperWhisper so I barely even touch the keyboard. I ask for the dumbest things like "decrease the padding on the sidebar by half" because I'm too lazy to find it. I "Accept All" always, I don't read the diffs anymore. When I get error messages I just copy paste them in with no comment, usually that fixes it. The code grows beyond my usual comprehension, I'd have to really read through it for a while. Sometimes the LLMs can't fix a bug so I just work around it or ask for random changes until it goes away. It's not too bad for throwaway weekend projects, but still quite amusing. I'm building a project or webapp, but it's not really coding - I just see stuff, say stuff, run stuff, and copy paste stuff, and it mostly works.
That is the "literal" widely understood and well defined meaning of what you wrote, but misspelled "vibecoding", according to the well known AI expert from OpenAI who originally defined and championed the term. Its meaning has not suddenly changed.
Only a vibe coder would vibe code a linter to vibe lint vibe coded code for them, without looking at ANY of that code themselves, by just hitting "Accept All" always. Because vibe coders by definition don't want to bother reading what they generated, don't care what it means, and still expect it to come out perfect.
Don't be a vibe coder, or a vibe writer, or a human stochastic parrot: read what you write and know the definitions of the words you use.
And if you're not a vibe coder, then don't claim to "vibecode": that's "not engineering" just "magical and wishful thinking", as you like to say.
>"I always have to correct its hallucinations during the day. Why would I ever let it run unsupervised overnight?"
And the answer to your question is: because you are using git, so you can look at the diffs before merging and deploying into production. You are using git and looking at the diffs, aren't you? Have you heard of PRs and code reviews? Or is that too much to ask of a "vibecoder"?
Because you keep vibe-reading what I write.
Here's what I literally started with: "I always have to correct its hallucinations during the day. Why would I ever let it run unsupervised overnight?"
Which means what? Oh. It says literally what it says.
Here's how I continued: "Which tests and optimizations do you propose to run after a night of supervised work when one of main things that all agents keep doing is 'load all records from db , and filter them in memory'?"
What does this mean? Oh, it means the literal meaning of the sentence. It also probably strongly implies that I actually look at the code produced by these things, as otherwise I wouldn't know things like "oh, the agent loads the whole db into memory and filters it into memory, we have to correct that".
You literally completely ignored all that and got irritated by just this sentence: "It's now become so bad, I had to literally vibecode a separate linter for this".
Because, see, I had to write a tool with the use of AI and run it against existing code, and correct it until it worked to my satisfaction so that I don't have to spend a lot of my time correcting a repeating error, so I automated the finding and the correction of a repearting error using a linter. And instead of writing that entire sentence I used the word "vibe-coded".
So no you've wasted a lot of breath arguing... what exactly is it that you arguing? Your inability to read what other people write? Functional illiteracy?
Edit. Since you're so keen on chasing me in other comments it's unsuprising that you again completely ignored the one comment https://news.ycombinator.com/item?id=48965638 where I link this: https://news.ycombinator.com/item?id=48962703 But sure. How dare I use the word "vibe-coding" incorrectly when it was coined by the Lord Our God Karpathy Himself.
Yes, surprisingly, this is something Google cannot do yet.
1/ Using GUI software. My agents are using headful Google Chrome and Figma. It helps a lot to have separate environment, which is not interfering my main machine.
2/ Running long processes (1h+), so I can leave main machine closed.
3/ Running intensive processes. I use Gemma, Whisper and Qwen, which could burn main machine CPU and resources.
I think I'm gonna be a late adopter on this one until the industry figures out a less cumbersome pricing model.
With the $200 Claude subscription I was able to get around $13-15k of API equivalent usage in one month (note: this was during the "+50% usage" promotion that they have kept extending since May). When you hit your usage limit for a given time period you get cut off until the time period resets; don't bother paying for additional usage credits, you will be disappointed.
My average API equivalent use is around $30-40/hr. I would just bite the bullet on a plan for one month, then use that to calibrate your expectations around usage and cost optimization. The plans are heavily subsidized.
I'm not entirely sure there's a big advantage to Claude remote control to be honest. Maybe it's just me being afraid of lock in (I do switch between Claude and Codex somewhat often) and/or the inertia of changing things up that keeps me from trying it.
Like I want the LLM to have a bank account and he can do ANYTHING with that bank account that he wants. But he can't fuck anything up that has to so with me. He only has 2 - 5k
Problem is I need the LLM to do this without an SS.
as one has
How? I mean what could be the ultimate usefulness of Claude if not to make money, just like NFT but with extra steps. Most people isn't using Claude to make massive paradigm shift breakthrough discoveries. Nobody's curing cancer nor solving climate change. Probably the most common use case for llms it's just speed up the grind and make money. Like nfts.
Or the bubble will implode and we see ourselves at the job queue anyway.
https://isaiprofitable.com/ (Is AI Profitable Yet?)
https://www.yahoo.com/entertainment/music/articles/largest-i... (“The largest IP theft in human history”)
https://www.recordinglaw.com/news/authors-protest-ai-empty-b... (Don't Steal This Book)
https://finance.yahoo.com/energy/articles/why-ai-boom-break-... (Why the AI Boom Is About to Break the U.S. Power Grid)
https://www.aljazeera.com/opinions/2026/1/21/ais-growing-thi... (AI’s growing thirst for water is becoming a public health risk)
I too confess to writing code by hand. It… works for me?
I'm using Qubes OS, where everything runs in VMs without GPU acceleration, and never experienced this.
A primary source of UI lag is how Apple's native virtualization framework processes multitouch trackpad events. There are other issues like mismatched resolutions and framerates, too. Ask your favorite ai to debug. You can try e.g. deselection of the trackpad setting and instead select basic mouse support, and that can clear things up..
0: https://gist.github.com/smith153/04b4068b5a2d7b234f1c3d5992d...
I've been wanting to set up something exactly like this for my own use, but... You know, time is limited.
This is just enough scaffolding to have a little project for Monday morning!
sudo useradd agent
sudo su agent
So it can blow up its own files, but not mine.It was also doing some kind of headless Chrome stuff in there. I don't know how that works, but it was taking screenshots iirc.
I did also set up VNC at some point but didn't find it worth using.
>If it makes a mess, I can dump and reinstall in seconds.
This is also true of a $3 VPS, where I found it very amusing to give my agent root. What's the worst that could happen ;)
Why not just use a VM in the cloud and just a CLI interface?
There has always been a trade off sold to consumers of security vs convenience and a belief that giving up a little security gets a lot of convenience.
It ends up being often about a convenience of adopting the new tech, not necessarily in a way that's in the best long term interests of each party.
In my case, most of my VMs live on a proxmox cluster of HP EliteDesk minis, with the laptop sitting outside that as the thing that manages the cluster and a few other odds and ends.
It's a glorified API gateway...it could run on a medium sized potato
Almost like an entire generation just grew up coding on macbooks as the obvious choice for coders and just can't conceptualize hardware outside of that walled garden
idle = doesn't have to be H200 speeds
uses macs = local LLM can use all the ram, be smarter
as for hosted models, potato is tasty