However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
LLMs don't retrain, it's too expensive atm.
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
The code is the documentation. There's no need for most of this stuff.
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
But that's one session. Isn't memory for… the next session?
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.
It's fascinating for local coding.
The slightest amount of guidance, input from experience can make a huge difference.
And it's quite fascinating. Still dont think there's trillions of TAM out there if Qwen3.8-Flash-Next runs fast enough on $3k (before memory cartel) pricing.
And it’s an experimental likely undertrained model. Wait til Qwen 4 Flash…
However, I found the same problems with agents that it has with humans (usually with humans it isn't TDD but code coverage requirements). The problem is the tests are now written to satisfy a bureaucracy rather than to properly verify the code. And there are real studies that point out the low value of these types of bureaucratic testing requirements. For example, when models are just required to write a test file, they don't produce a better result: https://arxiv.org/abs/2602.07900
I wrote a /verify skill that focuses on properly verifying code and I am finding it works a lot better. I do need to revise it now- in practice certain parts of the skill are doing all the heavy lifting and others are more dead weight. But it does seem to be properly orienting the agents towards finding defects. https://github.com/gregwebs/skills-sdlc/blob/main/skills/ver...
Context then is: docs, code and tests. Rather than try to build some omniscient agent, the context is built where the agent needs it.
TDD doesn't help the agent figure out what it is supposed to do now- it is already writing the code and now and it must know what it is supposed to do to write it.
I am trying to use evidence-based approaches. I am only observing better results and tweaking now, but I plan to benchmark when tweaking is done.
I'm working on javascript, and now I'm only doing this via typescript. I've successfully got this going:
1. Design a feature in plain language with the coding agen (opencode) and write a document for it.
2. Restart the context (or I use /compact to flush any errant details)
3. Pull the new plan into context and ask the agent to revise it follow Test Driven Development.
4. Depending on how big it is, the agent places it into multi stages, each with it's own document.
It then loops through the stages. It might be entirely based on the language you're using, but this loop seems strong enough.
Then when there's bugs, errors, anything, we update the doc, add more tests, then revise the code.
It's quite possible the harness you're using isn't setup properly. For me to do this locally, I had to hack on dynamic context pruning which I described here: https://news.ycombinator.com/item?id=49906637#49907641
I write README and docs by hand, thus only important stuff goes in there. If it is not worth my attention to write down, it's not worth writing down.
Smaller context, easier to consume.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
The github repo might be more helpful: https://github.com/Principle-Driven/pdd
I was wondering about this bit:
> "Code comments cite a versioned token where code depends on the rule."
What does a token look like? I went looking for an example, but I don't know what I'm looking for.
// The binding deliberately SKIPS the contradiction check: a form
// bound to an object the same document deletes is a legal stale
// state (PDD-3@v1)
Where PDD-3 deals with a specific type of race conditions that my agents were over-optimizing/spinning out on; Now if I ever change the rule the version would change to PDD-3@v2; which triggers a re-evaluation for all decisions that were made under the previous version.The questions are designed specifically to identify the state the agent is in, in order to assign its steering. Changing this tool changes how the agent works. You just need to make the use of this tool a necessity so it does not forget to call it.
Unlike principles and memories, questions are more openended and tend sometimes to trigger the right mentality in the agent even before it gets to the policy itself. In my opinion agents get lost in local work losing the big picture, my questions jog the big picture back into attention.
1. call ssp tool, no arg, it just reads features from the repo, like most other tools -> ssp locates state and sends 4-5 questions
2. agent responds the first batch -> state is further narrowed down -> ssp sends a second batch of questions
3. agent responds again to the interview -> state gets finally pinned down -> agent gets the steering assigned
4. a log is created of this ssp session, and reflection on the log used to refine the questions and steerings in the ssp tool (name comes from State-Space Policy)
It seems you tried to reinvent instruction files. What do you think is the difference between your approach and standard tools such as instruction files, AGENTS.md, and even skills?
If it helps, here's how you think about it, principles are not decisions, they're mental models and frameworks to allow making decisions. In my case, they're there to allow agents to make decisions without asking me. But the decisions cannot be recorded - the code serves that purpose. What can be recorded is why the commend for the block of code needs to cite what principle was used in the current shape, which is why there is a citation of the principle.
When the principle evolves, so does every decision, at the very least, every decision gets re-litigated to see if it needs to evolve.
Do you have any resources that guided your work?
A big part of it really comes to that I have led large product teams and platform engineering teams in the past, And I like giving feedback to my agents in the same way that I would plan things with my teams.
One of which is give them repeatable mental models which allow them to make decisions without relying on managers or me. The result is usually that the teams can work for weeks-to-months with autonomy.
I try achieving the same with my agents, so that they can work for 1-14 days without my intervention. I usually have them working on very very long running tasks/initiatives.
If you do want some resources, this is one of the pages that I've been using for nearly a decade to help people get started with mental models: https://fs.blog/mental-models/
But seriously, Opus has been garbage after 4.6 -- the public seems very easily fooled into thinking breadth == depth. Anthropic, I'll grant them, has been and is still an extraordinary data team. As for model alignment though... I have to wonder sometimes if they are even trying beyond just chat training.
"Guys, SI is just around the corner... Huh? What production DB? Anyway GPT2...I mean Mythos is too good it would be dangerous to release to the public!"
You’re confusing, I think, memory system with llm finetuning. Completely different concepts.
Every single project at work that has documentation/ADRs/etc ends up stale as fuck and it pollutes context more than it helps. Even the agent generated docs. Hell, especially the agent generated docs. It's like the LLMs can't just DELETE something, they always amend. So now there's stale details in there about something we abandoned polluting context.
The code is the documentation, none of this stuff is needed. I'm seeing WAY too many devs/engineers go down the rabbit hole of trying to build these tools and systems but not actually getting meaningful work done.
I've been doing my best to remove and delete all the documentation. I usually keep a super high level AGENTS.md/CLAUDE.md file and that's it. The llm/agent can figure it out. You might argue I waste 5 minutes and a dollar or two for the agent to "figure it out" each time, but I'm pretty certain people are wasting FAR more time and money building these things out and maintaining them in each of their projects (or NOT maintaining them and they're actively BAD for their projects).
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
It's not like it's being rewritten for efficiency. "Just because".
You haven't been paying attention then. I routinely see Claude and GPT models churning out python code to do stupid things like linting. Last week I even had a TypeScript project with prettier configured all over the place, including in a custom agent skill I added with the express purpose of getting the damned model to lint the code, being constantly prompted to run ad-hoc python code supposedly to format whitespaces. I even explicitly prompted one session to just use npm run lint, where I pointed out the exact line of code where prettier was invoked, and the session still churned python code to hande whitespaces.
Otherwise you can get a python one liner that execs a different script engine.
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.
I think there is a deeper problem emerging from this sort of behavior. Even when we bother to create agent skills with there own scripts that call tools like jq a specific way to achieve a goal, AI coding assistants and agents still go way out of their way to generate ad-hoc scripts to do the most absurdly stupid tasks such as parsing output in structured language, and even remove whitespaces from a markdown file. This means AI coding assistants and coding agents treat agent skills as mere suggestions of using a alternative option that more often than not the choose to ignore.
This has a very dangerous implication: your average user is trained to develop a pavlovian reflex to authorize agents to just execute their ad-hoc scripting code with our own permissions and credentials in our systems, which includes the ability to call anything over the internet.
https://gist.github.com/cynthiateeters/6868ca26c059a3106cd93...
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
That being said, the danger is measured in computer damage, which can be a lot personally and to a company, but less existential, so your mileage may vary as to how "scary" it is.
What scares me are the people who think that this software is the equivalent of an perfectly-smart elf in a box and use it blindly.
How, if at all, do you keep the architecture or a working theory of the code in your head?
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
I've been bitten by agent-written ADRs. Agents carelessly add extrapolated details and speculated nice-to-haves that I never asked for, and this becomes a source of bloat that keeps coming back like a boomerang.
Lack of this capability makes automated principles or patterns update a recipe for more bloat.
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- The folder name includes 'tripsy' and this is how I remember it.
- `jd` just parses my limited tree and `cd`s me to a folder.
- `jd 21.15` gets me there by number if preferred.
- `claude`- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
Legally and culinarily, they're vegetables. Botanically they're fruit.
Tomatoes are members of the nightshare family which includes tobacco, potato, and chili peppers.
They are used in popular recipes like Mexican salsa and Italian pasta sauce. Pasta sauce commonly uses the San Marzano variety. Here's a <picture> of San Marzano tomatoes from our Italy trip.
Rhett likes tomatoes in all dishes. Link only likes tomatoes if they're blended in a dish like pasta sauce or tomato soup. Frank is allergic and can't have tomatoes.
Tomato prices are up due to a bad <year> yield in <country>.
---
This is all memory. Where does it go?
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
https://backnotprop.com/blog/context-monorepos/
- Better models can keep the drift in check.
- What's super important these days is having all historical context, historical decision making, prototypes and whatnot.
- all projects and their worktrees located together
- all reference docs, code, etc - a folder away
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.I don't think you need more documentation or these memory systems. The code IS the documentation. Lots of this stuff is unnecessary. It's a bunch of people doing their special rain dances and then when it happens to rain they say "I did that"
https://cline.bot/blog/memory-bank-how-to-make-cline-an-ai-a...
If someone has good A/B evals of this being more effective I'll eat my hat, but the reason no one is publishing them is because well, evals are hard, and this is likely just magical thinking.
My current view is that an agent generally needs three things to work effectively: a way to discover information that isn’t obvious, such as a minimal `AGENTS.md` that points to more focused brief files; a clear way to verify that its work is correct; and some guidance on project-specific tastes. Everything else is noise.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
I take the article's point more directly. It's just a straighter line to have clear communication through documentation than to fuddle around with the perfect memory setup.
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
I'm using it to foster IOSP (integration operation segregation principle) for example.
https://youtu.be/6AgndHSkHFI?t=238
(You could of course argue that you don't like the direction the rules push the agents in given in the example - some people don't prefer small functions everywhere - but that's not the point: The lint-hooks work in pushing the agent in the desired direction; If one desires something else they'd simply need different rules)
Then for example you could write your own hook and convert existing code smell documentation which agents ignore into a format that works with the hook.
That's what I did; I did not use the library I linked to ~ It served only as an inspiration.
Instead, I let the agents create custom lint rules (using eslint, pylint, ...) and add custom coaching error messages based on where I want to take my codebase.
Or is that the secret sauce no one wants to share, the edge people see themselves having.
Just flagging big functions will make the agent write small functions ~ but not necessarliy in a good way
(for example the agent might just cut `doOneThingAndTheOther` in half and call the second half `doOneThingAndTheOther2`)
Indeed, AI fanatics lack critical thinking. They easily accept the idea that AI needs documentation, but merely disagree the method to approach it.
Isn't memory just documentation that agents create and update on the fly to fill in the documentation hole that your average user leaves open?
There was a time when the need to provide context and signal was a key topic in LLM and AI-assisted coding. People talked about MCPs AGENTS.md and README.md and agent skills and even comments, descriptive tests, and naming conventions. Supposedly the theory was that if you provide context, agents wouldn't misbehave so much. Then plan mode and spec-kit approaches stepped in with approaches aimed at providing that context when starting sessions, because users never bothered with docs or MCPs or AGENTS.md or anything. But then plan mode and spec-kit approaches were considered too laborious and requiring too much from users. Then, because users systematically failed to provide context, this task was finally given to agents. And things just worked.
You should ask yourself why in software engineering circles documentation is seen more as a problem than a solution to any problem.
The only "prompt" is in AGENTS.md saying that this thing exists and there's a map.md <- which is a one liner reference to whatever the agent stores in there.
And usually, I tackle a new feature, and at some point tell it to store to jot down notes in workbench if I'm comfortable with it (and if it needs to be stored in memory)
This makes it easy to just spawn other agents and such from a good point, I just point them that stuff is in workbench.
Token costs, seem decent and I can always just delete stuff in there as its purpose is ephemeral.
https://github.com/ArtRichards/docs-cli
and
My most recent example were some deprecated and archived repos that kept getting added to plans for patching issues.
We use a private Claude plugin marketplace for internal plugins and skills and I try to regularly prune my memory, migrating relevant stuff to a proper home (skills in the marketplace, docs in repos or Notion etc) and prune outdated information.
The idea of dumping everything into a big database and hoping the model will write the correct queries does work out to some extent. It's a very enchanting idea. However, it pales in comparison to having a dedicated tool per type. The outcomes seem to be much better when joins between low cardinality types occur within the token stream.
If your agent does need access to some enterprise knowledge base, I would give it a lexical search capability and not overthink it with vector shenanigans.
Tools are the only thing you need if you build them right. I don't even have a system prompt anymore aside from injecting the name of the robot and the current user's name. Keep in mind that all aspects of tools can be dynamic over time. I've got some where the description is composed by hundreds of lines of conditional string builder depending on the current state of the conversation.
I think what you are trying to say is that “give the LLMs access to information and it can figure out how to use it. Don’t get fancy about how to structure the data”.
If that’s the case it has largely turned out to be true. RAGs have fallen out of fashion.
Where does that put AGENTS.md though? Is it worth spending time to structure it or just dump it.
I've worked with docs that get _updated_, and progressive disclosure: make references to more obscure features in their own page, so they don't majorly bloat the context.
.agents/plans/<plan-name>/
.agents/notes/<topic>/
.agents/knowledge/<topic>/
I only commit knowledge and if knowledge gets big i add .agents/knowledge/INDEX.md
This is a good mix of human readable and agent fluent. Notes are ephemeral, knowledge is permanent.
I have a rule for knowledge that it has to be stable and mostly permanent (though updatable). And the agents are not allowed to post there unless docs are clean organized and with permission.
Notes are for jotting things down and handoffs, massaging a featureset. Agent can document at will.
Still WIP.
Organizing the markdown files by feature/issue/change makes it easier for the LLM to search for the appropriate documentation. Coupled with well-broken down Claude rules files, and Claude Code (or other harnesses) get better as you make more changes.
https://news.ycombinator.com/item?id=47300747
"We should revisit literate programming in the agent era" (silly.business) 292 points, 251 comments
This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.
This way the primary agent only has relevant information in their context to make decisions and take actions.
Context management is still under valued imo.
Not today. IMHO Documentation is just a form of structured memory and it’s all just context. Getting that context right is a hard problem and there’s a lot of different ways to skin that cat.
I don't know, but I have a pretty standard (I think?) setup, and Claude manages to find every relevant file every time. But I've also only used Claude for greenfield projects, where "documentation is primary, and code flows from documentation" is the philosophy.
I have CLAUDE.md describe all the types of documentation files and the directory structure. And then Claude is pretty aggressive (automatically) about always inserting cross-references everywhere. So a feature description will reference the ADR's that it implements, the ADR's say what feature implements them. A code file will make reference to the "implementation design" document that describes the motivation behind which iOS elements were chosen, how the animation is defined in a particular way that doesn't break another animation, and so forth. So I've really never run into a situation where Claude failed to read a file it should have. I've been pleasantly surprised.
I would say that the one really big thing I've had to learn is to teach Claude both in CLAUDE.md and in the header of every top-level design document, that keeping documentation current and in sync is paramount. Because its default seems to be to keep history and append, e.g. by default it will take a section of a document and mark it "[DEPRECATED]" and add the new version below. So my instructions are pretty clear in having it always be aggressive in maintaining current state only, always replace rather than append. And if there's anything we want to save from the previous approach (e.g. we did it X way previously and it failed because Y), then just add that as a new short note in the new current-state text, possibly with a pointer to a commit or tag or something.
So this seems to solve both recall and staleness in my projects at least.
The only thing I still haven't found a solution for is numbering. Claude is always giving everything numbers, like F23 for feature 23. But I'm always changing the order of things, inserting new things, deleting things, so I wind up with a sequence of development work that goes in order like "Phase 9", "Phase 9b", "Phase 9e", "Phase 11", "Phase 12". I'm halfway ready to abandon numbers entirely and just start giving things names from noun collections instead, so every feature is named after an animal, every ADR is named after a kitchen implement, or something. Or just four-digit hex codes chosen at random. Curious if anyone else has found what works.
When you have knowledge distributed in markdown files; finding them puts the path/filename into context as well as some indication of document size. (If its on line 1200 or line 20). This is extremely valuable for picking what ought to be focused on next.
RAG on the other hand creates the hardest challenge for these models. It instead puts 5 ideas with the highest similarity into the context in full.
Its the difference between having to remember a set of numbers when in a crowd that's talking about stuff, and having to remember them when the crowd is shouting out random numbers. The similarity in the task makes things harder. SoTA models work despite this, but its extra-gambling while you're already gambling.
In this implementation, Markdown should be considered harmful.
It looks like Operator Memory injects `.operator-shared/operator.md` and `.operator-shared/index/.md` directly into your agent's instructions before you even write the first prompt.
So if you clone a repo or review a PR where a bad actor put malicious instructions, now your agent executes those instructions automatically and silently.
It could exfil `.env` and `~/.ssh/`, change `~/.bashrc`, all kinds of dirty deeds.
Agents are pretty good now about not running prompt injections hidden in code and Markdown, but this plugin bypasses all of that, and puts the prompt injection right in the system prompt.
Seems bad.
"Code is the documentation" doesn't solve the problem in my experience. Because what happens is that you still need a lot of "why" comments in the code, and then these go stale, so you still have the same problem you have to solve. And so I find that markdown documentation is a lot easier to organize and review in a structured hierarchical way in one place, than code comments sprinkled across the repo.
It also doesn't seem to have a way to separate preference from factual memory which is useful when you have various humans interacting with the same agent. One human might prefer a certain output style over the other and that's something a memory framework can also address.
Like the MCP articles a few months ago, it also seems to assume all agents are cli coding harnesses running on your local machine. We have a handful of other things like chat bots, event-driven agents running on servers, chat driven agents running on servers in sandboxes--there's not a single filesystem and even if there were one, having multiple agents try to edit it at once would corrupt it.
My main issue with LLMs is that by construction they work primarily by addition, and are very task oriented. If you have documentation, it adds blobs corresponding to its task and that's it, and soon enough you need to break your documentation onto chapters and you're back at square one.
https://gist.github.com/karpathy/442a6bf555914893e9891c11519...
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.