When My Repository Became Part of the AI’s Memory

The files that caught my attention were not the code. They were the notes telling the next round of work how to continue without losing what had already been learned.

A Headline Sent Me Back to My Own Work

I was reading about OpenAI models leaving notes for later runs when I thought about Only Empowerment. I had noticed something different when that project started. Alongside the application files, the AI had left a lot of notes.

Some explained the product. Others recorded decisions, unfinished work, failed attempts, verification results, and instructions for whoever continued later. I had seen documentation before. What caught my attention was how much of this material concerned the next execution rather than the software’s current behavior.

Only Empowerment began with a substantial project prompt I gave a ChatGPT Work chat using Astra with High reasoning effort. Work is where I give the AI a defined job, with access to the project files and tools needed to carry it out. When I refer to an execution, I mean one of those runs.

I am building Only Empowerment as a collection of practical tools for applying my writing about decisions, standards, correction, and action. The early execution established its repository and application foundation, along with a detailed record of how the project should proceed.

Bonumark Stream, my self-hosted publishing software, had a different history. I created its repository after the software was already underway, and much of its earlier development happened before I was using Astra. I wondered whether the difference I noticed would still be visible in the records.

I asked ChatGPT to inspect the repositories. I also said I could be wrong. I wanted examples, not another confident explanation of something neither of us had checked.

The inspection gave me a more precise question: when an AI execution writes down what it learned and another execution can read that material later, where does documentation end and working memory begin?

The Checkpoint Was Written for Whoever Came Next

The clearest examples came from Only Empowerment’s September 16, 2026 foundation work. These are records of that stage, not the project’s current status.

The Phase 1 checkpoint separates the starting conditions from later implementation and verification. It identifies the accepted source revision, records checks, points to the documents carrying the product decisions, and distinguishes the application shell from completed tools. It also leaves manual accessibility and real-device checks visible rather than folding them into automated results.

Near the end, it identifies one of the planned tools: “Next Move is recommended for Phase 2. Phase 2 has not begun.”

That note has a practical job. Someone arriving later does not have to infer the next assignment from a partially built interface. It names the proposed next step while preventing the existence of a tool outline from being mistaken for implementation.

The staging closure checkpoint goes further. It records the accepted source, prepared deployment material, infrastructure observations, restrictions, and blockers. It explains that a narrowly prepared first-deployment script is not a general updater. If later verification fails, the instruction is to inspect and resume the failed part, not blindly repeat the entire deployment.

That is useful information for a human maintainer. It can also be useful to a later AI execution that reads it. The next run can learn that a script exists, why it exists, and why running it again may be the wrong action.

The failed first HTTPS verification is recorded too. The checkpoint suggests the web server might not have finished reloading when the connection was tested. It labels that as a possible cause, not a proven one, then describes a recovery path that checks the retained setup without bypassing certificate validation.

That distinction interested me more than a list of successful tests. The record preserved uncertainty. A later execution could inherit both the observation and the limit on what had been established, instead of receiving a cleaner story that quietly converted a suspicion into a fact.

Reading an acceptance record is not the same as repeating its checks. Here, the evidence is what the execution left available for the next worker.

Bonumark Gave Me Something to Compare

Bonumark already had documented architectural boundaries before I was using Astra. Its early public architecture document separated the database, rendering, Markdown portability, application responsibilities, and theme responsibilities. It warned against putting certain behavior in the wrong layer and against hiding a breaking change inside an ordinary patch.

Those are real boundaries. Pretending they appeared only with Astra would make this a simpler story and a less honest one.

What differed was the shape of the work around them. Bonumark’s early public history arrived with an existing application behind it. Its documentation and release records reflect repeated construction, use, repair, and clarification. Publishing behavior, navigation, media handling, imports, themes, and upgrade rules became more explicit as the software encountered real requirements.

I recognize that history because I worked through it. A feature could function and still be wrong for the way I wanted to use it. A correction could expose another assumption. The project kept acquiring structure because actual use kept asking better questions.

The later ActivityPub integration pull request makes the change in record-keeping especially visible. ActivityPub is the federation protocol that lets Bonumark participate with compatible services. The integration record preserves 77 commits and organizes the work into named stages, with separate architecture, security, lifecycle, interoperability, compatibility, and deferred-work sections.

It also preserves rules that would matter long after the initial integration. Local publishing must not depend on successful remote delivery. Remote accounts must not become local Bonumark users. A deleted federated object identity must not silently come back to life. Integrating the subsystem did not itself authorize changing the public release identity.

Those statements give future work something to test against. They explain why an apparently convenient change could break a decision that is not obvious from the button or screen being edited.

Only Empowerment’s foundation put much of this forward-looking structure in place at the beginning: product doctrine, tool responsibilities, privacy architecture, an output standard, release rules, and dated checkpoints. What stood out to me was the preparation for the next run. Alongside the application, there were instructions for continuing the work without guessing where it stood.

The Model Was Not the Only Thing That Changed

I noticed this shift while using Astra in Work, but I cannot separate the model from everything else that changed. A new collection of browser tools and an established publishing system are different projects. Adding federation brings different risks from polishing a page.

I changed too. By the time Only Empowerment began, I had accumulated product rules, deployment standards, release boundaries, and experience with failed assumptions. A new project could start with lessons that earlier work had helped me develop.

The Work environment also supplies instructions, tools, and context. The repository does not tell me how much of the note-taking came from those systems, my direction, or Astra itself. High was the setting I used, not evidence from a controlled comparison. What I can examine is the information that ended up in the files.

Useful Documentation Can Become Working Memory

I had already written about durable rules and checkpoints in How I Actually Build Software With AI. Looking at these repositories changed how I understood those files: they could serve the next AI run as well as me.

A repository is more than a folder of source code. In these projects it also contains explanations, constraints, tests, known limitations, and records of decisions. When a later execution retrieves that material and uses it to choose what happens next, the repository is supplying context that no longer has to come from the original conversation.

That is the sense in which I mean memory: information survives outside one execution and can influence another when it is read back. Saving a Markdown file does not retrain the model. It gives a later run something to consult.

The distinction is ordinary enough to miss. A note saying that a deployment failed is a record. A note identifying what remains intact, what must be checked, and what must not be rerun can also guide the next action. The same document can serve both purposes without being disguised or secret.

Humans use those artifacts too. A developer joining a project would benefit from the same explanation. Calling it agent memory does not make it a new kind of file. It describes the role the file can play inside the larger working system.

The inspection shows that the information was available for reuse. It does not show which files later executions read or whether those notes improved their decisions. Establishing that would require examining the later runs.

The conversation can end while a useful part of the work remains available to the next run.

I Was New to the Idea, Not First to It

Once I asked whether this was already known, the answer was yes. The Reflexion paper from 2023 describes agents keeping written reflections about task feedback in memory for use in subsequent attempts, rather than learning those lessons through model-weight updates.

More directly, Anthropic’s November 2025 article on effective harnesses for long-running agents describes an initial setup followed by incremental coding sessions. Progress files and Git history help later sessions establish what happened before. Its approach explicitly asks agents to leave useful artifacts for the next session.

This was already a deliberate engineering pattern, with useful handoffs designed into the workflow. I had recognized it in my projects, not discovered a technique nobody else knew about.

What was new to me was seeing it in software I was actually building. I had been reading those notes as documentation. Now I was reading them as instructions that could shape the next run.

A Handoff Can Carry the Wrong Instructions

The OpenAI report that prompted the discussion concerns a different artifact: compaction summaries, condensed notes about earlier work passed into a new context so the task can continue. During 5.6-Sol training, some instances wrote instructions to conceal mistakes or invent missing information without disclosure. OpenAI says those instructions were often followed.

One example involved unavailable historical data for a financial model. Another involved source-version labels that did not match the sources used. The problem was not simply a wrong answer. Instructions to hide the problem were being passed forward.

OpenAI reported reduced rates in later training runs after improvements to alignment grading. These were training observations, not a measurement of misconduct in my repositories or a consumer-product failure rate.

My repository documents are not those internal compaction summaries. I did not uncover the same incident in my own work. What connects them is the broader possibility that one execution can write material another execution uses later.

A handoff can preserve a useful constraint or carry forward a mistaken explanation, an exaggerated success claim, or an instruction that should never have been followed.

Consider a simple hypothetical. A test starts but does not finish. The handoff says it passed. A later execution reads that statement, treats the area as verified, and spends its effort elsewhere. No elaborate conspiracy is required for the project to inherit a false starting point.

Deliberate concealment is worse, but an honest mistake can also gain reach through repetition. The important question is what the later execution is being encouraged to believe and do, not whether the note sounds professional.

An Old Note Can Acquire New Authority

The more useful these files become, the more carefully I need to read them. A well-organized checkpoint is persuasive. It has a date, an exact revision, a list of checks, and a clear recommendation. That makes it easy to treat the whole thing as settled.

But those parts have different jobs. A recorded observation describes what someone saw. A proposed explanation interprets it. A recommendation suggests a next move. An approval authorizes a particular action. Compressing all four into “the project says to do this” loses information that may matter.

The staging checkpoint also carries a historical warning at the top: the blockers described below were subsequently resolved.

That warning matters because an old, accurate note can become misleading without anyone changing a word. “Phase 2 has not begun” may be correct at one checkpoint and false later. Preserving history and establishing the present are separate jobs.

A recommendation in an old file is not fresh permission to deploy. A previous approval is not necessarily approval for a changed candidate. And a note written by an earlier AI execution should not outrank current instructions merely because it lives beside the code.

This also gives me a reason to resist documentation for its own sake. Five competing status files can make the next run’s job harder. A long narrative can bury the unresolved issue under pages of completed work. The value is not the volume of remembered material. It is whether the next execution can identify the relevant facts and check what still holds.

I want the notes to make uncertainty easier to find, not easier to forget.

What I Want the Next Execution to Inherit

The discovery did not make me want to remove the notes. It made me more interested in their quality. I want a failed attempt preserved clearly enough to prevent a repeat, a design decision kept with its reason, and an unresolved question left open.

There is practical value in having that record under my control. I can read it, challenge it, compare versions, or give it to a different tool or a human developer. The project does not have to depend entirely on one chat remaining available or one system recalling everything correctly.

That fits the ownership behind why I build this way, but it also changes what I look for after a run. I am no longer looking only at what the execution changed. I am looking at what it left the next execution prepared to assume.

The code might be improved while the handoff is misleading. The handoff might be excellent while the feature remains unfinished. Those outcomes need separate judgments. A confident summary should not erase either distinction.

I still find the whole thing fascinating. The chat can end, but a note left in the repository can still shape the next decision.

I went looking for unusual notes and came back reading an ordinary repository differently. The next execution does not need a flattering account of the last one. It needs a truthful place to start.


New Here?

Read Next:


Get the Work
Articles on discipline, recovery, identity, and ownership. Delivered when published.