The idea
I tweeted a brief series of thoughts this weekend.
Commonsense docs feature that isn’t standard: align the [high level, not API ref] docs to the codebase and record them as having been written at some checksum of each file and (written at) git commit hash (since the primary failure mode is staleness)
LLM docs' failure mode is probably something like recency bias (they find the context window disproportionately salient) so one solution would probably be a docs writeup of some kind from scratch on each development session, and then have a docs process take stock in the long run
This would act like an expansion of a changelog, or a farsighted sliding window over myopic bursts. It overlaps with development journals (which in turn are internal-facing analogues to repo issues/PR documentation) but is written for user perspective not developer memory jogging
To expand on them a little:
Firstly, it should be standard practice to align docs to the codebase in some way, at a specific version (content checksum) of each file and edition (commit hash) of the repo. The reason being that as soon as any file changes we are liable to have documented a version of it which no longer exists, and we should be able to step back to a particular commit to compare against.
Secondly, as well as staleness there is a separate issue of a sort of ‘saliency bias’ or over-attendance to the context window (by definition in fact, that is all they have to attend to). This would be for example in the most recent task, say editing a particular feature. If docs were to be written off the back of such a context window, they would likely overrepresent the importance and saliency of said task over a neutral documentation process. This motivates some ‘view from nowhere’, or a fresh docs processor “from scratch” which the context window merely contributes to, and/or a long term process if not on every development session.
This felt analogous to a changelog in the sense of a document which records specific updates. However we don’t really have any proper mode of recording this sort of documentation update. We have development journals (which are mostly an engineering/state log but can incorporate elements of ADR, recording why decisions were made), but these are written for ‘internal’ developer use rather than simulating a user perspective.
So my vision here was in two parts:
-
Freshness pinning. Each high-level doc records the content checksum of every file it describes, and the commit hash of the repo when it was written. A checksum mismatch flags the doc for review. The commit hash gives a fixed point to diff against, showing what the docs were describing versus what exists now.
-
Saliency correction. Docs written in a development session overweight whatever that session was about. So session writeups should be treated as fragments, not as the docs themselves: user-facing notes written fresh each session. A slower, neutral docs process then periodically takes stock and folds them into the docs proper. This is the changelog-fragment pattern applied to conceptual docs, and the user-facing counterpart to a development journal.
So each development session would produce a documentation fragment: a user-facing note recording the changes and relevant intent from that session. These fragments are not themselves canonical docs; a consolidation process folds their content into the docs proper.
Prior art
Freshness pinning
This appears to already exist in Codocia, a docs drift checker aimed at coding agents. Doc frontmatter declares which files a page covers, and a snapshot file stores their content hashes. Content hashes decide staleness, and the commit hash is kept only as audit metadata. It flags the inverse case too, a changed source file with no doc covering it. doc-lattice does the same between documents rather than between docs and code: each doc records a hash of the upstream doc it derives from, and CI fails when that upstream changes. It calls this “traceability” and suggests you may have docs stemming from product briefs, or integration guides built on particular API designs.
Swimm is the commercial, patented version for dev team “knowledge management” since 2019. It couples docs to anything from folders down to single variables, verifies them on every commit and can auto-update them. Google takes a time-based rather than hash-based approach: Software Engineering at Google describes "freshness dates" on docs, which trigger reminder emails when a doc goes unreviewed for too long.
The premise that staleness is the primary failure mode has empirical support. Tan, Wagner and Treude analysed over 3,000 GitHub projects and found most contained an outdated code reference at some point in their history. A follow-up found the same in over a quarter of the 1,000 most popular repos. Their diagnosis is that developers don't notice when a code change makes the docs obsolete, and a recorded checksum is exactly the signal that would tell them.
It was also an oversight apparent at the conception of projects like llms.txt, which is only now (2 years in) receiving attention to the premise-undermining issue of freshly produced docs being conceptually stale (see #132).
Saliency correction
Part 2 exists in pieces. The fragment-then-consolidate structure is changesets and towncrier: a small Markdown file per PR, written while the change is fresh and collated at release time. This is my idea of "myopic bursts", but for release notes rather than conceptual docs.
RepoAgent (2024) regenerates LLM-written docs for the minimally affected scope on every commit via a pre-commit hook, but its output is object-level, close to API reference. Per-session agent memory files such as Cline's "Memory Bank" are rewritten each session, but they serve developer (or agent) memory, not users. They sit on the development journal side of the distinction.
What I haven't found is these parts in combination: user-perspective conceptual docs written as per-session fragments, pinned by checksum and commit, and reconciled by a long-horizon process.
Implications
Two implications follow then from the prior art.
First, whole-file checksums are too sensitive for high-level docs, since any edit trips them. Hashing something structural, like the exported interface, a normalised AST, or a catalogue of per-symbol intents as in MutaGReP (2025), would make a better staleness signal, with the raw hash kept for audit.
Second, these session fragments still carry the saliency bias of the sessions that produced them. So the consolidation process should work primarily from the actual git diff since each doc's recorded commit, and use the fragments as hints about intent. That way the pinning becomes the consolidator's input, not just its alarm.