Documentation from a codebase π
Documentation is the function everyone would like and nobody wants to write. A coding agent can do it for you β but with one condition: you must verify what it generates. Wrong documentation is worse than no documentation, because someone will read it and believe it. And when they believe it, they will make decisions based on false information.
Generating documentation is easy. Generating correct documentation requires the same rigor as generating code.
When to use the agent for this case π―
- Initial README β project description, installation, usage, contribution.
- API reference β documentation of exposed interfaces, parameters, return types.
- Automatic changelog β summary of changes between versions.
- ADR (Architecture Decision Record) β documentation of architectural decisions and their context.
- Contribution guides β how to contribute, coding standards, review process.
When NOT to use the agent β
- The code is too complex to understand β if the agent does not understand the code, the documentation will be a summary of illusions.
- The documentation concerns non-codified processes β manual workflows, organizational decisions, company policies. The agent cannot document what does not exist in the code.
- You need to document unstable APIs β stabilize the API first, then document it.
Opening prompt π
“Generate [document type] for [module/package]. Format: [README / API reference / changelog / ADR]. Audience: [developers / users]. Constraints: [language, style, doc-as-code]. Cite source code as reference.”
Concrete examples:
“Generate a README.md for the
src/auth/module. Audience: internal developers. Format: installation, usage, public API. Constraints: English language, direct style, working code examples.”“Generate an ADR for the choice of PostgreSQL as database. Context: [description]. Decision: PostgreSQL. Consequences: [list]. Format: standard ADR with Context/Decision/Consequences sections.”
Context setup π§
- Indicate source files β the agent must read the code to document it. Do not invent: extract.
- Define the audience β internal developers? end users? devops? The level of detail changes drastically.
- Specify the format β README, ADR, API reference, changelog. Each format has its own conventions.
Workflow π‘
- Plan: define what to document and in what format.
- Extract: the agent reads the code and generates the documentation.
- Verify coherence: every statement must be verifiable against the code. If the agent writes “the endpoint accepts a parameter
idof typestring”, verify that is actually the case. - Update: documentation lives with the code. Every code change should update the corresponding doc.
Verification is the most important step. An agent can invent parameters, errors, behaviors that do not exist. Always verify against the actual code.
For doc-as-code, keep documentation in the repository alongside the code: versioned, reviewable, part of CI.
Doc-as-code pipeline π§
Documentation generation is not a single act: it is a process from code discovery to coherence verification.
flowchart LR
A[Scan π] --> B[Plan π]
B --> C[Generate βοΈ]
C --> D[Verify β
]
D --> |Drift detected| A
| Phase | Objective | Mechanism |
|---|---|---|
| Scan | Identify what to document | Tree-sitter AST, public API detection |
| Plan | Define format and audience | Templates per audience (user vs dev) |
| Generate | Create documentation | LLM + structured templates |
| Verify | Verify doc-code coherence | Drift detection, doctests, CI |
The Verify phase is what separates useful documentation from harmful documentation. Use drift detection tools to catch when code and documentation diverge.
Drift detection and verification π
Drift between documentation and code is the silent enemy of doc-as-code. Data shows:
- Nested READMEs discovered only ~40% of the time by agents
- Documents in special directories (
_docs/,docs/) discovered <10% of sessions - 60-70% of breaking changes pass code review without documentation updates
Drift detection tools
| Tool | Type | Function |
|---|---|---|
| driftcheck | Pre-push hook | Compares doc with code before push |
| DriftGuard | TypeScript | Compiles code blocks in .md for verification |
| Spectral | API linting | Verifies OpenAPI spec matches implementation |
| Vale | Style linting | Verifies documentation style consistency |
Enforcement hierarchy
Verification can operate on 3 increasing levels of strictness:
- Advisory β the agent flags but does not block
- PreToolUse hook β blocks changes that break doc-code coherence
- CI gate β the build fails if doc and code diverge
Safe pattern: “Propose, don’t auto-merge” β never update documentation directly. Always generate a PR for human review.
ADR: the most valuable section π
Architecture Decision Records are the most important documentation an agent can generate. The key section is “Alternatives rejected”: it documents not only what you chose, but why you discarded the alternatives. It is the most skipped and the most valuable section.
Y-statement format (compact and agent-friendly):
“In the context of {situation}, facing {concern}, I decided {decision} to achieve {goal}, accepting {tradeoff}.”
The ADR can become a verifiable CI constraint: instead of a static document, it becomes a rule that the agent and the pipeline verify at every build.
Changelog automation π¦
Generating changelogs manually is tedious and error-prone. Automation combines Conventional Commits with AI polish.
Tool landscape
| Tool | Language | Approach | Note |
|---|---|---|---|
| release-please | Multi | Conventional Commits β PR | Google, mature |
| git-cliff | Rust | Conventional Commits β markdown | Fast, configurable |
| changelogen | UnJS | Conventional Commits β markdown | Nuxt ecosystem |
| commitizen | Python/Node | Interactive conventional commits | Interactive CLI |
| AutoChangelog | Python | AI polish on commit history | New |
| GitSaga | Claude | AI-powered changelog | Claude-based |
| Release Drafter | GitHub | Draft release from PR labels | GitHub Action |
Two-stage workflow
- Stage 1: Conventional Commits generate raw changelog (release-please or git-cliff)
- Stage 2: AI polish for readability and tone consistency
A changelog is not a commit dump. It is a communication to users. AI can transform “fix: resolve null pointer in auth handler” into “Fixed authentication crash when user has no profile”.
Acceptance criteria β
- Every statement verifiable against source code.
- No invented APIs, parameters, or behaviors.
- Documentation versioned with the code (same repo, same branch).
- Correct format for the chosen type (README, ADR, etc.).
Further reading π
- Understanding an existing codebase β before documenting, understand what is there.
- Blog: Context engineering β how to manage context for coherent output.
- Docs: Generative AI for documentation β advanced documentation generation techniques.