Skip to content
Documentation from codebase

Documentation from a codebase πŸ“–

Documentation is the function everyone would like and nobody wants to write. A coding agent can do it for you β€” but with one condition: you must verify what it generates. Wrong documentation is worse than no documentation, because someone will read it and believe it. And when they believe it, they will make decisions based on false information.

Generating documentation is easy. Generating correct documentation requires the same rigor as generating code.

When to use the agent for this case 🎯

  • Initial README β€” project description, installation, usage, contribution.
  • API reference β€” documentation of exposed interfaces, parameters, return types.
  • Automatic changelog β€” summary of changes between versions.
  • ADR (Architecture Decision Record) β€” documentation of architectural decisions and their context.
  • Contribution guides β€” how to contribute, coding standards, review process.

When NOT to use the agent β›”

  • The code is too complex to understand β€” if the agent does not understand the code, the documentation will be a summary of illusions.
  • The documentation concerns non-codified processes β€” manual workflows, organizational decisions, company policies. The agent cannot document what does not exist in the code.
  • You need to document unstable APIs β€” stabilize the API first, then document it.

Opening prompt πŸ“

“Generate [document type] for [module/package]. Format: [README / API reference / changelog / ADR]. Audience: [developers / users]. Constraints: [language, style, doc-as-code]. Cite source code as reference.”

Concrete examples:

“Generate a README.md for the src/auth/ module. Audience: internal developers. Format: installation, usage, public API. Constraints: English language, direct style, working code examples.”

“Generate an ADR for the choice of PostgreSQL as database. Context: [description]. Decision: PostgreSQL. Consequences: [list]. Format: standard ADR with Context/Decision/Consequences sections.”

Context setup πŸ”§

  1. Indicate source files β€” the agent must read the code to document it. Do not invent: extract.
  2. Define the audience β€” internal developers? end users? devops? The level of detail changes drastically.
  3. Specify the format β€” README, ADR, API reference, changelog. Each format has its own conventions.

Workflow πŸ’‘

  1. Plan: define what to document and in what format.
  2. Extract: the agent reads the code and generates the documentation.
  3. Verify coherence: every statement must be verifiable against the code. If the agent writes “the endpoint accepts a parameter id of type string”, verify that is actually the case.
  4. Update: documentation lives with the code. Every code change should update the corresponding doc.

Verification is the most important step. An agent can invent parameters, errors, behaviors that do not exist. Always verify against the actual code.

For doc-as-code, keep documentation in the repository alongside the code: versioned, reviewable, part of CI.

Doc-as-code pipeline πŸ”§

Documentation generation is not a single act: it is a process from code discovery to coherence verification.

    flowchart LR
    A[Scan πŸ”] --> B[Plan πŸ“‹]
    B --> C[Generate ✍️]
    C --> D[Verify βœ…]
    D --> |Drift detected| A
  
Phase Objective Mechanism
Scan Identify what to document Tree-sitter AST, public API detection
Plan Define format and audience Templates per audience (user vs dev)
Generate Create documentation LLM + structured templates
Verify Verify doc-code coherence Drift detection, doctests, CI

The Verify phase is what separates useful documentation from harmful documentation. Use drift detection tools to catch when code and documentation diverge.

Drift detection and verification πŸ“

Drift between documentation and code is the silent enemy of doc-as-code. Data shows:

  • Nested READMEs discovered only ~40% of the time by agents
  • Documents in special directories (_docs/, docs/) discovered <10% of sessions
  • 60-70% of breaking changes pass code review without documentation updates

Drift detection tools

Tool Type Function
driftcheck Pre-push hook Compares doc with code before push
DriftGuard TypeScript Compiles code blocks in .md for verification
Spectral API linting Verifies OpenAPI spec matches implementation
Vale Style linting Verifies documentation style consistency

Enforcement hierarchy

Verification can operate on 3 increasing levels of strictness:

  1. Advisory β€” the agent flags but does not block
  2. PreToolUse hook β€” blocks changes that break doc-code coherence
  3. CI gate β€” the build fails if doc and code diverge

Safe pattern: “Propose, don’t auto-merge” β€” never update documentation directly. Always generate a PR for human review.

ADR: the most valuable section πŸ“‹

Architecture Decision Records are the most important documentation an agent can generate. The key section is “Alternatives rejected”: it documents not only what you chose, but why you discarded the alternatives. It is the most skipped and the most valuable section.

Y-statement format (compact and agent-friendly):

“In the context of {situation}, facing {concern}, I decided {decision} to achieve {goal}, accepting {tradeoff}.”

The ADR can become a verifiable CI constraint: instead of a static document, it becomes a rule that the agent and the pipeline verify at every build.

Changelog automation πŸ“¦

Generating changelogs manually is tedious and error-prone. Automation combines Conventional Commits with AI polish.

Tool landscape

Tool Language Approach Note
release-please Multi Conventional Commits β†’ PR Google, mature
git-cliff Rust Conventional Commits β†’ markdown Fast, configurable
changelogen UnJS Conventional Commits β†’ markdown Nuxt ecosystem
commitizen Python/Node Interactive conventional commits Interactive CLI
AutoChangelog Python AI polish on commit history New
GitSaga Claude AI-powered changelog Claude-based
Release Drafter GitHub Draft release from PR labels GitHub Action

Two-stage workflow

  1. Stage 1: Conventional Commits generate raw changelog (release-please or git-cliff)
  2. Stage 2: AI polish for readability and tone consistency

A changelog is not a commit dump. It is a communication to users. AI can transform “fix: resolve null pointer in auth handler” into “Fixed authentication crash when user has no profile”.

Acceptance criteria βœ…

  • Every statement verifiable against source code.
  • No invented APIs, parameters, or behaviors.
  • Documentation versioned with the code (same repo, same branch).
  • Correct format for the chosen type (README, ADR, etc.).

Further reading πŸ“š

Last updated on