Skip to content
Understanding a codebase

Understanding an existing codebase πŸ”

Before touching a line of code, you need to understand what is already there. A coding agent is excellent at this: it can read the entire repository and return a map in minutes. But only if you ask the right questions and provide the right context. Otherwise you get a generic summary you could have read on Wikipedia.

An agent that understands a codebase is not an agent that has read everything: it is an agent that has read the right things in the right way.

When to use the agent for this case 🎯

  • Onboarding on a new project: you joined a team and need to understand the structure, stack, and patterns in use.
  • Architectural audit: you need to evaluate a project’s quality before a refactoring or a hiring decision.
  • Targeted exploration: “where is authentication handled?”, “which services call this endpoint?”.
  • First AGENTS.md draft: document stack, commands, and conventions in a living file.

When NOT to use the agent β›”

  • The project is a 500k-line monolith with no clear structure β€” do manual archaeology first, then delegate.
  • You need to modify code β€” this use case is read-only. If you need to intervene, move to adding features.
  • The context is sensitive (personal data, secrets, credentials) β€” make sure the model does not process data it should not see.

Opening prompt πŸ“

Use this template as a starting point and adapt it to your case:

“Analyze the repository [path]. Objective: [architectural understanding / onboarding / audit]. Expected output: [module map, call flows, risks, hotspots]. Do NOT modify any file.”

Concrete examples:

  • “Analyze the src/auth/ directory. Objective: understand the login flow. Expected output: Mermaid diagram of the flow, external dependencies, potential issues.”
  • “Audit the project in the current folder. Objective: evaluate architectural quality. Expected output: weak points, pattern violations, improvement suggestions.”

Context setup πŸ”§

To get useful results, prepare the ground:

  1. Indicate relevant directories β€” do not let the agent scroll through the entire monorepo. If you care about the payment/ module, start there.
  2. Make sure AGENTS.md exists β€” if it does not, ask the agent to generate a draft as a first task (see configuring the repository).
  3. Specify the level of detail β€” do you want a high-level overview or a deep analysis of a specific module?

Workflow πŸ’‘

Apply the plan→execute→validate cycle even in read-only mode:

  1. Plan: define what you want to understand (architecture? flow? dependencies?).
  2. Explore: the agent reads relevant files and generates a structured analysis.
  3. Validate: verify that the agent’s claims are correct. Check cited files, compare with the actual code.

For very large repositories, consider persistent codebase indexing tools β€” a static map or queryable AST that reduces token consumption in orientation sessions.

Validation is crucial: an agent can invent relationships between modules that do not exist. Always verify against the actual code.

    flowchart TD
    A[Start session] --> B[Initial scan: glob + grep]
    B --> C{Complex repo?}
    C --> No --> D[Load AGENTS.md + relevant files]
    C --> Yes --> E[Scoped exploration: target module]
    E --> F[AST indexing: Tree-sitter / ast-grep]
    F --> G[Map dependencies and interfaces]
    G --> D
    D --> H[Validate with tests / compilation]
    H --> I[Output: structured map]
  

Indexing and persistent indexing πŸ”Ž

For large repositories, token consumption becomes a real problem: every file loaded into the context window costs input tokens and degrades analysis quality. The solution is persistent indexing: a codebase map that survives across sessions and reduces the need to re-read everything from scratch.

Indexing mechanisms

Mechanism What it does Precisione Overhead Tool
Line Chunking Splits files into N-line blocks Low Minimal Any
AST (Tree-sitter) Parses syntactic structure Medium Low tree-sitter, ast-grep
Call Graph Maps function relationships High Medium code-review-graph, Graphify
SCIP/LSIF Deterministic cross-language index 100% High SCIP indexer

Tree-sitter (WASM version) is the ideal starting point: it parses code into atomic chunks of 50-1000 characters while maintaining syntactic structure. Aider RepoMap combines Tree-sitter + Ctags + PageRank to prioritize the most important symbols. ast-grep does structured pattern matching directly on the AST.

AGENTS.md: the golden rule

The ETH Zurich / LogicStar.ai study (Feb 2026, 138 tasks on SWE-bench Lite) demonstrated that:

  • LLM-generated AGENTS.md: -3% success rate, +20% inference cost, +14-22% reasoning tokens
  • Developer-written AGENTS.md: +4% success rate
  • Dynamic Adaptive Context (ACE): +10.6% success rate

The rule is simple: the root AGENTS.md must be under 50 lines. Details should be discovered progressively, not loaded all at startup.

Full-repo vs scoped exploration πŸ”€

Dimension Full-repo Scoped
Execution speed Baseline +28%
Context overhead High (+20% cost) Minimal
Result precision High noise Targeted
Cross-package contamination risk High None
Application Simple repos (<10k lines) Complex repos, monorepos

The practical rule: if the repo has more than 10k lines or is a monorepo, scoped exploration is always preferable. The agent does not need total visibility: it needs fast paths to the structure that matters.

Anti-patterns in exploration 🚫

  • Feeding dozens of files without scoping β†’ context decay and degraded model recall.
  • Accepting agent claims without verification β†’ hallucination: the agent invents relationships between modules that do not exist.
  • Incrementally adding rules to AGENTS.md without periodic review β†’ “ball-of-mud” bloat that makes the file useless.
  • “Repeated discovery” in every session for large projects β†’ use persistent indexing instead of re-reading everything from scratch.

Acceptance criteria βœ…

  • No modifications to project files (read-only).
  • Structured output: Mermaid diagram or module map with legend.
  • Every statement cites the reference file and line.
  • First AGENTS.md draft has been generated and manually validated.

Further reading πŸ“š

Last updated on