Understanding an existing codebase π
Before touching a line of code, you need to understand what is already there. A coding agent is excellent at this: it can read the entire repository and return a map in minutes. But only if you ask the right questions and provide the right context. Otherwise you get a generic summary you could have read on Wikipedia.
An agent that understands a codebase is not an agent that has read everything: it is an agent that has read the right things in the right way.
When to use the agent for this case π―
- Onboarding on a new project: you joined a team and need to understand the structure, stack, and patterns in use.
- Architectural audit: you need to evaluate a project’s quality before a refactoring or a hiring decision.
- Targeted exploration: “where is authentication handled?”, “which services call this endpoint?”.
- First AGENTS.md draft: document stack, commands, and conventions in a living file.
When NOT to use the agent β
- The project is a 500k-line monolith with no clear structure β do manual archaeology first, then delegate.
- You need to modify code β this use case is read-only. If you need to intervene, move to adding features.
- The context is sensitive (personal data, secrets, credentials) β make sure the model does not process data it should not see.
Opening prompt π
Use this template as a starting point and adapt it to your case:
“Analyze the repository [path]. Objective: [architectural understanding / onboarding / audit]. Expected output: [module map, call flows, risks, hotspots]. Do NOT modify any file.”
Concrete examples:
- “Analyze the
src/auth/directory. Objective: understand the login flow. Expected output: Mermaid diagram of the flow, external dependencies, potential issues.” - “Audit the project in the current folder. Objective: evaluate architectural quality. Expected output: weak points, pattern violations, improvement suggestions.”
Context setup π§
To get useful results, prepare the ground:
- Indicate relevant directories β do not let the agent scroll through the entire monorepo. If you care about the
payment/module, start there. - Make sure
AGENTS.mdexists β if it does not, ask the agent to generate a draft as a first task (see configuring the repository). - Specify the level of detail β do you want a high-level overview or a deep analysis of a specific module?
Workflow π‘
Apply the planβexecuteβvalidate cycle even in read-only mode:
- Plan: define what you want to understand (architecture? flow? dependencies?).
- Explore: the agent reads relevant files and generates a structured analysis.
- Validate: verify that the agent’s claims are correct. Check cited files, compare with the actual code.
For very large repositories, consider persistent codebase indexing tools β a static map or queryable AST that reduces token consumption in orientation sessions.
Validation is crucial: an agent can invent relationships between modules that do not exist. Always verify against the actual code.
flowchart TD
A[Start session] --> B[Initial scan: glob + grep]
B --> C{Complex repo?}
C --> No --> D[Load AGENTS.md + relevant files]
C --> Yes --> E[Scoped exploration: target module]
E --> F[AST indexing: Tree-sitter / ast-grep]
F --> G[Map dependencies and interfaces]
G --> D
D --> H[Validate with tests / compilation]
H --> I[Output: structured map]
Indexing and persistent indexing π
For large repositories, token consumption becomes a real problem: every file loaded into the context window costs input tokens and degrades analysis quality. The solution is persistent indexing: a codebase map that survives across sessions and reduces the need to re-read everything from scratch.
Indexing mechanisms
| Mechanism | What it does | Precisione | Overhead | Tool |
|---|---|---|---|---|
| Line Chunking | Splits files into N-line blocks | Low | Minimal | Any |
| AST (Tree-sitter) | Parses syntactic structure | Medium | Low | tree-sitter, ast-grep |
| Call Graph | Maps function relationships | High | Medium | code-review-graph, Graphify |
| SCIP/LSIF | Deterministic cross-language index | 100% | High | SCIP indexer |
Tree-sitter (WASM version) is the ideal starting point: it parses code into atomic chunks of 50-1000 characters while maintaining syntactic structure. Aider RepoMap combines Tree-sitter + Ctags + PageRank to prioritize the most important symbols. ast-grep does structured pattern matching directly on the AST.
AGENTS.md: the golden rule
The ETH Zurich / LogicStar.ai study (Feb 2026, 138 tasks on SWE-bench Lite) demonstrated that:
- LLM-generated AGENTS.md: -3% success rate, +20% inference cost, +14-22% reasoning tokens
- Developer-written AGENTS.md: +4% success rate
- Dynamic Adaptive Context (ACE): +10.6% success rate
The rule is simple: the root
AGENTS.mdmust be under 50 lines. Details should be discovered progressively, not loaded all at startup.
Full-repo vs scoped exploration π
| Dimension | Full-repo | Scoped |
|---|---|---|
| Execution speed | Baseline | +28% |
| Context overhead | High (+20% cost) | Minimal |
| Result precision | High noise | Targeted |
| Cross-package contamination risk | High | None |
| Application | Simple repos (<10k lines) | Complex repos, monorepos |
The practical rule: if the repo has more than 10k lines or is a monorepo, scoped exploration is always preferable. The agent does not need total visibility: it needs fast paths to the structure that matters.
Anti-patterns in exploration π«
- Feeding dozens of files without scoping β context decay and degraded model recall.
- Accepting agent claims without verification β hallucination: the agent invents relationships between modules that do not exist.
- Incrementally adding rules to
AGENTS.mdwithout periodic review β “ball-of-mud” bloat that makes the file useless. - “Repeated discovery” in every session for large projects β use persistent indexing instead of re-reading everything from scratch.
Acceptance criteria β
- No modifications to project files (read-only).
- Structured output: Mermaid diagram or module map with legend.
- Every statement cites the reference file and line.
- First
AGENTS.mddraft has been generated and manually validated.
Further reading π
- Configuring the repository β how to create an effective
AGENTS.md. - Managing context as a resource β context hygiene techniques for targeted explorations.
- Blog: Context engineering β the theoretical foundation of context management.