Every developer hits this wall: a repository lands in your lap, the person who wrote it left eight months ago, the README is three commands out of date, and someone expects a bug fix by Friday. Learning to understand someone else’s code quickly is not a talent, it is a procedure. Below is the exact 7-step method we use when we pick up an unfamiliar codebase, with the commands you can copy and paste on your first morning.
Why reading code feels harder than writing it
When you write code, you build the mental model first and the syntax second. When you read code, you get the syntax and have to reverse engineer the mental model. Worse, most people try to read a repository like a book: top to bottom, file by file, alphabetically. That approach fails because a codebase is not linear, it is a graph. How to read someone else’s code tackles the same question from another angle.
The fix is to stop reading and start tracing. You pick one behaviour that you know the software does, then follow it through the system. Everything else stays a black box until you need it.

The 7-step method at a glance
| Step | Goal | Time budget |
|---|---|---|
| 1. Run it before you read it | Working local environment | 45 to 90 min |
| 2. Find the real entry points | Know where execution starts | 20 min |
| 3. Interrogate the git history | Find hot files and owners | 20 min |
| 4. Build a dependency graph | See the shape of the project | 20 min |
| 5. Map one data flow on paper | Write the model before the code | 30 min |
| 6. Trace it with breakpoints | Confirm reality vs assumption | 60 min |
| 7. Ship a tiny safe change | Prove your understanding | 60 min |
Total: roughly one working day. That is the promise. Not full mastery, but enough context to be useful and to stop asking questions that the repository already answers. It is argued more carefully on towardsdatascience.com.
Step 1: Run it before you read a single line
You cannot trace behaviour you cannot reproduce. Your first hour is DevOps, not reading.
- Clone the repo and check what tooling it expects:
cat .nvmrc .tool-versions .python-version 2>/dev/null - Read the scripts block, which is the real documentation in a JavaScript project:
cat package.json | jq .scripts(ornpm runwith no arguments to list them). - Look for a Makefile,
docker-compose.yml, or.env.example. Copy the example env file:cp .env.example .env - Install and boot:
npm ci && npm run dev - Run the test suite once, even if it fails:
npm test. A failing suite tells you what the team tolerates.
Write down every undocumented step you had to figure out. Your first pull request can be a README fix, which is the cheapest way for a new hire or freelancer to earn trust.
What the scripts block tells you
devandstart: the entry command and therefore the entry filebuild: the bundler, the framework, the output targetmigrate,seed,prisma,knex: there is a database and it has statelintandformat: the style rules you will be judged one2e,playwright,cypress: a free, executable description of the critical user journeys
End-to-end tests are the single most underrated resource when you need to read code written by someone else. They list the flows the business actually cares about.
Step 2: Find the real entry points
Every application has a small number of doors. Find them and the rest of the maze becomes navigable.
| Stack | Where execution starts |
|---|---|
| Node / Express | server.js, src/index.ts, app.use() chains, routes/ |
| Next.js | app/ or pages/ folder, middleware.ts, layout.tsx |
| React SPA | main.tsx, the router config, top level providers |
| Laravel / PHP | public/index.php, routes/web.php, routes/api.php |
| Django | manage.py, urls.py, settings.py |
| WordPress plugin | Main plugin file header, add_action and add_filter hooks |
| Any CLI | bin/ folder plus the bin key in package.json |
Then grep the doors instead of browsing folders. Use ripgrep, it is fast and respects gitignore:
- Routes:
rg -n "router\.(get|post|put|delete)" --glob '!node_modules' - Background work:
rg -n "cron|queue|worker|schedule" -g '!*.lock' - External calls:
rg -n "fetch\(|axios\.|http\.request" - Feature flags and config switches:
rg -n "process\.env\." -o | sort -u
That last command is gold. The list of environment variables is a summary of every third party service, secret and toggle in the system.

Step 3: Interrogate the git history
The repository has a memory. Use it before you ask a human. Git history tells you which files matter, who to ask, and where the bugs live.
- Who owns what:
git shortlog -sne --since="18 months ago" - Change hotspots (files touched most often are usually the core, or the mess):
git log --since="12 months ago" --name-only --pretty=format: | grep -v '^$' | sort | uniq -c | sort -rn | head -30 - Recent direction of travel:
git log --oneline --no-merges -40 - Files that always change together (hidden coupling): pick a hot file and run
git log --format=%H -- path/to/file.ts, then inspect those commits. - Who last touched the line that confuses you:
git blame -L 40,80 path/to/file.ts - When a behaviour broke:
git bisect start, thengit bisect badandgit bisect good <tag>
Read the pull request titles too. Six months of PR titles is a better changelog than most changelogs.
Step 4: Build a dependency graph
Now get a picture of the shape. You are looking for layers, clusters and circular dependencies. How to Read and Understand Someone Else Code covers this in more depth.
- Size and language mix:
npx cloc . --exclude-dir=node_modules,distortokei - Module graph for JS and TS:
npx madge --extensions ts,tsx src --image graph.svg - Circular dependencies, the classic source of confusing behaviour:
npx madge --circular src - Direct third party dependencies:
npm ls --depth=0 - Python import graph:
pydeps yourpackage --max-bacon 2 - Unused code candidates:
npx knipornpx depcheck
Open the SVG and look for the nodes with the most arrows pointing at them. Those five or six files are the spine of the application. If you understand them, you understand roughly 60 percent of the project.
Step 5: Map the data flow before you touch any code
This is the step almost everyone skips, and it is the one that saves days. Pick one concrete user action, ideally the most valuable one in the product: a checkout, a login, a report export, an order submission. Then write the flow down as a numbered list, guessing where you have to guess.
Example for a quote request form:
- User submits form component (which file?)
- Client validation with a schema library (which schema?)
- POST to an API route (which path?)
- Auth or rate limiting middleware
- Service layer function, business rules applied
- Database write plus an outbound email or webhook
- Response shape returned to the UI, state updated, toast shown
Keep it in a scratch file called NOTES.md. You now have a hypothesis, and a hypothesis makes the next step surgical instead of exploratory. This is the difference between wandering a codebase and interrogating it.
Draw the three maps
- Request map: where input enters and how it is validated
- State map: what is stored in the database, in cache, in browser state, and who is the source of truth
- Boundary map: every external system: payment provider, CRM, email, analytics, file storage

Step 6: Trace the flow with breakpoints, not print statements
Now verify the hypothesis with a debugger. Reading tells you what the author intended. The debugger tells you what actually happens.
- Set a breakpoint at the entry point you identified in step 2, not deep inside the logic.
- Trigger the action in the browser or with
curl. - Step over, not into, until something surprises you. Then step in.
- Inspect the call stack panel. The call stack is the free architecture diagram nobody reads. Screenshot it and paste it into your notes.
- Watch the payload mutate between layers. Note every place the shape changes, because those are the seams where bugs live.
Useful specifics:
- Node:
node --inspect-brk ./node_modules/.bin/next dev, then openchrome://inspect - Node scripts: add
debugger;and runnode --inspect script.js - Python:
breakpoint()thenbt,args,wherein pdb - Frontend: React DevTools profiler plus a conditional breakpoint such as
id === 42 - SQL: enable query logging for one minute and read the actual queries issued by the ORM
- Network: the browser Network tab, sorted by initiator, maps UI to API in seconds
If a debugger is genuinely unavailable, use structured logging with a correlation id rather than scattered console.log calls, and remove them before committing.
Step 7: Prove your understanding with a tiny, safe change
You do not understand a codebase until it has responded to you. Close the loop with the smallest possible contribution:
- Write one characterisation test that asserts current behaviour of the flow you traced
- Rename one badly named local variable inside a single function
- Add a docblock summarising the flow at the entry point
- Fix the README steps you had to reverse engineer in step 1
- Commit your
NOTES.mdasdocs/architecture-notes.mdif the team allows it
Then open a small pull request. Reviewer comments on a five line diff will teach you more about the project’s conventions than a week of silent reading.
A realistic day-one timebox
| Time | Activity | Deliverable |
|---|---|---|
| 09:00 to 10:30 | Steps 1 and 2 | App running, entry points listed |
| 10:30 to 11:30 | Steps 3 and 4 | Hotspot list, dependency graph SVG |
| 11:30 to 12:00 | Step 5 | Written hypothesis of one flow |
| 13:00 to 14:30 | Step 6 | Verified flow plus call stack notes |
| 14:30 to 16:00 | Step 7 | First small pull request open |
| 16:00 to 17:00 | Questions | One batched list, not ten interruptions |

Common traps when you try to understand someone else’s code
- Reading alphabetically. Folder order has nothing to do with execution order.
- Refactoring on day one. Ugly code often encodes a real edge case. Understand first, judge later.
- Trusting comments and names. Comments rot. The debugger does not lie.
- Ignoring configuration. Half of modern behaviour lives in env vars, feature flags and infrastructure files.
- Going too deep too early. Black box anything that is not on the path of the flow you chose.
- Using an AI assistant as a substitute for tracing. Asking a model to summarise a file is a decent first pass, but it will confidently invent a data flow that does not exist. Verify with a breakpoint.
Tooling cheat sheet
| Need | Tool or command |
|---|---|
| Fast search | rg (ripgrep), fd |
| Repo size and languages | cloc, tokei |
| Module graph | madge, dependency-cruiser, pydeps |
| Dead code | knip, depcheck, vulture |
| History analysis | git shortlog, git blame, git bisect |
| Navigation in editor | Go to definition, find all references, call hierarchy |
| Runtime truth | Debugger breakpoints, Network tab, query logs |
The one sentence version
Run it, find the door, ask the history, draw the graph, guess the flow, verify with a breakpoint, then ship something small. Repeat that on every new project and unfamiliar repositories stop being intimidating. They become just another system waiting to be interrogated.
FAQ
How long should it take to understand someone else’s code?
One day for orientation and one traced flow. One to two weeks to be productive on small tickets. One to three months for genuine ownership on a medium sized codebase. Anyone promising instant mastery is selling something.
Should I read the code or run the debugger first?
Read just enough to form a hypothesis, then run the debugger to confirm it. Pure reading is slow and pure stepping is directionless. The combination is what makes the method fast.
What if the codebase has no tests and no documentation?
Then git history and the debugger are your documentation. Start by writing one characterisation test around the flow you traced. It becomes the first line of documentation and a safety net for your first real change.
How do I handle a codebase with terrible naming?
Do not fix names globally on day one. Rename only inside the function you are actively working on, keep a glossary in your notes mapping cryptic names to their real meaning, and use your editor’s rename symbol refactor so references stay consistent.
Can AI tools help me understand an unfamiliar repository?
Yes, for three things: summarising a single file, explaining an unfamiliar syntax or library, and generating a first draft of a call graph description. No, for authoritative answers on data flow, side effects or business rules. Always confirm with a breakpoint or a test.
What should a freelancer ask the client before starting?
Four questions: how do I run it locally, which flow generates the money, who was the last developer to touch it, and is there a staging environment. Everything else you can find in the repository yourself.

