← Back to homepage work

Case study / Summer 2026

CodeGraph

Explore a repository through syntax, relationships, and question-specific retrieval instead of sending all of its source to an agent.

System
Repository-exploration prototype
Interface
React Flow graph + source drawer
Retrieval
Seven LangGraph tools over Neo4j
View CodeGraph repository ↗

The problem

An unfamiliar repository presents two related problems: finding relevant code and understanding how it connects. A full-source prompt is large, while isolated snippets can lose caller and dependency context. CodeGraph builds a navigable structural index and lets an agent choose smaller observations for the current question.

System architecture

A React dashboard submits a GitHub repository to FastAPI, displays a React Flow graph, and fetches source separately when a function is selected. The backend combines Tree-sitter parsing, Neo4j relationships, metadata embeddings, Leiden communities, and a Claude-powered LangGraph ReAct agent.

Two paths through one repository index

  1. 01 / Interface

    React dashboard ↔ FastAPI

    Repository submission, graph exploration, source inspection, and questions.

    Ingestion path ↓

  2. 02 / Indexing

    GitHub archive → Tree-sitter → ETL

    Safe temporary extraction; supported-source discovery; syntax records; three-pass graph writes.

    Persist + enrich ↓

  3. 03 / Storage

    Neo4j graph · vectors · communities

    Repository → File → Function; inferred CALLS edges. OpenAI name/path embeddings and sampled community labels enrich the index.

    Question path: agent ↔ tools ↔ index ↓

  4. 04 / Retrieval

    LangGraph ReAct ↔ Claude

    Seven tools return structure, direct callers, outgoing calls, unresolved external names, semantic matches, and community summaries. Tool observations inform the final answer.

Ingestion creates the index; questions retrieve from it. Leiden groups the Function/CALLS graph before questions are asked. Semantic search embeds the query only when that tool is selected. Source inspection is a separate dashboard endpoint, not an eighth agent tool. The agent does not automatically expand communities into source-code context.

From ingestion to context

  1. Acquire and discover

    Validate an HTTPS GitHub owner/repository URL, download its default-branch archive, and extract with path-containment checks. Discover Python, JavaScript/JSX, TypeScript/TSX, Go, and Java source.

  2. Count and parse

    Count supported-source tokens and capture definitions, source spans, and lexical call records with Tree-sitter. Up to 32 per-file stages are active; a per-language lock serializes shared parser use. File failures are isolated, so completion can still mean incomplete coverage.

  3. Write and enrich

    Replace the old repository scope, then write files, functions with metadata embeddings, and calls in 100-record batches. Run undirected Leiden clustering and generate labels from sampled names and paths. These batches are separate transactions.

  4. Expose the index

    Stream newline-delimited progress records; the client buffers partial lines and requires a terminal completion record. Fetch graph data after success. The graph endpoint caps query rows at 200 and leaves raw source and vectors out of the bulk payload.

  5. Retrieve for the question

    Start a fresh system/user message pair. The agent chooses tools, receives formatted observations, and can call again before answering. Metadata-vector candidates are selected globally, then filtered by repository. There is no persistent conversation memory or enforced source-citation contract.

Engineering challenges

Syntax identity must survive storage

The parser now emits canonical IDs and lexical caller ownership, but the active writer still keys functions by repository plus simple name and assigns file-wide calls to every function. Duplicate names can merge, and false edges can contaminate source lookup, retrieval, and communities.

Progress is not a transaction

NDJSON makes a long ingestion visible, but HTTP 200 can carry a terminal error. Replacement deletes the prior scope before all enrichment finishes; later failure can leave partial graph state. The transport does not provide rollback or resumable jobs.

Compact evidence can still be wrong

Metadata embeddings describe names and paths rather than function behavior. Global top-k selection followed by repository filtering can lose relevant candidates. Direct caller queries represent one-hop graph evidence, not complete runtime impact.

Design decisions & tradeoffs

Tree-sitter for syntax

Benefit
Multi-language structure without executing repository code.
Tradeoff
Syntax captures do not resolve imports, dynamic dispatch, or every call target.

Graph, vectors, and GDS together

Benefit
Neo4j supports relationships, semantic candidates, and community enrichment in one store.
Tradeoff
Retrieval inherits graph inaccuracies; generated communities are exploratory groupings, not verified module boundaries.

Tools plus lazy source inspection

Benefit
Question-specific observations and a separate source drawer keep bulk graph responses smaller.
Tradeoff
The agent has no source tool or grounding validator; the drawer does not make its answers source-verified.

Bounded stages and batches

Benefit
Limit active per-file work and write in finite transaction units.
Tradeoff
Parsed results remain in memory, same-language parsing serializes, and partial commits remain visible. These limits do not prove a throughput gain.

Correctness & validation

The supplied audit establishes selected parser, API-helper, formatting, and scoping behavior. It also exposes limits that a successful build or mocked test cannot settle.

Backend / supplied audit
92 passed · 4 skipped

Temporary dependency overlay; TCP connections disabled. The four skipped cases require a live Neo4j database.

Frontend / supplied audit
13 Node tests passed

API helpers and utilities; no browser or component execution. The project’s Vite build also passed.

These are the manual’s September 24, 2026 results at revision 79fa768, not tests rerun for this portfolio. Live database, ANN, GDS, model-provider, browser, and Postman workflows were not exercised in that audit.

Path-safe extraction, parameterized queries, and raw-body HMAC verification are documented controls. Repository predicates are data selectors, not access control: the prototype has no authentication, and agent tool scope remains model-supplied.

Results & measurement

Reported LLM context / per query

≈737Kto≈18K

tokens per query

LLM context reduction
97.5%

Supplied résumé result · supported-source tokens compared with tool-returned context.

Baseline = token counts summed across discovered, supported source files. Ignored or unsupported files are excluded; a file counted before a parsing failure can still contribute.

Selected context = concatenated tool-message text in the final agent trace, counted with a proxy tiktoken encoding. The default encoding is for gpt-4o, with cl100k_base fallback.

Context reduction = (baseline − tool context) / baseline × 100

This comparison excludes system/user prompts, answer tokens, and repeated history sent across model rounds. It does not measure Claude billing, total cost, latency, or answer accuracy. An answer that calls no tools can show 100% reduction while having no retrieved evidence.

Evidence status / reported result

The supplied résumé reports these approximate counts. The manual found no reproducible run artifact: fixed repository revision, question set, raw traces, aggregation definition, and independent answer scores remain missing. Available evidence does not establish a general performance guarantee.

Engineering lessons

Carry provenance end to end

Parser IDs and caller ownership only improve downstream answers when writers, edges, queries, and source lookup preserve them. Schema migration is part of retrieval correctness.

Publish complete snapshots

A staged index with atomic activation would protect the last complete repository view. Streaming progress alone cannot supply consistency or recovery.

Evaluate evidence alongside context size

A fixed-revision, source-labeled question set should compare retrieval variants at matched budgets and score evidence recall, grounded answers, and appropriate uncertainty. Token reduction alone is insufficient.

Repository

Explore the parser, graph pipeline, retrieval tools, and dashboard in the supplied repository. This case study describes the manual’s inspected revision; it does not claim a hosted demo or production deployment.