Making legacy code understandable: How we reconcile code and documentation

By Mihaly Fodor (ERNI Romania)

Legacy code outlives the people who wrote it. Documentation drifts, knowledge disappears, and new team members pay the price. This article explores how LeRA uses deterministic analysis, LLMs and provenance to reconcile code and documentation – without letting AI make the decisions.

Why we built LeRA on atomic statements

Understanding legacy projects is a recurring theme for many of our clients. Code outlives the people who wrote it, documentation drifts away from what is running, and every new person on a project pays the ‘tax’ of rediscovering what everyone before them already knew but forgot to write down. It is not a glamorous problem, but it is one I encounter repeatedly, project after project. This summer gave me a chance to do something about it. We set up an internship for university students and were looking for a topic that would address a real need and produce a solution. I remember a call with a colleague around the same time, while I was already turning the idea over in my head, where he mentioned that a product owner had just left a project, and that a tool to help understand the existing state of play would be incredibly timely. That call is where the project, which we eventually named Legacy Requirements Assistant (LeRA), began.

Two audiences, one problem

The gap shows up differently depending on who hits it. For a new PO stepping into an unfamiliar project, the question is: what was documented, where it was documented, and how well that documentation is supported by the underlying code. For developers, the question is different: could this become an extension of their normal workflow, so their coding agent doesn’t have to hunt for context every time, and can reach deeper into the codebase than a quick grep would allow? Both angles pointed toward the same underlying need, so the project was born with a single goal: ingest documentation and code, break both down into atomic statements, and run a reconciliation phase that surfaces differences and gaps in understanding. Internally, we ended up calling those atomic statements ‘grunts’, and the name stuck.

Rules we set before writing code

Before the students wrote a single line of the pipeline, we agreed on a few ground rules, mostly because we had been burned by loose approaches before:

  • Rely on static analysis before using an LLM, as much as possible. This cuts down on hallucinations and prevents non-determinism from creeping into the results.
  • LLMs cannot decide anything. The workflow surfaces information; it does not make decisions on our behalf. That single constraint is what makes the output easy to trust, validate, and trace.
  • Get onto a real codebase as early as possible. We wanted every design choice checked against actual code rather than a toy example.
  • Build on atomic statements from the start. That was a lesson we had learned from previous initiatives.

After three months of intense work, we were doing LeRA demo presentations for our existing customers with genuinely good results and strong interest. Our plan now is to continue extending its capabilities, rather than calling it ‘finished’.

A deterministic core first

So, what makes this work? The part I am most proud of is the core, because it is boring in the best possible way. In the initial version, LeRA ingests .NET and Angular codebases through 37 extractors, identifying 142 types of facts—all without touching an LLM. Documentation is also parsed and linked to code deterministically: markdown, reStructuredText, AsciiDoc, plain text, and PDF, DOCX and PPTX files all go in.

This deterministic core sits inside a 16-stage pipeline. This covers discovery, inventory, the extractors, and the initial facts projection into template statements. The later stages – statement generation, embeddings, relations, and reconciliation – can optionally call an LLM, working on top of a scaffold that was built without it, rather than the LLM filling in blanks on its own.

Reconciliation without a black box

Reconciliation is where documentation and code are compared with each other. It is the stage I get asked about most, usually with some version of: “So it just finds where the docs are wrong?” It is more careful than that: the contradiction detectors that look for statements that disagree with one another are rule-based, and they always run, regardless of whether enrichment is enabled.

An LLM judge only gets involved to confirm candidate pairs when enrichment is enabled, and even then, it is confirming, not deciding; every conflict it finds still goes to a human to confirm or reject. LeRA never resolves a contradiction on its own – on purpose – in line with our rule about LLMs not making decisions. It is also worth being precise about what ‘gaps’ means here: we focus on documentation coverage (which functions, files or capabilities have a doc link and which do not), rather than a report of things documented but missing from the code – a much harder problem that we have deliberately not claimed to solve yet.

Provenance is not optional

Every statement that survives the pipeline carries its source file and span, a confidence value, and the method that produced it. Statements that cannot be grounded in something concrete are dropped by a gate before they reach a human. We did not want a tool that produces a confident-sounding sentence with no way to check its source; that is exactly the kind of thing that erodes trust the first time a user finds an error. On top of that, the entire pipeline is traced with OpenTelemetry: LLM calls, embedding calls and pipeline events all show up. When something looks off, we can see exactly what ran and in what order, instead of guessing.

What it costs to run

None of this is free, so we track it. Actual LLM and embedding spend is measured per scan against a budget cap, and before a scan even starts, a cost estimator provides a rough figure based on the size of the codebase and the volume of documentation. We do not have hard numbers on time saved or accuracy gained yet; the internal demos went well, but that is a qualitative statement, not a metric I am comfortable publishing.

Where this goes next

Right now, LeRA consists of a web UI for POs and developers, backed by a REST API. The ‘developer-harness’ angle from that original phone call – an agent querying LeRA directly instead of grepping around a codebase – is the direction we are heading in. I think it is the more interesting half of the original idea, and it is where most of the extension work will go. Three months in, what started as an internship topic sparked by a colleague’s bad week is now being shown around for existing and new customers.

Mentoring the interns through this was the most fun I have had at work this year. The result is a great starting point after just one summer, and a good reason for us to keep going.

Are you ready
for the digital tomorrow?
better ask ERNI

We empower people and businesses through innovation in software-based products and services.