docgen

Turn source artifacts — a codebase, a database, an API spec, a recording of real traffic — into one canonical graph format, then render that format to whatever the audience needs. The point is not the diagram. The point is the format in the middle.

What it is

Eight demos under semester/ draw their architecture with Graphviz. Every one of them hand-writes a .dot file with the palette inlined and commits the .svg beside it. Three different dark palettes between them, eight answers to the same question, and every one of those diagrams is out of date the moment somebody moves a file.

docgen is the one answer. It reads the source rather than being told about it, so a diagram cannot drift from the thing it describes, and it separates what a graph is from how it looks so one extraction feeds a diagram, an index, a notebook and a browsable site without being redone.

Nothing here writes a parser or a graph algorithm. Parsers are adopted (ast, tree-sitter, SQLAlchemy reflection via modelgen), algorithms are networkx's. What docgen owns is the adapters, the schema, the style tables and the emitters — all small, and all the places where the value is that we made the call.

The idea

N sources and M outputs need N×M converters if you join them directly, or N+M if you put a hub in the middle. The hub is an intermediate representation — the compiler term, and the same bargain: both sides depend on the IR and neither on the other.

It is lossy on purpose. It throws away every token of syntax and keeps "a class named User inherits from Base". That is the part that is worth versioning, worth diffing, and worth drawing.

The practical consequence is the thing to judge it on: adding a source costs one extractor and every emitter works on it unchanged; adding an output costs one emitter and every extractor feeds it unchanged. When the OpenAPI reader was written it emitted schemas using the same vocabulary the database reader uses — and the ER diagram drew an API's data model without anyone teaching it what an API was.

Five minutes

One command, if you want the whole thing — a book: what went in, every step, what came out, and a page to open.

make book SRC=/path/to/repo BOOK=out/book/mine
make check BOOK=out/book/mine
  larder   /path/to/repo — 45 files read, 2 failed, 12 packages
  book     312 nodes · 244 edges · 26 external · 8 artifacts
  ok       45 file(s) read produced 45 module(s)
  open     out/book/mine/site/index.html

Underneath it is three commands that compose, and each still works on its own. That is the interface, and the book does not replace it:

# 1. read something
python3 -m docgen.extractors.python --root ../station/tools/histgen -o ir.json

# 2. narrow it to a useful view
python3 -m docgen.ops ir.json --overview -o view.json

# 3. draw whatever its structure asks for
python3 -m docgen.emitters auto view.json -o out/

Or through the Makefile, which is a thin wrapper over exactly those:

make ir SRC=/path/to/repo OUT=out   # extract
make explore OUT=out                # the two-pane navigator
make site OUT=out                   # a docs site with a sidebar
make self                           # docgen's book of soleprint, then check it

Everything is offline and self-contained. No server, no CDN, no build step — the outputs open over file://.

The three concerns

DOT collapses three separate questions into one file format, which is why a hand-written .dot is never reusable: you cannot change the palette without editing the structure, and you cannot change the structure without re-deciding the layout. docgen keeps them apart.

concernquestionowner
structurewhat the graph isir/schema.json
meaningwhat things mean visuallystyle/*.json, keyed on kind
placementwhere things gothe emitter, and only there

An extractor has never heard of SVG, colours or layout. An emitter has never heard of Python, ast or SQL. Both halves of that are checked by parsing the source and looking at what it imports, because a rule nobody enforces is a rule that lasts about a month.

docgen's own module structure
docgen read by docgen. extractors/ reaches only ir; emitters/ reaches ir and style; ir/ reaches nothing outside itself; lab/ has no edges at all. Click to open the viewer — then click again for actual size.

The book

A book is one docgen operation, and it has a fixed shape: it begins by saying what went in and ends by saying what came out. Everything in between is an ordinary file that stands on its own.

larder ──► step ──► step ──► step ──► book
  what        each one usable          what
  came in     by itself                came out

The word comes from Atlas 1.0, where a book is a larder composed with a pattern into something published and served. docgen had already built all three parts under other names, so this is less an adoption than a renaming back.

Why both ends, rather than just the result

Because the two numbers are only worth having together. "1,505 nodes" is not a fact about anything. "225 files in, 1,505 nodes out, nothing lost" is. Before this, an extractor that read 45 of 47 files produced exactly the same document as one that read all 47, and the diagram looked complete either way — there was nowhere for the other two to be mentioned.

A clean diagram over an incomplete read is a lie by omission. It is also the failure mode the IR already guards against one level down: an unresolved name becomes an external node rather than being dropped, because silently losing a thing is worse than recording an unresolved one. The larder measure is that same rule applied to the input as a whole.

The larder — what came in

Deliberately not called a bucket. A bucket is somewhere bytes sit; a larder is stocked from outside, has an inventory, and goes stale. All three are worth measuring, and they are what the measure records:

"larder": {
  "kind":     "python",           // which extractor stocked it
  "identity": "../../station",    // path, or a DSN with the password masked
  "unit":     "file",             // file | table | path | entry | document
  "seen":     47,                 // what the larder offered
  "read":     45,                 // seen - len(failed), derived
  "failed":   [{"name": "a.py", "error": "syntax: line 3"}],
  "extra":    {"packages": 12}
}

read is derived and never stored. Stored, it invites the question "does that include the failures?" and every reader answers it differently; derived, there is nothing to get wrong — and ir/validate.py fails a document whose arithmetic disagrees with itself.

Failures are recorded by name, not counted. A count tells you a book is incomplete; a name tells you which part of it to distrust.

identity is the one field in docgen that could carry a secret — a database DSN has the password in it. It is masked at construction, and validate.py then sweeps for the mask having worked, using its own independent key list. A scrubber graded by its own word is not graded.

The book measure — what came out, and reconciled

Counts by kind, edges by kind, externals, and every artifact with its byte count. On its own that is just a summary. What makes it a measure is that it is reconciled against the larder:

  larder   ../../station — 45 files read, 2 failed, 12 packages
  book     312 nodes · 244 edges · 26 external · 8 artifacts
  ok       45 file(s) read produced 45 module(s)
  ok       2 file(s) could not be read
  ok       2 unreadable file(s) appear as 2 marked module(s)

The relation differs by source and is declared per extractor, because getting it wrong gives a check that passes for the wrong reason. One file becomes one module node — fewer means input was dropped. Four hundred HAR entries becoming twelve endpoints is not a loss, it is the point of the capture. And a file that failed to parse must still appear in the graph, carrying its error, or the picture is smaller than the source and says nothing about it.

python3 -m docgen.book exits 1 when a reconciliation fails. The book is still written — the evidence is the point — but a build that lost input should fail a pipeline rather than pass quietly.

The notebook is the sequence, the web is the last step

These two rule what gets generated, and each for its own reason.

The notebook is not one artifact; it is the sequence. Its first cell is the larder measure and its last cell is the book measure, which is what puts the two ends in the document rather than only in the tooling. Between them, one pair of cells per step: what the step did, and a cell that loads that step's artifact and prints one fact about it. That is the "usable by themselves" property made executable, and the test suite runs those cells.

The web output is last, so nothing depends on it, so it can be replaced wholesale without touching anything upstream. That is exactly what lets it rule the output without being a stable contract — the book measure is the promise, and the page displaying it is free to change drastically and often. It is also the artifact somebody definitely opens, which is why "2 of 47 files could not be read" has to appear there, above the diagram rather than below it.

The spine is scaffolding, not a gate. Running one step alone is still a book, just a short one — make ir works exactly as it did. An operation that cannot measure something says what it could not measure and carries on. Gating would destroy the property that makes the intermediate artifacts useful, which is the whole reason the sequence is worth having.

What a book looks like on disk

book/<slug>/
├── book.json        both measures, the steps, artifacts with byte counts
├── steps/           every intermediate — ir.json, view.json, graph.svg, …
├── notebook.ipynb   the sequence; first and last cells are the measures
├── overlay.json     hand-written, optional, re-applied every build
├── checks.py        this book's own assertions — optional
└── site/            the web output, both measures at the top
make book SRC=../station BOOK=out/book/station
make check BOOK=out/book/station

The IR

Plain JSON. Three keys, and it has survived four domains without gaining a fourth.

{
  "meta":  { "source": "python", "root": "app/", "schema_version": "1" },
  "nodes": [ { "id": "app.models.User", "kind": "class", "label": "User",
               "parent": "app.models",
               "attrs": { "file": "app/models.py", "line": 12, "lines": 40 } } ],
  "edges": [ { "source": "app.models.User", "target": "app.db.Base",
               "kind": "inherits", "attrs": {} } ]
}
fieldmeaning
id Fully qualified and stable across runs. Stability is what makes two extractions from two commits diffable; without it a diff reports noise and nobody trusts it.
kind The hinge of the whole system, and the only field style and layout may read. A small closed vocabulary per domain — module/class/function, table/column, endpoint, task.
parent Containment, and nothing else. A module contains a class. Relationships are edges.
attrs An open bag for whatever one domain cares about. file/line/lines are what let a box link to the line it came from, and what the minimap sizes by.
No visual information, ever. If a field would change between a light and a dark theme, it does not belong in the IR. shape: "cylinder" is not a field — it is kind: "datastore" plus a style rule, and that is exactly what lets the same IR render in a theme that has no cylinders. The test suite sweeps every emitted document for colour-like keys.

Stdlib dataclasses, not Pydantic

The IR's whole value is being a plain document anything can open. A format that needs a library installed to be read is an API, not a format. Validation is therefore a function called at the boundary rather than a property of the type, and it reads its field lists out of schema.json so the schema and the dataclasses cannot drift apart.

python3 -m docgen.ir ir.json

It catches what a schema cannot: an edge naming a node that does not exist, a containment cycle, a duplicate id, and a visual field smuggled into attrs.

Shape decides the drawing

A diagram that fights its layout engine is usually the wrong kind of diagram. The clearest evidence: the same 24-table database rendered 32034×136 through Graphviz — a 235:1 strip — and 1740×1860 through the ER emitter. Not because one engine is better, but because a schema is a set of peer entities with references, and laying it out in dependency ranks was never its shape.

So ops.classify() reads the structure and names the emitter, with the reason attached — advice without a reason gets overridden the first time it is inconvenient.

kinddrawn bywhen
erdcards in columnsentities with references
pipelineranks, left to righta chain with fan-out — an Airflow DAG, a build
layered / treeranks, top downranks genuinely suit it
sheetthe indexone level is wider than ~20 — a strip in any engine
flatthe indexmost nodes have no relationships: that is a list
$ python3 -m docgen.emitters auto view.json -o out/
  sheet    -> index
           109 nodes sit at one level; any layered engine draws that as a
           strip. Split it, scope it, or read it as an index

Why twenty

Measured, one diagram per subsystem: at or under 20 nodes the output lands around 1.6:1; at 70–106 nodes about 7:1; at 261 nodes 14:1. Aspect ratio is a property of the graph, not of the renderer — a layered engine puts one dependency level in one row, so the widest level is the width.

Every Graphviz lever was tried before concluding this. ratio=compress squashed a graph to an unreadable 1008×75; rankdir=LR merely rotated a 14:1 into a 1:6; packing disconnected components gained nothing. The fix was never a flag. It was to stop asking for one picture of everything — which is what explore does.

Extractors

Deterministic parsing only. No model in the structural path. A diagram built from an AST cannot be out of date with the code. A diagram built from a model's reading of the code is wrong the moment the model has a bad day — which is the problem this exists to fix.

readerreadsgives
pythona tree of .py modules, classes, functions; imports and inherits edges
code tree-sitterC#, TypeScript, TSX namespaces, classes, interfaces, methods — structure only
dba graphgen-compatible schema.json tables, columns, foreign keys
openapian OpenAPI / Swagger document endpoints and the shapes they carry
usagea HAR recording what was actually called, in what order

Two passes, because ast resolves nothing

Given class User(Base), Python's ast hands over the literal string "Base". It has no idea that came from from .db import Base three lines up. So pass one collects, per module, what it defines and what it imports; pass two resolves local names to fully qualified ids. The edge then points at app.db.Base — a real node — rather than at a box called Base that means nothing.

Unresolved names become nodes, never nothing. A third-party import or a dynamically-built base becomes a node of kind: "external" and keeps its edge. Dropping it would be the worse failure: the diagram would look complete and have quietly lost a dependency. Gathered up, those nodes are the project's real dependency surface.

C# and TypeScript

Handled by tree-sitter, which is why generics, nested types and a brace inside a string are non-events rather than special cases. The test suite asserts that last one specifically, because it is exactly where a hand-rolled scanner breaks.

This reader produces no edges. Resolving a C# using to the thing it names is a different and much larger job, and the consumer that needs this — the minimap — needs none of it. An extractor that quietly produced half a dependency graph would be worse than one producing none, because the half would look whole.

Usage, not just the spec

An OpenAPI document says what endpoints are. It does not say how to use them — least of all when they are not RESTful, or when a GraphQL endpoint sits alongside. So docgen also reads a HAR: the recording format that browser devtools, mitmproxy, Charles and Insomnia all export.

what traffic knowswhat a spec cannot
the order of callsa spec is a set; usage is a sequence
which parameters are always senta spec lists twenty optional ones
which statuses really happenthe 422 everybody hits is in no document
endpoints not in the documentGraphQL operations, found by body shape and named
which id formats a route takesnumeric and uuid on one route
No credential and no payload value reaches the IR — only field names and types. A HAR is full of live bearer tokens and cookies, and a generated document gets committed. The test suite plants a token in its fixture and fails if it appears anywhere in the output.

Two limits, stated rather than glossed: path templating is a guess (attrs.observed_paths keeps what was actually seen beside it), and consecutive is not caused-by — the edge weight is what separates a habit from an accident, and one recording will not tell you which.

Views

The first real diagram out of this pipeline was a 3000px-wide strip: four modules of actual content and sixty sys/json/typing boxes as their peers. The emitter was correct and the picture was useless. That is a missing view, not a broken renderer — and the fix belongs to every consumer at once, because the index, the diagram and the diff all want the same narrowing.

python3 -m docgen.ops ir.json --overview -o view.json
python3 -m docgen.ops ir.json --around docgen.ir --hops 2 -o view.json
python3 -m docgen.ops ir.json --split -o parts/
python3 -m docgen.ops ir.json --shape          # what will this look like?

All of them are IR→IR, all composable, and each produces a document that still validates. --overview is the default and dispatches on the source: a codebase reduces to its modules and its outside dependencies, a schema to its tables and their keys.

Edges are lifted when a view collapses detail, never dropped. A class in module A inheriting from a class in module B is a dependency of A on B. Collapsing docgen to its packages once kept 8 of 77 edges — those pictures were not simpler, they were wrong. Lifted edges carry a weight saying how many they stand for.

Depth is the tempting knob and the wrong one

A directory without an __init__.py is not a package, so its modules have no parent and sit at depth 0. soleprint has 173 such roots, and a depth-2 cut still held 566 functions and 142 classes. Selecting by kind does not care how the directories happen to be arranged.

Emitters

emitteroutputaudience
indexmarkdown, sidebar JSON anyone — no graph literacy required
dotDOT → Graphviz → SVGdependency structure
erdSVG, written directlya schema, as cards
minimapSVG, written directlywhat is where, at a glance
notebook.ipynba runnable walkthrough
sitea static docs sitereading
explorea two-pane navigatorfinding your way
autowhichever of the above fitsnot having to choose

The non-visual ones matter most for reach. A sorted, described list of what exists is readable by someone who will never open a diagram, and it also reports the dependency surface and any file that failed to parse. It is built second, not last — it is what proves the IR is not secretly diagram-shaped.

ERD — and where the layout came from

Not invented here. station/tools/graphgen/templates/index.html, the Supabase-style schema explorer already in this repo, had solved it:

const cols = Math.max(2, Math.ceil(Math.sqrt(sorted.length * 1.2)));

Columns from the square root of the table count. The aspect ratio is chosen rather than emergent, so the result stays near-square at 4 tables or 400. That is the one thing a rank-based engine cannot offer. Three more things it gets right: a table is a card with its columns; an edge leaves the column holding the key and lands on the target's primary key; and the geometry is computed rather than measured, so it renders identically on any machine.

an entity-relationship diagram
A schema from the sample room. Same emitter, same style file as every other diagram here.

Minimap

Sublime's minimap shrinks the characters. This draws the structure at full scale: one file is a column, one line is a fixed number of pixels, every construct a block sized by its span and coloured by what it is. No text inside a block — the shape is the message.

The claim is that the pattern comes from the colours alone, so nesting is drawn by inset rather than by hue. On soleprint you can see that modelgen is class-based, histgen is function-based and tester is mixed, without reading a line.

a structural minimap of docgen
docgen's own files. Blue class, amber interface, green function, dark for everything that is not a declaration — imports, constants, prose.

Explore

The minimap on its own shows shape and no meaning: a block says "a 30-line class", not which class or what it touches. So it is not the artifact. It is the selector.

make explore OUT=out     # then open out/explore/explore.html

Left — navigate

The whole thing at once. Scan by colour, click a block.

Right — explore

What that is, what it reaches, what reaches it, and the neighbourhood drawn small enough to read. Every neighbour is a link, so you walk outward from wherever you started.

This is what retires the 14:1 sheet. The whole graph is never drawn. The overview pane carries the overview, and only the neighbourhood of a selection is rendered — a handful of nodes, which lays out fine every time. Overview and detail stop competing for one picture.

The same split applies to a database: every table at once with no column detail, then click one to get its columns plus the tables its keys reach. The two differ exactly where they should — a module's neighbourhood deliberately leaves its contents out, because those are the hundred functions that made the sheet unreadable, while a table's brings them in, because a table without its columns is not a table.

The selection basket

Shift-click accumulates blocks. The basket is a copyable list of paths with a line count — enough to hand to distill, and enough to see that the selection got too big before spending the context on it. Navigating a tree quickly in order to decide what to feed a model is a real use, and this is the part that serves it.

Notebooks

A notebook is normally a source file somebody confects by hand: prose, code and stored output braided together, diffing badly, drifting from whatever it documents the moment either moves, with no way to tell by looking.

Here a notebook is a build artifact. The source is the OpenAPI document — the same file the server is built from — and the notebook is regenerated from it. Nobody edits the .ipynb, the same way nobody edits a .o. "Is this document current" stops being a question about somebody's diligence and becomes a question about whether the build ran.

This is the disagreement with jupytext. Jupytext fixes the diffing — it makes a notebook editable as text — and leaves the actual problem: you still hand-author it, so it still rots.

Generated base, hand-written overlay

Generation alone gives a document that is never stale and never says anything a parser could not work out. Hand-authoring alone gives insight and a document that rots. Two files is the only arrangement that gets both:

IR ──► spec ──(+ overlay)──► merged spec ──► .ipynb
   generated   hand-written       merged       emitted

The spec is an ordered list of steps with no Jupyter in it — a Swagger for notebooks, readable and diffable. The overlay is the only file anyone edits, and it is re-applied on every build. It can annotate, replace, insert, drop and order.

replace is the one that matters. It is how real usage gets into a document that a spec could not describe — the call that is always made with status=available, the GraphQL endpoint that is not in the OpenAPI file at all — and it keeps working unchanged once a usage recording supplies the same facts automatically.

Three properties hold it together: regenerating re-applies the overlay byte-for-byte; when the base moves underneath it the mismatch is reported, never silently dropped; and extraction works with the overlay absent — it is an addition, never a dependency.
python3 -m docgen.emitters notebook ir.json --scaffold overlay.json
python3 -m docgen.emitters notebook ir.json --overlay overlay.json -o walkthrough.ipynb

Style & colour

A style rule names a slot, never a colour. "border": "atlas" is the rule; a theme binds atlas to #43A047 in print and #15803d on the docs site.

That indirection is the whole point. common/theme/tokens.css, docs/graphs/themes/*.gvpr and style/lucid.json use the same slot names, so a diagram and the page around it match by construction — which is the rule docs/graphs/README.md already states. The dark theme's artery, atlas and station slots are exactly the --system-accent values the three system pages set, and the test suite fails if they drift apart.

Dark is the default, because a generated diagram lands in a dark docs page far more often than in a document. --theme lucid gives the print palette — and gives it to the page as well as the diagram, since both are baked from the same slots.

An unknown kind falls back to default rather than crashing. That matters more than it sounds: a new extractor with a new vocabulary renders plainly and legibly on day one, instead of requiring somebody to write a style file before they can see anything.

Where DOT stops

The emitter writes what DOT expresses natively and stops at the boundary rather than growing machinery. The limits are recorded in the style file itself: a cluster has a label and a fill but not a header bar; stroke-dasharray is not parameterised, so 4,4 and 5,5 collapse; rounded is binary, so 4px and 6px are identical. Those mark where a richer emitter would begin — and the style file carries the full specification regardless, so that emitter needs no re-authoring.

One limit was worth solving: DOT cannot use a cluster as an edge endpoint, so every module-to-module import silently vanished. The native answer is compound=true with lhead/ltail — draw between a representative leaf and clip the line at the cluster border.

Standalone

Copy the docgen/ folder anywhere and it works. The Makefile derives its own package name from where it sits, so it can be renamed too, and everything else resolves inside the directory.

cp -r docgen /somewhere/else
cd /somewhere/else/docgen
make doctor          # what this machine has
make check           # the suite, from the copy
make book SRC=/path/to/any/repo BOOK=out/book/theirs

One seam, and it is optional

Exactly one capability needs more than the folder: reading an OpenAPI document goes through station/tools/modelgen, which parses the spec and resolves $ref. It is deliberately not reimplemented here — a second OpenAPI reader in one repo is two things to keep correct.

So if you are using docgen standalone but keeping the repo alongside as reference, point at it:

export DOCGEN_REFERENCE=/path/to/repo
make doctor
# reference: /path/to/repo (from $DOCGEN_REFERENCE)

Resolution is $DOCGEN_REFERENCE first, then walking up from the package — so in place it needs no configuration, and an explicit path wins when set. Without it, the four other extractors and every emitter work unchanged; the OpenAPI reader reports what to set, and the suite skips rather than fails.

An env var rather than a config file, because there is one setting and it is a path. A config file for one path is a file to find, parse, document and validate, and the first question anyone asks of it is "where does it live" — which is the same question again.

It is asserted, not asserted-in-prose

A standalone claim decays the moment somebody adds a convenient import, and it decays silently, because the suite still passes inside the repo. So the suite reads its own source:

The folder was also literally copied to /tmp and run, which is how the one real bug here was found: a check asserting the reference repo is reachable, correct in place and wrong the moment there was nothing above. It now reports which case applies instead of assuming one.

contextchecksskipped
in the repo250tree-sitter (259 with it)
copied out242tree-sitter, OpenAPI, the in-place case
copied out, DOCGEN_REFERENCE set249 tree-sitter, the in-place case

Commands

Make

targetdoes
make book SRC=…one whole operation, measured at both ends
make checkdocgen's own suite, offline, nothing installed
make check BOOK=…one book's own level — generated and custom
make doctorwhat this machine has and what it is missing
make ir SRC=…extract Python into OUT/ir.json
make code SRC=…extract C#/TypeScript tree-sitter
make db SCHEMA=…extract a database schema
make viewthe default view for that source type
make graphdraw whatever the structure asks for
make indexmarkdown index and sidebar JSON
make minimapwhat is where, read from the colours
make explorethe two-pane navigator
make sitea self-contained docs site
make selfdocgen's book of soleprint, then check it

Variables: SRC, OUT, SCHEMA, OPENAPI, HAR, BOOK, SLUG, READER, OVERLAY, STYLE, THEME, SCALE, DEPTH, PY. The Makefile derives its own package name from where it sits, so the folder can be copied anywhere and renamed and still work.

READER rather than LANG because LANG is the shell's locale variable, so ?= inherits en_US.UTF-8 from the environment and the argument is rejected. Every target above is one step of a book and still works alone — that is the property the spine exists to preserve, not to replace.

Modules

python3 -m docgen.book --root SRC -o out/book/slug   # the whole operation
python3 -m docgen.book.checks out/book/slug          # that book's level

python3 -m docgen.extractors.python --root SRC -o ir.json
python3 -m docgen.extractors code    --root SRC -o ir.json
python3 -m docgen.extractors db      --schema schema.json -o ir.json
python3 -m docgen.extractors openapi --spec spec.yaml -o ir.json
python3 -m docgen.extractors usage   --har session.har -o ir.json

python3 -m docgen.ir ir.json                      # validate
python3 -m docgen.ops ir.json --overview -o view.json

python3 -m docgen.emitters auto     view.json -o out/
python3 -m docgen.emitters index    ir.json -o index.md
python3 -m docgen.emitters dot      view.json -o graph.svg --theme lucid
python3 -m docgen.emitters erd      ir.json -o schema.svg
python3 -m docgen.emitters minimap  ir.json -o map.svg --scale 0.5
python3 -m docgen.emitters notebook ir.json -o book.ipynb --overlay overlay.json
python3 -m docgen.emitters site     view.json -o site/
python3 -m docgen.emitters explore  ir.json -o explore/

Dependencies

Everything below is optional. The stdlib covers the whole structural path — see standalone for the one seam out of the folder.

The core is standard library only. Everything else is optional and reported by make doctor; when something is missing you lose exactly one capability and get told what to install.

needsforwithout it
graphviz (binary)rendering DOT to SVG ERD, minimap, index and notebooks still work
tree_sitter + grammarsC#, TypeScript, TSX Python only
networkxlab/ experiments nothing — nothing depends on it yet
nodetesting the browser pages those checks skip
psqlthe lab/ schema probe use modelgen's from-db instead
lab/ is where a dependency gets tried before anything depends on it. Nothing in ir/, extractors/, ops/ or emitters/ may import from it. When an experiment earns its place it graduates into ops/ behind an IR→IR signature, and then the dependency is declared.

Testing

Three levels, and they differ by what they assert about. The distinction decides what a failure means, which is why it is worth keeping:

commandasks aboutfails?
make doctorthe machine — what is installed never; it reports
make checkdocgen — 259 checks exit 1
make check BOOK=<dir>that one book exit 1

All 259 with both optional dependencies installed, 244 with neither — the suite skips rather than fails when tree-sitter, lxml or the OpenAPI reader is absent. See standalone for the counts outside the repo.

The book level, where custom checks live

The third level is the one that reaches a project docgen has never seen, and it is where framework and hand-written checks live together. The generated half is the spine's own assertions, identical for every book: both measures present, the reconciliation holding, every artifact where the ledger says it is, the notebook still executing. The custom half goes in the book's own checks.py and uses the same helpers, so a project's line and a framework line read identically and fail identically.

# out/book/station/checks.py
def checks(book, check, note, skip):
    note("what this project will not give up")

    # Payments moved once already and the move broke three dashboards.
    # If it is not here, something renamed it again.
    check("payments is still a module", True,
          any(n["id"] == "app.payments" for n in book.ir["nodes"]))

Same split as the notebook's base and overlay, for the same reason: generation alone cannot know what this project cares about, and hand-authoring alone rots.

Each check is one decision that has already been made, with the reason above it. Not coverage, and deliberately not an exhaustive sweep. A rule without its reason gets overridden the first time it is inconvenient, so failing a check should read as "you are about to undo this" rather than "something broke". The idiom is carried from rig's ctrl/selftest.sh, which is where the three-level split comes from.

The four that are the design

Of docgen's own checks, four assert the architecture rather than guard a regression, and they are the ones to keep if anything is ever cut:

Golden tests go on the IR, never on the SVG. Graphviz measures label text with the host's fonts to size nodes, so identical input produces different geometry on a machine with different fontconfig. The IR is deterministic; the SVG is not. Pinning the wrong one gives a suite that fails on somebody else's laptop for no reason anyone can act on.

The browser pages are JavaScript, so they are tested as JavaScript: a stub DOM under node drives the viewer's zoom and 1:1 toggle, and the explorer's select-and-walk. Both skip cleanly where node is absent.

Self-hosting is the honest end-to-end check, and it is where the real bugs came from — two name-resolution faults and a duplicate-id crash that no fixture had reached. make self builds docgen's book of soleprint and then runs that book's own level against it, so the two ends have to reconcile on 225 real files. If the index does not read like the system, something is wrong.

Limits & non-goals

Things deliberately not done, with the reason, so they are not re-litigated:

Known gaps, stated plainly: