Files
soleprint/soleprint/atlas2/docgen/README.md

11 KiB

docgen

Static analysis of a tree, and the artifacts that fall out of it.

The point is not the diagram. The point is the format in the middle — diagrams are one consumer of it, and not the one that reaches the most people.

extractors/        →   graph IR (JSON)   →   emitters/
(per source type)      (one schema)          (per output target)
                              ↑
                         style/*.json
                    (consumed by emitters only)
make sync                       # optional: .venv with every group, from uv.lock
make check                      # prove it, on a tree it builds itself
make book SRC=/path/to/repo     # one whole operation, measured at both ends
make run CONFIG=docgen.toml     # every book a run file lists
make help

The full documentation is docs/index.html — open it in a browser.

Layout

cli/            the command line — every command, and nothing else
book/           an operation: larder measure, steps, web output, run files
extractors/     source -> IR, one per source type
ir/             the contract: schema.json, the dataclasses, the validator
ops/            IR -> a smaller IR
emitters/       IR -> an artifact
notebook/       the notebook spec, before it is an .ipynb
style/          slots, themes, and harvesting a theme from real diagrams
fixtures/       docgen's own test inputs
lab/            sanctioned experiments; nothing imports it
reference.py    the one seam to the repo above, for reading OpenAPI

Library packages hold no command-line code — no argparse, no __main__.py, no cli_dot.py beside dot.py. It all lives in cli/, behind one entry point, and the selftest asserts it stays there.

pyproject.toml declares no required dependency: the structural path is the stdlib, Python 3.11+. The optional groups (code, openapi, harvest, lab) are pinned in uv.lock; [tool.uv] package = false, as dataconvert does, since this is a folder run in place rather than something to install.

Run files

# docgen.toml — beside the project it describes; paths relative to this file
[defaults]
out = "out/book"

[[book]]
name = "station"
root = "../soleprint/station"

[[book]]
name   = "shop"
schema = "schemas/shop.json"

make run CONFIG=docgen.toml [ONLY=station] [CHECK=1]. One failing book never stops the others; unknown keys are refused; a rebuild replaces the previous build's outputs and never touches a hand-written checks.py. docgen.example.toml runs against the shipped fixtures.

Or as three composable commands, which is what the Makefile is wrapping — run from the directory above this one:

python3 -m docgen extract python --root SRC -o ir.json
python3 -m docgen view ir.json --drop-stdlib -o view.json
python3 -m docgen emit dot view.json -o graph.svg --theme dark

DOT collapses three concerns; this separates them

concern question owner
structure what the graph is ir/schema.json — versioned, golden-tested
meaning what things mean visually style/*.json, keyed on kind
placement where things go Graphviz defaults. Phase two

An extractor has never heard of SVG, colours or layout. An emitter has never heard of Python, ast or SQL. The IR carries no visual information — if a field would change between light and dark theme, it does not belong in it. shape="cylinder" is not a field; it is kind="datastore" plus a style rule, which is what lets the same IR render in a theme that has no cylinders.

The selftest asserts all three of those, because they are the design rather than a nicety and they are exactly what erodes first.

The IR

{
  "meta":  { "source": "python", "root": "app/", "schema_version": "1" },
  "nodes": [ { "id": "app.models.User", "kind": "class", "label": "User",
               "parent": "app.models",
               "attrs": { "file": "app/models.py", "line": 12 } } ],
  "edges": [ { "source": "app.models.User", "target": "app.db.Base",
               "kind": "inherits", "attrs": {} } ]
}
  • id is fully qualified and stable across runs. That is what makes two graphs from two commits diffable.
  • kind is the hinge, and the only field style and layout may key on.
  • parent is containment. Relationships are edges.
  • attrs is an open bag; file/line let a UI link a box to a line.

Stdlib dataclasses, not Pydantic. A format that needs a library installed to be opened is not a format, it is an API. ir/validate.py is the check at the boundary, and it reads the field lists out of schema.json so the two cannot drift.

python3 -m docgen validate ir.json

It catches what a schema cannot: an edge naming a node that does not exist, a containment cycle, duplicate ids, and a visual field smuggled into attrs.

Extraction is deterministic

No LLM in the structural path. A diagram from an AST cannot be out of date with the code; one from a model's reading of the code is wrong the moment the model has a bad day, which is the problem this exists to fix.

ast resolves nothing on its own — class User(Base) yields the literal string "Base". So there are two passes: one collects each module's definitions and imports, the other resolves names against those tables.

from .db import Base ;  class User(Base)
    →  app.models.User --inherits--> app.db.Base        not "Base"

Unresolved names become kind: "external" nodes and keep their edges. Dropping them is the worse failure: the diagram looks complete and has quietly lost a dependency. Gathered by the index emitter, they are the project's dependency surface.

An unparseable file is recorded as a node with an error attr, not a crash — one bad file must not cost you the other four hundred.

calls edges are deliberately not attempted. Resolving self.foo() needs type inference, and a call graph that is quietly 60% right is worse than none because it reads as authoritative.

A second source

extractors/db.py reads the published {models, relationships, source} contract that modelgen already emits and graphgen already consumes. Tables become nodes, columns become contained nodes, foreign keys become edges — with no new top-level field, which was the checkpoint on whether the schema was right.

Connecting to a live database is not here. modelgen from-db --url ... does that and writes the schema this reads; the two-step also keeps credentials out of this pipeline entirely.

Views are not an emitter concern

The first real diagram out of this pipeline was a 3000px strip: four modules of content and sixty sys/json/typing boxes, all peers. The emitter was correct and the picture was useless. That is a missing view, and the fix belongs to every consumer at once — the index, the diagram and the diff all want "just this subsystem, two hops out, without the stdlib".

python3 -m docgen view ir.json --drop-stdlib --around docgen.ir --hops 2 -o view.json

drop_stdlib, drop_external, only_kinds, drop_kinds, subtree, neighbourhood, collapse_to_depth. All IR→IR, all composable, each producing a document that still validates.

Graph algorithms are not here. Transitive reduction, cycle detection and dominators are networkx's, and reimplementing them is the classic way to acquire a quiet bug. lab/ is where that dependency gets tried against real IRs before anything depends on it — the aim being to learn which part of it is actually attractive, rather than adopting all of it on faith.

One colour language

A style rule names a slot, never a colour. "border": "atlas" is the rule; the theme binds atlas to #43A047 in print and #15803d on the docs site.

That indirection is the whole point. common/theme/tokens.css, docs/graphs/themes/*.gvpr and style/lucid.json use the same slot names, so a diagram and the page around it match by construction — which is the rule docs/graphs/README.md already states. The dark theme's artery, atlas and station slots are exactly the --system-accent values set in artery/index.html:30, atlas/index.html:25 and station/index.html:29, and the selftest fails if they drift apart.

An unknown kind falls back to default rather than crashing, so a new extractor renders plainly and legibly on day one instead of needing a style file written first.

How a container picks its colour without the IR naming one: it does not. The IR says which spr model a group belongs to (attrs.domain — semantic), and domain_slots maps that to a slot. Same mechanism as --system-accent. With no domain, the emitter assigns by sorted id, so two runs agree.

Use DOT until it hits its limits

The emitter writes what DOT expresses natively and stops at the boundary rather than growing machinery. The limits are recorded in style/lucid.json under limits and reachable as Style.limits():

header bars a cluster has a label and a fill, not a 100%-width header rectangle
stroke-dasharray not parameterised — 4,4 and 5,5 collapse to one dash
corner radius rounded is binary, so 4px and 6px are identical
icon above label needs an HTML-like label table
sequence badges xlabel carries the number; the circle does not exist

Those mark where a richer emitter would begin. The style file carries the full spec regardless, so that emitter needs no re-authoring.

One limit that was worth solving: DOT cannot use a cluster as an edge endpoint, so every module-to-module import silently vanished. The native answer is compound=true with lhead/ltail — draw between a representative leaf and clip at the cluster border.

The output is addressable

id and kind pass through to the SVG as the element's id and class, and attrs.file/attrs.line become an href. A front end can bind behaviour to a box and a box can link to the line it came from, without the emitter knowing about either.

Testing

make check                  # docgen's own suite: 287 after make sync, 272 with nothing
make check BOOK=out/book/x  # one book's own level — generated and hand-written checks
make doctor                 # the machine; never fails

Golden tests go on the IR, never on the SVG. Graphviz measures label text with the host's fonts to size nodes, so identical input gives different geometry on a machine with different fontconfig. The IR is deterministic; the SVG is not.

Self-hosting is the honest end-to-end check, and it is where the real bugs came from — two name-resolution faults that no fixture had reached:

make self      # docgen's book of the widest tree it can see, then its checks

Where this sits

docgen belongs to Atlas — documentation is whose concern it is. It is not a station tool and is not under station/tools/; it may depend on station tools, which is the permitted direction.

Atlas 2 is a successor, not a replacement. soleprint/atlas/ is untouched: it carries client information and an idea still worth extracting — deriving frontend and backend tests from one source, which is the same shape as this pointed the other way.

Not here

No layout system, no positioning, no ELK. No HTML-like labels, no SVG post-pass. No LLM in the structural path — annotation (summarising a module, naming a cluster) is a later layer, cached to its own file keyed by node id, merged into attrs at emit time, and extraction must work with it absent. No configuration knobs until two real consumers disagree.