Compare commits

...

39 Commits

Author SHA1 Message Date
6dbc83a449 Merge branch 'rig-work' 2026-09-17 00:14:37 -03:00
730ebaff2f remove profile dependency 2026-09-17 00:14:27 -03:00
dd17021402 Merge branch 'rig-work' 2026-09-17 00:07:37 -03:00
19feac6d57 remove profile dependency 2026-09-17 00:07:24 -03:00
809a13eebe distill updates 2026-09-16 23:14:27 -03:00
004b397b94 Merge branch 'rig-work' 2026-09-16 19:14:54 -03:00
29e30693ea rig updates 2026-09-16 19:08:46 -03:00
3f7f6c988d Merge branch 'rig-work' 2026-09-16 14:21:29 -03:00
3e864aa919 clean up rig 2026-09-16 14:21:21 -03:00
c3fe4422f2 makefile standalone 2026-09-16 13:52:04 -03:00
99b1988504 updated contract 2026-09-16 13:44:02 -03:00
5219bd5edb Merge branch 'ui' 2026-09-16 09:38:08 -03:00
f7910bf42b ui framework extraction updates 2026-09-16 09:38:04 -03:00
24aeadde83 dataconvert updates 2026-09-16 09:33:03 -03:00
5391f50755 Merge branch 'rig-work' 2026-09-16 09:31:36 -03:00
fed5d92034 rig updates 2026-09-16 09:31:13 -03:00
7b70b7edc6 adapter updates 2026-09-16 09:13:47 -03:00
74e246e67a dataconvert updates 2026-09-16 09:07:07 -03:00
e05f8f1fca dataconvert updates 2026-09-16 08:57:43 -03:00
b102ab8de7 Merge branch 'docgen-graphgen' 2026-09-16 08:20:35 -03:00
2f3e9c2634 dataconvert: spreadsheets to SQL seeds, layouts from config, row cap, SCHEMA.md
Converts CSV, xlsx/xls/ods, directories, ZIPs and globs into one INSERT
file per table. Producer-specific layouts (header/data rows found by a
marker cell) and sheet naming live in a gitignored dataconvert.json, with
dataconvert-example.json as the template. --max-rows samples each table
and SCHEMA.md records columns, types, row counts and full sizes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 08:23:38 -03:00
86d051da48 add project, change cli file names, use config 2026-09-16 08:17:29 -03:00
6ce24586bd Merge branch 'berth' 2026-09-16 07:06:06 -03:00
7242b09e3a distill update 2026-09-16 05:27:53 -03:00
fbf47980d9 Merge branch 'rig-work' 2026-09-16 04:53:04 -03:00
b9238040a6 rig dep update 2026-09-16 04:52:12 -03:00
aba696df79 feature matching for clean extraction 2026-09-14 16:00:27 -03:00
974679a432 berth updates 2026-09-14 07:28:50 -03:00
26f99265ca Merge branch 'rig-work' 2026-09-14 07:12:19 -03:00
a29e0708e8 new conventions 2026-09-14 06:13:22 -03:00
a9df70cde0 berth init 2026-09-14 03:57:00 -03:00
e0426ecb01 self tests 2026-09-13 21:57:26 -03:00
37c4d588ea Merge branch 'main' into docgen-graphgen 2026-09-13 21:41:29 -03:00
160ee31b8c docgen: drop the superseded first pass, port the style harvester
The IR supersedes both intermediate designs (requirements D6, D7):
station/tools/docgen and the graph model that briefly lived in graphgen.
Shipping them beside atlas2/docgen would mean two graph models, which is the
thing the architecture argues against.

Carried across rather than lost:
  - style/extract.py and tokens.py, the offline theme harvester (R31). Rewritten
    to emit a *theme* — a slot-to-hex binding — rather than a whole style file,
    since what harvesting recovers is which colour a slot should be, not what a
    kind should look like.
  - graphgen/README.md, rewritten for what graphgen actually is now: the
    schema explorer. It fixes the blank station-index entry at run.py:304.

Two bugs found while doing it:
  - Canvas and ink are the two lightness extremes, not the two most common
    values. In a Graphviz SVG every label carries a fill, so the ink outnumbers
    the canvas 87 to 43 and the old rule produced a theme whose text was
    invisible against its own background.
  - A name defined in both branches of an if/else produced a duplicate id, which
    failed validation on docgen's own source. Disambiguated by line.
2026-09-13 21:41:29 -03:00
358b98f826 site emitter 2026-09-12 07:08:00 -03:00
7cb892ccfe add cases for code, dbs, and outline notebook generation 2026-09-12 06:50:05 -03:00
542d704da4 docgen iter 2 2026-09-12 06:42:49 -03:00
49a9f8ee57 sanitized rig 2026-09-12 02:44:40 -03:00
966f8fc821 attemp to develop rig in spr without an actual use case 2026-08-26 07:27:12 -03:00
217 changed files with 36142 additions and 4418 deletions

4
.gitignore vendored
View File

@@ -40,7 +40,5 @@ cfg/dlt/
# not land here. They are versioned in their own repo.
#
# Anchored at the ROOT on purpose: a copy is a SIBLING of rig/, so a rule inside
# rig/.gitignore cannot see it. The negation must name the full path for the same
# reason — `*-rig/` is unanchored and matches at any depth, including rig/sample-rig.
# rig/.gitignore cannot see it.
*-rig/
!rig/sample-rig/

View File

@@ -129,7 +129,7 @@ Every script stays runnable on its own — the standalone rule holds:
```bash
python build.py --cfg amar # -> gen/amar/
cd gen/standalone && python run.py # bare-metal
./ctrl/kind-up.sh # still works directly
./ctrl/cluster.sh up # still runs directly; rig builds the cluster
cd gen/<room> && ./ctrl/start.sh # each room owns its lifecycle scripts
```

View File

@@ -13,7 +13,7 @@
# make component ARGS="publish soleprint-ui /tmp/out --dist"
# make deploy ARGS="--build"
#
# Every script stays runnable on its own (./ctrl/kind-up.sh still works, and each
# Every script stays runnable on its own (./ctrl/cluster.sh up still works, and each
# built room keeps its own gen/<room>/ctrl/*.sh) — the standalone rule holds, and
# this only saves typing.
#
@@ -42,7 +42,7 @@ $(eval $(ARGS):;@:)
endif
.DEFAULT_GOAL := help
.PHONY: help build start stop dist docs cluster deploy component
.PHONY: help build start stop dist theme docs cluster deploy component
help: ## list targets
@grep -hE '^[a-z]+:.*?##' $(MAKEFILE_LIST) | sed 's/:.*##/\t/' | expand -t16
@@ -61,6 +61,11 @@ stop: ## stop a running room [<room>]
dist: ## compile the plexus UIs to single files [<room>]
bash ctrl/dist.sh $(or $(ARGS),$(ROOM))
# ── theme ──────────────────────────────────────────────────────────────────
theme: ## ad-hoc pages [new|parts|bake|check|export|run FILE]
bash ctrl/theme.sh $(or $(ARGS),bake)
# ── docs ───────────────────────────────────────────────────────────────────
docs: ## documentation [serve [port]|graphs [theme]] (default serve)

21
berth/.gitignore vendored Normal file
View File

@@ -0,0 +1,21 @@
# Machine-local config and credentials. Never committed.
ctrl/.env
# Rendered output. Regenerate with `make services render <target>`.
# Generated config is an artifact, not source — the estate file is the source.
#
# The path is ctrl/render/out/, NOT render/out/. A pattern containing a slash is
# anchored to the directory holding this .gitignore, so `render/out/` would mean
# berth/render/out/ — which does not exist, and the real output would have been
# committed. Caught by `git check-ignore -v`, which is the only way to be sure.
ctrl/render/out/
# The "default" scratch bucket: always gitignored, never versioned.
def/
# Key material. NEVER committed.
#
# Anchored to ctrl/ for the same reason as render/out/ above: a pattern with a
# slash resolves against this file's own directory. `git check-ignore -v` is the
# only way to confirm it, and ctrl/vpn.sh refuses to write a key until it does.
ctrl/.secrets/

78
berth/Makefile Normal file
View File

@@ -0,0 +1,78 @@
# One target per ctrl/ script; the subcommand is an argument, not a second
# target: `make estate show`, not `make estate-show`. The logic lives in the
# scripts, never here.
#
# Config layers, weakest first: ctrl/versions.env < ctrl/env.d/<target>.env <
# ctrl/.env < the environment. So `make estate plan TARGET=gcp` beats all.
#
# Every target defaults to its READ-ONLY verb, and the verbs that change a live
# estate are not reachable by a bare word. Rationale: README.md.
ESTATE := $(or $(shell sed -n 's/^ESTATE=//p' ctrl/.env 2>/dev/null),$(shell ls estate/*.json 2>/dev/null | head -1 | xargs -r basename | sed 's/\.json$$//'))
TARGET := $(or $(shell sed -n 's/^TARGET=//p' ctrl/.env 2>/dev/null),aws)
.PHONY: help check selftest estate services vpn dns certs host ports registry docs
help: ## list targets
@grep -hE '^[a-z][a-z-]*:.*?##' $(MAKEFILE_LIST) | sed 's/:.*##/\t/' | expand -t16
# ── preflight ──────────────────────────────────────────────────────────────
check: ## is this estate coherent? reports, never fixes
bash ctrl/check.sh
selftest: ## does berth still do what it says? exits 1 if not
bash ctrl/selftest.sh
ports: ## port map [show|verify] (default show)
bash ctrl/ports.sh $(or $(ARGS),show)
# ── the estate ─────────────────────────────────────────────────────────────
estate: ## the estate [show|list|plan|apply|destroy] (default show)
bash ctrl/estate.sh $(or $(ARGS),show)
services: ## gateway routes [list|render <target>|deploy] (default list)
bash ctrl/services.sh $(or $(ARGS),list)
# ── the network ────────────────────────────────────────────────────────────
vpn: ## overlays [list|show <ov>|check|render|keygen] (default list)
bash ctrl/vpn.sh $(or $(ARGS),list)
# ── names and trust ────────────────────────────────────────────────────────
dns: ## DNS records [list|add|add-wildcard|remove] (default list)
bash ctrl/dns.sh $(or $(ARGS),list)
certs: ## TLS [status|verify|renew|push] (default status)
bash ctrl/certs.sh $(or $(ARGS),status)
# ── the box ────────────────────────────────────────────────────────────────
host: ## the remote box [status|ports|services] (default status)
bash ctrl/host.sh $(or $(ARGS),status)
registry: ## the image registry [status] (default status)
bash ctrl/registry.sh $(or $(ARGS),status)
# ── docs ───────────────────────────────────────────────────────────────────
docs: ## documentation [serve|graphs] (default serve)
bash ctrl/docs.sh $(or $(ARGS),serve)
# ── swallowing the argument words — MUST BE LAST IN THIS FILE ──────────────
#
# Words after the target are arguments, but make reads each as a goal, so each
# gets a no-op rule. This block must come AFTER the real targets: when an
# argument names one (`make host ports`, `make vpn check`), the last definition
# wins, and it has to be the no-op. With it first, make ran both scripts.
#
# Make's "overriding recipe" warning is the swallow working as intended.
ARGS := $(wordlist 2,$(words $(MAKECMDGOALS)),$(MAKECMDGOALS))
ifneq ($(ARGS),)
$(eval $(ARGS):;@:)
# .PHONY too: some of those words name real directories (ctrl, estate, render),
# and make treats an existing directory as already built.
.PHONY: $(ARGS)
endif

250
berth/README.md Normal file
View File

@@ -0,0 +1,250 @@
# berth
spr's deploy half. A rig is a mobile installation; a **berth** is the allocated, paid-for
place where it is moored and actually operates. Local rig → remote berth.
```bash
make check # is this estate coherent? reports, never fixes
make estate show # the description, resolved
make vpn check # the overlay: addresses, routing, bindings, key hygiene
make services render aws
make ports verify # does the local map still agree with rig?
```
`rm -rf berth/` is the uninstall.
---
## What berth is
**One description of an estate, with a swappable executor.** The description is the artifact;
the tool that runs it is a rendering.
| layer | what berth uses | note |
| --- | --- | --- |
| overlay | **WireGuard**, config generated per peer | GPL-2.0, in-kernel, and **no coordination server** |
| infra | **OpenTofu** — one executor, not a pair | plain `.tf`; `terraform` works identically |
| gateway | Caddy locally, nginx on the box | a projection with per-target rules, not a format conversion |
| pipeline | Woodpecker | Actions/GitLab reachable from the same description; not built |
The estate's facts — domain, hosts, ports, instance size, firewall rules, services — live in
**one file every rendering reads** (`estate/<name>.json`), rather than being restated in
one tool's language and again in another's. **Swappability is bought by the description, not
by maintaining two renderings** — a rendering you *can* produce, not one you *must* keep in
step. Two live renderings cost you every resource twice, forever, with nothing enforcing that
they agree.
**The seam belongs in a script, not in a tool.** A tool's schema is a ceiling you do not
control.
*(This once leaned on a specific precedent, since withdrawn — `✖ B2` in [STALE.md](STALE.md).)*
**OpenTofu, and only OpenTofu.** Terraform has been BUSL-licensed since 2023; OpenTofu is the
CLI-compatible MPL-2.0 fork (Linux Foundation). Write plain `.tf` that runs under both; the
scripts say `tofu`, and `terraform` works identically.
**Pipelines are the second executor axis, and are not built.** Woodpecker is the
self-hosted rendering; Actions and GitLab CI are the standards that must be reachable from
the same description. Named here so the infra seam is not designed in a way that forecloses
it.
---
## The overlay is berth's network layer
Two instances in different clouds cannot share a VPC. A WireGuard overlay gives them one flat
address space that berth owns and can reproduce on any provider — and, under a flaky
environment, a second layer beneath whatever the provider offers.
That inverts the usual cloud pattern. Instead of a VPC with security groups and private
subnets, each instance gets a public IP, opens **only** the WireGuard port, and carries
everything else inside the tunnel. **The security boundary moves out of the provider's VPC
and into a layer that is identical on AWS, on GCP, and on a laptop behind NAT.**
**Plain WireGuard, not Tailscale/Headscale/NetBird.** Tailscale's client is open but its
coordination plane is proprietary SaaS; Headscale and NetBird are open but add a control plane
to run. Plain WireGuard needs no server at all.
**The cost is real and accepted:** no NAT traversal, no relay, no peer discovery. A peer behind
NAT must dial one with a public endpoint. Fine here — the instances have public IPs and the dev
box roams — but two roaming peers cannot reach each other. That is why `PersistentKeepalive` is
a checked invariant rather than a detail.
**Keys never enter the description.** Private keys are generated on the peer that owns them
(`make vpn keygen`) into `ctrl/.secrets/` and injected only at render time; public keys live in
the estate, because a config cannot be built without them. `vpn.sh` **refuses to write** either
a key or a rendered config until `git check-ignore` confirms the path is ignored — this repo
has already been bitten once by a `.gitignore` pattern anchoring to the wrong directory.
This also decides something about the IaC layer: **OpenTofu must never generate a WireGuard
private key**, because Terraform-lineage state stores every resource attribute in plaintext.
---
## Two rules that are not style preferences
### 1. Every default is the read-only verb
```
make estate -> show make certs -> status
make dns -> list make host -> status
make estate apply / destroy -> print the plan, then refuse without --yes
```
rig's `make cluster` defaults to `up`, because every rig verb is safe — a kind cluster is
disposable. berth's are not: `tofu destroy` costs money and takes live DNS with it. **A
tool where every verb is safe must not grow verbs that are not.**
The failure this prevents is not hypothetical. `ppl/ctrl/certs.sh:42` is `CMD="${1:-all}"`,
so a bare `./ctrl/certs.sh` there issues a real Let's Encrypt cert, rsyncs it to the gateway,
and reloads nginx. berth's `certs` defaults to `status`.
### 2. berth and rig share a convention, not code
Neither imports the other. They match on **shape** — the `make <noun> <verb>` dispatch, the
four-source config layering, the key names — and consistency is verified by
**recomputation**: `make ports verify` recomputes rig's `20000 + (cksum(name) % 200) * 10`
to check the local map, rather than sourcing rig's `lib/config.sh`.
Copying three stable lines is the whole cost of not coupling them. A shared library would
put something outside `rig/` on rig's path, and rig's promise is that
`grep -rIn -iE 'soleprint|\bspr\b'` across it returns nothing.
**rig is also unaware that berth exists.** `ppl/local/Caddyfile` is berth's to generate; rig
must not reference `local.ar` — its handover scrub refuses the string.
---
## The gateway doctrine
Practised across this codebase for a long time and never written down, so: written down.
- **Caddy where routing is dynamic and config-driven** — the in-cluster gateway that
multiplexes by Host header, and the host-side `.local.ar` name→port map.
- **nginx where it is a static server or a plain long-running compose service on the box.**
- **Envoy in `mpr`** — a deliberate one-off, not a third pattern.
- **`ingress-nginx` only as a kind addon** — a different thing again from either gateway.
### Rendering is a projection, not a format conversion
The local Caddyfile and the box's nginx are not two spellings of the same content:
| | local (Caddy) | cloud (nginx) |
| --- | --- | --- |
| granularity | one file | one file per vhost |
| blocks per service | one | two (`:80` redirect + `:443` server) |
| TLS | none; every address needs an explicit `:80` | one shared wildcard cert |
| upstream | `localhost:<port>` | container name + `resolver 127.0.0.11` |
| name depth | free | constrained by the cert |
| ambiguity | most specific wins | exact, else `default_server` (= load order) |
The `:80` is not decoration: without it Caddy 2 defaults each site to `:443` with auto-HTTPS,
which on `*.local.ar` means cert provisioning that fails and breaks the listener. The
variable upstream is not decoration either: naming the upstream in a variable forces runtime
DNS resolution, so nginx **starts when the upstream container is absent** — which is what
lets one nginx front a dozen independent compose stacks.
Because the two disambiguate by **opposite** rules, a name set that is unambiguous locally
can be ambiguous on the box. `make check` asserts against the projection, not the source.
### Installing generated vhosts — an order that is not optional
1. Generated config lands in `conf.d/generated/`, **not** `conf.d/`. `ppl/ctrl/deploy.sh`
rsyncs with `--delete`; sharing a directory means one set gets erased.
2. `nginx.conf` needs a **third** include line — its `conf.d/*.conf` glob does not recurse,
which is why `conf.d/soleprint/*.conf` already needs its own.
3. That include changes **load order**, and load order decides which `:443` block catches
unmatched names. So `default.conf`'s commented-out `:443 default_server` must be restored
**first**. `make check` fails on it deliberately: it is a gate, not a warning.
---
## What the checks assert, and why
The scripts are deliberately thin on comment — the reasoning lives here. Every check below
corresponds to something that is wrong, or was wrong, in a real estate.
### `make check` — the estate
| assertion | the failure it catches |
| --- | --- |
| every service name is covered by an issued cert SAN | **a wildcard matches exactly one label.** `*.d.com` covers `a.d.com` but not `a.b.d.com`, which needs its own SAN. Surfaces otherwise as a browser TLS warning, far from its cause |
| a `:443 default_server` exists | DNS and the cert are wildcard but nginx matches `server_name` exactly, so without one the fallback for any unknown name is whichever vhost loads first — alphabetically, by accident |
| firewall rules and listeners agree | a rule allowing a port nothing listens on is **dead config**; a service no compose file declares is **undocumented state**. Neither is visible from one side alone, which is why the inventory has two halves |
| `HOST` is an ssh alias, never a hostname | there is no `Host <domain>` block, so a bare hostname falls through to the global defaults, ssh offers every key in the agent in turn, and `MaxAuthTries` (6) trips with *"Too many authentication failures"* before reaching the right one. The aliases set `IdentitiesOnly yes` |
### `make vpn check` — the overlay
| assertion | the failure it catches |
| --- | --- |
| peer addresses unique and inside the subnet | a duplicate is a silent misroute, never an error |
| AllowedIPs do not overlap | AllowedIPs is **cryptokey routing** — the route table and the ACL at once. Overlapping ranges resolve to the last match, so an overlap is both a misroute and an unintended grant |
| something carries `PersistentKeepalive` if anything roams | without it a NAT mapping expires and the tunnel works only while traffic flows outward — *"works sometimes"*, the hardest failure to read. Note it is **not** a property of the roaming peer's own entry: the roaming machine sets it on the entry for the peer it **dials** |
| the listen port is in the firewall | otherwise no peer can be dialed at all |
| no private key in the description | public keys are *also* 44-char base64, so the shape proves nothing. The real assertions are **no field named `priv*`** and **no key outside a `public_key` field** |
| overlay-reached services bind a reachable address | **a tunnel cannot reach loopback.** A service on `127.0.0.1` is unreachable over the overlay; one on `0.0.0.0` is reachable but also exposed to the whole LAN |
### Capturing the overlay
```bash
sudo wg show | make vpn capture --write
```
`wg show` has three forms and **only the first is safe**:
| form | safe | why |
| --- | --- | --- |
| `wg show` | **yes** | prints `private key: (hidden)` |
| `wg show <if> dump` | **no** | field 1 of the first line *is* the private key |
| `wg showconf <if>` | **no** | prints `PrivateKey=` outright |
`capture` refuses the latter two by shape. It matches peers by **allowed-ips address, not
public key** — the keys are exactly what is missing at that point — and **drops a roaming
peer's endpoint in the parser**, since that value is a home ISP address and a roaming peer
has no stable endpoint anyway.
## Layout
```
berth/
├── Makefile one target per ctrl/ script; the verb is an argument
├── STALE.md withdrawn assumptions, each with a check that runs
├── estate/<name>.json THE ARTIFACT — one description, many renderings
└── ctrl/
├── check.sh reports and instructs; never fixes
├── estate.sh show | list | plan | apply --yes | destroy --yes
├── services.sh list | render <aws|gcp|local> | deploy
├── ports.sh show | verify (the rig coincidence check)
├── dns.sh certs.sh host.sh registry.sh docs.sh
├── versions.env pinned toolchain (weakest layer)
├── env.d/<target>.env provider shape: aws | gcp
├── .env.example -> ctrl/.env, machine-local (gitignored)
├── lib/config.sh the four-layer load, from rig
├── lib/estate.sh reading and projecting the estate
└── render/*.tmpl nginx vhost shapes, substituted with sed
```
Config layers, weakest first: `versions.env``env.d/<target>.env``ctrl/.env` → the
caller's environment. So `make estate plan TARGET=gcp` beats everything.
**Identity is explicit — the inversion of rig.** rig derives its name from its folder so that
copies never collide. berth refuses to guess, because a deployment has exactly one production
and a wrong guess acts on the wrong estate. `ESTATE` names a file; the only convenience is
that a single `estate/*.json` is used without being asked for.
**`python3`, not `jq`.** rig ships a pinned static `jq` because its floor is "docker and
nothing else" on a machine it does not control. berth's floor is already higher, so
`python3` is a dependency it *has* rather than one it *adds* — the same reasoning by which
rig chose `sed` over `envsubst`.
---
## Status
**B0 (this) is the shape.** `estate/mcrn.json` is marked `UNVERIFIED`: it records what the
repos *claim*, because `ppl/infra/` was written and never applied — no `~/.pulumi`, no
`venv`, no stack state, files dated `mar 6`. B1's read-only inventory is what replaces those
claims with observations. Until then, `estate plan` and `estate apply` refuse: there is
nothing truthful to compare against yet.
`make check` currently fails on two real things — see `def/plans/36.0/berth.md`.

104
berth/STALE.md Normal file
View File

@@ -0,0 +1,104 @@
# berth — withdrawn assumptions
**Everything in this file is no longer true.**
It exists so the live docs stay short and so a withdrawn assumption cannot quietly return:
each entry carries a **check**, and `ctrl/selftest.sh` runs every one of them. A retraction
that is only prose is a retraction nobody re-reads.
Kept rather than deleted for the reason the docgen thread already wrote down —
*a requirement that disappears without explanation comes back.*
### For agents
- **Do not restate these** in a plan, a README or a comment. One line pointing here is enough.
- **Ids are stable.** Cite `✖ B3`; do not re-explain it.
- **When you withdraw an assumption, move it here** — quoted claim, where it came from, what
superseded it and why, what changed in the code, and a check that proves it is gone.
- **Only withdrawn things belong here.** A warning that is still actionable is a live rule,
however historical it sounds, and stays where it is.
---
**✖ B1 — "Pulumi is the source in spr, and Terraform must match the config."**
*(INDEX §5, carried from 35.3)* Withdrawn 2026-09-12. berth uses **OpenTofu and only
OpenTofu**. The rule assumed *open source* and *industry standard* pull apart — Terraform
being BUSL, Pulumi being the open alternative. OpenTofu is both: the standard language, plain
Terraform-compatible HCL, under MPL-2.0. Swappability was never bought by keeping two
renderings; it is bought by `estate/*.json` being the artifact — a rendering you *can*
produce, not one you *must* maintain.
**Gone from:** `ctrl/versions.env` (no `PULUMI_VERSION`), `ctrl/estate.sh` (`plan` runs one
executor), `ctrl/lib/config.sh` (`PULUMI_STACK``TOFU_WORKSPACE`), `ctrl/check.sh`
(toolchain list).
**Deliberately kept:** `README.md` and `estate/mcrn.json` both record that `ppl/infra/` was
written and never applied — *"no `~/.pulumi`, no venv, no stack state"*. That is a historical
fact about the estate, not a live dependency.
**Check:** no `pulumi` in `ctrl/` or `Makefile`.
**✖ B2 — "ctlptl's rejection is the precedent for *the seam belongs in a script, not a tool*."**
*(INDEX §5, `berth/README.md`)* Withdrawn 2026-09-12. **ctlptl was reinstated** — pinned in
`rig/ctrl/versions.env` at v0.9.4 — and had been removed for the wrong reason. The rule may
still hold; it now has to stand on its own reasoning rather than that example.
**Gone from:** `README.md` — the argument is stated directly, with no borrowed evidence.
**Check:** `ctlptl` appears nowhere in berth.
**✖ B3 — "`wg show` is safe; `wg showconf` is not."** *(my own note, 2026-09-12)* Incomplete,
and the gap is the dangerous one. There are **three** forms, and `wg show <if> dump` puts the
**private key in field 1 of the first line**. Stated as a two-way distinction, the `dump` form
reads as safe.
**Now:** `wg show` plain is safe; `dump` and `showconf` are not. `ctrl/vpn.sh capture` refuses
the latter two **by shape**, rather than parsing around them.
**Check:** `vpn.sh` names all three forms, and `capture` rejects both unsafe ones.
**✖ B4 — "A roaming peer must set `PersistentKeepalive` on its own entry."**
*(`ctrl/vpn.sh`, first draft)* Wrong side, and wrong in the direction that looks fine:
`PersistentKeepalive` is set per-peer in a config, so the roaming machine sets it on the entry
for the peer it **dials**. The original check would have **warned on a correctly configured
overlay**.
**Now:** checked once per overlay — if anything roams, some peer entry must carry a keepalive.
**Check:** the invariant is not keyed on the roaming peer's own `keepalive` field.
**✖ B5 — "The Makefile's pass-through block goes near the top, with the other variables."**
*(`Makefile`, inherited from rig's layout)* Withdrawn 2026-09-12. When a subcommand **names a
real target**, make has two recipes for it and the **last definition wins** — so with the block
first, `make host ports` ran `ctrl/host.sh ports` *and* `ctrl/ports.sh ports`, the second
failing because `ports` is not one of its verbs. Same for `make host services`, `make vpn
check`, `make vpn show estate`.
**Now:** the `$(eval $(ARGS):;@:)` block is **last in the file**, so the no-op wins and the
word is swallowed — which is what an argument is. Make's *"overriding recipe"* warning is the
swallow working.
**Check:** every colliding invocation dispatches to exactly one script.
**✖ B6 — "A base64 key in the description can be caught by its shape."**
*(`ctrl/vpn.sh check`, first draft)* WireGuard **public** keys are also 44-char base64 and
legitimately live in the estate, so shape alone proves nothing and would flag correct data.
**Now:** two assertions instead — **no field named `priv*`**, and **no base64 key outside a
`public_key` field**.
**Check:** the estate's public keys do not trip the secret check.
**✖ B7 — "`network.wireguard` is where the overlay is described."** *(`estate/mcrn.json`)*
Superseded 2026-09-12: WireGuard is berth's network layer, not one service's transport, so it
is a top-level `vpn` block with named overlays and peers. `ctrl/check.sh` and
`ctrl/registry.sh` were repointed.
**Gone from:** the estate schema — a `wireguard_moved` tombstone marks the old key.
**Check:** nothing reads `network.wireguard.*`.
**✖ B8 — "berth and rig are related through their first uses."** *(early framing)* Withdrawn:
**rig and berth are peers — neither depends on the other.** They match on shape (dispatch,
config layering, key names) and consistency is verified by **recomputation**, never by
dependency. A shared library would put something outside `rig/` on rig's path.
**Check:** berth imports nothing from rig; `ports.sh` recomputes the port formula and agrees
with rig's golden values.
**✖ B9 — "`langfuse.mcrn.ar` is an exception that cannot be generated."**
*(`estate/mcrn.json`, `raw: true`)* Withdrawn 2026-09-14. It was filed as the one route a
template could not express — a static `upstream{}` to a WireGuard address, with no `resolver`
and no `set $var`. It was not an exception; it was **the first instance of the general case**.
Those three properties are not three decisions, they are one: *this service is reached by
address on the overlay, not by name on the docker network.* Naming that decision —
`placement` — makes the file renderable.
**Gone from:** `estate/mcrn.json``lng` and `langfuse` were **two entries for one socket**
and are now one service with `placement: local`, `peer: nrft`, and a `local_host` for the
name it answers to locally. `raw` is dropped.
**Check:** the generated vhost matches the hand-written one, normalised for comments and
whitespace — proof against a live route rather than an assertion.

18
berth/ctrl/.env.example Normal file
View File

@@ -0,0 +1,18 @@
# Machine-local config. Copy to ctrl/.env (gitignored) and edit.
#
# The estate's FACTS live in estate/<name>.json.
# The provider's SHAPE lives in ctrl/env.d/<target>.env.
# This file is only what differs between machines, plus credentials.
# ESTATE is required when estate/ holds more than one file.
# ESTATE=mcrn
# TARGET=aws
# ssh ALIASES, never hostnames — check.sh refuses a value containing a dot.
# HOST=mcrn # app user, no sudo
# HOST_ADMIN=mcrn-admin # sudo, only where genuinely required
# Credentials: names only, never values. The secrets stay in ~/.aws and
# ~/.config/gcloud where their own tooling manages them.
# AWS_PROFILE=default
# GCP_PROJECT=

58
berth/ctrl/certs.sh Normal file
View File

@@ -0,0 +1,58 @@
#!/usr/bin/env bash
# The wildcard TLS cert for the gateway.
#
# Usage:
# ./certs.sh status # SANs issued vs SANs the services need
# ./certs.sh verify # inspect the cert served on :443
# ./certs.sh renew # refuses: issues a real cert
# ./certs.sh push # refuses: ships to a live gateway
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
status() {
local issued; issued=$(estate_get "certs.issued" | python3 -c 'import json,sys
try: print("\n".join(json.load(sys.stdin)))
except Exception: pass')
echo "issued SANs (estate/${ESTATE}.json: certs.issued):"
echo "$issued" | sed 's/^/ /'
echo
echo "SANs the services NEED (derived from services[]):"
estate_sans | sed 's/^/ /'
echo
local missing=0 s
while IFS= read -r s; do
[ -z "$s" ] && continue
grep -qxF "$s" <<< "$issued" || { echo "MISSING: $s"; missing=1; }
done < <(estate_sans)
[ "$missing" = 0 ] && echo "the issued cert covers every derived name."
echo
echo "certbot image: $(eval echo "\$$CERTBOT_IMAGE_VAR") provider: $DNS_PROVIDER"
}
verify() {
echo "would run:"
echo " echo | openssl s_client -connect ${DOMAIN}:443 -servername ${DOMAIN} 2>/dev/null \\"
echo " | openssl x509 -noout -dates -ext subjectAltName"
echo
echo "read-only against a live host — announce and approve first (§7)."
}
refuse() {
echo "REFUSING: '$1' acts on a live cert and a live gateway." >&2
echo " renew issues a real Let's Encrypt cert (rate-limited)." >&2
echo " push rsyncs to the gateway and reloads nginx." >&2
echo " Neither runs without explicit approval." >&2
exit 1
}
case "${1:-status}" in
status) status ;;
verify) verify ;;
renew|push) refuse "$1" ;;
*) echo "usage: $0 [status|verify|renew|push]" >&2; exit 1 ;;
esac

153
berth/ctrl/check.sh Normal file
View File

@@ -0,0 +1,153 @@
#!/usr/bin/env bash
# Is this estate coherent? Reports and instructs; never fixes.
#
# Takes no subcommand — there is one question to ask.
# Every check corresponds to something wrong in the estate today.
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
WORST=0
note() { echo " $*"; }
warn() { echo " WARN $*"; [ "$WORST" -lt 1 ] && WORST=1; return 0; }
bad() { echo " FAIL $*"; WORST=2; return 0; }
echo "estate: $ESTATE target: $TARGET domain: $DOMAIN"
echo
# 1. cert coverage: a wildcard matches exactly one label.
echo "certs — does the cert cover every name the services serve?"
issued=$(estate_get "certs.issued" | python3 -c 'import json,sys
try: print("\n".join(json.load(sys.stdin)))
except Exception: pass')
if [ -z "$issued" ]; then
warn "no certs.issued in the estate file — cannot check coverage."
else
while IFS=$'\x1f' read -r name host up kind raw placement peer port lhost; do
[ -z "$host" ] && continue
fqdn="${host}.${DOMAIN}"
# A '*' host stands for "any single label here" — check the deepest
# name it can produce, which is the one that fails.
probe="$fqdn"
case "$host" in \*.*) probe="anyroom.${host#\*.}.${DOMAIN}" ;; esac
covered=""
while IFS= read -r san; do
[ -z "$san" ] && continue
if san_covers "$probe" "$san"; then covered=1; break; fi
done <<< "$issued"
if [ -z "$covered" ]; then
bad "$name: '$probe' is covered by NO issued SAN"
note " issued: $(echo "$issued" | tr '\n' ' ')"
note " a wildcard matches exactly ONE label — reissue with"
note " -d '*.${host#\*.}.${DOMAIN}' or move the name one level up"
fi
done < <(estate_services "$TARGET")
[ "$WORST" -lt 2 ] && note "every service name is covered."
fi
echo
# 2. unmatched names: DNS and the cert are wildcard, nginx is exact, so
# without a :443 default_server the fallback is whichever vhost loads first.
echo "gateway — is there a deliberate answer for unmatched names?"
DEFAULT_CONF="${PPL_DIR:-$HOME/wdir/semester/ppl}/gateway/nginx/conf.d/default.conf"
if [ ! -f "$DEFAULT_CONF" ]; then
note "ppl not on this machine at $DEFAULT_CONF — skipped."
elif grep -qE '^\s*listen\s+443.*default_server' "$DEFAULT_CONF"; then
note "default.conf has a :443 default_server."
else
bad "default.conf has NO :443 default_server."
note " Unmatched names fall through to the first-loaded vhost."
note " This is a PREREQUISITE for generating any config: adding a"
note " generated include changes load order, and load order is what"
note " currently decides the fallback."
fi
echo
# 3. a rule allowing a port nothing listens on is dead config; a service no
# compose file declares is undocumented state. Needs both halves to see.
echo "firewall — rules against listeners"
estate_get "firewall" | python3 -c '
import json,sys
try: fw = json.load(sys.stdin)
except Exception: fw = []
for r in fw:
n = r.get("note")
print(" %-6s %-5s %s" % (r["port"], r.get("proto","tcp"), r.get("desc","")))
if n: print(" UNRESOLVED: " + n)
'
note "listener side: unknown until captured (ss -ltnp over ssh $HOST)."
# Ask the structure, not the prose: the condition is "does a peer still lack a
# public key", not "is there a _status string". _status is ALWAYS non-empty —
# capture rewrites it to "CAPTURED ..." — so testing it for emptiness pinned
# this warning on permanently, including after the capture it asks for.
uncaptured=$(estate_get "vpn.overlays.estate.peers" 2>/dev/null | python3 -c '
import json, sys
try:
peers = json.load(sys.stdin)
except Exception:
sys.exit(0)
print(" ".join(n for n, p in peers.items() if not p.get("public_key")))
' 2>/dev/null)
if [ -n "$uncaptured" ]; then
warn "overlay: public keys not captured for:$uncaptured — see 'make vpn check'"
note " 10.8.0.1 carries the registry and woodpecker gRPC;"
note " 10.8.0.2 backs langfuse. Nothing in the tree creates the"
note " interface — a fresh box cannot start the gateway compose."
note " capture with: sudo wg show | make vpn capture --write"
else
note "overlay: $(estate_get 'vpn._status')"
fi
note "overlay detail: make vpn show estate"
echo
# 4. ssh aliases, never hostnames: a bare hostname offers every agent key and
# trips MaxAuthTries before reaching the right one.
echo "ssh — aliases, never hostnames"
for var in HOST HOST_ADMIN; do
val="${!var:-}"
if [ -z "$val" ]; then
warn "$var is unset."
elif [[ "$val" == *.* ]]; then
bad "$var='$val' looks like a hostname, not a ~/.ssh/config alias."
note " A bare hostname trips MaxAuthTries before reaching the key."
elif [ -f "$HOME/.ssh/config" ] && grep -qiE "^\s*Host\s+.*\b${val}\b" "$HOME/.ssh/config"; then
note "$var=$val — Host block present."
else
warn "$var='$val' has no matching Host block in ~/.ssh/config."
fi
done
for f in "$HOME/wdir/semester/ppl/ctrl/.env"; do
[ -f "$f" ] || continue
if grep -qE '^SERVER=.*\.' "$f"; then
warn "$f sets SERVER to a hostname, not an alias — every ppl script inherits it."
fi
done
echo
# ── 5. toolchain ───────────────────────────────────────────────────────────
echo "toolchain"
for t in python3 "$TOFU_BIN" aws gcloud ssh rsync wg; do
if command -v "$t" >/dev/null 2>&1; then
note "$(printf '%-8s' "$t") present"
else
note "$(printf '%-8s' "$t") MISSING — $(case $t in
tofu) echo 'blocks: estate plan/apply; terraform works identically' ;;
wg) echo 'blocks: vpn keygen and the overlay checks' ;;
aws) echo 'blocks: the control-plane inventory, dns on route53' ;;
gcloud) echo 'blocks: the gcp estate' ;;
*) echo 'blocks: most things' ;;
esac)"
fi
done
echo
case "$WORST" in
0) echo "OK" ;;
1) echo "OK, with warnings" ;;
2) echo "PROBLEMS FOUND — see FAIL lines above" ;;
esac
exit 0

81
berth/ctrl/dns.sh Normal file
View File

@@ -0,0 +1,81 @@
#!/usr/bin/env bash
# DNS records, over whichever provider the target names.
#
# Usage:
# ./dns.sh list
# ./dns.sh add <subdomain> # <sub>.<domain> -> <domain>
# ./dns.sh add-wildcard <sub>
# ./dns.sh remove <subdomain>
#
# Read-only verbs announce and wait; mutating ones refuse.
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
announce() {
echo "would run:"
printf ' %s\n' "$*"
}
list() {
case "$DNS_PROVIDER" in
route53)
announce "aws route53 list-resource-record-sets --hosted-zone-id $AWS_HOSTED_ZONE_ID" \
"--query 'ResourceRecordSets[].[Name,Type,TTL,ResourceRecords[0].Value]' --output table"
;;
google)
announce "gcloud dns record-sets list --zone=${ESTATE}-zone --project=${GCP_PROJECT:-<unset>}"
;;
*) echo "unknown DNS_PROVIDER: $DNS_PROVIDER" >&2; exit 1 ;;
esac
echo
echo "read-only, but NOT run: each batch is announced and"
echo "waited on. Approve it and it runs."
}
mutate() {
local verb="$1" sub="${2:-}"
[ -z "$sub" ] && { echo "usage: $0 $verb <subdomain>" >&2; exit 1; }
local name
case "$verb" in
add) name="${sub}.${DOMAIN}" ;;
add-wildcard) name="*.${sub}.${DOMAIN}" ;;
remove) name="${sub}.${DOMAIN}" ;;
esac
echo "$verb: $name -> $DOMAIN (provider: $DNS_PROVIDER)"
echo
# The wildcard-depth rule again, applied BEFORE the record is created rather
# than discovered in a browser afterwards.
if [ "$verb" = "add" ]; then
local issued; issued=$(estate_get "certs.issued" | python3 -c 'import json,sys
try: print("\n".join(json.load(sys.stdin)))
except Exception: pass')
local covered=""
while IFS= read -r san; do
[ -z "$san" ] && continue
san_covers "$name" "$san" && { covered=1; break; }
done <<< "$issued"
[ -z "$covered" ] && {
echo "WARNING: '$name' is covered by no issued SAN." >&2
echo " A wildcard matches ONE label. The record would" >&2
echo " resolve and then fail TLS. Reissue the cert first." >&2
echo >&2
}
fi
echo "REFUSING: this changes live DNS." >&2
echo " Nothing is created, modified or deleted on any account without" >&2
echo " explicit approval for that specific action." >&2
exit 1
}
case "${1:-list}" in
list) list ;;
add|add-wildcard|remove) mutate "$@" ;;
*) echo "usage: $0 [list|add <sub>|add-wildcard <sub>|remove <sub>]" >&2; exit 1 ;;
esac

27
berth/ctrl/docs.sh Normal file
View File

@@ -0,0 +1,27 @@
#!/usr/bin/env bash
# Documentation.
#
# Usage:
# ./docs.sh serve|graphs
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
case "${1:-serve}" in
serve)
echo "berth has no doc server yet — the README is the documentation."
echo " $(cd .. && pwd)/README.md"
echo
echo "the estate, resolved: make estate show"
;;
graphs)
echo "not implemented. The estate file is the graph's source; a"
echo "renderer belongs with docgen/graphgen, which is another thread's"
echo "another thread's — so this is a handoff, not a stub to fill in here."
;;
*) echo "usage: $0 [serve|graphs]" >&2; exit 1 ;;
esac

14
berth/ctrl/env.d/aws.env Normal file
View File

@@ -0,0 +1,14 @@
# Target: AWS — the estate that actually runs today (mcrn.ar).
TARGET_NAME=aws
CLOUD=aws
REGION=us-east-1
INSTANCE_TYPE=t3.small
# The Route53 hosted zone. It ALREADY EXISTS and is reused, never created —
# see estate.sh's refusal to create a zone on this target.
AWS_HOSTED_ZONE_ID=Z02279903503ZMIB5FC1N
SSH_KEY_NAME=mcrn
# certbot's DNS-01 plugin for this provider.
CERTBOT_IMAGE_VAR=CERTBOT_AWS_IMAGE
DNS_PROVIDER=route53

17
berth/ctrl/env.d/gcp.env Normal file
View File

@@ -0,0 +1,17 @@
# Target: GCP — the replica, on its own domain.
#
# The domain differs from AWS's on purpose: this target CREATES a DNS zone
# where aws reuses one, so a shared domain would create a second authoritative
# zone and break live DNS. The domain lives in estate/nrft.json, not here.
TARGET_NAME=gcp
CLOUD=gcp
REGION=us-central1
INSTANCE_TYPE=e2-small
GCP_PROJECT=
GCP_ZONE=us-central1-a
# Delegation order: create the zone and let it answer first, then set the
# nameservers at the registrar — it validates that they respond.
CERTBOT_IMAGE_VAR=CERTBOT_GCP_IMAGE
DNS_PROVIDER=google

101
berth/ctrl/estate.sh Normal file
View File

@@ -0,0 +1,101 @@
#!/usr/bin/env bash
# The estate: what it is, what would change, and — behind a gate — changing it.
#
# Usage:
# ./estate.sh # show
# ./estate.sh list # every estate/*.json
# ./estate.sh plan # tofu plan, read-only
# ./estate.sh apply --yes # refuses without --yes
# ./estate.sh destroy --yes
#
# Why the default is read-only: ../README.md
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
show() {
echo "estate: $ESTATE ($ESTATE_FILE)"
echo "target: $TARGET (cloud=$CLOUD region=$REGION)"
echo "domain: $DOMAIN"
echo "host: $HOST (sudo: $HOST_ADMIN)"
echo "workspace: $TOFU_WORKSPACE"
echo
local status; status=$(estate_get "_meta.status")
[ -n "$status" ] && echo " !! $status" && echo
echo "services in scope for '$TARGET':"
local name host up kind raw
while IFS=$'\x1f' read -r name host up kind raw placement peer port lhost; do
[ -z "$name" ] && continue
printf ' %-12s %-14s %-24s %s%s\n' \
"$name" "${host:--}" "${up:--}" "$kind" \
"$([ -n "$raw" ] && echo ' [hand-written]')"
done < <(estate_services "$TARGET")
echo
echo "cert SANs (derived, not listed):"
estate_sans | sed 's/^/ /'
}
list() {
local f n
printf '%-12s %-16s %s\n' ESTATE DOMAIN STATUS
for f in ../estate/*.json; do
[ -f "$f" ] || continue
n=$(basename "$f" .json)
printf '%-12s %-16s %s%s\n' "$n" \
"$(python3 -c 'import json,sys;print(json.load(open(sys.argv[1])).get("domain",""))' "$f")" \
"$(python3 -c 'import json,sys;print(json.load(open(sys.argv[1])).get("_meta",{}).get("status",""))' "$f")" \
"$([ "$n" = "$ESTATE" ] && echo ' <- this one')"
done
}
# Read-only. Meaningful only once state is imported: against empty state,
# plan reports "create N resources", which is not drift.
plan() {
echo "== $TOFU_BIN plan =="
if ! command -v "$TOFU_BIN" >/dev/null; then
echo " $TOFU_BIN not installed — skipped." >&2
echo " OpenTofu is the MPL-2.0 fork; 'terraform' works identically." >&2
else
echo " would run: $TOFU_BIN plan -var-file=<(estate)"
fi
echo
echo "NOTE: not wired up yet, and that is the point. tofu plan against empty"
echo " state reports \"create N resources\" — which is not drift, it is an"
echo " empty state. It becomes the check that proves the description"
echo " matches reality only once state is IMPORTED from the inventory."
echo " ppl/infra/ describes an aspiration: it was never applied."
}
# The gate. Two things have to be true: --yes present, AND the plan shown first.
require_yes() {
local verb="$1"; shift
local yes=""
for a in "$@"; do [ "$a" = "--yes" ] && yes=1; done
if [ -z "$yes" ]; then
echo "refusing to $verb without --yes." >&2
echo >&2
echo " $verb changes a live, billable estate and can take DNS with it." >&2
echo " Read the plan first: make estate plan" >&2
echo " Then: ./ctrl/estate.sh $verb --yes" >&2
exit 1
fi
echo "refusing to $verb: not implemented, and deliberately so." >&2
echo " The executor is not wired up, and nothing is imported yet, so" >&2
echo " there is nothing truthful to apply." >&2
exit 1
}
case "${1:-show}" in
show) show ;;
list) list ;;
plan) plan ;;
apply) shift; require_yes apply "$@" ;;
destroy) shift; require_yes destroy "$@" ;;
*) echo "usage: $0 [show|list|plan|apply --yes|destroy --yes]" >&2; exit 1 ;;
esac

63
berth/ctrl/host.sh Normal file
View File

@@ -0,0 +1,63 @@
#!/usr/bin/env bash
# The remote box — the other half of the inventory.
#
# Usage:
# ./host.sh status|ports|services
#
# Announces what it would run over `ssh $HOST` and does not run it. Always the
# ssh alias, never a hostname: there is no `Host mcrn.ar` block, so a bare
# hostname offers every agent key and trips MaxAuthTries.
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
guard_alias() {
if [[ "$HOST" == *.* ]]; then
echo "HOST='$HOST' is a hostname, not a ~/.ssh/config alias. Refusing." >&2
exit 1
fi
}
announce_batch() {
echo "would run, over 'ssh $HOST':"
printf ' %s\n' "$@"
echo
echo "read-only, and NOT run. Announce-first applies to EACH batch, not"
echo "once per session — so the commands can be read and"
echo "learned rather than scrolled past."
}
guard_alias
case "${1:-status}" in
status)
announce_batch \
"uname -a; uptime; df -h /" \
"docker ps --format '{{.Names}}\t{{.Image}}\t{{.Ports}}'" \
"docker network inspect gateway --format '{{range .Containers}}{{.Name}} {{end}}'" \
"systemctl list-units --type=service --state=running --no-pager" \
"systemctl list-timers --no-pager"
echo
echo "sudo-only, over 'ssh $HOST_ADMIN' and only where genuinely needed:"
echo " wg show # the WireGuard peers that exist nowhere in the tree"
;;
ports)
announce_batch "ss -ltnp"
echo "the estate declares these firewall rules:"
estate_get "firewall" | python3 -c '
import json,sys
for r in json.load(sys.stdin):
print(" %-6s %-5s %s" % (r["port"], r.get("proto","tcp"), r.get("desc","")))'
;;
services)
announce_batch "docker compose -f ~/ppl/gateway/docker-compose.yml ps"
echo "the estate declares $(estate_services "$TARGET" | grep -c . ) service(s) for target '$TARGET'."
echo "the gateway compose declares 8. The difference is sibling repos'"
echo "stacks joining the shared 'gateway' network — intended design, but"
echo "nothing in the tree lists it. The inventory produces that list."
;;
*) echo "usage: $0 [status|ports|services]" >&2; exit 1 ;;
esac

94
berth/ctrl/lib/config.sh Normal file
View File

@@ -0,0 +1,94 @@
# Shared config loading. Sourced, never executed. Run from ctrl/.
#
# Precedence, weakest first:
# ctrl/versions.env pinned toolchain (committed)
# ctrl/env.d/<target>.env provider shape: aws|gcp (committed)
# ctrl/.env machine-local + secrets (gitignored)
# the caller's env `make estate plan TARGET=gcp` (always wins)
CONFIG_OVERRIDABLE="TARGET ESTATE CLOUD REGION INSTANCE_TYPE
HOST HOST_ADMIN AWS_PROFILE AWS_HOSTED_ZONE_ID
GCP_PROJECT GCP_ZONE TOFU_WORKSPACE"
# rig's port formula, reproduced rather than imported. cksum because it is
# POSIX and gives the same value on every machine. Used to verify, not allocate.
derive_port_base() {
local h; h=$(printf '%s' "$1" | cksum | awk '{print $1}')
echo $((20000 + (h % 200) * 10))
}
_config_restore() {
local line
while IFS= read -r line; do
if [ -n "$line" ]; then
eval "export $line"
fi
done <<< "$1"
# A while loop returns its last body command's status; the trailing empty
# line would otherwise make this return 1 and trip `set -e` in the caller.
return 0
}
load_config() {
local k saved=""
for k in $CONFIG_OVERRIDABLE; do
# ${!k+x} distinguishes "set but empty" from "unset" — an explicit
# FOO= on the command line is a real choice and must survive.
if [ -n "${!k+x}" ]; then
saved+="$k=$(printf '%q' "${!k}")"$'\n'
fi
done
set -a
source ./versions.env
[ -f ./.env ] && source ./.env
set +a
# Re-apply overrides now so TARGET is the caller's before we pick the file.
_config_restore "$saved"
local target="${TARGET:-aws}"
if [ ! -f "./env.d/${target}.env" ]; then
echo "no such target: env.d/${target}.env" >&2
echo "available: $(ls env.d/*.env 2>/dev/null | xargs -n1 basename | sed 's/\.env$//' | tr '\n' ' ')" >&2
exit 1
fi
set -a
source "./env.d/${target}.env"
[ -f ./.env ] && source ./.env
set +a
_config_restore "$saved"
TARGET="$target"
# Identity is explicit: berth never guesses which estate it is acting on.
# The one convenience: a single estate/*.json is used without being asked.
if [ -z "${ESTATE:-}" ]; then
local n; n=$(ls ../estate/*.json 2>/dev/null | wc -l)
if [ "$n" = "1" ]; then
ESTATE=$(basename "$(ls ../estate/*.json)" .json)
else
echo "ESTATE is not set and estate/ holds $n candidates — refusing to guess." >&2
echo "available: $(ls ../estate/*.json 2>/dev/null | xargs -n1 basename | sed 's/\.json$//' | tr '\n' ' ')" >&2
echo "set it: make estate show ESTATE=<name>, or ESTATE= in ctrl/.env" >&2
exit 1
fi
fi
ESTATE_FILE="../estate/${ESTATE}.json"
if [ ! -f "$ESTATE_FILE" ]; then
echo "no such estate: estate/${ESTATE}.json" >&2
echo "available: $(ls ../estate/*.json 2>/dev/null | xargs -n1 basename | sed 's/\.json$//' | tr '\n' ' ')" >&2
exit 1
fi
# Facts come from the estate file, never restated in a target env.
DOMAIN=$(estate_get "domain")
HOST="${HOST:-$(estate_get "host")}"
HOST_ADMIN="${HOST_ADMIN:-$(estate_get "host_admin")}"
# Workspace == target, so the two can never mean different things.
TOFU_WORKSPACE="${TOFU_WORKSPACE:-$TARGET}"
}

147
berth/ctrl/lib/estate.sh Normal file
View File

@@ -0,0 +1,147 @@
# Reading and projecting estate/<name>.json. Sourced, never executed.
#
# python3 rather than jq: berth's floor already includes python3, so it is a
# dependency berth has rather than one it adds.
estate_get() {
python3 -c '
import json, sys
d = json.load(open(sys.argv[1]))
for k in sys.argv[2].split("."):
if isinstance(d, list):
try: k = int(k)
except ValueError: sys.exit(0)
try: d = d[k]
except Exception: sys.exit(0)
print("" if d is None else d if isinstance(d, str) else json.dumps(d))
' "$ESTATE_FILE" "$1"
}
# Services in scope for one target. A service names its targets; absent = all.
# Fields are US-separated (0x1f), not tab: tab is IFS whitespace, so bash
# collapses a run of them and an empty field would shift every later column.
estate_services() {
local target="${1:-$TARGET}"
python3 -c '
import json, sys
d = json.load(open(sys.argv[1]))
target = sys.argv[2]
for s in d.get("services", []):
tg = s.get("targets")
if tg is not None and target not in tg:
continue
print("\x1f".join([
s.get("name", ""),
s.get("host", ""),
str(s.get(target + "_upstream", s.get("upstream", "")) or ""),
s.get("kind", "proxy"),
"raw" if s.get("raw") else "",
s.get("placement", "box"),
s.get("peer", ""),
str(s.get("port", "") or ""),
s.get("local_host", s.get("host", "")),
]))
' "$ESTATE_FILE" "$target"
}
# The SAN list the services need, derived — never a literal list.
estate_sans() {
python3 -c '
import json, sys
d = json.load(open(sys.argv[1]))
domain = d["domain"]
sans = [domain]
depths = set()
for s in d.get("services", []):
h = s.get("host", "")
if not h:
continue
# A wildcard matches exactly ONE label. "git" needs *.domain; "dlt.spr"
# needs *.spr.domain. The parent of the leaf is what has to be covered.
parent = h.split(".", 1)[1] if "." in h else ""
depths.add(parent)
for p in sorted(depths):
sans.append("*." + (p + "." if p else "") + domain)
for s in sans:
print(s)
' "$ESTATE_FILE"
}
# Is <fqdn> covered by <san>? A wildcard matches exactly one label.
san_covers() {
local fqdn="$1" san="$2"
[ "$fqdn" = "$san" ] && return 0
case "$san" in
\*.*)
local suffix="${san#\*.}"
# Must end in .suffix AND have exactly one extra label.
case "$fqdn" in
*".$suffix") [ "${fqdn%".$suffix"}" = "${fqdn%%.*}" ] && return 0 ;;
esac
;;
esac
return 1
}
# ── the overlay ────────────────────────────────────────────────────────────
# Every overlay name, one per line.
overlay_names() {
python3 -c '
import json, sys
d = json.load(open(sys.argv[1]))
for n in d.get("vpn", {}).get("overlays", {}):
print(n)
' "$ESTATE_FILE"
}
# Peers of one overlay, US-separated:
# name, address, role, endpoint, public_key, allowed_ips, keepalive
overlay_peers() {
python3 -c '
import json, sys
d = json.load(open(sys.argv[1]))
ov = d.get("vpn", {}).get("overlays", {}).get(sys.argv[2], {})
for name, p in ov.get("peers", {}).items():
print("\x1f".join(str(x) if x is not None else "" for x in [
name, p.get("address"), p.get("role"), p.get("endpoint"),
p.get("public_key"), p.get("allowed_ips"), p.get("keepalive"),
]))
' "$ESTATE_FILE" "$1"
}
overlay_get() { estate_get "vpn.overlays.$1.$2"; }
# Is an address inside a CIDR? Pure python so there is no ipcalc dependency —
# berth's floor already includes python3 because the IaC side needs it.
addr_in_subnet() {
python3 -c '
import ipaddress, sys
try:
sys.exit(0 if ipaddress.ip_address(sys.argv[1]) in ipaddress.ip_network(sys.argv[2], strict=False) else 1)
except ValueError:
sys.exit(2)
' "$1" "$2"
}
# The upstream a service actually resolves to, as "host:port".
#
# A PLACED service has no literal `upstream` field: ✖ B9 replaced langfuse's
# hand-written `10.8.0.2:3000` with placement+peer+port, because being reached
# by address on the overlay is ONE decision, not three properties. Everything
# that asks "what does this service point at" must therefore resolve it the
# same way, or it silently sees an empty string and skips the service — which
# is exactly how vpn.sh's bindings invariant went quiet after B9 landed.
#
# usage: service_upstream <up> <placement> <peer> <port>
service_upstream() {
local up="$1" placement="$2" peer="$3" port="$4"
case "$placement" in
local|instance)
local addr; addr="$(overlay_get estate "peers.${peer}.address")"
[ -z "$addr" ] && return 1
printf '%s:%s' "$addr" "$port"
;;
*) printf '%s' "$up" ;;
esac
}

78
berth/ctrl/ports.sh Normal file
View File

@@ -0,0 +1,78 @@
#!/usr/bin/env bash
# The local port map, and whether it still agrees with rig.
#
# Usage:
# ./ports.sh show # DERIVED / ACTIVE / SOURCE
# ./ports.sh verify # recompute rig's formula, report drift
#
# berth recomputes rig's port formula rather than importing it, so neither
# depends on the other. See ../README.md.
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
# name<US>host<US>local_port for everything that has one.
_local_ports() {
python3 -c '
import json, sys
d = json.load(open(sys.argv[1]))
for s in d.get("services", []):
p = s.get("local_port")
if p:
print("\x1f".join([s.get("name",""), s.get("host",""), str(p)]))
' "$ESTATE_FILE"
}
show() {
local ld; ld=$(estate_get "local_domain"); : "${ld:=local.ar}"
printf '%-12s %-22s %-8s %-8s %s\n' NAME ADDRESS ACTIVE DERIVED SOURCE
local name host port base
while IFS=$'\x1f' read -r name host port; do
[ -z "$name" ] && continue
base=$(derive_port_base "$host")
if [ "$port" = "$base" ]; then
printf '%-12s %-22s %-8s %-8s %s\n' "$name" "${host}.${ld}" "$port" "$base" "derived"
else
printf '%-12s %-22s %-8s %-8s %s\n' "$name" "${host}.${ld}" "$port" "$base" "override"
fi
done < <(_local_ports)
echo
echo "DERIVED is what rig's formula gives for that name. ACTIVE is what the"
echo "estate records. 'override' is not an error — most of these were never"
echo "rigs. 'make ports verify' says which ones should have matched."
}
verify() {
local name host port base rc=0 checked=0
while IFS=$'\x1f' read -r name host port; do
[ -z "$name" ] && continue
# Only 20000-21999 is rig's to predict; anything else was never derived.
if [ "$port" -lt 20000 ] || [ "$port" -gt 21999 ]; then
continue
fi
checked=$((checked + 1))
base=$(derive_port_base "$host")
if [ "$port" != "$base" ]; then
echo "DRIFT $name (${host}): estate says $port, rig's formula gives $base"
echo " either the rig pinned HTTP_PORT in its ctrl/.env, or the"
echo " folder was renamed. The Caddy map is stale either way."
rc=1
else
echo "ok $name (${host}): $port"
fi
done < <(_local_ports)
echo
echo "checked $checked rig-shaped port(s) in 20000-21999."
[ "$rc" = 0 ] && echo "no drift." || echo "drift found — regenerate with 'make services render local'."
return 0
}
case "${1:-show}" in
show) show ;;
verify) verify ;;
*) echo "usage: $0 [show|verify]" >&2; exit 1 ;;
esac

33
berth/ctrl/registry.sh Normal file
View File

@@ -0,0 +1,33 @@
#!/usr/bin/env bash
# The image registry — remote, and reachable only over the overlay.
#
# Usage:
# ./registry.sh status
#
# One verb: berth reports on a registry running on someone else's box. Starting
# and stopping it is that box's business.
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
status() {
local wg_server; wg_server=$(overlay_get estate "peers.box.address")
echo "registry: registry.${DOMAIN} (public pull via /v2/)"
echo "push: ${wg_server}:5000 (WireGuard-only, not in the firewall)"
echo
echo "would run:"
echo " curl -s https://registry.${DOMAIN}/v2/_catalog"
echo
echo "NOTE: the push endpoint binds ${wg_server}, and $(estate_get 'vpn._status')."
echo " A freshly-provisioned box cannot start the gateway compose file"
echo " at all, because that bind fails."
}
case "${1:-status}" in
status) status ;;
*) echo "usage: $0 [status]" >&2; exit 1 ;;
esac

View File

@@ -0,0 +1,32 @@
# ${NAME} — GENERATED by berth from estate/${ESTATE}.json. Do not edit.
# Edit the estate file and re-run: make services render aws
server {
listen 80;
server_name ${FQDN};
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
server_name ${FQDN};
ssl_certificate /etc/nginx/certs/live/${DOMAIN}/fullchain.pem;
ssl_certificate_key /etc/nginx/certs/live/${DOMAIN}/privkey.pem;
# Docker's embedded DNS. Naming the upstream in a VARIABLE forces runtime
# resolution, so nginx STARTS even when the upstream container is absent.
# With a literal proxy_pass, one stopped container takes the whole gateway
# down at reload — which is what makes one nginx able to front a dozen
# independent compose stacks.
resolver 127.0.0.11 valid=30s;
location / {
set $upstream_${NAME} ${UPSTREAM_HOST};
proxy_pass http://$upstream_${NAME}:${UPSTREAM_PORT};
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}

View File

@@ -0,0 +1,23 @@
# ${NAME} — GENERATED by berth from estate/${ESTATE}.json. Do not edit.
# Edit the estate file and re-run: make services render aws
server {
listen 80;
server_name ${FQDN};
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
server_name ${FQDN};
ssl_certificate /etc/nginx/certs/live/${DOMAIN}/fullchain.pem;
ssl_certificate_key /etc/nginx/certs/live/${DOMAIN}/privkey.pem;
root /usr/share/nginx/html/${NAME};
index index.html;
location / {
try_files $uri $uri/ =404;
}
}

View File

@@ -0,0 +1,32 @@
# ${NAME} — GENERATED by berth from estate/${ESTATE}.json. Do not edit.
# Placement: ${PLACEMENT} (${PEER}) — reached over the overlay, not the docker network.
upstream ${NAME}_backend {
server ${UPSTREAM_HOST}:${UPSTREAM_PORT};
}
server {
listen 80;
server_name ${FQDN};
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
server_name ${FQDN};
ssl_certificate /etc/nginx/certs/live/${DOMAIN}/fullchain.pem;
ssl_certificate_key /etc/nginx/certs/live/${DOMAIN}/privkey.pem;
# No `resolver`, and no `set $var` indirection — deliberately. Those exist so
# nginx starts when a CONTAINER is absent; this upstream is a literal address
# on the overlay, which needs no DNS at all. The three properties are one
# decision, and placement is what decides them.
location / {
proxy_pass http://${NAME}_backend;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}

210
berth/ctrl/selftest.sh Normal file
View File

@@ -0,0 +1,210 @@
#!/usr/bin/env bash
# What berth has settled, and what it has withdrawn, written down as assertions.
#
# Two halves:
# - decisions that hold. Failing one means "you are about to undo this".
# - every entry in ../STALE.md. Failing one means a withdrawn assumption came
# back. That is the half that makes STALE.md an audit surface and not an
# archive — a retraction nobody re-reads is a retraction that decays.
#
# Scope: no cloud, no ssh, no sudo, no network. Cheap enough to actually run.
# `make check` reports on the world and never fails; this exits 1, like rig's.
#
# Usage: make selftest (or: bash ctrl/selftest.sh)
set -uo pipefail # NOT -e: one failing check must not abort the rest
cd "$(dirname "$0")"
source ./lib/config.sh
rc=0
passed=0
check() { # name, expected, actual
if [ "$2" = "$3" ]; then
printf ' ok %s\n' "$1"
passed=$((passed + 1))
else
printf ' FAIL %s\n expected: %s\n got: %s\n' "$1" "$2" "$3"
rc=1
fi
}
note() { printf '\n%s\n' "$1"; }
skip() { printf ' skip %s (%s)\n' "$1" "$2"; }
# An absence check must not match the files that RECORD the absence. STALE.md
# names every withdrawn thing by definition, and this file names them again to
# assert them — so both are excluded, or every check fails on itself. rig hits
# the same wall and assembles its pattern from fragments for the same reason.
NOSELF="--exclude=selftest.sh --exclude=STALE.md"
absent() { grep -rIl $NOSELF "$@" 2>/dev/null | wc -l; }
# A throwaway estate, for the checks that have to run berth rather than read it.
TMP_ESTATE=_selftest
cleanup() { rm -f "../estate/${TMP_ESTATE}.json"; }
trap cleanup EXIT
note "the withdrawn assumptions — ../STALE.md, one check each"
# B1 — Pulumi. The two surviving mentions are historical fact about ppl/infra
# and live in README.md and the estate, not in anything that runs.
check "B1 no pulumi in the code" "0" "$(absent -i pulumi . ../Makefile)"
# B2 — the ctlptl precedent. Withdrawn; the argument stands on its own now.
check "B2 the withdrawn precedent is cited nowhere" "0" "$(absent -i ctlptl ..)"
# B3 — `wg show <if> dump` leaks the private key in field 1. Stating only
# show-vs-showconf makes the dump form read as safe.
check "B3 all three wg forms are named" "yes" \
"$(grep -q 'dump' vpn.sh && grep -q 'showconf' vpn.sh && echo yes || echo no)"
check "B3 capture refuses showconf-shaped input" "1" \
"$(printf '[Interface]\nPrivateKey = x\n' | bash vpn.sh capture >/dev/null 2>&1; echo $?)"
check "B3 capture refuses dump-shaped input" "1" \
"$(printf 'priv\tpub\t51820\toff\n' | bash vpn.sh capture >/dev/null 2>&1; echo $?)"
# B4 — keepalive belongs to the peer that DIALS, not the one that roams. The
# first version warned on a correctly configured overlay, so the check is run
# against one: hub carries the keepalive, nrft roams.
python3 - <<'PY'
import json, collections
d = json.load(open("../estate/mcrn.json"), object_pairs_hook=collections.OrderedDict)
d["vpn"]["overlays"]["estate"]["peers"]["box"]["keepalive"] = 25
json.dump(d, open("../estate/_selftest.json", "w"), indent=2, ensure_ascii=False)
PY
check "B4 a correct overlay raises no keepalive warning" "0" \
"$(ESTATE=$TMP_ESTATE bash vpn.sh check 2>/dev/null | grep -ci 'no peer entry carries')"
# B6 — public keys are 44-char base64 too, so shape alone would flag correct
# data. Same fixture, with a real-shaped public key on a peer.
python3 - <<'PY'
import base64, collections, json, os
d = json.load(open("../estate/_selftest.json"), object_pairs_hook=collections.OrderedDict)
d["vpn"]["overlays"]["estate"]["peers"]["box"]["public_key"] = base64.b64encode(os.urandom(32)).decode()
json.dump(d, open("../estate/_selftest.json", "w"), indent=2, ensure_ascii=False)
PY
check "B6 a public key does not trip the secret check" "0" \
"$(ESTATE=$TMP_ESTATE bash vpn.sh check 2>/dev/null | grep -c 'FAIL.*key')"
cleanup # the fixture is done with; two estate files would make load_config
# refuse to guess below, which is right but reads as a config failure
# B5 — the pass-through block must be LAST, or a subcommand that names a real
# target runs that target too. Checked through make, not by reading the file.
note "B5 a subcommand that names a target dispatches once"
for combo in "host ports" "host services" "vpn check" "vpn show estate" "estate show"; do
check " make $combo" "1" \
"$(cd .. && make -n $combo 2>/dev/null | grep -c 'bash ctrl/')"
done
# B7 — the overlay moved out of network.wireguard into a top-level vpn block.
check "B7 nothing reads network.wireguard" "0" "$(absent 'network\.wireguard' .)"
# B8 — peers, not relatives. berth sources nothing from rig.
check "B8 berth sources nothing from rig" "0" "$(absent -E 'rig/ctrl|\.\./rig' .)"
# B9 — langfuse was filed as an exception a template could not express. It was
# the general case. The proof is a live route: render it and diff against the
# hand-written file, normalised for comments and whitespace.
LIVE=/home/mariano/wdir/semester/ppl/gateway/nginx/conf.d/langfuse.conf
if [ -f "$LIVE" ]; then
norm() { sed -e 's/#.*//' -e 's/[[:space:]]\+/ /g' -e 's/^ //' -e 's/ $//' -e '/^$/d' "$1"; }
bash services.sh render aws >/dev/null 2>&1
check "B9 the generated vhost reproduces the live one" "same" \
"$(diff -q <(norm ./render/out/aws/langfuse.conf) <(norm "$LIVE") >/dev/null 2>&1 \
&& echo same || echo different)"
else
skip "B9 generated vhost matches the live one" "ppl not on this machine"
fi
note "the safety contract — berth's verbs are not all safe"
check "estate defaults to show" "show" "$(cd .. && make -n estate 2>/dev/null | grep -oE 'estate\.sh [a-z]+' | awk '{print $2}')"
check "certs defaults to status" "status" "$(cd .. && make -n certs 2>/dev/null | grep -oE 'certs\.sh [a-z]+' | awk '{print $2}')"
check "dns defaults to list" "list" "$(cd .. && make -n dns 2>/dev/null | grep -oE 'dns\.sh [a-z]+' | awk '{print $2}')"
check "vpn defaults to list" "list" "$(cd .. && make -n vpn 2>/dev/null | grep -oE 'vpn\.sh [a-z]+' | awk '{print $2}')"
for verb in apply destroy; do
check "estate $verb refuses without --yes" "1" \
"$(bash estate.sh "$verb" >/dev/null 2>&1; echo $?)"
done
for verb in renew push; do
check "certs $verb refuses" "1" \
"$(bash certs.sh "$verb" >/dev/null 2>&1; echo $?)"
done
check "dns add refuses to change live DNS" "1" \
"$(bash dns.sh add selftest >/dev/null 2>&1; echo $?)"
check "vpn up refuses without --yes" "1" \
"$(bash vpn.sh up >/dev/null 2>&1; echo $?)"
note "config — the caller's env beats the files"
# Generated from CONFIG_OVERRIDABLE, so a new key enrols itself.
test_value() {
case "$1" in
TARGET) echo "gcp" ;;
ESTATE) echo "mcrn" ;;
*) echo "selftest-sentinel" ;;
esac
}
for key in $CONFIG_OVERRIDABLE; do
want="$(test_value "$key")"
got="$(export "$key=$want"; load_config >/dev/null 2>&1; echo "${!key}")"
check " caller's $key wins" "$want" "$got"
done
note "rig agreement — recomputed, never imported"
# rig pins these same constants in its own selftest. Both arrive at them from
# the same formula with no shared code, which is the coupling rule made testable.
check "derive_port_base rig" "20310" "$(derive_port_base rig)"
check "derive_port_base foo" "21690" "$(derive_port_base foo)"
check "derive_port_base my-proj" "21030" "$(derive_port_base my-proj)"
note "containment — berth writes nothing outside berth/"
check "no tracked change outside berth/" "0" \
"$(cd ../.. && git status --porcelain 2>/dev/null | grep -vc '^.. berth/')"
check "generated output is ignored" "yes" \
"$(cd .. && git check-ignore -q ctrl/render/out && echo yes || echo no)"
# A trailing-slash pattern matches directories only, so ask about a path
# inside it rather than the (not-yet-existing) directory itself.
check "key material is ignored" "yes" \
"$(cd .. && git check-ignore -q ctrl/.secrets/vpn/any.key && echo yes || echo no)"
note "capture and the checks that read it — three bugs found by running, 2026-09-14"
# 1. A placed service's upstream is DERIVED (✖ B9). Anything reading the raw
# `upstream` field sees "" and skips it — which is how vpn.sh's bindings
# invariant, the "my configurations broke" detector, went quiet the day
# placement landed while still printing OK. Vacuous passes are the failure
# mode this whole file exists to catch.
check "a placed service resolves to a real upstream" "10.8.0.2:3000" \
"$(bash -c 'source ./lib/config.sh; source ./lib/estate.sh; load_config >/dev/null;
service_upstream "" local nrft 3000')"
check "bindings actually inspects a service" "1" \
"$(bash ./vpn.sh check 2>/dev/null | grep -c 'no service currently has an overlay address' \
| awk '{print 1-$1}')"
# 2. A listen port belongs to a PEER. The roaming peer's is an ephemeral source
# port; writing it to the overlay renames the port the firewall rule is
# checked against — silently, since both are plausible integers.
check "a roaming peer's port is not the overlay's port" "51820" \
"$(python3 -c 'import json;print(json.load(open("../estate/mcrn.json"))["vpn"]["overlays"]["estate"]["listen_port"])')"
# 3. _status is always non-empty — capture rewrites it rather than clearing it —
# so a warning gated on "is it set" can never turn off, including after the
# capture it asks for. Gate on the structure instead.
check "the capture warning clears once keys are in" "0" \
"$(bash ./check.sh 2>/dev/null | grep -c 'public keys not captured')"
note "every STALE entry has a check here"
# Not "$0": line 15 cd's into this script's directory, so a relative $0 no
# longer resolves. After the cd the file is simply selftest.sh.
# Ids are counted wherever they appear — B5's sits in a note(), not a check name.
entries="$(grep -c '^\*\*✖ B' ../STALE.md)"
checked="$(grep -oE '\bB[1-9][0-9]?\b' selftest.sh | sort -u | wc -l)"
check "STALE.md entries are all covered" "$entries" "$checked"
printf '\n%d passed' "$passed"
[ "$rc" -ne 0 ] && printf ', SOME FAILED'
printf '\n'
exit "$rc"

183
berth/ctrl/services.sh Normal file
View File

@@ -0,0 +1,183 @@
#!/usr/bin/env bash
# Gateway routes, projected from the estate onto one target.
#
# Usage:
# ./services.sh # list
# ./services.sh render aws # -> render/out/aws/*.conf (nginx vhosts)
# ./services.sh render local # -> render/out/local/Caddyfile
# ./services.sh deploy # refuses; ppl/ctrl/deploy.sh ships config
#
# Each target is a projection with its own rules, not a format conversion.
# The nine axes they disagree on, and the install order: ../README.md
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
OUT_ROOT="./render/out"
list() {
printf '%-12s %-16s %-22s %-9s %s\n' NAME FQDN UPSTREAM PLACEMENT SOURCE
local name host up kind raw
while IFS=$'\x1f' read -r name host up kind raw placement peer port lhost; do
[ -z "$name" ] && continue
# A placed service has no literal `upstream` — it is derived from the
# peer's overlay address, so show what it actually resolves to.
local shown; shown="$(service_upstream "$up" "$placement" "$peer" "$port")" || shown=""
printf '%-12s %-16s %-22s %-9s %s\n' \
"$name" "${host}.${DOMAIN}" "${shown:--}" "$placement" \
"$([ -n "$raw" ] && echo 'hand-written' || echo 'generated')"
done < <(estate_services "$TARGET")
echo
echo "hand-written entries are NOT generated and NOT overwritten."
echo "run 'make estate show' to see why each one is an exception."
}
render_cloud() {
# Two statements: `local a="$1" b="$a"` expands all arguments before any
# assignment, so $a would still be unset.
local target="$1"
local out="$OUT_ROOT/$target"
rm -rf "$out"; mkdir -p "$out"
local name host up kind raw uhost uport n=0 skipped=0
while IFS=$'\x1f' read -r name host up kind raw placement peer port lhost; do
[ -z "$name" ] && continue
if [ -n "$raw" ]; then
skipped=$((skipped + 1))
continue
fi
# Placement picks the rendering. A container on the estate's own network
# is reached by NAME through docker's resolver; anything on the overlay
# is reached by ADDRESS and needs no DNS. That is one decision, not the
# three properties (upstream{}, no resolver, no set $var) it produces.
local tmpl
case "$placement" in
local|instance)
tmpl=./render/nginx-upstream.tmpl
uhost="$(overlay_get estate "peers.${peer}.address")"
uport="$port" # resolved via service_upstream's same rule
if [ -z "$uhost" ]; then
echo " ! $name: placement '$placement' names peer '$peer', which has no address" >&2
continue
fi
;;
hosted)
echo " ! $name: placement 'hosted' is declared but not rendered yet" >&2
continue
;;
*)
if [ "$kind" = "static" ]; then
tmpl=./render/nginx-static.tmpl
uhost=""; uport=""
else
tmpl=./render/nginx-proxy.tmpl
uhost="${up%%:*}"; uport="${up##*:}"
fi
;;
esac
sed -e "s|\${NAME}|${name}|g" \
-e "s|\${ESTATE}|${ESTATE}|g" \
-e "s|\${FQDN}|${host}.${DOMAIN}|g" \
-e "s|\${DOMAIN}|${DOMAIN}|g" \
-e "s|\${PLACEMENT}|${placement}|g" \
-e "s|\${PEER}|${peer}|g" \
-e "s|\${UPSTREAM_HOST}|${uhost}|g" \
-e "s|\${UPSTREAM_PORT}|${uport}|g" \
"$tmpl" > "$out/${name}.conf"
n=$((n + 1))
done < <(estate_services "$target")
echo "wrote $n vhost(s) to $out/ ($skipped hand-written, left alone)"
cat <<EONOTE
TO INSTALL THESE, THREE THINGS MUST HAPPEN IN THIS ORDER — and the order is the
whole reason this is not a one-liner:
1. These land in conf.d/generated/, NOT conf.d/. ppl/ctrl/deploy.sh rsyncs the
gateway with --delete; generated and hand-written config sharing one
directory means one of them gets erased.
2. nginx.conf needs a THIRD include line. Its conf.d/*.conf glob does not
recurse — which is exactly why conf.d/soleprint/*.conf already needs its
own line at nginx.conf:28-30.
3. That new include changes LOAD ORDER, and load order decides which :443
block catches unmatched names. So default.conf's commented-out
':443 default_server' must be restored FIRST. 'make check' fails on this
today, deliberately — it is a gate, not a warning.
EONOTE
}
render_local() {
local out="$OUT_ROOT/local"; mkdir -p "$out"
local ld; ld=$(estate_get "local_domain")
: "${ld:=local.ar}"
local name host up kind raw port drift=0
{
cat <<EOH
# GENERATED by berth from estate/${ESTATE}.json. Do not edit.
# Regenerate: make services render local
#
# Install: sudo ln -sf \$PWD/Caddyfile /etc/caddy/Caddyfile && sudo systemctl reload caddy
# All *.${ld} resolve to 127.0.0.1 via dnsmasq.
#
# Every site address carries an explicit :80. Without it Caddy 2 defaults to
# :443 with auto-HTTPS, which on *.${ld} means cert provisioning attempts that
# fail and break the listener. Plain HTTP only on this host.
#
# Caddy matches the MOST SPECIFIC site address, not the first — the opposite of
# nginx, which matches exactly and otherwise falls to default_server. A name set
# that is unambiguous here can be ambiguous on the box.
EOH
while IFS=$'\x1f' read -r name host up kind raw placement peer port lhost; do
[ -z "$name" ] && continue
port=$(python3 -c '
import json,sys
d=json.load(open(sys.argv[1]))
for s in d.get("services",[]):
if s.get("name")==sys.argv[2]:
print(s.get("local_port") or ""); break
' "$ESTATE_FILE" "$name")
[ -z "$port" ] && continue
echo
echo "${lhost}.${ld}:80, *.${lhost}.${ld}:80 {"
echo " reverse_proxy localhost:${port}"
echo "}"
done < <(estate_services local)
} > "$out/Caddyfile"
echo "wrote $out/Caddyfile"
echo
echo "rig is not consulted and does not know this exists — its handover"
echo "scrub refuses the string '${ld}'. Where a port belongs to a rig,"
echo "'make ports verify' RECOMPUTES rig's formula to check it rather than"
echo "importing rig's code. Convention, verified; not a dependency."
}
case "${1:-list}" in
list) list ;;
render)
shift
# `case "${1:-X}"` defaults the match but leaves $1 empty.
t="${1:-$TARGET}"
case "$t" in
local) render_local ;;
aws|gcp) render_cloud "$t" ;;
*) echo "usage: $0 render [aws|gcp|local]" >&2; exit 1 ;;
esac
;;
deploy)
echo "berth does not ship config; ppl/ctrl/deploy.sh does." >&2
echo " berth's half is the DESCRIPTION and the render. Shipping is" >&2
echo " rsync + compose against a live box, and it belongs where the" >&2
echo " credentials are: berth is the tool, ppl is the estate that" >&2
echo " holds the secrets." >&2
exit 1
;;
*) echo "usage: $0 [list|render [aws|gcp|local]|deploy]" >&2; exit 1 ;;
esac

11
berth/ctrl/versions.env Normal file
View File

@@ -0,0 +1,11 @@
# Pinned toolchain. Committed. The weakest config layer.
# The infra executor. One, not a pair — see README.md.
# This pin is a placeholder; set it from `tofu version` once installed.
TOFU_VERSION=1.9.0
TOFU_BIN=tofu
# certbot runs as a throwaway container so the DNS plugin's credentials never
# have to be installed on this machine.
CERTBOT_AWS_IMAGE=certbot/dns-route53:latest
CERTBOT_GCP_IMAGE=certbot/dns-google:latest

478
berth/ctrl/vpn.sh Normal file
View File

@@ -0,0 +1,478 @@
#!/usr/bin/env bash
# Overlays — WireGuard as berth's network layer.
#
# Usage:
# ./vpn.sh # list
# ./vpn.sh show <overlay> # topology
# ./vpn.sh check # invariants
# ./vpn.sh render <peer> # peer's wg0.conf -> render/out/vpn/
# ./vpn.sh keygen <peer> # keypair -> .secrets/; prints only the public key
# ./vpn.sh up|down --yes # refuses; the host operates its own tunnel
#
# sudo wg show | ./vpn.sh capture [--write]
#
# `wg show` is the only safe form: `wg show <if> dump` puts the private key in
# field 1, and `wg showconf` prints it outright. capture refuses both.
#
# Rationale, topology and key handling: ../README.md
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
source ./lib/estate.sh
load_config
SECRETS_DIR="./.secrets/vpn"
OUT_DIR="./render/out/vpn"
WORST=0
note() { echo " $*"; }
warn() { echo " WARN $*"; [ "$WORST" -lt 1 ] && WORST=1; return 0; }
bad() { echo " FAIL $*"; WORST=2; return 0; }
# Which addresses belong to THIS machine, so checks can distinguish what they
# can actually see from what needs capturing elsewhere.
my_overlay_addrs() { ip -4 -o addr show 2>/dev/null | awk '{split($4,a,"/"); print a[1]}'; }
is_me() { my_overlay_addrs | grep -qxF "$1"; }
list() {
local n sub port peers
for n in $(overlay_names); do
sub=$(overlay_get "$n" subnet)
port=$(overlay_get "$n" listen_port)
peers=$(overlay_peers "$n" | grep -c . || true)
printf '%-10s %-16s port %-7s %s peer(s)\n' "$n" "$sub" "$port" "$peers"
note "$(overlay_get "$n" purpose)"
done
local st; st=$(estate_get "vpn._status")
[ -n "$st" ] && { echo; echo " !! $st"; }
}
show() {
local ov="${1:-}"
[ -z "$ov" ] && { echo "usage: $0 show <overlay>" >&2; exit 1; }
overlay_names | grep -qxF "$ov" || {
echo "no such overlay: $ov" >&2
echo "available: $(overlay_names | tr '\n' ' ')" >&2; exit 1; }
echo "overlay: $ov subnet $(overlay_get "$ov" subnet) udp/$(overlay_get "$ov" listen_port)"
echo
printf '%-8s %-12s %-9s %-22s %s\n' PEER ADDRESS ROLE ENDPOINT PUBKEY
local name addr role ep pk aips ka
while IFS=$'\x1f' read -r name addr role ep pk aips ka; do
[ -z "$name" ] && continue
printf '%-8s %-12s %-9s %-22s %s%s\n' \
"$name" "$addr" "$role" "${ep:-}" "${pk:-}" \
"$(is_me "$addr" && echo ' <- this machine')"
done < <(overlay_peers "$ov")
}
check() {
local ov name addr role ep pk aips ka
for ov in $(overlay_names); do
local sub port
sub=$(overlay_get "$ov" subnet); port=$(overlay_get "$ov" listen_port)
echo "overlay '$ov' — $sub udp/$port"
# 1. addresses: unique, and inside the subnet. Two peers sharing an
# address is a silent misroute, never an error message.
local addrs; addrs=$(overlay_peers "$ov" | cut -d$'\x1f' -f2 | grep -v '^$' || true)
local dupes; dupes=$(echo "$addrs" | sort | uniq -d)
[ -n "$dupes" ] && bad "duplicate peer addresses: $(echo "$dupes" | tr '\n' ' ')"
while IFS= read -r a; do
[ -z "$a" ] && continue
addr_in_subnet "$a" "$sub" || bad "$a is outside $sub"
done <<< "$addrs"
while IFS=$'\x1f' read -r name addr role ep pk aips ka; do
[ -z "$name" ] && continue
# AllowedIPs is cryptokey routing — route table and ACL at once.
case "$aips" in
*0.0.0.0/0*) warn "$name: AllowedIPs includes 0.0.0.0/0 — full-tunnel. Deliberate?" ;;
esac
# A peer with no endpoint cannot be dialed; it must initiate.
if [ -z "$ep" ] && [ "$role" != "roaming" ]; then
warn "$name: role '$role' but no endpoint — nothing can dial it."
fi
# Keepalive is NOT on the roaming peer's own entry — it is set on
# the entry for the peer it dials. Checked per-overlay below.
[ -z "$pk" ] && note "$name: public_key not captured yet"
done < <(overlay_peers "$ov")
# If anything roams, some peer entry must carry a keepalive.
if overlay_peers "$ov" | cut -d$'\x1f' -f3 | grep -qx roaming; then
if ! overlay_peers "$ov" | cut -d$'\x1f' -f7 | grep -qE '^[0-9]+$'; then
warn "a peer roams but no peer entry carries PersistentKeepalive"
note " the roaming side sets it on the entry for the peer it dials;"
note " without it the NAT mapping expires and the tunnel works only"
note " while traffic flows outward — 'works sometimes'"
else
note "keepalive present on the dialed peer."
fi
fi
# 4. the listen port must be open wherever a peer is dialable.
local fwports; fwports=$(estate_get "firewall" | python3 -c '
import json,sys
try: print(" ".join(str(r.get("port")) for r in json.load(sys.stdin)))
except Exception: pass')
case " $fwports " in
*" $port "*) note "udp/$port present in the firewall description." ;;
*) bad "udp/$port is in no firewall rule — no peer could be dialed." ;;
esac
echo
done
# Public keys are also 44-char base64, so shape alone proves nothing. The
# assertions are: no field named private, no key outside a public_key field.
echo "secrets — the description must never carry a private key"
local leaked
leaked=$(python3 - estate/../../estate/*.json <<'PY' 2>/dev/null || true
import json, re, sys, glob
KEY = re.compile(r'^[A-Za-z0-9+/]{43}=$')
bad = []
for f in glob.glob("../estate/*.json"):
def walk(node, path):
if isinstance(node, dict):
for k, v in node.items():
if re.search(r'priv', k, re.I):
bad.append(f"{f}: field '{'.'.join(path+[k])}' is named private")
walk(v, path + [k])
elif isinstance(node, list):
for i, v in enumerate(node): walk(v, path + [str(i)])
elif isinstance(node, str) and KEY.match(node):
if not path or 'public' not in path[-1]:
bad.append(f"{f}: base64 key at '{'.'.join(path)}' is not a public_key field")
walk(json.load(open(f)), [])
print("\n".join(bad))
PY
)
if [ -n "$leaked" ]; then
echo "$leaked" | while IFS= read -r l; do [ -n "$l" ] && bad "$l"; done
else
note "clean — no private-named field, no stray key material."
fi
echo
# A service reached over the overlay must bind an address the tunnel can
# reach. Loopback cannot be reached through a tunnel.
echo "bindings — services reached over the overlay must bind a reachable address"
local checked=0
while IFS=$'\x1f' read -r name host up kind raw placement peer port lhost; do
# Resolve placement first: a placed service's upstream is derived, not
# literal, so reading `up` alone skips it and this invariant goes quiet.
up="$(service_upstream "$up" "$placement" "$peer" "$port")" || true
[ -z "$up" ] && continue
local uhost="${up%%:*}" uport="${up##*:}"
addr_in_subnet "$uhost" "$(overlay_get estate subnet)" 2>/dev/null || continue
checked=$((checked + 1))
if is_me "$uhost"; then
local binds; binds=$(ss -ltn 2>/dev/null | awk -v p=":$uport\$" '$4 ~ p {print $4}')
if [ -z "$binds" ]; then
bad "$name: nothing listens on :$uport here, but $uhost:$uport is its upstream"
elif echo "$binds" | grep -q '^127\.0\.0\.1:'; then
bad "$name: :$uport binds 127.0.0.1 — unreachable over the overlay"
note " the tunnel cannot reach loopback; bind 0.0.0.0 or $uhost"
else
note "$name: :$uport binds $(echo "$binds" | tr '\n' ' ')— reachable"
echo "$binds" | grep -q '^0\.0\.0\.0:' && \
note " (0.0.0.0 also exposes it to the LAN; $uhost alone would be tighter)"
fi
else
note "$name: upstream $uhost is another peer — needs capture there"
fi
done < <(estate_services "$TARGET")
[ "$checked" = 0 ] && note "no service currently has an overlay address as its upstream."
echo
case "$WORST" in
0) echo "OK" ;;
1) echo "OK, with warnings" ;;
2) echo "PROBLEMS FOUND — see FAIL lines above" ;;
esac
return 0
}
# A .gitignore pattern containing a slash anchors to its own directory, so the
# only way to know a path is ignored is to ask git.
assert_ignored() {
local path="$1"
if ! git check-ignore -q "$path" 2>/dev/null; then
echo "REFUSING: '$path' is not gitignored." >&2
echo " Writing key material there would stage it on the next 'git add'." >&2
echo " Verify with: git check-ignore -v $path" >&2
exit 1
fi
}
keygen() {
local peer="${1:-}"
[ -z "$peer" ] && { echo "usage: $0 keygen <peer>" >&2; exit 1; }
command -v wg >/dev/null || { echo "wg not installed." >&2; exit 1; }
mkdir -p "$SECRETS_DIR"
assert_ignored "$SECRETS_DIR"
local kf="$SECRETS_DIR/${peer}.key"
[ -e "$kf" ] && { echo "REFUSING: $kf exists. Delete it deliberately to rotate." >&2; exit 1; }
( umask 077; wg genkey > "$kf" )
echo "private key -> $kf (0600, gitignored, never leaves this machine)"
echo
echo "public key for the estate description:"
echo " $(wg pubkey < "$kf")"
echo
echo "Paste that into estate/*.json under vpn.overlays.<ov>.peers.${peer}.public_key."
echo "The private key stays here and is injected only at render time."
}
render() {
local peer="${1:-}"
[ -z "$peer" ] && { echo "usage: $0 render <peer>" >&2; exit 1; }
mkdir -p "$OUT_DIR"
assert_ignored "$OUT_DIR"
local ov=estate
local found=""
local name addr role ep pk aips ka
while IFS=$'\x1f' read -r name addr role ep pk aips ka; do
[ "$name" = "$peer" ] && found=1 && break
done < <(overlay_peers "$ov")
[ -z "$found" ] && { echo "no such peer '$peer' in overlay '$ov'" >&2; exit 1; }
local missing=""
while IFS=$'\x1f' read -r name addr role ep pk aips ka; do
[ -z "$pk" ] && missing="$missing $name"
done < <(overlay_peers "$ov")
if [ -n "$missing" ]; then
echo "REFUSING to render: public keys not captured for:$missing" >&2
echo " A config without every peer's public key is a config that silently" >&2
echo " drops those peers. Capture first: sudo wg show | $0 capture --write" >&2
exit 1
fi
echo "would write $OUT_DIR/${peer}.conf (all keys present)"
}
# Reads `wg show` on stdin. Peers match by allowed-ips address, not public key,
# because the keys are what is missing. A roaming peer's endpoint is a home
# address and has no stable value — dropped in the parser, not just unused.
capture() {
local write="" as_peer=""
while [ $# -gt 0 ]; do
case "$1" in
--write) write=1 ;;
--as) shift; as_peer="${1:-}"
[ -z "$as_peer" ] && { echo "--as needs a peer name" >&2; exit 1; } ;;
*) echo "capture: unknown argument '$1'" >&2; exit 1 ;;
esac
shift
done
local input; input=$(cat)
if [ -z "$input" ]; then
echo "nothing on stdin." >&2
echo " run: sudo wg show | $0 capture" >&2
exit 1
fi
# Refuse the unsafe forms outright rather than parsing around them.
if printf '%s' "$input" | grep -qiE '^\s*PrivateKey\s*=|^\[Interface\]'; then
echo "REFUSING: this looks like 'wg showconf' output — it contains a PRIVATE KEY." >&2
echo " Use 'sudo wg show' (plain). It prints 'private key: (hidden)'." >&2
exit 1
fi
if ! printf '%s' "$input" | grep -q 'interface:'; then
echo "REFUSING: this does not look like 'wg show' output." >&2
echo " If it was 'wg show <if> dump': that form's first field IS the" >&2
echo " private key. Use 'sudo wg show' with no subcommand." >&2
exit 1
fi
WRITE="$write" AS_PEER="$as_peer" INPUT="$input" python3 - "$ESTATE_FILE" <<'PYCAP'
import collections, ipaddress, json, os, re, sys
text = os.environ["INPUT"]
write = os.environ.get("WRITE") == "1"
path = sys.argv[1]
iface, peers, cur = {}, [], None
for line in text.splitlines():
st = line.strip()
if st.startswith("interface:"):
cur = iface; cur["name"] = st.split(":", 1)[1].strip(); continue
if st.startswith("peer:"):
cur = {"public_key": st.split(":", 1)[1].strip()}; peers.append(cur); continue
if cur is None or ":" not in st:
continue
k, v = st.split(":", 1)
k, v = k.strip().lower(), v.strip()
if k == "private key":
continue # never recorded, whatever it says
if k == "public key": cur["public_key"] = v
elif k == "listening port": cur["listen_port"] = v
elif k == "allowed ips": cur["allowed_ips"] = v
elif k == "endpoint": cur["endpoint"] = v
elif k == "persistent keepalive":
m = re.search(r"(\d+)", v)
if m: cur["keepalive"] = int(m.group(1))
d = json.load(open(path), object_pairs_hook=collections.OrderedDict)
ov = d["vpn"]["overlays"]["estate"]
# address -> peer name, from what the estate already declares
by_addr = {p["address"]: n for n, p in ov["peers"].items() if p.get("address")}
by_key = {p["public_key"]: n for n, p in ov["peers"].items() if p.get("public_key")}
subnet = ipaddress.ip_network(ov["subnet"]) if ov.get("subnet") else None
hubs = [n for n, p in ov["peers"].items() if p.get("role") == "hub"]
# Whose interface block is this? `--as` names it explicitly, and that is the only
# thing that works for output captured over ssh: the addresses on THIS machine
# say nothing about the machine the output came from.
as_peer = os.environ.get("AS_PEER") or ""
if as_peer:
if as_peer not in ov["peers"]:
print("no peer named %r in this overlay. known: %s"
% (as_peer, ", ".join(ov["peers"])))
raise SystemExit(1)
me = as_peer
else:
me = None
local = os.popen(
"ip -4 -o addr show 2>/dev/null | awk '{split($4,a,\"/\"); print a[1]}'"
).read().split()
for n, p in ov["peers"].items():
if p.get("address") and p["address"] in local:
me = n
changes = []
conflicts = []
staged = {}
def setf(peer, field, val, why=""):
p = ov["peers"][peer]
if val is None or p.get(field) == val:
return
# Two values for one field means the input is from another machine.
prev = staged.get((peer, field))
if prev is not None and prev != val:
conflicts.append((peer, field, prev, val))
return
staged[(peer, field)] = val
changes.append((peer, field, p.get(field), val, why))
if write:
p[field] = val
if me and iface.get("public_key"):
setf(me, "public_key", iface["public_key"], "(this machine's interface)")
def match(pr):
# 1. The public key IS the identity. Use it whenever the estate knows it.
n = by_key.get(pr["public_key"])
if n:
return n
nets = [a.strip() for a in pr.get("allowed_ips", "").split(",") if a.strip()]
# 2. An allowed-ip that is a declared peer address — the ordinary spoke case.
for a in nets:
if a.split("/")[0] in by_addr:
return by_addr[a.split("/")[0]]
# 3. A peer routing the WHOLE overlay is the hub seen from a spoke. Its
# allowed_ips is the subnet itself, so no single address ever matches it.
if subnet and len(hubs) == 1:
for a in nets:
try:
if ipaddress.ip_network(a, strict=False).supernet_of(subnet):
return hubs[0]
except ValueError:
continue
return None
for pr in peers:
name = match(pr)
if not name:
changes.append(("?", "UNMATCHED", None,
"allowed_ips=%s key=%s" % (pr.get("allowed_ips"), pr["public_key"][:12] + "..."),
"no estate peer has this key, this address, or this route"))
continue
setf(name, "public_key", pr.get("public_key"))
setf(name, "allowed_ips", pr.get("allowed_ips"))
setf(name, "keepalive", pr.get("keepalive"))
# endpoint: recorded ONLY for a non-roaming peer. For a roaming one the
# value is a home ISP address and is deliberately dropped here.
if pr.get("endpoint"):
if ov["peers"][name].get("role") == "roaming":
changes.append((name, "endpoint", None, "(dropped: roaming peer)",
"a home address is the one sensitive field; roaming peers have no stable endpoint"))
else:
setf(name, "endpoint", pr["endpoint"])
if me and iface.get("listen_port"):
try:
lp = int(iface["listen_port"])
except ValueError:
lp = None
if lp is not None:
# A listen port belongs to the PEER, not to the overlay. A roaming peer's
# is an ephemeral source port chosen by the kernel; writing it to the
# overlay would rename the port the firewall rule is checked against.
setf(me, "listen_port", lp)
if ov["peers"][me].get("role") == "hub" and ov.get("listen_port") != lp:
changes.append(("(overlay)", "listen_port", ov.get("listen_port"), lp,
"the hub's port is the overlay's port"))
if write:
ov["listen_port"] = lp
if conflicts:
print("REFUSING: the same field was reported twice with different values.\n")
for peer, field, a, b in conflicts:
print(" %s.%s: %s vs %s" % (peer, field, a, b))
print("\nThis usually means `wg show` output from one machine was piped into")
print("capture on another. Run capture on the machine the output came from.")
raise SystemExit(1)
if not changes:
print("nothing to record — the estate already matches what wg reports.")
else:
print("%-9s %-12s %-22s %s" % ("PEER", "FIELD", "WAS", "WOULD BE"))
for peer, field, was, val, why in changes:
print("%-9s %-12s %-22s %s" % (peer, field, was if was is not None else "—", val))
if why: print(" %s" % why)
if write:
still = [n for n, p in ov["peers"].items() if not p.get("public_key")]
if not still:
d["vpn"]["_status"] = ("CAPTURED %s — public keys, allowed-ips and keepalive read from "
"`wg show`. Roaming endpoints deliberately not recorded."
% __import__("datetime").date.today())
json.dump(d, open(path, "w"), indent=2, ensure_ascii=False)
open(path, "a").write("\n")
print("\nwritten to %s" % path)
else:
print("\nnothing written. Add --write to record it.")
PYCAP
}
refuse() {
local verb="$1"; shift
local yes=""
for a in "$@"; do [ "$a" = "--yes" ] && yes=1; done
[ -z "$yes" ] && {
echo "refusing to $verb without --yes." >&2
echo " $verb changes live networking — it can cut the path this session" >&2
echo " is reaching the estate through. Read 'make vpn check' first." >&2
exit 1; }
echo "refusing to $verb: not implemented. Bringing a tunnel up or down is" >&2
echo " the host's business, and the live one is systemd-managed" >&2
echo " (wg-quick@wg0). berth describes and renders; it does not operate." >&2
exit 1
}
case "${1:-list}" in
list) list ;;
show) shift; show "${1:-}" ;;
check) check ;;
render) shift; render "${1:-}" ;;
keygen) shift; keygen "${1:-}" ;;
capture) shift; capture "$@" ;;
up|down) v="$1"; shift; refuse "$v" "$@" ;;
*) echo "usage: $0 [list|show <ov>|check|render <peer>|keygen <peer>|capture [--as <peer>] [--write]|up --yes|down --yes]" >&2; exit 1 ;;
esac

326
berth/estate/mcrn.json Normal file
View File

@@ -0,0 +1,326 @@
{
"_meta": {
"status": "UNVERIFIED — derived from the repos, not from the estate",
"why": "ppl/infra/ was written and never applied: no ~/.pulumi, no infra/venv, no stack state, files dated 'mar 6'. The estate was built in the console and the IaC is aspirational. B1's inventory is what replaces these values with observed ones; until it runs, every field here is a CLAIM.",
"never_record": "credential values. Resource ids and settings only. nova's gateway secret is deliberately absent from this file even though it is committed in plaintext in ppl/gateway/nginx/conf.d/nova.conf — see services[].raw.",
"sources": [
"ppl/infra/__main__.py",
"ppl/ctrl/dns.sh",
"ppl/ctrl/certs.sh",
"ppl/gateway/docker-compose.yml",
"ppl/gateway/nginx/conf.d/",
"ppl/local/Caddyfile"
],
"placement": {
"box": "a container on the estate's own docker network — upstream is the container name",
"local": "a rig cluster on a peer, reached over the overlay — upstream is that peer's address",
"instance": "a dedicated cloud instance on the overlay — same rendering as `local`",
"hosted": "a managed endpoint. Declared so moving to one is a one-line change; unused.",
"_why": "A service says WHERE it runs. How it is reached follows from that, and the three properties of a static-upstream vhost — upstream{}, no resolver, no set $var — are one decision rather than three."
}
},
"domain": "mcrn.ar",
"local_domain": "local.ar",
"host": "mcrn",
"host_admin": "mcrn-admin",
"instance": {
"type": "t3.small",
"disk_gb": 30,
"disk_type": "gp3",
"image": "debian-12",
"user": "mariano"
},
"firewall": [
{
"port": 22,
"proto": "tcp",
"desc": "SSH"
},
{
"port": 80,
"proto": "tcp",
"desc": "HTTP"
},
{
"port": 443,
"proto": "tcp",
"desc": "HTTPS"
},
{
"port": 3022,
"proto": "tcp",
"desc": "Gitea SSH",
"note": "compose maps 3022:22 but GITEA__server__SSH_PORT=22, so gitea advertises :22 in clone URLs while listening on :3022. B1 confirms which is real."
},
{
"port": 51820,
"proto": "udp",
"desc": "WireGuard",
"note": "ABSENT from ppl/infra/__main__.py's four rules — but the tunnel is live (ping 10.8.0.1 succeeds), so the real security group must already allow it. The code therefore does not describe the estate. Confirm in V1."
}
],
"network": {
"docker_network": "gateway",
"docker_network_note": "A fixed, externally-joinable bridge name. Every unrelated app stack on the box joins it so nginx can resolve them by container name. This is why the gateway compose declares 8 services while nginx routes 20+ hostnames.",
"wireguard_moved": "superseded by the top-level `vpn` block"
},
"vpn": {
"_status": "CAPTURED 2026-09-14 — public keys, allowed-ips and keepalive read from `wg show`. Roaming endpoints deliberately not recorded.",
"_never_record": "private keys. `wg show` prints 'private key: (hidden)' and is the safe capture command. `wg showconf` dumps PrivateKey= in clear — never use it.",
"overlays": {
"estate": {
"purpose": "Connects the estate's machines across clouds without a shared VPC, and carries everything that does not need to be publicly reachable.",
"subnet": "10.8.0.0/24",
"listen_port": 51820,
"peers": {
"box": {
"address": "10.8.0.1",
"role": "hub",
"note": "mcrn.ar. Has a public IP, so it is the peer others dial. Carries the registry (:5000) and woodpecker's gRPC (:9000), both bound to this address and therefore overlay-only.",
"endpoint": "3.23.204.197:51820",
"public_key": "zVYCmi3xucuX7k/aDhrOUPyN4GRk96ffSDD6dUFQjh4=",
"allowed_ips": "10.8.0.0/24",
"keepalive": 25,
"listen_port": 51820
},
"nrft": {
"address": "10.8.0.2",
"role": "roaming",
"note": "The dev box. Behind NAT, so it must initiate and needs PersistentKeepalive. Verified: wg0 UP at 10.8.0.2/24, ping 10.8.0.1 0% loss at 153ms.",
"endpoint": null,
"public_key": "zlIBGs4y5rt6uVdmFBasHpafht6ErxG+R3ySCg5rh3s=",
"allowed_ips": "10.8.0.2/32, 192.168.1.0/24",
"keepalive": null,
"listen_port": 36145
},
"work": {
"address": "10.8.0.3",
"role": "roaming",
"note": "A work computer, granted access when it was needed. Identified by the user at capture time, 2026-09-14 — it was NOT in the description before, and the wire is where it was found. No handshake and no transfer have ever been recorded for it, so it is a standing grant rather than a live peer: it can connect, and never has. Whether to keep or revoke it is the host's call.",
"endpoint": null,
"public_key": "ruSZwKt/p60GVsTLSAhcKBIXKkSZsf0gWSmSH1+UgE0=",
"allowed_ips": "10.8.0.3/32",
"keepalive": null
}
}
}
}
},
"databases": [
"gitea",
"woodpecker",
"umami"
],
"certs": {
"issued": [
"mcrn.ar",
"*.mcrn.ar",
"*.spr.mcrn.ar"
],
"issued_source": "ppl/ctrl/certs.sh:92 — the -d flags passed to certbot",
"note": "What the cert ACTUALLY covers. estate_sans() derives what the services NEED. check.sh compares the two; the difference is the finding, not a restatement."
},
"services": [
{
"name": "gitea",
"host": "git",
"upstream": "gitea:3000",
"targets": [
"aws"
]
},
{
"name": "woodpecker",
"host": "ci",
"upstream": "woodpecker-server:8000",
"targets": [
"aws"
]
},
{
"name": "registry",
"host": "registry",
"upstream": "registry:5000",
"targets": [
"aws"
]
},
{
"name": "umami",
"host": "analytics",
"upstream": "umami:3000",
"targets": [
"aws"
]
},
{
"name": "docserve",
"host": "docs",
"upstream": "docserve:8020",
"targets": [
"aws"
]
},
{
"name": "ghost",
"host": "notes",
"upstream": "ghost:2368",
"targets": [
"aws"
]
},
{
"name": "deskmeter",
"host": "deskmeter",
"upstream": "dmweb:10000",
"targets": [
"aws"
],
"local_port": 10000
},
{
"name": "sysmonstm",
"host": "sysmonstm",
"upstream": "sysmonstm-edge:8080",
"targets": [
"aws"
],
"local_port": 8020
},
{
"name": "malvalava",
"host": "malvalava",
"upstream": "mlvclean-frontend:80",
"targets": [
"aws"
],
"local_port": 30090
},
{
"name": "soleprint",
"host": "soleprint",
"upstream": "soleprint:8000",
"targets": [
"aws"
],
"local_port": 12000
},
{
"name": "dlt",
"host": "dlt.spr",
"upstream": "dlt_spr:8000",
"targets": [
"aws"
]
},
{
"name": "sample",
"host": "sample.spr",
"upstream": "sample_spr:8000",
"targets": [
"aws"
]
},
{
"name": "mariano",
"host": "mariano",
"kind": "static",
"targets": [
"aws"
]
},
{
"name": "rigui",
"host": "rig",
"kind": "static",
"targets": [
"aws"
],
"local_port": 20310
},
{
"name": "unt",
"host": "unt",
"targets": [
"local"
],
"local_port": 8040
},
{
"name": "mpr",
"host": "mpr",
"targets": [
"local"
],
"local_port": 30080
},
{
"name": "nvi",
"host": "nvi",
"targets": [
"local"
],
"local_port": 8060
},
{
"name": "eth",
"host": "eth",
"targets": [
"local"
],
"local_port": 8050
},
{
"name": "amar",
"host": "amar",
"targets": [
"local"
],
"local_port": 8030
},
{
"name": "nova",
"host": "nova",
"upstream": "nova-ui:80",
"targets": [
"aws"
],
"raw": true,
"raw_why": "Gated on an X-Gateway-Secret header whose value is committed in plaintext. The value is NOT recorded here. Worse: stellarair.conf proxies to the SAME nova-ui:80 upstream WITHOUT the check, so the gate is bypassable by hostname. Stays hand-written until that is decided."
},
{
"name": "stellarair",
"host": "stellarair",
"upstream": "nova-ui:80",
"targets": [
"aws"
],
"raw": true,
"raw_why": "See nova. Same upstream, no header gate."
},
{
"name": "langfuse",
"host": "langfuse",
"local_host": "lng",
"placement": "local",
"peer": "nrft",
"port": 3000,
"targets": [
"aws",
"local"
],
"local_port": 3000,
"note": "One service, one socket, two names. It was two entries with one flagged `raw`; placement is what made the exception expressible, so it is generated now."
},
{
"name": "legacy",
"host": "*.soleprint",
"upstream": "soleprint:8000",
"targets": [
"aws"
],
"raw": true,
"raw_why": "A regex server_name with a named capture plus sub_filter injection — not expressible as a template. ALSO BROKEN: its /api/, /admin/, /static/ and / blocks proxy to 127.0.0.1, i.e. inside the nginx container where nothing listens, so every legacy room 502s. Only /wrapper/ uses the correct container-name form."
}
]
}

View File

@@ -6,16 +6,49 @@
# ./ctrl/cluster.sh down # delete it (drops every room's namespace)
# ./ctrl/cluster.sh status # what's running on it
#
# One target, one script — the variants live here. The kind-*.sh files stay
# exactly as they are and remain runnable on their own; this only dispatches.
# spr depends on rig, never the other way round. Building and deleting a cluster
# is rig's job, so up and down hand straight to rig/ctrl/cluster.sh, carrying the
# four things that make this cluster spr's rather than rig's defaults:
#
# CLUSTER=spr rooms deploy into the kind-spr context
# KIND_CONFIG spr's own shape, which maps the rooms' gateway NodePorts
# REGISTRY_MODE=none rooms load images straight into the node
# PROFILE=minimal pinned here, so a change to rig's own ctrl/.env can never
# quietly add addons to spr's cluster
#
# status stays here: it answers a question about rooms, not about the cluster.
set -e
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
RIG_CTRL="$SCRIPT_DIR/../rig/ctrl"
rig() {
CLUSTER=spr \
KIND_CONFIG="$SCRIPT_DIR/k8s/kind-config.yaml" \
REGISTRY_MODE=none \
PROFILE=minimal \
bash "$RIG_CTRL/cluster.sh" "$@"
}
case "${1:-status}" in
up) exec "$SCRIPT_DIR/kind-up.sh" ;;
down) exec "$SCRIPT_DIR/kind-down.sh" ;;
status) exec "$SCRIPT_DIR/kind-status.sh" ;;
up)
rig up
echo
echo "Per-room deploy:"
echo " cd gen/<room> && ./ctrl/k8s-up.sh"
;;
down)
rig down
;;
status)
if ! kind get clusters 2>/dev/null | grep -qx spr; then
echo "No 'spr' kind cluster — run: make cluster up"
exit 0
fi
kubectl --context kind-spr get namespaces -l soleprint-room
echo
kubectl --context kind-spr get pods -A -l soleprint-room
;;
*)
echo "Unknown subcommand: $1" >&2
echo "Usage: cluster.sh [up|down|status]" >&2

View File

@@ -3,9 +3,15 @@ apiVersion: kind.x-k8s.io/v1alpha4
# Single shared cluster for all soleprint rooms.
# Each room deploys into its own namespace; gateway Services pick a
# NodePort from the 30080-30099 range mapped here.
name: spr
#
# Built by rig, not by spr: ctrl/cluster.sh hands this file to
# rig/ctrl/cluster.sh, which substitutes CLUSTER and NODE_IMAGE (named without
# braces here so this comment survives the substitution). The shape is spr's —
# what its cluster needs is spr's business. Building it is rig's.
name: ${CLUSTER}
nodes:
- role: control-plane
image: ${NODE_IMAGE}
extraPortMappings:
# Room gateway NodePorts (one per active room).
- {containerPort: 30080, hostPort: 30080, protocol: TCP}

View File

@@ -1,12 +0,0 @@
#!/bin/bash
# Delete the shared `spr` kind cluster (drops every room's namespace too).
# Use `gen/<room>/ctrl/k8s-down.sh` instead if you only want to remove
# a single room's namespace.
set -e
if kind get clusters 2>/dev/null | grep -q '^spr$'; then
echo "Deleting kind cluster 'spr'..."
kind delete cluster --name spr
else
echo "No kind cluster 'spr' to delete."
fi

View File

@@ -1,12 +0,0 @@
#!/bin/bash
# Show what's running on the shared `spr` cluster.
set -e
if ! kind get clusters 2>/dev/null | grep -q '^spr$'; then
echo "No 'spr' kind cluster — run ctrl/kind-up.sh"
exit 0
fi
kubectl --context kind-spr get namespaces -l soleprint-room
echo
kubectl --context kind-spr get pods -A -l soleprint-room

View File

@@ -1,21 +0,0 @@
#!/bin/bash
# Create (or no-op) the single shared `spr` kind cluster used by every
# soleprint room. Per-room work happens inside namespaces — see
# `gen/<room>/ctrl/k8s-up.sh`.
set -e
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
KIND_CONFIG="$SCRIPT_DIR/k8s/kind-config.yaml"
if kind get clusters 2>/dev/null | grep -q '^spr$'; then
echo "Kind cluster 'spr' already exists."
else
echo "Creating kind cluster 'spr'..."
kind create cluster --config "$KIND_CONFIG"
fi
kubectl config use-context kind-spr >/dev/null
echo
echo "Cluster ready. Per-room deploy:"
echo " cd gen/<room> && ./ctrl/k8s-up.sh"

80
ctrl/theme.sh Executable file
View File

@@ -0,0 +1,80 @@
#!/usr/bin/env bash
# Bake the theme and its parts into the pages that use them.
#
# Usage:
# ./ctrl/theme.sh # bake — rewrite every generated block
# ./ctrl/theme.sh check # fail if any page is stale; changes nothing
# ./ctrl/theme.sh new [title] # a scaffold page to start from
# ./ctrl/theme.sh run FILE [--list|--check] [--only NAME]
# # every page a run file lists, each with a contract
# ./ctrl/theme.sh parts # what can be added, and the markup that adds it
# ./ctrl/theme.sh export [name...] # the contract for a subset, as one doc
#
# `export` is for handing a vetted LLM what it needs to write an ad-hoc page —
# the chosen parts, their markup, and the tokens resolved to literal values, so
# the document stands alone. Naming parts is the point: hand over everything and
# you get back a page built from Vue components that cannot run standalone.
#
# ./ctrl/theme.sh export panel split > /tmp/contract.md
#
# Call the script directly when piping; `make` echoes its recipe to stdout.
# For whole-repo context this is the wrong tool — station/tools/distill already
# flattens a tree to one budgeted document.
#
# A page that says `background: var(--bg)` and never gets `--bg` is UNSTYLED,
# not merely unbranded — the declaration is invalid at computed-value time. That
# is why every page carries a baked default, and why `check` is worth running.
#
# This exists because bake.py was reachable by no command at all: not from the
# Makefile, not from ctrl/, not from build.py. A drift check nobody runs is a
# drift check that reports nothing, and the evidence was already on disk —
# histgen's page linked /theme.css for months, was missing from the old
# hardcoded page list, and so was never baked once.
set -e
# Where the caller stood. A run file is named relative to there, and the cd below
# would otherwise make `run ./theme.toml` mean a different file.
CALLER_DIR="$PWD"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
ROOT_DIR="$(dirname "$SCRIPT_DIR")"
cd "$ROOT_DIR/soleprint"
PYTHON="${PYTHON:-python3}"
case "${1:-bake}" in
bake) exec "$PYTHON" common/theme/bake.py ;;
check) exec "$PYTHON" common/theme/bake.py --check ;;
parts)
exec "$PYTHON" common/theme/bake.py --parts
;;
run)
# A run file: every page a project has, its context, and a contract per
# page for the LLM. See soleprint/common/theme/theme.example.toml.
shift
args=() file="" prev=""
for a in "$@"; do
if [[ -z "$file" && "$a" != -* && "$prev" != "--only" ]]; then
file="$(cd "$CALLER_DIR" && realpath -m -- "$a")"
args+=("$file")
else
args+=("$a")
fi
prev="$a"
done
exec "$PYTHON" common/theme/bake.py --run "${args[@]}"
;;
new)
shift
exec "$PYTHON" common/theme/bake.py --new "$@"
;;
export)
shift
exec "$PYTHON" common/theme/bake.py --export "$@"
;;
*)
echo "Unknown: $1" >&2
echo "Usage: ./ctrl/theme.sh [new [title]|parts|bake|check|export [name...]|run FILE]" >&2
exit 1
;;
esac

11
rig/.gitignore vendored
View File

@@ -9,12 +9,13 @@ ctrl/.env
# generated: the .dot is a build artifact rendered from arch/*.json, never hand-edited.
# The .svg IS committed — onboarding material should render in a repo browser.
arch/*.dot
ctrl/Tiltfile.gen
# ctrl/Tiltfile.gen was here for a generator that no longer exists. ctrl/Tiltfile
# is now a real, committed file that derives its values when Tilt parses it, so
# there is nothing generated to ignore.
# binaries pulled by `make deps-bundle` for the air-gapped wizard image
# binaries pulled by `make deps-bundle` for the air-gapped installer image
vendor
# Client rigs are NOT ignored here. A copy is a SIBLING of this directory
# (spr/acme-rig), so a rule in this file cannot see it — the rules live in
# spr/.gitignore, anchored at spr's root, where `*-rig/` matches the siblings and
# `!rig/sample-rig/` keeps the committed stand-in.
# (../acme-rig), so a rule in this file cannot see it — the rules live in the
# parent repo's .gitignore, anchored at its root, where `*-rig/` matches them.

View File

@@ -4,9 +4,8 @@ The README says the prerequisite is Docker and nothing else. This is what that
actually looks like end to end: a bare Linux box, and a new project running under
Tilt at the end of it.
rig lives inside soleprint, at `spr/rig` — it is soleprint's cluster half, and
a client copy is a sibling (`spr/acme-rig`). Paths below are relative to
soleprint's checkout.
A copy of this directory is a sibling of it, named after the environment it
models (`acme-rig`). Paths below are relative to the parent checkout.
It spans three repos because the work does. **rig** prepares the machine — the
pinned toolchain, the cluster, the port arithmetic. **all** owns the shape a
@@ -50,7 +49,7 @@ isn't.
## Read the docs before installing anything
```bash
cd spr/rig
cd rig
make docs
```
@@ -67,11 +66,11 @@ persists; ctrl-c ends it.
## Ask what is wrong with this machine
```bash
make station
make check
cp ctrl/.env.example ctrl/.env
```
`station.sh` reports and instructs, and fixes nothing. It runs bare rather than
`check.sh` reports and instructs, and fixes nothing. It runs bare rather than
in a container because host detection only ever reads `/proc` and `/etc` — no
dependency beyond coreutils.
@@ -80,7 +79,7 @@ port rig binds derives from this directory's name, so the answer is specific to
this copy, and a clash here surfaces as an opaque `failed to bind host port` in
the middle of cluster creation if you skip it.
Copy the `.env` even though station only warns about it. It is gitignored, it is
Copy the `.env` even though the check only warns about it. It is gitignored, it is
where a machine-local override goes, and `ports.sh persist` expects it to exist.
@@ -88,33 +87,33 @@ where a machine-local override goes, and `ports.sh persist` expects it to exist.
This is the step where "nothing installed" stops being rhetorical.
`make deps` runs `ctrl/wizard.sh install` directly on the host, and the wizard
`make deps` runs `ctrl/deps.sh install` directly on the host, and the installer
fetches with `curl`. A stock `debian:trixie-slim` has no curl — detection runs
fine, then the first download dies with `curl: command not found` and an exit
code of 127. That is the bootstrap paradox `ctrl/Dockerfile.wizard` exists
to kill — the wizard carries its own toolchain so the host needs only Docker —
code of 127. That is the bootstrap paradox `ctrl/Dockerfile.deps` exists
to kill — the installer carries its own toolchain so the host needs only Docker —
but building the image and running it are two different things, and only the
build has a Makefile target today. **On a genuinely bare machine, run it by
hand:**
```bash
make wizard # builds rig-wizard:wizard
make deps-image # builds rig-deps:deps
mkdir -p ~/.local/bin
docker run --rm \
-v /:/host:ro \
-v /var/run/docker.sock:/var/run/docker.sock \
-v "$HOME/.local/bin:/out/bin" \
-e HOST_UID="$(id -u)" -e HOST_GID="$(id -g)" \
rig-wizard:wizard install dev
rig-deps:deps install dev
```
The image name follows the directory, like everything else here: in `spr/rig`
it is `rig-wizard`, in a copy called `spr/acme-rig` it is `acme-rig-wizard`. The
tag is `wizard` (or `full`, below), not `latest`.
The image name follows the directory, like everything else here: in `rig`
it is `rig-deps`, in a copy called `acme-rig` it is `acme-rig-deps`. The
tag is `deps` (or `full`, below), not `latest`.
None of the four arguments are guessable, so:
- **`/:/host:ro`** — the wizard reads the *host's* `/etc/os-release` and
- **`/:/host:ro`** — the installer reads the *host's* `/etc/os-release` and
`/etc/wsl.conf`, not the container's. `HOST_ROOT=/host` is already baked into
the image; this is what it points at. Read-only, and it is the only reason
detection inside a container tells you anything about the machine.
@@ -122,7 +121,7 @@ None of the four arguments are guessable, so:
and how it counts kind clusters already running.
- **`/out/bin`** — the image's `OUT_BIN`. Whatever you mount here is where the
four binaries land.
- **`HOST_UID` / `HOST_GID`** — the wizard runs as root so it can reach that
- **`HOST_UID` / `HOST_GID`** — the installer runs as root so it can reach that
socket, which means everything it writes into a mounted volume is root-owned
and useless to you. These drive the `chown` back. Omit them and the install
looks like it worked.
@@ -131,18 +130,18 @@ None of the four arguments are guessable, so:
tooling — which is the right answer on a managed or corporate-issued machine and
is why the split exists.
Then put them on PATH, which the wizard will remind you about because it cannot
Then put them on PATH, which the installer will remind you about because it cannot
edit your shell for you:
```bash
export PATH="$HOME/.local/bin:$PATH" # and add the same line to ~/.bashrc
```
If something else on this machine already provides `kubectl`, the wizard says so
If something else on this machine already provides `kubectl`, the installer says so
by name rather than shadowing it quietly. `OUT_BIN=$PWD/def/bin` installs
somewhere private instead.
**Two variants worth knowing before you need them.** `make wizard full` bakes
**Two variants worth knowing before you need them.** `make deps-image full` bakes
every pinned binary into the image at build time (`DEPS_SOURCE=baked`), so
`docker save` gives you the entire installer as one file to carry into an
air-gapped network. And `DEPS_SOURCE=artifactory` with `DEPS_ARTIFACTORY_URL`
@@ -195,8 +194,8 @@ follows is only the mechanical part.
```bash
SLUG=<slug> # short, lowercase, no separators
cp -r ~/wdir/all/projects/templates/broad ~/wdir/"$SLUG"
cd ~/wdir/"$SLUG"
cp -r ~/wdir/semester/all/projects/templates/broad ~/wdir/semester/"$SLUG"
cd ~/wdir/semester/"$SLUG"
grep -rl '<slug>' ctrl | xargs sed -i "s/<slug>/$SLUG/g"
cp ctrl/k8s/.env.example ctrl/k8s/.env
git init && git add -A && git commit -m "scaffold $SLUG from broad"
@@ -214,7 +213,7 @@ scaffold's Makefile from the directory, so there is nothing to edit for either.
is already in use, so copying it unchanged puts two projects on one port:
```bash
grep -h '^TILT_PORT=' ~/wdir/*/ctrl/k8s/.env 2>/dev/null | sort
grep -h '^TILT_PORT=' ~/wdir/semester/*/ctrl/k8s/.env 2>/dev/null | sort
```
Choose a free one in `1030010399` — the range ALL reserves in
@@ -244,12 +243,22 @@ The workload is an nginx placeholder so a fresh copy reaches something that
answers; replace it. Keep `30080` in step between the overlay patch and
`kind-config.yaml`'s `containerPort` — the hostPort is this project's to pick.
Reachability is a plain kind port mapping: no ingress controller and no MetalLB.
Caddy maps `<slug>.local.ar` onto the host port (`~/wdir/ppl/local/Caddyfile`),
Caddy maps `<slug>.local.ar` onto the host port (`~/wdir/semester/ppl/local/Caddyfile`),
with `*.local.ar` resolving to 127.0.0.1 through dnsmasq. That is the whole chain.
**The one file the scaffold still does not ship is `ctrl/Tiltfile`**`make
tilt-up` runs `cd ctrl && tilt up`, and there is nothing to run until you write
one. Copy it from a live project; `unt` and `nvi` are closest to the plain shape.
**For `ctrl/Tiltfile`, copy rig's** rather than a live project's. rig ships one
that derives its cluster, context, ports and manifest directory from
`ctrl/ports.sh active` instead of hardcoding a slug, and carries a catalogue of
the blocks every project here ends up needing. Copying from `unt` or `nvi` is
what the estate did until now, and it is why the same Tiltfile preamble exists
in six places with the slug typed in by hand five times each.
> **Two things in this document disagree with rig and are not settled.** It
> mandates Tilt ports in `1030010399`, while rig derives a block from the
> directory name at `20000+` so copies cannot collide — a rig-managed project
> takes rig's. And it names `ctrl/k8s/.env.example`, which is the `broad`
> scaffold's layout; rig's is `ctrl/.env.example`. Both are this document
> describing the house scaffold from inside rig's tree.
## Run it
@@ -270,9 +279,9 @@ delete-and-recreate for when a cluster wedges.
## Register it
The project exists; now it is findable. Add an entry to
`~/wdir/all/projects/index.json` and write its `projects/<slug>.md` beside the
`~/wdir/semester/all/projects/index.json` and write its `projects/<slug>.md` beside the
others. Structured fields in the index, prose in the markdown.
Putting it on the CI server and deploying it is `ppl`'s half, and it starts at
`~/wdir/ppl/ctrl/init-repo.sh` — gitea remote, then Woodpecker. That is a
`~/wdir/semester/ppl/ctrl/init-repo.sh` — gitea remote, then Woodpecker. That is a
different document.

View File

@@ -5,22 +5,37 @@
# bash file, and that file holds the variants:
#
# make cluster up -> ctrl/cluster.sh up
# make newbox destroy -> ctrl/newbox.sh destroy
#
# Config layers, weakest first: ctrl/versions.env (pinned toolchain) <
# ctrl/env.d/<profile>.env (cluster shape) < ctrl/.env (local, gitignored) <
# the environment. So `make cluster up PROFILE=client` beats everything.
# Config layers, weakest first: built-in defaults < ctrl/versions.env (pinned
# toolchain) < ctrl/env.d/<profile>.env (optional shape) < ctrl/.env (local,
# gitignored) < the environment. So `make cluster up PROFILE=<name>` beats them all.
#
# Start with: make setup (then: make cluster up && make docs)
# Identity follows the FOLDER NAME, so this directory can be copied elsewhere,
# renamed, and run as a separate environment with no edits. ctrl/.env overrides
# it when you want a name that differs from the directory.
#
# Asked once, of ctrl/ports.sh, which resolves it through lib/config.sh:
#
# CLUSTER KUBECONTEXT HTTP HTTPS TILT REGISTRY MANIFESTS_DIR
#
# Read positionally, so the order is a contract — ctrl/selftest.sh pins it.
#
# This used to be sed over ctrl/.env plus a slug computed here, which is a
# SECOND derivation of values lib/config.sh already owns — and the two could
# disagree about the port after `make ports persist`, or about the name for any
# directory whose sanitised form differs from its raw one. One source now; the
# Tiltfile reads the same line.
FACTS := $(shell bash ctrl/ports.sh active 2>/dev/null)
SLUG := $(shell echo '$(notdir $(CURDIR))' | tr '[:upper:]' '[:lower:]' | tr -c 'a-z0-9-' '-' | sed 's/^-*//; s/-*$$//')
CLUSTER := $(or $(shell sed -n 's/^CLUSTER=//p' ctrl/.env 2>/dev/null),$(SLUG))
KCTX := --context kind-$(CLUSTER)
TILT_PORT := $(shell sed -n 's/^TILT_PORT=//p' ctrl/.env 2>/dev/null)
WIZARD := $(SLUG)-wizard
# The fallback matters: ports.sh sources config.sh, and if a profile or .env is
# broken it exits non-zero. Losing the cluster name would send --context to the
# wrong place, so fall back to the folder rather than to empty.
CLUSTER := $(or $(word 1,$(FACTS)),$(SLUG))
KCTX := --context $(or $(word 2,$(FACTS)),kind-$(SLUG))
TILT_PORT := $(word 5,$(FACTS))
DEPSIMG := $(SLUG)-deps
# Words after the target become the script's subcommand. Make would otherwise
# treat them as goals of their own, so each gets a no-op rule.
@@ -35,8 +50,8 @@ $(eval $(ARGS):;@:)
.PHONY: $(ARGS)
endif
.PHONY: help setup station deps wizard cluster registry addons ports \
newbox dockerhost docs tilt \
.PHONY: help setup check selftest mem deps deps-image standalone cluster registry addons ports \
docs tilt \
kind-up kind-down kind-reset tilt-up tilt-down
help: ## list targets
@@ -44,19 +59,34 @@ help: ## list targets
# ── setup ──────────────────────────────────────────────────────────────────
setup: ## prepare this machine [core] [--share-docker] [--cluster]
setup: ## prepare this machine [core] [--cluster]
bash ctrl/setup.sh $(ARGS)
station: ## is this workstation ready? reports, never fixes
bash ctrl/station.sh
check: ## is this machine ready? reports, never fixes
bash ctrl/check.sh
# The counterpart to check: that one asks about the MACHINE and never fails,
# this one asks about RIG and exits 1, as `make standalone check` does. Checks are written
# as the decisions they defend, so a failure names what is being undone.
selftest: ## does rig still do what it says? exits 1 if not
bash ctrl/selftest.sh
mem: ## memory, its caps, what it survives [status|push|all|backup|restore]
bash ctrl/mem.sh $(or $(ARGS),status)
deps: ## install the toolchain [core|dev] (default dev)
bash ctrl/wizard.sh install $(or $(ARGS),dev)
bash ctrl/deps.sh install $(or $(ARGS),dev)
wizard: ## build the installer image [full]
docker build -f ctrl/Dockerfile.wizard \
--target $(if $(filter full,$(ARGS)),wizard-full,wizard) \
-t $(WIZARD):$(if $(filter full,$(ARGS)),full,wizard) .
# The one-file versions of rig's tools, one folder per profile, for machines the
# full rig is not going to. Generated from rig as it is, never edited by hand;
# `check` is what selftest runs to catch a kit left behind by a change to rig.
standalone: ## single-file kits [write|check|export DIR] (default write)
bash ctrl/standalone.sh $(or $(ARGS),write)
deps-image: ## build the installer image [full]
docker build -f ctrl/Dockerfile.deps \
--target $(if $(filter full,$(ARGS)),deps-full,deps) \
-t $(DEPSIMG):$(if $(filter full,$(ARGS)),full,deps) .
# ── cluster ────────────────────────────────────────────────────────────────
@@ -72,24 +102,21 @@ addons: ## profile addons [install|list] (default
ports: ## this environment's port block [show|persist]
bash ctrl/ports.sh $(or $(ARGS),show)
# ── host ───────────────────────────────────────────────────────────────────
newbox: ## throwaway environment [create|status|shell|destroy]
bash ctrl/newbox.sh $(or $(ARGS),status)
dockerhost: ## share Docker between distros [status|share|unshare]
$(if $(filter share unshare,$(ARGS)),sudo ,)bash ctrl/dockerhost.sh $(or $(ARGS),status)
# ── docs + dev loop ────────────────────────────────────────────────────────
docs: ## documentation [serve|graphs] (default serve)
bash ctrl/docs.sh $(or $(ARGS),serve)
# --port is only passed when TILT_PORT is actually set. It comes from ctrl/.env,
# which does NOT carry it by default — ports are derived at runtime in
# lib/config.sh unless `make ports persist` has written them. Without the guard
# tilt receives a bare `--port` with no value and fails on the flag rather than
# on anything real. `make ports show` prints the derived block.
# --port is only passed when TILT_PORT resolved. It normally does, since FACTS
# above asks ports.sh — but ports.sh can fail on a broken profile, and without
# the guard tilt receives a bare `--port` with no value and fails on the flag
# rather than on anything real. Tilt's own default is 10350, which is the number
# every project on this machine is trying not to collide on, so falling back to
# it silently is worse than not passing the flag.
#
# The Tiltfile asks ports.sh for the rest itself — cluster, registry and where
# the manifests are — so nothing needs passing here beyond what tilt's own flags
# require.
tilt: ## dev loop [up|down] (default up)
cd ctrl && tilt $(or $(ARGS),up) $(KCTX) $(if $(filter down,$(ARGS)),,$(if $(TILT_PORT),--port $(TILT_PORT)))
@@ -104,6 +131,11 @@ tilt: ## dev loop [up|down] (default
#
# `cluster list` and `cluster free` have no hyphenated twin on purpose — they
# are rig's own, with nothing to be consistent with.
#
# Nothing outside this file reads these names: the script is `ctrl/cluster.sh`
# and it takes the verb. So rename them, delete the ones you never type, or add
# the spelling your own projects use — an alias is two lines, and adding one
# costs nothing but a line in .PHONY above.
kind-up: ## alias for `cluster up`
bash ctrl/cluster.sh up
@@ -114,10 +146,10 @@ kind-down: ## alias for `cluster down`
kind-reset: ## alias for `cluster reset`
bash ctrl/cluster.sh reset
# These two match the other projects' spelling, but rig has no Tiltfile there
# is nothing to run yet, and they fail the same way `make tilt` does.
tilt-up: ## alias for `tilt up` (rig has no Tiltfile yet)
# These two match the other projects' spelling. rig ships ctrl/Tiltfile, so they
# run — it deploys the examples in k8s/base until you replace them.
tilt-up: ## alias for `tilt up`
cd ctrl && tilt up $(KCTX) $(if $(TILT_PORT),--port $(TILT_PORT))
tilt-down: ## alias for `tilt down` (rig has no Tiltfile yet)
tilt-down: ## alias for `tilt down`
cd ctrl && tilt down $(KCTX)

View File

@@ -9,6 +9,46 @@ topology, not the workloads.
**Docker.** Nothing else — no curl, no jq, no python, no apt repositories.
### Starting from plain Windows
Everything here is bash and runs *inside* a Linux shell, so on a Windows machine
that means WSL. Nothing in rig installs WSL, and nothing will: `wsl --install`
enables Windows features and requires a reboot, which is not something a script
should do to a machine on your behalf — and there is no tested undo for it.
From an elevated PowerShell or Command Prompt, once:
```powershell
wsl --install
```
Then reboot and open the Linux shell it installed.
**If you cloned this on the Windows side, copy it into WSL before carrying on.**
WSL can reach the Windows drives at `/mnt/c`, and working from there mostly
functions — slowly — but file watching does not: that filesystem raises no
inotify events, so anything watching for edits silently stops seeing them.
```bash
cp -r /mnt/c/Users/<you>/rig ~/rig
cd ~/rig
```
`make deps` reports it if you are running from `/mnt/...`. Then carry on below.
If it fails, the usual causes give unhelpful messages:
| symptom | cause |
| --- | --- |
| "the virtual machine could not be started" | virtualization disabled in BIOS/UEFI |
| the command is not recognised | Windows build too old — needs 2004 or later |
| the install starts, then nothing works | a reboot is still pending |
Running the scripts from **Git Bash, MSYS or Cygwin does not work** — those look
close enough to a Linux shell to get started and then fail without `/proc` or a
docker socket. `ctrl/deps.sh` detects that and says so rather than letting you
find out the slow way.
## Read the docs first
```bash
@@ -21,16 +61,65 @@ instructions for everything else. No cluster and no toolchain required.
## Then
```bash
make station # report host and config problems; changes nothing
make check # report host and config problems; changes nothing
make deps # install the toolchain (add `core` on a managed machine)
make cluster up # build the cluster for the active profile
```
`make cluster up` also starts this environment's local registry and wires it
into the node, so an image built locally is pullable by the cluster without
going near docker.io:
```bash
make registry status # prints: endpoint localhost:<port>
docker build -t localhost:<port>/app:1 .
docker push localhost:<port>/app:1
kubectl --context kind-$(basename $PWD) run app --image=localhost:<port>/app:1
```
The port block is derived from the directory name, so two copies of rig never
collide:
```bash
make ports show # HTTP / HTTPS / TILT / REGISTRY
make cluster list # every cluster on this machine, with memory
make cluster free # stop the others if memory is tight
make cluster down # remove this cluster and its registry
```
**The verbs are yours to change.** `cluster` is the script — `ctrl/cluster.sh`
and every spelling above is a `Makefile` target that calls it. `make kind-up` is
an alias for `make cluster up`, kept because the other projects on this machine
answer to that spelling and muscle memory spans repos rather than stopping at
one. Nothing outside the `Makefile` reads these names, so rename them, drop the
ones you never type, or add whatever your own projects already say: each alias
is two lines at the bottom of the file, calling the same script the canonical
target does.
**`make tilt` works on a fresh copy, unedited.** rig ships `ctrl/Tiltfile`, and
`k8s/base` already boots, so the dev loop comes up with the two examples running
and nothing to configure first.
It hardcodes nothing. It asks `ctrl/ports.sh active` for this environment's
cluster name, kube context, ports and manifest directory — the same values every
other rig script resolves through `ctrl/lib/config.sh` — so a copied and renamed
rig deploys into its own cluster with no edits. Every other project here writes
its slug into the Tiltfile five or six times by hand, which is exactly the
collision `kind-config.yaml.tpl` exists to avoid.
What it deploys is `MANIFESTS_DIR`, defaulting to rig's own `ctrl/k8s/overlays/dev`.
Point that at an overlay versioned elsewhere and rig stops owning the manifests.
Replace the examples, then add your images and resources in the two marked
sections. The catalogue below them holds the blocks that recur across every
project here — `docker_build`, resource ordering, gateway reload, port-forwards —
with the parts that are easy to get wrong already commented.
`make help` lists every target.
On a machine where Docker really is the only thing installed, `make deps` has
nothing to download with — see [BOOTSTRAP.md](BOOTSTRAP.md), which runs the
toolchain through the wizard container and carries on to scaffolding and running
toolchain through the installer container and carries on to scaffolding and running
a new project.
## One directory is one environment
@@ -39,24 +128,29 @@ Copy this directory, rename it, run it. Cluster name, kubectl context, image
tags and the host port block all derive from the directory name, so copies never
collide and neither one's teardown can touch the other.
rig lives inside soleprint, at `spr/rig` — it is soleprint's cluster half, and a
copy is a **sibling**: `spr/acme-rig`. That is why the ignore rules for client
rigs sit in `spr/.gitignore` rather than here; a rule in this directory cannot
see a directory beside it.
A copy of this directory is a **sibling** of it, named after the environment it
models (`acme-rig`). That is why the ignore rules for copies sit in the *parent*
repo's `.gitignore` rather than here: a rule in this directory cannot see a
directory beside it.
## Profiles
A profile is the shape of the cluster: how many nodes, which addons, whether the
apiserver audits. They live in `ctrl/env.d/`, and the active one is `PROFILE`.
**rig needs no profile.** With none named it runs on built-in defaults: one node,
no addons, a local registry, the newest Kubernetes version it pins. A profile is
an optional overlay — a file in `ctrl/env.d/`, named by `PROFILE` — for when you
want a different shape.
rig ships **examples**, not active profiles, because each one is a use case rather
than something every rig needs. Copy one to use it:
| Profile | For |
| example | shape |
| --- | --- |
| `minimal` | the default. One node, no addons, boots fast. |
| `client` | the regulated-estate shape — multi-node, audit on, registry mirror. |
| `offline` | air-gapped: everything from a preloaded local registry. |
| `data` | the dependency containers a soleprint room asks for. |
| `client.env.example` | multi-node, audit on, images through a mirror of a corporate registry |
| `offline.env.example` | air-gapped: everything from a preloaded local registry |
| `data.env.example` | databases and a scheduler: postgres, redis, airflow |
```bash
cp ctrl/env.d/data.env.example ctrl/env.d/data.env
PROFILE=data make cluster up
PROFILE=data make addons install
make addons # what the active profile wants, and what exists
@@ -67,9 +161,9 @@ restating node count and audit as variables:
| shape | nodes | audit | used by |
| --- | --- | --- | --- |
| `kind-config.yaml.tpl` | 1 | off | `minimal`, `data` |
| `kind-config.audit.yaml.tpl` | 1 | on | `offline` |
| `kind-config.client.yaml.tpl` | 3 | on | `client` |
| `kind-config.yaml.tpl` | 1 | off | the default; the `data` example |
| `kind-config.audit.yaml.tpl` | 1 | on | the `offline` example |
| `kind-config.client.yaml.tpl` | 3 | on | the `client` example |
Both numbers are read back out of the chosen file, so the YAML is the only place
that decides and there is nothing to drift. The layout under `ctrl/k8s/` is the
@@ -98,10 +192,10 @@ cluster does.
| `redis` | cache and broker |
| `airflow` | scheduled pipelines; needs postgres and redis |
The last three are the cluster half of **soleprint's cabinets**. A room declares
what it needs once, in `cfg/<room>/data/cabinets.json`; soleprint's `build.py`
composes those services into `docker-compose.yml` for a laptop, and these
install the same ones here. The names match on purpose each cabinet carries a
The last three are **cabinets**: a public service dropped in as-is, the upstream
image unmodified, reachable at a known address. A cabinet is declared once and
installs on either target — a `service.yml` composes it for a laptop, and these
install the same one here. The names match on purpose: each cabinet carries a
`rig_addon` field pointing at `ctrl/addons/<name>.sh`.
Plain manifests rather than helm charts, like every other addon: a chart repo is

View File

@@ -1,9 +1,10 @@
# Machine-local config. Copy to ctrl/.env (gitignored) and edit.
# Cluster SHAPE lives in ctrl/env.d/<profile>.env — not here.
# Cluster SHAPE: an optional profile in ctrl/env.d/ — see the *.env.example there.
# The architecture MODEL lives in arch/<name>.json — not here either.
# Which profile in ctrl/env.d/ to build. minimal | client | offline
PROFILE=minimal
# A profile in ctrl/env.d/ to build. Empty means rig's built-in defaults, which
# need no profile at all. Copy an env.d/*.env.example to <name>.env to add one.
PROFILE=
# Cluster name; the kubectl context becomes kind-<CLUSTER>.
# LEAVE THIS UNSET unless you need a name that differs from the directory —
@@ -26,10 +27,10 @@ PROFILE=minimal
# MANIFESTS_DIR=../platform-manifests/overlays/dev
MANIFESTS_DIR=ctrl/k8s/overlays/dev
# Where the wizard fetches the pinned binaries from.
# Where the installer fetches the pinned binaries from.
# upstream GitHub releases / dl.k8s.io (needs internet)
# artifactory a generic repo — what a locked-down client usually allows
# baked already inside the wizard image; no network at all
# baked already inside the installer image; no network at all
DEPS_SOURCE=upstream
DEPS_ARTIFACTORY_URL=
@@ -43,7 +44,7 @@ REGISTRY_PASSWORD=
# Corporate root CA, if Artifactory is fronted by an internal CA (it usually is).
# Trust has to reach THREE places and nothing does it for you: the host docker
# daemon, every kind node's containerd, and any in-cluster client. registry.sh
# handles the first two; station.sh reports when it's configured but not trusted.
# handles the first two; check.sh reports when it's configured but not trusted.
# Symptom when missing: x509: certificate signed by unknown authority
REGISTRY_CA_FILE=

View File

@@ -1,18 +1,18 @@
# The installation wizard. It does NOT run the cluster — it installs a toolchain
# The toolchain installer image. It does NOT run the cluster — it installs a toolchain
# onto the host and gets out of the way.
#
# This exists to kill a bootstrap paradox: a plain bash installer needs curl, jq
# and sha256sum to already be present, and a minimal Debian has none of them.
# The wizard carries its own toolchain, so the only host prerequisite is Docker.
# It carries its own toolchain, so the only host prerequisite is Docker.
#
# Two variants from one file:
# docker build -f ctrl/Dockerfile.wizard --target wizard -t <slug>-wizard .
# docker build -f ctrl/Dockerfile.wizard --target wizard-full -t <slug>-wizard:full .
# docker build -f ctrl/Dockerfile.deps --target deps -t <slug>-deps .
# docker build -f ctrl/Dockerfile.deps --target deps-full -t <slug>-deps:full .
#
# wizard-full bakes every pinned binary in at build time. `docker save` it and
# deps-full bakes every pinned binary in at build time. `docker save` it and
# you have the whole installer as one file to carry into an air-gapped network.
FROM debian:trixie-slim AS wizard
FROM debian:trixie-slim AS deps
# ca-certificates + curl: fetch and verify. graphviz + python3: render diagrams
# and validate the arch model, so the host never needs an apt package.
@@ -24,23 +24,27 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates curl jq graphviz python3 docker-cli \
&& rm -rf /var/lib/apt/lists/*
# The installer is the generated standalone kit, not deps.sh plus the files it
# reads. A kit is one file with its pins frozen in and is proven to run with
# nothing else from rig present — which is exactly what an image needs, and
# `make standalone` keeps it current. Pins are the same in every profile's kit.
ARG PROFILE=default
WORKDIR /work
COPY ctrl/versions.env /work/ctrl/versions.env
COPY ctrl/wizard.sh /work/ctrl/wizard.sh
RUN chmod +x /work/ctrl/wizard.sh
COPY standalone/${PROFILE}/rigdeps.sh /work/rigdeps.sh
RUN chmod +x /work/rigdeps.sh
# Defaults; every one is overridable with -e at run time.
ENV DEPS_SOURCE=upstream \
OUT_BIN=/out/bin \
HOST_ROOT=/host
ENTRYPOINT ["/work/ctrl/wizard.sh"]
ENTRYPOINT ["/work/rigdeps.sh"]
CMD ["install"]
# ---------------------------------------------------------------------------
# wizard-full — same wizard, binaries baked in, works with no network at all.
FROM wizard AS wizard-full
RUN /work/ctrl/wizard.sh fetch --to /opt/rig/bin
# deps-full — same image, binaries baked in, works with no network at all.
FROM deps AS deps-full
RUN /work/rigdeps.sh fetch --to /opt/rig/bin
ENV DEPS_SOURCE=baked \
BAKED_BIN=/opt/rig/bin

View File

@@ -0,0 +1,50 @@
# EXAMPLE — a component image. Copy, rename, replace.
#
# Named like the manifest it feeds and the resource it becomes:
#
# ctrl/Dockerfile.api -> image <cluster>-api -> image: in k8s/base/api.yaml
#
# That image string is the ONLY thing connecting the three. Nothing checks it;
# a typo shows up as a pod stuck in ImagePullBackOff pulling from the public
# index, which reads like a network problem and is not one.
#
# ── the one that catches everyone ──────────────────────────────────────────
# The Tiltfile passes two paths with DIFFERENT bases, in adjacent arguments:
#
# context='..' the REPO ROOT (the Tiltfile is in ctrl/)
# dockerfile='Dockerfile.api' relative to the TILTFILE, so ctrl/Dockerfile.api
#
# So every COPY below is resolved against the repo root, NOT against this file's
# directory. A file sitting right beside this one is still reached as `ctrl/`:
#
# COPY ctrl/nginx.conf /etc/nginx/conf.d/default.conf # correct
# COPY nginx.conf /etc/nginx/conf.d/default.conf # fails — no such file
#
# Nothing warns you. The build just cannot find a file that is visibly there.
FROM python:3.12-slim
WORKDIR /app
# Dependencies first, in their own layer: they change far less often than the
# code, so a source edit does not reinstall them on every rebuild.
COPY api/requirements.txt ./
RUN pip install --no-cache-dir -r requirements.txt
# Repo-root relative — see above.
COPY api/ ./api/
# Match this with the containerPort in the manifest and the target of the
# Service in front of it.
EXPOSE 8000
CMD ["python", "-m", "api"]
# ── live_update ────────────────────────────────────────────────────────────
# The sync in the Tiltfile's docker_build must land where this image expects it:
#
# live_update=[sync('../api', '/app/api')]
#
# matches `COPY api/ ./api/` with WORKDIR /app. If the two disagree, Tilt syncs
# into a path nothing reads and the container keeps serving the built copy —
# edits appear to do nothing, with no error anywhere.

139
rig/ctrl/Tiltfile Normal file
View File

@@ -0,0 +1,139 @@
# The dev loop. `make tilt` from the project root, or `cd ctrl && tilt up`.
#
# This file ships with rig and works unedited: rig's own k8s/base already boots,
# so `make tilt` comes up with a running cluster and no editing at all. What it
# deploys is two EXAMPLES — replace them, and add your own images and resources
# in the two marked sections near the bottom. The catalogue after them has the
# blocks to paste, with the parts that are easy to get wrong already commented.
#
# rig supplies this file; it does not own it. Nothing in rig reads it back, and
# nothing here is regenerated — edit it freely, the way you would edit
# k8s/base/example-mock.yaml. rig owns the machine, you own the workload.
#
# Nothing below is hardcoded to this directory, deliberately. Every other
# project here writes its slug into the Tiltfile five or six times by hand, so a
# copy of the project deploys into the original's cluster until someone
# remembers to edit all of them. A rig is meant to be copied and renamed, so it
# asks instead.
# ── who we are, and on which ports ─────────────────────────────────────────
# One question to rig, answered by ctrl/ports.sh, which resolves it through
# lib/config.sh — the same path every other rig script takes. That is the point:
# the cluster name is NOT the bare directory name (it is lowercased and reduced
# to a DNS label), and the ports honour anything pinned in ctrl/.env. Recomputing
# either of those here in Starlark is how two copies end up disagreeing about
# which cluster they are talking to.
_facts = str(local('bash ports.sh active', quiet=True)).split()
CLUSTER = _facts[0]
CTX = _facts[1]
HTTP = _facts[2]
HTTPS = _facts[3]
TILT = _facts[4]
REGISTRY = _facts[5]
# Where the manifests live. rig's own are the default; point MANIFESTS_DIR in
# ctrl/.env at an overlay versioned somewhere else and rig stops owning them —
# see k8s/README.md. Real manifests usually change on a different cadence, by
# different people, under different review.
#
# The value is REPO-ROOT relative, because that is the root everything else in
# rig is expressed against. This file runs in ctrl/, so prefix rather than
# assume: '../' + 'ctrl/k8s/overlays/dev' and '../' + '../platform/overlays/dev'
# are both right, where stripping a leading 'ctrl/' would only fix the first.
MANIFESTS = '../' + _facts[6]
# ── refuse to deploy into the wrong cluster ────────────────────────────────
# Tilt snapshots the kubectl context at startup, BEFORE parsing this file, so it
# cannot be switched from here — only refused. `make tilt` passes --context for
# you; this catches a bare `tilt up` after some other project moved the global
# context.
allow_k8s_contexts(CTX)
if k8s_context() != CTX:
fail("Wrong kubectl context: '%s'. This is %s — run: make tilt, or tilt up --context %s"
% (k8s_context(), CLUSTER, CTX))
# The namespace has to exist before anything lands in it, and kustomize does not
# guarantee ordering across resources. Creating it here is idempotent.
local('kubectl --context %s create namespace %s --dry-run=client -o yaml | kubectl --context %s apply -f -'
% (CTX, CLUSTER, CTX), quiet=True)
# ── images go to this environment's own registry ───────────────────────────
# Fail closed. Tilt can usually infer the kind registry on its own, but "usually"
# is an inference, and when it misses, an unqualified name like 'app' quietly
# means docker.io/library/app — a push to the public index instead of the
# registry two lines away. rig runs that registry; name it.
default_registry('localhost:' + REGISTRY)
k8s_yaml(kustomize(MANIFESTS))
# ── Images ─────────────────────────────────────────────────────────────────
# (nothing yet — rig's examples run upstream images. Add docker_build calls here.)
# ── Resources ──────────────────────────────────────────────────────────────
# (nothing yet — add k8s_resource calls here to name and order what you deploy.)
# Everything with no dev loop of its own, gathered so it does not clutter the UI.
k8s_resource(
objects=[CLUSTER + ':namespace'],
new_name='infra',
)
# ═══════════════════════════════════════════════════════════════════════════
# Catalogue — paste what you need, delete the rest.
#
# These are the shapes that recur across every project here, with the reasoning
# kept next to them. They are comments so this file runs as-is.
# ═══════════════════════════════════════════════════════════════════════════
#
# ── build an image ─────────────────────────────────────────────────────────
# The one genuinely non-obvious thing in the whole corpus: `context` and
# `dockerfile` are relative to DIFFERENT directories, in adjacent arguments,
# and nothing warns you.
#
# context= the REPO ROOT — this file is in ctrl/, so '..'
# dockerfile= relative to THIS file — so 'Dockerfile.api' is ctrl/Dockerfile.api
#
# Every COPY inside those Dockerfiles is therefore repo-root relative: a file
# sitting BESIDE the Dockerfile is still reached as `COPY ctrl/nginx.conf`.
#
# docker_build(
# CLUSTER + '-api', # must match `image:` in the manifest —
# context='..', # that string is the only thing
# dockerfile='Dockerfile.api', # connecting the two
# ignore=['.git', 'def', '.venv', 'node_modules', '__pycache__'],
# live_update=[sync('../api', '/app/api')],
# )
#
# ── name and order a resource ──────────────────────────────────────────────
# k8s_resource('api', resource_deps=['postgres'], labels=['app'])
# k8s_resource('gateway', resource_deps=['api', 'ui'], labels=['app'])
#
# ── reload the gateway when its config changes ─────────────────────────────
# A Caddyfile arriving via configMapGenerator with disableNameSuffixHash does
# NOT roll the pod — the ConfigMap name never changes, so nothing tells the
# Deployment anything happened. Without this you edit the routes and watch
# nothing take effect.
#
# local_resource(
# 'gateway-reload',
# cmd='kubectl --context %s -n %s rollout restart deployment/gateway' % (CTX, CLUSTER),
# deps=['k8s/base/Caddyfile'],
# resource_deps=['gateway'],
# auto_init=False,
# )
#
# ── an overlay whose secretGenerator reads outside its own directory ───────
# kustomize refuses to read above the kustomization root unless told to. Only
# add this if you actually have such a generator; it loosens a safety check.
#
# k8s_yaml(kustomize(MANIFESTS, flags=['--load-restrictor=LoadRestrictionsNone']))
#
# ── reach a service directly, bypassing the gateway ────────────────────────
# For a DB client or an admin UI. Prefer routing through the gateway: host ports
# are a single shared namespace across every project on this machine, which is
# why rig derives a block per environment in the first place. If you do need
# one, take it from this environment's own block rather than picking a number.
#
# k8s_resource('postgres', port_forwards=[str(int(HTTP) + 5) + ':5432'])

View File

@@ -1,5 +1,5 @@
#!/usr/bin/env bash
# Apache Airflow — the cluster half of soleprint's airflow cabinet.
# Apache Airflow — the cluster half of the airflow cabinet.
#
# Airflow needs a metadata database before it will start at all, so this refuses
# rather than rolls a pod that will CrashLoopBackOff while the real problem
@@ -7,7 +7,7 @@
#
# One pod on `standalone`, matching the compose cabinet: migration, admin user,
# scheduler and webserver in a single container. The official chart's five
# deployments model an installation; a room switching this on wants pipelines.
# deployments model an installation; switching this on means wanting pipelines.
set -euo pipefail
cd "$(dirname "$0")/.."

View File

@@ -1,13 +1,14 @@
#!/usr/bin/env bash
# PostgreSQL — the cluster half of soleprint's postgres cabinet.
# PostgreSQL — the cluster half of the postgres cabinet.
#
# A room declares the dependency once, in cfg/<room>/data/cabinets.json. On a
# laptop `build.py` composes it into docker-compose.yml; here it becomes a pod,
# so the same declaration works either way and nothing has to be remembered
# A cabinet is a public service dropped into the environment as-is — the
# upstream image, unmodified, reachable at a known address. This is the cluster
# half of it; the compose half is a `service.yml` beside a `cabinet.json`. The
# declaration is made once and both paths read it, so nothing is remembered
# twice.
#
# Plain manifests rather than a helm chart, matching the other addons: a chart
# repo is a network dependency, and the offline profile exists precisely so
# repo is a network dependency, and the offline example profile exists precisely so
# there is a path with none. The image is pinned in ctrl/versions.env and can be
# preloaded into a local registry like every other image here.
#
@@ -32,8 +33,8 @@ if $K get secret -n "$NS" postgres >/dev/null 2>&1; then
else
password=$(head -c 18 /dev/urandom | base64 | tr -d '/+=' | head -c 24)
$K create secret generic postgres -n "$NS" \
--from-literal=POSTGRES_DB="${POSTGRES_DB:-soleprint}" \
--from-literal=POSTGRES_USER="${POSTGRES_USER:-soleprint}" \
--from-literal=POSTGRES_DB="${POSTGRES_DB:-postgres}" \
--from-literal=POSTGRES_USER="${POSTGRES_USER:-postgres}" \
--from-literal=POSTGRES_PASSWORD="$password" >/dev/null
echo " generated a password (read it back with the command printed below)"
fi

View File

@@ -1,5 +1,5 @@
#!/usr/bin/env bash
# Redis — the cluster half of soleprint's redis cabinet.
# Redis — the cluster half of the redis cabinet.
#
# Cache, and the broker anything queue-shaped runs on. No persistence: a broker
# that loses its queue on restart is the honest local model, and a PVC here buys

218
rig/ctrl/check.sh Executable file
View File

@@ -0,0 +1,218 @@
#!/usr/bin/env bash
# Readiness check: is this machine ready to run rig?
#
# Reports and instructs; never silently fixes anything. Everything it finds is
# either already fine, or something a human has to decide on.
#
# Runs ctrl/deps.sh host detection in a container when Docker is the only thing
# installed, or directly when the toolchain is already present. Then adds the
# checks that need this repo's config: profile sanity, CA trust, port clashes.
set -euo pipefail
cd "$(dirname "$0")"
DEPS_IMAGE="${DEPS_IMAGE:-$(basename "$(cd .. && pwd)")-deps}"
# Host detection. Prefer running it bare — it needs no dependencies beyond
# coreutils — and fall back to the container only if this shell can't.
bash ./deps.sh detect
# ── repo-level checks ──────────────────────────────────────────────────────
source ./lib/config.sh
load_config
echo
echo "config"
echo " profile ${PROFILE_NAME} (nodes=${NODES} audit=${AUDIT})"
echo " cluster ${CLUSTER} (context ${KUBECONTEXT})"
echo " registry ${REGISTRY_MODE}"
echo " ingress ${INGRESS_MODE}"
if [ ! -f ./.env ]; then
echo " ! ctrl/.env missing — copy it: cp ctrl/.env.example ctrl/.env"
fi
# ── memory ─────────────────────────────────────────────────────────────────
#
# A profile on a box that is already full is the most common first failure, and
# it presents as pods stuck Pending rather than anything that says "memory".
# Warns; never blocks. Whether to try anyway is the user's call.
# A /proc/meminfo field in MB, 0 if absent. MEMINFO and OVERCOMMIT_FILE exist
# only so the tight and does-not-fit branches can be exercised against another
# machine's real numbers; in normal use they are the kernel's own files.
mb_of() {
awk -v k="$1:" '$1 == k { printf "%d", $2 / 1024; found = 1 }
END { if (!found) printf "0" }' "${MEMINFO:-/proc/meminfo}"
}
# NODE_MB — what one node costs — comes from load_config (lib/config.sh), where
# its measurement is recorded. It lives there, not here, because the memory tool
# and every standalone kit need the same number: a copy of it is how rigmini.sh
# came to say 2 GB per node long after rig had measured 800 MB.
# Every running container's working set in MB, tagged with the kind cluster it
# belongs to ('-' when it is not kind). docker stats reports usage minus page
# cache, which is what actually competes — cache is handed back under pressure.
# Counting only kind would hide the usual culprit on a managed workspace, where
# the memory is held by other containers entirely.
container_mb() {
docker info >/dev/null 2>&1 || return 0
awk -F'\t' '
FILENAME == ARGV[1] { cl[$1] = ($2 == "" ? "-" : $2); if ($2 != "") isc[$2] = 1; next }
{
grp = ($1 in cl ? cl[$1] : "-")
# A kind cluster'"'"'s local registry is a plain container with no kind
# label, named <cluster>-registry, so on its own it would read as a
# stranger. It belongs to its cluster — but only if that cluster exists:
# a registry whose cluster is gone is a genuine stray, and says so.
if (grp == "-" && $1 ~ /-registry$/) {
base = $1; sub(/-registry$/, "", base)
if (base in isc) grp = base
}
split($2, u, " "); v = u[1]; mb = 0
if (v ~ /GiB$/) { sub(/GiB$/, "", v); mb = v * 1024 }
else if (v ~ /MiB$/) { sub(/MiB$/, "", v); mb = v }
else if (v ~ /KiB$/) { sub(/KiB$/, "", v); mb = v / 1024 }
else if (v ~ /B$/) { sub(/B$/, "", v); mb = v / 1048576 }
printf "%d\t%s\t%s\n", mb, grp, $1
}
' <(docker ps --format '{{.Names}}\t{{.Label "io.x-k8s.kind.cluster"}}' 2>/dev/null) \
<(docker stats --no-stream --format '{{.Name}}\t{{.MemUsage}}' 2>/dev/null)
}
total_mb=$(mb_of MemTotal)
avail_mb=$(mb_of MemAvailable)
swap_used_mb=$(( $(mb_of SwapTotal) - $(mb_of SwapFree) ))
overcommit=$(cat "${OVERCOMMIT_FILE:-/proc/sys/vm/overcommit_memory}" 2>/dev/null || echo '?')
need_mb=$(( NODES * NODE_MB ))
rows=$(container_mb)
# Once this environment's own cluster is running, its real footprint is already
# out of MemAvailable and the per-node estimate stops being relevant. Subtracting
# the measurement from the estimate would count the same memory twice, and a
# running cluster that happens to sit under 800 MB would still "need" the gap.
ours_mb=$(awk -F'\t' -v c="$CLUSTER" '$2 == c { s += $1 } END { print s + 0 }' <<< "$rows")
still_mb=$(( ours_mb > 0 ? 0 : need_mb ))
echo
echo "memory"
printf " this profile ~%d MB %s node(s) x %d MB — the cluster alone, your workload on top\n" \
"$need_mb" "$NODES" "$NODE_MB"
if [ "$ours_mb" -gt 0 ]; then
printf " already held %d MB by '%s', which is up\n" "$ours_mb" "$CLUSTER"
fi
printf " available %d MB of %d MB\n" "$avail_mb" "$total_mb"
# The biggest things holding memory right now, other than this cluster: kind
# clusters summed per cluster, everything else by container name.
others=$(awk -F'\t' -v c="$CLUSTER" '
$2 != c && $2 != "-" && $2 != "" { k["kind cluster \x27" $2 "\x27"] += $1 }
$2 == "-" { k["container \x27" $3 "\x27"] += $1 }
END { for (n in k) printf "%d\t%s\n", k[n], n }' <<< "$rows" | sort -rn)
if [ -n "$others" ]; then
echo " held elsewhere:"
head -6 <<< "$others" | awk -F'\t' '{ printf " %6d MB %s\n", $1, $2 }'
n_others=$(wc -l <<< "$others")
if [ "$n_others" -gt 6 ]; then
echo " ... and $((n_others - 6)) more"
fi
fi
headroom=$(( avail_mb - still_mb ))
if [ "$still_mb" -eq 0 ]; then
if [ "$headroom" -ge 512 ]; then
printf " fits — already up; %d MB headroom for what you deploy\n" "$headroom"
else
printf " ! already up, but only %d MB headroom for anything you deploy\n" "$headroom"
fi
elif [ "$headroom" -ge 512 ]; then
printf " fits — %d MB headroom for what you deploy\n" "$headroom"
elif [ "$headroom" -ge 0 ]; then
printf " ! fits, but only %d MB headroom for anything you deploy\n" "$headroom"
else
printf " ! does not fit right now: ~%d MB needed, %d MB available\n" "$still_mb" "$avail_mb"
# Two failures with opposite fixes, and telling them apart is the point.
if [ "$still_mb" -le "$total_mb" ]; then
echo " The machine is big enough; something else is holding memory (above)."
echo " Stopping that is what helps — a bigger VM would not."
if grep -q 'kind cluster' <<< "$others"; then
echo " 'make cluster free' stops the other kind clusters. It stops, never deletes."
fi
else
echo " The machine itself is too small: ~${still_mb} MB needed, ${total_mb} MB total."
fi
fi
if [ "$swap_used_mb" -gt 0 ]; then
printf " ! %d MB already in swap — available memory does not count it, so expect a\n" "$swap_used_mb"
echo " cluster here to be slow well before it fails"
fi
if [ "$overcommit" = "1" ]; then
echo " ! overcommit=1: allocations never fail here, so read 'fits' as a ceiling."
echo " A cluster that starts cleanly can still lose processes to the OOM killer."
fi
# The CA reaches three places and only one of them is ours. Report the other two.
if [ -n "${REGISTRY_CA_FILE:-}" ]; then
echo
echo "registry CA"
if [ ! -r "$REGISTRY_CA_FILE" ]; then
echo " ! REGISTRY_CA_FILE not readable: $REGISTRY_CA_FILE"
else
echo " file $REGISTRY_CA_FILE"
host="${REGISTRY_REMOTE_URL#*://}"; host="${host%%/*}"
if [ -n "$host" ] && [ ! -f "/etc/docker/certs.d/${host}/ca.crt" ]; then
echo " ! the HOST docker daemon does not trust it yet:"
echo " sudo mkdir -p /etc/docker/certs.d/${host}"
echo " sudo cp ${REGISTRY_CA_FILE} /etc/docker/certs.d/${host}/ca.crt"
echo " (kind nodes are handled by registry.sh; in-cluster clients are the workload's job)"
fi
fi
fi
# Host ports this environment will try to bind. Checked before cluster creation
# because docker reports a clash halfway through, as an opaque
# "failed to bind host port ...: address already in use".
echo
echo "ports (block derived from the directory name — see 'make ports')"
port_busy() {
if command -v ss >/dev/null 2>&1; then
ss -ltn "sport = :$1" 2>/dev/null | grep -q LISTEN && return 0 || return 1
fi
# iproute2 is absent from a minimal Debian, so fall back to procfs rather
# than silently reporting everything as free.
local hex; hex=$(printf ':%04X' "$1")
grep -qi "^ *[0-9]*: [0-9A-F]*$hex " /proc/net/tcp /proc/net/tcp6 2>/dev/null
}
# A port held by THIS environment's own cluster is not a clash — it is the thing
# working. Reporting it as a problem every time the cluster is up would train
# people to ignore this section, which is the opposite of the point.
# Extract with a second grep rather than `tr -d ':->'`: in tr, ':->' is the
# character RANGE ':' to '>', which does not contain '-', so the trailing dash
# survives and nothing ever matches.
ours=$(docker ps --filter "label=io.x-k8s.kind.cluster=${CLUSTER}" \
--format '{{.Ports}}' 2>/dev/null | tr ',' '\n' \
| grep -oE ':[0-9]+->' | grep -oE '[0-9]+' || true)
clash=0
for entry in "HTTP:${HTTP_PORT}" "HTTPS:${HTTPS_PORT}" \
"TILT:${TILT_PORT}" "REGISTRY:${REGISTRY_PORT}"; do
name="${entry%%:*}"; p="${entry#*:}"
[ -n "$p" ] || continue
if ! port_busy "$p"; then
printf " %-9s %-6s free\n" "$name" "$p"
elif echo "$ours" | grep -qx "$p"; then
printf " %-9s %-6s in use by this environment's cluster\n" "$name" "$p"
else
printf " ! %-9s %-6s IN USE by something else\n" "$name" "$p"
clash=1
fi
done
if [ "$clash" -eq 1 ]; then
echo " override the clashing one in ctrl/.env, e.g. HTTP_PORT=21080"
echo " (or rename this directory — the whole block follows the name)"
fi

View File

@@ -25,7 +25,7 @@ up() {
# Say what this profile locks in BEFORE spending minutes building it:
# the audit policy is an apiserver flag and cannot be changed later.
echo "creating cluster '$CLUSTER' from profile '$PROFILE_NAME'"
echo " shape ctrl/k8s/$KIND_CONFIG"
echo " shape ${KIND_CONFIG_SHOWN}"
echo " nodes $NODES"
echo " image $NODE_IMAGE"
echo " audit $AUDIT"

858
rig/ctrl/deps.sh Executable file
View File

@@ -0,0 +1,858 @@
#!/usr/bin/env bash
# rig:standalone rigdeps detect
# Toolchain installer: detect the host, install a pinned toolchain onto it, then
# report what it could not do.
#
# It never runs the cluster, never uses sudo or apt, and writes only into
# $OUT_BIN (default ~/.local/bin). Everything that would touch the host proper —
# systemd, inotify limits, .wslconfig, docker group — is REPORTED for a human to
# decide on, never performed. That is what makes it safe to run on a machine that
# already has a working setup.
#
# Usage (normally via `make deps`, or directly):
# deps.sh detect # report host facts only, change nothing
# deps.sh list # the pinned versions
# deps.sh verify [core|dev] # run what is installed and see if it works
# deps.sh fetch [core|dev] [--to DIR] # download + verify into DIR
# deps.sh install [core|dev] # detect, fetch, install, report
#
# Tiers: 'core' is kubectl + jq (talk to a cluster); 'dev' adds kind and tilt
# Default is dev.
#
# Runs both inside the installer container and bare on a host. Inside the
# container, host files are read through $HOST_ROOT (mount / as :ro); bare, it
# falls back to /.
set -euo pipefail
# Keep the caller's cwd so a relative --to resolves where the user expects,
# not against ctrl/ once we've moved.
INVOKED_FROM="$PWD"
cd "$(dirname "$0")"
# Pins arrive through load_config like every other setting, not by sourcing
# versions.env here. That is what lets `make standalone` freeze them into a
# one-file installer: configuration has exactly one way in.
source ./lib/config.sh
load_config
# Resolve a possibly-relative path against the caller's original directory.
abspath() {
case "$1" in
/*) echo "$1" ;;
*) echo "$INVOKED_FROM/$1" ;;
esac
}
OUT_BIN="${OUT_BIN:-$HOME/.local/bin}"
HOST_ROOT="${HOST_ROOT:-/}"
DEPS_SOURCE="${DEPS_SOURCE:-upstream}"
DEPS_ARTIFACTORY_URL="${DEPS_ARTIFACTORY_URL:-}"
BAKED_BIN="${BAKED_BIN:-/opt/rig/bin}"
# Collected by detect(), printed by report_manual() at the very end.
MANUAL=()
# Host FILES (/etc/..., /mnt/c/...) must be read through the mount. Kernel-level
# facts (kernel version, meminfo, inotify) are shared with the container, so the
# container's own view is already the host's.
# A /proc/meminfo field in MB, 0 if the field is absent. MEMINFO exists so the
# tight and does-not-fit branches can be exercised against a real machine's
# numbers from somewhere else; in normal use it is always /proc/meminfo.
mb_of() {
awk -v k="$1:" '$1 == k { printf "%d", $2 / 1024; found = 1 }
END { if (!found) printf "0" }' "${MEMINFO:-/proc/meminfo}"
}
host_file() {
local p="${1#/}"
if [ "$HOST_ROOT" != "/" ] && [ -e "$HOST_ROOT/$p" ]; then
echo "$HOST_ROOT/$p"
else
echo "/$p"
fi
}
# ── the tools this script itself needs ─────────────────────────────────────
arch() {
case "$(uname -m)" in
x86_64|amd64) echo amd64 ;;
aarch64|arm64) echo arm64 ;;
*) uname -m ;;
esac
}
# The pins above are amd64. Rather than download something that cannot execute
# and let it fail as "cannot execute binary file: Exec format error", say so
# here and hand over the commands that produce the right checksums.
require_amd64() {
local a; a=$(arch)
[ "$a" = "amd64" ] && return 0
cat >&2 <<EOF
This machine is ${a} ($(uname -m)); every pin in this script is linux/amd64.
Nothing here would run, so it does not download. To make an ${a} version, the
URLs need the ${a} artifact and the checksums need to come from each project's
own published list — not from these values, and not from a download you did:
curl -sSL https://github.com/kubernetes-sigs/kind/releases/download/${KIND_VERSION}/checksums.txt
curl -sSL https://dl.k8s.io/release/${KUBECTL_VERSION}/bin/linux/${a}/kubectl.sha256
curl -sSL https://github.com/tilt-dev/tilt/releases/download/v${TILT_VERSION}/checksums.txt
curl -sSL https://github.com/tilt-dev/ctlptl/releases/download/v${CTLPTL_VERSION}/checksums.txt
curl -sSL https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/sha256sum.txt
Edit the pinned block at the top of this file with what those print.
EOF
exit 1
}
DL=""
pick_downloader() {
if command -v curl >/dev/null 2>&1; then DL=curl
elif command -v wget >/dev/null 2>&1; then DL=wget
else
echo "neither curl nor wget is installed, so nothing can be downloaded." >&2
echo "Install one first: $(pkg_install_cmd curl)" >&2
exit 1
fi
}
download() {
local url="$1" out="$2"
case "$DL" in
curl) curl -fsSL --retry 3 -o "$out" "$url" ;;
wget) wget -q --tries=3 -O "$out" "$url" ;;
esac
}
SHA=""
pick_sha() {
if command -v sha256sum >/dev/null 2>&1; then SHA=sha256sum
elif command -v shasum >/dev/null 2>&1; then SHA="shasum -a 256"
else
echo "no sha256sum and no shasum — downloads could not be verified." >&2
echo "Refusing to install unverified binaries." >&2
exit 1
fi
}
# ── package manager, for the instructions only ─────────────────────────────
# This never runs a package manager. It names one so the reported action is
# something you can paste, on the distro you are actually on — an apt line on
# Amazon Linux 2 is a wrong answer dressed up as help.
pkg_install_cmd() {
local pkg="$1"
if command -v apt-get >/dev/null 2>&1; then echo "sudo apt-get update && sudo apt-get install -y $pkg"
elif command -v dnf >/dev/null 2>&1; then echo "sudo dnf install -y $pkg"
elif command -v yum >/dev/null 2>&1; then echo "sudo yum install -y $pkg"
elif command -v zypper >/dev/null 2>&1; then echo "sudo zypper install -y $pkg"
elif command -v apk >/dev/null 2>&1; then echo "sudo apk add $pkg"
else echo "install '$pkg' with this system's package manager"
fi
}
docker_pkg() {
# Debian and Ubuntu call it docker.io; the RPM distros call it docker.
if command -v apt-get >/dev/null 2>&1; then echo docker.io; else echo docker; fi
}
# ── detect ─────────────────────────────────────────────────────────────────
# Windows outside WSL — Git Bash, MSYS, Cygwin — looks close enough to work and
# then fails in a pile of confusing ways: no /proc, no docker socket, none of
# the tooling. Detectable, so name it instead.
require_linux() {
case "$(uname -s)" in
MINGW*|MSYS*|CYGWIN*)
cat >&2 <<'EOF'
This has to run inside WSL, not Git Bash / MSYS / Cygwin.
If WSL is not installed yet, from an elevated PowerShell or Command Prompt:
wsl --install
That enables Windows features and needs a reboot, so it is not something this
script will do for you. Afterwards, open the Linux shell it installs and run
this from there.
See "Starting from plain Windows" in README.md.
EOF
exit 1 ;;
esac
}
is_wsl() { grep -qi microsoft /proc/version 2>/dev/null; }
detect() {
echo "host"
echo " kernel $(uname -r)"
echo " arch $(arch) ($(uname -m))"
local osr; osr=$(host_file /etc/os-release)
[ -r "$osr" ] && echo " distro $(sed -n 's/^PRETTY_NAME="\(.*\)"/\1/p' "$osr")"
# In MB. Whole gigabytes lose nearly half a GB on exactly the machines where
# it matters: 1874 MB available used to print as "1 GB". Facts only — whether
# that is enough depends on the profile, which check.sh knows and this does not.
local total_mb avail_mb swap_total_mb swap_used_mb om
total_mb=$(mb_of MemTotal)
avail_mb=$(mb_of MemAvailable)
swap_total_mb=$(mb_of SwapTotal)
swap_used_mb=$(( swap_total_mb - $(mb_of SwapFree) ))
printf " memory %d MB total, %d MB available\n" "$total_mb" "$avail_mb"
if [ "$swap_total_mb" -gt 0 ]; then
printf " swap %d MB used of %d MB\n" "$swap_used_mb" "$swap_total_mb"
fi
# How the kernel answers an allocation it cannot really satisfy. With 1 it
# always says yes and settles up later with the OOM killer, so a cluster that
# starts cleanly can still lose processes afterwards.
om=$(cat "${OVERCOMMIT_FILE:-/proc/sys/vm/overcommit_memory}" 2>/dev/null || echo '?')
case "$om" in
0) echo " overcommit 0 heuristic — allocations are granted on a guess" ;;
1) echo " overcommit 1 always — every allocation succeeds; the OOM killer is the only limit" ;;
2) echo " overcommit 2 strict — an allocation fails honestly instead of killing later" ;;
esac
echo " install to $OUT_BIN"
detect_libc
detect_prereqs
detect_wsl
detect_filesystem
detect_docker
detect_inotify
detect_toolchain
}
detect_wsl() {
if ! is_wsl; then
echo " platform native linux"
return
fi
echo " platform WSL"
# systemd is off by default in WSL, and the ingress/DNS paths that use a
# host service need it. Enabling it requires a Windows-side restart, which
# cannot be issued from inside the distro.
local wc; wc=$(host_file /etc/wsl.conf)
if [ -r "$wc" ] && grep -qE '^\s*systemd\s*=\s*true' "$wc"; then
echo " systemd enabled in wsl.conf"
else
echo " ! systemd not enabled in /etc/wsl.conf"
MANUAL+=("Enable systemd — add to /etc/wsl.conf:
[boot]
systemd=true
then from a WINDOWS terminal (not this shell): wsl --shutdown")
fi
# WSL regenerates /etc/resolv.conf on every boot, which silently reverts any
# local DNS setup.
if [ -r "$wc" ] && grep -qE '^\s*generateResolvConf\s*=\s*false' "$wc"; then
echo " resolv.conf pinned (generateResolvConf=false)"
else
echo " - resolv.conf is WSL-generated; DNS_MODE=dnsmasq would be reverted on reboot"
fi
local wcfg
wcfg=$(ls "$HOST_ROOT"/mnt/c/Users/*/.wslconfig 2>/dev/null | head -1 || true)
if [ -n "$wcfg" ] && grep -qE '^\s*memory\s*=' "$wcfg"; then
echo " wslconfig memory set: $(grep -E '^\s*memory\s*=' "$wcfg" | tr -d ' ')"
else
MANUAL+=("Cap/raise the WSL VM memory — see what is set versus what booted:
make mem status
It prints the edit to make and the command to apply it.")
fi
}
# Not a path check: /mnt is an ordinary mount point and an ext4 disk mounted
# there is perfectly fine. What matters is the filesystem. The Windows drives
# arrive as 9p (WSL2) or drvfs (WSL1); network and fuse mounts behave the same
# way. None of them deliver inotify events, so anything watching files goes
# quiet without saying why.
watch_hostile_fs() {
local dir="$1" fstype
fstype=$(findmnt -no FSTYPE --target "$dir" 2>/dev/null || true)
[ -n "$fstype" ] || fstype=$(stat -f -c %T "$dir" 2>/dev/null || true)
case "$fstype" in
9p|v9fs|drvfs|cifs|smb3|nfs|nfs4|fuse.sshfs|fuseblk) echo "$fstype" ;;
*) echo "" ;;
esac
}
detect_filesystem() {
local root fstype
root=$(cd .. && pwd -P)
fstype=$(watch_hostile_fs "$root")
if [ -n "$fstype" ]; then
echo " ! this directory is on $fstype — file watching will not work"
MANUAL+=("Move this onto the local disk. Nothing watching files sees changes
on a $fstype mount, and everything else is slower:
cp -r \"$root\" ~/ && cd ~/$(basename "$root")")
else
echo " filesystem $root ($(findmnt -no FSTYPE --target "$root" 2>/dev/null || echo local))"
fi
}
# tilt is the one binary here that needs a recent glibc. MEASURED, not guessed:
# tilt 0.37.6 on Amazon Linux 2 (glibc 2.26) fails with
#
# /lib64/libc.so.6: version `GLIBC_2.34' not found (required by .../tilt)
#
# which names a symbol rather than the problem. Amazon Linux 2 is a stock
# WorkSpaces bundle, so this is the likely case, not an exotic one. Report the
# version now; `verify` catches the actual failure after installing.
detect_libc() {
local v=""
if command -v ldd >/dev/null 2>&1; then
v=$(ldd --version 2>/dev/null | head -1 | grep -oE '[0-9]+\.[0-9]+$' || true)
fi
if [ -z "$v" ]; then
echo " libc unknown (no ldd) — 'verify' is the real test"
return 0
fi
echo " libc glibc $v"
if [ "$(printf '%s\n2.34\n' "$v" | sort -V | head -1)" != "2.34" ]; then
echo " ! older than glibc 2.34, which tilt needs. kubectl, kind, jq and"
echo " ctlptl are static or libc-only and work here; tilt will not start."
echo " Install the core tier, or run tilt from a container."
fi
return 0
}
# What this script needs to do its own job. Reported here so `detect` answers
# "will install work?" instead of leaving you to find out one download in.
# Amazon Linux 2 ships without tar, which is exactly the surprise this catches.
detect_prereqs() {
local missing=""
if command -v curl >/dev/null 2>&1; then echo " download curl"
elif command -v wget >/dev/null 2>&1; then echo " download wget"
else echo " ! no curl and no wget — nothing can be downloaded"; missing+=" curl"
fi
if command -v sha256sum >/dev/null 2>&1 || command -v shasum >/dev/null 2>&1; then
echo " checksums ok"
else
echo " ! no sha256sum or shasum — downloads could not be verified"
missing+=" coreutils"
fi
if command -v tar >/dev/null 2>&1 && command -v gzip >/dev/null 2>&1; then
echo " archives tar + gzip"
else
echo " ! no tar/gzip — tilt and ctlptl ship as tarballs, so the dev tier"
echo " cannot be unpacked. The core tier is two bare binaries and is fine."
missing+=" tar gzip"
fi
if [ -n "$missing" ]; then
MANUAL+=("Install what this script needs to run at all:
$(pkg_install_cmd "${missing# }")")
fi
return 0
}
detect_docker() {
# Reachability of the daemon is the real question, and the CLI is only how
# we ask it. Note that when this runs inside the installer container, Docker
# necessarily exists on the host — otherwise nothing would be executing —
# so a missing CLI in here is an installer packaging bug, not a host problem.
if ! command -v docker >/dev/null 2>&1; then
if [ -S /var/run/docker.sock ]; then
echo " docker socket present (no cli in this context)"
else
echo " ! docker not found and no socket at /var/run/docker.sock"
MANUAL+=("Install Docker — the one true prerequisite, and the only thing here
that needs root:
$(pkg_install_cmd "$(docker_pkg)")
sudo systemctl enable --now docker
sudo usermod -aG docker \"\$USER\"
then log out and back in, so the new group applies to your shell.")
fi
return
fi
if docker info >/dev/null 2>&1; then
echo " docker $(docker version --format '{{.Server.Version}}' 2>/dev/null)"
local n
n=$(docker ps --filter "label=io.x-k8s.kind.cluster" --format '{{.Names}}' 2>/dev/null | wc -l)
# Must be an `if`, not `[ ] && echo`: as the last statement in this
# function the latter returns 1 when the count is zero, and `set -e`
# then kills the caller. That is the fresh-machine case — no clusters
# yet — so the bug only ever shows up where it does most harm.
if [ "$n" -gt 0 ]; then
echo " - $n kind node container(s) already running; see 'make cluster list'"
fi
else
echo " ! docker cli present but the daemon is unreachable"
MANUAL+=("Start Docker, or add yourself to the docker group:
sudo usermod -aG docker \"\$USER\" # then log out and back in")
fi
}
# kind and Tilt both watch large trees. WSL ships defaults (8192/128) far too low,
# and the failure mode is silent: Tilt simply stops noticing file changes.
detect_inotify() {
local w i
w=$(cat /proc/sys/fs/inotify/max_user_watches 2>/dev/null || echo 0)
i=$(cat /proc/sys/fs/inotify/max_user_instances 2>/dev/null || echo 0)
echo " inotify watches=$w instances=$i"
if [ "$w" -lt 524288 ] || [ "$i" -lt 512 ]; then
echo " ! inotify limits are low — Tilt will silently stop noticing file changes"
MANUAL+=("Raise inotify limits (needs root on the host):
echo -e 'fs.inotify.max_user_watches=524288\\nfs.inotify.max_user_instances=512' \\
| sudo tee /etc/sysctl.d/99-rig.conf
sudo sysctl --system")
fi
}
# ── fetch ──────────────────────────────────────────────────────────────────
# Resolve where a given artifact comes from, honouring DEPS_SOURCE.
resolve_url() {
local upstream="$1"
case "$DEPS_SOURCE" in
upstream) echo "$upstream" ;;
artifactory)
if [ -z "$DEPS_ARTIFACTORY_URL" ]; then
echo "DEPS_SOURCE=artifactory but DEPS_ARTIFACTORY_URL is empty" >&2
exit 1
fi
echo "${DEPS_ARTIFACTORY_URL%/}/$(basename "$upstream")"
;;
*) echo "unsupported DEPS_SOURCE '$DEPS_SOURCE' for a download" >&2; exit 1 ;;
esac
}
verify() {
local file="$1" want="$2" name="$3" got
got=$($SHA "$file" | awk '{print $1}')
if [ "$got" != "$want" ]; then
echo "checksum mismatch for $name" >&2
echo " expected $want" >&2
echo " got $got" >&2
exit 1
fi
}
# fetch_bin <name> <url> <sha256> <dest-dir> — a bare binary
fetch_bin() {
local name="$1" url="$2" sha="$3" dest="$4"
local tmp="$dest/.$name.tmp"
echo " fetching $name"
download "$(resolve_url "$url")" "$tmp"
verify "$tmp" "$sha" "$name"
mv "$tmp" "$dest/$name"
chmod +x "$dest/$name"
}
# fetch_tgz <name> <url> <sha256> <dest-dir> <path-inside-archive> <strip>
# Archive layouts differ — tilt's is flat (the binary at the root, strip=0),
# others nest it a directory down — so the caller says which.
fetch_tgz() {
local name="$1" url="$2" sha="$3" dest="$4" inner="$5" strip="$6"
local tmp="$dest/.$name.tgz"
echo " fetching $name"
download "$(resolve_url "$url")" "$tmp"
verify "$tmp" "$sha" "$name"
# --no-same-owner: extracting as root would otherwise restore the uid/gid
# baked into the archive (some ship as uid 1001), leaving a binary the host
# user does not own.
tar -xzf "$tmp" -C "$dest" --strip-components="$strip" --no-same-owner "$inner"
rm -f "$tmp"
chmod +x "$dest/$name"
}
# The installer runs as root so it can reach the docker socket, which means
# everything it writes into a mounted volume lands root-owned and unusable from
# the host. Hand it back to whoever owns the mount point (the host user created
# that directory before mounting it).
fix_ownership() {
local dir="$1"
[ -d "$dir" ] || return 0
local owner="${HOST_UID:-}:${HOST_GID:-}"
if [ "$owner" = ":" ]; then
owner=$(stat -c '%u:%g' "$dir")
fi
[ "$owner" = "0:0" ] && return 0
chown -R "$owner" "$dir" 2>/dev/null || true
}
# Two tiers, because not every machine should get cluster tooling.
#
# core kubectl, jq — talk to a cluster someone else runs. Nothing that
# creates one. Appropriate on a managed or corporate-issued machine
# where development tools are not wanted by default.
# dev core plus kind and tilt — build clusters and hot-reload into them.
#
# The split exists because "install the toolchain" is not one decision: on a
# managed workspace the right answer is kubectl and nothing else.
CORE_TOOLS="kubectl jq"
# No helm: every addon installs with `kubectl apply -f <url>`, so nothing here
# has ever invoked it. Add it back the day something actually needs a chart.
#
# ctlptl is 'dev' rather than 'core' for the same reason kind is: core is "talk
# to a cluster someone else runs", and ctlptl builds them. It earns its place
# because it is what wires a cluster to a local registry — without one, an
# unqualified image name resolves to docker.io/library/<name> and there is
# nothing structural stopping a push there.
#
# docker-compose is 'dev' for the same reason, and is here because the distro
# docker packages ship the daemon and CLI but frequently not the compose
# plugin — so `docker compose up` fails with "unknown command" on an otherwise
# working Docker, and nothing about that message names the missing piece.
DEV_TOOLS="kind tilt ctlptl docker-compose"
# ── what is already on this machine ───────────────────────────────────────
#
# A tool already on PATH at its pinned version is left where it is. Without
# this, install downloads a second copy into OUT_BIN and then reports the first
# one as shadowed — noise, and wrong, when both are the same version. That is
# the normal state of any machine someone set up by hand, whatever directory
# they happened to choose.
pin_of() {
case "$1" in
kubectl) echo "$KUBECTL_VERSION" ;;
jq) echo "$JQ_VERSION" ;;
kind) echo "$KIND_VERSION" ;;
tilt) echo "$TILT_VERSION" ;;
ctlptl) echo "$CTLPTL_VERSION" ;;
docker-compose) echo "$COMPOSE_VERSION" ;;
esac
}
# The version string a binary reports. Each tool spells the question
# differently, and kubectl has to be told --client or it goes looking for a
# server to ask.
reported_version() {
local tool="$1" path="$2"
case "$tool" in
kubectl) "$path" version --client 2>/dev/null ;;
jq) "$path" --version 2>/dev/null ;;
*) "$path" version 2>/dev/null ;;
esac
}
# Does the binary at PATH report PIN? Matched as a whole version token, so
# 0.37.6 never matches 10.37.60, with the leading v optional either side: kind
# says v0.32.0, jq says jq-1.8.2, and tilt says v0.37.6 against a pin of 0.37.6.
#
# Bash's own regex rather than grep, deliberately. grep is not the same program
# on every machine — some builds reject patterns that others accept — and a
# failed grep inside a count reads exactly like a zero.
version_matches() {
local tool="$1" path="$2" pin="$3" out v re
out=$(reported_version "$tool" "$path") || return 1
v="${pin#v}"
v="${v//./\\.}"
re="(^|[^0-9.])v?${v}([^0-9.]|\$)"
[[ $out =~ $re ]]
}
# DEPS_ONLY narrows a fetch to the tools it names. Unset means the whole tier,
# which is what an explicit `deps.sh fetch` always gets: "download these into
# DIR" must not quietly skip something because this machine happens to have it.
# Only install() sets it, to what detect_toolchain found missing or mismatched.
want() { [ -z "${DEPS_ONLY:-}" ] || [[ " $DEPS_ONLY " == *" $1 "* ]]; }
# Every tool in the tier with its state, probed once and reported once. What
# still needs fetching is left in TOOLCHAIN_NEED for install() to act on.
TOOLCHAIN_NEED=""
detect_toolchain() {
local tier="${TIER:-dev}" b pin path found
TOOLCHAIN_NEED=""
echo
echo "toolchain (pinned, tier '$tier')"
for b in $(tier_tools "$tier"); do
pin=$(pin_of "$b")
path=$(command -v "$b" 2>/dev/null || true)
# compose is the one tool that is normally NOT a binary on PATH. It is a
# docker CLI plugin, so a machine where `docker compose` works perfectly
# has no `docker-compose` to find — and probing only PATH would report it
# missing and re-download a copy that is already there. That is the exact
# noise the version-aware skip exists to prevent, so ask docker instead.
if [ "$b" = docker-compose ] && [ -z "$path" ]; then
if found=$(docker compose version --short 2>/dev/null) && [ -n "$found" ]; then
if [ "${found#v}" = "${pin#v}" ]; then
printf " %-8s %-9s %s\n" "$b" "$pin" "docker cli plugin"
else
printf " ! %-8s wants %s, the docker cli plugin reports '%s'\n" \
"$b" "$pin" "$found"
TOOLCHAIN_NEED+="$b "
fi
continue
fi
fi
if [ -z "$path" ]; then
printf " - %-8s %-9s not found\n" "$b" "$pin"
TOOLCHAIN_NEED+="$b "
elif version_matches "$b" "$path" "$pin"; then
printf " %-8s %-9s %s\n" "$b" "$pin" "$path"
else
found=$(reported_version "$b" "$path" 2>/dev/null | head -1 || true)
printf " ! %-8s wants %s, %s reports '%s'\n" "$b" "$pin" "$path" "$found"
TOOLCHAIN_NEED+="$b "
fi
done
if [ -z "$TOOLCHAIN_NEED" ]; then
echo " every pinned tool is already on PATH — nothing to fetch"
else
echo " 'make deps' fetches only: ${TOOLCHAIN_NEED% }"
fi
}
fetch() {
local dest="$OUT_BIN" tier="${TIER:-dev}"
while [ $# -gt 0 ]; do
case "$1" in
--to) dest="$2"; shift 2 ;;
core|dev) tier="$1"; shift ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
dest="$(abspath "$dest")"
mkdir -p "$dest"
TIER="$tier"
if [ "$DEPS_SOURCE" = "baked" ]; then
echo "installing baked binaries from $BAKED_BIN"
cp -a "$BAKED_BIN"/. "$dest"/
fix_ownership "$dest"
return
fi
if [ -n "${DEPS_ONLY:-}" ]; then
echo "fetching ${DEPS_ONLY% } (source: $DEPS_SOURCE)"
else
echo "fetching '$tier' toolchain (source: $DEPS_SOURCE)"
fi
if want kubectl; then fetch_bin kubectl "$KUBECTL_URL" "$KUBECTL_SHA256" "$dest"; fi
if want jq; then fetch_bin jq "$JQ_URL" "$JQ_SHA256" "$dest"; fi
if [ "$tier" = "dev" ]; then
if want kind; then fetch_bin kind "$KIND_URL" "$KIND_SHA256" "$dest"; fi
if want tilt; then fetch_tgz tilt "$TILT_URL" "$TILT_SHA256" "$dest" tilt 0; fi
if want ctlptl; then fetch_tgz ctlptl "$CTLPTL_URL" "$CTLPTL_SHA256" "$dest" ctlptl 0; fi
if want docker-compose; then
fetch_bin docker-compose "$COMPOSE_URL" "$COMPOSE_SHA256" "$dest"
fi
fi
fix_ownership "$dest"
# kind writes the kubeconfig as root too; hand that back as well when it's
# a mounted host directory rather than container-local state.
fix_ownership "${KUBE_DIR:-/out/kube}"
}
# ── install ────────────────────────────────────────────────────────────────
report_manual() {
echo
if [ ${#MANUAL[@]} -eq 0 ]; then
echo "nothing left to do by hand."
return
fi
echo "host actions this cannot perform (${#MANUAL[@]}):"
echo
local n=1
for m in "${MANUAL[@]}"; do
echo " $n. $m"
echo
n=$((n + 1))
done
}
# Installing into a directory that sits early in PATH silently replaces whatever
# the machine was already using — which on a shared or client machine can break
# unrelated work (kubectl more than one minor away from a cluster is the common
# one). Say so; never decide it for them.
# Downloading a verified binary proves it is the right file, not that this
# machine can run it. On an old distro tilt fails here, with a linker error
# about a missing symbol, and finding that out now beats finding out during a
# first cluster build.
verify_tools() {
local tier="${1:-dev}" b bin out rc broke=0
echo "checking that each one actually runs"
for b in $(tier_tools "$tier"); do
bin="$OUT_BIN/$b"
if [ ! -x "$bin" ]; then
printf ' %-14s not installed\n' "$b"
continue
fi
# Not piped into `head`. With `pipefail` set, a tool that prints more
# than one line gets SIGPIPE when head closes the pipe, and the
# pipeline reports 141 — so a working kubectl was announced as "does
# not run here", with its own correct version string as the evidence.
# Take the first line afterwards, from the string.
rc=0
case "$b" in
kubectl) out=$("$bin" version --client 2>&1) || rc=$? ;;
jq) out=$("$bin" --version 2>&1) || rc=$? ;;
*) out=$("$bin" version 2>&1) || rc=$? ;;
esac
out=${out%%$'\n'*}
if [ "$rc" -eq 0 ]; then
printf ' %-14s %s\n' "$b" "$out"
else
printf ' ! %-12s does not run here: %s\n' "$b" "$out"
broke=1
fi
done
if [ "$broke" -eq 1 ]; then
echo
echo " A binary that downloads and verifies but will not start is almost"
echo " always this distro's libc being older than the release needs."
echo " 'detect' prints the glibc version. The core tier (kubectl + jq)"
echo " has no such dependency and will work regardless."
fi
return 0
}
list() {
echo "pinned, linux/amd64 only:"
printf ' %-14s %s\n' kubectl "$KUBECTL_VERSION"
printf ' %-14s %s\n' jq "$JQ_VERSION"
printf ' %-14s %s\n' kind "$KIND_VERSION"
printf ' %-14s %s\n' tilt "$TILT_VERSION"
printf ' %-14s %s\n' ctlptl "$CTLPTL_VERSION"
printf ' %-14s %s\n' docker-compose "$COMPOSE_VERSION"
echo
echo " core = $CORE_TOOLS"
echo " dev = $CORE_TOOLS $DEV_TOOLS"
echo
echo "Checksums are pinned in the block at the top of this file. To bump one,"
echo "take the new checksum from the publisher's own release list — the header"
echo "comment has the exact commands."
return 0
}
tier_tools() { [ "$1" = "core" ] && echo "$CORE_TOOLS" || echo "$CORE_TOOLS $DEV_TOOLS"; }
warn_shadowing() {
local b existing shadowed="" tier="${1:-dev}"
for b in $(tier_tools "$tier"); do
[ -x "$OUT_BIN/$b" ] || continue
# Where would this resolve if OUT_BIN weren't in the way?
existing=$(PATH=$(echo "$PATH" | tr ':' '\n' | grep -vx "$OUT_BIN" | paste -sd:) \
command -v "$b" 2>/dev/null || true)
[ -n "$existing" ] || continue
[ "$existing" = "$OUT_BIN/$b" ] && continue
# The same version in both places is not a conflict: nothing changes for
# any other project whichever copy PATH happens to find first.
if version_matches "$b" "$existing" "$(pin_of "$b")"; then continue; fi
shadowed+=" $b $existing"$'\n'
done
[ -n "$shadowed" ] || return 0
case ":${PATH}:" in
*":$OUT_BIN:"*) ;;
*) return 0 ;; # not on PATH yet, so nothing is being shadowed
esac
echo
echo " ! these were already installed elsewhere and are now shadowed by $OUT_BIN:"
printf '%s' "$shadowed"
echo " Other projects on this machine will pick up the new versions."
MANUAL+=("Decide which toolchain wins. To keep the previous one, remove what
was just installed:
rm -f $(for b in $(tier_tools "$tier"); do printf '%s ' "$OUT_BIN/$b"; done)
Or install somewhere private instead:
OUT_BIN=\$PWD/def/bin make deps # then put that dir first in PATH")
}
# A copy in OUT_BIN only gives you `docker-compose`. That hyphenated form is the
# retired v1 spelling; every compose file written in the last few years assumes
# `docker compose`, which resolves plugins BY NAME out of a plugin directory.
# So the binary is fetched like any other and then linked, in your own home —
# no root, and nothing outside it.
install_compose_plugin() {
local src="$OUT_BIN/docker-compose" dir="$HOME/.docker/cli-plugins"
[ -x "$src" ] || return 0
mkdir -p "$dir"
# Something else already owns that name — docker-desktop and some distro
# packages install a real file there. Overwriting it would take the plugin
# away from whatever put it there, so say so and let the user decide.
if [ -e "$dir/docker-compose" ] && [ ! -L "$dir/docker-compose" ]; then
MANUAL+=("Something already installs the compose plugin at
$dir/docker-compose
To use rig's pinned build instead:
ln -sf $src $dir/docker-compose")
return 0
fi
ln -sfn "$src" "$dir/docker-compose"
echo " compose plugin -> $dir/docker-compose"
return 0
}
install() {
local tier="${1:-dev}" b
TIER="$tier"
detect
# detect_toolchain has already probed PATH. Fetch only what it found missing
# or at the wrong version; a tool already present at its pin stays where it is.
if [ -n "$TOOLCHAIN_NEED" ]; then
echo
DEPS_ONLY="$TOOLCHAIN_NEED" fetch "$tier"
echo
echo "installed to $OUT_BIN ($tier):"
for b in $TOOLCHAIN_NEED; do
if [ -x "$OUT_BIN/$b" ]; then echo " $b"; fi
done
if [ "$tier" = "core" ]; then
echo " (no kind/tilt — 'make deps dev' adds them)"
fi
# Only when compose was one of the things fetched: linking a binary
# that is already satisfied elsewhere on PATH would point the plugin at
# a copy rig did not install.
case " $TOOLCHAIN_NEED " in
*" docker-compose "*) install_compose_plugin ;;
esac
# Only worth saying when something actually landed in OUT_BIN. When every
# tool was satisfied elsewhere, OUT_BIN may reasonably be off PATH, and
# telling the user to add it would be advice to fix nothing.
case ":${PATH}:" in
*":$OUT_BIN:"*) ;;
*) MANUAL+=("Put the toolchain on your PATH — add to ~/.bashrc:
export PATH=\"${OUT_BIN}:\$PATH\"") ;;
esac
fi
warn_shadowing "$tier"
report_manual
}
# ── main ───────────────────────────────────────────────────────────────────
require_linux
# Read the command, THEN shift — and shift only if there is something there.
# A bare `shift` with no positional parameters returns 1, and under `set -e`
# that ended the script before a single line was printed: running this with no
# arguments at all, the documented default, did nothing and said nothing.
cmd="${1:-install}"
[ $# -gt 0 ] && shift
# Baked mode copies binaries already in the image, so it needs no downloader.
need_downloads() {
require_amd64
if [ "$DEPS_SOURCE" != baked ]; then pick_downloader; fi
pick_sha
}
case "$cmd" in
detect) detect; report_manual ;;
list) list ;;
verify) verify_tools "${1:-dev}" ;;
fetch) need_downloads; fetch "$@" ;;
install) need_downloads; install "${1:-dev}" ;;
*) echo "usage: $0 [detect|list|verify|fetch|install]" >&2
echo " install [core|dev] (default dev)" >&2
echo " fetch [core|dev] [--to DIR]" >&2
echo " OUT_BIN=<dir> overrides the install directory" >&2
exit 1 ;;
esac

View File

@@ -1,272 +0,0 @@
#!/usr/bin/env bash
# Share ONE Docker daemon across WSL distros, instead of running one per distro.
#
# Why this exists
# ---------------
# WSL2 distros share a kernel and a network stack. Two dockerd instances then
# contend over docker0 and iptables, which can disturb the daemon you actually
# depend on. Docker Desktop avoids this by running a single daemon in a
# dedicated distro and sharing its socket — this is the same idea, without
# Docker Desktop.
#
# So a throwaway rig box does NOT install Docker. It borrows the daemon from
# whichever distro is the designated host. That also makes the test more honest:
# rig never installs Docker anyway — Docker is its documented prerequisite.
#
# How
# ---
# /mnt/wsl is a tmpfs with `shared` mount propagation, visible to every distro
# in the WSL VM. The owning distro exposes its socket there; guests point
# DOCKER_HOST at it. Two ways, with different costs:
#
# share bind-mount the existing socket onto the shared tmpfs.
# Instant, and dockerd is NEVER restarted. Lasts until the
# next WSL shutdown.
# share --persist additionally install a systemd drop-in so dockerd listens
# there itself. Survives restarts, but requires one Docker
# restart now — which stops every container that has no
# restart policy, since live-restore is off by default.
#
# The bind mount is the default precisely because the persistent version's cost
# is paid on a machine that is already working.
#
# Reversibility is the whole design
# ---------------------------------
# `unshare` removes the bind mount (no restart) and, if present, the drop-in.
# The original systemd unit is never edited — only an additive drop-in file is
# ever created — so undoing is deletion, not repair. `status` always states
# which of the three roles a distro is in, in those words.
#
# Nothing here runs automatically. It does nothing until invoked.
#
# Usage:
# dockerhost.sh status # which distro owns Docker; what this one uses
# dockerhost.sh share # share it (bind mount, no daemon restart)
# dockerhost.sh share --persist # ...and survive WSL restarts (restarts Docker)
# dockerhost.sh unshare # undo it; this distro owns its Docker again
# dockerhost.sh use [--persist] # point THIS distro at the shared socket
set -euo pipefail
SHARED_DIR=/mnt/wsl/shared-docker
SHARED_SOCK="$SHARED_DIR/docker.sock"
OWNER_FILE="$SHARED_DIR/OWNER"
DROPIN=/etc/systemd/system/docker.service.d/10-rig-shared-socket.conf
PROFILE_D=/etc/profile.d/rig-docker-host.sh
distro_name() { echo "${WSL_DISTRO_NAME:-$(hostname)}"; }
require_wsl() {
grep -qi microsoft /proc/version 2>/dev/null && return 0
echo "dockerhost is WSL-only: it relies on /mnt/wsl being shared between distros." >&2
exit 1
}
# ── status ─────────────────────────────────────────────────────────────────
status() {
require_wsl
echo "distro $(distro_name)"
if [ -f "$DROPIN" ] || mountpoint -q "$SHARED_SOCK" 2>/dev/null; then
echo "role SHARING — this distro's Docker is offered to other distros"
elif [ -n "${DOCKER_HOST:-}" ] && [ "${DOCKER_HOST}" = "unix://$SHARED_SOCK" ]; then
echo "role BORROWING — using another distro's Docker"
else
echo "role standalone — this WSL installation has the main host Docker"
fi
echo
if [ -S "$SHARED_SOCK" ]; then
echo "shared sock $SHARED_SOCK (present)"
[ -f "$OWNER_FILE" ] && sed 's/^/ /' "$OWNER_FILE"
else
echo "shared sock none — no distro is sharing right now"
fi
echo
echo "DOCKER_HOST ${DOCKER_HOST:-(unset — using /var/run/docker.sock)}"
if command -v docker >/dev/null 2>&1; then
echo "docker $(docker version --format '{{.Server.Version}}' 2>/dev/null || echo unreachable)"
else
echo "docker cli not installed"
fi
}
# ── share / unshare (run on the host distro) ───────────────────────────────
# Default: expose the EXISTING socket by bind-mounting it onto the shared tmpfs.
# /mnt/wsl has `shared` propagation, so the mount is visible in other distros.
#
# The point of doing it this way is that dockerd is never restarted. Restarting
# it stops every container that has no restart policy (live-restore is off by
# default), which on a working machine means quietly killing whatever you had
# running. Not a trade worth making just to expose a socket.
#
# Cost: a bind mount does not survive a WSL VM shutdown. `--persist` adds the
# systemd drop-in as well, which does survive but needs that one restart.
share_bind() {
mkdir -p "$SHARED_DIR"
chmod 0755 "$SHARED_DIR"
if mountpoint -q "$SHARED_SOCK" 2>/dev/null; then
echo "already bind-mounted at $SHARED_SOCK"
else
[ -S /var/run/docker.sock ] || { echo "no /var/run/docker.sock here" >&2; exit 1; }
# The target must exist as a file for a bind mount onto it.
[ -e "$SHARED_SOCK" ] || : > "$SHARED_SOCK"
mount --bind /var/run/docker.sock "$SHARED_SOCK"
echo "bind-mounted /var/run/docker.sock -> $SHARED_SOCK (no daemon restart)"
fi
cat > "$OWNER_FILE" <<EOF
owner distro: $(distro_name)
docker gid: $(getent group docker | cut -d: -f3)
socket: $SHARED_SOCK
method: bind-mount (until the next WSL shutdown)
EOF
}
share() {
require_wsl
[ "$(id -u)" -eq 0 ] || { echo "run with sudo: sudo bash ctrl/dockerhost.sh share" >&2; exit 1; }
share_bind
if [ "${1:-}" != "--persist" ]; then
echo
echo "This lasts until the next WSL shutdown. To make it survive, re-run with"
echo "--persist — but note that adds a systemd drop-in and RESTARTS Docker,"
echo "which stops any container that has no restart policy."
return 0
fi
if [ -f "$DROPIN" ]; then
echo "drop-in already present — sharing persists across restarts."
return 0
fi
echo
echo "--persist: installing a systemd drop-in and restarting Docker."
echo "Containers without a restart policy will stop and will NOT come back."
docker ps --format ' {{.Names}} restart={{.HostConfig.RestartPolicy.Name}}' 2>/dev/null \
|| docker ps --format ' {{.Names}}' 2>/dev/null || true
echo
local exec_line
exec_line=$(systemctl cat docker.service | grep -m1 '^ExecStart=')
if [ -z "$exec_line" ]; then
echo "could not read docker.service ExecStart — refusing to guess" >&2
exit 1
fi
mkdir -p "$(dirname "$DROPIN")" "$SHARED_DIR"
# Additive only: blank the inherited ExecStart, then restate it verbatim
# with one extra -H. Nothing about the original unit is edited.
cat > "$DROPIN" <<EOF
# Added by rig (ctrl/dockerhost.sh share).
#
# Adds a SECOND listening socket on the WSL-shared tmpfs so other distros can
# use this daemon instead of running their own. The original socket is
# untouched, so this distro behaves exactly as before.
#
# To undo: sudo bash ctrl/dockerhost.sh unshare
[Service]
ExecStartPre=-/bin/mkdir -p $SHARED_DIR
ExecStartPre=-/bin/chmod 0755 $SHARED_DIR
ExecStart=
${exec_line} -H unix://$SHARED_SOCK
EOF
systemctl daemon-reload
systemctl restart docker
# Guests need a group with a MATCHING GID to use the socket; GIDs are not
# consistent across distros, so record ours rather than assume.
cat > "$OWNER_FILE" <<EOF
owner distro: $(distro_name)
docker gid: $(getent group docker | cut -d: -f3)
socket: $SHARED_SOCK
EOF
echo "sharing from '$(distro_name)'"
echo " guests: export DOCKER_HOST=unix://$SHARED_SOCK"
echo " undo: sudo bash ctrl/dockerhost.sh unshare"
echo
echo "NOTE: /mnt/wsl is tmpfs and is cleared when the WSL VM shuts down."
echo " The drop-in recreates the directory on the next Docker start."
}
unshare_() {
require_wsl
[ "$(id -u)" -eq 0 ] || { echo "run with sudo: sudo bash ctrl/dockerhost.sh unshare" >&2; exit 1; }
local did=0
# The bind mount first: undoing it needs no restart, so a plain `share`
# is fully reversible without disturbing anything.
if mountpoint -q "$SHARED_SOCK" 2>/dev/null; then
umount "$SHARED_SOCK"
rm -f "$SHARED_SOCK"
echo " removed the bind mount (no restart needed)"
did=1
fi
rm -f "$OWNER_FILE"
rmdir "$SHARED_DIR" 2>/dev/null || true
if [ -f "$DROPIN" ]; then
rm -f "$DROPIN"
rmdir "$(dirname "$DROPIN")" 2>/dev/null || true
systemctl daemon-reload
systemctl restart docker
echo " removed the systemd drop-in and restarted Docker"
did=1
fi
if [ "$did" -eq 0 ]; then
echo "not sharing — this WSL installation already has the main host Docker."
return 0
fi
echo "restored: this WSL installation has the main host Docker again."
echo " (nothing else was changed; the original unit was never edited)"
}
# ── use (run on a guest distro) ────────────────────────────────────────────
use() {
require_wsl
if [ ! -S "$SHARED_SOCK" ]; then
echo "no shared socket at $SHARED_SOCK" >&2
echo "Run 'sudo bash ctrl/dockerhost.sh share' in the distro that owns Docker." >&2
exit 1
fi
# Align the local docker group GID with the owner's, or the socket is
# unreadable here even though it is visible.
if [ -f "$OWNER_FILE" ] && [ "$(id -u)" -eq 0 ]; then
local gid; gid=$(awk '/docker gid:/ {print $3}' "$OWNER_FILE")
if [ -n "$gid" ]; then
if getent group docker >/dev/null; then
[ "$(getent group docker | cut -d: -f3)" = "$gid" ] || groupmod -g "$gid" docker
else
groupadd -g "$gid" docker
fi
fi
fi
if [ "${1:-}" = "--persist" ]; then
[ "$(id -u)" -eq 0 ] || { echo "--persist needs root" >&2; exit 1; }
echo "export DOCKER_HOST=unix://$SHARED_SOCK" > "$PROFILE_D"
echo "persisted in $PROFILE_D"
fi
echo "export DOCKER_HOST=unix://$SHARED_SOCK"
}
case "${1:-status}" in
status) status ;;
share) shift; share "${1:-}" ;;
unshare) unshare_ ;;
use) shift; use "${1:-}" ;;
*) echo "usage: $0 [status|share|unshare|use [--persist]]" >&2; exit 1 ;;
esac

View File

@@ -1,3 +1,8 @@
# EXAMPLE PROFILE. rig needs none of these: with no profile it runs on its
# built-in defaults (lib/config.sh). To use this one, copy it to client.env in this
# directory and name it — PROFILE=client in ctrl/.env, or on the command line. It
# then overlays the defaults; anything it does not set, they still supply.
#
# client — the regulated-estate shape. Multi-node so taints, affinity and
# topology are real; apiserver audit on; images through a pull-through cache of
# the corporate registry.
@@ -19,7 +24,7 @@ DNS_MODE=hosts
# Opt in to the real ports below only when this is the ONLY environment and
# nothing else owns :80. They fail to bind otherwise, and docker reports it as an
# opaque "failed to bind host port 0.0.0.0:80/tcp: address already in use"
# halfway through cluster creation. `make station` checks before you spend the
# halfway through cluster creation. `make check` checks before you spend the
# time. Uncommenting also means only one environment can exist at a time.
# HTTP_PORT=80
# HTTPS_PORT=443

View File

@@ -1,11 +1,10 @@
# data — a cluster with the dependency containers a soleprint room asks for.
# EXAMPLE PROFILE. rig needs none of these: with no profile it runs on its
# built-in defaults (lib/config.sh). To use this one, copy it to data.env in this
# directory and name it — PROFILE=data in ctrl/.env, or on the command line. It
# then overlays the defaults; anything it does not set, they still supply.
#
# The point of this profile is that a room declares what it needs once, in
# cfg/<room>/data/cabinets.json, and gets it on either target: `build.py`
# composes those services into docker-compose.yml for a laptop, and the addons
# below install the same ones here. The names match deliberately —
# soleprint/station/cabinets/<name>/cabinet.json carries a `rig_addon` field
# pointing at ctrl/addons/<name>.sh.
# data — databases and a scheduler for an environment that needs them: postgres,
# redis and airflow, each an upstream image run unmodified.
#
# Everything lands in the `data` namespace (DATA_NAMESPACE to move it), so
# `make cluster reset` on the app namespace leaves the databases alone.
@@ -19,7 +18,7 @@ KIND_CONFIG=kind-config.yaml.tpl
# Order matters: addons.sh installs in the order listed, and airflow refuses to
# start without a metadata database, so postgres comes first.
ADDONS="metallb postgres redis airflow"
# local, not none — see minimal.env: `none` has no outward-push guard.
# local, not none — see the defaults in lib/config.sh: `none` has no outward-push guard.
REGISTRY_MODE=local
INGRESS_MODE=hostport
DNS_MODE=hosts
@@ -30,8 +29,8 @@ DATA_NAMESPACE=data
# Postgres identity. The password is not here: postgres.sh generates one on
# first install and keeps it across re-runs, so re-running the addon never
# rotates the credential out from under whatever is already connected.
POSTGRES_DB=soleprint
POSTGRES_USER=soleprint
POSTGRES_DB=app
POSTGRES_USER=app
POSTGRES_STORAGE=2Gi
AIRFLOW_ADMIN_USER=admin

View File

@@ -1,21 +0,0 @@
# minimal — the default. One node, no addons, no registry.
# Assumes nothing and boots fast. Start here; move to client.env when you need
# the regulated behaviours.
#
PROFILE_NAME=minimal
K8S_VERSION=v1_36
KIND_CONFIG=kind-config.yaml.tpl
ADDONS=""
# local, not none: `none` leaves the cluster with no registry to push to, and an
# unqualified image name then means docker.io/library/<name>. In a regulated
# estate that is a disclosure risk, not a convenience trade — so the default
# carries the guard even though it costs one container.
REGISTRY_MODE=local
INGRESS_MODE=hostport
DNS_MODE=hosts
# Ports are deliberately NOT set here. They derive from the directory name so
# several environments coexist — see ctrl/ports.sh, and `make ports` to see the
# block this one gets. A fixed default here would collide with whatever else the
# machine happens to be running; 8080 in particular is rarely free.

View File

@@ -1,5 +1,10 @@
# EXAMPLE PROFILE. rig needs none of these: with no profile it runs on its
# built-in defaults (lib/config.sh). To use this one, copy it to offline.env in this
# directory and name it — PROFILE=offline in ctrl/.env, or on the command line. It
# then overlays the defaults; anything it does not set, they still supply.
#
# offline — air-gapped. Everything comes from a local registry that was loaded
# ahead of time; nothing reaches the internet. Pair with the wizard-full image
# ahead of time; nothing reaches the internet. Pair with the deps-full image
# (DEPS_SOURCE=baked) so the toolchain install is offline too.
#
# The heavier addons are left out to keep first boot viable.

View File

@@ -1,15 +0,0 @@
# /etc/hosts block for this environment. Rendered by newbox.sh; ${CLUSTER} and
# ${HTTP_PORT} are substituted.
#
# Hostnames are a convenience, not a requirement — every service is reachable at
# localhost:<port> without any of this, which is why DNS is not touched by
# default. Add entries here as the model grows.
#
# On Windows the same block has to go in
# C:\Windows\System32\drivers\etc\hosts for a browser to resolve these. That
# file does NOT support wildcards, so every name must be listed explicitly.
# newbox.sh prints the block for you to paste rather than editing it.
127.0.0.1 ${CLUSTER}.local
127.0.0.1 api.${CLUSTER}.local
127.0.0.1 docs.${CLUSTER}.local

View File

@@ -1,8 +1,7 @@
# `ctrl/k8s` — cluster shape, and what runs on it
Same layout as every other project here (`unt`, `nvi`, `eth`, `mpr`, and
soleprint's generated rooms): a kind config, a kustomize `base/`, and an
`overlays/dev/` that patches it. See ALL `projects/templates/conventions.md`.
Same layout as every other project here: a kind config, a kustomize `base/`,
and an `overlays/dev/` that patches it.
```
kind-config*.yaml.tpl the cluster itself — nodes, ports, audit
@@ -29,12 +28,12 @@ not restate what the YAML already says.
| file | nodes | audit | profiles |
| --- | --- | --- | --- |
| `kind-config.yaml.tpl` | 1 | off | `minimal`, `data` |
| `kind-config.audit.yaml.tpl` | 1 | on | `offline` |
| `kind-config.client.yaml.tpl` | 3 | on | `client` |
| `kind-config.yaml.tpl` | 1 | off | the default; the `data` example |
| `kind-config.audit.yaml.tpl` | 1 | on | the `offline` example |
| `kind-config.client.yaml.tpl` | 3 | on | the `client` example |
A profile picks one with `KIND_CONFIG` in `ctrl/env.d/<profile>.env`. Adding a
shape is adding a file — there is no dispatcher to edit.
With no profile the default shape is used; a profile picks another with
`KIND_CONFIG`. Adding a shape is adding a file — there is no dispatcher to edit.
Audit is an apiserver flag and therefore fixed at creation: changing it is
`make cluster reset`, not a re-apply.

View File

@@ -44,7 +44,7 @@ nodes:
- role: control-plane
image: ${NODE_IMAGE}
# hostPath is resolved by the HOST dockerd, so this must be a host path even
# when cluster.sh runs inside the wizard container. HOST_WORKDIR says where
# when cluster.sh runs inside the installer container. HOST_WORKDIR says where
# this rig lives on the host; bare on a host it is just the repo root.
extraMounts:
- hostPath: ${HOST_WORKDIR}/ctrl/k8s/audit-policy.yaml

View File

@@ -5,10 +5,11 @@
# definition of how the config layers compose, which every script has to agree
# on exactly. Precedence, weakest first:
#
# built-in defaults below; fill only what nothing else set
# ctrl/versions.env pinned toolchain + image digests (committed)
# ctrl/env.d/<profile> cluster shape (committed)
# ctrl/env.d/<profile> cluster shape — OPTIONAL, examples ship as *.env.example
# ctrl/.env machine-local values and secrets (gitignored)
# the caller's env `make cluster up PROFILE=client` (always wins)
# the caller's env `make cluster up PROFILE=<name>` (always wins)
#
# That last rule is why this is more than a few `source` lines: .env sets
# PROFILE, so without snapshotting it would silently override the PROFILE the
@@ -21,9 +22,14 @@
# NODES and AUDIT are deliberately NOT here: they are properties of the chosen
# ctrl/k8s/kind-config*.yaml.tpl and are read back out of it below, so there is
# one place that decides the shape of the cluster rather than two that can drift.
#
# REGISTRY_PORT and MANIFESTS_DIR were missing here while ctrl/.env set them, so
# the caller's env silently LOST to the file for those two — the one precedence
# rule this header states. Both are now listed; the other twelve are unchanged.
CONFIG_OVERRIDABLE="PROFILE CLUSTER K8S_VERSION KIND_CONFIG ADDONS
REGISTRY_MODE INGRESS_MODE DNS_MODE TILT_PORT
SOURCE ARCH DEPS_SOURCE HTTP_PORT HTTPS_PORT"
SOURCE ARCH DEPS_SOURCE HTTP_PORT HTTPS_PORT
REGISTRY_PORT MANIFESTS_DIR"
# The containing folder's name, reduced to something kind accepts as a cluster
# name (a DNS label: lowercase alphanumerics and dashes). Run from ctrl/, so the
@@ -56,25 +62,48 @@ load_config() {
set -a
source ./versions.env
[ -f ./.env ] && source ./.env
# RIG_PORTABLE skips the machine-local layer. config_snapshot sets it, so a
# generated standalone kit never carries this machine's .env — which holds
# local values and, by its own description, secrets.
if [ -z "${RIG_PORTABLE:-}" ] && [ -f ./.env ]; then source ./.env; fi
set +a
# Re-apply overrides now so PROFILE is the caller's before we pick the file.
_config_restore "$saved"
local profile="${PROFILE:-minimal}"
if [ ! -f "./env.d/${profile}.env" ]; then
echo "no such profile: env.d/${profile}.env" >&2
echo "available: $(ls env.d/*.env 2>/dev/null | xargs -n1 basename | sed 's/\.env$//' | tr '\n' ' ')" >&2
exit 1
# A profile is an optional overlay, never a prerequisite. rig assumes no
# configuration: with no profile named — or no env.d/ at all — it runs on the
# built-in defaults below. What IS an error is naming a profile that does not
# exist, because a typo must not quietly fall back to something else.
local profile="${PROFILE:-}"
if [ -n "$profile" ] && [ "$profile" != default ]; then
if [ ! -f "./env.d/${profile}.env" ]; then
echo "no such profile: env.d/${profile}.env" >&2
echo "available: $(config_profiles | tr '\n' ' ')" >&2
exit 1
fi
set -a
source "./env.d/${profile}.env"
if [ -z "${RIG_PORTABLE:-}" ] && [ -f ./.env ]; then source ./.env; fi
set +a
_config_restore "$saved"
fi
set -a
source "./env.d/${profile}.env"
[ -f ./.env ] && source ./.env
set +a
_config_restore "$saved"
# The defaults a profile would otherwise have to supply. Weakest of all: a
# profile, ctrl/.env and the caller each override them.
PROFILE_NAME="${PROFILE_NAME:-default}"
ADDONS="${ADDONS-}"
# local, not none: with no registry an unqualified image name means
# docker.io/library/<name>, and a default must not make that disclosure.
REGISTRY_MODE="${REGISTRY_MODE:-local}"
INGRESS_MODE="${INGRESS_MODE:-hostport}"
DNS_MODE="${DNS_MODE:-hosts}"
# The newest node image versions.env pins, found rather than restated, so
# bumping the pins moves the default with them.
if [ -z "${K8S_VERSION:-}" ]; then
K8S_VERSION=$(compgen -v NODE_IMAGE_v | sort -V | tail -1)
K8S_VERSION="${K8S_VERSION#NODE_IMAGE_}"
fi
# Identity follows the FOLDER, so copying this directory somewhere else and
# renaming it yields a distinct environment with no further edits. Without
@@ -93,6 +122,12 @@ load_config() {
TILT_PORT="${TILT_PORT:-$((base + 2))}"
REGISTRY_PORT="${REGISTRY_PORT:-$((base + 3))}"
# Where the workload's manifests live, repo-root relative. Defaulted here so
# it is always resolved rather than sometimes-set: it is the seam that lets
# the real manifests be versioned away from the installer, and a consumer
# should not have to know whether anyone filled it in. See k8s/README.md.
MANIFESTS_DIR="${MANIFESTS_DIR:-ctrl/k8s/overlays/dev}"
# Profiles name a k8s minor (v1_36); versions.env holds the pinned digest.
local var="NODE_IMAGE_${K8S_VERSION}"
NODE_IMAGE="${!var:-}"
@@ -103,20 +138,41 @@ load_config() {
# The cluster's shape is a file in ctrl/k8s/, named by the profile. Adding a
# shape is adding a file; there is no dispatcher to edit.
#
# A host that needs its own shape — extra port mappings, more nodes — passes
# an absolute path instead, and rig renders it exactly like one of its own:
# ${CLUSTER} and ${NODE_IMAGE} are substituted either way. The shape stays in
# the host's tree, because what a host's cluster needs is the host's business;
# rig only knows how to build whatever it is handed.
KIND_CONFIG="${KIND_CONFIG:-kind-config.yaml.tpl}"
KIND_CONFIG_PATH="./k8s/${KIND_CONFIG}"
case "$KIND_CONFIG" in
/*) KIND_CONFIG_PATH="$KIND_CONFIG"; KIND_CONFIG_SHOWN="$KIND_CONFIG" ;;
*) KIND_CONFIG_PATH="./k8s/${KIND_CONFIG}"; KIND_CONFIG_SHOWN="ctrl/k8s/${KIND_CONFIG}" ;;
esac
if [ ! -f "$KIND_CONFIG_PATH" ]; then
echo "no such cluster shape: ctrl/k8s/${KIND_CONFIG}" >&2
echo "available: $(ls k8s/kind-config*.yaml.tpl 2>/dev/null | xargs -n1 basename | tr '\n' ' ')" >&2
echo "no such cluster shape: ${KIND_CONFIG_SHOWN}" >&2
echo "rig's own: $(ls k8s/kind-config*.yaml.tpl 2>/dev/null | xargs -n1 basename | tr '\n' ' ')" >&2
echo "or pass an absolute path to a shape of your own" >&2
exit 1
fi
# Read the shape back out of the YAML rather than trusting a profile to
# restate it. station.sh sizes the memory warning on NODES, and cluster.sh
# restate it. check.sh sizes the memory warning on NODES, and cluster.sh
# prints AUDIT before spending minutes building something that cannot be
# changed afterwards — both would mislead if the numbers drifted.
NODES=$(grep -c '^ - role:' "$KIND_CONFIG_PATH")
if grep -q 'audit-policy-file' "$KIND_CONFIG_PATH"; then AUDIT=on; else AUDIT=off; fi
# What one node costs, measured rather than guessed. On 2026-09-11 a minimal
# control-plane node ran at 620 MiB idle and ~728 MiB with a small mock, plus
# 16 MiB for the local registry — ~745 MiB of working set. 800 rounds that up,
# and agrees with the 800 MB observed independently on a larger rig. Worker
# nodes carry no etcd or apiserver and are lighter, so for a multi-node shape
# this errs high. It is the cluster alone: whatever you deploy comes on top.
#
# Here rather than in check.sh because the memory tool and every standalone
# kit need the same figure.
NODE_MB=800
}
# Render a cluster shape to stdout. sed rather than envsubst: envsubst is
@@ -125,7 +181,7 @@ load_config() {
# depending on something the caller does not set.
#
# hostPath entries are resolved by the HOST dockerd, so HOST_WORKDIR must stay a
# host path even when this runs inside the wizard container.
# host path even when this runs inside the installer container.
render_kind_config() {
local host_workdir="${HOST_WORKDIR:-$(cd .. && pwd)}"
sed -e "s|\${CLUSTER}|${CLUSTER}|g" \
@@ -146,3 +202,118 @@ _config_restore() {
# line would otherwise make this return 1 and trip `set -e` in the caller.
return 0
}
# ── what a standalone kit needs to know ────────────────────────────────────
# Two questions the kit generator (ctrl/standalone.sh) asks, so that it never has
# to know how configuration is stored. Where profiles live, which files are
# layered and what is derived are this file's business and can change freely;
# the generator only calls these.
# Every configuration rig can be run as, one per line: each profile file, or —
# when there are none — `default`, the built-in configuration load_config uses
# when no profile is named. Never empty, because rig never needs a profile.
config_profiles() {
local f found=""
for f in ./env.d/*.env; do
[ -e "$f" ] || continue
f=${f##*/}; echo "${f%.env}"; found=1
done
[ -n "$found" ] || echo default
}
# The resolved configuration, as `declare -p` lines — exactly what load_config
# leaves behind, minus the machine-local layer. A kit freezes this in place of
# load_config, so it carries rig's decisions and not this machine's secrets.
#
# config_snapshot <profile> that profile, as any machine would resolve it
# config_snapshot --current what THIS machine runs: every overridable key as
# resolved here, handed back in as if typed on the
# command line, over the same portable resolution.
# Values derived from those choices follow them;
# anything else the local layer set — credentials —
# is not carried. config_left_out names it.
#
# Found by difference, not by a list: whatever load_config sets today, it sets.
# A list here would be one more place to forget a variable.
config_snapshot() {
local _rig_snap_choices
if [ "$1" = --current ]; then
_rig_snap_choices=$( (
load_config >/dev/null || exit 1
for _rig_snap_n in $CONFIG_OVERRIDABLE; do
if [ -n "${!_rig_snap_n+x}" ]; then printf 'export %s=%q\n' "$_rig_snap_n" "${!_rig_snap_n}"; fi
done
) ) || return 1
else
_rig_snap_choices="export PROFILE=$(printf '%q' "$1")"
fi
(
# Nothing from the caller's shell may leak into a kit.
for _rig_snap_n in $CONFIG_OVERRIDABLE; do unset "$_rig_snap_n"; done
declare -A _rig_snap_was=()
for _rig_snap_n in $(compgen -v); do
_rig_snap_was[$_rig_snap_n]="${!_rig_snap_n-}"
done
eval "$_rig_snap_choices"
RIG_PORTABLE=1 load_config >/dev/null
for _rig_snap_n in $(compgen -v); do
case "$_rig_snap_n" in
_rig_snap_*|RIG_PORTABLE|BASH*|FUNCNAME|PIPESTATUS|LINENO|RANDOM|SRANDOM|\
SECONDS|EPOCH*|HISTCMD|COLUMNS|LINES|PWD|OLDPWD|_|SHLVL|OPTIND|OPTERR) continue ;;
esac
if [ -z "${_rig_snap_was[$_rig_snap_n]+x}" ] \
|| [ "${_rig_snap_was[$_rig_snap_n]}" != "${!_rig_snap_n-}" ]; then
declare -p "$_rig_snap_n"
fi
done
)
}
# The profile this machine runs, as load_config resolves it here.
config_current_profile() { ( load_config >/dev/null && echo "$PROFILE_NAME" ); }
# What an export of this machine's configuration does NOT carry, by name only:
# keys the machine-local layer sets that are not choices a caller may override.
# They are this machine's own — registry and mirror credentials, mostly — so the
# target has to be told to supply them. Values are never printed.
config_left_out() {
[ -f ./.env ] || return 0
local k
for k in $(sed -nE 's/^[[:space:]]*(export[[:space:]]+)?([A-Za-z_][A-Za-z0-9_]*)=.*/\2/p' ./.env | sort -u); do
case " $(echo $CONFIG_OVERRIDABLE) " in
*" $k "*) ;;
*) echo "$k" ;;
esac
done
}
# A replacement for load_config with a resolution frozen in (a profile, or
# --current — see config_snapshot), printed
# as a function definition for a standalone kit to carry. The generator embeds
# whatever this prints and interprets none of it, so what "frozen" means stays
# rig's decision.
#
# It keeps load_config's one stated rule: the caller's env wins for anything in
# CONFIG_OVERRIDABLE. A kit therefore behaves like rig — `OUT_BIN=... rigdeps.sh`
# still works — rather than like a copy with everything pinned.
#
# What freezing does give up, knowingly: values DERIVED from an overridable one
# are fixed at generation. Override CLUSTER and the ports stay the ones derived
# for the original name. Re-deriving would mean carrying the layering itself,
# which is exactly what a kit exists not to need.
config_freeze() {
local snap
snap=$(config_snapshot "$1") || return 1
cat <<'EOF'
load_config() {
local k saved=""
for k in $CONFIG_OVERRIDABLE; do
if [ -n "${!k+x}" ]; then saved+="$k=$(printf '%q' "${!k}")"$'\n'; fi
done
EOF
printf '%s\n' "$snap" | sed -E 's/^declare --* / declare -g /; s/^declare -([a-zA-Z]+) / declare -g\1 /'
cat <<'EOF'
_config_restore "$saved"
}
EOF
}

759
rig/ctrl/mem.sh Executable file
View File

@@ -0,0 +1,759 @@
#!/usr/bin/env bash
# rig:standalone rigmini status
# How much memory this machine will actually give you before something dies —
# rig's memory tool, and (generated from this file) the standalone rigmini.sh.
#
# There are two numbers and they are rarely the same. `status` reports what the
# machine ADVERTISES and what is quietly capping it. `push` finds what it will
# SURVIVE, by allocating until it stops. `all` does both and weighs the result
# against what this profile's cluster needs.
#
# The gap between them is the whole reason this exists. Under WSL the cap lives
# in .wslconfig; in a container or a managed workspace it is a cgroup limit, and
# there /proc/meminfo reports the HOST's memory while the kernel kills you at a
# fraction of it. A script that only read MemTotal would confidently report 32 GB
# on a box that OOMs at 2.
#
# Runs on native Linux and under WSL. On WSL the memory you see is a VM
# allocation that can be raised, and the commonest failure is raising it without
# restarting — so status compares what .wslconfig says with what actually booted.
#
# Reports and instructs. It never raises a limit, frees anything or installs a
# package. The one write it can make is `backup`, which copies .wslconfig beside
# itself, so that `restore` has something to put back after a hand edit.
#
# Usage:
# mem.sh status what it has, what caps it
# mem.sh push [--to GB] [--to-oom] climb until it stops
# mem.sh all [--budget GB] both, then the verdict
# mem.sh backup | restore .wslconfig, WSL only
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
# ── defaults ───────────────────────────────────────────────────────────────
STEP_MB=0 # per allocation; 0 means scale it to the ceiling. See push().
STEP_EXPLICIT=no # whether --step was given, which turns the scaling off.
TO_MB="" # --to: stop here regardless. Empty means no hard cap.
TO_OOM=no # --to-oom: opt in to running until the kernel intervenes.
BUDGET_GB="" # --budget; empty means what this profile's cluster needs, from rig.
BUDGET_EXPLICIT=no # whether --budget was given, which retires the guess below.
# ── platform ───────────────────────────────────────────────────────────────
# Windows outside WSL — Git Bash, MSYS, Cygwin — looks close enough to work and
# then fails in a pile of confusing ways: no /proc, no docker socket, none of
# the tooling. Detectable, so name it instead.
require_linux() {
case "$(uname -s)" in
MINGW*|MSYS*|CYGWIN*)
cat >&2 <<'EOF'
This has to run inside WSL, not Git Bash / MSYS / Cygwin.
If WSL is not installed yet, from an elevated PowerShell or Command Prompt:
wsl --install
That enables Windows features and needs a reboot, so it is not something this
script will do for you. Afterwards, open the Linux shell it installs and run
this from there.
EOF
exit 1 ;;
esac
# Everything below reads /proc. Without it there is nothing to measure, and
# failing here beats printing a page of empty fields.
if [ ! -r /proc/meminfo ]; then
echo "no readable /proc/meminfo — this needs a Linux kernel." >&2
echo "On macOS or a BSD none of the numbers below exist." >&2
exit 1
fi
}
is_wsl() { grep -qi microsoft /proc/version 2>/dev/null; }
is_container() {
[ -f /.dockerenv ] && return 0
grep -qE '(docker|containerd|kubepods|lxc|podman)' /proc/1/cgroup 2>/dev/null
}
platform() {
if is_wsl; then echo WSL
elif is_container; then echo container
else echo "native linux"
fi
}
# ── reading memory ─────────────────────────────────────────────────────────
mb() { echo $(( $(awk "/^$1:/{print \$2}" /proc/meminfo) / 1024 )); }
# MemAvailable arrived in kernel 3.14. Older kernels — and they turn up on
# corporate images — need the estimate it replaced, which is worse but not wrong.
avail_meminfo_mb() {
if grep -q '^MemAvailable:' /proc/meminfo; then
mb MemAvailable
else
awk '/^(MemFree|Buffers|Cached):/{t+=$2} END{print int(t/1024)}' /proc/meminfo
fi
}
# Where a cgroup records this cgroup's own limit and usage. Set once by
# find_cgroup, because every later reading needs both and hunting for the files
# on each call would be the slow part of the poll loop.
CG_MAX_FILE=""
CG_CUR_FILE=""
CG_VERSION=""
find_cgroup() {
local rel
# Inside a container the cgroup namespace makes the top of the tree BE the
# container's own cgroup, so the unqualified path is already the right one.
# On a host it is the root cgroup, which is never limited — hence the second
# attempt via /proc/self/cgroup, which names the slice this shell is in.
if [ -r /sys/fs/cgroup/memory.max ]; then
CG_VERSION=v2
CG_MAX_FILE=/sys/fs/cgroup/memory.max
CG_CUR_FILE=/sys/fs/cgroup/memory.current
elif [ -r /sys/fs/cgroup/memory/memory.limit_in_bytes ]; then
CG_VERSION=v1
CG_MAX_FILE=/sys/fs/cgroup/memory/memory.limit_in_bytes
CG_CUR_FILE=/sys/fs/cgroup/memory/memory.usage_in_bytes
fi
rel=$(awk -F: '$1=="0"{print $3; exit}' /proc/self/cgroup 2>/dev/null || true)
if [ -n "$rel" ] && [ "$rel" != "/" ] && [ -r "/sys/fs/cgroup${rel}/memory.max" ]; then
CG_VERSION=v2
CG_MAX_FILE="/sys/fs/cgroup${rel}/memory.max"
CG_CUR_FILE="/sys/fs/cgroup${rel}/memory.current"
return 0
fi
rel=$(awk -F: '$2 ~ /(^|,)memory(,|$)/{print $3; exit}' /proc/self/cgroup 2>/dev/null || true)
if [ -n "$rel" ] && [ "$rel" != "/" ] \
&& [ -r "/sys/fs/cgroup/memory${rel}/memory.limit_in_bytes" ]; then
CG_VERSION=v1
CG_MAX_FILE="/sys/fs/cgroup/memory${rel}/memory.limit_in_bytes"
CG_CUR_FILE="/sys/fs/cgroup/memory${rel}/memory.usage_in_bytes"
fi
return 0
}
# The cap in MB, or "" when there is none worth reporting. v2 spells unlimited
# "max"; v1 spells it as a number near 2^63, which is why this compares against
# MemTotal rather than testing for a magic value — a "limit" above the machine's
# own memory is not a limit, however it is written.
cgroup_cap_mb() {
local raw cap
[ -n "$CG_MAX_FILE" ] && [ -r "$CG_MAX_FILE" ] || { echo ""; return 0; }
raw=$(cat "$CG_MAX_FILE" 2>/dev/null || echo max)
[ "$raw" = "max" ] && { echo ""; return 0; }
case "$raw" in ''|*[!0-9]*) echo ""; return 0 ;; esac
cap=$((raw / 1024 / 1024))
[ "$cap" -ge "$(mb MemTotal)" ] && { echo ""; return 0; }
echo "$cap"
}
cgroup_used_mb() {
local raw
[ -n "$CG_CUR_FILE" ] && [ -r "$CG_CUR_FILE" ] || { echo ""; return 0; }
raw=$(cat "$CG_CUR_FILE" 2>/dev/null || echo "")
case "$raw" in ''|*[!0-9]*) echo ""; return 0 ;; esac
echo $((raw / 1024 / 1024))
}
# ulimit -v is a per-process address-space cap. It stops YOU long before the box
# does, and because it is inherited from a login shell it is easy to hit without
# knowing it is set.
ulimit_v_mb() {
local v; v=$(ulimit -v 2>/dev/null || echo unlimited)
[ "$v" = "unlimited" ] && { echo ""; return 0; }
case "$v" in ''|*[!0-9]*) echo ""; return 0 ;; esac
echo $((v / 1024))
}
# The number everything else is about: the lowest of the things that can stop
# you. Printed at the end of `status` and used as the sanity bound in `push`.
effective_ceiling_mb() {
local c; c=$(mb MemTotal)
local cap; cap=$(cgroup_cap_mb)
local ul; ul=$(ulimit_v_mb)
[ -n "$cap" ] && [ "$cap" -lt "$c" ] && c="$cap"
[ -n "$ul" ] && [ "$ul" -lt "$c" ] && c="$ul"
echo "$c"
}
# How much room is left RIGHT NOW, from whichever accounting actually governs.
# In a capped container /proc/meminfo describes the host and is worse than
# useless for this — it would report tens of gigabytes free on a box that is one
# allocation from being killed.
headroom_mb() {
local cap used
cap=$(cgroup_cap_mb)
used=$(cgroup_used_mb)
if [ -n "$cap" ] && [ -n "$used" ]; then
echo $(( cap - used ))
else
avail_meminfo_mb
fi
}
# ── status ─────────────────────────────────────────────────────────────────
# /mnt/c/Users can hold several real accounts — a renamed login leaves the old
# directory behind — so picking the first alphabetically is a coin toss. Ask
# Windows, then fall back to whichever profile actually owns a config.
wslconfig_path() {
local profile winpath found
profile=$(cmd.exe /c "echo %USERPROFILE%" 2>/dev/null | tr -d "\r\n" || true)
case "$profile" in
""|*%*) ;;
*) winpath=$(wslpath -u "$profile" 2>/dev/null || true)
if [ -n "$winpath" ] && [ -d "$winpath" ]; then
echo "$winpath/.wslconfig"; return 0
fi ;;
esac
found=$(ls -d /mnt/c/Users/*/.wslconfig 2>/dev/null | head -1 || true)
[ -n "$found" ] && echo "$found"
return 0
}
hogs() {
echo " holding the most:"
ps -eo rss,comm --sort=-rss 2>/dev/null \
| awk 'NR>1 && NR<=6 {printf " %6.0f MB %s\n", $1/1024, $2}'
return 0
}
status() {
local total avail swap_total swap_free cap ul cur
echo "host"
echo " platform $(platform)"
echo " kernel $(uname -r)"
[ -r /etc/os-release ] && \
echo " distro $(sed -n 's/^PRETTY_NAME="\(.*\)"/\1/p' /etc/os-release)"
echo " cpu $(getconf _NPROCESSORS_ONLN 2>/dev/null || echo '?') online, load $(cut -d' ' -f1-3 /proc/loadavg)"
# ── the caps first, because they decide what the totals below are worth ──
echo
echo "caps"
cap=$(cgroup_cap_mb)
if [ -n "$cap" ]; then
cur=$(cgroup_used_mb)
echo " cgroup ${cap} MB (${CG_VERSION}, ${CG_CUR_FILE##*/} says ${cur:-?} MB used)"
echo " ! /proc/meminfo below describes the HOST, not this cgroup."
echo " $(mb MemTotal) MB total is not yours; ${cap} MB is."
elif [ -n "$CG_VERSION" ]; then
echo " cgroup none (${CG_VERSION} present, no memory limit set)"
else
echo " cgroup no memory controller found"
fi
ul=$(ulimit_v_mb)
if [ -n "$ul" ]; then
echo " ! ulimit -v ${ul} MB — a per-process cap, inherited from your shell"
echo " it stops this process long before the machine runs out"
else
echo " ulimit -v unlimited"
fi
# overcommit_memory=0 is the default heuristic: a large allocation is
# granted on a guess, and the reckoning arrives later as an OOM kill rather
# than as a failed malloc. It is why `push` touches every page it asks for.
local om or_
om=$(cat /proc/sys/vm/overcommit_memory 2>/dev/null || echo '?')
or_=$(cat /proc/sys/vm/overcommit_ratio 2>/dev/null || echo '?')
case "$om" in
0) echo " overcommit 0 heuristic — allocations are granted on a guess," ;;
1) echo " overcommit 1 always — every allocation succeeds; the OOM killer is the only limit," ;;
2) echo " overcommit 2 strict (ratio ${or_}%) — allocation fails honestly instead of killing later," ;;
*) echo " overcommit ${om}" ;;
esac
[ "$om" != "?" ] && echo " so RSS is the number to trust, not what a process asked for"
# ── what it says it has ──
total=$(mb MemTotal); avail=$(avail_meminfo_mb)
swap_total=$(mb SwapTotal); swap_free=$(mb SwapFree)
echo
echo "memory"
echo " total ${total} MB"
echo " available ${avail} MB"
echo " swap ${swap_total} MB ($(( swap_total - swap_free )) MB used)"
if [ "$swap_total" -eq 0 ]; then
echo " - no swap: this box has no cushion. It goes from fine to OOM-killed"
echo " with nothing in between, which is the abrupt failure you get in a VM."
fi
# postgres puts its shared buffers in /dev/shm. Docker's default is 64 MB,
# and the resulting failure names neither shm nor the size.
if [ -d /dev/shm ]; then
local shm; shm=$(df -Pm /dev/shm 2>/dev/null | awk 'NR==2{print $2}')
if [ -n "$shm" ]; then
if [ "$shm" -le 64 ]; then
echo " ! /dev/shm ${shm} MB — postgres puts shared memory here and 64 MB"
echo " is docker's default. Raise it with --shm-size when postgres fails."
else
echo " /dev/shm ${shm} MB"
fi
fi
fi
echo
echo "disk"
local d
for d in / /tmp /var/lib/docker; do
[ -d "$d" ] || continue
df -Pm "$d" 2>/dev/null | awk -v p="$d" 'NR==2{printf " %-12s %s MB free of %s MB\n", p, $4, $2}'
done
# kind and Tilt both watch large trees, and the failure mode is silent:
# they simply stop noticing file changes. Cheap to report while we are here.
local w i
w=$(cat /proc/sys/fs/inotify/max_user_watches 2>/dev/null || echo 0)
i=$(cat /proc/sys/fs/inotify/max_user_instances 2>/dev/null || echo 0)
echo
echo "tooling"
echo " inotify watches=$w instances=$i"
if [ "$w" -lt 524288 ] || [ "$i" -lt 512 ]; then
echo " ! low — anything watching files will silently stop seeing changes"
fi
if ! command -v docker >/dev/null 2>&1; then
if [ -S /var/run/docker.sock ]; then
echo " docker socket present, no cli"
else
echo " docker not installed"
fi
elif docker info >/dev/null 2>&1; then
local n
n=$(docker ps -q 2>/dev/null | wc -l)
echo " docker $(docker version --format '{{.Server.Version}}' 2>/dev/null), ${n} container(s) running"
else
echo " ! docker cli present but the daemon is unreachable"
fi
# WSL keeps its cap on the Windows side, in a file this shell can read but
# not usefully apply — the change costs a full VM restart. Report it, and
# report the commonest mistake, which is editing it and not restarting.
if is_wsl; then
local cfg conf conf_mb n
cfg=$(wslconfig_path)
echo
echo "wsl"
if [ -z "$cfg" ]; then
echo " ! cannot tell which Windows profile owns .wslconfig"
else
echo " config $cfg"
conf=$(configured_memory "$cfg")
if [ -n "$conf" ]; then
conf_mb=$(to_mb "$conf")
echo " configured $conf (${conf_mb} MB), booted ${total} MB"
# The VM reports a little less than allocated; 15% covers the
# kernel without calling every healthy machine a mismatch.
if [ -n "$conf_mb" ] && [ "$total" -lt $(( conf_mb * 85 / 100 )) ]; then
echo " ! configured ${conf_mb} MB but booted ${total} MB — not applied yet."
echo " From a WINDOWS terminal: wsl --shutdown then start the distro again."
fi
else
echo " configured no memory= set (WSL defaults to 50% of host RAM, or 8 GB,"
echo " whichever is less). To raise it, add on the Windows side:"
echo " [wsl2]"
echo " memory=8GB"
echo " then from a WINDOWS terminal: wsl --shutdown"
fi
n=$(ls "$cfg".*.bak 2>/dev/null | wc -l)
if [ "$n" -gt 0 ]; then
echo " backups $n (newest: $(ls -t "$cfg".*.bak 2>/dev/null | head -1))"
fi
fi
else
echo
echo " - native linux: no VM allocation to raise. If memory is tight the levers"
echo " are freeing something or adding swap."
fi
echo
echo "effective ceiling $(effective_ceiling_mb) MB"
echo " the lowest of MemTotal, the cgroup cap and ulimit -v. What the box"
echo " claims. 'push' measures what it will actually hand over."
[ "$avail" -lt $(( total / 5 )) ] && { echo; hogs; }
return 0
}
# ── .wslconfig ─────────────────────────────────────────────────────────────
require_wsl() {
if ! is_wsl; then
echo "$1 acts on .wslconfig, which only exists under WSL." >&2
echo "This is native Linux — there is no VM allocation to save or roll back." >&2
echo "Use 'status' to see what the machine actually has." >&2
exit 1
fi
}
# backup and restore act on the file, so unlike status they must not guess.
wslconfig_required() {
local cfg; cfg=$(wslconfig_required)
if [ -z "$cfg" ]; then
echo "cannot tell which Windows profile owns .wslconfig. Candidates:" >&2
ls -d /mnt/c/Users/*/ 2>/dev/null \
| grep -viE "/(All Users|Default|Default User|Public)/$" | sed "s/^/ /" >&2
exit 1
fi
echo "$cfg"
}
configured_memory() {
[ -r "$1" ] || { echo ""; return; }
sed -n 's/^[[:space:]]*memory[[:space:]]*=[[:space:]]*//p' "$1" | tail -1 | tr -d '[:space:]'
}
# "9GB" / "8192MB" / "9G" -> MB, so it can be compared with /proc/meminfo.
to_mb() {
local v="${1^^}" n
n=$(echo "$v" | tr -dc '0-9')
[ -n "$n" ] || { echo ""; return; }
case "$v" in
*GB|*G) echo $(( n * 1024 )) ;;
*MB|*M) echo "$n" ;;
*) echo $(( n / 1024 / 1024 )) ;;
esac
}
backup() {
require_wsl backup
local cfg dest
cfg=$(wslconfig_required)
[ -r "$cfg" ] || { echo "nothing to back up: $cfg does not exist" >&2; exit 1; }
# Timestamped and never overwritten: a backup that can destroy itself on a
# second run is not a backup.
dest="${cfg}.$(date +%Y%m%d-%H%M%S).bak"
cp "$cfg" "$dest"
echo "backed up $dest"
echo
echo "Edit $cfg by hand, then from a WINDOWS terminal: wsl --shutdown"
}
restore() {
require_wsl restore
local cfg newest count
cfg=$(wslconfig_required)
newest=$(ls -t "$cfg".*.bak 2>/dev/null | head -1 || true)
[ -n "$newest" ] || { echo "no backups found beside $cfg" >&2; exit 1; }
echo "restoring $newest"
echo " -> $cfg"
echo
# Newest is the right default — undo the last edit — but if you backed up
# *after* editing, the state you want is older. Show the rest so a no-op
# restore is obviously a no-op rather than a mystery.
count=$(ls "$cfg".*.bak 2>/dev/null | wc -l)
if [ "$count" -gt 1 ]; then
echo "$count backups exist, newest first:"
ls -t "$cfg".*.bak | sed 's/^/ /'
echo " (restoring the newest; copy another by hand to pick an older one)"
echo
fi
if [ -r "$cfg" ]; then
echo "what changes:"
if diff "$cfg" "$newest" > /tmp/mem.diff 2>&1 && [ ! -s /tmp/mem.diff ]; then
echo " nothing — that backup is identical to the current config"
else
sed 's/^/ /' /tmp/mem.diff
fi
rm -f /tmp/mem.diff
echo
fi
printf "proceed? [y/N] "
read -r reply
case "$reply" in
y|Y|yes|Yes) ;;
*) echo "left alone"; return 0 ;;
esac
cp "$newest" "$cfg"
echo "restored. From a WINDOWS terminal: wsl --shutdown"
}
# ── push ───────────────────────────────────────────────────────────────────
STATE=""
CHILD=""
cleanup() {
if [ -n "$CHILD" ] && kill -0 "$CHILD" 2>/dev/null; then
kill -KILL "$CHILD" 2>/dev/null || true
wait "$CHILD" 2>/dev/null || true
fi
[ -n "$STATE" ] && rm -f "$STATE"
return 0
}
# The child allocates and stops itself; the parent only watches. That split is
# the point: under --to-oom the allocating process is expected to be killed, and
# something has to survive to say how far it got.
allocator() {
# Raise our own OOM score to the maximum so the kernel picks THIS process
# first. Raising needs no privilege (only lowering does). Without it, the
# kernel is free to choose your shell, your ssh session or dockerd — on a
# box you are still using, that is not an acceptable coin toss.
echo 1000 > "/proc/$BASHPID/oom_score_adj" 2>/dev/null || true
local arr=() held=0 i=0 rss swapped avail first_swap=0
local bytes=$((STEP_MB * 1024 * 1024))
local swap_used_start
swap_used_start=$(( $(mb SwapTotal) - $(mb SwapFree) ))
while :; do
# Written STRAIGHT INTO the array element. The obvious spelling —
# build one chunk and `arr+=("$chunk")` — costs three copies per step,
# not one: the template stays resident, expanding "$chunk" makes a
# temporary word, and the append makes the element. A 128 MB step then
# needs 384 MB transiently, and on a small box it is killed on the
# first append while reporting a third of the true ceiling.
#
# printf -v into a subscript also means every page is written, so it is
# resident rather than merely promised — the only kind of allocation
# that measures anything under heuristic overcommit.
printf -v "arr[$i]" '%*s' "$bytes" ''
i=$((i + 1)); held=$((held + STEP_MB))
rss=$(awk '/^VmRSS:/{print int($2/1024)}' "/proc/$BASHPID/status" 2>/dev/null || echo 0)
avail=$(headroom_mb)
swapped=$(( $(mb SwapTotal) - $(mb SwapFree) - swap_used_start ))
[ "$swapped" -lt 0 ] && swapped=0
printf '%8s MB held rss %7s MB headroom %7s MB swap +%s MB\n' \
"$held" "$rss" "$avail" "$swapped"
printf '%s %s %s %s\n' "$held" "$rss" "$avail" "$swapped" >> "$STATE"
# Worth calling out separately from the ceiling: this is where the box
# stops being fast and starts being unusable, which for a scheduler is
# a different and earlier problem than being killed.
if [ "$swapped" -gt 0 ] && [ "$first_swap" -eq 0 ]; then
first_swap=$held
echo " - first swap page at ${held} MB — past here it works but crawls"
echo "swapat $held" >> "$STATE"
fi
if [ -n "$TO_MB" ] && [ "$held" -ge "$TO_MB" ]; then
echo "stop reached-the-cap" >> "$STATE"; return 0
fi
if [ "$TO_OOM" = no ] && [ "$avail" -lt "$FLOOR_MB" ]; then
echo "stop floor" >> "$STATE"; return 0
fi
done
}
push() {
local total ceiling rc=0 last held rss swapat stop
total=$(mb MemTotal)
ceiling=$(effective_ceiling_mb)
# A step is worth about a sixty-fourth of the ceiling: enough resolution to
# find the edge, few enough lines to read, and small enough that the
# transient cost of one allocation never dominates a small box. A fixed
# size cannot do all three — 128 MB is fine on 16 GB and absurd on 512 MB.
if [ "$STEP_EXPLICIT" = no ]; then
STEP_MB=$(( ceiling / 64 ))
[ "$STEP_MB" -lt 4 ] && STEP_MB=4
[ "$STEP_MB" -gt 256 ] && STEP_MB=256
fi
# Stop with a cushion rather than riding it to the kill. How big a cushion
# depends on what it is protecting. Under a cgroup cap, running out kills
# only this container's own processes, so it need cover no more than the
# shell that prints the result — and a 512 MB cushion on a 1 GB box would
# halve the answer. On a host there is everything else to protect, and the
# OOM killer does not promise to pick the process that caused the problem.
if [ -n "$(cgroup_cap_mb)" ]; then FLOOR_MB=64; else FLOOR_MB=512; fi
[ $(( ceiling / 20 )) -gt "$FLOOR_MB" ] && FLOOR_MB=$(( ceiling / 20 ))
STATE=$(mktemp "${TMPDIR:-/tmp}/rigmini.XXXXXX")
trap cleanup EXIT
# INT kills the child and lets the summary below print anyway, so an
# impatient Ctrl-C still tells you how far it got — and, more importantly,
# still gives the memory back.
trap 'echo; echo " interrupted"; echo "stop interrupted" >> "$STATE"; [ -n "$CHILD" ] && kill -KILL "$CHILD" 2>/dev/null || true' INT
echo "push"
echo " step ${STEP_MB} MB per allocation, every page touched"
echo " ceiling ${ceiling} MB claimed"
if [ -n "$TO_MB" ]; then
echo " stopping at ${TO_MB} MB (--to)"
elif [ "$TO_OOM" = yes ]; then
echo " ! stopping only when the kernel stops it (--to-oom)"
echo " the allocating child is marked as the preferred OOM victim,"
echo " but nothing about an OOM kill is entirely polite. Not on a box"
echo " running anything you mind losing."
else
echo " stopping when headroom drops below ${FLOOR_MB} MB"
fi
echo
allocator &
CHILD=$!
wait "$CHILD" || rc=$?
CHILD=""
trap - INT
last=$(grep -E '^[0-9]' "$STATE" 2>/dev/null | tail -1 || true)
held=$(echo "$last" | awk '{print $1}')
rss=$(echo "$last" | awk '{print $2}')
swapat=$(awk '/^swapat/{print $2}' "$STATE" 2>/dev/null | head -1 || true)
stop=$(awk '/^stop/{print $2}' "$STATE" 2>/dev/null | head -1 || true)
echo
if [ -z "$held" ]; then
echo " ! nothing was allocated. Even one ${STEP_MB} MB chunk failed —"
echo " try a smaller --step, or check ulimit -v in 'status'."
return 1
fi
echo " reached ${rss:-$held} MB resident"
[ -n "$swapat" ] && echo " swapping from ${swapat} MB"
case "$stop" in
reached-the-cap)
echo " outcome stopped at the --to cap, not at a limit."
echo " The box held ${TO_MB} MB without complaint; there is more." ;;
floor)
echo " outcome stopped with a cushion intact, by choice."
echo " The real ceiling is higher — --to-oom finds it, at the"
echo " cost of an actual OOM kill." ;;
interrupted)
echo " outcome interrupted at ${rss:-$held} MB — where you stopped it,"
echo " not where the box did." ;;
*)
# No stop line means the child did not decide to stop: it was ended.
if [ "$rc" -ge 128 ]; then
echo " outcome the child was killed (signal $((rc - 128))) at ${rss:-$held} MB."
elif [ "$rc" -ne 0 ]; then
echo " outcome the allocation failed at ${rss:-$held} MB (exit ${rc})."
echo " bash could not get the next chunk — an honest malloc"
echo " failure rather than a kill. That is the strict-overcommit"
echo " or ulimit path."
else
echo " outcome ended at ${rss:-$held} MB."
fi
local ev
ev=$(dmesg 2>/dev/null | tail -80 | grep -iE 'oom-kill|killed process' | tail -1 || true)
if [ -n "$ev" ]; then
echo " kernel ${ev#*] }"
else
echo " - dmesg is unreadable here (dmesg_restrict, or no privilege),"
echo " so the kill cannot be confirmed from this side. The number stands."
fi ;;
esac
# The gap between the claim and the measurement is the finding — but only
# when the BOX chose where to stop. An empty $stop means the child was ended
# rather than deciding to end; anything else (--to, the floor) is a stop we
# asked for, and flagging those as short of the ceiling would put a warning
# on every deliberately small run.
local got="${rss:-$held}"
echo
if [ -z "$stop" ] && [ "$got" -lt $(( ceiling * 70 / 100 )) ]; then
echo " ! claimed ${ceiling} MB, gave up ${got} MB — under 70% of it."
echo " Something is taking the difference. 'status' names the candidates:"
echo " a cgroup cap, ulimit -v, or memory already resident."
fi
return 0
}
# ── all ────────────────────────────────────────────────────────────────────
all() {
status
echo
echo "────────────────────────────────────────────────────────────"
echo
push
local got budget_mb ceiling
load_config
if [ -n "$BUDGET_GB" ]; then
budget_mb=$(( BUDGET_GB * 1024 ))
else
budget_mb=$(( NODES * NODE_MB ))
fi
ceiling=$(effective_ceiling_mb)
got=$(grep -E '^[0-9]' "$STATE" 2>/dev/null | tail -1 | awk '{print $2}' || true)
[ -n "$got" ] || got=0
echo
echo "verdict"
if [ -n "$BUDGET_GB" ]; then
echo " budget ${budget_mb} MB (--budget)"
else
# rig's own figure for this profile: nodes times what one node costs.
# Addons carry no memory figure in rig yet, so this is the cluster alone
# and whatever you deploy comes on top. --budget once you know that too.
echo " budget ${budget_mb} MB — profile ${PROFILE_NAME}: ${NODES} node(s) x ${NODE_MB} MB,"
echo " the cluster alone; your workload comes on top (--budget GB)"
fi
echo " measured ${got} MB handed over"
if [ "$got" -ge "$budget_mb" ]; then
echo " fits, with $(( got - budget_mb )) MB spare."
if [ "$got" -lt $(( budget_mb * 130 / 100 )) ]; then
echo " - under 30% spare is thin once a workload runs on top: memory use"
echo " is spiky, and the spikes are what get killed."
fi
else
echo " ! short by $(( budget_mb - got )) MB."
if [ "$ceiling" -ge "$budget_mb" ]; then
echo " The box CLAIMS enough (${ceiling} MB) but did not deliver it."
echo " Free something, or read the caps section again."
else
echo " The box does not have it to give. A bigger machine, or a profile"
echo " with fewer nodes."
fi
fi
return 0
}
# ── main ───────────────────────────────────────────────────────────────────
parse_flags() {
while [ $# -gt 0 ]; do
case "$1" in
--to) TO_MB=$(( ${2:?--to needs a value in GB} * 1024 )); shift 2 ;;
--to-mb) TO_MB="${2:?--to-mb needs a value in MB}"; shift 2 ;;
--step) STEP_MB="${2:?--step needs a value in MB}"; STEP_EXPLICIT=yes; shift 2 ;;
--to-oom) TO_OOM=yes; shift ;;
--budget) BUDGET_GB="${2:?--budget needs a value in GB}"; BUDGET_EXPLICIT=yes; shift 2 ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
if [ "$TO_OOM" = yes ] && [ -n "$TO_MB" ]; then
echo "--to and --to-oom contradict each other: one stops early, the other" >&2
echo "refuses to stop at all. Pick one." >&2
exit 1
fi
return 0
}
require_linux
find_cgroup
cmd="${1:-status}"
[ $# -gt 0 ] && shift
case "$cmd" in
status) parse_flags "$@"; status ;;
push) parse_flags "$@"; push ;;
all) parse_flags "$@"; all ;;
backup) backup ;;
restore) restore ;;
*) echo "usage: $0 [status|push|all|backup|restore]" >&2
echo " push [--to GB] [--to-mb MB] [--step MB] [--to-oom]" >&2
echo " all [--budget GB]" >&2
exit 1 ;;
esac

View File

@@ -1,313 +0,0 @@
#!/usr/bin/env bash
# Create a disposable Linux environment to validate the installer from a
# genuinely clean slate — one that can be thrown away without touching the
# environment you actually work in.
#
# This is the ONLY host-aware file in the tree. Everything else needs just a
# Linux with Docker, which is what keeps other host types a later addition
# rather than a rewrite.
#
# On WSL it creates a second distro. There is no .bat and no PowerShell script:
# wsl.exe is callable from inside WSL, and wslpath converts the paths it wants.
# A machine with no WSL at all needs `wsl --install` run once by hand first —
# scripting a reboot-requiring Windows feature install is not worth it.
#
# Docker: borrowed by default, never installed twice
# --------------------------------------------------
# WSL2 distros share one kernel and one network stack, so two dockerd instances
# contend over docker0 and iptables and can disturb the daemon you depend on.
# (That is why Docker Desktop runs one daemon in a dedicated distro and shares
# its socket rather than installing one per distro.)
#
# REUSE_DOCKER=1 (default) borrow the host distro's daemon over /mnt/wsl.
# Nothing is installed; nothing can conflict.
# Requires `ctrl/dockerhost.sh share` once on the
# distro that owns Docker.
# REUSE_DOCKER=0 install a second daemon in the new distro. Only
# if you specifically want to test a from-scratch
# Docker install, and not on a machine you need.
#
# Borrowing is also the more honest test: rig never installs Docker anyway — it
# is the documented prerequisite — so a clean box does not need its own to
# exercise everything rig actually does.
#
# Usage: newbox.sh create | destroy [--purge] | status | shell
set -euo pipefail
cd "$(dirname "$0")"
source ./lib/config.sh
load_config
REPO="$(cd .. && pwd)"
# The distro is named after this environment, and that derived name is the ONLY
# thing this script will ever destroy. See guard_name().
BOX="${BOX:-${CLUSTER}box}"
BOX_USER="${BOX_USER:-dev}"
# Borrow the host distro's Docker rather than installing a second daemon.
REUSE_DOCKER="${REUSE_DOCKER:-1}"
SHARED_SOCK=/mnt/wsl/shared-docker/docker.sock
WSL_EXE=/mnt/c/Windows/System32/wsl.exe
# ── host detection ─────────────────────────────────────────────────────────
require_wsl() {
if ! grep -qi microsoft /proc/version 2>/dev/null; then
cat >&2 <<'EOF'
newbox is WSL-only for now.
On native Linux you do not need it: rig already isolates environments by
directory (own cluster, context, images and port block), so a second copy in a
second directory is the clean slate. To validate the installer itself against a
bare system, run the wizard against a stock Debian container instead.
EOF
exit 1
fi
if [ ! -x "$WSL_EXE" ]; then
echo "wsl.exe not found at $WSL_EXE" >&2
exit 1
fi
}
wsl_list() { "$WSL_EXE" -l -q 2>/dev/null | tr -d '\0\r'; }
box_exists() { wsl_list | grep -qx "$BOX"; }
# `wsl --unregister` permanently deletes a distro's filesystem. The whole safety
# story is this function: only the name derived from this directory can ever be
# a target, so a typo or a stray argument cannot destroy the distro you work in.
guard_name() {
local derived="${CLUSTER}box"
if [ "$BOX" != "$derived" ]; then
echo "refusing: BOX='$BOX' is not the name derived from this directory ('$derived')." >&2
echo "That guard exists because --unregister is irreversible." >&2
exit 1
fi
if [ -z "$CLUSTER" ] || [ "$BOX" = "box" ]; then
echo "refusing: empty environment name" >&2
exit 1
fi
}
# ── create ─────────────────────────────────────────────────────────────────
rootfs_path() {
local win_home; win_home=$(wslpath "$("$WSL_EXE" -d "$(wsl_list | head -1)" -e printf '%s' "$USERPROFILE" 2>/dev/null || true)" 2>/dev/null || true)
# Simpler and reliable: use the current user's Windows home via /mnt/c.
ls -d /mnt/c/Users/*/ 2>/dev/null | grep -viE '/(All Users|Default|Default User|Public)/$' | head -1
}
build_rootfs() {
local tar="$1"
if [ -f "$tar" ]; then
echo " rootfs cached: $(basename "$tar")"
return
fi
echo " exporting a stock Debian rootfs (cached for next time)"
local cid; cid=$(docker create debian:trixie-slim)
docker export "$cid" > "$tar"
docker rm -f "$cid" >/dev/null
}
provision() {
echo " provisioning (root)"
local hosts_block
hosts_block=$(CLUSTER="$CLUSTER" HTTP_PORT="$HTTP_PORT" \
envsubst < ./hosts.tmpl 2>/dev/null || sed "s/\${CLUSTER}/$CLUSTER/g" ./hosts.tmpl)
# Piped as stdin rather than a second script file, the same shape as any
# remote provisioning heredoc. Everything here is idempotent so a failed run
# can simply be repeated.
"$WSL_EXE" -d "$BOX" -u root -- bash -s <<PROVISION
set -euo pipefail
export DEBIAN_FRONTEND=noninteractive
apt-get update -qq
apt-get install -y -qq ca-certificates curl gnupg sudo >/dev/null
if [ "$REUSE_DOCKER" = "1" ]; then
# Borrow the host distro's daemon: CLI only, no dockerd, nothing to
# conflict with. The GID must match the owner's or the shared socket is
# unreadable here even though it is visible.
install -m 0755 -d /etc/apt/keyrings
if [ ! -f /etc/apt/keyrings/docker.asc ]; then
curl -fsSL https://download.docker.com/linux/debian/gpg -o /etc/apt/keyrings/docker.asc
chmod a+r /etc/apt/keyrings/docker.asc
fi
echo "deb [arch=\$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian \$(. /etc/os-release && echo \$VERSION_CODENAME) stable" \
> /etc/apt/sources.list.d/docker.list
apt-get update -qq
apt-get install -y -qq docker-ce-cli >/dev/null
echo "export DOCKER_HOST=unix://$SHARED_SOCK" > /etc/profile.d/rig-docker-host.sh
if [ -f /mnt/wsl/shared-docker/OWNER ]; then
gid=\$(awk '/docker gid:/ {print \$3}' /mnt/wsl/shared-docker/OWNER)
if [ -n "\$gid" ]; then
getent group docker >/dev/null && groupmod -g "\$gid" docker || groupadd -g "\$gid" docker
fi
fi
else
# A second daemon. Only when deliberately testing a from-scratch install.
install -m 0755 -d /etc/apt/keyrings
if [ ! -f /etc/apt/keyrings/docker.asc ]; then
curl -fsSL https://download.docker.com/linux/debian/gpg -o /etc/apt/keyrings/docker.asc
chmod a+r /etc/apt/keyrings/docker.asc
fi
echo "deb [arch=\$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian \$(. /etc/os-release && echo \$VERSION_CODENAME) stable" \
> /etc/apt/sources.list.d/docker.list
apt-get update -qq
apt-get install -y -qq docker-ce docker-ce-cli containerd.io >/dev/null
fi
id -u "$BOX_USER" >/dev/null 2>&1 || useradd -m -s /bin/bash "$BOX_USER"
usermod -aG sudo,docker "$BOX_USER"
echo "$BOX_USER ALL=(ALL) NOPASSWD:ALL" > /etc/sudoers.d/90-$BOX_USER
chmod 0440 /etc/sudoers.d/90-$BOX_USER
# systemd is off by default in WSL, and Docker needs it. Takes effect on the
# next start of this distro, which is why create() terminates it below.
cat > /etc/wsl.conf <<WSLCONF
[boot]
systemd=true
[user]
default=$BOX_USER
WSLCONF
# The default inotify limits are low enough that file watching silently stops
# working — no error, changes just stop being noticed. Fix it before it bites.
cat > /etc/sysctl.d/99-rig.conf <<SYSCTL
fs.inotify.max_user_watches=524288
fs.inotify.max_user_instances=512
SYSCTL
if ! grep -q 'rig environment' /etc/hosts 2>/dev/null; then
{ echo ""; echo "# rig environment"; cat <<'HOSTS'
$hosts_block
HOSTS
} >> /etc/hosts
fi
touch /etc/rig-provisioned
PROVISION
}
create() {
require_wsl
guard_name
local winhome; winhome=$(rootfs_path)
[ -n "$winhome" ] || { echo "could not locate the Windows user directory" >&2; exit 1; }
local tar="${winhome}rig-rootfs.tar"
local installdir="${winhome}WSL/${BOX}"
echo "creating '$BOX'"
if [ "$REUSE_DOCKER" = "1" ]; then
echo " docker: borrowing the host distro's daemon (nothing installed)"
if [ ! -S "$SHARED_SOCK" ]; then
echo
echo " No shared socket yet. In the distro that owns Docker, run once:"
echo " sudo bash ctrl/dockerhost.sh share"
echo " That adds one systemd drop-in and nothing else; undo with 'unshare'."
echo " Continuing — the box will be created, but Docker won't work in it"
echo " until you do that."
fi
else
echo
echo " REUSE_DOCKER=0: installing a SECOND Docker daemon."
echo " WSL distros share a network stack, so this can disturb Docker in"
echo " the distro you work in. Ctrl-C now if that is a bad trade today."
echo
sleep 4
fi
echo
if box_exists; then
echo " distro already registered"
else
build_rootfs "$tar"
mkdir -p "$installdir"
"$WSL_EXE" --import "$BOX" "$(wslpath -w "$installdir")" "$(wslpath -w "$tar")" --version 2
fi
# Resumable: a partially-created box is finished rather than restarted.
if "$WSL_EXE" -d "$BOX" -u root -- test -f /etc/rig-provisioned 2>/dev/null; then
echo " already provisioned"
else
provision
echo " restarting the distro so systemd and group membership apply"
"$WSL_EXE" --terminate "$BOX" # ONLY this distro; never --shutdown
fi
echo " copying rig in"
tar c -C "$REPO" --exclude=def --exclude=.git --exclude=ctrl/.env . \
| "$WSL_EXE" -d "$BOX" -u "$BOX_USER" -- bash -lc "mkdir -p ~/rig && tar x -C ~/rig"
echo
echo " docker: $("$WSL_EXE" -d "$BOX" -u "$BOX_USER" -- bash -lc 'systemctl is-active docker 2>/dev/null || echo inactive')"
echo
echo "next:"
echo " make newbox shell # a shell inside it"
echo " then: cd ~/rig && make station && make deps && make cluster up"
echo
echo "For a browser on Windows to resolve the hostnames, paste this into"
echo "C:\\Windows\\System32\\drivers\\etc\\hosts (it has no wildcard support):"
CLUSTER="$CLUSTER" envsubst < ./hosts.tmpl 2>/dev/null | grep -v '^#' | grep -v '^$' | sed 's/^/ /'
}
# ── the rest ───────────────────────────────────────────────────────────────
destroy() {
require_wsl
guard_name
if ! box_exists; then
echo "no distro '$BOX' to remove"
else
echo "about to PERMANENTLY delete the distro '$BOX' and its filesystem."
"$WSL_EXE" --terminate "$BOX" 2>/dev/null || true
"$WSL_EXE" --unregister "$BOX"
echo " unregistered"
fi
local winhome; winhome=$(rootfs_path)
rm -rf "${winhome}WSL/${BOX}" 2>/dev/null || true
if [ "${1:-}" = "--purge" ]; then
rm -f "${winhome}rig-rootfs.tar"
echo " cached rootfs removed"
fi
}
status() {
require_wsl
echo "environment $CLUSTER"
echo "distro $BOX"
if box_exists; then
echo "registered yes"
echo "provisioned $("$WSL_EXE" -d "$BOX" -u root -- test -f /etc/rig-provisioned 2>/dev/null && echo yes || echo no)"
echo "docker $("$WSL_EXE" -d "$BOX" -u root -- bash -lc 'systemctl is-active docker 2>/dev/null' || echo unknown)"
echo "rig copied $("$WSL_EXE" -d "$BOX" -u "$BOX_USER" -- bash -lc 'test -f ~/rig/Makefile && echo yes || echo no' 2>/dev/null)"
else
echo "registered no"
fi
echo
echo "all distros (this one is never touched unless it is '$BOX'):"
wsl_list | sed 's/^/ /'
}
shell() {
require_wsl
box_exists || { echo "no distro '$BOX' — run 'make newbox' first" >&2; exit 1; }
"$WSL_EXE" -d "$BOX" -u "$BOX_USER" --cd '~'
}
case "${1:-status}" in
create) create ;;
destroy) shift; destroy "${1:-}" ;;
status) status ;;
shell) shell ;;
*) echo "usage: $0 [create|destroy [--purge]|status|shell]" >&2; exit 1 ;;
esac

View File

@@ -19,7 +19,7 @@
# written into ctrl/.env, so it becomes pinned, visible and editable rather than
# a number that appears from nowhere. Anything already in ctrl/.env wins.
#
# Usage: ports.sh show | derive | persist
# Usage: ports.sh show | active | derive | persist
set -euo pipefail
cd "$(dirname "$0")"
@@ -38,6 +38,33 @@ derive() {
DERIVED_REGISTRY=$((base + 3))
}
# The resolved facts a consumer outside bash needs, machine-readable:
#
# CLUSTER KUBECONTEXT HTTP HTTPS TILT REGISTRY MANIFESTS_DIR
#
# Identity and ports together, because they are one fact set — the header above
# says so: both derive from the directory name so that copies never collide. A
# consumer needs all of them or none, and fetching them separately is how two
# end up disagreeing. MANIFESTS_DIR rides along because the one consumer that
# needs the addressing is the one that needs to know what to deploy.
#
# Space-separated, so MANIFESTS_DIR must not contain spaces. Everything else in
# rig already assumes that of paths — kind, docker and kubectl all do.
#
# `derive` answers a DIFFERENT question — what the directory name alone implies
# — and deliberately ignores ctrl/.env. Configuring anything from it would
# silently contradict this file's own rule that "anything already in ctrl/.env
# wins". `active` is what anything downstream should read.
#
# Why this exists at all: the cluster name is not the bare directory name.
# default_cluster_name() lowercases it and replaces every character outside
# [a-z0-9-], because it has to be a DNS label. Re-deriving that in another
# language is how a copy in `My_Project/` ends up guarding the wrong context.
active() {
load_config
echo "$CLUSTER $KUBECONTEXT $HTTP_PORT $HTTPS_PORT $TILT_PORT $REGISTRY_PORT $MANIFESTS_DIR"
}
show() {
derive
echo "environment $CLUSTER"
@@ -98,6 +125,7 @@ persist() {
case "${1:-show}" in
show) show ;;
derive) derive; echo "$DERIVED_HTTP $DERIVED_HTTPS $DERIVED_TILT $DERIVED_REGISTRY" ;;
active) active ;;
persist) persist ;;
*) echo "usage: $0 [show|derive|persist]" >&2; exit 1 ;;
*) echo "usage: $0 [show|active|derive|persist]" >&2; exit 1 ;;
esac

View File

@@ -43,7 +43,7 @@ K="kubectl --context ${KUBECONTEXT}"
# 2. every kind node's containerd — nodes do NOT inherit host trust
# 3. anything doing HTTPS from inside the cluster, in its own trust store
#
# We handle (2) here because it's ours to handle. (1) is reported by station.sh
# We handle (2) here because it's ours to handle. (1) is reported by check.sh
# since it needs root. (3) belongs to the workload.
install_ca_into_nodes() {
[ -n "${REGISTRY_CA_FILE:-}" ] || return 0

272
rig/ctrl/selftest.sh Executable file
View File

@@ -0,0 +1,272 @@
#!/usr/bin/env bash
# What rig has settled, written down as assertions.
#
# These are documentation that runs. Each check is ONE decision that has already
# been made, with the reason above it — not coverage, and deliberately not an
# exhaustive sweep of use cases. rig's own index says a rule without its reason
# gets overridden the first time it is inconvenient; a rule nobody can restate
# is worse. So the test says what was decided, and failing it should read as
# "you are about to undo this" rather than "something broke".
#
# Scope, on purpose:
# - no cluster, no docker, no network. It must be cheap enough to actually run.
# - it asserts about RIG. `make check` asserts about the MACHINE and never
# fails; this exits 1, the way `make standalone check` does.
# - what actually deploys is not testable here. `tilt ci` stays a manual step.
#
# Usage: make selftest (or: bash ctrl/selftest.sh)
set -uo pipefail # NOT -e: one failing check must not abort the rest
cd "$(dirname "$0")"
source ./lib/config.sh
rc=0
passed=0
check() { # name, expected, actual
if [ "$2" = "$3" ]; then
printf ' ok %s\n' "$1"
passed=$((passed + 1))
else
printf ' FAIL %s\n expected: %s\n got: %s\n' "$1" "$2" "$3"
rc=1
fi
}
note() { printf '\n%s\n' "$1"; }
# Resolve one key the way every rig script does, in a clean shell so the
# caller's exported value is the only thing in play.
resolved() {
bash -c 'source ./lib/config.sh; load_config >/dev/null 2>&1; printf "%s" "${!1}"' _ "$1"
}
note "rig needs no profile"
# rig assumes no configuration. A profile is an overlay on built-in defaults, so
# a rig with no env.d/ at all must resolve, report, and still generate a kit —
# and naming a profile that does not exist must still be an error, because a
# typo that silently fell back to the defaults would be worse than a failure.
NP="$(mktemp -d)"
cp -r .. "$NP/rig"; rm -rf "$NP/rig/ctrl/env.d"; sed -i '/^PROFILE=/d' "$NP/rig/ctrl/.env" 2>/dev/null
check "no env.d: config resolves" "default" \
"$(cd "$NP/rig/ctrl" && bash -c 'source ./lib/config.sh; load_config >/dev/null && echo "$PROFILE_NAME"' 2>&1)"
check "no env.d: the k8s version comes from the pins" "yes" \
"$(cd "$NP/rig/ctrl" && bash -c 'source ./lib/config.sh; load_config >/dev/null && [ -n "$NODE_IMAGE" ] && echo yes' 2>&1)"
check "no env.d: ports.sh active works" "7" \
"$(cd "$NP/rig/ctrl" && bash ports.sh active 2>/dev/null | wc -w)"
check "no env.d: a kit is generated for the defaults" "yes" \
"$( (cd "$NP/rig/ctrl" && rm -rf ../standalone/*/ && bash standalone.sh write >/dev/null 2>&1) && [ -f "$NP/rig/standalone/default/rigdeps.sh" ] && echo yes || echo no)"
check "a profile that does not exist is still an error" "yes" \
"$( (cd "$NP/rig/ctrl" && PROFILE=no-such-profile bash -c 'source ./lib/config.sh; load_config' >/dev/null 2>&1) && echo no || echo yes)"
rm -rf "$NP"
note "the ports.sh active contract"
# ports.sh active is read POSITIONALLY by two other files — the Makefile takes
# $(word 2) and $(word 5), the Tiltfile takes _facts[0]..[6]. Insert a field in
# the middle and nothing errors: Tilt simply guards on the wrong context or
# binds the wrong port. The field count and order are the contract, so they are
# pinned here rather than left to whoever edits ports.sh next.
FACTS="$(bash ports.sh active)"
check "active: exactly 7 fields" "7" "$(printf '%s' "$FACTS" | wc -w)"
read -r F_CLUSTER F_CTX F_HTTP F_HTTPS F_TILT F_REG F_MANIFESTS <<< "$FACTS"
check "active: field 2 is kind-<cluster>" "kind-$F_CLUSTER" "$F_CTX"
check "active: fields 3-6 are numeric" "yes" \
"$([[ "$F_HTTP$F_HTTPS$F_TILT$F_REG" =~ ^[0-9]+$ ]] && echo yes || echo no)"
check "active: field 7 is a path" "yes" \
"$([ -n "$F_MANIFESTS" ] && [ "${F_MANIFESTS#-}" = "$F_MANIFESTS" ] && echo yes || echo no)"
# derive answers a different question and must keep its own shape: it reports
# what the directory name implies, ignoring ctrl/.env, so nothing should
# configure itself from it.
check "derive: still 4 fields, not 7" "4" "$(bash ports.sh derive | wc -w)"
note "the caller's env beats the files"
# lib/config.sh states one precedence rule: versions.env < env.d/<profile> <
# ctrl/.env < the caller's env. It is enforced by CONFIG_OVERRIDABLE, a
# hand-maintained list — and a key missing from it loses to the file SILENTLY.
# REGISTRY_PORT and MANIFESTS_DIR were both missing on 2026-09-13 and were found
# by accident.
#
# So this loop is generated FROM the list: add a key to CONFIG_OVERRIDABLE and
# this test starts asking about it without anyone remembering to come here.
# Three keys name something that must exist and are validated at load, so they
# get a real alternative rather than a sentinel.
test_value() {
case "$1" in
# Picked from what exists, never named: rig must not need any particular
# profile, template or pinned version to be present for this to run.
PROFILE) config_profiles | head -1 ;;
K8S_VERSION) (set -a; source ./versions.env; compgen -v NODE_IMAGE_v | sort -V | head -1 | sed 's/^NODE_IMAGE_//') ;;
KIND_CONFIG) ls k8s/kind-config*.yaml.tpl 2>/dev/null | sort | head -1 | xargs -r basename ;;
*_PORT) echo "19999" ;;
CLUSTER) echo "selftest-name" ;;
MANIFESTS_DIR) echo "../elsewhere/overlays/dev" ;;
ADDONS) echo "metallb" ;;
*) echo "selftest-sentinel" ;;
esac
}
for key in $CONFIG_OVERRIDABLE; do
[ -n "$key" ] || continue
want="$(test_value "$key")"
if [ -z "$want" ]; then
check "precedence: $key has a test value" "yes" "no — add one to test_value()"
continue
fi
got="$(export "$key=$want"; resolved "$key")"
check "precedence: caller's $key wins" "$want" "$got"
done
note "one derivation, not three"
# The Makefile used to compute the cluster name itself and sed TILT_PORT out of
# ctrl/.env — a second derivation of values lib/config.sh already owns, which
# could disagree with it after `make ports persist`. It now reads ports.sh
# active. Nothing structurally prevents the sed coming back, so the agreement is
# asserted against the real `make -n` output rather than against the source.
# --no-print-directory and a grep, not `tail -1`: run from `make selftest` this
# is a RECURSIVE make, and the "Entering/Leaving directory" lines go to STDOUT.
# tail -1 then reads "make[1]: Leaving directory ..." and both checks below fail
# — but only when invoked through make, never when the script is run directly.
# A test that passes one way and fails the other is worse than no test.
MK="$(cd .. && make --no-print-directory -n tilt 2>/dev/null | grep -m1 'tilt ')"
check "Makefile: --context comes from active" "$F_CTX" \
"$(printf '%s' "$MK" | sed -n 's/.*--context \([^ ]*\).*/\1/p')"
check "Makefile: --port comes from active" "$F_TILT" \
"$(printf '%s' "$MK" | sed -n 's/.*--port \([^ ]*\).*/\1/p')"
note "identity follows the folder, safely"
# The cluster name is NOT the bare directory name: kind needs a DNS label, so
# default_cluster_name lowercases it and replaces everything outside [a-z0-9-].
# Re-deriving that anywhere else is how a copy ends up guarding the wrong
# context — which is exactly why the Tiltfile asks instead of computing.
TMP="$(mktemp -d)"
trap 'rm -rf "$TMP"' EXIT
mkdir -p "$TMP/My_Proj"
cp -r . "$TMP/My_Proj/ctrl"
# A pinned CLUSTER in .env would be an override, not a derivation, and this
# check is about the derivation.
sed -i '/^CLUSTER=/d' "$TMP/My_Proj/ctrl/.env" 2>/dev/null
COPY="$(cd "$TMP/My_Proj/ctrl" && bash ports.sh active)"
check "a dir named My_Proj derives a DNS label" "my-proj" "$(awk '{print $1}' <<< "$COPY")"
check "and a context to match" "kind-my-proj" "$(awk '{print $2}' <<< "$COPY")"
check "a renamed copy gets a DIFFERENT block" "different" \
"$([ "$(awk '{print $3}' <<< "$COPY")" != "$F_HTTP" ] && echo different || echo COLLIDES)"
note "ports are stable across versions"
# Not a change-detector. The block is derived, never stored, so if the
# derivation shifts then every EXISTING environment's ports move underneath it —
# a running cluster keeps its old ports while rig starts reporting new ones, and
# `make ports` stops describing reality. Anchored to three known names.
check "derive_port_base rig" "20310" "$(derive_port_base rig)"
check "derive_port_base foo" "21690" "$(derive_port_base foo)"
check "derive_port_base my-proj" "21030" "$(derive_port_base my-proj)"
note "rig stays standalone"
# rig sits inside a host project's tree but must be copyable straight out of it:
# no imports, no paths, no assumption the host is there. This grep is the whole
# test of that claim, and until now it lived only in prose and in whoever
# remembered to run it.
#
# The pattern is assembled from fragments so this file does not match ITSELF.
# Writing it literally would fail forever; excluding this file instead would put
# a blind spot in the one check that guards the boundary.
HOST_PAT="$(printf '%s' 'sole' 'print' '|\b' 'sp' 'r\b')"
check "no host-project references" "0" \
"$(cd .. && grep -rIl -iE "$HOST_PAT" . --exclude-dir=def 2>/dev/null | wc -l)"
note "the Tiltfile hardcodes nothing"
# Every other Tiltfile on this machine writes its slug in five or six times by
# hand, so a copied project deploys into the original's cluster until someone
# edits all of them. rig's asks ports.sh. A literal kind-<name> here would mean
# that has been undone.
check "no literal kind-<name>" "0" "$(grep -cE "['\"]kind-[a-z0-9]" Tiltfile)"
check "guards on the variable" "1" "$(grep -c 'allow_k8s_contexts(CTX)' Tiltfile)"
check "asks ports.sh for facts" "1" "$(grep -c "local('bash ports.sh active'" Tiltfile)"
note "standalone kits are generated, current, and call only real verbs"
# The kits under standalone/<profile>/ are rig flattened into single files, one
# per profile. A kit left behind by a change to rig is exactly the drift they
# replaced — rigmini.sh once said 2 GB per node long after rig measured 800 MB —
# so a stale kit fails here rather than waiting to be noticed on another machine.
check "every kit matches what rig generates now" "yes" \
"$(bash standalone.sh check >/dev/null 2>&1 && echo yes || echo "no — run make standalone")"
# Each kit's Makefile exists so nothing wrapping these scripts has to GUESS how to
# call them. A generated Makefile once did guess: `rigmini.sh on`, not a verb,
# and a bare `rigdeps.sh` for "check and report", which installs. So every
# target's default verb must be one its script's own dispatch accepts — read
# from that dispatch, not from a list here that could drift from it.
verbs_of() {
sed -n '/^case "\$cmd" in/,/^esac/p' "$1" | grep -oE '^ [a-z]+\)' | tr -d ' )'
}
kits=0
for mk in ../standalone/*/Makefile; do
[ -f "$mk" ] || continue
kit=$(dirname "$mk"); kits=$((kits + 1))
for target in $(grep -oE '^[a-z][a-z-]*:' "$mk" | tr -d ':' | grep -vx help); do
line="$(make --no-print-directory -s -n -f "$mk" "$target" 2>/dev/null | head -1)"
script=$(basename "$(printf '%s' "$line" | awk '{print $2}')")
verb=$(printf '%s' "$line" | awk '{print $NF}')
check "$(basename "$kit"): make $target -> $script $verb, a verb it accepts" "yes" \
"$(verbs_of "$kit/$script" | grep -qx "$verb" && echo yes || echo "no: '$verb'")"
done
check "$(basename "$kit"): no \`mini\` target, which already means minimal footprint" "0" \
"$(grep -cE '^mini:' "$mk")"
done
check "there is a kit for every profile" "$(config_profiles | wc -l)" "$kits"
# An export is "take the setup I have here somewhere else", so it carries this
# machine's CHOICES — profile, ports, manifest dir — and never its credentials:
# ctrl/.env can hold registry and mirror logins next to those choices. The
# committed per-profile kits carry neither, since they must be the same on any
# machine. Proven with sentinel values in a scratch copy, because the real
# ctrl/.env may have those keys empty — and an empty value proves nothing.
SX="$TMP/export-proof"; mkdir -p "$SX"; cp -r .. "$SX/rig"
cat >> "$SX/rig/ctrl/.env" <<'EOF'
REGISTRY_USER=selftest-sentinel-user
REGISTRY_PASSWORD=selftest-sentinel-password
MANIFESTS_DIR=../selftest-sentinel-choice/overlays/dev
EOF
( cd "$SX/rig/ctrl" && bash standalone.sh export "$SX/out" >/dev/null 2>&1 )
count_in() { grep -rcF -- "$1" "$2" 2>/dev/null | awk -F: '{s+=$2} END{print s+0}'; }
check "export: carries this machine's choices" "yes" \
"$([ "$(count_in selftest-sentinel-choice "$SX/out")" -gt 0 ] && echo yes || echo no)"
check "export: carries no credential" "0" \
"$(( $(count_in selftest-sentinel-user "$SX/out") + $(count_in selftest-sentinel-password "$SX/out") ))"
check "per-profile kits: carry neither, whatever this machine has" "0" \
"$( (cd "$SX/rig/ctrl" && source ./lib/config.sh && for p in $(config_profiles); do config_snapshot "$p"; done) \
| grep -cE 'selftest-sentinel-(choice|user|password)')"
check "export: refuses to write inside the repository" "yes" \
"$( (bash standalone.sh export ../standalone/selftest-mine >/dev/null 2>&1) && echo no || echo yes)"
note "optional — needs tilt and this rig's cluster"
# Parsing the Tiltfile for real is the only way to know it still evaluates, but
# Tilt snapshots a kubectl context before parsing, so it cannot run without a
# cluster. Skipped rather than failed when there is none, the same way docgen
# skips its graphgen section.
if ! command -v tilt >/dev/null; then
printf ' skip tilt is not installed\n'
elif ! kubectl config get-contexts -o name 2>/dev/null | grep -qx "$F_CTX"; then
printf " skip no %s context — run 'make cluster up' to include this\n" "$F_CTX"
else
out="$(tilt alpha tiltfile-result --context "$F_CTX" 2>&1)"
check "Tiltfile evaluates" "yes" \
"$(printf '%s' "$out" | grep -q '"Manifests"' && echo yes || echo "no: $(printf '%s' "$out" | tail -1)")"
fi
printf '\n'
if [ "$rc" -eq 0 ]; then
printf '%d checks passed — rig still does what it says\n' "$passed"
else
printf 'FAILED — a decision above has drifted; read the comment next to it\n' >&2
fi
exit "$rc"

View File

@@ -11,13 +11,9 @@
# an unfamiliar machine, the full picture is the whole point. Failures are
# collected and reported together, and the exit code reflects the worst outcome.
#
# The same script runs inside a fresh throwaway distro (newbox), so the
# provisioning path and the everyday path cannot drift apart.
#
# Usage:
# setup.sh # host checks + the dev toolchain
# setup.sh core # kubectl and jq only — no cluster tooling
# setup.sh --share-docker # ...and offer this distro's Docker to others
# setup.sh --cluster # ...and bring the cluster up
set -euo pipefail
cd "$(dirname "$0")"
@@ -25,7 +21,6 @@ cd "$(dirname "$0")"
source ./lib/config.sh
load_config
WITH_SHARE=0
WITH_CLUSTER=0
# Cluster tooling is not wanted everywhere: a managed or corporate-issued
# machine may legitimately want kubectl and nothing that builds clusters.
@@ -33,7 +28,6 @@ TIER=dev
for a in "$@"; do
case "$a" in
core|dev) TIER="$a" ;;
--share-docker) WITH_SHARE=1 ;;
--cluster) WITH_CLUSTER=1 ;;
*) echo "unknown option: $a" >&2; exit 1 ;;
esac
@@ -73,11 +67,11 @@ record() {
step_host() {
local out
if ! out=$(bash ./wizard.sh detect 2>&1); then
if ! out=$(bash ./deps.sh detect 2>&1); then
record host fail "detection failed"
return
fi
# Anything the wizard flagged with '!' needs a human; surface the count here
# Anything flagged with '!' needs a human; surface the count here
# and the detail below rather than burying it.
local warns; warns=$(echo "$out" | grep -c '^\s*!' || true)
HOST_DETAIL="$out"
@@ -102,7 +96,7 @@ step_toolchain() {
return
fi
if bash ./wizard.sh install "$TIER" >/tmp/rig-deps.$$ 2>&1; then
if bash ./deps.sh install "$TIER" >/tmp/rig-deps.$$ 2>&1; then
local still=""
for b in $want; do
[ -x "${OUT_BIN:-$HOME/.local/bin}/$b" ] || still="$still $b"
@@ -143,33 +137,6 @@ step_docker() {
fi
}
step_share_docker() {
if [ "$WITH_SHARE" -ne 1 ]; then
record docker-share skip "not requested (--share-docker)"
return
fi
if ! grep -qi microsoft /proc/version 2>/dev/null; then
record docker-share skip "not WSL — sharing only applies between WSL distros"
return
fi
if [ -f /etc/systemd/system/docker.service.d/10-rig-shared-socket.conf ]; then
record docker-share already "this distro is offering its Docker to others"
return
fi
# Needs root, and asking mid-script is worse than telling the user the
# single command to run.
if [ "$(id -u)" -ne 0 ] && ! sudo -n true 2>/dev/null; then
record docker-share manual "run: sudo bash ctrl/dockerhost.sh share"
return
fi
if sudo bash ./dockerhost.sh share >/tmp/rig-share.$$ 2>&1; then
record docker-share done "this distro now owns the shared Docker"
rm -f "/tmp/rig-share.$$"
else
record docker-share fail "see /tmp/rig-share.$$"
fi
}
step_ports() {
local busy=""
for entry in "HTTP:$HTTP_PORT" "HTTPS:$HTTPS_PORT" "TILT:$TILT_PORT" "REGISTRY:$REGISTRY_PORT"; do
@@ -215,7 +182,6 @@ step_host
step_toolchain
step_path
step_docker
step_share_docker
step_ports
step_cluster

394
rig/ctrl/standalone.sh Normal file
View File

@@ -0,0 +1,394 @@
#!/usr/bin/env bash
# Generate the standalone kits: single-file versions of rig's own tools, one
# folder per profile, for machines the full rig is not going to.
#
# A kit is a pure function of rig as it is right now. It gains nothing rig lacks
# and loses nothing rig has — improve rig, regenerate, and every kit follows.
# Nothing in standalone/<profile>/ is ever edited by hand.
#
# What this file does NOT know, on purpose: which tools rig has, what they are
# called, how its libraries are split, where configuration lives or what it
# contains. Rig will change shape — scripts get split, renamed and grow new
# libraries — and a generator that encoded today's layout would quietly produce
# a wrong kit the first time it did. So this works from a contract a script opts
# into, and from nothing else:
#
# 1. A marker comment, alone on a line near the top, declares an entry point:
# (hash) rig:standalone <kit-name> <default-verb>
# The default verb must only REPORT: it is run as a smoke test.
# 2. Every `source` an entry point makes names a .sh file by a path that
# resolves relative to the entry point. Libraries may source further
# libraries however they like — bash follows those itself.
# 3. Configuration enters through `load_config`, and the libraries provide
# `config_profiles`, `config_freeze <profile|--current>` — which prints a
# replacement load_config with that resolution frozen in — and, for an
# export, `config_current_profile` and `config_left_out`. How config is layered,
# stored, derived or frozen is rig's business; this only asks, and embeds
# the answer without interpreting it.
#
# Bash does the resolving, not a parser here. Libraries are sourced in a clean
# shell and read back with `declare -f` and `declare -p`, so any structure bash
# can load, this can flatten.
#
# And every kit is PROVEN to stand alone before it is written: no `source` left,
# no path into rig's tree in its code, `bash -n` clean, and its default verb run
# in an empty directory with nothing from rig present. A shape this has never
# seen either passes that, or generation stops and names the kit, the file, the
# line and what is wrong. It never writes a kit that only looks finished.
#
# Usage:
# standalone.sh write generate every kit into standalone/<profile>/
# standalone.sh check generate into a scratch dir and fail if any kit differs
# standalone.sh export DIR ONE kit for the configuration this machine runs —
# its profile plus the choices in its local config,
# WITHOUT its credentials — written outside the repo.
#
# write and check are what gets committed: one kit per profile, identical on any
# machine. export is the other question — "take the setup I have here somewhere
# else" — so it reflects this machine, and for exactly that reason it never lands
# in the repository.
set -euo pipefail
cd "$(dirname "$0")"
CTRL="$PWD"
ROOT="$(cd .. && pwd)"
OUT="$ROOT/standalone"
SELF_REL="ctrl/${0##*/}"
GENERATED_TAG="GENERATED by make standalone — do not edit"
# The contract's own functions: questions rig answers FOR this generator. They
# are never carried into a kit — load_config is replaced by the frozen one, and
# the rest mean nothing without rig's tree. The only names this file knows.
CONTRACT_FUNCS="load_config config_profiles config_snapshot config_freeze config_current_profile config_left_out"
FROZEN_OPEN="# ── configuration, frozen"
FROZEN_CLOSE="# ── end of frozen configuration"
refuse() { echo >&2; echo "standalone: refusing — $*" >&2; exit 1; }
# A clean bash with nothing from the caller's shell in it. What the kit carries
# must not depend on who ran the generator or what they had exported.
clean_bash() { env -i PATH="$PATH" HOME="$HOME" CONTRACT_FUNCS="$CONTRACT_FUNCS" bash --noprofile --norc "$@"; }
# ── 1. entry points ────────────────────────────────────────────────────────
entries() {
grep -rlE --include='*.sh' '^# rig:standalone [a-z0-9-]+ [a-z0-9-]+' . 2>/dev/null \
| sed 's|^\./||' | LC_ALL=C sort
}
marker_of() { # entry -> "kit verb"
sed -nE 's/^# rig:standalone ([a-z0-9-]+) ([a-z0-9-]+).*/\1 \2/p' "$1" | head -1
}
# ── 2. the libraries an entry point sources ────────────────────────────────
# Only the entry point's own `source` lines are read as text. Everything those
# libraries pull in is resolved by bash when they are sourced in step 3.
libs_of() { # entry -> one resolved lib path per line, relative to ctrl/
local entry="$1" dir line n path
dir=$(dirname "$entry")
while IFS=: read -r n line; do
path=$(printf '%s' "$line" | sed -E 's/^[[:space:]]*(source|\.)[[:space:]]+//; s/[[:space:]]+(#.*)?$//')
path=${path#\"}; path=${path%\"}; path=${path#\'}; path=${path%\'}
case "$path" in
*'$'*) refuse "$entry:$n sources '$path' — a path with a variable in it cannot be resolved; name the file" ;;
esac
case "$path" in
*.sh) ;;
*) refuse "$entry:$n sources '$path' directly — only libraries (.sh) may be sourced; configuration has to enter through load_config" ;;
esac
path="$dir/${path#./}"; path=${path#./}
[ -f "$path" ] || refuse "$entry:$n sources '$path', which does not exist"
printf '%s\n' "$path"
done < <(grep -nE '^[[:space:]]*(source|\.)[[:space:]]+[^=]' "$entry" || true)
}
# Into the global array `libs`. Not `mapfile < <(libs_of ...)`: a refusal inside
# a process substitution only ends that subshell, so generation would carry on
# past it and fail later with a message about something else entirely.
libs_into() {
local out
out=$(libs_of "$1") || exit 1
libs=()
[ -n "$out" ] && mapfile -t libs <<< "$out"
return 0
}
# ── 3. what the libraries define, read back from bash itself ───────────────
# The frozen config replaces load_config, and the generator's own two questions
# are useless inside a kit, so none of the three is carried.
lib_defs() { # entry lib... -> declare -p globals, then declare -f functions
local entry="$1"; shift
( cd "$(dirname "$entry")" && clean_bash -c '
skip_var() { case "$1" in CONTRACT_FUNCS|BASH*|FUNCNAME|PIPESTATUS|LINENO|RANDOM|SRANDOM|SECONDS|EPOCH*|HISTCMD|COLUMNS|LINES|PWD|OLDPWD|_|SHLVL|OPTIND|OPTERR|IFS|PS4|PATH|HOME|v|f|l|before_v|before_f) return 0 ;; esac; return 1; }
before_v=" $(compgen -v | tr "\n" " ") "
before_f=" $(compgen -A function | tr "\n" " ") "
for l in "$@"; do source "$l" || { echo "__FAIL__ sourcing $l" ; exit 1; }; done
for v in $(compgen -v); do
skip_var "$v" && continue
case "$before_v" in *" $v "*) continue ;; esac
declare -p "$v"
done
for f in $(compgen -A function); do
case "$before_f" in *" $f "*) continue ;; esac
case " skip_var $CONTRACT_FUNCS " in *" $f "*) continue ;; esac
declare -f "$f"
done
' _ "$@" ) || refuse "$entry: its libraries could not be sourced cleanly"
}
# ── 4. ask rig for profiles and resolved config ────────────────────────────
ask() { # entry lib... -- function args... -> that function's stdout
local entry="$1"; shift
local libs=() a
while [ $# -gt 0 ] && [ "$1" != -- ]; do libs+=("$1"); shift; done
shift
( cd "$(dirname "$entry")" && clean_bash -c '
n=0; for a in "$@"; do n=$((n+1)); [ "$a" = -- ] && break; done
for l in "${@:1:$((n-1))}"; do source "$l"; done
shift "$n"
declare -F "$1" >/dev/null || exit 3
"$@"
' _ "${libs[@]}" -- "$@" )
}
# ── 5. assemble one kit file ───────────────────────────────────────────────
assemble() { # entry profile out-file lib...
local entry="$1" profile="$2" dest="$3"; shift 3
local libs=("$@") calls_config=no
grep -qE '(^|[^A-Za-z0-9_])load_config([^A-Za-z0-9_]|$)' "$entry" && calls_config=yes
{
echo '#!/usr/bin/env bash'
echo "# $GENERATED_TAG"
echo "#"
echo "# $(basename "$dest") for ${KIT_LABEL:-profile '$profile'}, flattened from:"
echo "# ctrl/$entry"
local l; for l in ${libs[@]+"${libs[@]}"}; do echo "# ctrl/$l"; done
echo "# Edit those and run \`make standalone\`. Changes made here are lost, and"
echo "# \`make selftest\` fails while this file differs from what rig generates."
echo
if [ ${#libs[@]} -gt 0 ]; then
echo "# ── from the libraries ──"
lib_defs "$entry" "${libs[@]}"
echo
fi
if [ "$calls_config" = yes ]; then
local frozen
frozen=$(ask "$entry" ${libs[@]+"${libs[@]}"} -- config_freeze "${FREEZE_ARG:-$profile}") \
|| refuse "ctrl/$entry calls load_config, but its libraries do not answer config_freeze ${FREEZE_ARG:-$profile}"
echo "$FROZEN_OPEN for ${KIT_LABEL:-profile '$profile'} ──"
printf '%s\n' "$frozen"
echo "$FROZEN_CLOSE ──"
echo
fi
echo "# ── ctrl/$entry ──"
# The entry point itself, minus its shebang and marker, with each source
# line it made replaced by a note — what it sourced is already above.
awk '
NR == 1 && /^#!/ { next }
/^# rig:standalone / { next }
/^[[:space:]]*(source|\.)[[:space:]]+[^=]/ { print "# (sourced library inlined above)"; next }
{ print }
' "$entry"
} > "$dest"
chmod +x "$dest"
}
# ── 6. the kit's Makefile, from the markers ────────────────────────────────
verbs_of() { # entry -> its top-level dispatch arms
awk '/^case / { inb=1; next } /^esac/ { inb=0 } inb && match($0, /^ [a-z][a-z-]*\)/) { v=substr($0, 5, RLENGTH-5); printf "%s%s", (n++ ? "|" : ""), v }' "$1"
}
makefile() { # out-dir entry...
local dir="$1"; shift
local e kit verb target verbs targets=""
for e in "$@"; do targets+=" $(basename "$e" .sh)"; done
{
echo "# $GENERATED_TAG"
echo "#"
echo "# Shorthand for the scripts beside it; they run without it. Every target"
echo "# calls a verb its script accepts — read from that script's own dispatch."
echo
echo 'HERE := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))'
echo 'ARGS := $(wordlist 2,$(words $(MAKECMDGOALS)),$(MAKECMDGOALS))'
echo 'ifneq ($(ARGS),)'
echo '$(eval $(ARGS):;@:)'
echo '.PHONY: $(ARGS)'
echo 'endif'
echo
echo '.DEFAULT_GOAL := help'
echo ".PHONY: help$targets"
echo
echo 'help: ## list targets'
printf '\t%s\n' "@grep -hE '^[a-z][a-z-]*:.*?##' \$(MAKEFILE_LIST) | sed 's/:.*##/\\t/' | expand -t16"
for e in "$@"; do
read -r kit verb <<< "$(marker_of "$e")"
target=$(basename "$e" .sh)
verbs=$(verbs_of "$e")
echo
printf '%-30s ## %s.sh [%s] (default %s)\n' "$target:" "$kit" "${verbs:-?}" "$verb"
printf '\tbash $(HERE)%s.sh $(or $(ARGS),%s)\n' "$kit" "$verb"
done
} > "$dir/Makefile"
}
# ── 7. prove a kit stands alone ────────────────────────────────────────────
verify_kit() { # dir profile entry...
local dir="$1" profile="$2"; shift 2
local e kit verb f bad smoke rc
for e in "$@"; do
read -r kit verb <<< "$(marker_of "$e")"
f="$dir/$kit.sh"
bash -n "$f" 2>/dev/null || refuse "$profile/$kit.sh does not parse: $(bash -n "$f" 2>&1 | head -1)"
# Code only: comments are free to mention anything, and the frozen block
# is data — a value that happens to hold a path is harmless unless code
# opens it, and opening it is what the smoke run below would catch.
bad=$(awk -v fz_open="$FROZEN_OPEN" -v fz_close="$FROZEN_CLOSE" '
index($0, fz_open) == 1 { fz=1; next }
index($0, fz_close) == 1 { fz=0; next }
fz || /^[[:space:]]*#/ { next }
/^[[:space:]]*(source|\.)[[:space:]]+[^=]/ { printf "%d: still sources: %s\n", NR, $0; next }
# Rig-relative only. The preceding character may not be "/", so an
# absolute system path such as /var/lib/docker is not mistaken for
# rig lib/; an explicit ./ or ../ prefix is matched on its own.
/(^|[^A-Za-z0-9_.\/])(ctrl\/|lib\/|env\.d\/)|\.\.?\/(ctrl\/|lib\/|env\.d\/)|versions\.env|(^|[^A-Za-z0-9_])\.env([^A-Za-z0-9_]|$)/ {
printf "%d: refers into rig'"'"'s tree: %s\n", NR, $0
}' "$f" | head -3)
[ -z "$bad" ] || refuse "$profile/$kit.sh does not stand alone —"$'\n'"$(printf '%s\n' "$bad" | sed 's/^/ line /')"
done
# The real test: a folder holding only this kit, and nothing else from rig.
smoke=$(mktemp -d)
cp "$dir"/* "$smoke"/
for e in "$@"; do
read -r kit verb <<< "$(marker_of "$e")"
rc=0
out=$( (cd "$smoke" && timeout 120 bash "./$kit.sh" "$verb") 2>&1 ) || rc=$?
if [ "$rc" -ne 0 ]; then
rm -rf "$smoke"
refuse "$profile/$kit.sh $verb exits $rc in an empty directory:"$'\n'"$(printf '%s\n' "$out" | tail -5 | sed 's/^/ /')"
fi
done
( cd "$smoke" && make -s help >/dev/null ) || { rm -rf "$smoke"; refuse "$profile/Makefile: make help fails"; }
rm -rf "$smoke"
}
# ── generate ───────────────────────────────────────────────────────────────
generate() { # into-dir
local into="$1" e profiles="" p kit verb libs
local -a all_entries=()
while IFS= read -r e; do all_entries+=("$e"); done < <(entries)
[ ${#all_entries[@]} -gt 0 ] || refuse "no script under ctrl/ carries a '# rig:standalone <kit> <verb>' marker"
# Profiles come from whichever entry point's libraries can answer for them.
for e in "${all_entries[@]}"; do
libs_into "$e"
profiles=$(ask "$e" ${libs[@]+"${libs[@]}"} -- config_profiles 2>/dev/null) && [ -n "$profiles" ] && break
profiles=""
done
[ -n "$profiles" ] || refuse "no entry point's libraries answer config_profiles, so there is nothing to generate a kit per"
for p in $profiles; do
mkdir -p "$into/$p"
for e in "${all_entries[@]}"; do
read -r kit verb <<< "$(marker_of "$e")"
libs_into "$e"
assemble "$e" "$p" "$into/$p/$kit.sh" ${libs[@]+"${libs[@]}"}
done
makefile "$into/$p" "${all_entries[@]}"
verify_kit "$into/$p" "$p" "${all_entries[@]}"
echo " $p: $(cd "$into/$p" && ls | tr '\n' ' ')"
done
}
# A kit folder is ours if its Makefile says so. Anything else under standalone/
# is left alone, so a hand-written file there is never swept away.
is_generated_dir() { grep -qF "$GENERATED_TAG" "$1/Makefile" 2>/dev/null; }
cmd="${1:-write}"
[ $# -gt 0 ] && shift
case "$cmd" in
write)
tmp=$(mktemp -d); trap 'rm -rf "$tmp"' EXIT
echo "generating kits from rig's current tree"
generate "$tmp"
mkdir -p "$OUT"
for d in "$OUT"/*/; do
d=${d%/}; [ -d "$d" ] || continue
if is_generated_dir "$d" && [ ! -d "$tmp/${d##*/}" ]; then
echo " removed $(basename "$d") — no such profile any more"
rm -rf "$d"
fi
done
for d in "$tmp"/*/; do
d=${d%/}
rm -rf "$OUT/${d##*/}"
cp -r "$d" "$OUT/${d##*/}"
done
echo "wrote standalone/<profile>/ — every kit verified to stand alone"
;;
check)
tmp=$(mktemp -d); trap 'rm -rf "$tmp"' EXIT
generate "$tmp" >/dev/null
stale=0
for d in "$tmp"/*/; do
d=${d%/}; p=${d##*/}
if ! diff -rq "$d" "$OUT/$p" >/dev/null 2>&1; then
echo "stale: standalone/$p$(diff -rq "$d" "$OUT/$p" 2>&1 | head -1)"
stale=1
fi
done
for d in "$OUT"/*/; do
d=${d%/}; [ -d "$d" ] || continue
if is_generated_dir "$d" && [ ! -d "$tmp/${d##*/}" ]; then
echo "stale: standalone/${d##*/} — no such profile any more"; stale=1
fi
done
[ "$stale" -eq 0 ] || { echo "run: make standalone"; exit 1; }
echo "every kit is current"
;;
export)
dest="${1:-}"
[ -n "$dest" ] || refuse "export needs a directory, outside the repo: make standalone export ~/rig-kit"
dest=$(realpath -m "$dest")
top=$(git -C "$ROOT" rev-parse --show-toplevel 2>/dev/null || echo "$ROOT")
case "$dest/" in
"$top"/*) refuse "an export reflects this machine, so it does not go inside the repository — $dest is under $top. The committed per-profile kits are what standalone/ is for." ;;
esac
if [ -d "$dest" ] && [ -n "$(ls -A "$dest" 2>/dev/null)" ] && ! is_generated_dir "$dest"; then
refuse "$dest already holds something that is not a previous export — pick an empty directory"
fi
mapfile -t all_entries < <(entries)
[ ${#all_entries[@]} -gt 0 ] || refuse "no script under ctrl/ carries a '# rig:standalone <kit> <verb>' marker"
libs_into "${all_entries[0]}"
profile=$(ask "${all_entries[0]}" ${libs[@]+"${libs[@]}"} -- config_current_profile) \
|| refuse "this machine's configuration does not resolve — run make check"
left=$(ask "${all_entries[0]}" ${libs[@]+"${libs[@]}"} -- config_left_out | tr '\n' ' ')
tmp=$(mktemp -d); trap 'rm -rf "$tmp"' EXIT
echo "exporting the configuration this machine runs (profile '$profile')"
FREEZE_ARG=--current
KIT_LABEL="the configuration exported from $(hostname -s 2>/dev/null || echo this machine) (profile '$profile', local choices included, credentials not)"
for e in "${all_entries[@]}"; do
read -r kit verb <<< "$(marker_of "$e")"
libs_into "$e"
assemble "$e" "$profile" "$tmp/$kit.sh" ${libs[@]+"${libs[@]}"}
done
makefile "$tmp" "${all_entries[@]}"
verify_kit "$tmp" "export" "${all_entries[@]}"
rm -rf "$dest"; mkdir -p "$(dirname "$dest")"; cp -r "$tmp" "$dest"
echo " wrote $dest: $(cd "$dest" && ls | tr '\n' ' ')— verified to stand alone"
if [ -n "${left// /}" ]; then
echo
echo " NOT carried — this machine's own, set them on the target if it needs them:"
for k in $left; do echo " $k"; done
fi
;;
*) echo "usage: $SELF_REL [write|check|export DIR]" >&2; exit 1 ;;
esac

View File

@@ -1,106 +0,0 @@
#!/usr/bin/env bash
# Station check: is this workstation ready to run rig?
#
# Reports and instructs; never silently fixes anything. Everything it finds is
# either already fine, or something a human has to decide on.
#
# Runs the wizard's host detection in a container when Docker is the only thing
# installed, or directly when the toolchain is already present. Then adds the
# checks that need this repo's config: profile sanity, CA trust, port clashes.
set -euo pipefail
cd "$(dirname "$0")"
WIZARD_IMAGE="${WIZARD_IMAGE:-$(basename "$(cd .. && pwd)")-wizard}"
# Host detection. Prefer running it bare — it needs no dependencies beyond
# coreutils — and fall back to the container only if this shell can't.
bash ./wizard.sh detect
# ── repo-level checks ──────────────────────────────────────────────────────
source ./lib/config.sh
load_config
echo
echo "config"
echo " profile ${PROFILE_NAME} (nodes=${NODES} audit=${AUDIT})"
echo " cluster ${CLUSTER} (context ${KUBECONTEXT})"
echo " registry ${REGISTRY_MODE}"
echo " ingress ${INGRESS_MODE}"
if [ ! -f ./.env ]; then
echo " ! ctrl/.env missing — copy it: cp ctrl/.env.example ctrl/.env"
fi
# A 3-node profile on a box that's already full is the most common first
# failure, and it presents as pods stuck Pending rather than anything obvious.
avail=$(awk '/^MemAvailable:/{printf "%d", $2/1024/1024}' /proc/meminfo)
need=$((NODES * 2))
if [ "$avail" -lt "$need" ]; then
echo " ! profile '${PROFILE_NAME}' wants ~${need} GB, ${avail} GB available"
echo " 'make cluster list' shows what else is running; 'make cluster free' stops it"
fi
# The CA reaches three places and only one of them is ours. Report the other two.
if [ -n "${REGISTRY_CA_FILE:-}" ]; then
echo
echo "registry CA"
if [ ! -r "$REGISTRY_CA_FILE" ]; then
echo " ! REGISTRY_CA_FILE not readable: $REGISTRY_CA_FILE"
else
echo " file $REGISTRY_CA_FILE"
host="${REGISTRY_REMOTE_URL#*://}"; host="${host%%/*}"
if [ -n "$host" ] && [ ! -f "/etc/docker/certs.d/${host}/ca.crt" ]; then
echo " ! the HOST docker daemon does not trust it yet:"
echo " sudo mkdir -p /etc/docker/certs.d/${host}"
echo " sudo cp ${REGISTRY_CA_FILE} /etc/docker/certs.d/${host}/ca.crt"
echo " (kind nodes are handled by registry.sh; in-cluster clients are the workload's job)"
fi
fi
fi
# Host ports this environment will try to bind. Checked before cluster creation
# because docker reports a clash halfway through, as an opaque
# "failed to bind host port ...: address already in use".
echo
echo "ports (block derived from the directory name — see 'make ports')"
port_busy() {
if command -v ss >/dev/null 2>&1; then
ss -ltn "sport = :$1" 2>/dev/null | grep -q LISTEN && return 0 || return 1
fi
# iproute2 is absent from a minimal Debian, so fall back to procfs rather
# than silently reporting everything as free.
local hex; hex=$(printf ':%04X' "$1")
grep -qi "^ *[0-9]*: [0-9A-F]*$hex " /proc/net/tcp /proc/net/tcp6 2>/dev/null
}
# A port held by THIS environment's own cluster is not a clash — it is the thing
# working. Reporting it as a problem every time the cluster is up would train
# people to ignore this section, which is the opposite of the point.
# Extract with a second grep rather than `tr -d ':->'`: in tr, ':->' is the
# character RANGE ':' to '>', which does not contain '-', so the trailing dash
# survives and nothing ever matches.
ours=$(docker ps --filter "label=io.x-k8s.kind.cluster=${CLUSTER}" \
--format '{{.Ports}}' 2>/dev/null | tr ',' '\n' \
| grep -oE ':[0-9]+->' | grep -oE '[0-9]+' || true)
clash=0
for entry in "HTTP:${HTTP_PORT}" "HTTPS:${HTTPS_PORT}" \
"TILT:${TILT_PORT}" "REGISTRY:${REGISTRY_PORT}"; do
name="${entry%%:*}"; p="${entry#*:}"
[ -n "$p" ] || continue
if ! port_busy "$p"; then
printf " %-9s %-6s free\n" "$name" "$p"
elif echo "$ours" | grep -qx "$p"; then
printf " %-9s %-6s in use by this environment's cluster\n" "$name" "$p"
else
printf " ! %-9s %-6s IN USE by something else\n" "$name" "$p"
clash=1
fi
done
if [ "$clash" -eq 1 ]; then
echo " override the clashing one in ctrl/.env, e.g. HTTP_PORT=21080"
echo " (or rename this directory — the whole block follows the name)"
fi

View File

@@ -1,4 +1,4 @@
# Pinned toolchain — the single manifest the wizard installs from.
# Pinned toolchain — the single manifest ctrl/deps.sh installs from.
# Every entry is a single binary; none of them needs an apt repo.
# kubectl fully static
# kind libc only
@@ -6,8 +6,19 @@
# jq upstream static build (Debian's is linked against libjq/libonig)
#
# Checksums are the upstream-published SHA256 of the linux/amd64 artifact.
# To bump: change the version, then re-run `bash ctrl/versions-refresh.sh` and
# commit the result — never hand-edit a checksum.
#
# To bump: change the version, then take the checksum from the release's own
# published list — never hand-edit or hand-copy one from a download you did.
# For anything hosted on GitHub releases that is:
#
# curl -sSL https://github.com/<org>/<repo>/releases/download/<tag>/checksums.txt \
# | grep linux.x86_64
#
# (kubectl publishes its own instead: <KUBECTL_URL>.sha256.)
#
# There was a `ctrl/versions-refresh.sh` named here that has never existed. If
# bumping stops being rare enough to do by hand, write it — but a comment
# pointing at a missing script is worse than no comment.
KIND_VERSION=v0.32.0
KIND_SHA256=50030de23cf40a18505f20426f6a8506bedf13c6e509244bd1fa9463721b0f54
@@ -21,10 +32,27 @@ TILT_VERSION=0.37.6
TILT_SHA256=e9672b8a18d43501f35dcfe98465969a7db0e436b36cf0c50c7e6f8d40de5fe6
TILT_URL=https://github.com/tilt-dev/tilt/releases/download/v${TILT_VERSION}/tilt.${TILT_VERSION}.linux.x86_64.tar.gz
# ctlptl — creates a kind cluster WITH a local registry wired in, which is what
# keeps images off docker.io (an unqualified name means docker.io/library/<name>).
# Same publisher and same archive shape as tilt: binary at the archive root, so
# fetch_tgz handles it with strip=0 and no special case.
CTLPTL_VERSION=0.9.4
CTLPTL_SHA256=c63a1ec28e60bc3faf6becb76f53355c5cf5e0143dafdd27ad85db5584fa6b1e
CTLPTL_URL=https://github.com/tilt-dev/ctlptl/releases/download/v${CTLPTL_VERSION}/ctlptl.${CTLPTL_VERSION}.linux.x86_64.tar.gz
JQ_VERSION=1.8.2
JQ_SHA256=b1c22172dd303f3be49e935aa56aa48a8b7a46e0bc838b4997d3bb451495870f
JQ_URL=https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/jq-linux-amd64
# docker compose — the distro docker packages ship the daemon and the CLI but
# frequently not this, so `docker compose up` fails with "unknown command" on an
# otherwise working Docker. It is a CLI plugin, found by NAME in a plugin
# directory, so a copy in the bin dir alone only gives you the retired
# `docker-compose` v1 spelling; deps.sh links it into ~/.docker/cli-plugins.
COMPOSE_VERSION=5.5.1
COMPOSE_SHA256=db1889184726840f75c4f9c001048430d4f25b3be3cb084d3ddd762bc0aed576
COMPOSE_URL=https://github.com/docker/compose/releases/download/v${COMPOSE_VERSION}/docker-compose-linux-x86_64
# Node images shipped with KIND_VERSION above, pinned by digest so a kind upgrade
# can never silently move the k8s version. Profiles select one via K8S_VERSION.
# Older entries are kept deliberately: running a trailing-edge control plane is
@@ -44,9 +72,9 @@ CERT_MANAGER_VERSION=v1.21.1
METRICS_SERVER_VERSION=v0.9.0
METALLB_VERSION=v0.16.0
# Dependency containers. These mirror soleprint's cabinets
# (soleprint/station/cabinets/), so a room that declares postgres gets the same
# thing whether it runs on compose or in the cluster. Pinned by tag rather than
# Cabinets — public services dropped in as-is, the upstream image unmodified.
# The same declaration installs on compose or in the cluster, so a dependency is
# named once and works either way. Pinned by tag rather than
# digest because they are ordinary upstream images with no supply chain claim
# attached — bump freely, and preload them for the offline profile.
POSTGRES_IMAGE=postgres:16-alpine

View File

@@ -1,381 +0,0 @@
#!/usr/bin/env bash
# The installation wizard: detect the host, install a pinned toolchain onto it,
# then report what it could not do. It never runs the cluster and never mutates
# the host outside the directories mounted into it.
#
# Usage (normally via `make station` / `make deps`, or directly):
# wizard.sh detect # report host facts only, change nothing
# wizard.sh fetch [core|dev] [--to DIR] # download + verify into DIR
# wizard.sh install [core|dev] # detect, fetch, install, report
#
# Tiers: 'core' is kubectl + jq (talk to a cluster); 'dev' adds kind and tilt
# Default is dev.
#
# Runs both inside the wizard container and bare on a host. Inside the
# container, host files are read through $HOST_ROOT (mount / as :ro); bare, it
# falls back to /.
set -euo pipefail
# Keep the caller's cwd so a relative --to resolves where the user expects,
# not against ctrl/ once we've moved.
INVOKED_FROM="$PWD"
cd "$(dirname "$0")"
source ./versions.env
# Resolve a possibly-relative path against the caller's original directory.
abspath() {
case "$1" in
/*) echo "$1" ;;
*) echo "$INVOKED_FROM/$1" ;;
esac
}
OUT_BIN="${OUT_BIN:-$HOME/.local/bin}"
HOST_ROOT="${HOST_ROOT:-/}"
DEPS_SOURCE="${DEPS_SOURCE:-upstream}"
DEPS_ARTIFACTORY_URL="${DEPS_ARTIFACTORY_URL:-}"
BAKED_BIN="${BAKED_BIN:-/opt/rig/bin}"
# Collected by detect(), printed by report_manual() at the very end.
MANUAL=()
# Host FILES (/etc/..., /mnt/c/...) must be read through the mount. Kernel-level
# facts (kernel version, meminfo, inotify) are shared with the container, so the
# container's own view is already the host's.
host_file() {
local p="${1#/}"
if [ "$HOST_ROOT" != "/" ] && [ -e "$HOST_ROOT/$p" ]; then
echo "$HOST_ROOT/$p"
else
echo "/$p"
fi
}
# ── detect ─────────────────────────────────────────────────────────────────
is_wsl() { grep -qi microsoft /proc/version 2>/dev/null; }
detect() {
echo "host"
echo " kernel $(uname -r)"
local osr; osr=$(host_file /etc/os-release)
[ -r "$osr" ] && echo " distro $(sed -n 's/^PRETTY_NAME="\(.*\)"/\1/p' "$osr")"
local total_kb avail_kb
total_kb=$(awk '/^MemTotal:/{print $2}' /proc/meminfo)
avail_kb=$(awk '/^MemAvailable:/{print $2}' /proc/meminfo)
printf " memory %d GB total, %d GB available\n" \
$((total_kb / 1024 / 1024)) $((avail_kb / 1024 / 1024))
if [ $((avail_kb / 1024 / 1024)) -lt 4 ]; then
echo " ! under 4 GB available — a multi-node profile will struggle."
echo " 'make cluster list' shows the others; 'make cluster free' stops them."
fi
detect_wsl
detect_docker
detect_inotify
}
detect_wsl() {
if ! is_wsl; then
echo " platform native linux"
return
fi
echo " platform WSL"
# systemd is off by default in WSL, and the ingress/DNS paths that use a
# host service need it. Enabling it requires a Windows-side restart, which
# cannot be issued from inside the distro.
local wc; wc=$(host_file /etc/wsl.conf)
if [ -r "$wc" ] && grep -qE '^\s*systemd\s*=\s*true' "$wc"; then
echo " systemd enabled in wsl.conf"
else
echo " ! systemd not enabled in /etc/wsl.conf"
MANUAL+=("Enable systemd — add to /etc/wsl.conf:
[boot]
systemd=true
then from a WINDOWS terminal (not this shell): wsl --shutdown")
fi
# WSL regenerates /etc/resolv.conf on every boot, which silently reverts any
# local DNS setup.
if [ -r "$wc" ] && grep -qE '^\s*generateResolvConf\s*=\s*false' "$wc"; then
echo " resolv.conf pinned (generateResolvConf=false)"
else
echo " - resolv.conf is WSL-generated; DNS_MODE=dnsmasq would be reverted on reboot"
fi
local wcfg
wcfg=$(ls "$HOST_ROOT"/mnt/c/Users/*/.wslconfig 2>/dev/null | head -1 || true)
if [ -n "$wcfg" ] && grep -qE '^\s*memory\s*=' "$wcfg"; then
echo " wslconfig memory set: $(grep -E '^\s*memory\s*=' "$wcfg" | tr -d ' ')"
else
MANUAL+=("Cap/raise the WSL VM memory — in %USERPROFILE%\\.wslconfig on Windows:
[wsl2]
memory=8GB
then from a WINDOWS terminal: wsl --shutdown")
fi
}
detect_docker() {
# Reachability of the daemon is the real question, and the CLI is only how
# we ask it. Note that when this runs inside the wizard container, Docker
# necessarily exists on the host — otherwise nothing would be executing —
# so a missing CLI in here is a wizard packaging bug, not a host problem.
if ! command -v docker >/dev/null 2>&1; then
if [ -S /var/run/docker.sock ]; then
echo " docker socket present (no cli in this context)"
else
echo " ! docker not found and no socket at /var/run/docker.sock"
MANUAL+=("Install Docker — the one true prerequisite:
sudo apt-get install -y docker.io && sudo usermod -aG docker \"\$USER\"
then log out and back in.")
fi
return
fi
if docker info >/dev/null 2>&1; then
echo " docker $(docker version --format '{{.Server.Version}}' 2>/dev/null)"
local n
n=$(docker ps --filter "label=io.x-k8s.kind.cluster" --format '{{.Names}}' 2>/dev/null | wc -l)
# Must be an `if`, not `[ ] && echo`: as the last statement in this
# function the latter returns 1 when the count is zero, and `set -e`
# then kills the caller. That is the fresh-machine case — no clusters
# yet — so the bug only ever shows up where it does most harm.
if [ "$n" -gt 0 ]; then
echo " - $n kind node container(s) already running; see 'make cluster list'"
fi
else
echo " ! docker cli present but the daemon is unreachable"
MANUAL+=("Start Docker, or add yourself to the docker group:
sudo usermod -aG docker \"\$USER\" # then log out and back in")
fi
}
# kind and Tilt both watch large trees. WSL ships defaults (8192/128) far too low,
# and the failure mode is silent: Tilt simply stops noticing file changes.
detect_inotify() {
local w i
w=$(cat /proc/sys/fs/inotify/max_user_watches 2>/dev/null || echo 0)
i=$(cat /proc/sys/fs/inotify/max_user_instances 2>/dev/null || echo 0)
echo " inotify watches=$w instances=$i"
if [ "$w" -lt 524288 ] || [ "$i" -lt 512 ]; then
echo " ! inotify limits are low — Tilt will silently stop noticing file changes"
MANUAL+=("Raise inotify limits (needs root on the host):
echo -e 'fs.inotify.max_user_watches=524288\\nfs.inotify.max_user_instances=512' \\
| sudo tee /etc/sysctl.d/99-rig.conf
sudo sysctl --system")
fi
}
# ── fetch ──────────────────────────────────────────────────────────────────
# Resolve where a given artifact comes from, honouring DEPS_SOURCE.
resolve_url() {
local upstream="$1"
case "$DEPS_SOURCE" in
upstream) echo "$upstream" ;;
artifactory)
if [ -z "$DEPS_ARTIFACTORY_URL" ]; then
echo "DEPS_SOURCE=artifactory but DEPS_ARTIFACTORY_URL is empty" >&2
exit 1
fi
echo "${DEPS_ARTIFACTORY_URL%/}/$(basename "$upstream")"
;;
*) echo "unsupported DEPS_SOURCE '$DEPS_SOURCE' for a download" >&2; exit 1 ;;
esac
}
verify() {
local file="$1" want="$2" name="$3" got
got=$(sha256sum "$file" | awk '{print $1}')
if [ "$got" != "$want" ]; then
echo "checksum mismatch for $name" >&2
echo " expected $want" >&2
echo " got $got" >&2
exit 1
fi
}
# fetch_bin <name> <url> <sha256> <dest-dir> — a bare binary
fetch_bin() {
local name="$1" url="$2" sha="$3" dest="$4"
local tmp="$dest/.$name.tmp"
echo " fetching $name"
curl -fsSL --retry 3 -o "$tmp" "$(resolve_url "$url")"
verify "$tmp" "$sha" "$name"
mv "$tmp" "$dest/$name"
chmod +x "$dest/$name"
}
# fetch_tgz <name> <url> <sha256> <dest-dir> <path-inside-archive> <strip>
# Archive layouts differ — tilt's is flat (the binary at the root, strip=0),
# others nest it a directory down — so the caller says which.
fetch_tgz() {
local name="$1" url="$2" sha="$3" dest="$4" inner="$5" strip="$6"
local tmp="$dest/.$name.tgz"
echo " fetching $name"
curl -fsSL --retry 3 -o "$tmp" "$(resolve_url "$url")"
verify "$tmp" "$sha" "$name"
# --no-same-owner: extracting as root would otherwise restore the uid/gid
# baked into the archive (some ship as uid 1001), leaving a binary the host
# user does not own.
tar -xzf "$tmp" -C "$dest" --strip-components="$strip" --no-same-owner "$inner"
rm -f "$tmp"
chmod +x "$dest/$name"
}
# The wizard runs as root so it can reach the docker socket, which means
# everything it writes into a mounted volume lands root-owned and unusable from
# the host. Hand it back to whoever owns the mount point (the host user created
# that directory before mounting it).
fix_ownership() {
local dir="$1"
[ -d "$dir" ] || return 0
local owner="${HOST_UID:-}:${HOST_GID:-}"
if [ "$owner" = ":" ]; then
owner=$(stat -c '%u:%g' "$dir")
fi
[ "$owner" = "0:0" ] && return 0
chown -R "$owner" "$dir" 2>/dev/null || true
}
# Two tiers, because not every machine should get cluster tooling.
#
# core kubectl, jq — talk to a cluster someone else runs. Nothing that
# creates one. Appropriate on a managed or corporate-issued machine
# where development tools are not wanted by default.
# dev core plus kind and tilt — build clusters and hot-reload into them.
#
# The split exists because "install the toolchain" is not one decision: on a
# managed workspace the right answer is kubectl and nothing else.
CORE_TOOLS="kubectl jq"
# No helm: every addon installs with `kubectl apply -f <url>`, so nothing here
# has ever invoked it. Add it back the day something actually needs a chart.
DEV_TOOLS="kind tilt"
fetch() {
local dest="$OUT_BIN" tier="${TIER:-dev}"
while [ $# -gt 0 ]; do
case "$1" in
--to) dest="$2"; shift 2 ;;
core|dev) tier="$1"; shift ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
dest="$(abspath "$dest")"
mkdir -p "$dest"
TIER="$tier"
if [ "$DEPS_SOURCE" = "baked" ]; then
echo "installing baked binaries from $BAKED_BIN"
cp -a "$BAKED_BIN"/. "$dest"/
fix_ownership "$dest"
return
fi
echo "fetching '$tier' toolchain (source: $DEPS_SOURCE)"
fetch_bin kubectl "$KUBECTL_URL" "$KUBECTL_SHA256" "$dest"
fetch_bin jq "$JQ_URL" "$JQ_SHA256" "$dest"
if [ "$tier" = "dev" ]; then
fetch_bin kind "$KIND_URL" "$KIND_SHA256" "$dest"
fetch_tgz tilt "$TILT_URL" "$TILT_SHA256" "$dest" tilt 0
fi
fix_ownership "$dest"
# kind writes the kubeconfig as root too; hand that back as well when it's
# a mounted host directory rather than container-local state.
fix_ownership "${KUBE_DIR:-/out/kube}"
}
# ── install ────────────────────────────────────────────────────────────────
report_manual() {
echo
if [ ${#MANUAL[@]} -eq 0 ]; then
echo "nothing left to do by hand."
return
fi
echo "host actions the wizard cannot perform (${#MANUAL[@]}):"
echo
local n=1
for m in "${MANUAL[@]}"; do
echo " $n. $m"
echo
n=$((n + 1))
done
}
# Installing into a directory that sits early in PATH silently replaces whatever
# the machine was already using — which on a shared or client machine can break
# unrelated work (kubectl more than one minor away from a cluster is the common
# one). Say so; never decide it for them.
tier_tools() { [ "$1" = "core" ] && echo "$CORE_TOOLS" || echo "$CORE_TOOLS $DEV_TOOLS"; }
warn_shadowing() {
local b existing shadowed="" tier="${1:-dev}"
for b in $(tier_tools "$tier"); do
[ -x "$OUT_BIN/$b" ] || continue
# Where would this resolve if OUT_BIN weren't in the way?
existing=$(PATH=$(echo "$PATH" | tr ':' '\n' | grep -vx "$OUT_BIN" | paste -sd:) \
command -v "$b" 2>/dev/null || true)
[ -n "$existing" ] || continue
[ "$existing" = "$OUT_BIN/$b" ] && continue
shadowed+=" $b $existing"$'\n'
done
[ -n "$shadowed" ] || return 0
case ":${PATH}:" in
*":$OUT_BIN:"*) ;;
*) return 0 ;; # not on PATH yet, so nothing is being shadowed
esac
echo
echo " ! these were already installed elsewhere and are now shadowed by $OUT_BIN:"
printf '%s' "$shadowed"
echo " Other projects on this machine will pick up the new versions."
MANUAL+=("Decide which toolchain wins. To keep the previous one, remove what
was just installed:
rm -f $(for b in $(tier_tools "$tier"); do printf '%s ' "$OUT_BIN/$b"; done)
Or install somewhere private instead:
OUT_BIN=\$PWD/def/bin make deps # then put that dir first in PATH")
}
install() {
local tier="${1:-dev}"
detect
echo
fetch "$tier"
echo
echo "installed to $OUT_BIN ($tier):"
for b in $(tier_tools "$tier"); do
[ -x "$OUT_BIN/$b" ] && echo " $b"
done
if [ "$tier" = "core" ]; then
echo " (no kind/tilt — 'make deps dev' adds them)"
fi
warn_shadowing "$tier"
case ":${PATH}:" in
*":$OUT_BIN:"*) ;;
*) MANUAL+=("Put the toolchain on your PATH — add to ~/.bashrc:
export PATH=\"${OUT_BIN}:\$PATH\"") ;;
esac
report_manual
}
# ── main ───────────────────────────────────────────────────────────────────
case "${1:-install}" in
detect) detect; report_manual ;;
fetch) shift; fetch "$@" ;;
install) shift; install "${1:-dev}" ;;
*) echo "usage: $0 [detect|fetch|install]" >&2; exit 1 ;;
esac

View File

@@ -20,13 +20,13 @@ digraph rig_install {
bin [label="~/.local/bin\nkind · kubectl · tilt\njq" fillcolor="#121829" shape=cylinder]
}
subgraph cluster_wizard {
subgraph cluster_installer {
label="Installer container (transient)"
style=dashed
color="#1e2a4a"
fontcolor="#8892a8"
wizard [label="wizard\ncurl · jq · python · graphviz" fillcolor="#121829"]
installer [label="deps installer\ncurl · jq · python · graphviz" fillcolor="#121829"]
detect [label="detect host\nWSL · memory · inotify · docker" fillcolor="#121829"]
fetch [label="fetch + verify\nSHA256, pinned versions" fillcolor="#121829"]
}
@@ -34,14 +34,14 @@ digraph rig_install {
upstream [label="upstream\nreleases / corporate mirror" fillcolor="#1a3a1a" fontcolor="#00c853" shape=octagon]
report [label="report what it\nCANNOT do" fillcolor="#3a1a1a" fontcolor="#ffc107"]
docker -> wizard [label="docker run"]
wizard -> detect
docker -> installer [label="docker run"]
installer -> detect
detect -> fetch
fetch -> upstream [label="pinned + checksummed" color="#00c853"]
fetch -> bin [label="install"]
detect -> report [style=dashed label="sudo / Windows-side steps" color="#ffc107"]
// The container is gone after this; nothing depends on it at run time.
wizard -> gone [style=dotted label="exits"]
installer -> gone [style=dotted label="exits"]
gone [label="(container discarded)" fillcolor="#0a0e17" fontcolor="#4a5568" color="#1e2a4a" style="filled,dashed"]
}

View File

@@ -16,7 +16,7 @@
<text xml:space="preserve" text-anchor="middle" x="1011.15" y="-170.8" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#8892a8">Your machine</text>
</g>
<g id="clust2" class="cluster">
<title>cluster_wizard</title>
<title>cluster_installer</title>
<polygon fill="#0a0e17" stroke="#1e2a4a" stroke-dasharray="5,2" points="8,-95 8,-175 745.5,-175 745.5,-95 8,-95"/>
<text xml:space="preserve" text-anchor="middle" x="376.75" y="-155.8" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#8892a8">Installer container (transient)</text>
</g>
@@ -27,16 +27,16 @@
<text xml:space="preserve" text-anchor="middle" x="1011.15" y="-46.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#0066ff">Docker</text>
<text xml:space="preserve" text-anchor="middle" x="1011.15" y="-32.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#0066ff">(the one prerequisite)</text>
</g>
<!-- wizard -->
<!-- installer -->
<g id="node3" class="node">
<title>wizard</title>
<title>installer</title>
<polygon fill="#121829" stroke="#1e2a4a" points="181.25,-139 16,-139 16,-103 181.25,-103 181.25,-139"/>
<text xml:space="preserve" text-anchor="middle" x="98.62" y="-124.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">wizard</text>
<text xml:space="preserve" text-anchor="middle" x="98.62" y="-124.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">deps installer</text>
<text xml:space="preserve" text-anchor="middle" x="98.62" y="-110.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">curl · jq · python · graphviz</text>
</g>
<!-- docker&#45;&gt;wizard -->
<!-- docker&#45;&gt;installer -->
<g id="edge1" class="edge">
<title>docker&#45;&gt;wizard</title>
<title>docker&#45;&gt;installer</title>
<path fill="none" stroke="#4a5568" d="M906.12,-45.7C757.47,-50.51 475.98,-63.12 238.25,-94 223.38,-95.93 207.7,-98.48 192.45,-101.24"/>
<polygon fill="#4a5568" stroke="#4a5568" points="192.13,-97.74 182.93,-103.01 193.4,-104.63 192.13,-97.74"/>
<text xml:space="preserve" text-anchor="middle" x="506.62" y="-74.39" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">docker run</text>
@@ -57,9 +57,9 @@
<text xml:space="preserve" text-anchor="middle" x="333.62" y="-124.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">detect host</text>
<text xml:space="preserve" text-anchor="middle" x="333.62" y="-110.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">WSL · memory · inotify · docker</text>
</g>
<!-- wizard&#45;&gt;detect -->
<!-- installer&#45;&gt;detect -->
<g id="edge2" class="edge">
<title>wizard&#45;&gt;detect</title>
<title>installer&#45;&gt;detect</title>
<path fill="none" stroke="#4a5568" d="M181.44,-121C195.97,-121 211.28,-121 226.37,-121"/>
<polygon fill="#4a5568" stroke="#4a5568" points="226.33,-124.5 236.33,-121 226.33,-117.5 226.33,-124.5"/>
</g>
@@ -69,9 +69,9 @@
<polygon fill="#0a0e17" stroke="#1e2a4a" stroke-dasharray="5,2" points="400.5,-47 266.75,-47 266.75,-11 400.5,-11 400.5,-47"/>
<text xml:space="preserve" text-anchor="middle" x="333.62" y="-25.3" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#4a5568">(container discarded)</text>
</g>
<!-- wizard&#45;&gt;gone -->
<!-- installer&#45;&gt;gone -->
<g id="edge7" class="edge">
<title>wizard&#45;&gt;gone</title>
<title>installer&#45;&gt;gone</title>
<path fill="none" stroke="#4a5568" stroke-dasharray="1,5" d="M119.29,-102.62C138.26,-86 168.54,-62.29 199.25,-49.75 216.76,-42.6 236.49,-37.9 255.29,-34.81"/>
<polygon fill="#4a5568" stroke="#4a5568" points="255.67,-38.29 265.04,-33.35 254.64,-31.37 255.67,-38.29"/>
<text xml:space="preserve" text-anchor="middle" x="209.75" y="-52.45" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">exits</text>

Before

Width:  |  Height:  |  Size: 9.0 KiB

After

Width:  |  Height:  |  Size: 9.0 KiB

View File

@@ -20,7 +20,7 @@ digraph rig_environment {
cname [label="cluster name\nacmebank" fillcolor="#121829"]
ctx [label="kubectl context\nkind-acmebank" fillcolor="#121829"]
img [label="image tag\nacmebank-wizard" fillcolor="#121829"]
img [label="image tag\nacmebank-deps" fillcolor="#121829"]
ports [label="port block\n2130021309" fillcolor="#121829"]
reg [label="registry container\nacmebank-registry" fillcolor="#121829"]
}

View File

@@ -4,28 +4,28 @@
<!-- Generated by graphviz version 14.1.2 (0)
-->
<!-- Title: rig_environment Pages: 1 -->
<svg width="971pt" height="481pt"
viewBox="0.00 0.00 971.00 481.00" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink">
<svg width="962pt" height="481pt"
viewBox="0.00 0.00 962.00 481.00" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink">
<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 476.83)">
<title>rig_environment</title>
<polygon fill="#0a0e17" stroke="none" points="-4,4 -4,-476.83 967,-476.83 967,4 -4,4"/>
<text xml:space="preserve" text-anchor="middle" x="481.5" y="-453.63" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#0066ff">One environment per directory — copies never collide</text>
<polygon fill="#0a0e17" stroke="none" points="-4,4 -4,-476.83 958,-476.83 958,4 -4,4"/>
<text xml:space="preserve" text-anchor="middle" x="477" y="-453.63" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#0066ff">One environment per directory — copies never collide</text>
<g id="clust1" class="cluster">
<title>cluster_derived</title>
<polygon fill="#0a0e17" stroke="#1e2a4a" stroke-dasharray="5,2" points="8,-65 8,-144.5 605,-144.5 605,-65 8,-65"/>
<text xml:space="preserve" text-anchor="middle" x="306.5" y="-125.3" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#8892a8">Everything below is derived from it</text>
<polygon fill="#0a0e17" stroke="#1e2a4a" stroke-dasharray="5,2" points="8,-65 8,-144.5 596,-144.5 596,-65 8,-65"/>
<text xml:space="preserve" text-anchor="middle" x="302" y="-125.3" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#8892a8">Everything below is derived from it</text>
</g>
<g id="clust2" class="cluster">
<title>cluster_config</title>
<polygon fill="#0a0e17" stroke="#1e2a4a" stroke-dasharray="5,2" points="613,-65 613,-437.33 955,-437.33 955,-65 613,-65"/>
<text xml:space="preserve" text-anchor="middle" x="784" y="-418.13" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#8892a8">Configuration — weakest first, later wins</text>
<polygon fill="#0a0e17" stroke="#1e2a4a" stroke-dasharray="5,2" points="604,-65 604,-437.33 946,-437.33 946,-65 604,-65"/>
<text xml:space="preserve" text-anchor="middle" x="775" y="-418.13" font-family="Helvetica,sans-Serif" font-size="16.00" fill="#8892a8">Configuration — weakest first, later wins</text>
</g>
<!-- dirname -->
<g id="node1" class="node">
<title>dirname</title>
<polygon fill="#1f6feb" stroke="#1e2a4a" points="375.11,-197.44 375.11,-219.63 329.94,-235.33 266.06,-235.33 220.89,-219.63 220.89,-197.44 266.06,-181.75 329.94,-181.75 375.11,-197.44"/>
<text xml:space="preserve" text-anchor="middle" x="298" y="-211.59" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#ffffff">directory name</text>
<text xml:space="preserve" text-anchor="middle" x="298" y="-198.09" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#ffffff">e.g. acmebank/</text>
<polygon fill="#1f6feb" stroke="#1e2a4a" points="370.11,-197.44 370.11,-219.63 324.94,-235.33 261.06,-235.33 215.89,-219.63 215.89,-197.44 261.06,-181.75 324.94,-181.75 370.11,-197.44"/>
<text xml:space="preserve" text-anchor="middle" x="293" y="-211.59" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#ffffff">directory name</text>
<text xml:space="preserve" text-anchor="middle" x="293" y="-198.09" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#ffffff">e.g. acmebank/</text>
</g>
<!-- cname -->
<g id="node2" class="node">
@@ -37,8 +37,8 @@
<!-- dirname&#45;&gt;cname -->
<g id="edge1" class="edge">
<title>dirname&#45;&gt;cname</title>
<path fill="none" stroke="#4a5568" d="M232.76,-193.05C195.81,-183.01 149.78,-167.28 113,-144.5 101.37,-137.29 90.31,-127.09 81.33,-117.6"/>
<polygon fill="#4a5568" stroke="#4a5568" points="84.04,-115.37 74.73,-110.3 78.85,-120.07 84.04,-115.37"/>
<path fill="none" stroke="#4a5568" d="M228.68,-192.51C192.88,-182.36 148.52,-166.7 113,-144.5 101.4,-137.25 90.34,-127.04 81.36,-117.55"/>
<polygon fill="#4a5568" stroke="#4a5568" points="84.07,-115.32 74.76,-110.26 78.88,-120.02 84.07,-115.32"/>
</g>
<!-- ctx -->
<g id="node3" class="node">
@@ -50,120 +50,120 @@
<!-- dirname&#45;&gt;ctx -->
<g id="edge2" class="edge">
<title>dirname&#45;&gt;ctx</title>
<path fill="none" stroke="#4a5568" d="M269.64,-181.32C248.71,-161.98 220.42,-135.83 199.86,-116.83"/>
<polygon fill="#4a5568" stroke="#4a5568" points="202.4,-114.41 192.68,-110.19 197.65,-119.55 202.4,-114.41"/>
<path fill="none" stroke="#4a5568" d="M265.77,-181.32C245.77,-162.07 218.77,-136.06 199.05,-117.08"/>
<polygon fill="#4a5568" stroke="#4a5568" points="201.55,-114.63 191.91,-110.21 196.69,-119.67 201.55,-114.63"/>
</g>
<!-- img -->
<g id="node4" class="node">
<title>img</title>
<polygon fill="#121829" stroke="#1e2a4a" points="354,-109 242,-109 242,-73 354,-73 354,-109"/>
<text xml:space="preserve" text-anchor="middle" x="298" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">image tag</text>
<text xml:space="preserve" text-anchor="middle" x="298" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">acmebank&#45;wizard</text>
<polygon fill="#121829" stroke="#1e2a4a" points="344.12,-109 241.88,-109 241.88,-73 344.12,-73 344.12,-109"/>
<text xml:space="preserve" text-anchor="middle" x="293" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">image tag</text>
<text xml:space="preserve" text-anchor="middle" x="293" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">acmebank&#45;deps</text>
</g>
<!-- dirname&#45;&gt;img -->
<g id="edge3" class="edge">
<title>dirname&#45;&gt;img</title>
<path fill="none" stroke="#4a5568" d="M298,-181.32C298,-163.19 298,-139.07 298,-120.47"/>
<polygon fill="#4a5568" stroke="#4a5568" points="301.5,-120.67 298,-110.67 294.5,-120.67 301.5,-120.67"/>
<path fill="none" stroke="#4a5568" d="M293,-181.32C293,-163.19 293,-139.07 293,-120.47"/>
<polygon fill="#4a5568" stroke="#4a5568" points="296.5,-120.67 293,-110.67 289.5,-120.67 296.5,-120.67"/>
</g>
<!-- ports -->
<g id="node5" class="node">
<title>ports</title>
<polygon fill="#121829" stroke="#1e2a4a" points="460.38,-109 371.62,-109 371.62,-73 460.38,-73 460.38,-109"/>
<text xml:space="preserve" text-anchor="middle" x="416" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">port block</text>
<text xml:space="preserve" text-anchor="middle" x="416" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">2130021309</text>
<polygon fill="#121829" stroke="#1e2a4a" points="451.38,-109 362.62,-109 362.62,-73 451.38,-73 451.38,-109"/>
<text xml:space="preserve" text-anchor="middle" x="407" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">port block</text>
<text xml:space="preserve" text-anchor="middle" x="407" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">2130021309</text>
</g>
<!-- dirname&#45;&gt;ports -->
<g id="edge4" class="edge">
<title>dirname&#45;&gt;ports</title>
<path fill="none" stroke="#4a5568" d="M324.94,-181.53C336.66,-170.19 350.54,-156.71 363,-144.5 372.11,-135.57 382.06,-125.74 390.85,-117.02"/>
<polygon fill="#4a5568" stroke="#4a5568" points="393.15,-119.67 397.79,-110.14 388.22,-114.7 393.15,-119.67"/>
<path fill="none" stroke="#4a5568" d="M318.87,-181.32C337.78,-162.15 363.29,-136.3 382,-117.34"/>
<polygon fill="#4a5568" stroke="#4a5568" points="384.47,-119.81 389,-110.24 379.49,-114.9 384.47,-119.81"/>
</g>
<!-- reg -->
<g id="node6" class="node">
<title>reg</title>
<polygon fill="#121829" stroke="#1e2a4a" points="597.38,-109 478.62,-109 478.62,-73 597.38,-73 597.38,-109"/>
<text xml:space="preserve" text-anchor="middle" x="538" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">registry container</text>
<text xml:space="preserve" text-anchor="middle" x="538" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">acmebank&#45;registry</text>
<polygon fill="#121829" stroke="#1e2a4a" points="588.38,-109 469.62,-109 469.62,-73 588.38,-73 588.38,-109"/>
<text xml:space="preserve" text-anchor="middle" x="529" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">registry container</text>
<text xml:space="preserve" text-anchor="middle" x="529" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">acmebank&#45;registry</text>
</g>
<!-- dirname&#45;&gt;reg -->
<g id="edge5" class="edge">
<title>dirname&#45;&gt;reg</title>
<path fill="none" stroke="#4a5568" d="M356.87,-190.63C390.74,-179.74 433.49,-163.98 469,-144.5 483.3,-136.65 497.86,-126.04 509.9,-116.41"/>
<polygon fill="#4a5568" stroke="#4a5568" points="511.88,-119.31 517.39,-110.26 507.44,-113.9 511.88,-119.31"/>
<path fill="none" stroke="#4a5568" d="M350.86,-190.3C383.87,-179.35 425.43,-163.64 460,-144.5 474.27,-136.6 488.83,-125.98 500.87,-116.36"/>
<polygon fill="#4a5568" stroke="#4a5568" points="502.85,-119.26 508.36,-110.21 498.41,-113.84 502.85,-119.26"/>
</g>
<!-- cluster -->
<g id="node11" class="node">
<title>cluster</title>
<polygon fill="#1a1a3a" stroke="#1e2a4a" points="469.81,-10.54 469.81,-25.46 438.29,-36 393.71,-36 362.19,-25.46 362.19,-10.54 393.71,0 438.29,0 469.81,-10.54"/>
<text xml:space="preserve" text-anchor="middle" x="416" y="-14.3" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#0066ff">kind cluster</text>
<polygon fill="#1a1a3a" stroke="#1e2a4a" points="460.81,-10.54 460.81,-25.46 429.29,-36 384.71,-36 353.19,-25.46 353.19,-10.54 384.71,0 429.29,0 460.81,-10.54"/>
<text xml:space="preserve" text-anchor="middle" x="407" y="-14.3" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#0066ff">kind cluster</text>
</g>
<!-- cname&#45;&gt;cluster -->
<g id="edge9" class="edge">
<title>cname&#45;&gt;cluster</title>
<path fill="none" stroke="#4a5568" d="M93.06,-72.62C99.54,-69.72 106.38,-67.02 113,-65 192.78,-40.7 288.45,-28.91 350.65,-23.42"/>
<polygon fill="#4a5568" stroke="#4a5568" points="350.69,-26.93 360.36,-22.6 350.1,-19.96 350.69,-26.93"/>
<path fill="none" stroke="#4a5568" d="M93.35,-72.51C99.75,-69.67 106.48,-67 113,-65 189.55,-41.49 281.19,-29.58 341.58,-23.85"/>
<polygon fill="#4a5568" stroke="#4a5568" points="341.71,-27.35 351.35,-22.96 341.07,-20.38 341.71,-27.35"/>
</g>
<!-- ports&#45;&gt;cluster -->
<g id="edge10" class="edge">
<title>ports&#45;&gt;cluster</title>
<path fill="none" stroke="#4a5568" d="M416,-72.81C416,-65.23 416,-56.1 416,-47.54"/>
<polygon fill="#4a5568" stroke="#4a5568" points="419.5,-47.54 416,-37.54 412.5,-47.54 419.5,-47.54"/>
<path fill="none" stroke="#4a5568" d="M407,-72.81C407,-65.23 407,-56.1 407,-47.54"/>
<polygon fill="#4a5568" stroke="#4a5568" points="410.5,-47.54 407,-37.54 403.5,-47.54 410.5,-47.54"/>
</g>
<!-- versions -->
<g id="node7" class="node">
<title>versions</title>
<polygon fill="#121829" stroke="#1e2a4a" points="759.38,-401.83 652.62,-401.83 652.62,-365.83 759.38,-365.83 759.38,-401.83"/>
<text xml:space="preserve" text-anchor="middle" x="706" y="-386.88" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">versions.env</text>
<text xml:space="preserve" text-anchor="middle" x="706" y="-373.38" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">pinned toolchain</text>
<polygon fill="#121829" stroke="#1e2a4a" points="750.38,-401.83 643.62,-401.83 643.62,-365.83 750.38,-365.83 750.38,-401.83"/>
<text xml:space="preserve" text-anchor="middle" x="697" y="-386.88" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">versions.env</text>
<text xml:space="preserve" text-anchor="middle" x="697" y="-373.38" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">pinned toolchain</text>
</g>
<!-- profile -->
<g id="node8" class="node">
<title>profile</title>
<polygon fill="#121829" stroke="#1e2a4a" points="790.5,-318.58 621.5,-318.58 621.5,-282.58 790.5,-282.58 790.5,-318.58"/>
<text xml:space="preserve" text-anchor="middle" x="706" y="-303.63" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">env.d/&lt;profile&gt;.env</text>
<text xml:space="preserve" text-anchor="middle" x="706" y="-290.13" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">nodes · CNI · audit · addons</text>
<polygon fill="#121829" stroke="#1e2a4a" points="781.5,-318.58 612.5,-318.58 612.5,-282.58 781.5,-282.58 781.5,-318.58"/>
<text xml:space="preserve" text-anchor="middle" x="697" y="-303.63" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">env.d/&lt;profile&gt;.env</text>
<text xml:space="preserve" text-anchor="middle" x="697" y="-290.13" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">nodes · CNI · audit · addons</text>
</g>
<!-- versions&#45;&gt;profile -->
<g id="edge6" class="edge">
<title>versions&#45;&gt;profile</title>
<path fill="none" stroke="#4a5568" d="M706,-365.59C706,-355.32 706,-342.03 706,-330.21"/>
<polygon fill="#4a5568" stroke="#4a5568" points="709.5,-330.58 706,-320.58 702.5,-330.58 709.5,-330.58"/>
<text xml:space="preserve" text-anchor="middle" x="737.5" y="-339.28" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">overridden by</text>
<path fill="none" stroke="#4a5568" d="M697,-365.59C697,-355.32 697,-342.03 697,-330.21"/>
<polygon fill="#4a5568" stroke="#4a5568" points="700.5,-330.58 697,-320.58 693.5,-330.58 700.5,-330.58"/>
<text xml:space="preserve" text-anchor="middle" x="728.5" y="-339.28" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">overridden by</text>
</g>
<!-- localenv -->
<g id="node9" class="node">
<title>localenv</title>
<polygon fill="#121829" stroke="#1e2a4a" points="761.88,-226.54 646.12,-226.54 646.12,-190.54 761.88,-190.54 761.88,-226.54"/>
<text xml:space="preserve" text-anchor="middle" x="704" y="-211.59" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">ctrl/.env</text>
<text xml:space="preserve" text-anchor="middle" x="704" y="-198.09" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">secrets, overrides</text>
<polygon fill="#121829" stroke="#1e2a4a" points="752.88,-226.54 637.12,-226.54 637.12,-190.54 752.88,-190.54 752.88,-226.54"/>
<text xml:space="preserve" text-anchor="middle" x="695" y="-211.59" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">ctrl/.env</text>
<text xml:space="preserve" text-anchor="middle" x="695" y="-198.09" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#e8eaf0">secrets, overrides</text>
</g>
<!-- profile&#45;&gt;localenv -->
<g id="edge7" class="edge">
<title>profile&#45;&gt;localenv</title>
<path fill="none" stroke="#4a5568" d="M705.61,-282.22C705.34,-269.76 704.96,-252.69 704.64,-238.23"/>
<polygon fill="#4a5568" stroke="#4a5568" points="708.14,-238.28 704.42,-228.36 701.14,-238.43 708.14,-238.28"/>
<text xml:space="preserve" text-anchor="middle" x="736.68" y="-256.03" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">overridden by</text>
<path fill="none" stroke="#4a5568" d="M696.61,-282.22C696.34,-269.76 695.96,-252.69 695.64,-238.23"/>
<polygon fill="#4a5568" stroke="#4a5568" points="699.14,-238.28 695.42,-228.36 692.14,-238.43 699.14,-238.28"/>
<text xml:space="preserve" text-anchor="middle" x="727.68" y="-256.03" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">overridden by</text>
</g>
<!-- shell -->
<g id="node10" class="node">
<title>shell</title>
<polygon fill="#1a3a1a" stroke="#1e2a4a" points="772.38,-109 623.62,-109 623.62,-73 772.38,-73 772.38,-109"/>
<text xml:space="preserve" text-anchor="middle" x="698" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#00c853">the environment</text>
<text xml:space="preserve" text-anchor="middle" x="698" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#00c853">PROFILE=client make …</text>
<polygon fill="#1a3a1a" stroke="#1e2a4a" points="763.38,-109 614.62,-109 614.62,-73 763.38,-73 763.38,-109"/>
<text xml:space="preserve" text-anchor="middle" x="689" y="-94.05" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#00c853">the environment</text>
<text xml:space="preserve" text-anchor="middle" x="689" y="-80.55" font-family="Helvetica,sans-Serif" font-size="11.00" fill="#00c853">PROFILE=client make …</text>
</g>
<!-- localenv&#45;&gt;shell -->
<g id="edge8" class="edge">
<title>localenv&#45;&gt;shell</title>
<path fill="none" stroke="#00c853" d="M703.11,-190.49C702.16,-172.16 700.63,-142.72 699.49,-120.79"/>
<polygon fill="#00c853" stroke="#00c853" points="703,-120.81 698.99,-111.01 696.01,-121.18 703,-120.81"/>
<text xml:space="preserve" text-anchor="middle" x="733.21" y="-155.2" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">overridden by</text>
<path fill="none" stroke="#00c853" d="M694.11,-190.49C693.16,-172.16 691.63,-142.72 690.49,-120.79"/>
<polygon fill="#00c853" stroke="#00c853" points="694,-120.81 689.99,-111.01 687.01,-121.18 694,-120.81"/>
<text xml:space="preserve" text-anchor="middle" x="724.21" y="-155.2" font-family="Helvetica,sans-Serif" font-size="9.00" fill="#8892a8">overridden by</text>
</g>
<!-- shell&#45;&gt;cluster -->
<g id="edge11" class="edge">
<title>shell&#45;&gt;cluster</title>
<path fill="none" stroke="#4a5568" stroke-dasharray="5,2" d="M637.29,-72.6C627.83,-69.99 618.17,-67.38 609,-65 562.72,-52.99 509.89,-40.48 471.22,-31.54"/>
<polygon fill="#4a5568" stroke="#4a5568" points="472.2,-28.18 461.67,-29.35 470.63,-35 472.2,-28.18"/>
<path fill="none" stroke="#4a5568" stroke-dasharray="5,2" d="M628.29,-72.6C618.83,-69.99 609.17,-67.38 600,-65 553.72,-52.99 500.89,-40.48 462.22,-31.54"/>
<polygon fill="#4a5568" stroke="#4a5568" points="463.2,-28.18 452.67,-29.35 461.63,-35 463.2,-28.18"/>
</g>
</g>
</svg>

Before

Width:  |  Height:  |  Size: 11 KiB

After

Width:  |  Height:  |  Size: 11 KiB

View File

@@ -265,11 +265,11 @@
<h3>The only prerequisite</h3>
<p><b>Docker.</b> No curl, no jq, no python, no apt repositories to configure.</p>
<pre><code><span class="c"># then, in the environment directory:</span>
make station <span class="c"># is this workstation ready? reports, never fixes</span>
make check <span class="c"># is this machine ready? reports, never fixes</span>
make deps <span class="c"># install the pinned toolchain</span>
make cluster up <span class="c"># build the cluster for the active profile</span>
</code></pre>
<p>Read <code>make station</code> before <code>make deps</code>. It never changes
<p>Read <code>make check</code> before <code>make deps</code>. It never changes
anything — it prints what it found and, at the end, the steps it cannot perform
for you.</p>
</div>
@@ -280,13 +280,13 @@ make cluster up <span class="c"># build the cluster for the active profile</spa
<p class="lede">Start to finish, in order, with what each one actually does.</p>
<div class="prose">
<h3>1 &middot; make station</h3>
<p>Asks whether this workstation is ready. It <b>changes nothing</b> — it
<h3>1 &middot; make check</h3>
<p>Asks whether this machine is ready. It <b>changes nothing</b> — it
reports what it found and, at the end, the things only a human can do
(anything needing <code>sudo</code>, or a Windows-side restart). Read it
before installing anything; it is faster than discovering the same problems
one failure at a time.</p>
<pre><code>make station</code></pre>
<pre><code>make check</code></pre>
<h3>2 &middot; make setup</h3>
<p>Does the preparation that can be automated: installs the pinned
@@ -298,7 +298,6 @@ make cluster up <span class="c"># build the cluster for the active profile</spa
unfamiliar machine the complete list is the point. The tail of the output is
a to-do list of only what is outstanding.</p>
<pre><code>make setup <span class="c"># host + toolchain</span>
make setup --share-docker <span class="c"># ...and offer this machine's Docker to other distros</span>
</code></pre>
<h3>3 &middot; make cluster up</h3>
@@ -309,8 +308,8 @@ make setup --share-docker <span class="c"># ...and offer this machine's Docker
attempt was interrupted before the CNI was installed, running it again
finishes the job rather than reporting "already exists" and leaving every
node permanently NotReady.</p>
<pre><code>make cluster up <span class="c"># default profile</span>
make cluster up PROFILE=client <span class="c"># three nodes, audit on, cached registry</span>
<pre><code>make cluster up <span class="c"># built-in defaults — no profile needed</span>
make cluster up PROFILE=client <span class="c"># after copying env.d/client.env.example: three nodes, audit, cached registry</span>
make cluster reset <span class="c"># destroy and rebuild — the only way to change CNI or audit</span>
</code></pre>
@@ -325,7 +324,6 @@ make cluster reset <span class="c"># destroy and rebuild — the on
<dt>make cluster list</dt><dd>Every cluster on the machine, its memory cost and its port block. The usual reason a new one will not start is an old one you forgot about; <code>make cluster free</code> frees them without deleting.</dd>
<dt>make ports</dt><dd>This environment's port block, and whether each is derived or overridden.</dd>
<dt>make registry</dt><dd>Which of the four registry modes is active, and where it points.</dd>
<dt>make dockerhost</dt><dd>Which WSL distro owns Docker and what this one is using.</dd>
</dl>
<h3>Running more than one</h3>
@@ -381,10 +379,10 @@ make setup core <span class="c"># same distinction, via setup</span>
wants only Docker.</p>
<h3>Air-gapped</h3>
<pre><code>make wizard full <span class="c"># bakes every binary into the image</span>
docker save …-wizard:full | gzip &gt; rig.tgz
<pre><code>make deps-image full <span class="c"># bakes every binary into the image</span>
docker save …-deps:full | gzip &gt; rig.tgz
<span class="c"># carry that one file in, then:</span>
docker load &lt; rig.tgz &amp;&amp; make cluster up PROFILE=offline
docker load &lt; rig.tgz &amp;&amp; make cluster up PROFILE=offline <span class="c"># from env.d/offline.env.example</span>
</code></pre>
</div>
</section>
@@ -410,9 +408,9 @@ docker load &lt; rig.tgz &amp;&amp; make cluster up PROFILE=offline
<code>ctrl/.env</code> if you want it fixed rather than derived.</p>
<h3>Configuration layers</h3>
<p>Weakest first, later wins: pinned versions → the profile →
<code>ctrl/.env</code> → the environment. So
<code>make cluster up PROFILE=client</code> always beats every file.</p>
<p>Weakest first, later wins: built-in defaults → pinned versions → a
profile, if you name one → <code>ctrl/.env</code> → the environment. So
<code>make cluster up PROFILE=&lt;name&gt;</code> always beats every file.</p>
</div>
</section>
@@ -485,7 +483,7 @@ docker load &lt; rig.tgz &amp;&amp; make cluster up PROFILE=offline
registry is usually behind an internal CA, and trust has to reach
<b>three</b> places: the host Docker daemon, every cluster node's containerd
(nodes do <i>not</i> inherit host trust), and any in-cluster client. Set
<code>REGISTRY_CA_FILE</code> and <code>make station</code> reports which is
<code>REGISTRY_CA_FILE</code> and <code>make check</code> reports which is
still missing. The symptom otherwise is an opaque
<code>x509: certificate signed by unknown authority</code>.</p></div>
@@ -528,11 +526,11 @@ docker load &lt; rig.tgz &amp;&amp; make cluster up PROFILE=offline
<h3>Tilt stops noticing file changes</h3>
<p>Almost always <code>inotify</code> limits, and it fails <i>silently</i>
nothing errors, changes just stop being picked up. Defaults on WSL are far too
low. <code>make station</code> reports it and prints the fix.</p>
low. <code>make check</code> reports it and prints the fix.</p>
<h3>Cluster creation dies halfway with a port error</h3>
<p>Docker reports <code>failed to bind host port … address already in use</code>
partway through creating the cluster. Run <code>make station</code> first — it
partway through creating the cluster. Run <code>make check</code> first — it
checks every port in this environment's block before anything is built.</p>
<h3>Every node stays NotReady</h3>

View File

@@ -1,9 +0,0 @@
# NOTE: generated/ is deliberately NOT ignored. The artifact IS the deliverable —
# the whole point is a folder you copy, apply and boot without generating
# anything first. Regenerate it with `make manifest` after editing bundle.json
# or app/serve.py, and commit the result.
#
# (This differs from rig's ctrl/k8s/generated, which is a local build artifact.)
__pycache__/
*.pyc

View File

@@ -1,50 +0,0 @@
# Thin control Makefile — one target per ctrl/ script, subcommand as an
# argument. Same shape as rig's, for the same reason: the logic lives in the
# script, never here.
#
# make up -> ctrl/bundle.sh up
# make bundle down
#
# This directory is a BUNDLE, not an installer. It needs a cluster, which rig
# owns:
#
# cd .. && make cluster up # kind cluster for this environment
# make up # then deploy this bundle into it
#
# Start with: make up
ARGS := $(wordlist 2,$(words $(MAKECMDGOALS)),$(MAKECMDGOALS))
ifneq ($(ARGS),)
$(eval $(ARGS):;@:)
endif
.DEFAULT_GOAL := help
.PHONY: help bundle manifest up down status url list dev
help: ## list targets
@grep -hE '^[a-z]+:.*?##' $(MAKEFILE_LIST) | sed 's/:.*##/\t/' | expand -t16
bundle: ## the bundle [manifest|up|down|status|url|list|dev]
bash ctrl/bundle.sh $(or $(ARGS),status)
# Shorthands for the ones used constantly.
manifest: ## regenerate generated/<slug>.yaml — no cluster needed
bash ctrl/bundle.sh manifest
up: ## deploy this rig (installs MetalLB if absent)
bash ctrl/bundle.sh up
down: ## remove this rig (leaves cluster, MetalLB, siblings)
bash ctrl/bundle.sh down
status: ## what is deployed for this rig
bash ctrl/bundle.sh status
url: ## the address MetalLB assigned
bash ctrl/bundle.sh url
list: ## every rig in this cluster, with addresses
bash ctrl/bundle.sh list
dev: ## run the UI locally with vite — no cluster needed
bash ctrl/bundle.sh dev

View File

@@ -1,140 +0,0 @@
# sample-rig
A minimal, non-sensitive bundle that proves an installation works and shows what
shipped. Copy it, rename it, and you have another rig.
```bash
make manifest # generate the artifact — no cluster, no kubectl needed
make up # deploy it into the local cluster
make list # every rig in this cluster, with addresses
make dev # run the UI locally with vite, no cluster at all
```
`make up` prints an address. Open it and the page says **IT WORKS**, then lists
the tools and rigs in the bundle.
## What it is for
Three jobs, in the order you hit them:
1. **Prove the install.** kind is there, a cluster exists, MetalLB hands out an
address, a `type: LoadBalancer` Service actually resolves, and a pod serves.
If all of that works, the environment is sound.
2. **Say what shipped.** The page renders [`bundle.json`](bundle.json) —
standalone tools and rigs, flat, with none of soleprint's internal hierarchy.
Editing that file is the only step needed to change the listing.
3. **Stand in for the real thing.** Nothing here is sensitive. The real
architecture connects separately, against a setup already known to work.
## The UI is a complement, not the product
`rig-ui/` is just a vite app. It complements a rig; a rig is complete and useful
without it, and nothing depends on it being there. It is deliberately **not**
generated by kind or tilt — you copy the folder into a rig after that rig is
pulled, and apply one manifest:
```bash
kubectl apply -n <namespace> -f rig-ui/k8s.yaml
```
That file is the whole integration: one Pod running `npm run dev` on
`node:22-alpine`, one Service. A bare Pod rather than a Deployment because this
is a dev-loop convenience, not a workload to keep alive.
The app and `bundle.json` arrive as a ConfigMap, so nothing is baked into an
image and editing the bundle is the entire update cycle. The container runs
`npm install` at start, which needs egress to a registry — on a locked-down
cluster point npm at the internal one, or bake an image instead. Nothing else
changes if you do.
## One artifact, two destinations
`ctrl/manifest.py` emits `generated/<slug>.yaml` — namespace, the app and
bundle embedded in a ConfigMap, Pod, Service. It is self-contained and applies
unmodified anywhere:
```bash
kubectl apply -f generated/sample-rig.yaml # local kind, or an external cluster
```
`make up` applies **that same file**. There is no separate local path, so what
works here cannot quietly differ from what is applied elsewhere.
This is what `type: LoadBalancer` buys. MetalLB answers it on kind; the AWS load
balancer controller answers it on EKS. NodePort would not survive the trip — it
is a single cluster-wide port range, so two rigs would have to negotiate numbers.
**VPC-agnostic on purpose.** The target is EKS, but the Service carries no
annotations — no `aws-load-balancer-subnets`, no security groups, no `-scheme`,
no `-type: nlb`. Each of those encodes a specific network layout, and one of them
appearing here would pin the artifact to the account and VPC it was written
against, which is precisely what stops it also working on kind. Subnet discovery
is the cluster's business: EKS resolves it from the tags its own subnets carry.
That leaves one thing genuinely environment-specific — internal versus
internet-facing. A bare `LoadBalancer` provisions internet-facing, which a
regulated account will usually refuse, and should. That belongs in a
per-environment overlay applied on top, never inlined into this artifact.
**MetalLB only — no ingress-nginx.** Its controller supports a narrow window of
Kubernetes versions, so depending on it constrains which k8s a rig can be built
with. That undercuts running trailing-edge control planes to model a legacy
estate, which is the reason `versions.env` pins v1_33..v1_36. MetalLB carries no
such constraint, so reachability costs nothing in version coverage.
## Several rigs, one cluster
Identity follows the **folder name**, the same rule rig uses for cluster
identity. The namespace is the folder; resource names are generic, and names only
have to be unique within a namespace.
```bash
cp -r sample-rig corporate-rig
cd corporate-rig && make up # its own namespace, its own address
```
No edits, no collisions, both in the same local cluster. `make list` shows them
together. `make down` removes only this one — siblings, MetalLB and the cluster
are left alone.
Client rigs are gitignored (`*-rig/`, with `sample-rig/` the deliberate
exception): a rig's k8s files spell out a real architecture, and that is exactly
what must not land in this repo.
## Staging workstations
`ctrl/manifest.py` is stdlib-only on purpose: it runs on a bare machine before
anything is installed. The toolchain itself is rig's job — `make deps` installs
the pinned kind and tilt binaries, which is what makes a staging AWS workspace
reachable from the same commands as a laptop.
## Layout
```
sample-rig/
├── Makefile # thin — one target per ctrl/ script
├── bundle.json # what shipped; the UI renders THIS
├── rig-ui/ # the vite app — optional, copied into a rig to enable it
│ ├── k8s.yaml # how to plug it in: one Pod, one Service
│ ├── index.html
│ ├── package.json
│ ├── vite.config.js
│ └── src/{main.js,style.css}
├── ctrl/
│ ├── manifest.py # emits the artifact
│ └── bundle.sh # generate / deploy / inspect
└── generated/ # the artifact — committed, this is the deliverable
```
Editing `bundle.json` or anything in `rig-ui/` means re-running `make manifest`.
The ConfigMap carries a checksum of everything embedded, so a stale deployment is
visible rather than silent.
## Not built, but not foreclosed
Everything derives from `bundle.json` plus a target namespace. A Pulumi or
Terraform emitter would sit beside `ctrl/manifest.py` consuming the same inputs;
nothing above it assumes the artifact is YAML.
Licence terms for the compiled UI component belong in the soleprint-generated
bundle, not here — this sample carries no proprietary component.

View File

@@ -1,51 +0,0 @@
{
"_comment": "What this bundle contains. Single source of truth — the landing page renders THIS file, so adding an entry here is the only edit needed. Deliberately FLAT: standalone tools and rigs, with none of soleprint's internal hierarchy (no artery/atlas/station layering). Nothing here is sensitive; the real architecture connects separately.",
"bundle": {
"name": "sample-rig",
"description": "Non-sensitive sample bundle. Proves the kind install works and shows what ships.",
"sensitive": false
},
"tools": [
{
"name": "modelgen",
"summary": "Generate models from config",
"standalone": true
},
{
"name": "datagen",
"summary": "Generate test data from rig-owned generators",
"standalone": true
},
{
"name": "graphgen",
"summary": "Generate navigable model graphs",
"standalone": true
},
{
"name": "tester",
"summary": "HTTP contract test runner — one suite, any environment",
"standalone": true
},
{
"name": "databrowse",
"summary": "SQL data browser",
"standalone": true
},
{
"name": "sbwrapper",
"summary": "Sandbox wrapper",
"standalone": true
}
],
"rigs": [
{
"name": "sample-rig",
"summary": "This bundle — a minimal, copyable environment",
"active": true
}
],
"next": [
"Point MANIFESTS_DIR at the real manifests to connect the actual architecture.",
"Real k8s files are versioned separately and are not part of this bundle."
]
}

View File

@@ -1,40 +0,0 @@
{
"_comment": "MOCKED cluster state. Nothing here is read from a live cluster — it exists so the UI can be shown when there is no cluster at all (a locked-down machine, a laptop with no memory to spare, a demo where kind will not start). The page labels it as mocked; a demo that looks live but is not is worse than one that says so. When a real cluster is present the same shapes come from kubectl.",
"mocked": true,
"cluster": {
"name": "sample-rig",
"context": "kind-sample-rig",
"provider": "kind",
"profile": "minimal",
"k8s": "v1.36.1",
"nodes": 1
},
"workloads": [
{
"name": "rig-ui",
"summary": "Pod · node:22-alpine · vite on :5173",
"state": "Running"
},
{
"name": "metallb-system/controller",
"summary": "Deployment · assigns LoadBalancer addresses",
"state": "Running"
},
{
"name": "metallb-system/speaker",
"summary": "DaemonSet · answers ARP in layer 2 mode",
"state": "Running"
}
],
"services": [
{
"name": "rig-ui",
"summary": "LoadBalancer · 80 -> 5173 · no annotations, so it resolves on kind and on EKS alike",
"state": "172.18.255.200"
}
]
}

View File

@@ -1,223 +0,0 @@
#!/usr/bin/env bash
# The rig bundle: generate it, deploy it, tear it down, find it.
#
# Usage: bundle.sh manifest | up | down | status | url | list | dev
#
# What `up` proves, in order: kind installed and a cluster exists, MetalLB can
# hand out an address, a Service of type LoadBalancer actually resolves, and a
# pod serves the bundle listing. If all of that works the installation is sound,
# and the only thing missing is the real architecture.
#
# ONE ARTIFACT
# `up` applies generated/<slug>.yaml — the same self-contained file you would
# hand to an external cluster. There is no separate local path, so what works
# here cannot quietly differ from the master deployment applied elsewhere.
#
# ONE CLUSTER, SEVERAL RIGS
# Identity follows the FOLDER NAME, exactly as rig's cluster identity does. This
# directory deploys into a namespace named after itself, so copying it to
# corporate-rig/ yields a second rig in the SAME local cluster with no edits and
# no collisions — different namespace, its own MetalLB address. `list` shows all
# of them. The cluster itself is rig's business; this only ever owns a namespace.
#
# MetalLB is installed by calling rig's own addon script rather than
# reimplementing it — deriving the pool from the kind Docker network is the
# fiddly part and there should be exactly one copy of it.
set -euo pipefail
cd "$(dirname "$0")/.."
BUNDLE_ROOT="$(pwd)"
RIG_CTRL="$(cd .. && pwd)/ctrl"
# The containing folder's name, reduced to a DNS label (same rule as rig's
# default_cluster_name and ctrl/manifest.py, so all three agree on the slug).
slug() {
local n
n=$(basename "$BUNDLE_ROOT")
n=$(echo "$n" | tr '[:upper:]' '[:lower:]' | tr -c 'a-z0-9-' '-')
n=$(echo "$n" | sed 's/^-*//; s/-*$//')
echo "${n:-rig-bundle}"
}
NS="$(slug)"
ARTIFACT="generated/${NS}.yaml"
# Resolved lazily, not at load time: `manifest` and `dev` deliberately work
# with no cluster and no kubectl at all, and a top-level check would break that.
#
# Follows whatever context rig's cluster.sh selected, so this bundle works in a
# copied-and-renamed environment without being told which cluster it is in.
init_kube() {
KUBECONTEXT="${KUBECONTEXT:-$(kubectl config current-context 2>/dev/null || true)}"
if [ -z "$KUBECONTEXT" ]; then
echo "no kubectl context — bring a cluster up first: (cd .. && make cluster up)" >&2
exit 1
fi
KCTX="kubectl --context ${KUBECONTEXT}"
K="kubectl --context ${KUBECONTEXT} --namespace ${NS}"
}
require_cluster() {
if ! $KCTX cluster-info >/dev/null 2>&1; then
echo "context '$KUBECONTEXT' does not reach a cluster" >&2
echo "bring one up: (cd .. && make cluster up)" >&2
exit 1
fi
}
ensure_metallb() {
if $KCTX get deployment -n metallb-system controller >/dev/null 2>&1; then
echo "metallb: present"
return 0
fi
# Only kind needs it. On a real cluster the cloud load balancer answers a
# `type: LoadBalancer` Service, and installing MetalLB there would be wrong.
case "$KUBECONTEXT" in
kind-*) ;;
*)
echo "metallb: skipped — '$KUBECONTEXT' is not a kind context"
echo " (a cloud load balancer answers LoadBalancer services there)"
return 0
;;
esac
if [ ! -f "$RIG_CTRL/addons/metallb.sh" ]; then
echo "metallb is not installed and rig's addon script was not found at" >&2
echo " $RIG_CTRL/addons/metallb.sh" >&2
echo "a Service of type LoadBalancer will sit at <pending> without it." >&2
exit 1
fi
# rig's addons derive their target cluster from RIG'S OWN folder name via
# load_config, so left alone this bundle would install into `kind-rig` —
# a cluster that need not exist — while deploying everything else into the
# context actually selected. CLUSTER is in load_config's overridable set,
# so passing it here points the addon at the same cluster we are using.
local target="${KUBECONTEXT#kind-}"
echo "metallb: installing via rig's addon into '$target'"
CLUSTER="$target" bash "$RIG_CTRL/addons/metallb.sh"
}
# Regenerate the artifact. No cluster and no kubectl required — this is the step
# a staging workstation runs before anything is installed.
manifest() {
mkdir -p generated
python3 ctrl/manifest.py "$NS" > "$ARTIFACT"
echo "wrote $ARTIFACT ($(wc -l < "$ARTIFACT") lines)"
echo " applies as-is anywhere: kubectl apply -f ${BUNDLE_ROOT}/${ARTIFACT}"
}
up() {
manifest
init_kube
require_cluster
ensure_metallb
echo
echo "applying '${NS}' to context '${KUBECONTEXT}'"
$KCTX apply -f "$ARTIFACT"
# `rollout status` does not work on a bare Pod — it only understands
# Deployments, StatefulSets and DaemonSets. Wait on the condition instead.
# This is the slow step: the container npm-installs before vite serves.
echo "waiting for the pod to be ready (npm install runs first)..."
$K wait --for=condition=Ready pod/rig-ui --timeout=300s
echo
url
}
down() {
init_kube
# Delete the namespace and everything in it goes with it. Scoped to THIS
# rig — a sibling rig in the same cluster is untouched.
$KCTX delete namespace "$NS" --ignore-not-found
echo "'${NS}' removed (cluster, metallb and any sibling rig are left alone)"
}
status() {
init_kube
require_cluster
if ! $KCTX get namespace "$NS" >/dev/null 2>&1; then
echo "'${NS}' is not deployed — run: make up"
return 0
fi
$K get pod,svc,configmap -o wide
}
# Every rig in this cluster, not just this one — the point of the namespace
# split is that several coexist, so there has to be a way to see them together.
list() {
init_kube
require_cluster
local names
names=$($KCTX get namespace -l rig.bundle/name \
-o jsonpath='{.items[*].metadata.name}' 2>/dev/null || true)
if [ -z "$names" ]; then
echo "no rigs deployed in context '${KUBECONTEXT}'"
return 0
fi
printf "%-20s %-16s %s\n" RIG ADDRESS ""
local n ip
for n in $names; do
ip=$($KCTX -n "$n" get svc rig-ui \
-o jsonpath='{.status.loadBalancer.ingress[0].ip}' 2>/dev/null || true)
printf "%-20s %-16s %s\n" "$n" "${ip:-<pending>}" \
"$([ "$n" = "$NS" ] && echo '<- this one')"
done
}
# The address MetalLB (or a cloud load balancer) assigned. <pending> here is the
# classic silent failure: everything reports healthy and nothing is reachable.
url() {
init_kube
local ip
ip=$($K get svc rig-ui \
-o jsonpath='{.status.loadBalancer.ingress[0].ip}' 2>/dev/null || true)
if [ -z "$ip" ]; then
ip=$($K get svc rig-ui \
-o jsonpath='{.status.loadBalancer.ingress[0].hostname}' 2>/dev/null || true)
fi
if [ -z "$ip" ]; then
echo "no external address yet — nothing has assigned one."
echo "on kind: kubectl --context $KUBECONTEXT -n metallb-system get pods"
return 1
fi
echo "IT WORKS -> http://${ip}/"
echo " bundle http://${ip}/bundle.json"
}
# Run the UI locally with no cluster at all — the fast way to iterate on
# bundle.json. Same vite command the pod runs, so what you see here is what
# gets served there.
dev() {
if ! command -v npm >/dev/null 2>&1; then
echo "npm not found — the UI needs node locally for this." >&2
echo "(in-cluster it runs on the node:22-alpine image instead)" >&2
exit 1
fi
# bundle.json lives one level up so it stays the rig's data rather than the
# app's; vite serves public/ at the root, which is where the app fetches it.
mkdir -p rig-ui/public
cp bundle.json rig-ui/public/bundle.json
# The mocked cluster is a DEMO asset and is deliberately not embedded in the
# deployed artifact — on a real rig the UI would then show canned values
# beside a live cluster, which is precisely the lie its banner warns about.
# It is served here, and in the static build for the public UI-only page.
cp cluster.mock.json rig-ui/public/cluster.mock.json
cd rig-ui
[ -d node_modules ] || npm install --no-audit --no-fund
VITE_RIG_NAME="$NS" npm run dev
}
case "${1:-status}" in
manifest) manifest ;;
up) up ;;
down) down ;;
status) status ;;
url) url ;;
list) list ;;
dev) dev ;;
*) echo "usage: $0 [manifest|up|down|status|url|list|dev]" >&2; exit 1 ;;
esac

View File

@@ -1,141 +0,0 @@
#!/usr/bin/env python3
"""Emit the complete, self-contained deployment for this rig.
python3 ctrl/manifest.py [namespace] > generated/<slug>.yaml
The output is the ARTIFACT. It carries everything — namespace, the vite app and
bundle.json embedded in a ConfigMap, the Pod and the Service — so it applies
unmodified to any cluster:
kubectl apply -f generated/sample-rig.yaml
On kind, MetalLB answers the `type: LoadBalancer` Service. On a real external
cluster the cloud load balancer does. Same file, no edits, no branch — which is
the point: what runs locally is byte-identical to the deployment applied
elsewhere, so local success actually means something.
`ctrl/bundle.sh up` applies this same generated output rather than a separate
code path, so the local convenience wrapper can never drift from the artifact.
Stdlib only, deliberately: this must run on a bare staging workstation before
anything is installed, so it cannot depend on PyYAML or a template engine.
Open seam — not built: everything here derives from bundle.json plus a target
namespace. A Pulumi or Terraform emitter would sit beside this file consuming the
same inputs; nothing above it assumes the artifact is YAML.
"""
import json
import re
import sys
import zlib
from pathlib import Path
ROOT = Path(__file__).resolve().parent.parent
UI = ROOT / "rig-ui"
# Files embedded into the ConfigMap, mounted read-only at /src in the pod and
# copied into vite's layout at start (see rig-ui/k8s.yaml). Flat on purpose:
# ConfigMap keys cannot contain '/'.
EMBEDDED = {
"bundle.json": ROOT / "bundle.json",
"package.json": UI / "package.json",
"vite.config.js": UI / "vite.config.js",
"index.html": UI / "index.html",
"main.js": UI / "src" / "main.js",
"style.css": UI / "src" / "style.css",
}
def slug(name: str) -> str:
"""Reduce a folder name to a DNS label, matching ctrl/bundle.sh's rule."""
out = re.sub(r"[^a-z0-9-]", "-", name.lower()).strip("-")
return out or "rig-bundle"
def block(text: str, indent: int) -> str:
"""Indent a file's contents for a YAML literal block scalar.
Blank lines are emitted truly empty rather than as whitespace: trailing
spaces on an otherwise blank line are legal YAML but show up as diff noise
in a committed artifact.
"""
pad = " " * indent
return "\n".join(pad + line if line.strip() else "" for line in text.splitlines())
def checksum(parts: list[str]) -> str:
"""Stable content hash of everything embedded, stamped as a label.
A mounted ConfigMap updates in place without restarting anything, so without
a visible change nothing signals that the pod is serving stale content.
"""
return str(zlib.crc32("".join(parts).encode()) & 0xFFFFFFFF)
def build(namespace: str) -> str:
contents = {}
for key, path in EMBEDDED.items():
if not path.exists():
sys.exit(f"missing input: {path}")
contents[key] = path.read_text()
# Fail loudly here rather than shipping an artifact that renders an error.
try:
json.loads(contents["bundle.json"])
except json.JSONDecodeError as exc:
sys.exit(f"bundle.json is not valid JSON: {exc}")
app = (UI / "k8s.yaml").read_text()
app = app.replace("__RIG_NAME__", namespace)
data = "\n".join(
f" {key}: |\n{block(text, 4)}" for key, text in sorted(contents.items())
)
return f"""# GENERATED by ctrl/manifest.py — do not edit.
# Regenerate with: make manifest
#
# Self-contained: applies as-is to any cluster, local kind or external.
# kubectl apply -f this-file.yaml
#
# Namespace carries the identity, so several rigs coexist in one cluster.
apiVersion: v1
kind: Namespace
metadata:
name: {namespace}
labels:
rig.bundle/name: {namespace}
---
apiVersion: v1
kind: ConfigMap
metadata:
name: rig-ui
namespace: {namespace}
labels:
rig.bundle/checksum: "{checksum(list(contents.values()))}"
data:
{data}
---
{_namespaced(app, namespace)}
"""
def _namespaced(doc: str, namespace: str) -> str:
"""Add `namespace:` to each resource so the artifact applies without -n.
rig-ui/k8s.yaml omits it on purpose — applied by hand it should land in
whatever namespace you choose. Pinning it belongs to the generated artifact,
which has to be self-contained.
"""
return re.sub(
r"^(metadata:\n(?:[ \t]+.*\n)*?)([ \t]+)(name: rig-ui)$",
lambda m: f"{m.group(1)}{m.group(2)}{m.group(3)}\n{m.group(2)}namespace: {namespace}",
doc.strip(),
flags=re.MULTILINE,
)
if __name__ == "__main__":
target = sys.argv[1] if len(sys.argv) > 1 else slug(ROOT.name)
sys.stdout.write(build(target))

View File

@@ -1,406 +0,0 @@
# GENERATED by ctrl/manifest.py — do not edit.
# Regenerate with: make manifest
#
# Self-contained: applies as-is to any cluster, local kind or external.
# kubectl apply -f this-file.yaml
#
# Namespace carries the identity, so several rigs coexist in one cluster.
apiVersion: v1
kind: Namespace
metadata:
name: sample-rig
labels:
rig.bundle/name: sample-rig
---
apiVersion: v1
kind: ConfigMap
metadata:
name: rig-ui
namespace: sample-rig
labels:
rig.bundle/checksum: "2074194964"
data:
bundle.json: |
{
"_comment": "What this bundle contains. Single source of truth — the landing page renders THIS file, so adding an entry here is the only edit needed. Deliberately FLAT: standalone tools and rigs, with none of soleprint's internal hierarchy (no artery/atlas/station layering). Nothing here is sensitive; the real architecture connects separately.",
"bundle": {
"name": "sample-rig",
"description": "Non-sensitive sample bundle. Proves the kind install works and shows what ships.",
"sensitive": false
},
"tools": [
{
"name": "modelgen",
"summary": "Generate models from config",
"standalone": true
},
{
"name": "datagen",
"summary": "Generate test data from rig-owned generators",
"standalone": true
},
{
"name": "graphgen",
"summary": "Generate navigable model graphs",
"standalone": true
},
{
"name": "tester",
"summary": "HTTP contract test runner — one suite, any environment",
"standalone": true
},
{
"name": "databrowse",
"summary": "SQL data browser",
"standalone": true
},
{
"name": "sbwrapper",
"summary": "Sandbox wrapper",
"standalone": true
}
],
"rigs": [
{
"name": "sample-rig",
"summary": "This bundle — a minimal, copyable environment",
"active": true
}
],
"next": [
"Point MANIFESTS_DIR at the real manifests to connect the actual architecture.",
"Real k8s files are versioned separately and are not part of this bundle."
]
}
index.html: |
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>IT WORKS</title>
</head>
<body>
<main id="app"></main>
<script type="module" src="/src/main.js"></script>
</body>
</html>
main.js: |
import "./style.css";
/* The IT WORKS page: renders bundle.json as the list of what shipped.
*
* Plain vite, no framework — this is a complement to the rig, not part of it,
* and it should stay small enough that nobody has to adopt a stack to read it.
*
* bundle.json is fetched at runtime rather than imported, so the same built app
* serves whatever rig it was copied into. Editing the ConfigMap changes the page
* without rebuilding.
*
* Styling is a handful of rules on purpose. The real visual identity lives in
* the soleprint UI package; nothing here should grow into a theme.
*/
const esc = (s) =>
String(s).replace(/&/g, "&amp;").replace(/</g, "&lt;").replace(/>/g, "&gt;");
const tag = (text, on = false) =>
`<span class="tag${on ? " on" : ""}">${esc(text)}</span>`;
function items(list, activeKey) {
if (!list?.length) return `<li><span class="summary">nothing listed</span></li>`;
return list
.map((it) => {
const tags = [
it.standalone ? tag("standalone") : "",
it.state ? tag(it.state) : "",
activeKey && it[activeKey] ? tag("active", true) : "",
].join("");
return `<li><span class="name">${esc(it.name ?? "?")}</span>
<span class="summary">${esc(it.summary ?? "")}</span>${tags}</li>`;
})
.join("");
}
/* Cluster state, when there is any to show.
*
* Fetched separately and allowed to fail: the bundle listing is the point, and a
* rig with no cluster reachable is a normal state, not an error. Renders nothing
* at all when absent.
*
* When the payload says `mocked`, say so loudly. This exists to demo the UI on a
* machine where kind will not run — and a demo that looks live but is not is
* worse than one that admits it. */
function clusterSection(c) {
if (!c) return "";
const m = c.cluster ?? {};
const banner = c.mocked
? `<p class="mock">mocked — no cluster was queried; these are canned values</p>`
: "";
const meta = [m.context, m.k8s, m.profile ? `profile ${m.profile}` : "",
m.nodes ? `${m.nodes} node${m.nodes > 1 ? "s" : ""}` : ""]
.filter(Boolean).join(" · ");
return `
<h2>Cluster${c.mocked ? " (mocked)" : ""}</h2>
${banner}
${meta ? `<p class="sub">${esc(meta)}</p>` : ""}
<ul>${items(c.workloads)}</ul>
<h2>Services (${c.services?.length ?? 0})</h2>
<ul>${items(c.services)}</ul>`;
}
function render(b, name, cluster) {
const meta = b.bundle ?? {};
const next = (b.next ?? []).map((n) => `<li>${esc(n)}</li>`).join("");
return `
<h1><span class="ok">IT WORKS</span> — ${esc(name || meta.name || "rig")}</h1>
<p class="sub">${esc(meta.description ?? "")}</p>
<h2>Tools (${b.tools?.length ?? 0})</h2>
<ul>${items(b.tools)}</ul>
<h2>Rigs (${b.rigs?.length ?? 0})</h2>
<ul>${items(b.rigs, "active")}</ul>
${clusterSection(cluster)}
${next ? `<div class="next"><ul>${next}</ul></div>` : ""}`;
}
const app = document.getElementById("app");
const json = (path, required) =>
fetch(path).then((r) => {
if (r.ok) return r.json();
if (required) throw new Error(`${path} -> HTTP ${r.status}`);
return null; // optional: absent is a normal state, not an error
}).catch((err) => {
if (required) throw err;
return null;
});
Promise.all([json("/bundle.json", true), json("/cluster.mock.json", false)])
// RIG_NAME is injected by vite from the pod env, so two rigs sharing a
// cluster are distinguishable even if a copied bundle.json kept its old name.
.then(([b, cluster]) => {
app.innerHTML = render(b, import.meta.env.VITE_RIG_NAME, cluster);
})
.catch((err) => {
app.innerHTML = `<h1 class="err">bundle unavailable</h1>
<p class="sub">${esc(err.message)}</p>`;
});
package.json: |
{
"name": "rig-ui",
"private": true,
"version": "0.1.0",
"type": "module",
"scripts": {
"dev": "vite --host 0.0.0.0 --port 5173",
"build": "vite build",
"preview": "vite preview --host 0.0.0.0 --port 5173"
},
"devDependencies": {
"vite": "^6"
}
}
style.css: |
/* Minimal, self-contained. The real visual identity ships with the soleprint UI
package, which is a separate artifact — nothing here should grow into a theme. */
body {
margin: 0;
padding: 2.5rem 1.5rem;
background: #0d0d0f;
color: #e8e8f0;
font: 14px/1.6 ui-monospace, "JetBrains Mono", Menlo, monospace;
}
main { max-width: 52rem; margin: 0 auto; }
h1 { margin: 0; font-size: 1.6rem; letter-spacing: 0.02em; }
h1 .ok { color: #3ecf8e; }
h1.err { color: #f06565; }
.sub { color: #8888a0; margin: 0.35rem 0 2.25rem; }
h2 {
font-size: 0.75rem;
text-transform: uppercase;
letter-spacing: 0.12em;
color: #8888a0;
margin: 2rem 0 0.75rem;
font-weight: 600;
}
ul { list-style: none; margin: 0; padding: 0; }
li {
display: flex;
gap: 0.75rem;
align-items: baseline;
padding: 0.5rem 0.75rem;
border: 1px solid #2e2e38;
border-radius: 6px;
margin-bottom: 0.4rem;
background: #16161a;
}
.name { font-weight: 600; min-width: 9rem; }
.summary { color: #8888a0; flex: 1; }
.tag {
font-size: 0.7rem;
padding: 0.1rem 0.45rem;
border-radius: 3px;
background: #26262f;
color: #8888a0;
white-space: nowrap;
}
.tag.on { background: #3ecf8e; color: #0d0d0f; }
.next {
color: #555568;
font-size: 0.8rem;
margin-top: 2.5rem;
border-top: 1px solid #2e2e38;
padding-top: 1rem;
}
.next li {
display: list-item;
border: 0;
background: none;
padding: 0.15rem 0;
margin: 0 0 0 1.1rem;
list-style: disc;
}
/* Mocked-data banner. Deliberately loud: this only appears when the cluster
payload is canned, and a demo that looks live but is not is worse than one
that says so. */
.mock {
margin: 0 0 0.75rem;
padding: 0.4rem 0.75rem;
border: 1px dashed #f5a623;
border-radius: 6px;
color: #f5a623;
font-size: 0.8rem;
}
vite.config.js: |
import { defineConfig } from "vite";
/* Serves on 0.0.0.0 so the pod is reachable through the Service, and allows any
* Host header because the address is assigned at runtime (MetalLB locally, a
* cloud load balancer on EKS) and is never known at build time. */
export default defineConfig({
server: { host: "0.0.0.0", port: 5173, strictPort: true, allowedHosts: true },
preview: { host: "0.0.0.0", port: 5173, strictPort: true, allowedHosts: true },
});
---
# How to plug the UI into whatever k8s you generated. THIS IS THE WHOLE THING:
# one Pod running the vite app, one Service to reach it.
#
# Optional by design. The UI complements a rig; it is not part of the end
# product, and a rig is complete and useful without it. Apply this only when you
# want the listing:
#
# kubectl apply -n <your-namespace> -f rig-ui/k8s.yaml
#
# A bare Pod, not a Deployment — this is a dev-loop convenience, not a workload
# to keep alive. If it dies you re-apply it; nothing depends on it staying up.
#
# The app and bundle.json arrive as a ConfigMap named `rig-ui`, which
# ctrl/manifest.py generates from the folder. Nothing is baked into an image, so
# editing bundle.json and re-applying is the whole update cycle.
apiVersion: v1
kind: Pod
metadata:
name: rig-ui
namespace: sample-rig
labels:
app: rig-ui
spec:
containers:
- name: vite
image: node:22-alpine
workingDir: /app
# npm install at start: no image to build and no registry to publish to,
# which is the point of a minimal plug-in. It needs egress to a registry —
# on a locked-down cluster point npm at the internal one, or bake an image
# instead. Nothing else here changes if you do.
command: ["sh", "-c"]
# A ConfigMap mounts flat (keys cannot contain '/'), so the files are
# placed into vite's expected layout here. bundle.json goes to public/
# because that is what vite serves at /bundle.json, which is where the
# app fetches it.
args:
- |
mkdir -p /app/src /app/public &&
cp /src/package.json /src/vite.config.js /src/index.html /app/ &&
cp /src/main.js /src/style.css /app/src/ &&
cp /src/bundle.json /app/public/ &&
npm install --no-audit --no-fund &&
npm run dev
env:
# Rendered in the heading so two rigs sharing a cluster stay
# distinguishable. Set from the namespace by ctrl/manifest.py.
- name: VITE_RIG_NAME
value: sample-rig
ports:
- name: http
containerPort: 5173
volumeMounts:
# /src is read-only from the ConfigMap; the app is copied to a writable
# /app because npm install has to create node_modules.
- name: rig-ui
mountPath: /src
- name: app
mountPath: /app
readinessProbe:
httpGet: { path: /, port: 5173 }
# npm install decides how long this takes, and it is the slow part.
initialDelaySeconds: 15
periodSeconds: 5
failureThreshold: 30
resources:
requests: { memory: 128Mi, cpu: 50m }
limits: { memory: 512Mi }
volumes:
- name: rig-ui
configMap:
name: rig-ui
- name: app
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: rig-ui
namespace: sample-rig
labels:
app: rig-ui
# No annotations, deliberately — see k8s/app.yaml. The target is EKS but this
# stays VPC-agnostic: no subnets, no security groups, no -scheme, no -type.
# A bare LoadBalancer is what lets one manifest work on kind and on EKS.
spec:
type: LoadBalancer
selector:
app: rig-ui
ports:
- name: http
port: 80
targetPort: 5173
protocol: TCP
# Pinned, because a LoadBalancer Service also allocates a NodePort and
# this is the only address that works everywhere.
#
# On WSL the MetalLB address is on a docker bridge INSIDE the Linux VM,
# and Windows has no route to it — the page looks broken while the
# cluster is perfectly healthy. 30080 is what rig's `hostport` ingress
# mode publishes to the host, so this is reachable at
# localhost:$HTTP_PORT from a Windows browser with nothing configured.
#
# Costs nothing elsewhere: MetalLB still assigns an external IP on Linux,
# and on EKS the load balancer targets this NodePort anyway. One Service,
# no per-environment branch.
#
# A pinned NodePort is cluster-unique, so two rigs must live in separate
# clusters — which is how they are run anyway.
nodePort: 30080

View File

@@ -1,3 +0,0 @@
node_modules/
public/
dist/

View File

@@ -1,12 +0,0 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>IT WORKS</title>
</head>
<body>
<main id="app"></main>
<script type="module" src="/src/main.js"></script>
</body>
</html>

View File

@@ -1,108 +0,0 @@
# How to plug the UI into whatever k8s you generated. THIS IS THE WHOLE THING:
# one Pod running the vite app, one Service to reach it.
#
# Optional by design. The UI complements a rig; it is not part of the end
# product, and a rig is complete and useful without it. Apply this only when you
# want the listing:
#
# kubectl apply -n <your-namespace> -f rig-ui/k8s.yaml
#
# A bare Pod, not a Deployment — this is a dev-loop convenience, not a workload
# to keep alive. If it dies you re-apply it; nothing depends on it staying up.
#
# The app and bundle.json arrive as a ConfigMap named `rig-ui`, which
# ctrl/manifest.py generates from the folder. Nothing is baked into an image, so
# editing bundle.json and re-applying is the whole update cycle.
apiVersion: v1
kind: Pod
metadata:
name: rig-ui
labels:
app: rig-ui
spec:
containers:
- name: vite
image: node:22-alpine
workingDir: /app
# npm install at start: no image to build and no registry to publish to,
# which is the point of a minimal plug-in. It needs egress to a registry —
# on a locked-down cluster point npm at the internal one, or bake an image
# instead. Nothing else here changes if you do.
command: ["sh", "-c"]
# A ConfigMap mounts flat (keys cannot contain '/'), so the files are
# placed into vite's expected layout here. bundle.json goes to public/
# because that is what vite serves at /bundle.json, which is where the
# app fetches it.
args:
- |
mkdir -p /app/src /app/public &&
cp /src/package.json /src/vite.config.js /src/index.html /app/ &&
cp /src/main.js /src/style.css /app/src/ &&
cp /src/bundle.json /app/public/ &&
npm install --no-audit --no-fund &&
npm run dev
env:
# Rendered in the heading so two rigs sharing a cluster stay
# distinguishable. Set from the namespace by ctrl/manifest.py.
- name: VITE_RIG_NAME
value: __RIG_NAME__
ports:
- name: http
containerPort: 5173
volumeMounts:
# /src is read-only from the ConfigMap; the app is copied to a writable
# /app because npm install has to create node_modules.
- name: rig-ui
mountPath: /src
- name: app
mountPath: /app
readinessProbe:
httpGet: { path: /, port: 5173 }
# npm install decides how long this takes, and it is the slow part.
initialDelaySeconds: 15
periodSeconds: 5
failureThreshold: 30
resources:
requests: { memory: 128Mi, cpu: 50m }
limits: { memory: 512Mi }
volumes:
- name: rig-ui
configMap:
name: rig-ui
- name: app
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: rig-ui
labels:
app: rig-ui
# No annotations, deliberately — see k8s/app.yaml. The target is EKS but this
# stays VPC-agnostic: no subnets, no security groups, no -scheme, no -type.
# A bare LoadBalancer is what lets one manifest work on kind and on EKS.
spec:
type: LoadBalancer
selector:
app: rig-ui
ports:
- name: http
port: 80
targetPort: 5173
protocol: TCP
# Pinned, because a LoadBalancer Service also allocates a NodePort and
# this is the only address that works everywhere.
#
# On WSL the MetalLB address is on a docker bridge INSIDE the Linux VM,
# and Windows has no route to it — the page looks broken while the
# cluster is perfectly healthy. 30080 is what rig's `hostport` ingress
# mode publishes to the host, so this is reachable at
# localhost:$HTTP_PORT from a Windows browser with nothing configured.
#
# Costs nothing elsewhere: MetalLB still assigns an external IP on Linux,
# and on EKS the load balancer targets this NodePort anyway. One Service,
# no per-environment branch.
#
# A pinned NodePort is cluster-unique, so two rigs must live in separate
# clusters — which is how they are run anyway.
nodePort: 30080

File diff suppressed because it is too large Load Diff

View File

@@ -1,14 +0,0 @@
{
"name": "rig-ui",
"private": true,
"version": "0.1.0",
"type": "module",
"scripts": {
"dev": "vite --host 0.0.0.0 --port 5173",
"build": "vite build",
"preview": "vite preview --host 0.0.0.0 --port 5173"
},
"devDependencies": {
"vite": "^6"
}
}

View File

@@ -1,136 +0,0 @@
import "./style.css";
/* The IT WORKS page: renders bundle.json as the list of what shipped.
*
* Plain vite, no framework — this is a complement to the rig, not part of it,
* and it should stay small enough that nobody has to adopt a stack to read it.
*
* Laid out like soleprint's templated vein pages, because it does the same job:
* name each component, list what it exposes, show what comes back. Tool chrome
* and output are styled apart on purpose (see style.css) — that separation is
* what tells you whether you are reading the tool or its result.
*
* bundle.json is fetched at runtime rather than imported, so the same built app
* serves whatever rig it was copied into. Editing the ConfigMap changes the page
* without rebuilding.
*/
const esc = (s) =>
String(s).replace(/&/g, "&amp;").replace(/</g, "&lt;").replace(/>/g, "&gt;");
const tag = (text, on = false) =>
`<span class="tag${on ? " on" : ""}">${esc(text)}</span>`;
/* Tool chrome: one bordered card per component. */
function components(list, activeKey) {
if (!list?.length)
return `<div class="component"><p>nothing listed</p></div>`;
return list
.map((it) => {
const tags = [
it.standalone ? tag("standalone") : "",
it.state ? tag(it.state) : "",
activeKey && it[activeKey] ? tag("active", true) : "",
].join("");
return `<div class="component">
<h4>${esc(it.name ?? "?")} ${tags}</h4>
<p>${esc(it.summary ?? "")}</p>
</div>`;
})
.join("");
}
/* Endpoint rows: path on the left, what it returns on the right. */
function endpoints(list) {
return list
.map(
(e) => `<li><code>${esc(e.path)}</code>
<span class="desc">${esc(e.desc)}</span></li>`
)
.join("");
}
/* Output: what the endpoint above actually returns, so the page demonstrates
itself rather than describing what a demonstration would look like. */
function example(bundle) {
const sample = {
bundle: bundle.bundle?.name,
tools: (bundle.tools ?? []).map((t) => t.name),
rigs: (bundle.rigs ?? []).map((r) => r.name),
};
return `<pre class="output">${esc(JSON.stringify(sample, null, 2))}</pre>`;
}
/* Cluster state, when there is any to show.
*
* Fetched separately and allowed to fail: the bundle listing is the point, and a
* rig with no cluster reachable is a normal state, not an error. Renders nothing
* at all when absent. When the payload says `mocked`, say so loudly. */
function clusterSection(c) {
if (!c) return "";
const m = c.cluster ?? {};
const meta = [m.context, m.k8s, m.profile && `profile ${m.profile}`,
m.nodes && `${m.nodes} node${m.nodes > 1 ? "s" : ""}`]
.filter(Boolean).join(" · ");
return `
<h2>Cluster${c.mocked ? " (mocked)" : ""}</h2>
${c.mocked ? `<p class="mock">mocked — no cluster was queried; these are canned values</p>` : ""}
${meta ? `<p class="tagline">${esc(meta)}</p>` : ""}
<div class="components">${components(c.workloads)}</div>
<h2>Services</h2>
<div class="components">${components(c.services)}</div>`;
}
function render(b, name, cluster) {
const meta = b.bundle ?? {};
const next = (b.next ?? []).map((n) => `<li>${esc(n)}</li>`).join("");
return `
<h1><span class="ok">IT WORKS</span> — ${esc(name || meta.name || "rig")}</h1>
<p class="tagline">${esc(meta.description ?? "")}</p>
<h2>Tools (${b.tools?.length ?? 0})</h2>
<div class="components">${components(b.tools)}</div>
<h2>Rigs (${b.rigs?.length ?? 0})</h2>
<div class="components">${components(b.rigs, "active")}</div>
<h2>Endpoints</h2>
<ul class="endpoints">${endpoints([
{ path: "/", desc: "this page" },
{ path: "/bundle.json", desc: "the manifest it renders" },
])}</ul>
<h2>Example — GET /bundle.json</h2>
${example(b)}
${clusterSection(cluster)}
${next ? `<div class="next"><ul>${next}</ul></div>` : ""}`;
}
const app = document.getElementById("app");
const json = (path, required) =>
fetch(path)
.then((r) => {
if (r.ok) return r.json();
if (required) throw new Error(`${path} -> HTTP ${r.status}`);
return null; // optional: absent is a normal state, not an error
})
.catch((err) => {
if (required) throw err;
return null;
});
Promise.all([json("/bundle.json", true), json("/cluster.mock.json", false)])
// RIG_NAME is injected by vite from the pod env, so two rigs sharing a
// cluster are distinguishable even if a copied bundle.json kept its old name.
.then(([b, cluster]) => {
app.innerHTML = render(b, import.meta.env.VITE_RIG_NAME, cluster);
})
.catch((err) => {
app.innerHTML = `<h1 class="err">bundle unavailable</h1>
<p class="tagline">${esc(err.message)}</p>`;
});

View File

@@ -1,139 +0,0 @@
/* Minimal and self-contained — no framework dependency.
*
* The visual language follows soleprint's templated vein pages, because this
* page does the same job: say what a component is, list what it exposes, and
* show what comes back. Two treatments, deliberately distinct:
*
* TOOL CHROME bordered cards on the darker background, accent-coloured
* titles, endpoint rows separated by rules.
* OUTPUT a lighter raised block, monospace, pre-wrap and selectable —
* it is data, not furniture, and should read as a payload.
*
* Keeping them apart matters more than either looks: it is what tells you at a
* glance whether you are reading the tool or the thing it produced. */
:root {
--bg: #0d0d0f;
--surface: #16161a;
--surface-raised: #1e1e24;
--border: #2e2e38;
--border-strong: #3d3d4a;
--text: #e8e8f0;
--muted: #8888a0;
--accent: #3ecf8e;
--accent-dim: #f5a623;
--mono: "JetBrains Mono", "Cascadia Mono", Consolas, ui-monospace, monospace;
}
body {
margin: 0;
padding: 2.5rem 1.5rem;
background: var(--bg);
color: var(--text);
font: 14px/1.6 var(--mono);
}
main { max-width: 56rem; margin: 0 auto; }
h1 { margin: 0; font-size: 1.6rem; letter-spacing: 0.02em; }
h1 .ok { color: var(--accent); }
h1.err { color: #f06565; }
.tagline { color: var(--muted); margin: 0.35rem 0 2.25rem; }
h2 {
font-size: 0.75rem;
text-transform: uppercase;
letter-spacing: 0.12em;
color: var(--muted);
margin: 2.25rem 0 0.75rem;
font-weight: 600;
}
/* ── tool chrome ────────────────────────────────────────────────────────── */
.components {
display: grid;
grid-template-columns: repeat(auto-fit, minmax(200px, 1fr));
gap: 0.75rem;
}
.component {
background: var(--bg);
border: 1px solid var(--border-strong);
border-radius: 8px;
padding: 0.75rem;
}
.component h4 {
margin: 0 0 0.25rem;
font-size: 0.95rem;
color: var(--accent);
}
.component p {
margin: 0;
font-size: 0.85rem;
color: var(--muted);
}
.endpoints { list-style: none; margin: 0; padding: 0; }
.endpoints li {
display: flex;
justify-content: space-between;
flex-wrap: wrap;
gap: 0.5rem;
padding: 0.6rem 0;
border-bottom: 1px solid var(--border-strong);
}
.endpoints li:last-child { border-bottom: none; }
.endpoints code {
background: var(--surface);
color: var(--accent);
padding: 0.25rem 0.5rem;
border-radius: 4px;
}
.endpoints .desc { color: var(--muted); font-size: 0.9rem; }
.tag {
font-size: 0.7rem;
padding: 0.1rem 0.45rem;
border-radius: 3px;
background: var(--border);
color: var(--muted);
white-space: nowrap;
}
.tag.on { background: var(--accent); color: var(--bg); }
/* ── output ─────────────────────────────────────────────────────────────── */
.output {
background: var(--surface);
border: 1px solid var(--border);
border-radius: 8px;
padding: 1rem;
font-size: 0.85rem;
white-space: pre-wrap;
word-break: break-word;
user-select: text;
color: var(--text);
margin: 0;
}
.output .k { color: var(--accent); }
/* Mocked-data banner. Loud on purpose: it only appears when the payload is
canned, and a demo that looks live but is not is worse than one that says so. */
.mock {
margin: 0 0 0.75rem;
padding: 0.4rem 0.75rem;
border: 1px dashed var(--accent-dim);
border-radius: 6px;
color: var(--accent-dim);
font-size: 0.8rem;
}
.next {
color: #555568;
font-size: 0.8rem;
margin-top: 2.5rem;
border-top: 1px solid var(--border);
padding-top: 1rem;
}
.next ul { margin: 0; padding: 0 0 0 1.1rem; }
.next li { list-style: disc; padding: 0.15rem 0; }

View File

@@ -1,9 +0,0 @@
import { defineConfig } from "vite";
/* Serves on 0.0.0.0 so the pod is reachable through the Service, and allows any
* Host header because the address is assigned at runtime (MetalLB locally, a
* cloud load balancer on EKS) and is never known at build time. */
export default defineConfig({
server: { host: "0.0.0.0", port: 5173, strictPort: true, allowedHosts: true },
preview: { host: "0.0.0.0", port: 5173, strictPort: true, allowedHosts: true },
});

View File

@@ -0,0 +1,23 @@
# GENERATED by make standalone — do not edit
#
# Shorthand for the scripts beside it; they run without it. Every target
# calls a verb its script accepts — read from that script's own dispatch.
HERE := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
ARGS := $(wordlist 2,$(words $(MAKECMDGOALS)),$(MAKECMDGOALS))
ifneq ($(ARGS),)
$(eval $(ARGS):;@:)
.PHONY: $(ARGS)
endif
.DEFAULT_GOAL := help
.PHONY: help deps mem
help: ## list targets
@grep -hE '^[a-z][a-z-]*:.*?##' $(MAKEFILE_LIST) | sed 's/:.*##/\t/' | expand -t16
deps: ## rigdeps.sh [detect|list|verify|fetch|install] (default detect)
bash $(HERE)rigdeps.sh $(or $(ARGS),detect)
mem: ## rigmini.sh [status|push|all|backup|restore] (default status)
bash $(HERE)rigmini.sh $(or $(ARGS),status)

959
rig/standalone/default/rigdeps.sh Executable file
View File

@@ -0,0 +1,959 @@
#!/usr/bin/env bash
# GENERATED by make standalone — do not edit
#
# rigdeps.sh for profile 'default', flattened from:
# ctrl/deps.sh
# ctrl/lib/config.sh
# Edit those and run `make standalone`. Changes made here are lost, and
# `make selftest` fails while this file differs from what rig generates.
# ── from the libraries ──
declare -- CONFIG_OVERRIDABLE=$'PROFILE CLUSTER K8S_VERSION KIND_CONFIG ADDONS\n REGISTRY_MODE INGRESS_MODE DNS_MODE TILT_PORT\n SOURCE ARCH DEPS_SOURCE HTTP_PORT HTTPS_PORT\n REGISTRY_PORT MANIFESTS_DIR'
_config_restore ()
{
local line;
while IFS= read -r line; do
if [ -n "$line" ]; then
eval "export $line";
fi;
done <<< "$1";
return 0
}
default_cluster_name ()
{
local n;
n=$(basename "$(cd .. && pwd)");
n=$(echo "$n" | tr '[:upper:]' '[:lower:]' | tr -c 'a-z0-9-' '-');
n=$(echo "$n" | sed 's/^-*//; s/-*$//');
echo "${n:-rig}"
}
derive_port_base ()
{
local h;
h=$(printf '%s' "$1" | cksum | awk '{print $1}');
echo $((20000 + (h % 200) * 10))
}
render_kind_config ()
{
local host_workdir="${HOST_WORKDIR:-$(cd .. && pwd)}";
sed -e "s|\${CLUSTER}|${CLUSTER}|g" -e "s|\${NODE_IMAGE}|${NODE_IMAGE}|g" -e "s|\${HTTP_PORT}|${HTTP_PORT}|g" -e "s|\${HOST_WORKDIR}|${host_workdir}|g" "$KIND_CONFIG_PATH"
}
# ── configuration, frozen for profile 'default' ──
load_config() {
local k saved=""
for k in $CONFIG_OVERRIDABLE; do
if [ -n "${!k+x}" ]; then saved+="$k=$(printf '%q' "${!k}")"$'\n'; fi
done
declare -g ADDONS=""
declare -gx AIRFLOW_IMAGE="apache/airflow:2.10.4"
declare -g AUDIT="off"
declare -gx CERT_MANAGER_VERSION="v1.21.1"
declare -g CLUSTER="rig"
declare -gx COMPOSE_SHA256="db1889184726840f75c4f9c001048430d4f25b3be3cb084d3ddd762bc0aed576"
declare -gx COMPOSE_URL="https://github.com/docker/compose/releases/download/v5.5.1/docker-compose-linux-x86_64"
declare -gx COMPOSE_VERSION="5.5.1"
declare -gx CTLPTL_SHA256="c63a1ec28e60bc3faf6becb76f53355c5cf5e0143dafdd27ad85db5584fa6b1e"
declare -gx CTLPTL_URL="https://github.com/tilt-dev/ctlptl/releases/download/v0.9.4/ctlptl.0.9.4.linux.x86_64.tar.gz"
declare -gx CTLPTL_VERSION="0.9.4"
declare -g DNS_MODE="hosts"
declare -g HTTPS_PORT="20311"
declare -g HTTP_PORT="20310"
declare -g INGRESS_MODE="hostport"
declare -gx JQ_SHA256="b1c22172dd303f3be49e935aa56aa48a8b7a46e0bc838b4997d3bb451495870f"
declare -gx JQ_URL="https://github.com/jqlang/jq/releases/download/jq-1.8.2/jq-linux-amd64"
declare -gx JQ_VERSION="1.8.2"
declare -g K8S_VERSION="v1_36"
declare -g KIND_CONFIG="kind-config.yaml.tpl"
declare -g KIND_CONFIG_PATH="./k8s/kind-config.yaml.tpl"
declare -g KIND_CONFIG_SHOWN="ctrl/k8s/kind-config.yaml.tpl"
declare -gx KIND_SHA256="50030de23cf40a18505f20426f6a8506bedf13c6e509244bd1fa9463721b0f54"
declare -gx KIND_URL="https://github.com/kubernetes-sigs/kind/releases/download/v0.32.0/kind-linux-amd64"
declare -gx KIND_VERSION="v0.32.0"
declare -g KUBECONTEXT="kind-rig"
declare -gx KUBECTL_SHA256="ebbd080e7c2e275093b55915722043257eb24004363e20acb3c4d71919f88336"
declare -gx KUBECTL_URL="https://dl.k8s.io/release/v1.36.3/bin/linux/amd64/kubectl"
declare -gx KUBECTL_VERSION="v1.36.3"
declare -g MANIFESTS_DIR="ctrl/k8s/overlays/dev"
declare -gx METALLB_VERSION="v0.16.0"
declare -gx METRICS_SERVER_VERSION="v0.9.0"
declare -g NODES="1"
declare -g NODE_IMAGE="kindest/node:v1.36.1@sha256:3489c7674813ba5d8b1a9977baea8a6e553784dab7b84759d1014dbd78f7ebd5"
declare -gx NODE_IMAGE_v1_33="kindest/node:v1.33.12@sha256:3f5c8443c620245e4d355cfe09e96a91ead32ceaa569d3f1ca9edf0cb2fe2ff4"
declare -gx NODE_IMAGE_v1_34="kindest/node:v1.34.8@sha256:02722c2dedddcfc00febf5d27fbeb9b7b2c14294c82109ff4a85d89ac9ba3256"
declare -gx NODE_IMAGE_v1_35="kindest/node:v1.35.5@sha256:ce977ae6d65918d0b58a5f8b5e940429c2ce42fa3a5619ec2bbc60b949c0ac95"
declare -gx NODE_IMAGE_v1_36="kindest/node:v1.36.1@sha256:3489c7674813ba5d8b1a9977baea8a6e553784dab7b84759d1014dbd78f7ebd5"
declare -g NODE_MB="800"
declare -gx POSTGRES_IMAGE="postgres:16-alpine"
declare -gx PROFILE="default"
declare -g PROFILE_NAME="default"
declare -gx REDIS_IMAGE="redis:7-alpine"
declare -gx REGISTRY_IMAGE="registry:2"
declare -g REGISTRY_MODE="local"
declare -g REGISTRY_PORT="20313"
declare -gx STUB_IMAGE="python:3.12-slim"
declare -g TILT_PORT="20312"
declare -gx TILT_SHA256="e9672b8a18d43501f35dcfe98465969a7db0e436b36cf0c50c7e6f8d40de5fe6"
declare -gx TILT_URL="https://github.com/tilt-dev/tilt/releases/download/v0.37.6/tilt.0.37.6.linux.x86_64.tar.gz"
declare -gx TILT_VERSION="0.37.6"
_config_restore "$saved"
}
# ── end of frozen configuration ──
# ── ctrl/deps.sh ──
# Toolchain installer: detect the host, install a pinned toolchain onto it, then
# report what it could not do.
#
# It never runs the cluster, never uses sudo or apt, and writes only into
# $OUT_BIN (default ~/.local/bin). Everything that would touch the host proper —
# systemd, inotify limits, .wslconfig, docker group — is REPORTED for a human to
# decide on, never performed. That is what makes it safe to run on a machine that
# already has a working setup.
#
# Usage (normally via `make deps`, or directly):
# deps.sh detect # report host facts only, change nothing
# deps.sh list # the pinned versions
# deps.sh verify [core|dev] # run what is installed and see if it works
# deps.sh fetch [core|dev] [--to DIR] # download + verify into DIR
# deps.sh install [core|dev] # detect, fetch, install, report
#
# Tiers: 'core' is kubectl + jq (talk to a cluster); 'dev' adds kind and tilt
# Default is dev.
#
# Runs both inside the installer container and bare on a host. Inside the
# container, host files are read through $HOST_ROOT (mount / as :ro); bare, it
# falls back to /.
set -euo pipefail
# Keep the caller's cwd so a relative --to resolves where the user expects,
# not against ctrl/ once we've moved.
INVOKED_FROM="$PWD"
cd "$(dirname "$0")"
# Pins arrive through load_config like every other setting, not by sourcing
# versions.env here. That is what lets `make standalone` freeze them into a
# one-file installer: configuration has exactly one way in.
# (sourced library inlined above)
load_config
# Resolve a possibly-relative path against the caller's original directory.
abspath() {
case "$1" in
/*) echo "$1" ;;
*) echo "$INVOKED_FROM/$1" ;;
esac
}
OUT_BIN="${OUT_BIN:-$HOME/.local/bin}"
HOST_ROOT="${HOST_ROOT:-/}"
DEPS_SOURCE="${DEPS_SOURCE:-upstream}"
DEPS_ARTIFACTORY_URL="${DEPS_ARTIFACTORY_URL:-}"
BAKED_BIN="${BAKED_BIN:-/opt/rig/bin}"
# Collected by detect(), printed by report_manual() at the very end.
MANUAL=()
# Host FILES (/etc/..., /mnt/c/...) must be read through the mount. Kernel-level
# facts (kernel version, meminfo, inotify) are shared with the container, so the
# container's own view is already the host's.
# A /proc/meminfo field in MB, 0 if the field is absent. MEMINFO exists so the
# tight and does-not-fit branches can be exercised against a real machine's
# numbers from somewhere else; in normal use it is always /proc/meminfo.
mb_of() {
awk -v k="$1:" '$1 == k { printf "%d", $2 / 1024; found = 1 }
END { if (!found) printf "0" }' "${MEMINFO:-/proc/meminfo}"
}
host_file() {
local p="${1#/}"
if [ "$HOST_ROOT" != "/" ] && [ -e "$HOST_ROOT/$p" ]; then
echo "$HOST_ROOT/$p"
else
echo "/$p"
fi
}
# ── the tools this script itself needs ─────────────────────────────────────
arch() {
case "$(uname -m)" in
x86_64|amd64) echo amd64 ;;
aarch64|arm64) echo arm64 ;;
*) uname -m ;;
esac
}
# The pins above are amd64. Rather than download something that cannot execute
# and let it fail as "cannot execute binary file: Exec format error", say so
# here and hand over the commands that produce the right checksums.
require_amd64() {
local a; a=$(arch)
[ "$a" = "amd64" ] && return 0
cat >&2 <<EOF
This machine is ${a} ($(uname -m)); every pin in this script is linux/amd64.
Nothing here would run, so it does not download. To make an ${a} version, the
URLs need the ${a} artifact and the checksums need to come from each project's
own published list — not from these values, and not from a download you did:
curl -sSL https://github.com/kubernetes-sigs/kind/releases/download/${KIND_VERSION}/checksums.txt
curl -sSL https://dl.k8s.io/release/${KUBECTL_VERSION}/bin/linux/${a}/kubectl.sha256
curl -sSL https://github.com/tilt-dev/tilt/releases/download/v${TILT_VERSION}/checksums.txt
curl -sSL https://github.com/tilt-dev/ctlptl/releases/download/v${CTLPTL_VERSION}/checksums.txt
curl -sSL https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/sha256sum.txt
Edit the pinned block at the top of this file with what those print.
EOF
exit 1
}
DL=""
pick_downloader() {
if command -v curl >/dev/null 2>&1; then DL=curl
elif command -v wget >/dev/null 2>&1; then DL=wget
else
echo "neither curl nor wget is installed, so nothing can be downloaded." >&2
echo "Install one first: $(pkg_install_cmd curl)" >&2
exit 1
fi
}
download() {
local url="$1" out="$2"
case "$DL" in
curl) curl -fsSL --retry 3 -o "$out" "$url" ;;
wget) wget -q --tries=3 -O "$out" "$url" ;;
esac
}
SHA=""
pick_sha() {
if command -v sha256sum >/dev/null 2>&1; then SHA=sha256sum
elif command -v shasum >/dev/null 2>&1; then SHA="shasum -a 256"
else
echo "no sha256sum and no shasum — downloads could not be verified." >&2
echo "Refusing to install unverified binaries." >&2
exit 1
fi
}
# ── package manager, for the instructions only ─────────────────────────────
# This never runs a package manager. It names one so the reported action is
# something you can paste, on the distro you are actually on — an apt line on
# Amazon Linux 2 is a wrong answer dressed up as help.
pkg_install_cmd() {
local pkg="$1"
if command -v apt-get >/dev/null 2>&1; then echo "sudo apt-get update && sudo apt-get install -y $pkg"
elif command -v dnf >/dev/null 2>&1; then echo "sudo dnf install -y $pkg"
elif command -v yum >/dev/null 2>&1; then echo "sudo yum install -y $pkg"
elif command -v zypper >/dev/null 2>&1; then echo "sudo zypper install -y $pkg"
elif command -v apk >/dev/null 2>&1; then echo "sudo apk add $pkg"
else echo "install '$pkg' with this system's package manager"
fi
}
docker_pkg() {
# Debian and Ubuntu call it docker.io; the RPM distros call it docker.
if command -v apt-get >/dev/null 2>&1; then echo docker.io; else echo docker; fi
}
# ── detect ─────────────────────────────────────────────────────────────────
# Windows outside WSL — Git Bash, MSYS, Cygwin — looks close enough to work and
# then fails in a pile of confusing ways: no /proc, no docker socket, none of
# the tooling. Detectable, so name it instead.
require_linux() {
case "$(uname -s)" in
MINGW*|MSYS*|CYGWIN*)
cat >&2 <<'EOF'
This has to run inside WSL, not Git Bash / MSYS / Cygwin.
If WSL is not installed yet, from an elevated PowerShell or Command Prompt:
wsl --install
That enables Windows features and needs a reboot, so it is not something this
script will do for you. Afterwards, open the Linux shell it installs and run
this from there.
See "Starting from plain Windows" in README.md.
EOF
exit 1 ;;
esac
}
is_wsl() { grep -qi microsoft /proc/version 2>/dev/null; }
detect() {
echo "host"
echo " kernel $(uname -r)"
echo " arch $(arch) ($(uname -m))"
local osr; osr=$(host_file /etc/os-release)
[ -r "$osr" ] && echo " distro $(sed -n 's/^PRETTY_NAME="\(.*\)"/\1/p' "$osr")"
# In MB. Whole gigabytes lose nearly half a GB on exactly the machines where
# it matters: 1874 MB available used to print as "1 GB". Facts only — whether
# that is enough depends on the profile, which check.sh knows and this does not.
local total_mb avail_mb swap_total_mb swap_used_mb om
total_mb=$(mb_of MemTotal)
avail_mb=$(mb_of MemAvailable)
swap_total_mb=$(mb_of SwapTotal)
swap_used_mb=$(( swap_total_mb - $(mb_of SwapFree) ))
printf " memory %d MB total, %d MB available\n" "$total_mb" "$avail_mb"
if [ "$swap_total_mb" -gt 0 ]; then
printf " swap %d MB used of %d MB\n" "$swap_used_mb" "$swap_total_mb"
fi
# How the kernel answers an allocation it cannot really satisfy. With 1 it
# always says yes and settles up later with the OOM killer, so a cluster that
# starts cleanly can still lose processes afterwards.
om=$(cat "${OVERCOMMIT_FILE:-/proc/sys/vm/overcommit_memory}" 2>/dev/null || echo '?')
case "$om" in
0) echo " overcommit 0 heuristic — allocations are granted on a guess" ;;
1) echo " overcommit 1 always — every allocation succeeds; the OOM killer is the only limit" ;;
2) echo " overcommit 2 strict — an allocation fails honestly instead of killing later" ;;
esac
echo " install to $OUT_BIN"
detect_libc
detect_prereqs
detect_wsl
detect_filesystem
detect_docker
detect_inotify
detect_toolchain
}
detect_wsl() {
if ! is_wsl; then
echo " platform native linux"
return
fi
echo " platform WSL"
# systemd is off by default in WSL, and the ingress/DNS paths that use a
# host service need it. Enabling it requires a Windows-side restart, which
# cannot be issued from inside the distro.
local wc; wc=$(host_file /etc/wsl.conf)
if [ -r "$wc" ] && grep -qE '^\s*systemd\s*=\s*true' "$wc"; then
echo " systemd enabled in wsl.conf"
else
echo " ! systemd not enabled in /etc/wsl.conf"
MANUAL+=("Enable systemd — add to /etc/wsl.conf:
[boot]
systemd=true
then from a WINDOWS terminal (not this shell): wsl --shutdown")
fi
# WSL regenerates /etc/resolv.conf on every boot, which silently reverts any
# local DNS setup.
if [ -r "$wc" ] && grep -qE '^\s*generateResolvConf\s*=\s*false' "$wc"; then
echo " resolv.conf pinned (generateResolvConf=false)"
else
echo " - resolv.conf is WSL-generated; DNS_MODE=dnsmasq would be reverted on reboot"
fi
local wcfg
wcfg=$(ls "$HOST_ROOT"/mnt/c/Users/*/.wslconfig 2>/dev/null | head -1 || true)
if [ -n "$wcfg" ] && grep -qE '^\s*memory\s*=' "$wcfg"; then
echo " wslconfig memory set: $(grep -E '^\s*memory\s*=' "$wcfg" | tr -d ' ')"
else
MANUAL+=("Cap/raise the WSL VM memory — see what is set versus what booted:
make mem status
It prints the edit to make and the command to apply it.")
fi
}
# Not a path check: /mnt is an ordinary mount point and an ext4 disk mounted
# there is perfectly fine. What matters is the filesystem. The Windows drives
# arrive as 9p (WSL2) or drvfs (WSL1); network and fuse mounts behave the same
# way. None of them deliver inotify events, so anything watching files goes
# quiet without saying why.
watch_hostile_fs() {
local dir="$1" fstype
fstype=$(findmnt -no FSTYPE --target "$dir" 2>/dev/null || true)
[ -n "$fstype" ] || fstype=$(stat -f -c %T "$dir" 2>/dev/null || true)
case "$fstype" in
9p|v9fs|drvfs|cifs|smb3|nfs|nfs4|fuse.sshfs|fuseblk) echo "$fstype" ;;
*) echo "" ;;
esac
}
detect_filesystem() {
local root fstype
root=$(cd .. && pwd -P)
fstype=$(watch_hostile_fs "$root")
if [ -n "$fstype" ]; then
echo " ! this directory is on $fstype — file watching will not work"
MANUAL+=("Move this onto the local disk. Nothing watching files sees changes
on a $fstype mount, and everything else is slower:
cp -r \"$root\" ~/ && cd ~/$(basename "$root")")
else
echo " filesystem $root ($(findmnt -no FSTYPE --target "$root" 2>/dev/null || echo local))"
fi
}
# tilt is the one binary here that needs a recent glibc. MEASURED, not guessed:
# tilt 0.37.6 on Amazon Linux 2 (glibc 2.26) fails with
#
# /lib64/libc.so.6: version `GLIBC_2.34' not found (required by .../tilt)
#
# which names a symbol rather than the problem. Amazon Linux 2 is a stock
# WorkSpaces bundle, so this is the likely case, not an exotic one. Report the
# version now; `verify` catches the actual failure after installing.
detect_libc() {
local v=""
if command -v ldd >/dev/null 2>&1; then
v=$(ldd --version 2>/dev/null | head -1 | grep -oE '[0-9]+\.[0-9]+$' || true)
fi
if [ -z "$v" ]; then
echo " libc unknown (no ldd) — 'verify' is the real test"
return 0
fi
echo " libc glibc $v"
if [ "$(printf '%s\n2.34\n' "$v" | sort -V | head -1)" != "2.34" ]; then
echo " ! older than glibc 2.34, which tilt needs. kubectl, kind, jq and"
echo " ctlptl are static or libc-only and work here; tilt will not start."
echo " Install the core tier, or run tilt from a container."
fi
return 0
}
# What this script needs to do its own job. Reported here so `detect` answers
# "will install work?" instead of leaving you to find out one download in.
# Amazon Linux 2 ships without tar, which is exactly the surprise this catches.
detect_prereqs() {
local missing=""
if command -v curl >/dev/null 2>&1; then echo " download curl"
elif command -v wget >/dev/null 2>&1; then echo " download wget"
else echo " ! no curl and no wget — nothing can be downloaded"; missing+=" curl"
fi
if command -v sha256sum >/dev/null 2>&1 || command -v shasum >/dev/null 2>&1; then
echo " checksums ok"
else
echo " ! no sha256sum or shasum — downloads could not be verified"
missing+=" coreutils"
fi
if command -v tar >/dev/null 2>&1 && command -v gzip >/dev/null 2>&1; then
echo " archives tar + gzip"
else
echo " ! no tar/gzip — tilt and ctlptl ship as tarballs, so the dev tier"
echo " cannot be unpacked. The core tier is two bare binaries and is fine."
missing+=" tar gzip"
fi
if [ -n "$missing" ]; then
MANUAL+=("Install what this script needs to run at all:
$(pkg_install_cmd "${missing# }")")
fi
return 0
}
detect_docker() {
# Reachability of the daemon is the real question, and the CLI is only how
# we ask it. Note that when this runs inside the installer container, Docker
# necessarily exists on the host — otherwise nothing would be executing —
# so a missing CLI in here is an installer packaging bug, not a host problem.
if ! command -v docker >/dev/null 2>&1; then
if [ -S /var/run/docker.sock ]; then
echo " docker socket present (no cli in this context)"
else
echo " ! docker not found and no socket at /var/run/docker.sock"
MANUAL+=("Install Docker — the one true prerequisite, and the only thing here
that needs root:
$(pkg_install_cmd "$(docker_pkg)")
sudo systemctl enable --now docker
sudo usermod -aG docker \"\$USER\"
then log out and back in, so the new group applies to your shell.")
fi
return
fi
if docker info >/dev/null 2>&1; then
echo " docker $(docker version --format '{{.Server.Version}}' 2>/dev/null)"
local n
n=$(docker ps --filter "label=io.x-k8s.kind.cluster" --format '{{.Names}}' 2>/dev/null | wc -l)
# Must be an `if`, not `[ ] && echo`: as the last statement in this
# function the latter returns 1 when the count is zero, and `set -e`
# then kills the caller. That is the fresh-machine case — no clusters
# yet — so the bug only ever shows up where it does most harm.
if [ "$n" -gt 0 ]; then
echo " - $n kind node container(s) already running; see 'make cluster list'"
fi
else
echo " ! docker cli present but the daemon is unreachable"
MANUAL+=("Start Docker, or add yourself to the docker group:
sudo usermod -aG docker \"\$USER\" # then log out and back in")
fi
}
# kind and Tilt both watch large trees. WSL ships defaults (8192/128) far too low,
# and the failure mode is silent: Tilt simply stops noticing file changes.
detect_inotify() {
local w i
w=$(cat /proc/sys/fs/inotify/max_user_watches 2>/dev/null || echo 0)
i=$(cat /proc/sys/fs/inotify/max_user_instances 2>/dev/null || echo 0)
echo " inotify watches=$w instances=$i"
if [ "$w" -lt 524288 ] || [ "$i" -lt 512 ]; then
echo " ! inotify limits are low — Tilt will silently stop noticing file changes"
MANUAL+=("Raise inotify limits (needs root on the host):
echo -e 'fs.inotify.max_user_watches=524288\\nfs.inotify.max_user_instances=512' \\
| sudo tee /etc/sysctl.d/99-rig.conf
sudo sysctl --system")
fi
}
# ── fetch ──────────────────────────────────────────────────────────────────
# Resolve where a given artifact comes from, honouring DEPS_SOURCE.
resolve_url() {
local upstream="$1"
case "$DEPS_SOURCE" in
upstream) echo "$upstream" ;;
artifactory)
if [ -z "$DEPS_ARTIFACTORY_URL" ]; then
echo "DEPS_SOURCE=artifactory but DEPS_ARTIFACTORY_URL is empty" >&2
exit 1
fi
echo "${DEPS_ARTIFACTORY_URL%/}/$(basename "$upstream")"
;;
*) echo "unsupported DEPS_SOURCE '$DEPS_SOURCE' for a download" >&2; exit 1 ;;
esac
}
verify() {
local file="$1" want="$2" name="$3" got
got=$($SHA "$file" | awk '{print $1}')
if [ "$got" != "$want" ]; then
echo "checksum mismatch for $name" >&2
echo " expected $want" >&2
echo " got $got" >&2
exit 1
fi
}
# fetch_bin <name> <url> <sha256> <dest-dir> — a bare binary
fetch_bin() {
local name="$1" url="$2" sha="$3" dest="$4"
local tmp="$dest/.$name.tmp"
echo " fetching $name"
download "$(resolve_url "$url")" "$tmp"
verify "$tmp" "$sha" "$name"
mv "$tmp" "$dest/$name"
chmod +x "$dest/$name"
}
# fetch_tgz <name> <url> <sha256> <dest-dir> <path-inside-archive> <strip>
# Archive layouts differ — tilt's is flat (the binary at the root, strip=0),
# others nest it a directory down — so the caller says which.
fetch_tgz() {
local name="$1" url="$2" sha="$3" dest="$4" inner="$5" strip="$6"
local tmp="$dest/.$name.tgz"
echo " fetching $name"
download "$(resolve_url "$url")" "$tmp"
verify "$tmp" "$sha" "$name"
# --no-same-owner: extracting as root would otherwise restore the uid/gid
# baked into the archive (some ship as uid 1001), leaving a binary the host
# user does not own.
tar -xzf "$tmp" -C "$dest" --strip-components="$strip" --no-same-owner "$inner"
rm -f "$tmp"
chmod +x "$dest/$name"
}
# The installer runs as root so it can reach the docker socket, which means
# everything it writes into a mounted volume lands root-owned and unusable from
# the host. Hand it back to whoever owns the mount point (the host user created
# that directory before mounting it).
fix_ownership() {
local dir="$1"
[ -d "$dir" ] || return 0
local owner="${HOST_UID:-}:${HOST_GID:-}"
if [ "$owner" = ":" ]; then
owner=$(stat -c '%u:%g' "$dir")
fi
[ "$owner" = "0:0" ] && return 0
chown -R "$owner" "$dir" 2>/dev/null || true
}
# Two tiers, because not every machine should get cluster tooling.
#
# core kubectl, jq — talk to a cluster someone else runs. Nothing that
# creates one. Appropriate on a managed or corporate-issued machine
# where development tools are not wanted by default.
# dev core plus kind and tilt — build clusters and hot-reload into them.
#
# The split exists because "install the toolchain" is not one decision: on a
# managed workspace the right answer is kubectl and nothing else.
CORE_TOOLS="kubectl jq"
# No helm: every addon installs with `kubectl apply -f <url>`, so nothing here
# has ever invoked it. Add it back the day something actually needs a chart.
#
# ctlptl is 'dev' rather than 'core' for the same reason kind is: core is "talk
# to a cluster someone else runs", and ctlptl builds them. It earns its place
# because it is what wires a cluster to a local registry — without one, an
# unqualified image name resolves to docker.io/library/<name> and there is
# nothing structural stopping a push there.
#
# docker-compose is 'dev' for the same reason, and is here because the distro
# docker packages ship the daemon and CLI but frequently not the compose
# plugin — so `docker compose up` fails with "unknown command" on an otherwise
# working Docker, and nothing about that message names the missing piece.
DEV_TOOLS="kind tilt ctlptl docker-compose"
# ── what is already on this machine ───────────────────────────────────────
#
# A tool already on PATH at its pinned version is left where it is. Without
# this, install downloads a second copy into OUT_BIN and then reports the first
# one as shadowed — noise, and wrong, when both are the same version. That is
# the normal state of any machine someone set up by hand, whatever directory
# they happened to choose.
pin_of() {
case "$1" in
kubectl) echo "$KUBECTL_VERSION" ;;
jq) echo "$JQ_VERSION" ;;
kind) echo "$KIND_VERSION" ;;
tilt) echo "$TILT_VERSION" ;;
ctlptl) echo "$CTLPTL_VERSION" ;;
docker-compose) echo "$COMPOSE_VERSION" ;;
esac
}
# The version string a binary reports. Each tool spells the question
# differently, and kubectl has to be told --client or it goes looking for a
# server to ask.
reported_version() {
local tool="$1" path="$2"
case "$tool" in
kubectl) "$path" version --client 2>/dev/null ;;
jq) "$path" --version 2>/dev/null ;;
*) "$path" version 2>/dev/null ;;
esac
}
# Does the binary at PATH report PIN? Matched as a whole version token, so
# 0.37.6 never matches 10.37.60, with the leading v optional either side: kind
# says v0.32.0, jq says jq-1.8.2, and tilt says v0.37.6 against a pin of 0.37.6.
#
# Bash's own regex rather than grep, deliberately. grep is not the same program
# on every machine — some builds reject patterns that others accept — and a
# failed grep inside a count reads exactly like a zero.
version_matches() {
local tool="$1" path="$2" pin="$3" out v re
out=$(reported_version "$tool" "$path") || return 1
v="${pin#v}"
v="${v//./\\.}"
re="(^|[^0-9.])v?${v}([^0-9.]|\$)"
[[ $out =~ $re ]]
}
# DEPS_ONLY narrows a fetch to the tools it names. Unset means the whole tier,
# which is what an explicit `deps.sh fetch` always gets: "download these into
# DIR" must not quietly skip something because this machine happens to have it.
# Only install() sets it, to what detect_toolchain found missing or mismatched.
want() { [ -z "${DEPS_ONLY:-}" ] || [[ " $DEPS_ONLY " == *" $1 "* ]]; }
# Every tool in the tier with its state, probed once and reported once. What
# still needs fetching is left in TOOLCHAIN_NEED for install() to act on.
TOOLCHAIN_NEED=""
detect_toolchain() {
local tier="${TIER:-dev}" b pin path found
TOOLCHAIN_NEED=""
echo
echo "toolchain (pinned, tier '$tier')"
for b in $(tier_tools "$tier"); do
pin=$(pin_of "$b")
path=$(command -v "$b" 2>/dev/null || true)
# compose is the one tool that is normally NOT a binary on PATH. It is a
# docker CLI plugin, so a machine where `docker compose` works perfectly
# has no `docker-compose` to find — and probing only PATH would report it
# missing and re-download a copy that is already there. That is the exact
# noise the version-aware skip exists to prevent, so ask docker instead.
if [ "$b" = docker-compose ] && [ -z "$path" ]; then
if found=$(docker compose version --short 2>/dev/null) && [ -n "$found" ]; then
if [ "${found#v}" = "${pin#v}" ]; then
printf " %-8s %-9s %s\n" "$b" "$pin" "docker cli plugin"
else
printf " ! %-8s wants %s, the docker cli plugin reports '%s'\n" \
"$b" "$pin" "$found"
TOOLCHAIN_NEED+="$b "
fi
continue
fi
fi
if [ -z "$path" ]; then
printf " - %-8s %-9s not found\n" "$b" "$pin"
TOOLCHAIN_NEED+="$b "
elif version_matches "$b" "$path" "$pin"; then
printf " %-8s %-9s %s\n" "$b" "$pin" "$path"
else
found=$(reported_version "$b" "$path" 2>/dev/null | head -1 || true)
printf " ! %-8s wants %s, %s reports '%s'\n" "$b" "$pin" "$path" "$found"
TOOLCHAIN_NEED+="$b "
fi
done
if [ -z "$TOOLCHAIN_NEED" ]; then
echo " every pinned tool is already on PATH — nothing to fetch"
else
echo " 'make deps' fetches only: ${TOOLCHAIN_NEED% }"
fi
}
fetch() {
local dest="$OUT_BIN" tier="${TIER:-dev}"
while [ $# -gt 0 ]; do
case "$1" in
--to) dest="$2"; shift 2 ;;
core|dev) tier="$1"; shift ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
dest="$(abspath "$dest")"
mkdir -p "$dest"
TIER="$tier"
if [ "$DEPS_SOURCE" = "baked" ]; then
echo "installing baked binaries from $BAKED_BIN"
cp -a "$BAKED_BIN"/. "$dest"/
fix_ownership "$dest"
return
fi
if [ -n "${DEPS_ONLY:-}" ]; then
echo "fetching ${DEPS_ONLY% } (source: $DEPS_SOURCE)"
else
echo "fetching '$tier' toolchain (source: $DEPS_SOURCE)"
fi
if want kubectl; then fetch_bin kubectl "$KUBECTL_URL" "$KUBECTL_SHA256" "$dest"; fi
if want jq; then fetch_bin jq "$JQ_URL" "$JQ_SHA256" "$dest"; fi
if [ "$tier" = "dev" ]; then
if want kind; then fetch_bin kind "$KIND_URL" "$KIND_SHA256" "$dest"; fi
if want tilt; then fetch_tgz tilt "$TILT_URL" "$TILT_SHA256" "$dest" tilt 0; fi
if want ctlptl; then fetch_tgz ctlptl "$CTLPTL_URL" "$CTLPTL_SHA256" "$dest" ctlptl 0; fi
if want docker-compose; then
fetch_bin docker-compose "$COMPOSE_URL" "$COMPOSE_SHA256" "$dest"
fi
fi
fix_ownership "$dest"
# kind writes the kubeconfig as root too; hand that back as well when it's
# a mounted host directory rather than container-local state.
fix_ownership "${KUBE_DIR:-/out/kube}"
}
# ── install ────────────────────────────────────────────────────────────────
report_manual() {
echo
if [ ${#MANUAL[@]} -eq 0 ]; then
echo "nothing left to do by hand."
return
fi
echo "host actions this cannot perform (${#MANUAL[@]}):"
echo
local n=1
for m in "${MANUAL[@]}"; do
echo " $n. $m"
echo
n=$((n + 1))
done
}
# Installing into a directory that sits early in PATH silently replaces whatever
# the machine was already using — which on a shared or client machine can break
# unrelated work (kubectl more than one minor away from a cluster is the common
# one). Say so; never decide it for them.
# Downloading a verified binary proves it is the right file, not that this
# machine can run it. On an old distro tilt fails here, with a linker error
# about a missing symbol, and finding that out now beats finding out during a
# first cluster build.
verify_tools() {
local tier="${1:-dev}" b bin out rc broke=0
echo "checking that each one actually runs"
for b in $(tier_tools "$tier"); do
bin="$OUT_BIN/$b"
if [ ! -x "$bin" ]; then
printf ' %-14s not installed\n' "$b"
continue
fi
# Not piped into `head`. With `pipefail` set, a tool that prints more
# than one line gets SIGPIPE when head closes the pipe, and the
# pipeline reports 141 — so a working kubectl was announced as "does
# not run here", with its own correct version string as the evidence.
# Take the first line afterwards, from the string.
rc=0
case "$b" in
kubectl) out=$("$bin" version --client 2>&1) || rc=$? ;;
jq) out=$("$bin" --version 2>&1) || rc=$? ;;
*) out=$("$bin" version 2>&1) || rc=$? ;;
esac
out=${out%%$'\n'*}
if [ "$rc" -eq 0 ]; then
printf ' %-14s %s\n' "$b" "$out"
else
printf ' ! %-12s does not run here: %s\n' "$b" "$out"
broke=1
fi
done
if [ "$broke" -eq 1 ]; then
echo
echo " A binary that downloads and verifies but will not start is almost"
echo " always this distro's libc being older than the release needs."
echo " 'detect' prints the glibc version. The core tier (kubectl + jq)"
echo " has no such dependency and will work regardless."
fi
return 0
}
list() {
echo "pinned, linux/amd64 only:"
printf ' %-14s %s\n' kubectl "$KUBECTL_VERSION"
printf ' %-14s %s\n' jq "$JQ_VERSION"
printf ' %-14s %s\n' kind "$KIND_VERSION"
printf ' %-14s %s\n' tilt "$TILT_VERSION"
printf ' %-14s %s\n' ctlptl "$CTLPTL_VERSION"
printf ' %-14s %s\n' docker-compose "$COMPOSE_VERSION"
echo
echo " core = $CORE_TOOLS"
echo " dev = $CORE_TOOLS $DEV_TOOLS"
echo
echo "Checksums are pinned in the block at the top of this file. To bump one,"
echo "take the new checksum from the publisher's own release list — the header"
echo "comment has the exact commands."
return 0
}
tier_tools() { [ "$1" = "core" ] && echo "$CORE_TOOLS" || echo "$CORE_TOOLS $DEV_TOOLS"; }
warn_shadowing() {
local b existing shadowed="" tier="${1:-dev}"
for b in $(tier_tools "$tier"); do
[ -x "$OUT_BIN/$b" ] || continue
# Where would this resolve if OUT_BIN weren't in the way?
existing=$(PATH=$(echo "$PATH" | tr ':' '\n' | grep -vx "$OUT_BIN" | paste -sd:) \
command -v "$b" 2>/dev/null || true)
[ -n "$existing" ] || continue
[ "$existing" = "$OUT_BIN/$b" ] && continue
# The same version in both places is not a conflict: nothing changes for
# any other project whichever copy PATH happens to find first.
if version_matches "$b" "$existing" "$(pin_of "$b")"; then continue; fi
shadowed+=" $b $existing"$'\n'
done
[ -n "$shadowed" ] || return 0
case ":${PATH}:" in
*":$OUT_BIN:"*) ;;
*) return 0 ;; # not on PATH yet, so nothing is being shadowed
esac
echo
echo " ! these were already installed elsewhere and are now shadowed by $OUT_BIN:"
printf '%s' "$shadowed"
echo " Other projects on this machine will pick up the new versions."
MANUAL+=("Decide which toolchain wins. To keep the previous one, remove what
was just installed:
rm -f $(for b in $(tier_tools "$tier"); do printf '%s ' "$OUT_BIN/$b"; done)
Or install somewhere private instead:
OUT_BIN=\$PWD/def/bin make deps # then put that dir first in PATH")
}
# A copy in OUT_BIN only gives you `docker-compose`. That hyphenated form is the
# retired v1 spelling; every compose file written in the last few years assumes
# `docker compose`, which resolves plugins BY NAME out of a plugin directory.
# So the binary is fetched like any other and then linked, in your own home —
# no root, and nothing outside it.
install_compose_plugin() {
local src="$OUT_BIN/docker-compose" dir="$HOME/.docker/cli-plugins"
[ -x "$src" ] || return 0
mkdir -p "$dir"
# Something else already owns that name — docker-desktop and some distro
# packages install a real file there. Overwriting it would take the plugin
# away from whatever put it there, so say so and let the user decide.
if [ -e "$dir/docker-compose" ] && [ ! -L "$dir/docker-compose" ]; then
MANUAL+=("Something already installs the compose plugin at
$dir/docker-compose
To use rig's pinned build instead:
ln -sf $src $dir/docker-compose")
return 0
fi
ln -sfn "$src" "$dir/docker-compose"
echo " compose plugin -> $dir/docker-compose"
return 0
}
install() {
local tier="${1:-dev}" b
TIER="$tier"
detect
# detect_toolchain has already probed PATH. Fetch only what it found missing
# or at the wrong version; a tool already present at its pin stays where it is.
if [ -n "$TOOLCHAIN_NEED" ]; then
echo
DEPS_ONLY="$TOOLCHAIN_NEED" fetch "$tier"
echo
echo "installed to $OUT_BIN ($tier):"
for b in $TOOLCHAIN_NEED; do
if [ -x "$OUT_BIN/$b" ]; then echo " $b"; fi
done
if [ "$tier" = "core" ]; then
echo " (no kind/tilt — 'make deps dev' adds them)"
fi
# Only when compose was one of the things fetched: linking a binary
# that is already satisfied elsewhere on PATH would point the plugin at
# a copy rig did not install.
case " $TOOLCHAIN_NEED " in
*" docker-compose "*) install_compose_plugin ;;
esac
# Only worth saying when something actually landed in OUT_BIN. When every
# tool was satisfied elsewhere, OUT_BIN may reasonably be off PATH, and
# telling the user to add it would be advice to fix nothing.
case ":${PATH}:" in
*":$OUT_BIN:"*) ;;
*) MANUAL+=("Put the toolchain on your PATH — add to ~/.bashrc:
export PATH=\"${OUT_BIN}:\$PATH\"") ;;
esac
fi
warn_shadowing "$tier"
report_manual
}
# ── main ───────────────────────────────────────────────────────────────────
require_linux
# Read the command, THEN shift — and shift only if there is something there.
# A bare `shift` with no positional parameters returns 1, and under `set -e`
# that ended the script before a single line was printed: running this with no
# arguments at all, the documented default, did nothing and said nothing.
cmd="${1:-install}"
[ $# -gt 0 ] && shift
# Baked mode copies binaries already in the image, so it needs no downloader.
need_downloads() {
require_amd64
if [ "$DEPS_SOURCE" != baked ]; then pick_downloader; fi
pick_sha
}
case "$cmd" in
detect) detect; report_manual ;;
list) list ;;
verify) verify_tools "${1:-dev}" ;;
fetch) need_downloads; fetch "$@" ;;
install) need_downloads; install "${1:-dev}" ;;
*) echo "usage: $0 [detect|list|verify|fetch|install]" >&2
echo " install [core|dev] (default dev)" >&2
echo " fetch [core|dev] [--to DIR]" >&2
echo " OUT_BIN=<dir> overrides the install directory" >&2
exit 1 ;;
esac

860
rig/standalone/default/rigmini.sh Executable file
View File

@@ -0,0 +1,860 @@
#!/usr/bin/env bash
# GENERATED by make standalone — do not edit
#
# rigmini.sh for profile 'default', flattened from:
# ctrl/mem.sh
# ctrl/lib/config.sh
# Edit those and run `make standalone`. Changes made here are lost, and
# `make selftest` fails while this file differs from what rig generates.
# ── from the libraries ──
declare -- CONFIG_OVERRIDABLE=$'PROFILE CLUSTER K8S_VERSION KIND_CONFIG ADDONS\n REGISTRY_MODE INGRESS_MODE DNS_MODE TILT_PORT\n SOURCE ARCH DEPS_SOURCE HTTP_PORT HTTPS_PORT\n REGISTRY_PORT MANIFESTS_DIR'
_config_restore ()
{
local line;
while IFS= read -r line; do
if [ -n "$line" ]; then
eval "export $line";
fi;
done <<< "$1";
return 0
}
default_cluster_name ()
{
local n;
n=$(basename "$(cd .. && pwd)");
n=$(echo "$n" | tr '[:upper:]' '[:lower:]' | tr -c 'a-z0-9-' '-');
n=$(echo "$n" | sed 's/^-*//; s/-*$//');
echo "${n:-rig}"
}
derive_port_base ()
{
local h;
h=$(printf '%s' "$1" | cksum | awk '{print $1}');
echo $((20000 + (h % 200) * 10))
}
render_kind_config ()
{
local host_workdir="${HOST_WORKDIR:-$(cd .. && pwd)}";
sed -e "s|\${CLUSTER}|${CLUSTER}|g" -e "s|\${NODE_IMAGE}|${NODE_IMAGE}|g" -e "s|\${HTTP_PORT}|${HTTP_PORT}|g" -e "s|\${HOST_WORKDIR}|${host_workdir}|g" "$KIND_CONFIG_PATH"
}
# ── configuration, frozen for profile 'default' ──
load_config() {
local k saved=""
for k in $CONFIG_OVERRIDABLE; do
if [ -n "${!k+x}" ]; then saved+="$k=$(printf '%q' "${!k}")"$'\n'; fi
done
declare -g ADDONS=""
declare -gx AIRFLOW_IMAGE="apache/airflow:2.10.4"
declare -g AUDIT="off"
declare -gx CERT_MANAGER_VERSION="v1.21.1"
declare -g CLUSTER="rig"
declare -gx COMPOSE_SHA256="db1889184726840f75c4f9c001048430d4f25b3be3cb084d3ddd762bc0aed576"
declare -gx COMPOSE_URL="https://github.com/docker/compose/releases/download/v5.5.1/docker-compose-linux-x86_64"
declare -gx COMPOSE_VERSION="5.5.1"
declare -gx CTLPTL_SHA256="c63a1ec28e60bc3faf6becb76f53355c5cf5e0143dafdd27ad85db5584fa6b1e"
declare -gx CTLPTL_URL="https://github.com/tilt-dev/ctlptl/releases/download/v0.9.4/ctlptl.0.9.4.linux.x86_64.tar.gz"
declare -gx CTLPTL_VERSION="0.9.4"
declare -g DNS_MODE="hosts"
declare -g HTTPS_PORT="20311"
declare -g HTTP_PORT="20310"
declare -g INGRESS_MODE="hostport"
declare -gx JQ_SHA256="b1c22172dd303f3be49e935aa56aa48a8b7a46e0bc838b4997d3bb451495870f"
declare -gx JQ_URL="https://github.com/jqlang/jq/releases/download/jq-1.8.2/jq-linux-amd64"
declare -gx JQ_VERSION="1.8.2"
declare -g K8S_VERSION="v1_36"
declare -g KIND_CONFIG="kind-config.yaml.tpl"
declare -g KIND_CONFIG_PATH="./k8s/kind-config.yaml.tpl"
declare -g KIND_CONFIG_SHOWN="ctrl/k8s/kind-config.yaml.tpl"
declare -gx KIND_SHA256="50030de23cf40a18505f20426f6a8506bedf13c6e509244bd1fa9463721b0f54"
declare -gx KIND_URL="https://github.com/kubernetes-sigs/kind/releases/download/v0.32.0/kind-linux-amd64"
declare -gx KIND_VERSION="v0.32.0"
declare -g KUBECONTEXT="kind-rig"
declare -gx KUBECTL_SHA256="ebbd080e7c2e275093b55915722043257eb24004363e20acb3c4d71919f88336"
declare -gx KUBECTL_URL="https://dl.k8s.io/release/v1.36.3/bin/linux/amd64/kubectl"
declare -gx KUBECTL_VERSION="v1.36.3"
declare -g MANIFESTS_DIR="ctrl/k8s/overlays/dev"
declare -gx METALLB_VERSION="v0.16.0"
declare -gx METRICS_SERVER_VERSION="v0.9.0"
declare -g NODES="1"
declare -g NODE_IMAGE="kindest/node:v1.36.1@sha256:3489c7674813ba5d8b1a9977baea8a6e553784dab7b84759d1014dbd78f7ebd5"
declare -gx NODE_IMAGE_v1_33="kindest/node:v1.33.12@sha256:3f5c8443c620245e4d355cfe09e96a91ead32ceaa569d3f1ca9edf0cb2fe2ff4"
declare -gx NODE_IMAGE_v1_34="kindest/node:v1.34.8@sha256:02722c2dedddcfc00febf5d27fbeb9b7b2c14294c82109ff4a85d89ac9ba3256"
declare -gx NODE_IMAGE_v1_35="kindest/node:v1.35.5@sha256:ce977ae6d65918d0b58a5f8b5e940429c2ce42fa3a5619ec2bbc60b949c0ac95"
declare -gx NODE_IMAGE_v1_36="kindest/node:v1.36.1@sha256:3489c7674813ba5d8b1a9977baea8a6e553784dab7b84759d1014dbd78f7ebd5"
declare -g NODE_MB="800"
declare -gx POSTGRES_IMAGE="postgres:16-alpine"
declare -gx PROFILE="default"
declare -g PROFILE_NAME="default"
declare -gx REDIS_IMAGE="redis:7-alpine"
declare -gx REGISTRY_IMAGE="registry:2"
declare -g REGISTRY_MODE="local"
declare -g REGISTRY_PORT="20313"
declare -gx STUB_IMAGE="python:3.12-slim"
declare -g TILT_PORT="20312"
declare -gx TILT_SHA256="e9672b8a18d43501f35dcfe98465969a7db0e436b36cf0c50c7e6f8d40de5fe6"
declare -gx TILT_URL="https://github.com/tilt-dev/tilt/releases/download/v0.37.6/tilt.0.37.6.linux.x86_64.tar.gz"
declare -gx TILT_VERSION="0.37.6"
_config_restore "$saved"
}
# ── end of frozen configuration ──
# ── ctrl/mem.sh ──
# How much memory this machine will actually give you before something dies —
# rig's memory tool, and (generated from this file) the standalone rigmini.sh.
#
# There are two numbers and they are rarely the same. `status` reports what the
# machine ADVERTISES and what is quietly capping it. `push` finds what it will
# SURVIVE, by allocating until it stops. `all` does both and weighs the result
# against what this profile's cluster needs.
#
# The gap between them is the whole reason this exists. Under WSL the cap lives
# in .wslconfig; in a container or a managed workspace it is a cgroup limit, and
# there /proc/meminfo reports the HOST's memory while the kernel kills you at a
# fraction of it. A script that only read MemTotal would confidently report 32 GB
# on a box that OOMs at 2.
#
# Runs on native Linux and under WSL. On WSL the memory you see is a VM
# allocation that can be raised, and the commonest failure is raising it without
# restarting — so status compares what .wslconfig says with what actually booted.
#
# Reports and instructs. It never raises a limit, frees anything or installs a
# package. The one write it can make is `backup`, which copies .wslconfig beside
# itself, so that `restore` has something to put back after a hand edit.
#
# Usage:
# mem.sh status what it has, what caps it
# mem.sh push [--to GB] [--to-oom] climb until it stops
# mem.sh all [--budget GB] both, then the verdict
# mem.sh backup | restore .wslconfig, WSL only
set -euo pipefail
cd "$(dirname "$0")"
# (sourced library inlined above)
# ── defaults ───────────────────────────────────────────────────────────────
STEP_MB=0 # per allocation; 0 means scale it to the ceiling. See push().
STEP_EXPLICIT=no # whether --step was given, which turns the scaling off.
TO_MB="" # --to: stop here regardless. Empty means no hard cap.
TO_OOM=no # --to-oom: opt in to running until the kernel intervenes.
BUDGET_GB="" # --budget; empty means what this profile's cluster needs, from rig.
BUDGET_EXPLICIT=no # whether --budget was given, which retires the guess below.
# ── platform ───────────────────────────────────────────────────────────────
# Windows outside WSL — Git Bash, MSYS, Cygwin — looks close enough to work and
# then fails in a pile of confusing ways: no /proc, no docker socket, none of
# the tooling. Detectable, so name it instead.
require_linux() {
case "$(uname -s)" in
MINGW*|MSYS*|CYGWIN*)
cat >&2 <<'EOF'
This has to run inside WSL, not Git Bash / MSYS / Cygwin.
If WSL is not installed yet, from an elevated PowerShell or Command Prompt:
wsl --install
That enables Windows features and needs a reboot, so it is not something this
script will do for you. Afterwards, open the Linux shell it installs and run
this from there.
EOF
exit 1 ;;
esac
# Everything below reads /proc. Without it there is nothing to measure, and
# failing here beats printing a page of empty fields.
if [ ! -r /proc/meminfo ]; then
echo "no readable /proc/meminfo — this needs a Linux kernel." >&2
echo "On macOS or a BSD none of the numbers below exist." >&2
exit 1
fi
}
is_wsl() { grep -qi microsoft /proc/version 2>/dev/null; }
is_container() {
[ -f /.dockerenv ] && return 0
grep -qE '(docker|containerd|kubepods|lxc|podman)' /proc/1/cgroup 2>/dev/null
}
platform() {
if is_wsl; then echo WSL
elif is_container; then echo container
else echo "native linux"
fi
}
# ── reading memory ─────────────────────────────────────────────────────────
mb() { echo $(( $(awk "/^$1:/{print \$2}" /proc/meminfo) / 1024 )); }
# MemAvailable arrived in kernel 3.14. Older kernels — and they turn up on
# corporate images — need the estimate it replaced, which is worse but not wrong.
avail_meminfo_mb() {
if grep -q '^MemAvailable:' /proc/meminfo; then
mb MemAvailable
else
awk '/^(MemFree|Buffers|Cached):/{t+=$2} END{print int(t/1024)}' /proc/meminfo
fi
}
# Where a cgroup records this cgroup's own limit and usage. Set once by
# find_cgroup, because every later reading needs both and hunting for the files
# on each call would be the slow part of the poll loop.
CG_MAX_FILE=""
CG_CUR_FILE=""
CG_VERSION=""
find_cgroup() {
local rel
# Inside a container the cgroup namespace makes the top of the tree BE the
# container's own cgroup, so the unqualified path is already the right one.
# On a host it is the root cgroup, which is never limited — hence the second
# attempt via /proc/self/cgroup, which names the slice this shell is in.
if [ -r /sys/fs/cgroup/memory.max ]; then
CG_VERSION=v2
CG_MAX_FILE=/sys/fs/cgroup/memory.max
CG_CUR_FILE=/sys/fs/cgroup/memory.current
elif [ -r /sys/fs/cgroup/memory/memory.limit_in_bytes ]; then
CG_VERSION=v1
CG_MAX_FILE=/sys/fs/cgroup/memory/memory.limit_in_bytes
CG_CUR_FILE=/sys/fs/cgroup/memory/memory.usage_in_bytes
fi
rel=$(awk -F: '$1=="0"{print $3; exit}' /proc/self/cgroup 2>/dev/null || true)
if [ -n "$rel" ] && [ "$rel" != "/" ] && [ -r "/sys/fs/cgroup${rel}/memory.max" ]; then
CG_VERSION=v2
CG_MAX_FILE="/sys/fs/cgroup${rel}/memory.max"
CG_CUR_FILE="/sys/fs/cgroup${rel}/memory.current"
return 0
fi
rel=$(awk -F: '$2 ~ /(^|,)memory(,|$)/{print $3; exit}' /proc/self/cgroup 2>/dev/null || true)
if [ -n "$rel" ] && [ "$rel" != "/" ] \
&& [ -r "/sys/fs/cgroup/memory${rel}/memory.limit_in_bytes" ]; then
CG_VERSION=v1
CG_MAX_FILE="/sys/fs/cgroup/memory${rel}/memory.limit_in_bytes"
CG_CUR_FILE="/sys/fs/cgroup/memory${rel}/memory.usage_in_bytes"
fi
return 0
}
# The cap in MB, or "" when there is none worth reporting. v2 spells unlimited
# "max"; v1 spells it as a number near 2^63, which is why this compares against
# MemTotal rather than testing for a magic value — a "limit" above the machine's
# own memory is not a limit, however it is written.
cgroup_cap_mb() {
local raw cap
[ -n "$CG_MAX_FILE" ] && [ -r "$CG_MAX_FILE" ] || { echo ""; return 0; }
raw=$(cat "$CG_MAX_FILE" 2>/dev/null || echo max)
[ "$raw" = "max" ] && { echo ""; return 0; }
case "$raw" in ''|*[!0-9]*) echo ""; return 0 ;; esac
cap=$((raw / 1024 / 1024))
[ "$cap" -ge "$(mb MemTotal)" ] && { echo ""; return 0; }
echo "$cap"
}
cgroup_used_mb() {
local raw
[ -n "$CG_CUR_FILE" ] && [ -r "$CG_CUR_FILE" ] || { echo ""; return 0; }
raw=$(cat "$CG_CUR_FILE" 2>/dev/null || echo "")
case "$raw" in ''|*[!0-9]*) echo ""; return 0 ;; esac
echo $((raw / 1024 / 1024))
}
# ulimit -v is a per-process address-space cap. It stops YOU long before the box
# does, and because it is inherited from a login shell it is easy to hit without
# knowing it is set.
ulimit_v_mb() {
local v; v=$(ulimit -v 2>/dev/null || echo unlimited)
[ "$v" = "unlimited" ] && { echo ""; return 0; }
case "$v" in ''|*[!0-9]*) echo ""; return 0 ;; esac
echo $((v / 1024))
}
# The number everything else is about: the lowest of the things that can stop
# you. Printed at the end of `status` and used as the sanity bound in `push`.
effective_ceiling_mb() {
local c; c=$(mb MemTotal)
local cap; cap=$(cgroup_cap_mb)
local ul; ul=$(ulimit_v_mb)
[ -n "$cap" ] && [ "$cap" -lt "$c" ] && c="$cap"
[ -n "$ul" ] && [ "$ul" -lt "$c" ] && c="$ul"
echo "$c"
}
# How much room is left RIGHT NOW, from whichever accounting actually governs.
# In a capped container /proc/meminfo describes the host and is worse than
# useless for this — it would report tens of gigabytes free on a box that is one
# allocation from being killed.
headroom_mb() {
local cap used
cap=$(cgroup_cap_mb)
used=$(cgroup_used_mb)
if [ -n "$cap" ] && [ -n "$used" ]; then
echo $(( cap - used ))
else
avail_meminfo_mb
fi
}
# ── status ─────────────────────────────────────────────────────────────────
# /mnt/c/Users can hold several real accounts — a renamed login leaves the old
# directory behind — so picking the first alphabetically is a coin toss. Ask
# Windows, then fall back to whichever profile actually owns a config.
wslconfig_path() {
local profile winpath found
profile=$(cmd.exe /c "echo %USERPROFILE%" 2>/dev/null | tr -d "\r\n" || true)
case "$profile" in
""|*%*) ;;
*) winpath=$(wslpath -u "$profile" 2>/dev/null || true)
if [ -n "$winpath" ] && [ -d "$winpath" ]; then
echo "$winpath/.wslconfig"; return 0
fi ;;
esac
found=$(ls -d /mnt/c/Users/*/.wslconfig 2>/dev/null | head -1 || true)
[ -n "$found" ] && echo "$found"
return 0
}
hogs() {
echo " holding the most:"
ps -eo rss,comm --sort=-rss 2>/dev/null \
| awk 'NR>1 && NR<=6 {printf " %6.0f MB %s\n", $1/1024, $2}'
return 0
}
status() {
local total avail swap_total swap_free cap ul cur
echo "host"
echo " platform $(platform)"
echo " kernel $(uname -r)"
[ -r /etc/os-release ] && \
echo " distro $(sed -n 's/^PRETTY_NAME="\(.*\)"/\1/p' /etc/os-release)"
echo " cpu $(getconf _NPROCESSORS_ONLN 2>/dev/null || echo '?') online, load $(cut -d' ' -f1-3 /proc/loadavg)"
# ── the caps first, because they decide what the totals below are worth ──
echo
echo "caps"
cap=$(cgroup_cap_mb)
if [ -n "$cap" ]; then
cur=$(cgroup_used_mb)
echo " cgroup ${cap} MB (${CG_VERSION}, ${CG_CUR_FILE##*/} says ${cur:-?} MB used)"
echo " ! /proc/meminfo below describes the HOST, not this cgroup."
echo " $(mb MemTotal) MB total is not yours; ${cap} MB is."
elif [ -n "$CG_VERSION" ]; then
echo " cgroup none (${CG_VERSION} present, no memory limit set)"
else
echo " cgroup no memory controller found"
fi
ul=$(ulimit_v_mb)
if [ -n "$ul" ]; then
echo " ! ulimit -v ${ul} MB — a per-process cap, inherited from your shell"
echo " it stops this process long before the machine runs out"
else
echo " ulimit -v unlimited"
fi
# overcommit_memory=0 is the default heuristic: a large allocation is
# granted on a guess, and the reckoning arrives later as an OOM kill rather
# than as a failed malloc. It is why `push` touches every page it asks for.
local om or_
om=$(cat /proc/sys/vm/overcommit_memory 2>/dev/null || echo '?')
or_=$(cat /proc/sys/vm/overcommit_ratio 2>/dev/null || echo '?')
case "$om" in
0) echo " overcommit 0 heuristic — allocations are granted on a guess," ;;
1) echo " overcommit 1 always — every allocation succeeds; the OOM killer is the only limit," ;;
2) echo " overcommit 2 strict (ratio ${or_}%) — allocation fails honestly instead of killing later," ;;
*) echo " overcommit ${om}" ;;
esac
[ "$om" != "?" ] && echo " so RSS is the number to trust, not what a process asked for"
# ── what it says it has ──
total=$(mb MemTotal); avail=$(avail_meminfo_mb)
swap_total=$(mb SwapTotal); swap_free=$(mb SwapFree)
echo
echo "memory"
echo " total ${total} MB"
echo " available ${avail} MB"
echo " swap ${swap_total} MB ($(( swap_total - swap_free )) MB used)"
if [ "$swap_total" -eq 0 ]; then
echo " - no swap: this box has no cushion. It goes from fine to OOM-killed"
echo " with nothing in between, which is the abrupt failure you get in a VM."
fi
# postgres puts its shared buffers in /dev/shm. Docker's default is 64 MB,
# and the resulting failure names neither shm nor the size.
if [ -d /dev/shm ]; then
local shm; shm=$(df -Pm /dev/shm 2>/dev/null | awk 'NR==2{print $2}')
if [ -n "$shm" ]; then
if [ "$shm" -le 64 ]; then
echo " ! /dev/shm ${shm} MB — postgres puts shared memory here and 64 MB"
echo " is docker's default. Raise it with --shm-size when postgres fails."
else
echo " /dev/shm ${shm} MB"
fi
fi
fi
echo
echo "disk"
local d
for d in / /tmp /var/lib/docker; do
[ -d "$d" ] || continue
df -Pm "$d" 2>/dev/null | awk -v p="$d" 'NR==2{printf " %-12s %s MB free of %s MB\n", p, $4, $2}'
done
# kind and Tilt both watch large trees, and the failure mode is silent:
# they simply stop noticing file changes. Cheap to report while we are here.
local w i
w=$(cat /proc/sys/fs/inotify/max_user_watches 2>/dev/null || echo 0)
i=$(cat /proc/sys/fs/inotify/max_user_instances 2>/dev/null || echo 0)
echo
echo "tooling"
echo " inotify watches=$w instances=$i"
if [ "$w" -lt 524288 ] || [ "$i" -lt 512 ]; then
echo " ! low — anything watching files will silently stop seeing changes"
fi
if ! command -v docker >/dev/null 2>&1; then
if [ -S /var/run/docker.sock ]; then
echo " docker socket present, no cli"
else
echo " docker not installed"
fi
elif docker info >/dev/null 2>&1; then
local n
n=$(docker ps -q 2>/dev/null | wc -l)
echo " docker $(docker version --format '{{.Server.Version}}' 2>/dev/null), ${n} container(s) running"
else
echo " ! docker cli present but the daemon is unreachable"
fi
# WSL keeps its cap on the Windows side, in a file this shell can read but
# not usefully apply — the change costs a full VM restart. Report it, and
# report the commonest mistake, which is editing it and not restarting.
if is_wsl; then
local cfg conf conf_mb n
cfg=$(wslconfig_path)
echo
echo "wsl"
if [ -z "$cfg" ]; then
echo " ! cannot tell which Windows profile owns .wslconfig"
else
echo " config $cfg"
conf=$(configured_memory "$cfg")
if [ -n "$conf" ]; then
conf_mb=$(to_mb "$conf")
echo " configured $conf (${conf_mb} MB), booted ${total} MB"
# The VM reports a little less than allocated; 15% covers the
# kernel without calling every healthy machine a mismatch.
if [ -n "$conf_mb" ] && [ "$total" -lt $(( conf_mb * 85 / 100 )) ]; then
echo " ! configured ${conf_mb} MB but booted ${total} MB — not applied yet."
echo " From a WINDOWS terminal: wsl --shutdown then start the distro again."
fi
else
echo " configured no memory= set (WSL defaults to 50% of host RAM, or 8 GB,"
echo " whichever is less). To raise it, add on the Windows side:"
echo " [wsl2]"
echo " memory=8GB"
echo " then from a WINDOWS terminal: wsl --shutdown"
fi
n=$(ls "$cfg".*.bak 2>/dev/null | wc -l)
if [ "$n" -gt 0 ]; then
echo " backups $n (newest: $(ls -t "$cfg".*.bak 2>/dev/null | head -1))"
fi
fi
else
echo
echo " - native linux: no VM allocation to raise. If memory is tight the levers"
echo " are freeing something or adding swap."
fi
echo
echo "effective ceiling $(effective_ceiling_mb) MB"
echo " the lowest of MemTotal, the cgroup cap and ulimit -v. What the box"
echo " claims. 'push' measures what it will actually hand over."
[ "$avail" -lt $(( total / 5 )) ] && { echo; hogs; }
return 0
}
# ── .wslconfig ─────────────────────────────────────────────────────────────
require_wsl() {
if ! is_wsl; then
echo "$1 acts on .wslconfig, which only exists under WSL." >&2
echo "This is native Linux — there is no VM allocation to save or roll back." >&2
echo "Use 'status' to see what the machine actually has." >&2
exit 1
fi
}
# backup and restore act on the file, so unlike status they must not guess.
wslconfig_required() {
local cfg; cfg=$(wslconfig_required)
if [ -z "$cfg" ]; then
echo "cannot tell which Windows profile owns .wslconfig. Candidates:" >&2
ls -d /mnt/c/Users/*/ 2>/dev/null \
| grep -viE "/(All Users|Default|Default User|Public)/$" | sed "s/^/ /" >&2
exit 1
fi
echo "$cfg"
}
configured_memory() {
[ -r "$1" ] || { echo ""; return; }
sed -n 's/^[[:space:]]*memory[[:space:]]*=[[:space:]]*//p' "$1" | tail -1 | tr -d '[:space:]'
}
# "9GB" / "8192MB" / "9G" -> MB, so it can be compared with /proc/meminfo.
to_mb() {
local v="${1^^}" n
n=$(echo "$v" | tr -dc '0-9')
[ -n "$n" ] || { echo ""; return; }
case "$v" in
*GB|*G) echo $(( n * 1024 )) ;;
*MB|*M) echo "$n" ;;
*) echo $(( n / 1024 / 1024 )) ;;
esac
}
backup() {
require_wsl backup
local cfg dest
cfg=$(wslconfig_required)
[ -r "$cfg" ] || { echo "nothing to back up: $cfg does not exist" >&2; exit 1; }
# Timestamped and never overwritten: a backup that can destroy itself on a
# second run is not a backup.
dest="${cfg}.$(date +%Y%m%d-%H%M%S).bak"
cp "$cfg" "$dest"
echo "backed up $dest"
echo
echo "Edit $cfg by hand, then from a WINDOWS terminal: wsl --shutdown"
}
restore() {
require_wsl restore
local cfg newest count
cfg=$(wslconfig_required)
newest=$(ls -t "$cfg".*.bak 2>/dev/null | head -1 || true)
[ -n "$newest" ] || { echo "no backups found beside $cfg" >&2; exit 1; }
echo "restoring $newest"
echo " -> $cfg"
echo
# Newest is the right default — undo the last edit — but if you backed up
# *after* editing, the state you want is older. Show the rest so a no-op
# restore is obviously a no-op rather than a mystery.
count=$(ls "$cfg".*.bak 2>/dev/null | wc -l)
if [ "$count" -gt 1 ]; then
echo "$count backups exist, newest first:"
ls -t "$cfg".*.bak | sed 's/^/ /'
echo " (restoring the newest; copy another by hand to pick an older one)"
echo
fi
if [ -r "$cfg" ]; then
echo "what changes:"
if diff "$cfg" "$newest" > /tmp/mem.diff 2>&1 && [ ! -s /tmp/mem.diff ]; then
echo " nothing — that backup is identical to the current config"
else
sed 's/^/ /' /tmp/mem.diff
fi
rm -f /tmp/mem.diff
echo
fi
printf "proceed? [y/N] "
read -r reply
case "$reply" in
y|Y|yes|Yes) ;;
*) echo "left alone"; return 0 ;;
esac
cp "$newest" "$cfg"
echo "restored. From a WINDOWS terminal: wsl --shutdown"
}
# ── push ───────────────────────────────────────────────────────────────────
STATE=""
CHILD=""
cleanup() {
if [ -n "$CHILD" ] && kill -0 "$CHILD" 2>/dev/null; then
kill -KILL "$CHILD" 2>/dev/null || true
wait "$CHILD" 2>/dev/null || true
fi
[ -n "$STATE" ] && rm -f "$STATE"
return 0
}
# The child allocates and stops itself; the parent only watches. That split is
# the point: under --to-oom the allocating process is expected to be killed, and
# something has to survive to say how far it got.
allocator() {
# Raise our own OOM score to the maximum so the kernel picks THIS process
# first. Raising needs no privilege (only lowering does). Without it, the
# kernel is free to choose your shell, your ssh session or dockerd — on a
# box you are still using, that is not an acceptable coin toss.
echo 1000 > "/proc/$BASHPID/oom_score_adj" 2>/dev/null || true
local arr=() held=0 i=0 rss swapped avail first_swap=0
local bytes=$((STEP_MB * 1024 * 1024))
local swap_used_start
swap_used_start=$(( $(mb SwapTotal) - $(mb SwapFree) ))
while :; do
# Written STRAIGHT INTO the array element. The obvious spelling —
# build one chunk and `arr+=("$chunk")` — costs three copies per step,
# not one: the template stays resident, expanding "$chunk" makes a
# temporary word, and the append makes the element. A 128 MB step then
# needs 384 MB transiently, and on a small box it is killed on the
# first append while reporting a third of the true ceiling.
#
# printf -v into a subscript also means every page is written, so it is
# resident rather than merely promised — the only kind of allocation
# that measures anything under heuristic overcommit.
printf -v "arr[$i]" '%*s' "$bytes" ''
i=$((i + 1)); held=$((held + STEP_MB))
rss=$(awk '/^VmRSS:/{print int($2/1024)}' "/proc/$BASHPID/status" 2>/dev/null || echo 0)
avail=$(headroom_mb)
swapped=$(( $(mb SwapTotal) - $(mb SwapFree) - swap_used_start ))
[ "$swapped" -lt 0 ] && swapped=0
printf '%8s MB held rss %7s MB headroom %7s MB swap +%s MB\n' \
"$held" "$rss" "$avail" "$swapped"
printf '%s %s %s %s\n' "$held" "$rss" "$avail" "$swapped" >> "$STATE"
# Worth calling out separately from the ceiling: this is where the box
# stops being fast and starts being unusable, which for a scheduler is
# a different and earlier problem than being killed.
if [ "$swapped" -gt 0 ] && [ "$first_swap" -eq 0 ]; then
first_swap=$held
echo " - first swap page at ${held} MB — past here it works but crawls"
echo "swapat $held" >> "$STATE"
fi
if [ -n "$TO_MB" ] && [ "$held" -ge "$TO_MB" ]; then
echo "stop reached-the-cap" >> "$STATE"; return 0
fi
if [ "$TO_OOM" = no ] && [ "$avail" -lt "$FLOOR_MB" ]; then
echo "stop floor" >> "$STATE"; return 0
fi
done
}
push() {
local total ceiling rc=0 last held rss swapat stop
total=$(mb MemTotal)
ceiling=$(effective_ceiling_mb)
# A step is worth about a sixty-fourth of the ceiling: enough resolution to
# find the edge, few enough lines to read, and small enough that the
# transient cost of one allocation never dominates a small box. A fixed
# size cannot do all three — 128 MB is fine on 16 GB and absurd on 512 MB.
if [ "$STEP_EXPLICIT" = no ]; then
STEP_MB=$(( ceiling / 64 ))
[ "$STEP_MB" -lt 4 ] && STEP_MB=4
[ "$STEP_MB" -gt 256 ] && STEP_MB=256
fi
# Stop with a cushion rather than riding it to the kill. How big a cushion
# depends on what it is protecting. Under a cgroup cap, running out kills
# only this container's own processes, so it need cover no more than the
# shell that prints the result — and a 512 MB cushion on a 1 GB box would
# halve the answer. On a host there is everything else to protect, and the
# OOM killer does not promise to pick the process that caused the problem.
if [ -n "$(cgroup_cap_mb)" ]; then FLOOR_MB=64; else FLOOR_MB=512; fi
[ $(( ceiling / 20 )) -gt "$FLOOR_MB" ] && FLOOR_MB=$(( ceiling / 20 ))
STATE=$(mktemp "${TMPDIR:-/tmp}/rigmini.XXXXXX")
trap cleanup EXIT
# INT kills the child and lets the summary below print anyway, so an
# impatient Ctrl-C still tells you how far it got — and, more importantly,
# still gives the memory back.
trap 'echo; echo " interrupted"; echo "stop interrupted" >> "$STATE"; [ -n "$CHILD" ] && kill -KILL "$CHILD" 2>/dev/null || true' INT
echo "push"
echo " step ${STEP_MB} MB per allocation, every page touched"
echo " ceiling ${ceiling} MB claimed"
if [ -n "$TO_MB" ]; then
echo " stopping at ${TO_MB} MB (--to)"
elif [ "$TO_OOM" = yes ]; then
echo " ! stopping only when the kernel stops it (--to-oom)"
echo " the allocating child is marked as the preferred OOM victim,"
echo " but nothing about an OOM kill is entirely polite. Not on a box"
echo " running anything you mind losing."
else
echo " stopping when headroom drops below ${FLOOR_MB} MB"
fi
echo
allocator &
CHILD=$!
wait "$CHILD" || rc=$?
CHILD=""
trap - INT
last=$(grep -E '^[0-9]' "$STATE" 2>/dev/null | tail -1 || true)
held=$(echo "$last" | awk '{print $1}')
rss=$(echo "$last" | awk '{print $2}')
swapat=$(awk '/^swapat/{print $2}' "$STATE" 2>/dev/null | head -1 || true)
stop=$(awk '/^stop/{print $2}' "$STATE" 2>/dev/null | head -1 || true)
echo
if [ -z "$held" ]; then
echo " ! nothing was allocated. Even one ${STEP_MB} MB chunk failed —"
echo " try a smaller --step, or check ulimit -v in 'status'."
return 1
fi
echo " reached ${rss:-$held} MB resident"
[ -n "$swapat" ] && echo " swapping from ${swapat} MB"
case "$stop" in
reached-the-cap)
echo " outcome stopped at the --to cap, not at a limit."
echo " The box held ${TO_MB} MB without complaint; there is more." ;;
floor)
echo " outcome stopped with a cushion intact, by choice."
echo " The real ceiling is higher — --to-oom finds it, at the"
echo " cost of an actual OOM kill." ;;
interrupted)
echo " outcome interrupted at ${rss:-$held} MB — where you stopped it,"
echo " not where the box did." ;;
*)
# No stop line means the child did not decide to stop: it was ended.
if [ "$rc" -ge 128 ]; then
echo " outcome the child was killed (signal $((rc - 128))) at ${rss:-$held} MB."
elif [ "$rc" -ne 0 ]; then
echo " outcome the allocation failed at ${rss:-$held} MB (exit ${rc})."
echo " bash could not get the next chunk — an honest malloc"
echo " failure rather than a kill. That is the strict-overcommit"
echo " or ulimit path."
else
echo " outcome ended at ${rss:-$held} MB."
fi
local ev
ev=$(dmesg 2>/dev/null | tail -80 | grep -iE 'oom-kill|killed process' | tail -1 || true)
if [ -n "$ev" ]; then
echo " kernel ${ev#*] }"
else
echo " - dmesg is unreadable here (dmesg_restrict, or no privilege),"
echo " so the kill cannot be confirmed from this side. The number stands."
fi ;;
esac
# The gap between the claim and the measurement is the finding — but only
# when the BOX chose where to stop. An empty $stop means the child was ended
# rather than deciding to end; anything else (--to, the floor) is a stop we
# asked for, and flagging those as short of the ceiling would put a warning
# on every deliberately small run.
local got="${rss:-$held}"
echo
if [ -z "$stop" ] && [ "$got" -lt $(( ceiling * 70 / 100 )) ]; then
echo " ! claimed ${ceiling} MB, gave up ${got} MB — under 70% of it."
echo " Something is taking the difference. 'status' names the candidates:"
echo " a cgroup cap, ulimit -v, or memory already resident."
fi
return 0
}
# ── all ────────────────────────────────────────────────────────────────────
all() {
status
echo
echo "────────────────────────────────────────────────────────────"
echo
push
local got budget_mb ceiling
load_config
if [ -n "$BUDGET_GB" ]; then
budget_mb=$(( BUDGET_GB * 1024 ))
else
budget_mb=$(( NODES * NODE_MB ))
fi
ceiling=$(effective_ceiling_mb)
got=$(grep -E '^[0-9]' "$STATE" 2>/dev/null | tail -1 | awk '{print $2}' || true)
[ -n "$got" ] || got=0
echo
echo "verdict"
if [ -n "$BUDGET_GB" ]; then
echo " budget ${budget_mb} MB (--budget)"
else
# rig's own figure for this profile: nodes times what one node costs.
# Addons carry no memory figure in rig yet, so this is the cluster alone
# and whatever you deploy comes on top. --budget once you know that too.
echo " budget ${budget_mb} MB — profile ${PROFILE_NAME}: ${NODES} node(s) x ${NODE_MB} MB,"
echo " the cluster alone; your workload comes on top (--budget GB)"
fi
echo " measured ${got} MB handed over"
if [ "$got" -ge "$budget_mb" ]; then
echo " fits, with $(( got - budget_mb )) MB spare."
if [ "$got" -lt $(( budget_mb * 130 / 100 )) ]; then
echo " - under 30% spare is thin once a workload runs on top: memory use"
echo " is spiky, and the spikes are what get killed."
fi
else
echo " ! short by $(( budget_mb - got )) MB."
if [ "$ceiling" -ge "$budget_mb" ]; then
echo " The box CLAIMS enough (${ceiling} MB) but did not deliver it."
echo " Free something, or read the caps section again."
else
echo " The box does not have it to give. A bigger machine, or a profile"
echo " with fewer nodes."
fi
fi
return 0
}
# ── main ───────────────────────────────────────────────────────────────────
parse_flags() {
while [ $# -gt 0 ]; do
case "$1" in
--to) TO_MB=$(( ${2:?--to needs a value in GB} * 1024 )); shift 2 ;;
--to-mb) TO_MB="${2:?--to-mb needs a value in MB}"; shift 2 ;;
--step) STEP_MB="${2:?--step needs a value in MB}"; STEP_EXPLICIT=yes; shift 2 ;;
--to-oom) TO_OOM=yes; shift ;;
--budget) BUDGET_GB="${2:?--budget needs a value in GB}"; BUDGET_EXPLICIT=yes; shift 2 ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
if [ "$TO_OOM" = yes ] && [ -n "$TO_MB" ]; then
echo "--to and --to-oom contradict each other: one stops early, the other" >&2
echo "refuses to stop at all. Pick one." >&2
exit 1
fi
return 0
}
require_linux
find_cgroup
cmd="${1:-status}"
[ $# -gt 0 ] && shift
case "$cmd" in
status) parse_flags "$@"; status ;;
push) parse_flags "$@"; push ;;
all) parse_flags "$@"; all ;;
backup) backup ;;
restore) restore ;;
*) echo "usage: $0 [status|push|all|backup|restore]" >&2
echo " push [--to GB] [--to-mb MB] [--step MB] [--to-oom]" >&2
echo " all [--budget GB]" >&2
exit 1 ;;
esac

View File

@@ -1,296 +1,411 @@
<!DOCTYPE html>
<html lang="en">
<!--
MercadoPago shunt — config UI.
Brought onto the theme: it used to carry ~30 hardcoded hexes, its own reset,
its own button and input rules, and zero var(--…), so it was the one page in
the tree that could not follow a theme at all. Everything structural now comes
from the baked parts; what is left below is this shunt's own.
A shunt serves this on its own port and cannot fetch soleprint's /theme.css,
so the theme and the parts are baked in between markers by
common/theme/bake.py. Run `make theme bake` after changing a class.
Only `base` and `panel` are baked here — this page has no split pane, so it
carries no split.css and no split.js. That is the mechanism doing its job.
-->
<html lang="en" data-theme="soleprint">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>MercadoPago API (MOCK) - Configuration</title>
<!-- theme:here -->
<!-- theme:baked-defaults — generated by common/theme/bake.py; do not edit -->
<style>
* { box-sizing: border-box; margin: 0; padding: 0; }
body {
font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif;
background: #111827;
color: #e5e7eb;
padding: 20px;
}
.container { max-width: 1200px; margin: 0 auto; }
header {
background: #0071f2;
color: white;
padding: 20px;
border-radius: 8px;
margin-bottom: 24px;
}
h1 { font-size: 1.5rem; font-weight: 600; margin-bottom: 8px; }
.subtitle { opacity: 0.9; font-size: 0.875rem; }
:root {
--accent: #d4a574;
--font-mono: "JetBrains Mono", "Cascadia Mono", Consolas, monospace;
--font-size-base: 13px;
--font-size-sm: 11px;
--font-ui: Inter, "Segoe UI", system-ui, -apple-system, Arial, sans-serif;
--label-spacing: 0.04em;
--muted: #8888a0;
--panel-border: 1px solid #2e2e38;
--panel-header-height: 36px;
--panel-radius: 6px;
--space-1: 4px;
--space-2: 8px;
--space-3: 12px;
--space-4: 16px;
--space-6: 24px;
--status-error: #f06565;
--status-escalating: #f5a623;
--status-idle: #555568;
--status-live: #3ecf8e;
--status-processing: #4f9cf9;
--surface-0: #0d0d0f;
--surface-1: #16161a;
--surface-2: #1e1e24;
--surface-3: #2e2e38;
--text-dim: #555568;
--text-primary: #e8e8f0;
--text-secondary: #8888a0;
}
</style>
<!-- /theme:baked-defaults -->
<!-- theme:parts — generated by common/theme/bake.py; do not edit -->
<style>
/* part: base — common/theme/parts/base.css */
* {
box-sizing: border-box;
}
html,
body {
margin: 0;
}
body {
background: var(--surface-0);
color: var(--text-primary);
font-family: var(--font-ui);
font-size: var(--font-size-base);
}
/* Opt in on <html>, <body> and the top element when the page is an app shell
* that should fill the viewport. Left off, the page scrolls like a document —
* which is what most ad-hoc vein pages actually want. */
.fills {
height: 100%;
width: 100%;
}
button {
font-family: var(--font-ui);
font-size: var(--font-size-base);
color: var(--text-primary);
background: var(--surface-2);
border: var(--panel-border);
border-radius: var(--panel-radius);
padding: var(--space-1) var(--space-3);
cursor: pointer;
}
button:hover:not(:disabled) {
background: var(--surface-3);
}
button:disabled {
opacity: 0.5;
cursor: default;
}
input,
select,
textarea {
font-family: var(--font-ui);
font-size: var(--font-size-base);
color: var(--text-primary);
background: var(--surface-0);
border: var(--panel-border);
border-radius: 4px;
padding: var(--space-1) var(--space-2);
}
::-webkit-scrollbar {
width: 10px;
height: 10px;
}
::-webkit-scrollbar-thumb {
background: var(--surface-3);
border-radius: 5px;
}
::-webkit-scrollbar-track {
background: transparent;
}
/* part: panel — common/theme/parts/panel.css */
.panel {
position: relative;
background: var(--surface-1);
border: var(--panel-border);
border-radius: var(--panel-radius);
overflow: hidden;
display: flex;
flex-direction: column;
}
.panel-header {
display: flex;
align-items: center;
gap: var(--space-2);
height: var(--panel-header-height);
padding: 0 var(--space-3);
background: var(--surface-2);
border-bottom: var(--panel-border);
flex-shrink: 0;
}
.panel-title {
font-family: var(--font-ui);
font-size: var(--font-size-sm);
font-weight: 600;
color: var(--text-secondary);
text-transform: uppercase;
letter-spacing: 0.04em;
}
.panel-actions {
margin-left: auto;
display: flex;
align-items: center;
gap: var(--space-2);
}
.panel-status {
width: 8px;
height: 8px;
border-radius: 50%;
flex-shrink: 0;
background: var(--status-idle);
}
.panel-status.idle { background: var(--status-idle); }
.panel-status.live { background: var(--status-live); }
.panel-status.processing { background: var(--status-processing); }
.panel-status.error { background: var(--status-error); }
.panel-body {
flex: 1;
overflow: auto;
padding: var(--space-2);
min-height: 0;
}
</style>
<!-- /theme:parts -->
<style>
/* This page's own, and only its own. */
body { padding: var(--space-6); }
.container { max-width: 1100px; margin: 0 auto; }
header { margin-bottom: var(--space-6); }
h1 { margin: 0; font-size: 20px; font-family: var(--font-ui); }
.subtitle { color: var(--muted); font-size: 12px; margin-top: 4px; }
.mock-badge {
display: inline-block;
background: white;
color: #0071f2;
padding: 4px 12px;
border-radius: 4px;
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
margin-left: 12px;
}
.section {
background: #1f2937;
border-radius: 8px;
padding: 20px;
margin-bottom: 20px;
}
.section-header {
font-size: 1.1rem;
font-weight: 600;
margin-bottom: 16px;
color: #f9fafb;
}
.endpoint-list { display: flex; flex-direction: column; gap: 12px; }
.endpoint-card {
background: #374151;
border: 2px solid transparent;
border-radius: 6px;
padding: 16px;
cursor: pointer;
transition: all 0.2s;
}
.endpoint-card:hover { border-color: #0071f2; background: #4b5563; }
.endpoint-card.active { border-color: #0071f2; background: #4b5563; }
.endpoint-method {
display: inline-block;
padding: 2px 8px;
border-radius: 4px;
font-size: 0.75rem;
font-weight: 600;
margin-right: 8px;
}
.method-post { background: #10b981; color: white; }
.method-get { background: #3b82f6; color: white; }
.endpoint-path { font-family: monospace; font-size: 0.875rem; }
.endpoint-desc { font-size: 0.75rem; color: #9ca3af; margin-top: 6px; }
.form-group { margin-bottom: 16px; }
.form-label {
display: block;
font-size: 0.875rem;
font-weight: 500;
margin-bottom: 6px;
color: #f9fafb;
display: inline-block; margin-left: var(--space-2);
padding: 2px 8px; border-radius: 3px;
background: var(--status-escalating); color: var(--surface-0);
font-size: 10px; font-weight: 600; text-transform: uppercase;
letter-spacing: var(--label-spacing, .08em); vertical-align: middle;
}
.form-input, .form-textarea, .form-select {
width: 100%;
padding: 10px 12px;
background: #374151;
border: 1px solid #4b5563;
border-radius: 6px;
color: #e5e7eb;
font-size: 0.875rem;
}
.form-textarea { min-height: 200px; font-family: monospace; }
.form-input:focus, .form-textarea:focus, .form-select:focus {
outline: none;
border-color: #0071f2;
}
.btn {
padding: 10px 20px;
border: none;
border-radius: 6px;
font-size: 0.875rem;
font-weight: 600;
cursor: pointer;
transition: all 0.2s;
.panel { margin-bottom: var(--space-4); }
.lede { color: var(--muted); font-size: 12px; margin: 0 0 var(--space-3); }
.grid {
display: grid; gap: var(--space-2);
grid-template-columns: repeat(auto-fit, minmax(190px, 1fr));
}
.btn-primary {
background: #0071f2;
color: white;
/* Selectable cards — one treatment, used by both lists. */
.card {
padding: var(--space-3); cursor: pointer; text-align: left;
background: var(--surface-2); border: var(--panel-border);
border-radius: var(--panel-radius);
}
.btn-primary:hover { background: #005ac1; }
.btn-secondary {
background: #4b5563;
color: #e5e7eb;
margin-left: 8px;
.card:hover:not(:disabled) { border-color: var(--accent); background: var(--surface-3); }
.card.on { border-color: var(--accent); background: var(--surface-3); }
.card-name { font-weight: 600; color: var(--text-primary); margin-bottom: 4px; }
.card-desc { font-size: 11px; color: var(--muted); }
.verb {
font-size: 10px; font-weight: 600; padding: 1px 6px;
border-radius: 3px; border: 1px solid currentColor; margin-right: var(--space-2);
}
.btn-secondary:hover { background: #6b7280; }
.status-grid {
display: grid;
grid-template-columns: repeat(auto-fit, minmax(150px, 1fr));
gap: 12px;
.verb.POST { color: var(--status-live); }
.verb.GET { color: var(--status-processing); }
.path { font-family: var(--font-mono); font-size: 12px; }
.field { margin-bottom: var(--space-3); }
.field label {
display: block; margin-bottom: 4px; font-size: 11px;
color: var(--muted); text-transform: uppercase; letter-spacing: .04em;
}
.status-option {
background: #374151;
padding: 12px;
border-radius: 6px;
cursor: pointer;
border: 2px solid transparent;
.field input, .field textarea { width: 100%; font-family: var(--font-mono); font-size: 12px; }
.field textarea { min-height: 180px; resize: vertical; }
.actions { display: flex; gap: var(--space-2); }
.primary { background: var(--accent); color: var(--surface-0); border-color: transparent; }
.primary:hover:not(:disabled) { opacity: .9; background: var(--accent); }
.url {
font-family: var(--font-mono); font-size: 12px; user-select: all;
padding: var(--space-2) var(--space-3);
background: var(--surface-0); border: var(--panel-border);
border-radius: var(--panel-radius);
}
.status-option:hover { border-color: #0071f2; }
.status-option.selected { border-color: #0071f2; background: #4b5563; }
.status-name { font-weight: 600; color: #f9fafb; margin-bottom: 4px; }
.status-desc { font-size: 0.75rem; color: #9ca3af; }
.note { color: var(--text-dim); font-size: 11px; margin: var(--space-2) 0 0; }
</style>
</head>
<body>
<div class="container">
<header>
<h1>MercadoPago <span class="mock-badge">MOCK</span></h1>
<div class="subtitle">Configure mock payment responses and behavior</div>
</header>
<!-- Payment Status Configuration -->
<div class="section">
<div class="section-header">Default Payment Status</div>
<p style="color: #9ca3af; margin-bottom: 16px;">Choose what status new payments should return:</p>
<div class="status-grid">
<div class="status-option selected" onclick="selectStatus('approved')">
<div class="status-name">Approved</div>
<div class="status-desc">Payment successful</div>
</div>
<div class="status-option" onclick="selectStatus('rejected')">
<div class="status-name">Rejected</div>
<div class="status-desc">Payment failed</div>
</div>
<div class="status-option" onclick="selectStatus('pending')">
<div class="status-name">Pending</div>
<div class="status-desc">Awaiting confirmation</div>
</div>
<div class="status-option" onclick="selectStatus('in_process')">
<div class="status-name">In Process</div>
<div class="status-desc">Being processed</div>
</div>
</div>
<div class="container">
<header>
<h1>MercadoPago <span class="mock-badge">mock</span></h1>
<div class="subtitle">Configure mock payment responses and behaviour</div>
</header>
<div class="panel">
<div class="panel-header">
<span class="panel-title">Default payment status</span>
<span class="panel-status live"></span>
</div>
<div class="panel-body">
<p class="lede">What status new payments should return.</p>
<div class="grid" id="statuses"></div>
</div>
</div>
<!-- Endpoint Configuration -->
<div class="section">
<div class="section-header">Configure Endpoint Responses</div>
<div class="endpoint-list">
<div class="endpoint-card" onclick="selectEndpoint('POST', '/checkout/preferences', 'preference')">
<div>
<span class="endpoint-method method-post">POST</span>
<span class="endpoint-path">/checkout/preferences</span>
</div>
<div class="endpoint-desc">Create payment preference (Checkout Pro)</div>
</div>
<div class="endpoint-card" onclick="selectEndpoint('POST', '/v1/payments', 'payment')">
<div>
<span class="endpoint-method method-post">POST</span>
<span class="endpoint-path">/v1/payments</span>
</div>
<div class="endpoint-desc">Create payment (Checkout API)</div>
</div>
<div class="endpoint-card" onclick="selectEndpoint('GET', '/v1/payments/{id}', 'payment_get')">
<div>
<span class="endpoint-method method-get">GET</span>
<span class="endpoint-path">/v1/payments/{id}</span>
</div>
<div class="endpoint-desc">Get payment details</div>
</div>
<div class="endpoint-card" onclick="selectEndpoint('POST', '/oauth/token', 'oauth')">
<div>
<span class="endpoint-method method-post">POST</span>
<span class="endpoint-path">/oauth/token</span>
</div>
<div class="endpoint-desc">OAuth token exchange/refresh</div>
</div>
</div>
<div class="panel">
<div class="panel-header">
<span class="panel-title">Endpoint responses</span>
<span class="panel-actions"><span class="card-desc" id="count"></span></span>
</div>
<div class="panel-body">
<div class="grid" id="endpoints"></div>
</div>
</div>
<!-- Response Editor -->
<div class="section" id="responseEditor" style="display: none;">
<div class="section-header">Edit Response</div>
<div class="form-group">
<label class="form-label">Endpoint</label>
<input class="form-input" id="endpointDisplay" readonly>
<div class="panel" id="editor" hidden>
<div class="panel-header">
<span class="panel-title">Edit response</span>
<span class="panel-status processing"></span>
</div>
<div class="panel-body">
<div class="field">
<label for="endpointDisplay">Endpoint</label>
<input id="endpointDisplay" readonly>
</div>
<div class="form-group">
<label class="form-label">Mock Response (JSON)</label>
<textarea class="form-textarea" id="responseJson" placeholder='{"id": "123456", "status": "approved", "_mock": "MercadoPago"}'></textarea>
<div class="field">
<label for="responseJson">Mock response (JSON)</label>
<textarea id="responseJson"></textarea>
</div>
<div class="form-group">
<label class="form-label">HTTP Status Code</label>
<input type="number" class="form-input" id="statusCode" value="200">
<div class="field">
<label for="statusCode">HTTP status code</label>
<input type="number" id="statusCode" value="200">
</div>
<div class="form-group">
<label class="form-label">Delay (ms)</label>
<input type="number" class="form-input" id="delay" value="0">
<div class="field">
<label for="delay">Delay (ms)</label>
<input type="number" id="delay" value="0">
</div>
<div>
<button class="btn btn-primary" onclick="saveResponse()">Save Response</button>
<button class="btn btn-secondary" onclick="closeEditor()">Cancel</button>
<div class="actions">
<button class="primary" onclick="saveResponse()">Save response</button>
<button onclick="closeEditor()">Cancel</button>
</div>
<p class="note">Saving is not implemented yet — it was not implemented before this
page was rebrought onto the theme either, and pretending otherwise would be worse.</p>
</div>
</div>
<!-- Quick Test -->
<div class="section">
<div class="section-header">Quick Test</div>
<p style="color: #9ca3af; margin-bottom: 12px;">Test endpoint URL to hit for configured responses:</p>
<div class="form-input" style="background: #374151; user-select: all;">
http://localhost:8006/v1/payments
</div>
<div class="panel">
<div class="panel-header"><span class="panel-title">Quick test</span></div>
<div class="panel-body">
<p class="lede">Hit this URL to get the configured responses.</p>
<div class="url">http://localhost:8006/v1/payments</div>
</div>
</div>
</div>
<script>
let selectedEndpoint = null;
let selectedPaymentStatus = 'approved';
<script>
const STATUSES = [
{ key: 'approved', name: 'Approved', desc: 'Payment successful' },
{ key: 'rejected', name: 'Rejected', desc: 'Payment failed' },
{ key: 'pending', name: 'Pending', desc: 'Awaiting confirmation' },
{ key: 'in_process', name: 'In Process', desc: 'Being processed' },
];
function selectStatus(status) {
selectedPaymentStatus = status;
document.querySelectorAll('.status-option').forEach(opt => opt.classList.remove('selected'));
event.currentTarget.classList.add('selected');
}
const ENDPOINTS = [
{ verb: 'POST', path: '/checkout/preferences', type: 'preference', desc: 'Create payment preference (Checkout Pro)' },
{ verb: 'POST', path: '/v1/payments', type: 'payment', desc: 'Create payment (Checkout API)' },
{ verb: 'GET', path: '/v1/payments/{id}', type: 'payment_get', desc: 'Get payment details' },
{ verb: 'POST', path: '/oauth/token', type: 'oauth', desc: 'OAuth token exchange/refresh' },
];
function selectEndpoint(method, path, type) {
selectedEndpoint = {method, path, type};
document.querySelectorAll('.endpoint-card').forEach(c => c.classList.remove('active'));
event.currentTarget.classList.add('active');
document.getElementById('responseEditor').style.display = 'block';
document.getElementById('endpointDisplay').value = `${method} ${path}`;
document.getElementById('responseJson').value = getDefaultResponse(type);
}
let paymentStatus = 'approved';
let selected = null;
function getDefaultResponse(type) {
const defaults = {
preference: JSON.stringify({
"id": "123456-pref-id",
"init_point": "https://www.mercadopago.com.ar/checkout/v1/redirect?pref_id=123456",
"sandbox_init_point": "https://sandbox.mercadopago.com.ar/checkout/v1/redirect?pref_id=123456",
"_mock": "MercadoPago"
}, null, 2),
payment: JSON.stringify({
"id": 123456,
"status": selectedPaymentStatus,
"status_detail": selectedPaymentStatus === 'approved' ? 'accredited' : 'cc_rejected_other_reason',
"transaction_amount": 1500,
"currency_id": "ARS",
"_mock": "MercadoPago"
}, null, 2),
payment_get: JSON.stringify({
"id": 123456,
"status": "approved",
"status_detail": "accredited",
"transaction_amount": 1500,
"_mock": "MercadoPago"
}, null, 2),
oauth: JSON.stringify({
"access_token": "APP_USR-123456-mock-token",
"token_type": "Bearer",
"expires_in": 15552000,
"refresh_token": "TG-123456-mock-refresh",
"_mock": "MercadoPago"
}, null, 2)
};
return defaults[type] || '{}';
}
const $ = (id) => document.getElementById(id);
function saveResponse() {
alert('Mock response saved (feature pending implementation)');
}
function pick(container, el) {
container.querySelectorAll('.card').forEach((c) => c.classList.remove('on'));
el.classList.add('on');
}
function closeEditor() {
document.getElementById('responseEditor').style.display = 'none';
selectedEndpoint = null;
document.querySelectorAll('.endpoint-card').forEach(c => c.classList.remove('active'));
}
</script>
STATUSES.forEach((s, i) => {
const b = document.createElement('button');
b.className = 'card' + (i === 0 ? ' on' : '');
b.innerHTML = '<div class="card-name"></div><div class="card-desc"></div>';
b.firstChild.textContent = s.name;
b.lastChild.textContent = s.desc;
b.onclick = () => { paymentStatus = s.key; pick($('statuses'), b); };
$('statuses').appendChild(b);
});
ENDPOINTS.forEach((e) => {
const b = document.createElement('button');
b.className = 'card';
b.innerHTML = '<div><span class="verb ' + e.verb + '">' + e.verb + '</span>' +
'<span class="path"></span></div><div class="card-desc"></div>';
b.querySelector('.path').textContent = e.path;
b.lastChild.textContent = e.desc;
b.onclick = () => {
selected = e;
pick($('endpoints'), b);
$('editor').hidden = false;
$('endpointDisplay').value = e.verb + ' ' + e.path;
$('responseJson').value = defaultResponse(e.type);
};
$('endpoints').appendChild(b);
});
$('count').textContent = ENDPOINTS.length + ' endpoints';
function defaultResponse(type) {
const bodies = {
preference: {
id: '123456-pref-id',
init_point: 'https://www.mercadopago.com.ar/checkout/v1/redirect?pref_id=123456',
sandbox_init_point: 'https://sandbox.mercadopago.com.ar/checkout/v1/redirect?pref_id=123456',
_mock: 'MercadoPago',
},
payment: {
id: 123456,
status: paymentStatus,
status_detail: paymentStatus === 'approved' ? 'accredited' : 'cc_rejected_other_reason',
transaction_amount: 1500,
currency_id: 'ARS',
_mock: 'MercadoPago',
},
payment_get: {
id: 123456, status: 'approved', status_detail: 'accredited',
transaction_amount: 1500, _mock: 'MercadoPago',
},
oauth: {
access_token: 'APP_USR-123456-mock-token', token_type: 'Bearer',
expires_in: 15552000, refresh_token: 'TG-123456-mock-refresh', _mock: 'MercadoPago',
},
};
return JSON.stringify(bodies[type] || {}, null, 2);
}
function saveResponse() {
alert('Mock response saved (feature pending implementation)');
}
function closeEditor() {
$('editor').hidden = true;
selected = null;
document.querySelectorAll('#endpoints .card').forEach((c) => c.classList.remove('on'));
}
</script>
</body>
</html>

View File

@@ -2,13 +2,35 @@
Jira Vein - FastAPI app.
"""
from pathlib import Path
from fastapi import FastAPI
from fastapi.responses import FileResponse, JSONResponse
from .api.routes import router
from .core.config import settings
app = FastAPI(title="Jira Vein", version="0.1.0")
app.include_router(router)
UI = Path(__file__).parent / "ui" / "index.html"
@app.get("/ui", include_in_schema=False)
def ui():
"""The vein's ad-hoc interface.
One file, served as-is. Its theme and its parts are baked in by
common/theme/bake.py, so this is a plain FileResponse rather than a
template: there is nothing left to fill in at request time, and the same
bytes open from a double-click when nothing is serving them.
"""
if not UI.is_file():
return JSONResponse(
{"error": "no ui/index.html", "hint": "run `make theme bake` from the spr root"},
status_code=404,
)
return FileResponse(UI, media_type="text/html")
if __name__ == "__main__":
import uvicorn

View File

@@ -0,0 +1,496 @@
<!DOCTYPE html>
<!--
The jira vein's ad-hoc interface — the `ui/` slot veins/__init__.py has
declared since the beginning and no vein had filled.
It does the vein page job, which rig-ui states twice in its own comments after
rebuilding this look by hand: name what the thing exposes, and show what comes
back. Tool chrome and output are styled apart on purpose — that separation is
what tells you whether you are reading the tool or its result.
THREE RULES, all consequences of "a vein serves this on its own port, and it
must also open from a double-clicked file":
no /theme.css an absolute path assumes a server at the root
no build step no npm, no bundler, no Vue
no webfont a blocked stylesheet is a stall, not a fallback
So the theme and the parts are BAKED IN, between markers, by
common/theme/bake.py. Everything outside those markers is this page's, to edit
freely. Run `make theme bake` after changing a class; `make theme check` says
when a copy has gone stale.
The parts are chosen by what the markup uses — panel and split here, nothing
else. That is the whole point: this page carries no table, no log view, no
uplot, no vue-flow, because it uses none of them.
-->
<html lang="en" data-theme="soleprint">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>jira — vein</title>
<!-- theme:here -->
<!-- theme:baked-defaults — generated by common/theme/bake.py; do not edit -->
<style>
:root {
--font-mono: "JetBrains Mono", "Cascadia Mono", Consolas, monospace;
--font-size-base: 13px;
--font-size-sm: 11px;
--font-ui: Inter, "Segoe UI", system-ui, -apple-system, Arial, sans-serif;
--muted: #8888a0;
--panel-border: 1px solid #2e2e38;
--panel-header-height: 36px;
--panel-radius: 6px;
--space-1: 4px;
--space-2: 8px;
--space-3: 12px;
--space-4: 16px;
--status-error: #f06565;
--status-idle: #555568;
--status-live: #3ecf8e;
--status-processing: #4f9cf9;
--surface-0: #0d0d0f;
--surface-1: #16161a;
--surface-2: #1e1e24;
--surface-3: #2e2e38;
--text-dim: #555568;
--text-primary: #e8e8f0;
--text-secondary: #8888a0;
}
</style>
<!-- /theme:baked-defaults -->
<!-- theme:parts — generated by common/theme/bake.py; do not edit -->
<style>
/* part: base — common/theme/parts/base.css */
* {
box-sizing: border-box;
}
html,
body {
margin: 0;
}
body {
background: var(--surface-0);
color: var(--text-primary);
font-family: var(--font-ui);
font-size: var(--font-size-base);
}
/* Opt in on <html>, <body> and the top element when the page is an app shell
* that should fill the viewport. Left off, the page scrolls like a document —
* which is what most ad-hoc vein pages actually want. */
.fills {
height: 100%;
width: 100%;
}
button {
font-family: var(--font-ui);
font-size: var(--font-size-base);
color: var(--text-primary);
background: var(--surface-2);
border: var(--panel-border);
border-radius: var(--panel-radius);
padding: var(--space-1) var(--space-3);
cursor: pointer;
}
button:hover:not(:disabled) {
background: var(--surface-3);
}
button:disabled {
opacity: 0.5;
cursor: default;
}
input,
select,
textarea {
font-family: var(--font-ui);
font-size: var(--font-size-base);
color: var(--text-primary);
background: var(--surface-0);
border: var(--panel-border);
border-radius: 4px;
padding: var(--space-1) var(--space-2);
}
::-webkit-scrollbar {
width: 10px;
height: 10px;
}
::-webkit-scrollbar-thumb {
background: var(--surface-3);
border-radius: 5px;
}
::-webkit-scrollbar-track {
background: transparent;
}
/* part: panel — common/theme/parts/panel.css */
.panel {
position: relative;
background: var(--surface-1);
border: var(--panel-border);
border-radius: var(--panel-radius);
overflow: hidden;
display: flex;
flex-direction: column;
}
.panel-header {
display: flex;
align-items: center;
gap: var(--space-2);
height: var(--panel-header-height);
padding: 0 var(--space-3);
background: var(--surface-2);
border-bottom: var(--panel-border);
flex-shrink: 0;
}
.panel-title {
font-family: var(--font-ui);
font-size: var(--font-size-sm);
font-weight: 600;
color: var(--text-secondary);
text-transform: uppercase;
letter-spacing: 0.04em;
}
.panel-actions {
margin-left: auto;
display: flex;
align-items: center;
gap: var(--space-2);
}
.panel-status {
width: 8px;
height: 8px;
border-radius: 50%;
flex-shrink: 0;
background: var(--status-idle);
}
.panel-status.idle { background: var(--status-idle); }
.panel-status.live { background: var(--status-live); }
.panel-status.processing { background: var(--status-processing); }
.panel-status.error { background: var(--status-error); }
.panel-body {
flex: 1;
overflow: auto;
padding: var(--space-2);
min-height: 0;
}
/* part: split — common/theme/parts/split.css */
.split-pane {
display: flex;
width: 100%;
height: 100%;
min-height: 0;
min-width: 0;
overflow: hidden;
}
.split-pane.horizontal {
flex-direction: row;
}
.split-pane.vertical {
flex-direction: column;
}
.split-first,
.split-second {
min-height: 0;
min-width: 0;
overflow: hidden;
flex: 1;
}
/* Children fill their pane. */
.split-first > *,
.split-second > * {
width: 100%;
height: 100%;
}
.split-divider {
flex-shrink: 0;
background: transparent;
transition: background 0.15s;
touch-action: none;
z-index: 10;
}
.split-divider:hover,
.split-divider.dragging {
background: var(--text-dim);
}
.split-pane.horizontal > .split-divider {
width: 4px;
cursor: col-resize;
margin: 0 -2px;
}
.split-pane.vertical > .split-divider {
height: 4px;
cursor: row-resize;
margin: -2px 0;
}
</style>
<script>
/* part: split — common/theme/parts/split.js */
(function () {
'use strict'
function setup(root) {
var divider = root.querySelector(':scope > .split-divider')
if (!divider) return // no divider: a fixed split, deliberately
var first = root.querySelector(':scope > .split-first')
var second = root.querySelector(':scope > .split-second')
if (!first || !second) return
var horizontal = !root.classList.contains('vertical')
var mode = root.dataset.mode === 'px' ? 'px' : 'ratio'
var anchor = root.dataset.anchor === 'second' ? 'second' : 'first'
var size = parseFloat(root.dataset.size)
if (isNaN(size)) size = 1
var min = parseFloat(root.dataset.min)
if (isNaN(min)) min = mode === 'px' ? 0 : 0.1
var max = parseFloat(root.dataset.max)
if (isNaN(max)) max = mode === 'px' ? Infinity : 10
var sized = anchor === 'second' ? second : first
var flexed = anchor === 'second' ? first : second
var dragging = false
var startPos = 0
function apply() {
flexed.style.flex = '1'
if (mode === 'px') {
sized.style.flex = '0 0 auto'
sized.style[horizontal ? 'width' : 'height'] = size + 'px'
} else {
sized.style.flex = String(size)
}
}
divider.addEventListener('pointerdown', function (e) {
dragging = true
startPos = horizontal ? e.clientX : e.clientY
divider.classList.add('dragging')
divider.setPointerCapture(e.pointerId)
})
divider.addEventListener('pointermove', function (e) {
if (!dragging) return
var pos = horizontal ? e.clientX : e.clientY
var delta = pos - startPos
startPos = pos
// Dragging right/down grows the first pane. When the SECOND pane is the
// anchored one, that same gesture must shrink it, so invert.
if (anchor === 'second') delta = -delta
// Ratio mode is unitless, so pixels are scaled into it. The two constants
// are the SFC's, kept rather than re-derived: they are what the existing
// panes were tuned against, and vertical drags cover less travel.
var step = mode === 'px' ? delta : delta * (horizontal ? 0.01 : 0.02)
size = Math.max(min, Math.min(max, size + step))
apply()
})
function end(e) {
if (!dragging) return
dragging = false
divider.classList.remove('dragging')
if (e && e.pointerId !== undefined && divider.hasPointerCapture(e.pointerId)) {
divider.releasePointerCapture(e.pointerId)
}
}
divider.addEventListener('pointerup', end)
divider.addEventListener('pointercancel', end)
apply()
}
function start() {
var panes = document.querySelectorAll('[data-split]')
for (var i = 0; i < panes.length; i++) setup(panes[i])
}
if (document.readyState === 'loading') {
document.addEventListener('DOMContentLoaded', start)
} else {
start()
}
})()
</script>
<!-- /theme:parts -->
<style>
/* This page's own, and only its own. */
body { padding: var(--space-4); }
.page { display: flex; flex-direction: column; gap: var(--space-3); height: calc(100vh - 2 * var(--space-4)); }
header { display: flex; align-items: baseline; gap: var(--space-3); flex-wrap: wrap; }
h1 { margin: 0; font-size: 20px; font-family: var(--font-ui); }
.sub { color: var(--muted); font-family: var(--font-mono); font-size: 12px; }
.split-pane { border: var(--panel-border); border-radius: var(--panel-radius); }
.routes { display: flex; flex-direction: column; gap: 2px; }
.route {
display: flex; align-items: center; gap: var(--space-2);
padding: 4px var(--space-2); border-radius: var(--panel-radius);
font-family: var(--font-mono); font-size: 12px;
background: none; border: 0; width: 100%; text-align: left;
}
.route:hover:not(:disabled) { background: var(--surface-2); }
.route.on { background: var(--surface-3); }
.verb {
font-size: 10px; font-weight: 600; padding: 1px 6px; border-radius: 3px;
border: 1px solid currentColor; flex-shrink: 0;
}
.verb.GET { color: var(--status-processing); }
.verb.POST { color: var(--status-live); }
.controls { display: flex; gap: var(--space-2); margin-bottom: var(--space-2); }
.controls input { flex: 1; font-family: var(--font-mono); font-size: 12px; }
/* OUTPUT — a raised block, monospace, selectable. It is a payload, not
furniture, and should not read like the tool that produced it. */
.out {
margin: 0; padding: var(--space-3);
background: var(--surface-0); border: var(--panel-border);
border-radius: var(--panel-radius);
font-family: var(--font-mono); font-size: 12px; line-height: 1.5;
white-space: pre-wrap; word-break: break-word;
user-select: text; min-height: 100%;
}
.out.err { color: var(--status-error); }
.hint { color: var(--text-dim); }
</style>
</head>
<body>
<div class="page">
<header>
<h1>jira</h1>
<span class="sub">vein &middot; stateless API connector</span>
<span class="sub hint" id="base"></span>
</header>
<div class="split-pane horizontal" data-split data-size="1" data-min="0.4" data-max="3" style="flex:1; min-height:0;">
<div class="split-first">
<div class="panel">
<div class="panel-header">
<span class="panel-title">Exposes</span>
<span class="panel-actions"><span class="sub" id="count"></span></span>
</div>
<div class="panel-body">
<div class="routes" id="routes"></div>
</div>
</div>
</div>
<div class="split-divider"></div>
<div class="split-second">
<div class="panel">
<div class="panel-header">
<span class="panel-title">Returns</span>
<span class="panel-actions"><button id="run" disabled>send</button></span>
<span class="panel-status idle" id="dot"></span>
</div>
<div class="panel-body">
<div class="controls">
<input id="arg" placeholder="pick a route" disabled>
</div>
<pre class="out hint" id="out">Pick a route on the left, then send.
Opened from a file rather than served? Nothing will answer — the routes are
still the contract, which is half of what this page is for.</pre>
</div>
</div>
</div>
</div>
</div>
<script>
/* The vein's surface, as the page understands it. Kept here rather than fetched
so the page still says what the vein exposes when nothing is serving it. */
const ROUTES = [
{ verb: 'GET', path: '/health', arg: null, hint: 'connection check' },
{ verb: 'GET', path: '/mine', arg: null, hint: 'tickets assigned to you' },
{ verb: 'GET', path: '/backlog', arg: null, hint: 'the backlog' },
{ verb: 'GET', path: '/sprint', arg: null, hint: 'the current sprint' },
{ verb: 'GET', path: '/ticket/{key}', arg: 'key', hint: 'one ticket, e.g. PROJ-123' },
{ verb: 'POST', path: '/search', arg: 'jql', hint: 'a JQL query' },
{ verb: 'GET', path: '/epic/{key}/status', arg: 'key', hint: 'epic processing status' },
];
const $ = (id) => document.getElementById(id);
let active = null;
$('base').textContent = location.protocol === 'file:' ? 'not served — file://' : location.origin;
$('count').textContent = ROUTES.length + ' routes';
ROUTES.forEach((r, i) => {
const b = document.createElement('button');
b.className = 'route';
b.innerHTML = '<span class="verb ' + r.verb + '">' + r.verb + '</span>' +
'<span>' + r.path + '</span>';
b.title = r.hint;
b.onclick = () => select(i, b);
$('routes').appendChild(b);
});
function select(i, el) {
active = ROUTES[i];
document.querySelectorAll('.route').forEach((n) => n.classList.remove('on'));
el.classList.add('on');
$('arg').disabled = !active.arg;
$('arg').placeholder = active.arg ? active.hint : 'no argument';
$('arg').value = '';
$('run').disabled = false;
}
function status(state) { $('dot').className = 'panel-status ' + state; }
$('run').onclick = async () => {
if (!active) return;
status('processing');
$('out').className = 'out';
$('out').textContent = '…';
let path = active.path, init = { method: active.verb };
const v = $('arg').value.trim();
if (active.arg === 'key') path = path.replace('{key}', encodeURIComponent(v));
if (active.arg === 'jql') {
init.headers = { 'Content-Type': 'application/json' };
init.body = JSON.stringify({ jql: v });
}
try {
const res = await fetch(path, init);
const text = await res.text();
let body = text;
try { body = JSON.stringify(JSON.parse(text), null, 2); } catch (_) {}
$('out').textContent = res.status + ' ' + res.statusText + '\n\n' + body;
status(res.ok ? 'live' : 'error');
if (!res.ok) $('out').className = 'out err';
} catch (e) {
$('out').className = 'out err';
$('out').textContent = String(e) +
(location.protocol === 'file:' ? '\n\nThis page is not being served.' : '');
status('error');
}
};
</script>
</body>
</html>

5
soleprint/atlas2/docgen/.gitignore vendored Normal file
View File

@@ -0,0 +1,5 @@
# Everything this makes.
out/
__pycache__/
*.pyc
.venv/

View File

@@ -0,0 +1,179 @@
# docgen — code to diagram, and to everything else the IR can feed.
#
# Derived from where this file sits, so the folder can be copied anywhere and
# renamed and still work. The logic lives in the Python, never here: every target
# is one line calling `python3 -m docgen <command>`.
#
# make sync create .venv with every optional group (uv)
# make book SRC=../station the whole operation, measured at both ends
# make run CONFIG=docgen.toml every book a run file lists
# make check prove docgen, on a tree it builds itself
# make check BOOK=out/book/x prove one book — its own level
# make ir SRC=../station extract -> out/ir.json (one step, on its own)
# make self docgen's book of itself, then check it
# make doctor what this machine has
#
# Every step target still works alone — that is the property the book spine
# exists to preserve, not to replace. The steps compose by hand too:
#
# python3 -m docgen extract python --root SRC -o ir.json
# python3 -m docgen view ir.json --overview -o view.json
# python3 -m docgen emit dot view.json -o graph.svg --theme dark
HERE := $(patsubst %/,%,$(dir $(abspath $(lastword $(MAKEFILE_LIST)))))
PKG := $(notdir $(HERE))
PARENT := $(patsubst %/,%,$(dir $(HERE)))
VENV_PY := $(HERE)/.venv/bin/python
# The synced environment when there is one (`make sync`), the system Python when
# there is not — so `make check` still works with nothing installed at all.
PY ?= $(if $(wildcard $(VENV_PY)),$(VENV_PY),python3)
CLI := PYTHONPATH=$(PARENT) $(PY) -m $(PKG)
OUT ?= $(HERE)/out
SRC ?=
SCHEMA ?=
OPENAPI ?=
HAR ?=
STYLE ?= lucid
THEME ?=
SCALE ?= 0.55
BOOK ?=
SLUG ?=
# NOT `LANG`: that is the shell's locale variable, so `?=` inherits
# en_US.UTF-8 from the environment and --reader rejects it.
READER ?= python
OVERLAY ?=
CONFIG ?= docgen.toml
ONLY ?=
CHECK ?=
comma := ,
THEME_ARG := $(if $(THEME),--theme $(THEME))
STYLE_ARGS := --style $(STYLE) $(THEME_ARG)
SLUG_ARG := $(if $(SLUG),--slug $(SLUG))
OVER_ARG := $(if $(OVERLAY),--overlay $(OVERLAY))
ONLY_ARGS := $(foreach n,$(subst $(comma), ,$(ONLY)),--only $(n))
.PHONY: help sync lock book run check ir db code view graph index site minimap explore docs self doctor clean
help: ## List every target
@echo "docgen — static analysis of a tree, and the artifacts that fall out of it"
@echo
@grep -E '^[a-z-]+:.*?## .*$$' $(MAKEFILE_LIST) \
| awk 'BEGIN{FS=":.*?## "}{printf " \033[1m%-10s\033[0m %s\n", $$1, $$2}'
@echo
@echo " SRC=/path/to/tree what to read OUT=/path where output goes"
@echo " SCHEMA=schema.json a database instead STYLE=lucid THEME=dark|lucid"
@echo " OPENAPI=spec.yaml an API document HAR=session.har a recording"
@echo " BOOK=/path where a book goes, and which book to check"
@echo " READER=python|code ast, or tree-sitter SLUG=name what to call the book"
@echo " OVERLAY=overlay.json hand-written notebook additions, re-applied every build"
@echo " CONFIG=docgen.toml a run file ONLY=a,b just these books"
@echo " CHECK=1 with run: each book's own level after building it"
@echo
@echo " Three levels of test, by what they assert about:"
@echo " make doctor the machine. Never fails."
@echo " make check docgen. Exits 1."
@echo " make check BOOK=<dir> that book. Exits 1."
sync: ## Create .venv with every optional group, from uv.lock
@command -v uv >/dev/null || { echo "Error: uv is not installed — docgen still runs on the system python3" >&2; exit 1; }
@cd $(HERE) && uv sync --all-groups
lock: ## Re-resolve uv.lock after editing pyproject.toml
@cd $(HERE) && uv lock
book: ## SRC (or SCHEMA/OPENAPI/HAR) -> one operation, measured at both ends
@test -n "$(SRC)$(SCHEMA)$(OPENAPI)$(HAR)" \
|| { echo "Error: set SRC=/path/to/tree (or SCHEMA=, OPENAPI=, HAR=)" >&2; exit 1; }
@$(CLI) book \
$(if $(SRC),--root "$(SRC)" --reader $(READER)) \
$(if $(SCHEMA),--schema "$(SCHEMA)") \
$(if $(OPENAPI),--openapi "$(OPENAPI)") \
$(if $(HAR),--har "$(HAR)") \
-o "$(if $(BOOK),$(BOOK),$(OUT)/book)" \
$(STYLE_ARGS) $(SLUG_ARG) $(OVER_ARG)
run: ## CONFIG (a run file) -> every book it lists; ONLY=a,b for some
@$(CLI) run "$(CONFIG)" $(ONLY_ARGS) $(if $(CHECK),--check)
check: ## Prove docgen (or one book, with BOOK=<dir>)
@$(CLI) check $(if $(BOOK),"$(BOOK)")
ir: ## Extract SRC into OUT/ir.json
@test -n "$(SRC)" || { echo "Error: set SRC=/path/to/tree" >&2; exit 1; }
@$(CLI) extract python --root "$(SRC)" -o $(OUT)/ir.json
@$(CLI) validate $(OUT)/ir.json
code: ## Extract C#/TypeScript from SRC (needs the `code` group)
@test -n "$(SRC)" || { echo "Error: set SRC=/path/to/tree" >&2; exit 1; }
@$(CLI) extract code --root "$(SRC)" -o $(OUT)/ir.json
@$(CLI) validate $(OUT)/ir.json
db: ## Extract a graphgen-compatible SCHEMA into OUT/ir.json
@test -n "$(SCHEMA)" || { echo "Error: set SCHEMA=/path/to/schema.json" >&2; exit 1; }
@$(CLI) extract db --schema "$(SCHEMA)" -o $(OUT)/ir.json
@$(CLI) validate $(OUT)/ir.json
view: ## OUT/ir.json -> OUT/view.json, the default view for its source type
@$(CLI) view $(OUT)/ir.json --overview -o $(OUT)/view.json
graph: view ## OUT/view.json -> whatever its structure asks for
@$(CLI) emit auto $(OUT)/view.json -o $(OUT) $(STYLE_ARGS)
index: ## OUT/ir.json -> OUT/index.md and OUT/sidebar.json
@$(CLI) emit index $(OUT)/ir.json -o $(OUT)/index.md
@$(CLI) emit index $(OUT)/ir.json -o $(OUT)/sidebar.json
site: view ## OUT/view.json -> a self-contained docs site in OUT/site
@$(CLI) emit site $(OUT)/view.json -o $(OUT)/site $(STYLE_ARGS)
@echo " open $(OUT)/site/index.html"
minimap: ## OUT/ir.json -> OUT/minimap.svg — what is where, read from the colours
@$(CLI) emit minimap $(OUT)/ir.json -o $(OUT)/minimap.svg $(STYLE_ARGS) --scale $(SCALE)
explore: ## OUT/ir.json -> OUT/explore/ — navigate on one side, explore on the other
@$(CLI) emit explore $(OUT)/ir.json -o $(OUT)/explore $(STYLE_ARGS) --scale $(SCALE)
@echo " open $(OUT)/explore/explore.html"
docs: ## Regenerate the figures in docs/ — docgen documented by docgen
@mkdir -p $(HERE)/docs/img
@$(CLI) extract python --root $(HERE) -o /tmp/$(PKG)-docs.json >/dev/null
@$(CLI) view /tmp/$(PKG)-docs.json --overview -o /tmp/$(PKG)-docs-view.json >/dev/null
@$(CLI) emit dot /tmp/$(PKG)-docs-view.json -o $(HERE)/docs/img/architecture.svg -q
@$(CLI) emit minimap /tmp/$(PKG)-docs.json -o $(HERE)/docs/img/minimap.svg --scale 0.5 --width 860
@$(CLI) emit erd $(OUT)/ir.json -o $(HERE)/docs/img/erd.svg 2>/dev/null \
|| echo " (erd figure kept — needs a schema IR at $(OUT)/ir.json to refresh)"
@PYTHONPATH=$(PARENT) $(PY) -c "from $(PKG).emitters.site import VIEWER, _slots, _fill; \
from $(PKG).style import Style; import pathlib; \
pathlib.Path('$(HERE)/docs/viewer.html').write_text( \
_fill(VIEWER.replace('__TITLE__', 'docgen docs'), _slots(Style.load('lucid'))))"
@echo " open $(HERE)/docs/index.html"
self: ## docgen's book of the widest tree it can see, then check it
@$(eval SELF_SRC := $(shell PYTHONPATH=$(PARENT) $(PY) -c "from $(PKG) import reference; \
r = reference.root(); print(r if r else '$(HERE)')"))
@echo " self-hosting on $(SELF_SRC)"
@$(MAKE) --no-print-directory book SRC=$(SELF_SRC) SLUG=self \
BOOK=$(OUT)/book/self OUT=$(OUT)
@echo
@$(MAKE) --no-print-directory check BOOK=$(OUT)/book/self
doctor: ## Report whether this machine can run it
@printf 'python : %s ' '$(PY)'; $(PY) --version 2>&1 || echo MISSING
@printf 'uv : '; if ! command -v uv >/dev/null; then echo 'absent — fine; docgen runs on the system python3'; \
elif [ -x $(VENV_PY) ]; then echo "$$(uv --version), .venv synced"; \
else echo "$$(uv --version), .venv not synced — make sync for the optional groups"; fi
@printf 'dot : '; (dot -V 2>&1) || echo 'MISSING — sudo apt install graphviz (only to render)'
@printf 'tree-sit : '; $(PY) -c 'import tree_sitter, tree_sitter_c_sharp, tree_sitter_typescript; print("ok — C# and TypeScript available")' 2>/dev/null || echo 'absent — Python only. make sync, or the `code` group'
@printf 'lxml : '; $(PY) -c 'import lxml; print("ok — theme harvesting available")' 2>/dev/null || echo 'absent — only used to harvest a theme'
@printf 'yaml : '; $(PY) -c 'import yaml; print("ok — needed only to read OpenAPI")' 2>/dev/null || echo 'absent — only used by the OpenAPI reader'
@printf 'networkx : '; $(PY) -c 'import networkx; print(networkx.__version__ + " — for lab/ experiments")' 2>/dev/null || echo 'absent — only used in lab/'
@printf 'reference: '; PYTHONPATH=$(PARENT) $(PY) -c "from $(PKG) import reference; print(reference.describe())"
@printf 'package : %s (from %s)\n' '$(PKG)' '$(PARENT)'
@printf 'styles : '; PYTHONPATH=$(PARENT) $(PY) -c "from $(PKG).style import Style; print(', '.join(Style.available()))"
@PYTHONPATH=$(PARENT) $(PY) -c "import $(PKG).ir, $(PKG).emitters.dot, $(PKG).ops, $(PKG).cli" >/dev/null 2>&1 \
&& echo 'import : ok' || echo 'import : FAILED — is the folder intact?'
clean: ## Delete OUT. Nothing else is ever written to
@rm -rf "$(OUT)" && echo "Removed $(OUT)"

View File

@@ -0,0 +1,276 @@
# docgen
Static analysis of a tree, and the artifacts that fall out of it.
The point is not the diagram. The point is the format in the middle — diagrams
are one consumer of it, and not the one that reaches the most people.
```
extractors/ → graph IR (JSON) → emitters/
(per source type) (one schema) (per output target)
style/*.json
(consumed by emitters only)
```
```bash
make sync # optional: .venv with every group, from uv.lock
make check # prove it, on a tree it builds itself
make book SRC=/path/to/repo # one whole operation, measured at both ends
make run CONFIG=docgen.toml # every book a run file lists
make help
```
The full documentation is `docs/index.html` — open it in a browser.
## Layout
```
cli/ the command line — every command, and nothing else
book/ an operation: larder measure, steps, web output, run files
extractors/ source -> IR, one per source type
ir/ the contract: schema.json, the dataclasses, the validator
ops/ IR -> a smaller IR
emitters/ IR -> an artifact
notebook/ the notebook spec, before it is an .ipynb
style/ slots, themes, and harvesting a theme from real diagrams
fixtures/ docgen's own test inputs
lab/ sanctioned experiments; nothing imports it
reference.py the one seam to the repo above, for reading OpenAPI
```
Library packages hold **no command-line code** — no argparse, no `__main__.py`,
no `cli_dot.py` beside `dot.py`. It all lives in `cli/`, behind one entry point,
and the selftest asserts it stays there.
`pyproject.toml` declares no required dependency: the structural path is the
stdlib, Python 3.11+. The optional groups (`code`, `openapi`, `harvest`, `lab`)
are pinned in `uv.lock`; `[tool.uv] package = false`, as dataconvert does, since
this is a folder run in place rather than something to install.
## Run files
```toml
# docgen.toml — beside the project it describes; paths relative to this file
[defaults]
out = "out/book"
[[book]]
name = "station"
root = "../soleprint/station"
[[book]]
name = "shop"
schema = "schemas/shop.json"
```
`make run CONFIG=docgen.toml [ONLY=station] [CHECK=1]`. One failing book never
stops the others; unknown keys are refused; a rebuild replaces the previous
build's outputs and never touches a hand-written `checks.py`.
`docgen.example.toml` runs against the shipped fixtures.
Or as three composable commands, which is what the Makefile is wrapping — run
from the directory above this one:
```bash
python3 -m docgen extract python --root SRC -o ir.json
python3 -m docgen view ir.json --drop-stdlib -o view.json
python3 -m docgen emit dot view.json -o graph.svg --theme dark
```
## DOT collapses three concerns; this separates them
| concern | question | owner |
|---|---|---|
| **structure** | what the graph *is* | `ir/schema.json` — versioned, golden-tested |
| **meaning** | what things *mean visually* | `style/*.json`, keyed on `kind` |
| **placement** | where things *go* | Graphviz defaults. Phase two |
An extractor has never heard of SVG, colours or layout. An emitter has never
heard of Python, `ast` or SQL. **The IR carries no visual information** — if a
field would change between light and dark theme, it does not belong in it.
`shape="cylinder"` is not a field; it is `kind="datastore"` plus a style rule,
which is what lets the same IR render in a theme that has no cylinders.
The selftest asserts all three of those, because they are the design rather than
a nicety and they are exactly what erodes first.
## The IR
```json
{
"meta": { "source": "python", "root": "app/", "schema_version": "1" },
"nodes": [ { "id": "app.models.User", "kind": "class", "label": "User",
"parent": "app.models",
"attrs": { "file": "app/models.py", "line": 12 } } ],
"edges": [ { "source": "app.models.User", "target": "app.db.Base",
"kind": "inherits", "attrs": {} } ]
}
```
- **`id`** is fully qualified and **stable across runs**. That is what makes two
graphs from two commits diffable.
- **`kind`** is the hinge, and the only field style and layout may key on.
- **`parent`** is containment. Relationships are edges.
- **`attrs`** is an open bag; `file`/`line` let a UI link a box to a line.
Stdlib dataclasses, not Pydantic. A format that needs a library installed to be
opened is not a format, it is an API. `ir/validate.py` is the check at the
boundary, and it reads the field lists out of `schema.json` so the two cannot
drift.
```bash
python3 -m docgen validate ir.json
```
It catches what a schema cannot: an edge naming a node that does not exist, a
containment cycle, duplicate ids, and a visual field smuggled into `attrs`.
## Extraction is deterministic
**No LLM in the structural path.** A diagram from an AST cannot be out of date
with the code; one from a model's reading of the code is wrong the moment the
model has a bad day, which is the problem this exists to fix.
`ast` resolves nothing on its own — `class User(Base)` yields the literal string
`"Base"`. So there are two passes: one collects each module's definitions and
imports, the other resolves names against those tables.
```
from .db import Base ; class User(Base)
→ app.models.User --inherits--> app.db.Base not "Base"
```
**Unresolved names become `kind: "external"` nodes and keep their edges.**
Dropping them is the worse failure: the diagram looks complete and has quietly
lost a dependency. Gathered by the index emitter, they *are* the project's
dependency surface.
An unparseable file is recorded as a node with an `error` attr, not a crash —
one bad file must not cost you the other four hundred.
`calls` edges are deliberately **not** attempted. Resolving `self.foo()` needs
type inference, and a call graph that is quietly 60% right is worse than none
because it reads as authoritative.
### A second source
`extractors/db.py` reads the published `{models, relationships, source}`
contract that `modelgen` already emits and `graphgen` already consumes. Tables
become nodes, columns become contained nodes, foreign keys become edges — with
no new top-level field, which was the checkpoint on whether the schema was right.
Connecting to a live database is not here. `modelgen from-db --url ...` does
that and writes the schema this reads; the two-step also keeps credentials out
of this pipeline entirely.
## Views are not an emitter concern
The first real diagram out of this pipeline was a 3000px strip: four modules of
content and sixty `sys`/`json`/`typing` boxes, all peers. The emitter was
correct and the picture was useless. That is a **missing view**, and the fix
belongs to every consumer at once — the index, the diagram and the diff all want
"just this subsystem, two hops out, without the stdlib".
```bash
python3 -m docgen view ir.json --drop-stdlib --around docgen.ir --hops 2 -o view.json
```
`drop_stdlib`, `drop_external`, `only_kinds`, `drop_kinds`, `subtree`,
`neighbourhood`, `collapse_to_depth`. All IR→IR, all composable, each producing
a document that still validates.
Graph *algorithms* are not here. Transitive reduction, cycle detection and
dominators are `networkx`'s, and reimplementing them is the classic way to
acquire a quiet bug. `lab/` is where that dependency gets tried against real IRs
before anything depends on it — the aim being to learn which part of it is
actually attractive, rather than adopting all of it on faith.
## One colour language
A style rule names a **slot**, never a colour. `"border": "atlas"` is the rule;
the theme binds `atlas` to `#43A047` in print and `#15803d` on the docs site.
That indirection is the whole point. `common/theme/tokens.css`,
`docs/graphs/themes/*.gvpr` and `style/lucid.json` use the same slot names, so a
diagram and the page around it match by construction — which is the rule
`docs/graphs/README.md` already states. The `dark` theme's `artery`, `atlas` and
`station` slots are exactly the `--system-accent` values set in
`artery/index.html:30`, `atlas/index.html:25` and `station/index.html:29`, and
the selftest fails if they drift apart.
An unknown `kind` falls back to `default` rather than crashing, so a new
extractor renders plainly and legibly on day one instead of needing a style file
written first.
**How a container picks its colour without the IR naming one:** it does not. The
IR says which spr model a group belongs to (`attrs.domain` — semantic), and
`domain_slots` maps that to a slot. Same mechanism as `--system-accent`. With no
domain, the emitter assigns by sorted id, so two runs agree.
## Use DOT until it hits its limits
The emitter writes what DOT expresses natively and stops at the boundary rather
than growing machinery. The limits are recorded in `style/lucid.json` under
`limits` and reachable as `Style.limits()`:
| | |
|---|---|
| header bars | a cluster has a label and a fill, not a 100%-width header rectangle |
| `stroke-dasharray` | not parameterised — `4,4` and `5,5` collapse to one dash |
| corner radius | `rounded` is binary, so 4px and 6px are identical |
| icon above label | needs an HTML-like label table |
| sequence badges | `xlabel` carries the number; the circle does not exist |
Those mark where a richer emitter would begin. The style file carries the full
spec regardless, so that emitter needs no re-authoring.
One limit that *was* worth solving: DOT cannot use a cluster as an edge
endpoint, so every module-to-module import silently vanished. The native answer
is `compound=true` with `lhead`/`ltail` — draw between a representative leaf and
clip at the cluster border.
## The output is addressable
`id` and `kind` pass through to the SVG as the element's `id` and `class`, and
`attrs.file`/`attrs.line` become an `href`. A front end can bind behaviour to a
box and a box can link to the line it came from, without the emitter knowing
about either.
## Testing
```bash
make check # docgen's own suite: 287 after make sync, 272 with nothing
make check BOOK=out/book/x # one book's own level — generated and hand-written checks
make doctor # the machine; never fails
```
**Golden tests go on the IR, never on the SVG.** Graphviz measures label text
with the host's fonts to size nodes, so identical input gives different geometry
on a machine with different fontconfig. The IR is deterministic; the SVG is not.
Self-hosting is the honest end-to-end check, and it is where the real bugs came
from — two name-resolution faults that no fixture had reached:
```bash
make self # docgen's book of the widest tree it can see, then its checks
```
## Where this sits
`docgen` belongs to Atlas — documentation is whose concern it is. It is **not** a
station tool and is not under `station/tools/`; it *may depend on* station tools,
which is the permitted direction.
Atlas 2 is a successor, not a replacement. `soleprint/atlas/` is untouched: it
carries client information and an idea still worth extracting — deriving frontend
and backend tests from one source, which is the same shape as this pointed the
other way.
## Not here
No layout system, no positioning, no ELK. No HTML-like labels, no SVG post-pass.
No LLM in the structural path — annotation (summarising a module, naming a
cluster) is a later layer, cached to its own file keyed by node `id`, merged into
`attrs` at emit time, and extraction must work with it absent. No configuration
knobs until two real consumers disagree.

View File

@@ -0,0 +1 @@
"""Docgen — code to diagram. The IR is the product; diagrams are one consumer."""

View File

@@ -0,0 +1,8 @@
""" python3 -m docgen <command> [args] — see cli/__init__.py for the commands."""
import sys
from .cli import main
if __name__ == "__main__":
sys.exit(main())

View File

@@ -0,0 +1,315 @@
"""
A **book** — one docgen operation, measured at both ends.
larder ──► step ──► step ──► step ──► book
what each one usable what
came in by itself came out
The first step says what went into the operation. The last step says what came
out, and renders it for the web. Everything between is an ordinary artifact that
stands on its own — `ir.json` is still an `ir.json`, and `make ir` still works
without knowing this file exists.
## Why both ends, and why they are steps rather than a wrapper
Because the two numbers are only worth having **together**. "102 nodes" is not a
fact about anything; "49 files in, 102 nodes out, nothing lost" is. A book that
read 45 of 47 files and drew a clean diagram is lying by omission, and before
this there was no place for the other 2 to be mentioned.
They are steps, not a wrapper, because a wrapper is something you can forget to
apply. A step is in the sequence, and the sequence is the thing the notebook
emits — so the measure is in the document whether or not anyone remembered.
## The two things it is not
**Not a gate.** Running one step alone is still a book, just a short one. An
operation that cannot measure something says what it could not measure and
carries on. Gating would break the property that makes the intermediate
artifacts useful, and that property is the whole reason the sequence is worth
having.
**Not a new pipeline.** Everything here composes functions that already existed
— `ops.overview`, `emitters.dot`, `emitters.site`, `notebook.spec`. The spine
adds a ledger and two measurements. If it ever starts doing the work itself,
something has gone wrong.
## The web end is deliberately the loose one
It is last, so nothing depends on it, so it can be replaced wholesale without
touching a single thing upstream. That is what lets it rule what gets generated
without being a stable contract: the book measure is the promise, and the page
that displays it is free to change drastically and often.
"""
import json
from dataclasses import dataclass, field
from pathlib import Path
from .larder import Larder
# The two spine steps. Named here because the notebook, the site and the checks
# all have to agree on what they are called, and three string literals in three
# files is how they stop agreeing.
FIRST = "larder"
LAST = "book"
@dataclass
class Step:
"""One thing the operation did, and what it left behind."""
id: str
label: str
artifact: str | None = None # relative to the book directory
bytes: int = 0
note: str = ""
skipped: str = "" # why, when it did not run
def to_dict(self) -> dict:
out = {"id": self.id, "label": self.label}
if self.artifact:
out["artifact"] = self.artifact
out["bytes"] = self.bytes
if self.note:
out["note"] = self.note
if self.skipped:
out["skipped"] = self.skipped
return out
# How a larder's unit shows up on the output side. Per-kind because the relation
# genuinely differs, and because getting it wrong produces a check that passes
# for the wrong reason — which is what the first version of this did.
#
# relation "exact" one unit in, one node out. Fewer means input was dropped.
# "collapse" many units in, fewer nodes out, by design.
# noun what to call the output-side thing, in the claim
# represents whether a unit that FAILED still appears as a node. Where it does,
# that is checkable and is the whole-input form of the rule that an
# unresolved name becomes an `external` node rather than vanishing.
RELATION = {
"python": ("exact", "module", True),
"code": ("exact", "module", True),
"db": ("exact", "table", False),
"openapi": ("exact", "endpoint", False),
# A HAR entry with no URL has nothing to represent, so there is no node to
# look for. Its absence is the correct outcome and is recorded in `failed`.
"usage": ("collapse", "call", False),
}
def _unit_counts(ir: dict, larder: Larder) -> tuple[int, int]:
"""(nodes from units that were read, nodes from units that failed).
Split because a failed file still gets a module node — carrying
`attrs.error`, so the gap is visible in the graph rather than only in a log.
Counting them together made "2 files read produced 4 modules" pass a check
that was supposed to prove nothing had been dropped.
"""
nodes = ir.get("nodes") or []
if larder.kind in ("python", "code"):
mods = [n for n in nodes
if n["kind"] == "module" and (n.get("attrs") or {}).get("file")]
broken = {n["attrs"]["file"] for n in mods if (n.get("attrs") or {}).get("error")}
whole = {n["attrs"]["file"] for n in mods} - broken
return len(whole), len(broken)
if larder.kind == "db":
return sum(1 for n in nodes if n["kind"] == "table"), 0
if larder.kind == "openapi":
return len({
(n.get("attrs") or {}).get("path") for n in nodes
if n["kind"] == "endpoint" and (n.get("attrs") or {}).get("path")
}), 0
if larder.kind == "usage":
return sum(1 for n in nodes if n["kind"] in ("endpoint", "operation")), 0
return 0, 0
class Book:
"""A named operation, its ledger, and the two measures that bracket it."""
def __init__(self, slug: str, larder: Larder, out):
self.slug = slug
self.larder = larder
self.out = Path(out)
self.steps: list[Step] = []
self.ir: dict | None = None
self.notebook: str | None = None
# The first step, recorded before any work happens. Doing it here rather
# than at the end is the difference between a measure and a summary: it
# says what the operation *set out* to read, so a crash halfway leaves a
# book that still says what went in.
self.steps.append(Step(
id=FIRST,
label="what came in",
note=larder.line(),
))
# -- the ledger -------------------------------------------------------
def step(self, id: str, label: str, path=None, note: str = "",
skipped: str = "") -> Step:
"""Record a step. `path` is written already; this measures it."""
artifact, size = None, 0
if path is not None:
path = Path(path)
if path.exists():
artifact = str(path.relative_to(self.out)) if self.out in path.parents \
or path.parent == self.out else str(path)
size = path.stat().st_size if path.is_file() else _tree_bytes(path)
s = Step(id=id, label=label, artifact=artifact, bytes=size,
note=note, skipped=skipped)
self.steps.append(s)
return s
# -- the last step ----------------------------------------------------
def measure(self) -> dict:
"""What came out. Counted off the final IR and the ledger."""
ir = self.ir or {"nodes": [], "edges": []}
by_kind: dict[str, int] = {}
for n in ir.get("nodes") or []:
by_kind[n["kind"]] = by_kind.get(n["kind"], 0) + 1
edge_kinds: dict[str, int] = {}
for e in ir.get("edges") or []:
edge_kinds[e["kind"]] = edge_kinds.get(e["kind"], 0) + 1
artifacts = [s for s in self.steps if s.artifact]
return {
"nodes": sum(by_kind.values()),
"by_kind": {k: by_kind[k] for k in sorted(by_kind)},
"edges": sum(edge_kinds.values()),
"edges_by_kind": {k: edge_kinds[k] for k in sorted(edge_kinds)},
"external": by_kind.get("external", 0),
"steps": len(self.steps),
"artifacts": [{"path": s.artifact, "bytes": s.bytes} for s in artifacts],
"bytes": sum(s.bytes for s in artifacts),
}
def compare(self) -> list[dict]:
"""Reconcile the two ends. This is what having both is *for*.
Returns observations, each with a verdict, rather than raising: the book
is already built by the time anyone can compare, and refusing to write it
would destroy the evidence. `book/checks.py` turns these into pass/fail
at the book test level, and the CLI exits 1 when one of them is not ok.
"""
out = []
ir = self.ir or {}
produced, represented = _unit_counts(ir, self.larder)
relation, noun, represents = RELATION.get(
self.larder.kind, ("collapse", "node", False))
read, unit = self.larder.read, self.larder.unit
n_failed = len(self.larder.failed)
if relation == "exact":
ok = produced >= read
out.append({
"id": "units-accounted-for",
"ok": ok,
"claim": f"{read} {unit}(s) read produced {produced} {noun}(s)",
"why": (
"one unit in, one node out" if ok else
f"{read - produced} {unit}(s) were read and produced nothing — "
"the input was dropped between the extractor and the document, "
"which is the failure this measure exists to catch"
),
})
else:
out.append({
"id": "collapse-is-intended",
"ok": produced >= 1 or read == 0,
"claim": f"{read} {unit}(s) collapsed to {produced} {noun}(s)",
"why": "many recorded requests describe few endpoints — that is the point",
})
if n_failed:
out.append({
"id": "failures-surfaced",
"ok": True,
"claim": f"{n_failed} {unit}(s) could not be read",
"why": "named in the larder measure and on the landing page, not only in a log",
})
# The whole-input form of "an unresolved name becomes an `external` node".
# A file that failed to parse must still be in the graph, or the diagram
# shows a tree that is smaller than the tree on disk and says nothing.
if represents and n_failed:
out.append({
"id": "failures-still-in-the-graph",
"ok": represented >= n_failed,
"claim": f"{n_failed} unreadable {unit}(s) appear as {represented} "
f"marked {noun}(s)",
"why": (
"a gap that is drawn can be seen" if represented >= n_failed else
f"{n_failed - represented} unreadable {unit}(s) are missing from the "
"graph entirely — the picture is smaller than the source and does "
"not say so"
),
})
return out
# -- writing ----------------------------------------------------------
def close(self, *, site=None) -> Step:
"""Append the last step. Call once, after the web output is written.
Must be called BEFORE the notebook is built, because the notebook quotes
the book measure and the measure is not complete until this step exists.
The site is this step's *artifact*. The notebook is not: it is the whole
sequence's rendering rather than an item in it, and it is set on
`self.notebook` afterwards, by name only — a notebook cannot report its
own byte count without changing it.
"""
m = self.measure()
summary = " · ".join(f"{v} {k}" for k, v in
sorted(m["by_kind"].items(), key=lambda kv: -kv[1]))
artifact, size = None, 0
if site is not None:
site = Path(site)
if site.exists():
artifact, size = str(site.relative_to(self.out)), _tree_bytes(site)
self.steps.append(Step(
id=LAST,
label="what came out",
artifact=artifact,
bytes=size,
note=f"{summary} · {m['edges']} edges" if summary else "nothing was produced",
))
return self.steps[-1]
def to_dict(self) -> dict:
"""The book, as the ledger someone else can read.
`generated_at` is absent for the same reason the IR's is: an unchanged
larder must serialise to the same bytes, or nothing downstream can tell
a real change from a rebuild.
"""
out = {
"slug": self.slug,
"larder": self.larder.to_dict(),
"steps": [s.to_dict() for s in self.steps],
"book": self.measure(),
"reconciled": self.compare(),
}
if self.notebook:
out["notebook"] = self.notebook
return out
def write(self) -> Path:
self.out.mkdir(parents=True, exist_ok=True)
path = self.out / "book.json"
path.write_text(json.dumps(self.to_dict(), indent=2, sort_keys=False) + "\n")
return path
def _tree_bytes(root: Path) -> int:
return sum(p.stat().st_size for p in root.rglob("*") if p.is_file())

Some files were not shown because too many files have changed in this diff Show More