rig major updates

This commit is contained in:
2026-09-22 05:15:49 -03:00
parent 9c963514f1
commit 2a0a793f19
64 changed files with 1762 additions and 645 deletions

View File

@@ -0,0 +1,98 @@
# `examples/data` — an overlay with a database, a scheduler and a worked pipeline
postgres, redis and airflow, each an upstream image run unmodified, installed as
this overlay's own addons. rig ships the mechanism that finds and runs them; which
services a workload needs is the workload's business, so they live here rather
than in rig's `ctrl/addons/`. Copy what you need into your own overlay's `addons/`.
```bash
OVERLAY=examples/data make cluster up # cluster `data`: postgres, then airflow
OVERLAY=examples/data make tilt # the items-api simulator the DAG reads
```
```
rig.env ADDONS (metallb is rig's; the rest are here), namespace, identities, image pins
addons/postgres.sh one replica on a PVC; the password generated once and kept
addons/airflow.sh one `standalone` pod on its own `airflow` database; needs postgres
addons/redis.sh only for switching airflow to CeleryExecutor (not in ADDONS)
k8s/ the items-api simulator
dags/items_to_postgres.py the worked example: API client → adapter → postgres
Tiltfile names the simulator's resource
```
Everything lands in the `data` namespace (`DATA_NAMESPACE`, and the kustomization
names it too), so resetting an app's namespace leaves the databases alone. Costs
roughly 2 GB with airflow, under 1 without. Airflow's first boot runs the whole
metadata migration, so expect a few minutes before it is ready.
## The worked example: three links kept apart
- **Metadata DB (infra).** `addons/airflow.sh` creates an `airflow` database on the
same postgres and points `SQL_ALCHEMY_CONN` there — airflow's own tables never land
in the app's database.
- **DAG delivery.** The addon turns this overlay's `dags/` into the `airflow-dags`
ConfigMap, mounted at `/opt/airflow/dags`; re-run `make cluster up` after editing a
DAG. The faster path later: a kind `extraMount` of `dags/` plus a Tilt `sync` — noted,
not built.
- **Data connection (operational logic).** `AIRFLOW_CONN_APP_DB`, composed each run
from the postgres secret, gives DAGs the app's database as the `app_db` connection.
One password reaches both URLs, with nowhere to drift.
The DAG itself calls the simulator with its own HTTP client (the wire, as the API
returns it), renames the wire's fields into the app's names in `to_app_row` — **the
adapter, which belongs to whoever owns the app's model and so lives in the overlay** —
and upserts into `items`. Hourly, no backfill, one retry, idempotent on `item_id`.
```bash
kubectl --context kind-data -n data exec deploy/airflow -- airflow dags unpause items_to_postgres
kubectl --context kind-data -n data exec deploy/airflow -- airflow dags trigger items_to_postgres
kubectl --context kind-data -n data exec deploy/postgres -- psql -U app -d app -c 'select * from items'
```
## Addons in an overlay
Each one runs from rig's `ctrl/` (rig's `addons.sh` exports `RIG_CTRL`), so it
starts with `cd "${RIG_CTRL:?...}"`, sources `./lib/config.sh` and calls
`load_config` — and sees every key this overlay's `rig.env` sets. Run them through
rig (`bash ctrl/addons.sh install`, or `make cluster up`), not directly.
## postgres — plain manifests, one replica
Plain manifests rather than a helm chart: a chart repo is a network dependency,
and an offline machine needs a path with none. The image is pinned in `rig.env`
and can be preloaded into a local registry like every other image.
One replica on a PVC. This models a dependency for local work, not a
highly-available database, and pretending otherwise on a kind node would be a
more elaborate lie rather than a more useful one.
The password is not in `rig.env`: `addons/postgres.sh` generates one on first
install and keeps it across re-runs, so re-running the addon never rotates the
credential out from under whatever is already connected.
## redis
Cache, and the broker anything queue-shaped runs on. No persistence: a broker
that loses its queue on restart is the honest local model, and a PVC here buys
nothing but a volume to clean up.
## airflow
Airflow needs a metadata database before it will start at all, so the script
refuses rather than rolls a pod that will CrashLoopBackOff while the real problem
(postgres missing from `ADDONS`) stays invisible in the logs.
One pod on `standalone`: migration, admin user, scheduler and webserver in a
single container, on LocalExecutor, which needs no broker. The official chart's
five deployments model an installation; switching this on means wanting pipelines.
redis is here for the day it moves to CeleryExecutor, and not before.
## Reaching them
Reach the databases with port-forward rather than binding more host ports:
```bash
kubectl --context kind-data -n data port-forward svc/postgres 5432:5432
kubectl --context kind-data -n data port-forward svc/airflow 8080:8080
kubectl --context kind-data -n data get secret postgres -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d
```

View File

@@ -0,0 +1,6 @@
# The data overlay's half of the dev loop: the items-api simulator the example DAG reads.
# rig's ctrl/Tiltfile has already applied k8s/overlays/dev; paths here are relative to
# this folder. postgres and airflow are addons (make cluster up), not Tilt resources.
# Notes: README.md
k8s_resource('items-api', labels=['simulator'])

View File

@@ -0,0 +1,150 @@
#!/usr/bin/env bash
# Apache Airflow for this overlay: one `standalone` pod (LocalExecutor — no broker).
# Its own `airflow` database on the postgres addon; the app's data reaches DAGs as the
# `app_db` connection; DAGs from this overlay's dags/, as a ConfigMap.
# Requires the postgres addon; refuses to install without it.
# Notes: ../README.md
set -euo pipefail
cd "${RIG_CTRL:?run it through rig: bash ctrl/addons.sh install}"
source ./lib/config.sh
load_config
K="kubectl --context ${KUBECONTEXT}"
NS="${DATA_NAMESPACE:-data}"
if ! $K get deployment -n "$NS" postgres >/dev/null 2>&1; then
echo " ! airflow needs the postgres addon, and it is not installed" >&2
echo " add it before airflow in the overlay's ADDONS:" >&2
echo " ADDONS=\"... postgres airflow\"" >&2
exit 1
fi
# ── metadata DB: airflow's own tables, kept out of the app's database ──────
# Same postgres instance, separate database, created once (idempotent).
db_user=$($K get secret -n "$NS" postgres -o jsonpath='{.data.POSTGRES_USER}' | base64 -d)
db_pass=$($K get secret -n "$NS" postgres -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d)
db_name=$($K get secret -n "$NS" postgres -o jsonpath='{.data.POSTGRES_DB}' | base64 -d)
psql() { $K exec -n "$NS" deploy/postgres -- psql -U "$db_user" -d "$db_name" -tAc "$1"; }
if [ "$(psql "SELECT 1 FROM pg_database WHERE datname = 'airflow'")" = 1 ]; then
echo " database 'airflow' exists"
else
psql "CREATE DATABASE airflow" >/dev/null
echo " created database 'airflow' beside '${db_name}'"
fi
# ── what is generated once and kept: re-running never rotates these ────────
if $K get secret -n "$NS" airflow >/dev/null 2>&1; then
echo " secret exists, keeping the current admin password and fernet key"
else
admin_password=$(head -c 18 /dev/urandom | base64 | tr -d '/+=' | head -c 24)
# Airflow requires a 32-byte urlsafe-base64 key; without a fixed one every
# restart invalidates every stored connection.
fernet_key=$(head -c 32 /dev/urandom | base64 | tr '+/' '-_')
$K create secret generic airflow -n "$NS" \
--from-literal=ADMIN_USER="${AIRFLOW_ADMIN_USER:-admin}" \
--from-literal=ADMIN_PASSWORD="$admin_password" \
--from-literal=FERNET_KEY="$fernet_key" \
>/dev/null
echo " generated an admin password (read it back with the command below)"
fi
# ── connections: composed from what the postgres secret owns, every run ─────
# Nothing to drift: one password reaches the metadata DB and the data connection.
$K create secret generic airflow-connections -n "$NS" \
--from-literal=SQL_ALCHEMY_CONN="postgresql+psycopg2://${db_user}:${db_pass}@postgres:5432/airflow" \
--from-literal=AIRFLOW_CONN_APP_DB="postgres://${db_user}:${db_pass}@postgres:5432/${db_name}" \
--dry-run=client -o yaml | $K apply -f - >/dev/null
# ── DAG delivery: this overlay's dags/ as a ConfigMap ───────────────────────
# Edits land by re-running this addon (make cluster up). The later path — a kind
# extraMount of dags/ plus a Tilt sync — is noted in the README, not built.
# A ConfigMap volume is kubelet's ..data/..<timestamp> symlinks, and Airflow's DAG walker
# follows symlinks: without the .airflowignore it stops at "Detected recursive loop".
dags="$(_from_ctrl "$OVERLAY_DIR")/dags"
if [ -d "$dags" ]; then
$K create configmap airflow-dags -n "$NS" --from-file="$dags" \
--from-literal=.airflowignore='^\.\.' \
--dry-run=client -o yaml | $K apply -f - >/dev/null
echo " dags: $(ls "$dags" | grep -c '\.py$') file(s) from $(basename "$(_abs_from_ctrl "$OVERLAY_DIR")")/dags"
fi
echo " applying manifests"
$K apply -n "$NS" -f - >/dev/null <<YAML
apiVersion: v1
kind: Service
metadata:
name: airflow
spec:
selector:
app: airflow
ports:
- port: 8080
targetPort: 8080
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: airflow
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app: airflow
template:
metadata:
labels:
app: airflow
spec:
containers:
- name: airflow
image: ${AIRFLOW_IMAGE}
args: ["standalone"]
env:
- name: AIRFLOW__CORE__EXECUTOR
value: LocalExecutor
- name: AIRFLOW__CORE__LOAD_EXAMPLES
value: "false"
- name: AIRFLOW__DATABASE__SQL_ALCHEMY_CONN
valueFrom:
secretKeyRef: {name: airflow-connections, key: SQL_ALCHEMY_CONN}
- name: AIRFLOW_CONN_APP_DB
valueFrom:
secretKeyRef: {name: airflow-connections, key: AIRFLOW_CONN_APP_DB}
- name: AIRFLOW__CORE__FERNET_KEY
valueFrom:
secretKeyRef: {name: airflow, key: FERNET_KEY}
- name: _AIRFLOW_WWW_USER_USERNAME
valueFrom:
secretKeyRef: {name: airflow, key: ADMIN_USER}
- name: _AIRFLOW_WWW_USER_PASSWORD
valueFrom:
secretKeyRef: {name: airflow, key: ADMIN_PASSWORD}
ports:
- containerPort: 8080
volumeMounts:
- name: dags
mountPath: /opt/airflow/dags
readinessProbe:
httpGet:
path: /health
port: 8080
# First boot runs the whole migration before it serves anything.
initialDelaySeconds: 60
periodSeconds: 15
failureThreshold: 20
volumes:
- name: dags
configMap:
name: airflow-dags
optional: true
YAML
echo " waiting for airflow (the first boot migrates the database, so this is slow)..."
$K rollout status deployment/airflow -n "$NS" --timeout=600s
echo " in-cluster: http://airflow.${NS}.svc.cluster.local:8080"
echo " reach it: kubectl --context ${KUBECONTEXT} -n ${NS} port-forward svc/airflow 8080:8080"
echo " password: kubectl --context ${KUBECONTEXT} -n ${NS} get secret airflow -o jsonpath='{.data.ADMIN_PASSWORD}' | base64 -d"

View File

@@ -0,0 +1,105 @@
#!/usr/bin/env bash
# PostgreSQL for this overlay: plain manifests, one replica on a PVC, password generated once and kept.
# Notes: ../README.md
set -euo pipefail
cd "${RIG_CTRL:?run it through rig: bash ctrl/addons.sh install}"
source ./lib/config.sh
load_config
K="kubectl --context ${KUBECONTEXT}"
NS="${DATA_NAMESPACE:-data}"
$K get namespace "$NS" >/dev/null 2>&1 || $K create namespace "$NS"
# The password is generated once and then left alone, so re-running this does
# not rotate the credential out from under whatever is already connected.
if $K get secret -n "$NS" postgres >/dev/null 2>&1; then
echo " secret exists, keeping the current password"
else
password=$(head -c 18 /dev/urandom | base64 | tr -d '/+=' | head -c 24)
$K create secret generic postgres -n "$NS" \
--from-literal=POSTGRES_DB="${POSTGRES_DB:-postgres}" \
--from-literal=POSTGRES_USER="${POSTGRES_USER:-postgres}" \
--from-literal=POSTGRES_PASSWORD="$password" >/dev/null
echo " generated a password (read it back with the command printed below)"
fi
echo " applying manifests"
$K apply -n "$NS" -f - >/dev/null <<YAML
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: postgres-data
spec:
accessModes: [ReadWriteOnce]
resources:
requests:
storage: ${POSTGRES_STORAGE:-2Gi}
---
apiVersion: v1
kind: Service
metadata:
name: postgres
spec:
selector:
app: postgres
ports:
- port: 5432
targetPort: 5432
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: postgres
spec:
replicas: 1
# One volume, one writer. Rolling would start a second pod against the same
# PVC before the first exits, and Postgres refuses to share a data directory.
strategy:
type: Recreate
selector:
matchLabels:
app: postgres
template:
metadata:
labels:
app: postgres
spec:
containers:
- name: postgres
image: ${POSTGRES_IMAGE}
envFrom:
- secretRef:
name: postgres
env:
# The image initialises into the volume root otherwise, and a
# lost+found from the PVC makes it refuse to initdb.
- name: PGDATA
value: /var/lib/postgresql/data/pgdata
ports:
- containerPort: 5432
volumeMounts:
- name: data
mountPath: /var/lib/postgresql/data
readinessProbe:
exec:
command: ["sh", "-c", "pg_isready -U \$POSTGRES_USER -d \$POSTGRES_DB"]
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
exec:
command: ["sh", "-c", "pg_isready -U \$POSTGRES_USER -d \$POSTGRES_DB"]
initialDelaySeconds: 30
periodSeconds: 15
volumes:
- name: data
persistentVolumeClaim:
claimName: postgres-data
YAML
echo " waiting for postgres..."
$K rollout status deployment/postgres -n "$NS" --timeout=240s
echo " in-cluster: postgres.${NS}.svc.cluster.local:5432"
echo " password: kubectl --context ${KUBECONTEXT} -n ${NS} get secret postgres -o jsonpath='{.data.POSTGRES_PASSWORD}' | base64 -d"

View File

@@ -0,0 +1,57 @@
#!/usr/bin/env bash
# Redis for this overlay: cache and broker (Celery), no persistence.
# Notes: ../README.md
set -euo pipefail
cd "${RIG_CTRL:?run it through rig: bash ctrl/addons.sh install}"
source ./lib/config.sh
load_config
K="kubectl --context ${KUBECONTEXT}"
NS="${DATA_NAMESPACE:-data}"
$K get namespace "$NS" >/dev/null 2>&1 || $K create namespace "$NS"
echo " applying manifests"
$K apply -n "$NS" -f - >/dev/null <<YAML
apiVersion: v1
kind: Service
metadata:
name: redis
spec:
selector:
app: redis
ports:
- port: 6379
targetPort: 6379
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: redis
spec:
replicas: 1
selector:
matchLabels:
app: redis
template:
metadata:
labels:
app: redis
spec:
containers:
- name: redis
image: ${REDIS_IMAGE}
ports:
- containerPort: 6379
readinessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 3
periodSeconds: 5
YAML
echo " waiting for redis..."
$K rollout status deployment/redis -n "$NS" --timeout=180s
echo " in-cluster: redis://redis.${NS}.svc.cluster.local:6379/0"

View File

@@ -0,0 +1,81 @@
"""Pull items from the items API, rename them into the app's names, upsert into postgres.
Three links, kept apart on purpose:
- the API client (`fetch_items`) talks to the wire as it is — here the overlay's
own simulator, `items-api`, whose field names are the API's;
- the adapter (`to_app_row`) is the one place the wire's names become the app's:
`id` -> `item_id`, `name` -> `item_name`, `price.amount_cents` -> `price_cents`.
It belongs to whoever owns the app's model, so it lives in the overlay, not in rig;
- the load writes through the `app_db` connection (AIRFLOW_CONN_APP_DB, built by
addons/airflow.sh from the postgres secret) and is idempotent: an upsert keyed
on `item_id`, so a retry or a rerun never duplicates a row.
Operational logic is explicit and minimal: hourly, no backfill, one retry.
"""
import json
import urllib.request
from datetime import datetime, timedelta
from airflow import DAG
from airflow.operators.python import PythonOperator
ITEMS_URL = "http://items-api/v1/items"
CREATE = """
CREATE TABLE IF NOT EXISTS items (
item_id text PRIMARY KEY,
item_name text NOT NULL,
price_cents integer NOT NULL,
currency text NOT NULL,
loaded_at timestamptz NOT NULL DEFAULT now()
)
"""
UPSERT = """
INSERT INTO items (item_id, item_name, price_cents, currency)
VALUES (%(item_id)s, %(item_name)s, %(price_cents)s, %(currency)s)
ON CONFLICT (item_id) DO UPDATE
SET item_name = EXCLUDED.item_name,
price_cents = EXCLUDED.price_cents,
currency = EXCLUDED.currency,
loaded_at = now()
"""
def fetch_items():
"""The API client: the wire, as the API returns it."""
with urllib.request.urlopen(ITEMS_URL, timeout=10) as response:
return json.load(response)["items"]
def to_app_row(item):
"""The adapter: the API's names in, the app's names out."""
return {
"item_id": item["id"],
"item_name": item["name"],
"price_cents": item["price"]["amount_cents"],
"currency": item["price"]["currency"],
}
def load_items():
from airflow.providers.postgres.hooks.postgres import PostgresHook
rows = [to_app_row(item) for item in fetch_items()]
hook = PostgresHook(postgres_conn_id="app_db")
hook.run(CREATE)
for row in rows:
hook.run(UPSERT, parameters=row)
print(f"upserted {len(rows)} items")
with DAG(
dag_id="items_to_postgres",
schedule="@hourly",
start_date=datetime(2026, 1, 1),
catchup=False,
default_args={"retries": 1, "retry_delay": timedelta(minutes=1)},
tags=["example"],
) as dag:
PythonOperator(task_id="load_items", python_callable=load_items)

View File

@@ -0,0 +1,95 @@
# The simulator: a stub of the API the DAG reads, faithful to the wire (its field
# names are the API's, not the app's). Same shape as the starter's example-mock.
apiVersion: v1
kind: ConfigMap
metadata:
name: items-api-stub
data:
routes.json: |
{
"/health": {"status": 200, "body": {"status": "ok"}},
"/v1/items": {"status": 200, "body": {"items": [
{"id": "a-100", "name": "anvil", "price": {"amount_cents": 1999, "currency": "USD"}},
{"id": "b-200", "name": "bucket", "price": {"amount_cents": 450, "currency": "USD"}},
{"id": "c-300", "name": "crate", "price": {"amount_cents": 1200, "currency": "USD"}}
]}}
}
serve.py: |
import json, os
from http.server import BaseHTTPRequestHandler, HTTPServer
ROUTES = json.load(open("/etc/stub/routes.json"))
NAME = os.environ.get("STUB_NAME", "stub")
class H(BaseHTTPRequestHandler):
def do_GET(self):
r = ROUTES.get(self.path)
if r is None:
self.send_response(404)
self.end_headers()
self.wfile.write(json.dumps(
{"error": "no canned route", "stub": NAME, "path": self.path}
).encode())
return
body = json.dumps(r["body"]).encode()
self.send_response(r["status"])
self.send_header("Content-Type", "application/json")
self.send_header("X-Mocked-By", NAME)
self.end_headers()
self.wfile.write(body)
def log_message(self, fmt, *args):
print("%s %s" % (NAME, fmt % args), flush=True)
HTTPServer(("0.0.0.0", 8080), H).serve_forever()
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: items-api
labels:
app: items-api
rig.component/impl: mock
spec:
replicas: 1
selector:
matchLabels:
app: items-api
template:
metadata:
labels:
app: items-api
spec:
containers:
- name: stub
image: python:3.12-slim
command: ["python3", "/etc/stub/serve.py"]
env:
- name: STUB_NAME
value: items-api
ports:
- containerPort: 8080
volumeMounts:
- name: stub
mountPath: /etc/stub
readinessProbe:
httpGet: { path: /health, port: 8080 }
initialDelaySeconds: 2
resources:
requests: { memory: 32Mi, cpu: 10m }
limits: { memory: 64Mi }
volumes:
- name: stub
configMap:
name: items-api-stub
---
apiVersion: v1
kind: Service
metadata:
name: items-api
spec:
selector:
app: items-api
ports:
- port: 80
targetPort: 8080

View File

@@ -0,0 +1,10 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
# Beside the addons, in DATA_NAMESPACE — but no Namespace object: the addons created
# it, and a Namespace Tilt owned would be deleted by `tilt down`, taking postgres and
# airflow with it. rig's Tiltfile creates namespaces that are used and not declared.
namespace: data
resources:
- items-api.yaml

View File

@@ -0,0 +1,5 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ../../base

23
rig/examples/data/rig.env Normal file
View File

@@ -0,0 +1,23 @@
# This overlay's settings: layered over rig's defaults, under ctrl/.env and the caller.
# data — postgres, redis and airflow (upstream images, run unmodified) in their own namespace.
# Use it: OVERLAY=examples/data make cluster up Notes: README.md
# Order matters: addons install in the order listed, and airflow refuses to start
# without postgres, so postgres comes first. metallb is rig's own. Airflow runs
# LocalExecutor and needs no broker: add redis (before airflow) only to switch to Celery.
ADDONS="metallb postgres airflow"
# Namespace for the dependency containers (k8s/base/kustomization.yaml names it too).
DATA_NAMESPACE=data
# Postgres identity. The password is generated once by addons/postgres.sh and kept.
POSTGRES_DB=app
POSTGRES_USER=app
POSTGRES_STORAGE=2Gi
AIRFLOW_ADMIN_USER=admin
# Upstream images, pinned by tag; bump freely, and preload them for an offline machine.
POSTGRES_IMAGE=postgres:16-alpine
REDIS_IMAGE=redis:7-alpine
AIRFLOW_IMAGE=apache/airflow:2.10.4