Files
cross-py-builder/DESIGN.md
T
Brett Williams 18d8d2e398 Manage workers from the web UI, persisted in SQLite
Replace the CROSS_PY_WORKERS env list with a DB-backed worker store so agents
can be added/removed (and their online status verified) from the browser.
- db.py: add workers table + list_workers/add_worker/remove_worker.
- workers.py: configured_workers() reads the DB instead of settings.
- app.py: POST /api/workers (add, probes reachability) and
  DELETE /api/workers/<port>/<host>; startup log uses DB worker count.
- UI: Workers card with add form, per-worker live status + OS/CPU, remove.
- settings.py: drop CROSS_PY_WORKERS/_get_json_list; DB is single source.
- README/DESIGN updated to describe UI-managed workers.
2026-08-30 23:37:31 -05:00

219 lines
11 KiB
Markdown

# Cross-Py-Builder — Controller + Web UI Design
Target architecture: keep the existing worker (`build_agent.py`) HTTP API as-is, and add a
**controller** LXC that runs a web app. The controller stores every submitted request as a
**job row**, polls workers to pick a target, dispatches the build, records download artifacts,
and exposes a browser UI to submit, track progress, and download results.
The workers (`build_agent.py`) keep their HTTP API; only one small read-only endpoint
(`/progress/<job_id>`) is added for live build progress. Everything else is new controller code.
---
## 1. Components
```
Browser
│ (HTTP, port 8080)
┌───────────────────────────────── ─┐ ┌──────────────────────┐
│ CONTROLLER LXC │ │ WORKER A (ubuntu x64)│
│ │ │ │
│ Flask app (waitress) │ │ build_agent.py │
│ ├─ web UI (submit/track/download)│ │ /upload /status │
│ ├─ job API │ ───► │ /download /checkout │
│ ├─ scheduler (string executor) │ └──────────────────────┘
│ └─ SQLite (jobs.db) │ ┌──────────────────────┐
│ │ │ WORKER B (win arm64)│
└───────────────────────────────── ─┘ │ build_agent.py │
└──────────────────────┘
ctrl/ controller package
ctrl/static/ frontend (HTML/JS/CSS)
ctrl/jobs.db SQLite database (runtime)
builds/ downloaded artifacts (runtime)
```
**Worker discovery:** configurable via the web UI and persisted in the controller's SQLite
`workers` table (host + port per agent). On each job the controller queries every worker
`/status`, filters to `status=="ready"` and matching CPU/OS, and picks one. This replaces
Zeroconf on the controller; workers can stop advertising, and the Zeroconf code path can be
removed with the CLI retirement (§7).
**Worker management API** (adds/removes agents without touching config files):
| Endpoint | Method | Purpose |
|---|---|---|
| `GET /api/workers` | GET | live `/status` for each configured worker |
| `POST /api/workers` | POST | add a worker: JSON `{"spec": "host:port"}` (or `host`+`port`) |
| `DELETE /api/workers/<port>/<host>` | DELETE | remove a configured worker |
The workers table is seeded on demand from the UI; there is no `CROSS_PY_WORKERS`
environment bootstrap — the DB is the single source of truth.
We deliberately do **not** rewrite the agent's HTTP API, and do **not** add auth in v1 (same
trusted-LAN posture as today). The worker endpoint contract used by the controller:
| Worker endpoint | Method | Purpose |
|---|---|---|
| `/status` | GET | readiness, os, cpu, agent_version, hostname, ip |
| `/upload` | POST | multipart `file` zip → starts build, returns job JSON. `?async=1` returns `202` + `{"id": ...}` immediately, building in background |
| `/checkout_git` | POST | JSON `{"repo_url": ...}` → starts build. `?async=1` behaves like `/upload?async=1` |
| `/download/<job_id>` | GET | zip of `dist/` for a built job |
| `/delete_cache` | GET | clear cached jobs |
| `/healthz` | GET | liveness (already added) |
| `/progress/<job_id>` | GET | **NEW (small agent change):** live build snapshot (status, current step, elapsed, last log line) for SSE relay |
---
## 2. Data model
**SQLite table `jobs`** (one row per submitted build; single table keeps v1 simple):
| column | type | notes |
|---|---|---|
| `id` | text PK | controller job id (e.g. `JOB-<8 hex>`) — distinct from worker job id |
| `status` | text | `queued → dispatching → building → done \| failed \| cancelled` |
| `source_type` | text | `upload` \| `git` |
| `source` | text | uploaded filename or repo URL |
| `os_req` | text? | requested OS filter (nullable) |
| `cpu_req` | text? | requested CPU filter (nullable) |
| `worker_host` | text | worker chosen (host:port) |
| `worker_job_id` | text? | job id returned by the worker |
| `worker_url` | text? | base URL of worker (for download link) |
| `created_at` | text | ISO timestamp |
| `started_at` | text? | |
| `finished_at` | text? | |
| `log` | text | aggregated progress lines (appended, newline-delimited) |
| `error` | text? | last error detail |
| `artifacts` | text | JSON list of downloaded artifact paths/names |
Artifacts are downloaded by the controller into `builds/<job_id>/` (`<app>-<version>-<os>-<cpu>.zip`)
so results survive worker container restarts and remain available after the worker cleans up.
---
## 3. REST API (controller)
| Method | Path | Purpose |
|---|---|---|
| GET | `/` | web UI (single page) |
| GET | `/api/workers` | live `/status` for each configured worker |
| POST | `/api/workers` | add a worker: JSON `{"spec": "host:port"}` |
| DELETE | `/api/workers/<port>/<host>` | remove a configured worker |
| GET | `/api/capabilities` | distinct (os, cpu) across ready workers, for dropdowns |
| POST | `/api/jobs` | create job: multipart upload `file` OR JSON `{repo_url, os, cpu}` |
| GET | `/api/jobs` | list jobs (newest first, with status) |
| GET | `/api/jobs/<id>` | job detail incl. live `log` |
| POST | `/api/jobs/<id>/cancel` | set status → `cancelled` (best-effort) |
| GET | `/api/jobs/<id>/artifacts/<name>` | download a built artifact |
| GET | `/api/jobs/<id>/stream` | SSE log stream for live tail |
---
## 4. Scheduler (dispatch loop)
A single background thread owns all dispatch work.
1. Pop the oldest job with `status == "queued"` (FIFO).
2. From worker config, query `/status` (short timeout). Filter:
- `status` contains `ready` (not building/updating)
- if `os_req` set, worker `os` contains `os_req` (case-insensitive)
- if `cpu_req` set, worker `cpu` contains `cpu_req`
- exclude the local/controller LXC if present, to avoid self-builds
3. Choose first match (optionally: prefer workers with more free disk — v1: first match).
4. Mark job `dispatching`; set `worker_host`.
5. Send to worker (**async**): `POST /upload?async=1` (with the stored file) or
`POST /checkout_git?async=1` (with `{repo_url}`). The agent clones/saves the source, then
starts the build in a background thread (holding its `build_lock`) and returns `202`
immediately with `{"id": <worker_job_id>}`. On hard failure, mark `failed` (§ failure handling
below). Store `worker_job_id`, `worker_url`, set `building`.
6. **Progress relay** — while the build runs, poll `GET /progress/<worker_job_id>` on an interval,
appending each `last_log_line`/step snapshot to the job log and pushing to connected SSE
clients → live per-step progress in the UI. The build is complete when `/progress` returns
`404` (progress cleared) and the worker's `/status` reports `ready`. (This relies on the small
agent additions in §1.)
7. On completion: compute download URL `worker_url + /download/<worker_job_id>`, fetch the zip,
save to `builds/<job_id>/`, record artifact in the job row, set `done`.
**Concurrency/limits:** scheduler processes one job at a time (workers can already build only one
at a time — `409`). Multiple queued jobs simply wait. This makes NIC/disk behavior predictable
and avoids the current CLI dumping N builds in parallel.
**Failure handling:** if dispatch fails (worker down/409/500), mark job `failed` and record
`error` + log tail. No auto-retry in v1 (confirmed decision) — surface the error and let the
user resubmit to an eligible worker.
---
## 5. Frontend (single-page, no build-step)
- Ask for `/api/capabilities` → render OS + CPU dropdowns (or "any").
- Upload a zip **or** paste a git URL; submit.
- Job list refreshes via periodic `/api/jobs`.
- Detail view: status badge + log (poll, or SSE stream). Download button appears when `done`.
- No npm/JS toolchain — vanilla JS + `<link>`ed CSS; keeps the controller LXC dependency-light.
---
## 6. Project layout (new/changed files)
```
setup.py (add ctrl package + entry point `cross-py-controller`)
requirements.txt (add waitress; keep Flask/requests)
agent/ (worker agent: build_agent.py, zeroconf_server.py)
ctrl/
__init__.py
app.py (Flask factory, routes, waitress runner)
db.py (SQLite init + helpers; jobs + workers tables)
scheduler.py (dispatch + progress-poll loop, worker client)
workers.py (worker config load + /status probe)
settings.py (env-driven config: DB path, builds dir, worker list)
static/
index.html
app.js
style.css
```
Controller config example (env or JSON file):
```
CROSS_PY_DATA=/var/lib/cross-py-controller # holds jobs.db (jobs + workers) + builds/
CROSS_PY_PORT=8080
```
Workers are added/removed in the web UI and stored in the SQLite `workers` table (no
`CROSS_PY_WORKERS` env var).
---
## 7. What we reuse vs. retire
- **Reuse on workers:** agent HTTP API, plus the new `/progress/<job_id>` endpoint.
- **Retire the CLI entirely (confirmed decision):** `agent_manager.py` is removed. The
controller web UI is the sole frontend. This drops the Zeroconf-dependent manager code paths
(multi-parallel submits, hardcoded `DEFAULT_PORT`, the self-update flow) and leaves a single
interface to maintain.
## 8. Build environment model (per-build venv)
Workers keep creating a fresh venv and installing requirements per job. This is the isolated
model — no cross-job env contamination, and a bad or malicious `requirements.txt` can't poison
a shared environment. Tradeoff: each job pays ~2-5 min of pip installs (PyInstaller + deps)
and repeats downloads/disk churn.
Mitigation ladder, if wall-clock ever hurts (preserve isolation in all cases):
1. Persistent pip-cache volume on the worker (`~/.cache/pip`) — fast, no isolation loss.
2. Local wheel mirror / `--find-links` — same, plus offline-friendly.
3. Prebuilt PyInstaller baked into the worker LXC snapshot (skip its reinstall per job).
Do **not** move to a shared persistent venv until a concrete cross-job dependency problem
appears.
## 9. Deferred / next-phase (explicitly out of scope for v1)
- Auth / tokens on controller and worker endpoints
- Worker add/remove/update/restart/shutdown via the UI
- Auto-retry on a different worker; retry-with-backoff for queued jobs
- Persisting/streaming full-build logs from worker (`build-<name>.log` is on the worker today)
- Storing source zip centrally for re-runs
- Auto-scaling/provisioning of worker LXCs