Commit Graph

28 Commits (09cb8aea5c74ce6cafe99cbddaefcd02498e8ad3)

Author SHA1 Message Date
nexxo 5ae2dc41fc Ein Waechter, der aufraeumt — und die vier offenen Punkte aus dem Betrieb
DER WAECHTER. Ein abgebrochenes Update liess den Stapel halb unten liegen, den
Tunnel weg und die Seite eine Stunde lang mit 500 antworten — bis jemand von
Hand nachsah. Das wartet nie wieder auf einen Menschen.

deploy/watchdog.sh laeuft jede Minute als systemd-Dienst AUF DEM WIRT, nicht in
einem Container: ein Waechter im Container braeuchte den Docker-Socket
hineingereicht (Root auf dem Wirt fuer jeden, der je hineinkommt) und waere
genau dann tot, wenn man ihn braucht. Er nimmt dieselbe Sperre wie der
Update-Agent und fasst waehrend eines Updates nichts an.

Er kennt vier Fehlerbilder, alle vier heute wirklich passiert, und fuer jedes
genau einen Griff: fehlende Dienste starten; Container, die einander nicht mehr
finden, NEU ERZEUGEN (ein Neustart hilft da nicht); wg0 hochziehen; einen seit
ueber einer halben Stunde haengenden Wartungsmodus beenden. Was er nicht kennt,
protokolliert er und laesst es in Ruhe. `maintenance-hold` ist die Handbremse.
Beides nachgemessen: Dienst gestoppt -> wieder da; wg0 abgeraeumt -> wieder oben.

DREI LUECKEN, DIE DERSELBE VORFALL AUFGEDECKT HAT:

1. `phase()` machte `mkdir -p` ohne Fangnetz. Gehoerte storage/ nach einem
   frueheren Fehltritt root, starb das Update mit `set -e` an seiner ERSTEN
   Zeile, ohne eine einzige Ausgabe. Von aussen sah das aus wie
   "haengengeblieben"; in Wahrheit war es nach einer Millisekunde vorbei.
2. Ein zweiter Aufruf meldete "Already up to date" und tat NICHTS — der Checkout
   stand ja schon auf dem Ziel. Der Commit ist die falsche Frage; jetzt wird der
   Zustand gefragt: laeuft jeder Dienst, ist der Wartungsmodus aus.
3. Nach einer Netz-Umstellung reicht `up -d` nicht: nur neu gestartete Container
   haengen weiter am alten Netz, alle laufen, und trotzdem loest kein Name mehr
   auf. Jetzt `--force-recreate`.

DIE VIER PUNKTE AUS DEM BETRIEB:

- Provisioning zeigte 15/16 in der Liste und "16 von 16" in der Karte daneben,
  bei Status "Fertig" und 100 %. Die Liste rechnete `current_step + 1` — "dieser
  Schritt laeuft gerade" —, und bei einem fertigen Lauf laeuft keiner mehr.
- Mahngebuehren standen in CENT im Feld. Wer eine Gebuehr eintraegt, denkt in
  Euro und tippt "5"; daraus wurden fuenf Cent, ohne Widerspruch. Jetzt Euro im
  Feld, Cent in der Datenbank, gerundet statt abgeschnitten ((int)(19.99*100)
  ist 1998). Und Fristen und Geld stehen in zwei eigenen Bloecken mit eigener
  Ueberschrift, die Einheit im Feld statt in der Beschriftung.
- Die Bueroadresse stand in der Konsole neben dem Management-Netz mit dem
  Vermerk "nicht entfernbar" — sie kommt aber aus TRUSTED_RANGES in der .env und
  ist sehr wohl aenderbar. Beim naechsten Umzug des Anschlusses waere das ein
  Aussperren gewesen. Strukturell sind nur zwei Eintraege; alles andere steht
  jetzt als das da, was es ist, mit dem Weg heraus.
- Ein Host-Zugang hiess weiter "pve-fns-1", waehrend der Host laengst "fsn-01"
  hiess. Die Zeile zeigt jetzt den Namen des Hosts, nicht den einmal
  gespeicherten — damit traegt sie jede kuenftige Umbenennung von selbst.

Und die Frage "wofuer brauche ich Uptime Kuma": es ist benutzt — RegisterMonitoring
legt fuer jede Kunden-Instanz eine Ueberwachung an, SyncMonitoringStatus holt den
Stand, und ein Ausfall steht auf der Uebersicht. Ohne Kunden-Instanzen gibt es
nichts zu sehen, was wie "unbenutzt" aussieht. Das steht jetzt in den
Einstellungen, statt dort nur "API-Token und wo die Bruecke erreichbar ist".

2522 Tests gruen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 09:43:33 +02:00
nexxo 5f3d983658 Der Tunnel bekommt eine feste Adresse — damit ist die conntrack-Falle weg statt behandelt
Die Wurzel des ganzen Uebels, endlich an der Wurzel: Docker traegt den
veroeffentlichten UDP-Port als Weiterleitung auf die Container-Adresse ein, und
der Kernel merkt sich jeden laufenden Strom samt dieser Adresse. Bekam der
Container beim Neubau eine andere — und er bekam die naechste freie —, zeigten
alle gemerkten Stroeme ins Leere. Und sie verfallen nicht, weil WireGuards
Lebenszeichen sie alle 25 Sekunden auffrischt.

Deshalb kam ein Telefon nach Aus/Ein sofort zurueck und ein Host mit festem
Quellport ueberhaupt nicht.

Jetzt hat vpn-hub eine feste Adresse (172.18.0.240, ueber CLUPILOT_VPN_HUB_IP
aenderbar) in einem Netz mit erklaertem Subnetz. Nach einem Neubau entsteht exakt
dieselbe Weiterleitung, die gemerkten Stroeme bleiben gueltig, und es gibt nichts
mehr aufzuraeumen. Nachgemessen: 172.18.0.240 vor und nach
`up -d --force-recreate vpn-hub`.

UND DIE UMSTELLUNG SELBST, die mich beim Bauen fast den Stapel gekostet haette:
Docker kann ein bestehendes Netz nicht umdefinieren. Es muss neu angelegt werden,
und das scheitert, solange auch nur EIN Container daranhaengt — `up -d` bricht
dann mit "network … has active endpoints" ab und laesst den Stapel halb unten
stehen. Genau so hier passiert, mit einem Container, der gar nicht zum Stapel
gehoerte. Das Deployment vergleicht deshalb vorher das erklaerte Subnetz mit dem
tatsaechlichen und faehrt bei Abweichung EINMAL geordnet herunter, statt darueber
zu stolpern. Danach stimmen die Werte ueberein und der Block tut nie wieder etwas.

Das conntrack-Aufraeumen bleibt trotzdem drin: als Rueckfahrkarte fuer Server auf
aelterem Stand und fuer den Fall, dass jemand das Subnetz aendert.

2510 Tests gruen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 09:10:54 +02:00
nexxo e3c781f067 Der Tunnel bekommt einen eigenen Container
Bisher wohnte wg0 im Container des Provisionierungs-Arbeiters, also im Abbild
`clupilot-app`. Das wird bei fast jeder Freigabe neu gebaut — und ein neu
gebauter Container bekommt eine neue Adresse im Compose-Netz, womit die
Weiterleitung fuer UDP 51820 neu geschrieben wird und JEDE bestehende
WireGuard-Sitzung abreisst. Der Tunnel hing damit am Veroeffentlichungstakt der
Anwendung, und zusaetzlich daran, dass ein PHP-Prozess nicht abstuerzt.

Jetzt gehoert der Netz-Namensraum einem eigenen Container `vpn-hub` mit eigenem,
winzigem Abbild (Alpine plus wireguard-tools), das sich fast nie aendert.
Provisionierungs-Arbeiter, Terminal-Bruecke, interner DNS und internes Gateway
steigen dort ein, statt einer von ihnen den Namensraum zu besitzen.

NACHGEMESSEN, nicht angenommen: App-Abbild neu gebaut, Arbeiter, Bruecke und
Gateway per --force-recreate neu erzeugt — der Hub blieb Container 92e928cf53b0,
wg0 und beide Zugaenge unangetastet, und nginx erreichte die Bruecke weiter
(HTTP 426). Genau der Vorgang, der bisher jedes Mal alles abgerissen hat.

Nachgezogen:
- nginx spricht die Bruecke unter `vpn-hub:8082` an — dem Namen des
  Namensraum-Eigentuemers; ein Mitbewohner hat keinen eigenen DNS-Eintrag.
- update.sh baut vpn-hub mit und haengt Nachbar-Neustarts und den
  conntrack-Griff an die Frage, ob der Hub WIRKLICH neu gebaut wurde.
- update-agent.sh startet den Arbeiter nicht mehr neu, sondern signalisiert ihm.
  Diese Stelle laeuft unbeaufsichtigt hinter dem Knopf „Dienste neu starten" —
  wer den drueckt, rechnet nicht damit, sich selbst auszusperren.
- Vier Meldungen in der Konsole rieten dem Betreiber, genau den Befehl von Hand
  auszufuehren, der ihm den Tunnel abreisst. Auch die sind korrigiert.
- rescue-tunnel.sh und das Runbook zeigen auf den neuen Besitzer.

Eine Kleinigkeit unterwegs, die ich falsch angekuendigt hatte: `[[ … ]] && x=true`
bricht unter `set -e` NICHT ab — bash nimmt die linke Seite einer &&-Liste
ausdruecklich aus. Nachgeprueft; die if-Form bleibt trotzdem, aus Lesbarkeit, und
der Kommentar sagt jetzt den wahren Grund.

2509 Tests gruen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 07:49:29 +02:00
nexxo 4b3f1bb4ff Ein Update fasst den Tunnel nicht mehr an
Der Betreiber hat es dreimal hintereinander erlebt: Update gefahren, VPN weg.
Die Ursache dafuer, dass es JEDES Mal passierte, stand hier:

    docker compose restart queue queue-provisioning scheduler reverb

In dessen Netz-Namensraum lebt wg0. `restart` baut den Namensraum neu auf, also
riss jede WireGuard-Sitzung ab — die des Betreibers am Telefon wie die jedes
Hosts. Fuer ein Update, das nur PHP-Code aendert, ist das ein absurder Preis.

Seit der Arbeiter dort in einer Schleife laeuft (v1.4.4), geht es billiger:
`queue:restart` setzt ein Signal, der Arbeiter beendet sich nach dem laufenden
Auftrag, die Schleife startet ihn mit dem neuen Code neu. Der Container bleibt
stehen, wg0 bleibt oben, niemand merkt etwas.

Und fuer den Fall, dass `up -d` ihn doch neu baut (neues Abbild, geaenderte
Konfiguration): das Skript merkt sich die Container-ID vorher und nachher. Hat
sie sich geaendert, raeumt es die gemerkten UDP-Stroeme selbst weg —
`sudo -n conntrack -D -p udp --dport <port>`, und wenn es das nicht darf, steht
der Befehl als Warnung im Protokoll statt gar nichts.

Dieser Handgriff war bisher muendliche Ueberlieferung. Ohne ihn zeigen die
gemerkten Stroeme weiter auf den alten Container, und sie verfallen nicht:
WireGuard schickt alle 25 Sekunden ein Lebenszeichen und haelt den kaputten
Eintrag am Leben. Genau deshalb kam ein Telefon nach Aus- und Einschalten sofort
zurueck (neuer Quellport) und ein Host mit festem Port ueberhaupt nicht.

Ein Test haelt beides fest: queue-provisioning darf nicht in der Neustart-Liste
stehen, und der conntrack-Griff muss im Skript bleiben.

2509 Tests gruen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 07:09:12 +02:00
nexxo 4fc1ccd3db Terminal: die Bruecke ueberlebt jetzt ein Deployment
Aus dem Gesamt-Review, und der erste Befund ist der, der still zugeschlagen
haette.

`terminal` lebt im Netz-Namensraum von `queue-provisioning`, und ein Prozess
bleibt in dem Namensraum, in dem er gestartet ist. `update.sh` startet den Hub
neu — danach lauscht die Bruecke in einem, den es nicht mehr gibt. Nichts meldet
dabei einen Fehler: `docker compose ps` sagt weiter "healthy", weil die
Lebendpruefung ueber Loopback INNERHALB des verwaisten Namensraums laeuft. Nach
aussen antwortet nginx mit 502, und der Betreiber liest "Keine Verbindung —
laeuft der Terminal-Dienst?", waehrend der Dienst behauptet, es gehe ihm gut.
Genau dieselbe Falle, die zwei Bloecke tiefer schon fuer vpn-dns/vpn-gateway
behandelt ist; die Bruecke fehlte in der Behandlung.

Nachgemessen statt geglaubt: Hub neu gestartet -> Docker sagt "healthy", curl aus
dem Namensraum bekommt gar keine Antwort. Nach `restart terminal`: 200.

Und ein zweiter Ausrollfehler daneben: gebaut wurde nur `app`. `docker compose
up -d` baut nur Images, die es noch GAR NICHT gibt — beim ersten Ausrollen faellt
das nicht auf, danach nie wieder. Eine Aenderung an docker/terminal/ saehe
ausgeliefert aus, und es liefe das alte Image.

Ausserdem:
- Die Meldung zu 4502 zaehlte zwei Ursachen auf, der Code deckt fuenf. Die
  Bruecke schickt 4502 fuer JEDE gescheiterte Anmeldung, auch fuer einen
  abgewiesenen Schluessel — und das ist der wahrscheinlichste Fall, wenn ein Host
  neu aufgesetzt wurde. "antwortet nicht" war dort schlicht falsch: die Maschine
  hat geantwortet und abgelehnt. Titel und Text legen sich nicht mehr fest.
- R19: der Kommentar an der Kopfzeile der Spalte nannte "Berechtigung,
  Betriebsbereitschaft" als Grund, warum der Knopf nicht ueberall steht.
  Letzteres entscheidet seit dem Entsperren nichts mehr, und zwanzig Zeilen
  tiefer begruendete der Kommentar am Knopf ausfuehrlich das Gegenteil.
- Der Kommentar am Retry-Knopf erklaerte die Reihenfolge von .hidden gegen
  .inline-flex fuer zu unsicher, waehrend die Buehne dreissig Zeilen hoeher genau
  darauf baut. Tailwind gibt .hidden als letzte Display-Klasse aus; der wahre
  Grund fuer den Wrapper ist, dass die Klassen des Knopfes aus einem geteilten
  Bauteil kommen.
- REDIS_URL stand fest auf Datenbank 1, waehrend PHP REDIS_CACHE_DB liest. Wer
  die anfasst, legt auf der einen Seite ab, wo die andere nicht sucht.

2507 Tests gruen, compose config und bash -n sauber.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 19:41:33 +02:00
nexxo 07c51474d3 Update drehte sich im Kreis: es ersetzte sich selbst mitten im Lauf; dazu eine Seite fuer offene Punkte 2026-08-02 02:02:25 +02:00
nexxo acc4193071 www-data gehoerte sein eigenes Heimatverzeichnis nicht: npm ci brach das Deployment mit 243 ab 2026-08-02 01:03:23 +02:00
nexxo 006ce3568b Manage hostnames and watch their certificates from the console
files.clupilot.com is why this exists. DNS pointed at the right machine, the
env var was set, the release was deployed — and there was still no certificate,
because /etc/caddy/Caddyfile is maintained by hand and nobody thought of it as a
second, separate step. Nothing in the console would have said so. Three settings
looked correct and the address did not answer.

So the page holds two things side by side. The WISH — which names should be
served — and the REALITY: whether the name has a certificate and for how much
longer. The second is measured by opening a TLS connection and reading the
expiry, not by reading configuration, because the configuration is exactly what
looked right while the address was dead. verify_peer stays on: a certificate
that fails validation is not a certificate for this question, and a display that
called it valid would be the fake R19 records.

Applying goes through the existing agent, not a new channel. The console writes
a request, the path unit wakes the agent within a second, and the agent calls one
fixed command line of the root-owned helper.

What that helper is allowed to do is the careful part. It fetches the list
ITSELF rather than being handed one, and the list is HOSTNAMES, never Caddy
blocks — `php artisan clupilot:proxy-hosts` prints `<name> <purpose>` and nothing
else. Each name is matched against a strict pattern before it is used, and the
template around it lives in the helper, which the service account cannot touch.
install-agent.sh already states the principle for the sudoers grant: a grant is
only worth anything if the holder cannot change what it grants. A service account
that could write proxy configuration would have everything the proxy can do —
redirects anywhere, files from any directory.

Purpose is a column rather than a habit. A console name gets the network lock,
a public one does not, and a console name published without it looks exactly
like a working page.

Removing takes the name out of the list and NOT out of the running proxy. Two
decisions in one click, and the second one takes a site off the air.

There is deliberately no "renew" button. Caddy renews on its own at two thirds
of the lifetime; what an operator actually needs is a second attempt after an
issuance has failed, and that is a reload — which is what Apply does. Thirty days
is treated as a problem rather than a warning: at ninety days' lifetime, a
renewal should long since have run, so anything under it is not a tight
certificate but a renewal that is not happening.

The ACME contact falls back to the owner's address, because a contact nobody
reads is the step before expired customer certificates.

CONTRACT and HOST_STEP_NEEDS both move to 2, which is what tells a server
carrying the older helper to run the installer again — caught by the guard test
that compares the two halves.

2035 tests pass, assets build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 22:55:55 +02:00
nexxo a4626d2569 Stop root workers breaking every page, and let the panel be closed
tests / pest (push) Failing after 12m38s Details
tests / assets (push) Successful in 26s Details
tests / release (push) Has been skipped Details
── The 500 ─────────────────────────────────────────────────────────────────
"touch(): Utime failed: Operation not permitted", from BladeCompiler, on
every request for the affected view.

Blade compiles a view and then touch()es the compiled file to the source's
mtime. touch() with an explicit time needs ownership of the file. A compose
service without `user:` runs as ROOT — and queue, scheduler and reverb had
none — so the first worker to render a view wrote a root-owned compiled file
that the web process, as www-data, could never refresh again. The queue is
what renders mails, which is why this surfaced now.

queue, reverb and scheduler run as www-data. queue-provisioning stays root
and says why: it brings wg0 up and runs `wg set` for every peer change, which
needs NET_ADMIN on the running process. It renders no mail.

And update.sh normalises ownership at the END of a run as well as at the
start. The first call heals what a previous run left; the second heals what
this one made — `git checkout` rewrites the tree as the service account while
the old build is still serving.

── The panel that would not go away ─────────────────────────────────────────
The same 500 is why it kept coming back: every poll failed, the watcher
treats a failed request as "still restarting" (which it normally is), and the
overlay stayed up over a console nobody could then reach to find out why.

A restart is seconds. After two minutes of nothing the panel now says so and
offers a way out — for an operator only, behind the same flag as the step and
the log, so a customer on the 503 page is not shown a button suggesting they
can call the deployment off. A successful answer clears it again: one bad
minute must not leave a "something is wrong" notice sitting there for the
rest of the run.

Not an always-present close button. The deployment does not stop because
somebody dismissed a panel, and offering that at the wrong moment is a lie.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 15:43:54 +02:00
nexxo f8f0de4892 Let an update install host packages, without handing out root
Asked for directly: "später wenn ich etwas lokal erweitere will ich mich
nicht auf alle adminseiten einloggen müssen und extra rsync installieren".

An update runs as the service account, so anything root-owned on the host is
out of its reach and turns into "log into every server once and paste this".
The obvious fix — sudo on deploy/install-agent.sh — is not a grant at all:
the service account owns the checkout and can rewrite the very script it
would be allowed to run as root.

So the privileged part is written OUT of the checkout by install-agent.sh, to
/usr/local/sbin/clupilot-host-step, root:root 0755, from a quoted
here-document. The service account cannot influence a byte of it. sudoers
names one exact command line including its argument, so a step added to the
helper later is not covered by a grant written before it existed, and the
helper refuses anything not on its own list as well.

update.sh uses it only when rsync is actually missing, after the restart and
never fatally: the archive is collected hours later, and an update that died
over a package would be the bigger problem. A helper older than the steps the
updater wants is reported through the existing one-time-setup hint — silently
doing nothing is worse than an error.

Existing servers still need `sudo bash deploy/install-agent.sh` once, because
that is the run that puts the helper there. After it, they do not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 10:27:56 +02:00
nexxo 6998a22527 Repair ownership by looking inside, not at the door
tests / pest (push) Failing after 8m2s Details
tests / assets (push) Successful in 21s Details
tests / release (push) Has been skipped Details
The ownership repair added in 1.1.1 sampled the owner of each top-level
directory and skipped the recursion when it already matched. node_modules was
owned by www-data; node_modules/.vite-temp, left behind by an earlier root
build, was not. So the repair walked past it, `npm run build` failed with
EACCES, and the deployment stopped in maintenance mode — on the release whose
whole point was that this could not happen.

A directory's owner says nothing about its contents. It now walks each tree
once with find and changes only the entries that are wrong, which is also
cheaper than the detect-then-chown-everything it replaces.

Verified against the shape that actually failed: a root-owned file inside a
www-data-owned node_modules, repaired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 23:53:25 +02:00
nexxo 00b9107ee3 Deploy as the application's user, and take back what root already took
tests / pest (push) Failing after 7m19s Details
tests / assets (push) Successful in 19s Details
tests / release (push) Has been skipped Details
The 500 on the VPN config download was never in the VPN code. storage/logs/
laravel.log was owned by root, mode 644, since 27.07 04:16 — the moment a
failed `artisan optimize` wrote its error there during a deployment. From then
on the application could not append to its own log: Monolog threw on every
attempt, and a throw while logging is a 500 on any page that logs, with nothing
written down to say why. The error page said "Er wurde protokolliert". It was
not.

It surfaced on the config download because that is one of the few pages writing
a log line on its way through — on success and on a wrong password alike. Pages
that log nothing were unaffected, which is exactly why it looked like a VPN
fault for days.

`docker compose exec` is root unless told otherwise, and every deployment
script relied on that default — including the agent's allowlist sync, which
runs every minute. docker/entrypoint.sh had it right all along and drops to
www-data for precisely these commands. Now so does everything else, help text
included: telling an operator to run artisan as root is how the file ends up
owned by root in the first place.

update.sh also repairs what an earlier run left behind, before its first
unprivileged step rather than after — a root-owned log file is not
self-healing, the page that trips over it is nowhere near the deployment that
caused it, and with in_app unprivileged a root-owned vendor/ would break the
next composer step outright.

tests/Feature/DeploymentRunsAsTheAppUserTest.php holds the line. It was checked
against the previous commit and finds all nine places that were wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 22:52:09 +02:00
nexxo b85c12e8b9 Give the readiness probe long enough for a restart to finish
tests / pest (push) Successful in 7m13s Details
tests / assets (push) Successful in 19s Details
tests / release (push) Successful in 3s Details
The gateway is restarted twice in a deployment — once by `up -d`, then again
after the hub whose network namespace it lives in — and `docker compose exec`
into a container that is still coming up fails outright rather than waiting.
Twenty seconds did not cover that, so the first run of the new probe reported a
gateway that answers as down, and the resolver stayed out of client configs for
one more deployment.

The health site also now binds the hub address instead of only matching on it.
Nothing was published either way, but the file should do what its comment says.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 14:46:47 +02:00
nexxo 3699c7cdb9 Ask the tunnel gateway where it actually listens
tests / pest (push) Successful in 7m15s Details
tests / assets (push) Successful in 20s Details
tests / release (push) Successful in 5s Details
The console has been unreachable over the VPN, and the cause was the readiness
probe rather than the tunnel. The gateway binds to the WireGuard hub address
alone; the probe asked 127.0.0.1, where nothing has ever listened. It could only
fail, so VPN_READY stayed false, and on that basis the application withheld the
`DNS = 10.66.0.1` line from every client config it issued. A phone then built
the tunnel, asked its normal resolver for the console hostname, got the public
address and had its connection closed — indistinguishable from "this site does
not exist", which is exactly what it looked like.

Probing over TLS on the bare address would not have worked either: the site
matches on the console's hostname, so a request without SNI is offered no
certificate. The gateway now answers a plain-HTTP health port on the hub
address, which removes TLS, SNI and name resolution from the question and
answers only what is being asked — is this gateway listening, in this network
namespace, right now. Caddy refuses to start when the certificate is unreadable,
so a health port that answers still proves the whole file loaded.

Also here, all found while looking:

- Icons pushed their label onto a second line and rendered a size larger than
  asked for. Tailwind's preflight makes an svg display:block, and `.size-4` and
  `.size-5` have equal specificity, so stylesheet order decided — and it emits
  size-4 first. Every icon written as 16px was silently 20px. Recorded as R18.
- Four error pages printed `errors.404.hint` in place of a sentence: the lang
  files give those codes a null hint and Laravel returns the key for a null line.
- The Developer role had no label in either language, so the dropdown showed
  `admin_settings.role_developer`.
- The secrets area held one key, which is not worth a password gate. It now
  carries the credentials that actually stop the business when they expire —
  DNS, monitoring, SMTP — and the test button appears only where a checker
  exists, instead of reporting on Stripe whatever was being looked at.
- The update button said nothing about when a queued run would start or where a
  running one had got to; both are now shown, and a failure names its step.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 14:43:23 +02:00
Claude 2bb1d7cc3b Make the private hostnames look like nothing is there, and close the way past the proxy
tests / pest (push) Successful in 6m56s Details
tests / assets (push) Successful in 27s Details
tests / release (push) Successful in 4s Details
An empty 404 is itself information: it says something terminates TLS here and
chose not to answer. The console and the websocket endpoint now close the
connection instead, which is what a hostname that serves nothing looks like.
The websocket endpoint answers a genuine upgrade — verified live, 101 — and
nothing else; a browser opening it gets a closed connection.

Reviewing that turned up two holes that mattered more than the thing being
reviewed.

The compose defaults published the application and Reverb on every interface.
Docker publishes ports ahead of UFW, so those backends were reachable from the
internet with the firewall closed — and reaching a backend directly skips the
proxy, and with it every hostname and address rule keeping the console private.
The defaults are loopback now, and an update rebinds an existing installation
that is behind a proxy actually running and holding 443. A development box
without one is left alone, because taking its port away would look like the
machine had broken.

And the console's own "open" switch was emitting 0.0.0.0/0 into the proxy's
allowlist. One click would have put the console on the public internet, which
is the exact state the owner's rule exists to prevent. Switching it off relaxes
the application's check; the proxy keeps its list, always.

The reference proxy config carries the arrangement: QUIC early data off, since
an address-based decision taken on 0-RTT can answer 425 and some browsers do
not retry — an intermittent lockout from the console is the worst kind — and
explicit handling for plain HTTP, whose automatic redirect otherwise announces
both hidden names to anyone who asks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 11:26:48 +02:00
Claude c4649571df Wait for the tunnel interface, and prove the gateway answers
tests / pest (push) Successful in 7m23s Details
tests / assets (push) Successful in 25s Details
tests / release (push) Successful in 5s Details
Two things reported success while the tunnel console was unreachable.

The services were restarted the moment the hub container counted as started,
but the hub brings wg0 up as part of its start command — so they were binding an
address that did not exist yet. They now wait for the interface.

And readiness was a container-state check. A container can be up and running
while the process inside it listens in a network namespace that was torn down
underneath it: nothing errors, the address simply refuses connections, and
"running" reports everything fine. It asks the gateway now, and only a real
answer counts — because the consequence of getting this wrong is a client
config naming a resolver that is not there, which takes the device's whole name
resolution with it for as long as the tunnel is up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 10:24:29 +02:00
Claude d2d79b4f5f Restart the tunnel services after the hub they live inside
tests / pest (push) Successful in 10m42s Details
tests / assets (push) Successful in 18s Details
tests / release (push) Successful in 6s Details
Both VPN services share the provisioning container's network namespace, and a
process holds the namespace it started in. Restarting the hub — which every
deploy does, to pick up new code — therefore leaves them listening inside a
namespace that no longer exists.

Nothing errors. The containers stay up, Caddy's log still reports it is serving
on the tunnel address, and connections to that address are simply refused. It
looks precisely like a gateway that never worked, which is how it presented on
the live server after the first successful start.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 10:22:19 +02:00
Claude 457eeeeaef Serve the console inside the tunnel, without publishing the internal network
tests / pest (push) Successful in 10m12s Details
tests / assets (push) Successful in 21s Details
tests / release (push) Successful in 5s Details
The owner's phone shows a blank page on the VPN. That page is the reverse proxy
refusing the request, and it refuses because the phone is not on the tunnel: the
client config routes only the management subnet, so a request to the console's
public hostname leaves over the mobile network and arrives from a carrier
address. It is also why the peer reads "last contact: never" — WireGuard only
handshakes when it has traffic to send, and nothing is ever routed in.

The obvious fixes are both wrong. Adding the server's public address to
AllowedIPs routes the WireGuard endpoint into the tunnel it is trying to
establish. A public DNS record pointing at 10.66.0.1 publishes the internal
subnet to anyone enumerating the domain — the owner raised that himself, and he
was right.

So the console is served INSIDE the tunnel instead. A small resolver and a small
gateway share the hub container's network namespace, which is where wg0 lives,
and answer on 10.66.0.1 directly. A client therefore needs no route beyond the
subnet it already has, the host's docker bridge is never exposed, and the
request reaches the application with its real 10.66.0.x source — which is what
the console's own allowlist checks. The gateway reuses the certificate the
public Caddy already renews: a certificate is bound to the name, not to the
address serving it, so nothing new is issued and nothing internal reaches a
public zone.

Most of this commit is the arithmetic of not lying about it. Review found
fourteen ways the arrangement could report itself working while it was not:
enabling the compose profile after the services were started rather than before;
starting a gateway with empty certificate paths, which can only crash; treating
"container created" as "container running"; carrying a stale readiness marker
through a failed restart; reusing the previous hostname's certificate after a
rename; leaving the profile enabled when the hostname was cleared, in a script
that exits early precisely when nothing else would fix it; reading the renewal
timestamp as a user who cannot traverse Caddy's storage; and a `find | head`
that aborts the whole installer under pipefail the moment two issuers hold a
certificate for the same name.

One of them was mine and worth naming: to let the container read the
certificate, I had made the TLS private key world-readable. On a multi-user host
that hands the console's identity to every local account.

The client is told about the resolver only when both services are confirmed
running, because a config naming a resolver that does not exist takes the
device's entire name resolution with it for as long as the tunnel is up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 10:14:33 +02:00
Claude bc8bbc56a5 Make the console's access list reach the proxy, price in euros, and answer errors
tests / pest (push) Successful in 7m4s Details
tests / assets (push) Successful in 19s Details
tests / release (push) Successful in 4s Details
The console was unreachable, and not for the reason it looked like. Two causes,
both mine.

A failed `artisan optimize` — the one that hit duplicate route names — left a
broken route cache behind, so every request to the console host answered 500.
Rebuilt.

Underneath that: the reverse proxy has its own console allowlist, hard-coded,
and it runs BEFORE the application. Everything added on the console's own
access page was therefore ineffective, and when the owner's address changed
they were turned away by the proxy before reaching the page that would have
fixed it. The proxy now imports a fragment generated from the same list the
console manages, regenerated by the host agent and reloaded when it changes.

Getting that safe took most of the review. It refuses to rewrite an ambiguous
Caddyfile rather than replacing some other site's matcher; it never falls back
to a loopback-only list when the application cannot be reached, because that
list validates cleanly and locks out every remote operator; it retries a reload
that failed instead of assuming it worked; and an installation that upgrades
without rerunning the installer is told, because otherwise the whole mechanism
is invisibly absent.

Prices are entered in euros. The form asked for cents, so €799 was typed as
79900 and one slipped digit was a factor of ten on an invoice. Conversion
happens in one place, on the string rather than through a float — (float)
'79.90' × 100 is 7989.999… and casting truncates to 7989, one cent short on
exactly the prices people charge — and it refuses an amount the column cannot
hold instead of failing at the database.

Every error code has a page now, in the site's own language and typeface,
self-contained so it still renders when the asset manifest is the thing that
broke. There was only a bare white 404.

The VPN list stops being a seven-column table nobody could fit: names broke
across two lines and so did the headings. These are attributes of one access,
not quantities compared down a column, so each access is a row — identity on
one line, measurements in mono on the next.

The status page says what it measures rather than what it promises. "New orders
are delivered without failures" read as a marketing claim on a page whose only
job is to be believed, and said nothing about the last 24 hours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 06:51:05 +02:00
Claude 276f926d57 Add the update button, and give it a host-side agent that can actually update
tests / pest (push) Successful in 7m41s Details
tests / assets (push) Successful in 20s Details
tests / release (push) Successful in 4s Details
The panel had no way to update the installation, and could not have one on its
own: it runs as www-data inside a container, and the update restarts that
container. A button that shelled out would kill its own web server
mid-response, and would need the application to hold host-level credentials —
a far worse thing to own than an out-of-date checkout.

So the button is a mailbox. The panel writes a request into the checkout; a
systemd timer on the host, running as the service account, consumes it, runs
deploy/update.sh and writes back what happened. The same timer answers the
question the panel cannot answer alone — is there anything to update — which
needs a git fetch, and therefore credentials the application does not have.

The parts that were wrong before review, all of which would have shipped as a
button that looks fine and does nothing:

- The agent was only installed by install.sh, which the installed base never
  runs again. Every existing server would have shown the button permanently
  disabled. It has its own root entry point now, install.sh delegates to it,
  and update.sh says so when the unit is missing.
- On a release-pinned server the checkout is detached, so it followed
  "origin/HEAD" and then ran an update that deliberately stays on the pinned
  tag: updates advertised, nothing applied. It now compares TAGS in release
  mode and passes the target release through.
- On a branch other than main it advertised that branch's commits and then
  deployed main's, because update.sh defaults to main.
- A request written while the agent was down was executed whenever the agent
  next started, days later.
- A single status file meant the routine five-minute check overwrote a failed
  update with "idle" before anyone saw it.
- An agent that had been stopped left the button enabled forever, because a
  status file written once counted as an agent.
- The agent's error messages were German strings rendered into an English
  interface; it reports codes now and the panel translates them.

Also: the add-address button in the console-access panel was as tall as the
field plus its hint, because the hint was a sibling inside the flex row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 05:33:17 +02:00
nexxo e9240c1324 feat(deploy): releases you can pin to, and a version that tells the truth
VERSION at the repository root is the release number and the only place it is
written down — a file rather than `git describe`, because a source archive, a
shallow CI checkout and a container built without .git have no tags to
describe. Cutting a release is editing that file and merging it; CI creates the
annotated v<VERSION> tag once green, only on the commit that raised the number,
and never moves it. Servers are pinned to these.

Two modes, and the difference is the point. RELEASE=v1.0.0 pins a server to an
immutable tag, checked out detached: it moves when someone decides it moves.
BRANCH=main follows the edge, where CI tags a commit only after it is pushed.
The mode is remembered, so re-running the updater on a pinned box does not
quietly walk it back onto main.

Going backwards is refused, by commit ancestry rather than by comparing version
strings — a tag can be cut from anywhere, and only ancestry says whether history
is going back. The schema has already moved forward by then, and the migrations
needed to reverse it are not in the older checkout at all.

What the console reports comes from a manifest written atomically after every
step succeeded, never from live git: git says what the files are, not whether
the deployment came up. So a failed update keeps reporting the version that is
actually serving — including its version number, not the newer one the checkout
has already moved to. A deployment pinned to the tag reads "1.0.0 (abc1234)";
anything else reads "1.0.0-dev (abc1234) · main", because every commit after the
tag still carries VERSION=1.0.0 and is not that release. Reporting it as one is
how a bug gets filed against the wrong code.

Codex reviewed the design before it was written and the result six times after.
Its findings, all real: the version had to come from the manifest too, not just
the commit; branch names need JSON escaping or the manifest silently vanishes; a
`case` glob does not anchor and accepted 1x.2y.3garbage; pinning to the commit a
server already sits on skipped recording the mode, and then skipped detaching;
a manifest that failed to write would never be repaired; an unparseable
timestamp would 500 every admin page; and an empty tag list read as a
connectivity failure.

466 tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 15:21:38 +02:00
nexxo 8117630a65 feat(ops): update page, automatic 419 recovery, CI workflow
Three things that bit us on every deploy:

- artisan down now renders a branded page that says an update is running and
  reloads itself, instead of Laravel's bare 503.
- A page left open across a deploy carries a Livewire snapshot and CSRF token
  the new code rejects — the user got "419 Page Expired" and a dead interface.
  A 419 from a /livewire/ request now reloads the page; the session is still
  valid, so that is all it takes. Hooked at the fetch layer rather than through
  Livewire's request hook, whose failure callback is not invoked for this case
  in the installed version — verified against a real 419 in the browser, not a
  simulated one.
- Vite no longer empties its output directory: wiping it is what left an open
  page without its stylesheet mid-deploy, which is the "design is completely
  broken" symptom. update.sh also rebuilds the caches with optimize:clear +
  optimize instead of leaving them half-warm.

Plus CI: .gitea/workflows/tests.yml (Pest + asset build, and a tested- tag only
on a green main) and an opt-in act_runner service under the ci profile.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-26 00:58:27 +02:00
nexxo 2415dde288 fix(deploy,traffic): clear caches before restarts, hide the token, widen the lock
- update.sh restarted the long-lived services before clearing the caches, so
  they booted from the old cached config and kept it for the life of the
  process — the release's configuration never reached the workers.
- install.sh put the Gitea token in the clone URL, which is visible in the
  process list to every local user while the clone runs. It goes through a
  short-lived askpass helper now.
- The collector's uniqueness lock expired at exactly the collection interval,
  so a delayed queue could run two collectors on the same baseline and count
  the same delta twice — throttling a customer for traffic they never used.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 23:42:12 +02:00
nexxo 25c4c5df5a fix(deploy): restart the worker after the hub key, diff against what was deployed
- The provisioning worker boots before install.sh writes
  CLUPILOT_WG_HUB_PUBKEY and holds the empty value for the life of its process,
  so every host onboarded afterwards would have got an empty PublicKey in its
  wg0.conf.
- A failed update left the checkout at the target, so the rerun compared a
  commit with itself and skipped exactly the dependency and image steps that had
  not finished. The comparison base is now the last commit actually deployed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 23:38:32 +02:00
nexxo 4681b135db feat(deploy): run under a dedicated account instead of root
The installer still needs root for apt and the firewall, but it now creates a
clupilot service account that owns the checkout and runs Docker, maps the
container user to it (HOST_UID/GID), and update.sh refuses to run as root —
root-owned files in the checkout are files the application cannot write.

The account has no password login: docker group membership is root-equivalent
on the host, so it is reachable only via sudo -u or an SSH key.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 23:25:59 +02:00
nexxo 0219e3c987 fix(deploy): recreate the container, keep the root password secret, forget failures
- update.sh recreated the app container only at the very end, so an update that
  changes the PHP runtime would have installed dependencies and migrated inside
  the old one.
- install.sh randomised DB_PASSWORD but left DB_ROOT_PASSWORD at the value
  committed in .env.example — the same root credential on every installation,
  reachable from any container on the compose network.
- Settings cached the fallback after a database failure, so one blip could
  leave a hidden site publicly visible until someone cleared the cache. Only a
  successful read is cached now, proven by a test that drops the table and
  brings it back.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 23:25:06 +02:00
nexxo 773ef4bd5f fix(deploy): migrate behind maintenance mode, resume a failed update, install deps
Codex was right on all three, and the first was a design error on my part: the
checkout is bind-mounted into every container, so 'git merge' swaps the running
code instantly — 'migrations before traffic' was not achievable by ordering.
The updater now goes down → merge → dependencies → migrate → assets → restart →
up, trading a short deliberate outage for never serving code whose schema does
not exist yet. If a step fails the site STAYS down: coming back up with new code
on an old schema is worse than staying dark until someone looks.

It also records the last successfully deployed commit. A run that died halfway
left the checkout ahead, so the next run said 'already up to date' and skipped
the rest forever.

And it installs dependencies when a lockfile moved: vendor/ and node_modules/
live in the bind mount and shadow the image, so rebuilding the image never
updated them — migrations could run against stale packages.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 23:15:15 +02:00
nexxo 834abcec40 feat(deploy): installer and updater for a fresh server
install.sh sets up a bare Debian/Ubuntu server end to end: Docker, git,
WireGuard, clone, generated secrets, stack, migrations, hub keypair, the Owner
account and a closed firewall. Re-runnable: it keeps an existing .env and never
regenerates the hub key, which would disconnect every onboarded host.

update.sh pulls and applies. Migrations run before the new containers take
traffic, the image is rebuilt only when its definition changed, and the queue
workers are restarted — they are long-running processes that otherwise keep the
old classes in memory, which cost us an hour during development.

clupilot:create-admin creates or promotes an Owner, so re-running the installer
fixes a lost role instead of failing on a taken address.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 23:13:24 +02:00