Commit Graph

520 Commits (e031d4c20af9f67f77a5d3c00a6e81975c6755f8)

Author SHA1 Message Date
nexxo f8d0c2c353 Merge branch 'claude/angry-golick-187172' into feat/neue-pakete 2026-08-01 16:27:13 +02:00
nexxo ce7d0426cb Fix-Runde: preferredDatacenter() blind fuer Reservierung, Loeschen hob sie auf
preferredDatacenter() waehlte das Rechenzentrum ueber Host::availableGb() ohne
->unreserved() - ein Rechenzentrum mit einem grossen, aber komplett
reservierten Host sah geraeumiger aus als eines mit echtem allgemeinem
Bestand. Der Checkout haette die Bestellung dorthin gelegt, placeableIn()
haette dort niemanden gefunden, und sie waere geparkt, obwohl anderswo Platz
war. Jetzt ->unreserved(), derselbe Bestand wie largestPlaceableGb().

Die Migration nutzte nullOnDelete() und tat damit das Gegenteil der eigenen
Vorgabe: die Reservierung sollte einen Kundenaustritt nicht stillschweigend
ueberleben, loeste sich mit nullOnDelete() aber genau so auf, sobald der
Kunde verschwindet. restrictOnDelete() macht das Loeschen eines Kunden mit
eigener Maschine zum Fehler, bis ein Operator die Reservierung von Hand
gelöst hat - Migration und Modellkommentar sagen jetzt dasselbe.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:21:46 +02:00
nexxo 3df7434f7a Fix-Runde 1/1: Enterprise-Schwelle folgt dem Katalog statt einer festen Zahl
Zwei Betreiber-Entscheidungen aus der Durchsicht:

1. Keine zweite, feste Untergrenze mehr im Fliesstext ("ab 500 GB") -- die
   eigene Maschine beginnt dort, wo das groesste Paket endet, ihre Groesse ist
   Hardware-Frage einer Angebotsanfrage, kein zweiter Wert auf dem Preisblatt.
2. Die Zahl in der Ueberschrift ("Mehr als :quota?") kommt jetzt aus $plans --
   derselben Liste, die baseline() und comparison() schon lesen, statt fest im
   Sprachtext zu stehen. Eine spaetere Umschaltung der Paketleiter (Task 9)
   aendert sonst, was "am groessten" ist, und der Satz haette es nicht gemerkt.

max() auf einer leeren Plan-Liste (Katalog nicht lesbar) waere ein Fatal
gewesen -- enterprise() gibt in diesem Fall jetzt null zurueck, mit Test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:15:59 +02:00
nexxo 4ad1828f27 Ein reservierter Host gehoert seinem Kunden
placeableIn() nahm jeden aktiven Host im Rechenzentrum, und hasRoomFor() sowie
largestPlaceableGb() zaehlten eine exklusiv verkaufte Maschine obendrein mit —
der Shop versprach damit Platz, der bereits vergeben war. Ohne diese Markierung
ist 'eigener Server' ein Satz im Angebot und nirgends eine Tatsache.

hosts.reserved_for_customer_id (nullable, ueberlebt den Kunden) markiert die
Maschine; Host::placeableIn() und HostCapacity zaehlen sie nur noch fuer den
eigenen Mieter oder gar nicht mehr zum allgemeinen Bestand. Auf der
Host-Detailseite kann ein Operator reservieren und wieder loesen — Loesen
laeuft ueber ein eigenes Bestaetigungsmodal (R23), das selbst nichts aendert,
sondern nur an HostDetail::releaseReservation() zurueckmeldet. Host-Liste und
Kapazitaetsseite weisen eine reservierte Maschine als solche aus, statt sie
kommentarlos verschwinden zu lassen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:06:14 +02:00
nexxo 1692fd3977 Enterprise wird angefragt, nicht ausgepreist
Ab 500 GB ist es eine eigene Maschine — im Shop faende placeableIn() auf 388 GB
vergebbarem Platz keinen Host, und der Kunde erfuehre das nach der Zahlung.
Der Anfrage-Block erscheint als vierter Bestandteil neben der
Vergleichstabelle, sobald PlanCatalogue::sellable() Enterprise nicht mehr
liefert; die Familie selbst bleibt bestehen, nur sales_enabled entscheidet
(gesetzt wird das erst in Task 9).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 15:48:41 +02:00
nexxo 3b5dee8f0d Ein Ausweg wird nur angeboten, wenn er auch buchbar ist
DowngradeCheck baute seine Empfehlung aus packsToCover(), und die rechnet
bloss. BookAddon lehnt aus ZWEI Gruenden ab - der Deckel bei drei Bloecken
und der Ausschluss von Enterprise -, und die Empfehlung kannte keinen davon.
Solange ein Block 100 GB brachte, war das unerreichbar; seit er 20 GB bringt,
landet jede Ueberschreitung ueber 60 GB dort.

Business -> Team mit 600 GB belegt bot "5 x Zusatzspeicher buchen (+100 GB)"
an. Das Modal klemmte still auf drei, den vierten haette BookAddon abgelehnt,
und der Kunde stand nach 45 Euro im Monat genau dort, wo er vorher stand - auf
der Karte, die sein Abo billiger machen sollte. Enterprise -> Business empfahl
Bloecke, die es fuer dieses Paket ueberhaupt nicht gibt.

AddonCatalogue::bookableQuantity() antwortet jetzt auf beide Gruende. Es gab
einem Enterprise-Vertrag "3" auf die Frage, die sein eigener Docblock stellt.

DowngradeCheck stellt diese eine Frage und klemmt daran: `packs` ist die
gebrauchte Zahl nur, wenn der Vertrag sie auch buchen darf, sonst null - und
dann traegt `short`, was nach allen buchbaren Bloecken uebrig bliebe, gemessen
an DEREN Restmenge statt am Deckel, damit ein Kunde mit einem Block nicht mehr
zu loeschen bekommt als noetig. Kein halbes Angebot: eine Dauerbuchung, die
den Wechsel trotzdem nicht freigibt, ist kein Ausweg, sondern der ausgegraute
Knopf mit Preisschild.

packsToCover() bleibt reine Arithmetik, mit einem Kommentar, der sagt warum:
die kaufmaennische Grenze steht in AddonCatalogue, und ein Klemmen an dieser
Stelle zoege eine Vertragsabfrage in jeden Kontingent-Schritt und ins Portal,
die beide keine Verkaufsfrage stellen.

Der Knopf haengt schon an `packs > 0` und verschwindet von selbst; der Satz
wechselt auf downgrade_escape.capped, der die verbleibende Luecke nennt und
nicht den Grund - "hoechstens drei Bloecke" waere im Enterprise-Fall falsch,
wo es gar keine gibt.

Zwei bestehende Tests hingen an 600 und 800 GB aus der 100-GB-Zeit, beide
inzwischen nicht mehr deckbar; einer haette das Modal gesucht, das die Karte
zu Recht nicht mehr oeffnet. Auf 550 und 560 GB umgestellt, wo sie das pruefen,
wofuer sie geschrieben wurden.

Dazu ein Test, der belegt statt annimmt, warum das max(1, min(...)) in
ConfirmBookStorage stehen bleiben darf: ein Modal ist per openModal direkt
erreichbar, aber bookStoragePacks() legt weder am Deckel noch bei Enterprise
eine Bestellzeile an. Die bestehenden Tests deckten purchase() ab, nicht diese
Weiterleitung.

Voller Testlauf: 2405 bestanden. Pint sauber.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 15:40:26 +02:00
nexxo c3c05ff9f8 Interne Pakete verschwinden aus dem Laden und bleiben in der Konsole
sales_enabled waere der falsche Hebel gewesen: es schaltet auch das
Verschenken ab. Das Testpaket ist fuer Abnahmelaeufe da und muss verschenkbar
bleiben, also zwei Felder fuer zwei Fragen. Der Checkout lehnt einen internen
Schluessel ausdruecklich ab, ueber denselben Fang wie einen unbekannten — eine
URL ist keine Liste, und die Antwort darf keinen Unterschied verraten.

GrantPlan las bislang dieselbe sellable()-Liste wie Preisblatt und Warenkorb,
sowohl fuer sein Dropdown als auch fuer die Validierung des Formularfelds —
ein interner Schluessel waere dort ebenso abgelehnt worden wie im Checkout,
und das Verschenken haette sein einziges Tor verloren. PlanCatalogue bekommt
deshalb grantable() als zweiten Leser derselben Abfrage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 15:31:46 +02:00
nexxo afa6183903 Zusatzspeicher: 20 GB fuer 15 Euro, hoechstens drei, nicht bei Enterprise
100 GB fuer 10 Euro waren 0,10 Euro/GB — unter Einstandspreis und billiger als
jeder Aufstieg. Wer stapelt, belegt dann den knappsten Rohstoff zum niedrigsten
Preis. 0,75 Euro/GB liegen ueber beiden Aufstiegen (0,73 und 0,67), und der
Deckel bei drei liegt dort, wo Aufsteigen billiger UND besser wird.

BookAddon haelt den Ausschluss durch AddonCatalogue::availabilityRefusal()
(gleiches Muster wie CustomDomainAccess), damit greift er auch beim
Verschenken durch den Betreiber. Billing::purchase() und storageLimitNote
fragen dieselbe Regel VOR der Zahlung, sonst haette ein Enterprise-Kunde einen
zahlbaren Auftrag anlegen koennen, den BookAddon erst danach abgelehnt haette.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 15:09:02 +02:00
nexxo 4cd848eead Fix-Runde 2: Warnung fuer nicht reparierte Hosts, case-sichere Vorabpruefung, wiederholbare Migration
Vier Befunde aus dem Abschluss-Review ueber d62a2c8/c815aee/11ba7ee:

- Die Migration renamt nur Hosts mit dns_name; pve-fsn-1/pve-hel-1 blieben
  unrepariert und stumm. Sie werden jetzt gesammelt und gemeldet (Log +
  Konsole), ohne die Migration abzubrechen - diese Hosts laufen weiter.
- Die Vorabpruefung verglich in PHP byteweise, die Spalte liegt auf
  utf8mb4_unicode_ci. Umgestellt auf GROUP BY/HAVING in SQL, damit dieselbe
  Kollation entscheidet, die spaeter den Unique-Index baut.
- Ein Fehlschlag nach der Vorabpruefung liess einen zweiten Anlauf sofort an
  "Duplicate column name" sterben. Die beiden betroffenen Schema-Schritte
  stehen jetzt hinter Schema::hasColumn(), macht den Kommentar darueber wahr.
- Seeder (DatabaseSeeder, DemoCustomerSeeder) sind auf pve-*-Namen sitzen
  geblieben, weil sie dns_name nie benutzt hatten. Auf fsn-01/hel-01
  umgestellt, next_host_number entsprechend vorbelegt.

Dazu vier Kleinigkeiten: ein Test nagelte den falschen Config-Schluessel fest
(dns.zone statt platform_zone), HostName::claim() erzwingt jetzt wirklich
eine Transaktion statt es nur zu verlangen, down() vergisst nicht mehr den
Settings-Cache, und der Kommentar ueber HostName::free() nennt jetzt ehrlich
die Einschraenkung auf einen einzelnen Thread.

Alle vier Migrationslaeufe (Vorabpruefung-Kollision, Meldung fuer
unreparierte Hosts, Fehlschlag-und-erneuter-Anlauf, Rueckbau mit
Cache-Invalidierung) gegen echtes MariaDB auf einer Scratch-Datenbank
geprueft, um die parallele Billing-Session nicht zu beruehren. Voller
Testlauf: 2395 bestanden. Bericht mit allen Befehlen und Ausgaben unter
.superpowers/sdd/2026-08-01-hostname-vergabe/final-fix-report.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:52:15 +02:00
nexxo d6f2aa1c9e Fix-Runde 1: Formel-Test raus, eigene Testhilfe, Docblock nachgezogen
Der Formel-Test rechnete dieselbe Formel nach und wich dabei vom echten Code
ab (kein max(0, ...) um den neuen Summanden) — ein Test, der die Implementierung
nachrechnet und dabei abweicht, ist schlechter als keiner. Der Verhaltenstest
deckt die Anforderung bereits ab und bleibt als einziger Test der Datei.

StoragePackHeadroomTest.php hing außerdem an reservedRun() aus
CustomerStepsTest.php und brach einzeln gefahren mit einem PHP-Fatal ab, statt
mit einem ehrlichen Fehlschlag. Eigene, in sich geschlossene Hilfsfunktion
(packHeadroomRun()) nach dem Muster der Nachbardateien (hostRun(),
restartableInstance()) — die Datei läuft jetzt auch allein grün.

Der Docblock von growDisk() sprach noch vom Kopfraum, der "unchanged"
mitfährt, und widersprach damit dem neuen Summanden direkt darunter. Nachgezogen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:41:39 +02:00
nexxo 3ed80a05e0 Ein Block bringt seinen eigenen Kopfraum mit
Ein Block gibt 20 GB und belegt 22. Ohne das haette ein Start mit drei Bloecken
90 GB auf 100 GB Platte gehabt — zehn Gigabyte Kopfraum, wo die Regel bei dieser
Plattengroesse zwoelf verlangt, und der gestapelte Tarif waere genau der Fall
geworden, den die Regel verhindern soll.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:31:52 +02:00
nexxo af05ed6694 Kapazitaetspruefung hinter den Idempotenz-Kurzschluss verschoben
Durchsicht (R22): die Pruefung stand in __invoke(), vor book()s Kurzschluss
fuer einen wiederholt zugestellten Webhook ("Idempotent against a retried
webhook"). Damit fragte jede Wiederholung erneut "passt NOCH ein Block
drauf" - obwohl keiner hinzukommt - und ein laengst bezahlter, laengst
gebuchter Vorgang quittierte die Wiederholung mit einem Fehler, sobald der
Host zwischenzeitlich eng geworden war. Genau das Szenario, fuer das diese
Aufgabe gebaut wurde, nur gegen den eigenen Kunden gerichtet.

Jetzt sitzt die Pruefung in book(), hinter dem order_id+addon_key-Kurzschluss:
eine Wiederholung bekommt ihre bestehende Buchung zurueck, ohne die Frage
erneut zu stellen. Eine echte neue Buchung durchlaeuft die Pruefung wie
zuvor. Deckel (quantityRefusal) und Domain-Ausschluss bleiben unangetastet -
sie haben dasselbe Muster im Kleinen, sind aber nicht Gegenstand dieses
Befundes.

Neuer Test: derselbe Auftrag wird zweimal gebucht, der Host wird zwischen
den beiden Aufrufen eng - der zweite Aufruf gibt die vorhandene Buchung
zurueck statt zu werfen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:18:01 +02:00
nexxo 11ba7ee34f Hostdetailseite: die Adresse steht neben den IPs und ist ein Klick zur Proxmox-Oberflaeche 2026-08-01 14:12:03 +02:00
nexxo c815aee0b5 Fix-Runde 1: Vorabpruefung vor dem Unique-Index, zwei Praezisierungen, ein fehlender Test
Migration: die Reihenfolge in up() stellte den unwiderruflichsten Schritt
(app_settings loeschen, dns_name-Spalte weg) VOR den einzigen Schritt, der an
vorhandenen Daten scheitern kann (unique('name') - hosts.name trug noch nie
einen eindeutigen Index). Schlug der auf MariaDB fehl, war der Schaden nicht
mehr rueckgaengig zu machen: DDL committet dort implizit, die Migration gilt
aber mangels Eintrag in der Migrationstabelle als nicht gelaufen, und ein
zweiter up()-Versuch stirbt an der bereits fehlenden dns_name-Spalte. Jetzt
steht eine reine Vorabpruefung ganz am Anfang, die simuliert, was die
Uebertragung schreiben wuerde, und mit einer RuntimeException abbricht, bevor
irgendetwas angefasst ist; das Loeschen der app_settings-Zeilen steht jetzt
hinter dem Unique-Index, nicht davor. Gegen echtes MariaDB geprueft,
einschliesslich eines Laufs mit zwei absichtlich kollidierenden Hostnamen.

HostStepsTest: Titel und ein Kommentar praezisiert - der Schritt vergibt den
Namen nicht mehr, er veroeffentlicht ihn nur noch.

HostNamingTest: ungenutzten Import entfernt (Pint), Rueckfall-Test fuer
HostName::label() bei einem Code ohne gueltige Zeichen ergaenzt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:03:35 +02:00
nexxo 01bcc29d13 Ein Block wird nur gebucht, wenn der Host ihn tragen kann
HostCapacity wurde im Checkout, in der Bestellung und im Konsolenbereich
gefragt — beim Aufstocken nicht. Weil qcow2 duenn belegt ist, gelingt die
Ueberbuchung sofort und faellt erst auf, wenn die Gaeste wirklich schreiben.
Gefragt wird der Host der Instanz, denn eine laufende Instanz zieht nicht um.

StorageAllowanceTest: der Host in storageFixture() war mit 1000 GB (active()
Standard) schon fuer eine reine Business-Instanz (1050 GB disk_gb) zu klein —
placeableIn() haette sie nie dort platziert. Ohne Kapazitaetspruefung fiel das
nie auf; jetzt schon, deshalb auf 2000 GB angehoben.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:03:07 +02:00
nexxo c6ead1afc9 Deckel haelt auch am Rand: Warenkorb und Bestaetigungsdialog
Zwei Luecken aus der Durchsicht von Task 2, beide am selben Rand: genau am
erreichten Deckel.

purchase() klemmte $lines mit max(1, min(..., bookableQuantity, ...)) — bei
bookableQuantity=0 floort das auf eine Bestellzeile statt auf keine, und der
Speicher-Knopf war bedingungslos gerendert. Jetzt bricht purchase() fuer
type=storage frueh ab, wenn AddonCatalogue::quantityRefusal() ablehnt, und der
Knopf verschwindet zugunsten des Ablehnungssatzes.

ConfirmBookStorage klemmte an der absoluten Grenze (3), nicht an der
Restmenge — ein Vertrag mit einem gebuchten Block haette drei weitere
versprochen bekommen, obwohl nur zwei noch buchbar sind. Das Modal loest jetzt
seinen eigenen Vertrag auf (wie ConfirmCancelAddon und ConfirmRevokeSeat) und
klemmt an bookableQuantity().

Beide Faelle waren ungetestet; zwei neue Tests in StoragePackLimitTest.php
belegen sie und wurden vor dem Fix gegen den alten Stand als rot verifiziert.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:50:07 +02:00
nexxo d62a2c8ff8 Hostnamen vergibt CluPilot: ein Name statt zweier, und der Zaehler ueberlebt das Loeschen
HostName::claim() haengt den Namen jetzt an die Rechenzentrums-Zeile
(next_host_number), nicht an MAX(hosts.name)+1: der Zaehler ueberlebt so
das Loeschen des zuletzt angelegten Hosts. hosts.dns_name faellt weg -
name ist ab jetzt der einzige Name, den Konsole, DNS, /etc/hosts und
Proxmox-Node teilen. RegisterHostDns veroeffentlicht nur noch, was
StartHostOnboarding beim Anlegen vergeben hat, statt selbst zu
nummerieren; PrepareBaseSystem baut den FQDN ueber HostName::fqdn()
statt ueber die nirgends konfigurierte clupilot.net.

Zwei Testdateien ausserhalb der Aufgabenliste (DatacenterTest,
HostTakeoverPageTest) setzten ->set('name', ...) auf HostCreate, das
Feld jetzt aber nicht mehr hat - im vollen Testlauf nachgezogen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:45:43 +02:00
nexxo ad79ce44be Codex-Runde 2: der Gateway-Fix erreichte den Produktivpfad nicht
Zwei Befunde, beide echt, und der erste ist ein Patzer:
detect_network_style wurde auf default_gateway() umgestellt,
bridge-run.sh blieb auf awk '{ print  }'. Der Treiber reichte also
weiter den KARTENNAMEN als Gateway an build_bridge, und in der Strophe
stand 'gateway ens3' — der Fix half genau der Stelle nicht, fuer die er
gedacht war.

Meine Sandkiste konnte das nicht sehen: sie hatte immer ein via. Jetzt
ist sie parametrisiert, und ein Test faehrt bridge-run.sh mit einer
via-losen Standardroute durch und liest die geschriebene Strophe. Der
Beweis laeuft ueber den Produktivpfad, nicht ueber eine Einzelfunktion.

Zweitens: default_gateway suchte das erste via IRGENDWO in der Ausgabe,
detect_primary_interface das dev der ERSTEN Zeile. Bei zwei
Standardrouten baute das eine Bruecke ueber die Karte der einen mit dem
Gateway der anderen. Beide haengen jetzt an primary_default_route mit
head -1 — eine Route, eine Quelle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:34:12 +02:00
nexxo 1f32a2d5b3 Hoechstens drei Bloecke, geprueft in der Buchung
Die Ansicht durfte 50 in den Warenkorb legen, die Aktion pruefte nichts. Der
Deckel liegt kaufmaennisch dort, wo Aufsteigen billiger wird als Stapeln, und
gehoert deshalb dorthin, wo gebucht wird — gezaehlt ueber alle laufenden
Buchungen, denn drei Bestellungen a einem Block sind drei Bloecke.

StorageAllowanceTest schrieb die alte Notbremse (50) als erwartete Zahl fest;
angepasst auf die jetzt engere kaufmaennische Grenze (3), wie im Aufgabenblatt
fuer DowngradeTest vorgezeichnet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:33:36 +02:00
nexxo 98354716d0 Codex-Runde 1: drei P1 an der Rueckfahrkarte, plus ein Gateway ohne via
Die Ruecknahme konnte Erfolg melden, ohne einen zu haben. Drei Wege
dorthin, alle behoben:

- Sie schaltete den entmachteten Netzverwalter nicht wieder ein. Die
  Sicherung umfasst nur /etc/network/interfaces*; kam die Verbindung von
  cloud-init, networkd oder NetworkManager, spielte die Ruecknahme eine
  Datei zurueck, die die Maschine nie getragen hat, und liess den
  Verwalter abgeschaltet. Genau der tote Host, den der Zeitgeber
  verhindern soll. disown_network_manager hinterlaesst jetzt eine Notiz
  (WAS entmachtet, WELCHE Strophen verdraengt), die das Ruecknahme-Skript
  beim Feuern liest — aufgeschrieben statt eingebacken, weil der
  Zeitgeber vor dem Entmachten gestellt wird.
- Sie setzte 'rolled-back' auch, wenn tar oder ifreload scheiterten. Das
  urspruengliche network.sh hatte dafuer set -e; beim Umbau ist es
  verlorengegangen. Jetzt bricht jeder Fehlschlag ab, bevor die Marke
  entsteht — CluPilot pollt dann bis zur Frist statt 'ist zurueck' zu
  glauben.
- Eine verdraengte interfaces.d-Strophe wurde nur umbenannt. Der Stern in
  'source interfaces.d/*' fasst sie weiter; das versteckte die Kollision
  vor dem Leser, statt sie zu loesen. Jetzt wandert sie aus dem
  Verzeichnis heraus, und die Ruecknahme holt sie zurueck.

Dazu ein eigener Fund: das Gateway wurde mit awk '{print }' gelesen.
Bei 'default dev ens3 scope link' ist das der KARTENNAME, woraus
'gateway ens3' in der Strophe wuerde. default_gateway() liest jetzt
hinter dem via, und eine Standardroute ohne via wandert als eigene
up-Zeile mit, statt verlorenzugehen.

Zurueckgewiesen: der P2 zu extra_routes. 'ip route show dev X' laesst das
dev-Feld WEG (im Container nachgemessen), das angehaengte 'dev vmbr0' ist
also richtig.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:27:49 +02:00
nexxo 66bddd4bce EnsureNetworkBridge in die Pipeline: nach dem Neustart, vor der Nachpruefung
Der Schritt haengt jetzt zwischen RebootIntoPveKernel und
ConfigureProxmox. Nach dem Neustart, weil dort ifupdown2 steht und der
Tunnel gerade bewiesen hat, dass er einen Neustart ueberlebt. Vor
ConfigureProxmox, dessen refuseWithoutBridge() unveraendert stehen bleibt
und damit zur Nachpruefung wird — bauen UND pruefen, dasselbe Paar wie
BuildVmTemplate -> VerifyVmTemplate.

Beschriftung in beiden Sprachen, und ein Test verlangt das kuenftig von
JEDEM Pipeline-Schritt statt nur vom neuen: wer einen einhaengt und die
Sprachdateien vergisst, faellt im Test auf statt in der Konsole.

bridge.sh und bridge-run.sh gehoeren ins Bootstrap-Archiv — es ist der
einzige Weg, auf dem das Skript auf eine nackte Maschine kommt.

Der End-to-End-Test lief rot, und zu Recht: sein Attrappen-Host sagte
nichts ueber sein Netz, also lehnte der neue Schritt ab. Das war der
Beweis, dass er wirklich haengt. Er bekommt jetzt einen Host, der seine
Bruecke schon hat — den Bau-Pfad kann eine Attrappe nicht nachstellen,
er lebt davon, dass die Verbindung abreisst und wiederkommt. Dafuer gibt
es die Schritt-Tests und die drei Sandkasten-Laeufe.

scriptBridgeState/scriptBridgeStatus liegen in tests/Pest.php, nicht in
einer einzelnen Testdatei — zwei Dateien brauchen sie.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:14:40 +02:00
nexxo 9888f898b8 EnsureNetworkBridge: abbestellen nach der Nachpruefung, aufgeben ohne die Rueckfahrkarte wegzuwerfen
Drei Enden, und jedes hat eine Reihenfolge:

- Abbestellt wird erst, NACHDEM die Bruecke nachgeprueft ist. Dass der
  Zweig erreicht wird, ist der Beweis: SSH kam ueber den Tunnel an und
  vmbr0 traegt Standardroute und Adresse — bridge_proven, von aussen
  gefragt. Und nur, wenn DIESER Lauf den Zeitgeber auch gestellt hat;
  ohne Termin im Kontext gehoeren die Unit-Dateien jemand anderem.
- awaitRollback pollt, bis 'rolled-back' liegt, statt sofort zu
  scheitern. Solange der Zeitgeber aussteht, steckt die Maschine mitten
  in einer Umstellung, und ein fail() liesse sie dort liegen. Gedeckelt
  durch die Schritt-Frist, die das Dreifache der Zeitgeber-Frist ist.
- giveUp loescht den Termin (sonst wird jedes Retry nach einem
  Fehlschlag zum sofortigen zweiten Fehlschlag), behaelt bridge_attempts
  (sonst ist der Deckel von zwei Versuchen keiner) und bestellt den
  Zeitgeber NICHT ab — er ist die Rueckfahrkarte, und steht er noch, hat
  er seinen Grund.

Der sh-n-Test ist gegengeprobt: mit absichtlich kaputtem Quoting faellt
er.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:05:58 +02:00
nexxo 2e511899f8 EnsureNetworkBridge: Start abgekoppelt, und der Verbindungsabriss wird gepollt statt gezaehlt
Die tragende Zeile ist das try/catch um keyLogin. RunRunner:102
verwandelt jede geworfene Ausnahme in ein retry(), und retry verbraucht
das Versuchskonto — ungefangen brennt der erwartete Verbindungsabriss die
fuenf Versuche in wenigen Minuten durch und laesst den Lauf scheitern,
BEVOR der Host wieder da ist. Genau der Unterschied zwischen
'Wiederholung' und 'toter Server'.

Unterschieden wird am Termin im Run-Kontext: ohne ihn hat dieser Lauf
nichts angefasst, dann ist ein Verbindungsfehler ein gewoehnlicher und
darf einen Versuch kosten. Mit ihm laeuft gerade eine Umstellung, dann
wird gepollt.

bridge.sh und bridge-run.sh gehen wortgleich hoch (Test vergleicht Byte
fuer Byte gegen die Repo-Datei), env traegt den Hub-Schluessel mit, damit
der Treiber den Handshake gegen den RICHTIGEN Peer prueft. Start per
nohup setsid, PID-Datei abgewartet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:04:03 +02:00
nexxo 91ecb382f5 EnsureNetworkBridge: schon da heisst nichts anfassen, und keine Bruecke auf einer Bruecke
Der Schritt, aber erst sein billigster Teil. Zwei Zusicherungen:

- Traegt vmbr0 schon Standardroute UND Adresse, dann advance() ohne einen
  einzigen veraendernden Befehl. pve-fns-1 hat die Bruecke von Hand; ein
  Wiederanlauf, der sie umbaut, baut ein funktionierendes Netz um.
- Ist die Karte mit der Standardroute keine physische (Bond, Bridge,
  VLAN), dann fail() mit Klartext statt bauen. bridge_ports darauf waere
  falsch, und was dabei herauskommt, ist aus der Ferne nicht mehr zu
  reparieren.

readBridgeState fragt beides in einem Rundlauf und leitet aus dem
LAUFENDEN Zustand ab (ip, /sys/class/net), nie aus der Datei des
Anbieters. 'up' sind absichtlich alle drei Fakten: eine vmbr0 ohne
Adresse und ohne Standardroute ist eine Bruecke im Sinne von 'ip link'
und sonst nichts.

Start, Poll und Abbestellen folgen. Bis dahin steht dort ein fail() —
gefahrlos, weil der Schritt noch nicht in der Pipeline haengt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:02:23 +02:00
nexxo 5df2c1240e bridge-run.sh: der Treiber, und der Zeitgeber steht vor jeder Aenderung
Verdrahtet bridge.sh in der einen Reihenfolge, die stimmen muss:
sichern -> Zeitgeber -> uebernehmen -> umstellen -> nachsehen. Alles
davor stellt nur fest und veraendert nichts; ab 'sichern' gibt es einen
Weg zurueck, und erst ab dann darf ueberhaupt etwas angefasst werden.

Abgekoppelt ist hier Bedingung, nicht Optimierung: ifreload -a nimmt die
Leitung, ueber die der Befehl laeuft. PID als allererstes, damit ein
frueher Poll nicht 'running aber nicht lebendig' liest und einen gesunden
Lauf fuer tot erklaert.

Der Treiber bestellt den Zeitgeber NIE ab — das tut CluPilot nach dem
Wiederverbinden. Ein Test haelt das fest.

Dazu die Fremdverwalter-Erkennung in bridge.sh: cloud-init, networkd,
NetworkManager werden benannt und entmachtet, Unbekanntes fuehrt zum
Abbruch statt zu einem Versuch ins Blaue. Der Zeitgeber faengt diesen
Fall NICHT ab — zu seiner Zeit war alles in Ordnung, und die Bruecke
verschwaende erst beim naechsten Neustart, mit Kunden darauf.

Drei Tests fahren den Treiber in einer Sandkiste WIRKLICH durch, statt
nur seinen Text zu lesen: die ip-Attrappe antwortet vor dem ifreload
anders als danach. Sie belegen ok, failed-ohne-Bruecke und
failed-mit-schalem-Tunnel — und in allen drei Faellen, dass die
Rueckfahrkarte stehen bleibt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:56:16 +02:00
nexxo 4ba74e862a bridge.sh: nachsehen in beide Richtungen, und kein ping
Nur 'komme ich raus' zu pruefen ist notwendig und NICHT hinreichend.
Der Fehlerfall: ifreload bringt vmbr0 sauber hoch, die Maschine erreicht
das Internet, der Treiber waere zufrieden — aber wg0 kommt nicht zurueck.
Dann lebt der Host oeffentlich, CluPilot ist ausgesperrt, und wer hier
abbestellt, hat die Rueckfahrkarte weggeworfen.

wg0.conf enthaelt keine Geraetebindung; der Tunnel haengt an der
Quelladresse, die die Routing-Tabelle hergibt — und genau das ist die
Groesse, die der Umbau anfasst. Das macht den Fall nicht
unwahrscheinlicher, nur unauffaelliger: kein Fehler im Protokoll, nur
ein Handshake, der ausbleibt.

Ist der Handshake schal, wird wg-quick@wg0 EINMAL neu gestartet und
nochmal nachgesehen. Gefahrlos, weil ifreload die SSH-Sitzung ohnehin
schon mitgenommen hat — und es verwandelt einen haengenden Tunnel in
einen laufenden statt in eine Ruecknahme.

Kein ping: Hetzners Debian-Basis hat keins, PrepareBaseSystem
installiert es nicht, und network.sh:190 haette damit immer 'nicht
erreichbar' gesagt. Ein Test haelt bridge.sh ping-frei.

Nebenbei ein Fehler in den Tests selbst behoben: 'if gibtsnicht; then
… else echo NEIN; fi' ist in sh unwahr, also war jeder Test, der NEIN
erwartete, gruen SOLANGE die Funktion fehlte. verdictBody() meldet jetzt
FEHLT und trennt 'falsch' von 'gibt es nicht'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:47:02 +02:00
nexxo 8ac4a38e2e bridge.sh: die Rueckfahrkarte, und der Grund steht vor dem Zurueckspielen
Der Zeitgeber ist eine systemd-Einheit und kein 'sleep &': ein
Hintergrundlauf stirbt mit seiner Sitzung, und die Sitzung ist genau
das, was abreisst, wenn die Umstellung schiefgeht.

Zwei Dinge gegenueber der Fassung, die in network.sh stand:

- render_rollback_script ist vom Stellen getrennt, damit die Reihenfolge
  darin ohne systemd pruefbar ist. Und die Reihenfolge ist der Punkt:
  'failed' samt Grund wird geschrieben, BEVOR zurueckgespielt wird —
  ein halb gegluecktes Zurueckspielen soll das Urteil trotzdem
  hinterlassen. 'rolled-back' kommt zuletzt und ist das Signal, auf das
  CluPilot wartet.
- Der Zeitgeber raeumt seine Unit-Dateien nach dem Feuern selbst weg.
  Sonst sieht eine abgeschlossene Ruecknahme beim naechsten Hinsehen aus
  wie eine ausstehende.

Der Treiber bestellt NICHT ab. Das tut CluPilot, nachdem es sich ueber
den Tunnel neu verbunden hat — ein Skript auf dem Host kann ueber seine
eigene Erreichbarkeit von aussen nur raten.

Geprueft: der Zeitgeber steht vor dem ersten veraendernden Aufruf.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:39:44 +02:00
nexxo 587d1351d1 bridge.sh: die Strophe, gegen vier Anbieterfaelle geprueft
write_bridge_stanza ist von build_bridge getrennt: das Schreiben ist das,
was eine Maschine umbringt, und so ist es pruefbar, ohne dafuer ein Netz
neu laden zu muessen. Beide Pfade sind ueberschreibbar
(CLUPILOT_INTERFACES_FILE, CLUPILOT_IFRELOAD), dieselbe Technik wie
CLUPILOT_STORAGE_CFG beim Vorlagenbau.

Drei Dinge kann die Fassung mehr als die alte in network.sh:

- hwaddress festgenagelt. Eine Bruecke waehlt sonst die kleinste MAC
  ihrer Ports; bei einem Port ist das dieselbe, aber 'ist dieselbe' und
  'bleibt dieselbe' sind zweierlei.
- Zusatzrouten des Anbieters wandern mit. Ausgelassen bleiben die
  Kernel-Route zum eigenen Subnetz und die Link-Route zum Gateway —
  beide entstehen von selbst, und ein gescheitertes 'up' nimmt bei
  ifreload die ganze Strophe mit.
- IPv6, aber nur statisch und global. SLAAC/DHCPv6 werden bewusst nicht
  nachgebaut (forwarding=1 laesst den Kernel RAs ohne accept_ra=2
  verwerfen) — dafuer gibt es eine Zeile ins Protokoll statt eines
  stillen Verlusts.

Geprueft gegen echte sh: Hetzner /32 routed mit pointopoint, netcup
Subnetz ohne, DHCP ohne Adresse, statisches IPv6 mit fe80::1, SLAAC
faellt weg, Zusatzrouten kommen mit.

network.sh hat sein eigenes build_bridge abgegeben; ein Test haelt jetzt
jede der sieben Brueckenfunktionen auf genau einer Stelle im Repo fest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:37:17 +02:00
nexxo 6bda0e1d8a bridge.sh: die Erkennung herausgeloest und allein lauffaehig gemacht
Die Bruecken-Erkennung stand in network.sh, geschrieben fuer den
stillgelegten Rettungssystem-Weg, und borgte sich vier Dinge von
woanders: detect_primary_interface aus proxmox.sh, log, http_get und
CLUPILOT_PROBE_URL aus clupilot-bootstrap.sh.

Der Debian-Weg laedt die Bibliothek EINZELN auf den Host und faehrt sie
dort — geborgte Helfer waeren dann nicht da. Also allein lauffaehig, mit
log und http_get unter einem command-v-Schutz, damit der Bootstrap seine
eigenen behaelt.

Neu und aus dem laufenden Zustand abgeleitet:
- interface_is_physical — bridge_ports auf einem Bond oder einer
  bestehenden Bridge ist falsch, und aus der Ferne nicht reparierbar.
- address_is_dynamic — der Kernel markiert eine geleaste Adresse, das
  steht bei jedem Anbieter gleich da. Die Datei des Anbieters ist nur
  noch das Zweitsignal.

Geprueft gegen eine echte sh mit aufgezeichneten ip-Ausgaben, plus die
Zusicherung, dass es detect_primary_interface im Repo genau einmal gibt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:34:26 +02:00
nexxo 2ab6a92fd8 Die Packungsgroesse steht auf der Buchung, nicht in der Konfiguration
Bisher las StorageAllowance sie bei jeder Anzeige neu aus config. Solange es
eine Groesse gab, war das folgenlos — beim Zuschnitt von 100 GB auf 20 GB waere
es der stille Verlust von 80 GB je gekauftem Block gewesen, bei einem Kunden,
der bereits Daten darin liegen hat.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 12:25:30 +02:00
nexxo 55b749c129 Release v1.3.94 — die Kurven zeigen wieder, wie ausgelastet der Host ist
tests / pest (push) Failing after 9m19s Details
tests / assets (push) Successful in 21s Details
tests / release (push) Has been skipped Details
Drei Punkte des Besitzers, und der dritte hat den eigentlichen Fehler
freigelegt.

Die Verlaufslinien sahen komisch aus, weil sie auf ihr EIGENES Minimum und
Maximum skalierten. Ein Host mit 0,2 % CPU, der zwischen 0,1 und 0,3 schwankt,
zeichnete damit einen Seismographen über die volle Höhe — und widersprach flach
der Zahl direkt daneben. Dieselbe Regel stand im Chart.js-Entwurf schon
ausformuliert ("eine automatisch skalierte Achse lässt 3 % wie eine Wand
aussehen") und ging beim Umstieg auf die Kacheln verloren.

x-ui.spark bekommt deshalb `min`/`max`. Ohne Angabe bleibt alles wie bisher —
das ist, was jeder bestehende Aufrufer übergibt. Die Prozent-Kacheln geben
0 und 100 mit, die Netz-Kacheln nur die 0, weil MiB/s keine natürliche
Obergrenze haben. Ein untätiger Host zeichnet jetzt vier ruhige Linien statt
zweier wogender und zweier flacher — vorher wirkte CPU (0,2 %) belebter als RAM
(3,6 %), also genau verkehrt herum.

Dazu: Werte außerhalb der Grenzen werden geklemmt statt aus dem Kasten
gezeichnet, die Fläche ist ein Verlauf statt einer harten Kante, und die
Linien sind mit 120×40 statt 80×32 lesbar.

Ich hatte max=100 zwischendurch selbst verworfen, weil es "zu tot" aussah — auf
einem Bild, auf dem die Füllung wegen eines CSS-Fehlers gar nicht gezeichnet
wurde. Ein Vergleich mit einem kaputten Bild. Der Fehler: `.spark path
{ fill: none }` schlägt als CSS-Regel das Präsentationsattribut fill="url(#…)".
`fill: none` gehört an die Linie, nicht an jeden Pfad.

Die Ausstattung ist wieder einzeilig. Sechs Felder, sechs Spalten — und die
Bau-Kennung steht nur noch im Titel: als eigene Zeile zwang sie die ganze Tafel
in eine zweite Reihe, für eine Zeichenkette, die fast niemand liest.

Und die Wartezeit: die Seite stößt beim Öffnen eine Sammlung an, wenn noch
keine Messwerte da sind, statt bis zum nächsten minütlichen Lauf leer zu
bleiben. Ein leerer Kasten liest sich als "kaputt", nicht als "gleich". Der Job
ist ShouldBeUnique, zwei geöffnete Seiten reihen also keine zwei ein.

Codex-Befund dazu (P2): das galt auch für Hosts, die der Sammler ohnehin
überspringt — ein Host mitten in der Übernahme hat keinen Token, und seine
Seite hätte die ganze Flotte abgeklappert und die eigenen Kacheln trotzdem leer
gelassen. Wer gesammelt wird, steht jetzt einmal am Modell (Host::collectable),
gelesen von Job und Seite. Zwei Fassungen dieser Frage waren genau der Grund.

Geprüft: 2307 Tests grün, Pint sauber, Codex ohne Befund. Der Test für den
Maßstab prüft jetzt auch die echten Aufrufstellen und nicht nur das Bauteil —
gegengeprobt durch Entfernen von max=100, dann fällt er. Im Browser beide
Fälle angesehen: ruhiger Host vier flache Linien, ausgelasteter Host lesbare
Form, null Konsolenfehler.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 11:56:10 +02:00
nexxo dc35e5310f Release v1.3.93 — die Seite konnte den Host gar nicht fragen
tests / pest (push) Failing after 9m23s Details
tests / assets (push) Successful in 28s Details
tests / release (push) Has been skipped Details
Auf echter Hardware blieb jede Messkachel leer, während der Host sichtbar
online war und vor zwei Minuten geantwortet hatte. Kein Zufall und kein
Aussetzer: ein Entwurfsfehler.

Nur der queue-provisioning-Container hängt im WireGuard-Netz — NET_ADMIN,
/dev/net/tun, das wireguard-Volume. Der app-Container, der die Konsole rendert,
hat überhaupt keine Route zu 10.66.0.x. Ich hatte den Proxmox-Aufruf in
render() gelegt, also in den einen Container, der die Management-Adresse nicht
erreichen kann. Ein Aufruf von dort konnte nie etwas anderes sein als eine
Zeitüberschreitung.

PingHosts schreibt dieselbe Regel seit Langem in seinen Kopf: "runs on the
provisioning queue, which is where the Proxmox credentials are usable." Ich
habe sie gelesen und nicht angewendet.

Verschlimmert hat es mein eigenes catch (Throwable): der Grund wurde
verschluckt, und die Kachel sagte "keine Messwerte" — ununterscheidbar davon,
dass der Host schweigt. Auf dem Testhost fiel nichts auf, weil TEST-NET
ohnehin nie antwortet.

Jetzt zwei Hälften:

- HostLoadSeries::collect() holt und legt ab, aus Jobs\CollectHostLoad auf der
  provisioning-Warteschlange, minütlich — der Takt, in dem Proxmox einen
  frischen Messwert schreibt. Ein Fehlschlag wird protokolliert, mit Host,
  Node und Grund.
- HostLoadSeries::forHost() liest nur aus dem Zwischenspeicher und öffnet nie
  eine Verbindung. Ein Test hält das mit Http::assertNothingSent() fest.

Der Eintrag lebt fünf Minuten bei minütlichem Sammeln: länger als der Takt,
damit ein ausgefallener Lauf keine Seite leert, die eine Sekunde vorher in
Ordnung war — und kurz genug, dass ein stehengebliebener Sammler die Zahlen
mitnimmt, statt eine alte Stunde als aktuell stehenzulassen.

Der Sammler ist ShouldBeUnique (Codex-Befund, P1). Die provisioning-Warteschlange
ist DIESELBE, auf der Kunden-Bereitstellung läuft; ein stiller Host kostet den
vollen HTTP-Zeitablauf, und ohne diese Sperre stauten sich minütlich neue Läufe
hinter dem alten und verzögerten bezahlte Arbeit. Dasselbe Mittel, das
CollectInstanceTraffic nebenan schon benutzt.

Dazu: der Zustands-Punkt war mit 62 px so groß wie der Speicher-Ring nebenan.
Eine gefüllte Scheibe wiegt optisch weit mehr als ein dünner Ring und erschlug
die Kachel — jetzt 32 px.

Und ein Test, der aus Versehen recht behielt: die Kachel-Prüfung verglich mit
"50", was auch die 500 GB in der Instanzenliste darunter trifft. Sie prüft
jetzt Zahlen, die sonst nirgends auf der Seite vorkommen.

Noch offen, nicht hier angefasst: VmTemplateCheck fragt die Proxmox-API
ebenfalls aus dem app-Container heraus, von der Bereitschaftsseite aus. Selber
Fehler, Bestand, eigener Punkt.

Geprüft: 2299 Tests grün, Pint sauber, Codex ohne Befund. Im Browser mit
eingespielten Messwerten: sechs gefüllte Kacheln, Zustands-Scheibe in
Proportion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 11:23:00 +02:00
nexxo 589489a9fc Release v1.3.92 — sechs Kacheln statt dreier ungleicher Kästen
tests / pest (push) Failing after 10m15s Details
tests / assets (push) Successful in 28s Details
tests / release (push) Has been skipped Details
Nach der Vorlage des Besitzers. Zustand, Speicher und die gemeinsame Kurve
standen nebeneinander: zwei davon halb leer, die dritte überfüllt, und die
gemeinsame Kurve brauchte eine Legende, um zu sagen, welche Linie welche ist.

Eine Reihe je Kachel löst alle drei Beschwerden auf einmal. Die Beschriftung
der Kachel benennt die Reihe, also braucht es keine Legende mehr. Die Höhen
sind durch das Raster gleich statt zufällig. Und der leere Platz ist mit
Messwerten gefüllt, die es ohnehin schon gab: Proxmox' Aufzeichnung liefert
Netzdurchsatz in beide Richtungen mit, ungefragt.

Sechs Kacheln: CPU-Auslastung, RAM-Auslastung, Speicher (als Ring), Eingehend,
Ausgehend, Zustand. Gebaut mit x-ui.metric, x-ui.spark und x-ui.ring — die
gab es alle schon, und der Kopfkommentar von x-ui.metric sagt selbst "exactly
as the approved template draws it". Nichts daneben neu gebaut.

Die öffentliche IP steht jetzt auf der Seite. Sie stand vorher NUR klein unter
der Überschrift — die Adresse, unter der der Host wirklich erreichbar ist, war
in der Detailseite nirgends ein Feld. Sie führt jetzt die Ausstattungs-Tafel
an, und die Reserve-Eingabe ist mit dorthin gezogen: die Kacheln zeigen, was
gemessen wurde, die Tafel, was eingestellt ist. Ein Eingabefeld zwischen
Messwerten sähe aus, als ließe sich eine Messung ändern.

Alle vier Verlaufslinien tragen denselben Ton. Die Regel steht im Bauteil
selbst — "muted where the figure is observed, accent where it can be acted on"
—, und hier ist keine Zahl anzufassen. Vier verschiedene Töne nebeneinander
behaupten einen Unterschied, den es nicht gibt.

Zwei Funde aus der Prüfung
--------------------------

- x-ui.spark warf fehlende Messwerte per array_filter heraus und verband die
  Nachbarn. Zwei Fehler auf einmal: die Linie behauptete eine Messung, die es
  nicht gab, und alles danach rutschte nach links — eine Stunde mit zwei Lücken
  zeichnete sich als achtundfünfzig Minuten. Die x-Lage kommt jetzt aus dem
  Platz in der URSPRÜNGLICHEN Reihe, und zusammenhängende Messwerte werden als
  eigene Züge gezeichnet. Eine saubere Reihe ergibt genau einen Zug und
  dasselbe Bild wie vorher, was alle bisherigen Aufrufer liefern.

- Der Zwischenspeicher überlebt einen Deploy. Ein Eintrag aus v1.3.91 kennt
  netin/netout nicht, und die Host-Seite wäre 55 Sekunden lang an einem
  fehlenden Schlüssel gestorben — genau in der Minute, in der jemand nachsieht,
  ob das Update durch ist. Der Schlüssel heißt jetzt host-load:v2:<id> und
  wandert mit der Form mit.

Und einer, den kein Prüfer gemeldet hat: beim Zerlegen in Züge stand im
Flächenpfad ein `L` unmittelbar vor einem `M`. Gültig gelesen, nicht
gezeichnet — die Füllung verschwand still. Aufgefallen ist es beim Ansehen der
Seite, nicht durch eine Meldung; jetzt prüft ein Test, dass jeder Flächenpfad
mit M anfängt, mit Z endet und keinen Befehl direkt hinter einem anderen trägt.

x-ui.chart behält seinen update-on-Weg, obwohl diese Seite ihn nicht mehr
benutzt: er ist eine geprüfte Fähigkeit des gemeinsamen Bauteils, und der
darunterliegende Fix (Instanz aus dem reaktiven Alpine-Objekt) gilt für jeden
Chart.

Geprüft: 2291 Tests grün, Pint sauber, Codex ohne Befund. Im Browser mit
eingespielten Messwerten: sechs Kacheln, Lücke als echte Aussparung in Linie
UND Fläche, Leerzustand zeigt "—" statt einer Null, null Konsolenfehler über
einen vollen Poll-Zyklus.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 11:02:55 +02:00
nexxo 665a0c4e08 Release v1.3.91 — die Host-Seite zeigt Last statt Ausstattung
tests / pest (push) Failing after 9m29s Details
tests / assets (push) Successful in 20s Details
tests / release (push) Has been skipped Details
Die Karte "Rechenleistung" zeigte keine Leistung. 12 Kerne und 63 GB sind die
Ausstattung des Blechs und ändern sich nie — sie standen aber im selben
Kartenraster wie "Zustand" und "Speicher", die beide leben. Wer die Seite
öffnete, um zu sehen, wie es dem Host geht, las dort eine Zahl, die das nie
sagen konnte.

An ihrer Stelle steht jetzt die Last: CPU und RAM als Stundenkurve, beide in
Prozent auf EINER Achse, dazu die aktuellen Werte als beschriftete Zahlen.

Die Geschichte kommt aus Proxmox' eigener Aufzeichnung
(/nodes/{node}/rrddata), nicht aus einem eigenen Sampler. Ein Sampler hieße
neue Tabelle, minütlicher Job, Aufräum-Job und ~1440 Zeilen je Host und Tag —
um weniger genau nachzubauen, was ohnehin auf der Platte liegt. Die RRD ist ab
der ersten Sekunde gefüllt, auch für die Stunde vor dem ersten Hinsehen, und
kann nicht von dem abweichen, was Proxmox' eigene Oberfläche zeigt.

Eine Lücke bleibt eine Lücke: ein Punkt ohne Messwert wird null, nie 0, und
spanGaps steht auf false. Dieselbe Regel wie in instance_metrics. Antwortet der
Host gar nicht, sagt die Tafel das in einem Satz, statt eine ruhige Stunde zu
zeichnen — auf dem Testhost live bestätigt.

Der Fehler, den das ans Licht gebracht hat
------------------------------------------

x-ui.chart steht überall unter wire:ignore, sonst zerstört Livewire das Canvas.
Ein Poll erreicht den Chart also nie. Dafür bekam das Bauteil ein optionales
update-on: es hört auf ein Fenster-Ereignis und tauscht die Daten IM
bestehenden Chart.js-Objekt.

Das lief nicht. Die Zahlen neben der Kurve wanderten, die Kurve nicht, und
chart.update() starb still im Legenden-Plugin:

    TypeError: Cannot set properties of undefined (setting 'fullSize')

Grund: die Chart.js-Instanz lag als Eigenschaft im Alpine-Objekt und wurde
damit reaktiv umhüllt. Chart.js' Plugin-Innenleben überlebt das Proxy nicht.
Gemessen statt geschlossen: dieselbe Instanz wirft über das Proxy und läuft
über Alpine.raw(). Sie liegt jetzt in der Closure.

Aufgefallen ist es nie, weil bis zum ersten Live-Chart kein einziger Chart in
diesem Repo je update() gerufen hat — konstruieren und Erstzeichnen gehen durch
die Hülle noch. tests/Feature/ChartLiveUpdateTest.php hält die Regel fest,
damit der nächste Live-Chart nicht denselben Nachmittag kostet.

Der Rest
--------

- Version lesbar: "Proxmox VE 9.2.6" statt pve-manager/9.2.6/7f8d…, mitten in
  der Bau-Kennung abgeschnitten. Die Kennung steht klein darunter. Eine
  unerwartete Form wird unverändert durchgereicht statt verschluckt.
- Vier Kleinkarten (Mgmt-IP, Node, Version, Instanzen) sind eine
  Ausstattungs-Tafel geworden. Die Instanzen-Anzahl steht in der Überschrift
  der Liste, die sie ohnehin zeigt.
- Der Übernahme-Fortschritt klappt zu, sobald sie durch ist. Fünfzehn
  abgehakte Schritte sind auf einem laufenden Host kein Dauerinhalt —
  aufklappbar über <details>, ohne JavaScript.
- PlanVersion::requiredTemplateVmids() ersetzt die dritte Kopie derselben
  Fensterlogik.
- BuildVmTemplate sagt nicht mehr "this takes 10–20 minutes". Das war eine
  Schätzung vor dem ersten Lauf; gemessen waren es unter zwei. Eine Konsole,
  die falsch vorhersagt, erzieht dazu, sie zu ignorieren.

Geprüft: 2281 Tests grün, Pint sauber, Codex ohne Befund. Die Farbwahl gegen
den Validator gerechnet (ΔE 28,3 protan / 39,2 normal; der Akzent liegt unter
3:1 gegen die Fläche, deshalb tragen beide Reihen sichtbare Beschriftung). Im
Browser: null Konsolenfehler über einen vollen Poll-Zyklus, und die Kurve
wandert auf dem echten Poll ohne Neuladen — mit eingespielten Messwerten
belegt, weil die Testhosts in TEST-NET liegen und nie antworten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 10:24:21 +02:00
nexxo db125781ed Release v1.3.88 — die Vorlage baut sich selbst
tests / pest (push) Failing after 9m17s Details
tests / assets (push) Successful in 22s Details
tests / release (push) Has been skipped Details
Der letzte Handgriff in der Host-Übernahme fällt weg. VerifyVmTemplate meldete
bisher nur, dass eine Vorlage fehlt, weil niemand entschieden hatte, was in die
goldene Vorlage gehört. Entschieden ist es längst und steht in
deploy/bootstrap/lib/template.sh — der neue Schritt BuildVmTemplate lädt genau
diese Datei auf den Host und führt sie dort aus, statt ihre Prüfungen ein
zweites Mal in PHP zu haben.

Er läuft abgekoppelt und wird abgefragt: Abbild laden und drei
virt-customize-Läufe brauchen zehn bis zwanzig Minuten, ein einzelner
SSH-Aufruf liefe gegen den Befehlszeitablauf von 2000 s. "Läuft noch" heißt
dabei, dass der Prozess lebt (kill -0 gegen die hinterlegte PID) — in der
Statusdatei steht "running" auch dann noch, wenn niemand mehr da ist, der sie
ändert.

Fünf Fehler, die dabei aufgefallen sind und Geld gekostet hätten:

- qm importdisk hängte die Platte unter ${storage}:vm-9000-disk-0 ein. Der Name
  gilt nur bei Block-Ablagen; auf einer Verzeichnis-Ablage heißt sie
  local:9000/vm-9000-disk-0.qcow2 — also genau auf dem per Debian aufgesetzten
  Proxmox, um das es hier geht. Jetzt qm set --import-from, und Proxmox
  benennt selbst.
- growpart war nie installiert. GrowGuestFilesystem ruft es auf, und es lief
  bisher, weil Debians Cloud-Abbild es zufällig mitbringt. Fiele es heraus,
  läge jedes gekaufte Kontingent über einem Dateisystem, das nie gewachsen ist.
  Jetzt ausdrücklich eingebaut und als vierte Falle nachgewiesen.
- local nimmt ab Werk keine Platten an. Ohne das stirbt nicht nur der Bau,
  RegisterCapacity meldet danach Kapazität 0: ein Host, der fertig aussieht und
  nie einen Kunden tragen kann. ensure_image_storage greift nur ein, wenn keine
  Ablage Platten annimmt, hängt images an die vorhandene Liste an statt sie zu
  ersetzen, und schreibt über pvesm set statt in die pmxcfs-Datei.
- Ein abgebrochener Download blieb unter dem Zielnamen liegen und wäre beim
  nächsten Lauf ungeprüft weiterbenutzt worden. Jetzt .part, umbenannt erst
  nach geprüfter Summe.
- VerifyVmTemplate und VmTemplateCheck fragten nur, ob VMID 9000 existiert. Ein
  abgebrochener Bau hinterlässt eine gewöhnliche VM mit dieser Nummer, und
  beide sagten dazu "passt" — der Fehler kam beim ersten bezahlten Klon zurück.
  Jetzt template: 1.

isTemplate() stellt zwei Anfragen, weil die falsche Antwort hier etwas
zerstört: false heißt "Vorlage fehlt", und der Bau fängt mit qm destroy --purge
an. Proxmox beantwortet die Konfiguration einer nicht vorhandenen VM mit 500 —
demselben Code wie einen Knoten in Not. Die VM-Liste klärt deshalb die
Abwesenheit, alles darunter wirft und landet im Wiederholungs-Zweig.

Aufgeben beendet erst die Prozessgruppe, dann räumt es auf, und gebaut wird nur
die Fehlliste: create_proxmox_template räumt eine VMID weg, bevor es sie
anlegt, also hätte "alles Verlangte" eine gesunde zweite Vorlage auf dem Weg
zerstört.

Geprüft: 2267 Tests grün, Pint sauber, sh -n über alle drei Shell-Dateien, die
storage.cfg-Auswertung gegen eine echte Beispieldatei durchgespielt, und jeder
Befehl, den der Schritt absetzt, geht durch sh -n — keine andere Prüfung führt
diese Shell je aus. Drei Codex-Runden (R15), alle Befunde behoben.

Nicht geprüft: nichts davon lief je gegen echte Hardware. Die erste Übernahme
auf einem Proxmox-Host ist die Abnahme.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 05:01:44 +02:00
nexxo 97ac6c04e8 Release v1.3.87 — VM.Monitor gibt es in Proxmox 9 nicht mehr
tests / pest (push) Failing after 8m56s Details
tests / assets (push) Successful in 24s Details
tests / release (push) Has been skipped Details
Die Übernahme des ersten PVE-9-Hosts blieb in CreateAutomationToken stehen:
`pveum role modify` lehnt die GANZE Privilegienliste ab, sobald ein einziger
Name darin unbekannt ist — "invalid privilege 'VM.Monitor'". Die Liste stammt
aus der PVE-8-Zeit.

Ersatzlos gestrichen, nicht ersetzt: VM.Monitor gab Zugriff auf den
QEMU-Monitor, und den benutzt CluPilot nirgends. Geprüft, nicht vermutet — die
einzigen Treffer auf "monitor" im Provisioning betreffen Uptime Kuma. Auf PVE 8
fehlt damit ein Recht, das dort ohnehin niemand gebraucht hat; eine Liste passt
weiterhin auf beide Fassungen.

Dazu ein Test, der die angeforderten Namen gegen die gültigen hält — abgelesen
aus der Administrator-Rolle eines echten PVE 9. Die Liste ist die Verabredung
zwischen CluPilot und jedem Host, den es je übernimmt; sie hier zu prüfen ist
billiger als der Abbruch auf einer Maschine, die schon halb eingerichtet ist.

2245 Tests grün.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 03:42:37 +02:00
nexxo 4f04d9d29d Release v1.3.86 — der Grund von pveum, nicht nur die Feststellung
tests / pest (push) Failing after 8m49s Details
tests / assets (push) Successful in 21s Details
tests / release (push) Has been skipped Details
„could not converge the Proxmox automation role privileges", fünfmal
hintereinander. Die Meldung sagte, WAS scheiterte, und verschwieg das Einzige,
was weiterhilft: was `pveum` dazu geschrieben hat. Der Betreiber musste sich
auf den Host melden und den Befehl von Hand nachstellen, um eine Zeile zu
lesen, die daneben schon dastand.

Die Fehlerausgabe von `pveum role modify` steht jetzt in der Meldung — erste
Zeile, gekürzt: `pveum` stellt den Grund voran und breitet danach seine
Aufrufhilfe aus, und die wäre ein Bildschirm Text in einer Ereigniszeile.

2244 Tests grün.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 03:38:02 +02:00
nexxo b9fb401823 Release v1.3.85 — Handshake-Prüfung hing an einem Programm, das es nicht gibt
tests / pest (push) Failing after 8m54s Details
tests / assets (push) Successful in 22s Details
tests / release (push) Has been skipped Details
Die Übernahme brach in "WireGuard einrichten" ab: fünf Wiederholungen,
"WireGuard handshake not up yet". Geprüft wurde mit `ping -c1 -W2 <hub-ip>`.
Das fragte drei Dinge auf einmal und nannte nur eines: ob der Handshake steht,
ob ICMP durchkommt — und ob es `ping` auf der Maschine überhaupt GIBT.

Auf Hetzners `Debian-trixie-latest-amd64-base` gibt es das nicht. `iputils-ping`
ist im Image nicht dabei, und PrepareBaseSystem installiert
`curl gnupg ifupdown2 chrony`. Der Schritt scheiterte damit über einem Tunnel,
der stehen konnte. Auf dem alten Proxmox-Image war ping dabei, deshalb lief es
dort durch — dieselbe Pipeline, anderes Grundsystem.

Gefragt wird jetzt WireGuard selbst: `date +%s; wg show wg0 latest-handshakes`.
Die Uhr des HOSTS kommt in derselben Antwort mit, weil `latest-handshakes` eine
absolute Zeit ausgibt und ein Vergleich gegen UNSERE Uhr eine Zeitverschiebung
zwischen zwei Maschinen als Tunnelzustand läse. `date` ist in coreutils und
überall da.

Codex, zwei Runden:

- P1: `> 0` hiesse "hat jemals". WireGuard behält den Zeitstempel unbegrenzt,
  also meldete ein Wiederholungslauf über einem toten Tunnel "steht", und die
  Schritte danach wählten die Tunneladresse für SSH. Jetzt muss der Handshake
  frisch sein (180 s) und vom KONFIGURIERTEN Hub kommen — ein fremder Peer auf
  wg0 ist kein Beweis dafür, dass wir erreichbar sind.
- P1: Der neue Host-Versatz (.100) liess die Vergabe in einem Subnetz kleiner
  als /26 "erschöpft" melden, obwohl unten alles frei war. Der Versatz ist eine
  Bevorzugung, keine Bedingung: zweiter Durchgang von vorn.

Ausserdem, wie gewünscht: Hosts bekommen ihre Tunneladresse ab .100
(CLUPILOT_WG_HOST_OFFSET), Personen zählen weiter von unten. Fortlaufend
vergeben landete der erste Host zwischen zwei Notebooks, und wer eine Adresse
in einem Protokoll las, konnte nicht sagen, ob dahinter ein Mensch oder eine
Maschine steht.

2243 Tests grün.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 03:00:19 +02:00
nexxo 2f15d1e7a4 Release v1.3.84 — VPN-Zugänge nach Personen und Hosts getrennt
tests / pest (push) Failing after 9m6s Details
tests / assets (push) Successful in 20s Details
tests / release (push) Has been skipped Details
Personen und Hosts standen gemischt untereinander, nach Verbindungszustand
sortiert — ein Host-Zugang konnte also zwischen zwei Mitarbeitern stehen.
Beide sind Peers im selben Netz und haben dieselben Messwerte, aber es sind
zwei verschiedene Fragen: "wer von uns ist im Netz" und "welche Maschinen
hängen dran". Wer die eine stellt, liest die Antworten der anderen als
Rauschen.

Jetzt zwei Gruppen mit Überschrift und Anzahl. Eine Gruppe ohne Einträge wird
gar nicht erst gezeichnet — auf einer frischen Installation stünde sonst
"Hosts" über einer Lücke, bevor je einer angelegt wurde.

Getrennt wird nach der geladenen host-Beziehung, nicht nach `kind`: die
Plakette in der Zeile fragt dasselbe, und wer einen "Host-Zugang" liest, soll
ihn auch unter den Hosts finden. Ein adoptierter Peer (kind=system) mit
host_id gehört zu den Hosts, obwohl seine Art etwas anderes sagt.

Die Zeile selbst ist nach resources/views/components/admin/vpn-peer-row.blade.php
gewandert. Sie zweimal hinzuschreiben hätte geheissen, sie ab dem nächsten
Knopf an zwei Stellen zu pflegen — und die zweite fällt erst auf, wenn jemand
einen Host-Zugang sucht und ihn anders aussehen sieht als seinen eigenen.

Die Plakette bleibt trotz der Überschrift daneben: eine Zeile wandert beim
Suchen aus ihrer Überschrift heraus, und dann steht sie ohne sie da.

Codex: keine Befunde. 2237 Tests grün.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 02:34:15 +02:00
nexxo e0c0bb2c29 Release v1.3.83 — Fortschritt zeigte Übersetzungsschlüssel
tests / pest (push) Failing after 8m48s Details
tests / assets (push) Successful in 23s Details
tests / release (push) Has been skipped Details
Unter jedem Schritt stand `dashboard.step_state.done` statt eines Wortes —
im Adminbereich UND im Kundenportal, wo derselbe Stepper jemanden durch die
Einrichtung seiner Cloud begleitet.

Die Schlüssel gab es in keiner der beiden Sprachdateien. Die Paritätsprüfung
aus R16 konnte das nicht finden: sie vergleicht Deutsch mit Englisch, und die
beiden waren sich einig. Der Aufruf ist zusätzlich zusammengesetzt
(`__('dashboard.step_state.'.$state)`), also findet ihn auch keine Suche nach
festen Zeichenketten. Der neue Test rendert deshalb das Bauteil und prüft alle
vier Zustände in beiden Sprachen.

Dazu ein Test, der heute Nacht hochgegangen ist: PortalInvoicesTest verbot die
Zeichenkette '01.08.2026' — das Datum aus einer alten Attrappe. Heute ist der
1. August 2026, und die echten Rechnungen des Tests tragen das
Ausstellungsdatum von heute. Eine Zusicherung, die ein Datum verbietet,
verbietet auch den Tag, an dem es echt wird. Was die Attrappe ausmachte, war
ihr Satz, und den prüft der Nachbartest über 'Nächste Abbuchung' — stabil,
weil er nirgends sonst entsteht.

2235 Tests grün.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 02:08:28 +02:00
nexxo ac81bc5b98 Release v1.3.82 — Übernahme-Befehlszeile neu ausstellen
tests / pest (push) Failing after 8m50s Details
tests / assets (push) Successful in 23s Details
tests / release (push) Has been skipped Details
Die Zeile gab es genau einmal: auf der Seite direkt nach dem Anlegen. Wer sie
dort nicht kopierte — oder wem sie wegbrach, wie beim WireGuard-Fehler bis
1.3.80 — hatte danach einen Host in der Liste, für den es keinen Weg zu einer
Befehlszeile mehr gab. Ein zweites Anlegen scheitert an der eindeutigen IP,
also blieb nur Entfernen und von vorn.

HostTakeoverCommand behauptet in seinem Kopfkommentar, die Zeile werde "an
zwei Stellen gezeigt (beim Anlegen und beim Neuausstellen eines Codes)". Das
Zweite gab es nie — ein Kommentar, der eine Absicht beschreibt und wie eine
Beschreibung des Gebauten klingt.

Neu: ReissueTakeover als Modal auf der Host-Detailseite. Bestätigen zuerst
(R23), denn ein bisher ausgegebener Code wird dabei wertlos; danach steht die
ganze Zeile mit Kopieren-Knopf darin, nicht nur der Code.

Codex, zwei Runden, beide Male dieselbe Sorte Fehler von mir — die Ansicht
versteckt den Knopf, die Methode prüft nichts:

- P1: `issue()` erzwingt jetzt selbst, welcher Zustand zulässig ist. Ein
  Modal, das offen blieb, während der Host aktiv wurde, rief die Methode
  trotzdem; eine Livewire-Methode ist ohnehin von aussen aufrufbar.
- P2: `issue()` prüft die Tunnel-Einstellungen erneut. Sonst wird ein gültiger
  Code entwertet und dafür eine Zeile ausgegeben, der die WireGuard-Angaben
  fehlen — genau die kaputte Zeile, vor der die Warnung daneben steht.
- P1 der zweiten Runde: `disabled` gehörte nicht in die Liste der zulässigen
  Zustände. Es sieht aus wie "noch nicht fertig" und ist das Gegenteil —
  toggleMaintenance() schaltet einen LAUFENDEN Host so still. Der Knopf hätte
  eine Produktionsmaschine zur Neuinstallation angeboten.

Die Bedingung steht deshalb einmal am Bauteil (ReissueTakeover::eligible) und
wird von Ansicht und Methode gefragt: zwei Fassungen liefen auseinander,
sobald jemand eine ändert.

2231 Tests grün.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 01:00:53 +02:00
nexxo 8ab50f5650 Release v1.3.81 — Host anlegen scheiterte am WireGuard-Peer
tests / pest (push) Failing after 9m3s Details
tests / assets (push) Successful in 22s Details
tests / release (push) Has been skipped Details
Das Anlegen eines Hosts endete in der Konsole mit einem 500, bevor der
Betreiber den Einmal-Code je zu sehen bekam. Ursache war nicht das Anlegen,
sondern der Tunnel-Peer: HostEnrolment::issueWithKeys() rief `wg set wg0`
selbst auf — im Web-Request, also im `app`-Container. Der hat weder NET_ADMIN
noch /dev/net/tun; wg0 lebt im Provisioning-Container. Die Antwort war
"Unable to modify interface: Operation not permitted".

Der Peer geht jetzt als ApplyHostVpnPeer auf die Provisioning-Queue — genau
dorthin, wo ApplyVpnPeer es für die VPN-Zugänge längst richtig macht, mit
derselben `wireguard:hub`-Sperre und demselben Grundsatz: der Sollzustand
kommt beim Ausführen aus der Zeile, nicht aus dem beim Einreihen
festgehaltenen Wert. Der abgelöste Schlüssel reist als Wert mit, weil in der
Zeile zu diesem Zeitpunkt schon der neue steht.

Aufgefallen ist es nie, weil die Testsuite den Hub gegen FakeWireguardHub
tauscht — der bestehende Test blieb grün, während der echte Weg seit jeher
fehlschlug. Die zwei neuen Tests prüfen deshalb den WEG statt des Ergebnisses:
in der Anfrage bleibt der Hub unberührt, und der Auftrag liegt auf der
Provisioning-Queue.

Codex: 0 Fehler, 0 Sicherheitsbefunde. Dazu sein P2 — /.claude/ stand weder
im Index noch in .gitignore, ein `git add -A` hätte 153 MB als verschachteltes
Repo eingebettet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:43:10 +02:00
nexxo 7c4d4fd11c Zweite Fix-Runde: Regressionen der ersten (Trockenlauf, Doppelmail, Bestand, Isolierung) 2026-07-31 23:02:13 +02:00
nexxo b15fcc53a5 Fix-Runde: Reichweite der Sperre, keine doppelte Gebuehr, verlorene Nachricht 2026-07-31 22:57:56 +02:00
nexxo 4f037c8264 Mahnwesen: Einstellungen und die zwei Eingriffe von Hand 2026-07-31 22:43:52 +02:00
nexxo fb1ef3fc63 Rollen: aufklappbare Karten, gruppierte Rechte, Knopf im Raster 2026-07-31 22:39:46 +02:00
nexxo a64d41639b Mahnwesen: die sechs Kundenmails 2026-07-31 22:31:04 +02:00
nexxo a9f15449c6 Mahnwesen: Sperre und Freigabe, zusammen gebaut 2026-07-31 22:19:42 +02:00
nexxo 9899152d95 Mahnwesen: Tageslauf mit Gebuehrenrechnungen 2026-07-31 22:13:52 +02:00
nexxo 29c2718235 R20: Begruendung im Modal, und der Waechter sieht jetzt auch Listenzeilen 2026-07-31 22:09:05 +02:00
nexxo 7f7b654fcb Zahlungsprobleme: geplatzte Kaeufe aufzeichnen, beide Arten an einer Stelle 2026-07-31 21:58:52 +02:00
nexxo 9304382f62 Mahnwesen: Fall, Zeitplan und Gebuehrenstufen 2026-07-31 21:51:42 +02:00
nexxo 93e848d3f9 Rollen und Rechte in der Konsole zuweisbar 2026-07-31 21:46:22 +02:00
nexxo 6a6a672c11 Pruefergebnis als Modal statt am Seitenende 2026-07-31 21:36:43 +02:00
nexxo f91726d078 Testergebnis erscheint am Feld statt am Seitenende; DNS-Token bekommt einen Testknopf 2026-07-31 21:26:14 +02:00
nexxo b8a774b604 Portal: Karte tauschen, offene Rechnungen begleichen, Hinweis bei Faelligkeit 2026-07-31 21:12:31 +02:00
nexxo d6fd760a8b Offene Rechnungen bis zum Ende blaettern (Codex P2) 2026-07-31 20:58:21 +02:00
nexxo 0944de7cfa Passwortrichtlinie einstellbar, doppelte Anmelde-Mail behoben 2026-07-31 20:53:19 +02:00
nexxo 19ed6461cf Bereitschaft: das Absender-Postfach wird gefragt, ob es senden KANN 2026-07-31 20:35:58 +02:00
nexxo 82c6d62057 Stripe-Client: offene Rechnungen listen und einziehen 2026-07-31 20:04:09 +02:00
nexxo 0ae752d2ae Stripe-Client: SetupIntent und Vorgabe-Zahlungsmittel 2026-07-31 20:01:14 +02:00
nexxo 200df30b9d Hetzner-Cloud-DNS, Bereitschaftsseite aus sich heraus behebbar
Die alte Hetzner-DNS-API ist abgeschaltet (301 auf die Weboberflaeche).
Jede Bereitstellung starb in ConfigureDnsAndTls, nachdem der Kunde bezahlt
hatte. HttpHetznerDnsClient und DnsTokenCheck sprechen jetzt die Cloud-API:
Bearer statt Auth-API-Token, RRSets ueber {name}/{typ} statt Record-IDs, der
Zonen-Lookup entfaellt. Gegen das Live-Konto gemessen, nicht geraten.

Drei Fallen, die am Konto gemessen wurden: TXT-Werte muessen in
Anfuehrungszeichen (sonst 422, was als read_only gemeldet worden waere), ein
Name mit Zonensuffix wird STILL angenommen (201), und der Fake gab andere IDs
aus als der echte Client -- zwoelf Tests waren gruen ueber einem Abbau, der im
Betrieb geworfen haette.

Bereitschaftspunkte, die nur eine Shell beheben konnte, sind jetzt bedienbar:
SSH-Schluesselpaar erzeugen (Ed25519 ueber phpseclib, privater Teil direkt in
den Tresor), Stripe-Katalog abgleichen (Warteschlange, Trockenlauf, Ausgabe
wortgleich), Mailzustellung als Schalter statt MAIL_MAILER, Neustart der
Arbeiterprozesse. Stripe hat jetzt alle drei Werte in der Konsole: Secret Key,
Signatur-Secret (Tresor, je Modus getrennt) und Publishable Key (Klartext).

Codex-Review (R15), zwei P1 behoben: der Signaturschluessel faellt nicht mehr
vom Live- in den Testplatz, und eine Record-ID aus der alten API macht einen
Host nicht mehr unloeschbar -- was adressierbar ist, wandelt eine Migration um,
der Rest wird beim Loeschen laut uebergangen statt geworfen.

2109 Tests gruen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 19:50:01 +02:00
nexxo d321719180 Only a real zones array counts as an answer
Measured on the live server: a token with read AND write, the zone
clupilot.cloud present with fifteen records — and the console insisting the
account held no zones at all. Both screenshots contradict the message, so the
message was wrong, and the previous fix did not go far enough.

successful() is not enough. Something in the middle — a portal, a filter, a
proxy — answers with 200 and an HTML page. That body has no `zones` key, `??
[]` turned it into no zones, and the display concluded the Hetzner account was
empty. A 200 is not a promise about who answered.

So the body has to answer the question, not merely arrive: `{"zones": [...]}`
with an actual array, or it is reported as something else having spoken, with
the status and the first 120 characters of what came back. That last part is
what turns it from a verdict into a diagnosis — an operator who sees "Blocked by
policy" knows in one line what nothing else here could have told them.

A genuine `{"zones": []}` still means what it always meant.

2053 tests pass, assets build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 15:42:31 +02:00
nexxo c81a89ce5a A failed zone list is not an empty Hetzner account
The owner asked whether an empty token was being sent. It was not — an empty one
returns `missing` — but the question was worth following, and it found a fault
in the message I had just added.

The check classified 401 and 403 as `rejected` and then read the body. Every
OTHER unsuccessful response — 404, 429, 500, a cache's status page — has no
`zones` key, `?? []` turned that into no zones, and the console then stated "this
account holds no zones at all. The token probably belongs to a different Hetzner
project." A claim about somebody's account, derived from an error nobody looked
at, delivered with more confidence than the working case gets.

`successful()` is asked first now, and an unexpected status is reported as what
it is, with the number beside it: it says nothing about the zones, and it says
so. That is the distinction this check already draws between `unreachable` and
`read_only` — both are failures, only one of them tells you anything about the
token.

A genuinely empty list still means what it meant: 200 with zero zones is a token
for a project without zones.

The test uses a dataset rather than a loop. Http::fake() ADDS stubs instead of
replacing them, so a loop would have had the first status answer all four
iterations and the test would have proved one case three times over.

2050 tests pass, assets build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 15:36:47 +02:00
nexxo 0543a5a542 Let a failed measurement beat a green badge, and say which zone was missing
Three findings from the live readiness page, and the worst of them is the one
that looks like nothing.

Green "Erfüllt" sat directly above red "nicht in Ordnung", in the same row,
twice — for the DNS token and for the VM template. The badge came from
`satisfied` alone, the passive check that only establishes something IS
configured, and the measurement was rendered beside it without being allowed to
overrule it. Somebody scanning that list reads the badge, not the small print,
and walks away with "all green" while a measurement said it does not work. R19
names this exact shape — a call that reads as an assurance and is not one — as
worse than no check at all. The measurement wins now, for the badge, the icon
and the reason line.

zone_not_found was a dead end. It reads like "the token is wrong", so the
operator replaces the token — but a wrong token never gets that far: it comes
back as `rejected` from the 401 above. The token had just successfully listed
the zones. What is missing is the ZONE. The check now returns the zones it did
see, and the page puts them next to the one it wanted: looked for
clupilot.cloud, this account holds clupilot.com. The question answers itself.
An empty list says something else again, and gets its own sentence: the token
belongs to a different Hetzner project.

And two traffic tests were failing on main, unrelated to any of this, which is
why they were checked against a clean checkout before being touched. They build
"last month" as now()->subMonth()->format('Y-m'), and Carbon resolves that
calendrically: on 31 July it lands on 1 July, so the row meant to be last
period lands in the current one. Red on the 29th, 30th and 31st of every long
month, and today is the 31st. The production code does not have the trap —
currentPeriod() is now()->format('Y-m') with no arithmetic, and the two places
that do compute months already guard it — so this is the tests, and only the
tests.

2045 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 15:26:27 +02:00
nexxo 272cfe45e8 Keep IP addresses out of a certificate overview
127.0.0.1 was in the list on the live server, marked as console, showing "no
valid certificate: Connection refused" in red. It came from ADMIN_HOSTS, where a
bare IP is deliberately allowed — it is the way back into the console when a
name does not resolve. My filter only asked for a dot, and 127.0.0.1 has three.

Let's Encrypt does not issue for IP addresses, so that row was permanently red
and nobody could ever fix it. A red line that cannot be acted on is worse than
no line: it teaches the reader to skip the colour, which is the one thing the
overview needs them not to do.

Two checks now, not one: the address test, and a last label that is not numeric.
A TLD is never a number, and that also catches the forms FILTER_VALIDATE_IP lets
through.

The row already in the database goes away on the next sync, because otherwise my
mistake would sit on every installation that has already updated. Vanished
config rows are handled by what they carry: one that never had a certificate was
a mistake and is deleted; one that HAS a live certificate becomes a manual entry
instead, so nothing with a running expiry disappears without the operator
deciding. That is a change of mind from the previous commit, which said config
rows are never removed — this case showed the cost of that rule where the row
should never have existed.

2043 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 01:01:47 +02:00
nexxo cfaea7dc82 Fill the hostname page from the installation, not from a blank form
The page was empty, and that was a design fault rather than a missing button. I
built a register that starts blank — while the installation already serves half
a dozen names whose certificates were exactly what the operator wanted to see.
An overview you have to populate first does not answer "what do I have".

So the names are derived now, from the same configuration routes/web.php builds
its domain bindings from: SITE_HOSTS, APP_HOST, FILES_HOST, ADMIN_HOSTS. If a
name is in the environment, the application answers on it, and then it belongs
in this list without anybody typing it a second time. Opening the page syncs
them; the list is never empty again.

Syncing and measuring are deliberately separate. The sync costs nothing and runs
on page load. The measurement goes out over the network and runs on the button
or on a schedule — doing it on page load would mean waiting through half a dozen
TLS handshakes, and one of them is always the name that currently does not
resolve.

Which is the other half of what was missing: there was no overview because
nothing measured unless somebody pressed a button. A daily run at 04:17 fills
it, because the question that matters is not "is it valid right now" but "is
renewal running" — a certificate expiring in forty days is fine, the same one at
twenty means something has been broken for a week. Only a measurement taken
while nobody is looking can tell those apart.

The page now opens with four counts — total, valid, expiring soon, without a
certificate — and says when it last measured. A name that comes from the
environment is marked as such and cannot be removed here: it would come back at
the next sync, and a button that does nothing is worse than no button, because
it gets believed once.

2040 tests pass, assets build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 00:53:12 +02:00
nexxo 60c501e40d Put the page in the navigation, and stop files. handing out the portal
Two faults, both mine, both reported from the live server.

The page had a route and no navigation entry. A page reachable only by typing
its URL does not exist as far as the operator is concerned, and "Betrieb →
Hostnamen und Zertifikate" was exactly as findable as I had made it: not at all.
It sits under Betrieb rather than System because what is set there decides
whether an address ANSWERS — that is operations, not configuration.

And calling files.… without a path redirected to app.… The host-bound group only
claimed /bootstrap.tar.gz and /{file}; a bare / matches neither, so the request
fell through to the host-agnostic routes and landed on the portal. Two holes,
because / and a multi-segment path miss the placeholder for different reasons,
and both are closed now.

404 rather than a redirect, and that is the point rather than a detail. The
redirect told anybody who tried the name where the portal lives and that both
sit on the same machine — the one thing every other hostname in routes/web.php
is careful not to say. An address with nothing to offer has nothing to tell
either.

2037 tests pass, assets build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 23:21:15 +02:00
nexxo 006ce3568b Manage hostnames and watch their certificates from the console
files.clupilot.com is why this exists. DNS pointed at the right machine, the
env var was set, the release was deployed — and there was still no certificate,
because /etc/caddy/Caddyfile is maintained by hand and nobody thought of it as a
second, separate step. Nothing in the console would have said so. Three settings
looked correct and the address did not answer.

So the page holds two things side by side. The WISH — which names should be
served — and the REALITY: whether the name has a certificate and for how much
longer. The second is measured by opening a TLS connection and reading the
expiry, not by reading configuration, because the configuration is exactly what
looked right while the address was dead. verify_peer stays on: a certificate
that fails validation is not a certificate for this question, and a display that
called it valid would be the fake R19 records.

Applying goes through the existing agent, not a new channel. The console writes
a request, the path unit wakes the agent within a second, and the agent calls one
fixed command line of the root-owned helper.

What that helper is allowed to do is the careful part. It fetches the list
ITSELF rather than being handed one, and the list is HOSTNAMES, never Caddy
blocks — `php artisan clupilot:proxy-hosts` prints `<name> <purpose>` and nothing
else. Each name is matched against a strict pattern before it is used, and the
template around it lives in the helper, which the service account cannot touch.
install-agent.sh already states the principle for the sudoers grant: a grant is
only worth anything if the holder cannot change what it grants. A service account
that could write proxy configuration would have everything the proxy can do —
redirects anywhere, files from any directory.

Purpose is a column rather than a habit. A console name gets the network lock,
a public one does not, and a console name published without it looks exactly
like a working page.

Removing takes the name out of the list and NOT out of the running proxy. Two
decisions in one click, and the second one takes a site off the air.

There is deliberately no "renew" button. Caddy renews on its own at two thirds
of the lifetime; what an operator actually needs is a second attempt after an
issuance has failed, and that is a reload — which is what Apply does. Thirty days
is treated as a problem rather than a warning: at ninety days' lifetime, a
renewal should long since have run, so anything under it is not a tight
certificate but a renewal that is not happening.

The ACME contact falls back to the owner's address, because a contact nobody
reads is the step before expired customer certificates.

CONTRACT and HOST_STEP_NEEDS both move to 2, which is what tells a server
carrying the older helper to run the installer again — caught by the guard test
that compares the two halves.

2035 tests pass, assets build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 22:55:55 +02:00
nexxo 398028a57d Rebuild Add host as a page with one job, and let downloads through the gate
The page was a wall. Six paragraphs of procedure stacked above a form, so the
thing the page actually asks for — four fields — sat underneath an essay about
what would happen afterwards. It answered everything and showed nothing.

It is now built the way the settings page is built, because that page was
rebuilt for the same reason and there is no case for a second idiom: an eyebrow,
a title, a sticky rail on the left and panels of rows on the right.

The rail carries the six steps as two-word labels rather than paragraphs. That
is what the question at this moment actually is — where am I, how much is left —
and it fits in 232 pixels. The numbers carry the state; a tick beside them would
be a second sign for one statement, and a number can be counted.

What a step MEANS now appears in the row where it is due, not six times in
advance. The two provider steps get a panel of their own above the form, because
they have to be done before you save: the one-time code starts expiring the
moment you do. Everything after the form waits until there is something to say.

After saving the command becomes the page. Full width, its own framed block with
the warning in the header strip rather than floating above, and the two
remaining steps below it as ordinary rows.

Two token bugs went with it. `bg-canvas` does not exist — it was a class that
compiled to nothing, which is part of why the block looked wrong. And
`text-accent` on white is 2.9:1; the config says in as many words to use
accent-text for anything read, so the current step number does.

Also fixed, and it would have broken every takeover in production: PublicSiteGate
is appended to the whole web group, so while the site is hidden a server in a
rescue system fetching the archive would have received the 503 placeholder and
piped it into tar. The operator would have seen an unpack error on a machine
only reachable through the provider's console, with nothing pointing at a switch
in the admin area — and they are by definition neither on a management network
nor signed in, since leaving that state is the whole point. Exempted by ROUTE
NAME, not by hostname: the docblock rightly warns that exempting by host means
trusting a header the caller picks, but that warning is about the entire portal.
Behind these two routes are documents that are public anyway and an archive that
404s without a valid one-time code.

2026 tests pass, assets build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 21:51:30 +02:00
nexxo 156de12c8c Give the downloads a hostname of their own, with two halves
A domain that exists only to serve one installer is not worth a certificate.
One that also carries the AGB, the AV and the TOM is — those addresses go into
contracts and onto invoices and have to still resolve in three years, which an
address that moves with the next rebuild of the portal cannot promise. That
reason is what changed the answer.

files.… over cdn./storage./archiv.: a CDN is an edge network and this is not
one, so the name would be a lie the day a real CDN goes in front of it.
"storage" reads like object storage or customer data, and a customer seeing it
will wonder whether their files live there. "archiv" says superseded, which the
terms currently in force are not — and R13 keeps paths and names English
anyway.

Two halves on that host, with opposite rules, and that is the whole point of
giving it its own name:

Public — storage/app/files/public/, served to anyone, indexable, because
somebody looking for the terms should find them. Versioned filenames:
agb-2026-01.pdf, never agb.pdf, so a contract signed in January cannot come to
point at conditions written in July. The rule is written where somebody will
look for it rather than enforced, because a upload that rejects a filename
helps nobody.

Private — the installer, and it is not a file in that directory at all: it is
built from deploy/bootstrap on demand. The gate is the one-time enrolment code
that is ALREADY in the pasted line, resolved without being consumed, because
the code is still needed for every progress report and for the registration at
the end. No second secret: a dedicated download token would never expire, would
sit in shell histories forever, and would travel in the same line as the
WireGuard private key — protecting the least sensitive thing with the exposure
of the most sensitive one. 404 rather than 403 on a bad code, and noindex on
the response.

Path traversal is answered before it starts: basename() only, no directory tree
under public/ by design, and dotfiles refused. The test walks ../../.env three
different ways.

Empty FILES_HOST keeps the archive on the portal host exactly as before, so
nothing breaks in the window between setting the variable and the DNS record
existing.

The hostname had to move into phpunit.xml rather than a config()->set(): routes
are bound at boot, so a test that sets it afterwards is setting it too late.

2025 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 21:37:47 +02:00
nexxo 447d6fe0f3 Merge main into the host-takeover branch
One real conflict, in SyncStripeCatalogue::handle(), and both sides were right.

Main's parallel work resolves the adoption singletons and takes "before" counts
so a run can tell created from adopted. The operating-mode work refuses a
catalogue whose stored objects belong to the other mode, and records which mode
the catalogue now belongs to.

Kept both, with the mode guard first. It REFUSES, so it must not run behind
anything that has already built state; the counters only need to be in place
before the create loop, and they still are. Reversing that order would have the
command resolve singletons and take counts on a catalogue it is about to reject.

2017 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 21:21:45 +02:00
nexxo ba4996f316 Keep the platform on .com and the customers on .cloud
The owner caught me putting the archive on clupilot.cloud. That is the customer
zone, and looking into it turned up the same confusion already sitting in the
code — in the one place that matters most.

RegisterHostDns says "fsn-01.node.clupilot.com" in its own docblock. The
validation comment in Datacenters says it. ServicesTest writes it out verbatim.
The step itself built the name from config('provisioning.dns.zone') — the
CUSTOMER zone — so on this installation a host was actually called
fsn-01.node.clupilot.cloud. Three places asserting one thing and the code doing
another.

OfficialDomains explains why that matters and is worth not weakening: two
registrable domains by design, the company's for site, portal and console, the
instance zone for customer workloads. A Nextcloud is third-party software that
strangers sign into, and on the same registrable domain as the portal it shares
cookie scope with it. A host name in that zone does not break the separation,
but it puts it in question, and the next slip is more expensive.

So there is now a platform_zone, derived from APP_URL when unset, and
RegisterHostDns uses it.

My own archiveUrl was broken for a second reason. It fell back to url() when
APP_HOST is empty — and APP_HOST is empty on most installations, because empty
means "the portal answers on any hostname" and that is the default. Called from
the console, url() would have produced the CONSOLE hostname, and the line would
have 404'd on a machine that is not allowed to reach the admin area at all. It
takes the host from APP_URL now.

The script no longer guesses its own name. It used reverse DNS, then the tunnel
address, then a hard-coded clupilot.net — a third domain that appears nowhere
else in this project and was simply invented. PrepareBaseSystem has the same
invention. A guessed name does not stay guessed: it ends up in /etc/hostname, in
/etc/hosts, in every log line and in every certificate request the machine ever
makes. CluPilot knows the name because it just assigned it, so it passes --fqdn
and the script refuses without it. The value now also survives the reboot in the
arguments file, which it would not have.

The ACME contact moved to .com for the same reason it was wrong: the operator
does not live in the customer zone.

Open, and NOT decided here: the owner also wants the host to get a public DNS
record and a certificate on the .com name. RegisterHostDns deliberately writes
host names only into the tunnel's dnsmasq, and says why — publishing them hands
every scanner the internal subnet and roughly how many hosts sit behind it.
Nothing in the current design needs a public certificate for a host's own name;
Traefik serves customer domains, not this one. Reversing that is a security
decision and belongs to the owner, not to this commit.

1992 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 21:04:08 +02:00
nexxo 2194ab8079 Put the instructions where somebody reads them before they start
The previous commit shipped the guide AFTER the host was created. Its first two
steps happen at the provider — order the machine, boot the rescue system — and
whoever reads them there has the one-time code already in the clipboard with its
24-hour window running. The instructions were correct and in the wrong place,
which is its own kind of wrong.

Six steps now, not three, and the two provider ones are named rather than
assumed. They are marked "first" so the list does not read as something you can
start at the top of and work down: step 2 is half an hour of waiting, and doing
it inside a running code is exactly the mistake the page should prevent.

All six are visible the whole time, including the done ones and the ones still
to come. A guide that shows only the current step cannot answer "how much is
left", which is the question somebody has while a server they are paying for
sits in a rescue system.

It is on the hosts list too, collapsed. Somebody standing there may not have
ordered the machine yet — and that is step 1. Requiring them to click "add host"
to find out what the procedure is puts the answer behind the action it describes.

One component, two places, so the archive address in the instructions is the same
string the command line will carry. Two copies of that would drift, and the
difference would surface on a server that has already been paid for.

The step-4 command box moved into the guide rather than sitting beside it, so
after creating a host the whole procedure stays on screen. Somebody stuck at
step 5 should not be looking at a page that has become a single box.

1991 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 20:51:47 +02:00
nexxo 903ebdd2b2 Give the operator one line to copy and three steps around it
The console could describe a takeover it had no way to start. This is the
vertical slice that closes that: a one-time code, the archive the rescue system
fetches, and the page that says what to do with both.

The command carries EVERYTHING the script needs before the tunnel exists,
because there is nothing to fetch — that is the whole point of spec §5. Which
means CluPilot generates the WireGuard keypair and admits the peer at the hub
before the machine has ever booted, and hands the private half over in the line.
It is worthless within minutes: task 9 of the script replaces it with one
generated on the machine.

Shown exactly once. The database holds only the code's hash and never the
private key, so leaving the page does not bring it back — it mints a new code,
which invalidates the old one. That is deliberate: a glance at somebody's screen
should be worth nothing an hour later.

Which is also why save() no longer redirects. Sending the operator to the host
detail page sends them away from the only value they need, and an existing test
asserted that redirect — it now asserts the opposite, with the reason written
next to it.

SHA-256 rather than bcrypt for the code, and the reason is not speed. Both
endpoints have to FIND the host by the code; with bcrypt that means trying every
row. The code is 32 characters of CSPRNG output, so it has the entropy that
stretching exists to manufacture.

resolve() and claim() are separate because progress reports arrive BEFORE
registration. If reporting consumed the code, a host could never register after
its first message.

The archive URL is always the public hostname. The console runs under admin.…,
but this line executes on a machine that must not reach the admin area — it is
locked down for exactly that reason — so route() from the console would emit a
hostname that 404s on a server only reachable through the provider's console.

The page warns about missing tunnel settings BEFORE the host is created, not
after. An empty hub key produces a line that looks clean, copies fine, runs, and
ends in a tunnel that never handshakes — discovered on the machine, after
somebody has already paid for it.

The three steps lead with the rescue system, because that is the one nobody
knows by heart, and it says enabling is not the same as booting into it — the
script refuses a running production machine, which is what a half-done switch
looks like from the inside.

1986 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 20:40:19 +02:00
nexxo ecfced9257 Merge branch 'main' into claude/cool-sammet-368fee
tests / pest (push) Failing after 9m11s Details
tests / assets (push) Successful in 20s Details
tests / release (push) Has been skipped Details
2026-07-30 19:28:55 +02:00
nexxo 79654333a8 Name the recurring filter, and stop the spec contradicting itself
Three sentences left over from the wave before, all of the kind this branch
exists to end.

- "Every active Price" was one word short. activePricesFor() sends
  `type: recurring` as well as `active: true`, so a one-time Price is never
  listed — which made the hand-built-payment-link sentence false for a one-time
  link: neither listed nor archived. Both places say `recurring` now, and the
  docblock says why the filter is not a gap: createPrice() always sends
  `recurring[interval]` and is the only call that mints a Price at one of our
  Products, so the filter can only ever hide somebody else's one-time Price.
  A filter left out of a sentence is a gap the next reader goes looking for.
- The sweep test file still opened "The orphans nobody can adopt" — the exact
  narrowing the command's own docblock had to be corrected away from. It happens
  to describe the fixtures, which are all metadata-less, so it now says that
  instead of claiming it of the command.
- The design doc called --archive "harmlos" seven lines above the correction
  establishing that it is not, for an adoptable orphan. A document that
  contradicts itself seven lines apart teaches nobody anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 19:24:18 +02:00
nexxo a43be4186a Say what the sweep did not check, before telling anyone to archive
stripe:sweep-orphan-prices described itself as listing "what AdoptStripePrice
will not take over". It computes nothing of the sort: it lists every active
Price at one of our Products that no register row knows, adoptable or not — the
only question it asks is whether a row holds the id. And then it invited the
destructive path, with no word anywhere about running stripe:sync-catalogue
first.

Following that advice can make a plan unsellable. Archive an orphan the next
sync would have adopted and the sync can no longer see it: activePricesFor()
asks Stripe for active prices only, so adoption returns null and createPrice()
fires under the same key the crashed run used. For twenty-four hours Stripe
answers a repeated key by replaying the stored response — the archived Price's
id, in the body it had while it was live. The register writes it down as live,
and the checkout is refused. This is the same class of defect the branch exists
against: a comment that was eloquent and wrong.

- The docblock now says what the command lists, names the twenty-four-hour
  replay so nobody has to rediscover it, and the ordering requirement is in
  $description and printed on every path that found something, --archive
  included. The reassurance that archiving is harmless is now scoped to what
  THIS codebase sells; a payment link an operator built by hand against one of
  our Products is withdrawn all the same, and says so.
- stripe:sync-catalogue --dry-run says "created or adopted" again. A dry run
  counts intents without calling ensure(), so it cannot know which Stripe
  already holds — and a dry run is the first thing an operator runs after an
  interrupted one. The live line keeps both figures, where they are knowable.
- The sweep's plan side gets a test: an orphan at a family Product is reported
  and a known plan Price never is. Deleting either half of the lookup turns it
  red; the five tests it joins all stayed green for the products() half.
- The throttle test pinned "once" and neither "per price" nor "per day". A key
  mutated to a constant would have silenced every orphan after the first and
  kept it green — the silent no-op the branch is built against. It now plants a
  second orphan, and travels a day to hear the first one again.
- Two of my own documents corrected. The spec sentence claiming an adoptable
  orphan never appears in this list is what the implementation faithfully built;
  it holds only if a sync has run since the orphan appeared, which nothing
  enforced. And the two adoption classes do not carry the same run log:
  $adoptions on the Price side, $duplicates on the Product side.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 19:12:57 +02:00
nexxo 571acd0013 Check Stripe is configured on every path the sweep actually uses
The dry-run guard was copied from stripe:sync-catalogue, where it is right:
that command's --dry-run touches nothing, because everything it compares
already lives in our own tables. This command has no such local-only mode —
the report IS a live activePricesFor() call, dry-run or not — so `! $dryRun &&`
made the guard fire only on the one path that would have worked anyway. A
--dry-run against an unconfigured account fell straight into
HttpStripeClient's own exception instead of the clean error and self::FAILURE.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 18:48:49 +02:00
nexxo 277a6bab83 Give the operator something to do about an orphan nobody can adopt
The recognition step refuses a price that proves nothing about itself, and that
is right — adopting a stranger's price is worse than minting a second one. It
left an operator with a log line and no way to act on it.

Reports by default and touches nothing. With --archive the orphans stop being
sold, which is harmless for the reason it always is here: Stripe goes on billing
every subscription already on an archived price, and nothing is sold on a price
no row knows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 18:36:06 +02:00
nexxo 048e5ba81f Subtract only what was counted, not everything a singleton adopted
A figure may only subtract what it counted as an intent in the first place. The
family loop counts a Product before knowing whether it will be adopted or
minted, so an adoption there could honestly be subtracted back out — but
syncModules() has never counted a module's own Product as an intent at all; it
only counts a module's Prices. Reading AdoptStripeProduct::adoptions as one
run-wide delta could not tell the two apart, so a module Product adopted from
an interrupted run silently inflated "adopted" and shrank "created" by exactly
one, for a Product the command never claimed to have made in the first place.

The fix counts the family side locally, at the one call site that already
counts the intent, and leaves a module's Product out of both figures entirely —
adopted or minted, it was never counted, so neither number may move for it.
AdoptStripeProduct::$adoptions is gone with it: nothing reads a singleton-wide
total that cannot be attributed to one side or the other, and keeping it around
unread would be exactly the kind of state this codebase does not leave lying
about.

The duplicate-product report had the same shape of bug one line down: it read
straight off the singleton's list with no before/after snapshot, so a second
handle() call in one process would reprint a duplicate an earlier run already
named. Sliced to what this run itself added, the same way the counters beside
it already were.

Also: the unread `$charged` line in the adoption test that TDD had already
exercised as dead weight, and its now-unused PlanPrices import.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 18:20:27 +02:00
nexxo df7da334ee Say what was adopted, not that it was created
The count is taken before ensure() runs, so the sweep reported objects created in
Stripe when it had made none — and after the next interrupted run, that line is
the first thing a human reads. Telling the two apart is the whole point of the
recognition step.

Comparing the price id before and after the call does not distinguish them
either: there is no id before, in either case. AdoptStripePrice counts its own
adoptions instead, which is why it and its product sibling are now singletons —
resolved per call, a counter on the instance would never pass one.

Duplicate products are named in the report as well as the log. We deliberately do
not deactivate them, so only a person can resolve one, and a person reads this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:56:13 +02:00
nexxo 880b5f1998 Merge main into the operating-mode branch
Three real conflicts, and one file that did not conflict and mattered more.

Overview::notices(): both sides added notices. Kept all of them, and pointed
main's mail check at MailboxTransport::NON_DELIVERING — its own comment already
named the constant while the code carried a copy of the list.

billing.blade.php: main wrapped the page in a contract branch. The "payment is
not set up" error moved OUTSIDE it, because the customer most likely to meet
that message is the one buying for the first time, who has no contract yet and
would never have seen it from inside the @else.

ConsoleReportsRealDataTest: kept both sides rather than choosing. The renamed
test and its docblock explain why admin.systems_ok can no longer be asserted on
a bare install; main's mail pin still keeps "clean" clean in the mail
dimension. Picking one would have quietly weakened the other's claim.

HostStepsTest: git combined main's config()->set('provisioning.dns.zone',
'clupilot.com') with this branch's assertion on clupilot.cloud, and produced a
test that contradicted itself with no marker. Main's approach is the better one
— it pins the zone in the test and asserts the SHAPE of the name rather than
this box's domain — so its assertion stands.

HttpStripeClient merged silently and correctly: secret() still throws,
isConfigured() still reads the vault directly. Had it taken main's
filled($this->secret()), six callers that ask in order NOT to get an exception
would have become exception throwers, and the suite would have stayed green.
CheckoutWithoutStripeKeyTest is the lock that would have caught it.

StripeIdempotencyKeyTest (new on main) leaned on the environment fallback for
stripe.secret. That entry is strict now — no fallback in either direction, so
the .env cannot be a back door for a live key in test mode. It stores a real
vault row instead.

1966 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:53:20 +02:00
nexxo d6ec09a9b4 Recognise the product Stripe already has
The gap beside the guard. A run that creates a product and dies before storing
its id leaves an orphan exactly as a price does, and past the key's twenty-four
hours the next run makes a second one — after which activePricesFor() is asked
about the wrong product, price recognition goes blind for the whole family, and
every orphaned price on the first product is unreachable.

Two things are unlike the price side, deliberately. No money gate: a product
carries no amount, so its metadata is the whole of the proof. And a duplicate is
reported, never deactivated — archiving a price is provably harmless because
Stripe keeps billing subscriptions already on it, but deactivating a product
makes its prices unsellable and contracts can be running on those.

Silent about strangers, too. This asks for every active product on the account,
so an operator selling something else through it is an ordinary state, not a
finding worth a line on every run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:39:15 +02:00
nexxo 2a97d76b2d Rebuild the settings page out of panels and rows
**The withdrawal was invisible in two different ways.** It rendered only
while the fourteen days were still running, so the moment they expired the
whole block vanished — indistinguishable from a feature nobody built. And
the account it was looked for on has no contract at all: no package, no
subscription, nothing to withdraw from, so there was never anything for it
to be about.

It is now a row for every CONSUMER who has a contract, open or not: while
it runs, the deadline and the button; afterwards, the sentence saying why
it is over. A business never sees it — WithdrawalRight answers that, and
Settings::withdraw() refuses them again on the server. And an account with
no package says so and offers the packages, instead of a dead sentence in
a box.

**The page is built out of two pieces now.** x-ui.panel is a group: a card
with a header that names it. x-ui.row is a setting inside it: what it is
on the left, the control on the right, dividers between rows. That is the
shape every settings page worth copying uses, and it is the shape that
uses the width — a 280px label column with the control beside it fills the
line, where a stack of full-width inputs leaves two thirds of it empty and
a grid of cards of different heights walks down the screen as a staircase.

All four tabs are rebuilt on it, so the page reads as one designed thing
rather than four pages that happen to share a tab bar. Below `sm` the rows
stack, because on a telephone a label belongs above its field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:38:57 +02:00
nexxo 786318b6d4 Never tell anyone to delete a catalogue their contracts bill on
The refusal added in the previous commit had a state it read wrongly, and
the wrong reading was destructive.

record() sat at the END of handle() and in no try/finally, while
createProduct() and ensure() throw uncaught. A run that died after the
first object left hasStoredObjects() true and recorded() null. The same
state arises with no failure at all: CheckoutController → PlanPrices::
ensure() and BookAddon → SyncStripeAddonItems → AddonPrices::ensure()
mint missing prices and never call record().

In that state the next run — in the SAME mode — refused with the foreign
account message and its "Clear plan_families.stripe_product_id,
plan_prices.stripe_price_id and the … registers first", and
billing.catalogue_synced blocked with the same text. The objects were
from the account in force. The right move was to resume, which is what
the idempotency keys exist for; instead an operator was handed a delete
instruction for a catalogue live contracts are billed on.

Two changes:

  1. record() moves ahead of the creation loop, right behind the refusal.
     There it is already proved that either nothing is stored or what is
     stored belongs to the active account, so the moment carries the
     claim just as well — and a run that dies part-way can no longer
     leave a state that contradicts itself.
  2. "Origin never recorded" gets its own sentence and its own cure,
     separate from "established: other account".
     StripeCatalogueMode::matchesActiveMode() becomes
     belongsToAnotherMode(), which is only true where the other account
     is fact. The check still blocks — the origin cannot be proved — but
     it says "run the sync again", and it names no register to empty.

Red first:

  ⨯ it takes up a catalogue whose origin was never written down
      Failed asserting that 1 matches expected 0.
  ⨯ it leaves no half-built catalogue that contradicts itself when a run dies
      Failed asserting that null is identical to an object of class "App\Support\OperatingMode".
  ⨯ it tells an unrecorded origin apart from a foreign account

The two states are held apart by assertion, not by wording: only the
foreign-account sentence may name stripe_plan_prices, and the unrecorded
one must name stripe:sync-catalogue instead. The existing foreign-account
test keeps its teeth.

Full suite: 1817 passed, 6366 assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:30:56 +02:00
nexxo 18193a731d Stop asking for a signature the law does not want, and lay out the tab
**The owner is right about the agreement.** Art. 28 wants a CONTRACT, not
a ceremony — and a contract is concluded by incorporating the agreement
into the terms the customer accepts at checkout, which is exactly how
every hoster they have bought from does it. Nothing in the regulation asks
for a second, separate click. What it does ask is that the agreement is in
writing (electronic form included, Art. 28(9)), that it is the version in
force, and that the customer can obtain it.

So: the terms now say the agreement is part of the contract and needs no
separate signing, and name where it is. The card states that rather than
flagging the customer as outstanding — the warning chip is gone. The
button stays, reworded to "Zusätzlich bestätigen": a practice or a firm
that gets audited often wants an explicit record, and it costs a click.
The console's counter says how many confirmed rather than how many are
"open", because none of them are.

**The contract tab was a staircase** — a short card, a wide one spanning
both columns, then a short one again. It is one column of full-width
sections now, each with the same shape: a header row carrying the state
and its action, then the body. Nothing tracks sideways any more.

The package section states its own state in the header instead of a
paragraph in the body, the withdrawal sits under it as a row rather than a
second card, and the agreement's four files each have a button — reading
and keeping, agreement and measures, which the single "Herunterladen"
could not express.

Also: the generated version no longer carries "Aus dem Repository erzeugt
(clupilot:publish-dpa)" in the console's list. That is the command line,
not information.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:30:55 +02:00
nexxo fcea7ff3a8 Let the catalogue ask Stripe what products it already has
The price half of this question was answered; the product half was left open and
named as a known gap. It is the more consequential one: a second product makes
every activePricesFor() ask about the wrong one, so price recognition goes blind
for that whole family and every orphan on the first product is unreachable.

The whole active list comes back with no metadata filter, because Stripe cannot
search metadata and an account holds a handful of products. Deciding which are
ours belongs to the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:25:38 +02:00
nexxo 52aacddc56 Store the agreement where the web process can actually read it
Every link on the agreement page answered 404 — the download visibly, the
two inline ones just as dead — and nothing anywhere said why. The files
were there, whole and correct.

An artisan run inside a container is root; PHP-FPM is www-data. Laravel's
local disk creates a PRIVATE directory, which is 0700, so `dpa/` came out
as root-only. The web process could not enter it, `Storage::exists()`
answered false, and the routes did exactly what they were told: abort 404.
No exception, no log line, a page full of links to nothing.

Both writers — the command and the console's upload form — now store with
"public" visibility, which is the file MODE and nothing to do with the web:
0755/0644 instead of 0700/0600. The disk's root is outside the document
root either way, and the two routes still check who is asking. The command
also warns if a file it just wrote is not readable, because the alternative
is finding out from a customer.

The download itself now looks like one: the `download` attribute on both
anchors, and in the console a bordered control rather than a third link in
a row of two.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:19:39 +02:00
nexxo 4b99c6d5c6 Make the fake behave like Stripe where it matters
Three ways it did not. It replaced metadata where Stripe merges — and merging is
the property AdoptStripePrice's "write only when it differs" comparison is built
around. It handed metadata back uncast where the wire returns strings, so a test
could pass on an int that production's strict === rejects forever. And a tie on
`created` fell to insertion order, so two runs could adopt different prices.

Paging now throws instead of stopping when Stripe says there is another page and
the last item carries no id. invoiceLines() stops there and it costs a short
document; here it costs an unseen orphan and a second live price for one figure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:14:46 +02:00
nexxo 26cf1a72a0 Remember which Stripe account the catalogue was built in
The mode switches the credentials. It does not switch what was made from
them: a Price id and a Product id belong to the account that issued them,
and test and live are two accounts. Every credential got two slots on this
branch; the ids derived from them sit in single-valued columns.

So the planned sequence of this installation ended in the exact false
green this page exists to rule out — sync in test, store the live key,
switch, and billing.catalogue_synced went on reporting satisfied because
it only asked whether the column was filled. "Bereit für Livebetrieb",
and the first real order got "No such price".

Detection, not repair. Stripe does not put the account in the id —
prod_… and price_… look the same in both, only KEYS carry _test_/_live_
— so the origin cannot be read back out of a stored id, and asking
Stripe is out: this page reaches nothing over the network on a page load.
What is left is to write it down at sync time, which is what
App\Support\StripeCatalogueMode does.

One setting for the whole catalogue is only honest because
stripe:sync-catalogue now REFUSES a run into a catalogue that belongs to
the other account. Without that, the run would skip every row that already
carries an id, answer "already in step", and record an account it never
touched — the same lie one level down. The registers count as stored
objects too: inStep() takes a registered row as proof on its own for the
reverse-charge half.

The `breaks` sentence says what happens (checkout fails, no order) and
what actually helps, including the awkward half: re-running the sync is
not enough, the pointers and both registers have to be cleared first.

The slot migration backfills the one case it can prove: whatever is at
Stripe was made with the one key this installation has ever stored, so it
belongs to the account that key opens. Otherwise a long-synced catalogue
would read as "origin unknown" and the page would demand a re-sync nobody
needs.

Red first:

  ⨯ it does not call the catalogue synced when its ids belong to the other account
  ⨯ it says the sale is refused and a fresh sync is needed, not that a column is empty
  ⨯ it records the mode its objects were created in
  ⨯ it refuses to work into a catalogue that belongs to the other account
  ⨯ it syncs into the new account once the stale ids are cleared

Full suite: 1814 passed, 6357 assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:12:56 +02:00
nexxo d04ea712c7 Write the processing agreement and its annex, and let it be downloaded
The mechanism shipped without a document. Both are here now, generated
from this repository rather than uploaded — so the company is named from
CompanyProfile exactly as the invoices name it, the sub-processor list is
the three places customer data really leaves this application, and the
deletion deadlines are read off the commands that enforce them. A wording
change is a diff somebody can review, not a file that appears.

`clupilot:publish-dpa 1.0` renders both to PDF (TCPDF, the engine the
invoices already use), stores them on the private disk and puts them in
force; `--draft` stops short of that. The console's upload form is
untouched: a lawyer's revision arrives as a PDF and becomes a version like
any other.

**The annex says what this installation does, not what it would like to.**
Every measure in it was checked against the code first — the isolation is
one VM per customer because that is what provisioning builds, the backup
is a daily snapshot job at 02:00 because that is what RegisterBackup
creates, the deletion deadlines are the two prune commands. Claims the
system does not currently keep are NOT in the document, and they are named
in the handover instead. A TOM that promises more than the machine does is
the document an auditor reads before looking.

**Downloading** now carries the version in the filename —
CluPilot-AV-Vertrag-1.0.pdf — so "which fassung did I agree to" is
answerable from a downloads folder months later without opening anything.
Both sides get the button; inline stays the default, because a contract is
read before it is filed.

Placeholder register data is left off the document rather than printed: a
"FN 000000a" on a contract looks like a real number and is not one.

Also: the wordmark scan tripped on a paragraph whose first word is the
company name, for the second time. It now looks for the name as the whole
CONTENT of an element, which is what a lockup is, rather than for the name
after a tag, which is also how a sentence starts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:12:30 +02:00
nexxo 823eeaf413 Refuse a key that does not belong to the slot it sits in
put('stripe.secret', 'sk_live_REALMONEY') in test mode was accepted,
get() handed it out, billing.stripe_secret reported satisfied, and the
badge above it said "Testbetrieb". The strict rule closed the automatic
route into that state (no fallback); the typed one stayed open.

The prefix decides it without touching the network, and that rule now
lives in ONE place — OperatingMode::ofStripeKey() — called by the three
that were answering it separately: the slot migration (unchanged verdict,
`?? Live` for a value it cannot place), StripeCheck's `live` flag
(unchanged verdict, null stays false), and the readiness check, which
never asked at all. A key it cannot place is not reported as a
contradiction: this check only says what it can prove.

The two directions get their own `breaks` sentence, because the
consequences are opposite — real money moving while the console says
test, versus no money moving while the order looks paid.

The "Prüfen" button no longer contradicts the check either: the page
rendered only ok/reason, so a live key in the test slot answered
"Geprüft: in Ordnung". It now names the account the key belongs to.

Red first:

  ⨯ it refuses a live key sitting in the test slot
  ⨯ it refuses a test key sitting in the live slot
  ⨯ it says what the wrong key does, not that a field is empty

Guard tests (ConfirmInModal, ModalHeight, IconLayout, DisplayTimezone)
run with the blade change: green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:05:58 +02:00
nexxo 40532c6e02 Put the money gate in the class that promises it
AdoptStripePrice says in its own docblock that no adoption can move money, and
could compare three properties. The four that also decide what a Price charges
were filtered in HttpStripeClient — where FakeStripeClient cannot express them,
so no test of the class could reach the promise it makes, and a third client
implementation would drop the check in silence.

The client reports them now and the caller decides. Absent keys stay Stripe's
defaults, which is the direction that matters: read the other way round the gate
refuses every legitimate price and recognition becomes a no-op nothing alerts on.

And the unexplained-orphan warning is a state, not an event. It fired on every
sweep and every customer booking for as long as the orphan existed; once a day
per price is enough to act on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:04:25 +02:00
nexxo 5094f70c19 Say which slot the value in force is really coming from
Mode test, dns.token only in the :live slot, HETZNER_DNS_TOKEN also in the
.env — the state this installation is in the moment the migration runs:

  get()      -> "live-token-value"   (the stored live row is in force)
  source()   -> "environment"        -> card reads "Aus der Serverdatei"
  outline()  -> null                 -> no "Hinterlegt" block at all

So the console sent an operator to the .env to rotate a token that the
database supplies, where the edit would have had no effect. R19's dummy,
literally: a display that reads like an answer, is wrong, and stops the
next person looking.

source(), outline() and updatedAt() now ask the same one place get() does
(inForce()/fallback()), and the answer gets a THIRD value, 'stored_live'.
Not 'stored': that shows the Vergessen button, which would then have
pointed at the empty test slot. Strict entries (Stripe) still have no
fallback in either direction — pinned by its own test.

Red first (before the vault changed):

  ⨯ it names the live slot as the source when it is the one in force
    -'stored_live'
    +'environment'
  ⨯ it shows the outline and the date of the value that is actually in force

Two existing expectations moved from 'none' to 'stored_live' with the
reason written beside them: 'none' meant "nichts hinterlegt" about a
credential that was in use.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 17:01:33 +02:00
nexxo 507c976035 Pin the one question that must never throw
isConfigured() had no coverage at all: replacing its body with main's
`filled($this->secret())` left all 1795 tests green, while at runtime the
same mutation throws StripeNotConfigured. Six callers ask that question
precisely to avoid an exception — CheckoutController:87 would answer a
customer with a 500 instead of a sentence.

The merge makes it urgent rather than merely untidy: git auto-merges
HttpStripeClient.php without a conflict marker, so nothing forces anyone
to look at the one line where the two branches disagree.

Proved red by that exact mutation before it went in, then reverted:

  ⨯ it answers the configured question instead of throwing it
    StripeNotConfigured: No Stripe secret is stored for the [test] operating mode.
    at app/Services/Stripe/HttpStripeClient.php:376

Http::fake() + assertNothingSent(): the question is asked on every
checkout and may only read the vault.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:58:22 +02:00
nexxo 955f1b874f Deliver the processing agreement, and hold the proof it was accepted
Art. 28(3) GDPR wants a contract wherever personal data is processed on
somebody else's behalf, which is the whole of what this product does. "In
writing" there includes electronic form (Art. 28(9)), so a document the
customer can read plus a recorded acceptance is enough — no signature on
paper. The website already promises "AV-Vertrag inklusive", which means it
has to be obtainable without asking us for it. It was not obtainable at
all.

**The text is never this application's.** An operator uploads the document
their lawyer wrote, names the version, and publishes it; the measures ride
along as a second file, because they are an annex to the agreement and
"which measures applied when this customer accepted" has to have one
answer. Inventing the text here would have been worse than having none.

**Uploading and publishing are two acts.** Acceptance is per version, so
publishing leaves every customer who accepted the previous one outstanding
again — correct, and far too expensive to trigger by dropping a file on a
form. It goes through a confirmation modal that says exactly that (R23).

**The customer's side** is a card in the contract tab: read the agreement,
read the measures, one press to conclude it. What that press records is
what makes it evidence rather than a flag — the version, the moment, the
address it came from, and the login that pressed. Pressing twice is one
agreement (unique index, not a check somebody can forget), and a
superseded acceptance is kept rather than overwritten: it was true when it
was made, and the history is the point.

Nothing renders until a version is in force. A card offering an agreement
that does not exist is worse than the silence.

The files live on the private disk and are served through routes that
check who is asking — an agreement is not a public asset, and a guessable
URL to one would be a list of who our customers are. The customer route
takes no version parameter: which document applies is ours to say.

`dpa.manage` is its own capability on the OPERATOR guard. Whoever keeps
the platform running does not thereby decide what every customer is asked
to agree to — and a capability written under `web` since the 2026-07-29
move lands in a guard nothing authenticates against, which is how this one
first shipped answering 403 to a role that visibly had it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:56:57 +02:00
nexxo 1056dddc62 Take the address by field, fix the customer type, ask why they leave
**The billing address was one textarea**, so what landed in it was
whatever somebody typed: a postcode on the street line, a city with no
postcode, no country at all. It is four fields now — street, postcode,
town, country — in the customer's settings and in the console's customer
editor, which had its own free-text copy of the same field and would have
had an operator's correction silently reverted the next time the customer
saved.

`customers.billing_address` is neither dropped nor parsed: guessing which
line of an existing block is the street would put a postcode where a
street belongs on a document nobody can correct afterwards. It becomes the
COMPOSED form, rewritten from the fields on every save, so IssueInvoice
and everything else that reads it keep working, and a record nobody has
re-saved keeps the block it always had.

**The customer type is fixed once it is on record.** It decides the
fourteen-day right of withdrawal, and a business able to set itself to
"Privatperson" on this page is a business able to withdraw from a contract
it may not withdraw from. Registration asks the question; this page
reports the answer, and saveProfile refuses to answer it a second time —
in the component, not by hiding a radio, because a form that hides a
control has never stopped anybody who can post to /livewire/update. A
record created before the question existed may still be given one, once.

The panel is a fact rather than a form now, which is also what was odd
about it, and the sentence about who has the right of withdrawal is gone.

**Cancelling asks why.** A reason from a fixed list, required, with a note
beside it that "Sonstiges" cannot go without — the one fact worth having
about a departure, asked at the only moment the customer is looking at the
question. It lands on the contract (`cancel_reason`, `cancel_reason_note`).

And where the fourteen days are still running, the dialogue offers the
WITHDRAWAL as a way out of itself rather than as one more reason inside
it: withdrawing ends the service the same day and sends the whole amount
back, cancelling keeps the term that was paid for, and somebody entitled
to the first should not lose it by taking the second unasked. A business
is never offered it — WithdrawalRight answers that, not the template.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:41:32 +02:00
nexxo 865fd16f58 Put the price recognition back, all 3526 lines of it
2321854 removed it. That commit's own work — discarding an unconfirmed
registration and sweeping abandoned ones — is untouched and correct; what went
with it was every file of the Stripe price adoption merged an hour earlier:
AdoptStripePrice, IdempotencyKey, the unique-index migration, both test files,
the spec, the plan, and the edits in six more.

Nothing was lost. The commits stayed in history and the files stayed on disk,
untracked, which is what a staged deletion from a stale tree leaves behind.
Restored from 4eb90c8, the reviewed head, by explicit path — the fourteen paths
2321854 damaged and not one more, so the ten files that commit legitimately
added or changed keep exactly what it gave them.

This is the failure the repo's own rule exists to prevent: stage by path, never
`git add -A`, because a second session's index does not know what a first one
merged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:35:55 +02:00
nexxo 49d528dbea Fix round: point at the real tab, hide the button nobody can press, prove the quiet page load
Fixes four Important findings from the review of Task 12.

Corrected nine check `tab` values at the source (BillingChecks was already
correct; OnboardingChecks/ProvisioningChecks/DeliveryChecks all carried the
pre-redesign 'integrations' value, which is not a member of
Integrations::TABS): four onboarding checks now point at 'platform' or 'env'
depending on where their field actually is, five provisioning/delivery
checks point at 'services'. checkUrl()'s match-block resolver — praised as
correct for the six checks that point at a genuinely different admin page —
is unchanged; its `default` arm is now a pure safety net, not a route any
check actually relies on.

Wrapped the "Prüfen" button in @can('secrets.manage'): mount() admits
hosts.manage OR secrets.manage, but runCheck() requires secrets.manage alone,
so an Admin-role operator could see a button that 403s on press.

Fixed a docblock on HEARTBEAT_KEYS that asserted a staleness threshold was
kept in sync with OperationChecks::STALE_AFTER_MINUTES — that constant is a
key/settings-name map with no threshold of its own, and no such
synchronisation exists.

Three new tests, each confirmed red against a deliberately reintroduced
version of the bug it covers before being confirmed green: every check's tab
value against where its field actually lives, the run-check button hidden
from an operator who cannot press it, and Http::assertNothingSent() after a
plain page load.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 16:26:10 +02:00