3 Commits (6c92aa5dd7e2839b7a1268c8dfd3bfc00d01a6ad)
| Author | SHA1 | Message | Date |
|---|---|---|---|
|
|
6c92aa5dd7 |
Let an incident be deleted, and start measuring whether the hosts answer
tests / pest (push) Failing after 8m13s
Details
tests / assets (push) Successful in 21s
Details
tests / release (push) Has been skipped
Details
Three things, all from trying to actually use the incident page.
── Where the later entries go ───────────────────────────────────────────────
The form only ever creates the first one, and nothing said so — so it looked
as though the stage had to be chosen up front and there was only a dropdown
for the impact. The field now says what it becomes ("wird untersucht") and
where the rest are added: on the incident itself, once it exists.
── Deleting ────────────────────────────────────────────────────────────────
Asked for in order to test the thing at all, and right. A whole incident can
now be removed, with its entries, behind a confirmation in the product's own
dialog (R23).
Deliberately not the same as editing an entry, which there is still no way to
do: an update is a statement made at a time, and rewriting one afterwards is
what the record exists to prevent. Removing the whole incident is a different
act — it is how a test entry, or one published against the wrong service, is
taken back, and it leaves nothing half-standing behind. A test holds that
distinction so the delete button does not quietly become an exception to it.
── The hosts, which nothing was watching ────────────────────────────────────
`hosts.last_seen_at` drove the health dot in the console and was written
exactly once, at onboarding — RegisterCapacity stamped it and that was the
only write in the codebase. So every host read "offline" thirty minutes after
it was added, permanently, and the dot meant nothing. It is also why host
reachability could not appear on the status page: it was never being measured.
PingHosts asks every host's API the cheapest question it answers, once a
minute, and stamps the timestamp on success. Failure writes nothing on
purpose: absence IS the signal, and a second column counting failures would be
a second version of the same fact.
The status page has a fifth component for it. Counts only — never a host name,
never an address; the page is world-readable and the estate is not public
information. A host between five and thirty minutes silent reads as degraded,
not as healthy: "stale" is one of three answers healthState() gives and it is
not the good one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
|
|
0e3c50d9a1 |
Update on a schedule if the owner wants one, and say where incidents go
tests / pest (push) Failing after 8m22s
Details
tests / assets (push) Successful in 20s
Details
tests / release (push) Has been skipped
Details
── Where the status page gets its data ────────────────────────────────────── Asked in as many words, and fairly: the page for maintaining incidents exists (console → Betrieb → Störungen) and nothing anywhere said so. The page now carries the answer at the top — what is written there appears publicly at once, the four components measure themselves, and there is a link straight to the live status page. ── Automatic updates, inside a window, switchable ─────────────────────────── Tick it and name the days and hours; leave it alone and nothing changes. The button stays either way: manual is always possible, at any hour. It leaves exactly the same request the button leaves — same file, same agent, same run. A second path into a deployment would be a second set of rules to keep in step with the first, and the automation must not be able to do what an operator is prevented from doing by hand: it checks the same three things the button is disabled by (agent alive, something released, nothing running). Nothing guards against having already updated by hand, and nothing needs to: an installation that is current has nothing available, so pressing the button at four leaves the five o'clock window with nothing to do. That was the owner's own question and it answers itself. Defaults are Tuesday to Thursday, 05:00–06:00. Not Monday — the week starts and every fault costs double. Not Friday — nobody wants to repair a deployment on a Friday afternoon. Not three in the morning: if it goes wrong, the person who has to fix it is asleep. A window whose end is before its start wraps past midnight, because 23:00 to 02:00 is a reasonable thing to ask for and reading it as an empty range would silently mean "never". Switched on with no weekday selected is refused rather than corrected — it is a configuration that reads as working and does nothing. The clock is the wall clock (R19): an operator who writes 05:00 means five in the morning where they live, and a window that drifts by an hour twice a year is a window nobody trusts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
|
|
|
b85df4c141 |
Rebuild the status page as a status page
Asked for after looking at how real ones are built. The shape they all share is a convention, and a convention is what lets somebody find the answer without being taught the page: banner, components each with ninety days of uptime, then the written record underneath. This page had only the banner — the version that is of no use the morning after, because "Alle Dienste in Betrieb" is true and says nothing to somebody who was locked out the evening before. Three things are new. The ninety-day bar. It cannot be reconstructed after the fact: the monitoring table holds the last verdict per instance, not a history. So a sampler (clupilot:sample-status, every five minutes beside the monitoring sync) writes counts into one row per component per day, and the page divides them. A day with no row is drawn as a gap and says "nicht aufgezeichnet" — never green. Degraded samples count as reachable in the percentage but still colour the day: folding them into downtime reports a slow morning as an outage, ignoring them loses the only trace of it. The incident record. An operator reports a disruption in the console, posts updates as it develops, and the entry that says "behoben" is the same action that closes it — two separate steps is how a page ends up with a resolved note under an incident that still reports as ongoing. There is deliberately no way to edit an update: it is a statement made at a time, corrections are a further entry. An open incident may make a component look worse than the probes found it, never better. Scheduled maintenance, from the windows the console already keeps rather than a second list somebody has to remember to fill in. The measurement moved out of the controller into ServiceHealth, shared with the sampler, so the banner and the bars cannot disagree about what "healthy" means. The page now uses the shared design system and the site's header and footer. Its self-contained stylesheet was there to survive a broken asset build; it does not survive one — if the stylesheet cannot be served the application serving this page is already failing, and mid-deployment every route answers with the maintenance page. What the exemption bought was a third design nobody maintained. Found while building: durationMinutes had its operands the wrong way round, and diffInMinutes is signed, so the public page printed "Dauer -86 Minuten". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |