clusev/docker/restart-sentinel
boban 72b82df7a1 fix(update): surface update failures instead of spinning forever
A dashboard-triggered update that failed on the host left the /update-progress
page spinning to a vague 10-minute timeout — no error was ever shown, because
nothing wrote a failure marker for the page to read. Fix: make failures visible
and prevent the credential-prompt hang class.

- docker/restart-sentinel/watch.sh: run update.sh under `timeout -k 30 1800` with
  its output tee'd to run/update.log; on ANY non-zero exit (update.sh or its
  exec'd install.sh) — and on the early compose-missing / updater-missing paths —
  write {"stage":"error"} to run/update-phase.json via write_update_error().
- update.sh: export GIT_TERMINAL_PROMPT=0 + GIT_HTTP_LOW_SPEED_* and wrap the pull
  in `timeout 300`, so a private-repo credential miss fails fast instead of
  hanging on a non-interactive prompt.
- DeploymentService::updatePhase(): whitelist the 'error' stage.
- update-progress.blade.php: add an error state (#js-status-error) + showError();
  the status-feed poll now stops and shows it on stage=error, marking the active
  phase red — instead of looping.
- lang/{en,de}/update.php: error_heading / error_hint / back_button.
- UpdateProgressTest: feed surfaces stage=error; page renders the error branch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 18:58:32 +02:00
..
README.md feat(versions): one-click 'update now' from the dashboard 2026-06-19 18:13:34 +02:00
clusev-restart.path feat(system): auto-restart sentinel — one-click restart via host watcher (no docker socket) 2026-06-14 23:42:50 +02:00
clusev-restart.service feat(system): auto-restart sentinel — one-click restart via host watcher (no docker socket) 2026-06-14 23:42:50 +02:00
clusev-update.path feat(versions): one-click 'update now' from the dashboard 2026-06-19 18:13:34 +02:00
clusev-update.service feat(versions): one-click 'update now' from the dashboard 2026-06-19 18:13:34 +02:00
watch.sh fix(update): surface update failures instead of spinning forever 2026-07-03 18:58:32 +02:00

README.md

Clusev restart sentinel (host watcher)

One-click stack restart from the dashboard — without giving the container the Docker socket.

How it works

  Dashboard (app container)                     Host
  ─────────────────────────                     ────
  System → "Jetzt neu starten"
        │
        ▼
  DeploymentService::requestRestart()
  writes storage/app/restart-signal/restart.request
        │   (shared bind mount: ./run on the host)
        ▼
  ./run/restart.request appears  ──────────►  systemd clusev-restart.path
                                                    │  (PathExists=)
                                                    ▼
                                              clusev-restart.service (oneshot)
                                                    │  runs watch.sh once
                                                    ▼
                                              docker compose -f docker-compose.prod.yml up -d
                                              rm ./run/restart.request

The app only ever writes a marker file. A scoped host-side systemd unit does the restart. The file content is a non-sensitive ISO timestamp; the trigger is the file existing, not what it holds.

Files

File Role
watch.sh The actual logic: once (systemd) restarts the stack + deletes the sentinel; update runs update.sh (git pull + re-install); loop is a standalone inotify/poll fallback for hosts without systemd (restart only).
clusev-restart.path systemd .path unit — PathExists= the restart sentinel, triggers the restart service.
clusev-restart.service systemd oneshot — runs watch.sh once as the clusev user.
clusev-update.path systemd .path unit — PathExists= the update sentinel (run/update.request), triggers the update service.
clusev-update.service systemd oneshot — runs watch.sh update as root (update.sh → install.sh needs root).

The update sentinel (dashboard → Version & Releases → "Jetzt aktualisieren") works exactly like the restart one, but the watcher runs update.sh instead of a plain restart, and consumes the sentinel before running so a persistent pull failure can't loop. install.sh installs both pairs of units.

Install (systemd — preferred)

From the project root (default path /home/nexxo/clusev; adjust the units if yours differs):

sudo cp docker/restart-sentinel/clusev-restart.path    /etc/systemd/system/
sudo cp docker/restart-sentinel/clusev-restart.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now clusev-restart.path

Check it:

systemctl status clusev-restart.path
# Trigger a test restart by hand:
touch ./run/restart.request
journalctl -u clusev-restart.service -f

Both units hard-code /home/nexxo/clusev. If your project lives elsewhere, edit PathExists= (in .path), and WorkingDirectory= / Environment=CLUSEV_DIR= / ExecStart= (in .service) before copying. The User= in the service must be able to run docker compose (root, or a member of the docker group).

Install (no systemd — poll/inotify fallback)

Run the watcher as a long-lived loop (e.g. under your own supervisor, tmux, or nohup):

CLUSEV_DIR=/home/nexxo/clusev docker/restart-sentinel/watch.sh loop

It uses inotifywait if available (apt-get install inotify-tools), otherwise polls every CLUSEV_POLL seconds (default 5).

Security note

  • The container gets NO Docker socket — it cannot run docker at all. It only writes storage/app/restart-signal/restart.request on a shared bind mount.
  • The restart is performed by a scoped host systemd unit (clusev-restart.service), the only component permitted to drive Docker. A panel compromise therefore cannot run arbitrary docker commands — at worst it can request a restart of the same stack.
  • watch.sh only deletes the sentinel after a successful docker compose up -d, so a transient failure is retried on the next trigger rather than silently lost.

Config (watch.sh env)

Var Default Meaning
CLUSEV_DIR repo root (resolved from the script path) dir holding docker-compose.prod.yml
CLUSEV_SIGNAL $CLUSEV_DIR/run/restart.request absolute path of the sentinel file
CLUSEV_POLL 5 poll interval (seconds) for the no-inotify loop fallback