CluPilotCloud/docs/wireguard-hub.md

7.8 KiB

WireGuard management hub

Every Proxmox host CluPilot manages is reachable only over this tunnel — the hosts have no public SSH. The hub is the one endpoint that accepts peers.

Where it runs

Inside the queue-provisioning container, not on the VM host. That container is the only one that needs it: AdvanceRunJob, PurgeHost and RemoveWireguardPeer all onQueue('provisioning'), and it is the only worker consuming that queue. LocalWireguardHub shells out to wg / wg-quick, so the interface must exist in the same namespace as the PHP process.

Compose already grants what that needs: cap_add: NET_ADMIN, /dev/net/tun, net.ipv4.ip_forward=1, the published 51820/udp, and a named wireguard volume for /etc/wireguard. The container command runs wg-quick up wg0 before the worker starts, so the tunnel returns after every restart.

The host kernel module does not need modprobe: creating the first type wireguard link auto-loads wireguard.ko.

First-time setup

Only needed when the wireguard volume is empty (the config is not in git — it holds the private key).

docker compose exec queue-provisioning sh -c 'umask 077; wg genkey > /etc/wireguard/hub.key; wg pubkey < /etc/wireguard/hub.key'

Write /etc/wireguard/wg0.conf with that private key:

[Interface]
Address = 10.66.0.1/24
ListenPort = 51820
PrivateKey = <hub.key>

Then set in .envCLUPILOT_WG_ENDPOINT must be the address a Proxmox host can reach, not a container address:

CLUPILOT_WG_ENDPOINT=10.10.90.185:51820
CLUPILOT_WG_HUB_PUBKEY=<the pubkey printed above>

wg-quick save wg0 rewrites wg0.conf on every peer change, so keep manual edits to a minimum — comments in that file are lost.

Peer lifecycle

LocalWireguardHub::allocateIp() hands out the next free address in CLUPILOT_WG_SUBNET, skipping the hub IP, the network and the broadcast address, and skipping every hosts.wg_ip already taken. addPeer()/removePeer() run wg set and then wg-quick save wg0 — without the save, peers are lost on the next restart.

The console (/admin/vpn)

Create, block and remove accesses, and see who is actually connected.

Who sees and does what (designed with a Codex review; record-level rules live in VpnPeerPolicy, elevated ones are capabilities):

sees own sees all creates/blocks/revokes any downloads a stored config
every operator yes own only
Developer yes vpn.view.all own only
Admin / Owner yes vpn.view.all vpn.manage.all own only

Two decisions worth not undoing by accident:

  • Downloading is owner-only — deliberately not for Owner or Admin. A config is a credential, not metadata; sharing one destroys attribution, because WireGuard records the peer, not which person used a copied key. An administrator who needs access revokes that one and issues their own.
  • Issuing is not self-service. An access reaches the management network, so it needs vpn.manage.all even for oneself.

kind (staff|host|system) exists because a null owner previously had to mean two different things: "belongs to the admins" and "this is a host peer". A staff access without an owner is a bug; a peer merely found on the interface is system and belongs to nobody.

Losing the operator roles closes the owner's doors by itself — being staff is part of ownership, so a role edit outside revokeStaff() still cuts the tunnel paths. Revoking a staff member revokes their peers and purges their stored configs.

The web container has no wg0, so the page never talks to WireGuard. SyncVpnPeers runs on the provisioning queue (every minute from the scheduler, plus a throttled nudge from the page's five-second poll) and copies handshakes, traffic counters and source endpoints into vpn_peers. The page reads the database.

Two fields carry the state, and keeping them apart is what makes the page honest:

  • enabled — the operator's intent
  • present — what the hub actually reports

They disagree only while a change is in flight, which the UI shows as "wird angewendet". A sync never writes enabled.

Peers the host pipeline registers are adopted into the same list and named after their host, so "who is on the VPN" has no blind spots.

Things that are deliberate, and worth knowing before changing them:

  • Deleting is a soft delete. The row survives as a tombstone until the hub confirms the peer is gone. A hard delete would let the next sync see the still-present peer and adopt it back as a live access — silently restoring what was just revoked. The tombstone also keeps the key and tunnel address reserved meanwhile.
  • Every hub mutation takes Cache::lock('wireguard:hub') — ApplyVpnPeer, RemoveWireguardPeer and ConfigureWireguard. Without it a sync could read the interface around a mutation and then write what it saw.
  • Address allocation takes Cache::lock('wireguard:allocate'). Addresses come from one subnet but live in two tables (hosts.wg_ip, vpn_peers.allowed_ip), so no unique index can catch the other side's insert.
  • ApplyVpnPeer reads the current row, not the state captured at dispatch: a retried job must not undo a newer decision. A missing row can only mean the access is gone, so a stale payload may remove but never enable.
  • A host's key cannot be re-used for an operator access: wg set would rewrite that host's allowed-ip and cut its management tunnel.
  • Generated keypairs come from libsodium in PHP, not wg genkey — this container has no interface, and it keeps generation testable.
  • Storing a config is opt-in per access. Encrypted with VPN_CONFIG_KEY, never APP_KEY: a leaked application key must not also hand over the management network. Generate one with head -c 32 /dev/urandom | base64; empty means storing is switched off and the console says so. The secret is overwritten on revocation (secret_purged_at).
  • The password is asked on every retrieval, rate-limited to five attempts. Laravel's password.confirm keeps a 15-minute session stamp, which would authorise unlimited later downloads from an unattended browser.
  • The plaintext never enters a Livewire snapshot (ConfigHandoff hands over an opaque token), and the page stops polling while a config is on screen — otherwise every five-second poll would put the private key in a response.

Verifying it actually works

Mocked tests cannot prove a tunnel. Against the running stack:

  1. Start a throwaway peer that stands in for a Proxmox host: docker run -d --name wgpeer --cap-add NET_ADMIN --device /dev/net/tun clupilot-app:dev sleep 900
  2. Generate its keypair, register it through the app (not by hand) — app(WireguardHub::class)->addPeer('<peer pubkey>', '10.66.0.2').
  3. Bring up the peer with Endpoint = 10.10.90.185:51820, AllowedIPs = 10.66.0.0/24, PersistentKeepalive = 25.
  4. wg show wg0 on both sides must report a recent handshake.
  5. Prove the direction the pipeline needs — hub → peer TCP, which is what the SSH steps use: listen on 10.66.0.2:2222 in the peer and fsockopen to it from queue-provisioning. A handshake alone does not prove routed TCP.
  6. Clean up with removePeer() and delete the container.

Verified end-to-end on 2026-07-25: handshake established, hub→peer TCP banner received, peer persisted to wg0.conf by wg-quick save and removed again.

Before production

  • The endpoint becomes the production VM's address; regenerate the keypair there (the dev private key must never travel).
  • 51820/udp is the only port that has to be open inbound — the console and the Proxmox hosts stay off the public hostnames (see config/admin_access.php).
  • Back up /etc/wireguard/ (the wireguard volume): losing the hub private key means re-peering every host.