auto: docs sync 2026-08-14T12:00:03+00:00

Files changed: engine/.embcache.json engine/changelog.md engine/lint-report.md vault/.obsidian/workspace.json vault/docs/hardware/environment.md vault/docs/services/services.md vault/projects/fleet-storage-memory-upgrade.md vault/projects/navi-recon-separation.md vault/runbooks/omv-add-nfs-share.md vault/runbooks/peertube-sitemap-redis-oom.md vault/runbooks/pve-guest-park-and-adopt.md
This commit is contained in:
echo6-autocommit 2026-08-14 12:00:03 +00:00
commit afaf019076
11 changed files with 697 additions and 46 deletions

File diff suppressed because one or more lines are too long

View file

@ -241,3 +241,9 @@
## 2026-08-12T09:00:01Z — sweep complete (0 docs selected)
## 2026-08-13T09:00:01Z — sweep complete (0 docs selected)
## 2026-08-14T09:00:01Z — sweep run
- end: 2026-08-14T09:00:18Z
- mode: incremental
- docs selected: 7
- processed: 7 | written: 7 | flagged: 0 | errors: 0

View file

@ -1,6 +1,6 @@
# Vault Lint Report
Generated: 2026-08-14T00:00:12Z | Docs scanned: 107 | Elapsed: 0.0s
Generated: 2026-08-14T09:00:01Z | Docs scanned: 112 | Elapsed: 0.0s
## Summary
@ -8,7 +8,7 @@ Generated: 2026-08-14T00:00:12Z | Docs scanned: 107 | Elapsed: 0.0s
|----------|-------|
| ERROR (dead links) | 22 |
| WARN (schema) | 0 |
| INFO (orphans) | 39 |
| INFO (orphans) | 35 |
### WARN breakdown
- Missing frontmatter block: 0
@ -42,7 +42,6 @@ Generated: 2026-08-14T00:00:12Z | Docs scanned: 107 | Elapsed: 0.0s
## INFO — Orphan Notes (no incoming links, capped at 40)
- no incoming links: runbooks/add-peertube-channel.md
- no incoming links: runbooks/authentik-access-groups.md
- no incoming links: runbooks/authentik-create-invitation.md
- no incoming links: runbooks/authentik-oidc-application.md
@ -56,7 +55,6 @@ Generated: 2026-08-14T00:00:12Z | Docs scanned: 107 | Elapsed: 0.0s
- no incoming links: projects/fleet-platform-baseline.md
- no incoming links: docs/software/geo-tools.md
- no incoming links: glossary.md
- no incoming links: runbooks/headless-browser-page-verification.md
- no incoming links: runbooks/headscale-oidc-boot-order.md
- no incoming links: runbooks/ia-cli-reference.md
- no incoming links: runbooks/ia-download-mirror.md
@ -71,10 +69,8 @@ Generated: 2026-08-14T00:00:12Z | Docs scanned: 107 | Elapsed: 0.0s
- no incoming links: runbooks/meshtasticd-sim-nodes-runbook.md
- no incoming links: runbooks/nordvpn-lxc.md
- no incoming links: runbooks/pg-backup.md
- no incoming links: runbooks/pi-nas-omv-runbook.md
- no incoming links: runbooks/pipeline-patterns.md
- no incoming links: runbooks/proxmox-create-ubuntu-vm.md
- no incoming links: runbooks/proxmox-onboard-node.md
- no incoming links: runbooks/pymc-repeater-kiss-tnc-reenumeration.md
- no incoming links: runbooks/recon-operations.md
- no incoming links: runbooks/recon-service-integration.md
@ -108,17 +104,17 @@ Matt decides whether to create a real doc — when he does, future sweeps will l
| Term | Docs mentioning it |
|------|--------------------|
| `tailscale` | 40 |
| `docker` | 34 |
| `proxmox` | 31 |
| `docker` | 35 |
| `proxmox` | 34 |
| `headscale` | 24 |
| `peertube` | 24 |
| `meshtastic` | 19 |
| `peertube` | 19 |
| `mailcow` | 17 |
| `immich` | 15 |
| `nextcloud` | 14 |
| `element` | 14 |
| `forgejo` | 13 |
| `immich` | 13 |
| `nextcloud` | 13 |
| `aida-nebra` | 12 |
| `livesync` | 12 |
| `jellyfin` | 12 |
| `vaultwarden` | 12 |
| `aida-nebra` | 11 |
| `jellyfin` | 11 |

View file

@ -199,17 +199,22 @@
},
"active": "17bd4a6166f789d0",
"lastOpenFiles": [
"docs/services/services.md.tmp.2730138.e216805f7b3b",
"docs/hardware/environment.md.tmp.2730138.982fd86dec7c",
"docs/services/services.md.tmp.2730138.e01d11c7a6b5",
"docs/hardware/environment.md.tmp.2730138.c76433dc40b4",
"projects/navi-recon-separation.md",
"projects/navi-recon-separation.md.tmp.2730138.fb2f5474a2bc",
"projects/fleet-storage-memory-upgrade.md",
"projects/fleet-storage-memory-upgrade.md.tmp.2730138.f2694abef97e",
"runbooks/peertube-sitemap-redis-oom.md",
"runbooks/peertube-sitemap-redis-oom.md.tmp.2730138.acfcfa86751d",
"runbooks/omv-add-nfs-share.md",
"runbooks/omv-add-nfs-share.md.tmp.2730138.04a2d2fb1ec7",
"runbooks/pve-guest-park-and-adopt.md",
"runbooks/pve-guest-park-and-adopt.md.tmp.2730138.216ea270cee7",
"docs/services/services.md.tmp.1912821.a9007ef6f645",
"projects/fleet-patch-audit.md.tmp.3598161.730170652bdf",
"projects/fleet-patch-audit.md.tmp.3598161.d03a9fb9997e",
"projects/fleet-patch-audit.md.tmp.3598161.7f4127512795",
"projects/fleet-patch-audit.md.tmp.3598161.9e2c0bd025ae",
"projects/fleet-patch-audit.md.tmp.3598161.dc24c4c7df8b",
"projects/fleet-patch-audit.md.tmp.3598161.17f66ff46df8",
"docs/services/services.md.tmp.3598161.5e23f3a83462",
"docs/services/services.md.tmp.3598161.69bdec3ce793",
"glossary.md.tmp.3598161.3053e87ea7b8",
"glossary.md.tmp.3598161.9356baf31ea3",
"runbooks/conduit-operations.md",
"docs/software/conduit.md",
"archive/projects/vaultwarden-plan.md",
@ -231,11 +236,6 @@
"2026-06-19.md",
"Untitled.canvas",
"docs/hardware/environment.md",
"projects/fleet-patch-audit.md",
"rules/proxmox.md",
"docs/services/services.md",
"docs/software/central.md",
"docs/software/navi.md",
"assets/echo6yellow_logo_422x422_square.png",
"assets/echo6yellow_logo_422x81.png",
"assets/echo6_logo.png",

View file

@ -6,11 +6,11 @@ tags:
aliases: []
related:
- [[fleet-platform-baseline]]
- [[fleet-storage-memory-upgrade]]
- [[ip-allocation]]
- [[headscale-onboard-node]]
- [[toc-cortex-pve9.2-update]]
- [[proxmox-create-ubuntu-vm]]
updated: 2026-07-18
updated: 2026-08-14
---
# Echo6 Environment Reference
@ -23,7 +23,7 @@ Five nodes running Proxmox VE:
| data | 192.168.1.240 | 100.64.0.6 | AMD Ryzen 7 PRO 5750GE, 1TB NVMe + 1TB SATA SSD | 32GB DDR4-3200 | Database [[services]] |
| utility | 192.168.1.241 | 100.64.0.5 | AMD Ryzen 7 PRO 5750GE, 512GB NVMe | 32GB DDR4-3200 | Utility [[services]], monitoring |
| cloud | 192.168.1.242 | 100.64.0.4 | Intel i7-12700T, 512GB NVMe | 32GB DDR4-3200 | Cloud storage, personal [[services]] |
| media | 192.168.1.243 | 100.64.0.3 | Intel i7-14700T, 2x 512GB NVMe | 32GB DDR5-5600 | Media server, *arr stack |
| media | 192.168.1.243 | 100.64.0.3 | Intel i7-14700T, 2x 512GB NVMe | 32GB DDR5-4800 | Media server, *arr stack |
| toc | 192.168.1.244 | 100.64.0.13 | Workstation (i9-10900X) | 64GB DDR4 | GPU compute, AI/ML workloads |
### Node Storage Details
@ -35,6 +35,23 @@ Five nodes running Proxmox VE:
| media | 2x Intel SSDPEKNU512GZH 512GB (NVMe) | — |
| toc | 512GB NVMe | — |
**Free positions (verified live 2026-08-14):** data has none — only three external PCIe root ports exist and all are populated (NVMe, NIC, Wi-Fi), and the Wi-Fi M.2 is E-keyed so it cannot take a storage drive. utility and cloud each have an empty 2.5" SATA bay. media's second M.2 holds a leftover BitLocker Windows install (serial `PHKA142402U8512A`; the Proxmox drive is `PHKA142504HP512A`). toc has 8 unpopulated SATA ports and 4 free PCIe slots. pi-nas has one free SATA port (`ata5`). Placement plan is [[fleet-storage-memory-upgrade]].
**Memory ceilings:** data and utility are hard-capped at 32 GB and already there. cloud and media max at 64 GB. toc has 2 of 8 DIMM slots free but takes full-size DDR4 DIMMs, not SODIMM, and currently runs all modules at 2133 MT/s because rated speeds are mixed (3200/2666/2133).
### Cluster Backup Storage
`pinas-backup` — NFS `192.168.1.245:/export/pvebackup`, `vers=3`, content `backup`, mounted at `/mnt/pve/pinas-backup` on all five nodes. Backed by pi-nas `sdc1` (~19 TB free), deliberately not `sdd1` which carries PeerTube. Added 2026-08-14 for park-and-adopt guest relocation — see [[pve-guest-park-and-adopt]] and [[omv-add-nfs-share]].
### Non-Cluster Physical Hosts
| Host | Local IP | Tailscale | Hardware |
|------|----------|-----------|----------|
| pi-nas | 192.168.1.245 | 100.64.0.21 | Raspberry Pi 5, 8GB soldered, JMicron JMB58x 5-port SATA HBA, 2x 3TB btrfs RAID1 + 2x 24TB single ext4 |
| aida-nebra | 192.168.1.253 | 100.64.0.9 | Raspberry Pi Compute Module 3, Cortex-A53, 906MiB soldered, eMMC boot, RAK4631 on ttyACM1 |
Neither is upgradable — soldered RAM, no DIMM sockets.
### Node Hardware Identifiers
| Node | Make/Model | Serial/Service Tag |
@ -115,7 +132,7 @@ Five nodes running Proxmox VE:
| **edge1** (rebuilt Contabo VPS) | 5.189.158.149 | 100.64.0.40 | Debian 12 + Proxmox 8.4.19, **mail-only** — Mailcow in CT 101; host [[caddy]] + mailcow-dnat.service; rebuilt [[2026-06-19]] |
| edge2 | 184.174.35.153 | 100.64.0.26 | Contabo Cloud VPS 30 NVMe — Proxmox VE 8.4.19 (LXC-only), 8c/24GB/400GB — **permanent front door** for vault/forge/notes/auth/matrix/element/vpn/proxmox.echo6.co + idahomesh.com |
*Last updated: [[2026-06-19]] — Contabo VPS rebuilt as edge1 (mail-only, Debian 12 + Proxmox 8.4.19, 5.189.158.149 / tailnet 100.64.0.40); Mailcow CT 101 (10.10.10.2) on edge1; edge2 is now the permanent front door for all other services; Headscale node `contabo` moved to 100.64.0.40; previously added edge2 CT 107 (headscale), CT 106 (matrix), CT 105 ([[authentik]]), CT 104 (livesync), CT 103 (forgejo), CT 102 (vaultwarden)*
*Last updated: [[2026-06-19]] — Contabo VPS rebuilt as edge1 (mail-only, Debian 12 + Proxmox 8.4.19, 5.189.158.149 / tailnet 100.64.0.40); Mailcow CT 101 (10.10.10.2) on edge1; edge2 is now the permanent front door for all other [[services]]; Headscale node `contabo` moved to 100.64.0.40; previously added edge2 CT 107 (headscale), CT 106 (matrix), CT 105 ([[authentik]]), CT 104 (livesync), CT 103 (forgejo), CT 102 (vaultwarden)*
## LXC Containers
@ -144,7 +161,7 @@ Five nodes running Proxmox VE:
| matrix | edge2 (CT 106) | 10.10.10.24 | 100.64.0.37 | Matrix stack ([[synapse]] + MAS + Element + [[mautrix_signal]]; migrated from Contabo 2026-06-18) |
| headscale | edge2 (CT 107) | 10.10.10.25 | 100.64.0.38 | Headscale + Headplane tailnet control plane (migrated from Contabo [[2026-06-19]]) |
> **Note (2026-06-19):** edge2 CT placements CT 102107 confirmed; Forge git-SSH DNAT (`forgejo-ssh-dnat.service`) is a permanent systemd unit on edge2 host.
> **Note ([[2026-06-19]]):** edge2 CT placements CT 102107 confirmed; Forge git-SSH DNAT (`forgejo-ssh-dnat.service`) is a permanent systemd unit on edge2 host.
## IP Allocation Scheme
@ -172,7 +189,7 @@ Current registered nodes (25 total):
| utility | 100.64.0.5 | Proxmox |
| data | 100.64.0.6 | Proxmox |
| meshmonitor | 100.64.0.7 | LXC |
| caddy | 100.64.0.8 | LXC |
| [[caddy]] | 100.64.0.8 | LXC |
| aida-nebra | 100.64.0.9 | Pi |
| matt-desktop | 100.64.0.10 | Desktop |
| nextcloud | 100.64.0.11 | LXC |
@ -191,11 +208,11 @@ Current registered nodes (25 total):
| gl-a1300 | 100.64.0.29 | Router |
| bluefin | 100.64.0.30 | Desktop |
| wordpress | 100.64.0.31 | LXC (edge2 CT 101 — now runs Grav CMS; hostname unchanged) |
| meshai | 100.64.0.32 | LXC |
| [[meshai]] | 100.64.0.32 | LXC |
| vaultwarden | 100.64.0.33 | LXC (edge2 CT 102) |
| forgejo | 100.64.0.34 | LXC (edge2 CT 103) — node id 46 |
| livesync | 100.64.0.35 | LXC (edge2 CT 104) — migrated 2026-06-16 |
| authentik | 100.64.0.36 | LXC (edge2 CT 105) — node id 48, migrated 2026-06-18 |
| [[authentik]] | 100.64.0.36 | LXC (edge2 CT 105) — node id 48, migrated 2026-06-18 |
| matrix | 100.64.0.37 | LXC (edge2 CT 106) — migrated 2026-06-18 |
| headscale | 100.64.0.38 | LXC (edge2 CT 107) — migrated 2026-06-19 |

View file

@ -10,7 +10,7 @@ related:
- [[lxc-service-migration]]
- [[meshtastic-headscale-runbook]]
- [[expose-service-edge2]]
updated: 2026-08-05
updated: 2026-08-14
---
# Current Services Inventory
@ -32,12 +32,12 @@ updated: 2026-08-05
| [[argus]] | utility (CT 103) | 192.168.1.103:8080 | Internal | Python app on :8080 — OSINT intelligence gathering platform |
| [[central]] | utility (CT 104) | 192.168.1.104:8000 / 100.64.0.12 | central.echo6.mesh (mesh) | Data-hub spine — ~25 adapters → NATS/JetStream → TimescaleDB; serves traffic tiles to [[navi]] — see [[central]] |
| NATS/JetStream ([[central]]) | utility (CT 104) | 192.168.1.104:4222 / :8222 | Internal | [[central]] backend message bus (NATS :4222 client, :8222 monitoring) |
| TimescaleDB/PostGIS ([[central]]) | utility (CT 104) | 192.168.1.104:5432 | Internal | Central backend time-series + geospatial database (PostgreSQL 16 + TimescaleDB + PostGIS) |
| TimescaleDB/PostGIS ([[central]]) | utility (CT 104) | 192.168.1.104:5432 | Internal | [[central]] backend time-series + geospatial database (PostgreSQL 16 + TimescaleDB + PostGIS) |
| [[authentik]] | edge2 (CT 105) | 100.64.0.36:9000 | https://auth.echo6.co | SSO provider (Echo6 branded, custom CSS, dark theme) — fronted by edge2 host [[caddy]] (reverse_proxy 100.64.0.36:9000); **migrated from Contabo 2026-06-18** |
| Forge (Forgejo) | edge2 (CT 103) | 100.64.0.34:3001 HTTP / :2222 SSH (via edge2 DNAT) | https://forge.echo6.co | Git server — fronted by edge2 host [[caddy]] (reverse_proxy 100.64.0.34:3001); git SSH via iptables DNAT on edge2 (forgejo-ssh-dnat.service) — **migrated from Contabo 2026-06-16** |
| Headscale | edge2 (CT 107) | 100.64.0.38:8084 | https://vpn.echo6.co | Tailscale coordination (OIDC enabled) — fronted by edge2 host [[caddy]] — **migrated from Contabo [[2026-06-19]]** |
| Headplane | edge2 (CT 107) | 100.64.0.38:3100 | https://vpn.echo6.co/admin | Headscale web UI (OIDC via [[authentik]]) — fronted by edge2 host Caddy**migrated from Contabo [[2026-06-19]]** |
| Mailcow | **edge1 CT 101** (10.10.10.2) | 5.189.158.149 | https://mail.echo6.co | Email server (privileged LXC on rebuilt Contabo VPS, updated commit 52a41b4d / SOGo 5.12.8) — **rebuilt in-place 2026-06-19** |
| Headplane | edge2 (CT 107) | 100.64.0.38:3100 | https://vpn.echo6.co/admin | Headscale web UI (OIDC via [[authentik]]) — fronted by edge2 host [[caddy]]**migrated from Contabo [[2026-06-19]]** |
| Mailcow | **edge1 CT 101** (10.10.10.2) | 5.189.158.149 | https://mail.echo6.co | Email server (privileged LXC on rebuilt Contabo VPS, updated commit 52a41b4d / SOGo 5.12.8) — **rebuilt in-place [[2026-06-19]]** |
| Vaultwarden | edge2 (CT 102) | 100.64.0.33:8086 | https://vault.echo6.co | Password manager 1.37.1 (SSO enabled) — fronted by edge2 host Caddy (reverse_proxy 100.64.0.33:8086) |
| Grav | edge2 (CT 101) | 10.10.10.11:80 | https://idahomesh.com (+www) | Flat-file CMS 2.0.11, no database — Admin2 plugin at /admin; Apache 2.4.67 + mod_php + PHP 8.4.21; migrated from WordPress 2026-07-17 (MariaDB purged from the container); hostname still `wordpress` (unchanged) — fronted by edge2 host Caddy via **internal bridge IP** (reverse_proxy 10.10.10.11:80), unlike other edge2 services which proxy over tailnet |
| Syncthing | cortex | 100.64.0.14:22000 | Internal (Tailscale) | File sync — ~/.claude/, ~/projects/ (Syncthing on Contabo decommissioned 2026-06-19 with edge1 rebuild) |
@ -50,14 +50,14 @@ updated: 2026-08-05
| Radarr | media (VM 105) | 192.168.1.160:7878 | Internal | Movie automation (Docker) |
| Prowlarr | media (VM 105) | 192.168.1.160:9696 | Internal | Indexer manager (Docker) |
| SABnzbd | media (VM 105) | 192.168.1.160:8080 | Internal | [[usenet]] download client (Docker) |
| PeerTube | media (CT 110) | 192.168.1.170:9000 | https://stream.echo6.co | Video streaming (native, NFS on pi-nas, SSO) |
| PeerTube | media (CT 110) | 192.168.1.170:9000 | https://stream.echo6.co | Video streaming (native, NFS on pi-nas, SSO). CT raised to 8GB + redis capped 2026-08-13 — see [[peertube-sitemap-redis-oom]] |
| Open WebUI | cortex (VM 150) | 192.168.1.150:8080 | https://ai.echo6.co | AI chat interface (Docker, Ollama backend, SSO) |
| Qdrant | cortex (VM 150) | 192.168.1.150:6333 | Internal | Vector database (Docker, [[recon]] knowledge store) |
| TEI | cortex (VM 150) | 192.168.1.150:8090 | Internal | Text embeddings (Docker, bge-m3 1024-dim) |
| [[recon]] | data (VM 1130) | 192.168.1.130:8420 | https://recon.echo6.co | Knowledge extraction pipeline (systemd, dashboard+API) |
| navi-config | data (VM 1130) | 192.168.1.130:8422 | Internal | [[recon]] [[navi]] node config API |
| navi-contacts | data (VM 1130) | 192.168.1.130:8423 | Internal | [[recon]] [[navi]] contact enrichment API |
| navi-landclass | data (VM 1130) | 192.168.1.130:8424 | Internal | RECON navi land classification API |
| navi-landclass | data (VM 1130) | 192.168.1.130:8424 | Internal | [[recon]] [[navi]] land classification API |
| navi-places | data (VM 1130) | 192.168.1.130:8425 | Internal | RECON navi OSM place detail/enrichment |
| navi-geo | data (VM 1130) | 192.168.1.130:8426 | Internal | RECON navi geocode/reverse geocode API |
| navi-admin | data (VM 1130) | 192.168.1.130:8427 | Internal | RECON navi fleet admin-info aggregator |
@ -81,7 +81,7 @@ updated: 2026-08-05
| pt-transcoder | cortex (VM 150) | N/A | Internal | PeerTube H.265 NVENC transcoder (systemd, /opt/bulk-import/transcoder.py) |
| recon-sparse | cortex (VM 150) | 192.168.1.150:8091 | Internal | RECON sparse embedding service (systemd, bge-m3 model, port 8091) |
| obsidian-remote | cortex (VM 150) | 100.64.0.14:8082 → :3001 | Internal (Tailscale) | Headless web Obsidian (lscr.io/linuxserver/obsidian:latest, Docker) |
| mcc | media (CT 111) | 192.168.1.111:80/443 | Internal | Caddy + Postfix, pymc console web app; reverse-proxies /api,/auth,/ws → 192.168.1.253:8000 (aida-nebra) |
| mcc | media (CT 111) | 192.168.1.111:80 | Internal | Caddy + Postfix, OpenHop (pymc) console web app; serves static frontend, reverse-proxies /api,/auth,/ws → 192.168.1.253:8000 (aida-nebra). Port 80 only — no 443 listener |
| Samba | cortex (VM 150) | 192.168.1.150:445 | Internal | SMB file sharing — `//cortex/projects` → /home/zvx/projects (guest access) |
| Home Assistant | ha (cloud VM 151) | 192.168.1.151:8123 / 100.64.0.16 | Internal | Home automation platform (Docker, Ubuntu 24.04) |
@ -138,7 +138,7 @@ updated: 2026-08-05
- Homepage: centered Echo6 logo + pill search bar (Google-style, viewport-locked no-scroll)
- Results page: full-width two-column grid (results + sidebar), stretched search header
- Top nav bar: `.//photos`, `.//mail`, waffle app launcher (11 services), login avatar
- All nav links use Authentik launch URLs for seamless SSO pass-through
- All nav links use [[authentik]] launch URLs for seamless SSO pass-through
- search.echo6.co permanently redirects to echo6.co (301)
- Redis/Valkey cache (valkey container)
- Compose path: `/opt/searxng/docker-compose.yml`
@ -148,7 +148,7 @@ updated: 2026-08-05
- `img/echo6-logo.png` — Echo6 logo (replaces [[searxng]] logo)
- `img/favicon.png` — Echo6 favicon
- Config: `/opt/searxng/searxng-config/settings.yml` (instance_name: "Echo6", dark theme, center_alignment: false)
- SearXNG version: 2026.2.6 (Docker image: searxng/searxng:latest)
- [[searxng]] version: 2026.2.6 (Docker image: searxng/searxng:latest)
### utility - CT 104 (192.168.1.104 / Tailscale: 100.64.0.12)
- [[central]] data-hub spine (3 systemd units: central-supervisor, central-archive, central-gui)
@ -160,7 +160,7 @@ updated: 2026-08-05
### utility - CT 108 (192.168.1.144 / Tailscale: 100.64.0.32)
- [[meshai]] — LLM-powered Meshtastic mesh assistant (Docker)
- Bot name: AIDA, node ID !27780c47, channel 8 whitelist
- Image: work-meshai (local build, not ghcr.io/zvx-echo6/meshai:latest)
- Image: work-meshai (local build, not ghcr.io/zvx-echo6/[[meshai]]:latest)
- Backend: gemini-3.1-flash-lite with Google Search grounding
- Connects to meshtasticd **on aida-nebra** (192.168.1.253:4403) — the AIDA-N2 node !27780c47
- Exposes port 8080 (web UI)

View file

@ -0,0 +1,112 @@
---
title: Fleet Storage and Memory Upgrade
type: project
tags:
- storage
aliases: []
related:
- [[environment]]
- [[fleet-platform-baseline]]
- [[toc-cortex-pve9.2-update]]
- [[fleet-patch-audit]]
- [[pve-guest-park-and-adopt]]
updated: 2026-08-14
---
# Fleet Storage and Memory Upgrade
Placing a batch of acquired drives and memory across the fleet, and the storage-architecture decisions that came out of sizing it. All figures read live from the hardware 2026-08-13/14, not from spec sheets. Node reference is [[environment]].
---
## Parts on hand
| Part | Qty | Form factor | Where it fits | Gain |
|---|---|---|---|---|
| 32 GB DDR5 SODIMM | 2 | SODIMM | **media only** | 32 → 64 GB, its board maximum |
| 2 TB NVMe | 1 | M.2 2280 | media's freed M.2 | +2 TB fast local |
| 1 TB NVMe | 2 | M.2 2280 | utility, cloud, or toc | replaces a 512 GB, needs migration |
| 16 GB DDR4 SODIMM | 4 | SODIMM | **nowhere** | cold spares only |
The DDR4 has no home. data, utility and cloud each already run 2 × 16 GB DDR4-3200; data and utility are hard-capped at 32 GB, and cloud would need 32 GB sticks to gain anything. toc's two empty slots take **full-size DIMMs**, not SODIMM. Keep them as field spares for the three ThinkCentres.
Freed by the media work: one 512 GB Intel NVMe (0% wear, 9,791 hrs) and two 16 GB DDR5-4800 ADATA SODIMMs.
---
## Memory ceilings
| Node | Installed | Slots | Board max | Headroom |
|---|---|---|---|---|
| media | 32 GB DDR5-4800 | 2/2 | 64 GB | swap to 2 × 32 GB |
| cloud | 32 GB DDR4-3200 | 2/2 | 64 GB | needs 32 GB sticks |
| data | 32 GB DDR4-3200 | 2/2 | 32 GB | **at ceiling** |
| utility | 32 GB DDR4-3200 | 2/2 | 32 GB | **at ceiling** |
| toc | 64 GB DDR4 | 6/8 | 256 GB | 2 slots free, full-size DIMM |
| pi-nas | 8 GB soldered | — | fixed | none |
| aida-nebra | 906 MiB soldered | — | fixed | none |
**media runs DDR5-4800, not the 5600 previously recorded.**
**toc's memory runs at 2133 MT/s.** Installed modules are rated 3200, 2666 and 2133; DDR4 clocks the whole bus to the slowest module present. Matched sticks would buy speed as well as capacity. DIMM6 is an ECC module but ECC is not active on this board.
---
## Storage positions
| Node | Installed | Free positions |
|---|---|---|
| data | 1 TB NVMe + 1 TB SATA SSD | **none** |
| utility | 512 GB NVMe — 12% wear, 85.2 TB written | 1 × 2.5" SATA bay |
| cloud | 512 GB NVMe — 3% wear, 16.8 TB | 1 × 2.5" SATA bay |
| media | 2 × 512 GB NVMe (one = dead Windows) | 1 M.2 once pulled |
| toc | 512 GB NVMe — 26% wear, 85.6 TB | 8 SATA ports, 4 PCIe slots |
| pi-nas | 2 × 3 TB btrfs RAID1 + 2 × 24 TB single ext4 | 1 SATA port (`ata5`) |
**data has no second M.2.** Verified from PCIe topology, not spec sheets: only three external root ports exist and all are populated — NVMe, onboard NIC, Wi-Fi. The Wi-Fi slot is E-keyed and cannot take an M-keyed 2280 drive. Adding capacity to data means replacing a drive.
**media's second M.2 holds a BitLocker Windows install dated 2022-09-26**, untouched by Proxmox, 0% wear. It is the drive to pull.
Two identical Intel SSDPEKNU512GZH in media — go by serial:
- **Pull `PHKA142402U8512A`** (Windows)
- **Keep `PHKA142504HP512A`** (Proxmox boot and LVM)
Dell firmware reports a placeholder bus address for all three M.2 slots and Intel VMD remaps the PCI addresses, so the label serial is the only reliable identifier. Pulling the wrong one is non-destructive — the node just won't boot. Provisioning the wrong one is not.
**Highest-wear drives** are toc's (26%, 85.6 TB) and utility's (12%, 85.2 TB in only 3,459 hours — about 166 full drive-writes). Both are single drives with no redundancy.
---
## Sequence
1. **media, one shutdown, no migration.** Pull the Windows drive, fit the 2 TB, swap both DDR5 sticks for the 32s. Yields 64 GB and 2 TB of local NVMe. Takes down PeerTube, the *arr stack and mcc for the duration.
2. **utility.** 1 TB replaces the 512 GB SN740 — clone-and-swap, since its only M.2 is the boot drive. It is the fullest node (thin pool 45.4% vs cloud's 23.2%) and the fastest-wearing drive.
3. **toc.** Its lone M.2 holds the most-written drive in the fleet. With four free PCIe slots, an **M.2-to-PCIe adapter** lets the 1 TB be *added* rather than replacing the boot drive — no clone, no migration.
Open question: whether any of the spare 512 GB drives are 2.5" SATA rather than M.2. If so they drop straight into utility's and cloud's empty bays as pure additions with no migration at all.
---
## Ceph — evaluated, not adopted
Goal was migration freedom: move guests between nodes at will and rebuild nodes as needed.
Sizing, from 715 GB of guest data actually in use (~950 GB provisioned):
| Scope | Usable | `size=3` raw | With headroom | Per node × 5 |
|---|---|---|---|---|
| Guest disks only | 1 TB | 3 TB | ~4.7 TB | ~1 TB |
| Guests + nav/kiwix/library | 2 TB | 6 TB | ~9 TB | ~1.8 TB |
Capacity was never the blocker. **Free devices were.** Ceph wants a whole dedicated drive per OSD, and data has none free, while utility and cloud have only one M.2 each holding their boot drive — so the spare M.2 NVMe cannot become OSDs there. It also costs ~4 GB RAM per OSD (`osd_memory_target`) on two nodes hard-capped at 32 GB.
One unlock exists: if [[navi]] moves off data, the 1 TB SATA SSD empties and becomes data's OSD. See [[navi-recon-separation]].
Adopted instead: **park-and-adopt via vzdump** to NAS-backed storage, which achieves the migration-and-rebuild goal with hardware already owned. See [[pve-guest-park-and-adopt]]. Note that LXC cannot live-migrate in PVE 9 under any storage arrangement, so shared storage would have bought fast restart-migration, not zero downtime.
Running guests directly from NFS was also rejected: container volumes are raw, NFS snapshots need qcow2, so `pct snapshot` / `pct rollback` — the rollback path [[fleet-patch-audit]] depends on — would stop working, and pi-nas would become a single point of failure for every guest at once.
---
## Redundancy note
22 TB of bulk data on pi-nas sits on **single ext4 drives with no redundancy**, including PeerTube's 11 TB library. Only the small 3 TB pair is mirrored, and that pair has **73,119 power-on hours** — 8.3 years — while holding Immich and Nextcloud. SMART is clean with zero reallocated sectors on all four drives. One SATA port remains free on the controller. Build detail is [[pi-nas-omv-runbook]].

View file

@ -0,0 +1,96 @@
---
title: Separating navi from recon
type: project
tags:
- recon
aliases: []
related:
- [[navi]]
- [[deployment]]
- [[cc-rules]]
- [[fleet-storage-memory-upgrade]]
- [[recon-operations]]
updated: 2026-08-14
---
# Separating navi from recon
[[recon]] and [[navi]] are separate functions but they currently live in **one VM**, so neither can be relocated without the other. This documents what actually belongs to which, established 2026-08-13/14, and what a split would cost.
Platform docs: [[recon]], [[navi]].
---
## Current state
`recon-vm` (VMID 1130) on data — 4 cores, 24 GB allocated, 180 GB disk on local-lvm. It runs roughly twenty [[services]] across both platforms:
- **recon**`recon.service`, `recon-watchdog.service`
- **navi**`navi-admin`, `navi-config`, `navi-contacts`, `navi-geo`, `navi-landclass`, `navi-offroute`, `navi-places` (gunicorn, all bound to `127.0.0.1:84xx`)
- **geo stack**`photon.service`, `argus-resolver.service`, PostgreSQL 16, plus Docker running Nominatim v5 and Valhalla
- **content**`kiwix.service`
- `apache2` fronts the localhost-bound services
Calling a relocation of this VM "moving recon" badly undersells it.
---
## Ownership, measured
Three virtiofs shares are mounted from data's `/mnt/data` (the 1 TB SATA SSD at 93% full):
| Share | Size | Belongs to | Evidence |
|---|---|---|---|
| `nav` | 625 GB | **navi** | held open by `postgres`, `valhalla_`, `java` (Photon) |
| `kiwix` | 138 GB | kiwix-serve, also read by recon | `kiwix-ser` holds it; recon references `/mnt/kiwix/library` |
| `library` | 59 GB | **recon** | `/opt/recon` code references `/mnt/library/Acquired`, `_ingest`, `_acquired` |
`/mnt/nav` breakdown: `overture` 252 GB (contains PostgreSQL's `data_directory` at `/mnt/nav/overture/pgdata`), `tiles` 128 GB (Valhalla), `worldcover` 85 GB, `sources` 67 GB, `addresses` 36 GB, `nominatim-v5` 29 GB, `photon` 21 GB, `padus` 3.8 GB.
**recon's own footprint is 25 GB** — `/opt/recon`, of which `data` is 24 GB. It lives on the VM's disk, **not** on the full volume. Its [[environment]] holds only a profile setting, four Gemini keys and a PeerTube password. **recon uses no database at all** — the only `psycopg2` hits are files inside its venv. PostgreSQL belongs entirely to navi.
---
## Which one should move
**navi**, if either does.
- **Storage** — navi is 625 GB of the 821 GB on data's full SATA SSD. Moving recon frees *zero bytes* from the volume that is actually full. Moving navi takes it from 93% to roughly 21%.
- **Memory** — the 18 GB of page cache in that VM serves Nominatim, Valhalla, Photon and PostgreSQL. recon is a modest Python service. The VM uses 4.8 GB of 23 GB with the rest as working cache.
- **Speed** — navi's data sits on SATA SSD today. NVMe is materially faster for the random reads a geocoder and routing engine live on.
Counterpoint worth noting: recon holds a PeerTube password and feeds the PeerTube instance on media, so recon has a mild affinity *toward* media, not away from it.
---
## Why this is a rebuild, not a migration
**virtiofs is machine-local.** The shares are wired through raw QEMU `args:` to unix sockets on the data host, backed by three systemd units (`virtiofsd-nav/kiwix/library.service`) each running `/usr/libexec/virtiofsd --socket-path=/run/virtiofsd-<n>.sock --shared-dir=/mnt/data/<n> --cache=auto --announce-submounts`. It cannot cross hosts, and it blocks live migration outright. `vzdump` does not capture the contents either — see [[pve-guest-park-and-adopt]].
The VM config references socket *paths*, not host-specific IDs, so the same units on another host with `--shared-dir` repointed would work unchanged. Moving the whole VM is therefore tractable. Splitting the two platforms apart is not — it means standing up PostgreSQL with the overture and padus databases, the Nominatim and Valhalla containers, Photon, and the seven `navi-*` services on a new host, then removing them from the original.
Unresolved before planning: the `navi-*` services bind to `127.0.0.1` behind `apache2`, so the external entry point moves with navi. Whether recon calls navi's APIs has not been confirmed.
**Memory accounting caution:** the `virtiofsd` process for `nav` shows ~16.7 GB RSS, but `RssAnon` is only 20 MB — it is almost entirely `RssShmem`, the guest's own RAM mapped through the `memory-backend-memfd,share=on` object that virtiofs requires. Counting it on top of the VM's 24 GB is double-counting.
---
## Viewshed — planned, not built
No viewshed or line-of-sight code exists anywhere on the VM. Sizing for the planned feature, since it drives the memory budget:
Elevation data present is **SRTM 1 arc-second (~30 m)** — 48 `.hgt` tiles of 3601×3601 int16, 1.2 GB total, covering **4150°N, 110118°W** (roughly 620 × 400 miles, ~163,000 sq mi). Moving to 10 m 3DEP over the same footprint costs 722 GB depending on format; contour generation is where storage actually goes (1060 GB of vector tiles, plus 100150 GB of transient scratch during `gdal_contour`/`tippecanoe`, which is why cortex carries a 32 GB swapfile). 3DEP 10 m is US-only.
Memory depends entirely on algorithm:
| Approach | 250 mi radius at 10 m |
|---|---|
| Ray-cast, tile cache | ~8 GB (32-tile LRU) — algorithm itself under 200 MB |
| Full-grid raster | **~84 GB** across DEM + clutter + RSSI + working layers |
A 250-mile disc is 80,400² ≈ 6.5 billion cells; full-grid is quadratic in radius, ray-casting is linear. Tools in this space — including `xarray-spatial`, which mesh_terrain uses — implement full-grid, which is why they pull in `dask` for out-of-core work.
**Budget 32 GB** if building this. That covers full-grid at 100150 miles at 10 m, 250 miles at 30 m, or dask-chunked 250 miles at 10 m. Uncompromised 250 miles at 10 m in memory is not reachable on current hardware.
Physical note: with 4/3-Earth refraction, `d(km) = 4.12(√h₁ + √h₂)`. A 250-mile path needs ~2,380 m at both ends — summit-to-summit only. Realistic siting at 1,500 m gives about 200 miles station-to-station.
This budget does not fit on data, which is capped at 32 GB and already runs the whole geo stack. It fits comfortably on media at 64 GB. See [[fleet-storage-memory-upgrade]].

View file

@ -0,0 +1,147 @@
---
title: Add an NFS Share on pi-nas (OpenMediaVault RPC)
type: runbook
tags:
- storage
aliases: []
related:
- [[pi-nas-omv-runbook]]
- [[pve-guest-park-and-adopt]]
- [[proxmox-onboard-node]]
- [[proxmox-create-ubuntu-vm]]
- [[ct-runbook]]
updated: 2026-08-14
---
# Add an NFS Share on pi-nas (OpenMediaVault RPC)
pi-nas runs OpenMediaVault 8.4.0 (Synchrony). **Never hand-edit `/etc/exports` or `/etc/fstab` on this box.** Both are auto-generated from OMV's config database and carry the warning header; hand edits are silently reverted the next time any OMV config is applied, which will happen eventually whether or not you trigger it.
Create shares through OMV's RPC instead. The web UI does the same thing, but RPC is scriptable and works over SSH.
Initial NAS build is [[pi-nas-omv-runbook]].
---
## Key facts
The **new-object sentinel UUID** is how OMV signals "create this, don't update":
```
fa4b1c66-ef79-11e5-87a0-0002b3a176b4
```
Resolve it yourself if ever in doubt:
```bash
sudo php -r 'require_once("/usr/share/php/openmediavault/autoloader.inc");
echo \OMV\Environment::get("OMV_CONFIGOBJECT_NEW_UUID"),"\n";'
```
**Filesystem references** (`mntentref`) as of 2026-08-14:
| Disk | Size / role | mntentref |
|---|---|---|
| `sda`+`sdb` | 2.7 TB btrfs RAID1 — the only mirrored pool | `423365c2-513c-4900-bbde-a2e7b21e6bd5` |
| `sdc1` | 24 TB ext4, single | `a3c6d71c-527d-45a3-aab7-1868e858456f` |
| `sdd1` | 24 TB ext4, single — carries PeerTube | `79238c6b-1064-4ac9-9564-b12fc8608b74` |
Re-read them at any time:
```bash
sudo omv-rpc -u admin "ShareMgmt" "getList" '{"start":0,"limit":-1}'
```
---
## 1. Create the shared folder
`reldirpath` is relative to the chosen filesystem's mount point and must end in `/`. It may not contain `..`.
```bash
sudo omv-rpc -u admin "ShareMgmt" "set" '{
"uuid":"fa4b1c66-ef79-11e5-87a0-0002b3a176b4",
"name":"pvebackup",
"reldirpath":"pvebackup/",
"comment":"Proxmox vzdump archives",
"mntentref":"a3c6d71c-527d-45a3-aab7-1868e858456f",
"mode":"775"
}'
```
It returns the created object. **Keep the `uuid`** — the NFS share references it.
## 2. Create the NFS share, once per client network
Pass the sentinel for `mntentref` as well; OMV overwrites it with the bind-mount entry it creates for `/export/<name>`.
`no_root_squash` is required for anything Proxmox writes as root, including `vzdump`. Omit it for shares that do not need it.
```bash
SF=<uuid from step 1>
NEW=fa4b1c66-ef79-11e5-87a0-0002b3a176b4
# LAN
sudo omv-rpc -u admin "NFS" "setShare" "{
\"uuid\":\"$NEW\",\"sharedfolderref\":\"$SF\",\"mntentref\":\"$NEW\",
\"client\":\"192.168.1.0/24\",\"options\":\"rw\",
\"extraoptions\":\"subtree_check,insecure,no_root_squash\",
\"comment\":\"pvebackup - local\"}"
# tailnet
sudo omv-rpc -u admin "NFS" "setShare" "{
\"uuid\":\"$NEW\",\"sharedfolderref\":\"$SF\",\"mntentref\":\"$NEW\",
\"client\":\"100.64.0.0/10\",\"options\":\"rw\",
\"extraoptions\":\"subtree_check,insecure,no_root_squash\",
\"comment\":\"pvebackup - tailnet\"}"
```
`extraoptions` is pattern-validated: comma-separated words only, no spaces.
## 3. Apply
Scope it to the modules actually affected rather than applying everything:
```bash
sudo omv-rpc -u admin "Config" "applyChanges" '{"modules":["fstab","nfs"],"force":false}'
```
This regenerates `/etc/exports`, creates the bind mount, and re-runs `exportfs`.
## 4. Verify — including that you broke nothing
`exportfs` re-runs against live clients. Existing mounts survived this in practice, but confirm rather than assume:
```bash
# on pi-nas
sudo exportfs -v | grep -A1 pvebackup
df -h /export/pvebackup
sudo exportfs -s | grep -oE '^/export/[a-z]+' | sort -u
```
Then check every existing consumer still reads. As of 2026-08-14 those are media (`/mnt/peertube-storage`) and cloud (`/mnt/nfs-immich`, `/mnt/nfs-nextcloud`) — see [[services]].
---
## Mounting it in Proxmox
Storage config is cluster-wide, so run this once on any node and all five pick it up:
```bash
pvesm add nfs pinas-backup --server 192.168.1.245 \
--export /export/pvebackup --content backup --options vers=3
```
Confirm root can actually write, which is what `no_root_squash` buys:
```bash
dd if=/dev/zero of=/mnt/pve/pinas-backup/.writetest bs=1M count=32 && \
rm /mnt/pve/pinas-backup/.writetest && echo OK
```
Content types matter. `backup` is for vzdump archives; `images,rootdir` would let guests run directly from NFS, which was deliberately not done — see [[pve-guest-park-and-adopt]] for why.
---
## Redundancy warning
Only `sda`+`sdb` are mirrored (btrfs RAID1, 2.7 TB, ~2.1 TB free). `sdc` and `sdd` are single ext4 drives with no redundancy despite holding the bulk of the data, including PeerTube's library. Put anything that must survive a drive failure on the mirrored pool.

View file

@ -0,0 +1,143 @@
---
title: PeerTube Sitemap Redis OOM — Diagnosis and Fix
type: runbook
tags:
- media
aliases: []
related:
- [[add-peertube-channel]]
- [[caddy]]
- [[peertube-remote-runner]]
- [[central]]
- [[recon-operations]]
updated: 2026-08-14
---
# PeerTube Sitemap Redis OOM — Diagnosis and Fix
stream.echo6.co goes down and media's load average climbs past 100 while the CPU sits mostly idle. Root cause found and fixed 2026-08-13.
[[deployment]] layout is [[services]]; PeerTube lives in CT 110 on media.
---
## Signature
Recognise it by this combination — the idle CPU is the tell:
- media load average **>100** with CPU **~85% idle** and high iowait
- `pct exec 110` hangs and returns nothing
- Every process in the container — nginx, postgres, redis, sshd — stuck in `D` state
- Other home [[services]] (echo6.co, jellyfin, immich) respond normally in under 100 ms
- [[caddy]] on utility times out for stream.echo6.co only
It looks like a dead host. It is one wedged container on a host with plenty of free memory.
```bash
ssh ts-media 'uptime; ps -eo stat --no-headers | grep -c "^D"'
cat /sys/fs/cgroup/lxc/110/memory.current /sys/fs/cgroup/lxc/110/memory.max
grep -E '^oom' /sys/fs/cgroup/lxc/110/memory.events
cat /sys/fs/cgroup/lxc/110/io.pressure
```
At the time of the incident: memory 4.269 GB of a 4.294 GB limit, swap 100% full, **268 cgroup OOM kills**, `io.pressure` pinned at 99% — while the host itself had 13 GB free.
---
## Cause
`/sitemap.xml` is **108 MB** and takes ~45 s to generate, because the instance mirrors ~136K videos. PeerTube caches it in redis under:
```
redis-stream.echo6.co-api-cache-<epoch_ms>-/sitemap.xml
```
Each cached copy costs **402 MB** — roughly 4× the wire size, from Node string overhead — and carries a multi-hour TTL. Every cache miss mints another copy. Redis runs `maxmemory 0` with `noeviction`, so nothing ever evicts them. Generating the sitemap also spikes the PeerTube Node process to ~2.9 GB RSS on its own.
In a 4 GB container that is fatal. Memory and swap fill, the kernel OOM-kills redis every few hours, and every process ends up in uninterruptible sleep on major page faults.
Three cached copies accounted for 1.13 GB of redis's 1.14 GB. The genuine working set — bull job queues — is about **11 MB**.
---
## Fix
All three parts are reboot-safe and already applied.
### Container memory 4 GB → 8 GB
```bash
ssh ts-media
pct stop 110 # stops cleanly despite the D-state pile
pct set 110 -memory 8192
pct start 110
```
Check headroom first — media had 13.3 GB allocated of 31 GB.
### Cap redis
`volatile-lru`, **not** `allkeys-lru`. Bull job-queue keys are mostly TTL-less and must never be evicted; only the API cache entries carry TTLs.
```bash
R=$(grep -oP '(?<=auth: ")[^"]+' /var/www/peertube/config/production.yaml)
pct exec 110 -- redis-cli -a "$R" --no-auth-warning config set maxmemory 1gb
pct exec 110 -- redis-cli -a "$R" --no-auth-warning config set maxmemory-policy volatile-lru
pct exec 110 -- redis-cli -a "$R" --no-auth-warning config rewrite
```
`config rewrite` persists to `/etc/redis/redis.conf`.
### Stop serving the sitemap
The sitemap TTL is compiled into PeerTube 8.0.2 and is not settable in `production.yaml`, so the durable control is nginx. Add above `location / {` in `sites-available/peertube`:
```nginx
location = /sitemap.xml {
return 404;
}
```
Then `nginx -t && systemctl reload nginx`. Result: 404 in 43 ms instead of 108 MB over 45 s. This stops both the redis blob and the Node heap spike at source.
### Reclaim existing blobs
```bash
pct exec 110 -- bash -c "redis-cli -a $R --no-auth-warning --scan \
--pattern '*api-cache*sitemap.xml' | while read k; do \
redis-cli -a $R --no-auth-warning del \"\$k\"; done"
```
Dropped redis from 1.14 GB to 11 MB immediately.
---
## Editing nginx in this container — read this first
`sites-enabled/peertube` is a **symlink** to `sites-available/peertube`. Running `sed -i.bak` against it replaces the symlink with a regular file *and* leaves the `.bak` symlink inside `sites-enabled/`, so nginx loads the server block twice and warns:
```
conflicting server name "stream.echo6.co" on 0.0.0.0:80, ignored
```
Edit `sites-available/` directly and never leave backups inside `sites-enabled/`. A pristine pre-change copy is at `/root/peertube-nginx-orig-20260813.bak` inside CT 110.
---
## Verify
```bash
curl -o /dev/null -w '%{http_code} %{size_download}B %{time_total}s\n' https://stream.echo6.co/sitemap.xml # expect 404
curl -o /dev/null -w '%{http_code} %{time_total}s\n' https://stream.echo6.co/ # expect 200
```
Then confirm the page actually renders per [[headless-browser-page-verification]] — a 200 from nginx does not prove PeerTube is serving.
Post-fix: load 136 → 2.6, io.pressure 99% → 5%, swap 0, D-state 0, OOM kills 0.
---
## Still open
8 GB raises the ceiling; it does not stop growth. At ~136K videos and climbing, revisit if the library grows substantially. Fetching `/sitemap.xml` to measure it mints a fresh 402 MB redis entry, so do not casually curl it.
Related pipeline failure modes: [[peertube-remote-runner]], [[add-peertube-channel]].

View file

@ -0,0 +1,134 @@
---
title: Park a Proxmox Guest on the NAS and Re-Adopt It
type: runbook
tags:
- proxmox
- storage
aliases: []
related:
- [[omv-add-nfs-share]]
- [[environment]]
- [[proxmox-onboard-node]]
- [[proxmox-create-ubuntu-vm]]
- [[pi-nas-omv-runbook]]
updated: 2026-08-14
---
# Park a Proxmox Guest on the NAS and Re-Adopt It
Shut a guest down, archive the whole thing to pi-nas, rebuild or replace the node, then bring the guest back on **any** node in the cluster. This is the Proxmox equivalent of registering an orphaned VM from an ESXi datastore, and it is the supported way to free a node for rebuild without shared runtime storage.
The archive is self-contained and carries no reference to its origin host, so it can be restored to a different node, a different VMID, different storage, or a freshly reinstalled machine.
---
## Prerequisites
The `pinas-backup` storage must exist. It was added 2026-08-14 and is visible on all five cluster nodes:
| Setting | Value |
|---|---|
| Type | NFS, `vers=3` |
| Server | `192.168.1.245` (pi-nas) |
| Export | `/export/pvebackup` |
| Backing disk | `sdc1` — the 19 TB-free drive, deliberately not `sdd1` (PeerTube) |
| Content | `backup` |
| Mount point | `/mnt/pve/pinas-backup` |
Confirm before starting:
```bash
pvesm status | grep pinas-backup
```
Creating or recreating this share is covered in [[omv-add-nfs-share]]. General NAS build is [[pi-nas-omv-runbook]].
---
## 1. Park the guest
`--mode stop` shuts the guest down cleanly, archives it, and leaves it stopped.
```bash
# container
vzdump 111 --mode stop --storage pinas-backup --compress zstd
# VM
vzdump 105 --mode stop --storage pinas-backup --compress zstd
```
Note the archive name it prints. Files land in `/mnt/pve/pinas-backup/dump/`:
```
vzdump-lxc-111-2026_08_14-10_22_31.tar.zst # container
vzdump-qemu-105-2026_08_14-10_45_02.vma.zst # VM
```
Compression is significant — a 50 GB container typically lands near 1.5 GB.
List what is parked:
```bash
pvesm list pinas-backup
```
## 2. Remove the guest from the source node
Only once the archive is verified present and non-zero. This is the destructive step.
```bash
pct destroy 111 # container
qm destroy 105 # VM
```
Skip this entirely if the plan is to wipe the node anyway.
## 3. Rebuild the node
Node rebuild is [[proxmox-onboard-node]]. The archive is untouched by anything done to the node.
## 4. Adopt it back
Run this **on the node you want the guest to live on**. The VMID and target storage are free choices — they do not have to match the original.
```bash
# container
pct restore 111 pinas-backup:backup/vzdump-lxc-111-2026_08_14-10_22_31.tar.zst \
--storage local-lvm
# VM
qmrestore pinas-backup:backup/vzdump-qemu-105-2026_08_14-10_45_02.vma.zst 105 \
--storage local-lvm
```
To clone rather than move, restore under a new VMID and leave the original in place.
## 5. Start and verify
```bash
pct start 111 && pct status 111
qm start 105 && qm status 105
```
Verify the service itself, not just that the guest is running. For anything with a web front end, follow [[headless-browser-page-verification]].
---
## Things that will bite you
**Containers cannot live-migrate in PVE 9 at all.** This is a hard platform limitation, not a configuration gap. Restart migration is the only option for LXC. VMs live-migrate normally when storage is shared.
**Bind mounts and device mount points are not captured.** `vzdump` skips their contents. Anything reached through a bind mount, a `lxc.mount.entry`, or virtiofs must be moved separately. This applies directly to recon-vm, whose `nav`, `kiwix` and `library` shares are virtiofs from the host — see [[navi-recon-separation]].
**Passthrough does not survive relocation.** USB devices, PCI passthrough and custom `args:` reference host-specific paths. Guests using them need those recreated on the target node before they will start.
**`sdc` has no redundancy.** It is a single ext4 drive. An archive parked there dies with the drive — acceptable for a staging area you are actively pulling back from, not acceptable as the only copy of something. Only `sda`/`sdb` on pi-nas are mirrored.
**Space is shared.** `/export/pvebackup` and `/export/arr` sit on the same physical drive, so the ~19 TB free is not exclusively yours.
---
## Why not shared runtime storage
Running guests directly off NFS was considered and rejected: container volumes are raw images, NFS snapshots require qcow2, so `pct snapshot` / `pct rollback` stops working — and that rollback path is what [[fleet-patch-audit]] depends on for safe patching. It would also make pi-nas a single point of failure for every guest on every node simultaneously.
Ceph was considered for the same goal. It needs a dedicated whole drive per node, and three of five nodes have no free drive bay; it also costs ~4 GB RAM per OSD on two nodes that are hard-capped at 32 GB. See [[fleet-storage-memory-upgrade]].