echo6-docs/vault/projects/fleet-storage-memory-upgrade.md

146 lines
9.1 KiB
Markdown
Raw Normal View History

---
title: Fleet Storage and Memory Upgrade
type: project
tags:
- storage
aliases: []
related:
- [[environment]]
- [[navi-lift-to-media]]
- [[fleet-platform-baseline]]
- [[navi-recon-separation]]
- [[toc-cortex-pve9.2-update]]
updated: 2026-08-16
---
# Fleet Storage and Memory Upgrade
Placing a batch of acquired drives and memory across the fleet, and the storage-architecture decisions that came out of sizing it. All figures read live from the hardware 2026-08-13/14, not from spec sheets. Node reference is [[environment]].
---
## Parts on hand
Status as of 2026-08-15.
| Part | Qty | Form factor | Disposition |
|---|---|---|---|
| 32 GB DDR5 SODIMM | 2 | SODIMM | **INSTALLED** in media — 32 → 64 GB, board maximum reached |
| 2 TB NVMe (WD Green SN350) | 1 | M.2 2280 | **INSTALLED** in media as VG `tank``/mnt/nvme2tb` |
| 1 TB NVMe | 1 | M.2 2280 | **used elsewhere** — moved into a laptop |
| 1 TB NVMe | 1 | M.2 2280 | **spare, earmarked for utility** (no date set) |
| 16 GB DDR4 SODIMM | 2 | SODIMM | **used elsewhere** — moved into the same laptop |
| 16 GB DDR4 SODIMM | 2 | SODIMM | spare |
| 512 GB Intel NVMe | 1 | M.2 2280 | spare — pulled from media, 0% wear, 9,791 hrs, holds a 2022 BitLocker Windows install |
| 16 GB DDR5-4800 ADATA SODIMM | 2 | SODIMM | spare — displaced from media |
**Why utility is the target for the remaining 1 TB:** it has the fullest thin pool in the fleet (45% vs cloud's 23%) and the fastest-wearing drive — a 512 GB WD SN740 at **12% life used with 85 TB written in only 3,459 hours**, roughly 166 full drive-writes. Its single M.2 holds the boot drive, so this is a clone-and-swap, not an add. Its empty 2.5" SATA bay cannot take an M.2 drive.
The DDR4 SODIMMs had no home in the fleet — data, utility and cloud all already run 2 × 16 GB and the first two are hard-capped at 32 GB, while toc's empty slots need full-size DIMMs. Two went to the laptop instead.
---
## Memory ceilings
| Node | Installed | Slots | Board max | Headroom |
|---|---|---|---|---|
| media | 32 GB DDR5-4800 | 2/2 | 64 GB | swap to 2 × 32 GB |
| cloud | 32 GB DDR4-3200 | 2/2 | 64 GB | needs 32 GB sticks |
| data | 32 GB DDR4-3200 | 2/2 | 32 GB | **at ceiling** |
| utility | 32 GB DDR4-3200 | 2/2 | 32 GB | **at ceiling** |
| toc | 64 GB DDR4 | 6/8 | 256 GB | 2 slots free, full-size DIMM |
| pi-nas | 8 GB soldered | — | fixed | none |
| aida-nebra | 906 MiB soldered | — | fixed | none |
**media runs DDR5-4800, not the 5600 previously recorded.**
**toc's memory runs at 2133 MT/s.** Installed modules are rated 3200, 2666 and 2133; DDR4 clocks the whole bus to the slowest module present. Matched sticks would buy speed as well as capacity. DIMM6 is an ECC module but ECC is not active on this board.
---
## Storage positions
| Node | Installed | Free positions |
|---|---|---|
| data | 1 TB NVMe + 1 TB SATA SSD | **none** |
| utility | 512 GB NVMe — 12% wear, 85.2 TB written | 1 × 2.5" SATA bay |
| cloud | 512 GB NVMe — 3% wear, 16.8 TB | 1 × 2.5" SATA bay |
| media | 2 × 512 GB NVMe (one = dead Windows) | 1 M.2 once pulled |
| toc | 512 GB NVMe — 26% wear, 85.6 TB | 8 SATA ports, 4 PCIe slots |
| pi-nas | 2 × 3 TB btrfs RAID1 + 2 × 24 TB single ext4 | 1 SATA port (`ata5`) |
**data has no second M.2.** Verified from PCIe topology, not spec sheets: only three external root ports exist and all are populated — NVMe, onboard NIC, Wi-Fi. The Wi-Fi slot is E-keyed and cannot take an M-keyed 2280 drive. Adding capacity to data means replacing a drive.
**media's second M.2 holds a BitLocker Windows install dated 2022-09-26**, untouched by Proxmox, 0% wear. It is the drive to pull.
Two identical Intel SSDPEKNU512GZH in media — go by serial:
- **Pull `PHKA142402U8512A`** (Windows)
- **Keep `PHKA142504HP512A`** (Proxmox boot and LVM)
Dell firmware reports a placeholder bus address for all three M.2 slots and Intel VMD remaps the PCI addresses, so the label serial is the only reliable identifier. Pulling the wrong one is non-destructive — the node just won't boot. Provisioning the wrong one is not.
**Highest-wear drives** are toc's (26%, 85.6 TB) and utility's (12%, 85.2 TB in only 3,459 hours — about 166 full drive-writes). Both are single drives with no redundancy.
---
## Sequence
1. **media, one shutdown, no migration.** Pull the Windows drive, fit the 2 TB, swap both DDR5 sticks for the 32s. Yields 64 GB and 2 TB of local NVMe. Takes down PeerTube, the *arr stack and mcc for the duration.
2. **utility.** 1 TB replaces the 512 GB SN740 — clone-and-swap, since its only M.2 is the boot drive. It is the fullest node (thin pool 45.4% vs cloud's 23.2%) and the fastest-wearing drive.
3. **toc.** Its lone M.2 holds the most-written drive in the fleet. With four free PCIe slots, an **M.2-to-PCIe adapter** lets the 1 TB be *added* rather than replacing the boot drive — no clone, no migration.
Open question: whether any of the spare 512 GB drives are 2.5" SATA rather than M.2. If so they drop straight into utility's and cloud's empty bays as pure additions with no migration at all.
---
## Candidate node — Dell OptiPlex 3080 Micro (tag D9NPZB3)
Decoded from the Dell factory configuration 2026-08-14. A 2020-era 1-litre machine, same class as the ThinkCentres running utility and data.
| | As shipped |
|---|---|
| CPU | Core i5-10500T — 10th gen Comet Lake, 6C/12T, 12 MB cache, 2.3 → 3.8 GHz, **35 W** |
| RAM | 8 GB (2 x 4 GB) DDR4-2666 non-ECC SODIMM, 1Rx16 |
| Storage | 256 GB WD SN730 NVMe (M.2 2280) |
| Wireless | Intel Wi-Fi 6 AX200 2x2 + BT 5, internal antennas |
| Other | Discrete TPM enabled, 65 W adapter, Windows 10 Pro |
**Expansion:** 2 SODIMM slots, **64 GB** maximum, DDR4-2666 non-ECC. Both slots ship filled with 4 GB sticks, so any upgrade is a replacement, not an addition. One M.2 2230/2280 slot at **PCIe Gen3 x4** (up to 2 TB) — a Gen4 drive buys nothing here. One free 2.5" SATA bay, up to 2 TB.
With 32 GB and a 1 TB NVMe it becomes a legitimate Proxmox node, roughly on par with utility.
Three caveats before committing it:
- **Non-vPro** (`2RC2N : INFO,INTEL,N-VPRO,BASE`). No Intel AMT, so no out-of-band management — it lands in the same hole as the rest of the fleet, where a wedged node means a physical visit. Only in-band management is configured, and `817-BBSI` shows system monitoring was not selected either.
- **The 2.5" bay needs parts that did not ship.** The config carries `379-BBCY : No Additional Cable`, so the SATA cable and drive bracket are absent. Source them at purchase time; they are annoying to find later.
- **Single 1 GbE Realtek NIC** (`K7F14 : SRV,DRVR,REALTEK,LOM`). Not a concern in practice — utility already loads `r8169` and runs fine.
Dell service-tag lookups cannot be automated from here: their support site is behind Akamai bot protection and returns `Access Denied` to plain fetches, to their own JSON product-selector endpoints, and to a real headless Chromium with a full browser fingerprint. Their TechDirect API would work but needs an OAuth key we do not hold. Either read the tag in a browser, or skip Dell entirely and run `dmidecode` on the machine, which yields more than the support page does.
Service tags already read off the fleet: utility `MJ0LZNYT` and data `MZ010LPV` (both Lenovo 11JN002RUS), cloud `MJ0LQCGJ` (Lenovo 11T3000RUS), media `58SM6X3` (Dell OptiPlex Micro 7020).
---
## Ceph — evaluated, not adopted
Goal was migration freedom: move guests between nodes at will and rebuild nodes as needed.
Sizing, from 715 GB of guest data actually in use (~950 GB provisioned):
| Scope | Usable | `size=3` raw | With headroom | Per node × 5 |
|---|---|---|---|---|
| Guest disks only | 1 TB | 3 TB | ~4.7 TB | ~1 TB |
| Guests + nav/kiwix/library | 2 TB | 6 TB | ~9 TB | ~1.8 TB |
Capacity was never the blocker. **Free devices were.** Ceph wants a whole dedicated drive per OSD, and data has none free, while utility and cloud have only one M.2 each holding their boot drive — so the spare M.2 NVMe cannot become OSDs there. It also costs ~4 GB RAM per OSD (`osd_memory_target`) on two nodes hard-capped at 32 GB.
One unlock exists: if [[navi]] moves off data, the 1 TB SATA SSD empties and becomes data's OSD. See [[navi-recon-separation]].
Adopted instead: **park-and-adopt via vzdump** to NAS-backed storage, which achieves the migration-and-rebuild goal with hardware already owned. See [[pve-guest-park-and-adopt]]. Note that LXC cannot live-migrate in PVE 9 under any storage arrangement, so shared storage would have bought fast restart-migration, not zero downtime.
Running guests directly from NFS was also rejected: container volumes are raw, NFS snapshots need qcow2, so `pct snapshot` / `pct rollback` — the rollback path [[fleet-patch-audit]] depends on — would stop working, and pi-nas would become a single point of failure for every guest at once.
---
## Redundancy note
22 TB of bulk data on pi-nas sits on **single ext4 drives with no redundancy**, including PeerTube's 11 TB library. Only the small 3 TB pair is mirrored, and that pair has **73,119 power-on hours** — 8.3 years — while holding Immich and Nextcloud. SMART is clean with zero reallocated sectors on all four drives. One SATA port remains free on the controller. Build detail is [[pi-nas-omv-runbook]].