Check-in — Prox 24h memory soak start
TL;DR
Prox RAM maxed (60 GiB; swap nearly full). Deep attribution soak running on
infra-services. Owner judgments for Healthchecks keep, Komodo ntfy Alerter,
and Stage 7 pilot ntfy are done. Shrink/retire remains after soak window.
Phase status
| Item |
Status |
| Alert noise PR #255 |
Merged; AM blackhole + notify live |
| Prox mem soak tooling |
Live timers; PR #262 rebase onto main |
| Healthchecks dead-man |
Keep existing ping URL (timer already active) |
| Komodo Alerter |
Live → homelab-komodo-651b |
| Stage 7 pilot notify |
Merged #264; ntfy homelab-oneuptime-pilot-e11a |
What shipped
Mid-soak evidence (~2h, 6 samples)
Constant pressure: MemAvailable ~6–7 GiB; SwapFree often <200 MiB.
| Guest |
Avg RSS-ish |
Dominant processes (latest) |
| saltierpoop |
~20.5 GiB |
Saltbox core, video audit, Alloy, Tdarr |
| infra-services |
~20 GiB |
LiteLLM, Prometheus, LiteLLM workers |
| haos |
~4.9 GiB |
Home Assistant, influxd, supervisor |
| graylog |
~2.8 GiB |
OpenSearch datanode + Graylog JVM |
| sonarqube |
~1.1 GiB |
Sonar/ES JVMs |
Shrink/retire shortlist (after ≥24h, owner confirm destroy)
- graylog — highest consolidation leverage if syslog can move / stay
slim; SIEM is Wazuh.
- haos influxd — duplicate metrics plane vs infra-services Prometheus;
drop add-on or shrink VM memory if panels unused.
- infra-services / saltierpoop — do not retire; tune (LiteLLM workers,
Tdarr, balloon) only after smaller guests give headroom.
- sonarqube — keep; lower
maxmem if idle.
What remains
| Who |
Work |
| Agent |
Merge #262; finish 24h; open shrink PR with owner destroy confirms |
| Agent |
Finalize timer fires 2026-08-10 11:15 PDT → ntfy homelab-prox-mem-soak |
| Owner |
Subscribe that ntfy topic; reply with approve graylog / haos-influx / both / wait |
| Owner |
Stage 7 soak ≥1w then cutover |
flowchart TD
A[Max DIMMs] --> B[24h soak on infra-services]
B --> C[JSON under /var/log/homelab/prox-mem-soak]
C --> D[Shrink or retire guests]
D --> E[graylog / haos influx / balloon tune]