Skip to content

Check-in — Prox 24h memory soak start

TL;DR

Prox RAM maxed (60 GiB; swap nearly full). Deep attribution soak running on infra-services. Owner judgments for Healthchecks keep, Komodo ntfy Alerter, and Stage 7 pilot ntfy are done. Shrink/retire remains after soak window.

Phase status

Item Status
Alert noise PR #255 Merged; AM blackhole + notify live
Prox mem soak tooling Live timers; PR #262 rebase onto main
Healthchecks dead-man Keep existing ping URL (timer already active)
Komodo Alerter Live → homelab-komodo-651b
Stage 7 pilot notify Merged #264; ntfy homelab-oneuptime-pilot-e11a

What shipped

Mid-soak evidence (~2h, 6 samples)

Constant pressure: MemAvailable ~6–7 GiB; SwapFree often <200 MiB.

Guest Avg RSS-ish Dominant processes (latest)
saltierpoop ~20.5 GiB Saltbox core, video audit, Alloy, Tdarr
infra-services ~20 GiB LiteLLM, Prometheus, LiteLLM workers
haos ~4.9 GiB Home Assistant, influxd, supervisor
graylog ~2.8 GiB OpenSearch datanode + Graylog JVM
sonarqube ~1.1 GiB Sonar/ES JVMs

Shrink/retire shortlist (after ≥24h, owner confirm destroy)

  1. graylog — highest consolidation leverage if syslog can move / stay slim; SIEM is Wazuh.
  2. haos influxd — duplicate metrics plane vs infra-services Prometheus; drop add-on or shrink VM memory if panels unused.
  3. infra-services / saltierpoop — do not retire; tune (LiteLLM workers, Tdarr, balloon) only after smaller guests give headroom.
  4. sonarqube — keep; lower maxmem if idle.

What remains

Who Work
Agent Merge #262; finish 24h; open shrink PR with owner destroy confirms
Agent Finalize timer fires 2026-08-10 11:15 PDT → ntfy homelab-prox-mem-soak
Owner Subscribe that ntfy topic; reply with approve graylog / haos-influx / both / wait
Owner Stage 7 soak ≥1w then cutover
flowchart TD
  A[Max DIMMs] --> B[24h soak on infra-services]
  B --> C[JSON under /var/log/homelab/prox-mem-soak]
  C --> D[Shrink or retire guests]
  D --> E[graylog / haos influx / balloon tune]