Appliance Health and Monitoring
A School Server is healthy only when the appliance can safely perform its local role—not merely when the frontend responds.
Health dimensions
The runtime health model covers:
| Area | What is measured |
|---|---|
| Memory | total/free memory and utilization pressure |
| CPU | core count/load; Windows hardware task also records current CPU load |
| Disk | capacity/free space/utilization and Windows disk-protection thresholds |
| PostgreSQL | local database responsiveness |
| Redis | local coordination service responsiveness |
| Object storage | MinIO/object-storage readiness |
| Containers | host-side Compose state/health snapshot and freshness |
| Windows boot/watchdog | autonomous runtime/stack recovery state and repeated-failure diagnostics |
| Backup | latest complete appliance backup and restore-verification state |
| LAN certificate | validity and remaining certificate lifetime |
| Clock | measured NTP/source availability and Windows time-maintenance state |
| Hardware degradation | physical disk health/reliability evidence, battery health, memory pressure |
| Release integrity | checksum drift against the installed release |
| Update discovery | report-only posture against the configured approved HTTPS manifest |
| Bootstrap/setup | whether initial data and setup are complete |
| Write authority | whether writes are allowed, fenced, or authority is not ready |
The overall product health state remains the worst meaningful component state: healthy, attention, critical, or unknown. Windows appliance JSON provides additional host-control evidence for operations/support.
Windows autonomous state
A production Windows School Server writes persistent machine-readable files under its ProgramData-backed runtime directory, including:
boot-health.json
watchdog-health.json
disk-status.json
power-status.json
backup-status.json
backup-verify.json
certificate-status.json
time-status.json
integrity-status.json
hardware-status.json
update-availability.json
active-release.json
last-appliance-error.json
support\diagnostic-*.txt
Blocking operating conditions use hold-*.json. Boot/watchdog jobs intentionally respect holds. A held appliance is not “broken because it did not restart”; it may be protecting data during disk pressure, low-power shutdown, or failed/active supervised maintenance.
On Windows, inspect core state/holds with:
powershell -ExecutionPolicy Bypass -File .\windows-appliance.ps1 -Action Status
Also inspect task history when diagnosing autonomy:
Get-ScheduledTask -TaskName 'Makronexus-*' | Sort-Object TaskName | Select-Object TaskName, State
Get-ScheduledTaskInfo -TaskName 'Makronexus-Watchdog'
See Windows Autonomous Appliance for the twelve-task schedule and recovery model.
Resource thresholds
Current runtime thresholds include:
- product disk posture: attention at 85% used, critical at 95%;
- product memory posture: attention at 88%, critical at 95%;
- CPU: attention when five-minute load reaches one load unit per core, critical at two per core;
- LAN certificate: attention at 30 days remaining, critical at 7 days or expired;
- clock: default attention at 30 seconds absolute offset, critical at 120 seconds;
- backup age: compared to
APPLIANCE_BACKUP_MAX_AGE_MINUTES(default 1560 minutes), critical beyond twice that window; - Windows disk guard defaults: warning below 50 GiB free, write-protection hold below 20 GiB, recovery at 30 GiB;
- Windows hardware task defaults: memory pressure warning at 90% and battery-health warning at 60% when firmware data is available.
Product utilization percentages and Windows absolute-free-space guards solve different problems. Keep both: one describes health posture, the other reserves emergency headroom for PostgreSQL/WAL, Docker/WSL, updates and backups.
Container supervision and self-heal
Linux/systemd hosts can use monitor-host.sh plus install-host-monitoring.sh for recurring sanitized Compose health snapshots.
Windows adds host-level recovery beyond container restart: unless-stopped:
- the container-runtime task starts/reconciles Docker Desktop/Engine at startup/logon;
- boot reconciliation runs Compose and waits for required services;
- the watchdog runs every two minutes, recovering Docker/WSL where supported and reconciling unhealthy/missing containers;
- recovery is rate-limited;
- repeated failures capture sanitized diagnostics instead of looping without evidence.
Do not manually restart services repeatedly while a persistent hold exists.
LAN certificate monitoring
The appliance reports certificate expiry posture. On Windows, the daily native certificate task checks active NIC/IP state and renews when the certificate is missing/invalid, near expiry, or hostname/IP membership changes, then reloads nginx.
The appliance-local CA stays stable. If certificate health degrades:
- inspect
certificate-status.jsonand the Scheduled Task result; - verify persistent CA/key material still exists;
- verify the configured hostname/alias and current LAN addresses;
- verify LAN DNS/DHCP points clients to the appliance;
- never instruct users to bypass TLS warnings.
Clock/NTP monitoring
The runtime measures clock/NTP state. Windows additionally schedules Makronexus-Time-Sync daily to keep W32Time configured and request resynchronization.
WAN-offline operation can continue when an external NTP source is unavailable, but significant clock drift can invalidate replay-protected Cloud machine requests.
Backup and restore posture
A healthy backup posture needs more than a local file timestamp.
On Windows:
backup-status.jsonrecords the complete nightly appliance backup result;- SHA-256 checksums protect snapshot integrity;
backup-verify.jsonrecords the weekly restore proof into a temporary PostgreSQL database;- external backup copy, when configured, is verified after copy;
- a same-host backup still does not protect against total host loss.
Pair local autonomy with the governed DR policy and an independent failure domain where required.
Hardware degradation
Makronexus-Hardware-Health runs daily and records available evidence in hardware-status.json:
- memory pressure;
- CPU load;
- battery charge and estimated battery health where firmware exposes capacity data;
- physical disk health and operational state;
- reliability counters such as temperature, wear and errors where Windows/storage drivers expose them.
Clearly unhealthy/lost-communication storage is a blocking failure. High wear/temperature, low battery health and missing telemetry are warnings requiring review.
Release integrity and update posture
The daily integrity task re-hashes shipped bundle files against release-checksums.json. Unexpected mismatch is a release-integrity incident, not a normal configuration change.
The update-availability task is intentionally separate from activation. If APPLIANCE_UPDATE_MANIFEST_URL is blank it reports disabled. If configured, the URL must be HTTPS and the result is report-only (CURRENT, UPDATE_AVAILABLE, etc.). Do not interpret UPDATE_AVAILABLE as authorization to install immediately.
Cloud fleet telemetry
Authenticated appliance traffic/health reporting gives Cloud a fleet view. Depending on release support, fleet state can include application/protocol/schema version, release/channel, platform/uptime, heartbeat, bootstrap/readiness, replacement role, authority/fence state and update posture.
Use fresh telemetry to find drift/stale servers instead of relying on enrollment-time metadata.
The persistent local footer
The footer intentionally separates:
School Server reachability
+ local device/save status and write authority
+ School Server → Cloud replication status
Important states include Saved to school server, N changes held on this device, School server unavailable, Writes fenced for replacement, Write authority not ready, N changes waiting for cloud, Cloud unavailable, Cloud sync healthy, and Synchronization needs attention.
A Cloud outage is normally continuity/warning state; a missing local School Server, data-protection hold, or authority problem is more serious.
Operator response levels
Healthy
No action beyond normal observation.
Attention
Investigate disk/memory pressure, battery/storage degradation, certificate renewal window, aging/failed backup verification, stale supervisor state, conflicts, clock drift, or update/integrity warnings.
Critical
Treat database/storage unavailable, severe resource pressure, disk-protection hold, failed release integrity, unhealthy physical storage, expired/unreadable TLS, or missing required DR protection as an operational incident.
Unknown
Unknown is not healthy. Determine whether the signal is intentionally unavailable (for example WAN-offline NTP or unsupported SMART telemetry) or whether monitoring itself failed.
Handover checklist
- Cloud fleet shows the expected primary School Server.
- Health evidence is fresh.
- PostgreSQL/Redis/MinIO health is acceptable.
- Disk/memory have operational headroom.
- Windows boot/watchdog task state is good when applicable.
- No unexpected appliance hold exists.
- Hardware status has no blocking failure.
- LAN TLS certificate and clock state are healthy or explained.
- Complete backup and restore verification are within policy.
- Release integrity is acceptable.
- Bootstrap/setup are complete and write authority is ready.
- Footer states are understood by support staff.