Skip to main content
Version: Current

Appliance Health and Monitoring

A School Server is healthy only when the appliance can safely perform its local role—not merely when the frontend responds.

Health dimensions​

The runtime health model covers:

AreaWhat is measured
Memorytotal/free memory and utilization pressure
CPUcore count/load; Windows hardware task also records current CPU load
Diskcapacity/free space/utilization and Windows disk-protection thresholds
PostgreSQLlocal database responsiveness
Redislocal coordination service responsiveness
Object storageMinIO/object-storage readiness
Containershost-side Compose state/health snapshot and freshness
Windows boot/watchdogautonomous runtime/stack recovery state and repeated-failure diagnostics
Backuplatest complete appliance backup and restore-verification state
LAN certificatevalidity and remaining certificate lifetime
Clockmeasured NTP/source availability and Windows time-maintenance state
Hardware degradationphysical disk health/reliability evidence, battery health, memory pressure
Release integritychecksum drift against the installed release
Update discoveryreport-only posture against the configured approved HTTPS manifest
Bootstrap/setupwhether initial data and setup are complete
Write authoritywhether writes are allowed, fenced, or authority is not ready

The overall product health state remains the worst meaningful component state: healthy, attention, critical, or unknown. Windows appliance JSON provides additional host-control evidence for operations/support.

Windows autonomous state​

A production Windows School Server writes persistent machine-readable files under its ProgramData-backed runtime directory, including:

boot-health.json
watchdog-health.json
disk-status.json
power-status.json
backup-status.json
backup-verify.json
certificate-status.json
time-status.json
integrity-status.json
hardware-status.json
update-availability.json
active-release.json
last-appliance-error.json
support\diagnostic-*.txt

Blocking operating conditions use hold-*.json. Boot/watchdog jobs intentionally respect holds. A held appliance is not “broken because it did not restart”; it may be protecting data during disk pressure, low-power shutdown, or failed/active supervised maintenance.

On Windows, inspect core state/holds with:

powershell -ExecutionPolicy Bypass -File .\windows-appliance.ps1 -Action Status

Also inspect task history when diagnosing autonomy:

Get-ScheduledTask -TaskName 'Makronexus-*' | Sort-Object TaskName | Select-Object TaskName, State
Get-ScheduledTaskInfo -TaskName 'Makronexus-Watchdog'

See Windows Autonomous Appliance for the twelve-task schedule and recovery model.

Resource thresholds​

Current runtime thresholds include:

  • product disk posture: attention at 85% used, critical at 95%;
  • product memory posture: attention at 88%, critical at 95%;
  • CPU: attention when five-minute load reaches one load unit per core, critical at two per core;
  • LAN certificate: attention at 30 days remaining, critical at 7 days or expired;
  • clock: default attention at 30 seconds absolute offset, critical at 120 seconds;
  • backup age: compared to APPLIANCE_BACKUP_MAX_AGE_MINUTES (default 1560 minutes), critical beyond twice that window;
  • Windows disk guard defaults: warning below 50 GiB free, write-protection hold below 20 GiB, recovery at 30 GiB;
  • Windows hardware task defaults: memory pressure warning at 90% and battery-health warning at 60% when firmware data is available.

Product utilization percentages and Windows absolute-free-space guards solve different problems. Keep both: one describes health posture, the other reserves emergency headroom for PostgreSQL/WAL, Docker/WSL, updates and backups.

Container supervision and self-heal​

Linux/systemd hosts can use monitor-host.sh plus install-host-monitoring.sh for recurring sanitized Compose health snapshots.

Windows adds host-level recovery beyond container restart: unless-stopped:

  1. the container-runtime task starts/reconciles Docker Desktop/Engine at startup/logon;
  2. boot reconciliation runs Compose and waits for required services;
  3. the watchdog runs every two minutes, recovering Docker/WSL where supported and reconciling unhealthy/missing containers;
  4. recovery is rate-limited;
  5. repeated failures capture sanitized diagnostics instead of looping without evidence.

Do not manually restart services repeatedly while a persistent hold exists.

LAN certificate monitoring​

The appliance reports certificate expiry posture. On Windows, the daily native certificate task checks active NIC/IP state and renews when the certificate is missing/invalid, near expiry, or hostname/IP membership changes, then reloads nginx.

The appliance-local CA stays stable. If certificate health degrades:

  1. inspect certificate-status.json and the Scheduled Task result;
  2. verify persistent CA/key material still exists;
  3. verify the configured hostname/alias and current LAN addresses;
  4. verify LAN DNS/DHCP points clients to the appliance;
  5. never instruct users to bypass TLS warnings.

Clock/NTP monitoring​

The runtime measures clock/NTP state. Windows additionally schedules Makronexus-Time-Sync daily to keep W32Time configured and request resynchronization.

WAN-offline operation can continue when an external NTP source is unavailable, but significant clock drift can invalidate replay-protected Cloud machine requests.

Backup and restore posture​

A healthy backup posture needs more than a local file timestamp.

On Windows:

  • backup-status.json records the complete nightly appliance backup result;
  • SHA-256 checksums protect snapshot integrity;
  • backup-verify.json records the weekly restore proof into a temporary PostgreSQL database;
  • external backup copy, when configured, is verified after copy;
  • a same-host backup still does not protect against total host loss.

Pair local autonomy with the governed DR policy and an independent failure domain where required.

Hardware degradation​

Makronexus-Hardware-Health runs daily and records available evidence in hardware-status.json:

  • memory pressure;
  • CPU load;
  • battery charge and estimated battery health where firmware exposes capacity data;
  • physical disk health and operational state;
  • reliability counters such as temperature, wear and errors where Windows/storage drivers expose them.

Clearly unhealthy/lost-communication storage is a blocking failure. High wear/temperature, low battery health and missing telemetry are warnings requiring review.

Release integrity and update posture​

The daily integrity task re-hashes shipped bundle files against release-checksums.json. Unexpected mismatch is a release-integrity incident, not a normal configuration change.

The update-availability task is intentionally separate from activation. If APPLIANCE_UPDATE_MANIFEST_URL is blank it reports disabled. If configured, the URL must be HTTPS and the result is report-only (CURRENT, UPDATE_AVAILABLE, etc.). Do not interpret UPDATE_AVAILABLE as authorization to install immediately.

Cloud fleet telemetry​

Authenticated appliance traffic/health reporting gives Cloud a fleet view. Depending on release support, fleet state can include application/protocol/schema version, release/channel, platform/uptime, heartbeat, bootstrap/readiness, replacement role, authority/fence state and update posture.

Use fresh telemetry to find drift/stale servers instead of relying on enrollment-time metadata.

The footer intentionally separates:

School Server reachability
+ local device/save status and write authority
+ School Server → Cloud replication status

Important states include Saved to school server, N changes held on this device, School server unavailable, Writes fenced for replacement, Write authority not ready, N changes waiting for cloud, Cloud unavailable, Cloud sync healthy, and Synchronization needs attention.

A Cloud outage is normally continuity/warning state; a missing local School Server, data-protection hold, or authority problem is more serious.

Operator response levels​

Healthy​

No action beyond normal observation.

Attention​

Investigate disk/memory pressure, battery/storage degradation, certificate renewal window, aging/failed backup verification, stale supervisor state, conflicts, clock drift, or update/integrity warnings.

Critical​

Treat database/storage unavailable, severe resource pressure, disk-protection hold, failed release integrity, unhealthy physical storage, expired/unreadable TLS, or missing required DR protection as an operational incident.

Unknown​

Unknown is not healthy. Determine whether the signal is intentionally unavailable (for example WAN-offline NTP or unsupported SMART telemetry) or whether monitoring itself failed.

Handover checklist​

  • Cloud fleet shows the expected primary School Server.
  • Health evidence is fresh.
  • PostgreSQL/Redis/MinIO health is acceptable.
  • Disk/memory have operational headroom.
  • Windows boot/watchdog task state is good when applicable.
  • No unexpected appliance hold exists.
  • Hardware status has no blocking failure.
  • LAN TLS certificate and clock state are healthy or explained.
  • Complete backup and restore verification are within policy.
  • Release integrity is acceptable.
  • Bootstrap/setup are complete and write authority is ready.
  • Footer states are understood by support staff.

Next: Backup, Restore, and Server Replacement.