Validation, Updates, and Rollback
A School Server release is stateful infrastructure. Installation/update decisions must preserve school data, appliance identity, local trust material, machine credentials, synchronization state and write-authority safety.
Part 1 — Post-install acceptance
Infrastructure and trusted gateway
From the release bundle:
docker compose --env-file .env.school -f docker-compose.yml ps
./health-check.sh ./.env.school
health-check.sh requires both the HTTP bootstrap health endpoint and trusted HTTPS readiness using the appliance CA and canonical school.makronexus.local identity.
Do not certify the site because the frontend container alone is running.
Windows autonomous appliance acceptance
On Windows, also prove the host-level control plane is installed and operational:
Get-ScheduledTask -TaskName 'Makronexus-*' |
Sort-Object TaskName |
Select-Object TaskName, State
powershell -ExecutionPolicy Bypass -File .\windows-appliance.ps1 -Action Status
Production acceptance requires all thirteen Makronexus-* tasks, persistent state beneath C:\ProgramData\Makronexus\SchoolServer, no unexplained hold-*.json, a successful complete backup, successful restore verification, and a controlled Windows reboot that recovers Docker/Compose without manual repair.
The current task set includes Makronexus-mDNS-Responder; runtime\mdns-responder-status.json is part of LAN-discovery evidence.
See Windows Autonomous Appliance.
LAN trust
From a second managed LAN device:
- the appliance CA is trusted;
https://school.makronexus.localresolves;- there is no browser certificate warning;
/appliance-setupworks before commissioning and normal login works after setup.
The School Server itself also maintains a managed loopback hosts entry for canonical self-resolution. A server-side name-resolution failure and a LAN-client mDNS/DNS failure are different incidents and should be diagnosed separately.
Enrollment/bootstrap
Confirm:
- zero-touch enrollment discovered the correct school;
- stable site identity is persisted;
- bootstrap completed with its integrity checks;
- required file objects are usable;
- intended sync mode is configured;
- write authority is ready for the primary server;
- a replacement candidate, if present, remains fenced until cutover.
Appliance health
Confirm current posture for PostgreSQL, Redis, object storage, containers, disk/memory/CPU, backup/restore evidence, certificate and clock. On Windows also inspect hardware-status.json and any appliance holds. Resolve critical states before handover.
Synchronization and offline continuity
Run a controlled bidirectional test where authority permits it, then disconnect WAN while leaving LAN intact. Confirm local work continues, Cloud queue accumulates after local commit, and reconnect converges safely.
See LAN, Offline Operation, and Synchronization.
Part 2 — Repository/release certification is separate
Field acceptance proves a school installation. It does not prove the software release was built/tested correctly.
Use the exact repository release pipeline and record the evidence for the release candidate. If GitHub Actions/Vercel/another runner fails before jobs start, there is no test result. A startup_failure, runner-capacity problem, quota limit or skipped workflow must never be recorded as a passing build.
See Failure Certification for the destructive/recovery matrix that release owners should exercise on representative appliance environments.
Part 3 — Before updating
Treat every update as a controlled maintenance event.
Before running it:
- verify the new release is approved and its trust/checksum material is intact;
- retain the known-good previous release and release-specific rollback information;
- review migration/schema compatibility;
- protect a current approved recovery point according to operations policy;
- record current appliance health, authority, conflicts and replication queue;
- allow pending synchronization to drain when practical, but do not compromise local continuity to force it;
- preserve durable environment/runtime state, PostgreSQL/MinIO volumes and
DATA_ENCRYPTION_SECRET; - reconcile environment-template changes without overwriting discovered runtime identity;
- run the new release's preflight/validation.
On Windows, the canonical mutable environment and runtime state are under:
C:\ProgramData\Makronexus\SchoolServer\
.env.school
runtime\
backups\
active\
The release-local .env.school and runtime paths are links into that persistent state. Do not copy a new template over ProgramData merely because a new release was extracted.
Identity and trust that must persist
A normal update preserves the appliance rather than re-enrolling it. Do not intentionally regenerate:
- persisted tenant/school/site runtime identity;
DATA_ENCRYPTION_SECRET;- enrolled API/HMAC credential state in PostgreSQL;
- appliance-local CA/private key;
- LAN hostname policy;
- synchronization cursors/conflicts/idempotency state;
- local school data/files.
A changed software version is not a reason to create a new site identity.
If enrollment is already successful but initial bootstrap previously failed before any snapshot was applied, the supported recovery is to update the software in place, preserve the enrolled appliance identity, then resume bootstrap. Do not generate a new enrollment code merely because the bootstrap implementation was fixed in a later release.
Part 4 — Current packaged update flow
Linux / generic bundle
Use the new release's own updater:
./update.sh ./.env.school
The bundle prepares/preserves runtime state, runs release/environment checks, maintains time/LAN identity, pulls/uses exact release images, runs migrations, rolls the application services, and verifies trusted health.
Windows normal update
Run from the new signed release bundle in an elevated PowerShell session:
powershell -ExecutionPolicy Bypass -File .\update.ps1
The Windows wrapper prepares the persistent ProgramData environment/runtime links before invoking the normal bundle update. After the release is healthy, it refreshes the autonomous Scheduled Tasks and switches the stable C:\ProgramData\Makronexus\SchoolServer\active junction to the new release.
Scheduled tasks therefore do not depend on a technician leaving an arbitrary extracted release path untouched forever.
Windows supervised update
For a maintenance-window update that must prove a recovery point first, use:
powershell -ExecutionPolicy Bypass -File .\supervised-update.ps1
The supervisor:
- creates a persistent
maintenancehold so watchdog/boot reconciliation does not fight the update; - creates a complete pre-update appliance backup;
- verifies that backup by restoring its PostgreSQL dump into a temporary database;
- runs the packaged update;
- clears the hold only after successful completion.
If the update fails, the maintenance hold remains. That is intentional fail-closed behavior: collect evidence and resolve the failed release instead of allowing a restart storm around a partially updated appliance.
For a normal zero-touch/UI-enrolled appliance, CLOUD_BOOTSTRAP_MODE=skip preserves existing enrolled state; an update does not require re-enrollment.
A Windows appliance may check an explicitly configured HTTPS APPLIANCE_UPDATE_MANIFEST_URL, but that task only writes runtime\update-availability.json. It never downloads or activates a release automatically. Release activation remains a governed maintenance action.
Part 5 — Post-update validation
Repeat the important production checks:
- trusted HTTPS readiness passes;
- existing school data and files are present;
- runtime identity/site identity is unchanged;
- appliance CA continuity is intact;
- local login/navigation works;
- health posture is acceptable;
- Cloud fleet heartbeat/version posture updates;
- write authority still points to the intended primary;
- synchronization loads with no new blocking conflict/integrity error;
- local → Cloud and Cloud → local controlled tests succeed where authority allows;
- WAN-offline continuity still works for changes touching offline/runtime/sync infrastructure.
On Windows also confirm:
-
active-release.jsonpoints at the intended release; - the ProgramData
activejunction targets the intended release; - all thirteen
Makronexus-*tasks remain registered; -
Makronexus-mDNS-Responderis running/recoverable andmdns-responder-status.jsonis current when LAN discovery is exercised; - no unexpected maintenance/disk/power hold remains;
-
windows-appliance.ps1 -Action Statusis acceptable; - a controlled reboot still self-recovers after changes to the Windows control layer.
For an update performed to recover a previously failed initial bootstrap, also confirm enrollment remains the same appliance/site identity before resuming bootstrap.
Part 6 — Stop conditions
Stop and collect evidence when:
- release verification/checksum validation fails;
- environment validation fails unexpectedly;
- migrations fail;
- trusted HTTPS readiness fails;
- the appliance unexpectedly loses discovered identity or asks for a new enrollment;
- ProgramData environment/runtime continuity is broken;
- local CA/key material disappears;
- data/files appear missing;
DATA_ENCRYPTION_SECRETchanged or is unavailable;- write authority becomes ambiguous;
- synchronization reports new blocking compatibility/integrity failures;
- the Windows maintenance hold cannot be safely cleared because update outcome is uncertain.
Do not solve a failed update by deleting ProgramData state, deleting stateful volumes, removing holds without resolving their cause, or inventing a new appliance identity.
Part 7 — Rollback
The bundle includes rollback.sh / rollback.ps1. Use the rollback helper from the failed release according to the release handoff rather than mixing old scripts with new manifests.
Rollback has two dimensions:
- application/image rollback;
- database/data compatibility after any migrations that ran.
An older application is not automatically safe against a newer migrated schema. Before rollback, know whether the migration is backward compatible or whether an approved data restore/recovery point is required.
The Windows supervised updater supports automatic rollback only when explicitly invoked with a previous release path and -AllowAutomaticRollback. Do not enable that switch for a release unless schema backward compatibility is explicitly certified.
Part 8 — Evidence collection
Before destructive recovery, capture sanitized evidence:
docker compose --env-file .env.school -f docker-compose.yml ps
docker compose --env-file .env.school -f docker-compose.yml logs --tail=200 nginx backend frontend
On Windows also retain relevant persistent runtime JSON and runtime\support\diagnostic-*.txt. Never paste .env.school, enrollment codes, JWTs, API/HMAC secrets, private CA keys or sensitive student data into general support channels.
Production handover checklist
- Exact release evidence is recorded.
- Trusted LAN HTTPS is healthy.
- Zero-touch identity is correct/persistent.
- Bootstrap and file verification completed.
- Write authority is ready.
- Appliance health is acceptable.
- Windows autonomy/reboot/backup/restore acceptance passed when applicable.
- DR/backup responsibility is recorded.
- Bidirectional sync and WAN-offline behavior were tested.
- Known-good release/rollback plan is retained.
- Replacement/recovery procedure is understood.
Next read Security and Credential Lifecycle, Failure Certification, and keep Field Troubleshooting available during rollout.