Backup, Restore, and Server Replacement
Recovery is part of the appliance design, not an afterthought. A school must survive database corruption, host/disk loss, a bad update, or replacement hardware without creating two authoritative writers.
Recovery layers
Keep these separate:
- Local persistent data — PostgreSQL and MinIO volumes used by the live appliance.
- Autonomous appliance backup — on Windows, the scheduled complete snapshot plus checksum/restore-verification evidence.
- Known-good release + release-specific recovery point — supports safe rollback/recovery after a failed update.
- Disaster-recovery backup — must survive loss of the School Server failure domain.
- Cloud replicated copy — useful convergence/recovery input but not a substitute for a tested tenant/school backup and restore procedure.
A backup stored only on the same host/disk as the appliance does not protect against total host loss.
Windows autonomous backup layer
A production Windows School Server registers Makronexus-Nightly-Backup and Makronexus-Weekly-Restore-Test.
Default local path:
C:\ProgramData\Makronexus\SchoolServer\backups
The nightly job briefly quiesces application writers while PostgreSQL and MinIO remain online, then captures:
- PostgreSQL custom-format dump;
- MinIO object data;
- application signing/key material;
- persistent
.env.school; - runtime appliance identity and
.env.localwhen present; - appliance local CA and TLS material;
- release/image metadata;
- SHA-256 checksums and backup metadata.
It applies configured retention and can copy the completed snapshot to APPLIANCE_BACKUP_EXTERNAL_DIR. External copy is checksum-verified. Because snapshots contain school data and durable secrets, use only encrypted/protected removable/NAS storage.
The weekly restore task does not merely list a dump. It validates checksums, restores the latest PostgreSQL dump into a temporary database, requires a non-empty public schema, then destroys the temporary database. Its evidence is written to runtime\backup-verify.json.
Manual acceptance commands:
powershell -ExecutionPolicy Bypass -File .\windows-appliance.ps1 -Action Backup
powershell -ExecutionPolicy Bypass -File .\windows-appliance.ps1 -Action VerifyBackup
Do not commission a Windows appliance whose first complete backup or restore proof fails.
Disaster Recovery workspace
Makronexus also has the operator-facing Disaster Recovery workspace:
/<tenant>/operations/disaster-recovery
The product surface manages governed backup/restore jobs and their audit/status model. Use this for organization-level DR ownership rather than inventing an ad-hoc database-copy process.
A requested restore is not a completed restore. Follow the job through execution and verify the restored system before normal writes resume.
Backup policy
Operations should define:
- local autonomous backup frequency/retention;
- independent backup destination and failure domain;
- encryption/access control;
- checksum/integrity verification;
- restore-test cadence;
- owner/escalation when backup age or restore verification becomes unhealthy.
On Windows, the packaged nightly/weekly schedule provides the local baseline; it does not eliminate the need for an independent DR copy.
What recovery must preserve
A complete recovery plan accounts for:
- tenant/school PostgreSQL data;
- file/object data;
DATA_ENCRYPTION_SECRETand other durable cryptographic material required to decrypt protected local state;- appliance runtime identity when restoring the same logical appliance;
- ProgramData
.env.school, runtime state and local CA for same-appliance recovery; - release/image/schema compatibility;
- replication cursors/conflicts/idempotency state where restored with the database;
- audit/recovery evidence.
Restoring encrypted database state with the wrong DATA_ENCRYPTION_SECRET can produce a database that looks restored but cannot decrypt enrolled credentials. Preserve and verify the correct durable key material.
Restore decision tree
Do not improvise identity changes merely because new hardware is involved.
Restore verification
After any full restore, verify before accepting normal writes:
- backup/restore job evidence is complete;
- database is readable and schema matches the running release;
- expected tenant/school/site identity is correct;
- protected credential state can be decrypted;
- files referenced by business records are present/readable;
- LAN TLS identity/trust is valid;
- Windows ProgramData runtime/environment continuity is correct when restoring the same appliance;
- appliance health and hardware posture are acceptable;
- replication cursors/conflict state are understood;
- authority points to the intended primary;
- controlled sync/convergence is safe for the restored state.
Governed School Server replacement
Use the Cloud School Servers workspace when hardware is being replaced. Do not enroll an unrelated second primary.
Phase 1 — create replacement candidate
Cloud creates a replacement candidate and one-time code. The candidate enrolls through zero-touch setup, generates its own stable site identity, downloads/verifies school data, reports health/readiness, remains non-primary, and stays write-fenced while commissioning.
The old server remains authoritative during this phase.
Phase 2 — prepare cutover
For a reachable source, Cloud requests/records a source write fence. Verify the candidate is securely enrolled, bootstrap/setup complete, fresh in fleet telemetry, healthy enough for cutover, and still non-primary/fenced until promotion.
Phase 3 — finalize cutover
The final operation promotes the candidate, transfers School Server domain authority, and retires/revokes the former primary identity.
After cutover, the new server is primary, its replacement fence is removed, the old server cannot continue valid primary replication, and footer/fleet state must reflect new authority.
Failed-source override
If the old server is genuinely dead and cannot acknowledge a fence, use the explicit failed-source path only when source telemetry is stale/unavailable as expected, physical decommission is confirmed, the operator records a governed reason/confirmation, and the candidate is otherwise healthy and ready.
Physical decommission means the old server is powered off, removed from service, network-isolated, or otherwise controlled so it cannot resume school writes.
Do not use the failed-source override while an old primary might still serve users. Two active school-owned writers are a data-integrity incident.
Interrupted replacement
A replacement code can be reissued against the same unconfirmed candidate rather than creating uncontrolled duplicate candidates. If commissioning fails before cutover, keep the existing primary authoritative and repair/reissue the candidate path.
Post-replacement acceptance
- New appliance is the sole active primary.
- Former primary is revoked/retired and physically controlled.
- Authority records reference the intended primary.
- Candidate write fence is gone after promotion.
- Local login and governed writes work.
- Files are readable.
- Health/hardware state is fresh and acceptable.
- New appliance backup and restore proof succeed.
- Cloud replication converges without unexpected duplicates/conflicts.
- Old hardware cannot return to service without an explicit new lifecycle decision.
Continue with LAN, Offline Operation, and Synchronization.