Skip to main content
Version: Current

Backup, Restore, and Server Replacement

Recovery is part of the appliance design, not an afterthought. A school must survive database corruption, host/disk loss, a bad update, or replacement hardware without creating two authoritative writers.

Recovery layers​

Keep these separate:

  1. Local persistent data — PostgreSQL and MinIO volumes used by the live appliance.
  2. Autonomous appliance backup — on Windows, the scheduled complete snapshot plus checksum/restore-verification evidence.
  3. Known-good release + release-specific recovery point — supports safe rollback/recovery after a failed update.
  4. Disaster-recovery backup — must survive loss of the School Server failure domain.
  5. Cloud replicated copy — useful convergence/recovery input but not a substitute for a tested tenant/school backup and restore procedure.

A backup stored only on the same host/disk as the appliance does not protect against total host loss.

Windows autonomous backup layer​

A production Windows School Server registers Makronexus-Nightly-Backup and Makronexus-Weekly-Restore-Test.

Default local path:

C:\ProgramData\Makronexus\SchoolServer\backups

The nightly job briefly quiesces application writers while PostgreSQL and MinIO remain online, then captures:

  • PostgreSQL custom-format dump;
  • MinIO object data;
  • application signing/key material;
  • persistent .env.school;
  • runtime appliance identity and .env.local when present;
  • appliance local CA and TLS material;
  • release/image metadata;
  • SHA-256 checksums and backup metadata.

It applies configured retention and can copy the completed snapshot to APPLIANCE_BACKUP_EXTERNAL_DIR. External copy is checksum-verified. Because snapshots contain school data and durable secrets, use only encrypted/protected removable/NAS storage.

The weekly restore task does not merely list a dump. It validates checksums, restores the latest PostgreSQL dump into a temporary database, requires a non-empty public schema, then destroys the temporary database. Its evidence is written to runtime\backup-verify.json.

Manual acceptance commands:

powershell -ExecutionPolicy Bypass -File .\windows-appliance.ps1 -Action Backup
powershell -ExecutionPolicy Bypass -File .\windows-appliance.ps1 -Action VerifyBackup

Do not commission a Windows appliance whose first complete backup or restore proof fails.

Disaster Recovery workspace​

Makronexus also has the operator-facing Disaster Recovery workspace:

/<tenant>/operations/disaster-recovery

The product surface manages governed backup/restore jobs and their audit/status model. Use this for organization-level DR ownership rather than inventing an ad-hoc database-copy process.

A requested restore is not a completed restore. Follow the job through execution and verify the restored system before normal writes resume.

Backup policy​

Operations should define:

  • local autonomous backup frequency/retention;
  • independent backup destination and failure domain;
  • encryption/access control;
  • checksum/integrity verification;
  • restore-test cadence;
  • owner/escalation when backup age or restore verification becomes unhealthy.

On Windows, the packaged nightly/weekly schedule provides the local baseline; it does not eliminate the need for an independent DR copy.

What recovery must preserve​

A complete recovery plan accounts for:

  • tenant/school PostgreSQL data;
  • file/object data;
  • DATA_ENCRYPTION_SECRET and other durable cryptographic material required to decrypt protected local state;
  • appliance runtime identity when restoring the same logical appliance;
  • ProgramData .env.school, runtime state and local CA for same-appliance recovery;
  • release/image/schema compatibility;
  • replication cursors/conflicts/idempotency state where restored with the database;
  • audit/recovery evidence.
Encryption-key continuity

Restoring encrypted database state with the wrong DATA_ENCRYPTION_SECRET can produce a database that looks restored but cannot decrypt enrolled credentials. Preserve and verify the correct durable key material.

Restore decision tree​

Do not improvise identity changes merely because new hardware is involved.

Restore verification​

After any full restore, verify before accepting normal writes:

  • backup/restore job evidence is complete;
  • database is readable and schema matches the running release;
  • expected tenant/school/site identity is correct;
  • protected credential state can be decrypted;
  • files referenced by business records are present/readable;
  • LAN TLS identity/trust is valid;
  • Windows ProgramData runtime/environment continuity is correct when restoring the same appliance;
  • appliance health and hardware posture are acceptable;
  • replication cursors/conflict state are understood;
  • authority points to the intended primary;
  • controlled sync/convergence is safe for the restored state.

Governed School Server replacement​

Use the Cloud School Servers workspace when hardware is being replaced. Do not enroll an unrelated second primary.

Phase 1 — create replacement candidate​

Cloud creates a replacement candidate and one-time code. The candidate enrolls through zero-touch setup, generates its own stable site identity, downloads/verifies school data, reports health/readiness, remains non-primary, and stays write-fenced while commissioning.

The old server remains authoritative during this phase.

Phase 2 — prepare cutover​

For a reachable source, Cloud requests/records a source write fence. Verify the candidate is securely enrolled, bootstrap/setup complete, fresh in fleet telemetry, healthy enough for cutover, and still non-primary/fenced until promotion.

Phase 3 — finalize cutover​

The final operation promotes the candidate, transfers School Server domain authority, and retires/revokes the former primary identity.

After cutover, the new server is primary, its replacement fence is removed, the old server cannot continue valid primary replication, and footer/fleet state must reflect new authority.

Failed-source override​

If the old server is genuinely dead and cannot acknowledge a fence, use the explicit failed-source path only when source telemetry is stale/unavailable as expected, physical decommission is confirmed, the operator records a governed reason/confirmation, and the candidate is otherwise healthy and ready.

Physical decommission means the old server is powered off, removed from service, network-isolated, or otherwise controlled so it cannot resume school writes.

Never create split brain for convenience

Do not use the failed-source override while an old primary might still serve users. Two active school-owned writers are a data-integrity incident.

Interrupted replacement​

A replacement code can be reissued against the same unconfirmed candidate rather than creating uncontrolled duplicate candidates. If commissioning fails before cutover, keep the existing primary authoritative and repair/reissue the candidate path.

Post-replacement acceptance​

  • New appliance is the sole active primary.
  • Former primary is revoked/retired and physically controlled.
  • Authority records reference the intended primary.
  • Candidate write fence is gone after promotion.
  • Local login and governed writes work.
  • Files are readable.
  • Health/hardware state is fresh and acceptable.
  • New appliance backup and restore proof succeed.
  • Cloud replication converges without unexpected duplicates/conflicts.
  • Old hardware cannot return to service without an explicit new lifecycle decision.

Continue with LAN, Offline Operation, and Synchronization.