Skip to main content
Version: Current

Failure Certification

“Offline-first” is not certified by disconnecting Wi-Fi once. A production School Server release should be tested against the failure boundaries that can cause data loss, duplicate transactions, stale authority, corrupt files, or unsafe replacement.

This page defines the certification procedure and evidence expected for representative appliance environments. It does not claim a scenario has passed until the release team actually executes and records it.

Core invariants​

Every applicable failure test must preserve these invariants:

  1. Committed local data is not silently lost.
  2. A transport retry does not create a duplicate logical transaction.
  3. A blocking failure does not advance a replication/bootstrap cursor past unsafe work.
  4. A file metadata record must not be treated as usable while required bytes are missing/corrupt.
  5. Only the authorized School Server/Cloud side accepts governed writes.
  6. A replacement candidate cannot become a second primary accidentally.
  7. WAN loss does not become local ERP loss when the School Server/LAN remain healthy.
  8. Recovery is deterministic and leaves evidence an operator can understand.
  9. Restart/update preserves appliance identity, trust and encrypted machine state.
  10. The UI/footer describes the actual failure plane rather than a generic “offline” state.

Evidence record​

For every scenario record:

Release / image digests:
Deployment mode: local | hybrid
Test school / sanitized fixture:
Starting health:
Starting authority:
Starting replication cursors / pending counts:
Failure injected:
Exact timestamp:
User-visible state:
Backend/worker outcome:
Recovery action:
Ending data counts / integrity result:
Ending authority:
Ending sync/conflict state:
Pass / fail:
Evidence links / logs:

Never put real student secrets, API/HMAC secrets, private CA keys or raw production data into certification artifacts.

Matrix A — WAN and transport failures​

A1. WAN loss while idle​

  1. Start with a healthy primary School Server.
  2. Remove WAN while preserving LAN.
  3. Open the ERP from a second LAN device.
  4. Perform an eligible governed local write.
  5. Restore WAN.

Pass: local write commits; footer says Cloud unavailable/waiting for Cloud, not School Server unavailable; reconnect converges.

A2. WAN loss during School Server → Cloud push​

Inject loss after a batch is sent or while acknowledgement is uncertain.

Pass: retry is idempotent; no duplicate business transaction; push cursor advances only after safe acknowledgement.

A3. WAN loss during Cloud → School Server pull​

Interrupt while applying a batch.

Pass: already committed safe progress persists; blocking/unapplied work is retried; pull cursor does not skip work.

A4. Cloud/replay store unavailable​

Make Cloud machine authentication/replay protection unavailable.

Pass: machine sync fails closed; local School Server stays usable; disabling replay/HMAC is not required for recovery.

Matrix B — browser/LAN failures​

B1. Browser loses School Server while WAN still exists​

Pass: UI identifies School Server unavailable, not just “Cloud offline”; sensitive workflows requiring the local API are not falsely shown as committed.

B2. Device outbox vs Cloud queue​

Create a permitted device-held item, then separately create a committed local item waiting for Cloud.

Pass: footer/details distinguish “held on this device” from “waiting for cloud.”

Matrix C — bootstrap and enrollment interruptions​

C1. Power loss before one-time exchange​

Pass: no machine credential is created; fresh code can continue normally.

C2. Power loss after Cloud credential allocation but before local confirmation​

Pass: durable local credential/confirmation recovery or safe reissue can continue without creating an uncontrolled duplicate appliance.

C3. Browser/session loss after enrollment​

Pass: re-authorization/resume does not erase identity or bootstrap progress.

C4. WAN loss during verified snapshot bootstrap​

Pass: resume uses persisted checkpoint/progress; no DB wipe; final manifest/count/hash/file verification must still succeed.

C5. File corruption during bootstrap​

Alter/truncate a transferred object.

Pass: checksum/size verification rejects it; dependent file metadata is not accepted as successfully bootstrapped; retry can recover.

C6. Source state changes during bootstrap​

Create post-checkpoint changes while snapshot transfer is running.

Pass: bounded snapshot verification completes against its checkpoint, then post-checkpoint deltas converge without missing the new change.

Matrix D — service/resource failures​

D1. PostgreSQL unavailable/restart​

Pass: local write fails visibly without false commit; service recovery restores normal behavior; replication state remains coherent.

D2. Redis unavailable​

Pass: affected local coordination paths fail/recover according to policy; School Server health reports the service issue; no authority or replay bypass is introduced.

D3. MinIO/object storage unavailable​

Pass: file operations fail safely; metadata does not pretend a missing object is usable; health reflects storage failure.

D4. Disk pressure/full​

Test attention and critical thresholds, then a near/full-disk write condition.

Pass: health escalates; operations fail safely rather than corrupting state; recovery after capacity restoration is deterministic.

D5. CPU/memory pressure​

Pass: health reports pressure before/at critical thresholds and service recovery does not create duplicate sync work.

Matrix E — clock and TLS​

E1. Clock skew beyond warning threshold​

Pass: health reports attention while local ERP remains usable.

E2. Clock skew beyond machine replay tolerance​

Pass: signed Cloud requests fail; operator sees clock/NTP problem; no security control is disabled; correcting time restores sync.

E3. NTP source unavailable during WAN outage​

Pass: clock state can become unknown, local operation continues, and operator can distinguish unknown NTP from a local database outage.

E4. LAN certificate renewal window​

Run renewal check near configured threshold.

Pass: new server certificate validates under the same appliance CA and nginx reload/next use remains trusted by managed clients.

E5. Lost/unreadable certificate material​

Pass: readiness/health becomes critical instead of serving normal ERP over an untrusted/incorrect origin.

Matrix F — authority and concurrent writers​

F1. Cloud write to School-Server-owned domain​

Pass: Cloud edit is rejected by authority guard even when the user otherwise has CRUD permission.

F2. Non-primary appliance push​

Pass: replication is rejected before applying governed school data.

F3. Unclassified governed route/entity​

Pass: authority resolution fails closed rather than permitting the write.

F4. Replacement candidate local write before cutover​

Pass: candidate is write-fenced; footer reports replacement fencing.

F5. Simultaneous incompatible edit where both sides are legitimately editable by policy​

Pass: deterministic policy/conflict record is produced; no silent loss; administrator can review the conflict.

Matrix G — replacement​

G1. Healthy online replacement​

Commission candidate, bootstrap, verify health, prepare source fence, finalize cutover.

Pass: exactly one primary before and after; source stops governed writes before authority transfer; candidate becomes writable only after promotion; old identity is retired/revoked.

G2. Candidate becomes unhealthy before finalize​

Pass: cutover is rejected; old primary remains authoritative.

G3. Source is offline but not physically decommissioned​

Pass: stale/no heartbeat alone is insufficient for unsafe promotion; failed-source override requires explicit physical-offline/decommission confirmation.

G4. Dead-source override​

Physically isolate old host, provide required confirmation/reason, promote healthy candidate.

Pass: new primary becomes sole authority; old identity cannot resume Cloud replication if powered back on.

Matrix H — update/restart/recovery​

H1. Normal restart after zero-touch enrollment​

Pass: discovered tenant/school/site identity, machine credentials, CA and sync state survive; no re-enrollment required.

H2. Normal update​

Pass: trusted HTTPS, runtime identity, encrypted credentials, data/files, authority and sync state survive the packaged updater.

H3. Migration/update failure​

Pass: update stops with evidence; operator follows compatible rollback/recovery procedure; no blind old-app/new-schema combination is certified unless explicitly compatible.

H4. Rollback​

Pass: known-good release returns to a compatible data state and all field acceptance checks are repeated.

H5. Disaster restore with wrong DATA_ENCRYPTION_SECRET​

Pass: recovery is rejected/escalated rather than silently declaring success with undecryptable machine credentials.

Matrix I — mode semantics​

I1. Local only​

Pass: Cloud replication is intentionally disabled; local ERP and write authority behave as configured; UI does not report an outage merely because Cloud sync is off.

I2. Hybrid​

Pass: School Server shell/footer/setup and authority protections remain active; hybrid-specific Cloud capabilities do not bypass appliance ownership rules.

Certification verdict​

A release should not be labeled appliance-certified while a required scenario is unexecuted, unexplained or failing.

CI infrastructure failure is not certification

A GitHub Actions startup_failure, runner outage, Vercel quota error, skipped job or workflow that never creates jobs is no evidence about the code. Record it as infrastructure-blocked and obtain approved executable evidence before claiming the corresponding test/build gate passed.

Field teams do not need to run every destructive scenario at every school. Release engineering runs the destructive matrix on representative appliance environments; each school still runs the safe field acceptance subset described in Validation, Updates, and Rollback.