Failure Certification
“Offline-first” is not certified by disconnecting Wi-Fi once. A production School Server release should be tested against the failure boundaries that can cause data loss, duplicate transactions, stale authority, corrupt files, or unsafe replacement.
This page defines the certification procedure and evidence expected for representative appliance environments. It does not claim a scenario has passed until the release team actually executes and records it.
Core invariants
Every applicable failure test must preserve these invariants:
- Committed local data is not silently lost.
- A transport retry does not create a duplicate logical transaction.
- A blocking failure does not advance a replication/bootstrap cursor past unsafe work.
- A file metadata record must not be treated as usable while required bytes are missing/corrupt.
- Only the authorized School Server/Cloud side accepts governed writes.
- A replacement candidate cannot become a second primary accidentally.
- WAN loss does not become local ERP loss when the School Server/LAN remain healthy.
- Recovery is deterministic and leaves evidence an operator can understand.
- Restart/update preserves appliance identity, trust and encrypted machine state.
- The UI/footer describes the actual failure plane rather than a generic “offline” state.
Evidence record
For every scenario record:
Release / image digests:
Deployment mode: local | hybrid
Test school / sanitized fixture:
Starting health:
Starting authority:
Starting replication cursors / pending counts:
Failure injected:
Exact timestamp:
User-visible state:
Backend/worker outcome:
Recovery action:
Ending data counts / integrity result:
Ending authority:
Ending sync/conflict state:
Pass / fail:
Evidence links / logs:
Never put real student secrets, API/HMAC secrets, private CA keys or raw production data into certification artifacts.
Matrix A — WAN and transport failures
A1. WAN loss while idle
- Start with a healthy primary School Server.
- Remove WAN while preserving LAN.
- Open the ERP from a second LAN device.
- Perform an eligible governed local write.
- Restore WAN.
Pass: local write commits; footer says Cloud unavailable/waiting for Cloud, not School Server unavailable; reconnect converges.
A2. WAN loss during School Server → Cloud push
Inject loss after a batch is sent or while acknowledgement is uncertain.
Pass: retry is idempotent; no duplicate business transaction; push cursor advances only after safe acknowledgement.
A3. WAN loss during Cloud → School Server pull
Interrupt while applying a batch.
Pass: already committed safe progress persists; blocking/unapplied work is retried; pull cursor does not skip work.
A4. Cloud/replay store unavailable
Make Cloud machine authentication/replay protection unavailable.
Pass: machine sync fails closed; local School Server stays usable; disabling replay/HMAC is not required for recovery.
Matrix B — browser/LAN failures
B1. Browser loses School Server while WAN still exists
Pass: UI identifies School Server unavailable, not just “Cloud offline”; sensitive workflows requiring the local API are not falsely shown as committed.
B2. Device outbox vs Cloud queue
Create a permitted device-held item, then separately create a committed local item waiting for Cloud.
Pass: footer/details distinguish “held on this device” from “waiting for cloud.”
Matrix C — bootstrap and enrollment interruptions
C1. Power loss before one-time exchange
Pass: no machine credential is created; fresh code can continue normally.
C2. Power loss after Cloud credential allocation but before local confirmation
Pass: durable local credential/confirmation recovery or safe reissue can continue without creating an uncontrolled duplicate appliance.
C3. Browser/session loss after enrollment
Pass: re-authorization/resume does not erase identity or bootstrap progress.
C4. WAN loss during verified snapshot bootstrap
Pass: resume uses persisted checkpoint/progress; no DB wipe; final manifest/count/hash/file verification must still succeed.
C5. File corruption during bootstrap
Alter/truncate a transferred object.
Pass: checksum/size verification rejects it; dependent file metadata is not accepted as successfully bootstrapped; retry can recover.
C6. Source state changes during bootstrap
Create post-checkpoint changes while snapshot transfer is running.
Pass: bounded snapshot verification completes against its checkpoint, then post-checkpoint deltas converge without missing the new change.
Matrix D — service/resource failures
D1. PostgreSQL unavailable/restart
Pass: local write fails visibly without false commit; service recovery restores normal behavior; replication state remains coherent.
D2. Redis unavailable
Pass: affected local coordination paths fail/recover according to policy; School Server health reports the service issue; no authority or replay bypass is introduced.
D3. MinIO/object storage unavailable
Pass: file operations fail safely; metadata does not pretend a missing object is usable; health reflects storage failure.
D4. Disk pressure/full
Test attention and critical thresholds, then a near/full-disk write condition.
Pass: health escalates; operations fail safely rather than corrupting state; recovery after capacity restoration is deterministic.
D5. CPU/memory pressure
Pass: health reports pressure before/at critical thresholds and service recovery does not create duplicate sync work.
Matrix E — clock and TLS
E1. Clock skew beyond warning threshold
Pass: health reports attention while local ERP remains usable.
E2. Clock skew beyond machine replay tolerance
Pass: signed Cloud requests fail; operator sees clock/NTP problem; no security control is disabled; correcting time restores sync.
E3. NTP source unavailable during WAN outage
Pass: clock state can become unknown, local operation continues, and operator can distinguish unknown NTP from a local database outage.
E4. LAN certificate renewal window
Run renewal check near configured threshold.
Pass: new server certificate validates under the same appliance CA and nginx reload/next use remains trusted by managed clients.
E5. Lost/unreadable certificate material
Pass: readiness/health becomes critical instead of serving normal ERP over an untrusted/incorrect origin.
Matrix F — authority and concurrent writers
F1. Cloud write to School-Server-owned domain
Pass: Cloud edit is rejected by authority guard even when the user otherwise has CRUD permission.
F2. Non-primary appliance push
Pass: replication is rejected before applying governed school data.
F3. Unclassified governed route/entity
Pass: authority resolution fails closed rather than permitting the write.
F4. Replacement candidate local write before cutover
Pass: candidate is write-fenced; footer reports replacement fencing.
F5. Simultaneous incompatible edit where both sides are legitimately editable by policy
Pass: deterministic policy/conflict record is produced; no silent loss; administrator can review the conflict.
Matrix G — replacement
G1. Healthy online replacement
Commission candidate, bootstrap, verify health, prepare source fence, finalize cutover.
Pass: exactly one primary before and after; source stops governed writes before authority transfer; candidate becomes writable only after promotion; old identity is retired/revoked.
G2. Candidate becomes unhealthy before finalize
Pass: cutover is rejected; old primary remains authoritative.
G3. Source is offline but not physically decommissioned
Pass: stale/no heartbeat alone is insufficient for unsafe promotion; failed-source override requires explicit physical-offline/decommission confirmation.
G4. Dead-source override
Physically isolate old host, provide required confirmation/reason, promote healthy candidate.
Pass: new primary becomes sole authority; old identity cannot resume Cloud replication if powered back on.
Matrix H — update/restart/recovery
H1. Normal restart after zero-touch enrollment
Pass: discovered tenant/school/site identity, machine credentials, CA and sync state survive; no re-enrollment required.
H2. Normal update
Pass: trusted HTTPS, runtime identity, encrypted credentials, data/files, authority and sync state survive the packaged updater.
H3. Migration/update failure
Pass: update stops with evidence; operator follows compatible rollback/recovery procedure; no blind old-app/new-schema combination is certified unless explicitly compatible.
H4. Rollback
Pass: known-good release returns to a compatible data state and all field acceptance checks are repeated.
H5. Disaster restore with wrong DATA_ENCRYPTION_SECRET
Pass: recovery is rejected/escalated rather than silently declaring success with undecryptable machine credentials.
Matrix I — mode semantics
I1. Local only
Pass: Cloud replication is intentionally disabled; local ERP and write authority behave as configured; UI does not report an outage merely because Cloud sync is off.
I2. Hybrid
Pass: School Server shell/footer/setup and authority protections remain active; hybrid-specific Cloud capabilities do not bypass appliance ownership rules.
Certification verdict
A release should not be labeled appliance-certified while a required scenario is unexecuted, unexplained or failing.
A GitHub Actions startup_failure, runner outage, Vercel quota error, skipped job or workflow that never creates jobs is no evidence about the code. Record it as infrastructure-blocked and obtain approved executable evidence before claiming the corresponding test/build gate passed.
Field teams do not need to run every destructive scenario at every school. Release engineering runs the destructive matrix on representative appliance environments; each school still runs the safe field acceptance subset described in Validation, Updates, and Rollback.