Skip to main content
Version: Current

Field Troubleshooting

Troubleshoot by preserving state and identifying the exact failure plane. Do not wipe PostgreSQL, regenerate appliance identity, bypass TLS/authority, or disable signed machine authentication as a first response.

First response​

Record, where safely available:

  • school and release version;
  • exact step/state that failed;
  • whether the LAN School Server is reachable;
  • whether https://school.makronexus.local is trusted;
  • whether WAN/Cloud is reachable;
  • footer local-save, write-authority and Cloud-sync states;
  • appliance health state;
  • whether this is primary, replacement candidate, or retired server;
  • recent install/update/enrollment/replacement actions.

Do not copy secrets or real sensitive school records into a general support ticket.

Core diagnostics​

From the exact release bundle:

docker compose --env-file .env.school -f docker-compose.yml ps
./validate-env.sh ./.env.school
./health-check.sh ./.env.school

Targeted logs:

docker compose --env-file .env.school -f docker-compose.yml logs --tail=200 nginx backend frontend
docker compose --env-file .env.school -f docker-compose.yml logs --tail=200 postgres redis minio
docker compose --env-file .env.school -f docker-compose.yml logs --tail=200 notifications-worker events-outbox-worker

Never paste raw .env.school, runtime/.env.local, JWTs, enrollment codes, API/HMAC secrets, DATA_ENCRYPTION_SECRET, CA private keys or student/finance data into tickets/chat.

Preflight/release verification fails​

Likely causes: missing CLI, corrupt/mixed artifact, wrong trust key, changed release file, missing exact image.

Action: stop and obtain/fix the approved release. Do not edit checksum/signature metadata to force a pass and do not substitute an uncontrolled image tag.

Final release archive/path does not match school-bundle-release-*​

The Release Console finalization step renames intermediate packaging names into the public Makronexus artifact stem. A finalized offline handoff normally uses names like:

makronexus-school-server-<releaseVersion>-offline
makronexus-school-server-<releaseVersion>-offline.tar.gz
makronexus-school-server-<releaseVersion>-offline.tar.gz.sha256.json

Do not assume the intermediate school-bundle-release-<releaseVersion> path still exists.

Read the final names from RELEASE-REPORT.json:

$Report = Get-Content .\RELEASE-REPORT.json -Raw | ConvertFrom-Json
$Report.artifacts.releaseDirectory
$Report.artifacts.archive
$Report.artifacts.archiveChecksum

Require status: ready, distributable: true, no DO-NOT-DISTRIBUTE.txt, and then verify/extract using the report-named archive/checksum plus the extractor inside the report-named release directory.

If a command tries to execute scripts\extract-school-bundle-release.ps1 from a directory that has not been extracted or from a retired intermediate name, fix the operator command/runbook. Do not rename the finalized release artifacts by hand.

Environment validation fails on a fresh zero-touch install​

Normal fresh installs leave tenant, school and site identity unset. Do not “fix” a zero-touch profile by inventing UUIDs.

Check instead:

  • local secrets contain no placeholders;
  • required local secrets are strong/distinct;
  • the deployment profile is consistently zero-touch or consistently advanced pre-provisioned;
  • no partial CLOUD_SYNC_URL / API-key / HMAC configuration is present;
  • licensing fields match the intended enforcement policy.

A pre-provisioned profile must set its complete identity/credential contract consistently.

School URL does not resolve​

The canonical URL is:

https://school.makronexus.local

First identify where resolution is failing.

School Server itself cannot resolve the name​

On Windows, the appliance maintenance path keeps a managed self-resolution block in:

C:\Windows\System32\drivers\etc\hosts

Expected managed block:

# BEGIN MAKRONEXUS SCHOOL SERVER LAN IDENTITY
127.0.0.1 school.makronexus.local school.makronexus.lan
# END MAKRONEXUS SCHOOL SERVER LAN IDENTITY

Check:

Get-ScheduledTask -TaskName 'Makronexus-Certificate-Maintenance'
Get-ScheduledTaskInfo -TaskName 'Makronexus-Certificate-Maintenance'
Get-Content "$env:SystemRoot\System32\drivers\etc\hosts"
[System.Net.Dns]::GetHostAddresses('school.makronexus.local')

To prove HTTPS/backend health independently of DNS while preserving the hostname/SNI, use Windows curl with an explicit mapping. For a private appliance CA, Schannel may be unable to perform public CRL lookup; --ssl-no-revoke disables only that revocation lookup and does not disable hostname/trust validation:

curl.exe --ssl-no-revoke `
--resolve "school.makronexus.local:443:127.0.0.1" `
https://school.makronexus.local/api/v1/cloud-sync/appliances/platform/local/status

If that succeeds but Windows name resolution fails, diagnose the server self-resolution path rather than nginx/backend.

Other LAN clients cannot resolve the name​

Inspect the current mDNS task and evidence:

Get-ScheduledTask -TaskName 'Makronexus-mDNS-Responder'
Get-ScheduledTaskInfo -TaskName 'Makronexus-mDNS-Responder'
Get-Content C:\ProgramData\Makronexus\SchoolServer\active\runtime\mdns-responder-status.json -Raw

A recent PASS with the expected LAN IPv4 proves the responder has answered a query; it does not prove every client/VLAN permits mDNS. Also check:

  • server has the expected LAN address;
  • client and server are on a LAN/VLAN that permits client-to-server traffic and, where used, mDNS;
  • otherwise managed DHCP/DNS maps the canonical name to the server;
  • ports 80/443 are not blocked;
  • the configured appliance hostname matches the intended URL.

Use direct-IP/explicit-host diagnostics only to distinguish name resolution from application failure. Do not make the IP address the permanent user URL.

Browser certificate warning​

Do not click through as the production solution.

  1. obtain the appliance CA from:
http://<school-server-ip>/makronexus-root-ca.crt
  1. install/trust it through the managed-device policy;
  2. verify the browser uses https://school.makronexus.local;
  3. inspect runtime/tls/ and certificate health if the warning remains;
  4. rerun the supported release LAN/certificate maintenance path while preserving the existing CA unless the recovery plan intentionally changes trust.

If the appliance CA private key is lost, treat that as a trust/recovery incident rather than silently minting a new CA and surprising every client.

Trusted HTTPS readiness fails​

health-check.sh verifies both HTTP bootstrap health and HTTPS /readyz using the appliance CA and canonical host resolution.

Check nginx/backend/frontend, certificate files, configured hostname, and whether the server certificate is valid under the local CA. A running frontend container is not sufficient.

NTP/clock health is attention or critical​

Check host time-sync service and configured APPLIANCE_NTP_SERVER.

If WAN/NTP is temporarily unavailable, local ERP can continue. If clock drift is large, signed Cloud requests may fail until the clock is corrected.

Do not increase/disable the replay window or turn off HMAC protection as a routine clock fix.

One-time code expired/lost​

Generate/reissue a fresh code in Operations → Offline Sync → School Servers only when the appliance still needs enrollment authorization.

The hardened lifecycle reuses the same unconfirmed machine/candidate where appropriate and revokes abandoned temporary exchange state. Do not create a succession of unrelated appliances just because a browser session was lost.

If local status already reports enrollment.enrolled: true, do not generate a new enrollment code merely because bootstrap or a later setup step failed.

Enrollment code appears used after interrupted setup​

Do not wipe the local server or change its site identity.

Possible safe states include:

  • local credentials were already durably stored and machine confirmation can finish;
  • the commissioning record needs a governed code reissue;
  • setup session needs re-authorization.

Use the same appliance/replacement record and supported recovery path.

Zero-touch identity is missing after restart​

A successful enrollment should persist non-secret identity under the appliance runtime directory and mount/load it on later starts.

Check:

  • runtime/.env.local exists on the host;
  • the runtime directory is the same durable directory used by install/update;
  • file ownership permits backend access while keeping secrets protected;
  • update/redeploy did not replace the runtime directory with a blank path.

Unexpected re-enrollment after a normal restart/update is a stop/escalate condition.

Bootstrap interrupted​

Preserve the local database and runtime state. Restore connectivity and use the resume path.

Do not:

  • delete PostgreSQL;
  • reset checkpoint/cursors manually;
  • mark bootstrap complete by hand;
  • accept a file verification mismatch as success;
  • regenerate SITE_ID;
  • re-enroll an appliance that already reports enrollment.enrolled: true.

The verified bootstrap must still satisfy its checkpoint/manifest/payload/file evidence before setup can be considered complete.

Bootstrap fails before snapshot transfer​

If local status shows bootstrap failed with snapshotId absent/null and cursor/applied totals still zero, the failure occurred before authoritative snapshot data was applied.

For errors such as:

bootstrap_snapshot_local_preflight_query_failed:<entity>

collect the exact preflight error and verify the local schema/release contract. Do not create an ad-hoc table or mutate schema simply to bypass the preflight. If the root cause is a software defect fixed in a later signed School Server release, update the already-enrolled appliance in place, preserve all stateful volumes/identity, then resume bootstrap.

File record exists but object will not open​

Treat as an integrity/storage incident.

Check MinIO/object-storage health, file-transfer logs and checksum/size errors. Do not upload unrelated bytes just to make an object key exist.

School server unavailable​

This means the browser cannot reach the local API. Check LAN, nginx/backend, database dependencies and gateway health.

Do not confuse this with Cloud unavailable, where onsite operation can remain normal.

N changes held on this device​

Those changes have not reached the School Server database. Restore local API/LAN connectivity and let the permitted device outbox commit them. They are not Cloud-queue changes yet.

Cloud unavailable / N changes waiting for cloud​

The work is already committed locally. Check WAN, Cloud DNS/TLS, clock, machine credential/revocation state, protocol/schema compatibility and Cloud service health.

Do not switch to Local only simply to remove the warning.

Write authority not ready​

The School Server cannot positively establish governed write authority. Reads may remain available but writes fail closed.

Check:

  • enrolled appliance identity;
  • active primary role in Cloud;
  • authority policy for the school/domain;
  • control-plane refresh/Cloud reachability where required;
  • whether this server is an unfinished replacement candidate;
  • backend logs for authority classification/mismatch.

Do not bypass the guard at the API/database layer.

Writes fenced for replacement​

This is normally intentional during a replacement cutover. Do not treat it as a generic sync failure.

If this is the old primary, finish or cancel the governed replacement procedure before reopening writes. If it is the replacement candidate, it remains fenced until promotion.

Replacement will not finalize​

Common safe blockers:

  • candidate bootstrap incomplete;
  • setup incomplete;
  • health not healthy enough;
  • candidate telemetry stale;
  • source server still primary and not fenced;
  • failed-source override lacks explicit physical-decommission confirmation/reason.

Fix the blocker. Do not manually flip primary/authority fields in the database.

Source server is dead during replacement​

Physically isolate/decommission it first. Then use the Cloud failed-source override with the required explicit confirmation/reason and only after the candidate is healthy/ready.

Network silence alone is not proof that the old server cannot return and create split brain.

Cloud write rejected for authority violation​

Determine which side owns the domain for this school. If School Server authority is intentional, perform the change on the authoritative School Server/workflow. Do not loosen the middleware simply because a Cloud user has CRUD permission.

If the domain classification appears wrong or unknown, escalate it as an authority-map defect.

Synchronization conflict / attention​

A semantic conflict is not a transport retry. Use the governed conflict-review workflow. Do not delete local changes or manually advance cursors to make the warning disappear.

Protocol/app/schema compatibility rejection​

Update the stale/incompatible side according to the approved release path. Do not bypass compatibility headers/checks and do not apply a newer schema change stream to a runtime outside the supported contract.

Disk/memory/CPU health warning​

Use the health details to identify pressure. For disk, act before capacity reaches the critical range. Do not delete PostgreSQL/MinIO volumes or audit/recovery material indiscriminately to create space.

PostgreSQL / Redis / MinIO critical​

Inspect the specific service and host storage/resource condition. Preserve volumes. Restart/repair the narrow dependency according to operations policy.

Deleting a volume is a data-loss operation, not a generic service restart.

Backup age critical​

A same-host data volume does not clear the disaster-recovery warning. Restore off-machine backup scheduling/connectivity or execute the approved backup operation and verify completion.

If no completed DR backup exists, treat the protection posture as critical until resolved.

Update unexpectedly requests enrollment​

Stop. A normal update should preserve runtime/, PostgreSQL, enrolled credential state, CA and discovered identity.

Check whether the correct runtime/database volumes were reused and whether DATA_ENCRYPTION_SECRET changed. Do not blindly re-enroll against a blank/incorrect database because that can hide lost local state.

Update or migration fails​

Capture evidence and follow Validation, Updates, and Rollback. Do not run an older app against a newer migrated schema unless compatibility is explicitly known.

Appliance restore cannot decrypt credentials​

Verify the restored environment has the correct DATA_ENCRYPTION_SECRET. Do not overwrite encrypted credential fields or substitute a new key to force startup. Follow the approved restore/replacement procedure.

Escalation package​

Provide a sanitized package with:

  • school non-sensitive identifier;
  • release/version;
  • exact failed step and timestamp/timezone;
  • primary/candidate role;
  • docker compose ps output;
  • relevant health states;
  • footer authority/sync state;
  • targeted logs with secrets/data removed;
  • whether LAN, trusted HTTPS, WAN and NTP work;
  • bootstrap/sync/conflict/replacement state;
  • recent update/enrollment/cutover activity.

Never use these emergency shortcuts​

Do not delete data volumes without an approved recovery decision, edit runtime/.env.local to impersonate another school, reuse a retired site identity on new hardware, disable HMAC/replay protection, bypass write authority/fencing, manually promote a replacement in the database, advance sync cursors, ignore file checksum failures, accept browser TLS warnings as normal, rename finalized release artifacts to match an obsolete command, or mix release scripts/manifests.

After resolving an incident, rerun the applicable checks in Failure Certification or Validation, Updates, and Rollback.