ドキュメント
Operations handbook
Prepare observability, scheduling, incident diagnosis, recovery, and change procedures.
Observe the chain
Monitor scheduling, runtime instances, checkpoints, data freshness, quality exceptions, service calls, decisions, and supporting infrastructure rather than relying on a single health signal.
Diagnose with evidence
Preserve timestamps, identifiers, configuration, logs, runtime metrics, and the last known change. Engineering notes show examples of tracing symptoms to scheduler, catalog, and thread-level causes.
Recover deliberately
Document retry, replay, backfill, pause, rollback, backup, and escalation procedures. Test them with representative failures and named owners before production.
This guide consolidates repository README files, deployment notes, security guidance, and anonymized engineering records. Validate environment-specific decisions during architecture review.
Open source repositories