.wip-monitor.yml)Status: Active
First implementation: 2cld/wf/.wip-monitor.yml
Related: cross-platform-monitoring-pattern, site-status-page-pattern
Monitoring scripts accumulate hardcoded assumptions (drive letters, tar vs rsync, SSH users, port numbers). When infrastructure changes, the scripts break silently — producing false alerts or missing real problems.
The admin who changes the infrastructure is rarely the same context as the automation that monitors it. This creates drift.
Each project repo that Wip monitors publishes a .wip-monitor.yml file alongside its .wip-contract.md. This file is:
.wip-monitor.yml files┌─────────────────────────────────────────────────────────────────┐
│ Project Repos (admin-owned) │
│ │
│ repo/.wip-contract.md (human: goals, scope, contacts) │
│ repo/.wip-monitor.yml (machine: what to check) │
│ repo/site-config.yml (full infrastructure truth) │
└──────────────────────────────┬──────────────────────────────────┘
│ (coordinator reads at cron time)
▼
┌─────────────────────────────────────────────────────────────────┐
│ Coordination Node (e.g. nsdockerhv) │
│ │
│ monitor-runner.js │
│ - Fetches .wip-monitor.yml from each contracted repo │
│ - Validates schema │
│ - Executes checks per type/tier │
│ - Reports results (cron report, calendar event) │
└─────────────────────────────────────────────────────────────────┘
.wip-monitor.yml in same PR/commit (changes check method, adds drive)If admin forgets to update .wip-monitor.yml, coordinator detects drift:
| File | Audience | Purpose |
|---|---|---|
.wip-contract.md |
Humans + AI | Goals, scope, permissions, contacts |
.wip-monitor.yml |
Scripts + AI | Actionable check definitions |
site-config.yml |
Infrastructure docs | Full site truth (network, devices, services) |
The contract references the monitor file:
## Monitoring Scope
See [.wip-monitor.yml](./.wip-monitor.yml) for machine-readable check definitions.
schema: "1.0"
project: "<site-code>"
repo: "<org>/<repo>"
contact:
action: "<email>"
cc: "<email>"
access:
<hostname>:
zt_ip: "<zerotier IP>"
ssh_user: "<user>"
ssh_port: <port>
methods:
- name: "<method name>"
reliable: true/false
notes: "<context>"
checks:
- name: "<human label>"
type: <check_type>
tier: <operational|cold|glacial|scratch>
enabled: true/false
goal: "<which contract goal this serves>"
# ... type-specific fields
| Type | Purpose | Key Fields |
|---|---|---|
ping |
Node reachable | target (IP) |
http |
Service responding | target (URL), expect_status |
ssh_command |
Run command, check output | host, command, expect |
disk_space |
Volume capacity | host, volume_label, alert_below_gb, alert_below_pct |
file_age |
File freshness | host, path, max_age_hours |
backup_state |
Read .backup-state | path, key, expect |
api |
External API check | url, headers, expect |
process |
Service running | host, command (pgrep pattern) |
| Tier | Alert on low space? | Alert on unreachable? | Show in report? |
|---|---|---|---|
operational |
⚠️ YES | ❌ YES | Always |
cold |
ℹ️ info only | ⚠️ if expected UP | Always |
glacial |
No | No | On request |
scratch |
No | No | Never |
| Field | Purpose |
|---|---|
known_down: true |
Suppress alerts — tracked by issue |
issue: "<url>" |
Link to tracking issue (suppresses alerts until closed) |
volume_label |
Authoritative identifier (stable across drive letter changes) |
drive_letter |
Hint only (may change for USB/removable media) |
volume_label over drive_letter: USB drives get reassigned letters when re-plugged. Label is stable.tier controls alerting: Consistent vocabulary instead of per-check alert: true/falseknown_down + issue: Don’t nag about things already tracked. Alert resumes when issue is closed.access section: Scripts read connection info from YAML instead of hardcoding SSH users/ports.goal field: Ties each check back to .wip-contract.md goals without duplicating them.A. Runtime (recommended to start)
.wip-monitor.yml at execution timeB. Generated (future, for scale)
.wip-monitor.yml change.wip-monitor.yml for wf (proof of concept)