Category: docs/ops/deployments/
Purpose: Define abstract roles for federation compute nodes, their access boundaries, and how each site maps roles to actual usernames and hostnames.
Audience: Any federated infrastructure where multiple admins manage separate sites, and an AI coordination layer needs to observe without interfering.
When one person (or one AI) has access to everything, you get:
The pattern defines abstract roles with strict boundaries. Each site’s site-config.yml maps those roles to actual usernames and hostnames for that site.
These are the PATTERN names. They describe WHAT a role does, not WHO it is.
site-adminThe human who makes decisions and performs interactive work.
| Attribute | Definition |
|---|---|
| Type | Human, interactive |
| Access | Full (sudo, Docker, SSH, all nodes at this site) |
| Purpose | Deploy services, troubleshoot, approve changes |
| Cron | None (responds to findings, doesn’t run on schedule) |
| Repos | May push as personal identity |
| Communication | Receives action recommendations from coordination-agent |
Rule: site-admin DOES things. Deploys, configures, fixes. All infrastructure changes flow through this role.
backup-agentAutomated process that moves data on schedule. Very restrictive.
| Attribute | Definition |
|---|---|
| Type | Automated, non-interactive |
| Access | ONLY: read source data, write to backup targets, read/write .backup-state |
| Purpose | Backup execution ONLY (tar, scp, rsync, rotation) |
| Cron | Backup scripts only |
| Repos | None (no git access, no communication channels) |
| Cannot | sudo, restart services, install packages, push to git, send email |
Rule: backup-agent has the narrowest possible scope. It reads data from one place, writes it to another, and records that it did so. Nothing else.
On remote nodes: backup-agent is the SSH target that RECEIVES files. Restricted shell or forced-command that only allows SCP to designated paths.
coordination-agentAI or automation that observes and communicates. Never modifies infrastructure.
| Attribute | Definition |
|---|---|
| Type | Automated, non-interactive |
| Access | Read-only to: state files, logs, status output. Read-write to: repos (docs, issues) |
| Purpose | Monitoring, reporting, issue creation, contract management, split-brain detection |
| Cron | Status checks, morning reports |
| Repos | Push docs/issues to contracted repos |
| Cannot | sudo, Docker, SSH to remote infra nodes, modify running services |
Rule: coordination-agent is “eyes and mouth.” It READS state, REPORTS findings, WRITES documentation. If something needs fixing, it creates an issue for site-admin. It NEVER touches infrastructure.
Netstack defines abstract node roles. Each site maps these to actual hostnames.
| Role Code | Full Name | Purpose |
|---|---|---|
ng |
Network Gateway | Router, firewall, DHCP, DNS |
sg |
Storage Gateway | NAS, file server, ZFS pool host |
cg |
Compute Gateway | Primary compute (Docker, VMs, services) |
cg2 |
Secondary Compute | Additional compute capacity |
bu-0 |
Backup Target Primary | Off-site backup receiver #1 |
bu-1 |
Backup Target Secondary | Off-site backup receiver #2 |
Each site’s site-config.yml includes a roles: section that maps abstract roles to concrete names:
# site-config.yml — Role Mapping section
roles:
users:
site-admin:
username: "nsadmin" # cf uses nsadmin
auth: "SSH key (personal)"
email: "christrees@gmail.com"
notes: "Chris Trees — interactive admin"
backup-agent:
username: "buadmin" # could be any name
auth: "SSH key (id_backup)"
scope: "backup-daily.sh cron ONLY"
notes: "Restricted to backup operations"
coordination-agent:
username: "wip"
auth: "SSH key (ho-wip ed25519)"
github: "ho-wip"
email: "wip@horseoff.com"
scope: "read state, write repos"
# On REMOTE nodes (where this site receives connections FROM)
inbound:
backup-agent:
username: "buadmin" # receives SCP from federation controller
auth: "authorized_keys (nsdockerhv id_backup)"
scope: "write to ~/backups/ only"
coordination-agent:
username: "wip"
auth: "authorized_keys (nsdockerhv wip key)"
scope: "read ~/ops/site-status.json only"
nodes:
ng: "mikrotik" # wf uses MikroTik router
sg: "MediaVolume (ZFS on cg2)" # wf storage is ZFS pool
cg: "devwin10" # wf primary compute
cg2: "cg2 (Proxmox)" # wf secondary compute
bu-0: "slwin11ops" # primary backup target (sl)
bu-1: "devwin10 D:" # secondary backup target (wf local)
| Role | cf site | sl site | wf site |
|---|---|---|---|
| site-admin | nsadmin | ghadmin | ghadmin |
| backup-agent | buadmin | ghadmin (shared) | buadmin |
| coordination-agent | wip | wip (remote, on cf) | wip (remote, on cf) |
| ng | router (192.168.6.1) | Spectrum router | mikrotik |
| sg | CyberTruck D: | slwin11ops F: | MediaVolume (ZFS) |
| cg | nsdockerhv | slwin11ops WSL | LXC 100 on cg2 |
Note: sl uses ghadmin for both site-admin AND backup-agent because it’s a smaller site. The ROLE separation still applies (cron jobs run backup scripts, interactive work is separate) even when the same username fills both.
| Resource | site-admin | backup-agent | coordination-agent |
|---|---|---|---|
| sudo | ✅ | ❌ | ❌ |
| Docker | ✅ | ❌ | ❌ |
| Service start/stop | ✅ | ❌ | ❌ |
| SSH to other nodes | ✅ (manual) | ✅ (SCP only, cron) | ❌ |
| Read .backup-state | ✅ | ✅ write | 📖 read |
| Read site-status.json | ✅ | ❌ | 📖 read |
| Read/write backup dirs | ✅ | ✅ | ❌ |
| Push to git repos | ✅ (personal) | ❌ | ✅ (coordination identity) |
| Create issues | ✅ | ❌ | ✅ |
| Send email/calendar | ✅ (personal) | ❌ | ✅ (coordination identity) |
| Install packages | ✅ | ❌ | ❌ |
| Modify firewall/network | ✅ | ❌ | ❌ |
Sites grant access to each other’s coordination-agents via .wip-contract.md:
Site A admin grants → Site B's coordination-agent → read-only to state files
The contract specifies exactly what’s visible. No infrastructure access crosses site boundaries for coordination-agents.
Site A (cf) Site B (sl)
backup-agent writes state backup-agent writes state
├── .backup-state ├── site-status.json
└── logs/ └── logs/
│ (read-only) │ (read-only)
▼ ▼
┌──────────────────────────────────────────┐
│ coordination-agent (on cf) │
│ Reads both sites' state │
│ Creates issues on both repos │
│ Reports unified morning status │
└──────────────────────────────────────────┘
When the coordination-agent detects a problem:
Problem detected
├── Which site? → check site-config.yml roles.nodes mapping
├── Which role should fix it? → site-admin (infra change) or backup-agent (cron fix)
├── Where to communicate? → .wip-contract.md contact method
├── How? → Create issue on site repo with specific action recommendation
└── Who verifies? → coordination-agent reads state after fix
The coordination-agent NEVER fixes things directly. It:
The roles: section in site-config.yml enables ns-site-template to generate:
bootstrap.sh — creates the correct usernames for each roleops/scripts/ — cron scripts run as the backup-agent username| Property | How it’s achieved |
|---|---|
| Least privilege | Roles define maximum access — sites can restrict further |
| Auditability | Separate users = separate auth.log entries |
| Blast radius | Compromised coordination-agent can’t modify infra; compromised backup-agent can’t communicate |
| Revocability | Remove one key = revoke one role |
| Traceability | Git commits show which identity pushed |
| Federation-safe | Cross-site access is contract-driven and read-only |
| Site autonomy | Each site maps roles to names independently |