Applies to: Monitoring mixed Linux/Windows federation nodes from a single ops controller.
The ops controller (Linux) orchestrates all monitoring. Target nodes (Windows or Linux) answer queries via SSH. No monitoring agents needed on targets - just SSH + the target’s native shell.
Monitor the goal outcome, not every piece of infrastructure.
Don’t ask “is every container up?” Ask “are my goals being served?”
| Goal | What serves it | Monitor check |
|---|---|---|
| hwpc-rp backup | D: drive receives from cf | “Did backup arrive today?” (.backup-state fresh) |
| Sort piles | devwin10 online, drives mountable | “Can I work when I’m at wf?” (ping + SSH) |
| Federation storage | cg2 ZFS + wfMedia | “Is storage and media serving?” (plex active, ZFS healthy) |
Things NOT monitored (don’t serve a current goal):
When a new goal appears: add the monitoring check. When a goal is retired, remove its checks. Keep the signal-to-noise ratio high.
nsdockerhv (Linux, bash)
|
|-- SSH --> Linux target (bash commands)
|-- SSH --> Windows target (PowerShell commands)
|
v
.backup-state / status output / morning check-in
| Type | Runs on | Written in | Calls target via | Example |
|---|---|---|---|---|
| Monitoring | ops controller | bash | SSH + target shell | wf-status.sh |
| Setup/admin | target node | target’s native (PS1 or bash) | Run locally | setup-buadmin-windows.ps1 |
| Backup push | ops controller | bash | SSH + scp/rsync | backup-daily.sh |
#!/bin/bash
SSH_KEY="/home/buadmin/.ssh/id_backup"
SSH="ssh -i $SSH_KEY -o BatchMode=yes -o ConnectTimeout=5"
TARGET="buadmin@10.147.17.x"
# Single command
$SSH $TARGET "Get-Service sshd | Select Name,Status"
# Multi-line (use semicolons, not &&)
$SSH $TARGET "Get-Volume | Where-Object DriveLetter | Format-Table -AutoSize"
# Boolean check
RESULT=$($SSH $TARGET "Test-Connection 192.168.9.3 -Count 1 -Quiet")
if echo "$RESULT" | grep -qi "true"; then
echo "UP"
fi
| Bash | PowerShell (via SSH) | Notes |
|---|---|---|
cmd1 && cmd2 |
cmd1; cmd2 |
&& is not valid in PS |
echo "text" |
Write-Host "text" |
Or just output objects |
$VAR |
\$VAR |
Escape $ in bash heredocs |
grep pattern |
Select-String pattern |
Or pipe to findstr |
ls |
Get-ChildItem |
Or dir (alias) |
| User | Purpose | Has sudo? | SSH key |
|---|---|---|---|
| buadmin | Backup + monitoring (automated) | No | id_backup (ed25519) |
| nsadmin/ghadmin | Admin actions (manual) | Yes | Personal key |
Monitoring scripts should run as buadmin (least privilege). Until buadmin is deployed on all targets, nsadmin/ghadmin with id_backup key works.
Site status scripts are called by wip-daily-cron.sh:
# In wip-daily-cron.sh:
echo "## WF Site Status"
bash /home/nsadmin/code/wf/ops/scripts/wf-status.sh
Wip reads the output during morning standup and flags issues.