netstack

Compute Roles Pattern

Purpose: Define standard compute-role taxonomy for federation sites. Exemplar: wf compute architecture Related: netstack#17 (BMR), netstack#18


The Problem

Federation sites tend to collapse all workloads onto one machine: always-on infrastructure, cold storage, and ad-hoc workbench tasks all share hardware. This creates conflicts:

The Pattern: Role-Based Compute Assignment

Every site assigns hardware to one of four standard roles based on power profile and uptime requirements:

Role Power Profile Uptime BMR Target? Purpose
infra Always-on, low power (5-15W) 24/7 YES Tunnels, monitoring, WoL triggers, coordination
glacial Mostly OFF, WoL-scheduled Scheduled windows No Cold storage, backup receive, archival media
recovery Manual, powered during active recovery As-needed No Data recovery from old/damaged media
sort Manual, powered when physically present As-needed No Drive triage, indexing, data migration workbench

Not every site needs all four roles. A minimal site (like sl) might only have infra (WSL Docker on the one machine). A site with heavy storage (like wf) uses all four.

Role Definitions

infra (always-on gateway)

What it does:

Hardware requirements:

Candidate platforms:

MikroTik note: Only higher-end models support containers. Requirements: RouterOS v7.4+, ARM64 or x86 CPU, 1GB+ RAM. Models like RB951G (128MB RAM) CANNOT run containers - they remain network gateway (ng) role only.

BMR: This is the primary BMR target. bootstrap.sh from netstack#17 rebuilds this role from site-config.yml.

glacial (cold storage)

What it does:

Hardware requirements:

Candidate platforms:

Multi-unit pattern: Sites with multiple clients or isolation requirements can run multiple glacial units - one per client. Each unit has independent WoL schedule, independent failure domain.

Operational pattern:

infra sends WoL -> glacial boots -> mounts pool -> backup job runs
-> health check -> reports status -> shuts down -> infra logs result

recovery (temporary data extraction)

What it does:

Hardware requirements:

Lifecycle: Recovery is a temporary role. Once data is extracted and migrated to glacial storage, the hardware either becomes another glacial unit or gets decommissioned.

sort (workbench)

What it does:

Hardware requirements:

What happens to sorted data:

site-config.yml Integration

Sites declare compute roles in their site-config.yml:

compute:
  roles:
    infra:
      host: rpi5
      platform: docker
      always_on: true
      bmr_target: true
    glacial:
      - host: 1u-srv-01
        platform: linux
        wol_mac: "AA:BB:CC:DD:EE:01"
        schedule: "0 2 * * 0"
        client: "federation"
    recovery:
      host: 1u-srv-sg
      purpose: "Synology sg drive extraction"
      temporary: true
    sort:
      host: cg2
      purpose: "USB shelf + SAS drive triage"

Mapping Your Site

To apply this pattern:

  1. Inventory your hardware - what machines exist at the site?
  2. Identify the power profiles - what MUST be always-on? What can sleep?
  3. Assign roles - match hardware to roles based on power profile + capabilities
  4. Document in site-config.yml - add compute.roles section
  5. Plan migration - if everything is on one box today, plan phased separation

Examples

Site infra glacial recovery sort
wf LXC 100 (Phase 1) -> RPi 5 (Phase 2) 1U fleet (per-client) 1U with sg drives cg2 (SAS card)
sl slwin11ops (WSL Docker) slwin11ops (same machine, backup-receive only) – –
cf nsdockerhv (Docker VM on CyberTruck) CyberTruck local storage – –