Applies to: Any federation service that needs to be rebuilt after failure.
Every service must be recoverable from: backup data + repo config + documented steps. If you can’t recover a service by following a doc, the service isn’t properly documented.
Every service needs these three things documented:
1. DATA - what to restore (database, files, secrets)
2. CONFIG - how to rebuild the environment (repo, compose, scripts)
3. STEPS - ordered procedure to go from nothing to running
Each service’s repo (or site docs) should have a recovery section:
## Recovery: <service-name>
### What You Need
- [ ] Backup data from: <location> (describe what files/DB)
- [ ] Repo access: <repo URL> (code + config)
- [ ] Secrets: <where .env/tokens live> (vault, private backup)
- [ ] Host: <what to run it on> (specs, OS, network requirements)
### Backup Locations
| Data | Primary backup | Secondary backup |
|------|---------------|-----------------|
| <database> | <location> | <location> |
| <files> | <location> | <location> |
| <secrets> | <location> | <location> |
### Recovery Steps
1. Stand up host per [ops-node-setup](https://netstack.org/docs/ops/users/ops-node-setup/)
2. Clone repo: `git clone <repo>`
3. Restore secrets (.env) from <vault/backup location>
4. Restore data from backup: <specific commands>
5. Start service: <specific commands>
6. Verify: <how to confirm it's working>
7. Update DNS/access if host IP changed
### Recovery Time Estimate
- Host setup: X hours
- Data restore: X hours (depends on backup size + transfer speed)
- Verification: X minutes
- Total: X hours
### Last Tested
- Date: <when was this procedure last verified>
- Result: <pass/fail + notes>
For full site recovery (everything at a site is gone):
| Pattern | Role in recovery |
|---|---|
| federation-setup-guide | Node/OS standup |
| ops-node-setup | User environment |
| docker-service-pattern | Container rebuild |
| ssh-rsync-pattern | How backup data transfers |
| federation-backup-plan | Where backups live |
| service-lifecycle-pattern | Migration gates |
Site site-config.yml |
Service definitions + dependencies |
Per the recovery template, every service should have “Last Tested” documented. Schedule DR tests periodically:
A service without a tested recovery procedure is a risk, not an asset.