Sanitized single-commit public mirror of recipe-maintainer. - Removed test-ssh/.testenv (live creds); added test-ssh/.testenv.example placeholders. - Removed plans/ and planned-updates/ (deployment-planning docs) so no client/ deployment domains appear in the public repo. - All other secret stores were already gitignored. - docs.coopcloud.tech retained as a submodule (public upstream).
153 lines
11 KiB
Markdown
153 lines
11 KiB
Markdown
---
|
|
description: Test backing up and restoring a recipe's test instance
|
|
argument-hint: <recipe-name>
|
|
allowed-tools: [Bash, Read, Write, Glob, Grep, WebFetch]
|
|
---
|
|
|
|
# Recipe Backup Test
|
|
|
|
Test the full backup and restore cycle for a Co-op Cloud recipe: create a backup, tear down the test instance (undeploy + remove volumes), redeploy from scratch, restore from the backup, and verify the app still works.
|
|
|
|
**Important:** All `abra` commands that read the recipe (deploy, backup, restore, restart, ps) MUST use `--chaos` so they use the current local recipe checkout, including any uncommitted changes. This ensures we're testing the backup/restore labels as they exist in the working copy.
|
|
|
|
**TTY workaround:** Several `abra` subcommands (`backup create`, `backup snapshots`, `restore`, `volume remove`) fail with "the input device is not a TTY" in non-interactive environments. Wrap these with `script -qefc "..." /dev/null` to provide a pseudo-TTY.
|
|
|
|
The recipe name is: $ARGUMENTS
|
|
|
|
Read and follow the instructions in `.claude/commands/includes/logging.md`.
|
|
Read and follow the instructions in `.claude/commands/includes/guidelines.md`.
|
|
Read `~/.abra/recipes/backup-bot-two/README.md` for context on how backup-bot-two works (its CLI commands, backup/restore flow, and label conventions).
|
|
|
|
## Steps
|
|
|
|
1. **Free server resources** by running `/test-context-reset $ARGUMENTS` to undeploy unrelated apps from the test server while keeping this recipe and its dependencies running.
|
|
|
|
2. **Get the domain and server for this recipe**:
|
|
```
|
|
python3 scripts/get_test_instance.py --recipe $ARGUMENTS
|
|
```
|
|
This outputs DOMAIN and SERVER for the active instance. Also read `settings.toml` to get `DOMAIN_SUFFIX` (the `domain_suffix` field for the active instance).
|
|
- If the recipe has no `recipe-info/$ARGUMENTS/recipe.toml`, tell the user to run `/recipe-init $ARGUMENTS` first and stop.
|
|
|
|
3. **Read recipe-specific backup test notes** from `recipe-info/$ARGUMENTS/tests/backup-test.md`, if it exists.
|
|
- This optional file can contain recipe-specific instructions: extra data-seeding steps before the backup, services to check after restore, known caveats, additional verification commands, or anything else specific to testing this recipe's backup/restore cycle.
|
|
- Keep these notes in mind and apply them at the relevant steps below.
|
|
|
|
4. **Fetch the recipe and check for backup labels**:
|
|
- Fetch the recipe (check for local changes first — see guidelines).
|
|
- Read `~/.abra/recipes/$ARGUMENTS/compose.yml`.
|
|
- Search for `backupbot.backup` deploy labels.
|
|
- If **no backup labels** are found anywhere in the compose file:
|
|
- Tell the user this recipe does not have backup configured.
|
|
- Explain what's needed: at minimum, a `backupbot.backup=true` deploy label on the main service, plus appropriate pre/post hooks for databases (see the Co-op Cloud backup spec in `docs.coopcloud.tech/docs/specs/backup/`).
|
|
- Stop here.
|
|
- If backup labels are found, **summarise the backup configuration**:
|
|
- Which service has `backupbot.backup=true`
|
|
- Any `backupbot.backup.pre-hook` / `backupbot.backup.post-hook` commands
|
|
- Any `backupbot.restore.pre-hook` / `backupbot.restore.post-hook` commands
|
|
- Any volume/path restrictions (`backupbot.backup.path`, `backupbot.backup.volumes.*`)
|
|
|
|
5. **Ensure backup-bot-two is deployed on the server**:
|
|
- `abra app backup create` requires backup-bot-two to be running on the server. Check by running:
|
|
```
|
|
abra app ls --server <SERVER> --status 2>&1 | grep backup-bot
|
|
```
|
|
- If backup-bot-two is not deployed, deploy it:
|
|
- `abra recipe fetch backup-bot-two --force`
|
|
- `abra app new backup-bot-two --server <SERVER> --domain backupbot.<DOMAIN_SUFFIX> --secrets --no-input`
|
|
- `abra app deploy backupbot.<DOMAIN_SUFFIX> --force --no-input`
|
|
- If `app new` says it already exists, just check it's deployed and deploy if not.
|
|
- Confirm backup-bot-two is running before proceeding.
|
|
|
|
6. **Verify the app is deployed and healthy**:
|
|
- Run `abra app ps <DOMAIN> --chaos --no-input -m` to check deployment status (use `-m` for machine-readable output to avoid TTY issues).
|
|
- Run the existing health check script if available: `python3 recipe-info/$ARGUMENTS/tests/health_check.py`.
|
|
- If no health check script exists, curl `https://<DOMAIN>` and check for HTTP 200.
|
|
- If the app is not deployed, deploy it first: `abra app deploy <DOMAIN> --chaos --force --no-input`, then wait and re-check.
|
|
- If the backup labels were just added and the app was already deployed, **redeploy** so the running services pick up the new labels: `abra app deploy <DOMAIN> --chaos --force --no-input`.
|
|
- If `backup-test.md` specifies any data-seeding steps to perform before the backup (e.g. creating a test document, inserting a database record), do them now.
|
|
|
|
7. **Create a backup snapshot**:
|
|
- First, list existing snapshots to establish a baseline: `script -qefc "abra app backup snapshots <DOMAIN> --no-input" /dev/null`
|
|
- Note: `backup snapshots` does NOT support `--chaos`.
|
|
- Note the number of existing snapshots (may be zero).
|
|
- Run (with TTY wrapper): `script -qefc "abra app backup create <DOMAIN> --chaos --no-input" /dev/null`
|
|
- Confirm the output contains "backup finished" or similar success message.
|
|
- List snapshots again: `script -qefc "abra app backup snapshots <DOMAIN> --no-input" /dev/null`
|
|
- Confirm there is exactly one more snapshot than before.
|
|
- Note the newest snapshot ID from the output.
|
|
|
|
8. **Verify backup contents**:
|
|
- Run `backup ls` inside the backupbot container to list the files in the latest snapshot, filtered to this app:
|
|
```
|
|
script -qefc "abra app run backupbot.<DOMAIN_SUFFIX> app -- backup -h <DOMAIN> ls" /dev/null
|
|
```
|
|
- The `-h <DOMAIN>` flag filters the listing to only this app's snapshot. The `ls` command defaults to listing under `/var/lib/docker/volumes/`.
|
|
- Note: `abra app run` requires the TTY wrapper and does NOT support `--chaos`.
|
|
- Save the output, then use `grep` to confirm that expected files/paths are present:
|
|
- **Volume paths**: For each volume identified as backed up in step 3, grep the output for the volume name (the Docker volume name is `<stack_name>_<volume_name>`). Confirm that at least one file appears under each expected volume path.
|
|
- **Excluded volumes**: For any volumes explicitly excluded (e.g. `backupbot.backup.volumes.<name>=false`), confirm they do NOT appear in the listing.
|
|
- **Recipe-specific files**: If `backup-test.md` specifies particular files or patterns to look for in the backup, grep for those as well.
|
|
- If any expected volume paths are missing or any excluded volumes appear, flag as **FAIL** with details about what was missing or unexpectedly present.
|
|
|
|
9. **Undeploy the app**:
|
|
- Run: `abra app undeploy <DOMAIN> --no-input`
|
|
- Undeploy does not need `--chaos`.
|
|
|
|
10. **Remove volumes** to simulate complete data loss:
|
|
- Run (with TTY wrapper): `script -qefc "abra app volume remove <DOMAIN> --force --no-input" /dev/null`
|
|
- This ensures the restore test starts from a genuinely clean slate — no leftover data.
|
|
- **Do not skip this step** — testing restore onto existing data is not a valid backup test.
|
|
- If volume removal fails with "volume is in use" by dead/ghost containers that Docker can't remove, clear them manually via SSH:
|
|
1. Identify the ghost containers: `ssh <server> "docker ps -a --filter volume=<volume_name> --format '{{.ID}} {{.State}}'"` — they will show as `dead`
|
|
2. Remove their directories: `ssh <server> "sudo rm -rf /var/lib/docker/containers/<full_container_id>"`
|
|
3. Restart Docker: `ssh <server> "sudo systemctl restart docker"`
|
|
4. Wait ~15 seconds, then retry volume removal
|
|
|
|
11. **Redeploy from scratch**:
|
|
- Run: `abra app deploy <DOMAIN> --chaos --force --no-input`
|
|
- Wait for the app to come up (check with `abra app ps <DOMAIN> --chaos --no-input -m`).
|
|
- The app may not be fully healthy yet since it has no data — that's expected.
|
|
|
|
12. **Restore from backup**:
|
|
- Run (with TTY wrapper): `script -qefc "abra app restore <DOMAIN> --chaos --hooks --no-input" /dev/null`
|
|
- The `--hooks` flag ensures restore pre/post hooks run (e.g. database import commands).
|
|
- **Do NOT pass `--target`** to the restore command. The `--target` flag is a restic restore target *directory*, not a stack filter — passing a stack name as `--target` causes restic to restore files into a wrong subdirectory instead of the actual volume paths, and restore hooks will fail because the files won't be where the containers expect them.
|
|
- Confirm the output contains "Restoring Snapshot" or similar success message.
|
|
|
|
13. **Redeploy the app** to ensure it picks up the restored data:
|
|
- Do NOT use `abra app restart --all-services` — it tends to hang on stacks with many services.
|
|
- Instead, redeploy with force: `abra app deploy <DOMAIN> --chaos --force --no-input`
|
|
- Wait for all services to converge: `sleep 30`, then check with `abra app ps <DOMAIN> --chaos --no-input -m`. Allow up to 60 seconds for services to start. If abra reports a deploy timeout but `app ps` shows all services running and healthy, treat it as a pass.
|
|
|
|
14. **Run the test suite** to verify the app works after restore:
|
|
- Discover and run all test scripts from `recipe-info/$ARGUMENTS/tests/*.py`.
|
|
- For each script, record PASS (exit 0) or FAIL (non-zero).
|
|
- Read `recipe-info/$ARGUMENTS/test.md` and perform URL-based checks using `curl` or `WebFetch`.
|
|
- If `backup-test.md` specifies extra verification steps (e.g. "confirm the test document still exists", "check the database has records"), perform those too.
|
|
- If no tests exist at all, at minimum curl `https://<DOMAIN>` and check for HTTP 200.
|
|
|
|
15. **Summarise results**:
|
|
|
|
Report each phase of the backup/restore cycle:
|
|
|
|
| Phase | Result |
|
|
|-------|--------|
|
|
| Backup-bot-two check | PASS / FAIL |
|
|
| Initial health check | PASS / FAIL |
|
|
| Backup creation | PASS / FAIL |
|
|
| Backup contents verification | PASS / FAIL |
|
|
| Undeploy + volume removal | PASS / FAIL |
|
|
| Fresh redeploy | PASS / FAIL |
|
|
| Restore from backup | PASS / FAIL |
|
|
| Post-restore redeploy | PASS / FAIL |
|
|
| Test suite | PASS / FAIL (detail per test) |
|
|
|
|
- If all phases passed: confirm the recipe's backup/restore cycle is working correctly.
|
|
- If any phase failed: highlight which step failed, show relevant error output, and suggest troubleshooting:
|
|
- Check backup labels in compose.yml
|
|
- Check pre/post hook commands for errors
|
|
- Review app logs: `abra app logs <DOMAIN>`
|
|
- Check if all volumes are being backed up
|
|
- For database services, verify the dump/restore hooks are correct
|