Upgrades and backups
Three things carry versions: the server image, the CLI, and the actions. Postgres is the only state to back up.
What is versioned
| Component | Where | Versions |
|---|---|---|
| Server | ghcr.io/stackorder/stackorder | X.Y.Z and latest, from tags vX.Y.Z on stackorder/stackorder |
| CLI | Release assets on stackorder/stackorder | stackorder_X.Y.Z_<os>_<arch>.tar.gz with a SHA-256 checksum file |
| Actions and reusable workflows | stackorder/actions | v1 and v1.x.y; the v1 tag moves independently of server releases |
Workflows reference the reusable workflows at @v1. The stackorder-version input pins the CLI release that setup installs; it defaults to latest.
The API between the CLI and the server is v1. Fields are only ever added and both sides ignore fields they do not know, so the server and the CLI can be upgraded independently within v1.
Upgrading the server
- Read the release notes for the new version.
- Take a database snapshot if the release includes migrations.
- Deploy the new image tag.
- The new instance runs pending migrations at start-up, under a migration lock, so only one instance applies them.
- The previous instance keeps serving until the new one passes
/readyz.
On AWS with the Terraform module, this is a change of the image tag and an apply. See Deploy on AWS.
In-flight runs survive an upgrade. Work is queued in Postgres and every handler is idempotent: on SIGTERM an instance hands unstarted work back to the queue and gives running handlers 30 s, and anything a killed instance had claimed is claimed again after 10 minutes. The per-minute reconciliation catches any wave that finished during the switch. GitHub does not retry a failed webhook delivery by itself; redeliver it from the App's delivery log if needed.
To go back, deploy the previous image. The server never runs down migrations, so if the new version ran migrations the previous one does not understand, restore the snapshot from step 2.
Upgrading the CLI and actions
- Set
stackorder-versionin the calling workflows to pin the CLI, and bump it to move. Left atlatest, every run installs the newest release. - The
v1tag ofstackorder/actionsmoves on its own. Callingplan.yml@v1.x.yinstead of@v1pins the workflow file, but the reusable workflows still call the actions at@v1. - If the server sets
STACKORDER_REQUIRED_WORKFLOW_REFor the AWS roles pinjob_workflow_ref, keep the pinned pattern in step with the tags you use.refs/tags/v1*covers everyv1release.
Backups
Postgres holds the graphs, the run history, the locks, drift results, sessions, API keys and the work queue. Terraform state is never there.
- RDS or Aurora: turn on automated backups with point-in-time recovery.
- Elsewhere: use your platform's snapshots or
pg_dumpon a schedule.
What a lost database costs you: history, drift results and locks. Plans and applies keep working against your S3 state once a new database is in place, because nothing in the server is needed to operate Terraform.
Restoring
- Stop the server, or scale it to zero.
- Restore the database to the chosen point.
- Start the server. It runs any migrations the restored schema is missing.
- Check
GET /v1/overviewfor locks held, and compare them with open pull requests. A lock restored from before an apply finished, or a lock lost in the restore, needs a human decision: re-apply from the pull request, or release it withstackorder unlock. - Push to open pull requests, or comment
stackorder plan, to re-plan anything whose results were lost.
Rotating secrets
| Secret | How to rotate |
|---|---|
| App private key | Generate a new key in the App settings, update GITHUB_APP_PRIVATE_KEY, redeploy, then delete the old key. Runners never see it. |
| Webhook secret | Change it in the App settings and in GITHUB_WEBHOOK_SECRET together; deliveries in between fail signature checks and can be redelivered from the App's delivery log. |
| OAuth client secret | Generate a new one in the App settings, update GITHUB_OAUTH_CLIENT_SECRET, redeploy. |
| Session key | Change STACKORDER_SESSION_KEY on every instance at once. Everyone is signed out. |
| API keys | Only their SHA-256 is stored. Create a new key, switch the automation to it, then revoke the old one; see API keys. |
There are no runner-side secrets to rotate: runners authenticate with short-lived OIDC tokens.