Skip to content

Deploy on AWS ​

The repository ships a Terraform module in deploy/terraform that deploys the server on ECS Fargate behind an Application Load Balancer, with PostgreSQL on RDS or Aurora Serverless v2. It works with Terraform and OpenTofu, and Stackorder uses it to deploy itself, so it is also a working example of a stack Stackorder can manage.

hcl
module "stackorder" {
  source = "github.com/stackorder/stackorder//deploy/terraform?ref=v0.1.0"

  domain_name     = "stackorder.example.com"
  route53_zone_id = "Z0123456789ABCDEFGHIJ"
  image_tag       = "0.1.0"
}

Set exactly one of route53_zone_id, to have the module issue a DNS-validated certificate and create the alias record, or certificate_arn, to bring a certificate and point your own DNS at alb_dns_name. deploy/terraform/examples/ has three complete calls: complete (a new VPC, the artifact bucket and alarms), existing-vpc (your VPC and certificate) and self-hosted (the stack through which Stackorder deploys and upgrades itself).

What the module creates ​

ResourcePurpose
VPC with public and private subnets and NAT, unless create_vpc = falseTwo availability zones by default; one NAT gateway unless single_nat_gateway = false
Application Load Balancer, HTTPS listener with TLS 1.3, HTTP redirectForwards to port 8080; health checks on /readyz; 30 s deregistration delay; access logs to an S3 bucket (alb_access_logs_enabled) and a WAF web ACL (waf_web_acl_arn) are optional
ACM certificate and Route53 records, with route53_zone_idThe certificate and alias for domain_name
ECS cluster and Fargate service, 1 or 2 tasksRuns ghcr.io/stackorder/stackorder as user 65532, read-only root file system, with the deployment circuit breaker; the container health check runs stackorder-server healthcheck
RDS PostgreSQL 17 (db.t4g.micro, gp3, encrypted), or Aurora Serverless v2The database, with rds.force_ssl = 1, 7 days of point-in-time recovery, deletion protection and a final snapshot
Three Secrets Manager secretsDATABASE_URL; a JSON secret with the App credentials, STACKORDER_SESSION_KEY and STACKORDER_METRICS_TOKEN, injected through ECS secrets; and a copy of the metrics token for scrapers
Execution role and task roleThe task role has no permissions unless the artifact bucket or ECS Exec is enabled
Security groupsLoad balancer: 80 and 443 from ingress_cidrs. Tasks: 8080 from the load balancer, 443 out, Postgres to the database. Database: Postgres from the tasks
CloudWatch log groupThe server's logs, 30 days by default
Artifact bucket, with artifact_bucket_enabledFull plan text, expiring after artifact_retention_days
CloudWatch alarms, with alarms_enabledTarget 5xx, unhealthy targets, CPU, and database free storage or ACU utilization

The server needs nothing else. It holds no AWS credentials for your infrastructure.

Security groups filter by address, not by host name, so the tasks' egress is HTTPS to anywhere, through NAT: the tasks pull the image from ghcr.io and reach the GitHub API, GitHub's OIDC keys, Secrets Manager and CloudWatch Logs over their public endpoints. For stricter egress, put a proxy or firewall on the NAT path.

Inputs ​

The tables follow the descriptions in deploy/terraform/variables.tf and deploy/terraform/outputs.tf, and docs/test fails when they drift apart. Every input also has validation rules, which the descriptions summarise.

Naming ​

InputTypeDefaultDescription
namestring"stackorder"Name of the deployment, used as the name or prefix of every resource. Lowercase letters, digits and single hyphens, 2 to 24 characters.
tagsmap(string){}Tags added to every resource that supports them.

Network ​

InputTypeDefaultDescription
create_vpcbooltrueCreate a VPC with public and private subnets and NAT. When false, vpc_id, public_subnet_ids and private_subnet_ids are required.
vpc_cidrstring"10.0.0.0/16"IPv4 CIDR of the VPC created when create_vpc is true. Subnets are carved as eight equal blocks: public from the first four, private from the last four.
vpc_idstringnullID of an existing VPC. Required when create_vpc is false, must be null otherwise.
public_subnet_idslist(string)[]IDs of existing public subnets in at least two availability zones, for the load balancer. Required when create_vpc is false.
private_subnet_idslist(string)[]IDs of existing private subnets in at least two availability zones, with a route to the internet through NAT, for the tasks and the database. Required when create_vpc is false.
availability_zoneslist(string)[]Availability zones for the created VPC. Empty picks the first two available zones of the region.
single_nat_gatewaybooltrueUse one NAT gateway for all private subnets instead of one per availability zone.

Name, TLS and ingress ​

InputTypeDefaultDescription
domain_namestringrequiredFully qualified host name the server is reached at, such as stackorder.example.com.
route53_zone_idstringnullRoute53 hosted zone in which to create the ACM validation records and the alias record for domain_name. Exactly one of route53_zone_id and certificate_arn must be set.
certificate_arnstringnullARN of an existing ACM certificate covering domain_name. DNS for domain_name is then left to the caller. Exactly one of route53_zone_id and certificate_arn must be set.
base_urlstringnullPublic URL of the server when it differs from https://<domain_name>, for example behind another proxy. No trailing slash.
ssl_policystring"ELBSecurityPolicy-TLS13-1-2-2021-06"Security policy of the HTTPS listener. Must be a TLS 1.3 policy.
ingress_cidrslist(string)["0.0.0.0/0"]IPv4 or IPv6 CIDRs allowed to reach the load balancer on ports 80 and 443. GitHub webhooks and GitHub-hosted runners need the default. Ignored when github_webhook_ip_ranges_only is true.
github_webhook_ip_ranges_onlyboolfalseRestrict the load balancer to GitHub's webhook source ranges (the hooks list of the GitHub meta API, read at plan time) plus admin_cidrs, instead of ingress_cidrs.
admin_cidrslist(string)[]CIDRs of people and self-hosted runners that need the UI and API when github_webhook_ip_ranges_only is true.
alb_access_logs_enabledboolfalseWrite load balancer access logs to an S3 bucket the module creates, encrypted with SSE-S3 as ELB log delivery requires.
alb_access_logs_retention_daysnumber90Days after which objects in the access log bucket expire.
waf_web_acl_arnstringnullARN of a regional AWS WAFv2 web ACL in the module's region to associate with the load balancer. Null associates none.

Service ​

InputTypeDefaultDescription
imagestring"ghcr.io/stackorder/stackorder"Container image repository of the server.
image_tagstring"latest"Tag or digest (sha256:...) of the server image. Pin a release such as 1.2.3 so upgrades are explicit plans.
desired_countnumber1Number of server tasks. All coordination goes through Postgres, so a second task adds availability without any other change.
cpunumber256Fargate task CPU units.
memorynumber512Fargate task memory in MiB; must be a valid combination with cpu.
cpu_architecturestring"X86_64"CPU architecture of the task, X86_64 or ARM64.
enable_execute_commandboolfalseEnable ECS Exec. The SSM agent needs a writable root file system, so this also turns readonlyRootFilesystem off and grants the task role the ssmmessages permissions.
health_check_commandlist(string)["CMD", "/stackorder-server", "healthcheck"]Container health check command, starting with CMD or CMD-SHELL. The default runs the server's healthcheck subcommand, which GETs /healthz on the listen port. The distroless image has no shell or curl, so a replacement must be a command the image itself provides. Empty turns the container health check off and leaves task health to the load balancer check on /readyz.
stop_timeout_secondsnumber60Seconds ECS waits after SIGTERM before it kills the container (stopTimeout), 2 to 120 on Fargate. The server drains HTTP for up to 15 s, then its workers for up to 30 s plus 5 s for cancelled handlers, so a value under 50 can cut the drain short.
wait_for_steady_statebooltrueMake terraform apply wait until the new tasks pass /readyz, so an apply of an upgrade fails when the deployment rolls back.
log_retention_daysnumber30Retention of the server log group in days.
log_levelstring"info"Server log level (STACKORDER_LOG_LEVEL).
extra_environmentmap(string){}Additional environment variables for the server, such as STACKORDER_WORKERS or OTEL_EXPORTER_OTLP_ENDPOINT. Variables the module sets itself are rejected.

Database ​

InputTypeDefaultDescription
engine_versionstring"17"PostgreSQL major version, or major.minor. For Aurora a major version resolves to the AWS default minor of that major at plan time.
allow_major_version_upgradeboolfalseAllow a new major version in engine_version to upgrade the RDS instance or Aurora cluster in place. A major upgrade cannot be rolled back; take a snapshot first and set apply_immediately for the same apply.
apply_immediatelyboolfalseApply database changes, such as engine_version, instance_class or the parameter group, at once instead of in the next maintenance window. Changes that need a restart then cause a short outage.
instance_classstring"db.t4g.micro"RDS instance class. Ignored when use_aurora_serverless is true.
allocated_storagenumber20Initial RDS storage in GiB. Ignored when use_aurora_serverless is true.
max_allocated_storagenumber100Upper bound for RDS storage autoscaling in GiB; 0 disables autoscaling. Ignored when use_aurora_serverless is true.
multi_azboolfalseRun the RDS instance Multi-AZ, or add an Aurora reader in another zone.
deletion_protectionbooltrueProtect the database from deletion.
skip_final_snapshotboolfalseSkip the final database snapshot on destroy. Keep false outside of throwaway environments.
backup_retention_daysnumber7Automated backup retention in days; point-in-time recovery covers this window.
performance_insightsboolfalseEnable Performance Insights with the free 7 day retention. Not every instance class supports it.
use_aurora_serverlessboolfalseUse an Aurora PostgreSQL Serverless v2 cluster instead of an RDS instance. Switching an existing deployment replaces the database.
aurora_min_acunumber0.5Minimum Aurora Serverless v2 capacity in ACUs.
aurora_max_acunumber2Maximum Aurora Serverless v2 capacity in ACUs.
kms_key_arnstringnullCustomer managed KMS key for the Secrets Manager secrets and database storage. Null uses the AWS managed keys.

GitHub App and server settings ​

InputTypeDefaultDescription
github_app_idstringnullGitHub App id (GITHUB_APP_ID). Leave the App inputs null on the first deploy: the server starts in setup mode and /setup creates the App.
github_app_private_keystring (sensitive)nullPEM private key of the GitHub App (GITHUB_APP_PRIVATE_KEY).
github_webhook_secretstring (sensitive)nullWebhook secret of the GitHub App (GITHUB_WEBHOOK_SECRET).
github_oauth_client_idstringnullOAuth client id of the GitHub App, for human sign-in (GITHUB_OAUTH_CLIENT_ID).
github_oauth_client_secretstring (sensitive)nullOAuth client secret of the GitHub App (GITHUB_OAUTH_CLIENT_SECRET).
session_keystring (sensitive)null32 byte hex key for cookie signing (STACKORDER_SESSION_KEY). Null generates one.
metrics_tokenstring (sensitive)nullBearer token that GET /metrics requires (STACKORDER_METRICS_TOKEN), at least 16 printable ASCII characters without white space. Null generates 32 hexadecimal characters.
secret_recovery_window_daysnumber30Days Secrets Manager keeps a deleted secret recoverable; 0 deletes immediately.
github_api_urlstring"https://api.github.com"GitHub API base URL (GITHUB_API_URL); GitHub Enterprise Server uses https://<host>/api/v3.
required_workflow_refstringnullGlob that runner tokens' job_workflow_ref must match (STACKORDER_REQUIRED_WORKFLOW_REF), such as stackorder/actions/.github/workflows/.yml@refs/tags/v1. Null accepts any workflow.
oidc_audiencestringnullAudience runner OIDC tokens must carry (STACKORDER_OIDC_AUDIENCE). Null uses the public URL.

Artifact bucket and alarms ​

InputTypeDefaultDescription
artifact_bucket_enabledboolfalseCreate an S3 bucket for full plan text (STACKORDER_ARTIFACT_BUCKET) and grant the task role access to it. This is the only AWS permission the task role ever gets.
artifact_retention_daysnumber90Days after which objects in the artifact bucket expire.
alarms_enabledboolfalseCreate CloudWatch alarms for target 5xx responses, unhealthy targets and service CPU, plus database free storage (RDS) or ACU utilization (Aurora, whose storage grows on its own).
alarm_sns_topic_arnstringnullSNS topic notified when an alarm changes state. Required when alarms_enabled is true.
alarm_thresholdsobject{}Alarm thresholds: target 5xx responses per 5 minutes, average service CPU percent, RDS free storage in bytes, and Aurora ACU utilization percent.

What the module passes to the server ​

VariableSource
STACKORDER_BASE_URLbase_url, or https://<domain_name>
STACKORDER_LISTEN:8080
GITHUB_API_URLgithub_api_url
STACKORDER_OIDC_AUDIENCEoidc_audience, or the base URL
STACKORDER_REQUIRED_WORKFLOW_REFrequired_workflow_ref, when set
STACKORDER_ARTIFACT_BUCKETThe artifact bucket, when artifact_bucket_enabled
STACKORDER_LOG_LEVELlog_level
DATABASE_URLThe database secret: postgres://stackorder:<password>@<endpoint>:5432/stackorder?sslmode=require
GITHUB_APP_ID, GITHUB_APP_PRIVATE_KEY, GITHUB_WEBHOOK_SECRET, GITHUB_OAUTH_CLIENT_ID, GITHUB_OAUTH_CLIENT_SECRETThe App secret, when set
STACKORDER_SESSION_KEYThe App secret; session_key, or generated
STACKORDER_METRICS_TOKENThe App secret; metrics_token, or generated. Scrapers read the same value from the <name>/metrics-token secret

Everything else the server reads, such as STACKORDER_WORKERS, the retention durations, STACKORDER_LOG_FORMAT, GITHUB_WEB_URL, GITHUB_OIDC_ISSUER or OTEL_EXPORTER_OTLP_ENDPOINT, goes through extra_environment, which rejects the variables above. Values in extra_environment are plain text in the task definition, so keep secrets out of it; the metrics token has its own input, metrics_token, and scrapers read it from the secret named by the metrics_token_secret_arn output, see Security hardening. On GitHub Enterprise Server set github_api_url and add GITHUB_WEB_URL and GITHUB_OIDC_ISSUER to extra_environment.

Container health check ​

By default the container runs stackorder-server healthcheck every 30 s with a 5 s timeout, 3 retries and a 30 s start period. It asks /healthz, which does not touch the database, so the container check fails only when the process itself stops answering. The load balancer's /readyz check, database included, gates deployments and drives the circuit breaker, and ECS also replaces a task that fails it, so a database outage still makes ECS replace tasks, whatever health_check_command is. A replacement task then exits at start-up until the database answers again. Set health_check_command = [] to turn the container check off. stop_timeout_seconds (60 by default) gives the server time to drain HTTP and its workers before ECS sends SIGKILL.

Outputs ​

OutputDescription
urlPublic URL of the server (STACKORDER_BASE_URL).
alb_dns_nameDNS name of the load balancer; point a CNAME or alias here when DNS is not managed by the module.
alb_zone_idRoute53 zone id of the load balancer, for alias records.
setup_urlPage that creates the GitHub App from a manifest on the first deploy.
webhook_urlWebhook URL of the GitHub App.
ecs_cluster_nameName of the ECS cluster.
ecs_service_nameName of the ECS service.
task_definition_arnARN of the current task definition revision.
db_endpointHost name of the database writer endpoint.
db_secret_arnARN of the Secrets Manager secret holding DATABASE_URL.
app_secret_arnARN of the Secrets Manager secret holding the GitHub App credentials, the session key and the metrics token as JSON.
metrics_token_secret_arnARN of the Secrets Manager secret holding only the /metrics bearer token, as plain text; grant Prometheus read access to this one rather than to the app secret.
artifact_bucketName of the artifact bucket, or null when artifact_bucket_enabled is false.
alb_access_logs_bucketName of the load balancer access log bucket, or null when alb_access_logs_enabled is false.
security_group_idsSecurity group ids of the load balancer, the service and the database.
log_group_nameCloudWatch log group of the server.

First deployment ​

The server starts in setup mode while the GitHub App inputs are unset, serving only /setup, /healthz and /readyz. The first deployment uses that:

  1. Apply the module with the github_* inputs unset. The service comes up in setup mode.
  2. Open the setup_url output and create the App. The page prints the App id, private key, webhook secret and OAuth client id and secret once.
  3. Apply again with those five values. The App id, private key and webhook secret must be set together, as must the two OAuth values.
  4. Install the App on your repositories and continue with Getting started.

The values pass through Terraform state. Keep state encrypted and readable only by the roles that plan and apply this stack. The task definition carries the version ids of both secrets as Docker labels, so changing a secret value rolls the service onto it.

Upgrades ​

  1. Change image_tag to the new X.Y.Z, or a sha256: digest, and apply.
  2. ECS starts a new task while the old one keeps serving (100 % minimum healthy, up to 200 % during the deployment). The new task runs any database migrations at start-up, under a migration lock.
  3. The new task receives traffic once it passes /readyz; the old one drains for 30 s and stops. A task that never becomes healthy is rolled back by the circuit breaker, and wait_for_steady_state makes the apply fail.

Pin a specific version rather than latest, so an apply is the only thing that changes the running version. Take a database snapshot before an upgrade that includes migrations. See Upgrades and backups.

Scaling out ​

Set desired_count = 2 for availability. No other change is needed:

  • webhook handling, workers and the API are stateless;
  • all coordination goes through Postgres, and workers claim work with SKIP LOCKED;
  • the scheduler is single-leader through a Postgres advisory lock, so one task schedules and any task executes;
  • sessions live in Postgres, and nothing on local disk matters.

Both tasks read the same STACKORDER_SESSION_KEY from the App secret. Raise cpu and memory before adding tasks, and consider db.t4g.small or Aurora once the database is the bottleneck.

Testing the module ​

sh
cd deploy/terraform
terraform init -backend=false
terraform test

The tests plan against mock providers and need no AWS account. They need Terraform 1.11 or later; using the module as a child module works from Terraform 1.9.

Managing the module with Stackorder ​

Put the module call in its own stack, with its own backend "s3" block, in a repository that has Stackorder installed, as examples/self-hosted does. Plans for the server's own changes then run like any other stack. Applies depend on the server being up, so for an upgrade that might not come back healthy, keep a way to apply the stack by hand.