Environments and Guardrails: The Two New Gates at Stage Three
What a GitHub environment actually is, the wait timer that isn't a queue, and why cancelling a deploy mid-sync is only safe when the work has no partial side effects.
A security reviewer looks at the stable-stage pipeline and finds three things: long-lived keys, a public bucket, and actions pinned to tags that can silently change underneath you.
None of those get fixed by writing more YAML in the same two files. They get fixed by two new mechanisms this stage introduces: environments and concurrency control.
An environment is a named bundle of secrets and rules
A GitHub environment is exactly that: a name (production), a set of secrets scoped only to deploys targeting it, and a set of rules that apply before any job runs against it.
Adding one to a job is one line:
jobs:
deploy:
environment: production
runs-on: ubuntu-latestOnce that's there, the environment's rules apply automatically: required reviewers, who must approve before the job runs, on top of whatever the pull request already required. Prevent self-review, so the person who wrote the change can't also be the one who approves the deploy. And an optional wait timer.
The wait timer is a pause, not a queue
The wait timer holds a deploy for a set number of seconds, free on public repos, before it's allowed to run.
It's not there to throttle traffic or batch anything. It's a deliberate "are you sure" moment, a window to notice something's wrong and stop it before it ships. If a deploy touches data in a way that's hard to undo, a minute of dead air before it fires is a real, if blunt, safety net.
It's a strange feature to reach for often. Most teams won't need it. The ones who do are usually the ones where a bad deploy is expensive enough that even a one-minute pause is worth having.
Concurrency control: knowing which deploy actually wins
Two commits land seconds apart. Both trigger a deploy. Both start running.
They don't necessarily finish in the order they started. One request can lag, hang, or just run on a slower machine, and the second deploy can finish before the first one does. Without anything managing that, whichever one finishes last is what's actually live, and that's not always the one you expect.
concurrency:
group: deploy-production
cancel-in-progress: truecancel-in-progress: true means a new deploy cancels whatever's still running in the same group, so only the most recent commit's deploy actually completes. It also directly saves money: a company running a workflow that provisions over a hundred VMs per run doesn't want two of those runs alive at once just because someone pushed twice in a row.
When cancelling mid-run is not safe
Cancelling a job in progress isn't free of risk. A deploy cancelled halfway through a file sync can leave half the new files live and half the old ones still there. That's a broken site, not a rolled-back one.
The rule that actually matters: cancellation is safe when the work is idempotent, re-running it from scratch produces the same correct result, and unsafe when the work has side effects partway through that a second run won't cleanly undo or repeat. A full aws s3 sync restarted from zero is idempotent. A multi-step migration that writes to a database as it goes is not.
This is the exact same idea behind why a database transaction either fully commits or fully rolls back: partial completion is often worse than either finishing or never starting.
The Essentials
- An environment is a name plus scoped secrets plus rules, attached to a job with one
environment:key. Required reviewers, self-review prevention, and a wait timer all live there. - The wait timer is a pause to catch a mistake before it ships, not a queue or a rate limiter.
cancel-in-progresskeeps only the most recent deploy alive, which also saves real money on expensive pipelines.- Cancelling mid-run is only safe when the work is idempotent. A partial file sync or a partial migration can leave things in a worse state than either finishing or never starting.
Further Reading and Watching
Keep reading