A code error can bring a system down shortly after going live, making every software release a potential risk. Canary Deployment helps reduce this risk by gradually rolling out a new version to a small group of users before expanding it to the wider user base. But what is Canary Deployment, and how does it help maintain system stability throughout the CI/CD process? Let’s explore how this deployment strategy works and why it is widely used for safer software releases.
What Is a Canary Deployment?
A Canary Deployment is a software release strategy that runs a new version alongside the stable version in production. Instead of releasing the new version to all users at once, teams first route a small share of real traffic or production servers to it.
The two versions are then compared using predefined performance and reliability signals. If the new version performs as expected, it gradually receives more traffic. If issues occur, the deployment is then rolled back.
For example, a service is running on 10 pods. During a Canary Deployment, 9 pods continue running the current version while 1 pod runs the new version. The new pod receives roughly 10% of requests. Its error rate and latency are then compared with those of the other 9 pods before the team decides whether to increase its traffic.
The term “canary” comes from the phrase “canary in a coal mine”, where canaries served as early warnings of toxic gases. Similarly, in software deployment, if a new version shows problems, traffic can be quickly shifted back to the stable version, reducing the risk of a new release.
Canary Release vs. Canary Deployment
These terms are often used interchangeably, but they refer to different scopes and environments. Understanding the distinction helps teams define exactly what they are validating during a release.
| Approach | What it means | What it validates |
| Canary Release | Exposes a new feature to a small group of users first. | Primarily how users respond to the new feature. |
| Canary Deployment | Applies the canary approach to the entire change, including code, configuration, and infrastructure. | System stability and behavior under real production load. |
A key difference is that a staging server never serves production traffic, while a canary server does. Once the rollout is complete, the canary also remains part of the production fleet.
How a canary deployment works: Deploy, split, judge, promote
Every canary deployment strategy follows the same four-step loop, and teams differ mainly in how they split traffic and who they send to the canary.
- Deploy the new version beside the stable one, with no users routed to it yet.
- Split a small share of traffic to the new version.
- Judge the canary against the stable version over a fixed time window.
- Promote it to the next stage, or roll back by sending all traffic to stable.

There are 2 options for the canary deployment when splitting traffic to the new version:
| Two-step canary | Multi-stage (linear) canary | |
| How traffic moves | A small group gets the new version first, then everyone once it passes its checks | Traffic rises in equal increments until the new version handles all of it |
| Example sequence | 10% → 100% | 5% → 25% → 50% → 100% |
| Best for | Small, low-risk changes, such as a text fix or a minor configuration tweak | Riskier releases, such as a change to checkout logic or a new database query |
| Why it works | A single check tells you most of what you need to know | Each step exposes more users only after the previous share has passed its gate |
| Impact if it fails | Limited to the first group, as long as the check catches the problem | Limited to the current stage, so a failure at 25% reaches only a quarter of users |
| Trade-off | Fast, but with only one checkpoint | Safer, but every extra stage adds a bake period and needs enough traffic for a trustworthy result |
In practice, the choice comes down to how much a failed release would cost you and how much time you can give the rollout.
Beyond the number of steps, you also choose the infrastructure where the canary runs:
- In a rolling canary, you upgrade a subset of your existing servers and leave the rest on the stable version.
- In a side-by-side canary, you stand up a separate copy of the environment and split traffic at a load balancer or service mesh.
In short, rolling costs less, while side-by-side isolates the canary more cleanly.
The group you send to the canary also shapes what problems you can catch. Martin Fowler describes several ways to choose that group, such as a random sample of users or internal employees first.
- A random sample represents your whole user base, so it catches regressions that affect everyone.
- Employees, on the other hand, tolerate rough edges and report problems quickly, without putting real customers at risk.
For backend services with no direct users, it often makes more sense to send the canary to a few servers rather than a group of users.
Put together, these choices make up the last stage of the SDLC (software development life cycle), the point where tested code meets real traffic for the first time.
Canary vs blue-green vs rolling vs A/B testing vs feature flags
Canary deployment is most often compared with blue-green deployment. Rolling updates, A/B testing, and feature flags also enter the discussion, since each of them splits traffic in some way, which is why the five get confused.
The real difference lies in the question each one answers.
| Canary | Blue-green | Rolling update | A/B testing | Feature flags | |
| Question answered | Is this build safe? | Can we switch instantly? | Can we replace instances gradually? | Which variant performs better? | Who sees this feature? |
| Exposure pattern | Small share, then staged | 0% to 100% at once | Instance by instance | Split by user attribute | Per user, cohort, or flag |
| Rollback speed | Fast, reroute to stable | Instant, switch back | Slow, roll forward or back | Not a rollback tool | Instant, turn flag off |
| Extra infrastructure | Low to moderate | High, two full environments | None | Low | Low, plus a flag service |
| Best fit | Unknown behavior under real load | Fast cutover, simple rollback | Default for stateless services | Product experiments | Decoupling deploy from release |
| Main risk | Weak or missing signals | Cost, database sync | Little control over exposure | Mistaken for a safety check | Flag debt, config drift |
In practice, most teams combine these strategies rather than pick just one. They use rolling or blue-green deployments for infrastructure, a canary for each new build, and feature flags to turn on features once the build is stable. For a broader look, see our guide to deployment strategies.
Choose a canary deployment when your biggest risk is how the new version behaves in production. It works best when you have enough traffic to judge the result.
How to Evaluate a Canary Deployment: Key Metrics for Promotion
Once you’ve chosen a canary deployment, everything depends on how you judge it. A canary promotion is a statistical decision, and you can only trust it if four things hold:
- Canary deployment metrics. You compare the canary’s metrics against the stable version over the same time window.
- Canary promotion criteria. You fix the pass and fail rules before the rollout starts.
- Missing canary metrics. You treat missing data as a failed check, not a pass.
- Reliable canary analysis. You wait until the canary has served enough requests for the comparison to mean something.
If any step is skipped, the canary becomes a slower way of deploying to everyone.
1. Compare canary deployment metrics against a stable baseline
Absolute thresholds hide regressions. A gate of “fail above 1% errors” lets through a canary that pushes errors from 0.1% to 0.9%, which is a ninefold regression. The same gate fires during an unrelated incident that hits both versions equally.
A better approach is to measure the gap between canary and stable over the same time window. The most useful signals are these:
- Success and errors. Track success rate and error rate, split by HTTP status class, because a jump in 5xx responses is the clearest regression signal.
- Tail latency. Compare p95 and p99 as well as the median, since regressions often show up in the tail first.
- Resource saturation. Watch CPU and memory on canary instances against stable ones to catch leaks before they cause errors.
- Business outcome. Add one metric such as completed checkouts, because some regressions return a clean 200 and still break the product.
An absolute floor, such as a minimum success rate, still makes sense as a backstop. Even so, the main test should be whether the canary is worse than stable by more than a margin you chose in advance.
2. Set canary promotion criteria before you deploy
Knowing what to compare is only half the job. The other half is fixing the thresholds before you see any results, because criteria chosen during a rollout tend to drift toward “looks fine.”
To prevent that, write them into the rollout configuration, so your CI/CD pipeline makes the call rather than whoever happens to be watching the dashboard.
Most tools support this directly. Argo Rollouts runs metric queries at each step through an AnalysisTemplate, and Google Cloud Deploy lets you attach an analysis job to each phase and advance the rollout automatically when it passes.
For each outcome, decide in advance what happens next, including whether a person has to approve the next step.
3. Treat missing canary metrics as a failed check
Even well-chosen criteria fail when the data behind them is missing. A canary that reports no data has not passed. Usually the telemetry is broken, through a mislabelled pod or a query that matches nothing. Sometimes the canary never received traffic at all. Either way, the rollout has no evidence the new version works.
The danger is that a controller which reads an empty result as “healthy” will promote a blind canary. To avoid that, make the gate fail closed:
- Minimum request count. Require a set number of canary requests before any other check runs, so an idle canary can’t pass.
- Empty-query failure. Fail the check when a query returns no data instead of skipping it.
- Absence alerts. Alert when canary metrics stop arriving, separately from error thresholds.
- Default behavior. Check how your tool treats empty or errored queries, and override it if it passes by default.
With this rule in place, a canary that emits nothing stops at 5% instead of reaching everyone.
4. Get enough traffic for reliable canary analysis
Data can also be present but too thin to trust, because low traffic makes canary results unreliable. Suppose a canary handles 40 requests with zero errors. By the statistical rule of three, its true error rate could still be as high as 7.5% at 95% confidence.
For example, catching a regression from a 0.5% to a 1.0% error rate needs roughly 4,700 requests on each version, using a standard two-proportion sample-size calculation at 95% confidence and 80% power. A service handling 1,000 requests a minute sends about 1,000 requests to a 5% canary over a 20-minute stage. That’s enough to spot a severe regression, but too few for a subtle one.
For lower-traffic services, you have a few options:
- Longer bake time. Hold each stage until the canary has accumulated enough requests.
- Larger first share. Start at 20% instead of 5% so the canary reaches a useful sample sooner.
- Synthetic load. Send replayed or generated traffic to the canary to fill the gap.
- Request-count gates. Advance a stage after N requests instead of N minutes.
Below a certain volume, however, a canary adds little over a well-tested deploy.
3 things to do to avoid breaking canary deployments
Good metrics catch a bad build, but they can’t protect you from problems built into the release itself. Most canary deployment failures start there, because the old and new versions share a database and users, and the rollback path is rarely tested under pressure.
Keep database changes backward compatible (expand, migrate, contract)
During a canary deployment, the old and new versions read and write the same database, which makes the schema the easiest place to break a zero-downtime deployment. A migration that renames a column or changes a type breaks the stable version as soon as the canary applies it, and rolling back the code won’t undo the schema.
The safe pattern runs in three phases:
- Expand. Add new columns or tables in a form both versions tolerate.
- Migrate. Switch writes to the new structure and backfill existing data.
- Contract. Remove the old structures only after the new version has reached 100% and stayed there.
In other words, keep destructive migrations out of the release you are canarying. They belong in a later, separate change once nothing depends on the old structure.
Pin users to one version to avoid version skew
Shared users cause a similar problem to a shared database. A user whose requests alternate between versions can see the interface change between clicks, or trigger an API mismatch between a new frontend and an old backend.
To prevent this, keep each user on one version for the duration of the canary. Sticky sessions do this, as does header- or cookie-based routing in Istio or the Kubernetes Gateway API. Assigning whole cohorts to the canary works too.
The same holds between services. When two microservices are canarying at once, each one’s API has to stay compatible with both versions of the other.
Rehearse the rollback before you need it
Even with compatible data and pinned users, a canary is only as safe as its way back. Rerouting traffic is only half of a rollback, and the other half is dealing with what the canary already wrote, such as queued messages or rows in a new format.
For that reason, run rollback drills on a schedule and time how long a full rollback takes. Then document who triggers it and which signals should prompt it.
Feature flags also help here. When a problem is tied to one feature, turning off its flag fixes it without rolling back the whole build.
Implement a canary deployment on Kubernetes and the cloud
With the rules in place, the next step is choosing tools that enforce them. The table below compares the main tools for running a canary deployment on Kubernetes and the cloud:
| Option | Traffic precision | Automated analysis | Rollback | Best fit |
| Native Kubernetes | Approximate, by replica ratio | None built in | Manual scale-down | Getting started, simple services |
| Argo Rollouts | Precise with a mesh, ingress, or Gateway API | AnalysisTemplate with metric providers | Automatic abort | Kubernetes teams using Argo CD |
| Flagger | Precise via Istio, Linkerd, or ingress | Built-in metric checks | Automatic | Teams already on a service mesh |
| Google Cloud Deploy | Percentage phases on GKE and Cloud Run | Deploy analysis per phase | Roll back release | Google Cloud workloads |
| AWS CodeDeploy | Canary and linear configs for Lambda and ECS | CloudWatch alarms | Automatic on alarm | AWS-native services |
| Spinnaker | Depends on target platform | Automated canary analysis (Kayenta) | Automatic | Multi-cloud estates |
| Octopus Deploy | By deployment target, not percentage | Scripts or manual approval | Redeploy to targets | VM and on-prem fleets |
If you already use GitOps, an Argo CD canary deployment usually pairs Argo CD for sync with Argo Rollouts for the traffic steps. A minimal Rollout strategy looks like this:

Note that without a trafficRouting block, Argo Rollouts approximates the weights through replica counts. Add a mesh or ingress integration when the exact percentage matters.
Outside Kubernetes, managed cloud services cover the same ground.
- For a canary deployment on AWS, CodeDeploy offers predefined canary and linear configurations for Lambda and ECS. They shift traffic in timed increments and roll back when a CloudWatch alarm fires.
- On Google Cloud, Cloud Deploy handles GKE and Cloud Run targets, including parallel canaries across several regions.
Whichever tool you pick, configure the analysis gate to fail when the canary produces no data, and check what the tool does by default. The tool matters less than the gate you put in front of it, so choose the one that fits your platform and spend the effort on the promotion criteria.
If your team still promotes canaries by watching a dashboard, our cloud and DevSecOps team builds CI/CD pipelines with checks enforced at every gate. We also wire in the metrics and SLOs those gates read, and hand over the release and rollback runbooks in your own repositories.
Run canary deployments beyond web apps: Device fleets and mission-critical systems
If web traffic cannot be split by percentage, use cohort-based canary deployment instead. Each cohort, typically a single site or device group, must pass field telemetry checks and operator sign-off before the next wave begins.
Many mission-critical systems do not run behind a load balancer. For example, Eastgate Software spent six years building BagOS with Siemens Logistics, a standardized baggage-handling platform deployed across international airports and integrated with Siemens hardware and PLCs for real-time bag routing.
For systems like this, sending 5% of traffic to a new version is not meaningful. Each airport becomes a natural cohort, and release health is measured through operational outcomes such as baggage routing accuracy rather than HTTP errors.
Three patterns adapt the canary loop in this case:
- Device fleets. Roll out to one device, then one site, then the full fleet. A remote device can’t be rerouted, so updates need an A/B boot scheme that falls back automatically when health checks fail.
- On-prem sites. Ship to one customer site first and agree in advance on what it must report before the next site gets the update. Evidence comes from logs and operator sign-off rather than a live dashboard.
- Safety-critical systems. Run the new version in shadow mode, processing live inputs without acting on them. Compare its outputs with the stable version before any cutover.
Shadow and parallel running is where our work in regulated environments comes in. We ran old and new versions in parallel during the phased migration of a traffic control application for an ITS client, keeping 100% backward compatibility with existing workflows. The same approach kept a medical device data platform running continuously, as its regulators required, while the new platform took over.
In regulated programs, someone outside the build team should hold that final approval. Our independent QA and testing team signs off each release on evidence, with a trace from each requirement to its test and result that auditors can follow.
To see how these patterns fit together, consider a control-software update on a platform running at 12 sites. We’d structure the canary in four waves, and the plan below is an illustrative example rather than a record of a specific project.
- Wave 0. Run shadow mode at one low-traffic site for 72 hours. Promote only if the new version’s outputs match stable on at least 99.9% of decisions.
- Wave 1. Go live at that site during a staffed maintenance window, with a one-command rollback to the previous version on local hardware.
- Wave 2. Move to three sites of different sizes, to surface configuration differences early.
- Wave 3. Roll out to the remaining eight sites.
Each wave’s gate is zero safety-relevant alarms plus operator sign-off. Missing telemetry from a site counts as a failed gate, the same rule that applies to a web canary.
Know when canary deployment is the wrong choice
Cohorts and shadow mode stretch canary deployment a long way, but they don’t make it the right choice for every release.
In the six situations below, skip the canary deployment or pair it with another strategy:
- Low traffic. A handful of requests per stage can’t reveal a regression, so use blue-green with thorough pre-production testing instead.
- Irreversible changes. A data migration that can’t be split into expand and contract phases won’t roll back with the code, so canary the code and plan the migration separately.
- Client-installed software. Without server-side control, staged app-store rollouts or release channels do this job better.
- Missing observability. If you can’t tell canary metrics from stable metrics, build that labelling before you rely on a canary.
- All-or-nothing changes. A protocol cutover that both sides must adopt at once fits blue-green or a planned cutover.
- Regulated releases. Staged cohorts still work, but the canary becomes a formal pilot phase with documented approval at each wave.
If none of these apply and production behavior is your biggest unknown, a canary is usually the safest way to ship.
Final thoughts
Use a canary when production behavior is the risk you can’t test away, and only when you trust the gate that promotes it. That judgement holds whether the canary is 5% of API traffic or one airport out of twelve.
Start with the missing-metrics rule, because it’s the cheapest fix and it closes the gap that lets a blind canary reach everyone. Then write down, before your next rollout, what result would make you roll back. If the team can’t answer that, the canary is just a slow deploy.
If you’re designing release processes for systems where a bad deploy has operational consequences, talk to our engineers about how we approach staged rollouts on mission-critical platforms.


