Web & Mobile ¡ 5 minute read
Feature Flags: Separating Deployment From Release
Feature flags separate deploying code from enabling behaviour, which makes deployment routine and release a controlled decision. They also accumulate: flags that outlive their purpose become conditional paths nobody understands, so removal has to be part of the practice.
Feature flags separate deploying code from enabling behaviour, which changes how a team ships more than almost any other practice. They also accumulate quietly. This guide covers both halves, drawing on FISTA Solutions' web and mobile work.
What kinds of flag are there?
Four broad types with different lifetimes and different rules.
| Type | Lifetime and purpose |
|---|---|
| Release flag | Days to weeks; remove after rollout |
| Experiment flag | Tied to the test; remove at conclusion |
| Operational flag | Indefinite; kill switches, load shedding |
| Permission flag | Indefinite; really configuration |
| Ops override | Temporary; remove after the incident |
| Stale flag | None; should not exist |
How does this change deployment?
It makes deployment boring, which is the goal.
Code behind a disabled flag is inert. Deploying it carries the risk of the deployment mechanism rather than the risk of the feature, which means teams can deploy frequently without frequent anxiety.
Release then becomes a separate, reversible decision: enable for internal users, then a percentage, then everyone, with the ability to disable instantly at any point.
Why are kill switches the highest-value flags?
Because they convert incident response from a deployment into a configuration change.
A feature misbehaving in production can be disabled in seconds by whoever is on call, without a build, a pipeline run, or a release approval. That difference is frequently the difference between a brief degradation and a prolonged incident.
Put kill switches on anything with external dependencies, anything new, and anything expensive. They are cheap to add and they are what on-call responders reach for first. See how to set up AI on-call.
How do you control flag debt?
By giving every temporary flag an owner and a removal date at creation, and by reviewing the list regularly.
Stale flags leave conditional paths in the codebase that nobody understands and combinations nobody tests. Ten flags produce more states than any test suite covers, and most of those states have never executed.
The practical discipline is a periodic review that removes flags past their date or escalates them. Without it, the flag count only rises.
Where should evaluation happen?
Server-side for anything affecting behaviour, pricing, permissions, or security, with the resolved decision passed to the client.
Client-side evaluation exposes the whole flag set to anyone who looks, which leaks unreleased features and sometimes commercial information. It also produces flicker when the value resolves after the first render.
Where the client must evaluate â offline capability, for instance â send only the flags that user needs, already resolved.
How should gradual rollout work?
Consistently per user rather than per request, so a given person gets a stable experience.
Per-request evaluation means users see the feature intermittently, which is confusing and makes any measurement meaningless. Hash the user identifier to assign a stable bucket.
Roll out in stages with a watch period at each: internal, then a small percentage, then larger, watching error rates and the metrics the feature is meant to move. See AB testing implementation.
What does this cost in testing?
Combinatorial complexity, which is the real price of flags.
Each flag doubles the theoretical state space. In practice you test the current production configuration and the configuration you are rolling out to, and accept that other combinations are untested.
That acceptance is manageable with few flags and unmanageable with many, which is the strongest argument for removing them promptly.
What are the common mistakes?
No removal date or owner. Client-side evaluation of sensitive flags. Per-request rollout assignment. No kill switches on risky features. And flag counts that only grow.
How do you test it?
Test both states of any flag being rolled out, test the kill switch actually disables the feature, and include flag configuration in whatever you use to reproduce a production issue.
The last matters: a bug reproducible only under a specific flag combination is impossible to investigate without knowing which flags were on.
What does it cost to operate?
A managed flag service carries a subscription; a simple implementation is cheap to build and grows into a maintenance burden as requirements accumulate.
The larger cost is flag debt, which is paid in comprehension and testing rather than in money.
What should you measure?
Flag count and age distribution, flags past their removal date, kill switch usage during incidents, and the proportion of deployments that required a rollback rather than a flag change.
How do flags apply to AI features?
Particularly well. A kill switch on an AI feature lets on-call disable it when quality degrades, and a percentage rollout limits exposure while evaluation catches up with production behaviour.
Flags are also how you route between models during a migration, which makes the cutover reversible. See how to run a model migration.
When is this the wrong approach?
For a small team deploying infrequently with low risk, flags add indirection for a benefit they do not need. The practice earns its cost with frequent deployment, real user impact, or features worth staging.
What should you do first?
Add a kill switch to your riskiest feature and test that it works. That single flag is usually the highest-value one in the system.
How FISTA Solutions helps
FISTA Solutions builds and operates production systems through web and mobile, AI enablement, and staff augmentation: kill switches on anything risky so incidents become configuration changes, removal dates set at creation so flag debt does not accumulate, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To scope this work, message FISTA on WhatsApp, or read CI/CD for web apps.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What do feature flags actually give you?
The ability to deploy code that is inert and enable it separately. That makes deployment a low-risk routine event and release a decision you can limit, stage, and reverse without deploying again.
02What types of flag are there?
Release flags that are removed after rollout, experiment flags tied to a test's lifetime, operational flags such as kill switches that live indefinitely, and permission flags that are really configuration. Each has different rules.
03Why are kill switches valuable?
Because they turn an incident response into a configuration change. Disabling a misbehaving feature in seconds, without a deployment, is frequently the difference between a brief degradation and an outage.
04What is flag debt?
Flags that outlived their purpose, leaving conditional paths nobody understands and combinations nobody tests. It compounds, and a codebase with a hundred stale flags has states that have never been executed together.
05Where should flags be evaluated?
Server-side for anything affecting behaviour or security, with the decision passed to the client. Client-side evaluation exposes the flag set and risks flicker as the value resolves after render.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.