Releases should not happen by accident, they should happen with a plan and on schedule. I have worked under two release models that are imperfect in opposite directions and this is my best stab at how to do it better.
A DIP switch is a feature flag in hardware. It allows hardware to ship before its use is clear and the same hardware to fulfil different roles for different customers. These are the exact traits I was looking for when outlining the release process for kaniko.
Demand driven releases
In a demand driven model a release is cut because a customer asked for one. A bug was found in the field, it was fixed and now the customer wants an update with the bugfix. There is nothing wrong with that at small scale and it provides high velocity initially. The strain only shows up as you grow.
Over time a release for the customer turns into a product for the customer. The customer asks for a small deviation and you are going to release it specifically for that customer anyway, so it's no extra effort to just bake that setting in. Before you realise it you have a product catalogue that goes into the thousands and your support team loses oversight over what bugs are still in vogue and which ones have gone out of fashion already.
Of course most of those products will never get an update, because a customer who never complains never triggers a release. But once they do, a fix is cut from whatever the code looks like that day. It arrives bundled with changes nobody asked for and inadvertently triggers the next release. The software never amortises, because the effort scales with the number of customers rather than with the number of features.
Scheduled releases
In a calendar driven model the date and type of a release is defined in advance. Semver runs backwards here, the kind of work no longer decides the version, the version decides what kind of work is allowed.
On paper this looks much better, because now the chaos of the demand driven cycle is tamed and we have a clear plan for when each feature ships. But this also means that customers have to wait for the next release cycle to test drive their new feature and if a single change blocks their upgrade path, they can sit on the old version for years. Although LTS branches can keep that position viable, that puts a lot of strain on release management again.
Worse still, if a customer cannot be satisfied with the singular product, you can't just cut a dedicated release, or you risk derailing your entire plan. Where the first model said yes to everything and drowned in variants, this one can only stay afloat by saying no.
Two clocks
This reveals the fundamental tension that is baked into release management. On the one hand you want the features to ship as fast as possible to the customers that are waiting on it, on the other hand you want to be able to guarantee stability to everyone else. You want to be able to create bespoke solutions but you want to keep one generic product.
What bites customers that update frequently is the rapid and unpredictable activation of experimental features. Semver tries to box changes by how safe they are to take, but that is a crude signal. The customers that are stuck on LTS branches are usually stuck because a single change prevented them from upgrading initially and the longer they wait the riskier the upgrade becomes. The solution to both is to start thinking of releases in terms of behaviours and not code.
"Ship the code - release the feature" means that you resolve that tension within your codebase. Your codebase becomes bleeding edge and LTS simultaneously. You merge any changes as soon as possible, but you hide their activation behind a featureflag. Now you can ship the codebase with confidence that nothing changed. The feature sits in the binary like a sleeper cell, and the customers eager to test it can activate it effortlessly, whilst everyone else gets a stable experience. Your release cadence no longer dictates what code can be merged when, but simply says when the featureflags flip their default and get deprecated.
Lifecycle of a featureflag
In kaniko we adopted featureflags aggressively. Every behavioural change, even an obvious bugfix, is hidden behind one, because there is always a chance it breaks someone's setup.
FF_KANIKO_OCI_STAGES is a good example. It moved the storage type of intermediate stages from docker tarballs to an OCI layout, which can change the media type of the image you end up with, so it merged in October and shipped four days later switched off. For five months anyone who wanted the newer format turned it on themselves. In March the minor release made it the default, and in June the flag was deleted and the OCI layout became unconditional.
Enabling the feature is unspectacular. v1.28.0 activated nine flags at once, all of them dormant for months and already running in production for everyone who had opted in, which is a far better test than anything we could run ourselves.
The escape hatch then gets used for things nobody planned for. In June a dependency bump in v1.27.6 deadlocked builds that used --cache with metadata-only commands. The bug was in go-containerregistry rather than in us, and it only fired on the path that writes an OCI layout. FF_KANIKO_OCI_STAGES still had the previous implementation behind it, the docker tarball writer, which writes layers sequentially and cannot leak the limiter. Setting the flag back to false picked that backend instead, and the real fix shipped three weeks later in v1.28.0, the same release that deleted the flag.
Two people in that issue thread took different routes, one pinned the previous release and waited, the other set the flag, kept every other fix in the release, and carried on. That is the whole difference between a version and a behaviour as the unit you are allowed to refuse.
At osscontainertools we fully endorse the messiness of a codebase riddled with featureflags. We test one combination of them, and that only works because they have a defined expiration date. We ship the code on one clock and release the features on another. Which of the two you live on is yours to decide, not ours.