🚀 New: chi (χ) — an open-source autoresearch harness for fleets of LLM coding agents. Read the announcement.

Firebase's Kill Switches Shipped Globally, and a Two-Line Cleanup Crashed iOS Apps

Firebase's September 28 iOS SDK outage came from removing a remote config flag while a pointer to it stayed in the SDK. Notes on why kill-switch config is a global deployment, why the SDK crashed on a nil key instead of ignoring it, and what to validate in your own flag cleanup pipeline.

Contents

Firebase’s postmortem for the September 28 iOS outage contains one sentence that explains most of it: the flags the Google Analytics for Firebase SDK uses as kill switches are themselves “released globally”. A kill switch is supposed to be the safe, fast path for turning off a bad feature. Here the path that carries it had no staging, so the bad value went out to a wide population of apps.

The mechanics are in Firebase’s write-up and the pinned GitHub issue. Everything below comes from those two, plus the issue’s user-supplied crash logs. User reports are not Firebase’s numbers, and I flag them as such.

What actually broke

The SDK accumulates remote configuration flags that are no longer needed, and they are marked for cleanup. On September 28 at 17:38 PDT one was removed. Firebase calls it a 2-line change. A pointer to that flag stayed in the SDK. When a client fetched the payload and went looking for the missing flag, the SDK hit a fatal error. The postmortem later names the root cause: the SDK “missed validating that a flag’s name was not nil”.

The reporter who opened the issue gives the client side of the same failure. The crash is NSInvalidArgumentException, key cannot be nil from setObject:forKeyedSubscript:, and in all 42 crashes they checked it happened right after the SDK received the HTTP 200 response from its sdk-exp endpoint. The stack runs through APMETaskManager handleFetchingExperimentsResponse:data:error: into APMEExperiment copyWithZone:. The reporter notes the crashing thread shows only _dispatch_call_block_and_release, because the dictionary wrapper writes through dispatch_async and drops the caller’s frames.

That detail matters for a practical reason. The failure is on an SDK worker queue, in code the host app does not call directly. No app-level guard around your own analytics calls would have helped. The only defense available to app teams was not being on that SDK path, which for most apps is not a choice.

Why a 2h 11m outage did not become a crash loop

The postmortem says the crash was one-time per active user, “after which the SDK fell back to the previous config”, and that Firebase did not observe devices entering a permanent crash loop. So the client does have a recovery path, and it is the reason the incident is measured in a single crash per launch cycle rather than an unrecoverable state.

It is also why reporting kept looking bad after the rollback finished at 19:52. Crash reports are queued locally and uploaded on the next successful open, so developers saw elevated errors “for many hours”. Firebase’s pinned comment says they initially believed some instances were still crashing and now think those were delayed reports. If you were watching a crash dashboard that night, the lag was a property of the reporting pipeline, not evidence that the fix had failed.

The two integrity checks that were missing

The postmortem’s commitments are specific enough to read as a spec for the two missing checks.

Server side, they are adding validation to the config pipeline and pre-submit tests so that “if a flag cleanup or schema change results in broken dependencies or orphaned identifiers, the pipeline will automatically block the release”. The failure here was a referential integrity problem. The config said a flag did not exist, and code elsewhere still referred to it. “Existing tests passed on the change”, per the report, which is what you would expect if no test asserted that every flag the code reads is still defined.

Client side, the patch they promise within a week validates every incoming payload strictly. A flag that is corrupt or missing a required field will be ignored individually, and the SDK falls back to cached defaults. That is the narrower and better contract: one bad entry disables itself instead of taking down the parser. Firebase also says it is auditing its other iOS and Android SDKs for similar parsing weaknesses. That audit is the part I would watch, since a nil-key assumption in one config parser is rarely unique to it.

Dashboards were green for a reason

Both the Firebase and Google Ads status dashboards stayed green throughout. Firebase says they rely mainly on server-side health metrics and so did not register client-side crashes. The server side was healthy: it was serving a successful, well-formed HTTP 200. Updating the dashboards also took manual intervention that “took several hours”, partly because GA4F has no entry of its own on the Firebase dashboard and links to the Ads one. Firebase staff posted updates on the GitHub issue instead.

If your own service pushes config or code to clients you do not control, ask what your status signal looks like when the server is fine and every client dies after a successful response. A 200 is not evidence of a working release when the failure is in the consumer.

What I would take from this

Firebase’s fourth commitment is better monitoring and rollout phasing for configuration deployments. That is the part that applies beyond Firebase. Config that gates behavior, especially kill switches, tends to get a lighter release process than code because it feels like data. The postmortem is a case where it behaved like code with none of the staging.

Two checks are cheap to add to a flag lifecycle of your own:

  • When a flag is deleted from the config source, fail the change if any code path still reads that key. A grep for the key name in CI catches the case in this incident, and a typed flag registry catches it earlier.
  • In any client that parses remote config, treat each entry as independently skippable. A missing key, nil name or unknown type should log and fall back to the cached default for that entry, not throw.

The postmortem leaves some things unsaid. It does not give the number of affected apps or devices, and it does not say which tests passed or what they covered. For now the only counts are user-reported in the GitHub thread, for example one commenter’s roughly 20,000 crashes and another’s 110,000 events. They are unverified and cannot be summed into a blast radius. Firebase has the real figures in its own logs.

Sources