App performance and reliability

How to catch performance regressions before a release

Set a P90 budget per metric, compare each release candidate against the last production release, and block the rollout when a budget is exceeded.

12 min read

Quick answer

To catch performance regressions before a release, set a P90 budget for each metric you care about, compare every release candidate against the last production release on real devices, and block the release when a budget is exceeded. Compare per app version for native builds and per update for over-the-air JavaScript updates. Ship to a small percentage first so the comparison has data before everyone gets the change.

What a missed performance regression looks like

The regression shipped on Tuesday. A dependency bump added 400 milliseconds to bundle load, and the home feed started fetching before the auth token was ready, so every cold start ran a request that failed and retried. Nobody noticed. The dashboard is a line chart of P90 time to interactive over the last 30 days, and Tuesday looks like a small bump, the kind you see every week.

By Friday, adoption had passed 60% and the line was no longer a bump. Support had eleven tickets that said "slow since the update". The fix took an afternoon. Finding it took three days, most of them spent arguing about whether anything had changed.

The goal of this guide is to turn those three days into a failed check on Tuesday afternoon. You need a budget per metric and a comparison that isolates the release, plus a rollout shape that gives you data before the whole user base has the change.

How to set a performance budget

A target is a number you would like to hit. A budget is a number you refuse to exceed, and a release that exceeds it does not ship.

Set one budget per metric, at P90, from production data, not from a profiler run on your own phone. If you have no production numbers yet, start from the published thresholds and tighten after two weeks of data.

MetricPublished thresholdSource
Cold start5 s or more is slow (Android vitals); under 1.5 s recommended (Expo)Android vitals, Expo metrics reference
Warm start2 s or more is slow (Android vitals); under 0.5 s recommended (Expo)Android vitals, Expo
Bundle load timeUnder 0.3 s recommended (Expo, React Native only)Expo
Time to first renderUnder 2 s including cold launch (Expo)Expo
Time to interactiveUnder 3 s including cold launch (Expo). Firebase's default alert is 5 s at P90 for app startExpo, Firebase alerts
Frozen framesMore than 0.1% of frames over 700 ms is a frozen-frame screen (Firebase)Firebase screen traces

Budget the relative change too. A cold start of 1.4 seconds is under a 1.5 second budget. If the previous release was 0.9 seconds, you just shipped a 55% regression and your gate said nothing. Set an absolute budget and a relative one, for example "P90 TTI under 3 seconds and no more than 10% worse than the previous production release". The relative check catches regressions early. The absolute one stops the baseline drifting up one release at a time.

Budget the tail, not the middle. The median moves slowly and hides the users who complain. Android vitals, Firebase's default alert and the EAS Observe CLI all report P90 (the CLI takes --stat p90), so use it. Check P99 when you are looking for a specific class of device.

How to compare a new release against production

The Tuesday regression was invisible because the chart was indexed by time. On Tuesday, 5% of users had the new version, so the P90 across all users moved 5% of the way toward the regressed value. Time-indexed charts understate a regression until adoption is high, which is when it is too late.

Index by release instead. Put the previous production release in one column and the new release in the other, same metric, same percentile, same platform. The comparison is valid at 5% adoption as long as the new release has enough sessions to compute a stable P90; a few hundred is a reasonable floor. Firebase Performance Monitoring alerts will not fire on fewer than 100 samples in an hour, for the same reason.

The platforms already do part of this. Xcode Organizer shows launch time per app version at P50 and P90 and compares the current release with a previous one when you click its bar. Android vitals breaks startup down by version. Sentry Release Health lists crash-free sessions, crash-free users and adoption per release.

The gap for React Native teams is the JavaScript update. A team that ships over-the-air (OTA) JavaScript updates can change what users run several times a week with no new app version, and a per-version comparison lumps all of those updates into one column. You need a second dimension, the update ID, and the comparison should run at whichever level changed.

Three more dimensions matter when you read a comparison:

  • Platform. A regression on one platform only is a native module or a platform-specific branch. Never compare a blended number.
  • Device class and OS version. A staged rollout that lands in one region first can give the new release more sessions from older devices than the previous release had. Filter to the same cohort before you decide.
  • Environmental flags. Low-power mode, thermal state and network type all move startup numbers. EAS Observe attaches these to every TTI event as automatic event parameters, so you can filter them out.

How a staged rollout gives the comparison its data

A comparison needs data from the new release, and that data needs real users running it. That is the case for staged rollouts, independent of the safety argument.

For native builds the staging step already exists: TestFlight and Google Play's internal or closed testing tracks put the new build on real devices before store release. Collect metrics from those builds, tagged with their environment, so the comparison can run before you press release.

For OTA JavaScript updates, publish the update to a small percentage of users, wait for enough sessions, run the comparison, and then either complete the rollout or revert. Expo's production playbook for OTA updates describes the same loop: start at something like 10%, watch the signals, expand in steps.

OTA updates in React Native replace the JavaScript bundle only. Native code does not change, and the update path does not bypass store review; App Store Review Guideline 2.5.2 permits interpreted code that does not change the app's primary purpose. The staged rollout and the revert make that path safe, and they are why the comparison in this guide can run at 10% instead of 100%.

How to automate the regression check

The manual version of this guide is: open the dashboard the day after a release, pick the new release and the previous one, read two numbers, decide. That works for a team that releases every two weeks and never forgets. For everyone else the check has to run on its own.

A working gate has four steps.

  1. Publish the new release to a fraction of users. Native: internal testing track. Over-the-air: a rollout percentage.
  2. Wait for sessions. An hour is often enough for a large app; a day for a small one. Make the wait a schedule, not a sleep.
  3. Pull P90 for the new release and the previous one in machine-readable form and compare against the budget. Fail the job if either the absolute or the relative budget is exceeded.
  4. Act on the result. Pass: complete the rollout, or approve the store release. Fail: post to the release channel and hold. Revert if the regression is severe.

Step 3 is the one most teams have never had, because their metrics tool has no CLI. Firebase Performance Monitoring gives you alerts on thresholds but not a command you can call from CI. Xcode Organizer is a GUI. Only a tool with a CLI or an API can be wired into a gate.

Where Expo fits

EAS Observe collects startup, bundle load, first render, time to interactive and update download time from production devices, and the Observe dashboard groups every metric by app version, native build and EAS Update update ID. EAS Workflows runs the release pipeline, including staged rollouts and approval steps. Together they cover all four steps above.

The comparison

The App startup cards in the Observe dashboard show the metric for the latest and the previous release, and every chart carries a release marker for each native build and each update. The Release filter narrows to one app version, build or update. That is the manual check.

The same numbers come from the EAS CLI, grouped by app version, with a separate table per platform:

The --json output includes the update IDs under each version as an array, so a script can tell which updates shipped inside a given app version. To look at one update on its own, filter the individual samples or the per-route data by update ID:

The CLI docs list "gate a script or CI job on a metric" as a supported workflow, with --json --non-interactive as the flags for it. Some data is available only on certain plans; the command fails with an upgrade message when your plan does not include it.

The rollout

An EAS Workflows file can publish a JavaScript update to a percentage of users, hold for a human, and then complete the rollout. The file below is the documented pattern, unchanged:

Only one update can be rolled out on a branch at a time, and a rollout in progress must be completed or reverted before the next update to the same runtime version can be published. To back out, eas update:revert-update-rollout republishes the previous update so every client returns to it.

The scheduled check

A second workflow runs on a schedule, pulls the summary as JSON, compares it against the budget in a script you own, and posts to Slack when the check fails. Scheduled workflows run from the default branch in GMT and may be delayed at the top of the hour.

The script is yours because the budget is yours. It reads the newest version with at least a few hundred events and the version before it, and exits non-zero if P90 is over the absolute budget or more than your chosen percentage above the previous release. The after and if: ${{ failure() }} pattern is documented in the Workflows control flow syntax. To run the same check from GitHub Actions, authenticate with an EXPO_TOKEN secret and run the same npx eas-cli command, as described in Trigger builds from CI and How to integrate EAS Workflows with GitHub Actions.

Limitations. EAS Observe needs Expo SDK 55 or later and a development or production build; it does not run in Expo Go. It does not fail your workflow on its own; the budget comparison is a script you write against the JSON output. eas observe:metrics-summary groups by app version, not by update; per-update filtering is on observe:metrics, observe:routes and observe:events. Time to interactive is only recorded where you call markInteractive(). Metric data is retained for a minimum of 60 days. EAS Observe includes 100,000 events per month on the Free plan and 500,000 on paid plans, then usage-based pricing; EAS Workflows includes up to 60 minutes on the Free plan and is usage-based on paid plans. Scheduled workflows may be delayed or, in rare cases, skipped or run twice, so the check must be safe to repeat.

Release gate at a glance

StepNative buildOver-the-air JavaScript update
Stage the new releaseTestFlight or Play internal trackupdate job with rollout_percentage: 10
Wait for sessionsHours to a dayHours to a day
Compareeas observe:metrics-summary --json per app version, or Xcode Organizer per versioneas observe:metrics --update-id and eas observe:routes --update-id
PassApprove store releaserequire-approval then update-rollout to 100
FailHold the store releaseeas update:revert-update-rollout

Next step

Pick one metric, time to interactive, and write down its P90 for your current production release today. That number is your first budget. If your app is on Expo, the EAS Observe getting started guide gets the number into the CLI, and the update rollout job gives you the staged release to compare against it.

Verified on 12 September 2026.

Keep reading

Frequently Asked Questions