Building with AI agents

How to verify AI-written mobile code before you ship

Verify AI-written mobile code by checking user outcomes against the exact build you will ship, on iOS and Android, including saved data and failure cases.

12 min read

Quick answer

To verify AI-written mobile code, write down what the change must do for the user, then check each outcome against the exact build you plan to ship. Read the diff and any new dependencies, run the automated tests, and use the installed app on iOS and Android, including failure cases. Confirm saved data on the server, not only on screen, and list what stayed untested.

Step 1: Write down what the change has to do

The agent's summary reads "Photo upload implemented. All 14 tests pass." Its screenshot shows a green "Uploaded" toast. The staging bucket is empty, and the server log has one new line: 413 Payload Too Large. The app showed success before the server answered, and every one of those 14 tests mocked the network.

Write the expected user outcome before you review the code. For an upload, success might mean the server accepted the file, attached it to the right account, and returned a record the app can still fetch after a restart.

Keep that outcome separate from what the screen shows. A screenshot proves the app drew a success state. It cannot prove the state is true.

Pair each claim with the check that actually proves it:

ClaimCheck that proves itCheck that doesn't
User can upload a fileThe server accepted the record, and an authenticated read returns itA success toast
Other users cannot read itA server-side authorization test with a second accountThe button is hidden in the UI
A failed upload can recoverA forced failure followed by a successful retryA happy-path screen recording
The result survives a restartReopen the app and fetch the saved recordIn-memory state before closing
The native integration worksA run in the native environment the feature needsA browser preview

Spell out the product decisions you don't want the agent to invent, such as what happens to a pending upload when the user logs out. Vague acceptance criteria let a plausible implementation become the spec by accident.

Step 2: Pin down the commit and build you are testing

Record the commit hash and read the actual diff. Look hardest at changes outside the requested feature, especially new dependencies and build configuration.

A new dependency can bring native code or data collection with it. A config change can add a permission or point the app at a different backend. Review those effects next to the user-facing code.

Keep the lockfile with the build. An agent that reinstalls dependencies without the project's constraints can end up testing a different set of packages from the one you release.

If your app receives JavaScript updates separately from store builds, record both the native build and the update it ran. Runtime versions tell you which updates a build can load; they say nothing about whether the app behaves correctly.

Keep a short record like this with every test run. The field names are your own choice; no Expo or vendor tool requires this format.

Replace "latest build" with a real ID everywhere it appears. If the test job checked build 412 and the release job picks up whatever is newest, the passing test says nothing about what shipped.

Step 3: Read the tests the agent wrote

A test written by the same agent can repeat the implementation's wrong assumption. Read each assertion and ask whether a plausible bug would make it fail.

For a reported regression, keep a test that fails on the old code and passes with the fix. That pair tells you far more than a test written after the agent already changed the behavior.

Check what the test mocks. A test that swaps the network layer for a canned success response can check how the UI reacts to success. It cannot check the API contract or authorization.

Run the project's existing checks as well, such as TypeScript, ESLint and Jest with React Native Testing Library, or XCTest and JUnit in a native project. They catch nearby breakage cheaply. React Native's testing overview explains what each level covers and why tests against the running app still matter.

Diff the test files on their own. An agent can treat a failing assertion as something to turn green instead of a report about the feature, so a deleted or loosened expectation needs the same review as a production change.

Add a second reviewer when the risk justifies it. Another model agreeing with the first proves little; a separate run of the app with an outcome you can observe proves more.

Step 4: Run the installed app, including the failure paths

Run the build in the environments the change touches. A browser can't show you every native permission prompt or lifecycle event.

Pick device conditions on purpose. An image feature needs real hardware with realistic memory, and a sign-in change needs the redirect back from the identity provider. For background work, interrupt the app and resume it.

Test denied permissions and expired sessions wherever they affect the feature. Force the failures with a test backend instead of waiting for a real outage, and confirm the user can recover without losing work or creating duplicates.

Upgrade from the current public version with realistic saved data. A fresh install skips the migrations and the stored credentials that existing users already have.

UI test tools such as Detox, Espresso, Maestro and XCUITest drive the installed app and record what happened. Agent toolkits listed in Expo's AI agents guide, including Callstack's agent-device and Software Mansion's Argent, let an agent tap through flows and read logs on a simulator or emulator.

For each run, keep the starting state and the result, with logs or a server-side check for anything a screen recording can't show.

Physical and virtual devices answer different questions. Android's testing fundamentals recommend choosing the environment by the behavior under test. A cloud-hosted simulator is still a simulator, and a passing iOS run tells you nothing about Android.

If the environment you need isn't available, mark the check "not tested" and name who will run it. A partial result should never turn into a whole-app pass by omission.

Step 5: Check security and data handling

Review generated code with the same security bar as any other code, and spend the time where the feature crosses a trust line.

Confirm authorization on the server. Hiding a button in the app doesn't stop another client from sending the request, so test with a second account that should be refused.

Trace where secrets and personal data go. Anything in the app bundle ships to every user, so private server keys don't belong there, and diagnostic logs shouldn't carry credentials or personal data you don't need.

Check local storage and logout. Data left in a cache can leak between accounts on a shared device if nothing enforces ownership, so test account switches, not just one developer login.

Use the OWASP MASVS to organize the review. The standard gives you a structure for mobile security checks, and passing functional tests says nothing about whether those controls are in place.

For payments, health data or anything similar, get the right specialist to look. An agent can prepare the material and list its assumptions, but the decision belongs to someone who understands the app's real risk.

Step 6: Decide whether to ship

Write a short summary a release owner can read without opening the agent conversation. Name the change and the build ID. Link the checks that ran, explain any failures and retries, and list each remaining gap and why it does or doesn't block release.

Tie the decision to that build. If anything changes after verification, work out which checks need to run again; a new binary with different native dependencies needs more than last week's screenshot.

Match the rollout plan to the change. A data migration may need compatibility work before either a rollback or a forward fix is safe, and halting distribution doesn't undo data already changed on a device or the backend.

After release, watch the outcomes the feature changes in your crash and analytics tools, such as Sentry or Firebase Crashlytics, filtered by the release you shipped. Zero crashes means little for a feature that can fail quietly without killing the app.

When something slips through, fix the check that should have caught it. Treat the verification setup like product code you maintain.

How much do you rerun after a follow-up change?

Rerun what the change can affect. A small diff can still touch something sensitive.

Take two follow-ups to a build that already passed. One changes a button label. Check that the text still fits the supported layouts and that UI tests can still find the control. If the label describes a permission or a purchase, also check that it still matches what the button does, because a changed consent statement is not an ordinary wording edit.

The other follow-up upgrades the authentication library. The screens may look identical, but the app now depends on different native or redirect behavior. Make a new build and repeat sign-in, return-to-app, session expiry and logout on each affected platform.

Decide the rerun scope from the diff and its dependency effects. Don't ask the agent that wrote the change whether it is safe. You can still use an agent to find affected callers or draft a test plan.

Label results from earlier runs with the build ID they came from. The release summary should show which claims were checked against this build and which rest on review or an accepted gap.

Where Expo fits

EAS Build gives every Expo and React Native build an ID. In EAS Workflows, a build job's outputs include that build_id along with the git_commit_hash and fingerprint_hash it was made from. Later jobs take that build_id as input, so a Maestro test job (in alpha) and a submit job act on the same build instead of whatever is newest. Put a require-approval job between them and a person signs off on that build before it goes to the store.

The @expo/fingerprint library hashes your dependencies, custom native code, native project files and configuration. When an agent's follow-up change leaves the fingerprint unchanged, a JavaScript update may be enough. A changed fingerprint means a new build, and the native checks from Step 4 run again. One catch from the docs: editing a config plugin written as a raw or anonymous function can leave the hash unchanged, so give plugin functions stable names.

The Expo MCP Server gives Claude Code, Codex, Cursor and other agents tools to read EAS build logs and TestFlight crash reports. With the expo-mcp package and a local dev server, an agent can also take screenshots of the running app and tap views by testID.

Expo Simulators, currently presented as early access, can be evaluated for remote app interaction where the supported environment fits the test. Confirm access and build compatibility before you make it a required job. After release, EAS Observe adds production performance data; evaluate its preview error features separately from the checks you require before release.

Limitations

No service turns an agent's claim into proof on its own. You still define the assertion and check that the result belongs to the build you plan to ship. A successful build says nothing about hardware-specific behavior or security. The Maestro job in EAS Workflows is in alpha, and the Expo MCP Server's local iOS tools work only with simulators on a macOS host.

Next step

Take the OWASP mobile security verification requirements and your feature's acceptance criteria and write one acceptance table like the one in Step 1 before you ask the agent for the next change.

Verified on 12 September 2026.

Keep reading

Frequently Asked Questions