Building with AI agents

What is an agentic workflow for mobile development?

An agentic mobile workflow lets an AI coding agent read build and test results, change code and retry, while CI runs the builds and a person approves releases.

9 min read

Quick answer

An agentic workflow for mobile development is a loop in which an AI coding agent reads the result of a build or test, decides what to change, edits the code, and runs the check again until it passes or the agent reports a blocker. Fixed CI jobs still do the building and testing, and a person still approves what reaches users.

How is an agentic workflow different from a CI/CD pipeline?

Your iOS build fails at 6 p.m. with a linker error after a dependency bump. A CI pipeline posts a red X on the pull request and stops there. An agentic workflow hands the log to a coding agent such as Claude Code or Codex, which reads it, finds the mismatched native module version, fixes the dependency and requests another build.

The difference is who picks the next step. A pipeline runs the same steps in the same order on every push. An agent chooses its next action from the last result, while the build and test jobs it calls stay exactly as predictable as before.

Anthropic's post on building effective agents draws the same line between workflows that follow predefined code paths and agents that direct their own tool use. On mobile, keep the repeatable work in known jobs and use the agent where the next step depends on reading a result.

A job that always runs the same command after a push doesn't need a model. Adding an agent adds cost and one more way to be wrong, so put it where interpreting the output changes the work: a failing build, a flaky UI test, a crash report.

Give the agent a stopping rule. "Keep trying until it works" can produce twenty commits and no proof that any one of them fixed the original failure.

What happens in one loop, from failure to fix?

A useful loop starts with a signal you can reproduce, such as a failed test on a known build or a bug report with steps. The agent session reads that signal and calls build and test jobs to do narrow pieces of work.

PartWhat it doesExample
TriggerHands the agent a task and its contextA Maestro test fails on a specific build
Agent sessionInvestigates and picks the next actionReads the failure and proposes a narrow fix
Build and test jobsRun a defined operationBuild the app and rerun the failing test
Run recordKeeps what actually happenedBuild ID, test results, logs and screenshots
Release decisionApplies your release rulesApprove, reject or ask for another check

The agent session needs to know the repository and commit it is working on, along with the project's real build and test commands. Without them, an agent can mistake an environment problem, like a missing simulator runtime, for a bug in your app.

Let the agent work on a branch. Keep the original failure and every change made in response, so a reviewer can follow the result without replaying the whole conversation.

What should the agent get back from each build and test?

An exit code tells the agent whether a command succeeded by that command's own rules. A zero exit code from a build says nothing about whether a user can save a note.

Return results the agent can parse: which build was tested, which assertion failed, and the log lines around the failure. Strip credentials and customer data from those logs before the agent sees them.

For UI work, keep the starting screen and the screen after the action. A screenshot shows how the app looked at one moment. Only an assertion on saved data after a restart shows whether the data survived.

Keep "couldn't run" separate from "failed". If no iOS simulator was available, the agent should report iOS as untested instead of borrowing the Android result.

A scoped task for the agent can look like this:

Enforce the "Allowed" list in your tooling. A prompt can't make an overpowered token safe; the CI system or the token's scope has to block what the agent may not do.

Which decisions should stay with a person?

An agent can be wrong about why something failed. It can also edit a test until the wrong implementation passes.

Review changes to a test's expected behavior separately from changes to the feature code. Where you can, show the regression test failing before the fix and passing after it.

Use deterministic checks for anything you can state precisely. An agent looking at a screenshot can spot a broken layout, but an assertion on stored data is what tells you a record survived a restart.

Scope credentials to the job. A job that makes a test build doesn't need the App Store Connect API key that publishes to production, and code from an untrusted pull request should never run in a job that holds release secrets.

Set a retry budget and say when the agent should stop and report a blocker. Three identical failures from an expired signing certificate call for a person to renew the certificate; a fourth code edit won't help.

Tie each approval to a specific build and destination. If the agent rebuilds after you approve, your approval covered a build that no longer ships.

What does an agentic fix look like for a timeout bug?

Start with a narrow failure you can reproduce, such as an app that shows "Saved" after a request times out.

The agent reads the failing test and the logs, then reproduces the error on the identified build and proposes a change to the timeout handling. A build job makes a new build, and the test job reruns the test against it.

The result should show the original failure and the corrected behavior side by side. Check that a slow request that eventually succeeds still saves, so the fix doesn't turn every slow request into an error.

If the agent can't reproduce the report, it should say so. It can suggest extra logging, but it shouldn't claim a fix based only on an edited code path.

Widen the scope once you know how the loop fails. One well-watched repair loop for flaky UI tests teaches more than a plan for agents to run the whole app unsupervised.

Which tools make up an agentic mobile workflow?

Options in each layer, listed alphabetically:

LayerExamples
Coding agentClaude Code, Codex, Cursor, GitHub Copilot
Build and job runnerBitrise, CircleCI, Codemagic, EAS Workflows, GitHub Actions, Xcode Cloud
UI testsDetox, Espresso, Maestro, XCUITest
Agent control of a running appagent-device, Argent, Expo MCP Server
Production monitoringBugsnag, Datadog, EAS Observe, Embrace, Firebase Crashlytics, Sentry

Apple describes Xcode Cloud as a CI service built into Xcode for Apple developers, so if you pick it, plan a separate runner for your Android builds. Expo's agents guide lists agent-device from Callstack and Argent from Software Mansion as third-party toolkits that let an agent tap through flows, read logs and profile an app on a simulator or emulator.

Where Expo fits

Expo describes its build, test and release services as agentic mobile infrastructure: the jobs an agent calls, as opposed to a model running inside your app. EAS Workflows runs those builds and tests when the agent asks. Its pre-packaged jobs cover building, submitting, publishing updates, and running Maestro tests against an Android Emulator or iOS Simulator build (the Maestro job is in alpha). React Native gives the agent one codebase for iOS and Android, so it edits one project and the workflow builds both platforms.

The Expo MCP Server gives Claude Code, Codex, Cursor and VS Code tools to start a workflow run, list recent runs, and read the logs of a failed job or build. Expo Skills include an eas-workflows skill that teaches the agent the workflow YAML syntax.

The approval job pauses a workflow until a person approves or rejects it:

Put the approval job after your test jobs and make the release job depend on it. The approval job controls that one workflow; a token with wider access could still start a different release path, so scope tokens as well.

Expo Simulators is Expo's cloud simulator service. It is in early access with a waitlist, so not every account can use it yet. EAS Observe reports production performance, such as startup time from real user sessions, and an agent can query it with the eas observe: CLI commands.

Limitations

An agentic workflow is a way of working, and no single product supplies all of it. EAS Workflows runs jobs; it doesn't provide the reasoning agent, and it doesn't decide whether a change is safe to release. Expo Simulators is in early access, so check what your account can use today before you design a loop around it.

Next step

Use the EAS Workflows job reference to give your agent one scoped build-and-test workflow, and ask for a run record (build ID, test results, logs) from the first task.

Verified on 12 September 2026.

Keep reading

Frequently Asked Questions