---
title: Your skills need an evaluation mechanism
authors: Aditya Shukla, Kartik Gupta, David Tingle, Kudo Chien, Brent Vatne, David Mokos
published: October 8, 2026
categories: AI, Development
tags: agent skills, evals, AI development, agentic development
---

_This is a collaboration post with Aditya Shukla, Kartik Gupta, and David Tingle from the [Georgian AI Lab](https://georgian.io/). We partnered with them to measure how coding agents find and use Expo's agent skills._

## Key takeaways

- **Test whether agents can find skills, not only whether skills work when called.** Direct invocation does not show whether an agent will discover a skill during ordinary work. Use realistic task prompts and inspect the agent's trace to see what it finds.
- **Establish a clear entrypoint skill to help agents reach specialized guidance.** Across 10 app-building tasks, the share of agents loading a relevant Expo skill rose from 9% to 55% after we added `expo-overview` as an entrypoint skill.
- **Design for crowded skill catalogs.** When many skills share a catalog budget, descriptions can be shortened or omitted. Put the most important cues in your skill's title and opening sentence, and test it in a realistic catalog.

At Expo, we bundle our agent skills into [one plugin](https://github.com/expo/skills) for Claude Code and Codex. The plugin covers a fairly wide surface area: there are skills for building native interfaces, setting up navigation, fetching data, adding animations, working with [Expo Modules](https://docs.expo.dev/modules/overview/), and using [our cloud services](https://expo.dev/services). We are constantly refining these skills and adding the best context we can, so that they can augment what coding agents already know and help them make better decisions across a task.

For a while, though, we did not have a robust way to tell whether those improvements were actually helping. We would try a prompt ourselves, watch the agent do something sensible, and feel good about the result. But writing skills without evals started to feel a lot like writing code without tests. "It worked with my prompt" was becoming another version of "it worked on my machine."

This problem connects to a broader lesson from Georgian's experience with evaluation: a final answer cannot show us everything that happened along the way. In [Your AI Systems Need an Eval Loop](https://georgianailab.substack.com/p/your-ai-systems-need-an-eval-loop), Asna Shafiq, a member of the Georgian AI Lab, describes how traces can help teams observe agent behavior, but that you need evals to turn those observations into something you can measure and improve. We brought that same idea to our skills: rather than only checking whether a skill works in isolation, we needed to test whether an agent can find the right mix of skills during routine app-development tasks.

There was also a bias in the way we tested our work. As the people writing and maintaining these skills, we already knew the right usage patterns. If we wanted to test the Expo UI skill, we would invoke `/expo-ui` in Claude Code or `$expo-ui` in Codex and then see how the agent used it.

That tells us whether the skill is directionally useful once it has been loaded, but most users don't always invoke a skill by name. They're more likely to describe what they want: _"Build an iOS todo app with Expo and make it feel native,"_ or _"Add tabs and a settings screen to this app."_ Before any of the guidance _inside_ a skill can help, the agent first has to recognize that the skill is relevant and decide to invoke it.

We had no mechanism to measure the many variables across this process.

## Building an eval harness

To solve exactly this problem, we developed an evaluation harness with our collaborators at Georgian. Their earlier work on evaluation loops informed the design: 1) define what evidence would matter before a run, 2) capture the agent's trajectory on a task of interest, and 3) compare behavior under controlled conditions.

At a high level, the eval harness takes product requirements as a task, runs a coding agent against it, and evaluates both the process (the agent's trajectory and skill usage) and the product (the built mobile app).

### Define what to measure

The harness evaluates a coding agent using tasks drawn from a dataset of _scenarios_ that we have found to be representative of how real users build with Expo. A _scenario_ can be an end-to-end task such as creating a full mobile app in a one-shot setting, or more fine-grained, like debugging a dev build that keeps crashing or adding a feature to an existing app.

Any strong evaluation is built on top of a robust set of checks. Rather than inferring these checks post-hoc from the agent trajectory, each scenario in our dataset is annotated with:

- what skills we expect the agent to find helpful in performing the given task
- what evidence to look for if the guidance from the skills is followed
- what feature primitives an end-to-end task entails, that our agentic evaluator should QA

These checks give us three different layers of evidence:

1. **Skill triggering**: did the agent load the skills that were relevant to the task?
2. **Skill uptake**: can we see evidence that the agent followed the guidance it loaded?
3. **App behavior**: does the resulting app build and perform the behaviors described in the task PRD?

These questions are related, but they each shed light on a different slice of agent behavior that we have observed.

A skill cannot influence an agent if it is never loaded, so skill triggering must be measured. A skill can be loaded and then ignored, though, so measuring evidence of uptake matters too.

An agent can also build a working feature from its existing knowledge without ever consulting the skill we hoped it would use. This is perfectly fine! In fact, it's crucial to not penalize the agent for performing the task correctly even if it does not reach for Expo's curated wisdom. If anything, this is a signal that a skill may have been subsumed by increasing frontier capabilities. Constructing these as separate layers of evaluation has made it easier to understand where the coding agent can most be helped.

### Run the coding agent with your task

In the first iteration, we test coding agents against more ambitious, end-to-end scenarios: we start with a product requirements document (PRD) that captures a high-level user prompt about wanting to build a particular mobile app. Our eval harness then deploys a coding agent with this task and records its trajectory over this authoring process, including the skills it invokes. Once the agent completes generation, the harness zips the build artifacts up to be passed on downstream.

### Evaluate both the process and product

The harness then kicks off two parallel evaluation jobs that answer orthogonal questions:

**Static evaluation**: a set of fine-grained, deterministic checks that measure the build health of the produced app, whether any skills were invoked over the agent trajectory, and whether the guidance from these skills actually found its way into the produced source code.

**App evaluation**: an agentic evaluator that can build and spin up the mobile app that the coding agent produced, and then drive the app like a user would to validate the core feature primitives.

![The evaluation harness](https://cdn.sanity.io/images/9r24npb8/production/264760aacd68ad8f87e9e989e80c63af38a9041b-3280x1544.png)

### Run the eval harness in CI

The eval harness itself runs on [Expo’s CICD Workflows](https://expo.dev/services/workflows) and we connect it to the [Expo Skills repository](https://github.com/expo/skills) where Expo develops and maintains their plugin. Today, any time someone at Expo ships a new skill or updates existing skills, they can trigger an evaluation right from the pull request by adding an `eval` label. Then, the harness simply runs the selected scenarios against the version of the Expo plugin on `main` and the version in the pull request and finally surfaces the results of this eval directly in the PR.

This is now increasingly becoming part of the development lifecycle at Expo: skill authors own and maintain their tasks and ground truth checks just like unit tests, and therefore the plugin is itself developed like code: make an intervention, run it against a set of scenarios and inspect the evidence before deciding what to do next.

## How has this helped us?

One of the first things the authoring traces showed us was that the expected Expo skills were almost never loaded like we would have hoped.

This was useful because the applications could still look superficially successful. The agent already knew enough about Expo and React Native to produce code, so looking only at the final answer or scanning a few files would not necessarily reveal that our plugin had contributed nothing to the process. The trace made that visible.

We had a few ideas about why this may have been happening.

- The agent might already feel confident that it could complete the task and therefore never look for additional guidance.
- The descriptions in the skill frontmatter might not adequately describe the situations in which each skill should be used.
- As the plugin grew, the number of available skills might itself make the collection harder to navigate.

Previously, we could debate these possibilities, update a description, and try a few prompts by hand. Now we had a mechanism to test them.

We started with the most obvious intervention, improving the names and descriptions of individual skills. The description is what the agent sees when it is deciding whether a skill is relevant, so it seemed like the natural place to begin. These changes helped us clarify what each skill was for, but they did not get us far enough.

That led us back to the third hypothesis, which was actually a suspicion we'd harbored for a while…

## Giving the agent an entrypoint to Expo skills

The Expo ecosystem covers a lot of ground: React Native, the [Expo SDK](https://docs.expo.dev/versions/latest/) itself, [Expo Router](https://docs.expo.dev/router/introduction/), Expo Modules, native UI elements, and of course our Cloud Services. Together, they offer an incredibly powerful set of capabilities for app development, deployment and observability. Having focused skills for each lets us provide detailed guidance without putting the entire Expo ecosystem into the agent's context every time. But it also means the agent has to choose among more than twenty possible skills to start with.

We needed a single entrypoint: one that could help the agent enter the collection and then navigate it.

The result was [`expo-overview`](https://github.com/expo/skills/pull/107): an entrypoint skill and router for all Expo and EAS tasks. Its description establishes it as the starting point when any request involves Expo, EAS or a React Native task. Inside the skill, a map connects common goals to the more focused skills that contain the relevant guidance. It also owns a small set of setup rules that apply across the plugin, such as detecting the Expo SDK version and installing packages with `npx expo install`. Recently, Sentry and Convex arrived at a similar pattern for their own skill plugins: essentially, give the agent a clear place to begin, then route it toward more specialized guidance.

This overview skill does not change the fact that the agent still has to browse across 20+ skills; rather, it addresses the failure mode we had observed: that often, the agent would simply not invoke any Expo skill at all.

Why could this be the case? An app development task already endows the agent with a long list of decisions to make about architecture, dependencies, navigation, state, and implementation. If it also has to choose among more than twenty skills in the middle of doing everything else, it may simply continue with what it already knows and never reach for any of them. This is speculative, but the goal is simply to minimize the activation energy of reaching for a skill.

![The entrypoint](https://cdn.sanity.io/images/9r24npb8/production/25490b5862e2901295592769f975be44d9d60043-2550x1460.png)

## What did the harness show?

We added the `eval` label to the [pull request shipping this new skill](https://github.com/expo/skills/pull/107), and let our eval harness run the same three application scenarios against the existing plugin and the proposed change. The first encouraging signal was the agent verbalizing:

> "I'll start by loading the expo-overview skill, since this task involves building an Expo/React Native app..."

In another scenario, the agent explained the same decision even more directly:

> "I'll start by loading the Expo overview skill, since it must run first for any Expo/EAS task."

This was the behavior we were trying to create. The user had not named `expo-overview`, or any of the downstream skills, in the request. The agent recognized the broader Expo task, chose the entrypoint, and then used the routing information inside it to decide what else to load.

The traces then showed the following skills loaded in order, for three PRDs from our dataset:

- _Notes_: `expo-overview` → `expo-project-structure`
- _Hot Chocolate_: `expo-overview` → `expo-project-structure` → `expo-router`
- _Wiki Reader_: `expo-overview` → `expo-project-structure` → `expo-router` → `expo-ui`

The overview was loaded first in all three runs. Each run then continued from the overview into at least one relevant downstream skill. Once the agent crosses that first threshold of loading an Expo skill at all, it appears to be much more likely to make the second-order decision than it was when all of the possible skills were competing to be selected at once.

## From only 10% of agent sessions loading a skill to more than 50%

To see whether this initial promising result from the PR held more broadly, we reran Claude Code with `sonnet-5-high`, this time against 10 greenfield mobile app PRDs. The user-facing prompt mentioned Expo but never named a skill. We compared the plugin before and after the addition of `expo-overview`, and performed three runs apiece to iron out stochasticity in the agent trajectory.

In the baseline runs without `expo-overview`, only about 10% of the agent sessions decided to load any Expo skill. With `expo-overview`, that number jumped to 55% of all agent sessions loading 1 or more downstream skills.

![metrics](https://cdn.sanity.io/images/9r24npb8/production/2cf10d77c20a8780d6bd08b4c60ace7fd5f3dde2-3000x1144.png)

Every candidate session loaded `expo-overview`. This validated our goal of providing the agent a dependable entrypoint for an Expo-related task. More importantly, 55% continued to at least one more specialized skill, compared with only 9% of baseline sessions loading a skill directly.

The Relevant Skill Recall number gives the more demanding version of the same question. Within the dataset, we also specify a mapping of the subset of Expo skills relevant to each PRD. These are skills that would provide useful guidance for the corresponding task, and that we thus expect the agent to invoke. This Relevant Skill Recall rose from 2% to 18%, but remained incomplete: even after the addition of `expo-overview`, the agent was frequently invoking only a subset of the relevant skills for most tasks.

Looking at this number across individual skills gave us some more intuition:

![individual skills metrics](https://cdn.sanity.io/images/9r24npb8/production/43b5da6bf042ccc137d32451f50398620d3b82b4-3000x1374.png)

The skills that were loaded most often were tied to early, concrete decisions, such as scaffolding the codebase (`expo-project-structure`) or setting up the page routes (`expo-router`).

In one PRD, the agent loaded the overview skill and proceeded to _"(Let me) check the project-structure skill for scaffolding guidance."_ In another navigation-heavy task, it made the connection even more directly: _"Now let's convert it to use Expo Router (needed for tabs navigation) and check the router skill for setup steps."_ Both of these were examples of tasks the agent attended to immediately after perusing the overview skill, and thus often read the corresponding skill too.

The misses were equally intriguing: in one case, the agent loaded `expo-overview` and `expo-router`, then announced that it would build the core data layer, implemented RSS and Atom fetching and tested real feeds, all without ever loading `expo-data-fetching` (possibly the most relevant skill for this task), preferring to work from innate knowledge instead. In another instance, the agent loaded `expo-ui` for controls but did not load `expo-native-ui`, even though `expo-overview` recommends using them together. The generated code used `Dimensions.get()` and React Native's `SafeAreaView`, two patterns the skipped Native UI guidance would have steered it away from.

Across these instances, it appears that the agent invokes relevant skills much more frequently towards the beginning of a task, immediately after reading Overview. As it progresses through the rest of the work, it becomes less likely to invoke relevant skill(s), preferring to work from other information sources.

One possible explanation for this is that though `expo-overview` is the de-facto entrypoint, it is not quite perceived as a general router that the agent should repeatedly refer back to throughout the task.

This was an initial signal from a handful of datapoints, and more work needs to be done to more reliably route agents to the relevant skill wisdom. But we could now see a clear difference in the behavior we were trying to influence. That was the feedback loop we had been missing.

## The pitfalls of too many skills? Truncation!

Our initial experiments ran in Claude Code, but this result also made us curious about how skills are presented in other coding agents. We looked under the hood of the open-source [Codex CLI](https://github.com/openai/codex) to understand the surface where skill discovery happens.

Structured after the system prompt is a _Skill Catalog_ section containing each available skill's name, its description, and its file path. In Codex, that catalog has a fixed total length that must be shared by all skills available to the agent. For the version we inspected, this allowance was set to exactly 2% of the context window, or 5,440 tokens of 272k overall.

![Truncation](https://cdn.sanity.io/images/9r24npb8/production/53aae00fb7f1dff22ad7eee7b605a94f6fb19e16-3680x2600.png)

We had also long wondered about the effect of skill crowding: what happens when the skills from our plugin are competing with other non-Expo skills?

Through some controlled tests, we found that when the full skill catalog exceeded the allowance of 5.4k tokens, Codex started by shortening skill descriptions, preserving only their beginnings. With even more skills, the descriptions disappeared entirely, leaving just the skill names. This was followed eventually by entire skills being dropped from the catalog!

We saw this play out in one test environment with 175 skills installed, including 24 from the Expo plugin. With the 272k context window, we found all Expo skills listed, but only the first 60 or so characters of each description were visible.

![the pitfall of too many skills](https://cdn.sanity.io/images/9r24npb8/production/cbf5994eec637c447beb8fdebb4cfbb431ec802d-3680x1822.png)

So, as the catalog of available skills grows, the agent is choosing among more options while potentially receiving less information about each one. The name and opening words of a skill alone can make a difference.

**TL;DR:** _if there are too many skills fighting for the same 2% of your context window, some of their descriptions (or even the skills themselves!) might get cut! Make sure yours are robust to this!_

This also exposed an assumption in our setup: our eval harness deploys a coding agent in a clean environment with only the Expo plugin installed, under no competition with skills from other vendors. That isolation is useful when comparing two versions of the plugin, because it removes unrelated sources of variation. But it is not necessarily how users run coding agents. A real user may have many plugins installed, all competing for the same catalog budget and for the agent's attention.

In essence, skills need to be evaluated under the conditions in which we expect them to be discovered.

An isolated catalog can tell us whether a skill is legible on its own. It cannot tell us whether that skill remains visible and distinct inside a crowded catalog. Our initial experiment measured the first case. But the truncation we observed in the more representative local setting suggests that the second probably deserves its own eval condition too!

## What we learned

The first lesson was that a group of individually useful skills does not automatically compose into an equally effective plugin. Once it grows beyond a few skills, any plugin will likely need some form of information orchestration. An agent needs to understand not only what each skill says, but also how the skills relate to one another and where it should begin.

The second lesson was that we needed to evaluate the way users actually interact with the plugin. Direct invocation is still useful when we are developing the contents of one skill, but it cannot tell us whether an agent will discover that skill from a normal request. By starting from broad product requirements and reading the resulting traces, we were able to observe the selection problem directly.

This also included evaluating the setting in which users actually interact with the plugin: when many other skills may also be competing for attention, the effects of your design could look different. Therefore, _the environment is an equally important part of your evals_.

The third lesson was that process measurements can tell us something useful before we have a perfect end-to-end score. We are continuing to improve the harness's uptake checks and App Evaluator, because ultimately we care about whether the guidance produces a better application. But triggering is a necessary first step. If a skill never enters the agent's context, the quality of the instructions inside it does not matter for that run.

This is why connecting the harness to CI has been important for us. We do not want skill evaluation to be a one-time experiment that we run when we already expect a positive result. We want it to become part of how we develop the plugin: make a change, run realistic tasks, inspect the traces and artifacts, and use what we learn to choose the next intervention.

`expo-overview` was the first useful example of that loop. The harness showed us that agents were not finding our skills. We changed the architecture of the collection rather than continuing to tune every skill in isolation. The next run showed the agent reasoning about the new entrypoint in the way we had intended and loading more of the downstream guidance.

This gave us our first clear example of the process working: observe where agents struggle, make a targeted change, and measure whether their behavior changes.

There is more work to do. But we now have an evaluation mechanism. You need one too.