Until recently, I would say that Playwright has been leading AI-related innovation in testing. It regularly introduces something new: MCP servers, CLI tools, skills, and custom agent definitions, including approaches to self-healing tests. I have covered these developments in previous posts, and Playwright has given us quite a lot to experiment with.
Cypress also offers some interesting AI features, but they have not gained the same traction across the broader testing community. In my view, this is largely because many features in the Cypress ecosystem are still paid. The trend I see also remains the same: Playwright continues to grow, while Cypress stays broadly at the same level.
There was also Vibium. My impression is that it has yet to establish itself as a widely adopted alternative. And into this ecosystem, which, let's be honest, is still fairly static, comes a new player with a rather suggestive name: e2e.
This article includes a detailed, runnable demo, with the migrated tests, practical examples of the AI features and the measured costs. You can clone it and try the same experiment yourself.
In previous posts, I also discussed agentic testing: simply asking an agent to carry out tests. In my opinion, we have still not explored this thoroughly enough. My impression is that it has recently started working better than many people expect. If your opinion comes from older experiments, established habits, or a bad experience some time ago, I would encourage you to try it again and see whether that opinion still holds. From what I have seen, it already works quite well. But that is a subject I may return to in future posts.
Introducing E2E
e2e is a testing framework that, in practice, feels quite similar to Playwright. We write tests, interact with the application and check the results. What makes it interesting is that we can use an LLM for selected parts of a test, alongside ordinary actions and assertions.
There are four main AI operations: act, assert, waitFor and extract. I wanted to try all four in practice and show how they work in an existing test suite.
AI Actions: agent.act
We describe a goal, and the model decides which interactions are needed to complete it. A single goal can involve several clicks, filling fields or moving between pages. For example:
await agent.act('Open the cart and proceed to checkout.');
We still decide the order of steps in the test, but let the model choose how to perform this particular step. This is the difference from specifying every locator and click ourselves. The actions documentation explains how to supply test data and write goals.
AI Assertions: agent.assert
We describe what should be true, and the model checks whether the current screen supports that statement. This can be useful when the meaning matters more than the exact wording:
await agent.assert(
'Checkout explains that some items are unavailable ' +
'and that the customer’s cart was preserved.'
);
The assertion fails if the statement is false or the model cannot establish it from the screen. It judges the current page, without the history of earlier actions. An ordinary assertion is still a good choice for an exact price or quantity.
Waiting for a Result: agent.waitFor
Sometimes the expected result has not appeared yet. waitFor waits for a condition described in natural language, rather than checking it immediately:
await agent.waitFor('The order confirmation is visible.', {
timeout: 30_000,
});
It checks the screen periodically and asks the model to assess it when the observation changes. Unlike act, it does not perform the interactions needed to reach that result.
Reading Data: agent.extract
We can also ask the model to read information from the page and return it in a defined structure. For example, we could collect the visible product names:
import { z } from 'zod';
const products = await agent.extract('Read the visible product names.', {
schema: z.object({ names: z.array(z.string()) }),
});
The result is data we can use in subsequent checks. In my experiment, I used extraction to read product names and then compared their order with the expected list. I also used it to read the available inventory quantity, the stock movement and its reason, checking the returned values against the expected result. The assertions documentation covers assertions, waiting and extraction.
In my opinion, this makes for an interesting mix. We keep the familiar structure of an automated test, while moving selected actions and checks from deterministic code to probabilistic LLM behaviour. That places e2e somewhere between conventional deterministic Playwright tests and open-ended agentic testing, where an agent has much more freedom to decide what to do and how to test it.
Getting Started
Getting started is similar to working with any other npm tool. Run the setup wizard in your project directory:
npx e2e init
For browser tests, choose Web. The wizard adds the dependencies and creates a configuration file and an example test. You then set the application's URL and choose the model you want to use.
For the AI features in this experiment, I connected Sonnet through Anthropic's API. That requires an API key, supplied as an environment variable in the terminal where the tests run:
export ANTHROPIC_API_KEY="your-api-key"
npx e2e run
Ordinary tests without AI steps do not need a model or an API key. The official quickstart walks through the setup and also explains alternative ways to connect a model, including supported subscription logins.
Migrating the Existing Playwright Tests
The original Playwright suite contained 302 tests: 234 API tests and 68 UI tests. It covered login, registration, products, carts, checkout, orders and administration. All those tests are represented in the new project. There is also one small AI smoke test, giving 303 tests in total.
As usual, I ran the tests against my Awesome LocalStack, using its lightweight Docker Compose profile. The application runs locally at http://localhost:8081. Its own LLM endpoints use a local mock; the AI test steps call Anthropic's Sonnet API.
Codex performed the migration and fixed the problems found during execution. Looking at the local project and test-report timestamps, about 16 minutes passed between creating the new project and the first complete passing run. That run passed 303 tests in approximately 71 seconds, using four workers.
The 16 minutes covers that initial setup, rewrite and verification. It excludes trying the tool before creating the project, and the later work adding AI assertions, investigating failures and measuring costs. I also started with a known application and an existing working suite. A migration in a company project with unfamiliar dependencies could be a very different exercise.
Preparing the experiment and this article took longer, with work spread across two days. The larger recorded experiments alone account for about 35 minutes of test execution, including repeated runs and deliberately introduced errors. That does not include writing and editing, which I did not time separately.
There were some real adaptations. For example, the new runner prepared test fixtures differently from Playwright, so copying the original fixture structure would have created unnecessary test data. The API clients also needed a different HTTP implementation. Codex handled those changes, but this was more work than replacing an import.
The complete demo is public on GitHub. Its README explains how to start the local application, connect your own model API key and run the tests. For me, the useful result is that we could evaluate the tool against an existing suite. We did not need to invent a few easy scenarios and hope that the same approach would work elsewhere.
Adding 95 AI Assertions
I extended the 68 existing UI tests with 67 AI actions, 95 AI assertions, two AI waits and two structured-data reads. I kept the 234 API tests conventional. The additional smoke test is outside the measurements below.
The actions included editing products, adjusting inventory, changing order status, cancelling orders, modifying the cart and filling checkout details. The assertions checked things such as cart totals, saved profile instructions, order information and feedback after an unsuccessful checkout.
One of the cart assertions looks like this:
await agent.assert(
'The cart has two distinct product lines, one with quantity 3 ' +
'and the other with quantity 1. The summary combines them ' +
'as 4 units costing $42.50.'
);
The model reads the current page and decides whether the statement is true. We can also check meaning rather than exact wording: for example, whether checkout explains that some stock is unavailable and that the customer's cart was preserved.
I kept the original exact checks alongside these assertions. That made it possible to compare the AI result with ordinary checks of the UI and the data returned by the API. It also means that adding 95 AI assertions did not add 95 new test scenarios. These are additional checks inside the tests we already had.
Trying All Four Operations
I used the existing checkout, inventory and product-sorting tests to try all four operations. Across three consecutive live runs, all nine test executions passed. The extracted product order and inventory values also agreed with the exact expectations.
For waitFor, I wanted to check something more useful than asking it to recognise a result that was already there. I temporarily hid the order or inventory panel for six seconds, then made it visible again. All six attempts waited successfully, taking roughly nine seconds each. When I kept the panel hidden, all four attempts timed out at the eight-second limit.
These were controlled changes to the page, not application bugs. The results are encouraging, but a few repeated tests are not enough to establish a general reliability rate.
The waits and extraction worked, although I would not use them automatically. If waiting for one known element is enough, an ordinary wait is simpler. If a quantity has a stable locator, reading it directly is cheaper. I see more value when readiness depends on the meaning of several pieces of information, or when I need structured data from content whose layout varies.
Cost and Execution Time
Here are the results for the same 68 UI tests, run with four workers:
| How the tests ran | Passed | Time | Model API cost |
|---|---|---|---|
| Ordinary actions and assertions | 68/68 | 21 seconds | $0 |
| AI actions, ordinary assertions | 68/68 | 118 seconds | $1.21 |
| Ordinary actions, AI checks | 68/68 | 104 seconds | $0.74 |
| AI actions and checks | 68/68 | 200 seconds | $1.99 |
| AI actions and checks, with action replay | 68/68 | 182 seconds | $1.03 |
The AI-check rows include all 95 assertions, two waits and two extractions. These were separate runs, so the costs do not add up perfectly: model responses and the amount of cached input varied between runs. The 71-second full-suite result mentioned earlier was measured before this larger AI expansion.
In the run using AI actions and all four operations, the 95 assertions cost about $0.70. An assertion averaged $0.0073, or roughly 0.7 cents. At the measured average, 1,000 comparable assertions would cost approximately $7.32. The two waits cost $0.014 in total and the two extractions $0.019. That is less than one cent per operation in this run, although these were small examples with little data to read.
An AI action was more expensive: about 1.9 cents on average in that run. It sometimes needed several calls to the model to inspect the page, perform the interaction and decide it had finished. Asking AI to click an ordinary navigation link also has a price, even if there is little reasoning involved.
For a personal experiment, these amounts seem reasonable to me. The larger experiment cost $15.72 in total, including failed runs and the checks with deliberately introduced errors. The additional runs covering all four operations cost $5.24 of that total. The amount excludes the earlier, smaller experiments.
A company would have to multiply the cost by its actual usage. One thousand executions of this complete UI suite would cost about $1,992 at the measured rate, or about $1,035 with the replay behaviour seen here. Those figures assume similar pages, model usage and caching. They are examples of scale, not measured thousand-run results.
The time difference is more noticeable. Adding AI throughout the UI tests increased execution time from 21 seconds to 200 seconds, about nine and a half times as long. A few dollars may be easy to accept. Waiting longer for feedback on every change deserves more thought.
The prices are estimates from the model's reported token usage and Sonnet 5.5 API pricing, checked on 2 October 2026. They cover model calls, including failed attempts, rather than the cost of running the local application or developing the tests.
Caching Helps, but Assertions Still Cost Money
The runner can record successful actions and replay them later. In the replay-enabled run, 51 of the 67 actions completed without calling the model. Eight needed AI to continue after replay, and eight could not use a cached action.
That brought the total cost down from $1.99 to $1.03. The assertions still called Sonnet and cost about $0.71, with a little extra for waiting and extraction, so replaying actions does not make the whole test free.
Runtime improved less, from 200 to 182 seconds. Some actions spent time checking the replayed result before returning control to the model. One action returning to the product catalog took about 31 seconds. I would therefore measure both time and cost when evaluating caching, rather than assuming that saving tokens will make the suite proportionally faster.
There is also caching on Anthropic's side, which discounts repeated input. That was still available in the runs where I disabled action replay. The $1.99 result is the cost of running the AI actions live with provider caching available.
Do the Assertions Find Errors?
After fixing the test problems, all five ways of running the UI suite passed. That was encouraging, but I also wanted to know whether an AI assertion would reject an incorrect page.
Following the idea from my mutation testing article, I deliberately changed selected screens. I introduced 19 errors and tried each twice. Examples included a wrong cart total, an incorrect quantity, a paid order displayed as pending, a changed shipping address and an enabled purchase button for an out-of-stock product.
Initially, the model recognised 31 of the 38 faulty screens. It was unable to decide twice, which also fails a normal test. It accepted the other five.
The misses were quite useful. An order showed a correct line total of $25.00 but an incorrect overall total of $24.99. One attempt accepted the page; the next rejected it. Another assertion accepted an incorrect checkout total because a different summary on the same page still showed the correct amount.
I made two assertion descriptions more precise, naming the overall total and requiring all displayed checkout totals to agree. The model then rejected all 38 attempts with the same errors. I also tried four additional errors, twice each, without changing those descriptions again. It rejected all eight.
This is a small experiment, and I improved the descriptions after seeing the misses. I would not turn the final result into a claim that AI assertions are 100% reliable. What it showed me is that the wording matters, and that an assertion should be tried against an incorrect result as well as a correct one.
There was another clear limit. I changed inventory through the API while leaving the displayed quantity unchanged. The AI assertion accepted the screen, while an API check found the wrong quantity. The model was judging what it could see. It could not establish that the saved data agreed with the page.
These errors were introduced for the experiment. I did not discover a production application bug during this work.
Where I Would Use It
In my article about self-healing tests, I divided UI tests into two groups. First, we have local, isolated tests, usually with external dependencies mocked. These should be stable and deterministic, giving us fast, predictable feedback on changes to our own application. Then we have broader end-to-end tests. They are heavier and often slower, but they can provide substantial business value. Stakeholders expect a professional project to have some evidence that the whole system works together.
Looking at that distinction, I would keep probabilistic AI actions out of the local tests. I see much more potential for them in the broader end-to-end journeys.
Those journeys are more exposed to instability. They often involve third-party integrations, payment sandboxes, APIs maintained by other teams, or screens belonging to another team or even another company. We do not control all those dependencies, but we still need to use them to complete a business flow.
This is where AI actions look particularly useful to me. Imagine a test that needs to pass through an external provider's screens. Our goal is to complete that part of the journey and continue testing our product. Asking an agent to do that could reduce the amount of navigation code and the number of selectors we have to maintain whenever the provider changes its page.
In that situation, I would be comfortable giving the agent some freedom to choose how to proceed, while keeping exact checks on the business result we expect. The distinction from the broken Products button in this experiment matters: we need to be clear about which part of the journey we own and which behaviour the test is supposed to verify.
The money spent on model calls could turn out to be a good investment if it saves us from repeatedly updating selectors and investigating failures caused by changes to external screens. I would compare that cost with the time spent maintaining the existing tests. This seems like a particularly good use case for operating tools and interfaces over which we have limited control.
Another interesting use case is a visual redesign. Customers can be surprised by changes to a website, even when the functionality is still there. We could describe an existing user journey as an AI-driven test, verify that it works on the old version, and then run the same test against the new one. That would give us another way to check whether familiar tasks are still possible after changing the layout and navigation, without rewriting the test around the new selectors first.
These are two ideas I would like to explore further, and the list is certainly not complete. The framework is new, so I expect more use cases to emerge as people try it in their own projects. For me, this is a very interesting release, and I am curious to see how the testing community receives it.

Comments
Loading comments...