End-to-end tests have a bad reputation, and it is deserved. Most teams write a suite, watch it fail randomly for three months, then quietly stop running it. Playwright does not fix that on its own — but it removes most of the causes, provided you write tests the way it expects.
This guide covers getting a suite running, the patterns that keep it stable, and the specific habits that make E2E tests flaky.
Why Playwright Rather Than Selenium or Cypress
- Auto-waiting. Every action waits for the element to be actionable before interacting. This alone eliminates most flakiness, and it is why explicit sleeps almost never appear in a good Playwright suite.
- Real browser engines. Chromium, Firefox and WebKit, so you catch Safari-specific bugs without owning a Mac.
- Parallel by default. Tests run in isolated browser contexts across workers, which keeps a large suite to a sensible wall-clock time.
- Trace viewer. A failed run can be replayed with DOM snapshots, network activity and console output at every step. This is the feature that makes CI failures debuggable instead of infuriating.
Getting Started
npm init playwright@latest
That scaffolds a config, an example test, a GitHub Actions workflow, and downloads the browsers. Then:
# run everything, headless
npx playwright test
# one file, headed, so you can watch it
npx playwright test tests/login.spec.ts --headed
# open the trace for the last failure
npx playwright show-trace
A First Test Worth Keeping
import { test, expect } from '@playwright/test';
test('a signed-out visitor can reach the pricing page', async ({ page }) => {
await page.goto('/');
await page.getByRole('link', { name: 'Pricing' }).click();
await expect(page).toHaveURL(/\/pricing/);
await expect(page.getByRole('heading', { name: 'Plans' })).toBeVisible();
});
Two things here matter more than they look.
First, getByRole rather than a CSS selector. Role-based locators survive restyling and refactors, and they fail loudly when you break accessibility. A test suite built on .btn-primary > span:nth-child(2) breaks every time someone touches the markup.
Second, expect(...).toBeVisible() retries until it passes or times out. You do not need a wait before it.
Locators, in Order of Preference
| Locator | Example | Use when |
|---|---|---|
getByRole |
getByRole('button', { name: 'Save' }) |
Almost always — the default |
getByLabel |
getByLabel('Email address') |
Form fields |
getByText |
getByText('Order confirmed') |
Non-interactive content |
getByTestId |
getByTestId('cart-total') |
Last resort, when nothing else is stable |
| CSS / XPath | locator('.cart > li') |
Avoid — breaks on restyling |
Signing In Once Instead of Every Test
Logging in through the UI in every test is the most common reason suites are slow. Do it once, save the storage state, and reuse it.
// auth.setup.ts
import { test as setup } from '@playwright/test';
const authFile = 'playwright/.auth/user.json';
setup('authenticate', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill(process.env.TEST_EMAIL!);
await page.getByLabel('Password').fill(process.env.TEST_PASSWORD!);
await page.getByRole('button', { name: 'Sign in' }).click();
await page.waitForURL('/dashboard');
await page.context().storageState({ path: authFile });
});
Then point a project at it in playwright.config.ts so every test in that project starts signed in. A suite that logged in 80 times now logs in once.
Never hardcode credentials. Read them from environment variables and keep the auth file out of version control.
The Habits That Cause Flaky Tests
Nearly all E2E flakiness comes from a short list.
- Fixed sleeps.
waitForTimeout(2000)is either too short on a slow CI runner or wasted time locally. Wait for a condition instead: a URL, a visible element, a network response. - Tests that depend on each other. If test B needs the record test A created, running in parallel or re-running one test breaks it. Each test should create what it needs.
- Shared mutable state. Two tests using the same account and both editing the profile will collide. Give each worker its own data.
- Asserting on live third-party services. Payment sandboxes and email providers go down. Mock them at the network layer with
page.route(). - Animations. Elements that slide in can be clicked mid-flight. Disable animations in the test environment via CSS rather than sleeping around them.
Mocking the Network
test('shows an error when the API is down', async ({ page }) => {
await page.route('**/api/orders', route =>
route.fulfill({ status: 500, body: 'Server error' })
);
await page.goto('/orders');
await expect(page.getByText('Something went wrong')).toBeVisible();
});
Error states are the paths users hit at the worst moment and the paths nobody tests, because reproducing them by hand is tedious. Intercepting the request makes them trivial, and they are some of the highest-value tests you can write.
Running in CI
- name: Install Playwright browsers
run: npx playwright install --with-deps
- name: Run tests
run: npx playwright test
- uses: actions/upload-artifact@v4
if: always()
with:
name: playwright-report
path: playwright-report/
retention-days: 7
if: always() matters — without it the report only uploads when tests pass, which is precisely when you do not need it.
Set retries: 2 in CI only, and treat a test that passes on retry as a bug to fix rather than a success. Retries are for hiding infrastructure noise, not for tolerating a flaky test forever.
How Many E2E Tests Should You Have?
Fewer than you think. E2E tests are slow and expensive to maintain, so spend them on flows where a failure would be genuinely costly: signing up, signing in, checkout, whatever your product’s core action is.
Everything else — validation rules, formatting, edge cases in business logic — belongs in unit tests, which run in milliseconds and tell you exactly what broke. A suite of 30 well-chosen E2E tests that always passes is worth more than 300 that fail randomly.
Frequently Asked Questions
Is Playwright better than Cypress?
For most teams, yes. Playwright runs real WebKit for Safari coverage, parallelises across workers out of the box, handles multiple tabs and origins without workarounds, and is free at any scale. Cypress has a friendlier interactive runner, which some teams value highly.
How do I stop Playwright tests being flaky?
Remove every fixed sleep, make each test independent and self-provisioning, mock third-party services, and disable animations in the test environment. Those four changes fix the large majority of flakiness.
Should I use data-testid attributes?
Only when nothing better exists. Prefer getByRole and getByLabel — they survive refactors and double as an accessibility check. A test id is a reasonable fallback for something like a total that has no accessible name.
How long should an E2E suite take?
Under ten minutes in CI, ideally under five. Beyond that people stop waiting for it and start merging around it. Parallel workers and reusing authentication state are the two changes with the biggest effect.
Can Playwright test authenticated pages?
Yes — sign in once in a setup project, save the storage state to a file, and have the other projects load it. Every test then starts already authenticated without repeating the login flow.
Does Playwright work with React, Vue or Svelte?
It drives a real browser, so the framework is irrelevant. Anything the browser renders, Playwright can test. There are also component-testing modes if you want to test components in isolation rather than a full page.
The Bottom Line
Install it, write a handful of tests for the flows that would cost you money if they broke, use role-based locators, reuse authentication state, and never write a fixed sleep. That gets you a suite people trust — which is the only kind worth having.
✍️ Leave a Comment