Playwright End-to-End Testing Guide 2026: Tests That Do Not Flake

End-to-end tests have a bad reputation, and it is deserved. Most teams write a suite, watch it fail randomly for three months, then quietly stop running it. Playwright does not fix that on its own — but it removes most of the causes, provided you write tests the way it expects.

This guide covers getting a suite running, the patterns that keep it stable, and the specific habits that make E2E tests flaky.

Why Playwright Rather Than Selenium or Cypress

  • Auto-waiting. Every action waits for the element to be actionable before interacting. This alone eliminates most flakiness, and it is why explicit sleeps almost never appear in a good Playwright suite.
  • Real browser engines. Chromium, Firefox and WebKit, so you catch Safari-specific bugs without owning a Mac.
  • Parallel by default. Tests run in isolated browser contexts across workers, which keeps a large suite to a sensible wall-clock time.
  • Trace viewer. A failed run can be replayed with DOM snapshots, network activity and console output at every step. This is the feature that makes CI failures debuggable instead of infuriating.

Getting Started

npm init playwright@latest

That scaffolds a config, an example test, a GitHub Actions workflow, and downloads the browsers. Then:

# run everything, headless
npx playwright test

# one file, headed, so you can watch it
npx playwright test tests/login.spec.ts --headed

# open the trace for the last failure
npx playwright show-trace

A First Test Worth Keeping

import { test, expect } from '@playwright/test';

test('a signed-out visitor can reach the pricing page', async ({ page }) => {
  await page.goto('/');
  await page.getByRole('link', { name: 'Pricing' }).click();

  await expect(page).toHaveURL(/\/pricing/);
  await expect(page.getByRole('heading', { name: 'Plans' })).toBeVisible();
});

Two things here matter more than they look.

First, getByRole rather than a CSS selector. Role-based locators survive restyling and refactors, and they fail loudly when you break accessibility. A test suite built on .btn-primary > span:nth-child(2) breaks every time someone touches the markup.

Second, expect(...).toBeVisible() retries until it passes or times out. You do not need a wait before it.

Locators, in Order of Preference

Locator Example Use when
getByRole getByRole('button', { name: 'Save' }) Almost always — the default
getByLabel getByLabel('Email address') Form fields
getByText getByText('Order confirmed') Non-interactive content
getByTestId getByTestId('cart-total') Last resort, when nothing else is stable
CSS / XPath locator('.cart > li') Avoid — breaks on restyling

Signing In Once Instead of Every Test

Logging in through the UI in every test is the most common reason suites are slow. Do it once, save the storage state, and reuse it.

// auth.setup.ts
import { test as setup } from '@playwright/test';

const authFile = 'playwright/.auth/user.json';

setup('authenticate', async ({ page }) => {
  await page.goto('/login');
  await page.getByLabel('Email').fill(process.env.TEST_EMAIL!);
  await page.getByLabel('Password').fill(process.env.TEST_PASSWORD!);
  await page.getByRole('button', { name: 'Sign in' }).click();
  await page.waitForURL('/dashboard');
  await page.context().storageState({ path: authFile });
});

Then point a project at it in playwright.config.ts so every test in that project starts signed in. A suite that logged in 80 times now logs in once.

Never hardcode credentials. Read them from environment variables and keep the auth file out of version control.

The Habits That Cause Flaky Tests

Nearly all E2E flakiness comes from a short list.

  • Fixed sleeps. waitForTimeout(2000) is either too short on a slow CI runner or wasted time locally. Wait for a condition instead: a URL, a visible element, a network response.
  • Tests that depend on each other. If test B needs the record test A created, running in parallel or re-running one test breaks it. Each test should create what it needs.
  • Shared mutable state. Two tests using the same account and both editing the profile will collide. Give each worker its own data.
  • Asserting on live third-party services. Payment sandboxes and email providers go down. Mock them at the network layer with page.route().
  • Animations. Elements that slide in can be clicked mid-flight. Disable animations in the test environment via CSS rather than sleeping around them.

Mocking the Network

test('shows an error when the API is down', async ({ page }) => {
  await page.route('**/api/orders', route =>
    route.fulfill({ status: 500, body: 'Server error' })
  );

  await page.goto('/orders');
  await expect(page.getByText('Something went wrong')).toBeVisible();
});

Error states are the paths users hit at the worst moment and the paths nobody tests, because reproducing them by hand is tedious. Intercepting the request makes them trivial, and they are some of the highest-value tests you can write.

Running in CI

- name: Install Playwright browsers
  run: npx playwright install --with-deps

- name: Run tests
  run: npx playwright test

- uses: actions/upload-artifact@v4
  if: always()
  with:
    name: playwright-report
    path: playwright-report/
    retention-days: 7

if: always() matters — without it the report only uploads when tests pass, which is precisely when you do not need it.

Set retries: 2 in CI only, and treat a test that passes on retry as a bug to fix rather than a success. Retries are for hiding infrastructure noise, not for tolerating a flaky test forever.

How Many E2E Tests Should You Have?

Fewer than you think. E2E tests are slow and expensive to maintain, so spend them on flows where a failure would be genuinely costly: signing up, signing in, checkout, whatever your product’s core action is.

Everything else — validation rules, formatting, edge cases in business logic — belongs in unit tests, which run in milliseconds and tell you exactly what broke. A suite of 30 well-chosen E2E tests that always passes is worth more than 300 that fail randomly.

Frequently Asked Questions

Is Playwright better than Cypress?

For most teams, yes. Playwright runs real WebKit for Safari coverage, parallelises across workers out of the box, handles multiple tabs and origins without workarounds, and is free at any scale. Cypress has a friendlier interactive runner, which some teams value highly.

How do I stop Playwright tests being flaky?

Remove every fixed sleep, make each test independent and self-provisioning, mock third-party services, and disable animations in the test environment. Those four changes fix the large majority of flakiness.

Should I use data-testid attributes?

Only when nothing better exists. Prefer getByRole and getByLabel — they survive refactors and double as an accessibility check. A test id is a reasonable fallback for something like a total that has no accessible name.

How long should an E2E suite take?

Under ten minutes in CI, ideally under five. Beyond that people stop waiting for it and start merging around it. Parallel workers and reusing authentication state are the two changes with the biggest effect.

Can Playwright test authenticated pages?

Yes — sign in once in a setup project, save the storage state to a file, and have the other projects load it. Every test then starts already authenticated without repeating the login flow.

Does Playwright work with React, Vue or Svelte?

It drives a real browser, so the framework is irrelevant. Anything the browser renders, Playwright can test. There are also component-testing modes if you want to test components in isolation rather than a full page.

The Bottom Line

Install it, write a handful of tests for the flows that would cost you money if they broke, use role-based locators, reuse authentication state, and never write a fixed sleep. That gets you a suite people trust — which is the only kind worth having.

MD Rafikul Islam

Written by

MD Rafikul Islam is a software developer and editor of TechPulse. He writes about developer tools, hardware, AI, and practical technology decisions. Some articles are based on cited documentation and analysis rather than hands-on testing; readers should check each article for sources and testing disclosures. Corrections are welcome at rony.yf25@gmail.com.

✍️ Leave a Comment

Your email address will not be published. Required fields are marked *