────────────────────────────────────────
If you are building a SaaS product or any reasonably large web application, automated tests are essential. They protect the paths you have already defined and catch regressions on known flows.
The limitation is everything else. Complex products always contain edge cases, unusual sequences, and broken states that no one planned for. Exploratory testing exists to surface those unknowns.
Traditionally this work had to stay with a person. The value came from human curiosity and the ability to notice something unexpected. That approach is slow and expensive.
Playwright CLI combined with an AI coding agent changes the economics. You can now run structured exploration for hours while you continue working on other tasks.
────────────────────────────────────────
▶ What the Agent Does
The agent follows a mission you define:
• Opens the application
• Moves through the areas you specify
• Observes the current state of the page
• Records what it finds
• Stops only when it hits a hard blocker
You review the findings afterward and decide what deserves attention.
────────────────────────────────────────
▶ Setup: One Rules File
Create a file called `steps.md`. This becomes the mission brief. It should contain:
• How to use Playwright CLI
• Start URL and environment
• Scope: which areas or flows to explore and which to skip
• Credentials or test data (or where to find them)
• Rules for what to report and what to ignore
Example rules:
What to report:
• Functional problems (buttons that do nothing, wrong redirects, forms that fail to submit)
• Broken UI that blocks use
• Incorrect or missing copy on critical paths
• JavaScript errors that prevent actions
• Failed network requests on important APIs
• Accessibility blockers
• Dead ends (404s, infinite spinners, errors with no recovery)
What to ignore:
• Third-party analytics or tracking failures
• Benign console warnings from vendor scripts
• Purely cosmetic issues
• Cookie or consent banners unless they block the flow
Point the agent at `steps.md` and let it run. The agent creates and fills `report.md` as it works. ────────────────────────────────────────
▶ How the Agent Operates
For each step the agent:
1. Navigates to the relevant page
2. Takes a snapshot of the current state
3. Decides the next action
4. Performs the action
5. Takes another snapshot
7. Continues
It stops only on hard blockers such as missing credentials, captchas, unavailable data, or a completely down environment.
When the run finishes, you review `report.md`. ────────────────────────────────────────
▶ Why the Report File Matters
The agent writes to `report.md` after every meaningful step. You can open the file while the agent is still running and watch findings appear in real time. If the agent stops early, everything discovered up to that point is already saved. ────────────────────────────────────────
▶ Practical Tips
• Run multiple agents in parallel for different flows. Give each one its own rules file.
• Use headed mode on the first run of a new area. Later runs can be headless.
• Treat this as exploration, not as a replacement for deterministic automated tests. The goal is to surface problems you have not already encoded in your test suite.