By David Alves · September 30, 2026
Every team that runs automated tests against a web application knows the Monday-morning ritual. A developer renamed a button, moved a field or reorganized a page, and a stack of tests failed overnight. Nothing is actually broken. The tests just can't find what they are looking for, and someone spends the morning repairing scripts instead of testing anything new.
“Self-healing” tests promise to end that ritual, and many of them genuinely help. But a test that fixes itself can also fix away the very problem it was supposed to catch.
A quick word on terms. A locator (or selector) is the instruction a test uses to find something on the screen, such as “the button whose ID is submit-order.” A flaky test is one that passes and fails on the same code, with nothing changed.
When a locator stops matching, a self-healing tool looks for the element that most resembles the one it remembers. The open-source tool Healenium, for example, compares the page's current structure with the last version that worked, scores the likely candidates and uses the best match. It also produces a report so a person can confirm the fix.
Newer tools go further. Since version 1.56 (October 2025), Microsoft's Playwright testing framework includes AI “test agents”: a planner that explores an application and writes a test plan, a generator that turns the plan into tests, and a healer that runs failing tests and proposes repairs.
This is useful work. But notice the words: “most resembles,” “best match,” “proposes.” A heal is a well-informed guess, not a verification.
1. It finds the wrong element. Two buttons look alike. The “Submit Order” button was removed by mistake, and the tool happily clicks “Save Draft” instead. The test passes. No order was ever submitted.
2. It hides a real regression. Sometimes an element is missing because the release really did break it. Playwright's own documentation notes that its healer may skip a test when it believes the functionality itself is broken. That is sensible behavior, as long as a person sees it. A quietly healed or skipped test in a nightly report is easy to miss.
3. It wears down trust in the test suite. Google has reported that almost 16% of its tests showed some flakiness, and that developers sometimes dismissed real failures as “just flaky.” Silent heals give people one more reason to wave away a red result.
Underneath all three is an old problem testers call the test oracle problem: knowing what the correct result should be. AI can write test steps and repair locators. Deciding what “correct” means still comes from your requirements and your people.
GQP has built and run enterprise test automation for years. Our Fractional Agentic AI Team brings modern AI into that practice: we set up AI-assisted test creation and self-healing with review built in, connect every heal to your approval process, and keep the suite healthy as your application changes. We also help teams test new features as they are built, through Shift-Left Automation. Your people decide what “correct” means. We take on the maintenance.
Start with a 15-minute conversation. We'll ask how your test automation runs today and where it breaks, then tell you honestly whether AI would help.
Prefer the phone? Call us at 888-477-5580, or 888-GQP-5580. Alternatively, complete our Contact Us form here.
Sources: Healenium documentation, “How Healenium works.” Playwright documentation, “Test Agents,” and Playwright v1.56.0 release notes (October 6, 2025). J. Micco, “Flaky Tests at Google and How We Mitigate Them,” Google Testing Blog, May 27, 2016. E. T. Barr, M. Harman, P. McMinn, M. Shahbaz and S. Yoo, “The Oracle Problem in Software Testing: A Survey,” IEEE Transactions on Software Engineering 41(5), 2015. Stryker Mutator documentation.