Self-Healing Tests Can Heal the Wrong Thing

About This Post

By David Alves · September 30, 2026

Every team that runs automated tests against a web application knows the Monday-morning ritual. A developer renamed a button, moved a field or reorganized a page, and a stack of tests failed overnight. Nothing is actually broken. The tests just can't find what they are looking for, and someone spends the morning repairing scripts instead of testing anything new.

“Self-healing” tests promise to end that ritual, and many of them genuinely help. But a test that fixes itself can also fix away the very problem it was supposed to catch.

A quick word on terms. A locator (or selector) is the instruction a test uses to find something on the screen, such as “the button whose ID is submit-order.” A flaky test is one that passes and fails on the same code, with nothing changed.

How Self-Healing Actually Works

When a locator stops matching, a self-healing tool looks for the element that most resembles the one it remembers. The open-source tool Healenium, for example, compares the page's current structure with the last version that worked, scores the likely candidates and uses the best match. It also produces a report so a person can confirm the fix.

Newer tools go further. Since version 1.56 (October 2025), Microsoft's Playwright testing framework includes AI “test agents”: a planner that explores an application and writes a test plan, a generator that turns the plan into tests, and a healer that runs failing tests and proposes repairs.

This is useful work. But notice the words: “most resembles,” “best match,” “proposes.” A heal is a well-informed guess, not a verification.

Three Ways a Heal Goes Wrong

1. It finds the wrong element. Two buttons look alike. The “Submit Order” button was removed by mistake, and the tool happily clicks “Save Draft” instead. The test passes. No order was ever submitted.

2. It hides a real regression. Sometimes an element is missing because the release really did break it. Playwright's own documentation notes that its healer may skip a test when it believes the functionality itself is broken. That is sensible behavior, as long as a person sees it. A quietly healed or skipped test in a nightly report is easy to miss.

3. It wears down trust in the test suite. Google has reported that almost 16% of its tests showed some flakiness, and that developers sometimes dismissed real failures as “just flaky.” Silent heals give people one more reason to wave away a red result.

Underneath all three is an old problem testers call the test oracle problem: knowing what the correct result should be. AI can write test steps and repair locators. Deciding what “correct” means still comes from your requirements and your people.

Using AI in Test Automation Without Losing the Safety Net

  • Every heal gets reviewed. Treat a healed locator like any other code change: keep a record of the old locator, the new one and a screenshot, and have someone approve it before it becomes permanent.
  • Critical checks don't heal quietly. For checkout, payments or anything regulated, fail the build or require sign-off instead of healing automatically.
  • Flaky tests get fixed, not retried. Quarantine them and find the cause, rather than rerunning until the build turns green.
  • Test the tests. Mutation testing deliberately plants small bugs in the code and checks whether your tests catch them. It is a practical way to see whether AI-written tests actually test anything.
  • Let AI do the typing, not the judging. Drafting tests from user stories, suggesting locator repairs and spotting visual differences are all good uses. Approving what “right” looks like stays with your team.

How GQP Helps

GQP has built and run enterprise test automation for years. Our Fractional Agentic AI Team brings modern AI into that practice: we set up AI-assisted test creation and self-healing with review built in, connect every heal to your approval process, and keep the suite healthy as your application changes. We also help teams test new features as they are built, through Shift-Left Automation. Your people decide what “correct” means. We take on the maintenance.

How You'd Know It's Working

  • Fewer mornings spent repairing broken scripts.
  • Every healed test has a visible, approved record.
  • Real defects get caught, not healed.
  • Your team trusts a red build again.

Ready to Talk?

Start with a 15-minute conversation. We'll ask how your test automation runs today and where it breaks, then tell you honestly whether AI would help.

Prefer the phone? Call us at 888-477-5580, or 888-GQP-5580. Alternatively, complete our Contact Us form here.

Sources: Healenium documentation, “How Healenium works.” Playwright documentation, “Test Agents,” and Playwright v1.56.0 release notes (October 6, 2025). J. Micco, “Flaky Tests at Google and How We Mitigate Them,” Google Testing Blog, May 27, 2016. E. T. Barr, M. Harman, P. McMinn, M. Shahbaz and S. Yoo, “The Oracle Problem in Software Testing: A Survey,” IEEE Transactions on Software Engineering 41(5), 2015. Stryker Mutator documentation.