By Greg Margolin · September 24, 2026
The release passed its load test with room to spare. Two weeks later, on the first busy Monday of the quarter, the application slowed to a crawl. The team is puzzled: the test used more users than production ever sees.
This is one of the most common stories in performance engineering, and it usually isn't the tool's fault. The test measured something real. It just wasn't what production does.
A quick word on terms. A workload model describes the traffic a test imitates: how many requests arrive, of what kinds, at what pace. A virtual user (VU) is a simulated user in a load test. Latency percentiles such as the 95th and 99th (p95, p99) describe the slowest requests: p99 is the response time that 99% of requests beat.
1. The test's traffic slowed down when the system did. Many load tests use a fixed number of virtual users, each waiting for a response before sending the next request. This is called a closed workload model. When the system slows, the test automatically sends less traffic, exactly when real users would keep arriving. Real traffic behaves more like an open model, where arrivals don't wait for responses. The effect even has a name, coordinated omission, and it makes results look better than reality. Researchers at Carnegie Mellon showed back in 2006 how differently open and closed systems behave.
2. The average hid the problem. Google's Site Reliability Engineering book warns that averages obscure the slow tail of requests. Its example: with an average of 100 milliseconds, 1% of requests “might easily take 5 seconds.” Report percentiles, not averages.
3. The test data was too tidy. If every virtual user searches for the same few products, the caches do all the work. Production users are messier, and so is the load on the database.
4. The test was too short. Memory leaks, connection pools running dry and slowly filling queues show up over hours, not minutes. That's what a soak test, sustained load over a long period, is for.
Two more usual suspects: a test environment smaller than production, and third-party services replaced by stubs that always answer instantly.
AI is now built into the major performance tools. OpenText's Performance Engineering Aviator helps write and fix load test scripts and answers questions about test results in plain language. Tricentis NeoLoad connects AI assistants to its testing through the Model Context Protocol (MCP). AI can also help turn production logs into a first draft of a realistic workload model and flag unusual patterns in results.
The limits matter. OpenText's own documentation says there's no guarantee its AI answers are accurate. A workload model drawn from logs is a starting point to validate, not a prediction of the future. And an AI summary that points to a bottleneck is a hypothesis for an engineer to confirm.
Performance engineering has been at the core of GQP's work for years. Our Fractional Agentic AI Team brings AI into that practice through AI-Driven Performance Engineering: open workload models grounded in real traffic, percentile-based goals, soak tests where they matter, and AI-assisted analysis that experienced engineers confirm before anyone acts on it. See also our Performance Engineering service.
Start with a 15-minute conversation. We'll ask how you test performance today, then tell you honestly whether AI would help.
Prefer the phone? Call us at 888-477-5580, or 888-GQP-5580. Alternatively, complete our Contact Us form here.
Sources: Grafana k6 documentation, “Open and closed models” and “Load test types.” B. Schroeder, A. Wierman and M. Harchol-Balter, “Open Versus Closed: A Cautionary Tale,” USENIX NSDI 2006. G. Tene, “How NOT to Measure Latency.” Google, Site Reliability Engineering, chapters “Service Level Objectives” and “Monitoring Distributed Systems.” OpenText, Performance Engineering Aviator documentation. Tricentis, NeoLoad MCP announcement.