Skip to main content

Performance Testing Best Practices

This page lists good habits for performance testing. They make your numbers believable โ€” so that a "yes, it will handle the sale" is right โ€” and make your findings lead to fixes. They apply to SDETs, SREs and developers alike.

Each practice has:

  • Do โ€“ the good way.
  • Why โ€“ the reason in simple words.
  • An example, where it helps, from the practice lab.
How to use this page

Read it once after you finish Part 2 of the performance testing cheat sheet, then use the checklist at the end before sending any performance report.


Contentsโ€‹

  1. Agree Targets First
  2. Model Real Traffic
  3. Realistic Data
  4. Realistic Environments
  5. Healthy Load Generators
  6. Measure the Server, Not Just the Client
  7. Change One Thing at a Time
  8. Baselines and Trends
  9. Performance in CI
  10. Safety
  11. Short Checklist Before You Report

1. Agree Targets Firstโ€‹

In short: a result means nothing without a target agreed before the test.

  • Do write SLOs as numbers: load, percentile, time, error rate. Why: "p95 < 300 ms at 50 req/s" can pass or fail; "fast enough" can't.
  • Do agree targets with the business owner, not only engineers. Why: they decide what "good enough for the sale" means.
  • Do put the targets in the script as thresholds. Why: the test itself says pass or fail โ€” no debate after the run.

2. Model Real Trafficโ€‹

In short: load shaped like production gives results that predict production.

  • Do build the journey mix and peak rate from analytics or access logs. Why: 100% "add to cart" tests a system nobody uses.
  • Do include think time between steps. Why: without it, 100 virtual users act like thousands of real ones.
  • Do use an arrival-rate (open) executor for public web traffic. Why: real users keep arriving when the site is slow; a closed model backs off and hides the problem.
  • Do test the expected peak and the growth factor (ร—1.5, ร—2). Why: the plan has to survive next year's sale, too.

3. Realistic Dataโ€‹

In short: data size and variety change results more than most people expect.

  • Do test against production-like data volume (masked). Why: a query over 100 rows is instant; over 20 million it may need an index you forgot.
  • Do vary inputs: different users, products, search terms, including expensive ones. Why: repeated identical requests hit caches and look faster than reality.
  • Do give each virtual user its own account where the app locks per user. Why: 100 VUs logged in as one user test lock contention, not your workload.

4. Realistic Environmentsโ€‹

In short: know how the test environment differs from production, and say so.

  • Do match instance sizes, counts and config โ€” or state the ratio and scale carefully. Why: half the servers doesn't always mean half the capacity.
  • Do mock third parties that forbid load (payments, email, SMS). Why: their sandboxes throttle or ban you, and the results reflect them, not you.
  • Do keep the environment to yourself during the test. Why: someone else's batch job can look exactly like your bottleneck.

5. Healthy Load Generatorsโ€‹

In short: if the machine generating load is struggling, your numbers describe it, not the system.

  • Do watch CPU and memory on the load generator. Why: over ~80% CPU, the generator itself adds latency.
  • Do run generators close to the system (same region / network). Why: home Wi-Fi and the public internet add noise to every number.
  • Do distribute very large tests across several machines or a cloud service. Why: one laptop can't simulate a national sale.

6. Measure the Server, Not Just the Clientโ€‹

In short: client numbers say that it's slow; server numbers say why.

  • Do collect CPU, memory, connection pools, queue depths, GC and database metrics during every run. Why: in the lab, only /metrics showed 45 orders waiting for 5 connections.
  • Do keep graphs over time, not just end-of-run averages. Why: a leak or a slow start is a shape over time โ€” the soak test's memory went 4 โ†’ 153 MB while response times looked fine.
  • Do use traces for the slowest requests. Why: they show which call inside the request took the time.

7. Change One Thing at a Timeโ€‹

In short: tune in small steps and re-test after each, or you won't know what helped.

  • Do form a hypothesis first: "the pool limits orders to ~100/s". Why: a prediction you can test beats random tuning.
  • Do change one setting, re-run the same test, compare. Why: change pool size and caching together and you can't tell which fixed it.
  • Do repeat runs (at least 2โ€“3) before trusting a difference. Why: run-to-run noise of 5โ€“10% is normal.

In short: compare with the last result, not with an ideal.

  • Do store every run's summary (JSON) with the build/version. Why: "p95 went 52 โ†’ 180 ms in build 2.9" points straight at a change.
  • Do re-run the same baseline test after fixes. Why: proves the fix and becomes the new baseline.

9. Performance in CIโ€‹

In short: a short, stable performance gate on every deploy; big tests on a schedule.

  • Do run a few-minute gate test with thresholds after each staging deploy. Why: most regressions are caught the day they're introduced.
  • Do keep gate thresholds a little looser than SLOs and stable across runs. Why: a flaky gate gets switched off.
  • Do run full load, stress and soak tests before big releases and on a schedule. Why: they take too long for every change, but capacity drifts over time.

10. Safetyโ€‹

In short: a load test is a controlled attack โ€” have permission, a plan and a stop button.

  • Don't load-test systems you don't own or have written permission to test. Why: it can take a service down and may be illegal; to the provider it looks like a denial-of-service attack.
  • Do tell the on-call team and monitoring owners before big tests. Why: otherwise they'll treat your test as an incident.
  • Do use abortOnFail thresholds and ramp gradually. Why: you stop at the breaking point instead of far past it.
  • Don't test production without an agreed window, limits and a rollback plan.

11. Short Checklist Before You Reportโ€‹

  • Targets (SLOs) were agreed before the test and are in the script as thresholds
  • Workload model and think time come from real traffic data
  • Data volume and variety are production-like (and masked)
  • Environment differences from production are stated
  • Load generator stayed healthy (CPU, network) โ€” evidence kept
  • Server-side metrics collected for the whole run
  • Results repeated at least twice; noise understood
  • Bottleneck named with evidence; recommendations ranked
  • Compared with the previous baseline
  • Verdict is the first line of the report

Need more detail? Cheat sheet ยท Quick reference ยท Full guide