Skip to main content

Performance Testing Milestones & Mini-Projects

In short: six projects on the practice lab, each with tasks, expected results and a folded answer key. Every number below comes from a real run on a laptop; yours will differ a little, but the patterns โ€” where it breaks, what grows โ€” will match.

How to use this page

Start the lab (node --expose-gc app/server.js) before each milestone, and restart it between tests so every run starts clean. First read the milestone in the Roadmap. Write your prediction before each run โ€” comparing prediction with result is how you learn to reason about performance.

Contentsโ€‹


Milestone 1: Smoke testโ€‹

Practises: k6 basics, checks, thresholds, reading the summary, percentiles by hand.

#TaskExpected result
1Write tests/smoke.js: 1 VU, 10 s, products + search + order each iteration, sleep(1)Runs, all checks โœ“
2Thresholds: every check passes, no failed requestsBoth โœ“, exit code 0
3Find the slowest endpoint from the summaryOrders (~50 ms)
4Add a check that products returns 20 itemsโœ“
5By hand: average, p50 and p90 of 20 21 22 22 23 24 25 26 30 900 ms111.3 / 23.5 / 30
6Stop the lab and run againChecks โœ—, threshold โœ—, exit code 99
Answer key
tests/smoke.js
// tests/smoke.js โ€” 1 user, every endpoint once a second: "does it work at all?"
import http from 'k6/http';
import { check, sleep } from 'k6';

const BASE_URL = __ENV.BASE_URL || 'http://localhost:3333';

export const options = {
vus: 1,
duration: '10s',
thresholds: {
checks: ['rate==1.0'], // every check must pass
http_req_failed: ['rate==0'],
},
};

export default function () {
check(http.get(`${BASE_URL}/api/products`), { 'products 200': (r) => r.status === 200 });
check(http.get(`${BASE_URL}/api/search?q=product`), { 'search 200': (r) => r.status === 200 });
check(http.post(`${BASE_URL}/api/orders`), { 'order 201': (r) => r.status === 201 });
sleep(1);
}
    โœ“ products 200
โœ“ search 200
โœ“ order 201
http_req_duration..............: avg=24.01ms min=495ยตs med=19.68ms max=51.88ms p(90)=51.74ms p(95)=51.79ms
http_reqs......................: 30 2.794269/s

Task 3: http_req_duration mixes all three endpoints; max โ‰ˆ 52 ms comes from orders (each holds a "DB connection" for 50 ms). Tag requests (Milestone 2) to see each endpoint separately.

Task 4: '20 products': (r) => r.json().length === 20.

Watch out for: reading iteration_duration (~1.07 s) as response time โ€” it includes the sleep(1).

Try it: remove sleep(1). How many requests per second does one VU now make? Why is that unrealistic?


Milestone 2: Load test at peakโ€‹

Practises: workload models, arrival-rate executors, tags, per-endpoint thresholds, test data.

Workload model (given): peak 50 req/s โ€” browse 30/s, search 12/s, order 8/s. Targets: products p95 < 100 ms, search p95 < 300 ms, order p95 < 200 ms, errors < 1%.

#TaskExpected result
1One scenario per journey with constant-arrival-rate for 30 s3 scenarios run in parallel
2Search terms from data/search-terms.json via SharedArrayVaried queries
3Tag each request with name; one threshold per endpoint3 duration thresholds + 1 error threshold
4Run itAll 4 thresholds โœ“
5Double every rate and re-run; note which endpoint degrades firstSee answer key
Answer key
tests/load.js
// tests/load.js โ€” the expected peak: a realistic mix of journeys at a fixed arrival rate
import http from 'k6/http';
import { check, sleep } from 'k6';
import { SharedArray } from 'k6/data';

const BASE_URL = __ENV.BASE_URL || 'http://localhost:3333';
const terms = new SharedArray('terms', () => JSON.parse(open('../data/search-terms.json')));

export const options = {
scenarios: {
browse: { executor: 'constant-arrival-rate', exec: 'browse', rate: 30, timeUnit: '1s', duration: '30s', preAllocatedVUs: 20 },
search: { executor: 'constant-arrival-rate', exec: 'search', rate: 12, timeUnit: '1s', duration: '30s', preAllocatedVUs: 20 },
order: { executor: 'constant-arrival-rate', exec: 'order', rate: 8, timeUnit: '1s', duration: '30s', preAllocatedVUs: 20 },
},
thresholds: {
http_req_failed: ['rate<0.01'],
'http_req_duration{name:products}': ['p(95)<100'],
'http_req_duration{name:search}': ['p(95)<300'],
'http_req_duration{name:order}': ['p(95)<200'],
},
};

export function browse() {
const r = http.get(`${BASE_URL}/api/products`, { tags: { name: 'products' } });
check(r, { 'products 200': (res) => res.status === 200 });
}

export function search() {
const q = terms[Math.floor(Math.random() * terms.length)]; // test data from a file
const r = http.get(`${BASE_URL}/api/search?q=${encodeURIComponent(q)}`, { tags: { name: 'search' } });
check(r, { 'search 200': (res) => res.status === 200 });
}

export function order() {
const r = http.post(`${BASE_URL}/api/orders`, null, { tags: { name: 'order' } });
check(r, { 'order 201': (res) => res.status === 201 });
sleep(0.5); // think time after buying
}

Result of task 4:

    โœ“ 'p(95)<200' p(95)=68.66ms          โ† order
โœ“ 'p(95)<100' p(95)=11.33ms โ† products
โœ“ 'p(95)<300' p(95)=21.48ms โ† search
โœ“ 'rate<0.01' rate=0.00%

Task 5: at 16 orders/s the pool (100/s capacity) is still fine; search at 24/s adds CPU load but stays well under 300 ms on most laptops. The load test at ร—2 still passes โ€” which tells you there's headroom, but not how much. That's Milestone 3's job.

Watch out for: putting all three journeys in one default function with if (Math.random() < 0.6) and a VU-based executor โ€” it works, but the arrival rate then depends on response times (closed model).

Try it: change the order scenario to ramping-arrival-rate from 8 to 40/s. Does anything change?


Milestone 3: Find the orders capacityโ€‹

Practises: Little's Law, stress testing, custom metrics, abortOnFail, bottleneck evidence, re-testing a fix.

#TaskExpected result
1Predict the max orders/s from the code: 5 connections, 50 ms each~100/s
2Stress test: orders from 20 โ†’ 60 โ†’ 100 โ†’ 160/s, order_time Trend, abort when p95 > 500 msAborts partway
3During the run, watch /metricspoolInUse 5, poolWaiting climbing
4Write the finding: bottleneck, evidence, capacity3 sentences
5Predict, then set POOL_SIZE = 10, restart, re-runPasses all stages
Answer key
tests/stress.js
// tests/stress.js โ€” keep raising the order rate until the system breaks, and see HOW it breaks
import http from 'k6/http';
import { check } from 'k6';
import { Trend } from 'k6/metrics';

const BASE_URL = __ENV.BASE_URL || 'http://localhost:3333';
const orderTime = new Trend('order_time', true); // custom metric, in ms

export const options = {
scenarios: {
rising_orders: {
executor: 'ramping-arrival-rate',
startRate: 20, timeUnit: '1s',
preAllocatedVUs: 50, maxVUs: 400,
stages: [
{ target: 60, duration: '10s' }, // below capacity
{ target: 100, duration: '10s' }, // around capacity
{ target: 160, duration: '10s' }, // past capacity
],
},
},
thresholds: {
order_time: [{ threshold: 'p(95)<500', abortOnFail: true, delayAbortEval: '5s' }], // stop once it's clearly broken
},
};

export default function () {
const r = http.post(`${BASE_URL}/api/orders`);
orderTime.add(r.timings.duration);
check(r, { 'order 201': (res) => res.status === 201 });
}

Task 2 result (original lab):

    โœ— 'p(95)<500' p(95)=614.59ms
http_reqs......................: 1784 68.608476/s
level=error msg="thresholds on metrics 'order_time' were crossed; โ€ฆ stopping test prematurely"

Task 3 โ€” /metrics at the end: {"poolInUse":5,"poolWaiting":48,โ€ฆ}.

Task 4 โ€” a good finding: "Orders are limited to about 100 per second by the database connection pool (5 connections ร— 50 ms). Above that, requests queue: 48 were waiting and p95 reached 615 ms, so the test stopped. Expected peak is 8 orders/s, so there is about 12ร— headroom today."

Task 5 result with 10 connections โ€” no abort, all stages pass:

    โœ“ 'p(95)<500' p(95)=52.78ms
http_reqs......................: 2499 83.170404/s

New predicted capacity: 10 รท 0.05 = 200/s, above this test's 160/s top stage.

Watch out for: http_reqs of ~69/s in the aborted run is not the capacity โ€” it's an average over the whole run, including the slow start. The capacity is where latency started climbing (~100/s).

Try it: extend the stages to 250/s with 10 connections. Does it break close to 200/s, as predicted?


Milestone 4: A sale-start burstโ€‹

Practises: spike tests, scenarios with startTime, per-scenario thresholds, recovery.

#TaskExpected result
1Three scenarios on search: 5 VUs (10 s) โ†’ 150 VUs (10 s) โ†’ 5 VUs (10 s), sleep(0.2)Runs 30 s
2Thresholds: before and after p95 < 50 ms; during p95 < 1 s; errors < 1%See answer key
3Explain why search suffers so muchCPU-bound on one thread
4Recommend two fixesSee answer key
Answer key
tests/spike.js
// tests/spike.js โ€” a sudden burst (a sale starts), then back to normal: does it recover?
import http from 'k6/http';
import { check, sleep } from 'k6';

const BASE_URL = __ENV.BASE_URL || 'http://localhost:3333';

export const options = {
scenarios: {
before: { executor: 'constant-vus', vus: 5, duration: '10s' },
spike: { executor: 'constant-vus', vus: 150, duration: '10s', startTime: '10s' },
after: { executor: 'constant-vus', vus: 5, duration: '10s', startTime: '20s' },
},
thresholds: {
'http_req_duration{scenario:before}': ['p(95)<50'],
'http_req_duration{scenario:spike}': ['p(95)<1000'], // slower during the spike is acceptableโ€ฆ
'http_req_duration{scenario:after}': ['p(95)<50'], // โ€ฆbut it must recover afterwards
http_req_failed: ['rate<0.01'],
},
};

export default function () {
check(http.get(`${BASE_URL}/api/search?q=product`), { 'search 200': (r) => r.status === 200 });
sleep(0.2);
}
    โœ“ 'p(95)<50' p(95)=37.31ms       โ† before
โœ— 'p(95)<1000' p(95)=12.5s โ† during
โœ“ 'p(95)<50' p(95)=32.64ms โ† after โ€” recovered
โœ“ 'rate<0.01' rate=0.00%

Task 3: every search burns CPU for ~25โ€“30 ms, and Node.js runs this code on one thread. 150 users ร— 5 searches/s each would need far more CPU time per second than one thread has, so requests queue โ€” for up to 12 s.

Task 4: cache results for common queries; move search to a proper search index; run several processes (cluster) or instances behind a load balancer; rate-limit per user during a sale.

Watch out for: "0% errors" reads as success, but users waited 12 seconds โ€” most would have left. Always report latency and errors.

Try it: add a fourth scenario that runs GET /api/products during the spike. Is browsing also slow? (Hint: one thread serves everything.)


Milestone 5: Find and fix the memory leakโ€‹

Practises: soak tests, setup/teardown, resource metrics, proving a fix.

#TaskExpected result
1Soak test on /api/report: 5 VUs, 30 s; log heap before and after via /metricsHeap grows by ~150 MB
2Run it against the other endpoints instead โ€” which one leaks?Only /api/report
3Find the leak in server.jsThe cache array
4Fix it (keep only the newest 50 entries), restart, re-runHeap stays small
Answer key
tests/soak.js
// tests/soak.js โ€” steady, modest load for a long time; watch memory, not just speed
import http from 'k6/http';
import { check, sleep } from 'k6';

const BASE_URL = __ENV.BASE_URL || 'http://localhost:3333';

export const options = {
vus: 5,
duration: __ENV.DURATION || '30s', // real soak tests run for hours: DURATION=4h
};

export function setup() {
return { heapBefore: http.get(`${BASE_URL}/metrics`).json('heapMB') };
}

export default function () {
check(http.get(`${BASE_URL}/api/report`), { 'report 200': (r) => r.status === 200 });
sleep(0.2);
}

export function teardown(data) {
const heapAfter = http.get(`${BASE_URL}/metrics`).json('heapMB');
console.log(`heap before: ${data.heapBefore} MB, after: ${heapAfter} MB`);
}
leaky:  level=info msg="heap before: 4 MB, after: 153 MB"
fixed: level=info msg="heap before: 4 MB, after: 14 MB"

The fix in server.js:

      cache.push(new Array(25_000).fill(Math.random()));
if (cache.length > 50) cache.shift(); // fixed: keep only the newest 50

A real fix would use a proper cache with a size limit and expiry (an LRU cache).

Watch out for: without --expose-gc, heap numbers include garbage that hasn't been collected yet โ€” a fixed leak can still look like it grows. That's why the lab collects garbage before reporting.

Try it: run the leaky version with DURATION=2m. Extrapolate: how long until it reaches 2 GB?


Milestone 6: A deploy gate and a reportโ€‹

Practises: CI gates, handleSummary, baselines, regression detection, reporting.

#TaskExpected result
1tests/ci-gate.js: 5 VUs, 20 s; thresholds per endpoint; handleSummary writes results/summary.json and prints one lineOne-line result, JSON saved
2A GitHub Actions workflow that runs the gatePasses actionlint
3Simulate a regression: change dbQuery(50) to dbQuery(150); run the gateStill passes โ€” see answer key
4Change it to dbQuery(250); run the gateFails, exit code 99
5Write the full performance report for the lab (all milestones)Verdict first; outline in the quick reference
Answer key
tests/ci-gate.js
// tests/ci-gate.js โ€” a short performance gate for every deploy; writes a JSON summary for the pipeline
import http from 'k6/http';
import { check } from 'k6';

const BASE_URL = __ENV.BASE_URL || 'http://localhost:3333';

export const options = {
vus: 5,
duration: '20s',
thresholds: {
http_req_failed: ['rate<0.01'],
'http_req_duration{name:products}': ['p(95)<100'],
'http_req_duration{name:order}': ['p(95)<200'],
},
};

export default function () {
check(http.get(`${BASE_URL}/api/products`, { tags: { name: 'products' } }), { 'products 200': (r) => r.status === 200 });
check(http.post(`${BASE_URL}/api/orders`, null, { tags: { name: 'order' } }), { 'order 201': (r) => r.status === 201 });
}

export function handleSummary(data) {
const p95 = (name) => data.metrics[`http_req_duration{name:${name}}`].values['p(95)'].toFixed(1);
const line = `products p95=${p95('products')}ms, order p95=${p95('order')}ms, errors=${(data.metrics.http_req_failed.values.rate * 100).toFixed(2)}%`;
return {
stdout: `\n${line}\n`, // short line in the CI log
'results/summary.json': JSON.stringify(data, null, 2), // full data kept as an artifact
};
}
.github/workflows/perf.yml
# .github/workflows/perf.yml โ€” run the performance gate after each deploy to staging
name: performance-gate

on:
workflow_dispatch:
push:
branches: [main]

jobs:
k6:
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v5
- uses: grafana/setup-k6-action@v1
- name: Run the gate (fails the job if a threshold fails)
run: k6 run tests/ci-gate.js
env:
BASE_URL: ${{ vars.STAGING_URL }}
- uses: actions/upload-artifact@v4
if: ${{ !cancelled() }}
with:
name: k6-summary
path: results/summary.json
baseline:      products p95=0.6ms, order p95=52.0ms, errors=0.00%      exit 0
dbQuery(150): products p95=1.5ms, order p95=153.9ms, errors=0.00% exit 0 โ† 3ร— slower, still "passes"
dbQuery(250): products p95=2.7ms, order p95=254.6ms, errors=0.00% exit 99

Task 3 is the lesson: a threshold of 200 ms lets a 3ร— slowdown through. Keep the JSON from each run and compare with the previous baseline (for example, fail or warn if p95 grows by more than 25%) โ€” thresholds catch disasters; baselines catch regressions.

The report's first lines, for example:

Verdict:    Meets targets at expected peak (50 req/s); orders capacity โ‰ˆ 100/s (โ‰ˆ12ร— headroom).
Risks: Search is CPU-bound (p95 12.5 s during a 150-user burst).
Memory leak in /api/report (+150 MB in 30 s) โ€” fixed and verified.

Watch out for: handleSummary replaces k6's normal end-of-test summary on screen โ€” the โœ“/โœ— lines disappear, but the exit code still reports threshold failures.

Try it: add a step that compares results/summary.json with a stored baseline.json and fails if order p95 rose by more than 25%.


Final projectโ€‹

Performance-test an application you run yourself โ€” your own API, or an open source app started locally in Docker (for example OWASP Juice Shop on localhost:3000). Never load-test someone else's system without written permission.

Done when:

  • Workload model written down (journeys, mix, peak, think time) with its source
  • SLOs agreed with yourself in writing before testing, as thresholds
  • Smoke, load, stress, spike and soak scripts, each run at least twice
  • Server-side evidence for every finding (resource metrics, logs)
  • One bottleneck fixed (or configured) and the improvement proven with a re-run
  • CI gate workflow and a stored baseline
  • Report with the verdict first โ€” proof you can deliver the performance testing service