Load Testing a WebSocket Server with k6
WebSocket applications keep a connection open so that the server can push data as soon as it becomes available. That model suits chat, live dashboards, multiplayer games, trading screens, notifications and collaborative tools, but it also changes how performance needs to be measured. A test that sends one HTTP request per iteration cannot represent thousands of persistent clients.
k6 is a practical option for generating these connections from a script that can live beside the rest of a performance test suite. It can open sockets, authenticate users, send and receive messages, record custom metrics and apply thresholds in CI. The useful result is more than a connection count: it shows whether the server maintains acceptable latency and resource usage as concurrency grows.
WebSocket load testing also exposes problems that are easy to miss in functional testing. A server may accept 20,000 connections yet struggle when every client subscribes to several channels. It may deliver messages promptly during quiet periods but develop a backlog during a broadcast. Network limits, garbage collection, connection cleanup and load-balancer timeouts can all affect the outcome.
The examples below use a realistic approach for an Australian service, where customers may connect from Sydney, Melbourne, Brisbane, Perth and regional areas over different broadband and mobile networks. The goal is to build a repeatable test, choose meaningful measures and interpret results without confusing open sockets with useful capacity.
| Approach | Best for | Main measurement | Limitation |
|---|---|---|---|
| HTTP request load test | REST endpoints and polling clients | Request rate and HTTP latency | Does not model persistent sessions |
| WebSocket connection test | Handshake and idle connection capacity | Successful upgrades, open sockets, disconnects | Can hide message-processing pressure |
| Message workload test | Chat, events and live updates | Send-to-receive latency and delivery rate | Requires realistic message behaviour |
| Mixed protocol test | Applications using REST plus sockets | End-to-end user journey | More complex data and load coordination |
Model the real WebSocket session
Start by describing what one virtual user does. A typical session might obtain a token over HTTPS, connect to wss://, subscribe to one or more topics, send a heartbeat, receive events and close cleanly. Those actions should appear in the k6 script in roughly the same order as they occur in production.
The connection rate matters as much as the final number of virtual users. Ten thousand sockets opened over a minute create a very different workload from ten thousand sockets opened in ten seconds. The latter can expose TLS, authentication, load-balancer and connection-pool limits rather than normal application behaviour.
Use representative message sizes and frequencies. A support chat client may send only a few messages per minute, while a market-data screen may receive updates several times per second. If the Australian audience is spread across time zones, model the busy period in the relevant local zone rather than assuming that all customers arrive simultaneously at midnight UTC.
Build a k6 WebSocket script
The k6/ws module provides a straightforward event-driven API. A basic script can validate the HTTP 101 upgrade response, send a subscription message, count received events and close the socket after a controlled session.
import ws from 'k6/ws';
import { check } from 'k6';
import { Counter, Trend } from 'k6/metrics';
const received = new Counter('messages_received');
const deliveryLatency = new Trend('delivery_latency', true);
export const options = {
scenarios: {
clients: {
executor: 'constant-vus',
vus: 100,
duration: '2m',
},
},
thresholds: {
checks: ['rate>0.99'],
messages_received: ['count>1000'],
delivery_latency: ['p(95)<500'],
},
};
export default function () {
const url = 'wss://example.test/realtime';
const params = { tags: { service: 'realtime' } };
const response = ws.connect(url, params, function (socket) {
socket.on('open', function () {
socket.send(JSON.stringify({
type: 'subscribe',
channel: 'alerts',
token: __ENV.ACCESS_TOKEN,
}));
});
socket.on('message', function (payload) {
const event = JSON.parse(payload);
received.add(1);
if (event.sentAt) {
deliveryLatency.add(Date.now() - event.sentAt);
}
});
socket.on('error', function (error) {
console.error(`socket error: ${error.error()}`);
});
socket.setTimeout(function () {
socket.close();
}, 60000);
});
check(response, {
'WebSocket upgrade succeeded': (r) => r && r.status === 101,
});
}
The newer k6/websockets API is worth considering for projects that need the WebSocket standard API or more modern asynchronous patterns. Whichever module is selected, keep connection lifecycle logic clear: record the opening result, handle errors, track messages and close sockets deliberately rather than relying on the test runner to terminate them.
Choose scenarios that expose bottlenecks
A staged ramp is usually more informative than a single large burst. Begin with a small number of users, hold the load long enough for the application to settle, then increase connections in steps. k6 executors such as ramping-vus, constant-vus and arrival-rate executors support different shapes, although persistent sockets require careful thought about how long each virtual user remains active.
Separate the main workload into scenarios where practical. An idle-connection scenario measures memory and file-descriptor consumption. A subscription scenario measures topic registration and initial snapshots. A broadcast scenario measures fan-out, serialisation and delivery. A reconnect scenario tests what happens when a node or network path disappears and thousands of clients attempt to return.
For an Australian deployment, a test run from Sydney may make Melbourne traffic look healthy while understating Perth latency. Run load generators in more than one region when geography affects the service, or add controlled network delay to approximate remote clients. NBN connections, corporate proxies and mobile networks do not behave identically, so a single low-latency data-centre run should not represent every customer.
Measure more than open connections
The handshake success rate is essential, but it is only the first checkpoint. Track active sockets, failed upgrades, connection duration, unexpected closures, reconnects and authentication failures. Server-side metrics should include CPU, memory, garbage collection, event-loop delay, file descriptors, network throughput and the number of connections assigned to each node.
Message latency needs a precise definition. Add a timestamp when the server publishes an event and subtract it from the timestamp when the client receives it. Use percentiles such as p95 and p99 rather than averages, because a small number of very slow customers can be important in a real-time product.
Delivery counts also need context. A received-message counter cannot prove that every expected event arrived unless the payload includes a sequence number or event identifier. The test can then detect gaps, duplicates and out-of-order messages. This is especially important for notifications or operational alerts, where a socket remaining open does not mean that the application is functioning correctly.
Handle authentication, data and cleanup
Avoid putting one fixed token into every virtual user unless that is genuinely how the service works. Shared credentials can trigger rate limits or conceal per-user state problems. Generate test identities in advance, obtain short-lived tokens through an API setup step, or use a controlled pool of credentials with clear ownership.
Test data should support realistic subscriptions. If every user listens to the same channel, the result may describe a broadcast storm rather than normal usage. Distribute users across popular and less active topics, and include users who change subscriptions during a session. Keep payloads valid and varied enough to exercise parsing, authorisation and persistence paths.
Clean shutdown is part of the workload. Close sockets after a known duration, unsubscribe where the protocol requires it and test abrupt termination separately. A graceful close exercises normal application code; a dropped connection exercises cleanup, presence updates and broker removal. Both paths matter when users move between Wi-Fi and mobile networks during an arvo commute.
Run the test safely in delivery pipelines
A performance test should be repeatable enough to run before a release, while larger capacity tests can run on a schedule. Store the k6 script with application code, version its thresholds and label metrics by scenario, region, message type and service version. In continuous delivery, a short smoke load can catch a broken upgrade path without attempting to prove maximum capacity on every commit.
Protect the environment from accidental production traffic. Use a dedicated tenant, test accounts, an explicit user-agent and a recognisable message prefix. Agree on limits with operations before starting a long run. A test that unexpectedly triggers customer notifications or fills a shared broker is an operational incident, not a useful benchmark.
Treat the load generator as part of the system under observation. Check its CPU, network and socket limits so it does not become the bottleneck. Multiple smaller generators are often preferable to one large machine, particularly when simulating clients across Australia. Use a time source and dashboard that make it easy to align k6 data with server logs, broker metrics and cloud alerts.
Validate the surrounding operational path
WebSocket reliability depends on infrastructure outside the application process. Verify idle timeouts on reverse proxies, sticky-session behaviour where it is required, TLS termination, connection draining during deployment and broker failover. A rolling release may leave existing sockets attached to old nodes or close them all at once, creating a reconnect surge.
Certificate and domain checks belong in the same operational picture. A failed certificate can prevent every WebSocket handshake before the application receives a request; an automated SSL expiry checker is useful alongside the load test rather than as a replacement for it. Test the complete wss:// path, including the certificate chain and hostname used by clients.
Record the conditions for each run: k6 version, script revision, generator location, instance size, deployment version, connection ramp and message rate. If a run from Melbourne passes but one from Perth shows a p99 spike, that distinction should remain visible in the result. Capacity is a property of a particular architecture, region and workload, not a permanent number attached to the server.
Turn findings into a release decision
Define pass and fail criteria before generating load. For example, require 99.5% successful upgrades, p95 event delivery below 500 milliseconds, no unexplained message gaps and a bounded reconnect rate during a node restart. Select limits that reflect the product’s promise: a collaborative editor needs different latency from an occasional notification stream.
Use a progression of tests to locate the first meaningful failure. If idle connections consume memory but messages remain fast, investigate per-connection state. If upgrades fail while existing sessions remain healthy, inspect authentication, proxy limits and file descriptors. If delivery latency rises with stable CPU, examine broker throughput, network queues and backpressure.
The practical result should be a capacity statement that includes workload and conditions: “This deployment supports 8,000 connected clients receiving two 2 KB events per second, with p95 delivery below 400 milliseconds and a five-minute ramp in Sydney.” That statement is far more useful than saying the server handled 8,000 sockets.
A sound k6 test therefore combines connection lifecycle checks, realistic event traffic, regional awareness and server-side observability. Start with a small script, prove that its messages represent genuine user behaviour, then increase concurrency gradually until a measurable limit appears. The practical takeaway is to report WebSocket capacity in terms of connections, message rate, latency and failure behaviour together.