Reliable Retry Logic for Flaky HTTP Tests

Flaky HTTP tests pass and fail without a meaningful code change. They may time out when a service is briefly overloaded, receive a 502 from a reverse proxy, or race an eventually consistent database. A carefully designed retry policy can absorb these temporary faults, but a careless one can hide real defects and make a test suite slower and less trustworthy.

The useful distinction is between recovering from a transient transport problem and repeating a request that is already known to be invalid. Retry logic belongs around a small, well-understood part of the test, with clear limits, diagnostics, and rules for safe HTTP methods. It should improve signal in continuous integration rather than turn every failure into a delayed pass.

Find the Real Source of Flakiness

Before adding a retry, collect the status code, response body, elapsed time, request method, target host, correlation ID, and exception type. A failed assertion that says “expected 200, got 503” needs a different response from a connection reset during TLS negotiation. Repeating both failures indefinitely produces noise instead of evidence.

Look for environmental patterns in the test history. Failures concentrated during a deployment may indicate service readiness or connection draining. Failures on an asynchronous endpoint may indicate an incorrect polling interval. A test that fails only on a shared CI runner may be affected by CPU contention, DNS, ephemeral port exhaustion, or an overloaded test database.

Australian teams may see different latency profiles when a service runs in Sydney but a CI worker is in Melbourne, Brisbane, or Singapore. NBN congestion, mobile-network variability, and cross-region cloud calls can expose timing assumptions that remain invisible on a developer’s local connection. Record the region and execution environment so a retry does not conceal a location-specific problem.

Retry Only Transient Failures

A retry condition should be explicit. Connection resets, socket timeouts, DNS lookup failures, HTTP 408, 425, 429, and many 5xx responses can be temporary. A 401 caused by an expired token, a 403 caused by a permission error, a 404 for a missing fixture, and a 422 validation response usually require a test or application fix rather than another request.

The HTTP method matters because repeating a request can change data. GET, HEAD, and OPTIONS are normally safe to repeat, provided the endpoint follows HTTP semantics. PUT and DELETE can often be retried when the operation is designed to be idempotent. POST needs stronger protection: use an idempotency key, a unique operation token, or a test endpoint whose duplicate behaviour is explicitly defined.

A useful predicate separates transport errors from assertion failures:

boolean retryable(int status) {
    return status == 408 || status == 425 || status == 429
        || status == 500 || status == 502 || status == 503
        || status == 504;
}

Do not classify every unexpected status as transient. If a service returns a stable 500 because a required field is missing, three more requests will only create additional logs and possibly duplicate side effects.

Use Bounded Backoff and Jitter

Immediate retries create a thundering herd. If 200 CI jobs receive a 503 and all retry after exactly one second, they can overload the recovering service again. Exponential backoff spaces out attempts: a common sequence is 250 milliseconds, 500 milliseconds, 1 second, and 2 seconds, with a maximum delay and a maximum attempt count.

Jitter adds a random adjustment to each delay. Full jitter chooses a random value between zero and the calculated backoff; equal jitter keeps half the backoff and randomises the other half. Full jitter is straightforward for tests and avoids synchronised retries. A retry budget of three attempts and a total deadline of 10 to 20 seconds is usually safer than an unlimited loop.

Respect the Retry-After header for rate limiting when its value is reasonable, while still enforcing a local maximum. This matters for public or shared services, where repeated test traffic can affect other users. It also supports responsible load behaviour under Australian privacy and consumer expectations: a test suite should not generate an avoidable traffic spike against a production-like environment.

A Java helper can keep the policy visible:

<T> T withRetry(Supplier<T> operation, Predicate<T> retryable,
                int maxAttempts, Duration timeout) {
    long deadline = System.nanoTime() + timeout.toNanos();

    for (int attempt = 1; attempt <= maxAttempts; attempt++) {
        try {
            T result = operation.get();
            if (!retryable.test(result) || attempt == maxAttempts) {
                return result;
            }
        } catch (RuntimeException ex) {
            if (attempt == maxAttempts) throw ex;
        }

        long limit = Math.min(2000L, 250L << (attempt - 1));
        long delay = ThreadLocalRandom.current().nextLong(limit + 1);
        if (System.nanoTime() + Duration.ofMillis(delay).toNanos() > deadline) {
            break;
        }
        sleep(delay);
    }
    throw new AssertionError("Retry deadline exceeded");
}

The production version should preserve the final exception or response and attach earlier attempts as suppressed diagnostics. A test framework extension can provide this behaviour consistently, while a small helper is often easier to audit in a Groovy or Spock specification.

Handle Asynchronous HTTP Work Correctly

Retries are frequently misused for eventual consistency. Suppose a test creates an order and immediately requests its status. Retrying the GET may be valid, but the real contract is a state transition, so the test should poll until the order reaches READY or a deadline expires. Each poll needs a fresh response, a bounded wait, and a message that explains the last observed state.

Avoid sleeping inside a broad retry around the entire scenario. If setup, POST, polling, and assertions are all repeated, the test may create multiple records or lose the original failure. Separate the operation that can be repeated from the assertions that must run once. Persist the resource ID and use it to check the same entity on every poll.

For systems using queues or webhooks, wait for a domain event or a test-controlled callback where possible. Time-based sleeps are especially fragile across AEST and AEDT changes, overnight deployments, and different CI locations. A monotonic clock should measure deadlines, while timestamps in logs should include timezone and offset for human diagnosis.

When an endpoint is protected by a gateway or security appliance, test the surrounding automation separately from the business flow. For example, signature update automation can have its own retry and verification tests, rather than being accidentally retried as part of an unrelated API scenario.

Protect Test Data and Test Isolation

A retry is safe only when the test state remains understandable. Generate a unique correlation ID and, where supported, an idempotency key for each logical operation. Reuse that key across retries so the server can return the original result instead of creating duplicate payments, users, or bookings.

Clean-up must be resilient without becoming destructive. If a test creates a resource and the first response is lost, a second attempt may receive a conflict because the resource already exists. The test should locate the resource by its unique key, verify its state, and then delete it if appropriate. Never use a broad “delete all” recovery step in a shared environment.

Test data can expose regional and market-specific behaviour. Australian addresses may contain state abbreviations such as NSW, VIC, and QLD, while phone numbers, postcode validation, GST calculations, and daylight-saving rules can differ from fixtures written for another market. Keep these details explicit and deterministic, especially when APIs validate addresses or calculate prices.

The Privacy Act 1988 and the Australian Privacy Principles make careless retry logging a real risk if responses contain names, addresses, tokens, or payment information. Redact sensitive fields before recording each attempt. Correlation IDs, status codes, latency, and a hashed resource identifier usually provide enough evidence without copying personal data into CI logs.

Verify the Policy in Continuous Integration

A retry policy needs tests of its own. Use a stub server such as WireMock, MockWebServer, or Hoverfly to return a controlled sequence: timeout, 503, 503, then 200. Assert the number of calls, delay bounds, header handling, final response, and preservation of the original error. Add cases for 400, 401, 404, and a non-idempotent POST to prove that they do not retry.

In CI, report the number of attempts and classify the recovered failure. A test that passes only after two retries should remain visible as a warning or metric. Track retry rate by test, endpoint, service version, runner, and region. A rising rate often gives earlier warning of capacity or reliability degradation than a rising hard-failure rate.

Keep the default policy conservative for pull requests and allow a separate, clearly named resilience suite to exercise longer outages. Teams operating across Sydney, Perth, and overseas cloud regions should compare latency distributions instead of selecting a single timeout based on the fastest developer laptop. The same suite should run against representative staging infrastructure, with rate limits and access controls matching the intended environment.

Practical Review Lists

Before merging retry logic, check:

When diagnosing a flaky failure, inspect:

Strategy Suitable use Main risk Sensible default
Immediate retry Very short connection blips Synchronized load spikes Avoid except for narrowly scoped transport errors
Fixed delay Simple polling against a controlled stub Every job retries together Use only with jitter or low concurrency
Exponential backoff with jitter Transient HTTP and network faults Can extend test duration Three attempts, capped delay, total deadline
Deadline-based polling Eventual consistency and async workflows May hide a broken state transition Poll one resource until a clear terminal state
Unbounded retry Almost no CI test case Hangs builds and masks defects Never use in automated tests

A reliable HTTP test is designed around a contract, not a hope that another request will succeed. Classify the failure, retry only safe transient conditions, use bounded backoff with jitter, preserve test data, and expose recovered failures through metrics. The key point to remember is that retry logic should make temporary infrastructure noise easier to diagnose while leaving genuine application defects impossible to ignore.