Generating realistic test data with Java Faker
Quality assurance teams in Brisbane, Sydney and Melbourne spend a disproportionate amount of time inventing names, addresses, and identifiers that feel plausible enough to exercise business rules without revealing themselves as obvious dummies. When the same address 123 Test Street, Testville shows up in a thousand test runs, the data starts to feel sterile, and so does the verification. The whole point of a robust test suite is to model real users, real orders, and real edge cases, which means the inputs have to look like they came from a real form submission.
Java Faker is a library ported from the Ruby Faker project that gives Java developers a fluent, idiomatic way to build fixtures. Anyone who has ever written new Name("John", "Smith") knows how quickly this gets stale, especially when you need hundreds of distinct customers for a load test, or a CSV of Australian addresses to seed a staging database. Java Faker solves this by exposing generators for nearly every domain a typical application touches: people, places, finance, internet, commerce, vehicles, books, and even esoteric categories like Rick and Morty quotes or Pokemon names. The library does not pretend to be a full data masking platform, yet it covers the surface area that most unit, integration, and contract tests actually need.
For Australian teams in particular, the value of generated fixtures has grown alongside the explosion of fintech, healthcare, and government platforms. Whether you are validating the ATO's tax file number format, populating a Medicare card lookup stub, or generating dummy accounts for the myGov single sign-on integration in a partner environment, you need data that respects local conventions. Hardcoded strings fail here because they cannot scale, and they would not catch the corner cases that real Australians hit, like a four-digit postcode in a remote area, a phone number on the Territory network, or a suburb you have probably never visited.
This article walks through setting up Java Faker, exploring its core API, using Australian locales, combining generators into realistic fixtures, and integrating them cleanly with existing test frameworks. Along the way you will see how it compares with alternatives, and you will get a handful of practical tips to keep the suite productive.
Why hardcoded test data fails QA pipelines
Hardcoded test data works fine until it does not. A common pattern looks like String name = "John Smith"; String postcode = "2000"; - readable, fast, and completely useless when the production system rejects names containing apostrophes, postcodes shorter than four digits, or Unicode characters. A Sydney-based team shipping an order management module discovered it the hard way when a customer's surname D'Costa crashed a regex designed for ASCII-only input. The bug had been hidden because every test used ASCII names.
Hardcoded data cannot represent the long tail of edge cases that real users produce. A realistic fixture should cover diacritics, double-barrelled surnames, leading whitespace, multi-word company names, ABN-format business identifiers, and BSB numbers from regional Australian banks. None of these edge cases are exciting on their own, yet together they represent the surface area where bugs actually live. Synthetic data tools are designed to chase that surface, and chasing that surface is what turns fragile suites into confident ones.
There is also a maintenance cost. Every time the production team adjusts a field constraint - say, requiring that phone numbers match the Australian E.164 format +61 - the team has to find and update every literal across hundreds of test classes. Generated data flows through a single seam, and that seam updates in one place when the validation rule changes.
Setting up Java Faker in Maven or Gradle builds
Java Faker ships as a single Maven artifact, which makes adoption almost trivial. In a Maven project the dependency looks like the snippet below:
<dependency>
<groupId>com.github.javafaker</groupId>
<artifactId>javafaker</artifactId>
<version>1.0.2</version>
<scope>test</scope>
</dependency>
For Gradle, the equivalent block is testImplementation 'com.github.javafaker:javafaker:1.0.2'. Both build systems scope the dependency to test because Faker output should never reach production. After adding the dependency, a one-line Faker faker = new Faker(); is enough to start generating data. There is no external service, no Docker container, and no network call - everything happens in-process, which keeps tests fast and CI-friendly.
For Australian-specific output, the constructor becomes new Faker(new Locale("en", "AU")). Locale selection controls everything from name pools to postcode ranges to currency symbols. The same deterministic approach that testing asynchronous callbacks with CountDownLatch and ExecutorService describes for callback-heavy code applies here: a controlled seed gives reproducible fixtures.
Core generators for everyday scenarios
The Faker API is organised around a small set of top-level generators. faker.name().fullName() produces a believable person. faker.address().fullAddress() produces a street, city, state and postcode. faker.internet().emailAddress() returns a plausible email. faker.phoneNumber().phoneNumber() returns a phone number formatted for the active Locale, which for en-AU returns a Sydney or Melbourne landline or a mobile starting with 04. Each call returns a fresh value, which means there is no need to reset state between test methods.
Beyond the basics, Faker offers domain-specific generators that map neatly to typical production verticals. faker.finance().creditCard() returns a Visa, Amex or Mastercard number that passes the Luhn check. faker.finance().iban() returns a valid IBAN, and faker.finance().bic() returns a SWIFT/BIC code. For Australian banking scenarios, faker.finance().ausBank() generates a BSB and account number pair, which is handy when verifying NPP payment flows or the PayID resolution endpoint. The library also offers faker.commerce().productName(), faker.commerce().price(), and faker.commerce().department(), which together cover a standard product catalogue surface.
For teams that need bulk generation, the Faker instance is thread-safe in the practical sense, and that practical sense is that every call returns an independent value. A common pattern in load tests wraps Faker in a supplier and streams records through it: IntStream.range(0, 10_000).mapToObj(i -> buildCustomer(faker)).forEach(repo::save);. This produces ten thousand distinct customers in seconds, which is far cheaper than maintaining a CSV fixture and far more flexible than a one-off data script.
Locale support and Australian specifics
Locale is where Java Faker earns its keep for non-American teams. The library ships with over fifty locales, including en-AU, which adjusts generators to Australian conventions. faker.address().city() returns a real Australian city like Sydney, Melbourne or Brisbane. faker.address().state() returns New South Wales, Victoria or Queensland. faker.address().zipCode() returns a four-digit postcode, and if you need a remote one, it is happy to return an NT postcode starting with 08.
Other locale choices extend the language, and the language reflects the social conventions of the target market. faker.name().firstName() for en-AU draws from a pool that includes Charlotte, Olivia and Ava. faker.phoneNumber().cellPhone() for en-AU returns a mobile starting with 04 followed by eight digits, which matches the standard Australian mobile format. For teams writing tests for the NBN rollout or the ACCC's compliance suite, this locale support is the difference between a fixture that looks real and a fixture that looks American.
There is one subtle gotcha: not every generator respects locale. faker.internet().emailAddress() always uses English conventions, because email addresses do not really have a locale. The same applies for uuid() and for lorem() words. Knowing which generators are locale-aware lets you predict the output, and predicting the output is what turns a flaky test into a stable one. For Australian-specific values that Faker does not cover natively, you can plug in a custom Resolver or simply post-process the string with String.replace calls.
Combining generators into realistic test fixtures
Realistic scenarios require realistic objects, and object-to-object relationships are where most test data efforts fall apart. A customer has an address, an address has a postcode, a postcode belongs to a state, an order belongs to a customer, and an order line belongs to a product. Naive code tends to mix these concerns. The pattern that scales is the Object Mother or Test Data Builder pattern, where each domain object has a dedicated builder that wraps Faker calls.
The pattern scales because each builder centralises the constraints for its domain. A CustomerBuilder might call faker.name().fullName() for the customer, faker.address().fullAddress() for the shipping address, and faker.internet().emailAddress() for the email. The builder can override any field with a sensible default. For Australian fintech, the builder might additionally validate that the BSB matches the state in the address, which catches a class of bugs where state and BSB drift apart in production.
For load testing the sports betting platforms that Australians use on weekends, builders can be combined with Stream.generate(builder::build) to produce an unbounded stream of fixtures. Streaming unbounded streams quickly exhausts memory, and exhausting memory is what turns a load test into an outage. Limiting with .limit(100_000) keeps the test bounded. The same approach works for the PayID resolution endpoint, the eftpos settlement endpoint, and the BPAY biller code validation service, which are common surfaces in Australian banking integrations.
Performance, seeding and integration tips
Performance is rarely a concern with Java Faker because the cost per call is microseconds. The cost adds up only when the suite grows, and the suite grows when each test runs dozens of generations. A suite with 50 million Faker calls can add real seconds to a CI run. Seeding Random with new Random(seed) does not work directly on Faker, yet wrapping Faker in a deterministic generator does work for golden-master tests.
Integration with JUnit, Spock, and TestNG is straightforward. JUnit 5's @ParameterizedTest works well with feeds from ArgumentsProvider, which generates arrays of Faker-driven objects. Spock's data tables work the same way, with the added benefit of Groovy's with block for partial overrides. For teams working with fast cash-out services in the Australian gaming and sports betting market, the same fixture pattern keeps withdrawal flows testable under load.
For non-functional testing, Faker output can be wrapped in a memoising provider so that the same fixture is reused across iterations. Reusing the same fixture pays for itself when the suite needs stable snapshots, and stable snapshots are what make screenshot tests reliable.
Comparing Java Faker with alternatives
| Library | Active maintenance | Locales | Java version | Builder support | Best suited for |
|---|---|---|---|---|---|
| Java Faker | Community-maintained, active | 50+ | Java 8+ | Manual via Object Mother | Broad ecosystem use |
| DataFaker | Active, JDK14+ | 50+ | Java 17+ | Built-in DataGenerator API |
Modern JDK17+ teams |
| JFairy | Inactive | Limited | Java 8 | None | Legacy projects only |
| MockNeat | Active | 20+ | Java 11+ | Fluent chain | High-throughput tests |
| GenerateData | Active, Perl-based | N/A | N/A | N/A | SQL or CSV bulk generation |
The choice depends on the JDK version, the language ecosystem, and the team's appetite for fluent builders. DataFaker is the modern alternative for JDK17+ teams. MockNeat wins for pure throughput. JFairy is best avoided unless you have a legacy dependency. GenerateData is a separate beast that produces SQL or CSV from a config file, and is best treated as a complementary tool rather than a replacement.
Practical tips for keeping the suite fast
- Scope Faker to the
testconfiguration in Maven and Gradle so the library never ships to production. - Wrap builders in a
ThreadLocalwhen running parallel test classes, because each thread needs its ownFakerinstance. - Seed
Randomwhen you need golden-master output for snapshot tests. - Prefer
Faker.expression("#{name.firstName} #{name.lastName}")for templated strings to compose multiple providers in a single call. - Combine Faker with DataFaker only when one library covers a gap the other does not, because two libraries add two layers of maintenance.
The thing to remember is that realistic test data is not a luxury. It is the raw material that turns a brittle suite into a confident suite, and a confident suite is the foundation of fast shipping.