Testing JPA Repositories with @DataJpaTest and Testcontainers

JPA repositories sit close to the database, so tests that replace PostgreSQL with an in-memory H2 database can create misleading confidence. SQL generated by Hibernate, transaction boundaries, constraints, indexes, JSON columns, and database-specific functions may all behave differently in production. A repository test should verify the persistence behaviour that the application actually relies on.

Spring Boot’s @DataJpaTest provides a focused test slice, while Testcontainers supplies a disposable database running in a real container. Together, they create a practical middle ground between fast unit tests and slower full-system checks. This approach works well for Java and Groovy projects, including services deployed across Australian regions such as Sydney, Melbourne, and Brisbane.

Why Repository Tests Need A Real Database

A JPA repository is more than a Java interface. Its queries are interpreted by Hibernate and translated into SQL for a particular database engine. A JPQL query may pass against H2 but fail against PostgreSQL because of syntax, casting, case sensitivity, date functions, or an unsupported feature. Native queries are even more dependent on the production dialect.

Real persistence behaviour also includes details that mocks cannot prove. A mocked repository can verify that save() was called, but it cannot reveal whether a unique constraint rejects duplicate data, whether a lazy association is accessed after the session closes, or whether a transaction flushes changes at the expected point.

This matters when a team works with production data hosted in Australian cloud regions. A service in ap-southeast-2 may use PostgreSQL features, timezone settings, and collation rules that are absent from a local H2 setup. Running the same database engine during integration tests helps identify these differences before a release reaches customers in Sydney or Perth.

What @DataJpaTest Provides

@DataJpaTest loads a restricted Spring application context containing JPA repositories, entity mappings, Hibernate configuration, and related persistence components. It avoids starting unrelated web and messaging infrastructure, which keeps repository tests faster and easier to diagnose. By default, each test runs inside a transaction that is rolled back when the test completes.

A simple test might look like this:

@DataJpaTest
class OrderRepositoryTest {

    @Autowired
    private OrderRepository orderRepository;

    @Test
    void findsPendingOrdersForCustomer() {
        orderRepository.save(new Order("C-100", OrderStatus.PENDING));

        List<Order> orders =
                orderRepository.findByCustomerIdAndStatus("C-100", OrderStatus.PENDING);

        assertThat(orders).hasSize(1);
    }
}

The rollback behaviour is useful, but it can hide problems when application code depends on a commit. Events triggered after commit, database jobs, and visibility from another transaction require a more deliberate test. Calling flush() is also important when the assertion concerns database constraints or SQL execution rather than the in-memory persistence context.

@DataJpaTest may replace the configured data source with an embedded database when one is available. Testcontainers changes that arrangement by providing a real external database. The test remains focused on JPA while using the same database family as production.

Adding PostgreSQL With Testcontainers

A PostgreSQL container can be declared with the Testcontainers JUnit integration:

@Testcontainers
@DataJpaTest
class OrderRepositoryTest {

    @Container
    static PostgreSQLContainer<?> postgres =
            new PostgreSQLContainer<>("postgres:16-alpine")
                    .withDatabaseName("orders")
                    .withUsername("test")
                    .withPassword("test");

    @DynamicPropertySource
    static void databaseProperties(DynamicPropertyRegistry registry) {
        registry.add("spring.datasource.url", postgres::getJdbcUrl);
        registry.add("spring.datasource.username", postgres::getUsername);
        registry.add("spring.datasource.password", postgres::getPassword);
        registry.add("spring.datasource.driver-class-name",
                postgres::getDriverClassName);
    }
}

The container should use a version close to production. If the live service runs PostgreSQL 15, using PostgreSQL 16 in tests may conceal compatibility issues. Pinning the image tag also makes local runs and CI more reproducible than using a floating latest tag.

Recent Spring Boot versions can reduce boilerplate with service connection support, depending on the project’s Testcontainers and Boot versions. Whichever configuration style is used, the essential principle is the same: Spring’s datasource properties must point to the container before the JPA context is created.

Testcontainers requires a Docker-compatible runtime. Developers in Adelaide or Hobart may run Docker Desktop locally, while a CI agent in a Melbourne data centre might use a Linux Docker host. The test should fail with a clear infrastructure message when Docker is unavailable, rather than silently switching back to H2.

Keeping Schema And Data Predictable

Schema management is a major part of repository integration testing. If production uses Flyway or Liquibase, running the same migrations against the container is usually more valuable than relying on Hibernate’s automatic schema generation. It exercises table definitions, indexes, foreign keys, custom types, and migration ordering.

For a small test suite, @Sql scripts or repository fixtures can provide controlled data. Builders are often clearer than large JSON fixtures because they make the relevant fields visible:

Order order = OrderFixture.anOrder()
        .forCustomer("C-100")
        .withStatus(OrderStatus.PENDING)
        .createdAt(Instant.parse("2025-02-01T00:00:00Z"))
        .build();

Avoid sharing mutable entities between tests. A static fixture that is modified by one test can create order-dependent failures, especially when the persistence context retains managed objects. Save the fixture, call flush(), and clear the entity manager when the test needs to prove that data can be read back from the database rather than from Hibernate’s first-level cache.

Australia’s daylight-saving rules add another useful edge case. Sydney and Melbourne change clocks, while Brisbane does not, and Perth follows a different local time regime. Repository tests involving Instant, OffsetDateTime, reporting dates, or “today” queries should use an explicit clock and verify the database timezone assumptions.

Testing Approach Database Fidelity Speed Best Use Common Risk
Mocked repository None Very fast Service unit tests Persistence defects remain invisible
H2 with @DataJpaTest Low to moderate Fast Simple portable mappings SQL and dialect differences are missed
Testcontainers database High Moderate Repository and migration tests Requires container runtime and image management
Shared integration database High Variable Cross-service or staging checks Data isolation and test flakiness
Full application test High Slow Wiring and end-to-end behaviour Failures are harder to localise

Testing Queries, Constraints, And Transactions

Repository tests should describe observable persistence behaviour rather than repeat implementation details. A test for a derived query can verify filtering and ordering. A test for a native query should also cover null values, duplicate rows, pagination, and the shape of the returned projection where those cases matter.

Database constraints deserve explicit tests. Persist a valid entity, then attempt an invalid operation and call flush():

orderRepository.save(new Order("C-100", OrderStatus.PENDING));
orderRepository.save(new Order("C-100", OrderStatus.PENDING));

assertThatThrownBy(() -> orderRepository.flush())
        .isInstanceOf(DataIntegrityViolationException.class);

The exact exception chain can vary between Hibernate and the JDBC driver, so assertions should focus on the contract unless the application deliberately maps a specific vendor error. A test that only calls save() may pass because Hibernate has not sent the insert statement yet.

Transaction tests need special care because @DataJpaTest wraps each test in a transaction. If the production method uses REQUIRES_NEW, publishes an event after commit, or reads through a separate connection, a default slice test may not reproduce it. Use TestTransaction, an explicit transaction template, or a higher-level integration test for those cases.

Repository methods that use pessimistic locks, database functions, or concurrent updates should be tested with multiple transactions and a real database. These are poor candidates for mocks and often behave differently on H2. A focused concurrency test can reveal lock timeouts and isolation assumptions before a busy Friday arvo release window.

Avoiding Common Testcontainer Mistakes

Container startup is often the slowest part of the test suite. A static container shared by all tests in a class is generally more efficient than creating one for every method. Reuse across the entire developer machine can reduce startup time further, but reusable containers need careful cleanup and should not allow data from one test run to leak into another.

Context caching can also surprise teams. Spring may reuse an application context when test configuration matches, but changing dynamic properties or profiles can create several contexts and start more containers than expected. Keep container configuration in a shared base class or test configuration when that improves consistency, while preserving isolation where database state genuinely differs.

Do not depend on internet access during every test execution. Pull the required image in a controlled CI preparation step or use an approved internal registry where company policy requires it. Teams troubleshooting container startup should separate application failures from infrastructure failures; the same discipline used in application deployment troubleshooting helps identify whether the problem is Docker, networking, credentials, or the test itself.

Logs are part of diagnosis. Enable SQL logging temporarily when a query fails, inspect the container logs, and capture the database version in CI output. Excessive SQL logging should remain off for normal builds because it can expose sensitive fixture values and make failures difficult to read.

Integrating The Tests Into CI

A repository test suite should run locally and in continuous integration with as little environmental variation as possible. The build should declare the PostgreSQL image, Java version, migration tool, and Testcontainers dependencies explicitly. Maven and Gradle profiles can separate quick unit tests from container-backed persistence tests without allowing the latter to disappear from the standard verification lifecycle.

CI agents need enough CPU, memory, and disk for Docker images. This is particularly relevant for hosted runners serving teams across Australia, where a build may execute in a different region from the application’s primary deployment. The database container should still use UTC unless the application deliberately requires another setting; geographic location does not automatically justify local wall-clock assumptions.

A sensible pipeline runs unit tests first, then repository tests with Testcontainers, followed by broader API or end-to-end checks. When a repository test fails, its name should indicate the business rule, such as rejectsDuplicateExternalReference, rather than a vague label like databaseTest. Test reports, container logs, and migration output should be retained for failed builds.

A Practical Repository Testing Workflow

Start by listing the persistence behaviours that production depends on: query predicates, sorting, pagination, constraints, migrations, transaction boundaries, and database-specific features. Keep pure domain logic in fast unit tests, then use @DataJpaTest for repository behaviour and Testcontainers for the database fidelity that an embedded substitute cannot provide.

Use realistic but minimal fixtures, control time explicitly, flush when testing SQL effects, and choose a container image that matches production. Add a small number of tests for locks, commit behaviour, and timezone-sensitive queries rather than turning every repository test into a full system test.

The concrete next step is to replace one H2-backed repository test with @DataJpaTest, a pinned PostgreSQL Testcontainer, @DynamicPropertySource, and a flush() assertion against the real schema.