Core and Web-based Java
Hibernate, JPA, Mapping, Queries and Persistence Contexts
PGCP-AC
Object-relational mapping connects an object model to relational tables, but the two models have different ideas about identity, relationships, inheritance and navigation. JPA defines standard persistence concepts and APIs; Hibernate implements those contracts and adds provider-specific facilities. Productive use requires more than annotations: the developer must understand entity state, transaction scope, association ownership, SQL generation, fetching and the persistence context.
1. ORM and the impedance mismatch
Relational databases store rows in tables, relate them through keys and query sets. Java programs work with objects, references, inheritance, collections and behaviour. ORM describes mappings between these models:
Java entity Relational structure
------------ --------------------
class table
field/property column
object identity primary key
reference foreign key
collection related rows/join table
The mapping is not a replacement for SQL knowledge. Generated SQL still determines locks, indexes, constraints, round trips and performance. ORM is most useful when the domain model and transactional use cases are designed alongside the schema.
2. JPA and Hibernate
JPA, now Jakarta Persistence in newer platforms, defines annotations, interfaces, query language, lifecycle rules and mapping semantics. Hibernate ORM is a provider that implements those contracts and offers its own APIs.
| Standard concept | Hibernate-oriented equivalent |
|---|---|
EntityManagerFactory | SessionFactory |
EntityManager | Session |
| JPQL | HQL is a related Hibernate query language |
Factory objects are heavyweight, thread-safe application-level facilities and are normally created once per persistence unit. EntityManager or Session represents a short unit of interaction and must not be shared arbitrarily between concurrent threads.
Older applications use javax.persistence; newer Jakarta versions use jakarta.persistence. The source namespace and provider version must agree.
3. Entity requirements and identity
An entity represents persistent identity:
@Entity
@Table(name = "book")
public class Book {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(nullable = false)
private String title;
protected Book() {}
public Book(String title) {
this.title = title;
}
}
An entity needs an identifier and a no-argument constructor with at least protected visibility. Provider proxies and enhancement can impose further constraints on final classes or methods.
Database identity, object reference identity and Java equals are different notions. Equality based only on a database-generated ID is difficult before persistence because the ID may be null. Choose an equality strategy deliberately and avoid changing hash-relevant state while an entity is stored in a HashSet.
4. Field and property access
JPA determines the default access style from where mapping annotations appear. Annotations on fields select field access; annotations on getters select property access.
Field access lets the provider read state directly and keeps mapping close to fields. Property access invokes getters and setters, which can be useful but may unexpectedly include logic. Avoid mixing styles accidentally. @Access can override the strategy deliberately.
@Transient excludes a field or property from persistence. Java's transient keyword concerns serialization and is not the clearest way to express ORM mapping intent.
Column metadata can declare names, lengths, nullability, uniqueness hints, precision and scale, but actual database migrations and constraints remain authoritative.
5. Identifier strategies
@GeneratedValue strategies include AUTO, IDENTITY, SEQUENCE and TABLE. Choice affects portability, batching and when an ID becomes available.
- IDENTITY relies on an insert-generated value and can restrict insert batching.
- SEQUENCE obtains values from a database sequence and can preallocate ranges.
- TABLE simulates generation through a table and may contend.
- AUTO lets the provider choose based on capabilities.
Natural identifiers have business meaning; surrogate identifiers are storage identities without such meaning. Composite keys use @EmbeddedId or @IdClass and require stable equality.
Never assume IDs are consecutive or safe to expose as authorization. Gaps arise from allocation, rollback and concurrency.
6. Basic values, embeddables and converters
Basic Java values map to columns. An embeddable groups value fields without independent entity identity:
@Embeddable
public class Address {
private String city;
private String postalCode;
}
@Embedded
private Address address;
The Address columns live with the owning entity unless overrides specify otherwise. If Address needs its own lifecycle, references and ID, it should be an entity instead.
@ElementCollection maps a collection of basic or embeddable values to a dependent table. AttributeConverter<X,Y> converts a domain value to a supported database representation. Converter logic should be deterministic and preserve null and equality meaning.
7. Persistence context and identity map
An EntityManager contains a persistence context: a set of managed entities and their snapshots or tracking information. Within one context, loading the same entity type and primary key returns the same managed Java instance:
Book first = em.find(Book.class, 10L);
Book second = em.find(Book.class, 10L);
System.out.println(first == second); // normally true in this context
This first-level cache is mandatory and belongs to the context. It avoids duplicate managed representations and supports dirty checking. It is not a general query-result cache: a JPQL query may still execute SQL even when matching entities are already managed.
clear() detaches all managed entities, while detach(entity) detaches one. contains(entity) tests whether a particular instance is managed.
8. Entity lifecycle states
An entity moves among:
- transient/new: constructed but not managed and not represented as persistent state;
- managed/persistent: tracked by the active context;
- detached: has persistent identity but is no longer managed by that context;
- removed: managed and scheduled for deletion.
new object --persist--> managed --remove--> removed
|
close/clear/detach
↓
detached
detached --merge--> managed copy
State is relative to a particular context. An object detached from one EntityManager may have a managed counterpart in another.
9. Persist, find, reference, remove and merge
persist(entity) makes a new entity managed and schedules insertion. Calling it for an existing detached identity is conceptually incorrect and can fail.
find(Type.class, id) returns the entity or null and may consult the first-level cache. getReference can return a lazily initialized reference useful when only identity is needed, but accessing missing data may fail later.
remove(entity) schedules deletion of a managed entity. A detached entity generally must be found or merged first.
merge(detached) copies state into a managed instance and returns that instance:
Book managed = em.merge(detachedBook);
// continue with managed, not detachedBook
The argument does not become managed automatically. Ignoring the return value is a common source of later detached updates.
10. Dirty checking
The provider detects changes to managed state:
Book book = em.find(Book.class, id);
book.setTitle("Revised title");
No explicit update call is required. At flush, Hibernate compares or tracks changes and issues SQL. Dirty checking applies only while the entity remains managed.
Keep entity setters and lifecycle callbacks free from surprising external side effects. Bulk JPQL updates bypass normal managed-entity dirty checking and can leave context state stale; clear or synchronize deliberately afterward.
11. Flush versus commit
Flush synchronizes pending context changes with the database by issuing SQL. Commit completes the transaction and makes its effects durable according to database guarantees.
A flush can occur:
- explicitly through
flush(); - before transaction commit;
- before certain queries under automatic flush mode;
- at provider-defined points allowed by configuration.
Flush can reveal constraint violations before commit, but it does not commit. A later rollback still reverses uncommitted database changes. Do not equate “SQL appeared in logs” with “transaction succeeded.”
12. Transaction scope
Entity updates should occur inside a transaction:
EntityTransaction tx = em.getTransaction();
try {
tx.begin();
Book book = em.find(Book.class, id);
book.setTitle(title);
tx.commit();
} catch (RuntimeException ex) {
if (tx.isActive()) tx.rollback();
throw ex;
}
In Jakarta EE or Spring, declarative transaction management normally replaces manual demarcation. A transaction should surround one business unit, including all repository operations needed for its invariant.
Do not keep a transaction open while waiting for user input or remote calls unless carefully designed; long transactions hold connections, locks and snapshots.
13. Relationship mappings
Cardinality annotations include:
@ManyToOne: many child rows reference one parent;@OneToMany: one parent exposes many children;@OneToOne: one row corresponds to one related row;@ManyToMany: both sides can have many, usually through a join table.
@ManyToOne(fetch = FetchType.LAZY, optional = false)
@JoinColumn(name = "department_id", nullable = false)
private Department department;
The foreign-key-holding side is commonly the owning side. The ownership concept means “which mapping writes the relationship,” not which object has business importance.
Many-to-many mappings are convenient but often hide a relationship that has attributes such as enrollment date or status. Model such a join row as its own entity when it has business meaning.
14. Bidirectional ownership
The inverse collection uses mappedBy to name the owning property:
@OneToMany(mappedBy = "department",
cascade = CascadeType.ALL,
orphanRemoval = true)
private List<Employee> employees = new ArrayList<>();
Changing only the inverse collection may not update the foreign key. Maintain both sides:
public void addEmployee(Employee employee) {
employees.add(employee);
employee.setDepartment(this);
}
Helper methods keep the Java object graph consistent immediately and ensure the owning side carries the persistence change.
15. Cascades and orphan removal
JPA cascade options propagate operations such as PERSIST, MERGE, REMOVE, REFRESH and DETACH from one entity to associated entities. CascadeType.ALL combines them.
Cascade is an ORM operation rule. It is different from database ON DELETE CASCADE, which the database enforces independently.
orphanRemoval = true schedules removal when a dependent child is removed from its owning relationship. Use it only when the child cannot meaningfully exist independently. Cascading REMOVE from many-to-many associations can delete shared entities and is usually unsafe.
16. Fetching and lazy proxies
Lazy fetching defers association loading until accessed. Eager fetching requests immediate availability but does not guarantee one SQL join; it may still cause additional queries.
A lazy association accessed after its persistence context closes can cause a lazy-initialization failure:
Book book = repository.find(id); // context closes
book.getAuthors().size(); // may require unavailable load
Design fetch plans around the use case. Fetch needed data inside the transaction, use a fetch join, an entity graph or a DTO projection. Keeping the EntityManager open through view rendering hides query behaviour and broadens transaction/data-access boundaries.
17. The N+1 query problem
N+1 occurs when one query loads N parent rows and later access causes one additional query per parent:
1 query: select all departments
N queries: load employees for each department
It creates latency and load that grow with the result count. Detect it through SQL logging, query counters, tracing and realistic integration tests.
Possible solutions include fetch joins, entity graphs, batch fetching, subselect fetching or DTO projections. Loading every relationship eagerly is not a solution; it can create oversized joins, duplicate rows and unnecessary data.
18. JPQL and HQL
JPQL queries entity types and persistent attributes, not physical table and column names:
List<Book> books = em.createQuery(
"select b from Book b where b.title like :pattern",
Book.class)
.setParameter("pattern", prefix + "%")
.getResultList();
Parameters separate values from query syntax. HQL supports JPQL-style entity queries plus Hibernate extensions.
Path navigation may create joins according to the query. Use explicit joins when join type and aliases matter. A join fetch initializes an association in the same query but changes the fetched graph and can duplicate root rows when collections are joined.
19. Criteria and named queries
Criteria constructs typed query structure programmatically:
CriteriaBuilder cb = em.getCriteriaBuilder();
CriteriaQuery<Book> query = cb.createQuery(Book.class);
Root<Book> book = query.from(Book.class);
query.select(book)
.where(cb.equal(book.get("status"), status));
It is valuable for dynamic predicates, though verbose and still vulnerable to misspelled attribute strings unless a metamodel is used.
A named query associates stable query text with a name, commonly through annotations or metadata. It centralizes reuse and may be validated early. A named query is not automatically faster; execution still depends on SQL, plans, indexes, parameters and data.
20. Pagination and projections
setFirstResult and setMaxResults provide offset pagination. Large offsets can be expensive because the database still locates and skips earlier rows. Keyset pagination uses the last ordered key and often scales better.
Every page needs deterministic ordering, commonly with a unique tie-breaker:
order by createdAt desc, id desc
DTO projections select only required data:
select new com.acme.BookSummary(b.id, b.title)
from Book b
where b.status = :status
They reduce managed state and avoid accidental lazy traversal for read-only screens. Fetch-joining a collection together with pagination can yield incorrect or inefficient results; use a two-step ID query or another appropriate plan.
21. Locking and concurrent updates
Optimistic locking uses a version field:
@Version
private long version;
Updates include the old version in their condition. If another transaction changed the row, the affected count is zero and the provider reports an optimistic-lock failure. The application can report a conflict or retry an idempotent unit with a bounded policy.
Pessimistic locks ask the database to lock rows and can reduce conflicts but increase blocking and deadlock risk. Use them only when the use case justifies it. JVM synchronization cannot protect data from other processes or service instances.
22. First-level and second-level caches
The first-level cache is the persistence context and is always present. It provides identity and tracks managed state.
A second-level cache is optional and provider-configured. It can share entity or collection data across contexts. A query cache, when enabled, stores query result identifiers rather than replacing entity caching.
Caching helps read-mostly stable data but introduces invalidation, staleness, memory and cluster concerns. It does not repair inefficient N+1 traversal, missing database indexes or poor queries.
23. Inheritance mapping
JPA supports:
- SINGLE_TABLE: one table with a discriminator; fast polymorphic reads but many nullable columns;
- JOINED: base and subclass tables joined for subclass state; normalized but join-heavy;
- TABLE_PER_CLASS: separate concrete tables; polymorphic queries may require unions.
Choose from schema constraints and query patterns. Inheritance is not mandatory for code reuse; embeddables and composition may model data more clearly.
24. Lifecycle callbacks and auditing
Callbacks such as @PrePersist, @PostPersist, @PreUpdate and @PostLoad run around persistence events. They can set timestamps or validate local invariants.
Keep callbacks lightweight and deterministic. Avoid remote calls, hidden repository queries or major business workflows in callbacks because timing can depend on flush and providers. Auditing frameworks can populate created-by, updated-by and timestamp fields, but the source of user and clock data should remain testable.
Practical considerations
| Mistake | Correct approach |
|---|---|
| Sharing one EntityManager between threads | Scope it to one unit of work |
| Ignoring merge's return value | Continue with the returned managed instance |
| Assuming flush commits | Commit the transaction explicitly |
| Updating only the inverse association side | Maintain the owning side and both objects |
| Applying CascadeType.ALL everywhere | Choose operations from lifecycle ownership |
| Accessing lazy data after context closure | Fetch required data within the boundary |
| Making all mappings eager to fix N+1 | Design query-specific fetch plans |
| Querying table names in JPQL | Use entity and property names |
| Paginating an unordered query | Add deterministic unique ordering |
| Treating cache as a query cure | Measure SQL and fix access patterns |
Worked persistence trace
- A transaction begins and
findloads Book 10 into the context. - A second
findfor Book 10 returns the same managed instance. - Code changes its title; dirty checking records managed-state change.
- A JPQL query may trigger flush, sending UPDATE SQL.
- Flush has not committed; rollback can still undo it.
- Commit completes the transaction.
- Closing the EntityManager detaches the Book.
- Later
mergereturns a managed instance in a new context; the old argument remains detached.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.