Data Collection and DBMS
XML Data Models, Querying and Transformation
PGCP-BDA
XML document
An XML document is a hierarchical tree of elements, attributes, text and namespaces represented using tagged markup.
well-formed XML
Well-formed XML has one root element, properly nested and closed tags, quoted attribute values and legal names and character encoding.
XML namespace
A URI-qualified naming mechanism that prevents element and attribute name collisions between XML vocabularies.
DTD and XML Schema
DTD and XML Schema describe permitted XML structure; XML Schema also supports namespaces and rich data types.
XPath
XPath selects nodes and computes values from an XML tree using location paths, axes, node tests and predicates.
XQuery
XQuery is an expression language for querying and constructing XML values, with FLWOR expressions supporting iteration, filtering.
XSLT
XSLT transforms an XML source tree into XML, HTML, text or another representation using template-matching rules and XPath expressions.
XML parsing
Reading XML while checking well-formed structure and exposing nodes or events to an application.
XML database storage
Storage that preserves or maps XML hierarchy and supports retrieval through XML-aware paths or query languages.
XML transformation
Conversion of XML into another XML vocabulary, HTML or text, commonly by applying XSLT templates.
XML Documents and Schemas
XML represents information as a tree containing one document element, nested elements, attributes and text nodes. Names are case-sensitive and tags must close with proper nesting. Escapes represent markup characters in text. Namespaces qualify names with URI-bound prefixes so vocabularies can coexist.
A DTD or XML Schema validates allowed elements, attributes, order and constraints. Well-formedness only checks XML syntax; schema validity checks an application vocabulary. XML Schema can define numeric, date and restricted string types. Secure parsers should restrict external entity resolution because an untrusted document can otherwise expose files or cause excessive resource use.
XPath, XQuery and XSLT
XPath selects nodes through paths and predicates. /catalog/book follows an absolute child path while //book searches descendants. @id selects an attribute and [price > 500] filters nodes. The context node changes the meaning of a relative path. Query prefixes must be bound even if the source uses a default namespace.
XQuery builds on XPath and uses FLWOR expressions: for, let, where, order by and return. It can filter, join and construct XML. XSLT applies templates to transform a source tree into XML, HTML or text. xsl:value-of emits a value while xsl:apply-templates continues template processing. Transformation changes representation while schema validation checks conformance.
Relational and XML Mapping
Element-centric design represents most values as elements while attribute-centric design uses attributes for compact metadata. Attributes cannot contain nested structure and their order is not significant. Mixed content allows text interleaved with child elements and is common in documents but awkward for relational shredding.
Mapping XML into relations requires decisions about identity, order and repetition. A repeating element normally becomes a child table with a foreign key and sequence position. Optional elements map to nullable values only when absence and empty content have the same intended meaning. Storing the original XML alongside extracted columns can preserve information that a relational projection does not represent.
An XML index may accelerate path and value queries at added storage and update cost. Query plans should be examined for actual paths and selectivity. Serializing a result requires a declared character encoding and correct escaping. Canonicalization produces a standardized byte representation for comparison or signatures but does not make different information models equivalent.
DOM parsing builds a tree in memory and provides random navigation. Streaming parsers process events in document order with far lower memory use but cannot freely move backward. Large documents favor streaming when the transformation can be expressed in one pass. Parser limits should bound document size, nesting and entity expansion for untrusted input.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.