The XMLStreamReader allows forward, read-only access to XML. It is designed
to be the lowest level and most efficient way to read XML data.
The XMLStreamReader is designed to iterate over XML using
next() and hasNext(). The data can be accessed using methods such as getEventType(),
getNamespaceURI(), getLocalName() and getText();
The next() method causes the reader to read the next parse event.
The next() method returns an integer which identifies the type of event just read.
Parsing events are defined as the XML Declaration, a DTD,
start tag, character data, white space, end tag, comment,
or processing instruction. An attribute or namespace event may be encountered
at the root level of a document as the result of a query operation.
The following table describes which methods are valid in what state.
If a method is called in an invalid state the method will throw a
java.lang.IllegalStateException.
Valid methods for each state
Event Type
All States
START_ELEMENT
ATTRIBUTE
NAMESPACE
END_ELEMENT
CHARACTERS
CDATA
COMMENT
SPACE
START_DOCUMENT
END_DOCUMENT
PROCESSING_INSTRUCTION
ENTITY_REFERENCE
DTD
The following is a code sample to read an XML file containing multiple
“myobject” sub-elements. Only one myObject instance is kept in memory at
any given time to keep memory consumption low:
1var fileReader : FileReader = new FileReader(file, "UTF-8");2var xmlStreamReader : XMLStreamReader = new XMLStreamReader(fileReader);34while (xmlStreamReader.hasNext())5{6 if (xmlStreamReader.next() == XMLStreamConstants.START_ELEMENT)7 {8 var localElementName : String = xmlStreamReader.getLocalName();9 if (localElementName == "myobject")10 {11 // read single "myobject" as XML12 var myObject : XML = xmlStreamReader.getXMLObject();1314 // process myObject15 }16 }17}1819xmlStreamReader.close();20fileReader.close();
Returns the current value of the parse event as a string, this returns the string value of a CHARACTERS event, returns the value of a COMMENT, the replacement value for an ENTITY_REFERENCE, the string value of a CDATA section, the string value for a SPACE event, or the String value of the internal subset of the DTD.
Returns the normalized attribute value of the attribute with the namespace and localName If the namespaceURI is null the namespace is not checked for equality
Returns the current value of the parse event as a string, this returns the string value of a CHARACTERS event, returns the value of a COMMENT, the replacement value for an ENTITY_REFERENCE, the string value of a CDATA section, the string value for a SPACE event, or the String value of the internal subset of the DTD.
Reads a sub-tree of the XML document and parses it as XML object.
The stream must be positioned on a START_ELEMENT. Do not call the method
when the stream is positioned at document's root element. This would
cause the whole document to be parsed into a single XML what may lead to
an out-of-memory condition. Instead use #next() to navigate to
sub-elements and invoke getXMLObject() there. Do not keep references to
more than the currently processed XML to keep memory consumption low. The
method reads the stream up to the matching END_ELEMENT. When the method
returns the current event is the END_ELEMENT event.
Returns the count of attributes on this START_ELEMENT,
this method is only valid on a START_ELEMENT or ATTRIBUTE. This
count excludes namespace definitions. Attribute indices are
zero-based.
Returns the (local) name of the current event.
For START_ELEMENT or END_ELEMENT returns the (local) name of the current element.
For ENTITY_REFERENCE it returns entity name.
The current event must be START_ELEMENT or END_ELEMENT,
or ENTITY_REFERENCE.
Returns the count of namespaces declared on this START_ELEMENT or END_ELEMENT,
this method is only valid on a START_ELEMENT, END_ELEMENT or NAMESPACE. On
an END_ELEMENT the count is of the namespaces that are about to go
out of scope. This is the equivalent of the information reported
by SAX callback for an end element event.
If the current event is a START_ELEMENT or END_ELEMENT this method
returns the URI of the prefix or the default namespace.
Returns null if the event does not have a prefix.
Returns the current value of the parse event as a string,
this returns the string value of a CHARACTERS event,
returns the value of a COMMENT, the replacement value
for an ENTITY_REFERENCE, the string value of a CDATA section,
the string value for a SPACE event,
or the String value of the internal subset of the DTD.
If an ENTITY_REFERENCE has been resolved, any character data
will be reported as CHARACTERS events.
Returns the count of attributes on this START_ELEMENT,
this method is only valid on a START_ELEMENT or ATTRIBUTE. This
count excludes namespace definitions. Attribute indices are
zero-based.
Returns the normalized attribute value of the
attribute with the namespace and localName
If the namespaceURI is null the namespace
is not checked for equality
Parameters:
namespaceURI - the namespace of the attribute
localName - the local name of the attribute, cannot be null
Returns:
returns the value of the attribute or null if not found.
Returns the (local) name of the current event.
For START_ELEMENT or END_ELEMENT returns the (local) name of the current element.
For ENTITY_REFERENCE it returns entity name.
The current event must be START_ELEMENT or END_ELEMENT,
or ENTITY_REFERENCE.
Returns the count of namespaces declared on this START_ELEMENT or END_ELEMENT,
this method is only valid on a START_ELEMENT, END_ELEMENT or NAMESPACE. On
an END_ELEMENT the count is of the namespaces that are about to go
out of scope. This is the equivalent of the information reported
by SAX callback for an end element event.
Returns:
returns the number of namespace declarations on this specific element.
If the current event is a START_ELEMENT or END_ELEMENT this method
returns the URI of the prefix or the default namespace.
Returns null if the event does not have a prefix.
Returns:
the URI bound to this elements prefix, the default namespace, or null.
Returns the current value of the parse event as a string,
this returns the string value of a CHARACTERS event,
returns the value of a COMMENT, the replacement value
for an ENTITY_REFERENCE, the string value of a CDATA section,
the string value for a SPACE event,
or the String value of the internal subset of the DTD.
If an ENTITY_REFERENCE has been resolved, any character data
will be reported as CHARACTERS events.
Reads a sub-tree of the XML document and parses it as XML object.
The stream must be positioned on a START_ELEMENT. Do not call the method
when the stream is positioned at document's root element. This would
cause the whole document to be parsed into a single XML what may lead to
an out-of-memory condition. Instead use #next() to navigate to
sub-elements and invoke getXMLObject() there. Do not keep references to
more than the currently processed XML to keep memory consumption low. The
method reads the stream up to the matching END_ELEMENT. When the method
returns the current event is the END_ELEMENT event.
Returns true if there are more parsing events and false
if there are no more events. This method will return
false if the current state of the XMLStreamReader is
END_DOCUMENT
Get next parsing event - a processor may return all contiguous
character data in a single chunk, or it may split it into several chunks.
If the property javax.xml.stream.isCoalescing is set to true
element content must be coalesced and only one CHARACTERS event
must be returned for contiguous element content or
CDATA Sections.
By default entity references must be
expanded and reported transparently to the application.
An exception will be thrown if an entity reference cannot be expanded.
If element content is empty (i.e. content is "") then no CHARACTERS event will be reported.
The behavior of calling next() when being on foo will be:
1- the comment (COMMENT)
2- then the characters section (CHARACTERS)
3- then the CDATA section (another CHARACTERS)
4- then the next characters section (another CHARACTERS)
5- then the END_ELEMENT
NOTE: empty element (such as <tag/>) will be reported
with two separate events: START_ELEMENT, END_ELEMENT - This preserves
parsing equivalency of empty element to <tag></tag>.
This method will throw an IllegalStateException if it is called after hasNext() returns false.
Returns:
the integer code corresponding to the current parse event
Skips any white space (isWhiteSpace() returns true), COMMENT,
or PROCESSING_INSTRUCTION,
until a START_ELEMENT or END_ELEMENT is reached.
If other than white space characters, COMMENT, PROCESSING_INSTRUCTION, START_ELEMENT, END_ELEMENT
are encountered, an exception is thrown. This method should
be used when processing element-only content separated by white space.
Precondition: none
Postcondition: the current event is START_ELEMENT or END_ELEMENT
and cursor may have moved over any whitespace event.
Essentially it does the following (implementations are free to optimized
but must do equivalent processing):
Reads a sub-tree of the XML document and parses it as XML object.
The stream must be positioned on a START_ELEMENT. Do not call the method
when the stream is positioned at document's root element. This would
cause the whole document to be parsed into a single XML what may lead to
an out-of-memory condition. Instead use #next() to navigate to
sub-elements and invoke getXMLObject() there. Do not keep references to
more than the currently processed XML to keep memory consumption low. The
method reads the stream up to the matching END_ELEMENT. When the method
returns the current event is the END_ELEMENT event.
Test if the current event is of the given type and if the namespace and name match the current
namespace and name of the current event. If the namespaceURI is null it is not checked for equality,
if the localName is null it is not checked for equality.
Parameters:
type - the event type
namespaceURI - the uri of the event, may be null
localName - the localName of the event, may be null