Skip to content

xml.dom docs are missing important information necessary for usage #156388

Description

@dgelessus

Documentation

I recently tried to use xml.dom/xml.dom.minidom for the first time, as someone who has moderate experience with other Python XML libraries (xml.etree.ElementTree, lxml, BeautifulSoup), but almost no prior experience with DOM APIs in any programming language. I found that the Python docs for xml.dom and xml.dom.minidom are missing important bits of information, so I had trouble making sense of these APIs, until I started looking at the source code, typeshed stubs, and external documentation.

Some (but not all) of the missing information can be found in the W3C DOM spec. It's not clear to me if readers of the xml.dom docs are expected to know the DOM spec already. If so, it would be good to clearly say so at the start of the docs, and to link the spec more prominently than just in the "See also" box.

The specific missing/unclear/confusing info I noticed:

  • The docs for getDOMImplementation mention "well-known" implementation names, but these "well-known" names don't seem to be documented. I can only find a list in the xml.dom.domreg source code.
  • There is no explanation of the possible children of each node type. For some types, there is a short mention that they have no children, but the rest is unexplained. I had to refer to the DOM spec and experiment interactively to properly understand which node types can appear where in the tree.
  • Node.childNodes is documented only as "A list of nodes" when it's specifically a NodeList according to the DOM spec.
  • The docs for Node.nodeName and Node.nodeValue say that they correspond to other, type-specific attributes, but don't show the exact mappings for each node type. They do explicitly refer to the DOM spec, which is good, but it would be better to provide this information inline (or at least link to the right subsection of the spec).
  • The docs say that Node.nodeName can be None, but if I'm reading the DOM spec right, nodeName is non-null for every possible node type. This might be an error in the docs.
  • It would be good to explicitly point out that NodeList isn't a subclass of Node, the same way this is pointed out for NamedNodeMap.
  • The NodeList.item docs make it sound like the list can contain None elements and out-of-range indices are disallowed. In reality, out-of-range indices are allowed and result in None (and apparently that this is the only case where None is returned?).
  • The xml.dom.minidom docs say "NodeList objects are implemented using Python’s built-in list type.". This isn't useful information for a user of minidom, and makes it sound like minidom is using plain list instead of NodeList, which isn't the case.
  • DocumentType.publicId and DocumentType.systemId can be None, but it's not explained when that happens, in which combination, and what it means. I assume the None-ness of these attributes indicates whether the doctype is a PUBLIC or a SYSTEM one, as there's no other API that indicates this.
  • DocumentType.name is documented as "The name of the root element as given in the DOCTYPE declaration, if present." (emphasis mine). Can a DocumentType ever have no name? The DOM spec doesn't say so, and in the XML syntax, the name is a required part of a <!DOCTYPE ...>.
  • DocumentType.entities and DocumentType.notations are documented, but the Entity and Notation node types they contain are undocumented, except for their node type constants.
  • It would be good to point out that Attr nodes are never children of anything and only appear in attributes maps.
  • Attr.localName and Attr.prefix are documented under Attr, even though they're seemingly identical to the base Node attributes of the same names.
  • NamedNodeMap's ...NamedItem[NS] methods (from the DOM spec) are not documented, even though minidom implements them.
  • The NamedNodeMap docs mention "experimental methods that give this class more mapping behavior", but don't document them. I could only find these methods in the minidom source code and typeshed stubs. They also differ somewat from the normal Python Mapping methods, so it's not enough to document NamedNodeMap as a Python mapping.
    • minidom has two classes implementing the NamedNodeMap interface - NamedNodeMap and ReadOnlySequentialNamedNodeMap - and only the former implements the mapping-like methods. It's not documented when minidom uses which implementation (this is again only visible in the source code and stubs), so it's impossible to use the mapping-like methods reliably.
  • Comment and Text have a common base interface CharacterData, which is completely undocumented. (Duplicates [xml.dom] The CharacterData interface and other members are not documented #155631.)
  • Text and CDATASection are documented together, but it's never mentioned that CDATASection inherits from Text.
  • The documentation pages for both xml.dom and xml.dom.minidom have sections at the end explaining how the DOM spec and IDL are mapped to Python. These seem to be mostly redundant with each other, so it's unclear which bits are part of the "Python DOM API" and which are implementation details of minidom.

Sorry for the long list! I hope this is useful feedback.

I might be able to start fixing some of these issues, but I'm not sure how useful that would be, because I'm not confident that I've understood it all correctly yet.

Linked PRs

Activity

  1. picnixz commented on Aug 26, 2026

    @picnixz
    Member
  2. suhanemathur commented on Aug 27, 2026

    @suhanemathur

    Hi @dgelessus ! Thanks for the detailed writeup.

    I'd like to take on a subset of these points, specifically the items I've verified directly against the Lib/xml/dom/minidom.py and Lib/xml/dom/minicompat.py source code:

    -3: Document Node.childNodes specifically as a NodeList rather than just "a list of nodes."
    -6: Add a note clarifying that NodeList is not a subclass of Node (parallel to the existing note for NamedNodeMap).
    -7: Correct NodeList.item() docs in xml.dom.rst to clarify that out-of-range indices return None rather than being disallowed/raising.
    -8: Clean up stale language in xml.dom.minidom.rst regarding NodeList implementation details to better clarify how childNodes behaves.
    -12: Note that Attr nodes only appear in attribute maps and never as child nodes (enforced by HierarchyRequestErr via _child_node_types).

    I think the remaining items (#5 regarding nodeName/None, #14/#15 on NamedNodeMap methods, and the broader structural changes) would benefit from further maintainer discussion before working on them. I'm happy to leave those open for @serhiy-storchaka and the team to weigh in on while I can work on the points mentioned above.

  3. picnixz commented on Aug 27, 2026

    @picnixz
    Member

    @suhanemathur Please don't post bare LLM replies that look like "plans". It's not useful to the discussion. First, we need to determine what we want to expose, not just "change the docs". So for now, there is nothing to work on.

  4. serhiy-storchaka commented on Aug 27, 2026

    @serhiy-storchaka
    Member

    See #155632 which can address some of these points.

  5. dgelessus commented on Aug 27, 2026

    @dgelessus
    ContributorAuthor

    Wow, that's some timing :D Thank you for the link, that indeed adds multiple things I was looking for!

    I've edited the issue description and marked the items that will be resolved by your PR. The remaining items are still valid, I believe.

  6. added a commit that references this issue on Aug 28, 2026
  7. serhiy-storchaka commented on Aug 28, 2026

    @serhiy-storchaka
    Member

    I added few more fixes in #155632.

    #155641 fixes few other documentation points, together with implementation.

  8. dgelessus commented on Aug 31, 2026

    @dgelessus
    ContributorAuthor

    Thank you very much, that really helps! Updated the list again.

    The few remaining points seem relatively clear to me, so I'll try to submit a PR for those myself (unless someone else wants to do it).

    Just to be sure, have I understood correctly how the different combinations of DocumentType.publicId/.systemId map to different kinds of DOCTYPEs?

    publicId is not None publicId is None
    systemId is not None <!DOCTYPE root PUBLIC "publicId" "systemId"> <!DOCTYPE root SYSTEM "systemId">
    systemId is None impossible <!DOCTYPE root>
  9. serhiy-storchaka commented on Sep 1, 2026

    @serhiy-storchaka
    Member

    Yes, your table is correct for parsed documents, a parser can never produce a public identifier without a system identifier. createDocumentType() does not validate the identifiers, so other combinations can be created programmatically, but they do not make sense in XML.

  10. Neiland85 commented on Sep 16, 2026

    @Neiland85

    Thanks, Serhiy. I’ve incorporated this clarification into #157646, including the distinction between parsed documents and programmatically created DocumentType objects.

    The documentation now explicitly describes the PUBLIC, SYSTEM, and no-external-identifier cases, and notes that createDocumentType() does not validate identifier combinations.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

docsDocumentation in the Doc dirtopic-XML

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions