The component model¶
Three languages called XSD¶
When people say "XSD" they mean one of three things, and the difference is the reason this library exists.
- Schema documents. The XML you write:
<xs:element>,<xs:complexType>,xs:include,xs:import. This is a syntax, and it is deliberately flexible — the same schema can be spelled a dozen ways, split across forty files, or fold a definition inline instead of naming it. - Schema components. The abstract graph those documents assemble into: type definitions, element declarations, attribute uses, particles, model groups, wildcards. Named and inline definitions become the same kind of thing. Forty files become one graph.
- Validation semantics. The rules that decide whether a document is valid.
The specification defines every rule in the third layer against the second one, never against the first. "Element Locally Valid (Complex Type)" is a sentence about components, not about angle brackets.
So a tool that skips straight from documents to a verdict has to reconstruct
that middle layer internally — and then it throws it away. xsdkit builds it
and gives it to you.
The practical consequence
Questions like "what may appear inside a report?" are hard to answer
from documents (the children could come from a base type three
xs:extension steps up, or a xs:group reference, or a substitution
group) and easy to answer from components. Every awkward part has already
been resolved by the time you hold a Schemas.
The lifecycle¶
There are exactly two states, and they are two types.
flowchart LR
A["Schema documents<br/><small>.xsd files, strings, bytes</small>"]
B["SchemaSetBuilder<br/><small>reads, resolves, compiles</small>"]
C["Schemas / SchemaSet<br/><small>the component graph</small>"]
A --> B --> C
C -.->|query| D["children, types, facets,<br/>occurrence, automata"]
C -.->|validate| E["diagnostics + typed PSVI"]
A Schemas never exists in a half-resolved state. There is no Compile() you
can forget to call, and no accessor that returns None because the graph is
not ready yet — a design mistake that .NET's XmlSchemaSet and several others
make, where you can query an unresolved schema and get quiet nonsense.
Building is the expensive step; querying is cheap. Build once, keep it, ask it
many questions. It is Send + Sync, so one compiled schema can serve every
thread, and the Python bindings release the GIL around the build.
What a component graph contains¶
import xsdkit
schemas = xsdkit.SchemaSet.from_file("report.xsd")
schemas.counts
# {'types': 56, 'elements': 6, 'attributes': 12, 'particles': 7,
# 'model_groups': 0, 'attribute_groups': 0,
# 'identity_constraints': 0, 'notations': 0, 'annotations': 1}
len(schemas) # 6 — the globals *this* schema declares
Fifty-six types from a sixty-line schema, because the 50 built-ins are real
components too. xs:string is not a special case in a match arm somewhere —
it is a SimpleType with a variety, a primitive, facets and a base chain, and
it resolves exactly the way one of your own types does.
schemas.type("http://www.w3.org/2001/XMLSchema", "string")
# <Type simple {http://www.w3.org/2001/XMLSchema}string>
[t.qname for t in schemas.types["{urn:example}Money"].base_chain]
# ['{urn:example}Money',
# '{http://www.w3.org/2001/XMLSchema}decimal',
# '{http://www.w3.org/2001/XMLSchema}anyAtomicType',
# '{http://www.w3.org/2001/XMLSchema}anySimpleType',
# '{http://www.w3.org/2001/XMLSchema}anyType']
The built-ins are left out of schemas.types, though — they are in every
schema set and would bury what your documents actually declared.
schemas.type(...) still finds them.
Seven symbol spaces¶
XSD keeps seven separate namespaces for names. A type called Item and an
element called Item are unrelated components that never collide, and neither
does a model group of the same name.
| Symbol space | Declared by |
|---|---|
| Type | xs:simpleType, xs:complexType |
| Element | xs:element |
| Attribute | xs:attribute |
| Model group | xs:group |
| Attribute group | xs:attributeGroup |
| Notation | xs:notation |
| Identity constraint | xs:key, xs:keyref, xs:unique |
xsdkit keeps them apart. That is why lookups are per-kind — schemas.element(...),
schemas.type(...), schemas.attribute(...) — rather than one lookup that
would have to guess which Item you meant.
Subscripting follows the same rule. schemas["{urn:example}Item"] is always
an element, schemas.types["{urn:example}Item"] a type and
schemas.attributes[...] an attribute, so a name an element and a type share
reaches both, and a name looked up in the wrong place says where it is.
Global and local¶
Only global declarations — the direct children of xs:schema — are
addressable by name. A local element declared inside a complex type is a real
component with a real type, but it is scoped to that type, and two types may
each declare a price that means something different.
report = schemas["{urn:example}report"]
report.is_global # True
price = report["item"]["price"]
price.is_global # False — it exists only inside Item
price.qname # '{urn:example}price'
You reach locals by navigating to them, which is what element["child"] and
type.children are for.
Names¶
A name is a pair — namespace and local part — not a string. Prefixes (tns:,
xs:) belong to the document; they are resolved away at load time and never
appear in the model. Two documents that use different prefixes for the same
namespace produce identical components.
For display and lookup, the pair is written in Clark notation:
{urn:example}report. Anywhere xsdkit takes a name in Python it accepts
Clark notation, a bare local name, or an explicit (namespace, local) pair.
use xsdkit::Schemas;
fn names(schemas: &Schemas) -> Option<()> {
let report = schemas.element(Some("urn:example"), "report")?;
report.namespace(); // Some("urn:example")
report.local_name(); // "report"
report.display_name(); // "{urn:example}report"
// The same, for a `QName` held on its own.
let q = report.name();
schemas.namespace_of(q);
schemas.local_of(q);
Some(())
}
Internally names are interned to u32 ids, so comparing them is an integer
comparison rather than a string one — which matters when a content automaton
is comparing thousands of them per document. That is an implementation
decision, not one you are party to: nothing in either API asks you to hold an
interner to turn a name back into text.
Annotations survive¶
xs:documentation and xs:appinfo are kept, not discarded.
appinfo is kept verbatim, as XML text, because it is where schema
families put the machine-readable conventions the standard never specified —
database mappings, code lists, UI hints. Summarising it would destroy exactly the
information someone reaching for it needs. Each element in it declares the
namespaces in scope, so an XML parser takes it as it is.
Next¶
- Loading schemas — getting documents into a
SchemaSet. - Querying the model — what to ask it once you have one.