xmlrectr

From unfamiliar XML to analysis-ready R tables.

xmlrectr is an ambitious attempt at a universal XML rectangler for R.

It is not trying to be a universal XML parser. Excellent XML parsers already exist. The problem addressed here is different: how do you turn hierarchical XML into a table or tibble that is genuinely useful for analysis with the tidyverse, base R, Arrow, and other tools built around tabular data?

That problem becomes especially difficult when you are simply handed an XML file and know little or nothing about it. You may have no schema, no documentation, no knowledge of the XML vocabulary, and no predefined extraction rules. xmlrectr can inspect the actual document structure and, in many cases, construct a workable analyst-friendly tibble automatically.

When you do know the XML structure, have documentation or an XSD, or know exactly what you want to expose, the package becomes more explicit rather than less useful. You can review structural proposals, provide row and identifier choices, select and rename fields, use schema evidence, save reusable profiles, and obtain a rectangle closely aligned with your analytical purpose.

The package is deliberately generic. It contains no special rules for MARC, FHIR, EAD, ONIX, GPX, UBL, or any other XML vocabulary.


Installation

At present, xmlrectr can be installed from GitHub.

The recommended method is pak:

install.packages("pak")
pak::pak("larry77/xmlrectr")

Alternatively:

install.packages("remotes")
remotes::install_github("larry77/xmlrectr")

Because xmlrectr contains native C code, the GitHub version is compiled from source and requires a working build toolchain and libxml2 development files.

Windows build requirements

Windows is a first-class supported platform. Use the Rtools version matching your R installation and install it in the default location.

For the currently supported R series:

With a normal R/Rtools installation, manual PATH configuration should not usually be necessary.

If compilation fails, first check the toolchain:

install.packages("pkgbuild")
pkgbuild::has_build_tools(debug = TRUE)

Avoid mixing Rtools with unrelated MSYS2/MinGW/Strawberry Perl toolchains on PATH, as this can produce difficult-to-diagnose linking problems.

The GitHub Actions workflow compiles and tests the package on windows-latest.

Linux build requirements

On Debian/Ubuntu:

sudo apt install libxml2-dev pkg-config

Then install xmlrectr from R using pak or remotes.

macOS build requirements

With Homebrew:

brew install libxml2 pkg-config
export PKG_CONFIG_PATH="$(brew --prefix libxml2)/lib/pkgconfig:$PKG_CONFIG_PATH"

Then install xmlrectr from R.

For a local checkout:

pak::local_install(".")

Why XML rectangling is hard

XML is hierarchical. Analytical data is usually rectangular.

A naive “flatten everything” strategy can easily:

There is also an unavoidable conceptual limit: an arbitrary XML document does not have one mathematically unique tabular interpretation. Domain knowledge can always improve a rectangle.

xmlrectr therefore does not claim to infer the intended semantics of every XML document. Instead, it aims to make a useful, conservative structural interpretation when knowledge is scarce, while exposing progressively more control when the user knows more.

The core design is:

unknown XML
    |
    +--> automatic analyst-oriented rectangle
    |
    +--> structural proposal
             |
             v
        human review
             |
             v
        reusable profile
             |
             v
        compiled rectangle

XSD information can contribute evidence, but it is advisory rather than blindly authoritative.


Two kinds of automation

A major goal of xmlrectr is to automate both the analytical problem and the computational problem, while keeping both layers tunable.

1. Analytical automation: how should this XML become a table?

For exploration, rectangle_xml_analyst() can start from the XML file itself:

library(xmlrectr)

file <- system.file("extdata", "orders.xml", package = "xmlrectr")

analyst <- rectangle_xml_analyst(file)
analyst

The result is one self-contained tibble with:

This is the deliberately ambitious part of the package: starting from unfamiliar XML and attempting to produce something an R analyst can immediately inspect and work with.

For one-off exploration, this may be all you need.

For repeated or production use, once you understand the structure you will usually want to move to an explicit profile so that the intended rectangle becomes a stable, reviewable contract.

2. Computational automation: how should the work be executed?

Once a rectangle is defined, execution has its own set of choices: sequential or parallel processing, worker count, chunk size, task size, memory ownership, and whether a large document should be streamed instead of fully materialised.

Those choices are intentionally separated from the analytical meaning of the rectangle.

For the normal rectangling APIs, one argument can ask xmlrectr to choose whether parallel execution is worthwhile:

out <- rectangle_xml(
  file,
  profile,
  parallel = "auto"
)

Automatic execution planning is based on the observed structural workload and available resources, not on XML vocabulary names, filenames, or rules tuned to the validation corpus. The package does not apply one fixed worker/chunk recipe to every document.

Advanced controls remain available, but they are optional. The point of the architecture is that you should not need to turn every screw before getting useful work done on a large, nested or unfamiliar XML file.


Automatic when you need it, explicit when you want it

xmlrectr is designed to support a continuum of prior knowledge.

You know almost nothing about the XML

Start with the automatic analyst table:

analyst <- rectangle_xml_analyst(file)

This is the quickest route from an unfamiliar document to an R tibble.

You want to understand the structure before deciding

Ask the package for a proposal:

proposal <- propose_xml_profile(file)

review_xml_proposal(proposal, "rows")
review_xml_proposal(proposal, "ids")
review_xml_proposal(proposal, "fields")

The proposal is evidence, not an executable command. It lets the package inspect the XML and suggest plausible structural choices without pretending that software can know your analytical intent.

You know what you want to expose

Define it explicitly:

profile <- xml_profile(
  rows = "order",
  id = "id"
)

out <- rectangle_xml(file, profile)

A profile can also select and rename fields, specify types, provide namespace information, and choose the desired layout.

You can save the reviewed decision:

write_xml_profile(profile, "orders-profile.json")
profile2 <- read_xml_profile("orders-profile.json")

and reuse it across files belonging to the same XML family.

For repeated processing, compile the profile once:

spec <- compile_xml_profile(profile, file)

out <- rectangle_xml(
  file,
  spec,
  parallel = "auto"
)

This is where domain knowledge pays off: when you know the XML structure and the analytical question, the package can produce a rectangle much more closely aligned with your desiderata than any fully automatic method could infer.


XSD-assisted work

If an XSD is available, xmlrectr can use it as additional structural evidence:

xml <- system.file("extdata", "types.xml", package = "xmlrectr")
xsd <- system.file("extdata", "types.xsd", package = "xmlrectr")

proposal <- propose_xml_profile(xml, xsd = xsd)

review_xml_proposal(proposal, "xsd")

inspect_xsd() can expose useful declaration, occurrence, required-attribute and scalar-type information.

The important design choice is that XSD is advisory. A schema describes valid document structure, but it does not necessarily tell an analyst what should constitute a row, which repeated structures should become separate entities, or which fields are relevant to a particular analysis.

So the workflow remains:

XML structure + optional XSD + user knowledge
                    |
                    v
              reviewed profile
                    |
                    v
             analytical rectangle

What the analyst representation tries to preserve

The automatic analyst representation is designed as one self-contained atomic table per XML document.

Its structural contract is conservative:

The underlying canonical representation is even more explicit. It records document order, node identity, parentage, depth, namespaces, attributes, text/CDATA and other retained XML node types.

That canonical layer is the loss-aware structural foundation on which higher-level rectangles are built.


Built for real XML, not only toy examples

The ambition to work with arbitrary XML is useful only if the implementation can cope with XML as it exists in practice: large documents, deep nesting, repeated records, namespaces, irregular branches and substantial structural overhead.

A large part of the development of xmlrectr has therefore focused on algorithmic cost, memory behaviour, workload decomposition and semantic parity.

The package deliberately separates two questions:

  1. What should the rectangle mean?
  2. How should that rectangle be computed efficiently on this particular XML document?

Changing the computational engine must never change the first answer.

Native structural acceleration

Performance-critical canonical reading and structural/indexing operations have native C implementations using libxml2.

The R implementation remains the semantic reference. Compiled code is used to accelerate structural bottlenecks, not to introduce a second set of rectangling semantics.

This distinction matters: generic XML rectangling repeatedly performs structural operations for which interpreted R can become expensive on large trees. Moving those bottlenecks into compiled code makes the generic design practical without introducing vocabulary-specific shortcuts.

Bounded-memory streaming

For XML that should not be represented as one complete in-memory canonical table, xmlrectr provides a streaming path based on complete record subtrees.

batches <- list()

stats <- xml_stream_rectangle(
  "large.xml",
  spec,
  callback = function(batch) {
    batches[[length(batches) + 1L]] <<- batch
  },
  parallel = "auto"
)

The streaming architecture has an important correctness boundary:

The point is not merely to “use less RAM”. The package tries to keep parser state, record ownership, output ordering and rectangling semantics cleanly separated.

CSV and Parquet output

Large results can be written without first collecting the entire rectangle into one R object:

rectangle_xml_csv(
  "large.xml",
  spec,
  output = "large.csv",
  parallel = "auto"
)

For typed analytical output:

rectangle_xml_parquet(
  "large.xml",
  spec,
  output_dir = "large-parquet",
  parallel = "auto"
)

Parquet support requires the optional arrow package.

The same high-level execution controls are used by the main in-memory, streaming, CSV and Parquet interfaces. Parallel execution is therefore part of the normal rectangling architecture rather than a separate workflow bolted onto one output format.


Parallel execution without a second API

Parallelism is an execution choice of the same rectangling operation, not a separate family of user-facing functions.

# Exact sequential path
seq_out <- rectangle_xml(
  file,
  spec,
  parallel = FALSE
)

# Request parallel execution with tuned defaults
par_out <- rectangle_xml(
  file,
  spec,
  parallel = TRUE
)

# Let the engine decide whether parallel work is worthwhile
auto_out <- rectangle_xml(
  file,
  spec,
  parallel = "auto"
)

The sequential implementation is the semantic oracle. Parallel execution is required to preserve the same result:

identical(seq_out, par_out)

The three modes have deliberately simple meanings:

Lower-level *_parallel() functions exist for compatibility, testing and diagnostics. They are not the intended everyday API.

Parallelise records, not parser state

The parallel architecture follows a few strict rules:

  1. Only independent complete XML record subtrees are parallelised.
  2. XML is never divided by arbitrary byte ranges.
  3. Live xml2/libxml external pointers are never sent to workers.
  4. SAX parsing and record-boundary detection remain coordinator-side.
  5. Workers receive self-contained record material.
  6. Source order, identifiers and sequential semantics must remain exact.
  7. Coordinator-side responsibilities such as duplicate-ID checks and callbacks remain coordinator-side.
  8. The caller’s pre-existing Future plan is restored after execution.

These constraints are less flashy than a benchmark chart, but they are fundamental to making parallel execution a trustworthy implementation detail rather than a second semantics.


Why the scheduler is adaptive

Parallel XML rectangling is not simply a matter of running:

workers <- parallel::detectCores()

and dividing a file into equal pieces.

A useful execution plan depends on the actual structural work available:

For this reason, xmlrectr does not use filenames, XML vocabulary names, or rules learned specifically from the validation corpus to decide whether to parallelise.

The automatic policy is structural.

Automatic worker count: deliberately conservative

Benchmarks showed that increasing the number of workers beyond four can still reduce elapsed time, but efficiency falls and memory pressure increases.

The automatic/default worker policy therefore normally uses up to four workers, while preserving a core for the system where possible.

This is not a hard maximum.

Advanced users can explicitly request more workers:

rectangle_xml(
  file,
  spec,
  parallel = TRUE,
  workers = 8
)

The default is intended to be a balanced choice, not a claim that four workers are universally optimal.

How "auto" currently decides whether there is enough work

The current in-memory auto policy asks whether there is enough canonical record work per worker to justify process-level parallelism.

The policy includes a work floor of approximately:

25,000 canonical record nodes per worker

together with sufficient record/subtree structure, approximately:

at least 128 records per worker

or

a median record subtree of at least 1,000 canonical nodes

These are engineering defaults derived from broad structural benchmarking. They are not XML-vocabulary rules, and they should not be read as eternal constants or promises of a particular speedup.

There is also a cheap impossibility check. If the entire canonical table has fewer than roughly:

workers * 25,000

nodes, the workload cannot satisfy the per-worker work floor. Auto mode can then remain sequential immediately instead of performing a more expensive record-span analysis.

This fast path is important. An early version of automatic planning was semantically correct but could make small sequential jobs noticeably slower simply because planning repeated structural work that the sequential path would perform anyway. The current design reuses validation and rejects obviously too-small workloads before doing that extra work.

In other words, auto mode is designed not only to find parallel opportunities, but also to get out of the way when parallelism would be pointless.


Two parallel strategies, for two different constraints

xmlrectr retains two parallel execution strategies because throughput and memory pressure are different optimisation problems.

parallel_chunks: throughput-oriented

With parallel_chunks, workers own independent vectorised outer chunks.

Conceptually:

chunk 1 ---> worker 1
chunk 2 ---> worker 2
chunk 3 ---> worker 3
chunk 4 ---> worker 4

This strategy is designed primarily to:

For in-memory automatic execution, this is generally the preferred strategy.

A useful scheduling model is approximately one active owned outer chunk per worker.

shared_chunk: memory-oriented

With shared_chunk, one bounded outer chunk is shared and subdivided into multiple vectorised tasks coordinated through mori.

Conceptually:

             bounded shared outer chunk
                 /    |    |    \
              task  task  task  task
                |     |     |     |
              workers process coarse ranges

Its purpose is to reduce input-memory duplication while still preserving parallel work.

This is particularly attractive for streaming, CSV and Parquet workflows, where bounded memory is part of the reason for choosing the execution mode in the first place.

Automatic streaming/output execution therefore prefers shared_chunk when the required stack is available. If mori is unavailable, automatic execution can fall back to parallel_chunks.

The trade-off is intentional:

Neither strategy dominates the other on every machine and workload.


Chunk size and task size are different controls

The advanced API exposes:

workers
strategy
chunk_records
task_records

The last two parameters solve different problems.

chunk_records: working-set admission

chunk_records controls the size of the outer batch admitted at once.

It therefore influences:

For the current streaming defaults:

parallel_chunks:
    chunk_records = 512

For shared_chunk:

chunk_records = min(1024, max(256, workers * 256))

These are tuned defaults, not XML-specific rules.

task_records: scheduling granularity

Within a shared outer chunk, task_records controls how finely the work is subdivided for scheduling.

Benchmarks found a fairly broad useful plateau around 4 to 8 tasks per worker. The balanced automatic setting is therefore approximately 4 tasks per worker.

That gives workers enough independent work for load balancing without producing a large number of tiny tasks whose scheduling cost dominates useful computation.

So:

chunk_records -> controls admitted working-set size / memory
task_records  -> controls scheduling granularity inside that work

Keeping these concepts separate is especially important for shared_chunk.


What the benchmark evidence actually showed

Performance results are included here because the defaults were not chosen by intuition alone.

They should nevertheless be interpreted carefully:

Never compare raw elapsed times from different machines as though they belong to one benchmark series.

Absolute timings depend on processor, memory subsystem, operating system, R build, package versions and background load. The useful quantities are same-machine sequential/parallel speedup, worker efficiency, memory behaviour and exact semantic parity.

The following results are representative development measurements on one machine, referred to during development as einstein. They document why the current defaults exist; they are not runtime guarantees.

Worker-count experiment

A representative synthetic workload was approximately 16.266 MiB.

Sequential execution:

193.977 s

Parallel results:

Workers parallel_chunks Speedup shared_chunk Speedup
2 121.585 s 1.60x 107.070 s 1.81x
4 63.340 s 3.06x 68.869 s 2.82x
11 46.051 s 4.21x 55.525 s 3.49x

Several conclusions follow.

First, more than four workers can improve elapsed time. Four is therefore not a hard ceiling.

Second, scaling efficiency declines substantially at high worker counts. The extra processes are doing useful work, but the cost of coordination, memory traffic and finite task parallelism becomes increasingly important.

Third, four workers gave a strong compromise between speedup, efficiency and memory pressure. That is why automatic/default worker selection is normally capped there unless the user explicitly chooses otherwise.

Memory experiment

On the same approximate 16.266 MiB workload with four workers:

Strategy Peak PSS
parallel_chunks about 3618 MiB
shared_chunk about 3142 MiB

In that experiment, shared_chunk reduced peak proportional set size by roughly 13% and private memory by roughly 15%.

Depending on phase, median PSS could fall by considerably more.

That reduction is meaningful, even though shared_chunk can be slower on some workloads. It is the empirical reason the memory-oriented strategy remains part of the package rather than being removed in favour of the single fastest throughput result.

What this does not mean

These measurements do not imply:

The actual structural workload matters more than the byte size of the XML file.


Auto mode in real-world smoke tests

The final automatic policy was also checked against real XML on the same development machine.

Representative P6.1 smoke timings were:

XML Sequential Auto Auto decision
UBL 0.960 s 0.874 s sequential
Maven 1.111 s 1.126 s sequential
EAD 33.756 s 30.973 s parallel

The important observation for UBL and Maven is not the tiny timing difference. It is that automatic planning did not impose a material penalty on small jobs that should remain sequential.

The EAD document crossed the structural threshold and was sent to the parallel engine.

That particular EAD run should not be used to argue either that parallelism is spectacular or that it is useless. It happened to lie relatively close to the crossover region on that machine. Larger synthetic workloads demonstrated much stronger same-machine speedups.


Tuning for technically minded users

Most users should stop at:

rectangle_xml(file, spec, parallel = "auto")

The following controls exist for benchmarking, unusually constrained machines, or specialist tuning:

rectangle_xml(
  file,
  spec,
  parallel = TRUE,
  workers = 4,
  strategy = "shared_chunk",
  chunk_records = 1024,
  task_records = 64
)

Useful questions for an advanced tuning exercise are:

Do not tune from filenames or XML vocabulary names.

Do not assume that tiny chunk/task values used in semantic stress tests are production recommendations.

And do not optimise one specific XML file at the expense of the generic structural rules.


A reproducible way to benchmark your own XML

For performance work, compare strategies on the same machine and the same XML.

A simple reproducible pattern is:

spec <- compile_xml_profile(profile, file)

t_seq <- system.time(
  seq_out <- rectangle_xml(
    file,
    spec,
    parallel = FALSE
  )
)

t_auto <- system.time(
  auto_out <- rectangle_xml(
    file,
    spec,
    parallel = "auto"
  )
)

t_forced <- system.time(
  forced_out <- rectangle_xml(
    file,
    spec,
    parallel = TRUE
  )
)

stopifnot(
  identical(seq_out, auto_out),
  identical(seq_out, forced_out)
)

rbind(
  sequential = t_seq,
  auto = t_auto,
  forced_parallel = t_forced
)

For serious benchmarking:

  1. run multiple repetitions;
  2. compare within-machine speedups rather than raw times from different platforms;
  3. inspect memory as well as elapsed time when that matters;
  4. keep semantic checks in the benchmark harness;
  5. distinguish performance benchmarks from tests whose only purpose is to stress correctness.

Small XML documents may correctly be faster sequentially because process startup and scheduling have a cost. Larger, sufficiently coarse record workloads are where parallel execution can pay off.

This is why parallel = "auto" exists: parallelism is a tool, not a goal in itself.


Semantic testing is not performance benchmarking

Some of the harshest parallel tests deliberately used settings that would make poor production defaults.

For example, the 30-file forced-parallel semantic stress run used approximately:

workers       = 2
chunk_records = 64
task_records  = 8

Those small values were chosen to force many scheduling boundaries and expose correctness problems.

They were not selected for speed.

The result was:

This distinction is important when reading the project’s benchmark history. Different experiments answer different questions:

Conclusions from one class should not be casually transferred to another.


Battle-tested, not vocabulary-tuned

Before conversion into the R package, the frozen validated engine was exercised against a deliberately diverse corpus of 30 real-world XML documents.

The validation included:

In the final 30-file automatic-policy run:

Crucially, the engine was not modified with vocabulary-specific rules, filename-specific exceptions or thresholds tuned to make these files pass.

Current 30-file real-world validation corpus

The corpus covers very different XML domains and structures:

  1. UBL invoice
  2. FHIR patient
  3. GPX route
  4. KML places
  5. METS metadata
  6. MODS records
  7. JUnit report
  8. Nmap scan
  9. DocBook book
  10. PubMed articles
  11. BLAST result
  12. BioSample record
  13. RDF vocabulary
  14. GraphML graph
  15. RSS feed
  16. Atom feed
  17. XLIFF 2.0
  18. XLIFF 1.2
  19. MusicXML score
  20. OpenStreetMap
  21. SVG drawing
  22. SDMX Generic
  23. SDMX Structure
  24. OOXML shared strings
  25. Maven project
  26. Android manifest
  27. EAD finding aid
  28. SoapUI project
  29. TEI person data
  30. ONIX books

The purpose of this corpus is structural diversity, not optimisation for these particular vocabularies. General structural rules are preferred over corpus-specific special cases.

This validation does not mean that every arbitrary XML document has one objectively correct analyst table. It means that the package’s generic structural rules and execution engine have been exercised across a broad set of real-world XML shapes without resorting to vocabulary-specific parsers.


What the performance work tells us

The engineering conclusions behind the current defaults can be summarised compactly:

  1. Sequential execution is the semantic oracle.
  2. Automatic workers normally stop at four, but explicit higher counts are allowed.
  3. More workers can improve raw elapsed time, with declining efficiency.
  4. parallel_chunks is the primary throughput-oriented strategy.
  5. shared_chunk is the primary memory-oriented strategy.
  6. shared_chunk has shown materially lower peak/private memory in representative tests.
  7. Outer chunk_records and inner task_records solve different problems and are tuned separately.
  8. Shared execution benefits from a modest number of coarse tasks; excessive tiny tasks waste scheduling time.
  9. Auto mode estimates structural work, not filenames, XML types or file size alone.
  10. Fast auto rejection is important so small documents pay essentially no planning penalty.
  11. No performance optimisation is accepted merely because it helps the existing validation corpus.
  12. Any tuning change must preserve exact sequential parity before its speed or memory result matters.

The public API is simple because this machinery sits underneath it, not because the machinery does not exist.


Malformed XML

xmlrectr expects well-formed XML.

Malformed XML is detected and reported as an error or warning where appropriate. Repairing broken XML is intentionally outside the scope of the package: xmlrectr rectangles XML; it does not try to guess how a malformed source document should be rewritten.

Likewise, XSD inspection is intended to provide useful schema evidence for rectangling. xmlrectr is not a complete XSD validation or repair framework.


Documentation

This README is intended to be a self-contained introduction. You should not need to install the package or open a vignette merely to understand what xmlrectr is trying to do.

For readers who want more detail, the repository also contains technical material that can be read directly on GitHub:

The technical articles are intentionally more detailed than this README. They are the right place for readers interested in scheduler design, bounded-memory execution, native acceleration, benchmark interpretation, worker/chunk/task tuning, validation boundaries and the engineering decisions behind the simple public interface.


Design principles

A few principles define the project:

The intended experience is simple even though the implementation underneath is not:

Give xmlrectr an XML file. If you know nothing about it, start exploring immediately. If you know more, tell the package what you know. If the workload is large, let the execution engine do the heavy lifting.


Scope

xmlrectr aims to be a universal rectangler, not a universal semantic interpreter.

It is designed to take arbitrary well-formed XML and produce useful R-oriented tabular representations without requiring a vocabulary-specific parser. It can exploit schema information and user knowledge when they exist, but it does not require them for exploratory use.

That is a deliberately ambitious target. The package cannot know the domain meaning of every XML vocabulary, but it can do a great deal of the structural and computational work required to move from hierarchical XML to an analyst-friendly table.

That is the problem xmlrectr is built to solve.

mirror server hosted at Truenetwork, Russian Federation.