rsample
tidymodels' resampling package is retiring its old splitters for sliding windows.
A side-by-side editorial comparison of Docling and mlr3 — release velocity, themes, recent moves, and the top alternatives to consider.
Docling keeps widening the funnel: every release adds another format the parser can swallow.
Docling ships roughly weekly, and the shape of each release is consistent — one or two new input formats or model paths, then a dense list of parser corrections for DOCX, ODF, PDF and PPTX. The latest window adds Outlook .msg with optional attachment listing, an EBCDIC backend, VLM grounding output from Unlimited-OCR, and OpenAI logprobs exposed as generated tokens. Underneath, the OCR layer was restructured into a layout-driven pipeline with configurable modes and a RapidOCR refactor that resolves all PP-OCR languages by version and backbone.
mlr3 is hardening the seams where its abstractions meet real learners
Releases arrive every few weeks and read as a systematic audit of the Learner interface. Recent versions added a native_model binding and a predict_raw flag so users can reach the underlying package's model and raw prediction, gave encapsulated learners a wall-clock deadline alongside the existing timeout, and removed the deprecated Task$divide(). A run of fixes addresses correctness at the boundary - factor level ordering that inverted binary probabilities, fallback learners losing state, misaligned probability columns.
Docling ships roughly weekly, and the shape of each release is consistent — one or two new input formats or model paths, then a dense list of parser corrections for DOCX, ODF, PDF and PPTX. The latest window adds Outlook .msg with optional attachment listing, an EBCDIC backend, VLM grounding output from Unlimited-OCR, and OpenAI logprobs exposed as generated tokens. Underneath, the OCR layer was restructured into a layout-driven pipeline with configurable modes and a RapidOCR refactor that resolves all PP-OCR languages by version and backbone.
Two things are being built at once. The conversion surface keeps broadening toward whatever a document actually arrives as — email, mainframe encodings, scanned pages, audio via Whisper — while the service datamodel grows the knobs a hosted pipeline needs: chunking options and targets, PDF heading-level inference, batch connector sources, configurable stage shutdown timeouts. The steady drip of DOCX and ODF reading-order fixes says fidelity, not throughput, is where the hard problems still are.
Expect the format list to keep extending and the VLM and OCR paths to gain more configurability, with reading-order and list-numbering corrections continuing at the same rate. The agent-skills addition suggests more packaging for agent callers, though the entries show only a first step.
Releases arrive every few weeks and read as a systematic audit of the Learner interface. Recent versions added a native_model binding and a predict_raw flag so users can reach the underlying package's model and raw prediction, gave encapsulated learners a wall-clock deadline alongside the existing timeout, and removed the deprecated Task$divide(). A run of fixes addresses correctness at the boundary - factor level ordering that inverted binary probabilities, fallback learners losing state, misaligned probability columns.
The framework is maturing from wrapping models to being accountable for what happens when wrapping goes wrong. Structured Mlr3Error and Mlr3Warning classes, conditions stored on the learner log, and messages replaced by conditions all point at making failures programmatically inspectable rather than printed. In parallel, escape hatches to the upstream model are being formalised instead of left to users digging into internals.
Expect the remaining deprecated surface to follow Task$divide() out, and further work on encapsulation and fallback behaviour, which is where most recent fixes have clustered. The raw and native_model accessors suggest more of the upstream model will be surfaced deliberately.
Other ai-assistants products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either Docling or mlr3.
tidymodels' resampling package is retiring its old splitters for sliding windows.
tidymodels' preprocessing engine learned sparsity, then settled into deprecations.
The resampling companion to scikit-learn now ships mostly to stay compatible with it.
parsnip added a whole new regression type, then wired R models to JAX and PyTorch
Mem0 splits agent memory from user memory, then spends a week hardening the plumbing
Every post is a comparison page, and Pictory is always the answer.
See all Docling alternatives → · See all mlr3 alternatives →
Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.
They serve adjacent needs but don't currently overlap on shipped themes. Docling is currently shipping more aggressively (velocity 6.3 vs 0.0), with 0 editorial sparks in the last 30 days against 0. See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.
Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. Docling is currently shipping more aggressively (velocity 6.3 vs 0.0), with 0 editorial sparks in the last 30 days against 0. For your specific use case, the alternatives sections above list other ai-assistants products to evaluate alongside.
Top Docling alternatives in ai-assistants are ranked by recent ship velocity. Browse the "Docling alternatives" section above for the current picks, or visit /alternatives/docling for the full list with editorial commentary on each.
Top mlr3 alternatives in ai-assistants are ranked by recent ship velocity. Browse the "mlr3 alternatives" section above for the current picks, or visit /alternatives/mlr3 for the full list with editorial commentary on each.