recipes
tidymodels' preprocessing engine learned sparsity, then settled into deprecations.
A side-by-side editorial comparison of Docling and rsample — release velocity, themes, recent moves, and the top alternatives to consider.
Docling keeps widening the funnel: every release adds another format the parser can swallow.
Docling ships roughly weekly, and the shape of each release is consistent — one or two new input formats or model paths, then a dense list of parser corrections for DOCX, ODF, PDF and PPTX. The latest window adds Outlook .msg with optional attachment listing, an EBCDIC backend, VLM grounding output from Unlimited-OCR, and OpenAI logprobs exposed as generated tokens. Underneath, the OCR layer was restructured into a layout-driven pipeline with configurable modes and a RapidOCR refactor that resolves all PP-OCR languages by version and backbone.
tidymodels' resampling package is retiring its old splitters for sliding windows.
rsample is at 1.3.2, a small release covering spatialsample interoperability and a soft deprecation of the lag argument on initial_time_split(). The more consequential work sits behind it: 1.3.1 added internal_calibration_split() and a calibration() accessor so tune can fit a preprocessor and a post-processor on separate parts of the analysis set, and 1.3.0 superseded rolling_origin() with the sliding_* family.
Docling ships roughly weekly, and the shape of each release is consistent — one or two new input formats or model paths, then a dense list of parser corrections for DOCX, ODF, PDF and PPTX. The latest window adds Outlook .msg with optional attachment listing, an EBCDIC backend, VLM grounding output from Unlimited-OCR, and OpenAI logprobs exposed as generated tokens. Underneath, the OCR layer was restructured into a layout-driven pipeline with configurable modes and a RapidOCR refactor that resolves all PP-OCR languages by version and backbone.
Two things are being built at once. The conversion surface keeps broadening toward whatever a document actually arrives as — email, mainframe encodings, scanned pages, audio via Whisper — while the service datamodel grows the knobs a hosted pipeline needs: chunking options and targets, PDF heading-level inference, batch connector sources, configurable stage shutdown timeouts. The steady drip of DOCX and ODF reading-order fixes says fidelity, not throughput, is where the hard problems still are.
Expect the format list to keep extending and the VLM and OCR paths to gain more configurability, with reading-order and list-numbering corrections continuing at the same rate. The agent-skills addition suggests more packaging for agent callers, though the entries show only a first step.
rsample is at 1.3.2, a small release covering spatialsample interoperability and a soft deprecation of the lag argument on initial_time_split(). The more consequential work sits behind it: 1.3.1 added internal_calibration_split() and a calibration() accessor so tune can fit a preprocessor and a post-processor on separate parts of the analysis set, and 1.3.0 superseded rolling_origin() with the sliding_* family.
Two threads run through the window. Time-based resampling is migrating from rolling_origin() to sliding_window(), sliding_index() and sliding_period(), while validation_split() and its relatives have moved from soft deprecation to warning in favour of the three-way initial_validation_split(). Alongside that, rsample is growing infrastructure other tidymodels packages consume rather than user-facing splitters.
Given that validation_split() and friends now warn and initial_time_split()'s lag argument is soft-deprecated, the next release most likely escalates those deprecations rather than adding a resampling scheme.
Other ai-assistants products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either Docling or rsample.
tidymodels' preprocessing engine learned sparsity, then settled into deprecations.
The resampling companion to scikit-learn now ships mostly to stay compatible with it.
parsnip added a whole new regression type, then wired R models to JAX and PyTorch
mlr3 is hardening the seams where its abstractions meet real learners
Mem0 splits agent memory from user memory, then spends a week hardening the plumbing
Every post is a comparison page, and Pictory is always the answer.
See all Docling alternatives → · See all rsample alternatives →
Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.
They serve adjacent needs but don't currently overlap on shipped themes. Docling is currently shipping more aggressively (velocity 6.3 vs 0.0), with 0 editorial sparks in the last 30 days against 0. See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.
Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. Docling is currently shipping more aggressively (velocity 6.3 vs 0.0), with 0 editorial sparks in the last 30 days against 0. For your specific use case, the alternatives sections above list other ai-assistants products to evaluate alongside.
Top Docling alternatives in ai-assistants are ranked by recent ship velocity. Browse the "Docling alternatives" section above for the current picks, or visit /alternatives/docling for the full list with editorial commentary on each.
Top rsample alternatives in ai-assistants are ranked by recent ship velocity. Browse the "rsample alternatives" section above for the current picks, or visit /alternatives/rsample for the full list with editorial commentary on each.