rsample
tidymodels' resampling package is retiring its old splitters for sliding windows.
A side-by-side editorial comparison of Docling and imbalanced-learn — release velocity, themes, recent moves, and the top alternatives to consider.
Docling keeps widening the funnel: every release adds another format the parser can swallow.
Docling ships roughly weekly, and the shape of each release is consistent — one or two new input formats or model paths, then a dense list of parser corrections for DOCX, ODF, PDF and PPTX. The latest window adds Outlook .msg with optional attachment listing, an EBCDIC backend, VLM grounding output from Unlimited-OCR, and OpenAI logprobs exposed as generated tokens. Underneath, the OCR layer was restructured into a layout-driven pipeline with configurable modes and a RapidOCR refactor that resolves all PP-OCR languages by version and backbone.
The resampling companion to scikit-learn now ships mostly to stay compatible with it.
imbalanced-learn is at 0.14.2. Four of the six releases in the window exist to track a scikit-learn version — 1.5, 1.7, 1.8 and 1.9 in turn — or NumPy 2.0. The genuine additions are thin: InstanceHardnessCV in 0.14.0 and a clearer SMOTENC error when the categorical encoder collapses categories.
Docling ships roughly weekly, and the shape of each release is consistent — one or two new input formats or model paths, then a dense list of parser corrections for DOCX, ODF, PDF and PPTX. The latest window adds Outlook .msg with optional attachment listing, an EBCDIC backend, VLM grounding output from Unlimited-OCR, and OpenAI logprobs exposed as generated tokens. Underneath, the OCR layer was restructured into a layout-driven pipeline with configurable modes and a RapidOCR refactor that resolves all PP-OCR languages by version and backbone.
Two things are being built at once. The conversion surface keeps broadening toward whatever a document actually arrives as — email, mainframe encodings, scanned pages, audio via Whisper — while the service datamodel grows the knobs a hosted pipeline needs: chunking options and targets, PDF heading-level inference, batch connector sources, configurable stage shutdown timeouts. The steady drip of DOCX and ODF reading-order fixes says fidelity, not throughput, is where the hard problems still are.
Expect the format list to keep extending and the VLM and OCR paths to gain more configurability, with reading-order and list-numbering corrections continuing at the same rate. The agent-skills addition suggests more packaging for agent callers, though the entries show only a first step.
imbalanced-learn is at 0.14.2. Four of the six releases in the window exist to track a scikit-learn version — 1.5, 1.7, 1.8 and 1.9 in turn — or NumPy 2.0. The genuine additions are thin: InstanceHardnessCV in 0.14.0 and a clearer SMOTENC error when the categorical encoder collapses categories.
The project has settled into the role of a compatibility shim with a stable sampler catalogue. Release timing is set by upstream scikit-learn, not by its own roadmap, and the deprecations queued in 0.13.0 show the surface narrowing rather than growing.
The pattern points to the next release being another scikit-learn compatibility bump, with the Pipeline check_is_fitted deprecation scheduled to become an error in 0.15.
Other ai-assistants products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either Docling or imbalanced-learn.
tidymodels' resampling package is retiring its old splitters for sliding windows.
tidymodels' preprocessing engine learned sparsity, then settled into deprecations.
parsnip added a whole new regression type, then wired R models to JAX and PyTorch
mlr3 is hardening the seams where its abstractions meet real learners
Mem0 splits agent memory from user memory, then spends a week hardening the plumbing
Every post is a comparison page, and Pictory is always the answer.
See all Docling alternatives → · See all imbalanced-learn alternatives →
Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.
They serve adjacent needs but don't currently overlap on shipped themes. Docling is currently shipping more aggressively (velocity 6.3 vs 0.0), with 0 editorial sparks in the last 30 days against 0. See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.
Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. Docling is currently shipping more aggressively (velocity 6.3 vs 0.0), with 0 editorial sparks in the last 30 days against 0. For your specific use case, the alternatives sections above list other ai-assistants products to evaluate alongside.
Top Docling alternatives in ai-assistants are ranked by recent ship velocity. Browse the "Docling alternatives" section above for the current picks, or visit /alternatives/docling for the full list with editorial commentary on each.
Top imbalanced-learn alternatives in ai-assistants are ranked by recent ship velocity. Browse the "imbalanced-learn alternatives" section above for the current picks, or visit /alternatives/imbalanced-learn for the full list with editorial commentary on each.