← Back to analytics tools
Alternatives · analytics tools

textreuse alternatives

The best textreuse alternatives in analytics tools, ranked by Sparkpulse's velocity_score.

Updated Aug 14, 2026

Looking for the best alternatives to textreuse? Sparkpulse tracks and ranks 12 alternatives in analytics tools by shipping velocity — how frequently each ships meaningful updates, verified from official changelogs. For reference, textreuse shipped 0 meaningful updates in the last 30 days and carries a velocity score of 2.5 out of 10 in 2026. The alternatives below are ranked the same way, so you're comparing real release momentum, not marketing claims.

About textreuse

A dormant text-matching package revived, shipped as 1.0.0, and kept current with the tidyverse.

textreuse detects reused and quoted passages across document collections using minhash and locality-sensitive hashing, with local alignment for inspecting the matches it finds. After years of inactivity, the package reached a 1.0.0 CRAN release in May 2026 that folded accumulated feature work into one version — encoding control on corpus construction, deterministic skipped-document bookkeeping, and an align_local() that returns an empty alignment instead of erroring on non-matching texts. The 1.0.2 release since then is pure compatibility maintenance.

Velocity 2.5 · Last update 1h ago

Read the full textreuse trajectory →

Top 12 alternatives to textreuse

Ranked by recent ship velocity. Tap any card for the full editorial breakdown, or pivot to a head-to-head.

Browse all analytics tools products →

textreuse vs alternatives — shipping velocity at a glance

Velocity score (0–10) and meaningful releases shipped in the last 30 days, from official changelogs. Higher = shipping faster.

ProductVelocitySparks · 30dFocus areasLatest release
textreuse (baseline)2.50text-reuseminhashlsh
refsplitr5.00bibliometricsauthor-disambiguationgeoreferencing
pysparklyr3.81sparkdatabrickssnowflaketune_grid_spark() runs tidymodels tuning on Spark Connect
rbmi2.50clinical-trialsmissing-datamultiple-imputationrstan demoted to Suggests, Bayesian imputation now opt-in
tabnet2.50tabular-deep-learningtorchtidymodelsHierarchical multi-label classification via data.tree
osmapiR0.00openstreetmapapi-clientr-languageComplete OSM API coverage arrives in one release
ymlthis0.00r-markdownyamlretirementymlthis retired; Quarto covers the need
forestly0.00clinical-safetyadverse-eventsdata-visualizationStatic RTF forest plots join the interactive output
pharmaverseadam0.00pharmaverseclinical-dataadam
pkglite0.00r-packagespharma-submissionspackaging
gMCPLite0.00multiple-comparisonsclinical-trialsr-language
typst-gather0.00typsthermetic-buildsrustanalyze subcommand reports the import graph as JSON
naijR0.00nigeriageospatialreference-data

The 12 best textreuse alternatives, in depth

1. refsplitr · velocity 5.0

Author disambiguation for bibliometrics, still grinding on the hard part: which names are the same person.

Its velocity score of 5.0/10 reflects longer-term release cadence.

Where textreuse leans on text reuse, minhash and lsh, refsplitr focuses on bibliometrics, author disambiguation and georeferencing.

refsplitr and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

2. pysparklyr · velocity 3.8

Posit's Spark Connect bridge keeps adding backends — and now runs tidymodels tuning on the cluster.

Over the last 30 days pysparklyr shipped 1 meaningful update vs textreuse's 0, most recently “tune_grid_spark() runs tidymodels tuning on Spark Connect”. Its velocity score of 3.8/10 blends that with longer-term release cadence.

Where textreuse leans on text reuse, minhash and lsh, pysparklyr focuses on spark, databricks and snowflake.

Over the last 30 days pysparklyr has been shipping faster than textreuse — a point in its favour if release momentum matters to you.

3. rbmi · velocity 2.5

Reference-based multiple imputation for trials, now shipping without Bayesian support by default.

Its velocity score of 2.5/10 reflects longer-term release cadence; its most recent meaningful update was “rstan demoted to Suggests, Bayesian imputation now opt-in”.

Where textreuse leans on text reuse, minhash and lsh, rbmi focuses on clinical trials, missing data and multiple imputation.

rbmi and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

4. tabnet · velocity 2.5

A tabular deep-learning model in R that keeps widening what counts as a tabular task.

Its velocity score of 2.5/10 reflects longer-term release cadence; its most recent meaningful update was “Hierarchical multi-label classification via data.tree”.

Where textreuse leans on text reuse, minhash and lsh, tabnet focuses on tabular deep learning, torch and tidymodels.

tabnet and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

5. osmapiR · velocity 0.0

OsmapiR is the rare API client that tracks its server's wiki revision numbers in the changelog.

Its velocity score of 0.0/10 reflects longer-term release cadence; its most recent meaningful update was “Complete OSM API coverage arrives in one release”.

Where textreuse leans on text reuse, minhash and lsh, osmapiR focuses on openstreetmap, api client and r language.

osmapiR and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

6. ymlthis · velocity 0.0

Ymlthis retired itself, naming Quarto as the reason it no longer needs to exist.

Its velocity score of 0.0/10 reflects longer-term release cadence; its most recent meaningful update was “ymlthis retired; Quarto covers the need”.

Where textreuse leans on text reuse, minhash and lsh, ymlthis focuses on r markdown, yaml and retirement.

ymlthis and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

7. forestly · velocity 0.0

Forestly built an interactive safety review tool, then taught it to produce submission-ready RTF.

Its velocity score of 0.0/10 reflects longer-term release cadence; its most recent meaningful update was “Static RTF forest plots join the interactive output”.

Where textreuse leans on text reuse, minhash and lsh, forestly focuses on clinical safety, adverse events and data visualization.

forestly and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

8. pharmaverseadam · velocity 0.0

Pharmaverseadam is the pharmaverse's test-data mirror, and it now covers neurology.

Its velocity score of 0.0/10 reflects longer-term release cadence.

Where textreuse leans on text reuse, minhash and lsh, pharmaverseadam focuses on pharmaverse, clinical data and adam.

pharmaverseadam and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

9. pkglite · velocity 0.0

Pkglite's whole job is knowing which files in an R package are text — and it keeps getting better at guessing.

Its velocity score of 0.0/10 reflects longer-term release cadence.

Where textreuse leans on text reuse, minhash and lsh, pkglite focuses on r packages, pharma submissions and packaging.

pkglite and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

10. gMCPLite · velocity 0.0

GMCPLite exists to be gMCP without Java, and its releases guard that boundary rather than extend it.

Its velocity score of 0.0/10 reflects longer-term release cadence.

Where textreuse leans on text reuse, minhash and lsh, gMCPLite focuses on multiple comparisons, clinical trials and r language.

gMCPLite and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

11. typst-gather · velocity 0.0

Extracted from Quarto's CLI, typst-gather learned to explain a dependency tree before fetching it.

Its velocity score of 0.0/10 reflects longer-term release cadence; its most recent meaningful update was “analyze subcommand reports the import graph as JSON”.

Where textreuse leans on text reuse, minhash and lsh, typst-gather focuses on typst, hermetic builds and rust.

typst-gather and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

12. naijR · velocity 0.0

NaijR is assembling the Nigerian reference data R analysts otherwise hand-code every time.

Its velocity score of 0.0/10 reflects longer-term release cadence.

Where textreuse leans on text reuse, minhash and lsh, naijR focuses on nigeria, geospatial and reference data.

naijR and textreuse have shipped at a similar pace over the last 30 days, so the decision comes down to fit and feature depth.

Frequently asked questions

What are the best alternatives to textreuse?

The top textreuse alternatives we currently track in analytics tools are refsplitr, pysparklyr, rbmi, tabnet, osmapiR, ranked by recent ship velocity.

How is this list of textreuse alternatives ranked?

Alternatives are ranked by Sparkpulse's velocity_score — release cadence + 30-day spark count + sector-relative ship rate.

Can I compare textreuse directly with one of these alternatives?

Yes — every card has a "Compare with textreuse" link to a side-by-side /compare page.