← Back to home
Comparison · Analytics

refsplitr vs textreuse

A side-by-side editorial comparison of refsplitr and textreuse — release velocity, themes, recent moves, and the top alternatives to consider.

Shared themes:r-package

refsplitr vs textreuse: at a glance

Featurerefsplitrtextreuse
SectorAnalyticsAnalytics
Velocity score5.02.5
Sparks · 30d00
Top themesbibliometrics, author-disambiguation, georeferencing, ropenscitext-reuse, minhash, lsh, r-package
Last editorial update2h ago2h ago
WebsiteVisit →Visit →

What is refsplitr?

Author disambiguation for bibliometrics, still grinding on the hard part: which names are the same person.

refsplitr parses Web of Science reference records into tidy data and tries to resolve which author strings belong to the same researcher, then georeferences their institutional addresses for network and map visualizations. The active work is squarely on the disambiguation core: 1.2.3 continues refining the author grouping algorithm and 1.2.1 adjusted ORCID ID matching. Earlier, 1.2.0 replaced the address parsing algorithm and changed the default georeferencing option for author institutions.

Read the full refsplitr trajectory →

What is textreuse?

A dormant text-matching package revived, shipped as 1.0.0, and kept current with the tidyverse.

textreuse detects reused and quoted passages across document collections using minhash and locality-sensitive hashing, with local alignment for inspecting the matches it finds. After years of inactivity, the package reached a 1.0.0 CRAN release in May 2026 that folded accumulated feature work into one version — encoding control on corpus construction, deterministic skipped-document bookkeeping, and an align_local() that returns an empty alignment instead of erroring on non-matching texts. The 1.0.2 release since then is pure compatibility maintenance.

Read the full textreuse trajectory →

refsplitr vs textreuse: editorial side-by-side

R
refsplitr
ANALYTICS
5.0

Author disambiguation for bibliometrics, still grinding on the hard part: which names are the same person.

◆ Current state

refsplitr parses Web of Science reference records into tidy data and tries to resolve which author strings belong to the same researcher, then georeferences their institutional addresses for network and map visualizations. The active work is squarely on the disambiguation core: 1.2.3 continues refining the author grouping algorithm and 1.2.1 adjusted ORCID ID matching. Earlier, 1.2.0 replaced the address parsing algorithm and changed the default georeferencing option for author institutions.

◆ Where it's heading

Development has narrowed to the two operations that determine whether the output is usable — grouping author name variants and resolving addresses to coordinates. Everything else has been shedding: the maptools dependency was removed once that package was deprecated, and visualization changes are mostly about surfacing records the pipeline could not resolve, as with plot_net_country() returning fixable_countries so users can correct and rerun. Release notes are terse and defer to NEWS, so the changelog itself carries little detail.

◆ Prediction

Expect further incremental passes on author grouping and ORCID matching rather than new outputs; that algorithm is the package's accuracy ceiling and the last several releases have all touched it.

T
textreuse
ANALYTICS
2.5

A dormant text-matching package revived, shipped as 1.0.0, and kept current with the tidyverse.

◆ Current state

textreuse detects reused and quoted passages across document collections using minhash and locality-sensitive hashing, with local alignment for inspecting the matches it finds. After years of inactivity, the package reached a 1.0.0 CRAN release in May 2026 that folded accumulated feature work into one version — encoding control on corpus construction, deterministic skipped-document bookkeeping, and an align_local() that returns an empty alignment instead of erroring on non-matching texts. The 1.0.2 release since then is pure compatibility maintenance.

◆ Where it's heading

The arc here is restoration rather than expansion. The work has gone into making the package survivable — silencing deprecated dplyr and tidyr selection and many-to-many join warnings, moving from dead Travis and AppVeyor configs to GitHub Actions, and validating across five R platform and version combinations. Release notes now lead with verification evidence rather than features, which is the signature of a maintainer stabilizing an inherited codebase.

◆ Prediction

Expect continued compatibility releases tracking tidyverse deprecations; nothing in these entries indicates new hashing or alignment capability is planned.

Alternatives to refsplitr and textreuse

Other Analytics products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either refsplitr or textreuse.

See all refsplitr alternatives → · See all textreuse alternatives →

Recent activity from refsplitr and textreuse

Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.

  1. 20d agotextreuseCompatibility pass for current dplyr and tidyr
  2. 28d agorefsplitrFurther refinement of the author grouping algorithm
  3. 29d agorefsplitrMinor fixes and ORCID matching edits
  4. 3mo agotextreuseCRAN resubmission fixing a moved README URL
  5. 3mo agotextreuse1.0.0 consolidates years of accumulated feature work
  6. 1y agorefsplitrNew address parsing algorithm, changed georeferencing default
  7. 2y agorefsplitrUnresolved countries surfaced, maptools dependency dropped
  8. 6y agorefsplitrrOpenSci release v0.9.0
  9. 10y agotextreuseMinhashes split out from hashes in document objects

Frequently asked questions

What is the difference between refsplitr and textreuse?

Both compete on the same themes — r-package — within Analytics. refsplitr is currently shipping more aggressively (velocity 5.0 vs 2.5), with 0 editorial sparks in the last 30 days against 0. See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.

Is refsplitr better than textreuse?

Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. refsplitr is currently shipping more aggressively (velocity 5.0 vs 2.5), with 0 editorial sparks in the last 30 days against 0. For your specific use case, the alternatives sections above list other Analytics products to evaluate alongside.

What are the best alternatives to refsplitr?

Top refsplitr alternatives in Analytics are ranked by recent ship velocity. Browse the "refsplitr alternatives" section above for the current picks, or visit /alternatives/refsplitr for the full list with editorial commentary on each.

What are the best alternatives to textreuse?

Top textreuse alternatives in Analytics are ranked by recent ship velocity. Browse the "textreuse alternatives" section above for the current picks, or visit /alternatives/textreuse for the full list with editorial commentary on each.