← Back to home
Comparison · Analytics

datapack vs refsplitr

A side-by-side editorial comparison of datapack and refsplitr — release velocity, themes, recent moves, and the top alternatives to consider.

datapack vs refsplitr: at a glance

Featuredatapackrefsplitr
SectorAnalyticsAnalytics
Velocity score0.05.0
Sparks · 30d00
Top themesresearch-data, dataone, provenance, bagitbibliometrics, author-disambiguation, georeferencing, ropensci
Last editorial update42m ago2h ago
WebsiteVisit →Visit →

What is datapack?

The DataONE bundler learned to edit packages in 2017 and has coasted on that ever since

datapack assembles heterogeneous data files and metadata into a single transportable bundle, serialised as an OAI-ORE resource map and BagIt archive, for deposit into repositories like DataONE. Its functional surface settled with the 1.3.x line, which made assembled packages editable rather than write-once. Since then the releases have been sparse and defensive: SHA-256 as the default checksum in 1.4.0, BagIt spec conformance in 1.4.1, and a 2025 patch that states outright it contains no new features.

Read the full datapack trajectory →

What is refsplitr?

Author disambiguation for bibliometrics, still grinding on the hard part: which names are the same person.

refsplitr parses Web of Science reference records into tidy data and tries to resolve which author strings belong to the same researcher, then georeferences their institutional addresses for network and map visualizations. The active work is squarely on the disambiguation core: 1.2.3 continues refining the author grouping algorithm and 1.2.1 adjusted ORCID ID matching. Earlier, 1.2.0 replaced the address parsing algorithm and changed the default georeferencing option for author institutions.

Read the full refsplitr trajectory →

datapack vs refsplitr: editorial side-by-side

D
datapack
ANALYTICS
0.0

The DataONE bundler learned to edit packages in 2017 and has coasted on that ever since

◆ Current state

datapack assembles heterogeneous data files and metadata into a single transportable bundle, serialised as an OAI-ORE resource map and BagIt archive, for deposit into repositories like DataONE. Its functional surface settled with the 1.3.x line, which made assembled packages editable rather than write-once. Since then the releases have been sparse and defensive: SHA-256 as the default checksum in 1.4.0, BagIt spec conformance in 1.4.1, and a 2025 patch that states outright it contains no new features.

◆ Where it's heading

The arc runs from assembly to correctness of the resulting archive. Later releases keep tightening the metadata the resource map must carry — dc:creator always present, dcterms:modified always updated, the package correctly flagged as modified after any access-policy change — because a bundle whose provenance record is subtly wrong is worse than one that fails outright. The three-year gap between 1.4.1 and 1.4.2, and the latter's CRAN-note content, place this package firmly in preservation.

◆ Prediction

Expect the next release, if any, to be another CRAN-compliance patch rather than functional work. The 1.4.2 note that it contains no new features is the clearest statement in the feed about where this package sits.

R
refsplitr
ANALYTICS
5.0

Author disambiguation for bibliometrics, still grinding on the hard part: which names are the same person.

◆ Current state

refsplitr parses Web of Science reference records into tidy data and tries to resolve which author strings belong to the same researcher, then georeferences their institutional addresses for network and map visualizations. The active work is squarely on the disambiguation core: 1.2.3 continues refining the author grouping algorithm and 1.2.1 adjusted ORCID ID matching. Earlier, 1.2.0 replaced the address parsing algorithm and changed the default georeferencing option for author institutions.

◆ Where it's heading

Development has narrowed to the two operations that determine whether the output is usable — grouping author name variants and resolving addresses to coordinates. Everything else has been shedding: the maptools dependency was removed once that package was deprecated, and visualization changes are mostly about surfacing records the pipeline could not resolve, as with plot_net_country() returning fixable_countries so users can correct and rerun. Release notes are terse and defer to NEWS, so the changelog itself carries little detail.

◆ Prediction

Expect further incremental passes on author grouping and ORCID matching rather than new outputs; that algorithm is the package's accuracy ceiling and the last several releases have all touched it.

Alternatives to datapack and refsplitr

Other Analytics products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either datapack or refsplitr.

See all datapack alternatives → · See all refsplitr alternatives →

Recent activity from datapack and refsplitr

Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.

  1. 28d agorefsplitrFurther refinement of the author grouping algorithm
  2. 29d agorefsplitrMinor fixes and ORCID matching edits
  3. 10mo agodatapackCRAN documentation and CI cleanup
  4. 1y agorefsplitrNew address parsing algorithm, changed georeferencing default
  5. 2y agorefsplitrUnresolved countries surfaced, maptools dependency dropped
  6. 4y agodatapackBagIt serialisation brought in line with the current spec
  7. 5y agodatapackSHA-256 becomes the default checksum algorithm
  8. 6y agorefsplitrrOpenSci release v0.9.0
  9. 6y agodatapackResource map metadata guaranteed; removeRelationships() added
  10. 8y agodatapackupdateMetadata no longer drops package relationships
  11. 9y agodatapackAssembled data packages become editable in place

Frequently asked questions

What is the difference between datapack and refsplitr?

They serve adjacent needs but don't currently overlap on shipped themes. refsplitr is currently shipping more aggressively (velocity 5.0 vs 0.0), with 0 editorial sparks in the last 30 days against 0. See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.

Is datapack better than refsplitr?

Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. refsplitr is currently shipping more aggressively (velocity 5.0 vs 0.0), with 0 editorial sparks in the last 30 days against 0. For your specific use case, the alternatives sections above list other Analytics products to evaluate alongside.

What are the best alternatives to datapack?

Top datapack alternatives in Analytics are ranked by recent ship velocity. Browse the "datapack alternatives" section above for the current picks, or visit /alternatives/datapack for the full list with editorial commentary on each.

What are the best alternatives to refsplitr?

Top refsplitr alternatives in Analytics are ranked by recent ship velocity. Browse the "refsplitr alternatives" section above for the current picks, or visit /alternatives/refsplitr for the full list with editorial commentary on each.