← Back to home
Comparison · Analytics

collapse vs spanishoddata

A side-by-side editorial comparison of collapse and spanishoddata — release velocity, themes, recent moves, and the top alternatives to consider.

Shared themes:r-package

collapse vs spanishoddata: at a glance

Featurecollapsespanishoddata
SectorAnalyticsAnalytics
Velocity score0.00.0
Sparks · 30d00
Top themesdata-transformation, performance, simd, grouped-statisticsmobility-data, origin-destination, duckdb, open-data
Last editorial update1h ago1h ago
WebsiteVisit →Visit →

What is collapse?

collapse got a JSS paper and a 7x fmean speedup in the same release.

collapse provides fast grouped statistical computing and data transformation for R, built on a C backend with its own grouping, hashing and aggregation primitives. The 2.1.x line is a maintenance and optimization series: SIMD multiple-accumulator work delivering roughly 2x on fsum() and 7x on fmean() for systems without OpenMP, a custom internal unlist() with better attribute preservation, and a steady stream of correctness fixes in collap(), pivot() and roworderv().

Read the full collapse trajectory →

What is spanishoddata?

spanishoddata spent a year finding out its 2020-2021 data was quietly incomplete.

spanishoddata provides access to Spain's open mobility origin-destination datasets from the Ministry of Transport, converting them into DuckDB and parquet for analysis at scale. Nearly every release in this window is a data-fidelity fix rather than a feature: district-to-municipal reaggregation was wrong for the 2020-2021 vintage, literal 'NA' strings in the source CSVs broke DuckDB enum casting, and the Amazon S3 metadata bucket turned out to be truncated at March 2021, silently hiding data.

Read the full spanishoddata trajectory →

collapse vs spanishoddata: editorial side-by-side

C
collapse
ANALYTICS
0.0

collapse got a JSS paper and a 7x fmean speedup in the same release.

◆ Current state

collapse provides fast grouped statistical computing and data transformation for R, built on a C backend with its own grouping, hashing and aggregation primitives. The 2.1.x line is a maintenance and optimization series: SIMD multiple-accumulator work delivering roughly 2x on fsum() and 7x on fmean() for systems without OpenMP, a custom internal unlist() with better attribute preservation, and a steady stream of correctness fixes in collap(), pivot() and roworderv().

◆ Where it's heading

The package is consolidating institutionally as much as technically. The repository moved to the fastverse organization with multiple people granted access, the Journal of Statistical Software paper landed as the primary citation, and documentation now includes an AI-generated interactive layer. Technically the focus is the hashing and grouping core — the decision to treat -0 and 0 as equal across funique(), group(), fmatch(), fmode() and their derivatives was made in sync with an equivalent change in Rcpp, and accepted a measured 3% cost to get it. The last release with breaking changes sits outside this six-entry window.

◆ Prediction

Expect further targeted performance work on the grouped statistical functions and continued small correctness fixes; the governance move to fastverse suggests contribution volume rather than direction is what the maintainer is managing.

S
spanishoddata
ANALYTICS
0.0

spanishoddata spent a year finding out its 2020-2021 data was quietly incomplete.

◆ Current state

spanishoddata provides access to Spain's open mobility origin-destination datasets from the Ministry of Transport, converting them into DuckDB and parquet for analysis at scale. Nearly every release in this window is a data-fidelity fix rather than a feature: district-to-municipal reaggregation was wrong for the 2020-2021 vintage, literal 'NA' strings in the source CSVs broke DuckDB enum casting, and the Amazon S3 metadata bucket turned out to be truncated at March 2021, silently hiding data.

◆ Where it's heading

The package is in a trust-building phase. The pattern across 0.2.1 through 0.2.6 is the maintainers repeatedly discovering that upstream metadata and the package's own aggregation were misrepresenting what data existed, then fixing it and adding a check so it surfaces next time. That is now backed by infrastructure: comprehensive unit tests plus weekly live-data runs on GitHub workers that alert maintainers when the upstream ministry changes something. The last feature release sits outside the six-entry window, which is itself the story.

◆ Prediction

Expect continued upstream-tracking fixes as the ministry's API and S3 layout shift, with the experimental quick-access and checksum functions the most likely candidates for promotion to stable.

Alternatives to collapse and spanishoddata

Other Analytics products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either collapse or spanishoddata.

See all collapse alternatives → · See all spanishoddata alternatives →

Recent activity from collapse and spanishoddata

Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.

  1. 2mo agospanishoddataRedirected district files identified in v1 metadata
  2. 2mo agospanishoddataS3 metadata truncation at March 2021 bypassed via XML feed
  3. 2mo agocollapseSIMD accumulators give fmean a 7x speedup without OpenMP
  4. 4mo agospanishoddataLiteral NA strings no longer break DuckDB casting
  5. 4mo agospanishoddataLarge urban area zones can be reloaded again
  6. 5mo agospanishoddatatime_slot column removed; test coverage goes live-weekly
  7. 7mo agocollapseNegative zero now hashes equal to zero across the package
  8. 8mo agocollapsecollap() no longer double-aggregates external weights
  9. 9mo agocollapseCustom unlist() preserves attributes
  10. 0y agocollapseAssorted bug fixes
  11. 1y agospanishoddataDistrict-to-municipal reaggregation corrected for 2020-2021
  12. 1y agocollapsena_insert gains by-reference mode; gsplit and pivot speed up

Frequently asked questions

What is the difference between collapse and spanishoddata?

Both compete on the same themes — r-package — within Analytics. collapse and spanishoddata are shipping at a similar cadence (velocity 0.0 vs 0.0, both within Sparkpulse's "active" band). See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.

Is collapse better than spanishoddata?

Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. collapse and spanishoddata are shipping at a similar cadence (velocity 0.0 vs 0.0, both within Sparkpulse's "active" band). For your specific use case, the alternatives sections above list other Analytics products to evaluate alongside.

What are the best alternatives to collapse?

Top collapse alternatives in Analytics are ranked by recent ship velocity. Browse the "collapse alternatives" section above for the current picks, or visit /alternatives/collapse-r for the full list with editorial commentary on each.

What are the best alternatives to spanishoddata?

Top spanishoddata alternatives in Analytics are ranked by recent ship velocity. Browse the "spanishoddata alternatives" section above for the current picks, or visit /alternatives/spanishoddata for the full list with editorial commentary on each.