← Back to home
Comparison · Analytics

cleanepi vs nanoparquet

A side-by-side editorial comparison of cleanepi and nanoparquet — release velocity, themes, recent moves, and the top alternatives to consider.

cleanepi vs nanoparquet: at a glance

Featurecleanepinanoparquet
SectorAnalyticsAnalytics
Velocity score0.00.0
Sparks · 30d00
Top themesepiverse-trace, data-cleaning, line-list, reportingparquet, r-language, interoperability, data-formats
Last editorial update5h ago47m ago
WebsiteVisit →Visit →

What is cleanepi?

cleanepi is in the long tail of bug fixes that follows a 1.0 — and changed maintainers along the way.

cleanepi cleans and standardises epidemiological line list data — dates, subject IDs, missing values, duplicates — and produces a report of what it changed. Since 1.0.0 in mid-2024 the releases have been almost entirely corrective: date-guesser fixes, report structure fixes and matching behaviour corrections. Maintainership passed to Bubacarr Bah in 1.1.2.

Read the full cleanepi trajectory →

What is nanoparquet?

nanoparquet is chasing byte-level agreement with the Java and Rust Parquet readers, not feature count.

nanoparquet reads and writes Parquet from R with no Arrow dependency, which is its entire reason to exist. The 0.4.0 line renamed the reader API and added schema authoring plus `append_parquet()`, and the 0.5.x releases have gone after interoperability: definition and repetition level encodings the Apache Parquet Java library expects, flatbuffer alignment the Rust arrow-rs reader expects, 128-bit decimals, and Polars-written files that omit the dictionary page offset. The newest release adds `bit64::integer64` columns and writing to stdout.

Read the full nanoparquet trajectory →

cleanepi vs nanoparquet: editorial side-by-side

C
cleanepi
ANALYTICS
0.0

cleanepi is in the long tail of bug fixes that follows a 1.0 — and changed maintainers along the way.

◆ Current state

cleanepi cleans and standardises epidemiological line list data — dates, subject IDs, missing values, duplicates — and produces a report of what it changed. Since 1.0.0 in mid-2024 the releases have been almost entirely corrective: date-guesser fixes, report structure fixes and matching behaviour corrections. Maintainership passed to Bubacarr Bah in 1.1.2.

◆ Where it's heading

Work has concentrated on the report object and on making the cleaning functions behave predictably at the edges — case- and whitespace-insensitive missing-value matching, report elements returned as vectors instead of comma-separated strings, an argument to print a single operation's report. The underlying cleaning API has barely moved since 1.0.0, which suggests it is settled.

◆ Prediction

The report interface has been reworked repeatedly across these releases and is the most likely place for further change; the cleaning functions themselves look stable.

N
nanoparquet
ANALYTICS
0.0

nanoparquet is chasing byte-level agreement with the Java and Rust Parquet readers, not feature count.

◆ Current state

nanoparquet reads and writes Parquet from R with no Arrow dependency, which is its entire reason to exist. The 0.4.0 line renamed the reader API and added schema authoring plus `append_parquet()`, and the 0.5.x releases have gone after interoperability: definition and repetition level encodings the Apache Parquet Java library expects, flatbuffer alignment the Rust arrow-rs reader expects, 128-bit decimals, and Polars-written files that omit the dictionary page offset. The newest release adds `bit64::integer64` columns and writing to stdout.

◆ Where it's heading

Almost every entry since 0.4.0 names another engine — Java, arrow-rs, Polars, Arrow schema metadata — which tells you the maintainers are treating cross-reader fidelity as the product rather than R-side ergonomics. The type system is filling in from the edges: DECIMAL beyond 8 bytes, UUID, FLOAT16 and INTERVAL as raw lists, and now 64-bit integers with an explicit read-type option instead of a silent cast to double. Writing to `:stdout:` points at a second audience, shell pipelines rather than interactive R.

◆ Prediction

The remaining unmapped Parquet types the changelog has been parking in raw-vector lists — FLOAT16 and INTERVAL — are the obvious next targets, following the same pattern by which DECIMAL and UUID graduated to real R types.

Alternatives to cleanepi and nanoparquet

Other Analytics products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either cleanepi or nanoparquet.

See all cleanepi alternatives → · See all nanoparquet alternatives →

Recent activity from cleanepi and nanoparquet

Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.

  1. 3mo agonanoparquet64-bit integer columns and writing Parquet to stdout
  2. 4mo agonanoparquetFiles now readable by the Java and Rust Parquet libraries
  3. 9mo agocleanepiCase-insensitive missing-value matching and report fixes
  4. 1y agocleanepiReport elements become vectors; date parsing default restored
  5. 1y agocleanepiDate guesser corrected; empty-row indices fixed
  6. 1y agonanoparquetReads Polars files that omit the dictionary page offset
  7. 1y agonanoparquetDate, FLOAT, and mixed-encoding read fixes
  8. 1y agonanoparquetSchema authoring and append_parquet arrive with a renamed API
  9. 1y agonanoparquetFixes a write_parquet crash
  10. 2y agocleanepiFirst major release with cleaning and reporting

Frequently asked questions

What is the difference between cleanepi and nanoparquet?

They serve adjacent needs but don't currently overlap on shipped themes. cleanepi and nanoparquet are shipping at a similar cadence (velocity 0.0 vs 0.0, both within Sparkpulse's "active" band). See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.

Is cleanepi better than nanoparquet?

Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. cleanepi and nanoparquet are shipping at a similar cadence (velocity 0.0 vs 0.0, both within Sparkpulse's "active" band). For your specific use case, the alternatives sections above list other Analytics products to evaluate alongside.

What are the best alternatives to cleanepi?

Top cleanepi alternatives in Analytics are ranked by recent ship velocity. Browse the "cleanepi alternatives" section above for the current picks, or visit /alternatives/cleanepi for the full list with editorial commentary on each.

What are the best alternatives to nanoparquet?

Top nanoparquet alternatives in Analytics are ranked by recent ship velocity. Browse the "nanoparquet alternatives" section above for the current picks, or visit /alternatives/nanoparquet for the full list with editorial commentary on each.