Langfuse
Langfuse promotes Experiments out from under Datasets, making evaluation the primary workflow.
◆Recent moves
- 3mo ago
Experiments promoted to a top-level feature
A duplicate row for the Experiments rebuild already covered by the dated changelog entry from April 13. Same release, second listing, no additional detail.
- 3mo ago
Boolean scores for LLM-as-a-Judge evaluators
A duplicate row for the boolean LLM-as-a-Judge scores shipped on April 8. Same release, second listing.
- 4mo ago
Experiments as a First-Class Concept
⚡ SPARKExperiments stop being a mode of Datasets and become their own top-level feature, runnable without a dataset and comparable across runs. This is the structural change the score-type work has been feeding into — evaluation moves from a dataset chore to the primary workflow.
View source ↗ - 4mo ago
Boolean LLM-as-a-Judge Scores
LLM-as-a-Judge evaluators can return boolean true/false scores, arriving a week after categorical scores landed. Together they let a judge record a verdict rather than force every judgment onto a numeric scale.
View source ↗ - 4mo ago
Reference: dashboard behavior under Fast Preview
A documentation reference describing how dashboards differ under Fast Preview — trace counts, histograms, filters. Useful for interpreting the UI, but nothing shipped.
- 4mo ago
Roadmap threads1.1k
Not a release — a page fragment scraped from the site's navigation showing a roadmap thread count. The feed parser is picking up chrome alongside changelog items.