← Back to all sparks
L

Langfuse

INFRA · APIS
Velocity0.0

Langfuse promotes Experiments out from under Datasets, making evaluation the primary workflow.

llm-observabilityevaluationllm-as-a-judgeexperimentstracing
Current state
Langfuse's recent work is concentrated almost entirely on the evaluation surface. Experiments were rebuilt as a top-level feature that runs with or without a dataset attached, and can be compared across runs over time. The LLM-as-a-Judge evaluator gained categorical scores in late March and boolean true/false scores a week later, filling out the score types beyond plain numerics. Everything else in the window is documentation or scrape artifacts.
Where it's heading
The direction is evaluation as the product's centre of gravity rather than an appendage to tracing. Decoupling Experiments from Datasets removes the setup cost of running an eval, and the widening score types let judges express verdicts rather than only magnitudes — both point at teams running evals continuously against live traces instead of curated fixtures. Regional expansion shows up in the feed as Langfuse Cloud Japan. Cadence is the open question: nothing has published since April 21, so this arc is described from a three-month-old window.
Prediction
The score-type buildout and the run-comparison view are converging on scheduled or triggered evaluations against production traces, but the feed has been silent long enough that the next move cannot be called with confidence from these entries alone.

Recent moves

  1. 3mo ago

    Experiments promoted to a top-level feature

    A duplicate row for the Experiments rebuild already covered by the dated changelog entry from April 13. Same release, second listing, no additional detail.

  2. 3mo ago

    Boolean scores for LLM-as-a-Judge evaluators

    A duplicate row for the boolean LLM-as-a-Judge scores shipped on April 8. Same release, second listing.

  3. 4mo ago

    Experiments as a First-Class Concept

    ⚡ SPARK

    Experiments stop being a mode of Datasets and become their own top-level feature, runnable without a dataset and comparable across runs. This is the structural change the score-type work has been feeding into — evaluation moves from a dataset chore to the primary workflow.

    View source ↗
  4. 4mo ago

    Boolean LLM-as-a-Judge Scores

    LLM-as-a-Judge evaluators can return boolean true/false scores, arriving a week after categorical scores landed. Together they let a judge record a verdict rather than force every judgment onto a numeric scale.

    View source ↗
  5. 4mo ago

    Reference: dashboard behavior under Fast Preview

    A documentation reference describing how dashboards differ under Fast Preview — trace counts, histograms, filters. Useful for interpreting the UI, but nothing shipped.

  6. 4mo ago

    Roadmap threads1.1k

    Not a release — a page fragment scraped from the site's navigation showing a roadmap thread count. The feed parser is picking up chrome alongside changelog items.