← Back to all sparks
A

Arize Phoenix

AI-ML-PLATFORMS
Velocity2.6

AI observability and evaluation platform for LLM applications

Phoenix ships near-daily, steadily wiring evals and agent behavior into its traces.

llm observabilitytracingevaluationagentssdk
◆Current state
Phoenix releases the server, client, CLI and MCP packages almost daily, with many entries being dependency bumps. Substantive changes cluster around tracing: eval results in the trace tree, a new DECISION span kind, and bash tool spans flagged as errors on non-zero exits. The TypeScript client keeps gaining project and prompt management helpers.
◆Where it's heading
The trace is becoming the primary place to see evaluation outcomes and agent decisions, not just spans and latency. Management APIs are being filled out so teams can script project configuration rather than click through the UI.
◆Prediction
Expect the DECISION span kind to get dedicated UI or evaluation support in the next few minors, following how eval results reached the trace tree.

◆Recent moves

  1. 5d ago

    New DECISION span kind for agent traces

    Adds a DECISION span kind across the platform and routes datagen applications to their own projects. Gives agent branching points a first-class place in traces.

    View source ↗
  2. 6d ago

    MCP server picks up client 7.16.0

    Dependency bump to the latest client. No functional change to the MCP server.

    View source ↗
  3. 6d ago

    CLI picks up client 7.16.0

    Dependency bump to client 7.16.0. No CLI change.

    View source ↗
  4. 6d ago

    Client: manage annotation configs per project

    Adds helpers to list, assign and set annotation configs per project. Continues filling in scriptable project management.

    View source ↗
  5. 6d ago

    Eval results shown in the trace tree

    Eval results now appear inside the trace tree. Pulls evaluation into the debugging view rather than a separate screen.

    View source ↗
  6. 6d ago

    Failed bash tool spans flagged; experiment sorting

    Bash tool spans with non-zero exit codes are marked as errors, and experiments gain sorting and sequence numbers. Agent traces get more accurate failure signals.

    View source ↗