← Back to all sparks
C

Comet

AI-ASSISTANTS
Velocity5.0

ML experiment tracking and LLM observability platform, including Opik for evaluating LLM apps.

Comet writes the observability textbook while Opik quietly becomes the product.

llm-observabilityopikagent-tracingcost-intelligenceevaluationdeveloper-education
Current state
The feed is mostly educational and category content — what AI observability is, how to pick a model per task, which observability tools rank in 2026 — with real Opik engineering interleaved. The product work that does appear is specific: Agent Diagnostics for cross-trace analysis, MCP server cost and performance tuning, and Cost Intelligence built out of Comet's own token audit. The old experiment-tracking identity is barely visible.
Where it's heading
Comet has completed a pivot from classic ML experiment tracking to LLM and agent observability, and the content strategy is aimed at owning the category definition while Opik accumulates the features. The recurring theme in the engineering posts is cost — token spend, model selection, MCP optimization — which suggests the wedge is budget pressure rather than debugging alone. Buyer education is running ahead of shipped capability.
Prediction
Given how much of the writing now converges on spend, the next Opik features most likely tie evaluation and tracing directly to cost attribution, so model-selection decisions can be made from the same data that debugs them.

Recent moves

  1. 2d ago

    What is AI Observability? A Complete Guide to Debugging and Monitoring Modern AI Systems at Scale

    A category primer arguing that green infrastructure dashboards tell you nothing about why a model did what it did. Definitional content supporting the observability positioning, with no Opik change attached.

    View source ↗
  2. 4d ago

    LLM Model Selection: How to Pick the Right Model for Every Agentic Task

    Guidance on matching models to agentic tasks instead of defaulting every tool call to the most expensive one. It belongs to the cost thread that runs through the engineering posts, but ships nothing.

    View source ↗
  3. 4d ago

    Best LLM Observability Tools of 2026: Top Platforms & Features

    A ranked roundup of LLM observability platforms for 2026. Search-driven comparison content aimed at buyers evaluating the category Comet is defining.

    View source ↗
  4. 14d ago

    I Built a RAG Pipeline for F1 Team Radio, Then Made It Grade Itself

    A build log taking an F1 team-radio RAG demo to something usable via self-grading. A worked example of evaluation-driven development rather than a product change.

    View source ↗
  5. 29d ago

    One Prompt, 24 Versions: How Digibee Builds Prompts with Opik to Power Their AI-Native Integration Platform

    A customer story on Digibee versioning one prompt 24 times through Opik to run its integration platform. Proof of adoption for existing functionality.

    View source ↗
  6. 1mo ago

    Beyond the Single Trace: How We Built Agent Diagnostics for Opik

    Agent Diagnostics moves Opik past single-trace inspection to spotting patterns across many runs, which is the failure mode operators actually face. It is the clearest shipped capability in the window and sits at the centre of the agent-observability pivot.

    View source ↗