← Back to all sparks
B

Baseten

AI-ASSISTANTS
Velocity6.3

AI model deployment and inference platform for running ML models in production.

Baseten adds a throughput-tuned Fast tier while hardening agent controls and key governance.

model-inferenceagent-nativethroughputobservabilityenterprise-governance
Current state
Baseten is a model-inference platform pushing two fronts at once: agent-native operation (MCP server, CLI, coding-agent skill) and enterprise governance (org-scoped API keys, admin key visibility, workspace GPU accounting). Its Model APIs catalog rotates quickly - Inkling and GLM 5.2 in, four older models out.
Where it's heading
July's releases point at Baseten positioning as the serving layer for agentic workloads. The new Fast tier sells dedicated capacity on sustained per-user throughput, while the Management API and CLI make deploy, observe, and tune loops scriptable or agent-driven. Governance features - key scoping, GPU visibility - signal a push upmarket to teams that need audit and cost control.
Prediction
Expect the Fast tier to widen beyond GLM 5.2 to more high-demand models, and continued Management-API growth so a coding agent can run the full deploy/observe/tune loop without touching the console.

Recent moves

  1. 2d ago

    GLM 5.2 Fast available on Baseten

    ⚡ SPARK

    The Fast tier is Baseten's clearest bet on agentic serving: the same GLM 5.2 weights, but on dedicated capacity tuned for sustained per-user throughput rather than burst. It extends the agent-native arc from tooling (CLI, MCP) into the serving economics themselves.

    View source ↗
  2. 2d ago

    API key management keys

    An org-scoped WORKSPACE_MANAGE_API_KEYS key type lets teams create, list, and revoke keys programmatically. Part of the same governance push as admin key visibility and GPU accounting - enterprise plumbing, not a direction change.

    View source ↗
  3. 3d ago

    Observability APIs updates

    Logs, metrics, and audit logs now pull programmatically through the Management API across any deployment or environment. It rounds out the CLI/MCP story: the telemetry a human reads in the console is now scriptable or agent-driven.

    View source ↗
  4. 4d ago

    Workspace GPU usage

    A GPU usage tab gives org admins a workspace-wide view of consumption across models and deployments - cost and capacity visibility for teams scaling up, consistent with the month's upmarket governance additions.

    View source ↗
  5. 10d ago

    Inkling available on Baseten

    Inkling joins the Model APIs catalog via the OpenAI-compatible endpoint, with dedicated deployments for heavier workloads. Routine catalog expansion - Baseten keeps its hosted roster current rather than staking direction on any single model.

    View source ↗
  6. 17d ago

    Model API Deprecation (GLM 5.1, GLM 5, Kimi K2.5, Nemotron Super 120B)

    Baseten is retiring GLM 5.1, GLM 5, Kimi K2.5, and Nemotron Super 120B from its Model APIs. The counterpart to adding Inkling and GLM 5.2: the hosted catalog rotates toward current models, and teams on the older ones will need to migrate.

    View source ↗