Baseten
AI model deployment and inference platform for running ML models in production.
Baseten adds a throughput-tuned Fast tier while hardening agent controls and key governance.
◆Recent moves
- 2d ago
GLM 5.2 Fast available on Baseten
⚡ SPARKThe Fast tier is Baseten's clearest bet on agentic serving: the same GLM 5.2 weights, but on dedicated capacity tuned for sustained per-user throughput rather than burst. It extends the agent-native arc from tooling (CLI, MCP) into the serving economics themselves.
View source ↗ - 2d ago
API key management keys
An org-scoped WORKSPACE_MANAGE_API_KEYS key type lets teams create, list, and revoke keys programmatically. Part of the same governance push as admin key visibility and GPU accounting - enterprise plumbing, not a direction change.
View source ↗ - 3d ago
Observability APIs updates
Logs, metrics, and audit logs now pull programmatically through the Management API across any deployment or environment. It rounds out the CLI/MCP story: the telemetry a human reads in the console is now scriptable or agent-driven.
View source ↗ - 4d ago
Workspace GPU usage
A GPU usage tab gives org admins a workspace-wide view of consumption across models and deployments - cost and capacity visibility for teams scaling up, consistent with the month's upmarket governance additions.
View source ↗ - 10d ago
Inkling available on Baseten
Inkling joins the Model APIs catalog via the OpenAI-compatible endpoint, with dedicated deployments for heavier workloads. Routine catalog expansion - Baseten keeps its hosted roster current rather than staking direction on any single model.
View source ↗ - 17d ago
Model API Deprecation (GLM 5.1, GLM 5, Kimi K2.5, Nemotron Super 120B)
Baseten is retiring GLM 5.1, GLM 5, Kimi K2.5, and Nemotron Super 120B from its Model APIs. The counterpart to adding Inkling and GLM 5.2: the hosted catalog rotates toward current models, and teams on the older ones will need to migrate.
View source ↗