Baseten
AI model deployment and inference platform for running ML models in production.
Baseten adds AWS AssumeRole auth and a Viewer role — enterprise governance for model inference.
◆Recent moves
- 8d ago
Viewer role for read-only access
A new Viewer role lets teammates access Baseten resources — invoke models, inspect configurations — without deployment or write permissions. This fills the gap between full access and no access, enabling wider team visibility without security risk.
View source ↗ - 8d ago
AWS AssumeRole authentication
Baseten can now assume an IAM role in a customer's AWS account for ECR image pulls and S3 weight mirroring during builds, replacing the need for long-lived access keys. This is the AWS-native security pattern enterprise teams require before allowing a third party to touch their private model artifacts.
View source ↗ - 12d ago
GLM 5.3 available on Baseten
GLM 5.3, Z.ai's flagship model with a 1M-token context window, is now available through Baseten's OpenAI-compatible endpoint. The reasoning effort control (low/high/max) gives developers cost-vs-quality flexibility from the same model.
View source ↗ - 14d ago
GLM 5.3 Flash available on Baseten
GLM 5.3 Flash is available via the OpenAI-compatible endpoint — a faster, lower-cost variant of Z.ai's latest model. Expanding the GLM model family gives Baseten users a full performance tier to choose from within the same model series.
View source ↗ - 14d ago
GLM 5.3 Fast: third tier of the GLM 5.3 family
GLM 5.3 Fast (a third speed/cost tier in the GLM 5.3 family) is available on the same OpenAI-compatible endpoint. Rounds out the GLM 5.3 lineup already added this week.
View source ↗ - 15d ago
Autoscaling schedules
Autoscaling schedules are now generally available, letting environments automatically adjust capacity on a defined time window. Pre-warming capacity before predictable traffic spikes reduces cold-start latency without manual intervention before each high-demand period.
View source ↗