← Back to all sparks
G

Groq

INFRA · APIS
Velocity0.0

Ultra-fast AI inference API for running LLMs using custom hardware.

Groq is climbing off raw inference: hosted tools and MCP connectors now sit above the chips.

inferencemcpbuilt-in-toolsmodel-hostingttsenterprise
Current state
GroqCloud's changelog splits three ways: model availability (MiniMax M2.5, Qwen3-VL for enterprise, GPT-OSS variants), speech work consolidating on Orpheus after the PlayAI deprecation, and a newer line of server-side capabilities — browser search as a built-in tool, and MCP connectors for Google Workspace. SDK releases are maintenance-grade following December 2025's 1.0 GA. Cadence has thinned since Q1, with the most recent captured entries dating to April.
Where it's heading
The interesting movement is Groq offering things that execute rather than only generate: tools the model can call inside Groq's own runtime, and pre-built OAuth connectors to business applications. That is a bid to be the place agents run, not merely the fastest token endpoint, and it competes with the same agent-runtime pitch Bedrock and OpenAI's Responses API make. Speed remains the underlying claim, but the surface being sold is moving up the stack.
Prediction
Expect the connector catalog to widen beyond Google Workspace and more built-in tools to reach the general model lineup rather than only GPT-OSS.

Recent moves

  1. 3mo ago

    Browser Search (GPT OSS Models)

    Browser search becomes a built-in tool for GPT-OSS models, letting the model fetch from the web inside Groq's runtime instead of through caller-side plumbing. Captured from the docs page rather than a release note, so detail is thin, but it fits the move toward server-side execution.

    View source ↗
  2. 3mo ago

    Python SDK v1.2.0 and TypeScript SDK v1.1.2

    Quarterly roll-up of Python and TypeScript SDK fixes covering query-param merging, file upload handling, and general stability. Post-GA maintenance with no API surface change.

    View source ↗
  3. 3mo ago

    MiniMax M2.5 and Qwen3-VL 32B Instruct (Enterprise)

    MiniMax M2.5 and a 32B Qwen3-VL vision-language model land on GroqCloud for enterprise customers, gated behind account teams. Routine roster expansion, notable mainly for putting multimodal capacity behind the enterprise tier.

    View source ↗
  4. 3mo ago

    New Voices for Orpheus Arabic Saudi

    Two more voices, including a new default, bring the Arabic Saudi Orpheus model to six. Incremental, but it is the kind of coverage work that follows a platform-wide TTS migration.

    View source ↗
  5. 4mo ago

    Built-In Tools

    A docs index page for compound built-in tools captured as an entry, with no release content attached. Useful only as a marker that the built-in tool surface is being documented.

    View source ↗
  6. 6mo ago

    Platform-wide Migration from PlayAI to Orpheus TTS

    The platform completes its migration from PlayAI to Canopy Labs' Orpheus TTS, retiring playai-tts and playai-tts-arabic per the December 2025 deprecation. Orpheus brings vocal-direction controls, faster inference, and improved audio quality. A clean execution of a planned model swap rather than a directional change.

    View source ↗