← Back to all sparks
G

Groq

INFRA · APIS
Velocity0.0

Ultra-fast AI inference API for running LLMs using custom hardware.

Groq is climbing off raw inference: hosted tools and MCP connectors now sit above the chips.

inferencemcpbuilt-in-toolsmodel-hostingttsenterprise
◆Current state
GroqCloud's changelog splits three ways: model availability (MiniMax M2.5, Qwen3-VL for enterprise, GPT-OSS variants), speech work consolidating on Orpheus after the PlayAI deprecation, and a newer line of server-side capabilities — browser search as a built-in tool, and MCP connectors for Google Workspace. SDK releases are maintenance-grade following December 2025's 1.0 GA. Cadence has thinned since Q1, with the most recent captured entries dating to April.
◆Where it's heading
The interesting movement is Groq offering things that execute rather than only generate: tools the model can call inside Groq's own runtime, and pre-built OAuth connectors to business applications. That is a bid to be the place agents run, not merely the fastest token endpoint, and it competes with the same agent-runtime pitch Bedrock and OpenAI's Responses API make. Speed remains the underlying claim, but the surface being sold is moving up the stack.
◆Prediction
Expect the connector catalog to widen beyond Google Workspace and more built-in tools to reach the general model lineup rather than only GPT-OSS.

◆Recent moves

  1. 5mo ago

    Browser Search (GPT OSS Models)

    Browser search becomes a built-in tool for GPT-OSS models, letting the model fetch from the web inside Groq's runtime instead of through caller-side plumbing. Captured from the docs page rather than a release note, so detail is thin, but it fits the move toward server-side execution.

    View source ↗
  2. 5mo ago

    Python SDK v1.2.0 and TypeScript SDK v1.1.2

    Quarterly roll-up of Python and TypeScript SDK fixes covering query-param merging, file upload handling, and general stability. Post-GA maintenance with no API surface change.

    View source ↗
  3. 5mo ago

    MiniMax M2.5 and Qwen3-VL 32B Instruct (Enterprise)

    MiniMax M2.5 and a 32B Qwen3-VL vision-language model land on GroqCloud for enterprise customers, gated behind account teams. Routine roster expansion, notable mainly for putting multimodal capacity behind the enterprise tier.

    View source ↗
  4. 5mo ago

    New Voices for Orpheus Arabic Saudi

    Two more voices, including a new default, bring the Arabic Saudi Orpheus model to six. Incremental, but it is the kind of coverage work that follows a platform-wide TTS migration.

    View source ↗
  5. 5mo ago

    Built-In Tools

    A docs index page for compound built-in tools captured as an entry, with no release content attached. Useful only as a marker that the built-in tool surface is being documented.

    View source ↗
  6. 8mo ago

    Platform-wide Migration from PlayAI to Orpheus TTS

    The platform completes its migration from PlayAI to Canopy Labs' Orpheus TTS, retiring playai-tts and playai-tts-arabic per the December 2025 deprecation. Orpheus brings vocal-direction controls, faster inference, and improved audio quality. A clean execution of a planned model swap rather than a directional change.

    View source ↗