Ollama
Ollama becomes a gateway provider for Claude Desktop — and this feed missed the release that says so.
A side-by-side editorial comparison of Gemini and HeyGen — release velocity, themes, recent moves, and the top alternatives to consider.
Gemini stops watching video frame by frame and starts deciding what to watch.
Agentic video understanding is now available on Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite, switched on with an API config setting. Instead of ingesting video at a fixed frame rate, the model uses native video tools in an agentic loop to search, scan and inspect segments across frames, audio and transcript, cutting token consumption by up to 88 percent and cost by up to 66 percent while improving accuracy by up to 7 percent. It follows a fortnight of speech work — a dedicated 3.5 Transcribe model, voice-driven task delegation in Gemini Live — and a developer model refresh in Omni 1.1 Flash. Feed bodies are teaser length; this window was read from the source posts.
HeyGen is becoming a video studio that generates scenes rather than assembling templates.
HeyGen has moved past single-avatar clip generation. The Video Agent now builds a whole video from one prompt - script, voiceover, presenter, pacing, captions, music - and as of July it generates the motion graphics too instead of pulling them from a template library. Around that sit avatar-quality tools (Look Packs, Speech Cleanup, Edit Look), a two-host video podcast format, and a monthly enterprise release train covering audit logs, settings, and avatar controls.
Agentic video understanding is now available on Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite, switched on with an API config setting. Instead of ingesting video at a fixed frame rate, the model uses native video tools in an agentic loop to search, scan and inspect segments across frames, audio and transcript, cutting token consumption by up to 88 percent and cost by up to 66 percent while improving accuracy by up to 7 percent. It follows a fortnight of speech work — a dedicated 3.5 Transcribe model, voice-driven task delegation in Gemini Live — and a developer model refresh in Omni 1.1 Flash. Feed bodies are teaser length; this window was read from the source posts.
The same pattern keeps repeating at model level: give the model a native tool and a loop, and let it decide how to use its own context. Computer use in June, robotics in July, agentic vision for images, and now video. Google names agentic vision as the direct precedent for this release, so the technique is established and video is the modality it just reached — which is also why the efficiency numbers, not the capability, carry the announcement. The application layer is running a separate clock, converging on voice as the way input arrives across Live, Workspace and macOS.
Agentic vision covered images and this covers video, so audio-only and document processing are the modalities the same loop has not yet been pointed at. Whether the setting becomes the default rather than an opt-in config value is not stated.
HeyGen has moved past single-avatar clip generation. The Video Agent now builds a whole video from one prompt - script, voiceover, presenter, pacing, captions, music - and as of July it generates the motion graphics too instead of pulling them from a template library. Around that sit avatar-quality tools (Look Packs, Speech Cleanup, Edit Look), a two-host video podcast format, and a monthly enterprise release train covering audit logs, settings, and avatar controls.
Two directions run together. The platform is replacing retrieval with generation - templates give way to generated motion graphics, stitched clips give way to multiple avatars sharing one studio scene with automatic camera coverage. At the same time the company is packaging that engine for buyers who do not want a video tool at all, first as an API surface for Avatar V and now as a vertical product for real estate. The enterprise cadence became monthly, which is the tempo of a vendor selling to procurement rather than to creators.
Expect more vertical packs on the real-estate pattern, and for the generated-scene work to reach the API so other products can build on the same pipeline.
Other ai-assistants products tracked by Sparkpulse, ranked by recent ship velocity. Each card links to a full editorial trajectory and lets you pivot into a head-to-head comparison with either Gemini or HeyGen.
Ollama becomes a gateway provider for Claude Desktop — and this feed missed the release that says so.
Qodo is arguing that AI code review needs governance, not better instruction files
D-ID's feed is a content-marketing engine, not a changelog
Pictory publishes daily search content, not a changelog.
The Agents window now opens without a GitHub sign-in, if you bring an Anthropic key.
The OSS terms change takes effect today, under a steady drumbeat of agent-memory explainers.
See all Gemini alternatives → · See all HeyGen alternatives →
Latest ship moves from both products, interleaved chronologically. ⚡ = editorial spark.
They serve adjacent needs but don't currently overlap on shipped themes. Gemini and HeyGen are shipping at a similar cadence (velocity 6.3 vs 6.3, both within Sparkpulse's "active" band). See the at-a-glance table above for a side-by-side breakdown of velocity, recent sparks, and editorial themes.
Sparkpulse doesn't pick a winner — we score release velocity, not feature parity. Gemini and HeyGen are shipping at a similar cadence (velocity 6.3 vs 6.3, both within Sparkpulse's "active" band). For your specific use case, the alternatives sections above list other ai-assistants products to evaluate alongside.
Top Gemini alternatives in ai-assistants are ranked by recent ship velocity. Browse the "Gemini alternatives" section above for the current picks, or visit /alternatives/gemini for the full list with editorial commentary on each.
Top HeyGen alternatives in ai-assistants are ranked by recent ship velocity. Browse the "HeyGen alternatives" section above for the current picks, or visit /alternatives/heygen for the full list with editorial commentary on each.