Copilot and AWS turn agents on the data and tooling layer while Gemini widens into robotics and the browser
The week in ai-assistants
The center of gravity this week was not a new model but the plumbing agents run on. GitHub Copilot and AWS Machine Learning both posted the sector's top velocity, and both spent it on the unglamorous layer. Copilot moved its code reviewer from a closed feature to one that runs agent skills and MCP servers in general availability across every paid tier, then filled the rest of its changelog with the enterprise switches — user-based model policy, managed settings, device controls — that let admins constrain what it can reach. AWS pushed Bedrock AgentCore in the same direction and went one step further with the Agentic Catalog Experience in Amazon Quick, a curator agent that generates datasets and topics directly against the data catalog. The shared pattern is agents graduating from things you ask questions of to things that act on the tooling and data underneath them.
Gemini was the week's other high-velocity move, but along a different axis: surface rather than depth. Buried in a daily marketing feed were three genuine launches — Gemini Robotics ER 2 for embodied reasoning, Spark's integration into Chrome, and a new model tier — each extending the assistant into a place it did not previously live. Where Copilot and AWS are deepening one control plane, Gemini is spreading across the browser, the robot, and the desktop at once.
Leaders
GitHub Copilot leads on both velocity and spark count. The reviewer going generally available with agent skills and MCP is the entry that widens what Copilot can do rather than who is allowed to do it; the surrounding governance work — managed settings, a cloud agent for Linear — is the reach-then-constrain rhythm that has defined its year.
AWS Machine Learning matched that pace by assembling the parts of agent infrastructure nobody demos: identity, gateways that track the MCP 2026-07-28 spec, and the Agentic Catalog Experience pointing an agent at the catalog itself. It is competing on infrastructure completeness, not model quality.
Gemini shipped Robotics ER 2 as its furthest step from chat surfaces, targeting video understanding, tool orchestration, and multi-robot coordination. Read alongside Spark-in-Chrome, the direction is widening reach across form factors.
Baseten introduced Baseten for Model Labs, reselling its serving stack to the labs that build models rather than only the developers who call them — the clearest directional move in its batch, with a GLM 5.2 Fast tier landing alongside.
Dify shipped an experimental sandboxed shell agent assembled from uploaded Skills, the point where its agent runtime stops being infrastructure under the workflow engine and becomes something a user builds directly. Notably, it ships with a warning to expose it only to trusted users.
Wildcards
Lambda Labs registered zero product velocity but posted two of the week's most consequential entries: a near-total leadership change built around an outside infrastructure operator as CEO, paired with a $1 billion credit facility for gigawatt-scale expansion. The capital and the org chart point the same way — a GPU cloud restaffing to run AI factories at utility scale.
Snorkel AI released Senior SWE-Bench, an open benchmark built from real pull requests across twelve production repositories with half its tasks held private to blunt contamination. Snorkel keeps turning its blog into an evaluation lab, shipping the artifacts frontier labs hill-climb on rather than a conventional product.
Themes that compounded
- Agents move onto the data and tooling layer. AWS Machine Learning points a curator agent at the catalog and GitHub Copilot opens its reviewer to MCP servers and skills — the action is shifting from answering to acting.
- Reach first, then governance. Copilot's managed settings and DataRobot's agent-identity arc both frame constraint as the next layer once agents are everywhere.
- Selling to the other side of the market. Baseten resells its serving stack to model labs while OpenRouter's Classifiers turn a router into a spend-and-behavior control plane.
- Evaluation as the product. Snorkel AI ships Senior SWE-Bench and Promptfoo publishes its red-team skills to a marketplace — measurement is becoming the deliverable.
- Code review as a governance surface. Qodo's Rule Miner and Sourcegraph's Code Finder both reposition review and search as infrastructure sold to agents, not to engineers reading results.
Watch this week
The open question is whether the control-plane build-out translates into adoption or just optionality. GitHub Copilot and AWS Machine Learning are both betting that enterprises will pay for governance depth, but this week's entries are switches and previews, not usage. Watch whether Dify's sandboxed agent moves past its trusted-users-only warning, whether Baseten's Model Labs pitch draws a named lab, and whether Cline and InvokeAI — both mid-transition on desktop and video respectively — convert release candidates into stable ships.