← Back to ai-assistants
Weekly · ai-assistants · Week of September 28, 2026

Claude opened a developer plugin portal this week, formalizing its platform ambitions, while OpenRouter launched half-price batch inference and Copilot wired persistent memory into its agent pipeline.

Generated 11h agoDrawn from 6 products

The week in AI assistants

The AI assistant sector this week made two things clear: the products that have been building platform layers are now formalizing them, and the infrastructure underneath them is getting cheaper and more capable simultaneously. Claude opened a public developer plugin portal — a managed directory with review, analytics, and one-click install. OpenRouter launched a Batch API that cuts inference costs in half for async workloads. Both moves in the same week signal that the AI assistant layer is maturing from a feature into an ecosystem.

GitHub Copilot continued its accelerated release cadence — memory-aware security autofix, Claude Opus added to the model roster, local sandboxing, and enterprise opt-out defaults, all landing in a single weekly release window. The velocity is notable: Copilot is running at the pace of a product under competitive pressure.

Leaders

Claude made the most platform-defining move in the sector this week. A public developer plugin portal ships with submission, review, analytics, and one-click installation for end users — the infrastructure required to build and maintain an app ecosystem. In the same window, Claude eliminated the chat/Cowork distinction: every conversation is now a potential agentic session, and Claude Docs, Slides, and Design are available in every context without switching modes. Together these moves position Claude as a platform that developers build on and users stay inside, not just a model they query.

GitHub Copilot had its most feature-dense week of the current cycle. Agentic autofix became stateful — it reads Copilot Memory before generating security fixes, tailoring remediations to repository-specific conventions rather than treating each alert in isolation. Claude Opus joined the model roster, and local sandboxing shipped for the Copilot desktop app to constrain agentic execution. Enterprise accounts moved to opt-out defaults for new GA features. Each of these is directional; all four landing in a weekly release window signals deliberate acceleration.

OpenRouter launched the Batch API, cutting inference costs roughly in half for any workload that can wait up to 24 hours — nightly evals, bulk classification, document processing. Over 230,000 batches processed in beta before the GA launch. For teams running high-volume AI pipelines that don't need real-time responses, this changes the economics meaningfully. OpenRouter continues to position itself as the inference routing layer rather than any single model provider.

Ollama added thinking-level controls to its API: callers can now read per-model thinking defaults and override them per request. Until this, reasoning depth in local models was either on or off by model choice. Now it's a per-request parameter, making it possible to tune reasoning overhead for the specific task without changing the model. For teams running local models in production applications, this is the kind of control that matters at scale.

Baseten added web search as a server-side hosted tool via Grounded Inference. Models invoking it pull current web content from four provider networks without requiring client-side tool call handling or a separate web search API key. This moves Baseten past model hosting into augmented inference — the model, the retrieval, and the tool execution handled together at the serving layer.

Wildcards

KServe shipped Anthropic Messages API and OpenAI Completions API support in the same v0.20.0 release — a multi-protocol serving infrastructure that's now agnostic between the two most dominant AI API conventions. Confidential model serving and KV cache offloading also shipped. For enterprises running Kubernetes-native model serving, KServe is quietly becoming the infrastructure that doesn't require choosing a provider protocol.

Themes that compounded

  • Persistent context is becoming infrastructure: Claude's plugin ecosystem and GitHub Copilot's memory-aware autofix both shipped within the same week — in each case, the AI product remembers what it learned from prior sessions rather than starting cold.
  • Inference cost is dropping as a lever: OpenRouter's half-price Batch API and Baseten's server-side hosted tools both reduce the unit economics of AI workloads without requiring model changes.
  • Protocol convergence at the serving layer: KServe supporting both Anthropic and OpenAI protocols signals that infrastructure is abstracting provider differences rather than picking sides.
  • Every AI assistant is now evaluating the chat-to-agent threshold: Claude eliminated the mode split entirely this week. Others will face the same decision.

Watch this week

Claude's plugin portal is the most consequential structural move in the sector this week — watch which developers ship first and what categories they build in. The first wave of plugins will define the ecosystem's character. On the cost side, OpenRouter's Batch API passed 230,000 batches in beta; watch for teams to start routing high-volume evaluation and classification workloads through it as the economics land, which will show up in usage growth before any product announcement.