← Back to all sparks
P

Promptfoo

AI-ASSISTANTS
Velocity5.0

Open-source LLM evaluation, testing and red-teaming toolkit

Promptfoo tracks every new frontier model while its scanning action sheds an old runtime.

llm-evaluationprovider-coveragered-teamingsupply-chain-securitygithub-actions
Current state
The main line's rhythm is provider coverage — Claude Opus 5, Kimi K3, GPT-5.6, Grok 4.5 and current Azure and Gemini models all landed within a few releases, alongside websocket URL templating and a steady stream of assertion and red-team fixes. The newest entry is not the evaluator itself but its code-scan GitHub Action, which drops Node.js 20 as a breaking change and hardens its supply chain.
Where it's heading
Promptfoo's value is being the harness that already knows about whatever model shipped last week, and the release notes read accordingly: features are almost entirely new providers, while the bug-fix column carries the real engineering. The separate versioning of code-scan-action shows the security scanning side maturing into its own product with its own release cadence and its own breaking changes.
Prediction
Provider additions will keep arriving within days of each frontier launch, since that cadence is the product. The action dropping Node 20 while the previous action release moved to Node 24 suggests the supply-chain hardening work there is not finished.

Recent moves

  1. 17d ago

    Code-scan action drops Node 20 and hardens its supply chain

    The code-scan action drops Node.js 20 as a declared breaking change and patches undici alongside supply-chain hardening. Nothing new for evaluation users, but anyone pinning the action on an older runner has to move.

    View source ↗
  2. 1mo ago

    Promptfoo adds Claude Opus 5, Kimi K3 and websocket URL templating

    Claude Opus 5, Kimi K3 and refreshed Azure and Gemini catalogues arrive together, with websocket URL templating the one non-provider feature. The usual shape of a promptfoo release.

    View source ↗
  3. 2mo ago

    Promptfoo adds GPT-5.6, Grok 4.5 and an Open Interpreter provider

    GPT-5.6 on Bedrock and GA, Grok 4.5, an Open Interpreter provider and a new red-team goblin strategy. Token-usage accounting gets proper helpers for the first time.

    View source ↗
  4. 2mo ago

    Promptfoo adds per-test repeat and broad Bedrock model coverage

    Per-test repeat and broad Bedrock coverage — evaluation ergonomics alongside the standing provider expansion.

    View source ↗
  5. 2mo ago

    Code-scan action fixes mixed skip responses and moves to Node 24

    An earlier code-scan action patch fixing mixed skip responses and moving to Node 24. Maintenance on the action, not the evaluator.

    View source ↗
  6. 2mo ago

    Promptfoo unpins Docker Python and fixes over-redaction

    Docker Python unpinned and an over-redaction fix. Housekeeping.

    View source ↗