Snorkel AI
AI data development platform for enterprise model fine-tuning, evaluation, and curation.
Snorkel is building the scoreboard for agents that have to keep working, not just answer.
◆Recent moves
- 15h ago
Continual Learning Bench: measuring whether AI systems actually improve with experience
A named benchmark for whether systems improve with experience, extending the long-horizon evaluation thread that milestone-based scoring and the earlier continual-learning post opened. Note the feed excerpt here is the previous post's Train-to-Test summary with only the trailing link swapped, so the benchmark's actual scope cannot be read from the feed.
View source ↗ - 3d ago
Train-to-Test (T²) Scaling Laws: Why Reasoning Models Should Be Overtrained
A Reading Group presentation of Train-to-Test scaling laws, arguing reasoning models should be overtrained past Chinchilla-optimal ratios. Outside research relayed through Snorkel's series, with no artifact of its own.
View source ↗ - 16d ago
Milestone-Based Evaluation and Training for Long-Horizon AI Agents
An argument for scoring long-horizon agents at milestones rather than on final answers, since earlier decisions constrain later ones. It sets up the evaluation stance the benchmarks are built to serve.
View source ↗ - 17d ago
Enterprise environments and training AI agents for real-world workflows
A case that enterprise workflows need full environments rather than the single-task episodes most agent benchmarks score. Positioning for the expert-environment work rather than a released artifact.
View source ↗ - 24d ago
Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks
An error analysis of Claude Opus 5 on Senior SWE-Bench, where it places second overall and leads the bug-and-performance-investigation category. A results write-up on an existing benchmark, the same format as the earlier Grok run.
View source ↗ - 1mo ago
Senior SWE-Bench: Evaluating Coding Agents Like Senior Engineers
Senior SWE-Bench is Snorkel's own open-source, Harbor-compatible benchmark, with 100 senior-level tasks split between public and held-out halves. The held-out half is the part that makes it usable as a scoreboard rather than a training target.
View source ↗