Fine-tuning

AutoScientist: Adaption Automates Model Training Research

Adaption, Sara Hooker's lab, launched AutoScientist on May 13, 2026: it automates the model-training research loop, claims 35% average gains and win rates rising from 48% to 64%.

AutoScientist: Adaption Automates Model Training Research — article cover
On this page6 SECTIONS
  1. What AutoScientist Actually Does
  2. The Numbers: Win Rates From 48% to 64%
  3. From Adaptive Data to a Fully Adaptable Stack
  4. Why Standard Benchmarks Can’t Measure This
  5. What It Means for Developers and Researchers
  6. Sources

On May 13, 2026, Adaption — the AI lab founded by former Cohere VP of AI research Sara Hooker — released AutoScientist, a system that automates the full research loop behind model training and alignment. It co-optimizes training data and training recipes, iterating until the model converges on whatever objective the user describes. The company’s pitch: go from an idea to an adapted, fully owned model “in an afternoon.”

What makes it interesting is the positioning. Most fine-tuning tools attack one link in the chain — data cleaning, or hyperparameter search. AutoScientist sells automation of the entire training-research workflow, free for the first 30 days after launch. And Hooker’s framing to TechCrunch was bigger than the product: it “suggests we can finally allow for successful frontier AI trainings outside of these labs.”

What AutoScientist Actually Does

Adaption describes itself as a research-driven “neolab.” Hooker previously led AI research at Cohere, and in an October 2025 TechCrunch interview she was explicit about betting against the scaling race. AutoScientist is that bet turned into a product: not a one-shot fine-tuning utility, but a self-improving system that runs the research loop of training on its own.

The flow is simple to describe: a user states the capability they want, and the system simultaneously shapes the dataset composition and the training recipe until the model converges on that goal. Adaption names two target users — ML engineers who need fine-tuning without manual tuning work, and non-technical builders who until now could only prompt a model. The company’s tagline is blunt: “Intelligence should not arrive preconfigured, and building AI shouldn’t require a PhD.”

The Numbers: Win Rates From 48% to 64%

Adaption published three sets of figures. First, compared with training configurations recommended by human or AI researchers, AutoScientist’s training performs 35% better on average across all runs. Second, win rates climbed from 48% under researcher-recommended configurations to 64%. Third, the company says the gains held across 8 verticals, dataset sizes from 5k to 100k examples, and the model architectures available for fine-tuning through Together AI.

One caveat matters: the evaluation ran on Adaption’s own in-house, domain-specific benchmarks, not public leaderboards. TechCrunch noted the win-rate claims are hard to contextualize without external comparison points.

From Adaptive Data to a Fully Adaptable Stack

AutoScientist did not appear from nothing. Adaption’s earlier product, Adaptive Data, focuses on accumulating high-quality datasets over time; AutoScientist converts continuously improving data into continuously improving models. Hooker’s stated goal is making the whole stack “completely adaptable,” optimized on the fly for whatever task is at hand.

The roadmap points somewhere more aggressive: adaptation techniques that require no training at all. If that lands, the cost and latency of adapting a model drop by another order of magnitude.

Why Standard Benchmarks Can’t Measure This

Because the system produces models custom-built for a specific task, public benchmarks like SWE-Bench or ARC-AGI don’t apply — the same automated pipeline yields a different model for every objective. That is both the selling point and the verification problem. For a user, “better on my task” beats a leaderboard rank. For an outside observer, the only available evidence is the vendor’s own benchmarks and a hands-on trial — which is exactly why Adaption made the tool free for 30 days and let the results argue for themselves.

What It Means for Developers and Researchers

Three practical takeaways. First, the entry barrier for fine-tuning is dropping from “can write a training script” to “can describe a goal,” which meaningfully shortens the path to a proprietary model for non-technical teams. Second, the narrative of frontier training decoupling from the big labs now has its first shippable product behind it, echoing the accelerating model cadence that has defined 2026 so far. Third, the burden of verification shifts to the user: with no public benchmark to lean on, validating on your own eval set before rollout matters more than trusting a vendor’s win-rate chart.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL