Autonomous ML · 2026
The winning model changed when the data grew.
Two agent-run experiment phases tested meal-prediction models on 68 and 291 training sequences. The architecture ranking reversed when the history became 4.3× larger.
The question
A single model leaderboard hid the more useful question: which architecture wins under the data conditions a real product will have?
The outcome
A small MLP won the sparse regime. A Transformer won after the available history grew, pointing to a model strategy that changes with the evidence.
Why it matters
Data maturity is a product constraint, not an offline modeling footnote. A system that expects one architecture to stay best from onboarding through a year of history will ignore the change this experiment made visible.
- Role
- Experiment designer and investigator
- Collaborators
- Autonomous research loop built from Karpathy's autoresearch pattern
- Duration
- Two experiment phases
- Period
- 2026
More history. A different winner.
Validation MSE · lower is better
90 days 68 sequences
365 days 291 sequences
4.3× more training data — the MLP leads with sparse history; the Transformer leads with more.
The useful result was not a champion model. It was evidence that architecture choice should follow the operating regime instead of reputation or a paper trained at a different scale.
Seven improved the model on 68 training sequences.
Repeated after the history expanded to 291 sequences.
Enough to reverse the architecture ranking.
Evidence atlas
The conclusion changes with the data regime.
Explore the figures and the findings behind them.
Experiment evidence 01
The ranking flip
The MLP won with 68 sequences. With 291, the Transformer moved ahead while the MLP degraded.
FindingValidation MSE; lower is better. Dataset scale changed from 68 to 291 training sequences.
More history. A different winner.
Validation MSE · lower is better
90 days 68 sequences
365 days 291 sequences
4.3× more training data — the MLP leads with sparse history; the Transformer leads with more.
Experiment evidence 02
Small changes did most of the work
Loss, learning-rate, and noise changes improved the sparse-data baseline before architectural complexity helped.
FindingPhase-one experiment summary on the synthetic 90-day dataset.
Three small changes that helped.
Validation MSE · before each change = 100
Huber loss10% lower MSE
Learning-rate tuning42% lower MSE
Input noise15% lower MSE
Each pair uses its own starting point. Indexed from the original figure’s rounded reductions; these are separate comparisons, not a cumulative sequence.
Experiment evidence 03
Most ideas failed
The autonomous loop was valuable because it could discard poor runs cheaply and keep a comparable record of what had been tried.
FindingSelected failures from the 365-day experiment phase.
What the discarded runs taught us.
Error relative to the best run · lower is better
Blue line = best run (1×). Hatched ends mark bars capped at 30×; the labels show the reported result.
Consequential decisions
Where the work changed direction.
These choices shaped the architecture, the evidence, or the way the system could be used.
Hold the loop constant
Both phases used the same task and a five-minute budget per run. That made the change in data scale, rather than a shifting evaluation process, the consequential variable.
Keep the failures
Most runs made the model worse. Preserving those attempts showed which architecture, loss, regularization, and augmentation ideas failed under each regime instead of presenting a polished winner with no search history.
Let the product graduate models
The result suggests a practical sequence: begin with a simpler model while personal history is sparse, then reevaluate architecture after enough longitudinal data accumulates.