All work

Autonomous ML · 2026

The winning model changed when the data grew.

Two agent-run experiment phases tested meal-prediction models on 68 and 291 training sequences. The architecture ranking reversed when the history became 4.3× larger.

Machine learningGLP-1Experiment design

The question

A single model leaderboard hid the more useful question: which architecture wins under the data conditions a real product will have?

The outcome

A small MLP won the sparse regime. A Transformer won after the available history grew, pointing to a model strategy that changes with the evidence.

Why it matters

Data maturity is a product constraint, not an offline modeling footnote. A system that expects one architecture to stay best from onboarding through a year of history will ignore the change this experiment made visible.

Role
Experiment designer and investigator
Collaborators
Autonomous research loop built from Karpathy's autoresearch pattern
Duration
Two experiment phases
Period
2026
Experiment notes / 01

More history. A different winner.

Validation MSE · lower is better

90 days 68 sequences

MLPBest0.0099
LSTM0.0379
Transformer0.0439

365 days 291 sequences

MLP0.0189
LSTM0.0156
TransformerBest0.0123

4.3× more training data — the MLP leads with sparse history; the Transformer leads with more.

Primary artifact

The useful result was not a champion model. It was evidence that architecture choice should follow the operating regime instead of reputation or a paper trained at a different scale.

62phase-one runs

Seven improved the model on 68 training sequences.

80phase-two runs

Repeated after the history expanded to 291 sequences.

4.3×more data

Enough to reverse the architecture ranking.

Evidence atlas

The conclusion changes with the data regime.

Explore the figures and the findings behind them.

Experiment evidence 01

The ranking flip

The MLP won with 68 sequences. With 291, the Transformer moved ahead while the MLP degraded.

FindingValidation MSE; lower is better. Dataset scale changed from 68 to 291 training sequences.

Experiment notes / 01

More history. A different winner.

Validation MSE · lower is better

90 days 68 sequences

MLPBest0.0099
LSTM0.0379
Transformer0.0439

365 days 291 sequences

MLP0.0189
LSTM0.0156
TransformerBest0.0123

4.3× more training data — the MLP leads with sparse history; the Transformer leads with more.

Validation MSE; lower is better. Dataset scale changed from 68 to 291 training sequences.

Experiment evidence 02

Small changes did most of the work

Loss, learning-rate, and noise changes improved the sparse-data baseline before architectural complexity helped.

FindingPhase-one experiment summary on the synthetic 90-day dataset.

Experiment notes / 02

Three small changes that helped.

Validation MSE · before each change = 100

Huber loss10% lower MSE

Before100
After90

Learning-rate tuning42% lower MSE

Before100
After58

Input noise15% lower MSE

Before100
After85

Each pair uses its own starting point. Indexed from the original figure’s rounded reductions; these are separate comparisons, not a cumulative sequence.

Phase-one experiment summary on the synthetic 90-day dataset.

Experiment evidence 03

Most ideas failed

The autonomous loop was valuable because it could discard poor runs cheaply and keep a comparable record of what had been tried.

FindingSelected failures from the 365-day experiment phase.

Experiment notes / 03

What the discarded runs taught us.

Error relative to the best run · lower is better

No early stopping245× worse
Z-score normalization65× worse
Mixup augmentation67× worse
Transformer · 8 heads+70% worse
Lookahead optimizer+67% worse
Pre-norm Transformer+49% worse
GRU+47% worse
Bidirectional LSTM+40% worse
Snapshot ensemble+28% worse
3-model ensemble+13% worse

Blue line = best run (1×). Hatched ends mark bars capped at 30×; the labels show the reported result.

Selected failures from the 365-day experiment phase.

Consequential decisions

Where the work changed direction.

These choices shaped the architecture, the evidence, or the way the system could be used.

01

Hold the loop constant

Both phases used the same task and a five-minute budget per run. That made the change in data scale, rather than a shifting evaluation process, the consequential variable.

02

Keep the failures

Most runs made the model worse. Preserving those attempts showed which architecture, loss, regularization, and augmentation ideas failed under each regime instead of presenting a polished winner with no search history.

03

Let the product graduate models

The result suggests a practical sequence: begin with a simpler model while personal history is sparse, then reevaluate architecture after enough longitudinal data accumulates.

Next projectClinical-grade signal from a drop of blood.