Research · March to June 2026

Ship only what beats the baseline.

BrainstormLLM predicts which phases of a software-development pipeline a task actually needs, from specification through documentation, so AI agents can skip the rest. It was trained on 2,203 real Claude Code session turns and had to clear kill gates set before any training began.

2,203Training turns
0.796Mean F1
68%Lower cost
0.31 msInference
400K+RouterBench rows
93%Tier accuracy
/ Accuracy by phase(01)

Third approach, first pass.

A one-shot model reached a mean F1 of 0.52 and a plan cache reached 0.60. Both failed the gate. The version that passed conditions each phase's prediction on the outcome of the phases before it: the same kind of model, with better features.

Sequential predictor · F1 by phasen = 2,203
Mean0.796
Specify0.976
Implement0.916
Review0.723
Verify0.741
Deploy0.630
Document0.792

Kill gate: mean F1 of at least 0.75

/ Kill gates(02)

Set in advance. All passed.

Missing any one of these would have ended the project. The thresholds were written down before the first model was trained.

GateThresholdResultStatus
Enough data200+ examples, 3+ task types2,203 / 9Pass
Phase-skip accuracyMean F1 of at least 0.750.796Pass
Cost reductionAt least 25% against running every phase68%Pass
Inference timeUnder 10 ms with ONNX0.31 msPass
/ What I built(03)

Sequential predictors

Gradient-boosted models, one per pipeline phase, each conditioned on the phases before it.

Model-tier classifier

Trained separately on 400,000+ RouterBench data points; 93% cross-validated accuracy.

Inline inference

Both models export to ONNX, so BrainstormRouter can call them on every request.

Part of one loop.

BrainstormLLM learns from the trajectories the control plane records, and its predictions decide what the CLI runs next.