The challenge
Run serially on one box, a 100-trial hyperparameter sweep took ~two days, so the team ran fewer trials and shipped a worse model.
Understanding the workload
One profiled ~300 GB aggregate with high parallelism → burst, fanned out.
Best hardware for the job
Matched 8× NVIDIA A100 80GB. If a deadline demands it, One can match 8× H100 to cut wall-clock to ~51 min, matching to your priority, cost or speed.
Benchmark
| Pick | Hardware | Time | Cost | Verdict |
|---|---|---|---|---|
| One's match | 8× A100 80GB | ~122 min | $59.70 | best perf-per-dollar that fits |
| oversized | 8× H100 80GB | ~51 min | $66.76 | latency option for a hard deadline |
Completion
The whole sweep finished in one wall-clock window in the team's own cloud, then tore down.
What changed
The full sweep lands before standup now. We stopped rationing trials.
Placement, hardware matching and completion are real system behavior; the figures are editable model inputs.