What Makes a Good aeroTEST: Equipment, Road, Setups (2026)

What Makes a Good aeroTEST: Equipment, Road, Setups (2026)

Imagine two riders, same road, same helmet swap. One rides home confident the new helmet is worth a few watts. The other can't tell, and buys it anyway. The difference was not the helmet and not the road. It was how each of them tested. An aeroTEST is a measurement, and like every measurement it has a noise floor you can lower or raise with your own hands. This guide covers the factors you control, from the power meter to the order of your setups, including one relationship we quantified across 7,177 outdoor setups on our platform.

New to the protocol itself? Start with the step-by-step aeroTEST guide. This piece is about doing it well.

What "good" means

A good aeroTEST is one where the runs of a setup land on top of each other, and two setups that differ by a helmet's worth of drag show up as different. In numbers: the run-to-run scatter of CdA inside a setup, as a share of CdA. Every run on the platform carries its own uncertainty, reported as an error bar: the platform fits the run, compares it with the other runs of the setup, and reports the larger of the fit error and the run-to-run disagreement, as a share of CdA. Treat it as the platform's uncertainty indicator, not as a formal confidence interval. And it does not cover everything: a power meter that reads two percent high all day, or a position that changed between setups, shifts every run the same way and stays invisible to it. The point of this article is to keep both small: the bar, and the errors it cannot see.

Three things set the noise: the road, your equipment, and how you ride. The road sets the ceiling (we measured 977 of them). The other two decide where under that ceiling you land.

Three tiers of equipment

The same method returns very different numbers depending on what you bring. From our own test days with teams and from the platform's data, three tiers. Treat the table as a preparation checklist, not as a guarantee of any particular resolution.

LowMediumHigh
SpeedGPS onlyGPSGPS plus speed sensor, circumference measured
Power metersingle-sidedtotal power from both legstotal power from both legs, spider or crank based
Head unitanyanyGarmin Edge with the aeroAPP
Routewhatever is nearbyflat, under 5 m height difference, 500 mflat, under 5 m, 1,000 m, wind checked
Windignoredignoredfrom the side, road sheltered, two routes to choose from
Wheelsdeep rimsanyno deep rims for position tests
Clothingwind jacketjerseyrace kit, nothing loose
Positionrememberedrememberedfixed reference points, photographed
Tyre pressurewhatever it ischecked oncechecked before every setup

You do not need to be at High to learn something. A Medium session can resolve larger differences; whether it separates two helmets depends on how different they are and on the uncertainty of that session. A Low session tends to tell you that the aero helmet is slower on Tuesday and faster on Thursday, and both are the road talking.

Hold power like a metronome

This is the relationship we quantified. We took every outdoor aeroTEST setup on the platform with at least two clean runs and sorted the setups by how much the rider's power wandered within a run: the standard deviation of power during a run as a share of its mean, averaged over the setup's runs.

Power variation within a run (SD as share of mean)SetupsRun-to-run CdA scatter, medianError bar the platform reported, median
under 4 %1900.52 %2.3 %
4 to under 9 %4,2960.71 %2.9 %
9 to under 13 %1,8350.96 %3.0 %
13 % and over8561.52 %3.8 %

Run-to-run CdA scatter by how steadily power was held, four groups from under 4 % to 13 % and over: the median scatter rises from 0.52 % to 1.52 % across 7,177 outdoor setups

Setups with a power variation below 4 % showed a median run-to-run scatter of half a percent of CdA. Setups with a power variation of 13 % or more showed three times that, and the error bar the platform reported to those riders widened from 2.3 to 3.8 %.

Part of the gap is the number of runs: a setup that contains more runs shows more scatter by construction, and the unsteady group has slightly more of them. So we repeated the comparison within setups of the same length:

Power variationSetups with exactly 2 runsScatter, medianSetups with exactly 3 runsScatter, median
under 4 %1220.31 %511.51 %
4 to under 9 %2,4130.53 %1,4311.18 %
9 to under 13 %1,0060.63 %6211.61 %
13 % and over4130.74 %3242.14 %

At two runs, the least steady group scatters 2.4 times as much as the steadiest. At three runs, the steadiest group is only 51 setups and lands above the next one, so the sequence is not monotonic there; from the second group to the least steady the scatter still rises by 1.8 times. Steadier riders also rode faster, 44 against 39 km/h on average, and speed sharpens resolution on its own (why). So read this as an association across thousands of real sessions, not as a controlled experiment: steadiness, speed and experience travel together, and we have not separated them.

Why would steadiness go with scatter? Our best explanation, not something this analysis measured: a surge moves power into changing speed rather than overcoming drag, which makes the fit more sensitive to noise in the speed data, and a rider who surges is more likely to be a rider who moves, head up, shoulders open, a different frontal area for a few seconds. The platform shows you the steadiness after every run in the ride-quality panel, as power smoothness. Look at it before you look at the CdA.

What it means for your test power: choose the highest power you can hold steadily for every run in your planned session, not the highest you can hold. When the standard deviation of your power starts climbing in the second half of a session, your position is going with it. Stop, or accept wider bars.

Methodology: how we built this table

Every figure comes from stored test data on afasteryou.com. From 13,000+ outdoor sessions since 2018 we deliberately selected the valid out-and-back aeroTESTs and, inside them, every setup after the calibration run with at least two clean runs, carrying neither an error nor an outlier flag: 2,176 aeroTESTs on 812 roads, 7,190 setups, 7,177 of them with power data. Power variation is the standard deviation of power within a run divided by the run's mean, averaged over the setup's runs; the groups are under 4 %, 4 to under 9 %, 9 to under 13 %, and 13 % and over. Scatter is the standard deviation of CdA across the setup's runs divided by the setup's mean; we report the median across setups in each group, and repeat it for setups of exactly two and exactly three runs. Pooled residuals, the estimator used in the road article, give 2.7 %, 3.0 %, 3.2 % and 5.1 % for the four groups. Removing flagged runs lowers scatter mechanically, so the numbers describe what clean, structured testing on the platform produced, not every session ever ridden. As of September 2026.

The power meter

  • Total power beats one leg doubled. A single-sided meter measures one leg and doubles it, assuming a constant balance. If your balance drifts over a session, the doubled reading drifts with it, and the platform reads that drift as drag.
  • Calibrate once, at the start, warm. Then leave it. A zero-offset halfway through moves every run after it relative to every run before it.
  • Spider and crank systems that measure both legs (SRM, Quarq and Power2max are the ones our test recommendations have named since 2022) are our recommendation. Dual-sided pedals work too. What matters is total power from both legs, measured the same way all day.

Speed

Speed is half of the equation, and GPS speed is the noisier half. A speed sensor with a measured circumference substantially reduces that noise. Measure the circumference with a tape and enter it by hand; do not let the head unit guess, because its guess can change between runs, and keep it unchanged for the session. Check the tyre pressure before every setup, not once in the morning: pressure changes circumference, circumference changes speed, speed changes CdA.

The road, briefly

We wrote a whole article on this; the short version:

  • A straight kilometre, flat, less than 5 m height difference. Bends are fine if they stay under 30 degrees.
  • Wind from the side. Riding both directions gives the model opposing observations, which lets it separate your drag from a steady wind along the road. Gusts and a wind angle that changes between the legs still cost repeatability, so pick the road that puts the prevailing wind across your path.
  • Two routes, one roughly north-south, one east-west, and the forecast decides which you ride today.
  • Shelter: tree lines, embankments. But not a gap between two woods: a gap funnels.
  • No deep rims for position tests. A deep wheel sails in a crosswind, and it sails differently in every gust. Box-section wheels for anything that is not a wheel test.

Building the setups

The first setup earns your trust. On a new road, or the first day with a new bike, give the baseline four or five runs. If they agree within their reported uncertainty, you have a stability check for that moment: the road, the sensors and your position are behaving. It is not proof that every later difference is real; that still depends on each comparison and its uncertainty.

One change per setup. Helmet and skinsuit at once tells you the sum, and the sum is useless when the helmet was worse and the suit was better.

Wheels have an order. Test rear wheels first, with the speed sensor on the front wheel. Then front wheels, with the sensor moved to the rear. The wheel you are not testing should be a plain box-section rim, so it adds no sailing of its own.

Come back to the baseline. Ride it again at the end. If the closing baseline agrees with the opening one, your confidence in the session goes up, though matching endpoints cannot prove that every setup in between saw the same conditions. If it does not agree, conditions or execution changed somewhere in between. You cannot tell exactly when or how much, but you know how far to trust the comparisons in the middle. On long sessions, ride the baseline a third time halfway.

Nothing loose. No jacket, no flapping number, no unzipped jersey. Loose fabric is drag that changes every second; the method cannot average it.

What to test

The list riders bring to us, in the order that usually pays:

  • Your body first. Roughly 80 % of drag is you, not the bike (why, and what a good CdA is). Torso angle, head and shoulder position, arm width. A bike fitter decides whether a position is sustainable; an aeroTEST decides whether it is faster. Use both.
  • Helmet, and with it the visor-versus-glasses question, and the head angle you actually ride with. A helmet tested with the head up and raced with the head down was never tested.
  • Skinsuit: one-piece versus two-piece, short versus long sleeve, tri-suit versus aero suit. Then shoe covers.
  • Bottles: on the down tube, behind the saddle, on the aerobars, or not at all for a short race.
  • Stem, extensions, saddle height: small position changes you can make with an Allen key at the roadside.
  • Frame and wheels after position and clothing, unless the frame or the wheels are the decision you specifically need to make. Their gains depend on the rest of the setup, on speed and on wind angle, and wheels need the order above.

What a good setup looks like

Six runs of one setup, CdA in aero points with error bars: 18.9, 19.0, 18.8, 18.8, 19.1 and 18.7 aP, every run inside the others' error bars, mean 18.9

Six runs, one setup, from an anonymised outdoor aeroTEST on the platform, 2022. CdA is in aero points: 1 aP is 0.01 m². Every run sits inside the others' error bars, the whole setup within 0.4 aP, which is 0.004 m². That is a rider who held position and power. The average, 18.9 aP, is a precise baseline for every comparison in that session, as long as the other setups are ridden under the same conditions. Precise is not the same as absolutely right: a power meter reading high all day would shift all six runs together.

Effective wind per run for the same six runs, 0.1 to 0.4 metres per second, no spikes

The wind chart underneath is the reason. The platform reconstructs the effective wind from the two legs of each run; here it never leaves the tenths of a metre per second. A run that spikes while its neighbours sit flat points to a gust or to a problem with the run itself, and its CdA will often stand outside the others. Check it together with its error bar and with the outlier flag the platform sets. Do not delete it; ride another one.

A bad setup looks like the opposite: runs that stagger up and down by a full aero point, error bars that do not overlap, a wind chart with teeth. The fix is almost never in the software. It is a jacket, a gust, a tyre, or a rider who sat up for the last hundred metres.

Ready to test? Create your free account, plan one baseline and one change, and ride your first outdoor aeroTEST.

FAQ

Is a single-sided power meter good enough? It will produce a number. The number is one leg doubled, and your leg balance moves over a session. If it is what you have, keep the sessions short, and expect wider error bars.

Should I test at higher power to get a stronger signal? Only as high as you can hold steadily for the whole session. In our data, the pooled medians differed about threefold between the steadiest and the least steady setups, and the comparisons at fixed run count were smaller and not uniformly monotonic; an association, not a guaranteed effect. Still, a few extra watts do not buy back what unsteadiness costs. Choose race power and hold it like a metronome.

How many setups fit in one session? Six or seven is a full afternoon: calibration, an opening baseline, three or four changes, and a closing baseline. More than that and fatigue starts to show up as scatter in the later setups.

Can I test position with deep wheels? You can, but every gust will land in your position result as sailing. Put box-section wheels on for position and clothing tests; test the deep wheels as a wheel test, in their own setup.

Wind jacket for the cold? Between setups, yes. During a run, never. A flapping jacket is not a stable CdA, and no number from that run means anything.

Does the same setup give the same CdA on another day? Not necessarily the same absolute number: wind angle, temperature, how the suit sits, how well you reproduced the position, and plain measurement uncertainty all move it. Comparisons within one session are usually the more defensible ones, provided the baseline held and the difference is larger than the uncertainty. Between days, compare through the baseline you repeated in each session.


Sources: power-steadiness analysis from stored outdoor aeroTEST data on afasteryou.com, 2018 to September 2026 (7,177 setups; see the methodology box). Six-run example: anonymised outdoor aeroTEST on the platform, 2022. Protocol recommendations (runs per setup, wheel order, calibration, circumference): the aeroTEST test recommendations and the method as implemented on afasteryou.com, September 2026. Road-level scatter: Outdoor aero testing: why your road matters, 8,566 setups on 977 roads, August 2026. Resolution by speed and the meaning of the error bar: Velodrome aero testing. Share of drag from the rider's body and typical CdA by position: Cycling aerodynamics: how to test and improve your CdA. Physics: Martin, J.C. et al. (1998), Validation of a mathematical model for road cycling power, J Appl Biomech 14(3):276–291.

Pronto a diventare un Più veloce?

Inizia oggi la tua prova gratuita di 30 giorni. Scopri il piano di allenamento più intelligente al mondo.

30 giorni gratis
Cancella in qualsiasi momento
Inizia ora