Campaigns usually launch with no question attached, so they run while somebody watches, and six weeks later nobody can say what was learned. A registered question turns watching into an answer.
Testing happens in a hierarchy whether you acknowledge it or not. A subject line test on a badly targeted list measures nothing about subject lines. An offer test with the wrong buyer measures nothing about the offer. Each level only becomes measurable once the level above it is settled.
So the queue runs list, offer, subject, opener, call to action, timing, length. Most teams start around position three, which is why most outbound testing produces no durable conclusions.
A question needs enough sends per arm to separate a real difference from noise. Multiply by the number of arms and you have the cost of one question in volume. Divide your quarterly volume by that and you get the number of questions you can actually answer, which is frequently two or three rather than the dozen on the list.
Deciding which two matter is the real work here, and it is a strategy decision rather than a statistical one.
The question, the arms, who it runs on, how long it runs, the sample it needs, and the specific decision each possible outcome triggers. A question without the last part is an observation, and observations are what teams already have too many of.
Skills compound. These are the ones we usually install alongside it.