Most scoring models are invented in a meeting and never checked against reality, which is why the sales team stops trusting the tiers within a month.
Every dimension has to name the field it reads and where that field comes from. Applied honestly this removes a large part of any scoring model written in a workshop, because a great many dimensions turn out to be things everyone agrees matter and nobody can actually measure at scale.
Dropping them is the point. A model with four dimensions that are always populated beats a model with twelve that are populated a third of the time, and the second one produces tiers that look precise and mean nothing.
The model is run against a real sample of accounts, a human corrects the tiers that are wrong, and the model is adjusted and re-run. That loop repeats until two consecutive rounds need no corrections. Only then is it locked.
A model everyone agreed to in a meeting and nobody tested is the normal case, and it is why sales teams learn to ignore scores.
Some facts are a no rather than a deduction: wrong geography, a competitor, an existing customer, a size band you cannot serve. Modelling those as heavy negative weights lets a strong score elsewhere drag them back into view. They gate instead.
This defines the model. Applying it to lists is a separate skill, and re-fitting it against what actually replied is a third. Keeping them apart is what stops a model quietly re-tuning itself against the list it is scoring.
Skills compound. These are the ones we usually install alongside it.