A weighted scoring model ranks a set of options against criteria you fix before you look at the results, with each criterion carrying a declared share of the total. I have used them on roadmaps for years, and I have watched most of them fail the same way: the ranking comes out, someone senior does not like row four, row four moves, and nobody changes the weights. What is left is a spreadsheet that laundered a decision somebody had already made.
In August I built one for a project with no stakeholders at all, which turned out to be the useful part. I was drawing an icon set, Vigil Icons, and I had a list of candidate phrases far longer than the set I wanted to ship. So I scored the whole list before drawing a single icon, and then I did the thing I keep failing to do at work: I let it win.
What is a weighted scoring model?
A weighted scoring model ranks a set of options against a fixed list of criteria, where each criterion carries a percentage of the total score. The weights are set before the options are scored, so the ranking comes out of the criteria rather than out of whoever is arguing hardest in the room.
The mechanism is trivial. The discipline is not. Almost every failure I have seen comes from setting the weights after seeing a draft ranking, which is not a model, it is a justification. The sequence matters more than the arithmetic: decide what the thing is for, translate that into weights, then score. If you cannot write the weights down before you look at the list, you do not yet know what you are building, and the rubric will not tell you.
What criteria did I score 111 icon concepts against?
Five criteria, weighted to reflect what an icon in this set is for: getting sent to somebody. Sendability took 35%, visual metaphor 25%, durability 20%, and distinctiveness and survives-as-an-icon 10% each, with extra credit where the object itself performs the idea rather than illustrating it.
| Criterion | Weight | What it actually asks |
|---|---|---|
| Sendability | 35% | Would a person drop this into a group chat instead of typing the phrase? |
| Visual metaphor | 25% | Is there a concrete object that carries the idea without a caption? |
| Durability | 20% | Will this phrase still be in use in three years, or is it a season? |
| Distinctiveness | 10% | Does it look like anything else in the set? |
| Survives as an icon | 10% | Does it hold together at 24px, in one stroke weight, with no fill? |
That extra-credit rule did more work than the five criteria combined, and it is the reason the weights are lopsided. An icon set built for sending is not the same product as an icon set built for a design system, and pretending otherwise would have produced a tasteful, evenly-weighted, completely pointless ranking.
Notice what is missing. There is no criterion for how much I liked the phrase, how funny I found it, or how well it fit an aesthetic I already had in mind. Leaving that out is the whole point. A rubric that includes a row for personal preference has given itself permission to reach any conclusion.
What do you do when a scoring rubric disagrees with your judgement?
Either accept the result or change the weights and re-score everything, in public. What you cannot do is quietly override a single row, because a rubric you overrule whenever you dislike the answer is not a ranking, it is a record of preferences you already held.
This came up immediately. Weighting sendability at 35% seated four entries I would not have chosen by taste, two of them the kind of thing you would not put on a slide at work. They scored where they scored because people genuinely send them, the metaphor was available, and the criteria do not care what I think. The full ranking is public on the site, so anyone can check the arithmetic against my discomfort.
I kept them. Not because I am relaxed about it, but because the alternative was worse: a rubric with a silent veto is a rubric that produces exactly the list you would have written without it, at greater cost and with a spreadsheet attached for cover. If I had genuinely believed those entries did not belong, the honest move was available and I did not take it, which tells me something. Lower the sendability weight, raise durability, re-score all 111, and publish the new order. That is a real decision with a visible cost. Moving one row and leaving the weights alone is not.
The same test works on a roadmap. When a scoring model puts an unglamorous integration above the feature everyone wants to demo, you have two legitimate options and one illegitimate one, and most teams take the illegitimate one.
Should a scoring rubric be allowed to cut entire categories?
Yes, and deciding what the thing is not is usually the highest-leverage cut available. Removing a whole category early is cheaper than scoring every member of it and arguing about each one, and it sharpens the definition of what remains.
One whole category came out before scoring started: generational labels. Gen X and Gen Alpha are taxonomy, not jokes. A generational label sorts people; it does not accuse anyone of anything, and this set is built on accusation. Scoring them individually would have been busywork, because the problem was never any single entry — it was that the category answered a different question than the set was asking.
The second filter was harder to name and cut more: is there an object? If the best available drawing for a concept is a glowing rectangle, there is no icon there, only a label with a screen next to it. A couch wedged on a staircase is a better icon than a screen, because the couch is doing something and the screen is merely present. That test sits underneath the visual metaphor row, and it removed candidates the arithmetic would happily have seated.
Why set a stop condition before you start building?
A stop condition written in advance is a test the work either passes or fails. Written afterwards, it becomes a description of what you already produced, which is why so much finished work quietly redefines success to match its own result.
Mine was one sentence, fixed before the first icon: if the joke dies at 16px, redraw it. Everything else followed from it. Draw at 48 to 64px, but the icon has to read at 24. One 24-unit grid with a 2px safe area, 1.25px stroke, round caps and joins, no fills, and no words inside the frame — because a word inside an icon is an admission that the drawing failed.
The rule that came out of it: a recognizable noun, doing exactly one abnormal thing.
Gaslighting, rent-free, could've been an email. In each one the object is completely ordinary and exactly one thing about it is wrong.
The lantern is a normal lantern; the flame burns downward. The head is a normal head; there is a house living in it. The envelope is a normal envelope; there is a meeting inside. None of them need a caption, and none of them survive a second abnormal thing being added — which is what the 16px rule keeps catching, because detail is the first thing to die and the joke usually dies with it.
Twelve icons in the finished set clear that bar with nothing else attached. One is still flagged for another sketch round, which the rule also decided.
Where this transfers
None of this is unique to drawing. The pattern I keep returning to on product work is the same three moves in the same order: decide what the thing is for, write the criteria and the stop condition down before you can see the results, and then be willing to lose an argument to your own rubric.
I have written versions of this before from the other direction — starting with the decision rather than the model on an AI roadmap, and building a diligence checklist around questions rather than documents. Both are attempts at the same discipline: commit to the test before you can see whether you pass it.
The icon set was a good place to practice, precisely because nothing was at stake. There was no stakeholder to overrule me and no budget to defend, which meant the only person who could corrupt the rubric was me. I did not, quite, and the set is better for it in ways I can point at: it is more evenly useful and less a catalogue of things I personally find funny.
All 111 are free, in every format, at vigilicons.com. The ranking is on the page, weights included, so the arithmetic is checkable. If you disagree with the weights, that is the correct argument to be having.