In the public sector there is no revenue line to rank your backlog. So the definition of value has to be constructed and written down where everyone can see it, because if it isn’t, it gets constructed invisibly, usually by whoever shouts loudest.
I have worked in product across central government, regulators, local government and health. The failure pattern is the same everywhere. Without an explicit definition of value, everything becomes priority one. The deadline of the week wins. Seniority substitutes for evidence. The team ends up delivering the last conversation the product manager had, not the most valuable thing on the list.
A private-sector product manager can at least argue from projected revenue. We can’t. Our value is a blend: user need, the organisation’s aims, commitments already made to a portfolio, a board or a minister, and the practical question of whether the team can deliver the thing at all. The blend is different in every organisation. That is why no off-the-shelf score has ever worked for me, and why I built my own.
The model
I run my platform team’s roadmap backlog through six questions. Each is scored 1 to 5:
1. Do we think the users need or want this feature?
2. How important is this to the portfolio or organisation’s aims?
3. How quickly is this needed — is there a deadline or a committed delivery?
4. How big is the work? (5 is small — the aim is to deliver smaller items more quickly)
5. Has the team got the skills to do this, or do we need specialists or people from other departments?
6. Is this work difficult to deliver due to risk or things we don’t know yet?
The second half of the model is where value gets defined. Five weights say how much each question matters: user need 18, portfolio aims 22, deadlines 20, capacity to deliver 20, ability to deliver 20. They total 100, and each item’s scores are multiplied by them. The first three questions combine into a Product Priority score: how much this item matters. The last three combine into a Deliverability score: how realistic it is that this team ships it. The average of the two sorts the backlog into Now, Next and Later.
The arithmetic is deliberately simple. It fits on one spreadsheet tab and anyone in the room can check it.
An honest account of the weights
I wanted those weights set in a room with my stakeholders, arguing numbers onto the sheet. That meeting has not happened. On every platform team I have worked with, the work reads as a black art to people outside it, and a detailed spreadsheet session is a hard sell. Mine is no different. So the weights came out of my head, informed by what the portfolio says it cares about.
I am not going to dress that up as collective agreement. But it is still better than the alternative, and here is why: my weights are written down. Anyone can look at the sheet and say “user need at 18 is wrong”. Nobody has yet, though I don’t count that as endorsement. The real test comes the first time the model deprioritises something a stakeholder badly wants, and that test is still ahead of me. What I have until then is transparency rather than agreement, and transparency is the thing the loudest-voice method can never offer. An explicit, challengeable definition of value beats an implicit, unchallengeable one even when a single person wrote it.
You might notice the weights sit close to equal. I take that as roughly honest, since this team’s value genuinely is a blend. But I hold it loosely, precisely because no one has argued with it yet.
What the model told me that I didn’t expect
Two things, both slightly uncomfortable.
Almost every item on my backlog scores 1 on user need. Not because users don’t matter, but because a platform team’s work is mostly invisible to them. [Scott Colfer has estimated] that over two-thirds of digital work in government is internal or staff-facing capability work. My spreadsheet agrees with him. On a backlog like that, the user-need weight lies mostly dormant and barely moves the ordering. It is not dead, though. When a genuinely user-facing item does arrive, it jumps, which is exactly what I want. And the wall of 1s changed how I describe the team’s value to the portfolio: our product is stability, not features.
Second: a long run of my Later items carry identical scores. That is scoring fatigue, and I’m not going to pretend otherwise. Items more than two quarters out cannot be scored meaningfully, and forcing precise numbers onto them is fiction. The identical scores are the model telling me where my knowledge runs out. I treat that as a feature.
The obvious objection
Multiplying subjective guesses by subjective weights does not produce objectivity. It produces a number that looks more certain than it is. This objection is correct, and it misses what the model is for.
The number is not the output. The argument the number forces is the output. When an engineer says “there is no way that item is a 2 on risk”, that disagreement is the model working, and it happened before the quarter was planned rather than three sprints in. When a ranking looks wrong to someone, we check it against the weights together. Either the weights need changing or the instinct doesn’t survive contact with a written-down definition of value. Both results are useful. And if stakeholders keep overturning the output entirely, the problem is an unagreed strategy, and no spreadsheet fixes that.
The same reasoning explains why I didn’t use an off-the-shelf framework like RICE (reach, impact, confidence, effort) or SAFe’s WSJF (weighted shortest job first). They are fine frameworks, but each encodes someone else’s definition of value. RICE favours reach, which quietly undervalues statutory work affecting small groups of users. WSJF assumes you can estimate cost of delay honestly, which in my experience most teams cannot. One caution from my own sheet: scoring small items higher means big strategic work needs the deadline and portfolio scores to carry it, so watch that quick wins don’t permanently crowd out the big migration. Build your own weights and you at least know whose biases are in them. Steal the principle, not my spreadsheet.
Try it on your backlog
If you work in product in government, a council, a charity or anywhere else without a revenue line, try this. Write down the four or five things that constitute value where you are. Your list will not be mine, and it shouldn’t be. Put numbers on them, in the open, with your stakeholders if you can get them in the room and without them if you can’t. Then score your backlog and compare the ranking to your current roadmap. Where they differ, you have found either a flaw in the model or an assumption in your roadmap that was never examined. Both are worth an hour of your time.
One day I’d like the teams across a portfolio to run the same kind of model, so resourcing decisions between teams can be argued from something written down. I don’t think most portfolios, mine included, are mature enough for that yet. Team by team is how it starts. I’d like to hear where it breaks for you.
Personal views, not those of any client or department.
Mark Challinor is a product manager who has worked across UK central government, regulators, local government and health.
Questions this post answers
How do you prioritise a product backlog in the public sector?
Score every backlog item 1–5 against an explicit, written-down definition of value — in this model, six questions covering user need, organisational aims, urgency, size, skills and risk, combined through weights into a Product Priority score and a Deliverability score. The average of the two sorts the backlog into Now, Next and Later, and the arguments the scores provoke are where the real prioritisation happens.
Why not use an off-the-shelf prioritisation framework like RICE or WSJF?
Every off-the-shelf framework encodes someone else’s definition of value: RICE favours reach, which undervalues statutory work affecting small groups of users, and WSJF assumes teams can estimate cost of delay honestly, which most cannot. Building your own weights means you at least know whose biases are in the model.
Does weighted scoring make backlog prioritisation objective?
No — multiplying subjective guesses by subjective weights produces a number that looks more certain than it is. The value of the model is not the number but the argument it forces: disagreements about scores surface before the quarter is planned, against a written-down definition of value that anyone can challenge.




Great write-up, Mark. Seeing this used in practice has reinforced for me that the biggest benefit isn't the score itself, it's having an explicit definition of value that people can challenge and improve. It's certainly helped bring clarity to our roadmap, and I think there's a lot other teams could learn from in the way it makes priorities, trade-offs and assumptions visible.