eSourcingData - Source-to-Contract Procurement Software
Evaluation Management

Consistent, defensible evaluation. Every time.

eSourcing Data enforces independent scoring, applies weighting automatically and generates the evaluation report in seconds. Legally defensible by design.

Evaluation Matrix - Moderated ScoresCriteriaWeightSupplier ASupplier BSupplier CTechnical Quality40%383236Price / Value30%283026Social Value20%181417Experience10%989Total Score08488

Features

Independent panel scoring

Evaluators score independently - they cannot see other scores until moderation begins. Enforced by the system, not by process.

Automatic weighting application

Evaluation weightings applied to each score automatically. No spreadsheet, no manual calculation, no arithmetic errors.

Divergence flagging

When panel scores diverge significantly, the system flags it before moderation - ensuring the panel addresses it explicitly.

Automated evaluation report

Fully formatted evaluation report generated automatically after moderation. Board-ready, legally defensible, immediately available.

Conflict of interest gate

Evaluators cannot access submissions until they complete a conflict of interest declaration. All declarations logged and retained.

7-year audit trail

Every score, rationale and moderation decision logged with user ID and timestamp. Immutable record retained for 7 years.

Common questions

Can evaluators see each other's scores?

No. The system prevents evaluators from seeing other scores until individual scoring is complete. Enforced technically, not by process.

Is the evaluation report produced automatically?

Yes. Once moderation is complete, the evaluation report is generated automatically - formatted, weighted, with rationale, ranking and recommendation.

How long are records retained?

All evaluation records are retained for a minimum of 7 years in an immutable audit trail.

Available through G-Cloud 15

Evaluation management is available through RM1557.15 G-Cloud 15

This service can be procured through RM1557.15 G-Cloud 15 on the Digital Marketplace. Our team can help you identify the applicable service listing, define the implementation scope and prepare a written quotation.

See it in action.

Start with a free trial or pilot - no commitment required.

Request free pilot Book a demo

Evaluation is the part of a procurement that decides the outcome and the part most likely to be attacked afterwards. A challenge rarely argues that the wrong supplier won. It argues that the scoring was inconsistent, that the criteria were applied differently to different bidders, or that the reasons given do not match the record. Under the Procurement Act 2023 the answer to all three is the same: a documented, moderated, evidenced assessment that a reviewer can follow without needing anyone to explain it.

What evaluation management covers

Evaluation management is the discipline of turning a set of tender responses into a defensible award decision. It runs from designing the assessment methodology and building the scoring model, through appointing and briefing the panel, into independent scoring, moderation and consensus, and out to the award recommendation, the governance approval and the individual feedback letters. Each of those stages produces a record, and the record is the deliverable as much as the decision is.

The common misconception is that evaluation begins when bids arrive. In practice most evaluation problems are designed in months earlier, when criteria were written that cannot be scored consistently, or when weightings were set without modelling what they would do to the result. By the time responses are open, the room for correction is very small, because changing the methodology after bids are received is one of the few things that reliably turns a procurement into litigation.

Good evaluation management is therefore front loaded. Time spent testing the model against imagined responses, writing scoring guidance that a panel member can apply without interpretation, and rehearsing moderation, is time that removes risk from the phase where risk is most expensive.

The legal position under the Procurement Act 2023

The Act requires award criteria to relate to the subject matter of the contract, to be sufficiently clear, measurable and specific, and to be a proportionate means of assessing tenders. It also requires the authority to publish the criteria, their relative importance and the methodology for assessing tenders. In practice that means the scoring scale and the meaning of each score should be visible to bidders, not held back as internal guidance.

Assessment must be against the published criteria and nothing else. That sounds obvious and is broken constantly, usually in small ways: a panel member marking down a bidder for something known from a previous contract, or crediting a response for a feature nobody asked for. Both are unpublished criteria applied silently, and both are visible in a moderation record if the record is honest.

The Act also frames the most advantageous tender concept, moving away from a default assumption that price dominates. Authorities have wide latitude to weight quality, social value, and whole life cost, provided the weighting is published in advance and applied as published. The freedom is in the design stage. Once the notice is out, the model is fixed.

  • Criteria must relate to the subject matter and be clear and measurable
  • Weightings and assessment methodology must be published in advance
  • Score only against published criteria, using only the submitted response
  • Do not adjust the model after tenders are received

Building a scoring model that survives contact with real bids

A workable model has a small number of criteria, each with a clear question, a stated word or page limit, a defined scoring scale and descriptors that distinguish the scores. The descriptor is the load bearing part. If a five and a four are separated only by the words excellent and good, the panel will disagree and moderation will take days. If they are separated by what the response must demonstrate, the panel will converge quickly and the record will explain itself.

Weightings should be modelled before publication. Run three or four plausible bid profiles through the model: strong quality with high price, adequate quality with low price, and something in the middle. If a bidder can win on price alone despite a weak quality score, or if the quality weighting is so dominant that price is irrelevant, the model is not doing what the business case says it should. That is a design decision, not an accident to discover at award.

Pricing evaluation deserves the same scrutiny. Whether price is scored on a lowest price ratio, a mean based method or a benchmark, the method must be published and must be arithmetically stable when a very low or abnormally high bid arrives. Abnormally low tenders have their own process under the Act and cannot simply be excluded because the number looks wrong.

Panels: who evaluates and how they are briefed

A panel should be small enough to moderate and broad enough to cover the requirement. Three to five evaluators is typical, with technical, operational and, where relevant, service user perspectives represented. Every member should complete a conflict of interest declaration before seeing any response, and the declarations should be revisited if the bidder list contains surprises. Conflicts are not disqualifying in themselves, but undeclared ones are corrosive.

Briefing is not optional and should not be a five minute preamble. Panel members need the criteria, the scoring descriptors, a worked example, the rules on what they may and may not take into account, and a clear statement that scores must be justified in writing at the point they are given. Retro fitting rationale during moderation is the single most common way a scoring record loses credibility.

Independence during first pass scoring matters. Members should score without seeing each other's marks, because visible scores anchor the group and produce false consensus. The platform should enforce that, rather than relying on people not to look. Only after independent scores are locked should the panel see the spread and begin moderation.

Moderation and consensus in practice

Moderation is a structured conversation, chaired by someone who is not scoring, in which the panel reconciles differences and agrees a consensus score with a written rationale. The chair's job is to keep the discussion anchored to the response and the descriptors, to stop members introducing outside knowledge, and to ensure the rationale that gets recorded actually explains the score rather than restating it. A rationale reading good response, well evidenced is worthless in a challenge.

Consensus does not mean averaging. Averaging hides disagreement and produces scores nobody can defend individually. Where the panel genuinely cannot agree, the disagreement should be recorded along with the chair's decision and reasoning. An honest record of a difficult judgement is far stronger evidence than a tidy record that conceals one.

The output of moderation should be the same shape for every bidder: the consensus score, the strengths identified, the weaknesses or gaps, and the specific parts of the response relied on. That structure is what makes feedback letters straightforward to write and consistent across bidders, which in turn reduces the number of debriefs that escalate.

  • Independent scoring first, moderation second, never in reverse
  • Chair does not score and does not advocate
  • Record the reason, not a restatement of the score
  • Capture strengths and weaknesses in a consistent structure for every bidder

What goes wrong, and the challenge risk that follows

The recurring failures are predictable. Criteria that cannot be scored consistently. Panel members scoring on reputation rather than response. Moderation notes written after the award recommendation was drafted. Feedback letters that describe a weakness never mentioned in the moderation record. Late changes to the model to make the intended winner win. Each of these is discoverable, and each turns a defensible outcome into an indefensible one.

The Act sets a standstill period after the contract award notice, during which unsuccessful bidders receive an assessment summary explaining how their tender was assessed and, where relevant, how it compared with the winning tender. That summary is written from the moderation record. If the record is thin, the summary is thin, and a thin summary is an invitation to ask for more, which is often how challenges begin.

The practical defence is boring and effective: score against the published criteria, write the reason at the time, moderate properly, and let the feedback fall out of the record rather than being composed for the occasion. Authorities that do this rarely find themselves arguing about process, because the process explains itself.

Evaluating social value and other qualitative criteria

Social value is now a standard component of quality assessment and it is frequently the weakest scored section in a competition, because authorities ask for commitments without defining what a good commitment looks like. If a bidder offering twenty vague pledges scores the same as one offering three specific, resourced and measurable ones, the model is rewarding volume rather than value. Descriptors should reward specificity, deliverability and local relevance.

Whatever measurement framework is used, whether a TOMs style approach or a bespoke one, the evaluation should test whether the commitment is genuinely additional, whether it is proportionate to the contract, and whether the bidder has explained how it will be delivered and evidenced. Committing to outcomes that would have happened anyway is common and should not attract credit.

The same logic applies to other qualitative areas such as social and environmental sustainability, safeguarding, or accessibility. Ask a question that can be answered with evidence, publish descriptors that reward evidence, and make clear that commitments accepted at tender become contractual obligations that will be monitored after award.

Evidence, audit and the record you will need later

The evaluation record should be able to answer four questions without human assistance: who scored what, when they scored it, why they scored it that way, and what changed between the first pass and the final consensus. If any of those requires someone to remember, the record is incomplete. Version history matters as much as content, because the credible objection is not that a score was wrong but that it moved without explanation.

Internal audit and external scrutiny come at this from different angles. Audit typically asks whether the process followed the authority's own rules, including contract standing orders and delegated authority limits. A challenge asks whether it followed the published rules and treated bidders equally. A single record built during the process satisfies both. A record assembled afterwards usually satisfies neither.

Retention is part of this. Evaluation material may be needed long after award, for a freedom of information request, a contract dispute, or a follow on procurement. Keeping it in an accessible, structured form, rather than in a departed officer's mailbox, is a governance basic that is still routinely missed.

How eSourcing Data supports evaluation and moderation

eSourcing Data provides evaluation and moderation as part of a source to contract platform, so the scoring model sits with the tender documents, the responses and the eventual contract rather than in a separate spreadsheet. Panels score independently, scores are locked before the spread is visible, moderation is captured with rationale against each criterion, and the audit trail is produced as the work happens rather than reconstructed for an auditor.

Because governance, analytics and reporting are in the same system, heads of procurement can see evaluation progress across a portfolio of competitions, spot the ones that are stalling, and understand where panel capacity is the constraint. Data is held in the UK and the platform operates in line with GDPR, which usually settles the information governance conversation early.

Buyers who prefer to buy through a framework can access eSourcing Data software through RM1557.15 G-Cloud 15, where 28 software services are listed on the Digital Marketplace alongside cloud support services, purchased as call off contracts. Authorities that need help running the evaluation itself, rather than a system to run it in, can use consulting and outsourced procurement support instead.

Frequently asked questions

What is moderation in tender evaluation?

Moderation is the structured session where evaluators compare their independent scores, discuss differences against the published descriptors and agree a consensus score with a written rationale. It is chaired by someone who does not score. Moderation is not averaging: the point is to reach a justified position the whole panel can stand behind, and to record why that position was reached.

Do I have to publish my scoring methodology to bidders?

The Procurement Act 2023 requires award criteria, their relative importance and the assessment methodology to be published. In practice that means bidders should be able to see the scoring scale and what each score means, not just the weightings. Withholding scoring guidance while relying on it internally is a common and avoidable source of challenge.

How many people should be on an evaluation panel?

Three to five is typical for most competitions. You want enough breadth to cover the technical, operational and user perspectives relevant to the requirement, and few enough people that moderation is manageable and scheduling is realistic. Every member should declare conflicts of interest before seeing any response, and every member should be briefed on the criteria and descriptors.

Can evaluators use their own knowledge of a supplier?

No. Scores must be based on the published criteria and the submitted response. Prior experience of a supplier, good or bad, is an unpublished criterion applied unevenly, since it is only available for incumbents and local suppliers. Poor past performance may be relevant through the exclusion and past performance provisions, but that is a separate assessment with its own process and record.

What should an assessment summary contain?

It should explain how the tender was assessed against each criterion, the score awarded, the reasons for that score, and where relevant how it compared with the winning tender. It is written from the moderation record, so its quality is determined months earlier. Consistent structure across bidders reduces the number of debriefs that turn into disputes.

How do you evaluate social value fairly?

Ask questions that can be answered with evidence, and write descriptors that reward specificity, additionality, deliverability and local relevance rather than the number of pledges. Test whether the commitment is proportionate to the contract and whether the bidder has explained how it will be delivered and measured. Make clear that accepted commitments become contractual obligations monitored after award.

What is the standstill period for?

Standstill follows the contract award notice and gives unsuccessful bidders a window to review the assessment summary and decide whether to challenge before the contract is entered into. It exists so that problems can be raised while they are still fixable. Rushing or mishandling standstill is one of the more expensive procedural mistakes available to a buyer.

Can I change the evaluation model after receiving bids?

No. Changing criteria, weightings or scoring methodology after tenders are received undermines equal treatment and is very difficult to defend. If a genuine flaw emerges, the options are to seek clarification within the published rules or, in serious cases, to abandon and re run the procurement. Modelling weightings before publication is what prevents this situation.

Further reading

For buyerseSourcing and tenderingCompliance and auditSocial valueProcurement LibraryG-Cloud 15 service directoryBook a demo