Evaluation Methods Quarterly: Synthesizing an Evaluative Judgment
Classically, there are five steps in the logic of evaluation: we define the evaluand (e.g., program, personnel, product), set criteria and standards, gather performance data, compare that performance to our standards, and then synthesize an overall judgment of merit, worth, or significance. In a previous post I argued that setting standards is the hardest problem in evaluation. If that is the hardest, synthesis is the most neglected.
The final synthesis step makes the difference between a jumble of findings and a real evaluation. After assembling a dozen data sources and comparing each to our benchmarks, we should not hand the mess to the stakeholders like a toddler in the kitchen with a cheery “now it’s your turn to finish cooking!”
The best practice, by contrast, is to find a non-arbitrary way to combine all the results into a single evaluative conclusion. When we rush past this intellectual work, we usually end up with a research study instead of an evaluation. Enlightening perhaps, but not actionable. (To be sure, the “working logic” or evaluation may look different than what I’m describing, which is the “general logic” of an evaluation, to borrow a distinction from Deborah Fournier.)
My theory is that we avoid synthesis because gathering data and running comparisons feel like science but declaring a program or a product “good on balance” feels like opinion. Synthesis requires a lot of judgment, and we live in a society where asking “who am I to judge?” sounds empathetic instead of just confused. You’re a community stakeholder or a decision-maker or a voter – that’s who.
In a study of forty program evaluation reports, Marthe Hurteau and her colleagues found that only about half rendered an explicit judgment at all, and that the judgments offered, while generally legitimate, were rarely justified. The ingredients for a warranted judgment (the criteria, standards, evidence, all the good stuff), were present across almost all the reports. That is, the evaluators usually had what they needed to make a synthetic judgment but they choked. But passing the buck just leaves the hardest judgment to whoever reads the report and has to figure out how to weight competing sources of evidence downstream.
Chutes and Ladders
The approach to synthesis that we might want to use reflexively is almost always invalid. The default is to assign each criterion a numeric weight, score each on a common scale, multiply, and sum. Out comes a single score. This looks objective and is nice in a slide deck, but this “numeric weight-and-sum” comes with a big assumption. That assumption is that our criteria are fully compensatory, that is, strong performance on one criterion can always buy back weak performance on another. Imagine synthesizing ratings of a restaurant like this:
- Taste (30%) – chips and salsa were excellent, top marks.
- Menu (30%) – many options including those for specialty diets, perfect.
- Ambiance (20%) – lovely decor, no notes.
- Service (20%) – never brought main course to table so all I got were chips, zero points.
- TOTAL: this restaurant was 80% satisfactory, B-.
This is what we sound like when we use compensatory scoring for a housing program that is extraordinarily cost-efficient but reaches almost none of the highest-need families it was funded to serve. Average the “cost” criterion together with the “equity-of-access” criterion and you might get a respectable composite score. You can do that sort of thing with math if you want, but you really shouldn’t.
Compensatory criteria trade off against one another. A program can be a little weak on staff satisfaction if it is much stronger on some participant outcomes. Non-compensatory criteria cannot be exchanged at any price. We call these bars, hurdles, essentials, or must-haves. A program that harms the people it serves is not redeemed by being inexpensive and having good participant satisfaction – that’s cigarettes you’re thinking of. Some interventions that sound good, like specialized group therapy for people with traumatic experiences (Critical Incident Stress Debriefing), have turned out in meta-studies to actually increase PTSD (Litz, Gray, Bryant, and Adler, 2002). No amount of excellence elsewhere lifts the overall judgment above “inadequate” when you’re worsening the patient’s PTSD.
It’s interesting to ask why compensatory averaging is usually wrong. It is usually wrong because we go for it before checking whether certain conditions hold and often they do not. For one thing, we need to make sure that none of our criteria are pass-fail. The restaurant that does not bring the main course should fail automatically, not get a B- or even a D. Another question is whether the criteria interact in a purely additive way or whether something more interesting happens. Is someone really likely to give a top score on taste if they don’t like the menu? Perhaps there is a minimum level of menu performance that is necessary to open the gate to the highest scores on taste. If I were to diagram some of the real conversations I’ve had with stakeholders about what counts as good performance, it would look less like a set of parallel sliders on a soundboard than a game of Chutes and Ladders.(1)
Normative Midwifery
Planting a categorical hurdle is often necessary if we want to faithfully reflect the priorities of stakeholders. But, why not call everything a must-have?
I think we can set some ground rules for hurdles. First, there should only be a few. A hurdle is a floor below which the program’s core purpose fails or someone is harmed.(2) A criterion we care about a great deal but could trade off is still compensatory, but we assign it a heavy weight. Second, connect each hurdle to something outside mere preferences, such as to the evaluand’s stated purpose, a funder’s written requirement, a legal duty, or a demonstrable risk of harm. Compensatory criteria have lower stakes than categorical ones, so we need to be sure that these latter definitely aren’t arbitrary. Third, set the threshold before the data arrive. “Serve the priority population” is not yet a hurdle; “enroll at least X percent from the priority neighborhoods” is. A hurdle whose height is set too high after the race begins (or ends) invites accusations of bias. If you don’t know what a reasonable standard would be, say this, and the evaluation team can research it.
Michael Scriven’s “Qualitative Weight and Sum” (QWS) is a good default that orients the synthesis in the right direction without introducing false precision. In QWS, we rate the importance of each criterion, before seeing the data, on a short ordinal scale (e.g., essential, important, or minor) and rate performance on each dimension with qualitative grades (e.g., excellent to unacceptable). We then reason our way to a conclusion in words, using the ordinal standards to adjudicate the trade-offs. Notice that performance indicators can still be measured continuously as long as these indicators can be mapped to ordinal scores. The improvement here is that, because its weights and grades are ordinal labels rather than decimals, we can’t collapse the result into a single number and the trade-offs stay in view. Stakeholders can see exactly where they would have to disagree with us to reach a different conclusion.
One of the main motivations for QWS is that it makes trade-offs very noticeable compared to a single compensatory weighted sum. Which strengths come with which weaknesses? In cost-inclusive evaluation, one high-leverage recommendation is often to suggest the adoption of an intervention that is slightly less effective (perhaps not even noticeably so to the average participant) but orders of magnitude less expensive, allowing the intervention to reach far more people and move the needle. Trade-offs are unavoidable and critical, so we should use methods that make them very obvious.(3)
The issues we’ve been discussing with compensatory “numeric weight and sum” don’t follow directly from the fact that it is a form of mathematical calculation. In addition to Qualitative Weight and Sum, there are some statistical methods that preserve trade-offs and avoid reducing the synthesis judgment to a single score. Data Envelopment Analysis, which evaluators have used for decades to gauge the relative efficiency of programs, declines to impose a single weighting as well. It judges each evaluand (e.g., program, personnel, product) by the mix of weights most favorable to that specific evaluand, provided every other evaluand may be judged by that same standard.
The crux of the issue with assigning weights and categorical gates is that they embody normative commitments of stakeholders. A weight of 20% versus 10% says that the first thing is twice as valuable as the second thing to us. In the abstract that sounds fairly harmless, but in practice, that 20% may be “number of deaths prevented” and that 10% may be “avoids entrenchment of social inequality.” The choices really are that serious. What’s more – most people don’t know until they are asked in a structured fashion what the criteria are or how they compare.
Recently I found myself in conversation with an obstetrician at a social occasion. I asked her why she chose that branch of medicine over others she explored during her clinical rotation. She said she likes it because birth is always the same challenge that always needs to end the same way – regardless of how she gets there, that baby needs to be born.
One of the jobs of the evaluator is to midwife the process of helping people find out what they value. Ideas of what counts as “good enough” are in our minds and cultures, but usually there is labor involved in bringing them into the light. Sometimes the birth is complicated, but we stick with it because one way or another that baby needs to be born.
Getting It Right
Evaluative synthesis should be neither a mechanical calculation nor a gut intuition. It uses explicit rules, looks for categorical hurdles, and sets up a transparent decision-making method. When we do synthesis right, the whole evaluation reads as a logical argument running from standards and weights to a conclusion, with acknowledged trade-offs. When you reach this final step, a few questions help:
- Which of my criteria are actually compensatory, and which are hurdles that no amount of excellence can offset?
- For each hurdle, can I point to a purpose, requirement, or risk of harm outside my own preferences that justifies it, and did I fix its threshold before the data came in?
- Did I set my weights before I saw the performance data, or after?
- If I combined criteria into a single score, does that number contain trade-offs the stakeholders would actually endorse?
- Could a reader trace my overall judgment back through the standards, or have I asked them to take the overall verdict on faith?
Notes
(1) Chutes and Ladders is itself based on an ancient Indian game about karma called Snakes and Ladders. Just like evaluation criteria, the original game and its translated descendants are about how to be good.
(2) There are many real-world evaluands that demand we treat harm as compensatory rather than as a categorical zero-sum hurdle. For example, space exploration is inherently dangerous and some harm is considered regrettable but an acceptable loss. As astronaut Lisa Nowak said, “Of course risk is part of spaceflight. We accept some of that to achieve greater goals in exploration and find out more about ourselves and the universe.” Sometimes in policy and planning, there is no choice that does not cause harm – in those cases, harm is de facto compensatory.
(3) I recommend Scriven’s qualitative weight and sum in this post because it is vastly superior to compensatory weighted averages. However, there are of course several quantitative evaluation options that are natural extensions of Scriven’s logic. I am partial to visualizing the evaluation synthesis in multidimensional space, for example, since this is often a better portrayal of aspects of value that increase on a continuous basis, such as cost savings. This is compatible with categorical hurdles as well, which create no-go zones in the outcome space.
Read More from our Evaluation Methods Quarterly Series:
Evaluation Methods Quarterly is a blog series by Dr. Anthony Clairmont, EVALCORP’s Lead Methodologist. Dr. Clairmont explores key methodological concepts, challenges, and innovations in evaluation, offering insights to strengthen practice and advance the field.
- Evaluation Methods Quarterly: Synthesizing an Evaluative Judgment
- Evaluation Methods Quarterly: Asking Good Evaluation Questions
- Evaluation Methods Quarterly: Moving from Disparities to Standards
- Evaluation Methods Quarterly: Setting Standards is the Hardest Problem in Evaluation
- Evaluation Methods Quarterly: Randomization is Underused in Program Evaluation