The hardest sentence to write in this course is the one that estimates what your intervention will actually do.
The hardest sentence to write in this course is the one that estimates what your intervention will actually do. Students reach for a published effect size because it is the most defensible-looking number available, and that instinct is the source of the most common failure in the report. The effects reported in academic papers and the effects observed when the same class of intervention is run at scale are not the same size, and the gap has been measured. A recommendation built on the published figure is a promise the client cannot keep.
Note on the code: this guide describes the UNSW Sydney course. COMM3000 codes also exist at several North American universities with entirely different content, so check the course title on your own enrolment before reading further.
Author: MAAS Editorial Team · Reviewed by a Senior Economics mentor (PhD, Applied Economics)
Last updated: 2026-08-14
Category: writing-tips
How is the course built?
Direct answer: COMM3000 Evidence-based Intervention Design and Evaluation is a 6 unit-of-credit undergraduate course in the UNSW Business School, taught by the School of Economics in Term 3 at Kensington. The published course outline describes it as a fully problem-based learning experience in which you work as a management consultant or advisor, using insights from data analysis to address one of three problem types: motivating and designing an intervention, identifying and managing the risk attached to it, or evaluating its effectiveness and efficiency.
Evidence: The 2025 Term 3 course outline lists three assessment items: an individual presentation at 35 per cent, a group project at 40 per cent, and individual problem sets at 25 per cent. The outline also describes the course as a synthesis course, meaning you are expected to apply and integrate knowledge from across your whole degree rather than to demonstrate one new technique. That framing matters for how you are marked. In a synthesis course the marker is not checking whether you learned the content of this course; they are checking whether you can select the right tool from three years of study and justify the selection.
Example: A student who had done well in econometrics wrote a technically clean analysis and scored in the middle band. The method was not the problem. He had never explained why that method suited this client's decision, so the work read as a demonstration of skill rather than as advice.
Weightings and task types are revised between offerings, so treat the figures above as the shape of the assessment rather than as your own numbers, and confirm against the outline published for your term.
What does "evidence-based" actually demand?
Direct answer: A comparison. Evidence that an intervention worked requires some account of what would have happened without it, and a before-and-after measurement on its own does not supply that.
Evidence: The problem is not only that other things change over time. It is that the way programmes select their participants builds a false result into the data. Barnett, van der Pols and Dobson (2005) set out the mechanism: when unusually high or low measurements are followed by measurements closer to the average, natural variation looks like real change. Interventions are routinely targeted at exactly the cases that were unusually bad at the moment of measurement, which is the condition under which this effect is largest. A programme aimed at the worst-performing branches, the least engaged customers, or the lowest-scoring students will show improvement whether or not it did anything, and a report that presents that improvement as its result has measured the selection rule rather than the intervention.
| The claim in a weak report | What the marker asks | What the stronger report supplies |
|---|---|---|
| Uptake rose 12 points after the campaign | Compared with what? | A control or comparison group, or a stated reason none was feasible |
| We targeted the lowest performers and they improved | Would they have improved anyway? | Explicit treatment of regression to the mean |
| The result was statistically significant | Significant, and how large? | Effect size alongside the p-value, and what size would matter to the client |
Example: A team evaluated a reminder sent to customers who had missed a payment. Repayment rose sharply. The number did not survive the question of what the same customers would have done unprompted, since missing one payment is exactly the kind of unusual month that tends to be followed by a normal one. Adding a comparison group of similar customers who received nothing turned an unusable figure into a finding.
Why is the effect you cited six times too big?
Direct answer: Because the published literature is a filtered sample. Studies with large effects are more likely to be written up and accepted, and interventions run inside real organisations face frictions that a trial designed by researchers does not.
Evidence: DellaVigna and Linos (2022) assembled every trial run by two of the largest behavioural units in the United States, 126 randomised controlled trials covering 23 million people, and compared them with the published academic literature. In academic papers the average effect was 8.7 percentage points on take-up, a 33.4 per cent increase over the control condition. In the units running interventions at scale the average effect was 1.4 percentage points, an increase of 8.0 per cent. Both are real. They are not interchangeable, and quoting the first while proposing the second overstates your case by roughly a factor of six.
This is the single most useful piece of evidence you can bring to a report in this course, and it cuts both ways. It disciplines your forecast, and it also protects a modest projected effect from looking like a weak recommendation. An intervention delivering 1.4 points at negligible cost is not a disappointment; it is what the scaled evidence says a good one looks like.
Example: A group projected a 30 per cent improvement by citing a well-known trial. Asked what would happen if the effect came in at a fifth of that, they had no answer, because the whole business case rested on the headline figure. Rebuilding the case around the scaled estimate produced a smaller claim that survived questioning, which is the outcome the course is testing.
How do you design an intervention rather than pick one?
Direct answer: By starting from a diagnosis of why the behaviour is not happening, not from a list of tactics. Most weak proposals in this course choose the intervention first and construct the reasoning afterwards.
Evidence: Michie, van Stralen and West (2011) built their framework around exactly this problem, arguing that interventions are too often chosen by familiarity rather than by analysis. Their COM-B system holds that a behaviour requires capability, opportunity and motivation, and the practical value for your report is that the three point to different remedies. If people cannot do the thing, information will not help. If they can and want to but the environment blocks them, persuasion is wasted. Naming which of the three is binding is the diagnostic step that turns a proposal into an argument, and it is the step most often skipped.
Example: A team proposed an education campaign to raise enrolment in a workplace scheme. Their own interview data showed staff already understood the scheme and wanted it, but the sign-up form required a document most did not have to hand. The binding constraint was opportunity, so the education campaign would have spent the budget on the one thing that was not broken. The finding was already in their data; what was missing was the framework that made them look at it.
Where does efficiency come into the recommendation?
Direct answer: In the ratio, not in the total. The course description names effectiveness and efficiency separately, and a recommendation that reports only the effect has answered half the question.
Evidence: Benartzi and colleagues (2017) calculated impact-per-dollar ratios for behavioural interventions and compared them against traditional policy instruments such as tax incentives and financial inducements, and found the behavioural options often compared favourably. The comparison is what matters for your report. A large effect achieved expensively can be the worse option, and a small effect achieved almost for free can be the better one, which is why the effect size alone cannot carry a recommendation. The authors are careful that more calculations are needed before the comparison generalises, and that caution is worth reproducing rather than dropping when you cite them.
Example: Two options were put to a client. The first raised participation by 9 points through a subsidy. The second raised it by 3 points by changing a default. The second was recommended and accepted, because the report showed cost per additional participant rather than participation alone.
What does the individual presentation test that the group project cannot?
Direct answer: Whether the reasoning is yours. A group report can be strong while a member of the group cannot reconstruct why the analysis went the way it did, and at 35 per cent the individual component is where that shows.
Evidence: The course outline describes the presentation as thoughtful examination and assessment that goes beyond summarising content, drawing on both reflection and critical thinking, and maps it to five programme outcomes including business knowledge, problem solving and communication. In practice this means the questions afterwards are the assessment. A presenter who divided the work efficiently and owns only their section can present fluently and then stall on the first question about a choice someone else made.
Example: A student who had handled the data cleaning presented the team's recommendation clearly and was asked why the comparison group had been defined the way it was. She had not made that decision and could not defend it. The fix is unglamorous and cheap: before the presentation, have each member explain one decision they did not personally make.
Where do international students most often lose marks?
Direct answer: In the consultant register, which sits between two habits that many students bring with them. One is deference, where the recommendation is softened until it stops being a recommendation. The other is the descriptive report, where the analysis is presented thoroughly and the reader is left to draw the conclusion.
Evidence: The role the course assigns is advisor, and an advisor is assessed on the quality of a stated recommendation under uncertainty. Writing "the organisation may wish to consider several options" avoids being wrong by declining to say anything, and markers read it accordingly. The stronger move is to recommend one course of action, state the condition under which you would change your mind, and name what you would measure to find out. That is more cautious than hedging, because it exposes the reasoning rather than hiding it.
Example: A report ended with three options and no ranking. Adding two sentences, one naming the preferred option and one naming the evidence that would overturn it, moved the recommendation section a full band. Nothing in the analysis changed.
What do MAAS mentors actually do on a course like this?
On a course assessed by consulting output, the useful help arrives before the analysis is finalised rather than after. That means checking that your evaluation design can actually detect the effect you claim, before you collect anything; pressure-testing your projected effect against what scaled interventions really deliver, so the business case survives the first hard question; and rehearsing the questions that follow a presentation, which is where an individual mark is won or lost. The analysis is yours and so is the report. What a mentor adds is the awkward question, asked while there is still time to act on the answer.
If you want someone to stress-test an evaluation design before your team commits to it, send us the brief and the data you have so far.
Frequently asked questions
Do I need econometrics before taking COMM3000?
The course is described as a synthesis course drawing on skills built across your degree, and quantitative analysis is central to it. Check the conditions for enrolment in the handbook entry for your own year, since these are revised between offerings.
Is COMM3000 offered outside Term 3?
The class timetable lists Term 3 delivery in person at Kensington. Availability can change between years, so confirm against the current timetable before planning your enrolment sequence.
What if a randomised design is impossible for my client's problem?
Then say so and explain what you did instead. Naming the limits of your design and reasoning about what those limits do to your conclusion demonstrates the evaluation outcome far better than presenting a weak design as though it were a strong one.
How much does the group project mark depend on my individual contribution?
The group project is one item among three, and the two individual items together carry more weight than it does. Practically, this means a strong group cannot carry you and a weak group cannot sink you, but neither can you rely on your section alone at the presentation.
Can I recommend doing nothing?
Yes, if the evidence supports it and you show the reasoning. A recommendation not to intervene, backed by a cost comparison and a clear statement of what would change your mind, is a legitimate consulting output and is treated as one.
Related reading
- COMM1190: data insights and decisions assignment guide
- COMM1170: organisational resources assignment guide
- COMM2501: data visualisation and communication assignment guide
References
Barnett, A. G., van der Pols, J. C., & Dobson, A. J. (2005). Regression to the mean: What it is and how to deal with it. International Journal of Epidemiology, 34(1), 215–220. https://doi.org/10.1093/ije/dyh299
Benartzi, S., Beshears, J., Milkman, K. L., Sunstein, C. R., Thaler, R. H., Shankar, M., Tucker-Ray, W., Congdon, W. J., & Galing, S. (2017). Should governments invest more in nudging? Psychological Science, 28(8), 1041–1055. https://doi.org/10.1177/0956797617702501
DellaVigna, S., & Linos, E. (2022). RCTs to scale: Comprehensive evidence from two nudge units. Econometrica, 90(1), 81–116. https://doi.org/10.3982/ECTA18709
Michie, S., van Stralen, M. M., & West, R. (2011). The behaviour change wheel: A new method for characterising and designing behaviour change interventions. Implementation Science, 6, 42. https://doi.org/10.1186/1748-5908-6-42
