Almost every nonprofit impact report begins life as the same picture: five boxes in a row with arrows between them. Resources, activities, outputs, outcomes, impact. It comes from the W.K. Kellogg Foundation’s Logic Model Development Guide, the 2004 reference that most grant templates quietly descend from.
Inside that guide, in the chapter that teaches you to build the thing, is a sentence that contradicts the picture. “For that reason, Exercise 1 isn’t filled out from left to right. This exercise asks you to ‘do the outcomes first.'”
The diagram reads left to right. The manual says fill it in from the right. Almost nobody does, because the picture is the part that gets photocopied into the grant template and the instruction is the part that stays in the PDF. And the order you fill those boxes in is not a stylistic preference. It decides which numbers exist twelve months later, when a program officer asks what changed.
The columns are not the same kind of thing
Kellogg’s evaluator Beverly Anderson Parsons puts the case bluntly in the guide’s own margin: “I have become convinced that it makes a considerable difference if you do the outcomes before planning the activities. I definitely advocate doing the outcomes first! I find that people come up with much more effective activities when they do. Use the motto, ‘plan backward, implement forward.'”
The reason the order matters is that two of those five boxes are records of what you did and the rest are claims about what changed. The guide draws the line precisely. Outputs are “the direct results of program activities”, described in terms of size and scope: the number of classes taught, meetings held, or materials produced and distributed. Outcomes are “specific changes in attitudes, behaviors, knowledge, skills, status, or level of functioning”.
Outputs can be reconstructed from your files in an afternoon. Your calendar knows how many sessions you ran. Your sign-in sheets know how many people came. Outcomes cannot be reconstructed at all, because nobody wrote down what changed unless somebody decided in advance to ask.
The guide is unusually direct about how this goes wrong, and it is worth quoting because it describes the report almost everyone files: “Some grantees think activities are ends unto themselves. They report the numbers of participants they reach or the numbers of training sessions held as though they were results.” And then the line that settles the argument: “Conducting an activity is not the same as achieving results from the accomplishment of that activity. For example, being seen by a doctor is different from reducing the number of uninsured emergency room visits.”
Three rungs of paperwork and one rung of program design
Here is the part that reframes the whole exercise, and you can see it most clearly in the public transparency ladder that funders actually check.
Candid, the organization behind the nonprofit profiles most grantmakers look at, awards Seals of Transparency at four levels. Bronze is “Make sure donors can find you.” Silver is “Tell donors about your work to make the case for support.” Gold is “Share your financials and leadership demographics to gain funders’ trust and support.” Platinum is “Highlight your impact. Share your goals, strategies, and metrics to boost funding.”
Read those as a sequence and something jumps out. The first three ask for information that already exists somewhere in your building. Your address. A description of what you do. Your financials. Who is on your board. You can climb from nothing to Gold in an afternoon without changing one thing about how your programs run.
Platinum asks for something else. Candid’s own requirements state that “At least one recent impact metric is required to earn a Platinum Seal”. That is the rung you cannot climb in an afternoon, and not because the form is longer. It is because the number had to be collected while the program was still running, by someone who decided a year ago that it was worth collecting.
The ladder looks like four steps of escalating paperwork. It is really three steps of paperwork and one step of program design, wearing the same uniform. Every small organization that has stalled at Gold and assumed it was a documentation problem has actually run into the left-to-right habit from the first section.
One caveat on the incentive, stated plainly because it is the kind of number that gets repeated without its provenance. Candid says that “nonprofits with a Seal receive, on average, 62% more in donor contributions than organizations without a Seal.” That is Candid’s own figure about Candid’s own product, not an independent study, and organizations that keep their profile current are plausibly different from those that do not in ways that would move donations anyway. Treat it as a reason to look, not as a forecast.
Why this lands hardest on the smallest organizations
Most published advice on impact measurement is written as though a program evaluator exists. For the typical US public charity, that assumption is simply wrong. The Urban Institute’s National Center for Charitable Statistics found that “66.6 percent had less than $500,000 in expenses (211,782 organizations); they composed less than 2 percent of total public charity expenditures ($32.8 billion)”.
Two thirds of the sector, by headcount of organizations, runs on less than half a million dollars a year. In a body that size there is no evaluation staff, no data analyst, and frequently no full-time development role. The person writing the impact report is the person who ran the program, and they are writing it in the evening.
That constraint is not a reason to skip the work. It is the reason the order matters so much, because a small organization gets exactly one pass at collecting each year’s data and cannot go back for a second.
Four moves that produce a report you can actually write
Move 1: Write the outcome sentence first, and make it capable of being wrong
Before the activities, before the budget, write one sentence describing what will be different for a person because your program existed. Then apply a single test to it: what result, if you saw it next year, would prove this sentence false?
“We will strengthen our community” fails the test. Nothing could contradict it. “Families who complete the six-week course will be able to name three local services they did not know about before” passes, because a survey where nobody can name any would falsify it. Kellogg’s own definition is your checklist here: attitudes, behaviors, knowledge, skills, status, or level of functioning. If your sentence names none of those six, it is a mission statement, and mission statements cannot be reported on.
Write one. Two at most. The instinct to list eight outcomes is the instinct that guarantees none of them get measured.
Move 2: Pick the indicator by what you can observe, not by what sounds rigorous
An indicator is the specific thing you will count or ask. For each outcome sentence, choose one, and choose it by answering a practical question: who will observe this, at what moment, and will they still be in the room?
A food program can ask whether a household ran out of food before the end of the month. A job-training program can ask whether someone is working, and at what hourly wage. A tutoring program can compare a reading level at intake to the same measure in June. None of these needs statistical training. They need a decision, made early, about one number.
Resist the instinct to pick the indicator a large foundation would use. The best indicator is not the most sophisticated one; it is the one that will still be collected in month nine when everyone is tired.
Move 3: Put the ask inside the program, not after it
This is the move that most often fails, and it fails for a structural reason. Once a program ends, the participants leave and your response rate collapses. The follow-up survey sent in January to people you last saw in June is the single most common source of the sentence “we were not able to gather sufficient data.”
So attach the ask to a moment that already exists and that people already attend. The last fifteen minutes of the final session. The appointment where someone collects a certificate. The intake for the next cycle, where you ask returning participants about the previous one. If you need a follow-up at six months, schedule the specific date now and put it in the calendar of a named person, not of the organization.
This is the same discipline as booking a re-check date on a diagnosis rather than trusting yourself to remember, which is the habit we argued for in Customer Journey Map: Build It From the People Who Left. Evidence that depends on someone’s future goodwill is not evidence you have.
Move 4: Say what you are comparing to, or say that there is nothing
Your report will contain a number. That number means nothing without its comparison, and the comparison is where honest reporting separates from the other kind.
There are three respectable answers and one that is not. You can compare to the same people before the program, which is the usual small-organization answer and is genuinely useful. You can compare to a documented baseline for your area, if a public source provides one. You can compare to your own prior year. What you cannot do is state a change and let the sentence imply that your program caused it, when you have no comparison at all.
Write the comparison into the sentence: “Of 34 participants who completed the course, 21 reported being employed six months later, compared with 9 who reported employment at intake.” That sentence is modest, checkable, and considerably more persuasive to an experienced program officer than a confident claim with no basis under it. Naming the limit is what makes the rest of the report credible.
A worked example, and it is a hypothetical
The following is invented for illustration. It is not a real organization, a case study, or a result that anyone reported. It is here because the shape of the change is the point, not the numbers.
Suppose a small literacy nonprofit runs an after-school reading program and has filed the same report for three years: sessions held, volunteers recruited, children served. Renewals keep arriving, but the funder has started asking follow-up questions and the answers keep being estimates.
Working backwards this time, the outcome sentence is written first: children who attend at least twenty sessions will gain more than one grade level in reading over the school year. That sentence can be wrong, which is what makes it useful. The indicator follows from it directly: the reading assessment the partner school already administers in September and again in May, which nobody has to build.
Move 3 is where the plan changes shape. Rather than requesting scores from the school in June, when the school year is ending and nobody answers email, the program adds a line to the enrollment form that parents already sign, giving permission to receive both assessment results, and the September score is collected at the same session where the child is enrolled.
Twelve months later the report contains a sentence the previous three could not: a specific number of children, a specific gain, and a comparison to their own September starting point. It also contains a paragraph noting that children who attended fewer than twenty sessions were not assessed, so the figure describes the group that finished. That paragraph costs nothing and is the reason the rest of the page can be believed.
What it costs to actually collect this
Three real options, and the useful finding is at the bottom of the range rather than the top.
Google Workspace for Nonprofits is listed at “$0 USD / user / month” for eligible organizations, up to a maximum of 2,000 users, and includes Forms alongside the rest of the tools. For an organization that needs a survey and a spreadsheet, this is the entire requirement at no cost.
Jotform offers a free Starter plan with 5 active forms and 100 monthly submissions. Its Bronze plan is $39 per month, or $408 billed yearly which works out to $34 per month, and the pricing page states a nonprofit discount of “50% off Bronze/Silver/Gold”. Its advantage over a plain form is conditional logic and cleaner handling of the same survey repeated across cycles.
SurveyMonkey is the name most people reach for first, and its team pricing is where the trap is. Team Advantage is listed at “$30/ user / month” billed annually with a minimum of three users. That is $90 a month committed before a single participant answers a question.
Set that against the sector data above and the conclusion inverts the usual advice. Jotform’s free tier caps at 100 submissions a month. A great many of those 211,782 organizations under $500,000 in expenses serve fewer than 100 people in a program cycle, which means the free tier is not a trial they will outgrow. It is the finished tool. The thing that stops a small nonprofit producing an impact report has never once been the price of a survey form. It is whether anyone asked the question while the participants were still in the room.
Where an AI prompt genuinely helps, and where it cannot
The Measure and Communicate Your Impact prompt on our sibling site takes the objectives an organization is already carrying and turns them into a tracking framework, a set of candidate indicators, and draft narrative language for funders. You can find it at businessprompter.com/prompt/measure-and-communicate-your-impact, alongside the rest of the library at BusinessPrompter.com.
What it is good at is the work that gets skipped at 9pm: proposing indicators for an outcome you have written, catching that you have listed an output where you promised an outcome, and turning a page of notes into something a program officer can read in four minutes. That drafting pass is precisely the work a small organization would otherwise buy from a consultant or simply not do.
What it cannot do is Move 3. It cannot sit in the last session and ask the question. It has no access to your participants, your partner school, or the relationship that makes someone answer an email six months after a program ended. And there is a specific hazard worth naming: hand it an outcome with no data behind it and it will write you a fluent, well-organized paragraph about that outcome anyway. The polish is not evidence. The check on that is yours.
That division is the honest case for using it. The writing is the part a person is worst at after a long day, and the asking is the part only a person can do, because the reason anyone replies to your follow-up is that they know the staff member sending it.
The part we would argue for hardest
The sector’s advice on impact measurement is written for organizations that have a choice about what to measure. Two thirds of US public charities do not have that choice, and pretending otherwise produces the four-page activity report that nobody believes and everybody files.
The better move for a small organization is almost the opposite of what the guidance implies: measure one thing properly, and state clearly what you did not measure. A report carrying a single sourced number, its comparison, and an honest paragraph about its limits is more persuasive than four pages of session counts. It is also faster to write, which matters when the person writing it also ran the program.
There is a related discipline in deciding which work is worth this effort at all, and the scoring method in Impact Effort Matrix: Score It in Your Own Hours applies directly, because an evaluation plan that assumes hours you do not have will be abandoned in month four. And if the governance side is the part that is stalling, Nonprofit Board Responsibilities: Use Your Own Form 990 covers the documents a board is already accountable for, while Monthly Giving Program: Start With Your Smallest Donors takes the revenue question that usually sits behind the reporting pressure.
Take last year’s report and find its outcome sentence. Read it aloud, then ask what result would have proven it false. If no result could have, the problem was never the writing, and next year’s program is the only place it can be fixed.
Common questions
What is the difference between an output and an outcome?
An output is a record of what you did and an outcome is a change in the people you did it for. The Kellogg guide defines outputs as “the direct results of program activities” such as classes taught or materials distributed, and outcomes as “specific changes in attitudes, behaviors, knowledge, skills, status, or level of functioning”. If your number came out of your own calendar, it is an output.
We have no outcome data from last year. What do we put in this year’s report?
Report the outputs you have, label them as outputs rather than dressing them as results, and add a short section describing the one outcome you will measure this cycle and the exact moment you will collect it. Program officers read a great many reports and can tell the difference between an organization that is hiding a gap and one that has named it and scheduled the fix.
Do we have to prove our program caused the change?
No, and claiming you have when you have not is the fastest way to lose credibility. State the change and state your comparison, whether that is the same people at intake, a published baseline, or your own prior year. If you have no comparison, say so in the sentence rather than leaving the causal claim to be inferred.
What does it cost to start collecting outcome data?
It can be nothing. Google Workspace for Nonprofits is listed at “$0 USD / user / month” for eligible organizations and includes Forms, and Jotform‘s free Starter plan covers 5 forms and 100 monthly submissions, which is more than many small programs need in a cycle. By contrast SurveyMonkey‘s Team Advantage plan is “$30/ user / month” billed annually with a three-user minimum. The binding constraint is when you ask, not what you pay.
