Skip to content

How the evidence and change get made.

The innovative method that proves real change happened.

Four products, all built on a mix of methods suited to what's being measured. This page covers what that looks like in practice, and where AI fits into it.

The question every number here is answering

A login count says someone showed up. A satisfaction score says they didn't mind it. Neither says the programme changed anything for them.

What gets measured here is the harder thing underneath both: whether an innovation, AI-driven or not, serves the human goal it was built for. What that goal is changes with what's being tested. In Rowad-Tech, a hi-tech mindset programme for Arab-society students, it's whether a student can genuinely picture a hi-tech career for themselves and start acting on it: choosing advanced maths, joining a tech club, the kind of concrete step a satisfaction score never captures.

Sporting organisations are, more often than not, nonprofits too: resource-constrained, and increasingly looking to technology and AI to run more efficiently, same as an education nonprofit does.

The same nonprofit world also includes programmes that aren't about technology at all. A coexistence football programme like Tzav Pius or The Equalizer is one: the question there is whether shared training time changes how a Jewish teenager and an Arab teenager treat each other afterwards, off the pitch as much as on it.

On the sports-tech side specifically, the same question is what a training app, a wearable, or a performance platform is meant to do: change how an athlete trains or a coach decides, evidenced in real behaviour once there's a live rollout to measure.

Adoption is the easy number to report. This is the one that tells you if it worked. For a funding scheme, it's the second half of what a real audit checks: proof the call delivered what it promised, alongside the paper trail for where the money went.

Why AI matters here

AI is moving fast, and it matters. It's reshaping classrooms, training grounds, and how organisations of every kind get work done. Treating it as a fad or a checkbox item misses what's happening.

That's exactly why the question underneath it matters more, not less. AI isn't evaluated for its own sake. It's evaluated, alongside whatever other technology a programme runs on, for whether it moved something real: how a student learns, how a teacher teaches, how an athlete trains. Project 720 is a good example: personalised-learning technology across four different platforms, only some of it AI in the narrow sense most people mean by the word. What gets tested is the technology's actual effect, not whether it's labelled AI.

AI does a second job in this work too: validating and cross-checking instruments against each other, and coding open-ended answers, hundreds of them, in days instead of the weeks a manual pass takes.

The toolbox behind every number

The work is research-based, drawing on established methodology in both education and sport. On the education side, that means matched-cohort, pre/post designs built on validated instruments, Dweck's mindset scales and Hattie and Timperley's work on feedback among them, adapted rather than used off the shelf. On the sport side, it's grounded in the STRN Sports Technology Quality Framework. No single instrument, however well validated, tells the whole story on its own.

Self-report is in there: asking someone directly how they feel, what they believe, what they'd do. So are simulations, behavioural tasks that show what someone actually does under a specific condition, which often tell a different story than the self-report next to it. Where nothing existing fits, new instruments get built: a diary a participant fills in over weeks, a usage or login log, a piece of software that collects behavioural data quietly in the background while a programme runs. Some of this happens even earlier, inside product design: advising a startup on what user data its own product should be collecting, and how, before there's anything to evaluate at all.

None of this means more questionnaires. It means fewer, each one earning its place: a rushed teacher or a stretched nonprofit team won't finish a long one, and a long one measures worse anyway. Some of the custom-built and background tools above exist specifically to take that burden off the people being measured.

Mixed methods, pushed further than the term usually implies: qualitative and quantitative, self-report and simulation and observation, off-the-shelf and custom-built, combined and checked against each other, because that's what surfaces real behaviour instead of a single, flattering account of it.

Four points, as many conversations as it takes

Every engagement runs on four points, each one worked through in as many conversations as it takes.

The first is about understanding each other: what the programme is trying to do, who it's for, what's already being collected, and whether we're a good fit to work together. The second sets the scope and the timeline: what gets measured, over what period. The third builds the model: sitting down together to agree what real change would look like here, and what the concrete goals are. The fourth is delivery: the findings, the presentation, the conclusions and the report, walked through together.

Between those four points, the actual work happens: data collection, analysis, checking what holds up. Every engagement has a clear finish line, a defined deliverable handed over on an agreed date.

Book a call to see how this would be built around your programme, and whether the evidence backs it up.

Book a call