Where the evidence lives
What actually changed, and how we know it held up
Four evaluations below, each with the finding that held up and the one that didn't survive a closer look. The second kind is the reason to trust the first.
Furthest behind, moving fastest
An earlier read of this data showed a large shift among Arabic-speaking students. A cell-by-cell check caught the cause: an early coding pass had scored their end-of-year answers as a near-uniform "5" on most items, not their actual responses. Corrected, that finding didn't survive. What's below is what did.
Students who started furthest behind gained the most. The effect holds on all six measures tracked for the population with a valid comparison group, from p=.001 to p=.032, and survives every check run against it, including controlling for which school a student was in. In schools where teachers grew stronger as mentors over the year, students grew stronger too (rho≈0.63), strongest specifically on how supported students felt (rho≈0.70). The one place every instrument agrees: self-awareness and reflection are the weak link. Students' own ratings of it fell. Their answers to open scenarios showed thin planning from the start. Teachers turned a reflective eye on their own teaching in only 3% of responses. Classroom observation found no built-in reflection for four in ten students. Four separate measurements, one shared gap. Exactly what the programme should work on next.
Almost 2,000 seventh-graders across 28 classes and four different ed-tech platforms, one school year. The matched analysis behind these numbers covers 704 students and 65 teachers across 21 schools, tracked from the start of the year to the end.
The self-report had it backwards
On self-report scales, Arabic-speaking and Bedouin students rated themselves markedly higher than Hebrew-speaking students on resilience and future orientation. On the one task that's hard to answer flatteringly (where a student picks their own puzzle difficulty instead of rating themselves), that gap nearly disappeared, and reversed slightly. Most of the population difference is how each group uses a rating scale, not a real gap in resilience or mindset. It's exactly the kind of finding a self-report-only evaluation would have missed.
“Over the past two years, Iddo has led the measurement and evaluation of Mental Axis within the Ministry of Education's hi-tech programme. After getting a clear read on where the programme stood, and working inside real constraints, he built a serious measurement and evaluation plan for a genuinely hard area: students' mental frameworks, and understanding what actually drives their motivation to succeed on the hi-tech track. The insights he's gathered have helped us, and keep helping us, build more precise responses, and aim them at the students who actually need them, and at the teaching staff too.”
Yedidiya Green, Mental Axis Lead, Israeli Ministry of Education Hi-Tech ProgrammeInvolvement beat exposure
663 students at the end of the programme, measured against a comparison group in Jewish-society schools. The clearest gap: how concretely students could picture their own future (4.20 among high-schoolers and 4.15 among middle-schoolers, against 3.39 in the comparison group). Involvement beat exposure: being in an actual hi-tech activity, like a tech club, moved a student's intent to take advanced maths roughly three times as much as simply knowing someone in the field (d=0.45 vs d=0.16). One more finding worth naming plainly: girls outscored boys on almost every measure, yet in high school, 15% of girls ruled out advanced maths outright, twice the rate for boys. Whatever's driving that sits outside the mindset measures this instrument tracks. This is a single end-of-programme snapshot, not a before-and-after, so it describes where these students stood, not how far the programme moved them.
Motivation moved. Execution didn't.
Measured over one term, teachers left significantly more motivated to lead innovation in their own teaching and sharper at spotting real problems worth solving (p=.007 and p=.001). Their ability to actually analyse an opportunity and carry an idea through stayed flat (p=.613 and p=.112). The programme built the spark, not yet the follow-through. Precisely where its next version should focus.
“Iddo joined our EdTech accelerator team as an external advisor, leading the programme's evaluation and measurement process. He was present across both the micro and macro processes of the programme, and supported us through other work on bringing new technologies into use. Iddo has since joined further innovative projects, including AI-assisted measurement and evaluation of school education programmes, teachers and students, taken on similar work with a range of different populations, and consistently brings real, forward-moving value, whether as a team member or as project lead.”
Adi Dagan, Director of Technology Partnerships, Israeli Ministry of Education Innovation and Technology AdministrationWhat stands behind the sport-side work
The four case studies above are all education programmes. On the sport side, ThinkerBox's evaluation credibility rests on method and access: a working link to STRN's framework, a seat on the EU's AI-Skills-for-Sports working group, and assessment work inside two coexistence programmes reaching roughly 10,000 participants.
That's what a sports-tech client can check today, real method and real access, before commissioning the first delivered case study.
Want to see how this would look for your programme? Book a free call (just the evidence and an honest read on fit).
Book a call