I made this for kids around 10 to 14. A robot called Errol solves a math problem step by step and one of the steps is wrong. The kid has to find it and say what's wrong with it. Sometimes nothing is wrong, so just saying "there's a mistake" every time doesn't work. The user needs to enter an explanation if she finds an error to gain more XP; speed matters also for more points. There is no direct interaction or chatting with an LLM. Lathoa's harness is stable and has many evaluation steps to catch inconsistencies and prompt injections.
You can play one on the homepage without signing up.
The part that surprised me: it's hard to get an LLM to be wrong on purpose. Half the time it gives you the right answer and calls it wrong, or a "mistake" that's actually correct. So every case gets checked before a kid sees it. Where it can, a plain arithmetic check redoes the math exactly. A second model also solves the problem without seeing Errol's work. If anything disagrees, the case is thrown away.
The weak spot is that the second model can make the same mistake as the first. The arithmetic check is there for that, but it only works on English cases so far. German and Greek write decimals with a comma and I haven't got the parsing right yet.
What I'd really like to know: does finding someone else's mistake teach anything that solving the problem yourself doesn't? I'm not sure, and I'd like to hear from people who teach.
eichin 21 minutes ago
Reminds me of https://en.wikipedia.org/wiki/QAMA_Calculator - a calculator that won't reveal the precise answer until you supply an estimate (physical hardware in 2014, I think the current version is an app.) The website makes claims about the pedagogic value, but I didn't see any actual citations at https://qama.world/the-science/ just "everyone knows" kind of things... but if they had research supporting it that would probably inform your approach?
lemming an hour ago
This is an interesting idea, but I need more than one sample to get a feel for it. I did the example on the homepage, and then one more that it linked to, and both had extremely basic errors my kid would have no problem with (currently doing Math Academy prealgebra, approx year 9/8th grade level). I'd suggest allowing people to try more than one very basic error.
its-summertime 2 hours ago
I think speed would be irrelevant, or perhaps would discourage deep engagement with the problems.
I think presenting it as AI can be wrong is not as valuable as presenting it as developing the ability to question what one is told. (e.g. Verizon Math)
thanouil1411 2 hours ago
indeed I agree with you however studies have shown that kids have access to interfaces to chat with an LLM from a young age, and the traction this technology is taking is increasing. Indeed verizon math seems more whole when it comes to topic variety.
nanis 2 hours ago
> if she finds an error to gain more XP
What does this mean?
thanouil1411 2 hours ago
while solving a lathoa case, if you spot the error you get back an amount of points. if you score also the explanation exhibiting your understanding you gain more points.
satisfice 3 hours ago
It helps us practice critical thinking. What it doesn’t do is help us practice knowing WHEN to use critical thinking.
thanouil1411 3 hours ago
interesting perspective...when and how much also is a skill someone should have.