How NIFFLER has done on a fixed set of 500 made-up orders. See Getting Started for where this data comes from.
These came from real Razorpay test payments people made on this site. They are kept out of the numbers above on purpose, so the measured batch stays fixed and comparable.
Honest answer: it has been run on 200 orders, not lakhs, so anything else would be a guess. But the shape of the problem is clear enough to say where it would bend.
Finding the failed orders is plain code over the database. Checking the rules is plain code too. Both are fast at any size. Only the diagnosis needs a model, and only for orders that are worth chasing in the first place.
Each case costs several AI calls, and cases are handled one after another today. That was a choice for correctness, not speed. The real ceiling I hit was the free tier's daily token limit, not the code.
The rules comparison shows plain rules reach the same answer 73% of the time. Send the easy majority through rules and spend the model only on the rest, and the AI cost drops by roughly three quarters.
What is genuinely missing for that scale: cases are processed one at a time rather than in parallel, the run happens inside the web request instead of a proper queue, and none of it has been load tested. Those are real gaps, not ones I am going to pretend are solved.