Spending this week going deeper on the wildlife AI system instead of citing the topline number again. The part I keep coming back to is the confidence threshold: every image gets a classification and a score, and anything below the threshold gets routed to a person instead of logged as fact.
Setting that threshold is the part nobody talks about. Get it wrong in one direction and you're back to reviewing almost everything by hand, which defeats the point of building the system at all. Nobody hands you a formula for the right number. It comes down to a judgment call: how much risk of a missed or wrong classification is acceptable before a human needs to look.
I don't think there's a universally correct number here. What matters is picking one that's honest about how wrong the model still gets things. Curious how other people building anything with a human-review step think about where to set that line.