Key takeaways
- "Accuracy" in size recommendation is meaningless without a definition: the honest measure is business outcomes under A/B conditions, not a vendor's self-scored hit rate.
- The 2026 benchmark that matters is head-to-head: ETAM saw +5% revenue and −16% returns with Kleep; ba&sh saw +88% add-to-cart and +20% usage versus True Fit; Jules ran Kleep against Fitle and Kleep won.
- Cohort-based tools infer your size from people like you; measurement-based tools match your body to the actual garment. The gap shows up in borderline cases, which are most cases.
- A proper evaluation is an A/B test on your own catalogue and traffic, run over a full return cycle, with revenue and returns as the verdict.
- The right vendor questions are methodological: what data feeds the recommendation, how are new products handled, and will you accept a head-to-head test?
How accurate are size recommendations, really?
Every vendor in this category claims high accuracy. The claims cannot all be equally meaningful, because they rarely measure the same thing. One vendor counts a recommendation as accurate if the shopper kept the item; another if the shopper accepted the suggestion; another scores against a panel. Same word, three different tests.
A 2026 reality check starts by discarding self-declared accuracy altogether. The question a retailer actually needs answered is narrower and harder: on my catalogue, with my shoppers, does this tool produce more kept revenue and fewer returns than the alternative? That is measurable, and it has been measured.
What accuracy should actually mean
A size recommendation is correct when the shopper buys the recommended size, keeps it, and would not have been better served by another size. Everything upstream, model confidence scores, fit-survey agreement, is proxy. The kept order is the ground truth.
That definition immediately exposes the weakness of averages. A tool can be right on the easy middle of the bell curve and wrong exactly where shoppers need help: between two sizes, at the ends of the range, on garments that run large or small. Accuracy on the hard cases is the product; accuracy on the easy ones is decoration.
Cohort inference versus real measurements
Most incumbent tools lean on purchase-history cohorts: shoppers like you kept this size, so you probably will too. It scales cheaply, but it inherits every bias in the historical data and says nothing about the specific garment in front of the shopper.
Measurement-based systems like Kleep build a body profile, via a 30-second questionnaire or a two-photo scan, and match it against each garment's real spec. The difference is invisible on an average product and decisive on a borderline one, which is precisely where returns are born.
The A/B methodology that produces a real answer
- Split traffic randomly between tools (or tool versus control) on the same catalogue and period.
- Run through at least one full return cycle; conversion without keep-rate is half a verdict.
- Judge on revenue per session, return rate, and adoption, a tool shoppers ignore cannot be accurate in practice.
- Segment results by category and by borderline shoppers; the blended average flatters weak tools.
What the 2026 head-to-heads showed
The category now has genuine competitive evidence. ETAM measured Kleep at +5% revenue with a 16% reduction in returns. ba&sh evaluated Kleep against True Fit and saw +88% add-to-cart and +20% higher usage of the tool. Jules ran a direct A/B of Kleep against Fitle, and Kleep won the test.
Three retailers, three methodologies, one pattern: when accuracy is defined as outcomes under controlled conditions, the measurement-based approach carried the result. Usage matters as much as the algorithm, a recommendation shoppers actually invoke and trust is the one that changes the P&L.
Questions to ask any sizing vendor
- How exactly do you define and measure accuracy, and can I see the methodology?
- What powers the recommendation for a brand-new product with no purchase history?
- How do recommendations differ across garments within the same brand?
- Will you run a head-to-head A/B against my current tool, judged on revenue and returns?
- What adoption rate do you see, and how do you measure recommendation trust?
How accurate are AI size recommendations in 2026?
Accurate enough to move revenue and returns when measured properly. In controlled tests, ETAM saw +5% revenue and −16% returns with Kleep, and ba&sh recorded +88% add-to-cart versus True Fit. The honest measure is A/B outcomes, not vendor-declared hit rates.
What is the best way to evaluate a size recommendation vendor?
Run a head-to-head A/B on your own catalogue and traffic over a full return cycle, judged on revenue per session, return rate, and tool adoption. Any vendor confident in their accuracy should accept those terms; hesitation is itself a data point.
What is the difference between cohort-based and measurement-based sizing?
Cohort tools infer your size from what similar shoppers kept; measurement-based tools match your actual body profile to each garment's real measurements. Cohort inference struggles on new products and borderline bodies, exactly where fit help matters most.
Why does usage rate matter for sizing accuracy?
A recommendation only prevents a return if shoppers invoke and trust it. ba&sh saw 20% higher usage with Kleep versus True Fit, which compounds the accuracy advantage: a slightly better answer used far more often produces a much better business outcome.
Conclusion
The accuracy debate in size recommendation is settled the way all such debates should be: in production, under A/B conditions, on real revenue and real returns. ETAM, ba&sh, and Jules ran those tests. The lesson for any retailer is to demand the same standard, define accuracy as outcomes, insist on a head-to-head, and let your own traffic deliver the verdict.










