Does your model actually render size? Measure fit accuracy against real-world ground truth β real people, real garments, real physics. Never simulated.
FID, SSIM, and eyeballing output grids don't capture fit fidelity. A model that renders size M and size L as nearly identical garments still passes review under today's metrics β because none of them ask whether the fit actually changed.
In our own delta-matched synthetic samples, visible drape does not track garment ease (r = β0.10) β while real S/M/L captures track it monotonically. See how much fit signal survives into synthetic try-on images.
Grading a model against another model's mistakes β the hallucination just gets inherited
Physically captured, never simulated β a score you can trust
Fittings Labs scores virtual try-on fit fidelity against real-world ground truth β never simulated β so a passing score means the model actually gets fit right.
Every test case is scored against physically captured, real-world ground truth.
High-fidelity photos (Front, Side, Back) of the exact same individual wearing sequential sizes (S, M, L) β the ground truth a model's output is scored against for each transition.
Every reference is a physical capture of a real person in a real garment β tape-measured body and flat garment specs, zero simulation in the loop.
Specialized focus on high-difficulty categories like denim. We intentionally capture authentic "ill-fits" (e.g., fabric strain, jeans failing to button) to test models against real-world boundary thresholds.
Submit your model's outputs against the held-out benchmark set and receive a fit-fidelity score broken down by category and size transition.
Get the held-out benchmark set of size-transition test cases, or submit your model's outputs directly against our evaluation endpoint.
Outputs are scored against physically captured reference images and tape-measured specs β never a simulated reference.
Receive a category-level fit-fidelity score report, broken down by garment category and size transition.
We eliminate legal and privacy risks. Every single pixel in our repository is collected under strict ethical standards, built specifically for commercial enterprise AI deployment.
Ground truth is captured from vetted contributors via a controlled, multi-size photo protocol. Every capture is fully consented and anonymized before it enters the benchmark.
Every global contributor signs a comprehensive, explicit model release and data waiver. Fittings Labs holds 100% of the commercial intellectual property rights, guaranteeing zero copyright friction for your model training.
We strictly enforce privacy boundaries. All facial data and personally identifiable information (PII) are completely stripped and cropped via our automated backend pipeline before data ingestion. We sell 100% anonymized body-to-garment matrices.
Our data acquisition pipeline is fully compliant with global data privacy regulations, including GDPR (Europe), CCPA (California), and the latest AI governance acts.
All ground-truth assets are processed and securely stored in encrypted, access-controlled enterprise repository environments, preventing any data leaks or unauthorized access.
The benchmark set is held out, versioned, and never included in any training dataset β so a score reflects generalization, not memorization.
No reference image or label in the benchmark is synthetically generated. Every reference is a real physical capture β so the score can't inherit a simulator's mistakes.
Research, pipeline notes, and perspectives on building real-world training data for virtual try-on.

We checked whether FIT's bodyβgarment measurements survive into the rendered images. In our real S/M/L captures, visible drape tracks ease monotonically; across 75 delta-matched synthetic samples, it doesn't (r = β0.10) β and the gap isn't confined to tight garments.

State-of-the-art try-on models render nearly identical images for adjacent sizes of the same garment. We break down why the training data makes that inevitable, and what real, same-person, same-SKU, adjacent-size capture unlocks instead.
The same real-world, multi-size capture pipeline that powers the benchmark is available as training data licensing β paired sequential sizes, tape-measured ground-truth specs, and boundary-condition edge cases, ready to close the gap the benchmark just showed you.
Get immediate access to a curated teaser batch of our benchmark set β real-world, tape-measured, multi-size ground truth. Validate the 1-pixel accuracy of our normalization before you license the full benchmark or training data.