Table 4.
Overall rank of generative models for the use cases in the Model recommendation phase
| Use Case | Dataset | Overall model rank | |||||
|---|---|---|---|---|---|---|---|
| 1 | 2 | 3 | 4 | 5 | 6 | ||
| Education | UW | EMR-WGAN (5.6) | medBGAN (8.0) | Baseline (8.7) | medGAN (10.5) | WGAN (10.9) | DPGAN (13.2) |
| VUMC | WGAN (6.5) | EMR-WGAN (7.5) | medBGAN (10.1) | Baseline (10.7) | DPGAN (11.0) | medGAN (11.2) | |
| Medical AI Development | UW | EMR-WGAN (6.3) | WGAN (8.6) | medBGAN (8.7) | medGAN (9.9) | Baseline (11.2) | DPGAN (12.3) |
| VUMC | WGAN (8.1) | DPGAN (8.5) | EMR-WGAN (8.6) | medGAN (8.8) | medBGAN (10.8) | Baseline (12.0) | |
| System Development | UW | Baseline (7.9) | medBGAN (9.1) | EMR-WGAN (9.5) | medGAN (9.6) | WGAN (9.8) | DPGAN (11.1) |
| VUMC | Baseline (8.7) | medBGAN (9.3) | medGAN (9.3) | WGAN (9.4) | DPGAN (9.5) | EMR-WGAN (10.8) | |
Model ranks were based on the benchmarking framework scores (in parenthesis).
The fact that DPGAN and EMR-WGAN have the same score for the VUMC dataset in the Medical AI Development use case is due to precision loss instead of an actual tie.