Table 3.
Summary of external test accuracy in endoscopist–artificial intelligence (AI) interactions.
|
|
External test accuracy, % (95% CI) | P values | |||||
| Rater | First test: endoscopist alone | Second test: endoscopist with faulty AI | Third test: endoscopist with best performance AI | First test vs second test | First test vs third test | ||
| Expert endoscopist |
|
|
|
|
|
||
|
|
All answers | 90.8 (86.8-94.8), (187/206) | 90.3 (86.3-94.3), (186/206) | 90.3 (86.3-94.3), (186/206) | .87 | .87 | |
|
|
Confident answer | 93.8 (91.7-95.9), (181/193) | 90.6 (86.6-94.6), (182/201) | 90.6 (86.6-94.6), (184/203) | .23 | .24 | |
|
|
Unconfident answer | 46.2 (19.1-73.3), (6/13) | 80.0 (44.9-99.9), (4/5) | 66.7 (13.4-99.9), (2/3) | .20 | .52 | |
| Endoscopy trainee |
|
|
|
|
|
||
|
|
All answers | 68.4 (62.1-74.7), (141/206) | 66.5 (60.1-72.9), (137/206) | 81.6 (76.3-86.9), (168/206) | .67 | .002 | |
|
|
Confident answer | 80.6 (74.4-86.8), (125/155) | 71.1 (63.2-79.0), (91/128) | 92.2 (87.8-96.6), (130/141) | .06 | .004 | |
|
|
Unconfident answer | 31.4 (18.7-44.1), (16/51) | 59.0 (48.1-69.9), (46/78) | 58.5 (46.5-70.5), (38/65) | .002 | .004 | |
| General physician |
|
|
|
|
|
||
|
|
All answers | 65.0 (58.5-71.5), (134/206) | 62.1 (55.5-68.7), (128/206) | 64.1 (57.5-70.7), (132/206) | .38 | .84 | |
|
|
Confident answer | 77.4 (70.0-84.8), (96/124) | 71.6 (64.5-78.7), (111/155) | 74.2 (67.5-80.9), (121/163) | .27 | .53 | |
|
|
Unconfident answer | 46.3 (35.5-57.1), (38/82) | 33.3 (20.4-46.2), (17/51) | 25.6 (12.6-38.6), (11/43) | .14 | .02 | |