Publication: The Impact of Automated Tools on Clinical Decision-Making in Gestational Trophoblastic Disease
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Project 1 abstract: Objective: The hook effect is a limitation of assay-based laboratory tests utilized for hCG testing in GTD. A result of this effect is reporting of falsely low or inappropriately normal hCG values. However, apart from case reports and case series, no prior literature explores the possible clinical features associated with this phenomenon owing to its rarity. This study aims to bridge this gap in literature by investigating the association between the hook effect and clinical symptoms reported and using predictive models to create a clinical risk score.
Methods: Univariable and multivariable Firth modification of logistic regression were performed to determine clinical features associated with hook effect on hCG testing in a retrospective observational study of patients with GTD from the Rio de Janeiro Trophoblastic Disease Centre over a 10-year period of 2014 to 2024. Youden’s index was utilized to determine ideal cut points for hCG values, along with reporting sensitivity, specificity, PPV and NPV. Clinical features identified on association testing and hCG values identified on cut point estimation are utilized as predictors for the creation of a clinical risk score.
Results: Our study population consisted of 1,019 patients, with 34 (3.3%) reporting hook effect. The variables associated with an increased odds of hook effect include, vaginal bleeding (OR: 6.40, 95% CI: 2.75-17.96) increased uterine size for gestational age (OR=2.37, 95% CI: 1.20-4.71), history of previous mole (OR: 19.80, 95% CI: 4.42-77.90), theca-lutein cyst (OR: 7.62, 95% CI: 3.38-16.17), and pre-eclampsia. (OR: 18.84, 95% CI: 7.28-46.20). On cut point estimation, the clinically meaningful threshold for undiluted hCG values was determined at 1500 IU/L, with sensitivity of 91%, and specificity of 99%. A clinical risk score was created based on inferential testing to determine the risk of hook effect in patients with GTD. The variables in the predictive model included extremes of age, vaginal bleeding, history of previous mole and hCG ≤ 1500 IU/L. The model had high internal validity with an optimism corrected AUC (95% CI) of 0.99 (0.98-0.99) on bootstrapping with 1000 iterations.
Conclusions: Clinical features such as vaginal bleeding, increased uterine size for gestational age, history of previous mole, theca lutein cysts, and preeclampsia were associated with the presence of hook effect on hCG testing. Hook effect once promptly identified should necessitate dilution of samples to ascertain true hCG values to facilitate prompt treatment.
Project 2 abstract: Objective: Access to timely and appropriate medical information for the management of GTD is a challenge for physicians and patients alike due to the rarity of the disease and lack of trained physicians. Open source and freely available LLM chatbot ChatGPT can serve as a potential tool to supplement delivery of medical care in the absence of a GTD expert. In this paper, we evaluate whether ChatGPT V4.0 is able to provide accurate and complete responses to clinical, diagnostic and therapeutic clinical vignettes in GTD.
Methods: A cross-sectional survey-based study was conducted including an international multi-institutional panel of GTD experts inquiring about the agreement, accuracy and completeness levels on Likert scales for answers generated by ChatGPT v4.0 to ten clinical vignettes. To assess the proficiency of ChatGPT, a predetermined threshold of 75% was utilized based on Delphi based consensus studies. Globally and for per question exact binomial proportions with exact Clopper-Pearson confidence intervals were reported to assess whether 75% of experts responded in the positive categories across all 3 Likert scales. Krippendorf’s alpha across the three attributes was reported to compare observed agreement among raters along with Cronbach’s alpha for internal consistency of survey items. A GLMM model was fit to assess if academic rank of the experts was a predictor for their choice in the positive categories of Likert scales.
Results: We received 12 (40%) complete survey responses out of the 30 invitations sent out to GTD experts. The majority of experts, 50% (6/12) were appointed to the academic rank of full professors, and associate professors and practiced medicine in academic institutes (11/12). Most experts practiced in continents other than North America (9/12). On global binomial proportions ChatGPT failed to meet out threshold with 0.6 (0.26 - 0.87). It performed comparatively well on questions regarding molar pregnancies. Inter-rater reliability assessed via the Krippendorf’s α indicated low agreement across experts; agreement (α =.275), accuracy (α =.288) and completeness (α =.269), However, excellent homogeneity demonstrated across all attributes via Cronbach’s α; agreement (α = .88), accuracy (α = .84), and completeness (α = .79). Academic rank was not a significant predictor of Likert choices in positive categories for all attributes on global GLMM and subgroup analysis per Likert scale.
Conclusions: ChatGPT did not meet our predetermined performance thresholds at 75% and hence was determined as a poor supplement for physicians and patients seeking medical information on queries in GTD.