Large language models (LLMs) like ChatGPT are widely used for public health information retrieval. Prior urological studies have mainly validated ChatGPT performance based on AUA and EAU guidelines, while evidence aligned with CUA guidelines remains scarce. This study by Wyatt MacNevin’s team from Dalhousie University evaluated the appropriateness and reliability of ChatGPT-4.0 responses to common patient-oriented urological questions using CUA guidelines as the benchmark.
The research team selected 10 common urological questions covering nephrolithiasis, prostate cancer, benign prostatic hyperplasia (BPH), erectile dysfunction (ED), overactive bladder (OAB), urinary tract infections (UTI), andrology, hypogonadism, pediatric urology, and kidney cancer. Each question was phrased in plain language and submitted to ChatGPT‑4.0 (March 2025 version) in three independent sessions, generating 30 responses. Three evaluators—two practicing Canadian‑licensed urologists and one urology resident—graded each response on a four‑point Likert scale (0 = Completely incorrect, 1 = Some correct and some incorrect, 2 = Correct but inadequate, and 3 = Comprehensive), with an “appropriate” response defined as a score ≥ 2.00.
Of the 30 responses generated by ChatGPT‑4.0, only 40% (12/30) met the appropriateness threshold, with an overall mean score of 1.64 ± 0.85 falling between "some correct and some incorrect" and "correct but inadequate." When stratified by question difficulty, easy questions scored significantly higher than medium‑difficulty questions (1.87 ± 1.01 vs. 1.31 ± 0.47, p < 0.05), confirming that ChatGPT performs best on fact‑based, definitive queries. In domain‑specific analyses, prostate cancer, ED, andrology, and kidney cancer received perfect median scores of 3.00, likely due to the abundance of standardized online information on these topics, whereas UTI, OAB, nephrolithiasis, and hypogonadism scored lower—even when rated as “easy”—possibly reflecting inconsistent online content or discordance between CUA and other international guidelines. The mean variance across repeated questions was 0.27, indicating strong model consistency and reliability.
ChatGPT has gained popularity as a healthcare information source, yet most existing studies have referenced American or European guidelines. This is one of the few studies to assess its performance against Canadian standards. The 40% appropriateness rate signals potential, but also highlights that current outputs are insufficient for independent patient use.
Urologists should be aware that patients using ChatGPT may receive incomplete or inaccurate information. Proactive patient education about the model’s limitations is essential. The authors recommend mandatory disclaimers on medical responses and future integration of authoritative guidelines into LLM training—alongside prospective studies evaluating real‑world clinical impact.
ChatGPT produced appropriate answers for 40% of general urology questions when benchmarked against CUA guidelines. While promising, further refinement and rigorous validation are required before widespread clinical adoption can be recommended.
The work titled “Assessing the utility of a natural language processing model in answering common urological questions” was published in UroPrecision (published on August 20, 2025).
DOI:10.1002/uro2.70028