Introduction
This study aimed to evaluate the accuracy, reliability, and comprehensibility of information about surgically-assisted rapid palatal expansion provided by language models based on artificial intelligence (AI).
Methods
A cross-sectional content analysis was conducted on the responses to surgically-assisted rapid palatal expansion-related questions by ChatGPT-4 (OpenAI LLC, San Francisco, Calif), Gemini (Alphabet Inc, Mountain View, Calif), and Copilot (Microsoft, Redmond, Wash). In total, 115 questions (categorized into 11 domains) were created by 3 orthodontists and 1 oral and maxillofacial surgeon. The accuracy of the answers generated by the AI language models was independently evaluated by the same experts via a 5-point Likert scale. To test the relationships among categorical variables, when the sample size assumption was met, the Pearson chi-square test was used. However, when the sample size assumption was not met, Fisherās exact test was applied. Analyses were performed in SPSS (version 27; IBM, Armonk, NY).
Results
The responses of the AI types presented a general homogeneous distribution, with no statistically significant difference between the types of AI and the types of responses ( P >0.05). Although there were no significant differences, ChatGPT-4 had the highest objectively true rate. In contrast, Gemini produced answers with more balanced accuracy, whereas Copilot had the highest number of false answers.
Conclusions
These findings reveal that the accuracy of AI-supported language models in providing medical information may vary according to subject matter.
Highlights
-
ā¢
We examined homogeneous responses of AI types about SARPE.
-
ā¢
AI types had a moderate level of information about surgically-assisted rapid palatal expansion.
-
ā¢
The accuracy of AI-supported language models may vary according to the subject matter.
-
ā¢
ChatGPT had the highest objectively true responses, although not statistically significant.
-
ā¢
Copilot had the highest false responses, although not statistically significant.
Transversal malocclusion of the maxilla is a common type of malocclusion seen in 8%-22% of the population, , and it can have significant effects on jaw and dental esthetics and functionality. The treatment approach for patients with maxillary transverse deficiency varies depending on skeletal maturity. In subjects with skeletal immaturity, it is possible to separate the midpalatal suture by applying orthopedic forces. In contrast, in subjects with skeletal maturity, the resistance increases because of midpalatal suture fusion, whereas the effectiveness of the expansion procedure decreases. This contributes to undesirable effects of the applied orthopedic forces, resulting in alveolar bending, tilting of teeth, and minimal maxillary expansion. These complications can be minimized by surgical release of resistant bony structures. Therefore, surgically-assisted rapid palatal expansion (SARPE) is indicated in subjects with skeletally maturity.
The SARPE procedure is based on the principles of distraction osteogenesis and aims to enlarge the maxilla by performing osteotomies on resistant structures. During this surgical intervention, osteotomies are applied to the aperture piriformis, processus zygomaticus, processus pterygoideus, and sutura palatina mediana to reduce the resistance of the maxilla to expansion forces. ,, The distraction force is transferred to the maxilla through tooth-based, bone-based, or hybrid devices, and the desired expansion is achieved. ,, Numerous studies have shown that SARPE increases maxillary width both skeletally and dentally. ,,, As a safe and simple method of correcting existing malocclusion by eliminating maxillary constriction, , the SARPE procedure increases the predictability and success of expansion while reducing side effects (eg, a possible reduction in alveolar bone thickness and height, bone dehiscence, and gingival recession) and orthodontic recurrence. ,
Pain, bleeding, infection, nerve damage, apical root resorption, and tooth discoloration and devitalization have been reported as complications of SARPE. Inadequate or asymmetrical expansion and infection of the maxillary sinus are rare, whereas expansion of the alar floor, tinnitus, lacrimation, life-threatening epistaxis, and skull base fracture are extremely rare. ,,,
Although the SARPE procedure has been proposed as a safe and simple method for patients with maxillary transverse insufficiency, surgical intervention can sometimes pose difficulties for patient acceptance. Patients may be concerned about the procedure, process, and results. In these circumstances, it is very important to provide the necessary and correct information to patients. If adequate information about the treatment process and potential complications is not provided, patients may turn to various sources and alternative information to relieve their concerns.
Artificial intelligence (AI)-based natural language processing and natural language generation (NLG) models have recently become widely used tools to answer medical questions. NLG models based on AI are models that can generate human-like text, answer questions, and perform other language-related tasks with high accuracy. These models are capable of generating meaningful and coherent responses from large datasets. Patients can use these language models to access health information, medical advice, counseling, and answer their questions. Generative pretrained transformer 4 (GPT-4), the fourth-generation language model of the GPT series developed by OpenAI (OpenAI LLC, San Francisco, Calif), is capable of generating detailed answers from a large amount of data. Similarly, Gemini (Alphabet, Inc, Mountain View, Calif), developed by Google (Alphabet, Inc), and Copilot (Microsoft, Redmond, Wash), powered by Microsoftās AI innovation and using OpenAIās GPT-4 language model, were developed via large language models. These chatbots are trained using large amounts of data and deep learning algorithms, making them proficient at generating meaningful responses and predicting text by identifying relationships between words. Although these modeling-based platforms can provide fast and seemingly accurate information via advanced language models to answer health-related questions via a variety of sources, including medical databases, research studies, and health Web sites, there is some uncertainty about the accuracy and reliability of the information provided. ,, This study aimed to comparatively analyze the comprehensibility, accuracy, and comprehensiveness of responses to questions about SARPE asked by patients to 3 different chatbots.
Furthermore, this study aimed to evaluate the quality of information provided to patients by AI chatbots, including ChatGPT-4, Copilot, and Gemini, and to examine their ability to provide patients with access to accurate and comprehensible information. In addition, we evaluated the medical accuracy of each robotās responses and their effectiveness in terms of patient education. These findings can help determine how AI-based chatbots can be used more efficiently in health care.
Our hypothesis, in light of previous studies, was that the information provided by AI chatbots on SARPE would be accurate and that ChatGPT-4 would be the most accurate information provider, even if there would be differences among different chatbots.
Material and methods
In this study, content analysis was performed on the responses generated by ChatGPT-4 (OpenAI LLC), Gemini (Alphabet Inc), and Copilot (Microsoft) to questions related to SARPE. The questions were designed to cover all topics in which patients or laypeople may be asked and curious about SARPE. First, a question set consisting of 11 domains and 115 questions about SARPE was created by 3 authors (S.H., F.A.K.T., and O.O.Z.) ( Supplementary Table ). The question set was revised by a fourth author (E.C.O). AI chatbots were asked questions on January 6-8, 2025. Each question was asked individually and then deleted to avoid the learning process of the AI chatbots. The answer to each question was copied into a separate Word document and deleted from the AI chatbotsā memory. The evaluations were then made on separate Word documents created for each AI chatbot.
A customized Excel spreadsheet was created by 1 author (S.H.) to collect responses. For evaluation, the accuracy of answers provided by the AI robots was categorized into 5 categories: false, nonfacts, minimal facts, selected facts, and objectively true ( Table I ). A meeting was held to standardize the categorization and establish a common understanding of the evaluation system. The collected answers were independently evaluated by 3 orthodontists (S.H., F.A.K.T., and E.C.O.) and 1 oral and maxillofacial surgeon (O.O.Z.). In the case of a conflict among evaluators, a final decision was made by the principal investigator (S.H.).
Table I
Definition of the different accuracy categories
| Category (score) | Definition |
|---|---|
| Objectively true | A claim that is based on scientific evidence and presents all relevant information, whether positive or negative |
| Selected facts | A claim that presents some true selected facts based on scientific evidence, but omits important information related to a product |
| Minimal facts | A claim that exaggerates the benefit of the product, with an overemphasis on the benefit supported by poor-quality scientific evidence |
| Nonfacts | A claim that presents an intangible characteristic. Often, these claims are in the form of product opinions or lifestyle claims, leaving clinicians/patients to misinterpret the opinion as an objective product evaluation |
| False | A claim that is objectively false, either because of a lack of evidence to support it or contradicts available evidence |
The results evaluated in the analysis are not fixed and unalterable ābasic factsā because the quality of the responses is based on personal judgment. To evaluate the quality of each response more objectively by examining the quality of each response from various perspectives by different evaluators, an analysis based on the principles of the crowd (or ensemble) score strategy was conducted.
Because the study evaluated only the responses obtained from AI models, institutional ethics committee approval was not needed.
Statistical analysis
Descriptive statistics (number and percentage) of the data are presented. For categorical variables, a Fisherās exact test was applied when the sample size assumption was not met. Analyses were performed via the SPSS software (version 27; IBM, Armonk, NY).
Results
The distribution of the answers given according to AI type was given, and Fisherās exact test was applied to examine the relationships among them. The results of the analysis revealed that the responses given by AI types generally presented a homogeneous distribution, and there was no statistically significant difference between the types of AI and response types ( P >0.05) ( Fig 1 ; Table II ).
Distribution of answers according to AI type.
Table II
Distribution of responses by AI types and their relationships
| Answers | ChatGPT | Gemini | Copilot | Test statistic | P value | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| n | % | % AI | n | % | % AI | n | % | % AI | |||
| False | 6 | 27.3 | 5.2 | 5 | 22.7 | 4.3 | 11 | 50.0 | 9.6 | 10.509 | 0.203 |
| Nonfacts | 2 | 40.0 | 1.7 | 2 | 40.0 | 1.7 | 1 | 20.0 | 0.9 | ||
| Minimal facts | 9 | 50.0 | 7.8 | 6 | 33.3 | 5.2 | 3 | 16.7 | 2.6 | ||
| Selected facts | 18 | 23.7 | 15.7 | 27 | 35.5 | 23.5 | 31 | 40.8 | 27.0 | ||
| Objectively true | 80 | 35.7 | 69.6 | 75 | 33.5 | 65.2 | 69 | 30.8 | 60.0 | ||
For the domains, the distribution of the answers given according to AI type was provided, and Fisherās exact tests were used to examine the relationships among them. On the basis of analysis, no statistically significant relationships were found between AI types and responses in any of the domains ( P >0.05) ( Table III ). Although the rates of objectively true responses were 69.6% for ChatGPT, 65.2% for Gemini, and 60.0% for Copilot, the differences among these rates were not statistically significant ( Table II ). Similarly, the rates of false responses were 4.3% for Gemini, 5.2% for ChatGPT, and 9.6% for Copilot, but the differences among these rates were not statistically significant ( Table III ).
Table III
Distribution of responses according to AI types for domains and the relationships among them
| Answers | ChatGPT | Gemini | Copilot | Test statistic | P value | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| n | % | % AI | n | % | % AI | n | % | % AI | |||
| Knowledge and information | 6.149 | 0.657 | |||||||||
| False | 4 | 40.0 | 14.3 | 4 | 40.0 | 14.3 | 2 | 20.0 | 7.1 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 1 | 50.0 | 3.6 | 1 | 50.0 | 3.6 | ||
| Minimal facts | 2 | 66.7 | 7.1 | 1 | 33.3 | 3.6 | 0 | 0.0 | 0.0 | ||
| Selected facts | 4 | 20.0 | 14.3 | 8 | 40.0 | 28.6 | 8 | 40.0 | 28.6 | ||
| Objectively true | 18 | 36.7 | 64.3 | 14 | 28.6 | 50.0 | 17 | 34.7 | 60.7 | ||
| Compliance | 2.504 | 0.343 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Selected facts | 2 | 16.7 | 11.1 | 4 | 33.3 | 22.2 | 6 | 50.0 | 33.3 | ||
| Objectively true | 16 | 38.1 | 88.9 | 14 | 33.3 | 77.8 | 12 | 28.6 | 66.7 | ||
| Surgery | 8.682 | 0.101 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 2 | 100.0 | 28.6 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 4 | 57.1 | 57.1 | 3 | 42.9 | 42.9 | 0 | 0.0 | 0.0 | ||
| Selected facts | 1 | 50.0 | 14.3 | 0 | 0.0 | 0.0 | 1 | 50.0 | 14.3 | ||
| Objectively true | 2 | 20.0 | 28.6 | 4 | 40.0 | 57.1 | 4 | 40.0 | 57.1 | ||
| Hard and soft tissues | 7.022 | 0.650 | |||||||||
| False | 2 | 50.0 | 28.6 | 0 | 0.0 | 0.0 | 2 | 50.0 | 28.6 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 1 | 100.0 | 14.3 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 1 | 100.0 | 14.3 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Selected facts | 1 | 25.0 | 14.3 | 1 | 25.0 | 14.3 | 2 | 50.0 | 28.6 | ||
| Objectively true | 3 | 27.3 | 42.9 | 5 | 45.5 | 71.4 | 3 | 27.3 | 42.9 | ||
| Function | 3.737 | 1.000 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 0 | 0.0 | 0.0 | 1 | 100.0 | 14.3 | 0 | 0.0 | 0.0 | ||
| Selected facts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 1 | 100.0 | 14.3 | ||
| Objectively true | 7 | 36.8 | 100.0 | 6 | 31.6 | 85.7 | 6 | 31.6 | 85.7 | ||
| Satisfaction | 3.301 | 0.755 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 1 | 100.0 | 11.1 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Selected facts | 1 | 50.0 | 11.1 | 0 | 0.0 | 0.0 | 1 | 50.0 | 11.1 | ||
| Objectively true | 8 | 33.3 | 88.9 | 9 | 37.5 | 100.0 | 7 | 29.2 | 77.8 | ||
| Harms | 4.003 | 0.381 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 1 | 50.0 | 7.7 | 0 | 0.0 | 0.0 | 1 | 50.0 | 7.7 | ||
| Selected facts | 5 | 22.7 | 38.5 | 8 | 36.4 | 61.5 | 9 | 40.9 | 69.2 | ||
| Objectively true | 7 | 46.7 | 53.8 | 5 | 33.3 | 38.5 | 3 | 20.0 | 23.1 | ||
| Oral hygiene | 0.819 | 1.000 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Selected facts | 1 | 25.0 | 16.7 | 2 | 50.0 | 33.3 | 1 | 25.0 | 16.7 | ||
| Objectively true | 5 | 35.7 | 83.3 | 4 | 28.6 | 66.7 | 5 | 35.7 | 83.3 | ||
| Microbiological and physiological | 1.509 | 1.000 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Selected facts | 0 | 0.0 | 0.0 | 1 | 50.0 | 33.3 | 1 | 50.0 | 33.3 | ||
| Objectively true | 3 | 42.9 | 100.0 | 2 | 28.6 | 66.7 | 2 | 28.6 | 66.7 | ||
| Efficiency and cost effectiveness | 2.354 | 1.000 | |||||||||
| False | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Nonfacts | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 0 | 0.0 | 0.0 | 1 | 100.0 | 12.5 | 0 | 0.0 | 0.0 | ||
| Selected facts | 1 | 33.3 | 12.5 | 1 | 33.3 | 12.5 | 1 | 33.3 | 12.5 | ||
| Objectively true | 7 | 35.0 | 87.5 | 6 | 30.0 | 75.0 | 7 | 35.0 | 87.5 | ||
| Other | 11.665 | 0.072 | |||||||||
| False | 0 | 0.0 | 0.0 | 1 | 20.0 | 11.1 | 4 | 80.0 | 44.4 | ||
| Nonfacts | 2 | 100.0 | 22.2 | 0 | 0.0 | 0.0 | 0 | 0.0 | 0.0 | ||
| Minimal facts | 1 | 33.3 | 11.1 | 0 | 0.0 | 0.0 | 2 | 66.7 | 22.2 | ||
| Selected facts | 2 | 50.0 | 22.2 | 2 | 50.0 | 22.2 | 0 | 0.0 | 0.0 | ||
| Objectively true | 4 | 30.8 | 44.4 | 6 | 46.2 | 66.7 | 3 | 23.1 | 33.3 | ||
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree