The role of artificial intelligence in providing accurate and reliable information on surgically-assisted rapid palatal expansion: A cross-sectional study

Introduction

This study aimed to evaluate the accuracy, reliability, and comprehensibility of information about surgically-assisted rapid palatal expansion provided by language models based on artificial intelligence (AI).

Methods

A cross-sectional content analysis was conducted on the responses to surgically-assisted rapid palatal expansion-related questions by ChatGPT-4 (OpenAI LLC, San Francisco, Calif), Gemini (Alphabet Inc, Mountain View, Calif), and Copilot (Microsoft, Redmond, Wash). In total, 115 questions (categorized into 11 domains) were created by 3 orthodontists and 1 oral and maxillofacial surgeon. The accuracy of the answers generated by the AI language models was independently evaluated by the same experts via a 5-point Likert scale. To test the relationships among categorical variables, when the sample size assumption was met, the Pearson chi-square test was used. However, when the sample size assumption was not met, Fisher’s exact test was applied. Analyses were performed in SPSS (version 27; IBM, Armonk, NY).

Results

The responses of the AI types presented a general homogeneous distribution, with no statistically significant difference between the types of AI and the types of responses ( P >0.05). Although there were no significant differences, ChatGPT-4 had the highest objectively true rate. In contrast, Gemini produced answers with more balanced accuracy, whereas Copilot had the highest number of false answers.

Conclusions

These findings reveal that the accuracy of AI-supported language models in providing medical information may vary according to subject matter.

Highlights

  • •

    We examined homogeneous responses of AI types about SARPE.

  • •

    AI types had a moderate level of information about surgically-assisted rapid palatal expansion.

  • •

    The accuracy of AI-supported language models may vary according to the subject matter.

  • •

    ChatGPT had the highest objectively true responses, although not statistically significant.

  • •

    Copilot had the highest false responses, although not statistically significant.

Transversal malocclusion of the maxilla is a common type of malocclusion seen in 8%-22% of the population, , and it can have significant effects on jaw and dental esthetics and functionality. The treatment approach for patients with maxillary transverse deficiency varies depending on skeletal maturity. In subjects with skeletal immaturity, it is possible to separate the midpalatal suture by applying orthopedic forces. In contrast, in subjects with skeletal maturity, the resistance increases because of midpalatal suture fusion, whereas the effectiveness of the expansion procedure decreases. This contributes to undesirable effects of the applied orthopedic forces, resulting in alveolar bending, tilting of teeth, and minimal maxillary expansion. These complications can be minimized by surgical release of resistant bony structures. Therefore, surgically-assisted rapid palatal expansion (SARPE) is indicated in subjects with skeletally maturity.

The SARPE procedure is based on the principles of distraction osteogenesis and aims to enlarge the maxilla by performing osteotomies on resistant structures. During this surgical intervention, osteotomies are applied to the aperture piriformis, processus zygomaticus, processus pterygoideus, and sutura palatina mediana to reduce the resistance of the maxilla to expansion forces. ,, The distraction force is transferred to the maxilla through tooth-based, bone-based, or hybrid devices, and the desired expansion is achieved. ,, Numerous studies have shown that SARPE increases maxillary width both skeletally and dentally. ,,, As a safe and simple method of correcting existing malocclusion by eliminating maxillary constriction, , the SARPE procedure increases the predictability and success of expansion while reducing side effects (eg, a possible reduction in alveolar bone thickness and height, bone dehiscence, and gingival recession) and orthodontic recurrence. ,

Pain, bleeding, infection, nerve damage, apical root resorption, and tooth discoloration and devitalization have been reported as complications of SARPE. Inadequate or asymmetrical expansion and infection of the maxillary sinus are rare, whereas expansion of the alar floor, tinnitus, lacrimation, life-threatening epistaxis, and skull base fracture are extremely rare. ,,,

Although the SARPE procedure has been proposed as a safe and simple method for patients with maxillary transverse insufficiency, surgical intervention can sometimes pose difficulties for patient acceptance. Patients may be concerned about the procedure, process, and results. In these circumstances, it is very important to provide the necessary and correct information to patients. If adequate information about the treatment process and potential complications is not provided, patients may turn to various sources and alternative information to relieve their concerns.

Artificial intelligence (AI)-based natural language processing and natural language generation (NLG) models have recently become widely used tools to answer medical questions. NLG models based on AI are models that can generate human-like text, answer questions, and perform other language-related tasks with high accuracy. These models are capable of generating meaningful and coherent responses from large datasets. Patients can use these language models to access health information, medical advice, counseling, and answer their questions. Generative pretrained transformer 4 (GPT-4), the fourth-generation language model of the GPT series developed by OpenAI (OpenAI LLC, San Francisco, Calif), is capable of generating detailed answers from a large amount of data. Similarly, Gemini (Alphabet, Inc, Mountain View, Calif), developed by Google (Alphabet, Inc), and Copilot (Microsoft, Redmond, Wash), powered by Microsoft’s AI innovation and using OpenAI’s GPT-4 language model, were developed via large language models. These chatbots are trained using large amounts of data and deep learning algorithms, making them proficient at generating meaningful responses and predicting text by identifying relationships between words. Although these modeling-based platforms can provide fast and seemingly accurate information via advanced language models to answer health-related questions via a variety of sources, including medical databases, research studies, and health Web sites, there is some uncertainty about the accuracy and reliability of the information provided. ,, This study aimed to comparatively analyze the comprehensibility, accuracy, and comprehensiveness of responses to questions about SARPE asked by patients to 3 different chatbots.

Furthermore, this study aimed to evaluate the quality of information provided to patients by AI chatbots, including ChatGPT-4, Copilot, and Gemini, and to examine their ability to provide patients with access to accurate and comprehensible information. In addition, we evaluated the medical accuracy of each robot’s responses and their effectiveness in terms of patient education. These findings can help determine how AI-based chatbots can be used more efficiently in health care.

Our hypothesis, in light of previous studies, was that the information provided by AI chatbots on SARPE would be accurate and that ChatGPT-4 would be the most accurate information provider, even if there would be differences among different chatbots.

Material and methods

In this study, content analysis was performed on the responses generated by ChatGPT-4 (OpenAI LLC), Gemini (Alphabet Inc), and Copilot (Microsoft) to questions related to SARPE. The questions were designed to cover all topics in which patients or laypeople may be asked and curious about SARPE. First, a question set consisting of 11 domains and 115 questions about SARPE was created by 3 authors (S.H., F.A.K.T., and O.O.Z.) ( Supplementary Table ). The question set was revised by a fourth author (E.C.O). AI chatbots were asked questions on January 6-8, 2025. Each question was asked individually and then deleted to avoid the learning process of the AI chatbots. The answer to each question was copied into a separate Word document and deleted from the AI chatbots’ memory. The evaluations were then made on separate Word documents created for each AI chatbot.

A customized Excel spreadsheet was created by 1 author (S.H.) to collect responses. For evaluation, the accuracy of answers provided by the AI robots was categorized into 5 categories: false, nonfacts, minimal facts, selected facts, and objectively true ( Table I ). A meeting was held to standardize the categorization and establish a common understanding of the evaluation system. The collected answers were independently evaluated by 3 orthodontists (S.H., F.A.K.T., and E.C.O.) and 1 oral and maxillofacial surgeon (O.O.Z.). In the case of a conflict among evaluators, a final decision was made by the principal investigator (S.H.).

Table I

Definition of the different accuracy categories

Category (score) Definition
Objectively true A claim that is based on scientific evidence and presents all relevant information, whether positive or negative
Selected facts A claim that presents some true selected facts based on scientific evidence, but omits important information related to a product
Minimal facts A claim that exaggerates the benefit of the product, with an overemphasis on the benefit supported by poor-quality scientific evidence
Nonfacts A claim that presents an intangible characteristic. Often, these claims are in the form of product opinions or lifestyle claims, leaving clinicians/patients to misinterpret the opinion as an objective product evaluation
False A claim that is objectively false, either because of a lack of evidence to support it or contradicts available evidence

The results evaluated in the analysis are not fixed and unalterable ā€œbasic factsā€ because the quality of the responses is based on personal judgment. To evaluate the quality of each response more objectively by examining the quality of each response from various perspectives by different evaluators, an analysis based on the principles of the crowd (or ensemble) score strategy was conducted.

Because the study evaluated only the responses obtained from AI models, institutional ethics committee approval was not needed.

Statistical analysis

Descriptive statistics (number and percentage) of the data are presented. For categorical variables, a Fisher’s exact test was applied when the sample size assumption was not met. Analyses were performed via the SPSS software (version 27; IBM, Armonk, NY).

Results

The distribution of the answers given according to AI type was given, and Fisher’s exact test was applied to examine the relationships among them. The results of the analysis revealed that the responses given by AI types generally presented a homogeneous distribution, and there was no statistically significant difference between the types of AI and response types ( P >0.05) ( Fig 1 ; Table II ).

Fig 1

Distribution of answers according to AI type.

Table II

Distribution of responses by AI types and their relationships

Answers ChatGPT Gemini Copilot Test statistic P value
n % % AI n % % AI n % % AI
False 6 27.3 5.2 5 22.7 4.3 11 50.0 9.6 10.509 0.203
Nonfacts 2 40.0 1.7 2 40.0 1.7 1 20.0 0.9
Minimal facts 9 50.0 7.8 6 33.3 5.2 3 16.7 2.6
Selected facts 18 23.7 15.7 27 35.5 23.5 31 40.8 27.0
Objectively true 80 35.7 69.6 75 33.5 65.2 69 30.8 60.0

For the domains, the distribution of the answers given according to AI type was provided, and Fisher’s exact tests were used to examine the relationships among them. On the basis of analysis, no statistically significant relationships were found between AI types and responses in any of the domains ( P >0.05) ( Table III ). Although the rates of objectively true responses were 69.6% for ChatGPT, 65.2% for Gemini, and 60.0% for Copilot, the differences among these rates were not statistically significant ( Table II ). Similarly, the rates of false responses were 4.3% for Gemini, 5.2% for ChatGPT, and 9.6% for Copilot, but the differences among these rates were not statistically significant ( Table III ).

Table III

Distribution of responses according to AI types for domains and the relationships among them

Answers ChatGPT Gemini Copilot Test statistic P value
n % % AI n % % AI n % % AI
Knowledge and information 6.149 0.657
False 4 40.0 14.3 4 40.0 14.3 2 20.0 7.1
Nonfacts 0 0.0 0.0 1 50.0 3.6 1 50.0 3.6
Minimal facts 2 66.7 7.1 1 33.3 3.6 0 0.0 0.0
Selected facts 4 20.0 14.3 8 40.0 28.6 8 40.0 28.6
Objectively true 18 36.7 64.3 14 28.6 50.0 17 34.7 60.7
Compliance 2.504 0.343
False 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Selected facts 2 16.7 11.1 4 33.3 22.2 6 50.0 33.3
Objectively true 16 38.1 88.9 14 33.3 77.8 12 28.6 66.7
Surgery 8.682 0.101
False 0 0.0 0.0 0 0.0 0.0 2 100.0 28.6
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 4 57.1 57.1 3 42.9 42.9 0 0.0 0.0
Selected facts 1 50.0 14.3 0 0.0 0.0 1 50.0 14.3
Objectively true 2 20.0 28.6 4 40.0 57.1 4 40.0 57.1
Hard and soft tissues 7.022 0.650
False 2 50.0 28.6 0 0.0 0.0 2 50.0 28.6
Nonfacts 0 0.0 0.0 1 100.0 14.3 0 0.0 0.0
Minimal facts 1 100.0 14.3 0 0.0 0.0 0 0.0 0.0
Selected facts 1 25.0 14.3 1 25.0 14.3 2 50.0 28.6
Objectively true 3 27.3 42.9 5 45.5 71.4 3 27.3 42.9
Function 3.737 1.000
False 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 0 0.0 0.0 1 100.0 14.3 0 0.0 0.0
Selected facts 0 0.0 0.0 0 0.0 0.0 1 100.0 14.3
Objectively true 7 36.8 100.0 6 31.6 85.7 6 31.6 85.7
Satisfaction 3.301 0.755
False 0 0.0 0.0 0 0.0 0.0 1 100.0 11.1
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Selected facts 1 50.0 11.1 0 0.0 0.0 1 50.0 11.1
Objectively true 8 33.3 88.9 9 37.5 100.0 7 29.2 77.8
Harms 4.003 0.381
False 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 1 50.0 7.7 0 0.0 0.0 1 50.0 7.7
Selected facts 5 22.7 38.5 8 36.4 61.5 9 40.9 69.2
Objectively true 7 46.7 53.8 5 33.3 38.5 3 20.0 23.1
Oral hygiene 0.819 1.000
False 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Selected facts 1 25.0 16.7 2 50.0 33.3 1 25.0 16.7
Objectively true 5 35.7 83.3 4 28.6 66.7 5 35.7 83.3
Microbiological and physiological 1.509 1.000
False 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Selected facts 0 0.0 0.0 1 50.0 33.3 1 50.0 33.3
Objectively true 3 42.9 100.0 2 28.6 66.7 2 28.6 66.7
Efficiency and cost effectiveness 2.354 1.000
False 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Nonfacts 0 0.0 0.0 0 0.0 0.0 0 0.0 0.0
Minimal facts 0 0.0 0.0 1 100.0 12.5 0 0.0 0.0
Selected facts 1 33.3 12.5 1 33.3 12.5 1 33.3 12.5
Objectively true 7 35.0 87.5 6 30.0 75.0 7 35.0 87.5
Other 11.665 0.072
False 0 0.0 0.0 1 20.0 11.1 4 80.0 44.4
Nonfacts 2 100.0 22.2 0 0.0 0.0 0 0.0 0.0
Minimal facts 1 33.3 11.1 0 0.0 0.0 2 66.7 22.2
Selected facts 2 50.0 22.2 2 50.0 22.2 0 0.0 0.0
Objectively true 4 30.8 44.4 6 46.2 66.7 3 23.1 33.3
Only gold members can continue reading. Log In or Register to continue

Stay updated, free articles. Join our Telegram channel

Jun 27, 2026 | Posted by in CARDIOLOGY | Comments Off on The role of artificial intelligence in providing accurate and reliable information on surgically-assisted rapid palatal expansion: A cross-sectional study

Full access? Get Clinical Tree

Get Clinical Tree app for offline access