Introduction
Nasal septum deviation (NSD) is one of the contributing factors to impaired nasal function and dentofacial developmental abnormalities. Although cone-beam computed tomography (CBCT) is clinically valuable for NSD diagnosis, manual interpretation remains labor-intensive and expertise-dependent.
Methods
Our study included 330 CBCT scans diagnosed with either NSD or non-NSD to develop an automated 2-stage artificial intelligence (AI) framework integrating real-time detection and classification for NSD screening. In the first stage, the YOLOv11 (You Only Look Once) object detection algorithm was employed to detect the region of interest containing the nasal septum. In the second stage, 3 convolutional neural network architectures, ResNet, EfficientNet, and MobileNet, were evaluated for classifying CBCT images into NSD and normal categories.
Results
Among the YOLOv11 variants, YOLOv11n demonstrated superior performance with a precision of 0.996, a recall of 1.000, an mAP50 of 0.995, and an mAP50-95 of 0.873. For the classification task, Mobile_small emerged as the top-performing model, achieving an area under the curve of 0.817, an area under the precision-recall curve of 0.845, and an accuracy of 0.749. An AI-assisted diagnostic tool was developed based on YOLOv11n and MobileNet models and validated on 50 internal and 50 external CBCT scans. With AI assistance, orthodontists’ diagnostic accuracy increased by 20.12% and 21.49%, respectively, whereas average diagnosis time decreased by 23.75 seconds, improving efficiency by 53.92%.
Conclusions
The proposed system enables rapid NSD screening with diagnostic-level accuracy, demonstrating the viability of lightweight AI models for clinical CBCT analysis. AI-assisted diagnosis improves orthodontists’ accuracy and time efficiency in identifying NSD.
Highlights
-
•
An AI model was developed for nasal septum deviation detection using CBCT.
-
•
The YOLOv11n model achieved exceptional performance in region of interest extraction.
-
•
Lighter models can outperform more complex architectures for this specific dataset.
-
•
AI-aided detection of nasal septum deviation increased orthodontists’ diagnostic accuracy and efficiency.
-
•
The model’s robustness was confirmed through internal and external validation.
The nasal septum is a bone-cartilage structure located in the middle of the nasal cavity, dividing it into left and right chambers. During craniofacial development, the nasal septum, as a midline structure, experiences persistent superior and ventral pressure exerted by the surrounding facial and maxillary palatal bones. Such mechanical influences render the nasal septum susceptible to deformities such as nasal septum deviation (NSD), which refers to the condition in which the bony, cartilaginous septum, or both, deviates from the midline of the face. The etiology of NSD is multifactorial, including trauma, genetics, congenital malformations, infections, tumors, and neoplasia. Globally, the prevalence of NSD ranges 1%-20% in neonates, , 9.5%-41.6% in children and adolescents, ,, and up to 86.6% in adults. ,
Mild deviation of the nasal septum may be asymptomatic. However, moderate to severe deviation is a significant contributor to nasal obstruction, epistaxis, sinusitis, turbinate hypertrophy, reduced nasal secretions, recurrent upper respiratory tract infections, and a range of other associated consequences such as altered nasal morphology, headaches, snoring, tinnitus, altered respiratory function patterns, obstructive sleep apnea, and craniofacial malformations. Among children and adolescents during the peak period of craniofacial growth, NSD is recognized as a contributing factor to abnormal craniofacial development. Impaired nasal breathing often necessitates chronic mouth breathing, causing skeletal open bite, mandibular growth pattern featuring clockwise rotation, with or without mandibular retrognathia, transverse maxillary deficiency with posterior crossbite, high arched palate, and lip incompetence. ,
The most common treatment for adult NSD is septoplasty. However, in pediatric and adolescent patients, nasal septal surgery may pose a risk of disrupting growth centers. ,, Recent studies have shown that rapid maxillary expansion (RME) in growing patients, as a conservative treatment approach, can improve NSD , and increase the volumes of the nasal cavity, nasopharynx, and oropharynx, thereby enhancing nasal airway ventilation. , Thus, early screening of NSD in growing patients may allow for early intervention through RME, promoting the proper development of the nasal cavity, dentofacial structures, and overall growth.
The diagnosis of NSD can be made through clinical and radiological examinations. For example, after the relief of nasal congestion, the location and degree of NSD can be assessed through anterior rhinoscopy and nasal endoscopy. However, these procedures can be uncomfortable and carry a high interrater variability. Acoustic rhinometry, rhinomanometry, and nasal spectral sound analysis may assist in identifying NSD in the anterior nasal cavity, but their utility when used in isolation is limited, lacking sensitivity and specificity compared with anterior rhinoscopy, nasal endoscopy, and radiological examinations. Cone-beam computed tomography (CBCT) is considered reliable for diagnosing NSD, and is commonly used in orthodontic diagnostic processes. Patients with NSD often seek consultation in orthodontic clinics because of dentofacial deformities, making the CBCT an effective tool for preliminary screening of NSD.
During the diagnosis of NSD using CBCT, clinicians should systematically review all sectional images to assess whether the nasal septum exhibits deviations. This process is time-consuming, involves a substantial amount of repetitive work, and heavily depends on the clinician’s experience. Inaccurate landmark localization can lead to erroneous results. Therefore, it is necessary to develop an accurate and efficient algorithm to automatically classify whether the nasal septum is deviated on CBCT images.
Artificial intelligence (AI), particularly deep learning, has revolutionized automated image analysis by leveraging multilayered mathematical computations to learn and infer complex patterns in data. Although a prior study has applied Mask R-convolutional neural network (CNN) to NSD diagnosis using CBCT, the approach was constrained by limited datasets, partial diagnostic scope, high computational demands, and annotation burden, restricting clinical applicability. To address these limitations, this study introduces a 2-stage NSD recognition method based on YOLO (You Only Look Once) and MobileNet. A dataset of 330 full-volume CBCT scans from patients aged 8-43 years was established, and the model was validated on internal and external datasets with orthodontists of varying diagnostic experience, aiming to enhance accuracy, efficiency, and generalizability in NSD detection.
Material and methods
As a retrospective study using routinely collected data from health care activities, this study was approved by the Ethics Committee of Shanghai Ninth People’s Hospital (No. SH9H-2025-T102-2) to be conducted without patients’ informed consent.
A total of 330 full-volume CBCT scans from Shanghai Ninth People’s Hospital were included for model training and testing. All CBCT scans were performed using iCAT FLX (KaVo Dental, Biberach, Germany) and after exposure parameters below: 5 mA, 120 kV, 16 × 13 cm field of view, scanning time of 20 to 30 seconds, and pixel size of 0.25 mm. The inclusion criteria were as follows: (1) no restrictions on age or sex; and (2) high-resolution full-volume CBCT scans encompassing the entire nasal cavity region. The exclusion criteria were as follows: (1) history of nasal surgery, major nasal trauma, or nasal tumors; (2) severe craniofacial developmental abnormalities; and (3) poor image quality, including evident artifacts or distortion.
CBCT samples were analyzed using Dolphin 3D Imaging System (version 11.7.05.66; Dolphin Imaging and Management Solutions, Chatsworth, Calif). The NSD diagnosis followed a protocol described by Gregurić et al. A normal nasal septum is shown as having alignment between the bony septum and the midline across all coronal CBCT layers. NSD, as shown in Figure 1 , is defined as the bony septum deviating from the midline of the face in coronal CBCT images (because CBCT cannot precisely differentiate between mucosa and cartilage, this study only detects bony NSD). The nasal deviation angle (measured in coronal CBCT images as the angle between the most deviated point of the septum and the midline) is >0°.
NSD diagnosis. 1, nasal deviation angle.
Two evaluators (an orthodontist with 3 years of CBCT interpretation experience who has received extensive training in NSD diagnosis and an experienced otolaryngologist) were requested to analyze 330 CBCT scans individually to determine the presence or absence of NSD in each patient. In instances of disagreement, the final diagnosis was established through discussion between the 2 evaluators. If no agreement could be reached, a senior otolaryngologist was consulted. Two weeks after the initial examination, a random sample of 50 CBCT scans was selected from the total of 330 and reanalyzed independently by the 2 evaluators.
In this study, we propose a 2-stage model for the identification of NSD from CBCT images. The overall workflow of the proposed model is illustrated in Figure 2 . The first stage involves detecting the region of interest (ROI) using YOLOv11, followed by training and evaluating different CNN models in the second stage. The best model is then selected and used for prediction in the third stage.
An algorithm framework for CBCT image analysis based on a 2-stage deep learning model. In stage I, a CBCT file is taken as the input, and the YOLOv11 model is used to perform detection, generating ROI images. In stage II, the ROI images are further processed using the CNN models.
We use the YOLOv11 object detection algorithm to identify the ROI containing the nasal septum in coronal CBCT images. The digital imaging and communications in medicine files were first read and processed using the Python package SimpleITK, with a window level of 150 and a window width of 1500 to enhance visualization. These files were converted into 3-dimensional arrays of shape [N, H, W]. Each slice was then saved as a JPG image using OpenCV (Intel, Santa Clara, Calif, open-source community).
A total of 30 patients’ JPG images were randomly selected for manual annotation. The superior boundary of the annotated region was defined as the center of the crista galli, the inferior boundary as the nasal floor plane, and the lateral boundaries as the nasal cavity walls at the widest point of the nasal cavity on the coronal plane. On the basis of this annotated dataset, the YOLOv11 model was trained to detect nasal septum regions in JPG images.
We employed 5 variants of the YOLO object detection model, YOLOv11n, YOLOv11s, YOLOv11m, YOLOv11l, and YOLOv11x, to perform ROI detection. During the training process, we adhered to the official configuration settings. The model was trained for 10 epochs with an input size of 3 × 640 × 640 pixels. The stochastic gradient descent optimizer was employed with an initial learning rate of 0.01 and a momentum of 0.937.
In the second stage, CNN models were trained to classify CBCT images as NSD or normal. Three architectures were evaluated: ResNet, EfficientNet, and MobileNet, each with varying model sizes to explore their performance in our binary classification task. Specifically, ResNet variants included ResNet_10, ResNet_18, ResNet_50, and ResNet_101, differing in depth and parameters. MobileNetV3_Small (Mobile_small) and MobileNetV3_Large (Mobile_large) were selected to compare model complexity and computational efficiency. EfficientNet variants comprised EfficientNet_B1, EfficientNet_EL, and EfficientNet_Lite0, balancing accuracy and resource requirements.
To ensure the generalizability of the CNN classification models, the dataset was split based on patients rather than individual images. Specifically, the 330 patients were divided into training and testing sets at an approximate ratio of 7:3, with the training set comprising 3600 images with NSD and 2533 images without NSD, whereas the testing set included 968 images with NSD and 647 images without NSD. All images were resized to 3 × 224 × 224 pixels, then normalized and standardized to ensure consistent input across networks.
During the training process of image classification, we employed the AdamW optimizer with a 0.001 initial learning rate and Cross-Entropy Loss to optimize the model parameters. Networks were implemented in the PyTorch framework, trained for 50 epochs with a batch size of 32. For model validation, a 5-fold cross-validation strategy was employed. The training set was further partitioned into 5 equal subsets. In each iteration, 1 subset was designated as the validation set, whereas the remaining 4 subsets were combined for model training. This process was repeated 5 times, with each subset serving as the validation set once. The final performance of the model was evaluated based on the average of the 5 validation results.
We report precision, recall, mAP50 (mean average precision), and mAP50-95 as the standard evaluation metric for detection. Precision measures the proportion of true positive detections out of all positive predictions made by the model. Recall (sensitivity) measures the proportion of true positive detections out of all actual positive instances in the dataset. mAP50 evaluates the average precision across all classes at a fixed intersection over union (IoU) threshold of 0.5. The IoU measures the overlap between the predicted bounding box and the ground truth bounding box. A higher IoU indicates better localization. mAP50-95 extends the evaluation to multiple IoU thresholds ranging 0.50-0.95, with increments of 0.05. It calculates the mean average precision across these thresholds, providing a more stringent and comprehensive evaluation of the model’s performance.
The performance of each CNN model is evaluated using 5 metrics: receiver operating characteristic (ROC) curve, precision-recall (PR) curve, accuracy, F1 score, and specificity. The ROC curve provides a visual tool for evaluating the trade-off between sensitivity and specificity. A higher area under the curve (AUC) indicates better overall performance. The PR curve focuses on the performance of the model in identifying positive cases, making it more informative than the ROC curve when dealing with highly imbalanced data. The area under the PR curve (AUPR) provides a single scalar value summarizing the model’s ability to achieve high precision while maintaining high recall. Accuracy measures the proportion of correctly classified instances (both true positives and true negatives) out of the total number of instances. The F1 score is the harmonic mean of precision and recall, providing a balanced measure that takes into account both false positives and false negatives. Finally, specificity measures the proportion of actual negative cases that are correctly identified by the model. It reflects the model’s ability to avoid falsely labeling negative samples as positive. It is defined as the ratio of true negatives to the sum of true negatives and false positives.
The final model is deployed as an integrated clinical tool, designed to assist clinicians, especially orthodontists, in diagnosing NSD. In the validation phase of the AI-assisted diagnostic tool, both internal and external CBCT datasets were used. An additional 50 internal CBCT scans from Shanghai Ninth People’s Hospital and 50 external scans from Anhui Medical University Stomatological Hospital were incorporated for performance evaluation. The external CBCT scans were acquired from a Meyer software (mDX-13STSP1A, Hefei, Anhui) at the following settings: 5 mA, 120 kV, exposure time of 20 seconds, scanning area of 23 × 18 cm, and focal point nominal of 0.5 × 0.5 mm. The CBCT scans were selected by an experienced otolaryngologist according to predefined inclusion and exclusion criteria, with a balanced ratio of patients without NSD to patients with NSD (1:1). The reference diagnosis of NSD for these scans was determined by the same otolaryngologist.
We recruited 2 groups comprising a total of 10 orthodontists from Shanghai Ninth People’s Hospital to evaluate the diagnostic accuracy and time efficiency of the AI tool. Group 1 (clinicians 1-5) consisted of orthodontists with >5 years of CBCT interpretation experience, whereas group 2 (clinicians 6-10) included those with <2 years of experience. Ten clinicians performed 2 rounds of diagnostic evaluation to assess accuracy. MobileNet was used as the standard model in all AI-assisted diagnoses to ensure consistent and comparable results. In the first round (manual), clinicians independently assessed NSD without AI support; in the second session (manual + AI), the same task was repeated with AI assistance. Diagnostic outcomes were benchmarked against the reference standard to evaluate accuracy. To assess time efficiency, 10 clinicians recorded interpretation times for 4 randomly selected CBCT scans under both conditions, with the second assessment performed after a 2-week interval.
Statistical analysis
Statistical analysis was performed by SPSS (version 25; IBM, Armonk, NY) and Python (version 3.8). Weighted κ was examined to assess interevaluator agreement and intraevaluator reliability. Analysis of variance was employed to assess the significance of differences in predictive performance among the CNN models. The Mann-Whitney U test was used to compare the performance of the AI tool in differentiating NSD between the internal and external datasets. Paired t tests were conducted comparing the accuracy, F1 score, and diagnostic time of manual and manual + AI scenarios. A 2-sided P <0.05 was considered statistically significant.
Results
The demographic characteristics of 330 CBCT scans are summarized in Supplementary Table I . The consistency of 2 evaluators was satisfying (weighted κ coefficient, 0.86 [95% confidence interval (CI), 0.80-0.92]). The intraevaluator reliabilities of 2 evaluators were good (evaluator 1: weighted κ coefficient, 0.84 [95% CI, 0.69-0.99]; evaluator 2: weighted κ coefficient, 0.88 [95% CI, 0.75-1.00]).
The performance metrics for the different YOLOv11 variants are summarized in Table I . From the data, it can be observed that YOLOv11n stands out as the top-performing model in this study, achieving a precision of 0.996, a perfect recall of 1.000, an mAP50 of 0.995, and an mAP50-95 of 0.873. Its lightweight architecture also ensures efficient deployment on edge devices or systems with limited computational resources.
Table I
Comparison of detection accuracy among different YOLOv11 variants
| Model | Precision | Recall | mAP50 | mAP50-95 |
|---|---|---|---|---|
| YOLOv11n | 0.996 | 1.000 | 0.995 | 0.873 |
| YOLOv11s | 0.998 | 0.933 | 0.980 | 0.861 |
| YOLOv11m | 0.997 | 1.000 | 0.995 | 0.855 |
| YOLOv11l | 0.996 | 1.000 | 0.995 | 0.871 |
| YOLOv11x | 0.995 | 0.989 | 0.971 | 0.837 |
The comparison of ROC curves and AUC values is shown in Figure 3 , A . As presented, Mobile_small stands out as the top-performing model in this study, achieving an AUC of 0.817 with a 95% CI of 0.796-0.838. This result is particularly impressive given its relatively lightweight architecture, which suggests that it can achieve high performance while maintaining computational efficiency. Compared with Efficient_lite0, which has the second-highest AUC of 0.794, Mobile_small increases the AUC by 2.3%. ResNet_50, a widely used deep CNN, achieves an AUC of 0.745, which is lower than both Mobile_small and Efficient_lite0. This suggests that, despite its deeper architecture and more parameters, ResNet_50 may not be the best choice for this specific dataset.
Performance of different deep learning models on the test dataset: A, ROC curves; B, PR curves; C, Comparison of model accuracy across diverse deep learning architectures.
The PR curves and their corresponding AUPR values for models are shown in Figure 3 , B . Mobile_small again stands out as one of the top-performing models, achieving an AUPR of 0.845 with a 95% CI of 0.821-0.869. This result is consistent with its high performance in the ROC curve analysis, indicating that it maintains high precision across a wide range of recall values. Efficient_lite0 achieves the highest AUPR of 0.863, with a 95% CI of 0.844-0.882. This model demonstrates excellent performance in terms of both precision and recall, making it particularly suitable for scenarios in which both metrics are critical. Resnet_10 is the least performing model with an AUPR of 0.722, highlighting the limitations of very shallow networks in capturing the necessary features for accurate classification in medical images. The bar chart in Figure 3 , C, shows the accuracy of various deep learning models.
To address the need for quantitative assessment, performance metrics were systematically calculated for 9 distinct CNN classification models. The evaluation metrics included accuracy, recall, precision, sensitivity, specificity, AUC, and AUPR. The detailed results of these metrics for each model are presented in Table II . To further evaluate the differences in predictive performance among the models, we employed analysis of variance to assess the significance of these differences ( Supplementary Table II ).
Table II
Comparison of quantitative results
| Model | Accountability | Recall | Precision | Sensitivity | Specificity | AUC | AUPR |
|---|---|---|---|---|---|---|---|
| Mobile_large | 0.748 | 0.834 | 0.766 | 0.834 | 0.620 | 0.796 | 0.835 |
| Mobile_small | 0.749 | 0.804 | 0.783 | 0.804 | 0.666 | 0.817 | 0.845 |
| Resnet_10 | 0.627 | 0.732 | 0.674 | 0.732 | 0.470 | 0.638 | 0.722 |
| Resnet_18 | 0.685 | 0.787 | 0.716 | 0.787 | 0.533 | 0.732 | 0.808 |
| Resnet_50 | 0.712 | 0.838 | 0.724 | 0.838 | 0.522 | 0.745 | 0.808 |
| Resnet_101 | 0.651 | 0.753 | 0.692 | 0.753 | 0.498 | 0.700 | 0.795 |
| Efficient_b1 | 0.731 | 0.804 | 0.761 | 0.804 | 0.623 | 0.778 | 0.829 |
| Efficient_lite0 | 0.717 | 0.732 | 0.782 | 0.732 | 0.694 | 0.794 | 0.863 |
| Efficient_el | 0.720 | 0.847 | 0.729 | 0.847 | 0.529 | 0.756 | 0.795 |
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree