Clinical advances in curve of Spee assessment: Deep learning for automatic tooth landmark detection in Invisalign

Introduction

The curve of Spee (COS) is a key indicator of occlusal function and orthodontic outcomes. Its measurement traditionally relies on manual landmark identification from intraoral scans, which is time-consuming and operator-dependent. This study introduces a deep learning–based method for fully automated COS assessment based on intraoral scan data, aiming to improve measurement efficiency and support the evaluation of the predictability of COS leveling in patients with varying vertical skeletal patterns undergoing Invisalign treatment (Align Technology, Santa Clara, Calif).

Methods

In this retrospective study, a total of 194 mandibular arch models were used to train and validate an automated network for measuring COS. This network adopted a Structure-Aware Long Short-Term Memory framework, which employed a 2-stage method for detecting coarse and fine tooth landmarks to assess COS depth. The accuracy of the landmarks was evaluated using the mean radial error and success detection rate, whereas COS depth was assessed using the paired Wilcoxon test. Finally, 55 patients with different vertical skeletal patterns were selected to analyze differences in COS leveling effects. In addition, the extrusion of teeth relative to the occlusal plane was compared.

Results

The proposed network closely approximated the manual method, with the mean radial error for landmark detection being 0.32 ± 0.09 mm. The median measurement error for COS was 0.05 mm ( P <0.001). ClinCheck predicted an average of 0.24 mm higher COS leveling than the actual outcome ( P <0.01). The hyperdivergent group exhibited the highest predictability at 69%, whereas the hypodivergent group showed the lowest predictability at 58%. The accuracy of extrusion relative to the occlusal plane at the first molar was the lowest (65%).

Conclusions

Deep learning can aid in measuring COS. In ClinCheck, considering various vertical skeletal patterns is necessary when designing the leveling objectives for COS, with the first molar requiring particular attention.

Highlights

  • Invisalign reduces COS depth by about 0.56 mm, varying by vertical skeletal type.

  • A novel SA-LSTM framework enables efficient, automated landmark detection.

  • Initial COS depth impacts predictability, especially in hypodivergent patients.

  • Automated COS analysis supports clinical planning; hypodivergent cases need caution.

  • Future work: refine SA-LSTM, expand trials, and track COS leveling stability.refine SA-LSTM, expand trials, track COS leveling stability.

The curve of Spee (COS), a physiological mandibular curvature formed by the incisal edges of anterior teeth and buccal cusps of posterior teeth, serves as an essential anatomic landmark for both occlusal function assessment and mechanotherapy planning in orthodontic practice. , As a representative of clear aligner therapy (CAT), the Invisalign appliance (Align Technology, Santa Clara, Calif) exhibits unique clinical efficacy in COS leveling through its digital biomechanical simulation system. ,

Patients with different vertical skeletal patterns exhibit marked differences in alveolar bone remodeling capacity, biomechanical properties of masticatory muscles, temporomandibular joint adaptability, and initial COS depth. , Conventional fixed orthodontic appliances employ different leveling protocols based on vertical skeletal patterns. However, the full-arch coverage and posterior tooth eruption–inhibiting effects of clear aligners may limit the direct applicability of conventional biomechanical principles. Although Ciavarella et al identified significant differences in COS depth changes among patients with different vertical skeletal patterns, they did not examine differences in COS leveling efficiency.

Existing studies on COS leveling have mostly employed manual or semiautomated measurement methods, ,,,, which are operator-dependent and often have limited interexaminer consistency and measurement reproducibility.

The high-precision 3-dimensional (3D) intraoral scans (IOS) provide robust data for clinical research. , Unlike tooth segmentation, landmark regression is highly sensitive to shape details. , Compared with 2-dimensional imaging, processing 3D volumetric data imposes a higher burden on network parameters and computational resources and results in substantially higher memory and processing overhead. Moreover, the broader search space of 3D data increases the risk of prediction errors during model inference. To ensure accurate measurement and evaluation, landmark detection should not be limited to identifying individual points but should incorporate the spatial relationships among landmarks in 3D space.

Although existing deep learning models for 3D dental imaging, such as 3D U-Net, point-cloud networks, , and transformer-based architecture, have demonstrated acceptable performance in tooth segmentation, they exhibit notable inaccuracies in landmark localization in clinical practice. A key limitation is that projection or voxelization of 3D models often leads to blurring of critical anatomic details such as cusps and fissures, which results in displacement of landmarks from their actual anatomic locations. Furthermore, point-based methods cannot effectively distinguish between adjacent teeth with similar morphologic features, increasing the likelihood of placing a landmark on the wrong tooth. , In 2-stage frameworks that perform segmentation followed by landmark detection, any error in the initial segmentation step can propagate and amplify during the landmark regression phase, especially when the region of interest (ROI) is inaccurately defined. Moreover, processing high-resolution volumetric data using neural networks results in substantial computational overhead and often exhausts the memory capacity of standard graphics processing units.

This study aimed to automatically identify dental anatomic landmarks using deep learning to calculate COS depth and analyze the efficiency of COS leveling in patients with different vertical skeletal patterns.

Material and methods

This study was approved by the Research Ethics Committee of the Stomatological Hospital of Chongqing Medical University (2023 [LS No. 046]). Before data acquisition, all participants were confirmed to have fully dentate mandibular arches with no missing teeth except for the third molar. A total of 194 IOS models were obtained using an iTero intraoral scanner (Align Technology, Santa Clara, Calif) from the Department of Orthodontics at the Affiliated Stomatology Hospital of Chongqing Medical University. All datasets underwent thorough deidentification and anonymization before analysis and were stored in the standard tessellation language format with the following parameters: optical accuracy, 40 μm; wavelengths, 460 and 520 and 640 nm; scanning rate, 18 frames per second; depth of field, 0.1-20 mm.

Two orthodontists (S.M. and Y.W.), each with >5 years of experience, annotated anatomic landmarks on the collected 3D mandibular models using the Materialise Magics software (version 21.0; Materialise NV, Leuven, Belgium) ( Table I ). During annotation, the 2 clinicians meticulously identified relevant cusps and documented their corresponding coordinates (x, y, and z) in an Excel spreadsheet (version 2021; Microsoft Corporation, Redmond, Wash). The final results were obtained by calculating the mean value of their annotations. During annotation, each 3D model was manipulated using a mouse-driven graphic cursor to adjust viewing angles and spatial orientation for identifying anatomic landmarks. The software’s integrated spatial coordinate system automatically recorded 3D coordinates (x, y, and z), which were systematically documented in the Excel spreadsheet. The selection criteria for cusp landmark points were as follows: the teeth in the 3D models were represented by 3D surfaces composed of triangular meshes, with each vertex having 3D coordinates. The most prominent triangular vertex from each perspective was selected as a cusp landmark point by zooming in on the target tooth and rotating the field of view to obtain detailed images of the tooth surface ( Fig 1 ). The definition and selection criteria for cusp landmark points adhered to the standards established by Proffit and Nelson. ,

Table I

Anatomic landmarks and their corresponding definitions

Landmark Definitions
A The midpoint of the incisal edge of the mandibular left central incisor
B The midpoint of the incisal edge of the mandibular right central incisor
C The midpoint of line segment AB
D The distobuccal cusp of the second molar in the mandibular left jaw
E The distobuccal cusp of the second molar in the mandibular right jaw
F The mesiobuccal cusp of the second molar in the mandibular left jaw
G The mesiobuccal cusp of the second molar in the mandibular right jaw
H The mesiobuccal cusp of the first molar in the mandibular left jaw
I The mesiobuccal cusp of the first molar in the mandibular right jaw
J The buccal cusp of the mandibular right second premolar
K The buccal cusp of the mandibular left second premolar
L The buccal cusp of the mandibular left first premolar
M The buccal cusp of the mandibular right first premolar
N The cusp of the mandibular left canine
O The cusp of the mandibular right canine
Fig 1

Workflow of landmark annotation using Materialise Magics (version 21.0; Materialise, Leuven, Belgium).

The automatic detection of tooth landmarks was performed using a deep learning algorithm proposed by Chen (Open resource: https://github.com/runnanchen/SA-LSTM-3D-Landmark-Detection ). This algorithm was structured in 2 phases: (1) the first phase involved the use of a Structure-Aware Long Short-Term Memory (SA-LSTM) network to perform heatmap regression on down-sampled volumetric data obtained from IOS model to predict coarse landmarks and (2) the second phase involved generating multiresolution 3D predictions of lesions, along with detailed annotations, by cropping the regions annotated as coarse landmarks ( Fig 2 ).

Fig 2

Schematic representation of the proposed methodology. LLCI (A) , midpoint of the incisal edge of the mandibular left central incisor; RLCI (B) , midpoint of the incisal edge of the mandibular right central incisor.

In 3D models generated from IOS, significant variations may exist in the field of view and contextual details. To reduce computational costs and enhance algorithm efficiency, the network used a down-sampled version of the original IOS volume as the final input. This down-sampling may reduce the resolution of volumetric data, potentially making certain dental anatomic structures, such as the cusps and pits of molars, less distinct or unidentifiable. However, the critical contextual information embedded in the complex structure of human teeth, such as landmarks, is preserved in down-sampled data, which enables the network to approximate the locations of these landmarks accurately. By leveraging the global spatial context for identifying rough landmarks, the network enhances the efficiency of subsequent fine landmark identification.

During the training phase, a multiresolution cropping strategy was used for data augmentation, wherein the sizes of 3D patches varied across iterations in the fine-scale landmark detection process. This strategy allowed the SA-LSTM network to integrate outcomes from various perceptual fields, thereby improving prediction accuracy. In addition, to promote diversity among training samples, Gaussian noise was introduced at the center of iterative landmark crops, contributing to a more accurate ground truth for landmark identification.

After processing the down-sampled data, the network generated 3D heatmaps. Each heatmap showed a probability distribution for a specific landmark and indicated its approximate location. These heatmaps were integrated to obtain an initial set of feature points for fine-scale landmark detection.

The SA-LSTM network employs attention gates and a regression voting strategy to randomly generate a substantial number of patches, synthesizing diverse predictions to gradually produce a continuous output. To enhance the reliability of these predictions, the network integrated a graph attention module that encodes global shape structures. Within this framework, the attention gate assessed the relevance between the hidden feature states and the predicted landmarks. It iteratively removed irrelevant features while aggregating the most pertinent ones to form the final prediction. This strategy enhances processing efficiency in managing 3D volumetric data and minimizes the consumption of computational resources.

To mitigate accuracy loss associated with down-sampling volumetric data and improve the precision of fine-scale landmark detection, the network extracted 3D patches surrounding rough landmarks directly from the original volumetric data, thereby preserving essential geometric details.

The experimental framework was implemented on a Linux operating system using Python (version 3.8; Python Software Foundation, Wilmington, Del) within an Anaconda environment. Medical images were processed using SimpleITK (version 2.3.1; Insight Software Consortium, Chapel Hill, NC), with numerical operations managed through NumPy 1.24.4 and Pandas 2.0.3. Deep learning models were developed using PyTorch (version 2.4.1; Meta AI Research, Menlo Park, Calif) with Adam optimization (learning rate = 0.0002) and batch size of 1. Training occurred over 200 epochs on an NVIDIA GeForce RTX 3090 graphics processing unit (NVIDIA Corporation, Santa Clara, Calif).

After the training was completed, we implemented a 5-fold cross-validation approach as described by Chen et al to evaluate network performance. The dataset comprising 194 models was divided into 5 independent subsets (n = 35, 35, 35, 35, and 34) and an independent validation set (n = 20). During each validation round, 1 subset was designated as the test set, whereas the remaining 4 subsets served as training sets. This division adhered to a 7:2:1 ratio for training, test, and validation datasets. During 5 experimental iterations, each subset was sequentially assigned as the test set, with results aggregated across all rounds to derive landmark coordinates for individual teeth. The 5 models trained during the 5-fold cross-validation process were independently used to infer the validation set. Outcomes were consolidated to comprehensively evaluate the performance and stability of the model in the independent validation set, ensuring robust assessment of generalizability.

To assess the effectiveness of dental landmark detection, we used the mean radial error, standard deviation, and success detection rate in 5 target radii (1, 2, 3, 4, and 8 mm) as key metrics for measuring detection accuracy. Results were aggregated to calculate the overall error of the model across the entire dataset. Furthermore, predicted and ground truth landmarks in the JavaScript Object Notation format, along with the 3D model in the Neuroimaging Informatics Technology Initiative format, were imported into the 3D Slicer software (version 5.6.1; Brigham and Women’s Hospital, Boston, Mass). Spatial registration mapping was performed using the built-in “Markups to Model” module. A dual-view mode combining “Volume Rendering” and “Markups Fiducial” displays was used to visualize predicted ( red spheres) and ground truth ( green spheres) points with their spatial deviations on anatomic structures. The inference speed per model was used to evaluate the efficiency of the proposed network.

On the basis of an effect size (q) of 0.7, an α value of 0.05, and a β value of 0.1, the required sample size for the artificial intelligence (AI) and manual groups was estimated to be 46 participants each using G∗Power (version 3.1; Heinrich-Heine-Universität Düsseldorf, Düsseldorf, Germany). Furthermore, we retrospectively selected 55 patients from 2016 to 2024 based on established inclusion and exclusion criteria ( Table II ). The cohort comprised 45 women and 10 men, with a mean age of 24.1 ± 4.9 years. Pretreatment cephalograms were obtained using a Kodak 9000 system (Carestream Health, Rochester, NY) in the digital imaging and communications in medicine format with the following settings: 62 kV, 8 mA, 3.5-lp/mm resolution, and 0.8-second exposure. On the basis of the SN-MP angle recorded at initial examination, the patients were categorized into 3 vertical skeletal pattern groups: (1) 14 hyperdivergent (SN-MP, >35.5°), (2) 21 normodivergent (SN-MP, >30.5° and ≤35.5°), and (3) 20 hypodivergent (SN-MP, ≤30.5°).

Table II

Inclusion and exclusion criteria

Inclusion criteria Exclusion criteria
Adults, aged >18 y, skeletal age CS6 according to the cervical vertebral maturation method Use of auxiliaries in combination with clear aligners
No missing teeth, nonextraction orthodontic treatment Loose or visibly loose teeth or periodontal disease
Teeth with no noticeable wear Temporomandibular disorder
Sequence of ≥15 aligners Severe skeletal deformities or abnormal tooth shapes
During the orthodontic treatment, attachments were present and well-maintained. The appliance was worn for >22 h/d and replaced every 2 wk Patients exhibiting posterior crossbite malocclusion or significant vertical growth patterns
No history of maxillofacial trauma or surgery Systemic diseases that may affect tooth movement
No periodontal or temporomandibular joint disorders Abnormal molar crown morphology, particularly short crowns or insufficient eruption height
Mild or moderate crowding (4-6 mm), not addressed in the plan and untreated with IPR Implants, restorations, and nonremovable teeth or teeth with root canal treatment

The patients presented with Class Ⅰ or Class Ⅱ malocclusion and underwent the first series of Invisalign at the Orthodontic Department of the Affiliated Stomatological Hospital of Chongqing Medical University by the same orthodontist with extensive experience in CAT. To ensure the best clinical outcome for each patient’s ClinCheck (eg, overcorrection when necessary), the orthodontist made personalized adjustments. Attachments were used as an aid to treatment when necessary. The orthodontist accepted the default Align attachment placement if deemed appropriate; otherwise, additional attachments were added to optimize the treatment plan. In all patients, COS-level attachments were not removed from the Align default settings. In addition, no patient underwent predesigned or actual interproximal reduction (IPR) treatment.

Dental models were generated for each patient at 3 time points: initial scan (pretreatment), ClinCheck prediction, and actual outcome (clinical scans after treatment with the first set of Invisalign aligners). These models were stored in the standard tessellation language format. An occlusal plane (OP) was defined as the midpoint between the edges of the 2 mandibular central incisors, the distobuccal cusp tip of the left second molar, and the distobuccal cusp tip of the right second molar ( Fig 3 , A ). The vertical distances from the canine cusp tip, first premolar buccal cusp tip, second premolar buccal cusp tip, first molar mesiobuccal cusp tip, and second molar mesiobuccal cusp tip to the OP were measured, and the average of the left and right sides was recorded ( Fig 3 , B ). COS was calculated as the average of the largest vertical distances recorded from both sides.

Fig 3

Measurement of COS: A , Points A, B, and C correspond to the midpoints of the incisal edges of the mandibular left central incisor, mandibular right central incisor, and the line segment AB, respectively. Points D and E represent the distobuccal cusps of the second molars in the left and right mandibles, respectively; B , COS was calculated as the average of the largest vertical distance recorded between the left and right sides.

To minimize measurement errors, a trained orthodontist manually measured (S.M) COS depth for all 165 mandibular models as described earlier. The same orthodontist repeated the measurements after 2 weeks. In addition, the SA-LSTM network was used to assess COS depth in all models (AI group). To ensure the consistency and accuracy of the annotation, another dentist (Y.W.) with >15 years of experience reviewed all landmarks and made necessary modifications.

Intrarater reliability was assessed through duplicate measurements of 10 randomly selected mandibular models by a researcher, with a minimum 2-week interval among measurements. Interrater reliability was evaluated using the same set of 10 models measured independently by 2 researchers (S.M. and Y.W.).

To evaluate the efficiency and error of manual COS depth measurements, 10 randomly selected IOS models were used. The average measurements by 2 associate chief orthodontists served as the gold standard. Four dentists from different specialties (2 orthodontists, 1 periodontist, and 1 oral-maxillofacial surgeon) independently measured the models at 3-day intervals, with their time and results recorded and compared against the gold standard.

Statistical analysis was performed using SPSS (version 26.0; IBM Corp, Armonk, NY). Intrarater and interrater consistencies were evaluated using the intraclass correlation coefficient (ICC). The Kolmogorov-Smirnov and Shapiro-Wilk normality tests were used to assess data distribution. Pearson correlation coefficients were calculated to quantify linear relationships between measurement errors/operating time and number of attempts. The mean operating time with standard deviation was calculated for all practitioners and subgroups stratified by specialty and was visualized on time-series trend graphs. Differences in COS depth measurement and operating time across specialties were statistically analyzed using the paired t test or Wilcoxon signed rank test, which was selected based on normality. When data distribution was normal, the Bland-Altman test and paired t test were used to compare the data obtained through AI with that obtained through manual measurement. Continuous variables that did not fit a normal distribution were expressed as quartiles (median [25th percentile, 75th percentile]) and compared between groups using the Wilcoxon signed rank test or the Kruskal-Wallis test. The statistical significance level was established at P = 0.05.

Results

The measurement protocol demonstrated excellent intrarater reliability (ICC, 0.982 [95% confidence interval (CI), 0.953-0.994]) and strong interrater agreement (ICC, 0.943 [95% CI, 0.863-0.978]).

Among the 174 IOS models, a total of 2436 teeth were included in the test set. The accuracy rate of landmark detection in this dataset exceeded 97%, with an average deviation of 0.32 mm between the predicted and actual landmarks ( Fig 4 ; Supplementary Table I ). The detection time for a single model was 0.02 seconds. Figure 5 presents several representative examples of tooth landmark detection results.

Fig 4

Error of the SA-LSTM network across different dental landmarks during 5-fold cross-validation. Error bars represent the standard deviation. The mean radial error was 0.32 ± 0.09 mm ( purple ), demonstrating the stability and consistency of the network in predicting dental landmarks.

Fig 5

Representative landmark detection results: A, The automated network accurately detected all landmarks on a single mandibular arch, shown from the occlusal and lateral views; B, Four samples of different tooth types with accurate cusp detection. Green , ground truth landmarks; red , indicate predicted landmarks.

In the independent validation set (n = 20), the SA-LSTM network demonstrated robust prediction performance ( Supplementary Table II ). The mean radial error of landmark detection was 0.34 ± 0.35 mm, which was comparable to that in the test set, and the success detection rate of landmarks within 2 mm was 99%. The inference time for a single mandibular model was slightly higher, averaging only 0.8 seconds.

The operating time for COS measurement was negatively correlated with the number of attempts (r = −0.92, P <0.05); however, no significant correlation was observed between measurement error and the number of attempts (r = 0.05, P = 0.89) ( Supplementary Table III ). The training curve showed that the mean operating time of all dentists significantly decreased with an increase in the number of attempts ( Fig 6 , A ). The average measurement error of dentists with different specialties was 0.06 mm, and no significant difference in measurement errors among dentists from different specialties ( P = 0.88) ( Supplementary Table IV ). However, the average time spent by dentists with different specialties differed by 2.35 minutes ( P <0.01). In the first attempt, orthodontists completed the operation 6.67 minutes faster than general dentists; however, this time difference gradually decreased after the second attempt ( Fig 6 , B ).

Fig 6

Training curves of manual measurements: A, The average measurement time for a single mandibular model varied with the number of attempts; B, Average time divided by the number of attempts and the professional background of the operator.

A total of 55 patients met the inclusion criteria for this study. The demographic characteristics and treatment durations of these patients are shown in Table III . These data did not conform to a normal distribution ( Supplementary Table V ); consequently, intergroup comparisons were performed using either the Wilcoxon signed rank test or the Kruskal-Wallis test. The ICC value of 0.98 (95% CI, 0.92-0.99) suggested robust consistency.

Table III

Demographics and treatment duration

Variable Descriptive statistics
Age at start of treatment, y
Males 22.5 ± 3.7
Females 24.5 ± 5.0
Total 24.1 ± 4.9
Sex
Male 10 (18.2)
Female 45 (81.8)
Total 55 (100)
Treatment duration, mo
Males 20.9 ± 5.12
Females 23.3 ± 13.21
Age categories, y
18-25 35 (63.6)
26-35 18 (32.7)
>36 2 (3.6)

Note. Values are expressed as mean ± standard deviation or number (percentage).

Significant differences were found between COS depths measured using AI and manual methods ( P <0.001), with the median difference being 0.05 mm ( Fig 7 ; Supplementary Table VI ). The median error in measuring COS depth at all cusps using AI was <0.1 mm.

Fig 7

Violin plot demonstrating the errors associated with COS depth measurements using AI and manual methods at each designated point. L7, L6, L5, and L4 represent the average COS depths at the second molars, first molars, second premolars, and first premolars on both the left and right sides, respectively. COS denotes the average depth of the deepest points on both mandibular sides.

Normal distribution verification ( Supplementary Table VII ) revealed that all models generated at different time points conformed to a normal distribution. The Bland-Altman plot indicated significant differences between COS depths measured at ClinCheck prediction using the AI and manual methods, with measurement errors exceeding 0.1 mm ( Fig 8 ). Among AI measurements, the initial COS depth was the lowest in the hyperdivergent group ( Fig 9 ), whereas the hypodivergent group showed the lowest prediction rate for COS leveling ( Table IV ).

Fig 8

Bland-Altman plots showing agreement between COS depths measured using AI and manual methods at different time points: A, Before treatment; B, After treatment predicted by ClinCheck; C, After actual treatment; D, Predicted COS change; E, Actual COS change after treatment; F, Unachieved change.

Fig 9

COS leveling results obtained using AI.

Table IV

The COS leveling results obtained through an AI method

Group T0 T0-T1 T0-T2 Difference between T1 and T2 Predictability, %
Total (n = 55) 2.00 ± 0.08 0.81 ± 0.06 0.56 ± 0.06 0.24, P = 0.001 69
Hypodivergent (n = 20) 1.98 ± 0.66 0.82 ± 0.53 0.55 ± 0.58 0.27, P = 0.005 58
Normdivergent (n = 21) 2.21 ± 0.53 0.90 ± 0.10 0.66 ± 0.09 0.24, P = 0.02 67
Hyperdivergent (n = 14) 1.73 ± 0.57 0.64 ± 0.32 0.44 ± 0.39 0.20, P = 0.02 69

Note. Data are presented as mean ± standard deviation. The data for all subgroups are normally distributed. All measurements are in millimeters. A positive change indicates leveling of COS. Predictability was calculated by dividing T0-T1 by T0-T2.

T0 , initial scan; T1 , ClinCheck prediction; T2 , actual outcome.

The hyperdivergent group demonstrated the lowest COS depth before treatment (1.73 ± 0.57 mm; Table IV ). ClinCheck overestimated COS leveling compared with actual outcomes ( P = 0.001). Notably, the hyperdivergent group required the least planned leveling (0.64 mm) while exhibiting minimal clinical COS leveling (0.20 mm, P = 0.02). Conversely, the hypodivergent group exhibited the largest deviation between predicted and actual leveling outcomes (0.27 mm, P = 0.005). The overall predictability of COS leveling was 69%, with vertical dimension stratification showing higher accuracy in the hyperdivergent group (69%) than in the hypodivergent group (58%).

COS depth was the highest at the first molar before treatment, with a mean of 1.94 mm ( Table V ), whereas it was the lowest at the second molar. The second molar required the least planned leveling. A significant difference was observed between preset and actual extrusion at each tooth except the second molar ( P <0.01). In addition, the leveling efficiency at the first molar was the lowest, with a mean value of 65%.

Table Ⅴ

Extrusion of teeth relative to the OP by AI method

Tooth T0 T0-T1 T0-T2 T0-T1 vs T0-T2 Expression, %
Canines (cusp tip) 0.85 ± 0.51 0.53 ± 0.57 0.40 ± 0.62 0.12 ± 0.34, P <0.01 75
First premolars (buccal tip) 1.35 ± 0.64 0.88 ± 0.80 0.67 ± 0.80 0.21 ± 0.51, P <0.01 76
Second premolars (buccal tip) 1.68 ± 0.75 0.97 ± 0.88 0.63 ± 0.97 0.34 ± 0.43, P <0.01 65
First molars (mesiobuccal cusp tip) 1.94 ± 0.74 0.89 ± 0.91 0.58 ± 1.00 0.24 ± 0.45, P <0.01 65
Second molars (mesiobuccal cusp tip) 0.83 ± 0.41 0.30 ± 0.58 0.26 ± 0.59 0.03 ± 0.28, P = 0.22 87

Note. Data are presented as mean ± standard deviation. All measurements are in millimeters. A positive change indicates extrusion of teeth. Expression was calculated by dividing T0-T2 by T0-T1.

T0 , initial scan; T1 , ClinCheck prediction; T2 , actual outcome.

Discussion

COS leveling is a key stage in orthodontic treatment. The effectiveness of CAT in achieving COS leveling remains controversial, particularly among patients with different vertical skeletal patterns. Existing studies have relied on manual COS measurement, requiring approximately 8 minutes per model. ,, Our deep learning–based method automated landmark detection on IOS, completing COS measurement within 1 second. This approach significantly improves efficiency and enables large-scale assessment of COS leveling.

The key to evaluating the leveling effect of CAT lies in COS depth measurement; however, its complexity stems from multiple compound factors. Two-dimensional measurement involves generating a COS on a lateral cephalogram and measuring its depth and radius of curvature. However, this method neglects the 3D structural characteristics of teeth. In particular, the actual COS depth may be underestimated in hyperdivergent patients owing to the relatively upright vertical alignment of the dentition. Although cone-beam computed tomography digital models can effectively measure COS depth, their spatial resolution has limitations when capturing fine structures such as the intercuspal area. On the contrary, IOS models have extremely high precision and provide data for the application of AI in domains such as automated measurement and morphologic analysis. Most of COS measurement software based on IOS models are closed source, including semiautomatic ones such as OrthoCad (Align Technology, Santa Clara, Calif) and fully automatic ones such as 3Shape Ortho System (3D shape A/S, Copenhagen, Denmark). However, the measurement accuracy of this software remains unknown. In this study, manual COS measurements were performed using the open-source software Materialise Magics. Clinical training curves showed an insignificant mean error of 0.06 mm ( P = 0.88), which is comparable to previously reported errors in linear digital model measurements. In this study, manual COS measurement required approximately 8 minutes on average, which is similar to the time reported by Yu for Bolton ratio analysis. However, the latter involves marking more landmarks. This discrepancy arises because Bolton ratio analysis occurs on a 2-dimensional plane, simplifying landmark selection. On the contrary, COS measurement requires rotating the model in 3D space and precisely identifying landmarks from multiple angles, making it more complex and time-consuming.

Cutting-edge studies have primarily focused on tooth recognition on 3D scans ,,, ; however, dedicated studies on landmark detection remain relatively limited. ,,,,, Regressing coordinates directly from high-dimensional inputs, such as IOS, presents challenges in learning nonlinear mapping. A TS-MDL framework developed by Wu et al initially adopted iMeshSegNet and optimized the predicted segmentation masks through graph cutting, treating these masks as independent of ROIs. These ROIs were subsequently submitted to a series of regression networks known as PointNet-Reg, which generated heatmaps representing the encoded coordinates of dental landmarks. Expanding on this concept, Pfister et al encoded landmark position information into Gaussian heatmaps, effectively transforming the point localization task into a more manageable image-to-image/heatmap regression task. Furthermore, Kumar et al proposed an algorithm for automatic identification of certain dental features; however, this approach faced limitations in detecting landmarks outside its defined scope. To address this issue, we introduced a 2-stage SA-LSTM framework based on the structural characteristics of human teeth. This approach allowed the retention of contextual information in the down-sampled volume, effectively overcoming the challenges presented by large-scale initial IOS. Results revealed minimal discrepancies between AI and manual measurements (<0.1 mm) and stable performance in the independent validation set, supporting the reliability of the model. In addition, the proposed automated framework for COS measurement reduced single-model measurement time to 0.02 seconds, showing potential for clinical application.

Furthermore, we evaluated the feasibility of using deep learning to predict clinical outcomes and assess the impact of CAT on COS changes among patients with different vertical skeletal patterns. The Wilcoxon test and Bland-Altman plots indicated that most results generated by the AI algorithm were consistent with manual measurements; however, some discrepancies were noted. Therefore, although AI algorithms offer a reliable method for evaluating large datasets, it is essential to acknowledge and address potential errors relative to ground truth.

Our study excluded patients with IPR mainly because of the following reasons: clear aligners deliver orthodontic forces by closely enveloping the tooth crowns with elastic membranes, and their mechanical efficiency is significantly positively correlated with the interproximal contact area between teeth. The reduction in the mesiodistal diameter of the tooth crown caused by IPR may change the stress distribution pattern between the orthodontic appliance and tooth surface, affecting the objective assessment of the original efficacy of the appliance. In addition, clinical variability in the timing of IPR intervention may introduce uncontrollable mechanical confounding factors, interfering with the analysis of the mechanisms of vertical movement and mesiodistal space closure. In terms of vertical control strategies, this study selectively excluded auxiliary devices such as occlusal ramps and intermaxillary elastic traction, aiming to evaluate the inherent biomechanical potential of CAT in leveling COS through standardized treatment parameters.

In this study, the mean COS depth of 2 mm before treatment aligns with the findings reported previously , ; however, differences in COS depth were observed among patients with different vertical skeletal patterns. The mean initial COS depth in the hypodivergent group was 1.98 mm, which was lower than that in the normodivergent group (2.21 mm). These findings are inconsistent with those of studies indicating that hypodivergent patients exhibit the highest COS depth. This discrepancy may be attributed to the effects of sample distribution. Hyperdivergent patients typically present with a flatter initial COS, likely owing to larger mandibular plane angles leading to a more vertical alignment of mandibular teeth. ,

Invisalign led to an overall COS leveling of 0.56 mm after treatment ( P = 0.001); however, the effect varied among different vertical skeletal types. These findings are consistent with those reported by Goh et al, who observed an average COS leveling of 0.55 mm. A possible explanation is that clear aligners can effectively intrude mandibular anterior teeth to level the COS, achieving outcomes comparable to those of traditional fixed appliances. By fully covering the tooth surface, aligners centralize force application near the resistance center, promoting controlled and uniform tooth movement while preserving root angulation to prevent undesired anterior tipping. In this study, the predictability of COS leveling in hyperdivergent patients was 69%, likely owing to the lower initial COS depth and a preset minimum flattening requirement (0.64 mm) that reduced leveling difficulty. Conversely, hypodivergent patients, with a higher initial COS depth and larger preset flattening requirement (0.82 mm), exhibited higher flattening difficulty and lower predictability of COS leveling (58%). Overall, the efficiency of Invisalign in flattening the COS was 69%, which was significantly higher than that reported by Goh et al (35%). Its efficacy in leveling COS may improve with increasing facial angle.

Our study corroborates previous findings indicating that COS is deepest at the mesial buccal tip of the first molar before treatment, ,, with the mean value recorded at 1.94 mm. Stability and depth at this location may be influenced by the wear patterns of the teeth. Our analysis of individual tooth movements relative to the OP, as outlined in the ClinCheck plan, revealed favorable outcomes after treatment. Results indicated a gradual decrease in extrusion movement from the first premolar to the second molar, which is consistent with the findings reported by Goh et al. Notably, this study confirms that the extrusion rate for the first molar (65%) is lower than that for other teeth, likely owing to the significant chewing load at the first molars, which presents challenges in achieving vertical movements through CAT. Supporting this observation, Boyd and Waskalic noted that even with auxiliary attachments, aligners struggled to achieve extrusion of posterior teeth owing to insufficient contact efficiency between the aligner and attachments, which reduced mechanical effectiveness. In addition, the compressive effect on the molars may increase mandibular vertical height, particularly in patients with strong chewing muscles, increasing the risk of posttreatment recurrence.

Unlike previous studies that have investigated COS leveling outcomes using traditional methods, ,,, this study introduced a deep learning–based approach that demonstrated excellent reliability and validity. The findings support the use of AI as a suitable tool for comparing tooth movements predicted by ClinCheck with actual clinical outcomes. In addition, the sample size of this study significantly exceeds that of previous studies, , providing a broader understanding of COS variations during CAT.

However, it is noteworthy that this study included a higher proportion of female patients, which suggests that women are more concerned about esthetics and orthodontic appliance selection. In addition, the sample size of the hyperdivergent subgroup was notably smaller, primarily because these patients typically gain significant vertical control through extraction and are, therefore, more likely to opt for extraction. Consequently, our findings may not fully represent the effects of Invisalign treatment. Furthermore, the training dataset did not include dental models representing primary or mixed dentition, tooth loss, or dental wear. Therefore, future studies should focus on developing more precise multimodal models and investigating changes in clinical outcomes using larger and more diverse samples.

Altogether, the effectiveness of Invisalign in flattening the COS varies among patients with different vertical skeletal patterns. Clinicians should design treatment plans tailored to the unique skeletal characteristics of each patient and set realistic expectations accordingly. The initial COS depth significantly influences outcome predictability, necessitating more personalized intervention strategies for hypodivergent patients. Furthermore, special attention should be paid to the lower extrusion rate of the first molar to prevent adverse impacts on occlusion and enhance overall occlusal health.

Only gold members can continue reading. Log In or Register to continue

Stay updated, free articles. Join our Telegram channel

Jun 27, 2026 | Posted by in CARDIOLOGY | Comments Off on Clinical advances in curve of Spee assessment: Deep learning for automatic tooth landmark detection in Invisalign

Full access? Get Clinical Tree

Get Clinical Tree app for offline access