Development and internal validation of a multimodal MRI-FNAC radiomics-pathomics model for predicting cervical lymph node metastasis in papillary thyroid carcinoma
Highlight box
Key findings
• The multimodal magnetic resonance imaging (MRI)-fine-needle aspiration cytology (FNAC) radiomics-pathomics model based on ResNet-50 deep features significantly outperformed single-modality models. In the independent test set, the multimodal k-nearest neighbors model achieved the best performance, with an area under the curve (AUC) of 0.872 (95% confidence interval: 0.778–0.955), accuracy of 82.8%, specificity of 92.3%, and positive predictive value of 93.1%. DeLong tests confirmed statistical superiority over single-modality models, and the model demonstrated good calibration (Brier score =0.1059). Decision curve analysis confirmed its superior clinical net benefit.
What is known and what is new?
• It is well known that cervical lymph node metastasis (CLNM) significantly affects surgical strategy and prognosis in papillary thyroid carcinoma (PTC), yet accurate preoperative prediction remains challenging.
• This study integrates preoperative MRI radiomics with FNAC-based pathomics using deep learning features and demonstrates that the multimodal fusion model substantially improves predictive performance compared with MRI-only or pathomics-only approaches.
What is the implication, and what should change now?
• This non-invasive multimodal model shows promise in assisting thyroid surgeons with individualized decision-making regarding neck lymph node dissection. However, given that this was a single-center retrospective study with internal validation only, further prospective multicenter external validation is required before the model can be considered for clinical implementation. If validated, it may help reduce unnecessary prophylactic neck dissections and associated complications such as recurrent laryngeal nerve injury and hypoparathyroidism.
Introduction
Papillary thyroid carcinoma (PTC) is the most common histological subtype of thyroid cancer, accounting for approximately 80–90% of all thyroid malignancies (1). Cervical lymph node metastasis (CLNM) occurs in 40–60% of patients with PTC and represents a critical factor influencing local recurrence, distant metastasis, and overall prognosis (2,3). Therefore, accurate preoperative identification of CLNM is essential for determining the extent of surgery and optimizing individualized treatment strategies. Current preoperative assessment of lymph node status primarily depends on ultrasound and ultrasound-guided fine-needle aspiration cytology (FNAC) (4). Although widely used, these modalities have well-recognized limitations, including operator dependence and a relatively high false-negative rate, which can result in missed metastases. To avoid undertreatment, many surgeons perform prophylactic central neck lymph node dissection even in clinically node-negative patients (5). However, the benefit of routine prophylactic dissection remains controversial because it is associated with increased risks of recurrent laryngeal nerve injury, hypoparathyroidism, and prolonged operative time, while its impact on long-term oncologic outcomes is still debated. These challenges highlight the urgent need for more accurate, non-invasive preoperative tools to better stratify the risk of CLNM and guide surgical decision-making. In recent years, artificial intelligence and radiomics have been increasingly applied to thyroid cancer imaging. Multiple studies have demonstrated the potential of magnetic resonance imaging (MRI)-based radiomics for predicting CLNM, with reported area under the curve (AUC) values generally ranging from 0.75 to 0.85. Ultrasound radiomics and deep learning models have similarly achieved AUCs between 0.78 and 0.88 across different cohorts (6,7). However, most existing single-modality models capture only partial aspects of tumor biology. MRI radiomics can quantify macroscopic tumor morphology, enhancement heterogeneity, and spatial relationships but lacks direct microscopic pathological correlation (8,9). In contrast, pathomics derived from FNAC images enables quantitative analysis of nuclear atypia, chromatin texture, and cellular spatial architecture at the microscopic level, which may better reflect the biological aggressiveness and metastatic potential of tumor cells (10,11). The fusion of these two complementary scales—macroscopic imaging phenotypes from MRI and microscopic cellular features from FNAC—has the potential to provide a more comprehensive and biologically meaningful assessment of tumor invasiveness than either approach alone (12,13). Recent multimodal radiomics and pathomics studies have further demonstrated improved predictive performance in lymph node metastasis prediction (14,15). We hypothesized that a multimodal radiomics-pathomics model fusing deep features from preoperative MRI and FNAC images would achieve superior predictive performance (AUC >0.85 in the independent test set) compared with single-modality models by integrating complementary multi-scale tumor information. To identify the optimal modeling approach for this fused feature space, we systematically evaluated six classical machine learning algorithms—logistic regression (LR), support vector machine (SVM), Naive Bayes, k-nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGBoost). This comprehensive comparison was conducted because the best-performing algorithm often differs depending on the characteristics of radiomics versus pathomics features, as evidenced in recent machine learning studies on medical image-based prediction tasks (16).
In this study, we aimed to develop and internally validate a multimodal MRI-FNAC radiomics-pathomics model based on deep learning features for predicting CLNM in patients with PTC. We present this article in accordance with the TRIPOD reporting checklist (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0241/rc).
Methods
Patients and dataset
A total of 128 patients with PTC who underwent thyroidectomy and bilateral neck lymph node dissection at The First People’s Hospital of Zunyi (The Third Affiliated Hospital of Zunyi Medical University) between January 2021 and December 2025 were retrospectively enrolled. Patients were randomly divided into a training set (n=90) and a test set (n=38) at a 7:3 ratio. According to postoperative pathological results, patients were classified into CLNM-positive and CLNM-negative groups. Inclusion criteria were: (I) completion of MRI examination within 2 weeks before surgery; (II) preoperative ultrasound-guided FNAC of the thyroid classified as Bethesda category ≥ IV (using the 2017 Bethesda System for Reporting Thyroid Cytopathology) (17); and (III) availability of high-quality cytological images at ×400 magnification. Exclusion criteria were: (I) previous history of thyroid or neck surgery; (II) thyroid nodule diameter (mean of long and short diameters) <5 mm; and (III) poor quality of MRI or FNAC images. The diagnostic gold standard for CLNM was definitive postoperative histopathological examination of all resected lymph nodes from central and/or lateral neck dissection specimens, performed by experienced pathologists according to the 8th edition of the AJCC staging system. The sample size was determined by the available retrospective cohort meeting the inclusion criteria during the study period. With 78 patients having CLNM (events) and 16 final selected features after least absolute shrinkage and selection operator (LASSO) regression, the events-per-predictor ratio was approximately 4.9, which meets commonly recommended minimum thresholds for prediction modeling in radiomics and machine learning studies. No formal prospective power analysis was performed due to the retrospective nature of the study. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of The First People’s Hospital of Zunyi (The Third Affiliated Hospital of Zunyi Medical University) [Approval No. (2025)-1-341]. The requirement for informed consent was waived due to the retrospective nature of the study.
Multimodal data acquisition
All MRI examinations were performed using a Siemens MAGNETOM Vida 3.0 T scanner with a 64-channel neck coil. Scanning parameters were as follows: T1WI (TR =5.0 ms, TE =1.86 ms, FOV read =200 mm, FOV phase =100%, matrix =224×224, slice thickness =1.5 mm, number of excitations =1, scan time =2 min 24 s); fat-suppressed T2WI (TR =3,100 ms, TE =77 ms, FOV read =200 mm, FOV phase =87.5%, matrix =206×336, slice thickness =3 mm, number of excitations =3, scan time =2 min 51 s). Contrast-enhanced T1-weighted imaging (CE-T1WI) was acquired using the T1-VIBE sequence (TR =4.5 ms, TE =1.9 ms, FOV =200×200 mm, matrix =256×256, slice thickness =3.0 mm, no interslice gap, flip angle =12°, parallel acceleration factor =3). Single-phase scan time was approximately 18–22 s. Gadopentetate dimeglumine (Gd-DTPA, Bayer Schering Pharma, Germany) was administered intravenously at a dose of 0.1 mmol/kg and a rate of 3.0 mL/s, followed by a 15 mL saline flush. FNAC images were obtained from hematoxylin and eosin (H&E)-stained cytological slides reviewed by experienced pathologists. Representative tumor cell regions were selected under a fluorescence microscope. All images were captured at ×400 magnification with a resolution of 1,000×1,000 pixels and saved in TIFF format.
Feature extraction and selection
Deep features were extracted using a pretrained ResNet-50 model on the ImageNet dataset. To ensure a consistent 1,024-dimensional output from the last convolutional block before global average pooling, features were primarily extracted from layer 3 of the ResNet-50 architecture (1,024 channels). When layer 4 (2,048 channels) was used, channel reduction was applied via global average pooling to maintain dimensional consistency. All model parameters were frozen during extraction, with no fine-tuning performed. For MRI deep feature extraction, regions of interest (ROIs) were manually delineated on the arterial phase of contrast-enhanced T1-weighted imaging (CE-T1WI) using ITK-SNAP software (version 3.8.0) by a radiologist with more than 3 years of experience. Delineations were independently reviewed and corrected if necessary by a senior radiologist. Inter-observer reproducibility was evaluated on 30 randomly selected nodules, yielding intraclass correlation coefficients (ICC) >0.85 for all features. The ROI was rigidly registered and propagated to the fat-suppressed T2-weighted imaging (T2WI) sequence, with manual verification and adjustments to ensure spatial consistency. After 0–1 normalization of the mask, the three-dimensional lesion volume was sliced along the Z-axis. Features extracted from all slices using the pretrained ResNet-50 model were averaged to obtain a 2×1024-dimensional MRI deep feature vector per patient (denoted as M). These deep features capture high-level semantic representations of tumor morphology and heterogeneity but were not directly interpreted in a clinical radiological sense; instead, they were used as input for downstream machine learning models. For FNAC deep feature extraction, FNAC images were first converted to grayscale, followed by color deconvolution in QuPath software (version 0.4.4) to separate the H&E staining channels. Tumor cell regions were automatically segmented using QuPath’s nuclear segmentation tool with the following parameters: nuclear detection threshold =0.2, minimum nuclear area =20 µm2, maximum nuclear area =200 µm2, and sigma =1.5. Segmented non-overlapping image patches (256×256 pixels) were input into the pretrained ResNet-50 model, and features from all patches were averaged to generate a 1×1024-dimensional FNAC deep feature vector per patient (denoted as P). These pathomics features quantify microscopic nuclear and cellular characteristics that may reflect tumor biological behavior. For feature fusion and selection, the MRI (M) and FNAC (P) feature vectors were concatenated to form a multimodal deep feature vector (denoted as Z). To prevent data leakage, feature selection was performed exclusively within the training set (n=90) using a nested 5-fold stratified cross-validation framework. Univariate analysis (independent t-test or Mann-Whitney U test, P<0.05) was first applied to screen candidate features, followed by LASSO regression with L1 regularization. The optimal penalty parameter λ was determined via inner-loop 5-fold cross-validation. The selected features were then fixed and directly applied to the independent test set (n=38) without refitting.
Following nested cross-validation, 16 features were retained (11 from MRI and 5 from FNAC; range across folds: 14–18) for model construction. These features were standardized using Z-score normalization before being input into the machine learning classifiers.
Machine learning model construction and evaluation
Six machine learning algorithms—LR, SVM, Naive Bayes, KNN, RF, and XGBoost—were employed to construct classification models based on MRI features alone, FNAC pathomics features alone, and multimodal fusion features, respectively. Models were trained and internally validated using 5-fold stratified cross-validation in the training set. Independent validation was performed in the test set. Model performance was evaluated using the AUC, accuracy (ACC), sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Pairwise comparisons of AUCs between the best-performing models were performed using the DeLong test. Model calibration was assessed using calibration curves, Brier scores, and the Hosmer-Lemeshow test. Decision curve analysis (DCA) was conducted to quantify the clinical net benefit of the models across different threshold probabilities. Clinically acceptable performance thresholds (e.g., sensitivity >70% and specificity >80% for safely avoiding prophylactic neck dissection in low-risk patients) were discussed in the context of existing literature but were not predefined as formal decision thresholds in this study. The overall study workflow is illustrated in Figure 1.
Statistical analysis
All statistical analyses were performed using Python (version 3.8.3). Normally distributed continuous variables are presented as mean ± standard deviation and compared using independent samples t-tests. Non-normally distributed data are expressed as median (interquartile range) and compared using the chi-square test or Fisher’s exact test, as appropriate. Categorical data are presented as number (percentage). The ICC was used to evaluate the reproducibility of feature extraction. Receiver operating characteristic (ROC) curves were plotted to assess model performance. DeLong tests were used for statistical comparison of AUCs between models. Calibration was evaluated using Brier scores and Hosmer-Lemeshow tests. DCA was used to evaluate clinical utility. A two-sided P value < 0.05 was considered statistically significant.
Results
Clinical characteristics
A total of 128 patients with PTC were included in this study, with 90 patients in the training set and 38 patients in the test set. The mean age was 48.82±10.02 years (range, 24–75 years), and there were 91 females (71.1%) and 37 males (28.9%). According to postoperative histopathology, 78 patients (60.9%) had CLNM and 50 patients (39.1%) did not. No significant differences were observed between the CLNM-positive and CLNM-negative groups in age, sex, or maximum tumor diameter in either the training or test cohorts (all P>0.05). However, multifocality and extrathyroidal extension were significantly more prevalent in the CLNM-positive group (training set: multifocality P=0.004, extrathyroidal extension P=0.02; test set: multifocality P=0.01, extrathyroidal extension P=0.03; Table 1). These findings are consistent with established clinicopathological risk factors for CLNM in PTC.
Table 1
| Features | Training cohort (n=90) | Validation cohort (n=38) | |||||
|---|---|---|---|---|---|---|---|
| CLNM (+) (n=56) | CLNM (−) (n=34) | P | CLNM (+) (n=22) | CLNM (−) (n=16) | P | ||
| Age (years) | 48.35±10.72 | 50.83±9.29 | 0.27 | 47.88±9.12 | 45.26±9.46 | 0.40 | |
| Maximum tumor diameter (cm) | 1.41±0.68 | 1.34±0.72 | 0.64 | 1.36±0.81 | 1.27±0.72 | 0.73 | |
| Gender | 0.78 | >0.99 | |||||
| Male | 16 | 8 | 5 | 3 | |||
| Female | 40 | 26 | 17 | 13 | |||
| Number of tumor | 0.10 | 0.35 | |||||
| Solitary | 24 | 21 | 9 | 9 | |||
| Multifocal | 32 | 13 | 13 | 7 | |||
| Extrathyroidal extension | 0.02 | 0.03 | |||||
| No | 17 | 20 | 6 | 11 | |||
| Yes | 39 | 14 | 16 | 5 | |||
Data are presented as mean ± standard deviation or number. CLNM, cervical lymph node metastasis; PTC, papillary thyroid carcinoma.
Construction and performance of radiomics-pathomics models
A total of 2,048 MRI deep features and 1,024 pathomics deep features were initially extracted. After feature selection within the training set, 11 MRI features and 5 pathomics features were retained for model construction. Six machine learning algorithms were used to build prediction models based on MRI features alone, pathomics features alone, and multimodal fusion features. The MRI-only models showed moderate performance, with AUC values ranging from 0.626 to 0.711 (mean AUC, 0.695); the Naive Bayes classifier performed best among them. The pathomics-only models outperformed the MRI-only models, with a mean AUC of 0.771 (range, 0.758–0.799); the XGBoost model achieved the best performance in this category. The multimodal fusion models significantly outperformed the single-modality models, with AUC values generally above 0.81 (range, 0.811–0.872; mean AUC, 0.833). In the independent test set, the multimodal KNN model demonstrated the best balanced and superior performance, with an AUC of 0.872 [95% confidence interval (CI): 0.778–0.955], accuracy of 82.8%, sensitivity of 75.0%, specificity of 92.3%, PPV of 93.1%, and NPV of 74.3% (Figure 2). To statistically compare the AUCs of the best-performing models across modalities, DeLong tests were performed on the test set. The multimodal KNN model showed significantly higher AUC than the best pathomics-only model (XGBoost; P=0.046) and the best MRI-only model (Naive Bayes; P=0.003). In addition, model calibration was evaluated using calibration curves, Brier scores, and the Hosmer-Lemeshow test. The multimodal KNN model exhibited the best calibration performance, with the lowest Brier score (0.1059) and a non-significant Hosmer-Lemeshow P value (0.72), indicating good agreement between predicted probabilities and observed outcomes. Calibration curves for the top-performing models in each modality are presented in Figure 3. The detailed diagnostic performance of all models in the training and test sets is shown in Tables 2,3.
Table 2
| Model | AUC (95% CI) | Accuracy (%) | Sensitivity (%) | Specificity (%) | PPV (%) | NPV (%) |
|---|---|---|---|---|---|---|
| MRI | ||||||
| LR | 0.771 (0.549–0.854) | 70.1 | 72.9 | 67.3 | 74.5 | 68.1 |
| SVM | 0.761 (0.456–0.848) | 64.0 | 67.9 | 61.3 | 73.6 | 66.5 |
| Naive Bayes | 0.812 (0.702–0.889) | 76.8 | 71.2 | 82.4 | 80.3 | 73.9 |
| KNN | 0.804 (0.758–0.851) | 71.9 | 72.3 | 71.4 | 76.4 | 66.9 |
| Random forest | 0.833 (0.765–0.894) | 78.2 | 76.5 | 80.1 | 82.7 | 72.8 |
| XGBoost | 0.789 (0.663–0.852) | 73.5 | 79.3 | 67.8 | 75.1 | 72.3 |
| Pathology | ||||||
| LR | 0.856 (0.710–0.932) | 79.9 | 75.4 | 86.0 | 88.5 | 76.0 |
| SVM | 0.871 (0.812–0.918) | 81.7 | 74.8 | 89.2 | 89.6 | 74.5 |
| Naive Bayes | 0.847 (0.792–0.901) | 78.4 | 72.1 | 85.3 | 84.9 | 73.2 |
| KNN | 0.916 (0.884–0.949) | 86.0 | 79.2 | 94.7 | 95.2 | 77.9 |
| Random forest | 0.902 (0.851–0.935) | 84.1 | 78.6 | 90.3 | 91.8 | 76.4 |
| XGBoost | 0.923 (0.876–0.961) | 85.3 | 82.4 | 88.7 | 90.2 | 78.1 |
| MRI + pathology | ||||||
| LR | 0.969 (0.950–0.998) | 89.8 | 87.5 | 92.9 | 94.2 | 85.5 |
| SVM | 0.952 (0.912–0.972) | 87.4 | 81.2 | 94.1 | 94.8 | 80.3 |
| Bayes | 0.934 (0.834–0.967) | 87.6 | 83.6 | 92.7 | 93.5 | 83.9 |
| KNN | 0.988 (0.978–0.997) | 90.6 | 86.8 | 95.6 | 96.2 | 85.0 |
| Random forest | 0.961 (0.924–0.983) | 88.9 | 84.7 | 93.8 | 94.6 | 82.7 |
| XGBoost | 0.975 (0.948–0.989) | 89.2 | 85.3 | 93.5 | 94.3 | 83.6 |
AUC, area under the curve; CI, confidence interval; CLNM, cervical lymph node metastasis; KNN, k-nearest neighbors; LR, logistic regression; MRI, magnetic resonance imaging; NPV, negative predictive value; PPV, positive predictive value; PTC, papillary thyroid carcinoma; SVM, support vector machine; XGBoost, extreme gradient boosting.
Table 3
| Model | AUC (95% CI) | Accuracy (%) | Sensitivity (%) | Specificity (%) | PPV (%) | NPV (%) |
|---|---|---|---|---|---|---|
| MRI | ||||||
| LR | 0.689 (0.594–0.830) | 70.3 | 78.6 | 60.7 | 71.8 | 68.1 |
| SVM | 0.680 (0.586–0.781) | 65.6 | 72.2 | 57.1 | 68.4 | 61.5 |
| Naive Bayes | 0.731 (0.595–0.837) | 70.3 | 77.9 | 60.7 | 71.8 | 68.4 |
| KNN | 0.702 (0.597–0.829) | 71.9 | 80.2 | 60.7 | 72.5 | 70.8 |
| Random forest | 0.698 (0.628–0.803) | 67.2 | 75.0 | 58.3 | 69.2 | 64.0 |
| XGBoost | 0.667 (0.592–0.764) | 62.9 | 66.7 | 59.6 | 64.8 | 60.6 |
| Pathology | ||||||
| LR | 0.758 (0.632–0.886) | 73.4 | 70.7 | 75.0 | 78.8 | 67.7 |
| SVM | 0.763 (0.642–0.887) | 76.6 | 73.0 | 82.1 | 83.9 | 69.7 |
| Naive Bayes | 0.762 (0.642–0.886) | 74.3 | 72.6 | 92.8 | 92.9 | 72.2 |
| KNN | 0.765 (0.647–0.889) | 71.9 | 73.5 | 71.4 | 76.5 | 66.7 |
| Random forest | 0.779 (0.659–0.899) | 70.3 | 74.1 | 67.9 | 74.3 | 65.5 |
| XGBoost | 0.799 (0.679–0.914) | 75.0 | 75.0 | 75.9 | 79.4 | 70.0 |
| MRI + pathology | ||||||
| LR | 0.839 (0.738–0.936) | 78.8 | 77.8 | 71.4 | 77.8 | 71.4 |
| SVM | 0.811 (0.776–0.961) | 71.9 | 75.0 | 67.7 | 75.0 | 67.9 |
| Bayes | 0.819 (0.702–0.932) | 72.1 | 77.8 | 78.6 | 82.4 | 73.3 |
| KNN | 0.872 (0.778–0.955) | 82.8 | 75.0 | 92.3 | 93.1 | 74.3 |
| Random forest | 0.825 (0.728–0.923) | 74.5 | 76.0 | 64.3 | 72.9 | 66.7 |
| XGBoost | 0.831 (0.733–0.923) | 76.3 | 72.2 | 69.7 | 74.3 | 65.5 |
DeLong test for key pairwise AUC comparisons (test set); Multimodal KNN vs. Pathomics XGBoost: P=0.046; Multimodal KNN vs. MRI Naive Bayes: P=0.003. AUC, area under the curve; CI, confidence interval; CLNM, cervical lymph node metastasis; KNN, k-nearest neighbors; LR, logistic regression; MRI, magnetic resonance imaging; NPV, negative predictive value; PPV, positive predictive value; PTC, papillary thyroid carcinoma; SVM, support vector machine; XGBoost, extreme gradient boosting.
DCA
DCA in the test set (Figure 4) demonstrated that the net benefit curve of the multimodal KNN model was higher than those of the MRI-only and pathomics-only models across the clinically relevant threshold probability range of 0.05–0.60. The advantage was particularly evident in the 0.1–0.5 threshold range, indicating superior clinical net benefit. These findings suggest that the multimodal radiomics-pathomics model provides greater clinical utility for guiding surgical decision-making regarding CLNM in PTC patients.
Discussion
Lymph node metastasis is a key factor influencing local recurrence and distant metastasis in tumors. CLNM is particularly important for the prognosis of PTC patients. In this study, we developed MRI-based, FNAC pathomics-based, and multimodal fusion models using six machine learning classifiers by integrating preoperative MRI (T2WI and CE-T1WI) deep radiomics features with ResNet-50-extracted features from FNAC cytological images. The multimodal fusion models significantly outperformed single-modality models, with the KNN model achieving the highest AUC of 0.872 in the independent test set. DeLong tests confirmed that this performance was statistically superior to the best single-modality models (P=0.046 vs. pathomics XGBoost; P=0.003 vs. MRI Naive Bayes). DCA further confirmed its superior clinical net benefit. These results indicate that deep feature-level fusion of MRI radiomics and FNAC pathomics can substantially enhance preoperative prediction accuracy of CLNM in PTC and support more individualized surgical decision-making.
Consistent with previous studies, multifocality and extrathyroidal extension were significantly associated with CLNM. However, these features are typically confirmed only after surgery and cannot be reliably assessed preoperatively. The present study therefore focused exclusively on preoperative MRI and FNAC data to develop a fully non-invasive multimodal model.
Although ultrasound remains the preferred imaging modality for evaluating PTC (18), its performance for lymph node assessment is influenced by operator experience and anatomical factors (19,20). MRI offers superior soft-tissue resolution and multiplanar imaging, demonstrating higher sensitivity for detecting central compartment lymph node metastases (21). Previous MRI radiomics studies have reported promising results; however, the mean AUC of the MRI-only model in the present study was 0.695, possibly due to the use of only two sequences and deep features rather than handcrafted radiomics. Compared with single-modality studies (22), our integration of MRI with FNAC pathomics leverages MRI’s advantages in deeper anatomical assessment and achieved a higher test-set AUC (0.872), highlighting the benefit of deep feature fusion.
Notably, the optimal machine learning algorithm varied across modalities: Naive Bayes performed best for MRI features, XGBoost for pathomics features, and KNN for multimodal fusion. This variation likely reflects the distinct characteristics of each feature set and underscores the importance of model-data matching in multimodal radiomics-pathomics research (23-25).
This study has several important limitations that must be acknowledged. First, it was a single-center retrospective study with a relatively modest sample size (n=128 total; test set n=38). Although feature selection was performed exclusively within the training set using a nested 5-fold cross-validation framework to mitigate overfitting risk, and the final multimodal model retained only 16 features, the possibility of overfitting cannot be completely excluded when using high-dimensional deep features extracted from a frozen ResNet-50 model. Second, the random 7:3 split constitutes internal validation only; the test set is not geographically, temporally, or institutionally independent from the training set. Therefore, the generalizability of our findings to other centers, imaging protocols, or patient populations cannot be assumed and requires confirmation through prospective multicenter external validation studies. Third, while deep features from a frozen ResNet-50 reduce subjectivity compared with traditional handcrafted radiomics, their direct clinical interpretability remains limited; future studies should incorporate explainable AI techniques such as SHAP or Grad-CAM to improve biological understanding of the selected features. Fourth, the quality of FNAC images may vary with cellularity, smear preparation, and staining consistency, potentially affecting pathomics feature stability. Finally, although multifocality and extrathyroidal extension were associated with CLNM, these features are typically confirmed only postoperatively and were therefore not incorporated into the preoperative prediction model. These limitations are inherent to the current single-center retrospective design and indicate that, while the multimodal approach shows promise, the model should not yet be considered ready for routine clinical use without further rigorous external validation.
Future research with larger cohorts, multicenter validation, and multi-omics integration is warranted to improve model stability and clinical translational value. Prospective studies incorporating temporal or geographic external validation cohorts will be essential to establish the true generalizability of multimodal MRI-FNAC radiomics-pathomics models in PTC.
Conclusions
This study developed and internally validated a multimodal radiomics-pathomics model based on preoperative MRI and FNAC for predicting CLNM in patients with PTC. The multimodal fusion model significantly outperformed single-modality models, with the KNN model demonstrating the best performance in the independent test set (AUC =0.872). DeLong tests confirmed statistical superiority over the best single-modality models, and the multimodal model also showed good calibration (Brier score =0.1059). However, this was a single-center retrospective study with a relatively small sample size and internal validation only. The generalizability of the model to other institutions and populations remains to be confirmed through prospective multicenter external validation. Therefore, while the multimodal MRI-FNAC radiomics-pathomics approach shows promise as a non-invasive tool to assist preoperative risk stratification, it should not yet be considered ready for routine clinical implementation without further rigorous validation. Future multicenter studies with larger cohorts and external validation are essential to establish the clinical utility and generalizability of this model before it can be translated into surgical decision-making for patients with PTC.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0241/rc
Data Sharing Statement: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0241/dss
Peer Review File: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0241/prf
Funding: This work was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0241/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of The First People’s Hospital of Zunyi (The Third Affiliated Hospital of Zunyi Medical University) [Approval No. (2025)-1-341]. The requirement for informed consent was waived due to the retrospective nature of the study.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Zeng X, Wang Z, Gui Z, et al. High Incidence of Distant Metastasis Is Associated With Histopathological Subtype of Pediatric Papillary Thyroid Cancer - a Retrospective Analysis Based on SEER. Front Endocrinol (Lausanne) 2021;12:760901. [Crossref] [PubMed]
- Wu S, Liu Y, Ruan X, et al. Predictive factors for lymph node metastasis in papillary thyroid cancer patients undergoing neck dissection: insights from a large cohort study. Front Oncol 2024;14:1447903. [Crossref] [PubMed]
- Li M, Dal Maso L, Pizzato M, et al. Thyroid cancer in adolescents and young adults: a population-based study in 185 countries worldwide. Lancet Diabetes Endocrinol 2026;14:112-22. [Crossref] [PubMed]
- Oblak T, Perhavec A, Hocevar M, et al. Reduction of overtreatment without reduction of overdiagnosis in patients with differentiated thyroid cancer: mission impossible. Langenbecks Arch Surg 2021;406:2011-7. [Crossref] [PubMed]
- Salem FA, Bergenfelz A, Nordenström E, et al. Central lymph node dissection and permanent hypoparathyroidism after total thyroidectomy for papillary thyroid cancer: population-based study. Br J Surg 2021;108:684-90. [Crossref] [PubMed]
- Vasse T, Alhiyari Y, Evans LK, et al. Machine-learning-based tumor segmentation and classification using dynamic optical contrast imaging for thyroid cancer. Biophotonics Discov 2026;3:015001. [Crossref] [PubMed]
- Thomas J, Tessler FN. Artificial Intelligence Applications in Thyroid Cancer Diagnosis: 2026 Update. Thyroid 2026;36:133-40. [Crossref] [PubMed]
- Wang H, Song B, Ye N, et al. Machine learning-based multiparametric MRI radiomics for predicting the aggressiveness of papillary thyroid carcinoma. Eur J Radiol 2020;122:108755. [Crossref] [PubMed]
- Barzegar-Golmoghani E, Mohebi M, Gohari Z, et al. ELTIRADS framework for thyroid nodule classification integrating elastography, TIRADS, and radiomics with interpretable machine learning. Sci Rep 2025;15:8763. [Crossref] [PubMed]
- Zheng X, Yao Z, Huang Y, et al. Deep learning radiomics can predict axillary lymph node status in early-stage breast cancer. Nat Commun 2020;11:1236. [Crossref] [PubMed]
- Campanella G, Hanna MG, Geneslaw L, et al. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat Med 2019;25:1301-9. [Crossref] [PubMed]
- Li C, Chen X, Chen C, et al. Application of deep learning radiomics in oral squamous cell carcinoma-Extracting more information from medical images using advanced feature analysis. J Stomatol Oral Maxillofac Surg 2024;125:101840. [Crossref] [PubMed]
- Xiao W, Zhou W, Yuan H, et al. A radiopathomics model for predicting large-number cervical lymph node metastasis in clinical N0 papillary thyroid carcinoma. Eur Radiol 2025;35:4587-98. [Crossref] [PubMed]
- Wang W, Jin F, Song L, et al. Prediction of peripheral lymph node metastasis (LNM) in thyroid cancer using delta radiomics derived from enhanced CT combined with multiple machine learning algorithms. Eur J Med Res 2025;30:164. [Crossref] [PubMed]
- Li G, Yao J, Peng C, et al. Multimodal Nested Attention Network for Lymph Node Metastasis Prediction of Thyroid Carcinoma. Big Data Mining and Analytics 2026;9:178-97.
- Zhou L, Lu WP, Zhang HL, et al. Development and validation of an interpretable machine learning model for predicting central lymph node metastasis in papillary thyroid cancer. Front Oncol 2026;16:1839870. [Crossref] [PubMed]
- Cibas ES, Ali SZ. The 2017 Bethesda System for Reporting Thyroid Cytopathology. Thyroid 2017;27:1341-6. [Crossref] [PubMed]
- Scappaticcio L, Di Martino N, Caruso P, et al. The value of ACR, European, Korean, and ATA ultrasound risk stratification systems combined with RAS mutations for detecting thyroid carcinoma in cytologically indeterminate and suspicious for malignancy thyroid nodules. Hormones (Athens) 2024;23:687-97. [Crossref] [PubMed]
- Xue T, Liu C, Liu JJ, et al. Analysis of the Relevance of the Ultrasonographic Features of Papillary Thyroid Carcinoma and Cervical Lymph Node Metastasis on Conventional and Contrast-Enhanced Ultrasonography. Front Oncol 2021;11:794399. [Crossref] [PubMed]
- Lu C, Wang Y, Yu M. Is ultrasonographic evaluation sensitive enough to detect multicentric papillary thyroid carcinoma? Gland Surg 2020;9:737-46. [Crossref] [PubMed]
- Liu Z, Xun X, Wang Y, et al. MRI and ultrasonography detection of cervical lymph node metastases in differentiated thyroid carcinoma before reoperation. Am J Transl Res 2014;6:147-54.
- Qin H, Que Q, Lin P, et al. Magnetic resonance imaging (MRI) radiomics of papillary thyroid cancer (PTC): a comparison of predictive performance of multiple classifiers modeling to identify cervical lymph node metastases before surgery. Radiol Med 2021;126:1312-27. [Crossref] [PubMed]
- Zhang J, Hao L, Xu Q, et al. Radiomics and Clinical Characters Based Gaussian Naive Bayes (GNB) Model for Preoperative Differentiation of Pulmonary Pure Invasive Mucinous Adenocarcinoma From Mixed Mucinous Adenocarcinoma. Technol Cancer Res Treat 2024;23:15330338241258415. [Crossref] [PubMed]
- Shi Y, Zou Y, Liu J, et al. Ultrasound-based radiomics XGBoost model to assess the risk of central cervical lymph node metastasis in patients with papillary thyroid carcinoma: Individual application of SHAP. Front Oncol 2022;12:897596. [Crossref] [PubMed]
- Feng JW, Yang YX, Qin RJ, et al. Application and validation of the machine learning-based multimodal radiomics model for preoperative prediction of lateral lymph node metastasis in papillary thyroid carcinoma. Front Endocrinol (Lausanne) 2025;16:1618902. [Crossref] [PubMed]

