Clinical transportability, calibration drift, and local updating of a B-mode ultrasound radiomics-clinical prediction model for occult high-volume central lymph node metastasis in cN0 papillary thyroid carcinoma: development, temporal validation, and external validation
Original Article

Clinical transportability, calibration drift, and local updating of a B-mode ultrasound radiomics-clinical prediction model for occult high-volume central lymph node metastasis in cN0 papillary thyroid carcinoma: development, temporal validation, and external validation

Ziwei Zhang1#, Cailing Lin2#, Jing Ning1#, Xiaochen Liu1, Chenshan Dong1, Hang Ling3

1Department of Ultrasonography, Fujian Provincial Hospital, Fuzhou University Affiliated Fujian Provincial Hospital, Shengli Clinical Medical College of Fujian Medical University, Fuzhou, China; 2Department of Breast Surgery, Fujian Provincial Hospital, Fuzhou University Affiliated Fujian Provincial Hospital, Shengli Clinical Medical College of Fujian Medical University, Fuzhou, China; 3Department of Pathology, Fujian Provincial Hospital, Fuzhou University Affiliated Fujian Provincial Hospital, Shengli Clinical Medical College of Fujian Medical University, Fuzhou, China

Contributions: (I) Conception and design: Z Zhang, J Ning, H Ling; (II) Administrative support: H Ling, C Dong; (III) Provision of study materials or patients: C Lin, X Liu, C Dong; (IV) Collection and assembly of data: Z Zhang, C Lin, J Ning, X Liu; (V) Data analysis and interpretation: Z Zhang, J Ning, H Ling; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work.

Correspondence to: Hang Ling, PhD. Department of Pathology, Fujian Provincial Hospital, Fuzhou University Affiliated Fujian Provincial Hospital, Shengli Clinical Medical College of Fujian Medical University, No. 134 Dongjie, Fuzhou 350001, China. Email: ziwei_zhang2009@163.com.

Background: Occult high-volume central lymph node metastasis (CLNM) can remain undetected in clinically node-negative papillary thyroid carcinoma (cN0 PTC), yet nodal burden influences preoperative counseling and recurrence-oriented risk stratification. For clinical implementation, a preoperative radiomics-based risk estimate must remain transportable, locally calibrated, and interpretable after transfer to another center. This study evaluated the clinical transportability, calibration drift, threshold consequences, and local updating of a B-mode ultrasound radiomics-clinical prediction model for occult high-volume CLNM.

Methods: This three-center retrospective model-development, validation, and updating study used the source workbook containing 957 cN0 PTC patients. The endpoint was occult high-volume CLNM, defined as more than five metastatic central lymph nodes on postoperative pathology. Candidate predictors were restricted to preoperative clinical-ultrasound variables and a prespecified B-mode radiomics score. Clinical-only, B-mode score-only, and clinical plus B-mode score models were compared in the development cohort; the clinical plus B-mode score model was treated as the locked primary implementation model because it preserved routine clinical interpretability while incorporating the quantitative image signal. The locked model was evaluated in a temporal Center 1 cohort (n=210), Center 2 (n=139), and Center 3 (n=119). Discrimination, Brier score, calibration, fixed-threshold classification, decision curve analysis, threshold consequences, and post hoc intercept-only recalibration were assessed.

Results: High-volume CLNM was present in 399 of 957 patients (41.7%). Event rates were 44.8%, 45.2%, 41.0%, and 23.5% in the development, temporal validation, Center 2, and Center 3 cohorts. Development-cohort comparisons showed larger tumors, higher B-mode radiomics scores, higher Chinese Thyroid Imaging Reporting and Data System (C-TIRADS) subclass, and richer vascularity among endpoint-positive patients. The B-mode score accounted for most discrimination. Compared with the B-mode score-only model, adding clinical-ultrasound variables changed the area under the receiver operating characteristic curve (AUC) by +0.015 in development, +0.008 in temporal validation, −0.009 in Center 2, and −0.029 in Center 3. The locked clinical plus B-mode score model had AUC values of 0.823, 0.874, 0.866, and 0.870, respectively. Mean predicted risk aligned with observed risk in the development, temporal validation, and Center 2 cohorts but overestimated risk in Center 3 (36.4% predicted vs. 23.5% observed; calibration intercept, −0.902). At the locked threshold of 0.565, Center 3 retained specificity of 0.857 and negative predictive value (NPV) of 0.897, while intercept-only updating reduced the Center 3 Brier score from 0.138 to 0.121.

Conclusions: The locked radiomics-clinical model provided an interpretable preoperative risk-stratification framework whose deployment value depends on local calibration and threshold consequences. External deployment should evaluate spectrum shift, positive predictive value (PPV) and NPV, calibration, decision curves, and local updating rather than relying on AUC alone.

Keywords: Papillary thyroid carcinoma (PTC); high-volume central lymph node metastasis (high-volume CLNM); B-mode ultrasound; radiomics; external validation; calibration; model updating


Submitted Apr 12, 2026. Accepted for publication Jun 23, 2026. Published online Jul 24, 2026.

doi: 10.21037/gs-2026-0217


Highlight box

Key findings

• A B-mode radiomics score carried most of the discriminative information for occult high-volume central lymph node metastasis (CLNM) in clinically node-negative papillary thyroid carcinoma (cN0 PTC). The locked radiomics-clinical model maintained risk ranking across temporal and external validation cohorts, but calibration shifted in the external center with the lowest event rate and improved after intercept-only local updating.

What is known and what is new?

• Ultrasound radiomics can improve preoperative nodal-risk prediction in PTC, but many reports emphasize discrimination alone.

• This study shifts the focus from model construction to clinical transportability, showing that area under the receiver operating characteristic curve (AUC), calibration, local event rate, and threshold consequences can tell different parts of the implementation story.

What is the implication, and what should change now?

• Before local adoption, radiomics prediction tools should be evaluated for calibration and decision consequences in the target center. The model should be used as a calibrated risk-stratification adjunct for preoperative assessment, counseling, and follow-up planning, not as a stand-alone trigger for expanding surgery.


Introduction

Thyroid cancer has become one of the most frequently diagnosed endocrine malignancies worldwide, and papillary thyroid carcinoma (PTC) accounts for most incident cases (1,2). In routine clinical practice, the primary challenge has shifted from diagnosing PTC itself to the preoperative identification of apparently low-risk patients who harbor clinically significant regional metastasis. This distinction is especially important in clinically node-negative (cN0) PTC, where central compartment metastases may remain occult on preoperative imaging but still influence recurrence-oriented risk stratification, surgical planning, and postoperative surveillance (3-6). The central neck is also the nodal basin in which preoperative ultrasonography is least reliable, partly because lymph nodes are small, deep, and anatomically obscured by the thyroid gland, trachea, and adjacent soft tissues (6).

Current management frameworks emphasize individualized decision-making in differentiated thyroid cancer, including careful assessment of nodal disease burden rather than reliance on tumor diagnosis alone (2,3). High-volume central lymph node metastasis (CLNM), commonly operationalized as more than five metastatic central lymph nodes, has been associated with recurrence-related risk stratification and has been examined as a more decision-relevant endpoint than any-volume microscopic nodal positivity (4,5,7-13). This endpoint is clinically appealing because it aligns more closely with the practical question faced before initial surgery: which cN0 patients are likely to harbor a nodal burden large enough to justify heightened central-compartment assessment, closer perioperative counseling, or more intensive follow-up planning?

Several clinical, biochemical, conventional ultrasound, and radiomics-based models have been proposed for predicting CLNM or high-volume nodal disease in PTC (7-26). Recent studies have specifically addressed high-volume CLNM using multimodal ultrasonographic and clinicopathological models, conventional and contrast-enhanced ultrasound nomograms, and nomograms for predicting central and high-volume central lymph node metastases (7,8,11), while recent ultrasound-imaging machine-learning studies continue to evaluate CLNM prediction across local and multicenter settings (20-22,25,26). These studies have expanded the available predictor space beyond tumor size and visual ultrasound descriptors to include quantitative image phenotypes, elastography features, deep-learning outputs, and multimodal fusion models. At the same time, prediction-model evidence in thyroid imaging remains vulnerable to a familiar methodological limitation: strong development-cohort or internal-validation discrimination does not guarantee that the same model will produce calibrated and clinically interpretable probabilities after transfer to a different center (27-32). Disease prevalence, referral pattern, ultrasound acquisition, radiomics preprocessing, surgical extent, and pathological nodal yield can all change between institutions, altering the relationship between a risk score and the absolute probability observed in local practice.

This distinction between ranking and calibration is central for clinical translation. A model can have a favorable area under the receiver operating characteristic curve (AUC) while still overestimating or underestimating absolute risk in a target population, and such miscalibration can change positive predictive value (PPV), negative predictive value (NPV), high-risk labeling, and the consequences of a fixed surgical decision threshold (32,33). Recent reporting frameworks for regression and machine-learning prediction models, clustered validation designs, artificial intelligence in medical imaging, and radiomics standardization all emphasize transparent cohort definition, validation after model locking, calibration assessment, and clinically interpretable performance reporting (27-31,34). For radiomics models intended to inform preoperative surgical planning, these requirements are not merely technical reporting details; they determine whether an apparently accurate model can be interpreted safely when case mix changes.

The present study therefore reframed the analysis from model construction toward clinical transportability. Rather than asking whether clinical-radiomics fusion maximized AUC, we evaluated whether a locked B-mode ultrasound radiomics-clinical risk tool for occult high-volume CLNM in cN0 PTC maintained risk ranking, preserved calibrated absolute-risk estimates, and retained clinically interpretable threshold behavior after temporal and external transfer. The reporting framework combines discrimination, Brier score, calibration intercept and slope, fixed-threshold classification, PPV/NPV, decision curve analysis with default-strategy comparators, threshold-consequence plots, and local intercept recalibration to characterize the conditions under which model output may be used as clinical decision support across centers. We present this article in accordance with the TRIPOD reporting checklist (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0217/rc) (28).


Methods

Study design and reporting

This was a three-center retrospective prediction-model development, validation, and updating study. The analytic aim was to evaluate whether a B-mode ultrasound radiomics-clinical model developed in one center could retain discrimination, calibration, and clinically interpretable threshold behavior after temporal and external transfer. The modeling strategy was defined around a locked development model followed by temporal validation, two external validations, calibration assessment, decision-curve analysis, threshold-consequence analysis, and intercept-only local updating.

Participants and cohorts

Patients were screened from the source cohorts according to prespecified inclusion and exclusion criteria. The inclusion criteria were as follows: (I) age 18 years or older; (II) histopathologically confirmed PTC after initial thyroid surgery; (III) clinically node-negative (cN0) status before surgery, defined as no clinically suspicious central or lateral cervical lymph node documented on preoperative clinical or ultrasound assessment; (IV) available preoperative B-mode ultrasound images of the primary thyroid lesion suitable for radiomics analysis; (V) available preoperative clinical and ultrasound variables required for candidate predictor assessment; and (VI) available postoperative central compartment pathological lymph-node assessment for endpoint ascertainment. The exclusion criteria were as follows: (I) incomplete clinical, ultrasound, or pathological data; (II) unavailable or poor-quality B-mode ultrasound images that could not support lesion segmentation or radiomics-score extraction; (III) prior thyroid or neck surgery, cervical radiotherapy, or other preoperative treatment that could alter cervical lymph-node assessment; (IV) non-initial surgery for recurrent thyroid disease; (V) clinically evident cervical lymph-node metastasis before surgery; and (VI) coexisting non-PTC thyroid malignancy or another active malignancy. Candidate predictors were restricted to variables available preoperatively, while postoperative histopathological findings were used only for endpoint ascertainment. Fujian Provincial Hospital, Fuzhou University Affiliated Fujian Provincial Hospital (Center 1) contributed both a development cohort and a later temporal validation cohort, allowing assessment of temporal transport within the same institutional environment. Fujian Union Hospital (Center 2) and Fujian Cancer Hospital (Center 3) served as external validation cohorts, allowing evaluation across institutions with potentially different case mix, imaging workflow, and nodal yield. The study dataset contained 957 patients: 489 in the development cohort, 210 in the temporal validation cohort, 139 in Center 2, and 119 in Center 3. The patient-selection process is shown in Figure 1.

Figure 1 Patient selection flow diagram stratified by study center. CLNM, central lymph node metastasis; PTC, papillary thyroid carcinoma; US, ultrasound.

This retrospective study used de-identified clinical and imaging data and involved no additional patient contact or intervention. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Institutional Review Board of Fujian Provincial Hospital, Fuzhou University Affiliated Fujian Provincial Hospital (Approval No. K2025-04-002). The requirement for individual informed consent was waived by the approving institutional review board because of the retrospective design and use of de-identified data. All participating hospitals, including Fujian Union Hospital and Fujian Cancer Hospital, were informed of the study and agreed to participate.

Outcome

The endpoint was occult high-volume CLNM, defined as more than five metastatic central lymph nodes on postoperative pathology among patients classified as cN0 before surgery. This definition was chosen because nodal burden is more closely aligned with recurrence-oriented risk stratification than any-volume microscopic nodal positivity. Patients with no central nodal metastasis or five or fewer metastatic central nodes were endpoint-negative. The endpoint was not used as an input predictor during model development.

Predictors and model structure

The clinical predictor set comprised age, sex, tumor maximum diameter, margin contour, C-TIRADS 4 subclass, multifocality, and vascularity. These variables were selected because they are routinely available before surgery and can be extracted from clinical records or preoperative ultrasound reports. The radiomics predictor was a prespecified B-mode ultrasound radiomics score calculated from lesion-level primary-tumor B-mode image features. Preoperative B-mode images were acquired during routine clinical ultrasound using mainstream Philips and Mindray ultrasound platforms. Lesion-level regions of interest were segmented on primary-tumor B-mode images for radiomics-score generation, and the available project record documented acceptable segmentation repeatability (intraclass correlation coefficient, 0.82). In accordance with CLAIM and Image Biomarker Standardization Initiative (IBSI) reporting principles, device platform, acquisition parameters, image format, region-of-interest segmentation process, feature extractor and software version, preprocessing/discretization settings, and segmentation-repeatability assessment were treated as radiomics-source documentation items (31,34). The analytic dataset available for this revision contained the finalized B-mode score, ultrasound platform information, and segmentation-repeatability summary but not all feature-level provenance fields; these provenance descriptors were therefore not used as candidate predictors and are reported as reproducibility limitations. Postoperative pathological variables and any information unavailable before surgery were excluded from candidate predictors to prevent leakage. Three development-cohort models were fitted for comparison: a clinical-only model, a B-mode radiomics score-only model, and a clinical plus B-mode radiomics score model. The clinical plus B-mode score model was treated as the locked primary implementation model because it embeds the quantitative B-mode signal in a clinician-readable preoperative covariate set. External validation was used to evaluate transportability rather than to reselect the model. All preprocessing, imputation, scaling, encoding, and model fitting were performed inside the development cohort. The fitted model, preprocessing parameters, and decision threshold were then applied without refitting to the temporal and external validation cohorts.

Model evaluation

Model performance was evaluated by discrimination, overall accuracy of probabilistic prediction, calibration, fixed-threshold classification, and clinical utility. Discrimination was summarized by AUC. Overall probabilistic error was summarized by the Brier score. Calibration was described by mean predicted risk, calibration intercept, calibration slope, and calibration plots. Thresholded performance was summarized by sensitivity, specificity, PPV, NPV, accuracy, true-positive count, false-positive count, true-negative count, false-negative count, high-risk labeling rate, and event capture. The primary classification threshold was selected in the development cohort by the Youden index (0.565) and locked before validation. Decision curve analysis was used to compare the net benefit of model-guided decisions with default strategies across clinically relevant threshold probabilities. Bootstrap resampling was used to estimate 95% confidence intervals for AUC and Brier score.

Local intercept updating

For each validation cohort, intercept-only recalibration was evaluated as a post hoc local updating strategy (32,35). This method preserves predictor coefficients and shifts only the baseline log-odds so that the average predicted risk matches the observed event rate in the target cohort. It was used to distinguish loss of ranking from baseline-risk mismatch. A reduction in Brier score after intercept updating was interpreted as improved probabilistic fit in the target cohort, whereas the unchanged slope and AUC reflected preservation of the original predictor effects and patient ordering.

Statistical analysis

Continuous variables are summarized as median (IQR), and categorical variables as n (%). Development-cohort endpoint comparisons used Mann-Whitney U tests for continuous variables and chi-square or Fisher exact tests for categorical variables, as appropriate. Model comparisons were descriptive and were intended to characterize transportability rather than to select a new model after validation. All tests were two-sided. Analyses were performed in Python. The development cohort contained 219 endpoint events, supporting the prespecified compact logistic model under contemporary prediction-model sample-size guidance (36). All validation analyses were conducted after model locking, and validation cohorts were not used for feature selection, coefficient estimation, threshold selection, or preprocessing-parameter refitting.


Results

Patient selection, cohort spectrum, and endpoint burden

Figure 1 presents the patient-selection structure and cohort allocation. The analytic dataset contained 957 cN0 PTC patients, including 399 endpoint-positive cases (41.7%). The development cohort included 489 patients from Center 1, of whom 219 (44.8%) had occult high-volume CLNM. The temporal validation cohort included 210 later Center 1 patients, with 95 endpoint-positive cases (45.2%). The two external validation cohorts included 139 patients from Center 2 and 119 patients from Center 3, with endpoint-positive counts of 57 (41.0%) and 28 (23.5%), respectively. Thus, event rates were similar in the development, temporal validation, and Center 2 cohorts but lower in Center 3. Cohort-level demographic and ultrasound summaries are shown in Table 1.

Table 1

Cohort characteristics and center-level spectrum shift

Characteristic Development Temporal validation External validation 1 External validation 2
Center Center 1 Center 1 Center 2 Center 3
N 489 210 139 119
High-volume CLNM 219 (44.8) 95 (45.2) 57 (41.0) 28 (23.5)
Age, years 48.00 (36.00–56.00) 46.49 (38.00–54.00) 49.00 (37.50–55.00) 47.00 (39.50–56.00)
Female sex 342 (69.9) 148 (70.5) 106 (76.3) 83 (69.7)
Tumor maximum diameter, cm 0.70 (0.50–1.20) 0.70 (0.50–1.10) 0.70 (0.50–1.00) 0.70 (0.48–1.08)
Multifocality 136 (27.8) 69 (32.9) 43 (30.9) 35 (29.4)
Highest C-TIRADS subclass 317 (64.8) 135 (64.3) 88 (63.3) 67 (56.3)
Rich vascularity 25 (5.1) 16 (7.6) 10 (7.2) 3 (2.5)

Data are presented as median (interquartile range) or n (%), unless otherwise indicated. Center 1, Fujian Provincial Hospital; Center 2, Fujian Union Hospital; Center 3, Fujian Cancer Hospital. C-TIRADS, Chinese Thyroid Imaging Reporting and Data System; CLNM, central lymph node metastasis.

Across cohorts, median age ranged from 46.49 to 49.00 years, and the proportion of female patients ranged from 69.7% to 76.3%. Median tumor maximum diameter was 0.70 cm in each cohort, with interquartile ranges from 0.48−1.08 to 0.50−1.20 cm. The highest C-TIRADS subclass was present in 56.3% to 64.8% of patients across cohorts. Rich vascularity was recorded in 5.1% of the development cohort, 7.6% of the temporal validation cohort, 7.2% of Center 2, and 2.5% of Center 3. Mean predicted risk followed observed risk in the development, temporal validation, and Center 2 cohorts but remained higher than the observed event rate in Center 3 (36.4% predicted vs. 23.5% observed; Figure 2).

Figure 2 Study design and cohort spectrum. (A) The model-development, temporal-validation, external-validation, calibration, and local-updating workflow. (B) Center-level event-rate differences. (C) The distribution of predicted risks across cohorts. Center 1, Fujian Provincial Hospital; Center 2, Fujian Union Hospital; Center 3, Fujian Cancer Hospital.

Development-cohort endpoint comparisons

Endpoint-stratified comparisons in the development cohort are summarized in Table S1 and Figure 3. Endpoint-positive patients had a lower median age than endpoint-negative patients [45.00 (34.50−54.00) vs. 49.00 (38.00−58.00) years; P=0.005]. Median tumor maximum diameter was larger in endpoint-positive patients [0.83 (0.60−1.50) vs. 0.60 (0.40−1.00) cm; P<0.001]. The highest C-TIRADS subclass was more frequent in the endpoint-positive group (160/219, 73.1%) than in the endpoint-negative group (157/270, 58.1%; P<0.001). Rich vascularity was recorded in 17 endpoint-positive patients (7.8%) and 8 endpoint-negative patients (3.0%; P=0.029). The B-mode radiomics score was higher in endpoint-positive patients [median, 0.21 (IQR, −0.05 to 0.46)] than in endpoint-negative patients [median, −0.22 (IQR, −0.43 to −0.00); P<0.001]. Female sex, margin contour, and multifocality did not differ significantly in the development-cohort endpoint comparison.

Figure 3 Development-cohort endpoint signal before modeling. (A) Tumor maximum diameter distribution according to endpoint status. (B) B-mode radiomics score distribution according to endpoint status. (C) High-volume CLNM rate according to C-TIRADS subclass in the development cohort. (D) High-volume CLNM rate according to vascularity in the development cohort. C-TIRADS, Chinese Thyroid Imaging Reporting and Data System; CLNM, central lymph node metastasis.

Model specification and comparative performance

The locked primary implementation model combined the B-mode radiomics score with preoperative clinical-ultrasound variables. The largest positive coefficient in the fitted model was the B-mode radiomics score (coefficient, 1.535; odds ratio per encoded unit, 4.64), followed by rich vascularity category 1.0 (coefficient, 0.688; odds ratio, 1.99), margin contour category 2.0 (coefficient, 0.481; odds ratio, 1.62), and C-TIRADS subclass category 3.0 (coefficient, 0.211; odds ratio, 1.23). Full encoded coefficients are listed in Table S2.

Model comparisons are shown in Figure 4 and Table S3. The clinical-only model had AUCs of 0.664 (95% CI, 0.615−0.709) in development, 0.650 (0.574−0.727) in temporal validation, 0.631 (0.531−0.726) in Center 2, and 0.648 (0.534−0.756) in Center 3. The B-mode rad-score model had AUCs of 0.808 (0.771−0.846), 0.866 (0.821−0.912), 0.875 (0.814−0.930), and 0.899 (0.835−0.951), respectively. The clinical plus B-mode rad-score model had AUCs of 0.823 (0.788−0.860), 0.874 (0.830−0.921), 0.866 (0.804−0.924), and 0.870 (0.798−0.929), respectively. Thus, the B-mode score accounted for most of the discriminative performance, and the clinical plus B-mode model did not show a consistent external AUC advantage over the B-mode score-only model. Brier scores for the locked clinical plus B-mode model were 0.169 (0.151−0.187), 0.142 (0.117−0.165), 0.150 (0.120−0.182), and 0.138 (0.110−0.173) across the four cohorts (Table 2).

Figure 4 Model comparison and discrimination transportability. (A) ROC curves of the locked primary model across the development, temporal validation, and external validation cohorts. (B) AUC comparison across the clinical-only, B-mode radiomics score-only, and clinical plus B-mode radiomics score models. (C) Brier score comparison across the three models. (D) Summary interpretation of the model-comparison results. AUC, area under the ROC curve; C-TIRADS, Chinese Thyroid Imaging Reporting and Data System; ROC, receiver operating characteristic.

Table 2

Locked primary-model performance at the development-derived threshold

Metric Development (Center 1) Temporal validation (Center 1) External validation 1 (Center 2) External validation 2 (Center 3)
Event rate, % 44.8 45.2 41.0 23.5
AUC (95% CI) 0.823 (0.788–0.860) 0.874 (0.830–0.921) 0.866 (0.804–0.924) 0.870 (0.798–0.929)
Brier score (95% CI) 0.169 (0.151–0.187) 0.142 (0.117–0.165) 0.150 (0.120–0.182) 0.138 (0.110–0.173)
Mean predicted risk, % 44.8 44.1 44.7 36.4
Calibration intercept 0.006 0.123 −0.221 −0.902
Calibration slope 1.036 1.234 1.134 1.265
Sensitivity 0.644 0.674 0.702 0.679
Specificity 0.863 0.922 0.854 0.857
PPV 0.792 0.877 0.769 0.594
NPV 0.749 0.774 0.805 0.897

The fixed threshold was 0.565, selected in the development cohort by the Youden index and applied unchanged to all validation cohorts. Center 1, Fujian Provincial Hospital; Center 2, Fujian Union Hospital; Center 3, Fujian Cancer Hospital. AUC, area under the receiver operating characteristic curve; CI, confidence interval; NPV, negative predictive value; PPV, positive predictive value.

Calibration, local updating, and threshold consequences

Calibration differed across validation settings. In the development cohort, mean predicted risk was 44.8% and the calibration intercept was 0.006. In temporal validation, the observed event rate was 45.2%, mean predicted risk was 44.1%, and the calibration intercept was 0.123. In Center 2, the observed event rate was 41.0%, mean predicted risk was 44.7%, and the calibration intercept was −0.221. In Center 3, the observed event rate was 23.5%, mean predicted risk was 36.4%, and the calibration intercept was −0.902. Intercept-only updating aligned mean predicted risk with observed event rate in each validation cohort. The largest numeric change was observed in Center 3, where recalibrated mean predicted risk changed to 23.5% and Brier score decreased from 0.138 to 0.121 (Figure 5; Table S4).

Figure 5 Calibration drift and intercept-only updating. (A) Calibration curves before local intercept updating. (B) Calibration curves after intercept-only updating in the validation cohorts. (C) Observed event rate, original mean predicted risk, and recalibrated mean predicted risk across cohorts.

At the locked development-derived threshold of 0.565, sensitivity/specificity were 0.644/0.863 in development, 0.674/0.922 in temporal validation, 0.702/0.854 in Center 2, and 0.679/0.857 in Center 3. PPV/NPV were 0.792/0.749, 0.877/0.774, 0.769/0.805, and 0.594/0.897, respectively (Table 2). The corresponding true-positive/false-positive/true-negative/false-negative counts were 141/37/233/78 in development, 64/9/106/31 in temporal validation, 40/12/70/17 in Center 2, and 19/13/78/9 in Center 3 (Table S5). In Center 3, standard decision-curve analysis showed positive but modest net benefit for the locked model compared with treat-all and treat-none reference strategies across threshold probabilities from 0.05 to 0.60. At the locked threshold, 32 of 119 patients were labeled high risk and 67.9% of endpoint-positive patients were captured. Threshold-scenario analyses in Table S6 show how high-risk labels, event capture, PPV, and NPV changed across candidate thresholds from 0.200 to 0.600 (Figure 6).

Figure 6 Center 3 decision-curve and threshold consequences. (A) Standard decision-curve analysis for the locked primary model compared with treat-all and treat-none reference strategies; the vertical line marks the development-derived threshold of 0.565. (B-D) Focus on Center 3, where the local event rate was lowest, showing threshold-dependent sensitivity/specificity, PPV/NPV, and high-risk labeling. Center 3, Fujian Cancer Hospital. NPV, negative predictive value; PPV, positive predictive value.

Discussion

This three-center analysis evaluated a B-mode ultrasound radiomics-clinical prediction model for occult high-volume CLNM in cN0 PTC with a focus on transportability, calibration, and local updating. The primary finding is that the model maintained risk ranking across temporal and external validation cohorts, whereas absolute-risk calibration varied with center-level disease spectrum. A second important finding is that the B-mode radiomics score carried most of the discriminative signal, while the clinical plus B-mode model provided an interpretable locked implementation structure. In the lowest-event-rate external cohort, the locked model retained discrimination but overestimated average risk before recalibration. Intercept-only updating corrected the largest baseline-risk mismatch while preserving the original predictor effects. These findings separate three clinically distinct questions that are often combined in model reports: whether a model ranks patients correctly, whether its predicted probabilities are locally calibrated, and whether a fixed threshold produces acceptable clinical consequences in the target center.

The development-cohort comparisons showed that endpoint-positive patients had larger tumors, higher B-mode radiomics scores, more frequent highest C-TIRADS subclass, and richer vascularity. These variables correspond to biologically plausible imaging phenotypes supported by recent preoperative high-volume CLNM literature. Larger primary tumors may reflect greater tumor burden and a longer window for lymphatic invasion, a pattern repeatedly retained in recent high-volume CLNM models (7-13). Suspicious ultrasound categories, margin-related structural features, and vascular signals may represent invasive tumor morphology and perfusion-related characteristics associated with regional spread; recent high-volume and ultrasound-based studies have evaluated these same variables as preoperative correlates (7-11,14-16). The present analysis adds a quantitative B-mode radiomics score to this clinical-ultrasound context and reports how the combined model behaves after temporal and external transfer.

The radiomics score accounted for most of the discriminative gain compared with clinical variables alone. Radiomics features summarize gray-scale intensity distribution, texture heterogeneity, boundary complexity, and local acoustic patterns that may be invisible to visual scoring; recent radiomics, elastography, deep-learning, and multimodal studies in 2024–2026 support the relevance of such quantitative image phenotypes in PTC nodal prediction (14,15,18-23,25,26). Adding clinical-ultrasound variables to the B-mode score produced limited incremental discrimination. The rationale for retaining the clinical plus B-mode model is therefore practical and translational: clinical covariates keep the tool interpretable for surgeons and sonographers, while the radiomics score supplies the dominant quantitative signal. However, the external validation results also show that stronger discrimination does not automatically resolve calibration. A radiomics-derived score may preserve ranking while still requiring baseline-risk adjustment when the target center has a different referral pattern, case mix, or nodal yield.

The Center 3 results are particularly informative for implementation. The observed event rate was lower than in the development cohort, and the unrecalibrated model overestimated mean risk despite retaining high AUC. This pattern illustrates the calibration problem described in broader prediction-model literature: AUC measures ordering, whereas calibration describes agreement between predicted and observed risk (32). In settings where the clinical decision depends on an absolute probability threshold, calibration drift can change the number of patients labeled high risk, the PPV of a positive model result, and the reassurance provided by a negative result.

The local-updating analysis provides a pragmatic statistical response to this problem. Intercept-only updating does not alter the predictor weights and therefore does not claim that the imaging phenotype has changed across centers. Instead, it shifts the baseline log-odds to match the observed event rate of the target cohort. This distinction is useful in clinical prediction because loss of calibration may arise from baseline-risk mismatch even when relative ranking remains stable. In such a setting, rebuilding the model from a small local cohort could introduce instability, whereas intercept updating offers a transparent and parsimonious correction that can be evaluated before broader prospective testing (32,35).

For preoperative management of cN0 PTC, the model is best interpreted as a calibrated risk-stratification adjunct rather than a determinant of surgical extent. Its clinical value is to make an otherwise implicit preoperative risk estimate more explicit: patients with higher predicted risk may warrant closer central-compartment ultrasound review, expert sonographic reassessment, multidisciplinary discussion, and more careful postoperative surveillance planning. Conversely, a low-risk output in a locally calibrated model may support more conservative perioperative counseling, particularly in centers where prophylactic central neck dissection is not routine for cN0 disease. This implementation boundary is consistent with guideline-based individualized management and with recent prediction-model guidance requiring transparent validation, calibration, and clinical-impact assessment before deployment (3,27,30). The model should not be used to automatically expand surgery, because the acceptable balance between missed high-volume nodal disease and unnecessary escalation depends on surgeon experience, institutional policy, patient preference, and guideline-based care.

The fixed-threshold analysis clarifies the trade-offs that clinicians would face. A threshold optimized in the development cohort did not produce the same PPV/NPV balance in every validation cohort because the local event rate changed. This is clinically important because false negatives may leave high-volume nodal disease unrecognized preoperatively, whereas false positives may increase concern for a nodal burden that is not present. The decision-curve analysis in Center 3 shows that the model can offer net benefit compared with default strategies, but the magnitude is modest and must be read together with threshold consequences. A practical deployment pathway would therefore require local assessment of discrimination, calibration, and threshold consequences, followed by recalibration or threshold adjustment if the local event rate differs from the development cohort.

The model may also help standardize communication across clinicians. Central neck management in cN0 PTC involves balancing recurrence risk, surgical morbidity, operator experience, and patient preference. A calibrated probability estimate can create a common language for discussing risk, but it should not replace ultrasound review, operative judgment, pathological context, or guideline-based care. In this respect, the model is most appropriately positioned as a structured risk-stratification adjunct. Its value would be greatest if it reduces unstructured variability in preoperative assessment while keeping the final treatment decision anchored in multidisciplinary evaluation.

Several limitations should be considered. First, the study was retrospective, and selection, referral, imaging, and surgical practice patterns may have influenced the observed event rate and endpoint ascertainment. Second, although the study included temporal validation and two external validation cohorts, Center 3 contained only 28 endpoint-positive cases; the intercept-updating result in this cohort should therefore be treated as a calibration demonstration requiring prospective confirmation. Third, the endpoint depends on the extent of central lymph node dissection and pathological node retrieval, so variation in surgical or pathological practice could affect the measured nodal burden. Fourth, although the available project record documented mainstream Philips and Mindray ultrasound platforms and acceptable segmentation repeatability for radiomics-score extraction (intraclass correlation coefficient, 0.82), the analytic dataset did not contain complete radiomics-source documentation fields recommended by CLAIM and IBSI, including exact scanner models, probe frequencies, acquisition settings, original image-format records, feature extractor and software-version logs, and preprocessing/discretization settings. These limits detailed assessment of radiomics reproducibility across scanners, readers, and imaging protocols (31,34). Fifth, the analysis evaluated statistical performance and local updating but did not test whether model-informed management changes surgical decisions, complications, recurrence, or patient-reported outcomes.


Conclusions

A locked B-mode ultrasound radiomics-clinical model for occult high-volume CLNM in cN0 PTC maintained risk ranking across temporal and external validation cohorts, but absolute-risk calibration varied when the external disease spectrum changed. The B-mode score supplied most discriminative information, while the clinical plus B-mode model provided an interpretable implementation structure. Local intercept updating reduced the largest calibration mismatch without rebuilding the model. These findings support a transportability-oriented reporting strategy for thyroid radiomics prediction models in which AUC is reported together with calibration, threshold consequences, and local updating. Prospective multicenter validation and impact analysis are still required before routine clinical implementation.


Acknowledgments

The authors thank the Departments of Ultrasonography, Surgery, Pathology, and Medical Records of Fujian Provincial Hospital, Fujian Union Hospital, and Fujian Cancer Hospital for their support in case review, image review, and data collection. AI-assisted language drafting, formatting, and reference-organization support was used during preparation of this draft. The authors are responsible for verifying the data provenance, analyses, references, and final manuscript wording before submission.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0217/rc

Data Sharing Statement: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0217/dss

Peer Review File: Available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0217/prf

Funding: This study was supported by the Fujian Provincial Natural Science Foundation (No. 2025J08042). The funder had no role in study design, data analysis, manuscript preparation, or submission decisions.

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://gs.amegroups.com/article/view/10.21037/gs-2026-0217/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Institutional Review Board of Fujian Provincial Hospital, Fuzhou University Affiliated Fujian Provincial Hospital (Approval No. K2025-04-002). The requirement for individual informed consent was waived by the approving institutional review board because of the retrospective design and use of de-identified data. All participating hospitals, including Fujian Union Hospital and Fujian Cancer Hospital, were informed of the study and agreed to participate.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 2024;74:229-63. [Crossref] [PubMed]
  2. Chen DW, Lang BHH, McLeod DSA, et al. Thyroid cancer. Lancet 2023;401:1531-44. [Crossref] [PubMed]
  3. Haugen BR, Alexander EK, Bible KC, et al. 2015 American Thyroid Association Management Guidelines for Adult Patients with Thyroid Nodules and Differentiated Thyroid Cancer: The American Thyroid Association Guidelines Task Force on Thyroid Nodules and Differentiated Thyroid Cancer. Thyroid 2016;26:1-133. [Crossref] [PubMed]
  4. Randolph GW, Duh QY, Heller KS, et al. The prognostic significance of nodal metastases from papillary thyroid carcinoma can be stratified based on the size and number of metastatic lymph nodes, as well as the presence of extranodal extension. Thyroid 2012;22:1144-52. [Crossref] [PubMed]
  5. Adam MA, Pura J, Goffredo P, et al. Presence and Number of Lymph Node Metastases Are Associated With Compromised Survival for Patients Younger Than Age 45 Years With Papillary Thyroid Cancer. J Clin Oncol 2015;33:2370-5. [Crossref] [PubMed]
  6. Alabousi M, Alabousi A, Adham S, et al. Diagnostic Test Accuracy of Ultrasonography vs Computed Tomography for Papillary Thyroid Cancer Cervical Lymph Node Metastasis: A Systematic Review and Meta-analysis. JAMA Otolaryngol Head Neck Surg 2022;148:107-18. [Crossref] [PubMed]
  7. Qian J, Zhang Z, Chen Y, et al. Multimodal ultrasonographic and clinicopathological model for predicting high-volume lymph node metastasis in cN0 papillary thyroid carcinoma. Front Endocrinol (Lausanne) 2025;16:1613672. [Crossref] [PubMed]
  8. Yan X, Peng Q, Chen J, et al. A nomogram based on conventional and contrast-enhanced ultrasound for predicting high-volume central lymph node metastasis in papillary thyroid carcinoma. BMC Cancer 2025;25:1529. [Crossref] [PubMed]
  9. Zhu H, Zhang H, Wei P, et al. Development and validation of a clinical predictive model for high-volume lymph node metastasis of papillary thyroid carcinoma. Sci Rep 2024;14:15828. [Crossref] [PubMed]
  10. Liu XH, Yin HQ, Shen H, et al. A multivariable model of ultrasound and biochemical parameters for predicting high-volume lymph node metastases of papillary thyroid carcinoma with Hashimoto's thyroiditis. Front Endocrinol (Lausanne) 2024;15:1501142. [Crossref] [PubMed]
  11. Huang X, Gan X, Feng J, et al. Nomograms for predicting cervical central lymph node metastases and high-volume cervical central lymph node metastases in papillary thyroid carcinoma. Gland Surg 2025;14:421-35. [Crossref] [PubMed]
  12. Wang Z, Gui Z, Wang Z, et al. Clinical and ultrasonic risk factors for high-volume central lymph node metastasis in cN0 papillary thyroid microcarcinoma: A retrospective study and meta-analysis. Clin Endocrinol (Oxf) 2023;98:609-21. [Crossref] [PubMed]
  13. Feng JW, Ye J, Qi GF, et al. Nomograms for Prediction of High-Volume Lymph Node Metastasis in Papillary Thyroid Carcinoma Patients. Otolaryngol Head Neck Surg 2023;168:1054-66. [Crossref] [PubMed]
  14. Jia W, Cai Y, Wang S, et al. Predictive value of an ultrasound-based radiomics model for central lymph node metastasis of papillary thyroid carcinoma. Int J Med Sci 2024;21:1701-9. [Crossref] [PubMed]
  15. Feng JW, Liu SQ, Qi GF, et al. Development and Validation of Clinical-Radiomics Nomogram for Preoperative Prediction of Central Lymph Node Metastasis in Papillary Thyroid Carcinoma. Acad Radiol 2024;31:2292-305. [Crossref] [PubMed]
  16. Yan X, Mou X, Yang Y, et al. Predicting central lymph node metastasis in patients with papillary thyroid carcinoma based on ultrasound radiomic and morphological features analysis. BMC Med Imaging 2023;23:111. [Crossref] [PubMed]
  17. HajiEsmailPoor Z. Kargar Z, Tabnak P. Radiomics diagnostic performance in predicting lymph node metastasis of papillary thyroid carcinoma: A systematic review and meta-analysis. Eur J Radiol 2023;168:111129. [Crossref] [PubMed]
  18. Zhang XY, Zhang D, Zhou W, et al. Predicting lymph node metastasis in papillary thyroid carcinoma: radiomics using two types of ultrasound elastography. Cancer Imaging 2025;25:13. [Crossref] [PubMed]
  19. Gao L, Wen X, Yue G, et al. The Predictive Value of a Nomogram Based on Ultrasound Radiomics, Clinical Factors, and Enhanced Ultrasound Features for Central Lymph Node Metastasis in Papillary Thyroid Microcarcinoma. Ultrason Imaging 2025;47:93-103. [Crossref] [PubMed]
  20. Han H, Sun H, Zhou C, et al. Development and validation of a machine learning model for central compartmental lymph node metastasis in solitary papillary thyroid microcarcinoma via ultrasound imaging features and clinical parameters. BMC Med Imaging 2025;25:228. [Crossref] [PubMed]
  21. Liu W, Zheng J, Han L, et al. Clinical performance of a machine learning-based model for detecting lymph node metastasis in papillary thyroid carcinoma: A multicenter study. Int J Surg 2025;111:4062-7. [Crossref] [PubMed]
  22. Wang Z, Yang S, Li Y, et al. An Interpretable Machine-Learning Model for Predicting Occult Central Lymph Node Metastasis in Papillary Thyroid Cancer. J Clin Endocrinol Metab 2026;111:1232-47. [Crossref] [PubMed]
  23. Gao Y, Wang W, Yang Y, et al. An integrated model incorporating deep learning, hand-crafted radiomics and clinical and US features to diagnose central lymph node metastasis in patients with papillary thyroid cancer. BMC Cancer 2024;24:69. [Crossref] [PubMed]
  24. Chang L, Zhang Y, Zhu J, et al. An integrated nomogram combining deep learning, clinical characteristics and ultrasound features for predicting central lymph node metastasis in papillary thyroid cancer: A multicenter study. Front Endocrinol (Lausanne) 2023;14:964074. [Crossref] [PubMed]
  25. Peng X, Wu P, Li W, et al. AI-based multimodal prediction of lymph node metastasis and capsular invasion in cT1N0M0 papillary thyroid carcinoma. Front Endocrinol (Lausanne) 2025;16:1580885. [Crossref] [PubMed]
  26. Zhou W, Li L, Hao X, et al. Predicting central lymph node metastasis in papillary thyroid microcarcinoma: a breakthrough with interpretable machine learning. Front Endocrinol (Lausanne) 2025;16:1537386. [Crossref] [PubMed]
  27. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
  28. Collins GS, Reitsma JB, Altman DG, et al. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med 2015;162:55-63. [Crossref] [PubMed]
  29. Debray TPA, Collins GS, Riley RD, et al. Transparent reporting of multivariable prediction models developed or validated using clustered data: TRIPOD-Cluster checklist. BMJ 2023;380:e071018. [Crossref] [PubMed]
  30. Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 2025;388:e082505. [Crossref] [PubMed]
  31. Tejani AS, Klontzas ME, Gatti AA, et al. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 Update. Radiol Artif Intell 2024;6:e240300. [Crossref] [PubMed]
  32. Van Calster B, McLernon DJ, van Smeden M, et al. Calibration: the Achilles heel of predictive analytics. BMC Med 2019;17:230. [Crossref] [PubMed]
  33. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making 2006;26:565-74. [Crossref] [PubMed]
  34. Zwanenburg A, Vallières M, Abdalah MA, et al. The Image Biomarker Standardization Initiative: Standardized Quantitative Radiomics for High-Throughput Image-based Phenotyping. Radiology 2020;295:328-38. [Crossref] [PubMed]
  35. Janssen KJ, Moons KG, Kalkman CJ, et al. Updating methods improved the performance of a clinical prediction model in new patients. J Clin Epidemiol 2008;61:76-86. [Crossref] [PubMed]
  36. Riley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ 2020;368:m441. [Crossref] [PubMed]
Cite this article as: Zhang Z, Lin C, Ning J, Liu X, Dong C, Ling H. Clinical transportability, calibration drift, and local updating of a B-mode ultrasound radiomics-clinical prediction model for occult high-volume central lymph node metastasis in cN0 papillary thyroid carcinoma: development, temporal validation, and external validation. Gland Surg 2026;15(7):195. doi: 10.21037/gs-2026-0217

Download Citation