Deep learning-enabled monitoring of postoperative fracture healing on serial radiographs: a 150-patient study using an enhanced YOLOv11 framework.
To develop and validate a deep learning framework for classifying postoperative time-points as a proxy task for monitoring longitudinal fracture healing progression from serial radiographs. This retrospective study included 150 patients with paired pre-treatment and follow-up X-ray images. We built a detection-guided pipeline comprising (1) fracture-region localization using an enhanced YOLOv11 detector integrating attention mechanism, Focal-SIoU loss, and data augmentation, and (2) healing-status prediction from detected regions of interest by quantifying callus formation and fracture-line changes over time. Data were split at the patient level into training/validation/test cohorts. Performance was evaluated using accuracy, F1 score, ROC/AUC, and calibration, and compared with clinician readings. The YOLOv11-guided framework achieved reliable fracture localization and consistent healing assessment on serial radiographs. On the independent test set, it showed stable discriminative ability across follow-up stages and improved robustness over manual interpretation, particularly at early postoperative time points when radiographic changes are subtle. This single-center study demonstrates a technical framework for objective and scalable radiograph-based longitudinal fracture-healing monitoring. External, multi-center validation is required before broader clinical deployment. The proposed detection-enhanced YOLOv11 framework may support clinical follow-up and decision-making after fracture surgery.
- Research Article
48
- 10.1259/0007-1285-60-717-849
- Sep 1, 1987
- The British Journal of Radiology
The effects of therapy on the osteolytic bone lesions of Paget's disease have been assessed from serial bone radiographs. Changes in the rate of progression of lytic "wedge" lesions were measured and alterations in the texture of lytic "blade" lesions were graded on an empirical scale. Useful matching was possible using standard radiographs, although special care was needed to avoid artefacts from suboptimal positioning, magnification and variation in exposure. Serial radiographs were obtained of 57 lytic blade lesions in 54 patients receiving treatment with the bisphosphonate 1-hydroxyethylidene-1, 1-bisphosphonate (EHDP) and of 20 lesions in 20 patients treated with oral or intravenous 3-amino-1-hydroxypropylidene-1, 1-bisphosphonate (APD). Treatment with EHDP was associated with a significant deterioration in bone texture in 50% of lytic blade lesions, and with healing in only 20%. Deterioration was accompanied by an increase in local bone pain in 17% of these patients. In contrast, significant healing was observed in 17 of 20 lytic lesions (eight wedge, nine blade) within 6 months of beginning a course of intravenous or oral APD. In four of eight patients the progression of a lytic tibial wedge was arrested and in the remaining four the direction of wedge movement was reversed. In two patients the wedge had almost completely "filled in", making measurement difficult. Bone healing was usually accompanied by pain relief, reduction in skin temperature and rapid suppression of the urine hydroxyproline (uHP) into the normal range. However, in four patients who received intravenous APD, repair of lytic bone lesions was observed despite persisting elevation of uHP. These improvements with APD were sustained at 12 months, although in one patient whose biochemical indices were restored to normal the resorption front showed further progression, despite initial temporary reversal. The trends apparent in these short-term studies were also seen in four patients in whom wedge velocities were measured over periods of 6-10 years. These results confirm that after treatment of Paget's disease, bone healing or deterioration can be accurately assessed from serial standard radiographs. Reproducible matching is best achieved by ensuring that all radiographs are taken by the same radiographer. Minor alterations in radiological bone texture provide an important index of drug effect which is not always apparent from measurement of biochemical and other indices.
- Research Article
4
- 10.1186/s13018-024-04679-y
- Apr 3, 2024
- Journal of Orthopaedic Surgery and Research
BackgroundThe goal of this study is to propose a classification system with a common nomenclature for radiographic observations of periprosthetic bone changes following cTDR.MethodsAided by serial plain radiographs from recent cTDR cases (34 patients; 44 devices), a panel of experts assembled for the purpose of creating a classification system to aid in reproducibly and accurately identifying bony changes and assessing cTDR radiographic appearance. Subdividing the superior and inferior vertebral bodies into 3 equal sections, observed bone loss such as endplate rounding, cystic erosion adjacent to the endplate, and cystic erosion not adjacent to the endplate, is recorded. Determining if bone loss is progressive, based on serial radiographs, and estimating severity of bone loss (measured by the percentage of end plate involved) is recorded. Additional relevant bony changes and device observations include radiolucent lines, heterotopic ossification, vertebral body olisthesis, loss of core implant height, and presence of device migration, and subsidence.ResultsSerial radiographs from 19 patients (25 devices) implanted with a variety of cTDR designs were assessed by 6 investigators including clinicians and scientists experienced in cTDR or appendicular skeleton joint replacement. The overall agreement of assessments ranged from 49.9% (95% bootstrap confidence interval 45.1–73.1%) to 94.7% (95% CI 86.9–100.0%). There was reasonable agreement on the presence or absence of bone loss or radiolucencies (range: 58.4% (95% CI 51.5–82.7%) to 94.7% (95% CI 86.9–100.0%), as well as in the progression of radiolucent lines (82.9% (95% CI 74.4–96.5%)).ConclusionsThe novel classification system proposed demonstrated good concordance among experienced investigators in this field and represents a useful advancement for improving reporting in cTDR studies.
- Research Article
7
- 10.1097/01.blo.0000201163.97937.1a
- May 1, 2006
- Clinical Orthopaedics and Related Research
Hemostasis in bone is difficult to achieve because of the mineral content. Current techniques often are ineffective, can have systemic effects, or leave residual material in the wound. Our hypotheses were that a wand device coupling radiofrequency energy with a cooling conductive saline solution, applied topically to bone, could produce superior hemostasis compared with conventional electrocautery or no treatment, and not impede bone healing. Immediate hemostasis and subsequent bone healing for 6 and 12 weeks were evaluated in an iliac crest ostectomy (cancellous bone) and a drilled tibia defect (cortical bone) sheep model. Outcome variables were amount and intensity of bleeding, serial radiography, quantitative computed tomography, histology and mechanical testing. Control of bleeding was nearly complete (93%) and greater with the radiofrequency/saline treatment compared with electrocautery (56%) or no treatment (0%) in cancellous bone and cortical bone. Electrocautery induced surface char (black carbon debris) that could be seen at 6 and 12 weeks. There were no differences in bone healing between the radiofrequency and electrocautery device applications or untreated bone. At 12 weeks, all healing tibiae defects were as strong as undrilled tibiae. This may be an effective method to produce rapid hemostasis in bone without char or healing complications.
- Research Article
224
- 10.1378/chest.129.2.333
- Feb 1, 2006
- Chest
Pulmonary Cryptococcosis: Comparison of Clinical and Radiographic Characteristics in Immunocompetent and Immunocompromised Patients
- Research Article
8
- 10.3390/jcm14144929
- Jul 11, 2025
- Journal of clinical medicine
Background: Renal tumors, encompassing benign, malignant, and normal variants, represent a significant diagnostic challenge in radiology due to their overlapping visual characteristics on computed tomography (CT) scans. Manual interpretation is time consuming and susceptible to inter-observer variability, emphasizing the need for automated, reliable classification systems to support early and accurate diagnosis. Method and Materials: We propose KidneyNeXt, a custom convolutional neural network (CNN) architecture designed for the multi-class classification of renal tumors using CT imaging. The model integrates multi-branch convolutional pathways, grouped convolutions, and hierarchical feature extraction blocks to enhance representational capacity. Transfer learning with ImageNet 1K pretraining and fine tuning was employed to improve generalization across diverse datasets. Performance was evaluated on three CT datasets: a clinically curated retrospective dataset (3199 images), the Kaggle CT KIDNEY dataset (12,446 images), and the KAUH: Jordan dataset (7770 images). All images were preprocessed to 224 × 224 resolution without data augmentation and split into training, validation, and test subsets. Results: Across all datasets, KidneyNeXt demonstrated outstanding classification performance. On the clinical dataset, the model achieved 99.76% accuracy and a macro-averaged F1 score of 99.71%. On the Kaggle CT KIDNEY dataset, it reached 99.96% accuracy and a 99.94% F1 score. Finally, evaluation on the KAUH dataset yielded 99.74% accuracy and a 99.72% F1 score. The model showed strong robustness against class imbalance and inter-class similarity, with minimal misclassification rates and stable learning dynamics throughout training. Conclusions: The KidneyNeXt architecture offers a lightweight yet highly effective solution for the classification of renal tumors from CT images. Its consistently high performance across multiple datasets highlights its potential for real-world clinical deployment as a reliable decision support tool. Future work may explore the integration of clinical metadata and multimodal imaging to further enhance diagnostic precision and interpretability. Additionally, interpretability was addressed using Grad-CAM visualizations, which provided class-specific attention maps to highlight the regions contributing to the model's predictions.
- Research Article
48
- 10.1007/s00198-023-06900-w
- Sep 5, 2023
- Osteoporosis International
To explore the feasibility of using deep learning to establish a model for osteoporosis classification and bone density value prediction based on opportunistic CT scans and to verify its generalization and diagnostic ability using an independent test set. A total of 1219 cases of opportunistic CT scans were included in this study, with QCT results as the reference standard. The training set: test set: independent test set ratio was 703: 176: 340, and the independent test set data of 340 cases were from 3 different hospitals and 4 different CT scanners. The VB-Net structure automatic segmentation model was used to segment the trabecular bone, and DenseNet was used to establish a three-classification model and bone density value prediction regression model. The performance parameters of the models were calculated and evaluated. The ROC curves showed that the mean AUCs of the three-category classification model for categorizing cases into "normal," "osteopenia," and "osteoporosis" for the training set, test set, and independent test set were 0.999, 0.970, and 0.933, respectively. The F1 score, accuracy, precision, recall, precision, and specificity of the test set were 0.903, 0.909, 0.899, 0.908, and 0.956, respectively, and those of the independent test set were 0.798, 0.815, 0.792, 0.81, and 0.899, respectively. The MAEs of the bone density prediction regression model in the training set, test set, and independent test set were 3.15, 6.303, and 10.257, respectively, and the RMSEs were 4.127, 8.561, and 13.507, respectively. The R-squared values were 0.991, 0.962, and 0.878, respectively. The Pearson correlation coefficients were 0.996, 0.981, and 0.94, respectively, and the p values were all < 0.001. The predicted values and bone density values were highly positively correlated, and there was a significant linear relationship. Using deep learning neural networks to process opportunistic CT scan images of the body can accurately predict bone density values and perform bone density three-classification diagnosis, which can reduce the radiation risk, economic consumption, and time consumption brought by specialized bone density measurement, expand the scope of osteoporosis screening, and have broad application prospects.
- Research Article
- 10.7759/cureus.108377
- May 1, 2026
- Cureus
Introduction: The aim of this study is to evaluate bone healing after enucleation of cystic jaw lesions using region of interest (ROI)-based gray value analysis on panoramic radiographs. Additionally, this study aims to quantitatively assess early postoperative changes in the lesion area and to determine whether grayscale values approach those of the contralateral healthy bone. Furthermore, it seeks to investigate the applicability of panoramic radiography as a reliable and accessible tool for objective radiographic follow-up in routine clinical practice. Therefore, grayscale analysis provides an indirect, quantitative surrogate of bone healing based on relative pixel intensity changes, rather than a direct histological or volumetric assessment of bone healing.Methodology: In this retrospective study, 15 patients were selected from approximately 200 patients who presented to our clinic between June 1, 2025, and February 15, 2026, with cystic lesions ≥1.5 cm treated solely by enucleation. Postoperative panoramic radiographs were obtained three to six months after surgery. All images were acquired using the same digital panoramic system under fixed parameters (65 kVp, 5 mA, 17 s). Diagnosis was confirmed histopathologically, and all cases were radicular or dentigerous cysts. Image analysis was performed using ImageJ software (v1.54k; National Institutes of Health, Bethesda, MD). Standardized ROIs (60 × 60 pixels) were defined in the lesion area and the contralateral control area on pre- and postoperative radiographs. Gray values were measured twice and averaged.Results: The mean grayscale value of the lesion area increased from 77.54 ± 31.19 preoperatively to 94.57 ± 34.05 postoperatively, representing an approximate 22% increase (P < 0.001). In addition, postoperative grayscale values in the lesion region approached those of the contralateral control areas. The observed change demonstrated a large effect size, indicating that the radiographic improvement was not only statistically significant but also clinically meaningful.Conclusions: ROI-based grayscale analysis is a practical, reproducible, and cost-effective method for assessing postoperative radiographic changes related to bone healing after enucleation of cystic jaw lesions. This analysis, performed via panoramic radiographs, provides objective and reliable data for clinical follow-up. Furthermore, this approach may serve as a cost-effective and accessible alternative for quantitative monitoring of early bone healing in routine clinical practice, especially when advanced imaging modalities are not readily available.
- Research Article
1
- 10.3390/diagnostics15151862
- Jul 24, 2025
- Diagnostics
Objectives: Spinal diseases are commonly encountered health problems with a wide spectrum. In addition to degenerative changes, other common spinal pathologies include metastases and compression fractures. Benign tumors like hemangiomas and infections such as spondylodiscitis are also frequently observed. Although magnetic resonance imaging (MRI) is considered the gold standard in diagnostic imaging, the morphological similarities of lesions can pose significant challenges in differential diagnoses. In recent years, the use of artificial intelligence applications in medical imaging has become increasingly widespread. In this study, we aim to detect and classify vertebral body lesions using the YOLO-v8 (You Only Look Once, version 8) deep learning architecture. Materials and Methods: This study included MRI data from 235 patients with vertebral body lesions. The dataset comprised sagittal T1- and T2-weighted sequences. The diagnostic categories consisted of acute compression fractures, metastases, hemangiomas, atypical hemangiomas, and spondylodiscitis. For automated detection and classification of vertebral lesions, the YOLOv8 deep learning model was employed. Following image standardization and data augmentation, a total of 4179 images were generated. The dataset was randomly split into training (80%) and validation (20%) subsets. Additionally, an independent test set was constructed using MRI images from 54 patients who were not included in the training or validation phases to evaluate the model’s performance. Results: In the test, the YOLOv8 model achieved classification accuracies of 0.84 and 0.85 for T1- and T2-weighted MRI sequences, respectively. Among the diagnostic categories, spondylodiscitis had the highest accuracy in the T1 dataset (0.94), while acute compression fractures were most accurately detected in the T2 dataset (0.93). Hemangiomas exhibited the lowest classification accuracy in both modalities (0.73). The F1 scores were calculated as 0.83 for T1-weighted and 0.82 for T2-weighted sequences at optimal confidence thresholds. The model’s mean average precision (mAP) 0.5 values were 0.82 for T1 and 0.86 for T2 datasets, indicating high precision in lesion detection. Conclusions: The YOLO-v8 deep learning model we used demonstrates effective performance in distinguishing vertebral body metastases from different groups of benign pathologies.
- Research Article
- 10.1227/neu.0000000000002375_183
- Apr 1, 2023
- Neurosurgery
INTRODUCTION: Cervical Disc Replacement (CDR) is an increasingly utilized procedure to address cervical degenerative conditions, as an alternative to discectomy and fusion, for the purpose of preserving natural segmental motion. METHODS: Using serial radiographs from patients, a panel of experts assembled for the purpose of creating a classification system to aid in reproducibly and consistently identifying boney changes and assessing CDR device radiographic appearance. Subdividing the superior and inferior vertebral bodies into 3 equal sections, observed bone loss such as endplate rounding, cystic erosion adjacent to the endplate, and cystic erosion not adjacent to the endplate, is recorded. Determining if bone loss is progressive and estimating severity of bone loss (measured by the percentage of end plate involved) is recorded. Additional relevant boney changes and device observations include radiolucent lines, heterotopic ossification, loss of core implant height, and presence of device migration, subsidence, and olisthesis. RESULTS: Serial radiographs from 26 patients implanted with a variety of CDR products were assessed by 6 investigators including clinicians and scientists experienced in CDR or appendicular skeleton joint replacement. After several iterations, a classification system emerged that had improved concordance among the participating investigators. Next steps include assessing reliability with a broad group of clinicians, using serial radiographs from a significant number of patients, to confirm the proposed classification scheme. CONCLUSIONS: A standardized nomenclature for boney changes following CDR will facilitate accurate and reproducible scientific communications regarding the clinical outcomes of this procedure. The novel system proposed demonstrated good concordance among experienced investigators in this field and represents an important advancement.
- Research Article
63
- 10.1088/1741-2552/abb5be
- Oct 1, 2020
- Journal of Neural Engineering
Objective.Automatic sleep staging models suffer from an inherent class imbalance problem (CIP), which hinders the classifiers from achieving a better performance. To address this issue, we systematically studied sleep electroencephalogram data augmentation (DA) approaches. Furthermore, we modified and transferred novel DA approaches from related research fields, yielding new efficient ways to enhance sleep datasets. Approach. This study covers five DA methods, including repeating minority classes, morphological change, signal segmentation and recombination, dataset-to-dataset transfer, as well as generative adversarial network (GAN). We evaluated these mentioned DA methods by a sleep staging model on two datasets, the Montreal archive of sleep studies (MASS) and Sleep-EDF. We used a classification model with a typical convolutional neural network architecture to evaluate the effectiveness of the mentioned DA approaches. We also conducted a comprehensive analysis of these methods. Main results. The classification results showed that DA methods, especially DA by GAN, significantly improved the total classification performance in comparison with the baseline. The improvement of accuracy, F1 score and Cohen Kappa coefficient range from 0.90% to 3.79%, 0.73% to 3.48%, 2.61% to 5.43% on MASS and 1.36% to 4.79%, 1.47% to 4.23%, 2.22% to 4.04% on Sleep-EDF, respectively. DA methods improved the classification performance in most cases, whereas the performance of class N1 showed a subtle degradation in the F1 scores. Significance. Overall, our study proved that DA approaches are efficient in alleviating CIP lying in sleep staging tasks. Meanwhile, this study provided avenues for further improving the sleep staging accuracy using DA methods.
- Research Article
122
- 10.1159/000028694
- Aug 1, 1998
- Pediatric Neurosurgery
Certain CT and/or MRI abnormalities have been used medicolegally to time intracranial injuries from the infant shaken impact syndrome (ISIS). For example, parenchymal hypodensities on CT scans are said to arise only after 6–48 h have elapsed postinjury, and the presence of chronic or mixed subdural hematomas suggests injury that occured 1–4 weeks prior. However, these statements are based largely upon inference from data obtained in other conditions such as ischemic anoxic injury and chronic subdural hemorrhage in adults. Direct evidence about the evolution of intracranial injuries in infants with ISIS is sparse, and the radiographic changes following ISIS have never been systematically studied on serial imaging studies. One hundred-seventeen serial CT and MRI scans obtained from 33 infants with ISIS were reviewed retrospectively. The exact scan dates and times were obtained directly from the scans. Acute subdural hemorrhage was the most common intracranial abnormality and was present in 27 (81%) of the 33 infants. Other intracranial abnormalities included chronic subdural collections, subarachnoid hemorrhage, epidural hematomas, parenchymal hypodensities, edema and contusions, and atrophy and encephalomalacia. In 15 of the 33 infants, the injury could be timed with reasonable certainty, and the evolution of the radiographic changes followed over time. Six of the 15 infants had evidence of prior cranial trauma such as chronic subdural collections (5 infants) or mild atrophy (1 infant). Of the remaining 9 infants, parenchymal abnormalities such as hypodensities, edema and contusion appeared in virtually all of the initial scans performed approximately 3 h following the report of injury. One ‘chronic’ subdural collection was absent on the first scan performed 2.75 h postinjury, but appeared on a second scan performed 17 h later, suggesting that some ‘chronic’ subdural fluid collections may arise much sooner than previously thought. These findings challenge some of the current dogma about the timing of radiographic changes following abuse and are important in timing the alleged abuse for legal purposes.
- Research Article
1
- 10.1038/s41598-025-26228-1
- Nov 25, 2025
- Scientific Reports
Machine learning models are powerful tools for cardiovascular disease (CVD) prediction, but their performance is often limited by dataset size and class imbalance. While data augmentation techniques can address these issues, their impact on model interpretability and the relative importance of clinical predictors remains poorly understood. This study investigates how different data augmentation strategies affect the performance and feature importance hierarchy of an Extreme Gradient Boosting (XGBoost) model for CVD prediction. This study conducted an ablation study using a public CVD dataset. Three XGBoost models were developed and compared: a baseline model trained on original data, a model trained with data augmented by the Synthetic Minority Over-sampling Technique (SMOTE), and a model using a Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP). Model performance was evaluated using accuracy, F1-score, and AUC. Feature importance was quantified and compared across models using the Gain metric. All models demonstrated high predictive performance on the independent test set, with the SMOTE-augmented model achieving an accuracy and AUC of 1.0. Data augmentation fundamentally altered the model’s feature importance. In the baseline model, ‘oldpeak’ (Gain: 8.25) and ‘slope’ (Gain: 7.01) were the top predictors. In contrast, ‘slope’ became the single most dominant feature in both the SMOTE (Gain: 27.49) and WGAN-GP (Gain: 36.68) augmented models. Data augmentation can significantly reshape the predictive strategy of a high-performance machine learning model. For high-quality datasets, the primary effect of augmentation may be the re-prioritization of predictive features rather than a direct improvement in classification accuracy. These findings underscore the critical need to evaluate the impact of synthetic data on model interpretability before clinical application.
- Research Article
41
- 10.1016/s0268-0033(03)00082-2
- May 31, 2003
- Clinical Biomechanics
Enhancing mechanical strength during early fracture healing via shockwave treatment: an animal study
- Research Article
13
- 10.2196/46216
- Jun 1, 2023
- Journal of Medical Internet Research
BackgroundThe growing public interest and awareness regarding the significance of sleep is driving the demand for sleep monitoring at home. In addition to various commercially available wearable and nearable devices, sound-based sleep staging via deep learning is emerging as a decent alternative for their convenience and potential accuracy. However, sound-based sleep staging has only been studied using in-laboratory sound data. In real-world sleep environments (homes), there is abundant background noise, in contrast to quiet, controlled environments such as laboratories. The use of sound-based sleep staging at homes has not been investigated while it is essential for practical use on a daily basis. Challenges are the lack of and the expected huge expense of acquiring a sufficient size of home data annotated with sleep stages to train a large-scale neural network.ObjectiveThis study aims to develop and validate a deep learning method to perform sound-based sleep staging using audio recordings achieved from various uncontrolled home environments.MethodsTo overcome the limitation of lacking home data with known sleep stages, we adopted advanced training techniques and combined home data with hospital data. The training of the model consisted of 3 components: (1) the original supervised learning using 812 pairs of hospital polysomnography (PSG) and audio recordings, and the 2 newly adopted components; (2) transfer learning from hospital to home sounds by adding 829 smartphone audio recordings at home; and (3) consistency training using augmented hospital sound data. Augmented data were created by adding 8255 home noise data to hospital audio recordings. Besides, an independent test set was built by collecting 45 pairs of overnight PSG and smartphone audio recording at homes to examine the performance of the trained model.ResultsThe accuracy of the model was 76.2% (63.4% for wake, 64.9% for rapid-eye movement [REM], and 83.6% for non-REM) for our test set. The macro F1-score and mean per-class sensitivity were 0.714 and 0.706, respectively. The performance was robust across demographic groups such as age, gender, BMI, or sleep apnea severity (accuracy 73.4%-79.4%). In the ablation study, we evaluated the contribution of each component. While the supervised learning alone achieved accuracy of 69.2% on home sound data, adding consistency training to the supervised learning helped increase the accuracy to a larger degree (+4.3%) than adding transfer learning (+0.1%). The best performance was shown when both transfer learning and consistency training were adopted (+7.0%).ConclusionsThis study shows that sound-based sleep staging is feasible for home use. By adopting 2 advanced techniques (transfer learning and consistency training) the deep learning model robustly predicts sleep stages using sounds recorded at various uncontrolled home environments, without using any special equipment but smartphones only.
- Research Article
14
- 10.1007/s00167-021-06506-x
- Feb 28, 2021
- Knee Surgery, Sports Traumatology, Arthroscopy
The purpose of this study was to prospectively investigate osteotomy gap filling rates on serial plain radiographs, and to evaluate whether alignment correction is maintained after medial opening wedge high tibial osteotomy (MOWHTO) using a locking plate without bone graft. Between March 2014 and June 2017, MOWHTO was performed without bone graft regardless of gap size. Radiographs were taken preoperatively, postoperatively, at 1, 3, 6, 12, 18, and 24months after surgery. Radiographic examinations included a weight bearing long-standing anteroposterior (AP) view of the whole lower extremity, as well as, the AP, lateral, and both oblique views of the knee. Bone healing was measured on the medial oblique view of the knee. The postoperative alignment correction and its maintenance were assessed using the three radiologic parameters of the weight-bearing line (WBL) ratio, the hip-knee-ankle angle (HKAA), and the medial proximal tibial angle (MPTA) on the weight-bearing long-standing AP view of the lower extremity. Fifty-two consecutive patients underwent MOWHTO, but three patients failed to follow-up for more than 24months. A total of 49 patients were assessed in this study. The median opening gap height was 10.0mm (IQR, 8.0-12.0; range, 7-20). On immediate post-operative radiographs, the mean gap filling was 31.4 ± 3.6%. After 1, 3, 6, 12, 18, and 24months, the mean gap filling rates increased to 38.7 ± 4.4%, 51.4 ± 6.6%, 66.5 ± 5.1%, 84.8 ± 7.0%, 92.4 ± 5.6%, and 97.8 ± 2.3%, respectively. Statistical differences were observed between all the follow-up evaluations (P < 0.001). Statistical differences in the WBL ratio, HKAA, and MPTA were observed between preoperatively and 1month after surgery (P < 0.001). The mean PTSA increased significantly from preoperatively to postoperatively (P < 0.001). However, no statistical differences were found between the post-operative follow-up radiographs performed for these four values. MOWHTO using a locking plate without bone graft achieved at least 90% bone healing and had no loss in correction at 2 years postoperatively. III.