A novel study developed and validated a blood-based DNA methylation marker model (BBDMM) for predicting lung cancer risk. This model utilizes informative CpG sites identified through epigenome-wide association studies (EWAS). It was developed and internally validated in a German cohort of over 2,400 participants, then externally validated in Norwegian cohorts. The BBDMM demonstrated excellent discriminatory capacity, with AUCs of 0.84 and 0.85, and predictive stability over periods up to 18 years before diagnosis. These findings suggest significant potential for improving lung cancer screening.
Cancer epidemiology, biomarkers & prevention : a publication of the American Association for Cancer Research, cosponsored by the American Society of Preventive OncologyJun 02, 2026
Lung cancer is a lethal malignancy urgently requiring effective early detection strategies, as current cfDNA-based approaches often lack sensitivity in early stages. This study developed a novel computational feature, First-Order Transition Probability (FOTP), to capture nucleotide sequential dependencies within cfDNA fragments. Using low-pass whole-genome sequencing data, an SVM model trained with FOTP achieved 73.9% sensitivity for stage I and 81.8% for stage II lung cancer at 95% specificity. This method, which significantly outperforms existing fragmentomic features, is biologically interpretable and offers a scalable strategy for early cancer screening.
Lung cancer biopsies often yield limited material, complicating predictive molecular testing and sometimes necessitating re-biopsy. This study investigated the feasibility of repurposing routine diagnostic slides, specifically H&E-stained and immunostained sections, for molecular analysis. From 40 lung biopsy specimens, DNA extraction and targeted next-generation sequencing were successfully performed for 33 cases. The results demonstrated high concordance of detected variants, suggesting this approach is a viable alternative for molecular testing.
This study investigates the impact of non-independent and identically distributed (Non-IID) data on the performance of Federated Learning (FL) models for survival prediction using lung cancer data. Researchers compared Random Forest (RF) and AdaBoost algorithms across various scenarios of data distribution shifts among clients. The findings indicate that both algorithms experienced significant performance degradation under highly non-IID conditions. Notably, Random Forest consistently outperformed AdaBoost in these challenging scenarios. The developed FL algorithm holds potential for model personalization and fine-tuning, with broader applicability to other clinical datasets.