A Resource-Aware Data Processing Framework for Scalable Whole Exome Variant Analysis in an Under-Represented Asian Cohort
Abstract
Type 2 diabetes T2D is a rapidly growing metabolic and genetic disorder with Asia representing a major disease hotspot However Asian populations remain underrepresented in genomic studies limiting the identification of populationspecific risk variants Moreover largescale sequencing generates large volumes of data posing challenges for computational efficiency storage and variant interpretation This study presents a resourceefficient computational framework for whole exome sequencing WES analysis with integrated visual analytics A unified pipeline was developed by integrating automated quality filtering AFastQF Bowtie2 for alignment and variant detection using DeepVariant The framework is designed to minimize intermediate data generation while maintaining computational performance The framework was evaluated on WES datasets from T2D patients nonT2D individuals and healthy controls from an Asian cohort The proposed framework achieved about 24 reduction in intermediate storage requirements without increasing runtime The analysis identified both known T2Dassociated variants and those that are not in the database Three candidate variants in WFS1 HHEX and FTO genes were prioritized based on in silico pathogenicity predictions and absence from major population databases The study demonstrates that efficient and scalable genomic analysis can be achieved in resourceconstrained environments The proposed framework provides a practical solution for largescale variant discovery and interpretation in underrepresented populations.
Keywords
Asia, framework, resource-aware, type 2 diabetes, variants, whole exome sequencing.