Global genomic and health databases, crucial for AI-driven disease prediction and treatment, disproportionately rely on data from populations of European descent, marginalising South Asians. Despite South Asia experiencing high rates of diabetes and cardiovascular disease, tools like polygenic risk scores built on skewed data are less accurate for this demographic. The lack of representation highlights a critical need for regional data sovereignty and infrastructure, such as the , to ensure equitable healthcare outcomes.
The article highlights a profound data equity issue in global health governance. Because health policies and treatments are increasingly informed by big data and AI models trained on European populations, South Asians are at a disadvantage. This lack of diverse data means global health tools may not be effective for a massive portion of the world's population, impacting their right to health. The development of regional databases like the GenomeIndia Project reflects an effort toward data sovereignty, ensuring that health research addresses the specific genetic profiles and healthcare needs of the Indian population rather than relying on inferred data from the Global North. Furthermore, the lack of data sharing and harmonisation between different South Asian cohorts points to a need for better governance and collaboration in regional scientific research, similar to the frameworks suggested in the Lancet Regional Health.
The genetic diversity within South Asia, shaped by millennia of endogamy and distinct cultural practices, requires nuanced research. The article correctly identifies that treating South Asia as a single genetic monolith obscures vital differences in disease susceptibility. For example, traits like G6PD deficiency vary significantly across communities, and cardiometabolic risks manifest differently even within the same population. The reliance on European data for tools like polygenic risk scores (which estimate genetic risk based on multiple variants) means these tools are fundamentally flawed when applied to South Asians. This creates a disparity in healthcare quality, as treatments tailored for European genetics may be ineffective or even harmful to South Asians. Building diverse biobanks is therefore crucial for achieving health equity and addressing the disproportionately high burden of non-communicable diseases (NCDs) like type 2 diabetes in the region.
The current disparity in global health databases is rooted in historic inequities in research funding. As the article notes, despite low- and middle-income countries (LMICs) bearing the brunt of the global disease burden, research infrastructure has historically been concentrated in wealthier nations. This creates a cycle where LMICs lack the laboratory infrastructure and biobanking facilities necessary to generate large-scale data, further entrenching their reliance on Western models. Investing in domestic genomic research, as seen with initiatives like GenomeIndia Project and Phenome India, is not just a scientific necessity but an economic one. Developing accurate, localised predictive models can lead to earlier disease detection and more effective treatments, ultimately reducing the massive economic burden posed by managing chronic diseases like diabetes, which is projected to affect 125 million people in India by 2045. Furthermore, ensuring intellectual property rights and technology transfer for South Asian researchers is vital for retaining the economic value derived from regional genomic data.