South Asians are missing from global health databases: why this matters and what needs to change

South Asians are missing from global health databases: why this matters and what needs to change Advances in AI are allowing scientists to mine vast amounts of genomic data to detect disease earlier, predict risk and tailor treatments to individuals.

HealthNews Info Wire7 min read
South Asians are missing from global health databases: why this matters and what needs to change

South Asians are missing from global health databases: why this matters and what needs to change Advances in AI are allowing scientists to mine vast amounts of genomic data to detect disease earlier, predict risk and tailor treatments to individuals.

Article outline

  1. What happened
  2. The key numbers
  3. Why it matters
  4. Reaction
  5. Background
  6. The bottom line

Key points

  • (Rupsy Khurana is science communication and outreach lead at the National Centre for Biological Science, Bengaluru.
  • This person is an associate professor of genetics and genomic sciences, and AI in human health, at the Icahn School of Medicine, United States.
  • South Asians face higher rates of type 2 diabetes, cardiovascular disease and asthma than individuals of European ancestry.
  • Diagnostic thresholds, risk scores and prediction models developed predominantly from European populations should be validated and, where necessary, recalibrated using South Asian data.
  • Another study from Sri Lanka discovered that cardiometabolic risk did not fit into a single metabolic syndrome profile.

South Asians face higher rates of type 2 diabetes, cardiovascular disease and asthma than individuals of European ancestry. It means the tools built on European-heavy data are less accurate for the population that needs them most. Image applied for representational purposes only Photo Credit: Getty Images/iStockphoto.

Globally, more than one in 10 adults now live with diabetes. If you have South Asian roots, that risk is even higher, and it hits earlier than it does in plenty of other populations. In India alone, the number of individuals with diabetes is projected to reach 125 million by 2045.

When scientists try to understand why diseases such as diabetes and cardiovascular disease affect South Asians differently, they often have to rely on genetic data drawn from European populations, yet.

Advances in artificial intelligence and machine learning are allowing scientists to mine vast amounts of genomic and health data to detect disease earlier, predict risk, monitor patients and tailor treatments to individuals. But at the heart of this changing landscape lies an old, constant difficulty – the data employed to build these tools lack diversity.

Integrated biobanks such as the U.K. Biobank, which combine participants' genomic information with electronic health records, environmental exposures and lifestyle data, have transformed biomedical research. These repositories have accelerated drug development, informed clinical guidelines and supported shape public health policy throughout the world.

But South Asians remain largely absent from these datasets. While South Asians accounted for less than 1%, the NHGRI-EBI GWAS Catalogue, an online database of human genome-wide association studies, demonstrates that between 2005 and 2025, more than 86% of participants in these studies were of European ancestry.

"More than 20% of the world is being neglected in multi-modal data integration, " remarked Bhramar Mukherjee, senior associate dean of public health data science and data equity, Yale School of Public Health, United States. "It denies the human right and opportunity to attain the maximal possible health".

Notably, the pattern repeats in newer tools, too. A recent study published in Cell Genomics reviewed more than 13, 500 samples throughout three major single-cell resources-the Human Cell Atlas, the Human Tumour Atlas Network and the PsychAD Consortium and identified discovered a "striking, pervasive European overrepresentation and underrepresentation of Asian and Latino individuals."

"These single-cell atlases are becoming the reference maps for biology and medicine, and they are increasingly used to train the AI models that will shape future research and care, " remarked Kuan-lin Huang, senior author of the study. This person is an associate professor of genetics and genomic sciences, and AI in human health, at the Icahn School of Medicine, United States. The data gap is a health gap.

These data points are more than just statistics. South Asians face higher rates of type 2 diabetes, cardiovascular disease and asthma than individuals of European ancestry. It means the tools built on European-heavy data are less accurate for the population that needs them most.

Take polygenic risk scores. It combine the effects of plenty of genetic variants associated with a disease to estimate a person's overall genetic risk. A 2023 study discovered that polygenic risk scores for multiple sclerosis were less accurate when applied to South Asian populations.

"Most predictions regarding how variants affect gene expression or cell function are inferred from European datasets, and we don't know which of those predictions hold in South Asians. This limits our ability to understand disease mechanisms and identify drug targets relevant to South Asian populations, " remarked Shweta Ramdas, a geneticist based in Bengaluru.

Even the well-established measures of disease risk can vary between populations. Genetic traits such as G6PD deficiency. It can cause a type of anaemia, vary considerably throughout South Asia, with some ethnic groups in Pakistan and Afghanistan carrying the trait at much higher rates than others.

Another study from Sri Lanka discovered that cardiometabolic risk did not fit into a single metabolic syndrome profile. Even within the same population, men and women indicated distinct patterns of obesity, blood sugar, cholesterol and blood pressure.

"Diagnostic thresholds, risk scores and prediction models developed predominantly from European populations should be validated and, where necessary, recalibrated using South Asian data. Locally generated evidence is essential for equitable and accurate health care, " remarked Athula Sumathipala, director, Institute for Research and Development in Health and Social Care, Sri Lanka. South Asia is not one population.

South Asia constitutes one of the most diverse human populations in the world, shaped by thousands of years of migration, cultural diversity, endogamy and consanguineous marriages. Much of the existing research does not reflect this diversity. South Asians, Southeast Asians, West Asians and other Asian populations are often lumped together, obscuring significant differences between them.

India alone illustrates how much diversity can disappear when populations are treated as a single group. The GenomeIndia Project, introduced in 2020 to capture the country's genetic diversity and build a reference database for Indian populations, has already identified more than 40 million genetic variants unique to the Indian population.

"South Asia, and India in particular, cannot realistically be treated as one genetic block. We need to include distinct endogamous and tribal groups, not just a few urban cohorts, " remarked Dr. Ramdas. "A lot of these harmful variants aren't seen anywhere else".

Research funding, institutions, registries, biobanks and sizeable population cohorts have historically been built and sustained where the capital already was, leaving low- and middle-income countries with inadequate laboratory infrastructure, biobanking facilities and trained personnel to run comparable studies at scale.

"The global health landscape remains deeply unequal. While over 90% of the world's potential years of life lost occurred in low- and middle-income countries (LMICs), only regarding 10% of global health research funding addressed their health needs, " stated Dr. Sumathipala. "This imbalance goes beyond money, shaping whose problems are studied, whose questions are prioritised and whose evidence informs health policy and practice".

For most South Asian countries, genomic research can be tough to prioritise given more immediate and pressing public health needs such as infectious diseases, maternal and child health and non-communicable diseases.

But the region has created strides in the health landscape. Sustained investments in genomics have enabled a number of substantial prospective cohorts and population datasets, including GenomeIndia, Phenome India, Longevity India, the Sri Lankan Twin Registry Biobank and the Pakistan Genome Resource. But these independent cohorts and biobanks are mostly focused on individual diseases or specific populations, and often employ different systems for collecting and storing data. It makes it challenging to bring them together for sizeable genetic studies. India, for example, has a number of sizeable cohorts, but no harmonised system yet exists that lets researchers within and throughout borders work throughout them easily.

Notably, a recent perspective in the Lancet Regional Health – Southeast Asia, authored by scientists throughout India, Pakistan, Bangladesh and Sri Lanka, argues that the region risks being excluded from the genomic revolution unless it builds an infrastructure itself.

While ensuring that South Asian researchers and institutions retain a meaningful role in how their data are applied, the authors propose building greater regional collaboration between existing biobanks and cohorts. As well as giving South Asian researchers an initial period of priority access to their own data, that includes safeguards such as co-authorship for researchers contributing data, joint intellectual property, technology transfer, training and infrastructure backing. They propose to build a system in which existing datasets can speak to each other, populations that have historically been overlooked can be included, and the researchers generating the data can share in the scientific benefits.

This is vital since if the underlying data continues to stay skewed, the AI models and clinical tools built on top of it will reproduce and repeat those biases-only at a much larger scale.

"It may be late but still not too late, " remarked Dr. Sumathipala.

(Rupsy Khurana is science communication and outreach lead at the National Centre for Biological Science, Bengaluru. [email protected]).

For now, south Asians are missing from global health databases: why this matters and remains the part of the story worth watching, and further updates are likely as more details are confirmed.

Leave a Reply

Your email address will not be published. Required fields are marked *