The National Institutes of Health recently announced a remarkable achievement: the world’s largest database of human genomes, a repository of information that the agency says will help usher in a new era in personalized medicine.
Scientists are already developing clinical tools and techniques based on genomic information gathered from the agency’s All of Us program. One of the most promising? Polygenic risk scores. These can help predict a person’s likelihood of developing complex diseases, like cardiovascular disease and breast cancer, from as early as infancy.
But there’s a glaring problem: The tools were trained overwhelmingly on the DNA of people of European descent, so they often fail to accurately assess risks for anyone else. For some conditions, the forecasts for people of color are almost no better than flipping a coin.
Researchers are racing to correct the biases through advancements in modeling and widespread recruitment of minority groups. They fear that failing to close the gaps could worsen health care disparities — and prevent the technology from realizing its promise of helping to decrease chronic disease.
“It really warrants attention,” said Eimear Kenny, the director of the Institute for Genomic Health at Mount Sinai and a principal investigator in a national consortium that tests the genetic scores.
Geneticists measure a polygenic risk score by aggregating hundreds or thousands of tiny variations across the genome, many of which might nudge a person’s disease risk by only a fraction. By bundling them together, they can generate a cumulative score, ranking a person’s predisposition to common killers.
For patients at the extremely high end of the risk distribution, the scores can offer a medical advantage, potentially giving doctors a decades-long runway to intervene. Those in the highest percentiles for a cardiovascular disease score, for instance, can be prescribed cholesterol-lowering statins earlier, potentially helping prevent some heart attacks and strokes.
But the entire premise rests on the data used to train the model in the first place. The UK Biobank — established in the early 2000s and widely considered a leading genomic trove — is overwhelmingly homogenous; roughly 94 percent of its participants are white. The Million Veteran Program, a repository run by the U.S. Department of Veterans Affairs, is almost three-quarters white.
The cohort in the All of Us program, built by the N.I.H. with the goal of diversifying data, is still about 50 percent European, according to Alicia Martin, a statistical geneticist at the Broad Institute. The effort’s future is also at risk: One of its major funding streams, the 21st Century Cures Act, is set to expire at the end of this fiscal year. (The program’s budget has already been reduced by 72 percent since 2023.)
“You can’t just snap your fingers and overnight have a biobank from Africa, a biobank from India, a biobank from China that are all as large and open and comprehensive as the UK Biobank,” Dr. Martin said. “We can develop all the fancy statistical methods we want, but still not be able to overcome without more diverse representation.”
Other clinical tools could be affected by the ancestry bias, such as integrated artificial intelligence models that draw upon genomics and other biomedical data. Doctors who sequence a genome to look for rare mutations may also find genetic variants in patients of color that researchers have not studied enough to understand.
Genetic blind spots can also cause scientists to miss discoveries that benefit everyone. One example: the PCSK9 gene. Scientists found that a rare mutation that occurred mostly in a small proportion of Black people could naturally reduce the risk of developing coronary heart disease by 88 percent. That set off a race to develop drugs that mimic the mutation for people of all races and ethnicities. But scientists only found the variant because one study in Dallas had ensured that its participant pool was more than half Black.
Major demographic events in history, like mass migrations, explain why genetic findings from people in Africa or of African descent are vital to study. Communities there have accumulated great genetic diversity over hundreds of thousands of years. By contrast, present-day Europeans descend from a small group of Africans who expanded out of Africa about 50,000 years ago. As a result, they have much less diversity.
“If you’re going to pick a population to study for the sake of everybody’s good, the Europeans are the worst you could choose,” said Jay Kaufman, an epidemiologist at McGill University who has written about race and genomics.
The most obvious solution is to recruit more people of color, something the All of Us program says it has prioritized. The group works with local partner organizations like churches and community health centers to host discussion sessions and enroll rural communities through a mobile clinic. It also returns individualized health findings and genetic risks to the participants to establish a sense of reciprocity.
Still, recruitment can be an upstream endeavor, particularly in the United States, which has a history of mistreatment of people of color, from the Tuskegee Syphilis Study to discriminatory sickle-cell screening to the story of Henrietta Lacks, a Black patient whose cancer cells were used for research without her knowledge or consent.
“Why would you contribute if you’ve been wronged in that way in the past?” Dr. Martin said. “There is some earned mistrust that needs to be addressed — to have that level of trust to be willing to say, ‘Take my data, monitor it for decades, do whatever you want with it.’”
In the meantime, researchers designing polygenic risk scores have found technical ways to diversify data. Modeling tools can now pull weighted information from a variety of sources simultaneously, rather than drawing only from the largest European cohorts and then forcing their algorithms onto other groups. Mount Sinai Health System in New York and U.C.L.A. in Los Angeles, for example, have both established more diverse biobanks than the UK Biobank, with each of them containing tens of thousands of genomes.
Scientists also draw from rapidly growing biobanks in China, Japan, South Korea and Taiwan, which bolster the accuracy of polygenic risk scores for East Asian groups. Localized initiatives are underway in Peru, Mexico, Qatar and elsewhere. And a continentwide program called H3Africa is collecting genetic data in African populations and supporting efforts by scientists there to study how genes and environments cause diseases.
Academic networks have also tapped sources beyond biobanks, integrating data from large, disease-specific cohort projects, including the Atherosclerosis Risk in Communities Study, or ARIC, which intentionally enrolled thousands of Black participants from North Carolina and Mississippi.
For breast cancer, a project called Confluence aggregates data from over 300 individual studies across 62 countries, providing information from 400,000 breast cancer cases and more than 1.5 million controls. The influx of data from minority groups is being used to sharpen breast cancer polygenic risk scores for everyone.
Private genomics companies and start-ups looking to develop and sell polygenic risk scores do not typically have access to those vast networks. Until recently, they relied almost exclusively on data from the UK Biobank. But the All of Us program recently expanded its policies so that companies can use its more diverse data.
Will all of these efforts, taken together, be enough? It depends on whom you ask. Some experts argue that even an ethnically representative data pool is still imperfect, since variations in other factors — socioeconomic status, health access, age or even sex — can still cause a score’s accuracy to decay. Others believe that the statistical issue was overblown in the first place.
Dr. Kenny, who has been helping to test existing polygenic risk scores for 11 common conditions in racially and ethnically diverse patients, argues that the bias is too complex to resolve quickly.
“I don’t think these things are perfectly portable yet — and maybe never will be perfectly portable,” she said. “But there are things that are starting to narrow that gap.”
The post The Best Genetic Risk Tools Don’t Work Equally for Everyone appeared first on New York Times.




