
From big data to personalized medicine: bioinformatics perspectives and challenges
Humber College
On this page
Abstract
The rapid accumulation of various types of -omics data is opening new vistas for the development of novel therapeutic approaches grounded in big data. Particularly, personalized medicine stands out as one of the most promising goals. However, it faces significant logical and mathematical challenges. These include the ecological fallacy, which refers to making individual predictions based on average population data, and the Pareto principle, which suggests that approximately 80% of outcomes are driven by 20% of inputs. These raise the possibility of a dramatic discrepancy between our expectations for big data–based personalized medicine and its actual effectiveness. What are the practical issues our community must address to bridge the gap between the promise of big data and the reality of personalized medicine?
Keywords
- big data
- omics
- personalized medicine
- formal fallacies
- statistics
- Pareto principle
Ilya Ioshikhes
Former professor at Humber College Correspondence: iioschik@uottawa.ca Keywords: big data, omics, personalized medicine, formal fallacies, statistics, Pareto principle
Abstract
The rapid accumulation of various types of -omics data is opening new vistas for the development of novel therapeutic approaches grounded in big data. Particularly, personalized medicine stands out as one of the most promising goals. However, it faces significant logical and mathematical challenges. These include the ecological fallacy, which refers to making individual predictions based on average population data, and the Pareto principle, which suggests that approximately 80% of outcomes are driven by 20% of inputs. These raise the possibility of a dramatic discrepancy between our expectations for big data–based personalized medicine and its actual effectiveness. What are the practical issues our community must address to bridge the gap between the promise of big data and the reality of personalized medicine? Our era may be scientifically defined as the era of super-massive data analysis, in which incredibly large amounts of data from various fields of knowledge are examined, leading to profound conclusions through statistical inference. Generally speaking, statistical estimates are more reliable when they are supported by large amounts of data, in accordance with the Law of Large Numbers.1 For instance, in medical diagnostics, the use of specific disease symptoms and biochemical test results yields more reliable outcomes when based on a larger number of observations. Of course, it is not only the quantity of data that matters but also its quality, which may be influenced by factors such as data curation, processing, diversity, reproducibility, explainability, selection bias, noisy labels in image banks, genomic linkage errors, etc.
Thus, while the following perspective emphasizes the quantity of data, it also acknowledges quality as an indispensable component. Billions of individuals could potentially improve their health through tailored treatments.2 People would gain access to individualized medicine,3,4 with care plans designed specifically for them based on their unique genotype, other -omics data, and insights derived from statistical analyses of next-generation large-scale datasets. This approach would not only aid in curing diseases but also in preventing them in a manner best suited to each individual patient. But are statistical results truly applicable to individuals? To begin exploring this question, let us first consider how it applies in other areas of science. Statistical physics examines the distribution of parameters within a statistical ensemble (e.g., the energies of gas molecules). Although these parameters vary over time and from one molecule to another, the overall statistical distribution remains stable, resulting in consistent average values such as energy and temperature for the ensemble. Clearly, this is not the case for each individual molecule at any given moment; the behavior of average statistical parameters differs from that of the corresponding parameters at the level of individual molecules. Eukaryotic gene promoters contain a variety of DNA promoter elements, some more universal than others. Historically, the TATA box was the first general promoter element to be discovered, and all initially described eukaryotic promoters contained it. Later, TATA-less promoters were identified, leading to the hypothesis that another general promoter element must be present in these sequences. Such an element, known as the Initiator (Inr), was soon discovered. However, as more promoter sequences were collected, it became evident that many promoters lack both the TATA box (in over 80 % of cases) and the Inr. Although hundreds of additional promoter elements have since been identified, none are, on their own, necessary or sufficient for a genomic sequence to function as a promoter. Still, TATA box and Inr are the most important promoter elements from a statistical perspective. Many individual promoters don’t contain them, and other promoter elements are even less universal. This highlights a clear discrepancy between the general functioning of the gene transcription machinery and its operation at the level of individual promoters. More systematically, similar phenomena have been defined as the ecological (inference) fallacy.5, 6 This is a logical error that arises when conclusions about individual members of a group are drawn from aggregate data about the group as a whole. The fallacy lies in assuming that what holds true for a population must also apply to each of its individual members, an assumption that can lead to inaccurate or misleading conclusions. In the case of omics big data, the number of parameters associated with each individual is astronomical (and includes all nucleotides in their genome, the characteristics of their single cells, environmental conditions, and more). This number is much greater when considering the entirety of humankind, or other large communities such as soil bacterial populations. As a result, even with the most powerful computers, analyzing omics big data requires selecting only the most statistically significant features, usually achieved through dimensionality reduction.7 This involves omitting the majority of features from further analysis, based on the assumption that the remaining parameters still capture the most important features of the system, while making computational analysis cheaper and more feasible.8,9
Based on the key parameters selected in the previous step and the results of computational analysis for a given population (defined socially, geographically, etc.), each individual within that population may be assigned certain scores. These scores can guide the selection of medical treatments tailored to the individual’s unique characteristics. Such treatments would be (almost) perfectly suited to the person, informed by the vast amounts of relevant big data. We would require data on patient responses to therapies to achieve this level of stratification and personalized treatment. This approach paves the way for the development of tailored therapeutic strategies, an approach known as individualized medicine. The promise of individualized medicine should benefit all of us, but unfortunately, this is precisely where the ecological fallacy comes into play: “inferring about the nature of an entity (individual -II) based solely upon aggregate statistics collected for the group (population -II) to which that entity belongs”.10 The ecological fallacy is a type of formal fallacy that may arise in the interpretation of statistical data, as discussed above. A formal fallacy, in turn, is a line of reasoning invalidated by a flaw in its logical structure.11 So, where exactly was the flaw in the reasoning outlined in the previous paragraph? According to Freedman, “the ecological fallacy consists in thinking that relationships observed for groups necessarily hold for individuals”.12 Rather, parameter correlations for the population and for the individual are not necessarily the same, and perhaps are, in fact, usually not the same. To understand how significant this difference can be, let us consider the Pareto Principle, also known as the 80/20 rule. Originally, economist Vilfredo Pareto observed that 80 % of the land in the Kingdom of Italy was owned by 20 % of the population at that time.13 This observation was later generalized into what is now known as the Pareto Principle, that in many cases, 80 % of the effects result from 20 % of the causes.14 This principle has been validated across various fields, including business, economics, and quality control, e.g., the finding that 80 % of meaningful outcomes often come from just 20% of a company’s employees. Although this principle has not yet been widely confirmed in omics data, it has been demonstrated in at least some contexts.15-19 If applied to our case, it could suggest that 80 % of medical recommendations are derived from just 20% of the population. But what about the remaining 80 %? For them, these recommendations likely do not work so well. Otherwise, they would fall within that initial 20 %. As a result, only about 20 % of new patients might reasonably expect that big data-based individualized medicine will be effective for them. Clearly, this falls short of the efficiency gains we aspire to achieve through big data analysis. Thus, rather than relying solely on statistics, it might be imperative to use more profound machine learning (ML) and artificial intelligence (AI)-driven models that can capture individual patient traits and drive personalized predictions. These models can better stratify patient populations and perform cross-validations that improve generalizability to the 80 % portion of the population and help detect unseen trends. However, a fundamental problem remains unresolved: while training ML and AI models on big data may enhance diagnosis and treatment at the population level, will it truly improve outcomes for each individual? Furthermore, while the ecological fallacy and the Pareto principle illustrate fundamental errors in inference and the distribution of effects, they represent only a fraction of the obstacles that separate big data from personalized medicine (PM). In practice, the following technical issues must also be addressed (this list is not exhaustive):
1. Data quality and curation: Data curation, or data cleaning, is a basis of the data quality assessment and a key prerequisite of any data analytics services. Given the massive volumes of data to be analyzed, medicine would benefit greatly from automated data curation systems. The analysis of heterogeneous, multidisciplinary data requires the “detection and tracking of inconsistencies, missing values, outliers, and similarities,” as well as data standardization processes.20 2. Lack of collection standards, inconsistent labels, missing variables, and selection bias in electronic health records (EHRs) and biobanks: Since medical decisions are made based on heterogeneous information and criteria provided by patients, providers, and health systems, applying standard missing data techniques is often problematic.21 Standardized data provenance (the documentation of data origins, and its generation and processing history) based on blockchain technology shows promise. However, several challenges still remain, such as provenance security, communication overhead, scalability, revocation, and General Data Protection Regulation compliance.22 3. Need for robust cleaning, harmonization, and continuous audit pipelines: The continuous growth of digital datasets originating from diverse sources (sensor networks, social media, business transactions, etc.) demands the development of effective data management strategies. These must include robust data cleansing pipelines to detect and repair dirty data and to address inherent uncertainties caused by noise, missing values, inconsistencies, and other issues.23 4. Representativeness and equity: Many groups that are underrepresented or excluded in clinical research due to gender, age, race, or ethnicity have been shown to exhibit distinct disease presentations and responses to drugs or therapies. These differences may result from specific genotypes, epigenetic effects, or non-genetic lived experiences such as unconscious racism and variable socioeconomic and educational levels. Proper representativeness is critical for generalizability and population-specific relevance of medical findings.24 5. Underrepresentation of ethnic minorities: Research has shown that individuals from ethnic minority groups participate in medical research at lower rates than their white counterparts. This may result in invalid treatment generalization across all ethnic groups, without understanding how minority populations are affected, and may contribute to health inequalities and maltreatment for these populations.25 6. Data balancing strategies: Big data analytics in healthcare involves large-scale data from EHRs, wearable devices, and genomic studies, and can improve clinical decision-making, manage disease outbreak, and personalize treatment plans. However, it also introduces significant security and privacy challenges, with sensitive patient information becoming a target for cyber threats, data breaches, and unauthorized access. Hence, the benefits of data-driven medicine should be balanced with stringent data security, regulatory compliance, and ethical considerations.26 7. Multidimensional integration: The numerous benefits of big data analysis in healthcare settings are inhibited by multiple technological and cultural obstacles, privacy issues, and the critical need for standardized procedures to efficiently explore the huge volume of healthcare data27 coming from databases, EHRs and other documents, medical devices and sensors.28 Dimensionality reduction is indispensable for such analysis,29, 30 yet may result in loss of crucial individual information necessary for PM. 8. Challenge of combining genomics, proteomics, clinical and lifestyle data in real time: While the generation and integration of multi-omics data facilitate PM, the complexity of data integration
across different omics layers, the need for advanced computational tools, the high cost of comprehensive data generation, and issues related to data privacy, standardization, and the need for robust validation in diverse populations remain significant obstacles. Real-time integration of genomic data into clinical workflows will allow clinicians to use this data for personalized diagnosis, treatment planning, and preventive care.31 9. Interoperability and data governance: To transform medical decision-making, an unprecedented amount of data must be collected across diverse sectors. Standards for data sharing and interoperability, along with robust data stewardship and governance models, should be established to ensure the protection and integrity of the public health data system.32 10. Overfitting in datasets: ML models may outperform statistical regression algorithms in clinical predictions. Proper validation and careful tuning of these models are essential to avoid overfitting, where the model closely resembles the training dataset, and to ensure strong generalizability.33 11. Lack of explicit causality: In medicine, causality is generally expressed probabilistically, that is, a specific event carries a certain probability of leading to a specific outcome, as determined through randomized controlled trials (RCTs). For example, exposure to a risk factor (like cigarette smoking or diesel emissions) increases the probability of developing lung cancer, whereas certain treatment (such as a specific type of chemotherapy) increases the likelihood of survival in patients with a cancer diagnosis, although does not guarantee it. If RCTs are not feasible for practical or ethical reasons, observational studies should be used.34 12. Limited explainability: Although ML-based AI applications in healthcare can provide numerous benefits, they also present significant issues concerning explainability. These models often resemble a “black box,” offering no clear reasoning for a given conclusion. This makes understanding and justifying the models’ outcomes extremely difficult and adds an extra level of complexity to the decision-making process.35 13. Security, privacy, and regulation: The digitization of healthcare brings numerous benefits, including improved efficiency, enhanced patient care, and advanced data analytics. However, it also raises significant concerns regarding security and privacy, requiring the appropriate collection, use, and sharing of personal health information while preserving patient’s autonomy and confidentiality. The latter require robust authentication and access controls, encryption and data protection measures, continuous security monitoring, privacy-by-design principles, and effective training and awareness programs36 14. Hospital infrastructure barriers, staff training: The complexity of PM approaches and their associated challenges requires the development of effective engagement strategies, collaboration frameworks, infrastructure, education and training programs, ethical policies, resource allocation guidelines, regulatory compliance measures, and robust data management and privacy practices. All of these must be implemented by appropriately trained staff.37 Our growing community is focused on exploring the challenges and opportunities that arise at the intersection of big data and PM. As this field continues to evolve, contributions from statistics, data analysis, and biomedical research play a key role in shaping its direction. As we confront these intellectually demanding and practically impactful challenges, what perspectives (disciplinary, methodological, or philosophical) will help move the field forward?
