• Search by category

  • Show all

Data, Data Everywhere... II

August 13, 2026

I first read Samuel Taylor Coleridge's The Rime of the Ancient Mariner in infant school. As a child I was captivated by the spectral figures, the rotting sea, and the mariner's terrible isolation. I could not then have articulated the poem's philosophical weight, but I understood instinctively—as children often do—that actions have consequences. How one thoughtless act, a moment's impulse, can unravel the world.

That lesson has stayed with me across four decades in scientific research. And I find myself returning to it. It feels to me that we are surrounded by oceans of data, yet we appear increasingly uncertain how to think. The question that haunts me is whether we have become worshippers at the altar of data, sacrificing the very spirit of scientific inquiry in the process.

The Albatross of Indiscriminate Collection

The mariner's impulsive killing of the albatross was not malicious. It was thoughtless. He acted without considering the consequences, and his single, rash decision condemned his entire crew to suffering. One individual's failure of judgement affected everyone.

Has modern science shot its own albatross? It seems we have become obsessed with collecting ever-larger quantities of data, assuming that if we accumulate enough information—enough variables, enough endpoints, enough omics layers—artificial intelligence and sophisticated statistics will inevitably uncover hidden truths [1]. Historically condemned [2], the current philosophy is rarely examined, let alone justified.

In clinical research, this manifests as protocol bloat. Phase III protocols now average an astonishing 5.96 million data points [3]. A recent study by TransCelerate BioPharma and the Tufts Center for the Study of Drug Development found that nearly one-third of procedures collected in modern trials do not support primary, key secondary, or safety endpoints [4]. We are measuring for the sake of measuring. Each additional endpoint, each exploratory biomarker, each wearable output increases complexity, cost, site burden, and the risk of trial failure [5]. We have transformed elegant clinical investigation into an industrial process of indiscriminate accumulation.

The parallel with environmental destruction, political decision-making, and corporate ethics is unmistakable. In each domain, we see small, seemingly rational decisions aggregating into systemic, Monty Pythonesque insanity. We collect more carbon data but fail to act. We gather more patient data but lose sight of the patient. We accumulate more evidence but neglect wisdom.

The Seduction of Completeness

There is a growing belief that if we simply measure everything, uncertainty will disappear. Yet good science has traditionally embraced uncertainty. Every elegantly designed experiment deliberately ignores thousands of possible measurements because the scientist has already decided which ones matter. This is not laziness; it is a failure of intellectual discipline [6].

When I began my research career, indiscriminate data collection was criticised as a ‘fishing expedition.’ It was seen as a slap-dash approach to science, a substitute for genuine thought. We were trained to begin with a clearly defined hypothesis, to design carefully controlled experiments with minimal but meaningful measurements, and to interpret our results with philosophical clarity [7].

Today, the approach has inverted. We collect everything, analyse everything, and hope something statistically significant emerges. This is one short step from data dredging—what some call ‘p-hacking’ or ‘data torturing’, where we conduct ever more multivariate analyses in the desperate hope of finding 'something' (anything) to compensate for a weak philosophical foundation [8].

The danger is profound. Multiple testing, p-hacking, HARKing (Hypothesising After the Results are Known), and overfitted models generate false discoveries with alarming frequency [9][10]. The more data we collect, the more we find patterns that aren't really there [11]. The cognitive scientist Herbert A. Simon presciently observed: "A wealth of information creates a poverty of attention" [12]. We have filled our laboratories with data and emptied them of thought.

As the poet T.S. Eliot asked in The Rock (1934): "Where is the wisdom we have lost in knowledge? Where is the knowledge we have lost in information?" That progression, from information to knowledge to wisdom, is the journey the Ancient Mariner was forced to undertake. He knew how to sail; he lacked understanding. Modern science possesses unprecedented computational power, sequencing capability, and imaging technology. The question is whether we possess equal judgement in deciding what is worth measuring.

Artificial Intelligence: Accelerator or Crutch?

Let me be clear: I am not arguing against AI, machine learning, or advanced statistics. These tools are transforming biomedical science in genuinely exciting ways. AlphaFold's prediction of protein structures is one example of how AI can guide search over vast combinatorial spaces [13]. Machine learning can generate novel, interpretable hypotheses from high-dimensional datasets [14].

But we must distinguish between information, knowledge, and wisdom. AI can process information. Scientists generate understanding. Wisdom arises through experience, failure, and reflection [15]. A machine learning model trained on terabytes of clinical data can identify correlations that humans might miss [16]. It cannot tell us whether those correlations matter, whether they are causal, or whether they should guide patient care [17].

The danger is that AI encourages us to ask fewer original questions because it promises answers hidden within existing datasets [18]. Hypothesis-free science risks becoming philosophy-free science. We need to remember that AI is a tool for hypothesis generation, not a substitute for scientific judgement [19].

The Poverty of Attention

The mariner's redemption begins not through punishment but through humility. He does not consciously decide to repent; instead, he spontaneously admires the beauty of the sea snakes: "A spring of love gushed from my heart". Only when he learns to appreciate life for its own sake does the albatross fall from his neck. Redemption is achieved through a transformation of perspective.

We need a similar transformation. We need to rediscover the art of conducting science, to recognise that an experiment is about testing a thesis, not generating data. We must accept that negative results often teach us more than positive ones [20]. Failure, historically, sharpened theories. Today, with enough variables and enough statistical techniques, almost every dataset can be persuaded to produce an interesting signal [21]. We have stopped learning how to be wrong.

Karl Popper argued that science advances by exposing theories to the risk of falsification [22]. Data dredging reduces the opportunity for genuine falsification because it continually reformulates questions until something appears to be statistically significant [23]. We should not be squashing the data until they scream some hidden secret.

The Environmental Cost of Data

There is another dimension to this problem that we rarely discuss: the environmental cost of storing enormous biomedical datasets. Data storage systems generate carbon emissions through continuous energy consumption for operations and cooling, as well as through the full lifecycle of hardware [24].

One study estimated that storing staging CT scans for new endometrial cancer patients in the UK between 2020 and 2040 would generate 381 metric tons of CO2 equivalent—and that is for a single cancer type in a single country [25]. Data centres now represent 1% to 1.5% of total electricity use globally, and consumption is projected to increase dramatically with the rise of AI and big data analytics [26]. The endless accumulation of biomedical data carries ecological consequences that we can no longer ignore [27].

This returns us to the Ancient Mariner. The poem is, at its core, a warning about violating our relationship with the natural world. The mariner's suffering is ecological as much as psychological. As biologists and clinicians, we are part of that natural world. Endless data accumulation is not ecologically neutral [28].

The Importance of Storytelling

One of the poem's central Romantic ideas is that wisdom cannot simply be taught. The mariner becomes a wandering storyteller because his suffering has given him insight. His burden is to pass that wisdom on to others. His tale functions as a warning.

Storytelling remains fundamental to science. Scientific papers tell stories. Clinical study reports tell stories. Case reports tell stories. Medical writing transforms observations into understanding [29]. Stories preserve wisdom. Data alone do not. We need to ensure that our young scientists have the opportunity to learn through experience, to fail and recover, to sit quietly and think [30]. We are not giving them that time. And in doing so we ask where are the next generation of ‘thinking’ scientists coming from?

A Call for Humility

The lesson of The Rime of the Ancient Mariner is not that knowledge is dangerous. It is that knowledge divorced from wisdom courts disaster. Modern science stands on the brink of an age in which we can measure almost everything, simulate almost anything, and analyse almost infinitely. Yet our greatest discoveries will still depend upon the oldest scientific skill of all: asking the right question [31].

My feeling is that we need to restore scientific craftsmanship. We need experiments that test ideas, not simply generate datasets. We need to embrace failure as a teacher. We need to recognise that information grows exponentially, but wisdom grows slowly [32]. I am an old fuddy-duddy but for me the greatest discoveries often emerge not from collecting more information but from asking better questions.

Perhaps the greatest danger facing modern science is not that we have too little data—but that, surrounded by oceans of information, we may forget how to think [33].

References

  1. Agrawal A, McHale J, Oettl A. Artificial intelligence and scientific discovery: a model of prioritized search. Research Policy. 2024;53(4):104951.
  2. Mills JL. Data torturing. N Engl J Med. 1993 Oct 14;329(16):1196-9.
  3. Markey N, et al. Clinical trials are becoming more complex: a machine learning analysis of data from over 16,000 trials. Sci Rep. 2024 Feb 12;14(1):3514
  4. Getz K, et al. Insights Informing Strategies for Optimizing the Collection of Clinical Trial Data. Ther Innov Regul Sci. 2026 Mar;60(2):563-574.
  5. Moore TJ, Zhang H, Anderson G, Alexander GC. Estimated costs of pivotal trials for novel therapeutic agents approved by the US Food and Drug Administration, 2015-2016. JAMA Internal Medicine. 2018;178(11):1451-1457.
  6. Firestein S. Ignorance: How It Drives Science. Oxford University Press; 2012.
  7. Platt JR. Strong Inference. Science. 1964;146(3642):347-353.
  8. Head ML, Holman L, Lanfear R, Kahn AT, Jennions MD. The Extent and Consequences of P-Hacking in Science. PLoS Biology. 2015;13(3):e1002106.
  9. Ioannidis JPA. Why Most Published Research Findings Are False. PLoS Medicine. 2005;2(8):e124.
  10. Munafò MR, Nosek BA, Bishop DVM, et al. A manifesto for reproducible science. Nature Human Behaviour. 2017;1:0021.
  11. Smith GD, Ebrahim S. Data dredging, bias, or confounding. BMJ. 2002;325(7378):1437-1438.
  12. Simon HA. Designing Organizations for an Information-Rich World. In: Greenberger M, ed. Computers, Communications, and the Public Interest. Baltimore: Johns Hopkins University Press; 1971:37-72.
  13. Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583-589.
  14. Rajpurkar P, Chen E, Banerjee O, Topol EJ. AI in health and medicine. Nature Medicine. 2022;28:31-38.
  15. Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nature Medicine. 2019;25:44-56.
  16. Beam AL, Kohane IS. Big Data and Machine Learning in Health Care. JAMA. 2018;319(13):1317-1318.
  17. Zech JR, Badgeley MA, Liu M, et al. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Medicine. 2018;15(11):e1002683.
  18. Ludwig J, Mullainathan S. Machine Learning as a Tool for Hypothesis Generation. Quarterly Journal of Economics. 2024;139(2):751-827.
  19. Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health. 2021;3(11):e745-e750.
  20. Nosek BA, Errington TM. Making sense of replications. eLife. 2017;6:e23383.
  21. Simonsohn U, Nelson LD, Simmons JP. P-curve: a key to the file-drawer. Journal of Experimental Psychology: General. 2014;143(2):534-547.
  22. Popper K. The Logic of Scientific Discovery. London: Routledge; 1959.
  23. Mayo DG. Revisiting the Replication Crisis and the Untrustworthiness of Empirical Evidence. Stats. 2025;8(2). doi:10.3390/stats8020031
  24. Kotila, M, Pärssinen, M. Environmental impact assessment of online advertising.
  25. Jia Y, et al. Greenhouse gas emissions due to long-term data storage of CT with reformats and strategies for mitigation. Eur Radiol. 2026 Mar;36(3):2186-2197.
  26. Masanet E, Shehabi A, Lei N, Smith S, Koomey J. Recalibrating global data center energy-use estimates. Science. 2020;367(6481):984-986.
  27. Patil S, Tiew S, Carden SM. The Hidden Carbon Cost of Digital Ophthalmology: A Call for Sustainable Data Practices. Clinical & Experimental Ophthalmology. 2025.
  28. McCullough M. On Attention to Surroundings. Interactions. 2012;XIX(6).
  29. Greenhalgh T. How to Read a Paper: The Basics of Evidence-Based Medicine. 6th ed. Wiley-Blackwell; 2019.
  30. Feynman RP. "Cargo Cult Science" (1974 Caltech Commencement Address). Engineering and Science. 1974;37(7):10-13.
  31. Szucs D, Ioannidis JPA. A Tutorial on Hunting Statistical Significance by Chasing N. Frontiers in Psychology. 2016;7:1444.
  32. Sarewitz D. The pressure to publish pushes down quality. Nature. 2016;533:147.
  33. International Council for Harmonisation. ICH E6(R3) Guideline for Good Clinical Practice. 2025.

About the author

Tim Hardman
Managing Director
LinkedIn logo - blue square with white 'in' textView profile
Dr Tim Hardman is the Founder and Managing Director of Niche Science & Technology Ltd., the UK-based CRO he established in 1998 to deliver tailored, science-driven support to pharmaceutical and biotech companies. With 25+ years’ experience in clinical research, he has grown Niche from a specialist consultancy into a trusted early-phase development partner, helping both start-ups and established firms navigate complex clinical programmes with agility and confidence.

Tim is a prominent leader in the early development community. He serves as Chairman of the Association of Human Pharmacology in the Pharmaceutical Industry (AHPPI), championing best practice and strong industry–regulator dialogue in early-phase research. He ia also a Board member and ex-President of the European Federation for Exploratory Medicines Development (EUFEMED) from 2021 to 2023, promoting collaboration and harmonisation across Europe.

A scientist and entrepreneur at heart, Tim is an active commentator on regulatory innovation, AI in clinical research, and strategic outsourcing. He contributes to the Pharmaceutical Contract Management Group (PCMG) committee and holds an honorary fellowship at St George’s Medical School.

Throughout his career, Tim has combined scientific rigour with entrepreneurial drive—accelerating the journey from discovery to patient benefit.

Social Shares

Subscribe for updates

* indicates required

Get our latest news and publications

Sign up to our news letter

© 2025 Niche.org.uk     All rights reserved

HomePrivacy policy Corporate Social Responsibility