Science Is Leaking

My submission to the Astera Institute's essay contest.

Scientific meaning is badly packed for the trip we ask it to make. Like loose tomatoes in the back of a pickup truck, each turn from observation to model to paper to citation leaves more splattered on the road.

I learned this the hard way. Early in my PhD, I set up a network of sensors to measure sap flux, the rate at which water flows through trees, in a coastal forest affected by sea level rise. Sap flux sensors are basically four temperature sensors in a trench coat. They send heat pulses into a tree, time the pulse traveling a fixed distance, and use a thermal physics model to infer water flow. In commercial sensors, that model is hard-coded into closed-source hardware.

Maintaining sap flux sensors on Virginia's Eastern Shore
Wrangling sap flux sensors on Virginia's Eastern Shore.

The model and I did not get along. Several common estimation methods derive from the same thermal model, but I found that they disagreed with each other, sometimes by more than 70%. The field has learned to live with that disagreement. A bustling economy of post hoc corrections has grown up around it.

The disagreement was a clue. The fix turned out to be hiding in a signal the field has been throwing away for nearly 70 years: the drift in the heat-pulse trace. I derived an estimation method that uses that drift to recover missing parameters; standard methods average it out and call it noise. Recovering it requires the raw signal, which the sensors do not keep. The commercial sensors had already decided what information was worth remembering: only what their one equation needed. The experiment I needed to run had been made impossible before I even knew to ask for it.

Sensors are only the first leak. Sap flux estimates travel through many hands and publications into ecosystem water budgets, drought mortality models, and climate projections. At each step, choices about calibration, deployment, processing, and exclusion fall away from the values they shaped. By the time results arrive, the original thermal signal has become a number with no living connection to the assumptions that produced it. My new estimation method lives on a desert island.

This is not unique to sap flux, nor to plant ecophysiology. It is a general property of scientific memory: the substrate we write to when we produce new work and query when we build on what came before. Today's bottleneck is the relationship between scientists and that substrate. Researchers learn more than the record preserves, and the rest cannot be queried.

That gap matters most in fields where observation and definition happen together. Software has tight verification loops and shared infrastructure. Cellular and molecular biology have spent decades building shared handles for genes, proteins, pathways, mutants, assays, and phenotypes. Plant ecophysiology sits in a more open-ended region of science, alongside much of ecology, Earth systems science, and parts of neuroscience. The objects of study do not arrive already discretized. Researchers have to decide what a "tree," "stress," "health," or "sap flux" means for a particular instrument, model, and question before results can stack.

What our infrastructure can preserve about those decisions shapes our epistemology. When the information needed to answer precise questions leaks out of the pipeline, researchers are incentivized to ask questions the system can safely carry. Open-ended sciences have often absorbed that friction by loosening their claims rather than building the infrastructure to tighten them. Science can only accommodate the inquiry its memory can support.

At human speed, prose plus people has been a workable substrate for scientific memory: papers compress claims into narrative, datasets and code travel as attachments, and researchers use field norms to reconstruct what the record leaves implicit. LLMs acting as scientific agents put pressure on that arrangement. They feel like early coding agents: uneven, strange, sometimes brittle, but clearly able to learn tools and move through technical work.

Coding agents accelerated real work quickly in part because software had already externalized much of its memory into git, tests, package managers, APIs, type systems, and public standards. Science has stored more of its memory in people, prose, and conventions, so agents inherit artifacts that require them to reconstruct meaning after the fact: PDFs, prose methods, partial code, balkanized databases, and CSVs whose units live in an undergraduate's field notebook from three summers ago.

Scaling LLMs does not solve that problem alone. Today's substrate cannot supply meaning it never stored, no matter how many agents query it or how quickly they work. As LLMs become producers as well as consumers of science, that risk compounds: new artifacts inherit the lossiness of the records that produced them. Once those artifacts become part of the record, future inquiry is bounded by what the substrate preserved.

To close that gap, the record has to hold meaning in a durable form. I propose a type system for scientific memory: a substrate that preserves the meaning of measurements as they move from observation to model to claim. A type, here, is a contract for reuse. It says what a measurement can mean downstream, what transformations preserve that meaning, and what claims it can support.

Earlier efforts point in this direction. FAIR established principles for visibility: can data be found, accessed, and reused? Unified ontologies, the semantic web, and discourse graphs developed structures for organization: can concepts, entities, and claims be related explicitly? Both made important layers of the record more legible. The opportunity now is to keep meaning continuous across layers.

That continuity has to be vertical. Even if my sensor preserves the raw heat-pulse trace, meaning can still leak later if calibration choices, exclusions, and model assumptions disappear before the result reaches an ecosystem model. Every link has to stay typed, from raw signal to consuming claim, so standards at one layer are not undone at another.

The substrate also has to be plural and versioned. In open-ended science, definition is negotiated alongside observation, so a single field-level ontology is too rigid. The same observation should be able to carry competing interpretations, attributed and preserved rather than averaged into consensus. That disagreement is part of the science, not noise to be resolved.

Finally, the substrate has to accept natural language as a first-class input. Scientists were never going to write in a typed language on their own; the cost was higher than the value. LLMs change the bargain: they can help encode meaning as science happens, moving more of what researchers learn into the substrate without asking scientists to become programmers. That is where LLMs become more than accelerators of existing practice: the first practical mechanism for closing the gap between what scientists learn and what the record preserves. With more measurement and meaning available for recall, open fields can afford bolder epistemologies: sharper questions, tighter claims, and assumptions that remain attached to the evidence carrying them.

That claim is testable. Memory is tested by recall; recall is best evaluated by predictive power. The right experiment is an ambitious project possible within our collective understanding but poorly served by current infrastructure, using that gap as a high-signal test of whether better memory changes what science can explain.

I want to use my field's most challenging problem as a model system. Forest-scale responses to climate change remain poorly explained by coarse-grained data. Physiology-level studies show trees are stateful, responsive actors. Current models bridge that gap with fixed heuristics for allocation, acclimation, and mortality, which work locally but struggle across time, state, and scale. We can sidestep fragile heuristics by learning behavior directly from data inside trusted physical constraints, then testing whether tree-level predictions scale into forest outcomes.

But what data? Ecology produces sparse, expensive observations spread across tens of thousands of lossy publications. The bottleneck is composition: a behavioral model only becomes identifiable if evidence composes across datasets, instruments, sites, and scales. Community-scale signals, like canopy change captured by drones, are cheap enough to collect broadly and demanding enough to distinguish real explanatory power from brittle curve-fitting. That makes them high-signal tests for whether a memory substrate can connect measurements, mechanisms, and forest outcomes.

Aerial view of a coastal forest study site
A coastal forest dying back in response to rising sea levels.

I would run a controlled experiment with one prediction target and three memory regimes. First, use LLMs to extract as much existing literature data as possible and train on the raw extraction. Second, formalize the same extraction into the typed substrate before modeling. Third, generate new data end to end inside the substrate, preserving raw signal, calibration, assumptions, transforms, and provenance from observation forward. Each arm trains toward the same prediction target and is judged by the same forward predictions: how well can the tree model explain what happens to real forests threatened by climate change.

Memory regime Question it answers
Raw literature extraction Can LLMs recover enough meaning from human-mediated scientific memory to make hard predictions?
Typed literature extraction Does adding typed structure to the existing record meaningfully improve what can be predicted from it?
End-to-end typed memory Does structured memory from observation onward enable a faster, more predictive science?

If raw extraction produces meaningful predictive power, we should immediately aim LLMs at hard problems scattered through the literature. If typed extraction wins, we should autoformalize more of the existing record. If end-to-end memory wins, we should generalize the standards across open-ended fields. If none work, the bottleneck is upstream, and we should pivot toward making observation cheaper.

The pieces of this substrate are already coming together. Sap flux is my test case for the point of observation: to learn how raw readings, instrument assumptions, and type-level meaning can travel together, I designed an open logic board, assembled at JLCPCB, that cuts instrumentation cost from $1,000 to $20, drops idle power by 99%, and gives me ownership of how the data move.

Prototype sap flux logic board
v0.1 of my new sap flux logic board!

I have a working prototype for storing and retrieving operations on scientific data, and a draft language specification for attaching semantic meaning to observations and constructs. I initially built them separately, but now realize they're two halves of one system: the data layer needs types, and the type system needs provenance and reuse.

This work is better suited to a focused research organization than to a startup or a single academic lab. The right structure is a small engineering team building the substrate, paired with grants to field researchers who use it on real science and own the resulting findings. Hardware can spin out, and scientific results should belong to the people and sites that produce them. The organization's durable output is the public standard: closer to git than GitHub. Its job is to keep the loop tight: build the tool, run the field campaign, find the leak, change the standard, repeat.

I set out to understand what will become of forests on a changing planet, and found that it hinges on questions our infrastructure is not well-suited to answer. The bet is that this is fixable: build the memory while doing the science, measure it by prediction, and share the standards that survive contact with real fieldwork.

Comments

Loading comments...