Data, data, data: Everything is data
A larger dataset can feel like a clearer picture of the world. But what if all those extra observations come from the same places, seasons, or conditions? Is more always better? Does it mean we understand an ecosystem more fully? Where does quality come into play? And how do we know when we have enough data?
I’ve been thinking about these questions because, whether you’re a scientist or not, you’re a consumer of data.
Data inform the decisions we make (even if those data are articles we choose to read online, like this one!) and while some decisions may be small, others are large and lasting, such as policy support for new environmental regulations. While most of us don’t get to decide how a dataset is collected, we all can learn to be more informed about data sources and curation, and what questions those datasets allows us to ask. So, let’s dive in (SCUBA reference, maybe..).
How an ecosystem becomes data
Imagine a researcher surveying a coral reef for the first time. They might count the pufferfish they see or record the fishing boats that pass nearby. If neither appears in their notes, does that mean neither was there? Maybe. Or maybe they just weren’t looking at the right time.
Something missing from a dataset may tell us as much about where we directed our attention as it does about what was happening in nature. This is one way data can go “dark”: observations we don’t have, whose absence may matter more than we realize.
Now imagine that same researcher recording a fish feeding. Before that behavior becomes a data point, someone has to choose the site, the time of day, the species, the recording method, and what counts as “feeding.” Those choices aren’t made in a vacuum. We tend to study what we can reach, what interests us, what we can get funding for, and what our tools can measure.
The choices continue after the survey ends. Did someone remove missing values before sharing the dataset? If so, why were those values missing? Say an underwater camera can’t capture usable footage when the water turns cloudy after a storm. If we remove every cloudy recording, we may end up with plenty of footage of fish in clear water and very little idea of what happens after storms. Perhaps behavior differs in those conditions or a new species passes through from other geographical areas. The gap in the data itself is worth noticing and exploring.
That doesn’t mean we need perfect data before we can make a decision. It means we should be honest about what we know. Where and when were observations collected? Are the gaps random, or do they appear under particular conditions? If those conditions matter to the decision at hand, perhaps that tells us where to survey next.
You are the scientist of your own mind
Researchers make these kinds of choices every day. We might have a grant to study a particular question with a particular method, and we have to work within it. Sometimes our hands are tied. But those ordinary decisions shape our datasets, the conclusions we draw from them, and eventually how we see the world.
I’m not asking you to distrust science. I’m asking you to stay curious about how a claim was made. When you read about a new finding, consider what the researchers measured and what they might have missed. If something isn’t clear, look for another study or ask the researchers themselves. I’m a huge proponent of reaching out. It’s gratifying to hear that someone outside academia has read your work, and a good question can show us researchers where we haven’t explained ourselves well enough.
Many of us want our research to reach beyond other researchers. You can be part of that conversation. Be curious, ask questions, and be the scientist of your own mind.