The Data Didn’t Forget
A look back to the beginnings of environmental data and why environmental data is different
Jon Turner, Senior Project Manager, ddms, Inc.
This was my first introduction to the site: concrete, buildings, tanks, CAD files, paper, caution tape, and guessing!
It was the late 1990s, and the site was mostly active while a team of data scientists tried to understand what was happening underground. The geology was a mess. There were thin layers, lenses that pinched out, and groundwater moving just fast enough to be dangerous but not fast enough to fix anything quickly. Utilities crisscrossed the subsurface like a bad map drawn from memory (and some of the early map might have been). And the “plume” didn’t behave the way the conceptual model (and regulator) said it should.
Managing data before it was called data management
Back then, my job was primarily fieldwork, but I had a GIS background, so I knew about relational databases. We measured water levels by hand, logged flow rates from a treatment system that barely moved a few gallons per minute, and collected samples knowing full well that each one mattered because there simply wasn’t much water to capture. Pressure transducers told the story. Capture zones were small. The tolerance for error was smaller.
And nothing could happen without knowing exactly where things were.
GPS was still new enough to be frustrating. Bulky equipment; long wait times; satellites that didn’t always cooperate. A lot of folks struggled with mis-recorded points, mislabeled files, or gave up and sketched locations instead. I didn’t. I doublechecked every coordinate, waited for good fixes, and took notes in case something didn’t line up later.
Those coordinates (taken standing in the rain or next to a busy highway) are still used today. If you know what correction factors and WAAS are, you know it was done right—but you also know to take some early coordinates with a grain of salt.
At first, we put coordinates into spreadsheets. Then there was a small database, because the spreadsheets started to crack. We were importing six-hour transducer data from a dozen locations! I helped decide how wells would be named, how sample IDs would tie back to locations (why would we possibly need a unique ID?); how lab results would be lined up so trends could be seen; how to efficiently import the measurements and results; how to calculate the monthly totals or averages.
Nobody called it data management. It was just how you used technology to be more efficient and provide better service.
Reporting came next. Tables for quarterly reports. Figures showing drawdown. Graphs to see whether concentrations were changing or drifting. Did someone sample the wrong well? I made the first versions because someone had to, and because I understood the site, the GIS, and the database. Over time, those tables became “standard.” The figures were reused. The logic stuck.
As technology improved, what happened to the data?
Years passed. Regulations changed. Analytical methods improved. Detection limits dropped. Wells were added, modified, and sometimes quietly abandoned. The treatment system was tweaked repeatedly.
The data accumulated.
And it never went away.
Now, more than 25 years later, the site is still active, yet in a different phase. Pumping wells are being evaluated for removal. Insitu treatment is on the table. People want to know which data still matters, which trends are real, and which assumptions are outdated.
Today, I’m responsible for data management at the same site, in a different role.
Yes, some of the data “problems” we have now trace back to decisions I made decades ago. Naming conventions that made sense at the time. Tables designed for reports, not future analytics. Assumptions that were never written down because, back then, everyone just knew (yet at the time I felt we documented everything). Did I do those calculations correctly? Why didn’t I make those location IDs consistent like the others? But hindsight is 20:20, right?
And yet, so much of what still works today came from those early choices. Well location information (yes, we need screen intervals); database relationships; location and parameter groups. The definitions of task codes. The coordinates captured with care when it would have been easier not to. The summary tables. The logic.
Oh, the memories…
Environmental data has memory, whether you plan for it or not.
Most modern data systems (and maybe even human brains) are built for forgetting; it’s here today, gone tomorrow. Many data management strategies assume old rules can be overwritten, old context thrown away, old decisions quickly normalized into something cleaner and simpler without much care or context needed.
Long-term environmental data management doesn’t work like this. Every value exists in the shadow of what was known at the time. Every decision depends on understanding not just what the data says, but why it looked that way when it was first collected or imported.
Automation continued to expand at the site, and it helped. Rules now catch normalization mistakes earlier. Consistency has improved over time. Reporting has become faster. But automation has never answered the hardest questions, and some automation was never implemented. Someone still must “do things,” must decide whether the data makes sense, whether there was a mistake made in the field or lab, whether an anomaly is acceptable, whether a result supports the decision being made.
What changed over time wasn’t the need for judgment; it was where judgment lived.
In the early years, it lived in people’s heads, in emails, in margin notes on printed reports. Now, the lesson is clear: judgment has to be part of the system. Rules and automation should do what they do best—enforce consistency, reduce noise, flag problems—but expertise has to be captured, not hidden.
What we now realize (what would have helped to know earlier) isn’t that our issues weren’t going to be solved by a perfect standard or a single “right” platform. We needed architecture that accepted reality and that’s still the case: messy inputs, strict internal rules, and outputs tailored to project needs. Flexibility at the edges, disciplined at the core.
What we need is a system that doesn’t pretend that labs, field crews, sensors, and consultants will ever behave the same way. We need a governed core where data is preserved with context, many decisions are traceable, and history is respected, so that when a site changes direction, the data doesn’t fight back.
Some sites use data management software like Project Portal and/or EQuIS. Some don’t. Some sites need heavy integration with operations data. Others are driven almost entirely by regulatory analytical reporting. The right solution depends on the goals of our users.
What never changes is the fact that both environmental impacts and data stay with us, seemingly forever.
I know that better than anyone. I stood in the field when the first GPS points were collected; I’ve sampled wells on the shores of Lake Superior in June and at -8° in January; I’ve drilled some of the “geoprobes” that never got marked as “abandoned” in the database; I’ve matched up CAS numbers to parameter names countless times; and have seen “MW-1” spelled six different ways.
I was there when the systems were new, when the shortcuts still felt harmless. Now I feel responsible for making sure the next generation inherits data that can still be trusted yet make sure everyone has an appreciation for what these data have been through. And I am always downright excited to tell stories of the “olden days” and push our team at ddms to be more thoughtful, innovative, and better than me!
