Skip to main content

Command Palette

Search for a command to run...

Why Wastewater Compliance Is a Data Engineering Problem

How fragmented lab data and local rules shaped BrewClear’s architecture

Updated
•6 min read•View as Markdown

The first article in Building BrewClear: Data Engineering for Local Water Compliance

Imagine an environmental consultant opening a lab report for a small brewery. The report has a BOD result and a sample date. Somewhere else is the facility's permit. A sewer-use ordinance contains the receiving treatment authority's rules. The reporting deadline is on a calendar, and last month's corrective-action notes are in an email thread.

The number in the lab report matters. But it cannot answer the question the consultant actually has: What does this result mean for this facility, under the rules that apply to it, and what needs to happen next?

I am a data science engineer and a cofounder of BrewClear. We started building it because this question sits at the intersection of messy source data, local regulatory context, repeatable calculations, and human judgment. That makes it an engineering problem before it becomes a dashboard problem.

The same measurement needs different context

In the United States, an industrial facility that discharges to a municipal sewer may be subject to pretreatment requirements administered through a publicly owned treatment works, or POTW. The U.S. Environmental Protection Agency explains that the national program aims to protect treatment infrastructure and reduce industrial pollutants entering municipal systems. Local limits address the needs of a particular POTW, its workers, sludge, and receiving waters. They can be numeric or narrative.

So a BOD result is not a universal green or red number. Interpreting it requires the correct facility, discharge path, treatment authority, applicable limit, measurement unit, sample date, and sometimes a different threshold for a surcharge than for a violation. A result without that context is only a value in a file.

The engineering consequence is clear: we cannot hard-code one set of thresholds and call it a compliance platform. The rule itself must be data, connected to the facility and accompanied by a source and a verification status. We also need to know when a result was sampled and which version of a rule was used to assess it. BrewClear's proof of concept handles some of this context today; full historical rule versioning is still work ahead.

The inputs do not arrive as one clean dataset

Consultants receive data in forms created for different purposes. A lab may deliver one row per analyte, with sample identifiers, methods, units, qualifiers, and reporting limits. A simple spreadsheet may have one row per sample. A permit may be a PDF. A deadline may be in a calendar, while evidence of a corrective action lives in an email or document.

Those files cannot safely be merged by matching a date and hoping for the best. The pipeline needs to answer questions such as:

  • Which facility and sampling point does this record belong to?
  • Are two rows different analytes from one sample, or two separate samples?
  • Does the lab's name for an analyte match the parameter named in the local rule?
  • Are the units compatible, and what does a non-detect qualifier mean?
  • Which values were imported, which were extracted from a document, and which were entered by a person?

Our first approach in BrewClear is to normalize several input paths into a common reading record. The current proof of concept accepts manual entries, a simple CSV layout, a canonical lab electronic data deliverable, and document-assisted extraction. The lab parser groups analyte rows into samples and returns warnings for conditions such as unit mismatches or untracked analytes. Extracted document values prefill a form for review; they are not silently treated as confirmed measurements.

That distinction matters. A pipeline is useful only if it preserves the meaning and uncertainty of its inputs. A plausible-looking number with a lost unit or qualifier can be more dangerous than a missing value.

A decision should be traceable through the system

Once a reading is normalized, it still needs to be evaluated against the right local rule. In BrewClear, the core check is a deterministic function: it takes a reading and a locality record, then returns parameter-level results and an overall classification. A separate calculation produces a rough surcharge estimate where rates are configured. The same checking logic feeds reports, client packets, and the data tools used by the copilot.

That shared calculation is an architectural choice. If a dashboard, a report, and an AI assistant each implement their own threshold logic, they can disagree about the same sample. Reusing one calculation makes a discrepancy easier to detect and explain. It also gives us something concrete to test with known inputs and expected outputs.

Here is the flow we are building around:

Lab file / entry / document
          |
          v
Normalized reading + source + warnings
          |
          +-------- Applicable locality rule + provenance
          |                    |
          v                    v
             Deterministic check
                     |
          Classification + explanation
                     |
          Consultant review and action
                     |
             Draft report / packet

The diagram also shows a limit of the software. A classification is not the end of the workflow. Someone still needs to assess the source documents, investigate unexpected results, record corrective action when appropriate, and review what will be submitted. BrewClear does not file reports with an agency.

Why build BrewClear around this workflow?

A spreadsheet can calculate against a known threshold. It is much harder for a collection of spreadsheets, PDFs, and inboxes to maintain a reliable connection between a facility, its changing local rules, the original measurement, the calculation, and the action taken. That connection is the product we want to build.

Our current proof of concept demonstrates the workflow across Michigan, Washington, Oregon, and California with seeded facilities. It uses state and industry rule packs, locality records, reading ingestion, deterministic checks, reports, and a consultant workspace. This demonstrates the architecture; it does not mean every locality value is ready for unsupervised production use. Regulatory content must be checked against authoritative sources, and the person responsible for a report must review it.

The deeper engineering work is still ahead: stronger validation across varied lab formats, rule effective dates and versioning, end-to-end lineage, better storage of source documents, and production data-quality monitoring. Those are not peripheral improvements. They determine whether a future result can be explained and trusted.

That is why I describe wastewater compliance as a data engineering problem. The goal is not to automate judgment away. It is to give the people making decisions a consistent, traceable picture of the evidence and the rules.

Next in the series: how BrewClear represents state guidance and locality-specific limits as rule-pack data, so the same checking engine can work across different jurisdictions.