Contents
Operational data is useful when the people and systems relying on it can tell what it means, how current it is and what to do when it is wrong. Start with the business action the data supports. Then design the collection, transformation and delivery work around that action's tolerance for delay, uncertainty and failure.
Start with the action the data supports
Identify who consumes the data and what they do with it. A daily planning review and an immediate operational alert have different needs. Specify the required freshness, acceptable delay and consequence of an incomplete result before choosing a processing schedule or platform.
Define what the consumer should see when data is unavailable or stale. A last-known value may be useful if its age is visible, but misleading if presented as current. Agree which conditions should stop an automated action and which can continue with a clearly stated limitation.
Make the data contract explicit
Record the source, owner and meaning of important fields. Include identifiers, units, time zones and the treatment of missing or corrected values. Clarify whether a timestamp describes when an event occurred, when it was recorded or when it reached the receiving system.
Describe the transformations in business terms as well as code. If records are grouped, filtered or combined, explain which cases are included and excluded. Have a knowledgeable user review representative examples so a technically valid output does not conceal a different interpretation of the process.
Check quality at meaningful boundaries
Define checks where data enters the flow and before it is used. Consider missing identifiers, impossible values, duplicate records and unexpected changes in volume. Choose checks that reflect the action being supported rather than collecting a large set of measurements without a response plan.
Decide how rejected or suspicious records are handled. Keep enough context for an authorised owner to investigate, while limiting sensitive information in diagnostics. Distinguish a processing fault from a source-data issue so the right team receives the exception and can correct it.
Keep results explainable and reproducible
Retain appropriate references to the source records and transformation version behind an output. Agree retention and access requirements with the responsible organisation. The aim is to answer why a result changed without creating an uncontrolled copy of every sensitive field.
Plan how corrections and late-arriving records affect previous results. Decide whether consumers receive a revised value, an explicit correction or an updated dataset. Record these rules so historical comparisons do not silently mix different definitions or incomplete periods.
Test recovery before relying on the flow
Exercise interrupted processing, repeated input and unavailable destinations using controlled data. Verify that restarting or replaying work does not duplicate business effects. If a repair requires rebuilding derived records, define the scope and check that unrelated records remain unchanged.
Reconcile the output against representative source records after recovery. A process finishing successfully is only part of the evidence; the resulting data must match the agreed contract. Separate local tests from checks against actual provider interfaces and production configuration.
Assign an operating owner to the data product
Name who responds to stale data, failed checks and changes in the source system. Define which signals require action and how users learn about a known limitation. Keep a short operating guide that explains diagnosis, safe recovery and escalation without depending on the original developer being available.
Begin with a bounded flow and review its evidence with the people who use it. Expand once freshness, meaning and exception handling are understood. Maintain the contract as business rules change so the data remains useful for the operational decision it was built to support.
A data-quality exercise for an operational feed
Take a fictional daily work-order feed with 100 records. Three lack a stable order ID, two repeat an existing ID and five arrive after the planning meeting. Treat these as separate quality conditions, not eight interchangeable “bad rows”. The late records may be valid for historical reporting but unavailable for today’s planning decision. A complete record can still contain an incorrect status.
Define a response for each boundary: quarantine missing IDs for the source owner; reconcile duplicates using the agreed record/version rule; show the planning view’s data cut-off and excluded late arrivals. Keep event time, source update time and ingestion time distinct. Replaying the same file must reproduce the intended state without doubling work orders, and a later correction must remain explainable.
The Government Data Quality Framework distinguishes dimensions such as completeness, uniqueness and timeliness. Use those distinctions to select checks for this decision, then verify representative records with the operational owner. Do not claim accuracy merely because format validation passes. Record the source and transformation version behind a displayed total so a disagreement can lead to a specific investigation and correction.
How this relates to Veda Software’s work
Encounter’s managed trip information is relevant to making operational content maintainable. Its published narrative does not establish these feed statistics, quality tests or warehouse architecture.