Data lineage records where each dataset, column or value came from and which jobs transformed it, so any output can be traced back to its source rows and forward to everything derived from it. Orchestration tools can capture it at the job and dataset level, and SQL parsers at the column level; OpenLineage is an open specification for emitting it as run, job and dataset events.
In an FDE interview
It is how you answer the question a customer asks the first time your dashboard disagrees with their spreadsheet: which rows produced this number? A strong candidate designs for it from the start: each output row carries its source system, source primary key, extract time and transform version, and each aggregate records the snapshot and query version that produced it, so the number can be recomputed. In a retrieval system, each answer cites the document, version and chunk it used.
Lineage also runs forward, which is what makes deletion possible: when a person asks for their data to be erased, or a privacy officer asks exactly where it goes, lineage lists every derived copy of that personal data, including caches and embeddings, that a deletion must reach.
The lesson Data, PII and compliance covers lineage from source to answer.