Data Lineage Tracking for AI: Complete Guide | QuizBy Eyal Doron / December 6, 2025 / 1 minute of reading Data Lineage Tracking for AI: Complete Guide | Quiz 1 / 8 1. What does the EU AI Act require regarding training data according to the article? 1. Documentation is optional for all risk levels 2. Training data documentation for high-risk systems and demonstrable traceability requirements 3. Only the model output needs to be documented 4. No documentation is required for any AI systems Correct! Why: The EU AI Act requires training data documentation for high-risk AI systems demonstrating what data trained the model and its characteristics plus traceability requirements. Context: Lineage is the technical foundation for meeting these regulatory requirements. Remember: Document training data plus demonstrate traceability. 2 / 8 2. What Quick Win does the article recommend for starting lineage tracking? 1. Purchase an enterprise lineage platform immediately 2. Document all lineage manually in spreadsheets 3. Hire a dedicated lineage team 4. Mandate DVC or MLflow to version control training dataset and model file for highest-risk AI system Correct! Why: The article recommends mandating DVC or MLflow to version control the training dataset and model file for your highest-risk AI system this week. Context: This creates the essential model-to-data linkage required for basic compliance. Remember: Version control training data and model for highest-risk system. 3 / 8 3. What is the critical link for backward lineage according to the article? 1. Model-to-data linkage connecting each trained model to its training dataset versions 2. API authentication tokens 3. Network connection between servers 4. Database foreign keys Correct! Why: Model-to-data linkage explicitly connects each trained model to its training dataset versions – without it you cannot trace a prediction back to its training data. Context: Dataset version identification assigns unique identifiers to training data snapshots. Remember: No model-to-data link equals no backward traceability. 4 / 8 4. Why is transformation code versioning essential according to the article? 1. It makes the code run faster 2. It reduces storage costs 3. It is only needed for compliance audits 4. Capturing Git hash lets you know exactly which code version processed the data Correct! Why: Capturing the Git hash of the cleaning script lets you know exactly which code version processed the data enabling reproducibility. Context: This is part of documenting every transformation applied to raw data during preparation. Remember: Git hash equals reproducible transformations. 5 / 8 5. Why does feature engineering obscure data origins according to the article? 1. Derived features like ratios and aggregations create indirect connections to dozens of underlying data points 2. Engineering transforms data into unreadable formats 3. Feature engineering deletes the original data 4. Features are stored in different databases than source data Correct! Why: When you derive new features like ratios and aggregations and embeddings the connection to original data becomes indirect – a customer_risk_score might derive from dozens of underlying data points. Context: This is one of several factors that make AI lineage harder than traditional data lineage. Remember: Derived features hide their sources. 6 / 8 6. What is the difference between forward and backward lineage? 1. Forward is for new data while backward is for historical data 2. Forward is for training while backward is for inference only 3. Forward is automatic while backward requires manual effort 4. Forward traces source to output while backward traces output to source Correct! Why: Forward lineage traces data from source to output answering what happened to this data while backward lineage traces from output to source answering where did this prediction come from. Context: Both directions matter – forward supports compliance and auditing while backward enables debugging and explanation. Remember: Forward equals source to output – Backward equals output to source. 7 / 8 7. What three critical questions does lineage answer according to the article? 1. What data – what transformations – what model version 2. Where stored – when backed up – who owns it 3. Who accessed – when accessed – why accessed 4. How much – how fast – how accurate Correct! Why: The article states lineage answers what data and what transformations and what model version – if you cannot answer all three you have a lineage gap. Context: These questions form the foundation of traceability from prediction back to source. Remember: What data – What transformations – What model version. 8 / 8 8. According to the article – what analogy best describes data lineage for AI? 1. A firewall that protects data from unauthorized access 2. An encryption system that secures data at rest 3. A family tree for your data showing origin and transformations and destination 4. A backup system that stores copies of all data Correct! Why: The article describes data lineage as a family tree for your data showing where data came from and what happened to it along the way and where it ended up. Context: This is also compared to chain-of-custody for your AI pipeline documenting every transformation. Remember: Family tree plus chain-of-custody for data. Your score isThe average score is 0% Restart quiz Download PDF Please leave this field empty๐ The AI Security Manager's Newsletter Weekly insights on AI risk management, EU AI Act compliance, and practical security strategies. We donโt spam! Read our privacy policy for more info. Thank you! Please check your inbox to confirm your subscription.