Recently Published
1%
an R2 of 1% can be counter intuitive
Validity Issues RCA
The analysis of validity and outliers revealed that the data engineers failed to address invalid entries in the "surname" field of the dataset obtained from Kaggle. This oversight resulted in an increased occurrence of outliers, as the Bayesian and frequency-based methods employed were adversely affected by the dense concentration of erroneous surname data. Consequently, the dataset will be returned to the data owner for correction and remediation of the invalid surname entries. Upon successful rectification, the dataset can be reprocessed to obtain more reliable and accurate results, free from the distortions caused by the outliers stemming from the invalid surname data. The data quality issue and its identified root cause highlight the importance of rigorous data validation and cleansing processes to ensure the integrity and usability of datasets, particularly those sourced from external platforms. momentarily this will end my investigation, for further questions please email me.