RPubs will retire in June 2027. Your existing documents will stay accessible through December 31, 2031
and Connect Cloud is the recommended home for new publishing. Read the blog post

Recently Published

Validity Issues RCA
The analysis of validity and outliers revealed that the data engineers failed to address invalid entries in the "surname" field of the dataset obtained from Kaggle. This oversight resulted in an increased occurrence of outliers, as the Bayesian and frequency-based methods employed were adversely affected by the dense concentration of erroneous surname data. Consequently, the dataset will be returned to the data owner for correction and remediation of the invalid surname entries. Upon successful rectification, the dataset can be reprocessed to obtain more reliable and accurate results, free from the distortions caused by the outliers stemming from the invalid surname data. The data quality issue and its identified root cause highlight the importance of rigorous data validation and cleansing processes to ensure the integrity and usability of datasets, particularly those sourced from external platforms. momentarily this will end my investigation, for further questions please email me.
CONSORT 2010声明 フローチャート template
consort2010_flow.qmd 課題1:ファイルを読込ましてプロット 課題2:ラベルを90度回転させたい
The Deeper Discrepancies
The data quality assessment revealed the presence of outliers in the system, indicating potential validity issues that need to be addressed. Upon closer examination, it became evident that these outliers were not merely isolated occurrences but rather indicative of underlying data quality concerns. A thorough investigation into the root causes of these outliers was deemed necessary to ensure the integrity and reliability of the dataset. Outliers can arise from various sources, including data entry errors, system glitches, or even genuine but extreme observations. Distinguishing between valid and invalid outliers is crucial to avoid misguided decision-making or flawed analyses. The data stewards and owners must collaborate to establish clear guidelines and criteria for identifying and handling outliers effectively. Once identified, outliers can be addressed through various techniques, such as data cleansing, imputation, or removal, depending on their nature and the specific requirements of the analysis. Robust statistical methods may also be employed to mitigate the impact of outliers on analytical models and ensure accurate insights are derived from the data. Addressing validity issues and outliers is essential to maintain the overall quality and trustworthiness of the dataset, enabling stakeholders to make informed decisions based on reliable and accurate information, Next Question is; "what are those?
HTML
HTML Tutorial
A tutorial of HTML R Markdown
Correlation Matrix of Performance Variables from the 2020-21 EPL Season
Analysis of the correlation between Expected Goals (xG), Pass Completion (PassComp), Possession (Poss), Goals Scored (GF) and Points (Pts) in the 2020-21 EPL season
Module 12
HTML
Project1
An analysis of world cup winners and individual award correlations
HTML
test