RPubs will retire in June 2027. Your existing documents will stay accessible through December 31, 2031
and Connect Cloud is the recommended home for new publishing. Read the blog post

Recently Published

Clustering techniques
Concrete Jungle: Unraveling the Complexities of Apartment Pricing in Baku’s Urban Landscape
Through this analysis, we seek to uncover the roles of location, amenities, economic conditions, and other relevant variables that contribute to the fluctuation of property values in Baku. By leveraging data sourced from Kaggle and scraping information from the local real estate platform Bina.az, the project will employ advanced statistical methods to generate actionable insights. The outcome of this research will not only enhance our understanding of the Baku housing market but also provide critical policy recommendations for more informed urban planning and development strategies. Ultimately, this project aims to offer a comprehensive model that can guide future real estate investments and improve living conditions for residents in the city.
Taylor Swift’s Discography Analysis (Clustering)
This project explores Taylor Swift’s discography using clustering techniques to identify patterns in her music. By analyzing features like danceability, energy, and valence, the songs naturally split into two groups: high-energy pop anthems vs. emotional, introspective tracks. The analysis includes statistical validation, visualization inspired by Taylor’s album aesthetics, and some fun insights into the evolution of her sound.
Sharpe Ratio
The Sharpe Ratio is a widely used risk-adjusted performance metric in finance. It measures the excess return of an investment (or portfolio) compared to the risk-free rate per unit of volatility or standard deviation.
Association rules mining
The project explores the correlation of infant mortality rate, birth rate, healthcare expenditures, and gas emissions.
HTML
Joola Villages from RefLex Database; Groups from Sambou (2014: 30)
Correlation Analysis with Julia
Correlation Analysis with Julia
Plot
Ass31
Testing and Training Datasets
Several Julia packages offer functionality for splitting datasets into training and testing sets.
Dimension reduction and clustering of high-variance gene expression data
This project explores various dimensiona reduction and clustering techniques applied to high-variance gene expression data. The primary objective is to identify meaningful patterns in gene expression by reducing dimensions while preserving key structures. The analysis begins with data preprocessing, scaling, and feature selection, followed by Principal Component Analysis (PCA) and Factor Analysis (FA) to determine the optimal number of components or factors. Clustering methods, including K-Means, Hierarchical Clustering, and t-SNE with K-Means, are then applied to uncover potential subgroups within the data.