{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TP July 2022 EDA & Baseline Modeling","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we will explore the underlying data and build a baseline model.","metadata":{}},{"cell_type":"code","source":"library(mclust)\nlibrary(data.table)\nlibrary(magrittr)\nlibrary(skimr)\nlibrary(ggplot2)\nlibrary(caret)\nlibrary(tictoc)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## EDA","metadata":{"execution":{"iopub.status.busy":"2022-07-02T17:43:28.288496Z","iopub.execute_input":"2022-07-02T17:43:28.290044Z","iopub.status.idle":"2022-07-02T17:43:28.309985Z"}}},{"cell_type":"code","source":"DAT_DIR <- '../input/tabular-playground-series-jul-2022'\ndata_dt <- fread(file.path(DAT_DIR, \"data.csv\"))\ndim(data_dt)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T16:50:25.568743Z","iopub.execute_input":"2022-07-12T16:50:25.570212Z","iopub.status.idle":"2022-07-12T16:50:25.956836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data set has 30 columns and 98K rows. Now, Let us first review the column types.","metadata":{}},{"cell_type":"code","source":"table(unlist(lapply(data_dt[, -c(\"id\")], typeof)))","metadata":{"execution":{"iopub.status.busy":"2022-07-12T16:51:52.887485Z","iopub.execute_input":"2022-07-12T16:51:52.889021Z","iopub.status.idle":"2022-07-12T16:51:52.909678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(col_types <- unlist(lapply(data_dt[, -c(\"id\")], typeof)))","metadata":{"execution":{"iopub.status.busy":"2022-07-12T16:51:07.810089Z","iopub.execute_input":"2022-07-12T16:51:07.811568Z","iopub.status.idle":"2022-07-12T16:51:08.064078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Excluding the id column, the data set has 22 numeric columns and 7 integer columns. More specifically, variables $f_07$--$f_13$ are integer columns while the rest are numerical columns.","metadata":{}},{"cell_type":"code","source":"cat(\"Categorical variables: \\n\")\n(cat_cols <- names(col_types)[which(col_types == 'integer')])\ncat(\"\\n\")\ncat(\"Numerical variables: \\n\")\n(num_cols <- names(col_types)[which(col_types == 'double')])\ncat(\"\\n\")","metadata":{"execution":{"iopub.status.busy":"2022-07-12T16:53:20.333961Z","iopub.execute_input":"2022-07-12T16:53:20.335450Z","iopub.status.idle":"2022-07-12T16:53:20.367402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us take an overview of the content.","metadata":{}},{"cell_type":"markdown","source":"Now, let us zoom in on the distribution.","metadata":{}},{"cell_type":"code","source":"options(repr.plot.width=24, repr.plot.height=12)\n\ndata_dt[, -c(\"id\")] %>%\n  melt(id.vars=NULL) %>%\n  ggplot(aes(x=value)) +\n  stat_density() +\n  facet_wrap(~variable, scales=\"free\")","metadata":{"execution":{"iopub.status.busy":"2022-07-12T16:59:04.882713Z","iopub.execute_input":"2022-07-12T16:59:04.884221Z","iopub.status.idle":"2022-07-12T16:59:20.475838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We noticed the numerical columns have distribution close to normal while integer columns are not so close to normal. In addition, the numerical columns do contain some negative values.","metadata":{}},{"cell_type":"markdown","source":"The boxplots below provide an alternative view on the distribution.","metadata":{}},{"cell_type":"code","source":"data_dt[, -c(\"id\")] %>%\n  melt(id.vars=NULL) %>%\n  ggplot(aes(x=variable, y=value)) +\n  geom_boxplot() +\n  theme(axis.text.x=element_text(angle=90, vjust=0.5, hjust=1))","metadata":{"execution":{"iopub.status.busy":"2022-07-12T16:59:41.194147Z","iopub.execute_input":"2022-07-12T16:59:41.195745Z","iopub.status.idle":"2022-07-12T16:59:46.553907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see a few variables have significantly large variances than the rest and then let us drill down on it.","metadata":{}},{"cell_type":"code","source":"VAR_THRESH <- 1.5\ndata_dt[, -c(\"id\")][, lapply(.SD, function(x) var(x))] %>%\n  melt(id.vars=NULL) %>%\n  ggplot(aes(y=variable, x=value)) +\n  geom_col() +\n  geom_vline(xintercept=VAR_THRESH, color='red', linetype='dashed')","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:01:30.764411Z","iopub.execute_input":"2022-07-12T17:01:30.765914Z","iopub.status.idle":"2022-07-12T17:01:31.201634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see above, more than half of the variables have variance below the threshold of 1.5.","metadata":{}},{"cell_type":"markdown","source":"## Baseline Modeling","metadata":{}},{"cell_type":"markdown","source":"In this section, we will cover\n\n1. Data preprocessing.\n2. Identify optimal number of clusters.\n3. Clustering model with Gaussian Mixture Modeling (GMM).","metadata":{}},{"cell_type":"markdown","source":"### Data Preprocessing","metadata":{}},{"cell_type":"markdown","source":"For data processing, we will apply the following\n\n1. Drop columns with variance below the threshold.\n2. Scale the columns to remove mean and set standard deviation to 1.\n3. Apply YeoJohnson transformation to bring distribution closer to normal.","metadata":{}},{"cell_type":"code","source":"data_col_dt <-\n  data_dt[,-c(\"id\")][, lapply(.SD, function(x)\n    var(x))]\n\ndata_cols <- as.numeric(data_col_dt)\nnames(data_cols) <- names(data_col_dt)\n\ncat(\"Selected columns whose variance above the threshold\\n\")\n(selected_cols <- names(data_cols)[which(data_cols > VAR_THRESH)])\n\ndata_dt2 <-\n  data_dt[, .SD, .SDcols=c(\"id\", selected_cols)]","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:06:27.699540Z","iopub.execute_input":"2022-07-12T17:06:27.701009Z","iopub.status.idle":"2022-07-12T17:06:27.751592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let us scale the columns.","metadata":{}},{"cell_type":"code","source":"data_dt2[, setdiff(names(data_dt2), \"id\") := lapply(.SD, scale), .SDcols=setdiff(names(data_dt2), \"id\")]","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:06:55.783537Z","iopub.execute_input":"2022-07-12T17:06:55.785133Z","iopub.status.idle":"2022-07-12T17:06:55.854057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, let us apply power transformation with Yeo-Johnson algorithm as the data set contains some negative values.","metadata":{}},{"cell_type":"code","source":"pre_process <- preProcess(data_dt2[, -c(\"id\")], method=\"YeoJohnson\")\ndata_dt3 <- predict(pre_process, data_dt2[, -c(\"id\")])","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:07:45.943122Z","iopub.execute_input":"2022-07-12T17:07:45.944553Z","iopub.status.idle":"2022-07-12T17:07:46.954796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can check the distribution before and after of $f_09$, which has the largest variance.","metadata":{}},{"cell_type":"code","source":"par(mfrow=c(1,2))\nqqnorm(y=data_dt2$f_09, main=\"Normal Q-Q Plot before Transformation\")\nqqline(y=data_dt2$f_09, col='red')\nqqnorm(y=data_dt3$f_09, main=\"Normal Q-Q Plot after Transformation\")\nqqline(y=data_dt3$f_09, col='red')\npar(mfrow=c(1,1))","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:16:40.106855Z","iopub.execute_input":"2022-07-12T17:16:40.108368Z","iopub.status.idle":"2022-07-12T17:16:48.592215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The power transformation seems to help to bring the new distribution closer to normal-like.","metadata":{}},{"cell_type":"markdown","source":"Now, let us check the boxplots again.","metadata":{}},{"cell_type":"code","source":"data_dt3[, -c(\"id\")] %>%\n  melt(id.vars=NULL) %>%\n  ggplot(aes(x=variable, y=value)) +\n  geom_boxplot() +\n  theme(axis.text.x=element_text(angle=90, vjust=0.5, hjust=1))","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:17:44.011391Z","iopub.execute_input":"2022-07-12T17:17:44.013053Z","iopub.status.idle":"2022-07-12T17:17:46.601587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Determining Optimal Number of Clusters","metadata":{}},{"cell_type":"code","source":"set.seed(1234)\n\n# function to compute total within-cluster sum of square \nwss <- function(k) {\n  kmeans(data_dt3, k, nstart = 10)$tot.withinss\n}\n\n# Compute and plot wss for k = 1 to k = 10\nk.values <- 1:10\n\n# extract wss for 2-15 clusters\nwss_values <- unlist(lapply(k.values, wss))","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:19:06.818535Z","iopub.execute_input":"2022-07-12T17:19:06.820209Z","iopub.status.idle":"2022-07-12T17:21:04.955324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.table(k.values=k.values,\n          wss_values=wss_values) %>%\n    ggplot(aes(x=k.values, y=wss_values)) +\n    geom_line() +\n    geom_point() +\n    labs(x=\"Number of Clusters K\", y=\"Total within-cluster sum of squares\")\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:21:40.791462Z","iopub.execute_input":"2022-07-12T17:21:40.793349Z","iopub.status.idle":"2022-07-12T17:21:41.167289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the elbow test above, cluster numbers from 6 to 8 all seem OK. Therefore, we select 7 as the optimal as is used by the majority of the other notebooks.","metadata":{}},{"cell_type":"markdown","source":"## Baseline modeling","metadata":{}},{"cell_type":"markdown","source":"Now, let us use GMM to detect the clusters.","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:22:22.415567Z","iopub.execute_input":"2022-07-12T17:22:22.419435Z","iopub.status.idle":"2022-07-12T17:22:22.436860Z"}}},{"cell_type":"code","source":"num_clusters <- 7\nX <- data_dt3\n\ntic()\nmod <- Mclust(X, G=1:7)\ntoc()\nsummary(mod$BIC)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T17:22:56.551502Z","iopub.execute_input":"2022-07-12T17:22:56.553367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us review the model summary.","metadata":{}},{"cell_type":"code","source":"summary(mod)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can view the result of the clustering.","metadata":{}},{"cell_type":"code","source":"drmod <- MclustDR(mod, lambda=1)\nsummary(drmod)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(drmod, what=\"contour\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(drmod, what=\"boundaries\", ngrid=200)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, let us build the submission data.","metadata":{}},{"cell_type":"code","source":"final_submission_dt <-\n  data.table(id=data_dt$id,\n             Predicted=mod$classification-1)\n\nfinal_submission_dt","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fwrite(final_submission_dt, \"final_submission_dt.csv\")","metadata":{},"execution_count":null,"outputs":[]}]}