{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# AMEX Default Prediction - baseline model in R (Part 1)","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle"}},{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"AMEX Default Prediction has truly become my favorite hobby for the past month.  I have had a lot of fun reading innovative notebooks as well as insightful discussions. Meanwhile, I wish to see more R notebooks as currently all the high-score notebooks are exclusively in Python. ","metadata":{}},{"cell_type":"markdown","source":"As I primarily program in R, I have started a lonely journey. Based on my limited experiences, here are my top three obstacles to create a high score R notebook for this competition.\n\n1. The size of the data is huge and it is very challenging to build a notebook on Kaggle within 16 GB limit of memory size.  Python notebooks make innovative use of the data types to minimize memory footprint.  This is however, not an options for R.\n\n2. Data preprocessing takes a lot of time.  Several Python notebooks leverage existing packages in GPU computing to dramatically reduce the computing time. Unfortunately, R programmers have great challenges to follow this route as very limited GPU packages exist in R.\n\n3. Tree-boosting algorithms (e.g., xgboost, lightgbm, and catboost) generally provide better support for Python programmers. In particularly, releasing memory after training is quite challenging for the R libraries. \n","metadata":{}},{"cell_type":"markdown","source":"Fortunately, I have found good alternatives. More specifically,\n\n1. We will use excellent R packages ff and ffbase to map memory to hard drivers. As we shall demonstrate subsequently, these two packages provide a low maintenance solution to move data between hard drive and memory.\n\n2. We will use the power house of number crunching, i.e., the data.table package, to expedite data processing. Together with data.table's integration of multithreading, we can make the preprocessing time very fast.\n\n3. We will use Kaggle notebook's data save functionality to force release memory.  Memory units held by external data structures are forced to release once the current R session ends.","metadata":{}},{"cell_type":"markdown","source":"This is the first part of the notebook. In this part, we will cover data preprocessing and model fitting (with lightgbm). In the next part, we will cover the prediction.","metadata":{}},{"cell_type":"markdown","source":"## Data Preprocessing","metadata":{}},{"cell_type":"markdown","source":"We will use the well-known parquet data set prepared by [radar](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format).  We will use [arrow](https://cran.r-project.org/web/packages/arrow/index.html) to read the content and use [ff](https://cran.r-project.org/web/packages/ff/index.html) and [ffbase](https://github.com/edwindj/ffbase) packages to convert it to ffdf format.","metadata":{}},{"cell_type":"markdown","source":"Here are the list of libraries that we use.","metadata":{"execution":{"iopub.status.busy":"2022-06-19T03:38:39.658713Z","iopub.execute_input":"2022-06-19T03:38:39.66237Z","iopub.status.idle":"2022-06-19T03:38:39.944489Z"}}},{"cell_type":"code","source":"devtools::install_github(\"edwindj/ffbase\", subdir=\"pkg\")","metadata":{"execution":{"iopub.status.busy":"2022-07-29T00:59:23.209332Z","iopub.execute_input":"2022-07-29T00:59:23.293310Z","iopub.status.idle":"2022-07-29T00:59:47.225983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(ffbase)\nlibrary(magrittr)\nlibrary(data.table)\nlibrary(tictoc)\nlibrary(caret)\nlibrary(lightgbm)\n\noptions(DT.warn.size=FALSE)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T00:59:47.230071Z","iopub.execute_input":"2022-07-29T00:59:47.231700Z","iopub.status.idle":"2022-07-29T00:59:49.572893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are some of the global settings that we use for this notebook.","metadata":{}},{"cell_type":"code","source":"#~ Overall parameters\n\nPARQUET_DAT_DIR <- '../input/amex-data-integer-dtypes-parquet-format'\nCSV_DAT_DIR <- '../input/amex-default-prediction'\nSAVE_CV_MODEL_DIR <- '.'\nDEFAULT_DT_THREADS <- getDTthreads()\nDEFAULT_BATCHBYTES <- 536870912\ncat(paste(\"Number of threads for data.table: \", DEFAULT_DT_THREADS, \"\\n\", sep=\"\"))\ncat(paste(\"Default Batch Bytes for ffdfdply: \", DEFAULT_BATCHBYTES, \"\\n\", sep=\"\"))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T00:59:53.552000Z","iopub.execute_input":"2022-07-29T00:59:53.553913Z","iopub.status.idle":"2022-07-29T00:59:53.579422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will first create the training data by first reading into a data.table so that we can convert customer_ID column to factor.  This would enable us the save some storage as well.","metadata":{"execution":{"iopub.status.busy":"2022-06-19T01:54:23.132953Z","iopub.execute_input":"2022-06-19T01:54:23.135712Z","iopub.status.idle":"2022-06-19T01:54:23.160416Z"}}},{"cell_type":"code","source":"train_dt <- \n  arrow::read_parquet(file.path(PARQUET_DAT_DIR, \"train.parquet\")) \n\nsetDT(train_dt)\n\ntrain_dt[, S_2:=as.Date(S_2)]\ntrain_dt[, customer_ID:=factor(customer_ID)]\n\nsetorder(train_dt, customer_ID, S_2)\n\ncustomer_ID_levels <- levels(train_dt$customer_ID)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T00:59:57.344243Z","iopub.execute_input":"2022-07-29T00:59:57.345670Z","iopub.status.idle":"2022-07-29T01:00:34.581024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ff <- as.ffdf(train_dt)\nrm(train_dt); gc()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:00:42.676131Z","iopub.execute_input":"2022-07-29T01:00:42.677730Z","iopub.status.idle":"2022-07-29T01:01:07.570683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"In this section, we will follow the ideas in [@HUSEYINCOTEL's notebook](https://www.kaggle.com/code/huseyincot/amex-catboost-0-793) that are subsequently embraced by essentially all the notebooks, e.g., [this one](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793) and [this one](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart).","metadata":{}},{"cell_type":"code","source":"\ncat_cols <- c(\n  \"B_30\",\n  \"B_38\",\n  \"D_114\",\n  \"D_116\",\n  \"D_117\",\n  \"D_120\",\n  \"D_126\",\n  \"D_63\",\n  \"D_64\",\n  \"D_66\",\n  \"D_68\"\n)\n\ndate_cols <- c(\"S_2\")\nnum_cols <- setdiff(names(train_ff), c(\"customer_ID\", date_cols, cat_cols))","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:01:33.876323Z","iopub.execute_input":"2022-07-29T01:01:33.877827Z","iopub.status.idle":"2022-07-29T01:01:33.894057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We partition the data set into categorical and numerical. The function calc_all_cat_features() produces the resultant feature matrix for the input categorical columns.","metadata":{}},{"cell_type":"code","source":"calc_all_cat_features <- function(df, cat_cols) {\n  dat_dt <- as.data.table(df)\n  cust_dat_list <-\n    list(\n      N       = dat_dt[, .N, by=c(\"customer_ID\")],\n      \n      last    = dat_dt[, lapply(.SD, last), by=c(\"customer_ID\"), .SDcols=cat_cols],\n      \n      uniqueN = dat_dt[, lapply(.SD, uniqueN, na.rm=TRUE), by=c(\"customer_ID\"), .SDcols=cat_cols]\n    )\n\n  cust_dat_dt <-\n    setDT(unlist(cust_dat_list, recursive=FALSE))[]\n  \n  setnames(cust_dat_dt, \"N.customer_ID\", \"customer_ID\")\n  cust_dat_dt[, c(\"last.customer_ID\", \"uniqueN.customer_ID\"):=NULL]\n  \n  rm(cust_dat_list)\n  \n  # convert Inf/-Inf to NA so that tree boosting algorithms won't complain\n  invisible(lapply(names(cust_dat_dt),\n                   function(.name) set(cust_dat_dt, \n                                       which(is.infinite(cust_dat_dt[[.name]])), \n                                       j = .name,\n                                       value =NA)))  \n  \n  return(as.data.frame(cust_dat_dt))\n}\n","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:01:37.546961Z","iopub.execute_input":"2022-07-29T01:01:37.548339Z","iopub.status.idle":"2022-07-29T01:01:37.560370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"While less a RAM burden for generating the categorical features, we will still demo the practice of combining ffdfdply function in the ffbase package and the data.table package.","metadata":{}},{"cell_type":"code","source":"tic()\ncat_features_ff <-\n  ffdfdply(train_ff[c(\"customer_ID\", cat_cols)], split=train_ff$customer_ID, \n           FUN=calc_all_cat_features, cat_cols=cat_cols, BATCHBYTES = DEFAULT_BATCHBYTES)\ntoc()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:01:44.479247Z","iopub.execute_input":"2022-07-29T01:01:44.480686Z","iopub.status.idle":"2022-07-29T01:03:15.538675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, ffdfdply function does a fantastic job to hide the tedious steps to manage the files on the hard drive.  The difference is huge when I had to manually divide data file into several parts and read them one by one for my first few notebooks without it.","metadata":{}},{"cell_type":"markdown","source":"Now, let us proceed to the numerical features.","metadata":{"execution":{"iopub.status.busy":"2022-06-19T02:15:52.82114Z","iopub.execute_input":"2022-06-19T02:15:52.823735Z","iopub.status.idle":"2022-06-19T02:15:52.842847Z"}}},{"cell_type":"code","source":"calc_all_num_features <- function(df, num_cols) {\n  dat_dt <- as.data.table(df)\n  cust_dat_list <- list(\n    mean = dat_dt[, lapply(.SD, mean, na.rm=TRUE), by = c(\"customer_ID\"), .SDcols=num_cols],\n   \n    min  = dat_dt[, lapply(.SD, min, na.rm=TRUE), by = c(\"customer_ID\"), .SDcols=num_cols],\n    \n    max  = dat_dt[, lapply(.SD, max, na.rm=TRUE), by = c(\"customer_ID\"), .SDcols=num_cols],\n    \n    sd   = dat_dt[, lapply(.SD, sd, na.rm=TRUE), by = c(\"customer_ID\"), .SDcols=num_cols],\n      \n    q25  = dat_dt[, lapply(.SD, quantile, prob = c(0.25), na.rm=TRUE), by = c(\"customer_ID\"), .SDcols=c(\"P_2\", \"B_9\", \"R_1\", \"D_42\", \"D_44\", \"S_3\")], \n    q75  = dat_dt[, lapply(.SD, quantile, prob = c(0.75), na.rm=TRUE), by = c(\"customer_ID\"), .SDcols=c(\"P_2\", \"B_9\", \"R_1\", \"D_42\", \"D_44\", \"S_3\")],\n      \n    last = dat_dt[, lapply(.SD, last), by = c(\"customer_ID\"), .SDcols=num_cols]\n  )\n  \n  cust_dat_dt <-\n    setDT(unlist(cust_dat_list, recursive=FALSE))[]\n  \n  setnames(cust_dat_dt, 'mean.customer_ID', 'customer_ID')\n  #cust_dat_dt[, q25.customer_ID:=NULL]\n  #cust_dat_dt[, q75.customer_ID:=NULL]\n    \n  cust_dat_dt[, `:=`(min.customer_ID=NULL,\n                     max.customer_ID=NULL,\n                     sd.customer_ID=NULL,\n                     q25.customer_ID=NULL, q75.customer_ID=NULL,\n                     last.customer_ID=NULL)]\n  \n  rm(cust_dat_list)\n  \n  # convert Inf/-Inf to NA so that tree boosting algorithms won't complain\n  invisible(lapply(names(cust_dat_dt),\n                   function(.name) set(cust_dat_dt, \n                                       which(is.infinite(cust_dat_dt[[.name]])), \n                                       j = .name,\n                                       value =NA)))\n\n  return(as.data.frame(cust_dat_dt))\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:05:48.037215Z","iopub.execute_input":"2022-07-29T01:05:48.038670Z","iopub.status.idle":"2022-07-29T01:05:48.051333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For this next step, we will temporarily turn off warning in this script as the data set contains many NAs, which in turn triggers warnings of no values inside min(), max(), mean() functions when we generate features.","metadata":{}},{"cell_type":"code","source":"oldw <- getOption(\"warn\")\noptions(warn = -1)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:05:56.485683Z","iopub.execute_input":"2022-07-29T01:05:56.487091Z","iopub.status.idle":"2022-07-29T01:05:56.500489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tic()\nnum_features_ff <-\n  ffdfdply(train_ff[c(\"customer_ID\", num_cols)], split=train_ff$customer_ID, \n           FUN=calc_all_num_features, num_cols=num_cols, \n           BATCHBYTES = DEFAULT_BATCHBYTES)\ntoc()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:05:59.517526Z","iopub.execute_input":"2022-07-29T01:05:59.518927Z","iopub.status.idle":"2022-07-29T01:21:47.985829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"options(warn = oldw)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:22:18.677492Z","iopub.execute_input":"2022-07-29T01:22:18.678838Z","iopub.status.idle":"2022-07-29T01:22:18.710144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With only 2 threads, data.table manages to complete the feature generation in then 6 minutes. This is certainly quite impressive for a huge problem like this.","metadata":{}},{"cell_type":"markdown","source":"## Fitting Model","metadata":{}},{"cell_type":"markdown","source":"Let us first generate 5-Fold for cross validation.  We will use lightgbm for fitting the model. We will use the hyperparameter setting as in [this notebook](https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart). ","metadata":{}},{"cell_type":"code","source":"\ntrain_labels_dt <- fread(file.path(CSV_DAT_DIR, \"train_labels.csv\"), stringsAsFactors = TRUE)\ntrain_labels_dt[, customer_ID:=factor(customer_ID, customer_ID_levels)]\nsetorder(train_labels_dt, customer_ID)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:22:23.134896Z","iopub.execute_input":"2022-07-29T01:22:23.136405Z","iopub.status.idle":"2022-07-29T01:22:24.411564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"set.seed(1234)\nnum_folds <- 5\ntest_folds <- createFolds(train_labels_dt$target, k=num_folds, list=TRUE, returnTrain=FALSE)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:23:17.633176Z","iopub.execute_input":"2022-07-29T01:23:17.634575Z","iopub.status.idle":"2022-07-29T01:23:17.754721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We use our own implementation of [competition metric function in R](https://www.kaggle.com/code/stautxie/amex-metric-implementation-in-r).","metadata":{}},{"cell_type":"code","source":"\ntop_four_percent_captured <- function(y, pred) {\n  dt <- data.table(y=y,\n                   pred=pred)\n  setorder(dt, -pred, na.last=TRUE)\n  \n  dt[, weight:=ifelse(y == 0, 20, 1)]\n  four_pct_cutoff <- as.integer(0.04 * dt[, sum(weight)])\n  dt[, weight_cumsum:=cumsum(weight)]\n  dt_cutoff <- dt[weight_cumsum <= four_pct_cutoff]\n  return(dt_cutoff[y == 1, .N] / dt[y==1, .N])\n}\n\n#~\n\nweighted_gini <- function(y, pred) {\n  dt <- data.table(y=y,\n                   pred=pred)\n  setorder(dt, -pred, na.last=TRUE)\n  dt[, weight:=ifelse(y==0, 20, 1)]\n  dt[, random:=cumsum(weight / sum(weight))]\n  total_pos <- dt[, sum(weight * y)]\n  dt[, cum_pos_found:=cumsum(y * weight)]\n  dt[, lorenz:=cum_pos_found/total_pos]\n  dt[, gini:=(lorenz - random)*weight]\n  \n  return(as.numeric(dt[, sum(gini)]))\n}\n\n#~\n\nnormalized_weighted_gini <- function(y, pred) {\n  return(weighted_gini(y, pred) / weighted_gini(y, y))\n}\n\n#~\n\namex_metric <- function(y, pred) {\n  return (0.5 * normalized_weighted_gini(y, pred) +\n            0.5 * top_four_percent_captured(y, pred))\n}\n\n#~\n\namex_metric_lgb <- function(y_pred, dtrain) {\n  y_true <- getinfo(dtrain, \"label\")\n  amex_metric_val <- amex_metric(y_true, y_pred)\n  return(list(name=\"AMEX metric\", \n              value=amex_metric_val,\n              higher_better=TRUE))\n}\n","metadata":{"execution":{"iopub.status.busy":"2022-07-29T01:23:24.888579Z","iopub.execute_input":"2022-07-29T01:23:24.890084Z","iopub.status.idle":"2022-07-29T01:23:24.910010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let us kick off fitting the model.","metadata":{}},{"cell_type":"code","source":"fold_result_list <- list()\n\nfor (fold in 1:num_folds) {\n  # save the folds to continue model build because of Kaggle time limit\n  test_fold <- test_folds[[fold]]\n  write.csv(test_fold, paste(SAVE_CV_MODEL_DIR, \"/test_fold\", fold, \".csv\", sep=\"\")) \n}\n# test_fold<-list()\nfor (fold in 1:num_folds) {\n  # get current fold\n  cat(paste(\"Processing fold \", fold, \"\\n\", sep=\"\"))\n  test_fold <- test_folds[[fold]]\n  # file.exists(paste(SAVE_CV_MODEL_DIR, \"/test_fold\", fold, \".csv\", sep=\"\"))\n  # df<-read.csv(paste(\"../input/test-folds\", \"/test_fold\", fold, \".csv\", sep=\"\"))  \n  # test_fold<-df[,2]\n  cat(paste(\"Creating training data\\n\"))\n  cat_features_dt <- as.data.table(cat_features_ff)\n  setorder(cat_features_dt, customer_ID)\n   \n  num_features_dt <- as.data.table(num_features_ff)\n  setorder(num_features_dt, customer_ID)\n  \n  tr_cat_features_dt <- cat_features_dt[-test_fold, -c(\"customer_ID\")]\n  tr_num_features_dt <- num_features_dt[-test_fold, -c(\"customer_ID\")]\n  \n  va_cat_features_dt <- cat_features_dt[test_fold, -c(\"customer_ID\")]\n  va_num_features_dt <- num_features_dt[test_fold, -c(\"customer_ID\")]  \n    \n  cat_feature_names<-names(tr_cat_features_dt)\n  tr_target <- train_labels_dt[-test_fold]$target\n  rm(cat_features_dt, num_features_dt); gc()\n  \n  tr_mat_dt <- setDT(unlist(list(tr_cat_features_dt,\n                                 tr_num_features_dt), recursive=FALSE))\n  \n  rm(tr_cat_features_dt, tr_num_features_dt); gc()\n  \n  mat_cat_cols <- c(outer(c(\"last\"), cat_cols, FUN=paste, sep=\".\"))\n  \n  lgb_data_tr <- lgb.Dataset(data = tr_mat_dt %>% as.matrix(),\n                             label = tr_target)\n  \n  # lgb.Dataset.set.categorical(lgb_data_tr, mat_cat_cols)\n  \n  rm(tr_mat_dt); gc()\n  \n  # load validation data\n  \n  cat(paste(\"Creating validation data\\n\"))\n  # cat_features_dt <- as.data.table(cat_features_ff)\n  # setorder(cat_features_dt, customer_ID)\n  \n  # num_features_dt <- as.data.table(num_features_ff)\n  # setorder(num_features_dt, customer_ID)\n  \n  # va_cat_features_dt <- cat_features_dt[test_fold, -c(\"customer_ID\")]\n  # va_num_features_dt <- num_features_dt[test_fold, -c(\"customer_ID\")]\n  va_target <- train_labels_dt[test_fold]$target\n  # rm(cat_features_dt, num_features_dt); gc()\n  \n  va_mat_dt <- setDT(unlist(list(va_cat_features_dt,\n                                 va_num_features_dt), recursive=FALSE))\n  \n  # rm(va_cat_features_dt, va_num_features_dt); gc()\n  \n  lgb_data_va <- lgb.Dataset(data = va_mat_dt %>% as.matrix(),\n                             label = va_target)\n  \n  # lgb.Dataset.set.categorical(lgb_data_va, mat_cat_cols)\n  \n  # rm(va_mat_dt); gc()\n  \n  # Run lightgbm fitting\n  cat(paste(\"Predicting on the fold\\n\"))\n  set.seed(1234)\n  random_state <- 1234\n  \n  #lgb_params = list(objective = \"binary\",\n  #                  random_state=random_state,\n  #                  learning_rate=0.01,\n  #                  # boosting = \"dart\",\n  #                  n_estimators=1300,\n  #                  min_child_samples=2400,\n  #                  num_leaves=95,\n  #                  max_bin=511)\n  \n  #lgb_params = list(objective = \"binary\",  \n  #      metric = \"binary_logloss\",\n  #      boosting = \"goss\",\n  #      num_leaves = 100,\n  #      learning_rate = 0.001,\n  #      feature_fraction = 0.50,\n  #      n_jobs = 8,\n  #      lambda_l1 = 0.5,            \n  #      lambda_l2 = 0.5,\n  #      min_data_in_leaf = 200)\n  lgb_params = list(objective = \"binary\", \n        metric = \"binary_logloss\",\n        boosting = \"goss\",\n        num_leaves = 100,\n        learning_rate = 0.001,\n        max_depth = -1,\n        # feature_fraction = 0.041,\n        # bagging_freq = 5,\n        # bagging_fraction = 0.331,\n        n_jobs = 8,\n        min_data_in_leaf = 80,\n        min_sum_hessian_in_leaf =10.0,\n        mx_bin=512)\n\n  tic()\n  md_lgb <- lgb.train(params=lgb_params,\n                      data=lgb_data_tr,\n                      eval=amex_metric_lgb,\n                      # eval='binary',\n                      nrounds=10800,\n                      valids=list(tune=lgb_data_va),\n                      categorical_feature=cat_feature_names,\n                      early_stopping_rounds = 100,\n                      reset_data=TRUE,\n                      eval_freq = 100) \n                      \n  toc()\n  \n  # save the fitted model\n  \n  lgb.save(md_lgb, file.path(SAVE_CV_MODEL_DIR, paste(\"md_lgb_fold\", fold, \".txt\", sep=\"\")))\n  \n  # save weight for importance analysis\n  \n  importance <- lgb.importance(md_lgb, percentage = TRUE)\n  write.csv(importance, paste(SAVE_CV_MODEL_DIR, \"/importance\", fold, \".csv\", sep=\"\"))\n  # save prediction\n  \n  va_pred <- predict(md_lgb, data = va_mat_dt %>% as.matrix())\n  cat(paste('Fold ', fold, ' AMEX metric: ',\n            amex_metric(va_target, va_pred), \"\\n\", sep = \"\"))\n  \n  # save log-odd prediction\n  \n  va_pred_logodds <- predict(md_lgb, data = va_mat_dt %>% as.matrix(), rawscore = TRUE)\n  \n  # calculate AMEX metric\n  \n  model_score <- amex_metric(va_target, va_pred)\n  \n  # save to result object\n  cat(paste(\"Saving results and clean up\\n\"))\n  \n  fold_result_list[[fold]] <- \n    list(model_score = model_score,\n         va_target = va_target,\n         va_pred = va_pred,\n         va_pred_logodds = va_pred_logodds,\n         importance = importance)\n  \n  # clean up\n  lgb.unloader(restore=TRUE, wipe=TRUE, envir=environment())\n  rm(va_mat_dt); gc()\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Display feature lgb.importance","metadata":{}},{"cell_type":"code","source":"importance<- data.table::fread(\"./importance1.csv\")\nlgb.plot.importance(importance, top_n = 150L, measure = \"Gain\")\n\nimportance<- data.table::fread(\"./importance2.csv\")\nlgb.plot.importance(importance, top_n = 150L, measure = \"Gain\")\n\nimportance<- data.table::fread(\"./importance3.csv\")\nlgb.plot.importance(importance, top_n = 150L, measure = \"Gain\")\n\n#importance<- data.table::fread(\"./importance4.csv\")\n#lgb.plot.importance(importance, top_n = 150L, measure = \"Gain\")\n\n#importance<- data.table::fread(\"./importance5.csv\")\n#lgb.plot.importance(importance, top_n = 150L, measure = \"Gain\")\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-25T04:42:50.031798Z","iopub.execute_input":"2022-07-25T04:42:50.034087Z","iopub.status.idle":"2022-07-25T04:42:50.935718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In summary, we have fitted model with 5-fold CV with AMEX metrics between 0.789 and 0.796.","metadata":{"execution":{"iopub.status.busy":"2022-06-19T04:08:14.163209Z","iopub.execute_input":"2022-06-19T04:08:14.164896Z","iopub.status.idle":"2022-06-19T04:08:14.229657Z"}}}]}