{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"> This kernel is a XGB Tutorial with managing the memory that added code-comments and detailed descriptions based on [XGBoost Starter - LB 0.793](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793).","metadata":{}},{"cell_type":"markdown","source":"> TOC\n```\n1. Load Libraries\n2. Load Dataset and Manage the GPU Memory\n3. Feature Engineering\n4. Train XGB\n5. Save OOF Preds\n6. Feature Importance\n7. Data Processing and Feature Engineering for Test Data\n8. Infer Test\n9. Create Submission CSV\n```","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:42:35.078701Z","iopub.execute_input":"2022-07-17T15:42:35.079157Z","iopub.status.idle":"2022-07-17T15:42:35.087982Z","shell.execute_reply.started":"2022-07-17T15:42:35.079116Z","shell.execute_reply":"2022-07-17T15:42:35.086913Z"}}},{"cell_type":"markdown","source":"# 1. Load Libraries","metadata":{}},{"cell_type":"markdown","source":"When we train machine learning and deep learning, it is common to use the CPU to perform data preprocessing,etc. and to train the model with the GPU. And in this case, the process of copying(moving) the data loaded to the CPU to the GPU is required. For example, Working with dataframes with pandas, processing data on GPU memory with torch, and so on.\n\nInstead of doing that now, RAPIDS came out with the concept of 'Yo, Let's do the whole process on the GPU!'. RAPIDS is a CUDA process-based data science platform built and operated by NVIDIA.\n\nAll packages are named cuxx to emphasize that they are based on CUDA. By replacing pandas with cudf, numpy with cupy, and sklearn with cuml, it has been developed to use almost the same functions as before.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport cupy\nimport cudf\n\nimport matplotlib.pyplot as plt, gc, os\n\nprint('cudf version',cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:11.091036Z","iopub.execute_input":"2022-07-17T15:10:11.09141Z","iopub.status.idle":"2022-07-17T15:10:12.511764Z","shell.execute_reply.started":"2022-07-17T15:10:11.091379Z","shell.execute_reply":"2022-07-17T15:10:12.510966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Load Dataset and Manage the GPU Memory\n\nWhen using a GPU, we need to manage our GPU resources well with monitoring currently available GPU resources through the nvidia-smi shell command, etc.","metadata":{}},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:15.898746Z","iopub.execute_input":"2022-07-17T15:10:15.899407Z","iopub.status.idle":"2022-07-17T15:10:16.609671Z","shell.execute_reply.started":"2022-07-17T15:10:15.899371Z","shell.execute_reply":"2022-07-17T15:10:16.608704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One Tesla P100 is available and there are currently no processes running with that GPU.\n\nNow we will load the data onto GPU memory with cudf. The basic functions and usage are the same as in pandas.","metadata":{}},{"cell_type":"code","source":"# read_parquet() function reads a file in parquet format.\ndf = cudf.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:17.630807Z","iopub.execute_input":"2022-07-17T15:10:17.631205Z","iopub.status.idle":"2022-07-17T15:10:39.510018Z","shell.execute_reply.started":"2022-07-17T15:10:17.631172Z","shell.execute_reply":"2022-07-17T15:10:39.509199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:39.511671Z","iopub.execute_input":"2022-07-17T15:10:39.512042Z","iopub.status.idle":"2022-07-17T15:10:39.83058Z","shell.execute_reply.started":"2022-07-17T15:10:39.512014Z","shell.execute_reply":"2022-07-17T15:10:39.829831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:39.83174Z","iopub.execute_input":"2022-07-17T15:10:39.832581Z","iopub.status.idle":"2022-07-17T15:10:40.565054Z","shell.execute_reply.started":"2022-07-17T15:10:39.832541Z","shell.execute_reply":"2022-07-17T15:10:40.56392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Putting data on the GPU like this will take up 3669 MiB of memory in Memory-Usage. Given that the total available memory is 16280 MiB, it should be borne in mind that memory may be overwitten if the variable is copied several times or another dataset(eg, evaluation data) is loaded.","metadata":{"execution":{"iopub.status.busy":"2022-07-16T01:00:25.059036Z","iopub.execute_input":"2022-07-16T01:00:25.05959Z","iopub.status.idle":"2022-07-16T01:00:25.066313Z","shell.execute_reply.started":"2022-07-16T01:00:25.059548Z","shell.execute_reply":"2022-07-16T01:00:25.06519Z"}}},{"cell_type":"markdown","source":"So, if you want to cleanly convert `customer_ID` to a number, you can use the hex_to_int() and astype() functions.\n\nLet's apply that only for the last 16 characters. That's the unique characters.","metadata":{}},{"cell_type":"code","source":"df['customer_ID'] = df['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:40.567384Z","iopub.execute_input":"2022-07-17T15:10:40.56778Z","iopub.status.idle":"2022-07-17T15:10:41.059191Z","shell.execute_reply.started":"2022-07-17T15:10:40.567735Z","shell.execute_reply":"2022-07-17T15:10:41.058431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:41.063298Z","iopub.execute_input":"2022-07-17T15:10:41.065568Z","iopub.status.idle":"2022-07-17T15:10:41.932017Z","shell.execute_reply.started":"2022-07-17T15:10:41.06553Z","shell.execute_reply":"2022-07-17T15:10:41.931034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to get used to it, we will continue to track the memory as much as possible. Just truncating some of the data like this can reduce the memory being used.\n\nAbout 300 MiB has been reduced.","metadata":{}},{"cell_type":"markdown","source":"Column `S_2`contained time information. Like pandas, the data type is changed with the to_datetime() function.","metadata":{}},{"cell_type":"code","source":"df['S_2'] = cudf.to_datetime(df['S_2'])\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:41.935391Z","iopub.execute_input":"2022-07-17T15:10:41.935691Z","iopub.status.idle":"2022-07-17T15:10:42.218694Z","shell.execute_reply.started":"2022-07-17T15:10:41.935662Z","shell.execute_reply":"2022-07-17T15:10:42.217928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:42.220043Z","iopub.execute_input":"2022-07-17T15:10:42.220564Z","iopub.status.idle":"2022-07-17T15:10:42.946499Z","shell.execute_reply.started":"2022-07-17T15:10:42.220524Z","shell.execute_reply":"2022-07-17T15:10:42.945541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is also important to change it to an appropriate data type. In Python's memory operation process, string types use a lot of memory by default. So, if you can encode a categorical variable or change it to an integer type, it is advantageous for memory management.","metadata":{}},{"cell_type":"code","source":"df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:42.947881Z","iopub.execute_input":"2022-07-17T15:10:42.948189Z","iopub.status.idle":"2022-07-17T15:10:43.197592Z","shell.execute_reply.started":"2022-07-17T15:10:42.94816Z","shell.execute_reply":"2022-07-17T15:10:43.196689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a lot of empty cells. We will use the fillna() function to fill empty cells with -127. The lowest number that can be expressed in 1 byte(8-bits) is -129. A signed integer can be represented from -128 to 127, and it seems that the original author wanted to replace null values with minimal memory usage.\n\nHowever, since it is a processing that does not consider the distribution of each variable, there may be some disadvantages in model training.\n\nAnd in the original text, when executing the fillna() function with cudf, it was written as `df = df.fillna(NAN_VALUE)`, but we will write it as `df.fillna(NAN_VALUE, inplace=True)`. When written as in the original code, the GPU memory usage nearly doubles (3321MiB -> 6079MiB) as the null-padded df is copied and allocated to a new memory address. Since we are not working on a sufficient memory environment, we must set the `inplace=Tru` option to overwrite the currently allocated memory address.","metadata":{}},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:43.200149Z","iopub.execute_input":"2022-07-17T15:10:43.200767Z","iopub.status.idle":"2022-07-17T15:10:43.210091Z","shell.execute_reply.started":"2022-07-17T15:10:43.200727Z","shell.execute_reply":"2022-07-17T15:10:43.20917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"NAN_VALUE = -127 \ndf.fillna(NAN_VALUE, inplace=True)\ndf.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:43.212923Z","iopub.execute_input":"2022-07-17T15:10:43.213378Z","iopub.status.idle":"2022-07-17T15:10:43.717726Z","shell.execute_reply.started":"2022-07-17T15:10:43.213332Z","shell.execute_reply":"2022-07-17T15:10:43.716838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:43.71897Z","iopub.execute_input":"2022-07-17T15:10:43.719405Z","iopub.status.idle":"2022-07-17T15:10:44.449065Z","shell.execute_reply.started":"2022-07-17T15:10:43.719366Z","shell.execute_reply":"2022-07-17T15:10:44.448085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The read_file_GPU() function below performs this process at once.\n\nIn the original code, it was written only in such a way that GPU is used, but it is difficult to proceed with Kaggle servel. Therefore, we use the CPU in the process of loading the entire data, and in the process of learning or evaluating the data, So we will make another read_file_CPU() function.\n\nAnd we do not have enough CPU on Kaggle server. So if data is uploaded at once in the process of loading the test dataset later, the CPU memory will be exceeded and the Kaggle kernel is reset. Our CPU based function can load the data to CPU in batches and then, load the data to GPU in iteration form. \n\nIf you want to load data to the CPU, you can use pandas, and if you load data to the GPU,  you can use cudf.","metadata":{}},{"cell_type":"code","source":"from pyarrow.parquet import ParquetFile","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:19.901369Z","iopub.execute_input":"2022-07-17T15:11:19.902387Z","iopub.status.idle":"2022-07-17T15:11:19.907604Z","shell.execute_reply.started":"2022-07-17T15:11:19.902325Z","shell.execute_reply":"2022-07-17T15:11:19.906673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# original code\ndef read_file_GPU(path = '', usecols = None):\n    # read_parquet() function can read the parquet-type file.\n    # if you want to specify columns:\n    if usecols is not None: \n        df = cudf.read_parquet(path, columns=usecols)\n    # if you want to read all columns:\n    else: df = cudf.read_parquet(path)\n    \n    df['customer_ID'] = df['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\n    df.S_2 = cudf.to_datetime( df.S_2 )\n    df = df.fillna(NAN_VALUE) \n    print('shape of data:', df.shape)\n    \n    return df\n\n# modified code(CPU, batch load)\ndef read_file_CPU(path = '', iter_batch = None, usecols = None):\n    if usecols is not None:\n        # when retrieving only some columns(1~3), there is no problem even if all rows are retrieved.\n        df = pd.read_parquet(path, columns=usecols)\n    else:\n        # When importing all columns, data is imported in batch format.\n        df = iter_batch\n    \n    # it performs the same processing as it did with cudf.\n    df['customer_ID'] = df['customer_ID'].apply(lambda x : int(x[-16:],16)).astype('int64') \n    df.S_2 = pd.to_datetime( df.S_2 )\n    df.fillna(NAN_VALUE, inplace=True)\n    print('shape of data:', df.shape)\n    \n    return df\n\n# print('Reading train data...')\n# TRAIN_PATH = '../input/amex-data-integer-dtypes-parquet-format/train.parquet'\n# train = read_file(path = TRAIN_PATH)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:22.391593Z","iopub.execute_input":"2022-07-17T15:11:22.39202Z","iopub.status.idle":"2022-07-17T15:11:22.402457Z","shell.execute_reply.started":"2022-07-17T15:11:22.391986Z","shell.execute_reply":"2022-07-17T15:11:22.401622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The original kernel used variable name as train, not df. So we'll make it the same.\n\nHowever, the copy() function is not used here fore memory management. If you use the copy() function, it will be copied to another memory address.","metadata":{}},{"cell_type":"code","source":"train = df\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:30.826951Z","iopub.execute_input":"2022-07-17T15:11:30.827316Z","iopub.status.idle":"2022-07-17T15:11:31.045376Z","shell.execute_reply.started":"2022-07-17T15:11:30.827286Z","shell.execute_reply":"2022-07-17T15:11:31.044469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:34.351396Z","iopub.execute_input":"2022-07-17T15:11:34.351787Z","iopub.status.idle":"2022-07-17T15:11:35.10428Z","shell.execute_reply.started":"2022-07-17T15:11:34.351753Z","shell.execute_reply":"2022-07-17T15:11:35.103239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can see that additional memory is not used because only the variable name is changed(refer to memory) without copying.","metadata":{}},{"cell_type":"markdown","source":"# 3. Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"We do not put data as it is to the model, but transform it into statistical values before training and then train it.\n\nIn the process of data aggregation, the column takes on a multiindex format, but in order to include it in the model, a one-dimensional column must be maintained. So, Let's take a look at this task first and then run the whole process through a function.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T10:45:36.492633Z","iopub.execute_input":"2022-07-15T10:45:36.492999Z","iopub.status.idle":"2022-07-15T10:45:36.498767Z","shell.execute_reply.started":"2022-07-15T10:45:36.492968Z","shell.execute_reply":"2022-07-15T10:45:36.497787Z"}}},{"cell_type":"code","source":"multi_index_col_sample = train.groupby('customer_ID')[['B_30','B_38','D_114']].agg(['count','last','nunique']).columns\nmulti_index_col_sample","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:37.574266Z","iopub.execute_input":"2022-07-17T15:11:37.574645Z","iopub.status.idle":"2022-07-17T15:11:37.70543Z","shell.execute_reply.started":"2022-07-17T15:11:37.574614Z","shell.execute_reply":"2022-07-17T15:11:37.70469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"['_'.join(x) for x in multi_index_col_sample]","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:39.168928Z","iopub.execute_input":"2022-07-17T15:11:39.169292Z","iopub.status.idle":"2022-07-17T15:11:39.17522Z","shell.execute_reply.started":"2022-07-17T15:11:39.169262Z","shell.execute_reply":"2022-07-17T15:11:39.174351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can make it easier to see by attaching the double index as a one-dimensional index. Then, Let's check the entire function.","metadata":{}},{"cell_type":"code","source":"def process_and_feature_engineer(df):\n    # Put the remaining column names int the all_cols variable except for the customer_ID column and S_2 column in a list comprehension method.\n    # Columns in all_cols are the variables for training.\n    all_cols = [c for c in list(df.columns) if c not in ['customer_ID','S_2']]\n    \n    # Let's split the all_cols categorical variables and numerical variables.\n    # In the original code, categorical variables are splited like below but i find the others remained.\n    # But to maintain the countinuity and avoid confusion, we will use it as is.\n    cat_features = [\"B_30\",\"B_38\",\"D_114\",\"D_116\",\"D_117\",\"D_120\",\"D_126\",\"D_63\",\"D_64\",\"D_66\",\"D_68\"]\n    num_features = [col for col in all_cols if col not in cat_features]\n\n    # And aggregate numerical variables into statistics for each customer_ID.\n    test_num_agg = df.groupby(\"customer_ID\")[num_features].agg(['mean', 'std', 'min', 'max', 'last'])\n    # Aggregated columns are the type of MultiIndex. So connect them with '_' charactor like that we checked above.\n    test_num_agg.columns = ['_'.join(x) for x in test_num_agg.columns]\n\n    # Similary, Aggregate categorical variables for each customer_ID.\n    # 'count' is count the duplicate customer_ID's.\n    # 'last' is get the last value of each categorical variables.\n    # nunique() count the unique values of each categorical variables.\n    test_cat_agg = df.groupby(\"customer_ID\")[cat_features].agg(['count', 'last', 'nunique'])\n    test_cat_agg.columns = ['_'.join(x) for x in test_cat_agg.columns]\n\n    # Merge the numerical variables and categorical variables.\n    df = cudf.concat([test_num_agg, test_cat_agg], axis=1)\n    del test_num_agg, test_cat_agg\n    print('shape after engineering', df.shape)\n    \n    return df\n\n\ntrain = process_and_feature_engineer(train)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:41.268955Z","iopub.execute_input":"2022-07-17T15:11:41.269624Z","iopub.status.idle":"2022-07-17T15:11:42.444693Z","shell.execute_reply.started":"2022-07-17T15:11:41.26959Z","shell.execute_reply":"2022-07-17T15:11:42.443753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:43.4452Z","iopub.execute_input":"2022-07-17T15:11:43.445883Z","iopub.status.idle":"2022-07-17T15:11:44.275384Z","shell.execute_reply.started":"2022-07-17T15:11:43.445825Z","shell.execute_reply":"2022-07-17T15:11:44.274616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:45.052505Z","iopub.execute_input":"2022-07-17T15:11:45.053165Z","iopub.status.idle":"2022-07-17T15:11:45.940463Z","shell.execute_reply.started":"2022-07-17T15:11:45.053122Z","shell.execute_reply":"2022-07-17T15:11:45.939465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As the number of columns increased, the data capacity also increased.","metadata":{}},{"cell_type":"markdown","source":"We can load the target(label) data also. We will merge it into the train dataset created above.\n\nAt this time, customer_ID must be processed as an integer in the same way as before. \n\nThis task is to predict whether the customer will or will not pay the card expenses. Each customer's real repayment status is contained in train_labels.csv.","metadata":{}},{"cell_type":"code","source":"targets = cudf.read_csv('../input/amex-default-prediction/train_labels.csv')\ntargets['customer_ID'] = targets['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\ntargets = targets.set_index('customer_ID')\n\n# The index is the same as customer_ID, Merge based on the corresponding index.\ntrain = train.merge(targets, left_index=True, right_index=True, how='left')\n# Store the target data as an 8-bit integer.\ntrain.target = train.target.astype('int8')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:47.740869Z","iopub.execute_input":"2022-07-17T15:11:47.741264Z","iopub.status.idle":"2022-07-17T15:11:48.349517Z","shell.execute_reply.started":"2022-07-17T15:11:47.741232Z","shell.execute_reply":"2022-07-17T15:11:48.348642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = train.reset_index()\ntrain","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:49.931144Z","iopub.execute_input":"2022-07-17T15:11:49.931965Z","iopub.status.idle":"2022-07-17T15:11:51.098071Z","shell.execute_reply.started":"2022-07-17T15:11:49.931933Z","shell.execute_reply":"2022-07-17T15:11:51.09728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:52.002889Z","iopub.execute_input":"2022-07-17T15:11:52.003398Z","iopub.status.idle":"2022-07-17T15:11:52.726827Z","shell.execute_reply.started":"2022-07-17T15:11:52.003365Z","shell.execute_reply":"2022-07-17T15:11:52.725814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The targets are now merged into the train dataset, so deallocate them in memory.\n\nAlthough the capacity is small, it is better to always manage the memory after using the variable.","metadata":{}},{"cell_type":"code","source":"del targets","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:53.851445Z","iopub.execute_input":"2022-07-17T15:11:53.852014Z","iopub.status.idle":"2022-07-17T15:11:53.857696Z","shell.execute_reply.started":"2022-07-17T15:11:53.851976Z","shell.execute_reply":"2022-07-17T15:11:53.856865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:55.184144Z","iopub.execute_input":"2022-07-17T15:11:55.184725Z","iopub.status.idle":"2022-07-17T15:11:56.09629Z","shell.execute_reply.started":"2022-07-17T15:11:55.184683Z","shell.execute_reply":"2022-07-17T15:11:56.09504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Count the number of features to train the model, The first column is the id value and the last column is the label. Excluding 2 columns, the remaining number of columns is 198.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T11:15:43.092352Z","iopub.execute_input":"2022-07-15T11:15:43.092723Z","iopub.status.idle":"2022-07-15T11:15:43.098262Z","shell.execute_reply.started":"2022-07-15T11:15:43.092695Z","shell.execute_reply":"2022-07-15T11:15:43.097264Z"}}},{"cell_type":"code","source":"FEATURES = train.columns[1:-1]\nprint(f'There are {len(FEATURES)} features!')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:58.336092Z","iopub.execute_input":"2022-07-17T15:11:58.337168Z","iopub.status.idle":"2022-07-17T15:11:58.344696Z","shell.execute_reply.started":"2022-07-17T15:11:58.337117Z","shell.execute_reply":"2022-07-17T15:11:58.343892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Train XGB\n\nWhen training the model, we will utilize KFold cross-validation to avoid overfitting by using all data at lease once for traininig.\n\nThere is a problem in that the parts corresponding to 30 and 20 cannot be used for learning when the data is divided into learning/verification, such as 0:30 or 80:20 which are generally used for convenience. KFold splits training data and validation data(expressed as creating K-Folds) so that all datasets can be used for trainig.","metadata":{}},{"cell_type":"code","source":"# LOAD XGB LIBRARY\nfrom sklearn.model_selection import KFold\nimport xgboost as xgb\nprint('XGB Version',xgb.__version__)\n\n# XGB MODEL PARAMETERS\nxgb_parms = { \n    'max_depth':4, \n    'learning_rate':0.05, \n    'subsample':0.8,\n    'colsample_bytree':0.6, \n    'eval_metric':'logloss',\n    'objective':'binary:logistic',\n    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n    'random_state':42\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:00.3252Z","iopub.execute_input":"2022-07-17T15:12:00.32565Z","iopub.status.idle":"2022-07-17T15:12:00.445936Z","shell.execute_reply.started":"2022-07-17T15:12:00.325612Z","shell.execute_reply":"2022-07-17T15:12:00.445158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When traininig, we use DeviceQuantileDMatrix. All of usable-GPU memory is used at once for each calculation. And this DeviceQuantileDMatrix allows you to use GPU memory by dividing it into smaller units.\n\nIn order to use DeviceQuantileDMatrix, it is necessary to define a class that can pass data in batch form by repeatedly calling the iteration method, that is, the next() function.","metadata":{}},{"cell_type":"code","source":"class IterLoadForDMatrix(xgb.core.DataIter):\n    def __init__(self, df=None, features=None, target=None, batch_size=256*1024):\n        self.features = features\n        self.target = target\n        self.df = df\n        # It will start from 0 and increase by 1 until all of data is passed.\n        self.it = 0 \n        self.batch_size = batch_size\n        # np.ceil()은 a function that rounds up the float.\n        # Calculate the number of batches by dividing the data by the batch size\n        self.batches = int( np.ceil( len(df) / self.batch_size ) )\n        super().__init__()\n\n    def reset(self):\n        '''Reset the iterator'''\n        # If you need to perform iteration again from the beginning, \n        # you can initialize it with the reset() function.\n        self.it = 0\n\n    def next(self, input_data):\n        '''Yield next batch of data.'''\n        # self.batches defined at class instance creation. It contains the total number of batches that can be passed.\n        # End the iteration when self.it has reached the number of deliverable batches.\n        if self.it == self.batches:\n            # \n            return 0 \n        \n        # Get the start-end point for indexing.\n        a = self.it * self.batch_size\n        b = min( (self.it + 1) * self.batch_size, len(self.df) )\n        # Contain the data that will be passed by batch to the variable dt.\n        dt = cudf.DataFrame(self.df.iloc[a:b])\n        # Pass the feature and target from dt to input.\n        input_data(data=dt[self.features], label=dt[self.target]) \n        self.it += 1\n        return 1","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:02.080402Z","iopub.execute_input":"2022-07-17T15:12:02.080985Z","iopub.status.idle":"2022-07-17T15:12:02.097165Z","shell.execute_reply.started":"2022-07-17T15:12:02.08094Z","shell.execute_reply":"2022-07-17T15:12:02.096426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The function below is the logic to evaluate the model in the American Express - Default Prediction contest. It will be used as a criterion for optimization during model training.\n\nIn this tutorial, I expanded the original kernel that uses only CPUs for this calculation to select CPU and GPU. If you want to train the model using the train data on the GPU as it is, you must use the amex_metric_mod_GPU() function, which calculate optimization on the GPU. Here, we use amex_metric_mod_CPU().","metadata":{}},{"cell_type":"code","source":"def amex_metric_mod_CPU(y_true, y_pred):\n\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)\n\ndef amex_metric_mod_GPU(y_true, y_pred):\n\n    labels     = cupy.transpose(cupy.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = cupy.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[cupy.cumsum(weights) <= int(0.04 * cupy.sum(weights))]\n    top_four   = cupy.sum(cut_vals[:,0]) / cupy.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = cupy.transpose(cupy.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = cupy.where(labels[:,0]==0, 20, 1)\n        weight_random  = cupy.cumsum(weight / cupy.sum(weight))\n        total_pos      = cupy.sum(labels[:, 0] *  weight)\n        cum_pos_found  = cupy.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = cupy.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:04.090074Z","iopub.execute_input":"2022-07-17T15:12:04.090877Z","iopub.status.idle":"2022-07-17T15:12:04.106127Z","shell.execute_reply.started":"2022-07-17T15:12:04.090803Z","shell.execute_reply":"2022-07-17T15:12:04.105321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, we train the model, To make efficient use of the limited GPU, we will load the dataset down to the CPU first. Then, in the IterLoadForDMatrix() function, indexing in small units through cudf and loading memory to the GPU goes through.\n\nSince we have already loaded the training data into cudf and allocated it to the GPU memory, we will load it down to the CPU through the to_pandas() function. The to_pandas() function allows you to safely copy to CPU resources while transform cudf object into a pandas dataframe object.\n\nIn fact, if you use the RAPIDS platform, it is effective to perform all processes with a GPU, but there are many practical difficulties in applying it completely with a small amount of memory. Since we use Kaggle servers, it is important to make the best use of CPU and GPU memory like this.","metadata":{}},{"cell_type":"markdown","source":"To 'copy' means to remain in GPU memory. Let's check the resource status before and after copying together.","metadata":{}},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:05.985436Z","iopub.execute_input":"2022-07-17T15:12:05.985792Z","iopub.status.idle":"2022-07-17T15:12:06.713411Z","shell.execute_reply.started":"2022-07-17T15:12:05.985764Z","shell.execute_reply":"2022-07-17T15:12:06.712405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_cpu = train.to_pandas()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:07.574587Z","iopub.execute_input":"2022-07-17T15:12:07.57572Z","iopub.status.idle":"2022-07-17T15:12:11.795777Z","shell.execute_reply.started":"2022-07-17T15:12:07.575666Z","shell.execute_reply":"2022-07-17T15:12:11.794959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:11.797592Z","iopub.execute_input":"2022-07-17T15:12:11.798051Z","iopub.status.idle":"2022-07-17T15:12:12.55242Z","shell.execute_reply.started":"2022-07-17T15:12:11.798015Z","shell.execute_reply":"2022-07-17T15:12:12.551415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From now on, as multiple functions refer to multiple variables, and the allocated memory is called around, instantaneous memory usage will fluctuate. In Python, it is not necessary to manually manage memory because the garbage collector(which increases the available memory by deallocating the memory when the value allocated to a specific memory is not being referenced more than once) works internally.\n\nHowever, since optimization is not performed in real time for every execution, it is very helpful for optimization by manually operating the garbage collector through the gc.collect() function for every batch during model training and evaluation. So, let's run the function and move on.","metadata":{}},{"cell_type":"code","source":"import gc","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:12.554242Z","iopub.execute_input":"2022-07-17T15:12:12.554797Z","iopub.status.idle":"2022-07-17T15:12:12.558883Z","shell.execute_reply.started":"2022-07-17T15:12:12.554758Z","shell.execute_reply":"2022-07-17T15:12:12.558048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:12.560728Z","iopub.execute_input":"2022-07-17T15:12:12.561319Z","iopub.status.idle":"2022-07-17T15:12:12.73378Z","shell.execute_reply.started":"2022-07-17T15:12:12.561282Z","shell.execute_reply":"2022-07-17T15:12:12.732871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will set The K of KFold to 5. Then, for each of the 5 folds, we will have 4 learning folds and 1 validatio fold, and repeat learning a total of 5 times by moving the position of the valiation folds. As a result, the optimization is carried out through the average value of the 5 verification results.\n\nSEED can be specified with any number.","metadata":{}},{"cell_type":"code","source":"FOLDS = 5\nSEED = 42\nskf = KFold(n_splits=FOLDS, shuffle=True, random_state=SEED)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:14.935324Z","iopub.execute_input":"2022-07-17T15:12:14.937082Z","iopub.status.idle":"2022-07-17T15:12:14.945249Z","shell.execute_reply.started":"2022-07-17T15:12:14.937036Z","shell.execute_reply":"2022-07-17T15:12:14.944478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importances = []\noof = []\nTRAIN_SUBSAMPLE = 1.0\nVER = 1 # Version information to record when saving the model.\n\n# To perform cross-validation on the KFold object skf, iterative validation is performed after separating into training and validation folds.\nfor fold,(train_idx, valid_idx) in enumerate(skf.split(train_cpu, train_cpu.target)):\n    \n    # If you want to train model by only using some sample of the train data, set TRAIN_SUBSAMLE to less than 1.\n    # Then, you can configure the train set again by randomly extracting the ratio through the if statement below.\n    if TRAIN_SUBSAMPLE<1.0:\n        np.random.seed(SEED)\n        train_idx = np.random.choice(train_idx, int(len(train_idx)*TRAIN_SUBSAMPLE), replace=False)\n        np.random.seed(None)\n    \n    print('#'*25)\n    print('### Fold',fold+1)\n    print('### Train size',len(train_idx),'Valid size',len(valid_idx))\n    print(f'### Training with {int(TRAIN_SUBSAMPLE*100)}% fold data...')\n    print('#'*25)\n    \n    # Create the IterLoadForDMatrix instance. It will throw data in batches for training model. \n    Xy_train = IterLoadForDMatrix(train_cpu.loc[train_idx], FEATURES, 'target')\n    \n    # Define validation data to verify performance while training.\n    X_valid = train_cpu.loc[valid_idx, FEATURES]\n    y_valid = train_cpu.loc[valid_idx, 'target']\n    \n    dtrain = xgb.DeviceQuantileDMatrix(Xy_train, max_bin=256)\n    dvalid = xgb.DMatrix(data=X_valid, label=y_valid)\n    \n    # train the model\n    model = xgb.train(xgb_parms, \n                dtrain=dtrain,\n                evals=[(dtrain,'train'),(dvalid,'valid')],\n                num_boost_round=9999,\n                early_stopping_rounds=100,\n                verbose_eval=100) \n    # During cross-validation, the model is saved at each verification.\n    model.save_model(f'XGB_v{VER}_fold{fold}.xgb')\n    \n    # After training the model, we will check the feature importance. \n    # For this, importance is calculated for each training and stored in a variable dd.\n    dd = model.get_score(importance_type='weight')\n    df = pd.DataFrame({'feature':dd.keys(),f'importance_{fold}':dd.values()})\n    importances.append(df)\n    \n    # Validate the model. \n    oof_preds = model.predict(dvalid)\n    \n    # For calculating Accuracy, we uses the competition evaluation metric.\n    # In the original code, y_valied.values is put as it is,\n    # If not train_cpu but train that loaded in GPU memory was used as an argument,\n    # y_valid.values would be cupy._core.core.ndarray rather than np.ndarray.\n    # In this case, use cupy to let it compute directly on the GPU.\n    # -> acc = amex_metric_mod_GPU(y_valid.values, oof_preds)\n\n    # Since we use train_cpu loaded to the cpu as in the original code,\n    # The variable is np.ndarray(). So, we can just put it in amex_metric_mod_CPU().\n    acc = amex_metric_mod_CPU(y_valid.values, oof_preds)\n    print('Kaggle Metric =',acc,'\\n')\n    \n    # Also save the verification score(oof_pred) separately.\n    df = train_cpu.loc[valid_idx, ['customer_ID','target']].copy()\n    df['oof_pred'] = oof_preds\n    oof.append( df )\n    \n    # Let's free All variables used for training from memory.\n    del dtrain, Xy_train, dd, df\n    del X_valid, y_valid, dvalid, model\n    # And excecute garbage collection to completely remove the remaining non-referenced memory values even after the variable has been removed.\n    _ = gc.collect()\n    \nprint('#'*25)\n# When training is finished, the entire verification result is saved as a DataFrame.\n# Then, calculate the actual value and the verification value as an evaluation metric.\n# Here, if you use the train variable loaded in GPU memory as an argument like y_valid,\n# It would be not a pandas.core.frame.DataFrame, but a cudf.core.dataframe.DataFrame.\n# So, In this case, either merge them with cudf or replace all elements in oof with pandas's DataFrame objects.\n# Since we used train_cpu, we use pandas as is.\noof = pd.concat(oof,axis=0,ignore_index=True).set_index('customer_ID')\nacc = amex_metric_mod_GPU(oof.target.values, oof.oof_pred.values)\nprint('OVERALL CV Kaggle Metric =',acc)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-07-17T15:12:16.668349Z","iopub.execute_input":"2022-07-17T15:12:16.668702Z","iopub.status.idle":"2022-07-17T15:21:43.664459Z","shell.execute_reply.started":"2022-07-17T15:12:16.668674Z","shell.execute_reply":"2022-07-17T15:21:43.663557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now that training is over, the train_cpu dataset is no longer needed.\n# We will also clean up the memory referring to the train_cpu data. \ndel train, train_cpu\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:43.666493Z","iopub.execute_input":"2022-07-17T15:21:43.666902Z","iopub.status.idle":"2022-07-17T15:21:43.808528Z","shell.execute_reply.started":"2022-07-17T15:21:43.666863Z","shell.execute_reply":"2022-07-17T15:21:43.807331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:43.809943Z","iopub.execute_input":"2022-07-17T15:21:43.810975Z","iopub.status.idle":"2022-07-17T15:21:44.556624Z","shell.execute_reply.started":"2022-07-17T15:21:43.810936Z","shell.execute_reply":"2022-07-17T15:21:44.555604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Save OOF Preds","metadata":{}},{"cell_type":"markdown","source":"We have to save the prediction result with customer_ID. Just get the unique id information from the data file, change the hexadecimal number to an integer type as we did before, and merge the predicted values obtained during the training process.","metadata":{}},{"cell_type":"code","source":"TRAIN_PATH = '../input/amex-data-integer-dtypes-parquet-format/train.parquet'\noof_xgb = pd.read_parquet(TRAIN_PATH, columns=['customer_ID']).drop_duplicates()\noof_xgb['customer_ID_hash'] = oof_xgb['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')\noof_xgb = oof_xgb.set_index('customer_ID_hash')\noof_xgb = oof_xgb.merge(oof, left_index=True, right_index=True)\noof_xgb = oof_xgb.sort_index().reset_index(drop=True)\noof_xgb.to_csv(f'oof_xgb_v{VER}.csv',index=False)\noof_xgb.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:44.559029Z","iopub.execute_input":"2022-07-17T15:21:44.559779Z","iopub.status.idle":"2022-07-17T15:21:49.743266Z","shell.execute_reply.started":"2022-07-17T15:21:44.559734Z","shell.execute_reply":"2022-07-17T15:21:49.742467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Visualize the prediction results. The prediction result is a probability between 0 and 1. Since only 5% default(1) and the rest should be predicted as 0, the visualization is biased towards 0 and 1 in both directions, but there should be more values distributed at 0.","metadata":{}},{"cell_type":"code","source":"plt.hist(oof_xgb.oof_pred.values, bins=100)\nplt.title('OOF Predictions')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:49.744713Z","iopub.execute_input":"2022-07-17T15:21:49.745147Z","iopub.status.idle":"2022-07-17T15:21:50.10349Z","shell.execute_reply.started":"2022-07-17T15:21:49.745111Z","shell.execute_reply":"2022-07-17T15:21:50.102717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have the forecasts saved as a csv file, we remove the variables. The reason for repeating this process of saving to a file and freeing memory is that the RAM supported by Kaggle is not that large. In general, even in the local environment, RAM is not enough, so it may shut down while referencing or copying the memory. \n\nSo, to prevent this, it is recommended to save files to the hard-disk and keep the RAM lightly.","metadata":{}},{"cell_type":"code","source":"del oof_xgb, oof\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:50.104837Z","iopub.execute_input":"2022-07-17T15:21:50.105203Z","iopub.status.idle":"2022-07-17T15:21:50.27465Z","shell.execute_reply.started":"2022-07-17T15:21:50.105168Z","shell.execute_reply":"2022-07-17T15:21:50.27314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:50.276921Z","iopub.execute_input":"2022-07-17T15:21:50.277366Z","iopub.status.idle":"2022-07-17T15:21:51.074752Z","shell.execute_reply.started":"2022-07-17T15:21:50.277325Z","shell.execute_reply":"2022-07-17T15:21:51.073802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Feature Importance","metadata":{}},{"cell_type":"markdown","source":"Feature Importance is information about which variable is highly utilized in the model's task(prediction. here). \n\nAfter traininig the model, looking at this, if there are variables that are not important to the prediction, they will be removed, and the variables with excessive importance will go through a feedback process such as checking causality or correlation with the prediction target.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# While performing cross-validation, importances must have been accumulated as mush as the number of FOLDs(5).\n# So, Calculate the average importance per feature by merging into the df variable.\ndf = importances[0].copy()\nfor k in range(1,FOLDS):\n    df = df.merge(importances[k], on='feature', how='left')\ndf['importance'] = df.iloc[:,1:].mean(axis=1)\ndf = df.sort_values('importance',ascending=False)\ndf.to_csv(f'xgb_feature_importance_v{VER}.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.077352Z","iopub.execute_input":"2022-07-17T15:21:51.078026Z","iopub.status.idle":"2022-07-17T15:21:51.122603Z","shell.execute_reply.started":"2022-07-17T15:21:51.077985Z","shell.execute_reply":"2022-07-17T15:21:51.121766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.123801Z","iopub.execute_input":"2022-07-17T15:21:51.12421Z","iopub.status.idle":"2022-07-17T15:21:51.149812Z","shell.execute_reply.started":"2022-07-17T15:21:51.124172Z","shell.execute_reply":"2022-07-17T15:21:51.149088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Visualize only the top 20 features by importance with a bar chart.","metadata":{}},{"cell_type":"code","source":"NUM_FEATURES = 20\nplt.figure(figsize=(10,5*NUM_FEATURES//10))\nplt.barh(np.arange(NUM_FEATURES,0,-1), df.importance.values[:NUM_FEATURES])\nplt.yticks(np.arange(NUM_FEATURES,0,-1), df.feature.values[:NUM_FEATURES])\nplt.title(f'XGB Feature Importance - Top {NUM_FEATURES}')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.152665Z","iopub.execute_input":"2022-07-17T15:21:51.153012Z","iopub.status.idle":"2022-07-17T15:21:51.426444Z","shell.execute_reply.started":"2022-07-17T15:21:51.152985Z","shell.execute_reply":"2022-07-17T15:21:51.425687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Data Processing and Feature Engineering for Test Data","metadata":{}},{"cell_type":"markdown","source":"We wonder if a particular customer will repay the expense. So, let's get the unique customer_ID first.","metadata":{}},{"cell_type":"code","source":"TEST_PATH = '../input/amex-data-integer-dtypes-parquet-format/test.parquet'\ntest = read_file_CPU(path = TEST_PATH, usecols = ['customer_ID','S_2'])\ntest","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.427905Z","iopub.execute_input":"2022-07-17T15:21:51.428486Z","iopub.status.idle":"2022-07-17T15:23:29.377937Z","shell.execute_reply.started":"2022-07-17T15:21:51.428447Z","shell.execute_reply":"2022-07-17T15:23:29.377128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers = test[['customer_ID']].drop_duplicates().sort_index().values.flatten()\ncustomers","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:29.379397Z","iopub.execute_input":"2022-07-17T15:23:29.379978Z","iopub.status.idle":"2022-07-17T15:23:29.735532Z","shell.execute_reply.started":"2022-07-17T15:23:29.379941Z","shell.execute_reply":"2022-07-17T15:23:29.734691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this way, ID information is secured, and the default is predicted for each ID. Now we will load the test dataset. For now we needs only 2 columns, but like the train data, full dataset is very large. Therefore, we will load the dataset to the CPU in batch form as many as PART's, then pass each PART to the GPU so that the model can make predictions.\n\nTo do this, a criterion for dividing the PART is required. And the criterion requires two things: the row size in the test dataset before the duplicate customer_ID is removed, and the constant chunk size that after removing the duplicate.\n\nDo you remember the process_and_feature_engineer() function we created earlier? If the function passes, the duplicate of customer_ID is removed through group_by aggregation and it is replaced with statistical values. This is a method of performing prediction by passing the transformed dataset in chunk size.\n\nSo, let's create a function for this process and execute that.","metadata":{}},{"cell_type":"code","source":"def get_rows(customers, test, NUM_PARTS = 10, verbose = ''):\n    # One chunk(size of PART) can be obtained by dividing the size of the entire test-dataset by the number of PARTs.\n    # It is similar to finding the batch size when training the model.\n    chunk = len(customers)//NUM_PARTS\n    if verbose != '':\n        print(f'We will process {verbose} data as {NUM_PARTS} separate parts.')\n        print(f'There will be {chunk} customers in each part (except the last part).')\n        print('Below are number of rows in each part:')\n    rows = []\n\n    for k in range(NUM_PARTS):\n        # Keep the remain to cc if this PART is the last one.\n        if k==NUM_PARTS-1: \n            cc = customers[k*chunk:]\n        # If not the last PART, cut it in chunks from the front and put it in cc.\n        else: \n            cc = customers[k*chunk:(k+1)*chunk]\n        # Calculate the PART size by finding the number of customer_IDs included in the current PART.\n        s = test.loc[test.customer_ID.isin(cc)].shape[0]\n        # rows contain the size of 10 PARTs.\n        rows.append(s)\n    if verbose != '': print( rows )\n    return rows, chunk\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:29.737001Z","iopub.execute_input":"2022-07-17T15:23:29.737379Z","iopub.status.idle":"2022-07-17T15:23:29.745124Z","shell.execute_reply.started":"2022-07-17T15:23:29.737343Z","shell.execute_reply":"2022-07-17T15:23:29.74426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, with the function created above, we will divide customer_ID into a total of 10 groups(PARTs)","metadata":{"execution":{"iopub.status.busy":"2022-07-15T13:29:13.319054Z","iopub.execute_input":"2022-07-15T13:29:13.319608Z","iopub.status.idle":"2022-07-15T13:29:13.327749Z","shell.execute_reply.started":"2022-07-15T13:29:13.319567Z","shell.execute_reply":"2022-07-15T13:29:13.326253Z"}}},{"cell_type":"code","source":"# In the original code, the test dataset was divided into four.\n# But in case of using a kaggle server, you will get a GPU memory overflow error when you specify 4r, so we will divide it into 10.\nNUM_PARTS = 10\nrows,num_cust = get_rows(customers, test[['customer_ID']], NUM_PARTS = NUM_PARTS, verbose = 'test')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:29.746701Z","iopub.execute_input":"2022-07-17T15:23:29.747098Z","iopub.status.idle":"2022-07-17T15:23:30.991356Z","shell.execute_reply.started":"2022-07-17T15:23:29.747061Z","shell.execute_reply":"2022-07-17T15:23:30.990396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"On the last line, the size of each PART is printed. The size referred to here is the size including duplicates from the raw test dataset, not the chunk size.\n\nTherefore, it should be equal to the size of the test dataset when added together.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T13:34:33.019463Z","iopub.execute_input":"2022-07-15T13:34:33.019863Z","iopub.status.idle":"2022-07-15T13:34:33.026719Z","shell.execute_reply.started":"2022-07-15T13:34:33.019831Z","shell.execute_reply":"2022-07-15T13:34:33.025412Z"}}},{"cell_type":"code","source":"sum(rows) == len(test)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:30.992726Z","iopub.execute_input":"2022-07-17T15:23:30.993203Z","iopub.status.idle":"2022-07-17T15:23:30.999218Z","shell.execute_reply.started":"2022-07-17T15:23:30.993162Z","shell.execute_reply":"2022-07-17T15:23:30.998334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"and chunk size is,","metadata":{}},{"cell_type":"code","source":"num_cust","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:31.000686Z","iopub.execute_input":"2022-07-17T15:23:31.001303Z","iopub.status.idle":"2022-07-17T15:23:31.010702Z","shell.execute_reply.started":"2022-07-17T15:23:31.001267Z","shell.execute_reply":"2022-07-17T15:23:31.009824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Remove the loaded test data and deallocate memory.","metadata":{}},{"cell_type":"code","source":"del test\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:31.012462Z","iopub.execute_input":"2022-07-17T15:23:31.013193Z","iopub.status.idle":"2022-07-17T15:23:31.161038Z","shell.execute_reply.started":"2022-07-17T15:23:31.013099Z","shell.execute_reply":"2022-07-17T15:23:31.160076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:31.162702Z","iopub.execute_input":"2022-07-17T15:23:31.163083Z","iopub.status.idle":"2022-07-17T15:23:31.914814Z","shell.execute_reply.started":"2022-07-17T15:23:31.163048Z","shell.execute_reply":"2022-07-17T15:23:31.913827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. Infer Test","metadata":{}},{"cell_type":"markdown","source":"Through the code below, we divide the data into 10 PARTs through iteration, and then process the data for modeling and predict default using our model. The specific process is as follows.\n\n1. Devide the original test dataset into batches so that each PART can be included and load it to the CPU memory.\n2. The PART is loaded to the GPU for calculation again,\n3. Using the process_and_feature_engineering() function, convert raw data to statistical dataset for prediction.\n4. Then, the shape of dataset will be different. At this time, indexing can be performed with the chunk size obtained above.\n5. Index by chunk size and then predict default with the model.\n6. When prediction for all chunks is finished, merge and return the prediction result for the entire customer_ID.","metadata":{}},{"cell_type":"code","source":"skip_rows = 0\nskip_cust = 0\ntest_preds = []\n\n# Create an iteration object. By calling the object, you can raise as many rows as PART units to the CPU memory.\nTEST_PATH = '../input/amex-data-integer-dtypes-parquet-format/test.parquet'\nbatch = ParquetFile(TEST_PATH)\n# Since batch_size is fixed at one time call, it is not possible to get a different PARt size each time.\n# Because of this, the customer_ID that should be predicted in the current batch is possible to already be loaded in the previous batch.\n# So, set the batch size to the largest PART size, and merge it with the previous PART so that all customer_IDs can be indexed.\nprev_batch = pd.DataFrame()\nfor pres_batch, k in zip(batch.iter_batches(batch_size=max(rows)), range(NUM_PARTS)): \n    print(f'\\nReading test data...')\n    # Merge the dataset loaded from the previous batch and the current batch.\n    iter_batch = pd.concat([prev_batch, pres_batch.to_pandas()])\n    test_cpu = read_file_CPU(iter_batch = iter_batch, path = TEST_PATH)\n    # Then, Overwitten the previous batch with the current(present) batch.\n    # And clear the memory.\n    prev_batch = pres_batch.to_pandas()\n    del pres_batch, iter_batch\n    _ = gc.collect()\n    \n    # Load the PART into the GPU memory for pre-processing.\n    test_gpu = cudf.DataFrame(test_cpu)\n    skip_rows += rows[k]\n    print(f'=> Test part {k+1} has shape', test_gpu.shape)\n    \n    # With preprocessing, statistics are obtained for each customer_ID, and duplicate customer_IDs are removed.\n    # process_and_feature_engineer() uses cudf internally, not pandas.\n    # So, when preprocessing, the data must be placed on the GPU.\n    test_gpu = process_and_feature_engineer(test_gpu)\n    \n    # num_cust is the chunk size for de-duplicated customer_ID.\n    # We will do the credit default prediction for all customer_IDs by increasing the chunk size.\n    if k==NUM_PARTS-1: \n        test_gpu = test_gpu.loc[customers[skip_cust:]]\n    else: \n        test_gpu = test_gpu.loc[customers[skip_cust:skip_cust+num_cust]]\n    skip_cust += num_cust\n    print('shape after indexing(by chunk size)', test_gpu.shape)\n    \n    # Pass the features excluding the label from the test data to X_test.\n    X_test = test_gpu[FEATURES]\n    dtest = xgb.DMatrix(data=X_test)\n    del X_test\n    gc.collect()\n\n    # With our trained model, we can predict default.\n    # Our model was trained by cross-validation.\n    # Prediction is perfromed in the same way, and the prediction result can be obtained as the average value of all FOLD results.\n    model = xgb.Booster()\n    model.load_model(f'XGB_v{VER}_fold0.xgb')\n    preds = model.predict(dtest)\n    for f in range(1,FOLDS):\n        model.load_model(f'XGB_v{VER}_fold{f}.xgb')\n        preds += model.predict(dtest)\n    preds /= FOLDS\n    test_preds.append(preds)\n\n    # free the memory\n    del dtest, model\n    _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:28:45.079491Z","iopub.execute_input":"2022-07-17T15:28:45.080032Z","iopub.status.idle":"2022-07-17T15:34:25.417487Z","shell.execute_reply.started":"2022-07-17T15:28:45.079999Z","shell.execute_reply":"2022-07-17T15:34:25.416673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9. Create Submission CSV\n\nFinally, create a file for submission and visualize the test results.","metadata":{}},{"cell_type":"code","source":"len(test_preds)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:36:03.302426Z","iopub.execute_input":"2022-07-17T15:36:03.30283Z","iopub.status.idle":"2022-07-17T15:36:03.315063Z","shell.execute_reply.started":"2022-07-17T15:36:03.302799Z","shell.execute_reply":"2022-07-17T15:36:03.313723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# File for submission can be saved according to the format of the competition guide.\ntest_preds = np.concatenate(test_preds)\ntest = cudf.DataFrame(index=customers,data={'prediction':test_preds})\nsub = cudf.read_csv('../input/amex-default-prediction/sample_submission.csv')[['customer_ID']]\nsub['customer_ID_hash'] = sub['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\nsub = sub.set_index('customer_ID_hash')\nsub = sub.merge(test[['prediction']], left_index=True, right_index=True, how='left')\nsub = sub.reset_index(drop=True)\n\nsub.to_csv(f'submission_xgb_v{VER}.csv',index=False)\nprint('Submission file shape is', sub.shape )\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:34:25.424454Z","iopub.execute_input":"2022-07-17T15:34:25.426523Z","iopub.status.idle":"2022-07-17T15:34:26.533911Z","shell.execute_reply.started":"2022-07-17T15:34:25.426485Z","shell.execute_reply":"2022-07-17T15:34:26.533075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualize with hist plot.\nplt.hist(sub.to_pandas().prediction, bins=100)\nplt.title('Test Predictions')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:34:31.686057Z","iopub.execute_input":"2022-07-17T15:34:31.686514Z","iopub.status.idle":"2022-07-17T15:34:32.81205Z","shell.execute_reply.started":"2022-07-17T15:34:31.686474Z","shell.execute_reply":"2022-07-17T15:34:32.811173Z"},"trusted":true},"execution_count":null,"outputs":[]}]}