{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"> 본 커널은 [XGBoost Starter - LB 0.793](https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793)에 작성된 코드를 기반으로 한국어 설명을 추가한 XGBoost 튜토리얼입니다.","metadata":{}},{"cell_type":"markdown","source":"> TOC\n```\n1. Load Libraries\n2. Load Dataset and Manage the GPU Memory\n3. Feature Engineering\n4. Train XGB\n5. Save OOF Preds\n6. Feature Importance\n7. Data Processing and Feature Engineering for Test Data\n8. Infer Test\n9. Create Submission CSV\n```","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:42:35.078701Z","iopub.execute_input":"2022-07-17T15:42:35.079157Z","iopub.status.idle":"2022-07-17T15:42:35.087982Z","shell.execute_reply.started":"2022-07-17T15:42:35.079116Z","shell.execute_reply":"2022-07-17T15:42:35.086913Z"}}},{"cell_type":"markdown","source":"# 1. Load Libraries","metadata":{}},{"cell_type":"markdown","source":"우리가 머신러닝, 딥러닝을 학습시킬 때 CPU를 사용하여 데이터 전처리 등을 수행하고 GPU로 모델 학습을 하는 것이 일반적입니다. 그리고 이 때는 CPU에 올라간 데이터를 GPU로 복사(이동)하는 과정이 필요합니다. pandas로 데이터프레임을 다루고 torch로 GPU 메모리 상의 데이터를 처리하는 식이죠.\n\n이제는 그러지 말고, '전체 과정을 모두 GPU 위에서 진행하자' 라는 컨셉으로 나온게 RAPIDS입니다. RAPIDS는 엔비디아가 주도하여 구축하고 운영하고 있는 CUDA 프로세스 기반 데이터사이언스 플랫폼입니다.\n\nCUDA를 기반으로 한다는 것을 강조하듯 패키지 이름도 모두 cuxx로 지었습니다. pandas는 cudf로, numpy는 cupy로, sklearn은 cuml로 대체하여 거의 기존과 동일한 함수를 사용할 수 있게 개발되어 있습니다.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport cupy\nimport cudf\n\nimport matplotlib.pyplot as plt, gc, os\n\nprint('cudf version',cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:11.091036Z","iopub.execute_input":"2022-07-17T15:10:11.091410Z","iopub.status.idle":"2022-07-17T15:10:12.511764Z","shell.execute_reply.started":"2022-07-17T15:10:11.091379Z","shell.execute_reply":"2022-07-17T15:10:12.510966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Load Dataset and Manage the GPU Memory\n\ngpu를 사용할 때에는 gpu 자원을 잘 관리해줘야 합니다. nvidia-smi 쉘 명령을 통해 현재 가용가능한 GPU 리소스를 모니터링해줍니다.","metadata":{}},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:15.898746Z","iopub.execute_input":"2022-07-17T15:10:15.899407Z","iopub.status.idle":"2022-07-17T15:10:16.609671Z","shell.execute_reply.started":"2022-07-17T15:10:15.899371Z","shell.execute_reply":"2022-07-17T15:10:16.608704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tesla P100 1대를 사용할 수 있고, 현재 해당 GPU로 실행중인 프로세스는 없습니다.\n\n이제 cudf로 GPU 메모리 위에 데이터를 올려볼 것입니다. 기본적인 함수나 사용법은 pandas와 동일합니다. ","metadata":{}},{"cell_type":"code","source":"# read_parquet() 함수는 parquet 형식의 파일을 읽어옵니다.\ndf = cudf.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:17.630807Z","iopub.execute_input":"2022-07-17T15:10:17.631205Z","iopub.status.idle":"2022-07-17T15:10:39.510018Z","shell.execute_reply.started":"2022-07-17T15:10:17.631172Z","shell.execute_reply":"2022-07-17T15:10:39.509199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:39.511671Z","iopub.execute_input":"2022-07-17T15:10:39.512042Z","iopub.status.idle":"2022-07-17T15:10:39.830580Z","shell.execute_reply.started":"2022-07-17T15:10:39.512014Z","shell.execute_reply":"2022-07-17T15:10:39.829831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:39.831740Z","iopub.execute_input":"2022-07-17T15:10:39.832581Z","iopub.status.idle":"2022-07-17T15:10:40.565054Z","shell.execute_reply.started":"2022-07-17T15:10:39.832541Z","shell.execute_reply":"2022-07-17T15:10:40.563920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"이렇게 데이터를 GPU에 올려 놓으면 Memory-Usage에 3669MiB만큼의 메모리를 차지하게 됩니다. 전체 가용 메모리가 16280MiB임을 감안했을 때,\n해당 변수가 몇번 복사되거나 다른 데이터셋(예를들면 평가데이터)을 불러오면 메모리 Over될 수 있다는 것을 염두해두어야 합니다.","metadata":{"execution":{"iopub.status.busy":"2022-07-16T01:00:25.059036Z","iopub.execute_input":"2022-07-16T01:00:25.059590Z","iopub.status.idle":"2022-07-16T01:00:25.066313Z","shell.execute_reply.started":"2022-07-16T01:00:25.059548Z","shell.execute_reply":"2022-07-16T01:00:25.065190Z"}}},{"cell_type":"markdown","source":"다시 이어서 `customer_ID`를 숫자로 깔끔하게 변환해주고 싶다면, hex_to_int() 함수를 사용하고 astype()을 통해 처리할 수 있습니다.\n여기서는 마지막 16문자에 대해서 해당 처리를 수행합니다.","metadata":{}},{"cell_type":"code","source":"df['customer_ID'] = df['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:40.567384Z","iopub.execute_input":"2022-07-17T15:10:40.567780Z","iopub.status.idle":"2022-07-17T15:10:41.059191Z","shell.execute_reply.started":"2022-07-17T15:10:40.567735Z","shell.execute_reply":"2022-07-17T15:10:41.058431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:41.063298Z","iopub.execute_input":"2022-07-17T15:10:41.065568Z","iopub.status.idle":"2022-07-17T15:10:41.932017Z","shell.execute_reply.started":"2022-07-17T15:10:41.065530Z","shell.execute_reply":"2022-07-17T15:10:41.931034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"몸에 익히기 위해 조금 과할만큼 계속 메모리 추적을 하겠습니다. 데이터 일부를 이렇게 잘라주는 것만으로도 사용중인 메모리를 줄일 수 있습니다. \n약 300MiB나 줄였네요.","metadata":{}},{"cell_type":"markdown","source":"`S_2`열은 시간 정보를 담고 있습니다. pandas와 마찬가지로 to_datetime()함수로 데이터 타입을 변경해줍니다.","metadata":{}},{"cell_type":"code","source":"df['S_2'] = cudf.to_datetime(df['S_2'])\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:41.935391Z","iopub.execute_input":"2022-07-17T15:10:41.935691Z","iopub.status.idle":"2022-07-17T15:10:42.218694Z","shell.execute_reply.started":"2022-07-17T15:10:41.935662Z","shell.execute_reply":"2022-07-17T15:10:42.217928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:42.220043Z","iopub.execute_input":"2022-07-17T15:10:42.220564Z","iopub.status.idle":"2022-07-17T15:10:42.946499Z","shell.execute_reply.started":"2022-07-17T15:10:42.220524Z","shell.execute_reply":"2022-07-17T15:10:42.945541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"적절한 데이터 타입으로 변경해주는 것도 중요합니다. 파이썬의 메모리 운용 프로세스상 문자열 타입은 기본적으로 메모리 사용량이 많습니다. \n그래서 카테고리형 변수를 원한인코딩해주거나 정수형 등으로 변경할 수 있다면 변경해주는 것이 메모리 관리에 유리합니다.","metadata":{}},{"cell_type":"code","source":"df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:42.947881Z","iopub.execute_input":"2022-07-17T15:10:42.948189Z","iopub.status.idle":"2022-07-17T15:10:43.197592Z","shell.execute_reply.started":"2022-07-17T15:10:42.948160Z","shell.execute_reply":"2022-07-17T15:10:43.196689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"빈 셀이 매우 많습니다. fillna() 함수를 활용해 빈 셀은 -127로 채워주겠습니다. 참고로 1byte(8bit)로 표현가능한 가장 낮은 숫자는 -128입니다. signed integer(부호가 있는 정수)는 -128부터 127까지 표현가능하며, 원문 저자는 최소한의 메모리 사용량으로 null값을 대체하길 원했던 것 같습니다. \n다만, 변수별 분포를 고려하지 않은 처리 형태이므로 모델 학습에는 다소 불리한 부분이 생길 수 있습니다.\n\n그리고 원문에서는 cudf로 fillna()함수를 실행할때 `df = df.fillna(NAN_VALUE)`와 같이 작성했으나, 우리는 `df.fillna(NAN_VALUE, inplace=True)`로 작성해주겠습니다. 원문과 같이 작성하는 경우 null이 채워진 df가 새로운 메모리 주소에 복사되어 할당되면서 GPU 메모리 사용량은 2배(3321MiB -> 6079MiB) 가까이 증가합니다. 우리는 충분한 메모리 환경 위에서 작업하고 있지 않기 때문에 반드시 inplace=True 옵션을 주고, 현재 할당된 메모리주소에 덮어쓰도록 해줘야 합니다.","metadata":{}},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:43.200149Z","iopub.execute_input":"2022-07-17T15:10:43.200767Z","iopub.status.idle":"2022-07-17T15:10:43.210091Z","shell.execute_reply.started":"2022-07-17T15:10:43.200727Z","shell.execute_reply":"2022-07-17T15:10:43.209170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"NAN_VALUE = -127 \ndf.fillna(NAN_VALUE, inplace=True)\ndf.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:43.212923Z","iopub.execute_input":"2022-07-17T15:10:43.213378Z","iopub.status.idle":"2022-07-17T15:10:43.717726Z","shell.execute_reply.started":"2022-07-17T15:10:43.213332Z","shell.execute_reply":"2022-07-17T15:10:43.716838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:10:43.718970Z","iopub.execute_input":"2022-07-17T15:10:43.719405Z","iopub.status.idle":"2022-07-17T15:10:44.449065Z","shell.execute_reply.started":"2022-07-17T15:10:43.719366Z","shell.execute_reply":"2022-07-17T15:10:44.448085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"아래 read_file() 함수는 이 과정을 한번에 수행해줍니다. \n\n원문에서는 이렇게 GPU를 사용하는 방식으로만 작성되어 있으나 캐글 서버로 진행하기는 무리가 있습니다. 따라서 우리는 전체 데이터를 로드하는 과정에서는 CPU를 사용하고, 데이터를 학습 혹은 평가하는 과정에서 iteration 형태로 데이터를 쪼개서 GPU로 올려줄 것이기 때문에 CPU를 사용하는 read_file_CPU() 함수를 하나 더 만들어주겠습니다. CPU로 데이터를 올린다면 pandas를 사용하면 되고 GPU로 올린다면 cudf를 사용하면 됩니다. \n\n추가로, CPU로 데이터를 올리는 과정에서도 batch로 나눠 올려줘야 합니다. 추후 test 데이터셋을 로드하는 과정에서 한번에 데이터를 올리게 되면 CPU 메모리를 초과해 캐글 커널이 리셋됩니다.","metadata":{}},{"cell_type":"code","source":"from pyarrow.parquet import ParquetFile","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:19.901369Z","iopub.execute_input":"2022-07-17T15:11:19.902387Z","iopub.status.idle":"2022-07-17T15:11:19.907604Z","shell.execute_reply.started":"2022-07-17T15:11:19.902325Z","shell.execute_reply":"2022-07-17T15:11:19.906673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 원문 코드\ndef read_file_GPU(path = '', usecols = None):\n    # read_parquet() 함수는 parquet 형식의 파일을 읽어옵니다.\n    # 이 때, 컬럼을 지정해주고 싶다면\n    if usecols is not None: \n        df = cudf.read_parquet(path, columns=usecols)\n    # columns를 지정하지 않고 그대로 불러온다면,\n    else: df = cudf.read_parquet(path)\n    \n    df['customer_ID'] = df['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\n    df.S_2 = cudf.to_datetime( df.S_2 )\n    df.fillna(NAN_VALUE, inplace=True) \n    print('shape of data:', df.shape)\n    \n    return df\n\n# 수정된 코드(CPU, batch load)\ndef read_file_CPU(path = '', iter_batch = None, usecols = None):\n    if usecols is not None:\n        # 일부 컬럼(1~2개)만 가져올 때는 전체 row를 다 불러와도 문제가 생기지 않습니다.\n        df = pd.read_parquet(path, columns=usecols)\n    else:\n        # 전체 컬럼을 가져올 때는 batch 형태로 데이터를 가져옵니다.\n        df = iter_batch\n    \n    # cudf로 했던 것과 동일한 처리를 수행합니다.\n    df['customer_ID'] = df['customer_ID'].apply(lambda x : int(x[-16:],16)).astype('int64') \n    df.S_2 = pd.to_datetime( df.S_2 )\n    df.fillna(NAN_VALUE, inplace=True)\n    print('shape of data:', df.shape)\n    \n    return df\n\n# print('Reading train data...')\n# TRAIN_PATH = '../input/amex-data-integer-dtypes-parquet-format/train.parquet'\n# train = read_file(path = TRAIN_PATH)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:22.391593Z","iopub.execute_input":"2022-07-17T15:11:22.392020Z","iopub.status.idle":"2022-07-17T15:11:22.402457Z","shell.execute_reply.started":"2022-07-17T15:11:22.391986Z","shell.execute_reply":"2022-07-17T15:11:22.401622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"원본 커널에서는 df가 아닌 train 변수명을 사용했습니다. 동일하게 맞춰주겠습니다.\n\n단, 여기서 메모리 관리를 위해 copy()함수를 사용하지 않습니다. copy()함수를 사용하면 다른 메모리주소에 복사가 되어버립니다.","metadata":{}},{"cell_type":"code","source":"train = df\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:30.826951Z","iopub.execute_input":"2022-07-17T15:11:30.827316Z","iopub.status.idle":"2022-07-17T15:11:31.045376Z","shell.execute_reply.started":"2022-07-17T15:11:30.827286Z","shell.execute_reply":"2022-07-17T15:11:31.044469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:34.351396Z","iopub.execute_input":"2022-07-17T15:11:34.351787Z","iopub.status.idle":"2022-07-17T15:11:35.104280Z","shell.execute_reply.started":"2022-07-17T15:11:34.351753Z","shell.execute_reply":"2022-07-17T15:11:35.103239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"복사하지 않고 변수명만 변경(메모리 참조)했기 때문에 추가 메모리를 사용하지 않은 것을 확인할 수 있습니다.","metadata":{}},{"cell_type":"markdown","source":"# 3. Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"우리는 모델 학습에 데이터를 그대로 넣지 않고 통계값으로 변환 처리를 한 다음 학습시킬 것입니다.\n데이터를 집계하는 과정에서 컬럼이 multiindex 형식을 띠게 되는데, 모델에 넣기 위해서는 1차원 형태 컬럼을 유지해야 합니다.\n이 작업만 먼저 살펴보고 함수를 통해 전체 프로세스를 실행하겠습니다.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T10:45:36.492633Z","iopub.execute_input":"2022-07-15T10:45:36.492999Z","iopub.status.idle":"2022-07-15T10:45:36.498767Z","shell.execute_reply.started":"2022-07-15T10:45:36.492968Z","shell.execute_reply":"2022-07-15T10:45:36.497787Z"}}},{"cell_type":"code","source":"multi_index_col_sample = train.groupby('customer_ID')[['B_30','B_38','D_114']].agg(['count','last','nunique']).columns\nmulti_index_col_sample","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:37.574266Z","iopub.execute_input":"2022-07-17T15:11:37.574645Z","iopub.status.idle":"2022-07-17T15:11:37.705430Z","shell.execute_reply.started":"2022-07-17T15:11:37.574614Z","shell.execute_reply":"2022-07-17T15:11:37.704690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"['_'.join(x) for x in multi_index_col_sample]","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:39.168928Z","iopub.execute_input":"2022-07-17T15:11:39.169292Z","iopub.status.idle":"2022-07-17T15:11:39.175220Z","shell.execute_reply.started":"2022-07-17T15:11:39.169262Z","shell.execute_reply":"2022-07-17T15:11:39.174351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"이렇게 이중 인덱스를 1차원 인덱스로 붙여서 보기 쉽게 만들어줄 수 있습니다. 이어서 전체 함수를 확인합니다.","metadata":{}},{"cell_type":"code","source":"def process_and_feature_engineer(df):\n    # list comprehension 방식으로 customer_ID 컬럼과 S_2 컬럼을 제외한 나머지 컬럼명을 all_cols 변수에 담아줍니다. \n    # 해당 컬럼들은 인덱스와 ID값이 아닌 분석 대상 데이터를 가지고 있습니다.\n    all_cols = [c for c in list(df.columns) if c not in ['customer_ID','S_2']]\n    \n    # all_cols 중에서 카테고리형 변수와 수치형 변수를 나눠줍니다.\n    # 원본 커널에서 카테고리형 변수는 아래 11개 컬럼만 지정하고 있으나, 더 많은 컬럼이 있는 것으로 보입니다.\n    # 하지만 연속성을 유지하고 혼선을 방지하기 위해 그대로 사용하도록 하겠습니다.\n    cat_features = [\"B_30\",\"B_38\",\"D_114\",\"D_116\",\"D_117\",\"D_120\",\"D_126\",\"D_63\",\"D_64\",\"D_66\",\"D_68\"]\n    num_features = [col for col in all_cols if col not in cat_features]\n\n    # 각 customer_ID에 대해 수치형 변수들을 통계치로 집계해줍니다.\n    test_num_agg = df.groupby(\"customer_ID\")[num_features].agg(['mean', 'std', 'min', 'max', 'last'])\n    # 집계한 컬럼 명은 MultiIndex 타입입니다. '_' 문자로 컬럼명을 이어줍니다.\n    test_num_agg.columns = ['_'.join(x) for x in test_num_agg.columns]\n\n    # 각 customer_ID에 대해 카테고리형 변수를 집계해줍니다.\n    # count는 동일한 customer_ID가 몇번 등장하는지, last는 해당 customer_ID의 각 카테고리형 변수 중 가장 최근 값이 무엇인지 보여줍니다.\n    # nunique()는 각 카테고리형 변수에 등장하는 유일한 값을 세어줍니다.\n    test_cat_agg = df.groupby(\"customer_ID\")[cat_features].agg(['count', 'last', 'nunique'])\n    test_cat_agg.columns = ['_'.join(x) for x in test_cat_agg.columns]\n\n    # 기존 데이터와 통계치를 구한 값을 모두 병합해줍니다.\n    df = cudf.concat([test_num_agg, test_cat_agg], axis=1)\n    del test_num_agg, test_cat_agg\n    print('shape after engineering', df.shape)\n    \n    return df\n\n\ntrain = process_and_feature_engineer(train)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:41.268955Z","iopub.execute_input":"2022-07-17T15:11:41.269624Z","iopub.status.idle":"2022-07-17T15:11:42.444693Z","shell.execute_reply.started":"2022-07-17T15:11:41.269590Z","shell.execute_reply":"2022-07-17T15:11:42.443753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:43.445200Z","iopub.execute_input":"2022-07-17T15:11:43.445883Z","iopub.status.idle":"2022-07-17T15:11:44.275384Z","shell.execute_reply.started":"2022-07-17T15:11:43.445825Z","shell.execute_reply":"2022-07-17T15:11:44.274616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:45.052505Z","iopub.execute_input":"2022-07-17T15:11:45.053165Z","iopub.status.idle":"2022-07-17T15:11:45.940463Z","shell.execute_reply.started":"2022-07-17T15:11:45.053122Z","shell.execute_reply":"2022-07-17T15:11:45.939465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"컬럼이 많아지면서 데이터 용량도 늘어났습니다. ","metadata":{}},{"cell_type":"markdown","source":"데이터셋은 타겟(label) 데이터도 제공하고 있습니다. 위에서 생성한 train 데이터셋에 병합해주겠습니다.\n이 때, customer_ID는 앞에서 했던 것과 동일한 방식으로 정수형 처리를 해줘야 합니다. 본 테스크는 해당 customer가 대출을 갚을 것인가, 갚지 못할 것인가를 예측하는 과제입니다. customer 마다의 대출 상환 여부가 train_labels.csv에 담겨있습니다.","metadata":{}},{"cell_type":"code","source":"targets = cudf.read_csv('../input/amex-default-prediction/train_labels.csv')\ntargets['customer_ID'] = targets['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\ntargets = targets.set_index('customer_ID')\n\n# 인덱스는 customer_ID로 동일합니다. 해당 인덱스를 기준으로 merge합니다.\ntrain = train.merge(targets, left_index=True, right_index=True, how='left')\n# 타겟 데이터는 8비트 정수형으로 담아줍니다.\ntrain.target = train.target.astype('int8')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:47.740869Z","iopub.execute_input":"2022-07-17T15:11:47.741264Z","iopub.status.idle":"2022-07-17T15:11:48.349517Z","shell.execute_reply.started":"2022-07-17T15:11:47.741232Z","shell.execute_reply":"2022-07-17T15:11:48.348642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = train.reset_index()\ntrain","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:49.931144Z","iopub.execute_input":"2022-07-17T15:11:49.931965Z","iopub.status.idle":"2022-07-17T15:11:51.098071Z","shell.execute_reply.started":"2022-07-17T15:11:49.931933Z","shell.execute_reply":"2022-07-17T15:11:51.097280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:52.002889Z","iopub.execute_input":"2022-07-17T15:11:52.003398Z","iopub.status.idle":"2022-07-17T15:11:52.726827Z","shell.execute_reply.started":"2022-07-17T15:11:52.003365Z","shell.execute_reply":"2022-07-17T15:11:52.725814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"targets는 이제 train데이터셋에 병합되었으므로 메모리 상에서 할당 해제해줍니다.  \n용량은 얼마 되지 않지만 변수 사용이 끝나면 메모리 관리를 항상 해주는 편이 좋습니다.","metadata":{}},{"cell_type":"code","source":"del targets","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:53.851445Z","iopub.execute_input":"2022-07-17T15:11:53.852014Z","iopub.status.idle":"2022-07-17T15:11:53.857696Z","shell.execute_reply.started":"2022-07-17T15:11:53.851976Z","shell.execute_reply":"2022-07-17T15:11:53.856865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:55.184144Z","iopub.execute_input":"2022-07-17T15:11:55.184725Z","iopub.status.idle":"2022-07-17T15:11:56.096290Z","shell.execute_reply.started":"2022-07-17T15:11:55.184683Z","shell.execute_reply":"2022-07-17T15:11:56.095040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"모델에 학습시킬 피처 수를 세어봅니다. 첫번째 열은 id값이고 마지막 열은 label입니다. 2개 열을 제외한 나머지 열의 수는 198개입니다.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T11:15:43.092352Z","iopub.execute_input":"2022-07-15T11:15:43.092723Z","iopub.status.idle":"2022-07-15T11:15:43.098262Z","shell.execute_reply.started":"2022-07-15T11:15:43.092695Z","shell.execute_reply":"2022-07-15T11:15:43.097264Z"}}},{"cell_type":"code","source":"FEATURES = train.columns[1:-1]\nprint(f'There are {len(FEATURES)} features!')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:11:58.336092Z","iopub.execute_input":"2022-07-17T15:11:58.337168Z","iopub.status.idle":"2022-07-17T15:11:58.344696Z","shell.execute_reply.started":"2022-07-17T15:11:58.337117Z","shell.execute_reply":"2022-07-17T15:11:58.343892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Train XGB\n모델을 학습할 때는 KFold 교차검증을 활용해 모든 데이터를 최소 1회 이상 학습에 사용하면서 과적합을 막을 것입니다.\n일반적으로 편의상 사용되는 70:30 혹은 80:20 등으로 데이터를 학습/검증 용으로 나눠 학습시키는 방식은 30, 20에 해당하는 부분은 학습에 활용할 수 없다는 문제가 있습니다. KFold는 학습데이터와 검증데이터를 쪼개서(K개의 Fold를 만든다고 표현합니다) 모든 데이터셋이 학습에 활용될 수 있도록 합니다.","metadata":{}},{"cell_type":"code","source":"# LOAD XGB LIBRARY\nfrom sklearn.model_selection import KFold\nimport xgboost as xgb\nprint('XGB Version',xgb.__version__)\n\n# XGB MODEL PARAMETERS\nxgb_parms = { \n    'max_depth':4, \n    'learning_rate':0.05, \n    'subsample':0.8,\n    'colsample_bytree':0.6, \n    'eval_metric':'logloss',\n    'objective':'binary:logistic',\n    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n    'random_state':42\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:00.325200Z","iopub.execute_input":"2022-07-17T15:12:00.325650Z","iopub.status.idle":"2022-07-17T15:12:00.445936Z","shell.execute_reply.started":"2022-07-17T15:12:00.325612Z","shell.execute_reply":"2022-07-17T15:12:00.445158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"학습할 때, DeviceQuantileDMatrix를 사용합니다. GPU는 기본적으로 한번 사용시 모든 메모리가 일시에 연산에 사용되는데, 해당 함수를 사용하면 \nGPU 메모리를 작은 단위로 분할해서 사용하게 해줍니다. \n\nDeviceQuantileDMatrix를 사용하기 위해서는 iteration 방식, 즉 next() 함수를 반복 호출하며 배치 형태로 데이터를 넘겨줄 수 있는 클래스를 정의해줘야 합니다.","metadata":{}},{"cell_type":"code","source":"class IterLoadForDMatrix(xgb.core.DataIter):\n    def __init__(self, df=None, features=None, target=None, batch_size=256*1024):\n        self.features = features\n        self.target = target\n        self.df = df\n        # 0번부터 시작해 데이터를 모두 넘겨줄때까지 1씩 올려줄 것입니다.\n        self.it = 0 \n        self.batch_size = batch_size\n        # np.ceil()은 '올림'을 해주는 함수입니다. \n        # 데이터를 배치 사이즈로 나눠 배치 크기(갯수)를 계산합니다.\n        self.batches = int( np.ceil( len(df) / self.batch_size ) )\n        super().__init__()\n\n    def reset(self):\n        '''Reset the iterator'''\n        # iteration을 처음부터 다시 실시해줘야 한다면 reset() 함수로 초기화할 수 있습니다.\n        self.it = 0\n\n    def next(self, input_data):\n        '''Yield next batch of data.'''\n        # 클래스 인스턴스 생성시 self.batches가 정의되었습니다. 전달할 수 있는 총 배치 수를 담고 있습니다.\n        # self.it이 전달 가능한 배치 수에 도달했을 때, iteration을 종료합니다.\n        if self.it == self.batches:\n            # \n            return 0 \n        \n        # 인덱싱을 위한 start-end point를 구합니다.\n        a = self.it * self.batch_size\n        b = min( (self.it + 1) * self.batch_size, len(self.df) )\n        # 배치로 전달할 데이터를 인덱싱해서 dt 변수에 담아줍니다.\n        dt = cudf.DataFrame(self.df.iloc[a:b])\n        # dt로 받은 데이터에서 피처와 타겟을 넘겨줍니다.\n        input_data(data=dt[self.features], label=dt[self.target]) #, weight=dt['weight'])\n        \n        # iter를 1씩 올려주면서 다음 배치를 수행하기 위한 계산입니다.\n        self.it += 1\n        return 1","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:02.080402Z","iopub.execute_input":"2022-07-17T15:12:02.080985Z","iopub.status.idle":"2022-07-17T15:12:02.097165Z","shell.execute_reply.started":"2022-07-17T15:12:02.080940Z","shell.execute_reply":"2022-07-17T15:12:02.096426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"아래 함수는 본 데이터를 공유한 [American Express - Default Prediction](https://www.kaggle.com/competitions/amex-default-prediction) 대회에서 모델을 평가하는 로직입니다.([원문](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327534))\n\n모델 학습 과정에서 최적화하는 기준으로 사용하게 됩니다.\n\n본 튜토리얼에서는 CPU만 사용하는 기존 원문에서 CPU, GPU를 선택해 사용할 수 있도록 확장했습니다. 만약 GPU에 올라가있는 train 변수를 그대로 사용해 모델을 학습시키고자 한다면 반드시 GPU 상에서 연산을 수행해주는 amex_metric_mod_GPU() 함수를 사용해야 합니다. 여기서는 train 변수를 CPU로 내린 다음 학습시킬 것이므로 amex_metric_mod_CPU()를 사용합니다.\n\n","metadata":{}},{"cell_type":"code","source":"def amex_metric_mod_CPU(y_true, y_pred):\n\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)\n\ndef amex_metric_mod_GPU(y_true, y_pred):\n\n    labels     = cupy.transpose(cupy.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = cupy.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[cupy.cumsum(weights) <= int(0.04 * cupy.sum(weights))]\n    top_four   = cupy.sum(cut_vals[:,0]) / cupy.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = cupy.transpose(cupy.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = cupy.where(labels[:,0]==0, 20, 1)\n        weight_random  = cupy.cumsum(weight / cupy.sum(weight))\n        total_pos      = cupy.sum(labels[:, 0] *  weight)\n        cum_pos_found  = cupy.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = cupy.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:04.090074Z","iopub.execute_input":"2022-07-17T15:12:04.090877Z","iopub.status.idle":"2022-07-17T15:12:04.106127Z","shell.execute_reply.started":"2022-07-17T15:12:04.090803Z","shell.execute_reply":"2022-07-17T15:12:04.105321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"드디어 모델을 학습시킵니다. 제한된 GPU를 효율적으로 사용하기 위해 데이터셋을 CPU로 먼저 내려줄 것입니다. 그 다음, IterLoadForDMatrix() 함수에서 cudf를 통해 작은 단위로 인덱싱하여 GPU로 메모리를 올려주는 방식을 거치게 됩니다.  \n\n우리는 학습데이터를 이미 cudf로 불러와서 GPU 메모리에 할당시켰기 때문에 to_pandas() 함수를 통해 CPU로 내려주는 작업을 해주겠습니다.  to_pandas() 함수를 사용하면 pandas dataframe 객체로 들어가면서 CPU 자원으로 안전하게 복사할 수 있습니다.\n\n사실, RAPIDS 플랫폼을 사용한다면 모든 과정을 GPU로 수행하는 것이 효과적이지만 적은 메모리 가용량으로 온전하게 적용하기는 현실적으로 어려운 점이 많습니다. 우리는 캐글 서버를 사용하는 만큼 이렇게 CPU와 GPU 메모리를 최대한 잘 활용하는 것이 중요합니다.","metadata":{}},{"cell_type":"markdown","source":"'복사'한다는 것은 GPU 메모리에도 그대로 남아있다는 것을 의미합니다. 복사 전후의 리소스 상태를 함께 확인하겠습니다.","metadata":{}},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:05.985436Z","iopub.execute_input":"2022-07-17T15:12:05.985792Z","iopub.status.idle":"2022-07-17T15:12:06.713411Z","shell.execute_reply.started":"2022-07-17T15:12:05.985764Z","shell.execute_reply":"2022-07-17T15:12:06.712405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_cpu = train.to_pandas()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:07.574587Z","iopub.execute_input":"2022-07-17T15:12:07.575720Z","iopub.status.idle":"2022-07-17T15:12:11.795777Z","shell.execute_reply.started":"2022-07-17T15:12:07.575666Z","shell.execute_reply":"2022-07-17T15:12:11.794959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:11.797592Z","iopub.execute_input":"2022-07-17T15:12:11.798051Z","iopub.status.idle":"2022-07-17T15:12:12.552420Z","shell.execute_reply.started":"2022-07-17T15:12:11.798015Z","shell.execute_reply":"2022-07-17T15:12:12.551415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"이제부터 여러 함수가 여러 변수를 참조하고, 할당한 메모리가 이리저리 호출되면서 순간적인 메모리 사용량이 널뛰게 됩니다. 파이썬에는 내부적으로 가비지 콜렉터(특정 메모리에 할당된 값이 1회 이상 참조되고 있지 않은 경우 해당 메모리를 할당 해제해줌으로써 가용 메모리를 늘려주는 역할을 합니다.)가 돌아가기 때문에 수동으로 메모리 관리를 해줄 필요는 없습니다. 하지만 그렇다고 매 실행마다 실시간으로 최적화를 진행해주는 것은 아니기 때문에 모델 학습 및 평가 과정에서 매 배치마다 gc.collect() 함수를 통해 가비지 콜렉터를 수동으로 동작시켜주면 최적화에 많은 도움이 됩니다. 여기서도 한번 해당 함수를 실행해주고 넘어가겠습니다.","metadata":{}},{"cell_type":"code","source":"import gc","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:12.554242Z","iopub.execute_input":"2022-07-17T15:12:12.554797Z","iopub.status.idle":"2022-07-17T15:12:12.558883Z","shell.execute_reply.started":"2022-07-17T15:12:12.554758Z","shell.execute_reply":"2022-07-17T15:12:12.558048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:12.560728Z","iopub.execute_input":"2022-07-17T15:12:12.561319Z","iopub.status.idle":"2022-07-17T15:12:12.733780Z","shell.execute_reply.started":"2022-07-17T15:12:12.561282Z","shell.execute_reply":"2022-07-17T15:12:12.732871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"KFold의 K는 5개로 지정할 것입니다. 그러면 5개의 Fold에 대해 4개의 학습 폴드, 1개의 검증 폴드를 가지게 되고, 검증 폴드의 위치를 옮겨가며 총 5회 반복 학습을 시행합니다. 결과적으로 5번의 검증 결과에 대한 평균값을 통해 최적화를 해나가게 됩니다.\n\nSEED는 임의의 숫자로 지정하면 됩니다.","metadata":{}},{"cell_type":"code","source":"FOLDS = 5\nSEED = 42\nskf = KFold(n_splits=FOLDS, shuffle=True, random_state=SEED)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:12:14.935324Z","iopub.execute_input":"2022-07-17T15:12:14.937082Z","iopub.status.idle":"2022-07-17T15:12:14.945249Z","shell.execute_reply.started":"2022-07-17T15:12:14.937036Z","shell.execute_reply":"2022-07-17T15:12:14.944478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importances = []\noof = []\nTRAIN_SUBSAMPLE = 1.0\nVER = 1 # 모델을 저장할 때 기록할 버전 정보입니다.\n\n# KFold 객체 skf에 대해 교차검증을 수행하기 위하여 학습, 검증 폴드로 분리한 다음 순회하며 반복 검증을 실시합니다.\nfor fold,(train_idx, valid_idx) in enumerate(skf.split(train_cpu, train_cpu.target)):\n    \n    # train data 중에서도 일부만 샘플링해서 학습을 수행하고 싶다면 TRAIN_SUBSAMPLE을 1 미만으로 설정해줍니다.\n    # 그럼 아래 if문을 통해 그 비율만큼 랜덤으로 추출해서 다시 train set을 구성할 수 있습니다.\n    if TRAIN_SUBSAMPLE<1.0:\n        np.random.seed(SEED)\n        train_idx = np.random.choice(train_idx, int(len(train_idx)*TRAIN_SUBSAMPLE), replace=False)\n        np.random.seed(None)\n    \n    print('#'*25)\n    print('### Fold',fold+1)\n    print('### Train size',len(train_idx),'Valid size',len(valid_idx))\n    print(f'### Training with {int(TRAIN_SUBSAMPLE*100)}% fold data...')\n    print('#'*25)\n    \n    # 앞에서 만들어둔 IterLoadForDMatrix 객체 인스턴스를 생성해줍니다. 배치로 던져주며 학습하기 위함입니다.\n    Xy_train = IterLoadForDMatrix(train_cpu.loc[train_idx], FEATURES, 'target')\n    \n    # 학습을 수행하면서 성능을 검증해줄 검증 데이터도 정의해줍니다.\n    X_valid = train_cpu.loc[valid_idx, FEATURES]\n    y_valid = train_cpu.loc[valid_idx, 'target']\n    \n    # 이제 DeviceQuantileDMatrix를 사용해 iteration을 돌면서 학습할 수 있는 dtrain 객체를 만들어주고,\n    dtrain = xgb.DeviceQuantileDMatrix(Xy_train, max_bin=256)\n    # 마찬가지로 train 데이터와 같은 형태를 취해 비교연산이 가능하도록 DMatrix로 처리해줍니다.\n    dvalid = xgb.DMatrix(data=X_valid, label=y_valid)\n    \n    # 모델을 학습합니다.\n    model = xgb.train(xgb_parms, \n                dtrain=dtrain,\n                evals=[(dtrain,'train'),(dvalid,'valid')],\n                num_boost_round=9999,\n                early_stopping_rounds=100,\n                verbose_eval=100) \n    # 교차 검증 중 1회 검증시마다 모델을 저장해줍니다.\n    model.save_model(f'XGB_v{VER}_fold{fold}.xgb')\n    \n    # 모델 학습이 끝나고 feature importance를 확인할 것입니다. 이를 위해 매 학습마다 importance를 계산하여 변수에 저장해줍니다.\n    dd = model.get_score(importance_type='weight')\n    df = pd.DataFrame({'feature':dd.keys(),f'importance_{fold}':dd.values()})\n    importances.append(df)\n            \n    # 모델을 검증합니다. 정확도는 위에서 함수로 만들어둔 대회 평가지표를 사용합니다.\n    oof_preds = model.predict(dvalid)\n    \n    # 평가지표를 계산합니다. 이 때, 원문에서는 y_valid.values를 그대로 넣어주는데, \n    # 만약 train_cpu가 아니라 GPU 메모리에 올라가있는 train 변수를 그대로 인자로 사용했다면,\n    # 해당 객체는 np.ndarray가 아니라 cupy._core.core.ndarray일 것입니다.\n    # 이 때는, cupy를 사용해 GPU위에서 그대로 연산하도록 해줍니다.\n    # -> acc = amex_metric_mod_GPU(y_valid.values, oof_preds)\n\n    # 우리는 원문과 동일하게 cpu에 올린 train_cpu를 사용하므로,\n    # 객체는 np.ndarray() 입니다. 따라서 그대로 넣어주면 되겠습니다.\n    acc = amex_metric_mod_CPU(y_valid.values, oof_preds)\n    print('Kaggle Metric =',acc,'\\n')\n    \n    # 검증 스코어(oof_pred)도 따로 저장해줍니다.\n    df = train_cpu.loc[valid_idx, ['customer_ID','target']].copy()\n    df['oof_pred'] = oof_preds\n    oof.append( df )\n    \n    # 학습을 위해 사용한 변수들은 모두 메모리 해제해줍니다.\n    del dtrain, Xy_train, dd, df\n    del X_valid, y_valid, dvalid, model\n    # 변수를 제거했음에도 남아있는 비-참조 메모리 값들도 완전히 제거하기 위해 가비지 컬랙션을 실시해줍니다.\n    _ = gc.collect()\n    \nprint('#'*25)\n# 학습이 종료되면 전체 검증 결과를 데이터프레임으로 저장해주고, \n# 실제 값과 검증 값을 평가지표로 계산해줍니다.\n# 여기서 oof도 y_valid처럼, 만약 인자로 GPU 메모리에 올라가있는 train 변수를 사용했다면,\n# pandas.core.frame.DataFrame가 아닌 cudf.core.dataframe.DataFrame 타입일 것입니다.\n# 따라서 이 때는 cudf로 병합해주거나 oof 내 요소들을 모두 pandas 객체로 바꿔줘야 합니다.\n# 우리는 train_CPU를 사용했으므로 pandas를 그대로 사용합니다.\noof = pd.concat(oof,axis=0,ignore_index=True).set_index('customer_ID')\nacc = amex_metric_mod_GPU(oof.target.values, oof.oof_pred.values)\nprint('OVERALL CV Kaggle Metric =',acc)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-07-17T15:12:16.668349Z","iopub.execute_input":"2022-07-17T15:12:16.668702Z","iopub.status.idle":"2022-07-17T15:21:43.664459Z","shell.execute_reply.started":"2022-07-17T15:12:16.668674Z","shell.execute_reply":"2022-07-17T15:21:43.663557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 학습이 종료되었으니 이제 train_cpu 데이터셋도 필요가 없어졌습니다.\n# train_cpu 데이터를 참조하고 있는 메모리도 정리해주겠습니다.\ndel train, train_cpu\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:43.666493Z","iopub.execute_input":"2022-07-17T15:21:43.666902Z","iopub.status.idle":"2022-07-17T15:21:43.808528Z","shell.execute_reply.started":"2022-07-17T15:21:43.666863Z","shell.execute_reply":"2022-07-17T15:21:43.807331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:43.809943Z","iopub.execute_input":"2022-07-17T15:21:43.810975Z","iopub.status.idle":"2022-07-17T15:21:44.556624Z","shell.execute_reply.started":"2022-07-17T15:21:43.810936Z","shell.execute_reply":"2022-07-17T15:21:44.555604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Save OOF Preds","metadata":{}},{"cell_type":"markdown","source":"예측 결과를 customer_ID와 함께 저장해줄 것입니다. 데이터 파일에서 unique한 id 정보만 가져와 앞에서 수행했던 것처럼 16진수를 정수형으로 바꿔주고, 학습 과정에서 확보한 예측값을 병합해주면 됩니다.","metadata":{}},{"cell_type":"code","source":"TRAIN_PATH = '../input/amex-data-integer-dtypes-parquet-format/train.parquet'\noof_xgb = pd.read_parquet(TRAIN_PATH, columns=['customer_ID']).drop_duplicates()\noof_xgb['customer_ID_hash'] = oof_xgb['customer_ID'].apply(lambda x: int(x[-16:],16) ).astype('int64')\noof_xgb = oof_xgb.set_index('customer_ID_hash')\noof_xgb = oof_xgb.merge(oof, left_index=True, right_index=True)\noof_xgb = oof_xgb.sort_index().reset_index(drop=True)\noof_xgb.to_csv(f'oof_xgb_v{VER}.csv',index=False)\noof_xgb.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:44.559029Z","iopub.execute_input":"2022-07-17T15:21:44.559779Z","iopub.status.idle":"2022-07-17T15:21:49.743266Z","shell.execute_reply.started":"2022-07-17T15:21:44.559734Z","shell.execute_reply":"2022-07-17T15:21:49.742467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"예측 결과를 시각화해봅니다. 예측 결과는 0에서 1사이의 확률입니다. 단 5%만 채무 불이행(default, 1)이고 나머지는 0으로 예측해야 하기 때문에 시각화했을 때 0과 1쪽에 양방향으로 쏠려있되, 0에 더 많은 값이 분포되어 있어야 합니다.","metadata":{}},{"cell_type":"code","source":"plt.hist(oof_xgb.oof_pred.values, bins=100)\nplt.title('OOF Predictions')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:49.744713Z","iopub.execute_input":"2022-07-17T15:21:49.745147Z","iopub.status.idle":"2022-07-17T15:21:50.103490Z","shell.execute_reply.started":"2022-07-17T15:21:49.745111Z","shell.execute_reply":"2022-07-17T15:21:50.102717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"예측치를 csv 파일로 저장했으니 변수를 제거합니다. 이렇게 파일로 저장하고 메모리를 비우는 작업을 반복하는 이유는 캐글 상에서 지원되는 RAM이 크지 않기 때문입니다. 일반적으로 로컬 환경에서도 RAM은 여유롭지 않기 때문에 중간에 참조나 복사를 하면서 셧다운 되는 일이 발생할 수 있는데, 이를 방지해주기 위해 하드디스크로 파일을 저장하고 RAM은 가벼운 상태를 유지해주는 것이 좋습니다.","metadata":{}},{"cell_type":"code","source":"del oof_xgb, oof\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:50.104837Z","iopub.execute_input":"2022-07-17T15:21:50.105203Z","iopub.status.idle":"2022-07-17T15:21:50.274650Z","shell.execute_reply.started":"2022-07-17T15:21:50.105168Z","shell.execute_reply":"2022-07-17T15:21:50.273140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:50.276921Z","iopub.execute_input":"2022-07-17T15:21:50.277366Z","iopub.status.idle":"2022-07-17T15:21:51.074752Z","shell.execute_reply.started":"2022-07-17T15:21:50.277325Z","shell.execute_reply":"2022-07-17T15:21:51.073802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Feature Importance","metadata":{}},{"cell_type":"markdown","source":"Feature Importance는 피처 중요도, 즉 모델이 테스크를 수행(여기서는 예측)하는데에 어떤 변수가 활용 비중이 높은가에 대한 정보입니다.\n모델 학습이 끝난 다음 이렇게 살펴보면서 예측에 중요하지 않은 변수가 있다면 제거하고 중요도가 지나치게 변수들은 예측 대상과의 인과성 혹은 상관성을 확인하는 등의 피드백 과정을 거치게 됩니다.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# 교차검증을 수행하면서 FOLD 수(5)만큼 importances도 누적해서 구해졌을 것입니다.\n# df 변수에 병합하여 feature마다 평균 importance를 계산해줍니다.\ndf = importances[0].copy()\nfor k in range(1,FOLDS):\n    df = df.merge(importances[k], on='feature', how='left')\ndf['importance'] = df.iloc[:,1:].mean(axis=1)\ndf = df.sort_values('importance',ascending=False)\ndf.to_csv(f'xgb_feature_importance_v{VER}.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.077352Z","iopub.execute_input":"2022-07-17T15:21:51.078026Z","iopub.status.idle":"2022-07-17T15:21:51.122603Z","shell.execute_reply.started":"2022-07-17T15:21:51.077985Z","shell.execute_reply":"2022-07-17T15:21:51.121766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.123801Z","iopub.execute_input":"2022-07-17T15:21:51.124210Z","iopub.status.idle":"2022-07-17T15:21:51.149812Z","shell.execute_reply.started":"2022-07-17T15:21:51.124172Z","shell.execute_reply":"2022-07-17T15:21:51.149088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"bar 차트로 중요도 기준 상위 20개 feature만 시각화해줍니다.","metadata":{}},{"cell_type":"code","source":"NUM_FEATURES = 20\nplt.figure(figsize=(10,5*NUM_FEATURES//10))\nplt.barh(np.arange(NUM_FEATURES,0,-1), df.importance.values[:NUM_FEATURES])\nplt.yticks(np.arange(NUM_FEATURES,0,-1), df.feature.values[:NUM_FEATURES])\nplt.title(f'XGB Feature Importance - Top {NUM_FEATURES}')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.152665Z","iopub.execute_input":"2022-07-17T15:21:51.153012Z","iopub.status.idle":"2022-07-17T15:21:51.426444Z","shell.execute_reply.started":"2022-07-17T15:21:51.152985Z","shell.execute_reply":"2022-07-17T15:21:51.425687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Data Processing and Feature Engineering for Test Data","metadata":{}},{"cell_type":"markdown","source":"우리는 특정 customer가 지출액을 상환할 것인가가 궁금합니다. 따라서 unique한 customer_ID를 먼저 구하겠습니다.","metadata":{}},{"cell_type":"code","source":"TEST_PATH = '../input/amex-data-integer-dtypes-parquet-format/test.parquet'\ntest = read_file_CPU(path = TEST_PATH, usecols = ['customer_ID','S_2'])\ntest","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:21:51.427905Z","iopub.execute_input":"2022-07-17T15:21:51.428486Z","iopub.status.idle":"2022-07-17T15:23:29.377937Z","shell.execute_reply.started":"2022-07-17T15:21:51.428447Z","shell.execute_reply":"2022-07-17T15:23:29.377128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers = test[['customer_ID']].drop_duplicates().sort_index().values.flatten()\ncustomers","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:29.379397Z","iopub.execute_input":"2022-07-17T15:23:29.379978Z","iopub.status.idle":"2022-07-17T15:23:29.735532Z","shell.execute_reply.started":"2022-07-17T15:23:29.379941Z","shell.execute_reply":"2022-07-17T15:23:29.734691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"이렇게 ID 정보는 확보했고, 각각의 ID 마다 default(불이행) 여부를 예측하게 됩니다. 이제 test 데이터셋을 불러올텐데, 지금은 2개 컬럼만 가져왔지만 train 데이터와 마찬가지로 용량이 매우 큽니다. 따라서 CPU에 데이터셋을 PART만큼 배치 형태로 올린 다음, 해당 PART를 GPU로 넘겨주며 모델이 예측을 수행할 수 있도록 할 것입니다. \n\n그러기 위해서는 PART를 나누는 기준이 필요하며, 기준은 customer_ID의 중복이 제거되기 전 test 데이터셋에서의 row 사이즈와 중복을 제거한 후의 일정한 chunk 사이즈 2가지가 필요합니다. 앞에서 만든 process_and_feature_engineer() 함수를 기억하시나요? 해당 함수를 통과하면 group_by로 customer_ID의 중복이 제거되고 통계값을 가진 데이터셋으로 바뀝니다. 이렇게 변환된 데이터셋을 chunk 크기로 넘겨주며 예측을 수행하는 방식입니다.\n\n따라서 PART를 나누는 2가지 기준을 구할 수 있도록 함수를 만들고, 실행해줍니다.","metadata":{}},{"cell_type":"code","source":"def get_rows(customers, test, NUM_PARTS = 10, verbose = ''):\n    # 한 청크(1개 PART의 사이즈)는 전체 평가 데이터셋 크기를 PART 수로 나눠주면 구할 수 있습니다.\n    # 모델 학습할 때 배치 사이즈 구했던 것과 비슷합니다.\n    chunk = len(customers)//NUM_PARTS\n    if verbose != '':\n        print(f'We will process {verbose} data as {NUM_PARTS} separate parts.')\n        print(f'There will be {chunk} customers in each part (except the last part).')\n        print('Below are number of rows in each part:')\n    rows = []\n\n    for k in range(NUM_PARTS):\n        # 마지막 PART면 남은 부분을 모두 cc에 담습니다.\n        if k==NUM_PARTS-1: \n            cc = customers[k*chunk:]\n        # 마지막 PART가 아니면 앞에서부터 청크 단위로 잘라서 cc에 담습니다.\n        else: \n            cc = customers[k*chunk:(k+1)*chunk]\n        # 현재 PART에 포함된 customer_ID 수를 구하는 것을 통해 PART 사이즈를 계산합니다.\n        s = test.loc[test.customer_ID.isin(cc)].shape[0]\n        # rows에는 10개 PART의 크기가 담겨있습니다.\n        rows.append(s)\n    if verbose != '': print( rows )\n    return rows, chunk\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:29.737001Z","iopub.execute_input":"2022-07-17T15:23:29.737379Z","iopub.status.idle":"2022-07-17T15:23:29.745124Z","shell.execute_reply.started":"2022-07-17T15:23:29.737343Z","shell.execute_reply":"2022-07-17T15:23:29.744260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"이제 위에서 만든 함수로 customer_ID를 총 10개 그룹(PART)로 나누겠습니다.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T13:29:13.319054Z","iopub.execute_input":"2022-07-15T13:29:13.319608Z","iopub.status.idle":"2022-07-15T13:29:13.327749Z","shell.execute_reply.started":"2022-07-15T13:29:13.319567Z","shell.execute_reply":"2022-07-15T13:29:13.326253Z"}}},{"cell_type":"code","source":"# 참고로 원문에서는 테스트 데이터셋을 4개로 나눠주었습니다. \n# kaggle 서버를 사용하면 4개로 지정했을 때 GPU 메모리 초과 에러가 나기 때문에 10개로 분할하겠습니다.\nNUM_PARTS = 10\nrows,num_cust = get_rows(customers, test[['customer_ID']], NUM_PARTS = NUM_PARTS, verbose = 'test')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:29.746701Z","iopub.execute_input":"2022-07-17T15:23:29.747098Z","iopub.status.idle":"2022-07-17T15:23:30.991356Z","shell.execute_reply.started":"2022-07-17T15:23:29.747061Z","shell.execute_reply":"2022-07-17T15:23:30.990396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"마지막 줄에 각 PART별 크기가 출력되었습니다. 여기서 말하는 크기는 chunk size가 아닌 원본 테스트 데이터 기준에서의 중복을 포함한 크기입니다. \n따라서 모두 더했을 때 test 데이터셋의 사이즈와 같아야 합니다.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T13:34:33.019463Z","iopub.execute_input":"2022-07-15T13:34:33.019863Z","iopub.status.idle":"2022-07-15T13:34:33.026719Z","shell.execute_reply.started":"2022-07-15T13:34:33.019831Z","shell.execute_reply":"2022-07-15T13:34:33.025412Z"}}},{"cell_type":"code","source":"sum(rows) == len(test)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:30.992726Z","iopub.execute_input":"2022-07-17T15:23:30.993203Z","iopub.status.idle":"2022-07-17T15:23:30.999218Z","shell.execute_reply.started":"2022-07-17T15:23:30.993162Z","shell.execute_reply":"2022-07-17T15:23:30.998334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"chunk size는 다음과 같습니다","metadata":{}},{"cell_type":"code","source":"num_cust","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:31.000686Z","iopub.execute_input":"2022-07-17T15:23:31.001303Z","iopub.status.idle":"2022-07-17T15:23:31.010702Z","shell.execute_reply.started":"2022-07-17T15:23:31.001267Z","shell.execute_reply":"2022-07-17T15:23:31.009824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"불러왔던 테스트 데이터는 메모리 할당 해제해주겠습니다.","metadata":{}},{"cell_type":"code","source":"del test\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:31.012462Z","iopub.execute_input":"2022-07-17T15:23:31.013193Z","iopub.status.idle":"2022-07-17T15:23:31.161038Z","shell.execute_reply.started":"2022-07-17T15:23:31.013099Z","shell.execute_reply":"2022-07-17T15:23:31.160076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:23:31.162702Z","iopub.execute_input":"2022-07-17T15:23:31.163083Z","iopub.status.idle":"2022-07-17T15:23:31.914814Z","shell.execute_reply.started":"2022-07-17T15:23:31.163048Z","shell.execute_reply":"2022-07-17T15:23:31.913827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. Infer Test","metadata":{}},{"cell_type":"markdown","source":"아래 코드를 통해 우리는 데이터를 10개 파트로 분할해 처리하고, 예측하는 작업을 반복할 것입니다. 과정은 다음과 같습니다.\n1. 각 PART를 포함할 수 있도록 원본 테스트 데이터셋을 batch로 분할해서 CPU 메모리에 올려줍니다.\n2. 해당 PART는 다시 연산을 위해 GPU로 올려주고,\n3. process_and_feature_engineer() 함수로 모델에 넣기 위한 데이터(customer_ID별 통계치)로 바꿔줍니다.\n4. 그럼, 기존 데이터 shape과 달라지게 되는데 이때, 앞에서 확보한 chunk size로 인덱싱을 수행할 수 있습니다.\n5. chunk size로 인덱싱한 다음 모델로 예측을 하고, \n6. 모든 chunk에 대해 예측이 끝나면 병합해서 전체 customer_ID에 대해 예측한 결과를 돌려줍니다.","metadata":{}},{"cell_type":"code","source":"skip_rows = 0\nskip_cust = 0\ntest_preds = []\n\n# iteration 객체를 생성해줍니다. 해당 객체를 호출해가며 PART 단위 rows만큼 CPU 메모리로 올릴 수 있습니다.\nTEST_PATH = '../input/amex-data-integer-dtypes-parquet-format/test.parquet'\nbatch = ParquetFile(TEST_PATH)\n# batch_size는 한번 호출시 고정되므로 매번 다른 PART 크기만큼 가져올 수 없습니다.\n# 그러면, 현재 batch에서 예측해야 하는 customer_ID가 직전 batch 때 이미 로드되어서 현재 batch에서 누락되는 경우가 발생합니다.\n# 따라서 가장 큰 PART 크기만큼 가져오되, 직전 PART와 병합하여 모든 customer_ID가 인덱싱 될 수 있도록 합니다.\nprev_batch = pd.DataFrame()\nfor pres_batch, k in zip(batch.iter_batches(batch_size=max(rows)), range(NUM_PARTS)): \n    print(f'\\nReading test data...')\n    # 직전 batch와 현재 batch에서 로드된 데이터셋을 병합해 CPU 메모리에 올려줍니다.\n    iter_batch = pd.concat([prev_batch, pres_batch.to_pandas()])\n    test_cpu = read_file_CPU(iter_batch = iter_batch, path = TEST_PATH)\n    # 현재 batch는 직전 batch로 덮어쓰고,\n    # 메모리를 정리해줍니다.\n    prev_batch = pres_batch.to_pandas()\n    del pres_batch, iter_batch\n    _ = gc.collect()\n    \n    # 배치로 학습시킬 PART만 불러왔다면, 해당 데이터를 GPU 메모리에 올려줍니다.\n    test_gpu = cudf.DataFrame(test_cpu)\n    skip_rows += rows[k]\n    print(f'=> Test part {k+1} has shape', test_gpu.shape)\n    \n    # 전처리를 수행합니다. customer_ID별로 통계값이 구해지고, customer_ID 중복은 제거됩니다.\n    # 참고로, process_and_feature_engineer() 내부적으로 pandas가 아닌 cudf를 사용합니다.\n    # 따라서 전처리를 수행할 때 반드시 데이터는 GPU 위에 올려둔 상태여야 합니다.\n    test_gpu = process_and_feature_engineer(test_gpu)\n    \n    # num_cust는 중복 제거된 customer_ID에 대한 chunk size입니다.\n    # chunk size를 증가시켜가며 모든 customer_ID에 대해 신용 default 예측을 수행할 것입니다.\n    if k==NUM_PARTS-1: \n        test_gpu = test_gpu.loc[customers[skip_cust:]]\n    else: \n        test_gpu = test_gpu.loc[customers[skip_cust:skip_cust+num_cust]]\n    skip_cust += num_cust\n    print('shape after indexing(by chunk size)', test_gpu.shape)\n    \n    # 테스트 데이터에서 label을 제외한 feature들을 X_test에 넘겨줍니다.\n    X_test = test_gpu[FEATURES]\n    dtest = xgb.DMatrix(data=X_test)\n    del X_test\n    gc.collect()\n\n    # 학습한 모델을 가져와 예측을 수행합니다. \n    # 우리 모델은 교차검증 방식으로 학습했습니다.\n    # 예측도 동일한 방식으로 수행하며, 예측 결과는 전체 FOLD 결과의 평균값으로 얻을 수 있습니다.\n    model = xgb.Booster()\n    model.load_model(f'XGB_v{VER}_fold0.xgb')\n    preds = model.predict(dtest)\n    for f in range(1,FOLDS):\n        model.load_model(f'XGB_v{VER}_fold{f}.xgb')\n        preds += model.predict(dtest)\n    preds /= FOLDS\n    test_preds.append(preds)\n\n    # 메모리를 비워줍니다.\n    del dtest, model\n    _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:28:45.079491Z","iopub.execute_input":"2022-07-17T15:28:45.080032Z","iopub.status.idle":"2022-07-17T15:34:25.417487Z","shell.execute_reply.started":"2022-07-17T15:28:45.079999Z","shell.execute_reply":"2022-07-17T15:34:25.416673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9. Create Submission CSV\n\n마지막으로 제출용 파일을 만들고, 테스트 결과를 시각화해봅니다.","metadata":{}},{"cell_type":"code","source":"len(test_preds)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:36:03.302426Z","iopub.execute_input":"2022-07-17T15:36:03.302830Z","iopub.status.idle":"2022-07-17T15:36:03.315063Z","shell.execute_reply.started":"2022-07-17T15:36:03.302799Z","shell.execute_reply":"2022-07-17T15:36:03.313723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 제출용 파일은 대회 가이드에 맞게 형태를 맞춰 저장해주면 됩니다.\ntest_preds = np.concatenate(test_preds)\ntest = cudf.DataFrame(index=customers,data={'prediction':test_preds})\nsub = cudf.read_csv('../input/amex-default-prediction/sample_submission.csv')[['customer_ID']]\nsub['customer_ID_hash'] = sub['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\nsub = sub.set_index('customer_ID_hash')\nsub = sub.merge(test[['prediction']], left_index=True, right_index=True, how='left')\nsub = sub.reset_index(drop=True)\n\nsub.to_csv(f'submission_xgb_v{VER}.csv',index=False)\nprint('Submission file shape is', sub.shape )\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:34:25.424454Z","iopub.execute_input":"2022-07-17T15:34:25.426523Z","iopub.status.idle":"2022-07-17T15:34:26.533911Z","shell.execute_reply.started":"2022-07-17T15:34:25.426485Z","shell.execute_reply":"2022-07-17T15:34:26.533075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# hist plot으로 시각화해봅니다.\nplt.hist(sub.to_pandas().prediction, bins=100)\nplt.title('Test Predictions')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T15:34:31.686057Z","iopub.execute_input":"2022-07-17T15:34:31.686514Z","iopub.status.idle":"2022-07-17T15:34:32.812050Z","shell.execute_reply.started":"2022-07-17T15:34:31.686474Z","shell.execute_reply":"2022-07-17T15:34:32.811173Z"},"trusted":true},"execution_count":null,"outputs":[]}]}