{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"papermill":{"default_parameters":{},"duration":41.015305,"end_time":"2024-03-22T07:27:49.412065","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-03-22T07:27:08.396760","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### I recently started learning machine learning, and this competition has taught me a lot. Although Polars is very fast, I don't like its API incompatibility with Pandas. So, I used Pandas to read Parquet files for data loading, and it didn't result in any memory overflow. I hope this notebook is helpful for beginners.","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nimport numpy as np\nimport pandas as pd\nfrom sklearn.utils import shuffle\nfrom sklearn.preprocessing import LabelEncoder\nle = LabelEncoder()\n\ncsv_path = '/kaggle/input/home-credit-credit-risk-model-stability/csv_files/test/test_'\npar_path = '/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train/train_'","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:40:07.109257Z","iopub.execute_input":"2024-04-03T00:40:07.110085Z","iopub.status.idle":"2024-04-03T00:40:09.315856Z","shell.execute_reply.started":"2024-04-03T00:40:07.110051Z","shell.execute_reply":"2024-04-03T00:40:09.314922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base = pd.read_parquet(par_path + \"base.parquet\", columns=['case_id', 'target'])\nbase.shape","metadata":{"papermill":{"duration":0.165697,"end_time":"2024-03-22T07:27:26.729347","exception":false,"start_time":"2024-03-22T07:27:26.563650","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-03T00:40:09.317344Z","iopub.execute_input":"2024-04-03T00:40:09.317764Z","iopub.status.idle":"2024-04-03T00:40:09.566678Z","shell.execute_reply.started":"2024-04-03T00:40:09.317738Z","shell.execute_reply":"2024-04-03T00:40:09.565707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### In the following section, the features can be sorted by importance and saved to a CSV file named 'importances-gt0.csv'.","metadata":{}},{"cell_type":"code","source":"# useful_cols = pd.read_csv('/kaggle/input/lgbdata/importances-gt0.csv')\n# useful_cols = useful_cols[useful_cols['weight'] > 0]['name'].unique()\n# useful_cols.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:40:09.568314Z","iopub.execute_input":"2024-04-03T00:40:09.568719Z","iopub.status.idle":"2024-04-03T00:40:09.572580Z","shell.execute_reply.started":"2024-04-03T00:40:09.568684Z","shell.execute_reply":"2024-04-03T00:40:09.571804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def filter_df(fname):\n#     fcols = pd.read_csv(csv_path + fname + '.csv', nrows=2).columns\n#     columns_to_keep = [col for col in fcols if col in useful_cols]\n#     df = pd.read_parquet(par_path + fname + '.parquet', columns=columns_to_keep)\n    df = pd.read_parquet(par_path + fname + '.parquet')\n    numeric_cols = df.select_dtypes(include=['int64', 'float64']).columns\n    df[numeric_cols] = df[numeric_cols].astype('float32')\n    return df\n\ndef depth1_feats(df:pd.DataFrame):\n    numeric_cols = df.select_dtypes(include=['number']).columns.tolist()\n    numeric_cols.remove('case_id')\n    numeric_cols.remove('num_group1')\n    aggfeats = df.groupby('case_id')[numeric_cols].agg('sum').reset_index()\n\n    notnum_cols = df.select_dtypes(exclude=['number']).columns.tolist()\n    notnum_cols.append('case_id')\n    filfeats = df[df['num_group1'] == 0]\n    filfeats = filfeats.drop('num_group1', axis=1)\n    filfeats = filfeats.filter(items=notnum_cols)\n    return pd.merge(filfeats, aggfeats, how='left', on='case_id')\n\ndef depth2_feats(df:pd.DataFrame):\n    numeric_cols = df.select_dtypes(include=['number']).columns.tolist()\n    numeric_cols.remove('case_id')\n    numeric_cols.remove('num_group1')\n    numeric_cols.remove('num_group2')\n    aggfeats = df.groupby('case_id')[numeric_cols].agg('sum').reset_index()\n\n    notnum_cols = df.select_dtypes(exclude=['number']).columns.tolist()\n    notnum_cols.append('case_id')\n    df = df[df['num_group1'] == 0]\n    df = df[df['num_group2'] == 0]\n    filterdf = df.drop(['num_group1', 'num_group2'], axis=1)\n    filterdf = filterdf.filter(items=notnum_cols)\n    return pd.merge(filterdf, aggfeats, how='left', on='case_id')   \n","metadata":{"papermill":{"duration":0.028407,"end_time":"2024-03-22T07:27:26.764009","exception":false,"start_time":"2024-03-22T07:27:26.735602","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-03T00:40:09.574710Z","iopub.execute_input":"2024-04-03T00:40:09.575195Z","iopub.status.idle":"2024-04-03T00:40:09.585746Z","shell.execute_reply.started":"2024-04-03T00:40:09.575164Z","shell.execute_reply":"2024-04-03T00:40:09.584729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### tables of depth ==2 ","metadata":{}},{"cell_type":"markdown","source":"### looks like stupid, bu it won't cause Out of Memory","metadata":{}},{"cell_type":"code","source":"cache = None\nfor id in range(2):\n    bureau_a_2 = filter_df('credit_bureau_a_2_' + str(id))\n    if cache is None:\n        cache = bureau_a_2\n    else:\n        cache = pd.concat([cache, bureau_a_2])\ntmp = depth2_feats(cache)    \ndel cache\ngc.collect()\n\ncache = None\nfor id in range(2, 4):\n    bureau_a_2 = filter_df('credit_bureau_a_2_' + str(id))\n    if cache is None:\n        cache = bureau_a_2\n    else:\n        cache = pd.concat([cache, bureau_a_2])\ntmp = pd.concat([tmp, depth2_feats(cache)])    \ndel cache\ngc.collect()      \n        \ncache = None\nfor id in range(4, 6):\n    bureau_a_2 = filter_df('credit_bureau_a_2_' + str(id))\n    if cache is None:\n        cache = bureau_a_2\n    else:\n        cache = pd.concat([cache, bureau_a_2])\ntmp = pd.concat([tmp, depth2_feats(cache)])    \ndel cache\ngc.collect()            \n        \ncache = None\nfor id in range(6, 8):\n    bureau_a_2 = filter_df('credit_bureau_a_2_' + str(id))\n    if cache is None:\n        cache = bureau_a_2\n    else:\n        cache = pd.concat([cache, bureau_a_2])\ntmp = pd.concat([tmp, depth2_feats(cache)])    \ndel cache\ngc.collect()     \n    \ncache = None\nfor id in range(8, 10):\n    bureau_a_2 = filter_df('credit_bureau_a_2_' + str(id))\n    if cache is None:\n        cache = bureau_a_2\n    else:\n        cache = pd.concat([cache, bureau_a_2])\ntmp = pd.concat([tmp, depth2_feats(cache)])    \ndel cache\ngc.collect()     \n    \ncache = None\nfor id in range(10, 11):\n    bureau_a_2 = filter_df('credit_bureau_a_2_' + str(id))\n    if cache is None:\n        cache = bureau_a_2\n    else:\n        cache = pd.concat([cache, bureau_a_2])\ntmp = pd.concat([tmp, depth2_feats(cache)])    \ndel cache\ngc.collect()      \n\nprint('merging...')\ndata = pd.merge(base, tmp, how=\"left\", on=\"case_id\")\ndel tmp\ngc.collect() \ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:40:09.587098Z","iopub.execute_input":"2024-04-03T00:40:09.587674Z","iopub.status.idle":"2024-04-03T00:43:35.648823Z","shell.execute_reply.started":"2024-04-03T00:40:09.587646Z","shell.execute_reply":"2024-04-03T00:43:35.647445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth2_feats(filter_df(\"credit_bureau_b_2\")), how=\"left\", on=\"case_id\")\ngc.collect()\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:43:35.650355Z","iopub.execute_input":"2024-04-03T00:43:35.650689Z","iopub.status.idle":"2024-04-03T00:43:36.238444Z","shell.execute_reply.started":"2024-04-03T00:43:35.650661Z","shell.execute_reply":"2024-04-03T00:43:36.237187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth2_feats(filter_df(\"person_2\")), how='left', on=\"case_id\")\ngc.collect()\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:43:36.239850Z","iopub.execute_input":"2024-04-03T00:43:36.240285Z","iopub.status.idle":"2024-04-03T00:43:39.442975Z","shell.execute_reply.started":"2024-04-03T00:43:36.240247Z","shell.execute_reply":"2024-04-03T00:43:39.441931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### tables of depth ==1","metadata":{}},{"cell_type":"code","source":"cache = None\nfor id in range(2):\n    cache = pd.concat([cache, filter_df('applprev_1_' + str(id))])\n\ndata = pd.merge(base, depth1_feats(cache), how=\"left\", on=\"case_id\")\ndel cache\ngc.collect() \ndata.shape","metadata":{"papermill":{"duration":0.258298,"end_time":"2024-03-22T07:27:27.545099","exception":false,"start_time":"2024-03-22T07:27:27.286801","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-03T00:43:39.444371Z","iopub.execute_input":"2024-04-03T00:43:39.444701Z","iopub.status.idle":"2024-04-03T00:44:04.164382Z","shell.execute_reply.started":"2024-04-03T00:43:39.444675Z","shell.execute_reply":"2024-04-03T00:44:04.162632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cache = None\nfor id in range(4):\n    cache = pd.concat([cache, filter_df('credit_bureau_a_1_' + str(id))])\n\ndata = pd.merge(data, depth1_feats(cache), how=\"left\", on=\"case_id\")\ndel cache\ngc.collect()\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:44:04.165679Z","iopub.execute_input":"2024-04-03T00:44:04.165961Z","iopub.status.idle":"2024-04-03T00:45:30.944811Z","shell.execute_reply.started":"2024-04-03T00:44:04.165938Z","shell.execute_reply":"2024-04-03T00:45:30.943837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"credit_bureau_b_1\")), how=\"left\", on=\"case_id\")\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:45:30.948538Z","iopub.execute_input":"2024-04-03T00:45:30.948836Z","iopub.status.idle":"2024-04-03T00:45:35.644776Z","shell.execute_reply.started":"2024-04-03T00:45:30.948812Z","shell.execute_reply":"2024-04-03T00:45:35.643779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"debitcard_1\")), how=\"left\", on=\"case_id\")\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:45:35.645881Z","iopub.execute_input":"2024-04-03T00:45:35.646180Z","iopub.status.idle":"2024-04-03T00:45:40.672087Z","shell.execute_reply.started":"2024-04-03T00:45:35.646155Z","shell.execute_reply":"2024-04-03T00:45:40.671047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"deposit_1\")), how=\"left\", on=\"case_id\")\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:45:40.673134Z","iopub.execute_input":"2024-04-03T00:45:40.673431Z","iopub.status.idle":"2024-04-03T00:45:45.844731Z","shell.execute_reply.started":"2024-04-03T00:45:40.673408Z","shell.execute_reply":"2024-04-03T00:45:45.843597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"other_1\")), how='left', on=\"case_id\")\ngc.collect()\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:45:45.846062Z","iopub.execute_input":"2024-04-03T00:45:45.846378Z","iopub.status.idle":"2024-04-03T00:45:51.153009Z","shell.execute_reply.started":"2024-04-03T00:45:45.846353Z","shell.execute_reply":"2024-04-03T00:45:51.151824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"person_1\")), how='left', on=\"case_id\")\ngc.collect()\ndata.shape","metadata":{"papermill":{"duration":0.215245,"end_time":"2024-03-22T07:27:27.767696","exception":false,"start_time":"2024-03-22T07:27:27.552451","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-03T00:45:51.154358Z","iopub.execute_input":"2024-04-03T00:45:51.155066Z","iopub.status.idle":"2024-04-03T00:46:02.983103Z","shell.execute_reply.started":"2024-04-03T00:45:51.155037Z","shell.execute_reply":"2024-04-03T00:46:02.982297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"tax_registry_a_1\")), how=\"left\", on=\"case_id\")\ngc.collect()\ndata.shape","metadata":{"papermill":{"duration":0.199295,"end_time":"2024-03-22T07:27:28.202351","exception":false,"start_time":"2024-03-22T07:27:28.003056","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-03T00:46:02.984346Z","iopub.execute_input":"2024-04-03T00:46:02.984839Z","iopub.status.idle":"2024-04-03T00:46:11.905230Z","shell.execute_reply.started":"2024-04-03T00:46:02.984812Z","shell.execute_reply":"2024-04-03T00:46:11.904005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"tax_registry_b_1\")), how=\"left\", on=\"case_id\")\ndata.shape","metadata":{"papermill":{"duration":0.219161,"end_time":"2024-03-22T07:27:28.428976","exception":false,"start_time":"2024-03-22T07:27:28.209815","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-03T00:46:11.906553Z","iopub.execute_input":"2024-04-03T00:46:11.906922Z","iopub.status.idle":"2024-04-03T00:46:20.049099Z","shell.execute_reply.started":"2024-04-03T00:46:11.906847Z","shell.execute_reply":"2024-04-03T00:46:20.048016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.merge(data, depth1_feats(filter_df(\"tax_registry_c_1\")), how=\"left\", on=\"case_id\")\ndata.shape","metadata":{"papermill":{"duration":0.221613,"end_time":"2024-03-22T07:27:28.659501","exception":false,"start_time":"2024-03-22T07:27:28.437888","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-04-03T00:46:20.050392Z","iopub.execute_input":"2024-04-03T00:46:20.050705Z","iopub.status.idle":"2024-04-03T00:46:29.579607Z","shell.execute_reply.started":"2024-04-03T00:46:20.050680Z","shell.execute_reply":"2024-04-03T00:46:29.578482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### tables of depth ==0","metadata":{}},{"cell_type":"code","source":"cache = None\nfor id in range(2):\n    cache = pd.concat([cache, filter_df('static_0_' + str(id))])\n\ndata = pd.merge(data, cache, how=\"left\", on=\"case_id\")\ndel cache\ngc.collect()\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:46:29.582731Z","iopub.execute_input":"2024-04-03T00:46:29.583095Z","iopub.status.idle":"2024-04-03T00:46:54.115503Z","shell.execute_reply.started":"2024-04-03T00:46:29.583059Z","shell.execute_reply":"2024-04-03T00:46:54.114458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feats_static_cb_0 =  filter_df('static_cb_0')\n\n# trun riskassesment_302T to float\ndef convert_to_float(value):\n    if pd.isna(value):\n        return None\n\n    if '-' in value:\n        try:\n            values = value.split('-')\n            start = float(values[0].strip(' ').strip('%'))\n            end = float(values[1].strip('%').strip('%'))\n        except:\n            print(value)\n            pass\n        return max(start, end) / 100\n    else:\n        return float(value.strip('%')) / 100\n\nfeats_static_cb_0['riskassesment_302T'] = feats_static_cb_0['riskassesment_302T'].apply(convert_to_float)\n\ndata = pd.merge(data, feats_static_cb_0, how=\"left\", on=\"case_id\")\ndel feats_static_cb_0\ngc.collect()\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:46:54.117008Z","iopub.execute_input":"2024-04-03T00:46:54.117438Z","iopub.status.idle":"2024-04-03T00:47:10.911451Z","shell.execute_reply.started":"2024-04-03T00:46:54.117393Z","shell.execute_reply":"2024-04-03T00:47:10.909987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = data.drop(['case_id'], axis=1)\ndata.to_parquet('all_data.parquet', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:47:10.913606Z","iopub.execute_input":"2024-04-03T00:47:10.915460Z","iopub.status.idle":"2024-04-03T00:48:01.763173Z","shell.execute_reply.started":"2024-04-03T00:47:10.915416Z","shell.execute_reply":"2024-04-03T00:48:01.762290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now, we still have more than 20GB of memory. Let's start LGB","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score \nimport lightgbm as lgb","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:48:01.764748Z","iopub.execute_input":"2024-04-03T00:48:01.765527Z","iopub.status.idle":"2024-04-03T00:48:02.780792Z","shell.execute_reply.started":"2024-04-03T00:48:01.765489Z","shell.execute_reply":"2024-04-03T00:48:02.779727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# can read the parquet file that you saved before in your own datasts.\n# data = pd.read_parquet('/kaggle/working/all-data.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:48:02.782302Z","iopub.execute_input":"2024-04-03T00:48:02.783285Z","iopub.status.idle":"2024-04-03T00:48:02.787230Z","shell.execute_reply.started":"2024-04-03T00:48:02.783244Z","shell.execute_reply":"2024-04-03T00:48:02.786202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = data['target']\ndata = data.drop('target', axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:48:02.788498Z","iopub.execute_input":"2024-04-03T00:48:02.788800Z","iopub.status.idle":"2024-04-03T00:48:07.097836Z","shell.execute_reply.started":"2024-04-03T00:48:02.788775Z","shell.execute_reply":"2024-04-03T00:48:07.096782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numcols = []\ncatcols = []\ndef divide(df: pd.DataFrame) -> pd.DataFrame:\n    for col in df.columns:  \n        if df[col].dtype.name in ['int32', 'float32']:\n            numcols.append(col) \n        else:\n            catcols.append(col)\n    return df\ndata = divide(data)\nprint(f'catcols: {len(catcols)}, numcols: {len(numcols)}')","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:48:07.099015Z","iopub.execute_input":"2024-04-03T00:48:07.099355Z","iopub.status.idle":"2024-04-03T00:48:07.122960Z","shell.execute_reply.started":"2024-04-03T00:48:07.099321Z","shell.execute_reply":"2024-04-03T00:48:07.121800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This will take a few minutes\nfor col in catcols:\n    data[col] = LabelEncoder().fit_transform(data[col])","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:48:07.124200Z","iopub.execute_input":"2024-04-03T00:48:07.124509Z","iopub.status.idle":"2024-04-03T00:48:54.096789Z","shell.execute_reply.started":"2024-04-03T00:48:07.124484Z","shell.execute_reply":"2024-04-03T00:48:54.095434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.fillna(0, inplace=True)\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:48:54.098068Z","iopub.execute_input":"2024-04-03T00:48:54.098408Z","iopub.status.idle":"2024-04-03T00:49:01.902141Z","shell.execute_reply.started":"2024-04-03T00:48:54.098379Z","shell.execute_reply":"2024-04-03T00:49:01.900951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def gini_stability(preds, y_true):\n    aucs = roc_auc_score(y_true, preds)\n    # gini_in_time\n    batch_size = len(y_true)\n    x = np.arange(batch_size, dtype=np.float32)\n    # calculate a & b\n    a = np.sum((x - np.mean(x)) * (aucs - np.mean(aucs))) / np.sum(np.square(x - np.mean(x)))\n    b = np.mean(aucs) - a * np.mean(x)\n    # linear regression with a & b\n    y_hat = a * x + b\n    residuals = aucs - y_hat\n    res_std = np.sqrt(np.mean(np.square(residuals))) \n    avg_aucs = np.mean(aucs) \n    return avg_aucs * 2 - 1 + 88 * min(0, a) -0.5 * res_std","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:49:01.903636Z","iopub.execute_input":"2024-04-03T00:49:01.903943Z","iopub.status.idle":"2024-04-03T00:49:01.910455Z","shell.execute_reply.started":"2024-04-03T00:49:01.903918Z","shell.execute_reply":"2024-04-03T00:49:01.909392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {\n    'learning_rate':  0.1, \n    'num_leaves': 64,\n    'boosting':  'gbdt', \n    'objective': 'binary', \n    'metric': 'auc',\n    'feature_fraction': .9,\n    'seed': 42,\n    'max_depth': -1, \n    'boost_from_average': True,\n}","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:49:01.915925Z","iopub.execute_input":"2024-04-03T00:49:01.916319Z","iopub.status.idle":"2024-04-03T00:49:01.923692Z","shell.execute_reply.started":"2024-04-03T00:49:01.916287Z","shell.execute_reply":"2024-04-03T00:49:01.922554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"MAX_ROUNDS = 60\nlgb_iter1 = []\nlgb_ap1 = []\ngini = []\n\nmodel_name = None\n\nif os.path.exists('lgb_model.txt'):  \n    model_name = 'lgb_model.txt'\n\nkfold = KFold(n_splits=5, shuffle=True, random_state=42)\nfor train_idx, valid_idx in kfold.split(data, y):\n    X_train, y_train = data.iloc[train_idx], y.iloc[train_idx]\n    X_valid, y_valid = data.iloc[valid_idx], y.iloc[valid_idx]   \n    # LightGBM\n    lgb_gbm = lgb.train(params, \n                        lgb.Dataset(X_train, label=y_train), \n                        MAX_ROUNDS, \n                        lgb.Dataset(X_valid, label=y_valid), \n                        keep_training_booster=True,\n                        init_model=model_name\n                        )\n    print(\"Best iteration lgb = \", lgb_gbm.best_iteration)\n    lgb_iter1 = np.append(lgb_iter1, lgb_gbm.best_iteration)\n\n    pred = lgb_gbm.predict(X_valid, num_iteration=lgb_gbm.best_iteration)\n    ap = roc_auc_score (y_valid, pred)\n    print('valid AUC: ', ap)\n    lgb_ap1 = np.append(lgb_ap1, ap)  \n    \n    gi = gini_stability(pred, y_valid)\n    print('valid Gini: ', gi)\n    gini = np.append(gini, gi)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:49:01.925247Z","iopub.execute_input":"2024-04-03T00:49:01.925631Z","iopub.status.idle":"2024-04-03T00:56:50.637565Z","shell.execute_reply.started":"2024-04-03T00:49:01.925601Z","shell.execute_reply":"2024-04-03T00:56:50.636424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('lgb_iter1: ', lgb_iter1)\nprint('lgb_ap1: ', lgb_ap1)\nprint('gini: ', gini)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:56:50.639107Z","iopub.execute_input":"2024-04-03T00:56:50.639559Z","iopub.status.idle":"2024-04-03T00:56:50.648840Z","shell.execute_reply.started":"2024-04-03T00:56:50.639523Z","shell.execute_reply":"2024-04-03T00:56:50.647885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgb_gbm.save_model('lgb_model.txt', num_iteration=np.argmax(gini))","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:56:50.651008Z","iopub.execute_input":"2024-04-03T00:56:50.651717Z","iopub.status.idle":"2024-04-03T00:56:50.661615Z","shell.execute_reply.started":"2024-04-03T00:56:50.651653Z","shell.execute_reply":"2024-04-03T00:56:50.660455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_importance = lgb_gbm.feature_importance(importance_type='gain')\nfeats = pd.DataFrame({\n    \"name\": data.columns,\n    \"importance\": feature_importance\n})\nfeats","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:56:50.663364Z","iopub.execute_input":"2024-04-03T00:56:50.663788Z","iopub.status.idle":"2024-04-03T00:56:50.678377Z","shell.execute_reply.started":"2024-04-03T00:56:50.663754Z","shell.execute_reply":"2024-04-03T00:56:50.677274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from lightgbm import plot_importance\n\nplot_importance(lgb_gbm,max_num_features=60, figsize=(6, 12))\n","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:56:50.680083Z","iopub.execute_input":"2024-04-03T00:56:50.680431Z","iopub.status.idle":"2024-04-03T00:56:51.848436Z","shell.execute_reply.started":"2024-04-03T00:56:50.680397Z","shell.execute_reply":"2024-04-03T00:56:51.847337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"featsgz = feats[feats['importance'] > 0]\ndisplay(featsgz.shape)\nfeatsgz.to_csv('importances-gt0.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:56:51.849552Z","iopub.execute_input":"2024-04-03T00:56:51.849848Z","iopub.status.idle":"2024-04-03T00:56:51.863184Z","shell.execute_reply.started":"2024-04-03T00:56:51.849825Z","shell.execute_reply":"2024-04-03T00:56:51.862158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# lgb_gbm = lgb.Booster(model_file='/kaggle/input/lgbmodel/lgb_model.txt')\npreds = lgb_gbm.predict(data)","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:56:51.864403Z","iopub.execute_input":"2024-04-03T00:56:51.864811Z","iopub.status.idle":"2024-04-03T00:57:04.598412Z","shell.execute_reply.started":"2024-04-03T00:56:51.864777Z","shell.execute_reply":"2024-04-03T00:57:04.597419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({\n    \"case_id\": base[\"case_id\"].astype(int),\n    \"score\": preds\n})\n\nsubmission.to_csv('submission.csv', float_format='%.4f', index=False)\nprint('saved submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-04-03T00:57:04.599625Z","iopub.execute_input":"2024-04-03T00:57:04.599937Z","iopub.status.idle":"2024-04-03T00:57:08.334890Z","shell.execute_reply.started":"2024-04-03T00:57:04.599913Z","shell.execute_reply":"2024-04-03T00:57:08.333661Z"},"trusted":true},"execution_count":null,"outputs":[]}]}