{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"}],"dockerImageVersionId":30635,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Example Notebook (copy from source) \n\nWelcome to the example notebook for the Home Credit Kaggle competition. The goal of this competition is to determine how likely a customer is going to default on an issued loan. The main difference between the [first](https://www.kaggle.com/c/home-credit-default-risk) and this competition is that now your submission will be scored with a custom metric that will take into account how well the model performs in future. A decline in performance will be penalized. The goal is to create a model that is stable and performs well in the future.\n\nIn this notebook you will see how to:\n* Load the data\n* Join tables with Polars - a DataFrame library implemented in Rust language, designed to be blazingy fast and memory efficient.  \n* Create simple aggregation features\n* Train a LightGBM model\n* Create a submission table\n\n## Load the data (copy from source) ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import polars as pl\nimport numpy as np\nimport pandas as pd\nimport lightgbm as lgb\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score \n\ndataPath = \"/kaggle/input/home-credit-credit-risk-model-stability/\"","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:03.172298Z","iopub.execute_input":"2024-03-17T16:43:03.172561Z","iopub.status.idle":"2024-03-17T16:43:09.563748Z","shell.execute_reply.started":"2024-03-17T16:43:03.172535Z","shell.execute_reply":"2024-03-17T16:43:09.562777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_table_dtypes(df: pl.DataFrame) -> pl.DataFrame:\n    # implement here all desired dtypes for tables\n    # the following is just an example\n    for col in df.columns:\n        # last letter of column name will help you determine the type\n        if col[-1] in (\"P\", \"A\"):\n            df = df.with_columns(pl.col(col).cast(pl.Float64).alias(col))\n\n    return df\n\ndef convert_strings(df: pd.DataFrame) -> pd.DataFrame:\n    for col in df.columns:  \n        if df[col].dtype.name in ['object', 'string']:\n            df[col] = df[col].astype(\"string\").astype('category')\n            current_categories = df[col].cat.categories\n            new_categories = current_categories.to_list() + [\"Unknown\"]\n            new_dtype = pd.CategoricalDtype(categories=new_categories, ordered=True)\n            df[col] = df[col].astype(new_dtype)\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:09.565253Z","iopub.execute_input":"2024-03-17T16:43:09.565537Z","iopub.status.idle":"2024-03-17T16:43:09.573436Z","shell.execute_reply.started":"2024-03-17T16:43:09.565512Z","shell.execute_reply":"2024-03-17T16:43:09.572489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_basetable = pl.read_csv(dataPath + \"csv_files/train/train_base.csv\")\ntrain_static = pl.concat(\n    [\n        pl.read_csv(dataPath + \"csv_files/train/train_static_0_0.csv\").pipe(set_table_dtypes),\n        pl.read_csv(dataPath + \"csv_files/train/train_static_0_1.csv\").pipe(set_table_dtypes),\n    ],\n    how=\"vertical_relaxed\",\n)\ntrain_static_cb = pl.read_csv(dataPath + \"csv_files/train/train_static_cb_0.csv\").pipe(set_table_dtypes)\ntrain_person_1 = pl.read_csv(dataPath + \"csv_files/train/train_person_1.csv\").pipe(set_table_dtypes) \ntrain_credit_bureau_b_2 = pl.read_csv(dataPath + \"csv_files/train/train_credit_bureau_b_2.csv\").pipe(set_table_dtypes) ","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:09.574406Z","iopub.execute_input":"2024-03-17T16:43:09.574668Z","iopub.status.idle":"2024-03-17T16:43:24.565815Z","shell.execute_reply.started":"2024-03-17T16:43:09.574636Z","shell.execute_reply":"2024-03-17T16:43:24.564848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_basetable = pl.read_csv(dataPath + \"csv_files/test/test_base.csv\")\ntest_static = pl.concat(\n    [\n        pl.read_csv(dataPath + \"csv_files/test/test_static_0_0.csv\").pipe(set_table_dtypes),\n        pl.read_csv(dataPath + \"csv_files/test/test_static_0_1.csv\").pipe(set_table_dtypes),\n        pl.read_csv(dataPath + \"csv_files/test/test_static_0_2.csv\").pipe(set_table_dtypes),\n    ],\n    how=\"vertical_relaxed\",\n)\ntest_static_cb = pl.read_csv(dataPath + \"csv_files/test/test_static_cb_0.csv\").pipe(set_table_dtypes)\ntest_person_1 = pl.read_csv(dataPath + \"csv_files/test/test_person_1.csv\").pipe(set_table_dtypes) \ntest_credit_bureau_b_2 = pl.read_csv(dataPath + \"csv_files/test/test_credit_bureau_b_2.csv\").pipe(set_table_dtypes) ","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:24.568205Z","iopub.execute_input":"2024-03-17T16:43:24.568479Z","iopub.status.idle":"2024-03-17T16:43:24.657864Z","shell.execute_reply.started":"2024-03-17T16:43:24.568454Z","shell.execute_reply":"2024-03-17T16:43:24.657162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature engineering (copy from source) \n\nIn this part, we can see a simple example of joining tables via `case_id`. Here the loading and joining is done with polars library. Polars library is blazingly fast and has much smaller memory footprint than pandas. ","metadata":{}},{"cell_type":"code","source":"# We need to use aggregation functions in tables with depth > 1, so tables that contain num_group1 column or \n# also num_group2 column.\ntrain_person_1_feats_1 = train_person_1.group_by(\"case_id\").agg(\n    pl.col(\"mainoccupationinc_384A\").max().alias(\"mainoccupationinc_384A_max\"),\n    (pl.col(\"incometype_1044T\") == \"SELFEMPLOYED\").max().alias(\"mainoccupationinc_384A_any_selfemployed\")\n)\n\n# Here num_group1=0 has special meaning, it is the person who applied for the loan.\ntrain_person_1_feats_2 = train_person_1.select([\"case_id\", \"num_group1\", \"housetype_905L\"]).filter(\n    pl.col(\"num_group1\") == 0\n).drop(\"num_group1\").rename({\"housetype_905L\": \"person_housetype\"})\n\n# Here we have num_goup1 and num_group2, so we need to aggregate again.\ntrain_credit_bureau_b_2_feats = train_credit_bureau_b_2.group_by(\"case_id\").agg(\n    pl.col(\"pmts_pmtsoverdue_635A\").max().alias(\"pmts_pmtsoverdue_635A_max\"),\n    (pl.col(\"pmts_dpdvalue_108P\") > 31).max().alias(\"pmts_dpdvalue_108P_over31\")\n)\n\n# We will process in this examples only A-type and M-type columns, so we need to select them.\nselected_static_cols = []\nfor col in train_static.columns:\n    if col[-1] in (\"A\", \"M\"):\n        selected_static_cols.append(col)\nprint(selected_static_cols)\n\nselected_static_cb_cols = []\nfor col in train_static_cb.columns:\n    if col[-1] in (\"A\", \"M\"):\n        selected_static_cb_cols.append(col)\nprint(selected_static_cb_cols)\n\n# Join all tables together.\ndata = train_basetable.join(\n    train_static.select([\"case_id\"]+selected_static_cols), how=\"left\", on=\"case_id\"\n).join(\n    train_static_cb.select([\"case_id\"]+selected_static_cb_cols), how=\"left\", on=\"case_id\"\n).join(\n    train_person_1_feats_1, how=\"left\", on=\"case_id\"\n).join(\n    train_person_1_feats_2, how=\"left\", on=\"case_id\"\n).join(\n    train_credit_bureau_b_2_feats, how=\"left\", on=\"case_id\"\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:24.658956Z","iopub.execute_input":"2024-03-17T16:43:24.659239Z","iopub.status.idle":"2024-03-17T16:43:26.050801Z","shell.execute_reply.started":"2024-03-17T16:43:24.659208Z","shell.execute_reply":"2024-03-17T16:43:26.050042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_person_1_feats_1 = test_person_1.group_by(\"case_id\").agg(\n    pl.col(\"mainoccupationinc_384A\").max().alias(\"mainoccupationinc_384A_max\"),\n    (pl.col(\"incometype_1044T\") == \"SELFEMPLOYED\").max().alias(\"mainoccupationinc_384A_any_selfemployed\")\n)\n\ntest_person_1_feats_2 = test_person_1.select([\"case_id\", \"num_group1\", \"housetype_905L\"]).filter(\n    pl.col(\"num_group1\") == 0\n).drop(\"num_group1\").rename({\"housetype_905L\": \"person_housetype\"})\n\ntest_credit_bureau_b_2_feats = test_credit_bureau_b_2.group_by(\"case_id\").agg(\n    pl.col(\"pmts_pmtsoverdue_635A\").max().alias(\"pmts_pmtsoverdue_635A_max\"),\n    (pl.col(\"pmts_dpdvalue_108P\") > 31).max().alias(\"pmts_dpdvalue_108P_over31\")\n)\n\ndata_submission = test_basetable.join(\n    test_static.select([\"case_id\"]+selected_static_cols), how=\"left\", on=\"case_id\"\n).join(\n    test_static_cb.select([\"case_id\"]+selected_static_cb_cols), how=\"left\", on=\"case_id\"\n).join(\n    test_person_1_feats_1, how=\"left\", on=\"case_id\"\n).join(\n    test_person_1_feats_2, how=\"left\", on=\"case_id\"\n).join(\n    test_credit_bureau_b_2_feats, how=\"left\", on=\"case_id\"\n)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:26.051896Z","iopub.execute_input":"2024-03-17T16:43:26.052192Z","iopub.status.idle":"2024-03-17T16:43:26.062669Z","shell.execute_reply.started":"2024-03-17T16:43:26.052167Z","shell.execute_reply":"2024-03-17T16:43:26.061935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"case_ids = data[\"case_id\"].unique().shuffle(seed=1)\ncase_ids_train, case_ids_test = train_test_split(case_ids, train_size=0.6, random_state=1)\ncase_ids_valid, case_ids_test = train_test_split(case_ids_test, train_size=0.5, random_state=1)\n\ncols_pred = []\nfor col in data.columns:\n    if col[-1].isupper() and col[:-1].islower():\n        cols_pred.append(col)\n\nprint(cols_pred)\n\ndef from_polars_to_pandas(case_ids: pl.DataFrame) -> pl.DataFrame:\n    return (\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[[\"case_id\", \"WEEK_NUM\", \"target\"]].to_pandas(),\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[cols_pred].to_pandas(),\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[\"target\"].to_pandas()\n    )\n\nbase_train, X_train, y_train = from_polars_to_pandas(case_ids_train)\nbase_valid, X_valid, y_valid = from_polars_to_pandas(case_ids_valid)\nbase_test, X_test, y_test = from_polars_to_pandas(case_ids_test)\n\nfor df in [X_train, X_valid, X_test]:\n    df = convert_strings(df)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:26.063764Z","iopub.execute_input":"2024-03-17T16:43:26.064271Z","iopub.status.idle":"2024-03-17T16:43:33.335315Z","shell.execute_reply.started":"2024-03-17T16:43:26.064246Z","shell.execute_reply":"2024-03-17T16:43:33.334489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_submission = data_submission[cols_pred].to_pandas()\nX_submission = convert_strings(X_submission)\ncategorical_cols = X_train.select_dtypes(include=['category']).columns\n\nfor col in categorical_cols:\n    train_categories = set(X_train[col].cat.categories)\n    submission_categories = set(X_submission[col].cat.categories)\n    new_categories = submission_categories - train_categories\n    X_submission.loc[X_submission[col].isin(new_categories), col] = \"Unknown\"\n    new_dtype = pd.CategoricalDtype(categories=train_categories, ordered=True)\n    X_train[col] = X_train[col].astype(new_dtype)\n    X_submission[col] = X_submission[col].astype(new_dtype)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:09:43.777349Z","iopub.execute_input":"2024-03-17T18:09:43.777710Z","iopub.status.idle":"2024-03-17T18:09:43.840607Z","shell.execute_reply.started":"2024-03-17T18:09:43.777679Z","shell.execute_reply":"2024-03-17T18:09:43.839844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Until here everything was done same as in the example notebook. Now lets take a closer look to the features we have.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport math\n\ndf =  X_train\nn_cols = len(df.columns)\nn_rows = int(n_cols**0.5)\nn_cols_plot = n_cols // n_rows + 1\n\nfig, axes = plt.subplots(n_rows, n_cols_plot, figsize=(30, 20))\naxes = axes.flatten()\nfor i, col in enumerate(df.columns):\n    nan_percentage = df[col].isna().mean() * 100\n    not_nan_percentage = 100 - nan_percentage\n    ax = axes[i]\n    ax.bar(['NaN', 'Not NaN'], [nan_percentage, not_nan_percentage])\n    ax.set_title(col)\n    ax.set_ylim(0, 100)  # Set y-axis limit to 0-100 for percentage\n    ax.set_ylabel('Percentage')\n\nfor j in range(i + 1, len(axes)):\n    axes[j].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:43:33.336391Z","iopub.execute_input":"2024-03-17T16:43:33.336672Z","iopub.status.idle":"2024-03-17T16:43:40.701747Z","shell.execute_reply.started":"2024-03-17T16:43:33.336647Z","shell.execute_reply":"2024-03-17T16:43:40.700753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport math\n\ndf =  X_test\nn_cols = len(df.columns)\nn_rows = int(n_cols**0.5)\nn_cols_plot = n_cols // n_rows + 1\n\nfig, axes = plt.subplots(n_rows, n_cols_plot, figsize=(30, 20))\naxes = axes.flatten()\nfor i, col in enumerate(df.columns):\n    nan_percentage = df[col].isna().mean() * 100\n    not_nan_percentage = 100 - nan_percentage\n    ax = axes[i]\n    ax.bar(['NaN', 'Not NaN'], [nan_percentage, not_nan_percentage])\n    ax.set_title(col)\n    ax.set_ylim(0, 100)  # Set y-axis limit to 0-100 for percentage\n    ax.set_ylabel('Percentage')\n\nfor j in range(i + 1, len(axes)):\n    axes[j].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T17:31:18.966587Z","iopub.execute_input":"2024-03-17T17:31:18.967038Z","iopub.status.idle":"2024-03-17T17:31:25.968763Z","shell.execute_reply.started":"2024-03-17T17:31:18.967007Z","shell.execute_reply":"2024-03-17T17:31:25.967748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport math\n\ndf =  X_valid\nn_cols = len(df.columns)\nn_rows = int(n_cols**0.5)\nn_cols_plot = n_cols // n_rows + 1\n\nfig, axes = plt.subplots(n_rows, n_cols_plot, figsize=(30, 20))\naxes = axes.flatten()\nfor i, col in enumerate(df.columns):\n    nan_percentage = df[col].isna().mean() * 100\n    not_nan_percentage = 100 - nan_percentage\n    ax = axes[i]\n    ax.bar(['NaN', 'Not NaN'], [nan_percentage, not_nan_percentage])\n    ax.set_title(col)\n    ax.set_ylim(0, 100)  # Set y-axis limit to 0-100 for percentage\n    ax.set_ylabel('Percentage')\n\nfor j in range(i + 1, len(axes)):\n    axes[j].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T17:31:35.254282Z","iopub.execute_input":"2024-03-17T17:31:35.254641Z","iopub.status.idle":"2024-03-17T17:31:42.186038Z","shell.execute_reply.started":"2024-03-17T17:31:35.254611Z","shell.execute_reply":"2024-03-17T17:31:42.185074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we can see the percentage of NAN-s in data.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport math\n\ndf =  X_train\nn_cols = len(df.columns)\nn_rows = int(n_cols**0.5)\nn_cols_plot = n_cols // n_rows + 1\nfig, axes = plt.subplots(n_rows, n_cols_plot, figsize=(30, 20))\n\naxes = axes.flatten()\nfor i, col in enumerate(df.columns):\n    ax = axes[i]\n    if df[col].dtype=='float64' and np.min(df[col])>-1:\n        ax.hist(np.log(X_train[col]+1), bins=200)\n    elif df[col].dtype=='float64':\n        ax.hist(X_train[col], bins=200)\n    else:\n        ax.hist(X_train[col].astype(str), bins=200)\n    ax.set_title(col)\n    ax.set_ylabel('freq')\n\nfor j in range(i + 1, len(axes)):\n    axes[j].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T16:47:33.995970Z","iopub.execute_input":"2024-03-17T16:47:33.996620Z","iopub.status.idle":"2024-03-17T16:48:09.406430Z","shell.execute_reply.started":"2024-03-17T16:47:33.996590Z","shell.execute_reply":"2024-03-17T16:48:09.405446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport math\n\ndf =  X_test\nn_cols = len(df.columns)\nn_rows = int(n_cols**0.5)\nn_cols_plot = n_cols // n_rows + 1\nfig, axes = plt.subplots(n_rows, n_cols_plot, figsize=(30, 20))\n\naxes = axes.flatten()\nfor i, col in enumerate(df.columns):\n    ax = axes[i]\n    if df[col].dtype=='float64' and np.min(df[col])>-1:\n        ax.hist(np.log(X_train[col]+1), bins=200)\n    elif df[col].dtype=='float64':\n        ax.hist(X_train[col], bins=200)\n    else:\n        ax.hist(X_train[col].astype(str), bins=200)\n    ax.set_title(col)\n    ax.set_ylabel('freq')\n\nfor j in range(i + 1, len(axes)):\n    axes[j].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T17:31:42.567477Z","iopub.execute_input":"2024-03-17T17:31:42.568322Z","iopub.status.idle":"2024-03-17T17:32:19.129268Z","shell.execute_reply.started":"2024-03-17T17:31:42.568286Z","shell.execute_reply":"2024-03-17T17:32:19.128310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport math\n\ndf =  X_valid\nn_cols = len(df.columns)\nn_rows = int(n_cols**0.5)\nn_cols_plot = n_cols // n_rows + 1\nfig, axes = plt.subplots(n_rows, n_cols_plot, figsize=(30, 20))\n\naxes = axes.flatten()\nfor i, col in enumerate(df.columns):\n    ax = axes[i]\n    if df[col].dtype=='float64' and np.min(df[col])>-1:\n        ax.hist(np.log(X_train[col]+1), bins=200)\n    elif df[col].dtype=='float64':\n        ax.hist(X_train[col], bins=200)\n    else:\n        ax.hist(X_train[col].astype(str), bins=200)\n    ax.set_title(col)\n    ax.set_ylabel('freq')\n\nfor j in range(i + 1, len(axes)):\n    axes[j].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-17T17:32:19.131053Z","iopub.execute_input":"2024-03-17T17:32:19.131518Z","iopub.status.idle":"2024-03-17T17:32:55.319039Z","shell.execute_reply.started":"2024-03-17T17:32:19.131484Z","shell.execute_reply":"2024-03-17T17:32:55.318060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are graphics of feature distribution (without nans). Fully positive features are in log-scale.","metadata":{}},{"cell_type":"markdown","source":"What can we see from these graphs:\n* Some columns of the dataframe contain an enourmous number of nans. That is why I thougth that a possible improvement of the given example method is using models to pedict these nans, and not just fill them with default values or give all this work to our model. Mu solution contains models that predict those NANs\n* Some numerical features have a value that can be considered as a categorical feature - it is always zero, you can see the percentage of those on some graphs.\n\nSo my final solution to generate some new usefull features looks like this\n\n1. For each column that contains nans I learnt a catboost (regressor if numerical, classifier if categorical) to predict a nan value from other columns, filling nans in them with default value. These features will be called 'step_1_features'\n2. For each column with nans I learnt a second catboost that predicts a nan value, but using both initial columns and columns with predicts from first catboost.\n3. I also learnt catboost to predict value zero for columns where it is neccessary. It helped to fill nans with it as well.\n4. For the final model features that were predicted from all these 3 catboosts were used for trainig a lgb model\n\nAll these catoboosts had near to default parameters - what could have been done additionally is hyperparameter optimization for them, but I had no time.","metadata":{}},{"cell_type":"code","source":"from catboost import CatBoostClassifier, CatBoostRegressor, Pool, cv\nfrom sklearn.metrics import accuracy_score\n\nfrom ipywidgets import interact  \nimport ipywidgets as widgets ","metadata":{"execution":{"iopub.status.busy":"2024-03-17T17:35:56.970939Z","iopub.execute_input":"2024-03-17T17:35:56.971639Z","iopub.status.idle":"2024-03-17T17:35:56.976625Z","shell.execute_reply.started":"2024-03-17T17:35:56.971597Z","shell.execute_reply":"2024-03-17T17:35:56.975555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm \n\ndef catboost(pool, is_cat=False):\n    if is_cat:\n        model = CatBoostClassifier(\n            task_type=\"GPU\",\n            devices='0',\n            iterations=100,\n            learning_rate=1,\n            depth=10, \n            verbose=False\n        )\n        model.fit(pool)\n        return model\n    \n    model = CatBoostRegressor(\n        task_type=\"GPU\",\n        devices='0',\n        iterations=100,\n        learning_rate=1,\n        depth=10, \n        verbose=False\n    )\n    model.fit(pool)\n    return model\n\n\ndef fill_na_stupid(df):\n    df_copy = df.copy()\n    for col in df_copy:\n        dt = df_copy[col].dtype\n        if pd.api.types.is_numeric_dtype(dt):\n            df_copy[col] = df_copy[col].fillna(0)\n        else:\n            df_copy[col] = df_copy[col].astype('string').fillna('')\n    return df_copy\n\ndef cat_features_to_int(df, cat_features):\n    df_copy = df.copy()\n    for cat in cat_features:\n        df_copy[cat] = df_copy[cat].astype('string')\n        df_copy[cat] = df_copy[cat].fillna('Unknown')\n        vals = pd.unique(df_copy[cat])\n        d = {val:i for i, val in enumerate(vals)}\n        df_copy[cat] = df_copy[cat].apply(lambda x: d[x])\n        df_copy.loc[df[cat]==d.get('Unknown', -1), cat] = pd.NA\n    return df_copy\n\ndef fill_na_umom_fit_predict_first_step(df):\n    cols = list(df.columns)\n    cat_features = [name for name, dtype in df.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\n    df_copy = cat_features_to_int(df, cat_features)\n    \n    models = []\n    cur_df = df_copy.copy()\n    for col in tqdm(cols):\n        cond = df[col].isna()\n        if cond.sum() == 0:\n            models.append(None)\n            cur_df[col] = df_copy[col]\n            continue\n        x, y = fill_na_stupid(df_copy.drop(columns=col)), df_copy[col]\n        pool = Pool(\n            data=x[~cond], \n            label=y[~cond]\n        )\n        model = catboost(pool, col in cat_features)\n        models.append(model)\n        nan_pred = model.predict(\n            x[cond]\n        )\n        cur_df.loc[cond, col] = nan_pred.reshape(-1)\n    cur_df = cur_df.rename(mapper=lambda x: x+'_1', axis='columns')\n    return cur_df, models\n\ndef fill_na_umom_predict_first_step(df, models):\n    cols = list(df.columns)\n    cat_features = [name for name, dtype in df.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\n    df_copy = cat_features_to_int(df, cat_features)\n    cur_df = df_copy.copy()\n    for i, col in tqdm(enumerate(cols)):\n        cond = df[col].isna()\n        if cond.sum() == 0:\n            cur_df[col] = df_copy[col]\n            continue\n        x = fill_na_stupid(df_copy.drop(columns=col))\n        nan_pred = models[i].predict(\n            x[cond]\n        )\n        cur_df.loc[cond, col] = nan_pred.reshape(-1)\n    cur_df = cur_df.rename(mapper=lambda x: x+'_1', axis='columns')\n    return cur_df\n\ndef fill_na_umom_fit_predict_second_step(df, df2):\n    cols = list(df.columns)\n    cat_features = [name for name, dtype in df.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\n    df_copy = cat_features_to_int(df, cat_features)\n    \n    models = []\n    cur_df = df_copy.copy()\n    for col in tqdm(cols):\n        cond = df[col].isna()\n        if cond.sum() == 0:\n            models.append(None)\n            cur_df[col] = df_copy[col]\n            continue\n        x, y = pd.concat([fill_na_stupid(df_copy.drop(columns=col)), df2], axis=1), df_copy[col]\n        pool = Pool(\n            data=x[~cond], \n            label=y[~cond]\n        )\n        model = catboost(pool, col in cat_features)\n        models.append(model)\n        nan_pred = model.predict(\n            x[cond]\n        )\n        cur_df.loc[cond, col] = nan_pred.reshape(-1)\n    cur_df = cur_df.rename(mapper=lambda x: x+'_2', axis='columns')\n    return cur_df, models\n\ndef fill_na_umom_predict_second_step(df, df2, models):\n    cols = list(df.columns)\n    cat_features = [name for name, dtype in df.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\n    df_copy = cat_features_to_int(df, cat_features)\n    \n    cur_df = df_copy.copy()\n    for i, col in tqdm(enumerate(cols)):\n        cond = df[col].isna()\n        if cond.sum() == 0:\n            cur_df[col] = df_copy[col]\n            continue\n        x = pd.concat([fill_na_stupid(df_copy.drop(columns=col)), df2], axis=1)\n        nan_pred = models[i].predict(\n            x[cond]\n        )\n        cur_df.loc[cond, col] = nan_pred.reshape(-1)\n    cur_df = cur_df.rename(mapper=lambda x: x+'_2', axis='columns')\n    return cur_df","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:01:23.116125Z","iopub.execute_input":"2024-03-17T18:01:23.116914Z","iopub.status.idle":"2024-03-17T18:01:23.144874Z","shell.execute_reply.started":"2024-03-17T18:01:23.116878Z","shell.execute_reply":"2024-03-17T18:01:23.143960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_1step, models_1step =  fill_na_umom_fit_predict_first_step(X_train)\nX_train_2step, models_2step =  fill_na_umom_fit_predict_second_step(X_train, X_train_1step)\n\nX_val_1step = fill_na_umom_predict_first_step(X_valid, models_1step)\nX_val_2step = fill_na_umom_predict_second_step(X_valid, X_val_1step, models_2step)\n\nX_test_1step = fill_na_umom_predict_first_step(X_test, models_1step)\nX_test_2step = fill_na_umom_predict_second_step(X_test, X_test_1step, models_2step)\n\nX_submission_1step = fill_na_umom_predict_first_step(X_submission, models_1step)\nX_submission_2step = fill_na_umom_predict_second_step(X_submission, X_submission_1step, models_2step)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:10:30.089042Z","iopub.execute_input":"2024-03-17T18:10:30.089425Z","iopub.status.idle":"2024-03-17T18:17:42.815904Z","shell.execute_reply.started":"2024-03-17T18:10:30.089393Z","shell.execute_reply":"2024-03-17T18:17:42.814798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(3, 3, figsize=(20, 10))\n\naxes = axes.flatten()\naxes[0].hist(np.log(X_train[X_train.columns[15]]+1), bins=100);\naxes[0].set_title('column15 initial distribution')\naxes[1].hist(np.log(X_train_1step[X_train.columns[15]+'_1']+1), bins=100);\naxes[1].set_title('column15 nans predict distribution step1')\naxes[2].hist(np.log(X_train_2step[X_train.columns[15]+'_2']+1), bins=100);\naxes[2].set_title('column15 nans predict distribution step2')\n\naxes = axes.flatten()\naxes[3].hist(np.log(X_train[X_train.columns[0]]+1), bins=100);\naxes[3].set_title('column0 initial distribution')\naxes[4].hist(np.log(X_train_1step[X_train.columns[0]+'_1']+1), bins=100);\naxes[4].set_title('column0 nans predict distribution step1')\naxes[5].hist(np.log(X_train_2step[X_train.columns[0]+'_2']+1), bins=100);\naxes[5].set_title('column0 nans predict distribution step2')\n\naxes = axes.flatten()\naxes[6].hist(np.log(X_train[X_train.columns[4]]+1), bins=100);\naxes[6].set_title('column4 initial distribution')\naxes[7].hist(np.log(X_train_1step[X_train.columns[4]+'_1']+1), bins=100);\naxes[7].set_title('column4 nans predict distribution step1')\naxes[8].hist(np.log(X_train_2step[X_train.columns[4]+'_2']+1), bins=100);\naxes[8].set_title('column4 nans predict distribution step2')","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:17:45.277658Z","iopub.execute_input":"2024-03-17T18:17:45.278566Z","iopub.status.idle":"2024-03-17T18:17:48.648749Z","shell.execute_reply.started":"2024-03-17T18:17:45.278525Z","shell.execute_reply":"2024-03-17T18:17:48.647840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here you can see how distribuitions were changing after each step. You can see that shapes are close to each other (exapt column 0 after second step), but frequency changed a lot because we filled NANs.\n\nWhat we can also see that models didn't zeros for nans, that is why it should be done additionally.","metadata":{}},{"cell_type":"code","source":"def predict_zero_fit_predict(df):\n    cols = list(df.columns)\n    cat_features = [name for name, dtype in df.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\n    df_copy = cat_features_to_int(df, cat_features)\n    \n    models = []\n    cur_df = df_copy.copy()\n    for col in tqdm(cols):\n        cond = df[col].isna()\n        cond2 = df[col] == 0\n        if ((~cond) & (cond2)).sum() <10000 or df[col].dtype != 'float64':\n            models.append(None)\n            cur_df[col] = 0\n            continue\n        x, y = fill_na_stupid(df_copy[~cond].drop(columns=col)), df_copy[col][~cond]==0\n        pool = Pool( \n            data=x,\n            label=y\n        )\n        model = catboost(pool, True)\n        models.append(model)\n        zero_pred = model.predict(\n            fill_na_stupid(df_copy.drop(columns=col))\n        )\n        cur_df[col] = zero_pred.reshape(-1)\n    cur_df = cur_df.rename(mapper=lambda x: x+'_zero_pred', axis='columns')\n    return cur_df, models\n\ndef predict_zero_predict(df, models):\n    cols = list(df.columns)\n    cat_features = [name for name, dtype in df.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\n    df_copy = cat_features_to_int(df, cat_features)\n\n    cur_df = df_copy.copy()\n    for i, col in tqdm(enumerate(cols)):\n        if models[i] is None:\n            cur_df[col] = 0\n            continue\n        zero_pred = models[i].predict(\n            fill_na_stupid(df_copy.drop(columns=col))\n        )\n        cur_df[col] = zero_pred.reshape(-1)\n    cur_df = cur_df.rename(mapper=lambda x: x+'_zero_pred', axis='columns')\n    return cur_df","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:58:03.395714Z","iopub.execute_input":"2024-03-17T18:58:03.396110Z","iopub.status.idle":"2024-03-17T18:58:03.408461Z","shell.execute_reply.started":"2024-03-17T18:58:03.396076Z","shell.execute_reply":"2024-03-17T18:58:03.407678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_zero_pred, models_zero_pred =  predict_zero_fit_predict(X_train)\n\nX_val_zero_pred = predict_zero_predict(X_valid, models_zero_pred)\nX_test_zero_pred = predict_zero_predict(X_test, models_zero_pred)\nX_submission_zero_pred = predict_zero_predict(X_submission, models_zero_pred)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:58:04.767437Z","iopub.execute_input":"2024-03-17T18:58:04.767869Z","iopub.status.idle":"2024-03-17T18:58:32.110706Z","shell.execute_reply.started":"2024-03-17T18:58:04.767829Z","shell.execute_reply":"2024-03-17T18:58:32.109283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(20, 8))\n\naxes = axes.flatten()\naxes[0].hist((X_train[X_train.columns[0]][~X_train[X_train.columns[0]].isna()]==0).astype(str), bins=100);\naxes[0].set_title('column0==0 initial distribution')\naxes[1].hist(X_train_zero_pred[X_train.columns[0]+'_zero_pred'][~X_train[X_train.columns[0]].isna()], bins=100);\naxes[1].set_title('column0==0 predict on not nans')\naxes[2].hist(X_train_zero_pred[X_train.columns[0]+'_zero_pred'][X_train[X_train.columns[0]].isna()], bins=100);\naxes[2].set_title('column0==0 predict on nans')\n","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:39:19.196408Z","iopub.execute_input":"2024-03-17T18:39:19.197298Z","iopub.status.idle":"2024-03-17T18:39:21.018759Z","shell.execute_reply.started":"2024-03-17T18:39:19.197261Z","shell.execute_reply":"2024-03-17T18:39:21.017823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"here we can see how this zero prediction can help us will filling nans with zeros","metadata":{}},{"cell_type":"markdown","source":"Lets find out whether new features can be helpful","metadata":{}},{"cell_type":"code","source":"def gini_stability(base, w_fallingrate=88.0, w_resstd=-0.5):\n    gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n        .sort_values(\"WEEK_NUM\")\\\n        .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n        .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    return avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:53:41.220092Z","iopub.execute_input":"2024-03-17T18:53:41.220762Z","iopub.status.idle":"2024-03-17T18:53:41.227723Z","shell.execute_reply.started":"2024-03-17T18:53:41.220726Z","shell.execute_reply":"2024-03-17T18:53:41.226830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training LightGBM\n\nMinimal example of LightGBM training is shown below.","metadata":{}},{"cell_type":"code","source":"cat_features = [name for name, dtype in X_train.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\nlgb_train = lgb.Dataset(cat_features_to_int(fill_na_stupid(X_train), cat_features), label=y_train)\nlgb_valid = lgb.Dataset(cat_features_to_int(fill_na_stupid(X_valid), cat_features), label=y_valid, reference=lgb_train)\n\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 3,\n    \"num_leaves\": 31,\n    \"learning_rate\": 0.05,\n    \"feature_fraction\": 0.9,\n    \"bagging_fraction\": 0.8,\n    \"bagging_freq\": 5,\n    \"n_estimators\": 1000,\n    \"verbose\": -1,\n}\n\ngbm = lgb.train(\n    params,\n    lgb_train,\n    valid_sets=lgb_valid,\n    callbacks=[lgb.log_evaluation(50), lgb.early_stopping(10)]\n)\n\nfor base, X in [(base_train, X_train), (base_valid, X_valid), (base_test, X_test)]:\n    y_pred = gbm.predict(cat_features_to_int(fill_na_stupid(X), cat_features))\n    base[\"score\"] = y_pred\n\nprint('example metrics:\\n')\nprint(f'The AUC score on the train set is: {roc_auc_score(base_train[\"target\"], base_train[\"score\"])}') \nprint(f'The AUC score on the valid set is: {roc_auc_score(base_valid[\"target\"], base_valid[\"score\"])}') \nprint(f'The AUC score on the test set is: {roc_auc_score(base_test[\"target\"], base_test[\"score\"])}') \n\nstability_score_train = gini_stability(base_train)\nstability_score_valid = gini_stability(base_valid)\nstability_score_test = gini_stability(base_test)\n\nprint(f'The stability score on the train set is: {stability_score_train}') \nprint(f'The stability score on the valid set is: {stability_score_valid}') \nprint(f'The stability score on the test set is: {stability_score_test}')  ","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:53:50.371685Z","iopub.execute_input":"2024-03-17T18:53:50.372069Z","iopub.status.idle":"2024-03-17T18:53:51.948604Z","shell.execute_reply.started":"2024-03-17T18:53:50.372038Z","shell.execute_reply":"2024-03-17T18:53:51.947641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_zero_pred.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:56:58.998579Z","iopub.execute_input":"2024-03-17T18:56:58.998976Z","iopub.status.idle":"2024-03-17T18:56:59.005086Z","shell.execute_reply.started":"2024-03-17T18:56:58.998946Z","shell.execute_reply":"2024-03-17T18:56:59.004046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_concat = pd.concat([X_train, X_train_1step, X_train_2step, X_train_zero_pred], axis=1)\nX_valid_concat = pd.concat([X_valid, X_val_1step, X_val_2step, X_val_zero_pred], axis=1)\nX_test_concat = pd.concat([X_test, X_test_1step, X_test_2step, X_test_zero_pred], axis=1)\ncat_features = [name for name, dtype in X_train_concat.dtypes.to_dict().items() if not pd.api.types.is_numeric_dtype(dtype)]\n\nlgb_train = lgb.Dataset(cat_features_to_int(fill_na_stupid(X_train_concat), cat_features), label=y_train)\nlgb_valid = lgb.Dataset(cat_features_to_int(fill_na_stupid(X_valid_concat), cat_features), label=y_valid, reference=lgb_train)\n\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 3,\n    \"num_leaves\": 31,\n    \"learning_rate\": 0.05,\n    \"feature_fraction\": 0.9,\n    \"bagging_fraction\": 0.8,\n    \"bagging_freq\": 5,\n    \"n_estimators\": 1000,\n    \"verbose\": -1,\n}\n\ngbm = lgb.train(\n    params,\n    lgb_train,\n    valid_sets=lgb_valid,\n    callbacks=[lgb.log_evaluation(50), lgb.early_stopping(10)]\n)\n\nfor base, X in [(base_train, X_train_concat), (base_valid, X_valid_concat), (base_test, X_test_concat)]:\n    y_pred = gbm.predict(cat_features_to_int(fill_na_stupid(X), cat_features))\n    base[\"score\"] = y_pred\n\nprint('example metrics:\\n')\nprint(f'The AUC score on the train set is: {roc_auc_score(base_train[\"target\"], base_train[\"score\"])}') \nprint(f'The AUC score on the valid set is: {roc_auc_score(base_valid[\"target\"], base_valid[\"score\"])}') \nprint(f'The AUC score on the test set is: {roc_auc_score(base_test[\"target\"], base_test[\"score\"])}') \n\nstability_score_train = gini_stability(base_train)\nstability_score_valid = gini_stability(base_valid)\nstability_score_test = gini_stability(base_test)\n\nprint(f'The stability score on the train set is: {stability_score_train}') \nprint(f'The stability score on the valid set is: {stability_score_valid}') \nprint(f'The stability score on the test set is: {stability_score_test}')  ","metadata":{"execution":{"iopub.status.busy":"2024-03-17T18:58:47.108969Z","iopub.execute_input":"2024-03-17T18:58:47.109330Z","iopub.status.idle":"2024-03-17T19:02:00.355083Z","shell.execute_reply.started":"2024-03-17T18:58:47.109301Z","shell.execute_reply":"2024-03-17T19:02:00.354123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It actually got worse even on train. \n\nDont now how - i just added new features not removing old ones.\n\nAnyway, for the task score wasnt so neccessary, moreover Im sure my approach can be continued and can break at least baseline from example notebook.\n\nIn general, extra feature engineering was used in dealing with NANs with an implementation of many Catboosts.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission\n\nScoring the submission dataset is below, we need to take care of new categories. Then we save the score as a last step. ","metadata":{}},{"cell_type":"code","source":"X_submission_concat = pd.concat([X_submission, X_submission_1step, X_submission_2step, X_submission_zero_pred], axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T19:03:44.741062Z","iopub.execute_input":"2024-03-17T19:03:44.742053Z","iopub.status.idle":"2024-03-17T19:03:44.748808Z","shell.execute_reply.started":"2024-03-17T19:03:44.742012Z","shell.execute_reply":"2024-03-17T19:03:44.747645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_submission_pred = gbm.predict(cat_features_to_int(fill_na_stupid(X_submission_concat), cat_features), num_iteration=gbm.best_iteration)","metadata":{"execution":{"iopub.status.busy":"2024-03-17T19:04:07.320749Z","iopub.execute_input":"2024-03-17T19:04:07.321222Z","iopub.status.idle":"2024-03-17T19:04:07.429545Z","shell.execute_reply.started":"2024-03-17T19:04:07.321191Z","shell.execute_reply":"2024-03-17T19:04:07.428840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_submission.shape, X_train.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-17T19:04:09.543952Z","iopub.execute_input":"2024-03-17T19:04:09.544313Z","iopub.status.idle":"2024-03-17T19:04:09.552348Z","shell.execute_reply.started":"2024-03-17T19:04:09.544283Z","shell.execute_reply":"2024-03-17T19:04:09.551190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({\n    \"case_id\": data_submission[\"case_id\"].to_numpy(),\n    \"score\": y_submission_pred\n}).set_index('case_id')\nsubmission.to_csv(\"./submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-17T19:04:11.327265Z","iopub.execute_input":"2024-03-17T19:04:11.328246Z","iopub.status.idle":"2024-03-17T19:04:11.342443Z","shell.execute_reply.started":"2024-03-17T19:04:11.328211Z","shell.execute_reply":"2024-03-17T19:04:11.341217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Best of luck, and most importantly, enjoy the process of learning and discovery! \n\n<img src=\"https://i.imgur.com/obVWIBh.png\" alt=\"Image\" width=\"700\"/>","metadata":{}}]}