{"metadata":{"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"}],"dockerImageVersionId":30665,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.13"},"papermill":{"default_parameters":{},"duration":948.118349,"end_time":"2024-03-01T04:56:02.144870","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-03-01T04:40:14.026521","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Introduction\n\nThis competition is a challenging opportunity to develop a machine learning model that can accurately predict loan defaults while remaining stable over time. This is an important task for lenders, as it allows them to assess the risk of potential clients more accurately and make better lending decisions.\nThe competition is open to data scientists and machine learning practitioners of all levels.\n\n## The challenge\n\nThe goal of the competition is to develop a model that can predict whether a borrower will default on a loan. The model must be stable over time, meaning that its performance should not degrade significantly as new data becomes available.\n\nThe competition dataset includes a variety of features about borrowers, such as their demographics, financial history, and loan characteristics.\n\n## Installing Libraries\nWe install the keras3 and wandb libraries which are essential for this notebook.","metadata":{"papermill":{"duration":0.009056,"end_time":"2024-03-01T04:40:17.084040","exception":false,"start_time":"2024-03-01T04:40:17.074984","status":"completed"},"tags":[]}},{"cell_type":"code","source":"!pip install -q --upgrade keras wandb","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:40:17.103224Z","iopub.status.busy":"2024-03-01T04:40:17.102799Z","iopub.status.idle":"2024-03-01T04:40:32.433873Z","shell.execute_reply":"2024-03-01T04:40:32.432359Z"},"papermill":{"duration":15.343793,"end_time":"2024-03-01T04:40:32.436676","exception":false,"start_time":"2024-03-01T04:40:17.092883","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Importing Libraries\nNow we import the required libraries and also login to the Wandb library in order to save our logs. Weights and Biases (WandB) is a comprehensive platform dedicated to streamlining the machine learning experimentation and research process. It provides a range of features that empower data scientists and researchers to efficiently manage and analyze their experiments including experiment tracking, visualisation tools and hyperparameter optimisation.","metadata":{"papermill":{"duration":0.008545,"end_time":"2024-03-01T04:40:32.453420","exception":false,"start_time":"2024-03-01T04:40:32.444875","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import polars as pl\nimport numpy as np\nimport pandas as pd\nimport lightgbm as lgb\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import roc_auc_score \nimport keras\nfrom keras import layers\nfrom keras import ops\n\nimport math\nimport numpy as np\nimport pandas as pd\nfrom tensorflow import data as tf_data\nimport matplotlib.pyplot as plt\n\ndataPath = \"/kaggle/input/home-credit-credit-risk-model-stability/\"","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:40:32.473167Z","iopub.status.busy":"2024-03-01T04:40:32.472747Z","iopub.status.idle":"2024-03-01T04:40:51.393297Z","shell.execute_reply":"2024-03-01T04:40:51.392345Z"},"papermill":{"duration":18.933717,"end_time":"2024-03-01T04:40:51.395897","exception":false,"start_time":"2024-03-01T04:40:32.462180","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import wandb\nfrom kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nsecret_value_0 = user_secrets.get_secret(\"api_key\")\nwandb.login(key=secret_value_0)","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:40:51.414895Z","iopub.status.busy":"2024-03-01T04:40:51.414105Z","iopub.status.idle":"2024-03-01T04:40:54.234814Z","shell.execute_reply":"2024-03-01T04:40:54.233748Z"},"papermill":{"duration":2.832262,"end_time":"2024-03-01T04:40:54.237033","exception":false,"start_time":"2024-03-01T04:40:51.404771","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Generation\nNow we load and preprocess our data from the provided csv files for the competition.","metadata":{"papermill":{"duration":0.008728,"end_time":"2024-03-01T04:40:54.254729","exception":false,"start_time":"2024-03-01T04:40:54.246001","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def set_table_dtypes(df: pl.DataFrame) -> pl.DataFrame:\n    # implement here all desired dtypes for tables\n    # the following is just an example\n    for col in df.columns:\n        # last letter of column name will help you determine the type\n        if col[-1] in (\"P\", \"A\"):\n            df = df.with_columns(pl.col(col).cast(pl.Float64).alias(col))\n\n    return df\n\ndef convert_strings(df: pd.DataFrame) -> pd.DataFrame:\n    for col in df.columns:  \n        if df[col].dtype.name in ['object', 'string']:\n            df[col] = df[col].astype(\"string\").astype('category')\n            current_categories = df[col].cat.categories\n            new_categories = current_categories.to_list() + [\"Unknown\"]\n            new_dtype = pd.CategoricalDtype(categories=new_categories, ordered=True)\n            df[col] = df[col].astype(new_dtype)\n    return df","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:40:54.274441Z","iopub.status.busy":"2024-03-01T04:40:54.273603Z","iopub.status.idle":"2024-03-01T04:40:54.282438Z","shell.execute_reply":"2024-03-01T04:40:54.281480Z"},"papermill":{"duration":0.021431,"end_time":"2024-03-01T04:40:54.284660","exception":false,"start_time":"2024-03-01T04:40:54.263229","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_basetable = pl.read_csv(dataPath + \"csv_files/train/train_base.csv\")\ntrain_static = pl.concat(\n    [\n        pl.read_csv(dataPath + \"csv_files/train/train_static_0_0.csv\").pipe(set_table_dtypes),\n        pl.read_csv(dataPath + \"csv_files/train/train_static_0_1.csv\").pipe(set_table_dtypes),\n    ],\n    how=\"vertical_relaxed\",\n)\ntrain_static_cb = pl.read_csv(dataPath + \"csv_files/train/train_static_cb_0.csv\").pipe(set_table_dtypes)\ntrain_person_1 = pl.read_csv(dataPath + \"csv_files/train/train_person_1.csv\").pipe(set_table_dtypes) \ntrain_credit_bureau_b_2 = pl.read_csv(dataPath + \"csv_files/train/train_credit_bureau_b_2.csv\").pipe(set_table_dtypes) ","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:40:54.304892Z","iopub.status.busy":"2024-03-01T04:40:54.304581Z","iopub.status.idle":"2024-03-01T04:41:09.555666Z","shell.execute_reply":"2024-03-01T04:41:09.554391Z"},"papermill":{"duration":15.263829,"end_time":"2024-03-01T04:41:09.558513","exception":false,"start_time":"2024-03-01T04:40:54.294684","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature engineering\n\nIn this part, we can see a simple example of joining tables via `case_id`. Here the loading and joining is done with polars library. Polars library is blazingly fast and has much smaller memory footprint than pandas. ","metadata":{"papermill":{"duration":0.008795,"end_time":"2024-03-01T04:41:09.577016","exception":false,"start_time":"2024-03-01T04:41:09.568221","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# We need to use aggregation functions in tables with depth > 1, so tables that contain num_group1 column or \n# also num_group2 column.\ntrain_person_1_feats_1 = train_person_1.group_by(\"case_id\").agg(\n    pl.col(\"mainoccupationinc_384A\").max().alias(\"mainoccupationinc_384A_max\"),\n    (pl.col(\"incometype_1044T\") == \"SELFEMPLOYED\").max().alias(\"mainoccupationinc_384A_any_selfemployed\")\n)\n\n# Here num_group1=0 has special meaning, it is the person who applied for the loan.\ntrain_person_1_feats_2 = train_person_1.select([\"case_id\", \"num_group1\", \"housetype_905L\"]).filter(\n    pl.col(\"num_group1\") == 0\n).drop(\"num_group1\").rename({\"housetype_905L\": \"person_housetype\"})\n\n# Here we have num_goup1 and num_group2, so we need to aggregate again.\ntrain_credit_bureau_b_2_feats = train_credit_bureau_b_2.group_by(\"case_id\").agg(\n    pl.col(\"pmts_pmtsoverdue_635A\").max().alias(\"pmts_pmtsoverdue_635A_max\"),\n    (pl.col(\"pmts_dpdvalue_108P\") > 31).max().alias(\"pmts_dpdvalue_108P_over31\")\n)\n\n# We will process in this examples only A-type and M-type columns, so we need to select them.\nselected_static_cols = []\nfor col in train_static.columns:\n    if col[-1] in (\"A\", \"M\"):\n        selected_static_cols.append(col)\nprint(selected_static_cols)\n\nselected_static_cb_cols = []\nfor col in train_static_cb.columns:\n    if col[-1] in (\"A\", \"M\"):\n        selected_static_cb_cols.append(col)\nprint(selected_static_cb_cols)\n\n# Join all tables together.\ndata = train_basetable.join(\n    train_static.select([\"case_id\"]+selected_static_cols), how=\"left\", on=\"case_id\"\n).join(\n    train_static_cb.select([\"case_id\"]+selected_static_cb_cols), how=\"left\", on=\"case_id\"\n).join(\n    train_person_1_feats_1, how=\"left\", on=\"case_id\"\n).join(\n    train_person_1_feats_2, how=\"left\", on=\"case_id\"\n).join(\n    train_credit_bureau_b_2_feats, how=\"left\", on=\"case_id\"\n)","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:41:09.596484Z","iopub.status.busy":"2024-03-01T04:41:09.595634Z","iopub.status.idle":"2024-03-01T04:41:11.751426Z","shell.execute_reply":"2024-03-01T04:41:11.750471Z"},"papermill":{"duration":2.168309,"end_time":"2024-03-01T04:41:11.754017","exception":false,"start_time":"2024-03-01T04:41:09.585708","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"run = wandb.init(project=\"homecredit\")\ntable = wandb.Table(data=data.to_pandas()[:1000])\nrun.log({'data':table})\nrun.finish()","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:41:11.773515Z","iopub.status.busy":"2024-03-01T04:41:11.773088Z","iopub.status.idle":"2024-03-01T04:42:14.883181Z","shell.execute_reply":"2024-03-01T04:42:14.882172Z"},"papermill":{"duration":63.123205,"end_time":"2024-03-01T04:42:14.886120","exception":false,"start_time":"2024-03-01T04:41:11.762915","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we convert our generated data from polar to pandas dataframe for easy accessibility and to train our model","metadata":{"papermill":{"duration":0.00957,"end_time":"2024-03-01T04:42:14.906221","exception":false,"start_time":"2024-03-01T04:42:14.896651","status":"completed"},"tags":[]}},{"cell_type":"code","source":"case_ids = data[\"case_id\"].unique().shuffle(seed=1)\ncase_ids_train, case_ids_valid = train_test_split(case_ids, train_size=0.6, random_state=1)\n\ncols_pred = []\nfor col in data.columns:\n    if col[-1].isupper() and col[:-1].islower():\n        cols_pred.append(col)\n\nprint(cols_pred)\n\ndef from_polars_to_pandas(case_ids: pl.DataFrame) -> pl.DataFrame:\n    return (\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[[\"case_id\", \"WEEK_NUM\", \"target\"]].to_pandas(),\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[cols_pred].to_pandas(),\n        data.filter(pl.col(\"case_id\").is_in(case_ids))[\"target\"].to_pandas()\n    )\n\nbase_train, X_train, y_train = from_polars_to_pandas(case_ids_train)\nbase_valid, X_valid, y_valid = from_polars_to_pandas(case_ids_valid)\n\nfor df in [X_train, X_valid]:\n    df = convert_strings(df)","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:14.928791Z","iopub.status.busy":"2024-03-01T04:42:14.928044Z","iopub.status.idle":"2024-03-01T04:42:22.795966Z","shell.execute_reply":"2024-03-01T04:42:22.795081Z"},"papermill":{"duration":7.881887,"end_time":"2024-03-01T04:42:22.798513","exception":false,"start_time":"2024-03-01T04:42:14.916626","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Train: {X_train.shape}\")\nprint(f\"Valid: {X_valid.shape}\")","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:22.822666Z","iopub.status.busy":"2024-03-01T04:42:22.821912Z","iopub.status.idle":"2024-03-01T04:42:22.827448Z","shell.execute_reply":"2024-03-01T04:42:22.826294Z"},"papermill":{"duration":0.020392,"end_time":"2024-03-01T04:42:22.830012","exception":false,"start_time":"2024-03-01T04:42:22.809620","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"catcols = X_train.select_dtypes(include='category').columns.to_list()\nnumcols = X_train.select_dtypes(include='float64').columns.to_list()\nlen(catcols)+len(numcols)","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:22.851862Z","iopub.status.busy":"2024-03-01T04:42:22.851091Z","iopub.status.idle":"2024-03-01T04:42:22.963074Z","shell.execute_reply":"2024-03-01T04:42:22.961936Z"},"papermill":{"duration":0.125642,"end_time":"2024-03-01T04:42:22.965534","exception":false,"start_time":"2024-03-01T04:42:22.839892","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Configuration\nCreating the configuration to manage the dataset building, model creation and training process.","metadata":{"papermill":{"duration":0.009553,"end_time":"2024-03-01T04:42:22.985350","exception":false,"start_time":"2024-03-01T04:42:22.975797","status":"completed"},"tags":[]}},{"cell_type":"code","source":"LEARNING_RATE = 0.001\nWEIGHT_DECAY = 0.0001\nDROPOUT_RATE = 0.2\nBATCH_SIZE = 265\nNUM_EPOCHS = 2\n\nNUM_TRANSFORMER_BLOCKS = 3  # Number of transformer blocks.\nNUM_HEADS = 4  # Number of attention heads.\nEMBEDDING_DIMS = 16  # Embedding dimensions of the categorical features.\nMLP_HIDDEN_UNITS_FACTORS = [\n    2,\n    1,\n]  # MLP hidden layer units, as factors of the number of inputs.\nNUM_MLP_BLOCKS = 2  # Number of MLP blocks in the baseline model.","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:23.008268Z","iopub.status.busy":"2024-03-01T04:42:23.007538Z","iopub.status.idle":"2024-03-01T04:42:23.013545Z","shell.execute_reply":"2024-03-01T04:42:23.012476Z"},"papermill":{"duration":0.019916,"end_time":"2024-03-01T04:42:23.015493","exception":false,"start_time":"2024-03-01T04:42:22.995577","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we are encoding the categorical columns using the label encoder.","metadata":{"papermill":{"duration":0.009717,"end_time":"2024-03-01T04:42:23.035049","exception":false,"start_time":"2024-03-01T04:42:23.025332","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from sklearn.preprocessing import OneHotEncoder, LabelEncoder\nle = LabelEncoder()\n\n# Apply LabelEncoder to each categorical column\nfor col in catcols:\n    X_train[col] = le.fit_transform(X_train[col])\n    X_valid[col] = le.fit_transform(X_valid[col])\n","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:23.056800Z","iopub.status.busy":"2024-03-01T04:42:23.056138Z","iopub.status.idle":"2024-03-01T04:42:27.978784Z","shell.execute_reply":"2024-03-01T04:42:27.977958Z"},"papermill":{"duration":4.936115,"end_time":"2024-03-01T04:42:27.981087","exception":false,"start_time":"2024-03-01T04:42:23.044972","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p1 = pd.concat([X_train,X_valid])\nCATEGORICAL_FEATURES_WITH_VOCABULARY= {}\n\nfor col in catcols:\n    CATEGORICAL_FEATURES_WITH_VOCABULARY[col] = sorted(list(p1[col].unique()))","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:28.002527Z","iopub.status.busy":"2024-03-01T04:42:28.002138Z","iopub.status.idle":"2024-03-01T04:42:28.326684Z","shell.execute_reply":"2024-03-01T04:42:28.325827Z"},"papermill":{"duration":0.33743,"end_time":"2024-03-01T04:42:28.329068","exception":false,"start_time":"2024-03-01T04:42:27.991638","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"featcols = list(numcols) + list(catcols)\ndef create_model_inputs():\n    inputs = {}\n    for feature_name in featcols:\n        if feature_name in numcols:\n            inputs[feature_name] = layers.Input(\n                name=feature_name, shape=(), dtype=\"float32\"\n            )\n        else:\n            inputs[feature_name] = layers.Input(\n                name=feature_name, shape=(), dtype=\"float32\"\n            )\n    return inputs","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:28.350889Z","iopub.status.busy":"2024-03-01T04:42:28.350242Z","iopub.status.idle":"2024-03-01T04:42:28.356644Z","shell.execute_reply":"2024-03-01T04:42:28.355763Z"},"papermill":{"duration":0.019513,"end_time":"2024-03-01T04:42:28.358580","exception":false,"start_time":"2024-03-01T04:42:28.339067","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Encode features\nThe encode_inputs method returns encoded_categorical_feature_list and numerical_feature_list. We encode the categorical features as embeddings, using a fixed embedding_dims for all the features, regardless their vocabulary sizes. This is required for the Transformer model.","metadata":{"papermill":{"duration":0.009511,"end_time":"2024-03-01T04:42:28.377819","exception":false,"start_time":"2024-03-01T04:42:28.368308","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from functools import partial\ndef encode_inputs(inputs, embedding_dims):\n    encoded_categorical_feature_list = []\n    numerical_feature_list = []\n\n    for feature_name in inputs:\n        if feature_name in catcols:\n            vocabulary = CATEGORICAL_FEATURES_WITH_VOCABULARY[feature_name]\n            embedding = layers.Embedding(\n                input_dim=len(vocabulary), output_dim=embedding_dims\n            )\n\n            # Convert the index values to embedding representations.\n            encoded_categorical_feature = embedding(inputs[feature_name])\n            encoded_categorical_feature_list.append(encoded_categorical_feature)\n\n        else:\n            # Use the numerical features as-is.\n            numerical_feature = ops.expand_dims(inputs[feature_name], -1)\n            numerical_feature_list.append(numerical_feature)\n\n    return encoded_categorical_feature_list, numerical_feature_list\n","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:28.398963Z","iopub.status.busy":"2024-03-01T04:42:28.398679Z","iopub.status.idle":"2024-03-01T04:42:28.405411Z","shell.execute_reply":"2024-03-01T04:42:28.404636Z"},"papermill":{"duration":0.019548,"end_time":"2024-03-01T04:42:28.407490","exception":false,"start_time":"2024-03-01T04:42:28.387942","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Building\n\nThe TabTransformer model stands out as a novel and promising approach for tackling tabular data in supervised and semi-supervised learning tasks. The TabTransformer overcomes the limitations of traditional tabular models by leveraging the power of attention-based Transformers, a popular architecture in natural language processing (NLP). Here's how it works:\n\n* Embeddings: Tabular data is first converted into numerical representations called embeddings. This allows the model to understand the meaning behind each feature.\n* Transformer Layers: These layers use self-attention mechanisms to identify how important each feature is for every other feature, learning complex relationships and discovering hidden patterns.\n* MLP Layer: Finally, a Multi-Layer Perceptron (MLP) combines the information from the Transformer layers to make predictions.","metadata":{"papermill":{"duration":0.010212,"end_time":"2024-03-01T04:42:28.428188","exception":false,"start_time":"2024-03-01T04:42:28.417976","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\ndef create_mlp(hidden_units, dropout_rate, activation, normalization_layer, name=None):\n    mlp_layers = []\n    for units in hidden_units:\n        mlp_layers.append(normalization_layer()),\n        mlp_layers.append(layers.Dense(units, activation=activation))\n        mlp_layers.append(layers.Dropout(dropout_rate))\n\n    return keras.Sequential(mlp_layers, name=name)\n\ndef create_tabtransformer_classifier(\n    num_transformer_blocks,\n    num_heads,\n    embedding_dims,\n    mlp_hidden_units_factors,\n    dropout_rate,\n    use_column_embedding=False,\n):\n    # Create model inputs.\n    inputs = create_model_inputs()\n    # encode features.\n    encoded_categorical_feature_list, numerical_feature_list = encode_inputs(\n        inputs, embedding_dims\n    )\n    # Stack categorical feature embeddings for the Tansformer.\n    encoded_categorical_features = ops.stack(encoded_categorical_feature_list, axis=1)\n    # Concatenate numerical features.\n    numerical_features = layers.concatenate(numerical_feature_list)\n\n    # Add column embedding to categorical feature embeddings.\n    if use_column_embedding:\n        num_columns = encoded_categorical_features.shape[1]\n        column_embedding = layers.Embedding(\n            input_dim=num_columns, output_dim=embedding_dims\n        )\n        column_indices = ops.arange(start=0, stop=num_columns, step=1)\n        encoded_categorical_features = encoded_categorical_features + column_embedding(\n            column_indices\n        )\n\n    # Create multiple layers of the Transformer block.\n    for block_idx in range(num_transformer_blocks):\n        # Create a multi-head attention layer.\n        attention_output = layers.MultiHeadAttention(\n            num_heads=num_heads,\n            key_dim=embedding_dims,\n            dropout=dropout_rate,\n            name=f\"multihead_attention_{block_idx}\",\n        )(encoded_categorical_features, encoded_categorical_features)\n        # Skip connection 1.\n        x = layers.Add(name=f\"skip_connection1_{block_idx}\")(\n            [attention_output, encoded_categorical_features]\n        )\n        # Layer normalization 1.\n        x = layers.LayerNormalization(name=f\"layer_norm1_{block_idx}\", epsilon=1e-6)(x)\n        # Feedforward.\n        feedforward_output = create_mlp(\n            hidden_units=[embedding_dims],\n            dropout_rate=dropout_rate,\n            activation=keras.activations.gelu,\n            normalization_layer=partial(\n                layers.LayerNormalization, epsilon=1e-6\n            ),  # using partial to provide keyword arguments before initialization\n            name=f\"feedforward_{block_idx}\",\n        )(x)\n        # Skip connection 2.\n        x = layers.Add(name=f\"skip_connection2_{block_idx}\")([feedforward_output, x])\n        # Layer normalization 2.\n        encoded_categorical_features = layers.LayerNormalization(\n            name=f\"layer_norm2_{block_idx}\", epsilon=1e-6\n        )(x)\n\n    # Flatten the \"contextualized\" embeddings of the categorical features.\n    categorical_features = layers.Flatten()(encoded_categorical_features)\n    # Apply layer normalization to the numerical features.\n    numerical_features = layers.LayerNormalization(epsilon=1e-6)(numerical_features)\n    # Prepare the input for the final MLP block.\n    features = layers.concatenate([categorical_features, numerical_features])\n\n    # Compute MLP hidden_units.\n    mlp_hidden_units = [\n        factor * features.shape[-1] for factor in mlp_hidden_units_factors\n    ]\n    # Create final MLP.\n    features = create_mlp(\n        hidden_units=mlp_hidden_units,\n        dropout_rate=dropout_rate,\n        activation=keras.activations.selu,\n        normalization_layer=layers.BatchNormalization,\n        name=\"MLP\",\n    )(features)\n\n    # Add a sigmoid as a binary classifer.\n    outputs = layers.Dense(units=1, activation=\"sigmoid\", name=\"sigmoid\")(features)\n    model = keras.Model(inputs=inputs, outputs=outputs)\n    return model\n\n\ntabtransformer_model = create_tabtransformer_classifier(\n    num_transformer_blocks=NUM_TRANSFORMER_BLOCKS,\n    num_heads=NUM_HEADS,\n    embedding_dims=EMBEDDING_DIMS,\n    mlp_hidden_units_factors=MLP_HIDDEN_UNITS_FACTORS,\n    dropout_rate=DROPOUT_RATE,\n)","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:28.450550Z","iopub.status.busy":"2024-03-01T04:42:28.450113Z","iopub.status.idle":"2024-03-01T04:42:29.575406Z","shell.execute_reply":"2024-03-01T04:42:29.574619Z"},"papermill":{"duration":1.139742,"end_time":"2024-03-01T04:42:29.577734","exception":false,"start_time":"2024-03-01T04:42:28.437992","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Setting up training protocols\nAdamW: AdamW, standing for Adaptive Moment Estimation with Weight Decay, builds upon the popular Adam optimizer. It incorporates weight decay, a regularization technique that penalizes large weights during optimization, preventing overfitting and improving model generalizability. This \"W\" makes a significant difference, especially for tasks with sparse data or large models.\n\nBinary Crossentropy Loss: This loss function excels in binary classification tasks, where you aim to categorize outputs into only two classes (e.g., 0 or 1, cat or dog). It measures the difference between the model's predicted probabilities for each class and the actual label, penalizing predictions further away from the truth.","metadata":{"papermill":{"duration":0.009606,"end_time":"2024-03-01T04:42:29.597219","exception":false,"start_time":"2024-03-01T04:42:29.587613","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from wandb.keras import WandbMetricsLogger\nrun = wandb.init(project=\"homecredit\",name = 'model_training')\n\noptimizer = keras.optimizers.AdamW(\n        learning_rate=LEARNING_RATE, weight_decay=WEIGHT_DECAY\n\n    )\n\ntabtransformer_model.compile(\n    optimizer=optimizer,\n    loss=keras.losses.BinaryCrossentropy(),\n    metrics=[keras.metrics.BinaryAccuracy(name=\"accuracy\")],\n)\nprint(\"Start training the model...\")\nhistory = tabtransformer_model.fit(\n    [X_train[col] for col in X_train.columns],y_train, epochs=NUM_EPOCHS,batch_size=32, validation_data=([X_valid[col] for col in X_valid.columns],y_valid),callbacks=[WandbMetricsLogger()]\n)","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:42:29.617686Z","iopub.status.busy":"2024-03-01T04:42:29.617337Z","iopub.status.idle":"2024-03-01T04:55:52.894605Z","shell.execute_reply":"2024-03-01T04:55:52.893468Z"},"papermill":{"duration":804.27236,"end_time":"2024-03-01T04:55:53.878871","exception":false,"start_time":"2024-03-01T04:42:29.606511","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tabtransformer_model.save('model.keras')","metadata":{"execution":{"iopub.execute_input":"2024-03-01T04:55:55.761725Z","iopub.status.busy":"2024-03-01T04:55:55.760973Z","iopub.status.idle":"2024-03-01T04:55:56.047352Z","shell.execute_reply":"2024-03-01T04:55:56.046283Z"},"papermill":{"duration":1.254943,"end_time":"2024-03-01T04:55:56.049956","exception":false,"start_time":"2024-03-01T04:55:54.795013","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## References\n* https://keras.io/examples/structured_data/tabtransformer/\n* https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n* https://docs.wandb.ai/guides","metadata":{"papermill":{"duration":0.961175,"end_time":"2024-03-01T04:55:57.919818","exception":false,"start_time":"2024-03-01T04:55:56.958643","status":"completed"},"tags":[]}}]}