{
  "id": 331165,
  "title": "A checklist of ideas for Improving GBM Baselines",
  "url": "/competitions/amex-default-prediction/discussion/331165",
  "author_name": "",
  "post_date": "2022-06-16T02:11:01.012738600Z",
  "votes": 69,
  "comment_count": 9,
  "views": 0,
  "content": "<h1>Checklist for Improving GBM Baselines</h1>\n<p>The following is a checklist of things you should always do when after you got yourself a GBM (LightGBM, XGBoost, Catboost..) baseline for a tabular problem.</p>\n<h3>Data Prepossessing</h3>\n<p><em>Filtering/Handling Missing/Numerical/Categorical Values, Aggregation, Encoding, and Data Augmentation</em></p>\n<ul>\n<li>Filtering: In most of the competitions,  many successful solutions leave out some features excluded from the dataset. We could take the same approach by manually identifying features with a large number of missing values and a poor correlation with the target and removing them. This could be a good way to ensure that all the information in the dataset contributes positively to the model.</li>\n</ul>\n<pre><code>def feature_filter(data, threshold=0.1):\n    features = data.columns\n    filtered_features = []\n    for feature in features:\n        if data[feature].isnull().sum() &lt; threshold:\n            filtered_features.append(feature)\n    return filtered_features\n</code></pre>\n<pre><code>def feature_correlation(data, target, threshold=0.1):\n    correlations = data.corr()[target].drop(target)\n    # Filter the features with correlation to the target less than threshold\n    filtered_features = correlations[abs(correlations) &lt; threshold].index\n    return filtered_features\n</code></pre>\n<ul>\n<li>Handing Missing Values: The difference in the treatment of missing values might result in different outcomes and could make a big difference in the final results. we should experiment with different imputation methods and how we handle outliers. For example, imputing a numerical feature with the median instead of the mean will produce a dataset that is less affected by outliers. Imputing a categorical feature with the mode, on the other hand, might result in a dataset that preserves the distribution of categorical values. There are also additional techniques like adding a category for the missing values, considering missing values as a category, and taking the mode from the classes with similar features.</li>\n</ul>\n<pre><code>def fill_missing_values(data, imputation_method='median'):\n    data_copy = data.copy()\n    for column in data_copy.columns:\n        if data_copy[column].dtype == np.dtype('O'):\n            data_copy[column] = data_copy[column].fillna(data_copy[column].mode().iloc[0])\n        else:\n            if imputation_method == 'median':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].median())\n            elif imputation_method == 'mean':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].mean())\n    return data_copy\n</code></pre>\n<ul>\n<li>Check Numerical Features Scaling: Scaling is important for some models. KNN, Neural Networks, and some other models produce better results when the numerical features are scaled. There are different scaling techniques, like standardization, normalization, and so on. These techniques can be used to experiment with different models.</li>\n</ul>\n<pre><code>def standard_scaler(data):\n    scaler = StandardScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n</code></pre>\n<pre><code>def normalizer(data):\n    scaler = Normalizer()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n</code></pre>\n<pre><code>def min_max_scaler(data):\n    scaler = MinMaxScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n</code></pre>\n<p><strong>And for checking if you really need to even scale anything, check out this plotting function</strong></p>\n<pre><code>def plot_features(data, feature_list):\n    for feature in feature_list:\n        fig, ax = plt.subplots(figsize=(15, 5))\n        plt.subplot(1, 2, 1)\n        sns.distplot(data[feature], kde=False)\n        plt.subplot(1, 2, 2)\n        sns.boxplot(data[feature])\n        plt.show()\n</code></pre>\n<p><strong>Handling Categorical Features:</strong> Many of the top entries use LightGBM, XGBoost, or CatBoost, which do not require a lot of preprocessing. These models are able to handle categorical features. There are different ways to handle categorical features. We could transform the categorical features into numerical features by encoding them. We could also encode the features and keep them as categorical features. There are also other techniques that can be used to improve the way categorical features are handled.</p>\n<p><strong>Aggregation:</strong> Aggregation can be used to reduce the dimensionality of a dataset.  It can also be used to create new features by grouping the data and applying a mathematical function to the groups. We could create new features based on existing ones by grouping the data.</p>\n<pre><code>def aggregations(data, group_vars, num_vars):\n    \"\"\"\n    Create new features based on existing ones by grouping the data\n    \"\"\"\n    # Aggregation\n    data_copy = data.copy()\n    for group_var in group_vars:\n        for num_var in num_vars:\n            data_copy[group_var + '_' + num_var + '_mean'] = data_copy.groupby(group_var)[num_var].transform('mean')\n            data_copy[group_var + '_' + num_var + '_max'] = data_copy.groupby(group_var)[num_var].transform('max')\n            data_copy[group_var + '_' + num_var + '_min'] = data_copy.groupby(group_var)[num_var].transform('min')\n            data_copy[group_var + '_' + num_var + '_std'] = data_copy.groupby(group_var)[num_var].transform('std')\n            data_copy[group_var + '_' + num_var + '_sum'] = data_copy.groupby(group_var)[num_var].transform('sum')\n    return data_copy\n</code></pre>\n<p><strong>Data Augmentation:</strong> Data augmentation for tabular data?! Yes! Just like it is possible to do data augmentation for images and other types of data, it is also possible to do data augmentation for tabular data. Here are a few examples to help get started:</p>\n<ul>\n<li>Adding noise to the columns.</li>\n<li>Swapping values between rows on the same column.</li>\n<li>Generating synthetic rows based on linear interpolation of multiple rows.</li>\n<li>Noising the target</li>\n</ul>\n<h3>Preprocessing and Modeling</h3>\n<p><em>Scaling, Pipelines, and Trying Different Models</em></p>\n<p><strong>Scaling</strong> For some models, scaling is CRUCIAL (KNN, Neural Networks), and for some other models it might just produce better results when the numerical features are scaled. It might be a good idea to to try the scaling process in the model pipelines, to see if the model performance improves or not.</p>\n<p><strong>Try Different Models</strong> Most of the top entries for tabular competitions nowadays are using LightGBM, XGBoost, or CatBoost. There are other models here and there like KNN, Neural Networks, and so on but mostly it is GBMs.  It is worth noting other candidates to be taken into account for tabular datasets, for example:</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1908.07442\" target=\"_blank\">TabNet</a></li>\n</ul>\n<blockquote>\n  <p>TabNet uses sequential attention to choose which features to reason from at each decision step, enabling interpretability and more efficient learning as the learning capacity is used for the most salient features.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet\" target=\"_blank\">Good Notebook</a></p>\n<ul>\n<li><a href=\"https://www.tensorflow.org/decision_forests\" target=\"_blank\">TF Decision Forests</a></li>\n</ul>\n<blockquote>\n  <p>TensorFlow Decision Forests (TF-DF) is a collection of state-of-the-art algorithms for the training, serving and interpretation of Decision Forest models. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/code/usharengaraju/tensorflow-decision-forests-w-b\" target=\"_blank\">Good Notebook</a></p>\n<h3>Parameter Tuning and Feature Selection</h3>\n<p><em>Tuning the Model Parameters and Experimenting with Feature Selection Techniques</em></p>\n<p><strong>Parameters</strong> Baseline models are often not very well-tuned. The current approach is to start with the default parameters and change one parameter at a time. A better approach might be to use a hyperparameter optimizer to experiment with different parameter settings, while trying to find the optimal parameters.</p>\n<p><strong>For example, optuna:</strong></p>\n<pre><code>import optuna\ndef objective(trial):\n    depth = trial.suggest_int('depth', 1, 32)\n    lr = trial.suggest_uniform('lr', 0.001, 0.3)\n    bst = LGBMClassifier(max_depth=depth, learning_rate=lr, metric=\"auc\", \n                         verbose=0, random_state=21)\n    score = cross_val_score(bst, X_train, y_train, scoring='roc_auc', \n                            cv=3, n_jobs=-1).mean()\n    return score\n\nstudy = optuna.create_study()\nstudy.optimize(objective, n_trials=100)\n\ntrial = study.best_trial\nprint('Accuracy: {}'.format(trial.value))\nprint(\"Best hyperparameters: {}\".format(trial.params))\n</code></pre>\n<p><strong>GridSearchCV:</strong></p>\n<pre><code>from sklearn.model_selection import GridSearchCV\ngrid_search = GridSearchCV(estimator = model,\n                           param_grid = parameters,\n                           scoring = 'accuracy',\n                           cv = 10,\n                           n_jobs = -1)\ngrid_search = grid_search.fit(X_train, y_train)\nbest_accuracy = grid_search.best_score_\nbest_parameters = grid_search.best_params_\n</code></pre>\n<p><strong>Feature Selection</strong> It is also important to find ways to reduce the dimensionality of the dataset. Feature selection is one of the most important steps in the data analysis process. It can be used to reduce the dimensionality of the dataset, which improves the speed of the model and makes it easier to visualize. </p>\n<p><a href=\"https://www.kaggle.com/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">Oliver's best notebook ever about null importance</a> Is all you need</p>\n<p><strong>Recursive Feature Elimination:</strong></p>\n<pre><code>from sklearn.feature_selection import RFE\nestimator = RandomForestClassifier(n_estimators=100, random_state=21)\nselector = RFE(estimator, step=1)\nselector = selector.fit(X_train, y_train)\nselector.support_\nselector.ranking_\nX_train.columns[selector.support_]\nX_train_transformed = selector.transform(X_train)\nX_test_transformed = selector.transform(X_test)\n</code></pre>\n<p><strong>Permutation Feature Importance:</strong></p>\n<pre><code>import eli5\nfrom eli5.sklearn import PermutationImportance\nperm = PermutationImportance(clf, random_state=1).fit(X_test, y_test)\neli5.show_weights(perm, feature_names = X_test.columns.tolist())\n</code></pre>\n<h3>Evaluation and Prediction Submission</h3>\n<p><em>Selecting the Best Model, Testing, and Evaluating the Model</em></p>\n<p><strong>Selecting the Best Model</strong> This step is one of the most important steps. We could use cross-validation and hyperparameter tuning to produce the best possible model. </p>\n<p><strong>Testing</strong> We need to do additional tests to see if the model is able to make good predictions. The model needs to be tested against new data, not the train data, to avoid overfitting.</p>\n<p><strong>Evaluation</strong> The last step is to evaluate the model. We could try different metrics to measure the performance of the model. Some of the models work better for binary classification, some for multi-class classification, and some for regression.</p>\n<p><strong>Conclusion</strong> Overall, we need to find ways to improve the current model. The approaches I mentioned can be used to improve the model performance, in addition to providing insights into how the model works. The combination of different preprocessing techniques, feature selection, and modeling can be used to create a highly accurate model.</p>",
  "messages": [
    {
      "id": "1822002",
      "postDate": "06/16/2022 02:11:01",
      "content": "<h1>Checklist for Improving GBM Baselines</h1>\n<p>The following is a checklist of things you should always do when after you got yourself a GBM (LightGBM, XGBoost, Catboost..) baseline for a tabular problem.</p>\n<h3>Data Prepossessing</h3>\n<p><em>Filtering/Handling Missing/Numerical/Categorical Values, Aggregation, Encoding, and Data Augmentation</em></p>\n<ul>\n<li>Filtering: In most of the competitions,  many successful solutions leave out some features excluded from the dataset. We could take the same approach by manually identifying features with a large number of missing values and a poor correlation with the target and removing them. This could be a good way to ensure that all the information in the dataset contributes positively to the model.</li>\n</ul>\n<pre><code>def feature_filter(data, threshold=0.1):\n    features = data.columns\n    filtered_features = []\n    for feature in features:\n        if data[feature].isnull().sum() &lt; threshold:\n            filtered_features.append(feature)\n    return filtered_features\n</code></pre>\n<pre><code>def feature_correlation(data, target, threshold=0.1):\n    correlations = data.corr()[target].drop(target)\n    # Filter the features with correlation to the target less than threshold\n    filtered_features = correlations[abs(correlations) &lt; threshold].index\n    return filtered_features\n</code></pre>\n<ul>\n<li>Handing Missing Values: The difference in the treatment of missing values might result in different outcomes and could make a big difference in the final results. we should experiment with different imputation methods and how we handle outliers. For example, imputing a numerical feature with the median instead of the mean will produce a dataset that is less affected by outliers. Imputing a categorical feature with the mode, on the other hand, might result in a dataset that preserves the distribution of categorical values. There are also additional techniques like adding a category for the missing values, considering missing values as a category, and taking the mode from the classes with similar features.</li>\n</ul>\n<pre><code>def fill_missing_values(data, imputation_method='median'):\n    data_copy = data.copy()\n    for column in data_copy.columns:\n        if data_copy[column].dtype == np.dtype('O'):\n            data_copy[column] = data_copy[column].fillna(data_copy[column].mode().iloc[0])\n        else:\n            if imputation_method == 'median':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].median())\n            elif imputation_method == 'mean':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].mean())\n    return data_copy\n</code></pre>\n<ul>\n<li>Check Numerical Features Scaling: Scaling is important for some models. KNN, Neural Networks, and some other models produce better results when the numerical features are scaled. There are different scaling techniques, like standardization, normalization, and so on. These techniques can be used to experiment with different models.</li>\n</ul>\n<pre><code>def standard_scaler(data):\n    scaler = StandardScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n</code></pre>\n<pre><code>def normalizer(data):\n    scaler = Normalizer()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n</code></pre>\n<pre><code>def min_max_scaler(data):\n    scaler = MinMaxScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n</code></pre>\n<p><strong>And for checking if you really need to even scale anything, check out this plotting function</strong></p>\n<pre><code>def plot_features(data, feature_list):\n    for feature in feature_list:\n        fig, ax = plt.subplots(figsize=(15, 5))\n        plt.subplot(1, 2, 1)\n        sns.distplot(data[feature], kde=False)\n        plt.subplot(1, 2, 2)\n        sns.boxplot(data[feature])\n        plt.show()\n</code></pre>\n<p><strong>Handling Categorical Features:</strong> Many of the top entries use LightGBM, XGBoost, or CatBoost, which do not require a lot of preprocessing. These models are able to handle categorical features. There are different ways to handle categorical features. We could transform the categorical features into numerical features by encoding them. We could also encode the features and keep them as categorical features. There are also other techniques that can be used to improve the way categorical features are handled.</p>\n<p><strong>Aggregation:</strong> Aggregation can be used to reduce the dimensionality of a dataset.  It can also be used to create new features by grouping the data and applying a mathematical function to the groups. We could create new features based on existing ones by grouping the data.</p>\n<pre><code>def aggregations(data, group_vars, num_vars):\n    \"\"\"\n    Create new features based on existing ones by grouping the data\n    \"\"\"\n    # Aggregation\n    data_copy = data.copy()\n    for group_var in group_vars:\n        for num_var in num_vars:\n            data_copy[group_var + '_' + num_var + '_mean'] = data_copy.groupby(group_var)[num_var].transform('mean')\n            data_copy[group_var + '_' + num_var + '_max'] = data_copy.groupby(group_var)[num_var].transform('max')\n            data_copy[group_var + '_' + num_var + '_min'] = data_copy.groupby(group_var)[num_var].transform('min')\n            data_copy[group_var + '_' + num_var + '_std'] = data_copy.groupby(group_var)[num_var].transform('std')\n            data_copy[group_var + '_' + num_var + '_sum'] = data_copy.groupby(group_var)[num_var].transform('sum')\n    return data_copy\n</code></pre>\n<p><strong>Data Augmentation:</strong> Data augmentation for tabular data?! Yes! Just like it is possible to do data augmentation for images and other types of data, it is also possible to do data augmentation for tabular data. Here are a few examples to help get started:</p>\n<ul>\n<li>Adding noise to the columns.</li>\n<li>Swapping values between rows on the same column.</li>\n<li>Generating synthetic rows based on linear interpolation of multiple rows.</li>\n<li>Noising the target</li>\n</ul>\n<h3>Preprocessing and Modeling</h3>\n<p><em>Scaling, Pipelines, and Trying Different Models</em></p>\n<p><strong>Scaling</strong> For some models, scaling is CRUCIAL (KNN, Neural Networks), and for some other models it might just produce better results when the numerical features are scaled. It might be a good idea to to try the scaling process in the model pipelines, to see if the model performance improves or not.</p>\n<p><strong>Try Different Models</strong> Most of the top entries for tabular competitions nowadays are using LightGBM, XGBoost, or CatBoost. There are other models here and there like KNN, Neural Networks, and so on but mostly it is GBMs.  It is worth noting other candidates to be taken into account for tabular datasets, for example:</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1908.07442\" target=\"_blank\">TabNet</a></li>\n</ul>\n<blockquote>\n  <p>TabNet uses sequential attention to choose which features to reason from at each decision step, enabling interpretability and more efficient learning as the learning capacity is used for the most salient features.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet\" target=\"_blank\">Good Notebook</a></p>\n<ul>\n<li><a href=\"https://www.tensorflow.org/decision_forests\" target=\"_blank\">TF Decision Forests</a></li>\n</ul>\n<blockquote>\n  <p>TensorFlow Decision Forests (TF-DF) is a collection of state-of-the-art algorithms for the training, serving and interpretation of Decision Forest models. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/code/usharengaraju/tensorflow-decision-forests-w-b\" target=\"_blank\">Good Notebook</a></p>\n<h3>Parameter Tuning and Feature Selection</h3>\n<p><em>Tuning the Model Parameters and Experimenting with Feature Selection Techniques</em></p>\n<p><strong>Parameters</strong> Baseline models are often not very well-tuned. The current approach is to start with the default parameters and change one parameter at a time. A better approach might be to use a hyperparameter optimizer to experiment with different parameter settings, while trying to find the optimal parameters.</p>\n<p><strong>For example, optuna:</strong></p>\n<pre><code>import optuna\ndef objective(trial):\n    depth = trial.suggest_int('depth', 1, 32)\n    lr = trial.suggest_uniform('lr', 0.001, 0.3)\n    bst = LGBMClassifier(max_depth=depth, learning_rate=lr, metric=\"auc\", \n                         verbose=0, random_state=21)\n    score = cross_val_score(bst, X_train, y_train, scoring='roc_auc', \n                            cv=3, n_jobs=-1).mean()\n    return score\n\nstudy = optuna.create_study()\nstudy.optimize(objective, n_trials=100)\n\ntrial = study.best_trial\nprint('Accuracy: {}'.format(trial.value))\nprint(\"Best hyperparameters: {}\".format(trial.params))\n</code></pre>\n<p><strong>GridSearchCV:</strong></p>\n<pre><code>from sklearn.model_selection import GridSearchCV\ngrid_search = GridSearchCV(estimator = model,\n                           param_grid = parameters,\n                           scoring = 'accuracy',\n                           cv = 10,\n                           n_jobs = -1)\ngrid_search = grid_search.fit(X_train, y_train)\nbest_accuracy = grid_search.best_score_\nbest_parameters = grid_search.best_params_\n</code></pre>\n<p><strong>Feature Selection</strong> It is also important to find ways to reduce the dimensionality of the dataset. Feature selection is one of the most important steps in the data analysis process. It can be used to reduce the dimensionality of the dataset, which improves the speed of the model and makes it easier to visualize. </p>\n<p><a href=\"https://www.kaggle.com/ogrellier/feature-selection-with-null-importances\" target=\"_blank\">Oliver's best notebook ever about null importance</a> Is all you need</p>\n<p><strong>Recursive Feature Elimination:</strong></p>\n<pre><code>from sklearn.feature_selection import RFE\nestimator = RandomForestClassifier(n_estimators=100, random_state=21)\nselector = RFE(estimator, step=1)\nselector = selector.fit(X_train, y_train)\nselector.support_\nselector.ranking_\nX_train.columns[selector.support_]\nX_train_transformed = selector.transform(X_train)\nX_test_transformed = selector.transform(X_test)\n</code></pre>\n<p><strong>Permutation Feature Importance:</strong></p>\n<pre><code>import eli5\nfrom eli5.sklearn import PermutationImportance\nperm = PermutationImportance(clf, random_state=1).fit(X_test, y_test)\neli5.show_weights(perm, feature_names = X_test.columns.tolist())\n</code></pre>\n<h3>Evaluation and Prediction Submission</h3>\n<p><em>Selecting the Best Model, Testing, and Evaluating the Model</em></p>\n<p><strong>Selecting the Best Model</strong> This step is one of the most important steps. We could use cross-validation and hyperparameter tuning to produce the best possible model. </p>\n<p><strong>Testing</strong> We need to do additional tests to see if the model is able to make good predictions. The model needs to be tested against new data, not the train data, to avoid overfitting.</p>\n<p><strong>Evaluation</strong> The last step is to evaluate the model. We could try different metrics to measure the performance of the model. Some of the models work better for binary classification, some for multi-class classification, and some for regression.</p>\n<p><strong>Conclusion</strong> Overall, we need to find ways to improve the current model. The approaches I mentioned can be used to improve the model performance, in addition to providing insights into how the model works. The combination of different preprocessing techniques, feature selection, and modeling can be used to create a highly accurate model.</p>",
      "rawMarkdown": "# Checklist for Improving GBM Baselines\n\nThe following is a checklist of things you should always do when after you got yourself a GBM (LightGBM, XGBoost, Catboost..) baseline for a tabular problem.\n\n### Data Prepossessing\n\n*Filtering/Handling Missing/Numerical/Categorical Values, Aggregation, Encoding, and Data Augmentation*\n\n* Filtering: In most of the competitions,  many successful solutions leave out some features excluded from the dataset. We could take the same approach by manually identifying features with a large number of missing values and a poor correlation with the target and removing them. This could be a good way to ensure that all the information in the dataset contributes positively to the model.\n\n\n```python\ndef feature_filter(data, threshold=0.1):\n    features = data.columns\n    filtered_features = []\n    for feature in features:\n        if data[feature].isnull().sum() < threshold:\n            filtered_features.append(feature)\n    return filtered_features\n```\n\n```python\ndef feature_correlation(data, target, threshold=0.1):\n    correlations = data.corr()[target].drop(target)\n    # Filter the features with correlation to the target less than threshold\n    filtered_features = correlations[abs(correlations) < threshold].index\n    return filtered_features\n```\n\n* Handing Missing Values: The difference in the treatment of missing values might result in different outcomes and could make a big difference in the final results. we should experiment with different imputation methods and how we handle outliers. For example, imputing a numerical feature with the median instead of the mean will produce a dataset that is less affected by outliers. Imputing a categorical feature with the mode, on the other hand, might result in a dataset that preserves the distribution of categorical values. There are also additional techniques like adding a category for the missing values, considering missing values as a category, and taking the mode from the classes with similar features.\n\n```python\ndef fill_missing_values(data, imputation_method='median'):\n    data_copy = data.copy()\n    for column in data_copy.columns:\n        if data_copy[column].dtype == np.dtype('O'):\n            data_copy[column] = data_copy[column].fillna(data_copy[column].mode().iloc[0])\n        else:\n            if imputation_method == 'median':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].median())\n            elif imputation_method == 'mean':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].mean())\n    return data_copy\n```\n\n* Check Numerical Features Scaling: Scaling is important for some models. KNN, Neural Networks, and some other models produce better results when the numerical features are scaled. There are different scaling techniques, like standardization, normalization, and so on. These techniques can be used to experiment with different models.\n\n```python\ndef standard_scaler(data):\n    scaler = StandardScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n```\n\n```python\ndef normalizer(data):\n    scaler = Normalizer()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n```\n\n```python\ndef min_max_scaler(data):\n    scaler = MinMaxScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n```\n\n**And for checking if you really need to even scale anything, check out this plotting function**\n\n```python\ndef plot_features(data, feature_list):\n    for feature in feature_list:\n        fig, ax = plt.subplots(figsize=(15, 5))\n        plt.subplot(1, 2, 1)\n        sns.distplot(data[feature], kde=False)\n        plt.subplot(1, 2, 2)\n        sns.boxplot(data[feature])\n        plt.show()\n```\n\n**Handling Categorical Features:** Many of the top entries use LightGBM, XGBoost, or CatBoost, which do not require a lot of preprocessing. These models are able to handle categorical features. There are different ways to handle categorical features. We could transform the categorical features into numerical features by encoding them. We could also encode the features and keep them as categorical features. There are also other techniques that can be used to improve the way categorical features are handled.\n\n**Aggregation:** Aggregation can be used to reduce the dimensionality of a dataset.  It can also be used to create new features by grouping the data and applying a mathematical function to the groups. We could create new features based on existing ones by grouping the data.\n\n```python\ndef aggregations(data, group_vars, num_vars):\n    \"\"\"\n    Create new features based on existing ones by grouping the data\n    \"\"\"\n    # Aggregation\n    data_copy = data.copy()\n    for group_var in group_vars:\n        for num_var in num_vars:\n            data_copy[group_var + '_' + num_var + '_mean'] = data_copy.groupby(group_var)[num_var].transform('mean')\n            data_copy[group_var + '_' + num_var + '_max'] = data_copy.groupby(group_var)[num_var].transform('max')\n            data_copy[group_var + '_' + num_var + '_min'] = data_copy.groupby(group_var)[num_var].transform('min')\n            data_copy[group_var + '_' + num_var + '_std'] = data_copy.groupby(group_var)[num_var].transform('std')\n            data_copy[group_var + '_' + num_var + '_sum'] = data_copy.groupby(group_var)[num_var].transform('sum')\n    return data_copy\n```\n\n**Data Augmentation:** Data augmentation for tabular data?! Yes! Just like it is possible to do data augmentation for images and other types of data, it is also possible to do data augmentation for tabular data. Here are a few examples to help get started:\n\n- Adding noise to the columns.\n- Swapping values between rows on the same column.\n- Generating synthetic rows based on linear interpolation of multiple rows.\n- Noising the target\n\n\n### Preprocessing and Modeling\n\n*Scaling, Pipelines, and Trying Different Models*\n\n**Scaling** For some models, scaling is CRUCIAL (KNN, Neural Networks), and for some other models it might just produce better results when the numerical features are scaled. It might be a good idea to to try the scaling process in the model pipelines, to see if the model performance improves or not.\n\n**Try Different Models** Most of the top entries for tabular competitions nowadays are using LightGBM, XGBoost, or CatBoost. There are other models here and there like KNN, Neural Networks, and so on but mostly it is GBMs.  It is worth noting other candidates to be taken into account for tabular datasets, for example:\n\n- [TabNet](https://arxiv.org/abs/1908.07442)\n\n> TabNet uses sequential attention to choose which features to reason from at each decision step, enabling interpretability and more efficient learning as the learning capacity is used for the most salient features.\n\n[Good Notebook](https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet)\n\n- [TF Decision Forests](https://www.tensorflow.org/decision_forests)\n\n> TensorFlow Decision Forests (TF-DF) is a collection of state-of-the-art algorithms for the training, serving and interpretation of Decision Forest models. \n\n[Good Notebook](https://www.kaggle.com/code/usharengaraju/tensorflow-decision-forests-w-b)\n\n### Parameter Tuning and Feature Selection\n*Tuning the Model Parameters and Experimenting with Feature Selection Techniques*\n\n**Parameters** Baseline models are often not very well-tuned. The current approach is to start with the default parameters and change one parameter at a time. A better approach might be to use a hyperparameter optimizer to experiment with different parameter settings, while trying to find the optimal parameters.\n\n**For example, optuna:**\n```python\nimport optuna\ndef objective(trial):\n    depth = trial.suggest_int('depth', 1, 32)\n    lr = trial.suggest_uniform('lr', 0.001, 0.3)\n    bst = LGBMClassifier(max_depth=depth, learning_rate=lr, metric=\"auc\", \n                         verbose=0, random_state=21)\n    score = cross_val_score(bst, X_train, y_train, scoring='roc_auc', \n                            cv=3, n_jobs=-1).mean()\n    return score\n\nstudy = optuna.create_study()\nstudy.optimize(objective, n_trials=100)\n\ntrial = study.best_trial\nprint('Accuracy: {}'.format(trial.value))\nprint(\"Best hyperparameters: {}\".format(trial.params))\n```\n\n**GridSearchCV:**\n\n```python\nfrom sklearn.model_selection import GridSearchCV\ngrid_search = GridSearchCV(estimator = model,\n                           param_grid = parameters,\n                           scoring = 'accuracy',\n                           cv = 10,\n                           n_jobs = -1)\ngrid_search = grid_search.fit(X_train, y_train)\nbest_accuracy = grid_search.best_score_\nbest_parameters = grid_search.best_params_\n```\n\n**Feature Selection** It is also important to find ways to reduce the dimensionality of the dataset. Feature selection is one of the most important steps in the data analysis process. It can be used to reduce the dimensionality of the dataset, which improves the speed of the model and makes it easier to visualize. \n\n[Oliver's best notebook ever about null importance](https://www.kaggle.com/ogrellier/feature-selection-with-null-importances) Is all you need\n\n**Recursive Feature Elimination:**\n\n```python\nfrom sklearn.feature_selection import RFE\nestimator = RandomForestClassifier(n_estimators=100, random_state=21)\nselector = RFE(estimator, step=1)\nselector = selector.fit(X_train, y_train)\nselector.support_\nselector.ranking_\nX_train.columns[selector.support_]\nX_train_transformed = selector.transform(X_train)\nX_test_transformed = selector.transform(X_test)\n```\n\n**Permutation Feature Importance:**\n\n```python\nimport eli5\nfrom eli5.sklearn import PermutationImportance\nperm = PermutationImportance(clf, random_state=1).fit(X_test, y_test)\neli5.show_weights(perm, feature_names = X_test.columns.tolist())\n```\n\n### Evaluation and Prediction Submission\n\n*Selecting the Best Model, Testing, and Evaluating the Model*\n\n**Selecting the Best Model** This step is one of the most important steps. We could use cross-validation and hyperparameter tuning to produce the best possible model. \n\n**Testing** We need to do additional tests to see if the model is able to make good predictions. The model needs to be tested against new data, not the train data, to avoid overfitting.\n\n**Evaluation** The last step is to evaluate the model. We could try different metrics to measure the performance of the model. Some of the models work better for binary classification, some for multi-class classification, and some for regression.\n\n**Conclusion** Overall, we need to find ways to improve the current model. The approaches I mentioned can be used to improve the model performance, in addition to providing insights into how the model works. The combination of different preprocessing techniques, feature selection, and modeling can be used to create a highly accurate model.",
      "votes": null
    },
    {
      "id": "1822859",
      "postDate": "06/16/2022 18:32:49",
      "content": "<p>This info is DEVASTATING! thanks!</p>",
      "rawMarkdown": "This info is DEVASTATING! thanks!",
      "votes": null
    },
    {
      "id": "1823093",
      "postDate": "06/17/2022 03:37:52",
      "content": "<p>Great work! <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> Thanks for sharing</p>",
      "rawMarkdown": "Great work! @thedevastator Thanks for sharing",
      "votes": null
    },
    {
      "id": "1823208",
      "postDate": "06/17/2022 06:53:00",
      "content": "<p>Amazing  topic! very informative  <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> thanks for sharing 👍</p>",
      "rawMarkdown": "Amazing  topic! very informative  @thedevastator thanks for sharing 👍",
      "votes": null
    },
    {
      "id": "1823594",
      "postDate": "06/17/2022 14:06:49",
      "content": "<p>This is a great post. Thank you so much!</p>",
      "rawMarkdown": "This is a great post. Thank you so much!",
      "votes": null
    },
    {
      "id": "1824898",
      "postDate": "06/18/2022 17:55:08",
      "content": "<p>Interesting, good job <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "rawMarkdown": "Interesting, good job @thedevastator",
      "votes": null
    },
    {
      "id": "1826212",
      "postDate": "06/20/2022 06:52:31",
      "content": "<p>Amazing Thanku <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "rawMarkdown": "Amazing Thanku @thedevastator",
      "votes": null
    },
    {
      "id": "1826615",
      "postDate": "06/20/2022 13:44:03",
      "content": "<p>Detailed explanation, very helpful. Thanks for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>.</p>",
      "rawMarkdown": "Detailed explanation, very helpful. Thanks for sharing @thedevastator.",
      "votes": null
    },
    {
      "id": "1859039",
      "postDate": "07/17/2022 11:26:05",
      "content": "<p><a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> Thanks for sharing such valuable insights. After a month's struggle I have managed to work with huge datasets for this challenge and were able to train model. I tried LGBM and XGBoost but my scores are stuck at a point. 😨</p>\n<p>I was wondering wha else I can do with data feature to improve the score. I will try ideas in this post and see how it goes. 😁</p>",
      "rawMarkdown": "thedevastator Thanks for sharing such valuable insights. After a month's struggle I have managed to work with huge datasets for this challenge and were able to train model. I tried LGBM and XGBoost but my scores are stuck at a point. 😨\n\nI was wondering wha else I can do with data feature to improve the score. I will try ideas in this post and see how it goes. 😁",
      "votes": null
    },
    {
      "id": "1892223",
      "postDate": "08/10/2022 00:49:32",
      "content": "<p>Thanks for sharing! Learned a lot!</p>",
      "rawMarkdown": "Thanks for sharing! Learned a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1822859,
      "author_name": "wuuthraad",
      "author_url": "",
      "post_date": "06/16/2022 18:32:49",
      "content": "<p>This info is DEVASTATING! thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1823093,
      "author_name": "thamotharan",
      "author_url": "",
      "post_date": "06/17/2022 03:37:52",
      "content": "<p>Great work! <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1823208,
      "author_name": "summerakousar",
      "author_url": "",
      "post_date": "06/17/2022 06:53:00",
      "content": "<p>Amazing  topic! very informative  <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> thanks for sharing 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1823594,
      "author_name": "kevinoconnor956",
      "author_url": "",
      "post_date": "06/17/2022 14:06:49",
      "content": "<p>This is a great post. Thank you so much!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1824898,
      "author_name": "saberghaderi",
      "author_url": "",
      "post_date": "06/18/2022 17:55:08",
      "content": "<p>Interesting, good job <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1826212,
      "author_name": "sakshi20008",
      "author_url": "",
      "post_date": "06/20/2022 06:52:31",
      "content": "<p>Amazing Thanku <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1826615,
      "author_name": "naveenkonam1985",
      "author_url": "",
      "post_date": "06/20/2022 13:44:03",
      "content": "<p>Detailed explanation, very helpful. Thanks for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1859039,
      "author_name": "mirfanazam",
      "author_url": "",
      "post_date": "07/17/2022 11:26:05",
      "content": "<p><a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> Thanks for sharing such valuable insights. After a month's struggle I have managed to work with huge datasets for this challenge and were able to train model. I tried LGBM and XGBoost but my scores are stuck at a point. 😨</p>\n<p>I was wondering wha else I can do with data feature to improve the score. I will try ideas in this post and see how it goes. 😁</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1892223,
      "author_name": "arti1117",
      "author_url": "",
      "post_date": "08/10/2022 00:49:32",
      "content": "<p>Thanks for sharing! Learned a lot!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1822002": "# Checklist for Improving GBM Baselines\n\nThe following is a checklist of things you should always do when after you got yourself a GBM (LightGBM, XGBoost, Catboost..) baseline for a tabular problem.\n\n### Data Prepossessing\n\n*Filtering/Handling Missing/Numerical/Categorical Values, Aggregation, Encoding, and Data Augmentation*\n\n* Filtering: In most of the competitions,  many successful solutions leave out some features excluded from the dataset. We could take the same approach by manually identifying features with a large number of missing values and a poor correlation with the target and removing them. This could be a good way to ensure that all the information in the dataset contributes positively to the model.\n\n\n```python\ndef feature_filter(data, threshold=0.1):\n    features = data.columns\n    filtered_features = []\n    for feature in features:\n        if data[feature].isnull().sum() < threshold:\n            filtered_features.append(feature)\n    return filtered_features\n```\n\n```python\ndef feature_correlation(data, target, threshold=0.1):\n    correlations = data.corr()[target].drop(target)\n    # Filter the features with correlation to the target less than threshold\n    filtered_features = correlations[abs(correlations) < threshold].index\n    return filtered_features\n```\n\n* Handing Missing Values: The difference in the treatment of missing values might result in different outcomes and could make a big difference in the final results. we should experiment with different imputation methods and how we handle outliers. For example, imputing a numerical feature with the median instead of the mean will produce a dataset that is less affected by outliers. Imputing a categorical feature with the mode, on the other hand, might result in a dataset that preserves the distribution of categorical values. There are also additional techniques like adding a category for the missing values, considering missing values as a category, and taking the mode from the classes with similar features.\n\n```python\ndef fill_missing_values(data, imputation_method='median'):\n    data_copy = data.copy()\n    for column in data_copy.columns:\n        if data_copy[column].dtype == np.dtype('O'):\n            data_copy[column] = data_copy[column].fillna(data_copy[column].mode().iloc[0])\n        else:\n            if imputation_method == 'median':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].median())\n            elif imputation_method == 'mean':\n                data_copy[column] = data_copy[column].fillna(data_copy[column].mean())\n    return data_copy\n```\n\n* Check Numerical Features Scaling: Scaling is important for some models. KNN, Neural Networks, and some other models produce better results when the numerical features are scaled. There are different scaling techniques, like standardization, normalization, and so on. These techniques can be used to experiment with different models.\n\n```python\ndef standard_scaler(data):\n    scaler = StandardScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n```\n\n```python\ndef normalizer(data):\n    scaler = Normalizer()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n```\n\n```python\ndef min_max_scaler(data):\n    scaler = MinMaxScaler()\n    numerical = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    data_copy = data.copy()\n    data_copy[data_copy.select_dtypes(include=numerical).columns] = scaler.fit_transform(data_copy.select_dtypes(include=numerical))\n    return data_copy\n```\n\n**And for checking if you really need to even scale anything, check out this plotting function**\n\n```python\ndef plot_features(data, feature_list):\n    for feature in feature_list:\n        fig, ax = plt.subplots(figsize=(15, 5))\n        plt.subplot(1, 2, 1)\n        sns.distplot(data[feature], kde=False)\n        plt.subplot(1, 2, 2)\n        sns.boxplot(data[feature])\n        plt.show()\n```\n\n**Handling Categorical Features:** Many of the top entries use LightGBM, XGBoost, or CatBoost, which do not require a lot of preprocessing. These models are able to handle categorical features. There are different ways to handle categorical features. We could transform the categorical features into numerical features by encoding them. We could also encode the features and keep them as categorical features. There are also other techniques that can be used to improve the way categorical features are handled.\n\n**Aggregation:** Aggregation can be used to reduce the dimensionality of a dataset.  It can also be used to create new features by grouping the data and applying a mathematical function to the groups. We could create new features based on existing ones by grouping the data.\n\n```python\ndef aggregations(data, group_vars, num_vars):\n    \"\"\"\n    Create new features based on existing ones by grouping the data\n    \"\"\"\n    # Aggregation\n    data_copy = data.copy()\n    for group_var in group_vars:\n        for num_var in num_vars:\n            data_copy[group_var + '_' + num_var + '_mean'] = data_copy.groupby(group_var)[num_var].transform('mean')\n            data_copy[group_var + '_' + num_var + '_max'] = data_copy.groupby(group_var)[num_var].transform('max')\n            data_copy[group_var + '_' + num_var + '_min'] = data_copy.groupby(group_var)[num_var].transform('min')\n            data_copy[group_var + '_' + num_var + '_std'] = data_copy.groupby(group_var)[num_var].transform('std')\n            data_copy[group_var + '_' + num_var + '_sum'] = data_copy.groupby(group_var)[num_var].transform('sum')\n    return data_copy\n```\n\n**Data Augmentation:** Data augmentation for tabular data?! Yes! Just like it is possible to do data augmentation for images and other types of data, it is also possible to do data augmentation for tabular data. Here are a few examples to help get started:\n\n- Adding noise to the columns.\n- Swapping values between rows on the same column.\n- Generating synthetic rows based on linear interpolation of multiple rows.\n- Noising the target\n\n\n### Preprocessing and Modeling\n\n*Scaling, Pipelines, and Trying Different Models*\n\n**Scaling** For some models, scaling is CRUCIAL (KNN, Neural Networks), and for some other models it might just produce better results when the numerical features are scaled. It might be a good idea to to try the scaling process in the model pipelines, to see if the model performance improves or not.\n\n**Try Different Models** Most of the top entries for tabular competitions nowadays are using LightGBM, XGBoost, or CatBoost. There are other models here and there like KNN, Neural Networks, and so on but mostly it is GBMs.  It is worth noting other candidates to be taken into account for tabular datasets, for example:\n\n- [TabNet](https://arxiv.org/abs/1908.07442)\n\n> TabNet uses sequential attention to choose which features to reason from at each decision step, enabling interpretability and more efficient learning as the learning capacity is used for the most salient features.\n\n[Good Notebook](https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet)\n\n- [TF Decision Forests](https://www.tensorflow.org/decision_forests)\n\n> TensorFlow Decision Forests (TF-DF) is a collection of state-of-the-art algorithms for the training, serving and interpretation of Decision Forest models. \n\n[Good Notebook](https://www.kaggle.com/code/usharengaraju/tensorflow-decision-forests-w-b)\n\n### Parameter Tuning and Feature Selection\n*Tuning the Model Parameters and Experimenting with Feature Selection Techniques*\n\n**Parameters** Baseline models are often not very well-tuned. The current approach is to start with the default parameters and change one parameter at a time. A better approach might be to use a hyperparameter optimizer to experiment with different parameter settings, while trying to find the optimal parameters.\n\n**For example, optuna:**\n```python\nimport optuna\ndef objective(trial):\n    depth = trial.suggest_int('depth', 1, 32)\n    lr = trial.suggest_uniform('lr', 0.001, 0.3)\n    bst = LGBMClassifier(max_depth=depth, learning_rate=lr, metric=\"auc\", \n                         verbose=0, random_state=21)\n    score = cross_val_score(bst, X_train, y_train, scoring='roc_auc', \n                            cv=3, n_jobs=-1).mean()\n    return score\n\nstudy = optuna.create_study()\nstudy.optimize(objective, n_trials=100)\n\ntrial = study.best_trial\nprint('Accuracy: {}'.format(trial.value))\nprint(\"Best hyperparameters: {}\".format(trial.params))\n```\n\n**GridSearchCV:**\n\n```python\nfrom sklearn.model_selection import GridSearchCV\ngrid_search = GridSearchCV(estimator = model,\n                           param_grid = parameters,\n                           scoring = 'accuracy',\n                           cv = 10,\n                           n_jobs = -1)\ngrid_search = grid_search.fit(X_train, y_train)\nbest_accuracy = grid_search.best_score_\nbest_parameters = grid_search.best_params_\n```\n\n**Feature Selection** It is also important to find ways to reduce the dimensionality of the dataset. Feature selection is one of the most important steps in the data analysis process. It can be used to reduce the dimensionality of the dataset, which improves the speed of the model and makes it easier to visualize. \n\n[Oliver's best notebook ever about null importance](https://www.kaggle.com/ogrellier/feature-selection-with-null-importances) Is all you need\n\n**Recursive Feature Elimination:**\n\n```python\nfrom sklearn.feature_selection import RFE\nestimator = RandomForestClassifier(n_estimators=100, random_state=21)\nselector = RFE(estimator, step=1)\nselector = selector.fit(X_train, y_train)\nselector.support_\nselector.ranking_\nX_train.columns[selector.support_]\nX_train_transformed = selector.transform(X_train)\nX_test_transformed = selector.transform(X_test)\n```\n\n**Permutation Feature Importance:**\n\n```python\nimport eli5\nfrom eli5.sklearn import PermutationImportance\nperm = PermutationImportance(clf, random_state=1).fit(X_test, y_test)\neli5.show_weights(perm, feature_names = X_test.columns.tolist())\n```\n\n### Evaluation and Prediction Submission\n\n*Selecting the Best Model, Testing, and Evaluating the Model*\n\n**Selecting the Best Model** This step is one of the most important steps. We could use cross-validation and hyperparameter tuning to produce the best possible model. \n\n**Testing** We need to do additional tests to see if the model is able to make good predictions. The model needs to be tested against new data, not the train data, to avoid overfitting.\n\n**Evaluation** The last step is to evaluate the model. We could try different metrics to measure the performance of the model. Some of the models work better for binary classification, some for multi-class classification, and some for regression.\n\n**Conclusion** Overall, we need to find ways to improve the current model. The approaches I mentioned can be used to improve the model performance, in addition to providing insights into how the model works. The combination of different preprocessing techniques, feature selection, and modeling can be used to create a highly accurate model.",
    "1822859": "This info is DEVASTATING! thanks!",
    "1823093": "Great work! @thedevastator Thanks for sharing",
    "1823208": "Amazing  topic! very informative  @thedevastator thanks for sharing 👍",
    "1823594": "This is a great post. Thank you so much!",
    "1824898": "Interesting, good job @thedevastator",
    "1826212": "Amazing Thanku @thedevastator",
    "1826615": "Detailed explanation, very helpful. Thanks for sharing @thedevastator.",
    "1859039": "thedevastator Thanks for sharing such valuable insights. After a month's struggle I have managed to work with huge datasets for this challenge and were able to train model. I tried LGBM and XGBoost but my scores are stuck at a point. 😨\n\nI was wondering wha else I can do with data feature to improve the score. I will try ideas in this post and see how it goes. 😁",
    "1892223": "Thanks for sharing! Learned a lot!"
  },
  "source": "meta"
}