{
  "id": 478094,
  "title": "Improving model performance by handling data imbalance",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/478094",
  "author_name": "",
  "post_date": "2024-02-19T06:30:28.804033500Z",
  "votes": 40,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In this competition we have to deal with a highly imbalanced class distribution, as there are only about 3% of positive targets representing defaults:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>[k]</th>\n<th>[%]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Positive targets</td>\n<td>48</td>\n<td>3</td>\n</tr>\n<tr>\n<td>Negative targets</td>\n<td>1,479</td>\n<td>97</td>\n</tr>\n<tr>\n<td>Total cases</td>\n<td>1,527</td>\n<td>100</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>After grouping data by <code>WEEK_NUM</code>, one can see that the default ratio doesn't really change much over the time, even though the total number of cases per week could change dramatically:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F54e2ab7b4623137a7c2cbb00aaacf063%2FScreenshot%202024-02-18%20at%2010.19.08%20PM.png?generation=1708323569689791&amp;alt=media\"></p>\n<p>It means that when we stratify our folds by weeks, each of our models trains on 2-5% of positive targets and on 95-98% of negative ones. As a result, such an imbalance can affect model's ability to properly predict defaults (positive targets). Luckily, in this competition we mostly care about AUC, and not about actual probabilities, so can give a try to built-in class balancing features from popular gradient boosting libraries.</p>\n<h1>XGBoost</h1>\n<p>For <code>XGBClassifier</code> we can use <a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/param_tuning.html#handle-imbalanced-dataset\" target=\"_blank\">scale_pos_weight</a> option, that scales the positive class weight. According to <a href=\"https://xgboost.readthedocs.io/en/stable/parameter.html\" target=\"_blank\">documentation</a>, a typical value for this option is <code>sum(negative targets) / sum(positive targets)</code>, i.e. in our case it is around <code>31</code>. This value may need to be adjusted for each fold separately, as the default ratio varies over different time periods.</p>\n<p><strong>UPD:</strong> There is also a way to assign weight to each training sample separately. For that, one needs to convert training data to <a href=\"https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.core\" target=\"_blank\">DMatrix</a> object passing appropriate <code>weight</code> vector when creating it. Thanks to <a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> for pointing this out.</p>\n<h1>LightGBM</h1>\n<p>For <code>LGBMClassifier</code> there is a similar <code>scale_pos_weight</code> option to scale weight of positive class. In addition, there is a boolean flag <code>is_unbalanced</code>, that automatically assigns weights based on the ratio between the numbers of negative and positive classes. According to official <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters.html#is_unbalance\" target=\"_blank\">documentation</a> you can only use one of these options at a time, that makes sense as there is only one set of weights the model can use for training.</p>\n<h1>CatBoost</h1>\n<p>Just like the other libraries, <code>CatBoostClassifier</code> supports <a href=\"https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight\" target=\"_blank\">scale_pos_weight</a> option. There is also a more advanced <a href=\"https://catboost.ai/en/docs/references/training-parameters/common#auto_class_weights\" target=\"_blank\">auto_class_weights</a> feature, that can automatically adjust weights in <code>Balanced</code> and <code>SqrtBalanced</code> modes. You can use only one of these options to balance training.</p>\n<h1>Conclusions</h1>\n<p>Class weights, as all the other model parameters, may need to be tuned. In my experiments, setting <code>scale_pos_weight</code> to recommended typical values improved AUC by 0.01-0.02 for all the three libraries. The best LB improvement, however, was at the level of more humble 0.005 for LightGBM, but this is yet to be seen when the metric is fixed.</p>\n<p>Besides adjusting class weights to deal with imbalance, one can also employ other techniques like down-/upsampling of data or postprocessing. For more information on this please refer to the following notebooks and discussions:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/shahules/tackling-class-imbalance\" target=\"_blank\">tackling-class-imbalance</a></li>\n<li><a href=\"https://www.kaggle.com/code/kaanboke/xgboost-lightgbm-catboost-imbalanced-data\" target=\"_blank\">xgboost-lightgbm-catboost-imbalanced-data</a></li>\n<li><a href=\"https://www.kaggle.com/code/janiobachmann/credit-fraud-dealing-with-imbalanced-datasets\" target=\"_blank\">credit-fraud-dealing-with-imbalanced-datasets</a></li>\n<li><a href=\"https://www.kaggle.com/code/kailex/talkingdata-eda-and-class-imbalance\" target=\"_blank\">talkingdata-eda-and-class-imbalance</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icr-identify-age-related-conditions/discussion/412507\" target=\"_blank\">How To Balance Training And Boost CV and LB Score!</a></li>\n</ul>\n<p>Hope it helps and happy kaggling!</p>",
  "messages": [
    {
      "id": "2658363",
      "postDate": "02/19/2024 06:30:28",
      "content": "<p>In this competition we have to deal with a highly imbalanced class distribution, as there are only about 3% of positive targets representing defaults:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>[k]</th>\n<th>[%]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Positive targets</td>\n<td>48</td>\n<td>3</td>\n</tr>\n<tr>\n<td>Negative targets</td>\n<td>1,479</td>\n<td>97</td>\n</tr>\n<tr>\n<td>Total cases</td>\n<td>1,527</td>\n<td>100</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>After grouping data by <code>WEEK_NUM</code>, one can see that the default ratio doesn't really change much over the time, even though the total number of cases per week could change dramatically:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F54e2ab7b4623137a7c2cbb00aaacf063%2FScreenshot%202024-02-18%20at%2010.19.08%20PM.png?generation=1708323569689791&amp;alt=media\"></p>\n<p>It means that when we stratify our folds by weeks, each of our models trains on 2-5% of positive targets and on 95-98% of negative ones. As a result, such an imbalance can affect model's ability to properly predict defaults (positive targets). Luckily, in this competition we mostly care about AUC, and not about actual probabilities, so can give a try to built-in class balancing features from popular gradient boosting libraries.</p>\n<h1>XGBoost</h1>\n<p>For <code>XGBClassifier</code> we can use <a href=\"https://xgboost.readthedocs.io/en/stable/tutorials/param_tuning.html#handle-imbalanced-dataset\" target=\"_blank\">scale_pos_weight</a> option, that scales the positive class weight. According to <a href=\"https://xgboost.readthedocs.io/en/stable/parameter.html\" target=\"_blank\">documentation</a>, a typical value for this option is <code>sum(negative targets) / sum(positive targets)</code>, i.e. in our case it is around <code>31</code>. This value may need to be adjusted for each fold separately, as the default ratio varies over different time periods.</p>\n<p><strong>UPD:</strong> There is also a way to assign weight to each training sample separately. For that, one needs to convert training data to <a href=\"https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.core\" target=\"_blank\">DMatrix</a> object passing appropriate <code>weight</code> vector when creating it. Thanks to <a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> for pointing this out.</p>\n<h1>LightGBM</h1>\n<p>For <code>LGBMClassifier</code> there is a similar <code>scale_pos_weight</code> option to scale weight of positive class. In addition, there is a boolean flag <code>is_unbalanced</code>, that automatically assigns weights based on the ratio between the numbers of negative and positive classes. According to official <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters.html#is_unbalance\" target=\"_blank\">documentation</a> you can only use one of these options at a time, that makes sense as there is only one set of weights the model can use for training.</p>\n<h1>CatBoost</h1>\n<p>Just like the other libraries, <code>CatBoostClassifier</code> supports <a href=\"https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight\" target=\"_blank\">scale_pos_weight</a> option. There is also a more advanced <a href=\"https://catboost.ai/en/docs/references/training-parameters/common#auto_class_weights\" target=\"_blank\">auto_class_weights</a> feature, that can automatically adjust weights in <code>Balanced</code> and <code>SqrtBalanced</code> modes. You can use only one of these options to balance training.</p>\n<h1>Conclusions</h1>\n<p>Class weights, as all the other model parameters, may need to be tuned. In my experiments, setting <code>scale_pos_weight</code> to recommended typical values improved AUC by 0.01-0.02 for all the three libraries. The best LB improvement, however, was at the level of more humble 0.005 for LightGBM, but this is yet to be seen when the metric is fixed.</p>\n<p>Besides adjusting class weights to deal with imbalance, one can also employ other techniques like down-/upsampling of data or postprocessing. For more information on this please refer to the following notebooks and discussions:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/shahules/tackling-class-imbalance\" target=\"_blank\">tackling-class-imbalance</a></li>\n<li><a href=\"https://www.kaggle.com/code/kaanboke/xgboost-lightgbm-catboost-imbalanced-data\" target=\"_blank\">xgboost-lightgbm-catboost-imbalanced-data</a></li>\n<li><a href=\"https://www.kaggle.com/code/janiobachmann/credit-fraud-dealing-with-imbalanced-datasets\" target=\"_blank\">credit-fraud-dealing-with-imbalanced-datasets</a></li>\n<li><a href=\"https://www.kaggle.com/code/kailex/talkingdata-eda-and-class-imbalance\" target=\"_blank\">talkingdata-eda-and-class-imbalance</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/icr-identify-age-related-conditions/discussion/412507\" target=\"_blank\">How To Balance Training And Boost CV and LB Score!</a></li>\n</ul>\n<p>Hope it helps and happy kaggling!</p>",
      "rawMarkdown": "In this competition we have to deal with a highly imbalanced class distribution, as there are only about 3% of positive targets representing defaults:\n\n|                  | [k]    | [%] |\n|------------------|--------|-----|\n| Positive targets | 48     | 3   |\n| Negative targets | 1,479  | 97  |\n| Total cases      | 1,527  | 100 |\n\n<br>\n\nAfter grouping data by `WEEK_NUM`, one can see that the default ratio doesn't really change much over the time, even though the total number of cases per week could change dramatically:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F54e2ab7b4623137a7c2cbb00aaacf063%2FScreenshot%202024-02-18%20at%2010.19.08%20PM.png?generation=1708323569689791&alt=media)\n\nIt means that when we stratify our folds by weeks, each of our models trains on 2-5% of positive targets and on 95-98% of negative ones. As a result, such an imbalance can affect model's ability to properly predict defaults (positive targets). Luckily, in this competition we mostly care about AUC, and not about actual probabilities, so can give a try to built-in class balancing features from popular gradient boosting libraries.\n\n# XGBoost\n\nFor `XGBClassifier` we can use [scale_pos_weight](https://xgboost.readthedocs.io/en/stable/tutorials/param_tuning.html#handle-imbalanced-dataset) option, that scales the positive class weight. According to [documentation](https://xgboost.readthedocs.io/en/stable/parameter.html), a typical value for this option is `sum(negative targets) / sum(positive targets)`, i.e. in our case it is around `31`. This value may need to be adjusted for each fold separately, as the default ratio varies over different time periods.\n\n**UPD:** There is also a way to assign weight to each training sample separately. For that, one needs to convert training data to [DMatrix](https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.core) object passing appropriate `weight` vector when creating it. Thanks to @thomasmeiner for pointing this out.\n\n# LightGBM\n\nFor `LGBMClassifier` there is a similar `scale_pos_weight` option to scale weight of positive class. In addition, there is a boolean flag `is_unbalanced`, that automatically assigns weights based on the ratio between the numbers of negative and positive classes. According to official [documentation](https://lightgbm.readthedocs.io/en/latest/Parameters.html#is_unbalance) you can only use one of these options at a time, that makes sense as there is only one set of weights the model can use for training.\n\n# CatBoost\n\nJust like the other libraries, `CatBoostClassifier` supports [scale_pos_weight](https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight) option. There is also a more advanced [auto_class_weights](https://catboost.ai/en/docs/references/training-parameters/common#auto_class_weights) feature, that can automatically adjust weights in `Balanced` and `SqrtBalanced` modes. You can use only one of these options to balance training.\n\n# Conclusions\n\nClass weights, as all the other model parameters, may need to be tuned. In my experiments, setting `scale_pos_weight` to recommended typical values improved AUC by 0.01-0.02 for all the three libraries. The best LB improvement, however, was at the level of more humble 0.005 for LightGBM, but this is yet to be seen when the metric is fixed.\n\nBesides adjusting class weights to deal with imbalance, one can also employ other techniques like down-/upsampling of data or postprocessing. For more information on this please refer to the following notebooks and discussions:\n\n- [tackling-class-imbalance](https://www.kaggle.com/code/shahules/tackling-class-imbalance)\n- [xgboost-lightgbm-catboost-imbalanced-data](https://www.kaggle.com/code/kaanboke/xgboost-lightgbm-catboost-imbalanced-data)\n- [credit-fraud-dealing-with-imbalanced-datasets](https://www.kaggle.com/code/janiobachmann/credit-fraud-dealing-with-imbalanced-datasets)\n- [talkingdata-eda-and-class-imbalance](https://www.kaggle.com/code/kailex/talkingdata-eda-and-class-imbalance)\n- [How To Balance Training And Boost CV and LB Score!](https://www.kaggle.com/competitions/icr-identify-age-related-conditions/discussion/412507)\n\nHope it helps and happy kaggling!",
      "votes": null
    },
    {
      "id": "2665271",
      "postDate": "02/23/2024 15:13:12",
      "content": "<p>Thanks, that's interesting insights. Just curious, will results be similar if data is grouped not by weeks but months? Another thought is about whether ensembling different methods (e.g., XGBoost and Catboost) can imporove performance somehow. </p>",
      "rawMarkdown": "Thanks, that's interesting insights. Just curious, will results be similar if data is grouped not by weeks but months? Another thought is about whether ensembling different methods (e.g., XGBoost and Catboost) can imporove performance somehow.",
      "votes": null
    },
    {
      "id": "2665768",
      "postDate": "02/23/2024 20:22:10",
      "content": "<ul>\n<li><p>I guess grouping by months should result in a similar performance, but it is good to double check that on real data;</p></li>\n<li><p>Combining different models is almost always beneficial, when you ensemble diverse models. I didn’t work on ensembling in this competition yet, mostly trying to build a set of good single models first.</p></li>\n</ul>",
      "rawMarkdown": "I guess grouping by months should result in a similar performance, but it is good to double check that on real data;\n\n- Combining different models is almost always beneficial, when you ensemble diverse models. I didn’t work on ensembling in this competition yet, mostly trying to build a set of good single models first.",
      "votes": null
    },
    {
      "id": "2674453",
      "postDate": "02/29/2024 10:39:24",
      "content": "<p>What auc roc did you get after using scale_pos_weight? My score is about 0.84 in cv, but weight scaling does not work so well in my case. <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> </p>",
      "rawMarkdown": "What auc roc did you get after using scale_pos_weight? My score is about 0.84 in cv, but weight scaling does not work so well in my case. @kononenko",
      "votes": null
    },
    {
      "id": "2674465",
      "postDate": "02/29/2024 10:51:08",
      "content": "<p>Check out the first paragraph in the Conclusions section. I didn’t get better than that, but I only tested it for depth 1 features. From your CV it seems you are already using depth 2, right?</p>",
      "rawMarkdown": "Check out the first paragraph in the Conclusions section. I didn’t get better than that, but I only tested it for depth 1 features. From your CV it seems you are already using depth 2, right?",
      "votes": null
    },
    {
      "id": "2674483",
      "postDate": "02/29/2024 11:00:29",
      "content": "<p>Yes, I applied depth 2. </p>",
      "rawMarkdown": "Yes, I applied depth 2.",
      "votes": null
    },
    {
      "id": "2674496",
      "postDate": "02/29/2024 11:07:54",
      "content": "<p>Right, so for depth 1 it roughly changed AUC from 0.79-0.8 to 0.81-0.82 depending on a model.</p>",
      "rawMarkdown": "Right, so for depth 1 it roughly changed AUC from 0.79-0.8 to 0.81-0.82 depending on a model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2665271,
      "author_name": "ablaydosmaganbetov",
      "author_url": "",
      "post_date": "02/23/2024 15:13:12",
      "content": "<p>Thanks, that's interesting insights. Just curious, will results be similar if data is grouped not by weeks but months? Another thought is about whether ensembling different methods (e.g., XGBoost and Catboost) can imporove performance somehow. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2665768,
          "author_name": "kononenko",
          "author_url": "",
          "post_date": "02/23/2024 20:22:10",
          "content": "<ul>\n<li><p>I guess grouping by months should result in a similar performance, but it is good to double check that on real data;</p></li>\n<li><p>Combining different models is almost always beneficial, when you ensemble diverse models. I didn’t work on ensembling in this competition yet, mostly trying to build a set of good single models first.</p></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2674453,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "02/29/2024 10:39:24",
      "content": "<p>What auc roc did you get after using scale_pos_weight? My score is about 0.84 in cv, but weight scaling does not work so well in my case. <a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2674465,
          "author_name": "kononenko",
          "author_url": "",
          "post_date": "02/29/2024 10:51:08",
          "content": "<p>Check out the first paragraph in the Conclusions section. I didn’t get better than that, but I only tested it for depth 1 features. From your CV it seems you are already using depth 2, right?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2674483,
              "author_name": "jankowalski2000",
              "author_url": "",
              "post_date": "02/29/2024 11:00:29",
              "content": "<p>Yes, I applied depth 2. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2674496,
                  "author_name": "kononenko",
                  "author_url": "",
                  "post_date": "02/29/2024 11:07:54",
                  "content": "<p>Right, so for depth 1 it roughly changed AUC from 0.79-0.8 to 0.81-0.82 depending on a model.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2658363": "In this competition we have to deal with a highly imbalanced class distribution, as there are only about 3% of positive targets representing defaults:\n\n|                  | [k]    | [%] |\n|------------------|--------|-----|\n| Positive targets | 48     | 3   |\n| Negative targets | 1,479  | 97  |\n| Total cases      | 1,527  | 100 |\n\n<br>\n\nAfter grouping data by `WEEK_NUM`, one can see that the default ratio doesn't really change much over the time, even though the total number of cases per week could change dramatically:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2099265%2F54e2ab7b4623137a7c2cbb00aaacf063%2FScreenshot%202024-02-18%20at%2010.19.08%20PM.png?generation=1708323569689791&alt=media)\n\nIt means that when we stratify our folds by weeks, each of our models trains on 2-5% of positive targets and on 95-98% of negative ones. As a result, such an imbalance can affect model's ability to properly predict defaults (positive targets). Luckily, in this competition we mostly care about AUC, and not about actual probabilities, so can give a try to built-in class balancing features from popular gradient boosting libraries.\n\n# XGBoost\n\nFor `XGBClassifier` we can use [scale_pos_weight](https://xgboost.readthedocs.io/en/stable/tutorials/param_tuning.html#handle-imbalanced-dataset) option, that scales the positive class weight. According to [documentation](https://xgboost.readthedocs.io/en/stable/parameter.html), a typical value for this option is `sum(negative targets) / sum(positive targets)`, i.e. in our case it is around `31`. This value may need to be adjusted for each fold separately, as the default ratio varies over different time periods.\n\n**UPD:** There is also a way to assign weight to each training sample separately. For that, one needs to convert training data to [DMatrix](https://xgboost.readthedocs.io/en/stable/python/python_api.html#module-xgboost.core) object passing appropriate `weight` vector when creating it. Thanks to @thomasmeiner for pointing this out.\n\n# LightGBM\n\nFor `LGBMClassifier` there is a similar `scale_pos_weight` option to scale weight of positive class. In addition, there is a boolean flag `is_unbalanced`, that automatically assigns weights based on the ratio between the numbers of negative and positive classes. According to official [documentation](https://lightgbm.readthedocs.io/en/latest/Parameters.html#is_unbalance) you can only use one of these options at a time, that makes sense as there is only one set of weights the model can use for training.\n\n# CatBoost\n\nJust like the other libraries, `CatBoostClassifier` supports [scale_pos_weight](https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight) option. There is also a more advanced [auto_class_weights](https://catboost.ai/en/docs/references/training-parameters/common#auto_class_weights) feature, that can automatically adjust weights in `Balanced` and `SqrtBalanced` modes. You can use only one of these options to balance training.\n\n# Conclusions\n\nClass weights, as all the other model parameters, may need to be tuned. In my experiments, setting `scale_pos_weight` to recommended typical values improved AUC by 0.01-0.02 for all the three libraries. The best LB improvement, however, was at the level of more humble 0.005 for LightGBM, but this is yet to be seen when the metric is fixed.\n\nBesides adjusting class weights to deal with imbalance, one can also employ other techniques like down-/upsampling of data or postprocessing. For more information on this please refer to the following notebooks and discussions:\n\n- [tackling-class-imbalance](https://www.kaggle.com/code/shahules/tackling-class-imbalance)\n- [xgboost-lightgbm-catboost-imbalanced-data](https://www.kaggle.com/code/kaanboke/xgboost-lightgbm-catboost-imbalanced-data)\n- [credit-fraud-dealing-with-imbalanced-datasets](https://www.kaggle.com/code/janiobachmann/credit-fraud-dealing-with-imbalanced-datasets)\n- [talkingdata-eda-and-class-imbalance](https://www.kaggle.com/code/kailex/talkingdata-eda-and-class-imbalance)\n- [How To Balance Training And Boost CV and LB Score!](https://www.kaggle.com/competitions/icr-identify-age-related-conditions/discussion/412507)\n\nHope it helps and happy kaggling!",
    "2665271": "Thanks, that's interesting insights. Just curious, will results be similar if data is grouped not by weeks but months? Another thought is about whether ensembling different methods (e.g., XGBoost and Catboost) can imporove performance somehow.",
    "2665768": "I guess grouping by months should result in a similar performance, but it is good to double check that on real data;\n\n- Combining different models is almost always beneficial, when you ensemble diverse models. I didn’t work on ensembling in this competition yet, mostly trying to build a set of good single models first.",
    "2674453": "What auc roc did you get after using scale_pos_weight? My score is about 0.84 in cv, but weight scaling does not work so well in my case. @kononenko",
    "2674465": "Check out the first paragraph in the Conclusions section. I didn’t get better than that, but I only tested it for depth 1 features. From your CV it seems you are already using depth 2, right?",
    "2674483": "Yes, I applied depth 2.",
    "2674496": "Right, so for depth 1 it roughly changed AUC from 0.79-0.8 to 0.81-0.82 depending on a model."
  },
  "source": "meta"
}