{
  "id": 333953,
  "title": "What All we can try to improve our model performance?",
  "url": "/competitions/amex-default-prediction/discussion/333953",
  "author_name": "Jaswinder Singh",
  "post_date": "2022-06-29T06:12:16.292000",
  "votes": 34,
  "comment_count": 16,
  "views": 0,
  "content": "<p>As I have seen and tried lot of notebook public on this Competition. I am just trying to conclude what all i can try to improve my scores.<br>\n<strong>Algorithums</strong>- XGB,LGB, Catboost, Tabnet, RNN( GRU/LSTM), Pyboost , AutoMl models<br>\n<strong>Feature Engineering</strong>- Use Only given 188 features, ADD mean max min std type features, add A-B type features, add A/B or A*B type features or build some new features around that.<br>\n<strong>HP tuning</strong>- Specific to Models.<br>\n<strong>Different seed</strong><br>\n<strong>Ensemble multiple models</strong>     <br>\nTill now I have tried more than 20-30 iterations of above techniques  and getting a score of 0.797 ensembling 2 of my best models. and planning to try more combination of above things.</p>\n<p>Hope this will help others in this competition and Please let me know if a have missed any important thing to try to improve my results. and Don't forget to Upvote if you like this discussion.</p>",
  "messages": [
    {
      "id": 1836833,
      "postDate": "2022-06-29T06:12:16.293Z",
      "content": "<p>As I have seen and tried lot of notebook public on this Competition. I am just trying to conclude what all i can try to improve my scores.<br>\n<strong>Algorithums</strong>- XGB,LGB, Catboost, Tabnet, RNN( GRU/LSTM), Pyboost , AutoMl models<br>\n<strong>Feature Engineering</strong>- Use Only given 188 features, ADD mean max min std type features, add A-B type features, add A/B or A*B type features or build some new features around that.<br>\n<strong>HP tuning</strong>- Specific to Models.<br>\n<strong>Different seed</strong><br>\n<strong>Ensemble multiple models</strong>     <br>\nTill now I have tried more than 20-30 iterations of above techniques  and getting a score of 0.797 ensembling 2 of my best models. and planning to try more combination of above things.</p>\n<p>Hope this will help others in this competition and Please let me know if a have missed any important thing to try to improve my results. and Don't forget to Upvote if you like this discussion.</p>",
      "rawMarkdown": "As I have seen and tried lot of notebook public on this Competition. I am just trying to conclude what all i can try to improve my scores.\n**Algorithums**- XGB,LGB, Catboost, Tabnet, RNN( GRU/LSTM), Pyboost , AutoMl models\n**Feature Engineering**- Use Only given 188 features, ADD mean max min std type features, add A-B type features, add A/B or A*B type features or build some new features around that.\n**HP tuning**- Specific to Models.\n**Different seed**\n**Ensemble multiple models**     \nTill now I have tried more than 20-30 iterations of above techniques  and getting a score of 0.797 ensembling 2 of my best models. and planning to try more combination of above things.\n\nHope this will help others in this competition and Please let me know if a have missed any important thing to try to improve my results. and Don't forget to Upvote if you like this discussion.\n",
      "votes": 34
    },
    {
      "id": 1836937,
      "postDate": "2022-06-29T08:17:21.763Z",
      "content": "<p>also include lgbm DART boosting strategy. It is slow but improves the accuracy of your model.</p>",
      "rawMarkdown": "also include lgbm DART boosting strategy. It is slow but improves the accuracy of your model.",
      "votes": 7,
      "replies": [
        {
          "id": 1836955,
          "postDate": "2022-06-29T08:31:50.130Z",
          "content": "<p>Thanks for Idea <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> . But both of my current best models are LGB with dart. That's give a significant positive impact to scores. </p>",
          "rawMarkdown": "Thanks for Idea @mohammadrahmati . But both of my current best models are LGB with dart. That's give a significant positive impact to scores. ",
          "votes": 3
        },
        {
          "id": 1837638,
          "postDate": "2022-06-29T18:44:07.063Z",
          "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> If you have a non-DART lgbm model, can you expect an immediate improvement by changing to DART? Or would you also spend a lot of time readjusting the other hyperparameters?</p>",
          "rawMarkdown": "@mohammadrahmati If you have a non-DART lgbm model, can you expect an immediate improvement by changing to DART? Or would you also spend a lot of time readjusting the other hyperparameters?",
          "votes": 2
        },
        {
          "id": 1837733,
          "postDate": "2022-06-29T21:01:14.593Z",
          "content": "<p><a href=\"https://www.kaggle.com/datahobbit\" target=\"_blank\">@datahobbit</a> I would say yes. At least it worked for one of my models. I would suggest choosing the default parameters and let it run for a few thousands of rounds. </p>",
          "rawMarkdown": "@datahobbit I would say yes. At least it worked for one of my models. I would suggest choosing the default parameters and let it run for a few thousands of rounds. ",
          "votes": 1
        },
        {
          "id": 1837745,
          "postDate": "2022-06-29T21:09:23.437Z",
          "content": "<p>OK great thank you. At least with lgbm the slowdown from 'gbdt' is noticeable but not totally awful. I tried it first with xgboost but that switches from GPU to CPU and it became totally impractical to run enough rounds. </p>",
          "rawMarkdown": "OK great thank you. At least with lgbm the slowdown from 'gbdt' is noticeable but not totally awful. I tried it first with xgboost but that switches from GPU to CPU and it became totally impractical to run enough rounds. ",
          "votes": 3
        },
        {
          "id": 1837751,
          "postDate": "2022-06-29T21:29:37.493Z",
          "content": "<p>My system takes around one hour to finish 15000 rounds per fold on CPU (around 5 hours to finish all folds). <br>\nyou may find this useful: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575</a></p>",
          "rawMarkdown": "My system takes around one hour to finish 15000 rounds per fold on CPU (around 5 hours to finish all folds). \nyou may find this useful: https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575",
          "votes": 2
        },
        {
          "id": 1837799,
          "postDate": "2022-06-30T00:04:24.527Z",
          "content": "<blockquote>\n  <p>I tried it first with xgboost but that switches from GPU to CPU and it became totally impractical to run enough rounds.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/datahobbit\" target=\"_blank\">@datahobbit</a> How do you know that it switched to CPU? When i added the parameter <code>'booster': 'dart'</code> to my XGB model and then ran <code>nvidia-smi</code>, the GPU was being used and the CV improved.</p>\n<p>Additionally, i use parameters:</p>\n<pre><code>    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n</code></pre>",
          "rawMarkdown": ">I tried it first with xgboost but that switches from GPU to CPU and it became totally impractical to run enough rounds.\n\n @datahobbit How do you know that it switched to CPU? When i added the parameter `'booster': 'dart'` to my XGB model and then ran `nvidia-smi`, the GPU was being used and the CV improved.\n\nAdditionally, i use parameters:\n\n        'tree_method':'gpu_hist',\n        'predictor':'gpu_predictor',",
          "votes": 3
        },
        {
          "id": 1837917,
          "postDate": "2022-06-30T03:43:38.600Z",
          "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the advice. Chris, I don't know for certain it switched to GPU but the speed dropped by close to 100x so I just assumed. I will try the settings you provided - thanks again.</p>",
          "rawMarkdown": "@mohammadrahmati and @cdeotte thanks for the advice. Chris, I don't know for certain it switched to GPU but the speed dropped by close to 100x so I just assumed. I will try the settings you provided - thanks again.",
          "votes": 1
        },
        {
          "id": 1837965,
          "postDate": "2022-06-30T04:38:13.487Z",
          "content": "<p>It was much slower for me too. Maybe 10x or 20x slower, but i checked and it was using GPU at 85%. I'm not sure why DART is so slow.</p>",
          "rawMarkdown": "It was much slower for me too. Maybe 10x or 20x slower, but i checked and it was using GPU at 85%. I'm not sure why DART is so slow.",
          "votes": 5
        },
        {
          "id": 1837981,
          "postDate": "2022-06-30T05:03:43.293Z",
          "content": "<p>OK, thanks that's helpful. </p>",
          "rawMarkdown": "OK, thanks that's helpful. ",
          "votes": 1
        },
        {
          "id": 1854176,
          "postDate": "2022-07-13T13:31:21.523Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1839836,
      "postDate": "2022-07-01T18:06:05.093Z",
      "content": "<p>I would also add that trying different subsets of features can also be helpful (e.g. removing correlated features, or features that don't have much predictive power, etc.). Additionally, feature engineering is an important part of any machine learning competition, so spending time brainstorming new features to add can also be very beneficial. </p>",
      "rawMarkdown": "I would also add that trying different subsets of features can also be helpful (e.g. removing correlated features, or features that don't have much predictive power, etc.). Additionally, feature engineering is an important part of any machine learning competition, so spending time brainstorming new features to add can also be very beneficial. \n\n\n",
      "votes": 3
    },
    {
      "id": 1866531,
      "postDate": "2022-07-22T14:56:16.070Z",
      "content": "<p>hey~ what do you think about how much influence to the model s given by the seed,  try different is really cost time. </p>",
      "rawMarkdown": "hey~ what do you think about how much influence to the model s given by the seed,  try different is really cost time. "
    },
    {
      "id": 1840951,
      "postDate": "2022-07-02T17:07:14.160Z",
      "content": "<p>I have got the best results with an ANN model using Keras/TensorFlow. Also I have done oversampling since data is unbalanced.</p>",
      "rawMarkdown": "I have got the best results with an ANN model using Keras/TensorFlow. Also I have done oversampling since data is unbalanced."
    },
    {
      "id": 1837884,
      "postDate": "2022-06-30T02:59:46.140Z",
      "content": "<p>Hello, thanks for sharing.</p>",
      "rawMarkdown": "Hello, thanks for sharing.\n",
      "votes": 1
    },
    {
      "id": 1840005,
      "postDate": "2022-07-01T21:03:30.613Z",
      "content": "<p>Thanks for such a concise explanation.</p>",
      "rawMarkdown": "Thanks for such a concise explanation."
    }
  ],
  "comments": [
    {
      "id": 1836937,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-06-29T08:17:21.763000",
      "content": "<p>also include lgbm DART boosting strategy. It is slow but improves the accuracy of your model.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1836955,
          "author_name": "Jaswinder Singh",
          "author_url": "",
          "post_date": "2022-06-29T08:31:50.130000",
          "content": "<p>Thanks for Idea <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> . But both of my current best models are LGB with dart. That's give a significant positive impact to scores. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1837638,
          "author_name": "MichaelP",
          "author_url": "",
          "post_date": "2022-06-29T18:44:07.063000",
          "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> If you have a non-DART lgbm model, can you expect an immediate improvement by changing to DART? Or would you also spend a lot of time readjusting the other hyperparameters?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1837733,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-29T21:01:14.593000",
          "content": "<p><a href=\"https://www.kaggle.com/datahobbit\" target=\"_blank\">@datahobbit</a> I would say yes. At least it worked for one of my models. I would suggest choosing the default parameters and let it run for a few thousands of rounds. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1837745,
          "author_name": "MichaelP",
          "author_url": "",
          "post_date": "2022-06-29T21:09:23.437000",
          "content": "<p>OK great thank you. At least with lgbm the slowdown from 'gbdt' is noticeable but not totally awful. I tried it first with xgboost but that switches from GPU to CPU and it became totally impractical to run enough rounds. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1837751,
          "author_name": "1110Ra",
          "author_url": "",
          "post_date": "2022-06-29T21:29:37.493000",
          "content": "<p>My system takes around one hour to finish 15000 rounds per fold on CPU (around 5 hours to finish all folds). <br>\nyou may find this useful: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332575</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1837799,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-30T00:04:24.527000",
          "content": "<blockquote>\n  <p>I tried it first with xgboost but that switches from GPU to CPU and it became totally impractical to run enough rounds.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/datahobbit\" target=\"_blank\">@datahobbit</a> How do you know that it switched to CPU? When i added the parameter <code>'booster': 'dart'</code> to my XGB model and then ran <code>nvidia-smi</code>, the GPU was being used and the CV improved.</p>\n<p>Additionally, i use parameters:</p>\n<pre><code>    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1837917,
          "author_name": "MichaelP",
          "author_url": "",
          "post_date": "2022-06-30T03:43:38.600000",
          "content": "<p><a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the advice. Chris, I don't know for certain it switched to GPU but the speed dropped by close to 100x so I just assumed. I will try the settings you provided - thanks again.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1837965,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-30T04:38:13.487000",
          "content": "<p>It was much slower for me too. Maybe 10x or 20x slower, but i checked and it was using GPU at 85%. I'm not sure why DART is so slow.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1837981,
          "author_name": "MichaelP",
          "author_url": "",
          "post_date": "2022-06-30T05:03:43.293000",
          "content": "<p>OK, thanks that's helpful. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1854176,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-13T13:31:21.523000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1839836,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-01T18:06:05.093000",
      "content": "<p>I would also add that trying different subsets of features can also be helpful (e.g. removing correlated features, or features that don't have much predictive power, etc.). Additionally, feature engineering is an important part of any machine learning competition, so spending time brainstorming new features to add can also be very beneficial. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1866531,
      "author_name": "xianZ_waikato",
      "author_url": "",
      "post_date": "2022-07-22T14:56:16.070000",
      "content": "<p>hey~ what do you think about how much influence to the model s given by the seed,  try different is really cost time. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1840951,
      "author_name": "Johnny Torres",
      "author_url": "",
      "post_date": "2022-07-02T17:07:14.160000",
      "content": "<p>I have got the best results with an ANN model using Keras/TensorFlow. Also I have done oversampling since data is unbalanced.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1837884,
      "author_name": "Praise the Lord Jesus",
      "author_url": "",
      "post_date": "2022-06-30T02:59:46.140000",
      "content": "<p>Hello, thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1840005,
      "author_name": "Sohail Ahmed",
      "author_url": "",
      "post_date": "2022-07-01T21:03:30.613000",
      "content": "<p>Thanks for such a concise explanation.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1836833": "As I have seen and tried lot of notebook public on this Competition. I am just trying to conclude what all i can try to improve my scores.\n**Algorithums**- XGB,LGB, Catboost, Tabnet, RNN( GRU/LSTM), Pyboost , AutoMl models\n**Feature Engineering**- Use Only given 188 features, ADD mean max min std type features, add A-B type features, add A/B or A*B type features or build some new features around that.\n**HP tuning**- Specific to Models.\n**Different seed**\n**Ensemble multiple models**     \nTill now I have tried more than 20-30 iterations of above techniques  and getting a score of 0.797 ensembling 2 of my best models. and planning to try more combination of above things.\n\nHope this will help others in this competition and Please let me know if a have missed any important thing to try to improve my results. and Don't forget to Upvote if you like this discussion.\n",
    "1836937": "also include lgbm DART boosting strategy. It is slow but improves the accuracy of your model.",
    "1839836": "I would also add that trying different subsets of features can also be helpful (e.g. removing correlated features, or features that don't have much predictive power, etc.). Additionally, feature engineering is an important part of any machine learning competition, so spending time brainstorming new features to add can also be very beneficial. \n\n\n",
    "1866531": "hey~ what do you think about how much influence to the model s given by the seed,  try different is really cost time. ",
    "1840951": "I have got the best results with an ANN model using Keras/TensorFlow. Also I have done oversampling since data is unbalanced.",
    "1837884": "Hello, thanks for sharing.\n",
    "1840005": "Thanks for such a concise explanation."
  }
}