{
  "id": 347802,
  "title": "My first Kaggle competition: the learning is not over yet!",
  "url": "/competitions/amex-default-prediction/discussion/347802",
  "author_name": "",
  "post_date": "2022-08-25T12:47:59.903741900Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>It's nice to see that after some sleepless nights I managed to wrap up my first Kaggle competition on a positive note: jumped up 389 positions on the private leaderboard. Even though the final rank is 2067 (top 42%), I'm pretty stoked about my progress so far. </p>\n<p>Here are a couple of things I've learned during this competition:</p>\n<ul>\n<li>How to work with memory limitations</li>\n<li>How to work with RAPIDS</li>\n<li>Feature engineering ideas (my favorite trick was to combine max and min into a difference of max - min feature; it achieved the same score as having max and min as separate sets of features freeing up almost 300 data columns)</li>\n<li>I'm not so good at feature selection: I tried permutation importance, SHAP, Boruta, adversarial validation, tree feature importance and I couldn't get to a point where I felt comfortable with any of the results</li>\n<li>DART boosting - a slow but quite accurate boosting algorithm</li>\n<li>Ensembling works well: two weak models outscored my best single model</li>\n<li>Model validation - I managed to handpick the best model and the private score went up which is a positive note for me</li>\n</ul>\n<p>What I plan to learn from the top solutions once/if they become public:</p>\n<ul>\n<li>Knowledge distillation: seemed to have played an important role in this competition</li>\n<li>Become better at feature engineering:<ul>\n<li>check impact of dimensionality reduction techniques (umap/PCA etc.)</li>\n<li>find features that describe the 13-period time series more accurately (wasn't able to find the sweet spot during the competition)</li>\n<li>sky is the limit here, or maybe the RAM size :)</li></ul></li>\n<li>Feature selection with <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDS FIL</a></li>\n</ul>\n<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> for sharing your work and your findings!</p>\n<p>On to the next one!</p>",
  "messages": [
    {
      "id": "1913676",
      "postDate": "08/25/2022 12:47:59",
      "content": "<p>It's nice to see that after some sleepless nights I managed to wrap up my first Kaggle competition on a positive note: jumped up 389 positions on the private leaderboard. Even though the final rank is 2067 (top 42%), I'm pretty stoked about my progress so far. </p>\n<p>Here are a couple of things I've learned during this competition:</p>\n<ul>\n<li>How to work with memory limitations</li>\n<li>How to work with RAPIDS</li>\n<li>Feature engineering ideas (my favorite trick was to combine max and min into a difference of max - min feature; it achieved the same score as having max and min as separate sets of features freeing up almost 300 data columns)</li>\n<li>I'm not so good at feature selection: I tried permutation importance, SHAP, Boruta, adversarial validation, tree feature importance and I couldn't get to a point where I felt comfortable with any of the results</li>\n<li>DART boosting - a slow but quite accurate boosting algorithm</li>\n<li>Ensembling works well: two weak models outscored my best single model</li>\n<li>Model validation - I managed to handpick the best model and the private score went up which is a positive note for me</li>\n</ul>\n<p>What I plan to learn from the top solutions once/if they become public:</p>\n<ul>\n<li>Knowledge distillation: seemed to have played an important role in this competition</li>\n<li>Become better at feature engineering:<ul>\n<li>check impact of dimensionality reduction techniques (umap/PCA etc.)</li>\n<li>find features that describe the 13-period time series more accurately (wasn't able to find the sweet spot during the competition)</li>\n<li>sky is the limit here, or maybe the RAM size :)</li></ul></li>\n<li>Feature selection with <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDS FIL</a></li>\n</ul>\n<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> for sharing your work and your findings!</p>\n<p>On to the next one!</p>",
      "rawMarkdown": "It's nice to see that after some sleepless nights I managed to wrap up my first Kaggle competition on a positive note: jumped up 389 positions on the private leaderboard. Even though the final rank is 2067 (top 42%), I'm pretty stoked about my progress so far. \n\nHere are a couple of things I've learned during this competition:\n- How to work with memory limitations\n- How to work with RAPIDS\n- Feature engineering ideas (my favorite trick was to combine max and min into a difference of max - min feature; it achieved the same score as having max and min as separate sets of features freeing up almost 300 data columns)\n- I'm not so good at feature selection: I tried permutation importance, SHAP, Boruta, adversarial validation, tree feature importance and I couldn't get to a point where I felt comfortable with any of the results\n- DART boosting - a slow but quite accurate boosting algorithm\n- Ensembling works well: two weak models outscored my best single model\n- Model validation - I managed to handpick the best model and the private score went up which is a positive note for me\n\nWhat I plan to learn from the top solutions once/if they become public:\n - Knowledge distillation: seemed to have played an important role in this competition\n - Become better at feature engineering:\n \t- check impact of dimensionality reduction techniques (umap/PCA etc.)\n \t- find features that describe the 13-period time series more accurately (wasn't able to find the sweet spot during the competition)\n \t- sky is the limit here, or maybe the RAM size :)\n - Feature selection with [RAPIDS FIL](https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing)\n\nThanks @cdeotte @raddar @ragnar123 @illidan7 @roberthatch for sharing your work and your findings!\n\nOn to the next one!",
      "votes": null
    },
    {
      "id": "1913723",
      "postDate": "08/25/2022 13:18:55",
      "content": "<p>This was a great challenge and an invaluable learning experience for me too! Many thanks for the public datasets and discussion posts that helped me a lot, special mention to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <a href=\"https://www.kaggle.com/tilli7\" target=\"_blank\">@tilli7</a> <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> for their work!</p>",
      "rawMarkdown": "This was a great challenge and an invaluable learning experience for me too! Many thanks for the public datasets and discussion posts that helped me a lot, special mention to @cdeotte @raddar @ragnar123 @illidan7 @roberthatch @ambrosm @tilli7 @thedevastator for their work!",
      "votes": null
    },
    {
      "id": "1914076",
      "postDate": "08/25/2022 17:51:42",
      "content": "<p>My first time too! I had many of the same learnings :)</p>\n<p>Really glad to hear my notebooks were of some benefit to you :)</p>",
      "rawMarkdown": "My first time too! I had many of the same learnings :)\n\nReally glad to hear my notebooks were of some benefit to you :)",
      "votes": null
    },
    {
      "id": "1922173",
      "postDate": "09/01/2022 10:13:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/canonicalized\" target=\"_blank\">@canonicalized</a>, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @canonicalized, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    },
    {
      "id": "2311555",
      "postDate": "06/21/2023 08:32:56",
      "content": "<p><strong><em>Hi there! My name is Arif Patel and I am a developer from Dubai. I'm Also A Newbie Here</em></strong></p>",
      "rawMarkdown": "***Hi there! My name is Arif Patel and I am a developer from Dubai. I'm Also A Newbie Here***",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1913723,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/25/2022 13:18:55",
      "content": "<p>This was a great challenge and an invaluable learning experience for me too! Many thanks for the public datasets and discussion posts that helped me a lot, special mention to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/illidan7\" target=\"_blank\">@illidan7</a> <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> <a href=\"https://www.kaggle.com/tilli7\" target=\"_blank\">@tilli7</a> <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> for their work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1914076,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/25/2022 17:51:42",
      "content": "<p>My first time too! I had many of the same learnings :)</p>\n<p>Really glad to hear my notebooks were of some benefit to you :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1922173,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 10:13:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/canonicalized\" target=\"_blank\">@canonicalized</a>, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2311555,
      "author_name": "arifpatelpduk",
      "author_url": "",
      "post_date": "06/21/2023 08:32:56",
      "content": "<p><strong><em>Hi there! My name is Arif Patel and I am a developer from Dubai. I'm Also A Newbie Here</em></strong></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1913676": "It's nice to see that after some sleepless nights I managed to wrap up my first Kaggle competition on a positive note: jumped up 389 positions on the private leaderboard. Even though the final rank is 2067 (top 42%), I'm pretty stoked about my progress so far. \n\nHere are a couple of things I've learned during this competition:\n- How to work with memory limitations\n- How to work with RAPIDS\n- Feature engineering ideas (my favorite trick was to combine max and min into a difference of max - min feature; it achieved the same score as having max and min as separate sets of features freeing up almost 300 data columns)\n- I'm not so good at feature selection: I tried permutation importance, SHAP, Boruta, adversarial validation, tree feature importance and I couldn't get to a point where I felt comfortable with any of the results\n- DART boosting - a slow but quite accurate boosting algorithm\n- Ensembling works well: two weak models outscored my best single model\n- Model validation - I managed to handpick the best model and the private score went up which is a positive note for me\n\nWhat I plan to learn from the top solutions once/if they become public:\n - Knowledge distillation: seemed to have played an important role in this competition\n - Become better at feature engineering:\n \t- check impact of dimensionality reduction techniques (umap/PCA etc.)\n \t- find features that describe the 13-period time series more accurately (wasn't able to find the sweet spot during the competition)\n \t- sky is the limit here, or maybe the RAM size :)\n - Feature selection with [RAPIDS FIL](https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing)\n\nThanks @cdeotte @raddar @ragnar123 @illidan7 @roberthatch for sharing your work and your findings!\n\nOn to the next one!",
    "1913723": "This was a great challenge and an invaluable learning experience for me too! Many thanks for the public datasets and discussion posts that helped me a lot, special mention to @cdeotte @raddar @ragnar123 @illidan7 @roberthatch @ambrosm @tilli7 @thedevastator for their work!",
    "1914076": "My first time too! I had many of the same learnings :)\n\nReally glad to hear my notebooks were of some benefit to you :)",
    "1922173": "Hi @canonicalized, thanks for sharing! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "2311555": "***Hi there! My name is Arif Patel and I am a developer from Dubai. I'm Also A Newbie Here***"
  },
  "source": "meta"
}