{
  "id": 348352,
  "title": "Private Leaderboard 3442 - A Beginner's Simple Solution and Key Takeaways",
  "url": "/competitions/amex-default-prediction/writeups/shahilpravind-private-leaderboard-3442-a-beginner-",
  "author_name": "",
  "post_date": "2022-08-28T01:26:24.247217100Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n<p>Don't think this would be useful to anyone with experience on Kaggle but I thought I would share my rather basic solution. As this was my first actual project, I mostly ended up trying out all the things I had learned from tutorials in a real project rather than making a significant model. It proved a good experience and I look forward to becoming better at DS using the next steps for education identified through this competition.</p>\n<p>My best performing solution's notebook is available <a href=\"https://www.kaggle.com/code/shahilap96/private-leaderboard-3442-my-best-solution\" target=\"_blank\">here</a> for anyone that wishes to take a look.</p>\n<p>Final model:</p>\n<ol>\n<li>Denoised <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">dataset</a> by <strong><em>raddar</em></strong>.</li>\n<li>Last two transactions per customer as suggested by this <a href=\"https://www.kaggle.com/code/junjitakeshima/amex-try-to-improve-lgbm-starter-eng\" target=\"_blank\">notebook</a>. I tried all variations and found that two worked best for me.</li>\n<li>Dropped Customer_ID and S_2 columns. I found that keeping all other columns (including those with high number of nulls gave better results.</li>\n<li>scikit-learn's imputation and standard scaling.</li>\n<li>Two models -&gt; LGBM (n_estimators=300) and CatBoost (default parameters).</li>\n<li>Averaged my LGBM and CatBoost prediction probabilities for final submission. This gave the best solution.</li>\n</ol>\n<p>Things I tried:</p>\n<ol>\n<li>XGBoost - didn't perform as well as LGBM and CatBoost alone and when averaged with their scores.</li>\n<li>Simple NN - last minute trial, didn't perform well, probably due to lack of tuning.</li>\n<li>Random Forest, Decision Tree and other similar ML models.</li>\n<li>Removed columns with high number of nulls (50% null, 66% null, 75% null).</li>\n<li>Cross validation for measuring performance.</li>\n<li>Dask and Vaex for working with the large volume of data initially.</li>\n<li>Hyperparameter tuning using GridSearchCV.</li>\n</ol>\n<p>Things I need to look at next:</p>\n<ol>\n<li>Denoising data</li>\n<li>EDA</li>\n<li>Feature engineering</li>\n<li>Understanding higher ranking participants solutions</li>\n<li>Handling large volume of data when training models.</li>\n</ol>\n<p>I definitely have a long way to go but this competition was a great starting point and the community around it made it an awesome experience.</p>\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": "1916542",
      "postDate": "08/28/2022 01:26:24",
      "content": "<p>Hi everyone!</p>\n<p>Don't think this would be useful to anyone with experience on Kaggle but I thought I would share my rather basic solution. As this was my first actual project, I mostly ended up trying out all the things I had learned from tutorials in a real project rather than making a significant model. It proved a good experience and I look forward to becoming better at DS using the next steps for education identified through this competition.</p>\n<p>My best performing solution's notebook is available <a href=\"https://www.kaggle.com/code/shahilap96/private-leaderboard-3442-my-best-solution\" target=\"_blank\">here</a> for anyone that wishes to take a look.</p>\n<p>Final model:</p>\n<ol>\n<li>Denoised <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">dataset</a> by <strong><em>raddar</em></strong>.</li>\n<li>Last two transactions per customer as suggested by this <a href=\"https://www.kaggle.com/code/junjitakeshima/amex-try-to-improve-lgbm-starter-eng\" target=\"_blank\">notebook</a>. I tried all variations and found that two worked best for me.</li>\n<li>Dropped Customer_ID and S_2 columns. I found that keeping all other columns (including those with high number of nulls gave better results.</li>\n<li>scikit-learn's imputation and standard scaling.</li>\n<li>Two models -&gt; LGBM (n_estimators=300) and CatBoost (default parameters).</li>\n<li>Averaged my LGBM and CatBoost prediction probabilities for final submission. This gave the best solution.</li>\n</ol>\n<p>Things I tried:</p>\n<ol>\n<li>XGBoost - didn't perform as well as LGBM and CatBoost alone and when averaged with their scores.</li>\n<li>Simple NN - last minute trial, didn't perform well, probably due to lack of tuning.</li>\n<li>Random Forest, Decision Tree and other similar ML models.</li>\n<li>Removed columns with high number of nulls (50% null, 66% null, 75% null).</li>\n<li>Cross validation for measuring performance.</li>\n<li>Dask and Vaex for working with the large volume of data initially.</li>\n<li>Hyperparameter tuning using GridSearchCV.</li>\n</ol>\n<p>Things I need to look at next:</p>\n<ol>\n<li>Denoising data</li>\n<li>EDA</li>\n<li>Feature engineering</li>\n<li>Understanding higher ranking participants solutions</li>\n<li>Handling large volume of data when training models.</li>\n</ol>\n<p>I definitely have a long way to go but this competition was a great starting point and the community around it made it an awesome experience.</p>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Hi everyone!\n\nDon't think this would be useful to anyone with experience on Kaggle but I thought I would share my rather basic solution. As this was my first actual project, I mostly ended up trying out all the things I had learned from tutorials in a real project rather than making a significant model. It proved a good experience and I look forward to becoming better at DS using the next steps for education identified through this competition.\n\nMy best performing solution's notebook is available [here](https://www.kaggle.com/code/shahilap96/private-leaderboard-3442-my-best-solution) for anyone that wishes to take a look.\n\nFinal model:\n1. Denoised [dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) by ***raddar***.\n2. Last two transactions per customer as suggested by this [notebook](https://www.kaggle.com/code/junjitakeshima/amex-try-to-improve-lgbm-starter-eng). I tried all variations and found that two worked best for me.\n3. Dropped Customer_ID and S_2 columns. I found that keeping all other columns (including those with high number of nulls gave better results.\n4. scikit-learn's imputation and standard scaling.\n5. Two models -> LGBM (n_estimators=300) and CatBoost (default parameters).\n6. Averaged my LGBM and CatBoost prediction probabilities for final submission. This gave the best solution.\n\nThings I tried:\n1. XGBoost - didn't perform as well as LGBM and CatBoost alone and when averaged with their scores.\n2. Simple NN - last minute trial, didn't perform well, probably due to lack of tuning.\n3. Random Forest, Decision Tree and other similar ML models.\n4. Removed columns with high number of nulls (50% null, 66% null, 75% null).\n5. Cross validation for measuring performance.\n6. Dask and Vaex for working with the large volume of data initially.\n7. Hyperparameter tuning using GridSearchCV.\n\nThings I need to look at next:\n1. Denoising data\n2. EDA\n3. Feature engineering\n4. Understanding higher ranking participants solutions\n5. Handling large volume of data when training models.\n\nI definitely have a long way to go but this competition was a great starting point and the community around it made it an awesome experience.\n\nHappy Kaggling!",
      "votes": null
    },
    {
      "id": "1920239",
      "postDate": "08/31/2022 02:49:40",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/shahilap96\" target=\"_blank\">@shahilap96</a>. Let's grow together. I dont know if you performed some ensembling but, from the higher ranking participant solutions, it seemed to be vital. Wish you the best  </p>",
      "rawMarkdown": "Thanks for sharing @shahilap96. Let's grow together. I dont know if you performed some ensembling but, from the higher ranking participant solutions, it seemed to be vital. Wish you the best",
      "votes": null
    },
    {
      "id": "1920323",
      "postDate": "08/31/2022 04:41:02",
      "content": "<p>Thanks Felipe. I didn't do any complex ensembles like the high performing participants however, I did try stacking, max, min and average of my CatBoost and LGBM models. My final solution was average of prediction probabilities produced by the top CatBoost and LGBM models I made.</p>\n<p>All the best to you as well. Cheers!</p>",
      "rawMarkdown": "Thanks Felipe. I didn't do any complex ensembles like the high performing participants however, I did try stacking, max, min and average of my CatBoost and LGBM models. My final solution was average of prediction probabilities produced by the top CatBoost and LGBM models I made.\n\nAll the best to you as well. Cheers!",
      "votes": null
    },
    {
      "id": "1920876",
      "postDate": "08/31/2022 13:03:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shahilap96\" target=\"_blank\">@shahilap96</a>, thanks for sharing this! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @shahilap96, thanks for sharing this! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    },
    {
      "id": "1921486",
      "postDate": "08/31/2022 21:44:30",
      "content": "<p>\"This is not a scam\" is very suspicious however, I did end up doing the survey after taking a look at your discussion post regarding the survey.</p>\n<p>All the best with your study. Cheers!</p>",
      "rawMarkdown": "\"This is not a scam\" is very suspicious however, I did end up doing the survey after taking a look at your discussion post regarding the survey.\n\nAll the best with your study. Cheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1920239,
      "author_name": "clacflores",
      "author_url": "",
      "post_date": "08/31/2022 02:49:40",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/shahilap96\" target=\"_blank\">@shahilap96</a>. Let's grow together. I dont know if you performed some ensembling but, from the higher ranking participant solutions, it seemed to be vital. Wish you the best  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1920323,
          "author_name": "shahilap96",
          "author_url": "",
          "post_date": "08/31/2022 04:41:02",
          "content": "<p>Thanks Felipe. I didn't do any complex ensembles like the high performing participants however, I did try stacking, max, min and average of my CatBoost and LGBM models. My final solution was average of prediction probabilities produced by the top CatBoost and LGBM models I made.</p>\n<p>All the best to you as well. Cheers!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1920876,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "08/31/2022 13:03:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shahilap96\" target=\"_blank\">@shahilap96</a>, thanks for sharing this! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1921486,
          "author_name": "shahilap96",
          "author_url": "",
          "post_date": "08/31/2022 21:44:30",
          "content": "<p>\"This is not a scam\" is very suspicious however, I did end up doing the survey after taking a look at your discussion post regarding the survey.</p>\n<p>All the best with your study. Cheers!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1916542": "Hi everyone!\n\nDon't think this would be useful to anyone with experience on Kaggle but I thought I would share my rather basic solution. As this was my first actual project, I mostly ended up trying out all the things I had learned from tutorials in a real project rather than making a significant model. It proved a good experience and I look forward to becoming better at DS using the next steps for education identified through this competition.\n\nMy best performing solution's notebook is available [here](https://www.kaggle.com/code/shahilap96/private-leaderboard-3442-my-best-solution) for anyone that wishes to take a look.\n\nFinal model:\n1. Denoised [dataset](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) by ***raddar***.\n2. Last two transactions per customer as suggested by this [notebook](https://www.kaggle.com/code/junjitakeshima/amex-try-to-improve-lgbm-starter-eng). I tried all variations and found that two worked best for me.\n3. Dropped Customer_ID and S_2 columns. I found that keeping all other columns (including those with high number of nulls gave better results.\n4. scikit-learn's imputation and standard scaling.\n5. Two models -> LGBM (n_estimators=300) and CatBoost (default parameters).\n6. Averaged my LGBM and CatBoost prediction probabilities for final submission. This gave the best solution.\n\nThings I tried:\n1. XGBoost - didn't perform as well as LGBM and CatBoost alone and when averaged with their scores.\n2. Simple NN - last minute trial, didn't perform well, probably due to lack of tuning.\n3. Random Forest, Decision Tree and other similar ML models.\n4. Removed columns with high number of nulls (50% null, 66% null, 75% null).\n5. Cross validation for measuring performance.\n6. Dask and Vaex for working with the large volume of data initially.\n7. Hyperparameter tuning using GridSearchCV.\n\nThings I need to look at next:\n1. Denoising data\n2. EDA\n3. Feature engineering\n4. Understanding higher ranking participants solutions\n5. Handling large volume of data when training models.\n\nI definitely have a long way to go but this competition was a great starting point and the community around it made it an awesome experience.\n\nHappy Kaggling!",
    "1920239": "Thanks for sharing @shahilap96. Let's grow together. I dont know if you performed some ensembling but, from the higher ranking participant solutions, it seemed to be vital. Wish you the best",
    "1920323": "Thanks Felipe. I didn't do any complex ensembles like the high performing participants however, I did try stacking, max, min and average of my CatBoost and LGBM models. My final solution was average of prediction probabilities produced by the top CatBoost and LGBM models I made.\n\nAll the best to you as well. Cheers!",
    "1920876": "Hi @shahilap96, thanks for sharing this! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1921486": "\"This is not a scam\" is very suspicious however, I did end up doing the survey after taking a look at your discussion post regarding the survey.\n\nAll the best with your study. Cheers!"
  },
  "source": "meta"
}