{
  "id": 344695,
  "title": "Learning: Key learning or take away from this challenge.",
  "url": "/competitions/amex-default-prediction/discussion/344695",
  "author_name": "Sarang Pratap Chamola",
  "post_date": "2022-08-16T08:24:57.403000",
  "votes": 22,
  "comment_count": 37,
  "views": 0,
  "content": "<p>As the competition is going to meet its deadline soon, Everyone might have learned something <br>\nfrom this challenge.<br>\nCan you Share any key learning or suggestions? </p>",
  "messages": [
    {
      "id": 1900752,
      "postDate": "2022-08-16T08:24:57.403Z",
      "content": "<p>As the competition is going to meet its deadline soon, Everyone might have learned something <br>\nfrom this challenge.<br>\nCan you Share any key learning or suggestions? </p>",
      "rawMarkdown": "As the competition is going to meet its deadline soon, Everyone might have learned something \nfrom this challenge.\nCan you Share any key learning or suggestions? \n",
      "votes": 21
    },
    {
      "id": 1901491,
      "postDate": "2022-08-16T17:22:52.713Z",
      "content": "<p>For us it was a competition under slogan \"nothing works\".</p>\n<p>All previous experience became useless. All the things that in theory should work - didn't - modeling / fe / features selection / validation. Only private lb will show how good we did our homework but I am expecting significant shakeup due to specific metric and public/private test splits and luck.</p>",
      "rawMarkdown": "For us it was a competition under slogan \"nothing works\".\n\nAll previous experience became useless. All the things that in theory should work - didn't - modeling / fe / features selection / validation. Only private lb will show how good we did our homework but I am expecting significant shakeup due to specific metric and public/private test splits and luck.",
      "votes": 13,
      "replies": [
        {
          "id": 1901559,
          "postDate": "2022-08-16T18:39:30.200Z",
          "content": "<p>Good luck to you!</p>",
          "rawMarkdown": "Good luck to you!",
          "votes": 2
        },
        {
          "id": 1901569,
          "postDate": "2022-08-16T18:48:54.690Z",
          "content": "<p>I feel this a lot but would also probably amend to \"almost nothing works\" :p If in a typical competition maybe 30-40% of ideas work, here it seems like 10-20% to me (at least on CV/public LB, we'll see about private). I can at least say that my apparent improvements have consistently been from ideas that make sense instead of random chance or complete shots in the dark. I.e., my CV/LB improves because I built new features or a new model, and not because I switched my seed. I guess there could be enough randomness here that I'm deceiving myself, but I'm cautiously optimistic.</p>",
          "rawMarkdown": "I feel this a lot but would also probably amend to \"almost nothing works\" :p If in a typical competition maybe 30-40% of ideas work, here it seems like 10-20% to me (at least on CV/public LB, we'll see about private). I can at least say that my apparent improvements have consistently been from ideas that make sense instead of random chance or complete shots in the dark. I.e., my CV/LB improves because I built new features or a new model, and not because I switched my seed. I guess there could be enough randomness here that I'm deceiving myself, but I'm cautiously optimistic.",
          "votes": 7
        },
        {
          "id": 1904089,
          "postDate": "2022-08-18T00:21:44.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> <br>\n<a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a> <br>\nCould I know your best local CV score?</p>",
          "rawMarkdown": "@aquatic \n@kyakovlev \nCould I know your best local CV score?",
          "votes": 2
        },
        {
          "id": 1904091,
          "postDate": "2022-08-18T00:25:01.890Z",
          "content": "<p>LGBM CV</p>\n<pre><code>####################\nAMEX: 0.8013056008869829\n#################### \nAUC: 0.963721124463359\n</code></pre>",
          "rawMarkdown": "LGBM CV\n```\n####################\nAMEX: 0.8013056008869829\n#################### \nAUC: 0.963721124463359\n```",
          "votes": 9
        },
        {
          "id": 1904513,
          "postDate": "2022-08-18T08:34:49.550Z",
          "content": "<p>Nice! Good luck!</p>",
          "rawMarkdown": "Nice! Good luck!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1901823,
      "postDate": "2022-08-17T00:28:46.577Z",
      "content": "<p>Hi guys!</p>\n<p>I'm new to data science and this was my second project after the titanic tutorial so naturally, I learned a lot from this competition.</p>\n<p>Key learnings and experiences were:</p>\n<ol>\n<li>Working with datasets that don't fit into RAM</li>\n<li>Basic ensembling methods (like averaging scores from multiple models) can yield better results than using ensembling models from sklearn)</li>\n<li>CatBoost and LGBM</li>\n<li>GridSearch for hyperparameter tuning</li>\n<li>There were quite a few columns with more than 50% missing data. I found that removing these gave worse results than keeping them which was a surprise to me.</li>\n<li>Dask and Vaex</li>\n</ol>\n<p>Things I need to learn:</p>\n<ol>\n<li>EDA</li>\n<li>Noise reduction/removal</li>\n</ol>\n<p>On the path to a competitions bronze medal (not in this competition though 😅)</p>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Hi guys!\n\nI'm new to data science and this was my second project after the titanic tutorial so naturally, I learned a lot from this competition.\n\nKey learnings and experiences were:\n1. Working with datasets that don't fit into RAM\n2. Basic ensembling methods (like averaging scores from multiple models) can yield better results than using ensembling models from sklearn)\n3. CatBoost and LGBM\n4. GridSearch for hyperparameter tuning\n5. There were quite a few columns with more than 50% missing data. I found that removing these gave worse results than keeping them which was a surprise to me.\n6. Dask and Vaex\n\nThings I need to learn:\n1. EDA\n2. Noise reduction/removal\n\nOn the path to a competitions bronze medal (not in this competition though 😅)\n\nHappy Kaggling!",
      "votes": 9,
      "replies": [
        {
          "id": 1907864,
          "postDate": "2022-08-21T06:42:09.447Z",
          "content": "<p>Happy learnings!!</p>",
          "rawMarkdown": "Happy learnings!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1904045,
      "postDate": "2022-08-17T22:43:45.580Z",
      "content": "<p>It is the first competition I really tried, what I most had learned was how to deal with data larger than memory, in my personal projects I never needed to optimize data types. Also, I became better in vectorization and what I think is the most important, had learned to try every possible ideia, even if sounds stupid at first.</p>",
      "rawMarkdown": "It is the first competition I really tried, what I most had learned was how to deal with data larger than memory, in my personal projects I never needed to optimize data types. Also, I became better in vectorization and what I think is the most important, had learned to try every possible ideia, even if sounds stupid at first.",
      "votes": 7
    },
    {
      "id": 1901292,
      "postDate": "2022-08-16T14:54:21.837Z",
      "content": "<p>I'd never worked with DART before and it was interesting to see it shine on this dataset / metric. Seems like it may be especially good for noisy metrics and/or cases where very strong features (<code>P_2</code>) can drown out weaker ones.</p>",
      "rawMarkdown": "I'd never worked with DART before and it was interesting to see it shine on this dataset / metric. Seems like it may be especially good for noisy metrics and/or cases where very strong features (`P_2`) can drown out weaker ones.",
      "votes": 7,
      "replies": [
        {
          "id": 1922364,
          "postDate": "2022-09-01T12:40:36.037Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "rawMarkdown": "Hello @aquatic. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
        }
      ]
    },
    {
      "id": 1905380,
      "postDate": "2022-08-19T03:08:39.040Z",
      "content": "<p>Feature engineering is hard to do with a large number of features and a lot of correlations between features, bold use of column sampling will yield a decent result😂</p>",
      "rawMarkdown": "Feature engineering is hard to do with a large number of features and a lot of correlations between features, bold use of column sampling will yield a decent result😂",
      "votes": 5,
      "replies": [
        {
          "id": 1907862,
          "postDate": "2022-08-21T06:41:45.927Z",
          "content": "<p>All the best with your results</p>",
          "rawMarkdown": "All the best with your results"
        }
      ]
    },
    {
      "id": 1901294,
      "postDate": "2022-08-16T14:56:05.860Z",
      "content": "<p>wait for the end of the competition… possibly huge shake-up to come. </p>",
      "rawMarkdown": "wait for the end of the competition... possibly huge shake-up to come. ",
      "votes": 5
    },
    {
      "id": 1905764,
      "postDate": "2022-08-19T09:24:05.447Z",
      "content": "<p>GridSearchCV and joblib for the parallel usage</p>",
      "rawMarkdown": "GridSearchCV and joblib for the parallel usage",
      "votes": 3,
      "replies": [
        {
          "id": 1922362,
          "postDate": "2022-09-01T12:40:00.197Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/abysswalker1994\" target=\"_blank\">@abysswalker1994</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "rawMarkdown": "Hi @abysswalker1994 May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
        }
      ]
    },
    {
      "id": 1901740,
      "postDate": "2022-08-16T21:19:20.370Z",
      "content": "<p>I learned that a proper Feature Engineering and Selection is the key for this competition. More will come after the deadline… :)</p>",
      "rawMarkdown": "I learned that a proper Feature Engineering and Selection is the key for this competition. More will come after the deadline... :)",
      "votes": 3
    },
    {
      "id": 1922357,
      "postDate": "2022-09-01T12:37:00.773Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sarang210\" target=\"_blank\">@sarang210</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @sarang210. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": 1,
      "replies": [
        {
          "id": 1925612,
          "postDate": "2022-09-04T06:28:13.607Z",
          "content": "<p>Yes, sure why not!</p>",
          "rawMarkdown": "Yes, sure why not!"
        }
      ]
    },
    {
      "id": 1908740,
      "postDate": "2022-08-22T01:07:51.973Z",
      "content": "<p>This is my very first kaggle competition, and I can say that my two key learning from this competition are: feature engineering and LightGBM. Feature engineering can be challenging as the number of features and missing values increase. I never used LightGBM before, and I was very impressed on its speed (way much faster than XGBoost). </p>",
      "rawMarkdown": "This is my very first kaggle competition, and I can say that my two key learning from this competition are: feature engineering and LightGBM. Feature engineering can be challenging as the number of features and missing values increase. I never used LightGBM before, and I was very impressed on its speed (way much faster than XGBoost). ",
      "votes": 1,
      "replies": [
        {
          "id": 1908848,
          "postDate": "2022-08-22T04:58:15.520Z",
          "content": "<p>Yeah, Feature engineering was the key to this Challenge. </p>",
          "rawMarkdown": "Yeah, Feature engineering was the key to this Challenge. "
        },
        {
          "id": 1922358,
          "postDate": "2022-09-01T12:37:31.903Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/oscarm524\" target=\"_blank\">@oscarm524</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "rawMarkdown": "Hi @oscarm524. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
        }
      ]
    },
    {
      "id": 1904491,
      "postDate": "2022-08-18T08:14:27.723Z",
      "content": "<p>I think I learned a lot from the data processing and feature engineering, the gods were very good with the data, I could learn a lot of new knowledge from it, and this competition also made me understand dart more, dart shines in this competition.<br>\nThe competition is coming to an end and I hope everyone will do well.</p>",
      "rawMarkdown": "I think I learned a lot from the data processing and feature engineering, the gods were very good with the data, I could learn a lot of new knowledge from it, and this competition also made me understand dart more, dart shines in this competition.\nThe competition is coming to an end and I hope everyone will do well.",
      "votes": 1,
      "replies": [
        {
          "id": 1922361,
          "postDate": "2022-09-01T12:39:26.503Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/chal1ce\" target=\"_blank\">@chal1ce</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "rawMarkdown": "Hello @chal1ce. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
        }
      ]
    },
    {
      "id": 1901429,
      "postDate": "2022-08-16T16:37:43.093Z",
      "content": "<p>Using sklearn properly for the first time. Working with a large dataset and GPUs (thanks Chris for the XGBoost Starter!). Understanding custom objectives, but giving up on them for this competition. Realising that you can compete with FE (the opposite of what I do in my day job).</p>",
      "rawMarkdown": "Using sklearn properly for the first time. Working with a large dataset and GPUs (thanks Chris for the XGBoost Starter!). Understanding custom objectives, but giving up on them for this competition. Realising that you can compete with FE (the opposite of what I do in my day job).",
      "votes": 1,
      "replies": [
        {
          "id": 1901562,
          "postDate": "2022-08-16T18:42:14.610Z",
          "content": "<p>Good luck with Final standings <a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a> </p>",
          "rawMarkdown": "Good luck with Final standings @burritodan "
        }
      ]
    },
    {
      "id": 1905332,
      "postDate": "2022-08-19T00:57:24.800Z",
      "content": "<p>First competition. Very surprised by my own models that a basic logistic regression performed better than other \"advanced\" ML algorithms. </p>",
      "rawMarkdown": "First competition. Very surprised by my own models that a basic logistic regression performed better than other \"advanced\" ML algorithms. ",
      "votes": 2
    },
    {
      "id": 1905219,
      "postDate": "2022-08-18T20:32:33.753Z",
      "content": "<p>learn many About feature eng, dart LGBM and deal with large dataset. but still hope some body show me, what is the keep point to get in the top 1000, even top 100,  without having a 64GB RAM fancy machine…   when the comp finished</p>",
      "rawMarkdown": "learn many About feature eng, dart LGBM and deal with large dataset. but still hope some body show me, what is the keep point to get in the top 1000, even top 100,  without having a 64GB RAM fancy machine...   when the comp finished",
      "votes": 2
    },
    {
      "id": 1903217,
      "postDate": "2022-08-17T07:43:19.130Z",
      "content": "<p>Thank U for your topic</p>",
      "rawMarkdown": "Thank U for your topic",
      "votes": 2,
      "replies": [
        {
          "id": 1903236,
          "postDate": "2022-08-17T08:06:40.860Z",
          "content": "<p>Your welcome ✌️ and all the best for competition.</p>",
          "rawMarkdown": "Your welcome ✌️ and all the best for competition."
        },
        {
          "id": 1922359,
          "postDate": "2022-09-01T12:38:55.763Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/whatchocolatesgoing\" target=\"_blank\">@whatchocolatesgoing</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "rawMarkdown": "Hi @whatchocolatesgoing. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
        }
      ]
    },
    {
      "id": 1901236,
      "postDate": "2022-08-16T14:27:26.390Z",
      "content": "<p>LGBM once again looks like the sweet spot model for almost every possible problem that uses tabular data.</p>",
      "rawMarkdown": "LGBM once again looks like the sweet spot model for almost every possible problem that uses tabular data.",
      "votes": 2,
      "replies": [
        {
          "id": 1901274,
          "postDate": "2022-08-16T14:42:39.843Z",
          "content": "<p>LGBM and XGBoost seems like Kaggle favourites 😁!</p>",
          "rawMarkdown": "LGBM and XGBoost seems like Kaggle favourites 😁!"
        }
      ]
    },
    {
      "id": 1901209,
      "postDate": "2022-08-16T14:06:34.223Z",
      "content": "<p>Working with a very large dataset, feather and parquet format, using models other than ML in such datasets (I work in a bank where we are not allowed to use NN in such cases) are some learnings for me. <br>\nAll the best for this competition!!</p>",
      "rawMarkdown": "Working with a very large dataset, feather and parquet format, using models other than ML in such datasets (I work in a bank where we are not allowed to use NN in such cases) are some learnings for me. \nAll the best for this competition!!",
      "votes": 2,
      "replies": [
        {
          "id": 1901271,
          "postDate": "2022-08-16T14:41:32.983Z",
          "content": "<p>Surely, It was a Great experience to work with such a large dataset and applying unorthodox algorithms for prediction.<br>\nAll the best Ravi!! </p>",
          "rawMarkdown": "Surely, It was a Great experience to work with such a large dataset and applying unorthodox algorithms for prediction.\nAll the best Ravi!! \n\n"
        }
      ]
    },
    {
      "id": 1900798,
      "postDate": "2022-08-16T09:03:23.007Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1900823,
          "postDate": "2022-08-16T09:25:32.937Z",
          "content": "<p>Sure, all the best</p>",
          "rawMarkdown": "Sure, all the best"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1901491,
      "author_name": "Konstantin Yakovlev",
      "author_url": "",
      "post_date": "2022-08-16T17:22:52.713000",
      "content": "<p>For us it was a competition under slogan \"nothing works\".</p>\n<p>All previous experience became useless. All the things that in theory should work - didn't - modeling / fe / features selection / validation. Only private lb will show how good we did our homework but I am expecting significant shakeup due to specific metric and public/private test splits and luck.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 1901559,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-16T18:39:30.200000",
          "content": "<p>Good luck to you!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1901569,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2022-08-16T18:48:54.690000",
          "content": "<p>I feel this a lot but would also probably amend to \"almost nothing works\" :p If in a typical competition maybe 30-40% of ideas work, here it seems like 10-20% to me (at least on CV/public LB, we'll see about private). I can at least say that my apparent improvements have consistently been from ideas that make sense instead of random chance or complete shots in the dark. I.e., my CV/LB improves because I built new features or a new model, and not because I switched my seed. I guess there could be enough randomness here that I'm deceiving myself, but I'm cautiously optimistic.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1904089,
          "author_name": "Mohamed Eltayeb",
          "author_url": "",
          "post_date": "2022-08-18T00:21:44.230000",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> <br>\n<a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a> <br>\nCould I know your best local CV score?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1904091,
          "author_name": "Konstantin Yakovlev",
          "author_url": "",
          "post_date": "2022-08-18T00:25:01.890000",
          "content": "<p>LGBM CV</p>\n<pre><code>####################\nAMEX: 0.8013056008869829\n#################### \nAUC: 0.963721124463359\n</code></pre>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1904513,
          "author_name": "Mohamed Eltayeb",
          "author_url": "",
          "post_date": "2022-08-18T08:34:49.550000",
          "content": "<p>Nice! Good luck!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1901823,
      "author_name": "shahilpravind",
      "author_url": "",
      "post_date": "2022-08-17T00:28:46.577000",
      "content": "<p>Hi guys!</p>\n<p>I'm new to data science and this was my second project after the titanic tutorial so naturally, I learned a lot from this competition.</p>\n<p>Key learnings and experiences were:</p>\n<ol>\n<li>Working with datasets that don't fit into RAM</li>\n<li>Basic ensembling methods (like averaging scores from multiple models) can yield better results than using ensembling models from sklearn)</li>\n<li>CatBoost and LGBM</li>\n<li>GridSearch for hyperparameter tuning</li>\n<li>There were quite a few columns with more than 50% missing data. I found that removing these gave worse results than keeping them which was a surprise to me.</li>\n<li>Dask and Vaex</li>\n</ol>\n<p>Things I need to learn:</p>\n<ol>\n<li>EDA</li>\n<li>Noise reduction/removal</li>\n</ol>\n<p>On the path to a competitions bronze medal (not in this competition though 😅)</p>\n<p>Happy Kaggling!</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1907864,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-21T06:42:09.447000",
          "content": "<p>Happy learnings!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1904045,
      "author_name": "AG1998",
      "author_url": "",
      "post_date": "2022-08-17T22:43:45.580000",
      "content": "<p>It is the first competition I really tried, what I most had learned was how to deal with data larger than memory, in my personal projects I never needed to optimize data types. Also, I became better in vectorization and what I think is the most important, had learned to try every possible ideia, even if sounds stupid at first.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1901292,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2022-08-16T14:54:21.837000",
      "content": "<p>I'd never worked with DART before and it was interesting to see it shine on this dataset / metric. Seems like it may be especially good for noisy metrics and/or cases where very strong features (<code>P_2</code>) can drown out weaker ones.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1922364,
          "author_name": "Yang Liu",
          "author_url": "",
          "post_date": "2022-09-01T12:40:36.037000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1905380,
      "author_name": "Fading_Vay",
      "author_url": "",
      "post_date": "2022-08-19T03:08:39.040000",
      "content": "<p>Feature engineering is hard to do with a large number of features and a lot of correlations between features, bold use of column sampling will yield a decent result😂</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1907862,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-21T06:41:45.927000",
          "content": "<p>All the best with your results</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1901294,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-08-16T14:56:05.860000",
      "content": "<p>wait for the end of the competition… possibly huge shake-up to come. </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1905764,
      "author_name": "Abysswalker1994",
      "author_url": "",
      "post_date": "2022-08-19T09:24:05.447000",
      "content": "<p>GridSearchCV and joblib for the parallel usage</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1922362,
          "author_name": "Yang Liu",
          "author_url": "",
          "post_date": "2022-09-01T12:40:00.197000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/abysswalker1994\" target=\"_blank\">@abysswalker1994</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1901740,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2022-08-16T21:19:20.370000",
      "content": "<p>I learned that a proper Feature Engineering and Selection is the key for this competition. More will come after the deadline… :)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1922357,
      "author_name": "Yang Liu",
      "author_url": "",
      "post_date": "2022-09-01T12:37:00.773000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sarang210\" target=\"_blank\">@sarang210</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1925612,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-09-04T06:28:13.607000",
          "content": "<p>Yes, sure why not!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1908740,
      "author_name": "Oscar Aguilar",
      "author_url": "",
      "post_date": "2022-08-22T01:07:51.973000",
      "content": "<p>This is my very first kaggle competition, and I can say that my two key learning from this competition are: feature engineering and LightGBM. Feature engineering can be challenging as the number of features and missing values increase. I never used LightGBM before, and I was very impressed on its speed (way much faster than XGBoost). </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1908848,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-22T04:58:15.520000",
          "content": "<p>Yeah, Feature engineering was the key to this Challenge. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1922358,
          "author_name": "Yang Liu",
          "author_url": "",
          "post_date": "2022-09-01T12:37:31.903000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/oscarm524\" target=\"_blank\">@oscarm524</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1904491,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-18T08:14:27.723000",
      "content": "<p>I think I learned a lot from the data processing and feature engineering, the gods were very good with the data, I could learn a lot of new knowledge from it, and this competition also made me understand dart more, dart shines in this competition.<br>\nThe competition is coming to an end and I hope everyone will do well.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1922361,
          "author_name": "Yang Liu",
          "author_url": "",
          "post_date": "2022-09-01T12:39:26.503000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/chal1ce\" target=\"_blank\">@chal1ce</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1901429,
      "author_name": "Burrito Dan",
      "author_url": "",
      "post_date": "2022-08-16T16:37:43.093000",
      "content": "<p>Using sklearn properly for the first time. Working with a large dataset and GPUs (thanks Chris for the XGBoost Starter!). Understanding custom objectives, but giving up on them for this competition. Realising that you can compete with FE (the opposite of what I do in my day job).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1901562,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-16T18:42:14.610000",
          "content": "<p>Good luck with Final standings <a href=\"https://www.kaggle.com/burritodan\" target=\"_blank\">@burritodan</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1905332,
      "author_name": "Nate Shipley",
      "author_url": "",
      "post_date": "2022-08-19T00:57:24.800000",
      "content": "<p>First competition. Very surprised by my own models that a basic logistic regression performed better than other \"advanced\" ML algorithms. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1905219,
      "author_name": "Taylor_1224",
      "author_url": "",
      "post_date": "2022-08-18T20:32:33.753000",
      "content": "<p>learn many About feature eng, dart LGBM and deal with large dataset. but still hope some body show me, what is the keep point to get in the top 1000, even top 100,  without having a 64GB RAM fancy machine…   when the comp finished</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1903217,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-17T07:43:19.130000",
      "content": "<p>Thank U for your topic</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1903236,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-17T08:06:40.860000",
          "content": "<p>Your welcome ✌️ and all the best for competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1922359,
          "author_name": "Yang Liu",
          "author_url": "",
          "post_date": "2022-09-01T12:38:55.763000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/whatchocolatesgoing\" target=\"_blank\">@whatchocolatesgoing</a>. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1901236,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2022-08-16T14:27:26.390000",
      "content": "<p>LGBM once again looks like the sweet spot model for almost every possible problem that uses tabular data.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1901274,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-16T14:42:39.843000",
          "content": "<p>LGBM and XGBoost seems like Kaggle favourites 😁!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1901209,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-08-16T14:06:34.223000",
      "content": "<p>Working with a very large dataset, feather and parquet format, using models other than ML in such datasets (I work in a bank where we are not allowed to use NN in such cases) are some learnings for me. <br>\nAll the best for this competition!!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1901271,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-16T14:41:32.983000",
          "content": "<p>Surely, It was a Great experience to work with such a large dataset and applying unorthodox algorithms for prediction.<br>\nAll the best Ravi!! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1900798,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-16T09:03:23.007000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1900823,
          "author_name": "Sarang Pratap Chamola",
          "author_url": "",
          "post_date": "2022-08-16T09:25:32.937000",
          "content": "<p>Sure, all the best</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1900752": "As the competition is going to meet its deadline soon, Everyone might have learned something \nfrom this challenge.\nCan you Share any key learning or suggestions? \n",
    "1901491": "For us it was a competition under slogan \"nothing works\".\n\nAll previous experience became useless. All the things that in theory should work - didn't - modeling / fe / features selection / validation. Only private lb will show how good we did our homework but I am expecting significant shakeup due to specific metric and public/private test splits and luck.",
    "1901823": "Hi guys!\n\nI'm new to data science and this was my second project after the titanic tutorial so naturally, I learned a lot from this competition.\n\nKey learnings and experiences were:\n1. Working with datasets that don't fit into RAM\n2. Basic ensembling methods (like averaging scores from multiple models) can yield better results than using ensembling models from sklearn)\n3. CatBoost and LGBM\n4. GridSearch for hyperparameter tuning\n5. There were quite a few columns with more than 50% missing data. I found that removing these gave worse results than keeping them which was a surprise to me.\n6. Dask and Vaex\n\nThings I need to learn:\n1. EDA\n2. Noise reduction/removal\n\nOn the path to a competitions bronze medal (not in this competition though 😅)\n\nHappy Kaggling!",
    "1904045": "It is the first competition I really tried, what I most had learned was how to deal with data larger than memory, in my personal projects I never needed to optimize data types. Also, I became better in vectorization and what I think is the most important, had learned to try every possible ideia, even if sounds stupid at first.",
    "1901292": "I'd never worked with DART before and it was interesting to see it shine on this dataset / metric. Seems like it may be especially good for noisy metrics and/or cases where very strong features (`P_2`) can drown out weaker ones.",
    "1905380": "Feature engineering is hard to do with a large number of features and a lot of correlations between features, bold use of column sampling will yield a decent result😂",
    "1901294": "wait for the end of the competition... possibly huge shake-up to come. ",
    "1905764": "GridSearchCV and joblib for the parallel usage",
    "1901740": "I learned that a proper Feature Engineering and Selection is the key for this competition. More will come after the deadline... :)",
    "1922357": "Hi @sarang210. May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1908740": "This is my very first kaggle competition, and I can say that my two key learning from this competition are: feature engineering and LightGBM. Feature engineering can be challenging as the number of features and missing values increase. I never used LightGBM before, and I was very impressed on its speed (way much faster than XGBoost). ",
    "1904491": "I think I learned a lot from the data processing and feature engineering, the gods were very good with the data, I could learn a lot of new knowledge from it, and this competition also made me understand dart more, dart shines in this competition.\nThe competition is coming to an end and I hope everyone will do well.",
    "1901429": "Using sklearn properly for the first time. Working with a large dataset and GPUs (thanks Chris for the XGBoost Starter!). Understanding custom objectives, but giving up on them for this competition. Realising that you can compete with FE (the opposite of what I do in my day job).",
    "1905332": "First competition. Very surprised by my own models that a basic logistic regression performed better than other \"advanced\" ML algorithms. ",
    "1905219": "learn many About feature eng, dart LGBM and deal with large dataset. but still hope some body show me, what is the keep point to get in the top 1000, even top 100,  without having a 64GB RAM fancy machine...   when the comp finished",
    "1903217": "Thank U for your topic",
    "1901236": "LGBM once again looks like the sweet spot model for almost every possible problem that uses tabular data.",
    "1901209": "Working with a very large dataset, feather and parquet format, using models other than ML in such datasets (I work in a bank where we are not allowed to use NN in such cases) are some learnings for me. \nAll the best for this competition!!",
    "1900798": ""
  }
}