{
  "id": 348035,
  "title": "when reality hits!!",
  "url": "/competitions/amex-default-prediction/discussion/348035",
  "author_name": "",
  "post_date": "2022-08-26T13:59:59.083115700Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello 👋to all kagglers, and congratulations 🎉 to all winners 💪.</p>\n<p>It's my first Kaggle competition, and unfortunately didn't do very well. I recently completed my MS (Data Science) and decided to give Kaggle a go. I did very well during my MS projects and my initial thoughts were, \"How hard could it be,\" however when I started doing it and struggled to improve my scores, I realized, \"dude, wake up and face reality. 😟</p>\n<p>Here are my reflections (for newbies)…</p>\n<p><strong>Positives</strong>:</p>\n<ol>\n<li>The discussion threads and shared code by fellow Kagglers are the best part; The tips and techniques shared by generous grandmasters/masters are a real eye-opener. Something that one can never learn in the classroom or by reading books.</li>\n<li>In the Private leader board, my score improved, and I jumped 19 places up, which tells me that at least my model of not over-fitting. I did something right!</li>\n</ol>\n<p><strong>Areas to focus/improve</strong>:</p>\n<ol>\n<li>EDA and Feature engineering: This has been emphasized by ML gurus and my professor enough in the past, but as usual, newbies tend to focus on seemingly \"sexy\" stuff, i.e., models/algorithms.  </li>\n<li>Ensembling: Almost all winning solutions seems to use some model ensemble for final prediction. This is something not taught in formal ML/DS courses (at least I was not introduced to it). I need to research and learn the techniques before my next Kaggle competition.</li>\n<li>Pipeline: The ability to perform many quick experiments is the key. One needs to have a robust data-processing pipeline that can be used across projects with minor modifications. I had none; as part of this competition, I managed to put together a basic pipeline. </li>\n<li>Do not change path drastically: Enormous amount of ideas are shared and discussed in the competition, and it's tempting to jump from one to another, leaving your own plan aside. I did this a lot, got even more confused, and did not help. It's best to have a plan of action and stick with it until it's fully explored.</li>\n<li>Start early: It's better to start as early as possible and spend as much time as possible on data exploration, especially if you are new. Realized that making sense of masked data is not easy</li>\n<li>Make a team: A lone wolf (especially newbies) won't do, competitions such as this are way too overwhelming. Divide and rule would work better. </li>\n</ol>\n<p>Special thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for his curated dataset. I started with the raw csv dataset, it was a real pain and frustrating hardly could reach 0.6 score and I was on verge of giving up. Came across raddar's dataset and decided to give it a go and my scores jumped in range of 0.7. I was so excited that I called in sick next Monday 😜 and continued working. </p>\n<p>I hope to report a better score in my next competition, till then happy learning 👍</p>",
  "messages": [
    {
      "id": "1914881",
      "postDate": "08/26/2022 13:59:59",
      "content": "<p>Hello 👋to all kagglers, and congratulations 🎉 to all winners 💪.</p>\n<p>It's my first Kaggle competition, and unfortunately didn't do very well. I recently completed my MS (Data Science) and decided to give Kaggle a go. I did very well during my MS projects and my initial thoughts were, \"How hard could it be,\" however when I started doing it and struggled to improve my scores, I realized, \"dude, wake up and face reality. 😟</p>\n<p>Here are my reflections (for newbies)…</p>\n<p><strong>Positives</strong>:</p>\n<ol>\n<li>The discussion threads and shared code by fellow Kagglers are the best part; The tips and techniques shared by generous grandmasters/masters are a real eye-opener. Something that one can never learn in the classroom or by reading books.</li>\n<li>In the Private leader board, my score improved, and I jumped 19 places up, which tells me that at least my model of not over-fitting. I did something right!</li>\n</ol>\n<p><strong>Areas to focus/improve</strong>:</p>\n<ol>\n<li>EDA and Feature engineering: This has been emphasized by ML gurus and my professor enough in the past, but as usual, newbies tend to focus on seemingly \"sexy\" stuff, i.e., models/algorithms.  </li>\n<li>Ensembling: Almost all winning solutions seems to use some model ensemble for final prediction. This is something not taught in formal ML/DS courses (at least I was not introduced to it). I need to research and learn the techniques before my next Kaggle competition.</li>\n<li>Pipeline: The ability to perform many quick experiments is the key. One needs to have a robust data-processing pipeline that can be used across projects with minor modifications. I had none; as part of this competition, I managed to put together a basic pipeline. </li>\n<li>Do not change path drastically: Enormous amount of ideas are shared and discussed in the competition, and it's tempting to jump from one to another, leaving your own plan aside. I did this a lot, got even more confused, and did not help. It's best to have a plan of action and stick with it until it's fully explored.</li>\n<li>Start early: It's better to start as early as possible and spend as much time as possible on data exploration, especially if you are new. Realized that making sense of masked data is not easy</li>\n<li>Make a team: A lone wolf (especially newbies) won't do, competitions such as this are way too overwhelming. Divide and rule would work better. </li>\n</ol>\n<p>Special thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for his curated dataset. I started with the raw csv dataset, it was a real pain and frustrating hardly could reach 0.6 score and I was on verge of giving up. Came across raddar's dataset and decided to give it a go and my scores jumped in range of 0.7. I was so excited that I called in sick next Monday 😜 and continued working. </p>\n<p>I hope to report a better score in my next competition, till then happy learning 👍</p>",
      "rawMarkdown": "Hello 👋to all kagglers, and congratulations 🎉 to all winners 💪.\n\nIt's my first Kaggle competition, and unfortunately didn't do very well. I recently completed my MS (Data Science) and decided to give Kaggle a go. I did very well during my MS projects and my initial thoughts were, \"How hard could it be,\" however when I started doing it and struggled to improve my scores, I realized, \"dude, wake up and face reality. 😟\n\nHere are my reflections (for newbies)...\n\n**Positives**:\n1. The discussion threads and shared code by fellow Kagglers are the best part; The tips and techniques shared by generous grandmasters/masters are a real eye-opener. Something that one can never learn in the classroom or by reading books.\n2. In the Private leader board, my score improved, and I jumped 19 places up, which tells me that at least my model of not over-fitting. I did something right!\n\n**Areas to focus/improve**:\n1. EDA and Feature engineering: This has been emphasized by ML gurus and my professor enough in the past, but as usual, newbies tend to focus on seemingly \"sexy\" stuff, i.e., models/algorithms.  \n2. Ensembling: Almost all winning solutions seems to use some model ensemble for final prediction. This is something not taught in formal ML/DS courses (at least I was not introduced to it). I need to research and learn the techniques before my next Kaggle competition.\n3. Pipeline: The ability to perform many quick experiments is the key. One needs to have a robust data-processing pipeline that can be used across projects with minor modifications. I had none; as part of this competition, I managed to put together a basic pipeline. \n4. Do not change path drastically: Enormous amount of ideas are shared and discussed in the competition, and it's tempting to jump from one to another, leaving your own plan aside. I did this a lot, got even more confused, and did not help. It's best to have a plan of action and stick with it until it's fully explored.\n5. Start early: It's better to start as early as possible and spend as much time as possible on data exploration, especially if you are new. Realized that making sense of masked data is not easy\n6. Make a team: A lone wolf (especially newbies) won't do, competitions such as this are way too overwhelming. Divide and rule would work better. \n\nSpecial thanks to @raddar for his curated dataset. I started with the raw csv dataset, it was a real pain and frustrating hardly could reach 0.6 score and I was on verge of giving up. Came across raddar's dataset and decided to give it a go and my scores jumped in range of 0.7. I was so excited that I called in sick next Monday 😜 and continued working. \n\nI hope to report a better score in my next competition, till then happy learning 👍",
      "votes": null
    },
    {
      "id": "1915026",
      "postDate": "08/26/2022 16:03:13",
      "content": "<p>The more you compete, the more you'll realize how much of an art feature engineering is. Online courses barely even scratch the surface, the best place to kit up for your next competition would be the discussion forums once a contest ends. Happy kaggling!</p>",
      "rawMarkdown": "The more you compete, the more you'll realize how much of an art feature engineering is. Online courses barely even scratch the surface, the best place to kit up for your next competition would be the discussion forums once a contest ends. Happy kaggling!",
      "votes": null
    },
    {
      "id": "1915160",
      "postDate": "08/26/2022 18:02:43",
      "content": "<p>I totally agree with you, I am also planning to work on some top approaches and try to improve my score with this dataset. Happy learning and enjoy the ride <a href=\"https://www.kaggle.com/ol0fmeister\" target=\"_blank\">@ol0fmeister</a> </p>",
      "rawMarkdown": "I totally agree with you, I am also planning to work on some top approaches and try to improve my score with this dataset. Happy learning and enjoy the ride @ol0fmeister",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1915026,
      "author_name": "ol0fmeister",
      "author_url": "",
      "post_date": "08/26/2022 16:03:13",
      "content": "<p>The more you compete, the more you'll realize how much of an art feature engineering is. Online courses barely even scratch the surface, the best place to kit up for your next competition would be the discussion forums once a contest ends. Happy kaggling!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1915160,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "08/26/2022 18:02:43",
          "content": "<p>I totally agree with you, I am also planning to work on some top approaches and try to improve my score with this dataset. Happy learning and enjoy the ride <a href=\"https://www.kaggle.com/ol0fmeister\" target=\"_blank\">@ol0fmeister</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1914881": "Hello 👋to all kagglers, and congratulations 🎉 to all winners 💪.\n\nIt's my first Kaggle competition, and unfortunately didn't do very well. I recently completed my MS (Data Science) and decided to give Kaggle a go. I did very well during my MS projects and my initial thoughts were, \"How hard could it be,\" however when I started doing it and struggled to improve my scores, I realized, \"dude, wake up and face reality. 😟\n\nHere are my reflections (for newbies)...\n\n**Positives**:\n1. The discussion threads and shared code by fellow Kagglers are the best part; The tips and techniques shared by generous grandmasters/masters are a real eye-opener. Something that one can never learn in the classroom or by reading books.\n2. In the Private leader board, my score improved, and I jumped 19 places up, which tells me that at least my model of not over-fitting. I did something right!\n\n**Areas to focus/improve**:\n1. EDA and Feature engineering: This has been emphasized by ML gurus and my professor enough in the past, but as usual, newbies tend to focus on seemingly \"sexy\" stuff, i.e., models/algorithms.  \n2. Ensembling: Almost all winning solutions seems to use some model ensemble for final prediction. This is something not taught in formal ML/DS courses (at least I was not introduced to it). I need to research and learn the techniques before my next Kaggle competition.\n3. Pipeline: The ability to perform many quick experiments is the key. One needs to have a robust data-processing pipeline that can be used across projects with minor modifications. I had none; as part of this competition, I managed to put together a basic pipeline. \n4. Do not change path drastically: Enormous amount of ideas are shared and discussed in the competition, and it's tempting to jump from one to another, leaving your own plan aside. I did this a lot, got even more confused, and did not help. It's best to have a plan of action and stick with it until it's fully explored.\n5. Start early: It's better to start as early as possible and spend as much time as possible on data exploration, especially if you are new. Realized that making sense of masked data is not easy\n6. Make a team: A lone wolf (especially newbies) won't do, competitions such as this are way too overwhelming. Divide and rule would work better. \n\nSpecial thanks to @raddar for his curated dataset. I started with the raw csv dataset, it was a real pain and frustrating hardly could reach 0.6 score and I was on verge of giving up. Came across raddar's dataset and decided to give it a go and my scores jumped in range of 0.7. I was so excited that I called in sick next Monday 😜 and continued working. \n\nI hope to report a better score in my next competition, till then happy learning 👍",
    "1915026": "The more you compete, the more you'll realize how much of an art feature engineering is. Online courses barely even scratch the surface, the best place to kit up for your next competition would be the discussion forums once a contest ends. Happy kaggling!",
    "1915160": "I totally agree with you, I am also planning to work on some top approaches and try to improve my score with this dataset. Happy learning and enjoy the ride @ol0fmeister"
  },
  "source": "meta"
}