{
  "id": 497754,
  "title": "Exploring What We Can Learn from Kaggle Competitions",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/497754",
  "author_name": "eggachecat",
  "post_date": "2024-04-25T15:32:17.112000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello community, this discussion may not be directly related to the competition itself, but please consider it as a casual chat about learning and growth. Everyone's input is very welcome. 🫡</p>\n<p>As a Software Development Engineer (SDE), I often find myself feeling lost in Kaggle competitions. I'm looking to understand what specific things we can learn from such competitions, or what we should expect to learn. (with some experiences with pandas/plt/numpy/tree boost models)</p>\n<p>In my career as a SDE, learning new frameworks or languages typically involves mastering their benefits and the problems they solve through demos and by referencing mature open-source projects. However, when participating in Kaggle competitions, I find that this approach to learning seems inadequate. The main methodologies seem to revolve around feature engineering, tree models, and model ensembles, but I'm uncertain what concrete learning outcomes these activities really provide.</p>\n<p>Here are some doubts and observations I have encountered during the competitions:</p>\n<ul>\n<li>Data Imbalance / Skew / Feature Engineering: This is a common issue across various fields. Why isn't there a universal framework to address this problem? Why is manual EDA (Exploratory Data Analysis) still necessary?</li>\n<li>Actual Value of Exploratory Data Analysis (EDA): Despite the emphasis on the importance of EDA, I often find that the graphs produced using tools like shapely and plt, although visually appealing, either lack sufficient information or the conclusions are too obvious or unclear. What can we really learn from EDA? What techniques are genuinely useful?</li>\n<li>Effectiveness of Feature Engineering and the GOAT models: If tree models or neural networks are presumed to be very effective, what is the real use of feature engineering done during EDA? If models can address specific issues on the fly, then what is the purpose of manual feature engineering?</li>\n</ul>\n<p>I'm not sure if anyone can relate to what I'm saying, but if you have similar doubts or questions, feel free to add your thoughts. Discussions on these questions and any other relevant topics are highly encouraged. Please participate actively.  🫡</p>",
  "messages": [
    {
      "id": 2775286,
      "postDate": "2024-04-25T15:32:17.113Z",
      "content": "<p>Hello community, this discussion may not be directly related to the competition itself, but please consider it as a casual chat about learning and growth. Everyone's input is very welcome. 🫡</p>\n<p>As a Software Development Engineer (SDE), I often find myself feeling lost in Kaggle competitions. I'm looking to understand what specific things we can learn from such competitions, or what we should expect to learn. (with some experiences with pandas/plt/numpy/tree boost models)</p>\n<p>In my career as a SDE, learning new frameworks or languages typically involves mastering their benefits and the problems they solve through demos and by referencing mature open-source projects. However, when participating in Kaggle competitions, I find that this approach to learning seems inadequate. The main methodologies seem to revolve around feature engineering, tree models, and model ensembles, but I'm uncertain what concrete learning outcomes these activities really provide.</p>\n<p>Here are some doubts and observations I have encountered during the competitions:</p>\n<ul>\n<li>Data Imbalance / Skew / Feature Engineering: This is a common issue across various fields. Why isn't there a universal framework to address this problem? Why is manual EDA (Exploratory Data Analysis) still necessary?</li>\n<li>Actual Value of Exploratory Data Analysis (EDA): Despite the emphasis on the importance of EDA, I often find that the graphs produced using tools like shapely and plt, although visually appealing, either lack sufficient information or the conclusions are too obvious or unclear. What can we really learn from EDA? What techniques are genuinely useful?</li>\n<li>Effectiveness of Feature Engineering and the GOAT models: If tree models or neural networks are presumed to be very effective, what is the real use of feature engineering done during EDA? If models can address specific issues on the fly, then what is the purpose of manual feature engineering?</li>\n</ul>\n<p>I'm not sure if anyone can relate to what I'm saying, but if you have similar doubts or questions, feel free to add your thoughts. Discussions on these questions and any other relevant topics are highly encouraged. Please participate actively.  🫡</p>",
      "rawMarkdown": "Hello community, this discussion may not be directly related to the competition itself, but please consider it as a casual chat about learning and growth. Everyone's input is very welcome. 🫡\n\nAs a Software Development Engineer (SDE), I often find myself feeling lost in Kaggle competitions. I'm looking to understand what specific things we can learn from such competitions, or what we should expect to learn. (with some experiences with pandas/plt/numpy/tree boost models)\n\nIn my career as a SDE, learning new frameworks or languages typically involves mastering their benefits and the problems they solve through demos and by referencing mature open-source projects. However, when participating in Kaggle competitions, I find that this approach to learning seems inadequate. The main methodologies seem to revolve around feature engineering, tree models, and model ensembles, but I'm uncertain what concrete learning outcomes these activities really provide.\n\nHere are some doubts and observations I have encountered during the competitions:\n\n- Data Imbalance / Skew / Feature Engineering: This is a common issue across various fields. Why isn't there a universal framework to address this problem? Why is manual EDA (Exploratory Data Analysis) still necessary?\n- Actual Value of Exploratory Data Analysis (EDA): Despite the emphasis on the importance of EDA, I often find that the graphs produced using tools like shapely and plt, although visually appealing, either lack sufficient information or the conclusions are too obvious or unclear. What can we really learn from EDA? What techniques are genuinely useful?\n- Effectiveness of Feature Engineering and the GOAT models: If tree models or neural networks are presumed to be very effective, what is the real use of feature engineering done during EDA? If models can address specific issues on the fly, then what is the purpose of manual feature engineering?\n\nI'm not sure if anyone can relate to what I'm saying, but if you have similar doubts or questions, feel free to add your thoughts. Discussions on these questions and any other relevant topics are highly encouraged. Please participate actively.  🫡",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2775286": "Hello community, this discussion may not be directly related to the competition itself, but please consider it as a casual chat about learning and growth. Everyone's input is very welcome. 🫡\n\nAs a Software Development Engineer (SDE), I often find myself feeling lost in Kaggle competitions. I'm looking to understand what specific things we can learn from such competitions, or what we should expect to learn. (with some experiences with pandas/plt/numpy/tree boost models)\n\nIn my career as a SDE, learning new frameworks or languages typically involves mastering their benefits and the problems they solve through demos and by referencing mature open-source projects. However, when participating in Kaggle competitions, I find that this approach to learning seems inadequate. The main methodologies seem to revolve around feature engineering, tree models, and model ensembles, but I'm uncertain what concrete learning outcomes these activities really provide.\n\nHere are some doubts and observations I have encountered during the competitions:\n\n- Data Imbalance / Skew / Feature Engineering: This is a common issue across various fields. Why isn't there a universal framework to address this problem? Why is manual EDA (Exploratory Data Analysis) still necessary?\n- Actual Value of Exploratory Data Analysis (EDA): Despite the emphasis on the importance of EDA, I often find that the graphs produced using tools like shapely and plt, although visually appealing, either lack sufficient information or the conclusions are too obvious or unclear. What can we really learn from EDA? What techniques are genuinely useful?\n- Effectiveness of Feature Engineering and the GOAT models: If tree models or neural networks are presumed to be very effective, what is the real use of feature engineering done during EDA? If models can address specific issues on the fly, then what is the purpose of manual feature engineering?\n\nI'm not sure if anyone can relate to what I'm saying, but if you have similar doubts or questions, feel free to add your thoughts. Discussions on these questions and any other relevant topics are highly encouraged. Please participate actively.  🫡"
  }
}