{
  "id": 508326,
  "title": "Summary of my first kaggle competition",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508326",
  "author_name": "",
  "post_date": "2024-05-29T07:14:55.326209200Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>My main problems in this competition:<br>\n1 pursues LB scores too much, leading to overfitting.<br>\n2 did not find a suitable CV scheme. I've actually designed several CV solutions, but none of them work well.<br>\n3 Lack confidence. I found a lot of interesting patterns in the dataset, but only using LB scores to validate my ideas eventually made me give up on these ideas, and I actually found these features to be effective by reading other people's code.<br>\n4 Wrong commit strategy. Only the best LB scores are selected to submit, which is essentially an extension of problem 2- no reliable CV-.<br>\nHere's the takeaway:<br>\n1 improved the ability to process data, this was the first time I had to process data and build a model with limited memory.<br>\n2 Learn at least two pipeline frameworks to gain an overall understanding of the machine learning process.<br>\n3 Learn to use GPU acceleration.<br>\nWhat I hope to continue to learn:<br>\n1 Advanced feature engineering. I still couldn't judge the quality of my hand-crafted features and how to create robust new features.<br>\n2  New ways to aggregate data. My teammates and I tried using deep learning to generate aggregated features, but it didn't work well. I wanted to see how other big players did it and wanted to learn.<br>\n3 Robust CV. It was my worst defeat in this competition, so I want to see how others do it and learn from it.<br>\n4 Pseudo-labels. I tried pseudo tags but they didn't work well. I wanted to learn from someone else's code.<br>\nAlthough I was actually a bit angry and sad about my first competition, I will continue on my kaggle journey and I look forward to seeing your top solutions without hack. I will study and thank you.</p>",
  "messages": [
    {
      "id": "2842604",
      "postDate": "05/29/2024 07:14:55",
      "content": "<p>My main problems in this competition:<br>\n1 pursues LB scores too much, leading to overfitting.<br>\n2 did not find a suitable CV scheme. I've actually designed several CV solutions, but none of them work well.<br>\n3 Lack confidence. I found a lot of interesting patterns in the dataset, but only using LB scores to validate my ideas eventually made me give up on these ideas, and I actually found these features to be effective by reading other people's code.<br>\n4 Wrong commit strategy. Only the best LB scores are selected to submit, which is essentially an extension of problem 2- no reliable CV-.<br>\nHere's the takeaway:<br>\n1 improved the ability to process data, this was the first time I had to process data and build a model with limited memory.<br>\n2 Learn at least two pipeline frameworks to gain an overall understanding of the machine learning process.<br>\n3 Learn to use GPU acceleration.<br>\nWhat I hope to continue to learn:<br>\n1 Advanced feature engineering. I still couldn't judge the quality of my hand-crafted features and how to create robust new features.<br>\n2  New ways to aggregate data. My teammates and I tried using deep learning to generate aggregated features, but it didn't work well. I wanted to see how other big players did it and wanted to learn.<br>\n3 Robust CV. It was my worst defeat in this competition, so I want to see how others do it and learn from it.<br>\n4 Pseudo-labels. I tried pseudo tags but they didn't work well. I wanted to learn from someone else's code.<br>\nAlthough I was actually a bit angry and sad about my first competition, I will continue on my kaggle journey and I look forward to seeing your top solutions without hack. I will study and thank you.</p>",
      "rawMarkdown": "My main problems in this competition:\n\n1 pursues LB scores too much, leading to overfitting.\n2 did not find a suitable CV scheme. I've actually designed several CV solutions, but none of them work well.\n3 Lack confidence. I found a lot of interesting patterns in the dataset, but only using LB scores to validate my ideas eventually made me give up on these ideas, and I actually found these features to be effective by reading other people's code.\n4 Wrong commit strategy. Only the best LB scores are selected to submit, which is essentially an extension of problem 2- no reliable CV-.\n\nHere's the takeaway:\n\n1 improved the ability to process data, this was the first time I had to process data and build a model with limited memory.\n2 Learn at least two pipeline frameworks to gain an overall understanding of the machine learning process.\n3 Learn to use GPU acceleration.\n\nWhat I hope to continue to learn:\n\n1 Advanced feature engineering. I still couldn't judge the quality of my hand-crafted features and how to create robust new features.\n2  New ways to aggregate data. My teammates and I tried using deep learning to generate aggregated features, but it didn't work well. I wanted to see how other big players did it and wanted to learn.\n3 Robust CV. It was my worst defeat in this competition, so I want to see how others do it and learn from it.\n4 Pseudo-labels. I tried pseudo tags but they didn't work well. I wanted to learn from someone else's code.\n\nAlthough I was actually a bit angry and sad about my first competition, I will continue on my kaggle journey and I look forward to seeing your top solutions without hack. I will study and thank you.",
      "votes": null
    },
    {
      "id": "2842677",
      "postDate": "05/29/2024 08:00:21",
      "content": "<p>analysing your mistakes and creating a plan to learn from them is one of the best strategies to be successful in data science, well done keep it up</p>",
      "rawMarkdown": "analysing your mistakes and creating a plan to learn from them is one of the best strategies to be successful in data science, well done keep it up",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2842677,
      "author_name": "hassie9698",
      "author_url": "",
      "post_date": "05/29/2024 08:00:21",
      "content": "<p>analysing your mistakes and creating a plan to learn from them is one of the best strategies to be successful in data science, well done keep it up</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2842604": "My main problems in this competition:\n\n1 pursues LB scores too much, leading to overfitting.\n2 did not find a suitable CV scheme. I've actually designed several CV solutions, but none of them work well.\n3 Lack confidence. I found a lot of interesting patterns in the dataset, but only using LB scores to validate my ideas eventually made me give up on these ideas, and I actually found these features to be effective by reading other people's code.\n4 Wrong commit strategy. Only the best LB scores are selected to submit, which is essentially an extension of problem 2- no reliable CV-.\n\nHere's the takeaway:\n\n1 improved the ability to process data, this was the first time I had to process data and build a model with limited memory.\n2 Learn at least two pipeline frameworks to gain an overall understanding of the machine learning process.\n3 Learn to use GPU acceleration.\n\nWhat I hope to continue to learn:\n\n1 Advanced feature engineering. I still couldn't judge the quality of my hand-crafted features and how to create robust new features.\n2  New ways to aggregate data. My teammates and I tried using deep learning to generate aggregated features, but it didn't work well. I wanted to see how other big players did it and wanted to learn.\n3 Robust CV. It was my worst defeat in this competition, so I want to see how others do it and learn from it.\n4 Pseudo-labels. I tried pseudo tags but they didn't work well. I wanted to learn from someone else's code.\n\nAlthough I was actually a bit angry and sad about my first competition, I will continue on my kaggle journey and I look forward to seeing your top solutions without hack. I will study and thank you.",
    "2842677": "analysing your mistakes and creating a plan to learn from them is one of the best strategies to be successful in data science, well done keep it up"
  },
  "source": "meta"
}