{
  "id": 188590,
  "title": "Most realistic competition ever",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/188590",
  "author_name": "Gunes Evitan",
  "post_date": "2020-10-04T05:29:53.568000",
  "votes": 36,
  "comment_count": 13,
  "views": 0,
  "content": "<p>As some of you know, Kaggle competitions are trashed on some platforms for not reflecting the \"real  world data science projects\", but I can say that I found this competition very similar to my daily work.</p>\n<p>I think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set). Modelling aspect was important of course, but not as much important as data preparation. I would definitely want to see competitions like this in the future. </p>",
  "messages": [
    {
      "id": 1036641,
      "postDate": "2020-10-04T05:29:53.570Z",
      "content": "<p>As some of you know, Kaggle competitions are trashed on some platforms for not reflecting the \"real  world data science projects\", but I can say that I found this competition very similar to my daily work.</p>\n<p>I think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set). Modelling aspect was important of course, but not as much important as data preparation. I would definitely want to see competitions like this in the future. </p>",
      "rawMarkdown": "As some of you know, Kaggle competitions are trashed on some platforms for not reflecting the \"real  world data science projects\", but I can say that I found this competition very similar to my daily work.\n\nI think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set). Modelling aspect was important of course, but not as much important as data preparation. I would definitely want to see competitions like this in the future. \n\n\n",
      "votes": 36
    },
    {
      "id": 1036696,
      "postDate": "2020-10-04T07:16:54.680Z",
      "content": "<p>I think this competition requires certain level of data literacy and engineering skill. Copy-pasting data scientists have no chance to win.<br>\nAnd data is pretty noisy, dirty and small which makes this competition difficult.</p>",
      "rawMarkdown": "I think this competition requires certain level of data literacy and engineering skill. Copy-pasting data scientists have no chance to win.\nAnd data is pretty noisy, dirty and small which makes this competition difficult.",
      "votes": 5,
      "replies": [
        {
          "id": 1036708,
          "postDate": "2020-10-04T07:36:46.717Z",
          "content": "<p>You are right. I think teams using public kernels may end in a bad spot.</p>",
          "rawMarkdown": "You are right. I think teams using public kernels may end in a bad spot.",
          "votes": 4
        },
        {
          "id": 1036719,
          "postDate": "2020-10-04T07:51:17.850Z",
          "content": "<p>Well, I started from copying a public notebook and wasted a couple of days on flaws of notebook…<br>\nI was astonished that public notebooks have so many flaws and badly overfitting to LB.<br>\nThere's too many data scientists who don't understand about data.</p>",
          "rawMarkdown": "Well, I started from copying a public notebook and wasted a couple of days on flaws of notebook...\nI was astonished that public notebooks have so many flaws and badly overfitting to LB.\nThere's too many data scientists who don't understand about data.",
          "votes": 4
        },
        {
          "id": 1036886,
          "postDate": "2020-10-04T11:38:25.093Z",
          "content": "<p>I stopped starting from public notebooks long time ago since I noticed refactoring them actually takes more time than writing my own code.</p>",
          "rawMarkdown": "I stopped starting from public notebooks long time ago since I noticed refactoring them actually takes more time than writing my own code.",
          "votes": 7
        },
        {
          "id": 1038228,
          "postDate": "2020-10-05T16:56:32.687Z",
          "content": "<p>Yes, you are absolutely right, just like you, I try to write my own code, and sometimes I just use the public kernels to get new ideas, and this helps me.</p>",
          "rawMarkdown": "Yes, you are absolutely right, just like you, I try to write my own code, and sometimes I just use the public kernels to get new ideas, and this helps me.",
          "votes": 1
        },
        {
          "id": 1038480,
          "postDate": "2020-10-05T20:08:35.470Z",
          "content": "<p>Recently I've been recoding public notebooks \"in my own words\" or \"in my own code\".  I think 1) it shows that you understand what is being done 2) forces you to think about how you might do the same thing but potentially more efficient or just in a different way 3) you might find that some things are unnecessary due to all the forking and so you can slim the code down.  For me, this has proven really useful.  It's also great because it's interesting to see how others program.  Recently I took someone's code and was able to do the same thing in 20 lines vs. 100.  I've also seen the opposite, where some else is doing something in fewer lines then myself.  So it's cool to see how others solve problems.        </p>\n<p>It's hard when a lot of people just fork and blend.  For example, you could have 1 person that does 100% original work, and maybe they get a few bronze medals but at the same time maybe they end up with a lot of competitions less than 80% (looks bad) due to all the fork and blenders.  Then say you have someone else that is consistently greater than 20% leaderboard (looks good) but all their work is fork and blend (not good).  So Kaggler #2's profile (fork and blender) probably looks better than Kaggler #1's profile (original work) but I would be cautious to hire Kaggler #2.  At the same time, if all you do is original work, you're missing 1/2 the benefit of Kaggle, which is collaboration.  So it's a balance.  The pitfall is getting sucked into 100% fork &amp; blend work just because you find it hard to compete due to all the other fork and blenders.  On one hand, you might start doing better in the leaderboards, but in the end you're doing yourself a disservice as a data scientist because you're curbing your learning (which should really be the end goal, cause if it's just to win $$, there are better ways to earn a buck).         </p>",
          "rawMarkdown": "Recently I've been recoding public notebooks \"in my own words\" or \"in my own code\".  I think 1) it shows that you understand what is being done 2) forces you to think about how you might do the same thing but potentially more efficient or just in a different way 3) you might find that some things are unnecessary due to all the forking and so you can slim the code down.  For me, this has proven really useful.  It's also great because it's interesting to see how others program.  Recently I took someone's code and was able to do the same thing in 20 lines vs. 100.  I've also seen the opposite, where some else is doing something in fewer lines then myself.  So it's cool to see how others solve problems.        \n\nIt's hard when a lot of people just fork and blend.  For example, you could have 1 person that does 100% original work, and maybe they get a few bronze medals but at the same time maybe they end up with a lot of competitions less than 80% (looks bad) due to all the fork and blenders.  Then say you have someone else that is consistently greater than 20% leaderboard (looks good) but all their work is fork and blend (not good).  So Kaggler #2's profile (fork and blender) probably looks better than Kaggler #1's profile (original work) but I would be cautious to hire Kaggler #2.  At the same time, if all you do is original work, you're missing 1/2 the benefit of Kaggle, which is collaboration.  So it's a balance.  The pitfall is getting sucked into 100% fork & blend work just because you find it hard to compete due to all the other fork and blenders.  On one hand, you might start doing better in the leaderboards, but in the end you're doing yourself a disservice as a data scientist because you're curbing your learning (which should really be the end goal, cause if it's just to win $$, there are better ways to earn a buck).         ",
          "votes": 4
        }
      ]
    },
    {
      "id": 1037568,
      "postDate": "2020-10-05T06:06:44.990Z",
      "content": "<p>I agree. Just like in the real world, the winning models of this competition will bring no useful insight to the problem itself. The best linear approximation shall prevail ;)</p>",
      "rawMarkdown": "I agree. Just like in the real world, the winning models of this competition will bring no useful insight to the problem itself. The best linear approximation shall prevail ;)",
      "votes": 6,
      "replies": [
        {
          "id": 1040463,
          "postDate": "2020-10-07T06:43:44.320Z",
          "content": "<p>…aaaand, <a href=\"https://www.kaggle.com/jerryb11\" target=\"_blank\">@jerryb11</a> is right =)</p>",
          "rawMarkdown": "...aaaand, @jerryb11 is right =)"
        }
      ]
    },
    {
      "id": 1036652,
      "postDate": "2020-10-04T05:56:12.267Z",
      "content": "<p><code>I think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set)</code></p>\n<p>It's so true because I've seen public notebooks score very high on the public LB. But when I look at their predictions they're not realistic and are probably overfitting.</p>",
      "rawMarkdown": "`I think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set)`\n\nIt's so true because I've seen public notebooks score very high on the public LB. But when I look at their predictions they're not realistic and are probably overfitting.",
      "votes": 3
    },
    {
      "id": 1038950,
      "postDate": "2020-10-06T07:32:02.560Z",
      "content": "<p>Definitely agree with you <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. Data Preparation was the most important task in this competition rather than modelling. Our daily work is mostly surrounded by Data prep only so that will help a lot</p>",
      "rawMarkdown": "Definitely agree with you @gunesevitan. Data Preparation was the most important task in this competition rather than modelling. Our daily work is mostly surrounded by Data prep only so that will help a lot",
      "votes": 1
    },
    {
      "id": 1038315,
      "postDate": "2020-10-05T18:14:13.663Z",
      "content": "<p>Good comment. </p>",
      "rawMarkdown": "Good comment. "
    },
    {
      "id": 1036768,
      "postDate": "2020-10-04T09:03:12.310Z",
      "content": "<p>Yes, I concur. Challenging to try and focus on getting good CV - LB correlation, be creative and work outside the realms of public kernels, try innovate featuring and as <a href=\"https://www.kaggle.com/resistance0108\" target=\"_blank\">@resistance0108</a> said, data is noisy and dirty much like a \"real-world scenario\".</p>",
      "rawMarkdown": "Yes, I concur. Challenging to try and focus on getting good CV - LB correlation, be creative and work outside the realms of public kernels, try innovate featuring and as @resistance0108 said, data is noisy and dirty much like a \"real-world scenario\"."
    },
    {
      "id": 1038746,
      "postDate": "2020-10-06T02:57:08.520Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1036696,
      "author_name": "resistance0108",
      "author_url": "",
      "post_date": "2020-10-04T07:16:54.680000",
      "content": "<p>I think this competition requires certain level of data literacy and engineering skill. Copy-pasting data scientists have no chance to win.<br>\nAnd data is pretty noisy, dirty and small which makes this competition difficult.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1036708,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-10-04T07:36:46.717000",
          "content": "<p>You are right. I think teams using public kernels may end in a bad spot.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1036719,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "2020-10-04T07:51:17.850000",
          "content": "<p>Well, I started from copying a public notebook and wasted a couple of days on flaws of notebook…<br>\nI was astonished that public notebooks have so many flaws and badly overfitting to LB.<br>\nThere's too many data scientists who don't understand about data.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1036886,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2020-10-04T11:38:25.093000",
          "content": "<p>I stopped starting from public notebooks long time ago since I noticed refactoring them actually takes more time than writing my own code.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1038228,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-05T16:56:32.687000",
          "content": "<p>Yes, you are absolutely right, just like you, I try to write my own code, and sometimes I just use the public kernels to get new ideas, and this helps me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1038480,
          "author_name": "Matt Yates",
          "author_url": "",
          "post_date": "2020-10-05T20:08:35.470000",
          "content": "<p>Recently I've been recoding public notebooks \"in my own words\" or \"in my own code\".  I think 1) it shows that you understand what is being done 2) forces you to think about how you might do the same thing but potentially more efficient or just in a different way 3) you might find that some things are unnecessary due to all the forking and so you can slim the code down.  For me, this has proven really useful.  It's also great because it's interesting to see how others program.  Recently I took someone's code and was able to do the same thing in 20 lines vs. 100.  I've also seen the opposite, where some else is doing something in fewer lines then myself.  So it's cool to see how others solve problems.        </p>\n<p>It's hard when a lot of people just fork and blend.  For example, you could have 1 person that does 100% original work, and maybe they get a few bronze medals but at the same time maybe they end up with a lot of competitions less than 80% (looks bad) due to all the fork and blenders.  Then say you have someone else that is consistently greater than 20% leaderboard (looks good) but all their work is fork and blend (not good).  So Kaggler #2's profile (fork and blender) probably looks better than Kaggler #1's profile (original work) but I would be cautious to hire Kaggler #2.  At the same time, if all you do is original work, you're missing 1/2 the benefit of Kaggle, which is collaboration.  So it's a balance.  The pitfall is getting sucked into 100% fork &amp; blend work just because you find it hard to compete due to all the other fork and blenders.  On one hand, you might start doing better in the leaderboards, but in the end you're doing yourself a disservice as a data scientist because you're curbing your learning (which should really be the end goal, cause if it's just to win $$, there are better ways to earn a buck).         </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1037568,
      "author_name": "Jerry@B11",
      "author_url": "",
      "post_date": "2020-10-05T06:06:44.990000",
      "content": "<p>I agree. Just like in the real world, the winning models of this competition will bring no useful insight to the problem itself. The best linear approximation shall prevail ;)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1040463,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-10-07T06:43:44.320000",
          "content": "<p>…aaaand, <a href=\"https://www.kaggle.com/jerryb11\" target=\"_blank\">@jerryb11</a> is right =)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1036652,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2020-10-04T05:56:12.267000",
      "content": "<p><code>I think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set)</code></p>\n<p>It's so true because I've seen public notebooks score very high on the public LB. But when I look at their predictions they're not realistic and are probably overfitting.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1038950,
      "author_name": "VAIBHAV MATHUR",
      "author_url": "",
      "post_date": "2020-10-06T07:32:02.560000",
      "content": "<p>Definitely agree with you <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. Data Preparation was the most important task in this competition rather than modelling. Our daily work is mostly surrounded by Data prep only so that will help a lot</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1038315,
      "author_name": "Sophie",
      "author_url": "",
      "post_date": "2020-10-05T18:14:13.663000",
      "content": "<p>Good comment. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1036768,
      "author_name": "Trigram",
      "author_url": "",
      "post_date": "2020-10-04T09:03:12.310000",
      "content": "<p>Yes, I concur. Challenging to try and focus on getting good CV - LB correlation, be creative and work outside the realms of public kernels, try innovate featuring and as <a href=\"https://www.kaggle.com/resistance0108\" target=\"_blank\">@resistance0108</a> said, data is noisy and dirty much like a \"real-world scenario\".</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1038746,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-06T02:57:08.520000",
      "content": "",
      "votes": -4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1036641": "As some of you know, Kaggle competitions are trashed on some platforms for not reflecting the \"real  world data science projects\", but I can say that I found this competition very similar to my daily work.\n\nI think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set). Modelling aspect was important of course, but not as much important as data preparation. I would definitely want to see competitions like this in the future. \n\n\n",
    "1036696": "I think this competition requires certain level of data literacy and engineering skill. Copy-pasting data scientists have no chance to win.\nAnd data is pretty noisy, dirty and small which makes this competition difficult.",
    "1037568": "I agree. Just like in the real world, the winning models of this competition will bring no useful insight to the problem itself. The best linear approximation shall prevail ;)",
    "1036652": "`I think the main challenge of this competition was creating a robust bug-free data preparation pipeline for unseen real world data (private test set)`\n\nIt's so true because I've seen public notebooks score very high on the public LB. But when I look at their predictions they're not realistic and are probably overfitting.",
    "1038950": "Definitely agree with you @gunesevitan. Data Preparation was the most important task in this competition rather than modelling. Our daily work is mostly surrounded by Data prep only so that will help a lot",
    "1038315": "Good comment. ",
    "1036768": "Yes, I concur. Challenging to try and focus on getting good CV - LB correlation, be creative and work outside the realms of public kernels, try innovate featuring and as @resistance0108 said, data is noisy and dirty much like a \"real-world scenario\".",
    "1038746": ""
  }
}