{
  "id": 265447,
  "title": "Notebook sharing should be re-worked!!",
  "url": "/competitions/seti-breakthrough-listen/discussion/265447",
  "author_name": "Adriano Passos",
  "post_date": "2021-08-15T20:03:56.512000",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>This competition is one of the biggest examples of a deep problem Kaggle system has. Sharing notebooks should be done in a way to share knowledge, not for simple farming upvotes and easy medals. By far the most frustrating thing about kaggle is to work hard for a whole week to make a submission just to find out that 80% of the participants just copied and pasted a submission. That is not only unfair with the others but makes it near impossible to grasp your true progress.</p>\n<p>This happens over and over on all competitions but this specific one is absolutely bad. We have dozens of 'ensemble' notebooks that don't provide any useful insight or knowledge and yet they ruin the hard work of hundreds of people. Rules of kaggle comps should change! there shouldn't be allowed to share results after a certain deadline of like a month or so before the comp ends.</p>",
  "messages": [
    {
      "id": 1473970,
      "postDate": "2021-08-15T20:03:56.513Z",
      "content": "<p>This competition is one of the biggest examples of a deep problem Kaggle system has. Sharing notebooks should be done in a way to share knowledge, not for simple farming upvotes and easy medals. By far the most frustrating thing about kaggle is to work hard for a whole week to make a submission just to find out that 80% of the participants just copied and pasted a submission. That is not only unfair with the others but makes it near impossible to grasp your true progress.</p>\n<p>This happens over and over on all competitions but this specific one is absolutely bad. We have dozens of 'ensemble' notebooks that don't provide any useful insight or knowledge and yet they ruin the hard work of hundreds of people. Rules of kaggle comps should change! there shouldn't be allowed to share results after a certain deadline of like a month or so before the comp ends.</p>",
      "rawMarkdown": "This competition is one of the biggest examples of a deep problem Kaggle system has. Sharing notebooks should be done in a way to share knowledge, not for simple farming upvotes and easy medals. By far the most frustrating thing about kaggle is to work hard for a whole week to make a submission just to find out that 80% of the participants just copied and pasted a submission. That is not only unfair with the others but makes it near impossible to grasp your true progress.\n\nThis happens over and over on all competitions but this specific one is absolutely bad. We have dozens of 'ensemble' notebooks that don't provide any useful insight or knowledge and yet they ruin the hard work of hundreds of people. Rules of kaggle comps should change! there shouldn't be allowed to share results after a certain deadline of like a month or so before the comp ends.",
      "votes": 15
    },
    {
      "id": 1474564,
      "postDate": "2021-08-16T07:05:51.567Z",
      "content": "<p>This is a specific problem in CSV competitions, because people just crazily ensemble different csv outputs, so that it often is even completely unclear where the solutions come from. While kernel competitions definitely do not solve this fully, the public kernel sharing, and even ensembling, is quite better.</p>",
      "rawMarkdown": "This is a specific problem in CSV competitions, because people just crazily ensemble different csv outputs, so that it often is even completely unclear where the solutions come from. While kernel competitions definitely do not solve this fully, the public kernel sharing, and even ensembling, is quite better.",
      "votes": 6,
      "replies": [
        {
          "id": 1474896,
          "postDate": "2021-08-16T10:22:56.747Z",
          "content": "<p><a href=\"https://www.kaggle.com/Psi\" target=\"_blank\">@Psi</a> agree that it is a larger problem for specific CSV competitions….but it should also be no problem for Kaggle to follow up more on making sure that the rules are followed within the last week.</p>\n<p>Published high scoring kernels and their publishers could be acted upon according to those rules.</p>\n<p>I stopped counting the number of 'Ensemble of Ensemble of Ensemble…. ' kernels that were published in the last few days where some Kagglers show off their data science skills on being able to change a digit from for example 0.09 to 0.10…..</p>",
          "rawMarkdown": "@Psi agree that it is a larger problem for specific CSV competitions....but it should also be no problem for Kaggle to follow up more on making sure that the rules are followed within the last week.\n\nPublished high scoring kernels and their publishers could be acted upon according to those rules.\n\nI stopped counting the number of 'Ensemble of Ensemble of Ensemble.... ' kernels that were published in the last few days where some Kagglers show off their data science skills on being able to change a digit from for example 0.09 to 0.10.....",
          "votes": 3
        }
      ]
    },
    {
      "id": 1480111,
      "postDate": "2021-08-18T20:05:33.460Z",
      "content": "<p>If there is one single thing that you will get that 80% of those that only fork and tweak a bit, is knowledge and experience. You might not win quickly but over a longer period, you will be a much better data scientist. At least that's how I view my experience here. </p>\n<p>That being said, I get your frustration and \"medal farming\" should be discouraged but it is what comes with a popular platform. 👌 </p>",
      "rawMarkdown": "If there is one single thing that you will get that 80% of those that only fork and tweak a bit, is knowledge and experience. You might not win quickly but over a longer period, you will be a much better data scientist. At least that's how I view my experience here. \n\nThat being said, I get your frustration and \"medal farming\" should be discouraged but it is what comes with a popular platform. 👌 ",
      "votes": 1
    },
    {
      "id": 1475049,
      "postDate": "2021-08-16T12:20:55.020Z",
      "content": "<p>But… wouldn't using someone else's submission.csv in your pipeline get you disqualified in the end? Or is that only for \"in the money\" results where the actual code used to create the winning submission will be reviewed by the organizers?</p>",
      "rawMarkdown": "But... wouldn't using someone else's submission.csv in your pipeline get you disqualified in the end? Or is that only for \"in the money\" results where the actual code used to create the winning submission will be reviewed by the organizers?",
      "votes": 2
    },
    {
      "id": 1474479,
      "postDate": "2021-08-16T06:11:30.413Z",
      "content": "<p>I very much agree with you, and you should not be allowed to share high score notebooks in the week before the end of the competition.</p>",
      "rawMarkdown": "I very much agree with you, and you should not be allowed to share high score notebooks in the week before the end of the competition.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1474564,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-08-16T07:05:51.567000",
      "content": "<p>This is a specific problem in CSV competitions, because people just crazily ensemble different csv outputs, so that it often is even completely unclear where the solutions come from. While kernel competitions definitely do not solve this fully, the public kernel sharing, and even ensembling, is quite better.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1474896,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2021-08-16T10:22:56.747000",
          "content": "<p><a href=\"https://www.kaggle.com/Psi\" target=\"_blank\">@Psi</a> agree that it is a larger problem for specific CSV competitions….but it should also be no problem for Kaggle to follow up more on making sure that the rules are followed within the last week.</p>\n<p>Published high scoring kernels and their publishers could be acted upon according to those rules.</p>\n<p>I stopped counting the number of 'Ensemble of Ensemble of Ensemble…. ' kernels that were published in the last few days where some Kagglers show off their data science skills on being able to change a digit from for example 0.09 to 0.10…..</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1480111,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-08-18T20:05:33.460000",
      "content": "<p>If there is one single thing that you will get that 80% of those that only fork and tweak a bit, is knowledge and experience. You might not win quickly but over a longer period, you will be a much better data scientist. At least that's how I view my experience here. </p>\n<p>That being said, I get your frustration and \"medal farming\" should be discouraged but it is what comes with a popular platform. 👌 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1475049,
      "author_name": "Markus Frank",
      "author_url": "",
      "post_date": "2021-08-16T12:20:55.020000",
      "content": "<p>But… wouldn't using someone else's submission.csv in your pipeline get you disqualified in the end? Or is that only for \"in the money\" results where the actual code used to create the winning submission will be reviewed by the organizers?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1474479,
      "author_name": "liyuling1030",
      "author_url": "",
      "post_date": "2021-08-16T06:11:30.413000",
      "content": "<p>I very much agree with you, and you should not be allowed to share high score notebooks in the week before the end of the competition.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1473970": "This competition is one of the biggest examples of a deep problem Kaggle system has. Sharing notebooks should be done in a way to share knowledge, not for simple farming upvotes and easy medals. By far the most frustrating thing about kaggle is to work hard for a whole week to make a submission just to find out that 80% of the participants just copied and pasted a submission. That is not only unfair with the others but makes it near impossible to grasp your true progress.\n\nThis happens over and over on all competitions but this specific one is absolutely bad. We have dozens of 'ensemble' notebooks that don't provide any useful insight or knowledge and yet they ruin the hard work of hundreds of people. Rules of kaggle comps should change! there shouldn't be allowed to share results after a certain deadline of like a month or so before the comp ends.",
    "1474564": "This is a specific problem in CSV competitions, because people just crazily ensemble different csv outputs, so that it often is even completely unclear where the solutions come from. While kernel competitions definitely do not solve this fully, the public kernel sharing, and even ensembling, is quite better.",
    "1480111": "If there is one single thing that you will get that 80% of those that only fork and tweak a bit, is knowledge and experience. You might not win quickly but over a longer period, you will be a much better data scientist. At least that's how I view my experience here. \n\nThat being said, I get your frustration and \"medal farming\" should be discouraged but it is what comes with a popular platform. 👌 ",
    "1475049": "But... wouldn't using someone else's submission.csv in your pipeline get you disqualified in the end? Or is that only for \"in the money\" results where the actual code used to create the winning submission will be reviewed by the organizers?",
    "1474479": "I very much agree with you, and you should not be allowed to share high score notebooks in the week before the end of the competition."
  }
}