{
  "id": 516249,
  "title": "BROKEN Leaderboard! DISAPPOINTING Competition Experience!",
  "url": "/competitions/leash-BELKA/discussion/516249",
  "author_name": "",
  "post_date": "2024-07-01T22:26:51.953731Z",
  "votes": 9,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I'm a newbie to the Kaggle competition and disappointed with the experience. This morning, I reached out to several teams to ask for collaboration. After accepting one guy's invitation, I found that, in the submission history, the guy didn't have any original submission files. He just reran some shared notebooks that ensembles some uploaded prediction files by others, such as \"https://www.kaggle.com/code/andrey67/leash-bio-blender-0-440/notebook\", \"https://www.kaggle.com/code/saberghaderi/leash-bio-4149\". It achieves a very high LB score like 0.448, which destroys all the efforts most competitors have put into the competition.  When you look at the leaderboard carefully, especially with score 0.448, 0.433, 0.432, you will find a couple of teams reach such a high score with only one or two submission entries. It's possible that they also just ran the same shared notebooks to get the same score. Imagine, a guy, who didn't conduct any experiments for the competition, gets the top 10% ranking. How RIDICULOUS it is!</p>\n<p>Also, I believe the one who shares the ensemble notebook should also be responsible for the pollution. If someone really shares some ensemble techniques, please use dummy prediction files. That would be helpful and thankful for all the competitors.</p>\n<p>Anyway, BROKEN leaderboard! DISAPPOINTING competition Experience!</p>",
  "messages": [
    {
      "id": "2899878",
      "postDate": "07/01/2024 22:26:51",
      "content": "<p>I'm a newbie to the Kaggle competition and disappointed with the experience. This morning, I reached out to several teams to ask for collaboration. After accepting one guy's invitation, I found that, in the submission history, the guy didn't have any original submission files. He just reran some shared notebooks that ensembles some uploaded prediction files by others, such as \"https://www.kaggle.com/code/andrey67/leash-bio-blender-0-440/notebook\", \"https://www.kaggle.com/code/saberghaderi/leash-bio-4149\". It achieves a very high LB score like 0.448, which destroys all the efforts most competitors have put into the competition.  When you look at the leaderboard carefully, especially with score 0.448, 0.433, 0.432, you will find a couple of teams reach such a high score with only one or two submission entries. It's possible that they also just ran the same shared notebooks to get the same score. Imagine, a guy, who didn't conduct any experiments for the competition, gets the top 10% ranking. How RIDICULOUS it is!</p>\n<p>Also, I believe the one who shares the ensemble notebook should also be responsible for the pollution. If someone really shares some ensemble techniques, please use dummy prediction files. That would be helpful and thankful for all the competitors.</p>\n<p>Anyway, BROKEN leaderboard! DISAPPOINTING competition Experience!</p>",
      "rawMarkdown": "I'm a newbie to the Kaggle competition and disappointed with the experience. This morning, I reached out to several teams to ask for collaboration. After accepting one guy's invitation, I found that, in the submission history, the guy didn't have any original submission files. He just reran some shared notebooks that ensembles some uploaded prediction files by others, such as \"https://www.kaggle.com/code/andrey67/leash-bio-blender-0-440/notebook\", \"https://www.kaggle.com/code/saberghaderi/leash-bio-4149\". It achieves a very high LB score like 0.448, which destroys all the efforts most competitors have put into the competition.  When you look at the leaderboard carefully, especially with score 0.448, 0.433, 0.432, you will find a couple of teams reach such a high score with only one or two submission entries. It's possible that they also just ran the same shared notebooks to get the same score. Imagine, a guy, who didn't conduct any experiments for the competition, gets the top 10% ranking. How RIDICULOUS it is!\n\nAlso, I believe the one who shares the ensemble notebook should also be responsible for the pollution. If someone really shares some ensemble techniques, please use dummy prediction files. That would be helpful and thankful for all the competitors.\n\nAnyway, BROKEN leaderboard! DISAPPOINTING competition Experience!",
      "votes": null
    },
    {
      "id": "2900351",
      "postDate": "07/02/2024 08:11:14",
      "content": "<p>This is common in competitions that are not code competitions. In the open problems competition 6 months ago, the  lb/private index were shared and the notebooks were shared with manual lb overfitting (e.g. id1 +=0.2,id 2 -=0.1, etc.)<br>\nYou can consider teaming up by looking at the public lb, the person's lb, and the number of times the person has submitted.</p>\n<p>However, this is only the ranking of the lb, so your efforts will not be in vain depending on the result of the private data.<br>\nGood luck!</p>",
      "rawMarkdown": "This is common in competitions that are not code competitions. In the open problems competition 6 months ago, the  lb/private index were shared and the notebooks were shared with manual lb overfitting (e.g. id1 +=0.2,id 2 -=0.1, etc.)\nYou can consider teaming up by looking at the public lb, the person's lb, and the number of times the person has submitted.\n\nHowever, this is only the ranking of the lb, so your efforts will not be in vain depending on the result of the private data.\nGood luck!",
      "votes": null
    },
    {
      "id": "2900354",
      "postDate": "07/02/2024 08:13:58",
      "content": "<p><a href=\"https://www.kaggle.com/ganglii\" target=\"_blank\">@ganglii</a> Do remember that it is not the public LB ranking that decides the results of the competition. As discussed in multiple places in this forum, the private LB data are significantly different and a large shake-up is anticipated in the final private LB ranking that decides prizes and medals. Getting a high public ranking guarantees nothing.</p>\n<p>The normal ethical custom and practice of Kaggle competitions is that public notebooks will inevitably get submitted multiple times, both as people's developments of them and in the original form. It is fine to wish that the rules were different and to argue that they should be changed, but such a public spat is unfair on your easily identifiable teammate who has acted within the regulations of the competition.</p>",
      "rawMarkdown": "ganglii Do remember that it is not the public LB ranking that decides the results of the competition. As discussed in multiple places in this forum, the private LB data are significantly different and a large shake-up is anticipated in the final private LB ranking that decides prizes and medals. Getting a high public ranking guarantees nothing.\n\nThe normal ethical custom and practice of Kaggle competitions is that public notebooks will inevitably get submitted multiple times, both as people's developments of them and in the original form. It is fine to wish that the rules were different and to argue that they should be changed, but such a public spat is unfair on your easily identifiable teammate who has acted within the regulations of the competition.",
      "votes": null
    },
    {
      "id": "2900586",
      "postDate": "07/02/2024 11:33:20",
      "content": "<p>If you are so disappointed by this competition, what do you say about <a href=\"https://www.kaggle.com/competitions/birdclef-2024\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024</a>?<br>\nPlaces 32-245 are public notebook submits.</p>",
      "rawMarkdown": "If you are so disappointed by this competition, what do you say about https://www.kaggle.com/competitions/birdclef-2024?\nPlaces 32-245 are public notebook submits.",
      "votes": null
    },
    {
      "id": "2901016",
      "postDate": "07/02/2024 15:54:22",
      "content": "<p>If this kind of action is allowable and advocated in this community, I think exposure should not matter as people in the community understand him. And I didn't mean to blame someone who ran public notebooks. Instead, I think it is the action of sharing this kind of meaningless notebooks which only submits some uploaded prediction files that destroys the competition community. These notebooks don't bring any knowledge for other competitors to learn but pollute the whole lb. I didn't see any benefits of this kind of sharing.</p>",
      "rawMarkdown": "If this kind of action is allowable and advocated in this community, I think exposure should not matter as people in the community understand him. And I didn't mean to blame someone who ran public notebooks. Instead, I think it is the action of sharing this kind of meaningless notebooks which only submits some uploaded prediction files that destroys the competition community. These notebooks don't bring any knowledge for other competitors to learn but pollute the whole lb. I didn't see any benefits of this kind of sharing.",
      "votes": null
    },
    {
      "id": "2901045",
      "postDate": "07/02/2024 16:06:21",
      "content": "<p>If the public notebooks share real model training or data processing techniques, they are nice notebooks as they bring knowledge to others. If the public notebooks only teach people how to submit other uploaded prediction files, they are pure pollution to the lb. There are around 1000 teams in the competition you mentioned, I believe teams that just submitted shared prediction files don't even deserve a ranking of 250, top 25%. And I think the pollution should be attributed to someone who share these prediction files.</p>",
      "rawMarkdown": "If the public notebooks share real model training or data processing techniques, they are nice notebooks as they bring knowledge to others. If the public notebooks only teach people how to submit other uploaded prediction files, they are pure pollution to the lb. There are around 1000 teams in the competition you mentioned, I believe teams that just submitted shared prediction files don't even deserve a ranking of 250, top 25%. And I think the pollution should be attributed to someone who share these prediction files.",
      "votes": null
    },
    {
      "id": "2901057",
      "postDate": "07/02/2024 16:09:42",
      "content": "<p>Thanks for your suggestions about choosing teammates. I got a lesson this time.</p>",
      "rawMarkdown": "Thanks for your suggestions about choosing teammates. I got a lesson this time.",
      "votes": null
    },
    {
      "id": "2901185",
      "postDate": "07/02/2024 17:09:13",
      "content": "<p><a href=\"https://www.kaggle.com/ganglii\" target=\"_blank\">@ganglii</a>  This is a question with good arguments on both sides, and one which splits opinion on Kaggle.</p>\n<p>Sharing code gets upvotes on notebooks (sometimes, at least), and helps Kagglers progress through to higher tiers. It also gives people starting points for developing models. I would enter far fewer competitions if working code weren't provided as a starting point. Competitions that didn't allow shared models would be a lot less accessible. </p>\n<p>Generally the best public submissions act as a de facto baseline - if you want a medal, you have to do quite a lot better than that in most cases (but occasionally copy-and-submit has achieved this, as others have pointed out with regard to BirdCLEF). This one is probably going to be different, as we don't really we much idea what a good model is from the public LB scores alone.</p>\n<p>Certainly not all the shared notebooks are meaningless or useless. Those shared in this competitions include some helpful descriptor-based cheminformatics models and also tokenization-based approaches. There are some interesting solutions to the problems of trying to  load or otherwise handle a very large dataset too. </p>",
      "rawMarkdown": "ganglii  This is a question with good arguments on both sides, and one which splits opinion on Kaggle.\n\nSharing code gets upvotes on notebooks (sometimes, at least), and helps Kagglers progress through to higher tiers. It also gives people starting points for developing models. I would enter far fewer competitions if working code weren't provided as a starting point. Competitions that didn't allow shared models would be a lot less accessible. \n\nGenerally the best public submissions act as a de facto baseline - if you want a medal, you have to do quite a lot better than that in most cases (but occasionally copy-and-submit has achieved this, as others have pointed out with regard to BirdCLEF). This one is probably going to be different, as we don't really we much idea what a good model is from the public LB scores alone.\n\nCertainly not all the shared notebooks are meaningless or useless. Those shared in this competitions include some helpful descriptor-based cheminformatics models and also tokenization-based approaches. There are some interesting solutions to the problems of trying to  load or otherwise handle a very large dataset too.",
      "votes": null
    },
    {
      "id": "2903841",
      "postDate": "07/04/2024 03:19:16",
      "content": "<p>I agree. Sharing baseline code or a notebook that demonstrates a technique can help the community learn and improve. However, an inference notebook without any explanation seems more like a bid for upvotes.</p>",
      "rawMarkdown": "I agree. Sharing baseline code or a notebook that demonstrates a technique can help the community learn and improve. However, an inference notebook without any explanation seems more like a bid for upvotes.",
      "votes": null
    },
    {
      "id": "2905106",
      "postDate": "07/04/2024 17:41:23",
      "content": "<p>It's looking like a trend already. Another example ended several days ago </p>\n<p><a href=\"https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2\" target=\"_blank\">https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2</a>,</p>\n<p>where all (almost) silver and bronze zones were reached by using a public solution. The positive reason is that it seems like the quality of public notebooks has risen. On the other side, Kaggle is okay with that since such things attract newcomers.<br>\nConcerning LB pollution, even general Discussions Rankings are being corrupted because of the new Accomplishments forum, where folks are farming medals for nothing. Furthermore, notebook votes can be quietly collected from fake accounts. When can it be corrected? Simply speaking, welcome to modern Kaggle.</p>",
      "rawMarkdown": "It's looking like a trend already. Another example ended several days ago \n\nhttps://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2,\n\nwhere all (almost) silver and bronze zones were reached by using a public solution. The positive reason is that it seems like the quality of public notebooks has risen. On the other side, Kaggle is okay with that since such things attract newcomers.\nConcerning LB pollution, even general Discussions Rankings are being corrupted because of the new Accomplishments forum, where folks are farming medals for nothing. Furthermore, notebook votes can be quietly collected from fake accounts. When can it be corrected? Simply speaking, welcome to modern Kaggle.",
      "votes": null
    },
    {
      "id": "2909377",
      "postDate": "07/06/2024 22:08:35",
      "content": "<p>Perhaps the solution could be to set a cut off on the score that the largest number of users reached (mode) and award medals above that score?</p>",
      "rawMarkdown": "Perhaps the solution could be to set a cut off on the score that the largest number of users reached (mode) and award medals above that score?",
      "votes": null
    },
    {
      "id": "2909399",
      "postDate": "07/06/2024 23:49:45",
      "content": "<p>It could be unfair for those who do not just fork from the public notebook but also get the same cutoff score. </p>",
      "rawMarkdown": "It could be unfair for those who do not just fork from the public notebook but also get the same cutoff score.",
      "votes": null
    },
    {
      "id": "2911151",
      "postDate": "07/08/2024 06:01:34",
      "content": "<p>I agree. There is zero point in sharing some csv files to get certain scores.  You can't learn anything.  </p>",
      "rawMarkdown": "I agree. There is zero point in sharing some csv files to get certain scores.  You can't learn anything.",
      "votes": null
    },
    {
      "id": "2912512",
      "postDate": "07/09/2024 00:30:46",
      "content": "<p>So here we go again. Just a regular win of public notebook in this game as well.</p>",
      "rawMarkdown": "So here we go again. Just a regular win of public notebook in this game as well.",
      "votes": null
    },
    {
      "id": "2913121",
      "postDate": "07/09/2024 09:09:21",
      "content": "<p>A single submission to get a good generalized model is highly unlikely so it can be assumed that this solution is a replica of a public notebook. The medals should be assigned on the basis of the contributions made towards solving the problem. I understand its a hard problem but small steps can be taken to make the medal distribution more fairer. I don't have all the answers but a discussion can definitely help. Maybe a Kaggle competition to decide how to assign medals. :D</p>",
      "rawMarkdown": "A single submission to get a good generalized model is highly unlikely so it can be assumed that this solution is a replica of a public notebook. The medals should be assigned on the basis of the contributions made towards solving the problem. I understand its a hard problem but small steps can be taken to make the medal distribution more fairer. I don't have all the answers but a discussion can definitely help. Maybe a Kaggle competition to decide how to assign medals. :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2900351,
      "author_name": "masasato1999",
      "author_url": "",
      "post_date": "07/02/2024 08:11:14",
      "content": "<p>This is common in competitions that are not code competitions. In the open problems competition 6 months ago, the  lb/private index were shared and the notebooks were shared with manual lb overfitting (e.g. id1 +=0.2,id 2 -=0.1, etc.)<br>\nYou can consider teaming up by looking at the public lb, the person's lb, and the number of times the person has submitted.</p>\n<p>However, this is only the ranking of the lb, so your efforts will not be in vain depending on the result of the private data.<br>\nGood luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2901057,
          "author_name": "ganglii",
          "author_url": "",
          "post_date": "07/02/2024 16:09:42",
          "content": "<p>Thanks for your suggestions about choosing teammates. I got a lesson this time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2900354,
      "author_name": "jbomitchell",
      "author_url": "",
      "post_date": "07/02/2024 08:13:58",
      "content": "<p><a href=\"https://www.kaggle.com/ganglii\" target=\"_blank\">@ganglii</a> Do remember that it is not the public LB ranking that decides the results of the competition. As discussed in multiple places in this forum, the private LB data are significantly different and a large shake-up is anticipated in the final private LB ranking that decides prizes and medals. Getting a high public ranking guarantees nothing.</p>\n<p>The normal ethical custom and practice of Kaggle competitions is that public notebooks will inevitably get submitted multiple times, both as people's developments of them and in the original form. It is fine to wish that the rules were different and to argue that they should be changed, but such a public spat is unfair on your easily identifiable teammate who has acted within the regulations of the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2901016,
          "author_name": "ganglii",
          "author_url": "",
          "post_date": "07/02/2024 15:54:22",
          "content": "<p>If this kind of action is allowable and advocated in this community, I think exposure should not matter as people in the community understand him. And I didn't mean to blame someone who ran public notebooks. Instead, I think it is the action of sharing this kind of meaningless notebooks which only submits some uploaded prediction files that destroys the competition community. These notebooks don't bring any knowledge for other competitors to learn but pollute the whole lb. I didn't see any benefits of this kind of sharing.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2901185,
              "author_name": "jbomitchell",
              "author_url": "",
              "post_date": "07/02/2024 17:09:13",
              "content": "<p><a href=\"https://www.kaggle.com/ganglii\" target=\"_blank\">@ganglii</a>  This is a question with good arguments on both sides, and one which splits opinion on Kaggle.</p>\n<p>Sharing code gets upvotes on notebooks (sometimes, at least), and helps Kagglers progress through to higher tiers. It also gives people starting points for developing models. I would enter far fewer competitions if working code weren't provided as a starting point. Competitions that didn't allow shared models would be a lot less accessible. </p>\n<p>Generally the best public submissions act as a de facto baseline - if you want a medal, you have to do quite a lot better than that in most cases (but occasionally copy-and-submit has achieved this, as others have pointed out with regard to BirdCLEF). This one is probably going to be different, as we don't really we much idea what a good model is from the public LB scores alone.</p>\n<p>Certainly not all the shared notebooks are meaningless or useless. Those shared in this competitions include some helpful descriptor-based cheminformatics models and also tokenization-based approaches. There are some interesting solutions to the problems of trying to  load or otherwise handle a very large dataset too. </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2903841,
              "author_name": "faithk7u",
              "author_url": "",
              "post_date": "07/04/2024 03:19:16",
              "content": "<p>I agree. Sharing baseline code or a notebook that demonstrates a technique can help the community learn and improve. However, an inference notebook without any explanation seems more like a bid for upvotes.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2900586,
      "author_name": "ogurtsov",
      "author_url": "",
      "post_date": "07/02/2024 11:33:20",
      "content": "<p>If you are so disappointed by this competition, what do you say about <a href=\"https://www.kaggle.com/competitions/birdclef-2024\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024</a>?<br>\nPlaces 32-245 are public notebook submits.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2901045,
          "author_name": "ganglii",
          "author_url": "",
          "post_date": "07/02/2024 16:06:21",
          "content": "<p>If the public notebooks share real model training or data processing techniques, they are nice notebooks as they bring knowledge to others. If the public notebooks only teach people how to submit other uploaded prediction files, they are pure pollution to the lb. There are around 1000 teams in the competition you mentioned, I believe teams that just submitted shared prediction files don't even deserve a ranking of 250, top 25%. And I think the pollution should be attributed to someone who share these prediction files.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2905106,
          "author_name": "yekenot",
          "author_url": "",
          "post_date": "07/04/2024 17:41:23",
          "content": "<p>It's looking like a trend already. Another example ended several days ago </p>\n<p><a href=\"https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2\" target=\"_blank\">https://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2</a>,</p>\n<p>where all (almost) silver and bronze zones were reached by using a public solution. The positive reason is that it seems like the quality of public notebooks has risen. On the other side, Kaggle is okay with that since such things attract newcomers.<br>\nConcerning LB pollution, even general Discussions Rankings are being corrupted because of the new Accomplishments forum, where folks are farming medals for nothing. Furthermore, notebook votes can be quietly collected from fake accounts. When can it be corrected? Simply speaking, welcome to modern Kaggle.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2912512,
              "author_name": "yekenot",
              "author_url": "",
              "post_date": "07/09/2024 00:30:46",
              "content": "<p>So here we go again. Just a regular win of public notebook in this game as well.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2909377,
      "author_name": "zbigniewkirchner",
      "author_url": "",
      "post_date": "07/06/2024 22:08:35",
      "content": "<p>Perhaps the solution could be to set a cut off on the score that the largest number of users reached (mode) and award medals above that score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2909399,
          "author_name": "faithk7u",
          "author_url": "",
          "post_date": "07/06/2024 23:49:45",
          "content": "<p>It could be unfair for those who do not just fork from the public notebook but also get the same cutoff score. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2911151,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "07/08/2024 06:01:34",
      "content": "<p>I agree. There is zero point in sharing some csv files to get certain scores.  You can't learn anything.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2913121,
      "author_name": "ishgirwan",
      "author_url": "",
      "post_date": "07/09/2024 09:09:21",
      "content": "<p>A single submission to get a good generalized model is highly unlikely so it can be assumed that this solution is a replica of a public notebook. The medals should be assigned on the basis of the contributions made towards solving the problem. I understand its a hard problem but small steps can be taken to make the medal distribution more fairer. I don't have all the answers but a discussion can definitely help. Maybe a Kaggle competition to decide how to assign medals. :D</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2899878": "I'm a newbie to the Kaggle competition and disappointed with the experience. This morning, I reached out to several teams to ask for collaboration. After accepting one guy's invitation, I found that, in the submission history, the guy didn't have any original submission files. He just reran some shared notebooks that ensembles some uploaded prediction files by others, such as \"https://www.kaggle.com/code/andrey67/leash-bio-blender-0-440/notebook\", \"https://www.kaggle.com/code/saberghaderi/leash-bio-4149\". It achieves a very high LB score like 0.448, which destroys all the efforts most competitors have put into the competition.  When you look at the leaderboard carefully, especially with score 0.448, 0.433, 0.432, you will find a couple of teams reach such a high score with only one or two submission entries. It's possible that they also just ran the same shared notebooks to get the same score. Imagine, a guy, who didn't conduct any experiments for the competition, gets the top 10% ranking. How RIDICULOUS it is!\n\nAlso, I believe the one who shares the ensemble notebook should also be responsible for the pollution. If someone really shares some ensemble techniques, please use dummy prediction files. That would be helpful and thankful for all the competitors.\n\nAnyway, BROKEN leaderboard! DISAPPOINTING competition Experience!",
    "2900351": "This is common in competitions that are not code competitions. In the open problems competition 6 months ago, the  lb/private index were shared and the notebooks were shared with manual lb overfitting (e.g. id1 +=0.2,id 2 -=0.1, etc.)\nYou can consider teaming up by looking at the public lb, the person's lb, and the number of times the person has submitted.\n\nHowever, this is only the ranking of the lb, so your efforts will not be in vain depending on the result of the private data.\nGood luck!",
    "2900354": "ganglii Do remember that it is not the public LB ranking that decides the results of the competition. As discussed in multiple places in this forum, the private LB data are significantly different and a large shake-up is anticipated in the final private LB ranking that decides prizes and medals. Getting a high public ranking guarantees nothing.\n\nThe normal ethical custom and practice of Kaggle competitions is that public notebooks will inevitably get submitted multiple times, both as people's developments of them and in the original form. It is fine to wish that the rules were different and to argue that they should be changed, but such a public spat is unfair on your easily identifiable teammate who has acted within the regulations of the competition.",
    "2900586": "If you are so disappointed by this competition, what do you say about https://www.kaggle.com/competitions/birdclef-2024?\nPlaces 32-245 are public notebook submits.",
    "2901016": "If this kind of action is allowable and advocated in this community, I think exposure should not matter as people in the community understand him. And I didn't mean to blame someone who ran public notebooks. Instead, I think it is the action of sharing this kind of meaningless notebooks which only submits some uploaded prediction files that destroys the competition community. These notebooks don't bring any knowledge for other competitors to learn but pollute the whole lb. I didn't see any benefits of this kind of sharing.",
    "2901045": "If the public notebooks share real model training or data processing techniques, they are nice notebooks as they bring knowledge to others. If the public notebooks only teach people how to submit other uploaded prediction files, they are pure pollution to the lb. There are around 1000 teams in the competition you mentioned, I believe teams that just submitted shared prediction files don't even deserve a ranking of 250, top 25%. And I think the pollution should be attributed to someone who share these prediction files.",
    "2901057": "Thanks for your suggestions about choosing teammates. I got a lesson this time.",
    "2901185": "ganglii  This is a question with good arguments on both sides, and one which splits opinion on Kaggle.\n\nSharing code gets upvotes on notebooks (sometimes, at least), and helps Kagglers progress through to higher tiers. It also gives people starting points for developing models. I would enter far fewer competitions if working code weren't provided as a starting point. Competitions that didn't allow shared models would be a lot less accessible. \n\nGenerally the best public submissions act as a de facto baseline - if you want a medal, you have to do quite a lot better than that in most cases (but occasionally copy-and-submit has achieved this, as others have pointed out with regard to BirdCLEF). This one is probably going to be different, as we don't really we much idea what a good model is from the public LB scores alone.\n\nCertainly not all the shared notebooks are meaningless or useless. Those shared in this competitions include some helpful descriptor-based cheminformatics models and also tokenization-based approaches. There are some interesting solutions to the problems of trying to  load or otherwise handle a very large dataset too.",
    "2903841": "I agree. Sharing baseline code or a notebook that demonstrates a technique can help the community learn and improve. However, an inference notebook without any explanation seems more like a bid for upvotes.",
    "2905106": "It's looking like a trend already. Another example ended several days ago \n\nhttps://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2,\n\nwhere all (almost) silver and bronze zones were reached by using a public solution. The positive reason is that it seems like the quality of public notebooks has risen. On the other side, Kaggle is okay with that since such things attract newcomers.\nConcerning LB pollution, even general Discussions Rankings are being corrupted because of the new Accomplishments forum, where folks are farming medals for nothing. Furthermore, notebook votes can be quietly collected from fake accounts. When can it be corrected? Simply speaking, welcome to modern Kaggle.",
    "2909377": "Perhaps the solution could be to set a cut off on the score that the largest number of users reached (mode) and award medals above that score?",
    "2909399": "It could be unfair for those who do not just fork from the public notebook but also get the same cutoff score.",
    "2911151": "I agree. There is zero point in sharing some csv files to get certain scores.  You can't learn anything.",
    "2912512": "So here we go again. Just a regular win of public notebook in this game as well.",
    "2913121": "A single submission to get a good generalized model is highly unlikely so it can be assumed that this solution is a replica of a public notebook. The medals should be assigned on the basis of the contributions made towards solving the problem. I understand its a hard problem but small steps can be taken to make the medal distribution more fairer. I don't have all the answers but a discussion can definitely help. Maybe a Kaggle competition to decide how to assign medals. :D"
  },
  "source": "meta"
}