{
  "id": 277787,
  "title": "Shake-up is approaching, I can already feel it",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/277787",
  "author_name": "",
  "post_date": "2021-10-11T08:14:31.085147Z",
  "votes": 24,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Hi to all!<br>\nGiven all the features of the dataset and this competition in general, it was clear to me that the shake-up will definitely be. As of today, I understand that my place in the public list does not matter and I will definitely fall below, but I have no idea how much.<br>\nIn this competition I wanted to try several new methods and I have already tried them, so for me it is already successful, and what place I take does not matter.<br>\nI have already seen similar competitions, so I look forward to the results with fun.</p>\n<p>In what mood are you preparing for the end of this competition?<br>\nGood luck to all!</p>\n<p>P.S. I have tried most of the methods known to me to reduce overfiting of my models. But still I do not trust my results and believe that my model is bad, the only question is how bad it is). </p>",
  "messages": [
    {
      "id": "1541153",
      "postDate": "10/11/2021 08:14:31",
      "content": "<p>Hi to all!<br>\nGiven all the features of the dataset and this competition in general, it was clear to me that the shake-up will definitely be. As of today, I understand that my place in the public list does not matter and I will definitely fall below, but I have no idea how much.<br>\nIn this competition I wanted to try several new methods and I have already tried them, so for me it is already successful, and what place I take does not matter.<br>\nI have already seen similar competitions, so I look forward to the results with fun.</p>\n<p>In what mood are you preparing for the end of this competition?<br>\nGood luck to all!</p>\n<p>P.S. I have tried most of the methods known to me to reduce overfiting of my models. But still I do not trust my results and believe that my model is bad, the only question is how bad it is). </p>",
      "rawMarkdown": "Hi to all!\nGiven all the features of the dataset and this competition in general, it was clear to me that the shake-up will definitely be. As of today, I understand that my place in the public list does not matter and I will definitely fall below, but I have no idea how much.\nIn this competition I wanted to try several new methods and I have already tried them, so for me it is already successful, and what place I take does not matter.\nI have already seen similar competitions, so I look forward to the results with fun.\n\nIn what mood are you preparing for the end of this competition?\nGood luck to all!\n\nP.S. I have tried most of the methods known to me to reduce overfiting of my models. But still I do not trust my results and believe that my model is bad, the only question is how bad it is).",
      "votes": null
    },
    {
      "id": "1541208",
      "postDate": "10/11/2021 09:27:49",
      "content": "<p>What is really daunting is the fact that this competition is unfair.<br>\nThe rules say:<br>\n\"As a participant in the Competition, you agree … 3) not to attempt to probe the test set's labels\"</p>\n<p>Clearly, there are many teams that broke this rule so that they have access to more labeled data. Perhaps the competition hosts will verify the solutions of the winning teams - but how about the rest? What value does this rule have if it cannot be verified? If the hosts are unable to verify at least all the medal zones, this dummy constraint should be removed. Either respect the rules, or just remove them. These competitions require really a lot of time and finding out in the end that either everyone cheated except you, or that you are disqualified because everyone cheated and eventually you did it as well, is really demotivating. I didn't see any official statement of the competition hosts on that…<br>\nWhat should we expect? Is it better to do probing, or not?<br>\n<a href=\"https://www.kaggle.com/sbakas\" target=\"_blank\">@sbakas</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/ujjwalbaid\" target=\"_blank\">@ujjwalbaid</a> <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> </p>",
      "rawMarkdown": "What is really daunting is the fact that this competition is unfair.\nThe rules say:\n\"As a participant in the Competition, you agree ... 3) not to attempt to probe the test set's labels\"\n\nClearly, there are many teams that broke this rule so that they have access to more labeled data. Perhaps the competition hosts will verify the solutions of the winning teams - but how about the rest? What value does this rule have if it cannot be verified? If the hosts are unable to verify at least all the medal zones, this dummy constraint should be removed. Either respect the rules, or just remove them. These competitions require really a lot of time and finding out in the end that either everyone cheated except you, or that you are disqualified because everyone cheated and eventually you did it as well, is really demotivating. I didn't see any official statement of the competition hosts on that...\nWhat should we expect? Is it better to do probing, or not?\n@sbakas @juliaelliott @anthracene @ujjwalbaid @cdcarr",
      "votes": null
    },
    {
      "id": "1541359",
      "postDate": "10/11/2021 12:52:46",
      "content": "<p>It has already been addressed by the Kaggle staff in this thread -&gt; <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/256706\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/256706</a>. </p>\n<p>They will not take any action against hand labelers.</p>",
      "rawMarkdown": "It has already been addressed by the Kaggle staff in this thread -> https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/256706. \n\nThey will not take any action against hand labelers.",
      "votes": null
    },
    {
      "id": "1541366",
      "postDate": "10/11/2021 13:03:32",
      "content": "<p>From the experience of previous competitions, it seems to me that violators are usually disqualified after the deadline. <br>\nIf I am wrong, please write</p>",
      "rawMarkdown": "From the experience of previous competitions, it seems to me that violators are usually disqualified after the deadline. \nIf I am wrong, please write",
      "votes": null
    },
    {
      "id": "1541408",
      "postDate": "10/11/2021 13:37:01",
      "content": "<p>Do I understand correctly that your concern is related to the fact that some people may use the labeled public leaderboard to train their models?</p>",
      "rawMarkdown": "Do I understand correctly that your concern is related to the fact that some people may use the labeled public leaderboard to train their models?",
      "votes": null
    },
    {
      "id": "1541442",
      "postDate": "10/11/2021 13:57:44",
      "content": "<p>I have two concerns:</p>\n<ol>\n<li>Some people might use them as validation datasets locally, so that they can use all the training data and validate the models locally every epoch (or every n steps) instead of submitting them max. 5 times per day,</li>\n<li>They can also incorporate the test data and train their models with larger CV splits</li>\n</ol>\n<p>And it is not that I am saying this is wrong. I am just saying that if this is permitted, it shouldn't be officially forbidden.</p>\n<p>I have just opened the thread mentioned by <a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> and saw the comment \"We are not planning to take any action at this point\". As I understand, it means that the rule \"As a participant in the Competition, you agree … 3) not to attempt to probe the test set's labels\" does not longer apply. Given this fact, my team (and probably many more) is behind as we were following the official rules. I missed this laconic comment from the hosts. If an official rule does not apply, what's the point of keeping it and only confusing everyone?</p>",
      "rawMarkdown": "I have two concerns:\n1. Some people might use them as validation datasets locally, so that they can use all the training data and validate the models locally every epoch (or every n steps) instead of submitting them max. 5 times per day,\n2. They can also incorporate the test data and train their models with larger CV splits\n\nAnd it is not that I am saying this is wrong. I am just saying that if this is permitted, it shouldn't be officially forbidden.\n\nI have just opened the thread mentioned by @davidbroberts and saw the comment \"We are not planning to take any action at this point\". As I understand, it means that the rule \"As a participant in the Competition, you agree … 3) not to attempt to probe the test set's labels\" does not longer apply. Given this fact, my team (and probably many more) is behind as we were following the official rules. I missed this laconic comment from the hosts. If an official rule does not apply, what's the point of keeping it and only confusing everyone?",
      "votes": null
    },
    {
      "id": "1541448",
      "postDate": "10/11/2021 14:04:38",
      "content": "<p>The comment from the hosts in the mentioned thread: \"We are not planning to take any action at this point\"</p>",
      "rawMarkdown": "The comment from the hosts in the mentioned thread: \"We are not planning to take any action at this point\"",
      "votes": null
    },
    {
      "id": "1541463",
      "postDate": "10/11/2021 14:19:28",
      "content": "<p>Actually, I have the same concern as you <a href=\"https://www.kaggle.com/mikecho\" target=\"_blank\">@mikecho</a> . It is discouraging to learn that some people tried to probe LB and due to this some of us have no idea of our actual position in the LB and if no actions were taken for probing LB, then there is no transparency and fairness in this competition. </p>",
      "rawMarkdown": "Actually, I have the same concern as you @mikecho . It is discouraging to learn that some people tried to probe LB and due to this some of us have no idea of our actual position in the LB and if no actions were taken for probing LB, then there is no transparency and fairness in this competition.",
      "votes": null
    },
    {
      "id": "1541566",
      "postDate": "10/11/2021 16:12:44",
      "content": "<p>Between this and the general consensus that this task is not really solvable my motivation to work on it has collapsed.  I wasn't hoping to win, but I did want to genuinely compete, now I just want to move on to a new project 🙁</p>",
      "rawMarkdown": "Between this and the general consensus that this task is not really solvable my motivation to work on it has collapsed.  I wasn't hoping to win, but I did want to genuinely compete, now I just want to move on to a new project 🙁",
      "votes": null
    },
    {
      "id": "1541601",
      "postDate": "10/11/2021 16:58:28",
      "content": "<p>I think that the organizers will not take any action against hand labelers as well as against those who publish notebooks with a high rating at the end of the competition. <br>\nIn my opinion, we must remember that in order to hold a competition, the organizers (customers) must invest a large amount of money. And they want the best models in return. <br>\nAnd how can you increase competition between participants to get better models? <br>\nGive all participants good models, let some participants \"cheat a little\" still in the top 10 will be the best models without cheating. <br>\nFrom this point of view, I understand the organizers, but this is just my opinion.</p>",
      "rawMarkdown": "I think that the organizers will not take any action against hand labelers as well as against those who publish notebooks with a high rating at the end of the competition. \nIn my opinion, we must remember that in order to hold a competition, the organizers (customers) must invest a large amount of money. And they want the best models in return. \nAnd how can you increase competition between participants to get better models? \nGive all participants good models, let some participants \"cheat a little\" still in the top 10 will be the best models without cheating. \nFrom this point of view, I understand the organizers, but this is just my opinion.",
      "votes": null
    },
    {
      "id": "1541643",
      "postDate": "10/11/2021 17:59:48",
      "content": "<p>The solution is very simple - remove rules that are not to be followed. Why do they even exist? If they made it clear from the very beginning, they might have obtained already even better models because everyone would be training models on more data. It is in common interest of everyone.</p>",
      "rawMarkdown": "The solution is very simple - remove rules that are not to be followed. Why do they even exist? If they made it clear from the very beginning, they might have obtained already even better models because everyone would be training models on more data. It is in common interest of everyone.",
      "votes": null
    },
    {
      "id": "1541647",
      "postDate": "10/11/2021 18:04:57",
      "content": "<p>Is it realistic to get a score greater than 0.7 without hand labeling in public?</p>",
      "rawMarkdown": "Is it realistic to get a score greater than 0.7 without hand labeling in public?",
      "votes": null
    },
    {
      "id": "1541673",
      "postDate": "10/11/2021 18:21:03",
      "content": "<p>well, there are public kernels with even 0.73… But what is the value of them - we will see in a few days :)<br>\nMy solution did not use hand labeling but the validation score was below 0.7, so I expect it to be lower in the private LB.</p>",
      "rawMarkdown": "well, there are public kernels with even 0.73... But what is the value of them - we will see in a few days :)\nMy solution did not use hand labeling but the validation score was below 0.7, so I expect it to be lower in the private LB.",
      "votes": null
    },
    {
      "id": "1542003",
      "postDate": "10/12/2021 04:21:04",
      "content": "<p>Private LB will be lottery (super shake).</p>",
      "rawMarkdown": "Private LB will be lottery (super shake).",
      "votes": null
    },
    {
      "id": "1542475",
      "postDate": "10/12/2021 14:44:50",
      "content": "<p><strong>The 0.73 notebook shouldn't be a concern</strong>. If you look close at it,  this notebook blends the predictions of public LB ids and <strong>predict 0.5 for the ids of the private test dataset</strong>. After checking the ids, i noticed that there is <strong>roughly</strong> 1000 scans, 600 scans are given in train, 87 for the public LB and <strong>the remaing 400 will be for the private LB that determine the LB</strong>.  <strong>SO YOU SHOULD'NT WORRY</strong>. I predict that  <strong>shake up will be mad</strong>. . But i also don't think it is lottery at all, since I have some CV aligned with public LB before the coming of perfect score LB and public csv blends.</p>",
      "rawMarkdown": "**The 0.73 notebook shouldn't be a concern**. If you look close at it,  this notebook blends the predictions of public LB ids and **predict 0.5 for the ids of the private test dataset**. After checking the ids, i noticed that there is **roughly** 1000 scans, 600 scans are given in train, 87 for the public LB and **the remaing 400 will be for the private LB that determine the LB**.  **SO YOU SHOULD'NT WORRY**. I predict that  **shake up will be mad**. . But i also don't think it is lottery at all, since I have some CV aligned with public LB before the coming of perfect score LB and public csv blends.",
      "votes": null
    },
    {
      "id": "1542562",
      "postDate": "10/12/2021 16:42:46",
      "content": "<p>Let me tell you a story: Last year, my companion and I competed in the Google Landmark Recognition 2020 competition. We selected the wrong notebooks and dropped to 415 positions and finished in 550th place out of 736 teams. And with this result, I consider this competition one of the most successful for me because the knowledge gained in this competition allowed me to win many medals in subsequent competitions in various categories. So the new knowledge will stay with you and help you in the future, and people who did nothing and just copied the result will get nothing. I understand that it is unpleasant that you work hard to get the result, and the person who cheated gets better results than you, but I would not be limited to a place in the rankings when analyzing my results in the competition. Good luck!</p>",
      "rawMarkdown": "Let me tell you a story: Last year, my companion and I competed in the Google Landmark Recognition 2020 competition. We selected the wrong notebooks and dropped to 415 positions and finished in 550th place out of 736 teams. And with this result, I consider this competition one of the most successful for me because the knowledge gained in this competition allowed me to win many medals in subsequent competitions in various categories. So the new knowledge will stay with you and help you in the future, and people who did nothing and just copied the result will get nothing. I understand that it is unpleasant that you work hard to get the result, and the person who cheated gets better results than you, but I would not be limited to a place in the rankings when analyzing my results in the competition. Good luck!",
      "votes": null
    },
    {
      "id": "1543413",
      "postDate": "10/13/2021 13:18:52",
      "content": "<p>what will hand labelling have benefit in private leaderboard , if the test data is already unseen?</p>",
      "rawMarkdown": "what will hand labelling have benefit in private leaderboard , if the test data is already unseen?",
      "votes": null
    },
    {
      "id": "1543521",
      "postDate": "10/13/2021 15:11:04",
      "content": "<p>somebody pointed it out already. it increases your training dataset and can hand you fine generalization margin for private LB.</p>",
      "rawMarkdown": "somebody pointed it out already. it increases your training dataset and can hand you fine generalization margin for private LB.",
      "votes": null
    },
    {
      "id": "1543761",
      "postDate": "10/13/2021 19:25:54",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> , Just wondering if your current Lb score is without any hand labeling? </p>",
      "rawMarkdown": "Hey @ulrich07 , Just wondering if your current Lb score is without any hand labeling?",
      "votes": null
    },
    {
      "id": "1543812",
      "postDate": "10/13/2021 20:50:31",
      "content": "<p>It is a mix of quick lb probing and another trick I can't reveal yet 😎.  After all, this trick can be costly if shake up is not with me ! But as you know we usually choose a stable and a risky submission. So this score is a gamble.</p>",
      "rawMarkdown": "It is a mix of quick lb probing and another trick I can't reveal yet 😎.  After all, this trick can be costly if shake up is not with me ! But as you know we usually choose a stable and a risky submission. So this score is a gamble.",
      "votes": null
    },
    {
      "id": "1544139",
      "postDate": "10/14/2021 05:20:55",
      "content": "<p>Any guess what will be the AUC of the winner?<br>\nI am guessing 0.6-0.7</p>",
      "rawMarkdown": "Any guess what will be the AUC of the winner?\nI am guessing 0.6-0.7",
      "votes": null
    },
    {
      "id": "1544147",
      "postDate": "10/14/2021 05:28:40",
      "content": "<p>I guess it’ll be around 0.65, but possibly 0.7~ with randomness.</p>",
      "rawMarkdown": "I guess it’ll be around 0.65, but possibly 0.7~ with randomness.",
      "votes": null
    },
    {
      "id": "1544619",
      "postDate": "10/14/2021 13:49:15",
      "content": "<p>It's really hard for models to predict MGMT based on MR images. The models get the level of 0.6-0.7 on validation sets, however get the level of 0.4-0.8 on public sets.😑 This is meaningless in clinical practice.</p>",
      "rawMarkdown": "It's really hard for models to predict MGMT based on MR images. The models get the level of 0.6-0.7 on validation sets, however get the level of 0.4-0.8 on public sets.😑 This is meaningless in clinical practice.",
      "votes": null
    },
    {
      "id": "1544841",
      "postDate": "10/14/2021 17:37:45",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3407946%2Fb674080680317a509436948ed942db23%2Fkaggle.jpg?generation=1569486219538676&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3407946%2Fb674080680317a509436948ed942db23%2Fkaggle.jpg?generation=1569486219538676&alt=media)",
      "votes": null
    },
    {
      "id": "1545018",
      "postDate": "10/14/2021 22:14:56",
      "content": "<p>yes, indeed, in notebooks that gain more than 0.7, csv files or their derivatives are used, obtained not only as a result of the work of models, but also by random enumeration of values.</p>",
      "rawMarkdown": "yes, indeed, in notebooks that gain more than 0.7, csv files or their derivatives are used, obtained not only as a result of the work of models, but also by random enumeration of values.",
      "votes": null
    },
    {
      "id": "1545361",
      "postDate": "10/15/2021 06:34:52",
      "content": "<p>I am also guessing the range [0.7, 0.8] with some luck.</p>",
      "rawMarkdown": "I am also guessing the range [0.7, 0.8] with some luck.",
      "votes": null
    },
    {
      "id": "1545874",
      "postDate": "10/15/2021 16:49:41",
      "content": "<p>first kaggle competition, but just pray</p>",
      "rawMarkdown": "first kaggle competition, but just pray",
      "votes": null
    },
    {
      "id": "1546017",
      "postDate": "10/15/2021 18:28:15",
      "content": "<p>Don't select a public notebook that doesn't infer private data.<br>\nThis alone has the possibility of you shaking up.<br>\nFor example, The 0.73 notebook.</p>",
      "rawMarkdown": "Don't select a public notebook that doesn't infer private data.\nThis alone has the possibility of you shaking up.\nFor example, The 0.73 notebook.",
      "votes": null
    },
    {
      "id": "1546035",
      "postDate": "10/15/2021 18:32:14",
      "content": "<p>첫 대회부터 고통받고 있습니다…</p>",
      "rawMarkdown": "첫 대회부터 고통받고 있습니다...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1541208,
      "author_name": "mikecho",
      "author_url": "",
      "post_date": "10/11/2021 09:27:49",
      "content": "<p>What is really daunting is the fact that this competition is unfair.<br>\nThe rules say:<br>\n\"As a participant in the Competition, you agree … 3) not to attempt to probe the test set's labels\"</p>\n<p>Clearly, there are many teams that broke this rule so that they have access to more labeled data. Perhaps the competition hosts will verify the solutions of the winning teams - but how about the rest? What value does this rule have if it cannot be verified? If the hosts are unable to verify at least all the medal zones, this dummy constraint should be removed. Either respect the rules, or just remove them. These competitions require really a lot of time and finding out in the end that either everyone cheated except you, or that you are disqualified because everyone cheated and eventually you did it as well, is really demotivating. I didn't see any official statement of the competition hosts on that…<br>\nWhat should we expect? Is it better to do probing, or not?<br>\n<a href=\"https://www.kaggle.com/sbakas\" target=\"_blank\">@sbakas</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/ujjwalbaid\" target=\"_blank\">@ujjwalbaid</a> <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1541408,
          "author_name": "greylord1996",
          "author_url": "",
          "post_date": "10/11/2021 13:37:01",
          "content": "<p>Do I understand correctly that your concern is related to the fact that some people may use the labeled public leaderboard to train their models?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541442,
          "author_name": "mikecho",
          "author_url": "",
          "post_date": "10/11/2021 13:57:44",
          "content": "<p>I have two concerns:</p>\n<ol>\n<li>Some people might use them as validation datasets locally, so that they can use all the training data and validate the models locally every epoch (or every n steps) instead of submitting them max. 5 times per day,</li>\n<li>They can also incorporate the test data and train their models with larger CV splits</li>\n</ol>\n<p>And it is not that I am saying this is wrong. I am just saying that if this is permitted, it shouldn't be officially forbidden.</p>\n<p>I have just opened the thread mentioned by <a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> and saw the comment \"We are not planning to take any action at this point\". As I understand, it means that the rule \"As a participant in the Competition, you agree … 3) not to attempt to probe the test set's labels\" does not longer apply. Given this fact, my team (and probably many more) is behind as we were following the official rules. I missed this laconic comment from the hosts. If an official rule does not apply, what's the point of keeping it and only confusing everyone?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541463,
          "author_name": "wfarzana",
          "author_url": "",
          "post_date": "10/11/2021 14:19:28",
          "content": "<p>Actually, I have the same concern as you <a href=\"https://www.kaggle.com/mikecho\" target=\"_blank\">@mikecho</a> . It is discouraging to learn that some people tried to probe LB and due to this some of us have no idea of our actual position in the LB and if no actions were taken for probing LB, then there is no transparency and fairness in this competition. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541566,
          "author_name": "maxbaugh",
          "author_url": "",
          "post_date": "10/11/2021 16:12:44",
          "content": "<p>Between this and the general consensus that this task is not really solvable my motivation to work on it has collapsed.  I wasn't hoping to win, but I did want to genuinely compete, now I just want to move on to a new project 🙁</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1542562,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "10/12/2021 16:42:46",
          "content": "<p>Let me tell you a story: Last year, my companion and I competed in the Google Landmark Recognition 2020 competition. We selected the wrong notebooks and dropped to 415 positions and finished in 550th place out of 736 teams. And with this result, I consider this competition one of the most successful for me because the knowledge gained in this competition allowed me to win many medals in subsequent competitions in various categories. So the new knowledge will stay with you and help you in the future, and people who did nothing and just copied the result will get nothing. I understand that it is unpleasant that you work hard to get the result, and the person who cheated gets better results than you, but I would not be limited to a place in the rankings when analyzing my results in the competition. Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1541359,
      "author_name": "davidbroberts",
      "author_url": "",
      "post_date": "10/11/2021 12:52:46",
      "content": "<p>It has already been addressed by the Kaggle staff in this thread -&gt; <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/256706\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/256706</a>. </p>\n<p>They will not take any action against hand labelers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1541366,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "10/11/2021 13:03:32",
          "content": "<p>From the experience of previous competitions, it seems to me that violators are usually disqualified after the deadline. <br>\nIf I am wrong, please write</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541448,
          "author_name": "mikecho",
          "author_url": "",
          "post_date": "10/11/2021 14:04:38",
          "content": "<p>The comment from the hosts in the mentioned thread: \"We are not planning to take any action at this point\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541601,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "10/11/2021 16:58:28",
          "content": "<p>I think that the organizers will not take any action against hand labelers as well as against those who publish notebooks with a high rating at the end of the competition. <br>\nIn my opinion, we must remember that in order to hold a competition, the organizers (customers) must invest a large amount of money. And they want the best models in return. <br>\nAnd how can you increase competition between participants to get better models? <br>\nGive all participants good models, let some participants \"cheat a little\" still in the top 10 will be the best models without cheating. <br>\nFrom this point of view, I understand the organizers, but this is just my opinion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541643,
          "author_name": "mikecho",
          "author_url": "",
          "post_date": "10/11/2021 17:59:48",
          "content": "<p>The solution is very simple - remove rules that are not to be followed. Why do they even exist? If they made it clear from the very beginning, they might have obtained already even better models because everyone would be training models on more data. It is in common interest of everyone.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541647,
          "author_name": "greylord1996",
          "author_url": "",
          "post_date": "10/11/2021 18:04:57",
          "content": "<p>Is it realistic to get a score greater than 0.7 without hand labeling in public?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1541673,
          "author_name": "mikecho",
          "author_url": "",
          "post_date": "10/11/2021 18:21:03",
          "content": "<p>well, there are public kernels with even 0.73… But what is the value of them - we will see in a few days :)<br>\nMy solution did not use hand labeling but the validation score was below 0.7, so I expect it to be lower in the private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1542475,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "10/12/2021 14:44:50",
          "content": "<p><strong>The 0.73 notebook shouldn't be a concern</strong>. If you look close at it,  this notebook blends the predictions of public LB ids and <strong>predict 0.5 for the ids of the private test dataset</strong>. After checking the ids, i noticed that there is <strong>roughly</strong> 1000 scans, 600 scans are given in train, 87 for the public LB and <strong>the remaing 400 will be for the private LB that determine the LB</strong>.  <strong>SO YOU SHOULD'NT WORRY</strong>. I predict that  <strong>shake up will be mad</strong>. . But i also don't think it is lottery at all, since I have some CV aligned with public LB before the coming of perfect score LB and public csv blends.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1545018,
          "author_name": "zaakciiru",
          "author_url": "",
          "post_date": "10/14/2021 22:14:56",
          "content": "<p>yes, indeed, in notebooks that gain more than 0.7, csv files or their derivatives are used, obtained not only as a result of the work of models, but also by random enumeration of values.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1542003,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "10/12/2021 04:21:04",
      "content": "<p>Private LB will be lottery (super shake).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1543413,
      "author_name": "aniarya",
      "author_url": "",
      "post_date": "10/13/2021 13:18:52",
      "content": "<p>what will hand labelling have benefit in private leaderboard , if the test data is already unseen?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1543521,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "10/13/2021 15:11:04",
          "content": "<p>somebody pointed it out already. it increases your training dataset and can hand you fine generalization margin for private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1543761,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "10/13/2021 19:25:54",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> , Just wondering if your current Lb score is without any hand labeling? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1543812,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "10/13/2021 20:50:31",
          "content": "<p>It is a mix of quick lb probing and another trick I can't reveal yet 😎.  After all, this trick can be costly if shake up is not with me ! But as you know we usually choose a stable and a risky submission. So this score is a gamble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1544139,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "10/14/2021 05:20:55",
      "content": "<p>Any guess what will be the AUC of the winner?<br>\nI am guessing 0.6-0.7</p>",
      "votes": null,
      "replies": [
        {
          "id": 1544147,
          "author_name": "drtausamaru",
          "author_url": "",
          "post_date": "10/14/2021 05:28:40",
          "content": "<p>I guess it’ll be around 0.65, but possibly 0.7~ with randomness.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1545361,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "10/15/2021 06:34:52",
          "content": "<p>I am also guessing the range [0.7, 0.8] with some luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1544619,
      "author_name": "zerota",
      "author_url": "",
      "post_date": "10/14/2021 13:49:15",
      "content": "<p>It's really hard for models to predict MGMT based on MR images. The models get the level of 0.6-0.7 on validation sets, however get the level of 0.4-0.8 on public sets.😑 This is meaningless in clinical practice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1544841,
      "author_name": "jolasa",
      "author_url": "",
      "post_date": "10/14/2021 17:37:45",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3407946%2Fb674080680317a509436948ed942db23%2Fkaggle.jpg?generation=1569486219538676&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1545874,
      "author_name": "justhungryman",
      "author_url": "",
      "post_date": "10/15/2021 16:49:41",
      "content": "<p>first kaggle competition, but just pray</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1546017,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "10/15/2021 18:28:15",
      "content": "<p>Don't select a public notebook that doesn't infer private data.<br>\nThis alone has the possibility of you shaking up.<br>\nFor example, The 0.73 notebook.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1546035,
          "author_name": "justhungryman",
          "author_url": "",
          "post_date": "10/15/2021 18:32:14",
          "content": "<p>첫 대회부터 고통받고 있습니다…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1541153": "Hi to all!\nGiven all the features of the dataset and this competition in general, it was clear to me that the shake-up will definitely be. As of today, I understand that my place in the public list does not matter and I will definitely fall below, but I have no idea how much.\nIn this competition I wanted to try several new methods and I have already tried them, so for me it is already successful, and what place I take does not matter.\nI have already seen similar competitions, so I look forward to the results with fun.\n\nIn what mood are you preparing for the end of this competition?\nGood luck to all!\n\nP.S. I have tried most of the methods known to me to reduce overfiting of my models. But still I do not trust my results and believe that my model is bad, the only question is how bad it is).",
    "1541208": "What is really daunting is the fact that this competition is unfair.\nThe rules say:\n\"As a participant in the Competition, you agree ... 3) not to attempt to probe the test set's labels\"\n\nClearly, there are many teams that broke this rule so that they have access to more labeled data. Perhaps the competition hosts will verify the solutions of the winning teams - but how about the rest? What value does this rule have if it cannot be verified? If the hosts are unable to verify at least all the medal zones, this dummy constraint should be removed. Either respect the rules, or just remove them. These competitions require really a lot of time and finding out in the end that either everyone cheated except you, or that you are disqualified because everyone cheated and eventually you did it as well, is really demotivating. I didn't see any official statement of the competition hosts on that...\nWhat should we expect? Is it better to do probing, or not?\n@sbakas @juliaelliott @anthracene @ujjwalbaid @cdcarr",
    "1541359": "It has already been addressed by the Kaggle staff in this thread -> https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/256706. \n\nThey will not take any action against hand labelers.",
    "1541366": "From the experience of previous competitions, it seems to me that violators are usually disqualified after the deadline. \nIf I am wrong, please write",
    "1541408": "Do I understand correctly that your concern is related to the fact that some people may use the labeled public leaderboard to train their models?",
    "1541442": "I have two concerns:\n1. Some people might use them as validation datasets locally, so that they can use all the training data and validate the models locally every epoch (or every n steps) instead of submitting them max. 5 times per day,\n2. They can also incorporate the test data and train their models with larger CV splits\n\nAnd it is not that I am saying this is wrong. I am just saying that if this is permitted, it shouldn't be officially forbidden.\n\nI have just opened the thread mentioned by @davidbroberts and saw the comment \"We are not planning to take any action at this point\". As I understand, it means that the rule \"As a participant in the Competition, you agree … 3) not to attempt to probe the test set's labels\" does not longer apply. Given this fact, my team (and probably many more) is behind as we were following the official rules. I missed this laconic comment from the hosts. If an official rule does not apply, what's the point of keeping it and only confusing everyone?",
    "1541448": "The comment from the hosts in the mentioned thread: \"We are not planning to take any action at this point\"",
    "1541463": "Actually, I have the same concern as you @mikecho . It is discouraging to learn that some people tried to probe LB and due to this some of us have no idea of our actual position in the LB and if no actions were taken for probing LB, then there is no transparency and fairness in this competition.",
    "1541566": "Between this and the general consensus that this task is not really solvable my motivation to work on it has collapsed.  I wasn't hoping to win, but I did want to genuinely compete, now I just want to move on to a new project 🙁",
    "1541601": "I think that the organizers will not take any action against hand labelers as well as against those who publish notebooks with a high rating at the end of the competition. \nIn my opinion, we must remember that in order to hold a competition, the organizers (customers) must invest a large amount of money. And they want the best models in return. \nAnd how can you increase competition between participants to get better models? \nGive all participants good models, let some participants \"cheat a little\" still in the top 10 will be the best models without cheating. \nFrom this point of view, I understand the organizers, but this is just my opinion.",
    "1541643": "The solution is very simple - remove rules that are not to be followed. Why do they even exist? If they made it clear from the very beginning, they might have obtained already even better models because everyone would be training models on more data. It is in common interest of everyone.",
    "1541647": "Is it realistic to get a score greater than 0.7 without hand labeling in public?",
    "1541673": "well, there are public kernels with even 0.73... But what is the value of them - we will see in a few days :)\nMy solution did not use hand labeling but the validation score was below 0.7, so I expect it to be lower in the private LB.",
    "1542003": "Private LB will be lottery (super shake).",
    "1542475": "**The 0.73 notebook shouldn't be a concern**. If you look close at it,  this notebook blends the predictions of public LB ids and **predict 0.5 for the ids of the private test dataset**. After checking the ids, i noticed that there is **roughly** 1000 scans, 600 scans are given in train, 87 for the public LB and **the remaing 400 will be for the private LB that determine the LB**.  **SO YOU SHOULD'NT WORRY**. I predict that  **shake up will be mad**. . But i also don't think it is lottery at all, since I have some CV aligned with public LB before the coming of perfect score LB and public csv blends.",
    "1542562": "Let me tell you a story: Last year, my companion and I competed in the Google Landmark Recognition 2020 competition. We selected the wrong notebooks and dropped to 415 positions and finished in 550th place out of 736 teams. And with this result, I consider this competition one of the most successful for me because the knowledge gained in this competition allowed me to win many medals in subsequent competitions in various categories. So the new knowledge will stay with you and help you in the future, and people who did nothing and just copied the result will get nothing. I understand that it is unpleasant that you work hard to get the result, and the person who cheated gets better results than you, but I would not be limited to a place in the rankings when analyzing my results in the competition. Good luck!",
    "1543413": "what will hand labelling have benefit in private leaderboard , if the test data is already unseen?",
    "1543521": "somebody pointed it out already. it increases your training dataset and can hand you fine generalization margin for private LB.",
    "1543761": "Hey @ulrich07 , Just wondering if your current Lb score is without any hand labeling?",
    "1543812": "It is a mix of quick lb probing and another trick I can't reveal yet 😎.  After all, this trick can be costly if shake up is not with me ! But as you know we usually choose a stable and a risky submission. So this score is a gamble.",
    "1544139": "Any guess what will be the AUC of the winner?\nI am guessing 0.6-0.7",
    "1544147": "I guess it’ll be around 0.65, but possibly 0.7~ with randomness.",
    "1544619": "It's really hard for models to predict MGMT based on MR images. The models get the level of 0.6-0.7 on validation sets, however get the level of 0.4-0.8 on public sets.😑 This is meaningless in clinical practice.",
    "1544841": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3407946%2Fb674080680317a509436948ed942db23%2Fkaggle.jpg?generation=1569486219538676&alt=media)",
    "1545018": "yes, indeed, in notebooks that gain more than 0.7, csv files or their derivatives are used, obtained not only as a result of the work of models, but also by random enumeration of values.",
    "1545361": "I am also guessing the range [0.7, 0.8] with some luck.",
    "1545874": "first kaggle competition, but just pray",
    "1546017": "Don't select a public notebook that doesn't infer private data.\nThis alone has the possibility of you shaking up.\nFor example, The 0.73 notebook.",
    "1546035": "첫 대회부터 고통받고 있습니다..."
  },
  "source": "meta"
}