{
  "id": 308329,
  "title": "A Feedback: About Accuracy of This Competition’s Labels",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/308329",
  "author_name": "Bilzard",
  "post_date": "2022-02-18T07:08:04.134000",
  "votes": 33,
  "comment_count": 44,
  "views": 0,
  "content": "<p>To the Kaggle staff, the<br>\nSCIRO team.</p>\n<p>Thank you very much for organizing this competition.</p>\n<p>As a participant, I would like to give you some feedback about the accuracy of the labels in the competition data.</p>\n<p>As some of the other participants may have noticed, the training labels were inconsistent and there were many missing labels. My solution[1] showed an increase in the LB score by correcting the training labels, and the 3rd-place solution[2] suggests that there may be systematic errors in the fitness of the labels on the public and private datasets.</p>\n<p>This suggests that the evaluation labels may have been as inaccurate as the training labels. (Of course, we participants cannot see the test data set, so we can only speculate.)</p>\n<p>I’m not requesting re-evaluating this competitions result. Even if some of the other participants argue that we should, we shouldn’t. That would amount to an ex post facto application of the law.</p>\n<p>Instead, I would like to insist that the next time you hold a similar competition, I want you to pay full attention and cost to the accuracy of the labels.</p>\n<p>If one is a skilled kaggler, he/she know how to train models with good generalization performance even when the training dataset is inaccurate.<br>\nHowever, as we know, when the evaluation dataset is inaccurate, it is inevitable that the winning model will be suboptimal.</p>\n<p>This is not only unfair to the participants of the competition, but also disadvantageous to the host who will operate the model, as the model that will eventually be operated will not perform well enough.</p>\n<p>This competition had a high prize money at stake. Therefore, the cost could have been spent on making the labels for evaluation (and preferably for training) more accurate.</p>\n<p>So please think about this.</p>\n<p>Finally, I hope that this discussion will be used for constructive purposes.</p>\n<h2>Summary of Opinions</h2>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343</a></p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307871\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307871</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307707\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307707</a></li>\n</ul>",
  "messages": [
    {
      "id": 1695483,
      "postDate": "2022-02-18T07:08:04.133Z",
      "content": "<p>To the Kaggle staff, the<br>\nSCIRO team.</p>\n<p>Thank you very much for organizing this competition.</p>\n<p>As a participant, I would like to give you some feedback about the accuracy of the labels in the competition data.</p>\n<p>As some of the other participants may have noticed, the training labels were inconsistent and there were many missing labels. My solution[1] showed an increase in the LB score by correcting the training labels, and the 3rd-place solution[2] suggests that there may be systematic errors in the fitness of the labels on the public and private datasets.</p>\n<p>This suggests that the evaluation labels may have been as inaccurate as the training labels. (Of course, we participants cannot see the test data set, so we can only speculate.)</p>\n<p>I’m not requesting re-evaluating this competitions result. Even if some of the other participants argue that we should, we shouldn’t. That would amount to an ex post facto application of the law.</p>\n<p>Instead, I would like to insist that the next time you hold a similar competition, I want you to pay full attention and cost to the accuracy of the labels.</p>\n<p>If one is a skilled kaggler, he/she know how to train models with good generalization performance even when the training dataset is inaccurate.<br>\nHowever, as we know, when the evaluation dataset is inaccurate, it is inevitable that the winning model will be suboptimal.</p>\n<p>This is not only unfair to the participants of the competition, but also disadvantageous to the host who will operate the model, as the model that will eventually be operated will not perform well enough.</p>\n<p>This competition had a high prize money at stake. Therefore, the cost could have been spent on making the labels for evaluation (and preferably for training) more accurate.</p>\n<p>So please think about this.</p>\n<p>Finally, I hope that this discussion will be used for constructive purposes.</p>\n<h2>Summary of Opinions</h2>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343</a></p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307871\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307871</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307707\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307707</a></li>\n</ul>",
      "rawMarkdown": "To the Kaggle staff, the\nSCIRO team.\n\nThank you very much for organizing this competition.\n\nAs a participant, I would like to give you some feedback about the accuracy of the labels in the competition data.\n\nAs some of the other participants may have noticed, the training labels were inconsistent and there were many missing labels. My solution[1] showed an increase in the LB score by correcting the training labels, and the 3rd-place solution[2] suggests that there may be systematic errors in the fitness of the labels on the public and private datasets.\n\nThis suggests that the evaluation labels may have been as inaccurate as the training labels. (Of course, we participants cannot see the test data set, so we can only speculate.)\n\nI’m not requesting re-evaluating this competitions result. Even if some of the other participants argue that we should, we shouldn’t. That would amount to an ex post facto application of the law.\n\nInstead, I would like to insist that the next time you hold a similar competition, I want you to pay full attention and cost to the accuracy of the labels.\n\nIf one is a skilled kaggler, he/she know how to train models with good generalization performance even when the training dataset is inaccurate.\nHowever, as we know, when the evaluation dataset is inaccurate, it is inevitable that the winning model will be suboptimal.\n\nThis is not only unfair to the participants of the competition, but also disadvantageous to the host who will operate the model, as the model that will eventually be operated will not perform well enough.\n\nThis competition had a high prize money at stake. Therefore, the cost could have been spent on making the labels for evaluation (and preferably for training) more accurate.\n\nSo please think about this.\n\nFinally, I hope that this discussion will be used for constructive purposes.\n\n## Summary of Opinions\n\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343\n\n## Reference\n\n* [1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307871\n* [2] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307707",
      "votes": 33
    },
    {
      "id": 1695657,
      "postDate": "2022-02-18T09:36:05.007Z",
      "content": "<p>Thanks for opening this post, I wanted to open a similar one.</p>\n<p>I think this competition would have been better, and this also applies to other competitions, if either of these two scenarios would apply:</p>\n<ol>\n<li>Test data is similar to training data, meaning that public LB data is similar to both training data and private test data.</li>\n<li>Test data is not similar to training data, but in this case public LB needs to be similar to private LB.</li>\n</ol>\n<p>In the first case, we expect the test labels to be of similar quality as the training labels, this is a common setup for data science problems. This would allow to fully focus on building good models for the training labels with CV, and it will generalize well to test / production.</p>\n<p>In the second case, we expect the test labels to be different. In this case training was labeled really messy, and it is absolutely a reasonable goal to make test predictions more precise / better. So test being more precise makes total sense. This would test more our ability to train on messy data (i.e. cheap data) and generalize to tight / precise boxes.</p>\n<p>Unfortunately, in this competition the public LB was completely useless. So it was the worst combination of my elaborations for above. It neither tests our ability to generalize well to training data, nor to new types of test labels. This is also a complete unrealistic real-world setup, as in industry data-science projects you would never build a hold-out dataset (public LB) that neither represents training data, nor your goal of generalization (future test data).</p>\n<p>Actually, the whole metric only makes sense with more precise labels. The 0.8 IOU threshold has such a huge impact on the metric, and even with \"perfect\" predictions you cannot reach high scores on training data, as boxes are so inconsistent / messy. As soon as you have more consistent labels, the score goes up a lot, as we also saw on public LB. This allows us to train more consistent and precise models.</p>",
      "rawMarkdown": "Thanks for opening this post, I wanted to open a similar one.\n\nI think this competition would have been better, and this also applies to other competitions, if either of these two scenarios would apply:\n\n1. Test data is similar to training data, meaning that public LB data is similar to both training data and private test data.\n2. Test data is not similar to training data, but in this case public LB needs to be similar to private LB.\n\nIn the first case, we expect the test labels to be of similar quality as the training labels, this is a common setup for data science problems. This would allow to fully focus on building good models for the training labels with CV, and it will generalize well to test / production.\n\nIn the second case, we expect the test labels to be different. In this case training was labeled really messy, and it is absolutely a reasonable goal to make test predictions more precise / better. So test being more precise makes total sense. This would test more our ability to train on messy data (i.e. cheap data) and generalize to tight / precise boxes.\n\nUnfortunately, in this competition the public LB was completely useless. So it was the worst combination of my elaborations for above. It neither tests our ability to generalize well to training data, nor to new types of test labels. This is also a complete unrealistic real-world setup, as in industry data-science projects you would never build a hold-out dataset (public LB) that neither represents training data, nor your goal of generalization (future test data).\n\nActually, the whole metric only makes sense with more precise labels. The 0.8 IOU threshold has such a huge impact on the metric, and even with \"perfect\" predictions you cannot reach high scores on training data, as boxes are so inconsistent / messy. As soon as you have more consistent labels, the score goes up a lot, as we also saw on public LB. This allows us to train more consistent and precise models.",
      "votes": 10,
      "replies": [
        {
          "id": 1695806,
          "postDate": "2022-02-18T11:31:07.393Z",
          "content": "<p>I totally agree with that. Besides, you made the problem more focused. Thank you for the comment.</p>",
          "rawMarkdown": "I totally agree with that. Besides, you made the problem more focused. Thank you for the comment.",
          "votes": 1
        },
        {
          "id": 1695970,
          "postDate": "2022-02-18T14:07:45.477Z",
          "content": "<p>There's a nuance you're missing here, and that's dropout.</p>\n<p>For competition hosters / designers there is sometimes a motivation to obscure the private leaderboard as they don't want to bias contestants to overfit on some particular goal.  </p>\n<p>Take the Pawpularity contest as an example.  It's obvious certain biases (breeds, grooming requirements) exist in the dataset exist which can be overfit on.  The hosters know that, but might not want solutions which necessarily overfit on that data, as good and as effective as a signal as it might be.   So their holdout set is all the less desired pets, which they didn't put in the train/test.</p>\n<p>Designers may also want to hold contests to try different types of train/test and holdout through a series of competitions to better understand the relationships between the data.</p>\n<p>Yes, that's frustrating and luck seems to plays a greater factor, but it also gives the hosters an opportunity to look through the various solutions and see what works and what doesn't.  </p>\n<p>Sometimes you just can't control everything, as much as you wish you could.  We all have to accept being cogs to a degree.  I wouldn't be surprised to discover that all the people who get paid in a particular contest provide unusable solutions, and its lower ranked solutions (even as far down as silver) end up being actually used by the hosters.</p>\n<p>Maybe an idea here though .. we should get three submissions.  One for max LB, one for max CV, and one for LB/CV blend.  </p>\n<p>Encouraging a comment period from contestants might be a good idea as well before finalizing a contest.  I am pretty confident the hosters/kaggle staff would benefit greatly from this.</p>",
          "rawMarkdown": "There's a nuance you're missing here, and that's dropout.\n\nFor competition hosters / designers there is sometimes a motivation to obscure the private leaderboard as they don't want to bias contestants to overfit on some particular goal.  \n\nTake the Pawpularity contest as an example.  It's obvious certain biases (breeds, grooming requirements) exist in the dataset exist which can be overfit on.  The hosters know that, but might not want solutions which necessarily overfit on that data, as good and as effective as a signal as it might be.   So their holdout set is all the less desired pets, which they didn't put in the train/test.\n\nDesigners may also want to hold contests to try different types of train/test and holdout through a series of competitions to better understand the relationships between the data.\n\nYes, that's frustrating and luck seems to plays a greater factor, but it also gives the hosters an opportunity to look through the various solutions and see what works and what doesn't.  \n\nSometimes you just can't control everything, as much as you wish you could.  We all have to accept being cogs to a degree.  I wouldn't be surprised to discover that all the people who get paid in a particular contest provide unusable solutions, and its lower ranked solutions (even as far down as silver) end up being actually used by the hosters.\n\nMaybe an idea here though .. we should get three submissions.  One for max LB, one for max CV, and one for LB/CV blend.  \n\nEncouraging a comment period from contestants might be a good idea as well before finalizing a contest.  I am pretty confident the hosters/kaggle staff would benefit greatly from this."
        },
        {
          "id": 1696036,
          "postDate": "2022-02-18T14:55:22.183Z",
          "content": "<p>Maybe you didn't understand what I wrote, it is fine for test to be different, but making the public LB different to train and private, will lead to, on average, solutions. Because obviously people will try to account for the new peculiarities given in the public test data.</p>\n<p>If you would completely remove public test data from this competition, you would get much stronger solutions for the actual private data.</p>",
          "rawMarkdown": "Maybe you didn't understand what I wrote, it is fine for test to be different, but making the public LB different to train and private, will lead to, on average, solutions. Because obviously people will try to account for the new peculiarities given in the public test data.\n\nIf you would completely remove public test data from this competition, you would get much stronger solutions for the actual private data."
        },
        {
          "id": 1696213,
          "postDate": "2022-02-18T16:56:34.650Z",
          "content": "<p>Yes, and I  agree with you, but sometimes you have to weed through bad solutions to do some exploration versus exploitation.   It may seem like a lot of money to some folks to do these sorts of experiments, but it actually isn't.</p>\n<p>It's also possible they just screwed up.  As I mentioned, a comment period could help reduce that from happening.</p>",
          "rawMarkdown": "Yes, and I  agree with you, but sometimes you have to weed through bad solutions to do some exploration versus exploitation.   It may seem like a lot of money to some folks to do these sorts of experiments, but it actually isn't.\n\nIt's also possible they just screwed up.  As I mentioned, a comment period could help reduce that from happening.\n\n"
        },
        {
          "id": 1696282,
          "postDate": "2022-02-18T18:00:41.250Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1696285,
          "postDate": "2022-02-18T18:03:03.413Z",
          "content": "<p>I expected this competition to be meant as serious effort to help to prevent one big threat to Great Barrier Reef. And I was looking forward not only to seeing great solutions but also the biggest winner - GBR itself. <br>\nAs the both main posters here explained, the lack of labelling accuracy and unnatural data split didn't help creating the best possible and practically usable solutions. IMHO the Great Barrier Reef didn't win at all and I am only asking who and why chose such a strange data split. <br>\nCan anyone or anything persuade Kaggle and the hosts to learn from this example at least in nature conservation competitions where there is so much at stake?      </p>",
          "rawMarkdown": "I expected this competition to be meant as serious effort to help to prevent one big threat to Great Barrier Reef. And I was looking forward not only to seeing great solutions but also the biggest winner - GBR itself. \nAs the both main posters here explained, the lack of labelling accuracy and unnatural data split didn't help creating the best possible and practically usable solutions. IMHO the Great Barrier Reef didn't win at all and I am only asking who and why chose such a strange data split. \nCan anyone or anything persuade Kaggle and the hosts to learn from this example at least in nature conservation competitions where there is so much at stake?      "
        },
        {
          "id": 1696290,
          "postDate": "2022-02-18T18:07:44.897Z",
          "content": "<p>Money?  Paying for kaggling?  Wait, what?  </p>\n<p>I do have an agenda, for sure, but it's mostly just that people should feel free to share as much as they want.  I agree however, both sides of that discussion do not belong in competition forums.  I am happy to keep that to General if folks on the other side do as well.</p>\n<p>However, I don't think that is being discussed at all in thread so your post is very very confusing to me.</p>\n<p>But as to the 'real topic', I think it's fair to say that this competition may have been bungled.  My comments were more referring to the many comments insisting that all competitions should follow some predefined pattern regarding the split.  Please see my example above regarding Pawpularity for a reason why that wouldn't be ideal.</p>\n<p>*Also, it's not entirely unclear that nothing was learned.  It's very possible that really this was just a great opportunity for TF folks to confirm how far they are behind on PyTorch and money was very well spent.  It's also possible the TF folks didn't want a great solution, because they didn't want to look that bad..  *</p>\n<p>Folks need to accept that we're all just playing a part here.  An important part, for sure, but we don't control the game.  </p>",
          "rawMarkdown": "Money?  Paying for kaggling?  Wait, what?  \n\nI do have an agenda, for sure, but it's mostly just that people should feel free to share as much as they want.  I agree however, both sides of that discussion do not belong in competition forums.  I am happy to keep that to General if folks on the other side do as well.\n\nHowever, I don't think that is being discussed at all in thread so your post is very very confusing to me.\n\nBut as to the 'real topic', I think it's fair to say that this competition may have been bungled.  My comments were more referring to the many comments insisting that all competitions should follow some predefined pattern regarding the split.  Please see my example above regarding Pawpularity for a reason why that wouldn't be ideal.\n\n*Also, it's not entirely unclear that nothing was learned.  It's very possible that really this was just a great opportunity for TF folks to confirm how far they are behind on PyTorch and money was very well spent.  It's also possible the TF folks didn't want a great solution, because they didn't want to look that bad..  *\n\nFolks need to accept that we're all just playing a part here.  An important part, for sure, but we don't control the game.  "
        },
        {
          "id": 1696474,
          "postDate": "2022-02-18T21:36:51.753Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1696486,
          "postDate": "2022-02-18T22:06:27.907Z",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> Thank you for presenting an opposing view.</p>\n<p>However, it is a little difficult to understand the point of your argument.<br>\nAre you claiming that \"the designers of this competition deliberately chose dissimilar public and private datasets in order to get a good model\"?</p>\n<p>If so, I doubt that doing so will give the organizer and host a good model. How can you prove that the model finally chosen is not a model that happens to overfit the private dataset? I think relying on the luck factor never making things better.</p>",
          "rawMarkdown": "@kaggleqrdl Thank you for presenting an opposing view.\n\nHowever, it is a little difficult to understand the point of your argument.\nAre you claiming that \"the designers of this competition deliberately chose dissimilar public and private datasets in order to get a good model\"?\n\nIf so, I doubt that doing so will give the organizer and host a good model. How can you prove that the model finally chosen is not a model that happens to overfit the private dataset? I think relying on the luck factor never making things better."
        },
        {
          "id": 1696555,
          "postDate": "2022-02-18T23:41:19.913Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1696565,
          "postDate": "2022-02-19T00:01:22.337Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1696662,
          "postDate": "2022-02-19T02:13:09.967Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1696705,
          "postDate": "2022-02-19T03:38:16.480Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1696800,
          "postDate": "2022-02-19T05:45:53.257Z",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>As you and PSI ranked highly here, I'd tend to agree with your perspectives on this competition.  My point was only that I'm not sure we can extrapolate from this competition to future competitions.   Also, as I mentioned, we may want to look at this competition through the lens of analyzing the competitive status of TF vrs PyTorch rather than just finding COTS.</p>",
          "rawMarkdown": "@tatamikenn \n\nAs you and PSI ranked highly here, I'd tend to agree with your perspectives on this competition.  My point was only that I'm not sure we can extrapolate from this competition to future competitions.   Also, as I mentioned, we may want to look at this competition through the lens of analyzing the competitive status of TF vrs PyTorch rather than just finding COTS."
        },
        {
          "id": 1697540,
          "postDate": "2022-02-19T17:20:40.293Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1697673,
          "postDate": "2022-02-19T19:04:03.337Z",
          "content": "<p>Being kind to one another and showing mutual respect is the very essence of Kaggle, <a href=\"https://www.kaggle.com/blankaf\" target=\"_blank\">@blankaf</a></p>",
          "rawMarkdown": "Being kind to one another and showing mutual respect is the very essence of Kaggle, @blankaf"
        },
        {
          "id": 1697690,
          "postDate": "2022-02-19T19:27:42.540Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1697930,
          "postDate": "2022-02-20T01:04:18.450Z",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> </p>\n<blockquote>\n  <p>As you and PSI ranked highly here, I'd tend to agree with your perspectives on this competition.</p>\n</blockquote>\n<p>This discussion is independent of the LB score. You can share what you think.<br>\nSo, can you tell me about the below question?</p>\n<blockquote>\n  <p>Are you claiming that \"the designers of this competition deliberately chose dissimilar public and private datasets in order to get a good model\"?</p>\n</blockquote>",
          "rawMarkdown": "@kaggleqrdl \n\n> As you and PSI ranked highly here, I'd tend to agree with your perspectives on this competition.\n\nThis discussion is independent of the LB score. You can share what you think.\nSo, can you tell me about the below question?\n\n> Are you claiming that \"the designers of this competition deliberately chose dissimilar public and private datasets in order to get a good model\"?"
        },
        {
          "id": 1698157,
          "postDate": "2022-02-20T07:27:12.970Z",
          "content": "<p>This thread is too long for me to read all the replies, but I would also like to offer a possibility: if they reduce the percentage of public test data part to less than 5%, can we avoided such discussion as well? I remember that there were cases in past competitions where the percentage of public test data was very low.</p>",
          "rawMarkdown": "This thread is too long for me to read all the replies, but I would also like to offer a possibility: if they reduce the percentage of public test data part to less than 5%, can we avoided such discussion as well? I remember that there were cases in past competitions where the percentage of public test data was very low."
        },
        {
          "id": 1698246,
          "postDate": "2022-02-20T08:29:44.870Z",
          "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Thanks for the comment.</p>\n<blockquote>\n  <p>This thread is too long for me to read all the replies, but I would also like to offer a possibility</p>\n</blockquote>\n<p>Sorry, but Psi's initial comment summarizes the problem well. So you can read this first.</p>\n<blockquote>\n  <p>I would also like to offer a possibility: if they reduce the percentage of public test data part to less than 5%, can we avoided such discussion as well?</p>\n</blockquote>\n<p>I don't think so. One of our concern is about the dissimilarity of public and private test data. So I think changing the amount of the split never solve the problem.</p>",
          "rawMarkdown": "@haqishen Thanks for the comment.\n\n> This thread is too long for me to read all the replies, but I would also like to offer a possibility\n\nSorry, but Psi's initial comment summarizes the problem well. So you can read this first.\n\n> I would also like to offer a possibility: if they reduce the percentage of public test data part to less than 5%, can we avoided such discussion as well?\n\nI don't think so. One of our concern is about the dissimilarity of public and private test data. So I think changing the amount of the split never solve the problem."
        },
        {
          "id": 1698276,
          "postDate": "2022-02-20T08:59:32.110Z",
          "content": "<p>Ahh.. sorry, I should have made my meaning clearly.<br>\nI read PSI's comment, and want to add 3rd possibility on it.<br>\nIf public test data is only 5%, then everyone will know that public LB is unreliable.<br>\nThen everyone will focus more on optimizing local CV.<br>\nThere have been competitions in the past where this was the case.</p>\n<p>In addition, to be fair (although you may still think this is my bias) this competition dataset is in my opinion playing up to the kaggle average. To be honest, no leakage is already beyond 15-20% of the competitions.<br>\nIf the first competition you entered was the one below, you'd probably be so mad you'd just give up on kaggle ;)<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification</a></p>",
          "rawMarkdown": "Ahh.. sorry, I should have made my meaning clearly.\nI read PSI's comment, and want to add 3rd possibility on it.\nIf public test data is only 5%, then everyone will know that public LB is unreliable.\nThen everyone will focus more on optimizing local CV.\nThere have been competitions in the past where this was the case.\n\nIn addition, to be fair (although you may still think this is my bias) this competition dataset is in my opinion playing up to the kaggle average. To be honest, no leakage is already beyond 15-20% of the competitions.\nIf the first competition you entered was the one below, you'd probably be so mad you'd just give up on kaggle ;)\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification",
          "votes": 3
        },
        {
          "id": 1698307,
          "postDate": "2022-02-20T09:28:54.147Z",
          "content": "<blockquote>\n  <p>If public test data is only 5%, then everyone will know that public LB is unreliable.<br>\n  Then everyone will focus more on optimizing local CV.</p>\n</blockquote>\n<p>Oh, I understand your point.<br>\nYes, I think this also guides not to overfitting to public LB (although it seems kind of work-around).</p>\n<blockquote>\n  <p>In addition, to be fair (although you may still think this is my bias) this competition dataset is in my opinion playing up to the kaggle average.</p>\n</blockquote>\n<p>I see. So maybe Kaggle and host considers this quality is enough for their demanding level of the winning models.</p>",
          "rawMarkdown": "> If public test data is only 5%, then everyone will know that public LB is unreliable.\n> Then everyone will focus more on optimizing local CV.\n\nOh, I understand your point.\nYes, I think this also guides not to overfitting to public LB (although it seems kind of work-around).\n\n> In addition, to be fair (although you may still think this is my bias) this competition dataset is in my opinion playing up to the kaggle average.\n\nI see. So maybe Kaggle and host considers this quality is enough for their demanding level of the winning models."
        },
        {
          "id": 1698310,
          "postDate": "2022-02-20T09:29:35.023Z",
          "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>\n<p>The problem with that is you lose the ongoing competitive spirit.  Seeing a leaderboard and knowing that someone is higher ranked than what you're currently capable of doing can be inspiring.</p>",
          "rawMarkdown": "@haqishen \n\nThe problem with that is you lose the ongoing competitive spirit.  Seeing a leaderboard and knowing that someone is higher ranked than what you're currently capable of doing can be inspiring."
        },
        {
          "id": 1698320,
          "postDate": "2022-02-20T09:45:19.403Z",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> <br>\nThe truth is that many organizers don't run competitions for using winning solutions… Just for branding or hiring or something like that.<br>\nAnd, it turns out that some of the organizers' head directors don't even understand machine learning at all.<br>\nAnyway, congrats to your solo gold for your first competition! Hope to see you in future competitions.</p>\n<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> <br>\nOk I see your point, I was looking forward to learn how people push score to 0.8 before as well ;)</p>",
          "rawMarkdown": "@tatamikenn \nThe truth is that many organizers don't run competitions for using winning solutions... Just for branding or hiring or something like that.\nAnd, it turns out that some of the organizers' head directors don't even understand machine learning at all.\nAnyway, congrats to your solo gold for your first competition! Hope to see you in future competitions.\n\n@kaggleqrdl \nOk I see your point, I was looking forward to learn how people push score to 0.8 before as well ;)",
          "votes": 2
        },
        {
          "id": 1698366,
          "postDate": "2022-02-20T10:45:56.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Thanks! And congrats to your team’s 1st prize too.</p>",
          "rawMarkdown": "@haqishen Thanks! And congrats to your team’s 1st prize too."
        },
        {
          "id": 1698513,
          "postDate": "2022-02-20T12:55:25.840Z",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> <br>\nthank you for opening this discussion about an important topic.</p>",
          "rawMarkdown": "@tatamikenn \nthank you for opening this discussion about an important topic.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1699078,
      "postDate": "2022-02-20T22:37:24.420Z",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> . My own benchmarks (which I'm sure many of you have already seen from my discussion post <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607</a>) suggest that public LB's bboxes are systematically 10% tighter than that of the private LB.</p>\n<p>At IOU 0.8, assuming the centers of the bboxes are exact, each dimension must be within ~10% error to score a positive hit (as 0.9 x 0.9 = 0.81). If this is not the case, the correct inference would be counted as a false positive, and will actually impact the F2 score. This punishes high-recall models that systematically over/underestimate the dimensions of the bboxes compared with the bbox annotations, which goes against the competition organiser's stated aim that \"in this case it makes sense to tolerate some false positives in order to ensure very few starfish are missed.\"</p>\n<p>I think manual labelling is always going to be impacted by human error, yet auto-labelling will produce systematic bias and will favor one architecture / method over another. A simple solution would simply have been to alter the competition metric to be less sensitive to IOU (e.g. F2 @ IOU 0.3 to 0.6 with step 0.05), or simply F2 @ IOU 0.5.</p>",
      "rawMarkdown": "I agree with @tatamikenn and @philippsinger . My own benchmarks (which I'm sure many of you have already seen from my discussion post https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607) suggest that public LB's bboxes are systematically 10% tighter than that of the private LB.\n\nAt IOU 0.8, assuming the centers of the bboxes are exact, each dimension must be within ~10% error to score a positive hit (as 0.9 x 0.9 = 0.81). If this is not the case, the correct inference would be counted as a false positive, and will actually impact the F2 score. This punishes high-recall models that systematically over/underestimate the dimensions of the bboxes compared with the bbox annotations, which goes against the competition organiser's stated aim that \"in this case it makes sense to tolerate some false positives in order to ensure very few starfish are missed.\"\n\nI think manual labelling is always going to be impacted by human error, yet auto-labelling will produce systematic bias and will favor one architecture / method over another. A simple solution would simply have been to alter the competition metric to be less sensitive to IOU (e.g. F2 @ IOU 0.3 to 0.6 with step 0.05), or simply F2 @ IOU 0.5.",
      "votes": 3,
      "replies": [
        {
          "id": 1699243,
          "postDate": "2022-02-21T03:44:30.380Z",
          "content": "<blockquote>\n  <p>A simple solution would simply have been to alter the competition metric to be less sensitive to IOU (e.g. F2 @ IOU 0.3 to 0.6 with step 0.05), or simply F2 @ IOU 0.5</p>\n</blockquote>\n<p>I agree. This is another solution of the problem.<br>\nThanks for the comment!</p>",
          "rawMarkdown": "> A simple solution would simply have been to alter the competition metric to be less sensitive to IOU (e.g. F2 @ IOU 0.3 to 0.6 with step 0.05), or simply F2 @ IOU 0.5\n\nI agree. This is another solution of the problem.\nThanks for the comment!"
        }
      ]
    },
    {
      "id": 1699343,
      "postDate": "2022-02-21T05:52:01.937Z",
      "content": "<p>Thanks for posting a lot of opinions. I make summary of the opinions.<br>\n(Please point out if I missed any other opinions.)</p>\n<h2>Problems</h2>\n<ol>\n<li>Train labels are inaccurate -&gt; that is not a problem</li>\n<li>(Probably) test labels are inaccurate</li>\n<li>Private and Public labels are dissimilar</li>\n<li>Evaluation metrics are too sensitive compared to the accuracy of labels</li>\n</ol>\n<h2>Solutions</h2>\n<ol>\n<li>Make more accurate test labels (by <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> )</li>\n<li>Split public/private test data as they are similar[1] (by <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> ) </li>\n<li>Use less public LB samples (e.g. 5%) to make public LB more incredible, and let participants focus more on trusting their CVs[2] (by <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> )</li>\n<li>Make competition metrics less sensitive[3] (e.g. F2@0.3:0.6, F2@0.5) (by <a href=\"https://www.kaggle.com/alexchwong\" target=\"_blank\">@alexchwong</a> )</li>\n<li>Having pre contest period where contestants get a chance to comment and make suggestions about the data[4] (by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> )</li>\n</ol>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698276\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698276</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699078\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699078</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695551\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695551</a></li>\n</ul>",
      "rawMarkdown": "Thanks for posting a lot of opinions. I make summary of the opinions.\n(Please point out if I missed any other opinions.)\n\n## Problems\n\n1. Train labels are inaccurate -> that is not a problem\n2. (Probably) test labels are inaccurate\n3. Private and Public labels are dissimilar\n4. Evaluation metrics are too sensitive compared to the accuracy of labels\n\n## Solutions\n\n1. Make more accurate test labels (by @tatamikenn )\n2. Split public/private test data as they are similar[1] (by @philippsinger ) \n3. Use less public LB samples (e.g. 5%) to make public LB more incredible, and let participants focus more on trusting their CVs[2] (by @haqishen )\n4. Make competition metrics less sensitive[3] (e.g. F2@0.3:0.6, F2@0.5) (by @alexchwong )\n5. Having pre contest period where contestants get a chance to comment and make suggestions about the data[4] (by @kaggleqrdl )\n\n## Reference\n\n* [1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657\n* [2] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698276\n* [3] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699078\n* [4] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695551",
      "votes": 1,
      "replies": [
        {
          "id": 1699708,
          "postDate": "2022-02-21T11:24:44.850Z",
          "content": "<p>Good summary.  Hopefully Kaggle staff can give us an update on where they are with determining splits, even if just generally speaking.   Maybe we don't necesarily deserve concrete action, but I think we do deserve some updated insight into their thinking.</p>\n<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> You should tag Addison Howard to make sure this gets visibility.  </p>",
          "rawMarkdown": "Good summary.  Hopefully Kaggle staff can give us an update on where they are with determining splits, even if just generally speaking.   Maybe we don't necesarily deserve concrete action, but I think we do deserve some updated insight into their thinking.\n\n@tatamikenn You should tag Addison Howard to make sure this gets visibility.  "
        },
        {
          "id": 1699834,
          "postDate": "2022-02-21T13:27:53.997Z",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> </p>\n<blockquote>\n  <p>Maybe we don't necesarily deserve concrete action, but I think we do deserve some updated insight into their thinking.</p>\n</blockquote>\n<p>Yeah, it’s my hope too.</p>\n<blockquote>\n  <p>You should tag Addison Howard to make sure this gets visibility.</p>\n</blockquote>\n<p><br>\nWe had the reply from him:</p>\n<blockquote>\n  <p>Thank you! Certainly passing along to the impacted parties.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698237\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698237</a></p>",
          "rawMarkdown": "@kaggleqrdl \n\n> Maybe we don't necesarily deserve concrete action, but I think we do deserve some updated insight into their thinking.\n\nYeah, it’s my hope too.\n\n> You should tag Addison Howard to make sure this gets visibility.\n\n~~I already mention them in the below comment. However no reply is coming so far.~~\nWe had the reply from him:\n\n> Thank you! Certainly passing along to the impacted parties.\n\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698237"
        },
        {
          "id": 1699872,
          "postDate": "2022-02-21T13:49:35.600Z",
          "content": "<p>Just in case they missed, I also mentioned it in the finalization post of this competition.<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308248#1699868\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308248#1699868</a></p>",
          "rawMarkdown": "Just in case they missed, I also mentioned it in the finalization post of this competition.\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308248#1699868"
        },
        {
          "id": 1700645,
          "postDate": "2022-02-22T06:05:17.827Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1695611,
      "postDate": "2022-02-18T08:56:33.823Z",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> - was looking at TIDE during this competition and know you posted a topic on it. While reading on it, this article was along your point:</p>\n<p><a href=\"https://towardsdatascience.com/a-better-map-for-object-detection-32662767d424\" target=\"_blank\">https://towardsdatascience.com/a-better-map-for-object-detection-32662767d424</a><br>\n\"Recently, AI pioneer Andrew Ng launched a campaign for data-centric AI where his main goal is to shift the focus of AI practitioners from model/algorithm development to the quality of the data they use to train the models.\" </p>\n<p>Cannot remember the link but read that improvements to e.g. COCO become harder now because the underlying annotations have issues. So certainly an area to watch.  </p>",
      "rawMarkdown": "@tatamikenn - was looking at TIDE during this competition and know you posted a topic on it. While reading on it, this article was along your point:\n\nhttps://towardsdatascience.com/a-better-map-for-object-detection-32662767d424\n\"Recently, AI pioneer Andrew Ng launched a campaign for data-centric AI where his main goal is to shift the focus of AI practitioners from model/algorithm development to the quality of the data they use to train the models.\" \n\nCannot remember the link but read that improvements to e.g. COCO become harder now because the underlying annotations have issues. So certainly an area to watch.  ",
      "votes": 1,
      "replies": [
        {
          "id": 1695624,
          "postDate": "2022-02-18T09:15:45.673Z",
          "content": "<p>I see. I had heard of the keyword, but did not know about the trend.<br>\nThanks for sharing.</p>",
          "rawMarkdown": "I see. I had heard of the keyword, but did not know about the trend.\nThanks for sharing.\n\n"
        }
      ]
    },
    {
      "id": 1698237,
      "postDate": "2022-02-20T08:23:40.160Z",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a><br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Our feedback is basically summarize to this topics description and Psi's comment[2].<br>\nWe are happy if you use these feedbacks to make future competition good.</p>\n<p>Thanks</p>\n<p>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657</a></p>",
      "rawMarkdown": "@addisonhoward\n@sohier \n\nOur feedback is basically summarize to this topics description and Psi's comment[2].\nWe are happy if you use these feedbacks to make future competition good.\n\nThanks\n\n[1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657",
      "replies": [
        {
          "id": 1699350,
          "postDate": "2022-02-21T05:57:00.847Z",
          "content": "<p>I made summary of the opinions. Please check this also.<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343</a></p>",
          "rawMarkdown": "I made summary of the opinions. Please check this also.\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343"
        },
        {
          "id": 1700595,
          "postDate": "2022-02-22T05:30:55.397Z",
          "content": "<p>Thank you! Certainly passing along to the impacted parties.</p>",
          "rawMarkdown": "Thank you! Certainly passing along to the impacted parties.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1696538,
      "postDate": "2022-02-18T23:02:20.817Z",
      "content": "<p>Thanks. But are you sure commenting to the right post? This is not a solution post.</p>",
      "rawMarkdown": "Thanks. But are you sure commenting to the right post? This is not a solution post."
    },
    {
      "id": 1695492,
      "postDate": "2022-02-18T07:20:26.507Z",
      "content": "<p>It's a common complaint, and I'm fairly sure they pay very very close attention, it's just an immensely difficult problem to solve.  It may seem like these issues should be obvious, but appreciate that  it's usually after the fact and after the collective hive mind is hitting this data very hard with very sophisticated EDA techniques. </p>\n<p>We're usually given 2 or more submissions.  A baseline technique in most competitions is to use your highest CV and highest LB.  </p>",
      "rawMarkdown": "It's a common complaint, and I'm fairly sure they pay very very close attention, it's just an immensely difficult problem to solve.  It may seem like these issues should be obvious, but appreciate that  it's usually after the fact and after the collective hive mind is hitting this data very hard with very sophisticated EDA techniques. \n\nWe're usually given 2 or more submissions.  A baseline technique in most competitions is to use your highest CV and highest LB.  ",
      "replies": [
        {
          "id": 1695513,
          "postDate": "2022-02-18T07:33:19.190Z",
          "content": "<p>Thank you for commenting.</p>\n<p>I understand that this is a difficult problem to solve in general, but for this competition, I think it would have been possible to provide a higher quality data set.</p>\n<p>As you can see when you actually view the labels with the images, there are many frames where the labels are way off and distant COTS are not labeled.</p>\n<p>This kind of error is understandable if the labels were assigned automatically by a machine, but according to the paper of this competition, they were assigned by a human with assistance from a machine. I believe that human-assigned labels can be made more accurate.</p>\n<p>Also, I am not complaining. I am giving feedback so that the quality of this and other competitions will be higher.</p>",
          "rawMarkdown": "Thank you for commenting.\n\nI understand that this is a difficult problem to solve in general, but for this competition, I think it would have been possible to provide a higher quality data set.\n\nAs you can see when you actually view the labels with the images, there are many frames where the labels are way off and distant COTS are not labeled.\n\nThis kind of error is understandable if the labels were assigned automatically by a machine, but according to the paper of this competition, they were assigned by a human with assistance from a machine. I believe that human-assigned labels can be made more accurate.\n\nAlso, I am not complaining. I am giving feedback so that the quality of this and other competitions will be higher.",
          "votes": 1
        },
        {
          "id": 1695551,
          "postDate": "2022-02-18T07:56:54.640Z",
          "content": "<p>Yeah, maybe there should be a pre contest period where contestants get a chance to comment and make suggestions about the data.</p>",
          "rawMarkdown": "Yeah, maybe there should be a pre contest period where contestants get a chance to comment and make suggestions about the data."
        },
        {
          "id": 1695621,
          "postDate": "2022-02-18T09:07:46.783Z",
          "content": "<p>I agree. That is one of an idea for preventing mismatch of data accuracy.</p>",
          "rawMarkdown": "I agree. That is one of an idea for preventing mismatch of data accuracy."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1695657,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2022-02-18T09:36:05.007000",
      "content": "<p>Thanks for opening this post, I wanted to open a similar one.</p>\n<p>I think this competition would have been better, and this also applies to other competitions, if either of these two scenarios would apply:</p>\n<ol>\n<li>Test data is similar to training data, meaning that public LB data is similar to both training data and private test data.</li>\n<li>Test data is not similar to training data, but in this case public LB needs to be similar to private LB.</li>\n</ol>\n<p>In the first case, we expect the test labels to be of similar quality as the training labels, this is a common setup for data science problems. This would allow to fully focus on building good models for the training labels with CV, and it will generalize well to test / production.</p>\n<p>In the second case, we expect the test labels to be different. In this case training was labeled really messy, and it is absolutely a reasonable goal to make test predictions more precise / better. So test being more precise makes total sense. This would test more our ability to train on messy data (i.e. cheap data) and generalize to tight / precise boxes.</p>\n<p>Unfortunately, in this competition the public LB was completely useless. So it was the worst combination of my elaborations for above. It neither tests our ability to generalize well to training data, nor to new types of test labels. This is also a complete unrealistic real-world setup, as in industry data-science projects you would never build a hold-out dataset (public LB) that neither represents training data, nor your goal of generalization (future test data).</p>\n<p>Actually, the whole metric only makes sense with more precise labels. The 0.8 IOU threshold has such a huge impact on the metric, and even with \"perfect\" predictions you cannot reach high scores on training data, as boxes are so inconsistent / messy. As soon as you have more consistent labels, the score goes up a lot, as we also saw on public LB. This allows us to train more consistent and precise models.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1695806,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-18T11:31:07.393000",
          "content": "<p>I totally agree with that. Besides, you made the problem more focused. Thank you for the comment.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1695970,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-18T14:07:45.477000",
          "content": "<p>There's a nuance you're missing here, and that's dropout.</p>\n<p>For competition hosters / designers there is sometimes a motivation to obscure the private leaderboard as they don't want to bias contestants to overfit on some particular goal.  </p>\n<p>Take the Pawpularity contest as an example.  It's obvious certain biases (breeds, grooming requirements) exist in the dataset exist which can be overfit on.  The hosters know that, but might not want solutions which necessarily overfit on that data, as good and as effective as a signal as it might be.   So their holdout set is all the less desired pets, which they didn't put in the train/test.</p>\n<p>Designers may also want to hold contests to try different types of train/test and holdout through a series of competitions to better understand the relationships between the data.</p>\n<p>Yes, that's frustrating and luck seems to plays a greater factor, but it also gives the hosters an opportunity to look through the various solutions and see what works and what doesn't.  </p>\n<p>Sometimes you just can't control everything, as much as you wish you could.  We all have to accept being cogs to a degree.  I wouldn't be surprised to discover that all the people who get paid in a particular contest provide unusable solutions, and its lower ranked solutions (even as far down as silver) end up being actually used by the hosters.</p>\n<p>Maybe an idea here though .. we should get three submissions.  One for max LB, one for max CV, and one for LB/CV blend.  </p>\n<p>Encouraging a comment period from contestants might be a good idea as well before finalizing a contest.  I am pretty confident the hosters/kaggle staff would benefit greatly from this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696036,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-18T14:55:22.183000",
          "content": "<p>Maybe you didn't understand what I wrote, it is fine for test to be different, but making the public LB different to train and private, will lead to, on average, solutions. Because obviously people will try to account for the new peculiarities given in the public test data.</p>\n<p>If you would completely remove public test data from this competition, you would get much stronger solutions for the actual private data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696213,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-18T16:56:34.650000",
          "content": "<p>Yes, and I  agree with you, but sometimes you have to weed through bad solutions to do some exploration versus exploitation.   It may seem like a lot of money to some folks to do these sorts of experiments, but it actually isn't.</p>\n<p>It's also possible they just screwed up.  As I mentioned, a comment period could help reduce that from happening.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696282,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-18T18:00:41.250000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696285,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2022-02-18T18:03:03.413000",
          "content": "<p>I expected this competition to be meant as serious effort to help to prevent one big threat to Great Barrier Reef. And I was looking forward not only to seeing great solutions but also the biggest winner - GBR itself. <br>\nAs the both main posters here explained, the lack of labelling accuracy and unnatural data split didn't help creating the best possible and practically usable solutions. IMHO the Great Barrier Reef didn't win at all and I am only asking who and why chose such a strange data split. <br>\nCan anyone or anything persuade Kaggle and the hosts to learn from this example at least in nature conservation competitions where there is so much at stake?      </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696290,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-18T18:07:44.897000",
          "content": "<p>Money?  Paying for kaggling?  Wait, what?  </p>\n<p>I do have an agenda, for sure, but it's mostly just that people should feel free to share as much as they want.  I agree however, both sides of that discussion do not belong in competition forums.  I am happy to keep that to General if folks on the other side do as well.</p>\n<p>However, I don't think that is being discussed at all in thread so your post is very very confusing to me.</p>\n<p>But as to the 'real topic', I think it's fair to say that this competition may have been bungled.  My comments were more referring to the many comments insisting that all competitions should follow some predefined pattern regarding the split.  Please see my example above regarding Pawpularity for a reason why that wouldn't be ideal.</p>\n<p>*Also, it's not entirely unclear that nothing was learned.  It's very possible that really this was just a great opportunity for TF folks to confirm how far they are behind on PyTorch and money was very well spent.  It's also possible the TF folks didn't want a great solution, because they didn't want to look that bad..  *</p>\n<p>Folks need to accept that we're all just playing a part here.  An important part, for sure, but we don't control the game.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696474,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-18T21:36:51.753000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696486,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-18T22:06:27.907000",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> Thank you for presenting an opposing view.</p>\n<p>However, it is a little difficult to understand the point of your argument.<br>\nAre you claiming that \"the designers of this competition deliberately chose dissimilar public and private datasets in order to get a good model\"?</p>\n<p>If so, I doubt that doing so will give the organizer and host a good model. How can you prove that the model finally chosen is not a model that happens to overfit the private dataset? I think relying on the luck factor never making things better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696555,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-18T23:41:19.913000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696565,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-19T00:01:22.337000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696662,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-19T02:13:09.967000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696705,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-19T03:38:16.480000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696800,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-19T05:45:53.257000",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>As you and PSI ranked highly here, I'd tend to agree with your perspectives on this competition.  My point was only that I'm not sure we can extrapolate from this competition to future competitions.   Also, as I mentioned, we may want to look at this competition through the lens of analyzing the competitive status of TF vrs PyTorch rather than just finding COTS.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1697540,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-19T17:20:40.293000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1697673,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-19T19:04:03.337000",
          "content": "<p>Being kind to one another and showing mutual respect is the very essence of Kaggle, <a href=\"https://www.kaggle.com/blankaf\" target=\"_blank\">@blankaf</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1697690,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-19T19:27:42.540000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1697930,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-20T01:04:18.450000",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> </p>\n<blockquote>\n  <p>As you and PSI ranked highly here, I'd tend to agree with your perspectives on this competition.</p>\n</blockquote>\n<p>This discussion is independent of the LB score. You can share what you think.<br>\nSo, can you tell me about the below question?</p>\n<blockquote>\n  <p>Are you claiming that \"the designers of this competition deliberately chose dissimilar public and private datasets in order to get a good model\"?</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1698157,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2022-02-20T07:27:12.970000",
          "content": "<p>This thread is too long for me to read all the replies, but I would also like to offer a possibility: if they reduce the percentage of public test data part to less than 5%, can we avoided such discussion as well? I remember that there were cases in past competitions where the percentage of public test data was very low.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1698246,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-20T08:29:44.870000",
          "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Thanks for the comment.</p>\n<blockquote>\n  <p>This thread is too long for me to read all the replies, but I would also like to offer a possibility</p>\n</blockquote>\n<p>Sorry, but Psi's initial comment summarizes the problem well. So you can read this first.</p>\n<blockquote>\n  <p>I would also like to offer a possibility: if they reduce the percentage of public test data part to less than 5%, can we avoided such discussion as well?</p>\n</blockquote>\n<p>I don't think so. One of our concern is about the dissimilarity of public and private test data. So I think changing the amount of the split never solve the problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1698276,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2022-02-20T08:59:32.110000",
          "content": "<p>Ahh.. sorry, I should have made my meaning clearly.<br>\nI read PSI's comment, and want to add 3rd possibility on it.<br>\nIf public test data is only 5%, then everyone will know that public LB is unreliable.<br>\nThen everyone will focus more on optimizing local CV.<br>\nThere have been competitions in the past where this was the case.</p>\n<p>In addition, to be fair (although you may still think this is my bias) this competition dataset is in my opinion playing up to the kaggle average. To be honest, no leakage is already beyond 15-20% of the competitions.<br>\nIf the first competition you entered was the one below, you'd probably be so mad you'd just give up on kaggle ;)<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1698307,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-20T09:28:54.147000",
          "content": "<blockquote>\n  <p>If public test data is only 5%, then everyone will know that public LB is unreliable.<br>\n  Then everyone will focus more on optimizing local CV.</p>\n</blockquote>\n<p>Oh, I understand your point.<br>\nYes, I think this also guides not to overfitting to public LB (although it seems kind of work-around).</p>\n<blockquote>\n  <p>In addition, to be fair (although you may still think this is my bias) this competition dataset is in my opinion playing up to the kaggle average.</p>\n</blockquote>\n<p>I see. So maybe Kaggle and host considers this quality is enough for their demanding level of the winning models.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1698310,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-20T09:29:35.023000",
          "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>\n<p>The problem with that is you lose the ongoing competitive spirit.  Seeing a leaderboard and knowing that someone is higher ranked than what you're currently capable of doing can be inspiring.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1698320,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2022-02-20T09:45:19.403000",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> <br>\nThe truth is that many organizers don't run competitions for using winning solutions… Just for branding or hiring or something like that.<br>\nAnd, it turns out that some of the organizers' head directors don't even understand machine learning at all.<br>\nAnyway, congrats to your solo gold for your first competition! Hope to see you in future competitions.</p>\n<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> <br>\nOk I see your point, I was looking forward to learn how people push score to 0.8 before as well ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1698366,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-20T10:45:56.837000",
          "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Thanks! And congrats to your team’s 1st prize too.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1698513,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2022-02-20T12:55:25.840000",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> <br>\nthank you for opening this discussion about an important topic.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1699078,
      "author_name": "Alex Wong",
      "author_url": "",
      "post_date": "2022-02-20T22:37:24.420000",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> . My own benchmarks (which I'm sure many of you have already seen from my discussion post <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607</a>) suggest that public LB's bboxes are systematically 10% tighter than that of the private LB.</p>\n<p>At IOU 0.8, assuming the centers of the bboxes are exact, each dimension must be within ~10% error to score a positive hit (as 0.9 x 0.9 = 0.81). If this is not the case, the correct inference would be counted as a false positive, and will actually impact the F2 score. This punishes high-recall models that systematically over/underestimate the dimensions of the bboxes compared with the bbox annotations, which goes against the competition organiser's stated aim that \"in this case it makes sense to tolerate some false positives in order to ensure very few starfish are missed.\"</p>\n<p>I think manual labelling is always going to be impacted by human error, yet auto-labelling will produce systematic bias and will favor one architecture / method over another. A simple solution would simply have been to alter the competition metric to be less sensitive to IOU (e.g. F2 @ IOU 0.3 to 0.6 with step 0.05), or simply F2 @ IOU 0.5.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1699243,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-21T03:44:30.380000",
          "content": "<blockquote>\n  <p>A simple solution would simply have been to alter the competition metric to be less sensitive to IOU (e.g. F2 @ IOU 0.3 to 0.6 with step 0.05), or simply F2 @ IOU 0.5</p>\n</blockquote>\n<p>I agree. This is another solution of the problem.<br>\nThanks for the comment!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1699343,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-02-21T05:52:01.937000",
      "content": "<p>Thanks for posting a lot of opinions. I make summary of the opinions.<br>\n(Please point out if I missed any other opinions.)</p>\n<h2>Problems</h2>\n<ol>\n<li>Train labels are inaccurate -&gt; that is not a problem</li>\n<li>(Probably) test labels are inaccurate</li>\n<li>Private and Public labels are dissimilar</li>\n<li>Evaluation metrics are too sensitive compared to the accuracy of labels</li>\n</ol>\n<h2>Solutions</h2>\n<ol>\n<li>Make more accurate test labels (by <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> )</li>\n<li>Split public/private test data as they are similar[1] (by <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> ) </li>\n<li>Use less public LB samples (e.g. 5%) to make public LB more incredible, and let participants focus more on trusting their CVs[2] (by <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> )</li>\n<li>Make competition metrics less sensitive[3] (e.g. F2@0.3:0.6, F2@0.5) (by <a href=\"https://www.kaggle.com/alexchwong\" target=\"_blank\">@alexchwong</a> )</li>\n<li>Having pre contest period where contestants get a chance to comment and make suggestions about the data[4] (by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> )</li>\n</ol>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698276\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698276</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699078\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699078</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695551\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695551</a></li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1699708,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-21T11:24:44.850000",
          "content": "<p>Good summary.  Hopefully Kaggle staff can give us an update on where they are with determining splits, even if just generally speaking.   Maybe we don't necesarily deserve concrete action, but I think we do deserve some updated insight into their thinking.</p>\n<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> You should tag Addison Howard to make sure this gets visibility.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1699834,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-21T13:27:53.997000",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> </p>\n<blockquote>\n  <p>Maybe we don't necesarily deserve concrete action, but I think we do deserve some updated insight into their thinking.</p>\n</blockquote>\n<p>Yeah, it’s my hope too.</p>\n<blockquote>\n  <p>You should tag Addison Howard to make sure this gets visibility.</p>\n</blockquote>\n<p><br>\nWe had the reply from him:</p>\n<blockquote>\n  <p>Thank you! Certainly passing along to the impacted parties.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698237\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698237</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1699872,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-21T13:49:35.600000",
          "content": "<p>Just in case they missed, I also mentioned it in the finalization post of this competition.<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308248#1699868\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308248#1699868</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1700645,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-22T06:05:17.827000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1695611,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2022-02-18T08:56:33.823000",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> - was looking at TIDE during this competition and know you posted a topic on it. While reading on it, this article was along your point:</p>\n<p><a href=\"https://towardsdatascience.com/a-better-map-for-object-detection-32662767d424\" target=\"_blank\">https://towardsdatascience.com/a-better-map-for-object-detection-32662767d424</a><br>\n\"Recently, AI pioneer Andrew Ng launched a campaign for data-centric AI where his main goal is to shift the focus of AI practitioners from model/algorithm development to the quality of the data they use to train the models.\" </p>\n<p>Cannot remember the link but read that improvements to e.g. COCO become harder now because the underlying annotations have issues. So certainly an area to watch.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1695624,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-18T09:15:45.673000",
          "content": "<p>I see. I had heard of the keyword, but did not know about the trend.<br>\nThanks for sharing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1698237,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-02-20T08:23:40.160000",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a><br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Our feedback is basically summarize to this topics description and Psi's comment[2].<br>\nWe are happy if you use these feedbacks to make future competition good.</p>\n<p>Thanks</p>\n<p>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1699350,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-21T05:57:00.847000",
          "content": "<p>I made summary of the opinions. Please check this also.<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1700595,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2022-02-22T05:30:55.397000",
          "content": "<p>Thank you! Certainly passing along to the impacted parties.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1696538,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-02-18T23:02:20.817000",
      "content": "<p>Thanks. But are you sure commenting to the right post? This is not a solution post.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1695492,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-02-18T07:20:26.507000",
      "content": "<p>It's a common complaint, and I'm fairly sure they pay very very close attention, it's just an immensely difficult problem to solve.  It may seem like these issues should be obvious, but appreciate that  it's usually after the fact and after the collective hive mind is hitting this data very hard with very sophisticated EDA techniques. </p>\n<p>We're usually given 2 or more submissions.  A baseline technique in most competitions is to use your highest CV and highest LB.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1695513,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-18T07:33:19.190000",
          "content": "<p>Thank you for commenting.</p>\n<p>I understand that this is a difficult problem to solve in general, but for this competition, I think it would have been possible to provide a higher quality data set.</p>\n<p>As you can see when you actually view the labels with the images, there are many frames where the labels are way off and distant COTS are not labeled.</p>\n<p>This kind of error is understandable if the labels were assigned automatically by a machine, but according to the paper of this competition, they were assigned by a human with assistance from a machine. I believe that human-assigned labels can be made more accurate.</p>\n<p>Also, I am not complaining. I am giving feedback so that the quality of this and other competitions will be higher.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1695551,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-02-18T07:56:54.640000",
          "content": "<p>Yeah, maybe there should be a pre contest period where contestants get a chance to comment and make suggestions about the data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1695621,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-18T09:07:46.783000",
          "content": "<p>I agree. That is one of an idea for preventing mismatch of data accuracy.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1695483": "To the Kaggle staff, the\nSCIRO team.\n\nThank you very much for organizing this competition.\n\nAs a participant, I would like to give you some feedback about the accuracy of the labels in the competition data.\n\nAs some of the other participants may have noticed, the training labels were inconsistent and there were many missing labels. My solution[1] showed an increase in the LB score by correcting the training labels, and the 3rd-place solution[2] suggests that there may be systematic errors in the fitness of the labels on the public and private datasets.\n\nThis suggests that the evaluation labels may have been as inaccurate as the training labels. (Of course, we participants cannot see the test data set, so we can only speculate.)\n\nI’m not requesting re-evaluating this competitions result. Even if some of the other participants argue that we should, we shouldn’t. That would amount to an ex post facto application of the law.\n\nInstead, I would like to insist that the next time you hold a similar competition, I want you to pay full attention and cost to the accuracy of the labels.\n\nIf one is a skilled kaggler, he/she know how to train models with good generalization performance even when the training dataset is inaccurate.\nHowever, as we know, when the evaluation dataset is inaccurate, it is inevitable that the winning model will be suboptimal.\n\nThis is not only unfair to the participants of the competition, but also disadvantageous to the host who will operate the model, as the model that will eventually be operated will not perform well enough.\n\nThis competition had a high prize money at stake. Therefore, the cost could have been spent on making the labels for evaluation (and preferably for training) more accurate.\n\nSo please think about this.\n\nFinally, I hope that this discussion will be used for constructive purposes.\n\n## Summary of Opinions\n\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699343\n\n## Reference\n\n* [1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307871\n* [2] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307707",
    "1695657": "Thanks for opening this post, I wanted to open a similar one.\n\nI think this competition would have been better, and this also applies to other competitions, if either of these two scenarios would apply:\n\n1. Test data is similar to training data, meaning that public LB data is similar to both training data and private test data.\n2. Test data is not similar to training data, but in this case public LB needs to be similar to private LB.\n\nIn the first case, we expect the test labels to be of similar quality as the training labels, this is a common setup for data science problems. This would allow to fully focus on building good models for the training labels with CV, and it will generalize well to test / production.\n\nIn the second case, we expect the test labels to be different. In this case training was labeled really messy, and it is absolutely a reasonable goal to make test predictions more precise / better. So test being more precise makes total sense. This would test more our ability to train on messy data (i.e. cheap data) and generalize to tight / precise boxes.\n\nUnfortunately, in this competition the public LB was completely useless. So it was the worst combination of my elaborations for above. It neither tests our ability to generalize well to training data, nor to new types of test labels. This is also a complete unrealistic real-world setup, as in industry data-science projects you would never build a hold-out dataset (public LB) that neither represents training data, nor your goal of generalization (future test data).\n\nActually, the whole metric only makes sense with more precise labels. The 0.8 IOU threshold has such a huge impact on the metric, and even with \"perfect\" predictions you cannot reach high scores on training data, as boxes are so inconsistent / messy. As soon as you have more consistent labels, the score goes up a lot, as we also saw on public LB. This allows us to train more consistent and precise models.",
    "1699078": "I agree with @tatamikenn and @philippsinger . My own benchmarks (which I'm sure many of you have already seen from my discussion post https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607) suggest that public LB's bboxes are systematically 10% tighter than that of the private LB.\n\nAt IOU 0.8, assuming the centers of the bboxes are exact, each dimension must be within ~10% error to score a positive hit (as 0.9 x 0.9 = 0.81). If this is not the case, the correct inference would be counted as a false positive, and will actually impact the F2 score. This punishes high-recall models that systematically over/underestimate the dimensions of the bboxes compared with the bbox annotations, which goes against the competition organiser's stated aim that \"in this case it makes sense to tolerate some false positives in order to ensure very few starfish are missed.\"\n\nI think manual labelling is always going to be impacted by human error, yet auto-labelling will produce systematic bias and will favor one architecture / method over another. A simple solution would simply have been to alter the competition metric to be less sensitive to IOU (e.g. F2 @ IOU 0.3 to 0.6 with step 0.05), or simply F2 @ IOU 0.5.",
    "1699343": "Thanks for posting a lot of opinions. I make summary of the opinions.\n(Please point out if I missed any other opinions.)\n\n## Problems\n\n1. Train labels are inaccurate -> that is not a problem\n2. (Probably) test labels are inaccurate\n3. Private and Public labels are dissimilar\n4. Evaluation metrics are too sensitive compared to the accuracy of labels\n\n## Solutions\n\n1. Make more accurate test labels (by @tatamikenn )\n2. Split public/private test data as they are similar[1] (by @philippsinger ) \n3. Use less public LB samples (e.g. 5%) to make public LB more incredible, and let participants focus more on trusting their CVs[2] (by @haqishen )\n4. Make competition metrics less sensitive[3] (e.g. F2@0.3:0.6, F2@0.5) (by @alexchwong )\n5. Having pre contest period where contestants get a chance to comment and make suggestions about the data[4] (by @kaggleqrdl )\n\n## Reference\n\n* [1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657\n* [2] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1698276\n* [3] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1699078\n* [4] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695551",
    "1695611": "@tatamikenn - was looking at TIDE during this competition and know you posted a topic on it. While reading on it, this article was along your point:\n\nhttps://towardsdatascience.com/a-better-map-for-object-detection-32662767d424\n\"Recently, AI pioneer Andrew Ng launched a campaign for data-centric AI where his main goal is to shift the focus of AI practitioners from model/algorithm development to the quality of the data they use to train the models.\" \n\nCannot remember the link but read that improvements to e.g. COCO become harder now because the underlying annotations have issues. So certainly an area to watch.  ",
    "1698237": "@addisonhoward\n@sohier \n\nOur feedback is basically summarize to this topics description and Psi's comment[2].\nWe are happy if you use these feedbacks to make future competition good.\n\nThanks\n\n[1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/308329#1695657",
    "1696538": "Thanks. But are you sure commenting to the right post? This is not a solution post.",
    "1695492": "It's a common complaint, and I'm fairly sure they pay very very close attention, it's just an immensely difficult problem to solve.  It may seem like these issues should be obvious, but appreciate that  it's usually after the fact and after the collective hive mind is hitting this data very hard with very sophisticated EDA techniques. \n\nWe're usually given 2 or more submissions.  A baseline technique in most competitions is to use your highest CV and highest LB.  "
  }
}