{
  "id": 210304,
  "title": "Quality of Test Set",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210304",
  "author_name": "",
  "post_date": "2021-01-10T10:53:06.155854400Z",
  "votes": -3,
  "comment_count": 3,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/competition\" target=\"_blank\">@competition</a>_organizers, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> is there any information you could provide us on the test set? Eg. will the test set have better quality labels than the training set? It doesn't make any sense to evaluate a model on a noisy test set, as it would punish models that correctly classify samples.</p>\n<p>Edit #1: A similar question was asked <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209318\" target=\"_blank\">here</a> but I want to explicitly ask about the test set. </p>",
  "messages": [
    {
      "id": "1147192",
      "postDate": "01/10/2021 10:53:06",
      "content": "<p><a href=\"https://www.kaggle.com/competition\" target=\"_blank\">@competition</a>_organizers, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> is there any information you could provide us on the test set? Eg. will the test set have better quality labels than the training set? It doesn't make any sense to evaluate a model on a noisy test set, as it would punish models that correctly classify samples.</p>\n<p>Edit #1: A similar question was asked <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209318\" target=\"_blank\">here</a> but I want to explicitly ask about the test set. </p>",
      "rawMarkdown": "competition_organizers, @sohier is there any information you could provide us on the test set? Eg. will the test set have better quality labels than the training set? It doesn't make any sense to evaluate a model on a noisy test set, as it would punish models that correctly classify samples.\n\nEdit #1: A similar question was asked [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209318) but I want to explicitly ask about the test set.",
      "votes": null
    },
    {
      "id": "1148242",
      "postDate": "01/11/2021 02:53:27",
      "content": "<p>It would be one of the surprises of 2021 if the organizers provide information in this post regarding the noise level.  But the year is young so maybe they will throw you a little chat noise.</p>\n<p>Been here almost 3 years - participated in lots of competitions.  For the most part the test set has been what appeared to be a random sample selection of all available data where any difference between train and test is purely based on statistics and chance.  In a recent competition it did appear that an effort had been made by the organizers to insure cleaner labels in the test set.</p>\n<p>I don't expect to see that effort having been taken for this competition.  The 2019 version of Cassava had a very similar appearance.  </p>\n<p>This is the thing - the real world is noisy.  It probably does not make sense to award a model that needs a world without noise.</p>",
      "rawMarkdown": "It would be one of the surprises of 2021 if the organizers provide information in this post regarding the noise level.  But the year is young so maybe they will throw you a little chat noise.\n\nBeen here almost 3 years - participated in lots of competitions.  For the most part the test set has been what appeared to be a random sample selection of all available data where any difference between train and test is purely based on statistics and chance.  In a recent competition it did appear that an effort had been made by the organizers to insure cleaner labels in the test set.\n\nI don't expect to see that effort having been taken for this competition.  The 2019 version of Cassava had a very similar appearance.  \n\nThis is the thing - the real world is noisy.  It probably does not make sense to award a model that needs a world without noise.",
      "votes": null
    },
    {
      "id": "1148772",
      "postDate": "01/11/2021 11:27:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> <br>\nBut I want to point out that a less noisy test set would award a model that is accurate, not award a model that is less robust (needs a world without noise).<br>\nFor example, the test set might label a diseased leaf as healthy. This would punish a model that is accurate on that particular data point and reward one that guessed it correctly by chance.</p>",
      "rawMarkdown": "Thanks @pcjimmmy \nBut I want to point out that a less noisy test set would award a model that is accurate, not award a model that is less robust (needs a world without noise).\nFor example, the test set might label a diseased leaf as healthy. This would punish a model that is accurate on that particular data point and reward one that guessed it correctly by chance.",
      "votes": null
    },
    {
      "id": "1148961",
      "postDate": "01/11/2021 13:52:15",
      "content": "<p>Noisy data are more outliers than the majority of the data. Thus, a good model will perform better on the whole test data (clean +noisy data) than a model which overfitted on noise.</p>",
      "rawMarkdown": "Noisy data are more outliers than the majority of the data. Thus, a good model will perform better on the whole test data (clean +noisy data) than a model which overfitted on noise.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1148242,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/11/2021 02:53:27",
      "content": "<p>It would be one of the surprises of 2021 if the organizers provide information in this post regarding the noise level.  But the year is young so maybe they will throw you a little chat noise.</p>\n<p>Been here almost 3 years - participated in lots of competitions.  For the most part the test set has been what appeared to be a random sample selection of all available data where any difference between train and test is purely based on statistics and chance.  In a recent competition it did appear that an effort had been made by the organizers to insure cleaner labels in the test set.</p>\n<p>I don't expect to see that effort having been taken for this competition.  The 2019 version of Cassava had a very similar appearance.  </p>\n<p>This is the thing - the real world is noisy.  It probably does not make sense to award a model that needs a world without noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148772,
          "author_name": "dumbluck",
          "author_url": "",
          "post_date": "01/11/2021 11:27:47",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> <br>\nBut I want to point out that a less noisy test set would award a model that is accurate, not award a model that is less robust (needs a world without noise).<br>\nFor example, the test set might label a diseased leaf as healthy. This would punish a model that is accurate on that particular data point and reward one that guessed it correctly by chance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148961,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "01/11/2021 13:52:15",
          "content": "<p>Noisy data are more outliers than the majority of the data. Thus, a good model will perform better on the whole test data (clean +noisy data) than a model which overfitted on noise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1147192": "competition_organizers, @sohier is there any information you could provide us on the test set? Eg. will the test set have better quality labels than the training set? It doesn't make any sense to evaluate a model on a noisy test set, as it would punish models that correctly classify samples.\n\nEdit #1: A similar question was asked [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209318) but I want to explicitly ask about the test set.",
    "1148242": "It would be one of the surprises of 2021 if the organizers provide information in this post regarding the noise level.  But the year is young so maybe they will throw you a little chat noise.\n\nBeen here almost 3 years - participated in lots of competitions.  For the most part the test set has been what appeared to be a random sample selection of all available data where any difference between train and test is purely based on statistics and chance.  In a recent competition it did appear that an effort had been made by the organizers to insure cleaner labels in the test set.\n\nI don't expect to see that effort having been taken for this competition.  The 2019 version of Cassava had a very similar appearance.  \n\nThis is the thing - the real world is noisy.  It probably does not make sense to award a model that needs a world without noise.",
    "1148772": "Thanks @pcjimmmy \nBut I want to point out that a less noisy test set would award a model that is accurate, not award a model that is less robust (needs a world without noise).\nFor example, the test set might label a diseased leaf as healthy. This would punish a model that is accurate on that particular data point and reward one that guessed it correctly by chance.",
    "1148961": "Noisy data are more outliers than the majority of the data. Thus, a good model will perform better on the whole test data (clean +noisy data) than a model which overfitted on noise."
  },
  "source": "meta"
}