{
  "id": 201471,
  "title": "Noise in train and private test sets",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201471",
  "author_name": "",
  "post_date": "2020-12-05T07:03:52.437116600Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Can organizers please confirm how did they split images between the train and test sets. ?<br>\nGiven there were so many discussions about how noisy the training images are, it would be useful to know if test images were given extra processing/care or it was just a random split. <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>Sorry if this was already asked/provided before. </p>",
  "messages": [
    {
      "id": "1102670",
      "postDate": "12/05/2020 07:03:52",
      "content": "<p>Can organizers please confirm how did they split images between the train and test sets. ?<br>\nGiven there were so many discussions about how noisy the training images are, it would be useful to know if test images were given extra processing/care or it was just a random split. <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>Sorry if this was already asked/provided before. </p>",
      "rawMarkdown": "Can organizers please confirm how did they split images between the train and test sets. ?\nGiven there were so many discussions about how noisy the training images are, it would be useful to know if test images were given extra processing/care or it was just a random split. @juliaelliott \n\nSorry if this was already asked/provided before.",
      "votes": null
    },
    {
      "id": "1102866",
      "postDate": "12/05/2020 12:13:23",
      "content": "<p>You may doubt kaggle will disclose nothing about the data splitting while the competition is ongoing. </p>\n<p>IMHO, presence of noise make the competition fun, challenging and close to real life stuffs…compared to other oversimplified classification problems.   </p>\n<p>Keep in mind DL models right now have gain lot of  maturity , and a part of the goal here is to help to build robust models that can handle noisy/mislabelled datasets. </p>",
      "rawMarkdown": "You may doubt kaggle will disclose nothing about the data splitting while the competition is ongoing. \n\nIMHO, presence of noise make the competition fun, challenging and close to real life stuffs...compared to other oversimplified classification problems.   \n\nKeep in mind DL models right now have gain lot of  maturity , and a part of the goal here is to help to build robust models that can handle noisy/mislabelled datasets.",
      "votes": null
    },
    {
      "id": "1103180",
      "postDate": "12/05/2020 17:46:12",
      "content": "<p>Thanks for this viewpoint. I agree that real life datasets will be like this and this makes the compitition challenging. </p>\n<p>However, in real life you evaluate model performance on hand selected good quality data so I was trying to get some info.</p>",
      "rawMarkdown": "Thanks for this viewpoint. I agree that real life datasets will be like this and this makes the compitition challenging. \n\nHowever, in real life you evaluate model performance on hand selected good quality data so I was trying to get some info.",
      "votes": null
    },
    {
      "id": "1106592",
      "postDate": "12/09/2020 00:08:30",
      "content": "<p>This is correct, we will not be disclosing the data split method. Best of luck!</p>",
      "rawMarkdown": "This is correct, we will not be disclosing the data split method. Best of luck!",
      "votes": null
    },
    {
      "id": "1106726",
      "postDate": "12/09/2020 04:21:30",
      "content": "<p>\"However, in real life you evaluate model performance on hand selected good quality data \"</p>\n<p>Not true. In real application, sometimes clean data is impossible!</p>\n<p>it is possible to evaluation real performance of a learned model using noisy testset, if the noise is independent or if the noise characteristics are known </p>\n<p>Evaluating Classifiers by Means of Test Data with Noisy Labels <br>\n<a href=\"https://www.ijcai.org/Proceedings/03/Papers/076.pdf\" target=\"_blank\">https://www.ijcai.org/Proceedings/03/Papers/076.pdf</a></p>\n<p>it really depends on how you design your application and the requirement of the model for that application. then you decide what kind of test data is required …. and sometimes you have no choice to work with dirty data. that of course will also change the way how you interprete results, etc</p>",
      "rawMarkdown": "\"However, in real life you evaluate model performance on hand selected good quality data \"\n\nNot true. In real application, sometimes clean data is impossible!\n\nit is possible to evaluation real performance of a learned model using noisy testset, if the noise is independent or if the noise characteristics are known \n\nEvaluating Classifiers by Means of Test Data with Noisy Labels \nhttps://www.ijcai.org/Proceedings/03/Papers/076.pdf\n\nit really depends on how you design your application and the requirement of the model for that application. then you decide what kind of test data is required .... and sometimes you have no choice to work with dirty data. that of course will also change the way how you interprete results, etc",
      "votes": null
    },
    {
      "id": "1148959",
      "postDate": "01/11/2021 13:51:53",
      "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a>  atleast if you can confirm if the test data is more accurate  than train or similar noisy. In past we ahve seen competition where train data is noisy but test is not,so competition  was testing on how participants handle te noisy data. </p>",
      "rawMarkdown": "juliaelliott  atleast if you can confirm if the test data is more accurate  than train or similar noisy. In past we ahve seen competition where train data is noisy but test is not,so competition  was testing on how participants handle te noisy data.",
      "votes": null
    },
    {
      "id": "1149125",
      "postDate": "01/11/2021 15:53:06",
      "content": "<p>that is right. We are not complaining about the noise in the training data. <br>\nTo make the model really useful in the real world, it would be beneficial to know if the test data is more vetted/reviewed and can be expected to be of better quality than the training data. </p>",
      "rawMarkdown": "that is right. We are not complaining about the noise in the training data. \nTo make the model really useful in the real world, it would be beneficial to know if the test data is more vetted/reviewed and can be expected to be of better quality than the training data.",
      "votes": null
    },
    {
      "id": "1185718",
      "postDate": "02/04/2021 10:12:53",
      "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> Hello, is it possible to know if the testing data has better quality regarding the labellisation, or does it have noise such as the training set ?</p>",
      "rawMarkdown": "juliaelliott Hello, is it possible to know if the testing data has better quality regarding the labellisation, or does it have noise such as the training set ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1102866,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "12/05/2020 12:13:23",
      "content": "<p>You may doubt kaggle will disclose nothing about the data splitting while the competition is ongoing. </p>\n<p>IMHO, presence of noise make the competition fun, challenging and close to real life stuffs…compared to other oversimplified classification problems.   </p>\n<p>Keep in mind DL models right now have gain lot of  maturity , and a part of the goal here is to help to build robust models that can handle noisy/mislabelled datasets. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1103180,
          "author_name": "krisho007",
          "author_url": "",
          "post_date": "12/05/2020 17:46:12",
          "content": "<p>Thanks for this viewpoint. I agree that real life datasets will be like this and this makes the compitition challenging. </p>\n<p>However, in real life you evaluate model performance on hand selected good quality data so I was trying to get some info.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1106592,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "12/09/2020 00:08:30",
          "content": "<p>This is correct, we will not be disclosing the data split method. Best of luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1106726,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/09/2020 04:21:30",
          "content": "<p>\"However, in real life you evaluate model performance on hand selected good quality data \"</p>\n<p>Not true. In real application, sometimes clean data is impossible!</p>\n<p>it is possible to evaluation real performance of a learned model using noisy testset, if the noise is independent or if the noise characteristics are known </p>\n<p>Evaluating Classifiers by Means of Test Data with Noisy Labels <br>\n<a href=\"https://www.ijcai.org/Proceedings/03/Papers/076.pdf\" target=\"_blank\">https://www.ijcai.org/Proceedings/03/Papers/076.pdf</a></p>\n<p>it really depends on how you design your application and the requirement of the model for that application. then you decide what kind of test data is required …. and sometimes you have no choice to work with dirty data. that of course will also change the way how you interprete results, etc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1148959,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/11/2021 13:51:53",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a>  atleast if you can confirm if the test data is more accurate  than train or similar noisy. In past we ahve seen competition where train data is noisy but test is not,so competition  was testing on how participants handle te noisy data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149125,
          "author_name": "krisho007",
          "author_url": "",
          "post_date": "01/11/2021 15:53:06",
          "content": "<p>that is right. We are not complaining about the noise in the training data. <br>\nTo make the model really useful in the real world, it would be beneficial to know if the test data is more vetted/reviewed and can be expected to be of better quality than the training data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1185718,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "02/04/2021 10:12:53",
      "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> Hello, is it possible to know if the testing data has better quality regarding the labellisation, or does it have noise such as the training set ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1102670": "Can organizers please confirm how did they split images between the train and test sets. ?\nGiven there were so many discussions about how noisy the training images are, it would be useful to know if test images were given extra processing/care or it was just a random split. @juliaelliott \n\nSorry if this was already asked/provided before.",
    "1102866": "You may doubt kaggle will disclose nothing about the data splitting while the competition is ongoing. \n\nIMHO, presence of noise make the competition fun, challenging and close to real life stuffs...compared to other oversimplified classification problems.   \n\nKeep in mind DL models right now have gain lot of  maturity , and a part of the goal here is to help to build robust models that can handle noisy/mislabelled datasets.",
    "1103180": "Thanks for this viewpoint. I agree that real life datasets will be like this and this makes the compitition challenging. \n\nHowever, in real life you evaluate model performance on hand selected good quality data so I was trying to get some info.",
    "1106592": "This is correct, we will not be disclosing the data split method. Best of luck!",
    "1106726": "\"However, in real life you evaluate model performance on hand selected good quality data \"\n\nNot true. In real application, sometimes clean data is impossible!\n\nit is possible to evaluation real performance of a learned model using noisy testset, if the noise is independent or if the noise characteristics are known \n\nEvaluating Classifiers by Means of Test Data with Noisy Labels \nhttps://www.ijcai.org/Proceedings/03/Papers/076.pdf\n\nit really depends on how you design your application and the requirement of the model for that application. then you decide what kind of test data is required .... and sometimes you have no choice to work with dirty data. that of course will also change the way how you interprete results, etc",
    "1148959": "juliaelliott  atleast if you can confirm if the test data is more accurate  than train or similar noisy. In past we ahve seen competition where train data is noisy but test is not,so competition  was testing on how participants handle te noisy data.",
    "1149125": "that is right. We are not complaining about the noise in the training data. \nTo make the model really useful in the real world, it would be beneficial to know if the test data is more vetted/reviewed and can be expected to be of better quality than the training data.",
    "1185718": "juliaelliott Hello, is it possible to know if the testing data has better quality regarding the labellisation, or does it have noise such as the training set ?"
  },
  "source": "meta"
}