{
  "id": 71280,
  "title": "Techniques for mislabeled data",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71280",
  "author_name": "",
  "post_date": "2018-11-12T07:36:57.879252200Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am currently focusing on analyzing the given datasets, and have found some mislabeled datapoints, for example row 38, Why do Christian-conservative Americans call abortion murder? which label as sincere. The idea I have now is to drop them; however, it is difficult to find them all. What technique/method do you apply for this problem? </p>\n\n<p>Many thanks!</p>",
  "messages": [
    {
      "id": "419559",
      "postDate": "11/12/2018 07:36:57",
      "content": "<p>I am currently focusing on analyzing the given datasets, and have found some mislabeled datapoints, for example row 38, Why do Christian-conservative Americans call abortion murder? which label as sincere. The idea I have now is to drop them; however, it is difficult to find them all. What technique/method do you apply for this problem? </p>\n\n<p>Many thanks!</p>",
      "rawMarkdown": "I am currently focusing on analyzing the given datasets, and have found some mislabeled datapoints, for example row 38, Why do Christian-conservative Americans call abortion murder? which label as sincere. The idea I have now is to drop them; however, it is difficult to find them all. What technique/method do you apply for this problem? \n\nMany thanks!",
      "votes": null
    },
    {
      "id": "419576",
      "postDate": "11/12/2018 08:30:47",
      "content": "<p>Well, to begin with, how are you sure that this (or any other sentence) is miss-labeled? I mean it can certainly be insincere as you suggest but it is highly subjective and there exists no ground-truth...</p>",
      "rawMarkdown": "Well, to begin with, how are you sure that this (or any other sentence) is miss-labeled? I mean it can certainly be insincere as you suggest but it is highly subjective and there exists no ground-truth...",
      "votes": null
    },
    {
      "id": "420346",
      "postDate": "11/13/2018 13:34:25",
      "content": "<p>This example is at least potentially sincere... since the same labelling strategy is applied to the train and test set, I would imagine there is little to be gained by dropping it</p>",
      "rawMarkdown": "This example is at least potentially sincere... since the same labelling strategy is applied to the train and test set, I would imagine there is little to be gained by dropping it",
      "votes": null
    },
    {
      "id": "420748",
      "postDate": "11/14/2018 04:16:56",
      "content": "<p>I think there a lot of mislabeled data as I have make a detailed analysis here:\n<a href=\"https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru\">https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru</a></p>",
      "rawMarkdown": "I think there a lot of mislabeled data as I have make a detailed analysis here:\nhttps://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru",
      "votes": null
    },
    {
      "id": "420802",
      "postDate": "11/14/2018 06:31:17",
      "content": "<p>Indeed, I observed a math question been assigned as insincere on my t-SNE plot.  My current strategy for mislabeled data is using large batch size on nn model. </p>",
      "rawMarkdown": "Indeed, I observed a math question been assigned as insincere on my t-SNE plot.  My current strategy for mislabeled data is using large batch size on nn model.",
      "votes": null
    },
    {
      "id": "421012",
      "postDate": "11/14/2018 13:19:22",
      "content": "<p>How large is it, could you give some idea?  </p>\n\n<p>Reading your post above just see you mentioning about removing/correcting the mislabeled data. At first, I also agree. I think correcting them (if possible) will make the training easier. </p>\n\n<p>However, just realize that there is also a downside as the test dataset will also have noise. So, if we correcting the validation data, it will not have the same distribution as the test data, and so it will be difficult to judge our performance before submission.</p>",
      "rawMarkdown": "How large is it, could you give some idea?  \n\nReading your post above just see you mentioning about removing/correcting the mislabeled data. At first, I also agree. I think correcting them (if possible) will make the training easier. \n\nHowever, just realize that there is also a downside as the test dataset will also have noise. So, if we correcting the validation data, it will not have the same distribution as the test data, and so it will be difficult to judge our performance before submission.",
      "votes": null
    },
    {
      "id": "422247",
      "postDate": "11/16/2018 00:46:14",
      "content": "<p>I am using 256 to 1024 batch size.\nI also noticed that problem because the training set and the testing set both might gather in the same period. </p>\n\n<p>I am currently working on finding those (potential) mislabeled data and see why they get mislabeled. </p>",
      "rawMarkdown": "I am using 256 to 1024 batch size.\nI also noticed that problem because the training set and the testing set both might gather in the same period. \n\nI am currently working on finding those (potential) mislabeled data and see why they get mislabeled.",
      "votes": null
    },
    {
      "id": "422326",
      "postDate": "11/16/2018 04:22:11",
      "content": "<p>There are many mislabeled data points. But <code>Why do Christian-conservative Americans call abortion murder?</code> is a perfectly legit question. Abortion is a highly controversial topic in the U.S and this could be a highlight questions on Quora for many.</p>",
      "rawMarkdown": "There are many mislabeled data points. But `Why do Christian-conservative Americans call abortion murder?` is a perfectly legit question. Abortion is a highly controversial topic in the U.S and this could be a highlight questions on Quora for many.",
      "votes": null
    },
    {
      "id": "422361",
      "postDate": "11/16/2018 05:46:02",
      "content": "<p>Great! Looking forward to see that.</p>",
      "rawMarkdown": "Great! Looking forward to see that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 419576,
      "author_name": "georsara1",
      "author_url": "",
      "post_date": "11/12/2018 08:30:47",
      "content": "<p>Well, to begin with, how are you sure that this (or any other sentence) is miss-labeled? I mean it can certainly be insincere as you suggest but it is highly subjective and there exists no ground-truth...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420346,
      "author_name": "rdboyes",
      "author_url": "",
      "post_date": "11/13/2018 13:34:25",
      "content": "<p>This example is at least potentially sincere... since the same labelling strategy is applied to the train and test set, I would imagine there is little to be gained by dropping it</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420748,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "11/14/2018 04:16:56",
      "content": "<p>I think there a lot of mislabeled data as I have make a detailed analysis here:\n<a href=\"https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru\">https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 420802,
          "author_name": "konohayui",
          "author_url": "",
          "post_date": "11/14/2018 06:31:17",
          "content": "<p>Indeed, I observed a math question been assigned as insincere on my t-SNE plot.  My current strategy for mislabeled data is using large batch size on nn model. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 421012,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "11/14/2018 13:19:22",
          "content": "<p>How large is it, could you give some idea?  </p>\n\n<p>Reading your post above just see you mentioning about removing/correcting the mislabeled data. At first, I also agree. I think correcting them (if possible) will make the training easier. </p>\n\n<p>However, just realize that there is also a downside as the test dataset will also have noise. So, if we correcting the validation data, it will not have the same distribution as the test data, and so it will be difficult to judge our performance before submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422247,
          "author_name": "konohayui",
          "author_url": "",
          "post_date": "11/16/2018 00:46:14",
          "content": "<p>I am using 256 to 1024 batch size.\nI also noticed that problem because the training set and the testing set both might gather in the same period. </p>\n\n<p>I am currently working on finding those (potential) mislabeled data and see why they get mislabeled. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422361,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "11/16/2018 05:46:02",
          "content": "<p>Great! Looking forward to see that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 422326,
      "author_name": "sijunhe9248",
      "author_url": "",
      "post_date": "11/16/2018 04:22:11",
      "content": "<p>There are many mislabeled data points. But <code>Why do Christian-conservative Americans call abortion murder?</code> is a perfectly legit question. Abortion is a highly controversial topic in the U.S and this could be a highlight questions on Quora for many.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "419559": "I am currently focusing on analyzing the given datasets, and have found some mislabeled datapoints, for example row 38, Why do Christian-conservative Americans call abortion murder? which label as sincere. The idea I have now is to drop them; however, it is difficult to find them all. What technique/method do you apply for this problem? \n\nMany thanks!",
    "419576": "Well, to begin with, how are you sure that this (or any other sentence) is miss-labeled? I mean it can certainly be insincere as you suggest but it is highly subjective and there exists no ground-truth...",
    "420346": "This example is at least potentially sincere... since the same labelling strategy is applied to the train and test set, I would imagine there is little to be gained by dropping it",
    "420748": "I think there a lot of mislabeled data as I have make a detailed analysis here:\nhttps://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru",
    "420802": "Indeed, I observed a math question been assigned as insincere on my t-SNE plot.  My current strategy for mislabeled data is using large batch size on nn model.",
    "421012": "How large is it, could you give some idea?  \n\nReading your post above just see you mentioning about removing/correcting the mislabeled data. At first, I also agree. I think correcting them (if possible) will make the training easier. \n\nHowever, just realize that there is also a downside as the test dataset will also have noise. So, if we correcting the validation data, it will not have the same distribution as the test data, and so it will be difficult to judge our performance before submission.",
    "422247": "I am using 256 to 1024 batch size.\nI also noticed that problem because the training set and the testing set both might gather in the same period. \n\nI am currently working on finding those (potential) mislabeled data and see why they get mislabeled.",
    "422326": "There are many mislabeled data points. But `Why do Christian-conservative Americans call abortion murder?` is a perfectly legit question. Abortion is a highly controversial topic in the U.S and this could be a highlight questions on Quora for many.",
    "422361": "Great! Looking forward to see that."
  },
  "source": "meta"
}