{
  "id": 73159,
  "title": "How to deal with mislabelled problem",
  "url": "/competitions/quora-insincere-questions-classification/discussion/73159",
  "author_name": "",
  "post_date": "2018-11-30T11:07:26.663182400Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>In this competition, we can find the FP (false positive) and FN (false negative) samples in our trained models using cross-validation. Here, I find some high probability of FP and low probability of FN in my model.</p>\n\n<p>FP samples with high probability :\n- How can I seduce my aunt?\n- Why is Donald Trump being singled out and attacked when his actual policies are for the best of the country? Why do gay people attack him for not wanting Muslims in the country when Muslims believe that all gay people should be killed?\n- How do gay men have sex without getting feces on their penis?\n- How do I have sex with my maid servant?\n...</p>\n\n<p>FN samples with high probability:</p>\n\n<ul>\n<li>What is the difference between the words “kinetic” and “cinema”?</li>\n<li>What are the opportunities for an engineer on H4 EAD in Silicon Valley with 4 years of prior experience as an HR (Talent Management) in India? Is an MBA needed?</li>\n<li>How many SITA devices are present in Ramayana?</li>\n<li>What is the past tense of past tense?</li>\n<li>How can I connect with foreign engineering students? Who wants help on their assignments or wants homework help?\n...</li>\n</ul>\n\n<p>I think there are many mislabelled samples in our training set. How can we deal with them? I try to remove these texts, after considering. However, I cannot improve score in LB. I think that how to deal with mislabelled problem is key point to improve our score. My next step is find the difference between train and test distribution, so that I know test set with mislabelled samples similars with train set.</p>\n\n<p>Reference:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71429#latest-426109\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71429#latest-426109</a></li>\n<li><a href=\"https://www.kaggle.com/ratthachat/handle-overfitting-error-analysis-of-glove-gru\">https://www.kaggle.com/ratthachat/handle-overfitting-error-analysis-of-glove-gru</a></li>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72983\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72983</a></li>\n</ul>",
  "messages": [
    {
      "id": "430416",
      "postDate": "11/30/2018 11:07:26",
      "content": "<p>In this competition, we can find the FP (false positive) and FN (false negative) samples in our trained models using cross-validation. Here, I find some high probability of FP and low probability of FN in my model.</p>\n\n<p>FP samples with high probability :\n- How can I seduce my aunt?\n- Why is Donald Trump being singled out and attacked when his actual policies are for the best of the country? Why do gay people attack him for not wanting Muslims in the country when Muslims believe that all gay people should be killed?\n- How do gay men have sex without getting feces on their penis?\n- How do I have sex with my maid servant?\n...</p>\n\n<p>FN samples with high probability:</p>\n\n<ul>\n<li>What is the difference between the words “kinetic” and “cinema”?</li>\n<li>What are the opportunities for an engineer on H4 EAD in Silicon Valley with 4 years of prior experience as an HR (Talent Management) in India? Is an MBA needed?</li>\n<li>How many SITA devices are present in Ramayana?</li>\n<li>What is the past tense of past tense?</li>\n<li>How can I connect with foreign engineering students? Who wants help on their assignments or wants homework help?\n...</li>\n</ul>\n\n<p>I think there are many mislabelled samples in our training set. How can we deal with them? I try to remove these texts, after considering. However, I cannot improve score in LB. I think that how to deal with mislabelled problem is key point to improve our score. My next step is find the difference between train and test distribution, so that I know test set with mislabelled samples similars with train set.</p>\n\n<p>Reference:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71429#latest-426109\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71429#latest-426109</a></li>\n<li><a href=\"https://www.kaggle.com/ratthachat/handle-overfitting-error-analysis-of-glove-gru\">https://www.kaggle.com/ratthachat/handle-overfitting-error-analysis-of-glove-gru</a></li>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72983\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72983</a></li>\n</ul>",
      "rawMarkdown": "In this competition, we can find the FP (false positive) and FN (false negative) samples in our trained models using cross-validation. Here, I find some high probability of FP and low probability of FN in my model.\n\nFP samples with high probability :\n- How can I seduce my aunt?\n- Why is Donald Trump being singled out and attacked when his actual policies are for the best of the country? Why do gay people attack him for not wanting Muslims in the country when Muslims believe that all gay people should be killed?\n- How do gay men have sex without getting feces on their penis?\n- How do I have sex with my maid servant?\n...\n\nFN samples with high probability:\n\n- What is the difference between the words “kinetic” and “cinema”?\n- What are the opportunities for an engineer on H4 EAD in Silicon Valley with 4 years of prior experience as an HR (Talent Management) in India? Is an MBA needed?\n- How many SITA devices are present in Ramayana?\n- What is the past tense of past tense?\n- How can I connect with foreign engineering students? Who wants help on their assignments or wants homework help?\n...\n\nI think there are many mislabelled samples in our training set. How can we deal with them? I try to remove these texts, after considering. However, I cannot improve score in LB. I think that how to deal with mislabelled problem is key point to improve our score. My next step is find the difference between train and test distribution, so that I know test set with mislabelled samples similars with train set.\n\nReference:\n\n- https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71429#latest-426109\n- https://www.kaggle.com/ratthachat/handle-overfitting-error-analysis-of-glove-gru\n- https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72983",
      "votes": null
    },
    {
      "id": "430497",
      "postDate": "11/30/2018 13:26:08",
      "content": "<p>Good question! The next step I might correct these mislabelled samples artificially . But I doubt that there are also some wrong labels in the test data.</p>",
      "rawMarkdown": "Good question! The next step I might correct these mislabelled samples artificially . But I doubt that there are also some wrong labels in the test data.",
      "votes": null
    },
    {
      "id": "430519",
      "postDate": "11/30/2018 14:00:37",
      "content": "<p>I try to remove the \"insincere\" sample, but label as \"sincere\" such as \"someone sex with someone family members\". Then, I train the model again. However, this cannot improve the score in LB. I guess test set has mislabelled sample. Or, the mislabelled samples use as noise samples? </p>",
      "rawMarkdown": "I try to remove the \"insincere\" sample, but label as \"sincere\" such as \"someone sex with someone family members\". Then, I train the model again. However, this cannot improve the score in LB. I guess test set has mislabelled sample. Or, the mislabelled samples use as noise samples?",
      "votes": null
    },
    {
      "id": "430572",
      "postDate": "11/30/2018 16:12:21",
      "content": "<p>I agree, there are lots of mislabelled samples in the training set. I think we have these possible scenarios:</p>\n\n<ul>\n<li><p>The label noises are mistakes by humans. In this case, noise should be in the test set too (similar distribution). We probably could train a model to predict the noise too. But if the test set has noise too, then we won't make a better model for Quora, rather the winner will be the closest to their current \"bad\" solution.</p></li>\n<li><p>If the label noise is intentional by the sponsor, then their goal probably to force us to make a more robust model. There are some great papers about how to train with a noisy training set. Test set should clean in this case.</p></li>\n</ul>\n\n<p>I don't have time to do it myself, but I am really curious about the difference between the train and test distribution. <a href=\"/salonsai\">@salonsai</a>, a public kernel would be awesome :)</p>",
      "rawMarkdown": "I agree, there are lots of mislabelled samples in the training set. I think we have these possible scenarios:\n\n- The label noises are mistakes by humans. In this case, noise should be in the test set too (similar distribution). We probably could train a model to predict the noise too. But if the test set has noise too, then we won't make a better model for Quora, rather the winner will be the closest to their current \"bad\" solution.\n\n- If the label noise is intentional by the sponsor, then their goal probably to force us to make a more robust model. There are some great papers about how to train with a noisy training set. Test set should clean in this case.\n\nI don't have time to do it myself, but I am really curious about the difference between the train and test distribution. @salonsai, a public kernel would be awesome :)",
      "votes": null
    },
    {
      "id": "430585",
      "postDate": "11/30/2018 16:34:24",
      "content": "<p>Currently, I would assume that the private test set has similar properties to the current public test set. Otherwise the prediction problem would be quite different, e.g. removing noise from training then might make sense. In the end, what sense does it make to develop a model on a test set that you know is different to the real application? </p>",
      "rawMarkdown": "Currently, I would assume that the private test set has similar properties to the current public test set. Otherwise the prediction problem would be quite different, e.g. removing noise from training then might make sense. In the end, what sense does it make to develop a model on a test set that you know is different to the real application?",
      "votes": null
    },
    {
      "id": "431064",
      "postDate": "12/01/2018 15:10:11",
      "content": "<p>We knows that the training set has the mislabelled samples, but test set is uncertain whether it has mislabelled samples. I guess the samples of test set are manually labelled, so the distribution of training set and test set maybe different.</p>",
      "rawMarkdown": "We knows that the training set has the mislabelled samples, but test set is uncertain whether it has mislabelled samples. I guess the samples of test set are manually labelled, so the distribution of training set and test set maybe different.",
      "votes": null
    },
    {
      "id": "431492",
      "postDate": "12/02/2018 11:37:05",
      "content": "<p>Agreed, I was so comfused by the label, check this one, \"Why the [some F word] engineers see themselves superior than any other bachelor?\", the label is 0, R U kidding me, LOL...</p>",
      "rawMarkdown": "Agreed, I was so comfused by the label, check this one, \"Why the [some F word] engineers see themselves superior than any other bachelor?\", the label is 0, R U kidding me, LOL...",
      "votes": null
    },
    {
      "id": "432559",
      "postDate": "12/04/2018 03:29:56",
      "content": "<p>hi,any links of papers about how to train with a noisy training set</p>",
      "rawMarkdown": "hi,any links of papers about how to train with a noisy training set",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 430497,
      "author_name": "kaggleczs",
      "author_url": "",
      "post_date": "11/30/2018 13:26:08",
      "content": "<p>Good question! The next step I might correct these mislabelled samples artificially . But I doubt that there are also some wrong labels in the test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 430519,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "11/30/2018 14:00:37",
          "content": "<p>I try to remove the \"insincere\" sample, but label as \"sincere\" such as \"someone sex with someone family members\". Then, I train the model again. However, this cannot improve the score in LB. I guess test set has mislabelled sample. Or, the mislabelled samples use as noise samples? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 430572,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "11/30/2018 16:12:21",
      "content": "<p>I agree, there are lots of mislabelled samples in the training set. I think we have these possible scenarios:</p>\n\n<ul>\n<li><p>The label noises are mistakes by humans. In this case, noise should be in the test set too (similar distribution). We probably could train a model to predict the noise too. But if the test set has noise too, then we won't make a better model for Quora, rather the winner will be the closest to their current \"bad\" solution.</p></li>\n<li><p>If the label noise is intentional by the sponsor, then their goal probably to force us to make a more robust model. There are some great papers about how to train with a noisy training set. Test set should clean in this case.</p></li>\n</ul>\n\n<p>I don't have time to do it myself, but I am really curious about the difference between the train and test distribution. <a href=\"/salonsai\">@salonsai</a>, a public kernel would be awesome :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 430585,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/30/2018 16:34:24",
          "content": "<p>Currently, I would assume that the private test set has similar properties to the current public test set. Otherwise the prediction problem would be quite different, e.g. removing noise from training then might make sense. In the end, what sense does it make to develop a model on a test set that you know is different to the real application? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431064,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "12/01/2018 15:10:11",
          "content": "<p>We knows that the training set has the mislabelled samples, but test set is uncertain whether it has mislabelled samples. I guess the samples of test set are manually labelled, so the distribution of training set and test set maybe different.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 432559,
          "author_name": "mihailai",
          "author_url": "",
          "post_date": "12/04/2018 03:29:56",
          "content": "<p>hi,any links of papers about how to train with a noisy training set</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 431492,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "12/02/2018 11:37:05",
      "content": "<p>Agreed, I was so comfused by the label, check this one, \"Why the [some F word] engineers see themselves superior than any other bachelor?\", the label is 0, R U kidding me, LOL...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "430416": "In this competition, we can find the FP (false positive) and FN (false negative) samples in our trained models using cross-validation. Here, I find some high probability of FP and low probability of FN in my model.\n\nFP samples with high probability :\n- How can I seduce my aunt?\n- Why is Donald Trump being singled out and attacked when his actual policies are for the best of the country? Why do gay people attack him for not wanting Muslims in the country when Muslims believe that all gay people should be killed?\n- How do gay men have sex without getting feces on their penis?\n- How do I have sex with my maid servant?\n...\n\nFN samples with high probability:\n\n- What is the difference between the words “kinetic” and “cinema”?\n- What are the opportunities for an engineer on H4 EAD in Silicon Valley with 4 years of prior experience as an HR (Talent Management) in India? Is an MBA needed?\n- How many SITA devices are present in Ramayana?\n- What is the past tense of past tense?\n- How can I connect with foreign engineering students? Who wants help on their assignments or wants homework help?\n...\n\nI think there are many mislabelled samples in our training set. How can we deal with them? I try to remove these texts, after considering. However, I cannot improve score in LB. I think that how to deal with mislabelled problem is key point to improve our score. My next step is find the difference between train and test distribution, so that I know test set with mislabelled samples similars with train set.\n\nReference:\n\n- https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71429#latest-426109\n- https://www.kaggle.com/ratthachat/handle-overfitting-error-analysis-of-glove-gru\n- https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/72983",
    "430497": "Good question! The next step I might correct these mislabelled samples artificially . But I doubt that there are also some wrong labels in the test data.",
    "430519": "I try to remove the \"insincere\" sample, but label as \"sincere\" such as \"someone sex with someone family members\". Then, I train the model again. However, this cannot improve the score in LB. I guess test set has mislabelled sample. Or, the mislabelled samples use as noise samples?",
    "430572": "I agree, there are lots of mislabelled samples in the training set. I think we have these possible scenarios:\n\n- The label noises are mistakes by humans. In this case, noise should be in the test set too (similar distribution). We probably could train a model to predict the noise too. But if the test set has noise too, then we won't make a better model for Quora, rather the winner will be the closest to their current \"bad\" solution.\n\n- If the label noise is intentional by the sponsor, then their goal probably to force us to make a more robust model. There are some great papers about how to train with a noisy training set. Test set should clean in this case.\n\nI don't have time to do it myself, but I am really curious about the difference between the train and test distribution. @salonsai, a public kernel would be awesome :)",
    "430585": "Currently, I would assume that the private test set has similar properties to the current public test set. Otherwise the prediction problem would be quite different, e.g. removing noise from training then might make sense. In the end, what sense does it make to develop a model on a test set that you know is different to the real application?",
    "431064": "We knows that the training set has the mislabelled samples, but test set is uncertain whether it has mislabelled samples. I guess the samples of test set are manually labelled, so the distribution of training set and test set maybe different.",
    "431492": "Agreed, I was so comfused by the label, check this one, \"Why the [some F word] engineers see themselves superior than any other bachelor?\", the label is 0, R U kidding me, LOL...",
    "432559": "hi,any links of papers about how to train with a noisy training set"
  },
  "source": "meta"
}