{
  "id": 76909,
  "title": "False Analysis : Keywords Bias",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76909",
  "author_name": "",
  "post_date": "2019-01-07T20:05:32.305772900Z",
  "votes": 23,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I made a false analysis to understand where my model misses it the most, I found most of the wrong predictions related to two keywords after clean all stopwords [Trump and People]:</p>\n\n<p>For <strong>Sincere training data</strong> that the model marked True:</p>\n\n<ul>\n<li>People: 5282 <em>missed training label contains the word</em></li>\n<li>Trump: 2871 </li>\n</ul>\n\n<p>For <strong>InSincere training data</strong> that the model marked False:</p>\n\n<ul>\n<li>People: 2488</li>\n<li>Trump: 1148</li>\n</ul>\n\n<p><em>update: True means insincere, False means sincere</em></p>",
  "messages": [
    {
      "id": "451874",
      "postDate": "01/07/2019 20:05:32",
      "content": "<p>I made a false analysis to understand where my model misses it the most, I found most of the wrong predictions related to two keywords after clean all stopwords [Trump and People]:</p>\n\n<p>For <strong>Sincere training data</strong> that the model marked True:</p>\n\n<ul>\n<li>People: 5282 <em>missed training label contains the word</em></li>\n<li>Trump: 2871 </li>\n</ul>\n\n<p>For <strong>InSincere training data</strong> that the model marked False:</p>\n\n<ul>\n<li>People: 2488</li>\n<li>Trump: 1148</li>\n</ul>\n\n<p><em>update: True means insincere, False means sincere</em></p>",
      "rawMarkdown": "I made a false analysis to understand where my model misses it the most, I found most of the wrong predictions related to two keywords after clean all stopwords [Trump and People]:\n\nFor **Sincere training data** that the model marked True:\n\n - People: 5282 *missed training label contains the word*\n - Trump: 2871 \n\nFor **InSincere training data** that the model marked False:\n\n - People: 2488\n - Trump: 1148\n\n*update: True means insincere, False means sincere*",
      "votes": null
    },
    {
      "id": "451998",
      "postDate": "01/08/2019 02:40:51",
      "content": "<p>maybe these words are difficult to distinguish</p>",
      "rawMarkdown": "maybe these words are difficult to distinguish",
      "votes": null
    },
    {
      "id": "452010",
      "postDate": "01/08/2019 03:05:10",
      "content": "<p>I believe, the association of nearby words matters more than singe words. Right ?</p>",
      "rawMarkdown": "I believe, the association of nearby words matters more than singe words. Right ?",
      "votes": null
    },
    {
      "id": "452024",
      "postDate": "01/08/2019 03:44:06",
      "content": "<p>Yes, an association of words matters more, but the same word used in similar ways for both classifications, I checked many examples and found that RNN model classification was more sense for me than some of the provided targets. \nFor example, the model classified following insincere:</p>\n\n<ul>\n<li>( why are bootythongs on women so sexy ( to guys ) ?) </li>\n</ul>\n\n<p>However, Quora classification was sincere. \nI am trying to understand how really classification is done so at least I can work on some pre-processing to enhance.</p>",
      "rawMarkdown": "Yes, an association of words matters more, but the same word used in similar ways for both classifications, I checked many examples and found that RNN model classification was more sense for me than some of the provided targets. \nFor example, the model classified following insincere:\n\n - ( why are bootythongs on women so sexy ( to guys ) ?) \n\nHowever, Quora classification was sincere. \nI am trying to understand how really classification is done so at least I can work on some pre-processing to enhance.",
      "votes": null
    },
    {
      "id": "452048",
      "postDate": "01/08/2019 05:32:36",
      "content": "<p>The training set is very noisy, and I guess that's the case with the public test set too. And that's perfectly okay.</p>\n\n<p>However, the private set has to be clean, if it isn't then frankly speaking this whole competition is a waste of everyone's time.</p>\n\n<p>Also in my opinion the above question <em>is</em> sincere.</p>",
      "rawMarkdown": "The training set is very noisy, and I guess that's the case with the public test set too. And that's perfectly okay.\n\nHowever, the private set has to be clean, if it isn't then frankly speaking this whole competition is a waste of everyone's time.\n\nAlso in my opinion the above question *is* sincere.",
      "votes": null
    },
    {
      "id": "452056",
      "postDate": "01/08/2019 05:54:06",
      "content": "<p>Agree the private test data should be clean. another challenging thing in here the classification somehow opinion based not as direct as toxic classification. So hope the target generated with the same distribution and clear guidelines to labelers. \nIn the end, if generated with the same distribution it is ok.</p>",
      "rawMarkdown": "Agree the private test data should be clean. another challenging thing in here the classification somehow opinion based not as direct as toxic classification. So hope the target generated with the same distribution and clear guidelines to labelers. \nIn the end, if generated with the same distribution it is ok.",
      "votes": null
    },
    {
      "id": "452093",
      "postDate": "01/08/2019 07:19:55",
      "content": "<p>I would argue otherwise that the private data should not be clean and completely similarly distributed as public LB. Otherwise you would need to follow quite a different approach like removing wrongly labeled data and the whole 3 months are a waste of time.</p>",
      "rawMarkdown": "I would argue otherwise that the private data should not be clean and completely similarly distributed as public LB. Otherwise you would need to follow quite a different approach like removing wrongly labeled data and the whole 3 months are a waste of time.",
      "votes": null
    },
    {
      "id": "452610",
      "postDate": "01/09/2019 00:18:26",
      "content": "<p>Interesting thing, I once tried to add 2 cols in the auxiliary input to mark whether \"Trump\" and \"gay\" appears in a sentence (since they seem to have good discrimination to some extend, and the false analysis also). However, the result was a miserable story...</p>",
      "rawMarkdown": "Interesting thing, I once tried to add 2 cols in the auxiliary input to mark whether \"Trump\" and \"gay\" appears in a sentence (since they seem to have good discrimination to some extend, and the false analysis also). However, the result was a miserable story...",
      "votes": null
    },
    {
      "id": "453009",
      "postDate": "01/09/2019 13:53:47",
      "content": "<p>I agree with Psi, and if you read between the lines a bit on this</p>\n\n<blockquote>\n  <p>Note that the distribution of questions in the dataset should not be taken to be representative of the distribution of questions asked on Quora. This is, in part, because of the combination of sampling procedures and sanitization measures that have been applied to the final dataset.</p>\n</blockquote>\n\n<p>it sounds like the dataset used has been sanitized and then split up. Certainly it would be weird if that wasn't what was done.</p>\n\n<p>I am curious about this however</p>\n\n<blockquote>\n  <p>To date, Quora has employed both machine learning and manual review to address this problem.</p>\n</blockquote>\n\n<p>I hope this means we aren't trying to build a model to fit another model :/</p>",
      "rawMarkdown": "I agree with Psi, and if you read between the lines a bit on this\n\n&gt; Note that the distribution of questions in the dataset should not be taken to be representative of the distribution of questions asked on Quora. This is, in part, because of the combination of sampling procedures and sanitization measures that have been applied to the final dataset.\n\nit sounds like the dataset used has been sanitized and then split up. Certainly it would be weird if that wasn't what was done.\n\nI am curious about this however\n\n&gt; To date, Quora has employed both machine learning and manual review to address this problem.\n\nI hope this means we aren't trying to build a model to fit another model :/",
      "votes": null
    },
    {
      "id": "453578",
      "postDate": "01/10/2019 12:12:40",
      "content": "<p>Thank You for an interesting observation! Maybe You have some ideas on how to prevent this, maybe some kind of changing weights or post-processing?</p>",
      "rawMarkdown": "Thank You for an interesting observation! Maybe You have some ideas on how to prevent this, maybe some kind of changing weights or post-processing?",
      "votes": null
    },
    {
      "id": "453737",
      "postDate": "01/10/2019 17:38:01",
      "content": "<p>There are many things that can be done to overcome this, such as but not limited to 1- increase these values in the validation data and check wich network changes reduce the error. 2- augment more examples in the training data that include keywords with a higher error for the network to learn from it. 3- Remove or fix false labeled data (if any) in the training set.</p>",
      "rawMarkdown": "There are many things that can be done to overcome this, such as but not limited to 1- increase these values in the validation data and check wich network changes reduce the error. 2- augment more examples in the training data that include keywords with a higher error for the network to learn from it. 3- Remove or fix false labeled data (if any) in the training set.",
      "votes": null
    },
    {
      "id": "455139",
      "postDate": "01/13/2019 04:00:02",
      "content": "<p>I totally agree the Psi. Train set and test set (public and private) have noisy. So we should make our model robust. Anyway, I find my model is sensitive with \"gay\", \"sexy\"...</p>",
      "rawMarkdown": "I totally agree the Psi. Train set and test set (public and private) have noisy. So we should make our model robust. Anyway, I find my model is sensitive with \"gay\", \"sexy\"...",
      "votes": null
    },
    {
      "id": "455217",
      "postDate": "01/13/2019 09:45:55",
      "content": "<p>i spent 2 days to clean the dataset, such as replacing the strange character, correct misspelling and remove lots of strange symbols.  but it will harm the result under the same architecture with the same random seed.</p>",
      "rawMarkdown": "i spent 2 days to clean the dataset, such as replacing the strange character, correct misspelling and remove lots of strange symbols.  but it will harm the result under the same architecture with the same random seed.",
      "votes": null
    },
    {
      "id": "460178",
      "postDate": "01/23/2019 06:03:36",
      "content": "<p>Augmenting the training set works out to something similar to focal loss from what I understand. Is that a fair deduction?</p>",
      "rawMarkdown": "Augmenting the training set works out to something similar to focal loss from what I understand. Is that a fair deduction?",
      "votes": null
    },
    {
      "id": "460219",
      "postDate": "01/23/2019 08:07:22",
      "content": "<p>Maybe substituting the word Trump with 'president' and do similarly with other words, may bring better results though.</p>",
      "rawMarkdown": "Maybe substituting the word Trump with 'president' and do similarly with other words, may bring better results though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 451998,
      "author_name": "chenhao2334",
      "author_url": "",
      "post_date": "01/08/2019 02:40:51",
      "content": "<p>maybe these words are difficult to distinguish</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 452010,
      "author_name": "s4sarath",
      "author_url": "",
      "post_date": "01/08/2019 03:05:10",
      "content": "<p>I believe, the association of nearby words matters more than singe words. Right ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 452024,
      "author_name": "jaguar00",
      "author_url": "",
      "post_date": "01/08/2019 03:44:06",
      "content": "<p>Yes, an association of words matters more, but the same word used in similar ways for both classifications, I checked many examples and found that RNN model classification was more sense for me than some of the provided targets. \nFor example, the model classified following insincere:</p>\n\n<ul>\n<li>( why are bootythongs on women so sexy ( to guys ) ?) </li>\n</ul>\n\n<p>However, Quora classification was sincere. \nI am trying to understand how really classification is done so at least I can work on some pre-processing to enhance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 452048,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "01/08/2019 05:32:36",
          "content": "<p>The training set is very noisy, and I guess that's the case with the public test set too. And that's perfectly okay.</p>\n\n<p>However, the private set has to be clean, if it isn't then frankly speaking this whole competition is a waste of everyone's time.</p>\n\n<p>Also in my opinion the above question <em>is</em> sincere.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452056,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "01/08/2019 05:54:06",
          "content": "<p>Agree the private test data should be clean. another challenging thing in here the classification somehow opinion based not as direct as toxic classification. So hope the target generated with the same distribution and clear guidelines to labelers. \nIn the end, if generated with the same distribution it is ok.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452093,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "01/08/2019 07:19:55",
          "content": "<p>I would argue otherwise that the private data should not be clean and completely similarly distributed as public LB. Otherwise you would need to follow quite a different approach like removing wrongly labeled data and the whole 3 months are a waste of time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 453009,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "01/09/2019 13:53:47",
          "content": "<p>I agree with Psi, and if you read between the lines a bit on this</p>\n\n<blockquote>\n  <p>Note that the distribution of questions in the dataset should not be taken to be representative of the distribution of questions asked on Quora. This is, in part, because of the combination of sampling procedures and sanitization measures that have been applied to the final dataset.</p>\n</blockquote>\n\n<p>it sounds like the dataset used has been sanitized and then split up. Certainly it would be weird if that wasn't what was done.</p>\n\n<p>I am curious about this however</p>\n\n<blockquote>\n  <p>To date, Quora has employed both machine learning and manual review to address this problem.</p>\n</blockquote>\n\n<p>I hope this means we aren't trying to build a model to fit another model :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 455139,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "01/13/2019 04:00:02",
          "content": "<p>I totally agree the Psi. Train set and test set (public and private) have noisy. So we should make our model robust. Anyway, I find my model is sensitive with \"gay\", \"sexy\"...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 455217,
          "author_name": "zmzmzmzmzm",
          "author_url": "",
          "post_date": "01/13/2019 09:45:55",
          "content": "<p>i spent 2 days to clean the dataset, such as replacing the strange character, correct misspelling and remove lots of strange symbols.  but it will harm the result under the same architecture with the same random seed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460219,
          "author_name": "stefanobromuri",
          "author_url": "",
          "post_date": "01/23/2019 08:07:22",
          "content": "<p>Maybe substituting the word Trump with 'president' and do similarly with other words, may bring better results though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452610,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "01/09/2019 00:18:26",
      "content": "<p>Interesting thing, I once tried to add 2 cols in the auxiliary input to mark whether \"Trump\" and \"gay\" appears in a sentence (since they seem to have good discrimination to some extend, and the false analysis also). However, the result was a miserable story...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 453578,
      "author_name": "psyfaker",
      "author_url": "",
      "post_date": "01/10/2019 12:12:40",
      "content": "<p>Thank You for an interesting observation! Maybe You have some ideas on how to prevent this, maybe some kind of changing weights or post-processing?</p>",
      "votes": null,
      "replies": [
        {
          "id": 453737,
          "author_name": "jaguar00",
          "author_url": "",
          "post_date": "01/10/2019 17:38:01",
          "content": "<p>There are many things that can be done to overcome this, such as but not limited to 1- increase these values in the validation data and check wich network changes reduce the error. 2- augment more examples in the training data that include keywords with a higher error for the network to learn from it. 3- Remove or fix false labeled data (if any) in the training set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460178,
          "author_name": "isikkuntay",
          "author_url": "",
          "post_date": "01/23/2019 06:03:36",
          "content": "<p>Augmenting the training set works out to something similar to focal loss from what I understand. Is that a fair deduction?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "451874": "I made a false analysis to understand where my model misses it the most, I found most of the wrong predictions related to two keywords after clean all stopwords [Trump and People]:\n\nFor **Sincere training data** that the model marked True:\n\n - People: 5282 *missed training label contains the word*\n - Trump: 2871 \n\nFor **InSincere training data** that the model marked False:\n\n - People: 2488\n - Trump: 1148\n\n*update: True means insincere, False means sincere*",
    "451998": "maybe these words are difficult to distinguish",
    "452010": "I believe, the association of nearby words matters more than singe words. Right ?",
    "452024": "Yes, an association of words matters more, but the same word used in similar ways for both classifications, I checked many examples and found that RNN model classification was more sense for me than some of the provided targets. \nFor example, the model classified following insincere:\n\n - ( why are bootythongs on women so sexy ( to guys ) ?) \n\nHowever, Quora classification was sincere. \nI am trying to understand how really classification is done so at least I can work on some pre-processing to enhance.",
    "452048": "The training set is very noisy, and I guess that's the case with the public test set too. And that's perfectly okay.\n\nHowever, the private set has to be clean, if it isn't then frankly speaking this whole competition is a waste of everyone's time.\n\nAlso in my opinion the above question *is* sincere.",
    "452056": "Agree the private test data should be clean. another challenging thing in here the classification somehow opinion based not as direct as toxic classification. So hope the target generated with the same distribution and clear guidelines to labelers. \nIn the end, if generated with the same distribution it is ok.",
    "452093": "I would argue otherwise that the private data should not be clean and completely similarly distributed as public LB. Otherwise you would need to follow quite a different approach like removing wrongly labeled data and the whole 3 months are a waste of time.",
    "452610": "Interesting thing, I once tried to add 2 cols in the auxiliary input to mark whether \"Trump\" and \"gay\" appears in a sentence (since they seem to have good discrimination to some extend, and the false analysis also). However, the result was a miserable story...",
    "453009": "I agree with Psi, and if you read between the lines a bit on this\n\n&gt; Note that the distribution of questions in the dataset should not be taken to be representative of the distribution of questions asked on Quora. This is, in part, because of the combination of sampling procedures and sanitization measures that have been applied to the final dataset.\n\nit sounds like the dataset used has been sanitized and then split up. Certainly it would be weird if that wasn't what was done.\n\nI am curious about this however\n\n&gt; To date, Quora has employed both machine learning and manual review to address this problem.\n\nI hope this means we aren't trying to build a model to fit another model :/",
    "453578": "Thank You for an interesting observation! Maybe You have some ideas on how to prevent this, maybe some kind of changing weights or post-processing?",
    "453737": "There are many things that can be done to overcome this, such as but not limited to 1- increase these values in the validation data and check wich network changes reduce the error. 2- augment more examples in the training data that include keywords with a higher error for the network to learn from it. 3- Remove or fix false labeled data (if any) in the training set.",
    "455139": "I totally agree the Psi. Train set and test set (public and private) have noisy. So we should make our model robust. Anyway, I find my model is sensitive with \"gay\", \"sexy\"...",
    "455217": "i spent 2 days to clean the dataset, such as replacing the strange character, correct misspelling and remove lots of strange symbols.  but it will harm the result under the same architecture with the same random seed.",
    "460178": "Augmenting the training set works out to something similar to focal loss from what I understand. Is that a fair deduction?",
    "460219": "Maybe substituting the word Trump with 'president' and do similarly with other words, may bring better results though."
  },
  "source": "meta"
}