{
  "id": 79414,
  "title": "A ball park estimate of how pathological the dataset is",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79414",
  "author_name": "",
  "post_date": "2019-02-03T18:45:42.235762400Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Looking to the size and nature of the false positives of one of my models, I estimate that there are at least 30k questions which should be labelled as insincere, but aren't. Considering that in the whole dataset the number of insincere labelled questions are ~78k, you can get an idea of how bad the dataset is.</p>\n\n<p>You know the saying: \"Garbage in...\"</p>\n\n<p>Some Quora questions that my model marks as insincere, but are not marked as insincere (false positives):</p>\n\n<ul>\n<li>girls hate me , but they hate me even more when boys are around me , what do i do ?\nare muslims doing love jihad sex pervert ?</li>\n<li>will sociopaths have sex with women who are unattractive ?</li>\n<li>why do so many quora readers seem to be ignorant of web searching for answers ?</li>\n<li>how can a man with an md and a phd be mean to his patients and assault them for being transgender ?</li>\n<li>what percentage of the anti - trumpers here are russian bots ?</li>\n<li>are women attracted to men 's anus ?</li>\n<li>are [unk] stupid ?</li>\n</ul>\n\n<p>EDIT: The above questions have not been cherry-picked, they all come from a first sample.</p>",
  "messages": [
    {
      "id": "465671",
      "postDate": "02/03/2019 18:45:42",
      "content": "<p>Looking to the size and nature of the false positives of one of my models, I estimate that there are at least 30k questions which should be labelled as insincere, but aren't. Considering that in the whole dataset the number of insincere labelled questions are ~78k, you can get an idea of how bad the dataset is.</p>\n\n<p>You know the saying: \"Garbage in...\"</p>\n\n<p>Some Quora questions that my model marks as insincere, but are not marked as insincere (false positives):</p>\n\n<ul>\n<li>girls hate me , but they hate me even more when boys are around me , what do i do ?\nare muslims doing love jihad sex pervert ?</li>\n<li>will sociopaths have sex with women who are unattractive ?</li>\n<li>why do so many quora readers seem to be ignorant of web searching for answers ?</li>\n<li>how can a man with an md and a phd be mean to his patients and assault them for being transgender ?</li>\n<li>what percentage of the anti - trumpers here are russian bots ?</li>\n<li>are women attracted to men 's anus ?</li>\n<li>are [unk] stupid ?</li>\n</ul>\n\n<p>EDIT: The above questions have not been cherry-picked, they all come from a first sample.</p>",
      "rawMarkdown": "Looking to the size and nature of the false positives of one of my models, I estimate that there are at least 30k questions which should be labelled as insincere, but aren't. Considering that in the whole dataset the number of insincere labelled questions are ~78k, you can get an idea of how bad the dataset is.\n\nYou know the saying: \"Garbage in...\"\n\nSome Quora questions that my model marks as insincere, but are not marked as insincere (false positives):\n\n* girls hate me , but they hate me even more when boys are around me , what do i do ?\nare muslims doing love jihad sex pervert ?\n* will sociopaths have sex with women who are unattractive ?\n* why do so many quora readers seem to be ignorant of web searching for answers ?\n* how can a man with an md and a phd be mean to his patients and assault them for being transgender ?\n* what percentage of the anti - trumpers here are russian bots ?\n* are women attracted to men 's anus ?\n* are [unk] stupid ?\n\nEDIT: The above questions have not been cherry-picked, they all come from a first sample.",
      "votes": null
    },
    {
      "id": "465697",
      "postDate": "02/03/2019 19:53:03",
      "content": "<p>Your choice of false positive examples really shows off the subjective nature of this problem.  While these all seem mislabeled to you, at least one of them seem legitimately sincere questions to me.  </p>\n\n<p>\"how can a man with an md and a phd be mean to his patients and assault them for being transgender ?\"\nThis seems like a very sincere question from someone that just had a very bad experience at the doctors office.</p>\n\n<p>Other people probably have a different way to categorize these as well.  Instead of just a single binary classification, I would have liked to see this dataset with something like a percentage of people who label a question as insincere.  I think this could have led to models that could be better utilized in a production environment.  An even better dataset would have been to provide individual users opinions on how to labeled a subset of the questions and then predict their response on the rest.  This could produce a model where individual users are only seeing questions they would believe to be sincere vs. the aggregate model that shows a group's mean opinion.</p>\n\n<p>But that is the difficulty with this type of competition format.  A lot of what really goes into building this type of model (like designing the inputs and curating the data) is not included in the scope of the competition.</p>",
      "rawMarkdown": "Your choice of false positive examples really shows off the subjective nature of this problem.  While these all seem mislabeled to you, at least one of them seem legitimately sincere questions to me.  \n\n\"how can a man with an md and a phd be mean to his patients and assault them for being transgender ?\"\nThis seems like a very sincere question from someone that just had a very bad experience at the doctors office.\n\nOther people probably have a different way to categorize these as well.  Instead of just a single binary classification, I would have liked to see this dataset with something like a percentage of people who label a question as insincere.  I think this could have led to models that could be better utilized in a production environment.  An even better dataset would have been to provide individual users opinions on how to labeled a subset of the questions and then predict their response on the rest.  This could produce a model where individual users are only seeing questions they would believe to be sincere vs. the aggregate model that shows a group's mean opinion.\n\nBut that is the difficulty with this type of competition format.  A lot of what really goes into building this type of model (like designing the inputs and curating the data) is not included in the scope of the competition.",
      "votes": null
    },
    {
      "id": "465721",
      "postDate": "02/03/2019 20:59:25",
      "content": "<p>I understand the subjectivity of the task, but still, these examples were not cherry picked from my false positives... and you express doubts only in one of them. </p>\n\n<p>By the way, the policy is here:</p>\n\n<p><a href=\"https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora\">https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora</a></p>\n\n<p>And it says: \n\"Questions about individuals that are hurtful, mean-spirited or likely to make the person uncomfortable aren't allowed.\"\n\"Questions that constitute harassment and have the potential to make the experience of using Quora unpleasant or uncomfortable for group(s) of users may be deleted.\"</p>",
      "rawMarkdown": "I understand the subjectivity of the task, but still, these examples were not cherry picked from my false positives... and you express doubts only in one of them. \n\nBy the way, the policy is here:\n\nhttps://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora\n\nAnd it says: \n\"Questions about individuals that are hurtful, mean-spirited or likely to make the person uncomfortable aren't allowed.\"\n\"Questions that constitute harassment and have the potential to make the experience of using Quora unpleasant or uncomfortable for group(s) of users may be deleted.\"",
      "votes": null
    },
    {
      "id": "465735",
      "postDate": "02/03/2019 22:01:10",
      "content": "<p>You've created a new topic so I'll copy paste my previous comment :D</p>\n\n<p>\"... Sincerity is a very moral category itself so labeling the questions is a very subjective thing to do...\"</p>\n\n<p>For me some of questions above look very sincere btw</p>",
      "rawMarkdown": "You've created a new topic so I'll copy paste my previous comment :D\n\n\"... Sincerity is a very moral category itself so labeling the questions is a very subjective thing to do...\"\n\nFor me some of questions above look very sincere btw",
      "votes": null
    },
    {
      "id": "465744",
      "postDate": "02/03/2019 22:21:55",
      "content": "<p>Thanks Oleg, preferred to create it in a new topic.</p>\n\n<p>All this subjectivity reminds me to:</p>\n\n<p><a href=\"https://xkcd.com/114/\">https://xkcd.com/114/</a></p>",
      "rawMarkdown": "Thanks Oleg, preferred to create it in a new topic.\n\nAll this subjectivity reminds me to:\n\nhttps://xkcd.com/114/",
      "votes": null
    },
    {
      "id": "465851",
      "postDate": "02/04/2019 06:31:25",
      "content": "<p>'insincere' thredshold for different topics may be different.  when it comes to race and sex, it may be easier to cross the line.\nDid anyone tried grouping samples by topic and caculate different f1 thredshould? </p>",
      "rawMarkdown": "'insincere' thredshold for different topics may be different.  when it comes to race and sex, it may be easier to cross the line.\nDid anyone tried grouping samples by topic and caculate different f1 thredshould?",
      "votes": null
    },
    {
      "id": "477663",
      "postDate": "02/25/2019 03:52:59",
      "content": "<p>it is really a awesome  discussion!!  I create the features about sex and race for dealing with  the subjective nature of this problem.It  cost 704 private and 700 public.</p>",
      "rawMarkdown": "it is really a awesome  discussion!!  I create the features about sex and race for dealing with  the subjective nature of this problem.It  cost 704 private and 700 public.",
      "votes": null
    },
    {
      "id": "478194",
      "postDate": "02/25/2019 21:48:39",
      "content": "<p>Thanks! What do you mean by \"It cost 704 private and 700 public.\"?</p>",
      "rawMarkdown": "Thanks! What do you mean by \"It cost 704 private and 700 public.\"?",
      "votes": null
    },
    {
      "id": "482124",
      "postDate": "03/02/2019 11:18:39",
      "content": "<p>This process led to an increase in online performance , which I think may be related to the final data set.My final score has increased by four thousand points. 0.700 to 0.704.Thx</p>",
      "rawMarkdown": "This process led to an increase in online performance , which I think may be related to the final data set.My final score has increased by four thousand points. 0.700 to 0.704.Thx",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 465697,
      "author_name": "chris1974mi",
      "author_url": "",
      "post_date": "02/03/2019 19:53:03",
      "content": "<p>Your choice of false positive examples really shows off the subjective nature of this problem.  While these all seem mislabeled to you, at least one of them seem legitimately sincere questions to me.  </p>\n\n<p>\"how can a man with an md and a phd be mean to his patients and assault them for being transgender ?\"\nThis seems like a very sincere question from someone that just had a very bad experience at the doctors office.</p>\n\n<p>Other people probably have a different way to categorize these as well.  Instead of just a single binary classification, I would have liked to see this dataset with something like a percentage of people who label a question as insincere.  I think this could have led to models that could be better utilized in a production environment.  An even better dataset would have been to provide individual users opinions on how to labeled a subset of the questions and then predict their response on the rest.  This could produce a model where individual users are only seeing questions they would believe to be sincere vs. the aggregate model that shows a group's mean opinion.</p>\n\n<p>But that is the difficulty with this type of competition format.  A lot of what really goes into building this type of model (like designing the inputs and curating the data) is not included in the scope of the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 465721,
          "author_name": "manuelsh",
          "author_url": "",
          "post_date": "02/03/2019 20:59:25",
          "content": "<p>I understand the subjectivity of the task, but still, these examples were not cherry picked from my false positives... and you express doubts only in one of them. </p>\n\n<p>By the way, the policy is here:</p>\n\n<p><a href=\"https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora\">https://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora</a></p>\n\n<p>And it says: \n\"Questions about individuals that are hurtful, mean-spirited or likely to make the person uncomfortable aren't allowed.\"\n\"Questions that constitute harassment and have the potential to make the experience of using Quora unpleasant or uncomfortable for group(s) of users may be deleted.\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 465735,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "02/03/2019 22:01:10",
      "content": "<p>You've created a new topic so I'll copy paste my previous comment :D</p>\n\n<p>\"... Sincerity is a very moral category itself so labeling the questions is a very subjective thing to do...\"</p>\n\n<p>For me some of questions above look very sincere btw</p>",
      "votes": null,
      "replies": [
        {
          "id": 465744,
          "author_name": "manuelsh",
          "author_url": "",
          "post_date": "02/03/2019 22:21:55",
          "content": "<p>Thanks Oleg, preferred to create it in a new topic.</p>\n\n<p>All this subjectivity reminds me to:</p>\n\n<p><a href=\"https://xkcd.com/114/\">https://xkcd.com/114/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465851,
          "author_name": "tiopon",
          "author_url": "",
          "post_date": "02/04/2019 06:31:25",
          "content": "<p>'insincere' thredshold for different topics may be different.  when it comes to race and sex, it may be easier to cross the line.\nDid anyone tried grouping samples by topic and caculate different f1 thredshould? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 477663,
      "author_name": "shiyunlu",
      "author_url": "",
      "post_date": "02/25/2019 03:52:59",
      "content": "<p>it is really a awesome  discussion!!  I create the features about sex and race for dealing with  the subjective nature of this problem.It  cost 704 private and 700 public.</p>",
      "votes": null,
      "replies": [
        {
          "id": 478194,
          "author_name": "manuelsh",
          "author_url": "",
          "post_date": "02/25/2019 21:48:39",
          "content": "<p>Thanks! What do you mean by \"It cost 704 private and 700 public.\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 482124,
          "author_name": "shiyunlu",
          "author_url": "",
          "post_date": "03/02/2019 11:18:39",
          "content": "<p>This process led to an increase in online performance , which I think may be related to the final data set.My final score has increased by four thousand points. 0.700 to 0.704.Thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "465671": "Looking to the size and nature of the false positives of one of my models, I estimate that there are at least 30k questions which should be labelled as insincere, but aren't. Considering that in the whole dataset the number of insincere labelled questions are ~78k, you can get an idea of how bad the dataset is.\n\nYou know the saying: \"Garbage in...\"\n\nSome Quora questions that my model marks as insincere, but are not marked as insincere (false positives):\n\n* girls hate me , but they hate me even more when boys are around me , what do i do ?\nare muslims doing love jihad sex pervert ?\n* will sociopaths have sex with women who are unattractive ?\n* why do so many quora readers seem to be ignorant of web searching for answers ?\n* how can a man with an md and a phd be mean to his patients and assault them for being transgender ?\n* what percentage of the anti - trumpers here are russian bots ?\n* are women attracted to men 's anus ?\n* are [unk] stupid ?\n\nEDIT: The above questions have not been cherry-picked, they all come from a first sample.",
    "465697": "Your choice of false positive examples really shows off the subjective nature of this problem.  While these all seem mislabeled to you, at least one of them seem legitimately sincere questions to me.  \n\n\"how can a man with an md and a phd be mean to his patients and assault them for being transgender ?\"\nThis seems like a very sincere question from someone that just had a very bad experience at the doctors office.\n\nOther people probably have a different way to categorize these as well.  Instead of just a single binary classification, I would have liked to see this dataset with something like a percentage of people who label a question as insincere.  I think this could have led to models that could be better utilized in a production environment.  An even better dataset would have been to provide individual users opinions on how to labeled a subset of the questions and then predict their response on the rest.  This could produce a model where individual users are only seeing questions they would believe to be sincere vs. the aggregate model that shows a group's mean opinion.\n\nBut that is the difficulty with this type of competition format.  A lot of what really goes into building this type of model (like designing the inputs and curating the data) is not included in the scope of the competition.",
    "465721": "I understand the subjectivity of the task, but still, these examples were not cherry picked from my false positives... and you express doubts only in one of them. \n\nBy the way, the policy is here:\n\nhttps://www.quora.com/What-are-the-main-policies-and-guidelines-for-questions-on-Quora\n\nAnd it says: \n\"Questions about individuals that are hurtful, mean-spirited or likely to make the person uncomfortable aren't allowed.\"\n\"Questions that constitute harassment and have the potential to make the experience of using Quora unpleasant or uncomfortable for group(s) of users may be deleted.\"",
    "465735": "You've created a new topic so I'll copy paste my previous comment :D\n\n\"... Sincerity is a very moral category itself so labeling the questions is a very subjective thing to do...\"\n\nFor me some of questions above look very sincere btw",
    "465744": "Thanks Oleg, preferred to create it in a new topic.\n\nAll this subjectivity reminds me to:\n\nhttps://xkcd.com/114/",
    "465851": "'insincere' thredshold for different topics may be different.  when it comes to race and sex, it may be easier to cross the line.\nDid anyone tried grouping samples by topic and caculate different f1 thredshould?",
    "477663": "it is really a awesome  discussion!!  I create the features about sex and race for dealing with  the subjective nature of this problem.It  cost 704 private and 700 public.",
    "478194": "Thanks! What do you mean by \"It cost 704 private and 700 public.\"?",
    "482124": "This process led to an increase in online performance , which I think may be related to the final data set.My final score has increased by four thousand points. 0.700 to 0.704.Thx"
  },
  "source": "meta"
}