{
  "id": 79230,
  "title": "Quora would have done better by improving the data set",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79230",
  "author_name": "ManuelSH",
  "post_date": "2019-02-01T11:15:57.380000",
  "votes": 23,
  "comment_count": 10,
  "views": 0,
  "content": "<p>If Quora is making this competition to gain exposure or reputation , then forget what I am saying.</p>\n\n<p>But if they are actually making it to have a better model, then it would have been better for them to spend the competition money in some kind of crowd sourcing (e.g. Mechanical Turk) to get better tagged data.  25,000$ plus the commission paid to Kaggle easily will allow to correctly tag a few hundred thousand questions well (I estimate 800k with three persons checking each question).  If Quora have started with a higher quality dataset, even if it is smaller than current one, it will most probably allow to arrive to F1 much higher than the 0.7 that people are getting with a simpler model.</p>\n\n<p>The current approach Quora took (manually tagging some questions, then using an \"in-house\" model to tag the rest) is making a really bad dataset, as many have noticed. This competition is not about classifying insincere vs sincere questions. It is about classifying what a (rather bad) model classified as an insincere question.</p>\n\n<p>But hey, I am having fun with the contest...</p>",
  "messages": [
    {
      "id": 464720,
      "postDate": "2019-02-01T11:15:57.380Z",
      "content": "<p>If Quora is making this competition to gain exposure or reputation , then forget what I am saying.</p>\n\n<p>But if they are actually making it to have a better model, then it would have been better for them to spend the competition money in some kind of crowd sourcing (e.g. Mechanical Turk) to get better tagged data.  25,000$ plus the commission paid to Kaggle easily will allow to correctly tag a few hundred thousand questions well (I estimate 800k with three persons checking each question).  If Quora have started with a higher quality dataset, even if it is smaller than current one, it will most probably allow to arrive to F1 much higher than the 0.7 that people are getting with a simpler model.</p>\n\n<p>The current approach Quora took (manually tagging some questions, then using an \"in-house\" model to tag the rest) is making a really bad dataset, as many have noticed. This competition is not about classifying insincere vs sincere questions. It is about classifying what a (rather bad) model classified as an insincere question.</p>\n\n<p>But hey, I am having fun with the contest...</p>",
      "rawMarkdown": "If Quora is making this competition to gain exposure or reputation , then forget what I am saying.\n\nBut if they are actually making it to have a better model, then it would have been better for them to spend the competition money in some kind of crowd sourcing (e.g. Mechanical Turk) to get better tagged data.  25,000$ plus the commission paid to Kaggle easily will allow to correctly tag a few hundred thousand questions well (I estimate 800k with three persons checking each question).  If Quora have started with a higher quality dataset, even if it is smaller than current one, it will most probably allow to arrive to F1 much higher than the 0.7 that people are getting with a simpler model.\n\nThe current approach Quora took (manually tagging some questions, then using an \"in-house\" model to tag the rest) is making a really bad dataset, as many have noticed. This competition is not about classifying insincere vs sincere questions. It is about classifying what a (rather bad) model classified as an insincere question.\n\nBut hey, I am having fun with the contest...",
      "votes": 23
    },
    {
      "id": 465214,
      "postDate": "2019-02-02T16:26:05.143Z",
      "content": "<p>I agree with you, but  I believe also sincerity classification is a very challenging NLP problem itself. It is a very moral category so labeling the questions is a very subjective thing to do. Simply same question might look sincere/insincere for people from different cultural fields - we all know this from our every day life/news, no need for examples. Also sincerity in this task usually correlates with irony or sarcasm and that's even more complicated topic because irony depends not only on cultural field but on intelligence, education, etc of questioner.  Probably you need group of experts voting to label it somehow enough and still model won't be able to infer good because of lack of knowledge and complicated human reasoning. I believe itself it's even more complicated than intelligent Q/A system. But one day.. </p>",
      "rawMarkdown": "I agree with you, but  I believe also sincerity classification is a very challenging NLP problem itself. It is a very moral category so labeling the questions is a very subjective thing to do. Simply same question might look sincere/insincere for people from different cultural fields - we all know this from our every day life/news, no need for examples. Also sincerity in this task usually correlates with irony or sarcasm and that's even more complicated topic because irony depends not only on cultural field but on intelligence, education, etc of questioner.  Probably you need group of experts voting to label it somehow enough and still model won't be able to infer good because of lack of knowledge and complicated human reasoning. I believe itself it's even more complicated than intelligent Q/A system. But one day.. ",
      "votes": 6,
      "replies": [
        {
          "id": 465346,
          "postDate": "2019-02-02T22:04:30.403Z",
          "content": "<p>I think Quora has already an almost objective definition of insincere questions, which wouldn't make it hard for a source of annotators to use. It has published by someone already in the forum. And if you look to examples like the following, you will notice how badly labelled the dataset is:</p>\n\n<p>None of these were labelled as insincere:\n\"Longer dick more desirable?\"\n\"Can you get pregnant if he just sticks it in?\"\n\"Do intelligent people ask questions?\"\n\"How do grown men look in skinny jeans?\"\n\"Why does Ronaldinho looks like a girl?\"\n\"Does Java save your life?\"</p>",
          "rawMarkdown": "I think Quora has already an almost objective definition of insincere questions, which wouldn't make it hard for a source of annotators to use. It has published by someone already in the forum. And if you look to examples like the following, you will notice how badly labelled the dataset is:\n\nNone of these were labelled as insincere:\n\"Longer dick more desirable?\"\n\"Can you get pregnant if he just sticks it in?\"\n\"Do intelligent people ask questions?\"\n\"How do grown men look in skinny jeans?\"\n\"Why does Ronaldinho looks like a girl?\"\n\"Does Java save your life?\"",
          "votes": 1
        }
      ]
    },
    {
      "id": 465420,
      "postDate": "2019-02-03T04:12:18.780Z",
      "content": "<p>Seeing questions like \"Hello sir?\" and \"What meow?\" being labelled as sincere is a little annoying. There are 78 questions with 2 words or less. More than half are rated sincere. \nBut mislabelled examples seem to be a problem with the toxicity labelling in general. In the paper I linked to below (Challenges for Toxic Comment Detection), they write that they they found 10% of the sample in their dataset to be falsely labelled. For more than half of the false positive mistakes that their model made, they thought that the model was actually correct and that the true label was a mistake.</p>\n\n<p><a href=\"https://arxiv.org/pdf/1809.07572.pdf\">https://arxiv.org/pdf/1809.07572.pdf</a></p>",
      "rawMarkdown": "Seeing questions like \"Hello sir?\" and \"What meow?\" being labelled as sincere is a little annoying. There are 78 questions with 2 words or less. More than half are rated sincere. \nBut mislabelled examples seem to be a problem with the toxicity labelling in general. In the paper I linked to below (Challenges for Toxic Comment Detection), they write that they they found 10% of the sample in their dataset to be falsely labelled. For more than half of the false positive mistakes that their model made, they thought that the model was actually correct and that the true label was a mistake.\n\nhttps://arxiv.org/pdf/1809.07572.pdf",
      "votes": 2,
      "replies": [
        {
          "id": 465573,
          "postDate": "2019-02-03T14:08:19.067Z",
          "content": "<p>It's true, labelling is hard, but there are several methodologies to increase the quality of labelling through mechanical turk / crowd sourcing. Btw, nice paper!</p>",
          "rawMarkdown": "It's true, labelling is hard, but there are several methodologies to increase the quality of labelling through mechanical turk / crowd sourcing. Btw, nice paper!"
        }
      ]
    },
    {
      "id": 465296,
      "postDate": "2019-02-02T18:37:41.303Z",
      "content": "<p>Labeling indeed sucks. What I typically do in NLP projects is having at least 3 assessors label each sample. Here it looks like they preferred 3x larger but shittier dataset. \nI love tf-idf + logit and use it in a couple of projects in production. But what do we have here if we train this pipeline and interpret weights with eli5? just no sense at all :(</p>\n\n<p><img src=\"https://habrastorage.org/webt/vn/cx/tw/vncxtw5pm8b46j2xj26uh7dgjm4.png\" alt=\"enter image description here\">\n<img src=\"https://habrastorage.org/webt/qt/4k/yd/qt4kyd8fsjbnqfh-uowrvc9ihxw.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Labeling indeed sucks. What I typically do in NLP projects is having at least 3 assessors label each sample. Here it looks like they preferred 3x larger but shittier dataset. \nI love tf-idf + logit and use it in a couple of projects in production. But what do we have here if we train this pipeline and interpret weights with eli5? just no sense at all :(\n\n![enter image description here][1]\n![enter image description here][2]\n\n[1]:  https://habrastorage.org/webt/vn/cx/tw/vncxtw5pm8b46j2xj26uh7dgjm4.png\n[2]: https://habrastorage.org/webt/qt/4k/yd/qt4kyd8fsjbnqfh-uowrvc9ihxw.png",
      "votes": 2
    },
    {
      "id": 465207,
      "postDate": "2019-02-02T15:59:24.897Z",
      "content": "<p>At least I got the framework wrote then if there's another NLP competition I can just copy paste </p>",
      "rawMarkdown": "At least I got the framework wrote then if there's another NLP competition I can just copy paste ",
      "votes": 2
    },
    {
      "id": 465320,
      "postDate": "2019-02-02T20:16:16.677Z",
      "content": "<p>Or They wanna obtain model for future improving dataset (read 1M qourans too long, but model select 50k-70k), train it again and get better score.</p>",
      "rawMarkdown": "Or They wanna obtain model for future improving dataset (read 1M qourans too long, but model select 50k-70k), train it again and get better score."
    },
    {
      "id": 465667,
      "postDate": "2019-02-03T18:30:06.157Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 465322,
      "postDate": "2019-02-02T20:20:10.813Z",
      "content": "<p>Labelling to get ground-truth is a pain and time killer in any ML project. \nLabelling &gt; 1 million texts is a huge task, so I would say they did a good job here.\nLet's have fun with this dataset!</p>",
      "rawMarkdown": "Labelling to get ground-truth is a pain and time killer in any ML project. \nLabelling &gt; 1 million texts is a huge task, so I would say they did a good job here.\nLet's have fun with this dataset!",
      "votes": -2,
      "isDeleted": true,
      "replies": [
        {
          "id": 465349,
          "postDate": "2019-02-02T22:08:19.540Z",
          "content": "<p>I am not sure they did a good job, given the bad quality of the dataset... Re labelling, as I said, with $25k they could employ a crowd source to label several hundred thousand questions in less than a week, is not such a huge task.</p>\n\n<p>But all in all, I agree with you, let's have fun! (at least for 2 days more)</p>",
          "rawMarkdown": "I am not sure they did a good job, given the bad quality of the dataset... Re labelling, as I said, with $25k they could employ a crowd source to label several hundred thousand questions in less than a week, is not such a huge task.\n\nBut all in all, I agree with you, let's have fun! (at least for 2 days more)",
          "votes": 5
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 465214,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-02-02T16:26:05.143000",
      "content": "<p>I agree with you, but  I believe also sincerity classification is a very challenging NLP problem itself. It is a very moral category so labeling the questions is a very subjective thing to do. Simply same question might look sincere/insincere for people from different cultural fields - we all know this from our every day life/news, no need for examples. Also sincerity in this task usually correlates with irony or sarcasm and that's even more complicated topic because irony depends not only on cultural field but on intelligence, education, etc of questioner.  Probably you need group of experts voting to label it somehow enough and still model won't be able to infer good because of lack of knowledge and complicated human reasoning. I believe itself it's even more complicated than intelligent Q/A system. But one day.. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 465346,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-02-02T22:04:30.403000",
          "content": "<p>I think Quora has already an almost objective definition of insincere questions, which wouldn't make it hard for a source of annotators to use. It has published by someone already in the forum. And if you look to examples like the following, you will notice how badly labelled the dataset is:</p>\n\n<p>None of these were labelled as insincere:\n\"Longer dick more desirable?\"\n\"Can you get pregnant if he just sticks it in?\"\n\"Do intelligent people ask questions?\"\n\"How do grown men look in skinny jeans?\"\n\"Why does Ronaldinho looks like a girl?\"\n\"Does Java save your life?\"</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 465420,
      "author_name": "Julius Hochmuth",
      "author_url": "",
      "post_date": "2019-02-03T04:12:18.780000",
      "content": "<p>Seeing questions like \"Hello sir?\" and \"What meow?\" being labelled as sincere is a little annoying. There are 78 questions with 2 words or less. More than half are rated sincere. \nBut mislabelled examples seem to be a problem with the toxicity labelling in general. In the paper I linked to below (Challenges for Toxic Comment Detection), they write that they they found 10% of the sample in their dataset to be falsely labelled. For more than half of the false positive mistakes that their model made, they thought that the model was actually correct and that the true label was a mistake.</p>\n\n<p><a href=\"https://arxiv.org/pdf/1809.07572.pdf\">https://arxiv.org/pdf/1809.07572.pdf</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 465573,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-02-03T14:08:19.067000",
          "content": "<p>It's true, labelling is hard, but there are several methodologies to increase the quality of labelling through mechanical turk / crowd sourcing. Btw, nice paper!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 465296,
      "author_name": "Yury Kashnitsky",
      "author_url": "",
      "post_date": "2019-02-02T18:37:41.303000",
      "content": "<p>Labeling indeed sucks. What I typically do in NLP projects is having at least 3 assessors label each sample. Here it looks like they preferred 3x larger but shittier dataset. \nI love tf-idf + logit and use it in a couple of projects in production. But what do we have here if we train this pipeline and interpret weights with eli5? just no sense at all :(</p>\n\n<p><img src=\"https://habrastorage.org/webt/vn/cx/tw/vncxtw5pm8b46j2xj26uh7dgjm4.png\" alt=\"enter image description here\">\n<img src=\"https://habrastorage.org/webt/qt/4k/yd/qt4kyd8fsjbnqfh-uowrvc9ihxw.png\" alt=\"enter image description here\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 465207,
      "author_name": "Mekaveli.",
      "author_url": "",
      "post_date": "2019-02-02T15:59:24.897000",
      "content": "<p>At least I got the framework wrote then if there's another NLP competition I can just copy paste </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 465320,
      "author_name": "Leigh",
      "author_url": "",
      "post_date": "2019-02-02T20:16:16.677000",
      "content": "<p>Or They wanna obtain model for future improving dataset (read 1M qourans too long, but model select 50k-70k), train it again and get better score.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 465667,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-03T18:30:06.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 465322,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-02T20:20:10.813000",
      "content": "<p>Labelling to get ground-truth is a pain and time killer in any ML project. \nLabelling &gt; 1 million texts is a huge task, so I would say they did a good job here.\nLet's have fun with this dataset!</p>",
      "votes": -2,
      "replies": [
        {
          "id": 465349,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-02-02T22:08:19.540000",
          "content": "<p>I am not sure they did a good job, given the bad quality of the dataset... Re labelling, as I said, with $25k they could employ a crowd source to label several hundred thousand questions in less than a week, is not such a huge task.</p>\n\n<p>But all in all, I agree with you, let's have fun! (at least for 2 days more)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "464720": "If Quora is making this competition to gain exposure or reputation , then forget what I am saying.\n\nBut if they are actually making it to have a better model, then it would have been better for them to spend the competition money in some kind of crowd sourcing (e.g. Mechanical Turk) to get better tagged data.  25,000$ plus the commission paid to Kaggle easily will allow to correctly tag a few hundred thousand questions well (I estimate 800k with three persons checking each question).  If Quora have started with a higher quality dataset, even if it is smaller than current one, it will most probably allow to arrive to F1 much higher than the 0.7 that people are getting with a simpler model.\n\nThe current approach Quora took (manually tagging some questions, then using an \"in-house\" model to tag the rest) is making a really bad dataset, as many have noticed. This competition is not about classifying insincere vs sincere questions. It is about classifying what a (rather bad) model classified as an insincere question.\n\nBut hey, I am having fun with the contest...",
    "465214": "I agree with you, but  I believe also sincerity classification is a very challenging NLP problem itself. It is a very moral category so labeling the questions is a very subjective thing to do. Simply same question might look sincere/insincere for people from different cultural fields - we all know this from our every day life/news, no need for examples. Also sincerity in this task usually correlates with irony or sarcasm and that's even more complicated topic because irony depends not only on cultural field but on intelligence, education, etc of questioner.  Probably you need group of experts voting to label it somehow enough and still model won't be able to infer good because of lack of knowledge and complicated human reasoning. I believe itself it's even more complicated than intelligent Q/A system. But one day.. ",
    "465420": "Seeing questions like \"Hello sir?\" and \"What meow?\" being labelled as sincere is a little annoying. There are 78 questions with 2 words or less. More than half are rated sincere. \nBut mislabelled examples seem to be a problem with the toxicity labelling in general. In the paper I linked to below (Challenges for Toxic Comment Detection), they write that they they found 10% of the sample in their dataset to be falsely labelled. For more than half of the false positive mistakes that their model made, they thought that the model was actually correct and that the true label was a mistake.\n\nhttps://arxiv.org/pdf/1809.07572.pdf",
    "465296": "Labeling indeed sucks. What I typically do in NLP projects is having at least 3 assessors label each sample. Here it looks like they preferred 3x larger but shittier dataset. \nI love tf-idf + logit and use it in a couple of projects in production. But what do we have here if we train this pipeline and interpret weights with eli5? just no sense at all :(\n\n![enter image description here][1]\n![enter image description here][2]\n\n[1]:  https://habrastorage.org/webt/vn/cx/tw/vncxtw5pm8b46j2xj26uh7dgjm4.png\n[2]: https://habrastorage.org/webt/qt/4k/yd/qt4kyd8fsjbnqfh-uowrvc9ihxw.png",
    "465207": "At least I got the framework wrote then if there's another NLP competition I can just copy paste ",
    "465320": "Or They wanna obtain model for future improving dataset (read 1M qourans too long, but model select 50k-70k), train it again and get better score.",
    "465667": "",
    "465322": "Labelling to get ground-truth is a pain and time killer in any ML project. \nLabelling &gt; 1 million texts is a huge task, so I would say they did a good job here.\nLet's have fun with this dataset!"
  }
}