{
  "id": 71429,
  "title": "Funnily Mislabeled Samples",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71429",
  "author_name": "Joe Eddy",
  "post_date": "2018-11-13T16:27:01.317000",
  "votes": 29,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I'm not sure yet how often this sort of thing happens for longer questions, but for short questions (10 or less chars, of which there are only 37 samples in the training data) there are some blatantly mislabeled samples.</p>\n\n<p>All of these are labeled 0 (sincere):</p>\n\n<p>\"Is God 42?\"\n\"In Islam?\"\n\"How can I?\"\n\"Am I fake?\"\n\"Wat is 1A?\"\n\"Why is 16?\"\n\"Hello sir?\"\n\"IS 1+1 21?\"</p>\n\n<p>With any luck this is more funny than it is relevant to model training. But I can't help but wonder how often this mislabeling happens in the entire training corpus, given that even these obviously bad samples get mislabeled.</p>\n\n<p>Does anyone have findings to share on mislabeling in the broader corpus? I know the labels \"are not guaranteed to be perfect\", but there's a pretty broad spectrum between total noise and not quite perfect - clearly there's plenty of good signal, but I think it's important to keep in mind.</p>",
  "messages": [
    {
      "id": 420444,
      "postDate": "2018-11-13T16:27:01.317Z",
      "content": "<p>I'm not sure yet how often this sort of thing happens for longer questions, but for short questions (10 or less chars, of which there are only 37 samples in the training data) there are some blatantly mislabeled samples.</p>\n\n<p>All of these are labeled 0 (sincere):</p>\n\n<p>\"Is God 42?\"\n\"In Islam?\"\n\"How can I?\"\n\"Am I fake?\"\n\"Wat is 1A?\"\n\"Why is 16?\"\n\"Hello sir?\"\n\"IS 1+1 21?\"</p>\n\n<p>With any luck this is more funny than it is relevant to model training. But I can't help but wonder how often this mislabeling happens in the entire training corpus, given that even these obviously bad samples get mislabeled.</p>\n\n<p>Does anyone have findings to share on mislabeling in the broader corpus? I know the labels \"are not guaranteed to be perfect\", but there's a pretty broad spectrum between total noise and not quite perfect - clearly there's plenty of good signal, but I think it's important to keep in mind.</p>",
      "rawMarkdown": "I'm not sure yet how often this sort of thing happens for longer questions, but for short questions (10 or less chars, of which there are only 37 samples in the training data) there are some blatantly mislabeled samples.\n\nAll of these are labeled 0 (sincere):\n\n\"Is God 42?\"\n\"In Islam?\"\n\"How can I?\"\n\"Am I fake?\"\n\"Wat is 1A?\"\n\"Why is 16?\"\n\"Hello sir?\"\n\"IS 1+1 21?\"\n\nWith any luck this is more funny than it is relevant to model training. But I can't help but wonder how often this mislabeling happens in the entire training corpus, given that even these obviously bad samples get mislabeled.\n\nDoes anyone have findings to share on mislabeling in the broader corpus? I know the labels \"are not guaranteed to be perfect\", but there's a pretty broad spectrum between total noise and not quite perfect - clearly there's plenty of good signal, but I think it's important to keep in mind.",
      "votes": 29
    },
    {
      "id": 421124,
      "postDate": "2018-11-14T16:04:39.790Z",
      "content": "<p>I'm not sure what proportion of questions are mislabeled, but going through the first 200, here are all the questions that are mislabeled in my opinion (all labeled 0 when they should probably be labeled 1):</p>\n\n<p>\"Is Gaza slowly becoming Auschwitz, Dachau or Treblinka for Palestinians?\"\n\"Have you licked the skin of a corpse?\"\n\"Where is Muhammed now?\"\n\"I wear an insulin pump, and a lot of girls don't like it. Nick Jonas wears one and dated Selena Gomez. Is there difference between me and Nick several zeros missing in my bank account?\"\n\"Does the people who are rich and still claim reservation have any conscience? Why are they given reservation if are rich?\"</p>\n\n<p>From there, I can extrapolate that about 2.5% of questions are mislabeled. Given that about 6.2% of the questions in the training set are labeled 1, about a quarter of \"insincere\" questions are mislabeled.</p>",
      "rawMarkdown": "I'm not sure what proportion of questions are mislabeled, but going through the first 200, here are all the questions that are mislabeled in my opinion (all labeled 0 when they should probably be labeled 1):\n\n\"Is Gaza slowly becoming Auschwitz, Dachau or Treblinka for Palestinians?\"\n\"Have you licked the skin of a corpse?\"\n\"Where is Muhammed now?\"\n\"I wear an insulin pump, and a lot of girls don't like it. Nick Jonas wears one and dated Selena Gomez. Is there difference between me and Nick several zeros missing in my bank account?\"\n\"Does the people who are rich and still claim reservation have any conscience? Why are they given reservation if are rich?\"\n\nFrom there, I can extrapolate that about 2.5% of questions are mislabeled. Given that about 6.2% of the questions in the training set are labeled 1, about a quarter of \"insincere\" questions are mislabeled.\n",
      "votes": 7,
      "replies": [
        {
          "id": 422018,
          "postDate": "2018-11-15T17:13:23.983Z",
          "content": "<p>Interesting point! Some questions are hard to label even for a human. A lot of those are a matter of personal opinion. Concerning those questions for example:\n-  \"Have you licked the skin of a corpse?\"\n- \"Where is Muhammed now?\"</p>\n\n<p>I would rather tend to say they are sincere. Though not conventional for sure haha.</p>",
          "rawMarkdown": "Interesting point! Some questions are hard to label even for a human. A lot of those are a matter of personal opinion. Concerning those questions for example:\n-  \"Have you licked the skin of a corpse?\"\n- \"Where is Muhammed now?\"\n\nI would rather tend to say they are sincere. Though not conventional for sure haha."
        },
        {
          "id": 437402,
          "postDate": "2018-12-11T20:58:10.893Z",
          "content": "<p>As someone who occasionally asks, answers, and flags questions on Quora, but is not a Quora employee, I would judge all of those as sincere, though some seem to have been asked by people with a lack of perspective or background knowledge.</p>\n\n<p>For comparison, here are some questions I have asked (completely sincerely) that I think (and I may be wrong) have a similar (low, IMO) level of apparent absurdity:</p>\n\n<ul>\n<li><a href=\"https://www.quora.com/Are-there-science-fairs-for-adults\">Are there science fairs for adults?</a></li>\n<li><a href=\"https://www.quora.com/Whats-the-greatest-number-of-citations-youve-ever-seen-in-a-single-document\">What's the greatest number of citations you've ever seen in a single document?</a></li>\n<li><a href=\"https://www.quora.com/What-is-the-distribution-of-values-for-number-of-major-and-minor-characters-per-word-page-in-common-novels\">What is the distribution of values for number of major and minor characters per word/page in common novels?</a></li>\n<li><a href=\"https://www.quora.com/unanswered/What-data-storage-technology-can-store-data-for-12-000-years-with-intermittent-writing-and-reading-and-other-requirements\">What data storage technology can store data for 12,000 years with intermittent writing and reading, and other requirements?</a></li>\n<li><a href=\"https://www.quora.com/Why-doesnt-Quora-have-question-comments-anymore\">Why doesn't Quora have question comments anymore?</a></li>\n</ul>\n\n<p>The last one is interesting because its premise is false—Quora does still have question comments—which is a criterion for insincerity given in the competition rules, but I did ask it sincerely, because I didn't notice the button that you have to click to view question comments. (Immediately after asking the question, of course, I noticed the button and answered my own question.)</p>",
          "rawMarkdown": "As someone who occasionally asks, answers, and flags questions on Quora, but is not a Quora employee, I would judge all of those as sincere, though some seem to have been asked by people with a lack of perspective or background knowledge.\n\nFor comparison, here are some questions I have asked (completely sincerely) that I think (and I may be wrong) have a similar (low, IMO) level of apparent absurdity:\n\n* [Are there science fairs for adults?](https://www.quora.com/Are-there-science-fairs-for-adults)\n* [What's the greatest number of citations you've ever seen in a single document?](https://www.quora.com/Whats-the-greatest-number-of-citations-youve-ever-seen-in-a-single-document)\n* [What is the distribution of values for number of major and minor characters per word/page in common novels?](https://www.quora.com/What-is-the-distribution-of-values-for-number-of-major-and-minor-characters-per-word-page-in-common-novels)\n* [What data storage technology can store data for 12,000 years with intermittent writing and reading, and other requirements?](https://www.quora.com/unanswered/What-data-storage-technology-can-store-data-for-12-000-years-with-intermittent-writing-and-reading-and-other-requirements)\n* [Why doesn't Quora have question comments anymore?](https://www.quora.com/Why-doesnt-Quora-have-question-comments-anymore)\n\nThe last one is interesting because its premise is false—Quora does still have question comments—which is a criterion for insincerity given in the competition rules, but I did ask it sincerely, because I didn't notice the button that you have to click to view question comments. (Immediately after asking the question, of course, I noticed the button and answered my own question.)"
        }
      ]
    },
    {
      "id": 420749,
      "postDate": "2018-11-14T04:19:16.983Z",
      "content": "<p>Thanks Joe for sharing this!  I have found many mislabeled data in the longer sentences here as well : <a href=\"https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru\">https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru</a></p>\n\n<p>The question is, is there anyway to do about it ? Or we just have to leave it as a noise.</p>",
      "rawMarkdown": "Thanks Joe for sharing this!  I have found many mislabeled data in the longer sentences here as well : https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru\n\nThe question is, is there anyway to do about it ? Or we just have to leave it as a noise.",
      "votes": 7
    },
    {
      "id": 420526,
      "postDate": "2018-11-13T19:13:14.417Z",
      "content": "<p>The kernel page says <code>Quora has employed both machine learning and manual review</code> to identify sincerity in the past, this means that these questions could have been marked sincere by some ML model. Also, Quora questions have a descriptive paragraph after the title, so \"Why is 16?\" could've been followed by a description of some math problem. \"Hello sir?\" could've been followed by a question about etiquette in the scope of ESL. Quora's manual review would certainly see the question descriptions, and perhaps their ML algos would too. </p>\n\n<p>Unfortunately (and for no stated reason), they don't provide it to us in the data for this kernel, so we're at a notable disadvantage when it comes to titles like that. </p>",
      "rawMarkdown": "The kernel page says `Quora has employed both machine learning and manual review` to identify sincerity in the past, this means that these questions could have been marked sincere by some ML model. Also, Quora questions have a descriptive paragraph after the title, so \"Why is 16?\" could've been followed by a description of some math problem. \"Hello sir?\" could've been followed by a question about etiquette in the scope of ESL. Quora's manual review would certainly see the question descriptions, and perhaps their ML algos would too. \n\nUnfortunately (and for no stated reason), they don't provide it to us in the data for this kernel, so we're at a notable disadvantage when it comes to titles like that. ",
      "votes": 6,
      "replies": [
        {
          "id": 437388,
          "postDate": "2018-12-11T20:20:28.503Z",
          "content": "<blockquote>\n  <p>Also, Quora questions have a descriptive paragraph after the title</p>\n</blockquote>\n\n<p>Not anymore: Quora no longer allows question details, and used a bot to move all question details that existed before the change into a comment on the question. (Comments are hidden when you load the page, too—you have to click a button to see them.) The current rule is that the entire question must fit in the title (which has a limited length) and it must stand on its own. You are, however, allowed to include one URL as a \"context link\", to show the answerers what you're asking about, in case it's not obvious.</p>",
          "rawMarkdown": "&gt; Also, Quora questions have a descriptive paragraph after the title\n\nNot anymore: Quora no longer allows question details, and used a bot to move all question details that existed before the change into a comment on the question. (Comments are hidden when you load the page, too—you have to click a button to see them.) The current rule is that the entire question must fit in the title (which has a limited length) and it must stand on its own. You are, however, allowed to include one URL as a \"context link\", to show the answerers what you're asking about, in case it's not obvious."
        }
      ]
    },
    {
      "id": 420744,
      "postDate": "2018-11-14T04:12:19.557Z",
      "content": "<p>I don't know man. \"Is God 42?\" sounds like some Vsauce video title. </p>",
      "rawMarkdown": "I don't know man. \"Is God 42?\" sounds like some Vsauce video title. ",
      "votes": 3,
      "replies": [
        {
          "id": 426109,
          "postDate": "2018-11-22T16:20:23.417Z",
          "content": "<p>It is probably a reference to the Hitchhiker's guide to the galaxy:</p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy\">https://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy</a></p>",
          "rawMarkdown": "It is probably a reference to the Hitchhiker's guide to the galaxy:\n\nhttps://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy",
          "votes": 1
        }
      ]
    },
    {
      "id": 420542,
      "postDate": "2018-11-13T19:42:20.910Z",
      "content": "<p>Thats why they need a better model :D</p>",
      "rawMarkdown": "Thats why they need a better model :D",
      "votes": 3,
      "replies": [
        {
          "id": 420903,
          "postDate": "2018-11-14T09:54:08.030Z",
          "content": "<p>In your opinion, Is it better to correct those samples that are miss-labelled or to leave them as it is .?</p>",
          "rawMarkdown": "In your opinion, Is it better to correct those samples that are miss-labelled or to leave them as it is .?",
          "votes": 1
        },
        {
          "id": 421008,
          "postDate": "2018-11-14T13:15:05.997Z",
          "content": "<p>I think correcting them (if possible) will make the training easier; however, there is also a downside as the test dataset will also have noise. </p>\n\n<p>In that case, our validation data will not have the same distribution as the test data anymore, and so it will be difficult to judge our performance before submission.</p>",
          "rawMarkdown": "I think correcting them (if possible) will make the training easier; however, there is also a downside as the test dataset will also have noise. \n\nIn that case, our validation data will not have the same distribution as the test data anymore, and so it will be difficult to judge our performance before submission."
        },
        {
          "id": 421107,
          "postDate": "2018-11-14T15:44:50.380Z",
          "content": "<p>I think it's best to not correct the \"mislabeled\" questions. There might be patterns to which questions tend to be mislabeled (for example, very short questions might tended to be labeled \"sincere\" even if they are actually \"insincere\"). The test data would likely have the same patterns of mislabeling, so \"correcting\" the training data might actually decrease the accuracy of our models when applied to the test data.</p>",
          "rawMarkdown": "I think it's best to not correct the \"mislabeled\" questions. There might be patterns to which questions tend to be mislabeled (for example, very short questions might tended to be labeled \"sincere\" even if they are actually \"insincere\"). The test data would likely have the same patterns of mislabeling, so \"correcting\" the training data might actually decrease the accuracy of our models when applied to the test data.",
          "votes": 3
        },
        {
          "id": 421117,
          "postDate": "2018-11-14T15:53:26.003Z",
          "content": "<p>My point and Ben are exactly the same :)</p>",
          "rawMarkdown": "My point and Ben are exactly the same :)"
        },
        {
          "id": 421145,
          "postDate": "2018-11-14T16:37:26.687Z",
          "content": "<p>I agree, to achieve best performance on the holdout data, we shouldn't fix the labels. This is an oversight on Quora's part, since the data was labeled not only manually but also by their own, current ML algorithm.  </p>\n\n<p>You're essentially training a ML algorithm on the output of another algorithm, so there's already a hypothetical upper bound to accuracy on real, correctly labeled data. Based on their description of the competition and what they said about the data (in vague terms) I get the sense that there's a chance that the test data is labeled entirely manually vs the training data which is labeled by a mix of manual + ML (could be wrong, but there is a chance). Only way to know for sure is to train the same model first on the current data and then again on the data but with the labels manually fixed (perhaps we could get a running list of all known mislabeled rows) and see which one performs best on the holdout data. </p>\n\n<p>Another thing worth considering is that their initial algorithm (and manual labelers) almost certainly had access to not just the question title, but also the question description, which is (IMHO) necessary to determine the sincerity of a question in certain circumstances. So we're trying to fit an estimator to an existing model that was trained with more data than we have available now. Doesn't bode well, maybe that's why even the best algorithms are getting an F-score of around 0.7. </p>",
          "rawMarkdown": "I agree, to achieve best performance on the holdout data, we shouldn't fix the labels. This is an oversight on Quora's part, since the data was labeled not only manually but also by their own, current ML algorithm.  \n \nYou're essentially training a ML algorithm on the output of another algorithm, so there's already a hypothetical upper bound to accuracy on real, correctly labeled data. Based on their description of the competition and what they said about the data (in vague terms) I get the sense that there's a chance that the test data is labeled entirely manually vs the training data which is labeled by a mix of manual + ML (could be wrong, but there is a chance). Only way to know for sure is to train the same model first on the current data and then again on the data but with the labels manually fixed (perhaps we could get a running list of all known mislabeled rows) and see which one performs best on the holdout data. \n   \nAnother thing worth considering is that their initial algorithm (and manual labelers) almost certainly had access to not just the question title, but also the question description, which is (IMHO) necessary to determine the sincerity of a question in certain circumstances. So we're trying to fit an estimator to an existing model that was trained with more data than we have available now. Doesn't bode well, maybe that's why even the best algorithms are getting an F-score of around 0.7. ",
          "votes": 12
        },
        {
          "id": 421998,
          "postDate": "2018-11-15T16:54:40.460Z",
          "content": "<p>The possibility that the test set also have mislabeled questions is definitely to be taken into account. However there is no proof that there is a clear pattern in mislabeling. And definitely no proof that one's model -not built for that purpose- would catch it. Which would mean we do learn noise and it would be better to correct it.\nI don't think there is a clear answer to this question.\nA simple idea would be to say \"that option is too much time consuming and I'd rather focus on other parts of the process\".</p>",
          "rawMarkdown": "The possibility that the test set also have mislabeled questions is definitely to be taken into account. However there is no proof that there is a clear pattern in mislabeling. And definitely no proof that one's model -not built for that purpose- would catch it. Which would mean we do learn noise and it would be better to correct it.\nI don't think there is a clear answer to this question.\nA simple idea would be to say \"that option is too much time consuming and I'd rather focus on other parts of the process\"."
        }
      ]
    },
    {
      "id": 459139,
      "postDate": "2019-01-21T09:11:11.577Z",
      "content": "<p>You know the saying: \"garbage in...\"</p>",
      "rawMarkdown": "You know the saying: \"garbage in...\""
    },
    {
      "id": 458981,
      "postDate": "2019-01-21T00:16:19.657Z",
      "content": "<p>This problem actually seems to occur with other datasets as well. In the paper I linked to, they discuss problems that occur when trying to label toxic comments. They used data from wikipedia and twitter. In regard to the mislabelling of samples, they had this to say: \"We find that 23% of sampled comments in the false negatives of the Wikipedia dataset do not fulfill the toxic definition in our view.\" You can find them discussing this problem in section 6.1.</p>\n\n<p><a href=\"https://arxiv.org/pdf/1809.07572.pdf\">https://arxiv.org/pdf/1809.07572.pdf</a></p>\n\n<p>Anyways, here's a good one:\nWhy does America have to accept people from Sh1t hole countries?,0</p>",
      "rawMarkdown": "This problem actually seems to occur with other datasets as well. In the paper I linked to, they discuss problems that occur when trying to label toxic comments. They used data from wikipedia and twitter. In regard to the mislabelling of samples, they had this to say: \"We find that 23% of sampled comments in the false negatives of the Wikipedia dataset do not fulfill the toxic definition in our view.\" You can find them discussing this problem in section 6.1.\n\nhttps://arxiv.org/pdf/1809.07572.pdf\n\nAnyways, here's a good one:\nWhy does America have to accept people from Sh1t hole countries?,0\n"
    },
    {
      "id": 457491,
      "postDate": "2019-01-17T14:59:33.363Z",
      "content": "<p>The questions listed by OP have 2 or 3 words.</p>\n\n<p>I suspect that Quora has some logic behind (mis)labeling them as sincere for this challenge, and that our code should (mis)label all 2 or 3 words long questions as sincere.</p>",
      "rawMarkdown": "The questions listed by OP have 2 or 3 words.\n\nI suspect that Quora has some logic behind (mis)labeling them as sincere for this challenge, and that our code should (mis)label all 2 or 3 words long questions as sincere."
    },
    {
      "id": 439486,
      "postDate": "2018-12-15T15:27:34.447Z",
      "content": "<p>I have the intuition some of this mislabeling is done on purpose to avoid the influence of manual correction of predicted labels on the test set. If not, it would be way too easy to correct labels manually on the test set.</p>\n\n<p>About some insincere questions wrongly labeled as not insincere, I have tried to correct them and results have not improved. I guess there is too much data for them to be a problem to find the right minimum.</p>\n\n<p>I am confident some insincere questions are not actually detected in the question but in the answer (e.g., for spam). I guess it is very interesting for Quora to try to predict potential spam-related questions before the appearance fo the actual spam.</p>",
      "rawMarkdown": "I have the intuition some of this mislabeling is done on purpose to avoid the influence of manual correction of predicted labels on the test set. If not, it would be way too easy to correct labels manually on the test set.\n\nAbout some insincere questions wrongly labeled as not insincere, I have tried to correct them and results have not improved. I guess there is too much data for them to be a problem to find the right minimum.\n\nI am confident some insincere questions are not actually detected in the question but in the answer (e.g., for spam). I guess it is very interesting for Quora to try to predict potential spam-related questions before the appearance fo the actual spam.\n\n\n",
      "replies": [
        {
          "id": 439493,
          "postDate": "2018-12-15T15:42:30.327Z",
          "content": "<blockquote>\n  <p>I have the intuition some of this mislabeling is done on purpose to avoid the influence of manual correction of predicted labels on the test set. If not, it would be way too easy to correct labels manually on the test set.</p>\n</blockquote>\n\n<p>This is a 2 stages competition, with a new test set at 2nd stage.  Thus, that would be totally useless</p>",
          "rawMarkdown": "&gt; I have the intuition some of this mislabeling is done on purpose to avoid the influence of manual correction of predicted labels on the test set. If not, it would be way too easy to correct labels manually on the test set.\n\n\nThis is a 2 stages competition, with a new test set at 2nd stage.  Thus, that would be totally useless",
          "votes": 2
        }
      ]
    },
    {
      "id": 431082,
      "postDate": "2018-12-01T15:54:08.300Z",
      "content": "<p>I work in machine learning for automation, so most of my experience is in vision and not NLP so take what I am about to say with a grain of salt. In my experience, human judgement gets worse the longer they work on labeling something, and some people will judge things differently. So when working with labelled data where there is any room for subjectivity, there tends to be a chunk of the data (from my experience, it seems to be somewhere between 5-10% in most cases) that are labeled either incorrectly or are so ambiguous that no one could agree on what class it belongs to. If I can, I like to rip them from the training data and create my own special test set from them. Ambiguous samples can be particularly interesting because they can provide some insight into how your model is \"thinking\" about the data and what features map to what class.</p>",
      "rawMarkdown": "I work in machine learning for automation, so most of my experience is in vision and not NLP so take what I am about to say with a grain of salt. In my experience, human judgement gets worse the longer they work on labeling something, and some people will judge things differently. So when working with labelled data where there is any room for subjectivity, there tends to be a chunk of the data (from my experience, it seems to be somewhere between 5-10% in most cases) that are labeled either incorrectly or are so ambiguous that no one could agree on what class it belongs to. If I can, I like to rip them from the training data and create my own special test set from them. Ambiguous samples can be particularly interesting because they can provide some insight into how your model is \"thinking\" about the data and what features map to what class."
    },
    {
      "id": 423740,
      "postDate": "2018-11-19T01:01:38.113Z",
      "content": "<p>\"Wat is 1A?\" could be referring to the NPR program (just with a typo). Though <a href=\"https://www.quora.com/unanswered/What-is-1A\">from some research on Quora</a> it looks like it's tagged with stuff about Batteries. I <em>really</em> wish this dataset included the labels applied to a question.</p>",
      "rawMarkdown": "\"Wat is 1A?\" could be referring to the NPR program (just with a typo). Though [from some research on Quora](https://www.quora.com/unanswered/What-is-1A) it looks like it's tagged with stuff about Batteries. I *really* wish this dataset included the labels applied to a question."
    },
    {
      "id": 422375,
      "postDate": "2018-11-16T06:14:55.080Z",
      "content": "<p>the insincere ones  with shorter word length , looked at few samples  they seem to be rightly labelled </p>",
      "rawMarkdown": "the insincere ones  with shorter word length , looked at few samples  they seem to be rightly labelled "
    },
    {
      "id": 422365,
      "postDate": "2018-11-16T05:56:44.657Z",
      "content": "<p>Found these also In Islam?  What isOrganism? Can coffee?\nI 12? What is ergocalciferol Whatis diphthong? Does Bangladeshis?Whatis extempore?Human needs?\nWhat sexy?Whofound India?What cyber?Google jigsaw?Sykes–Picot Agreement?What meow?Maladaptive daydreaming?Whatis demobilisation?\nMIS reports?Semspa center?\nHello sir?Whatis synergy?\nExplain cryptocurrency?What nudist?\nAre rabbits?Why hospitality?\nWhatis computer?whorote gitanjali?\nWho.ismost powerful.man?What graphic?\nFree Sandeep?Nuclear weapons?\nWhich certification?VJTI TEXTILE?\nWhtis love?Whats nuclear?\nUTGST RATES?Whatis rpm?</p>\n\n<p>there is a definitely  a need for additional features like number of spelling mistakes , word count etc \nWhat do you think ?</p>\n\n<p>How are the ones in the insincere one's</p>",
      "rawMarkdown": "Found these also In Islam?  What isOrganism? Can coffee?\nI 12? What is ergocalciferol Whatis diphthong? Does Bangladeshis?Whatis extempore?Human needs?\nWhat sexy?Whofound India?What cyber?Google jigsaw?Sykes–Picot Agreement?What meow?Maladaptive daydreaming?Whatis demobilisation?\nMIS reports?Semspa center?\nHello sir?Whatis synergy?\nExplain cryptocurrency?What nudist?\nAre rabbits?Why hospitality?\nWhatis computer?whorote gitanjali?\nWho.ismost powerful.man?What graphic?\nFree Sandeep?Nuclear weapons?\nWhich certification?VJTI TEXTILE?\nWhtis love?Whats nuclear?\nUTGST RATES?Whatis rpm?\n\nthere is a definitely  a need for additional features like number of spelling mistakes , word count etc \nWhat do you think ?\n\nHow are the ones in the insincere one's\n"
    },
    {
      "id": 422181,
      "postDate": "2018-11-15T21:57:22.133Z",
      "content": "<p>Interesting findings, and an interesting point. Have you tried if question length is any indication of it getting labeled as insincere?</p>\n\n<p>I suppose many competitions, such as this, are really about just trying to build a classifier for the given dataset. The semantics of \"insincere\" are then left up to debate but not necessarily highly relevant here. Because the definition and true labels are so hard to specify to 100% agreement.</p>",
      "rawMarkdown": "Interesting findings, and an interesting point. Have you tried if question length is any indication of it getting labeled as insincere?\n\nI suppose many competitions, such as this, are really about just trying to build a classifier for the given dataset. The semantics of \"insincere\" are then left up to debate but not necessarily highly relevant here. Because the definition and true labels are so hard to specify to 100% agreement."
    },
    {
      "id": 457422,
      "postDate": "2019-01-17T11:26:32.467Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 421124,
      "author_name": "Emily Boyajian",
      "author_url": "",
      "post_date": "2018-11-14T16:04:39.790000",
      "content": "<p>I'm not sure what proportion of questions are mislabeled, but going through the first 200, here are all the questions that are mislabeled in my opinion (all labeled 0 when they should probably be labeled 1):</p>\n\n<p>\"Is Gaza slowly becoming Auschwitz, Dachau or Treblinka for Palestinians?\"\n\"Have you licked the skin of a corpse?\"\n\"Where is Muhammed now?\"\n\"I wear an insulin pump, and a lot of girls don't like it. Nick Jonas wears one and dated Selena Gomez. Is there difference between me and Nick several zeros missing in my bank account?\"\n\"Does the people who are rich and still claim reservation have any conscience? Why are they given reservation if are rich?\"</p>\n\n<p>From there, I can extrapolate that about 2.5% of questions are mislabeled. Given that about 6.2% of the questions in the training set are labeled 1, about a quarter of \"insincere\" questions are mislabeled.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 422018,
          "author_name": "Paul Sarmadi",
          "author_url": "",
          "post_date": "2018-11-15T17:13:23.983000",
          "content": "<p>Interesting point! Some questions are hard to label even for a human. A lot of those are a matter of personal opinion. Concerning those questions for example:\n-  \"Have you licked the skin of a corpse?\"\n- \"Where is Muhammed now?\"</p>\n\n<p>I would rather tend to say they are sincere. Though not conventional for sure haha.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 437402,
          "author_name": "Ian Oliver",
          "author_url": "",
          "post_date": "2018-12-11T20:58:10.893000",
          "content": "<p>As someone who occasionally asks, answers, and flags questions on Quora, but is not a Quora employee, I would judge all of those as sincere, though some seem to have been asked by people with a lack of perspective or background knowledge.</p>\n\n<p>For comparison, here are some questions I have asked (completely sincerely) that I think (and I may be wrong) have a similar (low, IMO) level of apparent absurdity:</p>\n\n<ul>\n<li><a href=\"https://www.quora.com/Are-there-science-fairs-for-adults\">Are there science fairs for adults?</a></li>\n<li><a href=\"https://www.quora.com/Whats-the-greatest-number-of-citations-youve-ever-seen-in-a-single-document\">What's the greatest number of citations you've ever seen in a single document?</a></li>\n<li><a href=\"https://www.quora.com/What-is-the-distribution-of-values-for-number-of-major-and-minor-characters-per-word-page-in-common-novels\">What is the distribution of values for number of major and minor characters per word/page in common novels?</a></li>\n<li><a href=\"https://www.quora.com/unanswered/What-data-storage-technology-can-store-data-for-12-000-years-with-intermittent-writing-and-reading-and-other-requirements\">What data storage technology can store data for 12,000 years with intermittent writing and reading, and other requirements?</a></li>\n<li><a href=\"https://www.quora.com/Why-doesnt-Quora-have-question-comments-anymore\">Why doesn't Quora have question comments anymore?</a></li>\n</ul>\n\n<p>The last one is interesting because its premise is false—Quora does still have question comments—which is a criterion for insincerity given in the competition rules, but I did ask it sincerely, because I didn't notice the button that you have to click to view question comments. (Immediately after asking the question, of course, I noticed the button and answered my own question.)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420749,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-11-14T04:19:16.983000",
      "content": "<p>Thanks Joe for sharing this!  I have found many mislabeled data in the longer sentences here as well : <a href=\"https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru\">https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru</a></p>\n\n<p>The question is, is there anyway to do about it ? Or we just have to leave it as a noise.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 420526,
      "author_name": "Alec Kulakowski",
      "author_url": "",
      "post_date": "2018-11-13T19:13:14.417000",
      "content": "<p>The kernel page says <code>Quora has employed both machine learning and manual review</code> to identify sincerity in the past, this means that these questions could have been marked sincere by some ML model. Also, Quora questions have a descriptive paragraph after the title, so \"Why is 16?\" could've been followed by a description of some math problem. \"Hello sir?\" could've been followed by a question about etiquette in the scope of ESL. Quora's manual review would certainly see the question descriptions, and perhaps their ML algos would too. </p>\n\n<p>Unfortunately (and for no stated reason), they don't provide it to us in the data for this kernel, so we're at a notable disadvantage when it comes to titles like that. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 437388,
          "author_name": "Ian Oliver",
          "author_url": "",
          "post_date": "2018-12-11T20:20:28.503000",
          "content": "<blockquote>\n  <p>Also, Quora questions have a descriptive paragraph after the title</p>\n</blockquote>\n\n<p>Not anymore: Quora no longer allows question details, and used a bot to move all question details that existed before the change into a comment on the question. (Comments are hidden when you load the page, too—you have to click a button to see them.) The current rule is that the entire question must fit in the title (which has a limited length) and it must stand on its own. You are, however, allowed to include one URL as a \"context link\", to show the answerers what you're asking about, in case it's not obvious.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 420744,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2018-11-14T04:12:19.557000",
      "content": "<p>I don't know man. \"Is God 42?\" sounds like some Vsauce video title. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 426109,
          "author_name": "kagSen",
          "author_url": "",
          "post_date": "2018-11-22T16:20:23.417000",
          "content": "<p>It is probably a reference to the Hitchhiker's guide to the galaxy:</p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy\">https://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 420542,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-11-13T19:42:20.910000",
      "content": "<p>Thats why they need a better model :D</p>",
      "votes": 3,
      "replies": [
        {
          "id": 420903,
          "author_name": "Sreeram",
          "author_url": "",
          "post_date": "2018-11-14T09:54:08.030000",
          "content": "<p>In your opinion, Is it better to correct those samples that are miss-labelled or to leave them as it is .?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421008,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-11-14T13:15:05.997000",
          "content": "<p>I think correcting them (if possible) will make the training easier; however, there is also a downside as the test dataset will also have noise. </p>\n\n<p>In that case, our validation data will not have the same distribution as the test data anymore, and so it will be difficult to judge our performance before submission.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421107,
          "author_name": "Emily Boyajian",
          "author_url": "",
          "post_date": "2018-11-14T15:44:50.380000",
          "content": "<p>I think it's best to not correct the \"mislabeled\" questions. There might be patterns to which questions tend to be mislabeled (for example, very short questions might tended to be labeled \"sincere\" even if they are actually \"insincere\"). The test data would likely have the same patterns of mislabeling, so \"correcting\" the training data might actually decrease the accuracy of our models when applied to the test data.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 421117,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2018-11-14T15:53:26.003000",
          "content": "<p>My point and Ben are exactly the same :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421145,
          "author_name": "Alec Kulakowski",
          "author_url": "",
          "post_date": "2018-11-14T16:37:26.687000",
          "content": "<p>I agree, to achieve best performance on the holdout data, we shouldn't fix the labels. This is an oversight on Quora's part, since the data was labeled not only manually but also by their own, current ML algorithm.  </p>\n\n<p>You're essentially training a ML algorithm on the output of another algorithm, so there's already a hypothetical upper bound to accuracy on real, correctly labeled data. Based on their description of the competition and what they said about the data (in vague terms) I get the sense that there's a chance that the test data is labeled entirely manually vs the training data which is labeled by a mix of manual + ML (could be wrong, but there is a chance). Only way to know for sure is to train the same model first on the current data and then again on the data but with the labels manually fixed (perhaps we could get a running list of all known mislabeled rows) and see which one performs best on the holdout data. </p>\n\n<p>Another thing worth considering is that their initial algorithm (and manual labelers) almost certainly had access to not just the question title, but also the question description, which is (IMHO) necessary to determine the sincerity of a question in certain circumstances. So we're trying to fit an estimator to an existing model that was trained with more data than we have available now. Doesn't bode well, maybe that's why even the best algorithms are getting an F-score of around 0.7. </p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 421998,
          "author_name": "Paul Sarmadi",
          "author_url": "",
          "post_date": "2018-11-15T16:54:40.460000",
          "content": "<p>The possibility that the test set also have mislabeled questions is definitely to be taken into account. However there is no proof that there is a clear pattern in mislabeling. And definitely no proof that one's model -not built for that purpose- would catch it. Which would mean we do learn noise and it would be better to correct it.\nI don't think there is a clear answer to this question.\nA simple idea would be to say \"that option is too much time consuming and I'd rather focus on other parts of the process\".</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 459139,
      "author_name": "ManuelSH",
      "author_url": "",
      "post_date": "2019-01-21T09:11:11.577000",
      "content": "<p>You know the saying: \"garbage in...\"</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 458981,
      "author_name": "Julius Hochmuth",
      "author_url": "",
      "post_date": "2019-01-21T00:16:19.657000",
      "content": "<p>This problem actually seems to occur with other datasets as well. In the paper I linked to, they discuss problems that occur when trying to label toxic comments. They used data from wikipedia and twitter. In regard to the mislabelling of samples, they had this to say: \"We find that 23% of sampled comments in the false negatives of the Wikipedia dataset do not fulfill the toxic definition in our view.\" You can find them discussing this problem in section 6.1.</p>\n\n<p><a href=\"https://arxiv.org/pdf/1809.07572.pdf\">https://arxiv.org/pdf/1809.07572.pdf</a></p>\n\n<p>Anyways, here's a good one:\nWhy does America have to accept people from Sh1t hole countries?,0</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 457491,
      "author_name": "Lalit",
      "author_url": "",
      "post_date": "2019-01-17T14:59:33.363000",
      "content": "<p>The questions listed by OP have 2 or 3 words.</p>\n\n<p>I suspect that Quora has some logic behind (mis)labeling them as sincere for this challenge, and that our code should (mis)label all 2 or 3 words long questions as sincere.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 439486,
      "author_name": "Daniel Cañueto",
      "author_url": "",
      "post_date": "2018-12-15T15:27:34.447000",
      "content": "<p>I have the intuition some of this mislabeling is done on purpose to avoid the influence of manual correction of predicted labels on the test set. If not, it would be way too easy to correct labels manually on the test set.</p>\n\n<p>About some insincere questions wrongly labeled as not insincere, I have tried to correct them and results have not improved. I guess there is too much data for them to be a problem to find the right minimum.</p>\n\n<p>I am confident some insincere questions are not actually detected in the question but in the answer (e.g., for spam). I guess it is very interesting for Quora to try to predict potential spam-related questions before the appearance fo the actual spam.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 439493,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-12-15T15:42:30.327000",
          "content": "<blockquote>\n  <p>I have the intuition some of this mislabeling is done on purpose to avoid the influence of manual correction of predicted labels on the test set. If not, it would be way too easy to correct labels manually on the test set.</p>\n</blockquote>\n\n<p>This is a 2 stages competition, with a new test set at 2nd stage.  Thus, that would be totally useless</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 431082,
      "author_name": "Max Hennick",
      "author_url": "",
      "post_date": "2018-12-01T15:54:08.300000",
      "content": "<p>I work in machine learning for automation, so most of my experience is in vision and not NLP so take what I am about to say with a grain of salt. In my experience, human judgement gets worse the longer they work on labeling something, and some people will judge things differently. So when working with labelled data where there is any room for subjectivity, there tends to be a chunk of the data (from my experience, it seems to be somewhere between 5-10% in most cases) that are labeled either incorrectly or are so ambiguous that no one could agree on what class it belongs to. If I can, I like to rip them from the training data and create my own special test set from them. Ambiguous samples can be particularly interesting because they can provide some insight into how your model is \"thinking\" about the data and what features map to what class.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 423740,
      "author_name": "Rob Rose",
      "author_url": "",
      "post_date": "2018-11-19T01:01:38.113000",
      "content": "<p>\"Wat is 1A?\" could be referring to the NPR program (just with a typo). Though <a href=\"https://www.quora.com/unanswered/What-is-1A\">from some research on Quora</a> it looks like it's tagged with stuff about Batteries. I <em>really</em> wish this dataset included the labels applied to a question.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422375,
      "author_name": "Sunil Vikram",
      "author_url": "",
      "post_date": "2018-11-16T06:14:55.080000",
      "content": "<p>the insincere ones  with shorter word length , looked at few samples  they seem to be rightly labelled </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422365,
      "author_name": "Sunil Vikram",
      "author_url": "",
      "post_date": "2018-11-16T05:56:44.657000",
      "content": "<p>Found these also In Islam?  What isOrganism? Can coffee?\nI 12? What is ergocalciferol Whatis diphthong? Does Bangladeshis?Whatis extempore?Human needs?\nWhat sexy?Whofound India?What cyber?Google jigsaw?Sykes–Picot Agreement?What meow?Maladaptive daydreaming?Whatis demobilisation?\nMIS reports?Semspa center?\nHello sir?Whatis synergy?\nExplain cryptocurrency?What nudist?\nAre rabbits?Why hospitality?\nWhatis computer?whorote gitanjali?\nWho.ismost powerful.man?What graphic?\nFree Sandeep?Nuclear weapons?\nWhich certification?VJTI TEXTILE?\nWhtis love?Whats nuclear?\nUTGST RATES?Whatis rpm?</p>\n\n<p>there is a definitely  a need for additional features like number of spelling mistakes , word count etc \nWhat do you think ?</p>\n\n<p>How are the ones in the insincere one's</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422181,
      "author_name": "averagemn",
      "author_url": "",
      "post_date": "2018-11-15T21:57:22.133000",
      "content": "<p>Interesting findings, and an interesting point. Have you tried if question length is any indication of it getting labeled as insincere?</p>\n\n<p>I suppose many competitions, such as this, are really about just trying to build a classifier for the given dataset. The semantics of \"insincere\" are then left up to debate but not necessarily highly relevant here. Because the definition and true labels are so hard to specify to 100% agreement.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 457422,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-17T11:26:32.467000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "420444": "I'm not sure yet how often this sort of thing happens for longer questions, but for short questions (10 or less chars, of which there are only 37 samples in the training data) there are some blatantly mislabeled samples.\n\nAll of these are labeled 0 (sincere):\n\n\"Is God 42?\"\n\"In Islam?\"\n\"How can I?\"\n\"Am I fake?\"\n\"Wat is 1A?\"\n\"Why is 16?\"\n\"Hello sir?\"\n\"IS 1+1 21?\"\n\nWith any luck this is more funny than it is relevant to model training. But I can't help but wonder how often this mislabeling happens in the entire training corpus, given that even these obviously bad samples get mislabeled.\n\nDoes anyone have findings to share on mislabeling in the broader corpus? I know the labels \"are not guaranteed to be perfect\", but there's a pretty broad spectrum between total noise and not quite perfect - clearly there's plenty of good signal, but I think it's important to keep in mind.",
    "421124": "I'm not sure what proportion of questions are mislabeled, but going through the first 200, here are all the questions that are mislabeled in my opinion (all labeled 0 when they should probably be labeled 1):\n\n\"Is Gaza slowly becoming Auschwitz, Dachau or Treblinka for Palestinians?\"\n\"Have you licked the skin of a corpse?\"\n\"Where is Muhammed now?\"\n\"I wear an insulin pump, and a lot of girls don't like it. Nick Jonas wears one and dated Selena Gomez. Is there difference between me and Nick several zeros missing in my bank account?\"\n\"Does the people who are rich and still claim reservation have any conscience? Why are they given reservation if are rich?\"\n\nFrom there, I can extrapolate that about 2.5% of questions are mislabeled. Given that about 6.2% of the questions in the training set are labeled 1, about a quarter of \"insincere\" questions are mislabeled.\n",
    "420749": "Thanks Joe for sharing this!  I have found many mislabeled data in the longer sentences here as well : https://www.kaggle.com/ratthachat/explore-limits-error-analysis-of-srk-s-glove-gru\n\nThe question is, is there anyway to do about it ? Or we just have to leave it as a noise.",
    "420526": "The kernel page says `Quora has employed both machine learning and manual review` to identify sincerity in the past, this means that these questions could have been marked sincere by some ML model. Also, Quora questions have a descriptive paragraph after the title, so \"Why is 16?\" could've been followed by a description of some math problem. \"Hello sir?\" could've been followed by a question about etiquette in the scope of ESL. Quora's manual review would certainly see the question descriptions, and perhaps their ML algos would too. \n\nUnfortunately (and for no stated reason), they don't provide it to us in the data for this kernel, so we're at a notable disadvantage when it comes to titles like that. ",
    "420744": "I don't know man. \"Is God 42?\" sounds like some Vsauce video title. ",
    "420542": "Thats why they need a better model :D",
    "459139": "You know the saying: \"garbage in...\"",
    "458981": "This problem actually seems to occur with other datasets as well. In the paper I linked to, they discuss problems that occur when trying to label toxic comments. They used data from wikipedia and twitter. In regard to the mislabelling of samples, they had this to say: \"We find that 23% of sampled comments in the false negatives of the Wikipedia dataset do not fulfill the toxic definition in our view.\" You can find them discussing this problem in section 6.1.\n\nhttps://arxiv.org/pdf/1809.07572.pdf\n\nAnyways, here's a good one:\nWhy does America have to accept people from Sh1t hole countries?,0\n",
    "457491": "The questions listed by OP have 2 or 3 words.\n\nI suspect that Quora has some logic behind (mis)labeling them as sincere for this challenge, and that our code should (mis)label all 2 or 3 words long questions as sincere.",
    "439486": "I have the intuition some of this mislabeling is done on purpose to avoid the influence of manual correction of predicted labels on the test set. If not, it would be way too easy to correct labels manually on the test set.\n\nAbout some insincere questions wrongly labeled as not insincere, I have tried to correct them and results have not improved. I guess there is too much data for them to be a problem to find the right minimum.\n\nI am confident some insincere questions are not actually detected in the question but in the answer (e.g., for spam). I guess it is very interesting for Quora to try to predict potential spam-related questions before the appearance fo the actual spam.\n\n\n",
    "431082": "I work in machine learning for automation, so most of my experience is in vision and not NLP so take what I am about to say with a grain of salt. In my experience, human judgement gets worse the longer they work on labeling something, and some people will judge things differently. So when working with labelled data where there is any room for subjectivity, there tends to be a chunk of the data (from my experience, it seems to be somewhere between 5-10% in most cases) that are labeled either incorrectly or are so ambiguous that no one could agree on what class it belongs to. If I can, I like to rip them from the training data and create my own special test set from them. Ambiguous samples can be particularly interesting because they can provide some insight into how your model is \"thinking\" about the data and what features map to what class.",
    "423740": "\"Wat is 1A?\" could be referring to the NPR program (just with a typo). Though [from some research on Quora](https://www.quora.com/unanswered/What-is-1A) it looks like it's tagged with stuff about Batteries. I *really* wish this dataset included the labels applied to a question.",
    "422375": "the insincere ones  with shorter word length , looked at few samples  they seem to be rightly labelled ",
    "422365": "Found these also In Islam?  What isOrganism? Can coffee?\nI 12? What is ergocalciferol Whatis diphthong? Does Bangladeshis?Whatis extempore?Human needs?\nWhat sexy?Whofound India?What cyber?Google jigsaw?Sykes–Picot Agreement?What meow?Maladaptive daydreaming?Whatis demobilisation?\nMIS reports?Semspa center?\nHello sir?Whatis synergy?\nExplain cryptocurrency?What nudist?\nAre rabbits?Why hospitality?\nWhatis computer?whorote gitanjali?\nWho.ismost powerful.man?What graphic?\nFree Sandeep?Nuclear weapons?\nWhich certification?VJTI TEXTILE?\nWhtis love?Whats nuclear?\nUTGST RATES?Whatis rpm?\n\nthere is a definitely  a need for additional features like number of spelling mistakes , word count etc \nWhat do you think ?\n\nHow are the ones in the insincere one's\n",
    "422181": "Interesting findings, and an interesting point. Have you tried if question length is any indication of it getting labeled as insincere?\n\nI suppose many competitions, such as this, are really about just trying to build a classifier for the given dataset. The semantics of \"insincere\" are then left up to debate but not necessarily highly relevant here. Because the definition and true labels are so hard to specify to 100% agreement.",
    "457422": ""
  }
}