{
  "id": 79606,
  "title": "A strategy that didn't work",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79606",
  "author_name": "",
  "post_date": "2019-02-05T22:23:41.379574500Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When I saw that only 6% of the data is insincere, I spent all my energies on augmenting the data, and then use the best public kernels architectures to get a boost in score.  </p>\n\n<p>It always made the F1 score worse.</p>\n\n<p>The strategy I used was :</p>\n\n<p>Using word similarity, it is possible to generate proper nouns for different peoples, swear words, countries that co-occur in GOOGLENEWS embeddings.  </p>\n\n<p>Now, an insincere question about one set of people would still be insincere if about another set of people, and in the occasional chance of historical inaccuracy,  the statement could still be read as insincere because of what has been said (which doesn't get substituted).  </p>\n\n<p>So, one can keep negative things said about Klingons and say negative things said about Cardassians, and that would still be insincere (usually).  Also, similar with swear words.  </p>\n\n<p>Then, it is possible to mix substitutions to get a multiplicative growth in augmentation based on original insincere questions.  </p>\n\n<p>I was able to increase the dataset to above 2 M samples with 40% insincere questions.  </p>\n\n<p>Not close.  No cigar.  Just burnt at both ends.</p>\n\n<p>Fun project, anyways. :)</p>",
  "messages": [
    {
      "id": "466756",
      "postDate": "02/05/2019 22:23:41",
      "content": "<p>When I saw that only 6% of the data is insincere, I spent all my energies on augmenting the data, and then use the best public kernels architectures to get a boost in score.  </p>\n\n<p>It always made the F1 score worse.</p>\n\n<p>The strategy I used was :</p>\n\n<p>Using word similarity, it is possible to generate proper nouns for different peoples, swear words, countries that co-occur in GOOGLENEWS embeddings.  </p>\n\n<p>Now, an insincere question about one set of people would still be insincere if about another set of people, and in the occasional chance of historical inaccuracy,  the statement could still be read as insincere because of what has been said (which doesn't get substituted).  </p>\n\n<p>So, one can keep negative things said about Klingons and say negative things said about Cardassians, and that would still be insincere (usually).  Also, similar with swear words.  </p>\n\n<p>Then, it is possible to mix substitutions to get a multiplicative growth in augmentation based on original insincere questions.  </p>\n\n<p>I was able to increase the dataset to above 2 M samples with 40% insincere questions.  </p>\n\n<p>Not close.  No cigar.  Just burnt at both ends.</p>\n\n<p>Fun project, anyways. :)</p>",
      "rawMarkdown": "When I saw that only 6% of the data is insincere, I spent all my energies on augmenting the data, and then use the best public kernels architectures to get a boost in score.  \n\nIt always made the F1 score worse.\n\nThe strategy I used was :\n\nUsing word similarity, it is possible to generate proper nouns for different peoples, swear words, countries that co-occur in GOOGLENEWS embeddings.  \n\nNow, an insincere question about one set of people would still be insincere if about another set of people, and in the occasional chance of historical inaccuracy,  the statement could still be read as insincere because of what has been said (which doesn't get substituted).  \n\nSo, one can keep negative things said about Klingons and say negative things said about Cardassians, and that would still be insincere (usually).  Also, similar with swear words.  \n\nThen, it is possible to mix substitutions to get a multiplicative growth in augmentation based on original insincere questions.  \n\nI was able to increase the dataset to above 2 M samples with 40% insincere questions.  \n\nNot close.  No cigar.  Just burnt at both ends.\n\nFun project, anyways. :)",
      "votes": null
    },
    {
      "id": "466758",
      "postDate": "02/05/2019 22:35:42",
      "content": "<p>Funny, I spent what little time I had on this project basically trying to safely do the reverse i.e. trying to drop as much data as possible that wasn't helpful/necessary for training to speed things up.</p>\n\n<p>Good luck in the shake-up.</p>",
      "rawMarkdown": "Funny, I spent what little time I had on this project basically trying to safely do the reverse i.e. trying to drop as much data as possible that wasn't helpful/necessary for training to speed things up.\n\nGood luck in the shake-up.",
      "votes": null
    },
    {
      "id": "466774",
      "postDate": "02/05/2019 22:55:24",
      "content": "<p>I heard the training data was 'noisy'.  I also wonder if stepping aside from ML, whether using a rule-based approach for swear words and other negativity would have made a difference.  Another thing would be the distance of the question for certain topics - politics, ethnicities, and so on.</p>",
      "rawMarkdown": "I heard the training data was 'noisy'.  I also wonder if stepping aside from ML, whether using a rule-based approach for swear words and other negativity would have made a difference.  Another thing would be the distance of the question for certain topics - politics, ethnicities, and so on.",
      "votes": null
    },
    {
      "id": "467407",
      "postDate": "02/07/2019 03:08:04",
      "content": "<p>That sounds like a very sensible approach to me. If it didn't work, maybe it's because of the noisy labels from Quora's own base model.</p>",
      "rawMarkdown": "That sounds like a very sensible approach to me. If it didn't work, maybe it's because of the noisy labels from Quora's own base model.",
      "votes": null
    },
    {
      "id": "468439",
      "postDate": "02/08/2019 21:26:49",
      "content": "<p>Thanks.  Next time, i will work in a team, so i have some bearings.</p>",
      "rawMarkdown": "Thanks.  Next time, i will work in a team, so i have some bearings.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 466758,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "02/05/2019 22:35:42",
      "content": "<p>Funny, I spent what little time I had on this project basically trying to safely do the reverse i.e. trying to drop as much data as possible that wasn't helpful/necessary for training to speed things up.</p>\n\n<p>Good luck in the shake-up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 466774,
          "author_name": "takeseven",
          "author_url": "",
          "post_date": "02/05/2019 22:55:24",
          "content": "<p>I heard the training data was 'noisy'.  I also wonder if stepping aside from ML, whether using a rule-based approach for swear words and other negativity would have made a difference.  Another thing would be the distance of the question for certain topics - politics, ethnicities, and so on.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 467407,
      "author_name": "gamerx",
      "author_url": "",
      "post_date": "02/07/2019 03:08:04",
      "content": "<p>That sounds like a very sensible approach to me. If it didn't work, maybe it's because of the noisy labels from Quora's own base model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 468439,
          "author_name": "takeseven",
          "author_url": "",
          "post_date": "02/08/2019 21:26:49",
          "content": "<p>Thanks.  Next time, i will work in a team, so i have some bearings.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "466756": "When I saw that only 6% of the data is insincere, I spent all my energies on augmenting the data, and then use the best public kernels architectures to get a boost in score.  \n\nIt always made the F1 score worse.\n\nThe strategy I used was :\n\nUsing word similarity, it is possible to generate proper nouns for different peoples, swear words, countries that co-occur in GOOGLENEWS embeddings.  \n\nNow, an insincere question about one set of people would still be insincere if about another set of people, and in the occasional chance of historical inaccuracy,  the statement could still be read as insincere because of what has been said (which doesn't get substituted).  \n\nSo, one can keep negative things said about Klingons and say negative things said about Cardassians, and that would still be insincere (usually).  Also, similar with swear words.  \n\nThen, it is possible to mix substitutions to get a multiplicative growth in augmentation based on original insincere questions.  \n\nI was able to increase the dataset to above 2 M samples with 40% insincere questions.  \n\nNot close.  No cigar.  Just burnt at both ends.\n\nFun project, anyways. :)",
    "466758": "Funny, I spent what little time I had on this project basically trying to safely do the reverse i.e. trying to drop as much data as possible that wasn't helpful/necessary for training to speed things up.\n\nGood luck in the shake-up.",
    "466774": "I heard the training data was 'noisy'.  I also wonder if stepping aside from ML, whether using a rule-based approach for swear words and other negativity would have made a difference.  Another thing would be the distance of the question for certain topics - politics, ethnicities, and so on.",
    "467407": "That sounds like a very sensible approach to me. If it didn't work, maybe it's because of the noisy labels from Quora's own base model.",
    "468439": "Thanks.  Next time, i will work in a team, so i have some bearings."
  },
  "source": "meta"
}