{
  "id": 72914,
  "title": "insincere words like bitcoin and cryptocurrency",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72914",
  "author_name": "Ryuki Mizoguchi",
  "post_date": "2018-11-28T09:35:48.337000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It seems that the words like 'cryptocurrency' and 'bitcoin'  have impact to insincere question.\nPlease tell me your idea to tackle with this problem</p>",
  "messages": [
    {
      "id": 429447,
      "postDate": "2018-11-28T21:26:09.093Z",
      "content": "<p>The challenge with some of these words is that they don’t exist in the pretrained embeddings. While «Bitcoin» exists in GloVe, «Blockchain» does not, and neither does «cryptocurrencies», all commonly used words in Quora questions.</p>\n\n<p>It can be tempting to map these unfound words to an umbrella term, like «Bitcoin», but that’s not necessarily a great idea. There’s a good chance that more specific words like «Bitcoin» tend to be used more often by trolls than neutral words like «blockchain». Glancing at the training set, that does indeed seem to be the case.</p>\n\n<p>One idea could be to generate some features based on some of the more common words in the training set, especially those with high insincerity rates. You have to be careful not to go overboard with this, or you might overfit.</p>\n\n<p>These are just some thoughts, and you have to validate every feature engineering strategy using the data anyway.</p>",
      "rawMarkdown": "The challenge with some of these words is that they don’t exist in the pretrained embeddings. While «Bitcoin» exists in GloVe, «Blockchain» does not, and neither does «cryptocurrencies», all commonly used words in Quora questions.\n\nIt can be tempting to map these unfound words to an umbrella term, like «Bitcoin», but that’s not necessarily a great idea. There’s a good chance that more specific words like «Bitcoin» tend to be used more often by trolls than neutral words like «blockchain». Glancing at the training set, that does indeed seem to be the case.\n\nOne idea could be to generate some features based on some of the more common words in the training set, especially those with high insincerity rates. You have to be careful not to go overboard with this, or you might overfit.\n\nThese are just some thoughts, and you have to validate every feature engineering strategy using the data anyway.",
      "votes": 4
    },
    {
      "id": 429058,
      "postDate": "2018-11-28T09:35:48.337Z",
      "content": "<p>It seems that the words like 'cryptocurrency' and 'bitcoin'  have impact to insincere question.\nPlease tell me your idea to tackle with this problem</p>",
      "rawMarkdown": "It seems that the words like 'cryptocurrency' and 'bitcoin'  have impact to insincere question.\nPlease tell me your idea to tackle with this problem",
      "votes": 2
    },
    {
      "id": 457413,
      "postDate": "2019-01-17T11:14:15.823Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 429447,
      "author_name": "Håkon Hapnes Strand",
      "author_url": "",
      "post_date": "2018-11-28T21:26:09.093000",
      "content": "<p>The challenge with some of these words is that they don’t exist in the pretrained embeddings. While «Bitcoin» exists in GloVe, «Blockchain» does not, and neither does «cryptocurrencies», all commonly used words in Quora questions.</p>\n\n<p>It can be tempting to map these unfound words to an umbrella term, like «Bitcoin», but that’s not necessarily a great idea. There’s a good chance that more specific words like «Bitcoin» tend to be used more often by trolls than neutral words like «blockchain». Glancing at the training set, that does indeed seem to be the case.</p>\n\n<p>One idea could be to generate some features based on some of the more common words in the training set, especially those with high insincerity rates. You have to be careful not to go overboard with this, or you might overfit.</p>\n\n<p>These are just some thoughts, and you have to validate every feature engineering strategy using the data anyway.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 457413,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-17T11:14:15.823000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "429447": "The challenge with some of these words is that they don’t exist in the pretrained embeddings. While «Bitcoin» exists in GloVe, «Blockchain» does not, and neither does «cryptocurrencies», all commonly used words in Quora questions.\n\nIt can be tempting to map these unfound words to an umbrella term, like «Bitcoin», but that’s not necessarily a great idea. There’s a good chance that more specific words like «Bitcoin» tend to be used more often by trolls than neutral words like «blockchain». Glancing at the training set, that does indeed seem to be the case.\n\nOne idea could be to generate some features based on some of the more common words in the training set, especially those with high insincerity rates. You have to be careful not to go overboard with this, or you might overfit.\n\nThese are just some thoughts, and you have to validate every feature engineering strategy using the data anyway.",
    "429058": "It seems that the words like 'cryptocurrency' and 'bitcoin'  have impact to insincere question.\nPlease tell me your idea to tackle with this problem",
    "457413": ""
  }
}