{
  "id": 75004,
  "title": "CNN vs Random Forest",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75004",
  "author_name": "",
  "post_date": "2018-12-18T00:05:59.984363400Z",
  "votes": -1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Evening Quora Kaggle Community. I am a new member and curious about competing in this competition. Currently running into run-time issues with my cleaning function for parsing through the data (beautiful soup). \nHowever, after browsing the discussions board often, I see a lot of users are discussion using a CNN for their programs and I am curious why there are not more conversations about implementing random forests. Is their O(n) runtime too long for the competition? are CNNs inherently more effective at NLP? I am curious and would love any feedback. </p>",
  "messages": [
    {
      "id": "440734",
      "postDate": "12/18/2018 00:05:59",
      "content": "<p>Evening Quora Kaggle Community. I am a new member and curious about competing in this competition. Currently running into run-time issues with my cleaning function for parsing through the data (beautiful soup). \nHowever, after browsing the discussions board often, I see a lot of users are discussion using a CNN for their programs and I am curious why there are not more conversations about implementing random forests. Is their O(n) runtime too long for the competition? are CNNs inherently more effective at NLP? I am curious and would love any feedback. </p>",
      "rawMarkdown": "Evening Quora Kaggle Community. I am a new member and curious about competing in this competition. Currently running into run-time issues with my cleaning function for parsing through the data (beautiful soup). \nHowever, after browsing the discussions board often, I see a lot of users are discussion using a CNN for their programs and I am curious why there are not more conversations about implementing random forests. Is their O(n) runtime too long for the competition? are CNNs inherently more effective at NLP? I am curious and would love any feedback.",
      "votes": null
    },
    {
      "id": "441247",
      "postDate": "12/18/2018 13:22:29",
      "content": "<p>In this competition, the Insincere question is often due to the combination of some words in the question. Therefore, CNN1D and max-pooling are very effective. Another reason is CNN runs very fast.</p>",
      "rawMarkdown": "In this competition, the Insincere question is often due to the combination of some words in the question. Therefore, CNN1D and max-pooling are very effective. Another reason is CNN runs very fast.",
      "votes": null
    },
    {
      "id": "441294",
      "postDate": "12/18/2018 14:01:36",
      "content": "<p>Random forests is a good method in applications where the order of features does not matter.\nIn NLP the sequence order matters :-)</p>",
      "rawMarkdown": "Random forests is a good method in applications where the order of features does not matter.\nIn NLP the sequence order matters :-)",
      "votes": null
    },
    {
      "id": "441415",
      "postDate": "12/18/2018 16:29:11",
      "content": "<p>Thank you Andrzej! that does make sense. I recently did a sentiment analysis project for my machine learning class and was curious if you could apply that method to this problem. I ran the training dataset provided in this competition thorugh beautiful_soup then vectorized the character array that was created afterwards. Then performed a PCA.fit function and ran it through the random forest. \nSo doing this method only accounts for individual features instead of a combination of them as suggest by user BAI above in this discussion? </p>",
      "rawMarkdown": "Thank you Andrzej! that does make sense. I recently did a sentiment analysis project for my machine learning class and was curious if you could apply that method to this problem. I ran the training dataset provided in this competition thorugh beautiful_soup then vectorized the character array that was created afterwards. Then performed a PCA.fit function and ran it through the random forest. \nSo doing this method only accounts for individual features instead of a combination of them as suggest by user BAI above in this discussion?",
      "votes": null
    },
    {
      "id": "441417",
      "postDate": "12/18/2018 16:31:39",
      "content": "<p>Thank you Bai, as mentioned with user Andrzej, I did a sentiment analysis project for a class and figured you could apply the same concept to this problem. However, as you stated, is the issue with using PCA's and random forests with this dataset is it does not analyze \"groups\" or multiple characters like a CNN would?\nMy apologies if this is an \"elementary\" question. I am still trying to distinguish from the dozens and dozens of different ways you can incorporate different algorithms into different sets of data. I cannot express how appreciative I am with your response. Thanks again! hope to hear from you soon.</p>",
      "rawMarkdown": "Thank you Bai, as mentioned with user Andrzej, I did a sentiment analysis project for a class and figured you could apply the same concept to this problem. However, as you stated, is the issue with using PCA's and random forests with this dataset is it does not analyze \"groups\" or multiple characters like a CNN would?\nMy apologies if this is an \"elementary\" question. I am still trying to distinguish from the dozens and dozens of different ways you can incorporate different algorithms into different sets of data. I cannot express how appreciative I am with your response. Thanks again! hope to hear from you soon.",
      "votes": null
    },
    {
      "id": "441807",
      "postDate": "12/19/2018 05:19:04",
      "content": "<p>I think this will result in the loss of semantic information, leading to the failure of the pre-training word vector.</p>",
      "rawMarkdown": "I think this will result in the loss of semantic information, leading to the failure of the pre-training word vector.",
      "votes": null
    },
    {
      "id": "442479",
      "postDate": "12/20/2018 02:26:02",
      "content": "<p>Theoretically if you could engineer some ingenious features, Random Forest would be just as effective. But in this particular domain, effective feature engineering is very difficult, so as a result most state of the art approaches utilize Neural Networks and CNNs are both fast and effective at modeling data with locality and order.</p>",
      "rawMarkdown": "Theoretically if you could engineer some ingenious features, Random Forest would be just as effective. But in this particular domain, effective feature engineering is very difficult, so as a result most state of the art approaches utilize Neural Networks and CNNs are both fast and effective at modeling data with locality and order.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 441247,
      "author_name": "xiaobai1123q",
      "author_url": "",
      "post_date": "12/18/2018 13:22:29",
      "content": "<p>In this competition, the Insincere question is often due to the combination of some words in the question. Therefore, CNN1D and max-pooling are very effective. Another reason is CNN runs very fast.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441417,
          "author_name": "nmarsh",
          "author_url": "",
          "post_date": "12/18/2018 16:31:39",
          "content": "<p>Thank you Bai, as mentioned with user Andrzej, I did a sentiment analysis project for a class and figured you could apply the same concept to this problem. However, as you stated, is the issue with using PCA's and random forests with this dataset is it does not analyze \"groups\" or multiple characters like a CNN would?\nMy apologies if this is an \"elementary\" question. I am still trying to distinguish from the dozens and dozens of different ways you can incorporate different algorithms into different sets of data. I cannot express how appreciative I am with your response. Thanks again! hope to hear from you soon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441807,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/19/2018 05:19:04",
          "content": "<p>I think this will result in the loss of semantic information, leading to the failure of the pre-training word vector.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441294,
      "author_name": "akuropatwinski",
      "author_url": "",
      "post_date": "12/18/2018 14:01:36",
      "content": "<p>Random forests is a good method in applications where the order of features does not matter.\nIn NLP the sequence order matters :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 441415,
          "author_name": "nmarsh",
          "author_url": "",
          "post_date": "12/18/2018 16:29:11",
          "content": "<p>Thank you Andrzej! that does make sense. I recently did a sentiment analysis project for my machine learning class and was curious if you could apply that method to this problem. I ran the training dataset provided in this competition thorugh beautiful_soup then vectorized the character array that was created afterwards. Then performed a PCA.fit function and ran it through the random forest. \nSo doing this method only accounts for individual features instead of a combination of them as suggest by user BAI above in this discussion? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442479,
      "author_name": "msf908",
      "author_url": "",
      "post_date": "12/20/2018 02:26:02",
      "content": "<p>Theoretically if you could engineer some ingenious features, Random Forest would be just as effective. But in this particular domain, effective feature engineering is very difficult, so as a result most state of the art approaches utilize Neural Networks and CNNs are both fast and effective at modeling data with locality and order.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "440734": "Evening Quora Kaggle Community. I am a new member and curious about competing in this competition. Currently running into run-time issues with my cleaning function for parsing through the data (beautiful soup). \nHowever, after browsing the discussions board often, I see a lot of users are discussion using a CNN for their programs and I am curious why there are not more conversations about implementing random forests. Is their O(n) runtime too long for the competition? are CNNs inherently more effective at NLP? I am curious and would love any feedback.",
    "441247": "In this competition, the Insincere question is often due to the combination of some words in the question. Therefore, CNN1D and max-pooling are very effective. Another reason is CNN runs very fast.",
    "441294": "Random forests is a good method in applications where the order of features does not matter.\nIn NLP the sequence order matters :-)",
    "441415": "Thank you Andrzej! that does make sense. I recently did a sentiment analysis project for my machine learning class and was curious if you could apply that method to this problem. I ran the training dataset provided in this competition thorugh beautiful_soup then vectorized the character array that was created afterwards. Then performed a PCA.fit function and ran it through the random forest. \nSo doing this method only accounts for individual features instead of a combination of them as suggest by user BAI above in this discussion?",
    "441417": "Thank you Bai, as mentioned with user Andrzej, I did a sentiment analysis project for a class and figured you could apply the same concept to this problem. However, as you stated, is the issue with using PCA's and random forests with this dataset is it does not analyze \"groups\" or multiple characters like a CNN would?\nMy apologies if this is an \"elementary\" question. I am still trying to distinguish from the dozens and dozens of different ways you can incorporate different algorithms into different sets of data. I cannot express how appreciative I am with your response. Thanks again! hope to hear from you soon.",
    "441807": "I think this will result in the loss of semantic information, leading to the failure of the pre-training word vector.",
    "442479": "Theoretically if you could engineer some ingenious features, Random Forest would be just as effective. But in this particular domain, effective feature engineering is very difficult, so as a result most state of the art approaches utilize Neural Networks and CNNs are both fast and effective at modeling data with locality and order."
  },
  "source": "meta"
}