{
  "id": 75745,
  "title": "Noob: Trouble Processing Large Data",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75745",
  "author_name": "Aaron ",
  "post_date": "2018-12-25T22:23:55.120000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hey all, I'm having some difficulty processing my data. Not sure, if this is a common beginniner struggle. I'm trying to create a featureset for each of the 1.3M test texts. Each text wouuld have 1500 features for the top 1,500 most common words. However, it crashes my computer/notebook when I run it. My code works fine for processing 2,000 of the texts. I've looked into sagemaker/ec2/colab, now am trying to use dask. Am I missing something? could someone point me in the right direction? ;(</p>",
  "messages": [
    {
      "id": 445473,
      "postDate": "2018-12-26T14:35:38.510Z",
      "content": "<pre>`\n<h1>The find_features function will determine which of the 1500 word features are contained in the review</h1>\n\ndef find_features(message):\n    words = word_tokenize(message)\n    features = {}\n    for word in word_features:\n        features[word] = (word in words)\n\n<code>return features\n</code></pre>\n\n<h1>Now lets do it for all the messages</h1>\n\n<p>messages = zip(processed, Y)</p>\n\n<h1>call find_features function for each message</h1>\n\n<p>featuresets = [(find_features(text), label) for (text, label) in messages]'</p>",
      "rawMarkdown": "<pre>`\n# The find_features function will determine which of the 1500 word features are contained in the review\ndef find_features(message):\n    words = word_tokenize(message)\n    features = {}\n    for word in word_features:\n        features[word] = (word in words)\n\n    return features\n# Now lets do it for all the messages\nmessages = zip(processed, Y)\n\n# call find_features function for each message\nfeaturesets = [(find_features(text), label) for (text, label) in messages]'</pre>"
    },
    {
      "id": 445194,
      "postDate": "2018-12-25T22:23:55.120Z",
      "content": "<p>Hey all, I'm having some difficulty processing my data. Not sure, if this is a common beginniner struggle. I'm trying to create a featureset for each of the 1.3M test texts. Each text wouuld have 1500 features for the top 1,500 most common words. However, it crashes my computer/notebook when I run it. My code works fine for processing 2,000 of the texts. I've looked into sagemaker/ec2/colab, now am trying to use dask. Am I missing something? could someone point me in the right direction? ;(</p>",
      "rawMarkdown": "Hey all, I'm having some difficulty processing my data. Not sure, if this is a common beginniner struggle. I'm trying to create a featureset for each of the 1.3M test texts. Each text wouuld have 1500 features for the top 1,500 most common words. However, it crashes my computer/notebook when I run it. My code works fine for processing 2,000 of the texts. I've looked into sagemaker/ec2/colab, now am trying to use dask. Am I missing something? could someone point me in the right direction? ;("
    },
    {
      "id": 445389,
      "postDate": "2018-12-26T11:03:02.403Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 445473,
      "author_name": "Aaron ",
      "author_url": "",
      "post_date": "2018-12-26T14:35:38.510000",
      "content": "<pre>`\n<h1>The find_features function will determine which of the 1500 word features are contained in the review</h1>\n\ndef find_features(message):\n    words = word_tokenize(message)\n    features = {}\n    for word in word_features:\n        features[word] = (word in words)\n\n<code>return features\n</code></pre>\n\n<h1>Now lets do it for all the messages</h1>\n\n<p>messages = zip(processed, Y)</p>\n\n<h1>call find_features function for each message</h1>\n\n<p>featuresets = [(find_features(text), label) for (text, label) in messages]'</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445389,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-26T11:03:02.403000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "445473": "<pre>`\n# The find_features function will determine which of the 1500 word features are contained in the review\ndef find_features(message):\n    words = word_tokenize(message)\n    features = {}\n    for word in word_features:\n        features[word] = (word in words)\n\n    return features\n# Now lets do it for all the messages\nmessages = zip(processed, Y)\n\n# call find_features function for each message\nfeaturesets = [(find_features(text), label) for (text, label) in messages]'</pre>",
    "445194": "Hey all, I'm having some difficulty processing my data. Not sure, if this is a common beginniner struggle. I'm trying to create a featureset for each of the 1.3M test texts. Each text wouuld have 1500 features for the top 1,500 most common words. However, it crashes my computer/notebook when I run it. My code works fine for processing 2,000 of the texts. I've looked into sagemaker/ec2/colab, now am trying to use dask. Am I missing something? could someone point me in the right direction? ;(",
    "445389": ""
  }
}