{
  "id": 74776,
  "title": "I am getting memory error while preprocessing training data",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74776",
  "author_name": "",
  "post_date": "2018-12-15T13:34:07.596233500Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>i am trying to convert text data like </p>\n\n<pre><code>train_X = tokenizer.texts_to_matrix(train_qid)\n</code></pre>\n\n<p>but i get MemoryError ,but using texts_to_martix on test data  and cross validation data works fine ,I think using texts_to_matrix on huge set like 1.3M sentances, it gets this kind of error,is there anyway to increase memory?.</p>\n\n<p>then i tried another method also,like spliting the train data and using texts_to_matrix sepratly on each of the splited chunks and appending them, but still i get memory error, is there any work around like this that actually works?</p>\n\n<pre><code>split_values=np.linspace(0,1306122,50)\ntrain_x=[]\nfor i in split_values:\n    temp=tokenizer.texts_to_matrix(train_qid[:int(i)])\n    train_x.append(temp)\n</code></pre>",
  "messages": [
    {
      "id": "439447",
      "postDate": "12/15/2018 13:34:07",
      "content": "<p>i am trying to convert text data like </p>\n\n<pre><code>train_X = tokenizer.texts_to_matrix(train_qid)\n</code></pre>\n\n<p>but i get MemoryError ,but using texts_to_martix on test data  and cross validation data works fine ,I think using texts_to_matrix on huge set like 1.3M sentances, it gets this kind of error,is there anyway to increase memory?.</p>\n\n<p>then i tried another method also,like spliting the train data and using texts_to_matrix sepratly on each of the splited chunks and appending them, but still i get memory error, is there any work around like this that actually works?</p>\n\n<pre><code>split_values=np.linspace(0,1306122,50)\ntrain_x=[]\nfor i in split_values:\n    temp=tokenizer.texts_to_matrix(train_qid[:int(i)])\n    train_x.append(temp)\n</code></pre>",
      "rawMarkdown": "i am trying to convert text data like \n\n    train_X = tokenizer.texts_to_matrix(train_qid)\n\nbut i get MemoryError ,but using texts_to_martix on test data  and cross validation data works fine ,I think using texts_to_matrix on huge set like 1.3M sentances, it gets this kind of error,is there anyway to increase memory?.\n\nthen i tried another method also,like spliting the train data and using texts_to_matrix sepratly on each of the splited chunks and appending them, but still i get memory error, is there any work around like this that actually works?\n\n    split_values=np.linspace(0,1306122,50)\n    train_x=[]\n    for i in split_values:\n        temp=tokenizer.texts_to_matrix(train_qid[:int(i)])\n        train_x.append(temp)",
      "votes": null
    },
    {
      "id": "439473",
      "postDate": "12/15/2018 14:53:13",
      "content": "<p>Use Apache Spark </p>",
      "rawMarkdown": "Use Apache Spark",
      "votes": null
    },
    {
      "id": "439485",
      "postDate": "12/15/2018 15:23:23",
      "content": "<p>i am using kaggle kernel</p>",
      "rawMarkdown": "i am using kaggle kernel",
      "votes": null
    },
    {
      "id": "439528",
      "postDate": "12/15/2018 18:18:47",
      "content": "<p>If you want to use BOW features you should go for scikit learns sparse implementations. What you want to do in keras is probably <code>texts_to_sequences</code>.</p>",
      "rawMarkdown": "If you want to use BOW features you should go for scikit learns sparse implementations. What you want to do in keras is probably `texts_to_sequences`.",
      "votes": null
    },
    {
      "id": "457408",
      "postDate": "01/17/2019 11:08:55",
      "content": "<p><a href=\"/praburocking\">@praburocking</a> collect unused objects to free memory. Use memory efficient tricks such as using <code>numpy</code> rather than list.</p>",
      "rawMarkdown": "praburocking collect unused objects to free memory. Use memory efficient tricks such as using `numpy` rather than list.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 439473,
      "author_name": "pankssid",
      "author_url": "",
      "post_date": "12/15/2018 14:53:13",
      "content": "<p>Use Apache Spark </p>",
      "votes": null,
      "replies": [
        {
          "id": 439485,
          "author_name": "praburocking",
          "author_url": "",
          "post_date": "12/15/2018 15:23:23",
          "content": "<p>i am using kaggle kernel</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 439528,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "12/15/2018 18:18:47",
      "content": "<p>If you want to use BOW features you should go for scikit learns sparse implementations. What you want to do in keras is probably <code>texts_to_sequences</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457408,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "01/17/2019 11:08:55",
      "content": "<p><a href=\"/praburocking\">@praburocking</a> collect unused objects to free memory. Use memory efficient tricks such as using <code>numpy</code> rather than list.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "439447": "i am trying to convert text data like \n\n    train_X = tokenizer.texts_to_matrix(train_qid)\n\nbut i get MemoryError ,but using texts_to_martix on test data  and cross validation data works fine ,I think using texts_to_matrix on huge set like 1.3M sentances, it gets this kind of error,is there anyway to increase memory?.\n\nthen i tried another method also,like spliting the train data and using texts_to_matrix sepratly on each of the splited chunks and appending them, but still i get memory error, is there any work around like this that actually works?\n\n    split_values=np.linspace(0,1306122,50)\n    train_x=[]\n    for i in split_values:\n        temp=tokenizer.texts_to_matrix(train_qid[:int(i)])\n        train_x.append(temp)",
    "439473": "Use Apache Spark",
    "439485": "i am using kaggle kernel",
    "439528": "If you want to use BOW features you should go for scikit learns sparse implementations. What you want to do in keras is probably `texts_to_sequences`.",
    "457408": "praburocking collect unused objects to free memory. Use memory efficient tricks such as using `numpy` rather than list."
  },
  "source": "meta"
}