{
  "id": 54112,
  "title": "Memory Error",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54112",
  "author_name": "",
  "post_date": "2018-04-09T19:54:07.623200500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Is there a workaround to load the whole training set into memory. I'm not able to figure out a way to train on the whole training dataset or should I just use a sample?</p>",
  "messages": [
    {
      "id": "311304",
      "postDate": "04/09/2018 19:54:07",
      "content": "<p>Is there a workaround to load the whole training set into memory. I'm not able to figure out a way to train on the whole training dataset or should I just use a sample?</p>",
      "rawMarkdown": "Is there a workaround to load the whole training set into memory. I'm not able to figure out a way to train on the whole training dataset or should I just use a sample?",
      "votes": null
    },
    {
      "id": "311315",
      "postDate": "04/09/2018 20:18:13",
      "content": "<p>If you're running your script through Kaggle kernels then <strong>kernels</strong>  section would be quite helpful. Solution is to process data depending on RAM. </p>",
      "rawMarkdown": "If you're running your script through Kaggle kernels then **kernels**  section would be quite helpful. Solution is to process data depending on RAM.",
      "votes": null
    },
    {
      "id": "311470",
      "postDate": "04/10/2018 05:53:31",
      "content": "<p>If you're asking for help, it helps if you say which language (Python? R? Scala?) you're using, which ML backend (pandas? dask? HDFS? other) , which machine you were running on, how much RAM and CPU it has.</p>",
      "rawMarkdown": "If you're asking for help, it helps if you say which language (Python? R? Scala?) you're using, which ML backend (pandas? dask? HDFS? other) , which machine you were running on, how much RAM and CPU it has.",
      "votes": null
    },
    {
      "id": "311658",
      "postDate": "04/10/2018 12:51:54",
      "content": "<p>Try Google compute engine, each gmail account has 300 usd credits with which you are able to build a decent VM equiping many cores and huge memory, say, 200GB. </p>",
      "rawMarkdown": "Try Google compute engine, each gmail account has 300 usd credits with which you are able to build a decent VM equiping many cores and huge memory, say, 200GB.",
      "votes": null
    },
    {
      "id": "311698",
      "postDate": "04/10/2018 14:30:18",
      "content": "<p>To answer your question, Python, Pandas, 16 GB RAM i7 </p>",
      "rawMarkdown": "To answer your question, Python, Pandas, 16 GB RAM i7",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 311315,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/09/2018 20:18:13",
      "content": "<p>If you're running your script through Kaggle kernels then <strong>kernels</strong>  section would be quite helpful. Solution is to process data depending on RAM. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311470,
      "author_name": "smcinerney",
      "author_url": "",
      "post_date": "04/10/2018 05:53:31",
      "content": "<p>If you're asking for help, it helps if you say which language (Python? R? Scala?) you're using, which ML backend (pandas? dask? HDFS? other) , which machine you were running on, how much RAM and CPU it has.</p>",
      "votes": null,
      "replies": [
        {
          "id": 311698,
          "author_name": "miteshyadav",
          "author_url": "",
          "post_date": "04/10/2018 14:30:18",
          "content": "<p>To answer your question, Python, Pandas, 16 GB RAM i7 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 311658,
      "author_name": "zhongzishi",
      "author_url": "",
      "post_date": "04/10/2018 12:51:54",
      "content": "<p>Try Google compute engine, each gmail account has 300 usd credits with which you are able to build a decent VM equiping many cores and huge memory, say, 200GB. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "311304": "Is there a workaround to load the whole training set into memory. I'm not able to figure out a way to train on the whole training dataset or should I just use a sample?",
    "311315": "If you're running your script through Kaggle kernels then **kernels**  section would be quite helpful. Solution is to process data depending on RAM.",
    "311470": "If you're asking for help, it helps if you say which language (Python? R? Scala?) you're using, which ML backend (pandas? dask? HDFS? other) , which machine you were running on, how much RAM and CPU it has.",
    "311658": "Try Google compute engine, each gmail account has 300 usd credits with which you are able to build a decent VM equiping many cores and huge memory, say, 200GB.",
    "311698": "To answer your question, Python, Pandas, 16 GB RAM i7"
  },
  "source": "meta"
}