{
  "id": 20504,
  "title": "Handling Large Dataset ",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20504",
  "author_name": "",
  "post_date": "2016-04-28T10:45:52.170Z",
  "votes": null,
  "comment_count": 4,
  "views": 1306,
  "content": "<p>Ever Since I started learning <strong>Data Science</strong>, this question has been bothering me for months. I searched it on various <em>Platforms</em>. But I could not get satisfactory answers. </p>\n\n<p>My question is, like other Kaggle Competition, this competition's dataset is also very large. The train file is itself <strong>511MB</strong>. How should I handle this? I mean to ask that, loading the dataset itself will take around 1GB of my RAM. So, Suppose I need to apply Random Forest Algo. How would I be able to do it as it requires huge memory.</p>\n\n<p>So, How do other Kagglers handle these types of huge dataset? Do they take samples? Do they have larger RAM? Do they use optimized code? <strong>Is Handling these types of Huge dataset also an art which I need to master</strong>?</p>\n\n<p>I would really appreciate if someone can give me a generalized answer. \nThanks in advance. :)</p>\n\n<p><strong>My System Specification - Ubuntu 14.04 with 6GB RAM.</strong></p>",
  "messages": [
    {
      "id": "117281",
      "postDate": "04/28/2016 10:45:52",
      "content": "<p>Ever Since I started learning <strong>Data Science</strong>, this question has been bothering me for months. I searched it on various <em>Platforms</em>. But I could not get satisfactory answers. </p>\n\n<p>My question is, like other Kaggle Competition, this competition's dataset is also very large. The train file is itself <strong>511MB</strong>. How should I handle this? I mean to ask that, loading the dataset itself will take around 1GB of my RAM. So, Suppose I need to apply Random Forest Algo. How would I be able to do it as it requires huge memory.</p>\n\n<p>So, How do other Kagglers handle these types of huge dataset? Do they take samples? Do they have larger RAM? Do they use optimized code? <strong>Is Handling these types of Huge dataset also an art which I need to master</strong>?</p>\n\n<p>I would really appreciate if someone can give me a generalized answer. \nThanks in advance. :)</p>\n\n<p><strong>My System Specification - Ubuntu 14.04 with 6GB RAM.</strong></p>",
      "rawMarkdown": "Ever Since I started learning **Data Science**, this question has been bothering me for months. I searched it on various *Platforms*. But I could not get satisfactory answers. \r\n\r\nMy question is, like other Kaggle Competition, this competition's dataset is also very large. The train file is itself **511MB**. How should I handle this? I mean to ask that, loading the dataset itself will take around 1GB of my RAM. So, Suppose I need to apply Random Forest Algo. How would I be able to do it as it requires huge memory.\r\n\r\nSo, How do other Kagglers handle these types of huge dataset? Do they take samples? Do they have larger RAM? Do they use optimized code? **Is Handling these types of Huge dataset also an art which I need to master**?\r\n \r\nI would really appreciate if someone can give me a generalized answer. \r\nThanks in advance. :)\r\n\r\n**My System Specification - Ubuntu 14.04 with 6GB RAM.**",
      "votes": null
    },
    {
      "id": "117294",
      "postDate": "04/28/2016 12:04:16",
      "content": "<p>Hi,</p>\n\n<p>there is a thread with more or less the same topic: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20251/how-are-you-handling-large-data-sets\">Lage data sets</a></p>\n\n<p>Sampling has been done. Use only is_booking lines. Use only 1% of the data etc. Online learning might be a possibility. Scripts parsing the data line by line will work also...</p>\n\n<p>You could also use dimensionality reduction.</p>\n\n<p>Regarding RAM: The more, the better. If I need more than I have I use a cloud computing instance.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nthere is a thread with more or less the same topic: [Lage data sets][1]\r\n\r\nSampling has been done. Use only is_booking lines. Use only 1% of the data etc. Online learning might be a possibility. Scripts parsing the data line by line will work also...\r\n\r\nYou could also use dimensionality reduction.\r\n\r\nRegarding RAM: The more, the better. If I need more than I have I use a cloud computing instance.\r\n\r\nGerhard\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20251/how-are-you-handling-large-data-sets",
      "votes": null
    },
    {
      "id": "117466",
      "postDate": "04/29/2016 07:52:15",
      "content": "<p>The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: <code>memory.limit(size=50000)</code>. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...</p>",
      "rawMarkdown": "The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: `memory.limit(size=50000)`. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...",
      "votes": null
    },
    {
      "id": "117482",
      "postDate": "04/29/2016 10:15:25",
      "content": "<p>[quote=Answer = rand();117466]</p>\n\n<p>The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: <code>memory.limit(size=50000)</code>. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...</p>\n\n<p>[/quote]</p>\n\n<p>memory.limit() gives me a warning. <code>It says it is windows-specific</code>.\nAlso, Does it all comes down to RAM Size? Is n't there another alternative? </p>",
      "rawMarkdown": "[quote=Answer = rand();117466]\r\n\r\nThe memory limit default in R is set based on your RAM, but you can increase it to 50gb using: `memory.limit(size=50000)`. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...\r\n\r\n[/quote]\r\n\r\nmemory.limit() gives me a warning. `It says it is windows-specific`.\r\nAlso, Does it all comes down to RAM Size? Is n't there another alternative?",
      "votes": null
    },
    {
      "id": "117682",
      "postDate": "04/30/2016 07:03:18",
      "content": "<p>[quote=Answer = rand();117466]</p>\n\n<p>The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: <code>memory.limit(size=50000)</code>. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...</p>\n\n<p>[/quote]</p>\n\n<p>you can set <code>memory.limit</code> to any size, if it is larger than available RAM it will user files, which means it is going to be slow, even on an SSD</p>",
      "rawMarkdown": "[quote=Answer = rand();117466]\r\n\r\nThe memory limit default in R is set based on your RAM, but you can increase it to 50gb using: `memory.limit(size=50000)`. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...\r\n\r\n[/quote]\r\n\r\nyou can set `memory.limit` to any size, if it is larger than available RAM it will user files, which means it is going to be slow, even on an SSD",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117294,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/28/2016 12:04:16",
      "content": "<p>Hi,</p>\n\n<p>there is a thread with more or less the same topic: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20251/how-are-you-handling-large-data-sets\">Lage data sets</a></p>\n\n<p>Sampling has been done. Use only is_booking lines. Use only 1% of the data etc. Online learning might be a possibility. Scripts parsing the data line by line will work also...</p>\n\n<p>You could also use dimensionality reduction.</p>\n\n<p>Regarding RAM: The more, the better. If I need more than I have I use a cloud computing instance.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117466,
      "author_name": "theanswerisrand",
      "author_url": "",
      "post_date": "04/29/2016 07:52:15",
      "content": "<p>The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: <code>memory.limit(size=50000)</code>. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117482,
      "author_name": "gau2112",
      "author_url": "",
      "post_date": "04/29/2016 10:15:25",
      "content": "<p>[quote=Answer = rand();117466]</p>\n\n<p>The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: <code>memory.limit(size=50000)</code>. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...</p>\n\n<p>[/quote]</p>\n\n<p>memory.limit() gives me a warning. <code>It says it is windows-specific</code>.\nAlso, Does it all comes down to RAM Size? Is n't there another alternative? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117682,
      "author_name": "alexlzzz",
      "author_url": "",
      "post_date": "04/30/2016 07:03:18",
      "content": "<p>[quote=Answer = rand();117466]</p>\n\n<p>The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: <code>memory.limit(size=50000)</code>. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...</p>\n\n<p>[/quote]</p>\n\n<p>you can set <code>memory.limit</code> to any size, if it is larger than available RAM it will user files, which means it is going to be slow, even on an SSD</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117281": "Ever Since I started learning **Data Science**, this question has been bothering me for months. I searched it on various *Platforms*. But I could not get satisfactory answers. \r\n\r\nMy question is, like other Kaggle Competition, this competition's dataset is also very large. The train file is itself **511MB**. How should I handle this? I mean to ask that, loading the dataset itself will take around 1GB of my RAM. So, Suppose I need to apply Random Forest Algo. How would I be able to do it as it requires huge memory.\r\n\r\nSo, How do other Kagglers handle these types of huge dataset? Do they take samples? Do they have larger RAM? Do they use optimized code? **Is Handling these types of Huge dataset also an art which I need to master**?\r\n \r\nI would really appreciate if someone can give me a generalized answer. \r\nThanks in advance. :)\r\n\r\n**My System Specification - Ubuntu 14.04 with 6GB RAM.**",
    "117294": "Hi,\r\n\r\nthere is a thread with more or less the same topic: [Lage data sets][1]\r\n\r\nSampling has been done. Use only is_booking lines. Use only 1% of the data etc. Online learning might be a possibility. Scripts parsing the data line by line will work also...\r\n\r\nYou could also use dimensionality reduction.\r\n\r\nRegarding RAM: The more, the better. If I need more than I have I use a cloud computing instance.\r\n\r\nGerhard\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20251/how-are-you-handling-large-data-sets",
    "117466": "The memory limit default in R is set based on your RAM, but you can increase it to 50gb using: `memory.limit(size=50000)`. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...",
    "117482": "[quote=Answer = rand();117466]\r\n\r\nThe memory limit default in R is set based on your RAM, but you can increase it to 50gb using: `memory.limit(size=50000)`. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...\r\n\r\n[/quote]\r\n\r\nmemory.limit() gives me a warning. `It says it is windows-specific`.\r\nAlso, Does it all comes down to RAM Size? Is n't there another alternative?",
    "117682": "[quote=Answer = rand();117466]\r\n\r\nThe memory limit default in R is set based on your RAM, but you can increase it to 50gb using: `memory.limit(size=50000)`. This comes at the expense of very slow virtual memory, which will affect performance. You ought to just buy more RAM - you can get 16gb for around $60...\r\n\r\n[/quote]\r\n\r\nyou can set `memory.limit` to any size, if it is larger than available RAM it will user files, which means it is going to be slow, even on an SSD"
  },
  "source": "meta"
}