{
  "id": 24132,
  "title": "How to handle such large datasets",
  "url": "/competitions/outbrain-click-prediction/discussion/24132",
  "author_name": "",
  "post_date": "2016-10-06T09:00:02.590Z",
  "votes": 7,
  "comment_count": 6,
  "views": 1589,
  "content": "<p>datasets around 30GBs cannot be handled in our normal desktops &amp; laptops. What other technology stackes data scientist are considering for this competition</p>",
  "messages": [
    {
      "id": "137955",
      "postDate": "10/06/2016 09:00:02",
      "content": "<p>datasets around 30GBs cannot be handled in our normal desktops &amp; laptops. What other technology stackes data scientist are considering for this competition</p>",
      "rawMarkdown": "datasets around 30GBs cannot be handled in our normal desktops & laptops. What other technology stackes data scientist are considering for this competition",
      "votes": null
    },
    {
      "id": "137959",
      "postDate": "10/06/2016 09:10:19",
      "content": "<p>I haven't yet checked this particular competition.</p>\n\n<p>But many models can be trained iteratively. You can load and unload stuff from memory as you go. You do not need to store the entire 30GB in memory if you cannot. Most neural networks packages allow you to do just that, including the new neural network stuff in sklearn.</p>\n\n<p>Also, if you use Linux, you can create a swap partition in a SSD card or other fast medium, and it automatically uses your pen as extra memory. </p>",
      "rawMarkdown": "I haven't yet checked this particular competition.\r\n\r\nBut many models can be trained iteratively. You can load and unload stuff from memory as you go. You do not need to store the entire 30GB in memory if you cannot. Most neural networks packages allow you to do just that, including the new neural network stuff in sklearn.\r\n\r\nAlso, if you use Linux, you can create a swap partition in a SSD card or other fast medium, and it automatically uses your pen as extra memory.",
      "votes": null
    },
    {
      "id": "147179",
      "postDate": "11/29/2016 08:45:23",
      "content": "<p>One way to go can be to use Apache Spark as it is capable of handling Big Data using it's Resilient Distributed Datasets (RDDs) and most important you can do machine learning using Mlib. </p>\n\n<p>If you know other ways to work with large datasets, please let me know, I'd love to leran more.</p>",
      "rawMarkdown": "One way to go can be to use Apache Spark as it is capable of handling Big Data using it's Resilient Distributed Datasets (RDDs) and most important you can do machine learning using Mlib. \r\n\r\nIf you know other ways to work with large datasets, please let me know, I'd love to leran more.",
      "votes": null
    },
    {
      "id": "147232",
      "postDate": "11/29/2016 16:07:42",
      "content": "<p>[quote=Ashwini Kumar;137955]</p>\n\n<p>datasets around 30GBs cannot be handled in our normal desktops &amp; laptops. What other technology stackes data scientist are considering for this competition</p>\n\n<p>[/quote]</p>\n\n<p>You can always split the large file in smaller files with split. </p>\n\n<blockquote>\n  <p>split -d -b 4GB 'page_views.csv' 'mini_views'\n  Perhaps this is a primitive way to do this, but it's ok to get you started.</p>\n</blockquote>\n\n<p>see: man split for more options </p>",
      "rawMarkdown": "[quote=Ashwini Kumar;137955]\r\n\r\ndatasets around 30GBs cannot be handled in our normal desktops & laptops. What other technology stackes data scientist are considering for this competition\r\n\r\n[/quote]\r\n\r\nYou can always split the large file in smaller files with split. \r\n>  split -d -b 4GB 'page_views.csv' 'mini_views'\r\nPerhaps this is a primitive way to do this, but it's ok to get you started.\r\n\r\nsee: man split for more options",
      "votes": null
    },
    {
      "id": "147245",
      "postDate": "11/29/2016 18:02:58",
      "content": "<p>I'm new to Data Science and not any kind of expert.\nBut I found it's might be good answer to use </p>\n\n<pre><code>pandas.read_csv(chunksize=CHUNKSIZE)\n</code></pre>\n\n<p>People also use </p>\n\n<pre><code>del(data_frame)\n</code></pre>\n\n<p>to save memory!!</p>",
      "rawMarkdown": "I'm new to Data Science and not any kind of expert.\r\nBut I found it's might be good answer to use \r\n\r\n    pandas.read_csv(chunksize=CHUNKSIZE)\r\n\r\nPeople also use \r\n\r\n    del(data_frame)\r\n\r\nto save memory!!",
      "votes": null
    },
    {
      "id": "147394",
      "postDate": "11/30/2016 11:27:17",
      "content": "<p>As Ricardo Cruz mentioned, training iteratively (row-wise or batch wise) is a very good way to handle such huge datasets. </p>\n\n<p>On the algorithms / tools front, apart from NN mentioned above, some other options are</p>\n\n<ol>\n<li>FTRL - <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\">a python implementation</a> </li>\n<li><a href=\"https://github.com/JohnLangford/vowpal_wabbit/wiki\">Vowpal Wabbit</a> </li>\n<li><a href=\"https://github.com/guestwalk/libffm\">FFM</a></li>\n</ol>\n\n<p>One another useful method is to create features from the larger files in an iterative fashion. Then load the new data into memory and use other in-memory algorithms.</p>\n\n<p>Hope this helps.!</p>",
      "rawMarkdown": "As Ricardo Cruz mentioned, training iteratively (row-wise or batch wise) is a very good way to handle such huge datasets. \r\n\r\nOn the algorithms / tools front, apart from NN mentioned above, some other options are\r\n\r\n1. FTRL - [a python implementation][1] \r\n2. [Vowpal Wabbit][2] \r\n3. [FFM][3]\r\n\r\nOne another useful method is to create features from the larger files in an iterative fashion. Then load the new data into memory and use other in-memory algorithms.\r\n\r\nHope this helps.!\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\r\n  [2]: https://github.com/JohnLangford/vowpal_wabbit/wiki\r\n  [3]: https://github.com/guestwalk/libffm",
      "votes": null
    },
    {
      "id": "147574",
      "postDate": "12/01/2016 12:16:33",
      "content": "<p>Another interesting approach, not yet mentioned, is to use an ensemble technique. I was told that the author of Random Forests, Leo Breiman, has an algorithm similar to AdaBoost but aimed for memory limitation situations.</p>\n\n<p>Breiman has a lot of work on bagging, and one algorithm of his (for this memory limited situations) is to load whatever you can to memory. Train your model with that subset of data. Then, keep the most difficult cases in memory and load some more of your data to memory, and add another model to your forest. Rinse and repeat. :)</p>\n\n<p>The algorithm has some of the advantages of gradient boosting trees and can use the entire data. I was told about this algorithm by a professor, I don't know the name of the algorithm or the paper, but I can found out if you are interested.</p>\n\n<p>Anyhow, like SRK said there are libraries that use online training for these situations when you have a single computer with limited memory. I have not used any myself, but they are worth checking out.</p>",
      "rawMarkdown": "Another interesting approach, not yet mentioned, is to use an ensemble technique. I was told that the author of Random Forests, Leo Breiman, has an algorithm similar to AdaBoost but aimed for memory limitation situations.\r\n\r\nBreiman has a lot of work on bagging, and one algorithm of his (for this memory limited situations) is to load whatever you can to memory. Train your model with that subset of data. Then, keep the most difficult cases in memory and load some more of your data to memory, and add another model to your forest. Rinse and repeat. :)\r\n\r\nThe algorithm has some of the advantages of gradient boosting trees and can use the entire data. I was told about this algorithm by a professor, I don't know the name of the algorithm or the paper, but I can found out if you are interested.\r\n\r\nAnyhow, like SRK said there are libraries that use online training for these situations when you have a single computer with limited memory. I have not used any myself, but they are worth checking out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 137959,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "10/06/2016 09:10:19",
      "content": "<p>I haven't yet checked this particular competition.</p>\n\n<p>But many models can be trained iteratively. You can load and unload stuff from memory as you go. You do not need to store the entire 30GB in memory if you cannot. Most neural networks packages allow you to do just that, including the new neural network stuff in sklearn.</p>\n\n<p>Also, if you use Linux, you can create a swap partition in a SSD card or other fast medium, and it automatically uses your pen as extra memory. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 147179,
      "author_name": "davidfumo",
      "author_url": "",
      "post_date": "11/29/2016 08:45:23",
      "content": "<p>One way to go can be to use Apache Spark as it is capable of handling Big Data using it's Resilient Distributed Datasets (RDDs) and most important you can do machine learning using Mlib. </p>\n\n<p>If you know other ways to work with large datasets, please let me know, I'd love to leran more.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 147232,
      "author_name": "marbel",
      "author_url": "",
      "post_date": "11/29/2016 16:07:42",
      "content": "<p>[quote=Ashwini Kumar;137955]</p>\n\n<p>datasets around 30GBs cannot be handled in our normal desktops &amp; laptops. What other technology stackes data scientist are considering for this competition</p>\n\n<p>[/quote]</p>\n\n<p>You can always split the large file in smaller files with split. </p>\n\n<blockquote>\n  <p>split -d -b 4GB 'page_views.csv' 'mini_views'\n  Perhaps this is a primitive way to do this, but it's ok to get you started.</p>\n</blockquote>\n\n<p>see: man split for more options </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 147245,
      "author_name": "mymkyt",
      "author_url": "",
      "post_date": "11/29/2016 18:02:58",
      "content": "<p>I'm new to Data Science and not any kind of expert.\nBut I found it's might be good answer to use </p>\n\n<pre><code>pandas.read_csv(chunksize=CHUNKSIZE)\n</code></pre>\n\n<p>People also use </p>\n\n<pre><code>del(data_frame)\n</code></pre>\n\n<p>to save memory!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 147394,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "11/30/2016 11:27:17",
      "content": "<p>As Ricardo Cruz mentioned, training iteratively (row-wise or batch wise) is a very good way to handle such huge datasets. </p>\n\n<p>On the algorithms / tools front, apart from NN mentioned above, some other options are</p>\n\n<ol>\n<li>FTRL - <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\">a python implementation</a> </li>\n<li><a href=\"https://github.com/JohnLangford/vowpal_wabbit/wiki\">Vowpal Wabbit</a> </li>\n<li><a href=\"https://github.com/guestwalk/libffm\">FFM</a></li>\n</ol>\n\n<p>One another useful method is to create features from the larger files in an iterative fashion. Then load the new data into memory and use other in-memory algorithms.</p>\n\n<p>Hope this helps.!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 147574,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "12/01/2016 12:16:33",
      "content": "<p>Another interesting approach, not yet mentioned, is to use an ensemble technique. I was told that the author of Random Forests, Leo Breiman, has an algorithm similar to AdaBoost but aimed for memory limitation situations.</p>\n\n<p>Breiman has a lot of work on bagging, and one algorithm of his (for this memory limited situations) is to load whatever you can to memory. Train your model with that subset of data. Then, keep the most difficult cases in memory and load some more of your data to memory, and add another model to your forest. Rinse and repeat. :)</p>\n\n<p>The algorithm has some of the advantages of gradient boosting trees and can use the entire data. I was told about this algorithm by a professor, I don't know the name of the algorithm or the paper, but I can found out if you are interested.</p>\n\n<p>Anyhow, like SRK said there are libraries that use online training for these situations when you have a single computer with limited memory. I have not used any myself, but they are worth checking out.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "137955": "datasets around 30GBs cannot be handled in our normal desktops & laptops. What other technology stackes data scientist are considering for this competition",
    "137959": "I haven't yet checked this particular competition.\r\n\r\nBut many models can be trained iteratively. You can load and unload stuff from memory as you go. You do not need to store the entire 30GB in memory if you cannot. Most neural networks packages allow you to do just that, including the new neural network stuff in sklearn.\r\n\r\nAlso, if you use Linux, you can create a swap partition in a SSD card or other fast medium, and it automatically uses your pen as extra memory.",
    "147179": "One way to go can be to use Apache Spark as it is capable of handling Big Data using it's Resilient Distributed Datasets (RDDs) and most important you can do machine learning using Mlib. \r\n\r\nIf you know other ways to work with large datasets, please let me know, I'd love to leran more.",
    "147232": "[quote=Ashwini Kumar;137955]\r\n\r\ndatasets around 30GBs cannot be handled in our normal desktops & laptops. What other technology stackes data scientist are considering for this competition\r\n\r\n[/quote]\r\n\r\nYou can always split the large file in smaller files with split. \r\n>  split -d -b 4GB 'page_views.csv' 'mini_views'\r\nPerhaps this is a primitive way to do this, but it's ok to get you started.\r\n\r\nsee: man split for more options",
    "147245": "I'm new to Data Science and not any kind of expert.\r\nBut I found it's might be good answer to use \r\n\r\n    pandas.read_csv(chunksize=CHUNKSIZE)\r\n\r\nPeople also use \r\n\r\n    del(data_frame)\r\n\r\nto save memory!!",
    "147394": "As Ricardo Cruz mentioned, training iteratively (row-wise or batch wise) is a very good way to handle such huge datasets. \r\n\r\nOn the algorithms / tools front, apart from NN mentioned above, some other options are\r\n\r\n1. FTRL - [a python implementation][1] \r\n2. [Vowpal Wabbit][2] \r\n3. [FFM][3]\r\n\r\nOne another useful method is to create features from the larger files in an iterative fashion. Then load the new data into memory and use other in-memory algorithms.\r\n\r\nHope this helps.!\r\n\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\r\n  [2]: https://github.com/JohnLangford/vowpal_wabbit/wiki\r\n  [3]: https://github.com/guestwalk/libffm",
    "147574": "Another interesting approach, not yet mentioned, is to use an ensemble technique. I was told that the author of Random Forests, Leo Breiman, has an algorithm similar to AdaBoost but aimed for memory limitation situations.\r\n\r\nBreiman has a lot of work on bagging, and one algorithm of his (for this memory limited situations) is to load whatever you can to memory. Train your model with that subset of data. Then, keep the most difficult cases in memory and load some more of your data to memory, and add another model to your forest. Rinse and repeat. :)\r\n\r\nThe algorithm has some of the advantages of gradient boosting trees and can use the entire data. I was told about this algorithm by a professor, I don't know the name of the algorithm or the paper, but I can found out if you are interested.\r\n\r\nAnyhow, like SRK said there are libraries that use online training for these situations when you have a single computer with limited memory. I have not used any myself, but they are worth checking out."
  },
  "source": "meta"
}