{
  "id": 27814,
  "title": "How do you work with the size of the data sets?",
  "url": "/competitions/outbrain-click-prediction/discussion/27814",
  "author_name": "",
  "post_date": "2017-01-17T09:30:08.900Z",
  "votes": null,
  "comment_count": 15,
  "views": 523,
  "content": "<p>Is my computer just not good enough? It takes absolutely forever to do anything with millions and millions of entries</p>",
  "messages": [
    {
      "id": "156611",
      "postDate": "01/17/2017 09:30:08",
      "content": "<p>Is my computer just not good enough? It takes absolutely forever to do anything with millions and millions of entries</p>",
      "rawMarkdown": "Is my computer just not good enough? It takes absolutely forever to do anything with millions and millions of entries",
      "votes": null
    },
    {
      "id": "156619",
      "postDate": "01/17/2017 10:43:51",
      "content": "<p>@ConnorHennen, it is hard to tell without knowing your computer specs. I have an i7 with16Go of RAM, computation times are very long because it is constantly switching between RAM and Swapfile.</p>",
      "rawMarkdown": "ConnorHennen, it is hard to tell without knowing your computer specs. I have an i7 with16Go of RAM, computation times are very long because it is constantly switching between RAM and Swapfile.",
      "votes": null
    },
    {
      "id": "156636",
      "postDate": "01/17/2017 12:46:04",
      "content": "<p>@ConnorHennen It will be much faster if you consider using a cluster on the cloud eg. Google Cloud Platform. </p>",
      "rawMarkdown": "ConnorHennen It will be much faster if you consider using a cluster on the cloud eg. Google Cloud Platform.",
      "votes": null
    },
    {
      "id": "156721",
      "postDate": "01/17/2017 17:53:12",
      "content": "<p>Ok, @Roshan Shetty could you direct me to a tutorial that would show me how to do such a thing? Either way thanks for your help</p>",
      "rawMarkdown": "Ok, @Roshan Shetty could you direct me to a tutorial that would show me how to do such a thing? Either way thanks for your help",
      "votes": null
    },
    {
      "id": "156735",
      "postDate": "01/17/2017 18:34:22",
      "content": "<p>And @kos_, my specs and storage are as attached. Does that help diagnose?</p>",
      "rawMarkdown": "And @kos_, my specs and storage are as attached. Does that help diagnose?",
      "votes": null
    },
    {
      "id": "156740",
      "postDate": "01/17/2017 18:49:31",
      "content": "<p>@ConnorHennen .. This might help <a href=\"https://cloud.google.com/dataproc/docs/tutorials/jupyter-notebook\">https://cloud.google.com/dataproc/docs/tutorials/jupyter-notebook</a></p>",
      "rawMarkdown": "ConnorHennen .. This might help https://cloud.google.com/dataproc/docs/tutorials/jupyter-notebook",
      "votes": null
    },
    {
      "id": "156744",
      "postDate": "01/17/2017 19:09:10",
      "content": "<p>@ConnorHennen, if you are using the whole train data set I guess you have the same issue with your RAM being full.</p>",
      "rawMarkdown": "ConnorHennen, if you are using the whole train data set I guess you have the same issue with your RAM being full.",
      "votes": null
    },
    {
      "id": "156749",
      "postDate": "01/17/2017 19:35:32",
      "content": "<p>In this competition I use only my notebook with 16Gb RAM. To work with such big dataset I do the following:</p>\n\n<ul>\n<li><p>Use ML tools which may work without loading all the data to memory\n(VW is a great example, FFM also has on-disk mode).</p></li>\n<li><p>Process data in streaming fashion - instead of loading all data to memory and transforming it, I load just relational tables (promoted_content, events, and so on) and transform click files line-by-line.</p></li>\n</ul>\n\n<p>Also, I write some code in the C++ to speedup things, but that's not requirement - I started to do so at relatively late stage of competitions, it's possible to get decent results without it.</p>",
      "rawMarkdown": "In this competition I use only my notebook with 16Gb RAM. To work with such big dataset I do the following:\r\n\r\n - Use ML tools which may work without loading all the data to memory\r\n   (VW is a great example, FFM also has on-disk mode).\r\n   \r\n -  Process data in streaming fashion - instead of loading all data to memory and transforming it, I load just relational tables (promoted_content, events, and so on) and transform click files line-by-line.\r\n\r\nAlso, I write some code in the C++ to speedup things, but that's not requirement - I started to do so at relatively late stage of competitions, it's possible to get decent results without it.",
      "votes": null
    },
    {
      "id": "156812",
      "postDate": "01/17/2017 23:11:44",
      "content": "<p>Thanks @Alexey for the feedback. That's what I also figured out during this competitions but I never managed to reduced my computation times with 16Gb RAM. How long does it take to train your model for example?</p>",
      "rawMarkdown": "Thanks @Alexey for the feedback. That's what I also figured out during this competitions but I never managed to reduced my computation times with 16Gb RAM. How long does it take to train your model for example?",
      "votes": null
    },
    {
      "id": "156814",
      "postDate": "01/17/2017 23:15:22",
      "content": "<p>About several hours. It depends on exact model, feature set and so on.</p>",
      "rawMarkdown": "About several hours. It depends on exact model, feature set and so on.",
      "votes": null
    },
    {
      "id": "156829",
      "postDate": "01/18/2017 01:07:31",
      "content": "<p>My notebook has an i7 with 12GB of ram. I'm using MySQL to processes the data, it takes about 1h to create the train and test sets. The train set has 13GB and the test has 4.3G. My model uses about 30 features.</p>\n\n<p>I'm using the FTRL model provided on the kernels because of the online training. To run the full model and build the submission file it takes about 50 minutes. The total memory used to train the model is 8GB.</p>\n\n<p>This is my first competition size wasn't a problem, the real challange was to find good features. My current LB score is: 0.67222</p>\n\n<p>I'm looking foward to the next competition I'm having a blast with kaggle</p>",
      "rawMarkdown": "My notebook has an i7 with 12GB of ram. I'm using MySQL to processes the data, it takes about 1h to create the train and test sets. The train set has 13GB and the test has 4.3G. My model uses about 30 features.\r\n\r\nI'm using the FTRL model provided on the kernels because of the online training. To run the full model and build the submission file it takes about 50 minutes. The total memory used to train the model is 8GB.\r\n\r\nThis is my first competition size wasn't a problem, the real challange was to find good features. My current LB score is: 0.67222\r\n\r\nI'm looking foward to the next competition I'm having a blast with kaggle",
      "votes": null
    },
    {
      "id": "156938",
      "postDate": "01/18/2017 17:26:29",
      "content": "<p>I rented out an AWS server w/ 32gb of ram. That was enough to do everything in memory. I was even able to train an xgboost model w/ 8 features on all the training data in a few minutes.</p>",
      "rawMarkdown": "I rented out an AWS server w/ 32gb of ram. That was enough to do everything in memory. I was even able to train an xgboost model w/ 8 features on all the training data in a few minutes.",
      "votes": null
    },
    {
      "id": "157032",
      "postDate": "01/19/2017 00:07:43",
      "content": "<p>How much does it cost to rent @vape naysh? I went with the free option initially, and it just took forever to even upload the .csv files.</p>",
      "rawMarkdown": "How much does it cost to rent @vape naysh? I went with the free option initially, and it just took forever to even upload the .csv files.",
      "votes": null
    },
    {
      "id": "157139",
      "postDate": "01/19/2017 12:17:11",
      "content": "<p>using LightGBM with 10,000,000 data * 500+ feature, cost about 5GB memory.  </p>",
      "rawMarkdown": "using LightGBM with 10,000,000 data * 500+ feature, cost about 5GB memory.",
      "votes": null
    },
    {
      "id": "157183",
      "postDate": "01/19/2017 15:21:04",
      "content": "<p>I used a Spark cluster to run full pre-processing (including page_views.csv) and feature engineering (1 master and 4 workers with 4 CPUs and 16 GB RAM each). See this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\">Kernel for an example EDA</a> on PySpark, with some info about running this cluster on Google Dataproc.</p>\n\n<p>To train models like XGBoost, LightGBM and LIBFFM, I used a server with 32 CPUs and 256 GB RAM, with SSD disk. These frameworks make good use of computer resources and finishes processing really fast under that configuration. FTRL and VowpalWabbit implement online algorithms, so they require low memory (8GB should be fine).</p>",
      "rawMarkdown": "I used a Spark cluster to run full pre-processing (including page_views.csv) and feature engineering (1 master and 4 workers with 4 CPUs and 16 GB RAM each). See this [Kernel for an example EDA][1] on PySpark, with some info about running this cluster on Google Dataproc.\r\n\r\nTo train models like XGBoost, LightGBM and LIBFFM, I used a server with 32 CPUs and 256 GB RAM, with SSD disk. These frameworks make good use of computer resources and finishes processing really fast under that configuration. FTRL and VowpalWabbit implement online algorithms, so they require low memory (8GB should be fine).\r\n\r\n\r\n  [1]: https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion",
      "votes": null
    },
    {
      "id": "157224",
      "postDate": "01/19/2017 18:47:27",
      "content": "<p>I am running on a desktop with a 4-core CPU and about 60GB RAM, and I use only python. For feature engineering, I find it useful to have the tables sorted (e.g. by uuid). With that, one could do a lot of things just streaming through the data.</p>\n\n<p>The training part, for some strange reason I couldn't get FFM working, so in the end I mostly used FM with sgd, but I suppose FFM would have worked in a similar way. I decided to keep as much in RAM as possible, and I felt the way to do it is to compress the data. I wrote some python code to do row and column sampling and shuffling, and the output is written to /dev/shm compressed using python's gzip library. I also changed the interface of libFM a little bit so it can read compressed data from pipes. To control how much RAM I consume, I can just tune the row sample rate and the compression level. This certainly comes with overhead, but in terms of engineering effort I'd say it's quite cheap.</p>",
      "rawMarkdown": "I am running on a desktop with a 4-core CPU and about 60GB RAM, and I use only python. For feature engineering, I find it useful to have the tables sorted (e.g. by uuid). With that, one could do a lot of things just streaming through the data.\r\n\r\nThe training part, for some strange reason I couldn't get FFM working, so in the end I mostly used FM with sgd, but I suppose FFM would have worked in a similar way. I decided to keep as much in RAM as possible, and I felt the way to do it is to compress the data. I wrote some python code to do row and column sampling and shuffling, and the output is written to /dev/shm compressed using python's gzip library. I also changed the interface of libFM a little bit so it can read compressed data from pipes. To control how much RAM I consume, I can just tune the row sample rate and the compression level. This certainly comes with overhead, but in terms of engineering effort I'd say it's quite cheap.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 156619,
      "author_name": "phansoks",
      "author_url": "",
      "post_date": "01/17/2017 10:43:51",
      "content": "<p>@ConnorHennen, it is hard to tell without knowing your computer specs. I have an i7 with16Go of RAM, computation times are very long because it is constantly switching between RAM and Swapfile.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156636,
      "author_name": "roshanshetty",
      "author_url": "",
      "post_date": "01/17/2017 12:46:04",
      "content": "<p>@ConnorHennen It will be much faster if you consider using a cluster on the cloud eg. Google Cloud Platform. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156721,
      "author_name": "connorvhennen",
      "author_url": "",
      "post_date": "01/17/2017 17:53:12",
      "content": "<p>Ok, @Roshan Shetty could you direct me to a tutorial that would show me how to do such a thing? Either way thanks for your help</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156735,
      "author_name": "connorvhennen",
      "author_url": "",
      "post_date": "01/17/2017 18:34:22",
      "content": "<p>And @kos_, my specs and storage are as attached. Does that help diagnose?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156740,
      "author_name": "roshanshetty",
      "author_url": "",
      "post_date": "01/17/2017 18:49:31",
      "content": "<p>@ConnorHennen .. This might help <a href=\"https://cloud.google.com/dataproc/docs/tutorials/jupyter-notebook\">https://cloud.google.com/dataproc/docs/tutorials/jupyter-notebook</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156744,
      "author_name": "phansoks",
      "author_url": "",
      "post_date": "01/17/2017 19:09:10",
      "content": "<p>@ConnorHennen, if you are using the whole train data set I guess you have the same issue with your RAM being full.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156749,
      "author_name": "alexeynoskov",
      "author_url": "",
      "post_date": "01/17/2017 19:35:32",
      "content": "<p>In this competition I use only my notebook with 16Gb RAM. To work with such big dataset I do the following:</p>\n\n<ul>\n<li><p>Use ML tools which may work without loading all the data to memory\n(VW is a great example, FFM also has on-disk mode).</p></li>\n<li><p>Process data in streaming fashion - instead of loading all data to memory and transforming it, I load just relational tables (promoted_content, events, and so on) and transform click files line-by-line.</p></li>\n</ul>\n\n<p>Also, I write some code in the C++ to speedup things, but that's not requirement - I started to do so at relatively late stage of competitions, it's possible to get decent results without it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156812,
      "author_name": "phansoks",
      "author_url": "",
      "post_date": "01/17/2017 23:11:44",
      "content": "<p>Thanks @Alexey for the feedback. That's what I also figured out during this competitions but I never managed to reduced my computation times with 16Gb RAM. How long does it take to train your model for example?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156814,
      "author_name": "alexeynoskov",
      "author_url": "",
      "post_date": "01/17/2017 23:15:22",
      "content": "<p>About several hours. It depends on exact model, feature set and so on.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156829,
      "author_name": "guilhermesantos",
      "author_url": "",
      "post_date": "01/18/2017 01:07:31",
      "content": "<p>My notebook has an i7 with 12GB of ram. I'm using MySQL to processes the data, it takes about 1h to create the train and test sets. The train set has 13GB and the test has 4.3G. My model uses about 30 features.</p>\n\n<p>I'm using the FTRL model provided on the kernels because of the online training. To run the full model and build the submission file it takes about 50 minutes. The total memory used to train the model is 8GB.</p>\n\n<p>This is my first competition size wasn't a problem, the real challange was to find good features. My current LB score is: 0.67222</p>\n\n<p>I'm looking foward to the next competition I'm having a blast with kaggle</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 156938,
      "author_name": "vapenaysh",
      "author_url": "",
      "post_date": "01/18/2017 17:26:29",
      "content": "<p>I rented out an AWS server w/ 32gb of ram. That was enough to do everything in memory. I was even able to train an xgboost model w/ 8 features on all the training data in a few minutes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157032,
      "author_name": "connorvhennen",
      "author_url": "",
      "post_date": "01/19/2017 00:07:43",
      "content": "<p>How much does it cost to rent @vape naysh? I went with the free option initially, and it just took forever to even upload the .csv files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157139,
      "author_name": "guolinke",
      "author_url": "",
      "post_date": "01/19/2017 12:17:11",
      "content": "<p>using LightGBM with 10,000,000 data * 500+ feature, cost about 5GB memory.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157183,
      "author_name": "gspmoreira",
      "author_url": "",
      "post_date": "01/19/2017 15:21:04",
      "content": "<p>I used a Spark cluster to run full pre-processing (including page_views.csv) and feature engineering (1 master and 4 workers with 4 CPUs and 16 GB RAM each). See this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\">Kernel for an example EDA</a> on PySpark, with some info about running this cluster on Google Dataproc.</p>\n\n<p>To train models like XGBoost, LightGBM and LIBFFM, I used a server with 32 CPUs and 256 GB RAM, with SSD disk. These frameworks make good use of computer resources and finishes processing really fast under that configuration. FTRL and VowpalWabbit implement online algorithms, so they require low memory (8GB should be fine).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157224,
      "author_name": "sangxia",
      "author_url": "",
      "post_date": "01/19/2017 18:47:27",
      "content": "<p>I am running on a desktop with a 4-core CPU and about 60GB RAM, and I use only python. For feature engineering, I find it useful to have the tables sorted (e.g. by uuid). With that, one could do a lot of things just streaming through the data.</p>\n\n<p>The training part, for some strange reason I couldn't get FFM working, so in the end I mostly used FM with sgd, but I suppose FFM would have worked in a similar way. I decided to keep as much in RAM as possible, and I felt the way to do it is to compress the data. I wrote some python code to do row and column sampling and shuffling, and the output is written to /dev/shm compressed using python's gzip library. I also changed the interface of libFM a little bit so it can read compressed data from pipes. To control how much RAM I consume, I can just tune the row sample rate and the compression level. This certainly comes with overhead, but in terms of engineering effort I'd say it's quite cheap.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "156611": "Is my computer just not good enough? It takes absolutely forever to do anything with millions and millions of entries",
    "156619": "ConnorHennen, it is hard to tell without knowing your computer specs. I have an i7 with16Go of RAM, computation times are very long because it is constantly switching between RAM and Swapfile.",
    "156636": "ConnorHennen It will be much faster if you consider using a cluster on the cloud eg. Google Cloud Platform.",
    "156721": "Ok, @Roshan Shetty could you direct me to a tutorial that would show me how to do such a thing? Either way thanks for your help",
    "156735": "And @kos_, my specs and storage are as attached. Does that help diagnose?",
    "156740": "ConnorHennen .. This might help https://cloud.google.com/dataproc/docs/tutorials/jupyter-notebook",
    "156744": "ConnorHennen, if you are using the whole train data set I guess you have the same issue with your RAM being full.",
    "156749": "In this competition I use only my notebook with 16Gb RAM. To work with such big dataset I do the following:\r\n\r\n - Use ML tools which may work without loading all the data to memory\r\n   (VW is a great example, FFM also has on-disk mode).\r\n   \r\n -  Process data in streaming fashion - instead of loading all data to memory and transforming it, I load just relational tables (promoted_content, events, and so on) and transform click files line-by-line.\r\n\r\nAlso, I write some code in the C++ to speedup things, but that's not requirement - I started to do so at relatively late stage of competitions, it's possible to get decent results without it.",
    "156812": "Thanks @Alexey for the feedback. That's what I also figured out during this competitions but I never managed to reduced my computation times with 16Gb RAM. How long does it take to train your model for example?",
    "156814": "About several hours. It depends on exact model, feature set and so on.",
    "156829": "My notebook has an i7 with 12GB of ram. I'm using MySQL to processes the data, it takes about 1h to create the train and test sets. The train set has 13GB and the test has 4.3G. My model uses about 30 features.\r\n\r\nI'm using the FTRL model provided on the kernels because of the online training. To run the full model and build the submission file it takes about 50 minutes. The total memory used to train the model is 8GB.\r\n\r\nThis is my first competition size wasn't a problem, the real challange was to find good features. My current LB score is: 0.67222\r\n\r\nI'm looking foward to the next competition I'm having a blast with kaggle",
    "156938": "I rented out an AWS server w/ 32gb of ram. That was enough to do everything in memory. I was even able to train an xgboost model w/ 8 features on all the training data in a few minutes.",
    "157032": "How much does it cost to rent @vape naysh? I went with the free option initially, and it just took forever to even upload the .csv files.",
    "157139": "using LightGBM with 10,000,000 data * 500+ feature, cost about 5GB memory.",
    "157183": "I used a Spark cluster to run full pre-processing (including page_views.csv) and feature engineering (1 master and 4 workers with 4 CPUs and 16 GB RAM each). See this [Kernel for an example EDA][1] on PySpark, with some info about running this cluster on Google Dataproc.\r\n\r\nTo train models like XGBoost, LightGBM and LIBFFM, I used a server with 32 CPUs and 256 GB RAM, with SSD disk. These frameworks make good use of computer resources and finishes processing really fast under that configuration. FTRL and VowpalWabbit implement online algorithms, so they require low memory (8GB should be fine).\r\n\r\n\r\n  [1]: https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion",
    "157224": "I am running on a desktop with a 4-core CPU and about 60GB RAM, and I use only python. For feature engineering, I find it useful to have the tables sorted (e.g. by uuid). With that, one could do a lot of things just streaming through the data.\r\n\r\nThe training part, for some strange reason I couldn't get FFM working, so in the end I mostly used FM with sgd, but I suppose FFM would have worked in a similar way. I decided to keep as much in RAM as possible, and I felt the way to do it is to compress the data. I wrote some python code to do row and column sampling and shuffling, and the output is written to /dev/shm compressed using python's gzip library. I also changed the interface of libFM a little bit so it can read compressed data from pipes. To control how much RAM I consume, I can just tune the row sample rate and the compression level. This certainly comes with overhead, but in terms of engineering effort I'd say it's quite cheap."
  },
  "source": "meta"
}