{
  "id": 20251,
  "title": "How are you handling large data sets?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20251",
  "author_name": "",
  "post_date": "2016-04-19T07:22:22.270Z",
  "votes": 5,
  "comment_count": 35,
  "views": 8802,
  "content": "<p>I am trying to do this in R or Python. I am interested to know how the kaggle community is handling big data sets. Below are few thoughts I have</p>\n\n<ol>\n<li>Take a subset of the database, build a model and then run the model\non the entire database using kaggle scripts</li>\n<li>Use packaged to handle big data sets such as &quot;<strong>ff</strong>&quot; for <strong>R</strong> or &quot;<strong>Blaze</strong>&quot; for <strong>python</strong> and do everything on your own system. </li>\n</ol>\n\n<p>I want do everything on my own system but so far I have found that these so called big data packages are not very helpful. Some are taking too long and some have memory errors or some other errors. I have been stuck in this phase for lmost a day. I haven't worked much on big data sets. I have't used kaggle scripts much either.</p>\n\n<p>What would be best way to work on this competition?</p>",
  "messages": [
    {
      "id": "115574",
      "postDate": "04/19/2016 07:22:22",
      "content": "<p>I am trying to do this in R or Python. I am interested to know how the kaggle community is handling big data sets. Below are few thoughts I have</p>\n\n<ol>\n<li>Take a subset of the database, build a model and then run the model\non the entire database using kaggle scripts</li>\n<li>Use packaged to handle big data sets such as &quot;<strong>ff</strong>&quot; for <strong>R</strong> or &quot;<strong>Blaze</strong>&quot; for <strong>python</strong> and do everything on your own system. </li>\n</ol>\n\n<p>I want do everything on my own system but so far I have found that these so called big data packages are not very helpful. Some are taking too long and some have memory errors or some other errors. I have been stuck in this phase for lmost a day. I haven't worked much on big data sets. I have't used kaggle scripts much either.</p>\n\n<p>What would be best way to work on this competition?</p>",
      "rawMarkdown": "I am trying to do this in R or Python. I am interested to know how the kaggle community is handling big data sets. Below are few thoughts I have\r\n\r\n 1. Take a subset of the database, build a model and then run the model\r\n    on the entire database using kaggle scripts\r\n 2. Use packaged to handle big data sets such as \"**ff**\" for **R** or \"**Blaze**\" for **python** and do everything on your own system. \r\n\r\nI want do everything on my own system but so far I have found that these so called big data packages are not very helpful. Some are taking too long and some have memory errors or some other errors. I have been stuck in this phase for lmost a day. I haven't worked much on big data sets. I have't used kaggle scripts much either.\r\n\r\nWhat would be best way to work on this competition?",
      "votes": null
    },
    {
      "id": "115590",
      "postDate": "04/19/2016 08:29:57",
      "content": "<p>Hi,</p>\n\n<p>I think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.</p>\n\n<p>Yesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two &quot;normal&quot; matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.</p>\n\n<p>And then we have not processed/used the &quot;click&quot; data.</p>\n\n<p>So my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.</p>\n\n<p>Another idea would be to use H2O with a local &quot;cluster&quot;! I think I will look into H2O for this one again.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nI think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.\r\n\r\nYesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two \"normal\" matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.\r\n\r\nAnd then we have not processed/used the \"click\" data.\r\n\r\nSo my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.\r\n\r\nAnother idea would be to use H2O with a local \"cluster\"! I think I will look into H2O for this one again.\r\n\r\nCheers\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "115595",
      "postDate": "04/19/2016 08:58:58",
      "content": "<p>[quote=MightyBird;115590]</p>\n\n<p>Hi,</p>\n\n<p>I think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.</p>\n\n<p>Yesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two &quot;normal&quot; matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.</p>\n\n<p>And then we have not processed/used the &quot;click&quot; data.</p>\n\n<p>So my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.</p>\n\n<p>Another idea would be to use H2O with a local &quot;cluster&quot;! I think I will look into H2O for this one again.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>\n\n<p>[/quote]</p>\n\n<p>I will start by working on a subset, while waiting for suggestions from the community. I think I have spent far too much time already.</p>",
      "rawMarkdown": "[quote=MightyBird;115590]\r\n\r\nHi,\r\n\r\nI think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.\r\n\r\nYesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two \"normal\" matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.\r\n\r\nAnd then we have not processed/used the \"click\" data.\r\n\r\nSo my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.\r\n\r\nAnother idea would be to use H2O with a local \"cluster\"! I think I will look into H2O for this one again.\r\n\r\nCheers\r\n\r\nGerhard\r\n\r\n[/quote]\r\n\r\nI will start by working on a subset, while waiting for suggestions from the community. I think I have spent far too much time already.",
      "votes": null
    },
    {
      "id": "115639",
      "postDate": "04/19/2016 13:01:39",
      "content": "<p>If you go to &quot;New Script&quot; in the Kaggle dash for this competition and enter the standard first line</p>\n\n<pre><code>train = pd.read_csv(&quot;../input/train.csv&quot;)\n</code></pre>\n\n<p>It will time out, &quot;Error: The script was killed, likely for trying to exceed the memory limit of 8192M.&quot;</p>\n\n<p>Soooooo is the Kaggle contest submission process different? I hope so. </p>",
      "rawMarkdown": "If you go to \"New Script\" in the Kaggle dash for this competition and enter the standard first line\r\n\r\n    train = pd.read_csv(\"../input/train.csv\")\r\n\r\nIt will time out, \"Error: The script was killed, likely for trying to exceed the memory limit of 8192M.\"\r\n\r\nSoooooo is the Kaggle contest submission process different? I hope so.",
      "votes": null
    },
    {
      "id": "115640",
      "postDate": "04/19/2016 13:12:34",
      "content": "<p>Some people use a database to vectorize train and test samples. It is unlikely that a processed dataset is the same size as the raw data.</p>\n\n<p>Subsampling only bookings, like suggested by MightyBird, makes a lot of sense.</p>\n\n<p>Finally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.</p>",
      "rawMarkdown": "Some people use a database to vectorize train and test samples. It is unlikely that a processed dataset is the same size as the raw data.\r\n\r\nSubsampling only bookings, like suggested by MightyBird, makes a lot of sense.\r\n\r\nFinally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.",
      "votes": null
    },
    {
      "id": "115652",
      "postDate": "04/19/2016 14:07:51",
      "content": "<p>I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\ndf = dd.read_csv('*-*train.csv')\ndf.groupby(df.value.mean().compute()</p>\n\n<p>Check it out <a href=\"http://dask.pydata.org/en/latest/\">http://dask.pydata.org/en/latest/</a></p>",
      "rawMarkdown": "I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\r\ndf = dd.read_csv('*-*train.csv')\r\ndf.groupby(df.value.mean().compute()\r\n\r\nCheck it out http://dask.pydata.org/en/latest/",
      "votes": null
    },
    {
      "id": "115659",
      "postDate": "04/19/2016 14:41:48",
      "content": "<p>One option is to process the dataset in chunks:</p>\n\n<ul>\n<li>For pandas, use <code>pd.read_csv('train.csv', chunksize=chunksize)</code>;</li>\n<li>For XGBoost use <code>model = xgb.train(..., xgb_model=model)</code>, assuming the returned model is used on the next train rounds.</li>\n</ul>\n\n<p>I'm not sure how it influences the model performance (as the train occurs in only a subset of the whole daatset, thus forcing a <code>subsample</code> parameter), but the memory footprint is reduced.</p>",
      "rawMarkdown": "One option is to process the dataset in chunks:\r\n\r\n- For pandas, use `pd.read_csv('train.csv', chunksize=chunksize)`;\r\n- For XGBoost use `model = xgb.train(..., xgb_model=model)`, assuming the returned model is used on the next train rounds.\r\n\r\nI'm not sure how it influences the model performance (as the train occurs in only a subset of the whole daatset, thus forcing a `subsample` parameter), but the memory footprint is reduced.",
      "votes": null
    },
    {
      "id": "115668",
      "postDate": "04/19/2016 15:10:17",
      "content": "<p>So, another experiment just took place...</p>\n\n<pre><code>2.5 M rows test\n3.0 M rows train\n98 columns each\n</code></pre>\n\n<p>32 GB of RAM was not enough for xgboost in R. Sad :-(</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "So, another experiment just took place...\r\n\r\n    2.5 M rows test\r\n    3.0 M rows train\r\n    98 columns each\r\n\r\n32 GB of RAM was not enough for xgboost in R. Sad :-(\r\n\r\nCheers\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "115669",
      "postDate": "04/19/2016 15:11:34",
      "content": "<p>What data types are you using? \nHow many cores?</p>",
      "rawMarkdown": "What data types are you using? \r\nHow many cores?",
      "votes": null
    },
    {
      "id": "115677",
      "postDate": "04/19/2016 15:35:00",
      "content": "<p>I was starting to think of using a different data representation right now. Used a normal matrix so far. 8 cores...</p>",
      "rawMarkdown": "I was starting to think of using a different data representation right now. Used a normal matrix so far. 8 cores...",
      "votes": null
    },
    {
      "id": "115682",
      "postDate": "04/19/2016 15:53:49",
      "content": "<p>A sparse matrix didn't help. It's the algorithm that's greedy. Will now downsample the train set. Starting with 50%.</p>",
      "rawMarkdown": "A sparse matrix didn't help. It's the algorithm that's greedy. Will now downsample the train set. Starting with 50%.",
      "votes": null
    },
    {
      "id": "115690",
      "postDate": "04/19/2016 16:20:25",
      "content": "<p>So, with 50% of train data I got it working. Uses ~20GB now. Algorithm is running - awfully slow although...</p>\n\n<p>It seems there is a problem. The xgboost results get worse with every round. Don't think I have seen this before. Maybe eta 0.5 is too high?</p>",
      "rawMarkdown": "So, with 50% of train data I got it working. Uses ~20GB now. Algorithm is running - awfully slow although...\r\n\r\nIt seems there is a problem. The xgboost results get worse with every round. Don't think I have seen this before. Maybe eta 0.5 is too high?",
      "votes": null
    },
    {
      "id": "115693",
      "postDate": "04/19/2016 16:34:55",
      "content": "<p>[quote=Triskelion;115640]</p>\n\n<p>Finally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\n[/quote]</p>\n\n<p>Are you referring to AWS or some other cloud computing services?</p>",
      "rawMarkdown": "[quote=Triskelion;115640]\r\n\r\nFinally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\r\n[/quote]\r\n\r\nAre you referring to AWS or some other cloud computing services?",
      "votes": null
    },
    {
      "id": "115694",
      "postDate": "04/19/2016 16:48:03",
      "content": "<p>[quote=William Rudebusch;115693]</p>\n\n<p>[quote=Triskelion;115640]</p>\n\n<p>Finally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\n[/quote]</p>\n\n<p>Are you referring to AWS or some other cloud computing services?</p>\n\n<p>[/quote]</p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Online_machine_learning\">https://en.wikipedia.org/wiki/Online_machine_learning</a></p>",
      "rawMarkdown": "[quote=William Rudebusch;115693]\r\n\r\n[quote=Triskelion;115640]\r\n\r\nFinally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\r\n[/quote]\r\n\r\nAre you referring to AWS or some other cloud computing services?\r\n\r\n[/quote]\r\n\r\nhttps://en.wikipedia.org/wiki/Online_machine_learning",
      "votes": null
    },
    {
      "id": "115792",
      "postDate": "04/20/2016 05:33:57",
      "content": "<p>[quote=ashish trehan;115652]</p>\n\n<p>I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\ndf = dd.read_csv('*-*train.csv')\ndf.groupby(df.value.mean().compute()</p>\n\n<p>Check it out <a href=\"http://dask.pydata.org/en/latest/\">http://dask.pydata.org/en/latest/</a></p>\n\n<p>[/quote]</p>\n\n<p>After loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?</p>",
      "rawMarkdown": "[quote=ashish trehan;115652]\r\n\r\nI recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\r\ndf = dd.read_csv('*-*train.csv')\r\ndf.groupby(df.value.mean().compute()\r\n\r\nCheck it out http://dask.pydata.org/en/latest/\r\n\r\n[/quote]\r\n\r\nAfter loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?",
      "votes": null
    },
    {
      "id": "115830",
      "postDate": "04/20/2016 11:55:14",
      "content": "<p>[quote=MightyBird;115590]</p>\n\n<p>Another idea would be to use H2O with a local &quot;cluster&quot;! I think I will look into H2O for this one again.</p>\n\n<p>[/quote]</p>\n\n<p>I tried that and even with a sizable amount of resources, this is still pretty huge data. H2O  could certainly do whatever you want it to with this size of data, but if you use the entire dataset and models with lots of trees then the competition will still be over by the time it finishes lol.</p>\n\n<p>So, I guess subsetting is the way to go. </p>",
      "rawMarkdown": "[quote=MightyBird;115590]\r\n\r\n\r\nAnother idea would be to use H2O with a local \"cluster\"! I think I will look into H2O for this one again.\r\n\r\n\r\n[/quote]\r\n\r\nI tried that and even with a sizable amount of resources, this is still pretty huge data. H2O  could certainly do whatever you want it to with this size of data, but if you use the entire dataset and models with lots of trees then the competition will still be over by the time it finishes lol.\r\n\r\nSo, I guess subsetting is the way to go.",
      "votes": null
    },
    {
      "id": "115832",
      "postDate": "04/20/2016 12:06:19",
      "content": "<p>Hi.</p>\n\n<p>I did a first xgboost model using only train data where is_booking == 1. I get stuck in the area of MAP@5 = 0.2</p>\n\n<p>I therefore think that we must also look at click data!</p>\n\n<p>My next step will thus be to random sample the whole train set.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi.\r\n\r\nI did a first xgboost model using only train data where is_booking == 1. I get stuck in the area of MAP@5 = 0.2\r\n\r\nI therefore think that we must also look at click data!\r\n\r\nMy next step will thus be to random sample the whole train set.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "115842",
      "postDate": "04/20/2016 12:49:12",
      "content": "<p>[quote=Brox;115792]</p>\n\n<p>[quote=ashish trehan;115652]</p>\n\n<p>I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\ndf = dd.read_csv('*-*train.csv')\ndf.groupby(df.value.mean().compute()</p>\n\n<p>Check it out <a href=\"http://dask.pydata.org/en/latest/\">http://dask.pydata.org/en/latest/</a></p>\n\n<p>[/quote]</p>\n\n<p>After loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "[quote=Brox;115792]\r\n\r\n[quote=ashish trehan;115652]\r\n\r\nI recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\r\ndf = dd.read_csv('*-*train.csv')\r\ndf.groupby(df.value.mean().compute()\r\n\r\nCheck it out http://dask.pydata.org/en/latest/\r\n\r\n[/quote]\r\n\r\nAfter loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "115844",
      "postDate": "04/20/2016 12:50:57",
      "content": "<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>",
      "rawMarkdown": "Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.",
      "votes": null
    },
    {
      "id": "115967",
      "postDate": "04/21/2016 03:14:13",
      "content": "<p>[quote=ashish trehan;115844]</p>\n\n<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, I loaded the entire training set.</p>\n\n<p>is loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)</p>",
      "rawMarkdown": "[quote=ashish trehan;115844]\r\n\r\nCan you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.\r\n\r\n[/quote]\r\n\r\nYes, I loaded the entire training set.\r\n\r\nis loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)",
      "votes": null
    },
    {
      "id": "115985",
      "postDate": "04/21/2016 07:20:21",
      "content": "<p>Coming back to number of cores, reduce it to half and see how it goes. As far as I remember for xgboost more cores =&gt; more memory needed.</p>",
      "rawMarkdown": "Coming back to number of cores, reduce it to half and see how it goes. As far as I remember for xgboost more cores => more memory needed.",
      "votes": null
    },
    {
      "id": "115986",
      "postDate": "04/21/2016 08:09:14",
      "content": "<p>I will try that. Would be good - on the memory side. Not so good on the runtime side though.</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "I will try that. Would be good - on the memory side. Not so good on the runtime side though.\r\n\r\nThanks\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "116059",
      "postDate": "04/21/2016 18:47:57",
      "content": "<p>[quote=Brox;115967]</p>\n\n<p>[quote=ashish trehan;115844]</p>\n\n<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, I loaded the entire training set.</p>\n\n<p>is loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)</p>\n\n<p>[/quote]</p>\n\n<p>Do you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.  </p>",
      "rawMarkdown": "[quote=Brox;115967]\r\n\r\n[quote=ashish trehan;115844]\r\n\r\nCan you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.\r\n\r\n[/quote]\r\n\r\nYes, I loaded the entire training set.\r\n\r\nis loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)\r\n\r\n[/quote]\r\n\r\n\r\nDo you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.",
      "votes": null
    },
    {
      "id": "116112",
      "postDate": "04/22/2016 04:17:59",
      "content": "<p>[quote=ashish trehan;116059]</p>\n\n<p>[quote=Brox;115967]</p>\n\n<p>[quote=ashish trehan;115844]</p>\n\n<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, I loaded the entire training set.</p>\n\n<p>is loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)</p>\n\n<p>[/quote]</p>\n\n<p>Do you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.  </p>\n\n<p>[/quote]</p>\n\n<p>That;s what I am doing now, either reading only few columns at a time or reading only few columns at a time.</p>",
      "rawMarkdown": "[quote=ashish trehan;116059]\r\n\r\n[quote=Brox;115967]\r\n\r\n[quote=ashish trehan;115844]\r\n\r\nCan you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.\r\n\r\n[/quote]\r\n\r\nYes, I loaded the entire training set.\r\n\r\nis loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)\r\n\r\n[/quote]\r\n\r\n\r\nDo you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.  \r\n\r\n\r\n[/quote]\r\n\r\nThat;s what I am doing now, either reading only few columns at a time or reading only few columns at a time.",
      "votes": null
    },
    {
      "id": "116121",
      "postDate": "04/22/2016 05:39:29",
      "content": "<p>Short update. Using only 1% of train and after feature engineering about 180 columns my cv in xgboost uses about 6GB on a two core machine. Still running...  I think this might well take 24 hours.</p>",
      "rawMarkdown": "Short update. Using only 1% of train and after feature engineering about 180 columns my cv in xgboost uses about 6GB on a two core machine. Still running...  I think this might well take 24 hours.",
      "votes": null
    },
    {
      "id": "117087",
      "postDate": "04/27/2016 10:55:35",
      "content": "<p>For xgboost I created dense feature set by subsampling training set using only is_booking =1 and some feature engineering techs and I could keep total memory usage with in 8GB.  I also did feature reduction on destination features.\nFor linear models such as logistic regression, sgd classifier or mlp, I onehot encoded the categorical features and generated a sparse matrix with dimension  size about 1200000. It is really time consuming to train a lr model or using mini batch sgd to train sgd classifier or mlp.\nNot yet tried any bagging things. It might work or might not.</p>",
      "rawMarkdown": "For xgboost I created dense feature set by subsampling training set using only is_booking =1 and some feature engineering techs and I could keep total memory usage with in 8GB.  I also did feature reduction on destination features.\r\nFor linear models such as logistic regression, sgd classifier or mlp, I onehot encoded the categorical features and generated a sparse matrix with dimension  size about 1200000. It is really time consuming to train a lr model or using mini batch sgd to train sgd classifier or mlp.\r\nNot yet tried any bagging things. It might work or might not.",
      "votes": null
    },
    {
      "id": "120077",
      "postDate": "05/15/2016 06:02:10",
      "content": "<p>I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?</p>",
      "rawMarkdown": "I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?",
      "votes": null
    },
    {
      "id": "120084",
      "postDate": "05/15/2016 08:12:11",
      "content": "<p>[quote=dun;120077]</p>\n\n<p>I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?</p>\n\n<p>[/quote]</p>\n\n<p>Here you go: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark</a></p>",
      "rawMarkdown": "[quote=dun;120077]\r\n\r\nI'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?\r\n\r\n[/quote]\r\n\r\nHere you go: https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark",
      "votes": null
    },
    {
      "id": "120104",
      "postDate": "05/15/2016 12:52:19",
      "content": "<p>has anyone tried online learning? does it work?</p>",
      "rawMarkdown": "has anyone tried online learning? does it work?",
      "votes": null
    },
    {
      "id": "120105",
      "postDate": "05/15/2016 12:54:19",
      "content": "<p>I have, via Vowpal. Nothing to write home about in terms of results.</p>",
      "rawMarkdown": "I have, via Vowpal. Nothing to write home about in terms of results.",
      "votes": null
    },
    {
      "id": "120127",
      "postDate": "05/15/2016 18:13:00",
      "content": "<p>[quote=dun;120077]</p>\n\n<p>I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?</p>\n\n<p>[/quote]\nBecause it doesn't work (in theory and in practice).</p>",
      "rawMarkdown": "[quote=dun;120077]\r\n\r\nI'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?\r\n\r\n[/quote]\r\nBecause it doesn't work (in theory and in practice).",
      "votes": null
    },
    {
      "id": "120132",
      "postDate": "05/15/2016 19:02:32",
      "content": "<p>Hi Mike,If you don't mind could you explain breifly why Spark doesn't fit in this competition?</p>",
      "rawMarkdown": "Hi Mike,If you don't mind could you explain breifly why Spark doesn't fit in this competition?",
      "votes": null
    },
    {
      "id": "120161",
      "postDate": "05/15/2016 22:37:51",
      "content": "<p>@Subhajit, thanks , I will take a look at this Spark solution.</p>",
      "rawMarkdown": "Subhajit, thanks , I will take a look at this Spark solution.",
      "votes": null
    },
    {
      "id": "120168",
      "postDate": "05/16/2016 00:52:45",
      "content": "<p>@Mike it would indeed be helpful if you could clarify how you mean.</p>",
      "rawMarkdown": "Mike it would indeed be helpful if you could clarify how you mean.",
      "votes": null
    },
    {
      "id": "120171",
      "postDate": "05/16/2016 01:35:26",
      "content": "<p>Spark has never worked for any Kaggle competitions to produce anything within the top 10% let alone a top 10 finish. You can forum search to see what kind of results Spark has produced in the past. Basically nothing of note except maybe a top 50% finish (maybe) or so using a cluster of hundred of fat machines costing hundreds of dollars on an AWS bill. </p>\n\n<p>Theoretically the source code of Spark shows it was not meant for machine learning on large datasets unlike say Vowpal, H2O, or XG. <a href=\"https://github.com/szilard/benchm-ml\">https://github.com/szilard/benchm-ml</a> Typically online models are used on Kaggle when datasets are large. XG is the one exception that seems to work regardless on a wide variety of problems.</p>\n\n<p>The philosophy behind Spark is probably more along the lines of keep operations up when nodes fail and easy to maintain code base versus pure, raw competitive machine learning performance. </p>",
      "rawMarkdown": "Spark has never worked for any Kaggle competitions to produce anything within the top 10% let alone a top 10 finish. You can forum search to see what kind of results Spark has produced in the past. Basically nothing of note except maybe a top 50% finish (maybe) or so using a cluster of hundred of fat machines costing hundreds of dollars on an AWS bill. \r\n\r\nTheoretically the source code of Spark shows it was not meant for machine learning on large datasets unlike say Vowpal, H2O, or XG. https://github.com/szilard/benchm-ml Typically online models are used on Kaggle when datasets are large. XG is the one exception that seems to work regardless on a wide variety of problems.\r\n\r\nThe philosophy behind Spark is probably more along the lines of keep operations up when nodes fail and easy to maintain code base versus pure, raw competitive machine learning performance.",
      "votes": null
    },
    {
      "id": "120174",
      "postDate": "05/16/2016 02:10:52",
      "content": "<p>Thanks a lot Mike for your reply.I understand clearly now.</p>\n\n<p>But as there are lot of enhancements going on spark,i hope spark will give competition to popular ML-libraries </p>\n\n<p>Thanks\nSiddhu</p>",
      "rawMarkdown": "Thanks a lot Mike for your reply.I understand clearly now.\r\n\r\nBut as there are lot of enhancements going on spark,i hope spark will give competition to popular ML-libraries \r\n\r\nThanks\r\nSiddhu",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115590,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/19/2016 08:29:57",
      "content": "<p>Hi,</p>\n\n<p>I think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.</p>\n\n<p>Yesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two &quot;normal&quot; matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.</p>\n\n<p>And then we have not processed/used the &quot;click&quot; data.</p>\n\n<p>So my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.</p>\n\n<p>Another idea would be to use H2O with a local &quot;cluster&quot;! I think I will look into H2O for this one again.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115595,
      "author_name": "",
      "author_url": "",
      "post_date": "04/19/2016 08:58:58",
      "content": "<p>[quote=MightyBird;115590]</p>\n\n<p>Hi,</p>\n\n<p>I think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.</p>\n\n<p>Yesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two &quot;normal&quot; matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.</p>\n\n<p>And then we have not processed/used the &quot;click&quot; data.</p>\n\n<p>So my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.</p>\n\n<p>Another idea would be to use H2O with a local &quot;cluster&quot;! I think I will look into H2O for this one again.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>\n\n<p>[/quote]</p>\n\n<p>I will start by working on a subset, while waiting for suggestions from the community. I think I have spent far too much time already.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115639,
      "author_name": "wrudebusch",
      "author_url": "",
      "post_date": "04/19/2016 13:01:39",
      "content": "<p>If you go to &quot;New Script&quot; in the Kaggle dash for this competition and enter the standard first line</p>\n\n<pre><code>train = pd.read_csv(&quot;../input/train.csv&quot;)\n</code></pre>\n\n<p>It will time out, &quot;Error: The script was killed, likely for trying to exceed the memory limit of 8192M.&quot;</p>\n\n<p>Soooooo is the Kaggle contest submission process different? I hope so. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115640,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "04/19/2016 13:12:34",
      "content": "<p>Some people use a database to vectorize train and test samples. It is unlikely that a processed dataset is the same size as the raw data.</p>\n\n<p>Subsampling only bookings, like suggested by MightyBird, makes a lot of sense.</p>\n\n<p>Finally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115652,
      "author_name": "ashishtrehan",
      "author_url": "",
      "post_date": "04/19/2016 14:07:51",
      "content": "<p>I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\ndf = dd.read_csv('*-*train.csv')\ndf.groupby(df.value.mean().compute()</p>\n\n<p>Check it out <a href=\"http://dask.pydata.org/en/latest/\">http://dask.pydata.org/en/latest/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115659,
      "author_name": "bguberfain",
      "author_url": "",
      "post_date": "04/19/2016 14:41:48",
      "content": "<p>One option is to process the dataset in chunks:</p>\n\n<ul>\n<li>For pandas, use <code>pd.read_csv('train.csv', chunksize=chunksize)</code>;</li>\n<li>For XGBoost use <code>model = xgb.train(..., xgb_model=model)</code>, assuming the returned model is used on the next train rounds.</li>\n</ul>\n\n<p>I'm not sure how it influences the model performance (as the train occurs in only a subset of the whole daatset, thus forcing a <code>subsample</code> parameter), but the memory footprint is reduced.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115668,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/19/2016 15:10:17",
      "content": "<p>So, another experiment just took place...</p>\n\n<pre><code>2.5 M rows test\n3.0 M rows train\n98 columns each\n</code></pre>\n\n<p>32 GB of RAM was not enough for xgboost in R. Sad :-(</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115669,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "04/19/2016 15:11:34",
      "content": "<p>What data types are you using? \nHow many cores?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115677,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/19/2016 15:35:00",
      "content": "<p>I was starting to think of using a different data representation right now. Used a normal matrix so far. 8 cores...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115682,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/19/2016 15:53:49",
      "content": "<p>A sparse matrix didn't help. It's the algorithm that's greedy. Will now downsample the train set. Starting with 50%.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115690,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/19/2016 16:20:25",
      "content": "<p>So, with 50% of train data I got it working. Uses ~20GB now. Algorithm is running - awfully slow although...</p>\n\n<p>It seems there is a problem. The xgboost results get worse with every round. Don't think I have seen this before. Maybe eta 0.5 is too high?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115693,
      "author_name": "wrudebusch",
      "author_url": "",
      "post_date": "04/19/2016 16:34:55",
      "content": "<p>[quote=Triskelion;115640]</p>\n\n<p>Finally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\n[/quote]</p>\n\n<p>Are you referring to AWS or some other cloud computing services?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115694,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "04/19/2016 16:48:03",
      "content": "<p>[quote=William Rudebusch;115693]</p>\n\n<p>[quote=Triskelion;115640]</p>\n\n<p>Finally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\n[/quote]</p>\n\n<p>Are you referring to AWS or some other cloud computing services?</p>\n\n<p>[/quote]</p>\n\n<p><a href=\"https://en.wikipedia.org/wiki/Online_machine_learning\">https://en.wikipedia.org/wiki/Online_machine_learning</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115792,
      "author_name": "",
      "author_url": "",
      "post_date": "04/20/2016 05:33:57",
      "content": "<p>[quote=ashish trehan;115652]</p>\n\n<p>I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\ndf = dd.read_csv('*-*train.csv')\ndf.groupby(df.value.mean().compute()</p>\n\n<p>Check it out <a href=\"http://dask.pydata.org/en/latest/\">http://dask.pydata.org/en/latest/</a></p>\n\n<p>[/quote]</p>\n\n<p>After loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115830,
      "author_name": "millerintllc",
      "author_url": "",
      "post_date": "04/20/2016 11:55:14",
      "content": "<p>[quote=MightyBird;115590]</p>\n\n<p>Another idea would be to use H2O with a local &quot;cluster&quot;! I think I will look into H2O for this one again.</p>\n\n<p>[/quote]</p>\n\n<p>I tried that and even with a sizable amount of resources, this is still pretty huge data. H2O  could certainly do whatever you want it to with this size of data, but if you use the entire dataset and models with lots of trees then the competition will still be over by the time it finishes lol.</p>\n\n<p>So, I guess subsetting is the way to go. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115832,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/20/2016 12:06:19",
      "content": "<p>Hi.</p>\n\n<p>I did a first xgboost model using only train data where is_booking == 1. I get stuck in the area of MAP@5 = 0.2</p>\n\n<p>I therefore think that we must also look at click data!</p>\n\n<p>My next step will thus be to random sample the whole train set.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115842,
      "author_name": "ashishtrehan",
      "author_url": "",
      "post_date": "04/20/2016 12:49:12",
      "content": "<p>[quote=Brox;115792]</p>\n\n<p>[quote=ashish trehan;115652]</p>\n\n<p>I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\ndf = dd.read_csv('*-*train.csv')\ndf.groupby(df.value.mean().compute()</p>\n\n<p>Check it out <a href=\"http://dask.pydata.org/en/latest/\">http://dask.pydata.org/en/latest/</a></p>\n\n<p>[/quote]</p>\n\n<p>After loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115844,
      "author_name": "ashishtrehan",
      "author_url": "",
      "post_date": "04/20/2016 12:50:57",
      "content": "<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115967,
      "author_name": "",
      "author_url": "",
      "post_date": "04/21/2016 03:14:13",
      "content": "<p>[quote=ashish trehan;115844]</p>\n\n<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, I loaded the entire training set.</p>\n\n<p>is loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115985,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "04/21/2016 07:20:21",
      "content": "<p>Coming back to number of cores, reduce it to half and see how it goes. As far as I remember for xgboost more cores =&gt; more memory needed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115986,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/21/2016 08:09:14",
      "content": "<p>I will try that. Would be good - on the memory side. Not so good on the runtime side though.</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116059,
      "author_name": "ashishtrehan",
      "author_url": "",
      "post_date": "04/21/2016 18:47:57",
      "content": "<p>[quote=Brox;115967]</p>\n\n<p>[quote=ashish trehan;115844]</p>\n\n<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, I loaded the entire training set.</p>\n\n<p>is loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)</p>\n\n<p>[/quote]</p>\n\n<p>Do you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116112,
      "author_name": "",
      "author_url": "",
      "post_date": "04/22/2016 04:17:59",
      "content": "<p>[quote=ashish trehan;116059]</p>\n\n<p>[quote=Brox;115967]</p>\n\n<p>[quote=ashish trehan;115844]</p>\n\n<p>Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, I loaded the entire training set.</p>\n\n<p>is loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)</p>\n\n<p>[/quote]</p>\n\n<p>Do you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.  </p>\n\n<p>[/quote]</p>\n\n<p>That;s what I am doing now, either reading only few columns at a time or reading only few columns at a time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116121,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/22/2016 05:39:29",
      "content": "<p>Short update. Using only 1% of train and after feature engineering about 180 columns my cv in xgboost uses about 6GB on a two core machine. Still running...  I think this might well take 24 hours.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117087,
      "author_name": "qqgeogor",
      "author_url": "",
      "post_date": "04/27/2016 10:55:35",
      "content": "<p>For xgboost I created dense feature set by subsampling training set using only is_booking =1 and some feature engineering techs and I could keep total memory usage with in 8GB.  I also did feature reduction on destination features.\nFor linear models such as logistic regression, sgd classifier or mlp, I onehot encoded the categorical features and generated a sparse matrix with dimension  size about 1200000. It is really time consuming to train a lr model or using mini batch sgd to train sgd classifier or mlp.\nNot yet tried any bagging things. It might work or might not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120077,
      "author_name": "dmatekenya",
      "author_url": "",
      "post_date": "05/15/2016 06:02:10",
      "content": "<p>I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120084,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "05/15/2016 08:12:11",
      "content": "<p>[quote=dun;120077]</p>\n\n<p>I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?</p>\n\n<p>[/quote]</p>\n\n<p>Here you go: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120104,
      "author_name": "masterliu",
      "author_url": "",
      "post_date": "05/15/2016 12:52:19",
      "content": "<p>has anyone tried online learning? does it work?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120105,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/15/2016 12:54:19",
      "content": "<p>I have, via Vowpal. Nothing to write home about in terms of results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120127,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "05/15/2016 18:13:00",
      "content": "<p>[quote=dun;120077]</p>\n\n<p>I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?</p>\n\n<p>[/quote]\nBecause it doesn't work (in theory and in practice).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120132,
      "author_name": "siddhu25",
      "author_url": "",
      "post_date": "05/15/2016 19:02:32",
      "content": "<p>Hi Mike,If you don't mind could you explain breifly why Spark doesn't fit in this competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120161,
      "author_name": "dmatekenya",
      "author_url": "",
      "post_date": "05/15/2016 22:37:51",
      "content": "<p>@Subhajit, thanks , I will take a look at this Spark solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120168,
      "author_name": "dmatekenya",
      "author_url": "",
      "post_date": "05/16/2016 00:52:45",
      "content": "<p>@Mike it would indeed be helpful if you could clarify how you mean.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120171,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "05/16/2016 01:35:26",
      "content": "<p>Spark has never worked for any Kaggle competitions to produce anything within the top 10% let alone a top 10 finish. You can forum search to see what kind of results Spark has produced in the past. Basically nothing of note except maybe a top 50% finish (maybe) or so using a cluster of hundred of fat machines costing hundreds of dollars on an AWS bill. </p>\n\n<p>Theoretically the source code of Spark shows it was not meant for machine learning on large datasets unlike say Vowpal, H2O, or XG. <a href=\"https://github.com/szilard/benchm-ml\">https://github.com/szilard/benchm-ml</a> Typically online models are used on Kaggle when datasets are large. XG is the one exception that seems to work regardless on a wide variety of problems.</p>\n\n<p>The philosophy behind Spark is probably more along the lines of keep operations up when nodes fail and easy to maintain code base versus pure, raw competitive machine learning performance. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120174,
      "author_name": "siddhu25",
      "author_url": "",
      "post_date": "05/16/2016 02:10:52",
      "content": "<p>Thanks a lot Mike for your reply.I understand clearly now.</p>\n\n<p>But as there are lot of enhancements going on spark,i hope spark will give competition to popular ML-libraries </p>\n\n<p>Thanks\nSiddhu</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115574": "I am trying to do this in R or Python. I am interested to know how the kaggle community is handling big data sets. Below are few thoughts I have\r\n\r\n 1. Take a subset of the database, build a model and then run the model\r\n    on the entire database using kaggle scripts\r\n 2. Use packaged to handle big data sets such as \"**ff**\" for **R** or \"**Blaze**\" for **python** and do everything on your own system. \r\n\r\nI want do everything on my own system but so far I have found that these so called big data packages are not very helpful. Some are taking too long and some have memory errors or some other errors. I have been stuck in this phase for lmost a day. I haven't worked much on big data sets. I have't used kaggle scripts much either.\r\n\r\nWhat would be best way to work on this competition?",
    "115590": "Hi,\r\n\r\nI think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.\r\n\r\nYesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two \"normal\" matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.\r\n\r\nAnd then we have not processed/used the \"click\" data.\r\n\r\nSo my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.\r\n\r\nAnother idea would be to use H2O with a local \"cluster\"! I think I will look into H2O for this one again.\r\n\r\nCheers\r\n\r\nGerhard",
    "115595": "[quote=MightyBird;115590]\r\n\r\nHi,\r\n\r\nI think there is a memory limitation of 8GB for kaggle scripts? And a runtime limintaion as well! So I think for good results this will not be enough.\r\n\r\nYesterday I tried the following - which might give you an idea of how your system needs to look like. I used only train data where is_booking = 1 and all test data, then destinations. Then I loaded this in R and did some feature engineering and some one hot encoding. Then I created two \"normal\" matrices for use in xgboost. And that's where I hit the top of my memory (8GB +4GB swap). So there is no memory left for the algorithm :-( Which will be considerable.\r\n\r\nAnd then we have not processed/used the \"click\" data.\r\n\r\nSo my guess is I will need at least 16GB, better 32GB. I intend to run this on a cloud computing instance.\r\n\r\nAnother idea would be to use H2O with a local \"cluster\"! I think I will look into H2O for this one again.\r\n\r\nCheers\r\n\r\nGerhard\r\n\r\n[/quote]\r\n\r\nI will start by working on a subset, while waiting for suggestions from the community. I think I have spent far too much time already.",
    "115639": "If you go to \"New Script\" in the Kaggle dash for this competition and enter the standard first line\r\n\r\n    train = pd.read_csv(\"../input/train.csv\")\r\n\r\nIt will time out, \"Error: The script was killed, likely for trying to exceed the memory limit of 8192M.\"\r\n\r\nSoooooo is the Kaggle contest submission process different? I hope so.",
    "115640": "Some people use a database to vectorize train and test samples. It is unlikely that a processed dataset is the same size as the raw data.\r\n\r\nSubsampling only bookings, like suggested by MightyBird, makes a lot of sense.\r\n\r\nFinally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.",
    "115652": "I recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\r\ndf = dd.read_csv('*-*train.csv')\r\ndf.groupby(df.value.mean().compute()\r\n\r\nCheck it out http://dask.pydata.org/en/latest/",
    "115659": "One option is to process the dataset in chunks:\r\n\r\n- For pandas, use `pd.read_csv('train.csv', chunksize=chunksize)`;\r\n- For XGBoost use `model = xgb.train(..., xgb_model=model)`, assuming the returned model is used on the next train rounds.\r\n\r\nI'm not sure how it influences the model performance (as the train occurs in only a subset of the whole daatset, thus forcing a `subsample` parameter), but the memory footprint is reduced.",
    "115668": "So, another experiment just took place...\r\n\r\n    2.5 M rows test\r\n    3.0 M rows train\r\n    98 columns each\r\n\r\n32 GB of RAM was not enough for xgboost in R. Sad :-(\r\n\r\nCheers\r\n\r\nGerhard",
    "115669": "What data types are you using? \r\nHow many cores?",
    "115677": "I was starting to think of using a different data representation right now. Used a normal matrix so far. 8 cores...",
    "115682": "A sparse matrix didn't help. It's the algorithm that's greedy. Will now downsample the train set. Starting with 50%.",
    "115690": "So, with 50% of train data I got it working. Uses ~20GB now. Algorithm is running - awfully slow although...\r\n\r\nIt seems there is a problem. The xgboost results get worse with every round. Don't think I have seen this before. Maybe eta 0.5 is too high?",
    "115693": "[quote=Triskelion;115640]\r\n\r\nFinally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\r\n[/quote]\r\n\r\nAre you referring to AWS or some other cloud computing services?",
    "115694": "[quote=William Rudebusch;115693]\r\n\r\n[quote=Triskelion;115640]\r\n\r\nFinally, have a look at online learning. You probably won't beat in-memory techniques, but you may learn how to do extremely big data problems with just a few MB of memory.\r\n[/quote]\r\n\r\nAre you referring to AWS or some other cloud computing services?\r\n\r\n[/quote]\r\n\r\nhttps://en.wikipedia.org/wiki/Online_machine_learning",
    "115792": "[quote=ashish trehan;115652]\r\n\r\nI recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\r\ndf = dd.read_csv('*-*train.csv')\r\ndf.groupby(df.value.mean().compute()\r\n\r\nCheck it out http://dask.pydata.org/en/latest/\r\n\r\n[/quote]\r\n\r\nAfter loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?",
    "115830": "[quote=MightyBird;115590]\r\n\r\n\r\nAnother idea would be to use H2O with a local \"cluster\"! I think I will look into H2O for this one again.\r\n\r\n\r\n[/quote]\r\n\r\nI tried that and even with a sizable amount of resources, this is still pretty huge data. H2O  could certainly do whatever you want it to with this size of data, but if you use the entire dataset and models with lots of trees then the competition will still be over by the time it finishes lol.\r\n\r\nSo, I guess subsetting is the way to go.",
    "115832": "Hi.\r\n\r\nI did a first xgboost model using only train data where is_booking == 1. I get stuck in the area of MAP@5 = 0.2\r\n\r\nI therefore think that we must also look at click data!\r\n\r\nMy next step will thus be to random sample the whole train set.\r\n\r\nGerhard",
    "115842": "[quote=Brox;115792]\r\n\r\n[quote=ashish trehan;115652]\r\n\r\nI recommend using Dask package to help work with this data. I deployed a dockerized ipython notebook server on Digital Ocean to move this off my own laptop. But Dask has proven to at least be able to work with this data. It has similar syntax to pandas here are some examples.\r\ndf = dd.read_csv('*-*train.csv')\r\ndf.groupby(df.value.mean().compute()\r\n\r\nCheck it out http://dask.pydata.org/en/latest/\r\n\r\n[/quote]\r\n\r\nAfter loading the dataset using dask.dataframe, my system can hardly show me the result of df.head(). it's taking so long. am I doing something wrong?\r\n\r\n[/quote]",
    "115844": "Can you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.",
    "115967": "[quote=ashish trehan;115844]\r\n\r\nCan you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.\r\n\r\n[/quote]\r\n\r\nYes, I loaded the entire training set.\r\n\r\nis loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)",
    "115985": "Coming back to number of cores, reduce it to half and see how it goes. As far as I remember for xgboost more cores => more memory needed.",
    "115986": "I will try that. Would be good - on the memory side. Not so good on the runtime side though.\r\n\r\nThanks\r\n\r\nGerhard",
    "116059": "[quote=Brox;115967]\r\n\r\n[quote=ashish trehan;115844]\r\n\r\nCan you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.\r\n\r\n[/quote]\r\n\r\nYes, I loaded the entire training set.\r\n\r\nis loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)\r\n\r\n[/quote]\r\n\r\n\r\nDo you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.",
    "116112": "[quote=ashish trehan;116059]\r\n\r\n[quote=Brox;115967]\r\n\r\n[quote=ashish trehan;115844]\r\n\r\nCan you share the code? So you loaded the whole file, while I would still recommend still passing only columns you are interested in.\r\n\r\n[/quote]\r\n\r\nYes, I loaded the entire training set.\r\n\r\nis loading few columns the only solution? we are likely to engineer more features in addition to given variables and run a classifier on them. no of columns are only going to increase (at least for the first iteration)\r\n\r\n[/quote]\r\n\r\n\r\nDo you require all the observations? Maybe you can use dd.sample(frac=0.1) to obtain a random sample of your data which hopefully gives you enough to build a model off of.  \r\n\r\n\r\n[/quote]\r\n\r\nThat;s what I am doing now, either reading only few columns at a time or reading only few columns at a time.",
    "116121": "Short update. Using only 1% of train and after feature engineering about 180 columns my cv in xgboost uses about 6GB on a two core machine. Still running...  I think this might well take 24 hours.",
    "117087": "For xgboost I created dense feature set by subsampling training set using only is_booking =1 and some feature engineering techs and I could keep total memory usage with in 8GB.  I also did feature reduction on destination features.\r\nFor linear models such as logistic regression, sgd classifier or mlp, I onehot encoded the categorical features and generated a sparse matrix with dimension  size about 1200000. It is really time consuming to train a lr model or using mini batch sgd to train sgd classifier or mlp.\r\nNot yet tried any bagging things. It might work or might not.",
    "120077": "I'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?",
    "120084": "[quote=dun;120077]\r\n\r\nI'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?\r\n\r\n[/quote]\r\n\r\nHere you go: https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20896/leakage-solution-with-spark-sql-pyspark",
    "120104": "has anyone tried online learning? does it work?",
    "120105": "I have, via Vowpal. Nothing to write home about in terms of results.",
    "120127": "[quote=dun;120077]\r\n\r\nI'm wondering why nobody has recommended Apache Spark as a solution for handling the large data. Is it a bad idea to use software like Spark for Kaggle competitions?\r\n\r\n[/quote]\r\nBecause it doesn't work (in theory and in practice).",
    "120132": "Hi Mike,If you don't mind could you explain breifly why Spark doesn't fit in this competition?",
    "120161": "Subhajit, thanks , I will take a look at this Spark solution.",
    "120168": "Mike it would indeed be helpful if you could clarify how you mean.",
    "120171": "Spark has never worked for any Kaggle competitions to produce anything within the top 10% let alone a top 10 finish. You can forum search to see what kind of results Spark has produced in the past. Basically nothing of note except maybe a top 50% finish (maybe) or so using a cluster of hundred of fat machines costing hundreds of dollars on an AWS bill. \r\n\r\nTheoretically the source code of Spark shows it was not meant for machine learning on large datasets unlike say Vowpal, H2O, or XG. https://github.com/szilard/benchm-ml Typically online models are used on Kaggle when datasets are large. XG is the one exception that seems to work regardless on a wide variety of problems.\r\n\r\nThe philosophy behind Spark is probably more along the lines of keep operations up when nodes fail and easy to maintain code base versus pure, raw competitive machine learning performance.",
    "120174": "Thanks a lot Mike for your reply.I understand clearly now.\r\n\r\nBut as there are lot of enhancements going on spark,i hope spark will give competition to popular ML-libraries \r\n\r\nThanks\r\nSiddhu"
  },
  "source": "meta"
}