{
  "id": 20557,
  "title": "How much we can do one single local computer?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20557",
  "author_name": "",
  "post_date": "2016-04-29T23:33:33.027Z",
  "votes": null,
  "comment_count": 16,
  "views": 1781,
  "content": "<p>Hello, Kagglers,</p>\n\n<p>It will be a nice experience to learn and discuss with kagglers in this competition. One character of this competition is the huge data size. It is my first time to play with data with this size (37,670,293 in train and 2,528,243 in test). I have a Mac with 24G RAM and 2.9GHz quad-core Intel Core i5 processor.</p>\n\n<p>I am curious on how much I can go with this single computer. Any people in top 40 using a single computer by far? Any suggestions are appreciated.</p>\n\n<p>Thanks,</p>",
  "messages": [
    {
      "id": "117640",
      "postDate": "04/29/2016 23:33:33",
      "content": "<p>Hello, Kagglers,</p>\n\n<p>It will be a nice experience to learn and discuss with kagglers in this competition. One character of this competition is the huge data size. It is my first time to play with data with this size (37,670,293 in train and 2,528,243 in test). I have a Mac with 24G RAM and 2.9GHz quad-core Intel Core i5 processor.</p>\n\n<p>I am curious on how much I can go with this single computer. Any people in top 40 using a single computer by far? Any suggestions are appreciated.</p>\n\n<p>Thanks,</p>",
      "rawMarkdown": "Hello, Kagglers,\r\n\r\nIt will be a nice experience to learn and discuss with kagglers in this competition. One character of this competition is the huge data size. It is my first time to play with data with this size (37,670,293 in train and 2,528,243 in test). I have a Mac with 24G RAM and 2.9GHz quad-core Intel Core i5 processor.\r\n\r\nI am curious on how much I can go with this single computer. Any people in top 40 using a single computer by far? Any suggestions are appreciated.\r\n\r\nThanks,",
      "votes": null
    },
    {
      "id": "117670",
      "postDate": "04/30/2016 05:06:47",
      "content": "<p>I have a fairly simple Laptop with 4GB and just using perl I got 0.443. Memory consumption was a max of 2GB I think. Runtime about 10-20 mins. The fun fact here is that I used all rows, though not yet all features that can be used / created.</p>\n\n<p>I also tried vowpal wabbit with horrible disk usage, about 50GB, a fair runtime of about 2 hours and a bad result of 0.02. Yes, couldn't believe it either. Maybe I did something wrong with data preparation. However, this needs no RAM either.</p>\n\n<p>I also tried xgboost but nearly at once run into memory problems, even on larger machines (32GB) and using only 1% of train. There is another thread about this topic!</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "I have a fairly simple Laptop with 4GB and just using perl I got 0.443. Memory consumption was a max of 2GB I think. Runtime about 10-20 mins. The fun fact here is that I used all rows, though not yet all features that can be used / created.\r\n\r\nI also tried vowpal wabbit with horrible disk usage, about 50GB, a fair runtime of about 2 hours and a bad result of 0.02. Yes, couldn't believe it either. Maybe I did something wrong with data preparation. However, this needs no RAM either.\r\n\r\nI also tried xgboost but nearly at once run into memory problems, even on larger machines (32GB) and using only 1% of train. There is another thread about this topic!\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117677",
      "postDate": "04/30/2016 06:27:45",
      "content": "<p>@MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?</p>",
      "rawMarkdown": "MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?",
      "votes": null
    },
    {
      "id": "117680",
      "postDate": "04/30/2016 06:55:22",
      "content": "<p>It's not a real model in terms of machine learning. I just count and average data and of course I use the information from the data leak. The trick is to read the files line by line and only memorize &quot;important&quot; things.</p>\n\n<p>Right now I again try to bring xgboost to good use. Will try to use portions of train (only booked) at a time, then average.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "It's not a real model in terms of machine learning. I just count and average data and of course I use the information from the data leak. The trick is to read the files line by line and only memorize \"important\" things.\r\n\r\nRight now I again try to bring xgboost to good use. Will try to use portions of train (only booked) at a time, then average.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117704",
      "postDate": "04/30/2016 12:58:16",
      "content": "<p>What are you guys doing to get good scores on the leaderboard? I'm confused about how to use the data leak and I can't figure out where to start...</p>",
      "rawMarkdown": "What are you guys doing to get good scores on the leaderboard? I'm confused about how to use the data leak and I can't figure out where to start...",
      "votes": null
    },
    {
      "id": "117715",
      "postDate": "04/30/2016 14:15:48",
      "content": "<p>The data leak is explained exactly in the data lead thread. You just need to do what's written there. And for the &quot;extra points&quot; I personally didn't do anything that was not mentioned in the forum or in the public scripts. Maybe I amended one thing or the other a bit. No magic involved!</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "The data leak is explained exactly in the data lead thread. You just need to do what's written there. And for the \"extra points\" I personally didn't do anything that was not mentioned in the forum or in the public scripts. Maybe I amended one thing or the other a bit. No magic involved!\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117750",
      "postDate": "04/30/2016 17:22:45",
      "content": "<p>It's a good time to get started with this competition. The leak has already been explained and now everybody is basically at the same starting point. </p>\n\n<p>As far as I know, nobody has done any meaningful machine learning with this competition yet. Not only is memory a serious constraint, the training time is also incredibly long because it's in the simplest form a 100-category classification problem. Even with only 1/30th of the observations and only two variables it takes me longer to train than I would like. Therefore, I think the most important step in this competition will be to figure out a way to sufficiently run machine learning algorithms. Otherwise it would just be perform decision trees by hand :) </p>",
      "rawMarkdown": "It's a good time to get started with this competition. The leak has already been explained and now everybody is basically at the same starting point. \r\n\r\nAs far as I know, nobody has done any meaningful machine learning with this competition yet. Not only is memory a serious constraint, the training time is also incredibly long because it's in the simplest form a 100-category classification problem. Even with only 1/30th of the observations and only two variables it takes me longer to train than I would like. Therefore, I think the most important step in this competition will be to figure out a way to sufficiently run machine learning algorithms. Otherwise it would just be perform decision trees by hand :)",
      "votes": null
    },
    {
      "id": "117769",
      "postDate": "04/30/2016 20:30:23",
      "content": "<p>Hi,</p>\n\n<p>just finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...</p>\n\n<p>Ran on 10 % of is_booking == 1 train data and obviously all of test. </p>\n\n<p>This is also a stretch for memory!</p>\n\n<p>Just so you know.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\njust finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...\r\n\r\nRan on 10 % of is_booking == 1 train data and obviously all of test. \r\n\r\nThis is also a stretch for memory!\r\n\r\nJust so you know.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117780",
      "postDate": "04/30/2016 21:29:28",
      "content": "<p>[quote=MightyBird;117769]</p>\n\n<p>Hi,</p>\n\n<p>just finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...</p>\n\n<p>Ran on 10 % of is_booking == 1 train data and obviously all of test. </p>\n\n<p>This is also a stretch for memory!</p>\n\n<p>Just so you know.</p>\n\n<p>Gerhard</p>\n\n<p>[/quote]</p>\n\n<p>What eval_metric are you using for xgboost?</p>",
      "rawMarkdown": "[quote=MightyBird;117769]\r\n\r\nHi,\r\n\r\njust finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...\r\n\r\nRan on 10 % of is_booking == 1 train data and obviously all of test. \r\n\r\nThis is also a stretch for memory!\r\n\r\nJust so you know.\r\n\r\nGerhard\r\n\r\n\r\n[/quote]\r\n\r\nWhat eval_metric are you using for xgboost?",
      "votes": null
    },
    {
      "id": "117811",
      "postDate": "05/01/2016 02:44:54",
      "content": "<p>@MightyBird\nThanks for your information. I run a xgboost model (no one hot encoding on features yet) on all the is_booking == 1 data, the local cv score ~ 0.29. Obviously, click data is very important.</p>",
      "rawMarkdown": "MightyBird\r\nThanks for your information. I run a xgboost model (no one hot encoding on features yet) on all the is_booking == 1 data, the local cv score ~ 0.29. Obviously, click data is very important.",
      "votes": null
    },
    {
      "id": "117813",
      "postDate": "05/01/2016 02:55:46",
      "content": "<p>@Qingchen\nThank you for sharing and congratulations on your high rank. :-)\nThis competition is really interesting due to the data size.</p>",
      "rawMarkdown": "Qingchen\r\nThank you for sharing and congratulations on your high rank. :-)\r\nThis competition is really interesting due to the data size.",
      "votes": null
    },
    {
      "id": "117827",
      "postDate": "05/01/2016 04:51:58",
      "content": "<p>I used logloss and got around 3.2 in cv.</p>\n\n<p>@Li Li did you use rank:pairwise? Me not.</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "I used logloss and got around 3.2 in cv.\r\n\r\n@Li Li did you use rank:pairwise? Me not.\r\n\r\nThanks\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "117839",
      "postDate": "05/01/2016 06:32:21",
      "content": "<p>@MightyBird</p>\n\n<p>I also used mlogloss  and got around 3.2 in CV.</p>\n\n<p>The MAP5 is calculated for the output, not in training process. </p>",
      "rawMarkdown": "MightyBird\r\n\r\nI also used mlogloss  and got around 3.2 in CV.\r\n\r\nThe MAP5 is calculated for the output, not in training process.",
      "votes": null
    },
    {
      "id": "117851",
      "postDate": "05/01/2016 07:33:08",
      "content": "<p>I am also struggling with the data size. So far the best I've got is from xgboost on 2.5% of the data with 300 trees for 11.5 hours on 8GB RAM, 4 threads. I got 0.28+. I am currently working on sgd classifier. Would any one of you be interested to team up? May be we can crack this together with a bigger RAM and better features.</p>",
      "rawMarkdown": "I am also struggling with the data size. So far the best I've got is from xgboost on 2.5% of the data with 300 trees for 11.5 hours on 8GB RAM, 4 threads. I got 0.28+. I am currently working on sgd classifier. Would any one of you be interested to team up? May be we can crack this together with a bigger RAM and better features.",
      "votes": null
    },
    {
      "id": "117858",
      "postDate": "05/01/2016 08:38:59",
      "content": "<p>@Subhajit Mandal</p>\n\n<p>Understand. That is painful but fun.</p>",
      "rawMarkdown": "Subhajit Mandal\r\n\r\nUnderstand. That is painful but fun.",
      "votes": null
    },
    {
      "id": "118297",
      "postDate": "05/03/2016 07:45:40",
      "content": "<p>[quote=Li Li;117677]</p>\n\n<p>@MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?</p>\n\n<p>[/quote]</p>\n\n<p>According to MightyBird comment, he wore that he could use only 1% of train data, not all data. Is there any one who could read all rows?</p>",
      "rawMarkdown": "[quote=Li Li;117677]\r\n\r\n@MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?\r\n\r\n[/quote]\r\n\r\nAccording to MightyBird comment, he wore that he could use only 1% of train data, not all data. Is there any one who could read all rows?",
      "votes": null
    },
    {
      "id": "118304",
      "postDate": "05/03/2016 08:00:45",
      "content": "<p>Hi,</p>\n\n<p>I think there are things mixed up here a bit. With xgboost I could only use a small part of train and got &quot;bad&quot; results. When I use perl I can make use of all 37.000.000 rows of train. And quite frankly, I think you need to look at all rows. Maybe not at once, but in the end.</p>\n\n<p>With the stated 2GB of RAM and a runtime of about 20 mins my perl script now scores 0.47x. I think I will give up on other models. I just don't have the compute power, time and patience.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nI think there are things mixed up here a bit. With xgboost I could only use a small part of train and got \"bad\" results. When I use perl I can make use of all 37.000.000 rows of train. And quite frankly, I think you need to look at all rows. Maybe not at once, but in the end.\r\n\r\nWith the stated 2GB of RAM and a runtime of about 20 mins my perl script now scores 0.47x. I think I will give up on other models. I just don't have the compute power, time and patience.\r\n\r\nGerhard",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117670,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/30/2016 05:06:47",
      "content": "<p>I have a fairly simple Laptop with 4GB and just using perl I got 0.443. Memory consumption was a max of 2GB I think. Runtime about 10-20 mins. The fun fact here is that I used all rows, though not yet all features that can be used / created.</p>\n\n<p>I also tried vowpal wabbit with horrible disk usage, about 50GB, a fair runtime of about 2 hours and a bad result of 0.02. Yes, couldn't believe it either. Maybe I did something wrong with data preparation. However, this needs no RAM either.</p>\n\n<p>I also tried xgboost but nearly at once run into memory problems, even on larger machines (32GB) and using only 1% of train. There is another thread about this topic!</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117677,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "04/30/2016 06:27:45",
      "content": "<p>@MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117680,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/30/2016 06:55:22",
      "content": "<p>It's not a real model in terms of machine learning. I just count and average data and of course I use the information from the data leak. The trick is to read the files line by line and only memorize &quot;important&quot; things.</p>\n\n<p>Right now I again try to bring xgboost to good use. Will try to use portions of train (only booked) at a time, then average.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117704,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "04/30/2016 12:58:16",
      "content": "<p>What are you guys doing to get good scores on the leaderboard? I'm confused about how to use the data leak and I can't figure out where to start...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117715,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/30/2016 14:15:48",
      "content": "<p>The data leak is explained exactly in the data lead thread. You just need to do what's written there. And for the &quot;extra points&quot; I personally didn't do anything that was not mentioned in the forum or in the public scripts. Maybe I amended one thing or the other a bit. No magic involved!</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117750,
      "author_name": "qwang88",
      "author_url": "",
      "post_date": "04/30/2016 17:22:45",
      "content": "<p>It's a good time to get started with this competition. The leak has already been explained and now everybody is basically at the same starting point. </p>\n\n<p>As far as I know, nobody has done any meaningful machine learning with this competition yet. Not only is memory a serious constraint, the training time is also incredibly long because it's in the simplest form a 100-category classification problem. Even with only 1/30th of the observations and only two variables it takes me longer to train than I would like. Therefore, I think the most important step in this competition will be to figure out a way to sufficiently run machine learning algorithms. Otherwise it would just be perform decision trees by hand :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117769,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "04/30/2016 20:30:23",
      "content": "<p>Hi,</p>\n\n<p>just finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...</p>\n\n<p>Ran on 10 % of is_booking == 1 train data and obviously all of test. </p>\n\n<p>This is also a stretch for memory!</p>\n\n<p>Just so you know.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117780,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "04/30/2016 21:29:28",
      "content": "<p>[quote=MightyBird;117769]</p>\n\n<p>Hi,</p>\n\n<p>just finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...</p>\n\n<p>Ran on 10 % of is_booking == 1 train data and obviously all of test. </p>\n\n<p>This is also a stretch for memory!</p>\n\n<p>Just so you know.</p>\n\n<p>Gerhard</p>\n\n<p>[/quote]</p>\n\n<p>What eval_metric are you using for xgboost?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117811,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/01/2016 02:44:54",
      "content": "<p>@MightyBird\nThanks for your information. I run a xgboost model (no one hot encoding on features yet) on all the is_booking == 1 data, the local cv score ~ 0.29. Obviously, click data is very important.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117813,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/01/2016 02:55:46",
      "content": "<p>@Qingchen\nThank you for sharing and congratulations on your high rank. :-)\nThis competition is really interesting due to the data size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117827,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/01/2016 04:51:58",
      "content": "<p>I used logloss and got around 3.2 in cv.</p>\n\n<p>@Li Li did you use rank:pairwise? Me not.</p>\n\n<p>Thanks</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117839,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/01/2016 06:32:21",
      "content": "<p>@MightyBird</p>\n\n<p>I also used mlogloss  and got around 3.2 in CV.</p>\n\n<p>The MAP5 is calculated for the output, not in training process. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117851,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "05/01/2016 07:33:08",
      "content": "<p>I am also struggling with the data size. So far the best I've got is from xgboost on 2.5% of the data with 300 trees for 11.5 hours on 8GB RAM, 4 threads. I got 0.28+. I am currently working on sgd classifier. Would any one of you be interested to team up? May be we can crack this together with a bigger RAM and better features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117858,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/01/2016 08:38:59",
      "content": "<p>@Subhajit Mandal</p>\n\n<p>Understand. That is painful but fun.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118297,
      "author_name": "bahramimaryam",
      "author_url": "",
      "post_date": "05/03/2016 07:45:40",
      "content": "<p>[quote=Li Li;117677]</p>\n\n<p>@MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?</p>\n\n<p>[/quote]</p>\n\n<p>According to MightyBird comment, he wore that he could use only 1% of train data, not all data. Is there any one who could read all rows?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118304,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/03/2016 08:00:45",
      "content": "<p>Hi,</p>\n\n<p>I think there are things mixed up here a bit. With xgboost I could only use a small part of train and got &quot;bad&quot; results. When I use perl I can make use of all 37.000.000 rows of train. And quite frankly, I think you need to look at all rows. Maybe not at once, but in the end.</p>\n\n<p>With the stated 2GB of RAM and a runtime of about 20 mins my perl script now scores 0.47x. I think I will give up on other models. I just don't have the compute power, time and patience.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117640": "Hello, Kagglers,\r\n\r\nIt will be a nice experience to learn and discuss with kagglers in this competition. One character of this competition is the huge data size. It is my first time to play with data with this size (37,670,293 in train and 2,528,243 in test). I have a Mac with 24G RAM and 2.9GHz quad-core Intel Core i5 processor.\r\n\r\nI am curious on how much I can go with this single computer. Any people in top 40 using a single computer by far? Any suggestions are appreciated.\r\n\r\nThanks,",
    "117670": "I have a fairly simple Laptop with 4GB and just using perl I got 0.443. Memory consumption was a max of 2GB I think. Runtime about 10-20 mins. The fun fact here is that I used all rows, though not yet all features that can be used / created.\r\n\r\nI also tried vowpal wabbit with horrible disk usage, about 50GB, a fair runtime of about 2 hours and a bad result of 0.02. Yes, couldn't believe it either. Maybe I did something wrong with data preparation. However, this needs no RAM either.\r\n\r\nI also tried xgboost but nearly at once run into memory problems, even on larger machines (32GB) and using only 1% of train. There is another thread about this topic!\r\n\r\nGerhard",
    "117677": "MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?",
    "117680": "It's not a real model in terms of machine learning. I just count and average data and of course I use the information from the data leak. The trick is to read the files line by line and only memorize \"important\" things.\r\n\r\nRight now I again try to bring xgboost to good use. Will try to use portions of train (only booked) at a time, then average.\r\n\r\nGerhard",
    "117704": "What are you guys doing to get good scores on the leaderboard? I'm confused about how to use the data leak and I can't figure out where to start...",
    "117715": "The data leak is explained exactly in the data lead thread. You just need to do what's written there. And for the \"extra points\" I personally didn't do anything that was not mentioned in the forum or in the public scripts. Maybe I amended one thing or the other a bit. No magic involved!\r\n\r\nGerhard",
    "117750": "It's a good time to get started with this competition. The leak has already been explained and now everybody is basically at the same starting point. \r\n\r\nAs far as I know, nobody has done any meaningful machine learning with this competition yet. Not only is memory a serious constraint, the training time is also incredibly long because it's in the simplest form a 100-category classification problem. Even with only 1/30th of the observations and only two variables it takes me longer to train than I would like. Therefore, I think the most important step in this competition will be to figure out a way to sufficiently run machine learning algorithms. Otherwise it would just be perform decision trees by hand :)",
    "117769": "Hi,\r\n\r\njust finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...\r\n\r\nRan on 10 % of is_booking == 1 train data and obviously all of test. \r\n\r\nThis is also a stretch for memory!\r\n\r\nJust so you know.\r\n\r\nGerhard",
    "117780": "[quote=MightyBird;117769]\r\n\r\nHi,\r\n\r\njust finished my xgboost model today. Took around 10 hrs on 8 core machine. Scored 0.24...\r\n\r\nRan on 10 % of is_booking == 1 train data and obviously all of test. \r\n\r\nThis is also a stretch for memory!\r\n\r\nJust so you know.\r\n\r\nGerhard\r\n\r\n\r\n[/quote]\r\n\r\nWhat eval_metric are you using for xgboost?",
    "117811": "MightyBird\r\nThanks for your information. I run a xgboost model (no one hot encoding on features yet) on all the is_booking == 1 data, the local cv score ~ 0.29. Obviously, click data is very important.",
    "117813": "Qingchen\r\nThank you for sharing and congratulations on your high rank. :-)\r\nThis competition is really interesting due to the data size.",
    "117827": "I used logloss and got around 3.2 in cv.\r\n\r\n@Li Li did you use rank:pairwise? Me not.\r\n\r\nThanks\r\n\r\nGerhard",
    "117839": "MightyBird\r\n\r\nI also used mlogloss  and got around 3.2 in CV.\r\n\r\nThe MAP5 is calculated for the output, not in training process.",
    "117851": "I am also struggling with the data size. So far the best I've got is from xgboost on 2.5% of the data with 300 trees for 11.5 hours on 8GB RAM, 4 threads. I got 0.28+. I am currently working on sgd classifier. Would any one of you be interested to team up? May be we can crack this together with a bigger RAM and better features.",
    "117858": "Subhajit Mandal\r\n\r\nUnderstand. That is painful but fun.",
    "118297": "[quote=Li Li;117677]\r\n\r\n@MightyBird. Congratulations on your score and thank you for reply. I am surprise that you used all rows and columns (am I correct?) with just 2GB RAM. What kind of model are you using?\r\n\r\n[/quote]\r\n\r\nAccording to MightyBird comment, he wore that he could use only 1% of train data, not all data. Is there any one who could read all rows?",
    "118304": "Hi,\r\n\r\nI think there are things mixed up here a bit. With xgboost I could only use a small part of train and got \"bad\" results. When I use perl I can make use of all 37.000.000 rows of train. And quite frankly, I think you need to look at all rows. Maybe not at once, but in the end.\r\n\r\nWith the stated 2GB of RAM and a runtime of about 20 mins my perl script now scores 0.47x. I think I will give up on other models. I just don't have the compute power, time and patience.\r\n\r\nGerhard"
  },
  "source": "meta"
}