{
  "id": 20330,
  "title": "Applying trained model on the huge test dataset",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20330",
  "author_name": "",
  "post_date": "2016-04-22T09:27:50.973Z",
  "votes": null,
  "comment_count": 8,
  "views": 1414,
  "content": "<p>Hello all, first of all I am new to kaggle compeditions, and i am using R for this compedition\nWe all are aware that the training and testing datasets are very huge to be read in. So I extracted only 10000 rows of the training dataset in the read.csv function with nrows=10000 and created a simple SVM model. But I cannot do the same with the test dataset because while submitting we got to submit the entire test dataset.\nWhen I try to  feed the test data(somewhere around 2500000+ records) to my SVM model it runs for sometime and ends up with cannot allocate memory of 93GB. I am using a machine intel i5 processor with 16GB of RAM. I know that R stores data only in RAM but I am not sure how to handle this situation. If any body could provide me a suggestion or how are you guys handling this situation, it would be really great and it will like a knowledge sharing for all of us. </p>",
  "messages": [
    {
      "id": "116159",
      "postDate": "04/22/2016 09:27:50",
      "content": "<p>Hello all, first of all I am new to kaggle compeditions, and i am using R for this compedition\nWe all are aware that the training and testing datasets are very huge to be read in. So I extracted only 10000 rows of the training dataset in the read.csv function with nrows=10000 and created a simple SVM model. But I cannot do the same with the test dataset because while submitting we got to submit the entire test dataset.\nWhen I try to  feed the test data(somewhere around 2500000+ records) to my SVM model it runs for sometime and ends up with cannot allocate memory of 93GB. I am using a machine intel i5 processor with 16GB of RAM. I know that R stores data only in RAM but I am not sure how to handle this situation. If any body could provide me a suggestion or how are you guys handling this situation, it would be really great and it will like a knowledge sharing for all of us. </p>",
      "rawMarkdown": "Hello all, first of all I am new to kaggle compeditions, and i am using R for this compedition\r\nWe all are aware that the training and testing datasets are very huge to be read in. So I extracted only 10000 rows of the training dataset in the read.csv function with nrows=10000 and created a simple SVM model. But I cannot do the same with the test dataset because while submitting we got to submit the entire test dataset.\r\nWhen I try to  feed the test data(somewhere around 2500000+ records) to my SVM model it runs for sometime and ends up with cannot allocate memory of 93GB. I am using a machine intel i5 processor with 16GB of RAM. I know that R stores data only in RAM but I am not sure how to handle this situation. If any body could provide me a suggestion or how are you guys handling this situation, it would be really great and it will like a knowledge sharing for all of us.",
      "votes": null
    },
    {
      "id": "116176",
      "postDate": "04/22/2016 12:18:51",
      "content": "<p>I'm also having this issue.  Right now, I'm trying something that's is not R optimized.  I'm trying a for loop for the 2.5+M rows.  What I'm doing is that in every predicted row, I'm compresing the 100 variables to just 1 variable (hotel cluster with top 5). </p>\n\n<p>This way I may get a 2.5M+ x 2 (including ID) instead of 2.5M+ rows x 101. (Big memory save)</p>\n\n<p>The  big issue is that R is no efficient with for loops, so far my computer has been running for 8 hours (just for the prediction).</p>\n\n<p>Probably a  more efficient way will be to deal with chunks of rows and the rbind the results.</p>",
      "rawMarkdown": "I'm also having this issue.  Right now, I'm trying something that's is not R optimized.  I'm trying a for loop for the 2.5+M rows.  What I'm doing is that in every predicted row, I'm compresing the 100 variables to just 1 variable (hotel cluster with top 5). \r\n\r\nThis way I may get a 2.5M+ x 2 (including ID) instead of 2.5M+ rows x 101. (Big memory save)\r\n\r\nThe  big issue is that R is no efficient with for loops, so far my computer has been running for 8 hours (just for the prediction).\r\n\r\nProbably a  more efficient way will be to deal with chunks of rows and the rbind the results.",
      "votes": null
    },
    {
      "id": "116701",
      "postDate": "04/25/2016 17:25:17",
      "content": "<p>I haven't started working with the dataset yet, but the prediction part should be trouble free (at least in Python), since all you have to do is break the test set in chunks that can be handled by your computer and predict each one individually and concatenate at the end. </p>\n\n<p>There should be no penalty doing this since the prediction is done one sample at a time any way.</p>",
      "rawMarkdown": "I haven't started working with the dataset yet, but the prediction part should be trouble free (at least in Python), since all you have to do is break the test set in chunks that can be handled by your computer and predict each one individually and concatenate at the end. \r\n\r\nThere should be no penalty doing this since the prediction is done one sample at a time any way.",
      "votes": null
    },
    {
      "id": "116948",
      "postDate": "04/26/2016 16:43:29",
      "content": "<p>Hi, What is your target parameter for the SVM?</p>",
      "rawMarkdown": "Hi, What is your target parameter for the SVM?",
      "votes": null
    },
    {
      "id": "116949",
      "postDate": "04/26/2016 16:44:12",
      "content": "<p>Don't batch gradient descent solves the memory problem?</p>",
      "rawMarkdown": "Don't batch gradient descent solves the memory problem?",
      "votes": null
    },
    {
      "id": "117105",
      "postDate": "04/27/2016 12:33:31",
      "content": "<p>@Dotank :My target variable is hotel_cluster and also can you please help or share code how to implement \n batch gradient descent in R ?</p>",
      "rawMarkdown": "Dotank :My target variable is hotel_cluster and also can you please help or share code how to implement \r\n batch gradient descent in R ?",
      "votes": null
    },
    {
      "id": "117111",
      "postDate": "04/27/2016 12:54:49",
      "content": "<p>I <a href=\"http://www.r-bloggers.com/regression-via-gradient-descent-in-r/\">http://www.r-bloggers.com/regression-via-gradient-descent-in-r/</a></p>",
      "rawMarkdown": "I http://www.r-bloggers.com/regression-via-gradient-descent-in-r/",
      "votes": null
    },
    {
      "id": "117112",
      "postDate": "04/27/2016 12:55:07",
      "content": "<p>I <a href=\"http://www.r-bloggers.com/regression-via-gradient-descent-in-r/\">http://www.r-bloggers.com/regression-via-gradient-descent-in-r/</a></p>",
      "rawMarkdown": "I http://www.r-bloggers.com/regression-via-gradient-descent-in-r/",
      "votes": null
    },
    {
      "id": "117296",
      "postDate": "04/28/2016 12:11:54",
      "content": "<p>Okay this piece of code in R helped in tackling the problem of handling huge test dataset of 2528243 records. But the for loop ran for about 12 hours in my machine which I felt it to be better than getting insufficient memory error. Though I have not tried any other algorithm yet, I wanted to get started with the SVM. But still handling the huge chunk of train data is still difficult(it throws insufficient memory error) for which I used only 10000 records for training the model</p>\n\n<pre><code>for (i in 1:nrow(test))\n{ #print(i)\n  SVM.model.predict.local[i] = predict(SVM.model, newdata=test[i,])\n}\n</code></pre>",
      "rawMarkdown": "Okay this piece of code in R helped in tackling the problem of handling huge test dataset of 2528243 records. But the for loop ran for about 12 hours in my machine which I felt it to be better than getting insufficient memory error. Though I have not tried any other algorithm yet, I wanted to get started with the SVM. But still handling the huge chunk of train data is still difficult(it throws insufficient memory error) for which I used only 10000 records for training the model\r\n\r\n    for (i in 1:nrow(test))\r\n    { #print(i)\r\n      SVM.model.predict.local[i] = predict(SVM.model, newdata=test[i,])\r\n    }",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116176,
      "author_name": "ikleiman",
      "author_url": "",
      "post_date": "04/22/2016 12:18:51",
      "content": "<p>I'm also having this issue.  Right now, I'm trying something that's is not R optimized.  I'm trying a for loop for the 2.5+M rows.  What I'm doing is that in every predicted row, I'm compresing the 100 variables to just 1 variable (hotel cluster with top 5). </p>\n\n<p>This way I may get a 2.5M+ x 2 (including ID) instead of 2.5M+ rows x 101. (Big memory save)</p>\n\n<p>The  big issue is that R is no efficient with for loops, so far my computer has been running for 8 hours (just for the prediction).</p>\n\n<p>Probably a  more efficient way will be to deal with chunks of rows and the rbind the results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116701,
      "author_name": "khaoticmind",
      "author_url": "",
      "post_date": "04/25/2016 17:25:17",
      "content": "<p>I haven't started working with the dataset yet, but the prediction part should be trouble free (at least in Python), since all you have to do is break the test set in chunks that can be handled by your computer and predict each one individually and concatenate at the end. </p>\n\n<p>There should be no penalty doing this since the prediction is done one sample at a time any way.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116948,
      "author_name": "dotank",
      "author_url": "",
      "post_date": "04/26/2016 16:43:29",
      "content": "<p>Hi, What is your target parameter for the SVM?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116949,
      "author_name": "dotank",
      "author_url": "",
      "post_date": "04/26/2016 16:44:12",
      "content": "<p>Don't batch gradient descent solves the memory problem?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117105,
      "author_name": "thanish",
      "author_url": "",
      "post_date": "04/27/2016 12:33:31",
      "content": "<p>@Dotank :My target variable is hotel_cluster and also can you please help or share code how to implement \n batch gradient descent in R ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117111,
      "author_name": "dotank",
      "author_url": "",
      "post_date": "04/27/2016 12:54:49",
      "content": "<p>I <a href=\"http://www.r-bloggers.com/regression-via-gradient-descent-in-r/\">http://www.r-bloggers.com/regression-via-gradient-descent-in-r/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117112,
      "author_name": "dotank",
      "author_url": "",
      "post_date": "04/27/2016 12:55:07",
      "content": "<p>I <a href=\"http://www.r-bloggers.com/regression-via-gradient-descent-in-r/\">http://www.r-bloggers.com/regression-via-gradient-descent-in-r/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117296,
      "author_name": "thanish",
      "author_url": "",
      "post_date": "04/28/2016 12:11:54",
      "content": "<p>Okay this piece of code in R helped in tackling the problem of handling huge test dataset of 2528243 records. But the for loop ran for about 12 hours in my machine which I felt it to be better than getting insufficient memory error. Though I have not tried any other algorithm yet, I wanted to get started with the SVM. But still handling the huge chunk of train data is still difficult(it throws insufficient memory error) for which I used only 10000 records for training the model</p>\n\n<pre><code>for (i in 1:nrow(test))\n{ #print(i)\n  SVM.model.predict.local[i] = predict(SVM.model, newdata=test[i,])\n}\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116159": "Hello all, first of all I am new to kaggle compeditions, and i am using R for this compedition\r\nWe all are aware that the training and testing datasets are very huge to be read in. So I extracted only 10000 rows of the training dataset in the read.csv function with nrows=10000 and created a simple SVM model. But I cannot do the same with the test dataset because while submitting we got to submit the entire test dataset.\r\nWhen I try to  feed the test data(somewhere around 2500000+ records) to my SVM model it runs for sometime and ends up with cannot allocate memory of 93GB. I am using a machine intel i5 processor with 16GB of RAM. I know that R stores data only in RAM but I am not sure how to handle this situation. If any body could provide me a suggestion or how are you guys handling this situation, it would be really great and it will like a knowledge sharing for all of us.",
    "116176": "I'm also having this issue.  Right now, I'm trying something that's is not R optimized.  I'm trying a for loop for the 2.5+M rows.  What I'm doing is that in every predicted row, I'm compresing the 100 variables to just 1 variable (hotel cluster with top 5). \r\n\r\nThis way I may get a 2.5M+ x 2 (including ID) instead of 2.5M+ rows x 101. (Big memory save)\r\n\r\nThe  big issue is that R is no efficient with for loops, so far my computer has been running for 8 hours (just for the prediction).\r\n\r\nProbably a  more efficient way will be to deal with chunks of rows and the rbind the results.",
    "116701": "I haven't started working with the dataset yet, but the prediction part should be trouble free (at least in Python), since all you have to do is break the test set in chunks that can be handled by your computer and predict each one individually and concatenate at the end. \r\n\r\nThere should be no penalty doing this since the prediction is done one sample at a time any way.",
    "116948": "Hi, What is your target parameter for the SVM?",
    "116949": "Don't batch gradient descent solves the memory problem?",
    "117105": "Dotank :My target variable is hotel_cluster and also can you please help or share code how to implement \r\n batch gradient descent in R ?",
    "117111": "I http://www.r-bloggers.com/regression-via-gradient-descent-in-r/",
    "117112": "I http://www.r-bloggers.com/regression-via-gradient-descent-in-r/",
    "117296": "Okay this piece of code in R helped in tackling the problem of handling huge test dataset of 2528243 records. But the for loop ran for about 12 hours in my machine which I felt it to be better than getting insufficient memory error. Though I have not tried any other algorithm yet, I wanted to get started with the SVM. But still handling the huge chunk of train data is still difficult(it throws insufficient memory error) for which I used only 10000 records for training the model\r\n\r\n    for (i in 1:nrow(test))\r\n    { #print(i)\r\n      SVM.model.predict.local[i] = predict(SVM.model, newdata=test[i,])\r\n    }"
  },
  "source": "meta"
}