{
  "id": 55449,
  "title": "LGB using R. Where did I go wrong?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55449",
  "author_name": "",
  "post_date": "2018-04-26T17:38:32.353606200Z",
  "votes": 2,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Disclaimer: I use R and Google Cloud Computing (Ubuntu 16.04) </p>\n\n<p>After going through most of the kernels using LGBM, I created my own crafted 52 features (I know there are people who use less than 10 features, but I want to find out these important features on my own). </p>\n\n<p>Everything went smoothly until the moment I start lgb.train \nOnce I start it, my RAM spikes and crashes everything (I have read python solutions, but the same does not work here).</p>\n\n<p>From 60 GB RAM, I slowly tweaked it up and now , I am at 500 GB RAM. Still no progress. It still crashes (I am not a fan of chunking data. If I can have the privilege of having these arsenal amounts of RAM at my disposal, why chunk/swap?)</p>\n\n<p>After 500 GB of RAM, I realized I am either doing something extremely wrong or lightGBM does not expect this kind of features (I have number of features which I have declared as categorical but am not sure how LGBM is processing it). </p>\n\n<p>Can anyone guide me?</p>",
  "messages": [
    {
      "id": "319720",
      "postDate": "04/26/2018 17:38:32",
      "content": "<p>Disclaimer: I use R and Google Cloud Computing (Ubuntu 16.04) </p>\n\n<p>After going through most of the kernels using LGBM, I created my own crafted 52 features (I know there are people who use less than 10 features, but I want to find out these important features on my own). </p>\n\n<p>Everything went smoothly until the moment I start lgb.train \nOnce I start it, my RAM spikes and crashes everything (I have read python solutions, but the same does not work here).</p>\n\n<p>From 60 GB RAM, I slowly tweaked it up and now , I am at 500 GB RAM. Still no progress. It still crashes (I am not a fan of chunking data. If I can have the privilege of having these arsenal amounts of RAM at my disposal, why chunk/swap?)</p>\n\n<p>After 500 GB of RAM, I realized I am either doing something extremely wrong or lightGBM does not expect this kind of features (I have number of features which I have declared as categorical but am not sure how LGBM is processing it). </p>\n\n<p>Can anyone guide me?</p>",
      "rawMarkdown": "Disclaimer: I use R and Google Cloud Computing (Ubuntu 16.04) \n\nAfter going through most of the kernels using LGBM, I created my own crafted 52 features (I know there are people who use less than 10 features, but I want to find out these important features on my own). \n\nEverything went smoothly until the moment I start lgb.train \nOnce I start it, my RAM spikes and crashes everything (I have read python solutions, but the same does not work here).\n\nFrom 60 GB RAM, I slowly tweaked it up and now , I am at 500 GB RAM. Still no progress. It still crashes (I am not a fan of chunking data. If I can have the privilege of having these arsenal amounts of RAM at my disposal, why chunk/swap?)\n\nAfter 500 GB of RAM, I realized I am either doing something extremely wrong or lightGBM does not expect this kind of features (I have number of features which I have declared as categorical but am not sure how LGBM is processing it). \n\nCan anyone guide me?",
      "votes": null
    },
    {
      "id": "319736",
      "postDate": "04/26/2018 18:31:36",
      "content": "<p>Try running the model with a few hundred rows to make sure the model actually works. It could be there is a bug in hyper parameters or data setup that is causing issues.</p>\n\n<p>If you get errors, then you know where to go. If there are no errors, use print(), summary(), str() and other functions to explore the model object.</p>\n\n<p>I don't use LGBM much but I know that some of these models have a hard time with categorical variables. It might be internally converting a categorical variable into a binary class matrix. That can take up a LOT of space. If that is the case, you can try converting your categorical variables into integer... something like as.integer(factor(my_categorical_variable)). Just be sure to use the same factor levels for train and test data sets.</p>\n\n<p>Good luck,\nJeff</p>",
      "rawMarkdown": "Try running the model with a few hundred rows to make sure the model actually works. It could be there is a bug in hyper parameters or data setup that is causing issues.\n\nIf you get errors, then you know where to go. If there are no errors, use print(), summary(), str() and other functions to explore the model object.\n\nI don't use LGBM much but I know that some of these models have a hard time with categorical variables. It might be internally converting a categorical variable into a binary class matrix. That can take up a LOT of space. If that is the case, you can try converting your categorical variables into integer... something like as.integer(factor(my_categorical_variable)). Just be sure to use the same factor levels for train and test data sets.\n\nGood luck,\nJeff",
      "votes": null
    },
    {
      "id": "320001",
      "postDate": "04/27/2018 09:07:55",
      "content": "<p>Thanks. The same works properly for 1 million sample. The whole thing blows up when all the data is used in the model</p>",
      "rawMarkdown": "Thanks. The same works properly for 1 million sample. The whole thing blows up when all the data is used in the model",
      "votes": null
    },
    {
      "id": "320018",
      "postDate": "04/27/2018 09:38:35",
      "content": "<p>What does your Input look like? \nDid you specify any categorical features inside lgb.Dataset?\nAny non numeric values inside your lgb.Dataset input-Matrix? </p>",
      "rawMarkdown": "What does your Input look like? \nDid you specify any categorical features inside lgb.Dataset?\nAny non numeric values inside your lgb.Dataset input-Matrix?",
      "votes": null
    },
    {
      "id": "320029",
      "postDate": "04/27/2018 10:17:03",
      "content": "<p>Categorical variables defined in lgb.dataset - app, device, os, channel, hour,minute,secs. + 50 numeric variables</p>",
      "rawMarkdown": "Categorical variables defined in lgb.dataset - app, device, os, channel, hour,minute,secs. + 50 numeric variables",
      "votes": null
    },
    {
      "id": "320036",
      "postDate": "04/27/2018 10:32:16",
      "content": "<p>It may be caused by the implementation of AUC formula used. I do not remember the name of R package, probably AUC, which works properly with one million samples but crashes while using few million.\nIf your ML method uses this kind of implementation of AUC, you have to use another evaluation function, e.g. logloss to work with whole train, and then use auc function from Metrics package which does not crash with 190 million sample.</p>",
      "rawMarkdown": "It may be caused by the implementation of AUC formula used. I do not remember the name of R package, probably AUC, which works properly with one million samples but crashes while using few million.\nIf your ML method uses this kind of implementation of AUC, you have to use another evaluation function, e.g. logloss to work with whole train, and then use auc function from Metrics package which does not crash with 190 million sample.",
      "votes": null
    },
    {
      "id": "320037",
      "postDate": "04/27/2018 10:32:37",
      "content": "<p>R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding, for example round(x * 100000) where x is a float number.\nrunning lgb on full data using ~ 20 features takes around 20GB of memory and the memory spike at training start stays under 60GB ( i use 32 RAM and another 30 GB swap space)</p>",
      "rawMarkdown": "R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding, for example round(x * 100000) where x is a float number.\nrunning lgb on full data using ~ 20 features takes around 20GB of memory and the memory spike at training start stays under 60GB ( i use 32 RAM and another 30 GB swap space)",
      "votes": null
    },
    {
      "id": "320068",
      "postDate": "04/27/2018 11:49:56",
      "content": "<p>Are those categoricals stored as numerics? </p>",
      "rawMarkdown": "Are those categoricals stored as numerics?",
      "votes": null
    },
    {
      "id": "320102",
      "postDate": "04/27/2018 14:03:24",
      "content": "<p>All categorical features are integers</p>",
      "rawMarkdown": "All categorical features are integers",
      "votes": null
    },
    {
      "id": "320103",
      "postDate": "04/27/2018 14:08:16",
      "content": "<p><a href=\"/malten\">@malten</a> The datatype I use are all integers. </p>",
      "rawMarkdown": "malten The datatype I use are all integers.",
      "votes": null
    },
    {
      "id": "320106",
      "postDate": "04/27/2018 14:23:40",
      "content": "<p>Hm, weird. Another thing you can try is to limit the number of bins with max_bin. I use 100. Using less than in standard (256) should also reduce the memory spike</p>",
      "rawMarkdown": "Hm, weird. Another thing you can try is to limit the number of bins with max_bin. I use 100. Using less than in standard (256) should also reduce the memory spike",
      "votes": null
    },
    {
      "id": "320202",
      "postDate": "04/27/2018 21:10:28",
      "content": "<p>Not sure what you mean but that is not correct,\nI run LGBM on R with AUC optimization on the entire data.</p>",
      "rawMarkdown": "Not sure what you mean but that is not correct,\nI run LGBM on R with AUC optimization on the entire data.",
      "votes": null
    },
    {
      "id": "320204",
      "postDate": "04/27/2018 21:13:32",
      "content": "<blockquote>\n  <p>LGB using R. Where did I go wrong?</p>\n</blockquote>\n\n<p>Using R maybe?  </p>\n\n<p>I know, I know, bad joke.\nBut I didn't know about this:</p>\n\n<blockquote>\n  <p>R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding</p>\n</blockquote>\n\n<p>Look like a bad joke to me actually.</p>",
      "rawMarkdown": "&gt; LGB using R. Where did I go wrong?\n\nUsing R maybe?  \n\nI know, I know, bad joke.\nBut I didn't know about this:\n\n&gt; R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding\n\nLook like a bad joke to me actually.",
      "votes": null
    },
    {
      "id": "320231",
      "postDate": "04/28/2018 00:12:03",
      "content": "<p><a href=\"/lalthan\">@lalthan</a>, If you use minutes and seconds as categorical features, it means your should build your model from scratch (random values from the range 0:59 would be as good or rather as bad as seconds and minutes). It is a method used by a few Kagglers to add a feature containing random integer values and then analyse the feature importance table with that blind marker. <br>\nEDIT: look at the positions of the features random60, minute and second in the feature importance table of the kernel: <a href=\"https://www.kaggle.com/sionek/testing-random-against-minutes-and-seconds\">https://www.kaggle.com/sionek/testing-random-against-minutes-and-seconds</a></p>",
      "rawMarkdown": "lalthan, If you use minutes and seconds as categorical features, it means your should build your model from scratch (random values from the range 0:59 would be as good or rather as bad as seconds and minutes). It is a method used by a few Kagglers to add a feature containing random integer values and then analyse the feature importance table with that blind marker.  \nEDIT: look at the positions of the features random60, minute and second in the feature importance table of the kernel: https://www.kaggle.com/sionek/testing-random-against-minutes-and-seconds",
      "votes": null
    },
    {
      "id": "320232",
      "postDate": "04/28/2018 00:16:58",
      "content": "<p>Have you tried load the data directly into lgb.Dataset function? This procedure saved some memory for me.</p>\n\n<pre><code>dtrain &lt;- lgb.Dataset('data/train_modified.csv',\n                  free_raw_data = FALSE,\n                  colnames = col_names[-1])     # col_names is c(\"is_attributed\", \"app\", \"device\", ...)\nlgb.Dataset.construct(dtrain)\n\nlgb.Dataset.set.categorical(dtrain, \n                        categorical_feature = c(\"app\", \"device\", \"os\",\n                                                \"channel\", \"hour\"))\n</code></pre>\n\n<p>the target variable should be the first column and there is no header in train_modified.csv.</p>",
      "rawMarkdown": "Have you tried load the data directly into lgb.Dataset function? This procedure saved some memory for me.\n\n    dtrain &lt;- lgb.Dataset('data/train_modified.csv',\n                      free_raw_data = FALSE,\n                      colnames = col_names[-1])     # col_names is c(\"is_attributed\", \"app\", \"device\", ...)\n    lgb.Dataset.construct(dtrain)\n\n    lgb.Dataset.set.categorical(dtrain, \n                            categorical_feature = c(\"app\", \"device\", \"os\",\n                                                    \"channel\", \"hour\"))\n\nthe target variable should be the first column and there is no header in train_modified.csv.",
      "votes": null
    },
    {
      "id": "320734",
      "postDate": "04/29/2018 16:39:06",
      "content": "<p>The lgb.Dataset function works fine. The problem occurs as  soon as lgb.train starts (Output from lgb.dataset is passed as dataset)</p>",
      "rawMarkdown": "The lgb.Dataset function works fine. The problem occurs as  soon as lgb.train starts (Output from lgb.dataset is passed as dataset)",
      "votes": null
    },
    {
      "id": "320736",
      "postDate": "04/29/2018 16:49:49",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> , if I am not mistaken, implementation of lightgbm is same in R and Python (unless you are working on some other API which is created on top of these libraries). Also, the only difference I see between R &amp; Python is how datatype is processed by R or Python before they are passed to LGBM. I am not concerned with this difference,  as mentioned I have plenty of RAM. \n<a href=\"/laurae2\">@laurae2</a> , are you there?\nI have the feeling that LGBM's memory requirement is exponentially related to the number of columns and rows (I have around 4000 bins created by LGBM internal categorical processing. </p>",
      "rawMarkdown": "cpmpml , if I am not mistaken, implementation of lightgbm is same in R and Python (unless you are working on some other API which is created on top of these libraries). Also, the only difference I see between R &amp; Python is how datatype is processed by R or Python before they are passed to LGBM. I am not concerned with this difference,  as mentioned I have plenty of RAM. \n@laurae2 , are you there?\nI have the feeling that LGBM's memory requirement is exponentially related to the number of columns and rows (I have around 4000 bins created by LGBM internal categorical processing.",
      "votes": null
    },
    {
      "id": "320742",
      "postDate": "04/29/2018 16:57:30",
      "content": "<p>I guess this may be the reason. My LGBM created around 4000 bins. But I still don't understand, why create a limit when I have around 500 GB RAM available</p>",
      "rawMarkdown": "I guess this may be the reason. My LGBM created around 4000 bins. But I still don't understand, why create a limit when I have around 500 GB RAM available",
      "votes": null
    },
    {
      "id": "320744",
      "postDate": "04/29/2018 16:59:52",
      "content": "<p><a href=\"/tpthegreat\">@tpthegreat</a>, how many features are you usingon the entire data?</p>",
      "rawMarkdown": "tpthegreat, how many features are you usingon the entire data?",
      "votes": null
    },
    {
      "id": "320745",
      "postDate": "04/29/2018 17:02:23",
      "content": "<p><a href=\"/sionek\">@sionek</a>, what you stated is a possibility. But 500 GB RAM should be able to overcome all these limitations</p>",
      "rawMarkdown": "sionek, what you stated is a possibility. But 500 GB RAM should be able to overcome all these limitations",
      "votes": null
    },
    {
      "id": "320755",
      "postDate": "04/29/2018 17:56:57",
      "content": "<p><a href=\"/lalthan\">@lalthan</a>\nAbout 30\nYou can share the code in here or in a kernel, and i'll see if i can assist you with it, see where it is different than my implementation in terms of syntax or whatever</p>",
      "rawMarkdown": "lalthan\nAbout 30\nYou can share the code in here or in a kernel, and i'll see if i can assist you with it, see where it is different than my implementation in terms of syntax or whatever",
      "votes": null
    },
    {
      "id": "320770",
      "postDate": "04/29/2018 19:13:38",
      "content": "<p>I have 33 features, I use all train data, and lgb with Python, and my process runs well under my 64GB + 64GB swap.  I don't think it goes above 64GB actually.   lgb memory is not exponential by any mean, it is proportional to the dataset size.</p>",
      "rawMarkdown": "I have 33 features, I use all train data, and lgb with Python, and my process runs well under my 64GB + 64GB swap.  I don't think it goes above 64GB actually.   lgb memory is not exponential by any mean, it is proportional to the dataset size.",
      "votes": null
    },
    {
      "id": "321490",
      "postDate": "05/01/2018 12:15:22",
      "content": "<p><a href=\"/lalthan\">@lalthan</a> LightGBM (both in R and Python) requires data to be in numeric format (not integer). If they are integer format, they could cause a crash (crash is pre 2.1.1).</p>\n\n<p>In R, it must be a numeric matrix or a dgCMatrix (auto converted if not the case). Integer matrices are not allowed.</p>\n\n<p>As for the number of bins, it will significantly increase memory usage to go over 255/256 bins (at least twice more if not more). Categorical features are automatically scanned to be splitted together to lower RAM usage while maintaining good enough performance (it doesn't do a one-hot encoding unless the cardinality is below 4, the default limit).</p>\n\n<p>50 features on this competition shouldn't take more than 60GB RAM though (assuming you get rid of the original data first and use data from binary files, and you use 256 bins at most per feature).</p>\n\n<p>It's better if you open an issue in LightGBM GitHub as I don't monitor Kaggle forums (found it randomly thanks to Kaggle notifications I don't check often): <a href=\"https://github.com/Microsoft/LightGBM\">https://github.com/Microsoft/LightGBM</a></p>",
      "rawMarkdown": "lalthan LightGBM (both in R and Python) requires data to be in numeric format (not integer). If they are integer format, they could cause a crash (crash is pre 2.1.1).\n\nIn R, it must be a numeric matrix or a dgCMatrix (auto converted if not the case). Integer matrices are not allowed.\n\nAs for the number of bins, it will significantly increase memory usage to go over 255/256 bins (at least twice more if not more). Categorical features are automatically scanned to be splitted together to lower RAM usage while maintaining good enough performance (it doesn't do a one-hot encoding unless the cardinality is below 4, the default limit).\n\n50 features on this competition shouldn't take more than 60GB RAM though (assuming you get rid of the original data first and use data from binary files, and you use 256 bins at most per feature).\n\nIt's better if you open an issue in LightGBM GitHub as I don't monitor Kaggle forums (found it randomly thanks to Kaggle notifications I don't check often): https://github.com/Microsoft/LightGBM",
      "votes": null
    },
    {
      "id": "321532",
      "postDate": "05/01/2018 13:43:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lalthan\">Lalthan</a>, you need to do garbage collection frequently to avoid memory blowup\nyou can either do</p>\n\n<pre><code>gc()\n</code></pre>\n\n<p>or equally, if you know which variable to delete </p>\n\n<pre><code>rm( variable_name)\n</code></pre>\n\n<p>for instance suppose you have created a temp variable, you can remove from memory as follows</p>\n\n<pre><code>tmp = ....\nrm(tmp)\n</code></pre>\n\n<p>gc() causes garbage collection. It is similar to python gc.collect()</p>",
      "rawMarkdown": "Hi [Lalthan](https://www.kaggle.com/lalthan), you need to do garbage collection frequently to avoid memory blowup\nyou can either do\n\n    gc()\n\nor equally, if you know which variable to delete \n\n    rm( variable_name)\n\nfor instance suppose you have created a temp variable, you can remove from memory as follows\n\n    tmp = ....\n    rm(tmp)\n\ngc() causes garbage collection. It is similar to python gc.collect()",
      "votes": null
    },
    {
      "id": "321593",
      "postDate": "05/01/2018 16:09:57",
      "content": "<p><a href=\"/laurae2\">@laurae2</a> . Thanks. Will recheck it again as per your suggestion and convert all data to numeric. Feature creation takes around 1 hour and due to this, I saved the files after feature creation to csv and feather format. However, same problem persists. I will change the datatype and revert back</p>",
      "rawMarkdown": "laurae2 . Thanks. Will recheck it again as per your suggestion and convert all data to numeric. Feature creation takes around 1 hour and due to this, I saved the files after feature creation to csv and feather format. However, same problem persists. I will change the datatype and revert back",
      "votes": null
    },
    {
      "id": "321637",
      "postDate": "05/01/2018 17:27:54",
      "content": "<p>How large is your depth/leaves values? If they are too large, RAM usage increases exponentially (like xgboost exponential RAM increase with maximum depth).</p>",
      "rawMarkdown": "How large is your depth/leaves values? If they are too large, RAM usage increases exponentially (like xgboost exponential RAM increase with maximum depth).",
      "votes": null
    },
    {
      "id": "322245",
      "postDate": "05/02/2018 15:59:20",
      "content": "<p>I did not set any limit on max_depth. I have now kept it down to a reasonably low number. num_leaves was low from the beginning (around 10 - 20).  I was under the assumption that max_depth would be insensitive to the other parameter settings and hence wanted to see how much it can go. Am I wrong?</p>",
      "rawMarkdown": "I did not set any limit on max_depth. I have now kept it down to a reasonably low number. num_leaves was low from the beginning (around 10 - 20).  I was under the assumption that max_depth would be insensitive to the other parameter settings and hence wanted to see how much it can go. Am I wrong?",
      "votes": null
    },
    {
      "id": "323584",
      "postDate": "05/05/2018 15:57:22",
      "content": "<p>Can you provide the hyperparameter combination which leads to your 500GB RAM usage?</p>\n\n<p>N.B: you do not have any negative values in your features, right?</p>",
      "rawMarkdown": "Can you provide the hyperparameter combination which leads to your 500GB RAM usage?\n\nN.B: you do not have any negative values in your features, right?",
      "votes": null
    },
    {
      "id": "323594",
      "postDate": "05/05/2018 16:21:14",
      "content": "<p>I do have some negative values in my features due to prior and next clicks features and ignored them as I do not expect them to have real impact. My parameter settings are all within acceptable ranges and strangely seems to works fine now after reinstalling LGBM. RAM usage just for preprocessing data is around 70 GB with garbage collection and removal of unnecessary data after each step is completed.  On a closer inspection, it seems that this is not only RAM related as R server just seem to crash the moment I increase the data size beyond a particular level and may have been fixed as I no longer get the crashes for the time being</p>",
      "rawMarkdown": "I do have some negative values in my features due to prior and next clicks features and ignored them as I do not expect them to have real impact. My parameter settings are all within acceptable ranges and strangely seems to works fine now after reinstalling LGBM. RAM usage just for preprocessing data is around 70 GB with garbage collection and removal of unnecessary data after each step is completed.  On a closer inspection, it seems that this is not only RAM related as R server just seem to crash the moment I increase the data size beyond a particular level and may have been fixed as I no longer get the crashes for the time being",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 319736,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "04/26/2018 18:31:36",
      "content": "<p>Try running the model with a few hundred rows to make sure the model actually works. It could be there is a bug in hyper parameters or data setup that is causing issues.</p>\n\n<p>If you get errors, then you know where to go. If there are no errors, use print(), summary(), str() and other functions to explore the model object.</p>\n\n<p>I don't use LGBM much but I know that some of these models have a hard time with categorical variables. It might be internally converting a categorical variable into a binary class matrix. That can take up a LOT of space. If that is the case, you can try converting your categorical variables into integer... something like as.integer(factor(my_categorical_variable)). Just be sure to use the same factor levels for train and test data sets.</p>\n\n<p>Good luck,\nJeff</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320001,
      "author_name": "lalthan",
      "author_url": "",
      "post_date": "04/27/2018 09:07:55",
      "content": "<p>Thanks. The same works properly for 1 million sample. The whole thing blows up when all the data is used in the model</p>",
      "votes": null,
      "replies": [
        {
          "id": 320036,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "04/27/2018 10:32:16",
          "content": "<p>It may be caused by the implementation of AUC formula used. I do not remember the name of R package, probably AUC, which works properly with one million samples but crashes while using few million.\nIf your ML method uses this kind of implementation of AUC, you have to use another evaluation function, e.g. logloss to work with whole train, and then use auc function from Metrics package which does not crash with 190 million sample.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320202,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "04/27/2018 21:10:28",
          "content": "<p>Not sure what you mean but that is not correct,\nI run LGBM on R with AUC optimization on the entire data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320744,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/29/2018 16:59:52",
          "content": "<p><a href=\"/tpthegreat\">@tpthegreat</a>, how many features are you usingon the entire data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320755,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "04/29/2018 17:56:57",
          "content": "<p><a href=\"/lalthan\">@lalthan</a>\nAbout 30\nYou can share the code in here or in a kernel, and i'll see if i can assist you with it, see where it is different than my implementation in terms of syntax or whatever</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320018,
      "author_name": "springmanndaniel",
      "author_url": "",
      "post_date": "04/27/2018 09:38:35",
      "content": "<p>What does your Input look like? \nDid you specify any categorical features inside lgb.Dataset?\nAny non numeric values inside your lgb.Dataset input-Matrix? </p>",
      "votes": null,
      "replies": [
        {
          "id": 320029,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/27/2018 10:17:03",
          "content": "<p>Categorical variables defined in lgb.dataset - app, device, os, channel, hour,minute,secs. + 50 numeric variables</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320068,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "04/27/2018 11:49:56",
          "content": "<p>Are those categoricals stored as numerics? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320102,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/27/2018 14:03:24",
          "content": "<p>All categorical features are integers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320231,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "04/28/2018 00:12:03",
          "content": "<p><a href=\"/lalthan\">@lalthan</a>, If you use minutes and seconds as categorical features, it means your should build your model from scratch (random values from the range 0:59 would be as good or rather as bad as seconds and minutes). It is a method used by a few Kagglers to add a feature containing random integer values and then analyse the feature importance table with that blind marker. <br>\nEDIT: look at the positions of the features random60, minute and second in the feature importance table of the kernel: <a href=\"https://www.kaggle.com/sionek/testing-random-against-minutes-and-seconds\">https://www.kaggle.com/sionek/testing-random-against-minutes-and-seconds</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320745,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/29/2018 17:02:23",
          "content": "<p><a href=\"/sionek\">@sionek</a>, what you stated is a possibility. But 500 GB RAM should be able to overcome all these limitations</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320037,
      "author_name": "malten",
      "author_url": "",
      "post_date": "04/27/2018 10:32:37",
      "content": "<p>R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding, for example round(x * 100000) where x is a float number.\nrunning lgb on full data using ~ 20 features takes around 20GB of memory and the memory spike at training start stays under 60GB ( i use 32 RAM and another 30 GB swap space)</p>",
      "votes": null,
      "replies": [
        {
          "id": 320103,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/27/2018 14:08:16",
          "content": "<p><a href=\"/malten\">@malten</a> The datatype I use are all integers. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320106,
          "author_name": "malten",
          "author_url": "",
          "post_date": "04/27/2018 14:23:40",
          "content": "<p>Hm, weird. Another thing you can try is to limit the number of bins with max_bin. I use 100. Using less than in standard (256) should also reduce the memory spike</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320742,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/29/2018 16:57:30",
          "content": "<p>I guess this may be the reason. My LGBM created around 4000 bins. But I still don't understand, why create a limit when I have around 500 GB RAM available</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320204,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/27/2018 21:13:32",
      "content": "<blockquote>\n  <p>LGB using R. Where did I go wrong?</p>\n</blockquote>\n\n<p>Using R maybe?  </p>\n\n<p>I know, I know, bad joke.\nBut I didn't know about this:</p>\n\n<blockquote>\n  <p>R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding</p>\n</blockquote>\n\n<p>Look like a bad joke to me actually.</p>",
      "votes": null,
      "replies": [
        {
          "id": 320736,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/29/2018 16:49:49",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> , if I am not mistaken, implementation of lightgbm is same in R and Python (unless you are working on some other API which is created on top of these libraries). Also, the only difference I see between R &amp; Python is how datatype is processed by R or Python before they are passed to LGBM. I am not concerned with this difference,  as mentioned I have plenty of RAM. \n<a href=\"/laurae2\">@laurae2</a> , are you there?\nI have the feeling that LGBM's memory requirement is exponentially related to the number of columns and rows (I have around 4000 bins created by LGBM internal categorical processing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320770,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/29/2018 19:13:38",
          "content": "<p>I have 33 features, I use all train data, and lgb with Python, and my process runs well under my 64GB + 64GB swap.  I don't think it goes above 64GB actually.   lgb memory is not exponential by any mean, it is proportional to the dataset size.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321490,
          "author_name": "laurae2",
          "author_url": "",
          "post_date": "05/01/2018 12:15:22",
          "content": "<p><a href=\"/lalthan\">@lalthan</a> LightGBM (both in R and Python) requires data to be in numeric format (not integer). If they are integer format, they could cause a crash (crash is pre 2.1.1).</p>\n\n<p>In R, it must be a numeric matrix or a dgCMatrix (auto converted if not the case). Integer matrices are not allowed.</p>\n\n<p>As for the number of bins, it will significantly increase memory usage to go over 255/256 bins (at least twice more if not more). Categorical features are automatically scanned to be splitted together to lower RAM usage while maintaining good enough performance (it doesn't do a one-hot encoding unless the cardinality is below 4, the default limit).</p>\n\n<p>50 features on this competition shouldn't take more than 60GB RAM though (assuming you get rid of the original data first and use data from binary files, and you use 256 bins at most per feature).</p>\n\n<p>It's better if you open an issue in LightGBM GitHub as I don't monitor Kaggle forums (found it randomly thanks to Kaggle notifications I don't check often): <a href=\"https://github.com/Microsoft/LightGBM\">https://github.com/Microsoft/LightGBM</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321593,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "05/01/2018 16:09:57",
          "content": "<p><a href=\"/laurae2\">@laurae2</a> . Thanks. Will recheck it again as per your suggestion and convert all data to numeric. Feature creation takes around 1 hour and due to this, I saved the files after feature creation to csv and feather format. However, same problem persists. I will change the datatype and revert back</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321637,
          "author_name": "laurae2",
          "author_url": "",
          "post_date": "05/01/2018 17:27:54",
          "content": "<p>How large is your depth/leaves values? If they are too large, RAM usage increases exponentially (like xgboost exponential RAM increase with maximum depth).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322245,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "05/02/2018 15:59:20",
          "content": "<p>I did not set any limit on max_depth. I have now kept it down to a reasonably low number. num_leaves was low from the beginning (around 10 - 20).  I was under the assumption that max_depth would be insensitive to the other parameter settings and hence wanted to see how much it can go. Am I wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323584,
          "author_name": "laurae2",
          "author_url": "",
          "post_date": "05/05/2018 15:57:22",
          "content": "<p>Can you provide the hyperparameter combination which leads to your 500GB RAM usage?</p>\n\n<p>N.B: you do not have any negative values in your features, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323594,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "05/05/2018 16:21:14",
          "content": "<p>I do have some negative values in my features due to prior and next clicks features and ignored them as I do not expect them to have real impact. My parameter settings are all within acceptable ranges and strangely seems to works fine now after reinstalling LGBM. RAM usage just for preprocessing data is around 70 GB with garbage collection and removal of unnecessary data after each step is completed.  On a closer inspection, it seems that this is not only RAM related as R server just seem to crash the moment I increase the data size beyond a particular level and may have been fixed as I no longer get the crashes for the time being</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320232,
      "author_name": "paulofelipe",
      "author_url": "",
      "post_date": "04/28/2018 00:16:58",
      "content": "<p>Have you tried load the data directly into lgb.Dataset function? This procedure saved some memory for me.</p>\n\n<pre><code>dtrain &lt;- lgb.Dataset('data/train_modified.csv',\n                  free_raw_data = FALSE,\n                  colnames = col_names[-1])     # col_names is c(\"is_attributed\", \"app\", \"device\", ...)\nlgb.Dataset.construct(dtrain)\n\nlgb.Dataset.set.categorical(dtrain, \n                        categorical_feature = c(\"app\", \"device\", \"os\",\n                                                \"channel\", \"hour\"))\n</code></pre>\n\n<p>the target variable should be the first column and there is no header in train_modified.csv.</p>",
      "votes": null,
      "replies": [
        {
          "id": 320734,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "04/29/2018 16:39:06",
          "content": "<p>The lgb.Dataset function works fine. The problem occurs as  soon as lgb.train starts (Output from lgb.dataset is passed as dataset)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321532,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/01/2018 13:43:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lalthan\">Lalthan</a>, you need to do garbage collection frequently to avoid memory blowup\nyou can either do</p>\n\n<pre><code>gc()\n</code></pre>\n\n<p>or equally, if you know which variable to delete </p>\n\n<pre><code>rm( variable_name)\n</code></pre>\n\n<p>for instance suppose you have created a temp variable, you can remove from memory as follows</p>\n\n<pre><code>tmp = ....\nrm(tmp)\n</code></pre>\n\n<p>gc() causes garbage collection. It is similar to python gc.collect()</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "319720": "Disclaimer: I use R and Google Cloud Computing (Ubuntu 16.04) \n\nAfter going through most of the kernels using LGBM, I created my own crafted 52 features (I know there are people who use less than 10 features, but I want to find out these important features on my own). \n\nEverything went smoothly until the moment I start lgb.train \nOnce I start it, my RAM spikes and crashes everything (I have read python solutions, but the same does not work here).\n\nFrom 60 GB RAM, I slowly tweaked it up and now , I am at 500 GB RAM. Still no progress. It still crashes (I am not a fan of chunking data. If I can have the privilege of having these arsenal amounts of RAM at my disposal, why chunk/swap?)\n\nAfter 500 GB of RAM, I realized I am either doing something extremely wrong or lightGBM does not expect this kind of features (I have number of features which I have declared as categorical but am not sure how LGBM is processing it). \n\nCan anyone guide me?",
    "319736": "Try running the model with a few hundred rows to make sure the model actually works. It could be there is a bug in hyper parameters or data setup that is causing issues.\n\nIf you get errors, then you know where to go. If there are no errors, use print(), summary(), str() and other functions to explore the model object.\n\nI don't use LGBM much but I know that some of these models have a hard time with categorical variables. It might be internally converting a categorical variable into a binary class matrix. That can take up a LOT of space. If that is the case, you can try converting your categorical variables into integer... something like as.integer(factor(my_categorical_variable)). Just be sure to use the same factor levels for train and test data sets.\n\nGood luck,\nJeff",
    "320001": "Thanks. The same works properly for 1 million sample. The whole thing blows up when all the data is used in the model",
    "320018": "What does your Input look like? \nDid you specify any categorical features inside lgb.Dataset?\nAny non numeric values inside your lgb.Dataset input-Matrix?",
    "320029": "Categorical variables defined in lgb.dataset - app, device, os, channel, hour,minute,secs. + 50 numeric variables",
    "320036": "It may be caused by the implementation of AUC formula used. I do not remember the name of R package, probably AUC, which works properly with one million samples but crashes while using few million.\nIf your ML method uses this kind of implementation of AUC, you have to use another evaluation function, e.g. logloss to work with whole train, and then use auc function from Metrics package which does not crash with 190 million sample.",
    "320037": "R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding, for example round(x * 100000) where x is a float number.\nrunning lgb on full data using ~ 20 features takes around 20GB of memory and the memory spike at training start stays under 60GB ( i use 32 RAM and another 30 GB swap space)",
    "320068": "Are those categoricals stored as numerics?",
    "320102": "All categorical features are integers",
    "320103": "malten The datatype I use are all integers.",
    "320106": "Hm, weird. Another thing you can try is to limit the number of bins with max_bin. I use 100. Using less than in standard (256) should also reduce the memory spike",
    "320202": "Not sure what you mean but that is not correct,\nI run LGBM on R with AUC optimization on the entire data.",
    "320204": "&gt; LGB using R. Where did I go wrong?\n\nUsing R maybe?  \n\nI know, I know, bad joke.\nBut I didn't know about this:\n\n&gt; R needs a lot of memory when using float matrices. What i did in this competition is to conver floats to integers by multiplying with a large number followed by rounding\n\nLook like a bad joke to me actually.",
    "320231": "lalthan, If you use minutes and seconds as categorical features, it means your should build your model from scratch (random values from the range 0:59 would be as good or rather as bad as seconds and minutes). It is a method used by a few Kagglers to add a feature containing random integer values and then analyse the feature importance table with that blind marker.  \nEDIT: look at the positions of the features random60, minute and second in the feature importance table of the kernel: https://www.kaggle.com/sionek/testing-random-against-minutes-and-seconds",
    "320232": "Have you tried load the data directly into lgb.Dataset function? This procedure saved some memory for me.\n\n    dtrain &lt;- lgb.Dataset('data/train_modified.csv',\n                      free_raw_data = FALSE,\n                      colnames = col_names[-1])     # col_names is c(\"is_attributed\", \"app\", \"device\", ...)\n    lgb.Dataset.construct(dtrain)\n\n    lgb.Dataset.set.categorical(dtrain, \n                            categorical_feature = c(\"app\", \"device\", \"os\",\n                                                    \"channel\", \"hour\"))\n\nthe target variable should be the first column and there is no header in train_modified.csv.",
    "320734": "The lgb.Dataset function works fine. The problem occurs as  soon as lgb.train starts (Output from lgb.dataset is passed as dataset)",
    "320736": "cpmpml , if I am not mistaken, implementation of lightgbm is same in R and Python (unless you are working on some other API which is created on top of these libraries). Also, the only difference I see between R &amp; Python is how datatype is processed by R or Python before they are passed to LGBM. I am not concerned with this difference,  as mentioned I have plenty of RAM. \n@laurae2 , are you there?\nI have the feeling that LGBM's memory requirement is exponentially related to the number of columns and rows (I have around 4000 bins created by LGBM internal categorical processing.",
    "320742": "I guess this may be the reason. My LGBM created around 4000 bins. But I still don't understand, why create a limit when I have around 500 GB RAM available",
    "320744": "tpthegreat, how many features are you usingon the entire data?",
    "320745": "sionek, what you stated is a possibility. But 500 GB RAM should be able to overcome all these limitations",
    "320755": "lalthan\nAbout 30\nYou can share the code in here or in a kernel, and i'll see if i can assist you with it, see where it is different than my implementation in terms of syntax or whatever",
    "320770": "I have 33 features, I use all train data, and lgb with Python, and my process runs well under my 64GB + 64GB swap.  I don't think it goes above 64GB actually.   lgb memory is not exponential by any mean, it is proportional to the dataset size.",
    "321490": "lalthan LightGBM (both in R and Python) requires data to be in numeric format (not integer). If they are integer format, they could cause a crash (crash is pre 2.1.1).\n\nIn R, it must be a numeric matrix or a dgCMatrix (auto converted if not the case). Integer matrices are not allowed.\n\nAs for the number of bins, it will significantly increase memory usage to go over 255/256 bins (at least twice more if not more). Categorical features are automatically scanned to be splitted together to lower RAM usage while maintaining good enough performance (it doesn't do a one-hot encoding unless the cardinality is below 4, the default limit).\n\n50 features on this competition shouldn't take more than 60GB RAM though (assuming you get rid of the original data first and use data from binary files, and you use 256 bins at most per feature).\n\nIt's better if you open an issue in LightGBM GitHub as I don't monitor Kaggle forums (found it randomly thanks to Kaggle notifications I don't check often): https://github.com/Microsoft/LightGBM",
    "321532": "Hi [Lalthan](https://www.kaggle.com/lalthan), you need to do garbage collection frequently to avoid memory blowup\nyou can either do\n\n    gc()\n\nor equally, if you know which variable to delete \n\n    rm( variable_name)\n\nfor instance suppose you have created a temp variable, you can remove from memory as follows\n\n    tmp = ....\n    rm(tmp)\n\ngc() causes garbage collection. It is similar to python gc.collect()",
    "321593": "laurae2 . Thanks. Will recheck it again as per your suggestion and convert all data to numeric. Feature creation takes around 1 hour and due to this, I saved the files after feature creation to csv and feather format. However, same problem persists. I will change the datatype and revert back",
    "321637": "How large is your depth/leaves values? If they are too large, RAM usage increases exponentially (like xgboost exponential RAM increase with maximum depth).",
    "322245": "I did not set any limit on max_depth. I have now kept it down to a reasonably low number. num_leaves was low from the beginning (around 10 - 20).  I was under the assumption that max_depth would be insensitive to the other parameter settings and hence wanted to see how much it can go. Am I wrong?",
    "323584": "Can you provide the hyperparameter combination which leads to your 500GB RAM usage?\n\nN.B: you do not have any negative values in your features, right?",
    "323594": "I do have some negative values in my features due to prior and next clicks features and ignored them as I do not expect them to have real impact. My parameter settings are all within acceptable ranges and strangely seems to works fine now after reinstalling LGBM. RAM usage just for preprocessing data is around 70 GB with garbage collection and removal of unnecessary data after each step is completed.  On a closer inspection, it seems that this is not only RAM related as R server just seem to crash the moment I increase the data size beyond a particular level and may have been fixed as I no longer get the crashes for the time being"
  },
  "source": "meta"
}