{
  "id": 201428,
  "title": " How much can we do in a \"limited\" environment?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201428",
  "author_name": "",
  "post_date": "2020-12-05T00:49:50.852378800Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm wondering how much we can perform better in this competition with Kaggle environment or free Google Colab environment. </p>\n<p>Now I managed to do basic feature engineering with cuDF on Kaggle notebook. This is very very fast. Amazing!<br>\nThen I'm trying to make some models with it. For the starter, I'm using CatBoost(<code>task_type='GPU'</code>) which is my favorite. Training is slower than LGBM, but inference is faster than LGBM, am I right? I still don't know it is the case in this competition as well.</p>\n<p>Anyway, I'm trying to make models.</p>\n<p>My instinct says 100M can't be used to train the model because of the memory limitation. <br>\nThe first try. I fed 30M records, training finished and… I got RAM OOM. Of course. <br>\nOkay then, 10M… Got OOM.<br>\nHmm, 6M?… OOM. Really? <br>\nFine, 1M?… Okay, got no errors.</p>\n<p>In some discussion, you guys say you use all the data. Does it mean you fed all the data to a model to train? <br>\nIf so, how can it be possible with Kaggle or free Colab environment? Or should I consider 'pay' for some kind of services? </p>\n<p>Or is there any tricks?<br>\nWhat I can think of for now are:</p>\n<ul>\n<li>Just throw away some data as 1M is already large enough</li>\n<li>Split the data into small parts and make many models to blend</li>\n<li>Some sort of online training algorithm can be applied? I guess it might be already discussed in other discussion, I just haven't caught up</li>\n</ul>\n<p>What do you think? I appreciate any opinions, ideas or comments. </p>\n<p>Thank you! </p>",
  "messages": [
    {
      "id": "1102461",
      "postDate": "12/05/2020 00:49:50",
      "content": "<p>I'm wondering how much we can perform better in this competition with Kaggle environment or free Google Colab environment. </p>\n<p>Now I managed to do basic feature engineering with cuDF on Kaggle notebook. This is very very fast. Amazing!<br>\nThen I'm trying to make some models with it. For the starter, I'm using CatBoost(<code>task_type='GPU'</code>) which is my favorite. Training is slower than LGBM, but inference is faster than LGBM, am I right? I still don't know it is the case in this competition as well.</p>\n<p>Anyway, I'm trying to make models.</p>\n<p>My instinct says 100M can't be used to train the model because of the memory limitation. <br>\nThe first try. I fed 30M records, training finished and… I got RAM OOM. Of course. <br>\nOkay then, 10M… Got OOM.<br>\nHmm, 6M?… OOM. Really? <br>\nFine, 1M?… Okay, got no errors.</p>\n<p>In some discussion, you guys say you use all the data. Does it mean you fed all the data to a model to train? <br>\nIf so, how can it be possible with Kaggle or free Colab environment? Or should I consider 'pay' for some kind of services? </p>\n<p>Or is there any tricks?<br>\nWhat I can think of for now are:</p>\n<ul>\n<li>Just throw away some data as 1M is already large enough</li>\n<li>Split the data into small parts and make many models to blend</li>\n<li>Some sort of online training algorithm can be applied? I guess it might be already discussed in other discussion, I just haven't caught up</li>\n</ul>\n<p>What do you think? I appreciate any opinions, ideas or comments. </p>\n<p>Thank you! </p>",
      "rawMarkdown": "I'm wondering how much we can perform better in this competition with Kaggle environment or free Google Colab environment. \n\nNow I managed to do basic feature engineering with cuDF on Kaggle notebook. This is very very fast. Amazing!\nThen I'm trying to make some models with it. For the starter, I'm using CatBoost(`task_type='GPU'`) which is my favorite. Training is slower than LGBM, but inference is faster than LGBM, am I right? I still don't know it is the case in this competition as well.\n\nAnyway, I'm trying to make models.\n\nMy instinct says 100M can't be used to train the model because of the memory limitation. \nThe first try. I fed 30M records, training finished and... I got RAM OOM. Of course. \nOkay then, 10M... Got OOM.\nHmm, 6M?... OOM. Really? \nFine, 1M?... Okay, got no errors.\n\nIn some discussion, you guys say you use all the data. Does it mean you fed all the data to a model to train? \nIf so, how can it be possible with Kaggle or free Colab environment? Or should I consider 'pay' for some kind of services? \n\nOr is there any tricks?\nWhat I can think of for now are:\n- Just throw away some data as 1M is already large enough\n- Split the data into small parts and make many models to blend\n- Some sort of online training algorithm can be applied? I guess it might be already discussed in other discussion, I just haven't caught up\n\nWhat do you think? I appreciate any opinions, ideas or comments. \n\nThank you!",
      "votes": null
    },
    {
      "id": "1102571",
      "postDate": "12/05/2020 04:26:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a>, I also ran into similar issues. Just created a thread on <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444\" target=\"_blank\">using Memory Profiler and reducing RAM overload</a>. I don't know if it helps solve your issue but the steps I described helped me out. </p>",
      "rawMarkdown": "Hi @kokitanisaka, I also ran into similar issues. Just created a thread on [using Memory Profiler and reducing RAM overload](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444). I don't know if it helps solve your issue but the steps I described helped me out.",
      "votes": null
    },
    {
      "id": "1103380",
      "postDate": "12/05/2020 21:16:27",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a>, thank you for the reply. And thank you for sharing your insight!</p>\n<p>Now I see what was going on. <br>\nI added much features in the post processing. Some of these features caused OOM.<br>\nAfter I removed some feats, now I can put much more records. At least I was able to put 10M records for now. </p>",
      "rawMarkdown": "Hey @manikanthr5, thank you for the reply. And thank you for sharing your insight!\n\nNow I see what was going on. \nI added much features in the post processing. Some of these features caused OOM.\nAfter I removed some feats, now I can put much more records. At least I was able to put 10M records for now.",
      "votes": null
    },
    {
      "id": "1103491",
      "postDate": "12/06/2020 00:41:58",
      "content": "<p>Yeah. Currently the same thing is happening with me. I am using all 100M records. I have few more ideas on how to solve it. Will share them if they work. </p>",
      "rawMarkdown": "Yeah. Currently the same thing is happening with me. I am using all 100M records. I have few more ideas on how to solve it. Will share them if they work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1102571,
      "author_name": "manikanthr5",
      "author_url": "",
      "post_date": "12/05/2020 04:26:15",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a>, I also ran into similar issues. Just created a thread on <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444\" target=\"_blank\">using Memory Profiler and reducing RAM overload</a>. I don't know if it helps solve your issue but the steps I described helped me out. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1103380,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "12/05/2020 21:16:27",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/manikanthr5\" target=\"_blank\">@manikanthr5</a>, thank you for the reply. And thank you for sharing your insight!</p>\n<p>Now I see what was going on. <br>\nI added much features in the post processing. Some of these features caused OOM.<br>\nAfter I removed some feats, now I can put much more records. At least I was able to put 10M records for now. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103491,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/06/2020 00:41:58",
          "content": "<p>Yeah. Currently the same thing is happening with me. I am using all 100M records. I have few more ideas on how to solve it. Will share them if they work. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1102461": "I'm wondering how much we can perform better in this competition with Kaggle environment or free Google Colab environment. \n\nNow I managed to do basic feature engineering with cuDF on Kaggle notebook. This is very very fast. Amazing!\nThen I'm trying to make some models with it. For the starter, I'm using CatBoost(`task_type='GPU'`) which is my favorite. Training is slower than LGBM, but inference is faster than LGBM, am I right? I still don't know it is the case in this competition as well.\n\nAnyway, I'm trying to make models.\n\nMy instinct says 100M can't be used to train the model because of the memory limitation. \nThe first try. I fed 30M records, training finished and... I got RAM OOM. Of course. \nOkay then, 10M... Got OOM.\nHmm, 6M?... OOM. Really? \nFine, 1M?... Okay, got no errors.\n\nIn some discussion, you guys say you use all the data. Does it mean you fed all the data to a model to train? \nIf so, how can it be possible with Kaggle or free Colab environment? Or should I consider 'pay' for some kind of services? \n\nOr is there any tricks?\nWhat I can think of for now are:\n- Just throw away some data as 1M is already large enough\n- Split the data into small parts and make many models to blend\n- Some sort of online training algorithm can be applied? I guess it might be already discussed in other discussion, I just haven't caught up\n\nWhat do you think? I appreciate any opinions, ideas or comments. \n\nThank you!",
    "1102571": "Hi @kokitanisaka, I also ran into similar issues. Just created a thread on [using Memory Profiler and reducing RAM overload](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444). I don't know if it helps solve your issue but the steps I described helped me out.",
    "1103380": "Hey @manikanthr5, thank you for the reply. And thank you for sharing your insight!\n\nNow I see what was going on. \nI added much features in the post processing. Some of these features caused OOM.\nAfter I removed some feats, now I can put much more records. At least I was able to put 10M records for now.",
    "1103491": "Yeah. Currently the same thing is happening with me. I am using all 100M records. I have few more ideas on how to solve it. Will share them if they work."
  },
  "source": "meta"
}