{
  "id": 56208,
  "title": "This competition is about RAM access",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56208",
  "author_name": "",
  "post_date": "2018-05-07T19:22:55.822958100Z",
  "votes": 1,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Even though there has been some very nice discussions on how to save RAM and remove memory spikes, I feel that this competition si about accessing machines and servers with enough RAM. You can imagine and spend time creating nice features but whenever you cannot load up the full database to test it, it makes the game harder for people like me that have limited resources. Let me know how you also feel.</p>\n\n<p>At the end of the day, it is very interesting as it teaches you to do systematically gc.collect() and del unused variables in your code as you make progress on the script. Coming from the C++ programming community, this is a very good habit. Language like python and R gives you the false impression that memory will be managed efficiently for you..</p>\n\n<p>This challenge shows you that all in all, if you want to manage your memory, you cannot avoid the tedious but necessary task to clean your variables... And enjoy the competition!</p>",
  "messages": [
    {
      "id": "324535",
      "postDate": "05/07/2018 19:22:55",
      "content": "<p>Even though there has been some very nice discussions on how to save RAM and remove memory spikes, I feel that this competition si about accessing machines and servers with enough RAM. You can imagine and spend time creating nice features but whenever you cannot load up the full database to test it, it makes the game harder for people like me that have limited resources. Let me know how you also feel.</p>\n\n<p>At the end of the day, it is very interesting as it teaches you to do systematically gc.collect() and del unused variables in your code as you make progress on the script. Coming from the C++ programming community, this is a very good habit. Language like python and R gives you the false impression that memory will be managed efficiently for you..</p>\n\n<p>This challenge shows you that all in all, if you want to manage your memory, you cannot avoid the tedious but necessary task to clean your variables... And enjoy the competition!</p>",
      "rawMarkdown": "Even though there has been some very nice discussions on how to save RAM and remove memory spikes, I feel that this competition si about accessing machines and servers with enough RAM. You can imagine and spend time creating nice features but whenever you cannot load up the full database to test it, it makes the game harder for people like me that have limited resources. Let me know how you also feel.\n\nAt the end of the day, it is very interesting as it teaches you to do systematically gc.collect() and del unused variables in your code as you make progress on the script. Coming from the C++ programming community, this is a very good habit. Language like python and R gives you the false impression that memory will be managed efficiently for you..\n\nThis challenge shows you that all in all, if you want to manage your memory, you cannot avoid the tedious but necessary task to clean your variables... And enjoy the competition!",
      "votes": null
    },
    {
      "id": "324547",
      "postDate": "05/07/2018 19:29:46",
      "content": "<p>not really, actually I spent a lot of money on AWS... But I still cant defeat those top guys. Still have too much to learn and I'm looking forward to their solutions.</p>",
      "rawMarkdown": "not really, actually I spent a lot of money on AWS... But I still cant defeat those top guys. Still have too much to learn and I'm looking forward to their solutions.",
      "votes": null
    },
    {
      "id": "324556",
      "postDate": "05/07/2018 19:37:17",
      "content": "<p>I personally think that there are ways around \"RAM\" issues at the cost of time ... </p>\n\n<p>Why do you need everything to be in memory? Ok, it's faster if everything is in memory ... but if you can't have everything in memory, aren't there any solutions? </p>\n\n<p>The competitions description states \"They handle 3 billion clicks per day\". Would anybody attempt to load all 3 billion clicks in memory to train a model? </p>\n\n<p>I personally don't have any experience with anything else but NN, and so ... batching is always the way to go. One decides on how to load the batches. I would imagine that any other kind of architectures should be able to do something similar. Even if computations are done in parallel I don't think all the data gets used in the same time ... hence data can be lazy loaded / unloaded ... </p>",
      "rawMarkdown": "I personally think that there are ways around \"RAM\" issues at the cost of time ... \n\nWhy do you need everything to be in memory? Ok, it's faster if everything is in memory ... but if you can't have everything in memory, aren't there any solutions? \n\nThe competitions description states \"They handle 3 billion clicks per day\". Would anybody attempt to load all 3 billion clicks in memory to train a model? \n\nI personally don't have any experience with anything else but NN, and so ... batching is always the way to go. One decides on how to load the batches. I would imagine that any other kind of architectures should be able to do something similar. Even if computations are done in parallel I don't think all the data gets used in the same time ... hence data can be lazy loaded / unloaded ...",
      "votes": null
    },
    {
      "id": "324657",
      "postDate": "05/07/2018 22:00:51",
      "content": "<p>Agree with @Snorlax; this is not about RAM.  </p>\n\n<p>I use GCP and have more than 72 gb RAM but I am nowhere close to the top 500.</p>\n\n<p>I started late and my guess is that my features are not good enough.  For example: </p>\n\n<p>My features dont capture and utilize the temporal information in the training data. </p>",
      "rawMarkdown": "Agree with @Snorlax; this is not about RAM.  \n\nI use GCP and have more than 72 gb RAM but I am nowhere close to the top 500.\n\nI started late and my guess is that my features are not good enough.  For example: \n\nMy features dont capture and utilize the temporal information in the training data.",
      "votes": null
    },
    {
      "id": "324663",
      "postDate": "05/07/2018 22:02:59",
      "content": "<p>The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once. Otherwise, you can do training only on a fraction of the database at the cost of getting a model that does not perform as well as the same model on the full db</p>",
      "rawMarkdown": "The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once. Otherwise, you can do training only on a fraction of the database at the cost of getting a model that does not perform as well as the same model on the full db",
      "votes": null
    },
    {
      "id": "324668",
      "postDate": "05/07/2018 22:04:22",
      "content": "<p>For NN, maybe memory is less an issue.. It is more GPU access as far I understand. But for ensemble methods, RAM is an issue on large database</p>",
      "rawMarkdown": "For NN, maybe memory is less an issue.. It is more GPU access as far I understand. But for ensemble methods, RAM is an issue on large database",
      "votes": null
    },
    {
      "id": "324671",
      "postDate": "05/07/2018 22:05:44",
      "content": "<p>Snorlax and Mehul Sampat, top guys have not only the RAM but also the computing power and the time to work on the competition!</p>",
      "rawMarkdown": "Snorlax and Mehul Sampat, top guys have not only the RAM but also the computing power and the time to work on the competition!",
      "votes": null
    },
    {
      "id": "324675",
      "postDate": "05/07/2018 22:12:33",
      "content": "<p>I think it will be interesting to do a competition where there is no way the data could fit in any machines so that people will have to do online learning/out-of-core training. And that is reflective of machine learning in the real world.</p>\n\n<p>For this competition, the size of the dataset falls awkwardly into the gap between \"RAM is not an issue\" and \"impossible to load in RAM\". 180 million rows is still manageable if you have a 16G laptop, you just need to increase your swap size. </p>",
      "rawMarkdown": "I think it will be interesting to do a competition where there is no way the data could fit in any machines so that people will have to do online learning/out-of-core training. And that is reflective of machine learning in the real world.\n\nFor this competition, the size of the dataset falls awkwardly into the gap between \"RAM is not an issue\" and \"impossible to load in RAM\". 180 million rows is still manageable if you have a 16G laptop, you just need to increase your swap size.",
      "votes": null
    },
    {
      "id": "324708",
      "postDate": "05/07/2018 22:40:17",
      "content": "<p>Sinjunhe, Agree that if you increse your swap size, you can manage a 180 million rows db but still if you have 96 Gb of RAM, you are better off!</p>",
      "rawMarkdown": "Sinjunhe, Agree that if you increse your swap size, you can manage a 180 million rows db but still if you have 96 Gb of RAM, you are better off!",
      "votes": null
    },
    {
      "id": "325081",
      "postDate": "05/08/2018 05:24:42",
      "content": "<p>I've actually took the time to look at Microsoft's LightGBM  (<a href=\"https://github.com/Microsoft/LightGBM\">https://github.com/Microsoft/LightGBM</a>) and was amazed to see that it DOESN'T uses all the data in the same time. (doh ... nothing new here, one has to iterate through data ...)</p>\n\n<p>This means that if one bothers to create a lazy loading Dataset the memory would not be a problem. </p>\n\n<p>Obviously I would have expected Microsoft to already provide a Dataset that can do that - I can't say for sure they don't because they claim \"Capable of handling large-scale data\" and haven't checked all the code ... (e.g. how they do this <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parallel-Learning-Guide.rst\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parallel-Learning-Guide.rst</a>)</p>\n\n<p>I understand most data scientist do not come from a software development background - if they are told use Pandas to load data they would do the pd.read_csv ... fill the memory with data ... and would not consider if there's an alternative solution that would not require as much memory. </p>\n\n<p>This doesn't mean that there aren't alternative solutions to be more efficient with memory. It means there are no obvious out-of-the-box libraries that are popular in the community. </p>\n\n<p>I wonder when \"data\" scientist forgot to use Database engines and do aggregation using better things than DataFrames :-)</p>\n\n<p>I would mention that even Pandas support chunking and not reading everything in memory - but I don't think it supports aggregations through chunks - in the end Pandas is not a real Database engine. Or let's say it's \"in-memory\" database engine. </p>\n\n<p>I personally think that loading all the data into a database engine, running all the \"features\" aggregations using a database engine and creating a data loader from the database engine is more efficient and less RAM demanding that having everything in memory. </p>",
      "rawMarkdown": "I've actually took the time to look at Microsoft's LightGBM  (https://github.com/Microsoft/LightGBM) and was amazed to see that it DOESN'T uses all the data in the same time. (doh ... nothing new here, one has to iterate through data ...)\n\nThis means that if one bothers to create a lazy loading Dataset the memory would not be a problem. \n\nObviously I would have expected Microsoft to already provide a Dataset that can do that - I can't say for sure they don't because they claim \"Capable of handling large-scale data\" and haven't checked all the code ... (e.g. how they do this https://github.com/Microsoft/LightGBM/blob/master/docs/Parallel-Learning-Guide.rst)\n\nI understand most data scientist do not come from a software development background - if they are told use Pandas to load data they would do the pd.read_csv ... fill the memory with data ... and would not consider if there's an alternative solution that would not require as much memory. \n\nThis doesn't mean that there aren't alternative solutions to be more efficient with memory. It means there are no obvious out-of-the-box libraries that are popular in the community. \n\nI wonder when \"data\" scientist forgot to use Database engines and do aggregation using better things than DataFrames :-)\n\nI would mention that even Pandas support chunking and not reading everything in memory - but I don't think it supports aggregations through chunks - in the end Pandas is not a real Database engine. Or let's say it's \"in-memory\" database engine. \n\nI personally think that loading all the data into a database engine, running all the \"features\" aggregations using a database engine and creating a data loader from the database engine is more efficient and less RAM demanding that having everything in memory.",
      "votes": null
    },
    {
      "id": "325568",
      "postDate": "05/08/2018 15:01:20",
      "content": "<p>Most image competitions have data too large to fit in memory.  </p>",
      "rawMarkdown": "Most image competitions have data too large to fit in memory.",
      "votes": null
    },
    {
      "id": "325569",
      "postDate": "05/08/2018 15:02:17",
      "content": "<blockquote>\n  <p>The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once.</p>\n</blockquote>\n\n<p>Not true, both support out of core learning, just read the doc ...</p>",
      "rawMarkdown": "&gt; The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once.\n\nNot true, both support out of core learning, just read the doc ...",
      "votes": null
    },
    {
      "id": "325593",
      "postDate": "05/08/2018 15:50:57",
      "content": "<p>@CPMP \"Not true\" :-) </p>\n\n<p>Reading some more into microsoft documentation about it I've found this <a href=\"https://lightgbm.readthedocs.io/en/latest/Features.html#data-parallel\">https://lightgbm.readthedocs.io/en/latest/Features.html#data-parallel</a> so if it supports \"data\" parallel it means it supports \"sequential\" data batches as well. </p>\n\n<p>This:</p>\n\n<ol>\n<li><p>Partition data horizontally</p></li>\n<li><p>Workers use local data to construct local histograms</p></li>\n<li><p>Merge global histograms from all local histograms</p></li>\n</ol>\n\n<p>4 .Find best split from merged global histograms, then perform splits</p>\n\n<p>Can be translated to:</p>\n\n<ol>\n<li><p>Partition data horizontally</p></li>\n<li><p>Iterate through each partition and construct <em>partion</em> historgams</p></li>\n<li><p>Merge <em>partition</em> histograms from all <em>partition</em> histograms</p></li>\n</ol>\n\n<p>4 .Find best split from merged global histograms, then perform splits</p>\n\n<p>I'm not saying that's supported out of the box. I'm not sure about that. I'm saying because of the nature of how normally data gets processed it should be possible to have a data loader that supports loading chunks of data and freeing already processed data, instead of having to load all the data in memory. Obviously it would come with a bit of overhead as data gets loaded into RAM etc ... but solves the issue of needing a machine with huge memory when data is big. If it's not out of the box it would be a matter of maybe changing what the iterator returns ... if it has the required data in memory ... return it ... if not ... load it and free some old data ... </p>\n\n<p>I'm sorry if I seem to be a bit over the top in regards to this. I come from a development background where memory and computational resources are limited and one has to come up with solutions to make things work. As an old school developer I find it awful that today's developers do not care that much about computer resources ... </p>\n\n<p>Anyway, if this is not supported out of the box by microsoft ... think what microsoft sells ... you might figure out why they didn't bother solving this \"memory\" issue ...</p>\n\n<p>Out of scope ... an inspirational video I think any ML should watch - <a href=\"https://www.youtube.com/watch?v=eZdOkDtYMoo\">https://www.youtube.com/watch?v=eZdOkDtYMoo</a></p>",
      "rawMarkdown": "CPMP \"Not true\" :-) \n\nReading some more into microsoft documentation about it I've found this https://lightgbm.readthedocs.io/en/latest/Features.html#data-parallel so if it supports \"data\" parallel it means it supports \"sequential\" data batches as well. \n\nThis:\n\n1. Partition data horizontally\n\n2. Workers use local data to construct local histograms\n\n3. Merge global histograms from all local histograms\n\n4 .Find best split from merged global histograms, then perform splits\n\nCan be translated to:\n\n1. Partition data horizontally\n\n2. Iterate through each partition and construct _partion_ historgams\n\n3. Merge _partition_ histograms from all _partition_ histograms\n\n4 .Find best split from merged global histograms, then perform splits\n\nI'm not saying that's supported out of the box. I'm not sure about that. I'm saying because of the nature of how normally data gets processed it should be possible to have a data loader that supports loading chunks of data and freeing already processed data, instead of having to load all the data in memory. Obviously it would come with a bit of overhead as data gets loaded into RAM etc ... but solves the issue of needing a machine with huge memory when data is big. If it's not out of the box it would be a matter of maybe changing what the iterator returns ... if it has the required data in memory ... return it ... if not ... load it and free some old data ... \n\nI'm sorry if I seem to be a bit over the top in regards to this. I come from a development background where memory and computational resources are limited and one has to come up with solutions to make things work. As an old school developer I find it awful that today's developers do not care that much about computer resources ... \n\nAnyway, if this is not supported out of the box by microsoft ... think what microsoft sells ... you might figure out why they didn't bother solving this \"memory\" issue ...\n\n\nOut of scope ... an inspirational video I think any ML should watch - https://www.youtube.com/watch?v=eZdOkDtYMoo",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 324547,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "05/07/2018 19:29:46",
      "content": "<p>not really, actually I spent a lot of money on AWS... But I still cant defeat those top guys. Still have too much to learn and I'm looking forward to their solutions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 324657,
          "author_name": "mpsampat",
          "author_url": "",
          "post_date": "05/07/2018 22:00:51",
          "content": "<p>Agree with @Snorlax; this is not about RAM.  </p>\n\n<p>I use GCP and have more than 72 gb RAM but I am nowhere close to the top 500.</p>\n\n<p>I started late and my guess is that my features are not good enough.  For example: </p>\n\n<p>My features dont capture and utilize the temporal information in the training data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324671,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/07/2018 22:05:44",
          "content": "<p>Snorlax and Mehul Sampat, top guys have not only the RAM but also the computing power and the time to work on the competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324556,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "05/07/2018 19:37:17",
      "content": "<p>I personally think that there are ways around \"RAM\" issues at the cost of time ... </p>\n\n<p>Why do you need everything to be in memory? Ok, it's faster if everything is in memory ... but if you can't have everything in memory, aren't there any solutions? </p>\n\n<p>The competitions description states \"They handle 3 billion clicks per day\". Would anybody attempt to load all 3 billion clicks in memory to train a model? </p>\n\n<p>I personally don't have any experience with anything else but NN, and so ... batching is always the way to go. One decides on how to load the batches. I would imagine that any other kind of architectures should be able to do something similar. Even if computations are done in parallel I don't think all the data gets used in the same time ... hence data can be lazy loaded / unloaded ... </p>",
      "votes": null,
      "replies": [
        {
          "id": 324663,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/07/2018 22:02:59",
          "content": "<p>The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once. Otherwise, you can do training only on a fraction of the database at the cost of getting a model that does not perform as well as the same model on the full db</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324668,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/07/2018 22:04:22",
          "content": "<p>For NN, maybe memory is less an issue.. It is more GPU access as far I understand. But for ensemble methods, RAM is an issue on large database</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325081,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/08/2018 05:24:42",
          "content": "<p>I've actually took the time to look at Microsoft's LightGBM  (<a href=\"https://github.com/Microsoft/LightGBM\">https://github.com/Microsoft/LightGBM</a>) and was amazed to see that it DOESN'T uses all the data in the same time. (doh ... nothing new here, one has to iterate through data ...)</p>\n\n<p>This means that if one bothers to create a lazy loading Dataset the memory would not be a problem. </p>\n\n<p>Obviously I would have expected Microsoft to already provide a Dataset that can do that - I can't say for sure they don't because they claim \"Capable of handling large-scale data\" and haven't checked all the code ... (e.g. how they do this <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parallel-Learning-Guide.rst\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parallel-Learning-Guide.rst</a>)</p>\n\n<p>I understand most data scientist do not come from a software development background - if they are told use Pandas to load data they would do the pd.read_csv ... fill the memory with data ... and would not consider if there's an alternative solution that would not require as much memory. </p>\n\n<p>This doesn't mean that there aren't alternative solutions to be more efficient with memory. It means there are no obvious out-of-the-box libraries that are popular in the community. </p>\n\n<p>I wonder when \"data\" scientist forgot to use Database engines and do aggregation using better things than DataFrames :-)</p>\n\n<p>I would mention that even Pandas support chunking and not reading everything in memory - but I don't think it supports aggregations through chunks - in the end Pandas is not a real Database engine. Or let's say it's \"in-memory\" database engine. </p>\n\n<p>I personally think that loading all the data into a database engine, running all the \"features\" aggregations using a database engine and creating a data loader from the database engine is more efficient and less RAM demanding that having everything in memory. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325569,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/08/2018 15:02:17",
          "content": "<blockquote>\n  <p>The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once.</p>\n</blockquote>\n\n<p>Not true, both support out of core learning, just read the doc ...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325593,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/08/2018 15:50:57",
          "content": "<p>@CPMP \"Not true\" :-) </p>\n\n<p>Reading some more into microsoft documentation about it I've found this <a href=\"https://lightgbm.readthedocs.io/en/latest/Features.html#data-parallel\">https://lightgbm.readthedocs.io/en/latest/Features.html#data-parallel</a> so if it supports \"data\" parallel it means it supports \"sequential\" data batches as well. </p>\n\n<p>This:</p>\n\n<ol>\n<li><p>Partition data horizontally</p></li>\n<li><p>Workers use local data to construct local histograms</p></li>\n<li><p>Merge global histograms from all local histograms</p></li>\n</ol>\n\n<p>4 .Find best split from merged global histograms, then perform splits</p>\n\n<p>Can be translated to:</p>\n\n<ol>\n<li><p>Partition data horizontally</p></li>\n<li><p>Iterate through each partition and construct <em>partion</em> historgams</p></li>\n<li><p>Merge <em>partition</em> histograms from all <em>partition</em> histograms</p></li>\n</ol>\n\n<p>4 .Find best split from merged global histograms, then perform splits</p>\n\n<p>I'm not saying that's supported out of the box. I'm not sure about that. I'm saying because of the nature of how normally data gets processed it should be possible to have a data loader that supports loading chunks of data and freeing already processed data, instead of having to load all the data in memory. Obviously it would come with a bit of overhead as data gets loaded into RAM etc ... but solves the issue of needing a machine with huge memory when data is big. If it's not out of the box it would be a matter of maybe changing what the iterator returns ... if it has the required data in memory ... return it ... if not ... load it and free some old data ... </p>\n\n<p>I'm sorry if I seem to be a bit over the top in regards to this. I come from a development background where memory and computational resources are limited and one has to come up with solutions to make things work. As an old school developer I find it awful that today's developers do not care that much about computer resources ... </p>\n\n<p>Anyway, if this is not supported out of the box by microsoft ... think what microsoft sells ... you might figure out why they didn't bother solving this \"memory\" issue ...</p>\n\n<p>Out of scope ... an inspirational video I think any ML should watch - <a href=\"https://www.youtube.com/watch?v=eZdOkDtYMoo\">https://www.youtube.com/watch?v=eZdOkDtYMoo</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324675,
      "author_name": "sijunhe9248",
      "author_url": "",
      "post_date": "05/07/2018 22:12:33",
      "content": "<p>I think it will be interesting to do a competition where there is no way the data could fit in any machines so that people will have to do online learning/out-of-core training. And that is reflective of machine learning in the real world.</p>\n\n<p>For this competition, the size of the dataset falls awkwardly into the gap between \"RAM is not an issue\" and \"impossible to load in RAM\". 180 million rows is still manageable if you have a 16G laptop, you just need to increase your swap size. </p>",
      "votes": null,
      "replies": [
        {
          "id": 324708,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/07/2018 22:40:17",
          "content": "<p>Sinjunhe, Agree that if you increse your swap size, you can manage a 180 million rows db but still if you have 96 Gb of RAM, you are better off!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325568,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/08/2018 15:01:20",
          "content": "<p>Most image competitions have data too large to fit in memory.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "324535": "Even though there has been some very nice discussions on how to save RAM and remove memory spikes, I feel that this competition si about accessing machines and servers with enough RAM. You can imagine and spend time creating nice features but whenever you cannot load up the full database to test it, it makes the game harder for people like me that have limited resources. Let me know how you also feel.\n\nAt the end of the day, it is very interesting as it teaches you to do systematically gc.collect() and del unused variables in your code as you make progress on the script. Coming from the C++ programming community, this is a very good habit. Language like python and R gives you the false impression that memory will be managed efficiently for you..\n\nThis challenge shows you that all in all, if you want to manage your memory, you cannot avoid the tedious but necessary task to clean your variables... And enjoy the competition!",
    "324547": "not really, actually I spent a lot of money on AWS... But I still cant defeat those top guys. Still have too much to learn and I'm looking forward to their solutions.",
    "324556": "I personally think that there are ways around \"RAM\" issues at the cost of time ... \n\nWhy do you need everything to be in memory? Ok, it's faster if everything is in memory ... but if you can't have everything in memory, aren't there any solutions? \n\nThe competitions description states \"They handle 3 billion clicks per day\". Would anybody attempt to load all 3 billion clicks in memory to train a model? \n\nI personally don't have any experience with anything else but NN, and so ... batching is always the way to go. One decides on how to load the batches. I would imagine that any other kind of architectures should be able to do something similar. Even if computations are done in parallel I don't think all the data gets used in the same time ... hence data can be lazy loaded / unloaded ...",
    "324657": "Agree with @Snorlax; this is not about RAM.  \n\nI use GCP and have more than 72 gb RAM but I am nowhere close to the top 500.\n\nI started late and my guess is that my features are not good enough.  For example: \n\nMy features dont capture and utilize the temporal information in the training data.",
    "324663": "The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once. Otherwise, you can do training only on a fraction of the database at the cost of getting a model that does not perform as well as the same model on the full db",
    "324668": "For NN, maybe memory is less an issue.. It is more GPU access as far I understand. But for ensemble methods, RAM is an issue on large database",
    "324671": "Snorlax and Mehul Sampat, top guys have not only the RAM but also the computing power and the time to work on the competition!",
    "324675": "I think it will be interesting to do a competition where there is no way the data could fit in any machines so that people will have to do online learning/out-of-core training. And that is reflective of machine learning in the real world.\n\nFor this competition, the size of the dataset falls awkwardly into the gap between \"RAM is not an issue\" and \"impossible to load in RAM\". 180 million rows is still manageable if you have a 16G laptop, you just need to increase your swap size.",
    "324708": "Sinjunhe, Agree that if you increse your swap size, you can manage a 180 million rows db but still if you have 96 Gb of RAM, you are better off!",
    "325081": "I've actually took the time to look at Microsoft's LightGBM  (https://github.com/Microsoft/LightGBM) and was amazed to see that it DOESN'T uses all the data in the same time. (doh ... nothing new here, one has to iterate through data ...)\n\nThis means that if one bothers to create a lazy loading Dataset the memory would not be a problem. \n\nObviously I would have expected Microsoft to already provide a Dataset that can do that - I can't say for sure they don't because they claim \"Capable of handling large-scale data\" and haven't checked all the code ... (e.g. how they do this https://github.com/Microsoft/LightGBM/blob/master/docs/Parallel-Learning-Guide.rst)\n\nI understand most data scientist do not come from a software development background - if they are told use Pandas to load data they would do the pd.read_csv ... fill the memory with data ... and would not consider if there's an alternative solution that would not require as much memory. \n\nThis doesn't mean that there aren't alternative solutions to be more efficient with memory. It means there are no obvious out-of-the-box libraries that are popular in the community. \n\nI wonder when \"data\" scientist forgot to use Database engines and do aggregation using better things than DataFrames :-)\n\nI would mention that even Pandas support chunking and not reading everything in memory - but I don't think it supports aggregations through chunks - in the end Pandas is not a real Database engine. Or let's say it's \"in-memory\" database engine. \n\nI personally think that loading all the data into a database engine, running all the \"features\" aggregations using a database engine and creating a data loader from the database engine is more efficient and less RAM demanding that having everything in memory.",
    "325568": "Most image competitions have data too large to fit in memory.",
    "325569": "&gt; The problem Mihai is that if you want to rely on model like lightgbm or xgboost, you need to load the full dataset at once.\n\nNot true, both support out of core learning, just read the doc ...",
    "325593": "CPMP \"Not true\" :-) \n\nReading some more into microsoft documentation about it I've found this https://lightgbm.readthedocs.io/en/latest/Features.html#data-parallel so if it supports \"data\" parallel it means it supports \"sequential\" data batches as well. \n\nThis:\n\n1. Partition data horizontally\n\n2. Workers use local data to construct local histograms\n\n3. Merge global histograms from all local histograms\n\n4 .Find best split from merged global histograms, then perform splits\n\nCan be translated to:\n\n1. Partition data horizontally\n\n2. Iterate through each partition and construct _partion_ historgams\n\n3. Merge _partition_ histograms from all _partition_ histograms\n\n4 .Find best split from merged global histograms, then perform splits\n\nI'm not saying that's supported out of the box. I'm not sure about that. I'm saying because of the nature of how normally data gets processed it should be possible to have a data loader that supports loading chunks of data and freeing already processed data, instead of having to load all the data in memory. Obviously it would come with a bit of overhead as data gets loaded into RAM etc ... but solves the issue of needing a machine with huge memory when data is big. If it's not out of the box it would be a matter of maybe changing what the iterator returns ... if it has the required data in memory ... return it ... if not ... load it and free some old data ... \n\nI'm sorry if I seem to be a bit over the top in regards to this. I come from a development background where memory and computational resources are limited and one has to come up with solutions to make things work. As an old school developer I find it awful that today's developers do not care that much about computer resources ... \n\nAnyway, if this is not supported out of the box by microsoft ... think what microsoft sells ... you might figure out why they didn't bother solving this \"memory\" issue ...\n\n\nOut of scope ... an inspirational video I think any ML should watch - https://www.youtube.com/watch?v=eZdOkDtYMoo"
  },
  "source": "meta"
}