{
  "id": 368170,
  "title": "💡 How to deal with this competition needing so much RAM -- a couple of things that worked for me",
  "url": "/competitions/otto-recommender-system/discussion/368170",
  "author_name": "",
  "post_date": "2022-11-23T23:57:05.668546900Z",
  "votes": 28,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>Addressing the question in the title is a big topic for this competition, so I thought I'd share</p>\n<p>• what I am doing<br>\n• what has been working for me so far</p>\n<p>but also have a question I would like to ask along the way.</p>\n<h3>My weapon of choice</h3>\n<p>My favorite discovery in this competition so far has been <code>polars</code>. Even when just reading <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">a highly memory-optimized dataset that I shared here</a> it takes up only half as much RAM as <code>pandas</code>!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fa1768bbb7d27f8cdd8244e09c55ecc33%2Fram_usage.png?generation=1669245965468803&amp;alt=media\" alt=\"\"></p>\n<p>Things only get better during operations or with more complex data. So <code>polars</code> for me has been a great starting point.</p>\n<h3>My secret weapon 😄</h3>\n<p>RAM is super expensive nowadays (I suspect because of supply chain issues). Before I moved, I had a rig with 196 GB of RAM, but I sold it 😭 and am now stuck with only 64 GB (and an mATX motherboard that supports only 128GB…).</p>\n<p>I could convince myself to spend the money on a new motherboard and get an ATX case, BUT the RAM prices are ridiculous!</p>\n<p>Another option are cloud VMs. I prefer to use GCP and I could get a VM with hundreds of gigs of RAM for around $6 per hour or $1 spot (but not sure what the availability is and whether that would work).</p>\n<p>Plus the nuisance of having to move data to the cloud… set everything up. I have been treating this competition as a fun past time so far (and am trying to move as fast as I can on this as an ML project that I discuss further in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000\" target=\"_blank\">💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition</a>, so spending money on this competition is not something I am particularly keen on.</p>\n<p>But there is a solution I've found that let's me tag along 🙂 And I believe it can still make me very competitive on this dataset if I hang around long enough.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F52ecc8b9df5cd96b501a3dc2b7e61277%2Fswap.png?generation=1669246554753428&amp;alt=media\" alt=\"\"></p>\n<p>I have an NVME drive, but SSD should also work. Please take a look at the output from <code>htop</code> above. I created a <code>swapfile</code> and am using it as swap! +200GB of RAM for free 🙂</p>\n<p>Now, of course, it is not even remotely as fast as RAM. But modern drives do get quite fast. Also, I really only need that RAM for data processing. Once I filter data before training (discarding sessions where I have no ground truth for), I can train my ranking model without issue!</p>\n<p>So I am quite happy to trade longer runtime for the need to have more RAM as I can run training overnight or throughout the day as I work on other stuff. There is no hurry for me in this competition 🙂</p>\n<p>So far I am using around 20 features (I shared nearly everything in the notebooks I published), I have not even optimized the hyperparams for my ranker due to lack of time, and while I realize this may sound naive, I think this approach of using your SSD as RAM can indeed take you very far in this competition with a) minimal time investment b) no need to spend additional money.</p>\n<h3>What are other things you could do?</h3>\n<p>In one of his comments, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> mentioned you could read data in chunks and train your model that way, but I am not fully sure how that would work with ranking models?</p>\n<p>Would you train a bunch of models on subsets of rows and ensemble, is that how this would work in practice?</p>\n<p>That might be quite good anyhow to do from a pure performance perspective, ensembling many weaker learners, but one downside here is that again it would require writing a lot of code, which would require time, and that is my constraint in this competition, that I essentially have limited time I can give to this competition as I have a full-time job and a family.</p>\n<h3>The solution I really wanted to go for 😄</h3>\n<p>Some time ago I learned about the market for older servers, in particular Xeon-based machines. They get retired by companies and then sold for ridiculously low prices. Just to put this in perspective, even 64GB of RAM  would cost me a small fortune</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F56f72a49c588f868d16e66213ed7f680%2FRAM_price.png?generation=1669247469446258&amp;alt=media\" alt=\"\"></p>\n<p>But for a comparable amount of money I could get a server with 2 physical CPUs and 512 GB of RAM! How cool is that?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F54c1ab5102b3fd3f1ef6afc7ed3bb18c%2FScreenshot%202022-11-23%20085129.png?generation=1669247509331921&amp;alt=media\" alt=\"\"></p>\n<p>Now I didn't end up getting this server, but given the ratio of value to price, it does seem very tempting! 🙂</p>\n<h3>Summary</h3>\n<p>So yeah, my hope in writing this post was to share some ideas of how to address one of the biggest challenges of this competition! My <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000\" target=\"_blank\">earlier post</a> didn't do so well, so I am not sure if stuff like this is genuinely helpful and of interest to others 🙂</p>\n<p>If there will be interest in such things I'll try to share more in the future as I come across other \"creative\" solutions 🙂 </p>\n<p>Thanks for reading!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2041425",
      "postDate": "11/23/2022 23:57:05",
      "content": "<p>Hey!</p>\n<p>Addressing the question in the title is a big topic for this competition, so I thought I'd share</p>\n<p>• what I am doing<br>\n• what has been working for me so far</p>\n<p>but also have a question I would like to ask along the way.</p>\n<h3>My weapon of choice</h3>\n<p>My favorite discovery in this competition so far has been <code>polars</code>. Even when just reading <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">a highly memory-optimized dataset that I shared here</a> it takes up only half as much RAM as <code>pandas</code>!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fa1768bbb7d27f8cdd8244e09c55ecc33%2Fram_usage.png?generation=1669245965468803&amp;alt=media\" alt=\"\"></p>\n<p>Things only get better during operations or with more complex data. So <code>polars</code> for me has been a great starting point.</p>\n<h3>My secret weapon 😄</h3>\n<p>RAM is super expensive nowadays (I suspect because of supply chain issues). Before I moved, I had a rig with 196 GB of RAM, but I sold it 😭 and am now stuck with only 64 GB (and an mATX motherboard that supports only 128GB…).</p>\n<p>I could convince myself to spend the money on a new motherboard and get an ATX case, BUT the RAM prices are ridiculous!</p>\n<p>Another option are cloud VMs. I prefer to use GCP and I could get a VM with hundreds of gigs of RAM for around $6 per hour or $1 spot (but not sure what the availability is and whether that would work).</p>\n<p>Plus the nuisance of having to move data to the cloud… set everything up. I have been treating this competition as a fun past time so far (and am trying to move as fast as I can on this as an ML project that I discuss further in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000\" target=\"_blank\">💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition</a>, so spending money on this competition is not something I am particularly keen on.</p>\n<p>But there is a solution I've found that let's me tag along 🙂 And I believe it can still make me very competitive on this dataset if I hang around long enough.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F52ecc8b9df5cd96b501a3dc2b7e61277%2Fswap.png?generation=1669246554753428&amp;alt=media\" alt=\"\"></p>\n<p>I have an NVME drive, but SSD should also work. Please take a look at the output from <code>htop</code> above. I created a <code>swapfile</code> and am using it as swap! +200GB of RAM for free 🙂</p>\n<p>Now, of course, it is not even remotely as fast as RAM. But modern drives do get quite fast. Also, I really only need that RAM for data processing. Once I filter data before training (discarding sessions where I have no ground truth for), I can train my ranking model without issue!</p>\n<p>So I am quite happy to trade longer runtime for the need to have more RAM as I can run training overnight or throughout the day as I work on other stuff. There is no hurry for me in this competition 🙂</p>\n<p>So far I am using around 20 features (I shared nearly everything in the notebooks I published), I have not even optimized the hyperparams for my ranker due to lack of time, and while I realize this may sound naive, I think this approach of using your SSD as RAM can indeed take you very far in this competition with a) minimal time investment b) no need to spend additional money.</p>\n<h3>What are other things you could do?</h3>\n<p>In one of his comments, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> mentioned you could read data in chunks and train your model that way, but I am not fully sure how that would work with ranking models?</p>\n<p>Would you train a bunch of models on subsets of rows and ensemble, is that how this would work in practice?</p>\n<p>That might be quite good anyhow to do from a pure performance perspective, ensembling many weaker learners, but one downside here is that again it would require writing a lot of code, which would require time, and that is my constraint in this competition, that I essentially have limited time I can give to this competition as I have a full-time job and a family.</p>\n<h3>The solution I really wanted to go for 😄</h3>\n<p>Some time ago I learned about the market for older servers, in particular Xeon-based machines. They get retired by companies and then sold for ridiculously low prices. Just to put this in perspective, even 64GB of RAM  would cost me a small fortune</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F56f72a49c588f868d16e66213ed7f680%2FRAM_price.png?generation=1669247469446258&amp;alt=media\" alt=\"\"></p>\n<p>But for a comparable amount of money I could get a server with 2 physical CPUs and 512 GB of RAM! How cool is that?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F54c1ab5102b3fd3f1ef6afc7ed3bb18c%2FScreenshot%202022-11-23%20085129.png?generation=1669247509331921&amp;alt=media\" alt=\"\"></p>\n<p>Now I didn't end up getting this server, but given the ratio of value to price, it does seem very tempting! 🙂</p>\n<h3>Summary</h3>\n<p>So yeah, my hope in writing this post was to share some ideas of how to address one of the biggest challenges of this competition! My <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000\" target=\"_blank\">earlier post</a> didn't do so well, so I am not sure if stuff like this is genuinely helpful and of interest to others 🙂</p>\n<p>If there will be interest in such things I'll try to share more in the future as I come across other \"creative\" solutions 🙂 </p>\n<p>Thanks for reading!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nAddressing the question in the title is a big topic for this competition, so I thought I'd share\n\n• what I am doing\n• what has been working for me so far\n\nbut also have a question I would like to ask along the way.\n\n### My weapon of choice\n\nMy favorite discovery in this competition so far has been `polars`. Even when just reading [a highly memory-optimized dataset that I shared here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) it takes up only half as much RAM as `pandas`!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fa1768bbb7d27f8cdd8244e09c55ecc33%2Fram_usage.png?generation=1669245965468803&alt=media)\n\nThings only get better during operations or with more complex data. So `polars` for me has been a great starting point.\n\n### My secret weapon 😄\n\nRAM is super expensive nowadays (I suspect because of supply chain issues). Before I moved, I had a rig with 196 GB of RAM, but I sold it 😭 and am now stuck with only 64 GB (and an mATX motherboard that supports only 128GB...).\n\nI could convince myself to spend the money on a new motherboard and get an ATX case, BUT the RAM prices are ridiculous!\n\nAnother option are cloud VMs. I prefer to use GCP and I could get a VM with hundreds of gigs of RAM for around $6 per hour or $1 spot (but not sure what the availability is and whether that would work).\n\nPlus the nuisance of having to move data to the cloud... set everything up. I have been treating this competition as a fun past time so far (and am trying to move as fast as I can on this as an ML project that I discuss further in [💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000), so spending money on this competition is not something I am particularly keen on.\n\nBut there is a solution I've found that let's me tag along 🙂 And I believe it can still make me very competitive on this dataset if I hang around long enough.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F52ecc8b9df5cd96b501a3dc2b7e61277%2Fswap.png?generation=1669246554753428&alt=media)\n\nI have an NVME drive, but SSD should also work. Please take a look at the output from `htop` above. I created a `swapfile` and am using it as swap! +200GB of RAM for free 🙂\n\nNow, of course, it is not even remotely as fast as RAM. But modern drives do get quite fast. Also, I really only need that RAM for data processing. Once I filter data before training (discarding sessions where I have no ground truth for), I can train my ranking model without issue!\n\nSo I am quite happy to trade longer runtime for the need to have more RAM as I can run training overnight or throughout the day as I work on other stuff. There is no hurry for me in this competition 🙂\n\nSo far I am using around 20 features (I shared nearly everything in the notebooks I published), I have not even optimized the hyperparams for my ranker due to lack of time, and while I realize this may sound naive, I think this approach of using your SSD as RAM can indeed take you very far in this competition with a) minimal time investment b) no need to spend additional money.\n\n### What are other things you could do?\n\nIn one of his comments, @cdeotte mentioned you could read data in chunks and train your model that way, but I am not fully sure how that would work with ranking models?\n\nWould you train a bunch of models on subsets of rows and ensemble, is that how this would work in practice?\n\nThat might be quite good anyhow to do from a pure performance perspective, ensembling many weaker learners, but one downside here is that again it would require writing a lot of code, which would require time, and that is my constraint in this competition, that I essentially have limited time I can give to this competition as I have a full-time job and a family.\n\n### The solution I really wanted to go for 😄\n\nSome time ago I learned about the market for older servers, in particular Xeon-based machines. They get retired by companies and then sold for ridiculously low prices. Just to put this in perspective, even 64GB of RAM  would cost me a small fortune\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F56f72a49c588f868d16e66213ed7f680%2FRAM_price.png?generation=1669247469446258&alt=media)\n\nBut for a comparable amount of money I could get a server with 2 physical CPUs and 512 GB of RAM! How cool is that?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F54c1ab5102b3fd3f1ef6afc7ed3bb18c%2FScreenshot%202022-11-23%20085129.png?generation=1669247509331921&alt=media)\n\nNow I didn't end up getting this server, but given the ratio of value to price, it does seem very tempting! 🙂\n\n### Summary\n\nSo yeah, my hope in writing this post was to share some ideas of how to address one of the biggest challenges of this competition! My [earlier post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000) didn't do so well, so I am not sure if stuff like this is genuinely helpful and of interest to others 🙂\n\nIf there will be interest in such things I'll try to share more in the future as I come across other \"creative\" solutions 🙂 \n\nThanks for reading!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2042808",
      "postDate": "11/25/2022 03:30:11",
      "content": "<p>Just a screenshot documenting using disk as RAM works 😄</p>\n<p>Just clean up a little bit before training and you are off to the races (that's what I believe the red bars indicate -- memory that is occupied but no longer used, was needed for data transformations but not needed anymore).</p>\n<p>BTW this is <a href=\"https://www.youtube.com/watch?v=kVy3-gMdViM\" target=\"_blank\">a really cool talk on polars</a> that explains why using it is so advantageous vs pandas from memory organization and speed of execution perspective, need to figure out how to use the lazy API with query optimization for an additional boost in performance 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fcbc6dcb8fa60c59d7a6ecb2c98e135b5%2Fdisk_as_RAM.png?generation=1669346852363330&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Just a screenshot documenting using disk as RAM works 😄\n\nJust clean up a little bit before training and you are off to the races (that's what I believe the red bars indicate -- memory that is occupied but no longer used, was needed for data transformations but not needed anymore).\n\nBTW this is [a really cool talk on polars](https://www.youtube.com/watch?v=kVy3-gMdViM) that explains why using it is so advantageous vs pandas from memory organization and speed of execution perspective, need to figure out how to use the lazy API with query optimization for an additional boost in performance 🙂\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fcbc6dcb8fa60c59d7a6ecb2c98e135b5%2Fdisk_as_RAM.png?generation=1669346852363330&alt=media)",
      "votes": null
    },
    {
      "id": "2044528",
      "postDate": "11/26/2022 15:33:26",
      "content": "<p>Thank you so much for another wonderful post! Can't imagine how you have done so much given a full-time demanding job and a family 👍🙏! Please don't get discouraged by the current upvotes, these posts are fantastic, people just need time to find them.</p>",
      "rawMarkdown": "Thank you so much for another wonderful post! Can't imagine how you have done so much given a full-time demanding job and a family 👍🙏! Please don't get discouraged by the current upvotes, these posts are fantastic, people just need time to find them.",
      "votes": null
    },
    {
      "id": "2044969",
      "postDate": "11/26/2022 21:57:13",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/danielliao\" target=\"_blank\">@danielliao</a> for your words of encouragement! Do appreciate it!!! 🙂</p>",
      "rawMarkdown": "Thank you @danielliao for your words of encouragement! Do appreciate it!!! 🙂",
      "votes": null
    },
    {
      "id": "2045736",
      "postDate": "11/27/2022 16:16:19",
      "content": "<p>Great insights! I can confirm that training a new model on disjoint partitions of the training set and using them as an ensemble of ensembles does improve performance when you can't create a single, large model. I used that technique in the SIGIR competition this past summer and got 2nd with an ensemble of 20 LightGBM models.</p>",
      "rawMarkdown": "Great insights! I can confirm that training a new model on disjoint partitions of the training set and using them as an ensemble of ensembles does improve performance when you can't create a single, large model. I used that technique in the SIGIR competition this past summer and got 2nd with an ensemble of 20 LightGBM models.",
      "votes": null
    },
    {
      "id": "2046047",
      "postDate": "11/27/2022 21:43:45",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">@dpalbrecht</a>! And congrats on your awesome performance in the SIGIR challenge! 🙂 </p>",
      "rawMarkdown": "Thanks @dpalbrecht! And congrats on your awesome performance in the SIGIR challenge! 🙂",
      "votes": null
    },
    {
      "id": "2046100",
      "postDate": "11/27/2022 23:19:08",
      "content": "<p>Thanks very much 😃</p>",
      "rawMarkdown": "Thanks very much 😃",
      "votes": null
    },
    {
      "id": "2046372",
      "postDate": "11/28/2022 06:49:53",
      "content": "<p>So this is about a feature engineering competition?  I have 256G memory, if I have time, I may participate in the competition</p>",
      "rawMarkdown": "So this is about a feature engineering competition?  I have 256G memory, if I have time, I may participate in the competition",
      "votes": null
    },
    {
      "id": "2046376",
      "postDate": "11/28/2022 06:51:56",
      "content": "<p>Yes, feature engineering is a component! But even processing submission files, combining them, can require creativity due to their size 😄 5 million+ rows with 20 entries per row concatenated to a string.</p>\n<p>Out of curiosity, what kind of system do you have the 256GB RAM in?</p>",
      "rawMarkdown": "Yes, feature engineering is a component! But even processing submission files, combining them, can require creativity due to their size 😄 5 million+ rows with 20 entries per row concatenated to a string.\n\nOut of curiosity, what kind of system do you have the 256GB RAM in?",
      "votes": null
    },
    {
      "id": "2056395",
      "postDate": "12/06/2022 04:51:17",
      "content": "<p>Hello Radek. Thanks for your brilliant posts. I've learned a lot from them. But this time I tried to load \"train.parquet\" with Polars, and the estimated memory usage was very close to the size with Pandas. Is there anything wrong?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3771204%2Fedc4563589544ec09226294bec9f40e2%2F_20221205235021.png?generation=1670302266909578&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello Radek. Thanks for your brilliant posts. I've learned a lot from them. But this time I tried to load \"train.parquet\" with Polars, and the estimated memory usage was very close to the size with Pandas. Is there anything wrong?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3771204%2Fedc4563589544ec09226294bec9f40e2%2F_20221205235021.png?generation=1670302266909578&alt=media)",
      "votes": null
    },
    {
      "id": "2056398",
      "postDate": "12/06/2022 04:56:04",
      "content": "<p>hey <a href=\"https://www.kaggle.com/hdong5\" target=\"_blank\">@hdong5</a>! Thank you, great to hear you are finding my work useful! 🙂</p>\n<p>Not really sure what could be happening there. Might be a good idea to restart the notebook between the runs? But guessing you probably already did this mhmmm…</p>\n<p>In my experience, the difference was quite big even with just loading the data (at least 1:2) and it became especially large with performing operations such as joins.</p>\n<p>Hard to really say what might be going there but I wouldn't worry too much 🙂 In practice, operations that wouldn't be possible with pandas become possible with polars, so that is the practical measurement that really matters!</p>",
      "rawMarkdown": "hey @hdong5! Thank you, great to hear you are finding my work useful! 🙂\n\nNot really sure what could be happening there. Might be a good idea to restart the notebook between the runs? But guessing you probably already did this mhmmm...\n\nIn my experience, the difference was quite big even with just loading the data (at least 1:2) and it became especially large with performing operations such as joins.\n\nHard to really say what might be going there but I wouldn't worry too much 🙂 In practice, operations that wouldn't be possible with pandas become possible with polars, so that is the practical measurement that really matters!",
      "votes": null
    },
    {
      "id": "2057111",
      "postDate": "12/06/2022 18:51:38",
      "content": "<p>Cool. I'd like to try Polars this time. Thanks.</p>",
      "rawMarkdown": "Cool. I'd like to try Polars this time. Thanks.",
      "votes": null
    },
    {
      "id": "2120436",
      "postDate": "01/29/2023 15:39:45",
      "content": "<p>Whereas I'm stuck with  kaggle kernel's 32GB RAM 😂. I really should invest in a good laptop !</p>\n<p>Personnally, I think that I lost almost all my time in this competition:</p>\n<ul>\n<li>Chunking everything.</li>\n<li>Having a dozens  notebook for each step --&gt; prone to errors.</li>\n</ul>\n<p>Great post <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>",
      "rawMarkdown": "Whereas I'm stuck with  kaggle kernel's 32GB RAM 😂. I really should invest in a good laptop !\n\nPersonnally, I think that I lost almost all my time in this competition:\n- Chunking everything.\n- Having a dozens  notebook for each step --> prone to errors.\n\nGreat post @radek1",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2042808,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/25/2022 03:30:11",
      "content": "<p>Just a screenshot documenting using disk as RAM works 😄</p>\n<p>Just clean up a little bit before training and you are off to the races (that's what I believe the red bars indicate -- memory that is occupied but no longer used, was needed for data transformations but not needed anymore).</p>\n<p>BTW this is <a href=\"https://www.youtube.com/watch?v=kVy3-gMdViM\" target=\"_blank\">a really cool talk on polars</a> that explains why using it is so advantageous vs pandas from memory organization and speed of execution perspective, need to figure out how to use the lazy API with query optimization for an additional boost in performance 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fcbc6dcb8fa60c59d7a6ecb2c98e135b5%2Fdisk_as_RAM.png?generation=1669346852363330&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2044528,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/26/2022 15:33:26",
      "content": "<p>Thank you so much for another wonderful post! Can't imagine how you have done so much given a full-time demanding job and a family 👍🙏! Please don't get discouraged by the current upvotes, these posts are fantastic, people just need time to find them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2044969,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/26/2022 21:57:13",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/danielliao\" target=\"_blank\">@danielliao</a> for your words of encouragement! Do appreciate it!!! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2045736,
      "author_name": "dpalbrecht",
      "author_url": "",
      "post_date": "11/27/2022 16:16:19",
      "content": "<p>Great insights! I can confirm that training a new model on disjoint partitions of the training set and using them as an ensemble of ensembles does improve performance when you can't create a single, large model. I used that technique in the SIGIR competition this past summer and got 2nd with an ensemble of 20 LightGBM models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2046047,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/27/2022 21:43:45",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dpalbrecht\" target=\"_blank\">@dpalbrecht</a>! And congrats on your awesome performance in the SIGIR challenge! 🙂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046100,
          "author_name": "dpalbrecht",
          "author_url": "",
          "post_date": "11/27/2022 23:19:08",
          "content": "<p>Thanks very much 😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2046372,
      "author_name": "mahluo",
      "author_url": "",
      "post_date": "11/28/2022 06:49:53",
      "content": "<p>So this is about a feature engineering competition?  I have 256G memory, if I have time, I may participate in the competition</p>",
      "votes": null,
      "replies": [
        {
          "id": 2046376,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/28/2022 06:51:56",
          "content": "<p>Yes, feature engineering is a component! But even processing submission files, combining them, can require creativity due to their size 😄 5 million+ rows with 20 entries per row concatenated to a string.</p>\n<p>Out of curiosity, what kind of system do you have the 256GB RAM in?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2056395,
      "author_name": "hdong5",
      "author_url": "",
      "post_date": "12/06/2022 04:51:17",
      "content": "<p>Hello Radek. Thanks for your brilliant posts. I've learned a lot from them. But this time I tried to load \"train.parquet\" with Polars, and the estimated memory usage was very close to the size with Pandas. Is there anything wrong?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3771204%2Fedc4563589544ec09226294bec9f40e2%2F_20221205235021.png?generation=1670302266909578&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2056398,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/06/2022 04:56:04",
          "content": "<p>hey <a href=\"https://www.kaggle.com/hdong5\" target=\"_blank\">@hdong5</a>! Thank you, great to hear you are finding my work useful! 🙂</p>\n<p>Not really sure what could be happening there. Might be a good idea to restart the notebook between the runs? But guessing you probably already did this mhmmm…</p>\n<p>In my experience, the difference was quite big even with just loading the data (at least 1:2) and it became especially large with performing operations such as joins.</p>\n<p>Hard to really say what might be going there but I wouldn't worry too much 🙂 In practice, operations that wouldn't be possible with pandas become possible with polars, so that is the practical measurement that really matters!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2057111,
          "author_name": "hdong5",
          "author_url": "",
          "post_date": "12/06/2022 18:51:38",
          "content": "<p>Cool. I'd like to try Polars this time. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2120436,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/29/2023 15:39:45",
      "content": "<p>Whereas I'm stuck with  kaggle kernel's 32GB RAM 😂. I really should invest in a good laptop !</p>\n<p>Personnally, I think that I lost almost all my time in this competition:</p>\n<ul>\n<li>Chunking everything.</li>\n<li>Having a dozens  notebook for each step --&gt; prone to errors.</li>\n</ul>\n<p>Great post <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2041425": "Hey!\n\nAddressing the question in the title is a big topic for this competition, so I thought I'd share\n\n• what I am doing\n• what has been working for me so far\n\nbut also have a question I would like to ask along the way.\n\n### My weapon of choice\n\nMy favorite discovery in this competition so far has been `polars`. Even when just reading [a highly memory-optimized dataset that I shared here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) it takes up only half as much RAM as `pandas`!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fa1768bbb7d27f8cdd8244e09c55ecc33%2Fram_usage.png?generation=1669245965468803&alt=media)\n\nThings only get better during operations or with more complex data. So `polars` for me has been a great starting point.\n\n### My secret weapon 😄\n\nRAM is super expensive nowadays (I suspect because of supply chain issues). Before I moved, I had a rig with 196 GB of RAM, but I sold it 😭 and am now stuck with only 64 GB (and an mATX motherboard that supports only 128GB...).\n\nI could convince myself to spend the money on a new motherboard and get an ATX case, BUT the RAM prices are ridiculous!\n\nAnother option are cloud VMs. I prefer to use GCP and I could get a VM with hundreds of gigs of RAM for around $6 per hour or $1 spot (but not sure what the availability is and whether that would work).\n\nPlus the nuisance of having to move data to the cloud... set everything up. I have been treating this competition as a fun past time so far (and am trying to move as fast as I can on this as an ML project that I discuss further in [💡 how to move faster on an ML project and achieve more with less compute/time/energy (relevant to this competition](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000), so spending money on this competition is not something I am particularly keen on.\n\nBut there is a solution I've found that let's me tag along 🙂 And I believe it can still make me very competitive on this dataset if I hang around long enough.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F52ecc8b9df5cd96b501a3dc2b7e61277%2Fswap.png?generation=1669246554753428&alt=media)\n\nI have an NVME drive, but SSD should also work. Please take a look at the output from `htop` above. I created a `swapfile` and am using it as swap! +200GB of RAM for free 🙂\n\nNow, of course, it is not even remotely as fast as RAM. But modern drives do get quite fast. Also, I really only need that RAM for data processing. Once I filter data before training (discarding sessions where I have no ground truth for), I can train my ranking model without issue!\n\nSo I am quite happy to trade longer runtime for the need to have more RAM as I can run training overnight or throughout the day as I work on other stuff. There is no hurry for me in this competition 🙂\n\nSo far I am using around 20 features (I shared nearly everything in the notebooks I published), I have not even optimized the hyperparams for my ranker due to lack of time, and while I realize this may sound naive, I think this approach of using your SSD as RAM can indeed take you very far in this competition with a) minimal time investment b) no need to spend additional money.\n\n### What are other things you could do?\n\nIn one of his comments, @cdeotte mentioned you could read data in chunks and train your model that way, but I am not fully sure how that would work with ranking models?\n\nWould you train a bunch of models on subsets of rows and ensemble, is that how this would work in practice?\n\nThat might be quite good anyhow to do from a pure performance perspective, ensembling many weaker learners, but one downside here is that again it would require writing a lot of code, which would require time, and that is my constraint in this competition, that I essentially have limited time I can give to this competition as I have a full-time job and a family.\n\n### The solution I really wanted to go for 😄\n\nSome time ago I learned about the market for older servers, in particular Xeon-based machines. They get retired by companies and then sold for ridiculously low prices. Just to put this in perspective, even 64GB of RAM  would cost me a small fortune\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F56f72a49c588f868d16e66213ed7f680%2FRAM_price.png?generation=1669247469446258&alt=media)\n\nBut for a comparable amount of money I could get a server with 2 physical CPUs and 512 GB of RAM! How cool is that?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F54c1ab5102b3fd3f1ef6afc7ed3bb18c%2FScreenshot%202022-11-23%20085129.png?generation=1669247509331921&alt=media)\n\nNow I didn't end up getting this server, but given the ratio of value to price, it does seem very tempting! 🙂\n\n### Summary\n\nSo yeah, my hope in writing this post was to share some ideas of how to address one of the biggest challenges of this competition! My [earlier post](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368000) didn't do so well, so I am not sure if stuff like this is genuinely helpful and of interest to others 🙂\n\nIf there will be interest in such things I'll try to share more in the future as I come across other \"creative\" solutions 🙂 \n\nThanks for reading!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2042808": "Just a screenshot documenting using disk as RAM works 😄\n\nJust clean up a little bit before training and you are off to the races (that's what I believe the red bars indicate -- memory that is occupied but no longer used, was needed for data transformations but not needed anymore).\n\nBTW this is [a really cool talk on polars](https://www.youtube.com/watch?v=kVy3-gMdViM) that explains why using it is so advantageous vs pandas from memory organization and speed of execution perspective, need to figure out how to use the lazy API with query optimization for an additional boost in performance 🙂\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fcbc6dcb8fa60c59d7a6ecb2c98e135b5%2Fdisk_as_RAM.png?generation=1669346852363330&alt=media)",
    "2044528": "Thank you so much for another wonderful post! Can't imagine how you have done so much given a full-time demanding job and a family 👍🙏! Please don't get discouraged by the current upvotes, these posts are fantastic, people just need time to find them.",
    "2044969": "Thank you @danielliao for your words of encouragement! Do appreciate it!!! 🙂",
    "2045736": "Great insights! I can confirm that training a new model on disjoint partitions of the training set and using them as an ensemble of ensembles does improve performance when you can't create a single, large model. I used that technique in the SIGIR competition this past summer and got 2nd with an ensemble of 20 LightGBM models.",
    "2046047": "Thanks @dpalbrecht! And congrats on your awesome performance in the SIGIR challenge! 🙂",
    "2046100": "Thanks very much 😃",
    "2046372": "So this is about a feature engineering competition?  I have 256G memory, if I have time, I may participate in the competition",
    "2046376": "Yes, feature engineering is a component! But even processing submission files, combining them, can require creativity due to their size 😄 5 million+ rows with 20 entries per row concatenated to a string.\n\nOut of curiosity, what kind of system do you have the 256GB RAM in?",
    "2056395": "Hello Radek. Thanks for your brilliant posts. I've learned a lot from them. But this time I tried to load \"train.parquet\" with Polars, and the estimated memory usage was very close to the size with Pandas. Is there anything wrong?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3771204%2Fedc4563589544ec09226294bec9f40e2%2F_20221205235021.png?generation=1670302266909578&alt=media)",
    "2056398": "hey @hdong5! Thank you, great to hear you are finding my work useful! 🙂\n\nNot really sure what could be happening there. Might be a good idea to restart the notebook between the runs? But guessing you probably already did this mhmmm...\n\nIn my experience, the difference was quite big even with just loading the data (at least 1:2) and it became especially large with performing operations such as joins.\n\nHard to really say what might be going there but I wouldn't worry too much 🙂 In practice, operations that wouldn't be possible with pandas become possible with polars, so that is the practical measurement that really matters!",
    "2057111": "Cool. I'd like to try Polars this time. Thanks.",
    "2120436": "Whereas I'm stuck with  kaggle kernel's 32GB RAM 😂. I really should invest in a good laptop !\n\nPersonnally, I think that I lost almost all my time in this competition:\n- Chunking everything.\n- Having a dozens  notebook for each step --> prone to errors.\n\nGreat post @radek1"
  },
  "source": "meta"
}