{
  "id": 368618,
  "title": "How much RAM do I need for this comp?",
  "url": "/competitions/otto-recommender-system/discussion/368618",
  "author_name": "",
  "post_date": "2022-11-26T16:19:54.248924800Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello, I am considering trying out this competition but have been reading a few posts discussing the data size and heavy RAM utilization. Does anyone have a ballpark about how much RAM would one need to comfortably work on this competition? </p>",
  "messages": [
    {
      "id": "2044572",
      "postDate": "11/26/2022 16:19:54",
      "content": "<p>Hello, I am considering trying out this competition but have been reading a few posts discussing the data size and heavy RAM utilization. Does anyone have a ballpark about how much RAM would one need to comfortably work on this competition? </p>",
      "rawMarkdown": "Hello, I am considering trying out this competition but have been reading a few posts discussing the data size and heavy RAM utilization. Does anyone have a ballpark about how much RAM would one need to comfortably work on this competition?",
      "votes": null
    },
    {
      "id": "2044705",
      "postDate": "11/26/2022 18:09:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raphael1123\" target=\"_blank\">@raphael1123</a> ! I think that it heavily depends on the strategy you want to adopt to face up the challenge. In <a href=\"https://www.kaggle.com/code/sbunzini/user-item-collaborative-filtering-ensemble\" target=\"_blank\">this</a> I used Sparse Matrices, a memory-efficient data structure, to run basic Recommender Systems algorithms; the same approach has been used <a href=\"https://www.kaggle.com/code/cocoshe/itemcf-with-item-item-similarity-matrix\" target=\"_blank\">here</a>.</p>\n<p>The most common approach so far is the creation of co-visitation matrices, like in <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/data\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">2</a> and similar. However, the starting point is almost always this <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">parquet format</a> from <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>",
      "rawMarkdown": "Hi @raphael1123 ! I think that it heavily depends on the strategy you want to adopt to face up the challenge. In [this](https://www.kaggle.com/code/sbunzini/user-item-collaborative-filtering-ensemble) I used Sparse Matrices, a memory-efficient data structure, to run basic Recommender Systems algorithms; the same approach has been used [here](https://www.kaggle.com/code/cocoshe/itemcf-with-item-item-similarity-matrix).\n\nThe most common approach so far is the creation of co-visitation matrices, like in [1](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/data), [2](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic) and similar. However, the starting point is almost always this [parquet format](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) from @radek1",
      "votes": null
    },
    {
      "id": "2044865",
      "postDate": "11/26/2022 19:37:37",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/raphael1123\" target=\"_blank\">@raphael1123</a>, I have been trying o process the data, and takes forever, I haven't been able to load everything this point so I can't tell but there is also already some memory efficientdatatsets available you can use.</p>",
      "rawMarkdown": "Hello @raphael1123, I have been trying o process the data, and takes forever, I haven't been able to load everything this point so I can't tell but there is also already some memory efficientdatatsets available you can use.",
      "votes": null
    },
    {
      "id": "2045018",
      "postDate": "11/27/2022 00:16:36",
      "content": "<p>A lot depends on the tools that you are using. I am finding with <code>polars</code> I can get away with much less RAM. With 64GB of RAM that I have on my local machine, I don't think I would have a great time using <code>pandas</code>. </p>\n<p>I wrote about more on the RAM situation here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368170\" target=\"_blank\">💡 How to deal with this competition needing so much RAM -- a couple of things that worked for me</a>, maybe this can shed some additional light for you.</p>\n<p>In essence, I don't think this competition is much different from any other on tabular data where you have a considerable number of rows and where you have to create your own features 🙂 As in, the size of data here is not spectacularly large, but as you have to perform joins, etc, RAM becomes a consideration.</p>\n<p>At the end of the day, I feel that this is the best way to put it: \"With less RAM, you have to compensate through additional work (you need to write more code), more RAM means you can take a simplified approach to not have to deal with things like reading data in chunks\". </p>\n<p>There is a corollary to this and that is if you don't pay any attention to your RAM usage, you will run out of memory no matter how much RAM you have 😄 Which I think makes this competition even more fun!</p>",
      "rawMarkdown": "A lot depends on the tools that you are using. I am finding with `polars` I can get away with much less RAM. With 64GB of RAM that I have on my local machine, I don't think I would have a great time using `pandas`. \n\nI wrote about more on the RAM situation here: [💡 How to deal with this competition needing so much RAM -- a couple of things that worked for me](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368170), maybe this can shed some additional light for you.\n\nIn essence, I don't think this competition is much different from any other on tabular data where you have a considerable number of rows and where you have to create your own features 🙂 As in, the size of data here is not spectacularly large, but as you have to perform joins, etc, RAM becomes a consideration.\n\nAt the end of the day, I feel that this is the best way to put it: \"With less RAM, you have to compensate through additional work (you need to write more code), more RAM means you can take a simplified approach to not have to deal with things like reading data in chunks\". \n\nThere is a corollary to this and that is if you don't pay any attention to your RAM usage, you will run out of memory no matter how much RAM you have 😄 Which I think makes this competition even more fun!",
      "votes": null
    },
    {
      "id": "2045646",
      "postDate": "11/27/2022 14:32:16",
      "content": "<p>i find polars to be much better than pandas regarding this!</p>",
      "rawMarkdown": "i find polars to be much better than pandas regarding this!",
      "votes": null
    },
    {
      "id": "2053041",
      "postDate": "12/02/2022 18:50:44",
      "content": "<p>For me I am using mostly numba, but 64 gb ram looks kind of enough only, I think you will need more</p>",
      "rawMarkdown": "For me I am using mostly numba, but 64 gb ram looks kind of enough only, I think you will need more",
      "votes": null
    },
    {
      "id": "2054465",
      "postDate": "12/04/2022 06:57:27",
      "content": "<p>At least 64GB i think,  otherwise you need add a lot of extra code for chunking processing.</p>",
      "rawMarkdown": "At least 64GB i think,  otherwise you need add a lot of extra code for chunking processing.",
      "votes": null
    },
    {
      "id": "2082888",
      "postDate": "01/02/2023 04:31:30",
      "content": "<p>hi radek, thanks for sharing your thoughts. I wonder if I clearly understand your points. So to the 2 steps: candidates generation and model training. the first step is super ram costing while the second is suitable to proper ram machine?</p>",
      "rawMarkdown": "hi radek, thanks for sharing your thoughts. I wonder if I clearly understand your points. So to the 2 steps: candidates generation and model training. the first step is super ram costing while the second is suitable to proper ram machine?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2044705,
      "author_name": "sbunzini",
      "author_url": "",
      "post_date": "11/26/2022 18:09:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raphael1123\" target=\"_blank\">@raphael1123</a> ! I think that it heavily depends on the strategy you want to adopt to face up the challenge. In <a href=\"https://www.kaggle.com/code/sbunzini/user-item-collaborative-filtering-ensemble\" target=\"_blank\">this</a> I used Sparse Matrices, a memory-efficient data structure, to run basic Recommender Systems algorithms; the same approach has been used <a href=\"https://www.kaggle.com/code/cocoshe/itemcf-with-item-item-similarity-matrix\" target=\"_blank\">here</a>.</p>\n<p>The most common approach so far is the creation of co-visitation matrices, like in <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/data\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">2</a> and similar. However, the starting point is almost always this <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">parquet format</a> from <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2044865,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "11/26/2022 19:37:37",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/raphael1123\" target=\"_blank\">@raphael1123</a>, I have been trying o process the data, and takes forever, I haven't been able to load everything this point so I can't tell but there is also already some memory efficientdatatsets available you can use.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2045018,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/27/2022 00:16:36",
      "content": "<p>A lot depends on the tools that you are using. I am finding with <code>polars</code> I can get away with much less RAM. With 64GB of RAM that I have on my local machine, I don't think I would have a great time using <code>pandas</code>. </p>\n<p>I wrote about more on the RAM situation here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368170\" target=\"_blank\">💡 How to deal with this competition needing so much RAM -- a couple of things that worked for me</a>, maybe this can shed some additional light for you.</p>\n<p>In essence, I don't think this competition is much different from any other on tabular data where you have a considerable number of rows and where you have to create your own features 🙂 As in, the size of data here is not spectacularly large, but as you have to perform joins, etc, RAM becomes a consideration.</p>\n<p>At the end of the day, I feel that this is the best way to put it: \"With less RAM, you have to compensate through additional work (you need to write more code), more RAM means you can take a simplified approach to not have to deal with things like reading data in chunks\". </p>\n<p>There is a corollary to this and that is if you don't pay any attention to your RAM usage, you will run out of memory no matter how much RAM you have 😄 Which I think makes this competition even more fun!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2082888,
          "author_name": "learnmore1",
          "author_url": "",
          "post_date": "01/02/2023 04:31:30",
          "content": "<p>hi radek, thanks for sharing your thoughts. I wonder if I clearly understand your points. So to the 2 steps: candidates generation and model training. the first step is super ram costing while the second is suitable to proper ram machine?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2045646,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "11/27/2022 14:32:16",
      "content": "<p>i find polars to be much better than pandas regarding this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2053041,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "12/02/2022 18:50:44",
      "content": "<p>For me I am using mostly numba, but 64 gb ram looks kind of enough only, I think you will need more</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2054465,
      "author_name": "evilpsycho42",
      "author_url": "",
      "post_date": "12/04/2022 06:57:27",
      "content": "<p>At least 64GB i think,  otherwise you need add a lot of extra code for chunking processing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2044572": "Hello, I am considering trying out this competition but have been reading a few posts discussing the data size and heavy RAM utilization. Does anyone have a ballpark about how much RAM would one need to comfortably work on this competition?",
    "2044705": "Hi @raphael1123 ! I think that it heavily depends on the strategy you want to adopt to face up the challenge. In [this](https://www.kaggle.com/code/sbunzini/user-item-collaborative-filtering-ensemble) I used Sparse Matrices, a memory-efficient data structure, to run basic Recommender Systems algorithms; the same approach has been used [here](https://www.kaggle.com/code/cocoshe/itemcf-with-item-item-similarity-matrix).\n\nThe most common approach so far is the creation of co-visitation matrices, like in [1](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575/data), [2](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic) and similar. However, the starting point is almost always this [parquet format](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) from @radek1",
    "2044865": "Hello @raphael1123, I have been trying o process the data, and takes forever, I haven't been able to load everything this point so I can't tell but there is also already some memory efficientdatatsets available you can use.",
    "2045018": "A lot depends on the tools that you are using. I am finding with `polars` I can get away with much less RAM. With 64GB of RAM that I have on my local machine, I don't think I would have a great time using `pandas`. \n\nI wrote about more on the RAM situation here: [💡 How to deal with this competition needing so much RAM -- a couple of things that worked for me](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368170), maybe this can shed some additional light for you.\n\nIn essence, I don't think this competition is much different from any other on tabular data where you have a considerable number of rows and where you have to create your own features 🙂 As in, the size of data here is not spectacularly large, but as you have to perform joins, etc, RAM becomes a consideration.\n\nAt the end of the day, I feel that this is the best way to put it: \"With less RAM, you have to compensate through additional work (you need to write more code), more RAM means you can take a simplified approach to not have to deal with things like reading data in chunks\". \n\nThere is a corollary to this and that is if you don't pay any attention to your RAM usage, you will run out of memory no matter how much RAM you have 😄 Which I think makes this competition even more fun!",
    "2045646": "i find polars to be much better than pandas regarding this!",
    "2053041": "For me I am using mostly numba, but 64 gb ram looks kind of enough only, I think you will need more",
    "2054465": "At least 64GB i think,  otherwise you need add a lot of extra code for chunking processing.",
    "2082888": "hi radek, thanks for sharing your thoughts. I wonder if I clearly understand your points. So to the 2 steps: candidates generation and model training. the first step is super ram costing while the second is suitable to proper ram machine?"
  },
  "source": "meta"
}