{
  "id": 373583,
  "title": "Not enough memory in Kaggle notebook!",
  "url": "/competitions/otto-recommender-system/discussion/373583",
  "author_name": "",
  "post_date": "2022-12-22T08:19:55.944123500Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm working on this competition using Kaggle notebbok.<br>\nData saved in parquet format is converted DataFrame and use it.<br>\nHowever, when trying to add a new columns, the session is cut off due to lack of memory.</p>\n<p>How do you use your data?<br>\nIs the only way to split the data into several pieces?</p>\n<p>I only want it to run on the Kaggle notebook.<br>\nPlease give me some good way!</p>",
  "messages": [
    {
      "id": "2072563",
      "postDate": "12/22/2022 08:19:55",
      "content": "<p>I'm working on this competition using Kaggle notebbok.<br>\nData saved in parquet format is converted DataFrame and use it.<br>\nHowever, when trying to add a new columns, the session is cut off due to lack of memory.</p>\n<p>How do you use your data?<br>\nIs the only way to split the data into several pieces?</p>\n<p>I only want it to run on the Kaggle notebook.<br>\nPlease give me some good way!</p>",
      "rawMarkdown": "I'm working on this competition using Kaggle notebbok.\nData saved in parquet format is converted DataFrame and use it.\nHowever, when trying to add a new columns, the session is cut off due to lack of memory.\n\nHow do you use your data?\nIs the only way to split the data into several pieces?\n\nI only want it to run on the Kaggle notebook.\nPlease give me some good way!",
      "votes": null
    },
    {
      "id": "2072565",
      "postDate": "12/22/2022 08:21:01",
      "content": "<p>If you don't want to split you can try to use polars.</p>",
      "rawMarkdown": "If you don't want to split you can try to use polars.",
      "votes": null
    },
    {
      "id": "2072579",
      "postDate": "12/22/2022 08:34:59",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> !<br>\nUpon investigation, I learned that polars is faster than pandas.<br>\nCan polars also use less memory??</p>",
      "rawMarkdown": "Thank you @simonveitner !\nUpon investigation, I learned that polars is faster than pandas.\nCan polars also use less memory??",
      "votes": null
    },
    {
      "id": "2072898",
      "postDate": "12/22/2022 13:59:58",
      "content": "<p>Consider working in chunks. Break the dataframe into parts where all of one customer is contained within the same part. Then work on a part and write result to disk. Then clear memory and work on next part.</p>\n<p>Also reduce each column to least data usage. For example convert <code>ts</code> column into seconds with <code>train['ts'] = (train['ts']//1000).astype('int32')</code>. Make sure each of your columns has least <code>dtype</code> possible.</p>\n<p>More tips are <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">here</a> in section titled Memory Management.</p>",
      "rawMarkdown": "Consider working in chunks. Break the dataframe into parts where all of one customer is contained within the same part. Then work on a part and write result to disk. Then clear memory and work on next part.\n\nAlso reduce each column to least data usage. For example convert `ts` column into seconds with `train['ts'] = (train['ts']//1000).astype('int32')`. Make sure each of your columns has least `dtype` possible.\n\nMore tips are [here][1] in section titled Memory Management.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210",
      "votes": null
    },
    {
      "id": "2072909",
      "postDate": "12/22/2022 14:08:38",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !<br>\nI'll keep an eye on your notebook and try to work on splitting up the data.</p>\n<p>I'll read your memory management article!</p>\n<p>You're very helpful, thank you so much!!!!!</p>",
      "rawMarkdown": "Thank you, @cdeotte !\nI'll keep an eye on your notebook and try to work on splitting up the data.\n\nI'll read your memory management article!\n\nYou're very helpful, thank you so much!!!!!",
      "votes": null
    },
    {
      "id": "2072926",
      "postDate": "12/22/2022 14:25:09",
      "content": "<p>Yes. But I suggest to use cudf as it's fester.</p>",
      "rawMarkdown": "Yes. But I suggest to use cudf as it's fester.",
      "votes": null
    },
    {
      "id": "2072966",
      "postDate": "12/22/2022 14:52:00",
      "content": "<p>Thank you!<br>\nI'll examine about cudf!</p>",
      "rawMarkdown": "Thank you!\nI'll examine about cudf!",
      "votes": null
    },
    {
      "id": "2079064",
      "postDate": "12/28/2022 23:35:28",
      "content": "<p>Memory optimization as explained by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">in a notebook from another competition</a> can go a really long way!</p>\n<p>It is a trick every Kaggler should have under their toolbelt, IMO 🙂</p>",
      "rawMarkdown": "Memory optimization as explained by @cdeotte [in a notebook from another competition](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635) can go a really long way!\n\nIt is a trick every Kaggler should have under their toolbelt, IMO 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2072565,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "12/22/2022 08:21:01",
      "content": "<p>If you don't want to split you can try to use polars.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2072579,
          "author_name": "takuma0306",
          "author_url": "",
          "post_date": "12/22/2022 08:34:59",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a> !<br>\nUpon investigation, I learned that polars is faster than pandas.<br>\nCan polars also use less memory??</p>",
          "votes": null,
          "replies": [
            {
              "id": 2072926,
              "author_name": "simonveitner",
              "author_url": "",
              "post_date": "12/22/2022 14:25:09",
              "content": "<p>Yes. But I suggest to use cudf as it's fester.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2072966,
                  "author_name": "takuma0306",
                  "author_url": "",
                  "post_date": "12/22/2022 14:52:00",
                  "content": "<p>Thank you!<br>\nI'll examine about cudf!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2072898,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/22/2022 13:59:58",
      "content": "<p>Consider working in chunks. Break the dataframe into parts where all of one customer is contained within the same part. Then work on a part and write result to disk. Then clear memory and work on next part.</p>\n<p>Also reduce each column to least data usage. For example convert <code>ts</code> column into seconds with <code>train['ts'] = (train['ts']//1000).astype('int32')</code>. Make sure each of your columns has least <code>dtype</code> possible.</p>\n<p>More tips are <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">here</a> in section titled Memory Management.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2072909,
          "author_name": "takuma0306",
          "author_url": "",
          "post_date": "12/22/2022 14:08:38",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !<br>\nI'll keep an eye on your notebook and try to work on splitting up the data.</p>\n<p>I'll read your memory management article!</p>\n<p>You're very helpful, thank you so much!!!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2079064,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "12/28/2022 23:35:28",
      "content": "<p>Memory optimization as explained by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">in a notebook from another competition</a> can go a really long way!</p>\n<p>It is a trick every Kaggler should have under their toolbelt, IMO 🙂</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2072563": "I'm working on this competition using Kaggle notebbok.\nData saved in parquet format is converted DataFrame and use it.\nHowever, when trying to add a new columns, the session is cut off due to lack of memory.\n\nHow do you use your data?\nIs the only way to split the data into several pieces?\n\nI only want it to run on the Kaggle notebook.\nPlease give me some good way!",
    "2072565": "If you don't want to split you can try to use polars.",
    "2072579": "Thank you @simonveitner !\nUpon investigation, I learned that polars is faster than pandas.\nCan polars also use less memory??",
    "2072898": "Consider working in chunks. Break the dataframe into parts where all of one customer is contained within the same part. Then work on a part and write result to disk. Then clear memory and work on next part.\n\nAlso reduce each column to least data usage. For example convert `ts` column into seconds with `train['ts'] = (train['ts']//1000).astype('int32')`. Make sure each of your columns has least `dtype` possible.\n\nMore tips are [here][1] in section titled Memory Management.\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210",
    "2072909": "Thank you, @cdeotte !\nI'll keep an eye on your notebook and try to work on splitting up the data.\n\nI'll read your memory management article!\n\nYou're very helpful, thank you so much!!!!!",
    "2072926": "Yes. But I suggest to use cudf as it's fester.",
    "2072966": "Thank you!\nI'll examine about cudf!",
    "2079064": "Memory optimization as explained by @cdeotte [in a notebook from another competition](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635) can go a really long way!\n\nIt is a trick every Kaggler should have under their toolbelt, IMO 🙂"
  },
  "source": "meta"
}