{
  "id": 374192,
  "title": "Has anyone encountered a MemoryError like this?",
  "url": "/competitions/otto-recommender-system/discussion/374192",
  "author_name": "",
  "post_date": "2022-12-25T23:48:39.659459500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I used <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\"><strong>Chris's</strong></a> baseline but used my own data, removing some of the outliers. But what was unexpected was the appearance of an unexpected MemoryError. Does anybody know why? Because it's been bothering me for a long time.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2F79ec7d9925b527c80896b19313f0b6e6%2F1.png?generation=1672012057586444&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2Ff8226780fa14d0403d7a59ed47b49bd8%2F2.png?generation=1672012095309584&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2075825",
      "postDate": "12/25/2022 23:48:39",
      "content": "<p>I used <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\"><strong>Chris's</strong></a> baseline but used my own data, removing some of the outliers. But what was unexpected was the appearance of an unexpected MemoryError. Does anybody know why? Because it's been bothering me for a long time.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2F79ec7d9925b527c80896b19313f0b6e6%2F1.png?generation=1672012057586444&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2Ff8226780fa14d0403d7a59ed47b49bd8%2F2.png?generation=1672012095309584&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I used [**Chris's**](https://www.kaggle.com/cdeotte) baseline but used my own data, removing some of the outliers. But what was unexpected was the appearance of an unexpected MemoryError. Does anybody know why? Because it's been bothering me for a long time.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2F79ec7d9925b527c80896b19313f0b6e6%2F1.png?generation=1672012057586444&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2Ff8226780fa14d0403d7a59ed47b49bd8%2F2.png?generation=1672012095309584&alt=media)",
      "votes": null
    },
    {
      "id": "2075863",
      "postDate": "12/26/2022 01:19:16",
      "content": "<p>It is because your own data is larger than kaggle's gpu memory. You can try to reduce your dataframe's size or use chunks to process.</p>",
      "rawMarkdown": "It is because your own data is larger than kaggle's gpu memory. You can try to reduce your dataframe's size or use chunks to process.",
      "votes": null
    },
    {
      "id": "2075947",
      "postDate": "12/26/2022 02:50:03",
      "content": "<p>I don't think it's a problem of exhausted memory, because the data I use is smaller than baseline data.😔</p>",
      "rawMarkdown": "I don't think it's a problem of exhausted memory, because the data I use is smaller than baseline data.😔",
      "votes": null
    },
    {
      "id": "2076002",
      "postDate": "12/26/2022 03:58:18",
      "content": "<p>You can try to increase the <code>DISK_PIECES</code> parameter. That will make each process chunk smaller and possibly avoid memory error.</p>\n<p>Also pay careful attention to the size of each session. In my notebook, i use tail 30 with</p>\n<pre><code>        df = df.reset_index(drop=True)\n        df['n'] = df.groupby('session').cumcount()\n        df = df.loc[df.n&lt;30].drop('n',axis=1)\n</code></pre>\n<p>This means when we <code>df.merge(df)</code> then each session grows to maximum of <code>30^2 = 900</code> rows. If you don't use tail and one user has 300 items, then <code>df.merge(df)</code> can create 90,000 new rows! After many users like this, it will easily grow to be too big.</p>\n<p>So keeping tail small or doing stuff like <code>df.merge(df.loc[df['type'].isin([1,2]))</code> and other tricks can make sure that the merge doesn't get too large for each user.</p>",
      "rawMarkdown": "You can try to increase the `DISK_PIECES` parameter. That will make each process chunk smaller and possibly avoid memory error.\n\nAlso pay careful attention to the size of each session. In my notebook, i use tail 30 with\n\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n\nThis means when we `df.merge(df)` then each session grows to maximum of `30^2 = 900` rows. If you don't use tail and one user has 300 items, then `df.merge(df)` can create 90,000 new rows! After many users like this, it will easily grow to be too big.\n\nSo keeping tail small or doing stuff like `df.merge(df.loc[df['type'].isin([1,2]))` and other tricks can make sure that the merge doesn't get too large for each user.",
      "votes": null
    },
    {
      "id": "2076398",
      "postDate": "12/26/2022 13:09:48",
      "content": "<p>Thank you for teaching😛</p>",
      "rawMarkdown": "Thank you for teaching😛",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2075863,
      "author_name": "carnozhao",
      "author_url": "",
      "post_date": "12/26/2022 01:19:16",
      "content": "<p>It is because your own data is larger than kaggle's gpu memory. You can try to reduce your dataframe's size or use chunks to process.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2075947,
          "author_name": "jasonhuangcn",
          "author_url": "",
          "post_date": "12/26/2022 02:50:03",
          "content": "<p>I don't think it's a problem of exhausted memory, because the data I use is smaller than baseline data.😔</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2076002,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/26/2022 03:58:18",
      "content": "<p>You can try to increase the <code>DISK_PIECES</code> parameter. That will make each process chunk smaller and possibly avoid memory error.</p>\n<p>Also pay careful attention to the size of each session. In my notebook, i use tail 30 with</p>\n<pre><code>        df = df.reset_index(drop=True)\n        df['n'] = df.groupby('session').cumcount()\n        df = df.loc[df.n&lt;30].drop('n',axis=1)\n</code></pre>\n<p>This means when we <code>df.merge(df)</code> then each session grows to maximum of <code>30^2 = 900</code> rows. If you don't use tail and one user has 300 items, then <code>df.merge(df)</code> can create 90,000 new rows! After many users like this, it will easily grow to be too big.</p>\n<p>So keeping tail small or doing stuff like <code>df.merge(df.loc[df['type'].isin([1,2]))</code> and other tricks can make sure that the merge doesn't get too large for each user.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2076398,
          "author_name": "jasonhuangcn",
          "author_url": "",
          "post_date": "12/26/2022 13:09:48",
          "content": "<p>Thank you for teaching😛</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2075825": "I used [**Chris's**](https://www.kaggle.com/cdeotte) baseline but used my own data, removing some of the outliers. But what was unexpected was the appearance of an unexpected MemoryError. Does anybody know why? Because it's been bothering me for a long time.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2F79ec7d9925b527c80896b19313f0b6e6%2F1.png?generation=1672012057586444&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5556447%2Ff8226780fa14d0403d7a59ed47b49bd8%2F2.png?generation=1672012095309584&alt=media)",
    "2075863": "It is because your own data is larger than kaggle's gpu memory. You can try to reduce your dataframe's size or use chunks to process.",
    "2075947": "I don't think it's a problem of exhausted memory, because the data I use is smaller than baseline data.😔",
    "2076002": "You can try to increase the `DISK_PIECES` parameter. That will make each process chunk smaller and possibly avoid memory error.\n\nAlso pay careful attention to the size of each session. In my notebook, i use tail 30 with\n\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n\nThis means when we `df.merge(df)` then each session grows to maximum of `30^2 = 900` rows. If you don't use tail and one user has 300 items, then `df.merge(df)` can create 90,000 new rows! After many users like this, it will easily grow to be too big.\n\nSo keeping tail small or doing stuff like `df.merge(df.loc[df['type'].isin([1,2]))` and other tricks can make sure that the merge doesn't get too large for each user.",
    "2076398": "Thank you for teaching😛"
  },
  "source": "meta"
}