{
  "id": 406084,
  "title": "Confusing Memory Limit Error",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406084",
  "author_name": "",
  "post_date": "2023-04-30T20:03:56.554719500Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>It'd very interesting, I get memory allocation limit error with the following code without applying the memory reduce function:</p>\n<p>df_train_defog=pd.concat([reader(f) for f in tqdm(train_defog)]).fillna(0)<br>\ndf_train_tdcsfog=pd.concat([reader(f) for f in tqdm(train_tdcsfog)]).fillna(0)<br>\ntrain=pd.concat([df_train_defog,df_train_tdcsfog])</p>\n<p>But not with the following code:<br>\ntrain = pd.concat([reader(f) for f in tqdm(train)]).fillna(0)</p>\n<p>I'm wondering why? Do you have any idea about the data frames concat overheads?</p>",
  "messages": [
    {
      "id": "2240740",
      "postDate": "04/30/2023 20:03:56",
      "content": "<p>It'd very interesting, I get memory allocation limit error with the following code without applying the memory reduce function:</p>\n<p>df_train_defog=pd.concat([reader(f) for f in tqdm(train_defog)]).fillna(0)<br>\ndf_train_tdcsfog=pd.concat([reader(f) for f in tqdm(train_tdcsfog)]).fillna(0)<br>\ntrain=pd.concat([df_train_defog,df_train_tdcsfog])</p>\n<p>But not with the following code:<br>\ntrain = pd.concat([reader(f) for f in tqdm(train)]).fillna(0)</p>\n<p>I'm wondering why? Do you have any idea about the data frames concat overheads?</p>",
      "rawMarkdown": "It'd very interesting, I get memory allocation limit error with the following code without applying the memory reduce function:\n\ndf_train_defog=pd.concat([reader(f) for f in tqdm(train_defog)]).fillna(0)\ndf_train_tdcsfog=pd.concat([reader(f) for f in tqdm(train_tdcsfog)]).fillna(0)\ntrain=pd.concat([df_train_defog,df_train_tdcsfog])\n\nBut not with the following code:\ntrain = pd.concat([reader(f) for f in tqdm(train)]).fillna(0)\n\nI'm wondering why? Do you have any idea about the data frames concat overheads?",
      "votes": null
    },
    {
      "id": "2241215",
      "postDate": "05/01/2023 09:57:16",
      "content": "<p>The reason you are experiencing a memory allocation error with the first code snippet is that you are concatenating two large data frames (df_train_defog and df_train_tdcsfog) and then concatenating the result with another data frame (train). The concatenation process creates a new data frame that is the concatenation of the input data frames. This means that you are creating and holding in memory three large data frames at the same time.</p>\n<p>The second code snippet avoids this issue by concatenating each data frame (reader(f)) directly into the final train data frame. This means that only one large data frame is being held in memory at any given time. This reduces the memory overhead of the concatenation process.</p>",
      "rawMarkdown": "The reason you are experiencing a memory allocation error with the first code snippet is that you are concatenating two large data frames (df_train_defog and df_train_tdcsfog) and then concatenating the result with another data frame (train). The concatenation process creates a new data frame that is the concatenation of the input data frames. This means that you are creating and holding in memory three large data frames at the same time.\n\nThe second code snippet avoids this issue by concatenating each data frame (reader(f)) directly into the final train data frame. This means that only one large data frame is being held in memory at any given time. This reduces the memory overhead of the concatenation process.",
      "votes": null
    },
    {
      "id": "2241232",
      "postDate": "05/01/2023 10:16:24",
      "content": "<p>Memory allocation errors may occur in the first block of code due to the creation of intermediate objects df_train_defog and df_train_tdcsfog, which increases the program's memory usage. In the second block of code, the data is concatenated directly, which may reduce memory usage and prevent the error. Concatenating data frames can be memory-intensive, so you can optimize the code to reduce memory usage by reading and processing data in batches, using more efficient data structures or algorithms.</p>",
      "rawMarkdown": "Memory allocation errors may occur in the first block of code due to the creation of intermediate objects df_train_defog and df_train_tdcsfog, which increases the program's memory usage. In the second block of code, the data is concatenated directly, which may reduce memory usage and prevent the error. Concatenating data frames can be memory-intensive, so you can optimize the code to reduce memory usage by reading and processing data in batches, using more efficient data structures or algorithms.",
      "votes": null
    },
    {
      "id": "2241752",
      "postDate": "05/01/2023 19:21:36",
      "content": "<p>Oh, thanks, you're right, it's actually at least 2 times larger in the first code</p>",
      "rawMarkdown": "Oh, thanks, you're right, it's actually at least 2 times larger in the first code",
      "votes": null
    },
    {
      "id": "2241754",
      "postDate": "05/01/2023 19:23:40",
      "content": "<p>Thanks, good idea and I realized the first method actually at least 2 times larger (I can delete the intermediate ones but I don't get to finish concating and then deleting the original ones…)</p>",
      "rawMarkdown": "Thanks, good idea and I realized the first method actually at least 2 times larger (I can delete the intermediate ones but I don't get to finish concating and then deleting the original ones...)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2241215,
      "author_name": "dbdmobile",
      "author_url": "",
      "post_date": "05/01/2023 09:57:16",
      "content": "<p>The reason you are experiencing a memory allocation error with the first code snippet is that you are concatenating two large data frames (df_train_defog and df_train_tdcsfog) and then concatenating the result with another data frame (train). The concatenation process creates a new data frame that is the concatenation of the input data frames. This means that you are creating and holding in memory three large data frames at the same time.</p>\n<p>The second code snippet avoids this issue by concatenating each data frame (reader(f)) directly into the final train data frame. This means that only one large data frame is being held in memory at any given time. This reduces the memory overhead of the concatenation process.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2241752,
          "author_name": "zahranikdel",
          "author_url": "",
          "post_date": "05/01/2023 19:21:36",
          "content": "<p>Oh, thanks, you're right, it's actually at least 2 times larger in the first code</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2241232,
      "author_name": "ericka42",
      "author_url": "",
      "post_date": "05/01/2023 10:16:24",
      "content": "<p>Memory allocation errors may occur in the first block of code due to the creation of intermediate objects df_train_defog and df_train_tdcsfog, which increases the program's memory usage. In the second block of code, the data is concatenated directly, which may reduce memory usage and prevent the error. Concatenating data frames can be memory-intensive, so you can optimize the code to reduce memory usage by reading and processing data in batches, using more efficient data structures or algorithms.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2241754,
          "author_name": "zahranikdel",
          "author_url": "",
          "post_date": "05/01/2023 19:23:40",
          "content": "<p>Thanks, good idea and I realized the first method actually at least 2 times larger (I can delete the intermediate ones but I don't get to finish concating and then deleting the original ones…)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2240740": "It'd very interesting, I get memory allocation limit error with the following code without applying the memory reduce function:\n\ndf_train_defog=pd.concat([reader(f) for f in tqdm(train_defog)]).fillna(0)\ndf_train_tdcsfog=pd.concat([reader(f) for f in tqdm(train_tdcsfog)]).fillna(0)\ntrain=pd.concat([df_train_defog,df_train_tdcsfog])\n\nBut not with the following code:\ntrain = pd.concat([reader(f) for f in tqdm(train)]).fillna(0)\n\nI'm wondering why? Do you have any idea about the data frames concat overheads?",
    "2241215": "The reason you are experiencing a memory allocation error with the first code snippet is that you are concatenating two large data frames (df_train_defog and df_train_tdcsfog) and then concatenating the result with another data frame (train). The concatenation process creates a new data frame that is the concatenation of the input data frames. This means that you are creating and holding in memory three large data frames at the same time.\n\nThe second code snippet avoids this issue by concatenating each data frame (reader(f)) directly into the final train data frame. This means that only one large data frame is being held in memory at any given time. This reduces the memory overhead of the concatenation process.",
    "2241232": "Memory allocation errors may occur in the first block of code due to the creation of intermediate objects df_train_defog and df_train_tdcsfog, which increases the program's memory usage. In the second block of code, the data is concatenated directly, which may reduce memory usage and prevent the error. Concatenating data frames can be memory-intensive, so you can optimize the code to reduce memory usage by reading and processing data in batches, using more efficient data structures or algorithms.",
    "2241752": "Oh, thanks, you're right, it's actually at least 2 times larger in the first code",
    "2241754": "Thanks, good idea and I realized the first method actually at least 2 times larger (I can delete the intermediate ones but I don't get to finish concating and then deleting the original ones...)"
  },
  "source": "meta"
}