{
  "id": 479671,
  "title": "Seeking Advice for First Kaggle Competition",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/479671",
  "author_name": "",
  "post_date": "2024-02-25T14:16:58.238210200Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>This is the first Kaggle competition I have entered in. I found the problem very exciting and so I have started working on processing the data to then run analysis on the parameters.</p>\n<p>However, I am running into some problems that I would like advice on. I have been trying to coalesce all of the data entries into one DataFrame to easily run my analysis on as many parameters at once as possible. However, I have been running into memory failures and am unsure on how to proceed.</p>\n<p>What has helped thus far is switching from Pandas to Polars, converting data types to more memory-efficient types, calling Python's garbage collection, and deleting all variables after use. However, I am still getting fatal errors after being able to load about half of the data on one DataFrame.</p>\n<p>(For reference, my computer has 32 Gigs of RAM.) I am unsure if, with proper coding practices, it is possible to load this much data into one DataFrame or rather that I am approaching the problem in an inefficient way.</p>\n<p>If someone could provide guidance, it would be much appreciated! :)</p>",
  "messages": [
    {
      "id": "2668117",
      "postDate": "02/25/2024 14:16:58",
      "content": "<p>Hello all,</p>\n<p>This is the first Kaggle competition I have entered in. I found the problem very exciting and so I have started working on processing the data to then run analysis on the parameters.</p>\n<p>However, I am running into some problems that I would like advice on. I have been trying to coalesce all of the data entries into one DataFrame to easily run my analysis on as many parameters at once as possible. However, I have been running into memory failures and am unsure on how to proceed.</p>\n<p>What has helped thus far is switching from Pandas to Polars, converting data types to more memory-efficient types, calling Python's garbage collection, and deleting all variables after use. However, I am still getting fatal errors after being able to load about half of the data on one DataFrame.</p>\n<p>(For reference, my computer has 32 Gigs of RAM.) I am unsure if, with proper coding practices, it is possible to load this much data into one DataFrame or rather that I am approaching the problem in an inefficient way.</p>\n<p>If someone could provide guidance, it would be much appreciated! :)</p>",
      "rawMarkdown": "Hello all,\n\nThis is the first Kaggle competition I have entered in. I found the problem very exciting and so I have started working on processing the data to then run analysis on the parameters.\n\nHowever, I am running into some problems that I would like advice on. I have been trying to coalesce all of the data entries into one DataFrame to easily run my analysis on as many parameters at once as possible. However, I have been running into memory failures and am unsure on how to proceed.\n\nWhat has helped thus far is switching from Pandas to Polars, converting data types to more memory-efficient types, calling Python's garbage collection, and deleting all variables after use. However, I am still getting fatal errors after being able to load about half of the data on one DataFrame.\n\n(For reference, my computer has 32 Gigs of RAM.) I am unsure if, with proper coding practices, it is possible to load this much data into one DataFrame or rather that I am approaching the problem in an inefficient way.\n\n If someone could provide guidance, it would be much appreciated! :)",
      "votes": null
    },
    {
      "id": "2668150",
      "postDate": "02/25/2024 14:35:01",
      "content": "<p>You can use chunking function. Instead of loading the entire dataset into memory at once, consider loading it in smaller chunks or batches using Pandas' read_csv() function with the chunksize. I hope it will help you! Good luck!</p>",
      "rawMarkdown": "You can use chunking function. Instead of loading the entire dataset into memory at once, consider loading it in smaller chunks or batches using Pandas' read_csv() function with the chunksize. I hope it will help you! Good luck!",
      "votes": null
    },
    {
      "id": "2668762",
      "postDate": "02/25/2024 21:43:23",
      "content": "<p>I've tried Polars rechunk but it doesn't seem to create any improvement.</p>",
      "rawMarkdown": "I've tried Polars rechunk but it doesn't seem to create any improvement.",
      "votes": null
    },
    {
      "id": "2668784",
      "postDate": "02/25/2024 22:25:13",
      "content": "<p>try loading files one at a time or a small group of files at a time, after processing them delete them to free the space.</p>",
      "rawMarkdown": "try loading files one at a time or a small group of files at a time, after processing them delete them to free the space.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2668150,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "02/25/2024 14:35:01",
      "content": "<p>You can use chunking function. Instead of loading the entire dataset into memory at once, consider loading it in smaller chunks or batches using Pandas' read_csv() function with the chunksize. I hope it will help you! Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2668762,
          "author_name": "timothycoffman1",
          "author_url": "",
          "post_date": "02/25/2024 21:43:23",
          "content": "<p>I've tried Polars rechunk but it doesn't seem to create any improvement.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2668784,
              "author_name": "shreyas9181",
              "author_url": "",
              "post_date": "02/25/2024 22:25:13",
              "content": "<p>try loading files one at a time or a small group of files at a time, after processing them delete them to free the space.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2668117": "Hello all,\n\nThis is the first Kaggle competition I have entered in. I found the problem very exciting and so I have started working on processing the data to then run analysis on the parameters.\n\nHowever, I am running into some problems that I would like advice on. I have been trying to coalesce all of the data entries into one DataFrame to easily run my analysis on as many parameters at once as possible. However, I have been running into memory failures and am unsure on how to proceed.\n\nWhat has helped thus far is switching from Pandas to Polars, converting data types to more memory-efficient types, calling Python's garbage collection, and deleting all variables after use. However, I am still getting fatal errors after being able to load about half of the data on one DataFrame.\n\n(For reference, my computer has 32 Gigs of RAM.) I am unsure if, with proper coding practices, it is possible to load this much data into one DataFrame or rather that I am approaching the problem in an inefficient way.\n\n If someone could provide guidance, it would be much appreciated! :)",
    "2668150": "You can use chunking function. Instead of loading the entire dataset into memory at once, consider loading it in smaller chunks or batches using Pandas' read_csv() function with the chunksize. I hope it will help you! Good luck!",
    "2668762": "I've tried Polars rechunk but it doesn't seem to create any improvement.",
    "2668784": "try loading files one at a time or a small group of files at a time, after processing them delete them to free the space."
  },
  "source": "meta"
}