{
  "id": 475492,
  "title": "Your notebook tried to allocate more memory than is available",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475492",
  "author_name": "",
  "post_date": "2024-02-08T17:13:58.554513500Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>encountering error when trying to merge every parquet files….</p>",
  "messages": [
    {
      "id": "2643179",
      "postDate": "02/08/2024 17:13:58",
      "content": "<p>encountering error when trying to merge every parquet files….</p>",
      "rawMarkdown": "encountering error when trying to merge every parquet files....",
      "votes": null
    },
    {
      "id": "2643274",
      "postDate": "02/08/2024 18:24:33",
      "content": "<p>Kaggle kernels have 30gb ram. If you don't merge those big parquet files efficiently it ll use all memory and restart the notebook.  </p>",
      "rawMarkdown": "Kaggle kernels have 30gb ram. If you don't merge those big parquet files efficiently it ll use all memory and restart the notebook.",
      "votes": null
    },
    {
      "id": "2643682",
      "postDate": "02/09/2024 03:23:36",
      "content": "<p>used polars and reduced mem but did not work.. any suggestions?</p>",
      "rawMarkdown": "used polars and reduced mem but did not work.. any suggestions?",
      "votes": null
    },
    {
      "id": "2643753",
      "postDate": "02/09/2024 04:28:40",
      "content": "<p>Was there some many to many join maybe?</p>",
      "rawMarkdown": "Was there some many to many join maybe?",
      "votes": null
    },
    {
      "id": "2643935",
      "postDate": "02/09/2024 07:23:37",
      "content": "<p>Use Lazy evaluation.</p>",
      "rawMarkdown": "Use Lazy evaluation.",
      "votes": null
    },
    {
      "id": "2643960",
      "postDate": "02/09/2024 07:38:21",
      "content": "<p>\"credit_bureau_a_1_*\" and \"credit_bureau_a_2_*\" files takes so much space alone. You can perform aggregation on these files one by one and then concatenate them</p>",
      "rawMarkdown": "\"credit_bureau_a_1_\\*\" and \"credit_bureau_a_2_\\*\" files takes so much space alone. You can perform aggregation on these files one by one and then concatenate them",
      "votes": null
    },
    {
      "id": "2645772",
      "postDate": "02/10/2024 13:05:22",
      "content": "<p><a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a> Are you aggregating for <code>depth &gt; 0</code>? </p>\n<blockquote>\n  <p>However, for tables with depth&gt;0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature.</p>\n</blockquote>\n<p>If not, the data won't fit into the memory available for kaggle notebooks.</p>",
      "rawMarkdown": "beckpro Are you aggregating for `depth > 0`? \n\n>However, for tables with depth>0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature.\n\nIf not, the data won't fit into the memory available for kaggle notebooks.",
      "votes": null
    },
    {
      "id": "2647113",
      "postDate": "02/11/2024 11:16:44",
      "content": "<p>I suggest you process the data to create features in batches and feed it by batches to a pretrained model.</p>",
      "rawMarkdown": "I suggest you process the data to create features in batches and feed it by batches to a pretrained model.",
      "votes": null
    },
    {
      "id": "2647202",
      "postDate": "02/11/2024 12:09:43",
      "content": "<p>thx for ur baseline and also thx for ur discussion. </p>",
      "rawMarkdown": "thx for ur baseline and also thx for ur discussion.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2643274,
      "author_name": "kishanvavdara",
      "author_url": "",
      "post_date": "02/08/2024 18:24:33",
      "content": "<p>Kaggle kernels have 30gb ram. If you don't merge those big parquet files efficiently it ll use all memory and restart the notebook.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2643682,
          "author_name": "beckpro",
          "author_url": "",
          "post_date": "02/09/2024 03:23:36",
          "content": "<p>used polars and reduced mem but did not work.. any suggestions?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2643753,
              "author_name": "thomasmeiner",
              "author_url": "",
              "post_date": "02/09/2024 04:28:40",
              "content": "<p>Was there some many to many join maybe?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2643935,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/09/2024 07:23:37",
              "content": "<p>Use Lazy evaluation.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2643960,
      "author_name": "greysky",
      "author_url": "",
      "post_date": "02/09/2024 07:38:21",
      "content": "<p>\"credit_bureau_a_1_*\" and \"credit_bureau_a_2_*\" files takes so much space alone. You can perform aggregation on these files one by one and then concatenate them</p>",
      "votes": null,
      "replies": [
        {
          "id": 2647202,
          "author_name": "beckpro",
          "author_url": "",
          "post_date": "02/11/2024 12:09:43",
          "content": "<p>thx for ur baseline and also thx for ur discussion. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2645772,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "02/10/2024 13:05:22",
      "content": "<p><a href=\"https://www.kaggle.com/beckpro\" target=\"_blank\">@beckpro</a> Are you aggregating for <code>depth &gt; 0</code>? </p>\n<blockquote>\n  <p>However, for tables with depth&gt;0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature.</p>\n</blockquote>\n<p>If not, the data won't fit into the memory available for kaggle notebooks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2647113,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/11/2024 11:16:44",
      "content": "<p>I suggest you process the data to create features in batches and feed it by batches to a pretrained model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2643179": "encountering error when trying to merge every parquet files....",
    "2643274": "Kaggle kernels have 30gb ram. If you don't merge those big parquet files efficiently it ll use all memory and restart the notebook.",
    "2643682": "used polars and reduced mem but did not work.. any suggestions?",
    "2643753": "Was there some many to many join maybe?",
    "2643935": "Use Lazy evaluation.",
    "2643960": "\"credit_bureau_a_1_\\*\" and \"credit_bureau_a_2_\\*\" files takes so much space alone. You can perform aggregation on these files one by one and then concatenate them",
    "2645772": "beckpro Are you aggregating for `depth > 0`? \n\n>However, for tables with depth>0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature.\n\nIf not, the data won't fit into the memory available for kaggle notebooks.",
    "2647113": "I suggest you process the data to create features in batches and feed it by batches to a pretrained model.",
    "2647202": "thx for ur baseline and also thx for ur discussion."
  },
  "source": "meta"
}