{
  "id": 134925,
  "title": "Your notebook tried to allocate more memory than is available. It has restarted.",
  "url": "/competitions/bengaliai-cv19/discussion/134925",
  "author_name": "",
  "post_date": "2020-03-11T06:07:11.007742500Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I keep getting this error during model.fit_generator()</p>\n\n<p><code>Your notebook tried to allocate more memory than is available. It has restarted.</code></p>\n\n<p>I am training the four parquet file within a for loop.... Training on one parquet file at one time.</p>\n\n<p>I am not doing any image preprocessing as of now.</p>\n\n<p>Can someone take a look at the <a href=\"https://www.kaggle.com/akrsrivastava/bengaliai\">notebook</a> and let me know what I am doing wrong.</p>",
  "messages": [
    {
      "id": "768735",
      "postDate": "03/11/2020 06:07:11",
      "content": "<p>I keep getting this error during model.fit_generator()</p>\n\n<p><code>Your notebook tried to allocate more memory than is available. It has restarted.</code></p>\n\n<p>I am training the four parquet file within a for loop.... Training on one parquet file at one time.</p>\n\n<p>I am not doing any image preprocessing as of now.</p>\n\n<p>Can someone take a look at the <a href=\"https://www.kaggle.com/akrsrivastava/bengaliai\">notebook</a> and let me know what I am doing wrong.</p>",
      "rawMarkdown": "I keep getting this error during model.fit_generator()\n\n`Your notebook tried to allocate more memory than is available. It has restarted.`\n\nI am training the four parquet file within a for loop.... Training on one parquet file at one time.\n\nI am not doing any image preprocessing as of now.\n\nCan someone take a look at the [notebook](https://www.kaggle.com/akrsrivastava/bengaliai) and let me know what I am doing wrong.",
      "votes": null
    },
    {
      "id": "769069",
      "postDate": "03/11/2020 13:58:37",
      "content": "<p><a href=\"/akrsrivastava\">@akrsrivastava</a> in the notebook you shared, you are trying to load all 4 parquet files into memory first and then using a batch data generator. Load 1 parquet file into memory at a time and train as you cannot read all 4 at once into memory. The notebook will crash even before you start training.</p>",
      "rawMarkdown": "akrsrivastava in the notebook you shared, you are trying to load all 4 parquet files into memory first and then using a batch data generator. Load 1 parquet file into memory at a time and train as you cannot read all 4 at once into memory. The notebook will crash even before you start training.",
      "votes": null
    },
    {
      "id": "769678",
      "postDate": "03/12/2020 06:23:21",
      "content": "<p>I am using the following\n`</p>\n\n<p><code>for train_file_idx in range(4):</code></p>\n\n<pre><code>print (f\"########## Training File: {train_file_idx} \\n\")\n\n\ntrain_image_df = pd.read_parquet(f\"/kaggle/input/bengaliai-cv19/train_image_data_{train_file_idx}.parquet\" ) \n</code></pre>\n\n<p>`\ntrain_image_df is just one parquet file.\nAnd then I use this to create train test splits . Even in the data generator, I am passing the train / test generated from thistrain_image_df. I am not sure where I am loading all parquet files</p>",
      "rawMarkdown": "I am using the following\n`\n\n`for train_file_idx in range(4): `\n\n\n    print (f\"########## Training File: {train_file_idx} \\n\")\n\n\n    train_image_df = pd.read_parquet(f\"/kaggle/input/bengaliai-cv19/train_image_data_{train_file_idx}.parquet\" ) \n`\ntrain\\_image\\_df is just one parquet file.\nAnd then I use this to create train test splits . Even in the data generator, I am passing the train / test generated from thistrain\\_image\\_df. I am not sure where I am loading all parquet files",
      "votes": null
    },
    {
      "id": "769750",
      "postDate": "03/12/2020 08:12:13",
      "content": "<p>In this loop which you are using in the beginning of your notebook, you are reading the 4 train parquet files into 1 pandas dataframe called <strong>train-image-df</strong>.  That is what i meant by you cannot load all 4 parquet files into memory all at once. When you read 1 train parquet file it takes roughly around 5GB.</p>\n\n<p>It does not matter if you are using a data generator later on in the notebook since you won't have enough memory to first fit the 4 parquet files into <strong>train-image-df</strong>. Hope this explanation helps.</p>",
      "rawMarkdown": "In this loop which you are using in the beginning of your notebook, you are reading the 4 train parquet files into 1 pandas dataframe called **train-image-df**.  That is what i meant by you cannot load all 4 parquet files into memory all at once. When you read 1 train parquet file it takes roughly around 5GB.\n\nIt does not matter if you are using a data generator later on in the notebook since you won't have enough memory to first fit the 4 parquet files into **train-image-df**. Hope this explanation helps.",
      "votes": null
    },
    {
      "id": "829564",
      "postDate": "05/01/2020 22:46:35",
      "content": "<p>I tried to save without making \"run all\" and I signed \" commit and save \" while ı am saving. It worked.</p>",
      "rawMarkdown": "I tried to save without making \"run all\" and I signed \" commit and save \" while ı am saving. It worked.",
      "votes": null
    },
    {
      "id": "875878",
      "postDate": "06/06/2020 08:25:09",
      "content": "<p>I got the <a href=\"https://www.kaggle.com/product-feedback/155897\">same issue</a> and have reported it under Product Feedback but so far no reply.</p>",
      "rawMarkdown": "I got the [same issue](https://www.kaggle.com/product-feedback/155897) and have reported it under Product Feedback but so far no reply.",
      "votes": null
    },
    {
      "id": "1486004",
      "postDate": "08/22/2021 15:19:33",
      "content": "<p>I see this is old post. <br>\nI loaded all parquet  file into single dataframe and deleted the dataframe with the command <strong>\"del df\"</strong>. Still I get the same error. Any suggestions?</p>",
      "rawMarkdown": "I see this is old post. \nI loaded all parquet  file into single dataframe and deleted the dataframe with the command **\"del df\"**. Still I get the same error. Any suggestions?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1486004,
      "author_name": "venkataratnamb",
      "author_url": "",
      "post_date": "08/22/2021 15:19:33",
      "content": "<p>I see this is old post. <br>\nI loaded all parquet  file into single dataframe and deleted the dataframe with the command <strong>\"del df\"</strong>. Still I get the same error. Any suggestions?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 769069,
      "author_name": "yovinyahathugoda",
      "author_url": "",
      "post_date": "03/11/2020 13:58:37",
      "content": "<p><a href=\"/akrsrivastava\">@akrsrivastava</a> in the notebook you shared, you are trying to load all 4 parquet files into memory first and then using a batch data generator. Load 1 parquet file into memory at a time and train as you cannot read all 4 at once into memory. The notebook will crash even before you start training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 769678,
          "author_name": "akrsrivastava",
          "author_url": "",
          "post_date": "03/12/2020 06:23:21",
          "content": "<p>I am using the following\n`</p>\n\n<p><code>for train_file_idx in range(4):</code></p>\n\n<pre><code>print (f\"########## Training File: {train_file_idx} \\n\")\n\n\ntrain_image_df = pd.read_parquet(f\"/kaggle/input/bengaliai-cv19/train_image_data_{train_file_idx}.parquet\" ) \n</code></pre>\n\n<p>`\ntrain_image_df is just one parquet file.\nAnd then I use this to create train test splits . Even in the data generator, I am passing the train / test generated from thistrain_image_df. I am not sure where I am loading all parquet files</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 769750,
          "author_name": "yovinyahathugoda",
          "author_url": "",
          "post_date": "03/12/2020 08:12:13",
          "content": "<p>In this loop which you are using in the beginning of your notebook, you are reading the 4 train parquet files into 1 pandas dataframe called <strong>train-image-df</strong>.  That is what i meant by you cannot load all 4 parquet files into memory all at once. When you read 1 train parquet file it takes roughly around 5GB.</p>\n\n<p>It does not matter if you are using a data generator later on in the notebook since you won't have enough memory to first fit the 4 parquet files into <strong>train-image-df</strong>. Hope this explanation helps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 829564,
      "author_name": "ftmozc",
      "author_url": "",
      "post_date": "05/01/2020 22:46:35",
      "content": "<p>I tried to save without making \"run all\" and I signed \" commit and save \" while ı am saving. It worked.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 875878,
      "author_name": "neomatrix369",
      "author_url": "",
      "post_date": "06/06/2020 08:25:09",
      "content": "<p>I got the <a href=\"https://www.kaggle.com/product-feedback/155897\">same issue</a> and have reported it under Product Feedback but so far no reply.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "768735": "I keep getting this error during model.fit_generator()\n\n`Your notebook tried to allocate more memory than is available. It has restarted.`\n\nI am training the four parquet file within a for loop.... Training on one parquet file at one time.\n\nI am not doing any image preprocessing as of now.\n\nCan someone take a look at the [notebook](https://www.kaggle.com/akrsrivastava/bengaliai) and let me know what I am doing wrong.",
    "769069": "akrsrivastava in the notebook you shared, you are trying to load all 4 parquet files into memory first and then using a batch data generator. Load 1 parquet file into memory at a time and train as you cannot read all 4 at once into memory. The notebook will crash even before you start training.",
    "769678": "I am using the following\n`\n\n`for train_file_idx in range(4): `\n\n\n    print (f\"########## Training File: {train_file_idx} \\n\")\n\n\n    train_image_df = pd.read_parquet(f\"/kaggle/input/bengaliai-cv19/train_image_data_{train_file_idx}.parquet\" ) \n`\ntrain\\_image\\_df is just one parquet file.\nAnd then I use this to create train test splits . Even in the data generator, I am passing the train / test generated from thistrain\\_image\\_df. I am not sure where I am loading all parquet files",
    "769750": "In this loop which you are using in the beginning of your notebook, you are reading the 4 train parquet files into 1 pandas dataframe called **train-image-df**.  That is what i meant by you cannot load all 4 parquet files into memory all at once. When you read 1 train parquet file it takes roughly around 5GB.\n\nIt does not matter if you are using a data generator later on in the notebook since you won't have enough memory to first fit the 4 parquet files into **train-image-df**. Hope this explanation helps.",
    "829564": "I tried to save without making \"run all\" and I signed \" commit and save \" while ı am saving. It worked.",
    "875878": "I got the [same issue](https://www.kaggle.com/product-feedback/155897) and have reported it under Product Feedback but so far no reply.",
    "1486004": "I see this is old post. \nI loaded all parquet  file into single dataframe and deleted the dataframe with the command **\"del df\"**. Still I get the same error. Any suggestions?"
  },
  "source": "meta"
}