{
  "id": 330754,
  "title": "Your notebook tried to allocate more memory than is available. It has restarted Error.",
  "url": "/competitions/amex-default-prediction/discussion/330754",
  "author_name": "",
  "post_date": "2022-06-14T05:45:13.878255900Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>My model is ready but i handled it with nrows restriction. When i need to upload to a variable my test data i would face with this \"Your notebook tried to allocate more memory than is available. It has restarted\" error. </p>\n<p>Is there anyonce experienced this?</p>\n<p>Thanks for support.</p>",
  "messages": [
    {
      "id": "1819762",
      "postDate": "06/14/2022 05:45:13",
      "content": "<p>My model is ready but i handled it with nrows restriction. When i need to upload to a variable my test data i would face with this \"Your notebook tried to allocate more memory than is available. It has restarted\" error. </p>\n<p>Is there anyonce experienced this?</p>\n<p>Thanks for support.</p>",
      "rawMarkdown": "My model is ready but i handled it with nrows restriction. When i need to upload to a variable my test data i would face with this \"Your notebook tried to allocate more memory than is available. It has restarted\" error. \n\nIs there anyonce experienced this?\n\nThanks for support.",
      "votes": null
    },
    {
      "id": "1819792",
      "postDate": "06/14/2022 06:16:48",
      "content": "<p>Yes, this is caused by too much memory in the dataset you train or too much memory in the dataset you infer. The solution is to optimize memory or code it to catch bugs.</p>",
      "rawMarkdown": "Yes, this is caused by too much memory in the dataset you train or too much memory in the dataset you infer. The solution is to optimize memory or code it to catch bugs.",
      "votes": null
    },
    {
      "id": "1819796",
      "postDate": "06/14/2022 06:20:39",
      "content": "<p><a href=\"https://www.kaggle.com/leewook\" target=\"_blank\">@leewook</a> firstly thanks,</p>\n<p>You have any script related this subject.</p>",
      "rawMarkdown": "leewook firstly thanks,\n\nYou have any script related this subject.",
      "votes": null
    },
    {
      "id": "1819871",
      "postDate": "06/14/2022 07:50:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/baycelik\" target=\"_blank\">@baycelik</a> This <a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-2-inference?scriptVersionId=96644290\" target=\"_blank\">old version of my Keras notebook</a> shows how to process the test data in chunks so that you never have the whole test data in memory at the same time.</p>",
      "rawMarkdown": "Hi @baycelik This [old version of my Keras notebook](https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-2-inference?scriptVersionId=96644290) shows how to process the test data in chunks so that you never have the whole test data in memory at the same time.",
      "votes": null
    },
    {
      "id": "1819878",
      "postDate": "06/14/2022 07:56:42",
      "content": "<p><a href=\"https://www.kaggle.com/am140307\" target=\"_blank\">@am140307</a> thank you</p>",
      "rawMarkdown": "am140307 thank you",
      "votes": null
    },
    {
      "id": "1820656",
      "postDate": "06/14/2022 20:14:59",
      "content": "<p>There are many post about how to work with this amount of data, using parket or feather, changing columns types to optimize space, etc.</p>\n<p>Also this post <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/330347\" target=\"_blank\"> A List of EDA tricks when memory is limited</a> have some nice info about this topic.</p>\n<p>For me deleting Arrays or Dataframe and spliting the data in files have solved the issue.</p>",
      "rawMarkdown": "There are many post about how to work with this amount of data, using parket or feather, changing columns types to optimize space, etc.\n\nAlso this post [ A List of EDA tricks when memory is limited](https://www.kaggle.com/competitions/amex-default-prediction/discussion/330347) have some nice info about this topic.\n\nFor me deleting Arrays or Dataframe and spliting the data in files have solved the issue.",
      "votes": null
    },
    {
      "id": "1820872",
      "postDate": "06/15/2022 04:28:50",
      "content": "<p><a href=\"https://www.kaggle.com/cucasf\" target=\"_blank\">@cucasf</a> thanks for support</p>",
      "rawMarkdown": "cucasf thanks for support",
      "votes": null
    },
    {
      "id": "1821811",
      "postDate": "06/15/2022 21:08:17",
      "content": "<p>You cant fit the original dataset into the kaggle notebook which supports only 16GB, but there are other things you can do:</p>\n<h3>Use size reduced dataset</h3>\n<p>You could use some of the reduced in size datasets published by the community (with different formats like parquet, feather, …) e.g.:</p>\n<p><a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">https://www.kaggle.com/datasets/munumbutt/amexfeather</a></p>\n<p>Then you can import it via pandas e.g.:</p>\n<p><code>train_csv = pd.read_feather(\"../input/amexfeather/train_data.ftr\")</code></p>\n<h3>Use chunks from dataset</h3>\n<p>You can also just use the original data and cut chunks out to start analyzing with them e.g.:</p>\n<p><code>train_csv = pd.read_csv(\"../input/amex-default-prediction/train_data.csv\", chunksize=500000)</code></p>\n<p>then you can use it like a normal dataframe after calling:</p>\n<p><code>train_csv_chunk = train_csv.get_chunk()</code></p>",
      "rawMarkdown": "You cant fit the original dataset into the kaggle notebook which supports only 16GB, but there are other things you can do:\n\n### Use size reduced dataset\nYou could use some of the reduced in size datasets published by the community (with different formats like parquet, feather, ...) e.g.:\n\nhttps://www.kaggle.com/datasets/munumbutt/amexfeather\n\nThen you can import it via pandas e.g.:\n\n`train_csv = pd.read_feather(\"../input/amexfeather/train_data.ftr\")`\n\n### Use chunks from dataset\n\nYou can also just use the original data and cut chunks out to start analyzing with them e.g.:\n\n`train_csv = pd.read_csv(\"../input/amex-default-prediction/train_data.csv\", chunksize=500000)`\n\nthen you can use it like a normal dataframe after calling:\n\n`train_csv_chunk = train_csv.get_chunk()`",
      "votes": null
    },
    {
      "id": "1822258",
      "postDate": "06/16/2022 07:25:28",
      "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> firstly thank you for your reply.</p>\n<p>Use chunks from dataset, this one is useful for train dataset and model but if i want to submit my test dataset submission will occur a problem. </p>",
      "rawMarkdown": "aliabdin1 firstly thank you for your reply.\n\nUse chunks from dataset, this one is useful for train dataset and model but if i want to submit my test dataset submission will occur a problem.",
      "votes": null
    },
    {
      "id": "1822712",
      "postDate": "06/16/2022 15:51:03",
      "content": "<p><a href=\"https://www.kaggle.com/baycelik\" target=\"_blank\">@baycelik</a> Yes, thats why you should use method #1 for testing and you could use method #2 to explore the train data or even test data but not to submit nor train a complete model.</p>",
      "rawMarkdown": "baycelik Yes, thats why you should use method #1 for testing and you could use method #2 to explore the train data or even test data but not to submit nor train a complete model.",
      "votes": null
    },
    {
      "id": "1823117",
      "postDate": "06/17/2022 04:34:11",
      "content": "<p>Thank you 👍</p>",
      "rawMarkdown": "Thank you 👍",
      "votes": null
    },
    {
      "id": "1823992",
      "postDate": "06/17/2022 21:22:49",
      "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> thank you</p>",
      "rawMarkdown": "aliabdin1 thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1819792,
      "author_name": "leewook",
      "author_url": "",
      "post_date": "06/14/2022 06:16:48",
      "content": "<p>Yes, this is caused by too much memory in the dataset you train or too much memory in the dataset you infer. The solution is to optimize memory or code it to catch bugs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1819796,
          "author_name": "baycelik",
          "author_url": "",
          "post_date": "06/14/2022 06:20:39",
          "content": "<p><a href=\"https://www.kaggle.com/leewook\" target=\"_blank\">@leewook</a> firstly thanks,</p>\n<p>You have any script related this subject.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1819871,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "06/14/2022 07:50:20",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/baycelik\" target=\"_blank\">@baycelik</a> This <a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-2-inference?scriptVersionId=96644290\" target=\"_blank\">old version of my Keras notebook</a> shows how to process the test data in chunks so that you never have the whole test data in memory at the same time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1819878,
          "author_name": "baycelik",
          "author_url": "",
          "post_date": "06/14/2022 07:56:42",
          "content": "<p><a href=\"https://www.kaggle.com/am140307\" target=\"_blank\">@am140307</a> thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1820656,
      "author_name": "cucasf",
      "author_url": "",
      "post_date": "06/14/2022 20:14:59",
      "content": "<p>There are many post about how to work with this amount of data, using parket or feather, changing columns types to optimize space, etc.</p>\n<p>Also this post <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/330347\" target=\"_blank\"> A List of EDA tricks when memory is limited</a> have some nice info about this topic.</p>\n<p>For me deleting Arrays or Dataframe and spliting the data in files have solved the issue.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1820872,
          "author_name": "baycelik",
          "author_url": "",
          "post_date": "06/15/2022 04:28:50",
          "content": "<p><a href=\"https://www.kaggle.com/cucasf\" target=\"_blank\">@cucasf</a> thanks for support</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1821811,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "06/15/2022 21:08:17",
      "content": "<p>You cant fit the original dataset into the kaggle notebook which supports only 16GB, but there are other things you can do:</p>\n<h3>Use size reduced dataset</h3>\n<p>You could use some of the reduced in size datasets published by the community (with different formats like parquet, feather, …) e.g.:</p>\n<p><a href=\"https://www.kaggle.com/datasets/munumbutt/amexfeather\" target=\"_blank\">https://www.kaggle.com/datasets/munumbutt/amexfeather</a></p>\n<p>Then you can import it via pandas e.g.:</p>\n<p><code>train_csv = pd.read_feather(\"../input/amexfeather/train_data.ftr\")</code></p>\n<h3>Use chunks from dataset</h3>\n<p>You can also just use the original data and cut chunks out to start analyzing with them e.g.:</p>\n<p><code>train_csv = pd.read_csv(\"../input/amex-default-prediction/train_data.csv\", chunksize=500000)</code></p>\n<p>then you can use it like a normal dataframe after calling:</p>\n<p><code>train_csv_chunk = train_csv.get_chunk()</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1822258,
          "author_name": "baycelik",
          "author_url": "",
          "post_date": "06/16/2022 07:25:28",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> firstly thank you for your reply.</p>\n<p>Use chunks from dataset, this one is useful for train dataset and model but if i want to submit my test dataset submission will occur a problem. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1822712,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "06/16/2022 15:51:03",
          "content": "<p><a href=\"https://www.kaggle.com/baycelik\" target=\"_blank\">@baycelik</a> Yes, thats why you should use method #1 for testing and you could use method #2 to explore the train data or even test data but not to submit nor train a complete model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1823117,
          "author_name": "pankajkumar2002",
          "author_url": "",
          "post_date": "06/17/2022 04:34:11",
          "content": "<p>Thank you 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1823992,
          "author_name": "baycelik",
          "author_url": "",
          "post_date": "06/17/2022 21:22:49",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1819762": "My model is ready but i handled it with nrows restriction. When i need to upload to a variable my test data i would face with this \"Your notebook tried to allocate more memory than is available. It has restarted\" error. \n\nIs there anyonce experienced this?\n\nThanks for support.",
    "1819792": "Yes, this is caused by too much memory in the dataset you train or too much memory in the dataset you infer. The solution is to optimize memory or code it to catch bugs.",
    "1819796": "leewook firstly thanks,\n\nYou have any script related this subject.",
    "1819871": "Hi @baycelik This [old version of my Keras notebook](https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-2-inference?scriptVersionId=96644290) shows how to process the test data in chunks so that you never have the whole test data in memory at the same time.",
    "1819878": "am140307 thank you",
    "1820656": "There are many post about how to work with this amount of data, using parket or feather, changing columns types to optimize space, etc.\n\nAlso this post [ A List of EDA tricks when memory is limited](https://www.kaggle.com/competitions/amex-default-prediction/discussion/330347) have some nice info about this topic.\n\nFor me deleting Arrays or Dataframe and spliting the data in files have solved the issue.",
    "1820872": "cucasf thanks for support",
    "1821811": "You cant fit the original dataset into the kaggle notebook which supports only 16GB, but there are other things you can do:\n\n### Use size reduced dataset\nYou could use some of the reduced in size datasets published by the community (with different formats like parquet, feather, ...) e.g.:\n\nhttps://www.kaggle.com/datasets/munumbutt/amexfeather\n\nThen you can import it via pandas e.g.:\n\n`train_csv = pd.read_feather(\"../input/amexfeather/train_data.ftr\")`\n\n### Use chunks from dataset\n\nYou can also just use the original data and cut chunks out to start analyzing with them e.g.:\n\n`train_csv = pd.read_csv(\"../input/amex-default-prediction/train_data.csv\", chunksize=500000)`\n\nthen you can use it like a normal dataframe after calling:\n\n`train_csv_chunk = train_csv.get_chunk()`",
    "1822258": "aliabdin1 firstly thank you for your reply.\n\nUse chunks from dataset, this one is useful for train dataset and model but if i want to submit my test dataset submission will occur a problem.",
    "1822712": "baycelik Yes, thats why you should use method #1 for testing and you could use method #2 to explore the train data or even test data but not to submit nor train a complete model.",
    "1823117": "Thank you 👍",
    "1823992": "aliabdin1 thank you"
  },
  "source": "meta"
}