{
  "id": 168169,
  "title": "Running out of memory",
  "url": "/competitions/birdsong-recognition/discussion/168169",
  "author_name": "Sepideh Doost",
  "post_date": "2020-07-19T14:03:22.971000",
  "votes": 7,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I'm constantly running out of RAM. So far, I have managed to save the spectrograms (from 1000 audio files only) in two separate datasets and read them in my notebook. But, after reading them and combining them into an numpy array, I face the RAM issue again. What do you do to avoid it?</p>\n\n<p>Thanks,\nSepideh </p>",
  "messages": [
    {
      "id": 935589,
      "postDate": "2020-07-19T14:03:22.973Z",
      "content": "<p>Hi,</p>\n\n<p>I'm constantly running out of RAM. So far, I have managed to save the spectrograms (from 1000 audio files only) in two separate datasets and read them in my notebook. But, after reading them and combining them into an numpy array, I face the RAM issue again. What do you do to avoid it?</p>\n\n<p>Thanks,\nSepideh </p>",
      "rawMarkdown": "Hi,\n\nI'm constantly running out of RAM. So far, I have managed to save the spectrograms (from 1000 audio files only) in two separate datasets and read them in my notebook. But, after reading them and combining them into an numpy array, I face the RAM issue again. What do you do to avoid it?\n\nThanks,\nSepideh ",
      "votes": 7
    },
    {
      "id": 935742,
      "postDate": "2020-07-19T15:47:16.320Z",
      "content": "<p>Sepideh, have you measured the size of your numpy arrays? You can do it invoking *array_name_here*.<strong>nbytes</strong>. This may give you some insight on how to deal with your data. (Note: I am not participating in this competition.)\nNot knowing your data in detail, some possible actions to explore include:\n(a) optimizing <strong>dtypes</strong>. Example: if you have an array of integers, declaring dtype=np.uint16 instead of dtype=np.uint64 would lead to a 75% memory allocation reduction. Same for float dtypes;\n(b) if you're dealing with sparse arrays (e.g. a binary array with a vast amount of zeros, few occurrences of ones) specific libraries like '<a href=\"https://sparse.pydata.org/en/latest/\">sparse</a>' might help you.</p>",
      "rawMarkdown": "Sepideh, have you measured the size of your numpy arrays? You can do it invoking *array_name_here*.**nbytes**. This may give you some insight on how to deal with your data. (Note: I am not participating in this competition.)\nNot knowing your data in detail, some possible actions to explore include:\n(a) optimizing **dtypes**. Example: if you have an array of integers, declaring dtype=np.uint16 instead of dtype=np.uint64 would lead to a 75% memory allocation reduction. Same for float dtypes;\n(b) if you're dealing with sparse arrays (e.g. a binary array with a vast amount of zeros, few occurrences of ones) specific libraries like '[sparse](https://sparse.pydata.org/en/latest/)' might help you.",
      "votes": 3,
      "replies": [
        {
          "id": 935799,
          "postDate": "2020-07-19T16:32:43.530Z",
          "content": "<p>Thank you Paulo for your help and suggestions. My main and biggest numpy array ((1000, 1251, 1025) with float32) is about 2.6GB.</p>\n\n<p>I'll follow your advice and see if it improves. :)</p>",
          "rawMarkdown": "Thank you Paulo for your help and suggestions. My main and biggest numpy array ((1000, 1251, 1025) with float32) is about 2.6GB.\n\nI'll follow your advice and see if it improves. :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 937588,
      "postDate": "2020-07-21T05:03:41.743Z",
      "content": "<p>I had a pretty painful time trying to understand TF docs but think Tom is correct in pointing out that it manages the input to the model for training quite nicely once done.</p>\n\n<p>I don't know whether everyone on the forum is generally pretty up to speed on these or whether others have also found the official docs quite difficult. There's a couple of good notebooks by Chris Deotte but for me those were already a couple of levels above my starting ability.</p>\n\n<p>I drafted a very quick notebook saving spectrogram clips and loading from TFRecords into a model and running in case it's helpful to anyone else getting a headache trying this. Tried to go line by line as I found it tough when I started. In particular I think it's not obvious sometimes in official examples how each data point has to be mapped with a data type and read in / out and how this can be adapted, and think errors in data mapping when loading often don't actually generate an error until the model runs.</p>\n\n<p>Disclaimer: I don't actually really know what I'm doing so if anyone wants to point out my mistakes, provide a better example etc then it's appreciated.</p>\n\n<p><a href=\"https://www.kaggle.com/davidedwards1/tf-records-newbie-test\">https://www.kaggle.com/davidedwards1/tf-records-newbie-test</a></p>",
      "rawMarkdown": "I had a pretty painful time trying to understand TF docs but think Tom is correct in pointing out that it manages the input to the model for training quite nicely once done.\n\nI don't know whether everyone on the forum is generally pretty up to speed on these or whether others have also found the official docs quite difficult. There's a couple of good notebooks by Chris Deotte but for me those were already a couple of levels above my starting ability.\n\nI drafted a very quick notebook saving spectrogram clips and loading from TFRecords into a model and running in case it's helpful to anyone else getting a headache trying this. Tried to go line by line as I found it tough when I started. In particular I think it's not obvious sometimes in official examples how each data point has to be mapped with a data type and read in / out and how this can be adapted, and think errors in data mapping when loading often don't actually generate an error until the model runs.\n\nDisclaimer: I don't actually really know what I'm doing so if anyone wants to point out my mistakes, provide a better example etc then it's appreciated.\n\nhttps://www.kaggle.com/davidedwards1/tf-records-newbie-test",
      "votes": 1,
      "replies": [
        {
          "id": 946667,
          "postDate": "2020-07-26T18:03:23.450Z",
          "content": "<p>Thanks Dave. I read your notebook. It was very helpful. 👍 </p>",
          "rawMarkdown": "Thanks Dave. I read your notebook. It was very helpful. 👍 "
        }
      ]
    },
    {
      "id": 936158,
      "postDate": "2020-07-20T03:03:02.323Z",
      "content": "<p>Think I've had issues with RAM - even after creating and saving numpy arrays and clearing out the stored objects, gc.collect() etc, the RAM has been maxed out. Feel like I'm missing some stupid mistake I'm making but can't seem to flush it out.</p>\n\n<p>Aside from RAM - with saved file size I feel like with this comp it's quite difficult to save training samples as numpy arrays and keep within reasonable / Kaggle-allowed file sizes. As per Paulo's comment I tried dtypes - not sure if my thinking is correct but I tried rounding all numbers to integers 0-255 to use np.uint8 data type. My understanding of numpy docs is that this is OK but someone can correct me if wrong. Think there is some loss of data in doing this as floats are obviously rounded to closest (1/255).</p>\n\n<p>I found with the spectrograms that the sparse matrix wasn't so useful as there's a fair amount of background noise. Even if this could be cleaned perfectly not sure whether having 'cleaned' training data is actually going to help train a robust model.</p>\n\n<p>I tried saving to TfRecords which was a whole separate splitting headache to understand but this is now better in terms of saved file size. But it doesn't feel like a very flexible format (no train test split function??) and not that easy for a novice like me. Still, after some fairly painful battling with documentation I have gotten a lot of data into Tfrecords and back out again into training a tensorflow model. Think I have around 5,800 arrays with (800 * 150 * 1) dimensions populated with 0-255 integers saved with total file size something like 1gb. Think this looks like maybe a 33% improvement vs 2.6gb for (1000 * 1251 * 1025) tho maybe that is just due to changing the dtype.</p>\n\n<p>I guess the other thing that potentially seemed helpful with tfrecords is that in loading the data,  I don't think it loads the whole dataset - it reads batches from the storage files. I wasn't sure how viable it was going to be to train a model while needing to load large numpy arrays into memory.</p>\n\n<p>If anyone has any bright ideas on more efficient storage (or on my issue with creating/saving/deleting numpy arrays seeming to drain RAM) I'm definitely interested.</p>",
      "rawMarkdown": "Think I've had issues with RAM - even after creating and saving numpy arrays and clearing out the stored objects, gc.collect() etc, the RAM has been maxed out. Feel like I'm missing some stupid mistake I'm making but can't seem to flush it out.\n\nAside from RAM - with saved file size I feel like with this comp it's quite difficult to save training samples as numpy arrays and keep within reasonable / Kaggle-allowed file sizes. As per Paulo's comment I tried dtypes - not sure if my thinking is correct but I tried rounding all numbers to integers 0-255 to use np.uint8 data type. My understanding of numpy docs is that this is OK but someone can correct me if wrong. Think there is some loss of data in doing this as floats are obviously rounded to closest (1/255).\n\nI found with the spectrograms that the sparse matrix wasn't so useful as there's a fair amount of background noise. Even if this could be cleaned perfectly not sure whether having 'cleaned' training data is actually going to help train a robust model.\n\nI tried saving to TfRecords which was a whole separate splitting headache to understand but this is now better in terms of saved file size. But it doesn't feel like a very flexible format (no train test split function??) and not that easy for a novice like me. Still, after some fairly painful battling with documentation I have gotten a lot of data into Tfrecords and back out again into training a tensorflow model. Think I have around 5,800 arrays with (800 * 150 * 1) dimensions populated with 0-255 integers saved with total file size something like 1gb. Think this looks like maybe a 33% improvement vs 2.6gb for (1000 * 1251 * 1025) tho maybe that is just due to changing the dtype.\n\nI guess the other thing that potentially seemed helpful with tfrecords is that in loading the data,  I don't think it loads the whole dataset - it reads batches from the storage files. I wasn't sure how viable it was going to be to train a model while needing to load large numpy arrays into memory.\n\nIf anyone has any bright ideas on more efficient storage (or on my issue with creating/saving/deleting numpy arrays seeming to drain RAM) I'm definitely interested.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 936941,
          "postDate": "2020-07-20T16:02:45.627Z",
          "content": "<p>One way could be data generator. <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator\">https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator</a></p>",
          "rawMarkdown": "One way could be data generator. https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator",
          "votes": 1
        },
        {
          "id": 937248,
          "postDate": "2020-07-20T21:06:30.583Z",
          "content": "<p>Hi Dave, thanks for sharing your experiences. This competition so far has been tough for me in terms of constant storage and memory issues. I've already got some improvement from dtype changing to float16. I'll try to avoid int8 as much as possible. I'm planning to look into other suggestions in this thread and hopefully it will be enough. Good luck!</p>",
          "rawMarkdown": "Hi Dave, thanks for sharing your experiences. This competition so far has been tough for me in terms of constant storage and memory issues. I've already got some improvement from dtype changing to float16. I'll try to avoid int8 as much as possible. I'm planning to look into other suggestions in this thread and hopefully it will be enough. Good luck!"
        }
      ]
    },
    {
      "id": 936722,
      "postDate": "2020-07-20T13:27:12.123Z",
      "content": "<p>Faced similar problem. That's what I did\n1. Split whole dataset into 6 small data set\n2. 6 notebooks parallel running- 7hrs run time\n3. Extracted features and  did normalisation/mean in same notebook- to reduced features size (imp)\n4. finally Assembled all 6 featured data in one notebook\nNow you are good to go :)</p>",
      "rawMarkdown": "Faced similar problem. That's what I did\n1. Split whole dataset into 6 small data set\n2. 6 notebooks parallel running- 7hrs run time\n3. Extracted features and  did normalisation/mean in same notebook- to reduced features size (imp)\n4. finally Assembled all 6 featured data in one notebook\nNow you are good to go :)",
      "votes": 2,
      "replies": [
        {
          "id": 937231,
          "postDate": "2020-07-20T20:51:18.213Z",
          "content": "<p>Great suggestion. I haven't worked on the full dataset yet. I'll do something similar when I get there. Thank you. </p>",
          "rawMarkdown": "Great suggestion. I haven't worked on the full dataset yet. I'll do something similar when I get there. Thank you. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 936192,
      "postDate": "2020-07-20T04:01:39.463Z",
      "content": "<p>For my own pipelines I usually use 'just in time' example creation. Tensorflow <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset\">datasets</a> are very good for this: batches are created and fed to the model as they are needed (and you can use tf.data.prefetch for a bit of buffering to ensure training never slows down). With a bit more effort, you can even read multiple filesets (say, one for noise and one for training examples) and use the <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#zip\">zip</a> and <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map\">map</a> functions to mix them on the fly. There's a guide for getting started <a href=\"https://www.tensorflow.org/guide/data\">here</a>.</p>\n\n<p>(I'm not sure what similar options exist for torch, I'm afraid.)</p>",
      "rawMarkdown": "For my own pipelines I usually use 'just in time' example creation. Tensorflow [datasets](https://www.tensorflow.org/api_docs/python/tf/data/Dataset) are very good for this: batches are created and fed to the model as they are needed (and you can use tf.data.prefetch for a bit of buffering to ensure training never slows down). With a bit more effort, you can even read multiple filesets (say, one for noise and one for training examples) and use the [zip](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#zip) and [map](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map) functions to mix them on the fly. There's a guide for getting started [here](https://www.tensorflow.org/guide/data).\n\n(I'm not sure what similar options exist for torch, I'm afraid.)",
      "votes": 2,
      "replies": [
        {
          "id": 937235,
          "postDate": "2020-07-20T20:53:22.783Z",
          "content": "<p>Thank you Tom. I should admit that I'm not much familiar with tensorflow datasets. This can be a good opportunity to start learning though. Thanks for your advice. </p>",
          "rawMarkdown": "Thank you Tom. I should admit that I'm not much familiar with tensorflow datasets. This can be a good opportunity to start learning though. Thanks for your advice. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 968507,
      "postDate": "2020-08-13T04:31:23.570Z",
      "content": "<p>ImageDataGenerator is one of the best ways to deal with huge image datasets according to me.<br>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator</a></p>\n<p>I hope you find this useful. <br>\nthank you</p>",
      "rawMarkdown": "ImageDataGenerator is one of the best ways to deal with huge image datasets according to me.\nhttps://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator\n\nI hope you find this useful. \nthank you"
    },
    {
      "id": 938982,
      "postDate": "2020-07-21T23:29:56.567Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 935742,
      "author_name": "Paulo Breviglieri",
      "author_url": "",
      "post_date": "2020-07-19T15:47:16.320000",
      "content": "<p>Sepideh, have you measured the size of your numpy arrays? You can do it invoking *array_name_here*.<strong>nbytes</strong>. This may give you some insight on how to deal with your data. (Note: I am not participating in this competition.)\nNot knowing your data in detail, some possible actions to explore include:\n(a) optimizing <strong>dtypes</strong>. Example: if you have an array of integers, declaring dtype=np.uint16 instead of dtype=np.uint64 would lead to a 75% memory allocation reduction. Same for float dtypes;\n(b) if you're dealing with sparse arrays (e.g. a binary array with a vast amount of zeros, few occurrences of ones) specific libraries like '<a href=\"https://sparse.pydata.org/en/latest/\">sparse</a>' might help you.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 935799,
          "author_name": "Sepideh Doost",
          "author_url": "",
          "post_date": "2020-07-19T16:32:43.530000",
          "content": "<p>Thank you Paulo for your help and suggestions. My main and biggest numpy array ((1000, 1251, 1025) with float32) is about 2.6GB.</p>\n\n<p>I'll follow your advice and see if it improves. :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 937588,
      "author_name": "Dave E",
      "author_url": "",
      "post_date": "2020-07-21T05:03:41.743000",
      "content": "<p>I had a pretty painful time trying to understand TF docs but think Tom is correct in pointing out that it manages the input to the model for training quite nicely once done.</p>\n\n<p>I don't know whether everyone on the forum is generally pretty up to speed on these or whether others have also found the official docs quite difficult. There's a couple of good notebooks by Chris Deotte but for me those were already a couple of levels above my starting ability.</p>\n\n<p>I drafted a very quick notebook saving spectrogram clips and loading from TFRecords into a model and running in case it's helpful to anyone else getting a headache trying this. Tried to go line by line as I found it tough when I started. In particular I think it's not obvious sometimes in official examples how each data point has to be mapped with a data type and read in / out and how this can be adapted, and think errors in data mapping when loading often don't actually generate an error until the model runs.</p>\n\n<p>Disclaimer: I don't actually really know what I'm doing so if anyone wants to point out my mistakes, provide a better example etc then it's appreciated.</p>\n\n<p><a href=\"https://www.kaggle.com/davidedwards1/tf-records-newbie-test\">https://www.kaggle.com/davidedwards1/tf-records-newbie-test</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 946667,
          "author_name": "Sepideh Doost",
          "author_url": "",
          "post_date": "2020-07-26T18:03:23.450000",
          "content": "<p>Thanks Dave. I read your notebook. It was very helpful. 👍 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 936158,
      "author_name": "Dave E",
      "author_url": "",
      "post_date": "2020-07-20T03:03:02.323000",
      "content": "<p>Think I've had issues with RAM - even after creating and saving numpy arrays and clearing out the stored objects, gc.collect() etc, the RAM has been maxed out. Feel like I'm missing some stupid mistake I'm making but can't seem to flush it out.</p>\n\n<p>Aside from RAM - with saved file size I feel like with this comp it's quite difficult to save training samples as numpy arrays and keep within reasonable / Kaggle-allowed file sizes. As per Paulo's comment I tried dtypes - not sure if my thinking is correct but I tried rounding all numbers to integers 0-255 to use np.uint8 data type. My understanding of numpy docs is that this is OK but someone can correct me if wrong. Think there is some loss of data in doing this as floats are obviously rounded to closest (1/255).</p>\n\n<p>I found with the spectrograms that the sparse matrix wasn't so useful as there's a fair amount of background noise. Even if this could be cleaned perfectly not sure whether having 'cleaned' training data is actually going to help train a robust model.</p>\n\n<p>I tried saving to TfRecords which was a whole separate splitting headache to understand but this is now better in terms of saved file size. But it doesn't feel like a very flexible format (no train test split function??) and not that easy for a novice like me. Still, after some fairly painful battling with documentation I have gotten a lot of data into Tfrecords and back out again into training a tensorflow model. Think I have around 5,800 arrays with (800 * 150 * 1) dimensions populated with 0-255 integers saved with total file size something like 1gb. Think this looks like maybe a 33% improvement vs 2.6gb for (1000 * 1251 * 1025) tho maybe that is just due to changing the dtype.</p>\n\n<p>I guess the other thing that potentially seemed helpful with tfrecords is that in loading the data,  I don't think it loads the whole dataset - it reads batches from the storage files. I wasn't sure how viable it was going to be to train a model while needing to load large numpy arrays into memory.</p>\n\n<p>If anyone has any bright ideas on more efficient storage (or on my issue with creating/saving/deleting numpy arrays seeming to drain RAM) I'm definitely interested.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 936941,
          "author_name": "Shikha Agrawal",
          "author_url": "",
          "post_date": "2020-07-20T16:02:45.627000",
          "content": "<p>One way could be data generator. <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator\">https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 937248,
          "author_name": "Sepideh Doost",
          "author_url": "",
          "post_date": "2020-07-20T21:06:30.583000",
          "content": "<p>Hi Dave, thanks for sharing your experiences. This competition so far has been tough for me in terms of constant storage and memory issues. I've already got some improvement from dtype changing to float16. I'll try to avoid int8 as much as possible. I'm planning to look into other suggestions in this thread and hopefully it will be enough. Good luck!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 936722,
      "author_name": "Arindam",
      "author_url": "",
      "post_date": "2020-07-20T13:27:12.123000",
      "content": "<p>Faced similar problem. That's what I did\n1. Split whole dataset into 6 small data set\n2. 6 notebooks parallel running- 7hrs run time\n3. Extracted features and  did normalisation/mean in same notebook- to reduced features size (imp)\n4. finally Assembled all 6 featured data in one notebook\nNow you are good to go :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 937231,
          "author_name": "Sepideh Doost",
          "author_url": "",
          "post_date": "2020-07-20T20:51:18.213000",
          "content": "<p>Great suggestion. I haven't worked on the full dataset yet. I'll do something similar when I get there. Thank you. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 936192,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2020-07-20T04:01:39.463000",
      "content": "<p>For my own pipelines I usually use 'just in time' example creation. Tensorflow <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset\">datasets</a> are very good for this: batches are created and fed to the model as they are needed (and you can use tf.data.prefetch for a bit of buffering to ensure training never slows down). With a bit more effort, you can even read multiple filesets (say, one for noise and one for training examples) and use the <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#zip\">zip</a> and <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map\">map</a> functions to mix them on the fly. There's a guide for getting started <a href=\"https://www.tensorflow.org/guide/data\">here</a>.</p>\n\n<p>(I'm not sure what similar options exist for torch, I'm afraid.)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 937235,
          "author_name": "Sepideh Doost",
          "author_url": "",
          "post_date": "2020-07-20T20:53:22.783000",
          "content": "<p>Thank you Tom. I should admit that I'm not much familiar with tensorflow datasets. This can be a good opportunity to start learning though. Thanks for your advice. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 968507,
      "author_name": "palak sood",
      "author_url": "",
      "post_date": "2020-08-13T04:31:23.570000",
      "content": "<p>ImageDataGenerator is one of the best ways to deal with huge image datasets according to me.<br>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator</a></p>\n<p>I hope you find this useful. <br>\nthank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 938982,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-21T23:29:56.567000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "935589": "Hi,\n\nI'm constantly running out of RAM. So far, I have managed to save the spectrograms (from 1000 audio files only) in two separate datasets and read them in my notebook. But, after reading them and combining them into an numpy array, I face the RAM issue again. What do you do to avoid it?\n\nThanks,\nSepideh ",
    "935742": "Sepideh, have you measured the size of your numpy arrays? You can do it invoking *array_name_here*.**nbytes**. This may give you some insight on how to deal with your data. (Note: I am not participating in this competition.)\nNot knowing your data in detail, some possible actions to explore include:\n(a) optimizing **dtypes**. Example: if you have an array of integers, declaring dtype=np.uint16 instead of dtype=np.uint64 would lead to a 75% memory allocation reduction. Same for float dtypes;\n(b) if you're dealing with sparse arrays (e.g. a binary array with a vast amount of zeros, few occurrences of ones) specific libraries like '[sparse](https://sparse.pydata.org/en/latest/)' might help you.",
    "937588": "I had a pretty painful time trying to understand TF docs but think Tom is correct in pointing out that it manages the input to the model for training quite nicely once done.\n\nI don't know whether everyone on the forum is generally pretty up to speed on these or whether others have also found the official docs quite difficult. There's a couple of good notebooks by Chris Deotte but for me those were already a couple of levels above my starting ability.\n\nI drafted a very quick notebook saving spectrogram clips and loading from TFRecords into a model and running in case it's helpful to anyone else getting a headache trying this. Tried to go line by line as I found it tough when I started. In particular I think it's not obvious sometimes in official examples how each data point has to be mapped with a data type and read in / out and how this can be adapted, and think errors in data mapping when loading often don't actually generate an error until the model runs.\n\nDisclaimer: I don't actually really know what I'm doing so if anyone wants to point out my mistakes, provide a better example etc then it's appreciated.\n\nhttps://www.kaggle.com/davidedwards1/tf-records-newbie-test",
    "936158": "Think I've had issues with RAM - even after creating and saving numpy arrays and clearing out the stored objects, gc.collect() etc, the RAM has been maxed out. Feel like I'm missing some stupid mistake I'm making but can't seem to flush it out.\n\nAside from RAM - with saved file size I feel like with this comp it's quite difficult to save training samples as numpy arrays and keep within reasonable / Kaggle-allowed file sizes. As per Paulo's comment I tried dtypes - not sure if my thinking is correct but I tried rounding all numbers to integers 0-255 to use np.uint8 data type. My understanding of numpy docs is that this is OK but someone can correct me if wrong. Think there is some loss of data in doing this as floats are obviously rounded to closest (1/255).\n\nI found with the spectrograms that the sparse matrix wasn't so useful as there's a fair amount of background noise. Even if this could be cleaned perfectly not sure whether having 'cleaned' training data is actually going to help train a robust model.\n\nI tried saving to TfRecords which was a whole separate splitting headache to understand but this is now better in terms of saved file size. But it doesn't feel like a very flexible format (no train test split function??) and not that easy for a novice like me. Still, after some fairly painful battling with documentation I have gotten a lot of data into Tfrecords and back out again into training a tensorflow model. Think I have around 5,800 arrays with (800 * 150 * 1) dimensions populated with 0-255 integers saved with total file size something like 1gb. Think this looks like maybe a 33% improvement vs 2.6gb for (1000 * 1251 * 1025) tho maybe that is just due to changing the dtype.\n\nI guess the other thing that potentially seemed helpful with tfrecords is that in loading the data,  I don't think it loads the whole dataset - it reads batches from the storage files. I wasn't sure how viable it was going to be to train a model while needing to load large numpy arrays into memory.\n\nIf anyone has any bright ideas on more efficient storage (or on my issue with creating/saving/deleting numpy arrays seeming to drain RAM) I'm definitely interested.\n\n",
    "936722": "Faced similar problem. That's what I did\n1. Split whole dataset into 6 small data set\n2. 6 notebooks parallel running- 7hrs run time\n3. Extracted features and  did normalisation/mean in same notebook- to reduced features size (imp)\n4. finally Assembled all 6 featured data in one notebook\nNow you are good to go :)",
    "936192": "For my own pipelines I usually use 'just in time' example creation. Tensorflow [datasets](https://www.tensorflow.org/api_docs/python/tf/data/Dataset) are very good for this: batches are created and fed to the model as they are needed (and you can use tf.data.prefetch for a bit of buffering to ensure training never slows down). With a bit more effort, you can even read multiple filesets (say, one for noise and one for training examples) and use the [zip](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#zip) and [map](https://www.tensorflow.org/api_docs/python/tf/data/Dataset#map) functions to mix them on the fly. There's a guide for getting started [here](https://www.tensorflow.org/guide/data).\n\n(I'm not sure what similar options exist for torch, I'm afraid.)",
    "968507": "ImageDataGenerator is one of the best ways to deal with huge image datasets according to me.\nhttps://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator\n\nI hope you find this useful. \nthank you",
    "938982": ""
  }
}