{
  "id": 181737,
  "title": "How to train data in parts?",
  "url": "/competitions/birdsong-recognition/discussion/181737",
  "author_name": "Arindam",
  "post_date": "2020-09-10T01:54:04.938000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I have spit whole data(audio files) into 10 equal parts and after data pre-processing, I have stored them into total 10 <strong>.npy files</strong> . Till now all good, but when I load them Ram goes out of memory, which is quite understandable as each \".npy\" file is around 3GB. So loading all together will need (3*10=)30 GB of Ram. (Kaggle ram is 16GB)</p>\n<blockquote>\n  <p>How I am loading: <code>audio_numy_batch=np.load( filename ,allow_pickle=True)</code> </p>\n  <p>What my one 'npy/numpy' files contain: list of pre-processed signals (of ~3000 audios) <br>\n  (this data of 3gb will be stored in Ram)</p>\n</blockquote>\n<p><strong>Question-Is there a way I can train on one .npy file and save the model and use that model again on next .npy file like that which will repeat 10 times?</strong></p>\n<p>In simple terms- How can we Train data in splits in different notebook?</p>\n<p>Any blog links or any suggestion will be highly helpful</p>",
  "messages": [
    {
      "id": 1008214,
      "postDate": "2020-09-12T20:37:39.097Z",
      "content": "<p>You don't need to have all the dataset loaded in memory at the same time.</p>\n<p>I think you might want to give <a href=\"https://www.tensorflow.org/guide/data\" target=\"_blank\">tf.data</a> a try. You can store many records in a single file (recommended size is 100-200 MB per file) and then use TFRecordsDataset which will handle the files automatically, so it won't open them all at the same time (no memory crash)</p>",
      "rawMarkdown": "You don't need to have all the dataset loaded in memory at the same time.\n\nI think you might want to give [tf.data](https://www.tensorflow.org/guide/data) a try. You can store many records in a single file (recommended size is 100-200 MB per file) and then use TFRecordsDataset which will handle the files automatically, so it won't open them all at the same time (no memory crash)",
      "votes": 1,
      "replies": [
        {
          "id": 1009689,
          "postDate": "2020-09-14T07:07:59.107Z",
          "content": "<p>After searching how to use tf.data I found this <br>\n<a href=\"https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline\" target=\"_blank\">https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline</a> <br>\nWill give it a shot. Thanks man <a href=\"https://www.kaggle.com/benayas\" target=\"_blank\">@benayas</a> </p>",
          "rawMarkdown": "After searching how to use tf.data I found this \nhttps://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline \nWill give it a shot. Thanks man @benayas ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1005335,
      "postDate": "2020-09-10T12:34:57.250Z",
      "content": "<p>You could try something like this:</p>\n<ol>\n<li>Load two files into memory when your dataset is created.</li>\n<li>Once all samples from the first file have been trained, load the 3rd file and remove samples from the 1st file to free up memory.</li>\n<li>Repeat for all the files.</li>\n</ol>",
      "rawMarkdown": "You could try something like this:\n1. Load two files into memory when your dataset is created.\n2. Once all samples from the first file have been trained, load the 3rd file and remove samples from the 1st file to free up memory.\n3. Repeat for all the files.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1005448,
          "postDate": "2020-09-10T13:54:14.237Z",
          "content": "<p>Thats great Idea <a href=\"https://www.kaggle.com/jackvial\" target=\"_blank\">@jackvial</a> </p>\n<pre><code>f1=np.load( f.npy , allow_pickle=True)  [3gb] (f1-&gt;features)\nl1=np.load( l.npy , allow_pickle=True)  [300 kb] (l1-&gt;labels)\n\ndel f1\ndel l1\ngc.collect()\n</code></pre>\n<p>However I tried to free up space by using above method but ram doesn't get free. Is there any other way?</p>",
          "rawMarkdown": "Thats great Idea @jackvial \n```\nf1=np.load( f.npy , allow_pickle=True)  [3gb] (f1->features)\nl1=np.load( l.npy , allow_pickle=True)  [300 kb] (l1->labels)\n\ndel f1\ndel l1\ngc.collect()\n```\nHowever I tried to free up space by using above method but ram doesn't get free. Is there any other way?"
        },
        {
          "id": 1008278,
          "postDate": "2020-09-13T00:22:57.600Z",
          "content": "<p>The way you are doing it <em>should</em> work. Are you running it in a notebook or script?</p>",
          "rawMarkdown": "The way you are doing it *should* work. Are you running it in a notebook or script?"
        },
        {
          "id": 1009646,
          "postDate": "2020-09-14T06:45:43.683Z",
          "content": "<p>currently in notebook, running in script can help in disk space maybe. Is there a difference</p>",
          "rawMarkdown": "currently in notebook, running in script can help in disk space maybe. Is there a difference"
        }
      ]
    },
    {
      "id": 1004746,
      "postDate": "2020-09-10T01:54:04.940Z",
      "content": "<p>I have spit whole data(audio files) into 10 equal parts and after data pre-processing, I have stored them into total 10 <strong>.npy files</strong> . Till now all good, but when I load them Ram goes out of memory, which is quite understandable as each \".npy\" file is around 3GB. So loading all together will need (3*10=)30 GB of Ram. (Kaggle ram is 16GB)</p>\n<blockquote>\n  <p>How I am loading: <code>audio_numy_batch=np.load( filename ,allow_pickle=True)</code> </p>\n  <p>What my one 'npy/numpy' files contain: list of pre-processed signals (of ~3000 audios) <br>\n  (this data of 3gb will be stored in Ram)</p>\n</blockquote>\n<p><strong>Question-Is there a way I can train on one .npy file and save the model and use that model again on next .npy file like that which will repeat 10 times?</strong></p>\n<p>In simple terms- How can we Train data in splits in different notebook?</p>\n<p>Any blog links or any suggestion will be highly helpful</p>",
      "rawMarkdown": "I have spit whole data(audio files) into 10 equal parts and after data pre-processing, I have stored them into total 10 **.npy files** . Till now all good, but when I load them Ram goes out of memory, which is quite understandable as each \".npy\" file is around 3GB. So loading all together will need (3*10=)30 GB of Ram. (Kaggle ram is 16GB)\n\n> How I am loading: `audio_numy_batch=np.load( filename ,allow_pickle=True)` \n\n> What my one 'npy/numpy' files contain: list of pre-processed signals (of ~3000 audios) \n(this data of 3gb will be stored in Ram)\n\n**Question-Is there a way I can train on one .npy file and save the model and use that model again on next .npy file like that which will repeat 10 times?**\n\nIn simple terms- How can we Train data in splits in different notebook?\n\nAny blog links or any suggestion will be highly helpful",
      "votes": 1
    },
    {
      "id": 1005543,
      "postDate": "2020-09-10T15:12:25.827Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1008214,
      "author_name": "Alberto Benayas",
      "author_url": "",
      "post_date": "2020-09-12T20:37:39.097000",
      "content": "<p>You don't need to have all the dataset loaded in memory at the same time.</p>\n<p>I think you might want to give <a href=\"https://www.tensorflow.org/guide/data\" target=\"_blank\">tf.data</a> a try. You can store many records in a single file (recommended size is 100-200 MB per file) and then use TFRecordsDataset which will handle the files automatically, so it won't open them all at the same time (no memory crash)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1009689,
          "author_name": "Arindam",
          "author_url": "",
          "post_date": "2020-09-14T07:07:59.107000",
          "content": "<p>After searching how to use tf.data I found this <br>\n<a href=\"https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline\" target=\"_blank\">https://stackoverflow.com/questions/48889482/feeding-npy-numpy-files-into-tensorflow-data-pipeline</a> <br>\nWill give it a shot. Thanks man <a href=\"https://www.kaggle.com/benayas\" target=\"_blank\">@benayas</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1005335,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2020-09-10T12:34:57.250000",
      "content": "<p>You could try something like this:</p>\n<ol>\n<li>Load two files into memory when your dataset is created.</li>\n<li>Once all samples from the first file have been trained, load the 3rd file and remove samples from the 1st file to free up memory.</li>\n<li>Repeat for all the files.</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1005448,
          "author_name": "Arindam",
          "author_url": "",
          "post_date": "2020-09-10T13:54:14.237000",
          "content": "<p>Thats great Idea <a href=\"https://www.kaggle.com/jackvial\" target=\"_blank\">@jackvial</a> </p>\n<pre><code>f1=np.load( f.npy , allow_pickle=True)  [3gb] (f1-&gt;features)\nl1=np.load( l.npy , allow_pickle=True)  [300 kb] (l1-&gt;labels)\n\ndel f1\ndel l1\ngc.collect()\n</code></pre>\n<p>However I tried to free up space by using above method but ram doesn't get free. Is there any other way?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1008278,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2020-09-13T00:22:57.600000",
          "content": "<p>The way you are doing it <em>should</em> work. Are you running it in a notebook or script?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1009646,
          "author_name": "Arindam",
          "author_url": "",
          "post_date": "2020-09-14T06:45:43.683000",
          "content": "<p>currently in notebook, running in script can help in disk space maybe. Is there a difference</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1005543,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-10T15:12:25.827000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1008214": "You don't need to have all the dataset loaded in memory at the same time.\n\nI think you might want to give [tf.data](https://www.tensorflow.org/guide/data) a try. You can store many records in a single file (recommended size is 100-200 MB per file) and then use TFRecordsDataset which will handle the files automatically, so it won't open them all at the same time (no memory crash)",
    "1005335": "You could try something like this:\n1. Load two files into memory when your dataset is created.\n2. Once all samples from the first file have been trained, load the 3rd file and remove samples from the 1st file to free up memory.\n3. Repeat for all the files.\n",
    "1004746": "I have spit whole data(audio files) into 10 equal parts and after data pre-processing, I have stored them into total 10 **.npy files** . Till now all good, but when I load them Ram goes out of memory, which is quite understandable as each \".npy\" file is around 3GB. So loading all together will need (3*10=)30 GB of Ram. (Kaggle ram is 16GB)\n\n> How I am loading: `audio_numy_batch=np.load( filename ,allow_pickle=True)` \n\n> What my one 'npy/numpy' files contain: list of pre-processed signals (of ~3000 audios) \n(this data of 3gb will be stored in Ram)\n\n**Question-Is there a way I can train on one .npy file and save the model and use that model again on next .npy file like that which will repeat 10 times?**\n\nIn simple terms- How can we Train data in splits in different notebook?\n\nAny blog links or any suggestion will be highly helpful",
    "1005543": ""
  }
}