{
  "id": 201541,
  "title": "RAM Overflow",
  "url": "/competitions/rfcx-species-audio-detection/discussion/201541",
  "author_name": "",
  "post_date": "2020-12-05T14:04:29.321122900Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hey Guys, I tried to run the code stated below for this competition but the problem I encountered is that even before completing 1216 files my CPU RAM overflows and Session Restarts. <br>\nAny Sugesstions?</p>\n<p>tp = []<br>\ntp_labels = []<br>\nhop_length = 128<br>\nn_fft = 512<br>\nfor i in tqdm(range(len(train_tp))):<br>\n    idx = train_tp['recording_id'][i]<br>\n    file_path = '../input/rfcx-species-audio-detection/train/' + idx + '.flac'<br>\n    signal, sr = sf.read(file_path)<br>\n    m = np.abs(np.fft.fft(signal)[:len(signal)//2])<br>\n    tp.append(m)<br>\n    tp_labels.append(train_tp['species_id'][i])</p>",
  "messages": [
    {
      "id": "1102949",
      "postDate": "12/05/2020 14:04:29",
      "content": "<p>Hey Guys, I tried to run the code stated below for this competition but the problem I encountered is that even before completing 1216 files my CPU RAM overflows and Session Restarts. <br>\nAny Sugesstions?</p>\n<p>tp = []<br>\ntp_labels = []<br>\nhop_length = 128<br>\nn_fft = 512<br>\nfor i in tqdm(range(len(train_tp))):<br>\n    idx = train_tp['recording_id'][i]<br>\n    file_path = '../input/rfcx-species-audio-detection/train/' + idx + '.flac'<br>\n    signal, sr = sf.read(file_path)<br>\n    m = np.abs(np.fft.fft(signal)[:len(signal)//2])<br>\n    tp.append(m)<br>\n    tp_labels.append(train_tp['species_id'][i])</p>",
      "rawMarkdown": "Hey Guys, I tried to run the code stated below for this competition but the problem I encountered is that even before completing 1216 files my CPU RAM overflows and Session Restarts. \nAny Sugesstions?\n\ntp = []\ntp_labels = []\nhop_length = 128\nn_fft = 512\nfor i in tqdm(range(len(train_tp))):\n    idx = train_tp['recording_id'][i]\n    file_path = '../input/rfcx-species-audio-detection/train/' + idx + '.flac'\n    signal, sr = sf.read(file_path)\n    m = np.abs(np.fft.fft(signal)[:len(signal)//2])\n    tp.append(m)\n    tp_labels.append(train_tp['species_id'][i])",
      "votes": null
    },
    {
      "id": "1103001",
      "postDate": "12/05/2020 14:55:30",
      "content": "<p>There is a lot of data to store everything in variables and your RAM. I think you could do one of two things to solve this issue:</p>\n<p>1- Don't save all your FFT data in variables. Create them and save all in your HD. Then access as needed.</p>\n<p>2- Don't create and save the FFTs before, but as needed in the training, doing the FFT for each audio your are using in training.</p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "There is a lot of data to store everything in variables and your RAM. I think you could do one of two things to solve this issue:\n\n1- Don't save all your FFT data in variables. Create them and save all in your HD. Then access as needed.\n\n2- Don't create and save the FFTs before, but as needed in the training, doing the FFT for each audio your are using in training.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "1103212",
      "postDate": "12/05/2020 18:13:21",
      "content": "<p>You can also see in this notebook of Giba how he does the FFT and reshapes it.</p>\n<p><a href=\"https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\" target=\"_blank\">https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast</a></p>\n<p>This could also work. </p>",
      "rawMarkdown": "You can also see in this notebook of Giba how he does the FFT and reshapes it.\n\nhttps://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\n\nThis could also work.",
      "votes": null
    },
    {
      "id": "1103638",
      "postDate": "12/06/2020 04:56:32",
      "content": "<p>I think that would definitely help.<br>\nThanks!!</p>",
      "rawMarkdown": "I think that would definitely help.\nThanks!!",
      "votes": null
    },
    {
      "id": "1104210",
      "postDate": "12/06/2020 17:50:53",
      "content": "<p>Thanks for mentioning the notebook, I run it and it worked completely fine. <br>\nBut then arises another question that the method this guy used in his notebook hardly differs from mine, like he used function calls to implement the same think and I used it directly. <br>\nAm I going wrong with something here? <br>\nHow does both method differs?<br>\nIs it because he used RAPIDS in his notebook?</p>",
      "rawMarkdown": "Thanks for mentioning the notebook, I run it and it worked completely fine. \nBut then arises another question that the method this guy used in his notebook hardly differs from mine, like he used function calls to implement the same think and I used it directly. \nAm I going wrong with something here? \nHow does both method differs?\nIs it because he used RAPIDS in his notebook?",
      "votes": null
    },
    {
      "id": "1104387",
      "postDate": "12/06/2020 22:09:43",
      "content": "<p>The notebook in question averages over every 1440 features and just keeps the average 1000 features per fft. That's why it fits into memory. </p>",
      "rawMarkdown": "The notebook in question averages over every 1440 features and just keeps the average 1000 features per fft. That's why it fits into memory.",
      "votes": null
    },
    {
      "id": "1104948",
      "postDate": "12/07/2020 11:52:17",
      "content": "<p>As <a href=\"https://www.kaggle.com/joergdietrich\" target=\"_blank\">@joergdietrich</a> said, the reason is because he does his FFTs smallers by reducing the features. That's why he could store all in RAM, but that's up to you what path to choose.</p>",
      "rawMarkdown": "As @joergdietrich said, the reason is because he does his FFTs smallers by reducing the features. That's why he could store all in RAM, but that's up to you what path to choose.",
      "votes": null
    },
    {
      "id": "1104980",
      "postDate": "12/07/2020 12:26:10",
      "content": "<p>You mean the line of code where he took mean across the column axis. Due to that RAM was utilized efficiently?</p>",
      "rawMarkdown": "You mean the line of code where he took mean across the column axis. Due to that RAM was utilized efficiently?",
      "votes": null
    },
    {
      "id": "1104989",
      "postDate": "12/07/2020 12:40:55",
      "content": "<p>Yes, where he did the average.. You can use his \"extract_fft\" function, as it is ready and working. As explained by Giba:</p>\n<blockquote>\n  <p>To reduce the dimension of original array. Instead of using 1440000 features, I take the average every 1440 features, decreasing the dimension to 1000.</p>\n  <p>similar of doing:<br>\n  result = np.zeros(1000)<br>\n  result[0] = np.mean( varfft[0:1440] )<br>\n  result[1] = np.mean( varfft[1440:2880] )<br>\n  …<br>\n  result[999] = np.mean( varfft[-1440:] )</p>\n</blockquote>",
      "rawMarkdown": "Yes, where he did the average.. You can use his \"extract_fft\" function, as it is ready and working. As explained by Giba:\n\n> To reduce the dimension of original array. Instead of using 1440000 features, I take the average every 1440 features, decreasing the dimension to 1000.\n\n> similar of doing:\nresult = np.zeros(1000)\nresult[0] = np.mean( varfft[0:1440] )\nresult[1] = np.mean( varfft[1440:2880] )\n…\nresult[999] = np.mean( varfft[-1440:] )",
      "votes": null
    },
    {
      "id": "1104995",
      "postDate": "12/07/2020 12:51:29",
      "content": "<p>Right, Now I got it. Thank You very much!!</p>",
      "rawMarkdown": "Right, Now I got it. Thank You very much!!",
      "votes": null
    },
    {
      "id": "1105006",
      "postDate": "12/07/2020 12:57:55",
      "content": "<p>You're welcome!</p>",
      "rawMarkdown": "You're welcome!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1103001,
      "author_name": "willsoares1",
      "author_url": "",
      "post_date": "12/05/2020 14:55:30",
      "content": "<p>There is a lot of data to store everything in variables and your RAM. I think you could do one of two things to solve this issue:</p>\n<p>1- Don't save all your FFT data in variables. Create them and save all in your HD. Then access as needed.</p>\n<p>2- Don't create and save the FFTs before, but as needed in the training, doing the FFT for each audio your are using in training.</p>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103638,
          "author_name": "utkarshg007",
          "author_url": "",
          "post_date": "12/06/2020 04:56:32",
          "content": "<p>I think that would definitely help.<br>\nThanks!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103212,
      "author_name": "willsoares1",
      "author_url": "",
      "post_date": "12/05/2020 18:13:21",
      "content": "<p>You can also see in this notebook of Giba how he does the FFT and reshapes it.</p>\n<p><a href=\"https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\" target=\"_blank\">https://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast</a></p>\n<p>This could also work. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1104210,
          "author_name": "utkarshg007",
          "author_url": "",
          "post_date": "12/06/2020 17:50:53",
          "content": "<p>Thanks for mentioning the notebook, I run it and it worked completely fine. <br>\nBut then arises another question that the method this guy used in his notebook hardly differs from mine, like he used function calls to implement the same think and I used it directly. <br>\nAm I going wrong with something here? <br>\nHow does both method differs?<br>\nIs it because he used RAPIDS in his notebook?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104387,
          "author_name": "joergdietrich",
          "author_url": "",
          "post_date": "12/06/2020 22:09:43",
          "content": "<p>The notebook in question averages over every 1440 features and just keeps the average 1000 features per fft. That's why it fits into memory. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104948,
          "author_name": "willsoares1",
          "author_url": "",
          "post_date": "12/07/2020 11:52:17",
          "content": "<p>As <a href=\"https://www.kaggle.com/joergdietrich\" target=\"_blank\">@joergdietrich</a> said, the reason is because he does his FFTs smallers by reducing the features. That's why he could store all in RAM, but that's up to you what path to choose.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104980,
          "author_name": "utkarshg007",
          "author_url": "",
          "post_date": "12/07/2020 12:26:10",
          "content": "<p>You mean the line of code where he took mean across the column axis. Due to that RAM was utilized efficiently?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104989,
          "author_name": "willsoares1",
          "author_url": "",
          "post_date": "12/07/2020 12:40:55",
          "content": "<p>Yes, where he did the average.. You can use his \"extract_fft\" function, as it is ready and working. As explained by Giba:</p>\n<blockquote>\n  <p>To reduce the dimension of original array. Instead of using 1440000 features, I take the average every 1440 features, decreasing the dimension to 1000.</p>\n  <p>similar of doing:<br>\n  result = np.zeros(1000)<br>\n  result[0] = np.mean( varfft[0:1440] )<br>\n  result[1] = np.mean( varfft[1440:2880] )<br>\n  …<br>\n  result[999] = np.mean( varfft[-1440:] )</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104995,
          "author_name": "utkarshg007",
          "author_url": "",
          "post_date": "12/07/2020 12:51:29",
          "content": "<p>Right, Now I got it. Thank You very much!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105006,
          "author_name": "willsoares1",
          "author_url": "",
          "post_date": "12/07/2020 12:57:55",
          "content": "<p>You're welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1102949": "Hey Guys, I tried to run the code stated below for this competition but the problem I encountered is that even before completing 1216 files my CPU RAM overflows and Session Restarts. \nAny Sugesstions?\n\ntp = []\ntp_labels = []\nhop_length = 128\nn_fft = 512\nfor i in tqdm(range(len(train_tp))):\n    idx = train_tp['recording_id'][i]\n    file_path = '../input/rfcx-species-audio-detection/train/' + idx + '.flac'\n    signal, sr = sf.read(file_path)\n    m = np.abs(np.fft.fft(signal)[:len(signal)//2])\n    tp.append(m)\n    tp_labels.append(train_tp['species_id'][i])",
    "1103001": "There is a lot of data to store everything in variables and your RAM. I think you could do one of two things to solve this issue:\n\n1- Don't save all your FFT data in variables. Create them and save all in your HD. Then access as needed.\n\n2- Don't create and save the FFTs before, but as needed in the training, doing the FFT for each audio your are using in training.\n\nHope this helps!",
    "1103212": "You can also see in this notebook of Giba how he does the FFT and reshapes it.\n\nhttps://www.kaggle.com/titericz/0-525-tabular-xgboost-gpu-fft-gpu-cuml-fast\n\nThis could also work.",
    "1103638": "I think that would definitely help.\nThanks!!",
    "1104210": "Thanks for mentioning the notebook, I run it and it worked completely fine. \nBut then arises another question that the method this guy used in his notebook hardly differs from mine, like he used function calls to implement the same think and I used it directly. \nAm I going wrong with something here? \nHow does both method differs?\nIs it because he used RAPIDS in his notebook?",
    "1104387": "The notebook in question averages over every 1440 features and just keeps the average 1000 features per fft. That's why it fits into memory.",
    "1104948": "As @joergdietrich said, the reason is because he does his FFTs smallers by reducing the features. That's why he could store all in RAM, but that's up to you what path to choose.",
    "1104980": "You mean the line of code where he took mean across the column axis. Due to that RAM was utilized efficiently?",
    "1104989": "Yes, where he did the average.. You can use his \"extract_fft\" function, as it is ready and working. As explained by Giba:\n\n> To reduce the dimension of original array. Instead of using 1440000 features, I take the average every 1440 features, decreasing the dimension to 1000.\n\n> similar of doing:\nresult = np.zeros(1000)\nresult[0] = np.mean( varfft[0:1440] )\nresult[1] = np.mean( varfft[1440:2880] )\n…\nresult[999] = np.mean( varfft[-1440:] )",
    "1104995": "Right, Now I got it. Thank You very much!!",
    "1105006": "You're welcome!"
  },
  "source": "meta"
}