{
  "id": 164164,
  "title": "SpeedUp Training....!",
  "url": "/competitions/birdsong-recognition/discussion/164164",
  "author_name": "Gopi Durgaprasad",
  "post_date": "2020-07-05T02:49:25.775000",
  "votes": -2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hai...,\nI tried different approaches for training, but the best it takes a min 1hour  for 1epoch.</p>\n\n<h3>Strategy 1</h3>\n\n<ul>\n<li><p>convert all .mp3 to .wav with sampling rate 16000\n<strong>failed :</strong> becasues of audio files are different lengths</p>\n\n<h3>Strategy 2</h3></li>\n<li><p>convert all .mp3 to .wav with sampling rate 44100\n<strong>failed :</strong> because of audio files are different lengths</p>\n\n<h3>Strategy 3</h3></li>\n<li><p>convert all .mp3 files into melspectrograms before traing\n<strong>failed :</strong> because of spectrogram lengths are different</p>\n\n<h3>Strategy 4</h3></li>\n<li><p>convert every file into 10sec interval and save as npy files\n<strong>failed :</strong> I don't know why its fail</p></li>\n</ul>\n\n<p><strong>Any suggestions</strong>\nThank you,</p>",
  "messages": [
    {
      "id": 915798,
      "postDate": "2020-07-05T05:20:10.213Z",
      "content": "<p>My current setup is fast enough (nearly 30 sec per epoch for complete dataset) I have saved melspectrograms as .npy files. and to handle different lengths checkout custom collect_fn function for pytorch dataloader. You can refer to my sample notebook <a href=\"https://www.kaggle.com/dhananjay3/simple-pytorch-starter\">here.</a></p>",
      "rawMarkdown": "My current setup is fast enough (nearly 30 sec per epoch for complete dataset) I have saved melspectrograms as .npy files. and to handle different lengths checkout custom collect_fn function for pytorch dataloader. You can refer to my sample notebook [here.](https://www.kaggle.com/dhananjay3/simple-pytorch-starter)",
      "votes": 1,
      "replies": [
        {
          "id": 915872,
          "postDate": "2020-07-05T07:05:55.843Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 916219,
          "postDate": "2020-07-05T13:32:06.300Z",
          "content": "<p>I guess it's 30 seconds because you took only 10 classes out of the approx 250 classes (to test the notebook), thus reducing the dataset size to more or less 800 elements...\nBTW I liked the idea of your collect_fn ;)</p>",
          "rawMarkdown": "I guess it's 30 seconds because you took only 10 classes out of the approx 250 classes (to test the notebook), thus reducing the dataset size to more or less 800 elements...\nBTW I liked the idea of your collect\\_fn ;)"
        },
        {
          "id": 916727,
          "postDate": "2020-07-06T01:11:07.637Z",
          "content": "<p><a href=\"/pranavkasela\">@pranavkasela</a> I am talking about complete dataset similar to that notebook. i.e. I saved 21k .npy files and loaded all of them in 30 seconds.  In the notebook 30 sec include training time (on cpu)  but the actual dataloading time is much lower for 10 species sample.</p>",
          "rawMarkdown": "@pranavkasela I am talking about complete dataset similar to that notebook. i.e. I saved 21k .npy files and loaded all of them in 30 seconds.  In the notebook 30 sec include training time (on cpu)  but the actual dataloading time is much lower for 10 species sample.",
          "votes": 1
        },
        {
          "id": 917080,
          "postDate": "2020-07-06T08:14:09.050Z",
          "content": "<p>Aah I thought you completed one epoch of training in less than 30 seconds on a CPU with all the data, you were talking about the loading time, my bad...</p>",
          "rawMarkdown": "Aah I thought you completed one epoch of training in less than 30 seconds on a CPU with all the data, you were talking about the loading time, my bad..."
        }
      ]
    },
    {
      "id": 915685,
      "postDate": "2020-07-05T02:49:25.777Z",
      "content": "<p>Hai...,\nI tried different approaches for training, but the best it takes a min 1hour  for 1epoch.</p>\n\n<h3>Strategy 1</h3>\n\n<ul>\n<li><p>convert all .mp3 to .wav with sampling rate 16000\n<strong>failed :</strong> becasues of audio files are different lengths</p>\n\n<h3>Strategy 2</h3></li>\n<li><p>convert all .mp3 to .wav with sampling rate 44100\n<strong>failed :</strong> because of audio files are different lengths</p>\n\n<h3>Strategy 3</h3></li>\n<li><p>convert all .mp3 files into melspectrograms before traing\n<strong>failed :</strong> because of spectrogram lengths are different</p>\n\n<h3>Strategy 4</h3></li>\n<li><p>convert every file into 10sec interval and save as npy files\n<strong>failed :</strong> I don't know why its fail</p></li>\n</ul>\n\n<p><strong>Any suggestions</strong>\nThank you,</p>",
      "rawMarkdown": "Hai...,\nI tried different approaches for training, but the best it takes a min 1hour  for 1epoch.\n### Strategy 1\n- convert all .mp3 to .wav with sampling rate 16000\n**failed :** becasues of audio files are different lengths\n### Strategy 2\n- convert all .mp3 to .wav with sampling rate 44100\n**failed :** because of audio files are different lengths\n### Strategy 3\n- convert all .mp3 files into melspectrograms before traing\n**failed :** because of spectrogram lengths are different\n### Strategy 4\n- convert every file into 10sec interval and save as npy files\n**failed :** I don't know why its fail\n\n**Any suggestions**\nThank you,",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 915798,
      "author_name": "Dhananjay Raut",
      "author_url": "",
      "post_date": "2020-07-05T05:20:10.213000",
      "content": "<p>My current setup is fast enough (nearly 30 sec per epoch for complete dataset) I have saved melspectrograms as .npy files. and to handle different lengths checkout custom collect_fn function for pytorch dataloader. You can refer to my sample notebook <a href=\"https://www.kaggle.com/dhananjay3/simple-pytorch-starter\">here.</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 915872,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-05T07:05:55.843000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916219,
          "author_name": "Pranav Kasela",
          "author_url": "",
          "post_date": "2020-07-05T13:32:06.300000",
          "content": "<p>I guess it's 30 seconds because you took only 10 classes out of the approx 250 classes (to test the notebook), thus reducing the dataset size to more or less 800 elements...\nBTW I liked the idea of your collect_fn ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916727,
          "author_name": "Dhananjay Raut",
          "author_url": "",
          "post_date": "2020-07-06T01:11:07.637000",
          "content": "<p><a href=\"/pranavkasela\">@pranavkasela</a> I am talking about complete dataset similar to that notebook. i.e. I saved 21k .npy files and loaded all of them in 30 seconds.  In the notebook 30 sec include training time (on cpu)  but the actual dataloading time is much lower for 10 species sample.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 917080,
          "author_name": "Pranav Kasela",
          "author_url": "",
          "post_date": "2020-07-06T08:14:09.050000",
          "content": "<p>Aah I thought you completed one epoch of training in less than 30 seconds on a CPU with all the data, you were talking about the loading time, my bad...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "915798": "My current setup is fast enough (nearly 30 sec per epoch for complete dataset) I have saved melspectrograms as .npy files. and to handle different lengths checkout custom collect_fn function for pytorch dataloader. You can refer to my sample notebook [here.](https://www.kaggle.com/dhananjay3/simple-pytorch-starter)",
    "915685": "Hai...,\nI tried different approaches for training, but the best it takes a min 1hour  for 1epoch.\n### Strategy 1\n- convert all .mp3 to .wav with sampling rate 16000\n**failed :** becasues of audio files are different lengths\n### Strategy 2\n- convert all .mp3 to .wav with sampling rate 44100\n**failed :** because of audio files are different lengths\n### Strategy 3\n- convert all .mp3 files into melspectrograms before traing\n**failed :** because of spectrogram lengths are different\n### Strategy 4\n- convert every file into 10sec interval and save as npy files\n**failed :** I don't know why its fail\n\n**Any suggestions**\nThank you,"
  }
}