{
  "id": 57597,
  "title": "(Help Needed) Preprocessing variable length audio data for FCN - without padding or cropping",
  "url": "/competitions/freesound-audio-tagging/discussion/57597",
  "author_name": "",
  "post_date": "2018-05-25T20:33:55.329413400Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm currently trying to implement a FCN using Keras to classify audio. So far I've created a basic CNN which takes the training set and either crop each sample to 2 second or pads it. However, I was reading a paper (<a href=\"https://arxiv.org/pdf/1707.02530.pdf\">https://arxiv.org/pdf/1707.02530.pdf</a>) in the paper they build an FCN which allows variable length input. I have created such a FCN using Keras and has a SpatialPyramidPooling layer.</p>\n\n<pre><code>batch_size = 64\nnum_channels = 3\nnum_classes = 10\n\nmodel = Sequential()\n\n# uses theano ordering. Note that we leave the image size as None to allow multiple image sizes\nmodel.add(Convolution2D(32, 3, 3, border_mode='same', input_shape=(3, None, None)))\nmodel.add(Activation('relu'))\nmodel.add(Convolution2D(32, 3, 3))\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\nmodel.add(Convolution2D(64, 3, 3, border_mode='same'))\nmodel.add(Activation('relu'))\nmodel.add(Convolution2D(64, 3, 3))\nmodel.add(Activation('relu'))\nmodel.add(SpatialPyramidPooling([1, 2, 4]))\nmodel.add(Dense(num_classes))\nmodel.add(Activation('softmax'))\nheremodel.compile(loss='categorical_crossentropy', optimizer='sgd')\n</code></pre>\n\n<p>However, I cannot seem to figure how to pre=process the audio in order for to be input into the FCN. Here is my function to prepare the data:</p>\n\n<pre><code>def prepare_data(df, config, data_dir, bands=128):\nx = np.empty(shape=(len(df.index), bands))\nlog_specgrams_2048 = []\nfor i, fname in enumerate(df.index):\n    file_path = data_dir + fname\n    data, _ = librosa.core.load(file_path, sr=16000, res_type=\"kaiser_fast\")\n    melspec = librosa.feature.melspectrogram(data, sr=16000, n_mels=bands)\n    logspec = librosa.core.power_to_db(melspec)\n    log_specgrams_2048.append(logspec)\nreturn log_specgrams_2048\n</code></pre>\n\n<p>And here is my train function:</p>\n\n<pre><code>  def run(config):\n    test = pd.read_csv(\"../input/sample_submission.csv\")\n    train = pd.read_csv(\"../input/train_1.csv\")\n\n    LABELS = list(train.label.unique())\n    label_idx = {label: i for i, label in enumerate(LABELS)}\n\n    train.set_index('fname', inplace=True)\n    test.set_index('fname', inplace=True)\n    train['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n\n    x_train = prepare_data(train, config, '../input/audio_train_1/')\n    # x_test = prepare_data(train, config, '../input/audio_test/')\n\n    y_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n\n    skf = StratifiedKFold(n_splits=2).split(np.zeros(len(train)), train.label_idx)\n    for i, (train_split, val_split) in enumerate(skf):\n        K.clear_session()\n        x, y, x_val, y_val = np.array(x_train[train_split]), y_train[train_split], np.array(x_train[val_split]), y_train[val_split]\n\n        checkpoint = ModelCheckpoint(config.job_dir + '/best_%d.h5' % i, monitor='val_loss', verbose=1,\n                                     save_best_only=True)\n        early = EarlyStopping(monitor='val_loss', mode='min', patience=5)\n        tb = TensorBoard(log_dir=os.path.join(config.job_dir, 'logs') + '/fold_%i' % i, write_graph=True)\n        callbacks_list = [checkpoint, early, tb]\n\n        print(('Fold: %d' % i) + '\\n' + '#' * 50)\n\n        curr_model = model_fn_aes(x_train.shape)\n\n        curr_model.fit(x, y, validation_data=(x_val, y_val), callbacks=callbacks_list,\n                       batch_size=64, epochs=config.max_epochs)\n\n        curr_model.load_weights(config.job_dir + '/best_%d.h5' % i)\n</code></pre>\n\n<p>I would greatly appreciate it if someone could help me fix the prepare_data function so that I can load the training set into an np array and input into the FCN whilst still using StratifiedKFold.</p>",
  "messages": [
    {
      "id": "333763",
      "postDate": "05/25/2018 20:33:55",
      "content": "<p>I'm currently trying to implement a FCN using Keras to classify audio. So far I've created a basic CNN which takes the training set and either crop each sample to 2 second or pads it. However, I was reading a paper (<a href=\"https://arxiv.org/pdf/1707.02530.pdf\">https://arxiv.org/pdf/1707.02530.pdf</a>) in the paper they build an FCN which allows variable length input. I have created such a FCN using Keras and has a SpatialPyramidPooling layer.</p>\n\n<pre><code>batch_size = 64\nnum_channels = 3\nnum_classes = 10\n\nmodel = Sequential()\n\n# uses theano ordering. Note that we leave the image size as None to allow multiple image sizes\nmodel.add(Convolution2D(32, 3, 3, border_mode='same', input_shape=(3, None, None)))\nmodel.add(Activation('relu'))\nmodel.add(Convolution2D(32, 3, 3))\nmodel.add(Activation('relu'))\nmodel.add(MaxPooling2D(pool_size=(2, 2)))\nmodel.add(Convolution2D(64, 3, 3, border_mode='same'))\nmodel.add(Activation('relu'))\nmodel.add(Convolution2D(64, 3, 3))\nmodel.add(Activation('relu'))\nmodel.add(SpatialPyramidPooling([1, 2, 4]))\nmodel.add(Dense(num_classes))\nmodel.add(Activation('softmax'))\nheremodel.compile(loss='categorical_crossentropy', optimizer='sgd')\n</code></pre>\n\n<p>However, I cannot seem to figure how to pre=process the audio in order for to be input into the FCN. Here is my function to prepare the data:</p>\n\n<pre><code>def prepare_data(df, config, data_dir, bands=128):\nx = np.empty(shape=(len(df.index), bands))\nlog_specgrams_2048 = []\nfor i, fname in enumerate(df.index):\n    file_path = data_dir + fname\n    data, _ = librosa.core.load(file_path, sr=16000, res_type=\"kaiser_fast\")\n    melspec = librosa.feature.melspectrogram(data, sr=16000, n_mels=bands)\n    logspec = librosa.core.power_to_db(melspec)\n    log_specgrams_2048.append(logspec)\nreturn log_specgrams_2048\n</code></pre>\n\n<p>And here is my train function:</p>\n\n<pre><code>  def run(config):\n    test = pd.read_csv(\"../input/sample_submission.csv\")\n    train = pd.read_csv(\"../input/train_1.csv\")\n\n    LABELS = list(train.label.unique())\n    label_idx = {label: i for i, label in enumerate(LABELS)}\n\n    train.set_index('fname', inplace=True)\n    test.set_index('fname', inplace=True)\n    train['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n\n    x_train = prepare_data(train, config, '../input/audio_train_1/')\n    # x_test = prepare_data(train, config, '../input/audio_test/')\n\n    y_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n\n    skf = StratifiedKFold(n_splits=2).split(np.zeros(len(train)), train.label_idx)\n    for i, (train_split, val_split) in enumerate(skf):\n        K.clear_session()\n        x, y, x_val, y_val = np.array(x_train[train_split]), y_train[train_split], np.array(x_train[val_split]), y_train[val_split]\n\n        checkpoint = ModelCheckpoint(config.job_dir + '/best_%d.h5' % i, monitor='val_loss', verbose=1,\n                                     save_best_only=True)\n        early = EarlyStopping(monitor='val_loss', mode='min', patience=5)\n        tb = TensorBoard(log_dir=os.path.join(config.job_dir, 'logs') + '/fold_%i' % i, write_graph=True)\n        callbacks_list = [checkpoint, early, tb]\n\n        print(('Fold: %d' % i) + '\\n' + '#' * 50)\n\n        curr_model = model_fn_aes(x_train.shape)\n\n        curr_model.fit(x, y, validation_data=(x_val, y_val), callbacks=callbacks_list,\n                       batch_size=64, epochs=config.max_epochs)\n\n        curr_model.load_weights(config.job_dir + '/best_%d.h5' % i)\n</code></pre>\n\n<p>I would greatly appreciate it if someone could help me fix the prepare_data function so that I can load the training set into an np array and input into the FCN whilst still using StratifiedKFold.</p>",
      "rawMarkdown": "I'm currently trying to implement a FCN using Keras to classify audio. So far I've created a basic CNN which takes the training set and either crop each sample to 2 second or pads it. However, I was reading a paper (https://arxiv.org/pdf/1707.02530.pdf) in the paper they build an FCN which allows variable length input. I have created such a FCN using Keras and has a SpatialPyramidPooling layer.\n\n    batch_size = 64\n    num_channels = 3\n    num_classes = 10\n    \n    model = Sequential()\n    \n    # uses theano ordering. Note that we leave the image size as None to allow multiple image sizes\n    model.add(Convolution2D(32, 3, 3, border_mode='same', input_shape=(3, None, None)))\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3))\n    model.add(Activation('relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    model.add(Convolution2D(64, 3, 3, border_mode='same'))\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3))\n    model.add(Activation('relu'))\n    model.add(SpatialPyramidPooling([1, 2, 4]))\n    model.add(Dense(num_classes))\n    model.add(Activation('softmax'))\n    heremodel.compile(loss='categorical_crossentropy', optimizer='sgd')\n\nHowever, I cannot seem to figure how to pre=process the audio in order for to be input into the FCN. Here is my function to prepare the data:\n\n    def prepare_data(df, config, data_dir, bands=128):\n    x = np.empty(shape=(len(df.index), bands))\n    log_specgrams_2048 = []\n    for i, fname in enumerate(df.index):\n        file_path = data_dir + fname\n        data, _ = librosa.core.load(file_path, sr=16000, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr=16000, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec)\n        log_specgrams_2048.append(logspec)\n    return log_specgrams_2048\n\n\nAnd here is my train function:\n\n      def run(config):\n        test = pd.read_csv(\"../input/sample_submission.csv\")\n        train = pd.read_csv(\"../input/train_1.csv\")\n        \n        LABELS = list(train.label.unique())\n        label_idx = {label: i for i, label in enumerate(LABELS)}\n        \n        train.set_index('fname', inplace=True)\n        test.set_index('fname', inplace=True)\n        train['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n        \n        x_train = prepare_data(train, config, '../input/audio_train_1/')\n        # x_test = prepare_data(train, config, '../input/audio_test/')\n        \n        y_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n        \n        skf = StratifiedKFold(n_splits=2).split(np.zeros(len(train)), train.label_idx)\n        for i, (train_split, val_split) in enumerate(skf):\n            K.clear_session()\n            x, y, x_val, y_val = np.array(x_train[train_split]), y_train[train_split], np.array(x_train[val_split]), y_train[val_split]\n        \n            checkpoint = ModelCheckpoint(config.job_dir + '/best_%d.h5' % i, monitor='val_loss', verbose=1,\n                                         save_best_only=True)\n            early = EarlyStopping(monitor='val_loss', mode='min', patience=5)\n            tb = TensorBoard(log_dir=os.path.join(config.job_dir, 'logs') + '/fold_%i' % i, write_graph=True)\n            callbacks_list = [checkpoint, early, tb]\n        \n            print(('Fold: %d' % i) + '\\n' + '#' * 50)\n        \n            curr_model = model_fn_aes(x_train.shape)\n        \n            curr_model.fit(x, y, validation_data=(x_val, y_val), callbacks=callbacks_list,\n                           batch_size=64, epochs=config.max_epochs)\n        \n            curr_model.load_weights(config.job_dir + '/best_%d.h5' % i)\n\nI would greatly appreciate it if someone could help me fix the prepare_data function so that I can load the training set into an np array and input into the FCN whilst still using StratifiedKFold.",
      "votes": null
    },
    {
      "id": "334135",
      "postDate": "05/26/2018 16:08:30",
      "content": "<p>Hi Henry,</p>\n\n<p>I come across this Japanese tech blog <a href=\"https://qiita.com/MuAuan/items/dc819c17bdb030c0e096\">SPP（SpatialPyramidPooling）for free size input...</a>.</p>\n\n<p>And it says that input_shape could be different order; change it from <code>(3, None, None)</code> to <code>( None, None, 3)</code>.</p>\n\n<p>But the blog handles image data. yours is 1 channel audio; 3 could be 1. And I think the paper seems to use 44,100Hz instead of 16,000Hz. Then I <em>guess</em>:</p>\n\n<ol>\n<li>Change <code>(3, None, None)</code> to <code>( None, None, 1)</code>.</li>\n<li>prepare_data could be like this:</li>\n</ol>\n\n<pre><code>sr = 44100\ndef prepare_data(df, config, data_dir, bands=128):\n    x = np.empty(shape=(len(df.index), bands))\n    log_specgrams_2048 = []\n    for i, fname in enumerate(df.index):\n        file_path = data_dir + fname\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr= sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        log_specgrams_2048.append(logspec)\n    return log_specgrams_2048\n</code></pre>\n\n<p>I hope it help some, anyway thanks for sharing great idea.</p>",
      "rawMarkdown": "Hi Henry,\n\nI come across this Japanese tech blog [SPP（SpatialPyramidPooling）for free size input...][1].\n\nAnd it says that input_shape could be different order; change it from `(3, None, None)` to `( None, None, 3)`.\n\nBut the blog handles image data. yours is 1 channel audio; 3 could be 1. And I think the paper seems to use 44,100Hz instead of 16,000Hz. Then I _guess_:\n\n1. Change `(3, None, None)` to `( None, None, 1)`.\n2. prepare_data could be like this:\n\n<pre><code>sr = 44100\ndef prepare_data(df, config, data_dir, bands=128):\n    x = np.empty(shape=(len(df.index), bands))\n    log_specgrams_2048 = []\n    for i, fname in enumerate(df.index):\n        file_path = data_dir + fname\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr= sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        log_specgrams_2048.append(logspec)\n    return log_specgrams_2048\n</code></pre>\n\nI hope it help some, anyway thanks for sharing great idea.\n\n  [1]: https://qiita.com/MuAuan/items/dc819c17bdb030c0e096",
      "votes": null
    },
    {
      "id": "336103",
      "postDate": "05/31/2018 04:17:27",
      "content": "<p>Hi daisukelab, </p>\n\n<p>Thank you very mcuh for your help! I just had another question, I've changed my run function to:</p>\n\n<pre><code>test = pd.read_csv(\"../input/sample_submission.csv\")\ntrain = pd.read_csv(\"../input/train_1.csv\")\n\nLABELS = list(train.label.unique())\nlabel_idx = {label: i for i, label in enumerate(LABELS)}\n\ntrain.set_index('fname', inplace=True)\ntest.set_index('fname', inplace=True)\ntrain['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n\nx_train = prepare_data(train, config, '../input/audio_train/')\nx_test = prepare_data(train, config, '../input/audio_test/')\n\ny_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n\nX_train, X_test, y_train, y_test = train_test_split(x_train, train.label_idx.values,\n                                                    test_size=0.2)\n\nmodel = model_fn_aes(config)\ncallbacks_list = []\nmodel.fit(X_train, np.array(y_train), validation_data=(X_test, np.array(y_test)), epochs=50, callbacks=callbacks_list)\n</code></pre>\n\n<p>Becuase I couldnt seem to get the StratifiedKFold function to work with regular lists. However I still cannot get the Keras fit function to work with the list or arrays and get the following error:\nExpected to see 1 array(s), but instead got the following list of 13 arrays</p>\n\n<p>I has hoping you might be able to help me get around this. Again I greatly appriciate your help!</p>",
      "rawMarkdown": "Hi daisukelab, \n\n\nThank you very mcuh for your help! I just had another question, I've changed my run function to:\n\n    test = pd.read_csv(\"../input/sample_submission.csv\")\n    train = pd.read_csv(\"../input/train_1.csv\")\n\n    LABELS = list(train.label.unique())\n    label_idx = {label: i for i, label in enumerate(LABELS)}\n\n    train.set_index('fname', inplace=True)\n    test.set_index('fname', inplace=True)\n    train['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n\n    x_train = prepare_data(train, config, '../input/audio_train/')\n    x_test = prepare_data(train, config, '../input/audio_test/')\n\n    y_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n\n    X_train, X_test, y_train, y_test = train_test_split(x_train, train.label_idx.values,\n                                                        test_size=0.2)\n\n    model = model_fn_aes(config)\n    callbacks_list = []\n    model.fit(X_train, np.array(y_train), validation_data=(X_test, np.array(y_test)), epochs=50, callbacks=callbacks_list)\n\n\nBecuase I couldnt seem to get the StratifiedKFold function to work with regular lists. However I still cannot get the Keras fit function to work with the list or arrays and get the following error:\nExpected to see 1 array(s), but instead got the following list of 13 arrays\n\nI has hoping you might be able to help me get around this. Again I greatly appriciate your help!",
      "votes": null
    },
    {
      "id": "336160",
      "postDate": "05/31/2018 06:45:47",
      "content": "<p>Hi Henry,</p>\n\n<p>I should share my conclusion that we cannot almost exploit benefit from SpatialPyramidPooling.</p>\n\n<ul>\n<li>What was created by prepare_data is a list that holds np.array with different length.</li>\n<li>Keras fit() accepts numpy array, not list. fit() will get to know what is the shape of given input, but it cannot get it from list; we have to feed numpy array.</li>\n<li>Numpy array cannot have variable length multi dimensional array.</li>\n<li>These restrictions prevent us from training with variable length input.</li>\n<li>Model itself would work fine with variable test input, if once it have trained with fixed length training inputs, but we cannot train by aforementioned reasons.</li>\n</ul>\n\n<p>You can start training by making all the training data into unified length by padding as following example:</p>\n\n<pre><code>def prepare_data(df, data_dir, sr=44100, bands=128):\n    X = []\n    # Convert once\n    for i, fname in enumerate(df.fname):\n        file_path = os.path.join(data_dir, fname)\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr=sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        X.append(logspec)\n    # Find longest\n    max_length = np.max([x.shape[1] for x in X])\n    # Pad zero to make them all the same length\n    X2 = [np.pad(x, ((0, 0), (0, max_length - x.shape[1]), (0, 0)), 'constant') for x in X]\n    return np.array(X2)</code></pre>\n\n<p>We still have way to make generator for feeding fixed-length-data-for-batch but variable-for-entire-training-set, but it takes time for both implementing and training.\nI personally abandon attempt for this approach...</p>",
      "rawMarkdown": "Hi Henry,\n\nI should share my conclusion that we cannot almost exploit benefit from SpatialPyramidPooling.\n\n- What was created by prepare_data is a list that holds np.array with different length.\n- Keras fit() accepts numpy array, not list. fit() will get to know what is the shape of given input, but it cannot get it from list; we have to feed numpy array.\n- Numpy array cannot have variable length multi dimensional array.\n- These restrictions prevent us from training with variable length input.\n- Model itself would work fine with variable test input, if once it have trained with fixed length training inputs, but we cannot train by aforementioned reasons.\n\nYou can start training by making all the training data into unified length by padding as following example:\n\n<pre><code>def prepare_data(df, data_dir, sr=44100, bands=128):\n    X = []\n    # Convert once\n    for i, fname in enumerate(df.fname):\n        file_path = os.path.join(data_dir, fname)\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr=sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        X.append(logspec)\n    # Find longest\n    max_length = np.max([x.shape[1] for x in X])\n    # Pad zero to make them all the same length\n    X2 = [np.pad(x, ((0, 0), (0, max_length - x.shape[1]), (0, 0)), 'constant') for x in X]\n    return np.array(X2)</code></pre>\n\nWe still have way to make generator for feeding fixed-length-data-for-batch but variable-for-entire-training-set, but it takes time for both implementing and training.\nI personally abandon attempt for this approach...",
      "votes": null
    },
    {
      "id": "336466",
      "postDate": "05/31/2018 19:28:21",
      "content": "<p>Thank you for your response again, I see the problem now, do you think it might be work setting the batch size to 1 or doing batched of similarly sized audio files, just some workaround I thought might work?​</p>",
      "rawMarkdown": "Thank you for your response again, I see the problem now, do you think it might be work setting the batch size to 1 or doing batched of similarly sized audio files, just some workaround I thought might work?​",
      "votes": null
    },
    {
      "id": "336479",
      "postDate": "05/31/2018 19:53:44",
      "content": "<p>Yes it would be possible, but:</p>\n\n<ul>\n<li>You will need to make simple data generator to feed [1, 128, varied-length, 1] shaped numpy array.</li>\n<li>It would take effort to tune LR, and also take very long time to train as explained in Deep Learning book by Goodfellow et al., see link below.</li>\n</ul>\n\n<p><a href=\"https://stackoverflow.com/questions/46654424/how-to-calculate-optimal-batch-size\">https://stackoverflow.com/questions/46654424/how-to-calculate-optimal-batch-size</a></p>\n\n<p>\"Small batches can offer a regularizing effect (Wilson and Martinez, 2003), perhaps due to the noise they add to the learning process. Generalization error is often best for a batch size of 1. Training with such a small batch size might require a small learning rate to maintain stability because of the high variance in the estimate of the gradient. The total runtime can be very high as a result of the need to make more steps, both because of the reduced learning rate and because it takes more steps to observe the entire training set.”</p>",
      "rawMarkdown": "Yes it would be possible, but:\n\n- You will need to make simple data generator to feed [1, 128, varied-length, 1] shaped numpy array.\n- It would take effort to tune LR, and also take very long time to train as explained in Deep Learning book by Goodfellow et al., see link below.\n\nhttps://stackoverflow.com/questions/46654424/how-to-calculate-optimal-batch-size\n\n\"Small batches can offer a regularizing effect (Wilson and Martinez, 2003), perhaps due to the noise they add to the learning process. Generalization error is often best for a batch size of 1. Training with such a small batch size might require a small learning rate to maintain stability because of the high variance in the estimate of the gradient. The total runtime can be very high as a result of the need to make more steps, both because of the reduced learning rate and because it takes more steps to observe the entire training set.”",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 334135,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "05/26/2018 16:08:30",
      "content": "<p>Hi Henry,</p>\n\n<p>I come across this Japanese tech blog <a href=\"https://qiita.com/MuAuan/items/dc819c17bdb030c0e096\">SPP（SpatialPyramidPooling）for free size input...</a>.</p>\n\n<p>And it says that input_shape could be different order; change it from <code>(3, None, None)</code> to <code>( None, None, 3)</code>.</p>\n\n<p>But the blog handles image data. yours is 1 channel audio; 3 could be 1. And I think the paper seems to use 44,100Hz instead of 16,000Hz. Then I <em>guess</em>:</p>\n\n<ol>\n<li>Change <code>(3, None, None)</code> to <code>( None, None, 1)</code>.</li>\n<li>prepare_data could be like this:</li>\n</ol>\n\n<pre><code>sr = 44100\ndef prepare_data(df, config, data_dir, bands=128):\n    x = np.empty(shape=(len(df.index), bands))\n    log_specgrams_2048 = []\n    for i, fname in enumerate(df.index):\n        file_path = data_dir + fname\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr= sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        log_specgrams_2048.append(logspec)\n    return log_specgrams_2048\n</code></pre>\n\n<p>I hope it help some, anyway thanks for sharing great idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 336103,
          "author_name": "hhargrea",
          "author_url": "",
          "post_date": "05/31/2018 04:17:27",
          "content": "<p>Hi daisukelab, </p>\n\n<p>Thank you very mcuh for your help! I just had another question, I've changed my run function to:</p>\n\n<pre><code>test = pd.read_csv(\"../input/sample_submission.csv\")\ntrain = pd.read_csv(\"../input/train_1.csv\")\n\nLABELS = list(train.label.unique())\nlabel_idx = {label: i for i, label in enumerate(LABELS)}\n\ntrain.set_index('fname', inplace=True)\ntest.set_index('fname', inplace=True)\ntrain['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n\nx_train = prepare_data(train, config, '../input/audio_train/')\nx_test = prepare_data(train, config, '../input/audio_test/')\n\ny_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n\nX_train, X_test, y_train, y_test = train_test_split(x_train, train.label_idx.values,\n                                                    test_size=0.2)\n\nmodel = model_fn_aes(config)\ncallbacks_list = []\nmodel.fit(X_train, np.array(y_train), validation_data=(X_test, np.array(y_test)), epochs=50, callbacks=callbacks_list)\n</code></pre>\n\n<p>Becuase I couldnt seem to get the StratifiedKFold function to work with regular lists. However I still cannot get the Keras fit function to work with the list or arrays and get the following error:\nExpected to see 1 array(s), but instead got the following list of 13 arrays</p>\n\n<p>I has hoping you might be able to help me get around this. Again I greatly appriciate your help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 336160,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/31/2018 06:45:47",
          "content": "<p>Hi Henry,</p>\n\n<p>I should share my conclusion that we cannot almost exploit benefit from SpatialPyramidPooling.</p>\n\n<ul>\n<li>What was created by prepare_data is a list that holds np.array with different length.</li>\n<li>Keras fit() accepts numpy array, not list. fit() will get to know what is the shape of given input, but it cannot get it from list; we have to feed numpy array.</li>\n<li>Numpy array cannot have variable length multi dimensional array.</li>\n<li>These restrictions prevent us from training with variable length input.</li>\n<li>Model itself would work fine with variable test input, if once it have trained with fixed length training inputs, but we cannot train by aforementioned reasons.</li>\n</ul>\n\n<p>You can start training by making all the training data into unified length by padding as following example:</p>\n\n<pre><code>def prepare_data(df, data_dir, sr=44100, bands=128):\n    X = []\n    # Convert once\n    for i, fname in enumerate(df.fname):\n        file_path = os.path.join(data_dir, fname)\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr=sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        X.append(logspec)\n    # Find longest\n    max_length = np.max([x.shape[1] for x in X])\n    # Pad zero to make them all the same length\n    X2 = [np.pad(x, ((0, 0), (0, max_length - x.shape[1]), (0, 0)), 'constant') for x in X]\n    return np.array(X2)</code></pre>\n\n<p>We still have way to make generator for feeding fixed-length-data-for-batch but variable-for-entire-training-set, but it takes time for both implementing and training.\nI personally abandon attempt for this approach...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 336466,
          "author_name": "hhargrea",
          "author_url": "",
          "post_date": "05/31/2018 19:28:21",
          "content": "<p>Thank you for your response again, I see the problem now, do you think it might be work setting the batch size to 1 or doing batched of similarly sized audio files, just some workaround I thought might work?​</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 336479,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/31/2018 19:53:44",
          "content": "<p>Yes it would be possible, but:</p>\n\n<ul>\n<li>You will need to make simple data generator to feed [1, 128, varied-length, 1] shaped numpy array.</li>\n<li>It would take effort to tune LR, and also take very long time to train as explained in Deep Learning book by Goodfellow et al., see link below.</li>\n</ul>\n\n<p><a href=\"https://stackoverflow.com/questions/46654424/how-to-calculate-optimal-batch-size\">https://stackoverflow.com/questions/46654424/how-to-calculate-optimal-batch-size</a></p>\n\n<p>\"Small batches can offer a regularizing effect (Wilson and Martinez, 2003), perhaps due to the noise they add to the learning process. Generalization error is often best for a batch size of 1. Training with such a small batch size might require a small learning rate to maintain stability because of the high variance in the estimate of the gradient. The total runtime can be very high as a result of the need to make more steps, both because of the reduced learning rate and because it takes more steps to observe the entire training set.”</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "333763": "I'm currently trying to implement a FCN using Keras to classify audio. So far I've created a basic CNN which takes the training set and either crop each sample to 2 second or pads it. However, I was reading a paper (https://arxiv.org/pdf/1707.02530.pdf) in the paper they build an FCN which allows variable length input. I have created such a FCN using Keras and has a SpatialPyramidPooling layer.\n\n    batch_size = 64\n    num_channels = 3\n    num_classes = 10\n    \n    model = Sequential()\n    \n    # uses theano ordering. Note that we leave the image size as None to allow multiple image sizes\n    model.add(Convolution2D(32, 3, 3, border_mode='same', input_shape=(3, None, None)))\n    model.add(Activation('relu'))\n    model.add(Convolution2D(32, 3, 3))\n    model.add(Activation('relu'))\n    model.add(MaxPooling2D(pool_size=(2, 2)))\n    model.add(Convolution2D(64, 3, 3, border_mode='same'))\n    model.add(Activation('relu'))\n    model.add(Convolution2D(64, 3, 3))\n    model.add(Activation('relu'))\n    model.add(SpatialPyramidPooling([1, 2, 4]))\n    model.add(Dense(num_classes))\n    model.add(Activation('softmax'))\n    heremodel.compile(loss='categorical_crossentropy', optimizer='sgd')\n\nHowever, I cannot seem to figure how to pre=process the audio in order for to be input into the FCN. Here is my function to prepare the data:\n\n    def prepare_data(df, config, data_dir, bands=128):\n    x = np.empty(shape=(len(df.index), bands))\n    log_specgrams_2048 = []\n    for i, fname in enumerate(df.index):\n        file_path = data_dir + fname\n        data, _ = librosa.core.load(file_path, sr=16000, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr=16000, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec)\n        log_specgrams_2048.append(logspec)\n    return log_specgrams_2048\n\n\nAnd here is my train function:\n\n      def run(config):\n        test = pd.read_csv(\"../input/sample_submission.csv\")\n        train = pd.read_csv(\"../input/train_1.csv\")\n        \n        LABELS = list(train.label.unique())\n        label_idx = {label: i for i, label in enumerate(LABELS)}\n        \n        train.set_index('fname', inplace=True)\n        test.set_index('fname', inplace=True)\n        train['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n        \n        x_train = prepare_data(train, config, '../input/audio_train_1/')\n        # x_test = prepare_data(train, config, '../input/audio_test/')\n        \n        y_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n        \n        skf = StratifiedKFold(n_splits=2).split(np.zeros(len(train)), train.label_idx)\n        for i, (train_split, val_split) in enumerate(skf):\n            K.clear_session()\n            x, y, x_val, y_val = np.array(x_train[train_split]), y_train[train_split], np.array(x_train[val_split]), y_train[val_split]\n        \n            checkpoint = ModelCheckpoint(config.job_dir + '/best_%d.h5' % i, monitor='val_loss', verbose=1,\n                                         save_best_only=True)\n            early = EarlyStopping(monitor='val_loss', mode='min', patience=5)\n            tb = TensorBoard(log_dir=os.path.join(config.job_dir, 'logs') + '/fold_%i' % i, write_graph=True)\n            callbacks_list = [checkpoint, early, tb]\n        \n            print(('Fold: %d' % i) + '\\n' + '#' * 50)\n        \n            curr_model = model_fn_aes(x_train.shape)\n        \n            curr_model.fit(x, y, validation_data=(x_val, y_val), callbacks=callbacks_list,\n                           batch_size=64, epochs=config.max_epochs)\n        \n            curr_model.load_weights(config.job_dir + '/best_%d.h5' % i)\n\nI would greatly appreciate it if someone could help me fix the prepare_data function so that I can load the training set into an np array and input into the FCN whilst still using StratifiedKFold.",
    "334135": "Hi Henry,\n\nI come across this Japanese tech blog [SPP（SpatialPyramidPooling）for free size input...][1].\n\nAnd it says that input_shape could be different order; change it from `(3, None, None)` to `( None, None, 3)`.\n\nBut the blog handles image data. yours is 1 channel audio; 3 could be 1. And I think the paper seems to use 44,100Hz instead of 16,000Hz. Then I _guess_:\n\n1. Change `(3, None, None)` to `( None, None, 1)`.\n2. prepare_data could be like this:\n\n<pre><code>sr = 44100\ndef prepare_data(df, config, data_dir, bands=128):\n    x = np.empty(shape=(len(df.index), bands))\n    log_specgrams_2048 = []\n    for i, fname in enumerate(df.index):\n        file_path = data_dir + fname\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr= sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        log_specgrams_2048.append(logspec)\n    return log_specgrams_2048\n</code></pre>\n\nI hope it help some, anyway thanks for sharing great idea.\n\n  [1]: https://qiita.com/MuAuan/items/dc819c17bdb030c0e096",
    "336103": "Hi daisukelab, \n\n\nThank you very mcuh for your help! I just had another question, I've changed my run function to:\n\n    test = pd.read_csv(\"../input/sample_submission.csv\")\n    train = pd.read_csv(\"../input/train_1.csv\")\n\n    LABELS = list(train.label.unique())\n    label_idx = {label: i for i, label in enumerate(LABELS)}\n\n    train.set_index('fname', inplace=True)\n    test.set_index('fname', inplace=True)\n    train['label_idx'] = train.label.apply(lambda elem: label_idx[elem])\n\n    x_train = prepare_data(train, config, '../input/audio_train/')\n    x_test = prepare_data(train, config, '../input/audio_test/')\n\n    y_train = to_categorical(train.label_idx, num_classes=config.n_classes)\n\n    X_train, X_test, y_train, y_test = train_test_split(x_train, train.label_idx.values,\n                                                        test_size=0.2)\n\n    model = model_fn_aes(config)\n    callbacks_list = []\n    model.fit(X_train, np.array(y_train), validation_data=(X_test, np.array(y_test)), epochs=50, callbacks=callbacks_list)\n\n\nBecuase I couldnt seem to get the StratifiedKFold function to work with regular lists. However I still cannot get the Keras fit function to work with the list or arrays and get the following error:\nExpected to see 1 array(s), but instead got the following list of 13 arrays\n\nI has hoping you might be able to help me get around this. Again I greatly appriciate your help!",
    "336160": "Hi Henry,\n\nI should share my conclusion that we cannot almost exploit benefit from SpatialPyramidPooling.\n\n- What was created by prepare_data is a list that holds np.array with different length.\n- Keras fit() accepts numpy array, not list. fit() will get to know what is the shape of given input, but it cannot get it from list; we have to feed numpy array.\n- Numpy array cannot have variable length multi dimensional array.\n- These restrictions prevent us from training with variable length input.\n- Model itself would work fine with variable test input, if once it have trained with fixed length training inputs, but we cannot train by aforementioned reasons.\n\nYou can start training by making all the training data into unified length by padding as following example:\n\n<pre><code>def prepare_data(df, data_dir, sr=44100, bands=128):\n    X = []\n    # Convert once\n    for i, fname in enumerate(df.fname):\n        file_path = os.path.join(data_dir, fname)\n        data, _ = librosa.core.load(file_path, sr= sr, res_type=\"kaiser_fast\")\n        melspec = librosa.feature.melspectrogram(data, sr=sr, n_mels=bands)\n        logspec = librosa.core.power_to_db(melspec) # shape would be [128, your_audio_length]\n        logspec = logspec[..., np.newaxis] # shape will be [128, your_audio_length, 1]\n        X.append(logspec)\n    # Find longest\n    max_length = np.max([x.shape[1] for x in X])\n    # Pad zero to make them all the same length\n    X2 = [np.pad(x, ((0, 0), (0, max_length - x.shape[1]), (0, 0)), 'constant') for x in X]\n    return np.array(X2)</code></pre>\n\nWe still have way to make generator for feeding fixed-length-data-for-batch but variable-for-entire-training-set, but it takes time for both implementing and training.\nI personally abandon attempt for this approach...",
    "336466": "Thank you for your response again, I see the problem now, do you think it might be work setting the batch size to 1 or doing batched of similarly sized audio files, just some workaround I thought might work?​",
    "336479": "Yes it would be possible, but:\n\n- You will need to make simple data generator to feed [1, 128, varied-length, 1] shaped numpy array.\n- It would take effort to tune LR, and also take very long time to train as explained in Deep Learning book by Goodfellow et al., see link below.\n\nhttps://stackoverflow.com/questions/46654424/how-to-calculate-optimal-batch-size\n\n\"Small batches can offer a regularizing effect (Wilson and Martinez, 2003), perhaps due to the noise they add to the learning process. Generalization error is often best for a batch size of 1. Training with such a small batch size might require a small learning rate to maintain stability because of the high variance in the estimate of the gradient. The total runtime can be very high as a result of the need to make more steps, both because of the reduced learning rate and because it takes more steps to observe the entire training set.”"
  },
  "source": "meta"
}