{
  "id": 47628,
  "title": "Summary of 25 Solution",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47628",
  "author_name": "",
  "post_date": "2018-01-17T05:14:31.255189100Z",
  "votes": 16,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Summary of the aproach to TF Speec Recognition Challenge</h1>\n\n<p>by ironbar</p>\n\n<h2>Input</h2>\n\n<p>I first normalized all the audios to have a duration of 1 second. I transformed the raw audio to an image using log mel power with 50 bank filters. The size of the window was 20ms and the step was 5s. I found that using a small step improved the cross-validation score (5ms better than 10ms, and 10ms better than 20ms). <br>\nI used the library librosa for computing the mel spectrum and also to transform from power to db. I normalized all the spectrograms using the max. I also tried without normalization but the results were worse. <br>\nAfter this I used the mean and std of the train dataset to normalize the spectrograms before feeding them to the network.</p>\n\n<p>At the end of the challenge I noticed that there were many mislabelled audios on the train dataset. I created candidates for cleaning using the predicitions on the hold folds and I made a simple app with jupyter notebook for listening to the audios and censure the mislabelled ones. I censured about 1.7% of the data. <br>\nAfter censuring this files the cross-validation scores raised from 0.967 up to 0.987.</p>\n\n<h2>Sampling</h2>\n\n<p>I used batches of size 36, having 3 samples of each of the standard labels.</p>\n\n<p>I divided the audios by user_id for cross-validation so a user was never on train and validation at the same time.</p>\n\n<h2>Data augmentation</h2>\n\n<p>On silence audios I applied random reverting and rolling. I tried mixing them but I decided it wasn't a good idea because there were big differences in volume between the silence audios.</p>\n\n<p>On unknown I tried with reverting the words and also repeatign the word many times in the audio. However it did not seem to improve. </p>\n\n<p>I believe more advanced data augmentation was needed in this challenge, but I only noticed this at the end of the challenge and there was no time. See the last section for more info.</p>\n\n<h2>Model</h2>\n\n<p>I used a customized version of Xception architecture. I played with the depth, number of filters and maxpool to adapt the architecture to the problem. It worked very well achieving cross-validation scores up to 0.99 on clean data.</p>\n\n<h2>Output</h2>\n\n<p>From the start of the challenge I used 3 outputs on the models:\n* word/silence\n* standard labels\n* expanded labels (standard + the other words provided)</p>\n\n<h2>Some thoughts</h2>\n\n<p>During all the challenge I saw that there were big differences between my cross-validation scores and the LB scores. At the end of the challenge I decided to hear some of the test audios and I discovered that there were new words on the test set, and that they were much more difficult that the ones on the train set.</p>\n\n<ul>\n<li>down vs town</li>\n<li>left vs learn</li>\n<li>off, on vs oh</li>\n</ul>\n\n<p>This explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset.</p>\n\n<p><strong>I don't understand why this kind of things are not clearly explained at the beginning of the challenge. I like Kaggle a lot but this is not the first time I feel I have been fooled by the host.</strong></p>\n\n<h3>A fresh new aproach</h3>\n\n<p>If at the beginning of the challenge I had know that our model need to be able to learn to distinguish very subttle differences between words I would have first trained a model for segmenting the audio into phonemes. Using that segmentation I would build a dictionary with all the phonemes in the train set. And by combining phonemes we can make useful data augmentation for this problem. For example we could create similar words to down: town, doll, own, call...\nUsing that data augmented dataset the model could be able to learn subttle differences between words.</p>",
  "messages": [
    {
      "id": "269678",
      "postDate": "01/17/2018 05:14:31",
      "content": "<h1>Summary of the aproach to TF Speec Recognition Challenge</h1>\n\n<p>by ironbar</p>\n\n<h2>Input</h2>\n\n<p>I first normalized all the audios to have a duration of 1 second. I transformed the raw audio to an image using log mel power with 50 bank filters. The size of the window was 20ms and the step was 5s. I found that using a small step improved the cross-validation score (5ms better than 10ms, and 10ms better than 20ms). <br>\nI used the library librosa for computing the mel spectrum and also to transform from power to db. I normalized all the spectrograms using the max. I also tried without normalization but the results were worse. <br>\nAfter this I used the mean and std of the train dataset to normalize the spectrograms before feeding them to the network.</p>\n\n<p>At the end of the challenge I noticed that there were many mislabelled audios on the train dataset. I created candidates for cleaning using the predicitions on the hold folds and I made a simple app with jupyter notebook for listening to the audios and censure the mislabelled ones. I censured about 1.7% of the data. <br>\nAfter censuring this files the cross-validation scores raised from 0.967 up to 0.987.</p>\n\n<h2>Sampling</h2>\n\n<p>I used batches of size 36, having 3 samples of each of the standard labels.</p>\n\n<p>I divided the audios by user_id for cross-validation so a user was never on train and validation at the same time.</p>\n\n<h2>Data augmentation</h2>\n\n<p>On silence audios I applied random reverting and rolling. I tried mixing them but I decided it wasn't a good idea because there were big differences in volume between the silence audios.</p>\n\n<p>On unknown I tried with reverting the words and also repeatign the word many times in the audio. However it did not seem to improve. </p>\n\n<p>I believe more advanced data augmentation was needed in this challenge, but I only noticed this at the end of the challenge and there was no time. See the last section for more info.</p>\n\n<h2>Model</h2>\n\n<p>I used a customized version of Xception architecture. I played with the depth, number of filters and maxpool to adapt the architecture to the problem. It worked very well achieving cross-validation scores up to 0.99 on clean data.</p>\n\n<h2>Output</h2>\n\n<p>From the start of the challenge I used 3 outputs on the models:\n* word/silence\n* standard labels\n* expanded labels (standard + the other words provided)</p>\n\n<h2>Some thoughts</h2>\n\n<p>During all the challenge I saw that there were big differences between my cross-validation scores and the LB scores. At the end of the challenge I decided to hear some of the test audios and I discovered that there were new words on the test set, and that they were much more difficult that the ones on the train set.</p>\n\n<ul>\n<li>down vs town</li>\n<li>left vs learn</li>\n<li>off, on vs oh</li>\n</ul>\n\n<p>This explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset.</p>\n\n<p><strong>I don't understand why this kind of things are not clearly explained at the beginning of the challenge. I like Kaggle a lot but this is not the first time I feel I have been fooled by the host.</strong></p>\n\n<h3>A fresh new aproach</h3>\n\n<p>If at the beginning of the challenge I had know that our model need to be able to learn to distinguish very subttle differences between words I would have first trained a model for segmenting the audio into phonemes. Using that segmentation I would build a dictionary with all the phonemes in the train set. And by combining phonemes we can make useful data augmentation for this problem. For example we could create similar words to down: town, doll, own, call...\nUsing that data augmented dataset the model could be able to learn subttle differences between words.</p>",
      "rawMarkdown": "# Summary of the aproach to TF Speec Recognition Challenge\nby ironbar\n\n## Input\nI first normalized all the audios to have a duration of 1 second. I transformed the raw audio to an image using log mel power with 50 bank filters. The size of the window was 20ms and the step was 5s. I found that using a small step improved the cross-validation score (5ms better than 10ms, and 10ms better than 20ms).  \nI used the library librosa for computing the mel spectrum and also to transform from power to db. I normalized all the spectrograms using the max. I also tried without normalization but the results were worse.  \nAfter this I used the mean and std of the train dataset to normalize the spectrograms before feeding them to the network.\n\nAt the end of the challenge I noticed that there were many mislabelled audios on the train dataset. I created candidates for cleaning using the predicitions on the hold folds and I made a simple app with jupyter notebook for listening to the audios and censure the mislabelled ones. I censured about 1.7% of the data.  \nAfter censuring this files the cross-validation scores raised from 0.967 up to 0.987.\n## Sampling\nI used batches of size 36, having 3 samples of each of the standard labels.\n\nI divided the audios by user_id for cross-validation so a user was never on train and validation at the same time.\n\n## Data augmentation\nOn silence audios I applied random reverting and rolling. I tried mixing them but I decided it wasn't a good idea because there were big differences in volume between the silence audios.\n\nOn unknown I tried with reverting the words and also repeatign the word many times in the audio. However it did not seem to improve. \n\nI believe more advanced data augmentation was needed in this challenge, but I only noticed this at the end of the challenge and there was no time. See the last section for more info.\n## Model\nI used a customized version of Xception architecture. I played with the depth, number of filters and maxpool to adapt the architecture to the problem. It worked very well achieving cross-validation scores up to 0.99 on clean data.\n\n## Output\nFrom the start of the challenge I used 3 outputs on the models:\n* word/silence\n* standard labels\n* expanded labels (standard + the other words provided)\n\n## Some thoughts\nDuring all the challenge I saw that there were big differences between my cross-validation scores and the LB scores. At the end of the challenge I decided to hear some of the test audios and I discovered that there were new words on the test set, and that they were much more difficult that the ones on the train set.\n\n* down vs town\n* left vs learn\n* off, on vs oh\n\nThis explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset.\n\n**I don't understand why this kind of things are not clearly explained at the beginning of the challenge. I like Kaggle a lot but this is not the first time I feel I have been fooled by the host.**\n\n### A fresh new aproach\nIf at the beginning of the challenge I had know that our model need to be able to learn to distinguish very subttle differences between words I would have first trained a model for segmenting the audio into phonemes. Using that segmentation I would build a dictionary with all the phonemes in the train set. And by combining phonemes we can make useful data augmentation for this problem. For example we could create similar words to down: town, doll, own, call...\nUsing that data augmented dataset the model could be able to learn subttle differences between words.",
      "votes": null
    },
    {
      "id": "285814",
      "postDate": "02/21/2018 02:51:16",
      "content": "<p>i think the take away message is:</p>\n\n<ul>\n<li><p>data (and labels) are never perfect</p></li>\n<li><p>understand the problem first, then understand the data (both train and LB data), then understand the algorithm</p></li>\n</ul>",
      "rawMarkdown": "i think the take away message is:\n\n- data (and labels) are never perfect\n\n- understand the problem first, then understand the data (both train and LB data), then understand the algorithm",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 285814,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/21/2018 02:51:16",
      "content": "<p>i think the take away message is:</p>\n\n<ul>\n<li><p>data (and labels) are never perfect</p></li>\n<li><p>understand the problem first, then understand the data (both train and LB data), then understand the algorithm</p></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "269678": "# Summary of the aproach to TF Speec Recognition Challenge\nby ironbar\n\n## Input\nI first normalized all the audios to have a duration of 1 second. I transformed the raw audio to an image using log mel power with 50 bank filters. The size of the window was 20ms and the step was 5s. I found that using a small step improved the cross-validation score (5ms better than 10ms, and 10ms better than 20ms).  \nI used the library librosa for computing the mel spectrum and also to transform from power to db. I normalized all the spectrograms using the max. I also tried without normalization but the results were worse.  \nAfter this I used the mean and std of the train dataset to normalize the spectrograms before feeding them to the network.\n\nAt the end of the challenge I noticed that there were many mislabelled audios on the train dataset. I created candidates for cleaning using the predicitions on the hold folds and I made a simple app with jupyter notebook for listening to the audios and censure the mislabelled ones. I censured about 1.7% of the data.  \nAfter censuring this files the cross-validation scores raised from 0.967 up to 0.987.\n## Sampling\nI used batches of size 36, having 3 samples of each of the standard labels.\n\nI divided the audios by user_id for cross-validation so a user was never on train and validation at the same time.\n\n## Data augmentation\nOn silence audios I applied random reverting and rolling. I tried mixing them but I decided it wasn't a good idea because there were big differences in volume between the silence audios.\n\nOn unknown I tried with reverting the words and also repeatign the word many times in the audio. However it did not seem to improve. \n\nI believe more advanced data augmentation was needed in this challenge, but I only noticed this at the end of the challenge and there was no time. See the last section for more info.\n## Model\nI used a customized version of Xception architecture. I played with the depth, number of filters and maxpool to adapt the architecture to the problem. It worked very well achieving cross-validation scores up to 0.99 on clean data.\n\n## Output\nFrom the start of the challenge I used 3 outputs on the models:\n* word/silence\n* standard labels\n* expanded labels (standard + the other words provided)\n\n## Some thoughts\nDuring all the challenge I saw that there were big differences between my cross-validation scores and the LB scores. At the end of the challenge I decided to hear some of the test audios and I discovered that there were new words on the test set, and that they were much more difficult that the ones on the train set.\n\n* down vs town\n* left vs learn\n* off, on vs oh\n\nThis explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset.\n\n**I don't understand why this kind of things are not clearly explained at the beginning of the challenge. I like Kaggle a lot but this is not the first time I feel I have been fooled by the host.**\n\n### A fresh new aproach\nIf at the beginning of the challenge I had know that our model need to be able to learn to distinguish very subttle differences between words I would have first trained a model for segmenting the audio into phonemes. Using that segmentation I would build a dictionary with all the phonemes in the train set. And by combining phonemes we can make useful data augmentation for this problem. For example we could create similar words to down: town, doll, own, call...\nUsing that data augmented dataset the model could be able to learn subttle differences between words.",
    "285814": "i think the take away message is:\n\n- data (and labels) are never perfect\n\n- understand the problem first, then understand the data (both train and LB data), then understand the algorithm"
  },
  "source": "meta"
}