{
  "id": 47841,
  "title": "FWIW, some code, 1D models, TF vs PyTorch, and Triplet Loss",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47841",
  "author_name": "RossWightman",
  "post_date": "2018-01-19T17:17:52.623000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I didn't do as well as I'd hoped in this competition, but as usual, I learned a lot.</p>\n\n<p>I started off enthusiastically in Tensorflow, basing my initial approach on the Tensorflow speech commands code base. I modified that code base to be compatible with TF Slim models and built some models of my own, including some 1D convolution models. I discovered at this point that there wasn't a huge variability from model to model, augmentation, dataset handling, and training hyperparams had had bigger impact. Most of my 1D models converged to something, with the 'conv1d_basic3' generally doing almost as well as the 2D mfcc models. The typical training resulted in 0.82-0.84 on the public LB, with one or two instances hitting 0.85 and low 0.86 (hard to reproduce). I got busy, frustrated and dropped things for quite a while with a few attempts here and there...</p>\n\n<p>With a week or so to go in the competition I decided I needed to at least get back into the top 10%. I dusted off some old PyTorch vision Kaggle competition codebases and hacked that together with a librosa based data augmentation pipeline I had started working on for the TF codebase. Within a day I had models consistently training with 0.85-0.87 after some manual silence class boosting, and 0.89 after basic ensembling. </p>\n\n<p>At this point, I switched back to (my) learning mode. I had really wanted to try metric learning via Triplet Loss in this competition as I thought it might be the 'proper' answer to the unknown class, but hadn't had a chance to. So, in the final days I worked on a Triplet Loss for the PyTorch models. Unfortunately I have yet to make the Triplet Loss converge. Still trying to get something out of it after the competition has ended. If anyone tried this and had (any) success with convergence to anything better than random feature vectors, would love to hear your approach... </p>\n\n<p><a href=\"https://github.com/rwightman/tensorflow-speech_commands\">https://github.com/rwightman/tensorflow-speech_commands</a></p>\n\n<p><a href=\"https://github.com/rwightman/pytorch-commands\">https://github.com/rwightman/pytorch-commands</a></p>",
  "messages": [
    {
      "id": 271092,
      "postDate": "2018-01-19T17:17:52.623Z",
      "content": "<p>I didn't do as well as I'd hoped in this competition, but as usual, I learned a lot.</p>\n\n<p>I started off enthusiastically in Tensorflow, basing my initial approach on the Tensorflow speech commands code base. I modified that code base to be compatible with TF Slim models and built some models of my own, including some 1D convolution models. I discovered at this point that there wasn't a huge variability from model to model, augmentation, dataset handling, and training hyperparams had had bigger impact. Most of my 1D models converged to something, with the 'conv1d_basic3' generally doing almost as well as the 2D mfcc models. The typical training resulted in 0.82-0.84 on the public LB, with one or two instances hitting 0.85 and low 0.86 (hard to reproduce). I got busy, frustrated and dropped things for quite a while with a few attempts here and there...</p>\n\n<p>With a week or so to go in the competition I decided I needed to at least get back into the top 10%. I dusted off some old PyTorch vision Kaggle competition codebases and hacked that together with a librosa based data augmentation pipeline I had started working on for the TF codebase. Within a day I had models consistently training with 0.85-0.87 after some manual silence class boosting, and 0.89 after basic ensembling. </p>\n\n<p>At this point, I switched back to (my) learning mode. I had really wanted to try metric learning via Triplet Loss in this competition as I thought it might be the 'proper' answer to the unknown class, but hadn't had a chance to. So, in the final days I worked on a Triplet Loss for the PyTorch models. Unfortunately I have yet to make the Triplet Loss converge. Still trying to get something out of it after the competition has ended. If anyone tried this and had (any) success with convergence to anything better than random feature vectors, would love to hear your approach... </p>\n\n<p><a href=\"https://github.com/rwightman/tensorflow-speech_commands\">https://github.com/rwightman/tensorflow-speech_commands</a></p>\n\n<p><a href=\"https://github.com/rwightman/pytorch-commands\">https://github.com/rwightman/pytorch-commands</a></p>",
      "rawMarkdown": "I didn't do as well as I'd hoped in this competition, but as usual, I learned a lot.\n\nI started off enthusiastically in Tensorflow, basing my initial approach on the Tensorflow speech commands code base. I modified that code base to be compatible with TF Slim models and built some models of my own, including some 1D convolution models. I discovered at this point that there wasn't a huge variability from model to model, augmentation, dataset handling, and training hyperparams had had bigger impact. Most of my 1D models converged to something, with the 'conv1d_basic3' generally doing almost as well as the 2D mfcc models. The typical training resulted in 0.82-0.84 on the public LB, with one or two instances hitting 0.85 and low 0.86 (hard to reproduce). I got busy, frustrated and dropped things for quite a while with a few attempts here and there...\n\nWith a week or so to go in the competition I decided I needed to at least get back into the top 10%. I dusted off some old PyTorch vision Kaggle competition codebases and hacked that together with a librosa based data augmentation pipeline I had started working on for the TF codebase. Within a day I had models consistently training with 0.85-0.87 after some manual silence class boosting, and 0.89 after basic ensembling. \n\nAt this point, I switched back to (my) learning mode. I had really wanted to try metric learning via Triplet Loss in this competition as I thought it might be the 'proper' answer to the unknown class, but hadn't had a chance to. So, in the final days I worked on a Triplet Loss for the PyTorch models. Unfortunately I have yet to make the Triplet Loss converge. Still trying to get something out of it after the competition has ended. If anyone tried this and had (any) success with convergence to anything better than random feature vectors, would love to hear your approach... \n\nhttps://github.com/rwightman/tensorflow-speech_commands\n\nhttps://github.com/rwightman/pytorch-commands",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "271092": "I didn't do as well as I'd hoped in this competition, but as usual, I learned a lot.\n\nI started off enthusiastically in Tensorflow, basing my initial approach on the Tensorflow speech commands code base. I modified that code base to be compatible with TF Slim models and built some models of my own, including some 1D convolution models. I discovered at this point that there wasn't a huge variability from model to model, augmentation, dataset handling, and training hyperparams had had bigger impact. Most of my 1D models converged to something, with the 'conv1d_basic3' generally doing almost as well as the 2D mfcc models. The typical training resulted in 0.82-0.84 on the public LB, with one or two instances hitting 0.85 and low 0.86 (hard to reproduce). I got busy, frustrated and dropped things for quite a while with a few attempts here and there...\n\nWith a week or so to go in the competition I decided I needed to at least get back into the top 10%. I dusted off some old PyTorch vision Kaggle competition codebases and hacked that together with a librosa based data augmentation pipeline I had started working on for the TF codebase. Within a day I had models consistently training with 0.85-0.87 after some manual silence class boosting, and 0.89 after basic ensembling. \n\nAt this point, I switched back to (my) learning mode. I had really wanted to try metric learning via Triplet Loss in this competition as I thought it might be the 'proper' answer to the unknown class, but hadn't had a chance to. So, in the final days I worked on a Triplet Loss for the PyTorch models. Unfortunately I have yet to make the Triplet Loss converge. Still trying to get something out of it after the competition has ended. If anyone tried this and had (any) success with convergence to anything better than random feature vectors, would love to hear your approach... \n\nhttps://github.com/rwightman/tensorflow-speech_commands\n\nhttps://github.com/rwightman/pytorch-commands"
  }
}