{
  "id": 62518,
  "title": "Using large amount of non-deep features",
  "url": "/competitions/freesound-audio-tagging/discussion/62518",
  "author_name": "",
  "post_date": "2018-08-02T16:40:46.456822100Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to the organizers for an interesting challenge!</p>\n\n<p>Coming into this competition, I was interested in using as little deep learning features as possible (just for the fun of it, and perhaps to see whether older approaches work as well).\nOn the public LB, my best non-deep learning approach got a score of 0.888; best approach with some CNN-based features added in got 0.926.</p>\n\n<p>For the deep learning features I used <a href=\"https://github.com/tensorflow/models/tree/master/research/audioset\">VGGish</a>, a VGG-based CNN trained on AudioSet (you can also check out the <a href=\"https://github.com/knstmrd/vggish-batch\">batch processing code</a> I wrote for it).\nBut mostly I focused on getting features from YAAFE and Essentia; for time-dependent features, I computed the following statistics and values: Position of minimum and maximum; Value of Minimum and maximum; Mean, standard deviation; 10th, 25th, 50th, 75th, 90th percentiles; Skew, kurtosis.</p>\n\n<p>I also computed autocorrelation function-based features (normalized value of first peak, peak position and ZCR of autocorrelation function), and also these features for the waveform passed through 48 Gammatone filters.</p>\n\n<p>A full description of the features I used can be found in my <a href=\"https://github.com/knstmrd/kagglefreesound\">Github repo</a>, and here are 20 top performing features (according to LightGBM importance in ascending order; numbers refer to the index if the feature is a vector)</p>\n\n<ol>\n<li>amplitudemod 1 max (YAAFE feature)</li>\n<li>wav argmin_rel</li>\n<li>wav mean</li>\n<li>energy basic max (YAAFE feature)</li>\n<li>wav skew</li>\n<li>TCToTotal (Essentia feature)</li>\n<li>obsir 8 min (YAAFE feature)</li>\n<li>autocorr peak position</li>\n<li>logattack 0 (Essentia feature)</li>\n<li>energy kurtosis (YAAFE feature)</li>\n<li>amplitudemod 5 max (YAAFE feature)</li>\n<li>spectralflux 10th percentile</li>\n<li>maxmag (Essentia feature)</li>\n<li>autocorr ZCR</li>\n<li>energy skew (YAAFE feature)</li>\n<li>obsir basic 0 max (YAAFE feature)</li>\n<li>logattack 1 (Essentia feature)</li>\n<li>wav kurtosis</li>\n<li>derivative SFX 1 (Essentia feature)</li>\n<li>salience (Essentia feature)</li>\n</ol>",
  "messages": [
    {
      "id": "365453",
      "postDate": "08/02/2018 16:40:46",
      "content": "<p>Thanks to the organizers for an interesting challenge!</p>\n\n<p>Coming into this competition, I was interested in using as little deep learning features as possible (just for the fun of it, and perhaps to see whether older approaches work as well).\nOn the public LB, my best non-deep learning approach got a score of 0.888; best approach with some CNN-based features added in got 0.926.</p>\n\n<p>For the deep learning features I used <a href=\"https://github.com/tensorflow/models/tree/master/research/audioset\">VGGish</a>, a VGG-based CNN trained on AudioSet (you can also check out the <a href=\"https://github.com/knstmrd/vggish-batch\">batch processing code</a> I wrote for it).\nBut mostly I focused on getting features from YAAFE and Essentia; for time-dependent features, I computed the following statistics and values: Position of minimum and maximum; Value of Minimum and maximum; Mean, standard deviation; 10th, 25th, 50th, 75th, 90th percentiles; Skew, kurtosis.</p>\n\n<p>I also computed autocorrelation function-based features (normalized value of first peak, peak position and ZCR of autocorrelation function), and also these features for the waveform passed through 48 Gammatone filters.</p>\n\n<p>A full description of the features I used can be found in my <a href=\"https://github.com/knstmrd/kagglefreesound\">Github repo</a>, and here are 20 top performing features (according to LightGBM importance in ascending order; numbers refer to the index if the feature is a vector)</p>\n\n<ol>\n<li>amplitudemod 1 max (YAAFE feature)</li>\n<li>wav argmin_rel</li>\n<li>wav mean</li>\n<li>energy basic max (YAAFE feature)</li>\n<li>wav skew</li>\n<li>TCToTotal (Essentia feature)</li>\n<li>obsir 8 min (YAAFE feature)</li>\n<li>autocorr peak position</li>\n<li>logattack 0 (Essentia feature)</li>\n<li>energy kurtosis (YAAFE feature)</li>\n<li>amplitudemod 5 max (YAAFE feature)</li>\n<li>spectralflux 10th percentile</li>\n<li>maxmag (Essentia feature)</li>\n<li>autocorr ZCR</li>\n<li>energy skew (YAAFE feature)</li>\n<li>obsir basic 0 max (YAAFE feature)</li>\n<li>logattack 1 (Essentia feature)</li>\n<li>wav kurtosis</li>\n<li>derivative SFX 1 (Essentia feature)</li>\n<li>salience (Essentia feature)</li>\n</ol>",
      "rawMarkdown": "Thanks to the organizers for an interesting challenge!\n\nComing into this competition, I was interested in using as little deep learning features as possible (just for the fun of it, and perhaps to see whether older approaches work as well).\nOn the public LB, my best non-deep learning approach got a score of 0.888; best approach with some CNN-based features added in got 0.926.\n\nFor the deep learning features I used [VGGish][1], a VGG-based CNN trained on AudioSet (you can also check out the [batch processing code][2] I wrote for it).\nBut mostly I focused on getting features from YAAFE and Essentia; for time-dependent features, I computed the following statistics and values: Position of minimum and maximum; Value of Minimum and maximum; Mean, standard deviation; 10th, 25th, 50th, 75th, 90th percentiles; Skew, kurtosis.\n\nI also computed autocorrelation function-based features (normalized value of first peak, peak position and ZCR of autocorrelation function), and also these features for the waveform passed through 48 Gammatone filters.\n\nA full description of the features I used can be found in my [Github repo][3], and here are 20 top performing features (according to LightGBM importance in ascending order; numbers refer to the index if the feature is a vector)\n\n 1. amplitudemod 1 max (YAAFE feature)\n 2. wav argmin_rel\n 3. wav mean\n 4. energy basic max (YAAFE feature)\n 5. wav skew\n 6. TCToTotal (Essentia feature)\n 7. obsir 8 min (YAAFE feature)\n 8. autocorr peak position\n 9. logattack 0 (Essentia feature)\n 10. energy kurtosis (YAAFE feature)\n 11. amplitudemod 5 max (YAAFE feature)\n 12. spectralflux 10th percentile\n 13. maxmag (Essentia feature)\n 14. autocorr ZCR\n 15. energy skew (YAAFE feature)\n 16. obsir basic 0 max (YAAFE feature)\n 17. logattack 1 (Essentia feature)\n 18. wav kurtosis\n 19. derivative SFX 1 (Essentia feature)\n 20. salience (Essentia feature)\n\n  [1]: https://github.com/tensorflow/models/tree/master/research/audioset\n  [2]: https://github.com/knstmrd/vggish-batch\n  [3]: https://github.com/knstmrd/kagglefreesound",
      "votes": null
    },
    {
      "id": "367361",
      "postDate": "08/07/2018 15:28:46",
      "content": "<p>Thanks for sharing, I especially wanted to use Essentia but cannot find tutorial there and didn't try.\nYour code seems great example especially for the people in this competition, it will be helpful for my future works. Thanks again!</p>",
      "rawMarkdown": "Thanks for sharing, I especially wanted to use Essentia but cannot find tutorial there and didn't try.\nYour code seems great example especially for the people in this competition, it will be helpful for my future works. Thanks again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 367361,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "08/07/2018 15:28:46",
      "content": "<p>Thanks for sharing, I especially wanted to use Essentia but cannot find tutorial there and didn't try.\nYour code seems great example especially for the people in this competition, it will be helpful for my future works. Thanks again!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "365453": "Thanks to the organizers for an interesting challenge!\n\nComing into this competition, I was interested in using as little deep learning features as possible (just for the fun of it, and perhaps to see whether older approaches work as well).\nOn the public LB, my best non-deep learning approach got a score of 0.888; best approach with some CNN-based features added in got 0.926.\n\nFor the deep learning features I used [VGGish][1], a VGG-based CNN trained on AudioSet (you can also check out the [batch processing code][2] I wrote for it).\nBut mostly I focused on getting features from YAAFE and Essentia; for time-dependent features, I computed the following statistics and values: Position of minimum and maximum; Value of Minimum and maximum; Mean, standard deviation; 10th, 25th, 50th, 75th, 90th percentiles; Skew, kurtosis.\n\nI also computed autocorrelation function-based features (normalized value of first peak, peak position and ZCR of autocorrelation function), and also these features for the waveform passed through 48 Gammatone filters.\n\nA full description of the features I used can be found in my [Github repo][3], and here are 20 top performing features (according to LightGBM importance in ascending order; numbers refer to the index if the feature is a vector)\n\n 1. amplitudemod 1 max (YAAFE feature)\n 2. wav argmin_rel\n 3. wav mean\n 4. energy basic max (YAAFE feature)\n 5. wav skew\n 6. TCToTotal (Essentia feature)\n 7. obsir 8 min (YAAFE feature)\n 8. autocorr peak position\n 9. logattack 0 (Essentia feature)\n 10. energy kurtosis (YAAFE feature)\n 11. amplitudemod 5 max (YAAFE feature)\n 12. spectralflux 10th percentile\n 13. maxmag (Essentia feature)\n 14. autocorr ZCR\n 15. energy skew (YAAFE feature)\n 16. obsir basic 0 max (YAAFE feature)\n 17. logattack 1 (Essentia feature)\n 18. wav kurtosis\n 19. derivative SFX 1 (Essentia feature)\n 20. salience (Essentia feature)\n\n  [1]: https://github.com/tensorflow/models/tree/master/research/audioset\n  [2]: https://github.com/knstmrd/vggish-batch\n  [3]: https://github.com/knstmrd/kagglefreesound",
    "367361": "Thanks for sharing, I especially wanted to use Essentia but cannot find tutorial there and didn't try.\nYour code seems great example especially for the people in this competition, it will be helpful for my future works. Thanks again!"
  },
  "source": "meta"
}