{
  "id": 90613,
  "title": "Best score using LSTMs?",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/90613",
  "author_name": "",
  "post_date": "2019-04-25T08:31:49.611291600Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I can get a LB: 0.495 with LSTMs. Whats your best score?</p>",
  "messages": [
    {
      "id": "522913",
      "postDate": "04/25/2019 08:31:49",
      "content": "<p>I can get a LB: 0.495 with LSTMs. Whats your best score?</p>",
      "rawMarkdown": "I can get a LB: 0.495 with LSTMs. Whats your best score?",
      "votes": null
    },
    {
      "id": "523981",
      "postDate": "04/27/2019 15:19:43",
      "content": "<p>I can get 0.65 with LSTM.</p>",
      "rawMarkdown": "I can get 0.65 with LSTM.",
      "votes": null
    },
    {
      "id": "524069",
      "postDate": "04/27/2019 20:35:49",
      "content": "<p>That's unbelievable! How? Is it rather some specific preprocessing, noisy/curated pretraining, data augmentation. or some fancy LSTM architectures? I admire your scores</p>",
      "rawMarkdown": "That's unbelievable! How? Is it rather some specific preprocessing, noisy/curated pretraining, data augmentation. or some fancy LSTM architectures? I admire your scores",
      "votes": null
    },
    {
      "id": "524075",
      "postDate": "04/27/2019 21:12:21",
      "content": "<p>Nothing really special, I used a default LSTM implementation. With my main model that doesn't use any RNNs, I can get a score of 0.7 which is my current LB score. I don't use the noisy subset for the moment being. I don't do any specific preprocessing (but I plan to experiment with some other audio representations later). </p>",
      "rawMarkdown": "Nothing really special, I used a default LSTM implementation. With my main model that doesn't use any RNNs, I can get a score of 0.7 which is my current LB score. I don't use the noisy subset for the moment being. I don't do any specific preprocessing (but I plan to experiment with some other audio representations later).",
      "votes": null
    },
    {
      "id": "524088",
      "postDate": "04/27/2019 22:18:17",
      "content": "<p>I am using melspectrogram with different GRU based networks (also very simple) and my CV scores are around 0.6 at best, with or without noisy data pretraining. That's why I don't get your post now - it looks like you do nothing and have a much better score than I am able to achieve. </p>\n\n<p>In this solution, the score is not that great too:\n<a href=\"https://www.kaggle.com/carlolepelaars/bidirectional-lstm-for-audio-labeling-with-keras\">https://www.kaggle.com/carlolepelaars/bidirectional-lstm-for-audio-labeling-with-keras</a></p>",
      "rawMarkdown": "I am using melspectrogram with different GRU based networks (also very simple) and my CV scores are around 0.6 at best, with or without noisy data pretraining. That's why I don't get your post now - it looks like you do nothing and have a much better score than I am able to achieve. \n\nIn this solution, the score is not that great too:\nhttps://www.kaggle.com/carlolepelaars/bidirectional-lstm-for-audio-labeling-with-keras",
      "votes": null
    },
    {
      "id": "524178",
      "postDate": "04/28/2019 06:28:14",
      "content": "<p>I didn't say I'm doing nothing :) I certainly did some improvements to a \"naive\" approach, but they are not really fancy (at least so far). </p>",
      "rawMarkdown": "I didn't say I'm doing nothing :) I certainly did some improvements to a \"naive\" approach, but they are not really fancy (at least so far).",
      "votes": null
    },
    {
      "id": "524179",
      "postDate": "04/28/2019 06:33:33",
      "content": "<p>My main advice would be to pay more attention on the problem statement and also to use domain knowledge (that's pretty obvious, right? :) )</p>",
      "rawMarkdown": "My main advice would be to pay more attention on the problem statement and also to use domain knowledge (that's pretty obvious, right? :) )",
      "votes": null
    },
    {
      "id": "524451",
      "postDate": "04/28/2019 19:17:50",
      "content": "<p>Seems obvious... Thanks for the tip. </p>\n\n<p>Domain knowledge for me is to understand mel-spectrogram or MFCC creation, choosing appropriate window size (I'm using 25 ms), resampling (I think can be good because don't need high frequencies to distinguish sounds), windowing, normalization (one mean and std based on train_cirated) - I don't see much impact using different parameters there.\nBut that's what you mean?</p>",
      "rawMarkdown": "Seems obvious... Thanks for the tip. \n\nDomain knowledge for me is to understand mel-spectrogram or MFCC creation, choosing appropriate window size (I'm using 25 ms), resampling (I think can be good because don't need high frequencies to distinguish sounds), windowing, normalization (one mean and std based on train_cirated) - I don't see much impact using different parameters there.\nBut that's what you mean?",
      "votes": null
    },
    {
      "id": "524465",
      "postDate": "04/28/2019 20:30:20",
      "content": "<p>Domain knowledge helps to avoid common mistakes in architecture choices, data preprocessing, etc.</p>",
      "rawMarkdown": "Domain knowledge helps to avoid common mistakes in architecture choices, data preprocessing, etc.",
      "votes": null
    },
    {
      "id": "528785",
      "postDate": "05/08/2019 14:58:53",
      "content": "<p>Just by intuition I want to say that LSTM should outperform CNN if features extracted correctly. At least current 0.63 LB solutions with spectrogram cropping are far from ideal since network has very hard time to differentiate similar sounds.</p>\n\n<p>However it was fun to read <a href=\"http://ceur-ws.org/Vol-2125/paper_134.pdf\">http://ceur-ws.org/Vol-2125/paper_134.pdf</a></p>\n\n<blockquote>\n  <p>\"Due to multiple technical issues and the limited time frame, we were not able\n  to show that an LSTM can perform as well as a state of the art CNN in the\n  BirdCLEF challenge.\"</p>\n</blockquote>",
      "rawMarkdown": "Just by intuition I want to say that LSTM should outperform CNN if features extracted correctly. At least current 0.63 LB solutions with spectrogram cropping are far from ideal since network has very hard time to differentiate similar sounds.\n\nHowever it was fun to read http://ceur-ws.org/Vol-2125/paper_134.pdf\n&gt; \"Due to multiple technical issues and the limited time frame, we were not able\nto show that an LSTM can perform as well as a state of the art CNN in the\nBirdCLEF challenge.\"",
      "votes": null
    },
    {
      "id": "528886",
      "postDate": "05/08/2019 21:02:47",
      "content": "<p>It looks like I was concentrating on mel spectrogram only.\nProviding additional features should help our model to correctly identify cropped samples - as proposed here (past year solution) <a href=\"https://github.com/sainathadapa/kaggle-freesound-audio-tagging\">https://github.com/sainathadapa/kaggle-freesound-audio-tagging</a></p>",
      "rawMarkdown": "It looks like I was concentrating on mel spectrogram only.\nProviding additional features should help our model to correctly identify cropped samples - as proposed here (past year solution) https://github.com/sainathadapa/kaggle-freesound-audio-tagging",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 523981,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "04/27/2019 15:19:43",
      "content": "<p>I can get 0.65 with LSTM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 524069,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "04/27/2019 20:35:49",
          "content": "<p>That's unbelievable! How? Is it rather some specific preprocessing, noisy/curated pretraining, data augmentation. or some fancy LSTM architectures? I admire your scores</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524075,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "04/27/2019 21:12:21",
          "content": "<p>Nothing really special, I used a default LSTM implementation. With my main model that doesn't use any RNNs, I can get a score of 0.7 which is my current LB score. I don't use the noisy subset for the moment being. I don't do any specific preprocessing (but I plan to experiment with some other audio representations later). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524088,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "04/27/2019 22:18:17",
          "content": "<p>I am using melspectrogram with different GRU based networks (also very simple) and my CV scores are around 0.6 at best, with or without noisy data pretraining. That's why I don't get your post now - it looks like you do nothing and have a much better score than I am able to achieve. </p>\n\n<p>In this solution, the score is not that great too:\n<a href=\"https://www.kaggle.com/carlolepelaars/bidirectional-lstm-for-audio-labeling-with-keras\">https://www.kaggle.com/carlolepelaars/bidirectional-lstm-for-audio-labeling-with-keras</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524178,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "04/28/2019 06:28:14",
          "content": "<p>I didn't say I'm doing nothing :) I certainly did some improvements to a \"naive\" approach, but they are not really fancy (at least so far). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524179,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "04/28/2019 06:33:33",
          "content": "<p>My main advice would be to pay more attention on the problem statement and also to use domain knowledge (that's pretty obvious, right? :) )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524451,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "04/28/2019 19:17:50",
          "content": "<p>Seems obvious... Thanks for the tip. </p>\n\n<p>Domain knowledge for me is to understand mel-spectrogram or MFCC creation, choosing appropriate window size (I'm using 25 ms), resampling (I think can be good because don't need high frequencies to distinguish sounds), windowing, normalization (one mean and std based on train_cirated) - I don't see much impact using different parameters there.\nBut that's what you mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524465,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "04/28/2019 20:30:20",
          "content": "<p>Domain knowledge helps to avoid common mistakes in architecture choices, data preprocessing, etc.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528785,
      "author_name": "vandalko",
      "author_url": "",
      "post_date": "05/08/2019 14:58:53",
      "content": "<p>Just by intuition I want to say that LSTM should outperform CNN if features extracted correctly. At least current 0.63 LB solutions with spectrogram cropping are far from ideal since network has very hard time to differentiate similar sounds.</p>\n\n<p>However it was fun to read <a href=\"http://ceur-ws.org/Vol-2125/paper_134.pdf\">http://ceur-ws.org/Vol-2125/paper_134.pdf</a></p>\n\n<blockquote>\n  <p>\"Due to multiple technical issues and the limited time frame, we were not able\n  to show that an LSTM can perform as well as a state of the art CNN in the\n  BirdCLEF challenge.\"</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 528886,
          "author_name": "vandalko",
          "author_url": "",
          "post_date": "05/08/2019 21:02:47",
          "content": "<p>It looks like I was concentrating on mel spectrogram only.\nProviding additional features should help our model to correctly identify cropped samples - as proposed here (past year solution) <a href=\"https://github.com/sainathadapa/kaggle-freesound-audio-tagging\">https://github.com/sainathadapa/kaggle-freesound-audio-tagging</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "522913": "I can get a LB: 0.495 with LSTMs. Whats your best score?",
    "523981": "I can get 0.65 with LSTM.",
    "524069": "That's unbelievable! How? Is it rather some specific preprocessing, noisy/curated pretraining, data augmentation. or some fancy LSTM architectures? I admire your scores",
    "524075": "Nothing really special, I used a default LSTM implementation. With my main model that doesn't use any RNNs, I can get a score of 0.7 which is my current LB score. I don't use the noisy subset for the moment being. I don't do any specific preprocessing (but I plan to experiment with some other audio representations later).",
    "524088": "I am using melspectrogram with different GRU based networks (also very simple) and my CV scores are around 0.6 at best, with or without noisy data pretraining. That's why I don't get your post now - it looks like you do nothing and have a much better score than I am able to achieve. \n\nIn this solution, the score is not that great too:\nhttps://www.kaggle.com/carlolepelaars/bidirectional-lstm-for-audio-labeling-with-keras",
    "524178": "I didn't say I'm doing nothing :) I certainly did some improvements to a \"naive\" approach, but they are not really fancy (at least so far).",
    "524179": "My main advice would be to pay more attention on the problem statement and also to use domain knowledge (that's pretty obvious, right? :) )",
    "524451": "Seems obvious... Thanks for the tip. \n\nDomain knowledge for me is to understand mel-spectrogram or MFCC creation, choosing appropriate window size (I'm using 25 ms), resampling (I think can be good because don't need high frequencies to distinguish sounds), windowing, normalization (one mean and std based on train_cirated) - I don't see much impact using different parameters there.\nBut that's what you mean?",
    "524465": "Domain knowledge helps to avoid common mistakes in architecture choices, data preprocessing, etc.",
    "528785": "Just by intuition I want to say that LSTM should outperform CNN if features extracted correctly. At least current 0.63 LB solutions with spectrogram cropping are far from ideal since network has very hard time to differentiate similar sounds.\n\nHowever it was fun to read http://ceur-ws.org/Vol-2125/paper_134.pdf\n&gt; \"Due to multiple technical issues and the limited time frame, we were not able\nto show that an LSTM can perform as well as a state of the art CNN in the\nBirdCLEF challenge.\"",
    "528886": "It looks like I was concentrating on mel spectrogram only.\nProviding additional features should help our model to correctly identify cropped samples - as proposed here (past year solution) https://github.com/sainathadapa/kaggle-freesound-audio-tagging"
  },
  "source": "meta"
}