{
  "id": 53340,
  "title": "External Data Thread",
  "url": "/competitions/freesound-audio-tagging/discussion/53340",
  "author_name": "",
  "post_date": "2018-03-29T14:36:22.882969900Z",
  "votes": 5,
  "comment_count": 16,
  "views": 0,
  "content": "<p>As per the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/rules\">Rules</a>, if you are using external data, please post it to this thread.</p>",
  "messages": [
    {
      "id": "305865",
      "postDate": "03/29/2018 14:36:22",
      "content": "<p>As per the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/rules\">Rules</a>, if you are using external data, please post it to this thread.</p>",
      "rawMarkdown": "As per the [Rules](https://www.kaggle.com/c/freesound-audio-tagging/rules), if you are using external data, please post it to this thread.",
      "votes": null
    },
    {
      "id": "324080",
      "postDate": "05/07/2018 05:26:02",
      "content": "<p>Urban Sound datasets\n<a href=\"https://serv.cusp.nyu.edu/projects/urbansounddataset/\">https://serv.cusp.nyu.edu/projects/urbansounddataset/</a></p>",
      "rawMarkdown": "Urban Sound datasets\nhttps://serv.cusp.nyu.edu/projects/urbansounddataset/",
      "votes": null
    },
    {
      "id": "325295",
      "postDate": "05/08/2018 09:28:38",
      "content": "<p>Interesting find. However,  the paper linked on that site states that the source of this data set is freesound.org:</p>\n\n<p>\"For each class, we started by downloading all sounds returned by the Freesound search engine when using the class name as a query\" </p>\n\n<p>And they also state it on the website itself:</p>\n\n<p>\" All files come from www.freesound.org.\"</p>\n\n<p>That seems to be in conflict with the competition rules: </p>\n\n<p>\"The use of any data coming from Freesound (including audio files and/or metadata) other than that provided in the Competition Website is forbidden to participate in the competition.\"</p>\n\n<p>That brings up some questions (at least for me): </p>\n\n<ul>\n<li>How deep do we have to dig to ensure we're not (accidentally) using freesound data? </li>\n<li>Could we get some clarification from the organizers on that? </li>\n<li>Can we use the urban sound data set?</li>\n</ul>",
      "rawMarkdown": "Interesting find. However,  the paper linked on that site states that the source of this data set is freesound.org:\n\n\"For each class, we started by downloading all sounds returned by the Freesound search engine when using the class name as a query\" \n\nAnd they also state it on the website itself:\n\n\" All files come from www.freesound.org.\"\n\nThat seems to be in conflict with the competition rules: \n\n\"The use of any data coming from Freesound (including audio files and/or metadata) other than that provided in the Competition Website is forbidden to participate in the competition.\"\n\nThat brings up some questions (at least for me): \n\n - How deep do we have to dig to ensure we're not (accidentally) using freesound data? \n - Could we get some clarification from the organizers on that? \n - Can we use the urban sound data set?",
      "votes": null
    },
    {
      "id": "327889",
      "postDate": "05/12/2018 20:59:04",
      "content": "<p>After discussing among the organizers, we have decided not to allow the use of Urban Sound datasets as external data.</p>\n\n<p>Urban Sound datasets have the categories 'dog_bark' and 'gun_shot', while the dataset used in this competition (<strong>FSDKaggle2018</strong>) has the categories 'Bark' and 'Gunshot_or_gunfire'. Since the source of audio content for all these datasets is Freesound, there may potentially be audio files that coincide in the test set of FSDKaggle2018 and in the Urban Sound datasets.</p>\n\n<p>Therefore, the use of Urban Sound datasets is not allowed. For similar reasons, it is also not allowed to use other datasets that are composed of Freesound content, for example:</p>\n\n<ul>\n<li>ESC-50: <a href=\"https://github.com/karoldvl/ESC-50\">https://github.com/karoldvl/ESC-50</a></li>\n<li>freefield1010: <a href=\"https://c4dm.eecs.qmul.ac.uk/rdr/handle/123456789/35\">https://c4dm.eecs.qmul.ac.uk/rdr/handle/123456789/35</a></li>\n<li>and the mentioned Urban Sound datasets: <a href=\"https://serv.cusp.nyu.edu/projects/urbansounddataset/\">https://serv.cusp.nyu.edu/projects/urbansounddataset/</a> </li>\n</ul>\n\n<p>Responding to <em>steiml</em> : to ensure you're not (accidentally) using freesound data, you can check the origin of the dataset, just as you did. Typically this info will (or should) be available from where you download the dataset. The three companion sites for the datasets listed above clearly state that they are based on Freesound content. Plus, as external data must be posted in this thread, we (organizers) can also check that it is a valid source.</p>\n\n<p>Hope this clarifies!</p>",
      "rawMarkdown": "After discussing among the organizers, we have decided not to allow the use of Urban Sound datasets as external data.\n\nUrban Sound datasets have the categories 'dog_bark' and 'gun_shot', while the dataset used in this competition (**FSDKaggle2018**) has the categories 'Bark' and 'Gunshot_or_gunfire'. Since the source of audio content for all these datasets is Freesound, there may potentially be audio files that coincide in the test set of FSDKaggle2018 and in the Urban Sound datasets.\n\nTherefore, the use of Urban Sound datasets is not allowed. For similar reasons, it is also not allowed to use other datasets that are composed of Freesound content, for example:\n\n- ESC-50: https://github.com/karoldvl/ESC-50\n- freefield1010: https://c4dm.eecs.qmul.ac.uk/rdr/handle/123456789/35\n- and the mentioned Urban Sound datasets: https://serv.cusp.nyu.edu/projects/urbansounddataset/ \n\nResponding to *steiml* : to ensure you're not (accidentally) using freesound data, you can check the origin of the dataset, just as you did. Typically this info will (or should) be available from where you download the dataset. The three companion sites for the datasets listed above clearly state that they are based on Freesound content. Plus, as external data must be posted in this thread, we (organizers) can also check that it is a valid source.\n\nHope this clarifies!",
      "votes": null
    },
    {
      "id": "327928",
      "postDate": "05/12/2018 23:47:28",
      "content": "<p>Thanks you </p>\n\n<p>I've got another one: \n<a href=\"https://research.google.com/audioset/dataset/index.html\">https://research.google.com/audioset/dataset/index.html</a></p>\n\n<p>They are mainly videos....</p>",
      "rawMarkdown": "Thanks you \n\nI've got another one: \nhttps://research.google.com/audioset/dataset/index.html\n\nThey are mainly videos....",
      "votes": null
    },
    {
      "id": "328101",
      "postDate": "05/13/2018 10:44:37",
      "content": "<p>Yep, AudioSet is fine!</p>",
      "rawMarkdown": "Yep, AudioSet is fine!",
      "votes": null
    },
    {
      "id": "342019",
      "postDate": "06/12/2018 18:02:26",
      "content": "<p>What about publicly available pretrained models trained on other datasets?  For example, I have noticed some papers that apply pretrained imagenet models to images generated from the mel spectograms of audio data.</p>",
      "rawMarkdown": "What about publicly available pretrained models trained on other datasets?  For example, I have noticed some papers that apply pretrained imagenet models to images generated from the mel spectograms of audio data.",
      "votes": null
    },
    {
      "id": "343552",
      "postDate": "06/15/2018 14:49:49",
      "content": "<p>Hi Peter, thanks for your question. </p>\n\n<p>Publicly available pre-trained models are a valid option as long as they are suitably documented in this thread. Please specify which data the model was trained on, and make sure the data has nothing to do with Freesound, as per the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/rules\">competition rules</a>.</p>\n\n<p>We take the opportunity to remind participants that <strong>any kind of data</strong> used in this competition other than the provided dataset <strong>FSDKaggle2018</strong>, must be specified in this thread. This also includes pre-trained models.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Peter, thanks for your question. \n\nPublicly available pre-trained models are a valid option as long as they are suitably documented in this thread. Please specify which data the model was trained on, and make sure the data has nothing to do with Freesound, as per the [competition rules][1].\n\nWe take the opportunity to remind participants that **any kind of data** used in this competition other than the provided dataset **FSDKaggle2018**, must be specified in this thread. This also includes pre-trained models.\n\nThanks!\n\n\n  [1]: https://www.kaggle.com/c/freesound-audio-tagging/rules",
      "votes": null
    },
    {
      "id": "343794",
      "postDate": "06/16/2018 06:05:13",
      "content": "<p>Thanks for the reply.  I'll certainly post here If I end up using pretrained models or external data.</p>",
      "rawMarkdown": "Thanks for the reply.  I'll certainly post here If I end up using pretrained models or external data.",
      "votes": null
    },
    {
      "id": "343962",
      "postDate": "06/16/2018 15:39:23",
      "content": "<p>Peter, I'm curious about the paper you mentioned that applies an ImageNet-pretrained model to audio spectrograms. Do you have a reference handy?</p>",
      "rawMarkdown": "Peter, I'm curious about the paper you mentioned that applies an ImageNet-pretrained model to audio spectrograms. Do you have a reference handy?",
      "votes": null
    },
    {
      "id": "344155",
      "postDate": "06/17/2018 05:12:57",
      "content": "<p>Here's a paper I was looking at: <a href=\"https://www.researchgate.net/profile/Shahin_Amiriparian/publication/318987452_Snore_Sound_Classification_Using_Image-Based_Deep_Spectrum_Features/links/599d6f6c0f7e9b892bb3c78b/Snore-Sound-Classification-Using-Image-Based-Deep-Spectrum-Features.pdf\">https://www.researchgate.net/profile/Shahin_Amiriparian/publication/318987452_Snore_Sound_Classification_Using_Image-Based_Deep_Spectrum_Features/links/599d6f6c0f7e9b892bb3c78b/Snore-Sound-Classification-Using-Image-Based-Deep-Spectrum-Features.pdf</a></p>",
      "rawMarkdown": "Here's a paper I was looking at: https://www.researchgate.net/profile/Shahin_Amiriparian/publication/318987452_Snore_Sound_Classification_Using_Image-Based_Deep_Spectrum_Features/links/599d6f6c0f7e9b892bb3c78b/Snore-Sound-Classification-Using-Image-Based-Deep-Spectrum-Features.pdf",
      "votes": null
    },
    {
      "id": "355201",
      "postDate": "07/11/2018 07:28:50",
      "content": "<p>Can I use pre-trained ImageNet-based model? Such as ResNet, DenseNet?</p>",
      "rawMarkdown": "Can I use pre-trained ImageNet-based model? Such as ResNet, DenseNet?",
      "votes": null
    },
    {
      "id": "355244",
      "postDate": "07/11/2018 09:21:39",
      "content": "<p>Hi, </p>\n\n<p>thanks for asking. Yes, publicly available pre-trained models are a valid option. Please document here the pre-trained model that you are using (briefly, just few lines) so that other participants can understand it (including the data that was originally used for training). A link to the model/implementation can be helpful. </p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi, \n\nthanks for asking. Yes, publicly available pre-trained models are a valid option. Please document here the pre-trained model that you are using (briefly, just few lines) so that other participants can understand it (including the data that was originally used for training). A link to the model/implementation can be helpful. \n\nThanks!",
      "votes": null
    },
    {
      "id": "355365",
      "postDate": "07/11/2018 14:56:00",
      "content": "<ul>\n<li>The pretrained model depends on ImageNet 2014 image datasets, </li>\n<li>It can be found here:\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></li>\n<li>Presently, we have tested VGG/Inception/ResNet/DenseNet. But we have\nonly used one pretrained CNN model in our current submission.</li>\n</ul>",
      "rawMarkdown": "The pretrained model depends on ImageNet 2014 image datasets, \n - It can be found here:\n   https://github.com/Cadene/pretrained-models.pytorch\n - Presently, we have tested VGG/Inception/ResNet/DenseNet. But we have\n   only used one pretrained CNN model in our current submission.",
      "votes": null
    },
    {
      "id": "360644",
      "postDate": "07/23/2018 00:11:57",
      "content": "<p>Hello,\nI am experimenting with augmenting the training data by mixing it with background noise.  As background noise, I am using sounds from the BBC sound effects archive (<a href=\"http://bbcsfx.acropolis.org.uk/\">http://bbcsfx.acropolis.org.uk</a>).  Here are the clips I am using:</p>\n\n<p>Clip #  Description</p>\n\n<p>07029153 birds</p>\n\n<p>07045019 birds</p>\n\n<p>07063060 birds</p>\n\n<p>07030031 birds</p>\n\n<p>07062077 birds</p>\n\n<p>07074152 crowd</p>\n\n<p>07056062 ocean</p>\n\n<p>07056065 ocean</p>\n\n<p>07059111 ocean</p>\n\n<p>07045258 siren</p>\n\n<p>You can find these clips by entering the clip # into the search box on the BBC sound effects web page.</p>\n\n<p>As additional background noise, I am also using the pink_noise and white_noise clips from the Google Speech Commands dataset (<a href=\"https://ai.googleblog.com/2017/08/launching-speech-commands-dataset.html\">https://ai.googleblog.com/2017/08/launching-speech-commands-dataset.html</a>)</p>",
      "rawMarkdown": "Hello,\nI am experimenting with augmenting the training data by mixing it with background noise.  As background noise, I am using sounds from the BBC sound effects archive (http://bbcsfx.acropolis.org.uk).  Here are the clips I am using:\n\nClip #  Description\n\n07029153 birds\n\n07045019 birds\n\n07063060 birds\n\n07030031 birds\n\n07062077 birds\n\n07074152 crowd\n\n07056062 ocean\n\n07056065 ocean\n\n07059111 ocean\n\n07045258 siren\n\nYou can find these clips by entering the clip # into the search box on the BBC sound effects web page.\n\nAs additional background noise, I am also using the pink_noise and white_noise clips from the Google Speech Commands dataset (https://ai.googleblog.com/2017/08/launching-speech-commands-dataset.html)",
      "votes": null
    },
    {
      "id": "363889",
      "postDate": "07/30/2018 08:59:34",
      "content": "<p>Hello,</p>\n\n<p>I am training my model based on the pre-trained model on the image net dataset. Here is the model link: \n<a href=\"https://github.com/fchollet/deep-learning-models/releases\">https://github.com/fchollet/deep-learning-models/releases</a></p>",
      "rawMarkdown": "Hello,\n\nI am training my model based on the pre-trained model on the image net dataset. Here is the model link: \nhttps://github.com/fchollet/deep-learning-models/releases",
      "votes": null
    },
    {
      "id": "364563",
      "postDate": "07/31/2018 19:16:07",
      "content": "<p>I'm using the weights from the MobileNet V2 model trained on the ImageNet dataset. The weights are available at <a href=\"https://github.com/fchollet/deep-learning-models/releases\">https://github.com/fchollet/deep-learning-models/releases</a>.</p>",
      "rawMarkdown": "I'm using the weights from the MobileNet V2 model trained on the ImageNet dataset. The weights are available at https://github.com/fchollet/deep-learning-models/releases.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 324080,
      "author_name": "ejlok1",
      "author_url": "",
      "post_date": "05/07/2018 05:26:02",
      "content": "<p>Urban Sound datasets\n<a href=\"https://serv.cusp.nyu.edu/projects/urbansounddataset/\">https://serv.cusp.nyu.edu/projects/urbansounddataset/</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 325295,
          "author_name": "steiml",
          "author_url": "",
          "post_date": "05/08/2018 09:28:38",
          "content": "<p>Interesting find. However,  the paper linked on that site states that the source of this data set is freesound.org:</p>\n\n<p>\"For each class, we started by downloading all sounds returned by the Freesound search engine when using the class name as a query\" </p>\n\n<p>And they also state it on the website itself:</p>\n\n<p>\" All files come from www.freesound.org.\"</p>\n\n<p>That seems to be in conflict with the competition rules: </p>\n\n<p>\"The use of any data coming from Freesound (including audio files and/or metadata) other than that provided in the Competition Website is forbidden to participate in the competition.\"</p>\n\n<p>That brings up some questions (at least for me): </p>\n\n<ul>\n<li>How deep do we have to dig to ensure we're not (accidentally) using freesound data? </li>\n<li>Could we get some clarification from the organizers on that? </li>\n<li>Can we use the urban sound data set?</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327889,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "05/12/2018 20:59:04",
      "content": "<p>After discussing among the organizers, we have decided not to allow the use of Urban Sound datasets as external data.</p>\n\n<p>Urban Sound datasets have the categories 'dog_bark' and 'gun_shot', while the dataset used in this competition (<strong>FSDKaggle2018</strong>) has the categories 'Bark' and 'Gunshot_or_gunfire'. Since the source of audio content for all these datasets is Freesound, there may potentially be audio files that coincide in the test set of FSDKaggle2018 and in the Urban Sound datasets.</p>\n\n<p>Therefore, the use of Urban Sound datasets is not allowed. For similar reasons, it is also not allowed to use other datasets that are composed of Freesound content, for example:</p>\n\n<ul>\n<li>ESC-50: <a href=\"https://github.com/karoldvl/ESC-50\">https://github.com/karoldvl/ESC-50</a></li>\n<li>freefield1010: <a href=\"https://c4dm.eecs.qmul.ac.uk/rdr/handle/123456789/35\">https://c4dm.eecs.qmul.ac.uk/rdr/handle/123456789/35</a></li>\n<li>and the mentioned Urban Sound datasets: <a href=\"https://serv.cusp.nyu.edu/projects/urbansounddataset/\">https://serv.cusp.nyu.edu/projects/urbansounddataset/</a> </li>\n</ul>\n\n<p>Responding to <em>steiml</em> : to ensure you're not (accidentally) using freesound data, you can check the origin of the dataset, just as you did. Typically this info will (or should) be available from where you download the dataset. The three companion sites for the datasets listed above clearly state that they are based on Freesound content. Plus, as external data must be posted in this thread, we (organizers) can also check that it is a valid source.</p>\n\n<p>Hope this clarifies!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327928,
      "author_name": "ejlok1",
      "author_url": "",
      "post_date": "05/12/2018 23:47:28",
      "content": "<p>Thanks you </p>\n\n<p>I've got another one: \n<a href=\"https://research.google.com/audioset/dataset/index.html\">https://research.google.com/audioset/dataset/index.html</a></p>\n\n<p>They are mainly videos....</p>",
      "votes": null,
      "replies": [
        {
          "id": 328101,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "05/13/2018 10:44:37",
          "content": "<p>Yep, AudioSet is fine!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 342019,
      "author_name": "paiforsyth",
      "author_url": "",
      "post_date": "06/12/2018 18:02:26",
      "content": "<p>What about publicly available pretrained models trained on other datasets?  For example, I have noticed some papers that apply pretrained imagenet models to images generated from the mel spectograms of audio data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 343552,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "06/15/2018 14:49:49",
          "content": "<p>Hi Peter, thanks for your question. </p>\n\n<p>Publicly available pre-trained models are a valid option as long as they are suitably documented in this thread. Please specify which data the model was trained on, and make sure the data has nothing to do with Freesound, as per the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/rules\">competition rules</a>.</p>\n\n<p>We take the opportunity to remind participants that <strong>any kind of data</strong> used in this competition other than the provided dataset <strong>FSDKaggle2018</strong>, must be specified in this thread. This also includes pre-trained models.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 343794,
          "author_name": "paiforsyth",
          "author_url": "",
          "post_date": "06/16/2018 06:05:13",
          "content": "<p>Thanks for the reply.  I'll certainly post here If I end up using pretrained models or external data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 343962,
          "author_name": "plakal",
          "author_url": "",
          "post_date": "06/16/2018 15:39:23",
          "content": "<p>Peter, I'm curious about the paper you mentioned that applies an ImageNet-pretrained model to audio spectrograms. Do you have a reference handy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344155,
          "author_name": "paiforsyth",
          "author_url": "",
          "post_date": "06/17/2018 05:12:57",
          "content": "<p>Here's a paper I was looking at: <a href=\"https://www.researchgate.net/profile/Shahin_Amiriparian/publication/318987452_Snore_Sound_Classification_Using_Image-Based_Deep_Spectrum_Features/links/599d6f6c0f7e9b892bb3c78b/Snore-Sound-Classification-Using-Image-Based-Deep-Spectrum-Features.pdf\">https://www.researchgate.net/profile/Shahin_Amiriparian/publication/318987452_Snore_Sound_Classification_Using_Image-Based_Deep_Spectrum_Features/links/599d6f6c0f7e9b892bb3c78b/Snore-Sound-Classification-Using-Image-Based-Deep-Spectrum-Features.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 355201,
      "author_name": "kelexu",
      "author_url": "",
      "post_date": "07/11/2018 07:28:50",
      "content": "<p>Can I use pre-trained ImageNet-based model? Such as ResNet, DenseNet?</p>",
      "votes": null,
      "replies": [
        {
          "id": 355244,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "07/11/2018 09:21:39",
          "content": "<p>Hi, </p>\n\n<p>thanks for asking. Yes, publicly available pre-trained models are a valid option. Please document here the pre-trained model that you are using (briefly, just few lines) so that other participants can understand it (including the data that was originally used for training). A link to the model/implementation can be helpful. </p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355365,
          "author_name": "kelexu",
          "author_url": "",
          "post_date": "07/11/2018 14:56:00",
          "content": "<ul>\n<li>The pretrained model depends on ImageNet 2014 image datasets, </li>\n<li>It can be found here:\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a></li>\n<li>Presently, we have tested VGG/Inception/ResNet/DenseNet. But we have\nonly used one pretrained CNN model in our current submission.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 360644,
      "author_name": "paiforsyth",
      "author_url": "",
      "post_date": "07/23/2018 00:11:57",
      "content": "<p>Hello,\nI am experimenting with augmenting the training data by mixing it with background noise.  As background noise, I am using sounds from the BBC sound effects archive (<a href=\"http://bbcsfx.acropolis.org.uk/\">http://bbcsfx.acropolis.org.uk</a>).  Here are the clips I am using:</p>\n\n<p>Clip #  Description</p>\n\n<p>07029153 birds</p>\n\n<p>07045019 birds</p>\n\n<p>07063060 birds</p>\n\n<p>07030031 birds</p>\n\n<p>07062077 birds</p>\n\n<p>07074152 crowd</p>\n\n<p>07056062 ocean</p>\n\n<p>07056065 ocean</p>\n\n<p>07059111 ocean</p>\n\n<p>07045258 siren</p>\n\n<p>You can find these clips by entering the clip # into the search box on the BBC sound effects web page.</p>\n\n<p>As additional background noise, I am also using the pink_noise and white_noise clips from the Google Speech Commands dataset (<a href=\"https://ai.googleblog.com/2017/08/launching-speech-commands-dataset.html\">https://ai.googleblog.com/2017/08/launching-speech-commands-dataset.html</a>)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 363889,
      "author_name": "miao007",
      "author_url": "",
      "post_date": "07/30/2018 08:59:34",
      "content": "<p>Hello,</p>\n\n<p>I am training my model based on the pre-trained model on the image net dataset. Here is the model link: \n<a href=\"https://github.com/fchollet/deep-learning-models/releases\">https://github.com/fchollet/deep-learning-models/releases</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 364563,
      "author_name": "sainathadapa",
      "author_url": "",
      "post_date": "07/31/2018 19:16:07",
      "content": "<p>I'm using the weights from the MobileNet V2 model trained on the ImageNet dataset. The weights are available at <a href=\"https://github.com/fchollet/deep-learning-models/releases\">https://github.com/fchollet/deep-learning-models/releases</a>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "305865": "As per the [Rules](https://www.kaggle.com/c/freesound-audio-tagging/rules), if you are using external data, please post it to this thread.",
    "324080": "Urban Sound datasets\nhttps://serv.cusp.nyu.edu/projects/urbansounddataset/",
    "325295": "Interesting find. However,  the paper linked on that site states that the source of this data set is freesound.org:\n\n\"For each class, we started by downloading all sounds returned by the Freesound search engine when using the class name as a query\" \n\nAnd they also state it on the website itself:\n\n\" All files come from www.freesound.org.\"\n\nThat seems to be in conflict with the competition rules: \n\n\"The use of any data coming from Freesound (including audio files and/or metadata) other than that provided in the Competition Website is forbidden to participate in the competition.\"\n\nThat brings up some questions (at least for me): \n\n - How deep do we have to dig to ensure we're not (accidentally) using freesound data? \n - Could we get some clarification from the organizers on that? \n - Can we use the urban sound data set?",
    "327889": "After discussing among the organizers, we have decided not to allow the use of Urban Sound datasets as external data.\n\nUrban Sound datasets have the categories 'dog_bark' and 'gun_shot', while the dataset used in this competition (**FSDKaggle2018**) has the categories 'Bark' and 'Gunshot_or_gunfire'. Since the source of audio content for all these datasets is Freesound, there may potentially be audio files that coincide in the test set of FSDKaggle2018 and in the Urban Sound datasets.\n\nTherefore, the use of Urban Sound datasets is not allowed. For similar reasons, it is also not allowed to use other datasets that are composed of Freesound content, for example:\n\n- ESC-50: https://github.com/karoldvl/ESC-50\n- freefield1010: https://c4dm.eecs.qmul.ac.uk/rdr/handle/123456789/35\n- and the mentioned Urban Sound datasets: https://serv.cusp.nyu.edu/projects/urbansounddataset/ \n\nResponding to *steiml* : to ensure you're not (accidentally) using freesound data, you can check the origin of the dataset, just as you did. Typically this info will (or should) be available from where you download the dataset. The three companion sites for the datasets listed above clearly state that they are based on Freesound content. Plus, as external data must be posted in this thread, we (organizers) can also check that it is a valid source.\n\nHope this clarifies!",
    "327928": "Thanks you \n\nI've got another one: \nhttps://research.google.com/audioset/dataset/index.html\n\nThey are mainly videos....",
    "328101": "Yep, AudioSet is fine!",
    "342019": "What about publicly available pretrained models trained on other datasets?  For example, I have noticed some papers that apply pretrained imagenet models to images generated from the mel spectograms of audio data.",
    "343552": "Hi Peter, thanks for your question. \n\nPublicly available pre-trained models are a valid option as long as they are suitably documented in this thread. Please specify which data the model was trained on, and make sure the data has nothing to do with Freesound, as per the [competition rules][1].\n\nWe take the opportunity to remind participants that **any kind of data** used in this competition other than the provided dataset **FSDKaggle2018**, must be specified in this thread. This also includes pre-trained models.\n\nThanks!\n\n\n  [1]: https://www.kaggle.com/c/freesound-audio-tagging/rules",
    "343794": "Thanks for the reply.  I'll certainly post here If I end up using pretrained models or external data.",
    "343962": "Peter, I'm curious about the paper you mentioned that applies an ImageNet-pretrained model to audio spectrograms. Do you have a reference handy?",
    "344155": "Here's a paper I was looking at: https://www.researchgate.net/profile/Shahin_Amiriparian/publication/318987452_Snore_Sound_Classification_Using_Image-Based_Deep_Spectrum_Features/links/599d6f6c0f7e9b892bb3c78b/Snore-Sound-Classification-Using-Image-Based-Deep-Spectrum-Features.pdf",
    "355201": "Can I use pre-trained ImageNet-based model? Such as ResNet, DenseNet?",
    "355244": "Hi, \n\nthanks for asking. Yes, publicly available pre-trained models are a valid option. Please document here the pre-trained model that you are using (briefly, just few lines) so that other participants can understand it (including the data that was originally used for training). A link to the model/implementation can be helpful. \n\nThanks!",
    "355365": "The pretrained model depends on ImageNet 2014 image datasets, \n - It can be found here:\n   https://github.com/Cadene/pretrained-models.pytorch\n - Presently, we have tested VGG/Inception/ResNet/DenseNet. But we have\n   only used one pretrained CNN model in our current submission.",
    "360644": "Hello,\nI am experimenting with augmenting the training data by mixing it with background noise.  As background noise, I am using sounds from the BBC sound effects archive (http://bbcsfx.acropolis.org.uk).  Here are the clips I am using:\n\nClip #  Description\n\n07029153 birds\n\n07045019 birds\n\n07063060 birds\n\n07030031 birds\n\n07062077 birds\n\n07074152 crowd\n\n07056062 ocean\n\n07056065 ocean\n\n07059111 ocean\n\n07045258 siren\n\nYou can find these clips by entering the clip # into the search box on the BBC sound effects web page.\n\nAs additional background noise, I am also using the pink_noise and white_noise clips from the Google Speech Commands dataset (https://ai.googleblog.com/2017/08/launching-speech-commands-dataset.html)",
    "363889": "Hello,\n\nI am training my model based on the pre-trained model on the image net dataset. Here is the model link: \nhttps://github.com/fchollet/deep-learning-models/releases",
    "364563": "I'm using the weights from the MobileNet V2 model trained on the ImageNet dataset. The weights are available at https://github.com/fchollet/deep-learning-models/releases."
  },
  "source": "meta"
}