{
  "id": 87960,
  "title": "Pretrained models and other recycling ideas",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/87960",
  "author_name": "Konrad Banachewicz",
  "post_date": "2019-04-04T18:05:13.381000",
  "votes": 16,
  "comment_count": 24,
  "views": 0,
  "content": "<p>The topic is going to surface pretty soon, so in order to save people some time a few quick references :\n* pretrained PyTorch models (mostly image-based, useful for working with spectrograms, STFT and the like) <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\n* pretrained Keras models: <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>\n* DeepSpeech might come in handy, I guess <a href=\"https://github.com/SeanNaren/deepspeech.pytorch\">https://github.com/SeanNaren/deepspeech.pytorch</a>\n* Baidu: <a href=\"https://github.com/baidu-research/warp-ctc\">https://github.com/baidu-research/warp-ctc</a>\n* speech recognition challenge <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge</a></p>\n\n<p>I'll try to add more if I find other interesting stuff.</p>\n\n<p>EDIT: Competition-specific rules prohibit pretrained models (h/t Addison for drawing my attention to that tiny detail), so use the list at your own risk :-) </p>",
  "messages": [
    {
      "id": 507463,
      "postDate": "2019-04-04T18:05:13.380Z",
      "content": "<p>The topic is going to surface pretty soon, so in order to save people some time a few quick references :\n* pretrained PyTorch models (mostly image-based, useful for working with spectrograms, STFT and the like) <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\n* pretrained Keras models: <a href=\"https://keras.io/applications/\">https://keras.io/applications/</a>\n* DeepSpeech might come in handy, I guess <a href=\"https://github.com/SeanNaren/deepspeech.pytorch\">https://github.com/SeanNaren/deepspeech.pytorch</a>\n* Baidu: <a href=\"https://github.com/baidu-research/warp-ctc\">https://github.com/baidu-research/warp-ctc</a>\n* speech recognition challenge <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge</a></p>\n\n<p>I'll try to add more if I find other interesting stuff.</p>\n\n<p>EDIT: Competition-specific rules prohibit pretrained models (h/t Addison for drawing my attention to that tiny detail), so use the list at your own risk :-) </p>",
      "rawMarkdown": "The topic is going to surface pretty soon, so in order to save people some time a few quick references :\n* pretrained PyTorch models (mostly image-based, useful for working with spectrograms, STFT and the like) https://github.com/Cadene/pretrained-models.pytorch\n* pretrained Keras models: https://keras.io/applications/\n* DeepSpeech might come in handy, I guess https://github.com/SeanNaren/deepspeech.pytorch\n* Baidu: https://github.com/baidu-research/warp-ctc\n* speech recognition challenge https://www.kaggle.com/c/tensorflow-speech-recognition-challenge\n\nI'll try to add more if I find other interesting stuff.\n\nEDIT: Competition-specific rules prohibit pretrained models (h/t Addison for drawing my attention to that tiny detail), so use the list at your own risk :-) ",
      "votes": 16
    },
    {
      "id": 507466,
      "postDate": "2019-04-04T18:10:37.010Z",
      "content": "<p>Hi Konrad,</p>\n\n<p>It should be noted that per the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/rules\">rules</a> of the competition, external data and pretrained models are not allowed.</p>",
      "rawMarkdown": "Hi Konrad,\n\nIt should be noted that per the [rules](https://www.kaggle.com/c/freesound-audio-tagging-2019/rules) of the competition, external data and pretrained models are not allowed.",
      "votes": 5,
      "replies": [
        {
          "id": 507477,
          "postDate": "2019-04-04T18:28:45.227Z",
          "content": "<p>My bad - I didn't read all the way to \"Competition-specific' ones. Should I delete the post?</p>",
          "rawMarkdown": "My bad - I didn't read all the way to \"Competition-specific' ones. Should I delete the post?"
        },
        {
          "id": 507479,
          "postDate": "2019-04-04T18:31:27.833Z",
          "content": "<p>I think it's fine to keep it - I'm sure you're not the only one who may have missed it so it's good for posterity!</p>",
          "rawMarkdown": "I think it's fine to keep it - I'm sure you're not the only one who may have missed it so it's good for posterity!"
        },
        {
          "id": 507499,
          "postDate": "2019-04-04T18:55:12.337Z",
          "content": "<p>Hi Addison, the Competition-Specific Rules talks about \"external data\", but does not make any explicit mention about pre-trained models. Or can both be technically considered as same? Apologies if the question sounds very naive!</p>",
          "rawMarkdown": "Hi Addison, the Competition-Specific Rules talks about \"external data\", but does not make any explicit mention about pre-trained models. Or can both be technically considered as same? Apologies if the question sounds very naive!",
          "votes": 1
        },
        {
          "id": 507529,
          "postDate": "2019-04-04T20:29:17.137Z",
          "content": "<p>Not a naive question - we've wrestled with the same topic on our team as well. We say they're considered as same, as the pre-trained models are themselves trained on external data. In some competitions we've prohibited one, but excluded the other, in which case we provide a whitelist for consideration</p>",
          "rawMarkdown": "Not a naive question - we've wrestled with the same topic on our team as well. We say they're considered as same, as the pre-trained models are themselves trained on external data. In some competitions we've prohibited one, but excluded the other, in which case we provide a whitelist for consideration",
          "votes": 1
        },
        {
          "id": 507540,
          "postDate": "2019-04-04T20:50:32.617Z",
          "content": "<p>Thank you for the clarification, Addison!</p>",
          "rawMarkdown": "Thank you for the clarification, Addison!"
        },
        {
          "id": 507582,
          "postDate": "2019-04-04T22:38:27.793Z",
          "content": "<p>Question about \"pre-trained models\" - I assume that one of the typical pre-trained models that IS NOT TRAINED can be used?</p>\n\n<p>Almost all of these pre-trained can easily be used without using the pre-trained weights and train from scratch with the allowed data.  </p>",
          "rawMarkdown": "Question about \"pre-trained models\" - I assume that one of the typical pre-trained models that IS NOT TRAINED can be used?\n\nAlmost all of these pre-trained can easily be used without using the pre-trained weights and train from scratch with the allowed data.  "
        },
        {
          "id": 507791,
          "postDate": "2019-04-05T07:46:48.253Z",
          "content": "<p>Just to be sure - we need to train models from scratch and make inference within 1 hour of GPU time?</p>",
          "rawMarkdown": "Just to be sure - we need to train models from scratch and make inference within 1 hour of GPU time?"
        },
        {
          "id": 508250,
          "postDate": "2019-04-05T21:28:51.670Z",
          "content": "<p>Hi Andrew, \nas explained in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data</a> section:</p>\n\n<p>Submissions must be made with inference models running in Kaggle Kernels. However, participants can decide to train also in the Kaggle Kernels or offline (see <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements\">Kernels Requirements</a> for details).</p>\n\n<p>So, you can train in your local server if you wish.</p>\n\n<p>thanks!</p>",
          "rawMarkdown": "Hi Andrew, \nas explained in the [Data](https://www.kaggle.com/c/freesound-audio-tagging-2019/data) section:\n\nSubmissions must be made with inference models running in Kaggle Kernels. However, participants can decide to train also in the Kaggle Kernels or offline (see [Kernels Requirements](https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements) for details).\n\nSo, you can train in your local server if you wish.\n\nthanks!",
          "votes": 1
        },
        {
          "id": 508252,
          "postDate": "2019-04-05T21:36:32.950Z",
          "content": "<p>Jimmy,  I'm not sure I understand the question. A pre-trained model that is not trained? do you refer to using standard openly-available architectures, like VGG? </p>\n\n<p>thanks!</p>",
          "rawMarkdown": "Jimmy,  I'm not sure I understand the question. A pre-trained model that is not trained? do you refer to using standard openly-available architectures, like VGG? \n\nthanks!"
        },
        {
          "id": 508641,
          "postDate": "2019-04-06T14:31:49.150Z",
          "content": "<p>Yes - almost every pre-trained model I have used had the option to NOT use the trained weights and start from scratch.  Can we use these if we don't use the already established weights?</p>",
          "rawMarkdown": "Yes - almost every pre-trained model I have used had the option to NOT use the trained weights and start from scratch.  Can we use these if we don't use the already established weights?"
        },
        {
          "id": 509002,
          "postDate": "2019-04-07T06:53:41.530Z",
          "content": "<p>Hi Eduardo (<a href=\"/eduardofonseca\">@eduardofonseca</a>), excuse me but let me confirm this to make it sure I will not violate rules.\nAs you've clarified as follows, <strong>we can train as much as possible only if</strong>:\n- Training from scratch.\n- Only using dataset (train_curated, train_noisy <strong>and test</strong>) provided.\n- Locally trained model can be used as a kaggle dataset.</p>\n\n<p>&gt; Submissions must be made with inference models running in Kaggle Kernels. However, participants &gt; can decide to train also in the Kaggle Kernels or offline (see Kernels Requirements for details).\n&gt; So, you can train in your local server if you wish.</p>\n\n<p>Thank you.</p>",
          "rawMarkdown": "Hi Eduardo (@eduardofonseca), excuse me but let me confirm this to make it sure I will not violate rules.\nAs you've clarified as follows, __we can train as much as possible only if__:\n- Training from scratch.\n- Only using dataset (train_curated, train_noisy __and test__) provided.\n- Locally trained model can be used as a kaggle dataset.\n\n&gt; Submissions must be made with inference models running in Kaggle Kernels. However, participants &gt; can decide to train also in the Kaggle Kernels or offline (see Kernels Requirements for details).\n&gt; So, you can train in your local server if you wish.\n\nThank you.",
          "votes": 1
        },
        {
          "id": 509358,
          "postDate": "2019-04-07T17:48:17.510Z",
          "content": "<p>Also re-iterating daisukelab's question:\nIf I train a model (say, a neural net) offline on the train data provided, and then upload the weights as a Kaggle dataset, I am allowed to use that in my Kernel?</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "Also re-iterating daisukelab's question:\nIf I train a model (say, a neural net) offline on the train data provided, and then upload the weights as a Kaggle dataset, I am allowed to use that in my Kernel?\n\nThanks!",
          "votes": 1
        },
        {
          "id": 510194,
          "postDate": "2019-04-08T19:50:09.610Z",
          "content": "<p>I also can't get this - so it's allowed to train model offline and just perform the inference in the kernel? What is the point of kernel competition then? Just to be sure a given amount of data can be classified in 4 CPU hours?</p>",
          "rawMarkdown": "I also can't get this - so it's allowed to train model offline and just perform the inference in the kernel? What is the point of kernel competition then? Just to be sure a given amount of data can be classified in 4 CPU hours?"
        },
        {
          "id": 510203,
          "postDate": "2019-04-08T20:03:23.590Z",
          "content": "<p>Hi George <a href=\"/kunstmord\">@kunstmord</a> and DavidS <a href=\"/davids1992\">@davids1992</a>, he already have clarified here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#508071\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#508071</a></p>\n\n<p>So the answer is yes, we can train locally as long as using the data provided only.\nWe are not supposed to use any other data, including models pretrained with other dataset like ImageNet.</p>",
          "rawMarkdown": "Hi George @kunstmord and DavidS @davids1992, he already have clarified here:\n\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#508071\n\nSo the answer is yes, we can train locally as long as using the data provided only.\nWe are not supposed to use any other data, including models pretrained with other dataset like ImageNet.",
          "votes": 1
        },
        {
          "id": 510221,
          "postDate": "2019-04-08T20:38:48.707Z",
          "content": "<p>as pointed out above by <a href=\"/daisukelab\">@daisukelab</a> (thanks!), please refer to this official thread for the discussion of the rules\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064</a></p>\n\n<p><a href=\"/davids1992\">@davids1992</a> last year we ran an offline challenge without kernels and we received some submissions with large models and ensembles. 30-40-way ensembles were not uncommon in top leaderboard positions with model sizes going over 55M weights (for comparison, a ResNet-50 is around 30-35M weights). Such large models and ensembles are not interesting from a research point of view since you can always play games with large models and ensembles to extract a few extra points but these don't generalize to other tasks or to real-world uses. Such models could not run in real-time on a mid-range smartphone or low-power Internet-of-things device in your home, for example, and are expensive to run even on a server. We'd like to gently push participants towards figuring out smarter models and better ways to use the data. The limits this year are very generous, and we think 4 hours of CPU time or 1 hour of GPU time is more than enough to run inference on a decent model, while hopefully making it harder to just submit an ensemble of 100 ResNets :)  As a point of comparison, the Challenge Baseline can generate predictions on the entire test set in under 5 GPU-minutes.</p>",
          "rawMarkdown": "as pointed out above by @daisukelab (thanks!), please refer to this official thread for the discussion of the rules\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064\n\n@davids1992 last year we ran an offline challenge without kernels and we received some submissions with large models and ensembles. 30-40-way ensembles were not uncommon in top leaderboard positions with model sizes going over 55M weights (for comparison, a ResNet-50 is around 30-35M weights). Such large models and ensembles are not interesting from a research point of view since you can always play games with large models and ensembles to extract a few extra points but these don't generalize to other tasks or to real-world uses. Such models could not run in real-time on a mid-range smartphone or low-power Internet-of-things device in your home, for example, and are expensive to run even on a server. We'd like to gently push participants towards figuring out smarter models and better ways to use the data. The limits this year are very generous, and we think 4 hours of CPU time or 1 hour of GPU time is more than enough to run inference on a decent model, while hopefully making it harder to just submit an ensemble of 100 ResNets :)  As a point of comparison, the Challenge Baseline can generate predictions on the entire test set in under 5 GPU-minutes.",
          "votes": 3
        },
        {
          "id": 510401,
          "postDate": "2019-04-09T03:21:05.083Z",
          "content": "<p>From a production point of view I actually agree with this competition's format. Inference time and model size are the things that matter the most. A model may take weeks to train, but as long as you can make it small and predict in real time it's still a wonderful trade off.</p>",
          "rawMarkdown": "From a production point of view I actually agree with this competition's format. Inference time and model size are the things that matter the most. A model may take weeks to train, but as long as you can make it small and predict in real time it's still a wonderful trade off."
        },
        {
          "id": 510522,
          "postDate": "2019-04-09T07:11:37.483Z",
          "content": "<p>I understand, that's indeed well thought out, thanks!</p>",
          "rawMarkdown": "I understand, that's indeed well thought out, thanks!"
        },
        {
          "id": 514867,
          "postDate": "2019-04-12T01:45:07.550Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> , quick comment: to train your system you can use\n- train_curated\n- train_noisy</p>\n\n<p>But <strong>the test set cannot</strong> be used to train the submitted system.  Please refer to this official thread for the discussion of the rules:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064</a></p>",
          "rawMarkdown": "@daisukelab , quick comment: to train your system you can use\n- train_curated\n- train_noisy\n\nBut **the test set cannot** be used to train the submitted system.  Please refer to this official thread for the discussion of the rules:\n[https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064)",
          "votes": 1
        },
        {
          "id": 514890,
          "postDate": "2019-04-12T02:27:41.260Z",
          "content": "<p><a href=\"/eduardofonseca\">@eduardofonseca</a> Thank you for clarification and link, I appreciate.</p>",
          "rawMarkdown": "@eduardofonseca Thank you for clarification and link, I appreciate."
        }
      ]
    },
    {
      "id": 507790,
      "postDate": "2019-04-05T07:45:32.380Z",
      "content": "<p>̶A̶̶l̶̶t̶̶h̶̶o̶̶u̶̶g̶̶h̶ ̶p̶̶r̶̶o̶̶h̶̶i̶̶b̶̶i̶̶t̶̶e̶̶d̶ ̶f̶̶r̶̶o̶̶m̶ ̶m̶̶y̶ ̶e̶̶x̶̶p̶̶e̶̶r̶̶i̶̶e̶̶n̶̶c̶̶e̶ ̶s̶̶o̶ ̶f̶̶a̶̶r̶ ̶I̶̶m̶̶a̶̶g̶̶e̶̶N̶̶e̶̶t̶ ̶m̶̶o̶̶d̶̶e̶̶l̶̶s̶ ̶h̶̶a̶̶v̶̶e̶ ̶n̶̶o̶ ̶a̶̶d̶̶v̶̶a̶̶n̶̶t̶̶a̶̶g̶̶e̶̶s̶ ̶o̶̶v̶̶e̶̶r̶ ̶t̶̶r̶̶a̶̶i̶̶n̶̶i̶̶n̶̶g̶ ̶f̶̶r̶̶o̶̶m̶ ̶s̶̶c̶̶r̶̶a̶̶t̶̶c̶̶h̶ ̶a̶̶n̶̶y̶̶w̶̶a̶̶y̶̶.̶</p>",
      "rawMarkdown": "̶A̶̶l̶̶t̶̶h̶̶o̶̶u̶̶g̶̶h̶ ̶p̶̶r̶̶o̶̶h̶̶i̶̶b̶̶i̶̶t̶̶e̶̶d̶ ̶f̶̶r̶̶o̶̶m̶ ̶m̶̶y̶ ̶e̶̶x̶̶p̶̶e̶̶r̶̶i̶̶e̶̶n̶̶c̶̶e̶ ̶s̶̶o̶ ̶f̶̶a̶̶r̶ ̶I̶̶m̶̶a̶̶g̶̶e̶̶N̶̶e̶̶t̶ ̶m̶̶o̶̶d̶̶e̶̶l̶̶s̶ ̶h̶̶a̶̶v̶̶e̶ ̶n̶̶o̶ ̶a̶̶d̶̶v̶̶a̶̶n̶̶t̶̶a̶̶g̶̶e̶̶s̶ ̶o̶̶v̶̶e̶̶r̶ ̶t̶̶r̶̶a̶̶i̶̶n̶̶i̶̶n̶̶g̶ ̶f̶̶r̶̶o̶̶m̶ ̶s̶̶c̶̶r̶̶a̶̶t̶̶c̶̶h̶ ̶a̶̶n̶̶y̶̶w̶̶a̶̶y̶̶.̶",
      "replies": [
        {
          "id": 508264,
          "postDate": "2019-04-05T22:12:33.733Z",
          "content": "<p>Out of curiosity, what architectures have you tried? Using DenseNet-201 with all layers fine-tuned always gave me a slight improvement over training from scratch (not necessarily constrained to DenseNet). Based on preliminary experiments, I've observed the same with this task too.</p>",
          "rawMarkdown": "Out of curiosity, what architectures have you tried? Using DenseNet-201 with all layers fine-tuned always gave me a slight improvement over training from scratch (not necessarily constrained to DenseNet). Based on preliminary experiments, I've observed the same with this task too."
        },
        {
          "id": 508291,
          "postDate": "2019-04-05T23:40:25.933Z",
          "content": "<p>I may want to take that statement back. I had a bug in my validation that made the score of the pretrained model lower that it really is. Looks like ImageNet models are much stronger for this task for some reason.</p>",
          "rawMarkdown": "I may want to take that statement back. I had a bug in my validation that made the score of the pretrained model lower that it really is. Looks like ImageNet models are much stronger for this task for some reason."
        }
      ]
    },
    {
      "id": 507497,
      "postDate": "2019-04-04T18:54:41.007Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 507466,
      "author_name": "Addison Howard",
      "author_url": "",
      "post_date": "2019-04-04T18:10:37.010000",
      "content": "<p>Hi Konrad,</p>\n\n<p>It should be noted that per the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/rules\">rules</a> of the competition, external data and pretrained models are not allowed.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 507477,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2019-04-04T18:28:45.227000",
          "content": "<p>My bad - I didn't read all the way to \"Competition-specific' ones. Should I delete the post?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 507479,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-04-04T18:31:27.833000",
          "content": "<p>I think it's fine to keep it - I'm sure you're not the only one who may have missed it so it's good for posterity!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 507499,
          "author_name": "Supratim Haldar",
          "author_url": "",
          "post_date": "2019-04-04T18:55:12.337000",
          "content": "<p>Hi Addison, the Competition-Specific Rules talks about \"external data\", but does not make any explicit mention about pre-trained models. Or can both be technically considered as same? Apologies if the question sounds very naive!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 507529,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-04-04T20:29:17.137000",
          "content": "<p>Not a naive question - we've wrestled with the same topic on our team as well. We say they're considered as same, as the pre-trained models are themselves trained on external data. In some competitions we've prohibited one, but excluded the other, in which case we provide a whitelist for consideration</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 507540,
          "author_name": "Supratim Haldar",
          "author_url": "",
          "post_date": "2019-04-04T20:50:32.617000",
          "content": "<p>Thank you for the clarification, Addison!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 507582,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2019-04-04T22:38:27.793000",
          "content": "<p>Question about \"pre-trained models\" - I assume that one of the typical pre-trained models that IS NOT TRAINED can be used?</p>\n\n<p>Almost all of these pre-trained can easily be used without using the pre-trained weights and train from scratch with the allowed data.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 507791,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-04-05T07:46:48.253000",
          "content": "<p>Just to be sure - we need to train models from scratch and make inference within 1 hour of GPU time?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 508250,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-04-05T21:28:51.670000",
          "content": "<p>Hi Andrew, \nas explained in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data</a> section:</p>\n\n<p>Submissions must be made with inference models running in Kaggle Kernels. However, participants can decide to train also in the Kaggle Kernels or offline (see <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements\">Kernels Requirements</a> for details).</p>\n\n<p>So, you can train in your local server if you wish.</p>\n\n<p>thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 508252,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-04-05T21:36:32.950000",
          "content": "<p>Jimmy,  I'm not sure I understand the question. A pre-trained model that is not trained? do you refer to using standard openly-available architectures, like VGG? </p>\n\n<p>thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 508641,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2019-04-06T14:31:49.150000",
          "content": "<p>Yes - almost every pre-trained model I have used had the option to NOT use the trained weights and start from scratch.  Can we use these if we don't use the already established weights?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 509002,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-04-07T06:53:41.530000",
          "content": "<p>Hi Eduardo (<a href=\"/eduardofonseca\">@eduardofonseca</a>), excuse me but let me confirm this to make it sure I will not violate rules.\nAs you've clarified as follows, <strong>we can train as much as possible only if</strong>:\n- Training from scratch.\n- Only using dataset (train_curated, train_noisy <strong>and test</strong>) provided.\n- Locally trained model can be used as a kaggle dataset.</p>\n\n<p>&gt; Submissions must be made with inference models running in Kaggle Kernels. However, participants &gt; can decide to train also in the Kaggle Kernels or offline (see Kernels Requirements for details).\n&gt; So, you can train in your local server if you wish.</p>\n\n<p>Thank you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 509358,
          "author_name": "George Oblapenko",
          "author_url": "",
          "post_date": "2019-04-07T17:48:17.510000",
          "content": "<p>Also re-iterating daisukelab's question:\nIf I train a model (say, a neural net) offline on the train data provided, and then upload the weights as a Kaggle dataset, I am allowed to use that in my Kernel?</p>\n\n<p>Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 510194,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-04-08T19:50:09.610000",
          "content": "<p>I also can't get this - so it's allowed to train model offline and just perform the inference in the kernel? What is the point of kernel competition then? Just to be sure a given amount of data can be classified in 4 CPU hours?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 510203,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-04-08T20:03:23.590000",
          "content": "<p>Hi George <a href=\"/kunstmord\">@kunstmord</a> and DavidS <a href=\"/davids1992\">@davids1992</a>, he already have clarified here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#508071\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#508071</a></p>\n\n<p>So the answer is yes, we can train locally as long as using the data provided only.\nWe are not supposed to use any other data, including models pretrained with other dataset like ImageNet.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 510221,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-08T20:38:48.707000",
          "content": "<p>as pointed out above by <a href=\"/daisukelab\">@daisukelab</a> (thanks!), please refer to this official thread for the discussion of the rules\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064</a></p>\n\n<p><a href=\"/davids1992\">@davids1992</a> last year we ran an offline challenge without kernels and we received some submissions with large models and ensembles. 30-40-way ensembles were not uncommon in top leaderboard positions with model sizes going over 55M weights (for comparison, a ResNet-50 is around 30-35M weights). Such large models and ensembles are not interesting from a research point of view since you can always play games with large models and ensembles to extract a few extra points but these don't generalize to other tasks or to real-world uses. Such models could not run in real-time on a mid-range smartphone or low-power Internet-of-things device in your home, for example, and are expensive to run even on a server. We'd like to gently push participants towards figuring out smarter models and better ways to use the data. The limits this year are very generous, and we think 4 hours of CPU time or 1 hour of GPU time is more than enough to run inference on a decent model, while hopefully making it harder to just submit an ensemble of 100 ResNets :)  As a point of comparison, the Challenge Baseline can generate predictions on the entire test set in under 5 GPU-minutes.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 510401,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2019-04-09T03:21:05.083000",
          "content": "<p>From a production point of view I actually agree with this competition's format. Inference time and model size are the things that matter the most. A model may take weeks to train, but as long as you can make it small and predict in real time it's still a wonderful trade off.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 510522,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-04-09T07:11:37.483000",
          "content": "<p>I understand, that's indeed well thought out, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 514867,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-04-12T01:45:07.550000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> , quick comment: to train your system you can use\n- train_curated\n- train_noisy</p>\n\n<p>But <strong>the test set cannot</strong> be used to train the submitted system.  Please refer to this official thread for the discussion of the rules:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 514890,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-04-12T02:27:41.260000",
          "content": "<p><a href=\"/eduardofonseca\">@eduardofonseca</a> Thank you for clarification and link, I appreciate.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 507790,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2019-04-05T07:45:32.380000",
      "content": "<p>̶A̶̶l̶̶t̶̶h̶̶o̶̶u̶̶g̶̶h̶ ̶p̶̶r̶̶o̶̶h̶̶i̶̶b̶̶i̶̶t̶̶e̶̶d̶ ̶f̶̶r̶̶o̶̶m̶ ̶m̶̶y̶ ̶e̶̶x̶̶p̶̶e̶̶r̶̶i̶̶e̶̶n̶̶c̶̶e̶ ̶s̶̶o̶ ̶f̶̶a̶̶r̶ ̶I̶̶m̶̶a̶̶g̶̶e̶̶N̶̶e̶̶t̶ ̶m̶̶o̶̶d̶̶e̶̶l̶̶s̶ ̶h̶̶a̶̶v̶̶e̶ ̶n̶̶o̶ ̶a̶̶d̶̶v̶̶a̶̶n̶̶t̶̶a̶̶g̶̶e̶̶s̶ ̶o̶̶v̶̶e̶̶r̶ ̶t̶̶r̶̶a̶̶i̶̶n̶̶i̶̶n̶̶g̶ ̶f̶̶r̶̶o̶̶m̶ ̶s̶̶c̶̶r̶̶a̶̶t̶̶c̶̶h̶ ̶a̶̶n̶̶y̶̶w̶̶a̶̶y̶̶.̶</p>",
      "votes": 0,
      "replies": [
        {
          "id": 508264,
          "author_name": "Turab Iqbal",
          "author_url": "",
          "post_date": "2019-04-05T22:12:33.733000",
          "content": "<p>Out of curiosity, what architectures have you tried? Using DenseNet-201 with all layers fine-tuned always gave me a slight improvement over training from scratch (not necessarily constrained to DenseNet). Based on preliminary experiments, I've observed the same with this task too.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 508291,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2019-04-05T23:40:25.933000",
          "content": "<p>I may want to take that statement back. I had a bug in my validation that made the score of the pretrained model lower that it really is. Looks like ImageNet models are much stronger for this task for some reason.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 507497,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-04T18:54:41.007000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "507463": "The topic is going to surface pretty soon, so in order to save people some time a few quick references :\n* pretrained PyTorch models (mostly image-based, useful for working with spectrograms, STFT and the like) https://github.com/Cadene/pretrained-models.pytorch\n* pretrained Keras models: https://keras.io/applications/\n* DeepSpeech might come in handy, I guess https://github.com/SeanNaren/deepspeech.pytorch\n* Baidu: https://github.com/baidu-research/warp-ctc\n* speech recognition challenge https://www.kaggle.com/c/tensorflow-speech-recognition-challenge\n\nI'll try to add more if I find other interesting stuff.\n\nEDIT: Competition-specific rules prohibit pretrained models (h/t Addison for drawing my attention to that tiny detail), so use the list at your own risk :-) ",
    "507466": "Hi Konrad,\n\nIt should be noted that per the [rules](https://www.kaggle.com/c/freesound-audio-tagging-2019/rules) of the competition, external data and pretrained models are not allowed.",
    "507790": "̶A̶̶l̶̶t̶̶h̶̶o̶̶u̶̶g̶̶h̶ ̶p̶̶r̶̶o̶̶h̶̶i̶̶b̶̶i̶̶t̶̶e̶̶d̶ ̶f̶̶r̶̶o̶̶m̶ ̶m̶̶y̶ ̶e̶̶x̶̶p̶̶e̶̶r̶̶i̶̶e̶̶n̶̶c̶̶e̶ ̶s̶̶o̶ ̶f̶̶a̶̶r̶ ̶I̶̶m̶̶a̶̶g̶̶e̶̶N̶̶e̶̶t̶ ̶m̶̶o̶̶d̶̶e̶̶l̶̶s̶ ̶h̶̶a̶̶v̶̶e̶ ̶n̶̶o̶ ̶a̶̶d̶̶v̶̶a̶̶n̶̶t̶̶a̶̶g̶̶e̶̶s̶ ̶o̶̶v̶̶e̶̶r̶ ̶t̶̶r̶̶a̶̶i̶̶n̶̶i̶̶n̶̶g̶ ̶f̶̶r̶̶o̶̶m̶ ̶s̶̶c̶̶r̶̶a̶̶t̶̶c̶̶h̶ ̶a̶̶n̶̶y̶̶w̶̶a̶̶y̶̶.̶",
    "507497": ""
  }
}