{
  "id": 238569,
  "title": "Previous competition's models as pretrained models?",
  "url": "/competitions/birdclef-2021/discussion/238569",
  "author_name": "",
  "post_date": "2021-05-12T16:06:50.343702200Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone, <br>\nIs it allowed to use previous competition's solution model as a pretrained model and finetune over it? <br>\nDoes it come under this point which is given in the Code requirements?</p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>\n<p>I tried to find if this topic has been discussed in any thread but could not find anything about it. Incase if there is already a thread where this is discussed and if someone can point it out it will be really helpful….</p>",
  "messages": [
    {
      "id": "1304435",
      "postDate": "05/12/2021 16:06:50",
      "content": "<p>Hi everyone, <br>\nIs it allowed to use previous competition's solution model as a pretrained model and finetune over it? <br>\nDoes it come under this point which is given in the Code requirements?</p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>\n<p>I tried to find if this topic has been discussed in any thread but could not find anything about it. Incase if there is already a thread where this is discussed and if someone can point it out it will be really helpful….</p>",
      "rawMarkdown": "Hi everyone, \nIs it allowed to use previous competition's solution model as a pretrained model and finetune over it? \nDoes it come under this point which is given in the Code requirements?\n>  Freely & publicly available external data is allowed, including pre-trained models\n\nI tried to find if this topic has been discussed in any thread but could not find anything about it. Incase if there is already a thread where this is discussed and if someone can point it out it will be really helpful....",
      "votes": null
    },
    {
      "id": "1306131",
      "postDate": "05/13/2021 16:13:53",
      "content": "<p>Hi, <br>\nyeah, it is allowed to use those models for training and tunning. I think this is not against the rule as any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.</p>",
      "rawMarkdown": "Hi, \nyeah, it is allowed to use those models for training and tunning. I think this is not against the rule as any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.",
      "votes": null
    },
    {
      "id": "1306319",
      "postDate": "05/13/2021 17:43:56",
      "content": "<blockquote>\n  <p>any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.</p>\n</blockquote>\n<p>Beware, if the dataset source has a license that limits its use then it carries forward.</p>",
      "rawMarkdown": "> any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.\n\nBeware, if the dataset source has a license that limits its use then it carries forward.",
      "votes": null
    },
    {
      "id": "1306327",
      "postDate": "05/13/2021 17:47:28",
      "content": "<p>Yeah, very true. Different license has different limitations so it's better to check once and then use them. But hopefully the publicly available models have that sort of license which doesn't restrict it's use in different kaggle competition.</p>",
      "rawMarkdown": "Yeah, very true. Different license has different limitations so it's better to check once and then use them. But hopefully the publicly available models have that sort of license which doesn't restrict it's use in different kaggle competition.",
      "votes": null
    },
    {
      "id": "1306369",
      "postDate": "05/13/2021 18:09:46",
      "content": "<p>Public model weigths are fine in general.  It is more datasets that may come with restrictions.</p>\n<p>If a model is finetuned on a dataset that has restrictions, then the restrictions carry forward to the fine tuned model.</p>",
      "rawMarkdown": "Public model weigths are fine in general.  It is more datasets that may come with restrictions.\n\nIf a model is finetuned on a dataset that has restrictions, then the restrictions carry forward to the fine tuned model.",
      "votes": null
    },
    {
      "id": "1306897",
      "postDate": "05/14/2021 06:13:23",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nwould u be able to throw some light on choosing <br>\nMelespecs param</p>\n<pre><code>SR = 32_000\nkwargs[\"n_fft\"] = kwargs.get(\"n_fft\", self.sr//10)\n kwargs[\"hop_length\"] = kwargs.get(\"hop_length\", self.sr//(10*4))\n\n  self.fmin = 0\n self.fmax = Nne or self.sr//2\n</code></pre>\n<p>DURATION = 5 # this is as per  test data clips duration<br>\n<code>self, data, sr=SR, n_mels=128, fmin=0, fmax=None, duration=DURATION, step=None, res_type=\"kaiser_fast\", resample=True</code></p>\n<p>MelSpecComputer(sr=self.sr, n_mels=self.n_mels, fmin=self.fmin,<br>\n                                                 fmax=self.fmax)</p>\n<p>Previous competition eg.</p>\n<pre><code>self.n_fft = 892\nself.sr = 21952 \nself.hop_length=245\nelf.n_mels =  224\n self.win_length = self.n_fft\n</code></pre>\n<p><code>spectr = librosa.feature.melspectrogram(y, sr=hp.sr, n_mels=hp.n_mels, n_fft=hp.n_fft, hop_length = hp.hop_length, win_length = hp.win_length, fmin = 300)</code></p>\n<p>I find theses to be very different from ones used in previous version.</p>\n<p>Are these decided based on test data or so .</p>",
      "rawMarkdown": "cpmpml \nwould u be able to throw some light on choosing \nMelespecs param\n```\nSR = 32_000\nkwargs[\"n_fft\"] = kwargs.get(\"n_fft\", self.sr//10)\n kwargs[\"hop_length\"] = kwargs.get(\"hop_length\", self.sr//(10*4))\n\n  self.fmin = 0\n self.fmax = Nne or self.sr//2\n```\nDURATION = 5 # this is as per  test data clips duration\n`self, data, sr=SR, n_mels=128, fmin=0, fmax=None, duration=DURATION, step=None, res_type=\"kaiser_fast\", resample=True`\n\n MelSpecComputer(sr=self.sr, n_mels=self.n_mels, fmin=self.fmin,\n                                                 fmax=self.fmax)\n\nPrevious competition eg.\n```\nself.n_fft = 892\nself.sr = 21952 \nself.hop_length=245\nelf.n_mels =  224\n self.win_length = self.n_fft\n```\n `spectr = librosa.feature.melspectrogram(y, sr=hp.sr, n_mels=hp.n_mels, n_fft=hp.n_fft, hop_length = hp.hop_length, win_length = hp.win_length, fmin = 300)`\n\nI find theses to be very different from ones used in previous version.\n\nAre these decided based on test data or so .",
      "votes": null
    },
    {
      "id": "1307022",
      "postDate": "05/14/2021 07:53:04",
      "content": "<p>I personally think there are two options that are possible: </p>\n<ul>\n<li>publicly available pretrained models</li>\n<li>you can pretrain new models on old data if the data is available</li>\n</ul>",
      "rawMarkdown": "I personally think there are two options that are possible: \n- publicly available pretrained models\n- you can pretrain new models on old data if the data is available",
      "votes": null
    },
    {
      "id": "1307174",
      "postDate": "05/14/2021 09:48:53",
      "content": "<p>I will not disclose what I am doing before end. The only thing I can tell you is to try and see what works best.</p>",
      "rawMarkdown": "I will not disclose what I am doing before end. The only thing I can tell you is to try and see what works best.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1306131,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "05/13/2021 16:13:53",
      "content": "<p>Hi, <br>\nyeah, it is allowed to use those models for training and tunning. I think this is not against the rule as any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1306319,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/13/2021 17:43:56",
          "content": "<blockquote>\n  <p>any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.</p>\n</blockquote>\n<p>Beware, if the dataset source has a license that limits its use then it carries forward.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1306327,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "05/13/2021 17:47:28",
          "content": "<p>Yeah, very true. Different license has different limitations so it's better to check once and then use them. But hopefully the publicly available models have that sort of license which doesn't restrict it's use in different kaggle competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1306369,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/13/2021 18:09:46",
          "content": "<p>Public model weigths are fine in general.  It is more datasets that may come with restrictions.</p>\n<p>If a model is finetuned on a dataset that has restrictions, then the restrictions carry forward to the fine tuned model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1306897,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "05/14/2021 06:13:23",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nwould u be able to throw some light on choosing <br>\nMelespecs param</p>\n<pre><code>SR = 32_000\nkwargs[\"n_fft\"] = kwargs.get(\"n_fft\", self.sr//10)\n kwargs[\"hop_length\"] = kwargs.get(\"hop_length\", self.sr//(10*4))\n\n  self.fmin = 0\n self.fmax = Nne or self.sr//2\n</code></pre>\n<p>DURATION = 5 # this is as per  test data clips duration<br>\n<code>self, data, sr=SR, n_mels=128, fmin=0, fmax=None, duration=DURATION, step=None, res_type=\"kaiser_fast\", resample=True</code></p>\n<p>MelSpecComputer(sr=self.sr, n_mels=self.n_mels, fmin=self.fmin,<br>\n                                                 fmax=self.fmax)</p>\n<p>Previous competition eg.</p>\n<pre><code>self.n_fft = 892\nself.sr = 21952 \nself.hop_length=245\nelf.n_mels =  224\n self.win_length = self.n_fft\n</code></pre>\n<p><code>spectr = librosa.feature.melspectrogram(y, sr=hp.sr, n_mels=hp.n_mels, n_fft=hp.n_fft, hop_length = hp.hop_length, win_length = hp.win_length, fmin = 300)</code></p>\n<p>I find theses to be very different from ones used in previous version.</p>\n<p>Are these decided based on test data or so .</p>",
      "votes": null,
      "replies": [
        {
          "id": 1307174,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/14/2021 09:48:53",
          "content": "<p>I will not disclose what I am doing before end. The only thing I can tell you is to try and see what works best.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1307022,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "05/14/2021 07:53:04",
      "content": "<p>I personally think there are two options that are possible: </p>\n<ul>\n<li>publicly available pretrained models</li>\n<li>you can pretrain new models on old data if the data is available</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1304435": "Hi everyone, \nIs it allowed to use previous competition's solution model as a pretrained model and finetune over it? \nDoes it come under this point which is given in the Code requirements?\n>  Freely & publicly available external data is allowed, including pre-trained models\n\nI tried to find if this topic has been discussed in any thread but could not find anything about it. Incase if there is already a thread where this is discussed and if someone can point it out it will be really helpful....",
    "1306131": "Hi, \nyeah, it is allowed to use those models for training and tunning. I think this is not against the rule as any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.",
    "1306319": "> any dataset on Kaggle which is made public has a licence under which anyone can use it for competitions or such purposes.\n\nBeware, if the dataset source has a license that limits its use then it carries forward.",
    "1306327": "Yeah, very true. Different license has different limitations so it's better to check once and then use them. But hopefully the publicly available models have that sort of license which doesn't restrict it's use in different kaggle competition.",
    "1306369": "Public model weigths are fine in general.  It is more datasets that may come with restrictions.\n\nIf a model is finetuned on a dataset that has restrictions, then the restrictions carry forward to the fine tuned model.",
    "1306897": "cpmpml \nwould u be able to throw some light on choosing \nMelespecs param\n```\nSR = 32_000\nkwargs[\"n_fft\"] = kwargs.get(\"n_fft\", self.sr//10)\n kwargs[\"hop_length\"] = kwargs.get(\"hop_length\", self.sr//(10*4))\n\n  self.fmin = 0\n self.fmax = Nne or self.sr//2\n```\nDURATION = 5 # this is as per  test data clips duration\n`self, data, sr=SR, n_mels=128, fmin=0, fmax=None, duration=DURATION, step=None, res_type=\"kaiser_fast\", resample=True`\n\n MelSpecComputer(sr=self.sr, n_mels=self.n_mels, fmin=self.fmin,\n                                                 fmax=self.fmax)\n\nPrevious competition eg.\n```\nself.n_fft = 892\nself.sr = 21952 \nself.hop_length=245\nelf.n_mels =  224\n self.win_length = self.n_fft\n```\n `spectr = librosa.feature.melspectrogram(y, sr=hp.sr, n_mels=hp.n_mels, n_fft=hp.n_fft, hop_length = hp.hop_length, win_length = hp.win_length, fmin = 300)`\n\nI find theses to be very different from ones used in previous version.\n\nAre these decided based on test data or so .",
    "1307022": "I personally think there are two options that are possible: \n- publicly available pretrained models\n- you can pretrain new models on old data if the data is available",
    "1307174": "I will not disclose what I am doing before end. The only thing I can tell you is to try and see what works best."
  },
  "source": "meta"
}