{
  "id": 313394,
  "title": "External data",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/313394",
  "author_name": "",
  "post_date": "2022-03-16T21:36:10.343052100Z",
  "votes": 3,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>a quick question regarding the external data rule just to be sure: does creating an embedding of the training set (and save it in a kaggle dataset for later use) count as using external data thereby disqualifying the user?</p>",
  "messages": [
    {
      "id": "1725166",
      "postDate": "03/16/2022 21:36:10",
      "content": "<p>Hi there,</p>\n<p>a quick question regarding the external data rule just to be sure: does creating an embedding of the training set (and save it in a kaggle dataset for later use) count as using external data thereby disqualifying the user?</p>",
      "rawMarkdown": "Hi there,\n\na quick question regarding the external data rule just to be sure: does creating an embedding of the training set (and save it in a kaggle dataset for later use) count as using external data thereby disqualifying the user?",
      "votes": null
    },
    {
      "id": "1725367",
      "postDate": "03/17/2022 03:21:20",
      "content": "<p>Hi Jacopo! </p>\n<p>I believe external data is not allowed-From the rules tab:</p>\n<blockquote>\n  <p>No external data. Only use the provided data to train your model.</p>\n</blockquote>",
      "rawMarkdown": "Hi Jacopo! \n\nI believe external data is not allowed-From the rules tab:\n\n> No external data. Only use the provided data to train your model.",
      "votes": null
    },
    {
      "id": "1725626",
      "postDate": "03/17/2022 09:10:06",
      "content": "<p>HI <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> and congrats for your GM title!</p>\n<p>Thanks for the clarification, however <a href=\"https://www.kaggle.com/c/kaggle-pog-series-s01e02/discussion/312554\" target=\"_blank\">here</a> it's also stated that </p>\n<blockquote>\n  <p>Training can be done offline and uploaded as a dataset</p>\n</blockquote>\n<p>that's why I asked as it's not fully clear to me. </p>",
      "rawMarkdown": "HI @init27 and congrats for your GM title!\n\nThanks for the clarification, however [here](https://www.kaggle.com/c/kaggle-pog-series-s01e02/discussion/312554) it's also stated that \n> Training can be done offline and uploaded as a dataset\n\nthat's why I asked as it's not fully clear to me.",
      "votes": null
    },
    {
      "id": "1725742",
      "postDate": "03/17/2022 11:55:41",
      "content": "<p>Thank you so much! 🙏</p>\n<p>This refers to a common process on Kaggle-inference only competition. </p>\n<p>You are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel. </p>\n<p>Please note: The other rules such as no usage of pretrained models applies-which means the model you train offline should be trained from scratch. </p>\n<p>I hope this helps! </p>",
      "rawMarkdown": "Thank you so much! 🙏\n\nThis refers to a common process on Kaggle-inference only competition. \n\nYou are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel. \n\nPlease note: The other rules such as no usage of pretrained models applies-which means the model you train offline should be trained from scratch. \n\nI hope this helps!",
      "votes": null
    },
    {
      "id": "1725818",
      "postDate": "03/17/2022 13:11:20",
      "content": "<p>So, going back to my question, am I allowed to do the following?</p>\n<ul>\n<li>extract features from audio files</li>\n<li>save those features in a dataset</li>\n<li>train a model from scratch based on that \"external\" dataset</li>\n<li>save model to another (or the same) external dataset</li>\n<li>create an inference notebook leveraging my pre-trained model and submit predictions working only with the test set</li>\n</ul>\n<p>I'll tag also <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for a final feedback, thank you!</p>",
      "rawMarkdown": "So, going back to my question, am I allowed to do the following?\n- extract features from audio files\n- save those features in a dataset\n- train a model from scratch based on that \"external\" dataset\n- save model to another (or the same) external dataset\n- create an inference notebook leveraging my pre-trained model and submit predictions working only with the test set\n\nI'll tag also @robikscube for a final feedback, thank you!",
      "votes": null
    },
    {
      "id": "1725855",
      "postDate": "03/17/2022 13:36:27",
      "content": "<p>Hi there. Sorry that there is confusion on this. Hopefully I can clear things up.</p>\n<p>The solution must not use any external data. That could be:</p>\n<ul>\n<li>Raw external data like labeled audio from an external source.</li>\n<li>Pretrained models that were trained on external data.</li>\n</ul>\n<p>So to answer your question I need to know what you mean by \"extract features from audio files\". If you are creating these features/embeddings without using a pretrained model to extract those features then you should be ok. Keep in mind I may ask for the code you used to create those features and confirm no pretrained models were used.</p>\n<p>Additionally, to qualify for the prize the submission needs to be done entirely in a kaggle notebook. So that means the test features need to be created in that notebook.</p>\n<p>Hope that helps. Let me know if I can clarify.</p>",
      "rawMarkdown": "Hi there. Sorry that there is confusion on this. Hopefully I can clear things up.\n\nThe solution must not use any external data. That could be:\n- Raw external data like labeled audio from an external source.\n- Pretrained models that were trained on external data.\n\nSo to answer your question I need to know what you mean by \"extract features from audio files\". If you are creating these features/embeddings without using a pretrained model to extract those features then you should be ok. Keep in mind I may ask for the code you used to create those features and confirm no pretrained models were used.\n\nAdditionally, to qualify for the prize the submission needs to be done entirely in a kaggle notebook. So that means the test features need to be created in that notebook.\n\nHope that helps. Let me know if I can clarify.",
      "votes": null
    },
    {
      "id": "1726090",
      "postDate": "03/17/2022 16:36:22",
      "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> <br>\nYou are right, my \"extract features from audio files\" is very generic, let me clarify: I'd like to extract features like loudness, spectral energy, bpm, danceability etc using a python library (Essentia). <br>\nThese features will be saved in a dataset then used for model creation.</p>",
      "rawMarkdown": "Thanks for the reply @robikscube \nYou are right, my \"extract features from audio files\" is very generic, let me clarify: I'd like to extract features like loudness, spectral energy, bpm, danceability etc using a python library (Essentia). \nThese features will be saved in a dataset then used for model creation.",
      "votes": null
    },
    {
      "id": "1726135",
      "postDate": "03/17/2022 17:14:09",
      "content": "<p>I don't know about the Essentia library specifically. But it does look like it has some features that use pretrained models. Using these to extract features would be against the rules. There may be other parts of essentia that don't use pretrained models but I'd only use them if you are 100% sure it doesn't use a pretrained model.</p>\n<p><a href=\"https://essentia.upf.edu/streaming_extractor_music.html#high-level-classifier-models\" target=\"_blank\">https://essentia.upf.edu/streaming_extractor_music.html#high-level-classifier-models</a></p>\n<p>From that page they say:</p>\n<p><em>We provide such models pretrained on our ground truth music collections for each version of the music extractor via a <a href=\"https://essentia.upf.edu/svm_models/\" target=\"_blank\">download page</a>.</em></p>\n<p>Using those models are not allowed.</p>",
      "rawMarkdown": "I don't know about the Essentia library specifically. But it does look like it has some features that use pretrained models. Using these to extract features would be against the rules. There may be other parts of essentia that don't use pretrained models but I'd only use them if you are 100% sure it doesn't use a pretrained model.\n\nhttps://essentia.upf.edu/streaming_extractor_music.html#high-level-classifier-models\n\nFrom that page they say:\n\n*We provide such models pretrained on our ground truth music collections for each version of the music extractor via a [download page](https://essentia.upf.edu/svm_models/).*\n\nUsing those models are not allowed.",
      "votes": null
    },
    {
      "id": "1726146",
      "postDate": "03/17/2022 17:24:26",
      "content": "<p>Thanks for pointing it out!<br>\nIt seems quite risky using that library, I'll try to see if there are parts that don't use pretrained models.</p>\n<p>Anyways, thanks for your time!</p>",
      "rawMarkdown": "Thanks for pointing it out!\nIt seems quite risky using that library, I'll try to see if there are parts that don't use pretrained models.\n\nAnyways, thanks for your time!",
      "votes": null
    },
    {
      "id": "1726216",
      "postDate": "03/17/2022 19:10:07",
      "content": "<p>That is a good point to clarify. Will it be the same with Librosa function for tempo? <a href=\"https://librosa.org/doc/latest/beat.html\" target=\"_blank\">https://librosa.org/doc/latest/beat.html</a>.  How about functions on this page? I made a dataset with those functions, but it might be invalid for this competition.</p>",
      "rawMarkdown": "That is a good point to clarify. Will it be the same with Librosa function for tempo? https://librosa.org/doc/latest/beat.html.  How about functions on this page? I made a dataset with those functions, but it might be invalid for this competition.",
      "votes": null
    },
    {
      "id": "1727337",
      "postDate": "03/17/2022 21:45:20",
      "content": "<p>ya that's a good point as well<br>\nI guess that tempo is more or less like calculating MFCC: yes, there is an algorithm but there's no training on external dataset involved.<br>\nAgain, I think only <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> can clarify without a shadow of a doubt!</p>",
      "rawMarkdown": "ya that's a good point as well\nI guess that tempo is more or less like calculating MFCC: yes, there is an algorithm but there's no training on external dataset involved.\nAgain, I think only @robikscube can clarify without a shadow of a doubt!",
      "votes": null
    },
    {
      "id": "1727345",
      "postDate": "03/17/2022 21:56:53",
      "content": "<p>The librosa beat function looks fine to me. It doesn't involve a pretrained model, as far as I can tell it's just an algorithm for detecting when the beats occur. Hope that helps.</p>",
      "rawMarkdown": "The librosa beat function looks fine to me. It doesn't involve a pretrained model, as far as I can tell it's just an algorithm for detecting when the beats occur. Hope that helps.",
      "votes": null
    },
    {
      "id": "1727376",
      "postDate": "03/17/2022 23:27:54",
      "content": "<p>Sorry again <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> , just another clarification. I'm sorry if this is becoming boring but better safe than sorry.</p>\n<p>Let's take this function as an example, the <a href=\"https://essentia.upf.edu/tutorial_pitch_melody.html\" target=\"_blank\">melody detection</a>. According to the docs, it uses an algo called PredominantPitchMelodia, which <a href=\"https://essentia.upf.edu/reference/std_PredominantPitchMelodia.html\" target=\"_blank\">in turn</a> estimates the fundamental frequency of the predominant melody using the MELODIA algorithm.</p>\n<p>Quite a mouthful, nevertheless it looks like there are a bunch of methods based on spectrograms, pitch calculations and other signal processing tricks.</p>\n<p>My questions then is how can we correctly define when we are facing a pretrained algo or not? For instance MFCC (and maybe the PitchMelodia algo above) is not a pretrained model but the formula for obtaining the coefficients may vaguely recall the concept of exploiting a pre-trained model, no?</p>\n<p>Just another bordeline example, what about this <a href=\"https://essentia.upf.edu/reference/streaming_Dissonance.html\" target=\"_blank\">algorithm here</a>, able to detect sensory dissonance, where the documentations says <code>the dissonance curves are based on perceptual experiments conducted in [paper cited]</code>?</p>\n<p>Thanks again for your time, I'm sure that other kagglers may find all this useful as well.</p>",
      "rawMarkdown": "Sorry again @robikscube , just another clarification. I'm sorry if this is becoming boring but better safe than sorry.\n\nLet's take this function as an example, the [melody detection](https://essentia.upf.edu/tutorial_pitch_melody.html). According to the docs, it uses an algo called PredominantPitchMelodia, which [in turn](https://essentia.upf.edu/reference/std_PredominantPitchMelodia.html) estimates the fundamental frequency of the predominant melody using the MELODIA algorithm.\n\nQuite a mouthful, nevertheless it looks like there are a bunch of methods based on spectrograms, pitch calculations and other signal processing tricks.\n\nMy questions then is how can we correctly define when we are facing a pretrained algo or not? For instance MFCC (and maybe the PitchMelodia algo above) is not a pretrained model but the formula for obtaining the coefficients may vaguely recall the concept of exploiting a pre-trained model, no?\n\nJust another bordeline example, what about this [algorithm here](https://essentia.upf.edu/reference/streaming_Dissonance.html), able to detect sensory dissonance, where the documentations says `the dissonance curves are based on perceptual experiments conducted in [paper cited]`?\n\nThanks again for your time, I'm sure that other kagglers may find all this useful as well.",
      "votes": null
    },
    {
      "id": "1727424",
      "postDate": "03/18/2022 01:02:11",
      "content": "<p>This is a great question and there certainly is some grey area here. I'm discussing this currently on my twitch stream and the short answer is: was the model trained on a corpus of data? In the cases you list above I believe the answer is no so they should be ok. Downloading a model that was trained on 100k audio files would not be ok. If you have more specific examples feel free to ask.</p>",
      "rawMarkdown": "This is a great question and there certainly is some grey area here. I'm discussing this currently on my twitch stream and the short answer is: was the model trained on a corpus of data? In the cases you list above I believe the answer is no so they should be ok. Downloading a model that was trained on 100k audio files would not be ok. If you have more specific examples feel free to ask.",
      "votes": null
    },
    {
      "id": "1728117",
      "postDate": "03/18/2022 15:42:17",
      "content": "<p>Sorry, I forgot to put the second page. I assume those functions are good to use for feature extractions. </p>\n<p><a href=\"https://librosa.org/doc/main/feature.html\" target=\"_blank\">https://librosa.org/doc/main/feature.html</a></p>",
      "rawMarkdown": "Sorry, I forgot to put the second page. I assume those functions are good to use for feature extractions. \n\nhttps://librosa.org/doc/main/feature.html",
      "votes": null
    },
    {
      "id": "1728392",
      "postDate": "03/18/2022 20:59:43",
      "content": "<pre><code>You are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel.\n</code></pre>\n<p>Does it mean I can train my model entirely offline and upload the weight just for inference on the test set? The code for training model will be public then <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> . Thanks</p>",
      "rawMarkdown": "```\nYou are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel.\n```\nDoes it mean I can train my model entirely offline and upload the weight just for inference on the test set? The code for training model will be public then @robikscube . Thanks",
      "votes": null
    },
    {
      "id": "1728410",
      "postDate": "03/18/2022 21:42:15",
      "content": "<p>Yes, librosa is fine to use.</p>",
      "rawMarkdown": "Yes, librosa is fine to use.",
      "votes": null
    },
    {
      "id": "1728412",
      "postDate": "03/18/2022 21:42:46",
      "content": "<p>Yes, you are allowed to train your model offline (only on the training data) and then upload to a kaggle dataset for inference.</p>",
      "rawMarkdown": "Yes, you are allowed to train your model offline (only on the training data) and then upload to a kaggle dataset for inference.",
      "votes": null
    },
    {
      "id": "1728512",
      "postDate": "03/19/2022 01:20:58",
      "content": "<p><a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> This is quite common on Kaggle, please checkout the BirdCLEF competition which IIRC follows a similar setup. </p>\n<p>I was actually really happy to see that GM Rob has setup a similar community competition for us, similar to a \"real\" (points) one.</p>",
      "rawMarkdown": "dienhoa This is quite common on Kaggle, please checkout the BirdCLEF competition which IIRC follows a similar setup. \n\nI was actually really happy to see that GM Rob has setup a similar community competition for us, similar to a \"real\" (points) one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1725367,
      "author_name": "init27",
      "author_url": "",
      "post_date": "03/17/2022 03:21:20",
      "content": "<p>Hi Jacopo! </p>\n<p>I believe external data is not allowed-From the rules tab:</p>\n<blockquote>\n  <p>No external data. Only use the provided data to train your model.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1725626,
          "author_name": "jacoporepossi",
          "author_url": "",
          "post_date": "03/17/2022 09:10:06",
          "content": "<p>HI <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> and congrats for your GM title!</p>\n<p>Thanks for the clarification, however <a href=\"https://www.kaggle.com/c/kaggle-pog-series-s01e02/discussion/312554\" target=\"_blank\">here</a> it's also stated that </p>\n<blockquote>\n  <p>Training can be done offline and uploaded as a dataset</p>\n</blockquote>\n<p>that's why I asked as it's not fully clear to me. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725742,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/17/2022 11:55:41",
          "content": "<p>Thank you so much! 🙏</p>\n<p>This refers to a common process on Kaggle-inference only competition. </p>\n<p>You are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel. </p>\n<p>Please note: The other rules such as no usage of pretrained models applies-which means the model you train offline should be trained from scratch. </p>\n<p>I hope this helps! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725818,
          "author_name": "jacoporepossi",
          "author_url": "",
          "post_date": "03/17/2022 13:11:20",
          "content": "<p>So, going back to my question, am I allowed to do the following?</p>\n<ul>\n<li>extract features from audio files</li>\n<li>save those features in a dataset</li>\n<li>train a model from scratch based on that \"external\" dataset</li>\n<li>save model to another (or the same) external dataset</li>\n<li>create an inference notebook leveraging my pre-trained model and submit predictions working only with the test set</li>\n</ul>\n<p>I'll tag also <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> for a final feedback, thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725855,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/17/2022 13:36:27",
          "content": "<p>Hi there. Sorry that there is confusion on this. Hopefully I can clear things up.</p>\n<p>The solution must not use any external data. That could be:</p>\n<ul>\n<li>Raw external data like labeled audio from an external source.</li>\n<li>Pretrained models that were trained on external data.</li>\n</ul>\n<p>So to answer your question I need to know what you mean by \"extract features from audio files\". If you are creating these features/embeddings without using a pretrained model to extract those features then you should be ok. Keep in mind I may ask for the code you used to create those features and confirm no pretrained models were used.</p>\n<p>Additionally, to qualify for the prize the submission needs to be done entirely in a kaggle notebook. So that means the test features need to be created in that notebook.</p>\n<p>Hope that helps. Let me know if I can clarify.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1726090,
          "author_name": "jacoporepossi",
          "author_url": "",
          "post_date": "03/17/2022 16:36:22",
          "content": "<p>Thanks for the reply <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> <br>\nYou are right, my \"extract features from audio files\" is very generic, let me clarify: I'd like to extract features like loudness, spectral energy, bpm, danceability etc using a python library (Essentia). <br>\nThese features will be saved in a dataset then used for model creation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1726135,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/17/2022 17:14:09",
          "content": "<p>I don't know about the Essentia library specifically. But it does look like it has some features that use pretrained models. Using these to extract features would be against the rules. There may be other parts of essentia that don't use pretrained models but I'd only use them if you are 100% sure it doesn't use a pretrained model.</p>\n<p><a href=\"https://essentia.upf.edu/streaming_extractor_music.html#high-level-classifier-models\" target=\"_blank\">https://essentia.upf.edu/streaming_extractor_music.html#high-level-classifier-models</a></p>\n<p>From that page they say:</p>\n<p><em>We provide such models pretrained on our ground truth music collections for each version of the music extractor via a <a href=\"https://essentia.upf.edu/svm_models/\" target=\"_blank\">download page</a>.</em></p>\n<p>Using those models are not allowed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1726146,
          "author_name": "jacoporepossi",
          "author_url": "",
          "post_date": "03/17/2022 17:24:26",
          "content": "<p>Thanks for pointing it out!<br>\nIt seems quite risky using that library, I'll try to see if there are parts that don't use pretrained models.</p>\n<p>Anyways, thanks for your time!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1727376,
          "author_name": "jacoporepossi",
          "author_url": "",
          "post_date": "03/17/2022 23:27:54",
          "content": "<p>Sorry again <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> , just another clarification. I'm sorry if this is becoming boring but better safe than sorry.</p>\n<p>Let's take this function as an example, the <a href=\"https://essentia.upf.edu/tutorial_pitch_melody.html\" target=\"_blank\">melody detection</a>. According to the docs, it uses an algo called PredominantPitchMelodia, which <a href=\"https://essentia.upf.edu/reference/std_PredominantPitchMelodia.html\" target=\"_blank\">in turn</a> estimates the fundamental frequency of the predominant melody using the MELODIA algorithm.</p>\n<p>Quite a mouthful, nevertheless it looks like there are a bunch of methods based on spectrograms, pitch calculations and other signal processing tricks.</p>\n<p>My questions then is how can we correctly define when we are facing a pretrained algo or not? For instance MFCC (and maybe the PitchMelodia algo above) is not a pretrained model but the formula for obtaining the coefficients may vaguely recall the concept of exploiting a pre-trained model, no?</p>\n<p>Just another bordeline example, what about this <a href=\"https://essentia.upf.edu/reference/streaming_Dissonance.html\" target=\"_blank\">algorithm here</a>, able to detect sensory dissonance, where the documentations says <code>the dissonance curves are based on perceptual experiments conducted in [paper cited]</code>?</p>\n<p>Thanks again for your time, I'm sure that other kagglers may find all this useful as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1727424,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/18/2022 01:02:11",
          "content": "<p>This is a great question and there certainly is some grey area here. I'm discussing this currently on my twitch stream and the short answer is: was the model trained on a corpus of data? In the cases you list above I believe the answer is no so they should be ok. Downloading a model that was trained on 100k audio files would not be ok. If you have more specific examples feel free to ask.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728392,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/18/2022 20:59:43",
          "content": "<pre><code>You are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel.\n</code></pre>\n<p>Does it mean I can train my model entirely offline and upload the weight just for inference on the test set? The code for training model will be public then <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> . Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728412,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/18/2022 21:42:46",
          "content": "<p>Yes, you are allowed to train your model offline (only on the training data) and then upload to a kaggle dataset for inference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728512,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/19/2022 01:20:58",
          "content": "<p><a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> This is quite common on Kaggle, please checkout the BirdCLEF competition which IIRC follows a similar setup. </p>\n<p>I was actually really happy to see that GM Rob has setup a similar community competition for us, similar to a \"real\" (points) one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1726216,
      "author_name": "satoshiss",
      "author_url": "",
      "post_date": "03/17/2022 19:10:07",
      "content": "<p>That is a good point to clarify. Will it be the same with Librosa function for tempo? <a href=\"https://librosa.org/doc/latest/beat.html\" target=\"_blank\">https://librosa.org/doc/latest/beat.html</a>.  How about functions on this page? I made a dataset with those functions, but it might be invalid for this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1727337,
          "author_name": "jacoporepossi",
          "author_url": "",
          "post_date": "03/17/2022 21:45:20",
          "content": "<p>ya that's a good point as well<br>\nI guess that tempo is more or less like calculating MFCC: yes, there is an algorithm but there's no training on external dataset involved.<br>\nAgain, I think only <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> can clarify without a shadow of a doubt!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1727345,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/17/2022 21:56:53",
          "content": "<p>The librosa beat function looks fine to me. It doesn't involve a pretrained model, as far as I can tell it's just an algorithm for detecting when the beats occur. Hope that helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728117,
          "author_name": "satoshiss",
          "author_url": "",
          "post_date": "03/18/2022 15:42:17",
          "content": "<p>Sorry, I forgot to put the second page. I assume those functions are good to use for feature extractions. </p>\n<p><a href=\"https://librosa.org/doc/main/feature.html\" target=\"_blank\">https://librosa.org/doc/main/feature.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1728410,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/18/2022 21:42:15",
          "content": "<p>Yes, librosa is fine to use.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1725166": "Hi there,\n\na quick question regarding the external data rule just to be sure: does creating an embedding of the training set (and save it in a kaggle dataset for later use) count as using external data thereby disqualifying the user?",
    "1725367": "Hi Jacopo! \n\nI believe external data is not allowed-From the rules tab:\n\n> No external data. Only use the provided data to train your model.",
    "1725626": "HI @init27 and congrats for your GM title!\n\nThanks for the clarification, however [here](https://www.kaggle.com/c/kaggle-pog-series-s01e02/discussion/312554) it's also stated that \n> Training can be done offline and uploaded as a dataset\n\nthat's why I asked as it's not fully clear to me.",
    "1725742": "Thank you so much! 🙏\n\nThis refers to a common process on Kaggle-inference only competition. \n\nYou are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel. \n\nPlease note: The other rules such as no usage of pretrained models applies-which means the model you train offline should be trained from scratch. \n\nI hope this helps!",
    "1725818": "So, going back to my question, am I allowed to do the following?\n- extract features from audio files\n- save those features in a dataset\n- train a model from scratch based on that \"external\" dataset\n- save model to another (or the same) external dataset\n- create an inference notebook leveraging my pre-trained model and submit predictions working only with the test set\n\nI'll tag also @robikscube for a final feedback, thank you!",
    "1725855": "Hi there. Sorry that there is confusion on this. Hopefully I can clear things up.\n\nThe solution must not use any external data. That could be:\n- Raw external data like labeled audio from an external source.\n- Pretrained models that were trained on external data.\n\nSo to answer your question I need to know what you mean by \"extract features from audio files\". If you are creating these features/embeddings without using a pretrained model to extract those features then you should be ok. Keep in mind I may ask for the code you used to create those features and confirm no pretrained models were used.\n\nAdditionally, to qualify for the prize the submission needs to be done entirely in a kaggle notebook. So that means the test features need to be created in that notebook.\n\nHope that helps. Let me know if I can clarify.",
    "1726090": "Thanks for the reply @robikscube \nYou are right, my \"extract features from audio files\" is very generic, let me clarify: I'd like to extract features like loudness, spectral energy, bpm, danceability etc using a python library (Essentia). \nThese features will be saved in a dataset then used for model creation.",
    "1726135": "I don't know about the Essentia library specifically. But it does look like it has some features that use pretrained models. Using these to extract features would be against the rules. There may be other parts of essentia that don't use pretrained models but I'd only use them if you are 100% sure it doesn't use a pretrained model.\n\nhttps://essentia.upf.edu/streaming_extractor_music.html#high-level-classifier-models\n\nFrom that page they say:\n\n*We provide such models pretrained on our ground truth music collections for each version of the music extractor via a [download page](https://essentia.upf.edu/svm_models/).*\n\nUsing those models are not allowed.",
    "1726146": "Thanks for pointing it out!\nIt seems quite risky using that library, I'll try to see if there are parts that don't use pretrained models.\n\nAnyways, thanks for your time!",
    "1726216": "That is a good point to clarify. Will it be the same with Librosa function for tempo? https://librosa.org/doc/latest/beat.html.  How about functions on this page? I made a dataset with those functions, but it might be invalid for this competition.",
    "1727337": "ya that's a good point as well\nI guess that tempo is more or less like calculating MFCC: yes, there is an algorithm but there's no training on external dataset involved.\nAgain, I think only @robikscube can clarify without a shadow of a doubt!",
    "1727345": "The librosa beat function looks fine to me. It doesn't involve a pretrained model, as far as I can tell it's just an algorithm for detecting when the beats occur. Hope that helps.",
    "1727376": "Sorry again @robikscube , just another clarification. I'm sorry if this is becoming boring but better safe than sorry.\n\nLet's take this function as an example, the [melody detection](https://essentia.upf.edu/tutorial_pitch_melody.html). According to the docs, it uses an algo called PredominantPitchMelodia, which [in turn](https://essentia.upf.edu/reference/std_PredominantPitchMelodia.html) estimates the fundamental frequency of the predominant melody using the MELODIA algorithm.\n\nQuite a mouthful, nevertheless it looks like there are a bunch of methods based on spectrograms, pitch calculations and other signal processing tricks.\n\nMy questions then is how can we correctly define when we are facing a pretrained algo or not? For instance MFCC (and maybe the PitchMelodia algo above) is not a pretrained model but the formula for obtaining the coefficients may vaguely recall the concept of exploiting a pre-trained model, no?\n\nJust another bordeline example, what about this [algorithm here](https://essentia.upf.edu/reference/streaming_Dissonance.html), able to detect sensory dissonance, where the documentations says `the dissonance curves are based on perceptual experiments conducted in [paper cited]`?\n\nThanks again for your time, I'm sure that other kagglers may find all this useful as well.",
    "1727424": "This is a great question and there certainly is some grey area here. I'm discussing this currently on my twitch stream and the short answer is: was the model trained on a corpus of data? In the cases you list above I believe the answer is no so they should be ok. Downloading a model that was trained on 100k audio files would not be ok. If you have more specific examples feel free to ask.",
    "1728117": "Sorry, I forgot to put the second page. I assume those functions are good to use for feature extractions. \n\nhttps://librosa.org/doc/main/feature.html",
    "1728392": "```\nYou are allowed to use offline compute where you train your models, the model weights can then be uploaded here and you need to submit via a kernel.\n```\nDoes it mean I can train my model entirely offline and upload the weight just for inference on the test set? The code for training model will be public then @robikscube . Thanks",
    "1728410": "Yes, librosa is fine to use.",
    "1728412": "Yes, you are allowed to train your model offline (only on the training data) and then upload to a kaggle dataset for inference.",
    "1728512": "dienhoa This is quite common on Kaggle, please checkout the BirdCLEF competition which IIRC follows a similar setup. \n\nI was actually really happy to see that GM Rob has setup a similar community competition for us, similar to a \"real\" (points) one."
  },
  "source": "meta"
}