{
  "id": 179766,
  "title": "What loss function are you using?",
  "url": "/competitions/birdsong-recognition/discussion/179766",
  "author_name": "",
  "post_date": "2020-09-02T18:40:57.400186100Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>Since the test data requires a multilabel classification i'm using the provided <code>secondary_labels</code> to train my model. I have used <code>BCEWithLogitsLoss</code> but my results are not improving. Training more epochs makes my results even worse. I'm using AudioSegment instead of Librosa. Does anyone know what is wrong or do you have bad experience with BCEWithLogitsLoss in this competition?<br>\nWhat loss function are you using?</p>\n<p>For Validation i'm using F1-Score with <code>average='samples'</code> it increased my score to 0.544 eventhough I know, that 0.544 is a full <code>nocall</code> submission. Tried it on <code>birdcall-check</code> the results looked good.</p>\n<p>Thank you in advance.</p>",
  "messages": [
    {
      "id": "995759",
      "postDate": "09/02/2020 18:40:57",
      "content": "<p>Hi,</p>\n<p>Since the test data requires a multilabel classification i'm using the provided <code>secondary_labels</code> to train my model. I have used <code>BCEWithLogitsLoss</code> but my results are not improving. Training more epochs makes my results even worse. I'm using AudioSegment instead of Librosa. Does anyone know what is wrong or do you have bad experience with BCEWithLogitsLoss in this competition?<br>\nWhat loss function are you using?</p>\n<p>For Validation i'm using F1-Score with <code>average='samples'</code> it increased my score to 0.544 eventhough I know, that 0.544 is a full <code>nocall</code> submission. Tried it on <code>birdcall-check</code> the results looked good.</p>\n<p>Thank you in advance.</p>",
      "rawMarkdown": "Hi,\n\nSince the test data requires a multilabel classification i'm using the provided `secondary_labels` to train my model. I have used `BCEWithLogitsLoss` but my results are not improving. Training more epochs makes my results even worse. I'm using AudioSegment instead of Librosa. Does anyone know what is wrong or do you have bad experience with BCEWithLogitsLoss in this competition?\nWhat loss function are you using?\n\nFor Validation i'm using F1-Score with `average='samples'` it increased my score to 0.544 eventhough I know, that 0.544 is a full `nocall` submission. Tried it on `birdcall-check` the results looked good.\n\nThank you in advance.",
      "votes": null
    },
    {
      "id": "995898",
      "postDate": "09/02/2020 23:30:15",
      "content": "<blockquote>\n  <p>Training more epochs makes my results even worse</p>\n</blockquote>\n<p>I guess at that point your model just tries to memorize the audio without learning any more patterns (A.k.a overfitting)</p>",
      "rawMarkdown": "> Training more epochs makes my results even worse\n\nI guess at that point your model just tries to memorize the audio without learning any more patterns (A.k.a overfitting)",
      "votes": null
    },
    {
      "id": "995920",
      "postDate": "09/03/2020 00:51:11",
      "content": "<p>I found that using secondary labels to train gave worse results than primary labels. As <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> has mentioned previously, secondary labels are very weak labels in high quality audio clips provided in this competition. I am thinking denoising, then doing mixup may provide better results for the multilabel classification problem but I have not tried it</p>",
      "rawMarkdown": "I found that using secondary labels to train gave worse results than primary labels. As @hidehisaarai1213 has mentioned previously, secondary labels are very weak labels in high quality audio clips provided in this competition. I am thinking denoising, then doing mixup may provide better results for the multilabel classification problem but I have not tried it",
      "votes": null
    },
    {
      "id": "995957",
      "postDate": "09/03/2020 01:48:06",
      "content": "<p>Btw I am also using BCEWithLogitsLoss</p>",
      "rawMarkdown": "Btw I am also using BCEWithLogitsLoss",
      "votes": null
    },
    {
      "id": "996214",
      "postDate": "09/03/2020 06:36:15",
      "content": "<p>Yes but I do provide a metric to detect overfitting, also by randomly choosing 5 Seconds from any clip it has alot of different data to read from. That means that I would need 120 epochs to go through every 5 seconds in each audiofile.</p>",
      "rawMarkdown": "Yes but I do provide a metric to detect overfitting, also by randomly choosing 5 Seconds from any clip it has alot of different data to read from. That means that I would need 120 epochs to go through every 5 seconds in each audiofile.",
      "votes": null
    },
    {
      "id": "997867",
      "postDate": "09/04/2020 09:46:37",
      "content": "<p>Thank you for your reply, this competiton does not make it easy to track what gives an improvement and what does not. Personally I have shaky results once its better then appying the same technique makes it worse then a previous technique etc.</p>\n<p>I will remove secondary labels then, thank you.</p>",
      "rawMarkdown": "Thank you for your reply, this competiton does not make it easy to track what gives an improvement and what does not. Personally I have shaky results once its better then appying the same technique makes it worse then a previous technique etc.\n\nI will remove secondary labels then, thank you.",
      "votes": null
    },
    {
      "id": "998776",
      "postDate": "09/05/2020 04:19:10",
      "content": "<p>Is the result CV or LB? I would not trust the LB since all nocall scores 54.4. If your CV improves by training more epochs then your model is not probably overfitting</p>",
      "rawMarkdown": "Is the result CV or LB? I would not trust the LB since all nocall scores 54.4. If your CV improves by training more epochs then your model is not probably overfitting",
      "votes": null
    },
    {
      "id": "999501",
      "postDate": "09/05/2020 17:37:58",
      "content": "<p>Bad results are on LB and I get &gt;0.60 as CV result. Yes but in general the test-set includes a lot of <code>nocall</code> classes still my LB is very bad, with increasing CV my LB gets worse and worse. I dont understand that phenomenon.</p>",
      "rawMarkdown": "Bad results are on LB and I get >0.60 as CV result. Yes but in general the test-set includes a lot of `nocall` classes still my LB is very bad, with increasing CV my LB gets worse and worse. I dont understand that phenomenon.",
      "votes": null
    },
    {
      "id": "999819",
      "postDate": "09/06/2020 04:00:27",
      "content": "<p>It could be that your model are better at predicting bird calls but worse at predicting nocall</p>",
      "rawMarkdown": "It could be that your model are better at predicting bird calls but worse at predicting nocall",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 995898,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "09/02/2020 23:30:15",
      "content": "<blockquote>\n  <p>Training more epochs makes my results even worse</p>\n</blockquote>\n<p>I guess at that point your model just tries to memorize the audio without learning any more patterns (A.k.a overfitting)</p>",
      "votes": null,
      "replies": [
        {
          "id": 996214,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/03/2020 06:36:15",
          "content": "<p>Yes but I do provide a metric to detect overfitting, also by randomly choosing 5 Seconds from any clip it has alot of different data to read from. That means that I would need 120 epochs to go through every 5 seconds in each audiofile.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998776,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/05/2020 04:19:10",
          "content": "<p>Is the result CV or LB? I would not trust the LB since all nocall scores 54.4. If your CV improves by training more epochs then your model is not probably overfitting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 999501,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/05/2020 17:37:58",
          "content": "<p>Bad results are on LB and I get &gt;0.60 as CV result. Yes but in general the test-set includes a lot of <code>nocall</code> classes still my LB is very bad, with increasing CV my LB gets worse and worse. I dont understand that phenomenon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 999819,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "09/06/2020 04:00:27",
          "content": "<p>It could be that your model are better at predicting bird calls but worse at predicting nocall</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 995920,
      "author_name": "alanchn31",
      "author_url": "",
      "post_date": "09/03/2020 00:51:11",
      "content": "<p>I found that using secondary labels to train gave worse results than primary labels. As <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> has mentioned previously, secondary labels are very weak labels in high quality audio clips provided in this competition. I am thinking denoising, then doing mixup may provide better results for the multilabel classification problem but I have not tried it</p>",
      "votes": null,
      "replies": [
        {
          "id": 995957,
          "author_name": "alanchn31",
          "author_url": "",
          "post_date": "09/03/2020 01:48:06",
          "content": "<p>Btw I am also using BCEWithLogitsLoss</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997867,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/04/2020 09:46:37",
          "content": "<p>Thank you for your reply, this competiton does not make it easy to track what gives an improvement and what does not. Personally I have shaky results once its better then appying the same technique makes it worse then a previous technique etc.</p>\n<p>I will remove secondary labels then, thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "995759": "Hi,\n\nSince the test data requires a multilabel classification i'm using the provided `secondary_labels` to train my model. I have used `BCEWithLogitsLoss` but my results are not improving. Training more epochs makes my results even worse. I'm using AudioSegment instead of Librosa. Does anyone know what is wrong or do you have bad experience with BCEWithLogitsLoss in this competition?\nWhat loss function are you using?\n\nFor Validation i'm using F1-Score with `average='samples'` it increased my score to 0.544 eventhough I know, that 0.544 is a full `nocall` submission. Tried it on `birdcall-check` the results looked good.\n\nThank you in advance.",
    "995898": "> Training more epochs makes my results even worse\n\nI guess at that point your model just tries to memorize the audio without learning any more patterns (A.k.a overfitting)",
    "995920": "I found that using secondary labels to train gave worse results than primary labels. As @hidehisaarai1213 has mentioned previously, secondary labels are very weak labels in high quality audio clips provided in this competition. I am thinking denoising, then doing mixup may provide better results for the multilabel classification problem but I have not tried it",
    "995957": "Btw I am also using BCEWithLogitsLoss",
    "996214": "Yes but I do provide a metric to detect overfitting, also by randomly choosing 5 Seconds from any clip it has alot of different data to read from. That means that I would need 120 epochs to go through every 5 seconds in each audiofile.",
    "997867": "Thank you for your reply, this competiton does not make it easy to track what gives an improvement and what does not. Personally I have shaky results once its better then appying the same technique makes it worse then a previous technique etc.\n\nI will remove secondary labels then, thank you.",
    "998776": "Is the result CV or LB? I would not trust the LB since all nocall scores 54.4. If your CV improves by training more epochs then your model is not probably overfitting",
    "999501": "Bad results are on LB and I get >0.60 as CV result. Yes but in general the test-set includes a lot of `nocall` classes still my LB is very bad, with increasing CV my LB gets worse and worse. I dont understand that phenomenon.",
    "999819": "It could be that your model are better at predicting bird calls but worse at predicting nocall"
  },
  "source": "meta"
}