{
  "id": 178753,
  "title": "Validation approach",
  "url": "/competitions/birdsong-recognition/discussion/178753",
  "author_name": "",
  "post_date": "2020-08-31T08:29:57.607998700Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Seen a notable event, validation scores doesn't not correlate well with the Public LB.<br>\nMay be local problem but worth highlighting.</p>\n<p>With the same training setup, picking the best val checkpoint gets worse Public LB.<br>\nMore training epochs with higher val loss beats a lower val loss with less epochs.<br>\nEven with same epochs and higher val loss can get a better Public LBs score.<br>\nHave also noticed that SWA from best ckps doesn't work well.</p>\n<p>Maybe the number of epochs, in relation to the amount of train data samples can be more important than the val loss, using fixed step scheduler and stop based on LB analyzes, worth trying.</p>\n<p>Or one just started a journey overfitting against the public test set! <br>\nMaybe time to implement a local F1 metric.</p>",
  "messages": [
    {
      "id": "992561",
      "postDate": "08/31/2020 08:29:57",
      "content": "<p>Seen a notable event, validation scores doesn't not correlate well with the Public LB.<br>\nMay be local problem but worth highlighting.</p>\n<p>With the same training setup, picking the best val checkpoint gets worse Public LB.<br>\nMore training epochs with higher val loss beats a lower val loss with less epochs.<br>\nEven with same epochs and higher val loss can get a better Public LBs score.<br>\nHave also noticed that SWA from best ckps doesn't work well.</p>\n<p>Maybe the number of epochs, in relation to the amount of train data samples can be more important than the val loss, using fixed step scheduler and stop based on LB analyzes, worth trying.</p>\n<p>Or one just started a journey overfitting against the public test set! <br>\nMaybe time to implement a local F1 metric.</p>",
      "rawMarkdown": "Seen a notable event, validation scores doesn't not correlate well with the Public LB.\nMay be local problem but worth highlighting.\n\nWith the same training setup, picking the best val checkpoint gets worse Public LB.\nMore training epochs with higher val loss beats a lower val loss with less epochs.\nEven with same epochs and higher val loss can get a better Public LBs score.\nHave also noticed that SWA from best ckps doesn't work well.\n\nMaybe the number of epochs, in relation to the amount of train data samples can be more important than the val loss, using fixed step scheduler and stop based on LB analyzes, worth trying.\n\nOr one just started a journey overfitting against the public test set! \nMaybe time to implement a local F1 metric.",
      "votes": null
    },
    {
      "id": "992599",
      "postDate": "08/31/2020 09:16:28",
      "content": "<p>I have also noticed the same obversations. It is suprising because it is completly different from the natural training observations and the fundamental theory of deep learning. Thats why I will personally focus on CV in this competition because I feel like this is going into the overfitting direction. I hope that the remaining 72% test data in private LB will cause a shake up.</p>\n<p>What was your validation metric, only val-loss or did you try F1-Score with options samples?</p>",
      "rawMarkdown": "I have also noticed the same obversations. It is suprising because it is completly different from the natural training observations and the fundamental theory of deep learning. Thats why I will personally focus on CV in this competition because I feel like this is going into the overfitting direction. I hope that the remaining 72% test data in private LB will cause a shake up.\n\nWhat was your validation metric, only val-loss or did you try F1-Score with options samples?",
      "votes": null
    },
    {
      "id": "992615",
      "postDate": "08/31/2020 09:30:23",
      "content": "<p>Yes only val-loss, time to impl. a local F1 metric. But I also have in mind, that in previous competitions' top team solutions with the same phenomenon used the approach to stop focus at local classification validation, although it is difficult to draw parallels. Too early to draw any conclusions in this case, more testing needed, but worth looking into. If the \"problem\" persists, I will choose one CV/local-val and one LB/fixed-epoch approach, just to be safe.</p>",
      "rawMarkdown": "Yes only val-loss, time to impl. a local F1 metric. But I also have in mind, that in previous competitions' top team solutions with the same phenomenon used the approach to stop focus at local classification validation, although it is difficult to draw parallels. Too early to draw any conclusions in this case, more testing needed, but worth looking into. If the \"problem\" persists, I will choose one CV/local-val and one LB/fixed-epoch approach, just to be safe.",
      "votes": null
    },
    {
      "id": "992830",
      "postDate": "08/31/2020 13:22:32",
      "content": "<p>Becareful since all-nocall submission can achieve 54.4 on public LB</p>",
      "rawMarkdown": "Becareful since all-nocall submission can achieve 54.4 on public LB",
      "votes": null
    },
    {
      "id": "995748",
      "postDate": "09/02/2020 18:31:55",
      "content": "<p>I have tried val-loss and I got worse scoring on LB then using F1-Score als metric. Still I trained for 65 Epochs and got a LB Score of: 0.028 training on F1-Score for 65 Epochs brings: 0.118 LB, now training on less epochs 25 and on F1-Score metric I score: 0.488 on LB</p>\n<p>It does not seem like it an overfitted LB score because my longer trained models are very bad.<br>\nI have no idea why this happens when training longer. Its just against the fundamentals of training a model. </p>",
      "rawMarkdown": "I have tried val-loss and I got worse scoring on LB then using F1-Score als metric. Still I trained for 65 Epochs and got a LB Score of: 0.028 training on F1-Score for 65 Epochs brings: 0.118 LB, now training on less epochs 25 and on F1-Score metric I score: 0.488 on LB\n\nIt does not seem like it an overfitted LB score because my longer trained models are very bad.\nI have no idea why this happens when training longer. Its just against the fundamentals of training a model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 992599,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/31/2020 09:16:28",
      "content": "<p>I have also noticed the same obversations. It is suprising because it is completly different from the natural training observations and the fundamental theory of deep learning. Thats why I will personally focus on CV in this competition because I feel like this is going into the overfitting direction. I hope that the remaining 72% test data in private LB will cause a shake up.</p>\n<p>What was your validation metric, only val-loss or did you try F1-Score with options samples?</p>",
      "votes": null,
      "replies": [
        {
          "id": 992615,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/31/2020 09:30:23",
          "content": "<p>Yes only val-loss, time to impl. a local F1 metric. But I also have in mind, that in previous competitions' top team solutions with the same phenomenon used the approach to stop focus at local classification validation, although it is difficult to draw parallels. Too early to draw any conclusions in this case, more testing needed, but worth looking into. If the \"problem\" persists, I will choose one CV/local-val and one LB/fixed-epoch approach, just to be safe.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 995748,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/02/2020 18:31:55",
          "content": "<p>I have tried val-loss and I got worse scoring on LB then using F1-Score als metric. Still I trained for 65 Epochs and got a LB Score of: 0.028 training on F1-Score for 65 Epochs brings: 0.118 LB, now training on less epochs 25 and on F1-Score metric I score: 0.488 on LB</p>\n<p>It does not seem like it an overfitted LB score because my longer trained models are very bad.<br>\nI have no idea why this happens when training longer. Its just against the fundamentals of training a model. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 992830,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/31/2020 13:22:32",
      "content": "<p>Becareful since all-nocall submission can achieve 54.4 on public LB</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "992561": "Seen a notable event, validation scores doesn't not correlate well with the Public LB.\nMay be local problem but worth highlighting.\n\nWith the same training setup, picking the best val checkpoint gets worse Public LB.\nMore training epochs with higher val loss beats a lower val loss with less epochs.\nEven with same epochs and higher val loss can get a better Public LBs score.\nHave also noticed that SWA from best ckps doesn't work well.\n\nMaybe the number of epochs, in relation to the amount of train data samples can be more important than the val loss, using fixed step scheduler and stop based on LB analyzes, worth trying.\n\nOr one just started a journey overfitting against the public test set! \nMaybe time to implement a local F1 metric.",
    "992599": "I have also noticed the same obversations. It is suprising because it is completly different from the natural training observations and the fundamental theory of deep learning. Thats why I will personally focus on CV in this competition because I feel like this is going into the overfitting direction. I hope that the remaining 72% test data in private LB will cause a shake up.\n\nWhat was your validation metric, only val-loss or did you try F1-Score with options samples?",
    "992615": "Yes only val-loss, time to impl. a local F1 metric. But I also have in mind, that in previous competitions' top team solutions with the same phenomenon used the approach to stop focus at local classification validation, although it is difficult to draw parallels. Too early to draw any conclusions in this case, more testing needed, but worth looking into. If the \"problem\" persists, I will choose one CV/local-val and one LB/fixed-epoch approach, just to be safe.",
    "992830": "Becareful since all-nocall submission can achieve 54.4 on public LB",
    "995748": "I have tried val-loss and I got worse scoring on LB then using F1-Score als metric. Still I trained for 65 Epochs and got a LB Score of: 0.028 training on F1-Score for 65 Epochs brings: 0.118 LB, now training on less epochs 25 and on F1-Score metric I score: 0.488 on LB\n\nIt does not seem like it an overfitted LB score because my longer trained models are very bad.\nI have no idea why this happens when training longer. Its just against the fundamentals of training a model."
  },
  "source": "meta"
}