{
  "id": 178297,
  "title": "Training time & CV metric",
  "url": "/competitions/birdsong-recognition/discussion/178297",
  "author_name": "",
  "post_date": "2020-08-29T12:44:25.632741300Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As the title, I think these questions are particularly relevant to this comp :).<br>\n1) What metric &amp; scheme are you using for local CV?<br>\n    -Personally I have experimented with thresholded f1, soft f1, mAP etc. Neither seems good indicator of LB score. Currently I am using good o' validation loss (BCE). This makes a bit of sense because validation BCE seems to capture our models' performance for both good &amp; positive predictions. There are submissions of models trained longer, with higher CV f1 and lower val BCE which scores terrible on the LB<br>\n2) How long (how many epochs) do you train for?<br>\n    -I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.</p>\n<p>What are your two cents? (wink</p>",
  "messages": [
    {
      "id": "990210",
      "postDate": "08/29/2020 12:44:25",
      "content": "<p>As the title, I think these questions are particularly relevant to this comp :).<br>\n1) What metric &amp; scheme are you using for local CV?<br>\n    -Personally I have experimented with thresholded f1, soft f1, mAP etc. Neither seems good indicator of LB score. Currently I am using good o' validation loss (BCE). This makes a bit of sense because validation BCE seems to capture our models' performance for both good &amp; positive predictions. There are submissions of models trained longer, with higher CV f1 and lower val BCE which scores terrible on the LB<br>\n2) How long (how many epochs) do you train for?<br>\n    -I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.</p>\n<p>What are your two cents? (wink</p>",
      "rawMarkdown": "As the title, I think these questions are particularly relevant to this comp :).\n1) What metric & scheme are you using for local CV?\n    -Personally I have experimented with thresholded f1, soft f1, mAP etc. Neither seems good indicator of LB score. Currently I am using good o' validation loss (BCE). This makes a bit of sense because validation BCE seems to capture our models' performance for both good & positive predictions. There are submissions of models trained longer, with higher CV f1 and lower val BCE which scores terrible on the LB\n2) How long (how many epochs) do you train for?\n    -I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.\n\nWhat are your two cents? (wink",
      "votes": null
    },
    {
      "id": "992017",
      "postDate": "08/30/2020 20:11:36",
      "content": "<p>1) I'm using thresholded F1 with the option <code>average='samples'</code>, its stated to be the competition metric:<br>\n<a href=\"https://www.kaggle.com/shonenkov/competition-metrics\" target=\"_blank\">https://www.kaggle.com/shonenkov/competition-metrics</a> i'm still playing around with it as it is not really working for me as a good indication for LB. Additonally I'm using <code>BCEWithLogitsLoss()</code> is using <code>BCELoss</code> better in any way?</p>\n<p>2) I'm using early stopping and usually train 40 epochs and with a low LR I train for 65 epochs, my results are as you stated, trained on higher epochs I score less than trained on a few epochs. It does confuse me and I did not figure out why that happens.</p>",
      "rawMarkdown": "1) I'm using thresholded F1 with the option `average='samples'`, its stated to be the competition metric:\nhttps://www.kaggle.com/shonenkov/competition-metrics i'm still playing around with it as it is not really working for me as a good indication for LB. Additonally I'm using `BCEWithLogitsLoss()` is using `BCELoss` better in any way?\n\n2) I'm using early stopping and usually train 40 epochs and with a low LR I train for 65 epochs, my results are as you stated, trained on higher epochs I score less than trained on a few epochs. It does confuse me and I did not figure out why that happens.",
      "votes": null
    },
    {
      "id": "992148",
      "postDate": "08/31/2020 01:43:46",
      "content": "<blockquote>\n  <p>I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.</p>\n</blockquote>\n<p>Was the CV's loss lower on epoch 41? I would 100% trust my CV if that's the case because generally longer training means better performance (assume that there's no overfitting takes place).</p>",
      "rawMarkdown": "> I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.\n\nWas the CV's loss lower on epoch 41? I would 100% trust my CV if that's the case because generally longer training means better performance (assume that there's no overfitting takes place).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 992017,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/30/2020 20:11:36",
      "content": "<p>1) I'm using thresholded F1 with the option <code>average='samples'</code>, its stated to be the competition metric:<br>\n<a href=\"https://www.kaggle.com/shonenkov/competition-metrics\" target=\"_blank\">https://www.kaggle.com/shonenkov/competition-metrics</a> i'm still playing around with it as it is not really working for me as a good indication for LB. Additonally I'm using <code>BCEWithLogitsLoss()</code> is using <code>BCELoss</code> better in any way?</p>\n<p>2) I'm using early stopping and usually train 40 epochs and with a low LR I train for 65 epochs, my results are as you stated, trained on higher epochs I score less than trained on a few epochs. It does confuse me and I did not figure out why that happens.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 992148,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/31/2020 01:43:46",
      "content": "<blockquote>\n  <p>I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.</p>\n</blockquote>\n<p>Was the CV's loss lower on epoch 41? I would 100% trust my CV if that's the case because generally longer training means better performance (assume that there's no overfitting takes place).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "990210": "As the title, I think these questions are particularly relevant to this comp :).\n1) What metric & scheme are you using for local CV?\n    -Personally I have experimented with thresholded f1, soft f1, mAP etc. Neither seems good indicator of LB score. Currently I am using good o' validation loss (BCE). This makes a bit of sense because validation BCE seems to capture our models' performance for both good & positive predictions. There are submissions of models trained longer, with higher CV f1 and lower val BCE which scores terrible on the LB\n2) How long (how many epochs) do you train for?\n    -I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.\n\nWhat are your two cents? (wink",
    "992017": "1) I'm using thresholded F1 with the option `average='samples'`, its stated to be the competition metric:\nhttps://www.kaggle.com/shonenkov/competition-metrics i'm still playing around with it as it is not really working for me as a good indication for LB. Additonally I'm using `BCEWithLogitsLoss()` is using `BCELoss` better in any way?\n\n2) I'm using early stopping and usually train 40 epochs and with a low LR I train for 65 epochs, my results are as you stated, trained on higher epochs I score less than trained on a few epochs. It does confuse me and I did not figure out why that happens.",
    "992148": "> I think there exists a point beyond which LB fails to match CV. One model at 18 epochs scored .567 on LB but the same modela t 41 epochs scored only .544.\n\nWas the CV's loss lower on epoch 41? I would 100% trust my CV if that's the case because generally longer training means better performance (assume that there's no overfitting takes place)."
  },
  "source": "meta"
}