{
  "id": 220727,
  "title": "Many thanks for making me broaden my knowledge !",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220727",
  "author_name": "",
  "post_date": "2021-02-19T10:12:34.697129500Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Have anyone of you experienced such phenomenon:<br>\nMy best cv was (5 folds) 0.904, 0.901, 0.905, 0.903, 0.900 , trusting it costed me 1100 places shake down 😂🙈🤐 I selected one highest LB  (0.904- average ensemble) and one highest CV, but didn't work for me…<br>\nI used Early Layer Regularization loss (noise: 0.4 , beta: 2), 5 TTAs, efficientnetb4_ns, light augmentations, Adam optimizer and step lr.</p>\n<p>For a first competition I am really grateful for the knowledge I have gained. Any comments for how to prevent such shakeups in the future are welcomed 😊🎉</p>",
  "messages": [
    {
      "id": "1210280",
      "postDate": "02/19/2021 10:12:34",
      "content": "<p>Have anyone of you experienced such phenomenon:<br>\nMy best cv was (5 folds) 0.904, 0.901, 0.905, 0.903, 0.900 , trusting it costed me 1100 places shake down 😂🙈🤐 I selected one highest LB  (0.904- average ensemble) and one highest CV, but didn't work for me…<br>\nI used Early Layer Regularization loss (noise: 0.4 , beta: 2), 5 TTAs, efficientnetb4_ns, light augmentations, Adam optimizer and step lr.</p>\n<p>For a first competition I am really grateful for the knowledge I have gained. Any comments for how to prevent such shakeups in the future are welcomed 😊🎉</p>",
      "rawMarkdown": "Have anyone of you experienced such phenomenon:\nMy best cv was (5 folds) 0.904, 0.901, 0.905, 0.903, 0.900 , trusting it costed me 1100 places shake down 😂🙈🤐 I selected one highest LB  (0.904- average ensemble) and one highest CV, but didn't work for me…\nI used Early Layer Regularization loss (noise: 0.4 , beta: 2), 5 TTAs, efficientnetb4_ns, light augmentations, Adam optimizer and step lr.\n\nFor a first competition I am really grateful for the knowledge I have gained. Any comments for how to prevent such shakeups in the future are welcomed 😊🎉",
      "votes": null
    },
    {
      "id": "1211896",
      "postDate": "02/20/2021 16:35:01",
      "content": "<p>Firstly, accuracy is not a great metric, it's very based on one category getting a slightly higher probability than the others. So, there's a certain amount of luck especially when lots of teams are pretty close to each other. </p>\n<p>Another consequence of the metric could be its shakiness as a validation metric, I also found it reassuring when other metrics like validation loss favoured a model. Additionally, it made be a little paranoid about relying on early stopping (much easier to overfit validation accuracy by early stopping than validation log-loss).</p>\n<p>Secondly, there's always a risk of leaking information (mostly if you used the 2019 data).</p>",
      "rawMarkdown": "Firstly, accuracy is not a great metric, it's very based on one category getting a slightly higher probability than the others. So, there's a certain amount of luck especially when lots of teams are pretty close to each other. \n\nAnother consequence of the metric could be its shakiness as a validation metric, I also found it reassuring when other metrics like validation loss favoured a model. Additionally, it made be a little paranoid about relying on early stopping (much easier to overfit validation accuracy by early stopping than validation log-loss).\n\nSecondly, there's always a risk of leaking information (mostly if you used the 2019 data).",
      "votes": null
    },
    {
      "id": "1211919",
      "postDate": "02/20/2021 17:01:51",
      "content": "<p>Thank you mate for the clarification. Maybe area under the precision - recall curve would be a good metric, or f1 score i quess?<br>\nIn fact, i didn't use the old data. Yea, maybe i did mistake for not trusting my validation loss… </p>",
      "rawMarkdown": "Thank you mate for the clarification. Maybe area under the precision - recall curve would be a good metric, or f1 score i quess?\nIn fact, i didn't use the old data. Yea, maybe i did mistake for not trusting my validation loss...",
      "votes": null
    },
    {
      "id": "1211933",
      "postDate": "02/20/2021 17:13:47",
      "content": "<p>I'm not sure that I truly figured out the answer for it. I relied on validation accuracy, but sanity checked it via label smoothed cross entropy on the validation set as a metric (and occasionally made some decisions based on that). With F1-score depending how you deal with multiple classes, I'd worry it would reward getting rare classes right too much (when accuracy cares less). I'm less sure how AuROC would behave in a multi class problem.</p>",
      "rawMarkdown": "I'm not sure that I truly figured out the answer for it. I relied on validation accuracy, but sanity checked it via label smoothed cross entropy on the validation set as a metric (and occasionally made some decisions based on that). With F1-score depending how you deal with multiple classes, I'd worry it would reward getting rare classes right too much (when accuracy cares less). I'm less sure how AuROC would behave in a multi class problem.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1211896,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "02/20/2021 16:35:01",
      "content": "<p>Firstly, accuracy is not a great metric, it's very based on one category getting a slightly higher probability than the others. So, there's a certain amount of luck especially when lots of teams are pretty close to each other. </p>\n<p>Another consequence of the metric could be its shakiness as a validation metric, I also found it reassuring when other metrics like validation loss favoured a model. Additionally, it made be a little paranoid about relying on early stopping (much easier to overfit validation accuracy by early stopping than validation log-loss).</p>\n<p>Secondly, there's always a risk of leaking information (mostly if you used the 2019 data).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211919,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/20/2021 17:01:51",
          "content": "<p>Thank you mate for the clarification. Maybe area under the precision - recall curve would be a good metric, or f1 score i quess?<br>\nIn fact, i didn't use the old data. Yea, maybe i did mistake for not trusting my validation loss… </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211933,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "02/20/2021 17:13:47",
          "content": "<p>I'm not sure that I truly figured out the answer for it. I relied on validation accuracy, but sanity checked it via label smoothed cross entropy on the validation set as a metric (and occasionally made some decisions based on that). With F1-score depending how you deal with multiple classes, I'd worry it would reward getting rare classes right too much (when accuracy cares less). I'm less sure how AuROC would behave in a multi class problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1210280": "Have anyone of you experienced such phenomenon:\nMy best cv was (5 folds) 0.904, 0.901, 0.905, 0.903, 0.900 , trusting it costed me 1100 places shake down 😂🙈🤐 I selected one highest LB  (0.904- average ensemble) and one highest CV, but didn't work for me…\nI used Early Layer Regularization loss (noise: 0.4 , beta: 2), 5 TTAs, efficientnetb4_ns, light augmentations, Adam optimizer and step lr.\n\nFor a first competition I am really grateful for the knowledge I have gained. Any comments for how to prevent such shakeups in the future are welcomed 😊🎉",
    "1211896": "Firstly, accuracy is not a great metric, it's very based on one category getting a slightly higher probability than the others. So, there's a certain amount of luck especially when lots of teams are pretty close to each other. \n\nAnother consequence of the metric could be its shakiness as a validation metric, I also found it reassuring when other metrics like validation loss favoured a model. Additionally, it made be a little paranoid about relying on early stopping (much easier to overfit validation accuracy by early stopping than validation log-loss).\n\nSecondly, there's always a risk of leaking information (mostly if you used the 2019 data).",
    "1211919": "Thank you mate for the clarification. Maybe area under the precision - recall curve would be a good metric, or f1 score i quess?\nIn fact, i didn't use the old data. Yea, maybe i did mistake for not trusting my validation loss...",
    "1211933": "I'm not sure that I truly figured out the answer for it. I relied on validation accuracy, but sanity checked it via label smoothed cross entropy on the validation set as a metric (and occasionally made some decisions based on that). With F1-score depending how you deal with multiple classes, I'd worry it would reward getting rare classes right too much (when accuracy cares less). I'm less sure how AuROC would behave in a multi class problem."
  },
  "source": "meta"
}