{
  "id": 134601,
  "title": "Confusing CV and LB",
  "url": "/competitions/bengaliai-cv19/discussion/134601",
  "author_name": "",
  "post_date": "2020-03-09T07:19:51.091002700Z",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Been doing some experiments and CV/LB gap is becoming quite frustrating.\nWhen in .97~.98 range the CV/LB is quite consistent. I found the gap breaking down then CV is high.</p>\n\n<p>Same approach, same fold\nEfficientnet-b4: CV.9967 LB .9887\nEfficientnet-b5: CV .9975 LB .9886\nEfficientnet-b6: CV .9979 LB .9886</p>\n\n<p>Wondering if I should trust my CV or LB...Any thoughts? Why would deeper models perform worse on LB? Anyone observing the same issues?</p>",
  "messages": [
    {
      "id": "767109",
      "postDate": "03/09/2020 07:19:51",
      "content": "<p>Been doing some experiments and CV/LB gap is becoming quite frustrating.\nWhen in .97~.98 range the CV/LB is quite consistent. I found the gap breaking down then CV is high.</p>\n\n<p>Same approach, same fold\nEfficientnet-b4: CV.9967 LB .9887\nEfficientnet-b5: CV .9975 LB .9886\nEfficientnet-b6: CV .9979 LB .9886</p>\n\n<p>Wondering if I should trust my CV or LB...Any thoughts? Why would deeper models perform worse on LB? Anyone observing the same issues?</p>",
      "rawMarkdown": "Been doing some experiments and CV/LB gap is becoming quite frustrating.\nWhen in .97~.98 range the CV/LB is quite consistent. I found the gap breaking down then CV is high.\n\nSame approach, same fold\nEfficientnet-b4: CV.9967 LB .9887\nEfficientnet-b5: CV .9975 LB .9886\nEfficientnet-b6: CV .9979 LB .9886\n\nWondering if I should trust my CV or LB...Any thoughts? Why would deeper models perform worse on LB? Anyone observing the same issues?",
      "votes": null
    },
    {
      "id": "767118",
      "postDate": "03/09/2020 07:26:10",
      "content": "<p>The most likely reason is poor performance on unseen graphemes in public (And possibly private). See discussion here <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134035</a></p>\n\n<p>For deeper models  one guess is it can learn the combinations better than generalizing because of higher capacity, try validating with this approach <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134434</a></p>",
      "rawMarkdown": "The most likely reason is poor performance on unseen graphemes in public (And possibly private). See discussion here https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\n\nFor deeper models  one guess is it can learn the combinations better than generalizing because of higher capacity, try validating with this approach https://www.kaggle.com/c/bengaliai-cv19/discussion/134434",
      "votes": null
    },
    {
      "id": "767173",
      "postDate": "03/09/2020 09:03:30",
      "content": "<p>Not related to this topic, but could I know what GPU are you use for running efficient?</p>",
      "rawMarkdown": "Not related to this topic, but could I know what GPU are you use for running efficient?",
      "votes": null
    },
    {
      "id": "767187",
      "postDate": "03/09/2020 09:26:56",
      "content": "<p>For E6, LB 99.86!  </p>",
      "rawMarkdown": "For E6, LB 99.86!",
      "votes": null
    },
    {
      "id": "767230",
      "postDate": "03/09/2020 10:35:35",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> Sorry that was a typo...Would've been on holiday if b6 got that score...thanks for the catch</p>",
      "rawMarkdown": "ipythonx Sorry that was a typo...Would've been on holiday if b6 got that score...thanks for the catch",
      "votes": null
    },
    {
      "id": "767231",
      "postDate": "03/09/2020 10:35:57",
      "content": "<p>I am using titan RTX and efficientnet-b6 takes about 3 days to run...so the LB turned out really disappointing</p>",
      "rawMarkdown": "I am using titan RTX and efficientnet-b6 takes about 3 days to run...so the LB turned out really disappointing",
      "votes": null
    },
    {
      "id": "767296",
      "postDate": "03/09/2020 12:55:11",
      "content": "<p>I have similar issues. Models scoring 0.998 on CV score 0.987 vs. 0.997 scoring 0.988. I think it might have to do with overfitting on seen graphemes which hurts performance on unseen graphemes. </p>",
      "rawMarkdown": "I have similar issues. Models scoring 0.998 on CV score 0.987 vs. 0.997 scoring 0.988. I think it might have to do with overfitting on seen graphemes which hurts performance on unseen graphemes.",
      "votes": null
    },
    {
      "id": "767322",
      "postDate": "03/09/2020 13:34:34",
      "content": "<p><a href=\"/vaillant\">@vaillant</a> wondering many epochs are you training? I am doing 300. I guess it also has to do with extended training.</p>",
      "rawMarkdown": "vaillant wondering many epochs are you training? I am doing 300. I guess it also has to do with extended training.",
      "votes": null
    },
    {
      "id": "767342",
      "postDate": "03/09/2020 13:55:58",
      "content": "<p>Ops, I was planning to visit China to meet you 😄 \nAnyway, though I still couldn't come a bit near such a score of yours, I am also facing such an issue. I surely need some advice to come close around LB: 97~98!</p>\n\n<p>I observe to cases:</p>\n\n<blockquote>\n  <p>E5: CV : 96.70 - LB: 96.5\n  E7: CV : 96.35 - LB: 96.71</p>\n</blockquote>\n\n<p>In addition, I also notice lb score improves if the validation score on grapheme_root improves promisingly. </p>",
      "rawMarkdown": "Ops, I was planning to visit China to meet you 😄 \nAnyway, though I still couldn't come a bit near such a score of yours, I am also facing such an issue. I surely need some advice to come close around LB: 97~98!\n\nI observe to cases:\n&gt; E5: CV : 96.70 - LB: 96.5\nE7: CV : 96.35 - LB: 96.71\n\nIn addition, I also notice lb score improves if the validation score on grapheme_root improves promisingly.",
      "votes": null
    },
    {
      "id": "767347",
      "postDate": "03/09/2020 13:58:21",
      "content": "<p>60 epochs. At least for CV, I don't see any increase in score for training longer. </p>",
      "rawMarkdown": "60 epochs. At least for CV, I don't see any increase in score for training longer.",
      "votes": null
    },
    {
      "id": "767441",
      "postDate": "03/09/2020 16:36:01",
      "content": "<p>These are my own LB vs. CV scores:\n```\nLB / CV</p>\n\n<p>0.9874 / 0.9986\n0.9883 / 0.9984\n0.9879 / 0.9953\n0.9884 / 0.9967\n0.9876 / 0.9981\n0.9875 / 0.9968\n```</p>",
      "rawMarkdown": "These are my own LB vs. CV scores:\n```\nLB / CV\n\n0.9874 / 0.9986\n0.9883 / 0.9984\n0.9879 / 0.9953\n0.9884 / 0.9967\n0.9876 / 0.9981\n0.9875 / 0.9968\n```",
      "votes": null
    },
    {
      "id": "767906",
      "postDate": "03/10/2020 08:20:56",
      "content": "<p>Thank you. It is really taking a long time to train. </p>",
      "rawMarkdown": "Thank you. It is really taking a long time to train.",
      "votes": null
    },
    {
      "id": "768099",
      "postDate": "03/10/2020 13:13:49",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I guess all the details that are needed for LB .98 are stated in the discussions; no magic. No cropping helped me the most, got me from low .97 to .985+. The rest are just usual cutmix/mixup and usual training. Best of luck! Thank you for your generous sharing in this competition.</p>",
      "rawMarkdown": "ipythonx I guess all the details that are needed for LB .98 are stated in the discussions; no magic. No cropping helped me the most, got me from low .97 to .985+. The rest are just usual cutmix/mixup and usual training. Best of luck! Thank you for your generous sharing in this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 767118,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "03/09/2020 07:26:10",
      "content": "<p>The most likely reason is poor performance on unseen graphemes in public (And possibly private). See discussion here <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134035</a></p>\n\n<p>For deeper models  one guess is it can learn the combinations better than generalizing because of higher capacity, try validating with this approach <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134434</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 767173,
      "author_name": "kasim0226",
      "author_url": "",
      "post_date": "03/09/2020 09:03:30",
      "content": "<p>Not related to this topic, but could I know what GPU are you use for running efficient?</p>",
      "votes": null,
      "replies": [
        {
          "id": 767231,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/09/2020 10:35:57",
          "content": "<p>I am using titan RTX and efficientnet-b6 takes about 3 days to run...so the LB turned out really disappointing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767906,
          "author_name": "kasim0226",
          "author_url": "",
          "post_date": "03/10/2020 08:20:56",
          "content": "<p>Thank you. It is really taking a long time to train. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 767187,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "03/09/2020 09:26:56",
      "content": "<p>For E6, LB 99.86!  </p>",
      "votes": null,
      "replies": [
        {
          "id": 767230,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/09/2020 10:35:35",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> Sorry that was a typo...Would've been on holiday if b6 got that score...thanks for the catch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767342,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "03/09/2020 13:55:58",
          "content": "<p>Ops, I was planning to visit China to meet you 😄 \nAnyway, though I still couldn't come a bit near such a score of yours, I am also facing such an issue. I surely need some advice to come close around LB: 97~98!</p>\n\n<p>I observe to cases:</p>\n\n<blockquote>\n  <p>E5: CV : 96.70 - LB: 96.5\n  E7: CV : 96.35 - LB: 96.71</p>\n</blockquote>\n\n<p>In addition, I also notice lb score improves if the validation score on grapheme_root improves promisingly. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768099,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/10/2020 13:13:49",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I guess all the details that are needed for LB .98 are stated in the discussions; no magic. No cropping helped me the most, got me from low .97 to .985+. The rest are just usual cutmix/mixup and usual training. Best of luck! Thank you for your generous sharing in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 767296,
      "author_name": "vaillant",
      "author_url": "",
      "post_date": "03/09/2020 12:55:11",
      "content": "<p>I have similar issues. Models scoring 0.998 on CV score 0.987 vs. 0.997 scoring 0.988. I think it might have to do with overfitting on seen graphemes which hurts performance on unseen graphemes. </p>",
      "votes": null,
      "replies": [
        {
          "id": 767322,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/09/2020 13:34:34",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> wondering many epochs are you training? I am doing 300. I guess it also has to do with extended training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767347,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "03/09/2020 13:58:21",
          "content": "<p>60 epochs. At least for CV, I don't see any increase in score for training longer. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 767441,
      "author_name": "vaillant",
      "author_url": "",
      "post_date": "03/09/2020 16:36:01",
      "content": "<p>These are my own LB vs. CV scores:\n```\nLB / CV</p>\n\n<p>0.9874 / 0.9986\n0.9883 / 0.9984\n0.9879 / 0.9953\n0.9884 / 0.9967\n0.9876 / 0.9981\n0.9875 / 0.9968\n```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "767109": "Been doing some experiments and CV/LB gap is becoming quite frustrating.\nWhen in .97~.98 range the CV/LB is quite consistent. I found the gap breaking down then CV is high.\n\nSame approach, same fold\nEfficientnet-b4: CV.9967 LB .9887\nEfficientnet-b5: CV .9975 LB .9886\nEfficientnet-b6: CV .9979 LB .9886\n\nWondering if I should trust my CV or LB...Any thoughts? Why would deeper models perform worse on LB? Anyone observing the same issues?",
    "767118": "The most likely reason is poor performance on unseen graphemes in public (And possibly private). See discussion here https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\n\nFor deeper models  one guess is it can learn the combinations better than generalizing because of higher capacity, try validating with this approach https://www.kaggle.com/c/bengaliai-cv19/discussion/134434",
    "767173": "Not related to this topic, but could I know what GPU are you use for running efficient?",
    "767187": "For E6, LB 99.86!",
    "767230": "ipythonx Sorry that was a typo...Would've been on holiday if b6 got that score...thanks for the catch",
    "767231": "I am using titan RTX and efficientnet-b6 takes about 3 days to run...so the LB turned out really disappointing",
    "767296": "I have similar issues. Models scoring 0.998 on CV score 0.987 vs. 0.997 scoring 0.988. I think it might have to do with overfitting on seen graphemes which hurts performance on unseen graphemes.",
    "767322": "vaillant wondering many epochs are you training? I am doing 300. I guess it also has to do with extended training.",
    "767342": "Ops, I was planning to visit China to meet you 😄 \nAnyway, though I still couldn't come a bit near such a score of yours, I am also facing such an issue. I surely need some advice to come close around LB: 97~98!\n\nI observe to cases:\n&gt; E5: CV : 96.70 - LB: 96.5\nE7: CV : 96.35 - LB: 96.71\n\nIn addition, I also notice lb score improves if the validation score on grapheme_root improves promisingly.",
    "767347": "60 epochs. At least for CV, I don't see any increase in score for training longer.",
    "767441": "These are my own LB vs. CV scores:\n```\nLB / CV\n\n0.9874 / 0.9986\n0.9883 / 0.9984\n0.9879 / 0.9953\n0.9884 / 0.9967\n0.9876 / 0.9981\n0.9875 / 0.9968\n```",
    "767906": "Thank you. It is really taking a long time to train.",
    "768099": "ipythonx I guess all the details that are needed for LB .98 are stated in the discussions; no magic. No cropping helped me the most, got me from low .97 to .985+. The rest are just usual cutmix/mixup and usual training. Best of luck! Thank you for your generous sharing in this competition."
  },
  "source": "meta"
}