{
  "id": 153348,
  "title": "High CV score but low LB score",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/153348",
  "author_name": "Sumon Kanti Dey",
  "post_date": "2020-05-24T11:00:21.787000",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F872089%2Fea1c71b7888214d923f7d5f7a832de8b%2FScreenshot%20from%202020-05-24%2004-57-28.png?generation=1590318006426292&amp;alt=media\" alt=\"\"></p>\n\n<p>This was my CV score but it was produced 0.934 LB score. I have used almost balanced dataset. what will be a good idea to get solution of the issue.</p>",
  "messages": [
    {
      "id": 859698,
      "postDate": "2020-05-24T16:50:38.487Z",
      "content": "<p>Hi,  I think we all get better scores on the validation dataset. I see two main reasons for that. \nFirst, for the 3 languages used in the validation set (tr/es/it) you have a small training data, while for the 3 other language in the test set you do not have any training data (ru/pt/fr) . \nSecond, the Turkish validation dataset is somehow very easy, compared to the Italian and the Spanish, maybe thats also pulling up the score. \nBut in my experience improvements in CV score mostly translate to improvements in LB score here, so it is useful.</p>",
      "rawMarkdown": "Hi,  I think we all get better scores on the validation dataset. I see two main reasons for that. \nFirst, for the 3 languages used in the validation set (tr/es/it) you have a small training data, while for the 3 other language in the test set you do not have any training data (ru/pt/fr) . \nSecond, the Turkish validation dataset is somehow very easy, compared to the Italian and the Spanish, maybe thats also pulling up the score. \nBut in my experience improvements in CV score mostly translate to improvements in LB score here, so it is useful.",
      "votes": 7,
      "replies": [
        {
          "id": 860230,
          "postDate": "2020-05-25T06:16:48.557Z",
          "content": "<p>Thanks <a href=\"/riblidezso\">@riblidezso</a> . Yeah i have used translated dataset for training. I understand your point.</p>",
          "rawMarkdown": "Thanks @riblidezso . Yeah i have used translated dataset for training. I understand your point."
        }
      ]
    },
    {
      "id": 859356,
      "postDate": "2020-05-24T11:00:21.787Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F872089%2Fea1c71b7888214d923f7d5f7a832de8b%2FScreenshot%20from%202020-05-24%2004-57-28.png?generation=1590318006426292&amp;alt=media\" alt=\"\"></p>\n\n<p>This was my CV score but it was produced 0.934 LB score. I have used almost balanced dataset. what will be a good idea to get solution of the issue.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F872089%2Fea1c71b7888214d923f7d5f7a832de8b%2FScreenshot%20from%202020-05-24%2004-57-28.png?generation=1590318006426292&amp;alt=media)\n\nThis was my CV score but it was produced 0.934 LB score. I have used almost balanced dataset. what will be a good idea to get solution of the issue.",
      "votes": 3
    },
    {
      "id": 860286,
      "postDate": "2020-05-25T07:07:03.547Z",
      "content": "<p>seen from your training loss and val loss. you may overfit?</p>",
      "rawMarkdown": "seen from your training loss and val loss. you may overfit?",
      "votes": 1,
      "replies": [
        {
          "id": 867415,
          "postDate": "2020-05-30T09:07:56.847Z",
          "content": "<p>In this prior comment <a href=\"/riblidezso\">@riblidezso</a> I have used translated train along with existing train data and also trying to increase validation acc with the translated data . Instead of using the validation set only I have used translated 0.05% data for validation. Also facing the same problem. </p>\n\n<p>loss: 0.1207 - accuracy: 0.9554 - val_loss: 0.1584 - val_accuracy: 0.9468</p>\n\n<p>But my LB score also .92 which lower than validation accuracy. Am i missed something. Any hints?</p>",
          "rawMarkdown": "In this prior comment @riblidezso I have used translated train along with existing train data and also trying to increase validation acc with the translated data . Instead of using the validation set only I have used translated 0.05% data for validation. Also facing the same problem. \n\nloss: 0.1207 - accuracy: 0.9554 - val_loss: 0.1584 - val_accuracy: 0.9468\n\nBut my LB score also .92 which lower than validation accuracy. Am i missed something. Any hints?"
        },
        {
          "id": 867653,
          "postDate": "2020-05-30T13:41:22.467Z",
          "content": "<p>I would not use translated training data for validation. Translation quality is not so good. Those comments are very different from the real validation/test comments.</p>\n\n<p>In my experience the 8000 validation sample is large enough to get stable validation results. For me, models which work better on that set are almost always better on the LB too. If you also train on validation data then you will need a cross validation scheme.</p>\n\n<p>Of course the validation score and the leaderboard score will not be the same, the validation will be higher (probably mainly due to the 2 reasons I mentioned above). But as long as the order of submissions is the same on your validation and the LB then the validation score is useful.</p>\n\n<p>Anyway given the 30hour TPU limit you can check almost all your different solutions on the leaderboard. In this competition the public leaderboard is probably the most trustworthy score.</p>",
          "rawMarkdown": "I would not use translated training data for validation. Translation quality is not so good. Those comments are very different from the real validation/test comments.\n\nIn my experience the 8000 validation sample is large enough to get stable validation results. For me, models which work better on that set are almost always better on the LB too. If you also train on validation data then you will need a cross validation scheme.\n\nOf course the validation score and the leaderboard score will not be the same, the validation will be higher (probably mainly due to the 2 reasons I mentioned above). But as long as the order of submissions is the same on your validation and the LB then the validation score is useful.\n\nAnyway given the 30hour TPU limit you can check almost all your different solutions on the leaderboard. In this competition the public leaderboard is probably the most trustworthy score.",
          "votes": 2
        },
        {
          "id": 868661,
          "postDate": "2020-05-31T12:08:37.680Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 868713,
          "postDate": "2020-05-31T12:35:34.977Z",
          "content": "<p>Thanks <a href=\"/riblidezso\">@riblidezso</a> </p>",
          "rawMarkdown": "Thanks @riblidezso \n"
        },
        {
          "id": 868763,
          "postDate": "2020-05-31T13:20:19.317Z",
          "content": "<p>A general rule of thumbs for me :  Data Augmentation is generally not a good idea for Validation data....And  given Translation has not really a good quality here. </p>",
          "rawMarkdown": "A general rule of thumbs for me :  Data Augmentation is generally not a good idea for Validation data....And  given Translation has not really a good quality here. "
        }
      ]
    },
    {
      "id": 868412,
      "postDate": "2020-05-31T07:49:56.637Z",
      "content": "<p>You'd expect close LB and CV scores ONLY when validation and test data come from identical distribution. So your gap depends on what your validation set is. I read from below it's competition validation set + some translated data - that's somewhat different than test set which has 6 foreign languages and no translation. So the gap which you are seeing may be reasonable. But if your validation data is still somewhat similar, at the very least, you should have your LB and CV moving in similar direction for any significant change. </p>",
      "rawMarkdown": "You'd expect close LB and CV scores ONLY when validation and test data come from identical distribution. So your gap depends on what your validation set is. I read from below it's competition validation set + some translated data - that's somewhat different than test set which has 6 foreign languages and no translation. So the gap which you are seeing may be reasonable. But if your validation data is still somewhat similar, at the very least, you should have your LB and CV moving in similar direction for any significant change. ",
      "replies": [
        {
          "id": 868717,
          "postDate": "2020-05-31T12:35:55.023Z",
          "content": "<p>Thanks <a href=\"/moizsaifee\">@moizsaifee</a> </p>",
          "rawMarkdown": "Thanks @moizsaifee "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 859698,
      "author_name": "Dezső Ribli",
      "author_url": "",
      "post_date": "2020-05-24T16:50:38.487000",
      "content": "<p>Hi,  I think we all get better scores on the validation dataset. I see two main reasons for that. \nFirst, for the 3 languages used in the validation set (tr/es/it) you have a small training data, while for the 3 other language in the test set you do not have any training data (ru/pt/fr) . \nSecond, the Turkish validation dataset is somehow very easy, compared to the Italian and the Spanish, maybe thats also pulling up the score. \nBut in my experience improvements in CV score mostly translate to improvements in LB score here, so it is useful.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 860230,
          "author_name": "Sumon Kanti Dey",
          "author_url": "",
          "post_date": "2020-05-25T06:16:48.557000",
          "content": "<p>Thanks <a href=\"/riblidezso\">@riblidezso</a> . Yeah i have used translated dataset for training. I understand your point.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 860286,
      "author_name": "MaChaogong",
      "author_url": "",
      "post_date": "2020-05-25T07:07:03.547000",
      "content": "<p>seen from your training loss and val loss. you may overfit?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 867415,
          "author_name": "Sumon Kanti Dey",
          "author_url": "",
          "post_date": "2020-05-30T09:07:56.847000",
          "content": "<p>In this prior comment <a href=\"/riblidezso\">@riblidezso</a> I have used translated train along with existing train data and also trying to increase validation acc with the translated data . Instead of using the validation set only I have used translated 0.05% data for validation. Also facing the same problem. </p>\n\n<p>loss: 0.1207 - accuracy: 0.9554 - val_loss: 0.1584 - val_accuracy: 0.9468</p>\n\n<p>But my LB score also .92 which lower than validation accuracy. Am i missed something. Any hints?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 867653,
          "author_name": "Dezső Ribli",
          "author_url": "",
          "post_date": "2020-05-30T13:41:22.467000",
          "content": "<p>I would not use translated training data for validation. Translation quality is not so good. Those comments are very different from the real validation/test comments.</p>\n\n<p>In my experience the 8000 validation sample is large enough to get stable validation results. For me, models which work better on that set are almost always better on the LB too. If you also train on validation data then you will need a cross validation scheme.</p>\n\n<p>Of course the validation score and the leaderboard score will not be the same, the validation will be higher (probably mainly due to the 2 reasons I mentioned above). But as long as the order of submissions is the same on your validation and the LB then the validation score is useful.</p>\n\n<p>Anyway given the 30hour TPU limit you can check almost all your different solutions on the leaderboard. In this competition the public leaderboard is probably the most trustworthy score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 868661,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-31T12:08:37.680000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 868713,
          "author_name": "Sumon Kanti Dey",
          "author_url": "",
          "post_date": "2020-05-31T12:35:34.977000",
          "content": "<p>Thanks <a href=\"/riblidezso\">@riblidezso</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 868763,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-05-31T13:20:19.317000",
          "content": "<p>A general rule of thumbs for me :  Data Augmentation is generally not a good idea for Validation data....And  given Translation has not really a good quality here. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 868412,
      "author_name": "Moiz Saifee",
      "author_url": "",
      "post_date": "2020-05-31T07:49:56.637000",
      "content": "<p>You'd expect close LB and CV scores ONLY when validation and test data come from identical distribution. So your gap depends on what your validation set is. I read from below it's competition validation set + some translated data - that's somewhat different than test set which has 6 foreign languages and no translation. So the gap which you are seeing may be reasonable. But if your validation data is still somewhat similar, at the very least, you should have your LB and CV moving in similar direction for any significant change. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 868717,
          "author_name": "Sumon Kanti Dey",
          "author_url": "",
          "post_date": "2020-05-31T12:35:55.023000",
          "content": "<p>Thanks <a href=\"/moizsaifee\">@moizsaifee</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "859698": "Hi,  I think we all get better scores on the validation dataset. I see two main reasons for that. \nFirst, for the 3 languages used in the validation set (tr/es/it) you have a small training data, while for the 3 other language in the test set you do not have any training data (ru/pt/fr) . \nSecond, the Turkish validation dataset is somehow very easy, compared to the Italian and the Spanish, maybe thats also pulling up the score. \nBut in my experience improvements in CV score mostly translate to improvements in LB score here, so it is useful.",
    "859356": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F872089%2Fea1c71b7888214d923f7d5f7a832de8b%2FScreenshot%20from%202020-05-24%2004-57-28.png?generation=1590318006426292&amp;alt=media)\n\nThis was my CV score but it was produced 0.934 LB score. I have used almost balanced dataset. what will be a good idea to get solution of the issue.",
    "860286": "seen from your training loss and val loss. you may overfit?",
    "868412": "You'd expect close LB and CV scores ONLY when validation and test data come from identical distribution. So your gap depends on what your validation set is. I read from below it's competition validation set + some translated data - that's somewhat different than test set which has 6 foreign languages and no translation. So the gap which you are seeing may be reasonable. But if your validation data is still somewhat similar, at the very least, you should have your LB and CV moving in similar direction for any significant change. "
  }
}