{
  "id": 128739,
  "title": "Large LB/CV difference",
  "url": "/competitions/bengaliai-cv19/discussion/128739",
  "author_name": "",
  "post_date": "2020-02-02T22:50:57.574265600Z",
  "votes": 2,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Hi!</p>\n\n<p>I'm using this kernel: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a></p>\n\n<p>When I trained without changing anything, I trained for 89 epochs but the best CV epoch was 20th, and I got CV 0.967 but LB only 0.9515. After using stepLR, I trained for 94 epoch, and CV converged at   59 epoch with CV score of 0.971; however, LB score is still as low as 0.9548. Both of those CV and LB differences are much larger than what other people have in the forum. I was wondering what might be the problem.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "735349",
      "postDate": "02/02/2020 22:50:57",
      "content": "<p>Hi!</p>\n\n<p>I'm using this kernel: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a></p>\n\n<p>When I trained without changing anything, I trained for 89 epochs but the best CV epoch was 20th, and I got CV 0.967 but LB only 0.9515. After using stepLR, I trained for 94 epoch, and CV converged at   59 epoch with CV score of 0.971; however, LB score is still as low as 0.9548. Both of those CV and LB differences are much larger than what other people have in the forum. I was wondering what might be the problem.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi!\n\nI'm using this kernel: https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\n\nWhen I trained without changing anything, I trained for 89 epochs but the best CV epoch was 20th, and I got CV 0.967 but LB only 0.9515. After using stepLR, I trained for 94 epoch, and CV converged at   59 epoch with CV score of 0.971; however, LB score is still as low as 0.9548. Both of those CV and LB differences are much larger than what other people have in the forum. I was wondering what might be the problem.\n\nThanks!",
      "votes": null
    },
    {
      "id": "735426",
      "postDate": "02/03/2020 02:46:53",
      "content": "<p>How did you split training set and validation set?\nMaybe this notebook could give you some inspiration:\n<a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">https://www.kaggle.com/yiheng/iterative-stratification</a></p>",
      "rawMarkdown": "How did you split training set and validation set?\nMaybe this notebook could give you some inspiration:\nhttps://www.kaggle.com/yiheng/iterative-stratification",
      "votes": null
    },
    {
      "id": "735440",
      "postDate": "02/03/2020 03:15:18",
      "content": "<p>I had 17% of val set. <a href=\"/haqishen\">@haqishen</a> </p>",
      "rawMarkdown": "I had 17% of val set. @haqishen",
      "votes": null
    },
    {
      "id": "736245",
      "postDate": "02/04/2020 01:38:39",
      "content": "<p>Now I have CV 0.9868 but LB only 0.9646. It is really confusing because other people's posts usually result in ~0.97 LB ~0.97-0.98 LB. I was wondering what might be the mistake. Thanks!</p>",
      "rawMarkdown": "Now I have CV 0.9868 but LB only 0.9646. It is really confusing because other people's posts usually result in ~0.97 LB ~0.97-0.98 LB. I was wondering what might be the mistake. Thanks!",
      "votes": null
    },
    {
      "id": "736247",
      "postDate": "02/04/2020 01:40:07",
      "content": "<p>Here is the macro_recall script that I used from  <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a> and I have 0.17 val split.</p>\n\n<p>And this is the predicton kernel that i am using without any change: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch</a></p>\n\n<p>```\ndef macro_recall(pred_y, y, n_grapheme=168, n_vowel=11, n_consonant=7):\n    pred_y = torch.split(pred_y, [n_grapheme, n_vowel, n_consonant], dim=1)\n    pred_labels = [torch.argmax(py, dim=1).cpu().numpy() for py in pred_y]</p>\n\n<pre><code>y = y.cpu().numpy()\n# pred_y = [p.cpu().numpy() for p in pred_y]\n\nrecall_grapheme = sklearn.metrics.recall_score(pred_labels[0], y[:, 0], average='macro')\nrecall_vowel = sklearn.metrics.recall_score(pred_labels[1], y[:, 1], average='macro')\nrecall_consonant = sklearn.metrics.recall_score(pred_labels[2], y[:, 2], average='macro')\nscores = [recall_grapheme, recall_vowel, recall_consonant]\nfinal_score = np.average(scores, weights=[2, 1, 1])\n# print(f'recall: grapheme {recall_grapheme}, vowel {recall_vowel}, consonant {recall_consonant}, '\n#       f'total {final_score}, y {y.shape}')\nreturn final_score\n</code></pre>\n\n<p>def calc_macro_recall(solution, submission):\n    # solution df, submission df\n    scores = []\n    for component in ['grapheme_root', 'consonant_diacritic', 'vowel_diacritic']:\n        y_true_subset = solution[solution[component] == component]['target'].values\n        y_pred_subset = submission[submission[component] == component]['target'].values\n        scores.append(sklearn.metrics.recall_score(\n            y_true_subset, y_pred_subset, average='macro'))\n    final_score = np.average(scores, weights=[2, 1, 1])\n    return final_score</p>\n\n<p>```</p>\n\n<p>and here is my val set generation code:</p>\n\n<p>```\nn_dataset = len(train_images)\ntrain_data_size = 200 if debug else int(n_dataset * 0.83)\nvalid_data_size = 100 if debug else int(n_dataset - train_data_size)</p>\n\n<p>perm = np.random.RandomState(777).permutation(n_dataset)\nprint('perm', perm)\ntrain_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(size=(image_size, image_size)),\n    indices=perm[:train_data_size], valid = False)\nvalid_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(affine=False, crop=True, size=(image_size, image_size)),\n    indices=perm[train_data_size:train_data_size+valid_data_size], valid = True)\nprint('train_dataset', len(train_dataset), 'valid_dataset', len(valid_dataset))\n```</p>",
      "rawMarkdown": "Here is the macro_recall script that I used from  https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch and I have 0.17 val split.\n\nAnd this is the predicton kernel that i am using without any change: https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch\n\n\n```\ndef macro_recall(pred_y, y, n_grapheme=168, n_vowel=11, n_consonant=7):\n    pred_y = torch.split(pred_y, [n_grapheme, n_vowel, n_consonant], dim=1)\n    pred_labels = [torch.argmax(py, dim=1).cpu().numpy() for py in pred_y]\n\n    y = y.cpu().numpy()\n    # pred_y = [p.cpu().numpy() for p in pred_y]\n\n    recall_grapheme = sklearn.metrics.recall_score(pred_labels[0], y[:, 0], average='macro')\n    recall_vowel = sklearn.metrics.recall_score(pred_labels[1], y[:, 1], average='macro')\n    recall_consonant = sklearn.metrics.recall_score(pred_labels[2], y[:, 2], average='macro')\n    scores = [recall_grapheme, recall_vowel, recall_consonant]\n    final_score = np.average(scores, weights=[2, 1, 1])\n    # print(f'recall: grapheme {recall_grapheme}, vowel {recall_vowel}, consonant {recall_consonant}, '\n    #       f'total {final_score}, y {y.shape}')\n    return final_score\n\n\ndef calc_macro_recall(solution, submission):\n    # solution df, submission df\n    scores = []\n    for component in ['grapheme_root', 'consonant_diacritic', 'vowel_diacritic']:\n        y_true_subset = solution[solution[component] == component]['target'].values\n        y_pred_subset = submission[submission[component] == component]['target'].values\n        scores.append(sklearn.metrics.recall_score(\n            y_true_subset, y_pred_subset, average='macro'))\n    final_score = np.average(scores, weights=[2, 1, 1])\n    return final_score\n\n```\n\nand here is my val set generation code:\n\n```\nn_dataset = len(train_images)\ntrain_data_size = 200 if debug else int(n_dataset * 0.83)\nvalid_data_size = 100 if debug else int(n_dataset - train_data_size)\n\nperm = np.random.RandomState(777).permutation(n_dataset)\nprint('perm', perm)\ntrain_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(size=(image_size, image_size)),\n    indices=perm[:train_data_size], valid = False)\nvalid_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(affine=False, crop=True, size=(image_size, image_size)),\n    indices=perm[train_data_size:train_data_size+valid_data_size], valid = True)\nprint('train_dataset', len(train_dataset), 'valid_dataset', len(valid_dataset))\n```",
      "votes": null
    },
    {
      "id": "736275",
      "postDate": "02/04/2020 02:29:02",
      "content": "<p>do NOT purely random split the data\nwhy not just give a read to the notebook i attached above</p>",
      "rawMarkdown": "do NOT purely random split the data\nwhy not just give a read to the notebook i attached above",
      "votes": null
    },
    {
      "id": "736286",
      "postDate": "02/04/2020 02:52:11",
      "content": "<p>Thank you! I'll take a look! <a href=\"/haqishen\">@haqishen</a> </p>",
      "rawMarkdown": "Thank you! I'll take a look! @haqishen",
      "votes": null
    },
    {
      "id": "736288",
      "postDate": "02/04/2020 02:55:42",
      "content": "<p>May I also ask why randomly splitting data is not good is this competition? <a href=\"/haqishen\">@haqishen</a>  Thanks!</p>",
      "rawMarkdown": "May I also ask why randomly splitting data is not good is this competition? @haqishen  Thanks!",
      "votes": null
    },
    {
      "id": "737092",
      "postDate": "02/04/2020 22:53:30",
      "content": "<p>You can check this out! <a href=\"http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf\">http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf</a> \n“ random distribution of multi-label training examples into sub- sets suffers from the following practical problem: it can lead to test subsets lacking even just one positive example of a rare label, which in turn causes cal- culation problems for a number of multi-label evaluation measures. The typical way these problems get by-passed in the literature is through complete removal of rare labels. This, however, implies that the performance of the learning sys- tems on rare labels is unimportant, which is seldom true. As an example consider that a multi-label learner is used for probabilistic indexing of a large multimedia collection, given a small annotated sample according to a multimedia ontology. Avoiding the evaluation of the multi-label learner for rare concepts of the on- tology, implies that we should not allow users to query the collection with such concepts, as the information retrieval performance level of the indexing system for these concepts would be uncertain. This limits the usefulness of the indexing system.”</p>",
      "rawMarkdown": "You can check this out! http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf \n“ random distribution of multi-label training examples into sub- sets suffers from the following practical problem: it can lead to test subsets lacking even just one positive example of a rare label, which in turn causes cal- culation problems for a number of multi-label evaluation measures. The typical way these problems get by-passed in the literature is through complete removal of rare labels. This, however, implies that the performance of the learning sys- tems on rare labels is unimportant, which is seldom true. As an example consider that a multi-label learner is used for probabilistic indexing of a large multimedia collection, given a small annotated sample according to a multimedia ontology. Avoiding the evaluation of the multi-label learner for rare concepts of the on- tology, implies that we should not allow users to query the collection with such concepts, as the information retrieval performance level of the indexing system for these concepts would be uncertain. This limits the usefulness of the indexing system.”",
      "votes": null
    },
    {
      "id": "737258",
      "postDate": "02/05/2020 05:27:50",
      "content": "<p>Hi, <a href=\"/tonychenxyz\">@tonychenxyz</a> </p>\n\n<p>How many epochs do you train to reach this <code>CV scores 0.9868</code> ? </p>",
      "rawMarkdown": "Hi, @tonychenxyz \n\nHow many epochs do you train to reach this `CV scores 0.9868` ?",
      "votes": null
    },
    {
      "id": "737269",
      "postDate": "02/05/2020 05:36:56",
      "content": "<p>102 <a href=\"/moximo13\">@moximo13</a> </p>",
      "rawMarkdown": "102 @moximo13",
      "votes": null
    },
    {
      "id": "739627",
      "postDate": "02/08/2020 04:38:46",
      "content": "<p>Update: now I am using stratified data split, but the lb/cv difference is still similarly large, around 0.013-0.02, so it's probably not the data split that caused the problem. I was wondering what might be other problems? <a href=\"/haqishen\">@haqishen</a> </p>",
      "rawMarkdown": "Update: now I am using stratified data split, but the lb/cv difference is still similarly large, around 0.013-0.02, so it's probably not the data split that caused the problem. I was wondering what might be other problems? @haqishen",
      "votes": null
    },
    {
      "id": "739866",
      "postDate": "02/08/2020 13:53:05",
      "content": "<p>I'm also having a cv/lb difference around 0.013. I'm using iterative stratification with 15% val set and wondering why they can have such little gap. :(</p>",
      "rawMarkdown": "I'm also having a cv/lb difference around 0.013. I'm using iterative stratification with 15% val set and wondering why they can have such little gap. :(",
      "votes": null
    },
    {
      "id": "739925",
      "postDate": "02/08/2020 15:05:25",
      "content": "<p>It’s possible be problem in inference code. Please check whether you can get a same probability for the same input when inferences.</p>",
      "rawMarkdown": "It’s possible be problem in inference code. Please check whether you can get a same probability for the same input when inferences.",
      "votes": null
    },
    {
      "id": "740069",
      "postDate": "02/08/2020 19:51:16",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> By probability, do you mean lb score or the output of model? Thanks!</p>",
      "rawMarkdown": "haqishen By probability, do you mean lb score or the output of model? Thanks!",
      "votes": null
    },
    {
      "id": "740198",
      "postDate": "02/09/2020 02:42:12",
      "content": "<p>I mean the raw output (logits) of the model</p>",
      "rawMarkdown": "I mean the raw output (logits) of the model",
      "votes": null
    },
    {
      "id": "741776",
      "postDate": "02/11/2020 00:09:14",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> I just checked. The outputs are the same.</p>",
      "rawMarkdown": "haqishen I just checked. The outputs are the same.",
      "votes": null
    },
    {
      "id": "741860",
      "postDate": "02/11/2020 01:15:16",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> Also, my maximum gap now becomes Cv 0.9886, Lb 0.9679. I was wondering what else could I try. Thank you very much!</p>",
      "rawMarkdown": "haqishen Also, my maximum gap now becomes Cv 0.9886, Lb 0.9679. I was wondering what else could I try. Thank you very much!",
      "votes": null
    },
    {
      "id": "741972",
      "postDate": "02/11/2020 02:47:40",
      "content": "<p>Another possible problem is that you use lower resolution as input. A lower resolution you use, a larger gap you get</p>",
      "rawMarkdown": "Another possible problem is that you use lower resolution as input. A lower resolution you use, a larger gap you get",
      "votes": null
    },
    {
      "id": "742002",
      "postDate": "02/11/2020 03:09:59",
      "content": "<p>I'm using 128x128. I'll try a larger size. Thank you! Also, do you think the CNN and linear tail layers would be an issue too? <a href=\"/haqishen\">@haqishen</a> Thanks!</p>",
      "rawMarkdown": "I'm using 128x128. I'll try a larger size. Thank you! Also, do you think the CNN and linear tail layers would be an issue too? @haqishen Thanks!",
      "votes": null
    },
    {
      "id": "742112",
      "postDate": "02/11/2020 04:41:52",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Maybe you can try another random split seed. I changed different seeds and the cv/lb gap can be made smaller but the lb scores are quite similar.</p>",
      "rawMarkdown": "tonychenxyz Maybe you can try another random split seed. I changed different seeds and the cv/lb gap can be made smaller but the lb scores are quite similar.",
      "votes": null
    },
    {
      "id": "742118",
      "postDate": "02/11/2020 04:48:50",
      "content": "<p><a href=\"/syoya1997\">@syoya1997</a> what is your cv/lb gap now?</p>",
      "rawMarkdown": "syoya1997 what is your cv/lb gap now?",
      "votes": null
    },
    {
      "id": "742126",
      "postDate": "02/11/2020 04:55:53",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Around 0.009 right now. Smaller than my previous gap 0.013. But lb score increases only by 0.001 because my cv scores went down with a different seed.</p>",
      "rawMarkdown": "tonychenxyz Around 0.009 right now. Smaller than my previous gap 0.013. But lb score increases only by 0.001 because my cv scores went down with a different seed.",
      "votes": null
    },
    {
      "id": "742141",
      "postDate": "02/11/2020 05:08:43",
      "content": "<p>Thanks for the information! I'll give another try. <a href=\"/syoya1997\">@syoya1997</a> </p>",
      "rawMarkdown": "Thanks for the information! I'll give another try. @syoya1997",
      "votes": null
    },
    {
      "id": "753967",
      "postDate": "02/22/2020 21:47:48",
      "content": "<p>I tried iterative stratification and it seems like I still have that gap, 98%+ val and 96.5% lb.. I'm not sure whats causing this please help!!</p>\n\n<p>Only way to reduce that is by adding a weight tensor in loss function</p>",
      "rawMarkdown": "I tried iterative stratification and it seems like I still have that gap, 98%+ val and 96.5% lb.. I'm not sure whats causing this please help!!\n\nOnly way to reduce that is by adding a weight tensor in loss function",
      "votes": null
    },
    {
      "id": "755389",
      "postDate": "02/24/2020 18:27:00",
      "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Hi! I was wondering if you have solved your problem. Also, what exactly do you mean by \"Only way to reduce that is by adding a weight tensor in loss function\"? In addition, are you using the kernel that I referred to (<a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a>)? Thank you very much!</p>",
      "rawMarkdown": "yannmajewski Hi! I was wondering if you have solved your problem. Also, what exactly do you mean by \"Only way to reduce that is by adding a weight tensor in loss function\"? In addition, are you using the kernel that I referred to (https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch)? Thank you very much!",
      "votes": null
    },
    {
      "id": "766440",
      "postDate": "03/08/2020 06:09:59",
      "content": "<p>I used a random split of 80-20. My CV score is ~ 0.98 but my LB score is 0.71. What could be the reason?\nI did not use any data augmentation</p>",
      "rawMarkdown": "I used a random split of 80-20. My CV score is ~ 0.98 but my LB score is 0.71. What could be the reason?\nI did not use any data augmentation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 735426,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "02/03/2020 02:46:53",
      "content": "<p>How did you split training set and validation set?\nMaybe this notebook could give you some inspiration:\n<a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">https://www.kaggle.com/yiheng/iterative-stratification</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 735440,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/03/2020 03:15:18",
          "content": "<p>I had 17% of val set. <a href=\"/haqishen\">@haqishen</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 736245,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "02/04/2020 01:38:39",
      "content": "<p>Now I have CV 0.9868 but LB only 0.9646. It is really confusing because other people's posts usually result in ~0.97 LB ~0.97-0.98 LB. I was wondering what might be the mistake. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 736275,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/04/2020 02:29:02",
          "content": "<p>do NOT purely random split the data\nwhy not just give a read to the notebook i attached above</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 736286,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/04/2020 02:52:11",
          "content": "<p>Thank you! I'll take a look! <a href=\"/haqishen\">@haqishen</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 736288,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/04/2020 02:55:42",
          "content": "<p>May I also ask why randomly splitting data is not good is this competition? <a href=\"/haqishen\">@haqishen</a>  Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737092,
          "author_name": "yuanlin08",
          "author_url": "",
          "post_date": "02/04/2020 22:53:30",
          "content": "<p>You can check this out! <a href=\"http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf\">http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf</a> \n“ random distribution of multi-label training examples into sub- sets suffers from the following practical problem: it can lead to test subsets lacking even just one positive example of a rare label, which in turn causes cal- culation problems for a number of multi-label evaluation measures. The typical way these problems get by-passed in the literature is through complete removal of rare labels. This, however, implies that the performance of the learning sys- tems on rare labels is unimportant, which is seldom true. As an example consider that a multi-label learner is used for probabilistic indexing of a large multimedia collection, given a small annotated sample according to a multimedia ontology. Avoiding the evaluation of the multi-label learner for rare concepts of the on- tology, implies that we should not allow users to query the collection with such concepts, as the information retrieval performance level of the indexing system for these concepts would be uncertain. This limits the usefulness of the indexing system.”</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737258,
          "author_name": "moximo13",
          "author_url": "",
          "post_date": "02/05/2020 05:27:50",
          "content": "<p>Hi, <a href=\"/tonychenxyz\">@tonychenxyz</a> </p>\n\n<p>How many epochs do you train to reach this <code>CV scores 0.9868</code> ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737269,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/05/2020 05:36:56",
          "content": "<p>102 <a href=\"/moximo13\">@moximo13</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 742112,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/11/2020 04:41:52",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Maybe you can try another random split seed. I changed different seeds and the cv/lb gap can be made smaller but the lb scores are quite similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 742118,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/11/2020 04:48:50",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a> what is your cv/lb gap now?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 742126,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/11/2020 04:55:53",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Around 0.009 right now. Smaller than my previous gap 0.013. But lb score increases only by 0.001 because my cv scores went down with a different seed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 742141,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/11/2020 05:08:43",
          "content": "<p>Thanks for the information! I'll give another try. <a href=\"/syoya1997\">@syoya1997</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 736247,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "02/04/2020 01:40:07",
      "content": "<p>Here is the macro_recall script that I used from  <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a> and I have 0.17 val split.</p>\n\n<p>And this is the predicton kernel that i am using without any change: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch</a></p>\n\n<p>```\ndef macro_recall(pred_y, y, n_grapheme=168, n_vowel=11, n_consonant=7):\n    pred_y = torch.split(pred_y, [n_grapheme, n_vowel, n_consonant], dim=1)\n    pred_labels = [torch.argmax(py, dim=1).cpu().numpy() for py in pred_y]</p>\n\n<pre><code>y = y.cpu().numpy()\n# pred_y = [p.cpu().numpy() for p in pred_y]\n\nrecall_grapheme = sklearn.metrics.recall_score(pred_labels[0], y[:, 0], average='macro')\nrecall_vowel = sklearn.metrics.recall_score(pred_labels[1], y[:, 1], average='macro')\nrecall_consonant = sklearn.metrics.recall_score(pred_labels[2], y[:, 2], average='macro')\nscores = [recall_grapheme, recall_vowel, recall_consonant]\nfinal_score = np.average(scores, weights=[2, 1, 1])\n# print(f'recall: grapheme {recall_grapheme}, vowel {recall_vowel}, consonant {recall_consonant}, '\n#       f'total {final_score}, y {y.shape}')\nreturn final_score\n</code></pre>\n\n<p>def calc_macro_recall(solution, submission):\n    # solution df, submission df\n    scores = []\n    for component in ['grapheme_root', 'consonant_diacritic', 'vowel_diacritic']:\n        y_true_subset = solution[solution[component] == component]['target'].values\n        y_pred_subset = submission[submission[component] == component]['target'].values\n        scores.append(sklearn.metrics.recall_score(\n            y_true_subset, y_pred_subset, average='macro'))\n    final_score = np.average(scores, weights=[2, 1, 1])\n    return final_score</p>\n\n<p>```</p>\n\n<p>and here is my val set generation code:</p>\n\n<p>```\nn_dataset = len(train_images)\ntrain_data_size = 200 if debug else int(n_dataset * 0.83)\nvalid_data_size = 100 if debug else int(n_dataset - train_data_size)</p>\n\n<p>perm = np.random.RandomState(777).permutation(n_dataset)\nprint('perm', perm)\ntrain_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(size=(image_size, image_size)),\n    indices=perm[:train_data_size], valid = False)\nvalid_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(affine=False, crop=True, size=(image_size, image_size)),\n    indices=perm[train_data_size:train_data_size+valid_data_size], valid = True)\nprint('train_dataset', len(train_dataset), 'valid_dataset', len(valid_dataset))\n```</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 739627,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "02/08/2020 04:38:46",
      "content": "<p>Update: now I am using stratified data split, but the lb/cv difference is still similarly large, around 0.013-0.02, so it's probably not the data split that caused the problem. I was wondering what might be other problems? <a href=\"/haqishen\">@haqishen</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 739925,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/08/2020 15:05:25",
          "content": "<p>It’s possible be problem in inference code. Please check whether you can get a same probability for the same input when inferences.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740069,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/08/2020 19:51:16",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> By probability, do you mean lb score or the output of model? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740198,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/09/2020 02:42:12",
          "content": "<p>I mean the raw output (logits) of the model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741776,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/11/2020 00:09:14",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> I just checked. The outputs are the same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741860,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/11/2020 01:15:16",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> Also, my maximum gap now becomes Cv 0.9886, Lb 0.9679. I was wondering what else could I try. Thank you very much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741972,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/11/2020 02:47:40",
          "content": "<p>Another possible problem is that you use lower resolution as input. A lower resolution you use, a larger gap you get</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 742002,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/11/2020 03:09:59",
          "content": "<p>I'm using 128x128. I'll try a larger size. Thank you! Also, do you think the CNN and linear tail layers would be an issue too? <a href=\"/haqishen\">@haqishen</a> Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 753967,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "02/22/2020 21:47:48",
          "content": "<p>I tried iterative stratification and it seems like I still have that gap, 98%+ val and 96.5% lb.. I'm not sure whats causing this please help!!</p>\n\n<p>Only way to reduce that is by adding a weight tensor in loss function</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755389,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/24/2020 18:27:00",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Hi! I was wondering if you have solved your problem. Also, what exactly do you mean by \"Only way to reduce that is by adding a weight tensor in loss function\"? In addition, are you using the kernel that I referred to (<a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a>)? Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 739866,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "02/08/2020 13:53:05",
      "content": "<p>I'm also having a cv/lb difference around 0.013. I'm using iterative stratification with 15% val set and wondering why they can have such little gap. :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 766440,
      "author_name": "venky2506",
      "author_url": "",
      "post_date": "03/08/2020 06:09:59",
      "content": "<p>I used a random split of 80-20. My CV score is ~ 0.98 but my LB score is 0.71. What could be the reason?\nI did not use any data augmentation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "735349": "Hi!\n\nI'm using this kernel: https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\n\nWhen I trained without changing anything, I trained for 89 epochs but the best CV epoch was 20th, and I got CV 0.967 but LB only 0.9515. After using stepLR, I trained for 94 epoch, and CV converged at   59 epoch with CV score of 0.971; however, LB score is still as low as 0.9548. Both of those CV and LB differences are much larger than what other people have in the forum. I was wondering what might be the problem.\n\nThanks!",
    "735426": "How did you split training set and validation set?\nMaybe this notebook could give you some inspiration:\nhttps://www.kaggle.com/yiheng/iterative-stratification",
    "735440": "I had 17% of val set. @haqishen",
    "736245": "Now I have CV 0.9868 but LB only 0.9646. It is really confusing because other people's posts usually result in ~0.97 LB ~0.97-0.98 LB. I was wondering what might be the mistake. Thanks!",
    "736247": "Here is the macro_recall script that I used from  https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch and I have 0.17 val split.\n\nAnd this is the predicton kernel that i am using without any change: https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch\n\n\n```\ndef macro_recall(pred_y, y, n_grapheme=168, n_vowel=11, n_consonant=7):\n    pred_y = torch.split(pred_y, [n_grapheme, n_vowel, n_consonant], dim=1)\n    pred_labels = [torch.argmax(py, dim=1).cpu().numpy() for py in pred_y]\n\n    y = y.cpu().numpy()\n    # pred_y = [p.cpu().numpy() for p in pred_y]\n\n    recall_grapheme = sklearn.metrics.recall_score(pred_labels[0], y[:, 0], average='macro')\n    recall_vowel = sklearn.metrics.recall_score(pred_labels[1], y[:, 1], average='macro')\n    recall_consonant = sklearn.metrics.recall_score(pred_labels[2], y[:, 2], average='macro')\n    scores = [recall_grapheme, recall_vowel, recall_consonant]\n    final_score = np.average(scores, weights=[2, 1, 1])\n    # print(f'recall: grapheme {recall_grapheme}, vowel {recall_vowel}, consonant {recall_consonant}, '\n    #       f'total {final_score}, y {y.shape}')\n    return final_score\n\n\ndef calc_macro_recall(solution, submission):\n    # solution df, submission df\n    scores = []\n    for component in ['grapheme_root', 'consonant_diacritic', 'vowel_diacritic']:\n        y_true_subset = solution[solution[component] == component]['target'].values\n        y_pred_subset = submission[submission[component] == component]['target'].values\n        scores.append(sklearn.metrics.recall_score(\n            y_true_subset, y_pred_subset, average='macro'))\n    final_score = np.average(scores, weights=[2, 1, 1])\n    return final_score\n\n```\n\nand here is my val set generation code:\n\n```\nn_dataset = len(train_images)\ntrain_data_size = 200 if debug else int(n_dataset * 0.83)\nvalid_data_size = 100 if debug else int(n_dataset - train_data_size)\n\nperm = np.random.RandomState(777).permutation(n_dataset)\nprint('perm', perm)\ntrain_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(size=(image_size, image_size)),\n    indices=perm[:train_data_size], valid = False)\nvalid_dataset = BengaliAIDataset(\n    train_images, train_labels, transform=Transform(affine=False, crop=True, size=(image_size, image_size)),\n    indices=perm[train_data_size:train_data_size+valid_data_size], valid = True)\nprint('train_dataset', len(train_dataset), 'valid_dataset', len(valid_dataset))\n```",
    "736275": "do NOT purely random split the data\nwhy not just give a read to the notebook i attached above",
    "736286": "Thank you! I'll take a look! @haqishen",
    "736288": "May I also ask why randomly splitting data is not good is this competition? @haqishen  Thanks!",
    "737092": "You can check this out! http://lpis.csd.auth.gr/publications/sechidis-ecmlpkdd-2011.pdf \n“ random distribution of multi-label training examples into sub- sets suffers from the following practical problem: it can lead to test subsets lacking even just one positive example of a rare label, which in turn causes cal- culation problems for a number of multi-label evaluation measures. The typical way these problems get by-passed in the literature is through complete removal of rare labels. This, however, implies that the performance of the learning sys- tems on rare labels is unimportant, which is seldom true. As an example consider that a multi-label learner is used for probabilistic indexing of a large multimedia collection, given a small annotated sample according to a multimedia ontology. Avoiding the evaluation of the multi-label learner for rare concepts of the on- tology, implies that we should not allow users to query the collection with such concepts, as the information retrieval performance level of the indexing system for these concepts would be uncertain. This limits the usefulness of the indexing system.”",
    "737258": "Hi, @tonychenxyz \n\nHow many epochs do you train to reach this `CV scores 0.9868` ?",
    "737269": "102 @moximo13",
    "739627": "Update: now I am using stratified data split, but the lb/cv difference is still similarly large, around 0.013-0.02, so it's probably not the data split that caused the problem. I was wondering what might be other problems? @haqishen",
    "739866": "I'm also having a cv/lb difference around 0.013. I'm using iterative stratification with 15% val set and wondering why they can have such little gap. :(",
    "739925": "It’s possible be problem in inference code. Please check whether you can get a same probability for the same input when inferences.",
    "740069": "haqishen By probability, do you mean lb score or the output of model? Thanks!",
    "740198": "I mean the raw output (logits) of the model",
    "741776": "haqishen I just checked. The outputs are the same.",
    "741860": "haqishen Also, my maximum gap now becomes Cv 0.9886, Lb 0.9679. I was wondering what else could I try. Thank you very much!",
    "741972": "Another possible problem is that you use lower resolution as input. A lower resolution you use, a larger gap you get",
    "742002": "I'm using 128x128. I'll try a larger size. Thank you! Also, do you think the CNN and linear tail layers would be an issue too? @haqishen Thanks!",
    "742112": "tonychenxyz Maybe you can try another random split seed. I changed different seeds and the cv/lb gap can be made smaller but the lb scores are quite similar.",
    "742118": "syoya1997 what is your cv/lb gap now?",
    "742126": "tonychenxyz Around 0.009 right now. Smaller than my previous gap 0.013. But lb score increases only by 0.001 because my cv scores went down with a different seed.",
    "742141": "Thanks for the information! I'll give another try. @syoya1997",
    "753967": "I tried iterative stratification and it seems like I still have that gap, 98%+ val and 96.5% lb.. I'm not sure whats causing this please help!!\n\nOnly way to reduce that is by adding a weight tensor in loss function",
    "755389": "yannmajewski Hi! I was wondering if you have solved your problem. Also, what exactly do you mean by \"Only way to reduce that is by adding a weight tensor in loss function\"? In addition, are you using the kernel that I referred to (https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch)? Thank you very much!",
    "766440": "I used a random split of 80-20. My CV score is ~ 0.98 but my LB score is 0.71. What could be the reason?\nI did not use any data augmentation"
  },
  "source": "meta"
}