{
  "id": 206488,
  "title": "Does the gap between CV and public LB score need to be very small?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206488",
  "author_name": "Yih-Dar SHIEH",
  "post_date": "2020-12-24T22:12:26.566000",
  "votes": 7,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I prepared a validation dataset for CV - it is almost about the same size as the hidden dataset that will be used for scoring our submissions (i.e it is 5x larger than the size used for public LB).</p>\n<p>At the beginning / middle of this competition, while my model was not good yet (the best at that moment is about 0.781), I have a tight correlation between CV / (public) LB - the differences ranges from <code>0.003</code> to <code>0.005</code>.</p>\n<p>Now, my model becomes better, but the CV / (public) LB has larger gap, diff ranges from <code>0.008</code> to <code>0.01</code>. After debugging as much as possible, (some bugs found and fixed), the gap didn't reduce much.</p>\n<p>Since the public LB only use 20% of the hidden dataset, so I decided to split my validation predictions to 5 parts, and calculate the auc score on each split.</p>\n<p>Here is what I got for the AUC:</p>\n<ul>\n<li><p>whole -  0.807128</p></li>\n<li><p>split 1 -   0.805363</p></li>\n<li><p>split 2 -   0.805951</p></li>\n<li><p>split 3 -   0.810864</p></li>\n<li><p>split 4 -   0.809288</p></li>\n<li><p>split 5 -   0.803919</p></li>\n</ul>\n<p>With another splitting:</p>\n<ul>\n<li><p>split 1 -   0.798685</p></li>\n<li><p>split 2 -   0.803263</p></li>\n<li><p>split 3 -   0.805438</p></li>\n<li><p>split 4 -   0.803936</p></li>\n<li><p>split 5 -   0.811451</p></li>\n</ul>\n<p>For the first splitting, the max auc is on split 3 with <code>0.8108</code> and the min auc is on split 5 with <code>0.8039</code>, and the diff is about <code>0.007</code>. For the second splitting, the gp is <code>0.0127</code> (auc <code>0.811451</code> vs <code>0.798685</code>).</p>\n<p>So it comes the question: does the large gap between my CV (on a valid dataset 5x larger than the one used for public LB) and the public LB really reflects there is some problem in my pipeline?<br>\nOr should I trust CV as long as the correlation is positive?</p>\n<p>ps: Maybe <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> have some thoughts about this situation :) ?</p>",
  "messages": [
    {
      "id": 1125633,
      "postDate": "2020-12-24T22:12:26.567Z",
      "content": "<p>I prepared a validation dataset for CV - it is almost about the same size as the hidden dataset that will be used for scoring our submissions (i.e it is 5x larger than the size used for public LB).</p>\n<p>At the beginning / middle of this competition, while my model was not good yet (the best at that moment is about 0.781), I have a tight correlation between CV / (public) LB - the differences ranges from <code>0.003</code> to <code>0.005</code>.</p>\n<p>Now, my model becomes better, but the CV / (public) LB has larger gap, diff ranges from <code>0.008</code> to <code>0.01</code>. After debugging as much as possible, (some bugs found and fixed), the gap didn't reduce much.</p>\n<p>Since the public LB only use 20% of the hidden dataset, so I decided to split my validation predictions to 5 parts, and calculate the auc score on each split.</p>\n<p>Here is what I got for the AUC:</p>\n<ul>\n<li><p>whole -  0.807128</p></li>\n<li><p>split 1 -   0.805363</p></li>\n<li><p>split 2 -   0.805951</p></li>\n<li><p>split 3 -   0.810864</p></li>\n<li><p>split 4 -   0.809288</p></li>\n<li><p>split 5 -   0.803919</p></li>\n</ul>\n<p>With another splitting:</p>\n<ul>\n<li><p>split 1 -   0.798685</p></li>\n<li><p>split 2 -   0.803263</p></li>\n<li><p>split 3 -   0.805438</p></li>\n<li><p>split 4 -   0.803936</p></li>\n<li><p>split 5 -   0.811451</p></li>\n</ul>\n<p>For the first splitting, the max auc is on split 3 with <code>0.8108</code> and the min auc is on split 5 with <code>0.8039</code>, and the diff is about <code>0.007</code>. For the second splitting, the gp is <code>0.0127</code> (auc <code>0.811451</code> vs <code>0.798685</code>).</p>\n<p>So it comes the question: does the large gap between my CV (on a valid dataset 5x larger than the one used for public LB) and the public LB really reflects there is some problem in my pipeline?<br>\nOr should I trust CV as long as the correlation is positive?</p>\n<p>ps: Maybe <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> have some thoughts about this situation :) ?</p>",
      "rawMarkdown": "I prepared a validation dataset for CV - it is almost about the same size as the hidden dataset that will be used for scoring our submissions (i.e it is 5x larger than the size used for public LB).\n\nAt the beginning / middle of this competition, while my model was not good yet (the best at that moment is about 0.781), I have a tight correlation between CV / (public) LB - the differences ranges from `0.003` to `0.005`.\n\nNow, my model becomes better, but the CV / (public) LB has larger gap, diff ranges from `0.008` to `0.01`. After debugging as much as possible, (some bugs found and fixed), the gap didn't reduce much.\n\nSince the public LB only use 20% of the hidden dataset, so I decided to split my validation predictions to 5 parts, and calculate the auc score on each split.\n\nHere is what I got for the AUC:\n\n - whole -  0.807128\n\n - split 1 -   0.805363\n\n - split 2 -   0.805951\n\n - split 3 -   0.810864\n\n - split 4 -   0.809288\n\n - split 5 -   0.803919\n\nWith another splitting:\n\n - split 1 -   0.798685\n\n - split 2 -   0.803263\n\n - split 3 -   0.805438\n\n - split 4 -   0.803936\n\n - split 5 -   0.811451\n\nFor the first splitting, the max auc is on split 3 with `0.8108` and the min auc is on split 5 with `0.8039`, and the diff is about `0.007`. For the second splitting, the gp is `0.0127` (auc `0.811451` vs `0.798685`).\n\nSo it comes the question: does the large gap between my CV (on a valid dataset 5x larger than the one used for public LB) and the public LB really reflects there is some problem in my pipeline?\nOr should I trust CV as long as the correlation is positive?\n\nps: Maybe @cdeotte have some thoughts about this situation :) ?",
      "votes": 7
    },
    {
      "id": 1126039,
      "postDate": "2020-12-25T09:32:38.847Z",
      "content": "<p>I'm experiencing the same thing, here are my splits (also ~500k each): 0.81233 / 0.80707 / 0.80514/ 0.80825 / 0.81681<br>\nBut when I shuffle the rows, split stabilizes a lot: 0.80916 / 81139 / 0.80973 / 0.81068 / 0.8092</p>\n<p>Although my CV-LB gap is much smaller ~ +0.0005</p>",
      "rawMarkdown": "I'm experiencing the same thing, here are my splits (also ~500k each): 0.81233 / 0.80707 / 0.80514/ 0.80825 / 0.81681\nBut when I shuffle the rows, split stabilizes a lot: 0.80916 / 81139 / 0.80973 / 0.81068 / 0.8092\n\nAlthough my CV-LB gap is much smaller ~ +0.0005\n",
      "votes": 1,
      "replies": [
        {
          "id": 1126042,
          "postDate": "2020-12-25T09:38:49.097Z",
          "content": "<p>Yes, shuffle makes the auc scores on each split more closer. Good for you to have a smaller gap</p>",
          "rawMarkdown": "Yes, shuffle makes the auc scores on each split more closer. Good for you to have a smaller gap"
        },
        {
          "id": 1128204,
          "postDate": "2020-12-27T08:58:17.060Z",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a>, I also tried to shuffling my validation dataset in user level (not shuffling the whole validation rows) - and I still observered the gap. The gap disappears only when the shuffling on the row lever (i.e. discarding the user information).</p>\n<p>You can see a summary below in my reply to <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> comment. </p>",
          "rawMarkdown": "@bacicnikola, I also tried to shuffling my validation dataset in user level (not shuffling the whole validation rows) - and I still observered the gap. The gap disappears only when the shuffling on the row lever (i.e. discarding the user information).\n\nYou can see a summary below in my reply to @lihaorocky comment. "
        }
      ]
    },
    {
      "id": 1129396,
      "postDate": "2020-12-28T10:07:31.710Z",
      "content": "<p>Well, after a lot of bug fixing work. Now the result seems more reasonable. Tested with some models.</p>\n<p>Local: 0.7903, LB: 0.792x<br>\nLocal: 0.7948, LB: 0.795x<br>\nLocal: 0.8027, LB: 0.805x</p>\n<p>Now I can focus on improving my model.</p>",
      "rawMarkdown": "Well, after a lot of bug fixing work. Now the result seems more reasonable. Tested with some models.\n\nLocal: 0.7903, LB: 0.792x\nLocal: 0.7948, LB: 0.795x\nLocal: 0.8027, LB: 0.805x\n\nNow I can focus on improving my model.",
      "votes": 2,
      "replies": [
        {
          "id": 1129502,
          "postDate": "2020-12-28T11:46:26.023Z",
          "content": "<p>Do you think you gap was due to the bugs (other than the wrong way of computing AUC), or it is mainly from the wrong way of computing AUC?</p>\n<p>Despite the fact I observed (the variance of AUC on each splilt/fold), I am still a bit worried. I also spent a lot of time to check my pipeline, but there is no inconsistency between the local val / submission pipeline.</p>\n<p>At least, your local CV dataset is the same size as the one used for the public LB, or the same size as the whole hidden dataset?</p>",
          "rawMarkdown": "Do you think you gap was due to the bugs (other than the wrong way of computing AUC), or it is mainly from the wrong way of computing AUC?\n\nDespite the fact I observed (the variance of AUC on each splilt/fold), I am still a bit worried. I also spent a lot of time to check my pipeline, but there is no inconsistency between the local val / submission pipeline.\n  \nAt least, your local CV dataset is the same size as the one used for the public LB, or the same size as the whole hidden dataset?"
        },
        {
          "id": 1129556,
          "postDate": "2020-12-28T12:32:57.307Z",
          "content": "<p>The gap I reported before is because there are bugs in my pipeline(wrong way to calculate one model input in train and inference pipeline…although I dropped it for now with some performance loss, I will add it later). The wrong way of computing AUC doesn't make too much difference. </p>\n<p>I use 250w rows as CV dataset.</p>",
          "rawMarkdown": "The gap I reported before is because there are bugs in my pipeline(wrong way to calculate one model input in train and inference pipeline...although I dropped it for now with some performance loss, I will add it later). The wrong way of computing AUC doesn't make too much difference. \n\nI use 250w rows as CV dataset."
        }
      ]
    },
    {
      "id": 1125680,
      "postDate": "2020-12-25T00:27:36.517Z",
      "content": "<p>Great post! I recently have very similar experience. I have the same cv strategy with the popular public shared one. When the local AUC is 0.803, the LB is 0.796. And I test with a model which gives me local AUC 0.7977, the LB is 0.790. The gap between them is similar to yours. Now a better model is training with a local AUC 0.810+, and we will see…But I was thinking the gap is because something wrong with my inference pipeline, since I just get this pipeline ready yesterday…But after seeing your post, well, maybe it's normal. Will recheck my pipeline and see what happens…</p>",
      "rawMarkdown": "Great post! I recently have very similar experience. I have the same cv strategy with the popular public shared one. When the local AUC is 0.803, the LB is 0.796. And I test with a model which gives me local AUC 0.7977, the LB is 0.790. The gap between them is similar to yours. Now a better model is training with a local AUC 0.810+, and we will see...But I was thinking the gap is because something wrong with my inference pipeline, since I just get this pipeline ready yesterday...But after seeing your post, well, maybe it's normal. Will recheck my pipeline and see what happens...",
      "votes": 2,
      "replies": [
        {
          "id": 1127995,
          "postDate": "2020-12-27T05:04:03.290Z",
          "content": "<p>After rechecking my inference pipeline for one whole day, I found some bugs. But the gap between local AUC (0.8160) and LB (0.801) is still very big…Will recheck the inference pipeline…</p>",
          "rawMarkdown": "After rechecking my inference pipeline for one whole day, I found some bugs. But the gap between local AUC (0.8160) and LB (0.801) is still very big...Will recheck the inference pipeline...",
          "votes": 1
        },
        {
          "id": 1128201,
          "postDate": "2020-12-27T08:55:24.410Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , I don't want to give conclusion for you, but since I have a vald dataset 5x larger than the one used for public LB, I decided to try 25 different ways of splittings (each splits the valid ds to 5 subsets, equal size, ratio of new users maintained), and here is the observation:</p>\n<p>Each way of splitting is called a <code>split</code>. On each split, there are 5 <code>folds</code>.</p>\n<pre><code>    \"global max auc\": 0.8149430894390665,\n    \"global min auc\": 0.8021997012189686,\n    \"max (over split) of average (over fold) auc\": 0.8080425274004627,\n    \"min (over split) of average (over fold) auc\": 0.8078732167252156,\n    \"average (over split) of max (over fold) auc\": 0.8113316524152409,\n    \"average (over split) of min auc\": 0.8049578651128897,\n    \"global max gap\": 0.012743388220097907,\n    \"average (over split) of max (over fold)  auc gap\": 0.006373787302351337,\n    \"gap of average (over fold) auc\": 0.00016931067524705856,\n    \"gap of average (over split) auc\": 0.006373787302351164,\n    \"average of average auc\": 0.8079555153180962\n</code></pre>",
          "rawMarkdown": "@lihaorocky , I don't want to give conclusion for you, but since I have a vald dataset 5x larger than the one used for public LB, I decided to try 25 different ways of splittings (each splits the valid ds to 5 subsets, equal size, ratio of new users maintained), and here is the observation:\n\nEach way of splitting is called a `split`. On each split, there are 5 `folds`.\n\n```\n    \"global max auc\": 0.8149430894390665,\n    \"global min auc\": 0.8021997012189686,\n    \"max (over split) of average (over fold) auc\": 0.8080425274004627,\n    \"min (over split) of average (over fold) auc\": 0.8078732167252156,\n    \"average (over split) of max (over fold) auc\": 0.8113316524152409,\n    \"average (over split) of min auc\": 0.8049578651128897,\n    \"global max gap\": 0.012743388220097907,\n    \"average (over split) of max (over fold)  auc gap\": 0.006373787302351337,\n    \"gap of average (over fold) auc\": 0.00016931067524705856,\n    \"gap of average (over split) auc\": 0.006373787302351164,\n    \"average of average auc\": 0.8079555153180962\n```"
        },
        {
          "id": 1128220,
          "postDate": "2020-12-27T09:06:32.937Z",
          "content": "<p>Thanks for the report. But I kind of know why the gap for me is so big. The local auc I referenced is the average calculated auc of each batch from each epoch during training in validation dataset not the auc of the overall prediction against the ground truth. And I just realized they should be different because  of the way the auc is calculated. But since my model is still training (using very limited resource…training lasts 3 days…). I will check if this is the reason after the model training is finished. </p>",
          "rawMarkdown": "Thanks for the report. But I kind of know why the gap for me is so big. The local auc I referenced is the average calculated auc of each batch from each epoch during training in validation dataset not the auc of the overall prediction against the ground truth. And I just realized they should be different because  of the way the auc is calculated. But since my model is still training (using very limited resource...training lasts 3 days...). I will check if this is the reason after the model training is finished. ",
          "votes": 1
        },
        {
          "id": 1128279,
          "postDate": "2020-12-27T09:45:12.043Z",
          "content": "<p>In this case, I would say yes - we should compute the AUC over the whole set of validation dataset, not the average of each AUC on the batches.</p>\n<p>You didn't save your predictions in a file? If you have such outputs, you can already verify if this is a problem.</p>",
          "rawMarkdown": "In this case, I would say yes - we should compute the AUC over the whole set of validation dataset, not the average of each AUC on the batches.\n\nYou didn't save your predictions in a file? If you have such outputs, you can already verify if this is a problem."
        },
        {
          "id": 1128289,
          "postDate": "2020-12-27T09:52:13.433Z",
          "content": "<p>Well, I should do that.😂</p>",
          "rawMarkdown": "Well, I should do that.😂",
          "votes": 2
        }
      ]
    },
    {
      "id": 1125640,
      "postDate": "2020-12-24T22:38:25.080Z",
      "content": "<p>Is the CV better than LB? Or is it the other way around?</p>\n<p>Initially I had have a very tight co-relation of CV and LB as well, but today I found out that I wasn't removing lectures from the CV. I've not yet checked what's the correlation they are removed but my guess is that now my CV would be worse than LB. </p>\n<p>From experiences I'd say to stick with a good validation set is better than trying to overfit on Public LB.</p>",
      "rawMarkdown": "Is the CV better than LB? Or is it the other way around?\n\nInitially I had have a very tight co-relation of CV and LB as well, but today I found out that I wasn't removing lectures from the CV. I've not yet checked what's the correlation they are removed but my guess is that now my CV would be worse than LB. \n\nFrom experiences I'd say to stick with a good validation set is better than trying to overfit on Public LB.",
      "replies": [
        {
          "id": 1126018,
          "postDate": "2020-12-25T09:01:01.310Z",
          "content": "<p>CV (based on a large validation dataset) is better than public LB - but as I mentioned, CV itself has some variations if I look the AUC scores on different splits of my validation dataset.</p>\n<blockquote>\n  <p>stick with a good validation set is better than trying to overfit on Public LB.</p>\n</blockquote>\n<p>That's true, but the large gap makes me think if my validation has some problem - that's the real question :)</p>",
          "rawMarkdown": "CV (based on a large validation dataset) is better than public LB - but as I mentioned, CV itself has some variations if I look the AUC scores on different splits of my validation dataset.\n\n> stick with a good validation set is better than trying to overfit on Public LB.\n\nThat's true, but the large gap makes me think if my validation has some problem - that's the real question :)"
        },
        {
          "id": 1126433,
          "postDate": "2020-12-25T15:52:56.100Z",
          "content": "<p>Finally completed a submission</p>\n<p>Val : 0.7863 (Without lectures)<br>\nLB  : 0.792</p>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> TPU model performs on-bar with GPU one on LB. So I think the whole effort was worth it. Now to do some upscaling :D</p>",
          "rawMarkdown": "Finally completed a submission\n\nVal : 0.7863 (Without lectures)\nLB  : 0.792\n\n@yihdarshieh TPU model performs on-bar with GPU one on LB. So I think the whole effort was worth it. Now to do some upscaling :D",
          "votes": 1
        },
        {
          "id": 1126455,
          "postDate": "2020-12-25T16:17:13.850Z",
          "content": "<p>Great! Good luck for the last 2 weeks</p>",
          "rawMarkdown": "Great! Good luck for the last 2 weeks",
          "votes": 1
        },
        {
          "id": 1126468,
          "postDate": "2020-12-25T16:29:55.220Z",
          "content": "<p>Thanks once again</p>",
          "rawMarkdown": "Thanks once again",
          "votes": 1
        },
        {
          "id": 1126731,
          "postDate": "2020-12-25T21:41:38.307Z",
          "content": "<p>Oh, you are at 40th now …  what did you do with TPU magic :)</p>",
          "rawMarkdown": "Oh, you are at 40th now ...  what did you do with TPU magic :)"
        },
        {
          "id": 1126924,
          "postDate": "2020-12-26T05:06:08.370Z",
          "content": "<p>Yeah. Just upscaled to d_model 256 for now :)</p>\n<p>This time my Val score is again same as LB public. I'm not sure what to make of it. I'll make a few more submission to check that out.</p>",
          "rawMarkdown": "Yeah. Just upscaled to d_model 256 for now :)\n\nThis time my Val score is again same as LB public. I'm not sure what to make of it. I'll make a few more submission to check that out."
        }
      ]
    },
    {
      "id": 1136972,
      "postDate": "2021-01-03T15:21:55.927Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1126039,
      "author_name": "Nikola Bacic",
      "author_url": "",
      "post_date": "2020-12-25T09:32:38.847000",
      "content": "<p>I'm experiencing the same thing, here are my splits (also ~500k each): 0.81233 / 0.80707 / 0.80514/ 0.80825 / 0.81681<br>\nBut when I shuffle the rows, split stabilizes a lot: 0.80916 / 81139 / 0.80973 / 0.81068 / 0.8092</p>\n<p>Although my CV-LB gap is much smaller ~ +0.0005</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1126042,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-25T09:38:49.097000",
          "content": "<p>Yes, shuffle makes the auc scores on each split more closer. Good for you to have a smaller gap</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128204,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-27T08:58:17.060000",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a>, I also tried to shuffling my validation dataset in user level (not shuffling the whole validation rows) - and I still observered the gap. The gap disappears only when the shuffling on the row lever (i.e. discarding the user information).</p>\n<p>You can see a summary below in my reply to <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> comment. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1129396,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2020-12-28T10:07:31.710000",
      "content": "<p>Well, after a lot of bug fixing work. Now the result seems more reasonable. Tested with some models.</p>\n<p>Local: 0.7903, LB: 0.792x<br>\nLocal: 0.7948, LB: 0.795x<br>\nLocal: 0.8027, LB: 0.805x</p>\n<p>Now I can focus on improving my model.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1129502,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-28T11:46:26.023000",
          "content": "<p>Do you think you gap was due to the bugs (other than the wrong way of computing AUC), or it is mainly from the wrong way of computing AUC?</p>\n<p>Despite the fact I observed (the variance of AUC on each splilt/fold), I am still a bit worried. I also spent a lot of time to check my pipeline, but there is no inconsistency between the local val / submission pipeline.</p>\n<p>At least, your local CV dataset is the same size as the one used for the public LB, or the same size as the whole hidden dataset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129556,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2020-12-28T12:32:57.307000",
          "content": "<p>The gap I reported before is because there are bugs in my pipeline(wrong way to calculate one model input in train and inference pipeline…although I dropped it for now with some performance loss, I will add it later). The wrong way of computing AUC doesn't make too much difference. </p>\n<p>I use 250w rows as CV dataset.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1125680,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2020-12-25T00:27:36.517000",
      "content": "<p>Great post! I recently have very similar experience. I have the same cv strategy with the popular public shared one. When the local AUC is 0.803, the LB is 0.796. And I test with a model which gives me local AUC 0.7977, the LB is 0.790. The gap between them is similar to yours. Now a better model is training with a local AUC 0.810+, and we will see…But I was thinking the gap is because something wrong with my inference pipeline, since I just get this pipeline ready yesterday…But after seeing your post, well, maybe it's normal. Will recheck my pipeline and see what happens…</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1127995,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2020-12-27T05:04:03.290000",
          "content": "<p>After rechecking my inference pipeline for one whole day, I found some bugs. But the gap between local AUC (0.8160) and LB (0.801) is still very big…Will recheck the inference pipeline…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1128201,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-27T08:55:24.410000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , I don't want to give conclusion for you, but since I have a vald dataset 5x larger than the one used for public LB, I decided to try 25 different ways of splittings (each splits the valid ds to 5 subsets, equal size, ratio of new users maintained), and here is the observation:</p>\n<p>Each way of splitting is called a <code>split</code>. On each split, there are 5 <code>folds</code>.</p>\n<pre><code>    \"global max auc\": 0.8149430894390665,\n    \"global min auc\": 0.8021997012189686,\n    \"max (over split) of average (over fold) auc\": 0.8080425274004627,\n    \"min (over split) of average (over fold) auc\": 0.8078732167252156,\n    \"average (over split) of max (over fold) auc\": 0.8113316524152409,\n    \"average (over split) of min auc\": 0.8049578651128897,\n    \"global max gap\": 0.012743388220097907,\n    \"average (over split) of max (over fold)  auc gap\": 0.006373787302351337,\n    \"gap of average (over fold) auc\": 0.00016931067524705856,\n    \"gap of average (over split) auc\": 0.006373787302351164,\n    \"average of average auc\": 0.8079555153180962\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128220,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2020-12-27T09:06:32.937000",
          "content": "<p>Thanks for the report. But I kind of know why the gap for me is so big. The local auc I referenced is the average calculated auc of each batch from each epoch during training in validation dataset not the auc of the overall prediction against the ground truth. And I just realized they should be different because  of the way the auc is calculated. But since my model is still training (using very limited resource…training lasts 3 days…). I will check if this is the reason after the model training is finished. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1128279,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-27T09:45:12.043000",
          "content": "<p>In this case, I would say yes - we should compute the AUC over the whole set of validation dataset, not the average of each AUC on the batches.</p>\n<p>You didn't save your predictions in a file? If you have such outputs, you can already verify if this is a problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1128289,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2020-12-27T09:52:13.433000",
          "content": "<p>Well, I should do that.😂</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1125640,
      "author_name": "AbdurRafae",
      "author_url": "",
      "post_date": "2020-12-24T22:38:25.080000",
      "content": "<p>Is the CV better than LB? Or is it the other way around?</p>\n<p>Initially I had have a very tight co-relation of CV and LB as well, but today I found out that I wasn't removing lectures from the CV. I've not yet checked what's the correlation they are removed but my guess is that now my CV would be worse than LB. </p>\n<p>From experiences I'd say to stick with a good validation set is better than trying to overfit on Public LB.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1126018,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-25T09:01:01.310000",
          "content": "<p>CV (based on a large validation dataset) is better than public LB - but as I mentioned, CV itself has some variations if I look the AUC scores on different splits of my validation dataset.</p>\n<blockquote>\n  <p>stick with a good validation set is better than trying to overfit on Public LB.</p>\n</blockquote>\n<p>That's true, but the large gap makes me think if my validation has some problem - that's the real question :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1126433,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-25T15:52:56.100000",
          "content": "<p>Finally completed a submission</p>\n<p>Val : 0.7863 (Without lectures)<br>\nLB  : 0.792</p>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> TPU model performs on-bar with GPU one on LB. So I think the whole effort was worth it. Now to do some upscaling :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1126455,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-25T16:17:13.850000",
          "content": "<p>Great! Good luck for the last 2 weeks</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1126468,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-25T16:29:55.220000",
          "content": "<p>Thanks once again</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1126731,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-25T21:41:38.307000",
          "content": "<p>Oh, you are at 40th now …  what did you do with TPU magic :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1126924,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-26T05:06:08.370000",
          "content": "<p>Yeah. Just upscaled to d_model 256 for now :)</p>\n<p>This time my Val score is again same as LB public. I'm not sure what to make of it. I'll make a few more submission to check that out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1136972,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-03T15:21:55.927000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125633": "I prepared a validation dataset for CV - it is almost about the same size as the hidden dataset that will be used for scoring our submissions (i.e it is 5x larger than the size used for public LB).\n\nAt the beginning / middle of this competition, while my model was not good yet (the best at that moment is about 0.781), I have a tight correlation between CV / (public) LB - the differences ranges from `0.003` to `0.005`.\n\nNow, my model becomes better, but the CV / (public) LB has larger gap, diff ranges from `0.008` to `0.01`. After debugging as much as possible, (some bugs found and fixed), the gap didn't reduce much.\n\nSince the public LB only use 20% of the hidden dataset, so I decided to split my validation predictions to 5 parts, and calculate the auc score on each split.\n\nHere is what I got for the AUC:\n\n - whole -  0.807128\n\n - split 1 -   0.805363\n\n - split 2 -   0.805951\n\n - split 3 -   0.810864\n\n - split 4 -   0.809288\n\n - split 5 -   0.803919\n\nWith another splitting:\n\n - split 1 -   0.798685\n\n - split 2 -   0.803263\n\n - split 3 -   0.805438\n\n - split 4 -   0.803936\n\n - split 5 -   0.811451\n\nFor the first splitting, the max auc is on split 3 with `0.8108` and the min auc is on split 5 with `0.8039`, and the diff is about `0.007`. For the second splitting, the gp is `0.0127` (auc `0.811451` vs `0.798685`).\n\nSo it comes the question: does the large gap between my CV (on a valid dataset 5x larger than the one used for public LB) and the public LB really reflects there is some problem in my pipeline?\nOr should I trust CV as long as the correlation is positive?\n\nps: Maybe @cdeotte have some thoughts about this situation :) ?",
    "1126039": "I'm experiencing the same thing, here are my splits (also ~500k each): 0.81233 / 0.80707 / 0.80514/ 0.80825 / 0.81681\nBut when I shuffle the rows, split stabilizes a lot: 0.80916 / 81139 / 0.80973 / 0.81068 / 0.8092\n\nAlthough my CV-LB gap is much smaller ~ +0.0005\n",
    "1129396": "Well, after a lot of bug fixing work. Now the result seems more reasonable. Tested with some models.\n\nLocal: 0.7903, LB: 0.792x\nLocal: 0.7948, LB: 0.795x\nLocal: 0.8027, LB: 0.805x\n\nNow I can focus on improving my model.",
    "1125680": "Great post! I recently have very similar experience. I have the same cv strategy with the popular public shared one. When the local AUC is 0.803, the LB is 0.796. And I test with a model which gives me local AUC 0.7977, the LB is 0.790. The gap between them is similar to yours. Now a better model is training with a local AUC 0.810+, and we will see...But I was thinking the gap is because something wrong with my inference pipeline, since I just get this pipeline ready yesterday...But after seeing your post, well, maybe it's normal. Will recheck my pipeline and see what happens...",
    "1125640": "Is the CV better than LB? Or is it the other way around?\n\nInitially I had have a very tight co-relation of CV and LB as well, but today I found out that I wasn't removing lectures from the CV. I've not yet checked what's the correlation they are removed but my guess is that now my CV would be worse than LB. \n\nFrom experiences I'd say to stick with a good validation set is better than trying to overfit on Public LB.",
    "1136972": ""
  }
}