{
  "id": 89312,
  "title": "Discrepancy between CV and LB",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/89312",
  "author_name": "Dmytro Danevskyi",
  "post_date": "2019-04-12T22:45:21.722000",
  "votes": 17,
  "comment_count": 32,
  "views": 0,
  "content": "<p>The data tab says that the curated train set and the test set came from the same dataset and both were labeled manually. However, I see that most public kernels and also my own pipeline (which differs from public kernels) have very high CV-LB difference that seems to be around 0.2. </p>\n\n<p>Are you getting the same results?</p>",
  "messages": [
    {
      "id": 515655,
      "postDate": "2019-04-12T22:45:21.723Z",
      "content": "<p>The data tab says that the curated train set and the test set came from the same dataset and both were labeled manually. However, I see that most public kernels and also my own pipeline (which differs from public kernels) have very high CV-LB difference that seems to be around 0.2. </p>\n\n<p>Are you getting the same results?</p>",
      "rawMarkdown": "The data tab says that the curated train set and the test set came from the same dataset and both were labeled manually. However, I see that most public kernels and also my own pipeline (which differs from public kernels) have very high CV-LB difference that seems to be around 0.2. \n\nAre you getting the same results?",
      "votes": 17
    },
    {
      "id": 521710,
      "postDate": "2019-04-23T10:08:47.610Z",
      "content": "<p>5folds, CV:0.67, LB:0.635<br>\nUpdate: 5folds, CV:0.685, LB:0.665</p>",
      "rawMarkdown": "5folds, CV:0.67, LB:0.635<br>\nUpdate: 5folds, CV:0.685, LB:0.665",
      "votes": 5,
      "replies": [
        {
          "id": 521715,
          "postDate": "2019-04-23T10:27:54.067Z",
          "content": "<p>Do you use noisy data when training?</p>",
          "rawMarkdown": "Do you use noisy data when training?"
        },
        {
          "id": 521723,
          "postDate": "2019-04-23T10:39:06.370Z",
          "content": "<p>Yes.</p>",
          "rawMarkdown": "Yes.",
          "votes": 1
        },
        {
          "id": 521747,
          "postDate": "2019-04-23T11:08:33.077Z",
          "content": "<p>May I ask if you use both curated and noisy data for validation?</p>",
          "rawMarkdown": "May I ask if you use both curated and noisy data for validation?"
        },
        {
          "id": 521753,
          "postDate": "2019-04-23T11:17:45.977Z",
          "content": "<p>Yes. Since I started this competition today, I had no time to make any effort regarding validation.<br>\nI have focused on modeling and loss function.</p>",
          "rawMarkdown": "Yes. Since I started this competition today, I had no time to make any effort regarding validation.<br>\nI have focused on modeling and loss function.",
          "votes": 1
        },
        {
          "id": 522826,
          "postDate": "2019-04-25T05:02:08.263Z",
          "content": "<p>Yeah I noticed if you validate with noisy data together then the CV more closely matches LB (you must check the lwlrap with noisy data included). Once you remove noisy data from validation after training together then you can see the lwlrap score is once again around 0.2 higher.</p>",
          "rawMarkdown": "Yeah I noticed if you validate with noisy data together then the CV more closely matches LB (you must check the lwlrap with noisy data included). Once you remove noisy data from validation after training together then you can see the lwlrap score is once again around 0.2 higher.",
          "votes": 2
        }
      ]
    },
    {
      "id": 518738,
      "postDate": "2019-04-17T17:52:39.173Z",
      "content": "<p><a href=\"/eduardofonseca\">@eduardofonseca</a> <a href=\"/plakal\">@plakal</a></p>\n\n<p>Hello, </p>\n\n<p>Just a little clarification. Can we expect that the stage2 test data will have the same distribution as the stage1 test data? Thanks.</p>",
      "rawMarkdown": "@eduardofonseca @plakal\n\nHello, \n\nJust a little clarification. Can we expect that the stage2 test data will have the same distribution as the stage1 test data? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 518834,
          "postDate": "2019-04-17T22:34:44.837Z",
          "content": "<p>The first and second stages were produced by a 1:3 random split of the original test dataset.</p>",
          "rawMarkdown": "The first and second stages were produced by a 1:3 random split of the original test dataset.",
          "votes": 7
        }
      ]
    },
    {
      "id": 517823,
      "postDate": "2019-04-16T15:24:24Z",
      "content": "<p>same for me. I do the predictions over the whole validation set (20% holdout) and afterwards i calculate the lwlrap. CV is about 0.73, i don't know exactly because i didn't do k-fold, just the final run over all the data with similar settings, the LB was then 0.505</p>",
      "rawMarkdown": "same for me. I do the predictions over the whole validation set (20% holdout) and afterwards i calculate the lwlrap. CV is about 0.73, i don't know exactly because i didn't do k-fold, just the final run over all the data with similar settings, the LB was then 0.505",
      "votes": 1
    },
    {
      "id": 516317,
      "postDate": "2019-04-14T01:43:47.673Z",
      "content": "<p>I have just made a few submissions and already seeing the similar relationships between cv and lb.</p>\n\n<ul>\n<li>cv 0.683, lb 0.436</li>\n<li>cv 0.739, lb 0.505</li>\n</ul>\n\n<p>I only use train_curated and 5-folds like op.</p>",
      "rawMarkdown": "I have just made a few submissions and already seeing the similar relationships between cv and lb.\n\n- cv 0.683, lb 0.436\n- cv 0.739, lb 0.505\n\nI only use train_curated and 5-folds like op.",
      "votes": 1
    },
    {
      "id": 515871,
      "postDate": "2019-04-13T08:35:38.903Z",
      "content": "<p>Yes, we also observe huge difference around 0.2 between CV and LB, but luckily its correlated, so when CV goes up LB goes up too.</p>",
      "rawMarkdown": "Yes, we also observe huge difference around 0.2 between CV and LB, but luckily its correlated, so when CV goes up LB goes up too.",
      "votes": 1,
      "replies": [
        {
          "id": 515891,
          "postDate": "2019-04-13T09:16:47.103Z",
          "content": "<p>Yeah, I also observe the correlation. However, such a huge gap is a bit disturbing, since it's stated that the train and test data have to be very similar.</p>",
          "rawMarkdown": "Yeah, I also observe the correlation. However, such a huge gap is a bit disturbing, since it's stated that the train and test data have to be very similar.",
          "votes": 1
        }
      ]
    },
    {
      "id": 515759,
      "postDate": "2019-04-13T04:58:59.637Z",
      "content": "<p>Yes I am seeing the same. On my local CV I get on average 0.76 lwlrap score which is close to 0.2 difference from LB. I used the lwlrap implementation that was provided and computed on the entire validation set.</p>",
      "rawMarkdown": "Yes I am seeing the same. On my local CV I get on average 0.76 lwlrap score which is close to 0.2 difference from LB. I used the lwlrap implementation that was provided and computed on the entire validation set.",
      "votes": 1,
      "replies": [
        {
          "id": 515766,
          "postDate": "2019-04-13T05:19:29.890Z",
          "content": "<p><a href=\"/eduardofonseca\">@eduardofonseca</a> if he has any thoughts.</p>\n\n<p>And just to be clear, your cross-validation set is a randomized subset of the curated train set, right (there are patterns in the order of rows in the CSV)? And you have all classes equally represented in your train set and your validation set? Are you sure you are not overfitting to your (now smaller) train set because you might be using a large model?</p>",
          "rawMarkdown": "@eduardofonseca if he has any thoughts.\n\nAnd just to be clear, your cross-validation set is a randomized subset of the curated train set, right (there are patterns in the order of rows in the CSV)? And you have all classes equally represented in your train set and your validation set? Are you sure you are not overfitting to your (now smaller) train set because you might be using a large model?\n"
        }
      ]
    },
    {
      "id": 515674,
      "postDate": "2019-04-12T23:45:55.683Z",
      "content": "<p>By CV, do you mean lwlrap computed on a held-out cross-validation set taken from the curated train set? Are you using the lwlrap implementation that we have provided, and are you computing it over the entire validation set? There was another thread about lwlrap where someone was computing lwlrap per batch and taking the average, which is not the way to combine lwlraps.</p>",
      "rawMarkdown": "By CV, do you mean lwlrap computed on a held-out cross-validation set taken from the curated train set? Are you using the lwlrap implementation that we have provided, and are you computing it over the entire validation set? There was another thread about lwlrap where someone was computing lwlrap per batch and taking the average, which is not the way to combine lwlraps.",
      "votes": 1,
      "replies": [
        {
          "id": 515821,
          "postDate": "2019-04-13T07:14:07.900Z",
          "content": "<p>Sure, I'm doing CV on the curated train set. I'm using the provided lwlrap implementation. The metric is computed on the entire validation set in a single pass.</p>",
          "rawMarkdown": "Sure, I'm doing CV on the curated train set. I'm using the provided lwlrap implementation. The metric is computed on the entire validation set in a single pass."
        },
        {
          "id": 515893,
          "postDate": "2019-04-13T09:19:10.337Z",
          "content": "<p>I've just tested my approach on a holdout set. The results are very close to my CV results and the gap is still around 0.2.</p>",
          "rawMarkdown": "I've just tested my approach on a holdout set. The results are very close to my CV results and the gap is still around 0.2.",
          "votes": 2
        }
      ]
    },
    {
      "id": 518604,
      "postDate": "2019-04-17T14:10:37.730Z",
      "content": "<p>5 folds   cv 0.803    lb  0.62</p>",
      "rawMarkdown": "5 folds   cv 0.803    lb  0.62",
      "votes": 2
    },
    {
      "id": 515709,
      "postDate": "2019-04-13T01:59:54.457Z",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a> Yes, it's same with me. As far as I see the statistics, train and test set are similar. So, I wonder how they are annotated, e.g. by different person and splitted?</p>",
      "rawMarkdown": "@ddanevskyi Yes, it's same with me. As far as I see the statistics, train and test set are similar. So, I wonder how they are annotated, e.g. by different person and splitted?",
      "votes": 2
    },
    {
      "id": 515851,
      "postDate": "2019-04-13T08:14:36.920Z",
      "content": "<p>When your CV improves, does your LB score improve too?</p>",
      "rawMarkdown": "When your CV improves, does your LB score improve too?",
      "votes": -1,
      "replies": [
        {
          "id": 515868,
          "postDate": "2019-04-13T08:29:00.057Z",
          "content": "<p>Yes. </p>",
          "rawMarkdown": "Yes. "
        }
      ]
    },
    {
      "id": 523696,
      "postDate": "2019-04-26T18:59:15.313Z",
      "content": "<p>Well, my last observations are that LB-CV gap is model-specific and also that as local score improves, the gap seems to become smaller. Note, however, that these facts may be specific to my approach and may not work for you.</p>",
      "rawMarkdown": "Well, my last observations are that LB-CV gap is model-specific and also that as local score improves, the gap seems to become smaller. Note, however, that these facts may be specific to my approach and may not work for you."
    },
    {
      "id": 521393,
      "postDate": "2019-04-22T21:18:42.897Z",
      "content": "<p>Has anyone managed to find a way to close the gap? :) It seems that training on the noisy part of the train set helps a bit, but the difference is still around 0.15...</p>",
      "rawMarkdown": "Has anyone managed to find a way to close the gap? :) It seems that training on the noisy part of the train set helps a bit, but the difference is still around 0.15...",
      "replies": [
        {
          "id": 521683,
          "postDate": "2019-04-23T09:06:18Z",
          "content": "<p>Same here, using both noisy and curated data, 5Fold CV\n0.747 CV  &lt;-&gt; 0.617 LB\na 0.13 difference</p>",
          "rawMarkdown": "Same here, using both noisy and curated data, 5Fold CV\n0.747 CV  &lt;-&gt; 0.617 LB\na 0.13 difference"
        }
      ]
    },
    {
      "id": 516143,
      "postDate": "2019-04-13T17:17:23.353Z",
      "content": "<p>Hi,</p>\n\n<p>just to confirm or deny some facts:</p>\n\n<ul>\n<li>the curated train set and the test set come from the currently under development FSD, yes. </li>\n<li>both are labeled manually, yes. Several people were involved in the annotation of both sets.</li>\n</ul>\n\n<p><a href=\"/ddanevskyi\">@ddanevskyi</a>, you mention:</p>\n\n<blockquote>\n  <p>it's stated that the train and test data have to be very similar.</p>\n</blockquote>\n\n<p>really? I'm afraid I don’t remember that being stated anywhere... Can you point us where it is stated?</p>\n\n<p>I am not sure I see the problem here. The fact that CV scores are higher than test scores can be just a feature, and it is not uncommon in recognition tasks.</p>\n\n<p>thanks!</p>",
      "rawMarkdown": "Hi,\n\njust to confirm or deny some facts:\n\n- the curated train set and the test set come from the currently under development FSD, yes. \n- both are labeled manually, yes. Several people were involved in the annotation of both sets.\n\n@ddanevskyi, you mention:\n&gt; it's stated that the train and test data have to be very similar.\n\nreally? I'm afraid I don’t remember that being stated anywhere... Can you point us where it is stated?\n\nI am not sure I see the problem here. The fact that CV scores are higher than test scores can be just a feature, and it is not uncommon in recognition tasks.\n\nthanks!",
      "replies": [
        {
          "id": 516153,
          "postDate": "2019-04-13T17:33:31.233Z",
          "content": "<p>Oh, maybe I got it wrong, but for some reason, I assumed that since train/test come from the same dataset and both were manually labeled, their distributions should be the same. If it's not the case, there is indeed no problem. Thank you for clarification.</p>",
          "rawMarkdown": "Oh, maybe I got it wrong, but for some reason, I assumed that since train/test come from the same dataset and both were manually labeled, their distributions should be the same. If it's not the case, there is indeed no problem. Thank you for clarification."
        }
      ]
    },
    {
      "id": 515790,
      "postDate": "2019-04-13T06:16:00.507Z",
      "content": "<p>As an alternative data point, I tried evaluating one of our baseline models that had been trained purely on the noisy data and got comparable lwlraps when evaluated on the test set and an equal-sized random subset of the curated train set (curated train was lower by ~0.004). I would double check that all classes are equally represented in your train and validation sets, and that your model has not gotten large enough that it overfits to your train and validation sets.</p>",
      "rawMarkdown": "As an alternative data point, I tried evaluating one of our baseline models that had been trained purely on the noisy data and got comparable lwlraps when evaluated on the test set and an equal-sized random subset of the curated train set (curated train was lower by ~0.004). I would double check that all classes are equally represented in your train and validation sets, and that your model has not gotten large enough that it overfits to your train and validation sets.",
      "replies": [
        {
          "id": 515823,
          "postDate": "2019-04-13T07:21:04.370Z",
          "content": "<p>It seems that almost everyone has this issue. My local scores are close to what <a href=\"/jamesrequa\">@jamesrequa</a> has reported and my LB score is also very close to his score. I really doubt that we have the same bug/leak since our pipelines are probably very different.</p>\n\n<p>What is the CV score for your baseline model? Currently, I'm interested in only the curated set performance.</p>\n\n<p>Thank you in advance!</p>",
          "rawMarkdown": "It seems that almost everyone has this issue. My local scores are close to what @jamesrequa has reported and my LB score is also very close to his score. I really doubt that we have the same bug/leak since our pipelines are probably very different.\n\nWhat is the CV score for your baseline model? Currently, I'm interested in only the curated set performance.\n\nThank you in advance!",
          "votes": 1
        },
        {
          "id": 516117,
          "postDate": "2019-04-13T16:25:35.317Z",
          "content": "<p>Our submitted baseline was trained using an internal copy of the train and test sets and was validated using the entire test set. I can re-run the best model using only the released data from Kaggle. Can you share the train and validation splits you used?</p>",
          "rawMarkdown": "Our submitted baseline was trained using an internal copy of the train and test sets and was validated using the entire test set. I can re-run the best model using only the released data from Kaggle. Can you share the train and validation splits you used?"
        },
        {
          "id": 516157,
          "postDate": "2019-04-13T17:41:37.463Z",
          "content": "<p>Sure, I used 5 fold cv with default KFold from <code>sklearn</code> with <code>random_state=42</code>. Currently, I'm using only curated data.</p>",
          "rawMarkdown": "Sure, I used 5 fold cv with default KFold from `sklearn` with `random_state=42`. Currently, I'm using only curated data."
        },
        {
          "id": 522246,
          "postDate": "2019-04-24T05:05:26.103Z",
          "content": "<p>To close out this sub-thread, as a rough point of comparison, I tried training the baseline with the same hyperparameters and number of epochs as the original submission, but with a 80%-20% train-validation split chosen by your choice of fold selector: sklearn's KFold with shuffle=True and random_state=42 (note that I tried just one of the 5 splits). The validation lwlrap was ~0.705 compared with public test lwlrap of ~0.537 and full (public + private) test lwlrap of ~0.544.</p>\n\n<p>I wouldn't get too alarmed by the difference in the lwlraps. As Eduardo said, train and test are taken from the same source of data and have been annotated in similar ways, but they are not identical in all respects. Perhaps you can figure out how to make use of the extra noisy data to your model's advantage. Also note that lwlrap is designed to be a weighted average of per-class lwlraps, so it might be time to look into how your models are doing for individual classes in order to raise your overall rank.</p>\n\n<p>BTW, you didn't mention it and the KFold seems to use a default shuffle of False, but I hope you are shuffling when making your folds? The lines in the training CSV are not in completely random order.</p>\n\n<p>Also worth checking out: splitting your data blindly may increase bias given that this is a multi-label problem. Perhaps take a look at <a href=\"http://scikit.ml/stratification.html\">http://scikit.ml/stratification.html</a> (although I don't have much experience with it myself)</p>",
          "rawMarkdown": "To close out this sub-thread, as a rough point of comparison, I tried training the baseline with the same hyperparameters and number of epochs as the original submission, but with a 80%-20% train-validation split chosen by your choice of fold selector: sklearn's KFold with shuffle=True and random_state=42 (note that I tried just one of the 5 splits). The validation lwlrap was ~0.705 compared with public test lwlrap of ~0.537 and full (public + private) test lwlrap of ~0.544.\n\nI wouldn't get too alarmed by the difference in the lwlraps. As Eduardo said, train and test are taken from the same source of data and have been annotated in similar ways, but they are not identical in all respects. Perhaps you can figure out how to make use of the extra noisy data to your model's advantage. Also note that lwlrap is designed to be a weighted average of per-class lwlraps, so it might be time to look into how your models are doing for individual classes in order to raise your overall rank.\n\nBTW, you didn't mention it and the KFold seems to use a default shuffle of False, but I hope you are shuffling when making your folds? The lines in the training CSV are not in completely random order.\n\nAlso worth checking out: splitting your data blindly may increase bias given that this is a multi-label problem. Perhaps take a look at http://scikit.ml/stratification.html (although I don't have much experience with it myself)",
          "votes": 5
        }
      ]
    },
    {
      "id": 522310,
      "postDate": "2019-04-24T07:54:02.883Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 521710,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2019-04-23T10:08:47.610000",
      "content": "<p>5folds, CV:0.67, LB:0.635<br>\nUpdate: 5folds, CV:0.685, LB:0.665</p>",
      "votes": 5,
      "replies": [
        {
          "id": 521715,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2019-04-23T10:27:54.067000",
          "content": "<p>Do you use noisy data when training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 521723,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2019-04-23T10:39:06.370000",
          "content": "<p>Yes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 521747,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-23T11:08:33.077000",
          "content": "<p>May I ask if you use both curated and noisy data for validation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 521753,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2019-04-23T11:17:45.977000",
          "content": "<p>Yes. Since I started this competition today, I had no time to make any effort regarding validation.<br>\nI have focused on modeling and loss function.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522826,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-04-25T05:02:08.263000",
          "content": "<p>Yeah I noticed if you validate with noisy data together then the CV more closely matches LB (you must check the lwlrap with noisy data included). Once you remove noisy data from validation after training together then you can see the lwlrap score is once again around 0.2 higher.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 518738,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-04-17T17:52:39.173000",
      "content": "<p><a href=\"/eduardofonseca\">@eduardofonseca</a> <a href=\"/plakal\">@plakal</a></p>\n\n<p>Hello, </p>\n\n<p>Just a little clarification. Can we expect that the stage2 test data will have the same distribution as the stage1 test data? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 518834,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-17T22:34:44.837000",
          "content": "<p>The first and second stages were produced by a 1:3 random split of the original test dataset.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 517823,
      "author_name": "Marek Wyborski",
      "author_url": "",
      "post_date": "2019-04-16T15:24:24",
      "content": "<p>same for me. I do the predictions over the whole validation set (20% holdout) and afterwards i calculate the lwlrap. CV is about 0.73, i don't know exactly because i didn't do k-fold, just the final run over all the data with similar settings, the LB was then 0.505</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 516317,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2019-04-14T01:43:47.673000",
      "content": "<p>I have just made a few submissions and already seeing the similar relationships between cv and lb.</p>\n\n<ul>\n<li>cv 0.683, lb 0.436</li>\n<li>cv 0.739, lb 0.505</li>\n</ul>\n\n<p>I only use train_curated and 5-folds like op.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 515871,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2019-04-13T08:35:38.903000",
      "content": "<p>Yes, we also observe huge difference around 0.2 between CV and LB, but luckily its correlated, so when CV goes up LB goes up too.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 515891,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-13T09:16:47.103000",
          "content": "<p>Yeah, I also observe the correlation. However, such a huge gap is a bit disturbing, since it's stated that the train and test data have to be very similar.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 515759,
      "author_name": "James Requa",
      "author_url": "",
      "post_date": "2019-04-13T04:58:59.637000",
      "content": "<p>Yes I am seeing the same. On my local CV I get on average 0.76 lwlrap score which is close to 0.2 difference from LB. I used the lwlrap implementation that was provided and computed on the entire validation set.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 515766,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-13T05:19:29.890000",
          "content": "<p><a href=\"/eduardofonseca\">@eduardofonseca</a> if he has any thoughts.</p>\n\n<p>And just to be clear, your cross-validation set is a randomized subset of the curated train set, right (there are patterns in the order of rows in the CSV)? And you have all classes equally represented in your train set and your validation set? Are you sure you are not overfitting to your (now smaller) train set because you might be using a large model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 515674,
      "author_name": "Manoj Plakal",
      "author_url": "",
      "post_date": "2019-04-12T23:45:55.683000",
      "content": "<p>By CV, do you mean lwlrap computed on a held-out cross-validation set taken from the curated train set? Are you using the lwlrap implementation that we have provided, and are you computing it over the entire validation set? There was another thread about lwlrap where someone was computing lwlrap per batch and taking the average, which is not the way to combine lwlraps.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 515821,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-13T07:14:07.900000",
          "content": "<p>Sure, I'm doing CV on the curated train set. I'm using the provided lwlrap implementation. The metric is computed on the entire validation set in a single pass.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 515893,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-13T09:19:10.337000",
          "content": "<p>I've just tested my approach on a holdout set. The results are very close to my CV results and the gap is still around 0.2.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 518604,
      "author_name": "qrfaction",
      "author_url": "",
      "post_date": "2019-04-17T14:10:37.730000",
      "content": "<p>5 folds   cv 0.803    lb  0.62</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 515709,
      "author_name": "AkiraSosa",
      "author_url": "",
      "post_date": "2019-04-13T01:59:54.457000",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a> Yes, it's same with me. As far as I see the statistics, train and test set are similar. So, I wonder how they are annotated, e.g. by different person and splitted?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 515851,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2019-04-13T08:14:36.920000",
      "content": "<p>When your CV improves, does your LB score improve too?</p>",
      "votes": -1,
      "replies": [
        {
          "id": 515868,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-13T08:29:00.057000",
          "content": "<p>Yes. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 523696,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-04-26T18:59:15.313000",
      "content": "<p>Well, my last observations are that LB-CV gap is model-specific and also that as local score improves, the gap seems to become smaller. Note, however, that these facts may be specific to my approach and may not work for you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 521393,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-04-22T21:18:42.897000",
      "content": "<p>Has anyone managed to find a way to close the gap? :) It seems that training on the noisy part of the train set helps a bit, but the difference is still around 0.15...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 521683,
          "author_name": "YLChan",
          "author_url": "",
          "post_date": "2019-04-23T09:06:18",
          "content": "<p>Same here, using both noisy and curated data, 5Fold CV\n0.747 CV  &lt;-&gt; 0.617 LB\na 0.13 difference</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 516143,
      "author_name": "Eduardo Fonseca",
      "author_url": "",
      "post_date": "2019-04-13T17:17:23.353000",
      "content": "<p>Hi,</p>\n\n<p>just to confirm or deny some facts:</p>\n\n<ul>\n<li>the curated train set and the test set come from the currently under development FSD, yes. </li>\n<li>both are labeled manually, yes. Several people were involved in the annotation of both sets.</li>\n</ul>\n\n<p><a href=\"/ddanevskyi\">@ddanevskyi</a>, you mention:</p>\n\n<blockquote>\n  <p>it's stated that the train and test data have to be very similar.</p>\n</blockquote>\n\n<p>really? I'm afraid I don’t remember that being stated anywhere... Can you point us where it is stated?</p>\n\n<p>I am not sure I see the problem here. The fact that CV scores are higher than test scores can be just a feature, and it is not uncommon in recognition tasks.</p>\n\n<p>thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 516153,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-13T17:33:31.233000",
          "content": "<p>Oh, maybe I got it wrong, but for some reason, I assumed that since train/test come from the same dataset and both were manually labeled, their distributions should be the same. If it's not the case, there is indeed no problem. Thank you for clarification.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 515790,
      "author_name": "Manoj Plakal",
      "author_url": "",
      "post_date": "2019-04-13T06:16:00.507000",
      "content": "<p>As an alternative data point, I tried evaluating one of our baseline models that had been trained purely on the noisy data and got comparable lwlraps when evaluated on the test set and an equal-sized random subset of the curated train set (curated train was lower by ~0.004). I would double check that all classes are equally represented in your train and validation sets, and that your model has not gotten large enough that it overfits to your train and validation sets.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 515823,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-13T07:21:04.370000",
          "content": "<p>It seems that almost everyone has this issue. My local scores are close to what <a href=\"/jamesrequa\">@jamesrequa</a> has reported and my LB score is also very close to his score. I really doubt that we have the same bug/leak since our pipelines are probably very different.</p>\n\n<p>What is the CV score for your baseline model? Currently, I'm interested in only the curated set performance.</p>\n\n<p>Thank you in advance!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 516117,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-13T16:25:35.317000",
          "content": "<p>Our submitted baseline was trained using an internal copy of the train and test sets and was validated using the entire test set. I can re-run the best model using only the released data from Kaggle. Can you share the train and validation splits you used?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 516157,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-13T17:41:37.463000",
          "content": "<p>Sure, I used 5 fold cv with default KFold from <code>sklearn</code> with <code>random_state=42</code>. Currently, I'm using only curated data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522246,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-24T05:05:26.103000",
          "content": "<p>To close out this sub-thread, as a rough point of comparison, I tried training the baseline with the same hyperparameters and number of epochs as the original submission, but with a 80%-20% train-validation split chosen by your choice of fold selector: sklearn's KFold with shuffle=True and random_state=42 (note that I tried just one of the 5 splits). The validation lwlrap was ~0.705 compared with public test lwlrap of ~0.537 and full (public + private) test lwlrap of ~0.544.</p>\n\n<p>I wouldn't get too alarmed by the difference in the lwlraps. As Eduardo said, train and test are taken from the same source of data and have been annotated in similar ways, but they are not identical in all respects. Perhaps you can figure out how to make use of the extra noisy data to your model's advantage. Also note that lwlrap is designed to be a weighted average of per-class lwlraps, so it might be time to look into how your models are doing for individual classes in order to raise your overall rank.</p>\n\n<p>BTW, you didn't mention it and the KFold seems to use a default shuffle of False, but I hope you are shuffling when making your folds? The lines in the training CSV are not in completely random order.</p>\n\n<p>Also worth checking out: splitting your data blindly may increase bias given that this is a multi-label problem. Perhaps take a look at <a href=\"http://scikit.ml/stratification.html\">http://scikit.ml/stratification.html</a> (although I don't have much experience with it myself)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 522310,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-24T07:54:02.883000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "515655": "The data tab says that the curated train set and the test set came from the same dataset and both were labeled manually. However, I see that most public kernels and also my own pipeline (which differs from public kernels) have very high CV-LB difference that seems to be around 0.2. \n\nAre you getting the same results?",
    "521710": "5folds, CV:0.67, LB:0.635<br>\nUpdate: 5folds, CV:0.685, LB:0.665",
    "518738": "@eduardofonseca @plakal\n\nHello, \n\nJust a little clarification. Can we expect that the stage2 test data will have the same distribution as the stage1 test data? Thanks.",
    "517823": "same for me. I do the predictions over the whole validation set (20% holdout) and afterwards i calculate the lwlrap. CV is about 0.73, i don't know exactly because i didn't do k-fold, just the final run over all the data with similar settings, the LB was then 0.505",
    "516317": "I have just made a few submissions and already seeing the similar relationships between cv and lb.\n\n- cv 0.683, lb 0.436\n- cv 0.739, lb 0.505\n\nI only use train_curated and 5-folds like op.",
    "515871": "Yes, we also observe huge difference around 0.2 between CV and LB, but luckily its correlated, so when CV goes up LB goes up too.",
    "515759": "Yes I am seeing the same. On my local CV I get on average 0.76 lwlrap score which is close to 0.2 difference from LB. I used the lwlrap implementation that was provided and computed on the entire validation set.",
    "515674": "By CV, do you mean lwlrap computed on a held-out cross-validation set taken from the curated train set? Are you using the lwlrap implementation that we have provided, and are you computing it over the entire validation set? There was another thread about lwlrap where someone was computing lwlrap per batch and taking the average, which is not the way to combine lwlraps.",
    "518604": "5 folds   cv 0.803    lb  0.62",
    "515709": "@ddanevskyi Yes, it's same with me. As far as I see the statistics, train and test set are similar. So, I wonder how they are annotated, e.g. by different person and splitted?",
    "515851": "When your CV improves, does your LB score improve too?",
    "523696": "Well, my last observations are that LB-CV gap is model-specific and also that as local score improves, the gap seems to become smaller. Note, however, that these facts may be specific to my approach and may not work for you.",
    "521393": "Has anyone managed to find a way to close the gap? :) It seems that training on the noisy part of the train set helps a bit, but the difference is still around 0.15...",
    "516143": "Hi,\n\njust to confirm or deny some facts:\n\n- the curated train set and the test set come from the currently under development FSD, yes. \n- both are labeled manually, yes. Several people were involved in the annotation of both sets.\n\n@ddanevskyi, you mention:\n&gt; it's stated that the train and test data have to be very similar.\n\nreally? I'm afraid I don’t remember that being stated anywhere... Can you point us where it is stated?\n\nI am not sure I see the problem here. The fact that CV scores are higher than test scores can be just a feature, and it is not uncommon in recognition tasks.\n\nthanks!",
    "515790": "As an alternative data point, I tried evaluating one of our baseline models that had been trained purely on the noisy data and got comparable lwlraps when evaluated on the test set and an equal-sized random subset of the curated train set (curated train was lower by ~0.004). I would double check that all classes are equally represented in your train and validation sets, and that your model has not gotten large enough that it overfits to your train and validation sets.",
    "522310": ""
  }
}