{
  "id": 308027,
  "title": "Question about CV",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/308027",
  "author_name": "",
  "post_date": "2022-02-16T19:06:43.872862400Z",
  "votes": 12,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>After reading most of the work done by the top scorers I still do not understand the way people are cross validating.</p>\n<p>Based on what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614</a> described in his discussion, given N folds, you build N independent models with the same architecture and then combine the outputs of their respective test sets into an OOF file, which you then compare to the groundtruth to create your CV metric.</p>\n<p>But what do you do in inference? Do you then keep all N models and ensemble them? Or do you just use CV to validate the performance of the model itself and then train a fresh model on the whole dataset?</p>\n<p>Thank you in advance!</p>",
  "messages": [
    {
      "id": "1693584",
      "postDate": "02/16/2022 19:06:43",
      "content": "<p>Hi everyone,</p>\n<p>After reading most of the work done by the top scorers I still do not understand the way people are cross validating.</p>\n<p>Based on what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614</a> described in his discussion, given N folds, you build N independent models with the same architecture and then combine the outputs of their respective test sets into an OOF file, which you then compare to the groundtruth to create your CV metric.</p>\n<p>But what do you do in inference? Do you then keep all N models and ensemble them? Or do you just use CV to validate the performance of the model itself and then train a fresh model on the whole dataset?</p>\n<p>Thank you in advance!</p>",
      "rawMarkdown": "Hi everyone,\n\nAfter reading most of the work done by the top scorers I still do not understand the way people are cross validating.\n\nBased on what @cdeotte https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614 described in his discussion, given N folds, you build N independent models with the same architecture and then combine the outputs of their respective test sets into an OOF file, which you then compare to the groundtruth to create your CV metric.\n\nBut what do you do in inference? Do you then keep all N models and ensemble them? Or do you just use CV to validate the performance of the model itself and then train a fresh model on the whole dataset?\n\nThank you in advance!",
      "votes": null
    },
    {
      "id": "1693623",
      "postDate": "02/16/2022 20:13:47",
      "content": "<p>This is a great question. There are two main approaches.</p>\n<ul>\n<li>Train N Fold models and ensemble (i.e. average) the N test predictions</li>\n<li>After discovering the best hyperparameters from CV, train 1 model using 100% data and infer test with 1 model</li>\n</ul>\n<p>In approach one, let's say you have 5-Fold models. Then you split the train data into 5 folds and train 5 models. Each model trains with 80% data and infers the other 20% data. When we combine all the 5x 20% train predictions, we have 1 prediction for each train row (i.e. movie frame). We compute our CV score on this (called OOF). During inference each of the 5 fold models predicts the test data. Then we have 5 predictions for each test row (movie frame). We ensemble these with WBFT.</p>\n<p>In approach two, we start with 5-fold models. We learn that learning rate 0.01 is best with batch size 8 and epochs 7. Afterward we ignore our 5-fold models. Now we train another model (a 6th model), that uses 100% train data. We train it with learning rate 0.01, batch size 8, and 7 epochs. Afterward we use this 1 model to predict the test data. We have 1 prediction for each test frame. This is our submission.</p>\n<p>(A third approach is to use \"approach one\", and then just infer 3 of the 5 fold models because of time limitation or something. Then we have 3 predictions for each test frame and use WBFT. And a fourth approach is to train multiple models with 100% train and then average them)</p>",
      "rawMarkdown": "This is a great question. There are two main approaches.\n* Train N Fold models and ensemble (i.e. average) the N test predictions\n* After discovering the best hyperparameters from CV, train 1 model using 100% data and infer test with 1 model\n\nIn approach one, let's say you have 5-Fold models. Then you split the train data into 5 folds and train 5 models. Each model trains with 80% data and infers the other 20% data. When we combine all the 5x 20% train predictions, we have 1 prediction for each train row (i.e. movie frame). We compute our CV score on this (called OOF). During inference each of the 5 fold models predicts the test data. Then we have 5 predictions for each test row (movie frame). We ensemble these with WBFT.\n\nIn approach two, we start with 5-fold models. We learn that learning rate 0.01 is best with batch size 8 and epochs 7. Afterward we ignore our 5-fold models. Now we train another model (a 6th model), that uses 100% train data. We train it with learning rate 0.01, batch size 8, and 7 epochs. Afterward we use this 1 model to predict the test data. We have 1 prediction for each test frame. This is our submission.\n\n(A third approach is to use \"approach one\", and then just infer 3 of the 5 fold models because of time limitation or something. Then we have 3 predictions for each test frame and use WBFT. And a fourth approach is to train multiple models with 100% train and then average them)",
      "votes": null
    },
    {
      "id": "1693762",
      "postDate": "02/16/2022 23:23:49",
      "content": "<p>Thank you so much for taking the time to explain it! It makes a lot of sense.</p>\n<p>We encountered some issues during the competition due to not having a robust way of evaluating our models and this will totally help us towards future competitions.</p>\n<p>PD: Also, thank you for your contributions to the community. You've are a huge inspiration and I'm learning a lot from you!</p>",
      "rawMarkdown": "Thank you so much for taking the time to explain it! It makes a lot of sense.\n\nWe encountered some issues during the competition due to not having a robust way of evaluating our models and this will totally help us towards future competitions.\n\nPD: Also, thank you for your contributions to the community. You've are a huge inspiration and I'm learning a lot from you!",
      "votes": null
    },
    {
      "id": "1693846",
      "postDate": "02/17/2022 02:15:31",
      "content": "<p>The same appreciation for your contributions from me :). </p>\n<p>I just want to clarify what I understood regarding OOF. Will OOF evaluation (CV score) be identical to the average of evaluation scores (i.e., evaluating with 5fold hold-out test sets) at train phase assuming best model saved at train phase and no TTA at CV score calculation phase?</p>",
      "rawMarkdown": "The same appreciation for your contributions from me :). \n\nI just want to clarify what I understood regarding OOF. Will OOF evaluation (CV score) be identical to the average of evaluation scores (i.e., evaluating with 5fold hold-out test sets) at train phase assuming best model saved at train phase and no TTA at CV score calculation phase?",
      "votes": null
    },
    {
      "id": "1693857",
      "postDate": "02/17/2022 02:45:09",
      "content": "<p>Yes. Generally speaking, if the train data comes from the same distribution as test data (and both are large enough), then CV score should match LB score.</p>",
      "rawMarkdown": "Yes. Generally speaking, if the train data comes from the same distribution as test data (and both are large enough), then CV score should match LB score.",
      "votes": null
    },
    {
      "id": "1694352",
      "postDate": "02/17/2022 11:07:22",
      "content": "<p>One more question.</p>\n<p>How do you choose the best splitting strategy? Do you use a fixed baseline model, use it to compute the OOF CV score over different splitting strategies (by video, creating sequences, etc) and choose the one that maximizes this CV score? </p>",
      "rawMarkdown": "One more question.\n\nHow do you choose the best splitting strategy? Do you use a fixed baseline model, use it to compute the OOF CV score over different splitting strategies (by video, creating sequences, etc) and choose the one that maximizes this CV score?",
      "votes": null
    },
    {
      "id": "1694585",
      "postDate": "02/17/2022 14:54:23",
      "content": "<p>We pick a strategy that <strong>mimics</strong> the relationship between train and test data. If train and test are just random subsets of some larger population then <code>random K-Fold</code> works. However, if there is a difference between train and test, like for example test uses different videos than train, then we should set up our CV to <strong>mimic</strong> this. We should put different videos in different CV folds.</p>\n<p>Another example might be predicting customer credit scores. If train data and test data have different customers, then we want to keep all rows pertaining to one customer within a single fold. To keep things in one fold we use <code>group K-Fold</code>.</p>\n<p>If the positive target is rare like 1% (or we have multi label and/or multi class and there are rare targets &lt;1%), then we should use <code>stratified K-Fold</code> to make sure that each validation fold includes some positive targets.</p>\n<p>After choosing our CV, the final test is to submit to LB. If our CV score is approximately equal to our LB score. And if our LB scores goes up when our CV score goes up, then we did a good job. Note that sometimes, the test data has some mystery so the best we can do is have LB go up and down when CV goes up and down</p>\n<p>Setting up a reliable CV scheme is the <strong>most important</strong> aspect of data science!</p>",
      "rawMarkdown": "We pick a strategy that **mimics** the relationship between train and test data. If train and test are just random subsets of some larger population then `random K-Fold` works. However, if there is a difference between train and test, like for example test uses different videos than train, then we should set up our CV to **mimic** this. We should put different videos in different CV folds.\n\nAnother example might be predicting customer credit scores. If train data and test data have different customers, then we want to keep all rows pertaining to one customer within a single fold. To keep things in one fold we use `group K-Fold`.\n\nIf the positive target is rare like 1% (or we have multi label and/or multi class and there are rare targets <1%), then we should use `stratified K-Fold` to make sure that each validation fold includes some positive targets.\n\nAfter choosing our CV, the final test is to submit to LB. If our CV score is approximately equal to our LB score. And if our LB scores goes up when our CV score goes up, then we did a good job. Note that sometimes, the test data has some mystery so the best we can do is have LB go up and down when CV goes up and down\n\nSetting up a reliable CV scheme is the **most important** aspect of data science!",
      "votes": null
    },
    {
      "id": "1695098",
      "postDate": "02/18/2022 00:21:16",
      "content": "<p>Great explanation!</p>\n<p>Would you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?</p>\n<p>I was surprised seeing the winners of this competition use the technique of training on all data - that could lead to over-fitting since you don't have a local CV score for the final model right?</p>\n<p>Maybe I'm overestimating the risk - and the advantage of having your inference time basically divided by the number of folds is really valuable</p>\n<p>Anyways they had to be really strict with not optimizing their solutions towards the public leaderboard which seems hard when you don't have a local CV for the final ensemble!</p>",
      "rawMarkdown": "Great explanation!\n\nWould you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?\n\nI was surprised seeing the winners of this competition use the technique of training on all data - that could lead to over-fitting since you don't have a local CV score for the final model right?\n\nMaybe I'm overestimating the risk - and the advantage of having your inference time basically divided by the number of folds is really valuable\n\nAnyways they had to be really strict with not optimizing their solutions towards the public leaderboard which seems hard when you don't have a local CV for the final ensemble!",
      "votes": null
    },
    {
      "id": "1695107",
      "postDate": "02/18/2022 00:43:05",
      "content": "<blockquote>\n  <p>Would you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?</p>\n</blockquote>\n<p>Using 100% data usually works great and boosts LB in all Kaggle competitions. From the K-Fold CV, we know what epoch to stop at (and what hyperparameters to use), so it won't overfit because CV didn't overfit. </p>\n<p>But, sometimes i don't use 100% when the local fold validation score fluctuate a lot from epoch to epoch. For example in this comp and CommonLit comp, my local validation score changes a lot from epoch to epoch. In this case, i don't trust using 100% (because i don't know if the stopping epoch will be the right epoch). </p>\n<p>Instead I pick a large fold number like 10 or 20. Then you just train 5 folds. Each fold will be trained on 90% or 95% (respectively) of train data. (Note you can use 20-Fold, but only train 5 of the 20 folds and submit 5 folds to LB. This is what our team did in CommonLit comp). </p>\n<p>Also note when you train with 100% data, you can still do this 5 times and submit an ensemble of 5 models. The LB score for NN's will usually also increase by training the same NN repeatedly and using different seeds to initialize the layer weights. Also training involves random augmentations and random batches which guarantee that two trained NN will never be the same.</p>",
      "rawMarkdown": ">Would you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?\n\nUsing 100% data usually works great and boosts LB in all Kaggle competitions. From the K-Fold CV, we know what epoch to stop at (and what hyperparameters to use), so it won't overfit because CV didn't overfit. \n\nBut, sometimes i don't use 100% when the local fold validation score fluctuate a lot from epoch to epoch. For example in this comp and CommonLit comp, my local validation score changes a lot from epoch to epoch. In this case, i don't trust using 100% (because i don't know if the stopping epoch will be the right epoch). \n\nInstead I pick a large fold number like 10 or 20. Then you just train 5 folds. Each fold will be trained on 90% or 95% (respectively) of train data. (Note you can use 20-Fold, but only train 5 of the 20 folds and submit 5 folds to LB. This is what our team did in CommonLit comp). \n\nAlso note when you train with 100% data, you can still do this 5 times and submit an ensemble of 5 models. The LB score for NN's will usually also increase by training the same NN repeatedly and using different seeds to initialize the layer weights. Also training involves random augmentations and random batches which guarantee that two trained NN will never be the same.",
      "votes": null
    },
    {
      "id": "1695135",
      "postDate": "02/18/2022 01:24:39",
      "content": "<p>Wow amazing!<br>\nIt makes so much more sense now</p>\n<p>To test if I got it right: <br>\nn fold-model ensembles (any ensemble that generates OOF predictions for the whole dataset) are only optimal when you make use of the OOF predictions. Otherwise you can use much more of the training data (even 100%) for your models as long as you are confident in not overfitting</p>",
      "rawMarkdown": "Wow amazing!\nIt makes so much more sense now\n\nTo test if I got it right: \nn fold-model ensembles (any ensemble that generates OOF predictions for the whole dataset) are only optimal when you make use of the OOF predictions. Otherwise you can use much more of the training data (even 100%) for your models as long as you are confident in not overfitting",
      "votes": null
    },
    {
      "id": "1695139",
      "postDate": "02/18/2022 01:31:20",
      "content": "<p>Exactly. If you want a dataframe of OOF, then you need to use K-Fold because the OOF are predictions on rows of train that are <strong>not</strong> included in your training process. But when it is time to submit to LB, you might as well use all 100% train if you can.</p>",
      "rawMarkdown": "Exactly. If you want a dataframe of OOF, then you need to use K-Fold because the OOF are predictions on rows of train that are **not** included in your training process. But when it is time to submit to LB, you might as well use all 100% train if you can.",
      "votes": null
    },
    {
      "id": "1695176",
      "postDate": "02/18/2022 02:02:02",
      "content": "<p>Thank you for your insights! Great job on the Grandmaster Series as well. I really enjoyed it</p>",
      "rawMarkdown": "Thank you for your insights! Great job on the Grandmaster Series as well. I really enjoyed it",
      "votes": null
    },
    {
      "id": "1695267",
      "postDate": "02/18/2022 03:53:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Your comment about ensembling with different nets made me remember the theoretical justification of attention heads in Transformers: they are supposed to capture different patterns to hopefully predict better.</p>\n<p>I usually attribute this to the infinite number phenomena in ML: there are must be infinite solutions to the same problem, so there must be some solutions better than others, by some criteria.</p>",
      "rawMarkdown": "Hi @cdeotte Your comment about ensembling with different nets made me remember the theoretical justification of attention heads in Transformers: they are supposed to capture different patterns to hopefully predict better.\n\nI usually attribute this to the infinite number phenomena in ML: there are must be infinite solutions to the same problem, so there must be some solutions better than others, by some criteria.",
      "votes": null
    },
    {
      "id": "1695980",
      "postDate": "02/18/2022 14:17:59",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I've seen you mention this idea is many posts over the last year or so.</p>\n<p>One thing I don't understand:<br>\nif you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?</p>",
      "rawMarkdown": "cdeotte, I've seen you mention this idea is many posts over the last year or so.\n\nOne thing I don't understand:\nif you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?",
      "votes": null
    },
    {
      "id": "1696070",
      "postDate": "02/18/2022 15:10:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> I want to understand you, why you think more training steps?</p>",
      "rawMarkdown": "Hi @jacob34 I want to understand you, why you think more training steps?",
      "votes": null
    },
    {
      "id": "1696085",
      "postDate": "02/18/2022 15:25:55",
      "content": "<blockquote>\n  <p>if you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?</p>\n</blockquote>\n<p>Yes and no. The \"extra training steps\" are using data that the model has never seen. Therefore if the train data has large variety, then the number of folds that worked for 5-Fold will also work for 100%.</p>\n<p>However if all the train data is very similar, then you are correct and sometimes, we need to decrease the number of epochs by 20% when going from 5-Fold to 100% (to compensate for the extra 25%). Another trick is to use 20-Fold (instead of 5-Fold) then when you go to 100% it is only 5% more (not 25% more). </p>\n<p>Lastly, i usually do this trick. If 5-Fold took 10 epochs, then i will train 100% with 10 epoch, 9 epoch, and 8 epoch and submit all 3 to the LB. If the LB is somewhat trustworthy, we can use the best one.</p>",
      "rawMarkdown": ">if you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?\n\nYes and no. The \"extra training steps\" are using data that the model has never seen. Therefore if the train data has large variety, then the number of folds that worked for 5-Fold will also work for 100%.\n\nHowever if all the train data is very similar, then you are correct and sometimes, we need to decrease the number of epochs by 20% when going from 5-Fold to 100% (to compensate for the extra 25%). Another trick is to use 20-Fold (instead of 5-Fold) then when you go to 100% it is only 5% more (not 25% more). \n\nLastly, i usually do this trick. If 5-Fold took 10 epochs, then i will train 100% with 10 epoch, 9 epoch, and 8 epoch and submit all 3 to the LB. If the LB is somewhat trustworthy, we can use the best one.",
      "votes": null
    },
    {
      "id": "1696260",
      "postDate": "02/18/2022 17:44:01",
      "content": "<p>That clears that up, then.<br>\nThanks!</p>",
      "rawMarkdown": "That clears that up, then.\nThanks!",
      "votes": null
    },
    {
      "id": "1698412",
      "postDate": "02/20/2022 11:17:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I have questions. During this competition, the relationship between CV and LB was non-linear. So, I realized that I was calculate it wrong but, I don't know how to calculate the CV correctly. In object detection competition like this, How do I calculate CV ? And what methods can I try to solve this relationship of non-linear problem ? </p>",
      "rawMarkdown": "Hi @cdeotte, I have questions. During this competition, the relationship between CV and LB was non-linear. So, I realized that I was calculate it wrong but, I don't know how to calculate the CV correctly. In object detection competition like this, How do I calculate CV ? And what methods can I try to solve this relationship of non-linear problem ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1693623,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/16/2022 20:13:47",
      "content": "<p>This is a great question. There are two main approaches.</p>\n<ul>\n<li>Train N Fold models and ensemble (i.e. average) the N test predictions</li>\n<li>After discovering the best hyperparameters from CV, train 1 model using 100% data and infer test with 1 model</li>\n</ul>\n<p>In approach one, let's say you have 5-Fold models. Then you split the train data into 5 folds and train 5 models. Each model trains with 80% data and infers the other 20% data. When we combine all the 5x 20% train predictions, we have 1 prediction for each train row (i.e. movie frame). We compute our CV score on this (called OOF). During inference each of the 5 fold models predicts the test data. Then we have 5 predictions for each test row (movie frame). We ensemble these with WBFT.</p>\n<p>In approach two, we start with 5-fold models. We learn that learning rate 0.01 is best with batch size 8 and epochs 7. Afterward we ignore our 5-fold models. Now we train another model (a 6th model), that uses 100% train data. We train it with learning rate 0.01, batch size 8, and 7 epochs. Afterward we use this 1 model to predict the test data. We have 1 prediction for each test frame. This is our submission.</p>\n<p>(A third approach is to use \"approach one\", and then just infer 3 of the 5 fold models because of time limitation or something. Then we have 3 predictions for each test frame and use WBFT. And a fourth approach is to train multiple models with 100% train and then average them)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1693762,
          "author_name": "dauriel",
          "author_url": "",
          "post_date": "02/16/2022 23:23:49",
          "content": "<p>Thank you so much for taking the time to explain it! It makes a lot of sense.</p>\n<p>We encountered some issues during the competition due to not having a robust way of evaluating our models and this will totally help us towards future competitions.</p>\n<p>PD: Also, thank you for your contributions to the community. You've are a huge inspiration and I'm learning a lot from you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695098,
          "author_name": "dollofcuty",
          "author_url": "",
          "post_date": "02/18/2022 00:21:16",
          "content": "<p>Great explanation!</p>\n<p>Would you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?</p>\n<p>I was surprised seeing the winners of this competition use the technique of training on all data - that could lead to over-fitting since you don't have a local CV score for the final model right?</p>\n<p>Maybe I'm overestimating the risk - and the advantage of having your inference time basically divided by the number of folds is really valuable</p>\n<p>Anyways they had to be really strict with not optimizing their solutions towards the public leaderboard which seems hard when you don't have a local CV for the final ensemble!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695107,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2022 00:43:05",
          "content": "<blockquote>\n  <p>Would you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?</p>\n</blockquote>\n<p>Using 100% data usually works great and boosts LB in all Kaggle competitions. From the K-Fold CV, we know what epoch to stop at (and what hyperparameters to use), so it won't overfit because CV didn't overfit. </p>\n<p>But, sometimes i don't use 100% when the local fold validation score fluctuate a lot from epoch to epoch. For example in this comp and CommonLit comp, my local validation score changes a lot from epoch to epoch. In this case, i don't trust using 100% (because i don't know if the stopping epoch will be the right epoch). </p>\n<p>Instead I pick a large fold number like 10 or 20. Then you just train 5 folds. Each fold will be trained on 90% or 95% (respectively) of train data. (Note you can use 20-Fold, but only train 5 of the 20 folds and submit 5 folds to LB. This is what our team did in CommonLit comp). </p>\n<p>Also note when you train with 100% data, you can still do this 5 times and submit an ensemble of 5 models. The LB score for NN's will usually also increase by training the same NN repeatedly and using different seeds to initialize the layer weights. Also training involves random augmentations and random batches which guarantee that two trained NN will never be the same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695135,
          "author_name": "dollofcuty",
          "author_url": "",
          "post_date": "02/18/2022 01:24:39",
          "content": "<p>Wow amazing!<br>\nIt makes so much more sense now</p>\n<p>To test if I got it right: <br>\nn fold-model ensembles (any ensemble that generates OOF predictions for the whole dataset) are only optimal when you make use of the OOF predictions. Otherwise you can use much more of the training data (even 100%) for your models as long as you are confident in not overfitting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695139,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2022 01:31:20",
          "content": "<p>Exactly. If you want a dataframe of OOF, then you need to use K-Fold because the OOF are predictions on rows of train that are <strong>not</strong> included in your training process. But when it is time to submit to LB, you might as well use all 100% train if you can.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695176,
          "author_name": "dollofcuty",
          "author_url": "",
          "post_date": "02/18/2022 02:02:02",
          "content": "<p>Thank you for your insights! Great job on the Grandmaster Series as well. I really enjoyed it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695267,
          "author_name": "maulberto3",
          "author_url": "",
          "post_date": "02/18/2022 03:53:33",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Your comment about ensembling with different nets made me remember the theoretical justification of attention heads in Transformers: they are supposed to capture different patterns to hopefully predict better.</p>\n<p>I usually attribute this to the infinite number phenomena in ML: there are must be infinite solutions to the same problem, so there must be some solutions better than others, by some criteria.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1695980,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "02/18/2022 14:17:59",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I've seen you mention this idea is many posts over the last year or so.</p>\n<p>One thing I don't understand:<br>\nif you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696070,
          "author_name": "maulberto3",
          "author_url": "",
          "post_date": "02/18/2022 15:10:48",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> I want to understand you, why you think more training steps?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696085,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2022 15:25:55",
          "content": "<blockquote>\n  <p>if you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?</p>\n</blockquote>\n<p>Yes and no. The \"extra training steps\" are using data that the model has never seen. Therefore if the train data has large variety, then the number of folds that worked for 5-Fold will also work for 100%.</p>\n<p>However if all the train data is very similar, then you are correct and sometimes, we need to decrease the number of epochs by 20% when going from 5-Fold to 100% (to compensate for the extra 25%). Another trick is to use 20-Fold (instead of 5-Fold) then when you go to 100% it is only 5% more (not 25% more). </p>\n<p>Lastly, i usually do this trick. If 5-Fold took 10 epochs, then i will train 100% with 10 epoch, 9 epoch, and 8 epoch and submit all 3 to the LB. If the LB is somewhat trustworthy, we can use the best one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1696260,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "02/18/2022 17:44:01",
          "content": "<p>That clears that up, then.<br>\nThanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1693846,
      "author_name": "enddl22",
      "author_url": "",
      "post_date": "02/17/2022 02:15:31",
      "content": "<p>The same appreciation for your contributions from me :). </p>\n<p>I just want to clarify what I understood regarding OOF. Will OOF evaluation (CV score) be identical to the average of evaluation scores (i.e., evaluating with 5fold hold-out test sets) at train phase assuming best model saved at train phase and no TTA at CV score calculation phase?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1693857,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/17/2022 02:45:09",
          "content": "<p>Yes. Generally speaking, if the train data comes from the same distribution as test data (and both are large enough), then CV score should match LB score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1694352,
      "author_name": "dauriel",
      "author_url": "",
      "post_date": "02/17/2022 11:07:22",
      "content": "<p>One more question.</p>\n<p>How do you choose the best splitting strategy? Do you use a fixed baseline model, use it to compute the OOF CV score over different splitting strategies (by video, creating sequences, etc) and choose the one that maximizes this CV score? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1694585,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/17/2022 14:54:23",
          "content": "<p>We pick a strategy that <strong>mimics</strong> the relationship between train and test data. If train and test are just random subsets of some larger population then <code>random K-Fold</code> works. However, if there is a difference between train and test, like for example test uses different videos than train, then we should set up our CV to <strong>mimic</strong> this. We should put different videos in different CV folds.</p>\n<p>Another example might be predicting customer credit scores. If train data and test data have different customers, then we want to keep all rows pertaining to one customer within a single fold. To keep things in one fold we use <code>group K-Fold</code>.</p>\n<p>If the positive target is rare like 1% (or we have multi label and/or multi class and there are rare targets &lt;1%), then we should use <code>stratified K-Fold</code> to make sure that each validation fold includes some positive targets.</p>\n<p>After choosing our CV, the final test is to submit to LB. If our CV score is approximately equal to our LB score. And if our LB scores goes up when our CV score goes up, then we did a good job. Note that sometimes, the test data has some mystery so the best we can do is have LB go up and down when CV goes up and down</p>\n<p>Setting up a reliable CV scheme is the <strong>most important</strong> aspect of data science!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1698412,
          "author_name": "seongwook93",
          "author_url": "",
          "post_date": "02/20/2022 11:17:26",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I have questions. During this competition, the relationship between CV and LB was non-linear. So, I realized that I was calculate it wrong but, I don't know how to calculate the CV correctly. In object detection competition like this, How do I calculate CV ? And what methods can I try to solve this relationship of non-linear problem ? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1693584": "Hi everyone,\n\nAfter reading most of the work done by the top scorers I still do not understand the way people are cross validating.\n\nBased on what @cdeotte https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175614 described in his discussion, given N folds, you build N independent models with the same architecture and then combine the outputs of their respective test sets into an OOF file, which you then compare to the groundtruth to create your CV metric.\n\nBut what do you do in inference? Do you then keep all N models and ensemble them? Or do you just use CV to validate the performance of the model itself and then train a fresh model on the whole dataset?\n\nThank you in advance!",
    "1693623": "This is a great question. There are two main approaches.\n* Train N Fold models and ensemble (i.e. average) the N test predictions\n* After discovering the best hyperparameters from CV, train 1 model using 100% data and infer test with 1 model\n\nIn approach one, let's say you have 5-Fold models. Then you split the train data into 5 folds and train 5 models. Each model trains with 80% data and infers the other 20% data. When we combine all the 5x 20% train predictions, we have 1 prediction for each train row (i.e. movie frame). We compute our CV score on this (called OOF). During inference each of the 5 fold models predicts the test data. Then we have 5 predictions for each test row (movie frame). We ensemble these with WBFT.\n\nIn approach two, we start with 5-fold models. We learn that learning rate 0.01 is best with batch size 8 and epochs 7. Afterward we ignore our 5-fold models. Now we train another model (a 6th model), that uses 100% train data. We train it with learning rate 0.01, batch size 8, and 7 epochs. Afterward we use this 1 model to predict the test data. We have 1 prediction for each test frame. This is our submission.\n\n(A third approach is to use \"approach one\", and then just infer 3 of the 5 fold models because of time limitation or something. Then we have 3 predictions for each test frame and use WBFT. And a fourth approach is to train multiple models with 100% train and then average them)",
    "1693762": "Thank you so much for taking the time to explain it! It makes a lot of sense.\n\nWe encountered some issues during the competition due to not having a robust way of evaluating our models and this will totally help us towards future competitions.\n\nPD: Also, thank you for your contributions to the community. You've are a huge inspiration and I'm learning a lot from you!",
    "1693846": "The same appreciation for your contributions from me :). \n\nI just want to clarify what I understood regarding OOF. Will OOF evaluation (CV score) be identical to the average of evaluation scores (i.e., evaluating with 5fold hold-out test sets) at train phase assuming best model saved at train phase and no TTA at CV score calculation phase?",
    "1693857": "Yes. Generally speaking, if the train data comes from the same distribution as test data (and both are large enough), then CV score should match LB score.",
    "1694352": "One more question.\n\nHow do you choose the best splitting strategy? Do you use a fixed baseline model, use it to compute the OOF CV score over different splitting strategies (by video, creating sequences, etc) and choose the one that maximizes this CV score?",
    "1694585": "We pick a strategy that **mimics** the relationship between train and test data. If train and test are just random subsets of some larger population then `random K-Fold` works. However, if there is a difference between train and test, like for example test uses different videos than train, then we should set up our CV to **mimic** this. We should put different videos in different CV folds.\n\nAnother example might be predicting customer credit scores. If train data and test data have different customers, then we want to keep all rows pertaining to one customer within a single fold. To keep things in one fold we use `group K-Fold`.\n\nIf the positive target is rare like 1% (or we have multi label and/or multi class and there are rare targets <1%), then we should use `stratified K-Fold` to make sure that each validation fold includes some positive targets.\n\nAfter choosing our CV, the final test is to submit to LB. If our CV score is approximately equal to our LB score. And if our LB scores goes up when our CV score goes up, then we did a good job. Note that sometimes, the test data has some mystery so the best we can do is have LB go up and down when CV goes up and down\n\nSetting up a reliable CV scheme is the **most important** aspect of data science!",
    "1695098": "Great explanation!\n\nWould you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?\n\nI was surprised seeing the winners of this competition use the technique of training on all data - that could lead to over-fitting since you don't have a local CV score for the final model right?\n\nMaybe I'm overestimating the risk - and the advantage of having your inference time basically divided by the number of folds is really valuable\n\nAnyways they had to be really strict with not optimizing their solutions towards the public leaderboard which seems hard when you don't have a local CV for the final ensemble!",
    "1695107": ">Would you agree that using multiple folds during inference should be more robust than the model trained on 100% data with parameters of 5-fold models?\n\nUsing 100% data usually works great and boosts LB in all Kaggle competitions. From the K-Fold CV, we know what epoch to stop at (and what hyperparameters to use), so it won't overfit because CV didn't overfit. \n\nBut, sometimes i don't use 100% when the local fold validation score fluctuate a lot from epoch to epoch. For example in this comp and CommonLit comp, my local validation score changes a lot from epoch to epoch. In this case, i don't trust using 100% (because i don't know if the stopping epoch will be the right epoch). \n\nInstead I pick a large fold number like 10 or 20. Then you just train 5 folds. Each fold will be trained on 90% or 95% (respectively) of train data. (Note you can use 20-Fold, but only train 5 of the 20 folds and submit 5 folds to LB. This is what our team did in CommonLit comp). \n\nAlso note when you train with 100% data, you can still do this 5 times and submit an ensemble of 5 models. The LB score for NN's will usually also increase by training the same NN repeatedly and using different seeds to initialize the layer weights. Also training involves random augmentations and random batches which guarantee that two trained NN will never be the same.",
    "1695135": "Wow amazing!\nIt makes so much more sense now\n\nTo test if I got it right: \nn fold-model ensembles (any ensemble that generates OOF predictions for the whole dataset) are only optimal when you make use of the OOF predictions. Otherwise you can use much more of the training data (even 100%) for your models as long as you are confident in not overfitting",
    "1695139": "Exactly. If you want a dataframe of OOF, then you need to use K-Fold because the OOF are predictions on rows of train that are **not** included in your training process. But when it is time to submit to LB, you might as well use all 100% train if you can.",
    "1695176": "Thank you for your insights! Great job on the Grandmaster Series as well. I really enjoyed it",
    "1695267": "Hi @cdeotte Your comment about ensembling with different nets made me remember the theoretical justification of attention heads in Transformers: they are supposed to capture different patterns to hopefully predict better.\n\nI usually attribute this to the infinite number phenomena in ML: there are must be infinite solutions to the same problem, so there must be some solutions better than others, by some criteria.",
    "1695980": "cdeotte, I've seen you mention this idea is many posts over the last year or so.\n\nOne thing I don't understand:\nif you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?",
    "1696070": "Hi @jacob34 I want to understand you, why you think more training steps?",
    "1696085": ">if you use the same amount of epochs on the 100% of the data as you do on 80% of the data, don't you end up taking 25% more training steps?\n\nYes and no. The \"extra training steps\" are using data that the model has never seen. Therefore if the train data has large variety, then the number of folds that worked for 5-Fold will also work for 100%.\n\nHowever if all the train data is very similar, then you are correct and sometimes, we need to decrease the number of epochs by 20% when going from 5-Fold to 100% (to compensate for the extra 25%). Another trick is to use 20-Fold (instead of 5-Fold) then when you go to 100% it is only 5% more (not 25% more). \n\nLastly, i usually do this trick. If 5-Fold took 10 epochs, then i will train 100% with 10 epoch, 9 epoch, and 8 epoch and submit all 3 to the LB. If the LB is somewhat trustworthy, we can use the best one.",
    "1696260": "That clears that up, then.\nThanks!",
    "1698412": "Hi @cdeotte, I have questions. During this competition, the relationship between CV and LB was non-linear. So, I realized that I was calculate it wrong but, I don't know how to calculate the CV correctly. In object detection competition like this, How do I calculate CV ? And what methods can I try to solve this relationship of non-linear problem ?"
  },
  "source": "meta"
}