{
  "id": 77949,
  "title": "cv strategy",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77949",
  "author_name": "",
  "post_date": "2019-01-18T01:57:36.000718600Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I'm curious how everyone is designing their cv. Here's mine.</p>\n\n<p>For each fold, perform threshold optimization after every epoch to get the highest validation f1 score. Use the checkpoint with the highest val f1 score to predict the test set and validation set. Repeat this for all folds. Now, with oof predictions for the entire train set, I run threshold optimization to get the best threshold, and use it on the test predictions. For test predictions, I'm currently averaging the output probabilities of each model then running threshold. However, I'm curious to see if anyone is running threshold on the output probabilities of each individual model then using majority voting to create the final solution.</p>\n\n<p>Current results:\n4-fold CV: 0.681, LB: 0.690\n5-fold CV: 0.677, LB: 0.691</p>",
  "messages": [
    {
      "id": "457732",
      "postDate": "01/18/2019 01:57:36",
      "content": "<p>I'm curious how everyone is designing their cv. Here's mine.</p>\n\n<p>For each fold, perform threshold optimization after every epoch to get the highest validation f1 score. Use the checkpoint with the highest val f1 score to predict the test set and validation set. Repeat this for all folds. Now, with oof predictions for the entire train set, I run threshold optimization to get the best threshold, and use it on the test predictions. For test predictions, I'm currently averaging the output probabilities of each model then running threshold. However, I'm curious to see if anyone is running threshold on the output probabilities of each individual model then using majority voting to create the final solution.</p>\n\n<p>Current results:\n4-fold CV: 0.681, LB: 0.690\n5-fold CV: 0.677, LB: 0.691</p>",
      "rawMarkdown": "I'm curious how everyone is designing their cv. Here's mine.\n\nFor each fold, perform threshold optimization after every epoch to get the highest validation f1 score. Use the checkpoint with the highest val f1 score to predict the test set and validation set. Repeat this for all folds. Now, with oof predictions for the entire train set, I run threshold optimization to get the best threshold, and use it on the test predictions. For test predictions, I'm currently averaging the output probabilities of each model then running threshold. However, I'm curious to see if anyone is running threshold on the output probabilities of each individual model then using majority voting to create the final solution.\n\nCurrent results:\n4-fold CV: 0.681, LB: 0.690\n5-fold CV: 0.677, LB: 0.691",
      "votes": null
    },
    {
      "id": "458717",
      "postDate": "01/20/2019 10:28:41",
      "content": "<p>I think I do exactly the same as you and trust more CV score than LB</p>\n\n<p>CV: 0.6812 LB: 0.690\nCV: 0.68075 LB: 0.691\nCV: 0.67906 LB: 0.693</p>",
      "rawMarkdown": "I think I do exactly the same as you and trust more CV score than LB\n\nCV: 0.6812 LB: 0.690\nCV: 0.68075 LB: 0.691\nCV: 0.67906 LB: 0.693",
      "votes": null
    },
    {
      "id": "458733",
      "postDate": "01/20/2019 11:44:44",
      "content": "<p>I use the same CV strategy and get cv results in the 0.689 range. Does anyone use a different cv strategy and get better results?</p>",
      "rawMarkdown": "I use the same CV strategy and get cv results in the 0.689 range. Does anyone use a different cv strategy and get better results?",
      "votes": null
    },
    {
      "id": "458993",
      "postDate": "01/21/2019 01:08:15",
      "content": "<p>I am also using the usual 5-fold CV to make OOF predictions and test predictions.\nbest CV 0.689, LB 0.691</p>",
      "rawMarkdown": "I am also using the usual 5-fold CV to make OOF predictions and test predictions.\nbest CV 0.689, LB 0.691",
      "votes": null
    },
    {
      "id": "459017",
      "postDate": "01/21/2019 02:56:31",
      "content": "<p>I think you might want to check this great kernel.\n<a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a></p>\n\n<p>Here's the short summary I gathered from the above kernel.\nBasically when we apply K-fold CV and perform averaging on each fold, the prediction we give is the ensemble of K models from each fold. And there are at least two factors affect the performance if an ensemble model, i.e. your LB score.\n1. The CV score of each fold (The performance of each base classifier), the higher the better.\n2. The divergence of each fold, the higher the better.</p>\n\n<p>So you should also check the second factor (the correlation between fold-predictions) in your validation framework. If you want, you can separate some data before K-fold CV as the validation data of the final prediction (the estimation of LB score). </p>",
      "rawMarkdown": "I think you might want to check this great kernel.\nhttps://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\n\nHere's the short summary I gathered from the above kernel.\nBasically when we apply K-fold CV and perform averaging on each fold, the prediction we give is the ensemble of K models from each fold. And there are at least two factors affect the performance if an ensemble model, i.e. your LB score.\n1. The CV score of each fold (The performance of each base classifier), the higher the better.\n2. The divergence of each fold, the higher the better.\n\nSo you should also check the second factor (the correlation between fold-predictions) in your validation framework. If you want, you can separate some data before K-fold CV as the validation data of the final prediction (the estimation of LB score).",
      "votes": null
    },
    {
      "id": "459142",
      "postDate": "01/21/2019 09:13:58",
      "content": "<p>Same approach.\n5-fold CV: 0.6943, LB: 0.699.</p>",
      "rawMarkdown": "Same approach.\n5-fold CV: 0.6943, LB: 0.699.",
      "votes": null
    },
    {
      "id": "460831",
      "postDate": "01/24/2019 13:55:18",
      "content": "<p>Does anyone have theoretical reasoning (or a link) to why K-fold CV is better than training a single larger model?</p>",
      "rawMarkdown": "Does anyone have theoretical reasoning (or a link) to why K-fold CV is better than training a single larger model?",
      "votes": null
    },
    {
      "id": "460848",
      "postDate": "01/24/2019 14:34:11",
      "content": "<p>k-fold creates k different models from different datasets, so each model will give slightly different prediction based on the data/signal it has seen. Therefore when you average the prediction from all K models, you are using bootstrap aggregation which helps your predictions.</p>\n\n<p>Also, with k-fold, you can train on the whole dataset. If you train a single larger model, you must use a holdout validation set (and lose the opportunity to train on it), or use no validation set at all.</p>",
      "rawMarkdown": "k-fold creates k different models from different datasets, so each model will give slightly different prediction based on the data/signal it has seen. Therefore when you average the prediction from all K models, you are using bootstrap aggregation which helps your predictions.\n\nAlso, with k-fold, you can train on the whole dataset. If you train a single larger model, you must use a holdout validation set (and lose the opportunity to train on it), or use no validation set at all.",
      "votes": null
    },
    {
      "id": "460927",
      "postDate": "01/24/2019 18:37:40",
      "content": "<p>I think when you have trained your model with train/dev and chosen suitable parameters, you can just train the same network again using 100% of your train+dev data? It shouldnt overfit, if it didn't overfit when using train/dev split and now having even more training data for same model and nr of epochs.</p>\n\n<p>In K-fold, instead of one model using 100% training data you will have 5 models using for example 80% of training data, so all of them should individually be weaker than one full model?</p>",
      "rawMarkdown": "I think when you have trained your model with train/dev and chosen suitable parameters, you can just train the same network again using 100% of your train+dev data? It shouldnt overfit, if it didn't overfit when using train/dev split and now having even more training data for same model and nr of epochs.\n\nIn K-fold, instead of one model using 100% training data you will have 5 models using for example 80% of training data, so all of them should individually be weaker than one full model?",
      "votes": null
    },
    {
      "id": "460930",
      "postDate": "01/24/2019 18:47:39",
      "content": "<ol>\n<li><p>You're right, if you have good parameter selection you can retrain on train+dev data in my opinion, as long as you have the correct round/epoch to stop training. And because it sees 100% of the data, it will fit closest to all that 100% of the data since that's what it's optimizing for! Kaggle grandmaster silogram recommends that you actually multiply the epochs by 1.1 when you are training on full dataset instead.</p></li>\n<li><p>The power of K Fold, in my opinion, is its built-in bootstrap aggregation. Using only 80% of the data for 5 models, then averaging them all, makes the final prediction more rugged, which helps prevent overfitting. Actually, in my opinion, the best CV strategy is to combine both. You should find the best parameters for a model using different feature representations, then combine all of this into a final prediction. I think more people would do this, but it's really difficult to \"find the best parameters for a model\" so that is why I and majority of people use K Fold. The other way also takes a lot longer too since you are training K times on full dataset instead of K times on 80%.</p></li>\n</ol>",
      "rawMarkdown": "1. You're right, if you have good parameter selection you can retrain on train+dev data in my opinion, as long as you have the correct round/epoch to stop training. And because it sees 100% of the data, it will fit closest to all that 100% of the data since that's what it's optimizing for! Kaggle grandmaster silogram recommends that you actually multiply the epochs by 1.1 when you are training on full dataset instead.\n\n2. The power of K Fold, in my opinion, is its built-in bootstrap aggregation. Using only 80% of the data for 5 models, then averaging them all, makes the final prediction more rugged, which helps prevent overfitting. Actually, in my opinion, the best CV strategy is to combine both. You should find the best parameters for a model using different feature representations, then combine all of this into a final prediction. I think more people would do this, but it's really difficult to \"find the best parameters for a model\" so that is why I and majority of people use K Fold. The other way also takes a lot longer too since you are training K times on full dataset instead of K times on 80%.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 458717,
      "author_name": "mnikita",
      "author_url": "",
      "post_date": "01/20/2019 10:28:41",
      "content": "<p>I think I do exactly the same as you and trust more CV score than LB</p>\n\n<p>CV: 0.6812 LB: 0.690\nCV: 0.68075 LB: 0.691\nCV: 0.67906 LB: 0.693</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 458733,
      "author_name": "bkkaggle",
      "author_url": "",
      "post_date": "01/20/2019 11:44:44",
      "content": "<p>I use the same CV strategy and get cv results in the 0.689 range. Does anyone use a different cv strategy and get better results?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 458993,
      "author_name": "m7catsue",
      "author_url": "",
      "post_date": "01/21/2019 01:08:15",
      "content": "<p>I am also using the usual 5-fold CV to make OOF predictions and test predictions.\nbest CV 0.689, LB 0.691</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459017,
      "author_name": "playif1",
      "author_url": "",
      "post_date": "01/21/2019 02:56:31",
      "content": "<p>I think you might want to check this great kernel.\n<a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a></p>\n\n<p>Here's the short summary I gathered from the above kernel.\nBasically when we apply K-fold CV and perform averaging on each fold, the prediction we give is the ensemble of K models from each fold. And there are at least two factors affect the performance if an ensemble model, i.e. your LB score.\n1. The CV score of each fold (The performance of each base classifier), the higher the better.\n2. The divergence of each fold, the higher the better.</p>\n\n<p>So you should also check the second factor (the correlation between fold-predictions) in your validation framework. If you want, you can separate some data before K-fold CV as the validation data of the final prediction (the estimation of LB score). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459142,
      "author_name": "kcostya",
      "author_url": "",
      "post_date": "01/21/2019 09:13:58",
      "content": "<p>Same approach.\n5-fold CV: 0.6943, LB: 0.699.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460831,
      "author_name": "jannen",
      "author_url": "",
      "post_date": "01/24/2019 13:55:18",
      "content": "<p>Does anyone have theoretical reasoning (or a link) to why K-fold CV is better than training a single larger model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 460848,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "01/24/2019 14:34:11",
          "content": "<p>k-fold creates k different models from different datasets, so each model will give slightly different prediction based on the data/signal it has seen. Therefore when you average the prediction from all K models, you are using bootstrap aggregation which helps your predictions.</p>\n\n<p>Also, with k-fold, you can train on the whole dataset. If you train a single larger model, you must use a holdout validation set (and lose the opportunity to train on it), or use no validation set at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460927,
          "author_name": "jannen",
          "author_url": "",
          "post_date": "01/24/2019 18:37:40",
          "content": "<p>I think when you have trained your model with train/dev and chosen suitable parameters, you can just train the same network again using 100% of your train+dev data? It shouldnt overfit, if it didn't overfit when using train/dev split and now having even more training data for same model and nr of epochs.</p>\n\n<p>In K-fold, instead of one model using 100% training data you will have 5 models using for example 80% of training data, so all of them should individually be weaker than one full model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460930,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "01/24/2019 18:47:39",
          "content": "<ol>\n<li><p>You're right, if you have good parameter selection you can retrain on train+dev data in my opinion, as long as you have the correct round/epoch to stop training. And because it sees 100% of the data, it will fit closest to all that 100% of the data since that's what it's optimizing for! Kaggle grandmaster silogram recommends that you actually multiply the epochs by 1.1 when you are training on full dataset instead.</p></li>\n<li><p>The power of K Fold, in my opinion, is its built-in bootstrap aggregation. Using only 80% of the data for 5 models, then averaging them all, makes the final prediction more rugged, which helps prevent overfitting. Actually, in my opinion, the best CV strategy is to combine both. You should find the best parameters for a model using different feature representations, then combine all of this into a final prediction. I think more people would do this, but it's really difficult to \"find the best parameters for a model\" so that is why I and majority of people use K Fold. The other way also takes a lot longer too since you are training K times on full dataset instead of K times on 80%.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "457732": "I'm curious how everyone is designing their cv. Here's mine.\n\nFor each fold, perform threshold optimization after every epoch to get the highest validation f1 score. Use the checkpoint with the highest val f1 score to predict the test set and validation set. Repeat this for all folds. Now, with oof predictions for the entire train set, I run threshold optimization to get the best threshold, and use it on the test predictions. For test predictions, I'm currently averaging the output probabilities of each model then running threshold. However, I'm curious to see if anyone is running threshold on the output probabilities of each individual model then using majority voting to create the final solution.\n\nCurrent results:\n4-fold CV: 0.681, LB: 0.690\n5-fold CV: 0.677, LB: 0.691",
    "458717": "I think I do exactly the same as you and trust more CV score than LB\n\nCV: 0.6812 LB: 0.690\nCV: 0.68075 LB: 0.691\nCV: 0.67906 LB: 0.693",
    "458733": "I use the same CV strategy and get cv results in the 0.689 range. Does anyone use a different cv strategy and get better results?",
    "458993": "I am also using the usual 5-fold CV to make OOF predictions and test predictions.\nbest CV 0.689, LB 0.691",
    "459017": "I think you might want to check this great kernel.\nhttps://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\n\nHere's the short summary I gathered from the above kernel.\nBasically when we apply K-fold CV and perform averaging on each fold, the prediction we give is the ensemble of K models from each fold. And there are at least two factors affect the performance if an ensemble model, i.e. your LB score.\n1. The CV score of each fold (The performance of each base classifier), the higher the better.\n2. The divergence of each fold, the higher the better.\n\nSo you should also check the second factor (the correlation between fold-predictions) in your validation framework. If you want, you can separate some data before K-fold CV as the validation data of the final prediction (the estimation of LB score).",
    "459142": "Same approach.\n5-fold CV: 0.6943, LB: 0.699.",
    "460831": "Does anyone have theoretical reasoning (or a link) to why K-fold CV is better than training a single larger model?",
    "460848": "k-fold creates k different models from different datasets, so each model will give slightly different prediction based on the data/signal it has seen. Therefore when you average the prediction from all K models, you are using bootstrap aggregation which helps your predictions.\n\nAlso, with k-fold, you can train on the whole dataset. If you train a single larger model, you must use a holdout validation set (and lose the opportunity to train on it), or use no validation set at all.",
    "460927": "I think when you have trained your model with train/dev and chosen suitable parameters, you can just train the same network again using 100% of your train+dev data? It shouldnt overfit, if it didn't overfit when using train/dev split and now having even more training data for same model and nr of epochs.\n\nIn K-fold, instead of one model using 100% training data you will have 5 models using for example 80% of training data, so all of them should individually be weaker than one full model?",
    "460930": "1. You're right, if you have good parameter selection you can retrain on train+dev data in my opinion, as long as you have the correct round/epoch to stop training. And because it sees 100% of the data, it will fit closest to all that 100% of the data since that's what it's optimizing for! Kaggle grandmaster silogram recommends that you actually multiply the epochs by 1.1 when you are training on full dataset instead.\n\n2. The power of K Fold, in my opinion, is its built-in bootstrap aggregation. Using only 80% of the data for 5 models, then averaging them all, makes the final prediction more rugged, which helps prevent overfitting. Actually, in my opinion, the best CV strategy is to combine both. You should find the best parameters for a model using different feature representations, then combine all of this into a final prediction. I think more people would do this, but it's really difficult to \"find the best parameters for a model\" so that is why I and majority of people use K Fold. The other way also takes a lot longer too since you are training K times on full dataset instead of K times on 80%."
  },
  "source": "meta"
}