{
  "id": 77406,
  "title": "What is the average correlation between your CV folds?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77406",
  "author_name": "",
  "post_date": "2019-01-12T10:58:42.784828100Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi</p>\n\n<p>As pointed out by <a href=\"/bminixhofer\">@bminixhofer</a> in his kernel <a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a> ,  lower correlations between test predictions of different folds should increase the score of the ensembled model. This might be helpful in deciding which kernels to choose for the stage 2 re-scoring.</p>\n\n<p>My best CV model:</p>\n\n<p>Public LB score: 0.691</p>\n\n<p>CV F1: 0.6889</p>\n\n<p>CV roc-auc: 0.9688</p>\n\n<p>Average correlations of test preds: 0.9680</p>",
  "messages": [
    {
      "id": "454843",
      "postDate": "01/12/2019 10:58:42",
      "content": "<p>Hi</p>\n\n<p>As pointed out by <a href=\"/bminixhofer\">@bminixhofer</a> in his kernel <a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed</a> ,  lower correlations between test predictions of different folds should increase the score of the ensembled model. This might be helpful in deciding which kernels to choose for the stage 2 re-scoring.</p>\n\n<p>My best CV model:</p>\n\n<p>Public LB score: 0.691</p>\n\n<p>CV F1: 0.6889</p>\n\n<p>CV roc-auc: 0.9688</p>\n\n<p>Average correlations of test preds: 0.9680</p>",
      "rawMarkdown": "Hi\n\nAs pointed out by @bminixhofer in his kernel https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed ,  lower correlations between test predictions of different folds should increase the score of the ensembled model. This might be helpful in deciding which kernels to choose for the stage 2 re-scoring.\n\nMy best CV model:\n\nPublic LB score: 0.691\n\nCV F1: 0.6889\n\nCV roc-auc: 0.9688\n\nAverage correlations of test preds: 0.9680",
      "votes": null
    },
    {
      "id": "454844",
      "postDate": "01/12/2019 11:01:10",
      "content": "<p>I agree that checking correlation can be helpful, but this statement </p>\n\n<blockquote>\n  <p>lower correlations between test predictions of different folds should\n  increase the score of the ensembled model</p>\n</blockquote>\n\n<p>is not generally applicable.</p>",
      "rawMarkdown": "I agree that checking correlation can be helpful, but this statement \n\n&gt; lower correlations between test predictions of different folds should\n&gt; increase the score of the ensembled model\n\n is not generally applicable.",
      "votes": null
    },
    {
      "id": "454855",
      "postDate": "01/12/2019 11:40:21",
      "content": "<p>Why not? Generally, the more diverse the models are the better ensembling works.</p>",
      "rawMarkdown": "Why not? Generally, the more diverse the models are the better ensembling works.",
      "votes": null
    },
    {
      "id": "454888",
      "postDate": "01/12/2019 12:38:12",
      "content": "<p>One reason is because ensemble performance depends on mainly two factors : \nbase individual performance <strong>and</strong> diversification among them.  </p>\n\n<p>Therefore, in some situations (<em>which usually not occur here, so I understand what Benjamin really meant</em>), we could have very bad K base classifiers, and although very diversified, the ensemble performance is still not great.</p>\n\n<p>Anyway, to reply this topic, average correlation of my current best model is around 0.91</p>",
      "rawMarkdown": "One reason is because ensemble performance depends on mainly two factors : \nbase individual performance **and** diversification among them.  \n\nTherefore, in some situations (*which usually not occur here, so I understand what Benjamin really meant*), we could have very bad K base classifiers, and although very diversified, the ensemble performance is still not great.\n\nAnyway, to reply this topic, average correlation of my current best model is around 0.91",
      "votes": null
    },
    {
      "id": "454915",
      "postDate": "01/12/2019 13:48:48",
      "content": "<p>How does your cv F1 score compare to your average correlation?</p>\n\n<p>Has anyone tried using a different number of folds (4 vs 5)? I tried using 4 folds but that gave me a lower cv F1 score. Would using fewer folds diversify the models enough to make up for the loss in cv F1?</p>",
      "rawMarkdown": "How does your cv F1 score compare to your average correlation?\n\n Has anyone tried using a different number of folds (4 vs 5)? I tried using 4 folds but that gave me a lower cv F1 score. Would using fewer folds diversify the models enough to make up for the loss in cv F1?",
      "votes": null
    },
    {
      "id": "454930",
      "postDate": "01/12/2019 14:28:27",
      "content": "<p>Because as already said if one model is bad, then it probably also won't correlate well with the others, and overall performance is bad. Also, if one model is super awesome on itself, and you ensemble it with worse models with lower correlation to it, then also the final prediction will be worse. </p>",
      "rawMarkdown": "Because as already said if one model is bad, then it probably also won't correlate well with the others, and overall performance is bad. Also, if one model is super awesome on itself, and you ensemble it with worse models with lower correlation to it, then also the final prediction will be worse.",
      "votes": null
    },
    {
      "id": "454931",
      "postDate": "01/12/2019 14:39:34",
      "content": "<p>Another small note to @bkkagle, even though using 4 folds could perhaps give you more diversification, but since the training data of each base classifier will be less (75% in 4 Folds vs.  80% in 5 Folds),  base performance will inevitably drop, and that affects the final ensemble performance.</p>\n\n<p>So I am not sure, based on logical reason, whether 4 Folds or 5 Folds are better in the trade-off, but, based on my own experiments, 5 Folds seem to be better here.  Also 6 folds also did not get good performance in my case.</p>\n\n<p>To reply your first question, since I don't think that \"pure\" average correlation is an appropriate metric by itself, and I have tried to do some few different kinds of CV-F1 measures, so I haven't tracked the avg. correlation with CV-f1 in my experiment. </p>",
      "rawMarkdown": "Another small note to @bkkagle, even though using 4 folds could perhaps give you more diversification, but since the training data of each base classifier will be less (75% in 4 Folds vs.  80% in 5 Folds),  base performance will inevitably drop, and that affects the final ensemble performance.\n\nSo I am not sure, based on logical reason, whether 4 Folds or 5 Folds are better in the trade-off, but, based on my own experiments, 5 Folds seem to be better here.  Also 6 folds also did not get good performance in my case.\n\nTo reply your first question, since I don't think that \"pure\" average correlation is an appropriate metric by itself, and I have tried to do some few different kinds of CV-F1 measures, so I haven't tracked the avg. correlation with CV-f1 in my experiment.",
      "votes": null
    },
    {
      "id": "455035",
      "postDate": "01/12/2019 19:42:04",
      "content": "<p>There is a really old kernel which shows corr between each fold prediction and the coeffients are even lower: <a href=\"https://www.kaggle.com/kagsen/k-fold-on-a-single-model-vs-a-model-salad\">https://www.kaggle.com/kagsen/k-fold-on-a-single-model-vs-a-model-salad</a></p>",
      "rawMarkdown": "There is a really old kernel which shows corr between each fold prediction and the coeffients are even lower: https://www.kaggle.com/kagsen/k-fold-on-a-single-model-vs-a-model-salad",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454844,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "01/12/2019 11:01:10",
      "content": "<p>I agree that checking correlation can be helpful, but this statement </p>\n\n<blockquote>\n  <p>lower correlations between test predictions of different folds should\n  increase the score of the ensembled model</p>\n</blockquote>\n\n<p>is not generally applicable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454855,
          "author_name": "bminixhofer",
          "author_url": "",
          "post_date": "01/12/2019 11:40:21",
          "content": "<p>Why not? Generally, the more diverse the models are the better ensembling works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454888,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "01/12/2019 12:38:12",
          "content": "<p>One reason is because ensemble performance depends on mainly two factors : \nbase individual performance <strong>and</strong> diversification among them.  </p>\n\n<p>Therefore, in some situations (<em>which usually not occur here, so I understand what Benjamin really meant</em>), we could have very bad K base classifiers, and although very diversified, the ensemble performance is still not great.</p>\n\n<p>Anyway, to reply this topic, average correlation of my current best model is around 0.91</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454915,
          "author_name": "bkkaggle",
          "author_url": "",
          "post_date": "01/12/2019 13:48:48",
          "content": "<p>How does your cv F1 score compare to your average correlation?</p>\n\n<p>Has anyone tried using a different number of folds (4 vs 5)? I tried using 4 folds but that gave me a lower cv F1 score. Would using fewer folds diversify the models enough to make up for the loss in cv F1?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454930,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "01/12/2019 14:28:27",
          "content": "<p>Because as already said if one model is bad, then it probably also won't correlate well with the others, and overall performance is bad. Also, if one model is super awesome on itself, and you ensemble it with worse models with lower correlation to it, then also the final prediction will be worse. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454931,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "01/12/2019 14:39:34",
          "content": "<p>Another small note to @bkkagle, even though using 4 folds could perhaps give you more diversification, but since the training data of each base classifier will be less (75% in 4 Folds vs.  80% in 5 Folds),  base performance will inevitably drop, and that affects the final ensemble performance.</p>\n\n<p>So I am not sure, based on logical reason, whether 4 Folds or 5 Folds are better in the trade-off, but, based on my own experiments, 5 Folds seem to be better here.  Also 6 folds also did not get good performance in my case.</p>\n\n<p>To reply your first question, since I don't think that \"pure\" average correlation is an appropriate metric by itself, and I have tried to do some few different kinds of CV-F1 measures, so I haven't tracked the avg. correlation with CV-f1 in my experiment. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 455035,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "01/12/2019 19:42:04",
      "content": "<p>There is a really old kernel which shows corr between each fold prediction and the coeffients are even lower: <a href=\"https://www.kaggle.com/kagsen/k-fold-on-a-single-model-vs-a-model-salad\">https://www.kaggle.com/kagsen/k-fold-on-a-single-model-vs-a-model-salad</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454843": "Hi\n\nAs pointed out by @bminixhofer in his kernel https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed ,  lower correlations between test predictions of different folds should increase the score of the ensembled model. This might be helpful in deciding which kernels to choose for the stage 2 re-scoring.\n\nMy best CV model:\n\nPublic LB score: 0.691\n\nCV F1: 0.6889\n\nCV roc-auc: 0.9688\n\nAverage correlations of test preds: 0.9680",
    "454844": "I agree that checking correlation can be helpful, but this statement \n\n&gt; lower correlations between test predictions of different folds should\n&gt; increase the score of the ensembled model\n\n is not generally applicable.",
    "454855": "Why not? Generally, the more diverse the models are the better ensembling works.",
    "454888": "One reason is because ensemble performance depends on mainly two factors : \nbase individual performance **and** diversification among them.  \n\nTherefore, in some situations (*which usually not occur here, so I understand what Benjamin really meant*), we could have very bad K base classifiers, and although very diversified, the ensemble performance is still not great.\n\nAnyway, to reply this topic, average correlation of my current best model is around 0.91",
    "454915": "How does your cv F1 score compare to your average correlation?\n\n Has anyone tried using a different number of folds (4 vs 5)? I tried using 4 folds but that gave me a lower cv F1 score. Would using fewer folds diversify the models enough to make up for the loss in cv F1?",
    "454930": "Because as already said if one model is bad, then it probably also won't correlate well with the others, and overall performance is bad. Also, if one model is super awesome on itself, and you ensemble it with worse models with lower correlation to it, then also the final prediction will be worse.",
    "454931": "Another small note to @bkkagle, even though using 4 folds could perhaps give you more diversification, but since the training data of each base classifier will be less (75% in 4 Folds vs.  80% in 5 Folds),  base performance will inevitably drop, and that affects the final ensemble performance.\n\nSo I am not sure, based on logical reason, whether 4 Folds or 5 Folds are better in the trade-off, but, based on my own experiments, 5 Folds seem to be better here.  Also 6 folds also did not get good performance in my case.\n\nTo reply your first question, since I don't think that \"pure\" average correlation is an appropriate metric by itself, and I have tried to do some few different kinds of CV-F1 measures, so I haven't tracked the avg. correlation with CV-f1 in my experiment.",
    "455035": "There is a really old kernel which shows corr between each fold prediction and the coeffients are even lower: https://www.kaggle.com/kagsen/k-fold-on-a-single-model-vs-a-model-salad"
  },
  "source": "meta"
}