{
  "id": 331148,
  "title": "How to Measure K-Fold Performance?",
  "url": "/competitions/amex-default-prediction/discussion/331148",
  "author_name": "",
  "post_date": "2022-06-15T23:21:23.268530Z",
  "votes": 12,
  "comment_count": 8,
  "views": 0,
  "content": "<p>There are two ways that I know of to measure k-fold performance:</p>\n<ol>\n<li>Calculate the metric for each out-of-fold prediction and average those metrics</li>\n<li>Keep track of all out-of-fold predictions and then calculate the metric once for the entire set of predictions</li>\n</ol>\n<p>What are the pros and cons of each approach?<br>\nThank you!</p>",
  "messages": [
    {
      "id": "1821919",
      "postDate": "06/15/2022 23:21:23",
      "content": "<p>There are two ways that I know of to measure k-fold performance:</p>\n<ol>\n<li>Calculate the metric for each out-of-fold prediction and average those metrics</li>\n<li>Keep track of all out-of-fold predictions and then calculate the metric once for the entire set of predictions</li>\n</ol>\n<p>What are the pros and cons of each approach?<br>\nThank you!</p>",
      "rawMarkdown": "There are two ways that I know of to measure k-fold performance:\n1. Calculate the metric for each out-of-fold prediction and average those metrics\n2. Keep track of all out-of-fold predictions and then calculate the metric once for the entire set of predictions\n\nWhat are the pros and cons of each approach?\nThank you!",
      "votes": null
    },
    {
      "id": "1822519",
      "postDate": "06/16/2022 13:18:44",
      "content": "<p>I would advice to use average over folds for this competitions.</p>\n<p>OOF predictions concatenation may give you weird results due to class probabilities ranges in different fold (due to different early stoppings rounds in each fold / the effect will be greater with higher LR). It could be solved partially by min max scaling but the end result will not be stable to use is as a decision making value.</p>",
      "rawMarkdown": "I would advice to use average over folds for this competitions.\n\nOOF predictions concatenation may give you weird results due to class probabilities ranges in different fold (due to different early stoppings rounds in each fold / the effect will be greater with higher LR). It could be solved partially by min max scaling but the end result will not be stable to use is as a decision making value.",
      "votes": null
    },
    {
      "id": "1822529",
      "postDate": "06/16/2022 13:29:32",
      "content": "<p>Both are just fine as long as you see the correlation with public score improvements.<br>\nBut generally 1. is more preferred for ranking metrics</p>",
      "rawMarkdown": "Both are just fine as long as you see the correlation with public score improvements.\nBut generally 1. is more preferred for ranking metrics",
      "votes": null
    },
    {
      "id": "1822658",
      "postDate": "06/16/2022 14:56:36",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>!</p>",
      "rawMarkdown": "Thank you @raddar!",
      "votes": null
    },
    {
      "id": "1822661",
      "postDate": "06/16/2022 14:57:23",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a>!</p>",
      "rawMarkdown": "Thank you @kyakovlev!",
      "votes": null
    },
    {
      "id": "1823543",
      "postDate": "06/17/2022 13:31:18",
      "content": "<p>Averaging is known to reduce the variance I would go with this approach, the second approach can be biased by the predictions on certain folds being better or worse than predictions on the remaining folds. </p>",
      "rawMarkdown": "Averaging is known to reduce the variance I would go with this approach, the second approach can be biased by the predictions on certain folds being better or worse than predictions on the remaining folds.",
      "votes": null
    },
    {
      "id": "1823698",
      "postDate": "06/17/2022 15:40:10",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/omarjarir\" target=\"_blank\">@omarjarir</a>!</p>",
      "rawMarkdown": "Thank you @omarjarir!",
      "votes": null
    },
    {
      "id": "1838954",
      "postDate": "07/01/2022 02:53:18",
      "content": "<p>calculating the metric for the validation set and averaging it works fine for this competition. </p>",
      "rawMarkdown": "calculating the metric for the validation set and averaging it works fine for this competition.",
      "votes": null
    },
    {
      "id": "1838958",
      "postDate": "07/01/2022 02:57:02",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a>!</p>",
      "rawMarkdown": "Thanks @mohammadrahmati!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1822519,
      "author_name": "kyakovlev",
      "author_url": "",
      "post_date": "06/16/2022 13:18:44",
      "content": "<p>I would advice to use average over folds for this competitions.</p>\n<p>OOF predictions concatenation may give you weird results due to class probabilities ranges in different fold (due to different early stoppings rounds in each fold / the effect will be greater with higher LR). It could be solved partially by min max scaling but the end result will not be stable to use is as a decision making value.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1822661,
          "author_name": "ryantran2165",
          "author_url": "",
          "post_date": "06/16/2022 14:57:23",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1822529,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/16/2022 13:29:32",
      "content": "<p>Both are just fine as long as you see the correlation with public score improvements.<br>\nBut generally 1. is more preferred for ranking metrics</p>",
      "votes": null,
      "replies": [
        {
          "id": 1822658,
          "author_name": "ryantran2165",
          "author_url": "",
          "post_date": "06/16/2022 14:56:36",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1823543,
      "author_name": "omarjarir",
      "author_url": "",
      "post_date": "06/17/2022 13:31:18",
      "content": "<p>Averaging is known to reduce the variance I would go with this approach, the second approach can be biased by the predictions on certain folds being better or worse than predictions on the remaining folds. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1823698,
          "author_name": "ryantran2165",
          "author_url": "",
          "post_date": "06/17/2022 15:40:10",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/omarjarir\" target=\"_blank\">@omarjarir</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1838954,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/01/2022 02:53:18",
      "content": "<p>calculating the metric for the validation set and averaging it works fine for this competition. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1838958,
          "author_name": "ryantran2165",
          "author_url": "",
          "post_date": "07/01/2022 02:57:02",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1821919": "There are two ways that I know of to measure k-fold performance:\n1. Calculate the metric for each out-of-fold prediction and average those metrics\n2. Keep track of all out-of-fold predictions and then calculate the metric once for the entire set of predictions\n\nWhat are the pros and cons of each approach?\nThank you!",
    "1822519": "I would advice to use average over folds for this competitions.\n\nOOF predictions concatenation may give you weird results due to class probabilities ranges in different fold (due to different early stoppings rounds in each fold / the effect will be greater with higher LR). It could be solved partially by min max scaling but the end result will not be stable to use is as a decision making value.",
    "1822529": "Both are just fine as long as you see the correlation with public score improvements.\nBut generally 1. is more preferred for ranking metrics",
    "1822658": "Thank you @raddar!",
    "1822661": "Thank you @kyakovlev!",
    "1823543": "Averaging is known to reduce the variance I would go with this approach, the second approach can be biased by the predictions on certain folds being better or worse than predictions on the remaining folds.",
    "1823698": "Thank you @omarjarir!",
    "1838954": "calculating the metric for the validation set and averaging it works fine for this competition.",
    "1838958": "Thanks @mohammadrahmati!"
  },
  "source": "meta"
}