{
  "id": 338001,
  "title": "Avg validation CV or overall validation CV?",
  "url": "/competitions/amex-default-prediction/discussion/338001",
  "author_name": "",
  "post_date": "2022-07-18T17:55:32.718257200Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>What are you looking at when you are checking your CV score? </p>\n<p>I've almost always used the overall way in past competitions, although both were giving almost the same results. However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly. I recall reading a comment saying that the average option is better than the overall (which makes sense to me) but I am not entirely sure.</p>",
  "messages": [
    {
      "id": "1861005",
      "postDate": "07/18/2022 17:55:32",
      "content": "<p>What are you looking at when you are checking your CV score? </p>\n<p>I've almost always used the overall way in past competitions, although both were giving almost the same results. However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly. I recall reading a comment saying that the average option is better than the overall (which makes sense to me) but I am not entirely sure.</p>",
      "rawMarkdown": "What are you looking at when you are checking your CV score? \n\nI've almost always used the overall way in past competitions, although both were giving almost the same results. However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly. I recall reading a comment saying that the average option is better than the overall (which makes sense to me) but I am not entirely sure.",
      "votes": null
    },
    {
      "id": "1861185",
      "postDate": "07/18/2022 20:14:55",
      "content": "<p>Like you said, in most cases they are near-identical. Since the whole train dataset is used for any downstream application (ensembling, etc), the overall CV is the relevant measure.</p>\n<blockquote>\n  <p>However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly.</p>\n</blockquote>\n<p>CV scores and their calculations have nothing to do with eventual leaderboard scores. Whether you calculate the average CV as 0.791532 or the overall CV as 0.791514, the leaderboard score for that particular prediction will be the same. CV scores are only four bookkeeping purposes, so it should be enough that they agree to the 4th decimal place. And even if they don't, I don't see a problem as long as we consistently use one way of calculating CV.</p>",
      "rawMarkdown": "Like you said, in most cases they are near-identical. Since the whole train dataset is used for any downstream application (ensembling, etc), the overall CV is the relevant measure.\n\n> However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly.\n\nCV scores and their calculations have nothing to do with eventual leaderboard scores. Whether you calculate the average CV as 0.791532 or the overall CV as 0.791514, the leaderboard score for that particular prediction will be the same. CV scores are only four bookkeeping purposes, so it should be enough that they agree to the 4th decimal place. And even if they don't, I don't see a problem as long as we consistently use one way of calculating CV.",
      "votes": null
    },
    {
      "id": "1861302",
      "postDate": "07/19/2022 00:35:57",
      "content": "<p>I look at both, mostly just as a way to help make sure that something strange isn't going on :) </p>\n<p>I would also recommend looking at all individual folds instead of just the summary. It's been discussed how noisy this problem/metric is, so an aggregate CV improvement might conceal an improvement concentrated in only 1 fold. That could be more about luck on that fold than real substance.  </p>",
      "rawMarkdown": "I look at both, mostly just as a way to help make sure that something strange isn't going on :) \n\nI would also recommend looking at all individual folds instead of just the summary. It's been discussed how noisy this problem/metric is, so an aggregate CV improvement might conceal an improvement concentrated in only 1 fold. That could be more about luck on that fold than real substance.",
      "votes": null
    },
    {
      "id": "1862033",
      "postDate": "07/19/2022 12:12:10",
      "content": "<p>I understand that in this case the metric cares about the order of predictions (such as AUC), so when you calculate the overall metric using oofs you are taking into account the order of predictions made by different models (a different model per fold) which I don't know if it makes much sense.</p>\n<p>On the other hand, (in my case) sometimes the Avg CV and the Overall CV give opposite results for 2 experiments, for example:</p>\n<p>Run a 5-fold CV with 3 different seeds (15 models in total):</p>\n<ul>\n<li>Experiment A: Avg 0.79694+-0.00012, Overall 0.79832</li>\n<li>Experiment B: Avg 0.79720+-0.00030, Overall 0.79830</li>\n</ul>\n<p>There is a difference of 112 features between one experiment and the other. Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.</p>\n<p>I calculate the Avg and the Overall as follows:</p>\n<ul>\n<li>Avg: Average of the score of all the validation folds (15 in total, 5 per seed) (and also the standard deviation)</li>\n<li>Overall: Average the oofs (this is an average of 3 elements, 1 per seed), and then calculate the metric with the entire oofs.</li>\n</ul>",
      "rawMarkdown": "I understand that in this case the metric cares about the order of predictions (such as AUC), so when you calculate the overall metric using oofs you are taking into account the order of predictions made by different models (a different model per fold) which I don't know if it makes much sense.\n\nOn the other hand, (in my case) sometimes the Avg CV and the Overall CV give opposite results for 2 experiments, for example:\n\nRun a 5-fold CV with 3 different seeds (15 models in total):\n- Experiment A: Avg 0.79694+-0.00012, Overall 0.79832\n- Experiment B: Avg 0.79720+-0.00030, Overall 0.79830\n\nThere is a difference of 112 features between one experiment and the other. Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.\n\nI calculate the Avg and the Overall as follows:\n- Avg: Average of the score of all the validation folds (15 in total, 5 per seed) (and also the standard deviation)\n- Overall: Average the oofs (this is an average of 3 elements, 1 per seed), and then calculate the metric with the entire oofs.",
      "votes": null
    },
    {
      "id": "1862541",
      "postDate": "07/19/2022 19:24:48",
      "content": "<blockquote>\n  <p>Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.</p>\n</blockquote>\n<p>I think so. When you account for StDev, the two experiments are indistinguishable based on averages, and they differ on 5th decimal place by overall CV. This metric is very unstable, and I am sure most of us have encountered similar discrepancy.</p>",
      "rawMarkdown": "> Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.\n\nI think so. When you account for StDev, the two experiments are indistinguishable based on averages, and they differ on 5th decimal place by overall CV. This metric is very unstable, and I am sure most of us have encountered similar discrepancy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1861185,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "07/18/2022 20:14:55",
      "content": "<p>Like you said, in most cases they are near-identical. Since the whole train dataset is used for any downstream application (ensembling, etc), the overall CV is the relevant measure.</p>\n<blockquote>\n  <p>However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly.</p>\n</blockquote>\n<p>CV scores and their calculations have nothing to do with eventual leaderboard scores. Whether you calculate the average CV as 0.791532 or the overall CV as 0.791514, the leaderboard score for that particular prediction will be the same. CV scores are only four bookkeeping purposes, so it should be enough that they agree to the 4th decimal place. And even if they don't, I don't see a problem as long as we consistently use one way of calculating CV.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1861302,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "07/19/2022 00:35:57",
      "content": "<p>I look at both, mostly just as a way to help make sure that something strange isn't going on :) </p>\n<p>I would also recommend looking at all individual folds instead of just the summary. It's been discussed how noisy this problem/metric is, so an aggregate CV improvement might conceal an improvement concentrated in only 1 fold. That could be more about luck on that fold than real substance.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1862033,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "07/19/2022 12:12:10",
      "content": "<p>I understand that in this case the metric cares about the order of predictions (such as AUC), so when you calculate the overall metric using oofs you are taking into account the order of predictions made by different models (a different model per fold) which I don't know if it makes much sense.</p>\n<p>On the other hand, (in my case) sometimes the Avg CV and the Overall CV give opposite results for 2 experiments, for example:</p>\n<p>Run a 5-fold CV with 3 different seeds (15 models in total):</p>\n<ul>\n<li>Experiment A: Avg 0.79694+-0.00012, Overall 0.79832</li>\n<li>Experiment B: Avg 0.79720+-0.00030, Overall 0.79830</li>\n</ul>\n<p>There is a difference of 112 features between one experiment and the other. Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.</p>\n<p>I calculate the Avg and the Overall as follows:</p>\n<ul>\n<li>Avg: Average of the score of all the validation folds (15 in total, 5 per seed) (and also the standard deviation)</li>\n<li>Overall: Average the oofs (this is an average of 3 elements, 1 per seed), and then calculate the metric with the entire oofs.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1862541,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/19/2022 19:24:48",
          "content": "<blockquote>\n  <p>Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.</p>\n</blockquote>\n<p>I think so. When you account for StDev, the two experiments are indistinguishable based on averages, and they differ on 5th decimal place by overall CV. This metric is very unstable, and I am sure most of us have encountered similar discrepancy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1861005": "What are you looking at when you are checking your CV score? \n\nI've almost always used the overall way in past competitions, although both were giving almost the same results. However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly. I recall reading a comment saying that the average option is better than the overall (which makes sense to me) but I am not entirely sure.",
    "1861185": "Like you said, in most cases they are near-identical. Since the whole train dataset is used for any downstream application (ensembling, etc), the overall CV is the relevant measure.\n\n> However, in this competition in which we are fighting for the last decimal place, the two ways of calculating the CV score sometimes differ significantly.\n\nCV scores and their calculations have nothing to do with eventual leaderboard scores. Whether you calculate the average CV as 0.791532 or the overall CV as 0.791514, the leaderboard score for that particular prediction will be the same. CV scores are only four bookkeeping purposes, so it should be enough that they agree to the 4th decimal place. And even if they don't, I don't see a problem as long as we consistently use one way of calculating CV.",
    "1861302": "I look at both, mostly just as a way to help make sure that something strange isn't going on :) \n\nI would also recommend looking at all individual folds instead of just the summary. It's been discussed how noisy this problem/metric is, so an aggregate CV improvement might conceal an improvement concentrated in only 1 fold. That could be more about luck on that fold than real substance.",
    "1862033": "I understand that in this case the metric cares about the order of predictions (such as AUC), so when you calculate the overall metric using oofs you are taking into account the order of predictions made by different models (a different model per fold) which I don't know if it makes much sense.\n\nOn the other hand, (in my case) sometimes the Avg CV and the Overall CV give opposite results for 2 experiments, for example:\n\nRun a 5-fold CV with 3 different seeds (15 models in total):\n- Experiment A: Avg 0.79694+-0.00012, Overall 0.79832\n- Experiment B: Avg 0.79720+-0.00030, Overall 0.79830\n\nThere is a difference of 112 features between one experiment and the other. Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.\n\nI calculate the Avg and the Overall as follows:\n- Avg: Average of the score of all the validation folds (15 in total, 5 per seed) (and also the standard deviation)\n- Overall: Average the oofs (this is an average of 3 elements, 1 per seed), and then calculate the metric with the entire oofs.",
    "1862541": "> Maybe the standard deviation related to the Avg could explain the discrepancy or the difference is too low and I am making a mountain out of a molehill.\n\nI think so. When you account for StDev, the two experiments are indistinguishable based on averages, and they differ on 5th decimal place by overall CV. This metric is very unstable, and I am sure most of us have encountered similar discrepancy."
  },
  "source": "meta"
}