{
  "id": 507926,
  "title": "Regarding the changing of the rule",
  "url": "/competitions/cvpr-metafood-3d-food-reconstruction-challenge/discussion/507926",
  "author_name": "",
  "post_date": "2024-05-27T19:58:28.810738500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>We have noticed the recent rule changes and would like to share some of our perspectives.</p>\n<p>The purpose of splitting the score calculation into private and public sections is to encourage the development of generalized algorithms rather than turning it into an numeric overfitting game. This is a common practice in academia, where benchmarks are divided into training, testing, and evaluation sets.</p>\n<p>So, we have the following concerns:</p>\n<p>Firstly, recalculating the scores for the first phase would make the distinction between private and public scores meaningless, akin to evaluate results using the training set, which is very uncommon. Especially with our limited data, this would lead to a competition based on submission adjustments rather than improvements in algorithms and innovations in methods.</p>\n<p>Secondly, if any rule modifications are necessary, it would be more reasonable to propose them before the results are announced. Making changes after both the results and the rules have been published impacts the fairness of the competition.</p>\n<p>Thirdly, rule changes should be based on broader discussion and consensus, rather than on a single argument.</p>\n<p>Fourthly, in the second phase, all 20 models have already been considered, addressing the issue of having too few models. Introducing all 20 models from the first phase would, for the reasons mentioned above, introduce significant unfairness.</p>\n<p>Therefore, I kindly suggest that we further discuss these rule changes to ensure the fairness and integrity of the competition. I believe that an open and thorough discussion will benefit all participants and uphold the standards of the competition.</p>\n<p>Thank you very much for your time and consideration. I look forward to your response.</p>\n<p>Best regards</p>",
  "messages": [
    {
      "id": "2839946",
      "postDate": "05/27/2024 19:58:28",
      "content": "<p>We have noticed the recent rule changes and would like to share some of our perspectives.</p>\n<p>The purpose of splitting the score calculation into private and public sections is to encourage the development of generalized algorithms rather than turning it into an numeric overfitting game. This is a common practice in academia, where benchmarks are divided into training, testing, and evaluation sets.</p>\n<p>So, we have the following concerns:</p>\n<p>Firstly, recalculating the scores for the first phase would make the distinction between private and public scores meaningless, akin to evaluate results using the training set, which is very uncommon. Especially with our limited data, this would lead to a competition based on submission adjustments rather than improvements in algorithms and innovations in methods.</p>\n<p>Secondly, if any rule modifications are necessary, it would be more reasonable to propose them before the results are announced. Making changes after both the results and the rules have been published impacts the fairness of the competition.</p>\n<p>Thirdly, rule changes should be based on broader discussion and consensus, rather than on a single argument.</p>\n<p>Fourthly, in the second phase, all 20 models have already been considered, addressing the issue of having too few models. Introducing all 20 models from the first phase would, for the reasons mentioned above, introduce significant unfairness.</p>\n<p>Therefore, I kindly suggest that we further discuss these rule changes to ensure the fairness and integrity of the competition. I believe that an open and thorough discussion will benefit all participants and uphold the standards of the competition.</p>\n<p>Thank you very much for your time and consideration. I look forward to your response.</p>\n<p>Best regards</p>",
      "rawMarkdown": "We have noticed the recent rule changes and would like to share some of our perspectives.\n\nThe purpose of splitting the score calculation into private and public sections is to encourage the development of generalized algorithms rather than turning it into an numeric overfitting game. This is a common practice in academia, where benchmarks are divided into training, testing, and evaluation sets.\n\nSo, we have the following concerns:\n\nFirstly, recalculating the scores for the first phase would make the distinction between private and public scores meaningless, akin to evaluate results using the training set, which is very uncommon. Especially with our limited data, this would lead to a competition based on submission adjustments rather than improvements in algorithms and innovations in methods.\n\nSecondly, if any rule modifications are necessary, it would be more reasonable to propose them before the results are announced. Making changes after both the results and the rules have been published impacts the fairness of the competition.\n\nThirdly, rule changes should be based on broader discussion and consensus, rather than on a single argument.\n\nFourthly, in the second phase, all 20 models have already been considered, addressing the issue of having too few models. Introducing all 20 models from the first phase would, for the reasons mentioned above, introduce significant unfairness.\n\nTherefore, I kindly suggest that we further discuss these rule changes to ensure the fairness and integrity of the competition. I believe that an open and thorough discussion will benefit all participants and uphold the standards of the competition.\n\nThank you very much for your time and consideration. I look forward to your response.\n\nBest regards",
      "votes": null
    },
    {
      "id": "2839976",
      "postDate": "05/27/2024 21:05:25",
      "content": "<p>Thank you for bringing this up. We made a quick clarification on the rule because it was our initial intention to evaluate all 20 models for the phase 1 as well. We didn't have the correct setting and the leaderboard released the scores before we make our final calculation. </p>\n<p>That being said, we do see there is potential problem associated with this setup involving the 15 public models into the score calculation. The key question to ask is that if the public score creates data leakage that give potential advantage. We also at the same time, want to ensure comprehensive evaluation. Feel free to provide your opinion. </p>\n<p>Another thing we want to mention is that we do check your models manually, and make sure they are the ones you used in the first phase and have valid shape. We hope this will somewhat prevent people from cheating.</p>\n<p>Please focus on the phase 2 and we are in the process of discussing the fairness of phase 1 evaluation, mainly on if we need to consider the public submissions or not. </p>",
      "rawMarkdown": "Thank you for bringing this up. We made a quick clarification on the rule because it was our initial intention to evaluate all 20 models for the phase 1 as well. We didn't have the correct setting and the leaderboard released the scores before we make our final calculation. \n\nThat being said, we do see there is potential problem associated with this setup involving the 15 public models into the score calculation. The key question to ask is that if the public score creates data leakage that give potential advantage. We also at the same time, want to ensure comprehensive evaluation. Feel free to provide your opinion. \n\nAnother thing we want to mention is that we do check your models manually, and make sure they are the ones you used in the first phase and have valid shape. We hope this will somewhat prevent people from cheating.\n\nPlease focus on the phase 2 and we are in the process of discussing the fairness of phase 1 evaluation, mainly on if we need to consider the public submissions or not.",
      "votes": null
    },
    {
      "id": "2854232",
      "postDate": "06/04/2024 07:24:21",
      "content": "<p>Thank you for noting this important point.<br>\nThe fairness of the evaluation methodology is crucial for research. As Yawei Jueluo mentioned, using all 20 objects for Phase 1 would be akin to evaluating on the train set (for 15 out of 20), which is clearly undesirable. This is especially severe if equal weight is given to each item, since that means the training set metrics have a 3 to 1 weight overall. The public scores do create data leakage: since the MAPE metric is relative to the GT volume, one can easily infer the exact ground truth volume for any single scene with one submission. It is therefore possible to even obtain a perfect score on the public leaderboard. Checking (in phase 2) that the models have a reasonable shape does not fix the problem, as one can predict a coarse 3D shape and then just scale it to match the known volume.<br>\nTherefore, while the public leaderboard is appreciated as it helped us develop better algorithms for the challenge, evaluating the solutions on this set for the final standings would be contrary to standard practice in computer vision research.<br>\nThank you in anticipation for your response.</p>",
      "rawMarkdown": "Thank you for noting this important point.\nThe fairness of the evaluation methodology is crucial for research. As Yawei Jueluo mentioned, using all 20 objects for Phase 1 would be akin to evaluating on the train set (for 15 out of 20), which is clearly undesirable. This is especially severe if equal weight is given to each item, since that means the training set metrics have a 3 to 1 weight overall. The public scores do create data leakage: since the MAPE metric is relative to the GT volume, one can easily infer the exact ground truth volume for any single scene with one submission. It is therefore possible to even obtain a perfect score on the public leaderboard. Checking (in phase 2) that the models have a reasonable shape does not fix the problem, as one can predict a coarse 3D shape and then just scale it to match the known volume.\nTherefore, while the public leaderboard is appreciated as it helped us develop better algorithms for the challenge, evaluating the solutions on this set for the final standings would be contrary to standard practice in computer vision research.\nThank you in anticipation for your response.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2839976,
      "author_name": "metafoodcvpr",
      "author_url": "",
      "post_date": "05/27/2024 21:05:25",
      "content": "<p>Thank you for bringing this up. We made a quick clarification on the rule because it was our initial intention to evaluate all 20 models for the phase 1 as well. We didn't have the correct setting and the leaderboard released the scores before we make our final calculation. </p>\n<p>That being said, we do see there is potential problem associated with this setup involving the 15 public models into the score calculation. The key question to ask is that if the public score creates data leakage that give potential advantage. We also at the same time, want to ensure comprehensive evaluation. Feel free to provide your opinion. </p>\n<p>Another thing we want to mention is that we do check your models manually, and make sure they are the ones you used in the first phase and have valid shape. We hope this will somewhat prevent people from cheating.</p>\n<p>Please focus on the phase 2 and we are in the process of discussing the fairness of phase 1 evaluation, mainly on if we need to consider the public submissions or not. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2854232,
      "author_name": "timitoc",
      "author_url": "",
      "post_date": "06/04/2024 07:24:21",
      "content": "<p>Thank you for noting this important point.<br>\nThe fairness of the evaluation methodology is crucial for research. As Yawei Jueluo mentioned, using all 20 objects for Phase 1 would be akin to evaluating on the train set (for 15 out of 20), which is clearly undesirable. This is especially severe if equal weight is given to each item, since that means the training set metrics have a 3 to 1 weight overall. The public scores do create data leakage: since the MAPE metric is relative to the GT volume, one can easily infer the exact ground truth volume for any single scene with one submission. It is therefore possible to even obtain a perfect score on the public leaderboard. Checking (in phase 2) that the models have a reasonable shape does not fix the problem, as one can predict a coarse 3D shape and then just scale it to match the known volume.<br>\nTherefore, while the public leaderboard is appreciated as it helped us develop better algorithms for the challenge, evaluating the solutions on this set for the final standings would be contrary to standard practice in computer vision research.<br>\nThank you in anticipation for your response.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2839946": "We have noticed the recent rule changes and would like to share some of our perspectives.\n\nThe purpose of splitting the score calculation into private and public sections is to encourage the development of generalized algorithms rather than turning it into an numeric overfitting game. This is a common practice in academia, where benchmarks are divided into training, testing, and evaluation sets.\n\nSo, we have the following concerns:\n\nFirstly, recalculating the scores for the first phase would make the distinction between private and public scores meaningless, akin to evaluate results using the training set, which is very uncommon. Especially with our limited data, this would lead to a competition based on submission adjustments rather than improvements in algorithms and innovations in methods.\n\nSecondly, if any rule modifications are necessary, it would be more reasonable to propose them before the results are announced. Making changes after both the results and the rules have been published impacts the fairness of the competition.\n\nThirdly, rule changes should be based on broader discussion and consensus, rather than on a single argument.\n\nFourthly, in the second phase, all 20 models have already been considered, addressing the issue of having too few models. Introducing all 20 models from the first phase would, for the reasons mentioned above, introduce significant unfairness.\n\nTherefore, I kindly suggest that we further discuss these rule changes to ensure the fairness and integrity of the competition. I believe that an open and thorough discussion will benefit all participants and uphold the standards of the competition.\n\nThank you very much for your time and consideration. I look forward to your response.\n\nBest regards",
    "2839976": "Thank you for bringing this up. We made a quick clarification on the rule because it was our initial intention to evaluate all 20 models for the phase 1 as well. We didn't have the correct setting and the leaderboard released the scores before we make our final calculation. \n\nThat being said, we do see there is potential problem associated with this setup involving the 15 public models into the score calculation. The key question to ask is that if the public score creates data leakage that give potential advantage. We also at the same time, want to ensure comprehensive evaluation. Feel free to provide your opinion. \n\nAnother thing we want to mention is that we do check your models manually, and make sure they are the ones you used in the first phase and have valid shape. We hope this will somewhat prevent people from cheating.\n\nPlease focus on the phase 2 and we are in the process of discussing the fairness of phase 1 evaluation, mainly on if we need to consider the public submissions or not.",
    "2854232": "Thank you for noting this important point.\nThe fairness of the evaluation methodology is crucial for research. As Yawei Jueluo mentioned, using all 20 objects for Phase 1 would be akin to evaluating on the train set (for 15 out of 20), which is clearly undesirable. This is especially severe if equal weight is given to each item, since that means the training set metrics have a 3 to 1 weight overall. The public scores do create data leakage: since the MAPE metric is relative to the GT volume, one can easily infer the exact ground truth volume for any single scene with one submission. It is therefore possible to even obtain a perfect score on the public leaderboard. Checking (in phase 2) that the models have a reasonable shape does not fix the problem, as one can predict a coarse 3D shape and then just scale it to match the known volume.\nTherefore, while the public leaderboard is appreciated as it helped us develop better algorithms for the challenge, evaluating the solutions on this set for the final standings would be contrary to standard practice in computer vision research.\nThank you in anticipation for your response."
  },
  "source": "meta"
}