{
  "id": 563292,
  "title": "About Fairness",
  "url": "/competitions/nexar-collision-prediction/discussion/563292",
  "author_name": "",
  "post_date": "2025-02-16T10:59:56.583526800Z",
  "votes": 3,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Does having all the the test videos available for participants not increase the risk of hard coding the solutions?</p>",
  "messages": [
    {
      "id": "3125552",
      "postDate": "02/16/2025 10:59:56",
      "content": "<p>Does having all the the test videos available for participants not increase the risk of hard coding the solutions?</p>",
      "rawMarkdown": "Does having all the the test videos available for participants not increase the risk of hard coding the solutions?",
      "votes": null
    },
    {
      "id": "3125735",
      "postDate": "02/16/2025 14:50:12",
      "content": "<p>I totally share the same concern. I was surprised to see all the test samples are provided, although the division between public and private leaderboard (LB) and the TTA groups remains unclear.</p>\n<p>Given that each day allows for three submissions (instead of two) and the test size is relatively small, I wouldn’t be surprised if the LB is compromised.</p>\n<p>I recommend to use a hidden test set for private evaluation. This would require additional efforts to change the current submission method from simply submitting a csv file to submitting a notebook for rerun on the hidden set. However, it would be worthwhile for the sake of ensuring a fair and meaningful competition.</p>",
      "rawMarkdown": "I totally share the same concern. I was surprised to see all the test samples are provided, although the division between public and private leaderboard (LB) and the TTA groups remains unclear.\n\nGiven that each day allows for three submissions (instead of two) and the test size is relatively small, I wouldn’t be surprised if the LB is compromised.\n\nI recommend to use a hidden test set for private evaluation. This would require additional efforts to change the current submission method from simply submitting a csv file to submitting a notebook for rerun on the hidden set. However, it would be worthwhile for the sake of ensuring a fair and meaningful competition.",
      "votes": null
    },
    {
      "id": "3125772",
      "postDate": "02/16/2025 15:51:52",
      "content": "<p>Leaderboard probing would backfire - winners must share full training code and reproduce results and explain method, so any gaming will be exposed and lead to disqualification.</p>",
      "rawMarkdown": "Leaderboard probing would backfire - winners must share full training code and reproduce results and explain method, so any gaming will be exposed and lead to disqualification.",
      "votes": null
    },
    {
      "id": "3125981",
      "postDate": "02/16/2025 20:28:13",
      "content": "<p>I believe cheaters can always find ways to make their code and results \"look good\", if they do win by such probing. It would also be too late to cache them if only after submitting the final results can we locate the cheaters. </p>\n<p>Anyways, since the 1344 test videos are already shared, at least everyone has the same chance to do probing. It is also a sort of fairness :)</p>",
      "rawMarkdown": "I believe cheaters can always find ways to make their code and results \"look good\", if they do win by such probing. It would also be too late to cache them if only after submitting the final results can we locate the cheaters. \n\nAnyways, since the 1344 test videos are already shared, at least everyone has the same chance to do probing. It is also a sort of fairness :)",
      "votes": null
    },
    {
      "id": "3125983",
      "postDate": "02/16/2025 20:31:31",
      "content": "<p>There are still ways participants could maneuver around it. One example is when the test samples are used to specifically improve the predictions (since predictions might be approximated possibly by driving experts); there would be no way of telling whether the participant came to those improvements by sheer intuition or whether they just abused the visibility of test target values.</p>",
      "rawMarkdown": "There are still ways participants could maneuver around it. One example is when the test samples are used to specifically improve the predictions (since predictions might be approximated possibly by driving experts); there would be no way of telling whether the participant came to those improvements by sheer intuition or whether they just abused the visibility of test target values.",
      "votes": null
    },
    {
      "id": "3125985",
      "postDate": "02/16/2025 20:33:49",
      "content": "<p>Other than fairness, another disadvantage is that the generalisation capability of the model cannot be properly evaluated if the full test set is simply given instead of hidden. One can easily do EDA on the test set and find the distribution of the test. Then the model  can be specifically trained on such distribution and thereby achieving a high score. </p>",
      "rawMarkdown": "Other than fairness, another disadvantage is that the generalisation capability of the model cannot be properly evaluated if the full test set is simply given instead of hidden. One can easily do EDA on the test set and find the distribution of the test. Then the model  can be specifically trained on such distribution and thereby achieving a high score.",
      "votes": null
    },
    {
      "id": "3125994",
      "postDate": "02/16/2025 20:49:43",
      "content": "<p>Per your comment above, \"wouldn't be surprised if the LB is compromised\" - I can assure you that I only used the training data (did an 80/20) train/val split, and results submitted is matching approximately what I get on my val set, and nearly each time the score on my val set improved it improved on my submission as well. My approach is not to focus so much on the model, more on \"which\" parts of the training videos/frames does the model really learn from - what is important, and which parts doesn't it learn from, using metrics collected per sample from each training run, and iterate on that - here is a screenshot from parts of my workflow to really understand whats going on based on the model I've chosen to use.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F1c49a89424b1459f20882091716f6de2%2FScreenshot%202025-02-16%20134557.png?generation=1739738844729652&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Per your comment above, \"wouldn't be surprised if the LB is compromised\" - I can assure you that I only used the training data (did an 80/20) train/val split, and results submitted is matching approximately what I get on my val set, and nearly each time the score on my val set improved it improved on my submission as well. My approach is not to focus so much on the model, more on \"which\" parts of the training videos/frames does the model really learn from - what is important, and which parts doesn't it learn from, using metrics collected per sample from each training run, and iterate on that - here is a screenshot from parts of my workflow to really understand whats going on based on the model I've chosen to use.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F1c49a89424b1459f20882091716f6de2%2FScreenshot%202025-02-16%20134557.png?generation=1739738844729652&alt=media)",
      "votes": null
    },
    {
      "id": "3126196",
      "postDate": "02/17/2025 05:39:14",
      "content": "<p>First of all, breathtaking visualizations. And the other thing is that no one is accusing you, I apologize if it appeared that way. It is a naive assumption to provide all of the data and expect EVERY participant to be fair and avoid cheating on the test set. </p>",
      "rawMarkdown": "First of all, breathtaking visualizations. And the other thing is that no one is accusing you, I apologize if it appeared that way. It is a naive assumption to provide all of the data and expect EVERY participant to be fair and avoid cheating on the test set.",
      "votes": null
    },
    {
      "id": "3126546",
      "postDate": "02/17/2025 13:23:51",
      "content": "<p>Almost offtopic. In other competitons I've seen retrain with pseudolabels on submission. Is pseudolabel (by model predictions, not human annotation) allowed?</p>",
      "rawMarkdown": "Almost offtopic. In other competitons I've seen retrain with pseudolabels on submission. Is pseudolabel (by model predictions, not human annotation) allowed?",
      "votes": null
    },
    {
      "id": "3126786",
      "postDate": "02/17/2025 17:06:11",
      "content": "<p>The hosting team shares similar concerns. These are the measures we have in place to try to mitigate the potential issues of having all test data available:</p>\n<ul>\n<li>We do not disclose which samples belong to the private and public sets. During the competition, you will receive a score based on the public part of the dataset. The final score will be based on the private part.</li>\n<li>We will examine the winning solutions (code and method description), and we should be able to replicate results using the provided code.</li>\n</ul>\n<p>We hope everybody plays fair and that the best team wins :-) </p>",
      "rawMarkdown": "The hosting team shares similar concerns. These are the measures we have in place to try to mitigate the potential issues of having all test data available:\n- We do not disclose which samples belong to the private and public sets. During the competition, you will receive a score based on the public part of the dataset. The final score will be based on the private part.\n- We will examine the winning solutions (code and method description), and we should be able to replicate results using the provided code.\n\nWe hope everybody plays fair and that the best team wins :-)",
      "votes": null
    },
    {
      "id": "3126790",
      "postDate": "02/17/2025 17:10:19",
      "content": "<p>I am not sure if I understand. If you want to run some kind of model to generate features that is OK. e.g. you can run a detector+tracker and use those detections in your inference. If you win, you should make all code available so that we can replicate the solution from scratch. </p>",
      "rawMarkdown": "I am not sure if I understand. If you want to run some kind of model to generate features that is OK. e.g. you can run a detector+tracker and use those detections in your inference. If you win, you should make all code available so that we can replicate the solution from scratch.",
      "votes": null
    },
    {
      "id": "3126875",
      "postDate": "02/17/2025 18:35:07",
      "content": "<p>Hi. Thanks for answer. I haven't started to train yet. That's my first video classification competition. I'm currently finishing the dataset and the augmentations I will apply. What I mean is to take the most confident predictions over test and retrain with them as pseudo labels. I've seen people win some decimals in other competitions with that. But since we have direct acces to test, the definition of confident predictions (the ones avobe certain threshold) can be suspicious.</p>",
      "rawMarkdown": "Hi. Thanks for answer. I haven't started to train yet. That's my first video classification competition. I'm currently finishing the dataset and the augmentations I will apply. What I mean is to take the most confident predictions over test and retrain with them as pseudo labels. I've seen people win some decimals in other competitions with that. But since we have direct acces to test, the definition of confident predictions (the ones avobe certain threshold) can be suspicious.",
      "votes": null
    },
    {
      "id": "3126894",
      "postDate": "02/17/2025 19:06:35",
      "content": "<p>Thank you for pointing this out!</p>\n<p>IMO the test data should never be part of the training (I will consult with the team to check if the rules should be updated). You can, however, split your training data into train/validation sets and use the technique you are suggesting (without touching the test set). </p>",
      "rawMarkdown": "Thank you for pointing this out!\n\nIMO the test data should never be part of the training (I will consult with the team to check if the rules should be updated). You can, however, split your training data into train/validation sets and use the technique you are suggesting (without touching the test set).",
      "votes": null
    },
    {
      "id": "3127502",
      "postDate": "02/18/2025 14:10:50",
      "content": "<p><a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a> We added the following to the rules to clarify that test data cannot be used in the training phase:</p>\n<blockquote>\n  <p>For the avoidance of doubt, it is strictly prohibited to use the test data for training the model in any way. Solutions found to have used test data during the evaluation phase will be disqualified.</p>\n</blockquote>",
      "rawMarkdown": "sacuscreed We added the following to the rules to clarify that test data cannot be used in the training phase:\n\n>For the avoidance of doubt, it is strictly prohibited to use the test data for training the model in any way. Solutions found to have used test data during the evaluation phase will be disqualified.",
      "votes": null
    },
    {
      "id": "3127605",
      "postDate": "02/18/2025 16:52:32",
      "content": "<p>maybe pin this message on top of the discussion forum? </p>",
      "rawMarkdown": "maybe pin this message on top of the discussion forum?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3125735,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "02/16/2025 14:50:12",
      "content": "<p>I totally share the same concern. I was surprised to see all the test samples are provided, although the division between public and private leaderboard (LB) and the TTA groups remains unclear.</p>\n<p>Given that each day allows for three submissions (instead of two) and the test size is relatively small, I wouldn’t be surprised if the LB is compromised.</p>\n<p>I recommend to use a hidden test set for private evaluation. This would require additional efforts to change the current submission method from simply submitting a csv file to submitting a notebook for rerun on the hidden set. However, it would be worthwhile for the sake of ensuring a fair and meaningful competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3125772,
          "author_name": "paulendresen76",
          "author_url": "",
          "post_date": "02/16/2025 15:51:52",
          "content": "<p>Leaderboard probing would backfire - winners must share full training code and reproduce results and explain method, so any gaming will be exposed and lead to disqualification.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3125981,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "02/16/2025 20:28:13",
              "content": "<p>I believe cheaters can always find ways to make their code and results \"look good\", if they do win by such probing. It would also be too late to cache them if only after submitting the final results can we locate the cheaters. </p>\n<p>Anyways, since the 1344 test videos are already shared, at least everyone has the same chance to do probing. It is also a sort of fairness :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3125994,
                  "author_name": "paulendresen76",
                  "author_url": "",
                  "post_date": "02/16/2025 20:49:43",
                  "content": "<p>Per your comment above, \"wouldn't be surprised if the LB is compromised\" - I can assure you that I only used the training data (did an 80/20) train/val split, and results submitted is matching approximately what I get on my val set, and nearly each time the score on my val set improved it improved on my submission as well. My approach is not to focus so much on the model, more on \"which\" parts of the training videos/frames does the model really learn from - what is important, and which parts doesn't it learn from, using metrics collected per sample from each training run, and iterate on that - here is a screenshot from parts of my workflow to really understand whats going on based on the model I've chosen to use.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F1c49a89424b1459f20882091716f6de2%2FScreenshot%202025-02-16%20134557.png?generation=1739738844729652&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3126196,
                      "author_name": "sabyrbazarymbetov",
                      "author_url": "",
                      "post_date": "02/17/2025 05:39:14",
                      "content": "<p>First of all, breathtaking visualizations. And the other thing is that no one is accusing you, I apologize if it appeared that way. It is a naive assumption to provide all of the data and expect EVERY participant to be fair and avoid cheating on the test set. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 3125983,
              "author_name": "sabyrbazarymbetov",
              "author_url": "",
              "post_date": "02/16/2025 20:31:31",
              "content": "<p>There are still ways participants could maneuver around it. One example is when the test samples are used to specifically improve the predictions (since predictions might be approximated possibly by driving experts); there would be no way of telling whether the participant came to those improvements by sheer intuition or whether they just abused the visibility of test target values.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3125985,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "02/16/2025 20:33:49",
      "content": "<p>Other than fairness, another disadvantage is that the generalisation capability of the model cannot be properly evaluated if the full test set is simply given instead of hidden. One can easily do EDA on the test set and find the distribution of the test. Then the model  can be specifically trained on such distribution and thereby achieving a high score. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3126546,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "02/17/2025 13:23:51",
      "content": "<p>Almost offtopic. In other competitons I've seen retrain with pseudolabels on submission. Is pseudolabel (by model predictions, not human annotation) allowed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3126790,
          "author_name": "danielcmoura",
          "author_url": "",
          "post_date": "02/17/2025 17:10:19",
          "content": "<p>I am not sure if I understand. If you want to run some kind of model to generate features that is OK. e.g. you can run a detector+tracker and use those detections in your inference. If you win, you should make all code available so that we can replicate the solution from scratch. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3126875,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "02/17/2025 18:35:07",
              "content": "<p>Hi. Thanks for answer. I haven't started to train yet. That's my first video classification competition. I'm currently finishing the dataset and the augmentations I will apply. What I mean is to take the most confident predictions over test and retrain with them as pseudo labels. I've seen people win some decimals in other competitions with that. But since we have direct acces to test, the definition of confident predictions (the ones avobe certain threshold) can be suspicious.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3126894,
                  "author_name": "danielcmoura",
                  "author_url": "",
                  "post_date": "02/17/2025 19:06:35",
                  "content": "<p>Thank you for pointing this out!</p>\n<p>IMO the test data should never be part of the training (I will consult with the team to check if the rules should be updated). You can, however, split your training data into train/validation sets and use the technique you are suggesting (without touching the test set). </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3127502,
                      "author_name": "danielcmoura",
                      "author_url": "",
                      "post_date": "02/18/2025 14:10:50",
                      "content": "<p><a href=\"https://www.kaggle.com/sacuscreed\" target=\"_blank\">@sacuscreed</a> We added the following to the rules to clarify that test data cannot be used in the training phase:</p>\n<blockquote>\n  <p>For the avoidance of doubt, it is strictly prohibited to use the test data for training the model in any way. Solutions found to have used test data during the evaluation phase will be disqualified.</p>\n</blockquote>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3127605,
                          "author_name": "shiyili",
                          "author_url": "",
                          "post_date": "02/18/2025 16:52:32",
                          "content": "<p>maybe pin this message on top of the discussion forum? </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3126786,
      "author_name": "danielcmoura",
      "author_url": "",
      "post_date": "02/17/2025 17:06:11",
      "content": "<p>The hosting team shares similar concerns. These are the measures we have in place to try to mitigate the potential issues of having all test data available:</p>\n<ul>\n<li>We do not disclose which samples belong to the private and public sets. During the competition, you will receive a score based on the public part of the dataset. The final score will be based on the private part.</li>\n<li>We will examine the winning solutions (code and method description), and we should be able to replicate results using the provided code.</li>\n</ul>\n<p>We hope everybody plays fair and that the best team wins :-) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3125552": "Does having all the the test videos available for participants not increase the risk of hard coding the solutions?",
    "3125735": "I totally share the same concern. I was surprised to see all the test samples are provided, although the division between public and private leaderboard (LB) and the TTA groups remains unclear.\n\nGiven that each day allows for three submissions (instead of two) and the test size is relatively small, I wouldn’t be surprised if the LB is compromised.\n\nI recommend to use a hidden test set for private evaluation. This would require additional efforts to change the current submission method from simply submitting a csv file to submitting a notebook for rerun on the hidden set. However, it would be worthwhile for the sake of ensuring a fair and meaningful competition.",
    "3125772": "Leaderboard probing would backfire - winners must share full training code and reproduce results and explain method, so any gaming will be exposed and lead to disqualification.",
    "3125981": "I believe cheaters can always find ways to make their code and results \"look good\", if they do win by such probing. It would also be too late to cache them if only after submitting the final results can we locate the cheaters. \n\nAnyways, since the 1344 test videos are already shared, at least everyone has the same chance to do probing. It is also a sort of fairness :)",
    "3125983": "There are still ways participants could maneuver around it. One example is when the test samples are used to specifically improve the predictions (since predictions might be approximated possibly by driving experts); there would be no way of telling whether the participant came to those improvements by sheer intuition or whether they just abused the visibility of test target values.",
    "3125985": "Other than fairness, another disadvantage is that the generalisation capability of the model cannot be properly evaluated if the full test set is simply given instead of hidden. One can easily do EDA on the test set and find the distribution of the test. Then the model  can be specifically trained on such distribution and thereby achieving a high score.",
    "3125994": "Per your comment above, \"wouldn't be surprised if the LB is compromised\" - I can assure you that I only used the training data (did an 80/20) train/val split, and results submitted is matching approximately what I get on my val set, and nearly each time the score on my val set improved it improved on my submission as well. My approach is not to focus so much on the model, more on \"which\" parts of the training videos/frames does the model really learn from - what is important, and which parts doesn't it learn from, using metrics collected per sample from each training run, and iterate on that - here is a screenshot from parts of my workflow to really understand whats going on based on the model I've chosen to use.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F1c49a89424b1459f20882091716f6de2%2FScreenshot%202025-02-16%20134557.png?generation=1739738844729652&alt=media)",
    "3126196": "First of all, breathtaking visualizations. And the other thing is that no one is accusing you, I apologize if it appeared that way. It is a naive assumption to provide all of the data and expect EVERY participant to be fair and avoid cheating on the test set.",
    "3126546": "Almost offtopic. In other competitons I've seen retrain with pseudolabels on submission. Is pseudolabel (by model predictions, not human annotation) allowed?",
    "3126786": "The hosting team shares similar concerns. These are the measures we have in place to try to mitigate the potential issues of having all test data available:\n- We do not disclose which samples belong to the private and public sets. During the competition, you will receive a score based on the public part of the dataset. The final score will be based on the private part.\n- We will examine the winning solutions (code and method description), and we should be able to replicate results using the provided code.\n\nWe hope everybody plays fair and that the best team wins :-)",
    "3126790": "I am not sure if I understand. If you want to run some kind of model to generate features that is OK. e.g. you can run a detector+tracker and use those detections in your inference. If you win, you should make all code available so that we can replicate the solution from scratch.",
    "3126875": "Hi. Thanks for answer. I haven't started to train yet. That's my first video classification competition. I'm currently finishing the dataset and the augmentations I will apply. What I mean is to take the most confident predictions over test and retrain with them as pseudo labels. I've seen people win some decimals in other competitions with that. But since we have direct acces to test, the definition of confident predictions (the ones avobe certain threshold) can be suspicious.",
    "3126894": "Thank you for pointing this out!\n\nIMO the test data should never be part of the training (I will consult with the team to check if the rules should be updated). You can, however, split your training data into train/validation sets and use the technique you are suggesting (without touching the test set).",
    "3127502": "sacuscreed We added the following to the rules to clarify that test data cannot be used in the training phase:\n\n>For the avoidance of doubt, it is strictly prohibited to use the test data for training the model in any way. Solutions found to have used test data during the evaluation phase will be disqualified.",
    "3127605": "maybe pin this message on top of the discussion forum?"
  },
  "source": "meta"
}