{
  "id": 456083,
  "title": "Something Interesting About Test Data.",
  "url": "/competitions/predict-ai-model-runtime/discussion/456083",
  "author_name": "",
  "post_date": "2023-11-18T00:56:57.681961300Z",
  "votes": 19,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Has anyone else encountered this situation? Using the prediction results of <code>random</code> as the <code>default</code> results can greatly improve both lb and pb. However, if the model is used to directly predict the <code>default</code>, the effect will be much worse. Something wrong with the test data? What is the reason for this situation?😮😮😮</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10671485%2Fccc914dcad5f07cf3c93a38fe4510537%2F_20231117165331.png?generation=1700268827943621&amp;alt=media\" alt=\"nlp random replaced by nlp default\"></p>\n<p>For the below one, <code>nlp random</code> was replaced by <code>nlp default</code>, but got higher score on both public leaderboard and private leaderboard.</p>",
  "messages": [
    {
      "id": "2529126",
      "postDate": "11/18/2023 00:56:57",
      "content": "<p>Has anyone else encountered this situation? Using the prediction results of <code>random</code> as the <code>default</code> results can greatly improve both lb and pb. However, if the model is used to directly predict the <code>default</code>, the effect will be much worse. Something wrong with the test data? What is the reason for this situation?😮😮😮</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10671485%2Fccc914dcad5f07cf3c93a38fe4510537%2F_20231117165331.png?generation=1700268827943621&amp;alt=media\" alt=\"nlp random replaced by nlp default\"></p>\n<p>For the below one, <code>nlp random</code> was replaced by <code>nlp default</code>, but got higher score on both public leaderboard and private leaderboard.</p>",
      "rawMarkdown": "Has anyone else encountered this situation? Using the prediction results of ``random`` as the ``default`` results can greatly improve both lb and pb. However, if the model is used to directly predict the ``default``, the effect will be much worse. Something wrong with the test data? What is the reason for this situation?😮😮😮\n\n![nlp random replaced by nlp default](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10671485%2Fccc914dcad5f07cf3c93a38fe4510537%2F_20231117165331.png?generation=1700268827943621&alt=media)\n\nFor the below one, ``nlp random`` was replaced by ``nlp default``, but got higher score on both public leaderboard and private leaderboard.",
      "votes": null
    },
    {
      "id": "2529129",
      "postDate": "11/18/2023 01:05:58",
      "content": "<p>Yes, my CV and LB were also better when I copied the test predictions for the random sets to the default sets.</p>\n<p>Maybe the random search strategy leads to more diverse configurations and is consequently better for training?</p>",
      "rawMarkdown": "Yes, my CV and LB were also better when I copied the test predictions for the random sets to the default sets.\n\nMaybe the random search strategy leads to more diverse configurations and is consequently better for training?",
      "votes": null
    },
    {
      "id": "2529134",
      "postDate": "11/18/2023 01:21:25",
      "content": "<p>Maybe, but the final result can be overfitting.😂</p>",
      "rawMarkdown": "Maybe, but the final result can be overfitting.😂",
      "votes": null
    },
    {
      "id": "2529136",
      "postDate": "11/18/2023 01:22:54",
      "content": "<p>i identify the common subset of config and node from xla:random and xla:default in the train set.<br>\nalthough the input is the same, the ground truth runtime is different. (the runtime ranking is also different)</p>",
      "rawMarkdown": "i identify the common subset of config and node from xla:random and xla:default in the train set.\nalthough the input is the same, the ground truth runtime is different. (the runtime ranking is also different)",
      "votes": null
    },
    {
      "id": "2529137",
      "postDate": "11/18/2023 01:23:39",
      "content": "<p>Thanks for sharing. We didn't notice it before your sharing. I used prediction(pb: 0.710) results of random as the default results to make a late submission.The improve is amazing. And I have tried to use random model to predict default, but cv didn't improve. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1069377%2Fe29b89ec1f335d4aeae5fe097fd9fcad%2F_20231118091949.png?generation=1700270405776237&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for sharing. We didn't notice it before your sharing. I used prediction(pb: 0.710) results of random as the default results to make a late submission.The improve is amazing. And I have tried to use random model to predict default, but cv didn't improve. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1069377%2Fe29b89ec1f335d4aeae5fe097fd9fcad%2F_20231118091949.png?generation=1700270405776237&alt=media)",
      "votes": null
    },
    {
      "id": "2529138",
      "postDate": "11/18/2023 01:27:40",
      "content": "<p>Congratulations on winning the gold medal！</p>",
      "rawMarkdown": "Congratulations on winning the gold medal！",
      "votes": null
    },
    {
      "id": "2529139",
      "postDate": "11/18/2023 01:28:00",
      "content": "<p>oops! 0.794 is very good!<br>\n(is the same trend observed in the validation data?)</p>",
      "rawMarkdown": "oops! 0.794 is very good!\n(is the same trend observed in the validation data?)",
      "votes": null
    },
    {
      "id": "2529140",
      "postDate": "11/18/2023 01:29:55",
      "content": "<p>Yes, that might be the reason. Congratulations on winning the gold medal！</p>",
      "rawMarkdown": "Yes, that might be the reason. Congratulations on winning the gold medal！",
      "votes": null
    },
    {
      "id": "2529143",
      "postDate": "11/18/2023 01:35:31",
      "content": "<p>I tried it too (just copied the nlp/random result directly to the nlp/default result)<br>\nI guess that the order of correct answers is the same in random and default.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fb0e581fcdd32428cac58eba953e0d737%2F2023-11-18%20103214.png?generation=1700271327130903&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I tried it too (just copied the nlp/random result directly to the nlp/default result)\nI guess that the order of correct answers is the same in random and default.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fb0e581fcdd32428cac58eba953e0d737%2F2023-11-18%20103214.png?generation=1700271327130903&alt=media)",
      "votes": null
    },
    {
      "id": "2529149",
      "postDate": "11/18/2023 01:45:36",
      "content": "<p>i think i read in some paper that says the sample configuration from random is derived from default set using genetic algorithm.<br>\nhence random set can be seen as an augmentation from default set (in the input space)?</p>\n<p>random set  spans larger congfiguration space and should be more stable?</p>\n<p>but there is one issue when we consider the output space, i think the runtime is measured and subjected to measurement noise.<br>\n(i.e. if you run the graph many times the measured runtime is different. i am not sure if the provided runtime ground truth has been averaged over many runs or just a single run results)</p>",
      "rawMarkdown": "i think i read in some paper that says the sample configuration from random is derived from default set using genetic algorithm.\nhence random set can be seen as an augmentation from default set (in the input space)?\n\nrandom set  spans larger congfiguration space and should be more stable?\n\nbut there is one issue when we consider the output space, i think the runtime is measured and subjected to measurement noise.\n(i.e. if you run the graph many times the measured runtime is different. i am not sure if the provided runtime ground truth has been averaged over many runs or just a single run results)",
      "votes": null
    },
    {
      "id": "2529150",
      "postDate": "11/18/2023 01:46:12",
      "content": "<p>Amazing😂!</p>",
      "rawMarkdown": "Amazing😂!",
      "votes": null
    },
    {
      "id": "2529162",
      "postDate": "11/18/2023 02:09:30",
      "content": "<p>My result 😂<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F73b51b2f637353fd60b751758c220bdc%2FF_Lns4dbcAAvDSn.jpeg?generation=1700273347843016&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "My result 😂\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F73b51b2f637353fd60b751758c220bdc%2FF_Lns4dbcAAvDSn.jpeg?generation=1700273347843016&alt=media)",
      "votes": null
    },
    {
      "id": "2529174",
      "postDate": "11/18/2023 02:28:24",
      "content": "<p>We also only found out about this now. Probably some data preparation issue like other people suggested.</p>\n<p>Just tried here and our score also improves a lot by doing it. Crazy that is unlikely that someone found this before the deadline, otherwise we probably would have seen scores above 0.8.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2Fff5732646fb6d386575d19989f0458e6%2Fss_score.png?generation=1700274112296160&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "We also only found out about this now. Probably some data preparation issue like other people suggested.\n\nJust tried here and our score also improves a lot by doing it. Crazy that is unlikely that someone found this before the deadline, otherwise we probably would have seen scores above 0.8.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2Fff5732646fb6d386575d19989f0458e6%2Fss_score.png?generation=1700274112296160&alt=media)",
      "votes": null
    },
    {
      "id": "2529206",
      "postDate": "11/18/2023 03:43:06",
      "content": "<p>I confirm this observation, as I am looking at my code that generated the final dataset files. Indeed, the <strong>indices</strong> that sort the runtime of any <code>nlp/random/test/&lt;filename&gt;.npz</code> are the same indices that sort its corresponding <code>nlp/default/test/&lt;filename&gt;.npz</code> -- this was by (mal)design -- we took 1000 configurations, sorted them, then shuffled them by initializing a random seed that is based on the <code>os.path.basename(&lt;filename&gt;)</code> :D</p>\n<p>Thank you for noticing it (we should've caught it ourselves) -- we should've shuffled on the entire relative path rather than only <code>basename</code>.</p>\n<p>For my own curiosity, I can check if any of the prize winners submitted the indices for random and for default.</p>",
      "rawMarkdown": "I confirm this observation, as I am looking at my code that generated the final dataset files. Indeed, the **indices** that sort the runtime of any `nlp/random/test/<filename>.npz` are the same indices that sort its corresponding `nlp/default/test/<filename>.npz` -- this was by (mal)design -- we took 1000 configurations, sorted them, then shuffled them by initializing a random seed that is based on the `os.path.basename(<filename>)` :D\n\nThank you for noticing it (we should've caught it ourselves) -- we should've shuffled on the entire relative path rather than only `basename`.\n\nFor my own curiosity, I can check if any of the prize winners submitted the indices for random and for default.",
      "votes": null
    },
    {
      "id": "2529212",
      "postDate": "11/18/2023 03:50:04",
      "content": "<p>Yes, you may check if any of the teams that won a prize used this method. This might indicate that their models are not necessarily robust enough, as it concerns the issue of prize money, which indeed requires some caution.</p>",
      "rawMarkdown": "Yes, you may check if any of the teams that won a prize used this method. This might indicate that their models are not necessarily robust enough, as it concerns the issue of prize money, which indeed requires some caution.",
      "votes": null
    },
    {
      "id": "2529219",
      "postDate": "11/18/2023 03:55:26",
      "content": "<p>There should also have been instances of data leakage in the history of Kaggle, but it seems that exploiting this is considered permissible.</p>",
      "rawMarkdown": "There should also have been instances of data leakage in the history of Kaggle, but it seems that exploiting this is considered permissible.",
      "votes": null
    },
    {
      "id": "2529783",
      "postDate": "11/18/2023 14:38:12",
      "content": "<p>Noted 👍. Thanks for sharing </p>",
      "rawMarkdown": "Noted 👍. Thanks for sharing",
      "votes": null
    },
    {
      "id": "2530054",
      "postDate": "11/18/2023 18:38:31",
      "content": "<p>Google is fast but not as human, because google created by humans</p>",
      "rawMarkdown": "Google is fast but not as human, because google created by humans",
      "votes": null
    },
    {
      "id": "2530078",
      "postDate": "11/18/2023 18:53:16",
      "content": "<p>Olala, that is quite something! Didn't know about it either and would raise my score quite a bit as well</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10842317%2Fe0f32fff6f939f8437ac32c134b38e1d%2Fnew_score.png?generation=1700333516259810&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Olala, that is quite something! Didn't know about it either and would raise my score quite a bit as well\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10842317%2Fe0f32fff6f939f8437ac32c134b38e1d%2Fnew_score.png?generation=1700333516259810&alt=media)",
      "votes": null
    },
    {
      "id": "2530091",
      "postDate": "11/18/2023 19:01:12",
      "content": "<p>Great work! Congratulations on winning the gold medal！<br>\nActually I think most people didn't find this interesting thing. 🤔</p>",
      "rawMarkdown": "Great work! Congratulations on winning the gold medal！\nActually I think most people didn't find this interesting thing. 🤔",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2529129,
      "author_name": "azyssd1",
      "author_url": "",
      "post_date": "11/18/2023 01:05:58",
      "content": "<p>Yes, my CV and LB were also better when I copied the test predictions for the random sets to the default sets.</p>\n<p>Maybe the random search strategy leads to more diverse configurations and is consequently better for training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2529134,
          "author_name": "lizhecheng",
          "author_url": "",
          "post_date": "11/18/2023 01:21:25",
          "content": "<p>Maybe, but the final result can be overfitting.😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529136,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2023 01:22:54",
      "content": "<p>i identify the common subset of config and node from xla:random and xla:default in the train set.<br>\nalthough the input is the same, the ground truth runtime is different. (the runtime ranking is also different)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2529140,
          "author_name": "lizhecheng",
          "author_url": "",
          "post_date": "11/18/2023 01:29:55",
          "content": "<p>Yes, that might be the reason. Congratulations on winning the gold medal！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529137,
      "author_name": "chenxin1991",
      "author_url": "",
      "post_date": "11/18/2023 01:23:39",
      "content": "<p>Thanks for sharing. We didn't notice it before your sharing. I used prediction(pb: 0.710) results of random as the default results to make a late submission.The improve is amazing. And I have tried to use random model to predict default, but cv didn't improve. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1069377%2Fe29b89ec1f335d4aeae5fe097fd9fcad%2F_20231118091949.png?generation=1700270405776237&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2529138,
          "author_name": "lizhecheng",
          "author_url": "",
          "post_date": "11/18/2023 01:27:40",
          "content": "<p>Congratulations on winning the gold medal！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2529139,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/18/2023 01:28:00",
          "content": "<p>oops! 0.794 is very good!<br>\n(is the same trend observed in the validation data?)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529143,
      "author_name": "shunrcn",
      "author_url": "",
      "post_date": "11/18/2023 01:35:31",
      "content": "<p>I tried it too (just copied the nlp/random result directly to the nlp/default result)<br>\nI guess that the order of correct answers is the same in random and default.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fb0e581fcdd32428cac58eba953e0d737%2F2023-11-18%20103214.png?generation=1700271327130903&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2529150,
          "author_name": "lizhecheng",
          "author_url": "",
          "post_date": "11/18/2023 01:46:12",
          "content": "<p>Amazing😂!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529149,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2023 01:45:36",
      "content": "<p>i think i read in some paper that says the sample configuration from random is derived from default set using genetic algorithm.<br>\nhence random set can be seen as an augmentation from default set (in the input space)?</p>\n<p>random set  spans larger congfiguration space and should be more stable?</p>\n<p>but there is one issue when we consider the output space, i think the runtime is measured and subjected to measurement noise.<br>\n(i.e. if you run the graph many times the measured runtime is different. i am not sure if the provided runtime ground truth has been averaged over many runs or just a single run results)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529162,
      "author_name": "knshnb",
      "author_url": "",
      "post_date": "11/18/2023 02:09:30",
      "content": "<p>My result 😂<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F73b51b2f637353fd60b751758c220bdc%2FF_Lns4dbcAAvDSn.jpeg?generation=1700273347843016&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529174,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "11/18/2023 02:28:24",
      "content": "<p>We also only found out about this now. Probably some data preparation issue like other people suggested.</p>\n<p>Just tried here and our score also improves a lot by doing it. Crazy that is unlikely that someone found this before the deadline, otherwise we probably would have seen scores above 0.8.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2Fff5732646fb6d386575d19989f0458e6%2Fss_score.png?generation=1700274112296160&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2529206,
      "author_name": "samihaija",
      "author_url": "",
      "post_date": "11/18/2023 03:43:06",
      "content": "<p>I confirm this observation, as I am looking at my code that generated the final dataset files. Indeed, the <strong>indices</strong> that sort the runtime of any <code>nlp/random/test/&lt;filename&gt;.npz</code> are the same indices that sort its corresponding <code>nlp/default/test/&lt;filename&gt;.npz</code> -- this was by (mal)design -- we took 1000 configurations, sorted them, then shuffled them by initializing a random seed that is based on the <code>os.path.basename(&lt;filename&gt;)</code> :D</p>\n<p>Thank you for noticing it (we should've caught it ourselves) -- we should've shuffled on the entire relative path rather than only <code>basename</code>.</p>\n<p>For my own curiosity, I can check if any of the prize winners submitted the indices for random and for default.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2529212,
          "author_name": "lizhecheng",
          "author_url": "",
          "post_date": "11/18/2023 03:50:04",
          "content": "<p>Yes, you may check if any of the teams that won a prize used this method. This might indicate that their models are not necessarily robust enough, as it concerns the issue of prize money, which indeed requires some caution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2529219,
          "author_name": "lizhecheng",
          "author_url": "",
          "post_date": "11/18/2023 03:55:26",
          "content": "<p>There should also have been instances of data leakage in the history of Kaggle, but it seems that exploiting this is considered permissible.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2529783,
      "author_name": "ankitanain",
      "author_url": "",
      "post_date": "11/18/2023 14:38:12",
      "content": "<p>Noted 👍. Thanks for sharing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2530054,
      "author_name": "vikrampython",
      "author_url": "",
      "post_date": "11/18/2023 18:38:31",
      "content": "<p>Google is fast but not as human, because google created by humans</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2530078,
      "author_name": "janisfluri",
      "author_url": "",
      "post_date": "11/18/2023 18:53:16",
      "content": "<p>Olala, that is quite something! Didn't know about it either and would raise my score quite a bit as well</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10842317%2Fe0f32fff6f939f8437ac32c134b38e1d%2Fnew_score.png?generation=1700333516259810&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2530091,
          "author_name": "lizhecheng",
          "author_url": "",
          "post_date": "11/18/2023 19:01:12",
          "content": "<p>Great work! Congratulations on winning the gold medal！<br>\nActually I think most people didn't find this interesting thing. 🤔</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2529126": "Has anyone else encountered this situation? Using the prediction results of ``random`` as the ``default`` results can greatly improve both lb and pb. However, if the model is used to directly predict the ``default``, the effect will be much worse. Something wrong with the test data? What is the reason for this situation?😮😮😮\n\n![nlp random replaced by nlp default](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10671485%2Fccc914dcad5f07cf3c93a38fe4510537%2F_20231117165331.png?generation=1700268827943621&alt=media)\n\nFor the below one, ``nlp random`` was replaced by ``nlp default``, but got higher score on both public leaderboard and private leaderboard.",
    "2529129": "Yes, my CV and LB were also better when I copied the test predictions for the random sets to the default sets.\n\nMaybe the random search strategy leads to more diverse configurations and is consequently better for training?",
    "2529134": "Maybe, but the final result can be overfitting.😂",
    "2529136": "i identify the common subset of config and node from xla:random and xla:default in the train set.\nalthough the input is the same, the ground truth runtime is different. (the runtime ranking is also different)",
    "2529137": "Thanks for sharing. We didn't notice it before your sharing. I used prediction(pb: 0.710) results of random as the default results to make a late submission.The improve is amazing. And I have tried to use random model to predict default, but cv didn't improve. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1069377%2Fe29b89ec1f335d4aeae5fe097fd9fcad%2F_20231118091949.png?generation=1700270405776237&alt=media)",
    "2529138": "Congratulations on winning the gold medal！",
    "2529139": "oops! 0.794 is very good!\n(is the same trend observed in the validation data?)",
    "2529140": "Yes, that might be the reason. Congratulations on winning the gold medal！",
    "2529143": "I tried it too (just copied the nlp/random result directly to the nlp/default result)\nI guess that the order of correct answers is the same in random and default.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fb0e581fcdd32428cac58eba953e0d737%2F2023-11-18%20103214.png?generation=1700271327130903&alt=media)",
    "2529149": "i think i read in some paper that says the sample configuration from random is derived from default set using genetic algorithm.\nhence random set can be seen as an augmentation from default set (in the input space)?\n\nrandom set  spans larger congfiguration space and should be more stable?\n\nbut there is one issue when we consider the output space, i think the runtime is measured and subjected to measurement noise.\n(i.e. if you run the graph many times the measured runtime is different. i am not sure if the provided runtime ground truth has been averaged over many runs or just a single run results)",
    "2529150": "Amazing😂!",
    "2529162": "My result 😂\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F73b51b2f637353fd60b751758c220bdc%2FF_Lns4dbcAAvDSn.jpeg?generation=1700273347843016&alt=media)",
    "2529174": "We also only found out about this now. Probably some data preparation issue like other people suggested.\n\nJust tried here and our score also improves a lot by doing it. Crazy that is unlikely that someone found this before the deadline, otherwise we probably would have seen scores above 0.8.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1648129%2Fff5732646fb6d386575d19989f0458e6%2Fss_score.png?generation=1700274112296160&alt=media)",
    "2529206": "I confirm this observation, as I am looking at my code that generated the final dataset files. Indeed, the **indices** that sort the runtime of any `nlp/random/test/<filename>.npz` are the same indices that sort its corresponding `nlp/default/test/<filename>.npz` -- this was by (mal)design -- we took 1000 configurations, sorted them, then shuffled them by initializing a random seed that is based on the `os.path.basename(<filename>)` :D\n\nThank you for noticing it (we should've caught it ourselves) -- we should've shuffled on the entire relative path rather than only `basename`.\n\nFor my own curiosity, I can check if any of the prize winners submitted the indices for random and for default.",
    "2529212": "Yes, you may check if any of the teams that won a prize used this method. This might indicate that their models are not necessarily robust enough, as it concerns the issue of prize money, which indeed requires some caution.",
    "2529219": "There should also have been instances of data leakage in the history of Kaggle, but it seems that exploiting this is considered permissible.",
    "2529783": "Noted 👍. Thanks for sharing",
    "2530054": "Google is fast but not as human, because google created by humans",
    "2530078": "Olala, that is quite something! Didn't know about it either and would raise my score quite a bit as well\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10842317%2Fe0f32fff6f939f8437ac32c134b38e1d%2Fnew_score.png?generation=1700333516259810&alt=media)",
    "2530091": "Great work! Congratulations on winning the gold medal！\nActually I think most people didn't find this interesting thing. 🤔"
  },
  "source": "meta"
}