{
  "id": 384967,
  "title": "Sort your submission by event_id",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/384967",
  "author_name": "",
  "post_date": "2023-02-10T12:43:39.276758300Z",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all</p>\n<p>I recommend always sort your submission by event_id.</p>\n<p>Apparently, the organizers do not merge predict and truth (do not take into account the  \"event_id\" column) when calculating the metric. If the test batch files are not ordered or the insertion does not go through the index, the order of the  \"event_id\"  may be violated. </p>\n<p>In this case, you will get a score far from the desired one (for me applying df.sort_values([\"event_id\"]) to submission improved the LB: 1.381 -&gt; 1.174).</p>\n<p>Happy computing…</p>",
  "messages": [
    {
      "id": "2137931",
      "postDate": "02/10/2023 12:43:39",
      "content": "<p>Hi all</p>\n<p>I recommend always sort your submission by event_id.</p>\n<p>Apparently, the organizers do not merge predict and truth (do not take into account the  \"event_id\" column) when calculating the metric. If the test batch files are not ordered or the insertion does not go through the index, the order of the  \"event_id\"  may be violated. </p>\n<p>In this case, you will get a score far from the desired one (for me applying df.sort_values([\"event_id\"]) to submission improved the LB: 1.381 -&gt; 1.174).</p>\n<p>Happy computing…</p>",
      "rawMarkdown": "Hi all\n\nI recommend always sort your submission by event_id.\n\nApparently, the organizers do not merge predict and truth (do not take into account the  \"event_id\" column) when calculating the metric. If the test batch files are not ordered or the insertion does not go through the index, the order of the  \"event_id\"  may be violated. \n\nIn this case, you will get a score far from the desired one (for me applying df.sort_values([\"event_id\"]) to submission improved the LB: 1.381 -> 1.174).\n\nHappy computing...",
      "votes": null
    },
    {
      "id": "2138299",
      "postDate": "02/10/2023 17:24:02",
      "content": "<p>I wasn't quite sure what you meant until I read this on the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data?select=sample_submission.parquet\" target=\"_blank\">Data page</a>, under the description of the 'sample_submission.parquet' file:</p>\n<blockquote>\n  <p>An example submission with the correct columns and <em>properly ordered</em> event IDs.</p>\n</blockquote>\n<p>In that sample file, the three sample events are stored in ascending numerical order by event_id: 2092, 7344 and finally 9482.</p>\n<p>Also remember that your final submission <strong>must be a csv</strong>.</p>",
      "rawMarkdown": "I wasn't quite sure what you meant until I read this on the [Data page](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data?select=sample_submission.parquet), under the description of the 'sample_submission.parquet' file:\n\n>An example submission with the correct columns and *properly ordered* event IDs.\n\nIn that sample file, the three sample events are stored in ascending numerical order by event_id: 2092, 7344 and finally 9482.\n\nAlso remember that your final submission **must be a csv**.",
      "votes": null
    },
    {
      "id": "2138320",
      "postDate": "02/10/2023 17:37:10",
      "content": "<p>Have not verified this finding yet - but I do know that on a few occasions over the years I have been doing kaggle competitions there have been a few were the order of the submission was an issue for me.  Events can easily get out of order if your doing multiprocessing.</p>",
      "rawMarkdown": "Have not verified this finding yet - but I do know that on a few occasions over the years I have been doing kaggle competitions there have been a few were the order of the submission was an issue for me.  Events can easily get out of order if your doing multiprocessing.",
      "votes": null
    },
    {
      "id": "2138335",
      "postDate": "02/10/2023 17:50:39",
      "content": "<p>I think this is a question for <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "I think this is a question for @sohier",
      "votes": null
    },
    {
      "id": "2140433",
      "postDate": "02/11/2023 18:38:27",
      "content": "<p>A fork of the <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">lighting fast baseline</a> does have a bad score (1.558) when the rows are sorted in reverse order by event.  The submission save used by polars does include the event_id - so it would appear that kaggle scoring is ignoring event_id in the submission file and your submission needs to be in the correct order.</p>",
      "rawMarkdown": "A fork of the [lighting fast baseline](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars) does have a bad score (1.558) when the rows are sorted in reverse order by event.  The submission save used by polars does include the event_id - so it would appear that kaggle scoring is ignoring event_id in the submission file and your submission needs to be in the correct order.",
      "votes": null
    },
    {
      "id": "2142297",
      "postDate": "02/13/2023 12:37:37",
      "content": "<p>It is always a good practice to sort your submissions by the event_id column <strong>to ensure that the order of the event_id values is not violated</strong>. <br>\nThis is because the organizers may not merge the predict and truth files based on the event_id column when calculating the metric. By sorting the submissions based on the event_id column, you can avoid getting a score far from the desired one and improve your leaderboard score.<br>\nMaybe I am wrong somehow….</p>",
      "rawMarkdown": "It is always a good practice to sort your submissions by the event_id column **to ensure that the order of the event_id values is not violated**. \nThis is because the organizers may not merge the predict and truth files based on the event_id column when calculating the metric. By sorting the submissions based on the event_id column, you can avoid getting a score far from the desired one and improve your leaderboard score.\nMaybe I am wrong somehow....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2138299,
      "author_name": "michaelbarrett",
      "author_url": "",
      "post_date": "02/10/2023 17:24:02",
      "content": "<p>I wasn't quite sure what you meant until I read this on the <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data?select=sample_submission.parquet\" target=\"_blank\">Data page</a>, under the description of the 'sample_submission.parquet' file:</p>\n<blockquote>\n  <p>An example submission with the correct columns and <em>properly ordered</em> event IDs.</p>\n</blockquote>\n<p>In that sample file, the three sample events are stored in ascending numerical order by event_id: 2092, 7344 and finally 9482.</p>\n<p>Also remember that your final submission <strong>must be a csv</strong>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2138320,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/10/2023 17:37:10",
      "content": "<p>Have not verified this finding yet - but I do know that on a few occasions over the years I have been doing kaggle competitions there have been a few were the order of the submission was an issue for me.  Events can easily get out of order if your doing multiprocessing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2140433,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "02/11/2023 18:38:27",
          "content": "<p>A fork of the <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">lighting fast baseline</a> does have a bad score (1.558) when the rows are sorted in reverse order by event.  The submission save used by polars does include the event_id - so it would appear that kaggle scoring is ignoring event_id in the submission file and your submission needs to be in the correct order.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2138335,
      "author_name": "pellerphys",
      "author_url": "",
      "post_date": "02/10/2023 17:50:39",
      "content": "<p>I think this is a question for <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2142297,
      "author_name": "giranntu",
      "author_url": "",
      "post_date": "02/13/2023 12:37:37",
      "content": "<p>It is always a good practice to sort your submissions by the event_id column <strong>to ensure that the order of the event_id values is not violated</strong>. <br>\nThis is because the organizers may not merge the predict and truth files based on the event_id column when calculating the metric. By sorting the submissions based on the event_id column, you can avoid getting a score far from the desired one and improve your leaderboard score.<br>\nMaybe I am wrong somehow….</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2137931": "Hi all\n\nI recommend always sort your submission by event_id.\n\nApparently, the organizers do not merge predict and truth (do not take into account the  \"event_id\" column) when calculating the metric. If the test batch files are not ordered or the insertion does not go through the index, the order of the  \"event_id\"  may be violated. \n\nIn this case, you will get a score far from the desired one (for me applying df.sort_values([\"event_id\"]) to submission improved the LB: 1.381 -> 1.174).\n\nHappy computing...",
    "2138299": "I wasn't quite sure what you meant until I read this on the [Data page](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data?select=sample_submission.parquet), under the description of the 'sample_submission.parquet' file:\n\n>An example submission with the correct columns and *properly ordered* event IDs.\n\nIn that sample file, the three sample events are stored in ascending numerical order by event_id: 2092, 7344 and finally 9482.\n\nAlso remember that your final submission **must be a csv**.",
    "2138320": "Have not verified this finding yet - but I do know that on a few occasions over the years I have been doing kaggle competitions there have been a few were the order of the submission was an issue for me.  Events can easily get out of order if your doing multiprocessing.",
    "2138335": "I think this is a question for @sohier",
    "2140433": "A fork of the [lighting fast baseline](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars) does have a bad score (1.558) when the rows are sorted in reverse order by event.  The submission save used by polars does include the event_id - so it would appear that kaggle scoring is ignoring event_id in the submission file and your submission needs to be in the correct order.",
    "2142297": "It is always a good practice to sort your submissions by the event_id column **to ensure that the order of the event_id values is not violated**. \nThis is because the organizers may not merge the predict and truth files based on the event_id column when calculating the metric. By sorting the submissions based on the event_id column, you can avoid getting a score far from the desired one and improve your leaderboard score.\nMaybe I am wrong somehow...."
  },
  "source": "meta"
}