{
  "id": 496903,
  "title": "Does the scoring code ignore row_id? (does row submission order matter?)",
  "url": "/competitions/birdclef-2024/discussion/496903",
  "author_name": "",
  "post_date": "2024-04-22T22:42:15.070509200Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>BirdCLEF 2024's homepage links to this notebook to describe scoring:<br>\n<a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">https://www.kaggle.com/code/metric/birdclef-roc-auc</a></p>\n<p>The very first thing the scoring code does - is delete \"row_id_column_name\":</p>\n<pre><code> () -&gt; :\n    \n     solution[row_id_column_name]\n     submission[row_id_column_name]\n</code></pre>\n<p>Which is fine - assuming everyone is in agreement on file order…  </p>\n<p>My submit notebook fetches the files like this:<br>\n<code>filenames_with_path = glob.glob(f\"{soundscapes_folder}/*.ogg\")</code></p>\n<p>Which produces a file list that has an order that corresponds to nothing in particular (which didn't concern me - as I assumed row_id was being used).</p>\n<p>Are the solution and submission dataframes re-sorted at some point so they are assured to have a matching order?</p>\n<p>(I'm kind of inclined to randomize the order of my rows on a submission - and see what happens…)</p>\n<p>-Rich</p>",
  "messages": [
    {
      "id": "2768523",
      "postDate": "04/22/2024 22:42:15",
      "content": "<p>BirdCLEF 2024's homepage links to this notebook to describe scoring:<br>\n<a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">https://www.kaggle.com/code/metric/birdclef-roc-auc</a></p>\n<p>The very first thing the scoring code does - is delete \"row_id_column_name\":</p>\n<pre><code> () -&gt; :\n    \n     solution[row_id_column_name]\n     submission[row_id_column_name]\n</code></pre>\n<p>Which is fine - assuming everyone is in agreement on file order…  </p>\n<p>My submit notebook fetches the files like this:<br>\n<code>filenames_with_path = glob.glob(f\"{soundscapes_folder}/*.ogg\")</code></p>\n<p>Which produces a file list that has an order that corresponds to nothing in particular (which didn't concern me - as I assumed row_id was being used).</p>\n<p>Are the solution and submission dataframes re-sorted at some point so they are assured to have a matching order?</p>\n<p>(I'm kind of inclined to randomize the order of my rows on a submission - and see what happens…)</p>\n<p>-Rich</p>",
      "rawMarkdown": "BirdCLEF 2024's homepage links to this notebook to describe scoring:\nhttps://www.kaggle.com/code/metric/birdclef-roc-auc\n\nThe very first thing the scoring code does - is delete \"row_id_column_name\":\n\n```python\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str) -> float:\n    '''\n    Version of macro-averaged ROC-AUC score that ignores all classes that have no true positive labels.\n    '''\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n```\n\nWhich is fine - assuming everyone is in agreement on file order...  \n\nMy submit notebook fetches the files like this:\n`filenames_with_path = glob.glob(f\"{soundscapes_folder}/*.ogg\")`\n\nWhich produces a file list that has an order that corresponds to nothing in particular (which didn't concern me - as I assumed row_id was being used).\n\nAre the solution and submission dataframes re-sorted at some point so they are assured to have a matching order?\n\n(I'm kind of inclined to randomize the order of my rows on a submission - and see what happens...)\n\n-Rich",
      "votes": null
    },
    {
      "id": "2768781",
      "postDate": "04/23/2024 03:12:04",
      "content": "<p>Just did 2 test submissions.</p>\n<p>1st one had random numbers for row_id - and it got rejected for having an invalid format.</p>\n<p>2nd test took a previous .61 LB notebook - and just resorted the predictions (maintaining column order, but reversing the rows).  The score dropped to .47 on the LB.</p>\n<p>So - there's presumably some kind of filter before the shared scoring code that deals with row_id.</p>",
      "rawMarkdown": "Just did 2 test submissions.\n\n1st one had random numbers for row_id - and it got rejected for having an invalid format.\n\n2nd test took a previous .61 LB notebook - and just resorted the predictions (maintaining column order, but reversing the rows).  The score dropped to .47 on the LB.\n\nSo - there's presumably some kind of filter before the shared scoring code that deals with row_id.",
      "votes": null
    },
    {
      "id": "2769715",
      "postDate": "04/23/2024 14:09:36",
      "content": "<p>Thanks, this is reassuring</p>",
      "rawMarkdown": "Thanks, this is reassuring",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2768781,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "04/23/2024 03:12:04",
      "content": "<p>Just did 2 test submissions.</p>\n<p>1st one had random numbers for row_id - and it got rejected for having an invalid format.</p>\n<p>2nd test took a previous .61 LB notebook - and just resorted the predictions (maintaining column order, but reversing the rows).  The score dropped to .47 on the LB.</p>\n<p>So - there's presumably some kind of filter before the shared scoring code that deals with row_id.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2769715,
          "author_name": "janmpia",
          "author_url": "",
          "post_date": "04/23/2024 14:09:36",
          "content": "<p>Thanks, this is reassuring</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2768523": "BirdCLEF 2024's homepage links to this notebook to describe scoring:\nhttps://www.kaggle.com/code/metric/birdclef-roc-auc\n\nThe very first thing the scoring code does - is delete \"row_id_column_name\":\n\n```python\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str) -> float:\n    '''\n    Version of macro-averaged ROC-AUC score that ignores all classes that have no true positive labels.\n    '''\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n```\n\nWhich is fine - assuming everyone is in agreement on file order...  \n\nMy submit notebook fetches the files like this:\n`filenames_with_path = glob.glob(f\"{soundscapes_folder}/*.ogg\")`\n\nWhich produces a file list that has an order that corresponds to nothing in particular (which didn't concern me - as I assumed row_id was being used).\n\nAre the solution and submission dataframes re-sorted at some point so they are assured to have a matching order?\n\n(I'm kind of inclined to randomize the order of my rows on a submission - and see what happens...)\n\n-Rich",
    "2768781": "Just did 2 test submissions.\n\n1st one had random numbers for row_id - and it got rejected for having an invalid format.\n\n2nd test took a previous .61 LB notebook - and just resorted the predictions (maintaining column order, but reversing the rows).  The score dropped to .47 on the LB.\n\nSo - there's presumably some kind of filter before the shared scoring code that deals with row_id.",
    "2769715": "Thanks, this is reassuring"
  },
  "source": "meta"
}