{
  "id": 567202,
  "title": "BYU_BioPhysics_91249 - Flaw in the metric?",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/567202",
  "author_name": "",
  "post_date": "2025-03-09T02:50:53.452516500Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The score function in metric.py checks not just the row count but also the exact <code>tomo_id</code> values row-by-row:</p>\n<pre><code>filename_equiv_array = solution[].eq(submission[])\n np.(filename_equiv_array) != (solution[]):\n     ValueError()\n</code></pre>\n<p>It requires <code>train_preds['tomo_id']</code> to match <code>train_solution['tomo_id']</code> exactly in order. Padding with -1 rows for missing predictions helps with row count, but if the tomo_ids don’t align, the check fails.</p>\n<p>Is there a flaw in the metric? It assumes the submission predicts the exact same number of motors per tomogram as the solution, in the same order, which is overly restrictive. A proper metric should handle varying numbers of predictions (e.g., fewer or more motors) by matching predicted motors to true motors based on proximity, not strict tomo_id sequence.</p>\n<p>The competition metric can be found here: <br>\n<a href=\"https://www.kaggle.com/code/metric/byu-biophysics-91249/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/byu-biophysics-91249/notebook</a></p>",
  "messages": [
    {
      "id": "3144889",
      "postDate": "03/09/2025 02:50:53",
      "content": "<p>The score function in metric.py checks not just the row count but also the exact <code>tomo_id</code> values row-by-row:</p>\n<pre><code>filename_equiv_array = solution[].eq(submission[])\n np.(filename_equiv_array) != (solution[]):\n     ValueError()\n</code></pre>\n<p>It requires <code>train_preds['tomo_id']</code> to match <code>train_solution['tomo_id']</code> exactly in order. Padding with -1 rows for missing predictions helps with row count, but if the tomo_ids don’t align, the check fails.</p>\n<p>Is there a flaw in the metric? It assumes the submission predicts the exact same number of motors per tomogram as the solution, in the same order, which is overly restrictive. A proper metric should handle varying numbers of predictions (e.g., fewer or more motors) by matching predicted motors to true motors based on proximity, not strict tomo_id sequence.</p>\n<p>The competition metric can be found here: <br>\n<a href=\"https://www.kaggle.com/code/metric/byu-biophysics-91249/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/byu-biophysics-91249/notebook</a></p>",
      "rawMarkdown": "The score function in metric.py checks not just the row count but also the exact `tomo_id` values row-by-row:\n\n```python\nfilename_equiv_array = solution['tomo_id'].eq(submission['tomo_id'])\nif np.sum(filename_equiv_array) != len(solution['tomo_id']):\n    raise ValueError('Submitted tomo_id values do not match the sample_submission file')\n```\n\nIt requires `train_preds['tomo_id']` to match `train_solution['tomo_id']` exactly in order. Padding with -1 rows for missing predictions helps with row count, but if the tomo_ids don’t align, the check fails.\n\nIs there a flaw in the metric? It assumes the submission predicts the exact same number of motors per tomogram as the solution, in the same order, which is overly restrictive. A proper metric should handle varying numbers of predictions (e.g., fewer or more motors) by matching predicted motors to true motors based on proximity, not strict tomo_id sequence.\n\nThe competition metric can be found here: \nhttps://www.kaggle.com/code/metric/byu-biophysics-91249/notebook",
      "votes": null
    },
    {
      "id": "3144914",
      "postDate": "03/09/2025 04:14:16",
      "content": "<p>That is not a flaw because both have been sorted by tomo_id before checking </p>\n<p><code>\nsolution = solution.sort_values('tomo_id').reset_index(drop=True)\nsubmission = submission.sort_values('tomo_id').reset_index(drop=True)\n</code></p>",
      "rawMarkdown": "That is not a flaw because both have been sorted by tomo_id before checking \n\n`\nsolution = solution.sort_values('tomo_id').reset_index(drop=True)\nsubmission = submission.sort_values('tomo_id').reset_index(drop=True)\n`",
      "votes": null
    },
    {
      "id": "3144981",
      "postDate": "03/09/2025 06:01:58",
      "content": "<p><strong>EDIT</strong>: I'm guessing since there is only one motor per tomogram in the test dataset, the competition metric is given for transparency, and is used solely to evaluate the test set. We will have to write our own FB-score metric for evaluating our performance on training and validation sets, where multiple motors (e.g. up to 10 motors) exist per tomogram.</p>\n<p>Ah okay. that makes sense. Thank you for the clarification 🙏</p>\n<p>The number of motors in the submission dataframe still has to be the same as the number of motors in the solution dataframe. </p>\n<p>I have tried examples where there are more motors in the solution dataframe than predicted, like the example below, where <code>tomo_id</code> 1 has two motors in the solution, and only one predicted in the submission, and it gives me the same error: <code>ValueError: Submitted tomo_id values do not match the sample_submission file</code>. </p>\n<p>Here is the code:</p>\n<pre><code> metric  score\n\nsolution = pd.DataFrame({\n     : [, , , , ],\n     : [-, , , , ],\n     : [-, , , , ],\n     : [-, , , , ],\n     : [, , , , ],\n     : [, , , , ]\n })\n\nsubmission = pd.DataFrame({\n     : [, , , ],\n     : [, , , -],\n     : [, , , -],\n     : [, , , -]\n })\nscore(solution, submission, , )\n</code></pre>",
      "rawMarkdown": "**EDIT**: I'm guessing since there is only one motor per tomogram in the test dataset, the competition metric is given for transparency, and is used solely to evaluate the test set. We will have to write our own FB-score metric for evaluating our performance on training and validation sets, where multiple motors (e.g. up to 10 motors) exist per tomogram.\n\nAh okay. that makes sense. Thank you for the clarification 🙏\n\nThe number of motors in the submission dataframe still has to be the same as the number of motors in the solution dataframe. \n\nI have tried examples where there are more motors in the solution dataframe than predicted, like the example below, where `tomo_id` 1 has two motors in the solution, and only one predicted in the submission, and it gives me the same error: `ValueError: Submitted tomo_id values do not match the sample_submission file`. \n\nHere is the code:\n\n```python\nfrom metric import score\n\nsolution = pd.DataFrame({\n     'tomo_id': [0, 1, 1, 2, 3],\n     'Motor axis 0': [-1, 250, 300, 100, 200],\n     'Motor axis 1': [-1, 250, 300, 100, 200],\n     'Motor axis 2': [-1, 250, 300, 100, 200],\n     'Voxel spacing': [10, 10, 10, 10, 10],\n     'Has motor': [0, 1, 1, 1, 1]\n })\n\nsubmission = pd.DataFrame({\n     'tomo_id': [0, 1, 2, 3],\n     'Motor axis 0': [100, 251, 600, -1],\n     'Motor axis 1': [100, 251, 600, -1],\n     'Motor axis 2': [100, 251, 600, -1]\n })\nscore(solution, submission, 1000, 2)\n```",
      "votes": null
    },
    {
      "id": "3145316",
      "postDate": "03/09/2025 17:34:44",
      "content": "<p>The competition metric works only with a maximum of one motor per tomo_id.</p>\n<p>That is also mentioned in the competition overview as well</p>",
      "rawMarkdown": "The competition metric works only with a maximum of one motor per tomo_id.\n\nThat is also mentioned in the competition overview as well",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3144914,
      "author_name": "alabibojesomo",
      "author_url": "",
      "post_date": "03/09/2025 04:14:16",
      "content": "<p>That is not a flaw because both have been sorted by tomo_id before checking </p>\n<p><code>\nsolution = solution.sort_values('tomo_id').reset_index(drop=True)\nsubmission = submission.sort_values('tomo_id').reset_index(drop=True)\n</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 3144981,
          "author_name": "rocktopus",
          "author_url": "",
          "post_date": "03/09/2025 06:01:58",
          "content": "<p><strong>EDIT</strong>: I'm guessing since there is only one motor per tomogram in the test dataset, the competition metric is given for transparency, and is used solely to evaluate the test set. We will have to write our own FB-score metric for evaluating our performance on training and validation sets, where multiple motors (e.g. up to 10 motors) exist per tomogram.</p>\n<p>Ah okay. that makes sense. Thank you for the clarification 🙏</p>\n<p>The number of motors in the submission dataframe still has to be the same as the number of motors in the solution dataframe. </p>\n<p>I have tried examples where there are more motors in the solution dataframe than predicted, like the example below, where <code>tomo_id</code> 1 has two motors in the solution, and only one predicted in the submission, and it gives me the same error: <code>ValueError: Submitted tomo_id values do not match the sample_submission file</code>. </p>\n<p>Here is the code:</p>\n<pre><code> metric  score\n\nsolution = pd.DataFrame({\n     : [, , , , ],\n     : [-, , , , ],\n     : [-, , , , ],\n     : [-, , , , ],\n     : [, , , , ],\n     : [, , , , ]\n })\n\nsubmission = pd.DataFrame({\n     : [, , , ],\n     : [, , , -],\n     : [, , , -],\n     : [, , , -]\n })\nscore(solution, submission, , )\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3145316,
              "author_name": "alabibojesomo",
              "author_url": "",
              "post_date": "03/09/2025 17:34:44",
              "content": "<p>The competition metric works only with a maximum of one motor per tomo_id.</p>\n<p>That is also mentioned in the competition overview as well</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3144889": "The score function in metric.py checks not just the row count but also the exact `tomo_id` values row-by-row:\n\n```python\nfilename_equiv_array = solution['tomo_id'].eq(submission['tomo_id'])\nif np.sum(filename_equiv_array) != len(solution['tomo_id']):\n    raise ValueError('Submitted tomo_id values do not match the sample_submission file')\n```\n\nIt requires `train_preds['tomo_id']` to match `train_solution['tomo_id']` exactly in order. Padding with -1 rows for missing predictions helps with row count, but if the tomo_ids don’t align, the check fails.\n\nIs there a flaw in the metric? It assumes the submission predicts the exact same number of motors per tomogram as the solution, in the same order, which is overly restrictive. A proper metric should handle varying numbers of predictions (e.g., fewer or more motors) by matching predicted motors to true motors based on proximity, not strict tomo_id sequence.\n\nThe competition metric can be found here: \nhttps://www.kaggle.com/code/metric/byu-biophysics-91249/notebook",
    "3144914": "That is not a flaw because both have been sorted by tomo_id before checking \n\n`\nsolution = solution.sort_values('tomo_id').reset_index(drop=True)\nsubmission = submission.sort_values('tomo_id').reset_index(drop=True)\n`",
    "3144981": "**EDIT**: I'm guessing since there is only one motor per tomogram in the test dataset, the competition metric is given for transparency, and is used solely to evaluate the test set. We will have to write our own FB-score metric for evaluating our performance on training and validation sets, where multiple motors (e.g. up to 10 motors) exist per tomogram.\n\nAh okay. that makes sense. Thank you for the clarification 🙏\n\nThe number of motors in the submission dataframe still has to be the same as the number of motors in the solution dataframe. \n\nI have tried examples where there are more motors in the solution dataframe than predicted, like the example below, where `tomo_id` 1 has two motors in the solution, and only one predicted in the submission, and it gives me the same error: `ValueError: Submitted tomo_id values do not match the sample_submission file`. \n\nHere is the code:\n\n```python\nfrom metric import score\n\nsolution = pd.DataFrame({\n     'tomo_id': [0, 1, 1, 2, 3],\n     'Motor axis 0': [-1, 250, 300, 100, 200],\n     'Motor axis 1': [-1, 250, 300, 100, 200],\n     'Motor axis 2': [-1, 250, 300, 100, 200],\n     'Voxel spacing': [10, 10, 10, 10, 10],\n     'Has motor': [0, 1, 1, 1, 1]\n })\n\nsubmission = pd.DataFrame({\n     'tomo_id': [0, 1, 2, 3],\n     'Motor axis 0': [100, 251, 600, -1],\n     'Motor axis 1': [100, 251, 600, -1],\n     'Motor axis 2': [100, 251, 600, -1]\n })\nscore(solution, submission, 1000, 2)\n```",
    "3145316": "The competition metric works only with a maximum of one motor per tomo_id.\n\nThat is also mentioned in the competition overview as well"
  },
  "source": "meta"
}