{
  "id": 508319,
  "title": "Inqury about the metric",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508319",
  "author_name": "",
  "post_date": "2024-05-29T06:15:36.965955500Z",
  "votes": 22,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thank you for organizing such a wonderful competition!</p>\n<p>I have a question regarding an excerpt from the metrics code (version 8 of <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\">https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549</a> ):</p>\n<pre><code>any_severe_spinal_labels = pd.Series(solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_weights = pd.Series(solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_predictions = pd.Series(solution.loc[solution[] == ].groupby()[].()) \nany_severe_spinal_loss = sklearn.metrics.log_loss(\n            y_true=any_severe_spinal_labels,\n            y_pred=any_severe_spinal_predictions,\n            sample_weight=any_severe_spinal_weights\n        )\n</code></pre>\n<p>In the code, both any_severe_spinal_predictions and any_severe_spinal_labels are derived from the same dataset (solution) and therefore have the same values. However, shouldn't any_severe_spinal_predictions be calculated from the submission dataset instead? </p>\n<p>For example:</p>\n<pre><code>any_severe_spinal_predictions = pd.Series(submission.loc[submission[] == ].groupby()[].())\n</code></pre>\n<p>Could you confirm if this is an oversight or if I am misunderstanding the code?</p>",
  "messages": [
    {
      "id": "2842538",
      "postDate": "05/29/2024 06:15:36",
      "content": "<p>Thank you for organizing such a wonderful competition!</p>\n<p>I have a question regarding an excerpt from the metrics code (version 8 of <a href=\"https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549\" target=\"_blank\">https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549</a> ):</p>\n<pre><code>any_severe_spinal_labels = pd.Series(solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_weights = pd.Series(solution.loc[solution[] == ].groupby()[].())\nany_severe_spinal_predictions = pd.Series(solution.loc[solution[] == ].groupby()[].()) \nany_severe_spinal_loss = sklearn.metrics.log_loss(\n            y_true=any_severe_spinal_labels,\n            y_pred=any_severe_spinal_predictions,\n            sample_weight=any_severe_spinal_weights\n        )\n</code></pre>\n<p>In the code, both any_severe_spinal_predictions and any_severe_spinal_labels are derived from the same dataset (solution) and therefore have the same values. However, shouldn't any_severe_spinal_predictions be calculated from the submission dataset instead? </p>\n<p>For example:</p>\n<pre><code>any_severe_spinal_predictions = pd.Series(submission.loc[submission[] == ].groupby()[].())\n</code></pre>\n<p>Could you confirm if this is an oversight or if I am misunderstanding the code?</p>",
      "rawMarkdown": "Thank you for organizing such a wonderful competition!\n\nI have a question regarding an excerpt from the metrics code (version 8 of https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549 ):\n\n```python\nany_severe_spinal_labels = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max())\nany_severe_spinal_weights = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\nany_severe_spinal_predictions = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max()) # same as the any_severe_spinal_labels\nany_severe_spinal_loss = sklearn.metrics.log_loss(\n            y_true=any_severe_spinal_labels,\n            y_pred=any_severe_spinal_predictions,\n            sample_weight=any_severe_spinal_weights\n        )\n```\n\nIn the code, both any_severe_spinal_predictions and any_severe_spinal_labels are derived from the same dataset (solution) and therefore have the same values. However, shouldn't any_severe_spinal_predictions be calculated from the submission dataset instead? \n\nFor example:\n\n```python\nany_severe_spinal_predictions = pd.Series(submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n```\n\nCould you confirm if this is an oversight or if I am misunderstanding the code?",
      "votes": null
    },
    {
      "id": "2842549",
      "postDate": "05/29/2024 06:29:40",
      "content": "<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>I would appreciate it if you could check it.</p>",
      "rawMarkdown": "felipekitamura @sohier \n\nI would appreciate it if you could check it.",
      "votes": null
    },
    {
      "id": "2843446",
      "postDate": "05/29/2024 15:21:35",
      "content": "<p>I also have the same question.</p>",
      "rawMarkdown": "I also have the same question.",
      "votes": null
    },
    {
      "id": "2844146",
      "postDate": "05/30/2024 01:29:55",
      "content": "<p>Thank you for catching this, I should not have let that bug slip into production. Please see <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522</a></p>",
      "rawMarkdown": "Thank you for catching this, I should not have let that bug slip into production. Please see https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2842549,
      "author_name": "abebe9849",
      "author_url": "",
      "post_date": "05/29/2024 06:29:40",
      "content": "<p><a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>I would appreciate it if you could check it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2843446,
      "author_name": "",
      "author_url": "",
      "post_date": "05/29/2024 15:21:35",
      "content": "<p>I also have the same question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844146,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "05/30/2024 01:29:55",
      "content": "<p>Thank you for catching this, I should not have let that bug slip into production. Please see <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2842538": "Thank you for organizing such a wonderful competition!\n\nI have a question regarding an excerpt from the metrics code (version 8 of https://www.kaggle.com/code/metric/rsna-lumbar-metric-71549 ):\n\n```python\nany_severe_spinal_labels = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max())\nany_severe_spinal_weights = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['sample_weight'].max())\nany_severe_spinal_predictions = pd.Series(solution.loc[solution['condition'] == 'spinal'].groupby('study_id')['severe'].max()) # same as the any_severe_spinal_labels\nany_severe_spinal_loss = sklearn.metrics.log_loss(\n            y_true=any_severe_spinal_labels,\n            y_pred=any_severe_spinal_predictions,\n            sample_weight=any_severe_spinal_weights\n        )\n```\n\nIn the code, both any_severe_spinal_predictions and any_severe_spinal_labels are derived from the same dataset (solution) and therefore have the same values. However, shouldn't any_severe_spinal_predictions be calculated from the submission dataset instead? \n\nFor example:\n\n```python\nany_severe_spinal_predictions = pd.Series(submission.loc[submission['condition'] == 'spinal'].groupby('study_id')['severe'].max())\n```\n\nCould you confirm if this is an oversight or if I am misunderstanding the code?",
    "2842549": "felipekitamura @sohier \n\nI would appreciate it if you could check it.",
    "2843446": "I also have the same question.",
    "2844146": "Thank you for catching this, I should not have let that bug slip into production. Please see https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/508522"
  },
  "source": "meta"
}