{
  "id": 523247,
  "title": "[NEW DATASET] Tiny Debug Dataset for RSNA 2024 Lumbar Spine Degenerative Classification",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523247",
  "author_name": "",
  "post_date": "2024-07-31T06:03:49.009193600Z",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hey Kagglers,</p>\n<p>I have created a tiny debug dataset for this competition which can be used as drop in replacement of original competition dataset.</p>\n<p><strong>Dataset Link</strong> : <a href=\"https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset\" target=\"_blank\">RSNA_LSDC_2024_submission_debug_dataset</a><br>\n<strong>Notebook Link</strong>: <a href=\"https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook\" target=\"_blank\">RSNA_LSDC_2024_submission_debug</a></p>\n<p><strong>Motivation:</strong></p>\n<ul>\n<li>I was struggling to validate my code while training on this huge dataset and it felt near impossible to debug when my notebook generated <code>submission.csv</code> successfully but scoring failed with either <code>Submission Scoring Error</code> or <code>Notebook Threw Exception</code>.</li>\n<li>I wanted to validate my submission code which loops through study_ids in test_images folder which is not possible with current test_images folder in competition dataset with single study_id instance.</li>\n<li>I had asked the organizers to have at least 2 folders in test_images for loop code validation [here](<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357#2917606\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357#2917606</a><br>\n) but organizers haven't responded since 20-days so I decided to do something about it myself.</li>\n<li>I also wanted to test my submission code against edge cases like less than or more than 3-folders etc which is not readily possible with current test_images folder.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/519612#2918053\" target=\"_blank\">here</a><br>\n<a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">Sergey Saharovskiy</a> suggested to run the inference pipeline on train folder which is good idea, but it would take long time to run inference on whole train folder which will again eat away valuable GPU resurces which are limited per week.</li>\n</ul>\n<p><strong>Solution:</strong></p>\n<ul>\n<li><p>I created a debug folder which basically is copy of original competition dataset but with very few data.</p></li>\n<li><p>The dataset includes 7-study_ids for train and test each with variations like 2,3,4-series per study_id which will validate your code for these variations.</p></li>\n<li><p>Happy coding !!!</p></li>\n</ul>\n<p>Example Usage Code for Drop In Replacement and Debugging:</p>\n<pre><code>DEBUG = \n DEBUG == :\n    base_dir = \n:\n    base_dir = \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10192946%2F5eeb4498cc87f034e150c2687b28a011%2FRSNA_LSDC_2024_submission_debug_dataset_%20RSNA%202024%20Lumbar%20Spine%20Degenerative%20Classification.JPG?generation=1722405792182836&amp;alt=media\" alt=\"dataset comparison\"></p>\n<p><strong>EDIT :</strong><br>\nupdated code to remove any series_ids in description csv which do not have corresponding image folder in train_images / test_images.</p>\n<p>Thanks   <a href=\"https://www.kaggle.com/deankelley\" target=\"_blank\">@deankelley</a> and <a href=\"https://www.kaggle.com/kaggleutata\" target=\"_blank\">@kaggleutata</a> for highlighting this issue with the dataset.<br>\nDue to your interest and enthusiasm, this debug dataset just got better for the new people entering in the competition.</p>",
  "messages": [
    {
      "id": "2941602",
      "postDate": "07/31/2024 06:03:49",
      "content": "<p>Hey Kagglers,</p>\n<p>I have created a tiny debug dataset for this competition which can be used as drop in replacement of original competition dataset.</p>\n<p><strong>Dataset Link</strong> : <a href=\"https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset\" target=\"_blank\">RSNA_LSDC_2024_submission_debug_dataset</a><br>\n<strong>Notebook Link</strong>: <a href=\"https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook\" target=\"_blank\">RSNA_LSDC_2024_submission_debug</a></p>\n<p><strong>Motivation:</strong></p>\n<ul>\n<li>I was struggling to validate my code while training on this huge dataset and it felt near impossible to debug when my notebook generated <code>submission.csv</code> successfully but scoring failed with either <code>Submission Scoring Error</code> or <code>Notebook Threw Exception</code>.</li>\n<li>I wanted to validate my submission code which loops through study_ids in test_images folder which is not possible with current test_images folder in competition dataset with single study_id instance.</li>\n<li>I had asked the organizers to have at least 2 folders in test_images for loop code validation [here](<a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357#2917606\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357#2917606</a><br>\n) but organizers haven't responded since 20-days so I decided to do something about it myself.</li>\n<li>I also wanted to test my submission code against edge cases like less than or more than 3-folders etc which is not readily possible with current test_images folder.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/519612#2918053\" target=\"_blank\">here</a><br>\n<a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">Sergey Saharovskiy</a> suggested to run the inference pipeline on train folder which is good idea, but it would take long time to run inference on whole train folder which will again eat away valuable GPU resurces which are limited per week.</li>\n</ul>\n<p><strong>Solution:</strong></p>\n<ul>\n<li><p>I created a debug folder which basically is copy of original competition dataset but with very few data.</p></li>\n<li><p>The dataset includes 7-study_ids for train and test each with variations like 2,3,4-series per study_id which will validate your code for these variations.</p></li>\n<li><p>Happy coding !!!</p></li>\n</ul>\n<p>Example Usage Code for Drop In Replacement and Debugging:</p>\n<pre><code>DEBUG = \n DEBUG == :\n    base_dir = \n:\n    base_dir = \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10192946%2F5eeb4498cc87f034e150c2687b28a011%2FRSNA_LSDC_2024_submission_debug_dataset_%20RSNA%202024%20Lumbar%20Spine%20Degenerative%20Classification.JPG?generation=1722405792182836&amp;alt=media\" alt=\"dataset comparison\"></p>\n<p><strong>EDIT :</strong><br>\nupdated code to remove any series_ids in description csv which do not have corresponding image folder in train_images / test_images.</p>\n<p>Thanks   <a href=\"https://www.kaggle.com/deankelley\" target=\"_blank\">@deankelley</a> and <a href=\"https://www.kaggle.com/kaggleutata\" target=\"_blank\">@kaggleutata</a> for highlighting this issue with the dataset.<br>\nDue to your interest and enthusiasm, this debug dataset just got better for the new people entering in the competition.</p>",
      "rawMarkdown": "Hey Kagglers,\n\nI have created a tiny debug dataset for this competition which can be used as drop in replacement of original competition dataset.\n\n**Dataset Link** : [RSNA_LSDC_2024_submission_debug_dataset](https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset)\n**Notebook Link**: [RSNA_LSDC_2024_submission_debug](https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook)\n\n\n**Motivation:**\n-  I was struggling to validate my code while training on this huge dataset and it felt near impossible to debug when my notebook generated `submission.csv` successfully but scoring failed with either `Submission Scoring Error` or `Notebook Threw Exception`.\n- I wanted to validate my submission code which loops through study_ids in test_images folder which is not possible with current test_images folder in competition dataset with single study_id instance.\n- I had asked the organizers to have at least 2 folders in test_images for loop code validation [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357#2917606\n) but organizers haven't responded since 20-days so I decided to do something about it myself.\n- I also wanted to test my submission code against edge cases like less than or more than 3-folders etc which is not readily possible with current test_images folder.\n- [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/519612#2918053)\n[Sergey Saharovskiy](https://www.kaggle.com/sergiosaharovskiy) suggested to run the inference pipeline on train folder which is good idea, but it would take long time to run inference on whole train folder which will again eat away valuable GPU resurces which are limited per week.\n\n\n**Solution:**\n- I created a debug folder which basically is copy of original competition dataset but with very few data.\n- The dataset includes 7-study_ids for train and test each with variations like 2,3,4-series per study_id which will validate your code for these variations.\n\n- Happy coding !!!\n\n\nExample Usage Code for Drop In Replacement and Debugging:\n```python\nDEBUG = True\nif DEBUG == True:\n    base_dir = '/kaggle/input/rsna-lsdc-2024-submission-debug-dataset/debug'\nelse:\n    base_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification'\n```\n![dataset comparison](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10192946%2F5eeb4498cc87f034e150c2687b28a011%2FRSNA_LSDC_2024_submission_debug_dataset_%20RSNA%202024%20Lumbar%20Spine%20Degenerative%20Classification.JPG?generation=1722405792182836&alt=media)\n\n\n**EDIT :**\nupdated code to remove any series_ids in description csv which do not have corresponding image folder in train_images / test_images.\n\nThanks   @deankelley and @kaggleutata for highlighting this issue with the dataset.\nDue to your interest and enthusiasm, this debug dataset just got better for the new people entering in the competition.",
      "votes": null
    },
    {
      "id": "2942117",
      "postDate": "07/31/2024 15:38:15",
      "content": "<p>I would recommend that anyone using this can use the flag:</p>\n<pre><code>ss_df = pd.read_csv()\nDEBUG =   (ss_df) &lt;=  \n</code></pre>\n<p>This means you can test out the notebook with the debug data during \"Run All and Save\" just submit the inference notebook as is and it will automatically switch to the hidden test data.</p>",
      "rawMarkdown": "I would recommend that anyone using this can use the flag:\n```python\nss_df = pd.read_csv(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/sample_submission.csv\")\nDEBUG = True and len(ss_df) <= 25 # or just DEBUG = len(ss_df) <= 25\n```\nThis means you can test out the notebook with the debug data during \"Run All and Save\" just submit the inference notebook as is and it will automatically switch to the hidden test data.",
      "votes": null
    },
    {
      "id": "2942234",
      "postDate": "07/31/2024 16:37:45",
      "content": "<p>Nice suggestion !<br>\nI have also verified that sample_submission.csv file is indeed changed while running on hidden set by using assertions on valid submission code.</p>",
      "rawMarkdown": "Nice suggestion !\nI have also verified that sample_submission.csv file is indeed changed while running on hidden set by using assertions on valid submission code.",
      "votes": null
    },
    {
      "id": "2946183",
      "postDate": "08/04/2024 04:21:05",
      "content": "<p>Rohit. The information you provided helped me understand the reason for my submission failure. I was feeling lost and lost, so it really helped me.</p>",
      "rawMarkdown": "Rohit. The information you provided helped me understand the reason for my submission failure. I was feeling lost and lost, so it really helped me.",
      "votes": null
    },
    {
      "id": "2946277",
      "postDate": "08/04/2024 07:14:35",
      "content": "<p><a href=\"https://www.kaggle.com/jinstat\" target=\"_blank\">@jinstat</a> happy to hear that i could contribute to your success.</p>\n<p>Please do upvote the dataset and notebook so they are more visible to other people with similar problems.</p>",
      "rawMarkdown": "jinstat happy to hear that i could contribute to your success.\n\nPlease do upvote the dataset and notebook so they are more visible to other people with similar problems.",
      "votes": null
    },
    {
      "id": "2967252",
      "postDate": "08/22/2024 17:18:35",
      "content": "<p>Thanks. You also answered my question. I was wondering whether each study will at least have 3 series that cover all three views in test dataset.</p>",
      "rawMarkdown": "Thanks. You also answered my question. I was wondering whether each study will at least have 3 series that cover all three views in test dataset.",
      "votes": null
    },
    {
      "id": "2969890",
      "postDate": "08/25/2024 14:10:14",
      "content": "<p><a href=\"https://www.kaggle.com/sychen52\" target=\"_blank\">@sychen52</a> , happy to help !<br>\nPlease do upvote the dataset and notebook to increase visibility.</p>",
      "rawMarkdown": "sychen52 , happy to help !\nPlease do upvote the dataset and notebook to increase visibility.",
      "votes": null
    },
    {
      "id": "3004554",
      "postDate": "10/02/2024 01:36:26",
      "content": "<p>thank you for your greate debug dataset!<br>\nI want to ask about a missing sagittal data(study_id: 3429409220, series_id:3559571772) in this debug dataset.<br>\nWhy is this sagittaldata missing ??</p>\n<p>3429409220/3559571772</p>",
      "rawMarkdown": "thank you for your greate debug dataset!\nI want to ask about a missing sagittal data(study_id: 3429409220, series_id:3559571772) in this debug dataset.\nWhy is this sagittaldata missing ??\n\n3429409220/3559571772",
      "votes": null
    },
    {
      "id": "3004692",
      "postDate": "10/02/2024 05:56:01",
      "content": "<p>Thanks for the observation <a href=\"https://www.kaggle.com/kaggleutata\" target=\"_blank\">@kaggleutata</a> ,</p>\n<p>This is edge case bug in my dataset.<br>\nthe hidden test set should not contain any series_id without corresponding image folder present.</p>\n<p>This issue was previously highlighted here  <a href=\"https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset/discussion/530786\" target=\"_blank\">Question about dataset and issue I'm having</a>.</p>\n<p>but I was a little busy to look at the code.</p>\n<p>I have updated the dataset for this case and removed any series_id rows that do not have matching image folder in the debug dataset.<br>\nhowever you still need to predict all 25-labels regardless of number of series_id folders in each study_id.</p>",
      "rawMarkdown": "Thanks for the observation @kaggleutata ,\n\nThis is edge case bug in my dataset.\nthe hidden test set should not contain any series_id without corresponding image folder present.\n\nThis issue was previously highlighted here  [Question about dataset and issue I'm having](https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset/discussion/530786).\n\nbut I was a little busy to look at the code.\n\nI have updated the dataset for this case and removed any series_id rows that do not have matching image folder in the debug dataset.\nhowever you still need to predict all 25-labels regardless of number of series_id folders in each study_id.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2942117,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "07/31/2024 15:38:15",
      "content": "<p>I would recommend that anyone using this can use the flag:</p>\n<pre><code>ss_df = pd.read_csv()\nDEBUG =   (ss_df) &lt;=  \n</code></pre>\n<p>This means you can test out the notebook with the debug data during \"Run All and Save\" just submit the inference notebook as is and it will automatically switch to the hidden test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2942234,
          "author_name": "rohitchaudhari25",
          "author_url": "",
          "post_date": "07/31/2024 16:37:45",
          "content": "<p>Nice suggestion !<br>\nI have also verified that sample_submission.csv file is indeed changed while running on hidden set by using assertions on valid submission code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2946183,
      "author_name": "jinstat",
      "author_url": "",
      "post_date": "08/04/2024 04:21:05",
      "content": "<p>Rohit. The information you provided helped me understand the reason for my submission failure. I was feeling lost and lost, so it really helped me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2946277,
          "author_name": "rohitchaudhari25",
          "author_url": "",
          "post_date": "08/04/2024 07:14:35",
          "content": "<p><a href=\"https://www.kaggle.com/jinstat\" target=\"_blank\">@jinstat</a> happy to hear that i could contribute to your success.</p>\n<p>Please do upvote the dataset and notebook so they are more visible to other people with similar problems.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2967252,
      "author_name": "sychen52",
      "author_url": "",
      "post_date": "08/22/2024 17:18:35",
      "content": "<p>Thanks. You also answered my question. I was wondering whether each study will at least have 3 series that cover all three views in test dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2969890,
          "author_name": "rohitchaudhari25",
          "author_url": "",
          "post_date": "08/25/2024 14:10:14",
          "content": "<p><a href=\"https://www.kaggle.com/sychen52\" target=\"_blank\">@sychen52</a> , happy to help !<br>\nPlease do upvote the dataset and notebook to increase visibility.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3004554,
      "author_name": "kaggleutata",
      "author_url": "",
      "post_date": "10/02/2024 01:36:26",
      "content": "<p>thank you for your greate debug dataset!<br>\nI want to ask about a missing sagittal data(study_id: 3429409220, series_id:3559571772) in this debug dataset.<br>\nWhy is this sagittaldata missing ??</p>\n<p>3429409220/3559571772</p>",
      "votes": null,
      "replies": [
        {
          "id": 3004692,
          "author_name": "rohitchaudhari25",
          "author_url": "",
          "post_date": "10/02/2024 05:56:01",
          "content": "<p>Thanks for the observation <a href=\"https://www.kaggle.com/kaggleutata\" target=\"_blank\">@kaggleutata</a> ,</p>\n<p>This is edge case bug in my dataset.<br>\nthe hidden test set should not contain any series_id without corresponding image folder present.</p>\n<p>This issue was previously highlighted here  <a href=\"https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset/discussion/530786\" target=\"_blank\">Question about dataset and issue I'm having</a>.</p>\n<p>but I was a little busy to look at the code.</p>\n<p>I have updated the dataset for this case and removed any series_id rows that do not have matching image folder in the debug dataset.<br>\nhowever you still need to predict all 25-labels regardless of number of series_id folders in each study_id.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2941602": "Hey Kagglers,\n\nI have created a tiny debug dataset for this competition which can be used as drop in replacement of original competition dataset.\n\n**Dataset Link** : [RSNA_LSDC_2024_submission_debug_dataset](https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset)\n**Notebook Link**: [RSNA_LSDC_2024_submission_debug](https://www.kaggle.com/code/rohitchaudhari25/rsna-lsdc-2024-submission-debug/notebook)\n\n\n**Motivation:**\n-  I was struggling to validate my code while training on this huge dataset and it felt near impossible to debug when my notebook generated `submission.csv` successfully but scoring failed with either `Submission Scoring Error` or `Notebook Threw Exception`.\n- I wanted to validate my submission code which loops through study_ids in test_images folder which is not possible with current test_images folder in competition dataset with single study_id instance.\n- I had asked the organizers to have at least 2 folders in test_images for loop code validation [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/516357#2917606\n) but organizers haven't responded since 20-days so I decided to do something about it myself.\n- I also wanted to test my submission code against edge cases like less than or more than 3-folders etc which is not readily possible with current test_images folder.\n- [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/519612#2918053)\n[Sergey Saharovskiy](https://www.kaggle.com/sergiosaharovskiy) suggested to run the inference pipeline on train folder which is good idea, but it would take long time to run inference on whole train folder which will again eat away valuable GPU resurces which are limited per week.\n\n\n**Solution:**\n- I created a debug folder which basically is copy of original competition dataset but with very few data.\n- The dataset includes 7-study_ids for train and test each with variations like 2,3,4-series per study_id which will validate your code for these variations.\n\n- Happy coding !!!\n\n\nExample Usage Code for Drop In Replacement and Debugging:\n```python\nDEBUG = True\nif DEBUG == True:\n    base_dir = '/kaggle/input/rsna-lsdc-2024-submission-debug-dataset/debug'\nelse:\n    base_dir = '/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification'\n```\n![dataset comparison](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10192946%2F5eeb4498cc87f034e150c2687b28a011%2FRSNA_LSDC_2024_submission_debug_dataset_%20RSNA%202024%20Lumbar%20Spine%20Degenerative%20Classification.JPG?generation=1722405792182836&alt=media)\n\n\n**EDIT :**\nupdated code to remove any series_ids in description csv which do not have corresponding image folder in train_images / test_images.\n\nThanks   @deankelley and @kaggleutata for highlighting this issue with the dataset.\nDue to your interest and enthusiasm, this debug dataset just got better for the new people entering in the competition.",
    "2942117": "I would recommend that anyone using this can use the flag:\n```python\nss_df = pd.read_csv(\"/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/sample_submission.csv\")\nDEBUG = True and len(ss_df) <= 25 # or just DEBUG = len(ss_df) <= 25\n```\nThis means you can test out the notebook with the debug data during \"Run All and Save\" just submit the inference notebook as is and it will automatically switch to the hidden test data.",
    "2942234": "Nice suggestion !\nI have also verified that sample_submission.csv file is indeed changed while running on hidden set by using assertions on valid submission code.",
    "2946183": "Rohit. The information you provided helped me understand the reason for my submission failure. I was feeling lost and lost, so it really helped me.",
    "2946277": "jinstat happy to hear that i could contribute to your success.\n\nPlease do upvote the dataset and notebook so they are more visible to other people with similar problems.",
    "2967252": "Thanks. You also answered my question. I was wondering whether each study will at least have 3 series that cover all three views in test dataset.",
    "2969890": "sychen52 , happy to help !\nPlease do upvote the dataset and notebook to increase visibility.",
    "3004554": "thank you for your greate debug dataset!\nI want to ask about a missing sagittal data(study_id: 3429409220, series_id:3559571772) in this debug dataset.\nWhy is this sagittaldata missing ??\n\n3429409220/3559571772",
    "3004692": "Thanks for the observation @kaggleutata ,\n\nThis is edge case bug in my dataset.\nthe hidden test set should not contain any series_id without corresponding image folder present.\n\nThis issue was previously highlighted here  [Question about dataset and issue I'm having](https://www.kaggle.com/datasets/rohitchaudhari25/rsna-lsdc-2024-submission-debug-dataset/discussion/530786).\n\nbut I was a little busy to look at the code.\n\nI have updated the dataset for this case and removed any series_id rows that do not have matching image folder in the debug dataset.\nhowever you still need to predict all 25-labels regardless of number of series_id folders in each study_id."
  },
  "source": "meta"
}