{
  "id": 529728,
  "title": "Submission exceeding time limit ",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/529728",
  "author_name": "",
  "post_date": "2024-08-22T15:00:21.358196500Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi to fellow kagglers,<br>\nI am stuck at the submission stage and hitting a wall, would greatly appreciate any pointers to resolution.</p>\n<p>Here is my submission notebook [<a href=\"https://www.kaggle.com/code/sumaniitm/suman-submission-v2\" target=\"_blank\">https://www.kaggle.com/code/sumaniitm/suman-submission-v2</a>]</p>\n<p>When I run with 400 studies from the training set, my entire notebook finishes in less than 10min. However, the submission continues past 2 hours.<br>\nI am not using GPU/TPU</p>",
  "messages": [
    {
      "id": "2967145",
      "postDate": "08/22/2024 15:00:21",
      "content": "<p>Hi to fellow kagglers,<br>\nI am stuck at the submission stage and hitting a wall, would greatly appreciate any pointers to resolution.</p>\n<p>Here is my submission notebook [<a href=\"https://www.kaggle.com/code/sumaniitm/suman-submission-v2\" target=\"_blank\">https://www.kaggle.com/code/sumaniitm/suman-submission-v2</a>]</p>\n<p>When I run with 400 studies from the training set, my entire notebook finishes in less than 10min. However, the submission continues past 2 hours.<br>\nI am not using GPU/TPU</p>",
      "rawMarkdown": "Hi to fellow kagglers,\nI am stuck at the submission stage and hitting a wall, would greatly appreciate any pointers to resolution.\n\nHere is my submission notebook [https://www.kaggle.com/code/sumaniitm/suman-submission-v2]\n\nWhen I run with 400 studies from the training set, my entire notebook finishes in less than 10min. However, the submission continues past 2 hours.\nI am not using GPU/TPU",
      "votes": null
    },
    {
      "id": "2967162",
      "postDate": "08/22/2024 15:26:08",
      "content": "<p>`    test_dict = {}<br>\n    image_files = []</p>\n<pre><code> dirname, _, filenames  .walk([]):\n     filename  filenames:\n        test_dict[..join(dirname, filename).split()[]] = image_files\n        image_files.append(..join(dirname, filename))`\n</code></pre>\n<p>Are not you accumulating paths into each test_dict[key]? So, won't be last test_dict[key], specially for hidden test that has more than one study_id and 3 series_id, contain all available paths? The total of paths increases geometrically.</p>\n<p>I think it should be… Sorry I'm not sure the structure you are trying to achieve, but the think is that you need to reset image_files as you did in debugging path.</p>",
      "rawMarkdown": "`    test_dict = {}\n    image_files = []\n\n    for dirname, _, filenames in os.walk(config['root_file_path']):\n        for filename in filenames:\n            test_dict[os.path.join(dirname, filename).split('/')[-3]] = image_files\n            image_files.append(os.path.join(dirname, filename))`\n\nAre not you accumulating paths into each test_dict[key]? So, won't be last test_dict[key], specially for hidden test that has more than one study_id and 3 series_id, contain all available paths? The total of paths increases geometrically.\n\nI think it should be... Sorry I'm not sure the structure you are trying to achieve, but the think is that you need to reset image_files as you did in debugging path.",
      "votes": null
    },
    {
      "id": "2967177",
      "postDate": "08/22/2024 15:43:57",
      "content": "<p>thanks for your response, however, the paths are accumulated in the list which is the value of each key in test_dict<br>\ne.g.<br>\n<code>len(test_dict['1298611933'])</code> gives 82<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1017353%2F7910e05f3a73b490d07387b1fc053cc8%2FScreenshot%202024-08-22%20at%209.13.28PM.png?generation=1724341426865590&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "thanks for your response, however, the paths are accumulated in the list which is the value of each key in test_dict\ne.g.\n`len(test_dict['1298611933'])` gives 82\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1017353%2F7910e05f3a73b490d07387b1fc053cc8%2FScreenshot%202024-08-22%20at%209.13.28PM.png?generation=1724341426865590&alt=media)",
      "votes": null
    },
    {
      "id": "2967186",
      "postDate": "08/22/2024 15:51:38",
      "content": "<p>So the keys are study_id. But on that example I only see one series_id. I think the code is overwriting series_id's. Anyway, your debug code resets that list on each study_id, the normal execution doesn't.</p>",
      "rawMarkdown": "So the keys are study_id. But on that example I only see one series_id. I think the code is overwriting series_id's. Anyway, your debug code resets that list on each study_id, the normal execution doesn't.",
      "votes": null
    },
    {
      "id": "2967213",
      "postDate": "08/22/2024 16:26:03",
      "content": "<p>thank you, I will investigate on this route further. But the screenshot I pasted is a cut down version, so the other series ids are not visible in the screenshot</p>",
      "rawMarkdown": "thank you, I will investigate on this route further. But the screenshot I pasted is a cut down version, so the other series ids are not visible in the screenshot",
      "votes": null
    },
    {
      "id": "2970039",
      "postDate": "08/25/2024 16:37:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sumaniitm\" target=\"_blank\">@sumaniitm</a> ,</p>\n<p>Please check out below discussion.<br>\nIn this dataset, there are 7-study_ids in test_images folder on which you can test you code.</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523247\" target=\"_blank\">[NEW DATASET] Tiny Debug Dataset for RSNA 2024 Lumbar Spine Degenerative Classification</a></p>",
      "rawMarkdown": "Hi @sumaniitm ,\n\nPlease check out below discussion.\nIn this dataset, there are 7-study_ids in test_images folder on which you can test you code.\n\n[[NEW DATASET] Tiny Debug Dataset for RSNA 2024 Lumbar Spine Degenerative Classification](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523247)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2967162,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "08/22/2024 15:26:08",
      "content": "<p>`    test_dict = {}<br>\n    image_files = []</p>\n<pre><code> dirname, _, filenames  .walk([]):\n     filename  filenames:\n        test_dict[..join(dirname, filename).split()[]] = image_files\n        image_files.append(..join(dirname, filename))`\n</code></pre>\n<p>Are not you accumulating paths into each test_dict[key]? So, won't be last test_dict[key], specially for hidden test that has more than one study_id and 3 series_id, contain all available paths? The total of paths increases geometrically.</p>\n<p>I think it should be… Sorry I'm not sure the structure you are trying to achieve, but the think is that you need to reset image_files as you did in debugging path.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2967177,
          "author_name": "sumaniitm",
          "author_url": "",
          "post_date": "08/22/2024 15:43:57",
          "content": "<p>thanks for your response, however, the paths are accumulated in the list which is the value of each key in test_dict<br>\ne.g.<br>\n<code>len(test_dict['1298611933'])</code> gives 82<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1017353%2F7910e05f3a73b490d07387b1fc053cc8%2FScreenshot%202024-08-22%20at%209.13.28PM.png?generation=1724341426865590&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2967186,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "08/22/2024 15:51:38",
              "content": "<p>So the keys are study_id. But on that example I only see one series_id. I think the code is overwriting series_id's. Anyway, your debug code resets that list on each study_id, the normal execution doesn't.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2967213,
                  "author_name": "sumaniitm",
                  "author_url": "",
                  "post_date": "08/22/2024 16:26:03",
                  "content": "<p>thank you, I will investigate on this route further. But the screenshot I pasted is a cut down version, so the other series ids are not visible in the screenshot</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2970039,
      "author_name": "rohitchaudhari25",
      "author_url": "",
      "post_date": "08/25/2024 16:37:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sumaniitm\" target=\"_blank\">@sumaniitm</a> ,</p>\n<p>Please check out below discussion.<br>\nIn this dataset, there are 7-study_ids in test_images folder on which you can test you code.</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523247\" target=\"_blank\">[NEW DATASET] Tiny Debug Dataset for RSNA 2024 Lumbar Spine Degenerative Classification</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2967145": "Hi to fellow kagglers,\nI am stuck at the submission stage and hitting a wall, would greatly appreciate any pointers to resolution.\n\nHere is my submission notebook [https://www.kaggle.com/code/sumaniitm/suman-submission-v2]\n\nWhen I run with 400 studies from the training set, my entire notebook finishes in less than 10min. However, the submission continues past 2 hours.\nI am not using GPU/TPU",
    "2967162": "`    test_dict = {}\n    image_files = []\n\n    for dirname, _, filenames in os.walk(config['root_file_path']):\n        for filename in filenames:\n            test_dict[os.path.join(dirname, filename).split('/')[-3]] = image_files\n            image_files.append(os.path.join(dirname, filename))`\n\nAre not you accumulating paths into each test_dict[key]? So, won't be last test_dict[key], specially for hidden test that has more than one study_id and 3 series_id, contain all available paths? The total of paths increases geometrically.\n\nI think it should be... Sorry I'm not sure the structure you are trying to achieve, but the think is that you need to reset image_files as you did in debugging path.",
    "2967177": "thanks for your response, however, the paths are accumulated in the list which is the value of each key in test_dict\ne.g.\n`len(test_dict['1298611933'])` gives 82\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1017353%2F7910e05f3a73b490d07387b1fc053cc8%2FScreenshot%202024-08-22%20at%209.13.28PM.png?generation=1724341426865590&alt=media)",
    "2967186": "So the keys are study_id. But on that example I only see one series_id. I think the code is overwriting series_id's. Anyway, your debug code resets that list on each study_id, the normal execution doesn't.",
    "2967213": "thank you, I will investigate on this route further. But the screenshot I pasted is a cut down version, so the other series ids are not visible in the screenshot",
    "2970039": "Hi @sumaniitm ,\n\nPlease check out below discussion.\nIn this dataset, there are 7-study_ids in test_images folder on which you can test you code.\n\n[[NEW DATASET] Tiny Debug Dataset for RSNA 2024 Lumbar Spine Degenerative Classification](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523247)"
  },
  "source": "meta"
}