{
  "id": 175244,
  "title": "Loading the private test set scans takes a long time",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/175244",
  "author_name": "",
  "post_date": "2020-08-17T16:04:54.187892400Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey folks, </p>\n<p>I'm wondering if there's something wrong with my code, or you might also have experienced a super slow (which results in a timeout error) loading of the private test set scans to memory? Notebook can run with no issues when creating a version. This happens when the notebook is submitted.</p>",
  "messages": [
    {
      "id": "973935",
      "postDate": "08/17/2020 16:04:54",
      "content": "<p>Hey folks, </p>\n<p>I'm wondering if there's something wrong with my code, or you might also have experienced a super slow (which results in a timeout error) loading of the private test set scans to memory? Notebook can run with no issues when creating a version. This happens when the notebook is submitted.</p>",
      "rawMarkdown": "Hey folks, \n\nI'm wondering if there's something wrong with my code, or you might also have experienced a super slow (which results in a timeout error) loading of the private test set scans to memory? Notebook can run with no issues when creating a version. This happens when the notebook is submitted.",
      "votes": null
    },
    {
      "id": "975471",
      "postDate": "08/18/2020 10:03:36",
      "content": "<p>Hello. Please, read again about the data: \"This is a synchronous rerun code competition. The provided test set is a small representative set of files (copied from the training set) to demonstrate the format of the private test set. When you submit your notebook, Kaggle will rerun your code on the test set, which contains unseen images.\"</p>",
      "rawMarkdown": "Hello. Please, read again about the data: \"This is a synchronous rerun code competition. The provided test set is a small representative set of files (copied from the training set) to demonstrate the format of the private test set. When you submit your notebook, Kaggle will rerun your code on the test set, which contains unseen images.\"",
      "votes": null
    },
    {
      "id": "975982",
      "postDate": "08/18/2020 15:04:58",
      "content": "<p>Thanks for the reply. I knew about this! I thought reading the unseen images should not take so long which ultimately ends with a <strong>timeout error</strong>.</p>",
      "rawMarkdown": "Thanks for the reply. I knew about this! I thought reading the unseen images should not take so long which ultimately ends with a **timeout error**.",
      "votes": null
    },
    {
      "id": "976137",
      "postDate": "08/18/2020 16:48:58",
      "content": "<p>I'm getting the same issue. Do you process all the scans at once ? I'm thinking about extracting features from the scans during data loader process (if you use PyTorch), .i.e on the fly, so that you can do your stuff in parallel. I'm working only with one scan per patient for time being so I did not try yet</p>",
      "rawMarkdown": "I'm getting the same issue. Do you process all the scans at once ? I'm thinking about extracting features from the scans during data loader process (if you use PyTorch), .i.e on the fly, so that you can do your stuff in parallel. I'm working only with one scan per patient for time being so I did not try yet",
      "votes": null
    },
    {
      "id": "976142",
      "postDate": "08/18/2020 16:53:48",
      "content": "<p>btw you should have a look on this <a href=\"https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram\" target=\"_blank\">kernel</a> which should fit your expectations</p>",
      "rawMarkdown": "btw you should have a look on this [kernel] (https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram) which should fit your expectations",
      "votes": null
    },
    {
      "id": "976203",
      "postDate": "08/18/2020 17:53:12",
      "content": "<p>Thanks! Will take a look at this.</p>",
      "rawMarkdown": "Thanks! Will take a look at this.",
      "votes": null
    },
    {
      "id": "976205",
      "postDate": "08/18/2020 17:57:05",
      "content": "<p>I preferably wanted to load and resize all the scans into a 3D volume, and save them for faster processing, but the code is stuck at loading stage and results in a timeout error. I'm trying another approach which does similar to what you suggested, i.e., loading the scan volumes, extracting features (latents), cache and return them as the <code>DataLoader</code> output. Still haven't been able to see if it works.</p>",
      "rawMarkdown": "I preferably wanted to load and resize all the scans into a 3D volume, and save them for faster processing, but the code is stuck at loading stage and results in a timeout error. I'm trying another approach which does similar to what you suggested, i.e., loading the scan volumes, extracting features (latents), cache and return them as the `DataLoader` output. Still haven't been able to see if it works.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 975471,
      "author_name": "eliasgreen",
      "author_url": "",
      "post_date": "08/18/2020 10:03:36",
      "content": "<p>Hello. Please, read again about the data: \"This is a synchronous rerun code competition. The provided test set is a small representative set of files (copied from the training set) to demonstrate the format of the private test set. When you submit your notebook, Kaggle will rerun your code on the test set, which contains unseen images.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 975982,
          "author_name": "keyvanl6",
          "author_url": "",
          "post_date": "08/18/2020 15:04:58",
          "content": "<p>Thanks for the reply. I knew about this! I thought reading the unseen images should not take so long which ultimately ends with a <strong>timeout error</strong>.</p>",
          "votes": null,
          "replies": [
            {
              "id": 976137,
              "author_name": "alexj21",
              "author_url": "",
              "post_date": "08/18/2020 16:48:58",
              "content": "<p>I'm getting the same issue. Do you process all the scans at once ? I'm thinking about extracting features from the scans during data loader process (if you use PyTorch), .i.e on the fly, so that you can do your stuff in parallel. I'm working only with one scan per patient for time being so I did not try yet</p>",
              "votes": null,
              "replies": [
                {
                  "id": 976142,
                  "author_name": "alexj21",
                  "author_url": "",
                  "post_date": "08/18/2020 16:53:48",
                  "content": "<p>btw you should have a look on this <a href=\"https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram\" target=\"_blank\">kernel</a> which should fit your expectations</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 976203,
                  "author_name": "keyvanl6",
                  "author_url": "",
                  "post_date": "08/18/2020 17:53:12",
                  "content": "<p>Thanks! Will take a look at this.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 976205,
              "author_name": "keyvanl6",
              "author_url": "",
              "post_date": "08/18/2020 17:57:05",
              "content": "<p>I preferably wanted to load and resize all the scans into a 3D volume, and save them for faster processing, but the code is stuck at loading stage and results in a timeout error. I'm trying another approach which does similar to what you suggested, i.e., loading the scan volumes, extracting features (latents), cache and return them as the <code>DataLoader</code> output. Still haven't been able to see if it works.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "973935": "Hey folks, \n\nI'm wondering if there's something wrong with my code, or you might also have experienced a super slow (which results in a timeout error) loading of the private test set scans to memory? Notebook can run with no issues when creating a version. This happens when the notebook is submitted.",
    "975471": "Hello. Please, read again about the data: \"This is a synchronous rerun code competition. The provided test set is a small representative set of files (copied from the training set) to demonstrate the format of the private test set. When you submit your notebook, Kaggle will rerun your code on the test set, which contains unseen images.\"",
    "975982": "Thanks for the reply. I knew about this! I thought reading the unseen images should not take so long which ultimately ends with a **timeout error**.",
    "976137": "I'm getting the same issue. Do you process all the scans at once ? I'm thinking about extracting features from the scans during data loader process (if you use PyTorch), .i.e on the fly, so that you can do your stuff in parallel. I'm working only with one scan per patient for time being so I did not try yet",
    "976142": "btw you should have a look on this [kernel] (https://www.kaggle.com/jameschapman19/pytorch-tabular-qr-histogram) which should fit your expectations",
    "976203": "Thanks! Will take a look at this.",
    "976205": "I preferably wanted to load and resize all the scans into a 3D volume, and save them for faster processing, but the code is stuck at loading stage and results in a timeout error. I'm trying another approach which does similar to what you suggested, i.e., loading the scan volumes, extracting features (latents), cache and return them as the `DataLoader` output. Still haven't been able to see if it works."
  },
  "source": "meta"
}