{
  "id": 174726,
  "title": "How can I save the complete CT pixel data?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/174726",
  "author_name": "",
  "post_date": "2020-08-15T02:15:07.822084700Z",
  "votes": -1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>When I try to iterate through each file and append the numpy array to a list, my notebook crashes from memory overuse.</p>",
  "messages": [
    {
      "id": "970920",
      "postDate": "08/15/2020 02:15:07",
      "content": "<p>When I try to iterate through each file and append the numpy array to a list, my notebook crashes from memory overuse.</p>",
      "rawMarkdown": "When I try to iterate through each file and append the numpy array to a list, my notebook crashes from memory overuse.",
      "votes": null
    },
    {
      "id": "970934",
      "postDate": "08/15/2020 02:45:40",
      "content": "<p>Maybe you could average some of the files' (numpy arrays) so that it fits in the memory?</p>",
      "rawMarkdown": "Maybe you could average some of the files' (numpy arrays) so that it fits in the memory?",
      "votes": null
    },
    {
      "id": "970994",
      "postDate": "08/15/2020 04:44:58",
      "content": "<p>Hmm, what exactly do you mean by that? Like store each patient CT as its own numpy array?</p>",
      "rawMarkdown": "Hmm, what exactly do you mean by that? Like store each patient CT as its own numpy array?",
      "votes": null
    },
    {
      "id": "971651",
      "postDate": "08/15/2020 19:14:35",
      "content": "<p>I thought you meant you ran out of RAM. So I was thinking if you could average some slices together it would fit in the memory?</p>",
      "rawMarkdown": "I thought you meant you ran out of RAM. So I was thinking if you could average some slices together it would fit in the memory?",
      "votes": null
    },
    {
      "id": "971705",
      "postDate": "08/15/2020 20:17:51",
      "content": "<p>Appending to a Numpy array makes a copy. That can eat up your memory. Resize the Dicom images as you go, and store as Int8 if possible. Even with those changes, you might not be able to fit everything in memory. Look at pipelines that read the dicom as it is needed.</p>",
      "rawMarkdown": "Appending to a Numpy array makes a copy. That can eat up your memory. Resize the Dicom images as you go, and store as Int8 if possible. Even with those changes, you might not be able to fit everything in memory. Look at pipelines that read the dicom as it is needed.",
      "votes": null
    },
    {
      "id": "971796",
      "postDate": "08/15/2020 22:27:22",
      "content": "<p>Sorry, I mean that I am currently appending arrays to a list, not appending to an array. But I think that you might be right a pipeline is probably the way to go.</p>",
      "rawMarkdown": "Sorry, I mean that I am currently appending arrays to a list, not appending to an array. But I think that you might be right a pipeline is probably the way to go.",
      "votes": null
    },
    {
      "id": "971797",
      "postDate": "08/15/2020 22:27:54",
      "content": "<p>Sorry, I'm not really sure what you mean. I want to keep as much of the data as possible so removing slices wouldn't be ideal for me.</p>",
      "rawMarkdown": "Sorry, I'm not really sure what you mean. I want to keep as much of the data as possible so removing slices wouldn't be ideal for me.",
      "votes": null
    },
    {
      "id": "971827",
      "postDate": "08/16/2020 00:30:05",
      "content": "<p><a href=\"https://www.kaggle.com/eladwar/20-seconds-or-less\">https://www.kaggle.com/eladwar/20-seconds-or-less</a> . </p>",
      "rawMarkdown": "https://www.kaggle.com/eladwar/20-seconds-or-less .",
      "votes": null
    },
    {
      "id": "994077",
      "postDate": "09/01/2020 10:59:02",
      "content": "<p>use, for example, tensorflow's tf.data API to achieve memory efficient data input pipelines. in huge datasets you'll be forced into caching and other mem efficient approaches anyways.  <a href=\"https://cs230.stanford.edu/blog/datapipeline/\" target=\"_blank\">tf.data input guide</a></p>",
      "rawMarkdown": "use, for example, tensorflow's tf.data API to achieve memory efficient data input pipelines. in huge datasets you'll be forced into caching and other mem efficient approaches anyways.  [tf.data input guide](https://cs230.stanford.edu/blog/datapipeline/)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 970934,
      "author_name": "jonykarki",
      "author_url": "",
      "post_date": "08/15/2020 02:45:40",
      "content": "<p>Maybe you could average some of the files' (numpy arrays) so that it fits in the memory?</p>",
      "votes": null,
      "replies": [
        {
          "id": 970994,
          "author_name": "akiroduey",
          "author_url": "",
          "post_date": "08/15/2020 04:44:58",
          "content": "<p>Hmm, what exactly do you mean by that? Like store each patient CT as its own numpy array?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971651,
          "author_name": "jonykarki",
          "author_url": "",
          "post_date": "08/15/2020 19:14:35",
          "content": "<p>I thought you meant you ran out of RAM. So I was thinking if you could average some slices together it would fit in the memory?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971797,
          "author_name": "akiroduey",
          "author_url": "",
          "post_date": "08/15/2020 22:27:54",
          "content": "<p>Sorry, I'm not really sure what you mean. I want to keep as much of the data as possible so removing slices wouldn't be ideal for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 971705,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "08/15/2020 20:17:51",
      "content": "<p>Appending to a Numpy array makes a copy. That can eat up your memory. Resize the Dicom images as you go, and store as Int8 if possible. Even with those changes, you might not be able to fit everything in memory. Look at pipelines that read the dicom as it is needed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 971796,
          "author_name": "akiroduey",
          "author_url": "",
          "post_date": "08/15/2020 22:27:22",
          "content": "<p>Sorry, I mean that I am currently appending arrays to a list, not appending to an array. But I think that you might be right a pipeline is probably the way to go.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 994077,
      "author_name": "dronych",
      "author_url": "",
      "post_date": "09/01/2020 10:59:02",
      "content": "<p>use, for example, tensorflow's tf.data API to achieve memory efficient data input pipelines. in huge datasets you'll be forced into caching and other mem efficient approaches anyways.  <a href=\"https://cs230.stanford.edu/blog/datapipeline/\" target=\"_blank\">tf.data input guide</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 971827,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "08/16/2020 00:30:05",
      "content": "<p><a href=\"https://www.kaggle.com/eladwar/20-seconds-or-less\">https://www.kaggle.com/eladwar/20-seconds-or-less</a> . </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "970920": "When I try to iterate through each file and append the numpy array to a list, my notebook crashes from memory overuse.",
    "970934": "Maybe you could average some of the files' (numpy arrays) so that it fits in the memory?",
    "970994": "Hmm, what exactly do you mean by that? Like store each patient CT as its own numpy array?",
    "971651": "I thought you meant you ran out of RAM. So I was thinking if you could average some slices together it would fit in the memory?",
    "971705": "Appending to a Numpy array makes a copy. That can eat up your memory. Resize the Dicom images as you go, and store as Int8 if possible. Even with those changes, you might not be able to fit everything in memory. Look at pipelines that read the dicom as it is needed.",
    "971796": "Sorry, I mean that I am currently appending arrays to a list, not appending to an array. But I think that you might be right a pipeline is probably the way to go.",
    "971797": "Sorry, I'm not really sure what you mean. I want to keep as much of the data as possible so removing slices wouldn't be ideal for me.",
    "971827": "https://www.kaggle.com/eladwar/20-seconds-or-less .",
    "994077": "use, for example, tensorflow's tf.data API to achieve memory efficient data input pipelines. in huge datasets you'll be forced into caching and other mem efficient approaches anyways.  [tf.data input guide](https://cs230.stanford.edu/blog/datapipeline/)"
  },
  "source": "meta"
}