{
  "id": 18257,
  "title": "Basic understanding of SAX, 2ch, 4ch",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18257",
  "author_name": "",
  "post_date": "2016-01-05T15:11:18.183Z",
  "votes": 1,
  "comment_count": 6,
  "views": 2191,
  "content": "<p>Let me preface my question with a quote from &quot;Ghostbusters II&quot; to set the theme:</p>\n\n<blockquote>\n  <p>Dr. Peter Venkman: Ray, pretend for a moment that I don't know\n  anything about metallurgy, engineering, or physics, and just tell me\n  what the hell is going on.</p>\n  \n  <p>Dr Ray Stantz: You never studied.</p>\n</blockquote>\n\n<p>OK, so, I read the Data description and watch the short video, but I'm still struggling to wrap my mind around this dataset.</p>\n\n<p>It seems that the sax_ series of folders contain images/slices which are ultimately more relevant, given the angle at which the images were taken.</p>\n\n<p>However, when trying to structure this data into a usable training dataset, I'm still confused by the facts that </p>\n\n<ol>\n<li>there are multiple, but inconsistent, numbers of sax_ folders per patient</li>\n<li>the numbers appended to the sax_ folders are not in common for all patients, e.g. some have sax_6 and others don't</li>\n<li>the numbers appended to 2ch_ and 4ch_ are also inconsistent between patients</li>\n<li>There are also some inconsistencies in the number of header (hdr) fields within the sax_ images</li>\n</ol>\n\n<p>I'm stronger in R than Python, which is perhaps part of the problem. The 2 official tutorials are in Python, so perhaps I can answer some of my own questions by getting past my R-bias and trying a little harder to follow those tuts. However, if some kind soul wouldn't mind helping me to get past this initial confusion, then I'd greatly appreciate it!</p>",
  "messages": [
    {
      "id": "103683",
      "postDate": "01/05/2016 15:11:18",
      "content": "<p>Let me preface my question with a quote from &quot;Ghostbusters II&quot; to set the theme:</p>\n\n<blockquote>\n  <p>Dr. Peter Venkman: Ray, pretend for a moment that I don't know\n  anything about metallurgy, engineering, or physics, and just tell me\n  what the hell is going on.</p>\n  \n  <p>Dr Ray Stantz: You never studied.</p>\n</blockquote>\n\n<p>OK, so, I read the Data description and watch the short video, but I'm still struggling to wrap my mind around this dataset.</p>\n\n<p>It seems that the sax_ series of folders contain images/slices which are ultimately more relevant, given the angle at which the images were taken.</p>\n\n<p>However, when trying to structure this data into a usable training dataset, I'm still confused by the facts that </p>\n\n<ol>\n<li>there are multiple, but inconsistent, numbers of sax_ folders per patient</li>\n<li>the numbers appended to the sax_ folders are not in common for all patients, e.g. some have sax_6 and others don't</li>\n<li>the numbers appended to 2ch_ and 4ch_ are also inconsistent between patients</li>\n<li>There are also some inconsistencies in the number of header (hdr) fields within the sax_ images</li>\n</ol>\n\n<p>I'm stronger in R than Python, which is perhaps part of the problem. The 2 official tutorials are in Python, so perhaps I can answer some of my own questions by getting past my R-bias and trying a little harder to follow those tuts. However, if some kind soul wouldn't mind helping me to get past this initial confusion, then I'd greatly appreciate it!</p>",
      "rawMarkdown": "Let me preface my question with a quote from \"Ghostbusters II\" to set the theme:\r\n\r\n> Dr. Peter Venkman: Ray, pretend for a moment that I don't know\r\n> anything about metallurgy, engineering, or physics, and just tell me\r\n> what the hell is going on.\r\n> \r\n> Dr Ray Stantz: You never studied.\r\n\r\nOK, so, I read the Data description and watch the short video, but I'm still struggling to wrap my mind around this dataset.\r\n\r\nIt seems that the sax_ series of folders contain images/slices which are ultimately more relevant, given the angle at which the images were taken.\r\n\r\nHowever, when trying to structure this data into a usable training dataset, I'm still confused by the facts that \r\n\r\n 1. there are multiple, but inconsistent, numbers of sax_ folders per patient\r\n 2. the numbers appended to the sax_ folders are not in common for all patients, e.g. some have sax_6 and others don't\r\n 3. the numbers appended to 2ch_ and 4ch_ are also inconsistent between patients\r\n 4. There are also some inconsistencies in the number of header (hdr) fields within the sax_ images\r\n\r\nI'm stronger in R than Python, which is perhaps part of the problem. The 2 official tutorials are in Python, so perhaps I can answer some of my own questions by getting past my R-bias and trying a little harder to follow those tuts. However, if some kind soul wouldn't mind helping me to get past this initial confusion, then I'd greatly appreciate it!",
      "votes": null
    },
    {
      "id": "103687",
      "postDate": "01/05/2016 15:28:21",
      "content": "<p>This is real clinical data, so it is messy and inconsistent. It would be nicer for computers (and the people programming them) if that was not the case, but it is unfortunately messy. Different patients have a different number of slices, etc. acquired. Sometimes the naming varies slightly, sometimes slices need to be repeated (often because the patient is unable to breath-hold). And so on. </p>\n\n<p>But generally the datasets consist of what we call a short axis stack. It is a set of images acquired in the &quot;short-axis&quot; view, i.e. perpendicular to the &quot;long-axis&quot; views. These slices are really what you need to determine the volumes of the LV at different points in time. Specifically, if you segment out the blood pool and multiply with the gap between slices and sum up for all slices, you will get the volume (Simpson's rule). The long axis views that we have tried to include are a 2-chamber view, which should contain a view of the left atrium and the left ventricle and the 4-chamber view, which should show you all 4 chambers of the heart. They can help guide you to figure out where the atrium ends and the ventricle begins. On the left side of the heart those two chambers are separated by a valve known as the mitral valve. The two long axis views can also assist you in the identification of the left ventricle, since it should be located where the 2-ch and the 4-ch views intersect. </p>\n\n<p>Maybe something like this is helpful:</p>\n\n<p><a href=\"http://www.scmr.org/assets/files/members/documents/Cardiac_views.pdf\">http://www.scmr.org/assets/files/members/documents/Cardiac_views.pdf</a></p>\n\n<p>Hope this helps a bit,</p>\n\n<p>Michael</p>",
      "rawMarkdown": "This is real clinical data, so it is messy and inconsistent. It would be nicer for computers (and the people programming them) if that was not the case, but it is unfortunately messy. Different patients have a different number of slices, etc. acquired. Sometimes the naming varies slightly, sometimes slices need to be repeated (often because the patient is unable to breath-hold). And so on. \r\n\r\nBut generally the datasets consist of what we call a short axis stack. It is a set of images acquired in the \"short-axis\" view, i.e. perpendicular to the \"long-axis\" views. These slices are really what you need to determine the volumes of the LV at different points in time. Specifically, if you segment out the blood pool and multiply with the gap between slices and sum up for all slices, you will get the volume (Simpson's rule). The long axis views that we have tried to include are a 2-chamber view, which should contain a view of the left atrium and the left ventricle and the 4-chamber view, which should show you all 4 chambers of the heart. They can help guide you to figure out where the atrium ends and the ventricle begins. On the left side of the heart those two chambers are separated by a valve known as the mitral valve. The two long axis views can also assist you in the identification of the left ventricle, since it should be located where the 2-ch and the 4-ch views intersect. \r\n\r\nMaybe something like this is helpful:\r\n\r\nhttp://www.scmr.org/assets/files/members/documents/Cardiac_views.pdf\r\n\r\nHope this helps a bit,\r\n\r\nMichael",
      "votes": null
    },
    {
      "id": "103693",
      "postDate": "01/05/2016 15:59:40",
      "content": "<p>Thank you very much!</p>\n\n<p>Yes, that does help a lot. Just to be 100% clear, what do the numbers appended to the folder names represent? Like the &quot;6&quot; in &quot;sax_6&quot;?</p>\n\n<p>Does that mean that it was the 6th time this patient had the MRI done (thus I can consider the different SAX folders to represent a kind of time series of MRI sessions) or something else?</p>\n\n<p>Thanks again.</p>",
      "rawMarkdown": "Thank you very much!\r\n\r\nYes, that does help a lot. Just to be 100% clear, what do the numbers appended to the folder names represent? Like the \"6\" in \"sax_6\"?\r\n\r\nDoes that mean that it was the 6th time this patient had the MRI done (thus I can consider the different SAX folders to represent a kind of time series of MRI sessions) or something else?\r\n\r\nThanks again.",
      "votes": null
    },
    {
      "id": "103694",
      "postDate": "01/05/2016 16:01:45",
      "content": "<p>Yes, different folder should be different time series, but probably from the same imaging session. Each imaging session will contain multiple scans.</p>",
      "rawMarkdown": "Yes, different folder should be different time series, but probably from the same imaging session. Each imaging session will contain multiple scans.",
      "votes": null
    },
    {
      "id": "103713",
      "postDate": "01/05/2016 18:45:22",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "103721",
      "postDate": "01/05/2016 20:41:23",
      "content": "<p>But don't the MRI time-series in  &quot;sax_6&quot; and &quot;sax_7&quot; , for example, correspond to consecutive slices of the same heart?</p>",
      "rawMarkdown": "But don't the MRI time-series in  \"sax_6\" and \"sax_7\" , for example, correspond to consecutive slices of the same heart?",
      "votes": null
    },
    {
      "id": "103722",
      "postDate": "01/05/2016 20:45:28",
      "content": "<p>That is not guaranteed in general. Some slices are repeated, etc. You have to look at the slice location in the meta data fields as described in the tutorials. </p>",
      "rawMarkdown": "That is not guaranteed in general. Some slices are repeated, etc. You have to look at the slice location in the meta data fields as described in the tutorials.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 103687,
      "author_name": "michaelhansen",
      "author_url": "",
      "post_date": "01/05/2016 15:28:21",
      "content": "<p>This is real clinical data, so it is messy and inconsistent. It would be nicer for computers (and the people programming them) if that was not the case, but it is unfortunately messy. Different patients have a different number of slices, etc. acquired. Sometimes the naming varies slightly, sometimes slices need to be repeated (often because the patient is unable to breath-hold). And so on. </p>\n\n<p>But generally the datasets consist of what we call a short axis stack. It is a set of images acquired in the &quot;short-axis&quot; view, i.e. perpendicular to the &quot;long-axis&quot; views. These slices are really what you need to determine the volumes of the LV at different points in time. Specifically, if you segment out the blood pool and multiply with the gap between slices and sum up for all slices, you will get the volume (Simpson's rule). The long axis views that we have tried to include are a 2-chamber view, which should contain a view of the left atrium and the left ventricle and the 4-chamber view, which should show you all 4 chambers of the heart. They can help guide you to figure out where the atrium ends and the ventricle begins. On the left side of the heart those two chambers are separated by a valve known as the mitral valve. The two long axis views can also assist you in the identification of the left ventricle, since it should be located where the 2-ch and the 4-ch views intersect. </p>\n\n<p>Maybe something like this is helpful:</p>\n\n<p><a href=\"http://www.scmr.org/assets/files/members/documents/Cardiac_views.pdf\">http://www.scmr.org/assets/files/members/documents/Cardiac_views.pdf</a></p>\n\n<p>Hope this helps a bit,</p>\n\n<p>Michael</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103693,
      "author_name": "millerintllc",
      "author_url": "",
      "post_date": "01/05/2016 15:59:40",
      "content": "<p>Thank you very much!</p>\n\n<p>Yes, that does help a lot. Just to be 100% clear, what do the numbers appended to the folder names represent? Like the &quot;6&quot; in &quot;sax_6&quot;?</p>\n\n<p>Does that mean that it was the 6th time this patient had the MRI done (thus I can consider the different SAX folders to represent a kind of time series of MRI sessions) or something else?</p>\n\n<p>Thanks again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103694,
      "author_name": "michaelhansen",
      "author_url": "",
      "post_date": "01/05/2016 16:01:45",
      "content": "<p>Yes, different folder should be different time series, but probably from the same imaging session. Each imaging session will contain multiple scans.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103713,
      "author_name": "millerintllc",
      "author_url": "",
      "post_date": "01/05/2016 18:45:22",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103721,
      "author_name": "rfarouni",
      "author_url": "",
      "post_date": "01/05/2016 20:41:23",
      "content": "<p>But don't the MRI time-series in  &quot;sax_6&quot; and &quot;sax_7&quot; , for example, correspond to consecutive slices of the same heart?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103722,
      "author_name": "michaelhansen",
      "author_url": "",
      "post_date": "01/05/2016 20:45:28",
      "content": "<p>That is not guaranteed in general. Some slices are repeated, etc. You have to look at the slice location in the meta data fields as described in the tutorials. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "103683": "Let me preface my question with a quote from \"Ghostbusters II\" to set the theme:\r\n\r\n> Dr. Peter Venkman: Ray, pretend for a moment that I don't know\r\n> anything about metallurgy, engineering, or physics, and just tell me\r\n> what the hell is going on.\r\n> \r\n> Dr Ray Stantz: You never studied.\r\n\r\nOK, so, I read the Data description and watch the short video, but I'm still struggling to wrap my mind around this dataset.\r\n\r\nIt seems that the sax_ series of folders contain images/slices which are ultimately more relevant, given the angle at which the images were taken.\r\n\r\nHowever, when trying to structure this data into a usable training dataset, I'm still confused by the facts that \r\n\r\n 1. there are multiple, but inconsistent, numbers of sax_ folders per patient\r\n 2. the numbers appended to the sax_ folders are not in common for all patients, e.g. some have sax_6 and others don't\r\n 3. the numbers appended to 2ch_ and 4ch_ are also inconsistent between patients\r\n 4. There are also some inconsistencies in the number of header (hdr) fields within the sax_ images\r\n\r\nI'm stronger in R than Python, which is perhaps part of the problem. The 2 official tutorials are in Python, so perhaps I can answer some of my own questions by getting past my R-bias and trying a little harder to follow those tuts. However, if some kind soul wouldn't mind helping me to get past this initial confusion, then I'd greatly appreciate it!",
    "103687": "This is real clinical data, so it is messy and inconsistent. It would be nicer for computers (and the people programming them) if that was not the case, but it is unfortunately messy. Different patients have a different number of slices, etc. acquired. Sometimes the naming varies slightly, sometimes slices need to be repeated (often because the patient is unable to breath-hold). And so on. \r\n\r\nBut generally the datasets consist of what we call a short axis stack. It is a set of images acquired in the \"short-axis\" view, i.e. perpendicular to the \"long-axis\" views. These slices are really what you need to determine the volumes of the LV at different points in time. Specifically, if you segment out the blood pool and multiply with the gap between slices and sum up for all slices, you will get the volume (Simpson's rule). The long axis views that we have tried to include are a 2-chamber view, which should contain a view of the left atrium and the left ventricle and the 4-chamber view, which should show you all 4 chambers of the heart. They can help guide you to figure out where the atrium ends and the ventricle begins. On the left side of the heart those two chambers are separated by a valve known as the mitral valve. The two long axis views can also assist you in the identification of the left ventricle, since it should be located where the 2-ch and the 4-ch views intersect. \r\n\r\nMaybe something like this is helpful:\r\n\r\nhttp://www.scmr.org/assets/files/members/documents/Cardiac_views.pdf\r\n\r\nHope this helps a bit,\r\n\r\nMichael",
    "103693": "Thank you very much!\r\n\r\nYes, that does help a lot. Just to be 100% clear, what do the numbers appended to the folder names represent? Like the \"6\" in \"sax_6\"?\r\n\r\nDoes that mean that it was the 6th time this patient had the MRI done (thus I can consider the different SAX folders to represent a kind of time series of MRI sessions) or something else?\r\n\r\nThanks again.",
    "103694": "Yes, different folder should be different time series, but probably from the same imaging session. Each imaging session will contain multiple scans.",
    "103713": "Thanks!",
    "103721": "But don't the MRI time-series in  \"sax_6\" and \"sax_7\" , for example, correspond to consecutive slices of the same heart?",
    "103722": "That is not guaranteed in general. Some slices are repeated, etc. You have to look at the slice location in the meta data fields as described in the tutorials."
  },
  "source": "meta"
}