{
  "id": 17955,
  "title": "Information on directory structure and naming",
  "url": "/competitions/second-annual-data-science-bowl/discussion/17955",
  "author_name": "",
  "post_date": "2015-12-16T11:44:44.973Z",
  "votes": 13,
  "comment_count": 10,
  "views": 1915,
  "content": "<p>I would appreciate some clarification on the naming of the directories/images we are given. Once I've extracted train.zip, I understand that each individual sample is located in its own directory (1, 2, ..., 500). Inside the 'study' directory of each of these samples there are several directories, I'm assuming from different angles/perspectives. My specific questions are:</p>\n\n<ul>\n<li>What does #ch_# mean? For example with 'train/1/study/2ch_21', what do the '2', 'ch', and '21' indicate?</li>\n<li>We have been told that images prefixed with 'sax_' are short axis cine images, but what do the suffixes to each of the directories mean? For example with 'train/1/study/sax_5', what does the '5' indicate?</li>\n<li>For the images in each of these directories, are we guaranteed to have 30 images? And do these 30 images always represent one heart beat or may they represent many heart beats? For example, 'train/1/study/sax_5/' has 30 images -- do all similar directories have 30 images and are they over an entire heart beat?</li>\n</ul>",
  "messages": [
    {
      "id": "101673",
      "postDate": "12/16/2015 11:44:44",
      "content": "<p>I would appreciate some clarification on the naming of the directories/images we are given. Once I've extracted train.zip, I understand that each individual sample is located in its own directory (1, 2, ..., 500). Inside the 'study' directory of each of these samples there are several directories, I'm assuming from different angles/perspectives. My specific questions are:</p>\n\n<ul>\n<li>What does #ch_# mean? For example with 'train/1/study/2ch_21', what do the '2', 'ch', and '21' indicate?</li>\n<li>We have been told that images prefixed with 'sax_' are short axis cine images, but what do the suffixes to each of the directories mean? For example with 'train/1/study/sax_5', what does the '5' indicate?</li>\n<li>For the images in each of these directories, are we guaranteed to have 30 images? And do these 30 images always represent one heart beat or may they represent many heart beats? For example, 'train/1/study/sax_5/' has 30 images -- do all similar directories have 30 images and are they over an entire heart beat?</li>\n</ul>",
      "rawMarkdown": "I would appreciate some clarification on the naming of the directories/images we are given. Once I've extracted train.zip, I understand that each individual sample is located in its own directory (1, 2, ..., 500). Inside the 'study' directory of each of these samples there are several directories, I'm assuming from different angles/perspectives. My specific questions are:\r\n\r\n- What does #ch_# mean? For example with 'train/1/study/2ch_21', what do the '2', 'ch', and '21' indicate?\r\n- We have been told that images prefixed with 'sax_' are short axis cine images, but what do the suffixes to each of the directories mean? For example with 'train/1/study/sax_5', what does the '5' indicate?\r\n- For the images in each of these directories, are we guaranteed to have 30 images? And do these 30 images always represent one heart beat or may they represent many heart beats? For example, 'train/1/study/sax_5/' has 30 images -- do all similar directories have 30 images and are they over an entire heart beat?",
      "votes": null
    },
    {
      "id": "101913",
      "postDate": "12/17/2015 20:44:54",
      "content": "<p>2ch and 4ch are the 2- and 4-chamber views.</p>",
      "rawMarkdown": "2ch and 4ch are the 2- and 4-chamber views.",
      "votes": null
    },
    {
      "id": "101957",
      "postDate": "12/18/2015 03:13:42",
      "content": "<p>megaminion asked:</p>\n\n<blockquote>\n  <p>For the images in each of these directories, are we guaranteed to have 30 images?</p>\n</blockquote>\n\n<p>That doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:</p>\n\n<ul>\n<li>(study count): (number of images)</li>\n<li>6274:  30</li>\n<li>26: 60</li>\n<li>9: 25</li>\n<li>3: 330</li>\n<li>2: 270</li>\n<li>2: 21</li>\n<li>1: 23</li>\n<li>1: 22</li>\n</ul>\n\n<p>Spot checking a couple directories with 30 images we have filenames like:\n<em>train/17/study/2ch_23/IM-5996-NNNN.dcm</em> - Where NNNN runs from 0000-0030</p>\n\n<p>The 60 image directories have filenames like:\n<em>train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm</em> - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002</p>\n\n<p>And the 330 image directory seems to follow this pattern:\n<em>train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm</em> - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011</p>\n\n<p>Without looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:</p>\n\n<blockquote>\n  <p>Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.</p>\n</blockquote>",
      "rawMarkdown": "megaminion asked:\r\n\r\n> For the images in each of these directories, are we guaranteed to have 30 images?\r\n\r\nThat doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:\r\n\r\n- (study count): (number of images)\r\n- 6274:  30\r\n- 26: 60\r\n- 9: 25\r\n- 3: 330\r\n- 2: 270\r\n- 2: 21\r\n- 1: 23\r\n- 1: 22\r\n\r\n\r\nSpot checking a couple directories with 30 images we have filenames like:\r\n*train/17/study/2ch_23/IM-5996-NNNN.dcm* - Where NNNN runs from 0000-0030\r\n\r\nThe 60 image directories have filenames like:\r\n*train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm* - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002\r\n\r\nAnd the 330 image directory seems to follow this pattern:\r\n*train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm* - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011\r\n\r\nWithout looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:\r\n\r\n> Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.",
      "votes": null
    },
    {
      "id": "101978",
      "postDate": "12/18/2015 08:20:59",
      "content": "<p>Tests with numer of images more or less than 30 in sax_* directories:</p>\n\n<p>Tests TRAIN: 123, 234, 279, 334, 416, 463, 499</p>\n\n<p>Tests VALID: 516, 590, 597, 619, 623</p>\n\n<p>I don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?</p>",
      "rawMarkdown": "Tests with numer of images more or less than 30 in sax_* directories:\r\n\r\nTests TRAIN: 123, 234, 279, 334, 416, 463, 499\r\n\r\nTests VALID: 516, 590, 597, 619, 623\r\n\r\nI don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?",
      "votes": null
    },
    {
      "id": "102006",
      "postDate": "12/18/2015 14:41:23",
      "content": "<p>[quote=ZFTurbo;101978]</p>\n\n<p>Tests with numer of images more or less than 30 in sax_* directories:</p>\n\n<p>Tests TRAIN: 123, 234, 279, 334, 416, 463, 499</p>\n\n<p>Tests VALID: 516, 590, 597, 619, 623</p>\n\n<p>I don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?</p>\n\n<p>[/quote]</p>\n\n<p>We suggest you attempt to automate the process of shaping the data into the format your algorithm wants. This is not just to be mean, but because you might benefit from catching other surprises lurking in the data. Who knows what is missing, duplicated, swapped, flipped, mirrored, etc.</p>\n\n<p>That said, we would not interpret fixing a few malformed directories as hand labeling. The spirit of that rule is to prevent an unfair scientific advantage, not to disqualify you because some outlier file format messed up one of your predictions.</p>",
      "rawMarkdown": "[quote=ZFTurbo;101978]\r\n\r\nTests with numer of images more or less than 30 in sax_* directories:\r\n\r\nTests TRAIN: 123, 234, 279, 334, 416, 463, 499\r\n\r\nTests VALID: 516, 590, 597, 619, 623\r\n\r\nI don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?\r\n\r\n[/quote]\r\n\r\nWe suggest you attempt to automate the process of shaping the data into the format your algorithm wants. This is not just to be mean, but because you might benefit from catching other surprises lurking in the data. Who knows what is missing, duplicated, swapped, flipped, mirrored, etc.\r\n\r\nThat said, we would not interpret fixing a few malformed directories as hand labeling. The spirit of that rule is to prevent an unfair scientific advantage, not to disqualify you because some outlier file format messed up one of your predictions.",
      "votes": null
    },
    {
      "id": "102030",
      "postDate": "12/18/2015 19:14:06",
      "content": "<p>Could we also get further clarification on the second question?</p>\n\n<blockquote>\n  <p>We have been told that images prefixed with 'sax_' are short axis cine\n  images, but what do the suffixes to each of the directories mean? For\n  example with 'train/1/study/sax_5', what does the '5' indicate?</p>\n</blockquote>",
      "rawMarkdown": "Could we also get further clarification on the second question?\r\n\r\n> We have been told that images prefixed with 'sax_' are short axis cine\r\n> images, but what do the suffixes to each of the directories mean? For\r\n> example with 'train/1/study/sax_5', what does the '5' indicate?",
      "votes": null
    },
    {
      "id": "102408",
      "postDate": "12/22/2015 11:15:39",
      "content": "<p>Does anybody know the answer to the second question posted by @megaminion?</p>",
      "rawMarkdown": "Does anybody know the answer to the second question posted by @megaminion?",
      "votes": null
    },
    {
      "id": "102639",
      "postDate": "12/24/2015 03:36:58",
      "content": "<p>[quote=Drew Farris;101957]</p>\n\n<p>Thanks for mentioning this. Here is what I did, maybe helpful to others:</p>\n\n<pre><code>#!/bin/sh\nfor image in `find data/train -regex '^.*/study/sax_[0-9]*/IM\\-[0-9]*\\-[0-9]*\\-[0-9]*\\..*'`; do\n   image2=`echo $image | sed 's/\\([0-9]*\\)\\/\\(IM-.*\\)\\-\\([0-9]*\\)/\\1_\\3\\/\\2/g'`\n   dir=$(dirname &quot;${image2}&quot;)\n   mkdir $dir\n   cp $image $image2\ndone\n</code></pre>\n\n<p>Convert <em>train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm</em>  to <em>train/463/study/sax_9_MMMM/IM-2776-NNNN.dcm</em></p>\n\n<p>megaminion asked:</p>\n\n<blockquote>\n  <p>For the images in each of these directories, are we guaranteed to have 30 images?</p>\n</blockquote>\n\n<p>That doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:</p>\n\n<ul>\n<li>(study count): (number of images)</li>\n<li>6274:  30</li>\n<li>26: 60</li>\n<li>9: 25</li>\n<li>3: 330</li>\n<li>2: 270</li>\n<li>2: 21</li>\n<li>1: 23</li>\n<li>1: 22</li>\n</ul>\n\n<p>Spot checking a couple directories with 30 images we have filenames like:\n<em>train/17/study/2ch_23/IM-5996-NNNN.dcm</em> - Where NNNN runs from 0000-0030</p>\n\n<p>The 60 image directories have filenames like:\n<em>train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm</em> - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002</p>\n\n<p>And the 330 image directory seems to follow this pattern:\n<em>train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm</em> - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011</p>\n\n<p>Without looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:</p>\n\n<blockquote>\n  <p>Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.</p>\n</blockquote>\n\n<p>[/quote]</p>",
      "rawMarkdown": "[quote=Drew Farris;101957]\r\n\r\nThanks for mentioning this. Here is what I did, maybe helpful to others:\r\n\r\n    #!/bin/sh\r\n    for image in `find data/train -regex '^.*/study/sax_[0-9]*/IM\\-[0-9]*\\-[0-9]*\\-[0-9]*\\..*'`; do\r\n       image2=`echo $image | sed 's/\\([0-9]*\\)\\/\\(IM-.*\\)\\-\\([0-9]*\\)/\\1_\\3\\/\\2/g'`\r\n   dir=$(dirname \"${image2}\")\r\n   mkdir $dir\r\n       cp $image $image2\r\n    done\r\n\r\n\r\nConvert *train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm*  to *train/463/study/sax_9_MMMM/IM-2776-NNNN.dcm*\r\n\r\nmegaminion asked:\r\n\r\n> For the images in each of these directories, are we guaranteed to have 30 images?\r\n\r\nThat doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:\r\n\r\n- (study count): (number of images)\r\n- 6274:  30\r\n- 26: 60\r\n- 9: 25\r\n- 3: 330\r\n- 2: 270\r\n- 2: 21\r\n- 1: 23\r\n- 1: 22\r\n\r\n\r\nSpot checking a couple directories with 30 images we have filenames like:\r\n*train/17/study/2ch_23/IM-5996-NNNN.dcm* - Where NNNN runs from 0000-0030\r\n\r\nThe 60 image directories have filenames like:\r\n*train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm* - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002\r\n\r\nAnd the 330 image directory seems to follow this pattern:\r\n*train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm* - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011\r\n\r\nWithout looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:\r\n\r\n> Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.\r\n\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "103129",
      "postDate": "12/29/2015 06:55:32",
      "content": "<p>@megaminion see the section on loading datasets at link below.</p>\n\n<p><a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/fourier-based-tutorial\">https://www.kaggle.com/c/second-annual-data-science-bowl/details/fourier-based-tutorial</a></p>",
      "rawMarkdown": "megaminion see the section on loading datasets at link below.\r\n\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/details/fourier-based-tutorial",
      "votes": null
    },
    {
      "id": "1078310",
      "postDate": "11/14/2020 16:07:13",
      "content": "<p>Thanks for the great content. I will also share with my friends &amp; once again Thanks a Lot<br>\n<a href=\"https://trimsnest.com/digitization/\">digitization</a></p>",
      "rawMarkdown": "Thanks for the great content. I will also share with my friends & once again Thanks a Lot\n<a href=\"https://trimsnest.com/digitization/\">digitization</a>",
      "votes": null
    },
    {
      "id": "1078314",
      "postDate": "11/14/2020 16:13:25",
      "content": "<p>Thanks for the great content. I will also share with my friends &amp; once again Thanks a Lot<br>\n<a href=\"https://trimsnest.com/digitization/\">digitization</a></p>",
      "rawMarkdown": "Thanks for the great content. I will also share with my friends & once again Thanks a Lot\n<a href=\"https://trimsnest.com/digitization/\">digitization</a>",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1078310,
      "author_name": "anthonysusai",
      "author_url": "",
      "post_date": "11/14/2020 16:07:13",
      "content": "<p>Thanks for the great content. I will also share with my friends &amp; once again Thanks a Lot<br>\n<a href=\"https://trimsnest.com/digitization/\">digitization</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1078314,
      "author_name": "anthonysusai",
      "author_url": "",
      "post_date": "11/14/2020 16:13:25",
      "content": "<p>Thanks for the great content. I will also share with my friends &amp; once again Thanks a Lot<br>\n<a href=\"https://trimsnest.com/digitization/\">digitization</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101913,
      "author_name": "operdeck",
      "author_url": "",
      "post_date": "12/17/2015 20:44:54",
      "content": "<p>2ch and 4ch are the 2- and 4-chamber views.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101957,
      "author_name": "drewfarris",
      "author_url": "",
      "post_date": "12/18/2015 03:13:42",
      "content": "<p>megaminion asked:</p>\n\n<blockquote>\n  <p>For the images in each of these directories, are we guaranteed to have 30 images?</p>\n</blockquote>\n\n<p>That doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:</p>\n\n<ul>\n<li>(study count): (number of images)</li>\n<li>6274:  30</li>\n<li>26: 60</li>\n<li>9: 25</li>\n<li>3: 330</li>\n<li>2: 270</li>\n<li>2: 21</li>\n<li>1: 23</li>\n<li>1: 22</li>\n</ul>\n\n<p>Spot checking a couple directories with 30 images we have filenames like:\n<em>train/17/study/2ch_23/IM-5996-NNNN.dcm</em> - Where NNNN runs from 0000-0030</p>\n\n<p>The 60 image directories have filenames like:\n<em>train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm</em> - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002</p>\n\n<p>And the 330 image directory seems to follow this pattern:\n<em>train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm</em> - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011</p>\n\n<p>Without looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:</p>\n\n<blockquote>\n  <p>Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101978,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "12/18/2015 08:20:59",
      "content": "<p>Tests with numer of images more or less than 30 in sax_* directories:</p>\n\n<p>Tests TRAIN: 123, 234, 279, 334, 416, 463, 499</p>\n\n<p>Tests VALID: 516, 590, 597, 619, 623</p>\n\n<p>I don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102006,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "12/18/2015 14:41:23",
      "content": "<p>[quote=ZFTurbo;101978]</p>\n\n<p>Tests with numer of images more or less than 30 in sax_* directories:</p>\n\n<p>Tests TRAIN: 123, 234, 279, 334, 416, 463, 499</p>\n\n<p>Tests VALID: 516, 590, 597, 619, 623</p>\n\n<p>I don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?</p>\n\n<p>[/quote]</p>\n\n<p>We suggest you attempt to automate the process of shaping the data into the format your algorithm wants. This is not just to be mean, but because you might benefit from catching other surprises lurking in the data. Who knows what is missing, duplicated, swapped, flipped, mirrored, etc.</p>\n\n<p>That said, we would not interpret fixing a few malformed directories as hand labeling. The spirit of that rule is to prevent an unfair scientific advantage, not to disqualify you because some outlier file format messed up one of your predictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102030,
      "author_name": "maxwang7",
      "author_url": "",
      "post_date": "12/18/2015 19:14:06",
      "content": "<p>Could we also get further clarification on the second question?</p>\n\n<blockquote>\n  <p>We have been told that images prefixed with 'sax_' are short axis cine\n  images, but what do the suffixes to each of the directories mean? For\n  example with 'train/1/study/sax_5', what does the '5' indicate?</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102408,
      "author_name": "francbracun",
      "author_url": "",
      "post_date": "12/22/2015 11:15:39",
      "content": "<p>Does anybody know the answer to the second question posted by @megaminion?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102639,
      "author_name": "amazingbob",
      "author_url": "",
      "post_date": "12/24/2015 03:36:58",
      "content": "<p>[quote=Drew Farris;101957]</p>\n\n<p>Thanks for mentioning this. Here is what I did, maybe helpful to others:</p>\n\n<pre><code>#!/bin/sh\nfor image in `find data/train -regex '^.*/study/sax_[0-9]*/IM\\-[0-9]*\\-[0-9]*\\-[0-9]*\\..*'`; do\n   image2=`echo $image | sed 's/\\([0-9]*\\)\\/\\(IM-.*\\)\\-\\([0-9]*\\)/\\1_\\3\\/\\2/g'`\n   dir=$(dirname &quot;${image2}&quot;)\n   mkdir $dir\n   cp $image $image2\ndone\n</code></pre>\n\n<p>Convert <em>train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm</em>  to <em>train/463/study/sax_9_MMMM/IM-2776-NNNN.dcm</em></p>\n\n<p>megaminion asked:</p>\n\n<blockquote>\n  <p>For the images in each of these directories, are we guaranteed to have 30 images?</p>\n</blockquote>\n\n<p>That doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:</p>\n\n<ul>\n<li>(study count): (number of images)</li>\n<li>6274:  30</li>\n<li>26: 60</li>\n<li>9: 25</li>\n<li>3: 330</li>\n<li>2: 270</li>\n<li>2: 21</li>\n<li>1: 23</li>\n<li>1: 22</li>\n</ul>\n\n<p>Spot checking a couple directories with 30 images we have filenames like:\n<em>train/17/study/2ch_23/IM-5996-NNNN.dcm</em> - Where NNNN runs from 0000-0030</p>\n\n<p>The 60 image directories have filenames like:\n<em>train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm</em> - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002</p>\n\n<p>And the 330 image directory seems to follow this pattern:\n<em>train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm</em> - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011</p>\n\n<p>Without looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:</p>\n\n<blockquote>\n  <p>Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.</p>\n</blockquote>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103129,
      "author_name": "lbollar",
      "author_url": "",
      "post_date": "12/29/2015 06:55:32",
      "content": "<p>@megaminion see the section on loading datasets at link below.</p>\n\n<p><a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/fourier-based-tutorial\">https://www.kaggle.com/c/second-annual-data-science-bowl/details/fourier-based-tutorial</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "101673": "I would appreciate some clarification on the naming of the directories/images we are given. Once I've extracted train.zip, I understand that each individual sample is located in its own directory (1, 2, ..., 500). Inside the 'study' directory of each of these samples there are several directories, I'm assuming from different angles/perspectives. My specific questions are:\r\n\r\n- What does #ch_# mean? For example with 'train/1/study/2ch_21', what do the '2', 'ch', and '21' indicate?\r\n- We have been told that images prefixed with 'sax_' are short axis cine images, but what do the suffixes to each of the directories mean? For example with 'train/1/study/sax_5', what does the '5' indicate?\r\n- For the images in each of these directories, are we guaranteed to have 30 images? And do these 30 images always represent one heart beat or may they represent many heart beats? For example, 'train/1/study/sax_5/' has 30 images -- do all similar directories have 30 images and are they over an entire heart beat?",
    "101913": "2ch and 4ch are the 2- and 4-chamber views.",
    "101957": "megaminion asked:\r\n\r\n> For the images in each of these directories, are we guaranteed to have 30 images?\r\n\r\nThat doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:\r\n\r\n- (study count): (number of images)\r\n- 6274:  30\r\n- 26: 60\r\n- 9: 25\r\n- 3: 330\r\n- 2: 270\r\n- 2: 21\r\n- 1: 23\r\n- 1: 22\r\n\r\n\r\nSpot checking a couple directories with 30 images we have filenames like:\r\n*train/17/study/2ch_23/IM-5996-NNNN.dcm* - Where NNNN runs from 0000-0030\r\n\r\nThe 60 image directories have filenames like:\r\n*train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm* - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002\r\n\r\nAnd the 330 image directory seems to follow this pattern:\r\n*train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm* - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011\r\n\r\nWithout looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:\r\n\r\n> Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.",
    "101978": "Tests with numer of images more or less than 30 in sax_* directories:\r\n\r\nTests TRAIN: 123, 234, 279, 334, 416, 463, 499\r\n\r\nTests VALID: 516, 590, 597, 619, 623\r\n\r\nI don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?",
    "102006": "[quote=ZFTurbo;101978]\r\n\r\nTests with numer of images more or less than 30 in sax_* directories:\r\n\r\nTests TRAIN: 123, 234, 279, 334, 416, 463, 499\r\n\r\nTests VALID: 516, 590, 597, 619, 623\r\n\r\nI don't think this was intentional. Since for case where 270 images in same dir - it's actually must be 9 different sax dir in my opinion. Is it violation of rules to split this dirs by hand (including VALID case)?\r\n\r\n[/quote]\r\n\r\nWe suggest you attempt to automate the process of shaping the data into the format your algorithm wants. This is not just to be mean, but because you might benefit from catching other surprises lurking in the data. Who knows what is missing, duplicated, swapped, flipped, mirrored, etc.\r\n\r\nThat said, we would not interpret fixing a few malformed directories as hand labeling. The spirit of that rule is to prevent an unfair scientific advantage, not to disqualify you because some outlier file format messed up one of your predictions.",
    "102030": "Could we also get further clarification on the second question?\r\n\r\n> We have been told that images prefixed with 'sax_' are short axis cine\r\n> images, but what do the suffixes to each of the directories mean? For\r\n> example with 'train/1/study/sax_5', what does the '5' indicate?",
    "102408": "Does anybody know the answer to the second question posted by @megaminion?",
    "102639": "[quote=Drew Farris;101957]\r\n\r\nThanks for mentioning this. Here is what I did, maybe helpful to others:\r\n\r\n    #!/bin/sh\r\n    for image in `find data/train -regex '^.*/study/sax_[0-9]*/IM\\-[0-9]*\\-[0-9]*\\-[0-9]*\\..*'`; do\r\n       image2=`echo $image | sed 's/\\([0-9]*\\)\\/\\(IM-.*\\)\\-\\([0-9]*\\)/\\1_\\3\\/\\2/g'`\r\n   dir=$(dirname \"${image2}\")\r\n   mkdir $dir\r\n       cp $image $image2\r\n    done\r\n\r\n\r\nConvert *train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm*  to *train/463/study/sax_9_MMMM/IM-2776-NNNN.dcm*\r\n\r\nmegaminion asked:\r\n\r\n> For the images in each of these directories, are we guaranteed to have 30 images?\r\n\r\nThat doesn't appear to be the case - looking at the table of contents for train.zip, it appears that train/XXX/study/YYY directories have anywhere between 330 and 30 images:\r\n\r\n- (study count): (number of images)\r\n- 6274:  30\r\n- 26: 60\r\n- 9: 25\r\n- 3: 330\r\n- 2: 270\r\n- 2: 21\r\n- 1: 23\r\n- 1: 22\r\n\r\n\r\nSpot checking a couple directories with 30 images we have filenames like:\r\n*train/17/study/2ch_23/IM-5996-NNNN.dcm* - Where NNNN runs from 0000-0030\r\n\r\nThe 60 image directories have filenames like:\r\n*train/463/study/sax_9/IM-2776-NNNN-MMMM.dcm* - Where NNNN also runs from 0000-0030 but MMMM seems to be 0001 or 0002\r\n\r\nAnd the 330 image directory seems to follow this pattern:\r\n*train/334/study/sax_29/IM-2360-NNNN-MMMM.dcm* - Where NNNN runs from 0000-00030 but MMMM runs from 00001 to 0011\r\n\r\nWithout looking at the images themselves, I can hypothesize that MMMM represent different series at the same slice location as referenced in the FAQ:\r\n\r\n> Generally, a slice location is repeated if there is an artifact on the images. You can use either slice but the odds are that the last slice at a given slice location is the best the technologist could acquire.\r\n\r\n\r\n[/quote]",
    "103129": "megaminion see the section on loading datasets at link below.\r\n\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/details/fourier-based-tutorial",
    "1078310": "Thanks for the great content. I will also share with my friends & once again Thanks a Lot\n<a href=\"https://trimsnest.com/digitization/\">digitization</a>",
    "1078314": "Thanks for the great content. I will also share with my friends & once again Thanks a Lot\n<a href=\"https://trimsnest.com/digitization/\">digitization</a>"
  },
  "source": "meta"
}