{
  "id": 18846,
  "title": "Possible problem with training volumes?",
  "url": "/competitions/second-annual-data-science-bowl/discussion/18846",
  "author_name": "",
  "post_date": "2016-02-09T13:56:01.277Z",
  "votes": 2,
  "comment_count": 14,
  "views": 1993,
  "content": "<p>In a <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18372/some-cases-are-quite-off-from-the-true-value\">thread some time ago</a>, @woshialex pointed out that the diastolic volume for study 429 seemed to be significantly larger than the volume given as ground truth. This was also supported by independent manual segmentation by @PaulG. </p>\n\n<p>@Julian de Wit noticed that one would get the correct volume if the volume calculation used SliceThickness (8mm) instead of Slice spacing extracted from SliceLocation data (10mm).</p>\n\n<p>Since then I noticed something similar in other cases. Here is an example of study 267 frame 29 (EDIT: last frame - files ending with 0030.dcm). I attach a screenshot with the slices and another one with the output of my segmentation algorithm. Below are the areas in pixels of the consecutive slices.</p>\n\n<pre><code>slice   pixels \n1        170\n2        428\n3        590\n4        768\n5        880\n6        977\n7       1081\n8       1101\n9       969\n10     862\n</code></pre>\n\n<p>The volume calculated using SliceLocation and the frustrum formula (from the Fourier tutorial) is 160.7 while the given value is 109.9. If I would calculate it using SLiceThickness I get 160.7*0.8=128.56 which is much closer - it is still too much but this is probably due to bad segmentation in the last slice. However I don't think that the error due to that could be so large as to explain the difference 160.7-109.9.</p>\n\n<p>I had similar suspicions about some other studies (with varying degrees of conviction).</p>\n\n<p>Is it really the case or I am doing something completely wrong with the segmentation? Could the organizers confirm (or deny) that some volumes were computed using SliceThickness instead of SliceLocation (possibly some program on a specific machine used such a default method?)? If it is so then is it possible to distinguish these cases using some DICOM metadata?</p>\n\n<p>If there are such erroneous data then this is a big problem for constructing learning algorithms of any kind...</p>\n\n<p><strong>EDIT:</strong> Sorry - I used my Python code numbering of frames starting from 0 - so these images correspond to files ending with 0030.dcm</p>",
  "messages": [
    {
      "id": "107375",
      "postDate": "02/09/2016 13:56:01",
      "content": "<p>In a <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18372/some-cases-are-quite-off-from-the-true-value\">thread some time ago</a>, @woshialex pointed out that the diastolic volume for study 429 seemed to be significantly larger than the volume given as ground truth. This was also supported by independent manual segmentation by @PaulG. </p>\n\n<p>@Julian de Wit noticed that one would get the correct volume if the volume calculation used SliceThickness (8mm) instead of Slice spacing extracted from SliceLocation data (10mm).</p>\n\n<p>Since then I noticed something similar in other cases. Here is an example of study 267 frame 29 (EDIT: last frame - files ending with 0030.dcm). I attach a screenshot with the slices and another one with the output of my segmentation algorithm. Below are the areas in pixels of the consecutive slices.</p>\n\n<pre><code>slice   pixels \n1        170\n2        428\n3        590\n4        768\n5        880\n6        977\n7       1081\n8       1101\n9       969\n10     862\n</code></pre>\n\n<p>The volume calculated using SliceLocation and the frustrum formula (from the Fourier tutorial) is 160.7 while the given value is 109.9. If I would calculate it using SLiceThickness I get 160.7*0.8=128.56 which is much closer - it is still too much but this is probably due to bad segmentation in the last slice. However I don't think that the error due to that could be so large as to explain the difference 160.7-109.9.</p>\n\n<p>I had similar suspicions about some other studies (with varying degrees of conviction).</p>\n\n<p>Is it really the case or I am doing something completely wrong with the segmentation? Could the organizers confirm (or deny) that some volumes were computed using SliceThickness instead of SliceLocation (possibly some program on a specific machine used such a default method?)? If it is so then is it possible to distinguish these cases using some DICOM metadata?</p>\n\n<p>If there are such erroneous data then this is a big problem for constructing learning algorithms of any kind...</p>\n\n<p><strong>EDIT:</strong> Sorry - I used my Python code numbering of frames starting from 0 - so these images correspond to files ending with 0030.dcm</p>",
      "rawMarkdown": "In a [thread some time ago][1], @woshialex pointed out that the diastolic volume for study 429 seemed to be significantly larger than the volume given as ground truth. This was also supported by independent manual segmentation by @PaulG. \r\n\r\n@Julian de Wit noticed that one would get the correct volume if the volume calculation used SliceThickness (8mm) instead of Slice spacing extracted from SliceLocation data (10mm).\r\n\r\nSince then I noticed something similar in other cases. Here is an example of study 267 frame 29 (EDIT: last frame - files ending with 0030.dcm). I attach a screenshot with the slices and another one with the output of my segmentation algorithm. Below are the areas in pixels of the consecutive slices.\r\n\r\n    slice   pixels \r\n    1        170\r\n    2        428\r\n    3        590\r\n    4        768\r\n    5        880\r\n    6        977\r\n    7       1081\r\n    8       1101\r\n    9       969\r\n    10     862\r\n\r\n\r\nThe volume calculated using SliceLocation and the frustrum formula (from the Fourier tutorial) is 160.7 while the given value is 109.9. If I would calculate it using SLiceThickness I get 160.7*0.8=128.56 which is much closer - it is still too much but this is probably due to bad segmentation in the last slice. However I don't think that the error due to that could be so large as to explain the difference 160.7-109.9.\r\n\r\nI had similar suspicions about some other studies (with varying degrees of conviction).\r\n\r\nIs it really the case or I am doing something completely wrong with the segmentation? Could the organizers confirm (or deny) that some volumes were computed using SliceThickness instead of SliceLocation (possibly some program on a specific machine used such a default method?)? If it is so then is it possible to distinguish these cases using some DICOM metadata?\r\n\r\nIf there are such erroneous data then this is a big problem for constructing learning algorithms of any kind...\r\n\r\n**EDIT:** Sorry - I used my Python code numbering of frames starting from 0 - so these images correspond to files ending with 0030.dcm\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18372/some-cases-are-quite-off-from-the-true-value",
      "votes": null
    },
    {
      "id": "107380",
      "postDate": "02/09/2016 14:33:06",
      "content": "<p>Hi rmldj,</p>\n\n<p>As a Kaggle Master you probably know what I'm about to say, but it bears repeating since we get a lot of questions on data integrity.</p>\n\n<p>I can't help you on the SliceThickness vs. SliceLocation question, but you should expect real data to have issues like this, often without an explanation (see <a href=\"https://www.kaggle.com/wiki/ANoteOnDataQuality\">https://www.kaggle.com/wiki/ANoteOnDataQuality</a>). These might stem from clerical typos, a tired doc, a confused trainee, differing protocols, a wrong assumption on a database join, you name it.</p>\n\n<p>The &quot;big problem&quot; for us occurs when the noise is systematic and widespread, else it's on the algorithms to be robust.</p>",
      "rawMarkdown": "Hi rmldj,\r\n\r\nAs a Kaggle Master you probably know what I'm about to say, but it bears repeating since we get a lot of questions on data integrity.\r\n\r\nI can't help you on the SliceThickness vs. SliceLocation question, but you should expect real data to have issues like this, often without an explanation (see https://www.kaggle.com/wiki/ANoteOnDataQuality). These might stem from clerical typos, a tired doc, a confused trainee, differing protocols, a wrong assumption on a database join, you name it.\r\n\r\nThe \"big problem\" for us occurs when the noise is systematic and widespread, else it's on the algorithms to be robust.",
      "votes": null
    },
    {
      "id": "107489",
      "postDate": "02/10/2016 11:12:06",
      "content": "<p>@rmldj</p>\n\n<p>What are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.</p>",
      "rawMarkdown": "rmldj\r\n\r\nWhat are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.",
      "votes": null
    },
    {
      "id": "107491",
      "postDate": "02/10/2016 11:29:47",
      "content": "<p>[quote=Paul Jurczak;107489]</p>\n\n<p>@rmldj</p>\n\n<p>What are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.</p>\n\n<p>[/quote]</p>\n\n<p>These are those files but in reverse order (from sax_15 to sax_6) - I order them according to increasing Slice Location... </p>",
      "rawMarkdown": "[quote=Paul Jurczak;107489]\r\n\r\n@rmldj\r\n\r\nWhat are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.\r\n\r\n[/quote]\r\n\r\nThese are those files but in reverse order (from sax_15 to sax_6) - I order them according to increasing Slice Location...",
      "votes": null
    },
    {
      "id": "107492",
      "postDate": "02/10/2016 11:55:58",
      "content": "<p>Sorry - I number frames from 0 (for easy iteration in Python) - so these files correspond to files ending with 0030.dcm! I will edit the topic to make it clear...</p>",
      "rawMarkdown": "Sorry - I number frames from 0 (for easy iteration in Python) - so these files correspond to files ending with 0030.dcm! I will edit the topic to make it clear...",
      "votes": null
    },
    {
      "id": "107493",
      "postDate": "02/10/2016 12:05:58",
      "content": "<p>The pixel counts I'm getting are (slice position, pixel count for image #29):</p>\n\n<p>-111.3, 134</p>\n\n<p>-101.3, 320</p>\n\n<p>-91.3,  420</p>\n\n<p>-81.3,  658</p>\n\n<p>-71.3,  613</p>\n\n<p>-61.3,  916</p>\n\n<p>-51.3,  971</p>\n\n<p>-41.3,  900</p>\n\n<p>-31.3,  888</p>\n\n<p>-21.3,  400</p>\n\n<p>which produces a volume of about 137 ml. My segmentation usually produces a bit smaller area than I would do manually, so I agree with your suspicion of large error in study #267.</p>\n\n<p>BTW, your segmentation of slice -21.3 is much too large (see attached image).</p>",
      "rawMarkdown": "The pixel counts I'm getting are (slice position, pixel count for image #29):\r\n\r\n-111.3,\t134\r\n\r\n-101.3,\t320\r\n\r\n-91.3,\t420\r\n\r\n-81.3,\t658\r\n\r\n-71.3,\t613\r\n\r\n-61.3,\t916\r\n\r\n-51.3,\t971\r\n\r\n-41.3,\t900\r\n\r\n-31.3,\t888\r\n\r\n-21.3,\t400\r\n\r\nwhich produces a volume of about 137 ml. My segmentation usually produces a bit smaller area than I would do manually, so I agree with your suspicion of large error in study #267.\r\n\r\nBTW, your segmentation of slice -21.3 is much too large (see attached image).",
      "votes": null
    },
    {
      "id": "107494",
      "postDate": "02/10/2016 12:10:44",
      "content": "<blockquote>\n  <p><em>I number frames from 0</em></p>\n</blockquote>\n\n<p>I was going to add that, at least for some slices, the frame #30 has slightly higher pixel count, which illustrates your point even better.</p>",
      "rawMarkdown": ">  *I number frames from 0*\r\n\r\nI was going to add that, at least for some slices, the frame #30 has slightly higher pixel count, which illustrates your point even better.",
      "votes": null
    },
    {
      "id": "107496",
      "postDate": "02/10/2016 12:20:30",
      "content": "<p>Thanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.</p>\n\n<p>However concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand? </p>",
      "rawMarkdown": "Thanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.\r\n\r\nHowever concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand?",
      "votes": null
    },
    {
      "id": "107499",
      "postDate": "02/10/2016 12:36:19",
      "content": "<p>[quote=rmldj;107496]</p>\n\n<p>Thanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.</p>\n\n<p>However concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand? </p>\n\n<p>[/quote]</p>\n\n<p>I didn't mean that image dimensions were larger, just the pixel count for LV segment was slightly larger. That's why you've chosen frames #30 to measure LV diastole end volume instead of other frames, right?</p>",
      "rawMarkdown": "[quote=rmldj;107496]\r\n\r\nThanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.\r\n\r\nHowever concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand? \r\n\r\n[/quote]\r\n\r\nI didn't mean that image dimensions were larger, just the pixel count for LV segment was slightly larger. That's why you've chosen frames #30 to measure LV diastole end volume instead of other frames, right?",
      "votes": null
    },
    {
      "id": "107500",
      "postDate": "02/10/2016 12:40:14",
      "content": "<p>Ok. I understand what you meant - indeed in this study frame #30 was largest..</p>",
      "rawMarkdown": "Ok. I understand what you meant - indeed in this study frame #30 was largest..",
      "votes": null
    },
    {
      "id": "108083",
      "postDate": "02/15/2016 12:40:12",
      "content": "<p>@William Cukierski <br>\nA 60ML error in ground truth might mean a difference of 0.0005+ on the LB depending on the # of patients if my calculations are correct. If someone took a gamble (which you could do for cases like 429) you could improve the LB score significanty with luck without having a better solution.. Personally I think this is a but too much for a 100.000+ competition.\n<br>\n<br>\nNext to this.. If we want to boast clinical significance of computerized volume measurement it would be good to have a clinical significant test-test to begin with. I'm pretty sure we can get at ~10ml deviations but cases like 429 spoil te party a bit..\n<br>\n<br>\nWouldn't it be possible to do a double check on the final test-set ?</p>",
      "rawMarkdown": "William Cukierski <br>\r\nA 60ML error in ground truth might mean a difference of 0.0005+ on the LB depending on the # of patients if my calculations are correct. If someone took a gamble (which you could do for cases like 429) you could improve the LB score significanty with luck without having a better solution.. Personally I think this is a but too much for a 100.000+ competition.\r\n<br>\r\n<br>\r\nNext to this.. If we want to boast clinical significance of computerized volume measurement it would be good to have a clinical significant test-test to begin with. I'm pretty sure we can get at ~10ml deviations but cases like 429 spoil te party a bit..\r\n<br>\r\n<br>\r\nWouldn't it be possible to do a double check on the final test-set ?",
      "votes": null
    },
    {
      "id": "108290",
      "postDate": "02/16/2016 15:45:51",
      "content": "<p>@Julian Even with perfect data competitions would still have an element of luck. It's the nature of having a large number of teams and a finite test set (one could take bootstrap subsets of the test set and, depending on how close the scores are, have a different leaderboard each time). On the plus side, noise for you is noise for everyone else.</p>\n\n<p>Why can't we fix noise in this specific case? Primarily, it's not economically reasonable to ask cardiologists to re-read the whole dataset to correct a few potential outliers. Even if the time and budget was there, I'm unconvinced that such a costly effort meaningfully changes the contributions of this project towards computerized volume measurement. We have to accept that &quot;the perfect is the enemy of the good&quot; and acknowledge the amount of &quot;good&quot; behind the data science bowl: a philanthropic commitment and involvement from the team at Booz + a forward-thinking research team at the NIH willing to put a large imaging data out there + a community of people like you who are capable of pushing the state of the art.</p>\n\n<p>Maybe this answer isn't the most satisfying, but I think it's important not to focus on small flaws at the expense of the larger opportunity.</p>",
      "rawMarkdown": "Julian Even with perfect data competitions would still have an element of luck. It's the nature of having a large number of teams and a finite test set (one could take bootstrap subsets of the test set and, depending on how close the scores are, have a different leaderboard each time). On the plus side, noise for you is noise for everyone else.\r\n\r\nWhy can't we fix noise in this specific case? Primarily, it's not economically reasonable to ask cardiologists to re-read the whole dataset to correct a few potential outliers. Even if the time and budget was there, I'm unconvinced that such a costly effort meaningfully changes the contributions of this project towards computerized volume measurement. We have to accept that \"the perfect is the enemy of the good\" and acknowledge the amount of \"good\" behind the data science bowl: a philanthropic commitment and involvement from the team at Booz + a forward-thinking research team at the NIH willing to put a large imaging data out there + a community of people like you who are capable of pushing the state of the art.\r\n\r\nMaybe this answer isn't the most satisfying, but I think it's important not to focus on small flaws at the expense of the larger opportunity.",
      "votes": null
    },
    {
      "id": "108299",
      "postDate": "02/16/2016 16:24:22",
      "content": "<p>Could it be possible for the organizers to tell us how is the measured volume for case 595 and 599 (which are in validation set so we can't check) calculated because they only have 3 slices and it apparently did not span the whole range of the heart, did the organizers use some method other than contour--&gt;area*thickness method for these cases because it will underestimate the true value significantly (but the provided number could have been calculated this way). We're building models to predict the provided numbers, not the true heart volume, so it affects how will we deal with cases with 3 slices only....</p>",
      "rawMarkdown": "Could it be possible for the organizers to tell us how is the measured volume for case 595 and 599 (which are in validation set so we can't check) calculated because they only have 3 slices and it apparently did not span the whole range of the heart, did the organizers use some method other than contour-->area*thickness method for these cases because it will underestimate the true value significantly (but the provided number could have been calculated this way). We're building models to predict the provided numbers, not the true heart volume, so it affects how will we deal with cases with 3 slices only....",
      "votes": null
    },
    {
      "id": "108381",
      "postDate": "02/17/2016 00:06:35",
      "content": "<p>@woshialex</p>\n\n<p>Good question! Technically, study #595 can be handled by extrapolating SAX slice volumes using the &quot;side&quot; view of LV from 2ch_6, but this is very hard to automatize. I'm planning to have a fallback option for incomplete data cases like this, where I would use only midsection of LV for volume prediction. There should be a strong correlation between LV midsection slice surface area and LV volume.</p>",
      "rawMarkdown": "woshialex\r\n\r\nGood question! Technically, study #595 can be handled by extrapolating SAX slice volumes using the \"side\" view of LV from 2ch_6, but this is very hard to automatize. I'm planning to have a fallback option for incomplete data cases like this, where I would use only midsection of LV for volume prediction. There should be a strong correlation between LV midsection slice surface area and LV volume.",
      "votes": null
    },
    {
      "id": "108820",
      "postDate": "02/20/2016 07:31:48",
      "content": "<p>@William Cukierski,<br>\nI see your point..</p>\n\n<p>First and foremost it's awesome that we are given the opportunity to solve these cool problems.<br>\nThanks for that.</p>",
      "rawMarkdown": "William Cukierski,<br>\r\nI see your point..\r\n\r\nFirst and foremost it's awesome that we are given the opportunity to solve these cool problems.<br>\r\nThanks for that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 107380,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "02/09/2016 14:33:06",
      "content": "<p>Hi rmldj,</p>\n\n<p>As a Kaggle Master you probably know what I'm about to say, but it bears repeating since we get a lot of questions on data integrity.</p>\n\n<p>I can't help you on the SliceThickness vs. SliceLocation question, but you should expect real data to have issues like this, often without an explanation (see <a href=\"https://www.kaggle.com/wiki/ANoteOnDataQuality\">https://www.kaggle.com/wiki/ANoteOnDataQuality</a>). These might stem from clerical typos, a tired doc, a confused trainee, differing protocols, a wrong assumption on a database join, you name it.</p>\n\n<p>The &quot;big problem&quot; for us occurs when the noise is systematic and widespread, else it's on the algorithms to be robust.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107489,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/10/2016 11:12:06",
      "content": "<p>@rmldj</p>\n\n<p>What are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107491,
      "author_name": "romualdj",
      "author_url": "",
      "post_date": "02/10/2016 11:29:47",
      "content": "<p>[quote=Paul Jurczak;107489]</p>\n\n<p>@rmldj</p>\n\n<p>What are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.</p>\n\n<p>[/quote]</p>\n\n<p>These are those files but in reverse order (from sax_15 to sax_6) - I order them according to increasing Slice Location... </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107492,
      "author_name": "romualdj",
      "author_url": "",
      "post_date": "02/10/2016 11:55:58",
      "content": "<p>Sorry - I number frames from 0 (for easy iteration in Python) - so these files correspond to files ending with 0030.dcm! I will edit the topic to make it clear...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107493,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/10/2016 12:05:58",
      "content": "<p>The pixel counts I'm getting are (slice position, pixel count for image #29):</p>\n\n<p>-111.3, 134</p>\n\n<p>-101.3, 320</p>\n\n<p>-91.3,  420</p>\n\n<p>-81.3,  658</p>\n\n<p>-71.3,  613</p>\n\n<p>-61.3,  916</p>\n\n<p>-51.3,  971</p>\n\n<p>-41.3,  900</p>\n\n<p>-31.3,  888</p>\n\n<p>-21.3,  400</p>\n\n<p>which produces a volume of about 137 ml. My segmentation usually produces a bit smaller area than I would do manually, so I agree with your suspicion of large error in study #267.</p>\n\n<p>BTW, your segmentation of slice -21.3 is much too large (see attached image).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107494,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/10/2016 12:10:44",
      "content": "<blockquote>\n  <p><em>I number frames from 0</em></p>\n</blockquote>\n\n<p>I was going to add that, at least for some slices, the frame #30 has slightly higher pixel count, which illustrates your point even better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107496,
      "author_name": "romualdj",
      "author_url": "",
      "post_date": "02/10/2016 12:20:30",
      "content": "<p>Thanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.</p>\n\n<p>However concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107499,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/10/2016 12:36:19",
      "content": "<p>[quote=rmldj;107496]</p>\n\n<p>Thanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.</p>\n\n<p>However concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand? </p>\n\n<p>[/quote]</p>\n\n<p>I didn't mean that image dimensions were larger, just the pixel count for LV segment was slightly larger. That's why you've chosen frames #30 to measure LV diastole end volume instead of other frames, right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 107500,
      "author_name": "romualdj",
      "author_url": "",
      "post_date": "02/10/2016 12:40:14",
      "content": "<p>Ok. I understand what you meant - indeed in this study frame #30 was largest..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108083,
      "author_name": "juliandewit",
      "author_url": "",
      "post_date": "02/15/2016 12:40:12",
      "content": "<p>@William Cukierski <br>\nA 60ML error in ground truth might mean a difference of 0.0005+ on the LB depending on the # of patients if my calculations are correct. If someone took a gamble (which you could do for cases like 429) you could improve the LB score significanty with luck without having a better solution.. Personally I think this is a but too much for a 100.000+ competition.\n<br>\n<br>\nNext to this.. If we want to boast clinical significance of computerized volume measurement it would be good to have a clinical significant test-test to begin with. I'm pretty sure we can get at ~10ml deviations but cases like 429 spoil te party a bit..\n<br>\n<br>\nWouldn't it be possible to do a double check on the final test-set ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108290,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "02/16/2016 15:45:51",
      "content": "<p>@Julian Even with perfect data competitions would still have an element of luck. It's the nature of having a large number of teams and a finite test set (one could take bootstrap subsets of the test set and, depending on how close the scores are, have a different leaderboard each time). On the plus side, noise for you is noise for everyone else.</p>\n\n<p>Why can't we fix noise in this specific case? Primarily, it's not economically reasonable to ask cardiologists to re-read the whole dataset to correct a few potential outliers. Even if the time and budget was there, I'm unconvinced that such a costly effort meaningfully changes the contributions of this project towards computerized volume measurement. We have to accept that &quot;the perfect is the enemy of the good&quot; and acknowledge the amount of &quot;good&quot; behind the data science bowl: a philanthropic commitment and involvement from the team at Booz + a forward-thinking research team at the NIH willing to put a large imaging data out there + a community of people like you who are capable of pushing the state of the art.</p>\n\n<p>Maybe this answer isn't the most satisfying, but I think it's important not to focus on small flaws at the expense of the larger opportunity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108299,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "02/16/2016 16:24:22",
      "content": "<p>Could it be possible for the organizers to tell us how is the measured volume for case 595 and 599 (which are in validation set so we can't check) calculated because they only have 3 slices and it apparently did not span the whole range of the heart, did the organizers use some method other than contour--&gt;area*thickness method for these cases because it will underestimate the true value significantly (but the provided number could have been calculated this way). We're building models to predict the provided numbers, not the true heart volume, so it affects how will we deal with cases with 3 slices only....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108381,
      "author_name": "pauljurczak",
      "author_url": "",
      "post_date": "02/17/2016 00:06:35",
      "content": "<p>@woshialex</p>\n\n<p>Good question! Technically, study #595 can be handled by extrapolating SAX slice volumes using the &quot;side&quot; view of LV from 2ch_6, but this is very hard to automatize. I'm planning to have a fallback option for incomplete data cases like this, where I would use only midsection of LV for volume prediction. There should be a strong correlation between LV midsection slice surface area and LV volume.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108820,
      "author_name": "juliandewit",
      "author_url": "",
      "post_date": "02/20/2016 07:31:48",
      "content": "<p>@William Cukierski,<br>\nI see your point..</p>\n\n<p>First and foremost it's awesome that we are given the opportunity to solve these cool problems.<br>\nThanks for that.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "107375": "In a [thread some time ago][1], @woshialex pointed out that the diastolic volume for study 429 seemed to be significantly larger than the volume given as ground truth. This was also supported by independent manual segmentation by @PaulG. \r\n\r\n@Julian de Wit noticed that one would get the correct volume if the volume calculation used SliceThickness (8mm) instead of Slice spacing extracted from SliceLocation data (10mm).\r\n\r\nSince then I noticed something similar in other cases. Here is an example of study 267 frame 29 (EDIT: last frame - files ending with 0030.dcm). I attach a screenshot with the slices and another one with the output of my segmentation algorithm. Below are the areas in pixels of the consecutive slices.\r\n\r\n    slice   pixels \r\n    1        170\r\n    2        428\r\n    3        590\r\n    4        768\r\n    5        880\r\n    6        977\r\n    7       1081\r\n    8       1101\r\n    9       969\r\n    10     862\r\n\r\n\r\nThe volume calculated using SliceLocation and the frustrum formula (from the Fourier tutorial) is 160.7 while the given value is 109.9. If I would calculate it using SLiceThickness I get 160.7*0.8=128.56 which is much closer - it is still too much but this is probably due to bad segmentation in the last slice. However I don't think that the error due to that could be so large as to explain the difference 160.7-109.9.\r\n\r\nI had similar suspicions about some other studies (with varying degrees of conviction).\r\n\r\nIs it really the case or I am doing something completely wrong with the segmentation? Could the organizers confirm (or deny) that some volumes were computed using SliceThickness instead of SliceLocation (possibly some program on a specific machine used such a default method?)? If it is so then is it possible to distinguish these cases using some DICOM metadata?\r\n\r\nIf there are such erroneous data then this is a big problem for constructing learning algorithms of any kind...\r\n\r\n**EDIT:** Sorry - I used my Python code numbering of frames starting from 0 - so these images correspond to files ending with 0030.dcm\r\n\r\n  [1]: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18372/some-cases-are-quite-off-from-the-true-value",
    "107380": "Hi rmldj,\r\n\r\nAs a Kaggle Master you probably know what I'm about to say, but it bears repeating since we get a lot of questions on data integrity.\r\n\r\nI can't help you on the SliceThickness vs. SliceLocation question, but you should expect real data to have issues like this, often without an explanation (see https://www.kaggle.com/wiki/ANoteOnDataQuality). These might stem from clerical typos, a tired doc, a confused trainee, differing protocols, a wrong assumption on a database join, you name it.\r\n\r\nThe \"big problem\" for us occurs when the noise is systematic and widespread, else it's on the algorithms to be robust.",
    "107489": "rmldj\r\n\r\nWhat are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.",
    "107491": "[quote=Paul Jurczak;107489]\r\n\r\n@rmldj\r\n\r\nWhat are the file names for slices 1 through 10 in your report? Are they 267\\study\\sax_6\\IM-7501-0029.dcm through 267\\study\\sax_15\\IM-7510-0029.dcm? If so, I'm getting substantially different pixel counts.\r\n\r\n[/quote]\r\n\r\nThese are those files but in reverse order (from sax_15 to sax_6) - I order them according to increasing Slice Location...",
    "107492": "Sorry - I number frames from 0 (for easy iteration in Python) - so these files correspond to files ending with 0030.dcm! I will edit the topic to make it clear...",
    "107493": "The pixel counts I'm getting are (slice position, pixel count for image #29):\r\n\r\n-111.3,\t134\r\n\r\n-101.3,\t320\r\n\r\n-91.3,\t420\r\n\r\n-81.3,\t658\r\n\r\n-71.3,\t613\r\n\r\n-61.3,\t916\r\n\r\n-51.3,\t971\r\n\r\n-41.3,\t900\r\n\r\n-31.3,\t888\r\n\r\n-21.3,\t400\r\n\r\nwhich produces a volume of about 137 ml. My segmentation usually produces a bit smaller area than I would do manually, so I agree with your suspicion of large error in study #267.\r\n\r\nBTW, your segmentation of slice -21.3 is much too large (see attached image).",
    "107494": ">  *I number frames from 0*\r\n\r\nI was going to add that, at least for some slices, the frame #30 has slightly higher pixel count, which illustrates your point even better.",
    "107496": "Thanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.\r\n\r\nHowever concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand?",
    "107499": "[quote=rmldj;107496]\r\n\r\nThanks! I definitely agree that slice 10 segmentation is wrong in my algorithm.\r\n\r\nHowever concerning your last comment - I did not seem to encounter a slice where frame #30 was larger than the rest (or I dropped these slices for a different reason??) - do you recall an example off hand? \r\n\r\n[/quote]\r\n\r\nI didn't mean that image dimensions were larger, just the pixel count for LV segment was slightly larger. That's why you've chosen frames #30 to measure LV diastole end volume instead of other frames, right?",
    "107500": "Ok. I understand what you meant - indeed in this study frame #30 was largest..",
    "108083": "William Cukierski <br>\r\nA 60ML error in ground truth might mean a difference of 0.0005+ on the LB depending on the # of patients if my calculations are correct. If someone took a gamble (which you could do for cases like 429) you could improve the LB score significanty with luck without having a better solution.. Personally I think this is a but too much for a 100.000+ competition.\r\n<br>\r\n<br>\r\nNext to this.. If we want to boast clinical significance of computerized volume measurement it would be good to have a clinical significant test-test to begin with. I'm pretty sure we can get at ~10ml deviations but cases like 429 spoil te party a bit..\r\n<br>\r\n<br>\r\nWouldn't it be possible to do a double check on the final test-set ?",
    "108290": "Julian Even with perfect data competitions would still have an element of luck. It's the nature of having a large number of teams and a finite test set (one could take bootstrap subsets of the test set and, depending on how close the scores are, have a different leaderboard each time). On the plus side, noise for you is noise for everyone else.\r\n\r\nWhy can't we fix noise in this specific case? Primarily, it's not economically reasonable to ask cardiologists to re-read the whole dataset to correct a few potential outliers. Even if the time and budget was there, I'm unconvinced that such a costly effort meaningfully changes the contributions of this project towards computerized volume measurement. We have to accept that \"the perfect is the enemy of the good\" and acknowledge the amount of \"good\" behind the data science bowl: a philanthropic commitment and involvement from the team at Booz + a forward-thinking research team at the NIH willing to put a large imaging data out there + a community of people like you who are capable of pushing the state of the art.\r\n\r\nMaybe this answer isn't the most satisfying, but I think it's important not to focus on small flaws at the expense of the larger opportunity.",
    "108299": "Could it be possible for the organizers to tell us how is the measured volume for case 595 and 599 (which are in validation set so we can't check) calculated because they only have 3 slices and it apparently did not span the whole range of the heart, did the organizers use some method other than contour-->area*thickness method for these cases because it will underestimate the true value significantly (but the provided number could have been calculated this way). We're building models to predict the provided numbers, not the true heart volume, so it affects how will we deal with cases with 3 slices only....",
    "108381": "woshialex\r\n\r\nGood question! Technically, study #595 can be handled by extrapolating SAX slice volumes using the \"side\" view of LV from 2ch_6, but this is very hard to automatize. I'm planning to have a fallback option for incomplete data cases like this, where I would use only midsection of LV for volume prediction. There should be a strong correlation between LV midsection slice surface area and LV volume.",
    "108820": "William Cukierski,<br>\r\nI see your point..\r\n\r\nFirst and foremost it's awesome that we are given the opportunity to solve these cool problems.<br>\r\nThanks for that."
  },
  "source": "meta"
}