{
  "id": 17885,
  "title": "Frequently Asked Questions",
  "url": "/competitions/second-annual-data-science-bowl/discussion/17885",
  "author_name": "",
  "post_date": "2015-12-14T18:32:34.470Z",
  "votes": 6,
  "comment_count": 24,
  "views": 9048,
  "content": "<p><strong>Q:</strong> Can you explain <em>topic X that requires subject matter expertise</em>?</p>\n\n<p><strong>A:</strong> Dr. Hansen has answered various questions regarding medical imaging and the dataset here: <a href=\"https://www.kaggle.com/michaelhansen/forum\">https://www.kaggle.com/michaelhansen/forum</a></p>\n\n<hr>\n\n<p><strong>Q:</strong> Can you explain what each file name means and how the files are related to each other?</p>\n\n<p><strong>A:</strong> The file naming convention is explained here: <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18257/basic-understanding-of-sax-2ch-4ch\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18257/basic-understanding-of-sax-2ch-4ch</a></p>\n\n<hr>\n\n<p><strong>Q:</strong> The Deep Learning tutorial is using an external data source, Sunnybrook dataset. Is this data source allowed in the competition?</p>\n\n<p><strong>A:</strong> Use of external data is not permitted, with the exception of the The Sunnybrook Cardiac Data (SCD). source: <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/rules\">https://www.kaggle.com/c/second-annual-data-science-bowl/rules</a> This means you may not pre-train models on datasets that are outside Kaggle DSB and Sunnybrook. You may, however, simulate original datasets via code if you wish. However, pretrained networks (even with an open license) are not allowed.</p>\n\n<hr>\n\n<p><strong>Q:</strong> I'm having trouble downloading the dataset. What are some good ways to download the data?</p>\n\n<p><strong>A:</strong> Try using a download manager such as: </p>\n\n<p><a href=\"https://addons.mozilla.org/en-US/firefox/addon/downthemall/developers\">https://addons.mozilla.org/en-US/firefox/addon/downthemall/developers</a></p>\n\n<p>Or try using wget with cookies from the terminal: </p>\n\n<p><a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line\">https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/96790\">https://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/96790</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/malware-classification/forums/t/12407/anyone-managed-to-use-wget-from-a-google-login/63584\">https://www.kaggle.com/c/malware-classification/forums/t/12407/anyone-managed-to-use-wget-from-a-google-login/63584</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/15313/downloading-data/85858\">https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/15313/downloading-data/85858</a></p>\n\n<hr>\n\n<p><strong>Q:</strong> Can Booz Allen Hamilton employees compete in the DSB?</p>\n\n<p><strong>A:</strong> Booz Allen Hamilton employees are not eligible to win any prize money. They, however, are free to compete for Kaggle points. A few employees, such as myself, are not allowed to compete, and act as hosts.</p>\n\n<hr>\n\n<p><strong>Q:</strong> Can I use language/software/package/library X for the DSB?</p>\n\n<p><strong>A:</strong> I'm making the call on this. You can use what you want. However, if a key component of your winning (top 3) model uses anything that is closed source (that lacks an open source alternative), that portion must be rewritten in an open language. This is to prevent black boxes that magically solve the DSB. A closed magic box, even if it solves the problem, is not useful to the sponsors or the broader community. We want to be able to replicate your approximate results using open source tools if necessary.</p>\n\n<hr>\n\n<p><strong>Q:</strong> Can I use meta data associated with folders names and DICOM files?</p>\n\n<p><strong>A:</strong> Yes but we can't confirm whether these features will be useful in the final private leaderboard scoring. This thread provides useful information on the DICOM files:\n<a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17906/other-dicom-aware-tools\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17906/other-dicom-aware-tools</a></p>\n\n<p>I will keep updating this, with the help of the other admins, as new questions come in.</p>",
  "messages": [
    {
      "id": "101199",
      "postDate": "12/14/2015 18:32:34",
      "content": "<p><strong>Q:</strong> Can you explain <em>topic X that requires subject matter expertise</em>?</p>\n\n<p><strong>A:</strong> Dr. Hansen has answered various questions regarding medical imaging and the dataset here: <a href=\"https://www.kaggle.com/michaelhansen/forum\">https://www.kaggle.com/michaelhansen/forum</a></p>\n\n<hr>\n\n<p><strong>Q:</strong> Can you explain what each file name means and how the files are related to each other?</p>\n\n<p><strong>A:</strong> The file naming convention is explained here: <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18257/basic-understanding-of-sax-2ch-4ch\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18257/basic-understanding-of-sax-2ch-4ch</a></p>\n\n<hr>\n\n<p><strong>Q:</strong> The Deep Learning tutorial is using an external data source, Sunnybrook dataset. Is this data source allowed in the competition?</p>\n\n<p><strong>A:</strong> Use of external data is not permitted, with the exception of the The Sunnybrook Cardiac Data (SCD). source: <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/rules\">https://www.kaggle.com/c/second-annual-data-science-bowl/rules</a> This means you may not pre-train models on datasets that are outside Kaggle DSB and Sunnybrook. You may, however, simulate original datasets via code if you wish. However, pretrained networks (even with an open license) are not allowed.</p>\n\n<hr>\n\n<p><strong>Q:</strong> I'm having trouble downloading the dataset. What are some good ways to download the data?</p>\n\n<p><strong>A:</strong> Try using a download manager such as: </p>\n\n<p><a href=\"https://addons.mozilla.org/en-US/firefox/addon/downthemall/developers\">https://addons.mozilla.org/en-US/firefox/addon/downthemall/developers</a></p>\n\n<p>Or try using wget with cookies from the terminal: </p>\n\n<p><a href=\"https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line\">https://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/96790\">https://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/96790</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/malware-classification/forums/t/12407/anyone-managed-to-use-wget-from-a-google-login/63584\">https://www.kaggle.com/c/malware-classification/forums/t/12407/anyone-managed-to-use-wget-from-a-google-login/63584</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/15313/downloading-data/85858\">https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/15313/downloading-data/85858</a></p>\n\n<hr>\n\n<p><strong>Q:</strong> Can Booz Allen Hamilton employees compete in the DSB?</p>\n\n<p><strong>A:</strong> Booz Allen Hamilton employees are not eligible to win any prize money. They, however, are free to compete for Kaggle points. A few employees, such as myself, are not allowed to compete, and act as hosts.</p>\n\n<hr>\n\n<p><strong>Q:</strong> Can I use language/software/package/library X for the DSB?</p>\n\n<p><strong>A:</strong> I'm making the call on this. You can use what you want. However, if a key component of your winning (top 3) model uses anything that is closed source (that lacks an open source alternative), that portion must be rewritten in an open language. This is to prevent black boxes that magically solve the DSB. A closed magic box, even if it solves the problem, is not useful to the sponsors or the broader community. We want to be able to replicate your approximate results using open source tools if necessary.</p>\n\n<hr>\n\n<p><strong>Q:</strong> Can I use meta data associated with folders names and DICOM files?</p>\n\n<p><strong>A:</strong> Yes but we can't confirm whether these features will be useful in the final private leaderboard scoring. This thread provides useful information on the DICOM files:\n<a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17906/other-dicom-aware-tools\">https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17906/other-dicom-aware-tools</a></p>\n\n<p>I will keep updating this, with the help of the other admins, as new questions come in.</p>",
      "rawMarkdown": "**Q:** Can you explain *topic X that requires subject matter expertise*?\r\n\r\n**A:** Dr. Hansen has answered various questions regarding medical imaging and the dataset here: https://www.kaggle.com/michaelhansen/forum\r\n\r\n----------\r\n\r\n**Q:** Can you explain what each file name means and how the files are related to each other?\r\n\r\n**A:** The file naming convention is explained here: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18257/basic-understanding-of-sax-2ch-4ch\r\n\r\n----------\r\n\r\n**Q:** The Deep Learning tutorial is using an external data source, Sunnybrook dataset. Is this data source allowed in the competition?\r\n\r\n**A:** Use of external data is not permitted, with the exception of the The Sunnybrook Cardiac Data (SCD). source: https://www.kaggle.com/c/second-annual-data-science-bowl/rules This means you may not pre-train models on datasets that are outside Kaggle DSB and Sunnybrook. You may, however, simulate original datasets via code if you wish. However, pretrained networks (even with an open license) are not allowed.\r\n\r\n----------\r\n\r\n\r\n**Q:** I'm having trouble downloading the dataset. What are some good ways to download the data?\r\n\r\n**A:** Try using a download manager such as: \r\n\r\nhttps://addons.mozilla.org/en-US/firefox/addon/downthemall/developers\r\n\r\nOr try using wget with cookies from the terminal: \r\n\r\nhttps://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line \r\n\r\nhttps://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/96790 \r\n\r\nhttps://www.kaggle.com/c/malware-classification/forums/t/12407/anyone-managed-to-use-wget-from-a-google-login/63584\r\n\r\nhttps://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/15313/downloading-data/85858\r\n\r\n----------\r\n**Q:** Can Booz Allen Hamilton employees compete in the DSB?\r\n\r\n**A:** Booz Allen Hamilton employees are not eligible to win any prize money. They, however, are free to compete for Kaggle points. A few employees, such as myself, are not allowed to compete, and act as hosts.\r\n\r\n\r\n----------\r\n**Q:** Can I use language/software/package/library X for the DSB?\r\n\r\n**A:** I'm making the call on this. You can use what you want. However, if a key component of your winning (top 3) model uses anything that is closed source (that lacks an open source alternative), that portion must be rewritten in an open language. This is to prevent black boxes that magically solve the DSB. A closed magic box, even if it solves the problem, is not useful to the sponsors or the broader community. We want to be able to replicate your approximate results using open source tools if necessary.\r\n\r\n\r\n----------\r\n\r\n**Q:** Can I use meta data associated with folders names and DICOM files?\r\n\r\n**A:** Yes but we can't confirm whether these features will be useful in the final private leaderboard scoring. This thread provides useful information on the DICOM files:\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17906/other-dicom-aware-tools\r\n\r\nI will keep updating this, with the help of the other admins, as new questions come in.",
      "votes": null
    },
    {
      "id": "101446",
      "postDate": "12/15/2015 12:53:34",
      "content": "<p>[quote=Mike Kim;101199]</p>\n\n<p>I will keep updating this, with the help of the other admins, as new questions come in.</p>\n\n<p>[/quote]</p>\n\n<p>Can you please tell how big is dataset after extracting the zip files.</p>",
      "rawMarkdown": "[quote=Mike Kim;101199]\r\n\r\nI will keep updating this, with the help of the other admins, as new questions come in.\r\n\r\n[/quote]\r\n\r\nCan you please tell how big is dataset after extracting the zip files.",
      "votes": null
    },
    {
      "id": "101451",
      "postDate": "12/15/2015 13:27:15",
      "content": "<p>[quote=DataGeek;101446]\nhow big is dataset after extracting the zip files.\n[/quote]</p>\n\n<p>$ du -s -h train validate</p>\n\n<p>32G    train</p>\n\n<p>13G    validate</p>\n\n<p>$</p>",
      "rawMarkdown": "[quote=DataGeek;101446]\r\nhow big is dataset after extracting the zip files.\r\n[/quote]\r\n\r\n$ du -s -h train validate\r\n\r\n 32G\ttrain\r\n\r\n 13G\tvalidate\r\n\r\n$",
      "votes": null
    },
    {
      "id": "101452",
      "postDate": "12/15/2015 13:29:29",
      "content": "<p>[quote=Kazuomi Kashii;101451]</p>\n\n<p>[quote=DataGeek;101446]\nhow big is dataset after extracting the zip files.\n[/quote]</p>\n\n<p>$ du -s -h train validate</p>\n\n<p>32G    train</p>\n\n<p>13G    validate</p>\n\n<p>$</p>\n\n<p>[/quote]</p>\n\n<p>Thanks :)</p>",
      "rawMarkdown": "[quote=Kazuomi Kashii;101451]\r\n\r\n[quote=DataGeek;101446]\r\nhow big is dataset after extracting the zip files.\r\n[/quote]\r\n\r\n$ du -s -h train validate\r\n\r\n 32G\ttrain\r\n\r\n 13G\tvalidate\r\n\r\n$\r\n\r\n[/quote]\r\n\r\nThanks :)",
      "votes": null
    },
    {
      "id": "101606",
      "postDate": "12/16/2015 03:31:02",
      "content": "<p>This question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : </p>\n\n<p>Are Vd and Vs in someway related to the blood outflow?</p>\n\n<p>I'm assuming that if the Vs or Vd are <em>different</em> than a specific metric, it impacts the rate of blood flow in and out of LV ??</p>",
      "rawMarkdown": "This question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : \r\n\r\nAre Vd and Vs in someway related to the blood outflow?\r\n\r\nI'm assuming that if the Vs or Vd are *different* than a specific metric, it impacts the rate of blood flow in and out of LV ??",
      "votes": null
    },
    {
      "id": "101671",
      "postDate": "12/16/2015 11:10:52",
      "content": "<p>[quote=UshaKota;101606]</p>\n\n<p>This question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : </p>\n\n<p>Are Vd and Vs in someway related to the blood outflow?</p>\n\n<p>I'm assuming that if the Vs or Vd are <em>different</em> than a specific metric, it impacts the rate of blood flow in and out of LV ??</p>\n\n<p>[/quote]</p>\n\n<p>The volumes are used to compute ejection fraction, which is the percentage of blood that's pumped out of the left ventricle with each beat. More description is available here: <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/data\">https://www.kaggle.com/c/second-annual-data-science-bowl/data</a></p>",
      "rawMarkdown": "[quote=UshaKota;101606]\r\n\r\nThis question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : \r\n\r\nAre Vd and Vs in someway related to the blood outflow?\r\n\r\nI'm assuming that if the Vs or Vd are *different* than a specific metric, it impacts the rate of blood flow in and out of LV ??\r\n\r\n\r\n[/quote]\r\n\r\nThe volumes are used to compute ejection fraction, which is the percentage of blood that's pumped out of the left ventricle with each beat. More description is available here: https://www.kaggle.com/c/second-annual-data-science-bowl/data",
      "votes": null
    },
    {
      "id": "101677",
      "postDate": "12/16/2015 12:14:23",
      "content": "<p>Thank you</p>\n\n<p>I read thru the LVEF and LV functionality and understood that identifying the LV and LV segments cross-sectional area is itself most significant  here in the problem.\nSince I do not have a clue about what data can be extracted from MRI scans, I thought at each segment view, the blood outflow is &quot;somehow&quot; captured.</p>\n\n<p>Thanks again</p>",
      "rawMarkdown": "Thank you\r\n\r\nI read thru the LVEF and LV functionality and understood that identifying the LV and LV segments cross-sectional area is itself most significant  here in the problem.\r\nSince I do not have a clue about what data can be extracted from MRI scans, I thought at each segment view, the blood outflow is \"somehow\" captured.\r\n\r\nThanks again",
      "votes": null
    },
    {
      "id": "101682",
      "postDate": "12/16/2015 13:31:16",
      "content": "<p>Mike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  </p>\n\n<p>thank you </p>",
      "rawMarkdown": "Mike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  \r\n\r\nthank you",
      "votes": null
    },
    {
      "id": "101754",
      "postDate": "12/16/2015 21:16:33",
      "content": "<p>[quote=Ekin;101682]</p>\n\n<p>Mike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  </p>\n\n<p>thank you </p>\n\n<p>[/quote]</p>\n\n<p>Hi Ekin. Yes, the data is available for research and academic pursuits. Please cite as &#8216;Data Science Bowl Cardiac Challenge Data&#8217;. And please share your work with us if you are so inclined!</p>",
      "rawMarkdown": "[quote=Ekin;101682]\r\n\r\n\r\nMike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  \r\n\r\nthank you \r\n\r\n[/quote]\r\n\r\nHi Ekin. Yes, the data is available for research and academic pursuits. Please cite as ‘Data Science Bowl Cardiac Challenge Data’. And please share your work with us if you are so inclined!",
      "votes": null
    },
    {
      "id": "101970",
      "postDate": "12/18/2015 06:13:32",
      "content": "<p>Hi, I am using a low-bandwidth, Could you please tell me the size of the test.zip ? I don't know whether&#8203; I am able to download this file and running my model within 7 days.</p>",
      "rawMarkdown": "Hi, I am using a low-bandwidth, Could you please tell me the size of the test.zip ? I don't know whether I am able to download this file and running my model within 7 days.",
      "votes": null
    },
    {
      "id": "101993",
      "postDate": "12/18/2015 10:56:28",
      "content": "<p>Can NVIDIA employees compete in the DSB? NVIDIA is not a sponsor but I wanna be on a safe side...</p>",
      "rawMarkdown": "Can NVIDIA employees compete in the DSB? NVIDIA is not a sponsor but I wanna be on a safe side...",
      "votes": null
    },
    {
      "id": "102171",
      "postDate": "12/20/2015 05:35:07",
      "content": "<p>EDIT: See Steve's post.</p>",
      "rawMarkdown": "EDIT: See Steve's post.",
      "votes": null
    },
    {
      "id": "102192",
      "postDate": "12/20/2015 10:49:49",
      "content": "<p>NVIDIA employees are welcomed and encouraged to participate. They are <strong>not</strong> eligible to win the cash prizes, however. They must adhere to the same rules that are applicable to Booz Allen staff. Specifically:</p>\n\n<ol>\n<li>You may not form a team with any prize-eligible participants. </li>\n<li>If you had any knowledge of the competition challenge prior to the December 14 launch, you are not\neligible to compete. </li>\n<li>When you create a team name, it must start with &#8220;NVD_&#8221;. This declares your company affiliation to all participants while allowing us to track your accomplishments.</li>\n</ol>",
      "rawMarkdown": "NVIDIA employees are welcomed and encouraged to participate. They are **not** eligible to win the cash prizes, however. They must adhere to the same rules that are applicable to Booz Allen staff. Specifically:\r\n\r\n 1. You may not form a team with any prize-eligible participants. \r\n 2. If you had any knowledge of the competition challenge prior to the December 14 launch, you are not\r\n    eligible to compete. \r\n 3. When you create a team name, it must start with “NVD_”. This declares your company affiliation to all participants while allowing us to track your accomplishments.",
      "votes": null
    },
    {
      "id": "102203",
      "postDate": "12/20/2015 12:12:27",
      "content": "<p>Mike, Steve, thanks a lot for clarification!</p>",
      "rawMarkdown": "Mike, Steve, thanks a lot for clarification!",
      "votes": null
    },
    {
      "id": "103085",
      "postDate": "12/28/2015 20:32:06",
      "content": "<p>NVIDIA and BAH employees may give me their code on the down low and then we'll split the prize money. KIDDING! </p>",
      "rawMarkdown": "NVIDIA and BAH employees may give me their code on the down low and then we'll split the prize money. KIDDING!",
      "votes": null
    },
    {
      "id": "103257",
      "postDate": "12/30/2015 19:48:07",
      "content": "<p>I would like to confirm if each volume in the training dataset corresponds to each sax folder/patient or is it the total volume comprising all folders/patient ?</p>",
      "rawMarkdown": "I would like to confirm if each volume in the training dataset corresponds to each sax folder/patient or is it the total volume comprising all folders/patient ?",
      "votes": null
    },
    {
      "id": "105480",
      "postDate": "01/23/2016 15:48:04",
      "content": "<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>",
      "rawMarkdown": "When the validation volumes are released 7th March, can those be used as part of the training for the final test set?",
      "votes": null
    },
    {
      "id": "105489",
      "postDate": "01/23/2016 18:20:20",
      "content": "<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>",
      "rawMarkdown": "[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes",
      "votes": null
    },
    {
      "id": "105494",
      "postDate": "01/23/2016 19:14:38",
      "content": "<p>Hi William,</p>\n\n<p>Timeline page - <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline\">https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline</a> - says &quot;Your model must be finalized and uploaded to Kaggle by this deadline&quot;. This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.</p>\n\n<p>And this makes sense. We can hardly expect to do the training for each patient at the hospital.</p>\n\n<p>Thanks,\nMax.</p>\n\n<p>[quote=William Cukierski;105489]</p>\n\n<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Hi William,\r\n\r\nTimeline page - https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline - says \"Your model must be finalized and uploaded to Kaggle by this deadline\". This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.\r\n\r\nAnd this makes sense. We can hardly expect to do the training for each patient at the hospital.\r\n\r\nThanks,\r\nMax.\r\n\r\n[quote=William Cukierski;105489]\r\n\r\n[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "105496",
      "postDate": "01/23/2016 19:15:30",
      "content": "<p>[quote=William Cukierski;105489]</p>\n\n<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>\n\n<p>[/quote]</p>\n\n<p>How can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?</p>",
      "rawMarkdown": "[quote=William Cukierski;105489]\r\n\r\n[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes\r\n\r\n[/quote]\r\n\r\nHow can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?",
      "votes": null
    },
    {
      "id": "105502",
      "postDate": "01/23/2016 20:23:32",
      "content": "<p>[quote=Maxim Milakov;105494]</p>\n\n<p>Hi William,</p>\n\n<p>Timeline page - <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline\">https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline</a> - says &quot;Your model must be finalized and uploaded to Kaggle by this deadline&quot;. This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.</p>\n\n<p>And this makes sense. We can hardly expect to do the training for each patient at the hospital.</p>\n\n<p>Thanks,\nMax.</p>\n\n<p>[quote=William Cukierski;105489]</p>\n\n<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>I actually have a different questions here. There is a large family of methods used in medical image analysis called multi-atlas segmentation. And the method does not have a regression or neural net model, but an algorithm with different components and a set of atlases.</p>\n\n<p>So the questions is, do you have to submit a &quot;model&quot;?</p>",
      "rawMarkdown": "[quote=Maxim Milakov;105494]\r\n\r\nHi William,\r\n\r\nTimeline page - https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline - says \"Your model must be finalized and uploaded to Kaggle by this deadline\". This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.\r\n\r\nAnd this makes sense. We can hardly expect to do the training for each patient at the hospital.\r\n\r\nThanks,\r\nMax.\r\n\r\n[quote=William Cukierski;105489]\r\n\r\n[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes\r\n\r\n[/quote]\r\n\r\n[/quote]\r\n\r\nI actually have a different questions here. There is a large family of methods used in medical image analysis called multi-atlas segmentation. And the method does not have a regression or neural net model, but an algorithm with different components and a set of atlases.\r\n\r\nSo the questions is, do you have to submit a \"model\"?",
      "votes": null
    },
    {
      "id": "105516",
      "postDate": "01/24/2016 01:43:07",
      "content": "<p>[quote=jstaker7;105496]</p>\n\n<p>How can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?</p>\n\n<p>[/quote]</p>\n\n<p>For our purposes, your &quot;model&quot; is the scientific approach for using a set of labeled images to make predictions on unlabeled images. While you are not allowed to modify the scientific approach during the final test phase of the competition, you are allowed to make changes such as:</p>\n\n<ul>\n<li>retraining the model on a set that includes the train + validation set</li>\n<li>pointing your code to the newly released test images</li>\n<li>MINOR bug fixes for test set surprises</li>\n</ul>\n\n<p>Why do we allow retraining on validation images? There are methods for &quot;reverse engineering&quot; labels from leaderboard scores (the current leaderboard uses 100% of the validation data). While we have no evidence of this happening in the competition, there is a potential that someone could reverse engineer the validation set labels and use them offline. To neutralize this potential unfair advantage, we release the labels and make the validation set fair game for training.</p>",
      "rawMarkdown": "[quote=jstaker7;105496]\r\n\r\nHow can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?\r\n\r\n[/quote]\r\n\r\nFor our purposes, your \"model\" is the scientific approach for using a set of labeled images to make predictions on unlabeled images. While you are not allowed to modify the scientific approach during the final test phase of the competition, you are allowed to make changes such as:\r\n\r\n- retraining the model on a set that includes the train + validation set\r\n- pointing your code to the newly released test images\r\n- MINOR bug fixes for test set surprises\r\n\r\nWhy do we allow retraining on validation images? There are methods for \"reverse engineering\" labels from leaderboard scores (the current leaderboard uses 100% of the validation data). While we have no evidence of this happening in the competition, there is a potential that someone could reverse engineer the validation set labels and use them offline. To neutralize this potential unfair advantage, we release the labels and make the validation set fair game for training.",
      "votes": null
    },
    {
      "id": "109853",
      "postDate": "03/01/2016 07:49:39",
      "content": "<p>hi \nI recently joined the bowl and i was wondering whether there will be continued development after the results . \nor the forums and data sets will be closed ?\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.</p>",
      "rawMarkdown": "hi \r\nI recently joined the bowl and i was wondering whether there will be continued development after the results . \r\nor the forums and data sets will be closed ?\r\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.",
      "votes": null
    },
    {
      "id": "111341",
      "postDate": "03/13/2016 22:11:24",
      "content": "<p>Hello,</p>\n\n<p>Could you please tell me - is it mandatory to keep the same order of patients' labels in our submission file as they are in sample_submission (starting with 1000-1140 and then 701-999)? Or we can sort them like 701-1140 order?</p>\n\n<p>Thanks in advance,\nAlex</p>",
      "rawMarkdown": "Hello,\r\n\r\nCould you please tell me - is it mandatory to keep the same order of patients' labels in our submission file as they are in sample_submission (starting with 1000-1140 and then 701-999)? Or we can sort them like 701-1140 order?\r\n\r\nThanks in advance,\r\nAlex",
      "votes": null
    },
    {
      "id": "111359",
      "postDate": "03/14/2016 02:29:14",
      "content": "<p>Any sort order works, as long as they're all there.</p>",
      "rawMarkdown": "Any sort order works, as long as they're all there.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 101446,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "12/15/2015 12:53:34",
      "content": "<p>[quote=Mike Kim;101199]</p>\n\n<p>I will keep updating this, with the help of the other admins, as new questions come in.</p>\n\n<p>[/quote]</p>\n\n<p>Can you please tell how big is dataset after extracting the zip files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101451,
      "author_name": "kazkus",
      "author_url": "",
      "post_date": "12/15/2015 13:27:15",
      "content": "<p>[quote=DataGeek;101446]\nhow big is dataset after extracting the zip files.\n[/quote]</p>\n\n<p>$ du -s -h train validate</p>\n\n<p>32G    train</p>\n\n<p>13G    validate</p>\n\n<p>$</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101452,
      "author_name": "thakurrajanand",
      "author_url": "",
      "post_date": "12/15/2015 13:29:29",
      "content": "<p>[quote=Kazuomi Kashii;101451]</p>\n\n<p>[quote=DataGeek;101446]\nhow big is dataset after extracting the zip files.\n[/quote]</p>\n\n<p>$ du -s -h train validate</p>\n\n<p>32G    train</p>\n\n<p>13G    validate</p>\n\n<p>$</p>\n\n<p>[/quote]</p>\n\n<p>Thanks :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101606,
      "author_name": "ushakota",
      "author_url": "",
      "post_date": "12/16/2015 03:31:02",
      "content": "<p>This question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : </p>\n\n<p>Are Vd and Vs in someway related to the blood outflow?</p>\n\n<p>I'm assuming that if the Vs or Vd are <em>different</em> than a specific metric, it impacts the rate of blood flow in and out of LV ??</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101671,
      "author_name": "shannonlantzy",
      "author_url": "",
      "post_date": "12/16/2015 11:10:52",
      "content": "<p>[quote=UshaKota;101606]</p>\n\n<p>This question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : </p>\n\n<p>Are Vd and Vs in someway related to the blood outflow?</p>\n\n<p>I'm assuming that if the Vs or Vd are <em>different</em> than a specific metric, it impacts the rate of blood flow in and out of LV ??</p>\n\n<p>[/quote]</p>\n\n<p>The volumes are used to compute ejection fraction, which is the percentage of blood that's pumped out of the left ventricle with each beat. More description is available here: <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/data\">https://www.kaggle.com/c/second-annual-data-science-bowl/data</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101677,
      "author_name": "ushakota",
      "author_url": "",
      "post_date": "12/16/2015 12:14:23",
      "content": "<p>Thank you</p>\n\n<p>I read thru the LVEF and LV functionality and understood that identifying the LV and LV segments cross-sectional area is itself most significant  here in the problem.\nSince I do not have a clue about what data can be extracted from MRI scans, I thought at each segment view, the blood outflow is &quot;somehow&quot; captured.</p>\n\n<p>Thanks again</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101682,
      "author_name": "ahmetekin",
      "author_url": "",
      "post_date": "12/16/2015 13:31:16",
      "content": "<p>Mike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  </p>\n\n<p>thank you </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101754,
      "author_name": "shannonlantzy",
      "author_url": "",
      "post_date": "12/16/2015 21:16:33",
      "content": "<p>[quote=Ekin;101682]</p>\n\n<p>Mike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  </p>\n\n<p>thank you </p>\n\n<p>[/quote]</p>\n\n<p>Hi Ekin. Yes, the data is available for research and academic pursuits. Please cite as &#8216;Data Science Bowl Cardiac Challenge Data&#8217;. And please share your work with us if you are so inclined!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101970,
      "author_name": "liubenyuan",
      "author_url": "",
      "post_date": "12/18/2015 06:13:32",
      "content": "<p>Hi, I am using a low-bandwidth, Could you please tell me the size of the test.zip ? I don't know whether&#8203; I am able to download this file and running my model within 7 days.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101993,
      "author_name": "milakov",
      "author_url": "",
      "post_date": "12/18/2015 10:56:28",
      "content": "<p>Can NVIDIA employees compete in the DSB? NVIDIA is not a sponsor but I wanna be on a safe side...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102171,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "12/20/2015 05:35:07",
      "content": "<p>EDIT: See Steve's post.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102192,
      "author_name": "sdmills",
      "author_url": "",
      "post_date": "12/20/2015 10:49:49",
      "content": "<p>NVIDIA employees are welcomed and encouraged to participate. They are <strong>not</strong> eligible to win the cash prizes, however. They must adhere to the same rules that are applicable to Booz Allen staff. Specifically:</p>\n\n<ol>\n<li>You may not form a team with any prize-eligible participants. </li>\n<li>If you had any knowledge of the competition challenge prior to the December 14 launch, you are not\neligible to compete. </li>\n<li>When you create a team name, it must start with &#8220;NVD_&#8221;. This declares your company affiliation to all participants while allowing us to track your accomplishments.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 102203,
      "author_name": "milakov",
      "author_url": "",
      "post_date": "12/20/2015 12:12:27",
      "content": "<p>Mike, Steve, thanks a lot for clarification!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103085,
      "author_name": "millerintllc",
      "author_url": "",
      "post_date": "12/28/2015 20:32:06",
      "content": "<p>NVIDIA and BAH employees may give me their code on the down low and then we'll split the prize money. KIDDING! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 103257,
      "author_name": "liveflow",
      "author_url": "",
      "post_date": "12/30/2015 19:48:07",
      "content": "<p>I would like to confirm if each volume in the training dataset corresponds to each sax folder/patient or is it the total volume comprising all folders/patient ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105480,
      "author_name": "wenzhe",
      "author_url": "",
      "post_date": "01/23/2016 15:48:04",
      "content": "<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105489,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "01/23/2016 18:20:20",
      "content": "<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105494,
      "author_name": "milakov",
      "author_url": "",
      "post_date": "01/23/2016 19:14:38",
      "content": "<p>Hi William,</p>\n\n<p>Timeline page - <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline\">https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline</a> - says &quot;Your model must be finalized and uploaded to Kaggle by this deadline&quot;. This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.</p>\n\n<p>And this makes sense. We can hardly expect to do the training for each patient at the hospital.</p>\n\n<p>Thanks,\nMax.</p>\n\n<p>[quote=William Cukierski;105489]</p>\n\n<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105496,
      "author_name": "jstaker7",
      "author_url": "",
      "post_date": "01/23/2016 19:15:30",
      "content": "<p>[quote=William Cukierski;105489]</p>\n\n<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>\n\n<p>[/quote]</p>\n\n<p>How can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105502,
      "author_name": "wenzhe",
      "author_url": "",
      "post_date": "01/23/2016 20:23:32",
      "content": "<p>[quote=Maxim Milakov;105494]</p>\n\n<p>Hi William,</p>\n\n<p>Timeline page - <a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline\">https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline</a> - says &quot;Your model must be finalized and uploaded to Kaggle by this deadline&quot;. This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.</p>\n\n<p>And this makes sense. We can hardly expect to do the training for each patient at the hospital.</p>\n\n<p>Thanks,\nMax.</p>\n\n<p>[quote=William Cukierski;105489]</p>\n\n<p>[quote=Imagic Mobile;105480]</p>\n\n<p>When the validation volumes are released 7th March, can those be used as part of the training for the final test set?</p>\n\n<p>[/quote]</p>\n\n<p>Yes</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>I actually have a different questions here. There is a large family of methods used in medical image analysis called multi-atlas segmentation. And the method does not have a regression or neural net model, but an algorithm with different components and a set of atlases.</p>\n\n<p>So the questions is, do you have to submit a &quot;model&quot;?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105516,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "01/24/2016 01:43:07",
      "content": "<p>[quote=jstaker7;105496]</p>\n\n<p>How can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?</p>\n\n<p>[/quote]</p>\n\n<p>For our purposes, your &quot;model&quot; is the scientific approach for using a set of labeled images to make predictions on unlabeled images. While you are not allowed to modify the scientific approach during the final test phase of the competition, you are allowed to make changes such as:</p>\n\n<ul>\n<li>retraining the model on a set that includes the train + validation set</li>\n<li>pointing your code to the newly released test images</li>\n<li>MINOR bug fixes for test set surprises</li>\n</ul>\n\n<p>Why do we allow retraining on validation images? There are methods for &quot;reverse engineering&quot; labels from leaderboard scores (the current leaderboard uses 100% of the validation data). While we have no evidence of this happening in the competition, there is a potential that someone could reverse engineer the validation set labels and use them offline. To neutralize this potential unfair advantage, we release the labels and make the validation set fair game for training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109853,
      "author_name": "waleedsial",
      "author_url": "",
      "post_date": "03/01/2016 07:49:39",
      "content": "<p>hi \nI recently joined the bowl and i was wondering whether there will be continued development after the results . \nor the forums and data sets will be closed ?\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111341,
      "author_name": "alexryzhkov",
      "author_url": "",
      "post_date": "03/13/2016 22:11:24",
      "content": "<p>Hello,</p>\n\n<p>Could you please tell me - is it mandatory to keep the same order of patients' labels in our submission file as they are in sample_submission (starting with 1000-1140 and then 701-999)? Or we can sort them like 701-1140 order?</p>\n\n<p>Thanks in advance,\nAlex</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 111359,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "03/14/2016 02:29:14",
      "content": "<p>Any sort order works, as long as they're all there.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "101199": "**Q:** Can you explain *topic X that requires subject matter expertise*?\r\n\r\n**A:** Dr. Hansen has answered various questions regarding medical imaging and the dataset here: https://www.kaggle.com/michaelhansen/forum\r\n\r\n----------\r\n\r\n**Q:** Can you explain what each file name means and how the files are related to each other?\r\n\r\n**A:** The file naming convention is explained here: https://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/18257/basic-understanding-of-sax-2ch-4ch\r\n\r\n----------\r\n\r\n**Q:** The Deep Learning tutorial is using an external data source, Sunnybrook dataset. Is this data source allowed in the competition?\r\n\r\n**A:** Use of external data is not permitted, with the exception of the The Sunnybrook Cardiac Data (SCD). source: https://www.kaggle.com/c/second-annual-data-science-bowl/rules This means you may not pre-train models on datasets that are outside Kaggle DSB and Sunnybrook. You may, however, simulate original datasets via code if you wish. However, pretrained networks (even with an open license) are not allowed.\r\n\r\n----------\r\n\r\n\r\n**Q:** I'm having trouble downloading the dataset. What are some good ways to download the data?\r\n\r\n**A:** Try using a download manager such as: \r\n\r\nhttps://addons.mozilla.org/en-US/firefox/addon/downthemall/developers\r\n\r\nOr try using wget with cookies from the terminal: \r\n\r\nhttps://www.kaggle.com/forums/f/15/kaggle-forum/t/6604/downloading-data-via-command-line \r\n\r\nhttps://www.kaggle.com/c/belkin-energy-disaggregation-competition/forums/t/5118/downloading-data-via-wget/96790 \r\n\r\nhttps://www.kaggle.com/c/malware-classification/forums/t/12407/anyone-managed-to-use-wget-from-a-google-login/63584\r\n\r\nhttps://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/15313/downloading-data/85858\r\n\r\n----------\r\n**Q:** Can Booz Allen Hamilton employees compete in the DSB?\r\n\r\n**A:** Booz Allen Hamilton employees are not eligible to win any prize money. They, however, are free to compete for Kaggle points. A few employees, such as myself, are not allowed to compete, and act as hosts.\r\n\r\n\r\n----------\r\n**Q:** Can I use language/software/package/library X for the DSB?\r\n\r\n**A:** I'm making the call on this. You can use what you want. However, if a key component of your winning (top 3) model uses anything that is closed source (that lacks an open source alternative), that portion must be rewritten in an open language. This is to prevent black boxes that magically solve the DSB. A closed magic box, even if it solves the problem, is not useful to the sponsors or the broader community. We want to be able to replicate your approximate results using open source tools if necessary.\r\n\r\n\r\n----------\r\n\r\n**Q:** Can I use meta data associated with folders names and DICOM files?\r\n\r\n**A:** Yes but we can't confirm whether these features will be useful in the final private leaderboard scoring. This thread provides useful information on the DICOM files:\r\nhttps://www.kaggle.com/c/second-annual-data-science-bowl/forums/t/17906/other-dicom-aware-tools\r\n\r\nI will keep updating this, with the help of the other admins, as new questions come in.",
    "101446": "[quote=Mike Kim;101199]\r\n\r\nI will keep updating this, with the help of the other admins, as new questions come in.\r\n\r\n[/quote]\r\n\r\nCan you please tell how big is dataset after extracting the zip files.",
    "101451": "[quote=DataGeek;101446]\r\nhow big is dataset after extracting the zip files.\r\n[/quote]\r\n\r\n$ du -s -h train validate\r\n\r\n 32G\ttrain\r\n\r\n 13G\tvalidate\r\n\r\n$",
    "101452": "[quote=Kazuomi Kashii;101451]\r\n\r\n[quote=DataGeek;101446]\r\nhow big is dataset after extracting the zip files.\r\n[/quote]\r\n\r\n$ du -s -h train validate\r\n\r\n 32G\ttrain\r\n\r\n 13G\tvalidate\r\n\r\n$\r\n\r\n[/quote]\r\n\r\nThanks :)",
    "101606": "This question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : \r\n\r\nAre Vd and Vs in someway related to the blood outflow?\r\n\r\nI'm assuming that if the Vs or Vd are *different* than a specific metric, it impacts the rate of blood flow in and out of LV ??",
    "101671": "[quote=UshaKota;101606]\r\n\r\nThis question is related to the measurement of Vd and Vs and may not  help to analyse the problem much, however from a naive perspective : \r\n\r\nAre Vd and Vs in someway related to the blood outflow?\r\n\r\nI'm assuming that if the Vs or Vd are *different* than a specific metric, it impacts the rate of blood flow in and out of LV ??\r\n\r\n\r\n[/quote]\r\n\r\nThe volumes are used to compute ejection fraction, which is the percentage of blood that's pumped out of the left ventricle with each beat. More description is available here: https://www.kaggle.com/c/second-annual-data-science-bowl/data",
    "101677": "Thank you\r\n\r\nI read thru the LVEF and LV functionality and understood that identifying the LV and LV segments cross-sectional area is itself most significant  here in the problem.\r\nSince I do not have a clue about what data can be extracted from MRI scans, I thought at each segment view, the blood outflow is \"somehow\" captured.\r\n\r\nThanks again",
    "101682": "Mike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  \r\n\r\nthank you",
    "101754": "[quote=Ekin;101682]\r\n\r\n\r\nMike, is it allowed to use the dataset for academic publication (with of course proper citation)? Also, is it allowed to use some of the MR images in the dataset as figures in such a publication?  \r\n\r\nthank you \r\n\r\n[/quote]\r\n\r\nHi Ekin. Yes, the data is available for research and academic pursuits. Please cite as ‘Data Science Bowl Cardiac Challenge Data’. And please share your work with us if you are so inclined!",
    "101970": "Hi, I am using a low-bandwidth, Could you please tell me the size of the test.zip ? I don't know whether I am able to download this file and running my model within 7 days.",
    "101993": "Can NVIDIA employees compete in the DSB? NVIDIA is not a sponsor but I wanna be on a safe side...",
    "102171": "EDIT: See Steve's post.",
    "102192": "NVIDIA employees are welcomed and encouraged to participate. They are **not** eligible to win the cash prizes, however. They must adhere to the same rules that are applicable to Booz Allen staff. Specifically:\r\n\r\n 1. You may not form a team with any prize-eligible participants. \r\n 2. If you had any knowledge of the competition challenge prior to the December 14 launch, you are not\r\n    eligible to compete. \r\n 3. When you create a team name, it must start with “NVD_”. This declares your company affiliation to all participants while allowing us to track your accomplishments.",
    "102203": "Mike, Steve, thanks a lot for clarification!",
    "103085": "NVIDIA and BAH employees may give me their code on the down low and then we'll split the prize money. KIDDING!",
    "103257": "I would like to confirm if each volume in the training dataset corresponds to each sax folder/patient or is it the total volume comprising all folders/patient ?",
    "105480": "When the validation volumes are released 7th March, can those be used as part of the training for the final test set?",
    "105489": "[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes",
    "105494": "Hi William,\r\n\r\nTimeline page - https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline - says \"Your model must be finalized and uploaded to Kaggle by this deadline\". This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.\r\n\r\nAnd this makes sense. We can hardly expect to do the training for each patient at the hospital.\r\n\r\nThanks,\r\nMax.\r\n\r\n[quote=William Cukierski;105489]\r\n\r\n[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes\r\n\r\n[/quote]",
    "105496": "[quote=William Cukierski;105489]\r\n\r\n[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes\r\n\r\n[/quote]\r\n\r\nHow can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?",
    "105502": "[quote=Maxim Milakov;105494]\r\n\r\nHi William,\r\n\r\nTimeline page - https://www.kaggle.com/c/second-annual-data-science-bowl/details/timeline - says \"Your model must be finalized and uploaded to Kaggle by this deadline\". This suggests that no further training should be done after March 7, moreover the model uploaded by March 7 should be used for generating predictions on the test set.\r\n\r\nAnd this makes sense. We can hardly expect to do the training for each patient at the hospital.\r\n\r\nThanks,\r\nMax.\r\n\r\n[quote=William Cukierski;105489]\r\n\r\n[quote=Imagic Mobile;105480]\r\n\r\nWhen the validation volumes are released 7th March, can those be used as part of the training for the final test set?\r\n\r\n[/quote]\r\n\r\nYes\r\n\r\n[/quote]\r\n\r\n[/quote]\r\n\r\nI actually have a different questions here. There is a large family of methods used in medical image analysis called multi-atlas segmentation. And the method does not have a regression or neural net model, but an algorithm with different components and a set of atlases.\r\n\r\nSo the questions is, do you have to submit a \"model\"?",
    "105516": "[quote=jstaker7;105496]\r\n\r\nHow can that be if our models must be finalized and submitted to Kaggle by the 7th? Perhaps I'm misunderstanding something?\r\n\r\n[/quote]\r\n\r\nFor our purposes, your \"model\" is the scientific approach for using a set of labeled images to make predictions on unlabeled images. While you are not allowed to modify the scientific approach during the final test phase of the competition, you are allowed to make changes such as:\r\n\r\n- retraining the model on a set that includes the train + validation set\r\n- pointing your code to the newly released test images\r\n- MINOR bug fixes for test set surprises\r\n\r\nWhy do we allow retraining on validation images? There are methods for \"reverse engineering\" labels from leaderboard scores (the current leaderboard uses 100% of the validation data). While we have no evidence of this happening in the competition, there is a potential that someone could reverse engineer the validation set labels and use them offline. To neutralize this potential unfair advantage, we release the labels and make the validation set fair game for training.",
    "109853": "hi \r\nI recently joined the bowl and i was wondering whether there will be continued development after the results . \r\nor the forums and data sets will be closed ?\r\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.",
    "111341": "Hello,\r\n\r\nCould you please tell me - is it mandatory to keep the same order of patients' labels in our submission file as they are in sample_submission (starting with 1000-1140 and then 701-999)? Or we can sort them like 701-1140 order?\r\n\r\nThanks in advance,\r\nAlex",
    "111359": "Any sort order works, as long as they're all there."
  },
  "source": "meta"
}