{
  "id": 19218,
  "title": "How hard is it really?",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19218",
  "author_name": "",
  "post_date": "2016-02-28T11:58:41.817Z",
  "votes": null,
  "comment_count": 20,
  "views": 2536,
  "content": "<p>Hello everyone!</p>\n\n<p>First of all, this topic is not directly related to the competition.</p>\n\n<p>I'm seeking the advice of the professionals on whether or not to choose the task of this competition as my masters thesis. I'm worried if the task is too advanced and I may not be able to do it in time (I only have a year), or may not be able to even come close to solving the problem at all! By the way,</p>\n\n<p>I have no background in image processing (I just took a course this semester) and I'm not an advanced data scientist, but I'm very enthusiastic and I'm willing to work day and night.</p>\n\n<p>Thank you all for your time and consideration.</p>",
  "messages": [
    {
      "id": "109599",
      "postDate": "02/28/2016 11:58:41",
      "content": "<p>Hello everyone!</p>\n\n<p>First of all, this topic is not directly related to the competition.</p>\n\n<p>I'm seeking the advice of the professionals on whether or not to choose the task of this competition as my masters thesis. I'm worried if the task is too advanced and I may not be able to do it in time (I only have a year), or may not be able to even come close to solving the problem at all! By the way,</p>\n\n<p>I have no background in image processing (I just took a course this semester) and I'm not an advanced data scientist, but I'm very enthusiastic and I'm willing to work day and night.</p>\n\n<p>Thank you all for your time and consideration.</p>",
      "rawMarkdown": "Hello everyone!\r\n\r\nFirst of all, this topic is not directly related to the competition.\r\n\r\nI'm seeking the advice of the professionals on whether or not to choose the task of this competition as my masters thesis. I'm worried if the task is too advanced and I may not be able to do it in time (I only have a year), or may not be able to even come close to solving the problem at all! By the way,\r\n\r\nI have no background in image processing (I just took a course this semester) and I'm not an advanced data scientist, but I'm very enthusiastic and I'm willing to work day and night.\r\n\r\nThank you all for your time and consideration.",
      "votes": null
    },
    {
      "id": "109602",
      "postDate": "02/28/2016 13:03:20",
      "content": "<p>In my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true.</p>\n\n<p>In my college, I had colleagues who solved problems very easily. But they rarely went beyond, and ironically even being able to solve problems more easily had a more superficial knowledge on the subject.</p>\n\n<p><strong>I believe the key ability is to understand the problem from a systematic point of view and put together solutions that apparently has no relation with each other.</strong></p>\n\n<p>For my part, I think that mathematical knowledge is essential to solve the problems, and I evaluate the facility of something by the amount of steps required to assemble the solution after it is understood.\nFrom my point of view, a model of 1000 lines is relatively easy, even if it was required 15 years of work to develop it.</p>\n\n<p>However,  do not forget of <strong>social and emotional skills</strong>, you will need to know how to deal with frustrations, failures and despairs.</p>\n\n<p>I also recommend other secondary skills that support their work style. In my case, understand the memory management allowed me to reduce processing time at 8x.\nI also recommend some knowledge on how to <strong>manage time</strong> and construction of software or model. It is relatively easy to make mistakes when you have millions of records or tens of versions of the same model.</p>\n\n<p>But this is my opinion, here you has others.</p>\n\n<p><a href=\"http://blog.kaggle.com/2016/02/22/profiling-top-kagglers-leustagos-current-7-highest-1/\">http://blog.kaggle.com/2016/02/22/profiling-top-kagglers-leustagos-current-7-highest-1/</a>\n<a href=\"http://blog.kaggle.com/2016/02/10/profiling-top-kagglers-kazanova-new-1-in-the-world/\">http://blog.kaggle.com/2016/02/10/profiling-top-kagglers-kazanova-new-1-in-the-world/</a>\n<a href=\"http://blog.kaggle.com/2016/02/04/noaa-right-whale-recognition-winners-interview-2nd-place-felix-lau/\">http://blog.kaggle.com/2016/02/04/noaa-right-whale-recognition-winners-interview-2nd-place-felix-lau/</a>\n<a href=\"http://blog.kaggle.com/2016/01/29/noaa-right-whale-recognition-winners-interview-1st-place-deepsense-io/\">http://blog.kaggle.com/2016/01/29/noaa-right-whale-recognition-winners-interview-1st-place-deepsense-io/</a></p>",
      "rawMarkdown": "In my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true.\r\n\r\nIn my college, I had colleagues who solved problems very easily. But they rarely went beyond, and ironically even being able to solve problems more easily had a more superficial knowledge on the subject.\r\n\r\n**I believe the key ability is to understand the problem from a systematic point of view and put together solutions that apparently has no relation with each other.**\r\n\r\nFor my part, I think that mathematical knowledge is essential to solve the problems, and I evaluate the facility of something by the amount of steps required to assemble the solution after it is understood.\r\nFrom my point of view, a model of 1000 lines is relatively easy, even if it was required 15 years of work to develop it.\r\n\r\nHowever,  do not forget of **social and emotional skills**, you will need to know how to deal with frustrations, failures and despairs.\r\n\r\nI also recommend other secondary skills that support their work style. In my case, understand the memory management allowed me to reduce processing time at 8x.\r\nI also recommend some knowledge on how to **manage time** and construction of software or model. It is relatively easy to make mistakes when you have millions of records or tens of versions of the same model.\r\n\r\nBut this is my opinion, here you has others.\r\n\r\nhttp://blog.kaggle.com/2016/02/22/profiling-top-kagglers-leustagos-current-7-highest-1/\r\nhttp://blog.kaggle.com/2016/02/10/profiling-top-kagglers-kazanova-new-1-in-the-world/\r\nhttp://blog.kaggle.com/2016/02/04/noaa-right-whale-recognition-winners-interview-2nd-place-felix-lau/\r\nhttp://blog.kaggle.com/2016/01/29/noaa-right-whale-recognition-winners-interview-1st-place-deepsense-io/",
      "votes": null
    },
    {
      "id": "109678",
      "postDate": "02/29/2016 11:16:08",
      "content": "<p>This contest is not easy, but I think it could be a good topic for the thesis.</p>",
      "rawMarkdown": "This contest is not easy, but I think it could be a good topic for the thesis.",
      "votes": null
    },
    {
      "id": "109688",
      "postDate": "02/29/2016 14:37:44",
      "content": "<p>[quote=Alvaro Osvaldo;109602]</p>\n\n<p>In my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true...</p>\n\n<p>[/quote]</p>\n\n<p>Thank you Alvaro for your reply. Honestly, the links were very inspiring specifically the interview with Leustagos.\nHowever, you didn't tell me your opinion about this competition being my masters thesis subject. I'm worried if it's too hard for a Masters thesis.</p>",
      "rawMarkdown": "[quote=Alvaro Osvaldo;109602]\r\n\r\nIn my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true...\r\n\r\n\r\n\r\n[/quote]\r\n\r\nThank you Alvaro for your reply. Honestly, the links were very inspiring specifically the interview with Leustagos.\r\nHowever, you didn't tell me your opinion about this competition being my masters thesis subject. I'm worried if it's too hard for a Masters thesis.",
      "votes": null
    },
    {
      "id": "109689",
      "postDate": "02/29/2016 14:48:14",
      "content": "<p>[quote=Jiming Ye;109678]</p>\n\n<p>This contest is not easy, but I think it could be a good topic for the thesis.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?</p>",
      "rawMarkdown": "[quote=Jiming Ye;109678]\r\n\r\nThis contest is not easy, but I think it could be a good topic for the thesis.\r\n\r\n[/quote]\r\n\r\nThank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?",
      "votes": null
    },
    {
      "id": "109740",
      "postDate": "02/29/2016 23:25:00",
      "content": "<p>Hi, The Ebili.</p>\n\n<p>In my opinion <strong>this competion is easy</strong>, in fact, i see this competition more easy than anothers based in numerical data.</p>\n\n<p>I think it because  a mathematical aproach i'm using.</p>\n\n<p>I believe so many people have so much problems because a mathematical gap between the techniques used and the mathematical nature of the problem. :O</p>",
      "rawMarkdown": "Hi, The Ebili.\r\n\r\nIn my opinion **this competion is easy**, in fact, i see this competition more easy than anothers based in numerical data.\r\n\r\nI think it because  a mathematical aproach i'm using.\r\n\r\nI believe so many people have so much problems because a mathematical gap between the techniques used and the mathematical nature of the problem. :O",
      "votes": null
    },
    {
      "id": "109854",
      "postDate": "03/01/2016 07:49:53",
      "content": "<p>hi \nI recently joined the bowl and i was wondering whether there will be continued development after the results . \nor the forums and data sets will be closed ?\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.</p>",
      "rawMarkdown": "hi \r\nI recently joined the bowl and i was wondering whether there will be continued development after the results . \r\nor the forums and data sets will be closed ?\r\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.",
      "votes": null
    },
    {
      "id": "109856",
      "postDate": "03/01/2016 08:38:05",
      "content": "<p>[quote=The Ebili;109689]</p>\n\n<p>Thank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?</p>\n\n<p>[/quote]</p>\n\n<p>LB is not everything. Our current LB just comes from tuning of the tutorial, it's not enough for a thesis. Our own methods are all failed, so I think this contest is not easy. However, it doesn't require a high performance machine. Just typical 4GB laptop with a decent GPU should be enough.</p>",
      "rawMarkdown": "[quote=The Ebili;109689]\r\n\r\nThank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?\r\n\r\n[/quote]\r\n\r\nLB is not everything. Our current LB just comes from tuning of the tutorial, it's not enough for a thesis. Our own methods are all failed, so I think this contest is not easy. However, it doesn't require a high performance machine. Just typical 4GB laptop with a decent GPU should be enough.",
      "votes": null
    },
    {
      "id": "109892",
      "postDate": "03/01/2016 12:28:16",
      "content": "<p>Hi Jiming,</p>\n\n<p>I'm the other way around - I couldn't get the mxnet tutorial to perform any better no matter what I tried, so I developed my own methods using image processing in R, and lasagne nolearn in Python. My latest submission doesn't use anything from the tutorials.</p>\n\n<p>Colin</p>",
      "rawMarkdown": "Hi Jiming,\r\n\r\nI'm the other way around - I couldn't get the mxnet tutorial to perform any better no matter what I tried, so I developed my own methods using image processing in R, and lasagne nolearn in Python. My latest submission doesn't use anything from the tutorials.\r\n\r\nColin",
      "votes": null
    },
    {
      "id": "109902",
      "postDate": "03/01/2016 13:17:24",
      "content": "<p>@waleedsial, i think that the leaderboard will remain open, like the anothers competitions, I believe the only difference will be the split in Public leaderboard and Private leaderboard.</p>\n\n<p>I believe that nobody prevents you to make a model even better of that will win this competition. :D</p>",
      "rawMarkdown": "waleedsial, i think that the leaderboard will remain open, like the anothers competitions, I believe the only difference will be the split in Public leaderboard and Private leaderboard.\r\n\r\nI believe that nobody prevents you to make a model even better of that will win this competition. :D",
      "votes": null
    },
    {
      "id": "109930",
      "postDate": "03/01/2016 15:39:08",
      "content": "<p>@Ebili:  I think a good related thesis topic would expand beyond the scope of this competition.  It may be difficult to gain acceptance of a black box (or at least very dark gray) returned by a convnet or to accept end-systolic and end-diastolic algorithm-returned volumes.  If I was a user (cardiologist), I would want visual confirmation that I had great information.  Just a couple thoughts:</p>\n\n<ul>\n<li>Based on a strong segmentation approach, could a dicom viewer clearly indicate the region identified as &quot;left-ventricle&quot; and the corresponding slice volume at end-systolic and end-diastolic states?  A user could then flip through slices and get visual confirmation that nothing odd is happening.</li>\n<li>In the case of erroneous segmentation, could a user make a parameter adjustment /modification to correct the region segmentation?  How could this play into an active learning scheme?</li>\n<li>Could the segments at end-systolic and end-diastolic be clearly rendered in 3D and give a fast indication of segmentation quality?</li>\n<li>Are there other diagnostic indicators that might be derived from the data and shared models?</li>\n</ul>",
      "rawMarkdown": "Ebili:  I think a good related thesis topic would expand beyond the scope of this competition.  It may be difficult to gain acceptance of a black box (or at least very dark gray) returned by a convnet or to accept end-systolic and end-diastolic algorithm-returned volumes.  If I was a user (cardiologist), I would want visual confirmation that I had great information.  Just a couple thoughts:\r\n\r\n - Based on a strong segmentation approach, could a dicom viewer clearly indicate the region identified as \"left-ventricle\" and the corresponding slice volume at end-systolic and end-diastolic states?  A user could then flip through slices and get visual confirmation that nothing odd is happening.\r\n - In the case of erroneous segmentation, could a user make a parameter adjustment /modification to correct the region segmentation?  How could this play into an active learning scheme?\r\n - Could the segments at end-systolic and end-diastolic be clearly rendered in 3D and give a fast indication of segmentation quality?\r\n - Are there other diagnostic indicators that might be derived from the data and shared models?",
      "votes": null
    },
    {
      "id": "110526",
      "postDate": "03/06/2016 09:44:48",
      "content": "<p>I found it useful and interesting to plot a 3-d graph of  slice areas, per slice over time, and a 2-d graph of LV volume over time. I imagine this must be clinically useful as well. For example attached.</p>",
      "rawMarkdown": "I found it useful and interesting to plot a 3-d graph of  slice areas, per slice over time, and a 2-d graph of LV volume over time. I imagine this must be clinically useful as well. For example attached.",
      "votes": null
    },
    {
      "id": "110533",
      "postDate": "03/06/2016 10:53:05",
      "content": "<p>Wow tracknut, you must have more processing power than me, or maybe you're using a more efficient algorithm. I didn't have the time to track and and every slice.</p>",
      "rawMarkdown": "Wow tracknut, you must have more processing power than me, or maybe you're using a more efficient algorithm. I didn't have the time to track and and every slice.",
      "votes": null
    },
    {
      "id": "110539",
      "postDate": "03/06/2016 11:17:39",
      "content": "<p>I don't think I did anything very special. My code is based on the Fourier tutorial, which calculates all the volumes. Saving to file takes a bit of time, and doing it on the full training set means a couple of hours on my Macbook Air, so I tend to use that sparsely.</p>\n\n<p>The 3-d graph was very useful in finding instances where identifying the slices at the base or apex went wrong. Also a nice smooth down-up volume graph must mean things are roughly OK </p>",
      "rawMarkdown": "I don't think I did anything very special. My code is based on the Fourier tutorial, which calculates all the volumes. Saving to file takes a bit of time, and doing it on the full training set means a couple of hours on my Macbook Air, so I tend to use that sparsely.\r\n\r\nThe 3-d graph was very useful in finding instances where identifying the slices at the base or apex went wrong. Also a nice smooth down-up volume graph must mean things are roughly OK",
      "votes": null
    },
    {
      "id": "110570",
      "postDate": "03/06/2016 18:15:19",
      "content": "<p><strong>A tip</strong>, if you divide all processing, for example, in 64 machines and get it done in 1 hour then it's likely that you process all files for between $ 10 to $ 100 depending on the instance in Amazon AWS.</p>",
      "rawMarkdown": "**A tip**, if you divide all processing, for example, in 64 machines and get it done in 1 hour then it's likely that you process all files for between $ 10 to $ 100 depending on the instance in Amazon AWS.",
      "votes": null
    },
    {
      "id": "110595",
      "postDate": "03/06/2016 21:48:15",
      "content": "<p>I could not even open dicom files and I regularly process satellite video images using python.</p>",
      "rawMarkdown": "I could not even open dicom files and I regularly process satellite video images using python.",
      "votes": null
    },
    {
      "id": "110596",
      "postDate": "03/06/2016 21:55:27",
      "content": "<p>Bad times &#176;.&#176; </p>\n\n<p>I used the ImageMagick to convert and extract the metadata.</p>\n\n<p>And a PHP script to organize all in a CSV data.</p>",
      "rawMarkdown": "Bad times °.° \r\n\r\nI used the ImageMagick to convert and extract the metadata.\r\n\r\nAnd a PHP script to organize all in a CSV data.",
      "votes": null
    },
    {
      "id": "110611",
      "postDate": "03/07/2016 00:28:24",
      "content": "<p>See my blog for an R script for opening and cleaning DICOM images. \n<a href=\"http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/\">http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/</a> </p>\n\n<p>Colin</p>",
      "rawMarkdown": "See my blog for an R script for opening and cleaning DICOM images. \r\nhttp://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/ \r\n\r\nColin",
      "votes": null
    },
    {
      "id": "110671",
      "postDate": "03/07/2016 13:09:46",
      "content": "<p>thanks colin that's great and I get to learn some R too</p>",
      "rawMarkdown": "thanks colin that's great and I get to learn some R too",
      "votes": null
    },
    {
      "id": "110673",
      "postDate": "03/07/2016 13:12:30",
      "content": "<p>As you may have noticed, I'm much more comfortable writing R than Python. But my next blog will have Python code because it uses theano / lasagne / nolearn which aren't available in R.</p>",
      "rawMarkdown": "As you may have noticed, I'm much more comfortable writing R than Python. But my next blog will have Python code because it uses theano / lasagne / nolearn which aren't available in R.",
      "votes": null
    },
    {
      "id": "110677",
      "postDate": "03/07/2016 13:26:14",
      "content": "<p>Nice Colin :D </p>",
      "rawMarkdown": "Nice Colin :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 109602,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "02/28/2016 13:03:20",
      "content": "<p>In my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true.</p>\n\n<p>In my college, I had colleagues who solved problems very easily. But they rarely went beyond, and ironically even being able to solve problems more easily had a more superficial knowledge on the subject.</p>\n\n<p><strong>I believe the key ability is to understand the problem from a systematic point of view and put together solutions that apparently has no relation with each other.</strong></p>\n\n<p>For my part, I think that mathematical knowledge is essential to solve the problems, and I evaluate the facility of something by the amount of steps required to assemble the solution after it is understood.\nFrom my point of view, a model of 1000 lines is relatively easy, even if it was required 15 years of work to develop it.</p>\n\n<p>However,  do not forget of <strong>social and emotional skills</strong>, you will need to know how to deal with frustrations, failures and despairs.</p>\n\n<p>I also recommend other secondary skills that support their work style. In my case, understand the memory management allowed me to reduce processing time at 8x.\nI also recommend some knowledge on how to <strong>manage time</strong> and construction of software or model. It is relatively easy to make mistakes when you have millions of records or tens of versions of the same model.</p>\n\n<p>But this is my opinion, here you has others.</p>\n\n<p><a href=\"http://blog.kaggle.com/2016/02/22/profiling-top-kagglers-leustagos-current-7-highest-1/\">http://blog.kaggle.com/2016/02/22/profiling-top-kagglers-leustagos-current-7-highest-1/</a>\n<a href=\"http://blog.kaggle.com/2016/02/10/profiling-top-kagglers-kazanova-new-1-in-the-world/\">http://blog.kaggle.com/2016/02/10/profiling-top-kagglers-kazanova-new-1-in-the-world/</a>\n<a href=\"http://blog.kaggle.com/2016/02/04/noaa-right-whale-recognition-winners-interview-2nd-place-felix-lau/\">http://blog.kaggle.com/2016/02/04/noaa-right-whale-recognition-winners-interview-2nd-place-felix-lau/</a>\n<a href=\"http://blog.kaggle.com/2016/01/29/noaa-right-whale-recognition-winners-interview-1st-place-deepsense-io/\">http://blog.kaggle.com/2016/01/29/noaa-right-whale-recognition-winners-interview-1st-place-deepsense-io/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109678,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "02/29/2016 11:16:08",
      "content": "<p>This contest is not easy, but I think it could be a good topic for the thesis.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109688,
      "author_name": "mheybpoodh",
      "author_url": "",
      "post_date": "02/29/2016 14:37:44",
      "content": "<p>[quote=Alvaro Osvaldo;109602]</p>\n\n<p>In my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true...</p>\n\n<p>[/quote]</p>\n\n<p>Thank you Alvaro for your reply. Honestly, the links were very inspiring specifically the interview with Leustagos.\nHowever, you didn't tell me your opinion about this competition being my masters thesis subject. I'm worried if it's too hard for a Masters thesis.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109689,
      "author_name": "mheybpoodh",
      "author_url": "",
      "post_date": "02/29/2016 14:48:14",
      "content": "<p>[quote=Jiming Ye;109678]</p>\n\n<p>This contest is not easy, but I think it could be a good topic for the thesis.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109740,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "02/29/2016 23:25:00",
      "content": "<p>Hi, The Ebili.</p>\n\n<p>In my opinion <strong>this competion is easy</strong>, in fact, i see this competition more easy than anothers based in numerical data.</p>\n\n<p>I think it because  a mathematical aproach i'm using.</p>\n\n<p>I believe so many people have so much problems because a mathematical gap between the techniques used and the mathematical nature of the problem. :O</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109854,
      "author_name": "waleedsial",
      "author_url": "",
      "post_date": "03/01/2016 07:49:53",
      "content": "<p>hi \nI recently joined the bowl and i was wondering whether there will be continued development after the results . \nor the forums and data sets will be closed ?\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109856,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "03/01/2016 08:38:05",
      "content": "<p>[quote=The Ebili;109689]</p>\n\n<p>Thank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?</p>\n\n<p>[/quote]</p>\n\n<p>LB is not everything. Our current LB just comes from tuning of the tutorial, it's not enough for a thesis. Our own methods are all failed, so I think this contest is not easy. However, it doesn't require a high performance machine. Just typical 4GB laptop with a decent GPU should be enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109892,
      "author_name": "colinpriest",
      "author_url": "",
      "post_date": "03/01/2016 12:28:16",
      "content": "<p>Hi Jiming,</p>\n\n<p>I'm the other way around - I couldn't get the mxnet tutorial to perform any better no matter what I tried, so I developed my own methods using image processing in R, and lasagne nolearn in Python. My latest submission doesn't use anything from the tutorials.</p>\n\n<p>Colin</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109902,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "03/01/2016 13:17:24",
      "content": "<p>@waleedsial, i think that the leaderboard will remain open, like the anothers competitions, I believe the only difference will be the split in Public leaderboard and Private leaderboard.</p>\n\n<p>I believe that nobody prevents you to make a model even better of that will win this competition. :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 109930,
      "author_name": "scsmith2",
      "author_url": "",
      "post_date": "03/01/2016 15:39:08",
      "content": "<p>@Ebili:  I think a good related thesis topic would expand beyond the scope of this competition.  It may be difficult to gain acceptance of a black box (or at least very dark gray) returned by a convnet or to accept end-systolic and end-diastolic algorithm-returned volumes.  If I was a user (cardiologist), I would want visual confirmation that I had great information.  Just a couple thoughts:</p>\n\n<ul>\n<li>Based on a strong segmentation approach, could a dicom viewer clearly indicate the region identified as &quot;left-ventricle&quot; and the corresponding slice volume at end-systolic and end-diastolic states?  A user could then flip through slices and get visual confirmation that nothing odd is happening.</li>\n<li>In the case of erroneous segmentation, could a user make a parameter adjustment /modification to correct the region segmentation?  How could this play into an active learning scheme?</li>\n<li>Could the segments at end-systolic and end-diastolic be clearly rendered in 3D and give a fast indication of segmentation quality?</li>\n<li>Are there other diagnostic indicators that might be derived from the data and shared models?</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110526,
      "author_name": "tracknut",
      "author_url": "",
      "post_date": "03/06/2016 09:44:48",
      "content": "<p>I found it useful and interesting to plot a 3-d graph of  slice areas, per slice over time, and a 2-d graph of LV volume over time. I imagine this must be clinically useful as well. For example attached.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110533,
      "author_name": "colinpriest",
      "author_url": "",
      "post_date": "03/06/2016 10:53:05",
      "content": "<p>Wow tracknut, you must have more processing power than me, or maybe you're using a more efficient algorithm. I didn't have the time to track and and every slice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110539,
      "author_name": "tracknut",
      "author_url": "",
      "post_date": "03/06/2016 11:17:39",
      "content": "<p>I don't think I did anything very special. My code is based on the Fourier tutorial, which calculates all the volumes. Saving to file takes a bit of time, and doing it on the full training set means a couple of hours on my Macbook Air, so I tend to use that sparsely.</p>\n\n<p>The 3-d graph was very useful in finding instances where identifying the slices at the base or apex went wrong. Also a nice smooth down-up volume graph must mean things are roughly OK </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110570,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "03/06/2016 18:15:19",
      "content": "<p><strong>A tip</strong>, if you divide all processing, for example, in 64 machines and get it done in 1 hour then it's likely that you process all files for between $ 10 to $ 100 depending on the instance in Amazon AWS.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110595,
      "author_name": "",
      "author_url": "",
      "post_date": "03/06/2016 21:48:15",
      "content": "<p>I could not even open dicom files and I regularly process satellite video images using python.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110596,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "03/06/2016 21:55:27",
      "content": "<p>Bad times &#176;.&#176; </p>\n\n<p>I used the ImageMagick to convert and extract the metadata.</p>\n\n<p>And a PHP script to organize all in a CSV data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110611,
      "author_name": "colinpriest",
      "author_url": "",
      "post_date": "03/07/2016 00:28:24",
      "content": "<p>See my blog for an R script for opening and cleaning DICOM images. \n<a href=\"http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/\">http://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/</a> </p>\n\n<p>Colin</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110671,
      "author_name": "",
      "author_url": "",
      "post_date": "03/07/2016 13:09:46",
      "content": "<p>thanks colin that's great and I get to learn some R too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110673,
      "author_name": "colinpriest",
      "author_url": "",
      "post_date": "03/07/2016 13:12:30",
      "content": "<p>As you may have noticed, I'm much more comfortable writing R than Python. But my next blog will have Python code because it uses theano / lasagne / nolearn which aren't available in R.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 110677,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "03/07/2016 13:26:14",
      "content": "<p>Nice Colin :D </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "109599": "Hello everyone!\r\n\r\nFirst of all, this topic is not directly related to the competition.\r\n\r\nI'm seeking the advice of the professionals on whether or not to choose the task of this competition as my masters thesis. I'm worried if the task is too advanced and I may not be able to do it in time (I only have a year), or may not be able to even come close to solving the problem at all! By the way,\r\n\r\nI have no background in image processing (I just took a course this semester) and I'm not an advanced data scientist, but I'm very enthusiastic and I'm willing to work day and night.\r\n\r\nThank you all for your time and consideration.",
    "109602": "In my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true.\r\n\r\nIn my college, I had colleagues who solved problems very easily. But they rarely went beyond, and ironically even being able to solve problems more easily had a more superficial knowledge on the subject.\r\n\r\n**I believe the key ability is to understand the problem from a systematic point of view and put together solutions that apparently has no relation with each other.**\r\n\r\nFor my part, I think that mathematical knowledge is essential to solve the problems, and I evaluate the facility of something by the amount of steps required to assemble the solution after it is understood.\r\nFrom my point of view, a model of 1000 lines is relatively easy, even if it was required 15 years of work to develop it.\r\n\r\nHowever,  do not forget of **social and emotional skills**, you will need to know how to deal with frustrations, failures and despairs.\r\n\r\nI also recommend other secondary skills that support their work style. In my case, understand the memory management allowed me to reduce processing time at 8x.\r\nI also recommend some knowledge on how to **manage time** and construction of software or model. It is relatively easy to make mistakes when you have millions of records or tens of versions of the same model.\r\n\r\nBut this is my opinion, here you has others.\r\n\r\nhttp://blog.kaggle.com/2016/02/22/profiling-top-kagglers-leustagos-current-7-highest-1/\r\nhttp://blog.kaggle.com/2016/02/10/profiling-top-kagglers-kazanova-new-1-in-the-world/\r\nhttp://blog.kaggle.com/2016/02/04/noaa-right-whale-recognition-winners-interview-2nd-place-felix-lau/\r\nhttp://blog.kaggle.com/2016/01/29/noaa-right-whale-recognition-winners-interview-1st-place-deepsense-io/",
    "109678": "This contest is not easy, but I think it could be a good topic for the thesis.",
    "109688": "[quote=Alvaro Osvaldo;109602]\r\n\r\nIn my opnion, easy and hard are point of views. And the hability to resolve something easly not allways is associated with a good production. The opposite can also be true...\r\n\r\n\r\n\r\n[/quote]\r\n\r\nThank you Alvaro for your reply. Honestly, the links were very inspiring specifically the interview with Leustagos.\r\nHowever, you didn't tell me your opinion about this competition being my masters thesis subject. I'm worried if it's too hard for a Masters thesis.",
    "109689": "[quote=Jiming Ye;109678]\r\n\r\nThis contest is not easy, but I think it could be a good topic for the thesis.\r\n\r\n[/quote]\r\n\r\nThank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?",
    "109740": "Hi, The Ebili.\r\n\r\nIn my opinion **this competion is easy**, in fact, i see this competition more easy than anothers based in numerical data.\r\n\r\nI think it because  a mathematical aproach i'm using.\r\n\r\nI believe so many people have so much problems because a mathematical gap between the techniques used and the mathematical nature of the problem. :O",
    "109854": "hi \r\nI recently joined the bowl and i was wondering whether there will be continued development after the results . \r\nor the forums and data sets will be closed ?\r\nI am newbie and i was trying to do some research on this topic but as it seems i will be late by the reults date.",
    "109856": "[quote=The Ebili;109689]\r\n\r\nThank you Jiming for your answer. I checked you on the leaderboard and you're doing really good, so I decided to bother you a little bit more and ask your opinion on the hardware that is required for this? What kind of hardware are you using (if I may ask)?\r\n\r\n[/quote]\r\n\r\nLB is not everything. Our current LB just comes from tuning of the tutorial, it's not enough for a thesis. Our own methods are all failed, so I think this contest is not easy. However, it doesn't require a high performance machine. Just typical 4GB laptop with a decent GPU should be enough.",
    "109892": "Hi Jiming,\r\n\r\nI'm the other way around - I couldn't get the mxnet tutorial to perform any better no matter what I tried, so I developed my own methods using image processing in R, and lasagne nolearn in Python. My latest submission doesn't use anything from the tutorials.\r\n\r\nColin",
    "109902": "waleedsial, i think that the leaderboard will remain open, like the anothers competitions, I believe the only difference will be the split in Public leaderboard and Private leaderboard.\r\n\r\nI believe that nobody prevents you to make a model even better of that will win this competition. :D",
    "109930": "Ebili:  I think a good related thesis topic would expand beyond the scope of this competition.  It may be difficult to gain acceptance of a black box (or at least very dark gray) returned by a convnet or to accept end-systolic and end-diastolic algorithm-returned volumes.  If I was a user (cardiologist), I would want visual confirmation that I had great information.  Just a couple thoughts:\r\n\r\n - Based on a strong segmentation approach, could a dicom viewer clearly indicate the region identified as \"left-ventricle\" and the corresponding slice volume at end-systolic and end-diastolic states?  A user could then flip through slices and get visual confirmation that nothing odd is happening.\r\n - In the case of erroneous segmentation, could a user make a parameter adjustment /modification to correct the region segmentation?  How could this play into an active learning scheme?\r\n - Could the segments at end-systolic and end-diastolic be clearly rendered in 3D and give a fast indication of segmentation quality?\r\n - Are there other diagnostic indicators that might be derived from the data and shared models?",
    "110526": "I found it useful and interesting to plot a 3-d graph of  slice areas, per slice over time, and a 2-d graph of LV volume over time. I imagine this must be clinically useful as well. For example attached.",
    "110533": "Wow tracknut, you must have more processing power than me, or maybe you're using a more efficient algorithm. I didn't have the time to track and and every slice.",
    "110539": "I don't think I did anything very special. My code is based on the Fourier tutorial, which calculates all the volumes. Saving to file takes a bit of time, and doing it on the full training set means a couple of hours on my Macbook Air, so I tend to use that sparsely.\r\n\r\nThe 3-d graph was very useful in finding instances where identifying the slices at the base or apex went wrong. Also a nice smooth down-up volume graph must mean things are roughly OK",
    "110570": "**A tip**, if you divide all processing, for example, in 64 machines and get it done in 1 hour then it's likely that you process all files for between $ 10 to $ 100 depending on the instance in Amazon AWS.",
    "110595": "I could not even open dicom files and I regularly process satellite video images using python.",
    "110596": "Bad times °.° \r\n\r\nI used the ImageMagick to convert and extract the metadata.\r\n\r\nAnd a PHP script to organize all in a CSV data.",
    "110611": "See my blog for an R script for opening and cleaning DICOM images. \r\nhttp://colinpriest.com/2016/03/07/second-annual-data-science-bowl-part-2/ \r\n\r\nColin",
    "110671": "thanks colin that's great and I get to learn some R too",
    "110673": "As you may have noticed, I'm much more comfortable writing R than Python. But my next blog will have Python code because it uses theano / lasagne / nolearn which aren't available in R.",
    "110677": "Nice Colin :D"
  },
  "source": "meta"
}