{
  "id": 15631,
  "title": "Critique of study design",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/15631",
  "author_name": "",
  "post_date": "2015-07-29T22:58:25.527Z",
  "votes": 3,
  "comment_count": 6,
  "views": 1922,
  "content": "<p>Because there are a lot of beginners who want to use this competition to learn about data science and medical image analysis, I think a critique of it as a scientific study  is needed. I have a background in medical imaging and have published many papers in the field and also directed a research project on computer-aided diagnosis (CAD). I think the study design has very serious problems and here are my reasons.</p>\n\n<p>The gold standard:\nAccording to the Data page, the goal is  &quot;to create an automated analysis system capable of assigning a score based on this scale.&quot; The scale, that is, the gold standard of the study, was a clinician's rating on a scale from 0 to 4. </p>\n\n<p>The gold standard is the most important part of the study. If the gold standard is not reliable, then the results of the study are meaningless. You may claim that the standard is what it is and an algorithm to replicate it has some use. But that is like claiming that replicating astrological forecasts is useful. </p>\n\n<p>The strongest gold standard is based on objective, reproducible evidence. For example, if you are reading chest x-ray  images to detect a lung nodule, the strongest standard is a direct biopsy that extracts the nodule. The next strongest is a more accurate examination that may be more expensive and/or invasive than the one you are studying: in the case of the lung nodules a thoracic computed tomography scan or for diabetic retinopathy, perhaps an optical coherence tomography scan. </p>\n\n<p>The next weakest gold standard is a panel of experts who diagnose your images and you combine their readings by averaging or majority vote for example. These are obviously subjective and in my experience there can be a huge difference between different experts and experts at different institutions. </p>\n\n<p>By far the weakest is the one used in this study; a single expert. This is almost totally subjective and worse there seems to have been no or little effort to spot check or do any quality assurance on the readings. For example, see my post on this forum on unexplainable ratings for blank or garbage images. </p>\n\n<p>The rating scale:\nThe study forces the rating to be one of the five levels 0 to 4. There is no &quot;not rated&quot; level for images that are too poor quality for a reliable diagnosis. In medical imaging you always need this option because it is unfair to the patient to provide a rating on a non-diagnostic quality image. If you rate unacceptable images as 0 you give the patient a false result if there is indeed disease. A rating of disease may lead to unnecessary and possible harmful additional tests. </p>\n\n<p>The image database:\nA good aspect of the study is the very large number of images. Unfortunately, this was achieved by combining images from many different sources some of which seemed to have poor quality. The medical system is analogous to a manufacturing process where you take images of the patients, process them through the system and produce a diagnosis. Any manufacturing engineer will tell you that you achieve the most reliable results by standardizing the inputs. The study takes the opposite approach and tries to make the medical system compensate for variations in inputs. I think this is a mistake.</p>\n\n<p>The contest organizers shrug off image quality by saying:\n    Like any real-world data set, you will encounter noise in both the images and labels. Images may contain artifacts, be out of focus, underexposed, or overexposed.  A major aim of this competition is to develop robust algorithms that can function in the presence of noise and variation.</p>\n\n<p>I disagree with this approach. I have spent a large amount of effort trying to compensate for poor image quality with software and I am sorry to say most of this effort was wasted. Companies that put their efforts into improving the image acquisition hardware and method instead of compensating for it with software always came out ahead. </p>\n\n<p>As an example, consider the Hubbell space telescope. You may recall that due to a monumental screw-up when the first images were made it was discovered that the main mirror focal length was incorrect resulting in blurred images. Image blur is very well understood with excellent mathematical models of its effects and many algorithms have been developed to &quot;de-blur&quot; images. Also NASA had unlimited computing power available to process the images. But they recognized that this would not provide the image quality required for their research so they risked the lives of astronauts and spent hundreds of millions of dollars to fly a mission to the satellite to replace the mirror. </p>\n\n<p>The image database used JPEG compressed data. The use of a lossy data compression algorithm like JPEG, which was designed to produce visually pleasing snapshots of people, would not be acceptable in other areas of medical imaging. The cost of data storage is so low that I do not think there is a justification for using lossy compression of medical data.</p>\n\n<p>If as a data analyst you are faced with poor quality data your first job is to improve the data acquisition. If your employer refuses to do this then let them know that this will severely impact the output. If they still refuse to improve it, consider going somewhere else because you will almost certainly get the blame when the project fails. </p>\n\n<p>Other data:\nAn advantage of computer-aided methods over human observers is the ability to handle additional data. In a study such as this one, something as simple as knowing the make and model of the camera hardware can improve the results dramatically. In addition, having the age, sex, fasting blood glucose level, BMI of the patient and so on could lead to much more reliable results. Part of the job of a data analyst is to get as much relevant data as you can particularly if it is already available on some other database as these patient data almost certainly are. </p>\n\n<p>Conclusion:\nI applaud the organizers for producing this contest that introduced many data analysts to medical image analysis. I think it is an important area that can benefit from the application of advanced data analysis techniques. However, beginners should be aware of the study limitations so they can improve their chances of success if they decide to continue with the field.  </p>",
  "messages": [
    {
      "id": "87511",
      "postDate": "07/29/2015 22:58:25",
      "content": "<p>Because there are a lot of beginners who want to use this competition to learn about data science and medical image analysis, I think a critique of it as a scientific study  is needed. I have a background in medical imaging and have published many papers in the field and also directed a research project on computer-aided diagnosis (CAD). I think the study design has very serious problems and here are my reasons.</p>\n\n<p>The gold standard:\nAccording to the Data page, the goal is  &quot;to create an automated analysis system capable of assigning a score based on this scale.&quot; The scale, that is, the gold standard of the study, was a clinician's rating on a scale from 0 to 4. </p>\n\n<p>The gold standard is the most important part of the study. If the gold standard is not reliable, then the results of the study are meaningless. You may claim that the standard is what it is and an algorithm to replicate it has some use. But that is like claiming that replicating astrological forecasts is useful. </p>\n\n<p>The strongest gold standard is based on objective, reproducible evidence. For example, if you are reading chest x-ray  images to detect a lung nodule, the strongest standard is a direct biopsy that extracts the nodule. The next strongest is a more accurate examination that may be more expensive and/or invasive than the one you are studying: in the case of the lung nodules a thoracic computed tomography scan or for diabetic retinopathy, perhaps an optical coherence tomography scan. </p>\n\n<p>The next weakest gold standard is a panel of experts who diagnose your images and you combine their readings by averaging or majority vote for example. These are obviously subjective and in my experience there can be a huge difference between different experts and experts at different institutions. </p>\n\n<p>By far the weakest is the one used in this study; a single expert. This is almost totally subjective and worse there seems to have been no or little effort to spot check or do any quality assurance on the readings. For example, see my post on this forum on unexplainable ratings for blank or garbage images. </p>\n\n<p>The rating scale:\nThe study forces the rating to be one of the five levels 0 to 4. There is no &quot;not rated&quot; level for images that are too poor quality for a reliable diagnosis. In medical imaging you always need this option because it is unfair to the patient to provide a rating on a non-diagnostic quality image. If you rate unacceptable images as 0 you give the patient a false result if there is indeed disease. A rating of disease may lead to unnecessary and possible harmful additional tests. </p>\n\n<p>The image database:\nA good aspect of the study is the very large number of images. Unfortunately, this was achieved by combining images from many different sources some of which seemed to have poor quality. The medical system is analogous to a manufacturing process where you take images of the patients, process them through the system and produce a diagnosis. Any manufacturing engineer will tell you that you achieve the most reliable results by standardizing the inputs. The study takes the opposite approach and tries to make the medical system compensate for variations in inputs. I think this is a mistake.</p>\n\n<p>The contest organizers shrug off image quality by saying:\n    Like any real-world data set, you will encounter noise in both the images and labels. Images may contain artifacts, be out of focus, underexposed, or overexposed.  A major aim of this competition is to develop robust algorithms that can function in the presence of noise and variation.</p>\n\n<p>I disagree with this approach. I have spent a large amount of effort trying to compensate for poor image quality with software and I am sorry to say most of this effort was wasted. Companies that put their efforts into improving the image acquisition hardware and method instead of compensating for it with software always came out ahead. </p>\n\n<p>As an example, consider the Hubbell space telescope. You may recall that due to a monumental screw-up when the first images were made it was discovered that the main mirror focal length was incorrect resulting in blurred images. Image blur is very well understood with excellent mathematical models of its effects and many algorithms have been developed to &quot;de-blur&quot; images. Also NASA had unlimited computing power available to process the images. But they recognized that this would not provide the image quality required for their research so they risked the lives of astronauts and spent hundreds of millions of dollars to fly a mission to the satellite to replace the mirror. </p>\n\n<p>The image database used JPEG compressed data. The use of a lossy data compression algorithm like JPEG, which was designed to produce visually pleasing snapshots of people, would not be acceptable in other areas of medical imaging. The cost of data storage is so low that I do not think there is a justification for using lossy compression of medical data.</p>\n\n<p>If as a data analyst you are faced with poor quality data your first job is to improve the data acquisition. If your employer refuses to do this then let them know that this will severely impact the output. If they still refuse to improve it, consider going somewhere else because you will almost certainly get the blame when the project fails. </p>\n\n<p>Other data:\nAn advantage of computer-aided methods over human observers is the ability to handle additional data. In a study such as this one, something as simple as knowing the make and model of the camera hardware can improve the results dramatically. In addition, having the age, sex, fasting blood glucose level, BMI of the patient and so on could lead to much more reliable results. Part of the job of a data analyst is to get as much relevant data as you can particularly if it is already available on some other database as these patient data almost certainly are. </p>\n\n<p>Conclusion:\nI applaud the organizers for producing this contest that introduced many data analysts to medical image analysis. I think it is an important area that can benefit from the application of advanced data analysis techniques. However, beginners should be aware of the study limitations so they can improve their chances of success if they decide to continue with the field.  </p>",
      "rawMarkdown": "Because there are a lot of beginners who want to use this competition to learn about data science and medical image analysis, I think a critique of it as a scientific study  is needed. I have a background in medical imaging and have published many papers in the field and also directed a research project on computer-aided diagnosis (CAD). I think the study design has very serious problems and here are my reasons.\r\n\r\nThe gold standard:\r\nAccording to the Data page, the goal is  \"to create an automated analysis system capable of assigning a score based on this scale.\" The scale, that is, the gold standard of the study, was a clinician's rating on a scale from 0 to 4. \r\n\r\nThe gold standard is the most important part of the study. If the gold standard is not reliable, then the results of the study are meaningless. You may claim that the standard is what it is and an algorithm to replicate it has some use. But that is like claiming that replicating astrological forecasts is useful. \r\n\r\nThe strongest gold standard is based on objective, reproducible evidence. For example, if you are reading chest x-ray  images to detect a lung nodule, the strongest standard is a direct biopsy that extracts the nodule. The next strongest is a more accurate examination that may be more expensive and/or invasive than the one you are studying: in the case of the lung nodules a thoracic computed tomography scan or for diabetic retinopathy, perhaps an optical coherence tomography scan. \r\n\r\nThe next weakest gold standard is a panel of experts who diagnose your images and you combine their readings by averaging or majority vote for example. These are obviously subjective and in my experience there can be a huge difference between different experts and experts at different institutions. \r\n\r\nBy far the weakest is the one used in this study; a single expert. This is almost totally subjective and worse there seems to have been no or little effort to spot check or do any quality assurance on the readings. For example, see my post on this forum on unexplainable ratings for blank or garbage images. \r\n\r\nThe rating scale:\r\nThe study forces the rating to be one of the five levels 0 to 4. There is no \"not rated\" level for images that are too poor quality for a reliable diagnosis. In medical imaging you always need this option because it is unfair to the patient to provide a rating on a non-diagnostic quality image. If you rate unacceptable images as 0 you give the patient a false result if there is indeed disease. A rating of disease may lead to unnecessary and possible harmful additional tests. \r\n\r\nThe image database:\r\nA good aspect of the study is the very large number of images. Unfortunately, this was achieved by combining images from many different sources some of which seemed to have poor quality. The medical system is analogous to a manufacturing process where you take images of the patients, process them through the system and produce a diagnosis. Any manufacturing engineer will tell you that you achieve the most reliable results by standardizing the inputs. The study takes the opposite approach and tries to make the medical system compensate for variations in inputs. I think this is a mistake.\r\n\r\nThe contest organizers shrug off image quality by saying:\r\n\tLike any real-world data set, you will encounter noise in both the images and labels. Images may contain artifacts, be out of focus, underexposed, or overexposed.  A major aim of this competition is to develop robust algorithms that can function in the presence of noise and variation.\r\n\r\nI disagree with this approach. I have spent a large amount of effort trying to compensate for poor image quality with software and I am sorry to say most of this effort was wasted. Companies that put their efforts into improving the image acquisition hardware and method instead of compensating for it with software always came out ahead. \r\n\r\nAs an example, consider the Hubbell space telescope. You may recall that due to a monumental screw-up when the first images were made it was discovered that the main mirror focal length was incorrect resulting in blurred images. Image blur is very well understood with excellent mathematical models of its effects and many algorithms have been developed to \"de-blur\" images. Also NASA had unlimited computing power available to process the images. But they recognized that this would not provide the image quality required for their research so they risked the lives of astronauts and spent hundreds of millions of dollars to fly a mission to the satellite to replace the mirror. \r\n\r\nThe image database used JPEG compressed data. The use of a lossy data compression algorithm like JPEG, which was designed to produce visually pleasing snapshots of people, would not be acceptable in other areas of medical imaging. The cost of data storage is so low that I do not think there is a justification for using lossy compression of medical data.\r\n\r\nIf as a data analyst you are faced with poor quality data your first job is to improve the data acquisition. If your employer refuses to do this then let them know that this will severely impact the output. If they still refuse to improve it, consider going somewhere else because you will almost certainly get the blame when the project fails. \r\n\r\nOther data:\r\nAn advantage of computer-aided methods over human observers is the ability to handle additional data. In a study such as this one, something as simple as knowing the make and model of the camera hardware can improve the results dramatically. In addition, having the age, sex, fasting blood glucose level, BMI of the patient and so on could lead to much more reliable results. Part of the job of a data analyst is to get as much relevant data as you can particularly if it is already available on some other database as these patient data almost certainly are. \r\n\r\nConclusion:\r\nI applaud the organizers for producing this contest that introduced many data analysts to medical image analysis. I think it is an important area that can benefit from the application of advanced data analysis techniques. However, beginners should be aware of the study limitations so they can improve their chances of success if they decide to continue with the field.",
      "votes": null
    },
    {
      "id": "87667",
      "postDate": "07/30/2015 20:31:34",
      "content": "<p>I disagree with most of this post.</p>\n\n<p>The one part I do agree with is having another label of &quot;Not Rated&quot; and/or &quot;Bad Input&quot;.  </p>\n\n<p>Was only a single expert used for the dataset?  I admit that this would be a weak spot as the algorithm would learn to rate images like that single person.  Multiple people labeling disjoint datasets should be fine, which is what I thought they had used.</p>\n\n<p>Having multiple sources of data (ie different imaging machines) is actually quite good.  As long as enough information is available to distinguish each level, then any variation will lead to a more robust outcome.  Plus it is way too expensive to replace all of the machines with 'good' equipment.  Having a label of &quot;Not rated&quot; would also help with determining if the data was good enough to diagnose or not. </p>\n\n<p>Lossy compression is actually crucial to getting better results.  Of course almost everyone reduced the data size massively for their models, so it probably didn't matter.  But if you tried to distribute 100,000 images in less than 200GB of space, you will get a lot more information if you use lossy compression.</p>\n\n<p>Finally, this algorithm is just one component.  Ideally you would like to see a progression of data for each patient and any and all related outcomes with each patient.  But I doubt that data is collected yet.</p>",
      "rawMarkdown": "I disagree with most of this post.\r\n\r\nThe one part I do agree with is having another label of \"Not Rated\" and/or \"Bad Input\".  \r\n\r\nWas only a single expert used for the dataset?  I admit that this would be a weak spot as the algorithm would learn to rate images like that single person.  Multiple people labeling disjoint datasets should be fine, which is what I thought they had used.\r\n\r\nHaving multiple sources of data (ie different imaging machines) is actually quite good.  As long as enough information is available to distinguish each level, then any variation will lead to a more robust outcome.  Plus it is way too expensive to replace all of the machines with 'good' equipment.  Having a label of \"Not rated\" would also help with determining if the data was good enough to diagnose or not. \r\n\r\nLossy compression is actually crucial to getting better results.  Of course almost everyone reduced the data size massively for their models, so it probably didn't matter.  But if you tried to distribute 100,000 images in less than 200GB of space, you will get a lot more information if you use lossy compression.\r\n\r\nFinally, this algorithm is just one component.  Ideally you would like to see a progression of data for each patient and any and all related outcomes with each patient.  But I doubt that data is collected yet.",
      "votes": null
    },
    {
      "id": "87759",
      "postDate": "07/31/2015 14:19:05",
      "content": "<p>I found bobalv's post to be fascinating, as it shows steps that can be taken to increase accuracy.  But, I agree with Ryan with the simple argument that given two models that match a gold standards accuracy, take the simpler one.   And more data, even with noise and redundancy removed, is far more important than analyzing a perfect image sensor reproduction.</p>",
      "rawMarkdown": "I found bobalv's post to be fascinating, as it shows steps that can be taken to increase accuracy.  But, I agree with Ryan with the simple argument that given two models that match a gold standards accuracy, take the simpler one.   And more data, even with noise and redundancy removed, is far more important than analyzing a perfect image sensor reproduction.",
      "votes": null
    },
    {
      "id": "87763",
      "postDate": "07/31/2015 15:05:58",
      "content": "<p>I also have a CAD background (and helped set up the competition :)</p>\n\n<p>Bear in that the problems we run are real-world problems subject to real-world constraints. The ideal diagnostic tool may indeed be the latest camera with the best lighting conditions using uncompressed tiffs and be trained on a panel of ratings and independently validated. In reality, you always make concessions on study size, design, cost, data availability.</p>\n\n<p>A major motivation of this study was to see how algorithms could function in messy, non-ideal conditions. Here, this meant a variety of camera conditions. In the future, it might be cell phone cameras, low cost equipment in pharmacies, old systems donated to clinics in needy areas. The use there is more a pragmatic screening tool (&quot;there's a good chance you need to see a doctor now&quot;) vs a diagnostic tool (&quot;we are certain you have a very serious disease and the FDA approval to say so&quot;).</p>\n\n<p>Footnote - we generally don't have a &quot;not rated&quot; category because it's not a natural fit for a competition format with ordinal/regression predictions (the &quot;distance&quot; between &quot;not rated&quot; and 0 doesn't have the same units as the &quot;distance&quot; between 0 and 1). In a production system, you would indeed need sanity checking and heuristics to ensure the algorithm isn't predicting on junk.</p>",
      "rawMarkdown": "I also have a CAD background (and helped set up the competition :)\r\n\r\nBear in that the problems we run are real-world problems subject to real-world constraints. The ideal diagnostic tool may indeed be the latest camera with the best lighting conditions using uncompressed tiffs and be trained on a panel of ratings and independently validated. In reality, you always make concessions on study size, design, cost, data availability.\r\n\r\nA major motivation of this study was to see how algorithms could function in messy, non-ideal conditions. Here, this meant a variety of camera conditions. In the future, it might be cell phone cameras, low cost equipment in pharmacies, old systems donated to clinics in needy areas. The use there is more a pragmatic screening tool (\"there's a good chance you need to see a doctor now\") vs a diagnostic tool (\"we are certain you have a very serious disease and the FDA approval to say so\").\r\n\r\nFootnote - we generally don't have a \"not rated\" category because it's not a natural fit for a competition format with ordinal/regression predictions (the \"distance\" between \"not rated\" and 0 doesn't have the same units as the \"distance\" between 0 and 1). In a production system, you would indeed need sanity checking and heuristics to ensure the algorithm isn't predicting on junk.",
      "votes": null
    },
    {
      "id": "87780",
      "postDate": "07/31/2015 18:04:50",
      "content": "<p>Some more thoughts given the comments:</p>\n\n<p>I agree that any study is limited by time and budget but we have to know what is required to get useful results and then make the &quot;so what test&quot; that if the study is done under the constraints whether it will produce them.</p>\n\n<p>Having multiple single experts evaluate different images for the gold standard does not increase reliability. Indeed, it increases the amount of training data required for an automated algorithm to produce reliable results exponentially by the curse of dimensionality. </p>\n\n<p>Similarly having more data does not necessarily improve accuracy. For example, combining results from a 10 megapixel 3000x3000 pixel camera with results from a 400 kilobyte 640x640 camera on different images will not increase the accuracy. As mentioned above it will increase the amount of training data required for automated evaluation exponentially. The reliability of automated algorithms is improved dramatically by reducing the range of input data, such as spoken numerals instead of general speech recognition.  </p>\n\n<p>As to using low quality images for screening, there are several issues. First, personnel costs are always by far the biggest expense. Even in third world countries having people waste their time with data that are not adequate to diagnose the medical condition is not cost effective for the people who make the exams. Not to mention the cost to the patient for a wrong diagnosis.</p>\n\n<p>For screening exams, we would have to compare diagnosis from low quality retinal image data with other low cost methods such as measuring the patient's field of vision and noting visual artifacts such as floaters combined with patient statistics such as blood glucose, blood pressure, and BMI.</p>\n\n<p>Perhaps having results from a test with high tech equipment such as a cell phone camera might get patients to do the simple things to improve their health such as doing more exercise, controlling their diet and so on but that leads us into an entirely different area. </p>",
      "rawMarkdown": "Some more thoughts given the comments:\r\n\r\nI agree that any study is limited by time and budget but we have to know what is required to get useful results and then make the \"so what test\" that if the study is done under the constraints whether it will produce them.\r\n\r\nHaving multiple single experts evaluate different images for the gold standard does not increase reliability. Indeed, it increases the amount of training data required for an automated algorithm to produce reliable results exponentially by the curse of dimensionality. \r\n\r\nSimilarly having more data does not necessarily improve accuracy. For example, combining results from a 10 megapixel 3000x3000 pixel camera with results from a 400 kilobyte 640x640 camera on different images will not increase the accuracy. As mentioned above it will increase the amount of training data required for automated evaluation exponentially. The reliability of automated algorithms is improved dramatically by reducing the range of input data, such as spoken numerals instead of general speech recognition.  \r\n\r\nAs to using low quality images for screening, there are several issues. First, personnel costs are always by far the biggest expense. Even in third world countries having people waste their time with data that are not adequate to diagnose the medical condition is not cost effective for the people who make the exams. Not to mention the cost to the patient for a wrong diagnosis.\r\n\r\nFor screening exams, we would have to compare diagnosis from low quality retinal image data with other low cost methods such as measuring the patient's field of vision and noting visual artifacts such as floaters combined with patient statistics such as blood glucose, blood pressure, and BMI.\r\n\r\nPerhaps having results from a test with high tech equipment such as a cell phone camera might get patients to do the simple things to improve their health such as doing more exercise, controlling their diet and so on but that leads us into an entirely different area.",
      "votes": null
    },
    {
      "id": "87786",
      "postDate": "07/31/2015 18:41:33",
      "content": "<p>An additional thought since I cannot edit my posts.</p>\n\n<p>As to compressed images, yes most people use a subset of the data but, with uncompressed data, the analyst gets to choose the subset that is best for the task, with JPEG compressed data, the subset is chosen by an algorithm that was not optimized to produce images for the data analysis task.</p>",
      "rawMarkdown": "An additional thought since I cannot edit my posts.\r\n\r\nAs to compressed images, yes most people use a subset of the data but, with uncompressed data, the analyst gets to choose the subset that is best for the task, with JPEG compressed data, the subset is chosen by an algorithm that was not optimized to produce images for the data analysis task.",
      "votes": null
    },
    {
      "id": "87793",
      "postDate": "07/31/2015 19:07:08",
      "content": "<p>I think this quote by Theodore Roosevelt is apt.\n<a href=\"http://www.goodreads.com/quotes/7-it-is-not-the-critic-who-counts-not-the-man\">http://www.goodreads.com/quotes/7-it-is-not-the-critic-who-counts-not-the-man</a></p>\n\n<p>Credit belongs to the organizers for putting together a very novel competition.</p>",
      "rawMarkdown": "I think this quote by Theodore Roosevelt is apt.\r\nhttp://www.goodreads.com/quotes/7-it-is-not-the-critic-who-counts-not-the-man\r\n\r\nCredit belongs to the organizers for putting together a very novel competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 87667,
      "author_name": "iamthep",
      "author_url": "",
      "post_date": "07/30/2015 20:31:34",
      "content": "<p>I disagree with most of this post.</p>\n\n<p>The one part I do agree with is having another label of &quot;Not Rated&quot; and/or &quot;Bad Input&quot;.  </p>\n\n<p>Was only a single expert used for the dataset?  I admit that this would be a weak spot as the algorithm would learn to rate images like that single person.  Multiple people labeling disjoint datasets should be fine, which is what I thought they had used.</p>\n\n<p>Having multiple sources of data (ie different imaging machines) is actually quite good.  As long as enough information is available to distinguish each level, then any variation will lead to a more robust outcome.  Plus it is way too expensive to replace all of the machines with 'good' equipment.  Having a label of &quot;Not rated&quot; would also help with determining if the data was good enough to diagnose or not. </p>\n\n<p>Lossy compression is actually crucial to getting better results.  Of course almost everyone reduced the data size massively for their models, so it probably didn't matter.  But if you tried to distribute 100,000 images in less than 200GB of space, you will get a lot more information if you use lossy compression.</p>\n\n<p>Finally, this algorithm is just one component.  Ideally you would like to see a progression of data for each patient and any and all related outcomes with each patient.  But I doubt that data is collected yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87759,
      "author_name": "deepdreaming",
      "author_url": "",
      "post_date": "07/31/2015 14:19:05",
      "content": "<p>I found bobalv's post to be fascinating, as it shows steps that can be taken to increase accuracy.  But, I agree with Ryan with the simple argument that given two models that match a gold standards accuracy, take the simpler one.   And more data, even with noise and redundancy removed, is far more important than analyzing a perfect image sensor reproduction.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87763,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "07/31/2015 15:05:58",
      "content": "<p>I also have a CAD background (and helped set up the competition :)</p>\n\n<p>Bear in that the problems we run are real-world problems subject to real-world constraints. The ideal diagnostic tool may indeed be the latest camera with the best lighting conditions using uncompressed tiffs and be trained on a panel of ratings and independently validated. In reality, you always make concessions on study size, design, cost, data availability.</p>\n\n<p>A major motivation of this study was to see how algorithms could function in messy, non-ideal conditions. Here, this meant a variety of camera conditions. In the future, it might be cell phone cameras, low cost equipment in pharmacies, old systems donated to clinics in needy areas. The use there is more a pragmatic screening tool (&quot;there's a good chance you need to see a doctor now&quot;) vs a diagnostic tool (&quot;we are certain you have a very serious disease and the FDA approval to say so&quot;).</p>\n\n<p>Footnote - we generally don't have a &quot;not rated&quot; category because it's not a natural fit for a competition format with ordinal/regression predictions (the &quot;distance&quot; between &quot;not rated&quot; and 0 doesn't have the same units as the &quot;distance&quot; between 0 and 1). In a production system, you would indeed need sanity checking and heuristics to ensure the algorithm isn't predicting on junk.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87780,
      "author_name": "bobalv",
      "author_url": "",
      "post_date": "07/31/2015 18:04:50",
      "content": "<p>Some more thoughts given the comments:</p>\n\n<p>I agree that any study is limited by time and budget but we have to know what is required to get useful results and then make the &quot;so what test&quot; that if the study is done under the constraints whether it will produce them.</p>\n\n<p>Having multiple single experts evaluate different images for the gold standard does not increase reliability. Indeed, it increases the amount of training data required for an automated algorithm to produce reliable results exponentially by the curse of dimensionality. </p>\n\n<p>Similarly having more data does not necessarily improve accuracy. For example, combining results from a 10 megapixel 3000x3000 pixel camera with results from a 400 kilobyte 640x640 camera on different images will not increase the accuracy. As mentioned above it will increase the amount of training data required for automated evaluation exponentially. The reliability of automated algorithms is improved dramatically by reducing the range of input data, such as spoken numerals instead of general speech recognition.  </p>\n\n<p>As to using low quality images for screening, there are several issues. First, personnel costs are always by far the biggest expense. Even in third world countries having people waste their time with data that are not adequate to diagnose the medical condition is not cost effective for the people who make the exams. Not to mention the cost to the patient for a wrong diagnosis.</p>\n\n<p>For screening exams, we would have to compare diagnosis from low quality retinal image data with other low cost methods such as measuring the patient's field of vision and noting visual artifacts such as floaters combined with patient statistics such as blood glucose, blood pressure, and BMI.</p>\n\n<p>Perhaps having results from a test with high tech equipment such as a cell phone camera might get patients to do the simple things to improve their health such as doing more exercise, controlling their diet and so on but that leads us into an entirely different area. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87786,
      "author_name": "bobalv",
      "author_url": "",
      "post_date": "07/31/2015 18:41:33",
      "content": "<p>An additional thought since I cannot edit my posts.</p>\n\n<p>As to compressed images, yes most people use a subset of the data but, with uncompressed data, the analyst gets to choose the subset that is best for the task, with JPEG compressed data, the subset is chosen by an algorithm that was not optimized to produce images for the data analysis task.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87793,
      "author_name": "acu192",
      "author_url": "",
      "post_date": "07/31/2015 19:07:08",
      "content": "<p>I think this quote by Theodore Roosevelt is apt.\n<a href=\"http://www.goodreads.com/quotes/7-it-is-not-the-critic-who-counts-not-the-man\">http://www.goodreads.com/quotes/7-it-is-not-the-critic-who-counts-not-the-man</a></p>\n\n<p>Credit belongs to the organizers for putting together a very novel competition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "87511": "Because there are a lot of beginners who want to use this competition to learn about data science and medical image analysis, I think a critique of it as a scientific study  is needed. I have a background in medical imaging and have published many papers in the field and also directed a research project on computer-aided diagnosis (CAD). I think the study design has very serious problems and here are my reasons.\r\n\r\nThe gold standard:\r\nAccording to the Data page, the goal is  \"to create an automated analysis system capable of assigning a score based on this scale.\" The scale, that is, the gold standard of the study, was a clinician's rating on a scale from 0 to 4. \r\n\r\nThe gold standard is the most important part of the study. If the gold standard is not reliable, then the results of the study are meaningless. You may claim that the standard is what it is and an algorithm to replicate it has some use. But that is like claiming that replicating astrological forecasts is useful. \r\n\r\nThe strongest gold standard is based on objective, reproducible evidence. For example, if you are reading chest x-ray  images to detect a lung nodule, the strongest standard is a direct biopsy that extracts the nodule. The next strongest is a more accurate examination that may be more expensive and/or invasive than the one you are studying: in the case of the lung nodules a thoracic computed tomography scan or for diabetic retinopathy, perhaps an optical coherence tomography scan. \r\n\r\nThe next weakest gold standard is a panel of experts who diagnose your images and you combine their readings by averaging or majority vote for example. These are obviously subjective and in my experience there can be a huge difference between different experts and experts at different institutions. \r\n\r\nBy far the weakest is the one used in this study; a single expert. This is almost totally subjective and worse there seems to have been no or little effort to spot check or do any quality assurance on the readings. For example, see my post on this forum on unexplainable ratings for blank or garbage images. \r\n\r\nThe rating scale:\r\nThe study forces the rating to be one of the five levels 0 to 4. There is no \"not rated\" level for images that are too poor quality for a reliable diagnosis. In medical imaging you always need this option because it is unfair to the patient to provide a rating on a non-diagnostic quality image. If you rate unacceptable images as 0 you give the patient a false result if there is indeed disease. A rating of disease may lead to unnecessary and possible harmful additional tests. \r\n\r\nThe image database:\r\nA good aspect of the study is the very large number of images. Unfortunately, this was achieved by combining images from many different sources some of which seemed to have poor quality. The medical system is analogous to a manufacturing process where you take images of the patients, process them through the system and produce a diagnosis. Any manufacturing engineer will tell you that you achieve the most reliable results by standardizing the inputs. The study takes the opposite approach and tries to make the medical system compensate for variations in inputs. I think this is a mistake.\r\n\r\nThe contest organizers shrug off image quality by saying:\r\n\tLike any real-world data set, you will encounter noise in both the images and labels. Images may contain artifacts, be out of focus, underexposed, or overexposed.  A major aim of this competition is to develop robust algorithms that can function in the presence of noise and variation.\r\n\r\nI disagree with this approach. I have spent a large amount of effort trying to compensate for poor image quality with software and I am sorry to say most of this effort was wasted. Companies that put their efforts into improving the image acquisition hardware and method instead of compensating for it with software always came out ahead. \r\n\r\nAs an example, consider the Hubbell space telescope. You may recall that due to a monumental screw-up when the first images were made it was discovered that the main mirror focal length was incorrect resulting in blurred images. Image blur is very well understood with excellent mathematical models of its effects and many algorithms have been developed to \"de-blur\" images. Also NASA had unlimited computing power available to process the images. But they recognized that this would not provide the image quality required for their research so they risked the lives of astronauts and spent hundreds of millions of dollars to fly a mission to the satellite to replace the mirror. \r\n\r\nThe image database used JPEG compressed data. The use of a lossy data compression algorithm like JPEG, which was designed to produce visually pleasing snapshots of people, would not be acceptable in other areas of medical imaging. The cost of data storage is so low that I do not think there is a justification for using lossy compression of medical data.\r\n\r\nIf as a data analyst you are faced with poor quality data your first job is to improve the data acquisition. If your employer refuses to do this then let them know that this will severely impact the output. If they still refuse to improve it, consider going somewhere else because you will almost certainly get the blame when the project fails. \r\n\r\nOther data:\r\nAn advantage of computer-aided methods over human observers is the ability to handle additional data. In a study such as this one, something as simple as knowing the make and model of the camera hardware can improve the results dramatically. In addition, having the age, sex, fasting blood glucose level, BMI of the patient and so on could lead to much more reliable results. Part of the job of a data analyst is to get as much relevant data as you can particularly if it is already available on some other database as these patient data almost certainly are. \r\n\r\nConclusion:\r\nI applaud the organizers for producing this contest that introduced many data analysts to medical image analysis. I think it is an important area that can benefit from the application of advanced data analysis techniques. However, beginners should be aware of the study limitations so they can improve their chances of success if they decide to continue with the field.",
    "87667": "I disagree with most of this post.\r\n\r\nThe one part I do agree with is having another label of \"Not Rated\" and/or \"Bad Input\".  \r\n\r\nWas only a single expert used for the dataset?  I admit that this would be a weak spot as the algorithm would learn to rate images like that single person.  Multiple people labeling disjoint datasets should be fine, which is what I thought they had used.\r\n\r\nHaving multiple sources of data (ie different imaging machines) is actually quite good.  As long as enough information is available to distinguish each level, then any variation will lead to a more robust outcome.  Plus it is way too expensive to replace all of the machines with 'good' equipment.  Having a label of \"Not rated\" would also help with determining if the data was good enough to diagnose or not. \r\n\r\nLossy compression is actually crucial to getting better results.  Of course almost everyone reduced the data size massively for their models, so it probably didn't matter.  But if you tried to distribute 100,000 images in less than 200GB of space, you will get a lot more information if you use lossy compression.\r\n\r\nFinally, this algorithm is just one component.  Ideally you would like to see a progression of data for each patient and any and all related outcomes with each patient.  But I doubt that data is collected yet.",
    "87759": "I found bobalv's post to be fascinating, as it shows steps that can be taken to increase accuracy.  But, I agree with Ryan with the simple argument that given two models that match a gold standards accuracy, take the simpler one.   And more data, even with noise and redundancy removed, is far more important than analyzing a perfect image sensor reproduction.",
    "87763": "I also have a CAD background (and helped set up the competition :)\r\n\r\nBear in that the problems we run are real-world problems subject to real-world constraints. The ideal diagnostic tool may indeed be the latest camera with the best lighting conditions using uncompressed tiffs and be trained on a panel of ratings and independently validated. In reality, you always make concessions on study size, design, cost, data availability.\r\n\r\nA major motivation of this study was to see how algorithms could function in messy, non-ideal conditions. Here, this meant a variety of camera conditions. In the future, it might be cell phone cameras, low cost equipment in pharmacies, old systems donated to clinics in needy areas. The use there is more a pragmatic screening tool (\"there's a good chance you need to see a doctor now\") vs a diagnostic tool (\"we are certain you have a very serious disease and the FDA approval to say so\").\r\n\r\nFootnote - we generally don't have a \"not rated\" category because it's not a natural fit for a competition format with ordinal/regression predictions (the \"distance\" between \"not rated\" and 0 doesn't have the same units as the \"distance\" between 0 and 1). In a production system, you would indeed need sanity checking and heuristics to ensure the algorithm isn't predicting on junk.",
    "87780": "Some more thoughts given the comments:\r\n\r\nI agree that any study is limited by time and budget but we have to know what is required to get useful results and then make the \"so what test\" that if the study is done under the constraints whether it will produce them.\r\n\r\nHaving multiple single experts evaluate different images for the gold standard does not increase reliability. Indeed, it increases the amount of training data required for an automated algorithm to produce reliable results exponentially by the curse of dimensionality. \r\n\r\nSimilarly having more data does not necessarily improve accuracy. For example, combining results from a 10 megapixel 3000x3000 pixel camera with results from a 400 kilobyte 640x640 camera on different images will not increase the accuracy. As mentioned above it will increase the amount of training data required for automated evaluation exponentially. The reliability of automated algorithms is improved dramatically by reducing the range of input data, such as spoken numerals instead of general speech recognition.  \r\n\r\nAs to using low quality images for screening, there are several issues. First, personnel costs are always by far the biggest expense. Even in third world countries having people waste their time with data that are not adequate to diagnose the medical condition is not cost effective for the people who make the exams. Not to mention the cost to the patient for a wrong diagnosis.\r\n\r\nFor screening exams, we would have to compare diagnosis from low quality retinal image data with other low cost methods such as measuring the patient's field of vision and noting visual artifacts such as floaters combined with patient statistics such as blood glucose, blood pressure, and BMI.\r\n\r\nPerhaps having results from a test with high tech equipment such as a cell phone camera might get patients to do the simple things to improve their health such as doing more exercise, controlling their diet and so on but that leads us into an entirely different area.",
    "87786": "An additional thought since I cannot edit my posts.\r\n\r\nAs to compressed images, yes most people use a subset of the data but, with uncompressed data, the analyst gets to choose the subset that is best for the task, with JPEG compressed data, the subset is chosen by an algorithm that was not optimized to produce images for the data analysis task.",
    "87793": "I think this quote by Theodore Roosevelt is apt.\r\nhttp://www.goodreads.com/quotes/7-it-is-not-the-critic-who-counts-not-the-man\r\n\r\nCredit belongs to the organizers for putting together a very novel competition."
  },
  "source": "meta"
}