{
  "id": 19839,
  "title": "A Medical Perspective on the Quality of the Left Ventricular Volume and Ejection Fraction Measurements in the 2016 Data Science Bowl Competition",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19839",
  "author_name": "",
  "post_date": "2016-03-30T14:26:18.757Z",
  "votes": 13,
  "comment_count": 1,
  "views": 3230,
  "content": "<h2>A Medical Perspective on the Quality of the Left Ventricular Volume and Ejection Fraction Measurements in the 2016 Data Science Bowl Competition</h2>\n\n<p><em>Andrew Arai, MD <br>\nNational Institutes of Health, <br>\nNational Heart, Lung, and Blood Institute, DHHS <br>\nBethesda, MD, USA</em></p>\n\n<p>Let me start by thanking all of the participating teams in this year&#8217;s Data Science Bowl (DSB) co-sponsored by Booz Allen and Kaggle. I had no idea what to expect from the results of this competition, particularly in light of the relatively short development time allotted. In brief, the results from the top 4 teams (which are all I have had access to so far) have been excellent.</p>\n\n<p>The measurement of left ventricular ejection fraction (LVEF) is probably one of the single most important numerical values determined on an adult patient with heart disease. It has been known for many years that a low ejection fraction predicts in patients that survive a heart attack are much more likely to die in the course of the next year than patients with a more normal LVEF. (1)  There are also diseases that cause a heart to enlarge before the LVEF changes such as an aortic valve that leaks severely.(2) Thus, measurement of both LV volumes and the LVEF provide complimentary information that helps in the diagnosis of many patients with heart disease.</p>\n\n<p>Cardiac MRI is considered a reference standard for measuring LVEF and LV volumes in patients. LVEF can be measured by many other tests including echocardiography, radionuclide ventriculography, cardiac computed tomography, and cardiac catheterization. One of the reasons cardiac MRI has been accepted as a reference standard for determining LVEF and LV volumes relates to the reproducibility of making these important measurements.(3)</p>\n\n<p>Let us review how well the 4 top teams in the DSB did with respect to measuring the LV volumes and LVEF. When comparing two measurements of the same images, our standard statistical tests are simple linear correlation and Bland Altman analysis.(4) The results from the top scoring team show an excellent correlation between their measurements the cardiologist&#8217;s measurement. The Bland Altman plot shows limits of agreement (1.96 * SD) at about 22 ml. </p>\n\n<p><img src=\"http://i.imgur.com/dJWWmtW.png?1\" alt=\"1st place team correlations, volumes\" title> <img src=\"http://i.imgur.com/NwaeXEn.png?1\" alt=\"1st place team Bland-Altman, volumes\" title>  </p>\n\n<p>The correlation between the experimental measurement and the cardiologist measurement of LVEF appears slightly lower than the correlations for LV volumes but these results are still very good. The Bland Altman plot probably better represents the level of agreement between the two measurements of LVEF since the range of values is dominated by many values in the normal or near-normal range. </p>\n\n<p><img src=\"http://i.imgur.com/tsAgtYI.png?1\" alt=\"1st place team EF correlation plot\" title> <img src=\"http://i.imgur.com/3QA28ix.png?1\" alt=\"1st place team Bland-Altman plot for EF\" title></p>\n\n<p>How do these results compare with published literature? As summarized in the following table from Bellenger et al,(3) there are different levels of agreement achieved when you ask one person to measure the same study twice (intraobserver). Agreement is usually a little worse when you ask two different people to measure the same study (interobserver). If you actually do two different MRI scans and have the same person measure the volumes and LVEF on each study, you get even worse levels of agreement (interstudy). In addition, inclusion of patients with heart disease, as opposed to normal subjects, also tends to degrade image quality as well as reproducibility of measurements. </p>\n\n<p><img src=\"http://i.imgur.com/nhnIDxr.png?2\" alt=\"Bellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278\" title></p>\n\n<p>Considering the patients included in the DSB were mostly referred for diagnostic studies, the most appropriate comparisons are probably derived from the inter observer measurements in patients with heart failure. In the Bellenger data, the Bland Altman limits of agreement were about 24 ml for LV end diastolic volume (EDV), about 19 ml for LV end systolic volume (ESV), and about 6.5 for LVEF. Thus the RMS error for LV end diastolic volume 12 ml, LV end systolic volume 10 ml, and LVEF of 4.7 by the winning team looks as good or better than the results published in 2000 when this sort of measurement was being actively studied in the primary medical literature. The following table indicates at least 4 teams achieved very similar levels of results.</p>\n\n<p><img src=\"http://i.imgur.com/7TuoavJ.png?1\" alt=\"RMS Errors\" title></p>\n\n<p>One should not over interpret these results and comparisons. The DSB and the Bellenger paper are two very different data sets that might affect the reproducibility of measurements. How the root mean square errors from the DSB statistics compare with the Belenger Bland Altman limits of agreement should be addressed more formally. We will be performing that next level of analysis soon. Nonetheless, the quality of the results definitely warrants further investigation. We would welcome contacts from participants and have potential next steps that could benefit from the hard work that the groups have put into this challenging analysis of cardiac MR images.</p>\n\n<p><strong>References:</strong></p>\n\n<p>(1) Risk stratification and survival after myocardial infarction. The New England journal of medicine. 1983;309:331-336</p>\n\n<p>(2) Bekeredjian R, Grayburn PA. Valvular heart disease: Aortic regurgitation. Circulation. 2005;112:125-134</p>\n\n<p>(3) Bellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278</p>\n\n<p>(4) Bland JM, Altman DG. Measuring agreement in method comparison studies. Statistical methods in medical research. 1999;8:135-160</p>\n\n<p><strong>Images above and additional analysis of the 2016 DSB results are available:</strong> Mulholland, J. Leading and Winning Team Submissions Analysis. Data Science Bowl website <a href=\"http://www.datasciencebowl.com/leading-and-winning-team-submissions-analysis/\">http://www.datasciencebowl.com/leading-and-winning-team-submissions-analysis/</a>. March 30, 2016. Accessed date March 30, 2016</p>",
  "messages": [
    {
      "id": "113323",
      "postDate": "03/30/2016 14:26:18",
      "content": "<h2>A Medical Perspective on the Quality of the Left Ventricular Volume and Ejection Fraction Measurements in the 2016 Data Science Bowl Competition</h2>\n\n<p><em>Andrew Arai, MD <br>\nNational Institutes of Health, <br>\nNational Heart, Lung, and Blood Institute, DHHS <br>\nBethesda, MD, USA</em></p>\n\n<p>Let me start by thanking all of the participating teams in this year&#8217;s Data Science Bowl (DSB) co-sponsored by Booz Allen and Kaggle. I had no idea what to expect from the results of this competition, particularly in light of the relatively short development time allotted. In brief, the results from the top 4 teams (which are all I have had access to so far) have been excellent.</p>\n\n<p>The measurement of left ventricular ejection fraction (LVEF) is probably one of the single most important numerical values determined on an adult patient with heart disease. It has been known for many years that a low ejection fraction predicts in patients that survive a heart attack are much more likely to die in the course of the next year than patients with a more normal LVEF. (1)  There are also diseases that cause a heart to enlarge before the LVEF changes such as an aortic valve that leaks severely.(2) Thus, measurement of both LV volumes and the LVEF provide complimentary information that helps in the diagnosis of many patients with heart disease.</p>\n\n<p>Cardiac MRI is considered a reference standard for measuring LVEF and LV volumes in patients. LVEF can be measured by many other tests including echocardiography, radionuclide ventriculography, cardiac computed tomography, and cardiac catheterization. One of the reasons cardiac MRI has been accepted as a reference standard for determining LVEF and LV volumes relates to the reproducibility of making these important measurements.(3)</p>\n\n<p>Let us review how well the 4 top teams in the DSB did with respect to measuring the LV volumes and LVEF. When comparing two measurements of the same images, our standard statistical tests are simple linear correlation and Bland Altman analysis.(4) The results from the top scoring team show an excellent correlation between their measurements the cardiologist&#8217;s measurement. The Bland Altman plot shows limits of agreement (1.96 * SD) at about 22 ml. </p>\n\n<p><img src=\"http://i.imgur.com/dJWWmtW.png?1\" alt=\"1st place team correlations, volumes\" title> <img src=\"http://i.imgur.com/NwaeXEn.png?1\" alt=\"1st place team Bland-Altman, volumes\" title>  </p>\n\n<p>The correlation between the experimental measurement and the cardiologist measurement of LVEF appears slightly lower than the correlations for LV volumes but these results are still very good. The Bland Altman plot probably better represents the level of agreement between the two measurements of LVEF since the range of values is dominated by many values in the normal or near-normal range. </p>\n\n<p><img src=\"http://i.imgur.com/tsAgtYI.png?1\" alt=\"1st place team EF correlation plot\" title> <img src=\"http://i.imgur.com/3QA28ix.png?1\" alt=\"1st place team Bland-Altman plot for EF\" title></p>\n\n<p>How do these results compare with published literature? As summarized in the following table from Bellenger et al,(3) there are different levels of agreement achieved when you ask one person to measure the same study twice (intraobserver). Agreement is usually a little worse when you ask two different people to measure the same study (interobserver). If you actually do two different MRI scans and have the same person measure the volumes and LVEF on each study, you get even worse levels of agreement (interstudy). In addition, inclusion of patients with heart disease, as opposed to normal subjects, also tends to degrade image quality as well as reproducibility of measurements. </p>\n\n<p><img src=\"http://i.imgur.com/nhnIDxr.png?2\" alt=\"Bellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278\" title></p>\n\n<p>Considering the patients included in the DSB were mostly referred for diagnostic studies, the most appropriate comparisons are probably derived from the inter observer measurements in patients with heart failure. In the Bellenger data, the Bland Altman limits of agreement were about 24 ml for LV end diastolic volume (EDV), about 19 ml for LV end systolic volume (ESV), and about 6.5 for LVEF. Thus the RMS error for LV end diastolic volume 12 ml, LV end systolic volume 10 ml, and LVEF of 4.7 by the winning team looks as good or better than the results published in 2000 when this sort of measurement was being actively studied in the primary medical literature. The following table indicates at least 4 teams achieved very similar levels of results.</p>\n\n<p><img src=\"http://i.imgur.com/7TuoavJ.png?1\" alt=\"RMS Errors\" title></p>\n\n<p>One should not over interpret these results and comparisons. The DSB and the Bellenger paper are two very different data sets that might affect the reproducibility of measurements. How the root mean square errors from the DSB statistics compare with the Belenger Bland Altman limits of agreement should be addressed more formally. We will be performing that next level of analysis soon. Nonetheless, the quality of the results definitely warrants further investigation. We would welcome contacts from participants and have potential next steps that could benefit from the hard work that the groups have put into this challenging analysis of cardiac MR images.</p>\n\n<p><strong>References:</strong></p>\n\n<p>(1) Risk stratification and survival after myocardial infarction. The New England journal of medicine. 1983;309:331-336</p>\n\n<p>(2) Bekeredjian R, Grayburn PA. Valvular heart disease: Aortic regurgitation. Circulation. 2005;112:125-134</p>\n\n<p>(3) Bellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278</p>\n\n<p>(4) Bland JM, Altman DG. Measuring agreement in method comparison studies. Statistical methods in medical research. 1999;8:135-160</p>\n\n<p><strong>Images above and additional analysis of the 2016 DSB results are available:</strong> Mulholland, J. Leading and Winning Team Submissions Analysis. Data Science Bowl website <a href=\"http://www.datasciencebowl.com/leading-and-winning-team-submissions-analysis/\">http://www.datasciencebowl.com/leading-and-winning-team-submissions-analysis/</a>. March 30, 2016. Accessed date March 30, 2016</p>",
      "rawMarkdown": "A Medical Perspective on the Quality of the Left Ventricular Volume and Ejection Fraction Measurements in the 2016 Data Science Bowl Competition\r\n------------------------------------------------------------------------\r\n\r\n\r\n*Andrew Arai, MD  \r\nNational Institutes of Health,   \r\nNational Heart, Lung, and Blood Institute, DHHS  \r\nBethesda, MD, USA*\r\n\r\nLet me start by thanking all of the participating teams in this year’s Data Science Bowl (DSB) co-sponsored by Booz Allen and Kaggle. I had no idea what to expect from the results of this competition, particularly in light of the relatively short development time allotted. In brief, the results from the top 4 teams (which are all I have had access to so far) have been excellent.\r\n\r\nThe measurement of left ventricular ejection fraction (LVEF) is probably one of the single most important numerical values determined on an adult patient with heart disease. It has been known for many years that a low ejection fraction predicts in patients that survive a heart attack are much more likely to die in the course of the next year than patients with a more normal LVEF. (1)  There are also diseases that cause a heart to enlarge before the LVEF changes such as an aortic valve that leaks severely.(2) Thus, measurement of both LV volumes and the LVEF provide complimentary information that helps in the diagnosis of many patients with heart disease.\r\n\r\nCardiac MRI is considered a reference standard for measuring LVEF and LV volumes in patients. LVEF can be measured by many other tests including echocardiography, radionuclide ventriculography, cardiac computed tomography, and cardiac catheterization. One of the reasons cardiac MRI has been accepted as a reference standard for determining LVEF and LV volumes relates to the reproducibility of making these important measurements.(3)\r\n\r\nLet us review how well the 4 top teams in the DSB did with respect to measuring the LV volumes and LVEF. When comparing two measurements of the same images, our standard statistical tests are simple linear correlation and Bland Altman analysis.(4) The results from the top scoring team show an excellent correlation between their measurements the cardiologist’s measurement. The Bland Altman plot shows limits of agreement (1.96 * SD) at about 22 ml. \r\n\r\n![1st place team correlations, volumes][1] ![1st place team Bland-Altman, volumes][2]  \r\n\r\nThe correlation between the experimental measurement and the cardiologist measurement of LVEF appears slightly lower than the correlations for LV volumes but these results are still very good. The Bland Altman plot probably better represents the level of agreement between the two measurements of LVEF since the range of values is dominated by many values in the normal or near-normal range. \r\n\r\n![1st place team EF correlation plot][3] ![1st place team Bland-Altman plot for EF][4]\r\n\r\nHow do these results compare with published literature? As summarized in the following table from Bellenger et al,(3) there are different levels of agreement achieved when you ask one person to measure the same study twice (intraobserver). Agreement is usually a little worse when you ask two different people to measure the same study (interobserver). If you actually do two different MRI scans and have the same person measure the volumes and LVEF on each study, you get even worse levels of agreement (interstudy). In addition, inclusion of patients with heart disease, as opposed to normal subjects, also tends to degrade image quality as well as reproducibility of measurements. \r\n\r\n![Bellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278][5]\r\n\r\nConsidering the patients included in the DSB were mostly referred for diagnostic studies, the most appropriate comparisons are probably derived from the inter observer measurements in patients with heart failure. In the Bellenger data, the Bland Altman limits of agreement were about 24 ml for LV end diastolic volume (EDV), about 19 ml for LV end systolic volume (ESV), and about 6.5 for LVEF. Thus the RMS error for LV end diastolic volume 12 ml, LV end systolic volume 10 ml, and LVEF of 4.7 by the winning team looks as good or better than the results published in 2000 when this sort of measurement was being actively studied in the primary medical literature. The following table indicates at least 4 teams achieved very similar levels of results.\r\n\r\n![RMS Errors][6]\r\n\r\nOne should not over interpret these results and comparisons. The DSB and the Bellenger paper are two very different data sets that might affect the reproducibility of measurements. How the root mean square errors from the DSB statistics compare with the Belenger Bland Altman limits of agreement should be addressed more formally. We will be performing that next level of analysis soon. Nonetheless, the quality of the results definitely warrants further investigation. We would welcome contacts from participants and have potential next steps that could benefit from the hard work that the groups have put into this challenging analysis of cardiac MR images.\r\n\r\n\r\n**References:**\r\n\r\n(1)\tRisk stratification and survival after myocardial infarction. The New England journal of medicine. 1983;309:331-336\r\n\r\n(2)\tBekeredjian R, Grayburn PA. Valvular heart disease: Aortic regurgitation. Circulation. 2005;112:125-134\r\n\r\n(3)\tBellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278\r\n\r\n(4)\tBland JM, Altman DG. Measuring agreement in method comparison studies. Statistical methods in medical research. 1999;8:135-160\r\n\r\n**Images above and additional analysis of the 2016 DSB results are available:** Mulholland, J. Leading and Winning Team Submissions Analysis. Data Science Bowl website http://www.datasciencebowl.com/leading-and-winning-team-submissions-analysis/. March 30, 2016. Accessed date March 30, 2016\r\n\r\n\r\n  [1]: http://i.imgur.com/dJWWmtW.png?1\r\n  [2]: http://i.imgur.com/NwaeXEn.png?1\r\n  [3]: http://i.imgur.com/tsAgtYI.png?1\r\n  [4]: http://i.imgur.com/3QA28ix.png?1\r\n  [5]: http://i.imgur.com/nhnIDxr.png?2\r\n  [6]: http://i.imgur.com/7TuoavJ.png?1",
      "votes": null
    },
    {
      "id": "220827",
      "postDate": "09/13/2017 08:07:08",
      "content": "<p>Excellent study, thank you</p>",
      "rawMarkdown": "Excellent study, thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 220827,
      "author_name": "blasco23",
      "author_url": "",
      "post_date": "09/13/2017 08:07:08",
      "content": "<p>Excellent study, thank you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "113323": "A Medical Perspective on the Quality of the Left Ventricular Volume and Ejection Fraction Measurements in the 2016 Data Science Bowl Competition\r\n------------------------------------------------------------------------\r\n\r\n\r\n*Andrew Arai, MD  \r\nNational Institutes of Health,   \r\nNational Heart, Lung, and Blood Institute, DHHS  \r\nBethesda, MD, USA*\r\n\r\nLet me start by thanking all of the participating teams in this year’s Data Science Bowl (DSB) co-sponsored by Booz Allen and Kaggle. I had no idea what to expect from the results of this competition, particularly in light of the relatively short development time allotted. In brief, the results from the top 4 teams (which are all I have had access to so far) have been excellent.\r\n\r\nThe measurement of left ventricular ejection fraction (LVEF) is probably one of the single most important numerical values determined on an adult patient with heart disease. It has been known for many years that a low ejection fraction predicts in patients that survive a heart attack are much more likely to die in the course of the next year than patients with a more normal LVEF. (1)  There are also diseases that cause a heart to enlarge before the LVEF changes such as an aortic valve that leaks severely.(2) Thus, measurement of both LV volumes and the LVEF provide complimentary information that helps in the diagnosis of many patients with heart disease.\r\n\r\nCardiac MRI is considered a reference standard for measuring LVEF and LV volumes in patients. LVEF can be measured by many other tests including echocardiography, radionuclide ventriculography, cardiac computed tomography, and cardiac catheterization. One of the reasons cardiac MRI has been accepted as a reference standard for determining LVEF and LV volumes relates to the reproducibility of making these important measurements.(3)\r\n\r\nLet us review how well the 4 top teams in the DSB did with respect to measuring the LV volumes and LVEF. When comparing two measurements of the same images, our standard statistical tests are simple linear correlation and Bland Altman analysis.(4) The results from the top scoring team show an excellent correlation between their measurements the cardiologist’s measurement. The Bland Altman plot shows limits of agreement (1.96 * SD) at about 22 ml. \r\n\r\n![1st place team correlations, volumes][1] ![1st place team Bland-Altman, volumes][2]  \r\n\r\nThe correlation between the experimental measurement and the cardiologist measurement of LVEF appears slightly lower than the correlations for LV volumes but these results are still very good. The Bland Altman plot probably better represents the level of agreement between the two measurements of LVEF since the range of values is dominated by many values in the normal or near-normal range. \r\n\r\n![1st place team EF correlation plot][3] ![1st place team Bland-Altman plot for EF][4]\r\n\r\nHow do these results compare with published literature? As summarized in the following table from Bellenger et al,(3) there are different levels of agreement achieved when you ask one person to measure the same study twice (intraobserver). Agreement is usually a little worse when you ask two different people to measure the same study (interobserver). If you actually do two different MRI scans and have the same person measure the volumes and LVEF on each study, you get even worse levels of agreement (interstudy). In addition, inclusion of patients with heart disease, as opposed to normal subjects, also tends to degrade image quality as well as reproducibility of measurements. \r\n\r\n![Bellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278][5]\r\n\r\nConsidering the patients included in the DSB were mostly referred for diagnostic studies, the most appropriate comparisons are probably derived from the inter observer measurements in patients with heart failure. In the Bellenger data, the Bland Altman limits of agreement were about 24 ml for LV end diastolic volume (EDV), about 19 ml for LV end systolic volume (ESV), and about 6.5 for LVEF. Thus the RMS error for LV end diastolic volume 12 ml, LV end systolic volume 10 ml, and LVEF of 4.7 by the winning team looks as good or better than the results published in 2000 when this sort of measurement was being actively studied in the primary medical literature. The following table indicates at least 4 teams achieved very similar levels of results.\r\n\r\n![RMS Errors][6]\r\n\r\nOne should not over interpret these results and comparisons. The DSB and the Bellenger paper are two very different data sets that might affect the reproducibility of measurements. How the root mean square errors from the DSB statistics compare with the Belenger Bland Altman limits of agreement should be addressed more formally. We will be performing that next level of analysis soon. Nonetheless, the quality of the results definitely warrants further investigation. We would welcome contacts from participants and have potential next steps that could benefit from the hard work that the groups have put into this challenging analysis of cardiac MR images.\r\n\r\n\r\n**References:**\r\n\r\n(1)\tRisk stratification and survival after myocardial infarction. The New England journal of medicine. 1983;309:331-336\r\n\r\n(2)\tBekeredjian R, Grayburn PA. Valvular heart disease: Aortic regurgitation. Circulation. 2005;112:125-134\r\n\r\n(3)\tBellenger NG, Davies LC, Francis JM, Coats AJ, Pennell DJ. Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance. J Cardiovasc Magn Reson. 2000;2:271-278\r\n\r\n(4)\tBland JM, Altman DG. Measuring agreement in method comparison studies. Statistical methods in medical research. 1999;8:135-160\r\n\r\n**Images above and additional analysis of the 2016 DSB results are available:** Mulholland, J. Leading and Winning Team Submissions Analysis. Data Science Bowl website http://www.datasciencebowl.com/leading-and-winning-team-submissions-analysis/. March 30, 2016. Accessed date March 30, 2016\r\n\r\n\r\n  [1]: http://i.imgur.com/dJWWmtW.png?1\r\n  [2]: http://i.imgur.com/NwaeXEn.png?1\r\n  [3]: http://i.imgur.com/tsAgtYI.png?1\r\n  [4]: http://i.imgur.com/3QA28ix.png?1\r\n  [5]: http://i.imgur.com/nhnIDxr.png?2\r\n  [6]: http://i.imgur.com/7TuoavJ.png?1",
    "220827": "Excellent study, thank you"
  },
  "source": "meta"
}