{
  "id": 284024,
  "title": "On the low performance of models in this challenge",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/284024",
  "author_name": "",
  "post_date": "2021-10-29T04:20:58.397411100Z",
  "votes": 30,
  "comment_count": 2,
  "views": 0,
  "content": "<p>On behalf of the organizers of this challenge at RSNA and MICCAI, we want to thank all competitors for participating in the Brain Tumor Radiogenomic Classification challenge.  The challenge broke new ground in several ways and, we believe, represents a significant contribution to research on the specific topic and on AI in medical imaging generally. The low scores achieved, by even the most successful algorithms, underscore some of the issues that must be addressed to successfully adopt AI in practice.</p>\n<p>While this challenge built on data collection efforts carried out to support a related Brain Tumor Segmentation (BraTS) challenge that MICCAI has sponsored for a decade, it included several innovations over previous challenges sponsored by RSNA on Kaggle. This was the first RSNA challenge to use magnetic resonance imaging (MRI) data. The dataset included multiple channels or image types/pulse sequences from each exam. The training dataset was sourced from 18 institutions internationally and is the largest resource of its kind. While some of the data comes from the Cancer Imaging Archive (TCIA), a public repository, and has been used in prior research, the majority has not previously been made publicly available.</p>\n<p>In another first, the predicted feature in the challenge was not a directly visible imaging finding, but rather a genomic marker determined by molecular analysis of biopsy specimens. </p>\n<p>Finally, the private test set was made more difficult through the inclusion of a significant proportion of cases from organizations not represented in the training dataset. The intent was to simulate real-world clinical scenarios and assess how well solutions would generalize to work with this data obtained at different sites.  </p>\n<p>The inherent difficulty in generalizing to new data was illustrated in a prior publication that reported a marked drop in performance when a model trained and validated using public data from the US to predict a different mutation in brain cancer (ATRX) tested poorly on an analogous Chinese dataset. (Eur Radiol (2018). <a href=\"https://doi.org/10.1007/s00330-017-5267-0)\" target=\"_blank\">https://doi.org/10.1007/s00330-017-5267-0)</a>. In our challenge, re-evaluation of the leaderboard using only private test data from sites represented in the training set showed no improvement in performance.  </p>\n<p>Interestingly, in prior published literature researchers have explored the same prediction problem posed in this challenge. Many of you have discovered publications that describe attempts to use deep learning to predict methylation (MGMT) status of brain tumors on MRI. These publications report impressive accuracies using smaller limited datasets, exceeding 0.8 AUC in at least one study. We are working to compare methods and share results with the authors of these studies to help determine whether their models can achieve comparable results with the challenge dataset.  </p>\n<p>The results of this challenge provide more questions than answers regarding the application of imaging AI for radiogenomics. Unfortunately, they do not provide evidence that imaging alone can be used to predict methylation (MGMT promoter status) reliably enough to deliver valuable prognostic information to patients and their treating physicians. Further analysis is needed to determine whether this approach is even feasible and what factors have contributed to the discrepancy between the performance seen in this challenge and that reported in the literature for similar tasks.  </p>\n<p>While, like many of you, we find the results of the challenge unexpected and disappointing, they show that there is much work to do to assess whether we can use medical imaging to reliably forecast genomic features of cancer. They suggest that published research showing successful use of imaging AI to predict methylation and several other brain cancer oncologic markers should be re-evaluated for their ability to generalize to multi-institutional data. We are grateful to the competition participants for helping to demonstrate the work to be done in this area.</p>\n<p>John Mongan<br>\nChair, RSNA Machine Learning Steering Subcommittee</p>",
  "messages": [
    {
      "id": "1564300",
      "postDate": "10/29/2021 04:20:58",
      "content": "<p>On behalf of the organizers of this challenge at RSNA and MICCAI, we want to thank all competitors for participating in the Brain Tumor Radiogenomic Classification challenge.  The challenge broke new ground in several ways and, we believe, represents a significant contribution to research on the specific topic and on AI in medical imaging generally. The low scores achieved, by even the most successful algorithms, underscore some of the issues that must be addressed to successfully adopt AI in practice.</p>\n<p>While this challenge built on data collection efforts carried out to support a related Brain Tumor Segmentation (BraTS) challenge that MICCAI has sponsored for a decade, it included several innovations over previous challenges sponsored by RSNA on Kaggle. This was the first RSNA challenge to use magnetic resonance imaging (MRI) data. The dataset included multiple channels or image types/pulse sequences from each exam. The training dataset was sourced from 18 institutions internationally and is the largest resource of its kind. While some of the data comes from the Cancer Imaging Archive (TCIA), a public repository, and has been used in prior research, the majority has not previously been made publicly available.</p>\n<p>In another first, the predicted feature in the challenge was not a directly visible imaging finding, but rather a genomic marker determined by molecular analysis of biopsy specimens. </p>\n<p>Finally, the private test set was made more difficult through the inclusion of a significant proportion of cases from organizations not represented in the training dataset. The intent was to simulate real-world clinical scenarios and assess how well solutions would generalize to work with this data obtained at different sites.  </p>\n<p>The inherent difficulty in generalizing to new data was illustrated in a prior publication that reported a marked drop in performance when a model trained and validated using public data from the US to predict a different mutation in brain cancer (ATRX) tested poorly on an analogous Chinese dataset. (Eur Radiol (2018). <a href=\"https://doi.org/10.1007/s00330-017-5267-0)\" target=\"_blank\">https://doi.org/10.1007/s00330-017-5267-0)</a>. In our challenge, re-evaluation of the leaderboard using only private test data from sites represented in the training set showed no improvement in performance.  </p>\n<p>Interestingly, in prior published literature researchers have explored the same prediction problem posed in this challenge. Many of you have discovered publications that describe attempts to use deep learning to predict methylation (MGMT) status of brain tumors on MRI. These publications report impressive accuracies using smaller limited datasets, exceeding 0.8 AUC in at least one study. We are working to compare methods and share results with the authors of these studies to help determine whether their models can achieve comparable results with the challenge dataset.  </p>\n<p>The results of this challenge provide more questions than answers regarding the application of imaging AI for radiogenomics. Unfortunately, they do not provide evidence that imaging alone can be used to predict methylation (MGMT promoter status) reliably enough to deliver valuable prognostic information to patients and their treating physicians. Further analysis is needed to determine whether this approach is even feasible and what factors have contributed to the discrepancy between the performance seen in this challenge and that reported in the literature for similar tasks.  </p>\n<p>While, like many of you, we find the results of the challenge unexpected and disappointing, they show that there is much work to do to assess whether we can use medical imaging to reliably forecast genomic features of cancer. They suggest that published research showing successful use of imaging AI to predict methylation and several other brain cancer oncologic markers should be re-evaluated for their ability to generalize to multi-institutional data. We are grateful to the competition participants for helping to demonstrate the work to be done in this area.</p>\n<p>John Mongan<br>\nChair, RSNA Machine Learning Steering Subcommittee</p>",
      "rawMarkdown": "On behalf of the organizers of this challenge at RSNA and MICCAI, we want to thank all competitors for participating in the Brain Tumor Radiogenomic Classification challenge.  The challenge broke new ground in several ways and, we believe, represents a significant contribution to research on the specific topic and on AI in medical imaging generally. The low scores achieved, by even the most successful algorithms, underscore some of the issues that must be addressed to successfully adopt AI in practice.\n\nWhile this challenge built on data collection efforts carried out to support a related Brain Tumor Segmentation (BraTS) challenge that MICCAI has sponsored for a decade, it included several innovations over previous challenges sponsored by RSNA on Kaggle. This was the first RSNA challenge to use magnetic resonance imaging (MRI) data. The dataset included multiple channels or image types/pulse sequences from each exam. The training dataset was sourced from 18 institutions internationally and is the largest resource of its kind. While some of the data comes from the Cancer Imaging Archive (TCIA), a public repository, and has been used in prior research, the majority has not previously been made publicly available.\n\nIn another first, the predicted feature in the challenge was not a directly visible imaging finding, but rather a genomic marker determined by molecular analysis of biopsy specimens. \n\nFinally, the private test set was made more difficult through the inclusion of a significant proportion of cases from organizations not represented in the training dataset. The intent was to simulate real-world clinical scenarios and assess how well solutions would generalize to work with this data obtained at different sites.  \n\nThe inherent difficulty in generalizing to new data was illustrated in a prior publication that reported a marked drop in performance when a model trained and validated using public data from the US to predict a different mutation in brain cancer (ATRX) tested poorly on an analogous Chinese dataset. (Eur Radiol (2018). https://doi.org/10.1007/s00330-017-5267-0). In our challenge, re-evaluation of the leaderboard using only private test data from sites represented in the training set showed no improvement in performance.  \n\nInterestingly, in prior published literature researchers have explored the same prediction problem posed in this challenge. Many of you have discovered publications that describe attempts to use deep learning to predict methylation (MGMT) status of brain tumors on MRI. These publications report impressive accuracies using smaller limited datasets, exceeding 0.8 AUC in at least one study. We are working to compare methods and share results with the authors of these studies to help determine whether their models can achieve comparable results with the challenge dataset.  \n\nThe results of this challenge provide more questions than answers regarding the application of imaging AI for radiogenomics. Unfortunately, they do not provide evidence that imaging alone can be used to predict methylation (MGMT promoter status) reliably enough to deliver valuable prognostic information to patients and their treating physicians. Further analysis is needed to determine whether this approach is even feasible and what factors have contributed to the discrepancy between the performance seen in this challenge and that reported in the literature for similar tasks.  \n\nWhile, like many of you, we find the results of the challenge unexpected and disappointing, they show that there is much work to do to assess whether we can use medical imaging to reliably forecast genomic features of cancer. They suggest that published research showing successful use of imaging AI to predict methylation and several other brain cancer oncologic markers should be re-evaluated for their ability to generalize to multi-institutional data. We are grateful to the competition participants for helping to demonstrate the work to be done in this area.\n\nJohn Mongan\nChair, RSNA Machine Learning Steering Subcommittee",
      "votes": null
    },
    {
      "id": "1565406",
      "postDate": "10/30/2021 12:58:45",
      "content": "<p>Thanks! I agree on many of your points. I applaud the significant effort of so many contributors and the attempt on such an ambitious task. I think you hit the nail on the head by saying that generalizability of AI models is one of the biggest challenge hampering the adoption of AI in medical imaging.</p>\n<p>You mentioned that a significant proportion of cases from organizations are not represented in the training dataset and having seen other challenges employ the same methodology, I agree with you that it is one of the reasons for the shake up and low performance across the board. I feel that generalizability is very hard to evaluate while the competition is ongoing, and as a result, I feel that the component of luck is significantly amplified in determining the winner. </p>\n<p>I do wonder if we can mitigate this by the design of the challenge and for the competitors, the way we train our models. For example, if the sources of the images are disclosed, it might offer the teams a means to evaluate generalizability while training the models. If I have images from 10 sources, I can choose to cross-validate my trained models very differently- train on data from 9 sources and validate on data from 1 source. I'm not sure if this information is provided in this competition, if it is, I feel silly now for not utilizing it! Also disclosing this aspect of the competition in the rules beforehand might nudge more of the teams to innovate new methods in the area of generalizability.  </p>",
      "rawMarkdown": "Thanks! I agree on many of your points. I applaud the significant effort of so many contributors and the attempt on such an ambitious task. I think you hit the nail on the head by saying that generalizability of AI models is one of the biggest challenge hampering the adoption of AI in medical imaging.\n\nYou mentioned that a significant proportion of cases from organizations are not represented in the training dataset and having seen other challenges employ the same methodology, I agree with you that it is one of the reasons for the shake up and low performance across the board. I feel that generalizability is very hard to evaluate while the competition is ongoing, and as a result, I feel that the component of luck is significantly amplified in determining the winner. \n\nI do wonder if we can mitigate this by the design of the challenge and for the competitors, the way we train our models. For example, if the sources of the images are disclosed, it might offer the teams a means to evaluate generalizability while training the models. If I have images from 10 sources, I can choose to cross-validate my trained models very differently- train on data from 9 sources and validate on data from 1 source. I'm not sure if this information is provided in this competition, if it is, I feel silly now for not utilizing it! Also disclosing this aspect of the competition in the rules beforehand might nudge more of the teams to innovate new methods in the area of generalizability.",
      "votes": null
    },
    {
      "id": "1569620",
      "postDate": "11/03/2021 15:36:45",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> please take a look into these two papers:</p>\n<ol>\n<li><a href=\"http://www.ajnr.org/content/early/2021/03/04/ajnr.A7029\" target=\"_blank\">http://www.ajnr.org/content/early/2021/03/04/ajnr.A7029</a></li>\n<li><a href=\"https://www.hindawi.com/journals/bmri/2020/9258649/\" target=\"_blank\">https://www.hindawi.com/journals/bmri/2020/9258649/</a></li>\n</ol>\n<p>I've tried both approaches (or imho close-enough approaches) to both of them and my models didn't learn anything. I've been training them outside the Kaggle though.</p>\n<p>I've reached out to the corresponding authors but they all just ignored my e-mails. </p>",
      "rawMarkdown": "Dear @anthracene please take a look into these two papers:\n1. http://www.ajnr.org/content/early/2021/03/04/ajnr.A7029\n2. https://www.hindawi.com/journals/bmri/2020/9258649/\n\nI've tried both approaches (or imho close-enough approaches) to both of them and my models didn't learn anything. I've been training them outside the Kaggle though.\n\nI've reached out to the corresponding authors but they all just ignored my e-mails.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1565406,
      "author_name": "yeeseng",
      "author_url": "",
      "post_date": "10/30/2021 12:58:45",
      "content": "<p>Thanks! I agree on many of your points. I applaud the significant effort of so many contributors and the attempt on such an ambitious task. I think you hit the nail on the head by saying that generalizability of AI models is one of the biggest challenge hampering the adoption of AI in medical imaging.</p>\n<p>You mentioned that a significant proportion of cases from organizations are not represented in the training dataset and having seen other challenges employ the same methodology, I agree with you that it is one of the reasons for the shake up and low performance across the board. I feel that generalizability is very hard to evaluate while the competition is ongoing, and as a result, I feel that the component of luck is significantly amplified in determining the winner. </p>\n<p>I do wonder if we can mitigate this by the design of the challenge and for the competitors, the way we train our models. For example, if the sources of the images are disclosed, it might offer the teams a means to evaluate generalizability while training the models. If I have images from 10 sources, I can choose to cross-validate my trained models very differently- train on data from 9 sources and validate on data from 1 source. I'm not sure if this information is provided in this competition, if it is, I feel silly now for not utilizing it! Also disclosing this aspect of the competition in the rules beforehand might nudge more of the teams to innovate new methods in the area of generalizability.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1569620,
      "author_name": "charzu",
      "author_url": "",
      "post_date": "11/03/2021 15:36:45",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> please take a look into these two papers:</p>\n<ol>\n<li><a href=\"http://www.ajnr.org/content/early/2021/03/04/ajnr.A7029\" target=\"_blank\">http://www.ajnr.org/content/early/2021/03/04/ajnr.A7029</a></li>\n<li><a href=\"https://www.hindawi.com/journals/bmri/2020/9258649/\" target=\"_blank\">https://www.hindawi.com/journals/bmri/2020/9258649/</a></li>\n</ol>\n<p>I've tried both approaches (or imho close-enough approaches) to both of them and my models didn't learn anything. I've been training them outside the Kaggle though.</p>\n<p>I've reached out to the corresponding authors but they all just ignored my e-mails. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1564300": "On behalf of the organizers of this challenge at RSNA and MICCAI, we want to thank all competitors for participating in the Brain Tumor Radiogenomic Classification challenge.  The challenge broke new ground in several ways and, we believe, represents a significant contribution to research on the specific topic and on AI in medical imaging generally. The low scores achieved, by even the most successful algorithms, underscore some of the issues that must be addressed to successfully adopt AI in practice.\n\nWhile this challenge built on data collection efforts carried out to support a related Brain Tumor Segmentation (BraTS) challenge that MICCAI has sponsored for a decade, it included several innovations over previous challenges sponsored by RSNA on Kaggle. This was the first RSNA challenge to use magnetic resonance imaging (MRI) data. The dataset included multiple channels or image types/pulse sequences from each exam. The training dataset was sourced from 18 institutions internationally and is the largest resource of its kind. While some of the data comes from the Cancer Imaging Archive (TCIA), a public repository, and has been used in prior research, the majority has not previously been made publicly available.\n\nIn another first, the predicted feature in the challenge was not a directly visible imaging finding, but rather a genomic marker determined by molecular analysis of biopsy specimens. \n\nFinally, the private test set was made more difficult through the inclusion of a significant proportion of cases from organizations not represented in the training dataset. The intent was to simulate real-world clinical scenarios and assess how well solutions would generalize to work with this data obtained at different sites.  \n\nThe inherent difficulty in generalizing to new data was illustrated in a prior publication that reported a marked drop in performance when a model trained and validated using public data from the US to predict a different mutation in brain cancer (ATRX) tested poorly on an analogous Chinese dataset. (Eur Radiol (2018). https://doi.org/10.1007/s00330-017-5267-0). In our challenge, re-evaluation of the leaderboard using only private test data from sites represented in the training set showed no improvement in performance.  \n\nInterestingly, in prior published literature researchers have explored the same prediction problem posed in this challenge. Many of you have discovered publications that describe attempts to use deep learning to predict methylation (MGMT) status of brain tumors on MRI. These publications report impressive accuracies using smaller limited datasets, exceeding 0.8 AUC in at least one study. We are working to compare methods and share results with the authors of these studies to help determine whether their models can achieve comparable results with the challenge dataset.  \n\nThe results of this challenge provide more questions than answers regarding the application of imaging AI for radiogenomics. Unfortunately, they do not provide evidence that imaging alone can be used to predict methylation (MGMT promoter status) reliably enough to deliver valuable prognostic information to patients and their treating physicians. Further analysis is needed to determine whether this approach is even feasible and what factors have contributed to the discrepancy between the performance seen in this challenge and that reported in the literature for similar tasks.  \n\nWhile, like many of you, we find the results of the challenge unexpected and disappointing, they show that there is much work to do to assess whether we can use medical imaging to reliably forecast genomic features of cancer. They suggest that published research showing successful use of imaging AI to predict methylation and several other brain cancer oncologic markers should be re-evaluated for their ability to generalize to multi-institutional data. We are grateful to the competition participants for helping to demonstrate the work to be done in this area.\n\nJohn Mongan\nChair, RSNA Machine Learning Steering Subcommittee",
    "1565406": "Thanks! I agree on many of your points. I applaud the significant effort of so many contributors and the attempt on such an ambitious task. I think you hit the nail on the head by saying that generalizability of AI models is one of the biggest challenge hampering the adoption of AI in medical imaging.\n\nYou mentioned that a significant proportion of cases from organizations are not represented in the training dataset and having seen other challenges employ the same methodology, I agree with you that it is one of the reasons for the shake up and low performance across the board. I feel that generalizability is very hard to evaluate while the competition is ongoing, and as a result, I feel that the component of luck is significantly amplified in determining the winner. \n\nI do wonder if we can mitigate this by the design of the challenge and for the competitors, the way we train our models. For example, if the sources of the images are disclosed, it might offer the teams a means to evaluate generalizability while training the models. If I have images from 10 sources, I can choose to cross-validate my trained models very differently- train on data from 9 sources and validate on data from 1 source. I'm not sure if this information is provided in this competition, if it is, I feel silly now for not utilizing it! Also disclosing this aspect of the competition in the rules beforehand might nudge more of the teams to innovate new methods in the area of generalizability.",
    "1569620": "Dear @anthracene please take a look into these two papers:\n1. http://www.ajnr.org/content/early/2021/03/04/ajnr.A7029\n2. https://www.hindawi.com/journals/bmri/2020/9258649/\n\nI've tried both approaches (or imho close-enough approaches) to both of them and my models didn't learn anything. I've been training them outside the Kaggle though.\n\nI've reached out to the corresponding authors but they all just ignored my e-mails."
  },
  "source": "meta"
}