{
  "id": 165044,
  "title": "Data quality control?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/165044",
  "author_name": "Johannes Hofmanninger",
  "post_date": "2020-07-08T11:08:07.547000",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The dataset has many substantial flaws:</p>\n\n<p>-corrupted files\n-scans have missing slices\n-there are at least two scans not encoded in proper HU\n-some scans have large areas of value 0 surrounding the image\n-for some series the order of the images is corrupted (missing patient position in the dicom header)\n-relevant meta data such as slice-thickness and reconstruction kernel are not present in the dicom header</p>\n\n<p>Will there be a revision of the data provided? And can the organizers ensure a minimum of data quality for the hidden test set? </p>",
  "messages": [
    {
      "id": 920126,
      "postDate": "2020-07-08T11:08:07.547Z",
      "content": "<p>The dataset has many substantial flaws:</p>\n\n<p>-corrupted files\n-scans have missing slices\n-there are at least two scans not encoded in proper HU\n-some scans have large areas of value 0 surrounding the image\n-for some series the order of the images is corrupted (missing patient position in the dicom header)\n-relevant meta data such as slice-thickness and reconstruction kernel are not present in the dicom header</p>\n\n<p>Will there be a revision of the data provided? And can the organizers ensure a minimum of data quality for the hidden test set? </p>",
      "rawMarkdown": "The dataset has many substantial flaws:\n\n-corrupted files\n-scans have missing slices\n-there are at least two scans not encoded in proper HU\n-some scans have large areas of value 0 surrounding the image\n-for some series the order of the images is corrupted (missing patient position in the dicom header)\n-relevant meta data such as slice-thickness and reconstruction kernel are not present in the dicom header\n\nWill there be a revision of the data provided? And can the organizers ensure a minimum of data quality for the hidden test set? \n\n\n\n",
      "votes": 10
    },
    {
      "id": 921538,
      "postDate": "2020-07-09T11:12:43.740Z",
      "content": "<p>Hi Johannes, </p>\n\n<p>This highlights the different angles data scientists and clinical scientists may have on imaging data. The data cannot be flawed because it is a real world cohort of patients with IPF - it exactly represents what it is supposed to. Part of the challenge is to develop tools that can accomodate variability in the image quality (even down to the acquisition parameters). This is not an easy thing to do; a major criticism to existing algorithms for classifying or quantifying disease in this field is that they have be developed on overly sanitized data and as you might expect, they do not generalise well in routine clinical practice. To give you an example of an algorithm developed and trained on what many would consider \"imperfect\" data see <a href=\"https://www.thelancet.com/journals/lanres/article/PIIS2213-2600%2818%2930286-8/fulltext\">here</a>. </p>\n\n<p>Some of the issues you mention are going to be rectified, so hold on!</p>\n\n<p>Many thanks for your interest in this challenge...and I wish you GOOD LUCK!</p>\n\n<p>Simon</p>",
      "rawMarkdown": "Hi Johannes, \n\nThis highlights the different angles data scientists and clinical scientists may have on imaging data. The data cannot be flawed because it is a real world cohort of patients with IPF - it exactly represents what it is supposed to. Part of the challenge is to develop tools that can accomodate variability in the image quality (even down to the acquisition parameters). This is not an easy thing to do; a major criticism to existing algorithms for classifying or quantifying disease in this field is that they have be developed on overly sanitized data and as you might expect, they do not generalise well in routine clinical practice. To give you an example of an algorithm developed and trained on what many would consider \"imperfect\" data see [here](https://www.thelancet.com/journals/lanres/article/PIIS2213-2600(18)30286-8/fulltext). \n\nSome of the issues you mention are going to be rectified, so hold on!\n\nMany thanks for your interest in this challenge...and I wish you GOOD LUCK!\n\nSimon",
      "votes": 1
    },
    {
      "id": 921601,
      "postDate": "2020-07-09T12:15:07.947Z",
      "content": "<p>Hi Simon,</p>\n\n<p>thanks for your reply! You are absolutely right, data diversity is key to generalization and I was not criticizing the diversity in acquisition parameters and image quality. Maybe this was a misunderstanding. In my opinion, there has to be a level of data quality (not particularly image quality) an algorithm can rely on such as the following of standards that were defined for the input modality. For example, if the algorithm is supposed to process a CT scan, should the algorithm rely on the assumption that the intensity values are encoded in HU? Should the algorithm rely on the specified orientation in the header (where is top, bottom, left, right)? Should the algorithm rely on the spacings specified (e.g. to calculate the volume of a structure)?  Should the algorithm rely on the fact that the provided image is a valid image (not corrupted)? </p>\n\n<p>In a real world scenario the algorithm would be used by trained personnel immediately spotting flawed input data (e.g. wrong orientation). It would most certainly only be certified for unmodified images as provided by the scanner (ensuring proper HU encoding, no artificial digitally added structures etc.). </p>\n\n<p>There are modifications to the data that seem completely arbitrary such as HU shifted by -1000 or -2000, corrupted files, the 0 value margins etc.</p>\n\n<p>I see risks specifically for a competition like this: \n(1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner). \n(2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development \nmay go into anticipating data corruptions and artifacts rather than the real thing.</p>\n\n<p>Best,\nJohannes</p>",
      "rawMarkdown": "Hi Simon,\n\nthanks for your reply! You are absolutely right, data diversity is key to generalization and I was not criticizing the diversity in acquisition parameters and image quality. Maybe this was a misunderstanding. In my opinion, there has to be a level of data quality (not particularly image quality) an algorithm can rely on such as the following of standards that were defined for the input modality. For example, if the algorithm is supposed to process a CT scan, should the algorithm rely on the assumption that the intensity values are encoded in HU? Should the algorithm rely on the specified orientation in the header (where is top, bottom, left, right)? Should the algorithm rely on the spacings specified (e.g. to calculate the volume of a structure)?  Should the algorithm rely on the fact that the provided image is a valid image (not corrupted)? \n\nIn a real world scenario the algorithm would be used by trained personnel immediately spotting flawed input data (e.g. wrong orientation). It would most certainly only be certified for unmodified images as provided by the scanner (ensuring proper HU encoding, no artificial digitally added structures etc.). \n\nThere are modifications to the data that seem completely arbitrary such as HU shifted by -1000 or -2000, corrupted files, the 0 value margins etc.\n\nI see risks specifically for a competition like this: \n(1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner). \n(2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development \nmay go into anticipating data corruptions and artifacts rather than the real thing.\n\nBest,\nJohannes\n\n  \n\n\n",
      "votes": 2
    },
    {
      "id": 920237,
      "postDate": "2020-07-08T12:59:11.253Z",
      "content": "<p><a href=\"/hojijoji\">@hojijoji</a> Thanks for raising these. I can confirm the slice thickness should have been present and looks to have been a victim of overeager scrubbing (as you observed, we aggressively anonymize/standardize what is in the headers since DICOM images can have lots of leaky metadata). I'll also look into the other issues with the host as well, but as with any real-world dataset that is sourced from many disparate channels, we advise you try to make your code robust to quality issues and caution against relying on the data.</p>\n\n<p>I'll post an update when we upload a corrected version. It almost certainly will be a non-breaking update, so that minimal changes are needed to incorporate the new data.</p>",
      "rawMarkdown": "@hojijoji Thanks for raising these. I can confirm the slice thickness should have been present and looks to have been a victim of overeager scrubbing (as you observed, we aggressively anonymize/standardize what is in the headers since DICOM images can have lots of leaky metadata). I'll also look into the other issues with the host as well, but as with any real-world dataset that is sourced from many disparate channels, we advise you try to make your code robust to quality issues and caution against relying on the data.\n\nI'll post an update when we upload a corrected version. It almost certainly will be a non-breaking update, so that minimal changes are needed to incorporate the new data.",
      "votes": 2
    },
    {
      "id": 924636,
      "postDate": "2020-07-11T14:52:39.187Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 921538,
      "author_name": "SimonWalsh",
      "author_url": "",
      "post_date": "2020-07-09T11:12:43.740000",
      "content": "<p>Hi Johannes, </p>\n\n<p>This highlights the different angles data scientists and clinical scientists may have on imaging data. The data cannot be flawed because it is a real world cohort of patients with IPF - it exactly represents what it is supposed to. Part of the challenge is to develop tools that can accomodate variability in the image quality (even down to the acquisition parameters). This is not an easy thing to do; a major criticism to existing algorithms for classifying or quantifying disease in this field is that they have be developed on overly sanitized data and as you might expect, they do not generalise well in routine clinical practice. To give you an example of an algorithm developed and trained on what many would consider \"imperfect\" data see <a href=\"https://www.thelancet.com/journals/lanres/article/PIIS2213-2600%2818%2930286-8/fulltext\">here</a>. </p>\n\n<p>Some of the issues you mention are going to be rectified, so hold on!</p>\n\n<p>Many thanks for your interest in this challenge...and I wish you GOOD LUCK!</p>\n\n<p>Simon</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 921601,
      "author_name": "Johannes Hofmanninger",
      "author_url": "",
      "post_date": "2020-07-09T12:15:07.947000",
      "content": "<p>Hi Simon,</p>\n\n<p>thanks for your reply! You are absolutely right, data diversity is key to generalization and I was not criticizing the diversity in acquisition parameters and image quality. Maybe this was a misunderstanding. In my opinion, there has to be a level of data quality (not particularly image quality) an algorithm can rely on such as the following of standards that were defined for the input modality. For example, if the algorithm is supposed to process a CT scan, should the algorithm rely on the assumption that the intensity values are encoded in HU? Should the algorithm rely on the specified orientation in the header (where is top, bottom, left, right)? Should the algorithm rely on the spacings specified (e.g. to calculate the volume of a structure)?  Should the algorithm rely on the fact that the provided image is a valid image (not corrupted)? </p>\n\n<p>In a real world scenario the algorithm would be used by trained personnel immediately spotting flawed input data (e.g. wrong orientation). It would most certainly only be certified for unmodified images as provided by the scanner (ensuring proper HU encoding, no artificial digitally added structures etc.). </p>\n\n<p>There are modifications to the data that seem completely arbitrary such as HU shifted by -1000 or -2000, corrupted files, the 0 value margins etc.</p>\n\n<p>I see risks specifically for a competition like this: \n(1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner). \n(2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development \nmay go into anticipating data corruptions and artifacts rather than the real thing.</p>\n\n<p>Best,\nJohannes</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 920237,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2020-07-08T12:59:11.253000",
      "content": "<p><a href=\"/hojijoji\">@hojijoji</a> Thanks for raising these. I can confirm the slice thickness should have been present and looks to have been a victim of overeager scrubbing (as you observed, we aggressively anonymize/standardize what is in the headers since DICOM images can have lots of leaky metadata). I'll also look into the other issues with the host as well, but as with any real-world dataset that is sourced from many disparate channels, we advise you try to make your code robust to quality issues and caution against relying on the data.</p>\n\n<p>I'll post an update when we upload a corrected version. It almost certainly will be a non-breaking update, so that minimal changes are needed to incorporate the new data.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 924636,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T14:52:39.187000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "920126": "The dataset has many substantial flaws:\n\n-corrupted files\n-scans have missing slices\n-there are at least two scans not encoded in proper HU\n-some scans have large areas of value 0 surrounding the image\n-for some series the order of the images is corrupted (missing patient position in the dicom header)\n-relevant meta data such as slice-thickness and reconstruction kernel are not present in the dicom header\n\nWill there be a revision of the data provided? And can the organizers ensure a minimum of data quality for the hidden test set? \n\n\n\n",
    "921538": "Hi Johannes, \n\nThis highlights the different angles data scientists and clinical scientists may have on imaging data. The data cannot be flawed because it is a real world cohort of patients with IPF - it exactly represents what it is supposed to. Part of the challenge is to develop tools that can accomodate variability in the image quality (even down to the acquisition parameters). This is not an easy thing to do; a major criticism to existing algorithms for classifying or quantifying disease in this field is that they have be developed on overly sanitized data and as you might expect, they do not generalise well in routine clinical practice. To give you an example of an algorithm developed and trained on what many would consider \"imperfect\" data see [here](https://www.thelancet.com/journals/lanres/article/PIIS2213-2600(18)30286-8/fulltext). \n\nSome of the issues you mention are going to be rectified, so hold on!\n\nMany thanks for your interest in this challenge...and I wish you GOOD LUCK!\n\nSimon",
    "921601": "Hi Simon,\n\nthanks for your reply! You are absolutely right, data diversity is key to generalization and I was not criticizing the diversity in acquisition parameters and image quality. Maybe this was a misunderstanding. In my opinion, there has to be a level of data quality (not particularly image quality) an algorithm can rely on such as the following of standards that were defined for the input modality. For example, if the algorithm is supposed to process a CT scan, should the algorithm rely on the assumption that the intensity values are encoded in HU? Should the algorithm rely on the specified orientation in the header (where is top, bottom, left, right)? Should the algorithm rely on the spacings specified (e.g. to calculate the volume of a structure)?  Should the algorithm rely on the fact that the provided image is a valid image (not corrupted)? \n\nIn a real world scenario the algorithm would be used by trained personnel immediately spotting flawed input data (e.g. wrong orientation). It would most certainly only be certified for unmodified images as provided by the scanner (ensuring proper HU encoding, no artificial digitally added structures etc.). \n\nThere are modifications to the data that seem completely arbitrary such as HU shifted by -1000 or -2000, corrupted files, the 0 value margins etc.\n\nI see risks specifically for a competition like this: \n(1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner). \n(2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development \nmay go into anticipating data corruptions and artifacts rather than the real thing.\n\nBest,\nJohannes\n\n  \n\n\n",
    "920237": "@hojijoji Thanks for raising these. I can confirm the slice thickness should have been present and looks to have been a victim of overeager scrubbing (as you observed, we aggressively anonymize/standardize what is in the headers since DICOM images can have lots of leaky metadata). I'll also look into the other issues with the host as well, but as with any real-world dataset that is sourced from many disparate channels, we advise you try to make your code robust to quality issues and caution against relying on the data.\n\nI'll post an update when we upload a corrected version. It almost certainly will be a non-breaking update, so that minimal changes are needed to incorporate the new data.",
    "924636": ""
  }
}