{
  "id": 221952,
  "title": "NIH Chest X-rays TFRerods and Usage Experience",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/221952",
  "author_name": "Nikita Kuzmenkov",
  "post_date": "2021-02-24T17:00:45.356000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<h3>Hello!</h3>\n<p><strong><a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-rays</a></strong>, one of the largest X-ray images dataset available, seems to be in the spotlight now.  As it's been recently reconfirmed by the host of the competition <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">here</a></strong>, the entire train (at least) and test (at most) data is 100% relabeled data from the NIH CXR. Therefore one should be careful to avoid overfitting thus making more harm than good from using this external data either for pretraining models or as an additional data source.</p>\n<p>Another point of concern is time and resources, as training a single <code>EfficientNetB4</code> on TPU with TFRecordDataset of <code>112,120</code> samples (and images downscaled to <code>600x600</code>) for 20 epochs must take over 2.5 hours. Skipping serialization and going with <code>from_tensor_slices</code> must be at least 3-4 more time-consuming. </p>\n<p>So I've decided to do a bit of preprocessing and serialized this dataset to TFRecords, which are now available in <code>600x600</code> image quality in this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/nih-chest-xrays-tfrecords\" target=\"_blank\">dataset</a></strong>. To make your custom TFRecords by filtering out the duplicates, tuning image quality, etc., please consider using this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-add-nih-chest-x-rays-tfrecords\" target=\"_blank\">starter notebook</a></strong>.</p>\n<p>I'd like to ask you to share your personal experience on using NIH CXR (or other external datasets) either for pretraining models or enlarging the training dataset and concerns about overfitting.</p>\n<p>Happy coding!</p>",
  "messages": [
    {
      "id": 1216973,
      "postDate": "2021-02-24T17:00:45.357Z",
      "content": "<h3>Hello!</h3>\n<p><strong><a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-rays</a></strong>, one of the largest X-ray images dataset available, seems to be in the spotlight now.  As it's been recently reconfirmed by the host of the competition <strong><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">here</a></strong>, the entire train (at least) and test (at most) data is 100% relabeled data from the NIH CXR. Therefore one should be careful to avoid overfitting thus making more harm than good from using this external data either for pretraining models or as an additional data source.</p>\n<p>Another point of concern is time and resources, as training a single <code>EfficientNetB4</code> on TPU with TFRecordDataset of <code>112,120</code> samples (and images downscaled to <code>600x600</code>) for 20 epochs must take over 2.5 hours. Skipping serialization and going with <code>from_tensor_slices</code> must be at least 3-4 more time-consuming. </p>\n<p>So I've decided to do a bit of preprocessing and serialized this dataset to TFRecords, which are now available in <code>600x600</code> image quality in this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/nih-chest-xrays-tfrecords\" target=\"_blank\">dataset</a></strong>. To make your custom TFRecords by filtering out the duplicates, tuning image quality, etc., please consider using this <strong><a href=\"https://www.kaggle.com/nickuzmenkov/ranzcr-clip-add-nih-chest-x-rays-tfrecords\" target=\"_blank\">starter notebook</a></strong>.</p>\n<p>I'd like to ask you to share your personal experience on using NIH CXR (or other external datasets) either for pretraining models or enlarging the training dataset and concerns about overfitting.</p>\n<p>Happy coding!</p>",
      "rawMarkdown": "### Hello!\n**[NIH Chest X-rays](https://www.kaggle.com/nih-chest-xrays/data)**, one of the largest X-ray images dataset available, seems to be in the spotlight now.  As it's been recently reconfirmed by the host of the competition **[here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808)**, the entire train (at least) and test (at most) data is 100% relabeled data from the NIH CXR. Therefore one should be careful to avoid overfitting thus making more harm than good from using this external data either for pretraining models or as an additional data source.\n\nAnother point of concern is time and resources, as training a single `EfficientNetB4` on TPU with TFRecordDataset of `112,120` samples (and images downscaled to `600x600`) for 20 epochs must take over 2.5 hours. Skipping serialization and going with `from_tensor_slices` must be at least 3-4 more time-consuming. \n\nSo I've decided to do a bit of preprocessing and serialized this dataset to TFRecords, which are now available in `600x600` image quality in this **[dataset](https://www.kaggle.com/nickuzmenkov/nih-chest-xrays-tfrecords)**. To make your custom TFRecords by filtering out the duplicates, tuning image quality, etc., please consider using this **[starter notebook](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-add-nih-chest-x-rays-tfrecords)**.\n\nI'd like to ask you to share your personal experience on using NIH CXR (or other external datasets) either for pretraining models or enlarging the training dataset and concerns about overfitting.\n\nHappy coding!\n\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1216973": "### Hello!\n**[NIH Chest X-rays](https://www.kaggle.com/nih-chest-xrays/data)**, one of the largest X-ray images dataset available, seems to be in the spotlight now.  As it's been recently reconfirmed by the host of the competition **[here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808)**, the entire train (at least) and test (at most) data is 100% relabeled data from the NIH CXR. Therefore one should be careful to avoid overfitting thus making more harm than good from using this external data either for pretraining models or as an additional data source.\n\nAnother point of concern is time and resources, as training a single `EfficientNetB4` on TPU with TFRecordDataset of `112,120` samples (and images downscaled to `600x600`) for 20 epochs must take over 2.5 hours. Skipping serialization and going with `from_tensor_slices` must be at least 3-4 more time-consuming. \n\nSo I've decided to do a bit of preprocessing and serialized this dataset to TFRecords, which are now available in `600x600` image quality in this **[dataset](https://www.kaggle.com/nickuzmenkov/nih-chest-xrays-tfrecords)**. To make your custom TFRecords by filtering out the duplicates, tuning image quality, etc., please consider using this **[starter notebook](https://www.kaggle.com/nickuzmenkov/ranzcr-clip-add-nih-chest-x-rays-tfrecords)**.\n\nI'd like to ask you to share your personal experience on using NIH CXR (or other external datasets) either for pretraining models or enlarging the training dataset and concerns about overfitting.\n\nHappy coding!\n\n"
  }
}