{
  "id": 224146,
  "title": "Remove black borders around some CXRs",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/224146",
  "author_name": "",
  "post_date": "2021-03-07T05:17:06.463585Z",
  "votes": 54,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Some CXRs from the dataset look like this: </p>\n<p><img src=\"https://i.ibb.co/1bcZJhd/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-1.jpg\" alt=\"\">(StudyInstanceUID 1.2.826.0.1.3680043.8.498.10043914748197350530333666211582267984)</p>\n<p>You can use a few lines of code to remove this black border.</p>\n<pre><code>        img = cv2.imread(img_file_path, cv2.IMREAD_GRAYSCALE)\n        mask = img &gt; 0\n        img = img[np.ix_(mask.any(1), mask.any(0))]\n</code></pre>\n<p>Result: <br>\n<img src=\"https://i.ibb.co/w0V53Pt/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-2.jpg\" alt=\"\"></p>\n<p>Hopefully it is useful for your preprocessing 👍</p>",
  "messages": [
    {
      "id": "1229153",
      "postDate": "03/07/2021 05:17:06",
      "content": "<p>Some CXRs from the dataset look like this: </p>\n<p><img src=\"https://i.ibb.co/1bcZJhd/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-1.jpg\" alt=\"\">(StudyInstanceUID 1.2.826.0.1.3680043.8.498.10043914748197350530333666211582267984)</p>\n<p>You can use a few lines of code to remove this black border.</p>\n<pre><code>        img = cv2.imread(img_file_path, cv2.IMREAD_GRAYSCALE)\n        mask = img &gt; 0\n        img = img[np.ix_(mask.any(1), mask.any(0))]\n</code></pre>\n<p>Result: <br>\n<img src=\"https://i.ibb.co/w0V53Pt/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-2.jpg\" alt=\"\"></p>\n<p>Hopefully it is useful for your preprocessing 👍</p>",
      "rawMarkdown": "Some CXRs from the dataset look like this: \n\n![](https://i.ibb.co/1bcZJhd/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-1.jpg)(StudyInstanceUID 1.2.826.0.1.3680043.8.498.10043914748197350530333666211582267984)\n\n\nYou can use a few lines of code to remove this black border.\n\n```\n        img = cv2.imread(img_file_path, cv2.IMREAD_GRAYSCALE)\n        mask = img > 0\n        img = img[np.ix_(mask.any(1), mask.any(0))]\n```\nResult: \n![](https://i.ibb.co/w0V53Pt/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-2.jpg)\n\nHopefully it is useful for your preprocessing 👍",
      "votes": null
    },
    {
      "id": "1230038",
      "postDate": "03/07/2021 18:39:47",
      "content": "<p><a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a> thanks a lot for sharing this! Looks very helpful for such images.</p>",
      "rawMarkdown": "reubenschmidt thanks a lot for sharing this! Looks very helpful for such images.",
      "votes": null
    },
    {
      "id": "1230408",
      "postDate": "03/08/2021 05:29:29",
      "content": "<p>Is there any method to find such images? or should we check manually.</p>",
      "rawMarkdown": "Is there any method to find such images? or should we check manually.",
      "votes": null
    },
    {
      "id": "1230452",
      "postDate": "03/08/2021 06:57:35",
      "content": "<p>run the snippet provided by <a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a> on all images and compare image size before and after ;)</p>",
      "rawMarkdown": "run the snippet provided by @reubenschmidt on all images and compare image size before and after ;)",
      "votes": null
    },
    {
      "id": "1230735",
      "postDate": "03/08/2021 12:03:18",
      "content": "<p>Use this during inference and you should see an increase in score. </p>",
      "rawMarkdown": "Use this during inference and you should see an increase in score.",
      "votes": null
    },
    {
      "id": "1231179",
      "postDate": "03/08/2021 18:20:09",
      "content": "<p><a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a>  how many such images.. around</p>",
      "rawMarkdown": "reubenschmidt  how many such images.. around",
      "votes": null
    },
    {
      "id": "1231236",
      "postDate": "03/08/2021 19:18:46",
      "content": "<p>As <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> said run the snippet and compare the size before and after.</p>",
      "rawMarkdown": "As @christofhenkel said run the snippet and compare the size before and after.",
      "votes": null
    },
    {
      "id": "1231503",
      "postDate": "03/09/2021 03:08:27",
      "content": "<p>i just add your code in inference where cv2. imread in LoadDataset for torch model mutil-head resnet，but it may do a fall on LB score(0.964-0.963). Maybe I has made a mistake on code. Could you give a specific preprocess with code.  Thanks a lot!</p>",
      "rawMarkdown": "i just add your code in inference where cv2. imread in LoadDataset for torch model mutil-head resnet，but it may do a fall on LB score(0.964-0.963). Maybe I has made a mistake on code. Could you give a specific preprocess with code.  Thanks a lot!",
      "votes": null
    },
    {
      "id": "1231509",
      "postDate": "03/09/2021 03:20:31",
      "content": "<p>That is interesting…I added it to inference and saw an increase <code>0.967 -&gt; 0.968</code> for a ResNet200D.</p>",
      "rawMarkdown": "That is interesting...I added it to inference and saw an increase `0.967 -> 0.968` for a ResNet200D.",
      "votes": null
    },
    {
      "id": "1231514",
      "postDate": "03/09/2021 03:32:51",
      "content": "<p>oh, it seems different for some models.  By the way, I should add this such code beyond read file path or other place， i wonder i miss the point</p>",
      "rawMarkdown": "oh, it seems different for some models.  By the way, I should add this such code beyond read file path or other place， i wonder i miss the point",
      "votes": null
    },
    {
      "id": "1231520",
      "postDate": "03/09/2021 03:39:16",
      "content": "<p>In PyTorch, should be something like this:</p>\n<pre><code>class TestDataset(Dataset):  \n    def __init__(self, df, transform=None):  \n        self.df = df  \n        self.file_names = df['StudyInstanceUID'].values  \n        self.transform = transform  \n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        file_name = self.file_names[idx]\n        file_path = f'{TEST_PATH}/{file_name}.jpg'\n\n        #preprocessing step\n        image = cv2.imread(file_path, cv2.IMREAD_GRAYSCALE)\n        mask = image &gt; 0\n        image = image[np.ix_(mask.any(1), mask.any(0))]\n\n        image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        if self.transform:\n            augmented = self.transform(image=image)\n            image = augmented['image']\n        return image\n</code></pre>\n<p>In this code I bump the channels back from 1 to 3. That's what <code>cv2.COLOR_GRAY2RGB</code> does. </p>",
      "rawMarkdown": "In PyTorch, should be something like this:\n\n\n```\nclass TestDataset(Dataset):  \n    def __init__(self, df, transform=None):  \n        self.df = df  \n        self.file_names = df['StudyInstanceUID'].values  \n        self.transform = transform  \n        \n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        file_name = self.file_names[idx]\n        file_path = f'{TEST_PATH}/{file_name}.jpg'\n        \n        #preprocessing step\n        image = cv2.imread(file_path, cv2.IMREAD_GRAYSCALE)\n        mask = image > 0\n        image = image[np.ix_(mask.any(1), mask.any(0))]\n        \n        image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        if self.transform:\n            augmented = self.transform(image=image)\n            image = augmented['image']\n        return image\n```\n\nIn this code I bump the channels back from 1 to 3. That's what `cv2.COLOR_GRAY2RGB` does.",
      "votes": null
    },
    {
      "id": "1231522",
      "postDate": "03/09/2021 03:42:52",
      "content": "<p>I set the same code like yours，hhh. So when you train the model，you also set this or just make inference？</p>",
      "rawMarkdown": "I set the same code like yours，hhh. So when you train the model，you also set this or just make inference？",
      "votes": null
    },
    {
      "id": "1231526",
      "postDate": "03/09/2021 03:48:27",
      "content": "<p>I have not tested this for training yet. I just used it with one of my saved models to see if it would help during inference and it did. </p>\n<p>Update (3/11/2021): I tested this during training and got worse results. My guess is that these black borders act as a crude sort of augmentation that helps the model generalize. </p>",
      "rawMarkdown": "I have not tested this for training yet. I just used it with one of my saved models to see if it would help during inference and it did. \n\nUpdate (3/11/2021): I tested this during training and got worse results. My guess is that these black borders act as a crude sort of augmentation that helps the model generalize.",
      "votes": null
    },
    {
      "id": "1231535",
      "postDate": "03/09/2021 03:56:23",
      "content": "<p>Yes, it can be a little different in inference for trained models.  Anyway，thank you for your great kindness！</p>",
      "rawMarkdown": "Yes, it can be a little different in inference for trained models.  Anyway，thank you for your great kindness！",
      "votes": null
    },
    {
      "id": "1231554",
      "postDate": "03/09/2021 04:15:36",
      "content": "<p>For resnet200d did you use image net weights by any chance :)</p>",
      "rawMarkdown": "For resnet200d did you use image net weights by any chance :)",
      "votes": null
    },
    {
      "id": "1231555",
      "postDate": "03/09/2021 04:18:03",
      "content": "<pre><code>For resnet200d did you use image net weights by any chance :)\n</code></pre>\n<p>I did not :) </p>",
      "rawMarkdown": "```\nFor resnet200d did you use image net weights by any chance :)\n```\n\nI did not :)",
      "votes": null
    },
    {
      "id": "1231562",
      "postDate": "03/09/2021 04:28:12",
      "content": "<p>I see! me neither :)</p>\n<p>If you are willing to, did you perform post processing like TTA or the likes? Any tips on what TTA are useful. </p>",
      "rawMarkdown": "I see! me neither :)\n\nIf you are willing to, did you perform post processing like TTA or the likes? Any tips on what TTA are useful.",
      "votes": null
    },
    {
      "id": "1231563",
      "postDate": "03/09/2021 04:29:26",
      "content": "<p>Asking because my single resnet200d can only hit 0.966 :)</p>",
      "rawMarkdown": "Asking because my single resnet200d can only hit 0.966 :)",
      "votes": null
    },
    {
      "id": "1231566",
      "postDate": "03/09/2021 04:36:13",
      "content": "<p>For TTA, I am only using simple horizontal flipping. In my experiments, this gives an average CV increase of about <code>0.002</code>. I have also tried using random crop TTA, but this only outperforms the horizontal flipping TTA for a high number of steps (5-10). I got an average CV increase of around <code>0.003</code> when using random crop TTA with 7 steps. In my opinion, it is not worth using more advanced TTA techniques, seeing as they barely outperform horizonal flipping TTA but take much longer to finish inference. </p>",
      "rawMarkdown": "For TTA, I am only using simple horizontal flipping. In my experiments, this gives an average CV increase of about `0.002`. I have also tried using random crop TTA, but this only outperforms the horizontal flipping TTA for a high number of steps (5-10). I got an average CV increase of around `0.003` when using random crop TTA with 7 steps. In my opinion, it is not worth using more advanced TTA techniques, seeing as they barely outperform horizonal flipping TTA but take much longer to finish inference.",
      "votes": null
    },
    {
      "id": "1231575",
      "postDate": "03/09/2021 04:48:39",
      "content": "<p>Thanks a lot. Same here I used sin’s horizontal flip! Just wondering on how to implement the TTA steps in pytorch. Cause by Default I only do once. </p>",
      "rawMarkdown": "Thanks a lot. Same here I used sin’s horizontal flip! Just wondering on how to implement the TTA steps in pytorch. Cause by Default I only do once.",
      "votes": null
    },
    {
      "id": "1231582",
      "postDate": "03/09/2021 04:58:52",
      "content": "<p>You would create the augmentations via albumentations and then use them in the test dataset and then just call the test data loader multiple times during inference. Something like:</p>\n<pre><code>def get_transforms(*, data, size=CFG.size):\n    if data == 'tta':\n        return Compose([\n            RandomResizedCrop(size, size, scale=(0.85, 1.0)),\n            HorizontalFlip(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225],\n            ),\n            ToTensorV2(),\n        ])\n\ndef tta_inference(model, loader, tta_steps=5):\n    all_probs = []\n    for i, step in enumerate(range(tta_steps)):\n        probs = []\n        for step, (images) in tqdm(enumerate(loader), total=len(loader)):\n            images = images.to(device)\n            with torch.no_grad():\n                y_preds = model(images)\n            y_preds = y_preds.sigmoid().to('cpu').numpy()\n            probs.append(y_preds)\n        all_probs.append(np.concatenate(probs))\n    avg_probs = np.mean(all_probs, axis=0)\n\n    return avg_probs\n</code></pre>\n<p>Then just create the test dataset and loader like:</p>\n<pre><code>test_dataset = TestDataset(test, transform=get_transforms(size=CFG.size, data='tta'))\ntest_loader = DataLoader(test_dataset, batch_size=CFG.batch_size, shuffle=False, \n                 num_workers=CFG.num_workers, pin_memory=True)\n</code></pre>\n<p>Then run inference:</p>\n<pre><code>predictions = tta_inference(model, test_loader, tta_steps=5)\n</code></pre>\n<p>Hope this is useful for your experiments.</p>",
      "rawMarkdown": "You would create the augmentations via albumentations and then use them in the test dataset and then just call the test data loader multiple times during inference. Something like:\n\n```\ndef get_transforms(*, data, size=CFG.size):\n    if data == 'tta':\n        return Compose([\n            RandomResizedCrop(size, size, scale=(0.85, 1.0)),\n            HorizontalFlip(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225],\n            ),\n            ToTensorV2(),\n        ])\n\ndef tta_inference(model, loader, tta_steps=5):\n    all_probs = []\n    for i, step in enumerate(range(tta_steps)):\n        probs = []\n        for step, (images) in tqdm(enumerate(loader), total=len(loader)):\n            images = images.to(device)\n            with torch.no_grad():\n                y_preds = model(images)\n            y_preds = y_preds.sigmoid().to('cpu').numpy()\n            probs.append(y_preds)\n        all_probs.append(np.concatenate(probs))\n    avg_probs = np.mean(all_probs, axis=0)\n    \n    return avg_probs\n```\n\nThen just create the test dataset and loader like:\n\n```\ntest_dataset = TestDataset(test, transform=get_transforms(size=CFG.size, data='tta'))\ntest_loader = DataLoader(test_dataset, batch_size=CFG.batch_size, shuffle=False, \n                 num_workers=CFG.num_workers, pin_memory=True)\n```\n\nThen run inference:\n\n```\npredictions = tta_inference(model, test_loader, tta_steps=5)\n```\n\nHope this is useful for your experiments.",
      "votes": null
    },
    {
      "id": "1231587",
      "postDate": "03/09/2021 05:01:12",
      "content": "<p>I tried following code in tensorflow, but it's giving error</p>\n<pre><code>import tensorflow.experimental.numpy as tnp\nimage = tf.image.rgb_to_grayscale(image)\nimage = tnp.array(image)\nmask = image &gt; 0\nmask = tnp.array(mask)\nimage = image[tnp.ix_(mask.any(1),mask.any(0))]\nimage = tf.image.grayscale_to_rgb(image)\n</code></pre>",
      "rawMarkdown": "I tried following code in tensorflow, but it's giving error\n\n```\nimport tensorflow.experimental.numpy as tnp\nimage = tf.image.rgb_to_grayscale(image)\nimage = tnp.array(image)\nmask = image > 0\nmask = tnp.array(mask)\nimage = image[tnp.ix_(mask.any(1),mask.any(0))]\nimage = tf.image.grayscale_to_rgb(image)\n```",
      "votes": null
    },
    {
      "id": "1231591",
      "postDate": "03/09/2021 05:06:06",
      "content": "<p>Really appreciative of your help. Huge props! Let’s connect after the competition through WhatsApp/telegram/LinkedIn after the end of the comp, if you don’t mind! Always great to learn from each other </p>",
      "rawMarkdown": "Really appreciative of your help. Huge props! Let’s connect after the competition through WhatsApp/telegram/LinkedIn after the end of the comp, if you don’t mind! Always great to learn from each other",
      "votes": null
    },
    {
      "id": "1232112",
      "postDate": "03/09/2021 13:43:55",
      "content": "<p>me too. 555</p>",
      "rawMarkdown": "me too. 555",
      "votes": null
    },
    {
      "id": "1234763",
      "postDate": "03/11/2021 15:03:03",
      "content": "<p>thank you for sharing this notebook</p>",
      "rawMarkdown": "thank you for sharing this notebook",
      "votes": null
    },
    {
      "id": "1234935",
      "postDate": "03/11/2021 17:36:02",
      "content": "<p>Hi, I have done the code you share above to realize 5x TTA, but some bug could be happened.  It is alright until the next model is loading which is after a first complete model predictions.  Did you ever meet this question? thanks!</p>",
      "rawMarkdown": "Hi, I have done the code you share above to realize 5x TTA, but some bug could be happened.  It is alright until the next model is loading which is after a first complete model predictions.  Did you ever meet this question? thanks!",
      "votes": null
    },
    {
      "id": "1234992",
      "postDate": "03/11/2021 18:26:54",
      "content": "<p>Could you share the error you are getting? What do you mean by 'the next model' - the above code is only for a single model. Are you able to get a full set of TTA predictions, i.e. did you get 5 predictions when you ran it or did it just do one step? </p>\n<p>I took the above code directly from one of my notebooks where everything ran smoothly, so I cannot say for certain what the problem is, but if you give me more details I will try to help. </p>",
      "rawMarkdown": "Could you share the error you are getting? What do you mean by 'the next model' - the above code is only for a single model. Are you able to get a full set of TTA predictions, i.e. did you get 5 predictions when you ran it or did it just do one step? \n\nI took the above code directly from one of my notebooks where everything ran smoothly, so I cannot say for certain what the problem is, but if you give me more details I will try to help.",
      "votes": null
    },
    {
      "id": "1235309",
      "postDate": "03/12/2021 03:35:55",
      "content": "<p>Yeah, it's kind of you. I just try the code below:</p>\n<pre><code> transforms_test = albumentations.Compose([\n     RandomResizedCrop(image_size, image_size, scale=(0.85, 1.0)),\n     HorizontalFlip(p=0.5),\n     Normalize(\n          mean=[0.485, 0.456, 0.406],\n          std=[0.229, 0.224, 0.225],\n      ),\n     ToTensorV2()\n ])\n\n class RANZCRDataset(Dataset):\n     def __init__(self, df, mode, transform=None):\n\n         self.df = df.reset_index(drop=True)\n         self.mode = mode\n         self.transform = transform\n         self.labels = df[target_cols].values\n\n     def __len__(self):\n         return len(self.df)\n\n    def __getitem__(self, index):\n         row = self.df.loc[index]\n         img = cv2.imread(row.file_path, cv2.IMREAD_GRAYSCALE)\n         mask = img &gt; 0\n         img = img[np.ix_(mask.any(1), mask.any(0))]\n         img = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n\n         if self.transform is not None:\n             res = self.transform(image=img)\n             img = res['image']\n         label = torch.tensor(self.labels[index]).float()\n         if self.mode == 'test':\n             return img\n         else:\n             return img, label\n\n test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n test['file_path'] = test.StudyInstanceUID.apply(lambda x: os.path.join('../input/ranzcr-clip-catheter- line-classification/test', f'{x}.jpg'))\n target_cols = test.iloc[:, 1:12].columns.tolist()\n\n test_dataset = RANZCRDataset(test, 'test', transform=transforms_test)\n test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False,  num_workers=24)\n\n def tta_inference_func(test_loader, tta_steps=5):\n     model.eval()\n     bar = tqdm(test_loader)\n     LOGITS = []\n     for i, step in enumerate(range(tta_steps)):\n         PREDS = []\n         for step, images in enumerate(bar):\n             x = images.to(device)\n             with torch.no_grad():\n                 logits = model(x)\n             logits = logits.sigmoid().detach().to('cpu').numpy()\n             PREDS.append(logits)\n         LOGITS.append(np.concatenate(PREDS))\n     avg_probs = np.mean(LOGITS, axis=0)\n\n     return avg_probs\n\n if submit:\n     test_preds = []\n     for i in range(len(enet_type)):\n         if enet_type[i] == 'resnet200d':\n             print('resnet200d loaded')\n             model = RANZCRResNet200D(enet_type[i], out_dim=len(target_cols))\n             model = model.to(device)\n         model.load_state_dict(torch.load(model_path[i], map_location='cuda:0'))\n         if tta:\n             test_preds += [tta_inference_func(test_loader, tta_steps=5)]\n         else:\n             test_preds += [inference_func(test_loader)]\n\n     sub2 = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n     sub2[target_cols] = np.mean(test_preds, axis=0)\n else:\n     pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv').to_csv('submission.csv', index=False)\n</code></pre>\n<p>And the bug happened when the first model is loaded, and predicted all the images. But the next model (cv-fold1) cannot load though i've wait for a long time.</p>",
      "rawMarkdown": "Yeah, it's kind of you. I just try the code below:\n\n     transforms_test = albumentations.Compose([\n         RandomResizedCrop(image_size, image_size, scale=(0.85, 1.0)),\n         HorizontalFlip(p=0.5),\n         Normalize(\n              mean=[0.485, 0.456, 0.406],\n              std=[0.229, 0.224, 0.225],\n          ),\n         ToTensorV2()\n     ])\n\n     class RANZCRDataset(Dataset):\n         def __init__(self, df, mode, transform=None):\n        \n             self.df = df.reset_index(drop=True)\n             self.mode = mode\n             self.transform = transform\n             self.labels = df[target_cols].values\n        \n         def __len__(self):\n             return len(self.df)\n    \n        def __getitem__(self, index):\n             row = self.df.loc[index]\n             img = cv2.imread(row.file_path, cv2.IMREAD_GRAYSCALE)\n             mask = img > 0\n             img = img[np.ix_(mask.any(1), mask.any(0))]\n             img = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n        \n             if self.transform is not None:\n                 res = self.transform(image=img)\n                 img = res['image']\n             label = torch.tensor(self.labels[index]).float()\n             if self.mode == 'test':\n                 return img\n             else:\n                 return img, label\n\n     test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n     test['file_path'] = test.StudyInstanceUID.apply(lambda x: os.path.join('../input/ranzcr-clip-catheter- line-classification/test', f'{x}.jpg'))\n     target_cols = test.iloc[:, 1:12].columns.tolist()\n\n     test_dataset = RANZCRDataset(test, 'test', transform=transforms_test)\n     test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False,  num_workers=24)\n\n     def tta_inference_func(test_loader, tta_steps=5):\n         model.eval()\n         bar = tqdm(test_loader)\n         LOGITS = []\n         for i, step in enumerate(range(tta_steps)):\n             PREDS = []\n             for step, images in enumerate(bar):\n                 x = images.to(device)\n                 with torch.no_grad():\n                     logits = model(x)\n                 logits = logits.sigmoid().detach().to('cpu').numpy()\n                 PREDS.append(logits)\n             LOGITS.append(np.concatenate(PREDS))\n         avg_probs = np.mean(LOGITS, axis=0)\n        \n         return avg_probs\n\n     if submit:\n         test_preds = []\n         for i in range(len(enet_type)):\n             if enet_type[i] == 'resnet200d':\n                 print('resnet200d loaded')\n                 model = RANZCRResNet200D(enet_type[i], out_dim=len(target_cols))\n                 model = model.to(device)\n             model.load_state_dict(torch.load(model_path[i], map_location='cuda:0'))\n             if tta:\n                 test_preds += [tta_inference_func(test_loader, tta_steps=5)]\n             else:\n                 test_preds += [inference_func(test_loader)]\n\n         sub2 = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n         sub2[target_cols] = np.mean(test_preds, axis=0)\n     else:\n         pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv').to_csv('submission.csv', index=False)\n\n\n\nAnd the bug happened when the first model is loaded, and predicted all the images. But the next model (cv-fold1) cannot load though i've wait for a long time.",
      "votes": null
    },
    {
      "id": "1235327",
      "postDate": "03/12/2021 03:56:36",
      "content": "<p>It shows: <br>\n     resnet200d loaded 100%<br>\n     3582/3582 [04:32&lt;00:00, 21.45it/s]<br>\nAnd i wait for 10 min…. nonthing out</p>",
      "rawMarkdown": "It shows: \n     resnet200d loaded 100%\n     3582/3582 [04:32<00:00, 21.45it/s]\nAnd i wait for 10 min.... nonthing out",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1230038,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "03/07/2021 18:39:47",
      "content": "<p><a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a> thanks a lot for sharing this! Looks very helpful for such images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1230408,
      "author_name": "shanmukh05",
      "author_url": "",
      "post_date": "03/08/2021 05:29:29",
      "content": "<p>Is there any method to find such images? or should we check manually.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1230452,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "03/08/2021 06:57:35",
          "content": "<p>run the snippet provided by <a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a> on all images and compare image size before and after ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1230735,
      "author_name": "tuckerarrants",
      "author_url": "",
      "post_date": "03/08/2021 12:03:18",
      "content": "<p>Use this during inference and you should see an increase in score. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1231179,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "03/08/2021 18:20:09",
      "content": "<p><a href=\"https://www.kaggle.com/reubenschmidt\" target=\"_blank\">@reubenschmidt</a>  how many such images.. around</p>",
      "votes": null,
      "replies": [
        {
          "id": 1231236,
          "author_name": "shanmukh05",
          "author_url": "",
          "post_date": "03/08/2021 19:18:46",
          "content": "<p>As <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> said run the snippet and compare the size before and after.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1231503,
      "author_name": "zhangyunsheng",
      "author_url": "",
      "post_date": "03/09/2021 03:08:27",
      "content": "<p>i just add your code in inference where cv2. imread in LoadDataset for torch model mutil-head resnet，but it may do a fall on LB score(0.964-0.963). Maybe I has made a mistake on code. Could you give a specific preprocess with code.  Thanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1231509,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/09/2021 03:20:31",
          "content": "<p>That is interesting…I added it to inference and saw an increase <code>0.967 -&gt; 0.968</code> for a ResNet200D.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231514,
          "author_name": "zhangyunsheng",
          "author_url": "",
          "post_date": "03/09/2021 03:32:51",
          "content": "<p>oh, it seems different for some models.  By the way, I should add this such code beyond read file path or other place， i wonder i miss the point</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231520,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/09/2021 03:39:16",
          "content": "<p>In PyTorch, should be something like this:</p>\n<pre><code>class TestDataset(Dataset):  \n    def __init__(self, df, transform=None):  \n        self.df = df  \n        self.file_names = df['StudyInstanceUID'].values  \n        self.transform = transform  \n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        file_name = self.file_names[idx]\n        file_path = f'{TEST_PATH}/{file_name}.jpg'\n\n        #preprocessing step\n        image = cv2.imread(file_path, cv2.IMREAD_GRAYSCALE)\n        mask = image &gt; 0\n        image = image[np.ix_(mask.any(1), mask.any(0))]\n\n        image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        if self.transform:\n            augmented = self.transform(image=image)\n            image = augmented['image']\n        return image\n</code></pre>\n<p>In this code I bump the channels back from 1 to 3. That's what <code>cv2.COLOR_GRAY2RGB</code> does. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231522,
          "author_name": "zhangyunsheng",
          "author_url": "",
          "post_date": "03/09/2021 03:42:52",
          "content": "<p>I set the same code like yours，hhh. So when you train the model，you also set this or just make inference？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231526,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/09/2021 03:48:27",
          "content": "<p>I have not tested this for training yet. I just used it with one of my saved models to see if it would help during inference and it did. </p>\n<p>Update (3/11/2021): I tested this during training and got worse results. My guess is that these black borders act as a crude sort of augmentation that helps the model generalize. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231535,
          "author_name": "zhangyunsheng",
          "author_url": "",
          "post_date": "03/09/2021 03:56:23",
          "content": "<p>Yes, it can be a little different in inference for trained models.  Anyway，thank you for your great kindness！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231554,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/09/2021 04:15:36",
          "content": "<p>For resnet200d did you use image net weights by any chance :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231555,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/09/2021 04:18:03",
          "content": "<pre><code>For resnet200d did you use image net weights by any chance :)\n</code></pre>\n<p>I did not :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231562,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/09/2021 04:28:12",
          "content": "<p>I see! me neither :)</p>\n<p>If you are willing to, did you perform post processing like TTA or the likes? Any tips on what TTA are useful. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231563,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/09/2021 04:29:26",
          "content": "<p>Asking because my single resnet200d can only hit 0.966 :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231566,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/09/2021 04:36:13",
          "content": "<p>For TTA, I am only using simple horizontal flipping. In my experiments, this gives an average CV increase of about <code>0.002</code>. I have also tried using random crop TTA, but this only outperforms the horizontal flipping TTA for a high number of steps (5-10). I got an average CV increase of around <code>0.003</code> when using random crop TTA with 7 steps. In my opinion, it is not worth using more advanced TTA techniques, seeing as they barely outperform horizonal flipping TTA but take much longer to finish inference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231575,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/09/2021 04:48:39",
          "content": "<p>Thanks a lot. Same here I used sin’s horizontal flip! Just wondering on how to implement the TTA steps in pytorch. Cause by Default I only do once. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231582,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/09/2021 04:58:52",
          "content": "<p>You would create the augmentations via albumentations and then use them in the test dataset and then just call the test data loader multiple times during inference. Something like:</p>\n<pre><code>def get_transforms(*, data, size=CFG.size):\n    if data == 'tta':\n        return Compose([\n            RandomResizedCrop(size, size, scale=(0.85, 1.0)),\n            HorizontalFlip(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225],\n            ),\n            ToTensorV2(),\n        ])\n\ndef tta_inference(model, loader, tta_steps=5):\n    all_probs = []\n    for i, step in enumerate(range(tta_steps)):\n        probs = []\n        for step, (images) in tqdm(enumerate(loader), total=len(loader)):\n            images = images.to(device)\n            with torch.no_grad():\n                y_preds = model(images)\n            y_preds = y_preds.sigmoid().to('cpu').numpy()\n            probs.append(y_preds)\n        all_probs.append(np.concatenate(probs))\n    avg_probs = np.mean(all_probs, axis=0)\n\n    return avg_probs\n</code></pre>\n<p>Then just create the test dataset and loader like:</p>\n<pre><code>test_dataset = TestDataset(test, transform=get_transforms(size=CFG.size, data='tta'))\ntest_loader = DataLoader(test_dataset, batch_size=CFG.batch_size, shuffle=False, \n                 num_workers=CFG.num_workers, pin_memory=True)\n</code></pre>\n<p>Then run inference:</p>\n<pre><code>predictions = tta_inference(model, test_loader, tta_steps=5)\n</code></pre>\n<p>Hope this is useful for your experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231591,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/09/2021 05:06:06",
          "content": "<p>Really appreciative of your help. Huge props! Let’s connect after the competition through WhatsApp/telegram/LinkedIn after the end of the comp, if you don’t mind! Always great to learn from each other </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1232112,
          "author_name": "redandhey",
          "author_url": "",
          "post_date": "03/09/2021 13:43:55",
          "content": "<p>me too. 555</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1234935,
          "author_name": "zhangyunsheng",
          "author_url": "",
          "post_date": "03/11/2021 17:36:02",
          "content": "<p>Hi, I have done the code you share above to realize 5x TTA, but some bug could be happened.  It is alright until the next model is loading which is after a first complete model predictions.  Did you ever meet this question? thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1234992,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/11/2021 18:26:54",
          "content": "<p>Could you share the error you are getting? What do you mean by 'the next model' - the above code is only for a single model. Are you able to get a full set of TTA predictions, i.e. did you get 5 predictions when you ran it or did it just do one step? </p>\n<p>I took the above code directly from one of my notebooks where everything ran smoothly, so I cannot say for certain what the problem is, but if you give me more details I will try to help. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1235309,
          "author_name": "zhangyunsheng",
          "author_url": "",
          "post_date": "03/12/2021 03:35:55",
          "content": "<p>Yeah, it's kind of you. I just try the code below:</p>\n<pre><code> transforms_test = albumentations.Compose([\n     RandomResizedCrop(image_size, image_size, scale=(0.85, 1.0)),\n     HorizontalFlip(p=0.5),\n     Normalize(\n          mean=[0.485, 0.456, 0.406],\n          std=[0.229, 0.224, 0.225],\n      ),\n     ToTensorV2()\n ])\n\n class RANZCRDataset(Dataset):\n     def __init__(self, df, mode, transform=None):\n\n         self.df = df.reset_index(drop=True)\n         self.mode = mode\n         self.transform = transform\n         self.labels = df[target_cols].values\n\n     def __len__(self):\n         return len(self.df)\n\n    def __getitem__(self, index):\n         row = self.df.loc[index]\n         img = cv2.imread(row.file_path, cv2.IMREAD_GRAYSCALE)\n         mask = img &gt; 0\n         img = img[np.ix_(mask.any(1), mask.any(0))]\n         img = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n\n         if self.transform is not None:\n             res = self.transform(image=img)\n             img = res['image']\n         label = torch.tensor(self.labels[index]).float()\n         if self.mode == 'test':\n             return img\n         else:\n             return img, label\n\n test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n test['file_path'] = test.StudyInstanceUID.apply(lambda x: os.path.join('../input/ranzcr-clip-catheter- line-classification/test', f'{x}.jpg'))\n target_cols = test.iloc[:, 1:12].columns.tolist()\n\n test_dataset = RANZCRDataset(test, 'test', transform=transforms_test)\n test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False,  num_workers=24)\n\n def tta_inference_func(test_loader, tta_steps=5):\n     model.eval()\n     bar = tqdm(test_loader)\n     LOGITS = []\n     for i, step in enumerate(range(tta_steps)):\n         PREDS = []\n         for step, images in enumerate(bar):\n             x = images.to(device)\n             with torch.no_grad():\n                 logits = model(x)\n             logits = logits.sigmoid().detach().to('cpu').numpy()\n             PREDS.append(logits)\n         LOGITS.append(np.concatenate(PREDS))\n     avg_probs = np.mean(LOGITS, axis=0)\n\n     return avg_probs\n\n if submit:\n     test_preds = []\n     for i in range(len(enet_type)):\n         if enet_type[i] == 'resnet200d':\n             print('resnet200d loaded')\n             model = RANZCRResNet200D(enet_type[i], out_dim=len(target_cols))\n             model = model.to(device)\n         model.load_state_dict(torch.load(model_path[i], map_location='cuda:0'))\n         if tta:\n             test_preds += [tta_inference_func(test_loader, tta_steps=5)]\n         else:\n             test_preds += [inference_func(test_loader)]\n\n     sub2 = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n     sub2[target_cols] = np.mean(test_preds, axis=0)\n else:\n     pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv').to_csv('submission.csv', index=False)\n</code></pre>\n<p>And the bug happened when the first model is loaded, and predicted all the images. But the next model (cv-fold1) cannot load though i've wait for a long time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1235327,
          "author_name": "zhangyunsheng",
          "author_url": "",
          "post_date": "03/12/2021 03:56:36",
          "content": "<p>It shows: <br>\n     resnet200d loaded 100%<br>\n     3582/3582 [04:32&lt;00:00, 21.45it/s]<br>\nAnd i wait for 10 min…. nonthing out</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1231587,
      "author_name": "shanmukh05",
      "author_url": "",
      "post_date": "03/09/2021 05:01:12",
      "content": "<p>I tried following code in tensorflow, but it's giving error</p>\n<pre><code>import tensorflow.experimental.numpy as tnp\nimage = tf.image.rgb_to_grayscale(image)\nimage = tnp.array(image)\nmask = image &gt; 0\nmask = tnp.array(mask)\nimage = image[tnp.ix_(mask.any(1),mask.any(0))]\nimage = tf.image.grayscale_to_rgb(image)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1234763,
      "author_name": "salmamohamed98",
      "author_url": "",
      "post_date": "03/11/2021 15:03:03",
      "content": "<p>thank you for sharing this notebook</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1229153": "Some CXRs from the dataset look like this: \n\n![](https://i.ibb.co/1bcZJhd/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-1.jpg)(StudyInstanceUID 1.2.826.0.1.3680043.8.498.10043914748197350530333666211582267984)\n\n\nYou can use a few lines of code to remove this black border.\n\n```\n        img = cv2.imread(img_file_path, cv2.IMREAD_GRAYSCALE)\n        mask = img > 0\n        img = img[np.ix_(mask.any(1), mask.any(0))]\n```\nResult: \n![](https://i.ibb.co/w0V53Pt/1-2-826-0-1-3680043-8-498-10043914748197350530333666211582267984-2.jpg)\n\nHopefully it is useful for your preprocessing 👍",
    "1230038": "reubenschmidt thanks a lot for sharing this! Looks very helpful for such images.",
    "1230408": "Is there any method to find such images? or should we check manually.",
    "1230452": "run the snippet provided by @reubenschmidt on all images and compare image size before and after ;)",
    "1230735": "Use this during inference and you should see an increase in score.",
    "1231179": "reubenschmidt  how many such images.. around",
    "1231236": "As @christofhenkel said run the snippet and compare the size before and after.",
    "1231503": "i just add your code in inference where cv2. imread in LoadDataset for torch model mutil-head resnet，but it may do a fall on LB score(0.964-0.963). Maybe I has made a mistake on code. Could you give a specific preprocess with code.  Thanks a lot!",
    "1231509": "That is interesting...I added it to inference and saw an increase `0.967 -> 0.968` for a ResNet200D.",
    "1231514": "oh, it seems different for some models.  By the way, I should add this such code beyond read file path or other place， i wonder i miss the point",
    "1231520": "In PyTorch, should be something like this:\n\n\n```\nclass TestDataset(Dataset):  \n    def __init__(self, df, transform=None):  \n        self.df = df  \n        self.file_names = df['StudyInstanceUID'].values  \n        self.transform = transform  \n        \n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        file_name = self.file_names[idx]\n        file_path = f'{TEST_PATH}/{file_name}.jpg'\n        \n        #preprocessing step\n        image = cv2.imread(file_path, cv2.IMREAD_GRAYSCALE)\n        mask = image > 0\n        image = image[np.ix_(mask.any(1), mask.any(0))]\n        \n        image = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n        if self.transform:\n            augmented = self.transform(image=image)\n            image = augmented['image']\n        return image\n```\n\nIn this code I bump the channels back from 1 to 3. That's what `cv2.COLOR_GRAY2RGB` does.",
    "1231522": "I set the same code like yours，hhh. So when you train the model，you also set this or just make inference？",
    "1231526": "I have not tested this for training yet. I just used it with one of my saved models to see if it would help during inference and it did. \n\nUpdate (3/11/2021): I tested this during training and got worse results. My guess is that these black borders act as a crude sort of augmentation that helps the model generalize.",
    "1231535": "Yes, it can be a little different in inference for trained models.  Anyway，thank you for your great kindness！",
    "1231554": "For resnet200d did you use image net weights by any chance :)",
    "1231555": "```\nFor resnet200d did you use image net weights by any chance :)\n```\n\nI did not :)",
    "1231562": "I see! me neither :)\n\nIf you are willing to, did you perform post processing like TTA or the likes? Any tips on what TTA are useful.",
    "1231563": "Asking because my single resnet200d can only hit 0.966 :)",
    "1231566": "For TTA, I am only using simple horizontal flipping. In my experiments, this gives an average CV increase of about `0.002`. I have also tried using random crop TTA, but this only outperforms the horizontal flipping TTA for a high number of steps (5-10). I got an average CV increase of around `0.003` when using random crop TTA with 7 steps. In my opinion, it is not worth using more advanced TTA techniques, seeing as they barely outperform horizonal flipping TTA but take much longer to finish inference.",
    "1231575": "Thanks a lot. Same here I used sin’s horizontal flip! Just wondering on how to implement the TTA steps in pytorch. Cause by Default I only do once.",
    "1231582": "You would create the augmentations via albumentations and then use them in the test dataset and then just call the test data loader multiple times during inference. Something like:\n\n```\ndef get_transforms(*, data, size=CFG.size):\n    if data == 'tta':\n        return Compose([\n            RandomResizedCrop(size, size, scale=(0.85, 1.0)),\n            HorizontalFlip(p=0.5),\n            Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225],\n            ),\n            ToTensorV2(),\n        ])\n\ndef tta_inference(model, loader, tta_steps=5):\n    all_probs = []\n    for i, step in enumerate(range(tta_steps)):\n        probs = []\n        for step, (images) in tqdm(enumerate(loader), total=len(loader)):\n            images = images.to(device)\n            with torch.no_grad():\n                y_preds = model(images)\n            y_preds = y_preds.sigmoid().to('cpu').numpy()\n            probs.append(y_preds)\n        all_probs.append(np.concatenate(probs))\n    avg_probs = np.mean(all_probs, axis=0)\n    \n    return avg_probs\n```\n\nThen just create the test dataset and loader like:\n\n```\ntest_dataset = TestDataset(test, transform=get_transforms(size=CFG.size, data='tta'))\ntest_loader = DataLoader(test_dataset, batch_size=CFG.batch_size, shuffle=False, \n                 num_workers=CFG.num_workers, pin_memory=True)\n```\n\nThen run inference:\n\n```\npredictions = tta_inference(model, test_loader, tta_steps=5)\n```\n\nHope this is useful for your experiments.",
    "1231587": "I tried following code in tensorflow, but it's giving error\n\n```\nimport tensorflow.experimental.numpy as tnp\nimage = tf.image.rgb_to_grayscale(image)\nimage = tnp.array(image)\nmask = image > 0\nmask = tnp.array(mask)\nimage = image[tnp.ix_(mask.any(1),mask.any(0))]\nimage = tf.image.grayscale_to_rgb(image)\n```",
    "1231591": "Really appreciative of your help. Huge props! Let’s connect after the competition through WhatsApp/telegram/LinkedIn after the end of the comp, if you don’t mind! Always great to learn from each other",
    "1232112": "me too. 555",
    "1234763": "thank you for sharing this notebook",
    "1234935": "Hi, I have done the code you share above to realize 5x TTA, but some bug could be happened.  It is alright until the next model is loading which is after a first complete model predictions.  Did you ever meet this question? thanks!",
    "1234992": "Could you share the error you are getting? What do you mean by 'the next model' - the above code is only for a single model. Are you able to get a full set of TTA predictions, i.e. did you get 5 predictions when you ran it or did it just do one step? \n\nI took the above code directly from one of my notebooks where everything ran smoothly, so I cannot say for certain what the problem is, but if you give me more details I will try to help.",
    "1235309": "Yeah, it's kind of you. I just try the code below:\n\n     transforms_test = albumentations.Compose([\n         RandomResizedCrop(image_size, image_size, scale=(0.85, 1.0)),\n         HorizontalFlip(p=0.5),\n         Normalize(\n              mean=[0.485, 0.456, 0.406],\n              std=[0.229, 0.224, 0.225],\n          ),\n         ToTensorV2()\n     ])\n\n     class RANZCRDataset(Dataset):\n         def __init__(self, df, mode, transform=None):\n        \n             self.df = df.reset_index(drop=True)\n             self.mode = mode\n             self.transform = transform\n             self.labels = df[target_cols].values\n        \n         def __len__(self):\n             return len(self.df)\n    \n        def __getitem__(self, index):\n             row = self.df.loc[index]\n             img = cv2.imread(row.file_path, cv2.IMREAD_GRAYSCALE)\n             mask = img > 0\n             img = img[np.ix_(mask.any(1), mask.any(0))]\n             img = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n        \n             if self.transform is not None:\n                 res = self.transform(image=img)\n                 img = res['image']\n             label = torch.tensor(self.labels[index]).float()\n             if self.mode == 'test':\n                 return img\n             else:\n                 return img, label\n\n     test = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n     test['file_path'] = test.StudyInstanceUID.apply(lambda x: os.path.join('../input/ranzcr-clip-catheter- line-classification/test', f'{x}.jpg'))\n     target_cols = test.iloc[:, 1:12].columns.tolist()\n\n     test_dataset = RANZCRDataset(test, 'test', transform=transforms_test)\n     test_loader = DataLoader(test_dataset, batch_size=batch_size, shuffle=False,  num_workers=24)\n\n     def tta_inference_func(test_loader, tta_steps=5):\n         model.eval()\n         bar = tqdm(test_loader)\n         LOGITS = []\n         for i, step in enumerate(range(tta_steps)):\n             PREDS = []\n             for step, images in enumerate(bar):\n                 x = images.to(device)\n                 with torch.no_grad():\n                     logits = model(x)\n                 logits = logits.sigmoid().detach().to('cpu').numpy()\n                 PREDS.append(logits)\n             LOGITS.append(np.concatenate(PREDS))\n         avg_probs = np.mean(LOGITS, axis=0)\n        \n         return avg_probs\n\n     if submit:\n         test_preds = []\n         for i in range(len(enet_type)):\n             if enet_type[i] == 'resnet200d':\n                 print('resnet200d loaded')\n                 model = RANZCRResNet200D(enet_type[i], out_dim=len(target_cols))\n                 model = model.to(device)\n             model.load_state_dict(torch.load(model_path[i], map_location='cuda:0'))\n             if tta:\n                 test_preds += [tta_inference_func(test_loader, tta_steps=5)]\n             else:\n                 test_preds += [inference_func(test_loader)]\n\n         sub2 = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv')\n         sub2[target_cols] = np.mean(test_preds, axis=0)\n     else:\n         pd.read_csv('../input/ranzcr-clip-catheter-line-classification/sample_submission.csv').to_csv('submission.csv', index=False)\n\n\n\nAnd the bug happened when the first model is loaded, and predicted all the images. But the next model (cv-fold1) cannot load though i've wait for a long time.",
    "1235327": "It shows: \n     resnet200d loaded 100%\n     3582/3582 [04:32<00:00, 21.45it/s]\nAnd i wait for 10 min.... nonthing out"
  },
  "source": "meta"
}