{
  "id": 168641,
  "title": "How to run functions from other libraries in tensorflow data pipeline",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168641",
  "author_name": "Sarthak khandelwal",
  "post_date": "2020-07-21T11:01:51.399000",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Recently I've been searching for new augmentation strategies to implement into my solution and came across <a href=\"https://www.kaggle.com/synked/really-realistic-hair-augmentations\" target=\"_blank\">this</a> work by <strong>Ayaan</strong>. The approach aims to introduce hairs like structure in the images and it really looks like as they are natural. It was implemented to deal with the image as a numpy array but I have a <code>tf.data</code> pipeline to perform data transformations and the images are \"Tensor\" objects and also the execution is in Graph mode so I cannot use <code>image.numpy()</code> here.</p>\n<p>The <code>tf.data</code> pipeline is same as described in <strong>Chris Deotte's</strong>  <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a>.</p>\n<p>The function which I want to implement -</p>\n<pre><code>def hair_augment(img,target):\n    img = tf.image.decode_jpeg(img,channels=3)\n    img = cv2.bitwise_and(img,img,mask=mask)\n    return img,target\n</code></pre>\n<p><br>\nI am calling the above function as  - <br>\n<code>dataset = datset.map( lambda img,target: (hair_augment(img),target))</code></p>\n<p>Maybe the reason for which this function raises exception is the use of \"OpenCV's\" function which are not built to deal with \"Tensors\" and hence I tried to use <strong>tf.bitwise.bitwise<strong>_and()</strong> in place of <strong>cv2.bitwise</strong>and()</strong> . That function does not overlap image with the mask but perform <strong>Bitwise-And</strong> operation and hence the results were dissatisfactory. </p>\n<p>I have also tried to use <strong>tf.py<em>function</em></strong> and <strong>tf.numpyfunction</strong>  but all of them raises exceptions and ultimately I am at the point where I started.</p>\n<p>One approach I have also tried (which is kind of naive) is that I extracted out \"images\" from the dataset, iterating over them using <strong>images.take(-1)</strong>  and passing them to <code>hair_augment</code> function which happened to be successful but now I had to combine \"images\" and \"targets\" again and so I used <code>tf.data.Dataset.zip()</code>  for that which also didn't proved to be a successful approach.</p>\n<p>It seems like similar issues have been asked over various platforms and no clear answer for has been obtained.</p>\n<p>The dataset and code for data pipeline is taken from Chris Deotte's <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a>.<br>\nRelevant Issues - </p>\n<ul>\n<li><a href=\"https://github.com/tensorflow/tensorflow/issues/38762\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/38762</a></li>\n<li><a href=\"https://github.com/tensorflow/tensorflow/issues/27519\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/27519</a></li>\n<li><a href=\"https://stackoverflow.com/questions/56665868/tensor-numpy-not-working-in-tensorflow-data-dataset-throws-the-error-attribu\" target=\"_blank\">https://stackoverflow.com/questions/56665868/tensor-numpy-not-working-in-tensorflow-data-dataset-throws-the-error-attribu</a></li>\n<li><a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316\" target=\"_blank\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316</a></li>\n</ul>",
  "messages": [
    {
      "id": 944335,
      "postDate": "2020-07-25T03:04:37.173Z",
      "content": "<p>I don't think TensorFlow <code>tf.data.Dataset</code> allows us to use any external libraries. I posted links about how to perform all the types of augmentation <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169721#944331\">here</a></p>",
      "rawMarkdown": "I don't think TensorFlow `tf.data.Dataset` allows us to use any external libraries. I posted links about how to perform all the types of augmentation [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169721#944331",
      "votes": 1,
      "replies": [
        {
          "id": 944384,
          "postDate": "2020-07-25T04:27:19.540Z",
          "content": "<p>Thanks for the info <a href=\"/cdeotte\">@cdeotte</a> :)</p>",
          "rawMarkdown": "Thanks for the info @cdeotte :)"
        }
      ]
    },
    {
      "id": 938612,
      "postDate": "2020-07-21T16:21:39.220Z",
      "content": "<p>I have had more success using opencv with tensorflow when I make a class rather than a def.</p>\n\n<p>The current leader of this competition shared a <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168371\">great augmentation code</a> </p>\n\n<p>I am going to paste my hair class - but I have almost no success in getting the code to paste looking correct (this time the first couple of lines are bugerred and the def_call should be tabbed over - all the code is in the class).  In my version of Chris notebook I would just call the class in the def where I am doing my other augmentations.</p>\n\n<p><strong>AdvancedHairAugmentation(image)</strong></p>\n\n<p>`class AdvancedHairAugmentation:\n    def <strong>init</strong>(self, hairs: int = 4, hairs_folder: str = \"\", p: float = 0.5):\n        self.hairs = hairs\n        self.hairs_folder = hairs_folder\n        self.p = p</p>\n\n<pre><code>def __call__(self, img):\n    if random.random() &lt; self.p:\n        n_hairs = random.randint(0, self.hairs)\n\n        if not n_hairs:\n            return img\n\n        # height, width, _ = img.shape  # target image width and height\n        hair_images = [im for im in os.listdir(self.hairs_folder) if 'png' in im]\n\n        for _ in range(n_hairs):\n            hair = cv2.imread(os.path.join(self.hairs_folder, random.choice(hair_images)))\n            hair = cv2.flip(hair, random.choice([-1, 0, 1]))\n            hair = cv2.rotate(hair, random.choice([0, 1, 2]))\n\n            h_height, h_width, _ = hair.shape  # hair image width and height\n            roi_ho = random.randint(0, img.shape[0] - hair.shape[0])\n            roi_wo = random.randint(0, img.shape[1] - hair.shape[1])\n            roi = img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width]\n\n            img2gray = cv2.cvtColor(hair, cv2.COLOR_BGR2GRAY)\n            ret, mask = cv2.threshold(img2gray, 10, 255, cv2.THRESH_BINARY)\n            mask_inv = cv2.bitwise_not(mask)\n            img_bg = cv2.bitwise_and(roi, roi, mask=mask_inv)\n            hair_fg = cv2.bitwise_and(hair, hair, mask=mask)\n\n            dst = cv2.add(img_bg, hair_fg)\n            img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width] = dst\n\n        return img`\n</code></pre>",
      "rawMarkdown": "I have had more success using opencv with tensorflow when I make a class rather than a def.\n\nThe current leader of this competition shared a [great augmentation code](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168371) \n\nI am going to paste my hair class - but I have almost no success in getting the code to paste looking correct (this time the first couple of lines are bugerred and the def_call should be tabbed over - all the code is in the class).  In my version of Chris notebook I would just call the class in the def where I am doing my other augmentations.\n\n\n\n**AdvancedHairAugmentation(image)**\n\n`class AdvancedHairAugmentation:\n    def __init__(self, hairs: int = 4, hairs_folder: str = \"\", p: float = 0.5):\n        self.hairs = hairs\n        self.hairs_folder = hairs_folder\n        self.p = p\n\n    def __call__(self, img):\n        if random.random() &lt; self.p:\n            n_hairs = random.randint(0, self.hairs)\n\n            if not n_hairs:\n                return img\n\n            # height, width, _ = img.shape  # target image width and height\n            hair_images = [im for im in os.listdir(self.hairs_folder) if 'png' in im]\n\n            for _ in range(n_hairs):\n                hair = cv2.imread(os.path.join(self.hairs_folder, random.choice(hair_images)))\n                hair = cv2.flip(hair, random.choice([-1, 0, 1]))\n                hair = cv2.rotate(hair, random.choice([0, 1, 2]))\n\n                h_height, h_width, _ = hair.shape  # hair image width and height\n                roi_ho = random.randint(0, img.shape[0] - hair.shape[0])\n                roi_wo = random.randint(0, img.shape[1] - hair.shape[1])\n                roi = img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width]\n\n                img2gray = cv2.cvtColor(hair, cv2.COLOR_BGR2GRAY)\n                ret, mask = cv2.threshold(img2gray, 10, 255, cv2.THRESH_BINARY)\n                mask_inv = cv2.bitwise_not(mask)\n                img_bg = cv2.bitwise_and(roi, roi, mask=mask_inv)\n                hair_fg = cv2.bitwise_and(hair, hair, mask=mask)\n\n                dst = cv2.add(img_bg, hair_fg)\n                img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width] = dst\n\n            return img`\n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 943704,
          "postDate": "2020-07-24T14:22:24.503Z",
          "content": "<p>Hey <a href=\"/pcjimmmy\">@pcjimmmy</a> a great implementation 😄. Did you called this class in <code>tf.data</code> pipeline (as an argument to <code>map</code> function)? And if yes then how did you convert \"img\" which is of type \"Tensor\" and which does not support opencv operations. Thanks and sorry for late reply. </p>",
          "rawMarkdown": "Hey @pcjimmmy a great implementation 😄. Did you called this class in `tf.data` pipeline (as an argument to `map` function)? And if yes then how did you convert \"img\" which is of type \"Tensor\" and which does not support opencv operations. Thanks and sorry for late reply. "
        },
        {
          "id": 944322,
          "postDate": "2020-07-25T02:44:39.950Z",
          "content": "<p>I have a def augment cell similiar to most that you see.  As a class I call the hair augment as shown below as part of my def.  </p>\n\n<p>I also have class functions for hair removal and microscope augmentations that are discussed in a number of posts.</p>\n\n<p>I call the augments in the map\ndataset = dataset.map(data_augment, num_parallel_calls=AUTO)</p>\n\n<p>`\ndef data_augment(image, label):</p>\n\n<pre><code>HairRemoval(image)\nAdvancedHairAugmentation(image)\nMicroscope(image)\n\nimage = tf.image.random_flip_left_right(image)\nimage = tf.image.random_flip_up_down(image)\nimage = tf.image.random_brightness(image, 0.3)\nimage = tf.image.random_contrast(image, 0.6, 2.0)\nimage = tf.image.random_hue(image, 0.5)\nimage = tf.image.random_saturation(image, 0.25, 2.5)\n\nreturn image, label   `\n</code></pre>\n\n<p>PS - Kaggle loves to kill underscores in pastes - so apparent typos are kaggle doing its thing.</p>",
          "rawMarkdown": "I have a def augment cell similiar to most that you see.  As a class I call the hair augment as shown below as part of my def.  \n\nI also have class functions for hair removal and microscope augmentations that are discussed in a number of posts.\n\nI call the augments in the map\ndataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n\n`\ndef data_augment(image, label):\n\n\n    HairRemoval(image)\n    AdvancedHairAugmentation(image)\n    Microscope(image)\n    \n    image = tf.image.random_flip_left_right(image)\n    image = tf.image.random_flip_up_down(image)\n    image = tf.image.random_brightness(image, 0.3)\n    image = tf.image.random_contrast(image, 0.6, 2.0)\n    image = tf.image.random_hue(image, 0.5)\n    image = tf.image.random_saturation(image, 0.25, 2.5)\n \n    return image, label   `\n\n\n\nPS - Kaggle loves to kill underscores in pastes - so apparent typos are kaggle doing its thing."
        },
        {
          "id": 944385,
          "postDate": "2020-07-25T04:29:16.017Z",
          "content": "<p>Thanks! But still, how did you managed to convert \"Tensor\" to ndarray?</p>",
          "rawMarkdown": "Thanks! But still, how did you managed to convert \"Tensor\" to ndarray?\n"
        },
        {
          "id": 944409,
          "postDate": "2020-07-25T05:03:42.483Z",
          "content": "<p>I have forked Chris triple on my local PC and ran it as shown.   </p>\n\n<p>Look at his transform def he has a note:\n     input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]</p>\n\n<p>My class should have the same note.</p>\n\n<p>His prepare_image  with my class useage:  (as usual my pastes go bad in kaggle)</p>\n\n<p>`def prepare_image(img, augment=True, dim=256): <br>\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.cast(img, tf.float32) / 255.0</p>\n\n<pre><code>img = tf.image.resize(img, [b, b],\n                            method=tf.image.ResizeMethod.NEAREST_NEIGHBOR)\nimg = transform(img,DIM=dim)\nimg = tf.image.random_flip_up_down(img)\nimg = tf.image.random_hue(img, 0.5)\nimg = tf.image.random_saturation(img, 0.25, 2.5)\nimg = tf.image.random_contrast(img, 0.6, 2.0)\nimg = tf.image.random_brightness(img, 0.3)\nHairRemoval(img)\nAdvancedHairAugmentation(img)\n\nimg = tf.reshape(img, [dim,dim, 3])\n\n\nreturn img`\n</code></pre>",
          "rawMarkdown": "I have forked Chris triple on my local PC and ran it as shown.   \n\nLook at his transform def he has a note:\n     input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n\nMy class should have the same note.\n\nHis prepare_image  with my class useage:  (as usual my pastes go bad in kaggle)\n    \n`def prepare_image(img, augment=True, dim=256):    \n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.cast(img, tf.float32) / 255.0\n\n    img = tf.image.resize(img, [b, b],\n                                method=tf.image.ResizeMethod.NEAREST_NEIGHBOR)\n    img = transform(img,DIM=dim)\n    img = tf.image.random_flip_up_down(img)\n    img = tf.image.random_hue(img, 0.5)\n    img = tf.image.random_saturation(img, 0.25, 2.5)\n    img = tf.image.random_contrast(img, 0.6, 2.0)\n    img = tf.image.random_brightness(img, 0.3)\n    HairRemoval(img)\n    AdvancedHairAugmentation(img)\n                      \n    img = tf.reshape(img, [dim,dim, 3])\n    \n            \n    return img`"
        },
        {
          "id": 944542,
          "postDate": "2020-07-25T07:16:59.760Z",
          "content": "<p>Thanks! I would surely try this and would let you know if stuck in implementing it.</p>",
          "rawMarkdown": "Thanks! I would surely try this and would let you know if stuck in implementing it."
        },
        {
          "id": 945061,
          "postDate": "2020-07-25T14:44:25.610Z",
          "content": "<p>FYI -I not currently using any of the 3 augmentations in my models.  While it works, it's expensive to add.  When I am a few days away from the end of this and have some models I like I will add them back in.  So my suggestion would be - getting it working - decide if the improvement in your model exists and than comment them out until the last week.</p>",
          "rawMarkdown": "FYI -I not currently using any of the 3 augmentations in my models.  While it works, it's expensive to add.  When I am a few days away from the end of this and have some models I like I will add them back in.  So my suggestion would be - getting it working - decide if the improvement in your model exists and than comment them out until the last week."
        }
      ]
    },
    {
      "id": 938116,
      "postDate": "2020-07-21T11:01:51.400Z",
      "content": "<p>Recently I've been searching for new augmentation strategies to implement into my solution and came across <a href=\"https://www.kaggle.com/synked/really-realistic-hair-augmentations\" target=\"_blank\">this</a> work by <strong>Ayaan</strong>. The approach aims to introduce hairs like structure in the images and it really looks like as they are natural. It was implemented to deal with the image as a numpy array but I have a <code>tf.data</code> pipeline to perform data transformations and the images are \"Tensor\" objects and also the execution is in Graph mode so I cannot use <code>image.numpy()</code> here.</p>\n<p>The <code>tf.data</code> pipeline is same as described in <strong>Chris Deotte's</strong>  <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a>.</p>\n<p>The function which I want to implement -</p>\n<pre><code>def hair_augment(img,target):\n    img = tf.image.decode_jpeg(img,channels=3)\n    img = cv2.bitwise_and(img,img,mask=mask)\n    return img,target\n</code></pre>\n<p><br>\nI am calling the above function as  - <br>\n<code>dataset = datset.map( lambda img,target: (hair_augment(img),target))</code></p>\n<p>Maybe the reason for which this function raises exception is the use of \"OpenCV's\" function which are not built to deal with \"Tensors\" and hence I tried to use <strong>tf.bitwise.bitwise<strong>_and()</strong> in place of <strong>cv2.bitwise</strong>and()</strong> . That function does not overlap image with the mask but perform <strong>Bitwise-And</strong> operation and hence the results were dissatisfactory. </p>\n<p>I have also tried to use <strong>tf.py<em>function</em></strong> and <strong>tf.numpyfunction</strong>  but all of them raises exceptions and ultimately I am at the point where I started.</p>\n<p>One approach I have also tried (which is kind of naive) is that I extracted out \"images\" from the dataset, iterating over them using <strong>images.take(-1)</strong>  and passing them to <code>hair_augment</code> function which happened to be successful but now I had to combine \"images\" and \"targets\" again and so I used <code>tf.data.Dataset.zip()</code>  for that which also didn't proved to be a successful approach.</p>\n<p>It seems like similar issues have been asked over various platforms and no clear answer for has been obtained.</p>\n<p>The dataset and code for data pipeline is taken from Chris Deotte's <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a>.<br>\nRelevant Issues - </p>\n<ul>\n<li><a href=\"https://github.com/tensorflow/tensorflow/issues/38762\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/38762</a></li>\n<li><a href=\"https://github.com/tensorflow/tensorflow/issues/27519\" target=\"_blank\">https://github.com/tensorflow/tensorflow/issues/27519</a></li>\n<li><a href=\"https://stackoverflow.com/questions/56665868/tensor-numpy-not-working-in-tensorflow-data-dataset-throws-the-error-attribu\" target=\"_blank\">https://stackoverflow.com/questions/56665868/tensor-numpy-not-working-in-tensorflow-data-dataset-throws-the-error-attribu</a></li>\n<li><a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316\" target=\"_blank\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316</a></li>\n</ul>",
      "rawMarkdown": "Recently I've been searching for new augmentation strategies to implement into my solution and came across [this](https://www.kaggle.com/synked/really-realistic-hair-augmentations) work by **Ayaan**. The approach aims to introduce hairs like structure in the images and it really looks like as they are natural. It was implemented to deal with the image as a numpy array but I have a `tf.data` pipeline to perform data transformations and the images are \"Tensor\" objects and also the execution is in Graph mode so I cannot use `image.numpy()` here.\n \nThe `tf.data` pipeline is same as described in **Chris Deotte's**  [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords).\n\nThe function which I want to implement -\n```\ndef hair_augment(img,target):\n    img = tf.image.decode_jpeg(img,channels=3)\n    img = cv2.bitwise_and(img,img,mask=mask)\n    return img,target\n``` \nI am calling the above function as  - \n`dataset = datset.map( lambda img,target: (hair_augment(img),target))`\n\nMaybe the reason for which this function raises exception is the use of \"OpenCV's\" function which are not built to deal with \"Tensors\" and hence I tried to use **tf.bitwise.bitwise___and()** in place of **cv2.bitwise__and()** . That function does not overlap image with the mask but perform **Bitwise-And** operation and hence the results were dissatisfactory. \n\nI have also tried to use **tf.py_function** and **tf.numpy_function**  but all of them raises exceptions and ultimately I am at the point where I started.\n\nOne approach I have also tried (which is kind of naive) is that I extracted out \"images\" from the dataset, iterating over them using **images.take(-1)**  and passing them to `hair_augment` function which happened to be successful but now I had to combine \"images\" and \"targets\" again and so I used `tf.data.Dataset.zip()`  for that which also didn't proved to be a successful approach.\n\nIt seems like similar issues have been asked over various platforms and no clear answer for has been obtained.\n\nThe dataset and code for data pipeline is taken from Chris Deotte's [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords).\nRelevant Issues - \n- https://github.com/tensorflow/tensorflow/issues/38762\n- https://github.com/tensorflow/tensorflow/issues/27519\n- https://stackoverflow.com/questions/56665868/tensor-numpy-not-working-in-tensorflow-data-dataset-throws-the-error-attribu\n- https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316",
      "votes": 2
    },
    {
      "id": 944527,
      "postDate": "2020-07-25T07:09:00.037Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 944539,
          "postDate": "2020-07-25T07:15:39.340Z",
          "content": "<p>Thanks <a href=\"/synked\">@synked</a>! Please have a look at the description above to save your time by avoiding the failed approaches that I used. </p>",
          "rawMarkdown": "Thanks @synked! Please have a look at the description above to save your time by avoiding the failed approaches that I used. "
        }
      ]
    },
    {
      "id": 944460,
      "postDate": "2020-07-25T06:11:32.147Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 944335,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-25T03:04:37.173000",
      "content": "<p>I don't think TensorFlow <code>tf.data.Dataset</code> allows us to use any external libraries. I posted links about how to perform all the types of augmentation <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169721#944331\">here</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 944384,
          "author_name": "Sarthak khandelwal",
          "author_url": "",
          "post_date": "2020-07-25T04:27:19.540000",
          "content": "<p>Thanks for the info <a href=\"/cdeotte\">@cdeotte</a> :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 938612,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2020-07-21T16:21:39.220000",
      "content": "<p>I have had more success using opencv with tensorflow when I make a class rather than a def.</p>\n\n<p>The current leader of this competition shared a <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168371\">great augmentation code</a> </p>\n\n<p>I am going to paste my hair class - but I have almost no success in getting the code to paste looking correct (this time the first couple of lines are bugerred and the def_call should be tabbed over - all the code is in the class).  In my version of Chris notebook I would just call the class in the def where I am doing my other augmentations.</p>\n\n<p><strong>AdvancedHairAugmentation(image)</strong></p>\n\n<p>`class AdvancedHairAugmentation:\n    def <strong>init</strong>(self, hairs: int = 4, hairs_folder: str = \"\", p: float = 0.5):\n        self.hairs = hairs\n        self.hairs_folder = hairs_folder\n        self.p = p</p>\n\n<pre><code>def __call__(self, img):\n    if random.random() &lt; self.p:\n        n_hairs = random.randint(0, self.hairs)\n\n        if not n_hairs:\n            return img\n\n        # height, width, _ = img.shape  # target image width and height\n        hair_images = [im for im in os.listdir(self.hairs_folder) if 'png' in im]\n\n        for _ in range(n_hairs):\n            hair = cv2.imread(os.path.join(self.hairs_folder, random.choice(hair_images)))\n            hair = cv2.flip(hair, random.choice([-1, 0, 1]))\n            hair = cv2.rotate(hair, random.choice([0, 1, 2]))\n\n            h_height, h_width, _ = hair.shape  # hair image width and height\n            roi_ho = random.randint(0, img.shape[0] - hair.shape[0])\n            roi_wo = random.randint(0, img.shape[1] - hair.shape[1])\n            roi = img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width]\n\n            img2gray = cv2.cvtColor(hair, cv2.COLOR_BGR2GRAY)\n            ret, mask = cv2.threshold(img2gray, 10, 255, cv2.THRESH_BINARY)\n            mask_inv = cv2.bitwise_not(mask)\n            img_bg = cv2.bitwise_and(roi, roi, mask=mask_inv)\n            hair_fg = cv2.bitwise_and(hair, hair, mask=mask)\n\n            dst = cv2.add(img_bg, hair_fg)\n            img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width] = dst\n\n        return img`\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 943704,
          "author_name": "Sarthak khandelwal",
          "author_url": "",
          "post_date": "2020-07-24T14:22:24.503000",
          "content": "<p>Hey <a href=\"/pcjimmmy\">@pcjimmmy</a> a great implementation 😄. Did you called this class in <code>tf.data</code> pipeline (as an argument to <code>map</code> function)? And if yes then how did you convert \"img\" which is of type \"Tensor\" and which does not support opencv operations. Thanks and sorry for late reply. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944322,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-25T02:44:39.950000",
          "content": "<p>I have a def augment cell similiar to most that you see.  As a class I call the hair augment as shown below as part of my def.  </p>\n\n<p>I also have class functions for hair removal and microscope augmentations that are discussed in a number of posts.</p>\n\n<p>I call the augments in the map\ndataset = dataset.map(data_augment, num_parallel_calls=AUTO)</p>\n\n<p>`\ndef data_augment(image, label):</p>\n\n<pre><code>HairRemoval(image)\nAdvancedHairAugmentation(image)\nMicroscope(image)\n\nimage = tf.image.random_flip_left_right(image)\nimage = tf.image.random_flip_up_down(image)\nimage = tf.image.random_brightness(image, 0.3)\nimage = tf.image.random_contrast(image, 0.6, 2.0)\nimage = tf.image.random_hue(image, 0.5)\nimage = tf.image.random_saturation(image, 0.25, 2.5)\n\nreturn image, label   `\n</code></pre>\n\n<p>PS - Kaggle loves to kill underscores in pastes - so apparent typos are kaggle doing its thing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944385,
          "author_name": "Sarthak khandelwal",
          "author_url": "",
          "post_date": "2020-07-25T04:29:16.017000",
          "content": "<p>Thanks! But still, how did you managed to convert \"Tensor\" to ndarray?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944409,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-25T05:03:42.483000",
          "content": "<p>I have forked Chris triple on my local PC and ran it as shown.   </p>\n\n<p>Look at his transform def he has a note:\n     input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]</p>\n\n<p>My class should have the same note.</p>\n\n<p>His prepare_image  with my class useage:  (as usual my pastes go bad in kaggle)</p>\n\n<p>`def prepare_image(img, augment=True, dim=256): <br>\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.cast(img, tf.float32) / 255.0</p>\n\n<pre><code>img = tf.image.resize(img, [b, b],\n                            method=tf.image.ResizeMethod.NEAREST_NEIGHBOR)\nimg = transform(img,DIM=dim)\nimg = tf.image.random_flip_up_down(img)\nimg = tf.image.random_hue(img, 0.5)\nimg = tf.image.random_saturation(img, 0.25, 2.5)\nimg = tf.image.random_contrast(img, 0.6, 2.0)\nimg = tf.image.random_brightness(img, 0.3)\nHairRemoval(img)\nAdvancedHairAugmentation(img)\n\nimg = tf.reshape(img, [dim,dim, 3])\n\n\nreturn img`\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944542,
          "author_name": "Sarthak khandelwal",
          "author_url": "",
          "post_date": "2020-07-25T07:16:59.760000",
          "content": "<p>Thanks! I would surely try this and would let you know if stuck in implementing it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 945061,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-25T14:44:25.610000",
          "content": "<p>FYI -I not currently using any of the 3 augmentations in my models.  While it works, it's expensive to add.  When I am a few days away from the end of this and have some models I like I will add them back in.  So my suggestion would be - getting it working - decide if the improvement in your model exists and than comment them out until the last week.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944527,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T07:09:00.037000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 944539,
          "author_name": "Sarthak khandelwal",
          "author_url": "",
          "post_date": "2020-07-25T07:15:39.340000",
          "content": "<p>Thanks <a href=\"/synked\">@synked</a>! Please have a look at the description above to save your time by avoiding the failed approaches that I used. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944460,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T06:11:32.147000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "944335": "I don't think TensorFlow `tf.data.Dataset` allows us to use any external libraries. I posted links about how to perform all the types of augmentation [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169721#944331",
    "938612": "I have had more success using opencv with tensorflow when I make a class rather than a def.\n\nThe current leader of this competition shared a [great augmentation code](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168371) \n\nI am going to paste my hair class - but I have almost no success in getting the code to paste looking correct (this time the first couple of lines are bugerred and the def_call should be tabbed over - all the code is in the class).  In my version of Chris notebook I would just call the class in the def where I am doing my other augmentations.\n\n\n\n**AdvancedHairAugmentation(image)**\n\n`class AdvancedHairAugmentation:\n    def __init__(self, hairs: int = 4, hairs_folder: str = \"\", p: float = 0.5):\n        self.hairs = hairs\n        self.hairs_folder = hairs_folder\n        self.p = p\n\n    def __call__(self, img):\n        if random.random() &lt; self.p:\n            n_hairs = random.randint(0, self.hairs)\n\n            if not n_hairs:\n                return img\n\n            # height, width, _ = img.shape  # target image width and height\n            hair_images = [im for im in os.listdir(self.hairs_folder) if 'png' in im]\n\n            for _ in range(n_hairs):\n                hair = cv2.imread(os.path.join(self.hairs_folder, random.choice(hair_images)))\n                hair = cv2.flip(hair, random.choice([-1, 0, 1]))\n                hair = cv2.rotate(hair, random.choice([0, 1, 2]))\n\n                h_height, h_width, _ = hair.shape  # hair image width and height\n                roi_ho = random.randint(0, img.shape[0] - hair.shape[0])\n                roi_wo = random.randint(0, img.shape[1] - hair.shape[1])\n                roi = img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width]\n\n                img2gray = cv2.cvtColor(hair, cv2.COLOR_BGR2GRAY)\n                ret, mask = cv2.threshold(img2gray, 10, 255, cv2.THRESH_BINARY)\n                mask_inv = cv2.bitwise_not(mask)\n                img_bg = cv2.bitwise_and(roi, roi, mask=mask_inv)\n                hair_fg = cv2.bitwise_and(hair, hair, mask=mask)\n\n                dst = cv2.add(img_bg, hair_fg)\n                img[roi_ho:roi_ho + h_height, roi_wo:roi_wo + h_width] = dst\n\n            return img`\n\n\n",
    "938116": "Recently I've been searching for new augmentation strategies to implement into my solution and came across [this](https://www.kaggle.com/synked/really-realistic-hair-augmentations) work by **Ayaan**. The approach aims to introduce hairs like structure in the images and it really looks like as they are natural. It was implemented to deal with the image as a numpy array but I have a `tf.data` pipeline to perform data transformations and the images are \"Tensor\" objects and also the execution is in Graph mode so I cannot use `image.numpy()` here.\n \nThe `tf.data` pipeline is same as described in **Chris Deotte's**  [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords).\n\nThe function which I want to implement -\n```\ndef hair_augment(img,target):\n    img = tf.image.decode_jpeg(img,channels=3)\n    img = cv2.bitwise_and(img,img,mask=mask)\n    return img,target\n``` \nI am calling the above function as  - \n`dataset = datset.map( lambda img,target: (hair_augment(img),target))`\n\nMaybe the reason for which this function raises exception is the use of \"OpenCV's\" function which are not built to deal with \"Tensors\" and hence I tried to use **tf.bitwise.bitwise___and()** in place of **cv2.bitwise__and()** . That function does not overlap image with the mask but perform **Bitwise-And** operation and hence the results were dissatisfactory. \n\nI have also tried to use **tf.py_function** and **tf.numpy_function**  but all of them raises exceptions and ultimately I am at the point where I started.\n\nOne approach I have also tried (which is kind of naive) is that I extracted out \"images\" from the dataset, iterating over them using **images.take(-1)**  and passing them to `hair_augment` function which happened to be successful but now I had to combine \"images\" and \"targets\" again and so I used `tf.data.Dataset.zip()`  for that which also didn't proved to be a successful approach.\n\nIt seems like similar issues have been asked over various platforms and no clear answer for has been obtained.\n\nThe dataset and code for data pipeline is taken from Chris Deotte's [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords).\nRelevant Issues - \n- https://github.com/tensorflow/tensorflow/issues/38762\n- https://github.com/tensorflow/tensorflow/issues/27519\n- https://stackoverflow.com/questions/56665868/tensor-numpy-not-working-in-tensorflow-data-dataset-throws-the-error-attribu\n- https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316",
    "944527": "",
    "944460": ""
  }
}