{
  "id": 169721,
  "title": "Coarse Dropout and Cutout Augmentation GPU/TPU",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169721",
  "author_name": "Chris Deotte",
  "post_date": "2020-07-24T23:21:05.752000",
  "votes": 114,
  "comment_count": 87,
  "views": 0,
  "content": "<p>When using TensorFlow's <code>tf.data.Dataset()</code> we cannot use Albumenations to do our augmentation. Instead we must write our own functions in TensorFlow. Previously I posted how to perform (1) Rotation, Sheer, Zoom, Shift for GPU/TPU <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a> and <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> (2) Cutmix and Mixup for GPU/TPU <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">here</a></p>\n\n<h1>Coarse Dropout and Cutout Augmentation for GPU/TPU</h1>\n\n<p>Coarse Dropout and Cutout augmentation are techniques to prevent overfitting and encourage generalization. They randomly remove rectangles from training images. By removing portions of the images, we challenge our models to pay attention to the entire image because it never knows what part of the image will be present. (This is similar and different to dropout layer within a CNN).</p>\n\n<ul>\n<li>Cutout is the technique of removing 1 large rectangle of random size </li>\n<li>Coarse dropout is the technique of removing many small rectanges of similar size. </li>\n</ul>\n\n<p>By changing the parameters below, we can have either coarse dropout or cutout. (For cutout, you'll need to add <code>tf.random.uniform</code> for random size. I leave this as an exercise for the reader).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8db206436734b5f0c7b2179c2e152eb0%2FScreen%20Shot%202020-07-22%20at%201.28.49%20PM.png?generation=1595632223386869&amp;alt=media\" alt=\"\"></p>\n\n<h1>Starter Notebook</h1>\n\n<p>I've posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a> showing how to do coarse dropout with TFRecords and <code>tf.data.Dataset</code>. The starter notebook also demonstrates upsampling (oversampling) the minority class.</p>\n\n<h1>Code</h1>\n\n<p>The reason that this code looks overly complicated just to turn a portion of an array to zeros is because you cannot do <code>image[ya:yb,xa:xb,:] = 0</code> in TensorFlow. TF does not allow index assignment. Instead we break an image into five pieces and then concatenate the five pieces while replacing the middle piece with all zeros.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F476c96c48b8b31969e00b55b26cef2f3%2Fcut.jpg?generation=1595632770269562&amp;alt=media\" alt=\"\"></p>\n\n<pre><code>def dropout(image, DIM=256, PROBABILITY = 0.75, CT = 8, SZ = 0.2):\n    # input - one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed\n\n    # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast( tf.random.uniform([],0,1) &amp;lt; PROBABILITY, tf.int32)\n    if (P == 0)|(CT == 0)|(SZ == 0): return image\n\n    for k in range( CT ):\n        # CHOOSE RANDOM LOCATION\n        x = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n        y = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n        # COMPUTE SQUARE \n        WIDTH = tf.cast( SZ*DIM,tf.int32) * P\n        ya = tf.math.maximum(0,y-WIDTH//2)\n        yb = tf.math.minimum(DIM,y+WIDTH//2)\n        xa = tf.math.maximum(0,x-WIDTH//2)\n        xb = tf.math.minimum(DIM,x+WIDTH//2)\n        # DROPOUT IMAGE\n        one = image[ya:yb,0:xa,:]\n        two = tf.zeros([yb-ya,xb-xa,3]) \n        three = image[ya:yb,xb:DIM,:]\n        middle = tf.concat([one,two,three],axis=1)\n        image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM,:,:]],axis=0)\n\n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR \n    image = tf.reshape(image,[DIM,DIM,3])\n    return image\n</code></pre>",
  "messages": [
    {
      "id": 944208,
      "postDate": "2020-07-24T23:21:05.753Z",
      "content": "<p>When using TensorFlow's <code>tf.data.Dataset()</code> we cannot use Albumenations to do our augmentation. Instead we must write our own functions in TensorFlow. Previously I posted how to perform (1) Rotation, Sheer, Zoom, Shift for GPU/TPU <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a> and <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\">here</a> (2) Cutmix and Mixup for GPU/TPU <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">here</a></p>\n\n<h1>Coarse Dropout and Cutout Augmentation for GPU/TPU</h1>\n\n<p>Coarse Dropout and Cutout augmentation are techniques to prevent overfitting and encourage generalization. They randomly remove rectangles from training images. By removing portions of the images, we challenge our models to pay attention to the entire image because it never knows what part of the image will be present. (This is similar and different to dropout layer within a CNN).</p>\n\n<ul>\n<li>Cutout is the technique of removing 1 large rectangle of random size </li>\n<li>Coarse dropout is the technique of removing many small rectanges of similar size. </li>\n</ul>\n\n<p>By changing the parameters below, we can have either coarse dropout or cutout. (For cutout, you'll need to add <code>tf.random.uniform</code> for random size. I leave this as an exercise for the reader).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8db206436734b5f0c7b2179c2e152eb0%2FScreen%20Shot%202020-07-22%20at%201.28.49%20PM.png?generation=1595632223386869&amp;alt=media\" alt=\"\"></p>\n\n<h1>Starter Notebook</h1>\n\n<p>I've posted a starter notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a> showing how to do coarse dropout with TFRecords and <code>tf.data.Dataset</code>. The starter notebook also demonstrates upsampling (oversampling) the minority class.</p>\n\n<h1>Code</h1>\n\n<p>The reason that this code looks overly complicated just to turn a portion of an array to zeros is because you cannot do <code>image[ya:yb,xa:xb,:] = 0</code> in TensorFlow. TF does not allow index assignment. Instead we break an image into five pieces and then concatenate the five pieces while replacing the middle piece with all zeros.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F476c96c48b8b31969e00b55b26cef2f3%2Fcut.jpg?generation=1595632770269562&amp;alt=media\" alt=\"\"></p>\n\n<pre><code>def dropout(image, DIM=256, PROBABILITY = 0.75, CT = 8, SZ = 0.2):\n    # input - one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed\n\n    # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast( tf.random.uniform([],0,1) &amp;lt; PROBABILITY, tf.int32)\n    if (P == 0)|(CT == 0)|(SZ == 0): return image\n\n    for k in range( CT ):\n        # CHOOSE RANDOM LOCATION\n        x = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n        y = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n        # COMPUTE SQUARE \n        WIDTH = tf.cast( SZ*DIM,tf.int32) * P\n        ya = tf.math.maximum(0,y-WIDTH//2)\n        yb = tf.math.minimum(DIM,y+WIDTH//2)\n        xa = tf.math.maximum(0,x-WIDTH//2)\n        xb = tf.math.minimum(DIM,x+WIDTH//2)\n        # DROPOUT IMAGE\n        one = image[ya:yb,0:xa,:]\n        two = tf.zeros([yb-ya,xb-xa,3]) \n        three = image[ya:yb,xb:DIM,:]\n        middle = tf.concat([one,two,three],axis=1)\n        image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM,:,:]],axis=0)\n\n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR \n    image = tf.reshape(image,[DIM,DIM,3])\n    return image\n</code></pre>",
      "rawMarkdown": "When using TensorFlow's `tf.data.Dataset()` we cannot use Albumenations to do our augmentation. Instead we must write our own functions in TensorFlow. Previously I posted how to perform (1) Rotation, Sheer, Zoom, Shift for GPU/TPU [here][1] and [here][3] (2) Cutmix and Mixup for GPU/TPU [here][2]\n\n# Coarse Dropout and Cutout Augmentation for GPU/TPU\nCoarse Dropout and Cutout augmentation are techniques to prevent overfitting and encourage generalization. They randomly remove rectangles from training images. By removing portions of the images, we challenge our models to pay attention to the entire image because it never knows what part of the image will be present. (This is similar and different to dropout layer within a CNN).\n\n* Cutout is the technique of removing 1 large rectangle of random size \n* Coarse dropout is the technique of removing many small rectanges of similar size. \n\nBy changing the parameters below, we can have either coarse dropout or cutout. (For cutout, you'll need to add `tf.random.uniform` for random size. I leave this as an exercise for the reader).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8db206436734b5f0c7b2179c2e152eb0%2FScreen%20Shot%202020-07-22%20at%201.28.49%20PM.png?generation=1595632223386869&amp;alt=media)\n\n\n# Starter Notebook\nI've posted a starter notebook [here][4] showing how to do coarse dropout with TFRecords and `tf.data.Dataset`. The starter notebook also demonstrates upsampling (oversampling) the minority class.\n\n# Code\nThe reason that this code looks overly complicated just to turn a portion of an array to zeros is because you cannot do `image[ya:yb,xa:xb,:] = 0` in TensorFlow. TF does not allow index assignment. Instead we break an image into five pieces and then concatenate the five pieces while replacing the middle piece with all zeros.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F476c96c48b8b31969e00b55b26cef2f3%2Fcut.jpg?generation=1595632770269562&amp;alt=media)\n\n    def dropout(image, DIM=256, PROBABILITY = 0.75, CT = 8, SZ = 0.2):\n        # input - one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n        # output - image with CT squares of side size SZ*DIM removed\n    \n        # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n        P = tf.cast( tf.random.uniform([],0,1) &lt; PROBABILITY, tf.int32)\n        if (P == 0)|(CT == 0)|(SZ == 0): return image\n    \n        for k in range( CT ):\n            # CHOOSE RANDOM LOCATION\n            x = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n            y = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n            # COMPUTE SQUARE \n            WIDTH = tf.cast( SZ*DIM,tf.int32) * P\n            ya = tf.math.maximum(0,y-WIDTH//2)\n            yb = tf.math.minimum(DIM,y+WIDTH//2)\n            xa = tf.math.maximum(0,x-WIDTH//2)\n            xb = tf.math.minimum(DIM,x+WIDTH//2)\n            # DROPOUT IMAGE\n            one = image[ya:yb,0:xa,:]\n            two = tf.zeros([yb-ya,xb-xa,3]) \n            three = image[ya:yb,xb:DIM,:]\n            middle = tf.concat([one,two,three],axis=1)\n            image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM,:,:]],axis=0)\n            \n        # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR \n        image = tf.reshape(image,[DIM,DIM,3])\n        return image\n\n[1]: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\n[2]: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\n[3]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[4]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
      "votes": 114
    },
    {
      "id": 953624,
      "postDate": "2020-07-31T23:34:25.127Z",
      "content": "<p>I've created a version that accepts batches of images, for anyone who is interested. I've tested it on a GPU setup, so I can't be sure this works on a TPU yet.</p>\n\n<p>edit: I see that the prob arg isn't actually utilized yet, I'll fix this tommorow.\nedit 2: probabilities now included :)\nedit 3: replaced a type in tf.scatter_nd where the image sizes were fixed to 128,128 -&gt; now adjusted for variable sizes</p>\n\n<p>```python\ndef coarse_dropout(images, square_size, num_squares, prob):\n    ''' Coarse image dropout \n    Coarse dropout masks are created by using tf.scatter_nd\n    Which can update matrices according to coordinates specified\n    by the user. </p>\n\n<pre><code>Parameters\n----------\nimages : tf Tensor\n    A batch of images of size [BS, H, W, C] \nsquare_size : float\n    A float between [0,1], a percentage of the image size\nnum_squares : int &amp;gt; 0\n    The number of dropout squares per single image, must be &amp;gt;0\nprob : float\n    A float between [0,1], probability that dropout is used for this batch\n\nYields\n------\nTYPE\n    DESCRIPTION.\n\n'''\n   # random dropout\nprob = tf.cast( tf.random.uniform([],0,1) &amp;lt; prob, tf.int32)\nif (prob == 0): return images\n\nimg_shape = images.shape\n_, h, w, c = img_shape[0], img_shape[1], img_shape[2], img_shape[3]\n\n# For some odd reason, the batch size is lost in the processing pipeline\n# seriously, how can it get lost ._., it literally says ds.batch()...\n# tensorflow shenanigans...\nbs = tf.cast(tf.reduce_sum(tf.ones_like(images)) / (images.shape[1] * images.shape[2] * images.shape[3]), tf.int32)\n\n\n# size of square in pixels (ssp), ASSUMING SQUARE IMAGES!\nssp = tf.cast(tf.math.ceil(h * square_size), tf.int32)\n\n# Create random start x coordinates\ncoords_x = tf.random.uniform([bs, num_squares], \n                             minval=0, \n                             maxval= (h - ssp), \n                             dtype=tf.int32) \n\n# Create ranges from start to the end of the line\n# [ssp, bs, num_squares]\ncoords_x = tf.linspace(coords_x, coords_x + ssp - 1, ssp)\n\n# [num_squares, bs, ssp]\ncoords_x = tf.cast(tf.transpose(coords_x), tf.int32)\n\n# Create random start y coordinates\ncoords_y = tf.random.uniform([bs, num_squares], \n                             minval=0, \n                             maxval= (h - ssp), \n                             dtype=tf.int32) \n\n# Create ranges from start to the end of the line\n# [ssp, bs, num_squares]\ncoords_y = tf.linspace(coords_y, coords_y + ssp - 1, ssp)\n\n# [num_squares, bs, ssp]\ncoords_y = tf.cast(tf.transpose(coords_y), tf.int32)\n\n\n# Create coordinate range combinations \n# and reshape to [bs, num_squares, 1, ssp * ssp]\ngrid_y = tf.reshape(tf.tile(coords_y, [1,1,ssp]), \n                    (bs, num_squares, 1, ssp * ssp))\n\n# and reshape to [bs, num_squares, ssp, ssp], transpose the inner matrices\ngrid_y = tf.transpose(tf.reshape(grid_y, \n                                 (bs, num_squares, ssp, ssp)), \n                      (0, 1, 3, 2))\n\n# Repeat for x coordinates\ngrid_x = tf.reshape(tf.tile(coords_x, [1,1,ssp]), \n                    (bs, num_squares, 1, ssp * ssp))\n\ngrid_x = tf.reshape(grid_x, (bs, num_squares, ssp, ssp))\n\n# Stack the grids into a single matrix\n# grid is [2, bs, num_squares, ssp, ssp]\ngrid = tf.stack([grid_y, grid_x], axis=0)\n\n# Transpose and reshape [ bs, ssp * ssp * num_squares, 2] \n# Creates an array of 2D coordinates ([[x1,y1], [x2, y2], ..., [xn, yn]])\n# over all squares (num_squares), for each combination (ssp*ssp) of coordinates \n# and each batch (bs)\ngrid = tf.reshape(tf.transpose(grid, (1, 4, 3, 2, 0)), \n                  (bs, ssp * ssp * num_squares, 2))\n\n# [bs, sz*sz*num_squares, 2]\n#grid = tf.reshape(grid, (bs, ssp*ssp*num_squares, 2))\n\n# create batch indices [0,..., bs] and reshape to [ssp * ssp * num_squares, bs]\nbatch_indices = tf.reshape(tf.tile(tf.range(0, bs), \n                                   [ssp * ssp * num_squares]), \n                           (ssp * ssp * num_squares, bs))\n\n# Transpose to get the right order, and reshape to match grid shape\nbatch_indices = tf.reshape(tf.transpose(batch_indices), \n                           (bs, ssp * ssp * num_squares, 1))\n\n# concatenate batch indices with the grid\n# these yield 3D coordinates like e.g.\n# [[bs0, x1, y1], [bs0, x2, y1], ..., [bsn, xn, yn]]\ngrid = tf.concat([batch_indices, grid], axis=2)\n\n# create a matrix of zeros, and update the matrix with the grid indices\n# this essentially creates a mask with coarse dropout squares\nmasks = tf.scatter_nd(grid[tf.newaxis,...], \n                        tf.ones([1,bs,ssp*ssp*num_squares]) * -1, \n                        shape=(bs, h, w)) +1\n\n# Due to overlap of squares, some get coordinates get updated twice\n# and result in values &amp;lt; -1, clip these values\nmasks = tf.clip_by_value(masks, 0, 1)\n\nreturn images * masks[..., tf.newaxis] \n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I've created a version that accepts batches of images, for anyone who is interested. I've tested it on a GPU setup, so I can't be sure this works on a TPU yet.\n\nedit: I see that the prob arg isn't actually utilized yet, I'll fix this tommorow.\nedit 2: probabilities now included :)\nedit 3: replaced a type in tf.scatter_nd where the image sizes were fixed to 128,128 -&gt; now adjusted for variable sizes\n\n```python\ndef coarse_dropout(images, square_size, num_squares, prob):\n    ''' Coarse image dropout \n    Coarse dropout masks are created by using tf.scatter_nd\n    Which can update matrices according to coordinates specified\n    by the user. \n\n    Parameters\n    ----------\n    images : tf Tensor\n        A batch of images of size [BS, H, W, C] \n    square_size : float\n        A float between [0,1], a percentage of the image size\n    num_squares : int &gt; 0\n        The number of dropout squares per single image, must be &gt;0\n    prob : float\n        A float between [0,1], probability that dropout is used for this batch\n\n    Yields\n    ------\n    TYPE\n        DESCRIPTION.\n\n    '''\n       # random dropout\n    prob = tf.cast( tf.random.uniform([],0,1) &lt; prob, tf.int32)\n    if (prob == 0): return images\n\n    img_shape = images.shape\n    _, h, w, c = img_shape[0], img_shape[1], img_shape[2], img_shape[3]\n    \n    # For some odd reason, the batch size is lost in the processing pipeline\n    # seriously, how can it get lost ._., it literally says ds.batch()...\n    # tensorflow shenanigans...\n    bs = tf.cast(tf.reduce_sum(tf.ones_like(images)) / (images.shape[1] * images.shape[2] * images.shape[3]), tf.int32)\n\n\n    # size of square in pixels (ssp), ASSUMING SQUARE IMAGES!\n    ssp = tf.cast(tf.math.ceil(h * square_size), tf.int32)\n\n    # Create random start x coordinates\n    coords_x = tf.random.uniform([bs, num_squares], \n                                 minval=0, \n                                 maxval= (h - ssp), \n                                 dtype=tf.int32) \n\n    # Create ranges from start to the end of the line\n    # [ssp, bs, num_squares]\n    coords_x = tf.linspace(coords_x, coords_x + ssp - 1, ssp)\n\n    # [num_squares, bs, ssp]\n    coords_x = tf.cast(tf.transpose(coords_x), tf.int32)\n\n    # Create random start y coordinates\n    coords_y = tf.random.uniform([bs, num_squares], \n                                 minval=0, \n                                 maxval= (h - ssp), \n                                 dtype=tf.int32) \n\n    # Create ranges from start to the end of the line\n    # [ssp, bs, num_squares]\n    coords_y = tf.linspace(coords_y, coords_y + ssp - 1, ssp)\n\n    # [num_squares, bs, ssp]\n    coords_y = tf.cast(tf.transpose(coords_y), tf.int32)\n\n\n    # Create coordinate range combinations \n    # and reshape to [bs, num_squares, 1, ssp * ssp]\n    grid_y = tf.reshape(tf.tile(coords_y, [1,1,ssp]), \n                        (bs, num_squares, 1, ssp * ssp))\n\n    # and reshape to [bs, num_squares, ssp, ssp], transpose the inner matrices\n    grid_y = tf.transpose(tf.reshape(grid_y, \n                                     (bs, num_squares, ssp, ssp)), \n                          (0, 1, 3, 2))\n\n    # Repeat for x coordinates\n    grid_x = tf.reshape(tf.tile(coords_x, [1,1,ssp]), \n                        (bs, num_squares, 1, ssp * ssp))\n\n    grid_x = tf.reshape(grid_x, (bs, num_squares, ssp, ssp))\n\n    # Stack the grids into a single matrix\n    # grid is [2, bs, num_squares, ssp, ssp]\n    grid = tf.stack([grid_y, grid_x], axis=0)\n\n    # Transpose and reshape [ bs, ssp * ssp * num_squares, 2] \n    # Creates an array of 2D coordinates ([[x1,y1], [x2, y2], ..., [xn, yn]])\n    # over all squares (num_squares), for each combination (ssp*ssp) of coordinates \n    # and each batch (bs)\n    grid = tf.reshape(tf.transpose(grid, (1, 4, 3, 2, 0)), \n                      (bs, ssp * ssp * num_squares, 2))\n\n    # [bs, sz*sz*num_squares, 2]\n    #grid = tf.reshape(grid, (bs, ssp*ssp*num_squares, 2))\n\n    # create batch indices [0,..., bs] and reshape to [ssp * ssp * num_squares, bs]\n    batch_indices = tf.reshape(tf.tile(tf.range(0, bs), \n                                       [ssp * ssp * num_squares]), \n                               (ssp * ssp * num_squares, bs))\n\n    # Transpose to get the right order, and reshape to match grid shape\n    batch_indices = tf.reshape(tf.transpose(batch_indices), \n                               (bs, ssp * ssp * num_squares, 1))\n\n    # concatenate batch indices with the grid\n    # these yield 3D coordinates like e.g.\n    # [[bs0, x1, y1], [bs0, x2, y1], ..., [bsn, xn, yn]]\n    grid = tf.concat([batch_indices, grid], axis=2)\n\n    # create a matrix of zeros, and update the matrix with the grid indices\n    # this essentially creates a mask with coarse dropout squares\n    masks = tf.scatter_nd(grid[tf.newaxis,...], \n                            tf.ones([1,bs,ssp*ssp*num_squares]) * -1, \n                            shape=(bs, h, w)) +1\n    \n    # Due to overlap of squares, some get coordinates get updated twice\n    # and result in values &lt; -1, clip these values\n    masks = tf.clip_by_value(masks, 0, 1)\n\n    return images * masks[..., tf.newaxis] \n```",
      "votes": 7,
      "replies": [
        {
          "id": 954270,
          "postDate": "2020-08-01T15:24:34.653Z",
          "content": "<p>Thanks <a href=\"/stephanovich\">@stephanovich</a> ! when i get time i will play with this. I'm curious to benchmark the speed and see if this is faster than \"per image\".</p>\n\n<p>I've noticed when using \"per image\", if we cutout more than 16 squares it will begin to slow down training, perhaps your \"per batch\" code can do more than 16 and not slow down training.</p>",
          "rawMarkdown": "Thanks @stephanovich ! when i get time i will play with this. I'm curious to benchmark the speed and see if this is faster than \"per image\".\n\nI've noticed when using \"per image\", if we cutout more than 16 squares it will begin to slow down training, perhaps your \"per batch\" code can do more than 16 and not slow down training.",
          "votes": 1
        },
        {
          "id": 955084,
          "postDate": "2020-08-02T10:17:17.990Z",
          "content": "<p>Edit: you might want to try the augmentation from this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171778\">discussion</a> and <a href=\"https://www.kaggle.com/benboren/tfrecord-progressive-sprinkles\">notebook</a>, it seems to be a more generalized version of coarse dropout, and does not use for loops either (i.e. vectorized).</p>\n\n<p>So I ran a small benchmark myself on a cpu with 32 threads, here's the results:\nThe experiments are run on 128*128 images with a variable number of squares each and a square size of 0.1 * the size of the image. The probability is set to one in both cases. I parsed through 60000 examples, which in my opinion was already enough to at least see a difference.\nI've included 3 runs: \n1. one without any sort of augmentation, which acts as the lower bound (or golden standard) for time benchmarking (it can never go faster than this).\n<code>python\nCPU times: user 11min 27s, sys: 30.9 s, total: 11min 58s\nWall time: 29.2 s\n</code></p>\n\n<p>2 . The new batch version. This version first prepares the image, then batches the images, and then augments them. Otherwise, we would not see any improvements.\n```  python</p>\n\n<h1>Using 8 squares</h1>\n\n<p>CPU times: user 28min 6s, sys: 1min 38s, total: 29min 44s\nWall time: 1min 12s</p>\n\n<h1>Using 16 squares</h1>\n\n<p>CPU times: user 33min, sys: 1min 36s, total: 34min 36s\nWall time: 1min 17s</p>\n\n<p>```</p>\n\n<p>3 . Your version. In the processing pipeline, the images are first prepared through prepare_image, then are augmented, and finally batched.\n```python</p>\n\n<h1>Using 8 squares</h1>\n\n<p>CPU times: user 50min 36s, sys: 2min 37s, total: 53min 14s\nWall time: 2min 4s</p>\n\n<h1>Using 16 squares</h1>\n\n<p>CPU times: user 1h 15min 50s, sys: 2min 41s, total: 1h 18min 32s\nWall time: 2min 48s\n```</p>\n\n<p>It seems the batch version does run faster, at the cost of some higher memory footprint (due to vectorization)</p>\n\n<p>Of course, the times can decrease/increase depending on the image sizes, number of squares which you already mentioned above, and possibly other factors.</p>",
          "rawMarkdown": "Edit: you might want to try the augmentation from this [discussion](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171778) and [notebook](https://www.kaggle.com/benboren/tfrecord-progressive-sprinkles), it seems to be a more generalized version of coarse dropout, and does not use for loops either (i.e. vectorized).\n\nSo I ran a small benchmark myself on a cpu with 32 threads, here's the results:\nThe experiments are run on 128*128 images with a variable number of squares each and a square size of 0.1 * the size of the image. The probability is set to one in both cases. I parsed through 60000 examples, which in my opinion was already enough to at least see a difference.\nI've included 3 runs: \n1. one without any sort of augmentation, which acts as the lower bound (or golden standard) for time benchmarking (it can never go faster than this).\n```python\nCPU times: user 11min 27s, sys: 30.9 s, total: 11min 58s\nWall time: 29.2 s\n```\n\n2 . The new batch version. This version first prepares the image, then batches the images, and then augments them. Otherwise, we would not see any improvements.\n```  python\n\n# Using 8 squares\nCPU times: user 28min 6s, sys: 1min 38s, total: 29min 44s\nWall time: 1min 12s\n\n# Using 16 squares\nCPU times: user 33min, sys: 1min 36s, total: 34min 36s\nWall time: 1min 17s\n\n```\n\n3 . Your version. In the processing pipeline, the images are first prepared through prepare_image, then are augmented, and finally batched.\n```python\n\n# Using 8 squares\nCPU times: user 50min 36s, sys: 2min 37s, total: 53min 14s\nWall time: 2min 4s\n\n# Using 16 squares\nCPU times: user 1h 15min 50s, sys: 2min 41s, total: 1h 18min 32s\nWall time: 2min 48s\n```\n\n\nIt seems the batch version does run faster, at the cost of some higher memory footprint (due to vectorization)\n\nOf course, the times can decrease/increase depending on the image sizes, number of squares which you already mentioned above, and possibly other factors.",
          "votes": 4
        }
      ]
    },
    {
      "id": 945580,
      "postDate": "2020-07-26T01:30:08.590Z",
      "content": "<p>Notice how experiment 1 <strong>without</strong> dropout has a larger gap between train AUC and validation AUC. While experiment 2 <strong>with</strong> dropout has a smaller gap. In the below example, dropout has helped the CNN generalize thus <strong>increasing validation</strong> AUC while preventing training AUC from overfitting (i.e. <strong>reducing training</strong> AUC). Notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a></p>\n\n<h2>WITHOUT coarse dropout</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F731ced48b76769659f1358d253f137af%2FScreen%20Shot%202020-07-25%20at%206.20.23%20PM.png?generation=1595726783638841&amp;alt=media\" alt=\"\"></p>\n\n<h2>WITH coarse dropout</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff8547390e3c0c420be6336ed92f44dae%2FScreen%20Shot%202020-07-25%20at%206.20.58%20PM.png?generation=1595726813782566&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Notice how experiment 1 **without** dropout has a larger gap between train AUC and validation AUC. While experiment 2 **with** dropout has a smaller gap. In the below example, dropout has helped the CNN generalize thus **increasing validation** AUC while preventing training AUC from overfitting (i.e. **reducing training** AUC). Notebook [here][1]\n## WITHOUT coarse dropout\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F731ced48b76769659f1358d253f137af%2FScreen%20Shot%202020-07-25%20at%206.20.23%20PM.png?generation=1595726783638841&amp;alt=media)\n## WITH coarse dropout\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff8547390e3c0c420be6336ed92f44dae%2FScreen%20Shot%202020-07-25%20at%206.20.58%20PM.png?generation=1595726813782566&amp;alt=media)\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\n\n",
      "votes": 5
    },
    {
      "id": 944950,
      "postDate": "2020-07-25T13:22:38.833Z",
      "content": "<p>This is really useful for tensorflow users in this competition <a href=\"/cdeotte\">@cdeotte</a> :D </p>",
      "rawMarkdown": "This is really useful for tensorflow users in this competition @cdeotte :D ",
      "votes": 3
    },
    {
      "id": 944692,
      "postDate": "2020-07-25T09:00:35.247Z",
      "content": "<p>UPDATE: If your forked version 15 or earlier of the starter notebook, note that it does <strong>not</strong> apply <strong>dropout</strong>. (It only applies <strong>upsample</strong>). During training, in the call to <code>get_dataset()</code>, I forgot to pass the dropout parameters (so they use the default of zero). In the 3 lines below, the last line was missing in notebook version 15 and earlier. I just committed notebook now. Versions 16 onward will have this line added and <strong>will apply</strong> dropout.</p>\n\n<pre><code>        get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n            dim=IMG_SIZES[fold], batch_size = BATCH_SIZES[fold],\n            droprate = DROP_FREQ[fold], dropct = DROP_CT[fold], dropsize = DROP_SIZE[fold])\n</code></pre>\n\n<p>We also need to add this last line to validation TTA prediction and test TTA prediction.</p>",
      "rawMarkdown": "UPDATE: If your forked version 15 or earlier of the starter notebook, note that it does **not** apply **dropout**. (It only applies **upsample**). During training, in the call to `get_dataset()`, I forgot to pass the dropout parameters (so they use the default of zero). In the 3 lines below, the last line was missing in notebook version 15 and earlier. I just committed notebook now. Versions 16 onward will have this line added and **will apply** dropout.\n\n            get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n                dim=IMG_SIZES[fold], batch_size = BATCH_SIZES[fold],\n                droprate = DROP_FREQ[fold], dropct = DROP_CT[fold], dropsize = DROP_SIZE[fold])\n\nWe also need to add this last line to validation TTA prediction and test TTA prediction.",
      "votes": 3
    },
    {
      "id": 954351,
      "postDate": "2020-08-01T16:46:06.593Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you for your great work! Please allow me one question.\nHow do you think about the difference between Coarse Dropout and GridMask ?</p>",
      "rawMarkdown": "@cdeotte Thank you for your great work! Please allow me one question.\nHow do you think about the difference between Coarse Dropout and GridMask ?",
      "votes": 1,
      "replies": [
        {
          "id": 954356,
          "postDate": "2020-08-01T16:52:39.810Z",
          "content": "<p>I have never used GridMask before. There is an implementation of GridMask for TensorFlow <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132986\">here</a> by Xie29. </p>\n\n<p>(Note: I have not tested Xie29's implementation. If you use it, i would compare training <strong>with</strong> and <strong>without</strong> it to make sure the code is efficient and doesn't slow down training).</p>",
          "rawMarkdown": "I have never used GridMask before. There is an implementation of GridMask for TensorFlow [here][1] by Xie29. \n\n(Note: I have not tested Xie29's implementation. If you use it, i would compare training **with** and **without** it to make sure the code is efficient and doesn't slow down training).\n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132986",
          "votes": 1
        },
        {
          "id": 954362,
          "postDate": "2020-08-01T17:05:21.667Z",
          "content": "<p>Thank you for your quick reply and suggestion. I plan to try both of them.\nAgain, I appreciate your nice works!</p>",
          "rawMarkdown": "Thank you for your quick reply and suggestion. I plan to try both of them.\nAgain, I appreciate your nice works!",
          "votes": 1
        }
      ]
    },
    {
      "id": 952933,
      "postDate": "2020-07-31T10:34:42.307Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Hi Chris! I really would like to try to find ensembling weights by using a machine learning algorithm and cv oof auc scores. Did you ever try this? Do you have an url or tip how to do this?\nBest Roman</p>",
      "rawMarkdown": "@cdeotte Hi Chris! I really would like to try to find ensembling weights by using a machine learning algorithm and cv oof auc scores. Did you ever try this? Do you have an url or tip how to do this?\nBest Roman",
      "votes": 1,
      "replies": [
        {
          "id": 953123,
          "postDate": "2020-07-31T14:47:45.470Z",
          "content": "<p>This is what all the top competitors do in every competition. This is the correct way to ensemble. (The incorrect way is trial and error using feedback from public LB because public test set contains less data than train set).</p>\n\n<p>You train multiple models using the same KFold folds and save all your OOF. Then you find weights with the OOF. For example if you have 3 models, then you find <code>w1, w2, w3</code> with <code>w1+w2+w3 = 1</code> where you maximize AUC below</p>\n\n<pre><code>oof = w1*oof1 + w2*oof2 + w3*oof3\nAUC = roc_auc_score(TRUE, oof)\n</code></pre>\n\n<p>Lastly you use these weights for your Kaggle submission as in</p>\n\n<pre><code>pred = w1*sub1 + w2*sub2 + w3*sub3\nsub = pd.DataFrame(dict(image_name=NAMES, target=PRED))\nsub.to_csv('submission.csv',index=False)\n</code></pre>\n\n<p>The simplest thing is to (1) find these weights by hand or grid search (i.e. with nested for-loops). Or (2) you can find these weights by fitting a linear or logistic regression model. (Since linear and logistic regression are simple models, you don't need holdout set, just <code>fit_transform</code> all the oof against true). Or (3) if you want to use a more complex stage 2 model like XGB, you need to set up a holdout set.</p>",
          "rawMarkdown": "This is what all the top competitors do in every competition. This is the correct way to ensemble. (The incorrect way is trial and error using feedback from public LB because public test set contains less data than train set).\n\nYou train multiple models using the same KFold folds and save all your OOF. Then you find weights with the OOF. For example if you have 3 models, then you find `w1, w2, w3` with `w1+w2+w3 = 1` where you maximize AUC below\n\n    oof = w1*oof1 + w2*oof2 + w3*oof3\n    AUC = roc_auc_score(TRUE, oof)\n\nLastly you use these weights for your Kaggle submission as in\n\n    pred = w1*sub1 + w2*sub2 + w3*sub3\n    sub = pd.DataFrame(dict(image_name=NAMES, target=PRED))\n    sub.to_csv('submission.csv',index=False)\n\nThe simplest thing is to (1) find these weights by hand or grid search (i.e. with nested for-loops). Or (2) you can find these weights by fitting a linear or logistic regression model. (Since linear and logistic regression are simple models, you don't need holdout set, just `fit_transform` all the oof against true). Or (3) if you want to use a more complex stage 2 model like XGB, you need to set up a holdout set.\n\n    ",
          "votes": 8,
          "replies": [
            {
              "id": 967043,
              "postDate": "2020-08-11T21:33:02.200Z",
              "content": "<p>Hello, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the above pointer.  I was trying to set weight values based on the CV score rather than LB feedback with different trails and errors.  However, I totally agree with <a href=\"https://www.kaggle.com/romanweilguny\" target=\"_blank\">@romanweilguny</a> that we should use ML algorithm to find an optimal weight for the ensemble. </p>\n<p>Later I've tried to use <a href=\"https://www.kaggle.com/ipythonx/efficientnet-b6-oof-weights-finder-seed-42\" target=\"_blank\">the following method</a> to find optimal weight, though few things are yet not clear to me, working on it. Do you have any suggestions for that? Basically we've kinda followed the above procedure that you've mentioned, the aim was to get <code>max(roc_auc_score(TRUE, oof))</code> after some iteration. </p>\n<p>Thank you -)</p>",
              "rawMarkdown": "Hello, @cdeotte thanks for the above pointer.  I was trying to set weight values based on the CV score rather than LB feedback with different trails and errors.  However, I totally agree with @romanweilguny that we should use ML algorithm to find an optimal weight for the ensemble. \n\nLater I've tried to use [the following method](https://www.kaggle.com/ipythonx/efficientnet-b6-oof-weights-finder-seed-42) to find optimal weight, though few things are yet not clear to me, working on it. Do you have any suggestions for that? Basically we've kinda followed the above procedure that you've mentioned, the aim was to get `max(roc_auc_score(TRUE, oof))` after some iteration. \n\nThank you -)"
            },
            {
              "id": 967258,
              "postDate": "2020-08-12T05:53:37.050Z",
              "content": "<p>I think an ML approach would be interesting.  Not sure how fast the compute is you have access to, but say you had 4 models you wanted to ensemble, you could brute force the weights with 10,000 iterations, which can totally be parallelized across cores, if you used values of .00, .10, .20, etc through .90.</p>",
              "rawMarkdown": "I think an ML approach would be interesting.  Not sure how fast the compute is you have access to, but say you had 4 models you wanted to ensemble, you could brute force the weights with 10,000 iterations, which can totally be parallelized across cores, if you used values of .00, .10, .20, etc through .90."
            }
          ]
        },
        {
          "id": 953320,
          "postDate": "2020-07-31T17:36:33.090Z",
          "content": "<p>Thx - thats what I have been looking for!</p>",
          "rawMarkdown": "Thx - thats what I have been looking for!",
          "votes": 1
        },
        {
          "id": 954193,
          "postDate": "2020-08-01T13:33:51.620Z",
          "content": "<p>I have noticed in your notebooks you select the best model in a fold based on the lowest loss.Would it be wrong to select the best model using max AUC, even if there is 0.3-0.5 difference between the lowest loss and the max AUC epoch's loss? Or is it  important to select a model based on loss</p>",
          "rawMarkdown": "I have noticed in your notebooks you select the best model in a fold based on the lowest loss.Would it be wrong to select the best model using max AUC, even if there is 0.3-0.5 difference between the lowest loss and the max AUC epoch's loss? Or is it  important to select a model based on loss"
        },
        {
          "id": 954268,
          "postDate": "2020-08-01T15:20:30.477Z",
          "content": "<p>There are basically 3 options.\n* select lowest val loss\n* select highest val auc\n* select after fixed number of epochs, like 12th each time</p>\n\n<p>There are pros and cons for each. The resultant LB AUC will be some random number added or subtracted to each Val AUC, so any of these 3 choices could produce the best LB AUC. We cannot know which is the best choice. Some people will even take all 3 and ensemble the 3 models (but that's not necessarily better either).</p>\n\n<p>In conclusion, i'm saying experiment will all, and do what you prefer. All will work well.</p>",
          "rawMarkdown": "There are basically 3 options.\n* select lowest val loss\n* select highest val auc\n* select after fixed number of epochs, like 12th each time\n\nThere are pros and cons for each. The resultant LB AUC will be some random number added or subtracted to each Val AUC, so any of these 3 choices could produce the best LB AUC. We cannot know which is the best choice. Some people will even take all 3 and ensemble the 3 models (but that's not necessarily better either).\n\nIn conclusion, i'm saying experiment will all, and do what you prefer. All will work well.",
          "votes": 1
        },
        {
          "id": 959717,
          "postDate": "2020-08-05T20:55:43.773Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Hi Chris! Do you use OOF with TTA or without TTA, selecting weights? </p>\n\n<blockquote>\n  <p>You train multiple models using the same KFold folds and save all your OOF. Then you find weights with the OOF. For example if you have 3 models, then you find w1, w2, w3 with w1+w2+w3 = 1 where you maximize AUC below</p>\n</blockquote>",
          "rawMarkdown": "@cdeotte Hi Chris! Do you use OOF with TTA or without TTA, selecting weights? \n\n&gt; You train multiple models using the same KFold folds and save all your OOF. Then you find weights with the OOF. For example if you have 3 models, then you find w1, w2, w3 with w1+w2+w3 = 1 where you maximize AUC below"
        },
        {
          "id": 967089,
          "postDate": "2020-08-11T23:12:53.580Z",
          "content": "<p>You should the same thing that was applied to your <code>submission.csv</code>. If your submission.csv has TTA applied then use OOF with TTA applied.</p>",
          "rawMarkdown": "You should the same thing that was applied to your `submission.csv`. If your submission.csv has TTA applied then use OOF with TTA applied."
        }
      ]
    },
    {
      "id": 952157,
      "postDate": "2020-07-30T17:05:00.073Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> a simple question: Do you have any ratio between image size and rest of the droprate, dropct and dropsize or you keep them same for all? Thanks for your great sources for this competition again...</p>",
      "rawMarkdown": "@cdeotte a simple question: Do you have any ratio between image size and rest of the droprate, dropct and dropsize or you keep them same for all? Thanks for your great sources for this competition again...",
      "votes": 1,
      "replies": [
        {
          "id": 952231,
          "postDate": "2020-07-30T18:12:41.463Z",
          "content": "<p>Good question. <code>DROP_SIZE</code> is a ratio between 0 and 1. So it scales your squares with your image size. (Square side to cutout out is <code>IMAGE_SIZE * DROP_SIZE</code>). Likewise <code>DROP_RATE</code> and <code>DROP_CT</code> naturally scale. Therefore once you find good numbers, you can use the same numbers for all image sizes.</p>\n\n<p>Here's an important note. The variable <code>DROP_CT</code> removes <code>CT</code> squares using a <code>for-loop</code>. I recommend using 16 or less and preferably 8 or less. If you want to remove more image, increase <code>DROP_RATE</code> and/or <code>DROP_SIZE</code>.</p>\n\n<p>If you turn this augmentation on and off, you will find that if <code>DROP_CT&amp;gt;=16</code> then it starts to slightly slow down your training time per epoch because the CPU can't augment (execute for-loop) fast enough (when doing all the other common stuff like rotation, sheer, zoom, etc etc).</p>",
          "rawMarkdown": "Good question. `DROP_SIZE` is a ratio between 0 and 1. So it scales your squares with your image size. (Square side to cutout out is `IMAGE_SIZE * DROP_SIZE`). Likewise `DROP_RATE` and `DROP_CT` naturally scale. Therefore once you find good numbers, you can use the same numbers for all image sizes.\n\nHere's an important note. The variable `DROP_CT` removes `CT` squares using a `for-loop`. I recommend using 16 or less and preferably 8 or less. If you want to remove more image, increase `DROP_RATE` and/or `DROP_SIZE`.\n\nIf you turn this augmentation on and off, you will find that if `DROP_CT&gt;=16` then it starts to slightly slow down your training time per epoch because the CPU can't augment (execute for-loop) fast enough (when doing all the other common stuff like rotation, sheer, zoom, etc etc).",
          "votes": 2
        },
        {
          "id": 952236,
          "postDate": "2020-07-30T18:18:15.123Z",
          "content": "<p>So for example if you set <code>DROP_FREQ = 1.0</code>, <code>DROP_CT = 8</code>, and <code>DROP_SIZE = 0.2</code> then regardless of image size, it probably removes about 20% of every image. (If the squares did not overlap it would be 25%).</p>",
          "rawMarkdown": "So for example if you set `DROP_FREQ = 1.0`, `DROP_CT = 8`, and `DROP_SIZE = 0.2` then regardless of image size, it probably removes about 20% of every image. (If the squares did not overlap it would be 25%).",
          "votes": 2
        },
        {
          "id": 952241,
          "postDate": "2020-07-30T18:31:41.527Z",
          "content": "<p>Thank you for the answer! I wanted to play with your approach today, implemented it on my model and some feedback here maybe you find them useful:</p>\n\n<p>In my case I didn't feel slowing down up to 16 CT's maybe 5% longer per epoch or something at high numbers, didn't go over 16 though.</p>\n\n<p>But it definitely slowed down my models learning process with increased randomness, had to increase epochs like 30% to reach dropout off levels of train auc.</p>\n\n<p>Haven't played enough to comment about cv, lb yet... But anyways thank you again! <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "rawMarkdown": "Thank you for the answer! I wanted to play with your approach today, implemented it on my model and some feedback here maybe you find them useful:\n\nIn my case I didn't feel slowing down up to 16 CT's maybe 5% longer per epoch or something at high numbers, didn't go over 16 though.\n\nBut it definitely slowed down my models learning process with increased randomness, had to increase epochs like 30% to reach dropout off levels of train auc.\n\nHaven't played enough to comment about cv, lb yet... But anyways thank you again! @cdeotte ",
          "votes": 1
        },
        {
          "id": 952284,
          "postDate": "2020-07-30T19:23:31.897Z",
          "content": "<p>Yes using image dropout or layer dropout means that you have to train for more epochs. The idea is that your model is learning slower but it will achieve a higher level of intelligence.</p>",
          "rawMarkdown": "Yes using image dropout or layer dropout means that you have to train for more epochs. The idea is that your model is learning slower but it will achieve a higher level of intelligence.",
          "votes": 1
        },
        {
          "id": 952613,
          "postDate": "2020-07-31T04:55:47.803Z",
          "content": "<p>Is there any method to speed up the for loop process. It seems to waste the incredible capabilities of TPU while doing for loop. </p>",
          "rawMarkdown": "Is there any method to speed up the for loop process. It seems to waste the incredible capabilities of TPU while doing for loop. ",
          "votes": 1
        },
        {
          "id": 952631,
          "postDate": "2020-07-31T05:17:12.977Z",
          "content": "<p>It does not waste TPU time. All augmentation of <code>tf.data.Dataset</code> happens on the TPU's CPU in parallel. So while the TPU is training one batch, the CPU is preparing the next batch. As long as <code>DROP_CT</code> is low enough it doesn't change train time at all (because the CPU batch time is less than the TPU batch time).</p>",
          "rawMarkdown": "It does not waste TPU time. All augmentation of `tf.data.Dataset` happens on the TPU's CPU in parallel. So while the TPU is training one batch, the CPU is preparing the next batch. As long as `DROP_CT` is low enough it doesn't change train time at all (because the CPU batch time is less than the TPU batch time).",
          "votes": 1
        }
      ]
    },
    {
      "id": 947100,
      "postDate": "2020-07-27T04:39:07.767Z",
      "content": "<p>Good post Chris!</p>",
      "rawMarkdown": "Good post Chris!",
      "votes": 1
    },
    {
      "id": 945624,
      "postDate": "2020-07-26T03:19:34.997Z",
      "content": "<p>In my understanding, we should not use dropout or cutout when predicting validation data or test data. Am I right?</p>",
      "rawMarkdown": "In my understanding, we should not use dropout or cutout when predicting validation data or test data. Am I right?",
      "votes": 1,
      "replies": [
        {
          "id": 945695,
          "postDate": "2020-07-26T04:53:30.620Z",
          "content": "<p>The best answer is try both and see what produces higher CV LB. Since we're using TTA in both our validation and test, I think it may be best to add dropout and/or cutout to the TTA.</p>\n\n<p>I will point out that simply removing dropout or cutout from predicting test or validation may not work. When you remove a dropout layer inside a CNN, you need to rescale the activations to adjust for the missing dropout. I'm not sure how you would \"adjust\" the validation or test predictions if we train with coarse dropout and then don't use coarse dropout for prediction.</p>",
          "rawMarkdown": "The best answer is try both and see what produces higher CV LB. Since we're using TTA in both our validation and test, I think it may be best to add dropout and/or cutout to the TTA.\n\nI will point out that simply removing dropout or cutout from predicting test or validation may not work. When you remove a dropout layer inside a CNN, you need to rescale the activations to adjust for the missing dropout. I'm not sure how you would \"adjust\" the validation or test predictions if we train with coarse dropout and then don't use coarse dropout for prediction.",
          "votes": 2
        },
        {
          "id": 946158,
          "postDate": "2020-07-26T11:59:09.063Z",
          "content": "<p>Thanks for your reply! I did a few experiments (I know it is not sufficient.) and found that as you said, removing dropout or cutout from prediction is worse. I'll report more reliable results asap! </p>",
          "rawMarkdown": "Thanks for your reply! I did a few experiments (I know it is not sufficient.) and found that as you said, removing dropout or cutout from prediction is worse. I'll report more reliable results asap! ",
          "votes": 1
        },
        {
          "id": 946176,
          "postDate": "2020-07-26T12:18:17.910Z",
          "content": "<p>Though I can understand the importance of TTA, I can't understand why using dropout works with validation or test data. My concern is that using these augmentations may lose important information. </p>",
          "rawMarkdown": "Though I can understand the importance of TTA, I can't understand why using dropout works with validation or test data. My concern is that using these augmentations may lose important information. ",
          "votes": 2
        },
        {
          "id": 946373,
          "postDate": "2020-07-26T14:41:27.497Z",
          "content": "<p>Keep in mind that the CNN is being trained with dropout. (And is smart enough to find malignant even with rectangles of images removed). If during TTA prediction there is no dropout, then the CNN could get confused.</p>\n\n<p>That's like teaching a tennis player to use a solid wood racket for months and during the tournament giving the tennis player a real racket with strings. You would think it would help, but the tennis player never learned how to use the real racket with strings.</p>",
          "rawMarkdown": "Keep in mind that the CNN is being trained with dropout. (And is smart enough to find malignant even with rectangles of images removed). If during TTA prediction there is no dropout, then the CNN could get confused.\n\nThat's like teaching a tennis player to use a solid wood racket for months and during the tournament giving the tennis player a real racket with strings. You would think it would help, but the tennis player never learned how to use the real racket with strings.",
          "votes": 3
        },
        {
          "id": 946457,
          "postDate": "2020-07-26T15:34:31.377Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> This is the first time I see someone using random cutout in TTA to be honest. I will try it, interesting idea. Usually it works well to train with cutout and don't do cutout in inference, it behaves very similar to dropout which you also dont use in inference. </p>\n\n<p>What it definitely does it that it adds randomness, but having randomness in inference can be risky. Random TTA is already something I dont feel comfortable with.</p>",
          "rawMarkdown": "@cdeotte This is the first time I see someone using random cutout in TTA to be honest. I will try it, interesting idea. Usually it works well to train with cutout and don't do cutout in inference, it behaves very similar to dropout which you also dont use in inference. \n\nWhat it definitely does it that it adds randomness, but having randomness in inference can be risky. Random TTA is already something I dont feel comfortable with."
        },
        {
          "id": 946461,
          "postDate": "2020-07-26T15:38:32.620Z",
          "content": "<p>For TTA (test time augmentation), isn't it best to use the augmentation that your network was trained with?</p>\n\n<p>If we remove image dropout from TTA, should we adjust predictions? Because if you add a dropout layer in your CNN like <code>x = Dropout(0.2)(x)</code>. Then when the network makes predictions it removes the layer dropout and it reduces the activations by 20%. (i.e. during inference Keras replaces layer dropout with <code>x = Lambda(x: 0.8*x)(x)</code>) What is the equivalent of \"reducing activations by 20%\" when using image dropout?</p>",
          "rawMarkdown": "For TTA (test time augmentation), isn't it best to use the augmentation that your network was trained with?\n\nIf we remove image dropout from TTA, should we adjust predictions? Because if you add a dropout layer in your CNN like `x = Dropout(0.2)(x)`. Then when the network makes predictions it removes the layer dropout and it reduces the activations by 20%. (i.e. during inference Keras replaces layer dropout with `x = Lambda(x: 0.8*x)(x)`) What is the equivalent of \"reducing activations by 20%\" when using image dropout?"
        },
        {
          "id": 946466,
          "postDate": "2020-07-26T15:43:55.413Z",
          "content": "<p>Doesn't the scaling happen during training and prediction is identity (is similar to what you say)? But you are right, haven't thought much about that when doing cutout. In practise I havent seen any issues with not doing cutout in inference, but also have never tried to keep it.</p>",
          "rawMarkdown": "Doesn't the scaling happen during training and prediction is identity (is similar to what you say)? But you are right, haven't thought much about that when doing cutout. In practise I havent seen any issues with not doing cutout in inference, but also have never tried to keep it."
        },
        {
          "id": 946467,
          "postDate": "2020-07-26T15:45:35.240Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> </p>\n\n<blockquote>\n  <p>Random TTA is already something I dont feel comfortable with.</p>\n</blockquote>\n\n<p>This competition takes TTA to a whole new level. People are using 25+ TTA steps. At that point, randomness gets averaged away.</p>",
          "rawMarkdown": "@philippsinger \n&gt; Random TTA is already something I dont feel comfortable with.\n\nThis competition takes TTA to a whole new level. People are using 25+ TTA steps. At that point, randomness gets averaged away."
        },
        {
          "id": 946474,
          "postDate": "2020-07-26T15:50:05.400Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Havent tested tta yet, but if tta outside of standard flips works this can be an approach indeed. Inference is fast.</p>\n\n<p>Do you just use train augs multiple times, or limit them?</p>",
          "rawMarkdown": "@cdeotte Havent tested tta yet, but if tta outside of standard flips works this can be an approach indeed. Inference is fast.\n\nDo you just use train augs multiple times, or limit them?"
        },
        {
          "id": 946481,
          "postDate": "2020-07-26T15:51:27.990Z",
          "content": "<blockquote>\n  <p>Doesn't the scaling happen during training and prediction is identity</p>\n</blockquote>\n\n<p>Oh yes, i think you're right. A layer of <code>x = dropout(0.2)(x)</code> simultaneously applies 20% dropout and increases activations by 25% internally with something like <code>x = Lambda(x: 1.25*x)(x)</code>. Then during inference it just removes the dropout layer and all is well.</p>",
          "rawMarkdown": "&gt; Doesn't the scaling happen during training and prediction is identity\n\nOh yes, i think you're right. A layer of `x = dropout(0.2)(x)` simultaneously applies 20% dropout and increases activations by 25% internally with something like `x = Lambda(x: 1.25*x)(x)`. Then during inference it just removes the dropout layer and all is well."
        },
        {
          "id": 946491,
          "postDate": "2020-07-26T15:55:48.483Z",
          "content": "<p>For TTA, i use all the training augmentations! I use rotation, sheer, scale, shift, contrast, saturation, brightness, flip, dropout, etc, etc. Using TTA 25+ is the powerhouse in this comp. It increases CV LB by something like 0.01+ or 0.02+!</p>\n\n<pre><code>    print('Predicting Test with TTA...')\n    ds_test = get_dataset(files_test,labeled=False,return_image_names=False,augment=True,\n        repeat=True,shuffle=False,dim=IMG_SIZES[fold],batch_size=BATCH_SIZE)\n    ct_test = count_data_items(files_test); STEPS = TTA * ct_test/BATCH_SIZE\n    pred = model.predict(ds_test,steps=STEPS,verbose=VERBOSE)[:TTA*ct_test,] \n    preds += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) / skf.n_splits\n</code></pre>\n\n<p>And <code>get_dataset(augment=True)</code> is the same pipeline that train uses.</p>\n\n<p>The power of TTA was first shown by AgentAuers' in his notebook <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">here</a></p>",
          "rawMarkdown": "For TTA, i use all the training augmentations! I use rotation, sheer, scale, shift, contrast, saturation, brightness, flip, dropout, etc, etc. Using TTA 25+ is the powerhouse in this comp. It increases CV LB by something like 0.01+ or 0.02+!\n\n        print('Predicting Test with TTA...')\n        ds_test = get_dataset(files_test,labeled=False,return_image_names=False,augment=True,\n            repeat=True,shuffle=False,dim=IMG_SIZES[fold],batch_size=BATCH_SIZE)\n        ct_test = count_data_items(files_test); STEPS = TTA * ct_test/BATCH_SIZE\n        pred = model.predict(ds_test,steps=STEPS,verbose=VERBOSE)[:TTA*ct_test,] \n        preds += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) / skf.n_splits\n\nAnd `get_dataset(augment=True)` is the same pipeline that train uses.\n\nThe power of TTA was first shown by AgentAuers' in his notebook [here][1]\n\n[1]: https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once",
          "votes": 1
        },
        {
          "id": 946501,
          "postDate": "2020-07-26T16:00:57.423Z",
          "content": "<p>Thank you, interesting, will try it.</p>",
          "rawMarkdown": "Thank you, interesting, will try it.",
          "votes": 1
        },
        {
          "id": 946509,
          "postDate": "2020-07-26T16:09:43.877Z",
          "content": "<p>I removed Cutout during my TTA. </p>\n\n<p>But I will try it to see if it improves. </p>",
          "rawMarkdown": "I removed Cutout during my TTA. \n\nBut I will try it to see if it improves. ",
          "votes": 1
        },
        {
          "id": 946531,
          "postDate": "2020-07-26T16:21:54.737Z",
          "content": "<p>Yes <a href=\"/serigne\">@serigne</a> , let us know. I have only recently added cutout to my augmentations. So I haven't done enough experiments myself to know whether it is best to include cutout or not during TTA inference.</p>",
          "rawMarkdown": "Yes @serigne , let us know. I have only recently added cutout to my augmentations. So I haven't done enough experiments myself to know whether it is best to include cutout or not during TTA inference.",
          "votes": 1
        },
        {
          "id": 946549,
          "postDate": "2020-07-26T16:36:20.617Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a>   Cutout definitely improved my training model.   But, as <a href=\"/philippsinger\">@philippsinger</a> ,  I was much reluctant to use Random Augmentation (specially those cutting the input images) during TTA. </p>\n\n<p>But, after <a href=\"/optimo\">@optimo</a> suggestion , I gave a try to +15 TTA  (with all my training augs except Cutout) and saw relatively huge improvement on LB. Next step is to add Cutout and let you know. </p>",
          "rawMarkdown": "@cdeotte   Cutout definitely improved my training model.   But, as @philippsinger ,  I was much reluctant to use Random Augmentation (specially those cutting the input images) during TTA. \n\nBut, after @optimo suggestion , I gave a try to +15 TTA  (with all my training augs except Cutout) and saw relatively huge improvement on LB. Next step is to add Cutout and let you know. ",
          "votes": 1
        },
        {
          "id": 946560,
          "postDate": "2020-07-26T16:41:15.190Z",
          "content": "<p><a href=\"/serigne\">@serigne</a> Does it also improve CV?</p>",
          "rawMarkdown": "@serigne Does it also improve CV?"
        },
        {
          "id": 946602,
          "postDate": "2020-07-26T17:10:28.570Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> for me it improves CV by about 0.01.</p>\n\n<p>To be honest I started with a lazy implementation, I was using coarse dropout and I left it as is in my TTA, but I was planning to remove it at some point. I agree that I would expect better results without masking part of my input during inference (I guess everyone would agree not to do mixcut during inference for example). It might also depend on the size of your dropout patches, if they are too big, you may be able to train your model but using it for inference would hurt your results. </p>\n\n<p>But I feel very comfortable with doing 20x TTA or even 100x TTA (even though I have better things to do haha), for me it would just converge to the best possible prediction for my model, it might be useless at some point but can't harm much. </p>",
          "rawMarkdown": "@philippsinger for me it improves CV by about 0.01.\n\nTo be honest I started with a lazy implementation, I was using coarse dropout and I left it as is in my TTA, but I was planning to remove it at some point. I agree that I would expect better results without masking part of my input during inference (I guess everyone would agree not to do mixcut during inference for example). It might also depend on the size of your dropout patches, if they are too big, you may be able to train your model but using it for inference would hurt your results. \n\nBut I feel very comfortable with doing 20x TTA or even 100x TTA (even though I have better things to do haha), for me it would just converge to the best possible prediction for my model, it might be useless at some point but can't harm much. ",
          "votes": 1
        },
        {
          "id": 946604,
          "postDate": "2020-07-26T17:12:03.947Z",
          "content": "<p>Oups I was eager to try it on test data but I didn't even check on validation data ( I used to try Random Augs only on training data so far)</p>\n\n<p>But you're right. These augmentations need to be checked on validation data first.  I will try to do it more systematically from now. </p>",
          "rawMarkdown": "Oups I was eager to try it on test data but I didn't even check on validation data ( I used to try Random Augs only on training data so far)\n\nBut you're right. These augmentations need to be checked on validation data first.  I will try to do it more systematically from now. ",
          "votes": 2
        },
        {
          "id": 947551,
          "postDate": "2020-07-27T10:40:42.523Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> You talk about using TTA in both your validation as well as test.  I have used TTA in test, but not tried it in validation and will give it a try.  So do you end up doing the same augmentations you do in train, on both validation and test?  Or do you do some on one and not on others?</p>",
          "rawMarkdown": "@cdeotte You talk about using TTA in both your validation as well as test.  I have used TTA in test, but not tried it in validation and will give it a try.  So do you end up doing the same augmentations you do in train, on both validation and test?  Or do you do some on one and not on others?"
        },
        {
          "id": 947753,
          "postDate": "2020-07-27T13:14:11.913Z",
          "content": "<p>I use TTA in both test and validation. That way I can choose the TTA parameters that maximize validation score. Currently, I use all augmentation techniques for train, validation, and test. (I have not experimented with using less augmentation on validation and test compared with train. Perhaps using less would increase CV score more).</p>",
          "rawMarkdown": "I use TTA in both test and validation. That way I can choose the TTA parameters that maximize validation score. Currently, I use all augmentation techniques for train, validation, and test. (I have not experimented with using less augmentation on validation and test compared with train. Perhaps using less would increase CV score more)."
        },
        {
          "id": 950170,
          "postDate": "2020-07-29T08:19:29.713Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Cutout on TTA improved further my LB score . </p>\n\n<p>I have run an experiment for fold0 on 256x256. <br>\n10 TTA random Augment w/o Cutout : LB 0.9357\n10 TTA random Augment w/ Cutout : LB 0.9392</p>\n\n<p>However I haven't check yet any random augment on validation data. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F974295%2F5b02ef384e321357eabe730a8e38a4ce%2FCapre.PNG?generation=1596010735794067&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@cdeotte Cutout on TTA improved further my LB score . \n\n\nI have run an experiment for fold0 on 256x256.   \n10 TTA random Augment w/o Cutout : LB 0.9357\n10 TTA random Augment w/ Cutout : LB 0.9392\n\nHowever I haven't check yet any random augment on validation data. \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F974295%2F5b02ef384e321357eabe730a8e38a4ce%2FCapre.PNG?generation=1596010735794067&amp;alt=media)\n",
          "votes": 5
        },
        {
          "id": 970107,
          "postDate": "2020-08-14T07:58:54.750Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks (you replied 18 days ago but just thinking again about this).  In your Triple Stratified notebook, if I am reading it right (I run PyTorch so I have not run that notebook), your validation dataset looks like it has <code>augment=False</code> or something along the lines of that, which is why I thought maybe in your approach you didn't do it in validation.  But I am possibly just reading your notebook wrong.</p>",
          "rawMarkdown": "@cdeotte Thanks (you replied 18 days ago but just thinking again about this).  In your Triple Stratified notebook, if I am reading it right (I run PyTorch so I have not run that notebook), your validation dataset looks like it has `augment=False` or something along the lines of that, which is why I thought maybe in your approach you didn't do it in validation.  But I am possibly just reading your notebook wrong."
        }
      ]
    },
    {
      "id": 944899,
      "postDate": "2020-07-25T12:40:13.797Z",
      "content": "<p>Can we also extend it for images with 4 channels? And also why compiler needs to know <code>image = tf.reshape(image,[DIM,DIM,3])</code>? </p>",
      "rawMarkdown": "Can we also extend it for images with 4 channels? And also why compiler needs to know `image = tf.reshape(image,[DIM,DIM,3])`? ",
      "votes": 1,
      "replies": [
        {
          "id": 944959,
          "postDate": "2020-07-25T13:28:43.517Z",
          "content": "<p>TensorFlow needs to determine the output shape of each function that operates on <code>tf.data.Dataset()</code>. If you remove <code>image = tf.reshape(image,[DIM,DIM,3])</code>, then i think TensorFlow complains that it can't determine the shape. (I'm not sure why it can't figure it out).</p>\n\n<p>If you have images with 4 channels, just change the 3 to a 4 in two places\n* <code>two = tf.zeros([yb-ya,xb-xa,3])</code>\n*  <code>image = tf.reshape(image,[DIM,DIM,3])</code></p>",
          "rawMarkdown": "TensorFlow needs to determine the output shape of each function that operates on `tf.data.Dataset()`. If you remove `image = tf.reshape(image,[DIM,DIM,3])`, then i think TensorFlow complains that it can't determine the shape. (I'm not sure why it can't figure it out).\n\nIf you have images with 4 channels, just change the 3 to a 4 in two places\n* `two = tf.zeros([yb-ya,xb-xa,3])`\n*  `image = tf.reshape(image,[DIM,DIM,3])`",
          "votes": 1
        },
        {
          "id": 944961,
          "postDate": "2020-07-25T13:30:05.877Z",
          "content": "<p>Or replace both 3 with <code>image.shape[-1]</code>. That should work and then it can handle whatever number of channels comes its way.</p>",
          "rawMarkdown": "Or replace both 3 with `image.shape[-1]`. That should work and then it can handle whatever number of channels comes its way."
        },
        {
          "id": 944972,
          "postDate": "2020-07-25T13:33:33.503Z",
          "content": "<p>Thanks. I'm bothering you too much these days I guess...</p>",
          "rawMarkdown": "Thanks. I'm bothering you too much these days I guess..."
        }
      ]
    },
    {
      "id": 944488,
      "postDate": "2020-07-25T06:42:59.263Z",
      "content": "<p>Can I use this function for JPEG image augmentation ?</p>",
      "rawMarkdown": "Can I use this function for JPEG image augmentation ?",
      "votes": 1,
      "replies": [
        {
          "id": 944575,
          "postDate": "2020-07-25T07:36:13.063Z",
          "content": "<p>For JPEG, you can use Albumentations' <code>albumentations.augmentations.transforms.Cutout</code> and <code>albumentations.augmentations.transforms.CoarseDropout</code>, the API is <a href=\"https://albumentations.readthedocs.io/en/latest/api/augmentations.html#module-albumentations.augmentations.transforms\">here</a></p>",
          "rawMarkdown": "For JPEG, you can use Albumentations' `albumentations.augmentations.transforms.Cutout` and `albumentations.augmentations.transforms.CoarseDropout`, the API is [here][1]\n\n[1]: https://albumentations.readthedocs.io/en/latest/api/augmentations.html#module-albumentations.augmentations.transforms",
          "votes": 1
        },
        {
          "id": 944587,
          "postDate": "2020-07-25T07:44:52.133Z",
          "content": "<p>Thank you so much.This is very helpful.</p>",
          "rawMarkdown": "Thank you so much.This is very helpful.",
          "votes": 1
        }
      ]
    },
    {
      "id": 944404,
      "postDate": "2020-07-25T04:54:44.833Z",
      "content": "<p>Nice work! </p>",
      "rawMarkdown": "Nice work! ",
      "votes": 1
    },
    {
      "id": 944383,
      "postDate": "2020-07-25T04:26:19.457Z",
      "content": "<p>Nice and precise implementations <a href=\"/cdeotte\">@cdeotte</a>! I was curious if <code>AutoAugment</code> can also be implemented in TF as it is proved to improve the model's performance. I found the pytorch implementation of <code>AutoAugmentation</code> <a href=\"https://github.com/DeepVoltaire/AutoAugment\">here</a>. </p>",
      "rawMarkdown": "Nice and precise implementations @cdeotte! I was curious if `AutoAugment` can also be implemented in TF as it is proved to improve the model's performance. I found the pytorch implementation of `AutoAugmentation` [here](https://github.com/DeepVoltaire/AutoAugment). ",
      "votes": 1,
      "replies": [
        {
          "id": 944577,
          "postDate": "2020-07-25T07:37:02.843Z",
          "content": "<p>Sure. Anything can be coded in TensorFlow. This would be interesting to try out.</p>",
          "rawMarkdown": "Sure. Anything can be coded in TensorFlow. This would be interesting to try out."
        }
      ]
    },
    {
      "id": 944232,
      "postDate": "2020-07-25T00:15:18.477Z",
      "content": "<p>Why makes it suitable for TPU and GPU? </p>",
      "rawMarkdown": "Why makes it suitable for TPU and GPU? ",
      "votes": 1,
      "replies": [
        {
          "id": 944236,
          "postDate": "2020-07-25T00:20:18.633Z",
          "content": "<p>This code can be executed directly on the GPU/TPU if you desire.</p>\n\n<p>TensorFlow runs their <code>tf.data.Dataset</code> pipeline on CPU in parallel with GPU/TPU training but once you have TensorFlow code, you can push it into a <code>tf.keras.layer</code> and make it part of your model. Then it will execute on the GPU/TPU inside your model. Sometimes this is faster than CPU parallel sometimes it is not. </p>\n\n<p>However, i don't suggest pushing it into a <code>tf.keras.layer</code>. I suggest just using it with <code>tf.data.Dataset</code> if you're using TensorFlow GPU/TPU.</p>",
          "rawMarkdown": "This code can be executed directly on the GPU/TPU if you desire.\n\nTensorFlow runs their `tf.data.Dataset` pipeline on CPU in parallel with GPU/TPU training but once you have TensorFlow code, you can push it into a `tf.keras.layer` and make it part of your model. Then it will execute on the GPU/TPU inside your model. Sometimes this is faster than CPU parallel sometimes it is not. \n\nHowever, i don't suggest pushing it into a `tf.keras.layer`. I suggest just using it with `tf.data.Dataset` if you're using TensorFlow GPU/TPU.",
          "votes": 3
        },
        {
          "id": 944317,
          "postDate": "2020-07-25T02:41:14.440Z",
          "content": "<p>It seems that we can't use albumentations in tf.data.Dataet of if we can somehow convert all the operations in tensorflow then they'll work fine?</p>",
          "rawMarkdown": "It seems that we can't use albumentations in tf.data.Dataet of if we can somehow convert all the operations in tensorflow then they'll work fine?",
          "votes": 1
        },
        {
          "id": 944331,
          "postDate": "2020-07-25T03:01:14.850Z",
          "content": "<p>Yeah, unfortunately we cannot use Albumentations with <code>tf.data.Dataset</code> nor use <code>tf.contrib</code> libraries nor any libraries. All augmentation needs to be written from scratch using TensorFlow language. </p>\n\n<p>I think we now have all the code to do the main augmentations with <code>tf.data.Dataset</code>. Rotation, Sheer, Zoom, Shift <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a>, Cutmix and mixup <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">here</a>, and now Cutout and sparse dropout in this discussion. And TensorFlow already has functions for color manipulations such as contrast, brightness, hue, sharpness, etc.</p>",
          "rawMarkdown": "Yeah, unfortunately we cannot use Albumentations with `tf.data.Dataset` nor use `tf.contrib` libraries nor any libraries. All augmentation needs to be written from scratch using TensorFlow language. \n\nI think we now have all the code to do the main augmentations with `tf.data.Dataset`. Rotation, Sheer, Zoom, Shift [here][1], Cutmix and mixup [here][2], and now Cutout and sparse dropout in this discussion. And TensorFlow already has functions for color manipulations such as contrast, brightness, hue, sharpness, etc.\n\n[1]: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\n[2]: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu",
          "votes": 2
        }
      ]
    },
    {
      "id": 944227,
      "postDate": "2020-07-25T00:04:22.700Z",
      "content": "<p>Thanks again for this Chris. I've been resisting the temptation to reduce the number and types of augmentations. I know that in the long run, I'll be better off. As you mentioned it's not as simple as one would hope to do custom augmentations in TF. I'm keen to try this out. </p>",
      "rawMarkdown": "Thanks again for this Chris. I've been resisting the temptation to reduce the number and types of augmentations. I know that in the long run, I'll be better off. As you mentioned it's not as simple as one would hope to do custom augmentations in TF. I'm keen to try this out. ",
      "votes": 1
    },
    {
      "id": 947848,
      "postDate": "2020-07-27T14:10:36.407Z",
      "content": "<p>Technically, you can use Albumentations through the tf dataset pipeline using tf.numpy_function. The downside is that TPU's do not support this feature (yet). On a GPU however, this should work.</p>",
      "rawMarkdown": "Technically, you can use Albumentations through the tf dataset pipeline using tf.numpy_function. The downside is that TPU's do not support this feature (yet). On a GPU however, this should work.",
      "votes": 2,
      "replies": [
        {
          "id": 947855,
          "postDate": "2020-07-27T14:15:53.297Z",
          "content": "<p>Thanks for letting me know that.</p>",
          "rawMarkdown": "Thanks for letting me know that.",
          "votes": 1
        }
      ]
    },
    {
      "id": 945900,
      "postDate": "2020-07-26T07:49:01.987Z",
      "content": "<p>You don't seem to be using hair_aug(), do you think we should use hair augmentation?</p>",
      "rawMarkdown": "You don't seem to be using hair_aug(), do you think we should use hair augmentation?",
      "votes": 2,
      "replies": [
        {
          "id": 946375,
          "postDate": "2020-07-26T14:42:52.723Z",
          "content": "<p>I have not tried hair augmentation yet. I would guess that it is helpful. Has anyone posted a TensorFlow <code>tf.data.Dataset</code> implementation? Please post the link here.</p>",
          "rawMarkdown": "I have not tried hair augmentation yet. I would guess that it is helpful. Has anyone posted a TensorFlow `tf.data.Dataset` implementation? Please post the link here.",
          "votes": 1
        },
        {
          "id": 946806,
          "postDate": "2020-07-26T21:18:55.550Z",
          "content": "<p>Here is my notebook implementing Roman's advance hair augmentation in TF:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-data-augmentation-in-tf-hair-batch-affine\">SIIM: Data Augmentation in TF: Hair + Batch Affine</a></p>\n\n<p>I have not converted it to the batch form, so it does slow down training. </p>",
          "rawMarkdown": "Here is my notebook implementing Roman's advance hair augmentation in TF:\n\n[SIIM: Data Augmentation in TF: Hair + Batch Affine](https://www.kaggle.com/graf10a/siim-data-augmentation-in-tf-hair-batch-affine)\n\nI have not converted it to the batch form, so it does slow down training. ",
          "votes": 4
        },
        {
          "id": 946943,
          "postDate": "2020-07-27T01:00:18.797Z",
          "content": "<p>Great thanks <a href=\"/graf10a\">@graf10a</a> , i will play with it.</p>",
          "rawMarkdown": "Great thanks @graf10a , i will play with it.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1182038,
      "postDate": "2021-02-02T09:15:43.760Z",
      "content": "<p>This is really very useful</p>",
      "rawMarkdown": "This is really very useful"
    },
    {
      "id": 1182039,
      "postDate": "2021-02-02T09:15:43.760Z",
      "content": "<p>This is really very useful</p>",
      "rawMarkdown": "This is really very useful"
    },
    {
      "id": 972499,
      "postDate": "2020-08-16T15:31:32.547Z",
      "content": "<p>is it same the cutmix ?</p>\n<p>def cutmix_aug(image, label=None, DIM=256, BATCH_SIZE=4, PROBABILITY=0.4):<br>\n    # input image - is a batch of images of size [n,dim,dim,3] not a single image of [dim,dim,3]<br>\n    # output - a batch of images with cutmix applied</p>\n<pre><code>imgs = []; labs = []\n\nfor j in range(BATCH_SIZE):\n\n    #random_uniform( shape, minval=0, maxval=None)        \n    # DO CUTMIX WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast(tf.random.uniform([], 0, 1) &lt;= PROBABILITY, tf.int32)\n\n    # CHOOSE RANDOM IMAGE TO CUTMIX WITH\n    k = tf.cast(tf.random.uniform([], 0, BATCH_SIZE), tf.int32)\n\n    # CHOOSE RANDOM LOCATION\n    x = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n    y = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n\n    # Beta(1, 1)\n    b = tf.random.uniform([], 0, 1) # this is beta dist with alpha=1.0\n\n\n    WIDTH = tf.cast(DIM * tf.math.sqrt(1-b),tf.int32) * P\n    ya = tf.math.maximum(0,y-WIDTH//2)\n    yb = tf.math.minimum(DIM,y+WIDTH//2)\n    xa = tf.math.maximum(0,x-WIDTH//2)\n    xb = tf.math.minimum(DIM,x+WIDTH//2)\n\n    # MAKE CUTMIX IMAGE\n    one = image[j,ya:yb,0:xa,:]\n    two = image[k,ya:yb,xa:xb,:]\n    three = image[j,ya:yb,xb:DIM,:]        \n    #ya:yb\n    middle = tf.concat([one,two,three],axis=1)\n\n    img = tf.concat([image[j,0:ya,:,:],middle,image[j,yb:DIM,:,:]],axis=0)\n    imgs.append(img)\n\n    # MAKE CUTMIX LABEL\n    a = tf.cast(WIDTH*WIDTH/DIM/DIM,tf.float32)\n    lab1 = label[j,]\n    lab2 = label[k,]\n    labs.append((1-a)*lab1 + a*lab2)\n\nimage2 = tf.reshape(tf.stack(imgs),(BATCH_SIZE, DIM, DIM, 3))\nlabel2 = tf.reshape(tf.stack(labs),(BATCH_SIZE, 1))\nreturn image2, label2\n</code></pre>\n<p>this script handle the label too, cast label to float32</p>",
      "rawMarkdown": "is it same the cutmix ?\n\ndef cutmix_aug(image, label=None, DIM=256, BATCH_SIZE=4, PROBABILITY=0.4):\n    # input image - is a batch of images of size [n,dim,dim,3] not a single image of [dim,dim,3]\n    # output - a batch of images with cutmix applied\n\n    imgs = []; labs = []\n    \n    for j in range(BATCH_SIZE):\n        \n        #random_uniform( shape, minval=0, maxval=None)        \n        # DO CUTMIX WITH PROBABILITY DEFINED ABOVE\n        P = tf.cast(tf.random.uniform([], 0, 1) <= PROBABILITY, tf.int32)\n        \n        # CHOOSE RANDOM IMAGE TO CUTMIX WITH\n        k = tf.cast(tf.random.uniform([], 0, BATCH_SIZE), tf.int32)\n        \n        # CHOOSE RANDOM LOCATION\n        x = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n        y = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n        \n        # Beta(1, 1)\n        b = tf.random.uniform([], 0, 1) # this is beta dist with alpha=1.0\n        \n\n        WIDTH = tf.cast(DIM * tf.math.sqrt(1-b),tf.int32) * P\n        ya = tf.math.maximum(0,y-WIDTH//2)\n        yb = tf.math.minimum(DIM,y+WIDTH//2)\n        xa = tf.math.maximum(0,x-WIDTH//2)\n        xb = tf.math.minimum(DIM,x+WIDTH//2)\n        \n        # MAKE CUTMIX IMAGE\n        one = image[j,ya:yb,0:xa,:]\n        two = image[k,ya:yb,xa:xb,:]\n        three = image[j,ya:yb,xb:DIM,:]        \n        #ya:yb\n        middle = tf.concat([one,two,three],axis=1)\n\n        img = tf.concat([image[j,0:ya,:,:],middle,image[j,yb:DIM,:,:]],axis=0)\n        imgs.append(img)\n        \n        # MAKE CUTMIX LABEL\n        a = tf.cast(WIDTH*WIDTH/DIM/DIM,tf.float32)\n        lab1 = label[j,]\n        lab2 = label[k,]\n        labs.append((1-a)*lab1 + a*lab2)\n\n    image2 = tf.reshape(tf.stack(imgs),(BATCH_SIZE, DIM, DIM, 3))\n    label2 = tf.reshape(tf.stack(labs),(BATCH_SIZE, 1))\n    return image2, label2\n\nthis script handle the label too, cast label to float32",
      "replies": [
        {
          "id": 972924,
          "postDate": "2020-08-17T01:25:30.137Z",
          "content": "<p>I published a CutMix notebook <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\" target=\"_blank\">here</a>. For coarse dropout, I used that CutMix notebook and just replace the second image with all zeros (to perform the dropout).</p>",
          "rawMarkdown": "I published a CutMix notebook [here][1]. For coarse dropout, I used that CutMix notebook and just replace the second image with all zeros (to perform the dropout).\n\n[1]: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu"
        },
        {
          "id": 972971,
          "postDate": "2020-08-17T02:22:35.937Z",
          "content": "<p>ah ok, so label still keep 👍</p>",
          "rawMarkdown": "ah ok, so label still keep 👍",
          "votes": 1
        }
      ]
    },
    {
      "id": 949312,
      "postDate": "2020-07-28T14:51:35.453Z",
      "content": "<p>This method is cool .\nBut I have a doubt\nas the dataset  is imbalanced , does by adding small black boxes to the image increases its performance because it is learning from the presence of small black boxes to predict the image as non malignant and as the number of non malignant class are more it may cause a delusion that the model is working better with the dropout  where as it is learning wrong thing which will affect its performance on the test data </p>",
      "rawMarkdown": "This method is cool .\nBut I have a doubt\nas the dataset  is imbalanced , does by adding small black boxes to the image increases its performance because it is learning from the presence of small black boxes to predict the image as non malignant and as the number of non malignant class are more it may cause a delusion that the model is working better with the dropout  where as it is learning wrong thing which will affect its performance on the test data \n\n",
      "replies": [
        {
          "id": 949342,
          "postDate": "2020-07-28T15:13:19.070Z",
          "content": "<p>As long as the black boxes are added to both benign and malignant images, all is good.</p>",
          "rawMarkdown": "As long as the black boxes are added to both benign and malignant images, all is good.",
          "votes": 2
        },
        {
          "id": 949400,
          "postDate": "2020-07-28T16:04:50.390Z",
          "content": "<p>Thank you for your time and response .\nI also wanted to know for this competition   up sampling the malignant data is better than using weights or do you advise on using both at the same time </p>",
          "rawMarkdown": "Thank you for your time and response .\nI also wanted to know for this competition   up sampling the malignant data is better than using weights or do you advise on using both at the same time "
        },
        {
          "id": 949433,
          "postDate": "2020-07-28T16:25:01.307Z",
          "content": "<p>Upsampling or using weights are very similar. For example you can double all the malignant images or you can use weights <code>{0:1,  1:2}</code> (malignant is 2x). One advantage of weights is less train data and faster training epochs. One advantage of upsample is that each batch now contains some malignant and the malignant learning process is more incremental.</p>\n\n<p>Additionally since we're using data augmentation and external data malignant images, then doing upsample adds more variety to the malignant class and hopefully encourages more malignant generalization.</p>",
          "rawMarkdown": "Upsampling or using weights are very similar. For example you can double all the malignant images or you can use weights `{0:1,  1:2}` (malignant is 2x). One advantage of weights is less train data and faster training epochs. One advantage of upsample is that each batch now contains some malignant and the malignant learning process is more incremental.\n\nAdditionally since we're using data augmentation and external data malignant images, then doing upsample adds more variety to the malignant class and hopefully encourages more malignant generalization."
        }
      ]
    },
    {
      "id": 945544,
      "postDate": "2020-07-26T00:02:59.827Z",
      "content": "<p>I tried Dropout in another Dataset and I got this error while training...\n<code>\nCompilation failure: Dynamic Spatial Convolution is not supported\n</code></p>",
      "rawMarkdown": "I tried Dropout in another Dataset and I got this error while training...\n```\nCompilation failure: Dynamic Spatial Convolution is not supported\n```",
      "replies": [
        {
          "id": 945550,
          "postDate": "2020-07-26T00:16:04.553Z",
          "content": "<p>hmm... i've seen that error before. It relates to TensorFlow knowing the size of stuff. Can your post your code that adds the dropout to your <code>tf.data.Dataset</code>? Make sure you give <code>dropout</code> the correct image size. And make sure that you add dropout <strong>before</strong> you batch, i.e. <code>ds = ds.map(dropout)</code> <strong>before</strong> <code>ds = ds.batch(BATCH_SIZE)</code></p>\n\n<pre><code>dropout(image, DIM=???, PROBABILITY = 0.75, CT = 8, SZ = 0.2)\n</code></pre>\n\n<p>replace ??? with correct size</p>",
          "rawMarkdown": "hmm... i've seen that error before. It relates to TensorFlow knowing the size of stuff. Can your post your code that adds the dropout to your `tf.data.Dataset`? Make sure you give `dropout` the correct image size. And make sure that you add dropout **before** you batch, i.e. `ds = ds.map(dropout)` **before** `ds = ds.batch(BATCH_SIZE)`\n\n    dropout(image, DIM=???, PROBABILITY = 0.75, CT = 8, SZ = 0.2)\n\nreplace ??? with correct size"
        },
        {
          "id": 946183,
          "postDate": "2020-07-26T12:23:43.453Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 946188,
          "postDate": "2020-07-26T12:26:32.670Z",
          "content": "<p>Same  code runs for GPU and CPU but shows error in TPU..\n```\ndef dropout_img(image,DIM=256, PROBABILITY = 0.5, CT = 50, SZ = 0.06):\n    # input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed</p>\n\n<pre><code># DO DROPOUT WITH PROBABILITY DEFINED ABOVE\nP = tf.cast( tf.random.uniform([],0,1)&lt;PROBABILITY, tf.int32)\nif (P==0)|(CT==0)|(SZ==0): \n    return image, mask\n\nfor k in range(CT):\n    # CHOOSE RANDOM LOCATION\n    x = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n    y = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n    # COMPUTE SQUARE \n    WIDTH = tf.cast( SZ*DIM,tf.int32) * P\n    ya = tf.math.maximum(0,y-WIDTH//2)\n    yb = tf.math.minimum(DIM,y+WIDTH//2)\n    xa = tf.math.maximum(0,x-WIDTH//2)\n    xb = tf.math.minimum(DIM,x+WIDTH//2)\n    # DROPOUT IMAGE\n    one = image[ya:yb,0:xa,:]\n    two = tf.zeros([yb-ya,xb-xa,4], dtype = image.dtype) \n    three = image[ya:yb,xb:DIM,:]\n    middle = tf.concat([one,two,three],axis=1)\n    image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM,:,:]],axis=0)\n\nimage = tf.reshape(image,[DIM,DIM,-1])\n\nreturn image\n</code></pre>\n\n<p><code>\n</code>\ndef augmentation_img(img, dim=256):   </p>\n\n<pre><code>seed = np.int(np.round(np.random.uniform()))\n\nimg = tf.image.random_flip_left_right(img, seed = seed)\n\nimg = transform_img(img, DIM =dim)\nimg = dropout_img(img, DIM=dim, PROBABILITY = 0.5, CT = 50, SZ = 0.06)\n\n\nimg = tf.image.random_contrast(img, 0.8, 1.2)\nimg = tf.image.random_brightness(img, 0.1)\n\nimg = tf.reshape(img, [dim, dim, -1])\n\nreturn img\n</code></pre>\n\n<p><code>\n</code>\n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO)\n    ds = ds.cache()</p>\n\n<pre><code>if repeat:\n    ds = ds.repeat()\n\nif shuffle: \n    ds = ds.shuffle(1024*8)\n    opt = tf.data.Options()\n    opt.experimental_deterministic = False\n    ds = ds.with_options(opt)\n\n ds = ds.map(lambda example: read_unlabeled_tfrecord(example,          return_image_names),num_parallel_calls=AUTO) \n\n    if augment:\n        ds = ds.map(lambda img, name: (augmentation_img(img, dim=dim), name), num_parallel_calls=AUTO)\n\nds = ds.batch(batch_size * REPLICAS)\nds = ds.prefetch(AUTO)\nreturn ds\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "Same  code runs for GPU and CPU but shows error in TPU..\n```\ndef dropout_img(image,DIM=256, PROBABILITY = 0.5, CT = 50, SZ = 0.06):\n    # input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed\n    \n    # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast( tf.random.uniform([],0,1)"
        }
      ]
    },
    {
      "id": 967894,
      "postDate": "2020-08-12T15:10:31.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 953250,
      "postDate": "2020-07-31T16:14:52.437Z",
      "content": "<p>Thank you !!</p>",
      "rawMarkdown": "Thank you !!",
      "votes": 1
    },
    {
      "id": 951260,
      "postDate": "2020-07-30T02:49:07.167Z",
      "content": "<p>Thanks for Sharing!</p>",
      "rawMarkdown": "Thanks for Sharing!",
      "votes": 1
    },
    {
      "id": 948203,
      "postDate": "2020-07-27T18:14:56.500Z",
      "content": "<p>Nice work, thanks for sharing!</p>",
      "rawMarkdown": "Nice work, thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 953624,
      "author_name": "Stephan",
      "author_url": "",
      "post_date": "2020-07-31T23:34:25.127000",
      "content": "<p>I've created a version that accepts batches of images, for anyone who is interested. I've tested it on a GPU setup, so I can't be sure this works on a TPU yet.</p>\n\n<p>edit: I see that the prob arg isn't actually utilized yet, I'll fix this tommorow.\nedit 2: probabilities now included :)\nedit 3: replaced a type in tf.scatter_nd where the image sizes were fixed to 128,128 -&gt; now adjusted for variable sizes</p>\n\n<p>```python\ndef coarse_dropout(images, square_size, num_squares, prob):\n    ''' Coarse image dropout \n    Coarse dropout masks are created by using tf.scatter_nd\n    Which can update matrices according to coordinates specified\n    by the user. </p>\n\n<pre><code>Parameters\n----------\nimages : tf Tensor\n    A batch of images of size [BS, H, W, C] \nsquare_size : float\n    A float between [0,1], a percentage of the image size\nnum_squares : int &amp;gt; 0\n    The number of dropout squares per single image, must be &amp;gt;0\nprob : float\n    A float between [0,1], probability that dropout is used for this batch\n\nYields\n------\nTYPE\n    DESCRIPTION.\n\n'''\n   # random dropout\nprob = tf.cast( tf.random.uniform([],0,1) &amp;lt; prob, tf.int32)\nif (prob == 0): return images\n\nimg_shape = images.shape\n_, h, w, c = img_shape[0], img_shape[1], img_shape[2], img_shape[3]\n\n# For some odd reason, the batch size is lost in the processing pipeline\n# seriously, how can it get lost ._., it literally says ds.batch()...\n# tensorflow shenanigans...\nbs = tf.cast(tf.reduce_sum(tf.ones_like(images)) / (images.shape[1] * images.shape[2] * images.shape[3]), tf.int32)\n\n\n# size of square in pixels (ssp), ASSUMING SQUARE IMAGES!\nssp = tf.cast(tf.math.ceil(h * square_size), tf.int32)\n\n# Create random start x coordinates\ncoords_x = tf.random.uniform([bs, num_squares], \n                             minval=0, \n                             maxval= (h - ssp), \n                             dtype=tf.int32) \n\n# Create ranges from start to the end of the line\n# [ssp, bs, num_squares]\ncoords_x = tf.linspace(coords_x, coords_x + ssp - 1, ssp)\n\n# [num_squares, bs, ssp]\ncoords_x = tf.cast(tf.transpose(coords_x), tf.int32)\n\n# Create random start y coordinates\ncoords_y = tf.random.uniform([bs, num_squares], \n                             minval=0, \n                             maxval= (h - ssp), \n                             dtype=tf.int32) \n\n# Create ranges from start to the end of the line\n# [ssp, bs, num_squares]\ncoords_y = tf.linspace(coords_y, coords_y + ssp - 1, ssp)\n\n# [num_squares, bs, ssp]\ncoords_y = tf.cast(tf.transpose(coords_y), tf.int32)\n\n\n# Create coordinate range combinations \n# and reshape to [bs, num_squares, 1, ssp * ssp]\ngrid_y = tf.reshape(tf.tile(coords_y, [1,1,ssp]), \n                    (bs, num_squares, 1, ssp * ssp))\n\n# and reshape to [bs, num_squares, ssp, ssp], transpose the inner matrices\ngrid_y = tf.transpose(tf.reshape(grid_y, \n                                 (bs, num_squares, ssp, ssp)), \n                      (0, 1, 3, 2))\n\n# Repeat for x coordinates\ngrid_x = tf.reshape(tf.tile(coords_x, [1,1,ssp]), \n                    (bs, num_squares, 1, ssp * ssp))\n\ngrid_x = tf.reshape(grid_x, (bs, num_squares, ssp, ssp))\n\n# Stack the grids into a single matrix\n# grid is [2, bs, num_squares, ssp, ssp]\ngrid = tf.stack([grid_y, grid_x], axis=0)\n\n# Transpose and reshape [ bs, ssp * ssp * num_squares, 2] \n# Creates an array of 2D coordinates ([[x1,y1], [x2, y2], ..., [xn, yn]])\n# over all squares (num_squares), for each combination (ssp*ssp) of coordinates \n# and each batch (bs)\ngrid = tf.reshape(tf.transpose(grid, (1, 4, 3, 2, 0)), \n                  (bs, ssp * ssp * num_squares, 2))\n\n# [bs, sz*sz*num_squares, 2]\n#grid = tf.reshape(grid, (bs, ssp*ssp*num_squares, 2))\n\n# create batch indices [0,..., bs] and reshape to [ssp * ssp * num_squares, bs]\nbatch_indices = tf.reshape(tf.tile(tf.range(0, bs), \n                                   [ssp * ssp * num_squares]), \n                           (ssp * ssp * num_squares, bs))\n\n# Transpose to get the right order, and reshape to match grid shape\nbatch_indices = tf.reshape(tf.transpose(batch_indices), \n                           (bs, ssp * ssp * num_squares, 1))\n\n# concatenate batch indices with the grid\n# these yield 3D coordinates like e.g.\n# [[bs0, x1, y1], [bs0, x2, y1], ..., [bsn, xn, yn]]\ngrid = tf.concat([batch_indices, grid], axis=2)\n\n# create a matrix of zeros, and update the matrix with the grid indices\n# this essentially creates a mask with coarse dropout squares\nmasks = tf.scatter_nd(grid[tf.newaxis,...], \n                        tf.ones([1,bs,ssp*ssp*num_squares]) * -1, \n                        shape=(bs, h, w)) +1\n\n# Due to overlap of squares, some get coordinates get updated twice\n# and result in values &amp;lt; -1, clip these values\nmasks = tf.clip_by_value(masks, 0, 1)\n\nreturn images * masks[..., tf.newaxis] \n</code></pre>\n\n<p>```</p>",
      "votes": 7,
      "replies": [
        {
          "id": 954270,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T15:24:34.653000",
          "content": "<p>Thanks <a href=\"/stephanovich\">@stephanovich</a> ! when i get time i will play with this. I'm curious to benchmark the speed and see if this is faster than \"per image\".</p>\n\n<p>I've noticed when using \"per image\", if we cutout more than 16 squares it will begin to slow down training, perhaps your \"per batch\" code can do more than 16 and not slow down training.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 955084,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-08-02T10:17:17.990000",
          "content": "<p>Edit: you might want to try the augmentation from this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171778\">discussion</a> and <a href=\"https://www.kaggle.com/benboren/tfrecord-progressive-sprinkles\">notebook</a>, it seems to be a more generalized version of coarse dropout, and does not use for loops either (i.e. vectorized).</p>\n\n<p>So I ran a small benchmark myself on a cpu with 32 threads, here's the results:\nThe experiments are run on 128*128 images with a variable number of squares each and a square size of 0.1 * the size of the image. The probability is set to one in both cases. I parsed through 60000 examples, which in my opinion was already enough to at least see a difference.\nI've included 3 runs: \n1. one without any sort of augmentation, which acts as the lower bound (or golden standard) for time benchmarking (it can never go faster than this).\n<code>python\nCPU times: user 11min 27s, sys: 30.9 s, total: 11min 58s\nWall time: 29.2 s\n</code></p>\n\n<p>2 . The new batch version. This version first prepares the image, then batches the images, and then augments them. Otherwise, we would not see any improvements.\n```  python</p>\n\n<h1>Using 8 squares</h1>\n\n<p>CPU times: user 28min 6s, sys: 1min 38s, total: 29min 44s\nWall time: 1min 12s</p>\n\n<h1>Using 16 squares</h1>\n\n<p>CPU times: user 33min, sys: 1min 36s, total: 34min 36s\nWall time: 1min 17s</p>\n\n<p>```</p>\n\n<p>3 . Your version. In the processing pipeline, the images are first prepared through prepare_image, then are augmented, and finally batched.\n```python</p>\n\n<h1>Using 8 squares</h1>\n\n<p>CPU times: user 50min 36s, sys: 2min 37s, total: 53min 14s\nWall time: 2min 4s</p>\n\n<h1>Using 16 squares</h1>\n\n<p>CPU times: user 1h 15min 50s, sys: 2min 41s, total: 1h 18min 32s\nWall time: 2min 48s\n```</p>\n\n<p>It seems the batch version does run faster, at the cost of some higher memory footprint (due to vectorization)</p>\n\n<p>Of course, the times can decrease/increase depending on the image sizes, number of squares which you already mentioned above, and possibly other factors.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 945580,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-26T01:30:08.590000",
      "content": "<p>Notice how experiment 1 <strong>without</strong> dropout has a larger gap between train AUC and validation AUC. While experiment 2 <strong>with</strong> dropout has a smaller gap. In the below example, dropout has helped the CNN generalize thus <strong>increasing validation</strong> AUC while preventing training AUC from overfitting (i.e. <strong>reducing training</strong> AUC). Notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a></p>\n\n<h2>WITHOUT coarse dropout</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F731ced48b76769659f1358d253f137af%2FScreen%20Shot%202020-07-25%20at%206.20.23%20PM.png?generation=1595726783638841&amp;alt=media\" alt=\"\"></p>\n\n<h2>WITH coarse dropout</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff8547390e3c0c420be6336ed92f44dae%2FScreen%20Shot%202020-07-25%20at%206.20.58%20PM.png?generation=1595726813782566&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 944950,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-07-25T13:22:38.833000",
      "content": "<p>This is really useful for tensorflow users in this competition <a href=\"/cdeotte\">@cdeotte</a> :D </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 944692,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-25T09:00:35.247000",
      "content": "<p>UPDATE: If your forked version 15 or earlier of the starter notebook, note that it does <strong>not</strong> apply <strong>dropout</strong>. (It only applies <strong>upsample</strong>). During training, in the call to <code>get_dataset()</code>, I forgot to pass the dropout parameters (so they use the default of zero). In the 3 lines below, the last line was missing in notebook version 15 and earlier. I just committed notebook now. Versions 16 onward will have this line added and <strong>will apply</strong> dropout.</p>\n\n<pre><code>        get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n            dim=IMG_SIZES[fold], batch_size = BATCH_SIZES[fold],\n            droprate = DROP_FREQ[fold], dropct = DROP_CT[fold], dropsize = DROP_SIZE[fold])\n</code></pre>\n\n<p>We also need to add this last line to validation TTA prediction and test TTA prediction.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 954351,
      "author_name": "cocoinit23",
      "author_url": "",
      "post_date": "2020-08-01T16:46:06.593000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thank you for your great work! Please allow me one question.\nHow do you think about the difference between Coarse Dropout and GridMask ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 954356,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T16:52:39.810000",
          "content": "<p>I have never used GridMask before. There is an implementation of GridMask for TensorFlow <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132986\">here</a> by Xie29. </p>\n\n<p>(Note: I have not tested Xie29's implementation. If you use it, i would compare training <strong>with</strong> and <strong>without</strong> it to make sure the code is efficient and doesn't slow down training).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954362,
          "author_name": "cocoinit23",
          "author_url": "",
          "post_date": "2020-08-01T17:05:21.667000",
          "content": "<p>Thank you for your quick reply and suggestion. I plan to try both of them.\nAgain, I appreciate your nice works!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 952933,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-07-31T10:34:42.307000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Hi Chris! I really would like to try to find ensembling weights by using a machine learning algorithm and cv oof auc scores. Did you ever try this? Do you have an url or tip how to do this?\nBest Roman</p>",
      "votes": 1,
      "replies": [
        {
          "id": 953123,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-31T14:47:45.470000",
          "content": "<p>This is what all the top competitors do in every competition. This is the correct way to ensemble. (The incorrect way is trial and error using feedback from public LB because public test set contains less data than train set).</p>\n\n<p>You train multiple models using the same KFold folds and save all your OOF. Then you find weights with the OOF. For example if you have 3 models, then you find <code>w1, w2, w3</code> with <code>w1+w2+w3 = 1</code> where you maximize AUC below</p>\n\n<pre><code>oof = w1*oof1 + w2*oof2 + w3*oof3\nAUC = roc_auc_score(TRUE, oof)\n</code></pre>\n\n<p>Lastly you use these weights for your Kaggle submission as in</p>\n\n<pre><code>pred = w1*sub1 + w2*sub2 + w3*sub3\nsub = pd.DataFrame(dict(image_name=NAMES, target=PRED))\nsub.to_csv('submission.csv',index=False)\n</code></pre>\n\n<p>The simplest thing is to (1) find these weights by hand or grid search (i.e. with nested for-loops). Or (2) you can find these weights by fitting a linear or logistic regression model. (Since linear and logistic regression are simple models, you don't need holdout set, just <code>fit_transform</code> all the oof against true). Or (3) if you want to use a more complex stage 2 model like XGB, you need to set up a holdout set.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 967043,
              "author_name": "Innat",
              "author_url": "",
              "post_date": "2020-08-11T21:33:02.200000",
              "content": "<p>Hello, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the above pointer.  I was trying to set weight values based on the CV score rather than LB feedback with different trails and errors.  However, I totally agree with <a href=\"https://www.kaggle.com/romanweilguny\" target=\"_blank\">@romanweilguny</a> that we should use ML algorithm to find an optimal weight for the ensemble. </p>\n<p>Later I've tried to use <a href=\"https://www.kaggle.com/ipythonx/efficientnet-b6-oof-weights-finder-seed-42\" target=\"_blank\">the following method</a> to find optimal weight, though few things are yet not clear to me, working on it. Do you have any suggestions for that? Basically we've kinda followed the above procedure that you've mentioned, the aim was to get <code>max(roc_auc_score(TRUE, oof))</code> after some iteration. </p>\n<p>Thank you -)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 967258,
              "author_name": "Signal",
              "author_url": "",
              "post_date": "2020-08-12T05:53:37.050000",
              "content": "<p>I think an ML approach would be interesting.  Not sure how fast the compute is you have access to, but say you had 4 models you wanted to ensemble, you could brute force the weights with 10,000 iterations, which can totally be parallelized across cores, if you used values of .00, .10, .20, etc through .90.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 953320,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-07-31T17:36:33.090000",
          "content": "<p>Thx - thats what I have been looking for!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954193,
          "author_name": "Rishiacharya",
          "author_url": "",
          "post_date": "2020-08-01T13:33:51.620000",
          "content": "<p>I have noticed in your notebooks you select the best model in a fold based on the lowest loss.Would it be wrong to select the best model using max AUC, even if there is 0.3-0.5 difference between the lowest loss and the max AUC epoch's loss? Or is it  important to select a model based on loss</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 954268,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T15:20:30.477000",
          "content": "<p>There are basically 3 options.\n* select lowest val loss\n* select highest val auc\n* select after fixed number of epochs, like 12th each time</p>\n\n<p>There are pros and cons for each. The resultant LB AUC will be some random number added or subtracted to each Val AUC, so any of these 3 choices could produce the best LB AUC. We cannot know which is the best choice. Some people will even take all 3 and ensemble the 3 models (but that's not necessarily better either).</p>\n\n<p>In conclusion, i'm saying experiment will all, and do what you prefer. All will work well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959717,
          "author_name": "ELEVEN",
          "author_url": "",
          "post_date": "2020-08-05T20:55:43.773000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Hi Chris! Do you use OOF with TTA or without TTA, selecting weights? </p>\n\n<blockquote>\n  <p>You train multiple models using the same KFold folds and save all your OOF. Then you find weights with the OOF. For example if you have 3 models, then you find w1, w2, w3 with w1+w2+w3 = 1 where you maximize AUC below</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 967089,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-11T23:12:53.580000",
          "content": "<p>You should the same thing that was applied to your <code>submission.csv</code>. If your submission.csv has TTA applied then use OOF with TTA applied.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 952157,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2020-07-30T17:05:00.073000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> a simple question: Do you have any ratio between image size and rest of the droprate, dropct and dropsize or you keep them same for all? Thanks for your great sources for this competition again...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 952231,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-30T18:12:41.463000",
          "content": "<p>Good question. <code>DROP_SIZE</code> is a ratio between 0 and 1. So it scales your squares with your image size. (Square side to cutout out is <code>IMAGE_SIZE * DROP_SIZE</code>). Likewise <code>DROP_RATE</code> and <code>DROP_CT</code> naturally scale. Therefore once you find good numbers, you can use the same numbers for all image sizes.</p>\n\n<p>Here's an important note. The variable <code>DROP_CT</code> removes <code>CT</code> squares using a <code>for-loop</code>. I recommend using 16 or less and preferably 8 or less. If you want to remove more image, increase <code>DROP_RATE</code> and/or <code>DROP_SIZE</code>.</p>\n\n<p>If you turn this augmentation on and off, you will find that if <code>DROP_CT&amp;gt;=16</code> then it starts to slightly slow down your training time per epoch because the CPU can't augment (execute for-loop) fast enough (when doing all the other common stuff like rotation, sheer, zoom, etc etc).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 952236,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-30T18:18:15.123000",
          "content": "<p>So for example if you set <code>DROP_FREQ = 1.0</code>, <code>DROP_CT = 8</code>, and <code>DROP_SIZE = 0.2</code> then regardless of image size, it probably removes about 20% of every image. (If the squares did not overlap it would be 25%).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 952241,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2020-07-30T18:31:41.527000",
          "content": "<p>Thank you for the answer! I wanted to play with your approach today, implemented it on my model and some feedback here maybe you find them useful:</p>\n\n<p>In my case I didn't feel slowing down up to 16 CT's maybe 5% longer per epoch or something at high numbers, didn't go over 16 though.</p>\n\n<p>But it definitely slowed down my models learning process with increased randomness, had to increase epochs like 30% to reach dropout off levels of train auc.</p>\n\n<p>Haven't played enough to comment about cv, lb yet... But anyways thank you again! <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952284,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-30T19:23:31.897000",
          "content": "<p>Yes using image dropout or layer dropout means that you have to train for more epochs. The idea is that your model is learning slower but it will achieve a higher level of intelligence.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952613,
          "author_name": "adigamer970",
          "author_url": "",
          "post_date": "2020-07-31T04:55:47.803000",
          "content": "<p>Is there any method to speed up the for loop process. It seems to waste the incredible capabilities of TPU while doing for loop. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952631,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-31T05:17:12.977000",
          "content": "<p>It does not waste TPU time. All augmentation of <code>tf.data.Dataset</code> happens on the TPU's CPU in parallel. So while the TPU is training one batch, the CPU is preparing the next batch. As long as <code>DROP_CT</code> is low enough it doesn't change train time at all (because the CPU batch time is less than the TPU batch time).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 947100,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-07-27T04:39:07.767000",
      "content": "<p>Good post Chris!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 945624,
      "author_name": "Hideaki Takahashi",
      "author_url": "",
      "post_date": "2020-07-26T03:19:34.997000",
      "content": "<p>In my understanding, we should not use dropout or cutout when predicting validation data or test data. Am I right?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 945695,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T04:53:30.620000",
          "content": "<p>The best answer is try both and see what produces higher CV LB. Since we're using TTA in both our validation and test, I think it may be best to add dropout and/or cutout to the TTA.</p>\n\n<p>I will point out that simply removing dropout or cutout from predicting test or validation may not work. When you remove a dropout layer inside a CNN, you need to rescale the activations to adjust for the missing dropout. I'm not sure how you would \"adjust\" the validation or test predictions if we train with coarse dropout and then don't use coarse dropout for prediction.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 946158,
          "author_name": "Hideaki Takahashi",
          "author_url": "",
          "post_date": "2020-07-26T11:59:09.063000",
          "content": "<p>Thanks for your reply! I did a few experiments (I know it is not sufficient.) and found that as you said, removing dropout or cutout from prediction is worse. I'll report more reliable results asap! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946176,
          "author_name": "Hideaki Takahashi",
          "author_url": "",
          "post_date": "2020-07-26T12:18:17.910000",
          "content": "<p>Though I can understand the importance of TTA, I can't understand why using dropout works with validation or test data. My concern is that using these augmentations may lose important information. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 946373,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T14:41:27.497000",
          "content": "<p>Keep in mind that the CNN is being trained with dropout. (And is smart enough to find malignant even with rectangles of images removed). If during TTA prediction there is no dropout, then the CNN could get confused.</p>\n\n<p>That's like teaching a tennis player to use a solid wood racket for months and during the tournament giving the tennis player a real racket with strings. You would think it would help, but the tennis player never learned how to use the real racket with strings.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 946457,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-26T15:34:31.377000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> This is the first time I see someone using random cutout in TTA to be honest. I will try it, interesting idea. Usually it works well to train with cutout and don't do cutout in inference, it behaves very similar to dropout which you also dont use in inference. </p>\n\n<p>What it definitely does it that it adds randomness, but having randomness in inference can be risky. Random TTA is already something I dont feel comfortable with.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946461,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T15:38:32.620000",
          "content": "<p>For TTA (test time augmentation), isn't it best to use the augmentation that your network was trained with?</p>\n\n<p>If we remove image dropout from TTA, should we adjust predictions? Because if you add a dropout layer in your CNN like <code>x = Dropout(0.2)(x)</code>. Then when the network makes predictions it removes the layer dropout and it reduces the activations by 20%. (i.e. during inference Keras replaces layer dropout with <code>x = Lambda(x: 0.8*x)(x)</code>) What is the equivalent of \"reducing activations by 20%\" when using image dropout?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946466,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-26T15:43:55.413000",
          "content": "<p>Doesn't the scaling happen during training and prediction is identity (is similar to what you say)? But you are right, haven't thought much about that when doing cutout. In practise I havent seen any issues with not doing cutout in inference, but also have never tried to keep it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946467,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T15:45:35.240000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> </p>\n\n<blockquote>\n  <p>Random TTA is already something I dont feel comfortable with.</p>\n</blockquote>\n\n<p>This competition takes TTA to a whole new level. People are using 25+ TTA steps. At that point, randomness gets averaged away.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946474,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-26T15:50:05.400000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Havent tested tta yet, but if tta outside of standard flips works this can be an approach indeed. Inference is fast.</p>\n\n<p>Do you just use train augs multiple times, or limit them?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946481,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T15:51:27.990000",
          "content": "<blockquote>\n  <p>Doesn't the scaling happen during training and prediction is identity</p>\n</blockquote>\n\n<p>Oh yes, i think you're right. A layer of <code>x = dropout(0.2)(x)</code> simultaneously applies 20% dropout and increases activations by 25% internally with something like <code>x = Lambda(x: 1.25*x)(x)</code>. Then during inference it just removes the dropout layer and all is well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946491,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T15:55:48.483000",
          "content": "<p>For TTA, i use all the training augmentations! I use rotation, sheer, scale, shift, contrast, saturation, brightness, flip, dropout, etc, etc. Using TTA 25+ is the powerhouse in this comp. It increases CV LB by something like 0.01+ or 0.02+!</p>\n\n<pre><code>    print('Predicting Test with TTA...')\n    ds_test = get_dataset(files_test,labeled=False,return_image_names=False,augment=True,\n        repeat=True,shuffle=False,dim=IMG_SIZES[fold],batch_size=BATCH_SIZE)\n    ct_test = count_data_items(files_test); STEPS = TTA * ct_test/BATCH_SIZE\n    pred = model.predict(ds_test,steps=STEPS,verbose=VERBOSE)[:TTA*ct_test,] \n    preds += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) / skf.n_splits\n</code></pre>\n\n<p>And <code>get_dataset(augment=True)</code> is the same pipeline that train uses.</p>\n\n<p>The power of TTA was first shown by AgentAuers' in his notebook <a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">here</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946501,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-26T16:00:57.423000",
          "content": "<p>Thank you, interesting, will try it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946509,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-26T16:09:43.877000",
          "content": "<p>I removed Cutout during my TTA. </p>\n\n<p>But I will try it to see if it improves. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946531,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T16:21:54.737000",
          "content": "<p>Yes <a href=\"/serigne\">@serigne</a> , let us know. I have only recently added cutout to my augmentations. So I haven't done enough experiments myself to know whether it is best to include cutout or not during TTA inference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946549,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-26T16:36:20.617000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a>   Cutout definitely improved my training model.   But, as <a href=\"/philippsinger\">@philippsinger</a> ,  I was much reluctant to use Random Augmentation (specially those cutting the input images) during TTA. </p>\n\n<p>But, after <a href=\"/optimo\">@optimo</a> suggestion , I gave a try to +15 TTA  (with all my training augs except Cutout) and saw relatively huge improvement on LB. Next step is to add Cutout and let you know. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946560,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-26T16:41:15.190000",
          "content": "<p><a href=\"/serigne\">@serigne</a> Does it also improve CV?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946602,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-26T17:10:28.570000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> for me it improves CV by about 0.01.</p>\n\n<p>To be honest I started with a lazy implementation, I was using coarse dropout and I left it as is in my TTA, but I was planning to remove it at some point. I agree that I would expect better results without masking part of my input during inference (I guess everyone would agree not to do mixcut during inference for example). It might also depend on the size of your dropout patches, if they are too big, you may be able to train your model but using it for inference would hurt your results. </p>\n\n<p>But I feel very comfortable with doing 20x TTA or even 100x TTA (even though I have better things to do haha), for me it would just converge to the best possible prediction for my model, it might be useless at some point but can't harm much. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946604,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-26T17:12:03.947000",
          "content": "<p>Oups I was eager to try it on test data but I didn't even check on validation data ( I used to try Random Augs only on training data so far)</p>\n\n<p>But you're right. These augmentations need to be checked on validation data first.  I will try to do it more systematically from now. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 947551,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-27T10:40:42.523000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> You talk about using TTA in both your validation as well as test.  I have used TTA in test, but not tried it in validation and will give it a try.  So do you end up doing the same augmentations you do in train, on both validation and test?  Or do you do some on one and not on others?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 947753,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-27T13:14:11.913000",
          "content": "<p>I use TTA in both test and validation. That way I can choose the TTA parameters that maximize validation score. Currently, I use all augmentation techniques for train, validation, and test. (I have not experimented with using less augmentation on validation and test compared with train. Perhaps using less would increase CV score more).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 950170,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-29T08:19:29.713000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Cutout on TTA improved further my LB score . </p>\n\n<p>I have run an experiment for fold0 on 256x256. <br>\n10 TTA random Augment w/o Cutout : LB 0.9357\n10 TTA random Augment w/ Cutout : LB 0.9392</p>\n\n<p>However I haven't check yet any random augment on validation data. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F974295%2F5b02ef384e321357eabe730a8e38a4ce%2FCapre.PNG?generation=1596010735794067&amp;alt=media\" alt=\"\"></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 970107,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-08-14T07:58:54.750000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks (you replied 18 days ago but just thinking again about this).  In your Triple Stratified notebook, if I am reading it right (I run PyTorch so I have not run that notebook), your validation dataset looks like it has <code>augment=False</code> or something along the lines of that, which is why I thought maybe in your approach you didn't do it in validation.  But I am possibly just reading your notebook wrong.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944899,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2020-07-25T12:40:13.797000",
      "content": "<p>Can we also extend it for images with 4 channels? And also why compiler needs to know <code>image = tf.reshape(image,[DIM,DIM,3])</code>? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 944959,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T13:28:43.517000",
          "content": "<p>TensorFlow needs to determine the output shape of each function that operates on <code>tf.data.Dataset()</code>. If you remove <code>image = tf.reshape(image,[DIM,DIM,3])</code>, then i think TensorFlow complains that it can't determine the shape. (I'm not sure why it can't figure it out).</p>\n\n<p>If you have images with 4 channels, just change the 3 to a 4 in two places\n* <code>two = tf.zeros([yb-ya,xb-xa,3])</code>\n*  <code>image = tf.reshape(image,[DIM,DIM,3])</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944961,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T13:30:05.877000",
          "content": "<p>Or replace both 3 with <code>image.shape[-1]</code>. That should work and then it can handle whatever number of channels comes its way.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944972,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-07-25T13:33:33.503000",
          "content": "<p>Thanks. I'm bothering you too much these days I guess...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944488,
      "author_name": "Khan Fashee Monowar (Sawrup)",
      "author_url": "",
      "post_date": "2020-07-25T06:42:59.263000",
      "content": "<p>Can I use this function for JPEG image augmentation ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 944575,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T07:36:13.063000",
          "content": "<p>For JPEG, you can use Albumentations' <code>albumentations.augmentations.transforms.Cutout</code> and <code>albumentations.augmentations.transforms.CoarseDropout</code>, the API is <a href=\"https://albumentations.readthedocs.io/en/latest/api/augmentations.html#module-albumentations.augmentations.transforms\">here</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944587,
          "author_name": "Khan Fashee Monowar (Sawrup)",
          "author_url": "",
          "post_date": "2020-07-25T07:44:52.133000",
          "content": "<p>Thank you so much.This is very helpful.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 944404,
      "author_name": "Luck",
      "author_url": "",
      "post_date": "2020-07-25T04:54:44.833000",
      "content": "<p>Nice work! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 944383,
      "author_name": "Sarthak khandelwal",
      "author_url": "",
      "post_date": "2020-07-25T04:26:19.457000",
      "content": "<p>Nice and precise implementations <a href=\"/cdeotte\">@cdeotte</a>! I was curious if <code>AutoAugment</code> can also be implemented in TF as it is proved to improve the model's performance. I found the pytorch implementation of <code>AutoAugmentation</code> <a href=\"https://github.com/DeepVoltaire/AutoAugment\">here</a>. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 944577,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T07:37:02.843000",
          "content": "<p>Sure. Anything can be coded in TensorFlow. This would be interesting to try out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944232,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2020-07-25T00:15:18.477000",
      "content": "<p>Why makes it suitable for TPU and GPU? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 944236,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T00:20:18.633000",
          "content": "<p>This code can be executed directly on the GPU/TPU if you desire.</p>\n\n<p>TensorFlow runs their <code>tf.data.Dataset</code> pipeline on CPU in parallel with GPU/TPU training but once you have TensorFlow code, you can push it into a <code>tf.keras.layer</code> and make it part of your model. Then it will execute on the GPU/TPU inside your model. Sometimes this is faster than CPU parallel sometimes it is not. </p>\n\n<p>However, i don't suggest pushing it into a <code>tf.keras.layer</code>. I suggest just using it with <code>tf.data.Dataset</code> if you're using TensorFlow GPU/TPU.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 944317,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-07-25T02:41:14.440000",
          "content": "<p>It seems that we can't use albumentations in tf.data.Dataet of if we can somehow convert all the operations in tensorflow then they'll work fine?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944331,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-25T03:01:14.850000",
          "content": "<p>Yeah, unfortunately we cannot use Albumentations with <code>tf.data.Dataset</code> nor use <code>tf.contrib</code> libraries nor any libraries. All augmentation needs to be written from scratch using TensorFlow language. </p>\n\n<p>I think we now have all the code to do the main augmentations with <code>tf.data.Dataset</code>. Rotation, Sheer, Zoom, Shift <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\">here</a>, Cutmix and mixup <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">here</a>, and now Cutout and sparse dropout in this discussion. And TensorFlow already has functions for color manipulations such as contrast, brightness, hue, sharpness, etc.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 944227,
      "author_name": "Bruce Young",
      "author_url": "",
      "post_date": "2020-07-25T00:04:22.700000",
      "content": "<p>Thanks again for this Chris. I've been resisting the temptation to reduce the number and types of augmentations. I know that in the long run, I'll be better off. As you mentioned it's not as simple as one would hope to do custom augmentations in TF. I'm keen to try this out. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 947848,
      "author_name": "Stephan",
      "author_url": "",
      "post_date": "2020-07-27T14:10:36.407000",
      "content": "<p>Technically, you can use Albumentations through the tf dataset pipeline using tf.numpy_function. The downside is that TPU's do not support this feature (yet). On a GPU however, this should work.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 947855,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-27T14:15:53.297000",
          "content": "<p>Thanks for letting me know that.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 945900,
      "author_name": "YujiAriyasu",
      "author_url": "",
      "post_date": "2020-07-26T07:49:01.987000",
      "content": "<p>You don't seem to be using hair_aug(), do you think we should use hair augmentation?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 946375,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T14:42:52.723000",
          "content": "<p>I have not tried hair augmentation yet. I would guess that it is helpful. Has anyone posted a TensorFlow <code>tf.data.Dataset</code> implementation? Please post the link here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946806,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-26T21:18:55.550000",
          "content": "<p>Here is my notebook implementing Roman's advance hair augmentation in TF:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-data-augmentation-in-tf-hair-batch-affine\">SIIM: Data Augmentation in TF: Hair + Batch Affine</a></p>\n\n<p>I have not converted it to the batch form, so it does slow down training. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 946943,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-27T01:00:18.797000",
          "content": "<p>Great thanks <a href=\"/graf10a\">@graf10a</a> , i will play with it.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1182038,
      "author_name": "Shruti",
      "author_url": "",
      "post_date": "2021-02-02T09:15:43.760000",
      "content": "<p>This is really very useful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1182039,
      "author_name": "Shruti",
      "author_url": "",
      "post_date": "2021-02-02T09:15:43.760000",
      "content": "<p>This is really very useful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972499,
      "author_name": "( ͡° ͜ʖ ͡°)",
      "author_url": "",
      "post_date": "2020-08-16T15:31:32.547000",
      "content": "<p>is it same the cutmix ?</p>\n<p>def cutmix_aug(image, label=None, DIM=256, BATCH_SIZE=4, PROBABILITY=0.4):<br>\n    # input image - is a batch of images of size [n,dim,dim,3] not a single image of [dim,dim,3]<br>\n    # output - a batch of images with cutmix applied</p>\n<pre><code>imgs = []; labs = []\n\nfor j in range(BATCH_SIZE):\n\n    #random_uniform( shape, minval=0, maxval=None)        \n    # DO CUTMIX WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast(tf.random.uniform([], 0, 1) &lt;= PROBABILITY, tf.int32)\n\n    # CHOOSE RANDOM IMAGE TO CUTMIX WITH\n    k = tf.cast(tf.random.uniform([], 0, BATCH_SIZE), tf.int32)\n\n    # CHOOSE RANDOM LOCATION\n    x = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n    y = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n\n    # Beta(1, 1)\n    b = tf.random.uniform([], 0, 1) # this is beta dist with alpha=1.0\n\n\n    WIDTH = tf.cast(DIM * tf.math.sqrt(1-b),tf.int32) * P\n    ya = tf.math.maximum(0,y-WIDTH//2)\n    yb = tf.math.minimum(DIM,y+WIDTH//2)\n    xa = tf.math.maximum(0,x-WIDTH//2)\n    xb = tf.math.minimum(DIM,x+WIDTH//2)\n\n    # MAKE CUTMIX IMAGE\n    one = image[j,ya:yb,0:xa,:]\n    two = image[k,ya:yb,xa:xb,:]\n    three = image[j,ya:yb,xb:DIM,:]        \n    #ya:yb\n    middle = tf.concat([one,two,three],axis=1)\n\n    img = tf.concat([image[j,0:ya,:,:],middle,image[j,yb:DIM,:,:]],axis=0)\n    imgs.append(img)\n\n    # MAKE CUTMIX LABEL\n    a = tf.cast(WIDTH*WIDTH/DIM/DIM,tf.float32)\n    lab1 = label[j,]\n    lab2 = label[k,]\n    labs.append((1-a)*lab1 + a*lab2)\n\nimage2 = tf.reshape(tf.stack(imgs),(BATCH_SIZE, DIM, DIM, 3))\nlabel2 = tf.reshape(tf.stack(labs),(BATCH_SIZE, 1))\nreturn image2, label2\n</code></pre>\n<p>this script handle the label too, cast label to float32</p>",
      "votes": 0,
      "replies": [
        {
          "id": 972924,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-17T01:25:30.137000",
          "content": "<p>I published a CutMix notebook <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\" target=\"_blank\">here</a>. For coarse dropout, I used that CutMix notebook and just replace the second image with all zeros (to perform the dropout).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972971,
          "author_name": "( ͡° ͜ʖ ͡°)",
          "author_url": "",
          "post_date": "2020-08-17T02:22:35.937000",
          "content": "<p>ah ok, so label still keep 👍</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 949312,
      "author_name": "Sagar",
      "author_url": "",
      "post_date": "2020-07-28T14:51:35.453000",
      "content": "<p>This method is cool .\nBut I have a doubt\nas the dataset  is imbalanced , does by adding small black boxes to the image increases its performance because it is learning from the presence of small black boxes to predict the image as non malignant and as the number of non malignant class are more it may cause a delusion that the model is working better with the dropout  where as it is learning wrong thing which will affect its performance on the test data </p>",
      "votes": 0,
      "replies": [
        {
          "id": 949342,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-07-28T15:13:19.070000",
          "content": "<p>As long as the black boxes are added to both benign and malignant images, all is good.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 949400,
          "author_name": "Sagar",
          "author_url": "",
          "post_date": "2020-07-28T16:04:50.390000",
          "content": "<p>Thank you for your time and response .\nI also wanted to know for this competition   up sampling the malignant data is better than using weights or do you advise on using both at the same time </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949433,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-28T16:25:01.307000",
          "content": "<p>Upsampling or using weights are very similar. For example you can double all the malignant images or you can use weights <code>{0:1,  1:2}</code> (malignant is 2x). One advantage of weights is less train data and faster training epochs. One advantage of upsample is that each batch now contains some malignant and the malignant learning process is more incremental.</p>\n\n<p>Additionally since we're using data augmentation and external data malignant images, then doing upsample adds more variety to the malignant class and hopefully encourages more malignant generalization.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 945544,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2020-07-26T00:02:59.827000",
      "content": "<p>I tried Dropout in another Dataset and I got this error while training...\n<code>\nCompilation failure: Dynamic Spatial Convolution is not supported\n</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 945550,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-26T00:16:04.553000",
          "content": "<p>hmm... i've seen that error before. It relates to TensorFlow knowing the size of stuff. Can your post your code that adds the dropout to your <code>tf.data.Dataset</code>? Make sure you give <code>dropout</code> the correct image size. And make sure that you add dropout <strong>before</strong> you batch, i.e. <code>ds = ds.map(dropout)</code> <strong>before</strong> <code>ds = ds.batch(BATCH_SIZE)</code></p>\n\n<pre><code>dropout(image, DIM=???, PROBABILITY = 0.75, CT = 8, SZ = 0.2)\n</code></pre>\n\n<p>replace ??? with correct size</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946183,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-26T12:23:43.453000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946188,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-07-26T12:26:32.670000",
          "content": "<p>Same  code runs for GPU and CPU but shows error in TPU..\n```\ndef dropout_img(image,DIM=256, PROBABILITY = 0.5, CT = 50, SZ = 0.06):\n    # input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed</p>\n\n<pre><code># DO DROPOUT WITH PROBABILITY DEFINED ABOVE\nP = tf.cast( tf.random.uniform([],0,1)&lt;PROBABILITY, tf.int32)\nif (P==0)|(CT==0)|(SZ==0): \n    return image, mask\n\nfor k in range(CT):\n    # CHOOSE RANDOM LOCATION\n    x = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n    y = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n    # COMPUTE SQUARE \n    WIDTH = tf.cast( SZ*DIM,tf.int32) * P\n    ya = tf.math.maximum(0,y-WIDTH//2)\n    yb = tf.math.minimum(DIM,y+WIDTH//2)\n    xa = tf.math.maximum(0,x-WIDTH//2)\n    xb = tf.math.minimum(DIM,x+WIDTH//2)\n    # DROPOUT IMAGE\n    one = image[ya:yb,0:xa,:]\n    two = tf.zeros([yb-ya,xb-xa,4], dtype = image.dtype) \n    three = image[ya:yb,xb:DIM,:]\n    middle = tf.concat([one,two,three],axis=1)\n    image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM,:,:]],axis=0)\n\nimage = tf.reshape(image,[DIM,DIM,-1])\n\nreturn image\n</code></pre>\n\n<p><code>\n</code>\ndef augmentation_img(img, dim=256):   </p>\n\n<pre><code>seed = np.int(np.round(np.random.uniform()))\n\nimg = tf.image.random_flip_left_right(img, seed = seed)\n\nimg = transform_img(img, DIM =dim)\nimg = dropout_img(img, DIM=dim, PROBABILITY = 0.5, CT = 50, SZ = 0.06)\n\n\nimg = tf.image.random_contrast(img, 0.8, 1.2)\nimg = tf.image.random_brightness(img, 0.1)\n\nimg = tf.reshape(img, [dim, dim, -1])\n\nreturn img\n</code></pre>\n\n<p><code>\n</code>\n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO)\n    ds = ds.cache()</p>\n\n<pre><code>if repeat:\n    ds = ds.repeat()\n\nif shuffle: \n    ds = ds.shuffle(1024*8)\n    opt = tf.data.Options()\n    opt.experimental_deterministic = False\n    ds = ds.with_options(opt)\n\n ds = ds.map(lambda example: read_unlabeled_tfrecord(example,          return_image_names),num_parallel_calls=AUTO) \n\n    if augment:\n        ds = ds.map(lambda img, name: (augmentation_img(img, dim=dim), name), num_parallel_calls=AUTO)\n\nds = ds.batch(batch_size * REPLICAS)\nds = ds.prefetch(AUTO)\nreturn ds\n</code></pre>\n\n<p>```</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 967894,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-12T15:10:31.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 953250,
      "author_name": "Aditya Baurai",
      "author_url": "",
      "post_date": "2020-07-31T16:14:52.437000",
      "content": "<p>Thank you !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 951260,
      "author_name": "Shravan Raikar",
      "author_url": "",
      "post_date": "2020-07-30T02:49:07.167000",
      "content": "<p>Thanks for Sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 948203,
      "author_name": "Sanjay K",
      "author_url": "",
      "post_date": "2020-07-27T18:14:56.500000",
      "content": "<p>Nice work, thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "944208": "When using TensorFlow's `tf.data.Dataset()` we cannot use Albumenations to do our augmentation. Instead we must write our own functions in TensorFlow. Previously I posted how to perform (1) Rotation, Sheer, Zoom, Shift for GPU/TPU [here][1] and [here][3] (2) Cutmix and Mixup for GPU/TPU [here][2]\n\n# Coarse Dropout and Cutout Augmentation for GPU/TPU\nCoarse Dropout and Cutout augmentation are techniques to prevent overfitting and encourage generalization. They randomly remove rectangles from training images. By removing portions of the images, we challenge our models to pay attention to the entire image because it never knows what part of the image will be present. (This is similar and different to dropout layer within a CNN).\n\n* Cutout is the technique of removing 1 large rectangle of random size \n* Coarse dropout is the technique of removing many small rectanges of similar size. \n\nBy changing the parameters below, we can have either coarse dropout or cutout. (For cutout, you'll need to add `tf.random.uniform` for random size. I leave this as an exercise for the reader).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8db206436734b5f0c7b2179c2e152eb0%2FScreen%20Shot%202020-07-22%20at%201.28.49%20PM.png?generation=1595632223386869&amp;alt=media)\n\n\n# Starter Notebook\nI've posted a starter notebook [here][4] showing how to do coarse dropout with TFRecords and `tf.data.Dataset`. The starter notebook also demonstrates upsampling (oversampling) the minority class.\n\n# Code\nThe reason that this code looks overly complicated just to turn a portion of an array to zeros is because you cannot do `image[ya:yb,xa:xb,:] = 0` in TensorFlow. TF does not allow index assignment. Instead we break an image into five pieces and then concatenate the five pieces while replacing the middle piece with all zeros.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F476c96c48b8b31969e00b55b26cef2f3%2Fcut.jpg?generation=1595632770269562&amp;alt=media)\n\n    def dropout(image, DIM=256, PROBABILITY = 0.75, CT = 8, SZ = 0.2):\n        # input - one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n        # output - image with CT squares of side size SZ*DIM removed\n    \n        # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n        P = tf.cast( tf.random.uniform([],0,1) &lt; PROBABILITY, tf.int32)\n        if (P == 0)|(CT == 0)|(SZ == 0): return image\n    \n        for k in range( CT ):\n            # CHOOSE RANDOM LOCATION\n            x = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n            y = tf.cast( tf.random.uniform([],0,DIM),tf.int32)\n            # COMPUTE SQUARE \n            WIDTH = tf.cast( SZ*DIM,tf.int32) * P\n            ya = tf.math.maximum(0,y-WIDTH//2)\n            yb = tf.math.minimum(DIM,y+WIDTH//2)\n            xa = tf.math.maximum(0,x-WIDTH//2)\n            xb = tf.math.minimum(DIM,x+WIDTH//2)\n            # DROPOUT IMAGE\n            one = image[ya:yb,0:xa,:]\n            two = tf.zeros([yb-ya,xb-xa,3]) \n            three = image[ya:yb,xb:DIM,:]\n            middle = tf.concat([one,two,three],axis=1)\n            image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM,:,:]],axis=0)\n            \n        # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR \n        image = tf.reshape(image,[DIM,DIM,3])\n        return image\n\n[1]: https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\n[2]: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\n[3]: https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\n[4]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
    "953624": "I've created a version that accepts batches of images, for anyone who is interested. I've tested it on a GPU setup, so I can't be sure this works on a TPU yet.\n\nedit: I see that the prob arg isn't actually utilized yet, I'll fix this tommorow.\nedit 2: probabilities now included :)\nedit 3: replaced a type in tf.scatter_nd where the image sizes were fixed to 128,128 -&gt; now adjusted for variable sizes\n\n```python\ndef coarse_dropout(images, square_size, num_squares, prob):\n    ''' Coarse image dropout \n    Coarse dropout masks are created by using tf.scatter_nd\n    Which can update matrices according to coordinates specified\n    by the user. \n\n    Parameters\n    ----------\n    images : tf Tensor\n        A batch of images of size [BS, H, W, C] \n    square_size : float\n        A float between [0,1], a percentage of the image size\n    num_squares : int &gt; 0\n        The number of dropout squares per single image, must be &gt;0\n    prob : float\n        A float between [0,1], probability that dropout is used for this batch\n\n    Yields\n    ------\n    TYPE\n        DESCRIPTION.\n\n    '''\n       # random dropout\n    prob = tf.cast( tf.random.uniform([],0,1) &lt; prob, tf.int32)\n    if (prob == 0): return images\n\n    img_shape = images.shape\n    _, h, w, c = img_shape[0], img_shape[1], img_shape[2], img_shape[3]\n    \n    # For some odd reason, the batch size is lost in the processing pipeline\n    # seriously, how can it get lost ._., it literally says ds.batch()...\n    # tensorflow shenanigans...\n    bs = tf.cast(tf.reduce_sum(tf.ones_like(images)) / (images.shape[1] * images.shape[2] * images.shape[3]), tf.int32)\n\n\n    # size of square in pixels (ssp), ASSUMING SQUARE IMAGES!\n    ssp = tf.cast(tf.math.ceil(h * square_size), tf.int32)\n\n    # Create random start x coordinates\n    coords_x = tf.random.uniform([bs, num_squares], \n                                 minval=0, \n                                 maxval= (h - ssp), \n                                 dtype=tf.int32) \n\n    # Create ranges from start to the end of the line\n    # [ssp, bs, num_squares]\n    coords_x = tf.linspace(coords_x, coords_x + ssp - 1, ssp)\n\n    # [num_squares, bs, ssp]\n    coords_x = tf.cast(tf.transpose(coords_x), tf.int32)\n\n    # Create random start y coordinates\n    coords_y = tf.random.uniform([bs, num_squares], \n                                 minval=0, \n                                 maxval= (h - ssp), \n                                 dtype=tf.int32) \n\n    # Create ranges from start to the end of the line\n    # [ssp, bs, num_squares]\n    coords_y = tf.linspace(coords_y, coords_y + ssp - 1, ssp)\n\n    # [num_squares, bs, ssp]\n    coords_y = tf.cast(tf.transpose(coords_y), tf.int32)\n\n\n    # Create coordinate range combinations \n    # and reshape to [bs, num_squares, 1, ssp * ssp]\n    grid_y = tf.reshape(tf.tile(coords_y, [1,1,ssp]), \n                        (bs, num_squares, 1, ssp * ssp))\n\n    # and reshape to [bs, num_squares, ssp, ssp], transpose the inner matrices\n    grid_y = tf.transpose(tf.reshape(grid_y, \n                                     (bs, num_squares, ssp, ssp)), \n                          (0, 1, 3, 2))\n\n    # Repeat for x coordinates\n    grid_x = tf.reshape(tf.tile(coords_x, [1,1,ssp]), \n                        (bs, num_squares, 1, ssp * ssp))\n\n    grid_x = tf.reshape(grid_x, (bs, num_squares, ssp, ssp))\n\n    # Stack the grids into a single matrix\n    # grid is [2, bs, num_squares, ssp, ssp]\n    grid = tf.stack([grid_y, grid_x], axis=0)\n\n    # Transpose and reshape [ bs, ssp * ssp * num_squares, 2] \n    # Creates an array of 2D coordinates ([[x1,y1], [x2, y2], ..., [xn, yn]])\n    # over all squares (num_squares), for each combination (ssp*ssp) of coordinates \n    # and each batch (bs)\n    grid = tf.reshape(tf.transpose(grid, (1, 4, 3, 2, 0)), \n                      (bs, ssp * ssp * num_squares, 2))\n\n    # [bs, sz*sz*num_squares, 2]\n    #grid = tf.reshape(grid, (bs, ssp*ssp*num_squares, 2))\n\n    # create batch indices [0,..., bs] and reshape to [ssp * ssp * num_squares, bs]\n    batch_indices = tf.reshape(tf.tile(tf.range(0, bs), \n                                       [ssp * ssp * num_squares]), \n                               (ssp * ssp * num_squares, bs))\n\n    # Transpose to get the right order, and reshape to match grid shape\n    batch_indices = tf.reshape(tf.transpose(batch_indices), \n                               (bs, ssp * ssp * num_squares, 1))\n\n    # concatenate batch indices with the grid\n    # these yield 3D coordinates like e.g.\n    # [[bs0, x1, y1], [bs0, x2, y1], ..., [bsn, xn, yn]]\n    grid = tf.concat([batch_indices, grid], axis=2)\n\n    # create a matrix of zeros, and update the matrix with the grid indices\n    # this essentially creates a mask with coarse dropout squares\n    masks = tf.scatter_nd(grid[tf.newaxis,...], \n                            tf.ones([1,bs,ssp*ssp*num_squares]) * -1, \n                            shape=(bs, h, w)) +1\n    \n    # Due to overlap of squares, some get coordinates get updated twice\n    # and result in values &lt; -1, clip these values\n    masks = tf.clip_by_value(masks, 0, 1)\n\n    return images * masks[..., tf.newaxis] \n```",
    "945580": "Notice how experiment 1 **without** dropout has a larger gap between train AUC and validation AUC. While experiment 2 **with** dropout has a smaller gap. In the below example, dropout has helped the CNN generalize thus **increasing validation** AUC while preventing training AUC from overfitting (i.e. **reducing training** AUC). Notebook [here][1]\n## WITHOUT coarse dropout\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F731ced48b76769659f1358d253f137af%2FScreen%20Shot%202020-07-25%20at%206.20.23%20PM.png?generation=1595726783638841&amp;alt=media)\n## WITH coarse dropout\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Ff8547390e3c0c420be6336ed92f44dae%2FScreen%20Shot%202020-07-25%20at%206.20.58%20PM.png?generation=1595726813782566&amp;alt=media)\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\n\n",
    "944950": "This is really useful for tensorflow users in this competition @cdeotte :D ",
    "944692": "UPDATE: If your forked version 15 or earlier of the starter notebook, note that it does **not** apply **dropout**. (It only applies **upsample**). During training, in the call to `get_dataset()`, I forgot to pass the dropout parameters (so they use the default of zero). In the 3 lines below, the last line was missing in notebook version 15 and earlier. I just committed notebook now. Versions 16 onward will have this line added and **will apply** dropout.\n\n            get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n                dim=IMG_SIZES[fold], batch_size = BATCH_SIZES[fold],\n                droprate = DROP_FREQ[fold], dropct = DROP_CT[fold], dropsize = DROP_SIZE[fold])\n\nWe also need to add this last line to validation TTA prediction and test TTA prediction.",
    "954351": "@cdeotte Thank you for your great work! Please allow me one question.\nHow do you think about the difference between Coarse Dropout and GridMask ?",
    "952933": "@cdeotte Hi Chris! I really would like to try to find ensembling weights by using a machine learning algorithm and cv oof auc scores. Did you ever try this? Do you have an url or tip how to do this?\nBest Roman",
    "952157": "@cdeotte a simple question: Do you have any ratio between image size and rest of the droprate, dropct and dropsize or you keep them same for all? Thanks for your great sources for this competition again...",
    "947100": "Good post Chris!",
    "945624": "In my understanding, we should not use dropout or cutout when predicting validation data or test data. Am I right?",
    "944899": "Can we also extend it for images with 4 channels? And also why compiler needs to know `image = tf.reshape(image,[DIM,DIM,3])`? ",
    "944488": "Can I use this function for JPEG image augmentation ?",
    "944404": "Nice work! ",
    "944383": "Nice and precise implementations @cdeotte! I was curious if `AutoAugment` can also be implemented in TF as it is proved to improve the model's performance. I found the pytorch implementation of `AutoAugmentation` [here](https://github.com/DeepVoltaire/AutoAugment). ",
    "944232": "Why makes it suitable for TPU and GPU? ",
    "944227": "Thanks again for this Chris. I've been resisting the temptation to reduce the number and types of augmentations. I know that in the long run, I'll be better off. As you mentioned it's not as simple as one would hope to do custom augmentations in TF. I'm keen to try this out. ",
    "947848": "Technically, you can use Albumentations through the tf dataset pipeline using tf.numpy_function. The downside is that TPU's do not support this feature (yet). On a GPU however, this should work.",
    "945900": "You don't seem to be using hair_aug(), do you think we should use hair augmentation?",
    "1182038": "This is really very useful",
    "1182039": "This is really very useful",
    "972499": "is it same the cutmix ?\n\ndef cutmix_aug(image, label=None, DIM=256, BATCH_SIZE=4, PROBABILITY=0.4):\n    # input image - is a batch of images of size [n,dim,dim,3] not a single image of [dim,dim,3]\n    # output - a batch of images with cutmix applied\n\n    imgs = []; labs = []\n    \n    for j in range(BATCH_SIZE):\n        \n        #random_uniform( shape, minval=0, maxval=None)        \n        # DO CUTMIX WITH PROBABILITY DEFINED ABOVE\n        P = tf.cast(tf.random.uniform([], 0, 1) <= PROBABILITY, tf.int32)\n        \n        # CHOOSE RANDOM IMAGE TO CUTMIX WITH\n        k = tf.cast(tf.random.uniform([], 0, BATCH_SIZE), tf.int32)\n        \n        # CHOOSE RANDOM LOCATION\n        x = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n        y = tf.cast(tf.random.uniform([], 0, DIM), tf.int32)\n        \n        # Beta(1, 1)\n        b = tf.random.uniform([], 0, 1) # this is beta dist with alpha=1.0\n        \n\n        WIDTH = tf.cast(DIM * tf.math.sqrt(1-b),tf.int32) * P\n        ya = tf.math.maximum(0,y-WIDTH//2)\n        yb = tf.math.minimum(DIM,y+WIDTH//2)\n        xa = tf.math.maximum(0,x-WIDTH//2)\n        xb = tf.math.minimum(DIM,x+WIDTH//2)\n        \n        # MAKE CUTMIX IMAGE\n        one = image[j,ya:yb,0:xa,:]\n        two = image[k,ya:yb,xa:xb,:]\n        three = image[j,ya:yb,xb:DIM,:]        \n        #ya:yb\n        middle = tf.concat([one,two,three],axis=1)\n\n        img = tf.concat([image[j,0:ya,:,:],middle,image[j,yb:DIM,:,:]],axis=0)\n        imgs.append(img)\n        \n        # MAKE CUTMIX LABEL\n        a = tf.cast(WIDTH*WIDTH/DIM/DIM,tf.float32)\n        lab1 = label[j,]\n        lab2 = label[k,]\n        labs.append((1-a)*lab1 + a*lab2)\n\n    image2 = tf.reshape(tf.stack(imgs),(BATCH_SIZE, DIM, DIM, 3))\n    label2 = tf.reshape(tf.stack(labs),(BATCH_SIZE, 1))\n    return image2, label2\n\nthis script handle the label too, cast label to float32",
    "949312": "This method is cool .\nBut I have a doubt\nas the dataset  is imbalanced , does by adding small black boxes to the image increases its performance because it is learning from the presence of small black boxes to predict the image as non malignant and as the number of non malignant class are more it may cause a delusion that the model is working better with the dropout  where as it is learning wrong thing which will affect its performance on the test data \n\n",
    "945544": "I tried Dropout in another Dataset and I got this error while training...\n```\nCompilation failure: Dynamic Spatial Convolution is not supported\n```",
    "967894": "",
    "953250": "Thank you !!",
    "951260": "Thanks for Sharing!",
    "948203": "Nice work, thanks for sharing!"
  }
}