{
  "id": 202560,
  "title": "Create Target for Mosaic Augmentation? ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202560",
  "author_name": "Innat",
  "post_date": "2020-12-10T18:57:35.442000",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<h2>Update</h2>\n<p>This is now supported in <code>keras-cv</code>.  The creation of labels is discussed here, <a href=\"https://github.com/keras-team/keras-cv/issues/250#issuecomment-1130256698\" target=\"_blank\">https://github.com/keras-team/keras-cv/issues/250#issuecomment-1130256698</a> . </p>\n<p><a href=\"https://github.com/keras-team/keras-cv\" target=\"_blank\">https://github.com/keras-team/keras-cv</a></p>\n<hr>\n<p>I am trying to do <strong>mosaic augmentation</strong> which first introduced in the <strong>Yolo-4</strong> paper-work. In <code>cutmix</code> or <code>mixup</code> or <code>fmix</code> we can use <code>beta</code> to weight the labels such as </p>\n<pre><code>label = label_one*beta + (1-beta)*label_two\n</code></pre>\n<p>However, unlike <code>cutmix</code> and others, a total of <strong>4</strong> samples is needed to create mosaic augmentation.  Such as below: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb0530886377eb3d7f0c29401ba069ffa%2Fdownload%20(1).png?generation=1607625768667914&amp;alt=media\" alt=\"![\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F80e35122f0fe72cc19127d28bfd1b20f%2Fdownload%20(2).png?generation=1607625836730889&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fdfce6cf87b2125c98affa7574ed38b18%2Fdownload%20(3).png?generation=1607625855803714&amp;alt=media\" alt=\"\"></p>\n<p>However, I'm not quite sure how to create a class label for this yet. One idea for that is, maybe we can use <strong>Dirichlet distribution</strong> which can give us <strong>4</strong> distributed values, such as </p>\n<pre><code>&gt;&gt; np.random.dirichlet((1, 1, 1, 1), 1) \narray([[0.38160767, 0.28880196, 0.29071801, 0.03887236]])\n</code></pre>\n<p>But it brings more complexity to the current setup. You can find the starter code from <a href=\"https://www.kaggle.com/ipythonx/tf-keras-sota-augmentation-in-sequence-generator?scriptVersionId=49012892\" target=\"_blank\">here</a>. Any efficient way to achieve this? </p>",
  "messages": [
    {
      "id": 1108532,
      "postDate": "2020-12-10T18:57:35.443Z",
      "content": "<h2>Update</h2>\n<p>This is now supported in <code>keras-cv</code>.  The creation of labels is discussed here, <a href=\"https://github.com/keras-team/keras-cv/issues/250#issuecomment-1130256698\" target=\"_blank\">https://github.com/keras-team/keras-cv/issues/250#issuecomment-1130256698</a> . </p>\n<p><a href=\"https://github.com/keras-team/keras-cv\" target=\"_blank\">https://github.com/keras-team/keras-cv</a></p>\n<hr>\n<p>I am trying to do <strong>mosaic augmentation</strong> which first introduced in the <strong>Yolo-4</strong> paper-work. In <code>cutmix</code> or <code>mixup</code> or <code>fmix</code> we can use <code>beta</code> to weight the labels such as </p>\n<pre><code>label = label_one*beta + (1-beta)*label_two\n</code></pre>\n<p>However, unlike <code>cutmix</code> and others, a total of <strong>4</strong> samples is needed to create mosaic augmentation.  Such as below: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb0530886377eb3d7f0c29401ba069ffa%2Fdownload%20(1).png?generation=1607625768667914&amp;alt=media\" alt=\"![\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F80e35122f0fe72cc19127d28bfd1b20f%2Fdownload%20(2).png?generation=1607625836730889&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fdfce6cf87b2125c98affa7574ed38b18%2Fdownload%20(3).png?generation=1607625855803714&amp;alt=media\" alt=\"\"></p>\n<p>However, I'm not quite sure how to create a class label for this yet. One idea for that is, maybe we can use <strong>Dirichlet distribution</strong> which can give us <strong>4</strong> distributed values, such as </p>\n<pre><code>&gt;&gt; np.random.dirichlet((1, 1, 1, 1), 1) \narray([[0.38160767, 0.28880196, 0.29071801, 0.03887236]])\n</code></pre>\n<p>But it brings more complexity to the current setup. You can find the starter code from <a href=\"https://www.kaggle.com/ipythonx/tf-keras-sota-augmentation-in-sequence-generator?scriptVersionId=49012892\" target=\"_blank\">here</a>. Any efficient way to achieve this? </p>",
      "rawMarkdown": "## Update \n\nThis is now supported in `keras-cv`.  The creation of labels is discussed here, https://github.com/keras-team/keras-cv/issues/250#issuecomment-1130256698 . \n\nhttps://github.com/keras-team/keras-cv\n\n\n---\n\nI am trying to do **mosaic augmentation** which first introduced in the **Yolo-4** paper-work. In `cutmix` or `mixup` or `fmix` we can use `beta` to weight the labels such as \n\n```\nlabel = label_one*beta + (1-beta)*label_two\n```\n\nHowever, unlike `cutmix` and others, a total of **4** samples is needed to create mosaic augmentation.  Such as below: \n\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb0530886377eb3d7f0c29401ba069ffa%2Fdownload%20(1).png?generation=1607625768667914&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F80e35122f0fe72cc19127d28bfd1b20f%2Fdownload%20(2).png?generation=1607625836730889&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fdfce6cf87b2125c98affa7574ed38b18%2Fdownload%20(3).png?generation=1607625855803714&alt=media)\n\nHowever, I'm not quite sure how to create a class label for this yet. One idea for that is, maybe we can use **Dirichlet distribution** which can give us **4** distributed values, such as \n\n```\n>> np.random.dirichlet((1, 1, 1, 1), 1) \narray([[0.38160767, 0.28880196, 0.29071801, 0.03887236]])\n```\n\nBut it brings more complexity to the current setup. You can find the starter code from [here](https://www.kaggle.com/ipythonx/tf-keras-sota-augmentation-in-sequence-generator?scriptVersionId=49012892). Any efficient way to achieve this? ",
      "votes": 5
    },
    {
      "id": 1108550,
      "postDate": "2020-12-10T19:24:21.080Z",
      "content": "<p>How about area weighted labels?</p>",
      "rawMarkdown": "How about area weighted labels?",
      "votes": 2,
      "replies": [
        {
          "id": 1108560,
          "postDate": "2020-12-10T19:37:32.567Z",
          "content": "<p>That actually my first assumption. Let's say if we can compute the distribution of spread of the <strong>4</strong> samples, like how much each of them <strong>covers the area</strong>, so that we can get some weights that bounds to 0 to 1 accordingly. </p>\n<p>Or we can use <strong>Dirichlet</strong> distribution, from it we can get <strong>4</strong> distributed values. Maybe then we can compute as follows:</p>\n<pre><code>image_list = [4 images]\nlabel_list = [4 label]\nnew_img = np.zeros((w, h))\n\nbeta_list = np.random.dirichlet((1, 1, 1, 1), 1)[0] \n# array([[0.59482542, 0.0185333 , 0.33322484, 0.05341645]])\n\nfor idx, beta in enumerate(beta_list):\n    x0, y0, w, h = get_crop_params(beta, full_img)  # something like this\n    new_img[x0, y0, w, h] = image_list[idx][x0, y0, w, h]\n    label_list[idx] = label_list[idx] * beta\n</code></pre>",
          "rawMarkdown": "That actually my first assumption. Let's say if we can compute the distribution of spread of the **4** samples, like how much each of them **covers the area**, so that we can get some weights that bounds to 0 to 1 accordingly. \n\nOr we can use **Dirichlet** distribution, from it we can get **4** distributed values. Maybe then we can compute as follows:\n\n```\nimage_list = [4 images]\nlabel_list = [4 label]\nnew_img = np.zeros((w, h))\n\nbeta_list = np.random.dirichlet((1, 1, 1, 1), 1)[0] \n# array([[0.59482542, 0.0185333 , 0.33322484, 0.05341645]])\n\nfor idx, beta in enumerate(beta_list):\n    x0, y0, w, h = get_crop_params(beta, full_img)  # something like this\n    new_img[x0, y0, w, h] = image_list[idx][x0, y0, w, h]\n    label_list[idx] = label_list[idx] * beta\n```",
          "votes": 1
        },
        {
          "id": 1108567,
          "postDate": "2020-12-10T19:44:40.227Z",
          "content": "<p>How do you choose each mosaic image size? Is it just random? Because if you could make them all equal you could choose balanced weights for the labels</p>",
          "rawMarkdown": "How do you choose each mosaic image size? Is it just random? Because if you could make them all equal you could choose balanced weights for the labels"
        },
        {
          "id": 1108568,
          "postDate": "2020-12-10T19:45:12.037Z",
          "content": "<p>I see you first create the weights and then crop or resize the images according to the weights.</p>",
          "rawMarkdown": "I see you first create the weights and then crop or resize the images according to the weights."
        },
        {
          "id": 1166200,
          "postDate": "2021-01-23T12:59:38.127Z",
          "content": "<p>Another way to look at this problem is by considering the lines of separation for both the width and height dimensions. When building the mosaic image, the goal is to combine 4 images into a single image. We can achieve this by randomly sampling midpoints (denoting the points of separation) in each dimension. This removes the rather complicated requirement of sampling 4 numbers summing up to 1. Instead the goal now is to sample 2 independent values from a uniform distribution - a much simpler and more intuitive alternative.</p>\n<p>So essentially, we sample two values:<br>\n<code>w = np.random.uniform(0, 1)</code><br>\n<code>h = np.random.uniform(0, 1)</code></p>\n<p><em>To generate realistic mosaics where each image has a noticeable contribution, we can sample values from [0.25 0.75], rather than from [0, 1]</em></p>\n<p>These two values are sufficient to parameterise the mosaic problem. Each image in the mosaic occupies areas spanned by the following coordinates: Consider that the mosaic image has dimensions WxH and the midpoints of each dimension is represented by w and h respectively.</p>\n<p>top left - (0, 0) to (w, h)<br>\ntop right - (w, 0) to (W, h)<br>\nbottom left - (0, h) to (w, H)<br>\nbottom right - (w, h) to (W, H)<br>\nThe sampled midpoints also help in calculating the class labels. Let's suppose we decide to use the area each image occupies within the mosaic as its corresponding contribution to the overall class label. For e.g Consider 4 images belonging to 4 classes {0, 1, 2, 3}. Now assume that the '0' image occupies the top left, '1' the top right, '2' the bottom left and '3' the bottom right. We can build the class label 'L' as follows<br>\n<a href=\"https://i.stack.imgur.com/h3w0X.png\" target=\"_blank\">you can view the equation at this link</a></p>",
          "rawMarkdown": "Another way to look at this problem is by considering the lines of separation for both the width and height dimensions. When building the mosaic image, the goal is to combine 4 images into a single image. We can achieve this by randomly sampling midpoints (denoting the points of separation) in each dimension. This removes the rather complicated requirement of sampling 4 numbers summing up to 1. Instead the goal now is to sample 2 independent values from a uniform distribution - a much simpler and more intuitive alternative.\n\nSo essentially, we sample two values:\n`w = np.random.uniform(0, 1)`\n`h = np.random.uniform(0, 1)`\n\n*To generate realistic mosaics where each image has a noticeable contribution, we can sample values from [0.25 0.75], rather than from [0, 1]*\n\nThese two values are sufficient to parameterise the mosaic problem. Each image in the mosaic occupies areas spanned by the following coordinates: Consider that the mosaic image has dimensions WxH and the midpoints of each dimension is represented by w and h respectively.\n\ntop left - (0, 0) to (w, h)\ntop right - (w, 0) to (W, h)\nbottom left - (0, h) to (w, H)\nbottom right - (w, h) to (W, H)\nThe sampled midpoints also help in calculating the class labels. Let's suppose we decide to use the area each image occupies within the mosaic as its corresponding contribution to the overall class label. For e.g Consider 4 images belonging to 4 classes {0, 1, 2, 3}. Now assume that the '0' image occupies the top left, '1' the top right, '2' the bottom left and '3' the bottom right. We can build the class label 'L' as follows\n[you can view the equation at this link](https://i.stack.imgur.com/h3w0X.png)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1109518,
      "postDate": "2020-12-11T19:25:21.473Z",
      "content": "<p>The leaves already have multiple diseases and you are trying to get more diseases there 😛</p>",
      "rawMarkdown": "The leaves already have multiple diseases and you are trying to get more diseases there 😛",
      "replies": [
        {
          "id": 1109562,
          "postDate": "2020-12-11T20:34:33.713Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1108550,
      "author_name": "Tolga",
      "author_url": "",
      "post_date": "2020-12-10T19:24:21.080000",
      "content": "<p>How about area weighted labels?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1108560,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-12-10T19:37:32.567000",
          "content": "<p>That actually my first assumption. Let's say if we can compute the distribution of spread of the <strong>4</strong> samples, like how much each of them <strong>covers the area</strong>, so that we can get some weights that bounds to 0 to 1 accordingly. </p>\n<p>Or we can use <strong>Dirichlet</strong> distribution, from it we can get <strong>4</strong> distributed values. Maybe then we can compute as follows:</p>\n<pre><code>image_list = [4 images]\nlabel_list = [4 label]\nnew_img = np.zeros((w, h))\n\nbeta_list = np.random.dirichlet((1, 1, 1, 1), 1)[0] \n# array([[0.59482542, 0.0185333 , 0.33322484, 0.05341645]])\n\nfor idx, beta in enumerate(beta_list):\n    x0, y0, w, h = get_crop_params(beta, full_img)  # something like this\n    new_img[x0, y0, w, h] = image_list[idx][x0, y0, w, h]\n    label_list[idx] = label_list[idx] * beta\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108567,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-12-10T19:44:40.227000",
          "content": "<p>How do you choose each mosaic image size? Is it just random? Because if you could make them all equal you could choose balanced weights for the labels</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1108568,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-10T19:45:12.037000",
          "content": "<p>I see you first create the weights and then crop or resize the images according to the weights.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1166200,
          "author_name": "Varun Anand",
          "author_url": "",
          "post_date": "2021-01-23T12:59:38.127000",
          "content": "<p>Another way to look at this problem is by considering the lines of separation for both the width and height dimensions. When building the mosaic image, the goal is to combine 4 images into a single image. We can achieve this by randomly sampling midpoints (denoting the points of separation) in each dimension. This removes the rather complicated requirement of sampling 4 numbers summing up to 1. Instead the goal now is to sample 2 independent values from a uniform distribution - a much simpler and more intuitive alternative.</p>\n<p>So essentially, we sample two values:<br>\n<code>w = np.random.uniform(0, 1)</code><br>\n<code>h = np.random.uniform(0, 1)</code></p>\n<p><em>To generate realistic mosaics where each image has a noticeable contribution, we can sample values from [0.25 0.75], rather than from [0, 1]</em></p>\n<p>These two values are sufficient to parameterise the mosaic problem. Each image in the mosaic occupies areas spanned by the following coordinates: Consider that the mosaic image has dimensions WxH and the midpoints of each dimension is represented by w and h respectively.</p>\n<p>top left - (0, 0) to (w, h)<br>\ntop right - (w, 0) to (W, h)<br>\nbottom left - (0, h) to (w, H)<br>\nbottom right - (w, h) to (W, H)<br>\nThe sampled midpoints also help in calculating the class labels. Let's suppose we decide to use the area each image occupies within the mosaic as its corresponding contribution to the overall class label. For e.g Consider 4 images belonging to 4 classes {0, 1, 2, 3}. Now assume that the '0' image occupies the top left, '1' the top right, '2' the bottom left and '3' the bottom right. We can build the class label 'L' as follows<br>\n<a href=\"https://i.stack.imgur.com/h3w0X.png\" target=\"_blank\">you can view the equation at this link</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1109518,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2020-12-11T19:25:21.473000",
      "content": "<p>The leaves already have multiple diseases and you are trying to get more diseases there 😛</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1109562,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-11T20:34:33.713000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1108532": "## Update \n\nThis is now supported in `keras-cv`.  The creation of labels is discussed here, https://github.com/keras-team/keras-cv/issues/250#issuecomment-1130256698 . \n\nhttps://github.com/keras-team/keras-cv\n\n\n---\n\nI am trying to do **mosaic augmentation** which first introduced in the **Yolo-4** paper-work. In `cutmix` or `mixup` or `fmix` we can use `beta` to weight the labels such as \n\n```\nlabel = label_one*beta + (1-beta)*label_two\n```\n\nHowever, unlike `cutmix` and others, a total of **4** samples is needed to create mosaic augmentation.  Such as below: \n\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fb0530886377eb3d7f0c29401ba069ffa%2Fdownload%20(1).png?generation=1607625768667914&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F80e35122f0fe72cc19127d28bfd1b20f%2Fdownload%20(2).png?generation=1607625836730889&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2Fdfce6cf87b2125c98affa7574ed38b18%2Fdownload%20(3).png?generation=1607625855803714&alt=media)\n\nHowever, I'm not quite sure how to create a class label for this yet. One idea for that is, maybe we can use **Dirichlet distribution** which can give us **4** distributed values, such as \n\n```\n>> np.random.dirichlet((1, 1, 1, 1), 1) \narray([[0.38160767, 0.28880196, 0.29071801, 0.03887236]])\n```\n\nBut it brings more complexity to the current setup. You can find the starter code from [here](https://www.kaggle.com/ipythonx/tf-keras-sota-augmentation-in-sequence-generator?scriptVersionId=49012892). Any efficient way to achieve this? ",
    "1108550": "How about area weighted labels?",
    "1109518": "The leaves already have multiple diseases and you are trying to get more diseases there 😛"
  }
}