{
  "id": 290402,
  "title": "New Data Augmentation Method for Instance Segmentation",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/290402",
  "author_name": "Neo",
  "post_date": "2021-11-24T13:06:19.984000",
  "votes": 44,
  "comment_count": 23,
  "views": 0,
  "content": "<p><a href=\"https://arxiv.org/pdf/2012.07177v2.pdf\" target=\"_blank\">Simple Copy-Paste is a Strong Data Augmentation Method\nfor Instance Segmentation</a></p>\n<p><img src=\"https://i.imgur.com/S5xBDPg.png\" alt=\" \"><br>\n<img src=\"https://i.imgur.com/VBLFUdO.png\" alt=\" \"></p>\n<p>Edited::</p>\n<p>SimpleCopy Paste Augmentation Implementation in MMDetection<br>\n<a href=\"https://github.com/open-mmlab/mmdetection/pull/6282\" target=\"_blank\">https://github.com/open-mmlab/mmdetection/pull/6282</a><br>\n<code>\ntrain_pipeline = [\n    dict(type='SimpleCopyPaste', prob=0),\n    dict(\n        type='Resize',\n        img_scale=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),\n                   (1333, 768), (1333, 800)],\n        multiscale_mode='value',\n        keep_ratio=True),\n    dict(type='RandomCrop', crop_size=(1024, 1024)),\n    dict(type='RandomFlip', flip_ratio=0.5),\n    dict(type='Normalize', **img_norm_cfg),\n    dict(type='Pad', size_divisor=32),\n    dict(type='DefaultFormatBundle'),\n    dict(type='Collect', keys=['img', 'gt_bboxes', 'gt_labels', 'gt_masks'])\n]</code></p>\n<p>To compose transforms with copy-paste augmentation with Albu::</p>\n<p><a href=\"https://github.com/conradry/copy-paste-aug\" target=\"_blank\">https://github.com/conradry/copy-paste-aug</a></p>\n<p>`import albumentations as A<br>\nfrom albumentations.pytorch.transforms import ToTensorV2<br>\nfrom copy_paste import CopyPaste</p>\n<p>transform = A.Compose([<br>\n      A.RandomScale(scale_limit=(-0.9, 1), p=1), #LargeScaleJitter from scale of 0.1 to 2<br>\n      A.PadIfNeeded(256, 256, border_mode=0), #constant 0 border<br>\n      A.RandomCrop(256, 256),<br>\n      A.HorizontalFlip(p=0.5),<br>\n      CopyPaste(blend=True, sigma=1, pct_objects_paste=0.5, p=1)<br>\n    ], bbox_params=A.BboxParams(format=\"coco\")<br>\n)`</p>",
  "messages": [
    {
      "id": 1593998,
      "postDate": "2021-11-24T13:06:19.983Z",
      "content": "<p><a href=\"https://arxiv.org/pdf/2012.07177v2.pdf\" target=\"_blank\">Simple Copy-Paste is a Strong Data Augmentation Method\nfor Instance Segmentation</a></p>\n<p><img src=\"https://i.imgur.com/S5xBDPg.png\" alt=\" \"><br>\n<img src=\"https://i.imgur.com/VBLFUdO.png\" alt=\" \"></p>\n<p>Edited::</p>\n<p>SimpleCopy Paste Augmentation Implementation in MMDetection<br>\n<a href=\"https://github.com/open-mmlab/mmdetection/pull/6282\" target=\"_blank\">https://github.com/open-mmlab/mmdetection/pull/6282</a><br>\n<code>\ntrain_pipeline = [\n    dict(type='SimpleCopyPaste', prob=0),\n    dict(\n        type='Resize',\n        img_scale=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),\n                   (1333, 768), (1333, 800)],\n        multiscale_mode='value',\n        keep_ratio=True),\n    dict(type='RandomCrop', crop_size=(1024, 1024)),\n    dict(type='RandomFlip', flip_ratio=0.5),\n    dict(type='Normalize', **img_norm_cfg),\n    dict(type='Pad', size_divisor=32),\n    dict(type='DefaultFormatBundle'),\n    dict(type='Collect', keys=['img', 'gt_bboxes', 'gt_labels', 'gt_masks'])\n]</code></p>\n<p>To compose transforms with copy-paste augmentation with Albu::</p>\n<p><a href=\"https://github.com/conradry/copy-paste-aug\" target=\"_blank\">https://github.com/conradry/copy-paste-aug</a></p>\n<p>`import albumentations as A<br>\nfrom albumentations.pytorch.transforms import ToTensorV2<br>\nfrom copy_paste import CopyPaste</p>\n<p>transform = A.Compose([<br>\n      A.RandomScale(scale_limit=(-0.9, 1), p=1), #LargeScaleJitter from scale of 0.1 to 2<br>\n      A.PadIfNeeded(256, 256, border_mode=0), #constant 0 border<br>\n      A.RandomCrop(256, 256),<br>\n      A.HorizontalFlip(p=0.5),<br>\n      CopyPaste(blend=True, sigma=1, pct_objects_paste=0.5, p=1)<br>\n    ], bbox_params=A.BboxParams(format=\"coco\")<br>\n)`</p>",
      "rawMarkdown": "[Simple Copy-Paste is a Strong Data Augmentation Method\nfor Instance Segmentation](https://arxiv.org/pdf/2012.07177v2.pdf)\n\n\n![ ](https://i.imgur.com/S5xBDPg.png)\n![ ] (https://i.imgur.com/VBLFUdO.png)\n\nEdited::\n\nSimpleCopy Paste Augmentation Implementation in MMDetection\nhttps://github.com/open-mmlab/mmdetection/pull/6282\n`\ntrain_pipeline = [\n    dict(type='SimpleCopyPaste', prob=0),\n    dict(\n        type='Resize',\n        img_scale=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),\n                   (1333, 768), (1333, 800)],\n        multiscale_mode='value',\n        keep_ratio=True),\n    dict(type='RandomCrop', crop_size=(1024, 1024)),\n    dict(type='RandomFlip', flip_ratio=0.5),\n    dict(type='Normalize', **img_norm_cfg),\n    dict(type='Pad', size_divisor=32),\n    dict(type='DefaultFormatBundle'),\n    dict(type='Collect', keys=['img', 'gt_bboxes', 'gt_labels', 'gt_masks'])\n]`\n\n\nTo compose transforms with copy-paste augmentation with Albu::\n\nhttps://github.com/conradry/copy-paste-aug\n\n`import albumentations as A\nfrom albumentations.pytorch.transforms import ToTensorV2\nfrom copy_paste import CopyPaste\n\ntransform = A.Compose([\n      A.RandomScale(scale_limit=(-0.9, 1), p=1), #LargeScaleJitter from scale of 0.1 to 2\n      A.PadIfNeeded(256, 256, border_mode=0), #constant 0 border\n      A.RandomCrop(256, 256),\n      A.HorizontalFlip(p=0.5),\n      CopyPaste(blend=True, sigma=1, pct_objects_paste=0.5, p=1)\n    ], bbox_params=A.BboxParams(format=\"coco\")\n)`",
      "votes": 43
    },
    {
      "id": 1601125,
      "postDate": "2021-12-01T03:23:05.697Z",
      "content": "<p>I actually tried copy-paste augmentation (I'm using Detectron2, it is quite easy to implement). I see a slight improvement, but not that much.</p>",
      "rawMarkdown": "I actually tried copy-paste augmentation (I'm using Detectron2, it is quite easy to implement). I see a slight improvement, but not that much.",
      "votes": 4,
      "replies": [
        {
          "id": 1601138,
          "postDate": "2021-12-01T03:54:23.010Z",
          "content": "<p>How do you integrate Simple Copy Paste with detectron2 ?</p>",
          "rawMarkdown": "How do you integrate Simple Copy Paste with detectron2 ?",
          "votes": 1
        },
        {
          "id": 1601162,
          "postDate": "2021-12-01T04:30:43.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> I create a custom mapper as in this documentation, and call copy-paste augmentation from that mapper: <a href=\"https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\" target=\"_blank\">https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html</a></p>\n<p>Here is the full code of my copy-paste augmentator using detectron2 datasets:<br>\n<a href=\"https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c\" target=\"_blank\">https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c</a></p>\n<p>The constructor takes the following parameters:</p>\n<ul>\n<li><code>d2_dataset</code> - basically a detectron2 dataset (returned by <code>DatasetCatalog.get(name)</code>. I set it to my <code>sartorius</code> train set (without <code>LiveCell</code> samples)</li>\n<li><code>paste_same_class</code> flag if set to <code>True</code> will only paste from a random sample of the same class.</li>\n<li><code>paste_density</code> dictates what portion of cells in that random sample we chose will be copied over to the sample that we are augmenting. I usually set it to no more than <code>0.4</code>. You can set a range.</li>\n<li><code>filter_area_thresh</code> basically says \"if after copy-pasting some cells will have area less than that amount of their original size, we remove the annotations for that cell\".</li>\n<li><code>p</code> is the probability of applying copy-paste augmentation.</li>\n</ul>\n<p>The <code>__call__</code> method is quite straightforward. It just takes a dataset sample (again, in Detectron2 format) and applies augmentation to it. Notice I don't apply augmentations to samples in <code>LiveCell</code> collection.</p>",
          "rawMarkdown": "@namgalielei I create a custom mapper as in this documentation, and call copy-paste augmentation from that mapper: https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\n\nHere is the full code of my copy-paste augmentator using detectron2 datasets:\nhttps://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c\n\nThe constructor takes the following parameters:\n- `d2_dataset` - basically a detectron2 dataset (returned by `DatasetCatalog.get(name)`. I set it to my `sartorius` train set (without `LiveCell` samples)\n- `paste_same_class` flag if set to `True` will only paste from a random sample of the same class.\n- `paste_density` dictates what portion of cells in that random sample we chose will be copied over to the sample that we are augmenting. I usually set it to no more than `0.4`. You can set a range.\n- `filter_area_thresh` basically says \"if after copy-pasting some cells will have area less than that amount of their original size, we remove the annotations for that cell\".\n- `p` is the probability of applying copy-paste augmentation.\n\nThe `__call__` method is quite straightforward. It just takes a dataset sample (again, in Detectron2 format) and applies augmentation to it. Notice I don't apply augmentations to samples in `LiveCell` collection.",
          "votes": 16
        },
        {
          "id": 1601348,
          "postDate": "2021-12-01T07:28:46.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a>  thanks for your works. You have a good engineering skill.</p>",
          "rawMarkdown": "@chankhavu  thanks for your works. You have a good engineering skill.",
          "votes": 3
        },
        {
          "id": 1601761,
          "postDate": "2021-12-01T15:06:23.210Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1601919,
          "postDate": "2021-12-01T16:38:05.810Z",
          "content": "<p>self.cls_indices = [<br>\n[<br>\nI for I, item in enumerate(d2_dataset)<br>\nIf the item [' annotations'] [0] [' category_id] = = cls_index<br>\n]<br>\nFor cls_index in range (3)<br>\n]<br>\nWhy the error occurs:<br>\nTypeError: string indices must be integers, not str</p>",
          "rawMarkdown": "self.cls_indices = [\n[\nI for I, item in enumerate(d2_dataset)\nIf the item [' annotations'] [0] [' category_id] = = cls_index\n]\nFor cls_index in range (3)\n]\nWhy the error occurs:\nTypeError: string indices must be integers, not str"
        },
        {
          "id": 1601924,
          "postDate": "2021-12-01T16:43:04.437Z",
          "content": "<p><code>for i, item in enumerate(d2_dataset)</code><br>\nI found that the contents of the loop were just keys in the dict. Is there a problem with the code?</p>",
          "rawMarkdown": "`for i, item in enumerate(d2_dataset)`\nI found that the contents of the loop were just keys in the dict. Is there a problem with the code?"
        },
        {
          "id": 1602069,
          "postDate": "2021-12-01T18:31:22.660Z",
          "content": "<p><a href=\"https://www.kaggle.com/joker\" target=\"_blank\">@joker</a> <code>d2_dataset</code> variable is supposed to be a list of dicts created using <code>register_coco_instances</code> as described in the documentation: <a href=\"https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html\" target=\"_blank\">https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html</a></p>\n<p>This notebook by <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> is the ideal example of how to read the Sartorius dataset in that format (in that notebook, it is <code>train_ds</code> variable): <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training\" target=\"_blank\">https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training</a></p>",
          "rawMarkdown": "@joker `d2_dataset` variable is supposed to be a list of dicts created using `register_coco_instances` as described in the documentation: https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html\n\nThis notebook by @slawekbiel is the ideal example of how to read the Sartorius dataset in that format (in that notebook, it is `train_ds` variable): https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training"
        },
        {
          "id": 1602426,
          "postDate": "2021-12-02T00:33:43.897Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1602521,
          "postDate": "2021-12-02T01:50:46.677Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1602573,
          "postDate": "2021-12-02T02:38:50.017Z",
          "content": "<p>I want to learn how your Mapper is defined?</p>",
          "rawMarkdown": "I want to learn how your Mapper is defined?"
        },
        {
          "id": 1603986,
          "postDate": "2021-12-03T02:06:37.023Z",
          "content": "<p>As a newbie, I'd also like to know how your Mapper is defined and able to run successfully.</p>",
          "rawMarkdown": "As a newbie, I'd also like to know how your Mapper is defined and able to run successfully."
        },
        {
          "id": 1603987,
          "postDate": "2021-12-03T02:12:13.630Z",
          "content": "<p>Hello, would you share your defined Mapper, looking for your reply</p>",
          "rawMarkdown": "Hello, would you share your defined Mapper, looking for your reply"
        },
        {
          "id": 1603996,
          "postDate": "2021-12-03T02:30:52.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/zaopolearning\" target=\"_blank\">@zaopolearning</a> hey, I added the mapper in this gist: <a href=\"https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c\" target=\"_blank\">https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c</a> - basically almost the same as the mapper in detectron documentation: <a href=\"https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\" target=\"_blank\">https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html</a></p>\n<p>(as you can see I use detectron2's transforms for geometric transforms like resizing and cropping, and albumentation for enhancing the image but…… augmentations doesn't seem to work here)</p>\n<p>sorry for messy code lol :) </p>",
          "rawMarkdown": "@zaopolearning hey, I added the mapper in this gist: https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c - basically almost the same as the mapper in detectron documentation: https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\n\n(as you can see I use detectron2's transforms for geometric transforms like resizing and cropping, and albumentation for enhancing the image but...... augmentations doesn't seem to work here)\n\nsorry for messy code lol :) "
        },
        {
          "id": 1604006,
          "postDate": "2021-12-03T02:57:01.307Z",
          "content": "<p>thank you!</p>",
          "rawMarkdown": "thank you!"
        },
        {
          "id": 1604012,
          "postDate": "2021-12-03T03:12:17.633Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1606219,
          "postDate": "2021-12-04T22:37:35.390Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a> ,thankyou for sharing this<br>\nI was trying to implement this and am getting the error in this part of the code:<br>\n<code>orig_img = cv2.imread(dataset_dict[\"file_name\"])</code></p>\n<p>TypeError: list indices must be integers or slices, not str<br>\nMy dataset_dict is also a list of dicts (as you gave example of train_ds).<br>\nI don't know how to resolve this.</p>",
          "rawMarkdown": "Hey @chankhavu ,thankyou for sharing this\nI was trying to implement this and am getting the error in this part of the code:\n` orig_img = cv2.imread(dataset_dict[\"file_name\"])`\n\nTypeError: list indices must be integers or slices, not str\nMy dataset_dict is also a list of dicts (as you gave example of train_ds).\nI don't know how to resolve this."
        },
        {
          "id": 1729420,
          "postDate": "2022-03-20T03:56:25.040Z",
          "content": "<p>It doesn't seem to work when my images are different sizes.</p>",
          "rawMarkdown": "It doesn't seem to work when my images are different sizes."
        }
      ]
    },
    {
      "id": 1607980,
      "postDate": "2021-12-06T07:01:15.433Z",
      "content": "<p>does anyone tried this method in mmdetection?</p>",
      "rawMarkdown": "does anyone tried this method in mmdetection?",
      "votes": 1
    },
    {
      "id": 1601005,
      "postDate": "2021-11-30T23:18:55.440Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mlneo07\" target=\"_blank\">@mlneo07</a>. Were you able to try this out? Looks like mmdet hasnt merged this feature yet</p>",
      "rawMarkdown": "Thanks @mlneo07. Were you able to try this out? Looks like mmdet hasnt merged this feature yet",
      "votes": 1,
      "replies": [
        {
          "id": 1601346,
          "postDate": "2021-12-01T07:23:05.787Z",
          "content": "<p>Yes, Slight improvement. I am using albumentations.  MMDet is not merged yet but we can try <a href=\"https://mmdetection.readthedocs.io/en/latest/tutorials/data_pipeline.html\" target=\"_blank\">@PIPELINES.register_module()</a> for cusotmize dataset. </p>",
          "rawMarkdown": "Yes, Slight improvement. I am using albumentations.  MMDet is not merged yet but we can try [@PIPELINES.register_module()](https://mmdetection.readthedocs.io/en/latest/tutorials/data_pipeline.html) for cusotmize dataset. ",
          "votes": 1
        },
        {
          "id": 1601374,
          "postDate": "2021-12-01T08:12:38.030Z",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/mlneo07\" target=\"_blank\">@mlneo07</a> . Do you have an example of the custom pipeline file? I tried adding the decorator to the <code>CopyPaste</code> class but I think that is not the right approach. </p>",
          "rawMarkdown": "thanks @mlneo07 . Do you have an example of the custom pipeline file? I tried adding the decorator to the `CopyPaste` class but I think that is not the right approach. "
        },
        {
          "id": 1677996,
          "postDate": "2022-02-06T07:49:33.973Z",
          "content": "<p>Do you use copy-paste in MMDetection?</p>",
          "rawMarkdown": "Do you use copy-paste in MMDetection?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1601125,
      "author_name": "Chan Kha Vu",
      "author_url": "",
      "post_date": "2021-12-01T03:23:05.697000",
      "content": "<p>I actually tried copy-paste augmentation (I'm using Detectron2, it is quite easy to implement). I see a slight improvement, but not that much.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1601138,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-12-01T03:54:23.010000",
          "content": "<p>How do you integrate Simple Copy Paste with detectron2 ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1601162,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2021-12-01T04:30:43.357000",
          "content": "<p><a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> I create a custom mapper as in this documentation, and call copy-paste augmentation from that mapper: <a href=\"https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\" target=\"_blank\">https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html</a></p>\n<p>Here is the full code of my copy-paste augmentator using detectron2 datasets:<br>\n<a href=\"https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c\" target=\"_blank\">https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c</a></p>\n<p>The constructor takes the following parameters:</p>\n<ul>\n<li><code>d2_dataset</code> - basically a detectron2 dataset (returned by <code>DatasetCatalog.get(name)</code>. I set it to my <code>sartorius</code> train set (without <code>LiveCell</code> samples)</li>\n<li><code>paste_same_class</code> flag if set to <code>True</code> will only paste from a random sample of the same class.</li>\n<li><code>paste_density</code> dictates what portion of cells in that random sample we chose will be copied over to the sample that we are augmenting. I usually set it to no more than <code>0.4</code>. You can set a range.</li>\n<li><code>filter_area_thresh</code> basically says \"if after copy-pasting some cells will have area less than that amount of their original size, we remove the annotations for that cell\".</li>\n<li><code>p</code> is the probability of applying copy-paste augmentation.</li>\n</ul>\n<p>The <code>__call__</code> method is quite straightforward. It just takes a dataset sample (again, in Detectron2 format) and applies augmentation to it. Notice I don't apply augmentations to samples in <code>LiveCell</code> collection.</p>",
          "votes": 16,
          "replies": []
        },
        {
          "id": 1601348,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-12-01T07:28:46.137000",
          "content": "<p><a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a>  thanks for your works. You have a good engineering skill.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1601761,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-12-01T15:06:23.210000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1601919,
          "author_name": "yerongg",
          "author_url": "",
          "post_date": "2021-12-01T16:38:05.810000",
          "content": "<p>self.cls_indices = [<br>\n[<br>\nI for I, item in enumerate(d2_dataset)<br>\nIf the item [' annotations'] [0] [' category_id] = = cls_index<br>\n]<br>\nFor cls_index in range (3)<br>\n]<br>\nWhy the error occurs:<br>\nTypeError: string indices must be integers, not str</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1601924,
          "author_name": "yerongg",
          "author_url": "",
          "post_date": "2021-12-01T16:43:04.437000",
          "content": "<p><code>for i, item in enumerate(d2_dataset)</code><br>\nI found that the contents of the loop were just keys in the dict. Is there a problem with the code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1602069,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2021-12-01T18:31:22.660000",
          "content": "<p><a href=\"https://www.kaggle.com/joker\" target=\"_blank\">@joker</a> <code>d2_dataset</code> variable is supposed to be a list of dicts created using <code>register_coco_instances</code> as described in the documentation: <a href=\"https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html\" target=\"_blank\">https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html</a></p>\n<p>This notebook by <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> is the ideal example of how to read the Sartorius dataset in that format (in that notebook, it is <code>train_ds</code> variable): <a href=\"https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training\" target=\"_blank\">https://www.kaggle.com/slawekbiel/positive-score-with-detectron-2-3-training</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1602426,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-12-02T00:33:43.897000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1602521,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-12-02T01:50:46.677000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1602573,
          "author_name": "yerongg",
          "author_url": "",
          "post_date": "2021-12-02T02:38:50.017000",
          "content": "<p>I want to learn how your Mapper is defined?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1603986,
          "author_name": "iwrbeverly",
          "author_url": "",
          "post_date": "2021-12-03T02:06:37.023000",
          "content": "<p>As a newbie, I'd also like to know how your Mapper is defined and able to run successfully.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1603987,
          "author_name": "sanmedbio.huajia",
          "author_url": "",
          "post_date": "2021-12-03T02:12:13.630000",
          "content": "<p>Hello, would you share your defined Mapper, looking for your reply</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1603996,
          "author_name": "Chan Kha Vu",
          "author_url": "",
          "post_date": "2021-12-03T02:30:52.027000",
          "content": "<p><a href=\"https://www.kaggle.com/zaopolearning\" target=\"_blank\">@zaopolearning</a> hey, I added the mapper in this gist: <a href=\"https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c\" target=\"_blank\">https://gist.github.com/hav4ik/2d80b63dcab69651ad7bf1495053e35c</a> - basically almost the same as the mapper in detectron documentation: <a href=\"https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html\" target=\"_blank\">https://detectron2.readthedocs.io/en/latest/tutorials/data_loading.html</a></p>\n<p>(as you can see I use detectron2's transforms for geometric transforms like resizing and cropping, and albumentation for enhancing the image but…… augmentations doesn't seem to work here)</p>\n<p>sorry for messy code lol :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1604006,
          "author_name": "yerongg",
          "author_url": "",
          "post_date": "2021-12-03T02:57:01.307000",
          "content": "<p>thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1604012,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-12-03T03:12:17.633000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1606219,
          "author_name": "Shreya Sajal",
          "author_url": "",
          "post_date": "2021-12-04T22:37:35.390000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/chankhavu\" target=\"_blank\">@chankhavu</a> ,thankyou for sharing this<br>\nI was trying to implement this and am getting the error in this part of the code:<br>\n<code>orig_img = cv2.imread(dataset_dict[\"file_name\"])</code></p>\n<p>TypeError: list indices must be integers or slices, not str<br>\nMy dataset_dict is also a list of dicts (as you gave example of train_ds).<br>\nI don't know how to resolve this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1729420,
          "author_name": "yerongg",
          "author_url": "",
          "post_date": "2022-03-20T03:56:25.040000",
          "content": "<p>It doesn't seem to work when my images are different sizes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1607980,
      "author_name": "shinning2021",
      "author_url": "",
      "post_date": "2021-12-06T07:01:15.433000",
      "content": "<p>does anyone tried this method in mmdetection?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1601005,
      "author_name": "Trushant Kalyanpur",
      "author_url": "",
      "post_date": "2021-11-30T23:18:55.440000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mlneo07\" target=\"_blank\">@mlneo07</a>. Were you able to try this out? Looks like mmdet hasnt merged this feature yet</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1601346,
          "author_name": "Neo",
          "author_url": "",
          "post_date": "2021-12-01T07:23:05.787000",
          "content": "<p>Yes, Slight improvement. I am using albumentations.  MMDet is not merged yet but we can try <a href=\"https://mmdetection.readthedocs.io/en/latest/tutorials/data_pipeline.html\" target=\"_blank\">@PIPELINES.register_module()</a> for cusotmize dataset. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1601374,
          "author_name": "Trushant Kalyanpur",
          "author_url": "",
          "post_date": "2021-12-01T08:12:38.030000",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/mlneo07\" target=\"_blank\">@mlneo07</a> . Do you have an example of the custom pipeline file? I tried adding the decorator to the <code>CopyPaste</code> class but I think that is not the right approach. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1677996,
          "author_name": "yerongg",
          "author_url": "",
          "post_date": "2022-02-06T07:49:33.973000",
          "content": "<p>Do you use copy-paste in MMDetection?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1593998": "[Simple Copy-Paste is a Strong Data Augmentation Method\nfor Instance Segmentation](https://arxiv.org/pdf/2012.07177v2.pdf)\n\n\n![ ](https://i.imgur.com/S5xBDPg.png)\n![ ] (https://i.imgur.com/VBLFUdO.png)\n\nEdited::\n\nSimpleCopy Paste Augmentation Implementation in MMDetection\nhttps://github.com/open-mmlab/mmdetection/pull/6282\n`\ntrain_pipeline = [\n    dict(type='SimpleCopyPaste', prob=0),\n    dict(\n        type='Resize',\n        img_scale=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),\n                   (1333, 768), (1333, 800)],\n        multiscale_mode='value',\n        keep_ratio=True),\n    dict(type='RandomCrop', crop_size=(1024, 1024)),\n    dict(type='RandomFlip', flip_ratio=0.5),\n    dict(type='Normalize', **img_norm_cfg),\n    dict(type='Pad', size_divisor=32),\n    dict(type='DefaultFormatBundle'),\n    dict(type='Collect', keys=['img', 'gt_bboxes', 'gt_labels', 'gt_masks'])\n]`\n\n\nTo compose transforms with copy-paste augmentation with Albu::\n\nhttps://github.com/conradry/copy-paste-aug\n\n`import albumentations as A\nfrom albumentations.pytorch.transforms import ToTensorV2\nfrom copy_paste import CopyPaste\n\ntransform = A.Compose([\n      A.RandomScale(scale_limit=(-0.9, 1), p=1), #LargeScaleJitter from scale of 0.1 to 2\n      A.PadIfNeeded(256, 256, border_mode=0), #constant 0 border\n      A.RandomCrop(256, 256),\n      A.HorizontalFlip(p=0.5),\n      CopyPaste(blend=True, sigma=1, pct_objects_paste=0.5, p=1)\n    ], bbox_params=A.BboxParams(format=\"coco\")\n)`",
    "1601125": "I actually tried copy-paste augmentation (I'm using Detectron2, it is quite easy to implement). I see a slight improvement, but not that much.",
    "1607980": "does anyone tried this method in mmdetection?",
    "1601005": "Thanks @mlneo07. Were you able to try this out? Looks like mmdet hasnt merged this feature yet"
  }
}