{
  "id": 297986,
  "title": "6th place solution. Higher resolution is all you need.",
  "url": "/competitions/sartorius-cell-instance-segmentation/writeups/carno-zhao-6th-place-solution-higher-resolution-is",
  "author_name": "",
  "post_date": "2021-12-31T02:24:31.463Z",
  "votes": 72,
  "comment_count": 56,
  "views": 0,
  "content": "<h1>Preprocessing</h1>\n<p>I only used <code>train</code> images, without semi-supervised or LiveCell.</p>\n<p>All train images were sliced into smaller size, with window size (208, 281), and stride (104, 140). So a (520, 704) image can be slices to 16 smaller images.</p>\n<h1>Training</h1>\n<h2>Model</h2>\n<p>It seems that model does not really matter in private LB. My res2net-101, resnext-101, detectoRS-r50 can get 0.350 in private LB, even though their public LB are very different.</p>\n<p>Ensemble cannot improve both public and private LB significantly. For me, single model is enough.</p>\n<h2>Hyperparams</h2>\n<p>I trained with mmdet, default 1x schedule and 1x swa training. Image scales were set as (1333, 1333)-(800,800).</p>\n<h1>Inferencing</h1>\n<p>Test images were also sliced as training images. The test image scales were [(1333,1333), (1024,1024), (800,800)].</p>\n<p>Masks were iterated from higher score to lower score. Scores lower than class-wise threshold and areas lower than class-wise pixel-threshold were removed. </p>\n<p>When dealing with overlaps, I remove mask whose overlapped part was more than 20% of itself.</p>\n<p>RCNN's NMS was replaced by Weighted cluster-NMS with DIoU.</p>\n<h1>Submission</h1>\n<p>My final submission was ensemble of HTC-Res2Net101 (trained with all data), HTC-ResNeXt101 (trained with all data) and HTC-Res2Net101 (trained with fold0)</p>\n<h1>Weakness</h1>\n<p>LONG inferencing time, 2h per model.</p>\n<h1>Github Repo</h1>\n<p><a href=\"https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution\" target=\"_blank\">https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution</a></p>",
  "messages": [
    {
      "id": "1633626",
      "postDate": "12/31/2021 00:27:39",
      "content": "<h1>Preprocessing</h1>\n<p>I only used <code>train</code> images, without semi-supervised or LiveCell.</p>\n<p>All train images were sliced into smaller size, with window size (208, 281), and stride (104, 140). So a (520, 704) image can be slices to 16 smaller images.</p>\n<h1>Training</h1>\n<h2>Model</h2>\n<p>It seems that model does not really matter in private LB. My res2net-101, resnext-101, detectoRS-r50 can get 0.350 in private LB, even though their public LB are very different.</p>\n<p>Ensemble cannot improve both public and private LB significantly. For me, single model is enough.</p>\n<h2>Hyperparams</h2>\n<p>I trained with mmdet, default 1x schedule and 1x swa training. Image scales were set as (1333, 1333)-(800,800).</p>\n<h1>Inferencing</h1>\n<p>Test images were also sliced as training images. The test image scales were [(1333,1333), (1024,1024), (800,800)].</p>\n<p>Masks were iterated from higher score to lower score. Scores lower than class-wise threshold and areas lower than class-wise pixel-threshold were removed. </p>\n<p>When dealing with overlaps, I remove mask whose overlapped part was more than 20% of itself.</p>\n<p>RCNN's NMS was replaced by Weighted cluster-NMS with DIoU.</p>\n<h1>Submission</h1>\n<p>My final submission was ensemble of HTC-Res2Net101 (trained with all data), HTC-ResNeXt101 (trained with all data) and HTC-Res2Net101 (trained with fold0)</p>\n<h1>Weakness</h1>\n<p>LONG inferencing time, 2h per model.</p>\n<h1>Github Repo</h1>\n<p><a href=\"https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution\" target=\"_blank\">https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution</a></p>",
      "rawMarkdown": "# Preprocessing\n\nI only used `train` images, without semi-supervised or LiveCell.\n\nAll train images were sliced into smaller size, with window size (208, 281), and stride (104, 140). So a (520, 704) image can be slices to 16 smaller images.\n\n\n\n# Training\n\n## Model\n\nIt seems that model does not really matter in private LB. My res2net-101, resnext-101, detectoRS-r50 can get 0.350 in private LB, even though their public LB are very different.\n\n\nEnsemble cannot improve both public and private LB significantly. For me, single model is enough.\n\n\n## Hyperparams\n\nI trained with mmdet, default 1x schedule and 1x swa training. Image scales were set as (1333, 1333)-(800,800).\n\n\n\n# Inferencing\n\nTest images were also sliced as training images. The test image scales were [(1333,1333), (1024,1024), (800,800)].\n\nMasks were iterated from higher score to lower score. Scores lower than class-wise threshold and areas lower than class-wise pixel-threshold were removed. \n\n\n\nWhen dealing with overlaps, I remove mask whose overlapped part was more than 20% of itself.\n\n\n\nRCNN's NMS was replaced by Weighted cluster-NMS with DIoU.\n\n# Submission\n\nMy final submission was ensemble of HTC-Res2Net101 (trained with all data), HTC-ResNeXt101 (trained with all data) and HTC-Res2Net101 (trained with fold0)\n\n# Weakness\n\nLONG inferencing time, 2h per model.\n\n# Github Repo\n\n[https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution](https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution)",
      "votes": null
    },
    {
      "id": "1633628",
      "postDate": "12/31/2021 00:31:55",
      "content": "<p>Congrats! Very neat and impressive solution, Carno Zhao is all we need: )</p>",
      "rawMarkdown": "Congrats! Very neat and impressive solution, Carno Zhao is all we need: )",
      "votes": null
    },
    {
      "id": "1633634",
      "postDate": "12/31/2021 00:36:21",
      "content": "<p>Congrats! One question about the cropping during training, how much of a boost did you get from doing that? From what I saw and tried, random cropping hurts performance, so this felt a bit odd to me.</p>",
      "rawMarkdown": "Congrats! One question about the cropping during training, how much of a boost did you get from doing that? From what I saw and tried, random cropping hurts performance, so this felt a bit odd to me.",
      "votes": null
    },
    {
      "id": "1633638",
      "postDate": "12/31/2021 00:39:30",
      "content": "<p>When using original images, my public LB was 0.320<br>\nWhen slicing images to 3x3=9 smaller overlapped images, my public LB was 0.331<br>\nWhen slicing images to 4x4=16 smaller overlapped images, my public LB was 0.334</p>\n<p>From 0.334 to 0.340 public LB, I used other tricks.</p>",
      "rawMarkdown": "When using original images, my public LB was 0.320\nWhen slicing images to 3x3=9 smaller overlapped images, my public LB was 0.331\nWhen slicing images to 4x4=16 smaller overlapped images, my public LB was 0.334\n\nFrom 0.334 to 0.340 public LB, I used other tricks.",
      "votes": null
    },
    {
      "id": "1633640",
      "postDate": "12/31/2021 00:41:32",
      "content": "<p>Very nice!!! congratulations !!!</p>",
      "rawMarkdown": "Very nice!!! congratulations !!!",
      "votes": null
    },
    {
      "id": "1633641",
      "postDate": "12/31/2021 00:41:48",
      "content": "<p>Thanks for the info!</p>",
      "rawMarkdown": "Thanks for the info!",
      "votes": null
    },
    {
      "id": "1633642",
      "postDate": "12/31/2021 00:42:23",
      "content": "<p>I think the key point is using higher resolution in inferencing. To do this, we need training models using higher resolution (Random crop or slicing as I did)</p>",
      "rawMarkdown": "I think the key point is using higher resolution in inferencing. To do this, we need training models using higher resolution (Random crop or slicing as I did)",
      "votes": null
    },
    {
      "id": "1633644",
      "postDate": "12/31/2021 00:47:19",
      "content": "<p>Unexpectedly,the offline slicing method of satellite image segmentation can also be applied to this competition！ congratulations!</p>",
      "rawMarkdown": "Unexpectedly,the offline slicing method of satellite image segmentation can also be applied to this competition！ congratulations!",
      "votes": null
    },
    {
      "id": "1633645",
      "postDate": "12/31/2021 00:47:20",
      "content": "<p>We found out about that too, though we did not train with higher resolution (Sadge). Initially we thought the performance boost was because of anchor size being too large, but for smaller anchor size model, higher res still give better performance, so we gave up trying to find out why and roll with it. Do you have a hypothesis on why this is the case?</p>",
      "rawMarkdown": "We found out about that too, though we did not train with higher resolution (Sadge). Initially we thought the performance boost was because of anchor size being too large, but for smaller anchor size model, higher res still give better performance, so we gave up trying to find out why and roll with it. Do you have a hypothesis on why this is the case?",
      "votes": null
    },
    {
      "id": "1633647",
      "postDate": "12/31/2021 00:50:59",
      "content": "<p>I also did some experiments about anchors. I think when using low resolution with small anchor size, the feature map downsampled from image is too small to provide useful information, especially in dense objects situation.</p>",
      "rawMarkdown": "I also did some experiments about anchors. I think when using low resolution with small anchor size, the feature map downsampled from image is too small to provide useful information, especially in dense objects situation.",
      "votes": null
    },
    {
      "id": "1633651",
      "postDate": "12/31/2021 00:54:06",
      "content": "<p>That make alot of sense, thank you!</p>",
      "rawMarkdown": "That make alot of sense, thank you!",
      "votes": null
    },
    {
      "id": "1633666",
      "postDate": "12/31/2021 01:12:44",
      "content": "<p>Congrats on solo gold medal <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, you don't you TTA?</p>",
      "rawMarkdown": "Congrats on solo gold medal @carnozhao, you don't you TTA?",
      "votes": null
    },
    {
      "id": "1633673",
      "postDate": "12/31/2021 01:16:23",
      "content": "<p>Yes I used Horizontal + Vertical flip, [(1333,1333), (1024,1024), (800,800)] multiscale test. All of them are provided in mmdet.</p>",
      "rawMarkdown": "Yes I used Horizontal + Vertical flip, [(1333,1333), (1024,1024), (800,800)] multiscale test. All of them are provided in mmdet.",
      "votes": null
    },
    {
      "id": "1633678",
      "postDate": "12/31/2021 01:19:58",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> for sharing, I also trained HTC but the performance not good. Can you share with us how did you train HTC-Res2Net101!</p>",
      "rawMarkdown": "Thanks @carnozhao for sharing, I also trained HTC but the performance not good. Can you share with us how did you train HTC-Res2Net101!",
      "votes": null
    },
    {
      "id": "1633680",
      "postDate": "12/31/2021 01:21:17",
      "content": "<p>I will update to my github repo soon!</p>",
      "rawMarkdown": "I will update to my github repo soon!",
      "votes": null
    },
    {
      "id": "1633691",
      "postDate": "12/31/2021 01:36:32",
      "content": "<p>Your solution is very good with you not using pseudo label and pretrain livecell , during training do you use augmentation, hyper parameters tuning or any special technique. I can't get such a high score with HTC</p>",
      "rawMarkdown": "Your solution is very good with you not using pseudo label and pretrain livecell , during training do you use augmentation, hyper parameters tuning or any special technique. I can't get such a high score with HTC",
      "votes": null
    },
    {
      "id": "1633707",
      "postDate": "12/31/2021 01:58:30",
      "content": "<p>I used H-flip, V-flip, Rot90 multiscale in training as augmentations. I didnt do any hyperparameter tuning because that's too time consuming. Again, the KEY points is smaller images :)</p>",
      "rawMarkdown": "I used H-flip, V-flip, Rot90 multiscale in training as augmentations. I didnt do any hyperparameter tuning because that's too time consuming. Again, the KEY points is smaller images :)",
      "votes": null
    },
    {
      "id": "1633708",
      "postDate": "12/31/2021 02:00:59",
      "content": "<p>Thanks for sharing and congrats.  </p>\n<p>How did you deal with the predicted masks where the test slices overlap?  </p>",
      "rawMarkdown": "Thanks for sharing and congrats.  \n\nHow did you deal with the predicted masks where the test slices overlap?",
      "votes": null
    },
    {
      "id": "1633724",
      "postDate": "12/31/2021 02:14:16",
      "content": "<p>github repo updated</p>",
      "rawMarkdown": "github repo updated",
      "votes": null
    },
    {
      "id": "1633726",
      "postDate": "12/31/2021 02:20:32",
      "content": "<p>Congratulations on your solo gold! Simple yet very effective! Great work!</p>",
      "rawMarkdown": "Congratulations on your solo gold! Simple yet very effective! Great work!",
      "votes": null
    },
    {
      "id": "1633733",
      "postDate": "12/31/2021 02:33:02",
      "content": "<p>Thanks for your explanation</p>",
      "rawMarkdown": "Thanks for your explanation",
      "votes": null
    },
    {
      "id": "1633740",
      "postDate": "12/31/2021 02:42:56",
      "content": "<p>Thank you, I see! Can you share the idea behind <code>sliced into smaller size</code>, why did you think that will work and tried?</p>",
      "rawMarkdown": "Thank you, I see! Can you share the idea behind `sliced into smaller size`, why did you think that will work and tried?",
      "votes": null
    },
    {
      "id": "1633741",
      "postDate": "12/31/2021 02:45:08",
      "content": "<p>Good question, thats also very important.<br>\nE.g, say slice A's x axis ranges from 0 to 100, and slice B's x axis ranges from 50 to 150. The middle point of overlap is 75, so I kept slice A's predictions whose box center ranges from 0-75, and slices B's predictions whose box center ranges from 75-150.</p>\n<p>This is only a simple exsample. You can apply it to x axis overlap, y axis overlap, and also multiple sequential overlaps.</p>\n<p>The detail of implementation can be found at <a href=\"https://www.kaggle.com/carnozhao/cell-submission?scriptVersionId=83795256&amp;cellId=3\" target=\"_blank\">my submission notebook</a></p>\n<pre><code>valid = (small_box[:,-1] &gt; THRESHOLDS_small[class_id]) &amp; \\\n            ~((i &lt;= 2) &amp; (small_box[:,[1,3]].mean(1) &gt; 3 * H / 10)) &amp; \\\n            ~((i &gt;= 1) &amp; (small_box[:,[1,3]].mean(1) &lt; H / 10)) &amp; \\\n            ~((j &lt;= 2) &amp; (small_box[:,[0,2]].mean(1) &gt; 3 * W / 10)) &amp; \\\n            ~((j &gt;= 1) &amp; (small_box[:,[0,2]].mean(1) &lt; W / 10))\n</code></pre>\n<p>Here, H/10, W/10, 3H/10, 3W/10 are overlap middle points of each small image.</p>",
      "rawMarkdown": "Good question, thats also very important.\nE.g, say slice A's x axis ranges from 0 to 100, and slice B's x axis ranges from 50 to 150. The middle point of overlap is 75, so I kept slice A's predictions whose box center ranges from 0-75, and slices B's predictions whose box center ranges from 75-150.\n\nThis is only a simple exsample. You can apply it to x axis overlap, y axis overlap, and also multiple sequential overlaps.\n\nThe detail of implementation can be found at [my submission notebook](https://www.kaggle.com/carnozhao/cell-submission?scriptVersionId=83795256&cellId=3)\n\n```python\nvalid = (small_box[:,-1] > THRESHOLDS_small[class_id]) & \\\n            ~((i <= 2) & (small_box[:,[1,3]].mean(1) > 3 * H / 10)) & \\\n            ~((i >= 1) & (small_box[:,[1,3]].mean(1) < H / 10)) & \\\n            ~((j <= 2) & (small_box[:,[0,2]].mean(1) > 3 * W / 10)) & \\\n            ~((j >= 1) & (small_box[:,[0,2]].mean(1) < W / 10))\n```\nHere, H/10, W/10, 3H/10, 3W/10 are overlap middle points of each small image.",
      "votes": null
    },
    {
      "id": "1633743",
      "postDate": "12/31/2021 02:49:06",
      "content": "<p>The biggest motivation is that the objects are small and densly distributed, which is different than other object detection dataset like COCO.  However, my GPU cannot train with really large images, so cutting them and resize cutted images to bigger one seems natural.</p>",
      "rawMarkdown": "The biggest motivation is that the objects are small and densly distributed, which is different than other object detection dataset like COCO.  However, my GPU cannot train with really large images, so cutting them and resize cutted images to bigger one seems natural.",
      "votes": null
    },
    {
      "id": "1633755",
      "postDate": "12/31/2021 03:09:49",
      "content": "<p>It's true that the size matters a lot. If you resize your result mask to a bigger size and test, the score can go high a little. If resize it to a smaller size it goes down a lot.I doubt really there is something inproper for the iou metric when the task target is all small objects.</p>",
      "rawMarkdown": "It's true that the size matters a lot. If you resize your result mask to a bigger size and test, the score can go high a little. If resize it to a smaller size it goes down a lot.I doubt really there is something inproper for the iou metric when the task target is all small objects.",
      "votes": null
    },
    {
      "id": "1633766",
      "postDate": "12/31/2021 03:33:09",
      "content": "<p>In my test, adding vertical flip to TTA did not improve the prediction.</p>",
      "rawMarkdown": "In my test, adding vertical flip to TTA did not improve the prediction.",
      "votes": null
    },
    {
      "id": "1633767",
      "postDate": "12/31/2021 03:42:06",
      "content": "<p>Yes, some of my models also meeted useless TTAs. It's wierd.</p>",
      "rawMarkdown": "Yes, some of my models also meeted useless TTAs. It's wierd.",
      "votes": null
    },
    {
      "id": "1633917",
      "postDate": "12/31/2021 06:49:01",
      "content": "<p>why did this(higher resolution matter most) happen even we change the biggest anchor size to  128 or 64 ?</p>",
      "rawMarkdown": "why did this(higher resolution matter most) happen even we change the biggest anchor size to  128 or 64 ?",
      "votes": null
    },
    {
      "id": "1633930",
      "postDate": "12/31/2021 07:12:06",
      "content": "<p>Sorry, I don't understand your question😂. Can you repeat it in Chinese?</p>",
      "rawMarkdown": "Sorry, I don't understand your question😂. Can you repeat it in Chinese?",
      "votes": null
    },
    {
      "id": "1633935",
      "postDate": "12/31/2021 07:30:37",
      "content": "<ol>\n<li><p>请问为什么增加inference时的分辨率如此的重要？<br>\n我把anchor size 改为了 8 16 32 64 128 时 如果infer分辨率还是800依然结果很差 lb 0.296，但是infer 分辨率改成 1300 就 lb 0.308了</p></li>\n<li><p>我的假设：<br>\n我在一开始一直认为这是因为 mask rcnn 前面几层feature map（p2 p3 p4）层数太浅，性能不如后面几层(p3 p4 p5)，  如果让分辨率增大，让大部分目标触发mask rcnn后面几层就能利用更深的网络产生的feature map 产生更好的结果。</p></li>\n<li><p>如果我上面这个假设成立：<br>\n那么调小anchor 让大部分目标出触发 128 的anchor 也能起到同样的效果</p></li>\n<li><p>验证：但是实际测试下来还是存在这个现象<br>\n仅调小anchor只能有很小提升，增加inference 时的分辨率可以极大提升精度</p></li>\n</ol>",
      "rawMarkdown": "1. 请问为什么增加inference时的分辨率如此的重要？\n  我把anchor size 改为了 8 16 32 64 128 时 如果infer分辨率还是800依然结果很差 lb 0.296，但是infer 分辨率改成 1300 就 lb 0.308了\n\n2.  我的假设：\n  我在一开始一直认为这是因为 mask rcnn 前面几层feature map（p2 p3 p4）层数太浅，性能不如后面几层(p3 p4 p5)，  如果让分辨率增大，让大部分目标触发mask rcnn后面几层就能利用更深的网络产生的feature map 产生更好的结果。\n\n3. 如果我上面这个假设成立：\n  那么调小anchor 让大部分目标出触发 128 的anchor 也能起到同样的效果\n\n\n4. 验证：但是实际测试下来还是存在这个现象\n  仅调小anchor只能有很小提升，增加inference 时的分辨率可以极大提升精度",
      "votes": null
    },
    {
      "id": "1633963",
      "postDate": "12/31/2021 08:13:55",
      "content": "<p>这是我的理解：小图像的下采样会损失大量的信息，proposal包围的位置特征点更少，不利于目标检测<br>\n但是增大图像可以增加信息量，proposal也可以包围更多特征点。</p>\n<p>尤其是这里的目标都偏小，模型偏向于在深层匹配，进一步加重了下采样的细节损失。</p>\n<p>对后面的RoIPooling也是同样的道理。</p>",
      "rawMarkdown": "这是我的理解：小图像的下采样会损失大量的信息，proposal包围的位置特征点更少，不利于目标检测\n但是增大图像可以增加信息量，proposal也可以包围更多特征点。\n\n尤其是这里的目标都偏小，模型偏向于在深层匹配，进一步加重了下采样的细节损失。\n\n对后面的RoIPooling也是同样的道理。",
      "votes": null
    },
    {
      "id": "1633993",
      "postDate": "12/31/2021 08:43:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>! Congratulations on the great solution and the gold medal. There is one thing I don't fully understand. Can you share how you handled masks that didn't fully fit into the slice? Have you tried combining masks that were split into multiple slice? Thank you. </p>",
      "rawMarkdown": "Hi @carnozhao! Congratulations on the great solution and the gold medal. There is one thing I don't fully understand. Can you share how you handled masks that didn't fully fit into the slice? Have you tried combining masks that were split into multiple slice? Thank you.",
      "votes": null
    },
    {
      "id": "1633997",
      "postDate": "12/31/2021 08:45:22",
      "content": "<p>In offline slicing, I used 25% as the threshold. If the mask area in slice is lower than 25% of its original area, it will be dropped.</p>\n<p>In inferencing, I used overlap mid point to decide whether a prediction should be kept. Details are mentioned at <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/297986#1633741\" target=\"_blank\">here</a>.</p>\n<p>Moreover, my slicing size is (208, 281), which is larger than most of cells, and stride is (104, 140), which is exactly the half of slice size. So, most of cells predictions will be fully covered by one slice. Those half predictions (at two different slices' edge) usually has lower score than a full prediction at slice center, which means that half predictions will be removed in postprocessing.</p>",
      "rawMarkdown": "In offline slicing, I used 25% as the threshold. If the mask area in slice is lower than 25% of its original area, it will be dropped.\n\nIn inferencing, I used overlap mid point to decide whether a prediction should be kept. Details are mentioned at [here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/297986#1633741).\n\nMoreover, my slicing size is (208, 281), which is larger than most of cells, and stride is (104, 140), which is exactly the half of slice size. So, most of cells predictions will be fully covered by one slice. Those half predictions (at two different slices' edge) usually has lower score than a full prediction at slice center, which means that half predictions will be removed in postprocessing.",
      "votes": null
    },
    {
      "id": "1634010",
      "postDate": "12/31/2021 09:03:31",
      "content": "<p><a href=\"https://www.kaggle.com/drzhuzhe\" target=\"_blank\">@drzhuzhe</a> <br>\n請問你改變anchor size 是在inference  還是training / inference都改變?</p>\n<p>提供一下我的想法<br>\n改變anchor size 可能會影響 每層feature map 對應回 input image 的覆蓋範圍<br>\n例如 P2 feature 是經過3次降採樣得到的<br>\n如果直接修改anchor sizee (32 -&gt; 8)<br>\n會很大程度影響proposal能覆蓋的範圍，進而影響效果<br>\n反之提高resolution不會影響proposal的覆蓋的範圍</p>",
      "rawMarkdown": "drzhuzhe \n請問你改變anchor size 是在inference  還是training / inference都改變?\n\n提供一下我的想法\n改變anchor size 可能會影響 每層feature map 對應回 input image 的覆蓋範圍\n例如 P2 feature 是經過3次降採樣得到的\n如果直接修改anchor sizee (32 -> 8)\n會很大程度影響proposal能覆蓋的範圍，進而影響效果\n反之提高resolution不會影響proposal的覆蓋的範圍",
      "votes": null
    },
    {
      "id": "1634013",
      "postDate": "12/31/2021 09:08:51",
      "content": "<ol>\n<li><p>anchor  是在 training 和 inference 时都改变了 <br>\ncfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256]]<br>\ncfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]</p></li>\n<li><p><code>例如 P2 feature 是經過3次降採樣得到的\n如果直接修改anchor sizee (32 -&gt; 8)</code>  我保留了 anchor 之间的 stride 未变</p></li>\n<li><p>以上都是在，数据增强时将 图片 scale 0.1 到 2 再截取 1024 x 1024 的区域来训练</p></li>\n</ol>\n<blockquote>\n  <p>image_size = 1024<br>\n  def build_customize_aug(cfg):    <br>\n      augs = [T.ResizeScale(<br>\n          min_scale=0.1, max_scale=2.0, target_height=image_size, target_width=image_size<br>\n      ),<br>\n      T.FixedSizeCrop(crop_size=(image_size, image_size)),</p>\n</blockquote>",
      "rawMarkdown": "1. anchor  是在 training 和 inference 时都改变了 \n   cfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256]]\n   cfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]\n\n2. `例如 P2 feature 是經過3次降採樣得到的\n如果直接修改anchor sizee (32 -> 8)`  我保留了 anchor 之间的 stride 未变\n\n3. 以上都是在，数据增强时将 图片 scale 0.1 到 2 再截取 1024 x 1024 的区域来训练\n\n> image_size = 1024\ndef build_customize_aug(cfg):    \n    augs = [T.ResizeScale(\n        min_scale=0.1, max_scale=2.0, target_height=image_size, target_width=image_size\n    ),\n    T.FixedSizeCrop(crop_size=(image_size, image_size)),",
      "votes": null
    },
    {
      "id": "1634015",
      "postDate": "12/31/2021 09:12:58",
      "content": "<p>Thanks for your explanation:)</p>",
      "rawMarkdown": "Thanks for your explanation:)",
      "votes": null
    },
    {
      "id": "1634033",
      "postDate": "12/31/2021 09:40:24",
      "content": "<p>我剛翻了下detectron裡的config<br>\nP2 的stride = 4 (那應該只有下採樣2次，這邊更正一下)</p>\n<p>另外上面的data augmentation 可能會有些不好的影響<br>\n如果原圖為520*704  而cell object size = 10 x 20<br>\n在scale 0.1x 會直接將information 直接丟光 (在data augmentation 階段)<br>\n因information 已經消失<br>\nmodel 也不太可能學習</p>",
      "rawMarkdown": "我剛翻了下detectron裡的config\nP2 的stride = 4 (那應該只有下採樣2次，這邊更正一下)\n\n另外上面的data augmentation 可能會有些不好的影響\n如果原圖為520*704  而cell object size = 10 x 20\n在scale 0.1x 會直接將information 直接丟光 (在data augmentation 階段)\n因information 已經消失\nmodel 也不太可能學習",
      "votes": null
    },
    {
      "id": "1634034",
      "postDate": "12/31/2021 09:44:25",
      "content": "<p>另外好奇一下你在aspect ration上的設定<br>\ncfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]<br>\n是有幫助的嗎?<br>\n在我這邊   <br>\n使用default的 [[0.5, 1.0, 2.0]] <br>\n跟修改後的[[1/2.5, 1.0, 2.5]]，[[1/3, 1/2, 1.0, 2.0, 3]] <br>\n在cv及lb上沒有差別</p>",
      "rawMarkdown": "另外好奇一下你在aspect ration上的設定\ncfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]\n是有幫助的嗎?\n在我這邊   \n使用default的 [[0.5, 1.0, 2.0]] \n跟修改後的[[1/2.5, 1.0, 2.5]]，[[1/3, 1/2, 1.0, 2.0, 3]] \n在cv及lb上沒有差別",
      "votes": null
    },
    {
      "id": "1634035",
      "postDate": "12/31/2021 09:45:18",
      "content": "<p>Your insights about model that does not really matter in private LB and that ensembling doesn't help are both very interesting and useful. I will take inspiration from your work. Compliments for your solo gold!</p>",
      "rawMarkdown": "Your insights about model that does not really matter in private LB and that ensembling doesn't help are both very interesting and useful. I will take inspiration from your work. Compliments for your solo gold!",
      "votes": null
    },
    {
      "id": "1634050",
      "postDate": "12/31/2021 10:15:25",
      "content": "<p>anchor 变小有帮助 , anchor ratio 没有大的差别</p>",
      "rawMarkdown": "anchor 变小有帮助 , anchor ratio 没有大的差别",
      "votes": null
    },
    {
      "id": "1634293",
      "postDate": "12/31/2021 14:42:33",
      "content": "<p>Congratulations on getting the individual gold batch in competition it not easy to achieve such a thing </p>",
      "rawMarkdown": "Congratulations on getting the individual gold batch in competition it not easy to achieve such a thing",
      "votes": null
    },
    {
      "id": "1634303",
      "postDate": "12/31/2021 14:54:42",
      "content": "<p>Congratulations :)</p>\n<p>Facts turned out that your method hight resolution with a black box training was a way to reach good enough result.</p>\n<p>During my state of art, I've read a 'Geometry-Aware Cell Detection with Deep Learning' method, it's interesting. I didn't realise it, so I would like to share this paper's idea (<a href=\"https://doi.org/10.1128/mSystems.00840-19)\" target=\"_blank\">https://doi.org/10.1128/mSystems.00840-19)</a>.</p>",
      "rawMarkdown": "Congratulations :)\n\nFacts turned out that your method hight resolution with a black box training was a way to reach good enough result.\n\nDuring my state of art, I've read a 'Geometry-Aware Cell Detection with Deep Learning' method, it's interesting. I didn't realise it, so I would like to share this paper's idea (https://doi.org/10.1128/mSystems.00840-19).",
      "votes": null
    },
    {
      "id": "1634345",
      "postDate": "12/31/2021 16:06:16",
      "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>  although we too used scales of higher image size, what i try to understand is how mmdet  gives final mask  output which is same as standard image size 520 704.</p>",
      "rawMarkdown": "carnozhao  although we too used scales of higher image size, what i try to understand is how mmdet  gives final mask  output which is same as standard image size 520 704.",
      "votes": null
    },
    {
      "id": "1634804",
      "postDate": "01/01/2022 07:25:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, cograts and thanks for your sharing.<br>\nDid you resize (208, 281) images into (1333, 1333)-(800,800) while training or the original (540, 704) ?</p>",
      "rawMarkdown": "Hi @carnozhao, cograts and thanks for your sharing.\nDid you resize (208, 281) images into (1333, 1333)-(800,800) while training or the original (540, 704) ?",
      "votes": null
    },
    {
      "id": "1634905",
      "postDate": "01/01/2022 09:49:52",
      "content": "<p>I resized 208,281</p>",
      "rawMarkdown": "I resized 208,281",
      "votes": null
    },
    {
      "id": "1634994",
      "postDate": "01/01/2022 11:34:44",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> It really opened my mind</p>",
      "rawMarkdown": "Thank you @carnozhao It really opened my mind",
      "votes": null
    },
    {
      "id": "1635310",
      "postDate": "01/01/2022 16:53:38",
      "content": "<p>Congrats! Very impressive that you have done it without using original data. </p>",
      "rawMarkdown": "Congrats! Very impressive that you have done it without using original data.",
      "votes": null
    },
    {
      "id": "1635373",
      "postDate": "01/01/2022 17:48:02",
      "content": "<p>Congratulations !great work</p>",
      "rawMarkdown": "Congratulations !great work",
      "votes": null
    },
    {
      "id": "1635724",
      "postDate": "01/02/2022 06:21:54",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> First off all Configurations! </p>\n<p>Y? You did not scale image resolution into 720x720 </p>\n<p>But your Solution is Impressive :)</p>",
      "rawMarkdown": "Hello, @carnozhao First off all Configurations! \n\nY? You did not scale image resolution into 720x720 \n\n\nBut your Solution is Impressive :)",
      "votes": null
    },
    {
      "id": "1635738",
      "postDate": "01/02/2022 06:32:11",
      "content": "<p>I tried train/test image scales: <br>\n(400,400)-(1333,1333)<br>\n(600,600)-(1333,1333)<br>\n(800,800)-(1333,1333)<br>\n(1024,1024)-(1333,1333)</p>\n<p>(800,800) performed best in local cross validation. (1024,1024) was slightly lower. (400,400) and (600,600) were much worse.</p>",
      "rawMarkdown": "I tried train/test image scales: \n(400,400)-(1333,1333)\n(600,600)-(1333,1333)\n(800,800)-(1333,1333)\n(1024,1024)-(1333,1333)\n\n(800,800) performed best in local cross validation. (1024,1024) was slightly lower. (400,400) and (600,600) were much worse.",
      "votes": null
    },
    {
      "id": "1635744",
      "postDate": "01/02/2022 06:36:43",
      "content": "<p>Sorry for my late reply. In mmet, input data including image array and its own \"original infomation\" named <code>img_meta</code> were passed to model. The model will resize output mask to image's original size.</p>",
      "rawMarkdown": "Sorry for my late reply. In mmet, input data including image array and its own \"original infomation\" named `img_meta` were passed to model. The model will resize output mask to image's original size.",
      "votes": null
    },
    {
      "id": "1635821",
      "postDate": "01/02/2022 08:14:27",
      "content": "<p>oh! Cleared Thank you for your information! Great Notebook Kernal </p>",
      "rawMarkdown": "oh! Cleared Thank you for your information! Great Notebook Kernal",
      "votes": null
    },
    {
      "id": "1636701",
      "postDate": "01/03/2022 07:20:42",
      "content": "<p>nice work<br>\nsimple is the best</p>",
      "rawMarkdown": "nice work\nsimple is the best",
      "votes": null
    },
    {
      "id": "1643444",
      "postDate": "01/09/2022 11:56:08",
      "content": "<p>hi sorry to be a so late comment with very specific question.<br>\nI tried to reproduce you work, and get a warning about \"DownSampleCocoDataset\"<br>\nKeyError: 'DownSampleCocoDataset is not in the dataset registry'<br>\nDoes that mean there is a customized dataset class? But I didn't find it in the repo…</p>\n<p>thank you and  congrats~</p>",
      "rawMarkdown": "hi sorry to be a so late comment with very specific question.\nI tried to reproduce you work, and get a warning about \"DownSampleCocoDataset\"\nKeyError: 'DownSampleCocoDataset is not in the dataset registry'\nDoes that mean there is a customized dataset class? But I didn't find it in the repo...\n\nthank you and  congrats~",
      "votes": null
    },
    {
      "id": "1643611",
      "postDate": "01/09/2022 15:16:37",
      "content": "<p>First, make sure that you are on \"sartorius\" branch. Then, \"DownSampleCocoDataset\" is implemented in <code>mmdet/dataset/dataset_wrappers.py</code>.</p>\n<p>It downsamples a dataset by a given ratio. E.g, a dataset including 1000 images with downsample ratio  2 will only use 500 random sampled images in training per epoch. It is easy to implement and you can try it yourself.</p>",
      "rawMarkdown": "First, make sure that you are on \"sartorius\" branch. Then, \"DownSampleCocoDataset\" is implemented in `mmdet/dataset/dataset_wrappers.py`.\n\nIt downsamples a dataset by a given ratio. E.g, a dataset including 1000 images with downsample ratio  2 will only use 500 random sampled images in training per epoch. It is easy to implement and you can try it yourself.",
      "votes": null
    },
    {
      "id": "1648543",
      "postDate": "01/13/2022 13:41:37",
      "content": "<p>Congratulations! Brevity is the soul of wit.</p>",
      "rawMarkdown": "Congratulations! Brevity is the soul of wit.",
      "votes": null
    },
    {
      "id": "1708236",
      "postDate": "03/01/2022 09:06:59",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1633628,
      "author_name": "charonwangg",
      "author_url": "",
      "post_date": "12/31/2021 00:31:55",
      "content": "<p>Congrats! Very neat and impressive solution, Carno Zhao is all we need: )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1633634,
      "author_name": "woprime",
      "author_url": "",
      "post_date": "12/31/2021 00:36:21",
      "content": "<p>Congrats! One question about the cropping during training, how much of a boost did you get from doing that? From what I saw and tried, random cropping hurts performance, so this felt a bit odd to me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1633638,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 00:39:30",
          "content": "<p>When using original images, my public LB was 0.320<br>\nWhen slicing images to 3x3=9 smaller overlapped images, my public LB was 0.331<br>\nWhen slicing images to 4x4=16 smaller overlapped images, my public LB was 0.334</p>\n<p>From 0.334 to 0.340 public LB, I used other tricks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633641,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "12/31/2021 00:41:48",
          "content": "<p>Thanks for the info!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633642,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 00:42:23",
          "content": "<p>I think the key point is using higher resolution in inferencing. To do this, we need training models using higher resolution (Random crop or slicing as I did)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633645,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "12/31/2021 00:47:20",
          "content": "<p>We found out about that too, though we did not train with higher resolution (Sadge). Initially we thought the performance boost was because of anchor size being too large, but for smaller anchor size model, higher res still give better performance, so we gave up trying to find out why and roll with it. Do you have a hypothesis on why this is the case?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633647,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 00:50:59",
          "content": "<p>I also did some experiments about anchors. I think when using low resolution with small anchor size, the feature map downsampled from image is too small to provide useful information, especially in dense objects situation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633651,
          "author_name": "woprime",
          "author_url": "",
          "post_date": "12/31/2021 00:54:06",
          "content": "<p>That make alot of sense, thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1633640,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "12/31/2021 00:41:32",
      "content": "<p>Very nice!!! congratulations !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1633644,
      "author_name": "aimanlim0",
      "author_url": "",
      "post_date": "12/31/2021 00:47:19",
      "content": "<p>Unexpectedly,the offline slicing method of satellite image segmentation can also be applied to this competition！ congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1633666,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "12/31/2021 01:12:44",
      "content": "<p>Congrats on solo gold medal <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, you don't you TTA?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1633673,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 01:16:23",
          "content": "<p>Yes I used Horizontal + Vertical flip, [(1333,1333), (1024,1024), (800,800)] multiscale test. All of them are provided in mmdet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633691,
          "author_name": "duykhanh99",
          "author_url": "",
          "post_date": "12/31/2021 01:36:32",
          "content": "<p>Your solution is very good with you not using pseudo label and pretrain livecell , during training do you use augmentation, hyper parameters tuning or any special technique. I can't get such a high score with HTC</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633707,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 01:58:30",
          "content": "<p>I used H-flip, V-flip, Rot90 multiscale in training as augmentations. I didnt do any hyperparameter tuning because that's too time consuming. Again, the KEY points is smaller images :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633733,
          "author_name": "duykhanh99",
          "author_url": "",
          "post_date": "12/31/2021 02:33:02",
          "content": "<p>Thanks for your explanation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633766,
          "author_name": "zaopolearning",
          "author_url": "",
          "post_date": "12/31/2021 03:33:09",
          "content": "<p>In my test, adding vertical flip to TTA did not improve the prediction.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633767,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 03:42:06",
          "content": "<p>Yes, some of my models also meeted useless TTAs. It's wierd.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634345,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/31/2021 16:06:16",
          "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>  although we too used scales of higher image size, what i try to understand is how mmdet  gives final mask  output which is same as standard image size 520 704.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1635744,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "01/02/2022 06:36:43",
          "content": "<p>Sorry for my late reply. In mmet, input data including image array and its own \"original infomation\" named <code>img_meta</code> were passed to model. The model will resize output mask to image's original size.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1633678,
      "author_name": "damtrongtuyen",
      "author_url": "",
      "post_date": "12/31/2021 01:19:58",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> for sharing, I also trained HTC but the performance not good. Can you share with us how did you train HTC-Res2Net101!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1633680,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 01:21:17",
          "content": "<p>I will update to my github repo soon!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633724,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 02:14:16",
          "content": "<p>github repo updated</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633740,
          "author_name": "damtrongtuyen",
          "author_url": "",
          "post_date": "12/31/2021 02:42:56",
          "content": "<p>Thank you, I see! Can you share the idea behind <code>sliced into smaller size</code>, why did you think that will work and tried?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633743,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 02:49:06",
          "content": "<p>The biggest motivation is that the objects are small and densly distributed, which is different than other object detection dataset like COCO.  However, my GPU cannot train with really large images, so cutting them and resize cutted images to bigger one seems natural.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1633708,
      "author_name": "jackchungchiehyu",
      "author_url": "",
      "post_date": "12/31/2021 02:00:59",
      "content": "<p>Thanks for sharing and congrats.  </p>\n<p>How did you deal with the predicted masks where the test slices overlap?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1633741,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 02:45:08",
          "content": "<p>Good question, thats also very important.<br>\nE.g, say slice A's x axis ranges from 0 to 100, and slice B's x axis ranges from 50 to 150. The middle point of overlap is 75, so I kept slice A's predictions whose box center ranges from 0-75, and slices B's predictions whose box center ranges from 75-150.</p>\n<p>This is only a simple exsample. You can apply it to x axis overlap, y axis overlap, and also multiple sequential overlaps.</p>\n<p>The detail of implementation can be found at <a href=\"https://www.kaggle.com/carnozhao/cell-submission?scriptVersionId=83795256&amp;cellId=3\" target=\"_blank\">my submission notebook</a></p>\n<pre><code>valid = (small_box[:,-1] &gt; THRESHOLDS_small[class_id]) &amp; \\\n            ~((i &lt;= 2) &amp; (small_box[:,[1,3]].mean(1) &gt; 3 * H / 10)) &amp; \\\n            ~((i &gt;= 1) &amp; (small_box[:,[1,3]].mean(1) &lt; H / 10)) &amp; \\\n            ~((j &lt;= 2) &amp; (small_box[:,[0,2]].mean(1) &gt; 3 * W / 10)) &amp; \\\n            ~((j &gt;= 1) &amp; (small_box[:,[0,2]].mean(1) &lt; W / 10))\n</code></pre>\n<p>Here, H/10, W/10, 3H/10, 3W/10 are overlap middle points of each small image.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1633726,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "12/31/2021 02:20:32",
      "content": "<p>Congratulations on your solo gold! Simple yet very effective! Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1633755,
      "author_name": "yichenwang1988",
      "author_url": "",
      "post_date": "12/31/2021 03:09:49",
      "content": "<p>It's true that the size matters a lot. If you resize your result mask to a bigger size and test, the score can go high a little. If resize it to a smaller size it goes down a lot.I doubt really there is something inproper for the iou metric when the task target is all small objects.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1633917,
      "author_name": "drzhuzhe",
      "author_url": "",
      "post_date": "12/31/2021 06:49:01",
      "content": "<p>why did this(higher resolution matter most) happen even we change the biggest anchor size to  128 or 64 ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1633930,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 07:12:06",
          "content": "<p>Sorry, I don't understand your question😂. Can you repeat it in Chinese?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633935,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "12/31/2021 07:30:37",
          "content": "<ol>\n<li><p>请问为什么增加inference时的分辨率如此的重要？<br>\n我把anchor size 改为了 8 16 32 64 128 时 如果infer分辨率还是800依然结果很差 lb 0.296，但是infer 分辨率改成 1300 就 lb 0.308了</p></li>\n<li><p>我的假设：<br>\n我在一开始一直认为这是因为 mask rcnn 前面几层feature map（p2 p3 p4）层数太浅，性能不如后面几层(p3 p4 p5)，  如果让分辨率增大，让大部分目标触发mask rcnn后面几层就能利用更深的网络产生的feature map 产生更好的结果。</p></li>\n<li><p>如果我上面这个假设成立：<br>\n那么调小anchor 让大部分目标出触发 128 的anchor 也能起到同样的效果</p></li>\n<li><p>验证：但是实际测试下来还是存在这个现象<br>\n仅调小anchor只能有很小提升，增加inference 时的分辨率可以极大提升精度</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1633963,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 08:13:55",
          "content": "<p>这是我的理解：小图像的下采样会损失大量的信息，proposal包围的位置特征点更少，不利于目标检测<br>\n但是增大图像可以增加信息量，proposal也可以包围更多特征点。</p>\n<p>尤其是这里的目标都偏小，模型偏向于在深层匹配，进一步加重了下采样的细节损失。</p>\n<p>对后面的RoIPooling也是同样的道理。</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634010,
          "author_name": "dragon229",
          "author_url": "",
          "post_date": "12/31/2021 09:03:31",
          "content": "<p><a href=\"https://www.kaggle.com/drzhuzhe\" target=\"_blank\">@drzhuzhe</a> <br>\n請問你改變anchor size 是在inference  還是training / inference都改變?</p>\n<p>提供一下我的想法<br>\n改變anchor size 可能會影響 每層feature map 對應回 input image 的覆蓋範圍<br>\n例如 P2 feature 是經過3次降採樣得到的<br>\n如果直接修改anchor sizee (32 -&gt; 8)<br>\n會很大程度影響proposal能覆蓋的範圍，進而影響效果<br>\n反之提高resolution不會影響proposal的覆蓋的範圍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634013,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "12/31/2021 09:08:51",
          "content": "<ol>\n<li><p>anchor  是在 training 和 inference 时都改变了 <br>\ncfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256]]<br>\ncfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]</p></li>\n<li><p><code>例如 P2 feature 是經過3次降採樣得到的\n如果直接修改anchor sizee (32 -&gt; 8)</code>  我保留了 anchor 之间的 stride 未变</p></li>\n<li><p>以上都是在，数据增强时将 图片 scale 0.1 到 2 再截取 1024 x 1024 的区域来训练</p></li>\n</ol>\n<blockquote>\n  <p>image_size = 1024<br>\n  def build_customize_aug(cfg):    <br>\n      augs = [T.ResizeScale(<br>\n          min_scale=0.1, max_scale=2.0, target_height=image_size, target_width=image_size<br>\n      ),<br>\n      T.FixedSizeCrop(crop_size=(image_size, image_size)),</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634033,
          "author_name": "dragon229",
          "author_url": "",
          "post_date": "12/31/2021 09:40:24",
          "content": "<p>我剛翻了下detectron裡的config<br>\nP2 的stride = 4 (那應該只有下採樣2次，這邊更正一下)</p>\n<p>另外上面的data augmentation 可能會有些不好的影響<br>\n如果原圖為520*704  而cell object size = 10 x 20<br>\n在scale 0.1x 會直接將information 直接丟光 (在data augmentation 階段)<br>\n因information 已經消失<br>\nmodel 也不太可能學習</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634034,
          "author_name": "dragon229",
          "author_url": "",
          "post_date": "12/31/2021 09:44:25",
          "content": "<p>另外好奇一下你在aspect ration上的設定<br>\ncfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]<br>\n是有幫助的嗎?<br>\n在我這邊   <br>\n使用default的 [[0.5, 1.0, 2.0]] <br>\n跟修改後的[[1/2.5, 1.0, 2.5]]，[[1/3, 1/2, 1.0, 2.0, 3]] <br>\n在cv及lb上沒有差別</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634050,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "12/31/2021 10:15:25",
          "content": "<p>anchor 变小有帮助 , anchor ratio 没有大的差别</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1633993,
      "author_name": "danjafish",
      "author_url": "",
      "post_date": "12/31/2021 08:43:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>! Congratulations on the great solution and the gold medal. There is one thing I don't fully understand. Can you share how you handled masks that didn't fully fit into the slice? Have you tried combining masks that were split into multiple slice? Thank you. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1633997,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "12/31/2021 08:45:22",
          "content": "<p>In offline slicing, I used 25% as the threshold. If the mask area in slice is lower than 25% of its original area, it will be dropped.</p>\n<p>In inferencing, I used overlap mid point to decide whether a prediction should be kept. Details are mentioned at <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/297986#1633741\" target=\"_blank\">here</a>.</p>\n<p>Moreover, my slicing size is (208, 281), which is larger than most of cells, and stride is (104, 140), which is exactly the half of slice size. So, most of cells predictions will be fully covered by one slice. Those half predictions (at two different slices' edge) usually has lower score than a full prediction at slice center, which means that half predictions will be removed in postprocessing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634015,
          "author_name": "danjafish",
          "author_url": "",
          "post_date": "12/31/2021 09:12:58",
          "content": "<p>Thanks for your explanation:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1634035,
      "author_name": "lucamassaron",
      "author_url": "",
      "post_date": "12/31/2021 09:45:18",
      "content": "<p>Your insights about model that does not really matter in private LB and that ensembling doesn't help are both very interesting and useful. I will take inspiration from your work. Compliments for your solo gold!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1634293,
      "author_name": "balavashan",
      "author_url": "",
      "post_date": "12/31/2021 14:42:33",
      "content": "<p>Congratulations on getting the individual gold batch in competition it not easy to achieve such a thing </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1634303,
      "author_name": "zhaoyuanyuanzyy",
      "author_url": "",
      "post_date": "12/31/2021 14:54:42",
      "content": "<p>Congratulations :)</p>\n<p>Facts turned out that your method hight resolution with a black box training was a way to reach good enough result.</p>\n<p>During my state of art, I've read a 'Geometry-Aware Cell Detection with Deep Learning' method, it's interesting. I didn't realise it, so I would like to share this paper's idea (<a href=\"https://doi.org/10.1128/mSystems.00840-19)\" target=\"_blank\">https://doi.org/10.1128/mSystems.00840-19)</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1634804,
      "author_name": "forcewithme",
      "author_url": "",
      "post_date": "01/01/2022 07:25:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, cograts and thanks for your sharing.<br>\nDid you resize (208, 281) images into (1333, 1333)-(800,800) while training or the original (540, 704) ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1634905,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "01/01/2022 09:49:52",
          "content": "<p>I resized 208,281</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1634994,
          "author_name": "forcewithme",
          "author_url": "",
          "post_date": "01/01/2022 11:34:44",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> It really opened my mind</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1635310,
      "author_name": "hsadeghian",
      "author_url": "",
      "post_date": "01/01/2022 16:53:38",
      "content": "<p>Congrats! Very impressive that you have done it without using original data. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1635373,
      "author_name": "aiswaryasivakumar",
      "author_url": "",
      "post_date": "01/01/2022 17:48:02",
      "content": "<p>Congratulations !great work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1635724,
      "author_name": "balasubramaniamv",
      "author_url": "",
      "post_date": "01/02/2022 06:21:54",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> First off all Configurations! </p>\n<p>Y? You did not scale image resolution into 720x720 </p>\n<p>But your Solution is Impressive :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1635738,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "01/02/2022 06:32:11",
          "content": "<p>I tried train/test image scales: <br>\n(400,400)-(1333,1333)<br>\n(600,600)-(1333,1333)<br>\n(800,800)-(1333,1333)<br>\n(1024,1024)-(1333,1333)</p>\n<p>(800,800) performed best in local cross validation. (1024,1024) was slightly lower. (400,400) and (600,600) were much worse.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1635821,
          "author_name": "balasubramaniamv",
          "author_url": "",
          "post_date": "01/02/2022 08:14:27",
          "content": "<p>oh! Cleared Thank you for your information! Great Notebook Kernal </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1636701,
      "author_name": "wyhsiao",
      "author_url": "",
      "post_date": "01/03/2022 07:20:42",
      "content": "<p>nice work<br>\nsimple is the best</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1643444,
      "author_name": "laineyzheng",
      "author_url": "",
      "post_date": "01/09/2022 11:56:08",
      "content": "<p>hi sorry to be a so late comment with very specific question.<br>\nI tried to reproduce you work, and get a warning about \"DownSampleCocoDataset\"<br>\nKeyError: 'DownSampleCocoDataset is not in the dataset registry'<br>\nDoes that mean there is a customized dataset class? But I didn't find it in the repo…</p>\n<p>thank you and  congrats~</p>",
      "votes": null,
      "replies": [
        {
          "id": 1643611,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "01/09/2022 15:16:37",
          "content": "<p>First, make sure that you are on \"sartorius\" branch. Then, \"DownSampleCocoDataset\" is implemented in <code>mmdet/dataset/dataset_wrappers.py</code>.</p>\n<p>It downsamples a dataset by a given ratio. E.g, a dataset including 1000 images with downsample ratio  2 will only use 500 random sampled images in training per epoch. It is easy to implement and you can try it yourself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1708236,
          "author_name": "laineyzheng",
          "author_url": "",
          "post_date": "03/01/2022 09:06:59",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1648543,
      "author_name": "martinid",
      "author_url": "",
      "post_date": "01/13/2022 13:41:37",
      "content": "<p>Congratulations! Brevity is the soul of wit.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1633626": "# Preprocessing\n\nI only used `train` images, without semi-supervised or LiveCell.\n\nAll train images were sliced into smaller size, with window size (208, 281), and stride (104, 140). So a (520, 704) image can be slices to 16 smaller images.\n\n\n\n# Training\n\n## Model\n\nIt seems that model does not really matter in private LB. My res2net-101, resnext-101, detectoRS-r50 can get 0.350 in private LB, even though their public LB are very different.\n\n\nEnsemble cannot improve both public and private LB significantly. For me, single model is enough.\n\n\n## Hyperparams\n\nI trained with mmdet, default 1x schedule and 1x swa training. Image scales were set as (1333, 1333)-(800,800).\n\n\n\n# Inferencing\n\nTest images were also sliced as training images. The test image scales were [(1333,1333), (1024,1024), (800,800)].\n\nMasks were iterated from higher score to lower score. Scores lower than class-wise threshold and areas lower than class-wise pixel-threshold were removed. \n\n\n\nWhen dealing with overlaps, I remove mask whose overlapped part was more than 20% of itself.\n\n\n\nRCNN's NMS was replaced by Weighted cluster-NMS with DIoU.\n\n# Submission\n\nMy final submission was ensemble of HTC-Res2Net101 (trained with all data), HTC-ResNeXt101 (trained with all data) and HTC-Res2Net101 (trained with fold0)\n\n# Weakness\n\nLONG inferencing time, 2h per model.\n\n# Github Repo\n\n[https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution](https://github.com/CarnoZhao/mmdetection/tree/sartorius_solution)",
    "1633628": "Congrats! Very neat and impressive solution, Carno Zhao is all we need: )",
    "1633634": "Congrats! One question about the cropping during training, how much of a boost did you get from doing that? From what I saw and tried, random cropping hurts performance, so this felt a bit odd to me.",
    "1633638": "When using original images, my public LB was 0.320\nWhen slicing images to 3x3=9 smaller overlapped images, my public LB was 0.331\nWhen slicing images to 4x4=16 smaller overlapped images, my public LB was 0.334\n\nFrom 0.334 to 0.340 public LB, I used other tricks.",
    "1633640": "Very nice!!! congratulations !!!",
    "1633641": "Thanks for the info!",
    "1633642": "I think the key point is using higher resolution in inferencing. To do this, we need training models using higher resolution (Random crop or slicing as I did)",
    "1633644": "Unexpectedly,the offline slicing method of satellite image segmentation can also be applied to this competition！ congratulations!",
    "1633645": "We found out about that too, though we did not train with higher resolution (Sadge). Initially we thought the performance boost was because of anchor size being too large, but for smaller anchor size model, higher res still give better performance, so we gave up trying to find out why and roll with it. Do you have a hypothesis on why this is the case?",
    "1633647": "I also did some experiments about anchors. I think when using low resolution with small anchor size, the feature map downsampled from image is too small to provide useful information, especially in dense objects situation.",
    "1633651": "That make alot of sense, thank you!",
    "1633666": "Congrats on solo gold medal @carnozhao, you don't you TTA?",
    "1633673": "Yes I used Horizontal + Vertical flip, [(1333,1333), (1024,1024), (800,800)] multiscale test. All of them are provided in mmdet.",
    "1633678": "Thanks @carnozhao for sharing, I also trained HTC but the performance not good. Can you share with us how did you train HTC-Res2Net101!",
    "1633680": "I will update to my github repo soon!",
    "1633691": "Your solution is very good with you not using pseudo label and pretrain livecell , during training do you use augmentation, hyper parameters tuning or any special technique. I can't get such a high score with HTC",
    "1633707": "I used H-flip, V-flip, Rot90 multiscale in training as augmentations. I didnt do any hyperparameter tuning because that's too time consuming. Again, the KEY points is smaller images :)",
    "1633708": "Thanks for sharing and congrats.  \n\nHow did you deal with the predicted masks where the test slices overlap?",
    "1633724": "github repo updated",
    "1633726": "Congratulations on your solo gold! Simple yet very effective! Great work!",
    "1633733": "Thanks for your explanation",
    "1633740": "Thank you, I see! Can you share the idea behind `sliced into smaller size`, why did you think that will work and tried?",
    "1633741": "Good question, thats also very important.\nE.g, say slice A's x axis ranges from 0 to 100, and slice B's x axis ranges from 50 to 150. The middle point of overlap is 75, so I kept slice A's predictions whose box center ranges from 0-75, and slices B's predictions whose box center ranges from 75-150.\n\nThis is only a simple exsample. You can apply it to x axis overlap, y axis overlap, and also multiple sequential overlaps.\n\nThe detail of implementation can be found at [my submission notebook](https://www.kaggle.com/carnozhao/cell-submission?scriptVersionId=83795256&cellId=3)\n\n```python\nvalid = (small_box[:,-1] > THRESHOLDS_small[class_id]) & \\\n            ~((i <= 2) & (small_box[:,[1,3]].mean(1) > 3 * H / 10)) & \\\n            ~((i >= 1) & (small_box[:,[1,3]].mean(1) < H / 10)) & \\\n            ~((j <= 2) & (small_box[:,[0,2]].mean(1) > 3 * W / 10)) & \\\n            ~((j >= 1) & (small_box[:,[0,2]].mean(1) < W / 10))\n```\nHere, H/10, W/10, 3H/10, 3W/10 are overlap middle points of each small image.",
    "1633743": "The biggest motivation is that the objects are small and densly distributed, which is different than other object detection dataset like COCO.  However, my GPU cannot train with really large images, so cutting them and resize cutted images to bigger one seems natural.",
    "1633755": "It's true that the size matters a lot. If you resize your result mask to a bigger size and test, the score can go high a little. If resize it to a smaller size it goes down a lot.I doubt really there is something inproper for the iou metric when the task target is all small objects.",
    "1633766": "In my test, adding vertical flip to TTA did not improve the prediction.",
    "1633767": "Yes, some of my models also meeted useless TTAs. It's wierd.",
    "1633917": "why did this(higher resolution matter most) happen even we change the biggest anchor size to  128 or 64 ?",
    "1633930": "Sorry, I don't understand your question😂. Can you repeat it in Chinese?",
    "1633935": "1. 请问为什么增加inference时的分辨率如此的重要？\n  我把anchor size 改为了 8 16 32 64 128 时 如果infer分辨率还是800依然结果很差 lb 0.296，但是infer 分辨率改成 1300 就 lb 0.308了\n\n2.  我的假设：\n  我在一开始一直认为这是因为 mask rcnn 前面几层feature map（p2 p3 p4）层数太浅，性能不如后面几层(p3 p4 p5)，  如果让分辨率增大，让大部分目标触发mask rcnn后面几层就能利用更深的网络产生的feature map 产生更好的结果。\n\n3. 如果我上面这个假设成立：\n  那么调小anchor 让大部分目标出触发 128 的anchor 也能起到同样的效果\n\n\n4. 验证：但是实际测试下来还是存在这个现象\n  仅调小anchor只能有很小提升，增加inference 时的分辨率可以极大提升精度",
    "1633963": "这是我的理解：小图像的下采样会损失大量的信息，proposal包围的位置特征点更少，不利于目标检测\n但是增大图像可以增加信息量，proposal也可以包围更多特征点。\n\n尤其是这里的目标都偏小，模型偏向于在深层匹配，进一步加重了下采样的细节损失。\n\n对后面的RoIPooling也是同样的道理。",
    "1633993": "Hi @carnozhao! Congratulations on the great solution and the gold medal. There is one thing I don't fully understand. Can you share how you handled masks that didn't fully fit into the slice? Have you tried combining masks that were split into multiple slice? Thank you.",
    "1633997": "In offline slicing, I used 25% as the threshold. If the mask area in slice is lower than 25% of its original area, it will be dropped.\n\nIn inferencing, I used overlap mid point to decide whether a prediction should be kept. Details are mentioned at [here](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/297986#1633741).\n\nMoreover, my slicing size is (208, 281), which is larger than most of cells, and stride is (104, 140), which is exactly the half of slice size. So, most of cells predictions will be fully covered by one slice. Those half predictions (at two different slices' edge) usually has lower score than a full prediction at slice center, which means that half predictions will be removed in postprocessing.",
    "1634010": "drzhuzhe \n請問你改變anchor size 是在inference  還是training / inference都改變?\n\n提供一下我的想法\n改變anchor size 可能會影響 每層feature map 對應回 input image 的覆蓋範圍\n例如 P2 feature 是經過3次降採樣得到的\n如果直接修改anchor sizee (32 -> 8)\n會很大程度影響proposal能覆蓋的範圍，進而影響效果\n反之提高resolution不會影響proposal的覆蓋的範圍",
    "1634013": "1. anchor  是在 training 和 inference 时都改变了 \n   cfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256]]\n   cfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]\n\n2. `例如 P2 feature 是經過3次降採樣得到的\n如果直接修改anchor sizee (32 -> 8)`  我保留了 anchor 之间的 stride 未变\n\n3. 以上都是在，数据增强时将 图片 scale 0.1 到 2 再截取 1024 x 1024 的区域来训练\n\n> image_size = 1024\ndef build_customize_aug(cfg):    \n    augs = [T.ResizeScale(\n        min_scale=0.1, max_scale=2.0, target_height=image_size, target_width=image_size\n    ),\n    T.FixedSizeCrop(crop_size=(image_size, image_size)),",
    "1634015": "Thanks for your explanation:)",
    "1634033": "我剛翻了下detectron裡的config\nP2 的stride = 4 (那應該只有下採樣2次，這邊更正一下)\n\n另外上面的data augmentation 可能會有些不好的影響\n如果原圖為520*704  而cell object size = 10 x 20\n在scale 0.1x 會直接將information 直接丟光 (在data augmentation 階段)\n因information 已經消失\nmodel 也不太可能學習",
    "1634034": "另外好奇一下你在aspect ration上的設定\ncfg.MODEL.ASPECT_RATIOS = [[0.25, 0.5, 1.0, 2.0, 4.0]]\n是有幫助的嗎?\n在我這邊   \n使用default的 [[0.5, 1.0, 2.0]] \n跟修改後的[[1/2.5, 1.0, 2.5]]，[[1/3, 1/2, 1.0, 2.0, 3]] \n在cv及lb上沒有差別",
    "1634035": "Your insights about model that does not really matter in private LB and that ensembling doesn't help are both very interesting and useful. I will take inspiration from your work. Compliments for your solo gold!",
    "1634050": "anchor 变小有帮助 , anchor ratio 没有大的差别",
    "1634293": "Congratulations on getting the individual gold batch in competition it not easy to achieve such a thing",
    "1634303": "Congratulations :)\n\nFacts turned out that your method hight resolution with a black box training was a way to reach good enough result.\n\nDuring my state of art, I've read a 'Geometry-Aware Cell Detection with Deep Learning' method, it's interesting. I didn't realise it, so I would like to share this paper's idea (https://doi.org/10.1128/mSystems.00840-19).",
    "1634345": "carnozhao  although we too used scales of higher image size, what i try to understand is how mmdet  gives final mask  output which is same as standard image size 520 704.",
    "1634804": "Hi @carnozhao, cograts and thanks for your sharing.\nDid you resize (208, 281) images into (1333, 1333)-(800,800) while training or the original (540, 704) ?",
    "1634905": "I resized 208,281",
    "1634994": "Thank you @carnozhao It really opened my mind",
    "1635310": "Congrats! Very impressive that you have done it without using original data.",
    "1635373": "Congratulations !great work",
    "1635724": "Hello, @carnozhao First off all Configurations! \n\nY? You did not scale image resolution into 720x720 \n\n\nBut your Solution is Impressive :)",
    "1635738": "I tried train/test image scales: \n(400,400)-(1333,1333)\n(600,600)-(1333,1333)\n(800,800)-(1333,1333)\n(1024,1024)-(1333,1333)\n\n(800,800) performed best in local cross validation. (1024,1024) was slightly lower. (400,400) and (600,600) were much worse.",
    "1635744": "Sorry for my late reply. In mmet, input data including image array and its own \"original infomation\" named `img_meta` were passed to model. The model will resize output mask to image's original size.",
    "1635821": "oh! Cleared Thank you for your information! Great Notebook Kernal",
    "1636701": "nice work\nsimple is the best",
    "1643444": "hi sorry to be a so late comment with very specific question.\nI tried to reproduce you work, and get a warning about \"DownSampleCocoDataset\"\nKeyError: 'DownSampleCocoDataset is not in the dataset registry'\nDoes that mean there is a customized dataset class? But I didn't find it in the repo...\n\nthank you and  congrats~",
    "1643611": "First, make sure that you are on \"sartorius\" branch. Then, \"DownSampleCocoDataset\" is implemented in `mmdet/dataset/dataset_wrappers.py`.\n\nIt downsamples a dataset by a given ratio. E.g, a dataset including 1000 images with downsample ratio  2 will only use 500 random sampled images in training per epoch. It is easy to implement and you can try it yourself.",
    "1648543": "Congratulations! Brevity is the soul of wit.",
    "1708236": "Thanks a lot!"
  },
  "source": "meta"
}