{
  "id": 298038,
  "title": "11th place solution",
  "url": "/competitions/sartorius-cell-instance-segmentation/writeups/deeplive-exe-v2-0-11th-place-solution",
  "author_name": "",
  "post_date": "2021-12-31T09:53:54.550Z",
  "votes": 77,
  "comment_count": 30,
  "views": 0,
  "content": "<p>First of all congratz to the winners and thanks to Sartorius and Kaggle for hosting the competition !</p>\n<p>Although this 11th place is a great finish, we're a bit disappointed since we spent a month at the top and the final weeks in the top 5. We knew we were not gonna win but this shake was quite a surprise for us, and we still haven't really understood why it happened. </p>\n<p>Our models were just better on Public than Private, there might be a small domain shift (more corts ? more small cells ? more large cells ? we will never now…) that emphasized a weakness of our pipeline.</p>\n<p>Anyways, here are few interesting points of the solution that I found were worth mentioning, the full code is on GitHub : <a href=\"https://github.com/TheoViel/kaggle_sartorius\" target=\"_blank\">https://github.com/TheoViel/kaggle_sartorius</a> <br>\nAlso, our inference code is here : <a href=\"https://www.kaggle.com/theoviel/sartorius-inference-final\" target=\"_blank\">https://www.kaggle.com/theoviel/sartorius-inference-final</a></p>\n<h3>Validation Scheme</h3>\n<p>Two weeks before the end of the competition, we switched to a validation set-up we thought would be reliable : splitting on plate &amp; well as done in the Livecell paper. </p>\n<p><img src=\"https://i.ibb.co/dDjvC6j/livecell-3.png\" alt=\"\"></p>\n<p>This might be one of the reason we shake down ? Our best private LB (0.349, public 0.339) was our best CV before switching to the above scheme. Still, our final score of 0.348 is our best CV with this setup (public 0.342)</p>\n<h3>Models</h3>\n<p>We only used machines with single RTX 2080 Ti so we had to be ingenious to be able to train on high resolution images and detect small cells. We used mask-rcnn based models and relied on mmdet but only for model definition and augmentations, the rest of the pipeline is hand-crafted, which made it more convenient for experimenting.</p>\n<p><img src=\"https://i.ibb.co/MkwRbVW/Ensembling-drawio-1.png\" alt=\"\"></p>\n<h5>Main points</h5>\n<ul>\n<li>Remove the stride of the first layer of the encoder to \"increase the resolution\" of the models without doing any resizing !</li>\n<li>Random Crops of size 256x256 for training</li>\n<li>Pretrain on Livecell</li>\n<li>4000 iterations of finetuning on the training data</li>\n<li>Backbones : resnet50, resnext101_32x4, resnext101_64x4, efficientnet_b4/b5/b6</li>\n<li>Models : MaskRCNN, Cascade, HTC</li>\n</ul>\n<h5>Ensembling</h5>\n<p>We average predictions of different models &amp; different flips at three stages, the stages are the boxes with thicker borders above.</p>\n<ul>\n<li>Proposals : For a given feature map output by the FPN, each of its pixel is assigned a score and a coordinates prediction by the convolutions. This is what we average. </li>\n<li>Boxes : We re-use the ensembled proposal and perform averaging of the class predictions and coordinates for each proposal. We used 4 flip TTAs.</li>\n<li>Masks : starting with the ensembled boxes, we average the masks - before the upsampling back to the original image size.</li>\n</ul>\n<p>This scheme doesn't really use NMS for ensembling which can be tricky to use. Hence we stacked a bunch of models. We used 6 models per cell type.</p>\n<h5>Post processing</h5>\n<ul>\n<li>NMS on boxes using high thresholds, then NMS on masks using low thresholds</li>\n<li>Corrupt back the astro masks as we trained on clean ones (+0.002 LB)</li>\n<li>Small masks removal</li>\n</ul>\n<p>We did a lot of hyper-parameters tweaking on CV : NMS thresholds, RPN and bbox_head params, confidence thresholds, minimum cell sizes, mask thresholds.</p>\n<h3>Few more words</h3>\n<ul>\n<li><p>Pseudo Labelling didn't really work for us, we used them in the finetuning phase with the original training data and progressively decayed their proportion.</p></li>\n<li><p>The RoiAlign layer from mmdet has implementation issues. Masks resulting from TTA appear shifted which hurt performances, especially when trying to use vertical flips. We had to shift the boxes by 0.5 to counter this. </p>\n<p>I will probably add more stuff later, and fix the typos and all. Feel free to ask any questions  !<br>\n<em>Thanks for reading !</em></p></li>\n</ul>",
  "messages": [
    {
      "id": "1634030",
      "postDate": "12/31/2021 09:33:58",
      "content": "<p>First of all congratz to the winners and thanks to Sartorius and Kaggle for hosting the competition !</p>\n<p>Although this 11th place is a great finish, we're a bit disappointed since we spent a month at the top and the final weeks in the top 5. We knew we were not gonna win but this shake was quite a surprise for us, and we still haven't really understood why it happened. </p>\n<p>Our models were just better on Public than Private, there might be a small domain shift (more corts ? more small cells ? more large cells ? we will never now…) that emphasized a weakness of our pipeline.</p>\n<p>Anyways, here are few interesting points of the solution that I found were worth mentioning, the full code is on GitHub : <a href=\"https://github.com/TheoViel/kaggle_sartorius\" target=\"_blank\">https://github.com/TheoViel/kaggle_sartorius</a> <br>\nAlso, our inference code is here : <a href=\"https://www.kaggle.com/theoviel/sartorius-inference-final\" target=\"_blank\">https://www.kaggle.com/theoviel/sartorius-inference-final</a></p>\n<h3>Validation Scheme</h3>\n<p>Two weeks before the end of the competition, we switched to a validation set-up we thought would be reliable : splitting on plate &amp; well as done in the Livecell paper. </p>\n<p><img src=\"https://i.ibb.co/dDjvC6j/livecell-3.png\" alt=\"\"></p>\n<p>This might be one of the reason we shake down ? Our best private LB (0.349, public 0.339) was our best CV before switching to the above scheme. Still, our final score of 0.348 is our best CV with this setup (public 0.342)</p>\n<h3>Models</h3>\n<p>We only used machines with single RTX 2080 Ti so we had to be ingenious to be able to train on high resolution images and detect small cells. We used mask-rcnn based models and relied on mmdet but only for model definition and augmentations, the rest of the pipeline is hand-crafted, which made it more convenient for experimenting.</p>\n<p><img src=\"https://i.ibb.co/MkwRbVW/Ensembling-drawio-1.png\" alt=\"\"></p>\n<h5>Main points</h5>\n<ul>\n<li>Remove the stride of the first layer of the encoder to \"increase the resolution\" of the models without doing any resizing !</li>\n<li>Random Crops of size 256x256 for training</li>\n<li>Pretrain on Livecell</li>\n<li>4000 iterations of finetuning on the training data</li>\n<li>Backbones : resnet50, resnext101_32x4, resnext101_64x4, efficientnet_b4/b5/b6</li>\n<li>Models : MaskRCNN, Cascade, HTC</li>\n</ul>\n<h5>Ensembling</h5>\n<p>We average predictions of different models &amp; different flips at three stages, the stages are the boxes with thicker borders above.</p>\n<ul>\n<li>Proposals : For a given feature map output by the FPN, each of its pixel is assigned a score and a coordinates prediction by the convolutions. This is what we average. </li>\n<li>Boxes : We re-use the ensembled proposal and perform averaging of the class predictions and coordinates for each proposal. We used 4 flip TTAs.</li>\n<li>Masks : starting with the ensembled boxes, we average the masks - before the upsampling back to the original image size.</li>\n</ul>\n<p>This scheme doesn't really use NMS for ensembling which can be tricky to use. Hence we stacked a bunch of models. We used 6 models per cell type.</p>\n<h5>Post processing</h5>\n<ul>\n<li>NMS on boxes using high thresholds, then NMS on masks using low thresholds</li>\n<li>Corrupt back the astro masks as we trained on clean ones (+0.002 LB)</li>\n<li>Small masks removal</li>\n</ul>\n<p>We did a lot of hyper-parameters tweaking on CV : NMS thresholds, RPN and bbox_head params, confidence thresholds, minimum cell sizes, mask thresholds.</p>\n<h3>Few more words</h3>\n<ul>\n<li><p>Pseudo Labelling didn't really work for us, we used them in the finetuning phase with the original training data and progressively decayed their proportion.</p></li>\n<li><p>The RoiAlign layer from mmdet has implementation issues. Masks resulting from TTA appear shifted which hurt performances, especially when trying to use vertical flips. We had to shift the boxes by 0.5 to counter this. </p>\n<p>I will probably add more stuff later, and fix the typos and all. Feel free to ask any questions  !<br>\n<em>Thanks for reading !</em></p></li>\n</ul>",
      "rawMarkdown": "First of all congratz to the winners and thanks to Sartorius and Kaggle for hosting the competition !\n\nAlthough this 11th place is a great finish, we're a bit disappointed since we spent a month at the top and the final weeks in the top 5. We knew we were not gonna win but this shake was quite a surprise for us, and we still haven't really understood why it happened. \n\nOur models were just better on Public than Private, there might be a small domain shift (more corts ? more small cells ? more large cells ? we will never now...) that emphasized a weakness of our pipeline.\n\nAnyways, here are few interesting points of the solution that I found were worth mentioning, the full code is on GitHub : https://github.com/TheoViel/kaggle_sartorius \nAlso, our inference code is here : https://www.kaggle.com/theoviel/sartorius-inference-final\n\n### Validation Scheme\n\nTwo weeks before the end of the competition, we switched to a validation set-up we thought would be reliable : splitting on plate & well as done in the Livecell paper. \n\n![](https://i.ibb.co/dDjvC6j/livecell-3.png)\n\nThis might be one of the reason we shake down ? Our best private LB (0.349, public 0.339) was our best CV before switching to the above scheme. Still, our final score of 0.348 is our best CV with this setup (public 0.342)\n\n### Models\n\nWe only used machines with single RTX 2080 Ti so we had to be ingenious to be able to train on high resolution images and detect small cells. We used mask-rcnn based models and relied on mmdet but only for model definition and augmentations, the rest of the pipeline is hand-crafted, which made it more convenient for experimenting.\n\n\n![](https://i.ibb.co/MkwRbVW/Ensembling-drawio-1.png)\n\n##### Main points\n- Remove the stride of the first layer of the encoder to \"increase the resolution\" of the models without doing any resizing !\n- Random Crops of size 256x256 for training\n- Pretrain on Livecell\n- 4000 iterations of finetuning on the training data\n- Backbones : resnet50, resnext101_32x4, resnext101_64x4, efficientnet_b4/b5/b6\n- Models : MaskRCNN, Cascade, HTC\n\n##### Ensembling\n\nWe average predictions of different models & different flips at three stages, the stages are the boxes with thicker borders above.\n- Proposals : For a given feature map output by the FPN, each of its pixel is assigned a score and a coordinates prediction by the convolutions. This is what we average. \n- Boxes : We re-use the ensembled proposal and perform averaging of the class predictions and coordinates for each proposal. We used 4 flip TTAs.\n- Masks : starting with the ensembled boxes, we average the masks - before the upsampling back to the original image size.\n\nThis scheme doesn't really use NMS for ensembling which can be tricky to use. Hence we stacked a bunch of models. We used 6 models per cell type.\n\n##### Post processing\n\n- NMS on boxes using high thresholds, then NMS on masks using low thresholds\n- Corrupt back the astro masks as we trained on clean ones (+0.002 LB)\n- Small masks removal\n\nWe did a lot of hyper-parameters tweaking on CV : NMS thresholds, RPN and bbox_head params, confidence thresholds, minimum cell sizes, mask thresholds.\n\n### Few more words\n\n- Pseudo Labelling didn't really work for us, we used them in the finetuning phase with the original training data and progressively decayed their proportion.\n- The RoiAlign layer from mmdet has implementation issues. Masks resulting from TTA appear shifted which hurt performances, especially when trying to use vertical flips. We had to shift the boxes by 0.5 to counter this. \n\n I will probably add more stuff later, and fix the typos and all. Feel free to ask any questions  !\n*Thanks for reading !*",
      "votes": null
    },
    {
      "id": "1634078",
      "postDate": "12/31/2021 10:41:25",
      "content": "<p>Great work! Congrats on results <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and team. </p>",
      "rawMarkdown": "Great work! Congrats on results @theoviel and team.",
      "votes": null
    },
    {
      "id": "1634160",
      "postDate": "12/31/2021 11:35:41",
      "content": "<p>Still a great result <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>! In spite of the slightly worse private leaderboard position than you expected, I loved reading about your cv strategy \"splitting on plate &amp; well as done in the Livecell paper\". A very smart indeed strategy, well reasoned and grounded.</p>",
      "rawMarkdown": "Still a great result @theoviel! In spite of the slightly worse private leaderboard position than you expected, I loved reading about your cv strategy \"splitting on plate & well as done in the Livecell paper\". A very smart indeed strategy, well reasoned and grounded.",
      "votes": null
    },
    {
      "id": "1634509",
      "postDate": "12/31/2021 18:34:26",
      "content": "<p>Amazing! What a useful and pratical skill when computation is not enough! Thanks and Congrats!</p>",
      "rawMarkdown": "Amazing! What a useful and pratical skill when computation is not enough! Thanks and Congrats!",
      "votes": null
    },
    {
      "id": "1634557",
      "postDate": "12/31/2021 19:51:38",
      "content": "<p>Thanks for your sharing!</p>",
      "rawMarkdown": "Thanks for your sharing!",
      "votes": null
    },
    {
      "id": "1634710",
      "postDate": "01/01/2022 04:10:34",
      "content": "<p>Congrats to u!<br>\nI involved small objects removing too, based on the prior distribution of the dataset, which actullay doesn't help too much as cellpose predicts very little FP ones. How do u do that?</p>",
      "rawMarkdown": "Congrats to u!\nI involved small objects removing too, based on the prior distribution of the dataset, which actullay doesn't help too much as cellpose predicts very little FP ones. How do u do that?",
      "votes": null
    },
    {
      "id": "1634728",
      "postDate": "01/01/2022 04:57:29",
      "content": "<p>Hi, I know you are disappointed, but your approach is truly novel, how you worked with effnets, and seeing your code shows how much hard work have you put.<br>\nI was also using mmdet, but using mostly inbuilt things of mmdet. I am surprised you can customise so many things. Do you refer something to learn how to do these customisations?<br>\nAlso I checked your GitHub ,I saw your custom effnets, custom head, but where do you apply your FPN? is this inbuilt somewhere I missed?</p>",
      "rawMarkdown": "Hi, I know you are disappointed, but your approach is truly novel, how you worked with effnets, and seeing your code shows how much hard work have you put.\nI was also using mmdet, but using mostly inbuilt things of mmdet. I am surprised you can customise so many things. Do you refer something to learn how to do these customisations?\nAlso I checked your GitHub ,I saw your custom effnets, custom head, but where do you apply your FPN? is this inbuilt somewhere I missed?",
      "votes": null
    },
    {
      "id": "1634818",
      "postDate": "01/01/2022 08:01:57",
      "content": "<p>Great work! Congrats on the results <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and team.</p>",
      "rawMarkdown": "Great work! Congrats on the results @theoviel and team.",
      "votes": null
    },
    {
      "id": "1635214",
      "postDate": "01/01/2022 15:34:09",
      "content": "<p>Congratulations !great work</p>",
      "rawMarkdown": "Congratulations !great work",
      "votes": null
    },
    {
      "id": "1635289",
      "postDate": "01/01/2022 16:46:00",
      "content": "<p>Really awesome work! Congrats. </p>",
      "rawMarkdown": "Really awesome work! Congrats.",
      "votes": null
    },
    {
      "id": "1635355",
      "postDate": "01/01/2022 17:35:58",
      "content": "<p>Thanks for the kind words <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> !<br>\nModels are defined in the configs and built by mmdet, FPN is defined here.<br>\nRegarding customisations you can refer to the mmdet doc : <a href=\"https://mmdetection.readthedocs.io/en/latest/\" target=\"_blank\">https://mmdetection.readthedocs.io/en/latest/</a> - and try looking at the code </p>",
      "rawMarkdown": "Thanks for the kind words @mrinath !\nModels are defined in the configs and built by mmdet, FPN is defined here.\nRegarding customisations you can refer to the mmdet doc : https://mmdetection.readthedocs.io/en/latest/ - and try looking at the code",
      "votes": null
    },
    {
      "id": "1635356",
      "postDate": "01/01/2022 17:37:19",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/lucamassaron\" target=\"_blank\">@lucamassaron</a> ! </p>",
      "rawMarkdown": "Thanks a lot @lucamassaron !",
      "votes": null
    },
    {
      "id": "1635357",
      "postDate": "01/01/2022 17:37:42",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> and congratz to your team as well !</p>",
      "rawMarkdown": "Thanks @duykhanh99 and congratz to your team as well !",
      "votes": null
    },
    {
      "id": "1635368",
      "postDate": "01/01/2022 17:46:47",
      "content": "<p>Congratulations !keep working</p>",
      "rawMarkdown": "Congratulations !keep working",
      "votes": null
    },
    {
      "id": "1637748",
      "postDate": "01/04/2022 06:52:42",
      "content": "<p>Really great result for you especially with single RTX 2080 Ti. This competition is so competitive you will be top1 next time. Congrats!</p>",
      "rawMarkdown": "Really great result for you especially with single RTX 2080 Ti. This competition is so competitive you will be top1 next time. Congrats!",
      "votes": null
    },
    {
      "id": "1638234",
      "postDate": "01/04/2022 14:59:30",
      "content": "<p>Don't be sorry for moving ut of the top place you held for so long in the competition. Think about the positive: you get a gold still (congrats!), and you probably learned something.</p>\n<p>Well, this is what I told myself after being in a similar position in last birdsong competition.</p>",
      "rawMarkdown": "Don't be sorry for moving ut of the top place you held for so long in the competition. Think about the positive: you get a gold still (congrats!), and you probably learned something.\n\nWell, this is what I told myself after being in a similar position in last birdsong competition.",
      "votes": null
    },
    {
      "id": "1638277",
      "postDate": "01/04/2022 15:30:23",
      "content": "<p>great work !</p>",
      "rawMarkdown": "great work !",
      "votes": null
    },
    {
      "id": "1638445",
      "postDate": "01/04/2022 18:24:18",
      "content": "<p>pretty useful </p>",
      "rawMarkdown": "pretty useful",
      "votes": null
    },
    {
      "id": "1640077",
      "postDate": "01/06/2022 06:50:04",
      "content": "<p>Congratulations !! Great work.</p>",
      "rawMarkdown": "Congratulations !! Great work.",
      "votes": null
    },
    {
      "id": "1687997",
      "postDate": "02/13/2022 09:50:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> congratulations on the gold medal!</p>\n<p>I've been trying to run <code>notebooks/Preparation.ipynb</code> in your code repository, but I got the warning <code>global /io/opencv/modules/imgcodecs/src/loadsave.cpp (239) findDecoder imread_('../input/hckfix/.png'):  can't open/read file: check file path/integrity</code>. I followed all the steps written in the repository. Although this warning doesn't stop executing the code, any ideas why it is happening?</p>\n<p>Also, how long are you expecting <code>Preparation.ipynb</code> to finish running? I'm running it on a VM of 16 vCPU and 60 GB of RAM, and it's been going for 24 hours, still not finished.</p>\n<p>Thank you in advance for your time!</p>",
      "rawMarkdown": "Hi @theoviel congratulations on the gold medal!\n\nI've been trying to run `notebooks/Preparation.ipynb` in your code repository, but I got the warning `global /io/opencv/modules/imgcodecs/src/loadsave.cpp (239) findDecoder imread_('../input/hckfix/.png'):  can't open/read file: check file path/integrity`. I followed all the steps written in the repository. Although this warning doesn't stop executing the code, any ideas why it is happening?\n\nAlso, how long are you expecting `Preparation.ipynb` to finish running? I'm running it on a VM of 16 vCPU and 60 GB of RAM, and it's been going for 24 hours, still not finished.\n\nThank you in advance for your time!",
      "votes": null
    },
    {
      "id": "1688031",
      "postDate": "02/13/2022 10:33:45",
      "content": "<p>The script should run quite quickly (&lt;1h) as far as I remember.<br>\nIt uses multiprocessing so you may want to turn that off.</p>\n<p>Regarding the error, you actually need to download the corrected masks here : <br>\n<a href=\"https://www.kaggle.com/hengck23/clean-astro-mask\" target=\"_blank\">https://www.kaggle.com/hengck23/clean-astro-mask</a> and put them in the <code>HCK_FIX_PATH</code> folder. I forgot to include this in the ReadMe, sorry about that</p>",
      "rawMarkdown": "The script should run quite quickly (<1h) as far as I remember.\nIt uses multiprocessing so you may want to turn that off.\n\nRegarding the error, you actually need to download the corrected masks here : \nhttps://www.kaggle.com/hengck23/clean-astro-mask and put them in the `HCK_FIX_PATH` folder. I forgot to include this in the ReadMe, sorry about that",
      "votes": null
    },
    {
      "id": "1688214",
      "postDate": "02/13/2022 13:35:22",
      "content": "<p>Thank you for replying, I'll give that a try! Again, congrats on the gold medal, and good luck in future competitions!</p>",
      "rawMarkdown": "Thank you for replying, I'll give that a try! Again, congrats on the gold medal, and good luck in future competitions!",
      "votes": null
    },
    {
      "id": "1688798",
      "postDate": "02/13/2022 20:32:13",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> congrats on the results of the competition! Just wondering, what specs is your pc for training models for this competition? You mentioned that it has RTX 2080 Ti, but how much RAM and how many cores for CPU?</p>\n<p>I've been trying to run the code in the repo, but everything is taking too long. Also, do I need GPU for running <code>Preparation.ipynb</code>?</p>",
      "rawMarkdown": "theoviel congrats on the results of the competition! Just wondering, what specs is your pc for training models for this competition? You mentioned that it has RTX 2080 Ti, but how much RAM and how many cores for CPU?\n\nI've been trying to run the code in the repo, but everything is taking too long. Also, do I need GPU for running `Preparation.ipynb`?",
      "votes": null
    },
    {
      "id": "1688847",
      "postDate": "02/13/2022 21:07:50",
      "content": "<p>My CPU is a 12-Core AMD Ryzen 9 3900XT - which is probably the reason why computations are fast on my setup. You don't need to use the GPU and I believe RAM is not an issue as well. <br>\nHaving a lot of RAM is required for inference though but this can be fixed by optimizing the code a bit. I sometimes ran into oom errors even with 64 Gb.</p>\n<p>The reason it is taking too long might be multiprocessing, you can also turn the <code>FIX</code> parameter to False (this shouldn't affect performance too much)</p>",
      "rawMarkdown": "My CPU is a 12-Core AMD Ryzen 9 3900XT - which is probably the reason why computations are fast on my setup. You don't need to use the GPU and I believe RAM is not an issue as well. \nHaving a lot of RAM is required for inference though but this can be fixed by optimizing the code a bit. I sometimes ran into oom errors even with 64 Gb.\n\nThe reason it is taking too long might be multiprocessing, you can also turn the `FIX` parameter to False (this shouldn't affect performance too much)",
      "votes": null
    },
    {
      "id": "1688851",
      "postDate": "02/13/2022 21:11:10",
      "content": "<p>Ok, how do I disable multiprocessing? Do I just change processes to 1 in <code>p = Pool(processes=1)</code>?</p>\n<p>Also (forgive me for asking so many questions!), how long are <code>Livecell.ipynb</code> and <code>Training.ipynb</code> supposed to take to run? I have 1 V100, would that be enough?</p>",
      "rawMarkdown": "Ok, how do I disable multiprocessing? Do I just change processes to 1 in `p = Pool(processes=1)`?\n\nAlso (forgive me for asking so many questions!), how long are `Livecell.ipynb` and `Training.ipynb` supposed to take to run? I have 1 V100, would that be enough?",
      "votes": null
    },
    {
      "id": "1688856",
      "postDate": "02/13/2022 21:14:11",
      "content": "<p>I'm not sure if using 1 process works, if it doesn't try replacing :</p>\n<pre><code>metas = []\nfor _, meta in tqdm(p.imap(prepare_mmdet_data_, range(len(df))), total=len(df)):\n    metas.append(meta)\n</code></pre>\n<p>With :</p>\n<pre><code>metas = []\nfor i in tqdm(range(len(df))):\n    metas.append(prepare_mmdet_data_(i))\n</code></pre>",
      "rawMarkdown": "I'm not sure if using 1 process works, if it doesn't try replacing :\n\n```\nmetas = []\nfor _, meta in tqdm(p.imap(prepare_mmdet_data_, range(len(df))), total=len(df)):\n    metas.append(meta)\n```\n\nWith :\n```\nmetas = []\nfor i in tqdm(range(len(df))):\n    metas.append(prepare_mmdet_data_(i))\n```",
      "votes": null
    },
    {
      "id": "1688928",
      "postDate": "02/13/2022 22:12:13",
      "content": "<p>ok, thanks very much, good luck with your next competition!!</p>",
      "rawMarkdown": "ok, thanks very much, good luck with your next competition!!",
      "votes": null
    },
    {
      "id": "1690310",
      "postDate": "02/14/2022 21:10:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, I downloaded the corrected masks, and everything worked like a charm! Thank you for your help yesterday!</p>\n<p>I'm now trying to run <code>Livecell.ipynb</code>, and I ran into another problem. <code>df = prepare_mmdet_data(name=\"livecell_shsy5y\")</code> returns <code>FileNotFoundError: [Errno 2] No such file or directory: '../output/livecell_shsy5y.csv'</code>.</p>\n<p>Any ideas for how to get <code>livecell_shsy5y.csv</code> file? Thanks in advance.</p>",
      "rawMarkdown": "Hi @theoviel, I downloaded the corrected masks, and everything worked like a charm! Thank you for your help yesterday!\n\nI'm now trying to run `Livecell.ipynb`, and I ran into another problem. `df = prepare_mmdet_data(name=\"livecell_shsy5y\")` returns `FileNotFoundError: [Errno 2] No such file or directory: '../output/livecell_shsy5y.csv'`.\n\nAny ideas for how to get `livecell_shsy5y.csv` file? Thanks in advance.",
      "votes": null
    },
    {
      "id": "1690320",
      "postDate": "02/14/2022 21:31:13",
      "content": "<p>The csv is computed in the first part of the notebook, you need to comment the line :<br>\n<code>annotations = []  # do not recompute</code></p>\n<p>The following cell will create the csvs depending on the values of <code>SHSY5Y_ONLY</code> and  <code>NO_SHSY5Y</code>. Note that the dataframe named <code>livecell.csv</code> is used for pretraining but you can change that in the first line of the pretrain function. </p>\n<p>Please use : </p>\n<pre><code>SHSY5Y_ONLY = False\nNO_SHSY5Y = False\nSINGLE_CLASS = False\n</code></pre>\n<p>to generate <code>livecell.csv</code></p>",
      "rawMarkdown": "The csv is computed in the first part of the notebook, you need to comment the line :\n`annotations = []  # do not recompute`\n\nThe following cell will create the csvs depending on the values of `SHSY5Y_ONLY` and  `NO_SHSY5Y`. Note that the dataframe named `livecell.csv` is used for pretraining but you can change that in the first line of the pretrain function. \n\nPlease use : \n```\nSHSY5Y_ONLY = False\nNO_SHSY5Y = False\nSINGLE_CLASS = False\n```\nto generate `livecell.csv`",
      "votes": null
    },
    {
      "id": "1690365",
      "postDate": "02/14/2022 22:26:50",
      "content": "<p>I've updated the repository and uploaded a script that was actually necessary to run the inference. There's a script to replace in the mmdet package otherwise you'll get an error, please check the ReadMe !<br>\nThanks for taking the time to run the code, this forces me to fix minor issues that should've been corrected much earlier.</p>",
      "rawMarkdown": "I've updated the repository and uploaded a script that was actually necessary to run the inference. There's a script to replace in the mmdet package otherwise you'll get an error, please check the ReadMe !\nThanks for taking the time to run the code, this forces me to fix minor issues that should've been corrected much earlier.",
      "votes": null
    },
    {
      "id": "1690394",
      "postDate": "02/14/2022 23:26:59",
      "content": "<p>Thanks for your time, Théo! It's great to learn from one of the top Kaggle Grandmasters!</p>",
      "rawMarkdown": "Thanks for your time, Théo! It's great to learn from one of the top Kaggle Grandmasters!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1634078,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "12/31/2021 10:41:25",
      "content": "<p>Great work! Congrats on results <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and team. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1635357,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "01/01/2022 17:37:42",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> and congratz to your team as well !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1634160,
      "author_name": "lucamassaron",
      "author_url": "",
      "post_date": "12/31/2021 11:35:41",
      "content": "<p>Still a great result <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>! In spite of the slightly worse private leaderboard position than you expected, I loved reading about your cv strategy \"splitting on plate &amp; well as done in the Livecell paper\". A very smart indeed strategy, well reasoned and grounded.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1635356,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "01/01/2022 17:37:19",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/lucamassaron\" target=\"_blank\">@lucamassaron</a> ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1634509,
      "author_name": "charonwangg",
      "author_url": "",
      "post_date": "12/31/2021 18:34:26",
      "content": "<p>Amazing! What a useful and pratical skill when computation is not enough! Thanks and Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1634557,
      "author_name": "thalesferraz",
      "author_url": "",
      "post_date": "12/31/2021 19:51:38",
      "content": "<p>Thanks for your sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1634710,
      "author_name": "yichenwang1988",
      "author_url": "",
      "post_date": "01/01/2022 04:10:34",
      "content": "<p>Congrats to u!<br>\nI involved small objects removing too, based on the prior distribution of the dataset, which actullay doesn't help too much as cellpose predicts very little FP ones. How do u do that?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1634728,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "01/01/2022 04:57:29",
      "content": "<p>Hi, I know you are disappointed, but your approach is truly novel, how you worked with effnets, and seeing your code shows how much hard work have you put.<br>\nI was also using mmdet, but using mostly inbuilt things of mmdet. I am surprised you can customise so many things. Do you refer something to learn how to do these customisations?<br>\nAlso I checked your GitHub ,I saw your custom effnets, custom head, but where do you apply your FPN? is this inbuilt somewhere I missed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1635355,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "01/01/2022 17:35:58",
          "content": "<p>Thanks for the kind words <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> !<br>\nModels are defined in the configs and built by mmdet, FPN is defined here.<br>\nRegarding customisations you can refer to the mmdet doc : <a href=\"https://mmdetection.readthedocs.io/en/latest/\" target=\"_blank\">https://mmdetection.readthedocs.io/en/latest/</a> - and try looking at the code </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1634818,
      "author_name": "fredericstudio",
      "author_url": "",
      "post_date": "01/01/2022 08:01:57",
      "content": "<p>Great work! Congrats on the results <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and team.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1635214,
      "author_name": "balavashan",
      "author_url": "",
      "post_date": "01/01/2022 15:34:09",
      "content": "<p>Congratulations !great work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1635289,
      "author_name": "hsadeghian",
      "author_url": "",
      "post_date": "01/01/2022 16:46:00",
      "content": "<p>Really awesome work! Congrats. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1635368,
      "author_name": "aiswaryasivakumar",
      "author_url": "",
      "post_date": "01/01/2022 17:46:47",
      "content": "<p>Congratulations !keep working</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1637748,
      "author_name": "mcggood",
      "author_url": "",
      "post_date": "01/04/2022 06:52:42",
      "content": "<p>Really great result for you especially with single RTX 2080 Ti. This competition is so competitive you will be top1 next time. Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1638234,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/04/2022 14:59:30",
      "content": "<p>Don't be sorry for moving ut of the top place you held for so long in the competition. Think about the positive: you get a gold still (congrats!), and you probably learned something.</p>\n<p>Well, this is what I told myself after being in a similar position in last birdsong competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1638277,
      "author_name": "sunilchoudhary1",
      "author_url": "",
      "post_date": "01/04/2022 15:30:23",
      "content": "<p>great work !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1638445,
      "author_name": "samymehdid",
      "author_url": "",
      "post_date": "01/04/2022 18:24:18",
      "content": "<p>pretty useful </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1640077,
      "author_name": "mdsufyan",
      "author_url": "",
      "post_date": "01/06/2022 06:50:04",
      "content": "<p>Congratulations !! Great work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1687997,
      "author_name": "coderrexe",
      "author_url": "",
      "post_date": "02/13/2022 09:50:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> congratulations on the gold medal!</p>\n<p>I've been trying to run <code>notebooks/Preparation.ipynb</code> in your code repository, but I got the warning <code>global /io/opencv/modules/imgcodecs/src/loadsave.cpp (239) findDecoder imread_('../input/hckfix/.png'):  can't open/read file: check file path/integrity</code>. I followed all the steps written in the repository. Although this warning doesn't stop executing the code, any ideas why it is happening?</p>\n<p>Also, how long are you expecting <code>Preparation.ipynb</code> to finish running? I'm running it on a VM of 16 vCPU and 60 GB of RAM, and it's been going for 24 hours, still not finished.</p>\n<p>Thank you in advance for your time!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1688031,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/13/2022 10:33:45",
          "content": "<p>The script should run quite quickly (&lt;1h) as far as I remember.<br>\nIt uses multiprocessing so you may want to turn that off.</p>\n<p>Regarding the error, you actually need to download the corrected masks here : <br>\n<a href=\"https://www.kaggle.com/hengck23/clean-astro-mask\" target=\"_blank\">https://www.kaggle.com/hengck23/clean-astro-mask</a> and put them in the <code>HCK_FIX_PATH</code> folder. I forgot to include this in the ReadMe, sorry about that</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688214,
          "author_name": "coderrexe",
          "author_url": "",
          "post_date": "02/13/2022 13:35:22",
          "content": "<p>Thank you for replying, I'll give that a try! Again, congrats on the gold medal, and good luck in future competitions!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690310,
          "author_name": "coderrexe",
          "author_url": "",
          "post_date": "02/14/2022 21:10:02",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, I downloaded the corrected masks, and everything worked like a charm! Thank you for your help yesterday!</p>\n<p>I'm now trying to run <code>Livecell.ipynb</code>, and I ran into another problem. <code>df = prepare_mmdet_data(name=\"livecell_shsy5y\")</code> returns <code>FileNotFoundError: [Errno 2] No such file or directory: '../output/livecell_shsy5y.csv'</code>.</p>\n<p>Any ideas for how to get <code>livecell_shsy5y.csv</code> file? Thanks in advance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690320,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/14/2022 21:31:13",
          "content": "<p>The csv is computed in the first part of the notebook, you need to comment the line :<br>\n<code>annotations = []  # do not recompute</code></p>\n<p>The following cell will create the csvs depending on the values of <code>SHSY5Y_ONLY</code> and  <code>NO_SHSY5Y</code>. Note that the dataframe named <code>livecell.csv</code> is used for pretraining but you can change that in the first line of the pretrain function. </p>\n<p>Please use : </p>\n<pre><code>SHSY5Y_ONLY = False\nNO_SHSY5Y = False\nSINGLE_CLASS = False\n</code></pre>\n<p>to generate <code>livecell.csv</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690365,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/14/2022 22:26:50",
          "content": "<p>I've updated the repository and uploaded a script that was actually necessary to run the inference. There's a script to replace in the mmdet package otherwise you'll get an error, please check the ReadMe !<br>\nThanks for taking the time to run the code, this forces me to fix minor issues that should've been corrected much earlier.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690394,
          "author_name": "coderrexe",
          "author_url": "",
          "post_date": "02/14/2022 23:26:59",
          "content": "<p>Thanks for your time, Théo! It's great to learn from one of the top Kaggle Grandmasters!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1688798,
      "author_name": "jgeoff",
      "author_url": "",
      "post_date": "02/13/2022 20:32:13",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> congrats on the results of the competition! Just wondering, what specs is your pc for training models for this competition? You mentioned that it has RTX 2080 Ti, but how much RAM and how many cores for CPU?</p>\n<p>I've been trying to run the code in the repo, but everything is taking too long. Also, do I need GPU for running <code>Preparation.ipynb</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1688847,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/13/2022 21:07:50",
          "content": "<p>My CPU is a 12-Core AMD Ryzen 9 3900XT - which is probably the reason why computations are fast on my setup. You don't need to use the GPU and I believe RAM is not an issue as well. <br>\nHaving a lot of RAM is required for inference though but this can be fixed by optimizing the code a bit. I sometimes ran into oom errors even with 64 Gb.</p>\n<p>The reason it is taking too long might be multiprocessing, you can also turn the <code>FIX</code> parameter to False (this shouldn't affect performance too much)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688851,
          "author_name": "jgeoff",
          "author_url": "",
          "post_date": "02/13/2022 21:11:10",
          "content": "<p>Ok, how do I disable multiprocessing? Do I just change processes to 1 in <code>p = Pool(processes=1)</code>?</p>\n<p>Also (forgive me for asking so many questions!), how long are <code>Livecell.ipynb</code> and <code>Training.ipynb</code> supposed to take to run? I have 1 V100, would that be enough?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688856,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/13/2022 21:14:11",
          "content": "<p>I'm not sure if using 1 process works, if it doesn't try replacing :</p>\n<pre><code>metas = []\nfor _, meta in tqdm(p.imap(prepare_mmdet_data_, range(len(df))), total=len(df)):\n    metas.append(meta)\n</code></pre>\n<p>With :</p>\n<pre><code>metas = []\nfor i in tqdm(range(len(df))):\n    metas.append(prepare_mmdet_data_(i))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688928,
          "author_name": "jgeoff",
          "author_url": "",
          "post_date": "02/13/2022 22:12:13",
          "content": "<p>ok, thanks very much, good luck with your next competition!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1634030": "First of all congratz to the winners and thanks to Sartorius and Kaggle for hosting the competition !\n\nAlthough this 11th place is a great finish, we're a bit disappointed since we spent a month at the top and the final weeks in the top 5. We knew we were not gonna win but this shake was quite a surprise for us, and we still haven't really understood why it happened. \n\nOur models were just better on Public than Private, there might be a small domain shift (more corts ? more small cells ? more large cells ? we will never now...) that emphasized a weakness of our pipeline.\n\nAnyways, here are few interesting points of the solution that I found were worth mentioning, the full code is on GitHub : https://github.com/TheoViel/kaggle_sartorius \nAlso, our inference code is here : https://www.kaggle.com/theoviel/sartorius-inference-final\n\n### Validation Scheme\n\nTwo weeks before the end of the competition, we switched to a validation set-up we thought would be reliable : splitting on plate & well as done in the Livecell paper. \n\n![](https://i.ibb.co/dDjvC6j/livecell-3.png)\n\nThis might be one of the reason we shake down ? Our best private LB (0.349, public 0.339) was our best CV before switching to the above scheme. Still, our final score of 0.348 is our best CV with this setup (public 0.342)\n\n### Models\n\nWe only used machines with single RTX 2080 Ti so we had to be ingenious to be able to train on high resolution images and detect small cells. We used mask-rcnn based models and relied on mmdet but only for model definition and augmentations, the rest of the pipeline is hand-crafted, which made it more convenient for experimenting.\n\n\n![](https://i.ibb.co/MkwRbVW/Ensembling-drawio-1.png)\n\n##### Main points\n- Remove the stride of the first layer of the encoder to \"increase the resolution\" of the models without doing any resizing !\n- Random Crops of size 256x256 for training\n- Pretrain on Livecell\n- 4000 iterations of finetuning on the training data\n- Backbones : resnet50, resnext101_32x4, resnext101_64x4, efficientnet_b4/b5/b6\n- Models : MaskRCNN, Cascade, HTC\n\n##### Ensembling\n\nWe average predictions of different models & different flips at three stages, the stages are the boxes with thicker borders above.\n- Proposals : For a given feature map output by the FPN, each of its pixel is assigned a score and a coordinates prediction by the convolutions. This is what we average. \n- Boxes : We re-use the ensembled proposal and perform averaging of the class predictions and coordinates for each proposal. We used 4 flip TTAs.\n- Masks : starting with the ensembled boxes, we average the masks - before the upsampling back to the original image size.\n\nThis scheme doesn't really use NMS for ensembling which can be tricky to use. Hence we stacked a bunch of models. We used 6 models per cell type.\n\n##### Post processing\n\n- NMS on boxes using high thresholds, then NMS on masks using low thresholds\n- Corrupt back the astro masks as we trained on clean ones (+0.002 LB)\n- Small masks removal\n\nWe did a lot of hyper-parameters tweaking on CV : NMS thresholds, RPN and bbox_head params, confidence thresholds, minimum cell sizes, mask thresholds.\n\n### Few more words\n\n- Pseudo Labelling didn't really work for us, we used them in the finetuning phase with the original training data and progressively decayed their proportion.\n- The RoiAlign layer from mmdet has implementation issues. Masks resulting from TTA appear shifted which hurt performances, especially when trying to use vertical flips. We had to shift the boxes by 0.5 to counter this. \n\n I will probably add more stuff later, and fix the typos and all. Feel free to ask any questions  !\n*Thanks for reading !*",
    "1634078": "Great work! Congrats on results @theoviel and team.",
    "1634160": "Still a great result @theoviel! In spite of the slightly worse private leaderboard position than you expected, I loved reading about your cv strategy \"splitting on plate & well as done in the Livecell paper\". A very smart indeed strategy, well reasoned and grounded.",
    "1634509": "Amazing! What a useful and pratical skill when computation is not enough! Thanks and Congrats!",
    "1634557": "Thanks for your sharing!",
    "1634710": "Congrats to u!\nI involved small objects removing too, based on the prior distribution of the dataset, which actullay doesn't help too much as cellpose predicts very little FP ones. How do u do that?",
    "1634728": "Hi, I know you are disappointed, but your approach is truly novel, how you worked with effnets, and seeing your code shows how much hard work have you put.\nI was also using mmdet, but using mostly inbuilt things of mmdet. I am surprised you can customise so many things. Do you refer something to learn how to do these customisations?\nAlso I checked your GitHub ,I saw your custom effnets, custom head, but where do you apply your FPN? is this inbuilt somewhere I missed?",
    "1634818": "Great work! Congrats on the results @theoviel and team.",
    "1635214": "Congratulations !great work",
    "1635289": "Really awesome work! Congrats.",
    "1635355": "Thanks for the kind words @mrinath !\nModels are defined in the configs and built by mmdet, FPN is defined here.\nRegarding customisations you can refer to the mmdet doc : https://mmdetection.readthedocs.io/en/latest/ - and try looking at the code",
    "1635356": "Thanks a lot @lucamassaron !",
    "1635357": "Thanks @duykhanh99 and congratz to your team as well !",
    "1635368": "Congratulations !keep working",
    "1637748": "Really great result for you especially with single RTX 2080 Ti. This competition is so competitive you will be top1 next time. Congrats!",
    "1638234": "Don't be sorry for moving ut of the top place you held for so long in the competition. Think about the positive: you get a gold still (congrats!), and you probably learned something.\n\nWell, this is what I told myself after being in a similar position in last birdsong competition.",
    "1638277": "great work !",
    "1638445": "pretty useful",
    "1640077": "Congratulations !! Great work.",
    "1687997": "Hi @theoviel congratulations on the gold medal!\n\nI've been trying to run `notebooks/Preparation.ipynb` in your code repository, but I got the warning `global /io/opencv/modules/imgcodecs/src/loadsave.cpp (239) findDecoder imread_('../input/hckfix/.png'):  can't open/read file: check file path/integrity`. I followed all the steps written in the repository. Although this warning doesn't stop executing the code, any ideas why it is happening?\n\nAlso, how long are you expecting `Preparation.ipynb` to finish running? I'm running it on a VM of 16 vCPU and 60 GB of RAM, and it's been going for 24 hours, still not finished.\n\nThank you in advance for your time!",
    "1688031": "The script should run quite quickly (<1h) as far as I remember.\nIt uses multiprocessing so you may want to turn that off.\n\nRegarding the error, you actually need to download the corrected masks here : \nhttps://www.kaggle.com/hengck23/clean-astro-mask and put them in the `HCK_FIX_PATH` folder. I forgot to include this in the ReadMe, sorry about that",
    "1688214": "Thank you for replying, I'll give that a try! Again, congrats on the gold medal, and good luck in future competitions!",
    "1688798": "theoviel congrats on the results of the competition! Just wondering, what specs is your pc for training models for this competition? You mentioned that it has RTX 2080 Ti, but how much RAM and how many cores for CPU?\n\nI've been trying to run the code in the repo, but everything is taking too long. Also, do I need GPU for running `Preparation.ipynb`?",
    "1688847": "My CPU is a 12-Core AMD Ryzen 9 3900XT - which is probably the reason why computations are fast on my setup. You don't need to use the GPU and I believe RAM is not an issue as well. \nHaving a lot of RAM is required for inference though but this can be fixed by optimizing the code a bit. I sometimes ran into oom errors even with 64 Gb.\n\nThe reason it is taking too long might be multiprocessing, you can also turn the `FIX` parameter to False (this shouldn't affect performance too much)",
    "1688851": "Ok, how do I disable multiprocessing? Do I just change processes to 1 in `p = Pool(processes=1)`?\n\nAlso (forgive me for asking so many questions!), how long are `Livecell.ipynb` and `Training.ipynb` supposed to take to run? I have 1 V100, would that be enough?",
    "1688856": "I'm not sure if using 1 process works, if it doesn't try replacing :\n\n```\nmetas = []\nfor _, meta in tqdm(p.imap(prepare_mmdet_data_, range(len(df))), total=len(df)):\n    metas.append(meta)\n```\n\nWith :\n```\nmetas = []\nfor i in tqdm(range(len(df))):\n    metas.append(prepare_mmdet_data_(i))\n```",
    "1688928": "ok, thanks very much, good luck with your next competition!!",
    "1690310": "Hi @theoviel, I downloaded the corrected masks, and everything worked like a charm! Thank you for your help yesterday!\n\nI'm now trying to run `Livecell.ipynb`, and I ran into another problem. `df = prepare_mmdet_data(name=\"livecell_shsy5y\")` returns `FileNotFoundError: [Errno 2] No such file or directory: '../output/livecell_shsy5y.csv'`.\n\nAny ideas for how to get `livecell_shsy5y.csv` file? Thanks in advance.",
    "1690320": "The csv is computed in the first part of the notebook, you need to comment the line :\n`annotations = []  # do not recompute`\n\nThe following cell will create the csvs depending on the values of `SHSY5Y_ONLY` and  `NO_SHSY5Y`. Note that the dataframe named `livecell.csv` is used for pretraining but you can change that in the first line of the pretrain function. \n\nPlease use : \n```\nSHSY5Y_ONLY = False\nNO_SHSY5Y = False\nSINGLE_CLASS = False\n```\nto generate `livecell.csv`",
    "1690365": "I've updated the repository and uploaded a script that was actually necessary to run the inference. There's a script to replace in the mmdet package otherwise you'll get an error, please check the ReadMe !\nThanks for taking the time to run the code, this forces me to fix minor issues that should've been corrected much earlier.",
    "1690394": "Thanks for your time, Théo! It's great to learn from one of the top Kaggle Grandmasters!"
  },
  "source": "meta"
}