{
  "id": 198382,
  "title": "Pytorch/fast.ai starter [0.905 LB]",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/198382",
  "author_name": "",
  "post_date": "2020-11-20T23:44:31.951531400Z",
  "votes": 80,
  "comment_count": 55,
  "views": 0,
  "content": "<p>Welcome to Human BioMolecular Atlas Program (HuBMAP) competition. I have prepared several kernels for people who use Pytorch and fast.ai to get started with this challenge and for others to borrow some ideas and explore them more. </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a> : the kernel used to cut the original images into tiles and save them to a dataset</li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-fast-ai-starter\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-fast-ai-starter</a> : the training kernel sharing my experience in the previous segmentation competitions and giving a Pytorch/fast.ai pipeline to start with</li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-fast-ai-starter-sub\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-fast-ai-starter-sub</a> : the starter inference kernel that got 0.905 LB on 4 times lower resolution images (256x256 tile dataset)</li>\n</ul>\n<p>I hope you will enjoy this competition, and I wish you the best luck. <br>\n03/13/2021 The datasets and kernels are updated to accommodate for the new data.</p>",
  "messages": [
    {
      "id": "1085502",
      "postDate": "11/20/2020 23:44:31",
      "content": "<p>Welcome to Human BioMolecular Atlas Program (HuBMAP) competition. I have prepared several kernels for people who use Pytorch and fast.ai to get started with this challenge and for others to borrow some ideas and explore them more. </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a> : the kernel used to cut the original images into tiles and save them to a dataset</li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-fast-ai-starter\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-fast-ai-starter</a> : the training kernel sharing my experience in the previous segmentation competitions and giving a Pytorch/fast.ai pipeline to start with</li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-fast-ai-starter-sub\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-fast-ai-starter-sub</a> : the starter inference kernel that got 0.905 LB on 4 times lower resolution images (256x256 tile dataset)</li>\n</ul>\n<p>I hope you will enjoy this competition, and I wish you the best luck. <br>\n03/13/2021 The datasets and kernels are updated to accommodate for the new data.</p>",
      "rawMarkdown": "Welcome to Human BioMolecular Atlas Program (HuBMAP) competition. I have prepared several kernels for people who use Pytorch and fast.ai to get started with this challenge and for others to borrow some ideas and explore them more. \n- https://www.kaggle.com/iafoss/256x256-images : the kernel used to cut the original images into tiles and save them to a dataset\n- https://www.kaggle.com/iafoss/hubmap-fast-ai-starter : the training kernel sharing my experience in the previous segmentation competitions and giving a Pytorch/fast.ai pipeline to start with\n-  https://www.kaggle.com/iafoss/hubmap-fast-ai-starter-sub : the starter inference kernel that got 0.905 LB on 4 times lower resolution images (256x256 tile dataset)\n\nI hope you will enjoy this competition, and I wish you the best luck. \n03/13/2021 The datasets and kernels are updated to accommodate for the new data.",
      "votes": null
    },
    {
      "id": "1085545",
      "postDate": "11/21/2020 00:50:20",
      "content": "<p>Very well done, thanks for sharing your pipeline</p>",
      "rawMarkdown": "Very well done, thanks for sharing your pipeline",
      "votes": null
    },
    {
      "id": "1085710",
      "postDate": "11/21/2020 06:18:39",
      "content": "<p>You are very welcome</p>",
      "rawMarkdown": "You are very welcome",
      "votes": null
    },
    {
      "id": "1085798",
      "postDate": "11/21/2020 07:54:30",
      "content": "<p>I have problems with inference to private test sets. I suspect it's because the private test set contains an image that's too big to read using the tifffile. I think we need to use a library that can load subregions of the image, such as pyvips. I would like to wait for an update of the kaggle environment.</p>",
      "rawMarkdown": "I have problems with inference to private test sets. I suspect it's because the private test set contains an image that's too big to read using the tifffile. I think we need to use a library that can load subregions of the image, such as pyvips. I would like to wait for an update of the kaggle environment.",
      "votes": null
    },
    {
      "id": "1085805",
      "postDate": "11/21/2020 08:00:23",
      "content": "<p>It is likely to be the case but the strange thing is that the notebook doesn't exit with \"out of memory\" error but rather runs to 9h limit and stops (actually it is likely trying to restart and face the same problem). Difficult to debug it without proper error message output.</p>",
      "rawMarkdown": "It is likely to be the case but the strange thing is that the notebook doesn't exit with \"out of memory\" error but rather runs to 9h limit and stops (actually it is likely trying to restart and face the same problem). Difficult to debug it without proper error message output.",
      "votes": null
    },
    {
      "id": "1085835",
      "postDate": "11/21/2020 08:16:00",
      "content": "<p>I have same problem. If an \"out of memory\" error occurs, there is case that it may run until the time expires instead of exiting immediately.</p>",
      "rawMarkdown": "I have same problem. If an \"out of memory\" error occurs, there is case that it may run until the time expires instead of exiting immediately.",
      "votes": null
    },
    {
      "id": "1086321",
      "postDate": "11/21/2020 15:32:40",
      "content": "<p>I've done successful commit on CPU instead of GPU which gives you +3 GB of RAM, so seems all is fine with file reading but problem with all memory usage.</p>",
      "rawMarkdown": "I've done successful commit on CPU instead of GPU which gives you +3 GB of RAM, so seems all is fine with file reading but problem with all memory usage.",
      "votes": null
    },
    {
      "id": "1086323",
      "postDate": "11/21/2020 15:34:44",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <br>\nSince they provided anatomical structure annotations in the public test set do you think they are also available in the private test set?</p>",
      "rawMarkdown": "Thanks for sharing @iafoss \nSince they provided anatomical structure annotations in the public test set do you think they are also available in the private test set?",
      "votes": null
    },
    {
      "id": "1086478",
      "postDate": "11/21/2020 18:03:58",
      "content": "<p>You are welcome. It is better to ask organizers to make sure that your LB position will remain after the end of the competition and u do not relay on data not given for private test set. However, how the scoring is working is the following. When u click commit, the environment is changed to one used for evaluation, which provides an extended test set with 12 images and maybe extended json files (if u use them). Next the script is executed on this data, and a scoring script is run based on the submission file extracted from the output. So, if u use json and get reasonable results at public LB, likely it is also working at private LB, while if u get an error it may mean that there are no json files provided in test environment. But to be safe, it is better to doublecheck with organizers</p>",
      "rawMarkdown": "You are welcome. It is better to ask organizers to make sure that your LB position will remain after the end of the competition and u do not relay on data not given for private test set. However, how the scoring is working is the following. When u click commit, the environment is changed to one used for evaluation, which provides an extended test set with 12 images and maybe extended json files (if u use them). Next the script is executed on this data, and a scoring script is run based on the submission file extracted from the output. So, if u use json and get reasonable results at public LB, likely it is also working at private LB, while if u get an error it may mean that there are no json files provided in test environment. But to be safe, it is better to doublecheck with organizers",
      "votes": null
    },
    {
      "id": "1086676",
      "postDate": "11/21/2020 23:25:18",
      "content": "<p>The CPU environment has more memory than the GPU environment, so I think submitting without the GPU is currently the only successful way.  But I want to use the GPU when submitting…</p>",
      "rawMarkdown": "The CPU environment has more memory than the GPU environment, so I think submitting without the GPU is currently the only successful way.  But I want to use the GPU when submitting...",
      "votes": null
    },
    {
      "id": "1087268",
      "postDate": "11/22/2020 15:00:34",
      "content": "<p>HI <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks! great work, little complicate hehe, lots to learn…  </p>\n<p>thanks</p>",
      "rawMarkdown": "HI @iafoss thanks! great work, little complicate hehe, lots to learn...  \n\nthanks",
      "votes": null
    },
    {
      "id": "1087393",
      "postDate": "11/22/2020 17:03:08",
      "content": "<p>you are welcome <a href=\"https://www.kaggle.com/oscarrangel\" target=\"_blank\">@oscarrangel</a> </p>",
      "rawMarkdown": "you are welcome @oscarrangel",
      "votes": null
    },
    {
      "id": "1088051",
      "postDate": "11/23/2020 09:17:55",
      "content": "<p>I've got the same problem… Still struggling on it</p>",
      "rawMarkdown": "I've got the same problem... Still struggling on it",
      "votes": null
    },
    {
      "id": "1089513",
      "postDate": "11/24/2020 15:14:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, was wondering how we can use one of these two to build a datablock?</p>\n<p>ds = HuBMAPDataset(tfms=get_aug())<br>\ndl = DataLoader(ds,batch_size=64,shuffle=False,num_workers=NUM_WORKERS)</p>\n<p>or what will be the way to build a datablock with the 256x256 images?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @iafoss, was wondering how we can use one of these two to build a datablock?\n\nds = HuBMAPDataset(tfms=get_aug())\ndl = DataLoader(ds,batch_size=64,shuffle=False,num_workers=NUM_WORKERS)\n\nor what will be the way to build a datablock with the 256x256 images?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1089652",
      "postDate": "11/24/2020 16:52:05",
      "content": "<p>The part u are referring is just for plotting several image examples, the data structures for training are created later on with ImageDataLoaders.from_dsets. I found that the native fast.ai way to treat the data is quite inconvenient, so I create ImageDataLoaders from datasets directly.</p>",
      "rawMarkdown": "The part u are referring is just for plotting several image examples, the data structures for training are created later on with ImageDataLoaders.from_dsets. I found that the native fast.ai way to treat the data is quite inconvenient, so I create ImageDataLoaders from datasets directly.",
      "votes": null
    },
    {
      "id": "1089714",
      "postDate": "11/24/2020 18:04:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, thanks a lot for the hint! your model is very complex, so I am trying to learn from the beginning….. so I am trying this, but is giving me an error. is looking for the number of classes in the dataset… but it has none.</p>\n<pre><code>data = ImageDataLoaders.from_dsets(ds_t,\n                                   ds_v,\n                                   bs=bs,\n                                   num_workers=NUM_WORKERS,\n                                   pin_memory=True).cuda()\n\nlearn = unet_learner(data, \n                    resnet34, \n                    loss_func=lovasz_hinge,\n                    metrics=[meanapv1]).to_fp16(clip=0.5)\n---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n&lt;ipython-input-15-db744f424c39&gt; in &lt;module&gt;\n----&gt; 1 learn = unet_learner(data, \n      2                     resnet34,\n      3                     loss_func=lovasz_hinge,\n      4                     metrics=[meanapv1]).to_fp16(clip=0.5)\n\n~/anaconda3/envs/torch/lib/python3.8/site-packages/fastai/vision/learner.py in unet_learner(dls, arch, loss_func, pretrained, cut, splitter, config, n_in, n_out, normalize, **kwargs)\n    193     size = dls.one_batch()[0].shape[-2:]\n    194     if n_out is None: n_out = get_c(dls)\n--&gt; 195     assert n_out, \"`n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\"\n    196     if normalize: _add_norm(dls, meta, pretrained)\n    197     model = models.unet.DynamicUnet(body, n_out, size, **config)\n\nAssertionError: `n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\n</code></pre>",
      "rawMarkdown": "Hi @iafoss, thanks a lot for the hint! your model is very complex, so I am trying to learn from the beginning..... so I am trying this, but is giving me an error. is looking for the number of classes in the dataset... but it has none.\n~~~\ndata = ImageDataLoaders.from_dsets(ds_t,\n                                   ds_v,\n                                   bs=bs,\n                                   num_workers=NUM_WORKERS,\n                                   pin_memory=True).cuda()\n\nlearn = unet_learner(data, \n                    resnet34, \n                    loss_func=lovasz_hinge,\n                    metrics=[meanapv1]).to_fp16(clip=0.5)\n---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n<ipython-input-15-db744f424c39> in <module>\n----> 1 learn = unet_learner(data, \n      2                     resnet34,\n      3                     loss_func=lovasz_hinge,\n      4                     metrics=[meanapv1]).to_fp16(clip=0.5)\n\n~/anaconda3/envs/torch/lib/python3.8/site-packages/fastai/vision/learner.py in unet_learner(dls, arch, loss_func, pretrained, cut, splitter, config, n_in, n_out, normalize, **kwargs)\n    193     size = dls.one_batch()[0].shape[-2:]\n    194     if n_out is None: n_out = get_c(dls)\n--> 195     assert n_out, \"`n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\"\n    196     if normalize: _add_norm(dls, meta, pretrained)\n    197     model = models.unet.DynamicUnet(body, n_out, size, **config)\n\nAssertionError: `n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\n\n~~~",
      "votes": null
    },
    {
      "id": "1089716",
      "postDate": "11/24/2020 18:09:27",
      "content": "<p>I think fast.ai Unet class requires n_out to build the corresponding object. You can check if u can pass it to unet_learner (using ?? command). Another option is just set this.c = 1 in the dataset. </p>",
      "rawMarkdown": "I think fast.ai Unet class requires n_out to build the corresponding object. You can check if u can pass it to unet_learner (using ?? command). Another option is just set this.c = 1 in the dataset.",
      "votes": null
    },
    {
      "id": "1089727",
      "postDate": "11/24/2020 18:17:09",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Thanks man! but did not work…. that is way I was trying to build something I can pass to different learners with different models, using the way you break the large images into pieces.</p>\n<pre><code>class HuBMAPDataset(Dataset):\n    def __init__(self, fold=fold, train=True, tfms=None):\n        ids = pd.read_csv(LABELS).id.values\n        kf = KFold(n_splits=nfolds,random_state=SEED,shuffle=True)\n        ids = set(ids[list(kf.split(ids))[fold][0 if train else 1]])\n        self.fnames = [fname for fname in os.listdir(TRAIN) if fname.split('_')[0] in ids]\n        self.train = train\n        self.tfms = tfms\n        self.c = 1\n\n-------------------------------------\nModuleAttributeError: 'DynamicUnet' object has no attribute 'enc0'\n</code></pre>",
      "rawMarkdown": "iafoss Thanks man! but did not work.... that is way I was trying to build something I can pass to different learners with different models, using the way you break the large images into pieces.\n~~~\nclass HuBMAPDataset(Dataset):\n    def __init__(self, fold=fold, train=True, tfms=None):\n        ids = pd.read_csv(LABELS).id.values\n        kf = KFold(n_splits=nfolds,random_state=SEED,shuffle=True)\n        ids = set(ids[list(kf.split(ids))[fold][0 if train else 1]])\n        self.fnames = [fname for fname in os.listdir(TRAIN) if fname.split('_')[0] in ids]\n        self.train = train\n        self.tfms = tfms\n        self.c = 1\n\n-------------------------------------\nModuleAttributeError: 'DynamicUnet' object has no attribute 'enc0'\n~~~",
      "votes": null
    },
    {
      "id": "1089736",
      "postDate": "11/24/2020 18:21:39",
      "content": "<p>the last error most likely is related to the way how u do the model split: u cannot use one from my kernel with a default Unet model</p>",
      "rawMarkdown": "the last error most likely is related to the way how u do the model split: u cannot use one from my kernel with a default Unet model",
      "votes": null
    },
    {
      "id": "1089759",
      "postDate": "11/24/2020 18:48:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, it worked!!!!! thanks a lot amigo, I really appreciate it!!! bless you! had to use 4 for the batch size on 1024x1024. amazing, I understand that it is not good to train models with such a low bz, as they wont generalized correctly, is this right ?</p>\n<p>ps.<br>\nis giving me another error, after the training starts<br>\nTypeError: list indices must be integers or slices, not str</p>",
      "rawMarkdown": "Hi @iafoss, it worked!!!!! thanks a lot amigo, I really appreciate it!!! bless you! had to use 4 for the batch size on 1024x1024. amazing, I understand that it is not good to train models with such a low bz, as they wont generalized correctly, is this right ?\n\nps.\nis giving me another error, after the training starts\nTypeError: list indices must be integers or slices, not str",
      "votes": null
    },
    {
      "id": "1089783",
      "postDate": "11/24/2020 19:19:19",
      "content": "<p>I think it is a bad idea to start from large images. And I'm not sure about the last error u got.</p>",
      "rawMarkdown": "I think it is a bad idea to start from large images. And I'm not sure about the last error u got.",
      "votes": null
    },
    {
      "id": "1089788",
      "postDate": "11/24/2020 19:27:42",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , Thanks! again. I see about the large images just wanted to test and see if I was going to be able to handle, (I have one RTX 24 Gigs) the size you recommended earlier.</p>\n<p>You got me  in the beginning of the road to start learning. studying the other notebooks you have about splitting large images. </p>\n<p>So much to learn, I been doing this for a year and I still dont see the light at the end of the tunnel…. hahaha</p>",
      "rawMarkdown": "iafoss , Thanks! again. I see about the large images just wanted to test and see if I was going to be able to handle, (I have one RTX 24 Gigs) the size you recommended earlier.\n\nYou got me  in the beginning of the road to start learning. studying the other notebooks you have about splitting large images. \n\nSo much to learn, I been doing this for a year and I still dont see the light at the end of the tunnel.... hahaha",
      "votes": null
    },
    {
      "id": "1089791",
      "postDate": "11/24/2020 19:36:42",
      "content": "<p>You are welcome</p>",
      "rawMarkdown": "You are welcome",
      "votes": null
    },
    {
      "id": "1091290",
      "postDate": "11/25/2020 23:17:04",
      "content": "<p>Folks, I myself had problems with memory in the private set. I solved it saving the image tiles to disk and then passing the data as a tf.Data.Dataset which reads serially from a file. Basically using the disk as working space. I am sure an equivalent approach can be used here:<br>\n<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk</a></p>",
      "rawMarkdown": "Folks, I myself had problems with memory in the private set. I solved it saving the image tiles to disk and then passing the data as a tf.Data.Dataset which reads serially from a file. Basically using the disk as working space. I am sure an equivalent approach can be used here:\n[https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk](https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk)",
      "votes": null
    },
    {
      "id": "1091496",
      "postDate": "11/26/2020 04:21:09",
      "content": "<p>I'm using the tifffile in my prediction pipeline. (TensorFlow + GPU)<br>\nLoad the whole image -&gt; Split to tiles -&gt; Predict -&gt; Combine the tiles -&gt; RLE encoder<br>\nI think tifffile is OK, just don't forget to use del and gc.collect().</p>",
      "rawMarkdown": "I'm using the tifffile in my prediction pipeline. (TensorFlow + GPU)\nLoad the whole image -> Split to tiles -> Predict -> Combine the tiles -> RLE encoder\nI think tifffile is OK, just don't forget to use del and gc.collect().",
      "votes": null
    },
    {
      "id": "1093446",
      "postDate": "11/27/2020 18:11:00",
      "content": "<p>hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , I was looking at images generated with the notebook 256x256 and I see that most of the tittles dont generate a mask, does that has matter with the classification job for the network to generalize?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "hi @iafoss , I was looking at images generated with the notebook 256x256 and I see that most of the tittles dont generate a mask, does that has matter with the classification job for the network to generalize?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1093479",
      "postDate": "11/27/2020 19:01:52",
      "content": "<p>You have just a few glomeruli, so it is not surprising. Also, that is why when u calculate the metric it is important to aggregate the Intersections and unions for the entire image, like the competition metric is evaluated. If you compute the metric in a tile wise manner, it is quite different task, and results may be misleading.</p>",
      "rawMarkdown": "You have just a few glomeruli, so it is not surprising. Also, that is why when u calculate the metric it is important to aggregate the Intersections and unions for the entire image, like the competition metric is evaluated. If you compute the metric in a tile wise manner, it is quite different task, and results may be misleading.",
      "votes": null
    },
    {
      "id": "1093545",
      "postDate": "11/27/2020 20:12:03",
      "content": "<p>do you think is better to eliminate those images tittles without mask ?</p>\n<p>thanks!</p>\n<p>PS.<br>\nwhat do you mean by this?<br>\n' If you compute the metric in a tile wise manner'</p>",
      "rawMarkdown": "do you think is better to eliminate those images tittles without mask ?\n\nthanks!\n\nPS.\nwhat do you mean by this?\n' If you compute the metric in a tile wise manner'",
      "votes": null
    },
    {
      "id": "1093557",
      "postDate": "11/27/2020 20:37:15",
      "content": "<p>I'd expect it may degrade the score because u need to have had negative examples to train the model.<br>\nFor the second question, it is the way when u compute Dice tile by tile.</p>",
      "rawMarkdown": "I'd expect it may degrade the score because u need to have had negative examples to train the model.\nFor the second question, it is the way when u compute Dice tile by tile.",
      "votes": null
    },
    {
      "id": "1093662",
      "postDate": "11/27/2020 22:25:10",
      "content": "<p>Thank you amigo.</p>",
      "rawMarkdown": "Thank you amigo.",
      "votes": null
    },
    {
      "id": "1093879",
      "postDate": "11/28/2020 05:19:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, I was wondering if you want to team up with me? lots of free time and good hardware</p>",
      "rawMarkdown": "Hi @iafoss, I was wondering if you want to team up with me? lots of free time and good hardware",
      "votes": null
    },
    {
      "id": "1093960",
      "postDate": "11/28/2020 07:09:05",
      "content": "<p>Currently I'm busy with other things, and it's too early stage of the competition for me to team up.</p>",
      "rawMarkdown": "Currently I'm busy with other things, and it's too early stage of the competition for me to team up.",
      "votes": null
    },
    {
      "id": "1096470",
      "postDate": "11/30/2020 14:20:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, I have asked this question in other chats but no one answered, I was wondering if you can tell me how many classes are there, as most of the pre-trained models need at least two, which makes sense, as I understand this is binary segmentation, but you tell me there is only one class….</p>\n<p>Can you explain.</p>\n<p>Thanks again for all your help.</p>\n<p>PS.<br>\nSorry it is my first segmentation exploration.</p>",
      "rawMarkdown": "Hi @iafoss, I have asked this question in other chats but no one answered, I was wondering if you can tell me how many classes are there, as most of the pre-trained models need at least two, which makes sense, as I understand this is binary segmentation, but you tell me there is only one class....\n\nCan you explain.\n\nThanks again for all your help.\n\nPS.\nSorry it is my first segmentation exploration.",
      "votes": null
    },
    {
      "id": "1096698",
      "postDate": "11/30/2020 17:47:34",
      "content": "<p>I think 2 is just a default option in fast.ai: they consider a binary classification as a multiclass classification with two classes. fast.ai , probably, try to use more uniform approach, but it makes things a little bit messier for binary classification( To get the predictions u will need to pass these two outputs through softmax (considering only one output as logits and taking a sigmoid wd be wrong). Though, it may require some coding to use single output with fast.ai. But overall those two things are nearly identical.</p>",
      "rawMarkdown": "I think 2 is just a default option in fast.ai: they consider a binary classification as a multiclass classification with two classes. fast.ai , probably, try to use more uniform approach, but it makes things a little bit messier for binary classification( To get the predictions u will need to pass these two outputs through softmax (considering only one output as logits and taking a sigmoid wd be wrong). Though, it may require some coding to use single output with fast.ai. But overall those two things are nearly identical.",
      "votes": null
    },
    {
      "id": "1096740",
      "postDate": "11/30/2020 18:30:25",
      "content": "<p>Thanks! <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> but your code considers only one, right ?</p>",
      "rawMarkdown": "Thanks! @iafoss but your code considers only one, right ?",
      "votes": null
    },
    {
      "id": "1096766",
      "postDate": "11/30/2020 18:55:57",
      "content": "<p>yes, I'm not pretending to write a code for multiclass problem here</p>",
      "rawMarkdown": "yes, I'm not pretending to write a code for multiclass problem here",
      "votes": null
    },
    {
      "id": "1096973",
      "postDate": "11/30/2020 22:54:29",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks for the information.</p>",
      "rawMarkdown": "Hello @iafoss thanks for the information.",
      "votes": null
    },
    {
      "id": "1096979",
      "postDate": "11/30/2020 23:02:20",
      "content": "<p>You are very welcome</p>",
      "rawMarkdown": "You are very welcome",
      "votes": null
    },
    {
      "id": "1099175",
      "postDate": "12/02/2020 06:26:36",
      "content": "<p>Does it work on private set? I use the same pipeline but the notebook just keep running 9h then report unknown error :(</p>",
      "rawMarkdown": "Does it work on private set? I use the same pipeline but the notebook just keep running 9h then report unknown error :(",
      "votes": null
    },
    {
      "id": "1101186",
      "postDate": "12/03/2020 17:40:59",
      "content": "<p>I had the same problem(Notebook Timeout).<br>\nAnd I was able to submit (Public and Private) if I tried to use less memory.</p>",
      "rawMarkdown": "I had the same problem(Notebook Timeout).\nAnd I was able to submit (Public and Private) if I tried to use less memory.",
      "votes": null
    },
    {
      "id": "1102414",
      "postDate": "12/04/2020 22:37:21",
      "content": "<p>I have an update on the issue with private LB sub: the later version loads images on tile by tile base using rasterio. Please consider this version to make a proper sub.</p>",
      "rawMarkdown": "I have an update on the issue with private LB sub: the later version loads images on tile by tile base using rasterio. Please consider this version to make a proper sub.",
      "votes": null
    },
    {
      "id": "1104588",
      "postDate": "12/07/2020 04:57:07",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!",
      "votes": null
    },
    {
      "id": "1109580",
      "postDate": "12/11/2020 21:09:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, what changes we need to make to the class UneXt50 in order to use it with higher resolutions?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @iafoss, what changes we need to make to the class UneXt50 in order to use it with higher resolutions?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1109625",
      "postDate": "12/11/2020 22:24:24",
      "content": "<p>You just need to give different input, no any changes are needed to the model itself</p>",
      "rawMarkdown": "You just need to give different input, no any changes are needed to the model itself",
      "votes": null
    },
    {
      "id": "1109721",
      "postDate": "12/12/2020 02:07:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks for the response, in what ways we can improve the model? adding more layers to the enconder/decoder ?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @iafoss thanks for the response, in what ways we can improve the model? adding more layers to the enconder/decoder ?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1109786",
      "postDate": "12/12/2020 04:22:36",
      "content": "<p>I think the best boost one could get by looking into the data</p>",
      "rawMarkdown": "I think the best boost one could get by looking into the data",
      "votes": null
    },
    {
      "id": "1110176",
      "postDate": "12/12/2020 13:48:09",
      "content": "<p>Coming from academia and the utopia of prepared datasets ready of modeling, a Senior Director of Artifical Intelligence at Tesla, he found that in the real world, the bread and butter of a deep learning system and where the blood, sweat, and tears would be shed, was in the data. Data acquisition, cleaning, preparation, and the day-to-day management thereof. This same sentiment can as much be inferred from any of you that watched Jeremy Howard's v2 walk through in late 20191… every single session was about getting your data modelable using the new v2 bits. That should tell you a lot!</p>",
      "rawMarkdown": "Coming from academia and the utopia of prepared datasets ready of modeling, a Senior Director of Artifical Intelligence at Tesla, he found that in the real world, the bread and butter of a deep learning system and where the blood, sweat, and tears would be shed, was in the data. Data acquisition, cleaning, preparation, and the day-to-day management thereof. This same sentiment can as much be inferred from any of you that watched Jeremy Howard's v2 walk through in late 20191... every single session was about getting your data modelable using the new v2 bits. That should tell you a lot!",
      "votes": null
    },
    {
      "id": "1110330",
      "postDate": "12/12/2020 16:21:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> can you create a dataloader from your dataset? I been looking around the internet, but nothing found.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @iafoss can you create a dataloader from your dataset? I been looking around the internet, but nothing found.\n\nThanks!",
      "votes": null
    },
    {
      "id": "1110665",
      "postDate": "12/12/2020 23:55:59",
      "content": "<p>If u are talking about Pytorch data loader, u can check my last submission kernel. You just create and pass the dataset u want to use to the DataLoader class. </p>",
      "rawMarkdown": "If u are talking about Pytorch data loader, u can check my last submission kernel. You just create and pass the dataset u want to use to the DataLoader class.",
      "votes": null
    },
    {
      "id": "1110677",
      "postDate": "12/13/2020 00:32:12",
      "content": "<p>At least I found that there are usually three main things to get a high performing model: (1) right concept for building the model for a particular task and data, (2) looking into the data and understanding what are the challenges and how to deal with them, and (3) proper validation to avoid getting \"shake up master\". Keep in mind that (1) requires (2). Other things, like backbones, hyperparameters, etc., are rather minor in many cases if one is using something reasonable enough here.</p>",
      "rawMarkdown": "At least I found that there are usually three main things to get a high performing model: (1) right concept for building the model for a particular task and data, (2) looking into the data and understanding what are the challenges and how to deal with them, and (3) proper validation to avoid getting \"shake up master\". Keep in mind that (1) requires (2). Other things, like backbones, hyperparameters, etc., are rather minor in many cases if one is using something reasonable enough here.",
      "votes": null
    },
    {
      "id": "1237246",
      "postDate": "03/13/2021 23:19:52",
      "content": "<p>I have updated datasets and starter kernels to accommodate for the new data. New LB is 0.905. Enjoy the competition, and the best luck.</p>",
      "rawMarkdown": "I have updated datasets and starter kernels to accommodate for the new data. New LB is 0.905. Enjoy the competition, and the best luck.",
      "votes": null
    },
    {
      "id": "1242630",
      "postDate": "03/17/2021 18:11:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a>,</p>\n<p>Thank you for the updates. I have noticed when you train your model with higher resolution, at some point it start producing NaN values. I tried to debug it but without any success. </p>\n<p>epoch     train_loss  valid_loss  dice_soft  dice_th   time <br>\n0         0.515302    0.431999    0.833299   0.878014  19:45      <br>\n…..<br>\n5         0.381877    0.357184    0.857088   0.897639  12:38     <br>\n6         nan         nan         None                   0.000000  12:15     <br>\n7         nan         nan         None                   0.000000   08:53     <br>\n8         nan         nan         None                   0.000000   07:15</p>",
      "rawMarkdown": "Hi @Iafoss,\n\nThank you for the updates. I have noticed when you train your model with higher resolution, at some point it start producing NaN values. I tried to debug it but without any success. \n\nepoch     train_loss  valid_loss  dice_soft  dice_th   time \n0         0.515302    0.431999    0.833299   0.878014  19:45      \n.....\n5         0.381877    0.357184    0.857088   0.897639  12:38     \n6         nan         nan         None                   0.000000  12:15     \n7         nan         nan         None                   0.000000   08:53     \n8         nan         nan         None                   0.000000   07:15",
      "votes": null
    },
    {
      "id": "1242913",
      "postDate": "03/17/2021 22:37:53",
      "content": "<p>Did you just rerun the kernel or changed something, like batch size? Check my recent reply regarding small bs <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Did you just rerun the kernel or changed something, like batch size? Check my recent reply regarding small bs [here](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter).",
      "votes": null
    },
    {
      "id": "1243122",
      "postDate": "03/18/2021 03:19:53",
      "content": "<p>The batch size 64 with image size 512x512 is not possible on my GPU. I tried with a batch size of 32 getting the same behaviour. I will read your reply, thanks though. </p>",
      "rawMarkdown": "The batch size 64 with image size 512x512 is not possible on my GPU. I tried with a batch size of 32 getting the same behaviour. I will read your reply, thanks though.",
      "votes": null
    },
    {
      "id": "1287552",
      "postDate": "04/29/2021 07:11:29",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!",
      "votes": null
    },
    {
      "id": "1287877",
      "postDate": "04/29/2021 13:42:20",
      "content": "<p>you are welcome</p>",
      "rawMarkdown": "you are welcome",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1085545,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "11/21/2020 00:50:20",
      "content": "<p>Very well done, thanks for sharing your pipeline</p>",
      "votes": null,
      "replies": [
        {
          "id": 1085710,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/21/2020 06:18:39",
          "content": "<p>You are very welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1085798,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "11/21/2020 07:54:30",
      "content": "<p>I have problems with inference to private test sets. I suspect it's because the private test set contains an image that's too big to read using the tifffile. I think we need to use a library that can load subregions of the image, such as pyvips. I would like to wait for an update of the kaggle environment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1085805,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/21/2020 08:00:23",
          "content": "<p>It is likely to be the case but the strange thing is that the notebook doesn't exit with \"out of memory\" error but rather runs to 9h limit and stops (actually it is likely trying to restart and face the same problem). Difficult to debug it without proper error message output.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1085835,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "11/21/2020 08:16:00",
          "content": "<p>I have same problem. If an \"out of memory\" error occurs, there is case that it may run until the time expires instead of exiting immediately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1086321,
          "author_name": "hypocrites",
          "author_url": "",
          "post_date": "11/21/2020 15:32:40",
          "content": "<p>I've done successful commit on CPU instead of GPU which gives you +3 GB of RAM, so seems all is fine with file reading but problem with all memory usage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1086676,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "11/21/2020 23:25:18",
          "content": "<p>The CPU environment has more memory than the GPU environment, so I think submitting without the GPU is currently the only successful way.  But I want to use the GPU when submitting…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088051,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "11/23/2020 09:17:55",
          "content": "<p>I've got the same problem… Still struggling on it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091290,
          "author_name": "marcosnovaes",
          "author_url": "",
          "post_date": "11/25/2020 23:17:04",
          "content": "<p>Folks, I myself had problems with memory in the private set. I solved it saving the image tiles to disk and then passing the data as a tf.Data.Dataset which reads serially from a file. Basically using the disk as working space. I am sure an equivalent approach can be used here:<br>\n<a href=\"https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk\" target=\"_blank\">https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091496,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "11/26/2020 04:21:09",
          "content": "<p>I'm using the tifffile in my prediction pipeline. (TensorFlow + GPU)<br>\nLoad the whole image -&gt; Split to tiles -&gt; Predict -&gt; Combine the tiles -&gt; RLE encoder<br>\nI think tifffile is OK, just don't forget to use del and gc.collect().</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1099175,
          "author_name": "yimyukei",
          "author_url": "",
          "post_date": "12/02/2020 06:26:36",
          "content": "<p>Does it work on private set? I use the same pipeline but the notebook just keep running 9h then report unknown error :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101186,
          "author_name": "yukkyo",
          "author_url": "",
          "post_date": "12/03/2020 17:40:59",
          "content": "<p>I had the same problem(Notebook Timeout).<br>\nAnd I was able to submit (Public and Private) if I tried to use less memory.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1086323,
      "author_name": "scherzo",
      "author_url": "",
      "post_date": "11/21/2020 15:34:44",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <br>\nSince they provided anatomical structure annotations in the public test set do you think they are also available in the private test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1086478,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/21/2020 18:03:58",
          "content": "<p>You are welcome. It is better to ask organizers to make sure that your LB position will remain after the end of the competition and u do not relay on data not given for private test set. However, how the scoring is working is the following. When u click commit, the environment is changed to one used for evaluation, which provides an extended test set with 12 images and maybe extended json files (if u use them). Next the script is executed on this data, and a scoring script is run based on the submission file extracted from the output. So, if u use json and get reasonable results at public LB, likely it is also working at private LB, while if u get an error it may mean that there are no json files provided in test environment. But to be safe, it is better to doublecheck with organizers</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1087268,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "11/22/2020 15:00:34",
      "content": "<p>HI <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks! great work, little complicate hehe, lots to learn…  </p>\n<p>thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1087393,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/22/2020 17:03:08",
          "content": "<p>you are welcome <a href=\"https://www.kaggle.com/oscarrangel\" target=\"_blank\">@oscarrangel</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1089513,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "11/24/2020 15:14:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, was wondering how we can use one of these two to build a datablock?</p>\n<p>ds = HuBMAPDataset(tfms=get_aug())<br>\ndl = DataLoader(ds,batch_size=64,shuffle=False,num_workers=NUM_WORKERS)</p>\n<p>or what will be the way to build a datablock with the 256x256 images?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1089652,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/24/2020 16:52:05",
          "content": "<p>The part u are referring is just for plotting several image examples, the data structures for training are created later on with ImageDataLoaders.from_dsets. I found that the native fast.ai way to treat the data is quite inconvenient, so I create ImageDataLoaders from datasets directly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089714,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/24/2020 18:04:40",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, thanks a lot for the hint! your model is very complex, so I am trying to learn from the beginning….. so I am trying this, but is giving me an error. is looking for the number of classes in the dataset… but it has none.</p>\n<pre><code>data = ImageDataLoaders.from_dsets(ds_t,\n                                   ds_v,\n                                   bs=bs,\n                                   num_workers=NUM_WORKERS,\n                                   pin_memory=True).cuda()\n\nlearn = unet_learner(data, \n                    resnet34, \n                    loss_func=lovasz_hinge,\n                    metrics=[meanapv1]).to_fp16(clip=0.5)\n---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n&lt;ipython-input-15-db744f424c39&gt; in &lt;module&gt;\n----&gt; 1 learn = unet_learner(data, \n      2                     resnet34,\n      3                     loss_func=lovasz_hinge,\n      4                     metrics=[meanapv1]).to_fp16(clip=0.5)\n\n~/anaconda3/envs/torch/lib/python3.8/site-packages/fastai/vision/learner.py in unet_learner(dls, arch, loss_func, pretrained, cut, splitter, config, n_in, n_out, normalize, **kwargs)\n    193     size = dls.one_batch()[0].shape[-2:]\n    194     if n_out is None: n_out = get_c(dls)\n--&gt; 195     assert n_out, \"`n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\"\n    196     if normalize: _add_norm(dls, meta, pretrained)\n    197     model = models.unet.DynamicUnet(body, n_out, size, **config)\n\nAssertionError: `n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089716,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/24/2020 18:09:27",
          "content": "<p>I think fast.ai Unet class requires n_out to build the corresponding object. You can check if u can pass it to unet_learner (using ?? command). Another option is just set this.c = 1 in the dataset. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089727,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/24/2020 18:17:09",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Thanks man! but did not work…. that is way I was trying to build something I can pass to different learners with different models, using the way you break the large images into pieces.</p>\n<pre><code>class HuBMAPDataset(Dataset):\n    def __init__(self, fold=fold, train=True, tfms=None):\n        ids = pd.read_csv(LABELS).id.values\n        kf = KFold(n_splits=nfolds,random_state=SEED,shuffle=True)\n        ids = set(ids[list(kf.split(ids))[fold][0 if train else 1]])\n        self.fnames = [fname for fname in os.listdir(TRAIN) if fname.split('_')[0] in ids]\n        self.train = train\n        self.tfms = tfms\n        self.c = 1\n\n-------------------------------------\nModuleAttributeError: 'DynamicUnet' object has no attribute 'enc0'\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089736,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/24/2020 18:21:39",
          "content": "<p>the last error most likely is related to the way how u do the model split: u cannot use one from my kernel with a default Unet model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089759,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/24/2020 18:48:53",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, it worked!!!!! thanks a lot amigo, I really appreciate it!!! bless you! had to use 4 for the batch size on 1024x1024. amazing, I understand that it is not good to train models with such a low bz, as they wont generalized correctly, is this right ?</p>\n<p>ps.<br>\nis giving me another error, after the training starts<br>\nTypeError: list indices must be integers or slices, not str</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089783,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/24/2020 19:19:19",
          "content": "<p>I think it is a bad idea to start from large images. And I'm not sure about the last error u got.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089788,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/24/2020 19:27:42",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , Thanks! again. I see about the large images just wanted to test and see if I was going to be able to handle, (I have one RTX 24 Gigs) the size you recommended earlier.</p>\n<p>You got me  in the beginning of the road to start learning. studying the other notebooks you have about splitting large images. </p>\n<p>So much to learn, I been doing this for a year and I still dont see the light at the end of the tunnel…. hahaha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1089791,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/24/2020 19:36:42",
          "content": "<p>You are welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1093446,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "11/27/2020 18:11:00",
      "content": "<p>hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , I was looking at images generated with the notebook 256x256 and I see that most of the tittles dont generate a mask, does that has matter with the classification job for the network to generalize?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1093479,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/27/2020 19:01:52",
          "content": "<p>You have just a few glomeruli, so it is not surprising. Also, that is why when u calculate the metric it is important to aggregate the Intersections and unions for the entire image, like the competition metric is evaluated. If you compute the metric in a tile wise manner, it is quite different task, and results may be misleading.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093545,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/27/2020 20:12:03",
          "content": "<p>do you think is better to eliminate those images tittles without mask ?</p>\n<p>thanks!</p>\n<p>PS.<br>\nwhat do you mean by this?<br>\n' If you compute the metric in a tile wise manner'</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093557,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/27/2020 20:37:15",
          "content": "<p>I'd expect it may degrade the score because u need to have had negative examples to train the model.<br>\nFor the second question, it is the way when u compute Dice tile by tile.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1093662,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/27/2020 22:25:10",
          "content": "<p>Thank you amigo.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1093879,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "11/28/2020 05:19:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, I was wondering if you want to team up with me? lots of free time and good hardware</p>",
      "votes": null,
      "replies": [
        {
          "id": 1093960,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/28/2020 07:09:05",
          "content": "<p>Currently I'm busy with other things, and it's too early stage of the competition for me to team up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1096470,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "11/30/2020 14:20:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, I have asked this question in other chats but no one answered, I was wondering if you can tell me how many classes are there, as most of the pre-trained models need at least two, which makes sense, as I understand this is binary segmentation, but you tell me there is only one class….</p>\n<p>Can you explain.</p>\n<p>Thanks again for all your help.</p>\n<p>PS.<br>\nSorry it is my first segmentation exploration.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1096698,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/30/2020 17:47:34",
          "content": "<p>I think 2 is just a default option in fast.ai: they consider a binary classification as a multiclass classification with two classes. fast.ai , probably, try to use more uniform approach, but it makes things a little bit messier for binary classification( To get the predictions u will need to pass these two outputs through softmax (considering only one output as logits and taking a sigmoid wd be wrong). Though, it may require some coding to use single output with fast.ai. But overall those two things are nearly identical.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096740,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "11/30/2020 18:30:25",
          "content": "<p>Thanks! <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> but your code considers only one, right ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1096766,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/30/2020 18:55:57",
          "content": "<p>yes, I'm not pretending to write a code for multiclass problem here</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1096973,
      "author_name": "puzuwe",
      "author_url": "",
      "post_date": "11/30/2020 22:54:29",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks for the information.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1096979,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/30/2020 23:02:20",
          "content": "<p>You are very welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1102414,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "12/04/2020 22:37:21",
      "content": "<p>I have an update on the issue with private LB sub: the later version loads images on tile by tile base using rasterio. Please consider this version to make a proper sub.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1104588,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "12/07/2020 04:57:07",
          "content": "<p>thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109580,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "12/11/2020 21:09:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, what changes we need to make to the class UneXt50 in order to use it with higher resolutions?</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109625,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/11/2020 22:24:24",
          "content": "<p>You just need to give different input, no any changes are needed to the model itself</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109721,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "12/12/2020 02:07:17",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> thanks for the response, in what ways we can improve the model? adding more layers to the enconder/decoder ?</p>\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109786,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/12/2020 04:22:36",
          "content": "<p>I think the best boost one could get by looking into the data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1110176,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "12/12/2020 13:48:09",
          "content": "<p>Coming from academia and the utopia of prepared datasets ready of modeling, a Senior Director of Artifical Intelligence at Tesla, he found that in the real world, the bread and butter of a deep learning system and where the blood, sweat, and tears would be shed, was in the data. Data acquisition, cleaning, preparation, and the day-to-day management thereof. This same sentiment can as much be inferred from any of you that watched Jeremy Howard's v2 walk through in late 20191… every single session was about getting your data modelable using the new v2 bits. That should tell you a lot!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1110677,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/13/2020 00:32:12",
          "content": "<p>At least I found that there are usually three main things to get a high performing model: (1) right concept for building the model for a particular task and data, (2) looking into the data and understanding what are the challenges and how to deal with them, and (3) proper validation to avoid getting \"shake up master\". Keep in mind that (1) requires (2). Other things, like backbones, hyperparameters, etc., are rather minor in many cases if one is using something reasonable enough here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1110330,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "12/12/2020 16:21:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> can you create a dataloader from your dataset? I been looking around the internet, but nothing found.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1110665,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/12/2020 23:55:59",
          "content": "<p>If u are talking about Pytorch data loader, u can check my last submission kernel. You just create and pass the dataset u want to use to the DataLoader class. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1237246,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "03/13/2021 23:19:52",
      "content": "<p>I have updated datasets and starter kernels to accommodate for the new data. New LB is 0.905. Enjoy the competition, and the best luck.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1242630,
          "author_name": "aramisvesal",
          "author_url": "",
          "post_date": "03/17/2021 18:11:48",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a>,</p>\n<p>Thank you for the updates. I have noticed when you train your model with higher resolution, at some point it start producing NaN values. I tried to debug it but without any success. </p>\n<p>epoch     train_loss  valid_loss  dice_soft  dice_th   time <br>\n0         0.515302    0.431999    0.833299   0.878014  19:45      <br>\n…..<br>\n5         0.381877    0.357184    0.857088   0.897639  12:38     <br>\n6         nan         nan         None                   0.000000  12:15     <br>\n7         nan         nan         None                   0.000000   08:53     <br>\n8         nan         nan         None                   0.000000   07:15</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242913,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "03/17/2021 22:37:53",
          "content": "<p>Did you just rerun the kernel or changed something, like batch size? Check my recent reply regarding small bs <a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243122,
          "author_name": "aramisvesal",
          "author_url": "",
          "post_date": "03/18/2021 03:19:53",
          "content": "<p>The batch size 64 with image size 512x512 is not possible on my GPU. I tried with a batch size of 32 getting the same behaviour. I will read your reply, thanks though. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1287552,
      "author_name": "sasakipage",
      "author_url": "",
      "post_date": "04/29/2021 07:11:29",
      "content": "<p>thanks a lot!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1287877,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "04/29/2021 13:42:20",
          "content": "<p>you are welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1085502": "Welcome to Human BioMolecular Atlas Program (HuBMAP) competition. I have prepared several kernels for people who use Pytorch and fast.ai to get started with this challenge and for others to borrow some ideas and explore them more. \n- https://www.kaggle.com/iafoss/256x256-images : the kernel used to cut the original images into tiles and save them to a dataset\n- https://www.kaggle.com/iafoss/hubmap-fast-ai-starter : the training kernel sharing my experience in the previous segmentation competitions and giving a Pytorch/fast.ai pipeline to start with\n-  https://www.kaggle.com/iafoss/hubmap-fast-ai-starter-sub : the starter inference kernel that got 0.905 LB on 4 times lower resolution images (256x256 tile dataset)\n\nI hope you will enjoy this competition, and I wish you the best luck. \n03/13/2021 The datasets and kernels are updated to accommodate for the new data.",
    "1085545": "Very well done, thanks for sharing your pipeline",
    "1085710": "You are very welcome",
    "1085798": "I have problems with inference to private test sets. I suspect it's because the private test set contains an image that's too big to read using the tifffile. I think we need to use a library that can load subregions of the image, such as pyvips. I would like to wait for an update of the kaggle environment.",
    "1085805": "It is likely to be the case but the strange thing is that the notebook doesn't exit with \"out of memory\" error but rather runs to 9h limit and stops (actually it is likely trying to restart and face the same problem). Difficult to debug it without proper error message output.",
    "1085835": "I have same problem. If an \"out of memory\" error occurs, there is case that it may run until the time expires instead of exiting immediately.",
    "1086321": "I've done successful commit on CPU instead of GPU which gives you +3 GB of RAM, so seems all is fine with file reading but problem with all memory usage.",
    "1086323": "Thanks for sharing @iafoss \nSince they provided anatomical structure annotations in the public test set do you think they are also available in the private test set?",
    "1086478": "You are welcome. It is better to ask organizers to make sure that your LB position will remain after the end of the competition and u do not relay on data not given for private test set. However, how the scoring is working is the following. When u click commit, the environment is changed to one used for evaluation, which provides an extended test set with 12 images and maybe extended json files (if u use them). Next the script is executed on this data, and a scoring script is run based on the submission file extracted from the output. So, if u use json and get reasonable results at public LB, likely it is also working at private LB, while if u get an error it may mean that there are no json files provided in test environment. But to be safe, it is better to doublecheck with organizers",
    "1086676": "The CPU environment has more memory than the GPU environment, so I think submitting without the GPU is currently the only successful way.  But I want to use the GPU when submitting...",
    "1087268": "HI @iafoss thanks! great work, little complicate hehe, lots to learn...  \n\nthanks",
    "1087393": "you are welcome @oscarrangel",
    "1088051": "I've got the same problem... Still struggling on it",
    "1089513": "Hi @iafoss, was wondering how we can use one of these two to build a datablock?\n\nds = HuBMAPDataset(tfms=get_aug())\ndl = DataLoader(ds,batch_size=64,shuffle=False,num_workers=NUM_WORKERS)\n\nor what will be the way to build a datablock with the 256x256 images?\n\nThanks!",
    "1089652": "The part u are referring is just for plotting several image examples, the data structures for training are created later on with ImageDataLoaders.from_dsets. I found that the native fast.ai way to treat the data is quite inconvenient, so I create ImageDataLoaders from datasets directly.",
    "1089714": "Hi @iafoss, thanks a lot for the hint! your model is very complex, so I am trying to learn from the beginning..... so I am trying this, but is giving me an error. is looking for the number of classes in the dataset... but it has none.\n~~~\ndata = ImageDataLoaders.from_dsets(ds_t,\n                                   ds_v,\n                                   bs=bs,\n                                   num_workers=NUM_WORKERS,\n                                   pin_memory=True).cuda()\n\nlearn = unet_learner(data, \n                    resnet34, \n                    loss_func=lovasz_hinge,\n                    metrics=[meanapv1]).to_fp16(clip=0.5)\n---------------------------------------------------------------------------\nAssertionError                            Traceback (most recent call last)\n<ipython-input-15-db744f424c39> in <module>\n----> 1 learn = unet_learner(data, \n      2                     resnet34,\n      3                     loss_func=lovasz_hinge,\n      4                     metrics=[meanapv1]).to_fp16(clip=0.5)\n\n~/anaconda3/envs/torch/lib/python3.8/site-packages/fastai/vision/learner.py in unet_learner(dls, arch, loss_func, pretrained, cut, splitter, config, n_in, n_out, normalize, **kwargs)\n    193     size = dls.one_batch()[0].shape[-2:]\n    194     if n_out is None: n_out = get_c(dls)\n--> 195     assert n_out, \"`n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\"\n    196     if normalize: _add_norm(dls, meta, pretrained)\n    197     model = models.unet.DynamicUnet(body, n_out, size, **config)\n\nAssertionError: `n_out` is not defined, and could not be inferred from data, set `dls.c` or pass `n_out`\n\n~~~",
    "1089716": "I think fast.ai Unet class requires n_out to build the corresponding object. You can check if u can pass it to unet_learner (using ?? command). Another option is just set this.c = 1 in the dataset.",
    "1089727": "iafoss Thanks man! but did not work.... that is way I was trying to build something I can pass to different learners with different models, using the way you break the large images into pieces.\n~~~\nclass HuBMAPDataset(Dataset):\n    def __init__(self, fold=fold, train=True, tfms=None):\n        ids = pd.read_csv(LABELS).id.values\n        kf = KFold(n_splits=nfolds,random_state=SEED,shuffle=True)\n        ids = set(ids[list(kf.split(ids))[fold][0 if train else 1]])\n        self.fnames = [fname for fname in os.listdir(TRAIN) if fname.split('_')[0] in ids]\n        self.train = train\n        self.tfms = tfms\n        self.c = 1\n\n-------------------------------------\nModuleAttributeError: 'DynamicUnet' object has no attribute 'enc0'\n~~~",
    "1089736": "the last error most likely is related to the way how u do the model split: u cannot use one from my kernel with a default Unet model",
    "1089759": "Hi @iafoss, it worked!!!!! thanks a lot amigo, I really appreciate it!!! bless you! had to use 4 for the batch size on 1024x1024. amazing, I understand that it is not good to train models with such a low bz, as they wont generalized correctly, is this right ?\n\nps.\nis giving me another error, after the training starts\nTypeError: list indices must be integers or slices, not str",
    "1089783": "I think it is a bad idea to start from large images. And I'm not sure about the last error u got.",
    "1089788": "iafoss , Thanks! again. I see about the large images just wanted to test and see if I was going to be able to handle, (I have one RTX 24 Gigs) the size you recommended earlier.\n\nYou got me  in the beginning of the road to start learning. studying the other notebooks you have about splitting large images. \n\nSo much to learn, I been doing this for a year and I still dont see the light at the end of the tunnel.... hahaha",
    "1089791": "You are welcome",
    "1091290": "Folks, I myself had problems with memory in the private set. I solved it saving the image tiles to disk and then passing the data as a tf.Data.Dataset which reads serially from a file. Basically using the disk as working space. I am sure an equivalent approach can be used here:\n[https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk](https://www.kaggle.com/marcosnovaes/hubmap-memory-efficient-submission-using-disk)",
    "1091496": "I'm using the tifffile in my prediction pipeline. (TensorFlow + GPU)\nLoad the whole image -> Split to tiles -> Predict -> Combine the tiles -> RLE encoder\nI think tifffile is OK, just don't forget to use del and gc.collect().",
    "1093446": "hi @iafoss , I was looking at images generated with the notebook 256x256 and I see that most of the tittles dont generate a mask, does that has matter with the classification job for the network to generalize?\n\nThanks!",
    "1093479": "You have just a few glomeruli, so it is not surprising. Also, that is why when u calculate the metric it is important to aggregate the Intersections and unions for the entire image, like the competition metric is evaluated. If you compute the metric in a tile wise manner, it is quite different task, and results may be misleading.",
    "1093545": "do you think is better to eliminate those images tittles without mask ?\n\nthanks!\n\nPS.\nwhat do you mean by this?\n' If you compute the metric in a tile wise manner'",
    "1093557": "I'd expect it may degrade the score because u need to have had negative examples to train the model.\nFor the second question, it is the way when u compute Dice tile by tile.",
    "1093662": "Thank you amigo.",
    "1093879": "Hi @iafoss, I was wondering if you want to team up with me? lots of free time and good hardware",
    "1093960": "Currently I'm busy with other things, and it's too early stage of the competition for me to team up.",
    "1096470": "Hi @iafoss, I have asked this question in other chats but no one answered, I was wondering if you can tell me how many classes are there, as most of the pre-trained models need at least two, which makes sense, as I understand this is binary segmentation, but you tell me there is only one class....\n\nCan you explain.\n\nThanks again for all your help.\n\nPS.\nSorry it is my first segmentation exploration.",
    "1096698": "I think 2 is just a default option in fast.ai: they consider a binary classification as a multiclass classification with two classes. fast.ai , probably, try to use more uniform approach, but it makes things a little bit messier for binary classification( To get the predictions u will need to pass these two outputs through softmax (considering only one output as logits and taking a sigmoid wd be wrong). Though, it may require some coding to use single output with fast.ai. But overall those two things are nearly identical.",
    "1096740": "Thanks! @iafoss but your code considers only one, right ?",
    "1096766": "yes, I'm not pretending to write a code for multiclass problem here",
    "1096973": "Hello @iafoss thanks for the information.",
    "1096979": "You are very welcome",
    "1099175": "Does it work on private set? I use the same pipeline but the notebook just keep running 9h then report unknown error :(",
    "1101186": "I had the same problem(Notebook Timeout).\nAnd I was able to submit (Public and Private) if I tried to use less memory.",
    "1102414": "I have an update on the issue with private LB sub: the later version loads images on tile by tile base using rasterio. Please consider this version to make a proper sub.",
    "1104588": "thanks a lot!",
    "1109580": "Hi @iafoss, what changes we need to make to the class UneXt50 in order to use it with higher resolutions?\n\nThanks!",
    "1109625": "You just need to give different input, no any changes are needed to the model itself",
    "1109721": "Hi @iafoss thanks for the response, in what ways we can improve the model? adding more layers to the enconder/decoder ?\n\nThanks!",
    "1109786": "I think the best boost one could get by looking into the data",
    "1110176": "Coming from academia and the utopia of prepared datasets ready of modeling, a Senior Director of Artifical Intelligence at Tesla, he found that in the real world, the bread and butter of a deep learning system and where the blood, sweat, and tears would be shed, was in the data. Data acquisition, cleaning, preparation, and the day-to-day management thereof. This same sentiment can as much be inferred from any of you that watched Jeremy Howard's v2 walk through in late 20191... every single session was about getting your data modelable using the new v2 bits. That should tell you a lot!",
    "1110330": "Hi @iafoss can you create a dataloader from your dataset? I been looking around the internet, but nothing found.\n\nThanks!",
    "1110665": "If u are talking about Pytorch data loader, u can check my last submission kernel. You just create and pass the dataset u want to use to the DataLoader class.",
    "1110677": "At least I found that there are usually three main things to get a high performing model: (1) right concept for building the model for a particular task and data, (2) looking into the data and understanding what are the challenges and how to deal with them, and (3) proper validation to avoid getting \"shake up master\". Keep in mind that (1) requires (2). Other things, like backbones, hyperparameters, etc., are rather minor in many cases if one is using something reasonable enough here.",
    "1237246": "I have updated datasets and starter kernels to accommodate for the new data. New LB is 0.905. Enjoy the competition, and the best luck.",
    "1242630": "Hi @Iafoss,\n\nThank you for the updates. I have noticed when you train your model with higher resolution, at some point it start producing NaN values. I tried to debug it but without any success. \n\nepoch     train_loss  valid_loss  dice_soft  dice_th   time \n0         0.515302    0.431999    0.833299   0.878014  19:45      \n.....\n5         0.381877    0.357184    0.857088   0.897639  12:38     \n6         nan         nan         None                   0.000000  12:15     \n7         nan         nan         None                   0.000000   08:53     \n8         nan         nan         None                   0.000000   07:15",
    "1242913": "Did you just rerun the kernel or changed something, like batch size? Check my recent reply regarding small bs [here](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter).",
    "1243122": "The batch size 64 with image size 512x512 is not possible on my GPU. I tried with a batch size of 32 getting the same behaviour. I will read your reply, thanks though.",
    "1287552": "thanks a lot!",
    "1287877": "you are welcome"
  },
  "source": "meta"
}