{
  "id": 118328,
  "title": "A middle-rank guys´ perspective & lessons from this competition",
  "url": "/competitions/understanding_cloud_organization/discussion/118328",
  "author_name": "",
  "post_date": "2019-11-20T19:16:29.763201Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi!</p>\n\n<p>This was our first segmentation competition, and I think both my team mate and I learned a lot in this context!</p>\n\n<p>EDIT: This actually turned into a pretty long post. Half of it is about organisation, the other half about pure data science.</p>\n\n<p>In this post I want to bring a few things forward that I have learned, from someone that is still learning a lot in computer vision. I think it's interesting to not only have the perspective of the best Kagglers, but also those in the middle &amp; bottom; everybody is learning something different and sometimes going back to the basics is helpful.</p>\n\n<p>First of all, a few notes on what I was using:\n- 1 Unet for each label. Same one, pre-trained on one label with 20 epochs then using those weights to learn with 5 epochs on the other 3 labels. Might not have been a great idea but did give interesting results while keeping the training time down.\n- Backbone: ResNet34\n- Batch size: 32\n- Resized to output size: 350, 525\n- Ended up not having time to add much augmentation nor threshold search for the predictions.\n- Implemented a Classifier (InceptionResnetV2 backbone) to try to improve results by finding images with no masks. Turned out to spend way to much time on that for what it produced.</p>\n\n<p>For the record, we were using Keras for everything, so not everything might be applicable.</p>\n\n<p>I'd like to separate this in a few categories:</p>\n\n<h1><strong>Personal take-aways</strong></h1>\n\n<h2><strong>Organisation &amp; handling the competition</strong></h2>\n\n<ul>\n<li>This competition was pretty tricky because of the labels. Some of them had some oddly specific regions, and as we all know, they were quite noisy / subjective. That did make for some very interesting reflecting though. Pretty interesting and great exercise for a first segmentation competition.</li>\n<li>On a personal note, I approach Kaggle competitions from a <strong>learning</strong> perspective over a <strong>competitive</strong> one. So while we didn't get a great ranking, I see this as a good success, as I learned a lot from it. This also means it gets tricky to balance out what <strong>I want to do myself</strong> against <strong>what I could use from others &amp; implement</strong>. It's just like how we are not using numpy to make out own models, where's the limit. I found it interesting to hard code my own UNet and then move towards the widely used <a href=\"https://github.com/qubvel/segmentation_models\">Segmentation model library</a>.</li>\n<li>As the images were pretty big and had 4 masks each, fitting everything in RAM wasn't doable on Kaggle notebooks. So this force us to dive deeper in making custom data generators. I started making my own (with a function instead of a class, that turned out to be a great learning experience, mostly about why <em>not</em> to do it and turn towards suing classes instead) before using one of the custom classes posted in the public notebooks and refactoring it to my own code.\n(For example we made our own encoding / decoding functions, which from a competition stand-point was probably dumping time out of the window, but from a learning experience was great).</li>\n<li>My workflow basically consisted of working on my own at first, then discussing ideas / issues with my team mate, and finally going to the forum / public notebook if we were still stuck. This is a basic yet very effective (IMHO) way of learning: trying stuff on your own, not getting influenced by outside factors, and only finding help when you're stuck. So there was a lot of re-inventing the wheel, but that's also great to learn AND understand why things are done one way or another.</li>\n<li>We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights, predictions, and images + labels.</li>\n<li>Having limited options and having to juggle between platforms pushed us to find ways around it and interesting solutions outside of just pure data science. We set up a bunch of Slack bots that would send us our training losses &amp; scores, and notify us when training was done. This way we could run training while working on something else, download weights &amp; results and shut down the kernel to maximize GPU time usage. Nothing mind blowing, but a good addition towards optimizing our limited resources.</li>\n</ul>\n\n<h2><strong>Technical aspects &amp; challenges of the competition</strong>:</h2>\n\n<ul>\n<li>Having limited hardware should have pushed us towards compressing the data more to be able to experiment and iterate faster. <a href=\"/cdeotte\">@cdeotte</a> shared some really interesting notebooks all along on how dividing the image sizes by 4 or even 5 helped him iterate much faster. This is definitely something we should have looked more into. Overall, thanks <a href=\"/cdeotte\">@cdeotte</a> for all the stuff you posted, there were some really nice tips for all the of us on limited hardware!</li>\n<li>Organization, keeping track of what was done and what needed to be done was better than our previous competitions, but could still see some improvements. Kaggle is terrible for version control and keeping track of notebooks, as it's basically non-existent. So we need to implement something ourselves to overcome this, and keep tracks of all the logs, weights &amp; preds. Other than learning about data science methods, I learned a lot about how to organize myself on a multi-month project a bit better than the previous times.</li>\n<li>Overall, I feel like we didn't get to try nearly as much things as we had in mind. I'm frustrated because what held us back wasn't lack of ideas and things we wanted to try, but loosing a lot of time implementing them &amp; slow training at times. I'd say the bottleneck came from 80% of slow implementations, and 20% on long training times. This means I have to improve a lot on implementing better implementations and reduce training times with different methods.</li>\n<li>Because of the slow implementations &amp; long training times, I sort of got lost into only trying to get one or two things done, when I should have tried to simplify to iterate faster.</li>\n<li>I tried prioritizing based on past experience and on the error analysis of my models. This helps figure out which was the most important out of: post-processing, fine-tuning the model, changing the model, making a classifier to eliminate images without labels, etc. </li>\n</ul>\n\n<h1><strong>Take-aways from top Kagglers after the competition</strong></h1>\n\n<ul>\n<li>There were some really good results even with (relatviely) simple pipelines.</li>\n<li>People tested a lot of different backbones, models, losses, to find which ones were producing the best results.</li>\n<li>Unet was used a lot, but FPN seems to be pretty good too.</li>\n<li>This competition had very noisy data. Ensembling seems to be pretty common to overcome this.</li>\n<li>There were some really interesting ideas:\n<ul><li>Doing some very selective pseduo-labelling to images that were carefully selected.</li>\n<li>Trying different types of resolutions / image augmentations techniques</li>\n<li>Public set probing. That's not really where I want to go personally as a data scientist, but it's always good to know what's the idea behind and how to do it.</li>\n<li>Weighted ensembling depending on the labels</li>\n<li>Using a classifier as a segmentation model itself. Or using a segmentation model to classify the images. </li></ul></li>\n</ul>\n\n<h1><strong>Conclusion</strong></h1>\n\n<p>This was a great learning experience.\nWhat I think is the biggest thing for me is being able to organize myself better &amp; be able to create a working pipeline as fast as possible. By that I not only mean over fitting to a small dataset, but also trying things like lowering the resolution of the images to try out multiple models, backbones, losses, etc. at an early stage.\nWorking only on free cloud computing resources is actually pretty fun. It's a great way to learn, trying to find as much tricks and ways around the limitations that are in place. One of my personal goals is to get a medal only with my laptop. I'm more interested in learning how to optimize my workflow rather than going up with raw compute power.</p>\n\n<p>I hope this is helpful for other people that might not be at the top of the leader board, but that are also trying to learn from these competitions! I'd also be curious to hear about any other middle or low scoring Kaggler with issues they have faced; as well as what higher ranking Kagglers might have to say about my debrief of this competition.</p>",
  "messages": [
    {
      "id": "677925",
      "postDate": "11/20/2019 19:16:29",
      "content": "<p>Hi!</p>\n\n<p>This was our first segmentation competition, and I think both my team mate and I learned a lot in this context!</p>\n\n<p>EDIT: This actually turned into a pretty long post. Half of it is about organisation, the other half about pure data science.</p>\n\n<p>In this post I want to bring a few things forward that I have learned, from someone that is still learning a lot in computer vision. I think it's interesting to not only have the perspective of the best Kagglers, but also those in the middle &amp; bottom; everybody is learning something different and sometimes going back to the basics is helpful.</p>\n\n<p>First of all, a few notes on what I was using:\n- 1 Unet for each label. Same one, pre-trained on one label with 20 epochs then using those weights to learn with 5 epochs on the other 3 labels. Might not have been a great idea but did give interesting results while keeping the training time down.\n- Backbone: ResNet34\n- Batch size: 32\n- Resized to output size: 350, 525\n- Ended up not having time to add much augmentation nor threshold search for the predictions.\n- Implemented a Classifier (InceptionResnetV2 backbone) to try to improve results by finding images with no masks. Turned out to spend way to much time on that for what it produced.</p>\n\n<p>For the record, we were using Keras for everything, so not everything might be applicable.</p>\n\n<p>I'd like to separate this in a few categories:</p>\n\n<h1><strong>Personal take-aways</strong></h1>\n\n<h2><strong>Organisation &amp; handling the competition</strong></h2>\n\n<ul>\n<li>This competition was pretty tricky because of the labels. Some of them had some oddly specific regions, and as we all know, they were quite noisy / subjective. That did make for some very interesting reflecting though. Pretty interesting and great exercise for a first segmentation competition.</li>\n<li>On a personal note, I approach Kaggle competitions from a <strong>learning</strong> perspective over a <strong>competitive</strong> one. So while we didn't get a great ranking, I see this as a good success, as I learned a lot from it. This also means it gets tricky to balance out what <strong>I want to do myself</strong> against <strong>what I could use from others &amp; implement</strong>. It's just like how we are not using numpy to make out own models, where's the limit. I found it interesting to hard code my own UNet and then move towards the widely used <a href=\"https://github.com/qubvel/segmentation_models\">Segmentation model library</a>.</li>\n<li>As the images were pretty big and had 4 masks each, fitting everything in RAM wasn't doable on Kaggle notebooks. So this force us to dive deeper in making custom data generators. I started making my own (with a function instead of a class, that turned out to be a great learning experience, mostly about why <em>not</em> to do it and turn towards suing classes instead) before using one of the custom classes posted in the public notebooks and refactoring it to my own code.\n(For example we made our own encoding / decoding functions, which from a competition stand-point was probably dumping time out of the window, but from a learning experience was great).</li>\n<li>My workflow basically consisted of working on my own at first, then discussing ideas / issues with my team mate, and finally going to the forum / public notebook if we were still stuck. This is a basic yet very effective (IMHO) way of learning: trying stuff on your own, not getting influenced by outside factors, and only finding help when you're stuck. So there was a lot of re-inventing the wheel, but that's also great to learn AND understand why things are done one way or another.</li>\n<li>We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights, predictions, and images + labels.</li>\n<li>Having limited options and having to juggle between platforms pushed us to find ways around it and interesting solutions outside of just pure data science. We set up a bunch of Slack bots that would send us our training losses &amp; scores, and notify us when training was done. This way we could run training while working on something else, download weights &amp; results and shut down the kernel to maximize GPU time usage. Nothing mind blowing, but a good addition towards optimizing our limited resources.</li>\n</ul>\n\n<h2><strong>Technical aspects &amp; challenges of the competition</strong>:</h2>\n\n<ul>\n<li>Having limited hardware should have pushed us towards compressing the data more to be able to experiment and iterate faster. <a href=\"/cdeotte\">@cdeotte</a> shared some really interesting notebooks all along on how dividing the image sizes by 4 or even 5 helped him iterate much faster. This is definitely something we should have looked more into. Overall, thanks <a href=\"/cdeotte\">@cdeotte</a> for all the stuff you posted, there were some really nice tips for all the of us on limited hardware!</li>\n<li>Organization, keeping track of what was done and what needed to be done was better than our previous competitions, but could still see some improvements. Kaggle is terrible for version control and keeping track of notebooks, as it's basically non-existent. So we need to implement something ourselves to overcome this, and keep tracks of all the logs, weights &amp; preds. Other than learning about data science methods, I learned a lot about how to organize myself on a multi-month project a bit better than the previous times.</li>\n<li>Overall, I feel like we didn't get to try nearly as much things as we had in mind. I'm frustrated because what held us back wasn't lack of ideas and things we wanted to try, but loosing a lot of time implementing them &amp; slow training at times. I'd say the bottleneck came from 80% of slow implementations, and 20% on long training times. This means I have to improve a lot on implementing better implementations and reduce training times with different methods.</li>\n<li>Because of the slow implementations &amp; long training times, I sort of got lost into only trying to get one or two things done, when I should have tried to simplify to iterate faster.</li>\n<li>I tried prioritizing based on past experience and on the error analysis of my models. This helps figure out which was the most important out of: post-processing, fine-tuning the model, changing the model, making a classifier to eliminate images without labels, etc. </li>\n</ul>\n\n<h1><strong>Take-aways from top Kagglers after the competition</strong></h1>\n\n<ul>\n<li>There were some really good results even with (relatviely) simple pipelines.</li>\n<li>People tested a lot of different backbones, models, losses, to find which ones were producing the best results.</li>\n<li>Unet was used a lot, but FPN seems to be pretty good too.</li>\n<li>This competition had very noisy data. Ensembling seems to be pretty common to overcome this.</li>\n<li>There were some really interesting ideas:\n<ul><li>Doing some very selective pseduo-labelling to images that were carefully selected.</li>\n<li>Trying different types of resolutions / image augmentations techniques</li>\n<li>Public set probing. That's not really where I want to go personally as a data scientist, but it's always good to know what's the idea behind and how to do it.</li>\n<li>Weighted ensembling depending on the labels</li>\n<li>Using a classifier as a segmentation model itself. Or using a segmentation model to classify the images. </li></ul></li>\n</ul>\n\n<h1><strong>Conclusion</strong></h1>\n\n<p>This was a great learning experience.\nWhat I think is the biggest thing for me is being able to organize myself better &amp; be able to create a working pipeline as fast as possible. By that I not only mean over fitting to a small dataset, but also trying things like lowering the resolution of the images to try out multiple models, backbones, losses, etc. at an early stage.\nWorking only on free cloud computing resources is actually pretty fun. It's a great way to learn, trying to find as much tricks and ways around the limitations that are in place. One of my personal goals is to get a medal only with my laptop. I'm more interested in learning how to optimize my workflow rather than going up with raw compute power.</p>\n\n<p>I hope this is helpful for other people that might not be at the top of the leader board, but that are also trying to learn from these competitions! I'd also be curious to hear about any other middle or low scoring Kaggler with issues they have faced; as well as what higher ranking Kagglers might have to say about my debrief of this competition.</p>",
      "rawMarkdown": "Hi!\n\nThis was our first segmentation competition, and I think both my team mate and I learned a lot in this context!\n\nEDIT: This actually turned into a pretty long post. Half of it is about organisation, the other half about pure data science.\n\nIn this post I want to bring a few things forward that I have learned, from someone that is still learning a lot in computer vision. I think it's interesting to not only have the perspective of the best Kagglers, but also those in the middle &amp; bottom; everybody is learning something different and sometimes going back to the basics is helpful.\n\nFirst of all, a few notes on what I was using:\n- 1 Unet for each label. Same one, pre-trained on one label with 20 epochs then using those weights to learn with 5 epochs on the other 3 labels. Might not have been a great idea but did give interesting results while keeping the training time down.\n- Backbone: ResNet34\n- Batch size: 32\n- Resized to output size: 350, 525\n- Ended up not having time to add much augmentation nor threshold search for the predictions.\n- Implemented a Classifier (InceptionResnetV2 backbone) to try to improve results by finding images with no masks. Turned out to spend way to much time on that for what it produced.\n\nFor the record, we were using Keras for everything, so not everything might be applicable.\n\nI'd like to separate this in a few categories:\n\n# **Personal take-aways**\n\n## **Organisation &amp; handling the competition**\n- This competition was pretty tricky because of the labels. Some of them had some oddly specific regions, and as we all know, they were quite noisy / subjective. That did make for some very interesting reflecting though. Pretty interesting and great exercise for a first segmentation competition.\n- On a personal note, I approach Kaggle competitions from a **learning** perspective over a **competitive** one. So while we didn't get a great ranking, I see this as a good success, as I learned a lot from it. This also means it gets tricky to balance out what **I want to do myself** against **what I could use from others &amp; implement**. It's just like how we are not using numpy to make out own models, where's the limit. I found it interesting to hard code my own UNet and then move towards the widely used [Segmentation model library](https://github.com/qubvel/segmentation_models).\n- As the images were pretty big and had 4 masks each, fitting everything in RAM wasn't doable on Kaggle notebooks. So this force us to dive deeper in making custom data generators. I started making my own (with a function instead of a class, that turned out to be a great learning experience, mostly about why *not* to do it and turn towards suing classes instead) before using one of the custom classes posted in the public notebooks and refactoring it to my own code.\n(For example we made our own encoding / decoding functions, which from a competition stand-point was probably dumping time out of the window, but from a learning experience was great).\n- My workflow basically consisted of working on my own at first, then discussing ideas / issues with my team mate, and finally going to the forum / public notebook if we were still stuck. This is a basic yet very effective (IMHO) way of learning: trying stuff on your own, not getting influenced by outside factors, and only finding help when you're stuck. So there was a lot of re-inventing the wheel, but that's also great to learn AND understand why things are done one way or another.\n- We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights, predictions, and images + labels.\n- Having limited options and having to juggle between platforms pushed us to find ways around it and interesting solutions outside of just pure data science. We set up a bunch of Slack bots that would send us our training losses &amp; scores, and notify us when training was done. This way we could run training while working on something else, download weights &amp; results and shut down the kernel to maximize GPU time usage. Nothing mind blowing, but a good addition towards optimizing our limited resources.\n\n## **Technical aspects &amp; challenges of the competition**:\n\n- Having limited hardware should have pushed us towards compressing the data more to be able to experiment and iterate faster. @cdeotte shared some really interesting notebooks all along on how dividing the image sizes by 4 or even 5 helped him iterate much faster. This is definitely something we should have looked more into. Overall, thanks @cdeotte for all the stuff you posted, there were some really nice tips for all the of us on limited hardware!\n- Organization, keeping track of what was done and what needed to be done was better than our previous competitions, but could still see some improvements. Kaggle is terrible for version control and keeping track of notebooks, as it's basically non-existent. So we need to implement something ourselves to overcome this, and keep tracks of all the logs, weights &amp; preds. Other than learning about data science methods, I learned a lot about how to organize myself on a multi-month project a bit better than the previous times.\n- Overall, I feel like we didn't get to try nearly as much things as we had in mind. I'm frustrated because what held us back wasn't lack of ideas and things we wanted to try, but loosing a lot of time implementing them &amp; slow training at times. I'd say the bottleneck came from 80% of slow implementations, and 20% on long training times. This means I have to improve a lot on implementing better implementations and reduce training times with different methods.\n- Because of the slow implementations &amp; long training times, I sort of got lost into only trying to get one or two things done, when I should have tried to simplify to iterate faster.\n- I tried prioritizing based on past experience and on the error analysis of my models. This helps figure out which was the most important out of: post-processing, fine-tuning the model, changing the model, making a classifier to eliminate images without labels, etc. \n\n\n# **Take-aways from top Kagglers after the competition**\n- There were some really good results even with (relatviely) simple pipelines.\n- People tested a lot of different backbones, models, losses, to find which ones were producing the best results.\n- Unet was used a lot, but FPN seems to be pretty good too.\n- This competition had very noisy data. Ensembling seems to be pretty common to overcome this.\n- There were some really interesting ideas:\n      - Doing some very selective pseduo-labelling to images that were carefully selected.\n      - Trying different types of resolutions / image augmentations techniques\n      - Public set probing. That's not really where I want to go personally as a data scientist, but it's always good to know what's the idea behind and how to do it.\n      - Weighted ensembling depending on the labels\n      - Using a classifier as a segmentation model itself. Or using a segmentation model to classify the images. \n\n\n# **Conclusion**\n\nThis was a great learning experience.\nWhat I think is the biggest thing for me is being able to organize myself better &amp; be able to create a working pipeline as fast as possible. By that I not only mean over fitting to a small dataset, but also trying things like lowering the resolution of the images to try out multiple models, backbones, losses, etc. at an early stage.\nWorking only on free cloud computing resources is actually pretty fun. It's a great way to learn, trying to find as much tricks and ways around the limitations that are in place. One of my personal goals is to get a medal only with my laptop. I'm more interested in learning how to optimize my workflow rather than going up with raw compute power.\n\nI hope this is helpful for other people that might not be at the top of the leader board, but that are also trying to learn from these competitions! I'd also be curious to hear about any other middle or low scoring Kaggler with issues they have faced; as well as what higher ranking Kagglers might have to say about my debrief of this competition.",
      "votes": null
    },
    {
      "id": "678615",
      "postDate": "11/21/2019 16:43:02",
      "content": "<p>I resonate everything that you said here . It's like someone has put words out of my mouth . I have learnt programming recently . Ideas are abundant in my mind because I read a lot but my slow implementation , getting stuck in implementing some idea for a week or two and then not getting success and moving on sometimes deters me a lot  . That also means I need lot of more time to digest a competition . I can't join a competition at last 20 days and do wonders like many people do . These are practical challenges and I think I should overcome them slowly and steadily . Each competition i try something new and they are only adding to what I can take forward in next completion . I believe all masters and grandmasters also had a fare share of the same . </p>\n\n<p>\"We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights\"\n- This was a nightmare to switch between Kaggle Kernel , Google colab , GitHub. Personal laptop and wasted lot of time and some money too in doing so .</p>",
      "rawMarkdown": "I resonate everything that you said here . It's like someone has put words out of my mouth . I have learnt programming recently . Ideas are abundant in my mind because I read a lot but my slow implementation , getting stuck in implementing some idea for a week or two and then not getting success and moving on sometimes deters me a lot  . That also means I need lot of more time to digest a competition . I can't join a competition at last 20 days and do wonders like many people do . These are practical challenges and I think I should overcome them slowly and steadily . Each competition i try something new and they are only adding to what I can take forward in next completion . I believe all masters and grandmasters also had a fare share of the same . \n\n\"We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights\"\n- This was a nightmare to switch between Kaggle Kernel , Google colab , GitHub. Personal laptop and wasted lot of time and some money too in doing so .",
      "votes": null
    },
    {
      "id": "678748",
      "postDate": "11/21/2019 20:00:32",
      "content": "<p>Thanks for the feedback!\nIt's always good to read we're not the only ones in that situation! And I do believe that indeed, a lot of the people at the top started modestly as well.\nData science blew up lately, so I'm sure there are a lot of us out there trying to catch up with people that have been at it for years. But it makes for interesting challenges!</p>\n\n<p>I also believe having a lot of people that are much better at these tasks than us is a great opportunity to learn.</p>\n\n<p>Though I see you're 26th in the competition, that's amazing already! Seems like you got things figured out pretty well!</p>",
      "rawMarkdown": "Thanks for the feedback!\nIt's always good to read we're not the only ones in that situation! And I do believe that indeed, a lot of the people at the top started modestly as well.\nData science blew up lately, so I'm sure there are a lot of us out there trying to catch up with people that have been at it for years. But it makes for interesting challenges!\n\nI also believe having a lot of people that are much better at these tasks than us is a great opportunity to learn.\n\nThough I see you're 26th in the competition, that's amazing already! Seems like you got things figured out pretty well!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 678615,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/21/2019 16:43:02",
      "content": "<p>I resonate everything that you said here . It's like someone has put words out of my mouth . I have learnt programming recently . Ideas are abundant in my mind because I read a lot but my slow implementation , getting stuck in implementing some idea for a week or two and then not getting success and moving on sometimes deters me a lot  . That also means I need lot of more time to digest a competition . I can't join a competition at last 20 days and do wonders like many people do . These are practical challenges and I think I should overcome them slowly and steadily . Each competition i try something new and they are only adding to what I can take forward in next completion . I believe all masters and grandmasters also had a fare share of the same . </p>\n\n<p>\"We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights\"\n- This was a nightmare to switch between Kaggle Kernel , Google colab , GitHub. Personal laptop and wasted lot of time and some money too in doing so .</p>",
      "votes": null,
      "replies": [
        {
          "id": 678748,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "11/21/2019 20:00:32",
          "content": "<p>Thanks for the feedback!\nIt's always good to read we're not the only ones in that situation! And I do believe that indeed, a lot of the people at the top started modestly as well.\nData science blew up lately, so I'm sure there are a lot of us out there trying to catch up with people that have been at it for years. But it makes for interesting challenges!</p>\n\n<p>I also believe having a lot of people that are much better at these tasks than us is a great opportunity to learn.</p>\n\n<p>Though I see you're 26th in the competition, that's amazing already! Seems like you got things figured out pretty well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "677925": "Hi!\n\nThis was our first segmentation competition, and I think both my team mate and I learned a lot in this context!\n\nEDIT: This actually turned into a pretty long post. Half of it is about organisation, the other half about pure data science.\n\nIn this post I want to bring a few things forward that I have learned, from someone that is still learning a lot in computer vision. I think it's interesting to not only have the perspective of the best Kagglers, but also those in the middle &amp; bottom; everybody is learning something different and sometimes going back to the basics is helpful.\n\nFirst of all, a few notes on what I was using:\n- 1 Unet for each label. Same one, pre-trained on one label with 20 epochs then using those weights to learn with 5 epochs on the other 3 labels. Might not have been a great idea but did give interesting results while keeping the training time down.\n- Backbone: ResNet34\n- Batch size: 32\n- Resized to output size: 350, 525\n- Ended up not having time to add much augmentation nor threshold search for the predictions.\n- Implemented a Classifier (InceptionResnetV2 backbone) to try to improve results by finding images with no masks. Turned out to spend way to much time on that for what it produced.\n\nFor the record, we were using Keras for everything, so not everything might be applicable.\n\nI'd like to separate this in a few categories:\n\n# **Personal take-aways**\n\n## **Organisation &amp; handling the competition**\n- This competition was pretty tricky because of the labels. Some of them had some oddly specific regions, and as we all know, they were quite noisy / subjective. That did make for some very interesting reflecting though. Pretty interesting and great exercise for a first segmentation competition.\n- On a personal note, I approach Kaggle competitions from a **learning** perspective over a **competitive** one. So while we didn't get a great ranking, I see this as a good success, as I learned a lot from it. This also means it gets tricky to balance out what **I want to do myself** against **what I could use from others &amp; implement**. It's just like how we are not using numpy to make out own models, where's the limit. I found it interesting to hard code my own UNet and then move towards the widely used [Segmentation model library](https://github.com/qubvel/segmentation_models).\n- As the images were pretty big and had 4 masks each, fitting everything in RAM wasn't doable on Kaggle notebooks. So this force us to dive deeper in making custom data generators. I started making my own (with a function instead of a class, that turned out to be a great learning experience, mostly about why *not* to do it and turn towards suing classes instead) before using one of the custom classes posted in the public notebooks and refactoring it to my own code.\n(For example we made our own encoding / decoding functions, which from a competition stand-point was probably dumping time out of the window, but from a learning experience was great).\n- My workflow basically consisted of working on my own at first, then discussing ideas / issues with my team mate, and finally going to the forum / public notebook if we were still stuck. This is a basic yet very effective (IMHO) way of learning: trying stuff on your own, not getting influenced by outside factors, and only finding help when you're stuck. So there was a lot of re-inventing the wheel, but that's also great to learn AND understand why things are done one way or another.\n- We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights, predictions, and images + labels.\n- Having limited options and having to juggle between platforms pushed us to find ways around it and interesting solutions outside of just pure data science. We set up a bunch of Slack bots that would send us our training losses &amp; scores, and notify us when training was done. This way we could run training while working on something else, download weights &amp; results and shut down the kernel to maximize GPU time usage. Nothing mind blowing, but a good addition towards optimizing our limited resources.\n\n## **Technical aspects &amp; challenges of the competition**:\n\n- Having limited hardware should have pushed us towards compressing the data more to be able to experiment and iterate faster. @cdeotte shared some really interesting notebooks all along on how dividing the image sizes by 4 or even 5 helped him iterate much faster. This is definitely something we should have looked more into. Overall, thanks @cdeotte for all the stuff you posted, there were some really nice tips for all the of us on limited hardware!\n- Organization, keeping track of what was done and what needed to be done was better than our previous competitions, but could still see some improvements. Kaggle is terrible for version control and keeping track of notebooks, as it's basically non-existent. So we need to implement something ourselves to overcome this, and keep tracks of all the logs, weights &amp; preds. Other than learning about data science methods, I learned a lot about how to organize myself on a multi-month project a bit better than the previous times.\n- Overall, I feel like we didn't get to try nearly as much things as we had in mind. I'm frustrated because what held us back wasn't lack of ideas and things we wanted to try, but loosing a lot of time implementing them &amp; slow training at times. I'd say the bottleneck came from 80% of slow implementations, and 20% on long training times. This means I have to improve a lot on implementing better implementations and reduce training times with different methods.\n- Because of the slow implementations &amp; long training times, I sort of got lost into only trying to get one or two things done, when I should have tried to simplify to iterate faster.\n- I tried prioritizing based on past experience and on the error analysis of my models. This helps figure out which was the most important out of: post-processing, fine-tuning the model, changing the model, making a classifier to eliminate images without labels, etc. \n\n\n# **Take-aways from top Kagglers after the competition**\n- There were some really good results even with (relatviely) simple pipelines.\n- People tested a lot of different backbones, models, losses, to find which ones were producing the best results.\n- Unet was used a lot, but FPN seems to be pretty good too.\n- This competition had very noisy data. Ensembling seems to be pretty common to overcome this.\n- There were some really interesting ideas:\n      - Doing some very selective pseduo-labelling to images that were carefully selected.\n      - Trying different types of resolutions / image augmentations techniques\n      - Public set probing. That's not really where I want to go personally as a data scientist, but it's always good to know what's the idea behind and how to do it.\n      - Weighted ensembling depending on the labels\n      - Using a classifier as a segmentation model itself. Or using a segmentation model to classify the images. \n\n\n# **Conclusion**\n\nThis was a great learning experience.\nWhat I think is the biggest thing for me is being able to organize myself better &amp; be able to create a working pipeline as fast as possible. By that I not only mean over fitting to a small dataset, but also trying things like lowering the resolution of the images to try out multiple models, backbones, losses, etc. at an early stage.\nWorking only on free cloud computing resources is actually pretty fun. It's a great way to learn, trying to find as much tricks and ways around the limitations that are in place. One of my personal goals is to get a medal only with my laptop. I'm more interested in learning how to optimize my workflow rather than going up with raw compute power.\n\nI hope this is helpful for other people that might not be at the top of the leader board, but that are also trying to learn from these competitions! I'd also be curious to hear about any other middle or low scoring Kaggler with issues they have faced; as well as what higher ranking Kagglers might have to say about my debrief of this competition.",
    "678615": "I resonate everything that you said here . It's like someone has put words out of my mouth . I have learnt programming recently . Ideas are abundant in my mind because I read a lot but my slow implementation , getting stuck in implementing some idea for a week or two and then not getting success and moving on sometimes deters me a lot  . That also means I need lot of more time to digest a competition . I can't join a competition at last 20 days and do wonders like many people do . These are practical challenges and I think I should overcome them slowly and steadily . Each competition i try something new and they are only adding to what I can take forward in next completion . I believe all masters and grandmasters also had a fare share of the same . \n\n\"We basically only have our laptops, and I personally am a student, I can't really afford to put money in cloud computing for now. So organisation was a big challenge in itself, trying to get the most out of free computing solutions. We used a combination of Kaggle Notebooks, Google Colab, Paperspace notebooks and some local computing (mostly for CPU only stages, that didn't require heavy GPU lifting). This made it particularly tricky to move around weights\"\n- This was a nightmare to switch between Kaggle Kernel , Google colab , GitHub. Personal laptop and wasted lot of time and some money too in doing so .",
    "678748": "Thanks for the feedback!\nIt's always good to read we're not the only ones in that situation! And I do believe that indeed, a lot of the people at the top started modestly as well.\nData science blew up lately, so I'm sure there are a lot of us out there trying to catch up with people that have been at it for years. But it makes for interesting challenges!\n\nI also believe having a lot of people that are much better at these tasks than us is a great opportunity to learn.\n\nThough I see you're 26th in the competition, that's amazing already! Seems like you got things figured out pretty well!"
  },
  "source": "meta"
}