{
  "id": 33788,
  "title": "Semantic Segmentation",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/33788",
  "author_name": "",
  "post_date": "2017-05-29T19:56:54.428534Z",
  "votes": 6,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I was debating on the semantic segmentation approach. SegNet is a popular deep network architecture. Is it okay to use such an architecture?</p>",
  "messages": [
    {
      "id": "186936",
      "postDate": "05/29/2017 19:56:54",
      "content": "<p>I was debating on the semantic segmentation approach. SegNet is a popular deep network architecture. Is it okay to use such an architecture?</p>",
      "rawMarkdown": "I was debating on the semantic segmentation approach. SegNet is a popular deep network architecture. Is it okay to use such an architecture?",
      "votes": null
    },
    {
      "id": "186938",
      "postDate": "05/29/2017 20:09:05",
      "content": "<p>Yes, it's totally fine to use semantic segmentation in this competition.</p>",
      "rawMarkdown": "Yes, it's totally fine to use semantic segmentation in this competition.",
      "votes": null
    },
    {
      "id": "186944",
      "postDate": "05/29/2017 20:39:30",
      "content": "<p>Yeah, I also use semantic segmentation</p>",
      "rawMarkdown": "Yeah, I also use semantic segmentation",
      "votes": null
    },
    {
      "id": "186948",
      "postDate": "05/29/2017 20:55:35",
      "content": "<p>Thank you! By the by, did you manually annotate the train data?</p>",
      "rawMarkdown": "Thank you! By the by, did you manually annotate the train data?",
      "votes": null
    },
    {
      "id": "186949",
      "postDate": "05/29/2017 20:56:46",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "186953",
      "postDate": "05/29/2017 21:13:57",
      "content": "<p>No, I'm making the net predict small squares centered at lion coordinates.</p>",
      "rawMarkdown": "No, I'm making the net predict small squares centered at lion coordinates.",
      "votes": null
    },
    {
      "id": "187116",
      "postDate": "05/30/2017 13:57:09",
      "content": "<p>If you don't mind sharing, do you upsample to get 1:1, or is your prediction done on the pooled layers?</p>",
      "rawMarkdown": "If you don't mind sharing, do you upsample to get 1:1, or is your prediction done on the pooled layers?",
      "votes": null
    },
    {
      "id": "187140",
      "postDate": "05/30/2017 15:16:51",
      "content": "<p>If I understood your question correctly then yes, I currently upsample to 1:1 - so the target has the same size as input. I considered making target say 4x smaller (removing two upsampling layers) for better performance, but did not get to try it.</p>",
      "rawMarkdown": "If I understood your question correctly then yes, I currently upsample to 1:1 - so the target has the same size as input. I considered making target say 4x smaller (removing two upsampling layers) for better performance, but did not get to try it.",
      "votes": null
    },
    {
      "id": "187887",
      "postDate": "06/01/2017 12:48:02",
      "content": "<p>I'll save you the trouble - it sucks :P</p>\n\n<p>Those skip connections, it appears, seem really important to getting the net to 'understand' the features it's looking for.</p>",
      "rawMarkdown": "I'll save you the trouble - it sucks :P\n\nThose skip connections, it appears, seem really important to getting the net to 'understand' the features it's looking for.",
      "votes": null
    },
    {
      "id": "187909",
      "postDate": "06/01/2017 14:01:09",
      "content": "<p>thanks for trying anyway, @authman!</p>",
      "rawMarkdown": "thanks for trying anyway, @authman!",
      "votes": null
    },
    {
      "id": "187965",
      "postDate": "06/01/2017 16:45:56",
      "content": "<p>I have tried sort of semantic segmentation but area of background is much larger that area of objects and network converges to 'all zeros'(all background) solution, but maybe it also depends on weights init, loss, etc.</p>",
      "rawMarkdown": "I have tried sort of semantic segmentation but area of background is much larger that area of objects and network converges to 'all zeros'(all background) solution, but maybe it also depends on weights init, loss, etc.",
      "votes": null
    },
    {
      "id": "188189",
      "postDate": "06/02/2017 07:02:33",
      "content": "<p>This can easily be remedied by modifying the loss function - it should \"punish\" mismatch of sealion pixels much more than mismatch of background pixels because the classes are not evenly distributed. It works pretty well... Haven't made it automatic yet though, I still need to tweak values by hand.</p>\n\n<p>I was wondering - do others try to predict sealions versus background and leave the classification (M/F/pups/...) for later steps, or do you try to do it in one step?</p>",
      "rawMarkdown": "This can easily be remedied by modifying the loss function - it should \"punish\" mismatch of sealion pixels much more than mismatch of background pixels because the classes are not evenly distributed. It works pretty well... Haven't made it automatic yet though, I still need to tweak values by hand.\n\nI was wondering - do others try to predict sealions versus background and leave the classification (M/F/pups/...) for later steps, or do you try to do it in one step?",
      "votes": null
    },
    {
      "id": "188193",
      "postDate": "06/02/2017 07:16:18",
      "content": "<p>@kglspl I do it in one steps, mostly for the sake of simplicity. Also since pups are basically stone-like things close to female lions, the net that has this information at detection stage might fare better.</p>",
      "rawMarkdown": "kglspl I do it in one steps, mostly for the sake of simplicity. Also since pups are basically stone-like things close to female lions, the net that has this information at detection stage might fare better.",
      "votes": null
    },
    {
      "id": "188992",
      "postDate": "06/04/2017 19:02:22",
      "content": "<p>@kglspl I even tried to predict number of sea lions in tile, but solution still degrade to all zero solution.</p>\n\n<p>Like:</p>\n\n<pre><code>y_true: [[0 0 0 0 0]\n [1 0 6 1 6]\n [0 0 3 0 2]\n [1 0 0 0 0]\n [1 0 4 0 3]\n [0 0 2 0 1]\n [0 0 0 0 0]\n [1 0 8 0 8]\n [1 0 2 0 2]\n [0 0 0 0 0]]\ny_pred: [[0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]]\n</code></pre>",
      "rawMarkdown": "kglspl I even tried to predict number of sea lions in tile, but solution still degrade to all zero solution.\n\nLike:\n\n    y_true: [[0 0 0 0 0]\n     [1 0 6 1 6]\n     [0 0 3 0 2]\n     [1 0 0 0 0]\n     [1 0 4 0 3]\n     [0 0 2 0 1]\n     [0 0 0 0 0]\n     [1 0 8 0 8]\n     [1 0 2 0 2]\n     [0 0 0 0 0]]\n    y_pred: [[0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]]",
      "votes": null
    },
    {
      "id": "189097",
      "postDate": "06/05/2017 08:31:45",
      "content": "<p>Great question, which I have been studying (as a newbie) over the last week.</p>\n\n<p>For any other people out there struggling to get their head around what all these detection and segmentation methods are, I found these two great videos which I am sure will be old news to most people, but were a big help to me...</p>\n\n<p><a href=\"https://youtu.be/Sb3b0ocD8mI\">CS231n Lecture 13 - Segmentation, soft attention, spatial transformers</a></p>\n\n<p><a href=\"https://youtu.be/_GfPYLNQank\">CS231n Lecture 8 - Localization and Detection</a></p>\n\n<p>In my humble and inexperienced opinion, this is not a Semantic Segmentation problem. It is possibly an Instance Segmentation problem (lecture 13), but primarily an Object Detection problem (lecture 8). If I am appearing pedantic, it is because of the google hours I have wasted looking at algorithms that would only tell me where the sea lions were, when what I actually wanted was algorithms that would identify individuals.</p>\n\n<p>I am aware, however, that the people who are second and third in this competition say they are both using semantic segmentation. Probably this is because they know far more than me. However it may also be a question of terminology (Gods can abuse words in a way that causes havoc for mortals). And may also be because so much of the SS ideas are based on Object Detection methods and that they are using SS cherries on top of an object detection cake. They will tell me if I am wrong.</p>\n\n<p>I don't want to put people off Semantic Segmentation - far from it - but for people taking their first steps in more advanced CNN techniques, I would have a good look at the Object Detection literature first. I spent hours reading SS literature that I couldn't understand because I hadn't read and understood the OD literature properly. Just trying to help people order their learning properly.</p>",
      "rawMarkdown": "Great question, which I have been studying (as a newbie) over the last week.\n\nFor any other people out there struggling to get their head around what all these detection and segmentation methods are, I found these two great videos which I am sure will be old news to most people, but were a big help to me...\n\n[CS231n Lecture 13 - Segmentation, soft attention, spatial transformers][1]\n\n[CS231n Lecture 8 - Localization and Detection][2]\n\n\n  [1]: https://youtu.be/Sb3b0ocD8mI\n  [2]: https://youtu.be/_GfPYLNQank\n\nIn my humble and inexperienced opinion, this is not a Semantic Segmentation problem. It is possibly an Instance Segmentation problem (lecture 13), but primarily an Object Detection problem (lecture 8). If I am appearing pedantic, it is because of the google hours I have wasted looking at algorithms that would only tell me where the sea lions were, when what I actually wanted was algorithms that would identify individuals.\n\nI am aware, however, that the people who are second and third in this competition say they are both using semantic segmentation. Probably this is because they know far more than me. However it may also be a question of terminology (Gods can abuse words in a way that causes havoc for mortals). And may also be because so much of the SS ideas are based on Object Detection methods and that they are using SS cherries on top of an object detection cake. They will tell me if I am wrong.\n\nI don't want to put people off Semantic Segmentation - far from it - but for people taking their first steps in more advanced CNN techniques, I would have a good look at the Object Detection literature first. I spent hours reading SS literature that I couldn't understand because I hadn't read and understood the OD literature properly. Just trying to help people order their learning properly.",
      "votes": null
    },
    {
      "id": "189099",
      "postDate": "06/05/2017 08:41:55",
      "content": "<p>@JamesEverard thanks for great pointers! I think object detection is also a totally valid approach to this problem, as is also density estimation (see e.g. here <a href=\"https://arxiv.org/abs/1608.06197\">https://arxiv.org/abs/1608.06197</a>). I think the beauty of this competition is that there are a lot of approaches possible, and no one knows which will work best.</p>\n\n<p>I'm really using semantic segmentation, not object detection. The reason is that I'm less familiar with object detection pipelines, and they seem more complex, with custom region proposal methods, custom layers, etc. - so I don't think I have enough time to learn them for this competition. Semantic segmentation on the other hand is just a neural network, but counting lions seems harder than with an object detection network - this is what I'm struggling with.</p>",
      "rawMarkdown": "JamesEverard thanks for great pointers! I think object detection is also a totally valid approach to this problem, as is also density estimation (see e.g. here https://arxiv.org/abs/1608.06197). I think the beauty of this competition is that there are a lot of approaches possible, and no one knows which will work best.\n\nI'm really using semantic segmentation, not object detection. The reason is that I'm less familiar with object detection pipelines, and they seem more complex, with custom region proposal methods, custom layers, etc. - so I don't think I have enough time to learn them for this competition. Semantic segmentation on the other hand is just a neural network, but counting lions seems harder than with an object detection network - this is what I'm struggling with.",
      "votes": null
    },
    {
      "id": "189998",
      "postDate": "06/07/2017 03:09:38",
      "content": "<p>@Konstantin Lopuhin When you make the net to predict small squares centered at sea lions, I assume you created your target images by drawing a small square around the provided dot. My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images), do you mind sharing some insights if not too much to ask.  </p>",
      "rawMarkdown": "Konstantin Lopuhin When you make the net to predict small squares centered at sea lions, I assume you created your target images by drawing a small square around the provided dot. My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images), do you mind sharing some insights if not too much to ask.",
      "votes": null
    },
    {
      "id": "190073",
      "postDate": "06/07/2017 06:20:57",
      "content": "<blockquote>\n  <p>My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images)</p>\n</blockquote>\n\n<p>@Samshipengs I used fixed size squares independent on the scale of the image or lion kind. I agree that handling different scales is very important here, so far I'm using scale augmentation when training the neural network.</p>",
      "rawMarkdown": "&gt; My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images)\n\n@Samshipengs I used fixed size squares independent on the scale of the image or lion kind. I agree that handling different scales is very important here, so far I'm using scale augmentation when training the neural network.",
      "votes": null
    },
    {
      "id": "190674",
      "postDate": "06/08/2017 07:16:26",
      "content": "<p>I think object detection will only work followed with fine-grained classification (something like <a href=\"http://www.vision.caltech.edu/visipedia/CUB-200-2011.html\">http://www.vision.caltech.edu/visipedia/CUB-200-2011.html</a> )</p>",
      "rawMarkdown": "I think object detection will only work followed with fine-grained classification (something like http://www.vision.caltech.edu/visipedia/CUB-200-2011.html )",
      "votes": null
    },
    {
      "id": "190736",
      "postDate": "06/08/2017 10:37:08",
      "content": "<p>@JamesEverard I can completely relate! As a newbie myself, it is hard to wrap your head around all these different approaches. However, I do see your point but I suppose it comes down to personal preferences.</p>",
      "rawMarkdown": "JamesEverard I can completely relate! As a newbie myself, it is hard to wrap your head around all these different approaches. However, I do see your point but I suppose it comes down to personal preferences.",
      "votes": null
    },
    {
      "id": "190898",
      "postDate": "06/08/2017 18:55:19",
      "content": "<p>I've been working a lot on this competition last month, using some kind of 'semantic segmentation' approach. The idea is: instead of generating binary masks, using Unet to regress a heatmap generated from the dotted annotation of each type of sea lions, as proposed in many related papers. The heatmap is made by putting gaussian shaped kernels centered at each sea lion. By either counting the number of peaks or just summing over the entire heatmap can get the final results (if normalized).</p>\n\n<p>One interesting finding is, even the model is not designed for semantic segmentation, during the first few epochs of training, the output clearly shows the shape of sea lions, as seen in the attachment. But I didn't look deep into that to see what I can get from it.</p>\n\n<p><del>The model takes a long time to train and tune, and it seems the scale variation is a big challenge as many have pointed out.  It took me too much time on this competition and I finally decide not to continue since I have some other important things to do :(  , share those ideas and findings to see if it helps</del></p>\n\n<p>updated 06/16/2017:\nI'm back</p>",
      "rawMarkdown": "I've been working a lot on this competition last month, using some kind of 'semantic segmentation' approach. The idea is: instead of generating binary masks, using Unet to regress a heatmap generated from the dotted annotation of each type of sea lions, as proposed in many related papers. The heatmap is made by putting gaussian shaped kernels centered at each sea lion. By either counting the number of peaks or just summing over the entire heatmap can get the final results (if normalized).\n\nOne interesting finding is, even the model is not designed for semantic segmentation, during the first few epochs of training, the output clearly shows the shape of sea lions, as seen in the attachment. But I didn't look deep into that to see what I can get from it.\n\n<del>The model takes a long time to train and tune, and it seems the scale variation is a big challenge as many have pointed out.  It took me too much time on this competition and I finally decide not to continue since I have some other important things to do :(  , share those ideas and findings to see if it helps</del>\n\nupdated 06/16/2017:\nI'm back",
      "votes": null
    },
    {
      "id": "190919",
      "postDate": "06/08/2017 20:23:04",
      "content": "<p>Liangkai,</p>\n\n<p>Thanks a lot for your comment. I have been contemplating to give up for a few days too.</p>\n\n<p>I have a similar approach using unet, unet modified in pixel classifier (doesn't perform on males due to the data imbalance) and a derivate of densenet. My masks are generated with some bounding boxes since I tried traditional approaches at first.\nWould you mind providing the link to the papers you are referring to regarding heat-map regression.</p>",
      "rawMarkdown": "Liangkai,\n\nThanks a lot for your comment. I have been contemplating to give up for a few days too.\n\nI have a similar approach using unet, unet modified in pixel classifier (doesn't perform on males due to the data imbalance) and a derivate of densenet. My masks are generated with some bounding boxes since I tried traditional approaches at first.\nWould you mind providing the link to the papers you are referring to regarding heat-map regression.",
      "votes": null
    },
    {
      "id": "190929",
      "postDate": "06/08/2017 20:59:54",
      "content": "<p>Hi eagle4,</p>\n\n<p>Those are some papers I referenced, good luck!</p>\n\n<p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf</a></p>\n\n<p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf</a></p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.10118.pdf\">https://arxiv.org/pdf/1705.10118.pdf</a></p>\n\n<p><a href=\"http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf\">http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf</a></p>",
      "rawMarkdown": "Hi eagle4,\n\nThose are some papers I referenced, good luck!\n\nhttps://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\n\nhttps://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\n\nhttps://arxiv.org/pdf/1705.10118.pdf\n\nhttp://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf",
      "votes": null
    },
    {
      "id": "190944",
      "postDate": "06/08/2017 21:52:18",
      "content": "<p>I quit too. After countless hours spent on this I just couldn't get a model to beat the benchmark and build from there. But it was a good experience and I have learned a lot along the way.</p>\n\n<p>Considering that the all zeroes benchmark score is 29, and the best soluition score is 12, I dare to say that there are no production ready solutions in this competition. I think the dataset is just too difficult.</p>",
      "rawMarkdown": "I quit too. After countless hours spent on this I just couldn't get a model to beat the benchmark and build from there. But it was a good experience and I have learned a lot along the way.\n\nConsidering that the all zeroes benchmark score is 29, and the best soluition score is 12, I dare to say that there are no production ready solutions in this competition. I think the dataset is just too difficult.",
      "votes": null
    },
    {
      "id": "191260",
      "postDate": "06/09/2017 18:20:12",
      "content": "<p>It's taking me too long to train and too long to generate prediction for all the test images, not sure if I can make it before the deadline. \nBut just regarding the density estimation method, do you have to build a model for each different sea lion classes? That's what I have been doing and seems like really inefficient, on the other hand, object detection looks a better fit (no need to have separate models for different classes, all predictions are made in one go).</p>",
      "rawMarkdown": "It's taking me too long to train and too long to generate prediction for all the test images, not sure if I can make it before the deadline. \nBut just regarding the density estimation method, do you have to build a model for each different sea lion classes? That's what I have been doing and seems like really inefficient, on the other hand, object detection looks a better fit (no need to have separate models for different classes, all predictions are made in one go).",
      "votes": null
    },
    {
      "id": "191262",
      "postDate": "06/09/2017 18:27:17",
      "content": "<p>@Samshipengs, yeah, huge test size is not easy to deal with. But it's definitely possible to apply an image segmentation approach with a model that predicts all classes at once.</p>",
      "rawMarkdown": "Samshipengs, yeah, huge test size is not easy to deal with. But it's definitely possible to apply an image segmentation approach with a model that predicts all classes at once.",
      "votes": null
    },
    {
      "id": "191327",
      "postDate": "06/09/2017 23:53:49",
      "content": "<p>@Konstantin Lopuhin\n Thanks for the tip. I'm not familiar with using segmentation approach and time seems like an issue now, so I will probably stick to what I have and try to make a submission if possible. But can you share any reference on that approach, I would like to read about it and learn.</p>",
      "rawMarkdown": "Konstantin Lopuhin\n Thanks for the tip. I'm not familiar with using segmentation approach and time seems like an issue now, so I will probably stick to what I have and try to make a submission if possible. But can you share any reference on that approach, I would like to read about it and learn.",
      "votes": null
    },
    {
      "id": "191385",
      "postDate": "06/10/2017 05:49:31",
      "content": "<p>Wish you luck with the submission!</p>\n\n<p>Re segmentation &amp; multiple classes: you can have a network output multiple classes like this: if you output just one class, then you have a convolution with 1 output channel, and using a sigmoid nonlinearity. To turn it into a multiclass segmentation net, there are two options:</p>\n\n<p>1) replace this last convolution with a conv with 5 output channels (= number of classes), and continue using sigmoid</p>\n\n<p>2) replace this last convolution with a conv with 6 output channels (= number of classes + background), and use softmax.</p>",
      "rawMarkdown": "Wish you luck with the submission!\n\nRe segmentation &amp; multiple classes: you can have a network output multiple classes like this: if you output just one class, then you have a convolution with 1 output channel, and using a sigmoid nonlinearity. To turn it into a multiclass segmentation net, there are two options:\n\n1) replace this last convolution with a conv with 5 output channels (= number of classes), and continue using sigmoid\n\n2) replace this last convolution with a conv with 6 output channels (= number of classes + background), and use softmax.",
      "votes": null
    },
    {
      "id": "191744",
      "postDate": "06/11/2017 16:39:27",
      "content": "<p>@Konstantin If it's ok with you how are you handling this imbalance? As for keras, it does note provide class weight for more than 3 dimenssion currently, I have two ideas:</p>\n\n<ol>\n<li>Changing the source code for sparse_categorical_entropy(its tiresome), need some pointers on these.</li>\n<li>Making 5 different models for each class.</li>\n</ol>\n\n<p>Thanks in advance. </p>",
      "rawMarkdown": "Konstantin If it's ok with you how are you handling this imbalance? As for keras, it does note provide class weight for more than 3 dimenssion currently, I have two ideas:\n\n 1. Changing the source code for sparse_categorical_entropy(its tiresome), need some pointers on these.\n 2. Making 5 different models for each class.\n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "191758",
      "postDate": "06/11/2017 17:16:58",
      "content": "<p>Hey @Daft Vader I'm not doing anything special to handle imbalance. In fact I tried setting a lower weight for background and that made the model work worse, maybe I was doing something wrong though.</p>",
      "rawMarkdown": "Hey @Daft Vader I'm not doing anything special to handle imbalance. In fact I tried setting a lower weight for background and that made the model work worse, maybe I was doing something wrong though.",
      "votes": null
    },
    {
      "id": "193432",
      "postDate": "06/16/2017 12:59:02",
      "content": "<p>In my opinion, second model (6 output + softmax) is better than first model (5 output + sigmoid). Because, one sea-lion should not belong to more than two classes.<br>I want to know whether first model have any advantages. </p>",
      "rawMarkdown": "In my opinion, second model (6 output + softmax) is better than first model (5 output + sigmoid). Because, one sea-lion should not belong to more than two classes.<br>I want to know whether first model have any advantages.",
      "votes": null
    },
    {
      "id": "193433",
      "postDate": "06/16/2017 13:05:00",
      "content": "<p>I agree, I'm using the second option. I have the first option implemented for some reason, but didn't try it seriously.</p>",
      "rawMarkdown": "I agree, I'm using the second option. I have the first option implemented for some reason, but didn't try it seriously.",
      "votes": null
    },
    {
      "id": "193605",
      "postDate": "06/17/2017 03:29:24",
      "content": "<p>@Konstantin, I was wondering what loss function are you using ( if you are okay with sharing it ) ? \nAccording to resources on internet, cross entropy doesnt seem to be an option since it slows down training when prediction is almost correct ( since gradients dec. when loss is less). \nDice coefficient can be used, but it doesnt have any established way of taking into account the class weights.</p>\n\n<p>Also, I guess the input tile size for semantic segmentation will significantly affect the results too. </p>",
      "rawMarkdown": "Konstantin, I was wondering what loss function are you using ( if you are okay with sharing it ) ? \nAccording to resources on internet, cross entropy doesnt seem to be an option since it slows down training when prediction is almost correct ( since gradients dec. when loss is less). \nDice coefficient can be used, but it doesnt have any established way of taking into account the class weights.\n\nAlso, I guess the input tile size for semantic segmentation will significantly affect the results too.",
      "votes": null
    },
    {
      "id": "193617",
      "postDate": "06/17/2017 05:34:46",
      "content": "<p>@KapilYadav I tried both cross entropy and cross entropy + dice. I'm not using class weights at the moment, tried them briefly and they didn't seem to help, but it may be that they help and I didn't notice. As for the dice loss, you'll be calculating dice coefficients per-class anyway, so it seems you could just multiply class losses by the weights. As for classes vs. background weight for dice, I think that's usually not required, but it's also possible, something along these lines, where alpha controls ratio between desired false positives (fp) and false negatives (fn). This is using PyTorch and not tested:</p>\n\n<pre><code>fp = y_pred * (1 - y)\nfn = (1 - y_pred) * y\n# fp + fn == y_pred.sum() + y.sum() - 2 * (y_pred * y)\ntp = y_pred * y\njaccard = tp / (fp * alpha + fn * (1 - apha) + tp)\n</code></pre>",
      "rawMarkdown": "KapilYadav I tried both cross entropy and cross entropy + dice. I'm not using class weights at the moment, tried them briefly and they didn't seem to help, but it may be that they help and I didn't notice. As for the dice loss, you'll be calculating dice coefficients per-class anyway, so it seems you could just multiply class losses by the weights. As for classes vs. background weight for dice, I think that's usually not required, but it's also possible, something along these lines, where alpha controls ratio between desired false positives (fp) and false negatives (fn). This is using PyTorch and not tested:\n\n    fp = y_pred * (1 - y)\n    fn = (1 - y_pred) * y\n    # fp + fn == y_pred.sum() + y.sum() - 2 * (y_pred * y)\n    tp = y_pred * y\n    jaccard = tp / (fp * alpha + fn * (1 - apha) + tp)",
      "votes": null
    },
    {
      "id": "193618",
      "postDate": "06/17/2017 05:35:26",
      "content": "<blockquote>\n  <p>Also, I guess the input tile size for semantic segmentation will significantly affect the results too.</p>\n</blockquote>\n\n<p>Yeah, I hope so - didn't experiment with it yet but plan to.</p>",
      "rawMarkdown": "&gt; Also, I guess the input tile size for semantic segmentation will significantly affect the results too.\n\nYeah, I hope so - didn't experiment with it yet but plan to.",
      "votes": null
    },
    {
      "id": "193668",
      "postDate": "06/17/2017 13:19:09",
      "content": "<p>@Konstantin, \nThanks for that explanation. I'll try without using background labels and with dice coeff.</p>",
      "rawMarkdown": "Konstantin, \nThanks for that explanation. I'll try without using background labels and with dice coeff.",
      "votes": null
    },
    {
      "id": "196264",
      "postDate": "06/26/2017 22:06:21",
      "content": "<blockquote>\n  <p><strong>Liangkai wrote</strong></p>\n  \n  <blockquote>\n    <p>Hi eagle4,</p>\n  </blockquote>\n  \n  <p>Those are some papers I referenced, good luck!</p>\n  \n  <p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf</a></p>\n  \n  <p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf</a></p>\n  \n  <p><a href=\"https://arxiv.org/pdf/1705.10118.pdf\">https://arxiv.org/pdf/1705.10118.pdf</a></p>\n  \n  <p><a href=\"http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf\">http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf</a></p>\n</blockquote>\n\n<p>That's a great reference. I started this problem with Lempitsky et al. papers and totally missed other good ones. Model described in Lempitsky (a modified one) performed very well on synthetic data, but failed on lions. I explain this by 3 factors: 1) multiple classes (original approach is for a single class), 2) diverse scales of objects, which resulted in huge overestimation of density =&gt; counts 3) extreme variation in object counts per image. In the end I added special post-processing model which resulted in rather complicated pipeline but made the work actually done. Should I have more time I would try other density approaches like the ones you quoted + segmentation.</p>",
      "rawMarkdown": "&gt; **Liangkai wrote**\n&gt; \n&gt; &gt; Hi eagle4,\n&gt; \n&gt; Those are some papers I referenced, good luck!\n&gt; \n&gt; https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\n&gt; \n&gt; https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\n&gt; \n&gt; https://arxiv.org/pdf/1705.10118.pdf\n&gt; \n&gt; http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf\n&gt; \n&gt; \n\nThat's a great reference. I started this problem with Lempitsky et al. papers and totally missed other good ones. Model described in Lempitsky (a modified one) performed very well on synthetic data, but failed on lions. I explain this by 3 factors: 1) multiple classes (original approach is for a single class), 2) diverse scales of objects, which resulted in huge overestimation of density =&gt; counts 3) extreme variation in object counts per image. In the end I added special post-processing model which resulted in rather complicated pipeline but made the work actually done. Should I have more time I would try other density approaches like the ones you quoted + segmentation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 186938,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "05/29/2017 20:09:05",
      "content": "<p>Yes, it's totally fine to use semantic segmentation in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 186949,
          "author_name": "azaftanveer",
          "author_url": "",
          "post_date": "05/29/2017 20:56:46",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 186944,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "05/29/2017 20:39:30",
      "content": "<p>Yeah, I also use semantic segmentation</p>",
      "votes": null,
      "replies": [
        {
          "id": 186948,
          "author_name": "azaftanveer",
          "author_url": "",
          "post_date": "05/29/2017 20:55:35",
          "content": "<p>Thank you! By the by, did you manually annotate the train data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 186953,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/29/2017 21:13:57",
          "content": "<p>No, I'm making the net predict small squares centered at lion coordinates.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 187116,
          "author_name": "authman",
          "author_url": "",
          "post_date": "05/30/2017 13:57:09",
          "content": "<p>If you don't mind sharing, do you upsample to get 1:1, or is your prediction done on the pooled layers?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 187140,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/30/2017 15:16:51",
          "content": "<p>If I understood your question correctly then yes, I currently upsample to 1:1 - so the target has the same size as input. I considered making target say 4x smaller (removing two upsampling layers) for better performance, but did not get to try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 187887,
          "author_name": "authman",
          "author_url": "",
          "post_date": "06/01/2017 12:48:02",
          "content": "<p>I'll save you the trouble - it sucks :P</p>\n\n<p>Those skip connections, it appears, seem really important to getting the net to 'understand' the features it's looking for.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 187909,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/01/2017 14:01:09",
          "content": "<p>thanks for trying anyway, @authman!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 189998,
          "author_name": "",
          "author_url": "",
          "post_date": "06/07/2017 03:09:38",
          "content": "<p>@Konstantin Lopuhin When you make the net to predict small squares centered at sea lions, I assume you created your target images by drawing a small square around the provided dot. My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images), do you mind sharing some insights if not too much to ask.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 190073,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/07/2017 06:20:57",
          "content": "<blockquote>\n  <p>My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images)</p>\n</blockquote>\n\n<p>@Samshipengs I used fixed size squares independent on the scale of the image or lion kind. I agree that handling different scales is very important here, so far I'm using scale augmentation when training the neural network.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 187965,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/01/2017 16:45:56",
      "content": "<p>I have tried sort of semantic segmentation but area of background is much larger that area of objects and network converges to 'all zeros'(all background) solution, but maybe it also depends on weights init, loss, etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 188189,
          "author_name": "kglspl",
          "author_url": "",
          "post_date": "06/02/2017 07:02:33",
          "content": "<p>This can easily be remedied by modifying the loss function - it should \"punish\" mismatch of sealion pixels much more than mismatch of background pixels because the classes are not evenly distributed. It works pretty well... Haven't made it automatic yet though, I still need to tweak values by hand.</p>\n\n<p>I was wondering - do others try to predict sealions versus background and leave the classification (M/F/pups/...) for later steps, or do you try to do it in one step?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188193,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/02/2017 07:16:18",
          "content": "<p>@kglspl I do it in one steps, mostly for the sake of simplicity. Also since pups are basically stone-like things close to female lions, the net that has this information at detection stage might fare better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188992,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/04/2017 19:02:22",
          "content": "<p>@kglspl I even tried to predict number of sea lions in tile, but solution still degrade to all zero solution.</p>\n\n<p>Like:</p>\n\n<pre><code>y_true: [[0 0 0 0 0]\n [1 0 6 1 6]\n [0 0 3 0 2]\n [1 0 0 0 0]\n [1 0 4 0 3]\n [0 0 2 0 1]\n [0 0 0 0 0]\n [1 0 8 0 8]\n [1 0 2 0 2]\n [0 0 0 0 0]]\ny_pred: [[0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]\n [0 0 0 0 0]]\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 191744,
          "author_name": "syeddanish",
          "author_url": "",
          "post_date": "06/11/2017 16:39:27",
          "content": "<p>@Konstantin If it's ok with you how are you handling this imbalance? As for keras, it does note provide class weight for more than 3 dimenssion currently, I have two ideas:</p>\n\n<ol>\n<li>Changing the source code for sparse_categorical_entropy(its tiresome), need some pointers on these.</li>\n<li>Making 5 different models for each class.</li>\n</ol>\n\n<p>Thanks in advance. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 191758,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/11/2017 17:16:58",
          "content": "<p>Hey @Daft Vader I'm not doing anything special to handle imbalance. In fact I tried setting a lower weight for background and that made the model work worse, maybe I was doing something wrong though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 189097,
      "author_name": "jinkos",
      "author_url": "",
      "post_date": "06/05/2017 08:31:45",
      "content": "<p>Great question, which I have been studying (as a newbie) over the last week.</p>\n\n<p>For any other people out there struggling to get their head around what all these detection and segmentation methods are, I found these two great videos which I am sure will be old news to most people, but were a big help to me...</p>\n\n<p><a href=\"https://youtu.be/Sb3b0ocD8mI\">CS231n Lecture 13 - Segmentation, soft attention, spatial transformers</a></p>\n\n<p><a href=\"https://youtu.be/_GfPYLNQank\">CS231n Lecture 8 - Localization and Detection</a></p>\n\n<p>In my humble and inexperienced opinion, this is not a Semantic Segmentation problem. It is possibly an Instance Segmentation problem (lecture 13), but primarily an Object Detection problem (lecture 8). If I am appearing pedantic, it is because of the google hours I have wasted looking at algorithms that would only tell me where the sea lions were, when what I actually wanted was algorithms that would identify individuals.</p>\n\n<p>I am aware, however, that the people who are second and third in this competition say they are both using semantic segmentation. Probably this is because they know far more than me. However it may also be a question of terminology (Gods can abuse words in a way that causes havoc for mortals). And may also be because so much of the SS ideas are based on Object Detection methods and that they are using SS cherries on top of an object detection cake. They will tell me if I am wrong.</p>\n\n<p>I don't want to put people off Semantic Segmentation - far from it - but for people taking their first steps in more advanced CNN techniques, I would have a good look at the Object Detection literature first. I spent hours reading SS literature that I couldn't understand because I hadn't read and understood the OD literature properly. Just trying to help people order their learning properly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 189099,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/05/2017 08:41:55",
          "content": "<p>@JamesEverard thanks for great pointers! I think object detection is also a totally valid approach to this problem, as is also density estimation (see e.g. here <a href=\"https://arxiv.org/abs/1608.06197\">https://arxiv.org/abs/1608.06197</a>). I think the beauty of this competition is that there are a lot of approaches possible, and no one knows which will work best.</p>\n\n<p>I'm really using semantic segmentation, not object detection. The reason is that I'm less familiar with object detection pipelines, and they seem more complex, with custom region proposal methods, custom layers, etc. - so I don't think I have enough time to learn them for this competition. Semantic segmentation on the other hand is just a neural network, but counting lions seems harder than with an object detection network - this is what I'm struggling with.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 190674,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/08/2017 07:16:26",
          "content": "<p>I think object detection will only work followed with fine-grained classification (something like <a href=\"http://www.vision.caltech.edu/visipedia/CUB-200-2011.html\">http://www.vision.caltech.edu/visipedia/CUB-200-2011.html</a> )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 190736,
          "author_name": "azaftanveer",
          "author_url": "",
          "post_date": "06/08/2017 10:37:08",
          "content": "<p>@JamesEverard I can completely relate! As a newbie myself, it is hard to wrap your head around all these different approaches. However, I do see your point but I suppose it comes down to personal preferences.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 190898,
      "author_name": "lzhang57",
      "author_url": "",
      "post_date": "06/08/2017 18:55:19",
      "content": "<p>I've been working a lot on this competition last month, using some kind of 'semantic segmentation' approach. The idea is: instead of generating binary masks, using Unet to regress a heatmap generated from the dotted annotation of each type of sea lions, as proposed in many related papers. The heatmap is made by putting gaussian shaped kernels centered at each sea lion. By either counting the number of peaks or just summing over the entire heatmap can get the final results (if normalized).</p>\n\n<p>One interesting finding is, even the model is not designed for semantic segmentation, during the first few epochs of training, the output clearly shows the shape of sea lions, as seen in the attachment. But I didn't look deep into that to see what I can get from it.</p>\n\n<p><del>The model takes a long time to train and tune, and it seems the scale variation is a big challenge as many have pointed out.  It took me too much time on this competition and I finally decide not to continue since I have some other important things to do :(  , share those ideas and findings to see if it helps</del></p>\n\n<p>updated 06/16/2017:\nI'm back</p>",
      "votes": null,
      "replies": [
        {
          "id": 190919,
          "author_name": "chabir",
          "author_url": "",
          "post_date": "06/08/2017 20:23:04",
          "content": "<p>Liangkai,</p>\n\n<p>Thanks a lot for your comment. I have been contemplating to give up for a few days too.</p>\n\n<p>I have a similar approach using unet, unet modified in pixel classifier (doesn't perform on males due to the data imbalance) and a derivate of densenet. My masks are generated with some bounding boxes since I tried traditional approaches at first.\nWould you mind providing the link to the papers you are referring to regarding heat-map regression.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 190929,
          "author_name": "lzhang57",
          "author_url": "",
          "post_date": "06/08/2017 20:59:54",
          "content": "<p>Hi eagle4,</p>\n\n<p>Those are some papers I referenced, good luck!</p>\n\n<p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf</a></p>\n\n<p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf</a></p>\n\n<p><a href=\"https://arxiv.org/pdf/1705.10118.pdf\">https://arxiv.org/pdf/1705.10118.pdf</a></p>\n\n<p><a href=\"http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf\">http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 190944,
          "author_name": "radustoicescu",
          "author_url": "",
          "post_date": "06/08/2017 21:52:18",
          "content": "<p>I quit too. After countless hours spent on this I just couldn't get a model to beat the benchmark and build from there. But it was a good experience and I have learned a lot along the way.</p>\n\n<p>Considering that the all zeroes benchmark score is 29, and the best soluition score is 12, I dare to say that there are no production ready solutions in this competition. I think the dataset is just too difficult.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 191260,
          "author_name": "",
          "author_url": "",
          "post_date": "06/09/2017 18:20:12",
          "content": "<p>It's taking me too long to train and too long to generate prediction for all the test images, not sure if I can make it before the deadline. \nBut just regarding the density estimation method, do you have to build a model for each different sea lion classes? That's what I have been doing and seems like really inefficient, on the other hand, object detection looks a better fit (no need to have separate models for different classes, all predictions are made in one go).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 191262,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/09/2017 18:27:17",
          "content": "<p>@Samshipengs, yeah, huge test size is not easy to deal with. But it's definitely possible to apply an image segmentation approach with a model that predicts all classes at once.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 191327,
          "author_name": "",
          "author_url": "",
          "post_date": "06/09/2017 23:53:49",
          "content": "<p>@Konstantin Lopuhin\n Thanks for the tip. I'm not familiar with using segmentation approach and time seems like an issue now, so I will probably stick to what I have and try to make a submission if possible. But can you share any reference on that approach, I would like to read about it and learn.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 191385,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/10/2017 05:49:31",
          "content": "<p>Wish you luck with the submission!</p>\n\n<p>Re segmentation &amp; multiple classes: you can have a network output multiple classes like this: if you output just one class, then you have a convolution with 1 output channel, and using a sigmoid nonlinearity. To turn it into a multiclass segmentation net, there are two options:</p>\n\n<p>1) replace this last convolution with a conv with 5 output channels (= number of classes), and continue using sigmoid</p>\n\n<p>2) replace this last convolution with a conv with 6 output channels (= number of classes + background), and use softmax.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193432,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "06/16/2017 12:59:02",
          "content": "<p>In my opinion, second model (6 output + softmax) is better than first model (5 output + sigmoid). Because, one sea-lion should not belong to more than two classes.<br>I want to know whether first model have any advantages. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193433,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/16/2017 13:05:00",
          "content": "<p>I agree, I'm using the second option. I have the first option implemented for some reason, but didn't try it seriously.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 196264,
          "author_name": "rakhlin",
          "author_url": "",
          "post_date": "06/26/2017 22:06:21",
          "content": "<blockquote>\n  <p><strong>Liangkai wrote</strong></p>\n  \n  <blockquote>\n    <p>Hi eagle4,</p>\n  </blockquote>\n  \n  <p>Those are some papers I referenced, good luck!</p>\n  \n  <p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf</a></p>\n  \n  <p><a href=\"https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\">https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf</a></p>\n  \n  <p><a href=\"https://arxiv.org/pdf/1705.10118.pdf\">https://arxiv.org/pdf/1705.10118.pdf</a></p>\n  \n  <p><a href=\"http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf\">http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf</a></p>\n</blockquote>\n\n<p>That's a great reference. I started this problem with Lempitsky et al. papers and totally missed other good ones. Model described in Lempitsky (a modified one) performed very well on synthetic data, but failed on lions. I explain this by 3 factors: 1) multiple classes (original approach is for a single class), 2) diverse scales of objects, which resulted in huge overestimation of density =&gt; counts 3) extreme variation in object counts per image. In the end I added special post-processing model which resulted in rather complicated pipeline but made the work actually done. Should I have more time I would try other density approaches like the ones you quoted + segmentation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193605,
      "author_name": "kapilyadav",
      "author_url": "",
      "post_date": "06/17/2017 03:29:24",
      "content": "<p>@Konstantin, I was wondering what loss function are you using ( if you are okay with sharing it ) ? \nAccording to resources on internet, cross entropy doesnt seem to be an option since it slows down training when prediction is almost correct ( since gradients dec. when loss is less). \nDice coefficient can be used, but it doesnt have any established way of taking into account the class weights.</p>\n\n<p>Also, I guess the input tile size for semantic segmentation will significantly affect the results too. </p>",
      "votes": null,
      "replies": [
        {
          "id": 193617,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/17/2017 05:34:46",
          "content": "<p>@KapilYadav I tried both cross entropy and cross entropy + dice. I'm not using class weights at the moment, tried them briefly and they didn't seem to help, but it may be that they help and I didn't notice. As for the dice loss, you'll be calculating dice coefficients per-class anyway, so it seems you could just multiply class losses by the weights. As for classes vs. background weight for dice, I think that's usually not required, but it's also possible, something along these lines, where alpha controls ratio between desired false positives (fp) and false negatives (fn). This is using PyTorch and not tested:</p>\n\n<pre><code>fp = y_pred * (1 - y)\nfn = (1 - y_pred) * y\n# fp + fn == y_pred.sum() + y.sum() - 2 * (y_pred * y)\ntp = y_pred * y\njaccard = tp / (fp * alpha + fn * (1 - apha) + tp)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193618,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/17/2017 05:35:26",
          "content": "<blockquote>\n  <p>Also, I guess the input tile size for semantic segmentation will significantly affect the results too.</p>\n</blockquote>\n\n<p>Yeah, I hope so - didn't experiment with it yet but plan to.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193668,
      "author_name": "kapilyadav",
      "author_url": "",
      "post_date": "06/17/2017 13:19:09",
      "content": "<p>@Konstantin, \nThanks for that explanation. I'll try without using background labels and with dice coeff.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "186936": "I was debating on the semantic segmentation approach. SegNet is a popular deep network architecture. Is it okay to use such an architecture?",
    "186938": "Yes, it's totally fine to use semantic segmentation in this competition.",
    "186944": "Yeah, I also use semantic segmentation",
    "186948": "Thank you! By the by, did you manually annotate the train data?",
    "186949": "Thanks!",
    "186953": "No, I'm making the net predict small squares centered at lion coordinates.",
    "187116": "If you don't mind sharing, do you upsample to get 1:1, or is your prediction done on the pooled layers?",
    "187140": "If I understood your question correctly then yes, I currently upsample to 1:1 - so the target has the same size as input. I considered making target say 4x smaller (removing two upsampling layers) for better performance, but did not get to try it.",
    "187887": "I'll save you the trouble - it sucks :P\n\nThose skip connections, it appears, seem really important to getting the net to 'understand' the features it's looking for.",
    "187909": "thanks for trying anyway, @authman!",
    "187965": "I have tried sort of semantic segmentation but area of background is much larger that area of objects and network converges to 'all zeros'(all background) solution, but maybe it also depends on weights init, loss, etc.",
    "188189": "This can easily be remedied by modifying the loss function - it should \"punish\" mismatch of sealion pixels much more than mismatch of background pixels because the classes are not evenly distributed. It works pretty well... Haven't made it automatic yet though, I still need to tweak values by hand.\n\nI was wondering - do others try to predict sealions versus background and leave the classification (M/F/pups/...) for later steps, or do you try to do it in one step?",
    "188193": "kglspl I do it in one steps, mostly for the sake of simplicity. Also since pups are basically stone-like things close to female lions, the net that has this information at detection stage might fare better.",
    "188992": "kglspl I even tried to predict number of sea lions in tile, but solution still degrade to all zero solution.\n\nLike:\n\n    y_true: [[0 0 0 0 0]\n     [1 0 6 1 6]\n     [0 0 3 0 2]\n     [1 0 0 0 0]\n     [1 0 4 0 3]\n     [0 0 2 0 1]\n     [0 0 0 0 0]\n     [1 0 8 0 8]\n     [1 0 2 0 2]\n     [0 0 0 0 0]]\n    y_pred: [[0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]\n     [0 0 0 0 0]]",
    "189097": "Great question, which I have been studying (as a newbie) over the last week.\n\nFor any other people out there struggling to get their head around what all these detection and segmentation methods are, I found these two great videos which I am sure will be old news to most people, but were a big help to me...\n\n[CS231n Lecture 13 - Segmentation, soft attention, spatial transformers][1]\n\n[CS231n Lecture 8 - Localization and Detection][2]\n\n\n  [1]: https://youtu.be/Sb3b0ocD8mI\n  [2]: https://youtu.be/_GfPYLNQank\n\nIn my humble and inexperienced opinion, this is not a Semantic Segmentation problem. It is possibly an Instance Segmentation problem (lecture 13), but primarily an Object Detection problem (lecture 8). If I am appearing pedantic, it is because of the google hours I have wasted looking at algorithms that would only tell me where the sea lions were, when what I actually wanted was algorithms that would identify individuals.\n\nI am aware, however, that the people who are second and third in this competition say they are both using semantic segmentation. Probably this is because they know far more than me. However it may also be a question of terminology (Gods can abuse words in a way that causes havoc for mortals). And may also be because so much of the SS ideas are based on Object Detection methods and that they are using SS cherries on top of an object detection cake. They will tell me if I am wrong.\n\nI don't want to put people off Semantic Segmentation - far from it - but for people taking their first steps in more advanced CNN techniques, I would have a good look at the Object Detection literature first. I spent hours reading SS literature that I couldn't understand because I hadn't read and understood the OD literature properly. Just trying to help people order their learning properly.",
    "189099": "JamesEverard thanks for great pointers! I think object detection is also a totally valid approach to this problem, as is also density estimation (see e.g. here https://arxiv.org/abs/1608.06197). I think the beauty of this competition is that there are a lot of approaches possible, and no one knows which will work best.\n\nI'm really using semantic segmentation, not object detection. The reason is that I'm less familiar with object detection pipelines, and they seem more complex, with custom region proposal methods, custom layers, etc. - so I don't think I have enough time to learn them for this competition. Semantic segmentation on the other hand is just a neural network, but counting lions seems harder than with an object detection network - this is what I'm struggling with.",
    "189998": "Konstantin Lopuhin When you make the net to predict small squares centered at sea lions, I assume you created your target images by drawing a small square around the provided dot. My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images), do you mind sharing some insights if not too much to ask.",
    "190073": "&gt; My question is, how did you create those small square on target images? Because I saw some of the images have different size for same class sea lions (for example, adult male appear to be in different sizes in some images)\n\n@Samshipengs I used fixed size squares independent on the scale of the image or lion kind. I agree that handling different scales is very important here, so far I'm using scale augmentation when training the neural network.",
    "190674": "I think object detection will only work followed with fine-grained classification (something like http://www.vision.caltech.edu/visipedia/CUB-200-2011.html )",
    "190736": "JamesEverard I can completely relate! As a newbie myself, it is hard to wrap your head around all these different approaches. However, I do see your point but I suppose it comes down to personal preferences.",
    "190898": "I've been working a lot on this competition last month, using some kind of 'semantic segmentation' approach. The idea is: instead of generating binary masks, using Unet to regress a heatmap generated from the dotted annotation of each type of sea lions, as proposed in many related papers. The heatmap is made by putting gaussian shaped kernels centered at each sea lion. By either counting the number of peaks or just summing over the entire heatmap can get the final results (if normalized).\n\nOne interesting finding is, even the model is not designed for semantic segmentation, during the first few epochs of training, the output clearly shows the shape of sea lions, as seen in the attachment. But I didn't look deep into that to see what I can get from it.\n\n<del>The model takes a long time to train and tune, and it seems the scale variation is a big challenge as many have pointed out.  It took me too much time on this competition and I finally decide not to continue since I have some other important things to do :(  , share those ideas and findings to see if it helps</del>\n\nupdated 06/16/2017:\nI'm back",
    "190919": "Liangkai,\n\nThanks a lot for your comment. I have been contemplating to give up for a few days too.\n\nI have a similar approach using unet, unet modified in pixel classifier (doesn't perform on males due to the data imbalance) and a derivate of densenet. My masks are generated with some bounding boxes since I tried traditional approaches at first.\nWould you mind providing the link to the papers you are referring to regarding heat-map regression.",
    "190929": "Hi eagle4,\n\nThose are some papers I referenced, good luck!\n\nhttps://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\n\nhttps://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\n\nhttps://arxiv.org/pdf/1705.10118.pdf\n\nhttp://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf",
    "190944": "I quit too. After countless hours spent on this I just couldn't get a model to beat the benchmark and build from there. But it was a good experience and I have learned a lot along the way.\n\nConsidering that the all zeroes benchmark score is 29, and the best soluition score is 12, I dare to say that there are no production ready solutions in this competition. I think the dataset is just too difficult.",
    "191260": "It's taking me too long to train and too long to generate prediction for all the test images, not sure if I can make it before the deadline. \nBut just regarding the density estimation method, do you have to build a model for each different sea lion classes? That's what I have been doing and seems like really inefficient, on the other hand, object detection looks a better fit (no need to have separate models for different classes, all predictions are made in one go).",
    "191262": "Samshipengs, yeah, huge test size is not easy to deal with. But it's definitely possible to apply an image segmentation approach with a model that predicts all classes at once.",
    "191327": "Konstantin Lopuhin\n Thanks for the tip. I'm not familiar with using segmentation approach and time seems like an issue now, so I will probably stick to what I have and try to make a submission if possible. But can you share any reference on that approach, I would like to read about it and learn.",
    "191385": "Wish you luck with the submission!\n\nRe segmentation &amp; multiple classes: you can have a network output multiple classes like this: if you output just one class, then you have a convolution with 1 output channel, and using a sigmoid nonlinearity. To turn it into a multiclass segmentation net, there are two options:\n\n1) replace this last convolution with a conv with 5 output channels (= number of classes), and continue using sigmoid\n\n2) replace this last convolution with a conv with 6 output channels (= number of classes + background), and use softmax.",
    "191744": "Konstantin If it's ok with you how are you handling this imbalance? As for keras, it does note provide class weight for more than 3 dimenssion currently, I have two ideas:\n\n 1. Changing the source code for sparse_categorical_entropy(its tiresome), need some pointers on these.\n 2. Making 5 different models for each class.\n\nThanks in advance.",
    "191758": "Hey @Daft Vader I'm not doing anything special to handle imbalance. In fact I tried setting a lower weight for background and that made the model work worse, maybe I was doing something wrong though.",
    "193432": "In my opinion, second model (6 output + softmax) is better than first model (5 output + sigmoid). Because, one sea-lion should not belong to more than two classes.<br>I want to know whether first model have any advantages.",
    "193433": "I agree, I'm using the second option. I have the first option implemented for some reason, but didn't try it seriously.",
    "193605": "Konstantin, I was wondering what loss function are you using ( if you are okay with sharing it ) ? \nAccording to resources on internet, cross entropy doesnt seem to be an option since it slows down training when prediction is almost correct ( since gradients dec. when loss is less). \nDice coefficient can be used, but it doesnt have any established way of taking into account the class weights.\n\nAlso, I guess the input tile size for semantic segmentation will significantly affect the results too.",
    "193617": "KapilYadav I tried both cross entropy and cross entropy + dice. I'm not using class weights at the moment, tried them briefly and they didn't seem to help, but it may be that they help and I didn't notice. As for the dice loss, you'll be calculating dice coefficients per-class anyway, so it seems you could just multiply class losses by the weights. As for classes vs. background weight for dice, I think that's usually not required, but it's also possible, something along these lines, where alpha controls ratio between desired false positives (fp) and false negatives (fn). This is using PyTorch and not tested:\n\n    fp = y_pred * (1 - y)\n    fn = (1 - y_pred) * y\n    # fp + fn == y_pred.sum() + y.sum() - 2 * (y_pred * y)\n    tp = y_pred * y\n    jaccard = tp / (fp * alpha + fn * (1 - apha) + tp)",
    "193618": "&gt; Also, I guess the input tile size for semantic segmentation will significantly affect the results too.\n\nYeah, I hope so - didn't experiment with it yet but plan to.",
    "193668": "Konstantin, \nThanks for that explanation. I'll try without using background labels and with dice coeff.",
    "196264": "&gt; **Liangkai wrote**\n&gt; \n&gt; &gt; Hi eagle4,\n&gt; \n&gt; Those are some papers I referenced, good luck!\n&gt; \n&gt; https://www.robots.ox.ac.uk/~vgg/publications/2010/Lempitsky10b/lempitsky10b.pdf\n&gt; \n&gt; https://www.robots.ox.ac.uk/~vgg/publications/2015/Xie15/weidi15.pdf\n&gt; \n&gt; https://arxiv.org/pdf/1705.10118.pdf\n&gt; \n&gt; http://www.cs.tau.ac.il/~wolf/papers/learning-count-cnn.pdf\n&gt; \n&gt; \n\nThat's a great reference. I started this problem with Lempitsky et al. papers and totally missed other good ones. Model described in Lempitsky (a modified one) performed very well on synthetic data, but failed on lions. I explain this by 3 factors: 1) multiple classes (original approach is for a single class), 2) diverse scales of objects, which resulted in huge overestimation of density =&gt; counts 3) extreme variation in object counts per image. In the end I added special post-processing model which resulted in rather complicated pipeline but made the work actually done. Should I have more time I would try other density approaches like the ones you quoted + segmentation."
  },
  "source": "meta"
}