{
  "id": 18299,
  "title": "Code Sharing",
  "url": "/competitions/noaa-right-whale-recognition/discussion/18299",
  "author_name": "",
  "post_date": "2016-01-08T06:09:08.503Z",
  "votes": 2,
  "comment_count": 9,
  "views": 2794,
  "content": "<p>Hi All,</p>\n\n<p>please share your code. we can learn new approaches from other teams.thank you.</p>",
  "messages": [
    {
      "id": "103952",
      "postDate": "01/08/2016 06:09:08",
      "content": "<p>Hi All,</p>\n\n<p>please share your code. we can learn new approaches from other teams.thank you.</p>",
      "rawMarkdown": "Hi All,\r\n\r\nplease share your code. we can learn new approaches from other teams.thank you.",
      "votes": null
    },
    {
      "id": "104002",
      "postDate": "01/08/2016 14:36:55",
      "content": "<p>My approach was very simple. I wasn't planning to enter until Anil posted his code, because the prospect of hand labeling so many images to train a face detector was daunting. Ultimately I ran the neon code and produced some decent crops, and the first run of the classifier gave me a leaderboard score of 4.11725 with 60 iterations. I then tried to run the classifier in CXXNET since I'm more familiar with that tool, but for some reason I could never get a decent result; validation scores dropped to around 5.8 and then went flat. So I went back to Neon, increased image width to 432, and augmented the dataset by adding a top/bottom flip of each image and three contrast adjusted images using PIL.ImageEnhance, with factors of 0, 1, and 2. This gave my best LB score of 3.19730.</p>\n\n<p>At that point I stopped; I couldn't figure out how to do on the fly data augmentation in Neon and for some reason the other tools weren't working. I'm really looking forward to studying the approaches of the top ranking teams.</p>",
      "rawMarkdown": "My approach was very simple. I wasn't planning to enter until Anil posted his code, because the prospect of hand labeling so many images to train a face detector was daunting. Ultimately I ran the neon code and produced some decent crops, and the first run of the classifier gave me a leaderboard score of 4.11725 with 60 iterations. I then tried to run the classifier in CXXNET since I'm more familiar with that tool, but for some reason I could never get a decent result; validation scores dropped to around 5.8 and then went flat. So I went back to Neon, increased image width to 432, and augmented the dataset by adding a top/bottom flip of each image and three contrast adjusted images using PIL.ImageEnhance, with factors of 0, 1, and 2. This gave my best LB score of 3.19730.\r\n\r\nAt that point I stopped; I couldn't figure out how to do on the fly data augmentation in Neon and for some reason the other tools weren't working. I'm really looking forward to studying the approaches of the top ranking teams.",
      "votes": null
    },
    {
      "id": "104003",
      "postDate": "01/08/2016 14:45:48",
      "content": "<p>Thanks for sharing your approach Mr.james King.</p>",
      "rawMarkdown": "Thanks for sharing your approach Mr.james King.",
      "votes": null
    },
    {
      "id": "104019",
      "postDate": "01/08/2016 15:56:12",
      "content": "<p>First I used a whale detector to cut out most of the image. Then I used Anil's annotations and localiser to obtain face crops. I tried different variations of the localiser to obtain 7 different sets of oriented head portrait. I then trained a simple CNN to distinguish good from bad crops and use only the good ones. This worked really well so at this stage I had good crops for all the images except the ones that the initial whale detection has failed (around 400 images which means a penalty around 0.4 logloss).</p>\n\n<p>For the face recognition I used a variant of the VGG network. 10 3x3 convolutional layers, no fully connected layer, number of channels starting from 64 and doubling up to 512 every two layers, max pooling every two layers, leaky relus. What it was crucial for decent performance of the recogniser was on line augmentation. Particular helpfull was RGB color jitter and slight rotations (-10 to 10 degree)  as well as vertical flipping. For the last one I  am not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis.  Nevertheless probably worked because it discouraged learning of useless features which in my  case was more performance hindering than obscuring whale identifiers such as the non centre callosities. Ultimately I believe better augmentation, a more optimised network architecture, a better whale detector (or just skipping this part if I had 12 GB gpu) and possible some model resembling could lead to a logloss around 1.  The 0.59 of the winning team is really impressive and I am eagerly waiting to see their approach.</p>",
      "rawMarkdown": "First I used a whale detector to cut out most of the image. Then I used Anil's annotations and localiser to obtain face crops. I tried different variations of the localiser to obtain 7 different sets of oriented head portrait. I then trained a simple CNN to distinguish good from bad crops and use only the good ones. This worked really well so at this stage I had good crops for all the images except the ones that the initial whale detection has failed (around 400 images which means a penalty around 0.4 logloss).\r\n\r\nFor the face recognition I used a variant of the VGG network. 10 3x3 convolutional layers, no fully connected layer, number of channels starting from 64 and doubling up to 512 every two layers, max pooling every two layers, leaky relus. What it was crucial for decent performance of the recogniser was on line augmentation. Particular helpfull was RGB color jitter and slight rotations (-10 to 10 degree)  as well as vertical flipping. For the last one I  am not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis.  Nevertheless probably worked because it discouraged learning of useless features which in my  case was more performance hindering than obscuring whale identifiers such as the non centre callosities. Ultimately I believe better augmentation, a more optimised network architecture, a better whale detector (or just skipping this part if I had 12 GB gpu) and possible some model resembling could lead to a logloss around 1.  The 0.59 of the winning team is really impressive and I am eagerly waiting to see their approach.",
      "votes": null
    },
    {
      "id": "104020",
      "postDate": "01/08/2016 15:58:19",
      "content": "<p>My path was similar to @James King. Thank you @Anil for the great idea to crop the image.</p>\n\n<p>Since I don't have a maxwell GPU, I borrowed one from a friend to run the localizer to get the crop. After some parameter tuning, I figured the quality is good enough to start classification, I switched to use my own laptop which has a GT755M GPU. I changed some parameters in Anil's crop code to make a smaller crop, then convert the crop to 64*192 gray-scale image since for colored image, I can only train at most 8 layer cnn, but for gray-scale image I can fit a  12-15 layer network in the 2G memory of my GPU. I first tried a 8 layer network in caffe and keras, both with real time augmentation, somehow the keras one runs faster, so I start to use keras exclusively. </p>\n\n<p>I used 5%(about 228 images) as validation set to tune networks and parameters. The 8 layer network give me score of 3.5 on LB. Then I start to add layers in the network, one with 10 cnn layers and 4*4 filter size, another with 13 cnn layer and 3*3 filter size, both have 2 fully connected layer in the end, I also find leakyReLU works better for both of my networks, they score 2.6 and 2.4 on public LB, I get 2.2 when I ensemble them 30%/70%.   I also tried a 15 layer vgg network, but the decrease of validation logloss is too slow, when the competition end,  the validation logloss goes down to 3.2,  and seems to be still decreasing. </p>\n\n<p>I am looking forward to learning the approaches from top players!</p>",
      "rawMarkdown": "My path was similar to @James King. Thank you @Anil for the great idea to crop the image.\r\n\r\nSince I don't have a maxwell GPU, I borrowed one from a friend to run the localizer to get the crop. After some parameter tuning, I figured the quality is good enough to start classification, I switched to use my own laptop which has a GT755M GPU. I changed some parameters in Anil's crop code to make a smaller crop, then convert the crop to 64*192 gray-scale image since for colored image, I can only train at most 8 layer cnn, but for gray-scale image I can fit a  12-15 layer network in the 2G memory of my GPU. I first tried a 8 layer network in caffe and keras, both with real time augmentation, somehow the keras one runs faster, so I start to use keras exclusively. \r\n\r\nI used 5%(about 228 images) as validation set to tune networks and parameters. The 8 layer network give me score of 3.5 on LB. Then I start to add layers in the network, one with 10 cnn layers and 4*4 filter size, another with 13 cnn layer and 3*3 filter size, both have 2 fully connected layer in the end, I also find leakyReLU works better for both of my networks, they score 2.6 and 2.4 on public LB, I get 2.2 when I ensemble them 30%/70%.   I also tried a 15 layer vgg network, but the decrease of validation logloss is too slow, when the competition end,  the validation logloss goes down to 3.2,  and seems to be still decreasing. \r\n\r\nI am looking forward to learning the approaches from top players!",
      "votes": null
    },
    {
      "id": "104074",
      "postDate": "01/08/2016 23:47:21",
      "content": "<p>Our main approach was online data augmentation and specifically trying to match the variation we saw in the data: including water-like noise and white balance to preserve the subtle color differences between the water and the whale.  We also applied this augmentation to the keypoint detection, with more emphasis on geometric augmentation.  In the end, our validation error was still consistently decreasing and matched the LB, which we created through stratified sampling of the whale frequencies.</p>\n\n<p>One path that didn't end up working was joint keypoint prediction via regression of coordinates with @Anil's annotations.  There seemed to be quite the relationship between the blowhole and bonnet, but it ended up performing worse than the separate probability masks of Anil's method.  Anyone have any intuition why? Perhaps the mask yields much more signal per image (since there is image_length^2 labels)?</p>\n\n<p>Props to @Anil on his code, neon looks like a pretty great library.  Also looking forward to the winner reports! Wonder if anyone used the deep-residual approach from the latest ImageNet winner.</p>\n\n<p>[quote=Tsakalis Kostas;104019]\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \n[/quote]</p>\n\n<p>Nice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?</p>\n\n<p>[quote=Tsakalis Kostas;104019]\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \n[/quote]</p>\n\n<p>Perhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always <a href=\"http://arxiv.org/abs/1311.2901\">poke and visualize</a> it if you were really curious.</p>",
      "rawMarkdown": "Our main approach was online data augmentation and specifically trying to match the variation we saw in the data: including water-like noise and white balance to preserve the subtle color differences between the water and the whale.  We also applied this augmentation to the keypoint detection, with more emphasis on geometric augmentation.  In the end, our validation error was still consistently decreasing and matched the LB, which we created through stratified sampling of the whale frequencies.\r\n\r\nOne path that didn't end up working was joint keypoint prediction via regression of coordinates with @Anil's annotations.  There seemed to be quite the relationship between the blowhole and bonnet, but it ended up performing worse than the separate probability masks of Anil's method.  Anyone have any intuition why? Perhaps the mask yields much more signal per image (since there is image_length^2 labels)?\r\n\r\nProps to @Anil on his code, neon looks like a pretty great library.  Also looking forward to the winner reports! Wonder if anyone used the deep-residual approach from the latest ImageNet winner.\r\n\r\n[quote=Tsakalis Kostas;104019]\r\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \r\n[/quote]\r\n\r\nNice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?\r\n\r\n[quote=Tsakalis Kostas;104019]\r\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \r\n[/quote]\r\n\r\nPerhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always [poke and visualize][1] it if you were really curious.\r\n\r\n\r\n  [1]: http://arxiv.org/abs/1311.2901",
      "votes": null
    },
    {
      "id": "104118",
      "postDate": "01/09/2016 07:31:26",
      "content": "<p>I would like to thank here <em>eduardofv</em> for his contribution. Using <em>cv2</em>-based histograms it is possible to achieve level of <strong>4.65</strong> without <em>GPU</em> or any form of hand labeling.</p>",
      "rawMarkdown": "I would like to thank here *eduardofv* for his contribution. Using *cv2*-based histograms it is possible to achieve level of **4.65** without *GPU* or any form of hand labeling.",
      "votes": null
    },
    {
      "id": "104161",
      "postDate": "01/09/2016 20:30:03",
      "content": "<p>[quote=JonathanKChang;104074]</p>\n\n<p>[quote=Tsakalis Kostas;104019]\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \n[/quote]</p>\n\n<p>Nice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?\n[/quote]</p>\n\n<p>I first trained 7 different localisers (slightly different architectures, different target width etc).  I din't used the original images because in order to fit the localiser to 4GB I had to reduce the resolution to 256 which for the whole image looses too much detail. Instead I used eduardofv's unsupervised whale detector to crop a region around whale's body. So if this step fail it cannot be recovered in subsequent steeps. \nNow the surprisingly fact is that for each of the test images at least one of the 7 variants had got the crop right (except if the initial whale detection has failed whereby there was not even a whale in the photo) so what a did was train a binary classification network (architecture similar with the classifier) where I feed the good traincrops as class 0 and from the training images I produce several bad crops by adding noise to the key point coordinates(class 1).\nThen on classification time for each of the 7 crops I run the good-bad classifier and ignore those that are bad.\n[quote=JonathanKChang;104074]</p>\n\n<p>[quote=Tsakalis Kostas;104019]\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \n[/quote]</p>\n\n<p>Perhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always <a href=\"http://arxiv.org/abs/1311.2901\">poke and visualize</a> it if you were really curious.</p>\n\n<p>[/quote]\nTrue, I wish I had more time...</p>",
      "rawMarkdown": "[quote=JonathanKChang;104074]\r\n\r\n[quote=Tsakalis Kostas;104019]\r\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \r\n[/quote]\r\n\r\nNice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?\r\n[/quote]\r\n\r\nI first trained 7 different localisers (slightly different architectures, different target width etc).  I din't used the original images because in order to fit the localiser to 4GB I had to reduce the resolution to 256 which for the whole image looses too much detail. Instead I used eduardofv's unsupervised whale detector to crop a region around whale's body. So if this step fail it cannot be recovered in subsequent steeps. \r\nNow the surprisingly fact is that for each of the test images at least one of the 7 variants had got the crop right (except if the initial whale detection has failed whereby there was not even a whale in the photo) so what a did was train a binary classification network (architecture similar with the classifier) where I feed the good traincrops as class 0 and from the training images I produce several bad crops by adding noise to the key point coordinates(class 1).\r\nThen on classification time for each of the 7 crops I run the good-bad classifier and ignore those that are bad.\r\n[quote=JonathanKChang;104074]\r\n\r\n[quote=Tsakalis Kostas;104019]\r\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \r\n[/quote]\r\n\r\nPerhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always [poke and visualize][1] it if you were really curious.\r\n\r\n\r\n  [1]: http://arxiv.org/abs/1311.2901\r\n\r\n[/quote]\r\nTrue, I wish I had more time...",
      "votes": null
    },
    {
      "id": "105284",
      "postDate": "01/21/2016 19:22:54",
      "content": "<p>Hi everyone,</p>\n\n<p>I'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a <a href=\"https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/\">blog post</a> about my approach and I also share some of my code there. I'd appreciate your feedback and comments!</p>\n\n<p>Andras</p>",
      "rawMarkdown": "Hi everyone,\r\n\r\nI'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a [blog post][1] about my approach and I also share some of my code there. I'd appreciate your feedback and comments!\r\n\r\nAndras\r\n\r\n\r\n  [1]: https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/",
      "votes": null
    },
    {
      "id": "105288",
      "postDate": "01/21/2016 19:56:09",
      "content": "<p>[quote=Andy;105284]</p>\n\n<p>Hi everyone,</p>\n\n<p>I'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a <a href=\"https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/\">blog post</a> about my approach and I also share some of my code there. I'd appreciate your feedback and comments!</p>\n\n<p>Andras</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for sharing. </p>\n\n<p>You're spot on about ensemble methods, and needing to learn them. They are a wonderful addition to the data scientists tool belt. If you haven't already seen this, it's a great place to start:</p>\n\n<p><a href=\"http://mlwave.com/kaggle-ensembling-guide/\">http://mlwave.com/kaggle-ensembling-guide/</a></p>",
      "rawMarkdown": "[quote=Andy;105284]\r\n\r\nHi everyone,\r\n\r\nI'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a [blog post][1] about my approach and I also share some of my code there. I'd appreciate your feedback and comments!\r\n\r\nAndras\r\n\r\n\r\n  [1]: https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/\r\n\r\n[/quote]\r\n\r\nThanks for sharing. \r\n\r\nYou're spot on about ensemble methods, and needing to learn them. They are a wonderful addition to the data scientists tool belt. If you haven't already seen this, it's a great place to start:\r\n\r\nhttp://mlwave.com/kaggle-ensembling-guide/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 104002,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "01/08/2016 14:36:55",
      "content": "<p>My approach was very simple. I wasn't planning to enter until Anil posted his code, because the prospect of hand labeling so many images to train a face detector was daunting. Ultimately I ran the neon code and produced some decent crops, and the first run of the classifier gave me a leaderboard score of 4.11725 with 60 iterations. I then tried to run the classifier in CXXNET since I'm more familiar with that tool, but for some reason I could never get a decent result; validation scores dropped to around 5.8 and then went flat. So I went back to Neon, increased image width to 432, and augmented the dataset by adding a top/bottom flip of each image and three contrast adjusted images using PIL.ImageEnhance, with factors of 0, 1, and 2. This gave my best LB score of 3.19730.</p>\n\n<p>At that point I stopped; I couldn't figure out how to do on the fly data augmentation in Neon and for some reason the other tools weren't working. I'm really looking forward to studying the approaches of the top ranking teams.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104003,
      "author_name": "niranjanmudhiraj",
      "author_url": "",
      "post_date": "01/08/2016 14:45:48",
      "content": "<p>Thanks for sharing your approach Mr.james King.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104019,
      "author_name": "epinephelus",
      "author_url": "",
      "post_date": "01/08/2016 15:56:12",
      "content": "<p>First I used a whale detector to cut out most of the image. Then I used Anil's annotations and localiser to obtain face crops. I tried different variations of the localiser to obtain 7 different sets of oriented head portrait. I then trained a simple CNN to distinguish good from bad crops and use only the good ones. This worked really well so at this stage I had good crops for all the images except the ones that the initial whale detection has failed (around 400 images which means a penalty around 0.4 logloss).</p>\n\n<p>For the face recognition I used a variant of the VGG network. 10 3x3 convolutional layers, no fully connected layer, number of channels starting from 64 and doubling up to 512 every two layers, max pooling every two layers, leaky relus. What it was crucial for decent performance of the recogniser was on line augmentation. Particular helpfull was RGB color jitter and slight rotations (-10 to 10 degree)  as well as vertical flipping. For the last one I  am not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis.  Nevertheless probably worked because it discouraged learning of useless features which in my  case was more performance hindering than obscuring whale identifiers such as the non centre callosities. Ultimately I believe better augmentation, a more optimised network architecture, a better whale detector (or just skipping this part if I had 12 GB gpu) and possible some model resembling could lead to a logloss around 1.  The 0.59 of the winning team is really impressive and I am eagerly waiting to see their approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104020,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "01/08/2016 15:58:19",
      "content": "<p>My path was similar to @James King. Thank you @Anil for the great idea to crop the image.</p>\n\n<p>Since I don't have a maxwell GPU, I borrowed one from a friend to run the localizer to get the crop. After some parameter tuning, I figured the quality is good enough to start classification, I switched to use my own laptop which has a GT755M GPU. I changed some parameters in Anil's crop code to make a smaller crop, then convert the crop to 64*192 gray-scale image since for colored image, I can only train at most 8 layer cnn, but for gray-scale image I can fit a  12-15 layer network in the 2G memory of my GPU. I first tried a 8 layer network in caffe and keras, both with real time augmentation, somehow the keras one runs faster, so I start to use keras exclusively. </p>\n\n<p>I used 5%(about 228 images) as validation set to tune networks and parameters. The 8 layer network give me score of 3.5 on LB. Then I start to add layers in the network, one with 10 cnn layers and 4*4 filter size, another with 13 cnn layer and 3*3 filter size, both have 2 fully connected layer in the end, I also find leakyReLU works better for both of my networks, they score 2.6 and 2.4 on public LB, I get 2.2 when I ensemble them 30%/70%.   I also tried a 15 layer vgg network, but the decrease of validation logloss is too slow, when the competition end,  the validation logloss goes down to 3.2,  and seems to be still decreasing. </p>\n\n<p>I am looking forward to learning the approaches from top players!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104074,
      "author_name": "jonathankchang",
      "author_url": "",
      "post_date": "01/08/2016 23:47:21",
      "content": "<p>Our main approach was online data augmentation and specifically trying to match the variation we saw in the data: including water-like noise and white balance to preserve the subtle color differences between the water and the whale.  We also applied this augmentation to the keypoint detection, with more emphasis on geometric augmentation.  In the end, our validation error was still consistently decreasing and matched the LB, which we created through stratified sampling of the whale frequencies.</p>\n\n<p>One path that didn't end up working was joint keypoint prediction via regression of coordinates with @Anil's annotations.  There seemed to be quite the relationship between the blowhole and bonnet, but it ended up performing worse than the separate probability masks of Anil's method.  Anyone have any intuition why? Perhaps the mask yields much more signal per image (since there is image_length^2 labels)?</p>\n\n<p>Props to @Anil on his code, neon looks like a pretty great library.  Also looking forward to the winner reports! Wonder if anyone used the deep-residual approach from the latest ImageNet winner.</p>\n\n<p>[quote=Tsakalis Kostas;104019]\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \n[/quote]</p>\n\n<p>Nice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?</p>\n\n<p>[quote=Tsakalis Kostas;104019]\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \n[/quote]</p>\n\n<p>Perhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always <a href=\"http://arxiv.org/abs/1311.2901\">poke and visualize</a> it if you were really curious.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104118,
      "author_name": "sunland",
      "author_url": "",
      "post_date": "01/09/2016 07:31:26",
      "content": "<p>I would like to thank here <em>eduardofv</em> for his contribution. Using <em>cv2</em>-based histograms it is possible to achieve level of <strong>4.65</strong> without <em>GPU</em> or any form of hand labeling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 104161,
      "author_name": "epinephelus",
      "author_url": "",
      "post_date": "01/09/2016 20:30:03",
      "content": "<p>[quote=JonathanKChang;104074]</p>\n\n<p>[quote=Tsakalis Kostas;104019]\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \n[/quote]</p>\n\n<p>Nice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?\n[/quote]</p>\n\n<p>I first trained 7 different localisers (slightly different architectures, different target width etc).  I din't used the original images because in order to fit the localiser to 4GB I had to reduce the resolution to 256 which for the whole image looses too much detail. Instead I used eduardofv's unsupervised whale detector to crop a region around whale's body. So if this step fail it cannot be recovered in subsequent steeps. \nNow the surprisingly fact is that for each of the test images at least one of the 7 variants had got the crop right (except if the initial whale detection has failed whereby there was not even a whale in the photo) so what a did was train a binary classification network (architecture similar with the classifier) where I feed the good traincrops as class 0 and from the training images I produce several bad crops by adding noise to the key point coordinates(class 1).\nThen on classification time for each of the 7 crops I run the good-bad classifier and ignore those that are bad.\n[quote=JonathanKChang;104074]</p>\n\n<p>[quote=Tsakalis Kostas;104019]\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \n[/quote]</p>\n\n<p>Perhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always <a href=\"http://arxiv.org/abs/1311.2901\">poke and visualize</a> it if you were really curious.</p>\n\n<p>[/quote]\nTrue, I wish I had more time...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105284,
      "author_name": "andraszsom",
      "author_url": "",
      "post_date": "01/21/2016 19:22:54",
      "content": "<p>Hi everyone,</p>\n\n<p>I'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a <a href=\"https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/\">blog post</a> about my approach and I also share some of my code there. I'd appreciate your feedback and comments!</p>\n\n<p>Andras</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 105288,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "01/21/2016 19:56:09",
      "content": "<p>[quote=Andy;105284]</p>\n\n<p>Hi everyone,</p>\n\n<p>I'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a <a href=\"https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/\">blog post</a> about my approach and I also share some of my code there. I'd appreciate your feedback and comments!</p>\n\n<p>Andras</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for sharing. </p>\n\n<p>You're spot on about ensemble methods, and needing to learn them. They are a wonderful addition to the data scientists tool belt. If you haven't already seen this, it's a great place to start:</p>\n\n<p><a href=\"http://mlwave.com/kaggle-ensembling-guide/\">http://mlwave.com/kaggle-ensembling-guide/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "103952": "Hi All,\r\n\r\nplease share your code. we can learn new approaches from other teams.thank you.",
    "104002": "My approach was very simple. I wasn't planning to enter until Anil posted his code, because the prospect of hand labeling so many images to train a face detector was daunting. Ultimately I ran the neon code and produced some decent crops, and the first run of the classifier gave me a leaderboard score of 4.11725 with 60 iterations. I then tried to run the classifier in CXXNET since I'm more familiar with that tool, but for some reason I could never get a decent result; validation scores dropped to around 5.8 and then went flat. So I went back to Neon, increased image width to 432, and augmented the dataset by adding a top/bottom flip of each image and three contrast adjusted images using PIL.ImageEnhance, with factors of 0, 1, and 2. This gave my best LB score of 3.19730.\r\n\r\nAt that point I stopped; I couldn't figure out how to do on the fly data augmentation in Neon and for some reason the other tools weren't working. I'm really looking forward to studying the approaches of the top ranking teams.",
    "104003": "Thanks for sharing your approach Mr.james King.",
    "104019": "First I used a whale detector to cut out most of the image. Then I used Anil's annotations and localiser to obtain face crops. I tried different variations of the localiser to obtain 7 different sets of oriented head portrait. I then trained a simple CNN to distinguish good from bad crops and use only the good ones. This worked really well so at this stage I had good crops for all the images except the ones that the initial whale detection has failed (around 400 images which means a penalty around 0.4 logloss).\r\n\r\nFor the face recognition I used a variant of the VGG network. 10 3x3 convolutional layers, no fully connected layer, number of channels starting from 64 and doubling up to 512 every two layers, max pooling every two layers, leaky relus. What it was crucial for decent performance of the recogniser was on line augmentation. Particular helpfull was RGB color jitter and slight rotations (-10 to 10 degree)  as well as vertical flipping. For the last one I  am not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis.  Nevertheless probably worked because it discouraged learning of useless features which in my  case was more performance hindering than obscuring whale identifiers such as the non centre callosities. Ultimately I believe better augmentation, a more optimised network architecture, a better whale detector (or just skipping this part if I had 12 GB gpu) and possible some model resembling could lead to a logloss around 1.  The 0.59 of the winning team is really impressive and I am eagerly waiting to see their approach.",
    "104020": "My path was similar to @James King. Thank you @Anil for the great idea to crop the image.\r\n\r\nSince I don't have a maxwell GPU, I borrowed one from a friend to run the localizer to get the crop. After some parameter tuning, I figured the quality is good enough to start classification, I switched to use my own laptop which has a GT755M GPU. I changed some parameters in Anil's crop code to make a smaller crop, then convert the crop to 64*192 gray-scale image since for colored image, I can only train at most 8 layer cnn, but for gray-scale image I can fit a  12-15 layer network in the 2G memory of my GPU. I first tried a 8 layer network in caffe and keras, both with real time augmentation, somehow the keras one runs faster, so I start to use keras exclusively. \r\n\r\nI used 5%(about 228 images) as validation set to tune networks and parameters. The 8 layer network give me score of 3.5 on LB. Then I start to add layers in the network, one with 10 cnn layers and 4*4 filter size, another with 13 cnn layer and 3*3 filter size, both have 2 fully connected layer in the end, I also find leakyReLU works better for both of my networks, they score 2.6 and 2.4 on public LB, I get 2.2 when I ensemble them 30%/70%.   I also tried a 15 layer vgg network, but the decrease of validation logloss is too slow, when the competition end,  the validation logloss goes down to 3.2,  and seems to be still decreasing. \r\n\r\nI am looking forward to learning the approaches from top players!",
    "104074": "Our main approach was online data augmentation and specifically trying to match the variation we saw in the data: including water-like noise and white balance to preserve the subtle color differences between the water and the whale.  We also applied this augmentation to the keypoint detection, with more emphasis on geometric augmentation.  In the end, our validation error was still consistently decreasing and matched the LB, which we created through stratified sampling of the whale frequencies.\r\n\r\nOne path that didn't end up working was joint keypoint prediction via regression of coordinates with @Anil's annotations.  There seemed to be quite the relationship between the blowhole and bonnet, but it ended up performing worse than the separate probability masks of Anil's method.  Anyone have any intuition why? Perhaps the mask yields much more signal per image (since there is image_length^2 labels)?\r\n\r\nProps to @Anil on his code, neon looks like a pretty great library.  Also looking forward to the winner reports! Wonder if anyone used the deep-residual approach from the latest ImageNet winner.\r\n\r\n[quote=Tsakalis Kostas;104019]\r\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \r\n[/quote]\r\n\r\nNice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?\r\n\r\n[quote=Tsakalis Kostas;104019]\r\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \r\n[/quote]\r\n\r\nPerhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always [poke and visualize][1] it if you were really curious.\r\n\r\n\r\n  [1]: http://arxiv.org/abs/1311.2901",
    "104118": "I would like to thank here *eduardofv* for his contribution. Using *cv2*-based histograms it is possible to achieve level of **4.65** without *GPU* or any form of hand labeling.",
    "104161": "[quote=JonathanKChang;104074]\r\n\r\n[quote=Tsakalis Kostas;104019]\r\n I then trained a simple CNN to distinguish good from bad crops and use only the good ones. \r\n[/quote]\r\n\r\nNice idea!  How many labels and what architecture did you choose for that?  For ones where face-detection failed, were the initial whale-detections useful?  Did you learn a separate whale classifier on those?\r\n[/quote]\r\n\r\nI first trained 7 different localisers (slightly different architectures, different target width etc).  I din't used the original images because in order to fit the localiser to 4GB I had to reduce the resolution to 256 which for the whole image looses too much detail. Instead I used eduardofv's unsupervised whale detector to crop a region around whale's body. So if this step fail it cannot be recovered in subsequent steeps. \r\nNow the surprisingly fact is that for each of the test images at least one of the 7 variants had got the crop right (except if the initial whale detection has failed whereby there was not even a whale in the photo) so what a did was train a binary classification network (architecture similar with the classifier) where I feed the good traincrops as class 0 and from the training images I produce several bad crops by adding noise to the key point coordinates(class 1).\r\nThen on classification time for each of the 7 crops I run the good-bad classifier and ignore those that are bad.\r\n[quote=JonathanKChang;104074]\r\n\r\n[quote=Tsakalis Kostas;104019]\r\nvertical flipping... not sure why it worked. It should have made things worse as the whale's mark on their faces are not symmetric across the body axis. \r\n[/quote]\r\n\r\nPerhaps in part it was due to the white water and other occlusions being random and not symmetric, while the whale face shape (not markings) being symmetric.  That kind of whale markings vs water-noise invariance could be learned in other filters alongside the marker-recognizer filters, which indeed may loose a bit of discriminative power by learning flipped invariance.  You could always [poke and visualize][1] it if you were really curious.\r\n\r\n\r\n  [1]: http://arxiv.org/abs/1311.2901\r\n\r\n[/quote]\r\nTrue, I wish I had more time...",
    "105284": "Hi everyone,\r\n\r\nI'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a [blog post][1] about my approach and I also share some of my code there. I'd appreciate your feedback and comments!\r\n\r\nAndras\r\n\r\n\r\n  [1]: https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/",
    "105288": "[quote=Andy;105284]\r\n\r\nHi everyone,\r\n\r\nI'm an astrophysicist who recently turned into a data scientist. This kaggle competition was my first machine learning project so I still have a lot to learn and I hope I whale do better on the next competition! :) I wrote a [blog post][1] about my approach and I also share some of my code there. I'd appreciate your feedback and comments!\r\n\r\nAndras\r\n\r\n\r\n  [1]: https://andraszsom.wordpress.com/2016/01/11/right-whale-recognition-kaggle/\r\n\r\n[/quote]\r\n\r\nThanks for sharing. \r\n\r\nYou're spot on about ensemble methods, and needing to learn them. They are a wonderful addition to the data scientists tool belt. If you haven't already seen this, it's a great place to start:\r\n\r\nhttp://mlwave.com/kaggle-ensembling-guide/"
  },
  "source": "meta"
}