{
  "id": 22018,
  "title": "ideas for moving from LB0.20 to LB0.10",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22018",
  "author_name": "",
  "post_date": "2016-07-03T01:51:10.067Z",
  "votes": 10,
  "comment_count": 16,
  "views": 2524,
  "content": "<p>Time is running short. I have a few ideas to move from LB 0.20. I would be grateful if anyone can help to advise on which of the followings are  most likely to work:</p>\n\n<ul>\n<li>[strategy.1] focus on tweaking the performance of my current solution (e.g ensemble of vgg16, vgg19, googlenet). Sweep hyperparamters like learning rate, batch size. Try relu, prelu, maxout, elu. Try max,ave,max+ave,fractional pool, etc Try heavy augmentation. </li>\n</ul>\n\n<p>e.g <a href=\"https://arxiv.org/abs/1606.02228\">https://arxiv.org/abs/1606.02228</a>\n&quot;Systematic evaluation of CNN advances on the ImageNet&quot;, code: <a href=\"https://github.com/ducha-aiki/caffenet-benchmark\">https://github.com/ducha-aiki/caffenet-benchmark</a></p>\n\n<p>**the problem is that it takes too much time and i estimate best performance to be LB 0.18</p>\n\n<ul>\n<li>[strategy.2]  One problem is lack of data. Try to use traditional classifiers on deep features. E.g. boosting trees/ random forests on features of all layers of deep CNN network. One end-to-end solution is combine traditional + deep is &quot;deep neural forest&quot;:\n<a href=\"http://www.cv-foundation.org/openaccess/content_iccv_2015/html/Kontschieder_Deep_Neural_Decision_ICCV_2015_paper.html\">http://www.cv-foundation.org/openaccess/content_iccv_2015/html/Kontschieder_Deep_Neural_Decision_ICCV_2015_paper.html</a></li>\n</ul>\n\n<p>**I worked with boosting tree+ vgg16 (<a href=\"https://arxiv.org/pdf/1504.07339.pdf\">https://arxiv.org/pdf/1504.07339.pdf</a>, <a href=\"http://arxiv.org/pdf/1603.00124.pdf\">http://arxiv.org/pdf/1603.00124.pdf</a>) and know it outperforms deep learning in 2-class caltech pedestrian detection which has limited training data. But i am not sure if it works for 10-class image classification here.</p>\n\n<ul>\n<li><p>[strategy.3]  restructure the problem. e.g. use hierarchical classification. Group the confusing class (e.g. c0-norm and c9-talk) to one class first, then refine it.e.g. HD-CNN:\n<a href=\"https://arxiv.org/abs/1410.0736\">https://arxiv.org/abs/1410.0736</a></p></li>\n<li><p>[strategy.4]  Change to large margin loss. Since the evaluation criteria is logloss and not accuracy, this should help. Large margin softmax seems promising : <a href=\"http://jmlr.org/proceedings/papers/v48/liud16.pdf\">http://jmlr.org/proceedings/papers/v48/liud16.pdf</a></p></li>\n<li><p>[strategy.5]  Believe that resNet should work, afterall it's is the imagenet winner 2015. Continue to change the resNet32, resNet50 architecture to find one that fits the current kaggle driver data. There are many modified  resNet like wide resNet, pre-activation resNet, etc</p></li>\n</ul>\n\n<p>** I already make several attempts with resNet but couldn't get it work as well as vgg16.</p>\n\n<ul>\n<li>[strategy.6]  Find the best &quot;fine-gradient categorization&quot; paper and implement. Bilinear model seems promising: <a href=\"http://arxiv.org/abs/1504.07889\">http://arxiv.org/abs/1504.07889</a>. These fine grain action recognition papers seems promising too: <a href=\"https://sites.google.com/site/amirrosenfeld/\">https://sites.google.com/site/amirrosenfeld/</a>\n(Hand-Object Interaction and Precise Localization in Transitive Action Recognition)\n(Visual Concept Recognition and Localization via Iterative Introspection)</li>\n</ul>\n\n<p>** my concern is that these may not work well if fundamental issue is lack of data</p>\n\n<ul>\n<li>[strategy.7]  Build and train a network from scratch. Forget about pretrained model. This gives more room to design more different network architecture.</li>\n</ul>\n\n<p>** personally, i like this solution the best. But I not sure if the kaggle train data is enough for training from scratch.</p>",
  "messages": [
    {
      "id": "125823",
      "postDate": "07/03/2016 01:51:10",
      "content": "<p>Time is running short. I have a few ideas to move from LB 0.20. I would be grateful if anyone can help to advise on which of the followings are  most likely to work:</p>\n\n<ul>\n<li>[strategy.1] focus on tweaking the performance of my current solution (e.g ensemble of vgg16, vgg19, googlenet). Sweep hyperparamters like learning rate, batch size. Try relu, prelu, maxout, elu. Try max,ave,max+ave,fractional pool, etc Try heavy augmentation. </li>\n</ul>\n\n<p>e.g <a href=\"https://arxiv.org/abs/1606.02228\">https://arxiv.org/abs/1606.02228</a>\n&quot;Systematic evaluation of CNN advances on the ImageNet&quot;, code: <a href=\"https://github.com/ducha-aiki/caffenet-benchmark\">https://github.com/ducha-aiki/caffenet-benchmark</a></p>\n\n<p>**the problem is that it takes too much time and i estimate best performance to be LB 0.18</p>\n\n<ul>\n<li>[strategy.2]  One problem is lack of data. Try to use traditional classifiers on deep features. E.g. boosting trees/ random forests on features of all layers of deep CNN network. One end-to-end solution is combine traditional + deep is &quot;deep neural forest&quot;:\n<a href=\"http://www.cv-foundation.org/openaccess/content_iccv_2015/html/Kontschieder_Deep_Neural_Decision_ICCV_2015_paper.html\">http://www.cv-foundation.org/openaccess/content_iccv_2015/html/Kontschieder_Deep_Neural_Decision_ICCV_2015_paper.html</a></li>\n</ul>\n\n<p>**I worked with boosting tree+ vgg16 (<a href=\"https://arxiv.org/pdf/1504.07339.pdf\">https://arxiv.org/pdf/1504.07339.pdf</a>, <a href=\"http://arxiv.org/pdf/1603.00124.pdf\">http://arxiv.org/pdf/1603.00124.pdf</a>) and know it outperforms deep learning in 2-class caltech pedestrian detection which has limited training data. But i am not sure if it works for 10-class image classification here.</p>\n\n<ul>\n<li><p>[strategy.3]  restructure the problem. e.g. use hierarchical classification. Group the confusing class (e.g. c0-norm and c9-talk) to one class first, then refine it.e.g. HD-CNN:\n<a href=\"https://arxiv.org/abs/1410.0736\">https://arxiv.org/abs/1410.0736</a></p></li>\n<li><p>[strategy.4]  Change to large margin loss. Since the evaluation criteria is logloss and not accuracy, this should help. Large margin softmax seems promising : <a href=\"http://jmlr.org/proceedings/papers/v48/liud16.pdf\">http://jmlr.org/proceedings/papers/v48/liud16.pdf</a></p></li>\n<li><p>[strategy.5]  Believe that resNet should work, afterall it's is the imagenet winner 2015. Continue to change the resNet32, resNet50 architecture to find one that fits the current kaggle driver data. There are many modified  resNet like wide resNet, pre-activation resNet, etc</p></li>\n</ul>\n\n<p>** I already make several attempts with resNet but couldn't get it work as well as vgg16.</p>\n\n<ul>\n<li>[strategy.6]  Find the best &quot;fine-gradient categorization&quot; paper and implement. Bilinear model seems promising: <a href=\"http://arxiv.org/abs/1504.07889\">http://arxiv.org/abs/1504.07889</a>. These fine grain action recognition papers seems promising too: <a href=\"https://sites.google.com/site/amirrosenfeld/\">https://sites.google.com/site/amirrosenfeld/</a>\n(Hand-Object Interaction and Precise Localization in Transitive Action Recognition)\n(Visual Concept Recognition and Localization via Iterative Introspection)</li>\n</ul>\n\n<p>** my concern is that these may not work well if fundamental issue is lack of data</p>\n\n<ul>\n<li>[strategy.7]  Build and train a network from scratch. Forget about pretrained model. This gives more room to design more different network architecture.</li>\n</ul>\n\n<p>** personally, i like this solution the best. But I not sure if the kaggle train data is enough for training from scratch.</p>",
      "rawMarkdown": "Time is running short. I have a few ideas to move from LB 0.20. I would be grateful if anyone can help to advise on which of the followings are  most likely to work:\r\n\r\n - [strategy.1] focus on tweaking the performance of my current solution (e.g ensemble of vgg16, vgg19, googlenet). Sweep hyperparamters like learning rate, batch size. Try relu, prelu, maxout, elu. Try max,ave,max+ave,fractional pool, etc Try heavy augmentation. \r\n\r\ne.g https://arxiv.org/abs/1606.02228\r\n\"Systematic evaluation of CNN advances on the ImageNet\", code: https://github.com/ducha-aiki/caffenet-benchmark\r\n\r\n**the problem is that it takes too much time and i estimate best performance to be LB 0.18\r\n\r\n\r\n\r\n - [strategy.2]  One problem is lack of data. Try to use traditional classifiers on deep features. E.g. boosting trees/ random forests on features of all layers of deep CNN network. One end-to-end solution is combine traditional + deep is \"deep neural forest\":\r\n http://www.cv-foundation.org/openaccess/content_iccv_2015/html/Kontschieder_Deep_Neural_Decision_ICCV_2015_paper.html\r\n\r\n**I worked with boosting tree+ vgg16 (https://arxiv.org/pdf/1504.07339.pdf, http://arxiv.org/pdf/1603.00124.pdf) and know it outperforms deep learning in 2-class caltech pedestrian detection which has limited training data. But i am not sure if it works for 10-class image classification here.\r\n \r\n - [strategy.3]  restructure the problem. e.g. use hierarchical classification. Group the confusing class (e.g. c0-norm and c9-talk) to one class first, then refine it.e.g. HD-CNN:\r\nhttps://arxiv.org/abs/1410.0736\r\n\r\n\r\n - [strategy.4]  Change to large margin loss. Since the evaluation criteria is logloss and not accuracy, this should help. Large margin softmax seems promising : http://jmlr.org/proceedings/papers/v48/liud16.pdf\r\n\r\n - [strategy.5]  Believe that resNet should work, afterall it's is the imagenet winner 2015. Continue to change the resNet32, resNet50 architecture to find one that fits the current kaggle driver data. There are many modified  resNet like wide resNet, pre-activation resNet, etc\r\n\r\n** I already make several attempts with resNet but couldn't get it work as well as vgg16.\r\n\r\n - [strategy.6]  Find the best \"fine-gradient categorization\" paper and implement. Bilinear model seems promising: http://arxiv.org/abs/1504.07889. These fine grain action recognition papers seems promising too: https://sites.google.com/site/amirrosenfeld/\r\n(Hand-Object Interaction and Precise Localization in Transitive Action Recognition)\r\n(Visual Concept Recognition and Localization via Iterative Introspection)\r\n\r\n\r\n** my concern is that these may not work well if fundamental issue is lack of data\r\n\r\n - [strategy.7]  Build and train a network from scratch. Forget about pretrained model. This gives more room to design more different network architecture.\r\n\r\n** personally, i like this solution the best. But I not sure if the kaggle train data is enough for training from scratch.",
      "votes": null
    },
    {
      "id": "125848",
      "postDate": "07/03/2016 12:08:30",
      "content": "<p>Hi Heng,</p>\n\n<p>I think your ideas may help us a little.<br>\nHowever, we need more innovative approach to get LB 0.10.</p>",
      "rawMarkdown": "Hi Heng,\r\n\r\nI think your ideas may help us a little.<br>\r\nHowever, we need more innovative approach to get LB 0.10.",
      "votes": null
    },
    {
      "id": "125850",
      "postDate": "07/03/2016 12:22:04",
      "content": "<p>here is one idea to generate more train synthetic train samples.</p>",
      "rawMarkdown": "here is one idea to generate more train synthetic train samples.",
      "votes": null
    },
    {
      "id": "125856",
      "postDate": "07/03/2016 14:48:14",
      "content": "<p>Please follow this post for this idea:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125855#post125855\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125855#post125855</a></p>\n\n<p>[quote=Heng CherKeng;125850]</p>\n\n<p>here is one idea to generate more train synthetic train samples.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Please follow this post for this idea:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125855#post125855\r\n\r\n[quote=Heng CherKeng;125850]\r\n\r\nhere is one idea to generate more train synthetic train samples.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "125993",
      "postDate": "07/05/2016 10:59:04",
      "content": "<p>What about using Triplet loss?  Which seems to fit in this task, where some categories are not well classified. </p>",
      "rawMarkdown": "What about using Triplet loss?  Which seems to fit in this task, where some categories are not well classified.",
      "votes": null
    },
    {
      "id": "125994",
      "postDate": "07/05/2016 11:15:09",
      "content": "<p>@SecondPlan\nThank you for the suggestion. Metric learning layer will also be one of my option. The problem of triple loss is how to sample pairs of train samples.  I will report results if I implement it later. Thanks!</p>\n\n<p>** I am also looking at related metric loss  (eqn[1],[2]) from this paper:\n<a href=\"https://arxiv.org/pdf/1605.07270.pdf\">https://arxiv.org/pdf/1605.07270.pdf</a></p>",
      "rawMarkdown": "SecondPlan\r\nThank you for the suggestion. Metric learning layer will also be one of my option. The problem of triple loss is how to sample pairs of train samples.  I will report results if I implement it later. Thanks!\r\n\r\n** I am also looking at related metric loss  (eqn[1],[2]) from this paper:\r\nhttps://arxiv.org/pdf/1605.07270.pdf",
      "votes": null
    },
    {
      "id": "125997",
      "postDate": "07/05/2016 12:00:06",
      "content": "<p>there is torch code for TripletNet, <a href=\"https://github.com/eladhoffer/TripletNet\">https://github.com/eladhoffer/TripletNet</a>.  Do you tried other features, like Fisher vector?</p>",
      "rawMarkdown": "there is torch code for TripletNet, https://github.com/eladhoffer/TripletNet.  Do you tried other features, like Fisher vector?",
      "votes": null
    },
    {
      "id": "125999",
      "postDate": "07/05/2016 12:21:00",
      "content": "<p>@SecondPlan\nThank you for the link. For fisher vector, I am trying bilinear layer now.</p>\n\n<p>&quot;We propose bilinear models, a recognition architecture that consists of two feature extractors whose outputs are multiplied using outer product at each location of the image and pooled to obtain an image descriptor. This architecture can model local pairwise feature interactions in a translationally invariant manner which is particularly useful for fine-grained categorization. It also generalizes various orderless texture descriptors such as the Fisher vector, VLAD and O2P&quot;</p>\n\n<p><a href=\"http://arxiv.org/abs/1504.07889\">http://arxiv.org/abs/1504.07889</a>\nBilinear CNN Models for Fine-grained Visual Recognition</p>\n\n<p>see also\n<a href=\"https://github.com/gy20073/compact_bilinear_pooling/tree/master/caffe-20160312\">https://github.com/gy20073/compact_bilinear_pooling/tree/master/caffe-20160312</a></p>\n\n<p>[quote=SecondPlan;125997]</p>\n\n<p>there is torch code for TripletNet, <a href=\"https://github.com/eladhoffer/TripletNet\">https://github.com/eladhoffer/TripletNet</a>.  Do you tried other features, like Fisher vector?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "SecondPlan\r\nThank you for the link. For fisher vector, I am trying bilinear layer now.\r\n\r\n\"We propose bilinear models, a recognition architecture that consists of two feature extractors whose outputs are multiplied using outer product at each location of the image and pooled to obtain an image descriptor. This architecture can model local pairwise feature interactions in a translationally invariant manner which is particularly useful for fine-grained categorization. It also generalizes various orderless texture descriptors such as the Fisher vector, VLAD and O2P\"\r\n\r\nhttp://arxiv.org/abs/1504.07889\r\nBilinear CNN Models for Fine-grained Visual Recognition\r\n\r\nsee also\r\nhttps://github.com/gy20073/compact_bilinear_pooling/tree/master/caffe-20160312\r\n\r\n[quote=SecondPlan;125997]\r\n\r\nthere is torch code for TripletNet, https://github.com/eladhoffer/TripletNet.  Do you tried other features, like Fisher vector?\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "126018",
      "postDate": "07/05/2016 15:24:35",
      "content": "<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>",
      "rawMarkdown": "anyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.",
      "votes": null
    },
    {
      "id": "126041",
      "postDate": "07/05/2016 19:14:43",
      "content": "<p>[quote=Heng CherKeng;126018]</p>\n\n<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>\n\n<p>[/quote]</p>\n\n<p>Hey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. </p>",
      "rawMarkdown": "[quote=Heng CherKeng;126018]\r\n\r\nanyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.\r\n\r\n[/quote]\r\n\r\nHey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning.",
      "votes": null
    },
    {
      "id": "126048",
      "postDate": "07/05/2016 20:16:29",
      "content": "<p>@SecondPlan</p>\n\n<p>are you familiar with YOLO framework?\nyolo is being used to detect objects in image.\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\n(It also estimates the correct bounding box size if the object is present).</p>\n\n<p>In our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.</p>\n\n<p>SSD is multi scale yolo.</p>\n\n<ul>\n<li><a href=\"http://pjreddie.com/darknet/\">http://pjreddie.com/darknet/</a></li>\n<li><a href=\"https://github.com/xingwangsfu/caffe-yolo\">https://github.com/xingwangsfu/caffe-yolo</a></li>\n<li><a href=\"https://github.com/weiliu89/caffe/tree/ssd\">https://github.com/weiliu89/caffe/tree/ssd</a></li>\n</ul>\n\n<p>[quote=SecondPlan;126041]</p>\n\n<p>[quote=Heng CherKeng;126018]</p>\n\n<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>\n\n<p>[/quote]</p>\n\n<p>Hey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "SecondPlan\r\n\r\nare you familiar with YOLO framework?\r\nyolo is being used to detect objects in image.\r\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\r\n(It also estimates the correct bounding box size if the object is present).\r\n\r\n\r\nIn our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.\r\n\r\n\r\nSSD is multi scale yolo.\r\n\r\n - http://pjreddie.com/darknet/\r\n - https://github.com/xingwangsfu/caffe-yolo\r\n - https://github.com/weiliu89/caffe/tree/ssd\r\n \r\n\r\n[quote=SecondPlan;126041]\r\n\r\n[quote=Heng CherKeng;126018]\r\n\r\nanyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.\r\n\r\n[/quote]\r\n\r\nHey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. \r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "126105",
      "postDate": "07/06/2016 08:48:32",
      "content": "<p>[quote=Heng CherKeng;126048]</p>\n\n<p>@SecondPlan</p>\n\n<p>are you familiar with YOLO framework?\nyolo is being used to detect objects in image.\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\n(It also estimates the correct bounding box size if the object is present).</p>\n\n<p>In our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.</p>\n\n<p>SSD is multi scale yolo.</p>\n\n<ul>\n<li><a href=\"http://pjreddie.com/darknet/\">http://pjreddie.com/darknet/</a></li>\n<li><a href=\"https://github.com/xingwangsfu/caffe-yolo\">https://github.com/xingwangsfu/caffe-yolo</a></li>\n<li><a href=\"https://github.com/weiliu89/caffe/tree/ssd\">https://github.com/weiliu89/caffe/tree/ssd</a></li>\n</ul>\n\n<p>[quote=SecondPlan;126041]</p>\n\n<p>[quote=Heng CherKeng;126018]</p>\n\n<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>\n\n<p>[/quote]</p>\n\n<p>Hey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]\nThanks for the info, but for our task, if we identify a face, we need to know the direction of the face facing to, which is valuable information that yolo miss. I will try more naive approach, like the hierarchical classification</p>",
      "rawMarkdown": "[quote=Heng CherKeng;126048]\r\n\r\n@SecondPlan\r\n\r\nare you familiar with YOLO framework?\r\nyolo is being used to detect objects in image.\r\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\r\n(It also estimates the correct bounding box size if the object is present).\r\n\r\n\r\nIn our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.\r\n\r\n\r\nSSD is multi scale yolo.\r\n\r\n - http://pjreddie.com/darknet/\r\n - https://github.com/xingwangsfu/caffe-yolo\r\n - https://github.com/weiliu89/caffe/tree/ssd\r\n \r\n\r\n[quote=SecondPlan;126041]\r\n\r\n[quote=Heng CherKeng;126018]\r\n\r\nanyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.\r\n\r\n[/quote]\r\n\r\nHey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. \r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\nThanks for the info, but for our task, if we identify a face, we need to know the direction of the face facing to, which is valuable information that yolo miss. I will try more naive approach, like the hierarchical classification",
      "votes": null
    },
    {
      "id": "126532",
      "postDate": "07/09/2016 10:00:47",
      "content": "<p>A related paper in cvpr 2016\nMultiple Scale Faster-RCNN Approach to\nDriver&#8217;s Cell-phone Usage and Hands on Steering Wheel Detection</p>\n\n<p><a href=\"http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\">http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf</a></p>",
      "rawMarkdown": "A related paper in cvpr 2016\r\nMultiple Scale Faster-RCNN Approach to\r\nDriver’s Cell-phone Usage and Hands on Steering Wheel Detection\r\n\r\nhttp://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf",
      "votes": null
    },
    {
      "id": "126572",
      "postDate": "07/09/2016 23:03:04",
      "content": "<p>[quote=Heng CherKeng;126532]</p>\n\n<p>A related paper in cvpr 2016\nMultiple Scale Faster-RCNN Approach to\nDriver&#8217;s Cell-phone Usage and Hands on Steering Wheel Detection</p>\n\n<p><a href=\"http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\">http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf</a></p>\n\n<p>[/quote]</p>\n\n<p>what about this one? <a href=\"http://arxiv.org/pdf/1512.08086v1.pdf\">http://arxiv.org/pdf/1512.08086v1.pdf</a></p>",
      "rawMarkdown": "[quote=Heng CherKeng;126532]\r\n\r\nA related paper in cvpr 2016\r\nMultiple Scale Faster-RCNN Approach to\r\nDriver’s Cell-phone Usage and Hands on Steering Wheel Detection\r\n\r\nhttp://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\r\n\r\n[/quote]\r\n\r\nwhat about this one? http://arxiv.org/pdf/1512.08086v1.pdf",
      "votes": null
    },
    {
      "id": "126573",
      "postDate": "07/09/2016 23:03:23",
      "content": "<p>[quote=Heng CherKeng;126532]</p>\n\n<p>A related paper in cvpr 2016\nMultiple Scale Faster-RCNN Approach to\nDriver&#8217;s Cell-phone Usage and Hands on Steering Wheel Detection</p>\n\n<p><a href=\"http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\">http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf</a></p>\n\n<p>[/quote]</p>\n\n<p>what about this one? <a href=\"http://arxiv.org/pdf/1512.08086v1.pdf\">http://arxiv.org/pdf/1512.08086v1.pdf</a></p>",
      "rawMarkdown": "[quote=Heng CherKeng;126532]\r\n\r\nA related paper in cvpr 2016\r\nMultiple Scale Faster-RCNN Approach to\r\nDriver’s Cell-phone Usage and Hands on Steering Wheel Detection\r\n\r\nhttp://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\r\n\r\n[/quote]\r\n\r\nwhat about this one? http://arxiv.org/pdf/1512.08086v1.pdf",
      "votes": null
    },
    {
      "id": "126580",
      "postDate": "07/10/2016 05:12:56",
      "content": "<p>Summary of part of CVPR 2016 Fine-Grained Visual Categorization papers: <a href=\"http://mp.weixin.qq.com/s?plg_nld=1&plg_uin=1&mid=2650325020&idx=1&plg_nld=1&scene=23&plg_auth=1&__biz=MzI1NTE4NTUwOQ%3D%3D&plg_dev=1&srcid=0708FIXOyiOw9Wf3e79bsnT5&plg_usr=1&plg_vkey=1&sn=d0eb308a95f4dfc90de10ca08e3f9e47#rd&appinstall=1\">http://mp.weixin.qq.com/s?plg_nld=1&amp;plg_uin=1&amp;mid=2650325020&amp;idx=1&amp;plg_nld=1&amp;scene=23&amp;plg_auth=1&amp;__biz=MzI1NTE4NTUwOQ%3D%3D&amp;plg_dev=1&amp;srcid=0708FIXOyiOw9Wf3e79bsnT5&amp;plg_usr=1&amp;plg_vkey=1&amp;sn=d0eb308a95f4dfc90de10ca08e3f9e47#rd&amp;appinstall=1</a></p>",
      "rawMarkdown": "Summary of part of CVPR 2016 Fine-Grained Visual Categorization papers: http://mp.weixin.qq.com/s?plg_nld=1&plg_uin=1&mid=2650325020&idx=1&plg_nld=1&scene=23&plg_auth=1&__biz=MzI1NTE4NTUwOQ%3D%3D&plg_dev=1&srcid=0708FIXOyiOw9Wf3e79bsnT5&plg_usr=1&plg_vkey=1&sn=d0eb308a95f4dfc90de10ca08e3f9e47#rd&appinstall=1",
      "votes": null
    },
    {
      "id": "126581",
      "postDate": "07/10/2016 05:27:41",
      "content": "<p>Just a thought, actually more of a question to the experts here: </p>\n\n<p>To account for the limited training data, can we create additional images (with new drivers in a similar setting) and append to the training set? The participants here could contribute to this public database.</p>\n\n<p>Even with a low resolution (say 64x64) of a large training set, a better performing model could probably be learned. </p>",
      "rawMarkdown": "Just a thought, actually more of a question to the experts here: \r\n\r\nTo account for the limited training data, can we create additional images (with new drivers in a similar setting) and append to the training set? The participants here could contribute to this public database.\r\n\r\nEven with a low resolution (say 64x64) of a large training set, a better performing model could probably be learned.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 125848,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "07/03/2016 12:08:30",
      "content": "<p>Hi Heng,</p>\n\n<p>I think your ideas may help us a little.<br>\nHowever, we need more innovative approach to get LB 0.10.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125850,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/03/2016 12:22:04",
      "content": "<p>here is one idea to generate more train synthetic train samples.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125856,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/03/2016 14:48:14",
      "content": "<p>Please follow this post for this idea:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125855#post125855\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125855#post125855</a></p>\n\n<p>[quote=Heng CherKeng;125850]</p>\n\n<p>here is one idea to generate more train synthetic train samples.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125993,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "07/05/2016 10:59:04",
      "content": "<p>What about using Triplet loss?  Which seems to fit in this task, where some categories are not well classified. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125994,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/05/2016 11:15:09",
      "content": "<p>@SecondPlan\nThank you for the suggestion. Metric learning layer will also be one of my option. The problem of triple loss is how to sample pairs of train samples.  I will report results if I implement it later. Thanks!</p>\n\n<p>** I am also looking at related metric loss  (eqn[1],[2]) from this paper:\n<a href=\"https://arxiv.org/pdf/1605.07270.pdf\">https://arxiv.org/pdf/1605.07270.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125997,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "07/05/2016 12:00:06",
      "content": "<p>there is torch code for TripletNet, <a href=\"https://github.com/eladhoffer/TripletNet\">https://github.com/eladhoffer/TripletNet</a>.  Do you tried other features, like Fisher vector?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125999,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/05/2016 12:21:00",
      "content": "<p>@SecondPlan\nThank you for the link. For fisher vector, I am trying bilinear layer now.</p>\n\n<p>&quot;We propose bilinear models, a recognition architecture that consists of two feature extractors whose outputs are multiplied using outer product at each location of the image and pooled to obtain an image descriptor. This architecture can model local pairwise feature interactions in a translationally invariant manner which is particularly useful for fine-grained categorization. It also generalizes various orderless texture descriptors such as the Fisher vector, VLAD and O2P&quot;</p>\n\n<p><a href=\"http://arxiv.org/abs/1504.07889\">http://arxiv.org/abs/1504.07889</a>\nBilinear CNN Models for Fine-grained Visual Recognition</p>\n\n<p>see also\n<a href=\"https://github.com/gy20073/compact_bilinear_pooling/tree/master/caffe-20160312\">https://github.com/gy20073/compact_bilinear_pooling/tree/master/caffe-20160312</a></p>\n\n<p>[quote=SecondPlan;125997]</p>\n\n<p>there is torch code for TripletNet, <a href=\"https://github.com/eladhoffer/TripletNet\">https://github.com/eladhoffer/TripletNet</a>.  Do you tried other features, like Fisher vector?</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126018,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/05/2016 15:24:35",
      "content": "<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126041,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "07/05/2016 19:14:43",
      "content": "<p>[quote=Heng CherKeng;126018]</p>\n\n<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>\n\n<p>[/quote]</p>\n\n<p>Hey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126048,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/05/2016 20:16:29",
      "content": "<p>@SecondPlan</p>\n\n<p>are you familiar with YOLO framework?\nyolo is being used to detect objects in image.\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\n(It also estimates the correct bounding box size if the object is present).</p>\n\n<p>In our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.</p>\n\n<p>SSD is multi scale yolo.</p>\n\n<ul>\n<li><a href=\"http://pjreddie.com/darknet/\">http://pjreddie.com/darknet/</a></li>\n<li><a href=\"https://github.com/xingwangsfu/caffe-yolo\">https://github.com/xingwangsfu/caffe-yolo</a></li>\n<li><a href=\"https://github.com/weiliu89/caffe/tree/ssd\">https://github.com/weiliu89/caffe/tree/ssd</a></li>\n</ul>\n\n<p>[quote=SecondPlan;126041]</p>\n\n<p>[quote=Heng CherKeng;126018]</p>\n\n<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>\n\n<p>[/quote]</p>\n\n<p>Hey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. </p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126105,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "07/06/2016 08:48:32",
      "content": "<p>[quote=Heng CherKeng;126048]</p>\n\n<p>@SecondPlan</p>\n\n<p>are you familiar with YOLO framework?\nyolo is being used to detect objects in image.\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\n(It also estimates the correct bounding box size if the object is present).</p>\n\n<p>In our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.</p>\n\n<p>SSD is multi scale yolo.</p>\n\n<ul>\n<li><a href=\"http://pjreddie.com/darknet/\">http://pjreddie.com/darknet/</a></li>\n<li><a href=\"https://github.com/xingwangsfu/caffe-yolo\">https://github.com/xingwangsfu/caffe-yolo</a></li>\n<li><a href=\"https://github.com/weiliu89/caffe/tree/ssd\">https://github.com/weiliu89/caffe/tree/ssd</a></li>\n</ul>\n\n<p>[quote=SecondPlan;126041]</p>\n\n<p>[quote=Heng CherKeng;126018]</p>\n\n<p>anyone try to to attack the problem using structured learning?\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.</p>\n\n<p>[/quote]</p>\n\n<p>Hey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]\nThanks for the info, but for our task, if we identify a face, we need to know the direction of the face facing to, which is valuable information that yolo miss. I will try more naive approach, like the hierarchical classification</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126532,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/09/2016 10:00:47",
      "content": "<p>A related paper in cvpr 2016\nMultiple Scale Faster-RCNN Approach to\nDriver&#8217;s Cell-phone Usage and Hands on Steering Wheel Detection</p>\n\n<p><a href=\"http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\">http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126572,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "07/09/2016 23:03:04",
      "content": "<p>[quote=Heng CherKeng;126532]</p>\n\n<p>A related paper in cvpr 2016\nMultiple Scale Faster-RCNN Approach to\nDriver&#8217;s Cell-phone Usage and Hands on Steering Wheel Detection</p>\n\n<p><a href=\"http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\">http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf</a></p>\n\n<p>[/quote]</p>\n\n<p>what about this one? <a href=\"http://arxiv.org/pdf/1512.08086v1.pdf\">http://arxiv.org/pdf/1512.08086v1.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126573,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "07/09/2016 23:03:23",
      "content": "<p>[quote=Heng CherKeng;126532]</p>\n\n<p>A related paper in cvpr 2016\nMultiple Scale Faster-RCNN Approach to\nDriver&#8217;s Cell-phone Usage and Hands on Steering Wheel Detection</p>\n\n<p><a href=\"http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\">http://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf</a></p>\n\n<p>[/quote]</p>\n\n<p>what about this one? <a href=\"http://arxiv.org/pdf/1512.08086v1.pdf\">http://arxiv.org/pdf/1512.08086v1.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126580,
      "author_name": "xuguozhi",
      "author_url": "",
      "post_date": "07/10/2016 05:12:56",
      "content": "<p>Summary of part of CVPR 2016 Fine-Grained Visual Categorization papers: <a href=\"http://mp.weixin.qq.com/s?plg_nld=1&plg_uin=1&mid=2650325020&idx=1&plg_nld=1&scene=23&plg_auth=1&__biz=MzI1NTE4NTUwOQ%3D%3D&plg_dev=1&srcid=0708FIXOyiOw9Wf3e79bsnT5&plg_usr=1&plg_vkey=1&sn=d0eb308a95f4dfc90de10ca08e3f9e47#rd&appinstall=1\">http://mp.weixin.qq.com/s?plg_nld=1&amp;plg_uin=1&amp;mid=2650325020&amp;idx=1&amp;plg_nld=1&amp;scene=23&amp;plg_auth=1&amp;__biz=MzI1NTE4NTUwOQ%3D%3D&amp;plg_dev=1&amp;srcid=0708FIXOyiOw9Wf3e79bsnT5&amp;plg_usr=1&amp;plg_vkey=1&amp;sn=d0eb308a95f4dfc90de10ca08e3f9e47#rd&amp;appinstall=1</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126581,
      "author_name": "opraveen",
      "author_url": "",
      "post_date": "07/10/2016 05:27:41",
      "content": "<p>Just a thought, actually more of a question to the experts here: </p>\n\n<p>To account for the limited training data, can we create additional images (with new drivers in a similar setting) and append to the training set? The participants here could contribute to this public database.</p>\n\n<p>Even with a low resolution (say 64x64) of a large training set, a better performing model could probably be learned. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "125823": "Time is running short. I have a few ideas to move from LB 0.20. I would be grateful if anyone can help to advise on which of the followings are  most likely to work:\r\n\r\n - [strategy.1] focus on tweaking the performance of my current solution (e.g ensemble of vgg16, vgg19, googlenet). Sweep hyperparamters like learning rate, batch size. Try relu, prelu, maxout, elu. Try max,ave,max+ave,fractional pool, etc Try heavy augmentation. \r\n\r\ne.g https://arxiv.org/abs/1606.02228\r\n\"Systematic evaluation of CNN advances on the ImageNet\", code: https://github.com/ducha-aiki/caffenet-benchmark\r\n\r\n**the problem is that it takes too much time and i estimate best performance to be LB 0.18\r\n\r\n\r\n\r\n - [strategy.2]  One problem is lack of data. Try to use traditional classifiers on deep features. E.g. boosting trees/ random forests on features of all layers of deep CNN network. One end-to-end solution is combine traditional + deep is \"deep neural forest\":\r\n http://www.cv-foundation.org/openaccess/content_iccv_2015/html/Kontschieder_Deep_Neural_Decision_ICCV_2015_paper.html\r\n\r\n**I worked with boosting tree+ vgg16 (https://arxiv.org/pdf/1504.07339.pdf, http://arxiv.org/pdf/1603.00124.pdf) and know it outperforms deep learning in 2-class caltech pedestrian detection which has limited training data. But i am not sure if it works for 10-class image classification here.\r\n \r\n - [strategy.3]  restructure the problem. e.g. use hierarchical classification. Group the confusing class (e.g. c0-norm and c9-talk) to one class first, then refine it.e.g. HD-CNN:\r\nhttps://arxiv.org/abs/1410.0736\r\n\r\n\r\n - [strategy.4]  Change to large margin loss. Since the evaluation criteria is logloss and not accuracy, this should help. Large margin softmax seems promising : http://jmlr.org/proceedings/papers/v48/liud16.pdf\r\n\r\n - [strategy.5]  Believe that resNet should work, afterall it's is the imagenet winner 2015. Continue to change the resNet32, resNet50 architecture to find one that fits the current kaggle driver data. There are many modified  resNet like wide resNet, pre-activation resNet, etc\r\n\r\n** I already make several attempts with resNet but couldn't get it work as well as vgg16.\r\n\r\n - [strategy.6]  Find the best \"fine-gradient categorization\" paper and implement. Bilinear model seems promising: http://arxiv.org/abs/1504.07889. These fine grain action recognition papers seems promising too: https://sites.google.com/site/amirrosenfeld/\r\n(Hand-Object Interaction and Precise Localization in Transitive Action Recognition)\r\n(Visual Concept Recognition and Localization via Iterative Introspection)\r\n\r\n\r\n** my concern is that these may not work well if fundamental issue is lack of data\r\n\r\n - [strategy.7]  Build and train a network from scratch. Forget about pretrained model. This gives more room to design more different network architecture.\r\n\r\n** personally, i like this solution the best. But I not sure if the kaggle train data is enough for training from scratch.",
    "125848": "Hi Heng,\r\n\r\nI think your ideas may help us a little.<br>\r\nHowever, we need more innovative approach to get LB 0.10.",
    "125850": "here is one idea to generate more train synthetic train samples.",
    "125856": "Please follow this post for this idea:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125855#post125855\r\n\r\n[quote=Heng CherKeng;125850]\r\n\r\nhere is one idea to generate more train synthetic train samples.\r\n\r\n[/quote]",
    "125993": "What about using Triplet loss?  Which seems to fit in this task, where some categories are not well classified.",
    "125994": "SecondPlan\r\nThank you for the suggestion. Metric learning layer will also be one of my option. The problem of triple loss is how to sample pairs of train samples.  I will report results if I implement it later. Thanks!\r\n\r\n** I am also looking at related metric loss  (eqn[1],[2]) from this paper:\r\nhttps://arxiv.org/pdf/1605.07270.pdf",
    "125997": "there is torch code for TripletNet, https://github.com/eladhoffer/TripletNet.  Do you tried other features, like Fisher vector?",
    "125999": "SecondPlan\r\nThank you for the link. For fisher vector, I am trying bilinear layer now.\r\n\r\n\"We propose bilinear models, a recognition architecture that consists of two feature extractors whose outputs are multiplied using outer product at each location of the image and pooled to obtain an image descriptor. This architecture can model local pairwise feature interactions in a translationally invariant manner which is particularly useful for fine-grained categorization. It also generalizes various orderless texture descriptors such as the Fisher vector, VLAD and O2P\"\r\n\r\nhttp://arxiv.org/abs/1504.07889\r\nBilinear CNN Models for Fine-grained Visual Recognition\r\n\r\nsee also\r\nhttps://github.com/gy20073/compact_bilinear_pooling/tree/master/caffe-20160312\r\n\r\n[quote=SecondPlan;125997]\r\n\r\nthere is torch code for TripletNet, https://github.com/eladhoffer/TripletNet.  Do you tried other features, like Fisher vector?\r\n\r\n[/quote]",
    "126018": "anyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.",
    "126041": "[quote=Heng CherKeng;126018]\r\n\r\nanyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.\r\n\r\n[/quote]\r\n\r\nHey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning.",
    "126048": "SecondPlan\r\n\r\nare you familiar with YOLO framework?\r\nyolo is being used to detect objects in image.\r\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\r\n(It also estimates the correct bounding box size if the object is present).\r\n\r\n\r\nIn our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.\r\n\r\n\r\nSSD is multi scale yolo.\r\n\r\n - http://pjreddie.com/darknet/\r\n - https://github.com/xingwangsfu/caffe-yolo\r\n - https://github.com/weiliu89/caffe/tree/ssd\r\n \r\n\r\n[quote=SecondPlan;126041]\r\n\r\n[quote=Heng CherKeng;126018]\r\n\r\nanyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.\r\n\r\n[/quote]\r\n\r\nHey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. \r\n\r\n[/quote]",
    "126105": "[quote=Heng CherKeng;126048]\r\n\r\n@SecondPlan\r\n\r\nare you familiar with YOLO framework?\r\nyolo is being used to detect objects in image.\r\nit divides the image into grid and classifies if each grid contain an object of certain class or not.\r\n(It also estimates the correct bounding box size if the object is present).\r\n\r\n\r\nIn our case, we can use the same training and testing code. we divide the image into grids and classify if each grid is part of the object of certain class or background instead. we ignore the box estimation.\r\n\r\n\r\nSSD is multi scale yolo.\r\n\r\n - http://pjreddie.com/darknet/\r\n - https://github.com/xingwangsfu/caffe-yolo\r\n - https://github.com/weiliu89/caffe/tree/ssd\r\n \r\n\r\n[quote=SecondPlan;126041]\r\n\r\n[quote=Heng CherKeng;126018]\r\n\r\nanyone try to to attack the problem using structured learning?\r\ne.g. see attachment picture.: the target ground truth to be learned is a grid of labels per image, instead of a single label per image.\r\n\r\n[/quote]\r\n\r\nHey, Heng CherKeng. That's something I was thinking. We humans do the finegrained classification by seeing the locations where they different. These kind of outputs shown in the picture served as categorical variables that  can be feed into a decision tree or random forest to find a good decision boundary. But I am not familiar with structure learning. \r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\nThanks for the info, but for our task, if we identify a face, we need to know the direction of the face facing to, which is valuable information that yolo miss. I will try more naive approach, like the hierarchical classification",
    "126532": "A related paper in cvpr 2016\r\nMultiple Scale Faster-RCNN Approach to\r\nDriver’s Cell-phone Usage and Hands on Steering Wheel Detection\r\n\r\nhttp://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf",
    "126572": "[quote=Heng CherKeng;126532]\r\n\r\nA related paper in cvpr 2016\r\nMultiple Scale Faster-RCNN Approach to\r\nDriver’s Cell-phone Usage and Hands on Steering Wheel Detection\r\n\r\nhttp://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\r\n\r\n[/quote]\r\n\r\nwhat about this one? http://arxiv.org/pdf/1512.08086v1.pdf",
    "126573": "[quote=Heng CherKeng;126532]\r\n\r\nA related paper in cvpr 2016\r\nMultiple Scale Faster-RCNN Approach to\r\nDriver’s Cell-phone Usage and Hands on Steering Wheel Detection\r\n\r\nhttp://www.cv-foundation.org//openaccess/content_cvpr_2016_workshops/w3/papers/Le_Multiple_Scale_Faster-RCNN_CVPR_2016_paper.pdf\r\n\r\n[/quote]\r\n\r\nwhat about this one? http://arxiv.org/pdf/1512.08086v1.pdf",
    "126580": "Summary of part of CVPR 2016 Fine-Grained Visual Categorization papers: http://mp.weixin.qq.com/s?plg_nld=1&plg_uin=1&mid=2650325020&idx=1&plg_nld=1&scene=23&plg_auth=1&__biz=MzI1NTE4NTUwOQ%3D%3D&plg_dev=1&srcid=0708FIXOyiOw9Wf3e79bsnT5&plg_usr=1&plg_vkey=1&sn=d0eb308a95f4dfc90de10ca08e3f9e47#rd&appinstall=1",
    "126581": "Just a thought, actually more of a question to the experts here: \r\n\r\nTo account for the limited training data, can we create additional images (with new drivers in a similar setting) and append to the training set? The participants here could contribute to this public database.\r\n\r\nEven with a low resolution (say 64x64) of a large training set, a better performing model could probably be learned."
  },
  "source": "meta"
}