{
  "id": 22906,
  "title": "A brief summary ",
  "url": "/competitions/state-farm-distracted-driver-detection/writeups/jacobkie-a-brief-summary",
  "author_name": "",
  "post_date": "2016-08-17T12:32:09.063Z",
  "votes": 57,
  "comment_count": 11,
  "views": 2229,
  "content": "<p>Firstly, I would like to thank Kaggle and State Farm to host such an interesting competition. I would also like to thank ZF Turbo, Dong Jiao for their starting code and many others who have kindly share their valuable ideas and experiences on the forum. I have learned a lot during this competition.</p>\n\n<p>Also, sorry for the late summary, I was quite occupied in the past two weeks.</p>\n\n<p>I understand that data set in this competition displays two unique characteristics: </p>\n\n<ol>\n<li>The test set is much bigger than the train set, so it is easy to get overfit.</li>\n<li>The images are correlated as they are produced as frames from a video clip.  Images taken from a single video segment are put into the same category, although if judged individually, a few of them might be labelled differently.</li>\n</ol>\n\n<p>My solution focuses mainly on these two characteristics.</p>\n\n<ul>\n<li><strong>Models</strong></li>\n</ul>\n\n<p>I used two models in this competition: pre-trained Vgg16, and Vgg16_3, a modified version of Vgg16 I devised to deal with overfit.</p>\n\n<p>The best LB score I got with a single Vgg16 model is around 0.3 as a result of overfit. It seemed to me that the model memorized some other features rather than the desired ones. To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver&#8217;s right hand cross is highly correlated with c5.</p>\n\n<p>The original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3. </p>\n\n<p>The original vgg16 network except the last output layer are shared cross the three inputs. The three outputs of the second-to-last layer (the second full layer), which correspond to the three inputs, are connected with three separate final output layers. </p>\n\n<p>I used Vgg16_3 to fine-tune saved model files from Vgg16. It improves the single model score by around 0.05.    </p>\n\n<ul>\n<li><strong>K_Nearest_Neighbor Average</strong></li>\n</ul>\n\n<p>I used output of the last Maxpool layer (pool5) of Vgg16 with pertained weights to map each test image to a 512*7*7-dimensional coordinate, and use distances in this space (pool5-feature space) to define similarity among test images. A weighted average of the predictions of each image, together with their 10 nearest neighbors in pool5-feature space, would generally improve a single model score by 0.10-0.12. </p>\n\n<ul>\n<li><strong>Ensemble average</strong></li>\n</ul>\n\n<p>Ensemble average was applied for each category separately. <br>\nFrom confusion matrices of the CV set, I noticed that some trained models did a decent job at distinguishing a certain category from the rest, while quite confused among the rest categories. So for each category, models among top-ten percent of cross-entropy loss associated with this specific category are chosen and used for calculating ensemble average. This category-specific average outperforms significantly compared with simple arithmetic/geometric average according to my experience.</p>\n\n<ul>\n<li><strong>Segment average</strong></li>\n</ul>\n\n<p>The last step, segment average, the same as K-nearest-neighbor-average, focuses on the second characteristics of the competition as I mentioned above. I divide the test images into small groups according to their similarities in the 512*7*7-dimensional pool5-feature space. If images in one group display consistent and also confident predictions, I receive this as a strong positive signal and renormalize all the images in that group to share the same predictions. </p>\n\n<p>As I read in the forum, some of you have achieved a single-model LB score below 0.15, which is admirable. I believe you can reduce that by around 40% by using K_Nearest_Neighbor Average alone, single model. I hope all of you had enjoyed this competition as much as I did, thanks again. </p>",
  "messages": [
    {
      "id": "131467",
      "postDate": "08/17/2016 12:32:09",
      "content": "<p>Firstly, I would like to thank Kaggle and State Farm to host such an interesting competition. I would also like to thank ZF Turbo, Dong Jiao for their starting code and many others who have kindly share their valuable ideas and experiences on the forum. I have learned a lot during this competition.</p>\n\n<p>Also, sorry for the late summary, I was quite occupied in the past two weeks.</p>\n\n<p>I understand that data set in this competition displays two unique characteristics: </p>\n\n<ol>\n<li>The test set is much bigger than the train set, so it is easy to get overfit.</li>\n<li>The images are correlated as they are produced as frames from a video clip.  Images taken from a single video segment are put into the same category, although if judged individually, a few of them might be labelled differently.</li>\n</ol>\n\n<p>My solution focuses mainly on these two characteristics.</p>\n\n<ul>\n<li><strong>Models</strong></li>\n</ul>\n\n<p>I used two models in this competition: pre-trained Vgg16, and Vgg16_3, a modified version of Vgg16 I devised to deal with overfit.</p>\n\n<p>The best LB score I got with a single Vgg16 model is around 0.3 as a result of overfit. It seemed to me that the model memorized some other features rather than the desired ones. To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver&#8217;s right hand cross is highly correlated with c5.</p>\n\n<p>The original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3. </p>\n\n<p>The original vgg16 network except the last output layer are shared cross the three inputs. The three outputs of the second-to-last layer (the second full layer), which correspond to the three inputs, are connected with three separate final output layers. </p>\n\n<p>I used Vgg16_3 to fine-tune saved model files from Vgg16. It improves the single model score by around 0.05.    </p>\n\n<ul>\n<li><strong>K_Nearest_Neighbor Average</strong></li>\n</ul>\n\n<p>I used output of the last Maxpool layer (pool5) of Vgg16 with pertained weights to map each test image to a 512*7*7-dimensional coordinate, and use distances in this space (pool5-feature space) to define similarity among test images. A weighted average of the predictions of each image, together with their 10 nearest neighbors in pool5-feature space, would generally improve a single model score by 0.10-0.12. </p>\n\n<ul>\n<li><strong>Ensemble average</strong></li>\n</ul>\n\n<p>Ensemble average was applied for each category separately. <br>\nFrom confusion matrices of the CV set, I noticed that some trained models did a decent job at distinguishing a certain category from the rest, while quite confused among the rest categories. So for each category, models among top-ten percent of cross-entropy loss associated with this specific category are chosen and used for calculating ensemble average. This category-specific average outperforms significantly compared with simple arithmetic/geometric average according to my experience.</p>\n\n<ul>\n<li><strong>Segment average</strong></li>\n</ul>\n\n<p>The last step, segment average, the same as K-nearest-neighbor-average, focuses on the second characteristics of the competition as I mentioned above. I divide the test images into small groups according to their similarities in the 512*7*7-dimensional pool5-feature space. If images in one group display consistent and also confident predictions, I receive this as a strong positive signal and renormalize all the images in that group to share the same predictions. </p>\n\n<p>As I read in the forum, some of you have achieved a single-model LB score below 0.15, which is admirable. I believe you can reduce that by around 40% by using K_Nearest_Neighbor Average alone, single model. I hope all of you had enjoyed this competition as much as I did, thanks again. </p>",
      "rawMarkdown": "Firstly, I would like to thank Kaggle and State Farm to host such an interesting competition. I would also like to thank ZF Turbo, Dong Jiao for their starting code and many others who have kindly share their valuable ideas and experiences on the forum. I have learned a lot during this competition.\r\n\r\nAlso, sorry for the late summary, I was quite occupied in the past two weeks.\r\n\t\r\nI understand that data set in this competition displays two unique characteristics: \r\n\r\n 1. The test set is much bigger than the train set, so it is easy to get overfit.\r\n 2. The images are correlated as they are produced as frames from a video clip.  Images taken from a single video segment are put into the same category, although if judged individually, a few of them might be labelled differently.\r\n\r\nMy solution focuses mainly on these two characteristics.\r\n\t\r\n\r\n - **Models**\r\n\r\nI used two models in this competition: pre-trained Vgg16, and Vgg16_3, a modified version of Vgg16 I devised to deal with overfit.\r\n\r\nThe best LB score I got with a single Vgg16 model is around 0.3 as a result of overfit. It seemed to me that the model memorized some other features rather than the desired ones. To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver’s right hand cross is highly correlated with c5.\r\n\r\nThe original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3. \r\n\r\nThe original vgg16 network except the last output layer are shared cross the three inputs. The three outputs of the second-to-last layer (the second full layer), which correspond to the three inputs, are connected with three separate final output layers. \r\n\r\nI used Vgg16_3 to fine-tune saved model files from Vgg16. It improves the single model score by around 0.05.\t\r\n\r\n - **K_Nearest_Neighbor Average**\r\n\r\nI used output of the last Maxpool layer (pool5) of Vgg16 with pertained weights to map each test image to a 512*7*7-dimensional coordinate, and use distances in this space (pool5-feature space) to define similarity among test images. A weighted average of the predictions of each image, together with their 10 nearest neighbors in pool5-feature space, would generally improve a single model score by 0.10-0.12. \r\n \r\n\r\n - **Ensemble average**\r\n\r\nEnsemble average was applied for each category separately.   \r\nFrom confusion matrices of the CV set, I noticed that some trained models did a decent job at distinguishing a certain category from the rest, while quite confused among the rest categories. So for each category, models among top-ten percent of cross-entropy loss associated with this specific category are chosen and used for calculating ensemble average. This category-specific average outperforms significantly compared with simple arithmetic/geometric average according to my experience.\r\n\r\n - **Segment average**\r\n\r\nThe last step, segment average, the same as K-nearest-neighbor-average, focuses on the second characteristics of the competition as I mentioned above. I divide the test images into small groups according to their similarities in the 512*7*7-dimensional pool5-feature space. If images in one group display consistent and also confident predictions, I receive this as a strong positive signal and renormalize all the images in that group to share the same predictions. \r\n\r\nAs I read in the forum, some of you have achieved a single-model LB score below 0.15, which is admirable. I believe you can reduce that by around 40% by using K_Nearest_Neighbor Average alone, single model. I hope all of you had enjoyed this competition as much as I did, thanks again.",
      "votes": null
    },
    {
      "id": "131692",
      "postDate": "08/19/2016 04:09:36",
      "content": "<p>thanks for sharing and congratulations for the awesome result!</p>",
      "rawMarkdown": "thanks for sharing and congratulations for the awesome result!",
      "votes": null
    },
    {
      "id": "134678",
      "postDate": "09/07/2016 21:23:32",
      "content": "<p>what is the distance are u using in the 512x7x7 space?</p>",
      "rawMarkdown": "what is the distance are u using in the 512x7x7 space?",
      "votes": null
    },
    {
      "id": "135025",
      "postDate": "09/10/2016 09:51:04",
      "content": "<p>@Ellen Gao Jian, Thanks\n@SecondPlan, Euclidian distance</p>",
      "rawMarkdown": "Ellen Gao Jian, Thanks\r\n@SecondPlan, Euclidian distance",
      "votes": null
    },
    {
      "id": "135026",
      "postDate": "09/10/2016 09:54:25",
      "content": "<p>@Ellen Gao Jian, Thanks\n@SecondPlan, Euclidian distance</p>",
      "rawMarkdown": "Ellen Gao Jian, Thanks\r\n@SecondPlan, Euclidian distance",
      "votes": null
    },
    {
      "id": "219170",
      "postDate": "09/07/2017 08:04:10",
      "content": "<p>Great. Thanks for sharing this. Very much appreciated.</p>",
      "rawMarkdown": "Great. Thanks for sharing this. Very much appreciated.",
      "votes": null
    },
    {
      "id": "256113",
      "postDate": "12/11/2017 07:11:46",
      "content": "<p>can you share your demo ?</p>",
      "rawMarkdown": "can you share your demo ?",
      "votes": null
    },
    {
      "id": "270008",
      "postDate": "01/17/2018 16:41:42",
      "content": "<p>Thanks for the write-up.  Super helpful!</p>",
      "rawMarkdown": "Thanks for the write-up.  Super helpful!",
      "votes": null
    },
    {
      "id": "366983",
      "postDate": "08/06/2018 20:56:41",
      "content": "<p>Hi <a href=\"/jacobkie\">@jacobkie</a>\nCan you share your codes for playing around ? how I can contact you ? can you send me your contact details through messages here ? </p>\n\n<p>thx </p>",
      "rawMarkdown": "Hi @jacobkie\nCan you share your codes for playing around ? how I can contact you ? can you send me your contact details through messages here ? \n\nthx",
      "votes": null
    },
    {
      "id": "723004",
      "postDate": "01/19/2020 11:41:54",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1012370",
      "postDate": "09/16/2020 03:51:03",
      "content": "<p>hi thank you for sharing! i'm doing final year project in the university and interested in this topic. do you mind sending e your code so that i can learn from it? thank you very much!!!</p>",
      "rawMarkdown": "hi thank you for sharing! i'm doing final year project in the university and interested in this topic. do you mind sending e your code so that i can learn from it? thank you very much!!!",
      "votes": null
    },
    {
      "id": "3430964",
      "postDate": "03/29/2026 05:57:24",
      "content": "<p>hlo bro can you send me the dataset +code so that i can use this in my project included this as my first step so if you can send please</p>",
      "rawMarkdown": "hlo bro can you send me the dataset +code so that i can use this in my project included this as my first step so if you can send please",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012370,
      "author_name": "txing001",
      "author_url": "",
      "post_date": "09/16/2020 03:51:03",
      "content": "<p>hi thank you for sharing! i'm doing final year project in the university and interested in this topic. do you mind sending e your code so that i can learn from it? thank you very much!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3430964,
      "author_name": "shubham45675",
      "author_url": "",
      "post_date": "03/29/2026 05:57:24",
      "content": "<p>hlo bro can you send me the dataset +code so that i can use this in my project included this as my first step so if you can send please</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 131692,
      "author_name": "ellengaojian",
      "author_url": "",
      "post_date": "08/19/2016 04:09:36",
      "content": "<p>thanks for sharing and congratulations for the awesome result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 134678,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "09/07/2016 21:23:32",
      "content": "<p>what is the distance are u using in the 512x7x7 space?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 135025,
      "author_name": "jacobkie",
      "author_url": "",
      "post_date": "09/10/2016 09:51:04",
      "content": "<p>@Ellen Gao Jian, Thanks\n@SecondPlan, Euclidian distance</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 135026,
      "author_name": "jacobkie",
      "author_url": "",
      "post_date": "09/10/2016 09:54:25",
      "content": "<p>@Ellen Gao Jian, Thanks\n@SecondPlan, Euclidian distance</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 219170,
      "author_name": "chmaxx",
      "author_url": "",
      "post_date": "09/07/2017 08:04:10",
      "content": "<p>Great. Thanks for sharing this. Very much appreciated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 256113,
      "author_name": "penmily",
      "author_url": "",
      "post_date": "12/11/2017 07:11:46",
      "content": "<p>can you share your demo ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270008,
      "author_name": "joshuacheatham",
      "author_url": "",
      "post_date": "01/17/2018 16:41:42",
      "content": "<p>Thanks for the write-up.  Super helpful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 366983,
      "author_name": "hannasolo",
      "author_url": "",
      "post_date": "08/06/2018 20:56:41",
      "content": "<p>Hi <a href=\"/jacobkie\">@jacobkie</a>\nCan you share your codes for playing around ? how I can contact you ? can you send me your contact details through messages here ? </p>\n\n<p>thx </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 723004,
      "author_name": "donggyunhong",
      "author_url": "",
      "post_date": "01/19/2020 11:41:54",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "131467": "Firstly, I would like to thank Kaggle and State Farm to host such an interesting competition. I would also like to thank ZF Turbo, Dong Jiao for their starting code and many others who have kindly share their valuable ideas and experiences on the forum. I have learned a lot during this competition.\r\n\r\nAlso, sorry for the late summary, I was quite occupied in the past two weeks.\r\n\t\r\nI understand that data set in this competition displays two unique characteristics: \r\n\r\n 1. The test set is much bigger than the train set, so it is easy to get overfit.\r\n 2. The images are correlated as they are produced as frames from a video clip.  Images taken from a single video segment are put into the same category, although if judged individually, a few of them might be labelled differently.\r\n\r\nMy solution focuses mainly on these two characteristics.\r\n\t\r\n\r\n - **Models**\r\n\r\nI used two models in this competition: pre-trained Vgg16, and Vgg16_3, a modified version of Vgg16 I devised to deal with overfit.\r\n\r\nThe best LB score I got with a single Vgg16 model is around 0.3 as a result of overfit. It seemed to me that the model memorized some other features rather than the desired ones. To help the model to concentrate on more essential features, I manually selected two regions of interests for each train image. One is the head region, which I believe is the most informative part of the whole image. And the other is the bottom-right quarter where the appearance of a driver’s right hand cross is highly correlated with c5.\r\n\r\nThe original train image, together with the two selected regions of interests, are the three inputs of Vgg16_3. \r\n\r\nThe original vgg16 network except the last output layer are shared cross the three inputs. The three outputs of the second-to-last layer (the second full layer), which correspond to the three inputs, are connected with three separate final output layers. \r\n\r\nI used Vgg16_3 to fine-tune saved model files from Vgg16. It improves the single model score by around 0.05.\t\r\n\r\n - **K_Nearest_Neighbor Average**\r\n\r\nI used output of the last Maxpool layer (pool5) of Vgg16 with pertained weights to map each test image to a 512*7*7-dimensional coordinate, and use distances in this space (pool5-feature space) to define similarity among test images. A weighted average of the predictions of each image, together with their 10 nearest neighbors in pool5-feature space, would generally improve a single model score by 0.10-0.12. \r\n \r\n\r\n - **Ensemble average**\r\n\r\nEnsemble average was applied for each category separately.   \r\nFrom confusion matrices of the CV set, I noticed that some trained models did a decent job at distinguishing a certain category from the rest, while quite confused among the rest categories. So for each category, models among top-ten percent of cross-entropy loss associated with this specific category are chosen and used for calculating ensemble average. This category-specific average outperforms significantly compared with simple arithmetic/geometric average according to my experience.\r\n\r\n - **Segment average**\r\n\r\nThe last step, segment average, the same as K-nearest-neighbor-average, focuses on the second characteristics of the competition as I mentioned above. I divide the test images into small groups according to their similarities in the 512*7*7-dimensional pool5-feature space. If images in one group display consistent and also confident predictions, I receive this as a strong positive signal and renormalize all the images in that group to share the same predictions. \r\n\r\nAs I read in the forum, some of you have achieved a single-model LB score below 0.15, which is admirable. I believe you can reduce that by around 40% by using K_Nearest_Neighbor Average alone, single model. I hope all of you had enjoyed this competition as much as I did, thanks again.",
    "131692": "thanks for sharing and congratulations for the awesome result!",
    "134678": "what is the distance are u using in the 512x7x7 space?",
    "135025": "Ellen Gao Jian, Thanks\r\n@SecondPlan, Euclidian distance",
    "135026": "Ellen Gao Jian, Thanks\r\n@SecondPlan, Euclidian distance",
    "219170": "Great. Thanks for sharing this. Very much appreciated.",
    "256113": "can you share your demo ?",
    "270008": "Thanks for the write-up.  Super helpful!",
    "366983": "Hi @jacobkie\nCan you share your codes for playing around ? how I can contact you ? can you send me your contact details through messages here ? \n\nthx",
    "723004": "Thanks for sharing!",
    "1012370": "hi thank you for sharing! i'm doing final year project in the university and interested in this topic. do you mind sending e your code so that i can learn from it? thank you very much!!!",
    "3430964": "hlo bro can you send me the dataset +code so that i can use this in my project included this as my first step so if you can send please"
  },
  "source": "meta"
}