{
  "id": 22618,
  "title": "Image Clustering",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22618",
  "author_name": "",
  "post_date": "2016-08-02T02:27:38.310Z",
  "votes": 2,
  "comment_count": 3,
  "views": 349,
  "content": "<p>Hi all,</p>\n\n<p>I'm trying to understand why clustering the images did improve significantly my public LB score but had almost no effect on the private LB.</p>\n\n<p>What I have tried was:\n- Extract layer 29 from VGG-16 for all the images.\n- Cluster the images using the extracted data (to about 10000 clusters, some of the cluster had only few images and other even more than 20).\n- Checking the clusters showed mostly good clusters (the same driver doing the same thing).\n- Last I averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average\n- My score go from ~0.18 to ~0.16</p>\n\n<p>I guess it's related to how the split between private/public LB was done.</p>\n\n<p>What do you think?</p>",
  "messages": [
    {
      "id": "129741",
      "postDate": "08/02/2016 02:27:38",
      "content": "<p>Hi all,</p>\n\n<p>I'm trying to understand why clustering the images did improve significantly my public LB score but had almost no effect on the private LB.</p>\n\n<p>What I have tried was:\n- Extract layer 29 from VGG-16 for all the images.\n- Cluster the images using the extracted data (to about 10000 clusters, some of the cluster had only few images and other even more than 20).\n- Checking the clusters showed mostly good clusters (the same driver doing the same thing).\n- Last I averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average\n- My score go from ~0.18 to ~0.16</p>\n\n<p>I guess it's related to how the split between private/public LB was done.</p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "Hi all,\r\n\r\nI'm trying to understand why clustering the images did improve significantly my public LB score but had almost no effect on the private LB.\r\n\r\nWhat I have tried was:\r\n- Extract layer 29 from VGG-16 for all the images.\r\n- Cluster the images using the extracted data (to about 10000 clusters, some of the cluster had only few images and other even more than 20).\r\n- Checking the clusters showed mostly good clusters (the same driver doing the same thing).\r\n- Last I averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average\r\n- My score go from ~0.18 to ~0.16\r\n\r\nI guess it's related to how the split between private/public LB was done.\r\n\r\nWhat do you think?",
      "votes": null
    },
    {
      "id": "129743",
      "postDate": "08/02/2016 02:42:17",
      "content": "<p>It's interesting to use clustering to post-process the predictions. Thank you for sharing this approach.\nBut why  &quot;averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average&quot; should give better predictions in the clusters? </p>",
      "rawMarkdown": "It's interesting to use clustering to post-process the predictions. Thank you for sharing this approach.\r\nBut why  \"averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average\" should give better predictions in the clusters?",
      "votes": null
    },
    {
      "id": "129746",
      "postDate": "08/02/2016 03:11:17",
      "content": "<p>Thanks for your response. </p>\n\n<p>My thought was that similar to the 10 samples crops that have been used in the resnet solution, I can use the fact that many of the images are very similar (and sometime even almost identical) to do something similar. </p>\n\n<p>I wasn't sure it will work but I thought it's worth trying. Especially because even training without augmentation was very slow for me so generating more images was not an option. </p>\n\n<p>To my surprise it worked great but only on the public LB.  :(</p>",
      "rawMarkdown": "Thanks for your response. \r\n\r\nMy thought was that similar to the 10 samples crops that have been used in the resnet solution, I can use the fact that many of the images are very similar (and sometime even almost identical) to do something similar. \r\n\r\nI wasn't sure it will work but I thought it's worth trying. Especially because even training without augmentation was very slow for me so generating more images was not an option. \r\n\r\nTo my surprise it worked great but only on the public LB.  :(",
      "votes": null
    },
    {
      "id": "129760",
      "postDate": "08/02/2016 05:36:51",
      "content": "<p>Would it be median, or average that am I thinking?</p>",
      "rawMarkdown": "Would it be median, or average that am I thinking?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129743,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "08/02/2016 02:42:17",
      "content": "<p>It's interesting to use clustering to post-process the predictions. Thank you for sharing this approach.\nBut why  &quot;averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average&quot; should give better predictions in the clusters? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129746,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "08/02/2016 03:11:17",
      "content": "<p>Thanks for your response. </p>\n\n<p>My thought was that similar to the 10 samples crops that have been used in the resnet solution, I can use the fact that many of the images are very similar (and sometime even almost identical) to do something similar. </p>\n\n<p>I wasn't sure it will work but I thought it's worth trying. Especially because even training without augmentation was very slow for me so generating more images was not an option. </p>\n\n<p>To my surprise it worked great but only on the public LB.  :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129760,
      "author_name": "",
      "author_url": "",
      "post_date": "08/02/2016 05:36:51",
      "content": "<p>Would it be median, or average that am I thinking?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129741": "Hi all,\r\n\r\nI'm trying to understand why clustering the images did improve significantly my public LB score but had almost no effect on the private LB.\r\n\r\nWhat I have tried was:\r\n- Extract layer 29 from VGG-16 for all the images.\r\n- Cluster the images using the extracted data (to about 10000 clusters, some of the cluster had only few images and other even more than 20).\r\n- Checking the clusters showed mostly good clusters (the same driver doing the same thing).\r\n- Last I averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average\r\n- My score go from ~0.18 to ~0.16\r\n\r\nI guess it's related to how the split between private/public LB was done.\r\n\r\nWhat do you think?",
    "129743": "It's interesting to use clustering to post-process the predictions. Thank you for sharing this approach.\r\nBut why  \"averaged the predictions of all the images from the same cluster and changed the prediction of each image in the cluster to that average\" should give better predictions in the clusters?",
    "129746": "Thanks for your response. \r\n\r\nMy thought was that similar to the 10 samples crops that have been used in the resnet solution, I can use the fact that many of the images are very similar (and sometime even almost identical) to do something similar. \r\n\r\nI wasn't sure it will work but I thought it's worth trying. Especially because even training without augmentation was very slow for me so generating more images was not an option. \r\n\r\nTo my surprise it worked great but only on the public LB.  :(",
    "129760": "Would it be median, or average that am I thinking?"
  },
  "source": "meta"
}