{
  "id": 47976,
  "title": "Transfer learning code -- 50 percent leaderboard accuracy",
  "url": "/competitions/sp-society-camera-model-identification/discussion/47976",
  "author_name": "",
  "post_date": "2018-01-21T19:35:06.013373Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The code is here: <a href=\"https://github.com/yuanqing811/CID\">https://github.com/yuanqing811/CID</a></p>\n\n<p>The code achieves about 90 percent validation accuracy, but only 50% leaderboard accuracy. Hope you find it useful as a starting point. </p>",
  "messages": [
    {
      "id": "271881",
      "postDate": "01/21/2018 19:35:06",
      "content": "<p>The code is here: <a href=\"https://github.com/yuanqing811/CID\">https://github.com/yuanqing811/CID</a></p>\n\n<p>The code achieves about 90 percent validation accuracy, but only 50% leaderboard accuracy. Hope you find it useful as a starting point. </p>",
      "rawMarkdown": "The code is here: https://github.com/yuanqing811/CID\n\nThe code achieves about 90 percent validation accuracy, but only 50% leaderboard accuracy. Hope you find it useful as a starting point.",
      "votes": null
    },
    {
      "id": "272188",
      "postDate": "01/22/2018 13:06:25",
      "content": "<p>Any idea why the jump from validation to test accuracy? (I have the same issue)</p>",
      "rawMarkdown": "Any idea why the jump from validation to test accuracy? (I have the same issue)",
      "votes": null
    },
    {
      "id": "272233",
      "postDate": "01/22/2018 15:37:10",
      "content": "<p>Hi, not really, but couple of pointers that may help. The code as posted used 25 patches from each training image from the central 1120x1120 crop. With this, we got 90% validation accuracy and 50% LB accuracy. When we used 4 patches from each image from the central 448x448 crop, we got 75% validation accuracy but 55% LB accuracy. Also, our best submission of 62% LB accuracy is from a self trained network (4 CNN + 2 FC layers) which had a even smaller gap between validation and LB accuracy.</p>\n\n<p>Hope that helps.</p>",
      "rawMarkdown": "Hi, not really, but couple of pointers that may help. The code as posted used 25 patches from each training image from the central 1120x1120 crop. With this, we got 90% validation accuracy and 50% LB accuracy. When we used 4 patches from each image from the central 448x448 crop, we got 75% validation accuracy but 55% LB accuracy. Also, our best submission of 62% LB accuracy is from a self trained network (4 CNN + 2 FC layers) which had a even smaller gap between validation and LB accuracy.\n\nHope that helps.",
      "votes": null
    },
    {
      "id": "274358",
      "postDate": "01/26/2018 12:22:19",
      "content": "<p>Hey QingYuan,\nmaybe there is some bias in the composition of test data. \nI am thinking if your model performs relatively poor on a specific category and the test data contains a much higher fraction of this category.</p>",
      "rawMarkdown": "Hey QingYuan,\nmaybe there is some bias in the composition of test data. \nI am thinking if your model performs relatively poor on a specific category and the test data contains a much higher fraction of this category.",
      "votes": null
    },
    {
      "id": "274895",
      "postDate": "01/27/2018 16:51:14",
      "content": "<p>hi FabSchreiber, thanks for the comment. It is certainly a possibility. Based on Andres Torrubia's breakthroughs it seems like that using larger/diverse dataset (like Gleb's dataset) may be the key to getting better results. Not sure if Andres figured out any other reason for the discrepancy.  </p>",
      "rawMarkdown": "hi FabSchreiber, thanks for the comment. It is certainly a possibility. Based on Andres Torrubia's breakthroughs it seems like that using larger/diverse dataset (like Gleb's dataset) may be the key to getting better results. Not sure if Andres figured out any other reason for the discrepancy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 272188,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/22/2018 13:06:25",
      "content": "<p>Any idea why the jump from validation to test accuracy? (I have the same issue)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 272233,
      "author_name": "yuanqing",
      "author_url": "",
      "post_date": "01/22/2018 15:37:10",
      "content": "<p>Hi, not really, but couple of pointers that may help. The code as posted used 25 patches from each training image from the central 1120x1120 crop. With this, we got 90% validation accuracy and 50% LB accuracy. When we used 4 patches from each image from the central 448x448 crop, we got 75% validation accuracy but 55% LB accuracy. Also, our best submission of 62% LB accuracy is from a self trained network (4 CNN + 2 FC layers) which had a even smaller gap between validation and LB accuracy.</p>\n\n<p>Hope that helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274358,
      "author_name": "fabsta",
      "author_url": "",
      "post_date": "01/26/2018 12:22:19",
      "content": "<p>Hey QingYuan,\nmaybe there is some bias in the composition of test data. \nI am thinking if your model performs relatively poor on a specific category and the test data contains a much higher fraction of this category.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274895,
      "author_name": "yuanqing",
      "author_url": "",
      "post_date": "01/27/2018 16:51:14",
      "content": "<p>hi FabSchreiber, thanks for the comment. It is certainly a possibility. Based on Andres Torrubia's breakthroughs it seems like that using larger/diverse dataset (like Gleb's dataset) may be the key to getting better results. Not sure if Andres figured out any other reason for the discrepancy.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "271881": "The code is here: https://github.com/yuanqing811/CID\n\nThe code achieves about 90 percent validation accuracy, but only 50% leaderboard accuracy. Hope you find it useful as a starting point.",
    "272188": "Any idea why the jump from validation to test accuracy? (I have the same issue)",
    "272233": "Hi, not really, but couple of pointers that may help. The code as posted used 25 patches from each training image from the central 1120x1120 crop. With this, we got 90% validation accuracy and 50% LB accuracy. When we used 4 patches from each image from the central 448x448 crop, we got 75% validation accuracy but 55% LB accuracy. Also, our best submission of 62% LB accuracy is from a self trained network (4 CNN + 2 FC layers) which had a even smaller gap between validation and LB accuracy.\n\nHope that helps.",
    "274358": "Hey QingYuan,\nmaybe there is some bias in the composition of test data. \nI am thinking if your model performs relatively poor on a specific category and the test data contains a much higher fraction of this category.",
    "274895": "hi FabSchreiber, thanks for the comment. It is certainly a possibility. Based on Andres Torrubia's breakthroughs it seems like that using larger/diverse dataset (like Gleb's dataset) may be the key to getting better results. Not sure if Andres figured out any other reason for the discrepancy."
  },
  "source": "meta"
}