{
  "id": 35089,
  "title": "Summary of rank 6 solution",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/writeups/kubilai-summary-of-rank-6-solution",
  "author_name": "",
  "post_date": "2017-06-22T00:41:38.707647200Z",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations for the winners for a great effort!  Glad to see the top solutions differ significantly in score, look forward to hearing your solutions！</p>\n\n<p>The most significant gain for my solution is using SSD to create bounding boxes for the Os.  Huge thanks to <a href=\"https://www.kaggle.com/deveaup\">Paul</a> for providing the bounding box annotations.</p>\n\n<p>My model is an ensemble of vgg-19, xception, resnet 50 and inception v3 models.  The top two models are: vgg-19 fine-tuned with 7 frozen layers and xception with 2 frozen layers.  They are similar in score.  The rest, various fine-tuned and bottle-necked models, are significantly weaker but added slightly to the ensemble's effectiveness.  All models ran with 3-fold CV, and are combined with linear regression to produce the final result.</p>\n\n<p>I attempted to include some \"leak\" features in a second submission, but its result is much worse.  As far as I can see, the final results are straight.  Great thanks to the organizers and admins for making this a useful competition!</p>\n\n<p>The whole pipeline ran on a gtx970 for five days, barely making the submission deadline.  The top two models ran for 1.5 and 2 days each.</p>",
  "messages": [
    {
      "id": "194849",
      "postDate": "06/22/2017 00:41:38",
      "content": "<p>Congratulations for the winners for a great effort!  Glad to see the top solutions differ significantly in score, look forward to hearing your solutions！</p>\n\n<p>The most significant gain for my solution is using SSD to create bounding boxes for the Os.  Huge thanks to <a href=\"https://www.kaggle.com/deveaup\">Paul</a> for providing the bounding box annotations.</p>\n\n<p>My model is an ensemble of vgg-19, xception, resnet 50 and inception v3 models.  The top two models are: vgg-19 fine-tuned with 7 frozen layers and xception with 2 frozen layers.  They are similar in score.  The rest, various fine-tuned and bottle-necked models, are significantly weaker but added slightly to the ensemble's effectiveness.  All models ran with 3-fold CV, and are combined with linear regression to produce the final result.</p>\n\n<p>I attempted to include some \"leak\" features in a second submission, but its result is much worse.  As far as I can see, the final results are straight.  Great thanks to the organizers and admins for making this a useful competition!</p>\n\n<p>The whole pipeline ran on a gtx970 for five days, barely making the submission deadline.  The top two models ran for 1.5 and 2 days each.</p>",
      "rawMarkdown": "Congratulations for the winners for a great effort!  Glad to see the top solutions differ significantly in score, look forward to hearing your solutions！\n\nThe most significant gain for my solution is using SSD to create bounding boxes for the Os.  Huge thanks to [Paul][1] for providing the bounding box annotations.\n\nMy model is an ensemble of vgg-19, xception, resnet 50 and inception v3 models.  The top two models are: vgg-19 fine-tuned with 7 frozen layers and xception with 2 frozen layers.  They are similar in score.  The rest, various fine-tuned and bottle-necked models, are significantly weaker but added slightly to the ensemble's effectiveness.  All models ran with 3-fold CV, and are combined with linear regression to produce the final result.\n\nI attempted to include some \"leak\" features in a second submission, but its result is much worse.  As far as I can see, the final results are straight.  Great thanks to the organizers and admins for making this a useful competition!\n\nThe whole pipeline ran on a gtx970 for five days, barely making the submission deadline.  The top two models ran for 1.5 and 2 days each.\n\n  [1]: https://www.kaggle.com/deveaup",
      "votes": null
    },
    {
      "id": "194850",
      "postDate": "06/22/2017 00:55:05",
      "content": "<p>Thanks a lot for sharing your approach Kubilai.One quick question, the bb annotations were not provided for additional train right ? Did you had to predict BB for additional train as well ?</p>",
      "rawMarkdown": "Thanks a lot for sharing your approach Kubilai.One quick question, the bb annotations were not provided for additional train right ? Did you had to predict BB for additional train as well ?",
      "votes": null
    },
    {
      "id": "194858",
      "postDate": "06/22/2017 01:39:51",
      "content": "<p>I didn't have bb annotations for the additional training set and didn't need it.  SSD works pretty well with just the original set.  I used 3-fold CV on that as well and applied all three models to all images to produce the bb list, and then used DBS to combine the bb list into one.  The best submission uses only crops to train.</p>",
      "rawMarkdown": "I didn't have bb annotations for the additional training set and didn't need it.  SSD works pretty well with just the original set.  I used 3-fold CV on that as well and applied all three models to all images to produce the bb list, and then used DBS to combine the bb list into one.  The best submission uses only crops to train.",
      "votes": null
    },
    {
      "id": "194860",
      "postDate": "06/22/2017 01:48:28",
      "content": "<p>Thanks for your sharing! What does DBS mean?</p>",
      "rawMarkdown": "Thanks for your sharing! What does DBS mean?",
      "votes": null
    },
    {
      "id": "194864",
      "postDate": "06/22/2017 01:52:49",
      "content": "<p>dbscan, a clustering algorithm in sklearn.  I used it to good effect in the nature conservancy fisheries competition, and just applied it as is to this comp.</p>",
      "rawMarkdown": "dbscan, a clustering algorithm in sklearn.  I used it to good effect in the nature conservancy fisheries competition, and just applied it as is to this comp.",
      "votes": null
    },
    {
      "id": "194922",
      "postDate": "06/22/2017 07:56:48",
      "content": "<p>so you use dbscan to eliminate outliers, then calculate mean of remained bbxs?</p>",
      "rawMarkdown": "so you use dbscan to eliminate outliers, then calculate mean of remained bbxs?",
      "votes": null
    },
    {
      "id": "195162",
      "postDate": "06/22/2017 21:40:49",
      "content": "<p>More to pick out the top 1 (cancer comp) or few (fisheries comp).  Higher confidence if several models pick out bb's that cluster together.</p>",
      "rawMarkdown": "More to pick out the top 1 (cancer comp) or few (fisheries comp).  Higher confidence if several models pick out bb's that cluster together.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 194850,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "06/22/2017 00:55:05",
      "content": "<p>Thanks a lot for sharing your approach Kubilai.One quick question, the bb annotations were not provided for additional train right ? Did you had to predict BB for additional train as well ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 194858,
          "author_name": "kubilai",
          "author_url": "",
          "post_date": "06/22/2017 01:39:51",
          "content": "<p>I didn't have bb annotations for the additional training set and didn't need it.  SSD works pretty well with just the original set.  I used 3-fold CV on that as well and applied all three models to all images to produce the bb list, and then used DBS to combine the bb list into one.  The best submission uses only crops to train.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194860,
          "author_name": "wangg12",
          "author_url": "",
          "post_date": "06/22/2017 01:48:28",
          "content": "<p>Thanks for your sharing! What does DBS mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194864,
          "author_name": "kubilai",
          "author_url": "",
          "post_date": "06/22/2017 01:52:49",
          "content": "<p>dbscan, a clustering algorithm in sklearn.  I used it to good effect in the nature conservancy fisheries competition, and just applied it as is to this comp.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194922,
          "author_name": "wubalubadubdub",
          "author_url": "",
          "post_date": "06/22/2017 07:56:48",
          "content": "<p>so you use dbscan to eliminate outliers, then calculate mean of remained bbxs?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195162,
          "author_name": "kubilai",
          "author_url": "",
          "post_date": "06/22/2017 21:40:49",
          "content": "<p>More to pick out the top 1 (cancer comp) or few (fisheries comp).  Higher confidence if several models pick out bb's that cluster together.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "194849": "Congratulations for the winners for a great effort!  Glad to see the top solutions differ significantly in score, look forward to hearing your solutions！\n\nThe most significant gain for my solution is using SSD to create bounding boxes for the Os.  Huge thanks to [Paul][1] for providing the bounding box annotations.\n\nMy model is an ensemble of vgg-19, xception, resnet 50 and inception v3 models.  The top two models are: vgg-19 fine-tuned with 7 frozen layers and xception with 2 frozen layers.  They are similar in score.  The rest, various fine-tuned and bottle-necked models, are significantly weaker but added slightly to the ensemble's effectiveness.  All models ran with 3-fold CV, and are combined with linear regression to produce the final result.\n\nI attempted to include some \"leak\" features in a second submission, but its result is much worse.  As far as I can see, the final results are straight.  Great thanks to the organizers and admins for making this a useful competition!\n\nThe whole pipeline ran on a gtx970 for five days, barely making the submission deadline.  The top two models ran for 1.5 and 2 days each.\n\n  [1]: https://www.kaggle.com/deveaup",
    "194850": "Thanks a lot for sharing your approach Kubilai.One quick question, the bb annotations were not provided for additional train right ? Did you had to predict BB for additional train as well ?",
    "194858": "I didn't have bb annotations for the additional training set and didn't need it.  SSD works pretty well with just the original set.  I used 3-fold CV on that as well and applied all three models to all images to produce the bb list, and then used DBS to combine the bb list into one.  The best submission uses only crops to train.",
    "194860": "Thanks for your sharing! What does DBS mean?",
    "194864": "dbscan, a clustering algorithm in sklearn.  I used it to good effect in the nature conservancy fisheries competition, and just applied it as is to this comp.",
    "194922": "so you use dbscan to eliminate outliers, then calculate mean of remained bbxs?",
    "195162": "More to pick out the top 1 (cancer comp) or few (fisheries comp).  Higher confidence if several models pick out bb's that cluster together."
  },
  "source": "meta"
}