{
  "id": 82369,
  "title": "5th solution blog post + code ",
  "url": "/competitions/humpback-whale-identification/writeups/zfturbo-weimin-5th-solution-blog-post-code",
  "author_name": "",
  "post_date": "2019-03-12T18:31:00.943Z",
  "votes": 65,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Congrats to the winners! </p>\n\n<p>Updated github code is <a href=\"https://github.com/aaxwaz/Humpback-whale-identification-challenge\">here</a>. </p>\n\n<p>Blog post sharing the solution is <a href=\"https://weiminwang.blog/2019/03/01/whale-identification-5th-place-approach-using-siamese-networks-with-adversarial-training/\">here</a>. </p>",
  "messages": [
    {
      "id": "481058",
      "postDate": "03/01/2019 02:26:27",
      "content": "<p>Congrats to the winners! </p>\n\n<p>Updated github code is <a href=\"https://github.com/aaxwaz/Humpback-whale-identification-challenge\">here</a>. </p>\n\n<p>Blog post sharing the solution is <a href=\"https://weiminwang.blog/2019/03/01/whale-identification-5th-place-approach-using-siamese-networks-with-adversarial-training/\">here</a>. </p>",
      "rawMarkdown": "Congrats to the winners! \n\nUpdated github code is [here](https://github.com/aaxwaz/Humpback-whale-identification-challenge). \n\nBlog post sharing the solution is [here](https://weiminwang.blog/2019/03/01/whale-identification-5th-place-approach-using-siamese-networks-with-adversarial-training/).",
      "votes": null
    },
    {
      "id": "481145",
      "postDate": "03/01/2019 05:01:59",
      "content": "<p>Thanks for such a nice blog post!</p>",
      "rawMarkdown": "Thanks for such a nice blog post!",
      "votes": null
    },
    {
      "id": "481200",
      "postDate": "03/01/2019 06:35:55",
      "content": "<p>Stellar blog !</p>",
      "rawMarkdown": "Stellar blog !",
      "votes": null
    },
    {
      "id": "481245",
      "postDate": "03/01/2019 07:31:23",
      "content": "<p>Thanks for sharing your blog post <a href=\"/weimin\">@weimin</a></p>",
      "rawMarkdown": "Thanks for sharing your blog post @weimin",
      "votes": null
    },
    {
      "id": "481302",
      "postDate": "03/01/2019 08:19:57",
      "content": "<p>Congrats <a href=\"/weimin\">@weimin</a> and thanks for sharing.</p>",
      "rawMarkdown": "Congrats @weimin and thanks for sharing.",
      "votes": null
    },
    {
      "id": "481350",
      "postDate": "03/01/2019 09:44:12",
      "content": "<p>Congrats to your win and Thanks for the blog !!</p>",
      "rawMarkdown": "Congrats to your win and Thanks for the blog !!",
      "votes": null
    },
    {
      "id": "481377",
      "postDate": "03/01/2019 10:20:09",
      "content": "<p>Congratulations Weimin! Nice blog post!</p>",
      "rawMarkdown": "Congratulations Weimin! Nice blog post!",
      "votes": null
    },
    {
      "id": "481396",
      "postDate": "03/01/2019 10:49:39",
      "content": "<p>Congrats on the result! And thx for sharing your approach in the blog post - was a very good read.</p>",
      "rawMarkdown": "Congrats on the result! And thx for sharing your approach in the blog post - was a very good read.",
      "votes": null
    },
    {
      "id": "481414",
      "postDate": "03/01/2019 11:29:50",
      "content": "<p>Thank you for sharing detail in blog, I've enjoyed a lot :) And congratulations!</p>",
      "rawMarkdown": "Thank you for sharing detail in blog, I've enjoyed a lot :) And congratulations!",
      "votes": null
    },
    {
      "id": "481492",
      "postDate": "03/01/2019 13:30:25",
      "content": "<p>hi weiim..Congrats for your position \nI was. curious to know which loss did you use out here ?\nI tried using siemese earlier but was not getting any where .. \n1) densenet121 2) used a resizing squish based method no bounding boxes  3) loss Contrastive loss</p>",
      "rawMarkdown": "hi weiim..Congrats for your position \nI was. curious to know which loss did you use out here ?\nI tried using siemese earlier but was not getting any where .. \n1) densenet121 2) used a resizing squish based method no bounding boxes  3) loss Contrastive loss",
      "votes": null
    },
    {
      "id": "481564",
      "postDate": "03/01/2019 15:05:29",
      "content": "<p>just BCE. As I understand, diff between Contrastive and BCE is square vs log. I believe in this case log will do better. </p>",
      "rawMarkdown": "just BCE. As I understand, diff between Contrastive and BCE is square vs log. I believe in this case log will do better.",
      "votes": null
    },
    {
      "id": "481827",
      "postDate": "03/01/2019 22:18:24",
      "content": "<p>Thanks for sharing and big congrats!</p>",
      "rawMarkdown": "Thanks for sharing and big congrats!",
      "votes": null
    },
    {
      "id": "481976",
      "postDate": "03/02/2019 06:04:37",
      "content": "<p>1) in your page you have provided siamese architecture ... where you fitted BCE loss there... \n2) have you used any metric layer at the end of each parallel path.. what was formula for layer if used...</p>",
      "rawMarkdown": "1) in your page you have provided siamese architecture ... where you fitted BCE loss there... \n2) have you used any metric layer at the end of each parallel path.. what was formula for layer if used...",
      "votes": null
    },
    {
      "id": "482104",
      "postDate": "03/02/2019 10:39:19",
      "content": "<p>Hi,  Weimin Wang, It's very kind of you to write such a neat and understandable blog like that! I was surprised to see your stacking method, with 15 train by train score matrices, we can learn to predict from prediction! I learn very much from you!\nI have some questions:\n1.   How much improvement of the stacking(from where to where) and, how you merge the stacking prediction on the similarity score matrix, to set the score to be zero if stacking prediction is 0 and keep the score if prediction is 1, or just multiply elements of the  two matrix (score &amp; stacking prediction matrix) with respect to its position to be the ultimate <strong>score matrix</strong>?\n2.  4-fold CV seems very computational expensive, especially when you have to train 15 large models, each with 4 fold, so what the batch size is, and how many GPUs you used during the competition\n3.  what's your adversarial sample pair mining method? lapjv or its variant (how you tackle the annoying time consuming problem)or other linear assignment algorithm ?\n4.   with higher performance, the threshold of <strong>new_whale</strong> will be lower, i am interested to know what's your threshold of single model (e.g. single Densenet121) and of your ensemble of 15 models :)\nTHX!</p>",
      "rawMarkdown": "Hi,  Weimin Wang, It's very kind of you to write such a neat and understandable blog like that! I was surprised to see your stacking method, with 15 train by train score matrices, we can learn to predict from prediction! I learn very much from you!\nI have some questions:\n1.   How much improvement of the stacking(from where to where) and, how you merge the stacking prediction on the similarity score matrix, to set the score to be zero if stacking prediction is 0 and keep the score if prediction is 1, or just multiply elements of the  two matrix (score &amp; stacking prediction matrix) with respect to its position to be the ultimate **score matrix**?\n2.  4-fold CV seems very computational expensive, especially when you have to train 15 large models, each with 4 fold, so what the batch size is, and how many GPUs you used during the competition\n3.  what's your adversarial sample pair mining method? lapjv or its variant (how you tackle the annoying time consuming problem)or other linear assignment algorithm ?\n4.   with higher performance, the threshold of **new_whale** will be lower, i am interested to know what's your threshold of single model (e.g. single Densenet121) and of your ensemble of 15 models :)\nTHX!",
      "votes": null
    },
    {
      "id": "482219",
      "postDate": "03/02/2019 14:27:59",
      "content": "<p>1 unfortunately stacking didn't show much improvement on LB for this comp. We ensembeld stacking in our final so it helped a bit overall (like a diff model). \n2 We started early so we slowly trained those models with very limited GPU resources, and batchsize used was between 32 and 64. \n3. We used lap to solve, and cut the matrix into 3~4 sub-blocks to estimate and reduce computation time\n4. We chose threshold that gives us around 2200 new whales on test dataset. So single model thre was 0.99 whereas final ensemble thre was down to around 0.5</p>",
      "rawMarkdown": "1 unfortunately stacking didn't show much improvement on LB for this comp. We ensembeld stacking in our final so it helped a bit overall (like a diff model). \n2 We started early so we slowly trained those models with very limited GPU resources, and batchsize used was between 32 and 64. \n3. We used lap to solve, and cut the matrix into 3~4 sub-blocks to estimate and reduce computation time\n4. We chose threshold that gives us around 2200 new whales on test dataset. So single model thre was 0.99 whereas final ensemble thre was down to around 0.5",
      "votes": null
    },
    {
      "id": "482257",
      "postDate": "03/02/2019 15:39:40",
      "content": "<p>thanks for replying!  well,  every epoch you use <strong>lap</strong>,  do you make different sub-blocks, (e.g. shuffle the order of the score matrix  <strong>rows</strong>  for ensuring every pair of the samples have the same chance for being selected) ?  I think this will give us a near-optimal solution.</p>",
      "rawMarkdown": "thanks for replying!  well,  every epoch you use **lap**,  do you make different sub-blocks, (e.g. shuffle the order of the score matrix  **rows**  for ensuring every pair of the samples have the same chance for being selected) ?  I think this will give us a near-optimal solution.",
      "votes": null
    },
    {
      "id": "482299",
      "postDate": "03/02/2019 16:54:40",
      "content": "<p>yes we shuffle rows every few epochs. </p>",
      "rawMarkdown": "yes we shuffle rows every few epochs.",
      "votes": null
    },
    {
      "id": "482300",
      "postDate": "03/02/2019 16:55:43",
      "content": "<ol>\n<li>It is end when comparing two images - match (1) or unmatch (0)</li>\n<li>not sure what you asked but don;t think we used</li>\n</ol>",
      "rawMarkdown": "1. It is end when comparing two images - match (1) or unmatch (0)\n2. not sure what you asked but don;t think we used",
      "votes": null
    },
    {
      "id": "483761",
      "postDate": "03/05/2019 04:57:49",
      "content": "<p>Congratulations and thanks for wonderful sharing,  which CNN visualization tools you use to visualize activation area?</p>",
      "rawMarkdown": "Congratulations and thanks for wonderful sharing,  which CNN visualization tools you use to visualize activation area?",
      "votes": null
    },
    {
      "id": "484145",
      "postDate": "03/05/2019 16:11:41",
      "content": "<p>恭喜你！请问你的网络结构跟Martin一样吗？</p>",
      "rawMarkdown": "恭喜你！请问你的网络结构跟Martin一样吗？",
      "votes": null
    },
    {
      "id": "485971",
      "postDate": "03/08/2019 06:39:47",
      "content": "<p>what i meant was how did you compared the outputs of siam network  to produce single FC output for BCE...\ni.e computation in distance layer... if you can put high level formula or modules for that..</p>",
      "rawMarkdown": "what i meant was how did you compared the outputs of siam network  to produce single FC output for BCE...\ni.e computation in distance layer... if you can put high level formula or modules for that..",
      "votes": null
    },
    {
      "id": "1488261",
      "postDate": "08/24/2021 07:23:14",
      "content": "<p>thxxx a lot!</p>",
      "rawMarkdown": "thxxx a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1488261,
      "author_name": "jackyzhuo",
      "author_url": "",
      "post_date": "08/24/2021 07:23:14",
      "content": "<p>thxxx a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481145,
      "author_name": "demonplus",
      "author_url": "",
      "post_date": "03/01/2019 05:01:59",
      "content": "<p>Thanks for such a nice blog post!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481200,
      "author_name": "priteshshrivastava",
      "author_url": "",
      "post_date": "03/01/2019 06:35:55",
      "content": "<p>Stellar blog !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481245,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "03/01/2019 07:31:23",
      "content": "<p>Thanks for sharing your blog post <a href=\"/weimin\">@weimin</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481302,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/01/2019 08:19:57",
      "content": "<p>Congrats <a href=\"/weimin\">@weimin</a> and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481350,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "03/01/2019 09:44:12",
      "content": "<p>Congrats to your win and Thanks for the blog !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481377,
      "author_name": "mihaskalic",
      "author_url": "",
      "post_date": "03/01/2019 10:20:09",
      "content": "<p>Congratulations Weimin! Nice blog post!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481396,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "03/01/2019 10:49:39",
      "content": "<p>Congrats on the result! And thx for sharing your approach in the blog post - was a very good read.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481414,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "03/01/2019 11:29:50",
      "content": "<p>Thank you for sharing detail in blog, I've enjoyed a lot :) And congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481492,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "03/01/2019 13:30:25",
      "content": "<p>hi weiim..Congrats for your position \nI was. curious to know which loss did you use out here ?\nI tried using siemese earlier but was not getting any where .. \n1) densenet121 2) used a resizing squish based method no bounding boxes  3) loss Contrastive loss</p>",
      "votes": null,
      "replies": [
        {
          "id": 481564,
          "author_name": "weimin",
          "author_url": "",
          "post_date": "03/01/2019 15:05:29",
          "content": "<p>just BCE. As I understand, diff between Contrastive and BCE is square vs log. I believe in this case log will do better. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 481976,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "03/02/2019 06:04:37",
          "content": "<p>1) in your page you have provided siamese architecture ... where you fitted BCE loss there... \n2) have you used any metric layer at the end of each parallel path.. what was formula for layer if used...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 482300,
          "author_name": "weimin",
          "author_url": "",
          "post_date": "03/02/2019 16:55:43",
          "content": "<ol>\n<li>It is end when comparing two images - match (1) or unmatch (0)</li>\n<li>not sure what you asked but don;t think we used</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 485971,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "03/08/2019 06:39:47",
          "content": "<p>what i meant was how did you compared the outputs of siam network  to produce single FC output for BCE...\ni.e computation in distance layer... if you can put high level formula or modules for that..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 481827,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/01/2019 22:18:24",
      "content": "<p>Thanks for sharing and big congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 482104,
      "author_name": "gengshi",
      "author_url": "",
      "post_date": "03/02/2019 10:39:19",
      "content": "<p>Hi,  Weimin Wang, It's very kind of you to write such a neat and understandable blog like that! I was surprised to see your stacking method, with 15 train by train score matrices, we can learn to predict from prediction! I learn very much from you!\nI have some questions:\n1.   How much improvement of the stacking(from where to where) and, how you merge the stacking prediction on the similarity score matrix, to set the score to be zero if stacking prediction is 0 and keep the score if prediction is 1, or just multiply elements of the  two matrix (score &amp; stacking prediction matrix) with respect to its position to be the ultimate <strong>score matrix</strong>?\n2.  4-fold CV seems very computational expensive, especially when you have to train 15 large models, each with 4 fold, so what the batch size is, and how many GPUs you used during the competition\n3.  what's your adversarial sample pair mining method? lapjv or its variant (how you tackle the annoying time consuming problem)or other linear assignment algorithm ?\n4.   with higher performance, the threshold of <strong>new_whale</strong> will be lower, i am interested to know what's your threshold of single model (e.g. single Densenet121) and of your ensemble of 15 models :)\nTHX!</p>",
      "votes": null,
      "replies": [
        {
          "id": 482219,
          "author_name": "weimin",
          "author_url": "",
          "post_date": "03/02/2019 14:27:59",
          "content": "<p>1 unfortunately stacking didn't show much improvement on LB for this comp. We ensembeld stacking in our final so it helped a bit overall (like a diff model). \n2 We started early so we slowly trained those models with very limited GPU resources, and batchsize used was between 32 and 64. \n3. We used lap to solve, and cut the matrix into 3~4 sub-blocks to estimate and reduce computation time\n4. We chose threshold that gives us around 2200 new whales on test dataset. So single model thre was 0.99 whereas final ensemble thre was down to around 0.5</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 482257,
          "author_name": "gengshi",
          "author_url": "",
          "post_date": "03/02/2019 15:39:40",
          "content": "<p>thanks for replying!  well,  every epoch you use <strong>lap</strong>,  do you make different sub-blocks, (e.g. shuffle the order of the score matrix  <strong>rows</strong>  for ensuring every pair of the samples have the same chance for being selected) ?  I think this will give us a near-optimal solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 482299,
          "author_name": "weimin",
          "author_url": "",
          "post_date": "03/02/2019 16:54:40",
          "content": "<p>yes we shuffle rows every few epochs. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 483761,
      "author_name": "fuxungao",
      "author_url": "",
      "post_date": "03/05/2019 04:57:49",
      "content": "<p>Congratulations and thanks for wonderful sharing,  which CNN visualization tools you use to visualize activation area?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 484145,
      "author_name": "nostayup",
      "author_url": "",
      "post_date": "03/05/2019 16:11:41",
      "content": "<p>恭喜你！请问你的网络结构跟Martin一样吗？</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "481058": "Congrats to the winners! \n\nUpdated github code is [here](https://github.com/aaxwaz/Humpback-whale-identification-challenge). \n\nBlog post sharing the solution is [here](https://weiminwang.blog/2019/03/01/whale-identification-5th-place-approach-using-siamese-networks-with-adversarial-training/).",
    "481145": "Thanks for such a nice blog post!",
    "481200": "Stellar blog !",
    "481245": "Thanks for sharing your blog post @weimin",
    "481302": "Congrats @weimin and thanks for sharing.",
    "481350": "Congrats to your win and Thanks for the blog !!",
    "481377": "Congratulations Weimin! Nice blog post!",
    "481396": "Congrats on the result! And thx for sharing your approach in the blog post - was a very good read.",
    "481414": "Thank you for sharing detail in blog, I've enjoyed a lot :) And congratulations!",
    "481492": "hi weiim..Congrats for your position \nI was. curious to know which loss did you use out here ?\nI tried using siemese earlier but was not getting any where .. \n1) densenet121 2) used a resizing squish based method no bounding boxes  3) loss Contrastive loss",
    "481564": "just BCE. As I understand, diff between Contrastive and BCE is square vs log. I believe in this case log will do better.",
    "481827": "Thanks for sharing and big congrats!",
    "481976": "1) in your page you have provided siamese architecture ... where you fitted BCE loss there... \n2) have you used any metric layer at the end of each parallel path.. what was formula for layer if used...",
    "482104": "Hi,  Weimin Wang, It's very kind of you to write such a neat and understandable blog like that! I was surprised to see your stacking method, with 15 train by train score matrices, we can learn to predict from prediction! I learn very much from you!\nI have some questions:\n1.   How much improvement of the stacking(from where to where) and, how you merge the stacking prediction on the similarity score matrix, to set the score to be zero if stacking prediction is 0 and keep the score if prediction is 1, or just multiply elements of the  two matrix (score &amp; stacking prediction matrix) with respect to its position to be the ultimate **score matrix**?\n2.  4-fold CV seems very computational expensive, especially when you have to train 15 large models, each with 4 fold, so what the batch size is, and how many GPUs you used during the competition\n3.  what's your adversarial sample pair mining method? lapjv or its variant (how you tackle the annoying time consuming problem)or other linear assignment algorithm ?\n4.   with higher performance, the threshold of **new_whale** will be lower, i am interested to know what's your threshold of single model (e.g. single Densenet121) and of your ensemble of 15 models :)\nTHX!",
    "482219": "1 unfortunately stacking didn't show much improvement on LB for this comp. We ensembeld stacking in our final so it helped a bit overall (like a diff model). \n2 We started early so we slowly trained those models with very limited GPU resources, and batchsize used was between 32 and 64. \n3. We used lap to solve, and cut the matrix into 3~4 sub-blocks to estimate and reduce computation time\n4. We chose threshold that gives us around 2200 new whales on test dataset. So single model thre was 0.99 whereas final ensemble thre was down to around 0.5",
    "482257": "thanks for replying!  well,  every epoch you use **lap**,  do you make different sub-blocks, (e.g. shuffle the order of the score matrix  **rows**  for ensuring every pair of the samples have the same chance for being selected) ?  I think this will give us a near-optimal solution.",
    "482299": "yes we shuffle rows every few epochs.",
    "482300": "1. It is end when comparing two images - match (1) or unmatch (0)\n2. not sure what you asked but don;t think we used",
    "483761": "Congratulations and thanks for wonderful sharing,  which CNN visualization tools you use to visualize activation area?",
    "484145": "恭喜你！请问你的网络结构跟Martin一样吗？",
    "485971": "what i meant was how did you compared the outputs of siam network  to produce single FC output for BCE...\ni.e computation in distance layer... if you can put high level formula or modules for that..",
    "1488261": "thxxx a lot!"
  },
  "source": "meta"
}