{
  "id": 82356,
  "title": "4th Place Solution: SIFT + Siamese",
  "url": "/competitions/humpback-whale-identification/writeups/david-4th-place-solution-sift-siamese",
  "author_name": "",
  "post_date": "2019-03-01T00:19:03.431610Z",
  "votes": 70,
  "comment_count": 23,
  "views": 0,
  "content": "<p>My goal in this competition was to learn more about low-shot learning problems and to try to get to GM, so I’ll share what I learned.</p>\n\n<p>I find it useful to try to think like the sponsor and ask why they would host the competition and if I were them what would want to get out of it.  There was already a playground competition, so why release it again?  My thoughts were that 1. maybe Kaggle wanted to show the difference in quality of solutions for a free playground vs a prize value based solution competition, or 2. the sponsor wanted to get more out of the really challenging part of the problem, namely how to identify new_whale (N=0) and N=1 samples.  So my focus was on the latter and specifically how to identify as many N=1 samples as possible.</p>\n\n<p>There are three main components to my pipeline:</p>\n\n<ul>\n<li><strong>Keypoint matching</strong> – old school approach with a few new school tricks</li>\n<li><strong>Siamese network</strong> – like many, <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">Martin’s previous work</a> formed the basis here</li>\n<li><strong>Post-processing</strong> – to give low sample classes a fair shake</li>\n</ul>\n\n<p><strong>Keypoint matching</strong>\nThis accounted for &gt;80% of my final predictions, and was classic keypoint matching, one of the original low-shot methods.  I tried SIFT, ROOTSIFT, and a host of binary descriptors and matchers, there wasn’t a lot of difference between the different techniques.</p>\n\n<p>The dataset here was in the sweet spot where brute force keypoint matching came into play.  7960 test images vs 15,697 train images is within the realm of reason.  I chose the pure brute-force method at full image resolution, all test images vs all train images, no bag-of-words or knn clustering of the keypoints.  There were a couple big challenges I had to overcome:</p>\n\n<ol>\n<li><em>Speed</em>.  Keypoint descriptors/matching can take up to 1-2s per image depending on your HW setup, but I used several tricks like indexing all keypoints to a hdf5 file, storing all keypoints into RAM during matching, and use of the great <a href=\"https://github.com/facebookresearch/faiss\">faiss library</a>.  Across two systems I could finish a brute force run of the full dataset in ~12 hours.</li>\n<li><em>False positives</em>.  The main issue with kp matching on this dataset was the false positives which had two sources: the background ocean and many of the bright points on the whale flukes.  I addressed this by using a unet to segment only the whale tail, and a custom xgboost model of the homography matrix to classify the final homography between image pairs as valid or not.</li>\n</ol>\n\n<p>Final kp matching pipeline:\n - Extract all kps from train and test (raw images, full scale) into hdf5 files, restricting keypoints to unet predicted mask area of whale fluke.  Extracting from CLAHE preprocessed images worked best.\n - Matching:\n    a. Loop through all test/train pairs\n    b. Match keypoints using faiss\n    c. Double homography filtering of keypoints (LMEDS followed by RANSAC)\n    c. xgboost prediction to validate homography matrix\n    d. if # of matches &gt; threshold, then use prediction</p>\n\n<p><strong>Siamese network</strong>\nThis is the weakest part of my pipeline, there are other posts indicating much stronger networks than what I used.  I just adapted Martin’s code a bit, and used an ensemble of InceptionResNetV2, InceptionV3, and ResNet50.  I didn’t add in any augmentations and stuck with grayscale images, nothing fancy.\nTo help training move on a little quicker, I did a fair amount of pretraining of the backbone network before feeding it in the Siamese network, which seemed to help.  My pretraining pipeline was:\n - train classification on top 200 classes\n - fine-tune on all classes where N&gt;8 (~576 classes)\n - fine-tune on all classes\n - fine-tune on all classes + mixup + image size 384x384 </p>\n\n<p><strong>Post processing</strong>\nI found some similarities in the data between this competition and the Statoil Iceberg challenge, so I was able to use some of the same tricks from Weimin and my <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge\">winning solution</a> there, mainly that there were insights from test predictions that could be used to further enhance the test predictions. </p>\n\n<p>When analyzing the resulting prediction matrix from the Siamese network, I noticed that there was always a handful of the same train images that disproportionately dominated the top-5 positions.  This got me thinking that I needed to find a way to either suppress the dominate predictions or figure out how to get the N=1 classes a more fair chance to rise to the top of the prediction pool.</p>\n\n<p>The idea I came up was pretty simple: instead of looking at the prediction matrix in the traditional way of “which train image is closest to my test image”, I transposed the matrix to now look at “which test image is closest to my train images”.  When I limited the transposed matrix to the N=1 samples, I found that I could use a new threshold along the train axis for the N=1 train samples.  This was highly effective at generating many more of the correct N=1 samples in my top-1 prediction.  I’m sure there are better ways of accomplishing the same goal.</p>\n\n<p>I was surprised by the number of mislabels other competitors found, and thanks to <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/81885\">Alex Mokin and the contributors to this post</a> I took advantage of making sure the redundant classes were accounted for appropriately.</p>\n\n<p><strong>Pipleline weaknesses:</strong>\nAgain, thinking like the sponsor, they may not love my solution for a couple of reasons: 1. very computationally expensive, especially the keypoint matching pipeline, and 2. the difficulty to convert the pipleine into an easy way to do single image inference due to the post processing.</p>\n\n<p>I would probably take someone else’s solution who has a strong siamese network and drop it into my pipeline as a pure replacement.  This would require retuning of the post processing pipeline but it’s possible.</p>\n\n<p><strong>Pipleline strengths:</strong>\nI think the keypoint pipeline is pretty strong, without a lot of opportunity to squeeze more if using traditional keypoint algorithms.  The unet and xgb model incorporation into the pipeline really helps cut the false positives to be negligible.</p>",
  "messages": [
    {
      "id": "480981",
      "postDate": "03/01/2019 00:19:03",
      "content": "<p>My goal in this competition was to learn more about low-shot learning problems and to try to get to GM, so I’ll share what I learned.</p>\n\n<p>I find it useful to try to think like the sponsor and ask why they would host the competition and if I were them what would want to get out of it.  There was already a playground competition, so why release it again?  My thoughts were that 1. maybe Kaggle wanted to show the difference in quality of solutions for a free playground vs a prize value based solution competition, or 2. the sponsor wanted to get more out of the really challenging part of the problem, namely how to identify new_whale (N=0) and N=1 samples.  So my focus was on the latter and specifically how to identify as many N=1 samples as possible.</p>\n\n<p>There are three main components to my pipeline:</p>\n\n<ul>\n<li><strong>Keypoint matching</strong> – old school approach with a few new school tricks</li>\n<li><strong>Siamese network</strong> – like many, <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">Martin’s previous work</a> formed the basis here</li>\n<li><strong>Post-processing</strong> – to give low sample classes a fair shake</li>\n</ul>\n\n<p><strong>Keypoint matching</strong>\nThis accounted for &gt;80% of my final predictions, and was classic keypoint matching, one of the original low-shot methods.  I tried SIFT, ROOTSIFT, and a host of binary descriptors and matchers, there wasn’t a lot of difference between the different techniques.</p>\n\n<p>The dataset here was in the sweet spot where brute force keypoint matching came into play.  7960 test images vs 15,697 train images is within the realm of reason.  I chose the pure brute-force method at full image resolution, all test images vs all train images, no bag-of-words or knn clustering of the keypoints.  There were a couple big challenges I had to overcome:</p>\n\n<ol>\n<li><em>Speed</em>.  Keypoint descriptors/matching can take up to 1-2s per image depending on your HW setup, but I used several tricks like indexing all keypoints to a hdf5 file, storing all keypoints into RAM during matching, and use of the great <a href=\"https://github.com/facebookresearch/faiss\">faiss library</a>.  Across two systems I could finish a brute force run of the full dataset in ~12 hours.</li>\n<li><em>False positives</em>.  The main issue with kp matching on this dataset was the false positives which had two sources: the background ocean and many of the bright points on the whale flukes.  I addressed this by using a unet to segment only the whale tail, and a custom xgboost model of the homography matrix to classify the final homography between image pairs as valid or not.</li>\n</ol>\n\n<p>Final kp matching pipeline:\n - Extract all kps from train and test (raw images, full scale) into hdf5 files, restricting keypoints to unet predicted mask area of whale fluke.  Extracting from CLAHE preprocessed images worked best.\n - Matching:\n    a. Loop through all test/train pairs\n    b. Match keypoints using faiss\n    c. Double homography filtering of keypoints (LMEDS followed by RANSAC)\n    c. xgboost prediction to validate homography matrix\n    d. if # of matches &gt; threshold, then use prediction</p>\n\n<p><strong>Siamese network</strong>\nThis is the weakest part of my pipeline, there are other posts indicating much stronger networks than what I used.  I just adapted Martin’s code a bit, and used an ensemble of InceptionResNetV2, InceptionV3, and ResNet50.  I didn’t add in any augmentations and stuck with grayscale images, nothing fancy.\nTo help training move on a little quicker, I did a fair amount of pretraining of the backbone network before feeding it in the Siamese network, which seemed to help.  My pretraining pipeline was:\n - train classification on top 200 classes\n - fine-tune on all classes where N&gt;8 (~576 classes)\n - fine-tune on all classes\n - fine-tune on all classes + mixup + image size 384x384 </p>\n\n<p><strong>Post processing</strong>\nI found some similarities in the data between this competition and the Statoil Iceberg challenge, so I was able to use some of the same tricks from Weimin and my <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge\">winning solution</a> there, mainly that there were insights from test predictions that could be used to further enhance the test predictions. </p>\n\n<p>When analyzing the resulting prediction matrix from the Siamese network, I noticed that there was always a handful of the same train images that disproportionately dominated the top-5 positions.  This got me thinking that I needed to find a way to either suppress the dominate predictions or figure out how to get the N=1 classes a more fair chance to rise to the top of the prediction pool.</p>\n\n<p>The idea I came up was pretty simple: instead of looking at the prediction matrix in the traditional way of “which train image is closest to my test image”, I transposed the matrix to now look at “which test image is closest to my train images”.  When I limited the transposed matrix to the N=1 samples, I found that I could use a new threshold along the train axis for the N=1 train samples.  This was highly effective at generating many more of the correct N=1 samples in my top-1 prediction.  I’m sure there are better ways of accomplishing the same goal.</p>\n\n<p>I was surprised by the number of mislabels other competitors found, and thanks to <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/81885\">Alex Mokin and the contributors to this post</a> I took advantage of making sure the redundant classes were accounted for appropriately.</p>\n\n<p><strong>Pipleline weaknesses:</strong>\nAgain, thinking like the sponsor, they may not love my solution for a couple of reasons: 1. very computationally expensive, especially the keypoint matching pipeline, and 2. the difficulty to convert the pipleine into an easy way to do single image inference due to the post processing.</p>\n\n<p>I would probably take someone else’s solution who has a strong siamese network and drop it into my pipeline as a pure replacement.  This would require retuning of the post processing pipeline but it’s possible.</p>\n\n<p><strong>Pipleline strengths:</strong>\nI think the keypoint pipeline is pretty strong, without a lot of opportunity to squeeze more if using traditional keypoint algorithms.  The unet and xgb model incorporation into the pipeline really helps cut the false positives to be negligible.</p>",
      "rawMarkdown": "My goal in this competition was to learn more about low-shot learning problems and to try to get to GM, so I’ll share what I learned.\n\nI find it useful to try to think like the sponsor and ask why they would host the competition and if I were them what would want to get out of it.  There was already a playground competition, so why release it again?  My thoughts were that 1. maybe Kaggle wanted to show the difference in quality of solutions for a free playground vs a prize value based solution competition, or 2. the sponsor wanted to get more out of the really challenging part of the problem, namely how to identify new_whale (N=0) and N=1 samples.  So my focus was on the latter and specifically how to identify as many N=1 samples as possible.\n\nThere are three main components to my pipeline:\n\n\n \n\n - **Keypoint matching** – old school approach with a few new school tricks\n - **Siamese network** – like many, [Martin’s previous work][1] formed the basis here\n -  **Post-processing** – to give low sample classes a fair shake\n\n  \n\n**Keypoint matching**\nThis accounted for &gt;80% of my final predictions, and was classic keypoint matching, one of the original low-shot methods.  I tried SIFT, ROOTSIFT, and a host of binary descriptors and matchers, there wasn’t a lot of difference between the different techniques.\n\nThe dataset here was in the sweet spot where brute force keypoint matching came into play.  7960 test images vs 15,697 train images is within the realm of reason.  I chose the pure brute-force method at full image resolution, all test images vs all train images, no bag-of-words or knn clustering of the keypoints.  There were a couple big challenges I had to overcome:\n\n1. *Speed*.  Keypoint descriptors/matching can take up to 1-2s per image depending on your HW setup, but I used several tricks like indexing all keypoints to a hdf5 file, storing all keypoints into RAM during matching, and use of the great [faiss library][2].  Across two systems I could finish a brute force run of the full dataset in ~12 hours.\n2. *False positives*.  The main issue with kp matching on this dataset was the false positives which had two sources: the background ocean and many of the bright points on the whale flukes.  I addressed this by using a unet to segment only the whale tail, and a custom xgboost model of the homography matrix to classify the final homography between image pairs as valid or not.\n\nFinal kp matching pipeline:\n - Extract all kps from train and test (raw images, full scale) into hdf5 files, restricting keypoints to unet predicted mask area of whale fluke.  Extracting from CLAHE preprocessed images worked best.\n - Matching:\n\ta. Loop through all test/train pairs\n\tb. Match keypoints using faiss\n\tc. Double homography filtering of keypoints (LMEDS followed by RANSAC)\n\tc. xgboost prediction to validate homography matrix\n\td. if # of matches &gt; threshold, then use prediction\n\n\n\n**Siamese network**\nThis is the weakest part of my pipeline, there are other posts indicating much stronger networks than what I used.  I just adapted Martin’s code a bit, and used an ensemble of InceptionResNetV2, InceptionV3, and ResNet50.  I didn’t add in any augmentations and stuck with grayscale images, nothing fancy.\nTo help training move on a little quicker, I did a fair amount of pretraining of the backbone network before feeding it in the Siamese network, which seemed to help.  My pretraining pipeline was:\n - train classification on top 200 classes\n - fine-tune on all classes where N&gt;8 (~576 classes)\n - fine-tune on all classes\n - fine-tune on all classes + mixup + image size 384x384 \n\n\n**Post processing**\nI found some similarities in the data between this competition and the Statoil Iceberg challenge, so I was able to use some of the same tricks from Weimin and my [winning solution][3] there, mainly that there were insights from test predictions that could be used to further enhance the test predictions. \n\n\nWhen analyzing the resulting prediction matrix from the Siamese network, I noticed that there was always a handful of the same train images that disproportionately dominated the top-5 positions.  This got me thinking that I needed to find a way to either suppress the dominate predictions or figure out how to get the N=1 classes a more fair chance to rise to the top of the prediction pool.\n\n\nThe idea I came up was pretty simple: instead of looking at the prediction matrix in the traditional way of “which train image is closest to my test image”, I transposed the matrix to now look at “which test image is closest to my train images”.  When I limited the transposed matrix to the N=1 samples, I found that I could use a new threshold along the train axis for the N=1 train samples.  This was highly effective at generating many more of the correct N=1 samples in my top-1 prediction.  I’m sure there are better ways of accomplishing the same goal.\n\nI was surprised by the number of mislabels other competitors found, and thanks to [Alex Mokin and the contributors to this post][4] I took advantage of making sure the redundant classes were accounted for appropriately.\n\n**Pipleline weaknesses:**\nAgain, thinking like the sponsor, they may not love my solution for a couple of reasons: 1. very computationally expensive, especially the keypoint matching pipeline, and 2. the difficulty to convert the pipleine into an easy way to do single image inference due to the post processing.\n\nI would probably take someone else’s solution who has a strong siamese network and drop it into my pipeline as a pure replacement.  This would require retuning of the post processing pipeline but it’s possible.\n\n**Pipleline strengths:**\nI think the keypoint pipeline is pretty strong, without a lot of opportunity to squeeze more if using traditional keypoint algorithms.  The unet and xgb model incorporation into the pipeline really helps cut the false positives to be negligible.\n\n\n  [1]: https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\n  [2]: https://github.com/facebookresearch/faiss\n  [3]: https://www.kaggle.com/c/statoil-iceberg-classifier-challenge\n  [4]: https://www.kaggle.com/c/humpback-whale-identification/discussion/81885",
      "votes": null
    },
    {
      "id": "480985",
      "postDate": "03/01/2019 00:22:06",
      "content": "<p>Wow, respect for old school matching. Feeding homography to classifier is new way of validating it for me</p>",
      "rawMarkdown": "Wow, respect for old school matching. Feeding homography to classifier is new way of validating it for me",
      "votes": null
    },
    {
      "id": "480995",
      "postDate": "03/01/2019 00:33:47",
      "content": "<p>I'm happy to see you become a Grand Master, David!</p>\n\n<p>Couple of questions:\n- Did you try any other features besides SIFT (e.g. DL feature extractors)?\n- How does keypoint matching performance compare to siamese networks, was it better?</p>",
      "rawMarkdown": "I'm happy to see you become a Grand Master, David!\n\nCouple of questions:\n- Did you try any other features besides SIFT (e.g. DL feature extractors)?\n- How does keypoint matching performance compare to siamese networks, was it better?",
      "votes": null
    },
    {
      "id": "480996",
      "postDate": "03/01/2019 00:34:02",
      "content": "<p>Congratulation!  May I know your single model scores ? </p>",
      "rawMarkdown": "Congratulation!  May I know your single model scores ?",
      "votes": null
    },
    {
      "id": "481001",
      "postDate": "03/01/2019 00:48:13",
      "content": "<p>Thanks Miha!\nI didn't really explore keypoints outside of the traditional descriptors provided in OpenCV.  I did a single run using DELF descriptors but the results were a bit worse so I just stuck to the basics.</p>\n\n<p>Overall keypoint matching provided me better top-1 performance than siamese networks, for the images they classified.  The problem with keypoints was on the lower resolution images where not many keypoints are detected.  I definitely needed the siamese networks to complement keypoint detection.</p>",
      "rawMarkdown": "Thanks Miha!\nI didn't really explore keypoints outside of the traditional descriptors provided in OpenCV.  I did a single run using DELF descriptors but the results were a bit worse so I just stuck to the basics.\n\nOverall keypoint matching provided me better top-1 performance than siamese networks, for the images they classified.  The problem with keypoints was on the lower resolution images where not many keypoints are detected.  I definitely needed the siamese networks to complement keypoint detection.",
      "votes": null
    },
    {
      "id": "481004",
      "postDate": "03/01/2019 00:52:32",
      "content": "<p>congratulations become a GM  </p>",
      "rawMarkdown": "congratulations become a GM",
      "votes": null
    },
    {
      "id": "481018",
      "postDate": "03/01/2019 01:11:10",
      "content": "<p>Congratulations and thanks for sharing the solution...</p>",
      "rawMarkdown": "Congratulations and thanks for sharing the solution...",
      "votes": null
    },
    {
      "id": "481021",
      "postDate": "03/01/2019 01:11:56",
      "content": "<p>SIFT matching on hi-res images is a really neat idea! Congrats!</p>",
      "rawMarkdown": "SIFT matching on hi-res images is a really neat idea! Congrats!",
      "votes": null
    },
    {
      "id": "481249",
      "postDate": "03/01/2019 07:37:09",
      "content": "<p>Congratulations on becoming competitions grandmaster and for getting gold in this competition. <a href=\"/tivfrvqhs5\">@tivfrvqhs5</a></p>",
      "rawMarkdown": "Congratulations on becoming competitions grandmaster and for getting gold in this competition. @tivfrvqhs5",
      "votes": null
    },
    {
      "id": "481295",
      "postDate": "03/01/2019 08:17:57",
      "content": "<p>Congrats <a href=\"/tivfrvqhs5\">@tivfrvqhs5</a> and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @tivfrvqhs5 and thanks for sharing your solution overview.",
      "votes": null
    },
    {
      "id": "481759",
      "postDate": "03/01/2019 20:10:24",
      "content": "<p>Very cool! Great job!</p>",
      "rawMarkdown": "Very cool! Great job!",
      "votes": null
    },
    {
      "id": "481819",
      "postDate": "03/01/2019 22:12:37",
      "content": "<p>Great approach, congrats!</p>",
      "rawMarkdown": "Great approach, congrats!",
      "votes": null
    },
    {
      "id": "482022",
      "postDate": "03/02/2019 07:52:08",
      "content": "<p>The core idea of key point matching is that samples from the same whale is just two photos taken from two perspective, great! <br>\nI wonder if the prediction of kp matching dominate  (only) top 1 position of your prediction whenever  # of matches &gt; threshold? And siamese is just a way to complement the rest top 2~5, (if kp fails, top 1~5) , right?\nCongratulations on becoming competitions GrandMaster : )</p>",
      "rawMarkdown": "The core idea of key point matching is that samples from the same whale is just two photos taken from two perspective, great!  \nI wonder if the prediction of kp matching dominate  (only) top 1 position of your prediction whenever  # of matches &gt; threshold? And siamese is just a way to complement the rest top 2~5, (if kp fails, top 1~5) , right?\nCongratulations on becoming competitions GrandMaster : )",
      "votes": null
    },
    {
      "id": "482026",
      "postDate": "03/02/2019 07:55:15",
      "content": "<p>yes, you have the idea right.  kp matching was used for top1 replacement only, and complemented a full prediction matrix from a siamese network.</p>",
      "rawMarkdown": "yes, you have the idea right.  kp matching was used for top1 replacement only, and complemented a full prediction matrix from a siamese network.",
      "votes": null
    },
    {
      "id": "485619",
      "postDate": "03/07/2019 16:54:37",
      "content": "<p>Congrats!</p>\n\n<p>Achieving this rank with this approach is unbelievable!\nYou can use this pipeline for upcoming Google Landmark Recognition Challenge =D</p>\n\n<p>By the way, what SIFT implementation did you use? (is the detector also SIFT?)</p>",
      "rawMarkdown": "Congrats!\n\nAchieving this rank with this approach is unbelievable!\nYou can use this pipeline for upcoming Google Landmark Recognition Challenge =D\n\nBy the way, what SIFT implementation did you use? (is the detector also SIFT?)",
      "votes": null
    },
    {
      "id": "485620",
      "postDate": "03/07/2019 16:57:37",
      "content": "<p>How did you treat new_whale predictions?\nThey appear only in top-1 results where kp fails?</p>",
      "rawMarkdown": "How did you treat new_whale predictions?\nThey appear only in top-1 results where kp fails?",
      "votes": null
    },
    {
      "id": "486533",
      "postDate": "03/08/2019 23:32:09",
      "content": "<p>SIFT detector and RootSIFT extractor</p>",
      "rawMarkdown": "SIFT detector and RootSIFT extractor",
      "votes": null
    },
    {
      "id": "486534",
      "postDate": "03/08/2019 23:33:27",
      "content": "<p>new_whale was only predicted from siamese network.  kps matches via sift became top-1 predictions where a match exists, otherwise default is siamese</p>",
      "rawMarkdown": "new_whale was only predicted from siamese network.  kps matches via sift became top-1 predictions where a match exists, otherwise default is siamese",
      "votes": null
    },
    {
      "id": "487081",
      "postDate": "03/10/2019 05:10:13",
      "content": "<p>Thank you for your reply!</p>",
      "rawMarkdown": "Thank you for your reply!",
      "votes": null
    },
    {
      "id": "487083",
      "postDate": "03/10/2019 05:11:59",
      "content": "<p>I see. Sounds reasonable.</p>",
      "rawMarkdown": "I see. Sounds reasonable.",
      "votes": null
    },
    {
      "id": "490191",
      "postDate": "03/14/2019 11:04:53",
      "content": "<p>Congratulations and thanks for sharing the solution.\nCould you share your code?I want to learn your solution.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing the solution.\nCould you share your code?I want to learn your solution.",
      "votes": null
    },
    {
      "id": "490867",
      "postDate": "03/14/2019 23:44:41",
      "content": "<p>Posted:\n<a href=\"https://github.com/daustingm1/humpback-whale-4th-place\">https://github.com/daustingm1/humpback-whale-4th-place</a></p>",
      "rawMarkdown": "Posted:\nhttps://github.com/daustingm1/humpback-whale-4th-place",
      "votes": null
    },
    {
      "id": "496460",
      "postDate": "03/22/2019 08:25:34",
      "content": "<p><a href=\"/tivfrvqhs5\">@tivfrvqhs5</a> I see you use <code>mask_predictions_test_4</code>.zip and <code>mask_predictions_train_known_4.zip</code> in your repo, but I can't find where mentioned in this repo. Can you explain it (what is it, how to construct it?). Thanks!</p>",
      "rawMarkdown": "tivfrvqhs5 I see you use `mask_predictions_test_4`.zip and `mask_predictions_train_known_4.zip` in your repo, but I can't find where mentioned in this repo. Can you explain it (what is it, how to construct it?). Thanks!",
      "votes": null
    },
    {
      "id": "583131",
      "postDate": "07/24/2019 04:28:44",
      "content": "<p>Thank you very much for the sharing code!</p>\n\n<p>If you don't mind, I'm glad if you show us how to train/create xgboost prediction model (<code>final_model_gb.pkl</code>)</p>\n\n<p>I was really surprised that keypoint matching approach is such powerful approach.</p>",
      "rawMarkdown": "Thank you very much for the sharing code!\n\nIf you don't mind, I'm glad if you show us how to train/create xgboost prediction model (`final_model_gb.pkl`)\n\nI was really surprised that keypoint matching approach is such powerful approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 480985,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "03/01/2019 00:22:06",
      "content": "<p>Wow, respect for old school matching. Feeding homography to classifier is new way of validating it for me</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 480995,
      "author_name": "mihaskalic",
      "author_url": "",
      "post_date": "03/01/2019 00:33:47",
      "content": "<p>I'm happy to see you become a Grand Master, David!</p>\n\n<p>Couple of questions:\n- Did you try any other features besides SIFT (e.g. DL feature extractors)?\n- How does keypoint matching performance compare to siamese networks, was it better?</p>",
      "votes": null,
      "replies": [
        {
          "id": 481001,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "03/01/2019 00:48:13",
          "content": "<p>Thanks Miha!\nI didn't really explore keypoints outside of the traditional descriptors provided in OpenCV.  I did a single run using DELF descriptors but the results were a bit worse so I just stuck to the basics.</p>\n\n<p>Overall keypoint matching provided me better top-1 performance than siamese networks, for the images they classified.  The problem with keypoints was on the lower resolution images where not many keypoints are detected.  I definitely needed the siamese networks to complement keypoint detection.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 480996,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "03/01/2019 00:34:02",
      "content": "<p>Congratulation!  May I know your single model scores ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481004,
      "author_name": "withli",
      "author_url": "",
      "post_date": "03/01/2019 00:52:32",
      "content": "<p>congratulations become a GM  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481018,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "03/01/2019 01:11:10",
      "content": "<p>Congratulations and thanks for sharing the solution...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481021,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "03/01/2019 01:11:56",
      "content": "<p>SIFT matching on hi-res images is a really neat idea! Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481249,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "03/01/2019 07:37:09",
      "content": "<p>Congratulations on becoming competitions grandmaster and for getting gold in this competition. <a href=\"/tivfrvqhs5\">@tivfrvqhs5</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481295,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/01/2019 08:17:57",
      "content": "<p>Congrats <a href=\"/tivfrvqhs5\">@tivfrvqhs5</a> and thanks for sharing your solution overview.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481759,
      "author_name": "noonv13",
      "author_url": "",
      "post_date": "03/01/2019 20:10:24",
      "content": "<p>Very cool! Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481819,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/01/2019 22:12:37",
      "content": "<p>Great approach, congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 482022,
      "author_name": "gengshi",
      "author_url": "",
      "post_date": "03/02/2019 07:52:08",
      "content": "<p>The core idea of key point matching is that samples from the same whale is just two photos taken from two perspective, great! <br>\nI wonder if the prediction of kp matching dominate  (only) top 1 position of your prediction whenever  # of matches &gt; threshold? And siamese is just a way to complement the rest top 2~5, (if kp fails, top 1~5) , right?\nCongratulations on becoming competitions GrandMaster : )</p>",
      "votes": null,
      "replies": [
        {
          "id": 482026,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "03/02/2019 07:55:15",
          "content": "<p>yes, you have the idea right.  kp matching was used for top1 replacement only, and complemented a full prediction matrix from a siamese network.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 485620,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "03/07/2019 16:57:37",
          "content": "<p>How did you treat new_whale predictions?\nThey appear only in top-1 results where kp fails?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 486534,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "03/08/2019 23:33:27",
          "content": "<p>new_whale was only predicted from siamese network.  kps matches via sift became top-1 predictions where a match exists, otherwise default is siamese</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 487083,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "03/10/2019 05:11:59",
          "content": "<p>I see. Sounds reasonable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 485619,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "03/07/2019 16:54:37",
      "content": "<p>Congrats!</p>\n\n<p>Achieving this rank with this approach is unbelievable!\nYou can use this pipeline for upcoming Google Landmark Recognition Challenge =D</p>\n\n<p>By the way, what SIFT implementation did you use? (is the detector also SIFT?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 486533,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "03/08/2019 23:32:09",
          "content": "<p>SIFT detector and RootSIFT extractor</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 487081,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "03/10/2019 05:10:13",
          "content": "<p>Thank you for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 490191,
      "author_name": "durantdp",
      "author_url": "",
      "post_date": "03/14/2019 11:04:53",
      "content": "<p>Congratulations and thanks for sharing the solution.\nCould you share your code?I want to learn your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 490867,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "03/14/2019 23:44:41",
          "content": "<p>Posted:\n<a href=\"https://github.com/daustingm1/humpback-whale-4th-place\">https://github.com/daustingm1/humpback-whale-4th-place</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496460,
          "author_name": "",
          "author_url": "",
          "post_date": "03/22/2019 08:25:34",
          "content": "<p><a href=\"/tivfrvqhs5\">@tivfrvqhs5</a> I see you use <code>mask_predictions_test_4</code>.zip and <code>mask_predictions_train_known_4.zip</code> in your repo, but I can't find where mentioned in this repo. Can you explain it (what is it, how to construct it?). Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 583131,
          "author_name": "smatsumoto",
          "author_url": "",
          "post_date": "07/24/2019 04:28:44",
          "content": "<p>Thank you very much for the sharing code!</p>\n\n<p>If you don't mind, I'm glad if you show us how to train/create xgboost prediction model (<code>final_model_gb.pkl</code>)</p>\n\n<p>I was really surprised that keypoint matching approach is such powerful approach.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "480981": "My goal in this competition was to learn more about low-shot learning problems and to try to get to GM, so I’ll share what I learned.\n\nI find it useful to try to think like the sponsor and ask why they would host the competition and if I were them what would want to get out of it.  There was already a playground competition, so why release it again?  My thoughts were that 1. maybe Kaggle wanted to show the difference in quality of solutions for a free playground vs a prize value based solution competition, or 2. the sponsor wanted to get more out of the really challenging part of the problem, namely how to identify new_whale (N=0) and N=1 samples.  So my focus was on the latter and specifically how to identify as many N=1 samples as possible.\n\nThere are three main components to my pipeline:\n\n\n \n\n - **Keypoint matching** – old school approach with a few new school tricks\n - **Siamese network** – like many, [Martin’s previous work][1] formed the basis here\n -  **Post-processing** – to give low sample classes a fair shake\n\n  \n\n**Keypoint matching**\nThis accounted for &gt;80% of my final predictions, and was classic keypoint matching, one of the original low-shot methods.  I tried SIFT, ROOTSIFT, and a host of binary descriptors and matchers, there wasn’t a lot of difference between the different techniques.\n\nThe dataset here was in the sweet spot where brute force keypoint matching came into play.  7960 test images vs 15,697 train images is within the realm of reason.  I chose the pure brute-force method at full image resolution, all test images vs all train images, no bag-of-words or knn clustering of the keypoints.  There were a couple big challenges I had to overcome:\n\n1. *Speed*.  Keypoint descriptors/matching can take up to 1-2s per image depending on your HW setup, but I used several tricks like indexing all keypoints to a hdf5 file, storing all keypoints into RAM during matching, and use of the great [faiss library][2].  Across two systems I could finish a brute force run of the full dataset in ~12 hours.\n2. *False positives*.  The main issue with kp matching on this dataset was the false positives which had two sources: the background ocean and many of the bright points on the whale flukes.  I addressed this by using a unet to segment only the whale tail, and a custom xgboost model of the homography matrix to classify the final homography between image pairs as valid or not.\n\nFinal kp matching pipeline:\n - Extract all kps from train and test (raw images, full scale) into hdf5 files, restricting keypoints to unet predicted mask area of whale fluke.  Extracting from CLAHE preprocessed images worked best.\n - Matching:\n\ta. Loop through all test/train pairs\n\tb. Match keypoints using faiss\n\tc. Double homography filtering of keypoints (LMEDS followed by RANSAC)\n\tc. xgboost prediction to validate homography matrix\n\td. if # of matches &gt; threshold, then use prediction\n\n\n\n**Siamese network**\nThis is the weakest part of my pipeline, there are other posts indicating much stronger networks than what I used.  I just adapted Martin’s code a bit, and used an ensemble of InceptionResNetV2, InceptionV3, and ResNet50.  I didn’t add in any augmentations and stuck with grayscale images, nothing fancy.\nTo help training move on a little quicker, I did a fair amount of pretraining of the backbone network before feeding it in the Siamese network, which seemed to help.  My pretraining pipeline was:\n - train classification on top 200 classes\n - fine-tune on all classes where N&gt;8 (~576 classes)\n - fine-tune on all classes\n - fine-tune on all classes + mixup + image size 384x384 \n\n\n**Post processing**\nI found some similarities in the data between this competition and the Statoil Iceberg challenge, so I was able to use some of the same tricks from Weimin and my [winning solution][3] there, mainly that there were insights from test predictions that could be used to further enhance the test predictions. \n\n\nWhen analyzing the resulting prediction matrix from the Siamese network, I noticed that there was always a handful of the same train images that disproportionately dominated the top-5 positions.  This got me thinking that I needed to find a way to either suppress the dominate predictions or figure out how to get the N=1 classes a more fair chance to rise to the top of the prediction pool.\n\n\nThe idea I came up was pretty simple: instead of looking at the prediction matrix in the traditional way of “which train image is closest to my test image”, I transposed the matrix to now look at “which test image is closest to my train images”.  When I limited the transposed matrix to the N=1 samples, I found that I could use a new threshold along the train axis for the N=1 train samples.  This was highly effective at generating many more of the correct N=1 samples in my top-1 prediction.  I’m sure there are better ways of accomplishing the same goal.\n\nI was surprised by the number of mislabels other competitors found, and thanks to [Alex Mokin and the contributors to this post][4] I took advantage of making sure the redundant classes were accounted for appropriately.\n\n**Pipleline weaknesses:**\nAgain, thinking like the sponsor, they may not love my solution for a couple of reasons: 1. very computationally expensive, especially the keypoint matching pipeline, and 2. the difficulty to convert the pipleine into an easy way to do single image inference due to the post processing.\n\nI would probably take someone else’s solution who has a strong siamese network and drop it into my pipeline as a pure replacement.  This would require retuning of the post processing pipeline but it’s possible.\n\n**Pipleline strengths:**\nI think the keypoint pipeline is pretty strong, without a lot of opportunity to squeeze more if using traditional keypoint algorithms.  The unet and xgb model incorporation into the pipeline really helps cut the false positives to be negligible.\n\n\n  [1]: https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\n  [2]: https://github.com/facebookresearch/faiss\n  [3]: https://www.kaggle.com/c/statoil-iceberg-classifier-challenge\n  [4]: https://www.kaggle.com/c/humpback-whale-identification/discussion/81885",
    "480985": "Wow, respect for old school matching. Feeding homography to classifier is new way of validating it for me",
    "480995": "I'm happy to see you become a Grand Master, David!\n\nCouple of questions:\n- Did you try any other features besides SIFT (e.g. DL feature extractors)?\n- How does keypoint matching performance compare to siamese networks, was it better?",
    "480996": "Congratulation!  May I know your single model scores ?",
    "481001": "Thanks Miha!\nI didn't really explore keypoints outside of the traditional descriptors provided in OpenCV.  I did a single run using DELF descriptors but the results were a bit worse so I just stuck to the basics.\n\nOverall keypoint matching provided me better top-1 performance than siamese networks, for the images they classified.  The problem with keypoints was on the lower resolution images where not many keypoints are detected.  I definitely needed the siamese networks to complement keypoint detection.",
    "481004": "congratulations become a GM",
    "481018": "Congratulations and thanks for sharing the solution...",
    "481021": "SIFT matching on hi-res images is a really neat idea! Congrats!",
    "481249": "Congratulations on becoming competitions grandmaster and for getting gold in this competition. @tivfrvqhs5",
    "481295": "Congrats @tivfrvqhs5 and thanks for sharing your solution overview.",
    "481759": "Very cool! Great job!",
    "481819": "Great approach, congrats!",
    "482022": "The core idea of key point matching is that samples from the same whale is just two photos taken from two perspective, great!  \nI wonder if the prediction of kp matching dominate  (only) top 1 position of your prediction whenever  # of matches &gt; threshold? And siamese is just a way to complement the rest top 2~5, (if kp fails, top 1~5) , right?\nCongratulations on becoming competitions GrandMaster : )",
    "482026": "yes, you have the idea right.  kp matching was used for top1 replacement only, and complemented a full prediction matrix from a siamese network.",
    "485619": "Congrats!\n\nAchieving this rank with this approach is unbelievable!\nYou can use this pipeline for upcoming Google Landmark Recognition Challenge =D\n\nBy the way, what SIFT implementation did you use? (is the detector also SIFT?)",
    "485620": "How did you treat new_whale predictions?\nThey appear only in top-1 results where kp fails?",
    "486533": "SIFT detector and RootSIFT extractor",
    "486534": "new_whale was only predicted from siamese network.  kps matches via sift became top-1 predictions where a match exists, otherwise default is siamese",
    "487081": "Thank you for your reply!",
    "487083": "I see. Sounds reasonable.",
    "490191": "Congratulations and thanks for sharing the solution.\nCould you share your code?I want to learn your solution.",
    "490867": "Posted:\nhttps://github.com/daustingm1/humpback-whale-4th-place",
    "496460": "tivfrvqhs5 I see you use `mask_predictions_test_4`.zip and `mask_predictions_train_known_4.zip` in your repo, but I can't find where mentioned in this repo. Can you explain it (what is it, how to construct it?). Thanks!",
    "583131": "Thank you very much for the sharing code!\n\nIf you don't mind, I'm glad if you show us how to train/create xgboost prediction model (`final_model_gb.pkl`)\n\nI was really surprised that keypoint matching approach is such powerful approach."
  },
  "source": "meta"
}