{
  "id": 328281,
  "title": "1st place solution - New augmentation method + 5 model ensemble",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328281",
  "author_name": "David Austin",
  "post_date": "2022-05-31T17:57:21.937000",
  "votes": 30,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks to the sponsors and Kaggle for organizing this important research challenge.</p>\n<p><strong><em>TLDR</em></strong>: Solution architecture includes 5 model ensemble at two different image sizes with embedding dimension reduced by PCA.  Secret sauce is a newly developed augmentation method designed to deal with large occlusions while maintaining semantic cohesion of the overall scene.</p>\n<h2>Augmentation - Introducing BlendFlip:</h2>\n<p>The core problem in this challenge is dealing with large masked occlusions in test/query images.  For this my goal was to inpaint the masked region while attempting to:</p>\n<ul>\n<li>maintain semantic cohesion with the overall scene</li>\n<li>maintain continuity around masked edges</li>\n<li>fill in visually realistic content</li>\n</ul>\n<p>When looking at hotel room images, a common pattern is many vertical and horizontal features against repeating backgrounds.  For instance wall corners that run vertical, door frames that run vertical and horizontal, pictures on walls that are square, etc.  This observation gave me the idea to apply image patches from the same image directly adjacent to the masked region and essentially flip them into the masked region, and then blend the flips from the vertical and horizontal direction, hence the name BlendFlip.</p>\n<p><img src=\"https://i.imgur.com/q8ZAdPK.jpg\" alt=\"BlendFlip\"></p>\n<p>BlendFlip algorithm:</p>\n<ol>\n<li>Determine max distance between masked region edges in the vertical direction (image top to mask top, mask bottom to image bottom).  Do same for horizontal.  Max direction becomes the regions to be used for inpainting.</li>\n<li>Flip max vertical region into masked region.  If initial flip doesn’t fill entire masked region, keep flipping in vertical region until masked region is filled, cut off at edge of masked region.</li>\n<li>Repeat step 2 for horizontal region.</li>\n<li>Take 50/50 blend of step2 and step3 (cv2.addWeighted) as final inpainted region.</li>\n</ol>\n<p>BlendFlip helps address 2 1/2 of the 3 problems listed above.  It helps maintain semantic cohesion with visually realistic content.  It maintains continuity on 2 of the 4 sides of the masked edges.  I tried GAN's, traditional CV inpainting methods, variations of CutMix, but BlendFlip significantly outperformed them all.</p>\n<p><img src=\"https://i.imgur.com/sl6d9IK.jpg\" alt=\"Example 1\"><br>\n<img src=\"https://i.imgur.com/i18vrXt.jpg\" alt=\"Example 2\"><br>\n<img src=\"https://i.imgur.com/BxLju54.jpg\" alt=\"Example 3\"></p>\n<p>The goal of blendflip is NOT to predict what might be in the occluded region, but rather maintain the integrity of the image embedding from the non-occluded regions.</p>\n<h2>Solution architecture</h2>\n<p>The overall solution architecture is shown in the diagram below.  Images of sizes 1024x1024 (longest max side scaled to 1024, aspect ratio maintained) or 384x384 are fed into 5 models of various architecture.  Each model is trained with ArcFace loss and the embedding size of each is 1536D.  The embeddings are then concatenated and reduced to a total of 3072D via PCA (size determined by maintaining 99% of variance of embedding concat), and then KNN performed.  Each model is trained with a 50% probability of BlendFlip on each image.  BlendFlip is performed on all test images prior to inference.  No post processing or re-ranking was applied. </p>\n<p><img src=\"https://i.imgur.com/0vbOz3Z.jpg\" alt=\"Solution architecture\"></p>\n<p>For the CVPR paper I will include proper ablation studies and relative impact of the solution ingredients.  BlendFlip accounted for ~0.03-0.04 mAP.</p>",
  "messages": [
    {
      "id": 1807094,
      "postDate": "2022-05-31T17:57:21.937Z",
      "content": "<p>Thanks to the sponsors and Kaggle for organizing this important research challenge.</p>\n<p><strong><em>TLDR</em></strong>: Solution architecture includes 5 model ensemble at two different image sizes with embedding dimension reduced by PCA.  Secret sauce is a newly developed augmentation method designed to deal with large occlusions while maintaining semantic cohesion of the overall scene.</p>\n<h2>Augmentation - Introducing BlendFlip:</h2>\n<p>The core problem in this challenge is dealing with large masked occlusions in test/query images.  For this my goal was to inpaint the masked region while attempting to:</p>\n<ul>\n<li>maintain semantic cohesion with the overall scene</li>\n<li>maintain continuity around masked edges</li>\n<li>fill in visually realistic content</li>\n</ul>\n<p>When looking at hotel room images, a common pattern is many vertical and horizontal features against repeating backgrounds.  For instance wall corners that run vertical, door frames that run vertical and horizontal, pictures on walls that are square, etc.  This observation gave me the idea to apply image patches from the same image directly adjacent to the masked region and essentially flip them into the masked region, and then blend the flips from the vertical and horizontal direction, hence the name BlendFlip.</p>\n<p><img src=\"https://i.imgur.com/q8ZAdPK.jpg\" alt=\"BlendFlip\"></p>\n<p>BlendFlip algorithm:</p>\n<ol>\n<li>Determine max distance between masked region edges in the vertical direction (image top to mask top, mask bottom to image bottom).  Do same for horizontal.  Max direction becomes the regions to be used for inpainting.</li>\n<li>Flip max vertical region into masked region.  If initial flip doesn’t fill entire masked region, keep flipping in vertical region until masked region is filled, cut off at edge of masked region.</li>\n<li>Repeat step 2 for horizontal region.</li>\n<li>Take 50/50 blend of step2 and step3 (cv2.addWeighted) as final inpainted region.</li>\n</ol>\n<p>BlendFlip helps address 2 1/2 of the 3 problems listed above.  It helps maintain semantic cohesion with visually realistic content.  It maintains continuity on 2 of the 4 sides of the masked edges.  I tried GAN's, traditional CV inpainting methods, variations of CutMix, but BlendFlip significantly outperformed them all.</p>\n<p><img src=\"https://i.imgur.com/sl6d9IK.jpg\" alt=\"Example 1\"><br>\n<img src=\"https://i.imgur.com/i18vrXt.jpg\" alt=\"Example 2\"><br>\n<img src=\"https://i.imgur.com/BxLju54.jpg\" alt=\"Example 3\"></p>\n<p>The goal of blendflip is NOT to predict what might be in the occluded region, but rather maintain the integrity of the image embedding from the non-occluded regions.</p>\n<h2>Solution architecture</h2>\n<p>The overall solution architecture is shown in the diagram below.  Images of sizes 1024x1024 (longest max side scaled to 1024, aspect ratio maintained) or 384x384 are fed into 5 models of various architecture.  Each model is trained with ArcFace loss and the embedding size of each is 1536D.  The embeddings are then concatenated and reduced to a total of 3072D via PCA (size determined by maintaining 99% of variance of embedding concat), and then KNN performed.  Each model is trained with a 50% probability of BlendFlip on each image.  BlendFlip is performed on all test images prior to inference.  No post processing or re-ranking was applied. </p>\n<p><img src=\"https://i.imgur.com/0vbOz3Z.jpg\" alt=\"Solution architecture\"></p>\n<p>For the CVPR paper I will include proper ablation studies and relative impact of the solution ingredients.  BlendFlip accounted for ~0.03-0.04 mAP.</p>",
      "rawMarkdown": "Thanks to the sponsors and Kaggle for organizing this important research challenge.\n\n***TLDR***: Solution architecture includes 5 model ensemble at two different image sizes with embedding dimension reduced by PCA.  Secret sauce is a newly developed augmentation method designed to deal with large occlusions while maintaining semantic cohesion of the overall scene.\n\n## Augmentation - Introducing BlendFlip:\nThe core problem in this challenge is dealing with large masked occlusions in test/query images.  For this my goal was to inpaint the masked region while attempting to:\n- maintain semantic cohesion with the overall scene\n- maintain continuity around masked edges\n- fill in visually realistic content\n\nWhen looking at hotel room images, a common pattern is many vertical and horizontal features against repeating backgrounds.  For instance wall corners that run vertical, door frames that run vertical and horizontal, pictures on walls that are square, etc.  This observation gave me the idea to apply image patches from the same image directly adjacent to the masked region and essentially flip them into the masked region, and then blend the flips from the vertical and horizontal direction, hence the name BlendFlip.\n\n![BlendFlip](https://i.imgur.com/q8ZAdPK.jpg)\n\nBlendFlip algorithm:\n1. Determine max distance between masked region edges in the vertical direction (image top to mask top, mask bottom to image bottom).  Do same for horizontal.  Max direction becomes the regions to be used for inpainting.\n2. Flip max vertical region into masked region.  If initial flip doesn’t fill entire masked region, keep flipping in vertical region until masked region is filled, cut off at edge of masked region.\n3. Repeat step 2 for horizontal region.\n4. Take 50/50 blend of step2 and step3 (cv2.addWeighted) as final inpainted region.\n\nBlendFlip helps address 2 1/2 of the 3 problems listed above.  It helps maintain semantic cohesion with visually realistic content.  It maintains continuity on 2 of the 4 sides of the masked edges.  I tried GAN's, traditional CV inpainting methods, variations of CutMix, but BlendFlip significantly outperformed them all.\n\n![Example 1](https://i.imgur.com/sl6d9IK.jpg)\n![Example 2](https://i.imgur.com/i18vrXt.jpg)\n![Example 3](https://i.imgur.com/BxLju54.jpg)\n\nThe goal of blendflip is NOT to predict what might be in the occluded region, but rather maintain the integrity of the image embedding from the non-occluded regions.\n\n## Solution architecture\nThe overall solution architecture is shown in the diagram below.  Images of sizes 1024x1024 (longest max side scaled to 1024, aspect ratio maintained) or 384x384 are fed into 5 models of various architecture.  Each model is trained with ArcFace loss and the embedding size of each is 1536D.  The embeddings are then concatenated and reduced to a total of 3072D via PCA (size determined by maintaining 99% of variance of embedding concat), and then KNN performed.  Each model is trained with a 50% probability of BlendFlip on each image.  BlendFlip is performed on all test images prior to inference.  No post processing or re-ranking was applied. \n\n![Solution architecture](https://i.imgur.com/0vbOz3Z.jpg)\n\nFor the CVPR paper I will include proper ablation studies and relative impact of the solution ingredients.  BlendFlip accounted for ~0.03-0.04 mAP.",
      "votes": 30
    },
    {
      "id": 1807918,
      "postDate": "2022-06-01T12:41:47.677Z",
      "content": "<p>Congrats to the 1st place, amazing results and nice solution! BlendFlip is very interesting idea, so glad to see some innovative approach for the occlusions, I believe it may become a thing how to deal with them in similar problems (I will definitely give it a try next time).</p>\n<p>Can't wait to see the relative impact of the solution ingredients. Would you mind sharing the effect of PCA here? I am pretty curious.</p>\n<p>I can't see in the description if you used any external data, does it mean you managed to get the results with only competition data? That would be really impressive.</p>",
      "rawMarkdown": "Congrats to the 1st place, amazing results and nice solution! BlendFlip is very interesting idea, so glad to see some innovative approach for the occlusions, I believe it may become a thing how to deal with them in similar problems (I will definitely give it a try next time).\n\nCan't wait to see the relative impact of the solution ingredients. Would you mind sharing the effect of PCA here? I am pretty curious.\n\nI can't see in the description if you used any external data, does it mean you managed to get the results with only competition data? That would be really impressive.",
      "votes": 1,
      "replies": [
        {
          "id": 1808228,
          "postDate": "2022-06-01T16:14:58.990Z",
          "content": "<blockquote>\n  <p>Would you mind sharing the effect of PCA here?</p>\n</blockquote>\n<p>Performing PCA gives about a 0.01 mAP improvement  (vs simple concat of embeddings) while reducing dimensionality at the same time.</p>\n<blockquote>\n  <p>I can't see in the description if you used any external data, does it mean you managed to get the results with only competition data?</p>\n</blockquote>\n<p>I used a very small amount of data from FGVC8, only from classes where &lt;8 samples existed in FGVC9.</p>",
          "rawMarkdown": "> Would you mind sharing the effect of PCA here?\n\nPerforming PCA gives about a 0.01 mAP improvement  (vs simple concat of embeddings) while reducing dimensionality at the same time.\n\n> I can't see in the description if you used any external data, does it mean you managed to get the results with only competition data?\n\nI used a very small amount of data from FGVC8, only from classes where <8 samples existed in FGVC9.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1807330,
      "postDate": "2022-05-31T23:27:45.093Z",
      "content": "<p>A huge Congrats David in such an important competition. Beautiful augmentation and Solution architecture.</p>",
      "rawMarkdown": "A huge Congrats David in such an important competition. Beautiful augmentation and Solution architecture.",
      "votes": 1
    },
    {
      "id": 1807822,
      "postDate": "2022-06-01T11:01:38.233Z",
      "content": "<p>Hi, I am a beginner. I just partly understand your solution. Will you public your code?</p>",
      "rawMarkdown": "Hi, I am a beginner. I just partly understand your solution. Will you public your code?",
      "votes": 2
    },
    {
      "id": 1807108,
      "postDate": "2022-05-31T18:10:45.823Z",
      "content": "<p>That's a really cool approach, thanks for sharing!</p>",
      "rawMarkdown": "That's a really cool approach, thanks for sharing!",
      "votes": 2
    },
    {
      "id": 1904823,
      "postDate": "2022-08-18T14:19:01.697Z",
      "content": "<p>Super cool!</p>",
      "rawMarkdown": "Super cool!"
    },
    {
      "id": 1808069,
      "postDate": "2022-06-01T14:13:14.510Z",
      "content": "<p>Brilliant <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>. Thanks for explaining.</p>",
      "rawMarkdown": "Brilliant @tivfrvqhs5. Thanks for explaining."
    },
    {
      "id": 1811206,
      "postDate": "2022-06-04T12:49:26.543Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1807918,
      "author_name": "Michal",
      "author_url": "",
      "post_date": "2022-06-01T12:41:47.677000",
      "content": "<p>Congrats to the 1st place, amazing results and nice solution! BlendFlip is very interesting idea, so glad to see some innovative approach for the occlusions, I believe it may become a thing how to deal with them in similar problems (I will definitely give it a try next time).</p>\n<p>Can't wait to see the relative impact of the solution ingredients. Would you mind sharing the effect of PCA here? I am pretty curious.</p>\n<p>I can't see in the description if you used any external data, does it mean you managed to get the results with only competition data? That would be really impressive.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1808228,
          "author_name": "David Austin",
          "author_url": "",
          "post_date": "2022-06-01T16:14:58.990000",
          "content": "<blockquote>\n  <p>Would you mind sharing the effect of PCA here?</p>\n</blockquote>\n<p>Performing PCA gives about a 0.01 mAP improvement  (vs simple concat of embeddings) while reducing dimensionality at the same time.</p>\n<blockquote>\n  <p>I can't see in the description if you used any external data, does it mean you managed to get the results with only competition data?</p>\n</blockquote>\n<p>I used a very small amount of data from FGVC8, only from classes where &lt;8 samples existed in FGVC9.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1807330,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-05-31T23:27:45.093000",
      "content": "<p>A huge Congrats David in such an important competition. Beautiful augmentation and Solution architecture.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1807822,
      "author_name": "Hien Tq",
      "author_url": "",
      "post_date": "2022-06-01T11:01:38.233000",
      "content": "<p>Hi, I am a beginner. I just partly understand your solution. Will you public your code?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1807108,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2022-05-31T18:10:45.823000",
      "content": "<p>That's a really cool approach, thanks for sharing!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1904823,
      "author_name": "bobam",
      "author_url": "",
      "post_date": "2022-08-18T14:19:01.697000",
      "content": "<p>Super cool!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1808069,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-06-01T14:13:14.510000",
      "content": "<p>Brilliant <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>. Thanks for explaining.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1811206,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T12:49:26.543000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1807094": "Thanks to the sponsors and Kaggle for organizing this important research challenge.\n\n***TLDR***: Solution architecture includes 5 model ensemble at two different image sizes with embedding dimension reduced by PCA.  Secret sauce is a newly developed augmentation method designed to deal with large occlusions while maintaining semantic cohesion of the overall scene.\n\n## Augmentation - Introducing BlendFlip:\nThe core problem in this challenge is dealing with large masked occlusions in test/query images.  For this my goal was to inpaint the masked region while attempting to:\n- maintain semantic cohesion with the overall scene\n- maintain continuity around masked edges\n- fill in visually realistic content\n\nWhen looking at hotel room images, a common pattern is many vertical and horizontal features against repeating backgrounds.  For instance wall corners that run vertical, door frames that run vertical and horizontal, pictures on walls that are square, etc.  This observation gave me the idea to apply image patches from the same image directly adjacent to the masked region and essentially flip them into the masked region, and then blend the flips from the vertical and horizontal direction, hence the name BlendFlip.\n\n![BlendFlip](https://i.imgur.com/q8ZAdPK.jpg)\n\nBlendFlip algorithm:\n1. Determine max distance between masked region edges in the vertical direction (image top to mask top, mask bottom to image bottom).  Do same for horizontal.  Max direction becomes the regions to be used for inpainting.\n2. Flip max vertical region into masked region.  If initial flip doesn’t fill entire masked region, keep flipping in vertical region until masked region is filled, cut off at edge of masked region.\n3. Repeat step 2 for horizontal region.\n4. Take 50/50 blend of step2 and step3 (cv2.addWeighted) as final inpainted region.\n\nBlendFlip helps address 2 1/2 of the 3 problems listed above.  It helps maintain semantic cohesion with visually realistic content.  It maintains continuity on 2 of the 4 sides of the masked edges.  I tried GAN's, traditional CV inpainting methods, variations of CutMix, but BlendFlip significantly outperformed them all.\n\n![Example 1](https://i.imgur.com/sl6d9IK.jpg)\n![Example 2](https://i.imgur.com/i18vrXt.jpg)\n![Example 3](https://i.imgur.com/BxLju54.jpg)\n\nThe goal of blendflip is NOT to predict what might be in the occluded region, but rather maintain the integrity of the image embedding from the non-occluded regions.\n\n## Solution architecture\nThe overall solution architecture is shown in the diagram below.  Images of sizes 1024x1024 (longest max side scaled to 1024, aspect ratio maintained) or 384x384 are fed into 5 models of various architecture.  Each model is trained with ArcFace loss and the embedding size of each is 1536D.  The embeddings are then concatenated and reduced to a total of 3072D via PCA (size determined by maintaining 99% of variance of embedding concat), and then KNN performed.  Each model is trained with a 50% probability of BlendFlip on each image.  BlendFlip is performed on all test images prior to inference.  No post processing or re-ranking was applied. \n\n![Solution architecture](https://i.imgur.com/0vbOz3Z.jpg)\n\nFor the CVPR paper I will include proper ablation studies and relative impact of the solution ingredients.  BlendFlip accounted for ~0.03-0.04 mAP.",
    "1807918": "Congrats to the 1st place, amazing results and nice solution! BlendFlip is very interesting idea, so glad to see some innovative approach for the occlusions, I believe it may become a thing how to deal with them in similar problems (I will definitely give it a try next time).\n\nCan't wait to see the relative impact of the solution ingredients. Would you mind sharing the effect of PCA here? I am pretty curious.\n\nI can't see in the description if you used any external data, does it mean you managed to get the results with only competition data? That would be really impressive.",
    "1807330": "A huge Congrats David in such an important competition. Beautiful augmentation and Solution architecture.",
    "1807822": "Hi, I am a beginner. I just partly understand your solution. Will you public your code?",
    "1807108": "That's a really cool approach, thanks for sharing!",
    "1904823": "Super cool!",
    "1808069": "Brilliant @tivfrvqhs5. Thanks for explaining.",
    "1811206": ""
  }
}