{
  "id": 110935,
  "title": "13th place solution summary",
  "url": "/competitions/open-images-2019-visual-relationship/writeups/appian-13th-place-solution-summary",
  "author_name": "",
  "post_date": "2019-10-02T11:31:36.577Z",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all. Conglatulations to the winners and thank you Google AI to host this interesting competition this year again.</p>\n\n<p>Like most of previous years solutions, I split the problem into two parts, <code>non-is</code> and <code>is</code> relationships as they have quite different characteristics.</p>\n\n<h3>1. non-is relationship</h3>\n\n<p>For this task, I focused on the relationship between two objects such as <code>Man on Chair</code>, <code>Cat under Table</code>. \nMy approach has two steps. Detect objects and then find out relationships for every possible triplets.</p>\n\n<p><strong>Object Detection</strong></p>\n\n<p>There are 57 objects such as <code>Man</code>, <code>Oven</code> to be a part of triplets. I used cascade-rcnn from <a href=\"https://github.com/open-mmlab/mmdetection/tree/master/mmdet\">mmdet</a> for these 57 classes with a bit of modification such as adding test time augumentations.</p>\n\n<p><a href=\"https://github.com/appian42/kaggle/blob/master/openimages/cascade-rcnn.conf\">This</a> is the .conf file I feeded to mmdet. Noticeble changes from default parametersz are </p>\n\n<ul>\n<li>2 more anchor boxes in RPN to capture high aspect ratio objects.</li>\n<li>CosineAnnealing instead of StepLR.</li>\n<li>score threshold 0.0001 for RCNN instead of 0.05.</li>\n<li>NMS treshold 0.4 instead of 0.5.</li>\n<li>max_per_image 400 instead of 100.</li>\n</ul>\n\n<p>I under-sampled frequent classes such as Man, Woman, Chair and used only 150,000 images to shorten the train time at the cost of accuracy.</p>\n\n<p><strong>triplet relationships</strong></p>\n\n<p>There are 287 triplet relationships.\nI took a similar approach to <a href=\"https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64630\">anokas's solution</a> last year as it's extremely simple and easy to implement.</p>\n\n<p>Some changes I made was</p>\n\n<ul>\n<li>Merged similar triplets into the same class based on some engineered features such as IOU, IOF so that rare triplets can be trained thanks to more frequent triplets.</li>\n<li>After merging, there are 90 classes out of 287 triplet relationships and 4-fold LightGBM models were separetely trained for each of these classes.</li>\n</ul>\n\n<p>Averaged AUC is 0.9641. After separating merged classes into original classes, averaged AUC is 0.9623.</p>\n\n<p><strong>Submission</strong></p>\n\n<ul>\n<li>0.31692 (public)</li>\n<li>0.24060 (private)</li>\n</ul>\n\n<h3>2. is relationship</h3>\n\n<p><strong>Object detection</strong></p>\n\n<p>I used cascade-rcnn to directly detect attributed objects such as <code>Table Wooden</code>, <code>Bench Plastic</code>. There are 42 classes for this task. This approach is similar to <a href=\"https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64642\">toshif's solution</a> last year but I failed to make the model as good as he did. </p>\n\n<p><strong>Submission</strong></p>\n\n<ul>\n<li>0.07346 (public)</li>\n<li>0.07130 (private)</li>\n</ul>\n\n<h3>3. submissions combined</h3>\n\n<p>I just put them together.</p>\n\n<ul>\n<li>0.38469 (public)</li>\n<li>0.30781 (private)</li>\n</ul>\n\n<h3>4. Possible improvements</h3>\n\n<ul>\n<li>Use full image size.</li>\n<li>Use all dataset.</li>\n<li>Use external dataset (such as COCO, Objects365 if it improves any).</li>\n<li>Train CNN model for triplet relationships just like <a href=\"https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64651\">tito's solution</a> last year. Some triplet relationships such as <code>Man holds Violin</code>, <code>Man plays Violin</code> are hard to differentiate because simple features such as IoU, IoF could not really capture the difference but CNN could.</li>\n</ul>\n\n<p>Thanks for reading!</p>",
  "messages": [
    {
      "id": "638736",
      "postDate": "10/02/2019 11:15:49",
      "content": "<p>Hi all. Conglatulations to the winners and thank you Google AI to host this interesting competition this year again.</p>\n\n<p>Like most of previous years solutions, I split the problem into two parts, <code>non-is</code> and <code>is</code> relationships as they have quite different characteristics.</p>\n\n<h3>1. non-is relationship</h3>\n\n<p>For this task, I focused on the relationship between two objects such as <code>Man on Chair</code>, <code>Cat under Table</code>. \nMy approach has two steps. Detect objects and then find out relationships for every possible triplets.</p>\n\n<p><strong>Object Detection</strong></p>\n\n<p>There are 57 objects such as <code>Man</code>, <code>Oven</code> to be a part of triplets. I used cascade-rcnn from <a href=\"https://github.com/open-mmlab/mmdetection/tree/master/mmdet\">mmdet</a> for these 57 classes with a bit of modification such as adding test time augumentations.</p>\n\n<p><a href=\"https://github.com/appian42/kaggle/blob/master/openimages/cascade-rcnn.conf\">This</a> is the .conf file I feeded to mmdet. Noticeble changes from default parametersz are </p>\n\n<ul>\n<li>2 more anchor boxes in RPN to capture high aspect ratio objects.</li>\n<li>CosineAnnealing instead of StepLR.</li>\n<li>score threshold 0.0001 for RCNN instead of 0.05.</li>\n<li>NMS treshold 0.4 instead of 0.5.</li>\n<li>max_per_image 400 instead of 100.</li>\n</ul>\n\n<p>I under-sampled frequent classes such as Man, Woman, Chair and used only 150,000 images to shorten the train time at the cost of accuracy.</p>\n\n<p><strong>triplet relationships</strong></p>\n\n<p>There are 287 triplet relationships.\nI took a similar approach to <a href=\"https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64630\">anokas's solution</a> last year as it's extremely simple and easy to implement.</p>\n\n<p>Some changes I made was</p>\n\n<ul>\n<li>Merged similar triplets into the same class based on some engineered features such as IOU, IOF so that rare triplets can be trained thanks to more frequent triplets.</li>\n<li>After merging, there are 90 classes out of 287 triplet relationships and 4-fold LightGBM models were separetely trained for each of these classes.</li>\n</ul>\n\n<p>Averaged AUC is 0.9641. After separating merged classes into original classes, averaged AUC is 0.9623.</p>\n\n<p><strong>Submission</strong></p>\n\n<ul>\n<li>0.31692 (public)</li>\n<li>0.24060 (private)</li>\n</ul>\n\n<h3>2. is relationship</h3>\n\n<p><strong>Object detection</strong></p>\n\n<p>I used cascade-rcnn to directly detect attributed objects such as <code>Table Wooden</code>, <code>Bench Plastic</code>. There are 42 classes for this task. This approach is similar to <a href=\"https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64642\">toshif's solution</a> last year but I failed to make the model as good as he did. </p>\n\n<p><strong>Submission</strong></p>\n\n<ul>\n<li>0.07346 (public)</li>\n<li>0.07130 (private)</li>\n</ul>\n\n<h3>3. submissions combined</h3>\n\n<p>I just put them together.</p>\n\n<ul>\n<li>0.38469 (public)</li>\n<li>0.30781 (private)</li>\n</ul>\n\n<h3>4. Possible improvements</h3>\n\n<ul>\n<li>Use full image size.</li>\n<li>Use all dataset.</li>\n<li>Use external dataset (such as COCO, Objects365 if it improves any).</li>\n<li>Train CNN model for triplet relationships just like <a href=\"https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64651\">tito's solution</a> last year. Some triplet relationships such as <code>Man holds Violin</code>, <code>Man plays Violin</code> are hard to differentiate because simple features such as IoU, IoF could not really capture the difference but CNN could.</li>\n</ul>\n\n<p>Thanks for reading!</p>",
      "rawMarkdown": "Hi all. Conglatulations to the winners and thank you Google AI to host this interesting competition this year again.\n\nLike most of previous years solutions, I split the problem into two parts, `non-is` and `is` relationships as they have quite different characteristics.\n\n### 1. non-is relationship\n\nFor this task, I focused on the relationship between two objects such as `Man on Chair`, `Cat under Table`. \nMy approach has two steps. Detect objects and then find out relationships for every possible triplets.\n\n**Object Detection**\n\nThere are 57 objects such as `Man`, `Oven` to be a part of triplets. I used cascade-rcnn from [mmdet](https://github.com/open-mmlab/mmdetection/tree/master/mmdet) for these 57 classes with a bit of modification such as adding test time augumentations.\n\n[This](https://github.com/appian42/kaggle/blob/master/openimages/cascade-rcnn.conf) is the .conf file I feeded to mmdet. Noticeble changes from default parametersz are \n\n- 2 more anchor boxes in RPN to capture high aspect ratio objects.\n- CosineAnnealing instead of StepLR.\n- score threshold 0.0001 for RCNN instead of 0.05.\n- NMS treshold 0.4 instead of 0.5.\n- max\\_per\\_image 400 instead of 100.\n\nI under-sampled frequent classes such as Man, Woman, Chair and used only 150,000 images to shorten the train time at the cost of accuracy.\n\n**triplet relationships**\n\nThere are 287 triplet relationships.\nI took a similar approach to [anokas's solution](https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64630) last year as it's extremely simple and easy to implement.\n\nSome changes I made was\n\n- Merged similar triplets into the same class based on some engineered features such as IOU, IOF so that rare triplets can be trained thanks to more frequent triplets.\n- After merging, there are 90 classes out of 287 triplet relationships and 4-fold LightGBM models were separetely trained for each of these classes.\n\nAveraged AUC is 0.9641. After separating merged classes into original classes, averaged AUC is 0.9623.\n\n\n**Submission**\n\n- 0.31692 (public)\n- 0.24060 (private)\n\n\n### 2. is relationship\n\n**Object detection**\n\nI used cascade-rcnn to directly detect attributed objects such as `Table Wooden`, `Bench Plastic`. There are 42 classes for this task. This approach is similar to [toshif's solution](https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64642) last year but I failed to make the model as good as he did. \n\n**Submission**\n\n- 0.07346 (public)\n- 0.07130 (private)\n\n\n### 3. submissions combined\n\nI just put them together.\n\n- 0.38469 (public)\n- 0.30781 (private)\n\n\n### 4. Possible improvements\n\n- Use full image size.\n- Use all dataset.\n- Use external dataset (such as COCO, Objects365 if it improves any).\n- Train CNN model for triplet relationships just like [tito's solution](https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64651) last year. Some triplet relationships such as `Man holds Violin`, `Man plays Violin` are hard to differentiate because simple features such as IoU, IoF could not really capture the difference but CNN could.\n\n\nThanks for reading!",
      "votes": null
    },
    {
      "id": "638856",
      "postDate": "10/02/2019 13:49:05",
      "content": "<p>Shame about your LB shakeup, congrats anyway!</p>",
      "rawMarkdown": "Shame about your LB shakeup, congrats anyway!",
      "votes": null
    },
    {
      "id": "638881",
      "postDate": "10/02/2019 14:23:17",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights… <a href=\"/appian\">@appian</a> </p>",
      "rawMarkdown": "Congrats\nThank you for Sharing your Approach &amp; Insights… @appian",
      "votes": null
    },
    {
      "id": "638900",
      "postDate": "10/02/2019 14:35:50",
      "content": "<p>Congratulations! :-). do you have any plan to open-source your code solutions?</p>",
      "rawMarkdown": "Congratulations! :-). do you have any plan to open-source your code solutions?",
      "votes": null
    },
    {
      "id": "640229",
      "postDate": "10/04/2019 00:12:21",
      "content": "<p>Congratulations and thank you for sharing approach. \nIt is interesting for me to know the approach for Visual Relationship tasks!</p>",
      "rawMarkdown": "Congratulations and thank you for sharing approach. \nIt is interesting for me to know the approach for Visual Relationship tasks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 638856,
      "author_name": "artyomp",
      "author_url": "",
      "post_date": "10/02/2019 13:49:05",
      "content": "<p>Shame about your LB shakeup, congrats anyway!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638881,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "10/02/2019 14:23:17",
      "content": "<p>Congrats\nThank you for Sharing your Approach &amp; Insights… <a href=\"/appian\">@appian</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 638900,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "10/02/2019 14:35:50",
      "content": "<p>Congratulations! :-). do you have any plan to open-source your code solutions?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 640229,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "10/04/2019 00:12:21",
      "content": "<p>Congratulations and thank you for sharing approach. \nIt is interesting for me to know the approach for Visual Relationship tasks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "638736": "Hi all. Conglatulations to the winners and thank you Google AI to host this interesting competition this year again.\n\nLike most of previous years solutions, I split the problem into two parts, `non-is` and `is` relationships as they have quite different characteristics.\n\n### 1. non-is relationship\n\nFor this task, I focused on the relationship between two objects such as `Man on Chair`, `Cat under Table`. \nMy approach has two steps. Detect objects and then find out relationships for every possible triplets.\n\n**Object Detection**\n\nThere are 57 objects such as `Man`, `Oven` to be a part of triplets. I used cascade-rcnn from [mmdet](https://github.com/open-mmlab/mmdetection/tree/master/mmdet) for these 57 classes with a bit of modification such as adding test time augumentations.\n\n[This](https://github.com/appian42/kaggle/blob/master/openimages/cascade-rcnn.conf) is the .conf file I feeded to mmdet. Noticeble changes from default parametersz are \n\n- 2 more anchor boxes in RPN to capture high aspect ratio objects.\n- CosineAnnealing instead of StepLR.\n- score threshold 0.0001 for RCNN instead of 0.05.\n- NMS treshold 0.4 instead of 0.5.\n- max\\_per\\_image 400 instead of 100.\n\nI under-sampled frequent classes such as Man, Woman, Chair and used only 150,000 images to shorten the train time at the cost of accuracy.\n\n**triplet relationships**\n\nThere are 287 triplet relationships.\nI took a similar approach to [anokas's solution](https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64630) last year as it's extremely simple and easy to implement.\n\nSome changes I made was\n\n- Merged similar triplets into the same class based on some engineered features such as IOU, IOF so that rare triplets can be trained thanks to more frequent triplets.\n- After merging, there are 90 classes out of 287 triplet relationships and 4-fold LightGBM models were separetely trained for each of these classes.\n\nAveraged AUC is 0.9641. After separating merged classes into original classes, averaged AUC is 0.9623.\n\n\n**Submission**\n\n- 0.31692 (public)\n- 0.24060 (private)\n\n\n### 2. is relationship\n\n**Object detection**\n\nI used cascade-rcnn to directly detect attributed objects such as `Table Wooden`, `Bench Plastic`. There are 42 classes for this task. This approach is similar to [toshif's solution](https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64642) last year but I failed to make the model as good as he did. \n\n**Submission**\n\n- 0.07346 (public)\n- 0.07130 (private)\n\n\n### 3. submissions combined\n\nI just put them together.\n\n- 0.38469 (public)\n- 0.30781 (private)\n\n\n### 4. Possible improvements\n\n- Use full image size.\n- Use all dataset.\n- Use external dataset (such as COCO, Objects365 if it improves any).\n- Train CNN model for triplet relationships just like [tito's solution](https://www.kaggle.com/c/google-ai-open-images-visual-relationship-track/discussion/64651) last year. Some triplet relationships such as `Man holds Violin`, `Man plays Violin` are hard to differentiate because simple features such as IoU, IoF could not really capture the difference but CNN could.\n\n\nThanks for reading!",
    "638856": "Shame about your LB shakeup, congrats anyway!",
    "638881": "Congrats\nThank you for Sharing your Approach &amp; Insights… @appian",
    "638900": "Congratulations! :-). do you have any plan to open-source your code solutions?",
    "640229": "Congratulations and thank you for sharing approach. \nIt is interesting for me to know the approach for Visual Relationship tasks!"
  },
  "source": "meta"
}