{
  "id": 64630,
  "title": "7th Place Solution Summary - anokas",
  "url": "/competitions/google-ai-open-images-visual-relationship-track/discussion/64630",
  "author_name": "anokas",
  "post_date": "2018-08-31T00:06:32.965000",
  "votes": 39,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Congratulations to the winners and thanks to Google AI for hosting this fun and interesting competition. My solution was relatively simple, taking a submission file from the object detection competition as an input and building some heuristic models based off bounding boxes.</p>\n\n<p><strong>Object Detection</strong></p>\n\n<p>For object detection bounding boxes I used Google’s <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\">pre-trained Faster R-CNN model</a> trained on V2 of the Open Images dataset - the labels they were trained on were actually a superset of the labels in this dataset so it was just a case of mapping the labels and filtering out ones we weren’t interested in. After expanding the label hierarchy, this model scored about 0.37 in the object detection competition. I also predicted on flipped images and combined the predictions, netting me <strong>0.375</strong> overall.</p>\n\n<p>One thing that led to a huge increase in score in object detection was to not threshold the predictions and leave the low confidence predictions in the submission file. Because of the way average precision works, you cannot be penalised for adding additional false positives with a lower confidence than all your other predictions, however you can still improve your recall if you find additional objects that weren’t previously detected.</p>\n\n<p><strong>Visual Relationships</strong></p>\n\n<p>This challenge can actually be split up into two subproblems. The first subproblem involves the <code>is</code> label (for example <code>chair is wooden</code>), and in fact only involves a single object and an attribute. This subproblem spans 57 relationships across 5 different descriptive attributes.</p>\n\n<p>The second subproblem is indeed based around visual relationships, and involves a pair of objects as well as a proverb, such as <code>chair at table</code> - there are about 250 different triplets that occur within the dataset.</p>\n\n<p>While the two types of problems have about equal occurences in the dataset, the triplet relationship detection has about 5x more distinct relationships - meaning it has about 5x more weight on the LB score, due to 80% of the metric being the mean of relationship-wise AP.</p>\n\n<p><strong>Attribute Classification</strong></p>\n\n<p>For this subproblem, I built a training dataset by getting the bounding boxes of all objects which could have a descriptive attribute from the bbox ground truth file. I then matched these objects with the boxes in the relationships ground truth file, which gave me a ground truth (for each object example, which attributes are present).</p>\n\n<p>On this dataset, I then trained a single pre-trained DenseNet121 network on the cropped objects resized to 224x224, to predict the probability of all 5 classes - this achieved about 0.95-0.99 AUC on all the classes.</p>\n\n<p>The pre-trained model is then used to predict the probability of the 5 attributes for all the applicable bounding boxes in the object detection challenge, with the final predicted probability as follows:</p>\n\n<pre><code>P(chair is wooden) = P(bounding box is chair) * P(bounding box is wooden)\n</code></pre>\n\n<p>Where the <code>bounding box is chair</code> probability is the output of the object detection model.</p>\n\n<p><strong>Triplet Relationships</strong></p>\n\n<p>For the visual relationship triplets, I also built a dataset similarly to attribute classification - all potential pairs of bounding boxes of the two objects, and whether they have the given relationship.</p>\n\n<p>This gave me around 250 training sets in total for the different triplets (of varying size).\nFor the first 100 relationships, I built a 5-fold XGBoost model on each one with the following features:</p>\n\n<ul>\n<li>% box 1 inside box 2, % box 2 inside box 1</li>\n<li>IoU (intersection over union between the boxes)</li>\n<li>Horizontal/vertical offset of the box centres</li>\n<li>Euclidean distance between the two box centres</li>\n<li>Euclidean distance normalised by the size of the boxes (zoom invariance)</li>\n</ul>\n\n<p>And then using the models, the submitted probability of each relationship is for example:</p>\n\n<pre><code>P(chair at table) = P(bounding box is chair) * P(bounding box is table) * P(chair at table XGBoost)\n</code></pre>\n\n<p>It seemed that building a separate model for each relationship worked well, as I noticed they each had very different feature importances (eg. sometimes the model was looking at the overlap, while for others the model was looking at the difference in height between the objects). Most classes also had &gt;0.95 AUC, showing the XGBoost model was very good at classifying whether a relationship existed.</p>\n\n<p>For the other relationships (with a few hundred samples or less) I replaced the xgboost model with a simple prior - what proportion of pair occurences had the relationship in the training data. I tried using eg. linear models for these but didn’t see an improvement.</p>\n\n<p>Overall, the solution runs from start to finish in under 24 hours on an i7+single GPU - I am happy with the performance given the simplicity. I observed a linear increase in performance in visual relationships given an improvement in the bounding boxes scores, so this solution could have potentially scored much higher if I had a better object detector! :)</p>\n\n<p>I am curious to see how other competitors approached the problem, as this is quite a new problem for Kaggle.</p>\n\n<p>- anokas</p>",
  "messages": [
    {
      "id": 379148,
      "postDate": "2018-08-31T00:06:32.967Z",
      "content": "<p>Congratulations to the winners and thanks to Google AI for hosting this fun and interesting competition. My solution was relatively simple, taking a submission file from the object detection competition as an input and building some heuristic models based off bounding boxes.</p>\n\n<p><strong>Object Detection</strong></p>\n\n<p>For object detection bounding boxes I used Google’s <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\">pre-trained Faster R-CNN model</a> trained on V2 of the Open Images dataset - the labels they were trained on were actually a superset of the labels in this dataset so it was just a case of mapping the labels and filtering out ones we weren’t interested in. After expanding the label hierarchy, this model scored about 0.37 in the object detection competition. I also predicted on flipped images and combined the predictions, netting me <strong>0.375</strong> overall.</p>\n\n<p>One thing that led to a huge increase in score in object detection was to not threshold the predictions and leave the low confidence predictions in the submission file. Because of the way average precision works, you cannot be penalised for adding additional false positives with a lower confidence than all your other predictions, however you can still improve your recall if you find additional objects that weren’t previously detected.</p>\n\n<p><strong>Visual Relationships</strong></p>\n\n<p>This challenge can actually be split up into two subproblems. The first subproblem involves the <code>is</code> label (for example <code>chair is wooden</code>), and in fact only involves a single object and an attribute. This subproblem spans 57 relationships across 5 different descriptive attributes.</p>\n\n<p>The second subproblem is indeed based around visual relationships, and involves a pair of objects as well as a proverb, such as <code>chair at table</code> - there are about 250 different triplets that occur within the dataset.</p>\n\n<p>While the two types of problems have about equal occurences in the dataset, the triplet relationship detection has about 5x more distinct relationships - meaning it has about 5x more weight on the LB score, due to 80% of the metric being the mean of relationship-wise AP.</p>\n\n<p><strong>Attribute Classification</strong></p>\n\n<p>For this subproblem, I built a training dataset by getting the bounding boxes of all objects which could have a descriptive attribute from the bbox ground truth file. I then matched these objects with the boxes in the relationships ground truth file, which gave me a ground truth (for each object example, which attributes are present).</p>\n\n<p>On this dataset, I then trained a single pre-trained DenseNet121 network on the cropped objects resized to 224x224, to predict the probability of all 5 classes - this achieved about 0.95-0.99 AUC on all the classes.</p>\n\n<p>The pre-trained model is then used to predict the probability of the 5 attributes for all the applicable bounding boxes in the object detection challenge, with the final predicted probability as follows:</p>\n\n<pre><code>P(chair is wooden) = P(bounding box is chair) * P(bounding box is wooden)\n</code></pre>\n\n<p>Where the <code>bounding box is chair</code> probability is the output of the object detection model.</p>\n\n<p><strong>Triplet Relationships</strong></p>\n\n<p>For the visual relationship triplets, I also built a dataset similarly to attribute classification - all potential pairs of bounding boxes of the two objects, and whether they have the given relationship.</p>\n\n<p>This gave me around 250 training sets in total for the different triplets (of varying size).\nFor the first 100 relationships, I built a 5-fold XGBoost model on each one with the following features:</p>\n\n<ul>\n<li>% box 1 inside box 2, % box 2 inside box 1</li>\n<li>IoU (intersection over union between the boxes)</li>\n<li>Horizontal/vertical offset of the box centres</li>\n<li>Euclidean distance between the two box centres</li>\n<li>Euclidean distance normalised by the size of the boxes (zoom invariance)</li>\n</ul>\n\n<p>And then using the models, the submitted probability of each relationship is for example:</p>\n\n<pre><code>P(chair at table) = P(bounding box is chair) * P(bounding box is table) * P(chair at table XGBoost)\n</code></pre>\n\n<p>It seemed that building a separate model for each relationship worked well, as I noticed they each had very different feature importances (eg. sometimes the model was looking at the overlap, while for others the model was looking at the difference in height between the objects). Most classes also had &gt;0.95 AUC, showing the XGBoost model was very good at classifying whether a relationship existed.</p>\n\n<p>For the other relationships (with a few hundred samples or less) I replaced the xgboost model with a simple prior - what proportion of pair occurences had the relationship in the training data. I tried using eg. linear models for these but didn’t see an improvement.</p>\n\n<p>Overall, the solution runs from start to finish in under 24 hours on an i7+single GPU - I am happy with the performance given the simplicity. I observed a linear increase in performance in visual relationships given an improvement in the bounding boxes scores, so this solution could have potentially scored much higher if I had a better object detector! :)</p>\n\n<p>I am curious to see how other competitors approached the problem, as this is quite a new problem for Kaggle.</p>\n\n<p>- anokas</p>",
      "rawMarkdown": "Congratulations to the winners and thanks to Google AI for hosting this fun and interesting competition. My solution was relatively simple, taking a submission file from the object detection competition as an input and building some heuristic models based off bounding boxes.\n\n**Object Detection**\n\nFor object detection bounding boxes I used Google’s [pre-trained Faster R-CNN model](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md) trained on V2 of the Open Images dataset - the labels they were trained on were actually a superset of the labels in this dataset so it was just a case of mapping the labels and filtering out ones we weren’t interested in. After expanding the label hierarchy, this model scored about 0.37 in the object detection competition. I also predicted on flipped images and combined the predictions, netting me **0.375** overall.\n\nOne thing that led to a huge increase in score in object detection was to not threshold the predictions and leave the low confidence predictions in the submission file. Because of the way average precision works, you cannot be penalised for adding additional false positives with a lower confidence than all your other predictions, however you can still improve your recall if you find additional objects that weren’t previously detected.\n\n**Visual Relationships**\n\nThis challenge can actually be split up into two subproblems. The first subproblem involves the `is` label (for example `chair is wooden`), and in fact only involves a single object and an attribute. This subproblem spans 57 relationships across 5 different descriptive attributes.\n\nThe second subproblem is indeed based around visual relationships, and involves a pair of objects as well as a proverb, such as `chair at table` - there are about 250 different triplets that occur within the dataset.\n\nWhile the two types of problems have about equal occurences in the dataset, the triplet relationship detection has about 5x more distinct relationships - meaning it has about 5x more weight on the LB score, due to 80% of the metric being the mean of relationship-wise AP.\n\n**Attribute Classification**\n\nFor this subproblem, I built a training dataset by getting the bounding boxes of all objects which could have a descriptive attribute from the bbox ground truth file. I then matched these objects with the boxes in the relationships ground truth file, which gave me a ground truth (for each object example, which attributes are present).\n\nOn this dataset, I then trained a single pre-trained DenseNet121 network on the cropped objects resized to 224x224, to predict the probability of all 5 classes - this achieved about 0.95-0.99 AUC on all the classes.\n\nThe pre-trained model is then used to predict the probability of the 5 attributes for all the applicable bounding boxes in the object detection challenge, with the final predicted probability as follows:\n\n    P(chair is wooden) = P(bounding box is chair) * P(bounding box is wooden)\n\nWhere the `bounding box is chair` probability is the output of the object detection model.\n\n**Triplet Relationships**\n\nFor the visual relationship triplets, I also built a dataset similarly to attribute classification - all potential pairs of bounding boxes of the two objects, and whether they have the given relationship.\n\nThis gave me around 250 training sets in total for the different triplets (of varying size).\nFor the first 100 relationships, I built a 5-fold XGBoost model on each one with the following features:\n\n- % box 1 inside box 2, % box 2 inside box 1\n- IoU (intersection over union between the boxes)\n- Horizontal/vertical offset of the box centres\n- Euclidean distance between the two box centres\n- Euclidean distance normalised by the size of the boxes (zoom invariance)\n\nAnd then using the models, the submitted probability of each relationship is for example:\n\n    P(chair at table) = P(bounding box is chair) * P(bounding box is table) * P(chair at table XGBoost)\n\nIt seemed that building a separate model for each relationship worked well, as I noticed they each had very different feature importances (eg. sometimes the model was looking at the overlap, while for others the model was looking at the difference in height between the objects). Most classes also had &gt;0.95 AUC, showing the XGBoost model was very good at classifying whether a relationship existed.\n\nFor the other relationships (with a few hundred samples or less) I replaced the xgboost model with a simple prior - what proportion of pair occurences had the relationship in the training data. I tried using eg. linear models for these but didn’t see an improvement.\n\nOverall, the solution runs from start to finish in under 24 hours on an i7+single GPU - I am happy with the performance given the simplicity. I observed a linear increase in performance in visual relationships given an improvement in the bounding boxes scores, so this solution could have potentially scored much higher if I had a better object detector! :)\n\nI am curious to see how other competitors approached the problem, as this is quite a new problem for Kaggle.\n\n\\- anokas\n",
      "votes": 38
    },
    {
      "id": 379158,
      "postDate": "2018-08-31T00:31:48.837Z",
      "content": "<p>Congratulations, Youngest Grandmaster! 😎</p>",
      "rawMarkdown": "Congratulations, Youngest Grandmaster! 😎",
      "votes": 1
    },
    {
      "id": 584276,
      "postDate": "2019-07-25T16:43:04.887Z",
      "content": "<p><a href=\"/anokas\">@anokas</a> would you be publishing your code anytime? </p>",
      "rawMarkdown": "@anokas would you be publishing your code anytime? "
    },
    {
      "id": 543737,
      "postDate": "2019-06-04T18:36:30.073Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 379508,
      "postDate": "2018-08-31T13:33:54.007Z",
      "content": "<p>Grandmaster status !!\nCongrats since you never picked the traditional ML competition.  </p>",
      "rawMarkdown": "Grandmaster status !!\nCongrats since you never picked the traditional ML competition.  "
    },
    {
      "id": 379486,
      "postDate": "2018-08-31T12:54:28.020Z",
      "content": "<p>Congrats <a href=\"/anokas\">@anokas</a>! Very interesting solution.</p>\n\n<p>Will also try to do a write up of mine over the next day or two. Like yourself I am also very much looking forward to other write ups - seems we have all taken very differing approaches.</p>",
      "rawMarkdown": "Congrats @anokas! Very interesting solution.\n\nWill also try to do a write up of mine over the next day or two. Like yourself I am also very much looking forward to other write ups - seems we have all taken very differing approaches."
    },
    {
      "id": 379457,
      "postDate": "2018-08-31T11:34:59.977Z",
      "content": "<p>Congratulations <a href=\"/anokas\">@anokas</a>! Grandmaster is coming :)</p>",
      "rawMarkdown": "Congratulations @anokas! Grandmaster is coming :)"
    },
    {
      "id": 379157,
      "postDate": "2018-08-31T00:31:28.753Z",
      "content": "<p>Congratulations! I started this competition pretty late, so I didn't have time to prepare an object detector that is good enough for this competition......</p>",
      "rawMarkdown": "Congratulations! I started this competition pretty late, so I didn't have time to prepare an object detector that is good enough for this competition......",
      "replies": [
        {
          "id": 379160,
          "postDate": "2018-08-31T00:34:24.853Z",
          "content": "<p>Thanks Alexander. I guess you missed Google's pre-trained model then (like it seems many did), as it basically worked out of the box for this competition - I only spent a few hours working on object detection. The fun stuff was the rest of it! :D</p>",
          "rawMarkdown": "Thanks Alexander. I guess you missed Google's pre-trained model then (like it seems many did), as it basically worked out of the box for this competition - I only spent a few hours working on object detection. The fun stuff was the rest of it! :D",
          "votes": 2
        },
        {
          "id": 379166,
          "postDate": "2018-08-31T00:43:07.363Z",
          "content": "<p>I tried to re-train the model and it failed miserably. So I used the out-of-shelf version of the Faster-RCNN v2 atrous, which gave me 0.23 public LB. I guess that I didn't have time to try out other tricks... But not thresholding, flipping image, and re-mapping, that was really clever! Guess that I will try those out next time...</p>",
          "rawMarkdown": "I tried to re-train the model and it failed miserably. So I used the out-of-shelf version of the Faster-RCNN v2 atrous, which gave me 0.23 public LB. I guess that I didn't have time to try out other tricks... But not thresholding, flipping image, and re-mapping, that was really clever! Guess that I will try those out next time..."
        },
        {
          "id": 379168,
          "postDate": "2018-08-31T00:44:05.847Z",
          "content": "<p>Removing the thresholding is what gave me the boost from 0.27 -&gt; 0.37 :)</p>",
          "rawMarkdown": "Removing the thresholding is what gave me the boost from 0.27 -&gt; 0.37 :)"
        }
      ]
    },
    {
      "id": 379152,
      "postDate": "2018-08-31T00:13:41.960Z",
      "content": "<p>congratulations on competitions grandmaster status!</p>",
      "rawMarkdown": "congratulations on competitions grandmaster status!",
      "replies": [
        {
          "id": 379153,
          "postDate": "2018-08-31T00:14:33.383Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        },
        {
          "id": 379165,
          "postDate": "2018-08-31T00:43:03.893Z",
          "content": "<p>Just to verify, the relationship prediction did not actually use any CNNs, except for classifying the bounding boxes.</p>",
          "rawMarkdown": "Just to verify, the relationship prediction did not actually use any CNNs, except for classifying the bounding boxes."
        },
        {
          "id": 379169,
          "postDate": "2018-08-31T00:45:30.610Z",
          "content": "<p>This is true for the triplet relationships. For the attribute relationships (<code>chair is wooden</code>, <code>bottle is plastic</code>, etc) I trained a single CNN multi-label classifier.</p>",
          "rawMarkdown": "This is true for the triplet relationships. For the attribute relationships (`chair is wooden`, `bottle is plastic`, etc) I trained a single CNN multi-label classifier."
        },
        {
          "id": 379171,
          "postDate": "2018-08-31T00:55:11.307Z",
          "content": "<blockquote>\n  <p>This gave me around 250 training sets in total for the different triplets (of varying size)</p>\n</blockquote>\n\n<p>So the target variable for the triplet relationship detection is one of the 250 unique triplets?</p>",
          "rawMarkdown": "&gt; This gave me around 250 training sets in total for the different triplets (of varying size)\n\nSo the target variable for the triplet relationship detection is one of the 250 unique triplets?"
        }
      ]
    },
    {
      "id": 565561,
      "postDate": "2019-07-01T05:42:47.190Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 379158,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2018-08-31T00:31:48.837000",
      "content": "<p>Congratulations, Youngest Grandmaster! 😎</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 584276,
      "author_name": "Rohit Midha",
      "author_url": "",
      "post_date": "2019-07-25T16:43:04.887000",
      "content": "<p><a href=\"/anokas\">@anokas</a> would you be publishing your code anytime? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543737,
      "author_name": "Karan Sindwani",
      "author_url": "",
      "post_date": "2019-06-04T18:36:30.073000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 379508,
      "author_name": "eagle4",
      "author_url": "",
      "post_date": "2018-08-31T13:33:54.007000",
      "content": "<p>Grandmaster status !!\nCongrats since you never picked the traditional ML competition.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 379486,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2018-08-31T12:54:28.020000",
      "content": "<p>Congrats <a href=\"/anokas\">@anokas</a>! Very interesting solution.</p>\n\n<p>Will also try to do a write up of mine over the next day or two. Like yourself I am also very much looking forward to other write ups - seems we have all taken very differing approaches.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 379457,
      "author_name": "Meyk",
      "author_url": "",
      "post_date": "2018-08-31T11:34:59.977000",
      "content": "<p>Congratulations <a href=\"/anokas\">@anokas</a>! Grandmaster is coming :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 379157,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2018-08-31T00:31:28.753000",
      "content": "<p>Congratulations! I started this competition pretty late, so I didn't have time to prepare an object detector that is good enough for this competition......</p>",
      "votes": 0,
      "replies": [
        {
          "id": 379160,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2018-08-31T00:34:24.853000",
          "content": "<p>Thanks Alexander. I guess you missed Google's pre-trained model then (like it seems many did), as it basically worked out of the box for this competition - I only spent a few hours working on object detection. The fun stuff was the rest of it! :D</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 379166,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2018-08-31T00:43:07.363000",
          "content": "<p>I tried to re-train the model and it failed miserably. So I used the out-of-shelf version of the Faster-RCNN v2 atrous, which gave me 0.23 public LB. I guess that I didn't have time to try out other tricks... But not thresholding, flipping image, and re-mapping, that was really clever! Guess that I will try those out next time...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 379168,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2018-08-31T00:44:05.847000",
          "content": "<p>Removing the thresholding is what gave me the boost from 0.27 -&gt; 0.37 :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 379152,
      "author_name": "Master",
      "author_url": "",
      "post_date": "2018-08-31T00:13:41.960000",
      "content": "<p>congratulations on competitions grandmaster status!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 379153,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2018-08-31T00:14:33.383000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 379165,
          "author_name": "Master",
          "author_url": "",
          "post_date": "2018-08-31T00:43:03.893000",
          "content": "<p>Just to verify, the relationship prediction did not actually use any CNNs, except for classifying the bounding boxes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 379169,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2018-08-31T00:45:30.610000",
          "content": "<p>This is true for the triplet relationships. For the attribute relationships (<code>chair is wooden</code>, <code>bottle is plastic</code>, etc) I trained a single CNN multi-label classifier.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 379171,
          "author_name": "Master",
          "author_url": "",
          "post_date": "2018-08-31T00:55:11.307000",
          "content": "<blockquote>\n  <p>This gave me around 250 training sets in total for the different triplets (of varying size)</p>\n</blockquote>\n\n<p>So the target variable for the triplet relationship detection is one of the 250 unique triplets?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 565561,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-01T05:42:47.190000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "379148": "Congratulations to the winners and thanks to Google AI for hosting this fun and interesting competition. My solution was relatively simple, taking a submission file from the object detection competition as an input and building some heuristic models based off bounding boxes.\n\n**Object Detection**\n\nFor object detection bounding boxes I used Google’s [pre-trained Faster R-CNN model](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md) trained on V2 of the Open Images dataset - the labels they were trained on were actually a superset of the labels in this dataset so it was just a case of mapping the labels and filtering out ones we weren’t interested in. After expanding the label hierarchy, this model scored about 0.37 in the object detection competition. I also predicted on flipped images and combined the predictions, netting me **0.375** overall.\n\nOne thing that led to a huge increase in score in object detection was to not threshold the predictions and leave the low confidence predictions in the submission file. Because of the way average precision works, you cannot be penalised for adding additional false positives with a lower confidence than all your other predictions, however you can still improve your recall if you find additional objects that weren’t previously detected.\n\n**Visual Relationships**\n\nThis challenge can actually be split up into two subproblems. The first subproblem involves the `is` label (for example `chair is wooden`), and in fact only involves a single object and an attribute. This subproblem spans 57 relationships across 5 different descriptive attributes.\n\nThe second subproblem is indeed based around visual relationships, and involves a pair of objects as well as a proverb, such as `chair at table` - there are about 250 different triplets that occur within the dataset.\n\nWhile the two types of problems have about equal occurences in the dataset, the triplet relationship detection has about 5x more distinct relationships - meaning it has about 5x more weight on the LB score, due to 80% of the metric being the mean of relationship-wise AP.\n\n**Attribute Classification**\n\nFor this subproblem, I built a training dataset by getting the bounding boxes of all objects which could have a descriptive attribute from the bbox ground truth file. I then matched these objects with the boxes in the relationships ground truth file, which gave me a ground truth (for each object example, which attributes are present).\n\nOn this dataset, I then trained a single pre-trained DenseNet121 network on the cropped objects resized to 224x224, to predict the probability of all 5 classes - this achieved about 0.95-0.99 AUC on all the classes.\n\nThe pre-trained model is then used to predict the probability of the 5 attributes for all the applicable bounding boxes in the object detection challenge, with the final predicted probability as follows:\n\n    P(chair is wooden) = P(bounding box is chair) * P(bounding box is wooden)\n\nWhere the `bounding box is chair` probability is the output of the object detection model.\n\n**Triplet Relationships**\n\nFor the visual relationship triplets, I also built a dataset similarly to attribute classification - all potential pairs of bounding boxes of the two objects, and whether they have the given relationship.\n\nThis gave me around 250 training sets in total for the different triplets (of varying size).\nFor the first 100 relationships, I built a 5-fold XGBoost model on each one with the following features:\n\n- % box 1 inside box 2, % box 2 inside box 1\n- IoU (intersection over union between the boxes)\n- Horizontal/vertical offset of the box centres\n- Euclidean distance between the two box centres\n- Euclidean distance normalised by the size of the boxes (zoom invariance)\n\nAnd then using the models, the submitted probability of each relationship is for example:\n\n    P(chair at table) = P(bounding box is chair) * P(bounding box is table) * P(chair at table XGBoost)\n\nIt seemed that building a separate model for each relationship worked well, as I noticed they each had very different feature importances (eg. sometimes the model was looking at the overlap, while for others the model was looking at the difference in height between the objects). Most classes also had &gt;0.95 AUC, showing the XGBoost model was very good at classifying whether a relationship existed.\n\nFor the other relationships (with a few hundred samples or less) I replaced the xgboost model with a simple prior - what proportion of pair occurences had the relationship in the training data. I tried using eg. linear models for these but didn’t see an improvement.\n\nOverall, the solution runs from start to finish in under 24 hours on an i7+single GPU - I am happy with the performance given the simplicity. I observed a linear increase in performance in visual relationships given an improvement in the bounding boxes scores, so this solution could have potentially scored much higher if I had a better object detector! :)\n\nI am curious to see how other competitors approached the problem, as this is quite a new problem for Kaggle.\n\n\\- anokas\n",
    "379158": "Congratulations, Youngest Grandmaster! 😎",
    "584276": "@anokas would you be publishing your code anytime? ",
    "543737": "Congratulations",
    "379508": "Grandmaster status !!\nCongrats since you never picked the traditional ML competition.  ",
    "379486": "Congrats @anokas! Very interesting solution.\n\nWill also try to do a write up of mine over the next day or two. Like yourself I am also very much looking forward to other write ups - seems we have all taken very differing approaches.",
    "379457": "Congratulations @anokas! Grandmaster is coming :)",
    "379157": "Congratulations! I started this competition pretty late, so I didn't have time to prepare an object detector that is good enough for this competition......",
    "379152": "congratulations on competitions grandmaster status!",
    "565561": ""
  }
}