{
  "id": 64635,
  "title": "45th place solution (an economical one)",
  "url": "/competitions/google-ai-open-images-visual-relationship-track/writeups/kent-ai-lab-alexander-liao-45th-place-solution-an-",
  "author_name": "",
  "post_date": "2018-08-31T01:36:31.846805100Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I was busy preparing for the TrackML contest (and other things in life), so I only had around three weeks to prepare for this year's Open Image. I bet that there are other solutions that would result in way better score, but this solution might be the best tradeoff between time, cost, and performance. Apart from the (stupid) attempt of trying to retrain the pretrained network on V4 dataset, it also does not use much GPU resources since I don't have access to any dedicated server.</p>\n\n<p>In short, I tried to re-create results obtained from the paper \"Visual Relationship Detection with Deep Structural Ranking.\" (AAAI 2018) Their GitHub repo can be found at <a href=\"https://github.com/GriffinLiang/vrd-dsr\">https://github.com/GriffinLiang/vrd-dsr</a>. First, I used the <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\">faster_rcnn_inception_resnet_v2_atrous_oid</a> pretrained model, found in the tensorflow model detection zoo. I then tweaked the inference script to generate kaggle prediction strings, and ran in parallel on 8 Nvidia K80 GPUs. Considering the current hourly rate on GCP of $0.45 USD per GPU per hour, the cost is around 35 US dollars. This on-the-shelf inference gave me 0.23 public LB on the detection track, but according to <a href=\"https://www.kaggle.com/anokas\"></a><a href=\"/anokas\">@anokas</a> a few other techniques can boost the LB drastically to 0.37.</p>\n\n<p>Then, I scp'd the detection track back to my personal workstation and closed the cloud instance. Here's the network architecture of the vrd-dsr network:\n<img src=\"https://preview.ibb.co/mtsHL9/net.png\" alt=\"\"></p>\n\n<p>Since we've already had spatial features (provided by the bounding boxes), we can construct a training pipeline with clever code-tweaking to avoid using the memory-expensive object detector. I didn't get the word2vec semantic embedding of words to work, but if I had more time it's definitely worth a try. The pretrained vgg16 CNN is also available on the GitHub repo, so all we need are the two following files: train.pkl and so_prior.pkl</p>\n\n<p>so_prior stands for (subject-object prior). According to the author, \"so_prior.pkl contains the conditional probability given subject and object. For example, if we want to calculate p(ride|person, horse), we should find all the relationships describing person and horse and the exact relationships of person-ride-horse. We use N and M to denote the number of the above two relationships, and p(ride|person, horse) can be simply acquired as N/M.\" So we can create such a file for this google open image challenge, and it has a shape of (subject, object, predicate).</p>\n\n<p>train.pkl description from the repo:\npython list\neach item is a dictionary with the following keys: {'img_path', 'classes', 'boxes', 'ix1', 'ix2', 'rel_classes'}\n'classes' and 'boxes' describe the objects contained in a single image.\n'ix1': subject index.\n'ix2': object index.\n'rel_classes': relationship for a subject-object pair.</p>\n\n<p>So, I just needed to re-format the training data into this train.pkl. Since 'classes' and 'rel_classes' had to be integers, I simply used a not-so-fancy numeric encoding by a python dictionary, and use the inverted version in the end to decode. Since I needed to filter samples in the official train csv according to the classes I put in so_prior.pkl, making the so_prior from vrd classes didn't really work because there were too few samples that fit inside the vrd class. I don't know if it's a bug in my code or not. In the end, I created the so_prior from boxable classes, added with visual attribute classes (in a later attempt), so the shape is either (601,601,9) or (606,606,10).</p>\n\n<p>It takes around 4 hours to train the network on my Nvidia Quadro M2000M with 4GB ram, so it's pretty achievable. I could've create a separate vrd-dsr network just for visual attributes, but I didn't have time to work on another network. I will probably release the tweaked code next week if anyone's interesed, but now the working directory is just too messy.</p>\n\n<p>I know that a 45/234 place isn't that impressive, but I believe that this is the simplest solution you can find that gives you a silver. It also doesn't use much resources (I would be really happy if I could use a dedicated 1080Ti server to try on more experiments), so I believe that it's worth a look. This competition has the largest dataset I've ever seen, and I learned a lot from trying to work it out with limited time and resource. Again, great work guys, and congratulate <a href=\"/anokas\">@anokas</a> for being the youngest GM in the competition category!</p>",
  "messages": [
    {
      "id": "379187",
      "postDate": "08/31/2018 01:36:31",
      "content": "<p>I was busy preparing for the TrackML contest (and other things in life), so I only had around three weeks to prepare for this year's Open Image. I bet that there are other solutions that would result in way better score, but this solution might be the best tradeoff between time, cost, and performance. Apart from the (stupid) attempt of trying to retrain the pretrained network on V4 dataset, it also does not use much GPU resources since I don't have access to any dedicated server.</p>\n\n<p>In short, I tried to re-create results obtained from the paper \"Visual Relationship Detection with Deep Structural Ranking.\" (AAAI 2018) Their GitHub repo can be found at <a href=\"https://github.com/GriffinLiang/vrd-dsr\">https://github.com/GriffinLiang/vrd-dsr</a>. First, I used the <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\">faster_rcnn_inception_resnet_v2_atrous_oid</a> pretrained model, found in the tensorflow model detection zoo. I then tweaked the inference script to generate kaggle prediction strings, and ran in parallel on 8 Nvidia K80 GPUs. Considering the current hourly rate on GCP of $0.45 USD per GPU per hour, the cost is around 35 US dollars. This on-the-shelf inference gave me 0.23 public LB on the detection track, but according to <a href=\"https://www.kaggle.com/anokas\"></a><a href=\"/anokas\">@anokas</a> a few other techniques can boost the LB drastically to 0.37.</p>\n\n<p>Then, I scp'd the detection track back to my personal workstation and closed the cloud instance. Here's the network architecture of the vrd-dsr network:\n<img src=\"https://preview.ibb.co/mtsHL9/net.png\" alt=\"\"></p>\n\n<p>Since we've already had spatial features (provided by the bounding boxes), we can construct a training pipeline with clever code-tweaking to avoid using the memory-expensive object detector. I didn't get the word2vec semantic embedding of words to work, but if I had more time it's definitely worth a try. The pretrained vgg16 CNN is also available on the GitHub repo, so all we need are the two following files: train.pkl and so_prior.pkl</p>\n\n<p>so_prior stands for (subject-object prior). According to the author, \"so_prior.pkl contains the conditional probability given subject and object. For example, if we want to calculate p(ride|person, horse), we should find all the relationships describing person and horse and the exact relationships of person-ride-horse. We use N and M to denote the number of the above two relationships, and p(ride|person, horse) can be simply acquired as N/M.\" So we can create such a file for this google open image challenge, and it has a shape of (subject, object, predicate).</p>\n\n<p>train.pkl description from the repo:\npython list\neach item is a dictionary with the following keys: {'img_path', 'classes', 'boxes', 'ix1', 'ix2', 'rel_classes'}\n'classes' and 'boxes' describe the objects contained in a single image.\n'ix1': subject index.\n'ix2': object index.\n'rel_classes': relationship for a subject-object pair.</p>\n\n<p>So, I just needed to re-format the training data into this train.pkl. Since 'classes' and 'rel_classes' had to be integers, I simply used a not-so-fancy numeric encoding by a python dictionary, and use the inverted version in the end to decode. Since I needed to filter samples in the official train csv according to the classes I put in so_prior.pkl, making the so_prior from vrd classes didn't really work because there were too few samples that fit inside the vrd class. I don't know if it's a bug in my code or not. In the end, I created the so_prior from boxable classes, added with visual attribute classes (in a later attempt), so the shape is either (601,601,9) or (606,606,10).</p>\n\n<p>It takes around 4 hours to train the network on my Nvidia Quadro M2000M with 4GB ram, so it's pretty achievable. I could've create a separate vrd-dsr network just for visual attributes, but I didn't have time to work on another network. I will probably release the tweaked code next week if anyone's interesed, but now the working directory is just too messy.</p>\n\n<p>I know that a 45/234 place isn't that impressive, but I believe that this is the simplest solution you can find that gives you a silver. It also doesn't use much resources (I would be really happy if I could use a dedicated 1080Ti server to try on more experiments), so I believe that it's worth a look. This competition has the largest dataset I've ever seen, and I learned a lot from trying to work it out with limited time and resource. Again, great work guys, and congratulate <a href=\"/anokas\">@anokas</a> for being the youngest GM in the competition category!</p>",
      "rawMarkdown": "I was busy preparing for the TrackML contest (and other things in life), so I only had around three weeks to prepare for this year's Open Image. I bet that there are other solutions that would result in way better score, but this solution might be the best tradeoff between time, cost, and performance. Apart from the (stupid) attempt of trying to retrain the pretrained network on V4 dataset, it also does not use much GPU resources since I don't have access to any dedicated server.\n\nIn short, I tried to re-create results obtained from the paper \"Visual Relationship Detection with Deep Structural Ranking.\" (AAAI 2018) Their GitHub repo can be found at [https://github.com/GriffinLiang/vrd-dsr][1]. First, I used the [faster_rcnn_inception_resnet_v2_atrous_oid][2] pretrained model, found in the tensorflow model detection zoo. I then tweaked the inference script to generate kaggle prediction strings, and ran in parallel on 8 Nvidia K80 GPUs. Considering the current hourly rate on GCP of $0.45 USD per GPU per hour, the cost is around 35 US dollars. This on-the-shelf inference gave me 0.23 public LB on the detection track, but according to [@anokas][3] a few other techniques can boost the LB drastically to 0.37.\n\nThen, I scp'd the detection track back to my personal workstation and closed the cloud instance. Here's the network architecture of the vrd-dsr network:\n![][4]\n\nSince we've already had spatial features (provided by the bounding boxes), we can construct a training pipeline with clever code-tweaking to avoid using the memory-expensive object detector. I didn't get the word2vec semantic embedding of words to work, but if I had more time it's definitely worth a try. The pretrained vgg16 CNN is also available on the GitHub repo, so all we need are the two following files: train.pkl and so_prior.pkl\n\nso_prior stands for (subject-object prior). According to the author, \"so_prior.pkl contains the conditional probability given subject and object. For example, if we want to calculate p(ride|person, horse), we should find all the relationships describing person and horse and the exact relationships of person-ride-horse. We use N and M to denote the number of the above two relationships, and p(ride|person, horse) can be simply acquired as N/M.\" So we can create such a file for this google open image challenge, and it has a shape of (subject, object, predicate).\n\ntrain.pkl description from the repo:\npython list\neach item is a dictionary with the following keys: {'img_path', 'classes', 'boxes', 'ix1', 'ix2', 'rel_classes'}\n'classes' and 'boxes' describe the objects contained in a single image.\n'ix1': subject index.\n'ix2': object index.\n'rel_classes': relationship for a subject-object pair.\n\nSo, I just needed to re-format the training data into this train.pkl. Since 'classes' and 'rel_classes' had to be integers, I simply used a not-so-fancy numeric encoding by a python dictionary, and use the inverted version in the end to decode. Since I needed to filter samples in the official train csv according to the classes I put in so_prior.pkl, making the so_prior from vrd classes didn't really work because there were too few samples that fit inside the vrd class. I don't know if it's a bug in my code or not. In the end, I created the so_prior from boxable classes, added with visual attribute classes (in a later attempt), so the shape is either (601,601,9) or (606,606,10).\n\nIt takes around 4 hours to train the network on my Nvidia Quadro M2000M with 4GB ram, so it's pretty achievable. I could've create a separate vrd-dsr network just for visual attributes, but I didn't have time to work on another network. I will probably release the tweaked code next week if anyone's interesed, but now the working directory is just too messy.\n\nI know that a 45/234 place isn't that impressive, but I believe that this is the simplest solution you can find that gives you a silver. It also doesn't use much resources (I would be really happy if I could use a dedicated 1080Ti server to try on more experiments), so I believe that it's worth a look. This competition has the largest dataset I've ever seen, and I learned a lot from trying to work it out with limited time and resource. Again, great work guys, and congratulate @anokas for being the youngest GM in the competition category!\n\n  [1]: https://github.com/GriffinLiang/vrd-dsr\n  [2]: https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\n  [3]: https://www.kaggle.com/anokas\n  [4]: https://preview.ibb.co/mtsHL9/net.png",
      "votes": null
    },
    {
      "id": "379204",
      "postDate": "08/31/2018 02:24:20",
      "content": "<p>Interesting; don't know why deep neural networks did poorly here.</p>",
      "rawMarkdown": "Interesting; don't know why deep neural networks did poorly here.",
      "votes": null
    },
    {
      "id": "872955",
      "postDate": "06/03/2020 16:25:24",
      "content": "<p>Your approach is interesting. Will you be releasing the code ?</p>",
      "rawMarkdown": "Your approach is interesting. Will you be releasing the code ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 379204,
      "author_name": "nzabcd",
      "author_url": "",
      "post_date": "08/31/2018 02:24:20",
      "content": "<p>Interesting; don't know why deep neural networks did poorly here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 872955,
      "author_name": "varungadre0910",
      "author_url": "",
      "post_date": "06/03/2020 16:25:24",
      "content": "<p>Your approach is interesting. Will you be releasing the code ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "379187": "I was busy preparing for the TrackML contest (and other things in life), so I only had around three weeks to prepare for this year's Open Image. I bet that there are other solutions that would result in way better score, but this solution might be the best tradeoff between time, cost, and performance. Apart from the (stupid) attempt of trying to retrain the pretrained network on V4 dataset, it also does not use much GPU resources since I don't have access to any dedicated server.\n\nIn short, I tried to re-create results obtained from the paper \"Visual Relationship Detection with Deep Structural Ranking.\" (AAAI 2018) Their GitHub repo can be found at [https://github.com/GriffinLiang/vrd-dsr][1]. First, I used the [faster_rcnn_inception_resnet_v2_atrous_oid][2] pretrained model, found in the tensorflow model detection zoo. I then tweaked the inference script to generate kaggle prediction strings, and ran in parallel on 8 Nvidia K80 GPUs. Considering the current hourly rate on GCP of $0.45 USD per GPU per hour, the cost is around 35 US dollars. This on-the-shelf inference gave me 0.23 public LB on the detection track, but according to [@anokas][3] a few other techniques can boost the LB drastically to 0.37.\n\nThen, I scp'd the detection track back to my personal workstation and closed the cloud instance. Here's the network architecture of the vrd-dsr network:\n![][4]\n\nSince we've already had spatial features (provided by the bounding boxes), we can construct a training pipeline with clever code-tweaking to avoid using the memory-expensive object detector. I didn't get the word2vec semantic embedding of words to work, but if I had more time it's definitely worth a try. The pretrained vgg16 CNN is also available on the GitHub repo, so all we need are the two following files: train.pkl and so_prior.pkl\n\nso_prior stands for (subject-object prior). According to the author, \"so_prior.pkl contains the conditional probability given subject and object. For example, if we want to calculate p(ride|person, horse), we should find all the relationships describing person and horse and the exact relationships of person-ride-horse. We use N and M to denote the number of the above two relationships, and p(ride|person, horse) can be simply acquired as N/M.\" So we can create such a file for this google open image challenge, and it has a shape of (subject, object, predicate).\n\ntrain.pkl description from the repo:\npython list\neach item is a dictionary with the following keys: {'img_path', 'classes', 'boxes', 'ix1', 'ix2', 'rel_classes'}\n'classes' and 'boxes' describe the objects contained in a single image.\n'ix1': subject index.\n'ix2': object index.\n'rel_classes': relationship for a subject-object pair.\n\nSo, I just needed to re-format the training data into this train.pkl. Since 'classes' and 'rel_classes' had to be integers, I simply used a not-so-fancy numeric encoding by a python dictionary, and use the inverted version in the end to decode. Since I needed to filter samples in the official train csv according to the classes I put in so_prior.pkl, making the so_prior from vrd classes didn't really work because there were too few samples that fit inside the vrd class. I don't know if it's a bug in my code or not. In the end, I created the so_prior from boxable classes, added with visual attribute classes (in a later attempt), so the shape is either (601,601,9) or (606,606,10).\n\nIt takes around 4 hours to train the network on my Nvidia Quadro M2000M with 4GB ram, so it's pretty achievable. I could've create a separate vrd-dsr network just for visual attributes, but I didn't have time to work on another network. I will probably release the tweaked code next week if anyone's interesed, but now the working directory is just too messy.\n\nI know that a 45/234 place isn't that impressive, but I believe that this is the simplest solution you can find that gives you a silver. It also doesn't use much resources (I would be really happy if I could use a dedicated 1080Ti server to try on more experiments), so I believe that it's worth a look. This competition has the largest dataset I've ever seen, and I learned a lot from trying to work it out with limited time and resource. Again, great work guys, and congratulate @anokas for being the youngest GM in the competition category!\n\n  [1]: https://github.com/GriffinLiang/vrd-dsr\n  [2]: https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/detection_model_zoo.md\n  [3]: https://www.kaggle.com/anokas\n  [4]: https://preview.ibb.co/mtsHL9/net.png",
    "379204": "Interesting; don't know why deep neural networks did poorly here.",
    "872955": "Your approach is interesting. Will you be releasing the code ?"
  },
  "source": "meta"
}