{
  "id": 95393,
  "title": "[Solution] Public: 0.657 => Private: Error (OH MY GOD...)",
  "url": "/competitions/imet-2019-fgvc6/discussion/95393",
  "author_name": "",
  "post_date": "2019-06-12T00:06:05.772969200Z",
  "votes": 45,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Congrats to  all the teams got the medal, and thanks to Kaggle and hosting team for this  interesting competition !</p>\n\n<p>Unfortunately, I made a mistake and lost my place in leaderboard, I'm so sad. <br>\nBut I really enjoyed competing with you kagglers for this 2 months :)</p>\n\n<p>Here, I'd like to share my solution.</p>\n\n<h2>Basic idea</h2>\n\n<p>This competition's  dataset includes various size images. So, I trained multiple models which input sizes have <strong>different aspect ratios</strong>, then made ensemble of them.</p>\n\n<h2>level 1 (pure classification models)</h2>\n\n<h3><strong>Model</strong></h3>\n\n<p>SE-ResNeXt50(1 model) and SE-ResNeXt101(8 model).  Global average pooling  layer is followed by FC, ReLU, Dropout and FC.</p>\n\n<h3><strong>Data Augmentation</strong></h3>\n\n<p>I applied <a href=\"https://chainercv.readthedocs.io/en/stable/reference/links/ssd.html#random-distort\">random_distort</a> (brightness, contrast, saturation, hue), horizontal flip, random crop, random erase, and pca lighting.</p>\n\n<h3><strong>Loss</strong></h3>\n\n<p>I used <a href=\"https://arxiv.org/abs/1708.02002\">FocalLoss</a> with <strong>alpha=0.75</strong> and gamma=1.5. This alpha setting is because <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/overview/evaluation\">F2 score weights <strong>recall</strong> higher than precision</a>.</p>\n\n<h3><strong>Optimization</strong></h3>\n\n<ul>\n<li>Optimizer : SGD + NesterovAG\n<ul><li>momentum = 0.9, weight decay=1e-04</li></ul></li>\n<li>learning rate: differs by model. These are decided by <a href=\"https://arxiv.org/abs/1506.01186\">LR-Range Test</a>.</li>\n<li>batch size:  differs by model.</li>\n<li>epoch: 20</li>\n<li>learning schedule: cosine anealing (one cycle) with linear warmup(4 epoch)</li>\n</ul>\n\n<h3>Threshold Decision</h3>\n\n<p>As I wrote in <a href=\"https://www.kaggle.com/ttahara/eda-compare-number-of-culture-and-tag-attributes\">this kernel</a>, there is difference between number of culture attributes and one of tag attributes each art has. \nThen, I prepared two separated thresholds for culture attribute and tag attribute.</p>\n\n<h3>summary of each model</h3>\n\n<p>|     model      |image size  |aspect ratio| learning rate |   5foldcv    | Public | \n|:--------------:|:----------:|:----------:|:-------------:|:------------:|:------:|\n| SE-ResNeXt-50  |  288 x 288 | 1 : 1      |   2.15e-02    |    0.6132    |  0.637 |\n| SE-ResNeXt-101 |  368 x 368 | 1 : 1      |   1.8e-02     |    0.6233    |    -   |\n| SE-ResNeXt-101 |  448 x 448 | 1 : 1      |   1.5e-02     |    0.6243    |  0.647 |\n| SE-ResNeXt-101 |  368 x 560 | 2 : 3      |   1.3e-02     |    0.6228    |    -   |\n| SE-ResNeXt-101 |  560 x 368 | 3 : 2      |   1.3e-02     |    0.6232    |    -   |\n| SE-ResNeXt-101 |  320 x 640 | 1 : 2      |   1.2e-02     |    0.6190    |    -   |\n| SE-ResNeXt-101 |  640 x 320 | 2 : 1      |   1.2e-02     |    0.6197    |    -   |\n| SE-ResNeXt-101 |  256 x 784 | 1 : 3      |   1.8e-02     |    0.6127    |    -   |\n| SE-ResNeXt-101 |  784 x 256 | 3 : 1      |   1.8e-02     |    0.6147    |    -   |</p>\n\n<p>When finished training all the models,  there were only a few days left for me. So I got  public scores of only 288x288 and 448x448.</p>\n\n<h2>level2 (ensemble models)</h2>\n\n<p>In the last few days, I tried several ensemble method. I explain them in chronological order.</p>\n\n<h3>0. [CV:0.6456, Public:0.652] averaging (2days to go)</h3>\n\n<p>Simply averaging probabirities output by 9 models.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_2days_to_go.png\" alt=\"solution  2 days to go\"></p>\n\n<h3>1. [CV:0.6473, Public:0.654] stacking by <em>class-wise</em> MLPs shaing weights(1 days to go):</h3>\n\n<p>I applied a MLP (which have one hidden layer) to multiple models' outputs for each class <strong>separately</strong>. These MLPs share their weights.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_1day_to_go.png\" alt=\"solution 1 day to go\"></p>\n\n<h3>2. [CV:0.6483, Public:0.656] stacking by <em>GNN</em> and class-wise MLPs shaing weights(the last day)</h3>\n\n<p>Each <em>class-wise</em> MLP, mentioned above, <strong>cannot consider</strong> outputs for other classes.  Got an inspiration from <a href=\"https://arxiv.org/abs/1902.09720\">this paper</a>, I tried applying GNN (Graph Neural Network).\nThis paper applied GNN for <strong>single</strong> model's outputs for each class, however, I applied it for <strong>multiple</strong> models' outputs.</p>\n\n<p>Logits for each class obtained from level1 models are assigned to one node. After several iterations of updating messages and hidden states, each class node's hidden state is feeded into the class-wise MLP.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_last_day.png\" alt=\"solution the last day\"></p>\n\n<p><strong><em>Note</em></strong>: This GNN stacking model is incomplete because <strong>the update functions for messages and hidden states are shared among all classes</strong>, for simple implementation (I've wanted to revise, but the deadline came earlier).  This means <strong>all edges have the same weight</strong>. I'll improve this model and try it in next other multi-class classification competition.</p>\n\n<h3>3. [CV:0.6487, Public:0.657] stacking by <strong>GNN</strong> and class-wise MLPs shaing weights <em>with image size info</em>(the additional day)</h3>\n\n<p>Additionally, I used meta features of images related with image size, as I made ensemble of models whose input have various aspect ratios. This slightly improved my score.  </p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_additional_day.png\" alt=\"solution additional day\"></p>\n\n<p>Sorry for my poor english, and feel free to ask me if you have any questions.</p>\n\n<p>Thank you for reading!</p>",
  "messages": [
    {
      "id": "550701",
      "postDate": "06/12/2019 00:06:05",
      "content": "<p>Congrats to  all the teams got the medal, and thanks to Kaggle and hosting team for this  interesting competition !</p>\n\n<p>Unfortunately, I made a mistake and lost my place in leaderboard, I'm so sad. <br>\nBut I really enjoyed competing with you kagglers for this 2 months :)</p>\n\n<p>Here, I'd like to share my solution.</p>\n\n<h2>Basic idea</h2>\n\n<p>This competition's  dataset includes various size images. So, I trained multiple models which input sizes have <strong>different aspect ratios</strong>, then made ensemble of them.</p>\n\n<h2>level 1 (pure classification models)</h2>\n\n<h3><strong>Model</strong></h3>\n\n<p>SE-ResNeXt50(1 model) and SE-ResNeXt101(8 model).  Global average pooling  layer is followed by FC, ReLU, Dropout and FC.</p>\n\n<h3><strong>Data Augmentation</strong></h3>\n\n<p>I applied <a href=\"https://chainercv.readthedocs.io/en/stable/reference/links/ssd.html#random-distort\">random_distort</a> (brightness, contrast, saturation, hue), horizontal flip, random crop, random erase, and pca lighting.</p>\n\n<h3><strong>Loss</strong></h3>\n\n<p>I used <a href=\"https://arxiv.org/abs/1708.02002\">FocalLoss</a> with <strong>alpha=0.75</strong> and gamma=1.5. This alpha setting is because <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/overview/evaluation\">F2 score weights <strong>recall</strong> higher than precision</a>.</p>\n\n<h3><strong>Optimization</strong></h3>\n\n<ul>\n<li>Optimizer : SGD + NesterovAG\n<ul><li>momentum = 0.9, weight decay=1e-04</li></ul></li>\n<li>learning rate: differs by model. These are decided by <a href=\"https://arxiv.org/abs/1506.01186\">LR-Range Test</a>.</li>\n<li>batch size:  differs by model.</li>\n<li>epoch: 20</li>\n<li>learning schedule: cosine anealing (one cycle) with linear warmup(4 epoch)</li>\n</ul>\n\n<h3>Threshold Decision</h3>\n\n<p>As I wrote in <a href=\"https://www.kaggle.com/ttahara/eda-compare-number-of-culture-and-tag-attributes\">this kernel</a>, there is difference between number of culture attributes and one of tag attributes each art has. \nThen, I prepared two separated thresholds for culture attribute and tag attribute.</p>\n\n<h3>summary of each model</h3>\n\n<p>|     model      |image size  |aspect ratio| learning rate |   5foldcv    | Public | \n|:--------------:|:----------:|:----------:|:-------------:|:------------:|:------:|\n| SE-ResNeXt-50  |  288 x 288 | 1 : 1      |   2.15e-02    |    0.6132    |  0.637 |\n| SE-ResNeXt-101 |  368 x 368 | 1 : 1      |   1.8e-02     |    0.6233    |    -   |\n| SE-ResNeXt-101 |  448 x 448 | 1 : 1      |   1.5e-02     |    0.6243    |  0.647 |\n| SE-ResNeXt-101 |  368 x 560 | 2 : 3      |   1.3e-02     |    0.6228    |    -   |\n| SE-ResNeXt-101 |  560 x 368 | 3 : 2      |   1.3e-02     |    0.6232    |    -   |\n| SE-ResNeXt-101 |  320 x 640 | 1 : 2      |   1.2e-02     |    0.6190    |    -   |\n| SE-ResNeXt-101 |  640 x 320 | 2 : 1      |   1.2e-02     |    0.6197    |    -   |\n| SE-ResNeXt-101 |  256 x 784 | 1 : 3      |   1.8e-02     |    0.6127    |    -   |\n| SE-ResNeXt-101 |  784 x 256 | 3 : 1      |   1.8e-02     |    0.6147    |    -   |</p>\n\n<p>When finished training all the models,  there were only a few days left for me. So I got  public scores of only 288x288 and 448x448.</p>\n\n<h2>level2 (ensemble models)</h2>\n\n<p>In the last few days, I tried several ensemble method. I explain them in chronological order.</p>\n\n<h3>0. [CV:0.6456, Public:0.652] averaging (2days to go)</h3>\n\n<p>Simply averaging probabirities output by 9 models.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_2days_to_go.png\" alt=\"solution  2 days to go\"></p>\n\n<h3>1. [CV:0.6473, Public:0.654] stacking by <em>class-wise</em> MLPs shaing weights(1 days to go):</h3>\n\n<p>I applied a MLP (which have one hidden layer) to multiple models' outputs for each class <strong>separately</strong>. These MLPs share their weights.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_1day_to_go.png\" alt=\"solution 1 day to go\"></p>\n\n<h3>2. [CV:0.6483, Public:0.656] stacking by <em>GNN</em> and class-wise MLPs shaing weights(the last day)</h3>\n\n<p>Each <em>class-wise</em> MLP, mentioned above, <strong>cannot consider</strong> outputs for other classes.  Got an inspiration from <a href=\"https://arxiv.org/abs/1902.09720\">this paper</a>, I tried applying GNN (Graph Neural Network).\nThis paper applied GNN for <strong>single</strong> model's outputs for each class, however, I applied it for <strong>multiple</strong> models' outputs.</p>\n\n<p>Logits for each class obtained from level1 models are assigned to one node. After several iterations of updating messages and hidden states, each class node's hidden state is feeded into the class-wise MLP.</p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_last_day.png\" alt=\"solution the last day\"></p>\n\n<p><strong><em>Note</em></strong>: This GNN stacking model is incomplete because <strong>the update functions for messages and hidden states are shared among all classes</strong>, for simple implementation (I've wanted to revise, but the deadline came earlier).  This means <strong>all edges have the same weight</strong>. I'll improve this model and try it in next other multi-class classification competition.</p>\n\n<h3>3. [CV:0.6487, Public:0.657] stacking by <strong>GNN</strong> and class-wise MLPs shaing weights <em>with image size info</em>(the additional day)</h3>\n\n<p>Additionally, I used meta features of images related with image size, as I made ensemble of models whose input have various aspect ratios. This slightly improved my score.  </p>\n\n<p><img src=\"https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_additional_day.png\" alt=\"solution additional day\"></p>\n\n<p>Sorry for my poor english, and feel free to ask me if you have any questions.</p>\n\n<p>Thank you for reading!</p>",
      "rawMarkdown": "Congrats to  all the teams got the medal, and thanks to Kaggle and hosting team for this  interesting competition !\n\nUnfortunately, I made a mistake and lost my place in leaderboard, I'm so sad.  \nBut I really enjoyed competing with you kagglers for this 2 months :)\n\nHere, I'd like to share my solution.\n\n## Basic idea\n\nThis competition's  dataset includes various size images. So, I trained multiple models which input sizes have **different aspect ratios**, then made ensemble of them.\n\n## level 1 (pure classification models)\n\n### **Model**\nSE-ResNeXt50(1 model) and SE-ResNeXt101(8 model).  Global average pooling  layer is followed by FC, ReLU, Dropout and FC.\n\n### **Data Augmentation**\nI applied [random_distort](https://chainercv.readthedocs.io/en/stable/reference/links/ssd.html#random-distort) (brightness, contrast, saturation, hue), horizontal flip, random crop, random erase, and pca lighting.\n\n### **Loss**\n\nI used [FocalLoss](https://arxiv.org/abs/1708.02002) with **alpha=0.75** and gamma=1.5. This alpha setting is because [F2 score weights **recall** higher than precision](https://www.kaggle.com/c/imet-2019-fgvc6/overview/evaluation).\n\n### **Optimization**\n* Optimizer : SGD + NesterovAG\n    * momentum = 0.9, weight decay=1e-04\n* learning rate: differs by model. These are decided by [LR-Range Test](https://arxiv.org/abs/1506.01186).\n* batch size:  differs by model.\n* epoch: 20\n* learning schedule: cosine anealing (one cycle) with linear warmup(4 epoch)\n\n### Threshold Decision\nAs I wrote in [this kernel](https://www.kaggle.com/ttahara/eda-compare-number-of-culture-and-tag-attributes), there is difference between number of culture attributes and one of tag attributes each art has. \nThen, I prepared two separated thresholds for culture attribute and tag attribute.\n\n### summary of each model\n\n|     model      |image size  |aspect ratio| learning rate |   5foldcv    | Public | \n|:--------------:|:----------:|:----------:|:-------------:|:------------:|:------:|\n| SE-ResNeXt-50  |  288 x 288 | 1 : 1      |   2.15e-02    |    0.6132    |  0.637 |\n| SE-ResNeXt-101 |  368 x 368 | 1 : 1      |   1.8e-02     |    0.6233    |    -   |\n| SE-ResNeXt-101 |  448 x 448 | 1 : 1      |   1.5e-02     |    0.6243    |  0.647 |\n| SE-ResNeXt-101 |  368 x 560 | 2 : 3      |   1.3e-02     |    0.6228    |    -   |\n| SE-ResNeXt-101 |  560 x 368 | 3 : 2      |   1.3e-02     |    0.6232    |    -   |\n| SE-ResNeXt-101 |  320 x 640 | 1 : 2      |   1.2e-02     |    0.6190    |    -   |\n| SE-ResNeXt-101 |  640 x 320 | 2 : 1      |   1.2e-02     |    0.6197    |    -   |\n| SE-ResNeXt-101 |  256 x 784 | 1 : 3      |   1.8e-02     |    0.6127    |    -   |\n| SE-ResNeXt-101 |  784 x 256 | 3 : 1      |   1.8e-02     |    0.6147    |    -   |\n\n  \nWhen finished training all the models,  there were only a few days left for me. So I got  public scores of only 288x288 and 448x448.\n\n## level2 (ensemble models)\nIn the last few days, I tried several ensemble method. I explain them in chronological order.\n\n### 0. [CV:0.6456, Public:0.652] averaging (2days to go)\n\nSimply averaging probabirities output by 9 models.\n  \n  \n![solution  2 days to go](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_2days_to_go.png)\n  \n  \n### 1. [CV:0.6473, Public:0.654] stacking by _class-wise_ MLPs shaing weights(1 days to go): \n  \nI applied a MLP (which have one hidden layer) to multiple models' outputs for each class **separately**. These MLPs share their weights.\n  \n  \n![solution 1 day to go](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_1day_to_go.png)\n  \n  \n### 2. [CV:0.6483, Public:0.656] stacking by _GNN_ and class-wise MLPs shaing weights(the last day)\n\nEach _class-wise_ MLP, mentioned above, **cannot consider** outputs for other classes.  Got an inspiration from [this paper](https://arxiv.org/abs/1902.09720), I tried applying GNN (Graph Neural Network).\nThis paper applied GNN for **single** model's outputs for each class, however, I applied it for **multiple** models' outputs.\n\n\nLogits for each class obtained from level1 models are assigned to one node. After several iterations of updating messages and hidden states, each class node's hidden state is feeded into the class-wise MLP.\n\n![solution the last day](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_last_day.png)\n\n**_Note_**: This GNN stacking model is incomplete because **the update functions for messages and hidden states are shared among all classes**, for simple implementation (I've wanted to revise, but the deadline came earlier).  This means **all edges have the same weight**. I'll improve this model and try it in next other multi-class classification competition.\n\n### 3. [CV:0.6487, Public:0.657] stacking by **GNN** and class-wise MLPs shaing weights _with image size info_(the additional day)\n\nAdditionally, I used meta features of images related with image size, as I made ensemble of models whose input have various aspect ratios. This slightly improved my score.  \n\n![solution additional day](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_additional_day.png)\n  \n  \n  \n  \n  \nSorry for my poor english, and feel free to ask me if you have any questions.\n\nThank you for reading!",
      "votes": null
    },
    {
      "id": "550710",
      "postDate": "06/12/2019 00:40:36",
      "content": "<p>Thank you for sharing! GNN stacking looks interesting. I'm not familiar with GNN, but is there some theoretical advantage of GNN for handling multiple models' outputs? Even if not strictly proved, it's ok for me. If you have any insights, it would be nice if you could share it.</p>",
      "rawMarkdown": "Thank you for sharing! GNN stacking looks interesting. I'm not familiar with GNN, but is there some theoretical advantage of GNN for handling multiple models' outputs? Even if not strictly proved, it's ok for me. If you have any insights, it would be nice if you could share it.",
      "votes": null
    },
    {
      "id": "550711",
      "postDate": "06/12/2019 00:41:34",
      "content": "<p>Thank you for sharing the solution.\nIn the next competition, I look forward to competing with you again!!</p>",
      "rawMarkdown": "Thank you for sharing the solution.\nIn the next competition, I look forward to competing with you again!!",
      "votes": null
    },
    {
      "id": "550714",
      "postDate": "06/12/2019 00:44:44",
      "content": "<p>Very impressive and interesting way for using GNN from classwise feature!\nIs this familiar way? (as you cited the paper)\nAnd what information did you expect to extract from this (GNN)?\nThank you for sharing : )</p>",
      "rawMarkdown": "Very impressive and interesting way for using GNN from classwise feature!\nIs this familiar way? (as you cited the paper)\nAnd what information did you expect to extract from this (GNN)?\nThank you for sharing : )",
      "votes": null
    },
    {
      "id": "550752",
      "postDate": "06/12/2019 02:11:42",
      "content": "<p>Nice presentation! Thank you for your contribution!</p>",
      "rawMarkdown": "Nice presentation! Thank you for your contribution!",
      "votes": null
    },
    {
      "id": "550770",
      "postDate": "06/12/2019 02:36:44",
      "content": "<p><a href=\"/ttahara\">@ttahara</a> \nThank you for sharing this nice solution! <br>\nAlthough I didn't have participated in this competition, your solution is interesting for me.   </p>\n\n<p>Just for checking my understanding, one question.  </p>\n\n<p>In your figure, GNN looks an Octagram ( something like a black magic ! ). <br>\n<a href=\"https://en.wikipedia.org/wiki/Octagram\">https://en.wikipedia.org/wiki/Octagram</a> <br>\n<img src=\"https://cdn-ak.f.st-hatena.com/images/fotolife/g/greenwind120170/20190612/20190612113324.jpg\" alt=\"Black Magic\"></p>\n\n<p>But if I'm right, GNN has the edges whose numbers equals to the kinds of classes, right? <br>\nSorry for a trivial question.  </p>\n\n<p>Thanks in advance.  </p>",
      "rawMarkdown": "ttahara \nThank you for sharing this nice solution!  \nAlthough I didn't have participated in this competition, your solution is interesting for me.   \n  \nJust for checking my understanding, one question.  \n  \nIn your figure, GNN looks an Octagram ( something like a black magic ! ).  \n[https://en.wikipedia.org/wiki/Octagram](https://en.wikipedia.org/wiki/Octagram)  \n![Black Magic](https://cdn-ak.f.st-hatena.com/images/fotolife/g/greenwind120170/20190612/20190612113324.jpg)\n\nBut if I'm right, GNN has the edges whose numbers equals to the kinds of classes, right?  \nSorry for a trivial question.  \n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "550799",
      "postDate": "06/12/2019 03:28:34",
      "content": "<p>This is the most elegant solution!!! Thank you for sharing!</p>",
      "rawMarkdown": "This is the most elegant solution!!! Thank you for sharing!",
      "votes": null
    },
    {
      "id": "550815",
      "postDate": "06/12/2019 04:04:09",
      "content": "<p>I am so sorry that you are not at Private Leader Board...\nBut your solution is very incredible! I will find something useful for a future competition.\nThank you for sharing your solution!!</p>",
      "rawMarkdown": "I am so sorry that you are not at Private Leader Board...\nBut your solution is very incredible! I will find something useful for a future competition.\nThank you for sharing your solution!!",
      "votes": null
    },
    {
      "id": "551042",
      "postDate": "06/12/2019 09:26:31",
      "content": "<p>I'm interested in what software did you use to draw these pictures? Could you share about that?</p>",
      "rawMarkdown": "I'm interested in what software did you use to draw these pictures? Could you share about that?",
      "votes": null
    },
    {
      "id": "551115",
      "postDate": "06/12/2019 11:05:45",
      "content": "<p><a href=\"/sishihara\">@sishihara</a> <a href=\"/triwave33\">@triwave33</a> </p>\n\n<p>Thank you for question about GNN !</p>\n\n<p>&gt;  theoretical advantage </p>\n\n<p>Sorry, I'm not familiar with GNN, too. I cannot explain any theoretical advantage.\nBut I share my intuition and thought about applying GNN stacking to multi-label(class) classification in this comment.</p>\n\n<p>&gt; Is this familiar way?</p>\n\n<p>After reading the paper(<a href=\"https://arxiv.org/abs/1902.09720\">Learning a Deep ConvNet for Multi-label Classification with Partial Labels</a>), I searched other papers which assign CNN outputs for each class to nodes of GNN. But I've not found.\nMoreover, I've not found stacking model  by GNN.</p>\n\n<p>FYI: the following papers may be related.\n* <a href=\"http://openaccess.thecvf.com/content_iccv_2017/html/Qi_3D_Graph_Neural_ICCV_2017_paper.html\">3D Graph Neural Networks for RGBD Semantic Segmentation</a></p>\n\n<p>This paper assigns pixels to nodes.</p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1902.09720\">Multi-Label Image Recognition with Graph Convolutional Networks</a></li>\n</ul>\n\n<p>This paper assigns classes to nodes. But node features are word embedding, not CNN outputs.</p>\n\n<p>&gt; what information did you expect to extract from this (GNN)?</p>\n\n<p>IMO, GNN stacking have two advantages,</p>\n\n<ul>\n<li>It deals with each class's <em>packed</em> feature.</li>\n</ul>\n\n<p>When stacking, I have 9(model) * 1103(class) features.  If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN. </p>\n\n<ul>\n<li>(may be,) It can extract dependency of classes</li>\n</ul>\n\n<p>I think GNN can extract relationships between classes such as \"when class A is labeled, class B is also Labeled\". \nHowever,  for simple implementation, <strong>update functions of this solution's GNN are shared among all classes.</strong>\nSo, I don't understand extracted features in this solution well. </p>\n\n<p>Very sorry for poor English...\nThanks!</p>",
      "rawMarkdown": "sishihara @triwave33 \n\nThank you for question about GNN !\n\n&gt;  theoretical advantage \n\nSorry, I'm not familiar with GNN, too. I cannot explain any theoretical advantage.\nBut I share my intuition and thought about applying GNN stacking to multi-label(class) classification in this comment.\n\n&gt; Is this familiar way?\n\nAfter reading the paper([Learning a Deep ConvNet for Multi-label Classification with Partial Labels](https://arxiv.org/abs/1902.09720)), I searched other papers which assign CNN outputs for each class to nodes of GNN. But I've not found.\nMoreover, I've not found stacking model  by GNN.\n\nFYI: the following papers may be related.\n* [3D Graph Neural Networks for RGBD Semantic Segmentation](http://openaccess.thecvf.com/content_iccv_2017/html/Qi_3D_Graph_Neural_ICCV_2017_paper.html)\n\nThis paper assigns pixels to nodes.\n\n* [Multi-Label Image Recognition with Graph Convolutional Networks](https://arxiv.org/abs/1902.09720)\n\nThis paper assigns classes to nodes. But node features are word embedding, not CNN outputs.\n\n\n&gt; what information did you expect to extract from this (GNN)?\n  \nIMO, GNN stacking have two advantages,\n  \n*  It deals with each class's _packed_ feature.\n  \nWhen stacking, I have 9(model) * 1103(class) features.  If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN. \n  \n* (may be,) It can extract dependency of classes\n  \nI think GNN can extract relationships between classes such as \"when class A is labeled, class B is also Labeled\". \nHowever,  for simple implementation, **update functions of this solution's GNN are shared among all classes.**\nSo, I don't understand extracted features in this solution well. \n\nVery sorry for poor English...\nThanks!",
      "votes": null
    },
    {
      "id": "551118",
      "postDate": "06/12/2019 11:12:11",
      "content": "<p>Yes, I used a black magic (stacking) by a magic circle (GNN), ha-ha!</p>\n\n<p>Your understanding is right. Each class's node is connected to all the other nodes.</p>",
      "rawMarkdown": "Yes, I used a black magic (stacking) by a magic circle (GNN), ha-ha!\n\nYour understanding is right. Each class's node is connected to all the other nodes.",
      "votes": null
    },
    {
      "id": "551120",
      "postDate": "06/12/2019 11:12:46",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "551123",
      "postDate": "06/12/2019 11:14:46",
      "content": "<p>These figures were created by Microsoft PowerPoint.</p>",
      "rawMarkdown": "These figures were created by Microsoft PowerPoint.",
      "votes": null
    },
    {
      "id": "551126",
      "postDate": "06/12/2019 11:16:29",
      "content": "<p>Thank you. I'm very grad to share my idea in this research competition!</p>",
      "rawMarkdown": "Thank you. I'm very grad to share my idea in this research competition!",
      "votes": null
    },
    {
      "id": "551128",
      "postDate": "06/12/2019 11:23:03",
      "content": "<p>Thanks!\nI'm very grad If my idea gives you some inspiration !</p>",
      "rawMarkdown": "Thanks!\nI'm very grad If my idea gives you some inspiration !",
      "votes": null
    },
    {
      "id": "551133",
      "postDate": "06/12/2019 11:31:17",
      "content": "<p>nice, thank you.</p>",
      "rawMarkdown": "nice, thank you.",
      "votes": null
    },
    {
      "id": "551153",
      "postDate": "06/12/2019 11:51:46",
      "content": "<p>Sorry to hear that. \nAnyway both your solution and write-up are beautiful and great. \nI have no doubt you will smash next competitions. </p>",
      "rawMarkdown": "Sorry to hear that. \nAnyway both your solution and write-up are beautiful and great. \nI have no doubt you will smash next competitions.",
      "votes": null
    },
    {
      "id": "551191",
      "postDate": "06/12/2019 12:40:51",
      "content": "<p>Thank you for sharing additional information and citations. I’ll do read them.</p>\n\n<p>From your explanation, I understand the merit of GNN in multiclass competition such that 1) feeding and treating data with retaining their structure and 2) getting dependency between classes.</p>\n\n<p>I’m very very looking forward to your further implementation and getting insight from that. But I believe your idea and solution must be brilliant even now. Thank you again!</p>",
      "rawMarkdown": "Thank you for sharing additional information and citations. I’ll do read them.\n\nFrom your explanation, I understand the merit of GNN in multiclass competition such that 1) feeding and treating data with retaining their structure and 2) getting dependency between classes.\n\nI’m very very looking forward to your further implementation and getting insight from that. But I believe your idea and solution must be brilliant even now. Thank you again!",
      "votes": null
    },
    {
      "id": "551205",
      "postDate": "06/12/2019 13:04:48",
      "content": "<p>Many thanks for your comprehensive yet beautiful solutions! Sorry to hear that you missed private LB. would you mind to share your source code solutions? I'm interested in your GNN integrations with other models! :-)</p>",
      "rawMarkdown": "Many thanks for your comprehensive yet beautiful solutions! Sorry to hear that you missed private LB. would you mind to share your source code solutions? I'm interested in your GNN integrations with other models! :-)",
      "votes": null
    },
    {
      "id": "551257",
      "postDate": "06/12/2019 14:07:01",
      "content": "<p>Thanks!\nWell clarified!</p>",
      "rawMarkdown": "Thanks!\nWell clarified!",
      "votes": null
    },
    {
      "id": "551458",
      "postDate": "06/12/2019 18:38:13",
      "content": "<p>Thanks! See you in the next competition.</p>",
      "rawMarkdown": "Thanks! See you in the next competition.",
      "votes": null
    },
    {
      "id": "551633",
      "postDate": "06/13/2019 00:50:30",
      "content": "<p>Your summarization is so good ! Thank you.</p>",
      "rawMarkdown": "Your summarization is so good ! Thank you.",
      "votes": null
    },
    {
      "id": "551780",
      "postDate": "06/13/2019 05:30:44",
      "content": "<p>I also fail in stage2, sad, too...\nreally impressed by the GNN part, nice work.</p>",
      "rawMarkdown": "I also fail in stage2, sad, too...\nreally impressed by the GNN part, nice work.",
      "votes": null
    },
    {
      "id": "552450",
      "postDate": "06/14/2019 02:13:01",
      "content": "<p>Very nice solution.. I am sorry for the error happened but look a it as a priceless learning opportunities. Next comp you will be stronger..</p>\n\n<p>One question: what is the intuition behind using stacking by class-wise MLPs sharing weights. Isn't the deeper NN for multi-label classification is basically almost works just like that after all?</p>",
      "rawMarkdown": "Very nice solution.. I am sorry for the error happened but look a it as a priceless learning opportunities. Next comp you will be stronger..\n\nOne question: what is the intuition behind using stacking by class-wise MLPs sharing weights. Isn't the deeper NN for multi-label classification is basically almost works just like that after all?",
      "votes": null
    },
    {
      "id": "552458",
      "postDate": "06/14/2019 02:20:09",
      "content": "<p>Very interesting!</p>\n\n<p>You said:\n&gt; When stacking, I have 9(model) * 1103(class) features. If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN.</p>\n\n<p>How about using 2D convnet + linear dense layer, just like any other traditional NN. 2D input will not lose dimensional info, and you can leverage that by regarding X: classes , Y: predictions of the models..</p>",
      "rawMarkdown": "Very interesting!\n\nYou said:\n&gt; When stacking, I have 9(model) * 1103(class) features. If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN.\n\nHow about using 2D convnet + linear dense layer, just like any other traditional NN. 2D input will not lose dimensional info, and you can leverage that by regarding X: classes , Y: predictions of the models..",
      "votes": null
    },
    {
      "id": "552491",
      "postDate": "06/14/2019 03:45:00",
      "content": "<p>I want to get a medal in the next competition using this experience. Thanks !</p>",
      "rawMarkdown": "I want to get a medal in the next competition using this experience. Thanks !",
      "votes": null
    },
    {
      "id": "552643",
      "postDate": "06/14/2019 09:02:02",
      "content": "<p><a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/95225#550467\">Host implies that we might be able to make late submission someday.</a> If that, I'd like to make it and share the kernel. Thanks :)</p>",
      "rawMarkdown": "[Host implies that we might be able to make late submission someday.](https://www.kaggle.com/c/imet-2019-fgvc6/discussion/95225#550467) If that, I'd like to make it and share the kernel. Thanks :)",
      "votes": null
    },
    {
      "id": "552885",
      "postDate": "06/14/2019 18:35:57",
      "content": "<p>Nice solution! Sorry to see you also had problems in stage 2 :(</p>",
      "rawMarkdown": "Nice solution! Sorry to see you also had problems in stage 2 :(",
      "votes": null
    },
    {
      "id": "553093",
      "postDate": "06/15/2019 05:34:36",
      "content": "<p>At that time, I applied the class-wise MLPs as the extension of weighted averaging of which weights are shared among the all classes.</p>\n\n<p>I tried different depth:\n* having no hidden layer (equivalent to weighted averaging)\n* having one hidden layer (used in my solution)\n* having two hidden layers\n* having three hidden layers</p>\n\n<p>In my experiment, MLPs having one hidden layer performed good. </p>",
      "rawMarkdown": "At that time, I applied the class-wise MLPs as the extension of weighted averaging of which weights are shared among the all classes.\n\nI tried different depth:\n* having no hidden layer (equivalent to weighted averaging)\n* having one hidden layer (used in my solution)\n* having two hidden layers\n* having three hidden layers\n\nIn my experiment, MLPs having one hidden layer performed good.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 550710,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "06/12/2019 00:40:36",
      "content": "<p>Thank you for sharing! GNN stacking looks interesting. I'm not familiar with GNN, but is there some theoretical advantage of GNN for handling multiple models' outputs? Even if not strictly proved, it's ok for me. If you have any insights, it would be nice if you could share it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550711,
      "author_name": "owruby",
      "author_url": "",
      "post_date": "06/12/2019 00:41:34",
      "content": "<p>Thank you for sharing the solution.\nIn the next competition, I look forward to competing with you again!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551458,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/12/2019 18:38:13",
          "content": "<p>Thanks! See you in the next competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550714,
      "author_name": "triwave33",
      "author_url": "",
      "post_date": "06/12/2019 00:44:44",
      "content": "<p>Very impressive and interesting way for using GNN from classwise feature!\nIs this familiar way? (as you cited the paper)\nAnd what information did you expect to extract from this (GNN)?\nThank you for sharing : )</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550752,
      "author_name": "codingcliff",
      "author_url": "",
      "post_date": "06/12/2019 02:11:42",
      "content": "<p>Nice presentation! Thank you for your contribution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551126,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/12/2019 11:16:29",
          "content": "<p>Thank you. I'm very grad to share my idea in this research competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550770,
      "author_name": "maxwell110",
      "author_url": "",
      "post_date": "06/12/2019 02:36:44",
      "content": "<p><a href=\"/ttahara\">@ttahara</a> \nThank you for sharing this nice solution! <br>\nAlthough I didn't have participated in this competition, your solution is interesting for me.   </p>\n\n<p>Just for checking my understanding, one question.  </p>\n\n<p>In your figure, GNN looks an Octagram ( something like a black magic ! ). <br>\n<a href=\"https://en.wikipedia.org/wiki/Octagram\">https://en.wikipedia.org/wiki/Octagram</a> <br>\n<img src=\"https://cdn-ak.f.st-hatena.com/images/fotolife/g/greenwind120170/20190612/20190612113324.jpg\" alt=\"Black Magic\"></p>\n\n<p>But if I'm right, GNN has the edges whose numbers equals to the kinds of classes, right? <br>\nSorry for a trivial question.  </p>\n\n<p>Thanks in advance.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 551118,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/12/2019 11:12:11",
          "content": "<p>Yes, I used a black magic (stacking) by a magic circle (GNN), ha-ha!</p>\n\n<p>Your understanding is right. Each class's node is connected to all the other nodes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 551257,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "06/12/2019 14:07:01",
          "content": "<p>Thanks!\nWell clarified!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550799,
      "author_name": "haradataman",
      "author_url": "",
      "post_date": "06/12/2019 03:28:34",
      "content": "<p>This is the most elegant solution!!! Thank you for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551120,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/12/2019 11:12:46",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550815,
      "author_name": "yoshitaka1105",
      "author_url": "",
      "post_date": "06/12/2019 04:04:09",
      "content": "<p>I am so sorry that you are not at Private Leader Board...\nBut your solution is very incredible! I will find something useful for a future competition.\nThank you for sharing your solution!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551128,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/12/2019 11:23:03",
          "content": "<p>Thanks!\nI'm very grad If my idea gives you some inspiration !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551042,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "06/12/2019 09:26:31",
      "content": "<p>I'm interested in what software did you use to draw these pictures? Could you share about that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 551123,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/12/2019 11:14:46",
          "content": "<p>These figures were created by Microsoft PowerPoint.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 551133,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "06/12/2019 11:31:17",
          "content": "<p>nice, thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551115,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "06/12/2019 11:05:45",
      "content": "<p><a href=\"/sishihara\">@sishihara</a> <a href=\"/triwave33\">@triwave33</a> </p>\n\n<p>Thank you for question about GNN !</p>\n\n<p>&gt;  theoretical advantage </p>\n\n<p>Sorry, I'm not familiar with GNN, too. I cannot explain any theoretical advantage.\nBut I share my intuition and thought about applying GNN stacking to multi-label(class) classification in this comment.</p>\n\n<p>&gt; Is this familiar way?</p>\n\n<p>After reading the paper(<a href=\"https://arxiv.org/abs/1902.09720\">Learning a Deep ConvNet for Multi-label Classification with Partial Labels</a>), I searched other papers which assign CNN outputs for each class to nodes of GNN. But I've not found.\nMoreover, I've not found stacking model  by GNN.</p>\n\n<p>FYI: the following papers may be related.\n* <a href=\"http://openaccess.thecvf.com/content_iccv_2017/html/Qi_3D_Graph_Neural_ICCV_2017_paper.html\">3D Graph Neural Networks for RGBD Semantic Segmentation</a></p>\n\n<p>This paper assigns pixels to nodes.</p>\n\n<ul>\n<li><a href=\"https://arxiv.org/abs/1902.09720\">Multi-Label Image Recognition with Graph Convolutional Networks</a></li>\n</ul>\n\n<p>This paper assigns classes to nodes. But node features are word embedding, not CNN outputs.</p>\n\n<p>&gt; what information did you expect to extract from this (GNN)?</p>\n\n<p>IMO, GNN stacking have two advantages,</p>\n\n<ul>\n<li>It deals with each class's <em>packed</em> feature.</li>\n</ul>\n\n<p>When stacking, I have 9(model) * 1103(class) features.  If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN. </p>\n\n<ul>\n<li>(may be,) It can extract dependency of classes</li>\n</ul>\n\n<p>I think GNN can extract relationships between classes such as \"when class A is labeled, class B is also Labeled\". \nHowever,  for simple implementation, <strong>update functions of this solution's GNN are shared among all classes.</strong>\nSo, I don't understand extracted features in this solution well. </p>\n\n<p>Very sorry for poor English...\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551191,
          "author_name": "triwave33",
          "author_url": "",
          "post_date": "06/12/2019 12:40:51",
          "content": "<p>Thank you for sharing additional information and citations. I’ll do read them.</p>\n\n<p>From your explanation, I understand the merit of GNN in multiclass competition such that 1) feeding and treating data with retaining their structure and 2) getting dependency between classes.</p>\n\n<p>I’m very very looking forward to your further implementation and getting insight from that. But I believe your idea and solution must be brilliant even now. Thank you again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 551633,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/13/2019 00:50:30",
          "content": "<p>Your summarization is so good ! Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 552458,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "06/14/2019 02:20:09",
          "content": "<p>Very interesting!</p>\n\n<p>You said:\n&gt; When stacking, I have 9(model) * 1103(class) features. If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN.</p>\n\n<p>How about using 2D convnet + linear dense layer, just like any other traditional NN. 2D input will not lose dimensional info, and you can leverage that by regarding X: classes , Y: predictions of the models..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551153,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "06/12/2019 11:51:46",
      "content": "<p>Sorry to hear that. \nAnyway both your solution and write-up are beautiful and great. \nI have no doubt you will smash next competitions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 552491,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/14/2019 03:45:00",
          "content": "<p>I want to get a medal in the next competition using this experience. Thanks !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551205,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "06/12/2019 13:04:48",
      "content": "<p>Many thanks for your comprehensive yet beautiful solutions! Sorry to hear that you missed private LB. would you mind to share your source code solutions? I'm interested in your GNN integrations with other models! :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 552643,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/14/2019 09:02:02",
          "content": "<p><a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/95225#550467\">Host implies that we might be able to make late submission someday.</a> If that, I'd like to make it and share the kernel. Thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551780,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "06/13/2019 05:30:44",
      "content": "<p>I also fail in stage2, sad, too...\nreally impressed by the GNN part, nice work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 552450,
      "author_name": "hwasiti",
      "author_url": "",
      "post_date": "06/14/2019 02:13:01",
      "content": "<p>Very nice solution.. I am sorry for the error happened but look a it as a priceless learning opportunities. Next comp you will be stronger..</p>\n\n<p>One question: what is the intuition behind using stacking by class-wise MLPs sharing weights. Isn't the deeper NN for multi-label classification is basically almost works just like that after all?</p>",
      "votes": null,
      "replies": [
        {
          "id": 553093,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "06/15/2019 05:34:36",
          "content": "<p>At that time, I applied the class-wise MLPs as the extension of weighted averaging of which weights are shared among the all classes.</p>\n\n<p>I tried different depth:\n* having no hidden layer (equivalent to weighted averaging)\n* having one hidden layer (used in my solution)\n* having two hidden layers\n* having three hidden layers</p>\n\n<p>In my experiment, MLPs having one hidden layer performed good. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 552885,
      "author_name": "artyomp",
      "author_url": "",
      "post_date": "06/14/2019 18:35:57",
      "content": "<p>Nice solution! Sorry to see you also had problems in stage 2 :(</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "550701": "Congrats to  all the teams got the medal, and thanks to Kaggle and hosting team for this  interesting competition !\n\nUnfortunately, I made a mistake and lost my place in leaderboard, I'm so sad.  \nBut I really enjoyed competing with you kagglers for this 2 months :)\n\nHere, I'd like to share my solution.\n\n## Basic idea\n\nThis competition's  dataset includes various size images. So, I trained multiple models which input sizes have **different aspect ratios**, then made ensemble of them.\n\n## level 1 (pure classification models)\n\n### **Model**\nSE-ResNeXt50(1 model) and SE-ResNeXt101(8 model).  Global average pooling  layer is followed by FC, ReLU, Dropout and FC.\n\n### **Data Augmentation**\nI applied [random_distort](https://chainercv.readthedocs.io/en/stable/reference/links/ssd.html#random-distort) (brightness, contrast, saturation, hue), horizontal flip, random crop, random erase, and pca lighting.\n\n### **Loss**\n\nI used [FocalLoss](https://arxiv.org/abs/1708.02002) with **alpha=0.75** and gamma=1.5. This alpha setting is because [F2 score weights **recall** higher than precision](https://www.kaggle.com/c/imet-2019-fgvc6/overview/evaluation).\n\n### **Optimization**\n* Optimizer : SGD + NesterovAG\n    * momentum = 0.9, weight decay=1e-04\n* learning rate: differs by model. These are decided by [LR-Range Test](https://arxiv.org/abs/1506.01186).\n* batch size:  differs by model.\n* epoch: 20\n* learning schedule: cosine anealing (one cycle) with linear warmup(4 epoch)\n\n### Threshold Decision\nAs I wrote in [this kernel](https://www.kaggle.com/ttahara/eda-compare-number-of-culture-and-tag-attributes), there is difference between number of culture attributes and one of tag attributes each art has. \nThen, I prepared two separated thresholds for culture attribute and tag attribute.\n\n### summary of each model\n\n|     model      |image size  |aspect ratio| learning rate |   5foldcv    | Public | \n|:--------------:|:----------:|:----------:|:-------------:|:------------:|:------:|\n| SE-ResNeXt-50  |  288 x 288 | 1 : 1      |   2.15e-02    |    0.6132    |  0.637 |\n| SE-ResNeXt-101 |  368 x 368 | 1 : 1      |   1.8e-02     |    0.6233    |    -   |\n| SE-ResNeXt-101 |  448 x 448 | 1 : 1      |   1.5e-02     |    0.6243    |  0.647 |\n| SE-ResNeXt-101 |  368 x 560 | 2 : 3      |   1.3e-02     |    0.6228    |    -   |\n| SE-ResNeXt-101 |  560 x 368 | 3 : 2      |   1.3e-02     |    0.6232    |    -   |\n| SE-ResNeXt-101 |  320 x 640 | 1 : 2      |   1.2e-02     |    0.6190    |    -   |\n| SE-ResNeXt-101 |  640 x 320 | 2 : 1      |   1.2e-02     |    0.6197    |    -   |\n| SE-ResNeXt-101 |  256 x 784 | 1 : 3      |   1.8e-02     |    0.6127    |    -   |\n| SE-ResNeXt-101 |  784 x 256 | 3 : 1      |   1.8e-02     |    0.6147    |    -   |\n\n  \nWhen finished training all the models,  there were only a few days left for me. So I got  public scores of only 288x288 and 448x448.\n\n## level2 (ensemble models)\nIn the last few days, I tried several ensemble method. I explain them in chronological order.\n\n### 0. [CV:0.6456, Public:0.652] averaging (2days to go)\n\nSimply averaging probabirities output by 9 models.\n  \n  \n![solution  2 days to go](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_2days_to_go.png)\n  \n  \n### 1. [CV:0.6473, Public:0.654] stacking by _class-wise_ MLPs shaing weights(1 days to go): \n  \nI applied a MLP (which have one hidden layer) to multiple models' outputs for each class **separately**. These MLPs share their weights.\n  \n  \n![solution 1 day to go](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_1day_to_go.png)\n  \n  \n### 2. [CV:0.6483, Public:0.656] stacking by _GNN_ and class-wise MLPs shaing weights(the last day)\n\nEach _class-wise_ MLP, mentioned above, **cannot consider** outputs for other classes.  Got an inspiration from [this paper](https://arxiv.org/abs/1902.09720), I tried applying GNN (Graph Neural Network).\nThis paper applied GNN for **single** model's outputs for each class, however, I applied it for **multiple** models' outputs.\n\n\nLogits for each class obtained from level1 models are assigned to one node. After several iterations of updating messages and hidden states, each class node's hidden state is feeded into the class-wise MLP.\n\n![solution the last day](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_last_day.png)\n\n**_Note_**: This GNN stacking model is incomplete because **the update functions for messages and hidden states are shared among all classes**, for simple implementation (I've wanted to revise, but the deadline came earlier).  This means **all edges have the same weight**. I'll improve this model and try it in next other multi-class classification competition.\n\n### 3. [CV:0.6487, Public:0.657] stacking by **GNN** and class-wise MLPs shaing weights _with image size info_(the additional day)\n\nAdditionally, I used meta features of images related with image size, as I made ensemble of models whose input have various aspect ratios. This slightly improved my score.  \n\n![solution additional day](https://raw.githubusercontent.com/tawatawara/imet-2019-collection-FGVC6/master/solution_additional_day.png)\n  \n  \n  \n  \n  \nSorry for my poor english, and feel free to ask me if you have any questions.\n\nThank you for reading!",
    "550710": "Thank you for sharing! GNN stacking looks interesting. I'm not familiar with GNN, but is there some theoretical advantage of GNN for handling multiple models' outputs? Even if not strictly proved, it's ok for me. If you have any insights, it would be nice if you could share it.",
    "550711": "Thank you for sharing the solution.\nIn the next competition, I look forward to competing with you again!!",
    "550714": "Very impressive and interesting way for using GNN from classwise feature!\nIs this familiar way? (as you cited the paper)\nAnd what information did you expect to extract from this (GNN)?\nThank you for sharing : )",
    "550752": "Nice presentation! Thank you for your contribution!",
    "550770": "ttahara \nThank you for sharing this nice solution!  \nAlthough I didn't have participated in this competition, your solution is interesting for me.   \n  \nJust for checking my understanding, one question.  \n  \nIn your figure, GNN looks an Octagram ( something like a black magic ! ).  \n[https://en.wikipedia.org/wiki/Octagram](https://en.wikipedia.org/wiki/Octagram)  \n![Black Magic](https://cdn-ak.f.st-hatena.com/images/fotolife/g/greenwind120170/20190612/20190612113324.jpg)\n\nBut if I'm right, GNN has the edges whose numbers equals to the kinds of classes, right?  \nSorry for a trivial question.  \n\nThanks in advance.",
    "550799": "This is the most elegant solution!!! Thank you for sharing!",
    "550815": "I am so sorry that you are not at Private Leader Board...\nBut your solution is very incredible! I will find something useful for a future competition.\nThank you for sharing your solution!!",
    "551042": "I'm interested in what software did you use to draw these pictures? Could you share about that?",
    "551115": "sishihara @triwave33 \n\nThank you for question about GNN !\n\n&gt;  theoretical advantage \n\nSorry, I'm not familiar with GNN, too. I cannot explain any theoretical advantage.\nBut I share my intuition and thought about applying GNN stacking to multi-label(class) classification in this comment.\n\n&gt; Is this familiar way?\n\nAfter reading the paper([Learning a Deep ConvNet for Multi-label Classification with Partial Labels](https://arxiv.org/abs/1902.09720)), I searched other papers which assign CNN outputs for each class to nodes of GNN. But I've not found.\nMoreover, I've not found stacking model  by GNN.\n\nFYI: the following papers may be related.\n* [3D Graph Neural Networks for RGBD Semantic Segmentation](http://openaccess.thecvf.com/content_iccv_2017/html/Qi_3D_Graph_Neural_ICCV_2017_paper.html)\n\nThis paper assigns pixels to nodes.\n\n* [Multi-Label Image Recognition with Graph Convolutional Networks](https://arxiv.org/abs/1902.09720)\n\nThis paper assigns classes to nodes. But node features are word embedding, not CNN outputs.\n\n\n&gt; what information did you expect to extract from this (GNN)?\n  \nIMO, GNN stacking have two advantages,\n  \n*  It deals with each class's _packed_ feature.\n  \nWhen stacking, I have 9(model) * 1103(class) features.  If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN. \n  \n* (may be,) It can extract dependency of classes\n  \nI think GNN can extract relationships between classes such as \"when class A is labeled, class B is also Labeled\". \nHowever,  for simple implementation, **update functions of this solution's GNN are shared among all classes.**\nSo, I don't understand extracted features in this solution well. \n\nVery sorry for poor English...\nThanks!",
    "551118": "Yes, I used a black magic (stacking) by a magic circle (GNN), ha-ha!\n\nYour understanding is right. Each class's node is connected to all the other nodes.",
    "551120": "Thanks!",
    "551123": "These figures were created by Microsoft PowerPoint.",
    "551126": "Thank you. I'm very grad to share my idea in this research competition!",
    "551128": "Thanks!\nI'm very grad If my idea gives you some inspiration !",
    "551133": "nice, thank you.",
    "551153": "Sorry to hear that. \nAnyway both your solution and write-up are beautiful and great. \nI have no doubt you will smash next competitions.",
    "551191": "Thank you for sharing additional information and citations. I’ll do read them.\n\nFrom your explanation, I understand the merit of GNN in multiclass competition such that 1) feeding and treating data with retaining their structure and 2) getting dependency between classes.\n\nI’m very very looking forward to your further implementation and getting insight from that. But I believe your idea and solution must be brilliant even now. Thank you again!",
    "551205": "Many thanks for your comprehensive yet beautiful solutions! Sorry to hear that you missed private LB. would you mind to share your source code solutions? I'm interested in your GNN integrations with other models! :-)",
    "551257": "Thanks!\nWell clarified!",
    "551458": "Thanks! See you in the next competition.",
    "551633": "Your summarization is so good ! Thank you.",
    "551780": "I also fail in stage2, sad, too...\nreally impressed by the GNN part, nice work.",
    "552450": "Very nice solution.. I am sorry for the error happened but look a it as a priceless learning opportunities. Next comp you will be stronger..\n\nOne question: what is the intuition behind using stacking by class-wise MLPs sharing weights. Isn't the deeper NN for multi-label classification is basically almost works just like that after all?",
    "552458": "Very interesting!\n\nYou said:\n&gt; When stacking, I have 9(model) * 1103(class) features. If flatten and directory input them into a model (e,g, lightGBM), the information that features belong to the same class group is lost. As I hesitated to do this , tried using GNN.\n\nHow about using 2D convnet + linear dense layer, just like any other traditional NN. 2D input will not lose dimensional info, and you can leverage that by regarding X: classes , Y: predictions of the models..",
    "552491": "I want to get a medal in the next competition using this experience. Thanks !",
    "552643": "[Host implies that we might be able to make late submission someday.](https://www.kaggle.com/c/imet-2019-fgvc6/discussion/95225#550467) If that, I'd like to make it and share the kernel. Thanks :)",
    "552885": "Nice solution! Sorry to see you also had problems in stage 2 :(",
    "553093": "At that time, I applied the class-wise MLPs as the extension of weighted averaging of which weights are shared among the all classes.\n\nI tried different depth:\n* having no hidden layer (equivalent to weighted averaging)\n* having one hidden layer (used in my solution)\n* having two hidden layers\n* having three hidden layers\n\nIn my experiment, MLPs having one hidden layer performed good."
  },
  "source": "meta"
}