{
  "id": 90173,
  "title": "What could be the next step?",
  "url": "/competitions/imet-2019-fgvc6/discussion/90173",
  "author_name": "",
  "post_date": "2019-04-21T12:21:01.334085600Z",
  "votes": 18,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I have tried different models for this competition(ResNet50, DenseNet121, SEResNet) and rewrote the ResNet50, however, as a newbie, I'm stuck now. What could be the next step I could do? Can someone help me and give me any guidance?</p>",
  "messages": [
    {
      "id": "520611",
      "postDate": "04/21/2019 12:21:01",
      "content": "<p>I have tried different models for this competition(ResNet50, DenseNet121, SEResNet) and rewrote the ResNet50, however, as a newbie, I'm stuck now. What could be the next step I could do? Can someone help me and give me any guidance?</p>",
      "rawMarkdown": "I have tried different models for this competition(ResNet50, DenseNet121, SEResNet) and rewrote the ResNet50, however, as a newbie, I'm stuck now. What could be the next step I could do? Can someone help me and give me any guidance?",
      "votes": null
    },
    {
      "id": "520632",
      "postDate": "04/21/2019 13:03:45",
      "content": "<p>First, I guess that you can do more EDA (explore data analysis) and find more wierd things, which is really help.\n<a href=\"https://www.kaggle.com/chewzy/eda-weird-images-with-new-updates\">this kernel</a> is a very good example. So, the next step after EDA is to find the way to use your discovery. For example, take the size of images and the relationship between labels into consideration. And do better data augmentation or preprocessing. <strong>Rethinking the loss function</strong> is also a important method to get good grades. They all depend on a good EDA.</p>\n\n<p>Second,  you should build better CV(cross validation). It is helpful to balance distributions of multilabel data across splits for cross validation. You can reference <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819\">this</a>.  \"A good CV is half of success. I won’t go to the next step if I can’t find a good way to evaluate my model.\" bestfitting said this, which is very insight. you can read <a href=\"http://blog.kaggle.com/2018/05/07/profiling-top-kagglers-bestfitting-currently-1-in-the-world/\">Profiling Top Kagglers: Bestfitting, Currently #1 in the World</a> .</p>\n\n<p>Third, you can use attention model or other networks to do fine-grain classification. You really need network structures based on <a href=\"https://www.kaggle.com/seefun/you-really-need-attention-pytorch\">attention</a>. And you should also read some papers on the few-shot learning.</p>\n\n<p>Last, you can try more training tricks, like better weight initialization, better optimizer, better learning rate scheduler and so on.  And more inference tricks, like <a href=\"https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it\">TTA(test time augmentation)</a> and model ensemble. <a href=\"https://www.kaggle.com/seefun/ensemble-top-3-models-in-public-kernels-0-9781\">This</a> is a good example of the power of ensemble. </p>\n\n<p>If you have any questions，feel free to ask me. GLHF in Kaggle!</p>",
      "rawMarkdown": "First, I guess that you can do more EDA (explore data analysis) and find more wierd things, which is really help.\n[this kernel](https://www.kaggle.com/chewzy/eda-weird-images-with-new-updates) is a very good example. So, the next step after EDA is to find the way to use your discovery. For example, take the size of images and the relationship between labels into consideration. And do better data augmentation or preprocessing. **Rethinking the loss function** is also a important method to get good grades. They all depend on a good EDA.\n\nSecond,  you should build better CV(cross validation). It is helpful to balance distributions of multilabel data across splits for cross validation. You can reference [this](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819).  \"A good CV is half of success. I won’t go to the next step if I can’t find a good way to evaluate my model.\" bestfitting said this, which is very insight. you can read [Profiling Top Kagglers: Bestfitting, Currently #1 in the World](http://blog.kaggle.com/2018/05/07/profiling-top-kagglers-bestfitting-currently-1-in-the-world/) .\n\nThird, you can use attention model or other networks to do fine-grain classification. You really need network structures based on [attention](https://www.kaggle.com/seefun/you-really-need-attention-pytorch). And you should also read some papers on the few-shot learning.\n\nLast, you can try more training tricks, like better weight initialization, better optimizer, better learning rate scheduler and so on.  And more inference tricks, like [TTA(test time augmentation)](https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it) and model ensemble. [This](https://www.kaggle.com/seefun/ensemble-top-3-models-in-public-kernels-0-9781) is a good example of the power of ensemble. \n\nIf you have any questions，feel free to ask me. GLHF in Kaggle!",
      "votes": null
    },
    {
      "id": "520639",
      "postDate": "04/21/2019 13:14:30",
      "content": "<p>If you encounter bottleneck，go back to step1, rethinking the data and your EDA.</p>",
      "rawMarkdown": "If you encounter bottleneck，go back to step1, rethinking the data and your EDA.",
      "votes": null
    },
    {
      "id": "520640",
      "postDate": "04/21/2019 13:17:02",
      "content": "<p>Thanks for your elaborate guidance!!! Wish you have a great score in this competition.</p>",
      "rawMarkdown": "Thanks for your elaborate guidance!!! Wish you have a great score in this competition.",
      "votes": null
    },
    {
      "id": "520710",
      "postDate": "04/21/2019 16:09:22",
      "content": "<p>Just to add, you can use this package <a href=\"https://github.com/trent-b/iterative-stratification\">Iterative stratification </a> to help making your cv splits. </p>\n\n<p>It was used quite successfully by many in the hpa competition. Just convert your target to hot vectors to use it. </p>\n\n<p>I actually found that quite confusing as a total noob during hpa, and I think I might have discovered the least efficient way possible to do that. But it’s not hard, if you, or anyone, would find it helpful I can post a little snippet on how to turn the attribute ids into a vector to use with this package. </p>",
      "rawMarkdown": "Just to add, you can use this package [Iterative stratification ](https://github.com/trent-b/iterative-stratification) to help making your cv splits. \n\nIt was used quite successfully by many in the hpa competition. Just convert your target to hot vectors to use it. \n\nI actually found that quite confusing as a total noob during hpa, and I think I might have discovered the least efficient way possible to do that. But it’s not hard, if you, or anyone, would find it helpful I can post a little snippet on how to turn the attribute ids into a vector to use with this package.",
      "votes": null
    },
    {
      "id": "520929",
      "postDate": "04/22/2019 02:22:52",
      "content": "<p>sklearn.preprocessing.MultiLabelBinarizer</p>",
      "rawMarkdown": "sklearn.preprocessing.MultiLabelBinarizer",
      "votes": null
    },
    {
      "id": "520948",
      "postDate": "04/22/2019 03:39:19",
      "content": "<p>For sure, faster than np.eye</p>\n\n<p>But just for fun</p>\n\n<pre><code>y = [[int(i) for i in s.split()] for s in train_df['attribute_ids']]\noh = np.zeros((len(y), 1103))\nfor i in tqdm(range(len(y))):\n    oh[i] = np.eye((1103),dtype=np.float)[y[i]].sum(axis=0)\n\nfrom iterstrat.ml_stratifiers import MultilabelStratifiedKFold\nimport numpy as np\n\nX = np.array(train_df['id'])\ny = oh #np.array(train_df['attribute_ids'])\n\nmskf = MultilabelStratifiedKFold(n_splits=2, random_state=0)\n\nfor train_index, test_index in mskf.split(X, y):\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    X_train, X_test = X[train_index], X[test_index]\n    y_train, y_test = y[train_index], y[test_index]\n</code></pre>",
      "rawMarkdown": "For sure, faster than np.eye\n\nBut just for fun\n\n    y = [[int(i) for i in s.split()] for s in train_df['attribute_ids']]\n    oh = np.zeros((len(y), 1103))\n    for i in tqdm(range(len(y))):\n        oh[i] = np.eye((1103),dtype=np.float)[y[i]].sum(axis=0)\n    \n    from iterstrat.ml_stratifiers import MultilabelStratifiedKFold\n    import numpy as np\n\n    X = np.array(train_df['id'])\n    y = oh #np.array(train_df['attribute_ids'])\n\n    mskf = MultilabelStratifiedKFold(n_splits=2, random_state=0)\n\n    for train_index, test_index in mskf.split(X, y):\n        print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n        X_train, X_test = X[train_index], X[test_index]\n        y_train, y_test = y[train_index], y[test_index]",
      "votes": null
    },
    {
      "id": "520962",
      "postDate": "04/22/2019 04:35:44",
      "content": "<p>For the detail of the last step, training tricks. You can refer to this <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/90250\">discussion</a>.</p>",
      "rawMarkdown": "For the detail of the last step, training tricks. You can refer to this [discussion](https://www.kaggle.com/c/imet-2019-fgvc6/discussion/90250).",
      "votes": null
    },
    {
      "id": "524272",
      "postDate": "04/28/2019 11:16:09",
      "content": "<p>Many thanks</p>",
      "rawMarkdown": "Many thanks",
      "votes": null
    },
    {
      "id": "530178",
      "postDate": "05/12/2019 02:18:26",
      "content": "<p>wondering how much it will boost by using correct stratified kfold?</p>",
      "rawMarkdown": "wondering how much it will boost by using correct stratified kfold?",
      "votes": null
    },
    {
      "id": "536548",
      "postDate": "05/24/2019 16:39:36",
      "content": "<p>Hi, I find the image labels is imbalanced in the dataset. May I ask, do you use undersampling or oversampling？And is the weight in loss function useful? Thank you.</p>",
      "rawMarkdown": "Hi, I find the image labels is imbalanced in the dataset. May I ask, do you use undersampling or oversampling？And is the weight in loss function useful? Thank you.",
      "votes": null
    },
    {
      "id": "536722",
      "postDate": "05/25/2019 04:54:26",
      "content": "<p>Many people use focal loss or other loss function to deal with the imbalanced dataset. And sorry for that, I don't understand what do you mean by ''undersample or oversample'' ?</p>",
      "rawMarkdown": "Many people use focal loss or other loss function to deal with the imbalanced dataset. And sorry for that, I don't understand what do you mean by ''undersample or oversample'' ?",
      "votes": null
    },
    {
      "id": "536742",
      "postDate": "05/25/2019 06:38:03",
      "content": "<p><a href=\"https://en.wikipedia.org/wiki/Oversampling_and_undersampling_in_data_analysis\">Oversampling and undersampling</a>. I have tried the focal loss. But it did not work well. The categorical_crossentropy is better than focal loss in my model. Sorry for wrongly written word.</p>",
      "rawMarkdown": "[Oversampling and undersampling](https://en.wikipedia.org/wiki/Oversampling_and_undersampling_in_data_analysis). I have tried the focal loss. But it did not work well. The categorical_crossentropy is better than focal loss in my model. Sorry for wrongly written word.",
      "votes": null
    },
    {
      "id": "539048",
      "postDate": "05/29/2019 12:46:42",
      "content": "<p>focal loss is really good! It really works in my code!</p>",
      "rawMarkdown": "focal loss is really good! It really works in my code!",
      "votes": null
    },
    {
      "id": "539413",
      "postDate": "05/30/2019 03:17:52",
      "content": "<p>focal loss does not work in my se-resnext101(pytorch), but it work in my densenet121(keras). I am wondering how to do that same thing in se-resnext101. May I ask, would you share you kernel after the competition. Thanks a lot!</p>",
      "rawMarkdown": "focal loss does not work in my se-resnext101(pytorch), but it work in my densenet121(keras). I am wondering how to do that same thing in se-resnext101. May I ask, would you share you kernel after the competition. Thanks a lot!",
      "votes": null
    },
    {
      "id": "539438",
      "postDate": "05/30/2019 03:39:08",
      "content": "<p>Interesting, focal loss never works for me...</p>",
      "rawMarkdown": "Interesting, focal loss never works for me...",
      "votes": null
    },
    {
      "id": "539485",
      "postDate": "05/30/2019 05:30:53",
      "content": "<p>I will be glad to do that if all of my teammates agree.</p>",
      "rawMarkdown": "I will be glad to do that if all of my teammates agree.",
      "votes": null
    },
    {
      "id": "539839",
      "postDate": "05/30/2019 15:35:29",
      "content": "<p>Why not share codes after competition? We both need to learn a lot.</p>",
      "rawMarkdown": "Why not share codes after competition? We both need to learn a lot.",
      "votes": null
    },
    {
      "id": "540129",
      "postDate": "05/31/2019 04:37:26",
      "content": "<p>Focal loss doesn't work for me with seresnext101 in Pytorch.</p>",
      "rawMarkdown": "Focal loss doesn't work for me with seresnext101 in Pytorch.",
      "votes": null
    },
    {
      "id": "540256",
      "postDate": "05/31/2019 08:36:40",
      "content": "<p>In my experiment，Focal loss can give you a slight increase in scores but it accelerated convergence rate.</p>",
      "rawMarkdown": "In my experiment，Focal loss can give you a slight increase in scores but it accelerated convergence rate.",
      "votes": null
    },
    {
      "id": "554164",
      "postDate": "06/17/2019 04:54:30",
      "content": "<p>Make one of our training kernel public, <a href=\"https://www.kaggle.com/jionie/training-fold-0-ibnresnext101-lovasz-focal\">https://www.kaggle.com/jionie/training-fold-0-ibnresnext101-lovasz-focal</a>. Could finish within 9h and get around 0.606 cv for one single fold.</p>",
      "rawMarkdown": "Make one of our training kernel public, https://www.kaggle.com/jionie/training-fold-0-ibnresnext101-lovasz-focal. Could finish within 9h and get around 0.606 cv for one single fold.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 520632,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "04/21/2019 13:03:45",
      "content": "<p>First, I guess that you can do more EDA (explore data analysis) and find more wierd things, which is really help.\n<a href=\"https://www.kaggle.com/chewzy/eda-weird-images-with-new-updates\">this kernel</a> is a very good example. So, the next step after EDA is to find the way to use your discovery. For example, take the size of images and the relationship between labels into consideration. And do better data augmentation or preprocessing. <strong>Rethinking the loss function</strong> is also a important method to get good grades. They all depend on a good EDA.</p>\n\n<p>Second,  you should build better CV(cross validation). It is helpful to balance distributions of multilabel data across splits for cross validation. You can reference <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819\">this</a>.  \"A good CV is half of success. I won’t go to the next step if I can’t find a good way to evaluate my model.\" bestfitting said this, which is very insight. you can read <a href=\"http://blog.kaggle.com/2018/05/07/profiling-top-kagglers-bestfitting-currently-1-in-the-world/\">Profiling Top Kagglers: Bestfitting, Currently #1 in the World</a> .</p>\n\n<p>Third, you can use attention model or other networks to do fine-grain classification. You really need network structures based on <a href=\"https://www.kaggle.com/seefun/you-really-need-attention-pytorch\">attention</a>. And you should also read some papers on the few-shot learning.</p>\n\n<p>Last, you can try more training tricks, like better weight initialization, better optimizer, better learning rate scheduler and so on.  And more inference tricks, like <a href=\"https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it\">TTA(test time augmentation)</a> and model ensemble. <a href=\"https://www.kaggle.com/seefun/ensemble-top-3-models-in-public-kernels-0-9781\">This</a> is a good example of the power of ensemble. </p>\n\n<p>If you have any questions，feel free to ask me. GLHF in Kaggle!</p>",
      "votes": null,
      "replies": [
        {
          "id": 520639,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "04/21/2019 13:14:30",
          "content": "<p>If you encounter bottleneck，go back to step1, rethinking the data and your EDA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520640,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "04/21/2019 13:17:02",
          "content": "<p>Thanks for your elaborate guidance!!! Wish you have a great score in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520962,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "04/22/2019 04:35:44",
          "content": "<p>For the detail of the last step, training tricks. You can refer to this <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/90250\">discussion</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520710,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "04/21/2019 16:09:22",
      "content": "<p>Just to add, you can use this package <a href=\"https://github.com/trent-b/iterative-stratification\">Iterative stratification </a> to help making your cv splits. </p>\n\n<p>It was used quite successfully by many in the hpa competition. Just convert your target to hot vectors to use it. </p>\n\n<p>I actually found that quite confusing as a total noob during hpa, and I think I might have discovered the least efficient way possible to do that. But it’s not hard, if you, or anyone, would find it helpful I can post a little snippet on how to turn the attribute ids into a vector to use with this package. </p>",
      "votes": null,
      "replies": [
        {
          "id": 520929,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "04/22/2019 02:22:52",
          "content": "<p>sklearn.preprocessing.MultiLabelBinarizer</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520948,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/22/2019 03:39:19",
          "content": "<p>For sure, faster than np.eye</p>\n\n<p>But just for fun</p>\n\n<pre><code>y = [[int(i) for i in s.split()] for s in train_df['attribute_ids']]\noh = np.zeros((len(y), 1103))\nfor i in tqdm(range(len(y))):\n    oh[i] = np.eye((1103),dtype=np.float)[y[i]].sum(axis=0)\n\nfrom iterstrat.ml_stratifiers import MultilabelStratifiedKFold\nimport numpy as np\n\nX = np.array(train_df['id'])\ny = oh #np.array(train_df['attribute_ids'])\n\nmskf = MultilabelStratifiedKFold(n_splits=2, random_state=0)\n\nfor train_index, test_index in mskf.split(X, y):\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    X_train, X_test = X[train_index], X[test_index]\n    y_train, y_test = y[train_index], y[test_index]\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524272,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "04/28/2019 11:16:09",
          "content": "<p>Many thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530178,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "05/12/2019 02:18:26",
          "content": "<p>wondering how much it will boost by using correct stratified kfold?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 536548,
      "author_name": "saladjay",
      "author_url": "",
      "post_date": "05/24/2019 16:39:36",
      "content": "<p>Hi, I find the image labels is imbalanced in the dataset. May I ask, do you use undersampling or oversampling？And is the weight in loss function useful? Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 536722,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "05/25/2019 04:54:26",
          "content": "<p>Many people use focal loss or other loss function to deal with the imbalanced dataset. And sorry for that, I don't understand what do you mean by ''undersample or oversample'' ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536742,
          "author_name": "saladjay",
          "author_url": "",
          "post_date": "05/25/2019 06:38:03",
          "content": "<p><a href=\"https://en.wikipedia.org/wiki/Oversampling_and_undersampling_in_data_analysis\">Oversampling and undersampling</a>. I have tried the focal loss. But it did not work well. The categorical_crossentropy is better than focal loss in my model. Sorry for wrongly written word.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539048,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "05/29/2019 12:46:42",
          "content": "<p>focal loss is really good! It really works in my code!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539413,
          "author_name": "saladjay",
          "author_url": "",
          "post_date": "05/30/2019 03:17:52",
          "content": "<p>focal loss does not work in my se-resnext101(pytorch), but it work in my densenet121(keras). I am wondering how to do that same thing in se-resnext101. May I ask, would you share you kernel after the competition. Thanks a lot!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539438,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "05/30/2019 03:39:08",
          "content": "<p>Interesting, focal loss never works for me...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539485,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "05/30/2019 05:30:53",
          "content": "<p>I will be glad to do that if all of my teammates agree.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539839,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "05/30/2019 15:35:29",
          "content": "<p>Why not share codes after competition? We both need to learn a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540129,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "05/31/2019 04:37:26",
          "content": "<p>Focal loss doesn't work for me with seresnext101 in Pytorch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540256,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "05/31/2019 08:36:40",
          "content": "<p>In my experiment，Focal loss can give you a slight increase in scores but it accelerated convergence rate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 554164,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "06/17/2019 04:54:30",
          "content": "<p>Make one of our training kernel public, <a href=\"https://www.kaggle.com/jionie/training-fold-0-ibnresnext101-lovasz-focal\">https://www.kaggle.com/jionie/training-fold-0-ibnresnext101-lovasz-focal</a>. Could finish within 9h and get around 0.606 cv for one single fold.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "520611": "I have tried different models for this competition(ResNet50, DenseNet121, SEResNet) and rewrote the ResNet50, however, as a newbie, I'm stuck now. What could be the next step I could do? Can someone help me and give me any guidance?",
    "520632": "First, I guess that you can do more EDA (explore data analysis) and find more wierd things, which is really help.\n[this kernel](https://www.kaggle.com/chewzy/eda-weird-images-with-new-updates) is a very good example. So, the next step after EDA is to find the way to use your discovery. For example, take the size of images and the relationship between labels into consideration. And do better data augmentation or preprocessing. **Rethinking the loss function** is also a important method to get good grades. They all depend on a good EDA.\n\nSecond,  you should build better CV(cross validation). It is helpful to balance distributions of multilabel data across splits for cross validation. You can reference [this](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819).  \"A good CV is half of success. I won’t go to the next step if I can’t find a good way to evaluate my model.\" bestfitting said this, which is very insight. you can read [Profiling Top Kagglers: Bestfitting, Currently #1 in the World](http://blog.kaggle.com/2018/05/07/profiling-top-kagglers-bestfitting-currently-1-in-the-world/) .\n\nThird, you can use attention model or other networks to do fine-grain classification. You really need network structures based on [attention](https://www.kaggle.com/seefun/you-really-need-attention-pytorch). And you should also read some papers on the few-shot learning.\n\nLast, you can try more training tricks, like better weight initialization, better optimizer, better learning rate scheduler and so on.  And more inference tricks, like [TTA(test time augmentation)](https://www.kaggle.com/andrewkh/test-time-augmentation-tta-worth-it) and model ensemble. [This](https://www.kaggle.com/seefun/ensemble-top-3-models-in-public-kernels-0-9781) is a good example of the power of ensemble. \n\nIf you have any questions，feel free to ask me. GLHF in Kaggle!",
    "520639": "If you encounter bottleneck，go back to step1, rethinking the data and your EDA.",
    "520640": "Thanks for your elaborate guidance!!! Wish you have a great score in this competition.",
    "520710": "Just to add, you can use this package [Iterative stratification ](https://github.com/trent-b/iterative-stratification) to help making your cv splits. \n\nIt was used quite successfully by many in the hpa competition. Just convert your target to hot vectors to use it. \n\nI actually found that quite confusing as a total noob during hpa, and I think I might have discovered the least efficient way possible to do that. But it’s not hard, if you, or anyone, would find it helpful I can post a little snippet on how to turn the attribute ids into a vector to use with this package.",
    "520929": "sklearn.preprocessing.MultiLabelBinarizer",
    "520948": "For sure, faster than np.eye\n\nBut just for fun\n\n    y = [[int(i) for i in s.split()] for s in train_df['attribute_ids']]\n    oh = np.zeros((len(y), 1103))\n    for i in tqdm(range(len(y))):\n        oh[i] = np.eye((1103),dtype=np.float)[y[i]].sum(axis=0)\n    \n    from iterstrat.ml_stratifiers import MultilabelStratifiedKFold\n    import numpy as np\n\n    X = np.array(train_df['id'])\n    y = oh #np.array(train_df['attribute_ids'])\n\n    mskf = MultilabelStratifiedKFold(n_splits=2, random_state=0)\n\n    for train_index, test_index in mskf.split(X, y):\n        print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n        X_train, X_test = X[train_index], X[test_index]\n        y_train, y_test = y[train_index], y[test_index]",
    "520962": "For the detail of the last step, training tricks. You can refer to this [discussion](https://www.kaggle.com/c/imet-2019-fgvc6/discussion/90250).",
    "524272": "Many thanks",
    "530178": "wondering how much it will boost by using correct stratified kfold?",
    "536548": "Hi, I find the image labels is imbalanced in the dataset. May I ask, do you use undersampling or oversampling？And is the weight in loss function useful? Thank you.",
    "536722": "Many people use focal loss or other loss function to deal with the imbalanced dataset. And sorry for that, I don't understand what do you mean by ''undersample or oversample'' ?",
    "536742": "[Oversampling and undersampling](https://en.wikipedia.org/wiki/Oversampling_and_undersampling_in_data_analysis). I have tried the focal loss. But it did not work well. The categorical_crossentropy is better than focal loss in my model. Sorry for wrongly written word.",
    "539048": "focal loss is really good! It really works in my code!",
    "539413": "focal loss does not work in my se-resnext101(pytorch), but it work in my densenet121(keras). I am wondering how to do that same thing in se-resnext101. May I ask, would you share you kernel after the competition. Thanks a lot!",
    "539438": "Interesting, focal loss never works for me...",
    "539485": "I will be glad to do that if all of my teammates agree.",
    "539839": "Why not share codes after competition? We both need to learn a lot.",
    "540129": "Focal loss doesn't work for me with seresnext101 in Pytorch.",
    "540256": "In my experiment，Focal loss can give you a slight increase in scores but it accelerated convergence rate.",
    "554164": "Make one of our training kernel public, https://www.kaggle.com/jionie/training-fold-0-ibnresnext101-lovasz-focal. Could finish within 9h and get around 0.606 cv for one single fold."
  },
  "source": "meta"
}