{
  "id": 155167,
  "title": "Yet another starter [training + inference]",
  "url": "/competitions/alaska2-image-steganalysis/discussion/155167",
  "author_name": "",
  "post_date": "2020-05-31T15:02:11.056779100Z",
  "votes": 23,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>Thank you all authors for creating baseline kernels.\nI would like to share with you, my friends, training pipeline too. I have used PyTorch GPU, and:</p>\n\n<ul>\n<li>4 Classes</li>\n<li>GroupKFold splitting</li>\n<li>Class Balance</li>\n<li>Flips</li>\n<li>Label Smoothing</li>\n<li>EfficientNetB2</li>\n<li>ReduceLROnPlateau</li>\n</ul>\n\n<p>Hope it helps you for \"jump start\" ;)</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">[Train + Inference] GPU Baseline</a></p>\n\n<p>Welcome!</p>",
  "messages": [
    {
      "id": "868901",
      "postDate": "05/31/2020 15:02:11",
      "content": "<p>Hi everyone!</p>\n\n<p>Thank you all authors for creating baseline kernels.\nI would like to share with you, my friends, training pipeline too. I have used PyTorch GPU, and:</p>\n\n<ul>\n<li>4 Classes</li>\n<li>GroupKFold splitting</li>\n<li>Class Balance</li>\n<li>Flips</li>\n<li>Label Smoothing</li>\n<li>EfficientNetB2</li>\n<li>ReduceLROnPlateau</li>\n</ul>\n\n<p>Hope it helps you for \"jump start\" ;)</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">[Train + Inference] GPU Baseline</a></p>\n\n<p>Welcome!</p>",
      "rawMarkdown": "Hi everyone!\n\nThank you all authors for creating baseline kernels.\nI would like to share with you, my friends, training pipeline too. I have used PyTorch GPU, and:\n\n- 4 Classes\n- GroupKFold splitting\n- Class Balance\n- Flips\n- Label Smoothing\n- EfficientNetB2\n- ReduceLROnPlateau\n\nHope it helps you for \"jump start\" ;)\n\n[[Train + Inference] GPU Baseline](https://www.kaggle.com/shonenkov/train-inference-gpu-baseline)\n\nWelcome!",
      "votes": null
    },
    {
      "id": "869093",
      "postDate": "05/31/2020 18:25:21",
      "content": "<p>Have you tried different parameters for label smoothing? For me the best eps is 0.1. And question for all: Does splitting into 9 classes instead of 4 really improve validation score? Since for now i am using only 4 classes. </p>",
      "rawMarkdown": "Have you tried different parameters for label smoothing? For me the best eps is 0.1. And question for all: Does splitting into 9 classes instead of 4 really improve validation score? Since for now i am using only 4 classes.",
      "votes": null
    },
    {
      "id": "869105",
      "postDate": "05/31/2020 18:34:38",
      "content": "<p>No, I didn't use other values of eps yet :) thank you for sharing information: will try eps 0.1 also</p>",
      "rawMarkdown": "No, I didn't use other values of eps yet :) thank you for sharing information: will try eps 0.1 also",
      "votes": null
    },
    {
      "id": "869141",
      "postDate": "05/31/2020 18:55:26",
      "content": "<p>and i am not sure about disk quota here, but if you have an opportunity, you can covert images to np.arrays and read them directly, it speeds up learning about 30% since the bottleneck for me was the DataLoader class.</p>",
      "rawMarkdown": "and i am not sure about disk quota here, but if you have an opportunity, you can covert images to np.arrays and read them directly, it speeds up learning about 30% since the bottleneck for me was the DataLoader class.",
      "votes": null
    },
    {
      "id": "869146",
      "postDate": "05/31/2020 18:59:27",
      "content": "<p>Thank you. I am afraid here (in kaggle) not enough RAM for this trick ;)</p>",
      "rawMarkdown": "Thank you. I am afraid here (in kaggle) not enough RAM for this trick ;)",
      "votes": null
    },
    {
      "id": "869147",
      "postDate": "05/31/2020 19:03:13",
      "content": "<p>i am talking about disk quota) since np.load works 10+ times faster than cv2.imread. Just resave images as np.array:)</p>",
      "rawMarkdown": "i am talking about disk quota) since np.load works 10+ times faster than cv2.imread. Just resave images as np.array:)",
      "votes": null
    },
    {
      "id": "869151",
      "postDate": "05/31/2020 19:10:06",
      "content": "<p>Aa, it is right :) we can create public datasets only &lt; 20GB, so we can split by several datasets and use npy per image. But I wouldn't like to solve this competition with kaggle gpu quota. In my machine simple way is creating big npy matrix and using RAM if need speed. </p>",
      "rawMarkdown": "Aa, it is right :) we can create public datasets only &lt; 20GB, so we can split by several datasets and use npy per image. But I wouldn't like to solve this competition with kaggle gpu quota. In my machine simple way is creating big npy matrix and using RAM if need speed.",
      "votes": null
    },
    {
      "id": "869154",
      "postDate": "05/31/2020 19:17:19",
      "content": "<p>what amount of RAM do you have? 0_o</p>",
      "rawMarkdown": "what amount of RAM do you have? 0_o",
      "votes": null
    },
    {
      "id": "869163",
      "postDate": "05/31/2020 19:27:34",
      "content": "<p>enough for loading 300k npy images 512x512 (~220gb) and 256x256 (only ~55 gb), but most of the tasks here can be solved using colab pro, but not this :(</p>",
      "rawMarkdown": "enough for loading 300k npy images 512x512 (~220gb) and 256x256 (only ~55 gb), but most of the tasks here can be solved using colab pro, but not this :(",
      "votes": null
    },
    {
      "id": "869940",
      "postDate": "06/01/2020 11:42:43",
      "content": "<p>Thanks for nice sharing.\nI have one question.\nWhich main idea do you think affects the high score?\nThere are many interesting approaches such as 4 Classes, Class Balance, Label Smoothing.</p>",
      "rawMarkdown": "Thanks for nice sharing.\nI have one question.\nWhich main idea do you think affects the high score?\nThere are many interesting approaches such as 4 Classes, Class Balance, Label Smoothing.",
      "votes": null
    },
    {
      "id": "869963",
      "postDate": "06/01/2020 12:05:09",
      "content": "<p>In any case I have added original (that used for training) \"csv\" with my GroupKFold splitting in <a href=\"https://www.kaggle.com/shonenkov/alaska2-public-baseline\">dataset</a> and update logs for 35 epoch</p>",
      "rawMarkdown": "In any case I have added original (that used for training) \"csv\" with my GroupKFold splitting in [dataset](https://www.kaggle.com/shonenkov/alaska2-public-baseline) and update logs for 35 epoch",
      "votes": null
    },
    {
      "id": "869981",
      "postDate": "06/01/2020 12:15:45",
      "content": "<p>thank you! I haven't studied deep other solutions. 4 classes is very important when compared with binary for me</p>",
      "rawMarkdown": "thank you! I haven't studied deep other solutions. 4 classes is very important when compared with binary for me",
      "votes": null
    },
    {
      "id": "870041",
      "postDate": "06/01/2020 13:00:29",
      "content": "<p>You are on fire!!!</p>",
      "rawMarkdown": "You are on fire!!!",
      "votes": null
    },
    {
      "id": "870043",
      "postDate": "06/01/2020 13:01:06",
      "content": "<p>Thank you!\nDataset is large and takes a long time to train.\nI wanted to know the core part.</p>",
      "rawMarkdown": "Thank you!\nDataset is large and takes a long time to train.\nI wanted to know the core part.",
      "votes": null
    },
    {
      "id": "872361",
      "postDate": "06/03/2020 06:01:56",
      "content": "<p>Thx Alex, it looks like there are jpeg files with different 3 quality factors in each folders, have you consider splitting each algorithms into 3 classes? That would make it total of 9 or 12 classes.</p>",
      "rawMarkdown": "Thx Alex, it looks like there are jpeg files with different 3 quality factors in each folders, have you consider splitting each algorithms into 3 classes? That would make it total of 9 or 12 classes.",
      "votes": null
    },
    {
      "id": "1118852",
      "postDate": "12/19/2020 13:08:31",
      "content": "<p>Alex, first of all, thanks for great notebook! <br>\nBut i have some troubles with yours weights (which you uploaded, after 23 and 33 epochs)<br>\nWith your weights, your dataset (holded on yours csv file), on yours validation fold (0) i have alaska-roc-auc 0.69<br>\nBut you had 0.92<br>\nDo you have any suggestions why this could happen?<br>\nI haven't changed absolutely anything, maybe you loaded the wrong weights?</p>",
      "rawMarkdown": "Alex, first of all, thanks for great notebook! \nBut i have some troubles with yours weights (which you uploaded, after 23 and 33 epochs)\nWith your weights, your dataset (holded on yours csv file), on yours validation fold (0) i have alaska-roc-auc 0.69\nBut you had 0.92\nDo you have any suggestions why this could happen?\nI haven't changed absolutely anything, maybe you loaded the wrong weights?",
      "votes": null
    },
    {
      "id": "1119049",
      "postDate": "12/19/2020 17:10:59",
      "content": "<p>Update!  Error fixed by downgrading efficientnet_pytorch on version 0.6.3</p>",
      "rawMarkdown": "Update!  Error fixed by downgrading efficientnet_pytorch on version 0.6.3",
      "votes": null
    },
    {
      "id": "1121179",
      "postDate": "12/21/2020 12:33:39",
      "content": "<p>Hi , I am aslo working on this recent, Have you tried training this at least a single epoch ?</p>",
      "rawMarkdown": "Hi , I am aslo working on this recent, Have you tried training this at least a single epoch ?",
      "votes": null
    },
    {
      "id": "1121195",
      "postDate": "12/21/2020 12:44:57",
      "content": "<p>No, i used weights only. But I'm sure everything will be ok with the correct library versions</p>",
      "rawMarkdown": "No, i used weights only. But I'm sure everything will be ok with the correct library versions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 869093,
      "author_name": "vovanf98",
      "author_url": "",
      "post_date": "05/31/2020 18:25:21",
      "content": "<p>Have you tried different parameters for label smoothing? For me the best eps is 0.1. And question for all: Does splitting into 9 classes instead of 4 really improve validation score? Since for now i am using only 4 classes. </p>",
      "votes": null,
      "replies": [
        {
          "id": 869105,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "05/31/2020 18:34:38",
          "content": "<p>No, I didn't use other values of eps yet :) thank you for sharing information: will try eps 0.1 also</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869141,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/31/2020 18:55:26",
          "content": "<p>and i am not sure about disk quota here, but if you have an opportunity, you can covert images to np.arrays and read them directly, it speeds up learning about 30% since the bottleneck for me was the DataLoader class.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869146,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "05/31/2020 18:59:27",
          "content": "<p>Thank you. I am afraid here (in kaggle) not enough RAM for this trick ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869147,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/31/2020 19:03:13",
          "content": "<p>i am talking about disk quota) since np.load works 10+ times faster than cv2.imread. Just resave images as np.array:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869151,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "05/31/2020 19:10:06",
          "content": "<p>Aa, it is right :) we can create public datasets only &lt; 20GB, so we can split by several datasets and use npy per image. But I wouldn't like to solve this competition with kaggle gpu quota. In my machine simple way is creating big npy matrix and using RAM if need speed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869154,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/31/2020 19:17:19",
          "content": "<p>what amount of RAM do you have? 0_o</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869163,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "05/31/2020 19:27:34",
          "content": "<p>enough for loading 300k npy images 512x512 (~220gb) and 256x256 (only ~55 gb), but most of the tasks here can be solved using colab pro, but not this :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 869940,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "06/01/2020 11:42:43",
      "content": "<p>Thanks for nice sharing.\nI have one question.\nWhich main idea do you think affects the high score?\nThere are many interesting approaches such as 4 Classes, Class Balance, Label Smoothing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 869981,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "06/01/2020 12:15:45",
          "content": "<p>thank you! I haven't studied deep other solutions. 4 classes is very important when compared with binary for me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 870043,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "06/01/2020 13:01:06",
          "content": "<p>Thank you!\nDataset is large and takes a long time to train.\nI wanted to know the core part.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 869963,
      "author_name": "shonenkov",
      "author_url": "",
      "post_date": "06/01/2020 12:05:09",
      "content": "<p>In any case I have added original (that used for training) \"csv\" with my GroupKFold splitting in <a href=\"https://www.kaggle.com/shonenkov/alaska2-public-baseline\">dataset</a> and update logs for 35 epoch</p>",
      "votes": null,
      "replies": [
        {
          "id": 1118852,
          "author_name": "pixml333",
          "author_url": "",
          "post_date": "12/19/2020 13:08:31",
          "content": "<p>Alex, first of all, thanks for great notebook! <br>\nBut i have some troubles with yours weights (which you uploaded, after 23 and 33 epochs)<br>\nWith your weights, your dataset (holded on yours csv file), on yours validation fold (0) i have alaska-roc-auc 0.69<br>\nBut you had 0.92<br>\nDo you have any suggestions why this could happen?<br>\nI haven't changed absolutely anything, maybe you loaded the wrong weights?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119049,
          "author_name": "pixml333",
          "author_url": "",
          "post_date": "12/19/2020 17:10:59",
          "content": "<p>Update!  Error fixed by downgrading efficientnet_pytorch on version 0.6.3</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121179,
          "author_name": "avishkathushan",
          "author_url": "",
          "post_date": "12/21/2020 12:33:39",
          "content": "<p>Hi , I am aslo working on this recent, Have you tried training this at least a single epoch ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121195,
          "author_name": "pixml333",
          "author_url": "",
          "post_date": "12/21/2020 12:44:57",
          "content": "<p>No, i used weights only. But I'm sure everything will be ok with the correct library versions</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 870041,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "06/01/2020 13:00:29",
      "content": "<p>You are on fire!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 872361,
      "author_name": "vincentpoont2",
      "author_url": "",
      "post_date": "06/03/2020 06:01:56",
      "content": "<p>Thx Alex, it looks like there are jpeg files with different 3 quality factors in each folders, have you consider splitting each algorithms into 3 classes? That would make it total of 9 or 12 classes.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "868901": "Hi everyone!\n\nThank you all authors for creating baseline kernels.\nI would like to share with you, my friends, training pipeline too. I have used PyTorch GPU, and:\n\n- 4 Classes\n- GroupKFold splitting\n- Class Balance\n- Flips\n- Label Smoothing\n- EfficientNetB2\n- ReduceLROnPlateau\n\nHope it helps you for \"jump start\" ;)\n\n[[Train + Inference] GPU Baseline](https://www.kaggle.com/shonenkov/train-inference-gpu-baseline)\n\nWelcome!",
    "869093": "Have you tried different parameters for label smoothing? For me the best eps is 0.1. And question for all: Does splitting into 9 classes instead of 4 really improve validation score? Since for now i am using only 4 classes.",
    "869105": "No, I didn't use other values of eps yet :) thank you for sharing information: will try eps 0.1 also",
    "869141": "and i am not sure about disk quota here, but if you have an opportunity, you can covert images to np.arrays and read them directly, it speeds up learning about 30% since the bottleneck for me was the DataLoader class.",
    "869146": "Thank you. I am afraid here (in kaggle) not enough RAM for this trick ;)",
    "869147": "i am talking about disk quota) since np.load works 10+ times faster than cv2.imread. Just resave images as np.array:)",
    "869151": "Aa, it is right :) we can create public datasets only &lt; 20GB, so we can split by several datasets and use npy per image. But I wouldn't like to solve this competition with kaggle gpu quota. In my machine simple way is creating big npy matrix and using RAM if need speed.",
    "869154": "what amount of RAM do you have? 0_o",
    "869163": "enough for loading 300k npy images 512x512 (~220gb) and 256x256 (only ~55 gb), but most of the tasks here can be solved using colab pro, but not this :(",
    "869940": "Thanks for nice sharing.\nI have one question.\nWhich main idea do you think affects the high score?\nThere are many interesting approaches such as 4 Classes, Class Balance, Label Smoothing.",
    "869963": "In any case I have added original (that used for training) \"csv\" with my GroupKFold splitting in [dataset](https://www.kaggle.com/shonenkov/alaska2-public-baseline) and update logs for 35 epoch",
    "869981": "thank you! I haven't studied deep other solutions. 4 classes is very important when compared with binary for me",
    "870041": "You are on fire!!!",
    "870043": "Thank you!\nDataset is large and takes a long time to train.\nI wanted to know the core part.",
    "872361": "Thx Alex, it looks like there are jpeg files with different 3 quality factors in each folders, have you consider splitting each algorithms into 3 classes? That would make it total of 9 or 12 classes.",
    "1118852": "Alex, first of all, thanks for great notebook! \nBut i have some troubles with yours weights (which you uploaded, after 23 and 33 epochs)\nWith your weights, your dataset (holded on yours csv file), on yours validation fold (0) i have alaska-roc-auc 0.69\nBut you had 0.92\nDo you have any suggestions why this could happen?\nI haven't changed absolutely anything, maybe you loaded the wrong weights?",
    "1119049": "Update!  Error fixed by downgrading efficientnet_pytorch on version 0.6.3",
    "1121179": "Hi , I am aslo working on this recent, Have you tried training this at least a single epoch ?",
    "1121195": "No, i used weights only. But I'm sure everything will be ok with the correct library versions"
  },
  "source": "meta"
}