{
  "id": 77731,
  "title": "5th Place Solution",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/vpp-5th-place-solution",
  "author_name": "",
  "post_date": "2019-01-17T07:51:56.313Z",
  "votes": 31,
  "comment_count": 21,
  "views": 0,
  "content": "<p>First of all, congratulations to all the winners! Thanks to Kaggle and HPA team for hosting such an interesting competition and thanks to  <a href=\"https://www.kaggle.com/tomomimoriyama\">TomomiMoriyama</a>, <a href=\"https://www.kaggle.com/hengck23\">Heng CherKeng</a>, <a href=\"https://www.kaggle.com/manyfoldcv\">ManyFoldCV</a> and <a href=\"https://www.kaggle.com/spytensor\">Spytensor</a>.  </p>\n\n<p>Here is a brief summary of our solution.</p>\n\n<h3>DataSet</h3>\n\n<p>Like most other competitors, we used both official (both PNG and TIFF) and external data. To deal with class-imbalance, we used WeightedRandomSampler （method in pytorch） during training and <a href=\"https://github.com/trent-b/iterative-stratification\">MultilabelStratifiedShuffleSplit</a> to split the data into training and validation. We constructed 10 folds cross validation sets with 8% for validation.</p>\n\n<h3>Image Preprocessing</h3>\n\n<p>The HPA dataset has four dyeing modes each of which is an RGB image of its own, so we took only one channel （r=r,g=g,b=b,y=b） to form a 4-channel input for training.</p>\n\n<p>All PNG images are kept at their original 512 size, whereas the TIFF images are resized to 1024.</p>\n\n<h3>Augmentation</h3>\n\n<p>Rotation, Flip, and Shear.</p>\n\n<p>We didn't use random cropping. Instead we trained 5 models using crop5 (method in pytorch) and found it to be more effective.</p>\n\n<h3>Models</h3>\n\n<p>For our base networks, we mainly used Inception-v3，-v4, and Xception. We have also tried DenseNet, SENet and ResNet, but the results were suboptimal.</p>\n\n<p>We used three different scales during training (512 for PNG images and 650, 800 for TIFF images) with different random seeds for the 10-folds CV.</p>\n\n<p>Modifications</p>\n\n<ol>\n<li>Changed the last pooling layer to global pooling.</li>\n<li>Appended an additional fully connected layer with output dimension 128 after the global pooling.</li>\n<li>We also divided the training process into two stages where the first stage used size 512 with model trained on ImageNet, and the second stage used size 650 or 800 with model trained from the first stage. We found this to be slightly better than training with fixed size all the way.</li>\n</ol>\n\n<h3>Training</h3>\n\n<ul>\n<li>loss: <a href=\"https://pytorch.org/docs/stable/nn.html?highlight=multilabelsoftmarginloss#torch.nn.MultiLabelSoftMarginLoss\">MultiLabelSoftMarginLoss</a></li>\n<li>lr: 0.05 (for size 512, pretrained on ImageNet)，0.01 (for size 650 and 800，pretrained using size 512); lrscheduler: steplr(gamma=0.1,step=6)</li>\n<li>optimizer: SGD</li>\n<li>epochs: 25, early stopping for training with size 650 or 800 (around 15 epochs), model selected based on loss (instead of F1 score)</li>\n<li>sampling weights for different classes: [1.0, 5.97, 2.89, 5.75, 4.64, 4.27, 5.46, 3.2, 14.48, 14.84, 15.14, 6.92, 6.86, 8.12, 6.32, 19.24, 8.48, 11.93, 7.32, 5.48, 11.99, 2.39, 6.3, 3.0, 12.06, 1.0, 10.39, 16.5]</li>\n</ul>\n\n<h3>Multi-Thresholds</h3>\n\n<p>We used the validation sets to search for threshold for each class by optimizing the F1 score begining with 0.15 for all classes.</p>\n\n<h3>Test</h3>\n\n<p>(with multi-thresholds)</p>\n\n<p><img src=\"https://s2.ax1x.com/2019/01/17/kpZ7lV.png\" alt=\"\"></p>\n\n<h3>Ensembling</h3>\n\n<p>Final prediction is ensemble of above methods: Size 800, 10-fold for Inception-v3; Size 650 and 800, 10-fold for Inception-v4; Size 800, 10-fold, Size 650, 1-fold, Size 512, 5-fold for Xception (the reason for 5-fold instead of 10 was simply because we didn't have enough submissions to check the performances of all models, so we simply took the best ones).</p>\n\n<h3>Things that did not work for us</h3>\n\n<ul>\n<li>Training with larger input size (&gt;= 1024), which forced us reduce the batch size.</li>\n<li>3-channel input</li>\n<li>focal loss</li>\n<li>C3D</li>\n<li>TTA: unlike a lot of other competitors, TTA during test time actually didn't work for us.</li>\n<li>Other traditional machine learning methods such as DecisionTree, RandomForest, and SVM.</li>\n</ul>",
  "messages": [
    {
      "id": "456591",
      "postDate": "01/16/2019 06:17:37",
      "content": "<p>First of all, congratulations to all the winners! Thanks to Kaggle and HPA team for hosting such an interesting competition and thanks to  <a href=\"https://www.kaggle.com/tomomimoriyama\">TomomiMoriyama</a>, <a href=\"https://www.kaggle.com/hengck23\">Heng CherKeng</a>, <a href=\"https://www.kaggle.com/manyfoldcv\">ManyFoldCV</a> and <a href=\"https://www.kaggle.com/spytensor\">Spytensor</a>.  </p>\n\n<p>Here is a brief summary of our solution.</p>\n\n<h3>DataSet</h3>\n\n<p>Like most other competitors, we used both official (both PNG and TIFF) and external data. To deal with class-imbalance, we used WeightedRandomSampler （method in pytorch） during training and <a href=\"https://github.com/trent-b/iterative-stratification\">MultilabelStratifiedShuffleSplit</a> to split the data into training and validation. We constructed 10 folds cross validation sets with 8% for validation.</p>\n\n<h3>Image Preprocessing</h3>\n\n<p>The HPA dataset has four dyeing modes each of which is an RGB image of its own, so we took only one channel （r=r,g=g,b=b,y=b） to form a 4-channel input for training.</p>\n\n<p>All PNG images are kept at their original 512 size, whereas the TIFF images are resized to 1024.</p>\n\n<h3>Augmentation</h3>\n\n<p>Rotation, Flip, and Shear.</p>\n\n<p>We didn't use random cropping. Instead we trained 5 models using crop5 (method in pytorch) and found it to be more effective.</p>\n\n<h3>Models</h3>\n\n<p>For our base networks, we mainly used Inception-v3，-v4, and Xception. We have also tried DenseNet, SENet and ResNet, but the results were suboptimal.</p>\n\n<p>We used three different scales during training (512 for PNG images and 650, 800 for TIFF images) with different random seeds for the 10-folds CV.</p>\n\n<p>Modifications</p>\n\n<ol>\n<li>Changed the last pooling layer to global pooling.</li>\n<li>Appended an additional fully connected layer with output dimension 128 after the global pooling.</li>\n<li>We also divided the training process into two stages where the first stage used size 512 with model trained on ImageNet, and the second stage used size 650 or 800 with model trained from the first stage. We found this to be slightly better than training with fixed size all the way.</li>\n</ol>\n\n<h3>Training</h3>\n\n<ul>\n<li>loss: <a href=\"https://pytorch.org/docs/stable/nn.html?highlight=multilabelsoftmarginloss#torch.nn.MultiLabelSoftMarginLoss\">MultiLabelSoftMarginLoss</a></li>\n<li>lr: 0.05 (for size 512, pretrained on ImageNet)，0.01 (for size 650 and 800，pretrained using size 512); lrscheduler: steplr(gamma=0.1,step=6)</li>\n<li>optimizer: SGD</li>\n<li>epochs: 25, early stopping for training with size 650 or 800 (around 15 epochs), model selected based on loss (instead of F1 score)</li>\n<li>sampling weights for different classes: [1.0, 5.97, 2.89, 5.75, 4.64, 4.27, 5.46, 3.2, 14.48, 14.84, 15.14, 6.92, 6.86, 8.12, 6.32, 19.24, 8.48, 11.93, 7.32, 5.48, 11.99, 2.39, 6.3, 3.0, 12.06, 1.0, 10.39, 16.5]</li>\n</ul>\n\n<h3>Multi-Thresholds</h3>\n\n<p>We used the validation sets to search for threshold for each class by optimizing the F1 score begining with 0.15 for all classes.</p>\n\n<h3>Test</h3>\n\n<p>(with multi-thresholds)</p>\n\n<p><img src=\"https://s2.ax1x.com/2019/01/17/kpZ7lV.png\" alt=\"\"></p>\n\n<h3>Ensembling</h3>\n\n<p>Final prediction is ensemble of above methods: Size 800, 10-fold for Inception-v3; Size 650 and 800, 10-fold for Inception-v4; Size 800, 10-fold, Size 650, 1-fold, Size 512, 5-fold for Xception (the reason for 5-fold instead of 10 was simply because we didn't have enough submissions to check the performances of all models, so we simply took the best ones).</p>\n\n<h3>Things that did not work for us</h3>\n\n<ul>\n<li>Training with larger input size (&gt;= 1024), which forced us reduce the batch size.</li>\n<li>3-channel input</li>\n<li>focal loss</li>\n<li>C3D</li>\n<li>TTA: unlike a lot of other competitors, TTA during test time actually didn't work for us.</li>\n<li>Other traditional machine learning methods such as DecisionTree, RandomForest, and SVM.</li>\n</ul>",
      "rawMarkdown": "First of all, congratulations to all the winners! Thanks to Kaggle and HPA team for hosting such an interesting competition and thanks to  [TomomiMoriyama](https://www.kaggle.com/tomomimoriyama), [Heng CherKeng](https://www.kaggle.com/hengck23), [ManyFoldCV](https://www.kaggle.com/manyfoldcv) and [Spytensor](https://www.kaggle.com/spytensor).  \n\nHere is a brief summary of our solution.\n\n### DataSet\n\nLike most other competitors, we used both official (both PNG and TIFF) and external data. To deal with class-imbalance, we used WeightedRandomSampler （method in pytorch） during training and [MultilabelStratifiedShuffleSplit](https://github.com/trent-b/iterative-stratification) to split the data into training and validation. We constructed 10 folds cross validation sets with 8% for validation.\n\n### Image Preprocessing\n\nThe HPA dataset has four dyeing modes each of which is an RGB image of its own, so we took only one channel （r=r,g=g,b=b,y=b） to form a 4-channel input for training.\n\nAll PNG images are kept at their original 512 size, whereas the TIFF images are resized to 1024.\n\n### Augmentation\n\nRotation, Flip, and Shear.\n\nWe didn't use random cropping. Instead we trained 5 models using crop5 (method in pytorch) and found it to be more effective.\n\n### Models\n\nFor our base networks, we mainly used Inception-v3，-v4, and Xception. We have also tried DenseNet, SENet and ResNet, but the results were suboptimal.\n\nWe used three different scales during training (512 for PNG images and 650, 800 for TIFF images) with different random seeds for the 10-folds CV.\n\nModifications\n\n1. Changed the last pooling layer to global pooling.\n2. Appended an additional fully connected layer with output dimension 128 after the global pooling.\n3. We also divided the training process into two stages where the first stage used size 512 with model trained on ImageNet, and the second stage used size 650 or 800 with model trained from the first stage. We found this to be slightly better than training with fixed size all the way.\n\n### Training\n\n* loss: [MultiLabelSoftMarginLoss](https://pytorch.org/docs/stable/nn.html?highlight=multilabelsoftmarginloss#torch.nn.MultiLabelSoftMarginLoss)\n* lr: 0.05 (for size 512, pretrained on ImageNet)，0.01 (for size 650 and 800，pretrained using size 512); lrscheduler: steplr(gamma=0.1,step=6)\n* optimizer: SGD\n* epochs: 25, early stopping for training with size 650 or 800 (around 15 epochs), model selected based on loss (instead of F1 score)\n* sampling weights for different classes: [1.0, 5.97, 2.89, 5.75, 4.64, 4.27, 5.46, 3.2, 14.48, 14.84, 15.14, 6.92, 6.86, 8.12, 6.32, 19.24, 8.48, 11.93, 7.32, 5.48, 11.99, 2.39, 6.3, 3.0, 12.06, 1.0, 10.39, 16.5]\n  \n### Multi-Thresholds\n\nWe used the validation sets to search for threshold for each class by optimizing the F1 score begining with 0.15 for all classes.\n\n\n### Test\n\n(with multi-thresholds)\n\n![](https://s2.ax1x.com/2019/01/17/kpZ7lV.png)\n\n### Ensembling\n\nFinal prediction is ensemble of above methods: Size 800, 10-fold for Inception-v3; Size 650 and 800, 10-fold for Inception-v4; Size 800, 10-fold, Size 650, 1-fold, Size 512, 5-fold for Xception (the reason for 5-fold instead of 10 was simply because we didn't have enough submissions to check the performances of all models, so we simply took the best ones).\n\n### Things that did not work for us\n* Training with larger input size (&gt;= 1024), which forced us reduce the batch size.\n* 3-channel input\n* focal loss\n* C3D\n* TTA: unlike a lot of other competitors, TTA during test time actually didn't work for us.\n* Other traditional machine learning methods such as DecisionTree, RandomForest, and SVM.",
      "votes": null
    },
    {
      "id": "456859",
      "postDate": "01/16/2019 17:04:39",
      "content": "<p>hi,thanks for posting the solution..\n1) could you explain what is scale here...\n2) When i tried to used weights for classes then it was resulting into too high val loss,did you also find similar things for early epochs ?</p>",
      "rawMarkdown": "hi,thanks for posting the solution..\n1) could you explain what is scale here...\n2) When i tried to used weights for classes then it was resulting into too high val loss,did you also find similar things for early epochs ?",
      "votes": null
    },
    {
      "id": "457304",
      "postDate": "01/17/2019 07:51:14",
      "content": "<ol>\n<li>Scale means the size of the input image(base size is 512x512).</li>\n<li>I didn't find the situation you said.  （BN and dropout after full connection，SGD,  BCE or MultiLabelSoftMarginLoss,  reasonable learning rate and lrscheduler may be helpful)</li>\n</ol>",
      "rawMarkdown": "1. Scale means the size of the input image(base size is 512x512).\n2. I didn't find the situation you said.  （BN and dropout after full connection，SGD,  BCE or MultiLabelSoftMarginLoss,  reasonable learning rate and lrscheduler may be helpful)",
      "votes": null
    },
    {
      "id": "457433",
      "postDate": "01/17/2019 12:22:14",
      "content": "<p>THanks\n1) how did u do ensembling ,any easy example .  say model 1 prediction probs [0.99.....0.11] ,model2 [0.8....0.5]\n2) did you  also use external data,if yes could u please provide the script</p>",
      "rawMarkdown": "THanks\n1) how did u do ensembling ,any easy example .  say model 1 prediction probs [0.99.....0.11] ,model2 [0.8....0.5]\n2) did you  also use external data,if yes could u please provide the script",
      "votes": null
    },
    {
      "id": "457673",
      "postDate": "01/17/2019 22:57:15",
      "content": "<p>You must have nice hardware. I had Ti1080x2 and, for size 768, mt batch size was 8.</p>",
      "rawMarkdown": "You must have nice hardware. I had Ti1080x2 and, for size 768, mt batch size was 8.",
      "votes": null
    },
    {
      "id": "457854",
      "postDate": "01/18/2019 07:58:25",
      "content": "<ol>\n<li>We simply did weighted arithmetic mean on sigmoids.  </li>\n<li>Yes, we used the external data. The following is our script. <br>\n<code>\nimport os\nimport pandas as pd\nimport cv2\nimport multiprocessing\nimport tqdm\nimport requests\ncolors = ['red', 'green', 'blue', 'yellow']\nBASE_DATASET_PATH = './external_data/HPAv18/'\nDIR = BASE_DATASET_PATH + \"jpg/\"\nGray_DIR = BASE_DATASET_PATH + \"rgby_1024_png/\"\nv18_url = 'http://v18.proteinatlas.org/images/'\nif not os.path.exists(Gray_DIR):\nos.mkdir(Gray_DIR)\ndef download_img(item_name):\nimg = item_name.split('_')\nfor color in colors:\n    img_path = img[0] + '/' + \"_\".join(img[1:]) + \"_\" + color + \".jpg\"\n    img_name = item_name + \"_\" + color + \".jpg\"\n    img_url = v18_url + img_path\n    try:\n        r = requests.get(img_url, allow_redirects=True)\n        open(DIR + img_name, 'wb').write(r.content)\n    except Exception, e:\n        print e\n        print 'Error,{},{}'.format(img_url,img_name)\ndef rgb_to_gray(item_name):\nfor color in colors:\n    img_name = item_name + \"_\" + color + \".jpg\"\n    img_path = DIR + img_name\n    save_path = Gray_DIR + img_name[:-4] + '.png'\n    if os.path.exists(save_path):\n        continue\n    img = cv2.imread(img_path)\n    index = 0\n    if color == 'blue':\n        index = 0\n    elif color == 'green':\n        index = 1\n    elif color == 'red':\n        index = 2\n    elif color == 'yellow':\n        index = 1\n    img_gray = img[..., index]\n    if img_gray.shape[0] != 1024 or img_gray.shape[1] != 1024:\n        img_gray = cv2.resize(img_gray, (1024, 1024))\n    cv2.imwrite(save_path, img_gray)\nif __name__ == '__main__':\nimgList = pd.read_csv(BASE_DATASET_PATH + \"HPAv18RBGY_wodpl.csv\")\npool = multiprocessing.Pool(processes=50)\npool.map(download_img, imgList['Id'])\npBar = tqdm.tqdm(total=len(imgList))\nfor i, item_name in enumerate(imgList['Id']):\n    rgb_to_gray(item_name)\n    if i % 100 == 0:\n        pBar.update(100)\npBar.close()\n</code></li>\n</ol>",
      "rawMarkdown": "1. We simply did weighted arithmetic mean on sigmoids.  \n2. Yes, we used the external data. The following is our script.    \n```\nimport os\nimport pandas as pd\nimport cv2\nimport multiprocessing\nimport tqdm\nimport requests\ncolors = ['red', 'green', 'blue', 'yellow']\nBASE_DATASET_PATH = './external_data/HPAv18/'\nDIR = BASE_DATASET_PATH + \"jpg/\"\nGray_DIR = BASE_DATASET_PATH + \"rgby_1024_png/\"\nv18_url = 'http://v18.proteinatlas.org/images/'\nif not os.path.exists(Gray_DIR):\n    os.mkdir(Gray_DIR)\ndef download_img(item_name):\n    img = item_name.split('_')\n    for color in colors:\n        img_path = img[0] + '/' + \"_\".join(img[1:]) + \"_\" + color + \".jpg\"\n        img_name = item_name + \"_\" + color + \".jpg\"\n        img_url = v18_url + img_path\n        try:\n            r = requests.get(img_url, allow_redirects=True)\n            open(DIR + img_name, 'wb').write(r.content)\n        except Exception, e:\n            print e\n            print 'Error,{},{}'.format(img_url,img_name)\ndef rgb_to_gray(item_name):\n    for color in colors:\n        img_name = item_name + \"_\" + color + \".jpg\"\n        img_path = DIR + img_name\n        save_path = Gray_DIR + img_name[:-4] + '.png'\n        if os.path.exists(save_path):\n            continue\n        img = cv2.imread(img_path)\n        index = 0\n        if color == 'blue':\n            index = 0\n        elif color == 'green':\n            index = 1\n        elif color == 'red':\n            index = 2\n        elif color == 'yellow':\n            index = 1\n        img_gray = img[..., index]\n        if img_gray.shape[0] != 1024 or img_gray.shape[1] != 1024:\n            img_gray = cv2.resize(img_gray, (1024, 1024))\n        cv2.imwrite(save_path, img_gray)\nif __name__ == '__main__':\n    imgList = pd.read_csv(BASE_DATASET_PATH + \"HPAv18RBGY_wodpl.csv\")\n    pool = multiprocessing.Pool(processes=50)\n    pool.map(download_img, imgList['Id'])\n    pBar = tqdm.tqdm(total=len(imgList))\n    for i, item_name in enumerate(imgList['Id']):\n        rgb_to_gray(item_name)\n        if i % 100 == 0:\n            pBar.update(100)\n    pBar.close()\n```",
      "votes": null
    },
    {
      "id": "457864",
      "postDate": "01/18/2019 08:26:21",
      "content": "<p>HA HA! We had p100x4, batch size was 10 (inception-v4, 800, single GPU)</p>",
      "rawMarkdown": "HA HA! We had p100x4, batch size was 10 (inception-v4, 800, single GPU)",
      "votes": null
    },
    {
      "id": "458025",
      "postDate": "01/18/2019 15:49:42",
      "content": "<p>Thanku</p>\n\n<p>1) there were some posts which mentioned there are images in external hpa that are duplicates in kaggles train set ,how u filtered the duplicated data ??\n2) Are the images in the file HPAv18RBGY_wodpl.csv   different from the ones that shown and downloaded here   <a href=\"https://www.kaggle.com/artemtprv/load-external-data/data\">https://www.kaggle.com/artemtprv/load-external-data/data</a>\n3) how you appended the labels from the files,\n4 ) finaly do we had to use the different normalization for external hpa</p>",
      "rawMarkdown": "Thanku\n\n1) there were some posts which mentioned there are images in external hpa that are duplicates in kaggles train set ,how u filtered the duplicated data ??\n2) Are the images in the file HPAv18RBGY_wodpl.csv   different from the ones that shown and downloaded here   https://www.kaggle.com/artemtprv/load-external-data/data\n3) how you appended the labels from the files,\n4 ) finaly do we had to use the different normalization for external hpa",
      "votes": null
    },
    {
      "id": "458137",
      "postDate": "01/18/2019 21:44:56",
      "content": "<p>I was using resnet50, sz=768, one Ti1080 with batch size 8. When it failed at 16, I went directly to 8 so it might have been a little higher. With resnet34, I had batch size 16.</p>",
      "rawMarkdown": "I was using resnet50, sz=768, one Ti1080 with batch size 8. When it failed at 16, I went directly to 8 so it might have been a little higher. With resnet34, I had batch size 16.",
      "votes": null
    },
    {
      "id": "458865",
      "postDate": "01/20/2019 17:11:43",
      "content": "<p>I just used this file <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/432870/10816/HPAv18RBGY_wodpl.csv\">HPAv18RBGY_wodpl.csv</a> , added to the training set.  No other operation has been taken.</p>",
      "rawMarkdown": "I just used this file [HPAv18RBGY_wodpl.csv](https://storage.googleapis.com/kaggle-forum-message-attachments/432870/10816/HPAv18RBGY_wodpl.csv) , added to the training set.  No other operation has been taken.",
      "votes": null
    },
    {
      "id": "459133",
      "postDate": "01/21/2019 08:52:10",
      "content": "<p>thanks..\ndid you do any differential normalization i.e different normalization for HPA and different for competition, as some people mentioning that HPA data has got different pixel intensity than competition  </p>",
      "rawMarkdown": "thanks..\ndid you do any differential normalization i.e different normalization for HPA and different for competition, as some people mentioning that HPA data has got different pixel intensity than competition",
      "votes": null
    },
    {
      "id": "459861",
      "postDate": "01/22/2019 14:03:55",
      "content": "<p>Hi lingyundev \nI used HPA data along with regular data but i get high val losses. Note that i used only selective data based on class rarity and original image resize to 512 .Total image downloaded 32k.\nWhat could be the reason for high val losses.... any more processing is that needed before we can use HPA.</p>\n\n<p>Appreciate your input here. </p>",
      "rawMarkdown": "Hi lingyundev \nI used HPA data along with regular data but i get high val losses. Note that i used only selective data based on class rarity and original image resize to 512 .Total image downloaded 32k.\nWhat could be the reason for high val losses.... any more processing is that needed before we can use HPA.\n\nAppreciate your input here.",
      "votes": null
    },
    {
      "id": "460617",
      "postDate": "01/24/2019 05:08:39",
      "content": "<p>hi jaideep <br>\nDuring the competition, we merged the HPAv18 and official datasets directly.\ndid you tried the parameters in my solution? especially the loss, optimizer  and learning rate.</p>",
      "rawMarkdown": "hi jaideep  \nDuring the competition, we merged the HPAv18 and official datasets directly.\ndid you tried the parameters in my solution? especially the loss, optimizer  and learning rate.",
      "votes": null
    },
    {
      "id": "460635",
      "postDate": "01/24/2019 05:37:28",
      "content": "<p>Make sure you preprocess the external data correctly. Note that unlike the competition data, each image has 3 channels. If you read it as 1 color image, it will be very wrong. You need to read it as 3ch and save only the correct channel as grayimage</p>",
      "rawMarkdown": "Make sure you preprocess the external data correctly. Note that unlike the competition data, each image has 3 channels. If you read it as 1 color image, it will be very wrong. You need to read it as 3ch and save only the correct channel as grayimage",
      "votes": null
    },
    {
      "id": "460695",
      "postDate": "01/24/2019 08:57:13",
      "content": "<p>hi ling m yet to change the loss,i use focal loss the images i downloaded contained all the channels .</p>",
      "rawMarkdown": "hi ling m yet to change the loss,i use focal loss the images i downloaded contained all the channels .",
      "votes": null
    },
    {
      "id": "460698",
      "postDate": "01/24/2019 09:01:43",
      "content": "<p>hi moshel\nimages i downloaded looked to contain all the channels not just the RGB. it also had yellow channel data. so i dint do any thing after download ,just appended the list of downloaded images to regular image list.\nI used david silva script in official \n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a>\ndef download(pid, image_list, base_url, save_dir, image_size=(512, 512)):\n    colors = ['red', 'green', 'blue', 'yellow']\n    for i in tqdm(image_list, postfix=pid):\n        img_id = i.split('_', 1)\n        for color in colors:\n            img_path = img_id[0] + '/' + img_id[1] + '_' + color + '.jpg'\n            img_name = i + '_' + color + '.png'\n            img_url = base_url + img_path</p>\n\n<pre><code>        # Get the raw response from the url\n        r = requests.get(img_url, allow_redirects=True, stream=True)\n        r.raw.decode_content = True\n\n        # Use PIL to resize the image and to convert it to L\n        # (8-bit pixels, black and white)\n        im = Image.open(r.raw)\n        im = im.resize(image_size, Image.LANCZOS).convert('L')\n        im.save(os.path.join(save_dir, img_name), 'PNG')\n</code></pre>\n\n<p>let me know if there is any miss between download and use of image for training...</p>",
      "rawMarkdown": "hi moshel\nimages i downloaded looked to contain all the channels not just the RGB. it also had yellow channel data. so i dint do any thing after download ,just appended the list of downloaded images to regular image list.\nI used david silva script in official \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\ndef download(pid, image_list, base_url, save_dir, image_size=(512, 512)):\n    colors = ['red', 'green', 'blue', 'yellow']\n    for i in tqdm(image_list, postfix=pid):\n        img_id = i.split('_', 1)\n        for color in colors:\n            img_path = img_id[0] + '/' + img_id[1] + '_' + color + '.jpg'\n            img_name = i + '_' + color + '.png'\n            img_url = base_url + img_path\n\n            # Get the raw response from the url\n            r = requests.get(img_url, allow_redirects=True, stream=True)\n            r.raw.decode_content = True\n\n            # Use PIL to resize the image and to convert it to L\n            # (8-bit pixels, black and white)\n            im = Image.open(r.raw)\n            im = im.resize(image_size, Image.LANCZOS).convert('L')\n            im.save(os.path.join(save_dir, img_name), 'PNG')\nlet me know if there is any miss between download and use of image for training...",
      "votes": null
    },
    {
      "id": "460710",
      "postDate": "01/24/2019 09:27:56",
      "content": "<p>yet, it is confusing. if you do :</p>\n\n<p>&gt; file hpa_site_data/10580_1610_C1_1_green.jpg </p>\n\n<p>you will get this output:</p>\n\n<p><code>\nhpa_site_data/10580_1610_C1_1_green.jpg: JPEG image data, JFIF standard 1.01, resolution (DPCM), density 59326x59326, segment length 16, comment: \"ImageJ=1.48v\", baseline, precision 8, 2048x2048, frames 3\n</code></p>\n\n<p>Note the \"frames 3\", it means that there are 3 channels in the image, although 2 are more or less empty</p>\n\n<p>I think the L conversion is wrong - the discussion was pretty long, but you might want to read it all. i didn't use PIL, but here is my code snippet for converting it to single colour:</p>\n\n<p>```\n    def process_file(fn):</p>\n\n<pre><code>        colour = fn.split(\"/\")[-1].split(\".\")[0].split(\"_\")[-1]\n\n        if colour == 'yellow':\n\n                return\n\n        img = cv2.imread(fn)\n\n        if img.data == None:\n\n                print(\"defective image {}\".format(fn))\n\n                return\n\n        ofn = \"512_images/\"+fn.split(\"/\")[-1].split(\".\")[0]+\".png\"\n\n        img = cv2.resize(img, (512,512))[:,:,clr_idx[colour]]\n\n        cv2.imwrite(ofn, img)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "yet, it is confusing. if you do :\n\n&gt; file hpa\\_site\\_data/10580\\_1610\\_C1\\_1\\_green.jpg \n\nyou will get this output:\n\n```\nhpa_site_data/10580_1610_C1_1_green.jpg: JPEG image data, JFIF standard 1.01, resolution (DPCM), density 59326x59326, segment length 16, comment: \"ImageJ=1.48v\", baseline, precision 8, 2048x2048, frames 3\n```\n\n\nNote the \"frames 3\", it means that there are 3 channels in the image, although 2 are more or less empty\n\n\nI think the L conversion is wrong - the discussion was pretty long, but you might want to read it all. i didn't use PIL, but here is my code snippet for converting it to single colour:\n\n```\n    def process_file(fn):\n\n            colour = fn.split(\"/\")[-1].split(\".\")[0].split(\"_\")[-1]\n\n            if colour == 'yellow':\n\n                    return\n\n            img = cv2.imread(fn)\n\n            if img.data == None:\n\n                    print(\"defective image {}\".format(fn))\n\n                    return\n\n            ofn = \"512_images/\"+fn.split(\"/\")[-1].split(\".\")[0]+\".png\"\n\n            img = cv2.resize(img, (512,512))[:,:,clr_idx[colour]]\n\n            cv2.imwrite(ofn, img)\n```",
      "votes": null
    },
    {
      "id": "460907",
      "postDate": "01/24/2019 16:57:53",
      "content": "<p>1) Ah! as per doc L converts RGB to G ,could u let me know the catch in this case if we merely do this</p>\n\n<p>2) If i read your code correctly , you downloaded first each channel of image as RGB ,during resize you are reading 0 or 1 or 2  index of RGB depending upon the channel you read  eg. if green then clr_idx will return 1 ,if red it would return  0 or 2 , by doing just so is your image getting converted to gray , standard methodology is to do weighted sum of each channel in RGB scale to get gray version of image from RGB. </p>\n\n<p>3) You trained up your model using RGB or RGBA ,i see you are returning nothing for yellow</p>",
      "rawMarkdown": "1) Ah! as per doc L converts RGB to G ,could u let me know the catch in this case if we merely do this\n\n2) If i read your code correctly , you downloaded first each channel of image as RGB ,during resize you are reading 0 or 1 or 2  index of RGB depending upon the channel you read  eg. if green then clr_idx will return 1 ,if red it would return  0 or 2 , by doing just so is your image getting converted to gray , standard methodology is to do weighted sum of each channel in RGB scale to get gray version of image from RGB. \n\n3) You trained up your model using RGB or RGBA ,i see you are returning nothing for yellow",
      "votes": null
    },
    {
      "id": "460935",
      "postDate": "01/24/2019 19:02:56",
      "content": "<p>Yes, this is why your loss skyrocket. These images are synthetic. The channel is actually grayscale padded with empty channels. If you do weighted L you will get something very unlike what the competition images are like.\nJust try viewing the converted images, you'll see the problem right away. \nYes, i found yellow useless, but your milage may vary. Yellow has two channels not one, red and green iirc</p>",
      "rawMarkdown": "Yes, this is why your loss skyrocket. These images are synthetic. The channel is actually grayscale padded with empty channels. If you do weighted L you will get something very unlike what the competition images are like.\nJust try viewing the converted images, you'll see the problem right away. \nYes, i found yellow useless, but your milage may vary. Yellow has two channels not one, red and green iirc",
      "votes": null
    },
    {
      "id": "461144",
      "postDate": "01/25/2019 10:56:38",
      "content": "<p>I find that images i view are very dark with very low structure visibility ... \nWould it be fine if i read the image using cv2 and then split\nimg = cv2.imread(\"/path/to/your/image.jpg\")\n(channel_b, channel_g, channel_r) = cv2.split(img)\nfor green color i will take channel_g,blue  for yellow r/2+g/2</p>",
      "rawMarkdown": "I find that images i view are very dark with very low structure visibility ... \nWould it be fine if i read the image using cv2 and then split\nimg = cv2.imread(\"/path/to/your/image.jpg\")\n(channel_b, channel_g, channel_r) = cv2.split(img)\nfor green color i will take channel_g,blue  for yellow r/2+g/2",
      "votes": null
    },
    {
      "id": "461150",
      "postDate": "01/25/2019 11:13:41",
      "content": "<p>Never used split... Just give it a try and see if the images are clear again. </p>",
      "rawMarkdown": "Never used split... Just give it a try and see if the images are clear again.",
      "votes": null
    },
    {
      "id": "470101",
      "postDate": "02/12/2019 11:24:51",
      "content": "<p>hi moshel/lingyundev ...\nThanks for the help ,with the help of you i am able to get a private lb score .5123 using max ensemble of dense121/resnet50   &amp; los multilabel soft margin.</p>\n\n<p>@Lingyundev\ncould you please elaborate what you meant by  10 fold ensemble</p>",
      "rawMarkdown": "hi moshel/lingyundev ...\nThanks for the help ,with the help of you i am able to get a private lb score .5123 using max ensemble of dense121/resnet50   &amp; los multilabel soft margin.\n\n@Lingyundev\ncould you please elaborate what you meant by  10 fold ensemble",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 456859,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "01/16/2019 17:04:39",
      "content": "<p>hi,thanks for posting the solution..\n1) could you explain what is scale here...\n2) When i tried to used weights for classes then it was resulting into too high val loss,did you also find similar things for early epochs ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 457304,
          "author_name": "lingyundev",
          "author_url": "",
          "post_date": "01/17/2019 07:51:14",
          "content": "<ol>\n<li>Scale means the size of the input image(base size is 512x512).</li>\n<li>I didn't find the situation you said.  （BN and dropout after full connection，SGD,  BCE or MultiLabelSoftMarginLoss,  reasonable learning rate and lrscheduler may be helpful)</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 457433,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/17/2019 12:22:14",
          "content": "<p>THanks\n1) how did u do ensembling ,any easy example .  say model 1 prediction probs [0.99.....0.11] ,model2 [0.8....0.5]\n2) did you  also use external data,if yes could u please provide the script</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 457854,
          "author_name": "lingyundev",
          "author_url": "",
          "post_date": "01/18/2019 07:58:25",
          "content": "<ol>\n<li>We simply did weighted arithmetic mean on sigmoids.  </li>\n<li>Yes, we used the external data. The following is our script. <br>\n<code>\nimport os\nimport pandas as pd\nimport cv2\nimport multiprocessing\nimport tqdm\nimport requests\ncolors = ['red', 'green', 'blue', 'yellow']\nBASE_DATASET_PATH = './external_data/HPAv18/'\nDIR = BASE_DATASET_PATH + \"jpg/\"\nGray_DIR = BASE_DATASET_PATH + \"rgby_1024_png/\"\nv18_url = 'http://v18.proteinatlas.org/images/'\nif not os.path.exists(Gray_DIR):\nos.mkdir(Gray_DIR)\ndef download_img(item_name):\nimg = item_name.split('_')\nfor color in colors:\n    img_path = img[0] + '/' + \"_\".join(img[1:]) + \"_\" + color + \".jpg\"\n    img_name = item_name + \"_\" + color + \".jpg\"\n    img_url = v18_url + img_path\n    try:\n        r = requests.get(img_url, allow_redirects=True)\n        open(DIR + img_name, 'wb').write(r.content)\n    except Exception, e:\n        print e\n        print 'Error,{},{}'.format(img_url,img_name)\ndef rgb_to_gray(item_name):\nfor color in colors:\n    img_name = item_name + \"_\" + color + \".jpg\"\n    img_path = DIR + img_name\n    save_path = Gray_DIR + img_name[:-4] + '.png'\n    if os.path.exists(save_path):\n        continue\n    img = cv2.imread(img_path)\n    index = 0\n    if color == 'blue':\n        index = 0\n    elif color == 'green':\n        index = 1\n    elif color == 'red':\n        index = 2\n    elif color == 'yellow':\n        index = 1\n    img_gray = img[..., index]\n    if img_gray.shape[0] != 1024 or img_gray.shape[1] != 1024:\n        img_gray = cv2.resize(img_gray, (1024, 1024))\n    cv2.imwrite(save_path, img_gray)\nif __name__ == '__main__':\nimgList = pd.read_csv(BASE_DATASET_PATH + \"HPAv18RBGY_wodpl.csv\")\npool = multiprocessing.Pool(processes=50)\npool.map(download_img, imgList['Id'])\npBar = tqdm.tqdm(total=len(imgList))\nfor i, item_name in enumerate(imgList['Id']):\n    rgb_to_gray(item_name)\n    if i % 100 == 0:\n        pBar.update(100)\npBar.close()\n</code></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458025,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/18/2019 15:49:42",
          "content": "<p>Thanku</p>\n\n<p>1) there were some posts which mentioned there are images in external hpa that are duplicates in kaggles train set ,how u filtered the duplicated data ??\n2) Are the images in the file HPAv18RBGY_wodpl.csv   different from the ones that shown and downloaded here   <a href=\"https://www.kaggle.com/artemtprv/load-external-data/data\">https://www.kaggle.com/artemtprv/load-external-data/data</a>\n3) how you appended the labels from the files,\n4 ) finaly do we had to use the different normalization for external hpa</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458865,
          "author_name": "lingyundev",
          "author_url": "",
          "post_date": "01/20/2019 17:11:43",
          "content": "<p>I just used this file <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/432870/10816/HPAv18RBGY_wodpl.csv\">HPAv18RBGY_wodpl.csv</a> , added to the training set.  No other operation has been taken.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 459133,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/21/2019 08:52:10",
          "content": "<p>thanks..\ndid you do any differential normalization i.e different normalization for HPA and different for competition, as some people mentioning that HPA data has got different pixel intensity than competition  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 459861,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/22/2019 14:03:55",
          "content": "<p>Hi lingyundev \nI used HPA data along with regular data but i get high val losses. Note that i used only selective data based on class rarity and original image resize to 512 .Total image downloaded 32k.\nWhat could be the reason for high val losses.... any more processing is that needed before we can use HPA.</p>\n\n<p>Appreciate your input here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460617,
          "author_name": "lingyundev",
          "author_url": "",
          "post_date": "01/24/2019 05:08:39",
          "content": "<p>hi jaideep <br>\nDuring the competition, we merged the HPAv18 and official datasets directly.\ndid you tried the parameters in my solution? especially the loss, optimizer  and learning rate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460635,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/24/2019 05:37:28",
          "content": "<p>Make sure you preprocess the external data correctly. Note that unlike the competition data, each image has 3 channels. If you read it as 1 color image, it will be very wrong. You need to read it as 3ch and save only the correct channel as grayimage</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460695,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/24/2019 08:57:13",
          "content": "<p>hi ling m yet to change the loss,i use focal loss the images i downloaded contained all the channels .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460698,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/24/2019 09:01:43",
          "content": "<p>hi moshel\nimages i downloaded looked to contain all the channels not just the RGB. it also had yellow channel data. so i dint do any thing after download ,just appended the list of downloaded images to regular image list.\nI used david silva script in official \n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a>\ndef download(pid, image_list, base_url, save_dir, image_size=(512, 512)):\n    colors = ['red', 'green', 'blue', 'yellow']\n    for i in tqdm(image_list, postfix=pid):\n        img_id = i.split('_', 1)\n        for color in colors:\n            img_path = img_id[0] + '/' + img_id[1] + '_' + color + '.jpg'\n            img_name = i + '_' + color + '.png'\n            img_url = base_url + img_path</p>\n\n<pre><code>        # Get the raw response from the url\n        r = requests.get(img_url, allow_redirects=True, stream=True)\n        r.raw.decode_content = True\n\n        # Use PIL to resize the image and to convert it to L\n        # (8-bit pixels, black and white)\n        im = Image.open(r.raw)\n        im = im.resize(image_size, Image.LANCZOS).convert('L')\n        im.save(os.path.join(save_dir, img_name), 'PNG')\n</code></pre>\n\n<p>let me know if there is any miss between download and use of image for training...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460710,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/24/2019 09:27:56",
          "content": "<p>yet, it is confusing. if you do :</p>\n\n<p>&gt; file hpa_site_data/10580_1610_C1_1_green.jpg </p>\n\n<p>you will get this output:</p>\n\n<p><code>\nhpa_site_data/10580_1610_C1_1_green.jpg: JPEG image data, JFIF standard 1.01, resolution (DPCM), density 59326x59326, segment length 16, comment: \"ImageJ=1.48v\", baseline, precision 8, 2048x2048, frames 3\n</code></p>\n\n<p>Note the \"frames 3\", it means that there are 3 channels in the image, although 2 are more or less empty</p>\n\n<p>I think the L conversion is wrong - the discussion was pretty long, but you might want to read it all. i didn't use PIL, but here is my code snippet for converting it to single colour:</p>\n\n<p>```\n    def process_file(fn):</p>\n\n<pre><code>        colour = fn.split(\"/\")[-1].split(\".\")[0].split(\"_\")[-1]\n\n        if colour == 'yellow':\n\n                return\n\n        img = cv2.imread(fn)\n\n        if img.data == None:\n\n                print(\"defective image {}\".format(fn))\n\n                return\n\n        ofn = \"512_images/\"+fn.split(\"/\")[-1].split(\".\")[0]+\".png\"\n\n        img = cv2.resize(img, (512,512))[:,:,clr_idx[colour]]\n\n        cv2.imwrite(ofn, img)\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460907,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/24/2019 16:57:53",
          "content": "<p>1) Ah! as per doc L converts RGB to G ,could u let me know the catch in this case if we merely do this</p>\n\n<p>2) If i read your code correctly , you downloaded first each channel of image as RGB ,during resize you are reading 0 or 1 or 2  index of RGB depending upon the channel you read  eg. if green then clr_idx will return 1 ,if red it would return  0 or 2 , by doing just so is your image getting converted to gray , standard methodology is to do weighted sum of each channel in RGB scale to get gray version of image from RGB. </p>\n\n<p>3) You trained up your model using RGB or RGBA ,i see you are returning nothing for yellow</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 460935,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/24/2019 19:02:56",
          "content": "<p>Yes, this is why your loss skyrocket. These images are synthetic. The channel is actually grayscale padded with empty channels. If you do weighted L you will get something very unlike what the competition images are like.\nJust try viewing the converted images, you'll see the problem right away. \nYes, i found yellow useless, but your milage may vary. Yellow has two channels not one, red and green iirc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461144,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/25/2019 10:56:38",
          "content": "<p>I find that images i view are very dark with very low structure visibility ... \nWould it be fine if i read the image using cv2 and then split\nimg = cv2.imread(\"/path/to/your/image.jpg\")\n(channel_b, channel_g, channel_r) = cv2.split(img)\nfor green color i will take channel_g,blue  for yellow r/2+g/2</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461150,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/25/2019 11:13:41",
          "content": "<p>Never used split... Just give it a try and see if the images are clear again. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 470101,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/12/2019 11:24:51",
          "content": "<p>hi moshel/lingyundev ...\nThanks for the help ,with the help of you i am able to get a private lb score .5123 using max ensemble of dense121/resnet50   &amp; los multilabel soft margin.</p>\n\n<p>@Lingyundev\ncould you please elaborate what you meant by  10 fold ensemble</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 457673,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "01/17/2019 22:57:15",
      "content": "<p>You must have nice hardware. I had Ti1080x2 and, for size 768, mt batch size was 8.</p>",
      "votes": null,
      "replies": [
        {
          "id": 457864,
          "author_name": "lingyundev",
          "author_url": "",
          "post_date": "01/18/2019 08:26:21",
          "content": "<p>HA HA! We had p100x4, batch size was 10 (inception-v4, 800, single GPU)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458137,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "01/18/2019 21:44:56",
          "content": "<p>I was using resnet50, sz=768, one Ti1080 with batch size 8. When it failed at 16, I went directly to 8 so it might have been a little higher. With resnet34, I had batch size 16.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "456591": "First of all, congratulations to all the winners! Thanks to Kaggle and HPA team for hosting such an interesting competition and thanks to  [TomomiMoriyama](https://www.kaggle.com/tomomimoriyama), [Heng CherKeng](https://www.kaggle.com/hengck23), [ManyFoldCV](https://www.kaggle.com/manyfoldcv) and [Spytensor](https://www.kaggle.com/spytensor).  \n\nHere is a brief summary of our solution.\n\n### DataSet\n\nLike most other competitors, we used both official (both PNG and TIFF) and external data. To deal with class-imbalance, we used WeightedRandomSampler （method in pytorch） during training and [MultilabelStratifiedShuffleSplit](https://github.com/trent-b/iterative-stratification) to split the data into training and validation. We constructed 10 folds cross validation sets with 8% for validation.\n\n### Image Preprocessing\n\nThe HPA dataset has four dyeing modes each of which is an RGB image of its own, so we took only one channel （r=r,g=g,b=b,y=b） to form a 4-channel input for training.\n\nAll PNG images are kept at their original 512 size, whereas the TIFF images are resized to 1024.\n\n### Augmentation\n\nRotation, Flip, and Shear.\n\nWe didn't use random cropping. Instead we trained 5 models using crop5 (method in pytorch) and found it to be more effective.\n\n### Models\n\nFor our base networks, we mainly used Inception-v3，-v4, and Xception. We have also tried DenseNet, SENet and ResNet, but the results were suboptimal.\n\nWe used three different scales during training (512 for PNG images and 650, 800 for TIFF images) with different random seeds for the 10-folds CV.\n\nModifications\n\n1. Changed the last pooling layer to global pooling.\n2. Appended an additional fully connected layer with output dimension 128 after the global pooling.\n3. We also divided the training process into two stages where the first stage used size 512 with model trained on ImageNet, and the second stage used size 650 or 800 with model trained from the first stage. We found this to be slightly better than training with fixed size all the way.\n\n### Training\n\n* loss: [MultiLabelSoftMarginLoss](https://pytorch.org/docs/stable/nn.html?highlight=multilabelsoftmarginloss#torch.nn.MultiLabelSoftMarginLoss)\n* lr: 0.05 (for size 512, pretrained on ImageNet)，0.01 (for size 650 and 800，pretrained using size 512); lrscheduler: steplr(gamma=0.1,step=6)\n* optimizer: SGD\n* epochs: 25, early stopping for training with size 650 or 800 (around 15 epochs), model selected based on loss (instead of F1 score)\n* sampling weights for different classes: [1.0, 5.97, 2.89, 5.75, 4.64, 4.27, 5.46, 3.2, 14.48, 14.84, 15.14, 6.92, 6.86, 8.12, 6.32, 19.24, 8.48, 11.93, 7.32, 5.48, 11.99, 2.39, 6.3, 3.0, 12.06, 1.0, 10.39, 16.5]\n  \n### Multi-Thresholds\n\nWe used the validation sets to search for threshold for each class by optimizing the F1 score begining with 0.15 for all classes.\n\n\n### Test\n\n(with multi-thresholds)\n\n![](https://s2.ax1x.com/2019/01/17/kpZ7lV.png)\n\n### Ensembling\n\nFinal prediction is ensemble of above methods: Size 800, 10-fold for Inception-v3; Size 650 and 800, 10-fold for Inception-v4; Size 800, 10-fold, Size 650, 1-fold, Size 512, 5-fold for Xception (the reason for 5-fold instead of 10 was simply because we didn't have enough submissions to check the performances of all models, so we simply took the best ones).\n\n### Things that did not work for us\n* Training with larger input size (&gt;= 1024), which forced us reduce the batch size.\n* 3-channel input\n* focal loss\n* C3D\n* TTA: unlike a lot of other competitors, TTA during test time actually didn't work for us.\n* Other traditional machine learning methods such as DecisionTree, RandomForest, and SVM.",
    "456859": "hi,thanks for posting the solution..\n1) could you explain what is scale here...\n2) When i tried to used weights for classes then it was resulting into too high val loss,did you also find similar things for early epochs ?",
    "457304": "1. Scale means the size of the input image(base size is 512x512).\n2. I didn't find the situation you said.  （BN and dropout after full connection，SGD,  BCE or MultiLabelSoftMarginLoss,  reasonable learning rate and lrscheduler may be helpful)",
    "457433": "THanks\n1) how did u do ensembling ,any easy example .  say model 1 prediction probs [0.99.....0.11] ,model2 [0.8....0.5]\n2) did you  also use external data,if yes could u please provide the script",
    "457673": "You must have nice hardware. I had Ti1080x2 and, for size 768, mt batch size was 8.",
    "457854": "1. We simply did weighted arithmetic mean on sigmoids.  \n2. Yes, we used the external data. The following is our script.    \n```\nimport os\nimport pandas as pd\nimport cv2\nimport multiprocessing\nimport tqdm\nimport requests\ncolors = ['red', 'green', 'blue', 'yellow']\nBASE_DATASET_PATH = './external_data/HPAv18/'\nDIR = BASE_DATASET_PATH + \"jpg/\"\nGray_DIR = BASE_DATASET_PATH + \"rgby_1024_png/\"\nv18_url = 'http://v18.proteinatlas.org/images/'\nif not os.path.exists(Gray_DIR):\n    os.mkdir(Gray_DIR)\ndef download_img(item_name):\n    img = item_name.split('_')\n    for color in colors:\n        img_path = img[0] + '/' + \"_\".join(img[1:]) + \"_\" + color + \".jpg\"\n        img_name = item_name + \"_\" + color + \".jpg\"\n        img_url = v18_url + img_path\n        try:\n            r = requests.get(img_url, allow_redirects=True)\n            open(DIR + img_name, 'wb').write(r.content)\n        except Exception, e:\n            print e\n            print 'Error,{},{}'.format(img_url,img_name)\ndef rgb_to_gray(item_name):\n    for color in colors:\n        img_name = item_name + \"_\" + color + \".jpg\"\n        img_path = DIR + img_name\n        save_path = Gray_DIR + img_name[:-4] + '.png'\n        if os.path.exists(save_path):\n            continue\n        img = cv2.imread(img_path)\n        index = 0\n        if color == 'blue':\n            index = 0\n        elif color == 'green':\n            index = 1\n        elif color == 'red':\n            index = 2\n        elif color == 'yellow':\n            index = 1\n        img_gray = img[..., index]\n        if img_gray.shape[0] != 1024 or img_gray.shape[1] != 1024:\n            img_gray = cv2.resize(img_gray, (1024, 1024))\n        cv2.imwrite(save_path, img_gray)\nif __name__ == '__main__':\n    imgList = pd.read_csv(BASE_DATASET_PATH + \"HPAv18RBGY_wodpl.csv\")\n    pool = multiprocessing.Pool(processes=50)\n    pool.map(download_img, imgList['Id'])\n    pBar = tqdm.tqdm(total=len(imgList))\n    for i, item_name in enumerate(imgList['Id']):\n        rgb_to_gray(item_name)\n        if i % 100 == 0:\n            pBar.update(100)\n    pBar.close()\n```",
    "457864": "HA HA! We had p100x4, batch size was 10 (inception-v4, 800, single GPU)",
    "458025": "Thanku\n\n1) there were some posts which mentioned there are images in external hpa that are duplicates in kaggles train set ,how u filtered the duplicated data ??\n2) Are the images in the file HPAv18RBGY_wodpl.csv   different from the ones that shown and downloaded here   https://www.kaggle.com/artemtprv/load-external-data/data\n3) how you appended the labels from the files,\n4 ) finaly do we had to use the different normalization for external hpa",
    "458137": "I was using resnet50, sz=768, one Ti1080 with batch size 8. When it failed at 16, I went directly to 8 so it might have been a little higher. With resnet34, I had batch size 16.",
    "458865": "I just used this file [HPAv18RBGY_wodpl.csv](https://storage.googleapis.com/kaggle-forum-message-attachments/432870/10816/HPAv18RBGY_wodpl.csv) , added to the training set.  No other operation has been taken.",
    "459133": "thanks..\ndid you do any differential normalization i.e different normalization for HPA and different for competition, as some people mentioning that HPA data has got different pixel intensity than competition",
    "459861": "Hi lingyundev \nI used HPA data along with regular data but i get high val losses. Note that i used only selective data based on class rarity and original image resize to 512 .Total image downloaded 32k.\nWhat could be the reason for high val losses.... any more processing is that needed before we can use HPA.\n\nAppreciate your input here.",
    "460617": "hi jaideep  \nDuring the competition, we merged the HPAv18 and official datasets directly.\ndid you tried the parameters in my solution? especially the loss, optimizer  and learning rate.",
    "460635": "Make sure you preprocess the external data correctly. Note that unlike the competition data, each image has 3 channels. If you read it as 1 color image, it will be very wrong. You need to read it as 3ch and save only the correct channel as grayimage",
    "460695": "hi ling m yet to change the loss,i use focal loss the images i downloaded contained all the channels .",
    "460698": "hi moshel\nimages i downloaded looked to contain all the channels not just the RGB. it also had yellow channel data. so i dint do any thing after download ,just appended the list of downloaded images to regular image list.\nI used david silva script in official \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\ndef download(pid, image_list, base_url, save_dir, image_size=(512, 512)):\n    colors = ['red', 'green', 'blue', 'yellow']\n    for i in tqdm(image_list, postfix=pid):\n        img_id = i.split('_', 1)\n        for color in colors:\n            img_path = img_id[0] + '/' + img_id[1] + '_' + color + '.jpg'\n            img_name = i + '_' + color + '.png'\n            img_url = base_url + img_path\n\n            # Get the raw response from the url\n            r = requests.get(img_url, allow_redirects=True, stream=True)\n            r.raw.decode_content = True\n\n            # Use PIL to resize the image and to convert it to L\n            # (8-bit pixels, black and white)\n            im = Image.open(r.raw)\n            im = im.resize(image_size, Image.LANCZOS).convert('L')\n            im.save(os.path.join(save_dir, img_name), 'PNG')\nlet me know if there is any miss between download and use of image for training...",
    "460710": "yet, it is confusing. if you do :\n\n&gt; file hpa\\_site\\_data/10580\\_1610\\_C1\\_1\\_green.jpg \n\nyou will get this output:\n\n```\nhpa_site_data/10580_1610_C1_1_green.jpg: JPEG image data, JFIF standard 1.01, resolution (DPCM), density 59326x59326, segment length 16, comment: \"ImageJ=1.48v\", baseline, precision 8, 2048x2048, frames 3\n```\n\n\nNote the \"frames 3\", it means that there are 3 channels in the image, although 2 are more or less empty\n\n\nI think the L conversion is wrong - the discussion was pretty long, but you might want to read it all. i didn't use PIL, but here is my code snippet for converting it to single colour:\n\n```\n    def process_file(fn):\n\n            colour = fn.split(\"/\")[-1].split(\".\")[0].split(\"_\")[-1]\n\n            if colour == 'yellow':\n\n                    return\n\n            img = cv2.imread(fn)\n\n            if img.data == None:\n\n                    print(\"defective image {}\".format(fn))\n\n                    return\n\n            ofn = \"512_images/\"+fn.split(\"/\")[-1].split(\".\")[0]+\".png\"\n\n            img = cv2.resize(img, (512,512))[:,:,clr_idx[colour]]\n\n            cv2.imwrite(ofn, img)\n```",
    "460907": "1) Ah! as per doc L converts RGB to G ,could u let me know the catch in this case if we merely do this\n\n2) If i read your code correctly , you downloaded first each channel of image as RGB ,during resize you are reading 0 or 1 or 2  index of RGB depending upon the channel you read  eg. if green then clr_idx will return 1 ,if red it would return  0 or 2 , by doing just so is your image getting converted to gray , standard methodology is to do weighted sum of each channel in RGB scale to get gray version of image from RGB. \n\n3) You trained up your model using RGB or RGBA ,i see you are returning nothing for yellow",
    "460935": "Yes, this is why your loss skyrocket. These images are synthetic. The channel is actually grayscale padded with empty channels. If you do weighted L you will get something very unlike what the competition images are like.\nJust try viewing the converted images, you'll see the problem right away. \nYes, i found yellow useless, but your milage may vary. Yellow has two channels not one, red and green iirc",
    "461144": "I find that images i view are very dark with very low structure visibility ... \nWould it be fine if i read the image using cv2 and then split\nimg = cv2.imread(\"/path/to/your/image.jpg\")\n(channel_b, channel_g, channel_r) = cv2.split(img)\nfor green color i will take channel_g,blue  for yellow r/2+g/2",
    "461150": "Never used split... Just give it a try and see if the images are clear again.",
    "470101": "hi moshel/lingyundev ...\nThanks for the help ,with the help of you i am able to get a private lb score .5123 using max ensemble of dense121/resnet50   &amp; los multilabel soft margin.\n\n@Lingyundev\ncould you please elaborate what you meant by  10 fold ensemble"
  },
  "source": "meta"
}