{
  "id": 75691,
  "title": "Some information that may help you reach 0.59+...",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/75691",
  "author_name": "Kulbear",
  "post_date": "2018-12-25T04:04:36.534000",
  "votes": 96,
  "comment_count": 63,
  "views": 0,
  "content": "<ol>\n<li>External data is not a must for 0.58+ (based on my teammate's result), but it does help (make your life easier).</li>\n<li>Image size matters, use larger image size if you can.</li>\n<li>Personally found Adam works well in our case... You can play with it...</li>\n<li>Augmentation (rotation, flip and shear seem to be enough, maybe colorjitter?)</li>\n<li>Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere...)</li>\n<li>Reduce your learning rate, with a proper schedule, to escape from local minima...</li>\n<li>I personally always got some worse results when I tried to combine simple BCE with some other fancy loss (Focal, Soft F1 loss...)...</li>\n<li>Optimal threshold searching does not work for me, up to now.</li>\n<li>Ensemble your many folds (and different models if possible).</li>\n<li>I tried to ensemble my 256 5-fold and got 0.556, then I try to ensemble my 512 5-fold and got 586. I added the 256 result to the 512 result I can reach 0.59X (forgot) ...</li>\n<li>The above score can be reached with BN-Inception...</li>\n<li>I use <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812</a> as my starter template as I joined this competition pretty late, like, about 2 weeks ago.</li>\n</ol>\n\n<p>Bonus: a weird observation --&gt; using varied batch size, when the model converge to some BCE (say, 0.7), the exact same settings and model, the one with larger batch size will give better local f1 validation score, though the LB scores are not very different, potentially there are some noise labels (in the external?) ?</p>\n\n<p>Other info:\nI use one 1080Ti + 1080, and some GCP credits to use P100.</p>\n\n<p>My question:\n1. Using sigmoid threshold 0.5 cannot produce decent results, I am using some smaller value but I'd like to understand why this happens, any help is appreciated.</p>\n\n<p>Update 12-25:\nFind the \"correct\" way of processing the external data is critical!</p>",
  "messages": [
    {
      "id": 444881,
      "postDate": "2018-12-25T04:04:36.533Z",
      "content": "<ol>\n<li>External data is not a must for 0.58+ (based on my teammate's result), but it does help (make your life easier).</li>\n<li>Image size matters, use larger image size if you can.</li>\n<li>Personally found Adam works well in our case... You can play with it...</li>\n<li>Augmentation (rotation, flip and shear seem to be enough, maybe colorjitter?)</li>\n<li>Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere...)</li>\n<li>Reduce your learning rate, with a proper schedule, to escape from local minima...</li>\n<li>I personally always got some worse results when I tried to combine simple BCE with some other fancy loss (Focal, Soft F1 loss...)...</li>\n<li>Optimal threshold searching does not work for me, up to now.</li>\n<li>Ensemble your many folds (and different models if possible).</li>\n<li>I tried to ensemble my 256 5-fold and got 0.556, then I try to ensemble my 512 5-fold and got 586. I added the 256 result to the 512 result I can reach 0.59X (forgot) ...</li>\n<li>The above score can be reached with BN-Inception...</li>\n<li>I use <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812</a> as my starter template as I joined this competition pretty late, like, about 2 weeks ago.</li>\n</ol>\n\n<p>Bonus: a weird observation --&gt; using varied batch size, when the model converge to some BCE (say, 0.7), the exact same settings and model, the one with larger batch size will give better local f1 validation score, though the LB scores are not very different, potentially there are some noise labels (in the external?) ?</p>\n\n<p>Other info:\nI use one 1080Ti + 1080, and some GCP credits to use P100.</p>\n\n<p>My question:\n1. Using sigmoid threshold 0.5 cannot produce decent results, I am using some smaller value but I'd like to understand why this happens, any help is appreciated.</p>\n\n<p>Update 12-25:\nFind the \"correct\" way of processing the external data is critical!</p>",
      "rawMarkdown": "1. External data is not a must for 0.58+ (based on my teammate's result), but it does help (make your life easier).\n2. Image size matters, use larger image size if you can.\n3. Personally found Adam works well in our case... You can play with it...\n4. Augmentation (rotation, flip and shear seem to be enough, maybe colorjitter?)\n5. Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere...)\n6. Reduce your learning rate, with a proper schedule, to escape from local minima...\n7. I personally always got some worse results when I tried to combine simple BCE with some other fancy loss (Focal, Soft F1 loss...)...\n8. Optimal threshold searching does not work for me, up to now.\n9. Ensemble your many folds (and different models if possible).\n10. I tried to ensemble my 256 5-fold and got 0.556, then I try to ensemble my 512 5-fold and got 586. I added the 256 result to the 512 result I can reach 0.59X (forgot) ...\n11. The above score can be reached with BN-Inception...\n12. I use https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812 as my starter template as I joined this competition pretty late, like, about 2 weeks ago.\n\nBonus: a weird observation --&gt; using varied batch size, when the model converge to some BCE (say, 0.7), the exact same settings and model, the one with larger batch size will give better local f1 validation score, though the LB scores are not very different, potentially there are some noise labels (in the external?) ?\n\nOther info:\nI use one 1080Ti + 1080, and some GCP credits to use P100.\n\nMy question:\n1. Using sigmoid threshold 0.5 cannot produce decent results, I am using some smaller value but I'd like to understand why this happens, any help is appreciated.\n\n\nUpdate 12-25:\nFind the \"correct\" way of processing the external data is critical!",
      "votes": 94
    },
    {
      "id": 444910,
      "postDate": "2018-12-25T06:24:05.597Z",
      "content": "<p>@ManyFoldCV Thanks for the hints!</p>\n\n<p>What public/validation score do you get (or you estimate) if we only use kaggle train data without leak, without external for a single model without TTA, without ensemble?</p>",
      "rawMarkdown": "@ManyFoldCV Thanks for the hints!\n\nWhat public/validation score do you get (or you estimate) if we only use kaggle train data without leak, without external for a single model without TTA, without ensemble?",
      "votes": 3,
      "replies": [
        {
          "id": 445134,
          "postDate": "2018-12-25T18:11:50.310Z",
          "content": "<p>I've learned so much from you Heng :D</p>\n\n<p>Using only the official data I got a public score 0.472. At that time my local f1 formula had a problem so I don't have a local stat.</p>",
          "rawMarkdown": "I've learned so much from you Heng :D\n\nUsing only the official data I got a public score 0.472. At that time my local f1 formula had a problem so I don't have a local stat.",
          "votes": 1
        }
      ]
    },
    {
      "id": 446001,
      "postDate": "2018-12-27T10:34:03.447Z",
      "content": "<p>How do you ensemble models? Caculating average of outputs? Voting?</p>",
      "rawMarkdown": "How do you ensemble models? Caculating average of outputs? Voting?",
      "votes": 4,
      "replies": [
        {
          "id": 446010,
          "postDate": "2018-12-27T10:44:43.953Z",
          "content": "<p>for me，simply averaging the output worse than single fold model</p>",
          "rawMarkdown": "for me，simply averaging the output worse than single fold model"
        },
        {
          "id": 446013,
          "postDate": "2018-12-27T10:47:14.070Z",
          "content": "<p>There are two options:\n1. Your k-fold split is bad (you probably need more subtle stratification or something like that).\n2. You single fold model is by chance overfitted to the public LB, but average of folds is still more robust, so it will be better on private LB.</p>",
          "rawMarkdown": "There are two options:\n1. Your k-fold split is bad (you probably need more subtle stratification or something like that).\n2. You single fold model is by chance overfitted to the public LB, but average of folds is still more robust, so it will be better on private LB.",
          "votes": 2
        },
        {
          "id": 446033,
          "postDate": "2018-12-27T11:27:45.773Z",
          "content": "<p>Thanks for your reply\nAbout the k fold split，I use the Multilabel Stratification  python package by @Trent ,so may be that no the problem.\nhope my model can survive in private LB. </p>",
          "rawMarkdown": "Thanks for your reply\nAbout the k fold split，I use the Multilabel Stratification  python package by @Trent ,so may be that no the problem.\nhope my model can survive in private LB. "
        },
        {
          "id": 446166,
          "postDate": "2018-12-27T16:29:07.460Z",
          "content": "<p>Average or weighted average the proba</p>",
          "rawMarkdown": "Average or weighted average the proba",
          "votes": 2
        }
      ]
    },
    {
      "id": 445922,
      "postDate": "2018-12-27T08:12:37.350Z",
      "content": "<p>I think this is the one you are looking for:</p>\n\n<pre>train_df_orig=train_df.copy()    \nlows = [15,15,15,8,9,10,8,9,10,8,9,10,17,20,24,26,15,27,15,20,24,17,8,15,27,27,27]\nfor i in lows:\n    target = str(i)\n    indicies = train_df_orig.loc[train_df_orig['Target'] == target].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.startswith(target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.endswith(\" \"+target)].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.contains(\" \"+target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n</pre>",
      "rawMarkdown": "I think this is the one you are looking for:\n<pre>train_df_orig=train_df.copy()    \nlows = [15,15,15,8,9,10,8,9,10,8,9,10,17,20,24,26,15,27,15,20,24,17,8,15,27,27,27]\nfor i in lows:\n    target = str(i)\n    indicies = train_df_orig.loc[train_df_orig['Target'] == target].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.startswith(target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.endswith(\" \"+target)].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.contains(\" \"+target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n</pre>",
      "votes": 4,
      "replies": [
        {
          "id": 446169,
          "postDate": "2018-12-27T16:29:53.103Z",
          "content": "<p>Thanks <a href=\"/ldm314\">@ldm314</a>, that's the one.</p>",
          "rawMarkdown": "Thanks @ldm314, that's the one."
        },
        {
          "id": 446460,
          "postDate": "2018-12-28T05:22:54.550Z",
          "content": "<p>Thanks for the snippet. Out of curiosity how did you choose the frequency to oversample the classes in your 'lows' list?</p>",
          "rawMarkdown": "Thanks for the snippet. Out of curiosity how did you choose the frequency to oversample the classes in your 'lows' list?",
          "votes": 1
        }
      ]
    },
    {
      "id": 446234,
      "postDate": "2018-12-27T18:46:17.470Z",
      "content": "<p>So what is the 'correct' way of processing the external data?</p>",
      "rawMarkdown": "So what is the 'correct' way of processing the external data?",
      "votes": 1,
      "replies": [
        {
          "id": 446302,
          "postDate": "2018-12-27T20:52:19.863Z",
          "content": "<p>I processed it in the way we discussed on previous threads - basically a straight copy of the corresponding channels, adding R+G for yellow. When I added the external data to the challenge data (a LOT of data) my results were not optimal. So I tried adding the rare classes first and kept adding until the results started to degrade. That helped. </p>",
          "rawMarkdown": "I processed it in the way we discussed on previous threads - basically a straight copy of the corresponding channels, adding R+G for yellow. When I added the external data to the challenge data (a LOT of data) my results were not optimal. So I tried adding the rare classes first and kept adding until the results started to degrade. That helped. ",
          "votes": 7
        },
        {
          "id": 446317,
          "postDate": "2018-12-27T21:32:04.460Z",
          "content": "<p>Hi <a href=\"/maw501\">@maw501</a>, pete is giving the correct way :DDD</p>",
          "rawMarkdown": "Hi @maw501, pete is giving the correct way :DDD"
        },
        {
          "id": 446322,
          "postDate": "2018-12-27T21:54:45.140Z",
          "content": "<p>Thanks Pete and ManyFoldCV.</p>\n\n<p>So basically download using something like the following (note the line marked <em>*</em> which just extracts a single channel):</p>\n\n<pre><code>def download(id_):\n    try:\n        image_dir, _, image_id = id_.partition('_')\n        arr_out = np.zeros((512,512,3))\n        for i, c in enumerate(['red', 'green', 'blue']):\n            url = f'http://v18.proteinatlas.org/images/{image_dir}/{image_id}_{c}.jpg'\n            r = requests.get(url)\n            image = Image.open(BytesIO(r.content)).resize((512, 512), PIL.Image.LANCZOS)\n            arr_out[:,:,i] = (np.array(image)[:,:,i]).astype('uint8') # ***\n        img_pil = Image.fromarray(arr_out.astype('uint8')) \n        img_pil.save(f'{hpa_dir}/{id_}.png')\n    except:\n        print(f'{id_} broke...')\n</code></pre>\n\n<p>Are you adding hpa to your validation sets as well?</p>",
          "rawMarkdown": "Thanks Pete and ManyFoldCV.\n\nSo basically download using something like the following (note the line marked *** which just extracts a single channel):\n\n    def download(id_):\n        try:\n            image_dir, _, image_id = id_.partition('_')\n            arr_out = np.zeros((512,512,3))\n            for i, c in enumerate(['red', 'green', 'blue']):\n                url = f'http://v18.proteinatlas.org/images/{image_dir}/{image_id}_{c}.jpg'\n                r = requests.get(url)\n                image = Image.open(BytesIO(r.content)).resize((512, 512), PIL.Image.LANCZOS)\n                arr_out[:,:,i] = (np.array(image)[:,:,i]).astype('uint8') # ***\n            img_pil = Image.fromarray(arr_out.astype('uint8')) \n            img_pil.save(f'{hpa_dir}/{id_}.png')\n        except:\n            print(f'{id_} broke...')\n\nAre you adding hpa to your validation sets as well?"
        },
        {
          "id": 446481,
          "postDate": "2018-12-28T06:43:00.210Z",
          "content": "<p>I did. I worried about possible duplicates but did not act on it.</p>",
          "rawMarkdown": "I did. I worried about possible duplicates but did not act on it.",
          "votes": 1
        },
        {
          "id": 446557,
          "postDate": "2018-12-28T09:25:48.530Z",
          "content": "<p>Here is what I use to download and store the images for protein of interest(green) and the other 3 markers(red/blue/yellow).  I am not combining them into one image like you do, rather I am saving them as separate images just like in the train and the test set.</p>\n\n<pre><code>def downloadImages(id):\n    colors = ['red','green','blue','yellow']\n    img_path = v18_url + id.replace('_', '/', 1)\n    for i, color in enumerate(colors):\n        response = requests.get(img_path + \"_\" + color + \".jpg\", allow_redirects=True)\n        im = Image.open(BytesIO(response.content)).resize((512,512))\n        arr = np.asarray(im)\n        if i == 3: # Yellow\n            im = Image.fromarray(((arr[:,:,0] + arr[:,:,1])/2).astype(np.uint8))\n       else :\n           im = Image.fromarray(arr[:,:,i])\n        im.save(DATA_DIR + id + \"_\" + color + \".png\")\n</code></pre>",
          "rawMarkdown": "Here is what I use to download and store the images for protein of interest(green) and the other 3 markers(red/blue/yellow).  I am not combining them into one image like you do, rather I am saving them as separate images just like in the train and the test set.\n\n\n    def downloadImages(id):\n        colors = ['red','green','blue','yellow']\n        img_path = v18_url + id.replace('_', '/', 1)\n        for i, color in enumerate(colors):\n            response = requests.get(img_path + \"_\" + color + \".jpg\", allow_redirects=True)\n            im = Image.open(BytesIO(response.content)).resize((512,512))\n            arr = np.asarray(im)\n            if i == 3: # Yellow\n                im = Image.fromarray(((arr[:,:,0] + arr[:,:,1])/2).astype(np.uint8))\n           else :\n               im = Image.fromarray(arr[:,:,i])\n            im.save(DATA_DIR + id + \"_\" + color + \".png\")\n\n   "
        },
        {
          "id": 446788,
          "postDate": "2018-12-28T17:38:10.737Z",
          "content": "<p><a href=\"/maw501\">@maw501</a> I suggest to save the RGB space JPEG image to the disk first then process it if you have enough space. It will be about ~60G, 280k images. </p>",
          "rawMarkdown": "@maw501 I suggest to save the RGB space JPEG image to the disk first then process it if you have enough space. It will be about ~60G, 280k images. "
        },
        {
          "id": 446813,
          "postDate": "2018-12-28T18:21:09.643Z",
          "content": "<p>Hi <a href=\"/manyfoldcv\">@manyfoldcv</a> - I've already got the hpa...just finding training pretty unstable with it atm. </p>",
          "rawMarkdown": "Hi @manyfoldcv - I've already got the hpa...just finding training pretty unstable with it atm. "
        },
        {
          "id": 450287,
          "postDate": "2019-01-04T16:27:17.060Z",
          "content": "<p>which classes do you keep? From hpa, I've added Nuclear membrane and every class that has less samples which gave 0.55 for 5-fold model</p>",
          "rawMarkdown": "which classes do you keep? From hpa, I've added Nuclear membrane and every class that has less samples which gave 0.55 for 5-fold model"
        },
        {
          "id": 450295,
          "postDate": "2019-01-04T16:38:58.980Z",
          "content": "<p>Keeping all at the minute but using a sampling strategy to even out the distributions a little.</p>",
          "rawMarkdown": "Keeping all at the minute but using a sampling strategy to even out the distributions a little."
        },
        {
          "id": 451038,
          "postDate": "2019-01-06T10:31:06.223Z",
          "content": "<p>Yellow=R+G or Yellow=(R+G)/2. I've tried using later for hpa only. Is this correct?</p>",
          "rawMarkdown": "Yellow=R+G or Yellow=(R+G)/2. I've tried using later for hpa only. Is this correct?"
        },
        {
          "id": 451455,
          "postDate": "2019-01-07T06:03:58.950Z",
          "content": "<p><a href=\"/maw501\">@maw501</a> thanks not sure if weighted sampling is worth ceasing my iterative stratified kfold runs at this point. <a href=\"/valanm\">@valanm</a> i am using the yellow channel image provided in the hpa, and none that don't have a y-channel image. It seems pete is doing R+G=Y and and Shiv is doing (R+G)/2=Y. Dividing by 2 seems to be more correct, as an attempt to stay relatively close in scale to the other channels. However, pete's score may beg to differ, assuming he's not dividing by 2</p>",
          "rawMarkdown": "@maw501 thanks not sure if weighted sampling is worth ceasing my iterative stratified kfold runs at this point. @valanm i am using the yellow channel image provided in the hpa, and none that don't have a y-channel image. It seems pete is doing R+G=Y and and Shiv is doing (R+G)/2=Y. Dividing by 2 seems to be more correct, as an attempt to stay relatively close in scale to the other channels. However, pete's score may beg to differ, assuming he's not dividing by 2"
        }
      ]
    },
    {
      "id": 445269,
      "postDate": "2018-12-26T04:57:36.627Z",
      "content": "<p>I just wanted to say thanks. It's my first time trying CNN's and I only just started a couple days ago. With under 2 weeks to go, I'm still hoping to do well by learning from people like you. If only training didn't take forever haha.</p>",
      "rawMarkdown": "I just wanted to say thanks. It's my first time trying CNN's and I only just started a couple days ago. With under 2 weeks to go, I'm still hoping to do well by learning from people like you. If only training didn't take forever haha.",
      "votes": 1,
      "replies": [
        {
          "id": 445293,
          "postDate": "2018-12-26T06:30:33.457Z",
          "content": "<p>Thanks. To get a better score you need to do some more experiments. The current sharing from me (and my team) are just helping to save some time on doing duplicated jobs. (Try loss1, try loss2, try loss3... These might not be very helpful for learning and practice purpose...)</p>",
          "rawMarkdown": "Thanks. To get a better score you need to do some more experiments. The current sharing from me (and my team) are just helping to save some time on doing duplicated jobs. (Try loss1, try loss2, try loss3... These might not be very helpful for learning and practice purpose...)\n\n"
        },
        {
          "id": 446414,
          "postDate": "2018-12-28T03:01:58.827Z",
          "content": "<p>Thnx for sharing!\n  Did u train only using the HPA data (containing 70k+ samples). My  LB score went down after training using the HPA data... I wonder whether there are some additional process before training using the HPA data ?</p>",
          "rawMarkdown": "Thnx for sharing!\n  Did u train only using the HPA data (containing 70k+ samples). My  LB score went down after training using the HPA data... I wonder whether there are some additional process before training using the HPA data ?"
        }
      ]
    },
    {
      "id": 444999,
      "postDate": "2018-12-25T10:18:41.787Z",
      "content": "<p>Thanks for sharing, ManyFoldCV!<br>\nI have one question.<br>\n\"External data is not a must for 0.58+ (based on my teammate's result)\"<br>\nis this score model ensemble or single model?</p>",
      "rawMarkdown": "Thanks for sharing, ManyFoldCV!<br>\nI have one question.<br>\n\"External data is not a must for 0.58+ (based on my teammate's result)\"<br>\nis this score model ensemble or single model?",
      "votes": 1,
      "replies": [
        {
          "id": 445132,
          "postDate": "2018-12-25T18:09:31.640Z",
          "content": "<p>It should be a single model with 1024 input... It's from my teammate Pavel I think.</p>",
          "rawMarkdown": "It should be a single model with 1024 input... It's from my teammate Pavel I think.",
          "votes": 1
        },
        {
          "id": 445147,
          "postDate": "2018-12-25T18:49:34.143Z",
          "content": "<p>Single model with only train data reached 0.58+....it is so amazing.</p>",
          "rawMarkdown": "Single model with only train data reached 0.58+....it is so amazing.",
          "votes": 1
        },
        {
          "id": 445888,
          "postDate": "2018-12-27T07:05:59.227Z",
          "content": "<p>It is an ensembled result, but without the external data. Sorry about the wrong information.</p>",
          "rawMarkdown": "It is an ensembled result, but without the external data. Sorry about the wrong information.",
          "votes": 1
        },
        {
          "id": 445891,
          "postDate": "2018-12-27T07:14:55.737Z",
          "content": "<p>Don't worry, no problem.<br>\nThanks ManyFoldCV!<br>\nYour information is very helpful for me.</p>",
          "rawMarkdown": "Don't worry, no problem.<br>\nThanks ManyFoldCV!<br>\nYour information is very helpful for me."
        }
      ]
    },
    {
      "id": 445057,
      "postDate": "2018-12-25T14:06:08.290Z",
      "content": "<blockquote>\n  <ol>\n  <li>Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere…)  </li>\n  </ol>\n</blockquote>\n\n<p>I guess this is what you mentioned <br>\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374#437548\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374#437548</a></p>",
      "rawMarkdown": "&gt; 5. Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere…)  \n\nI guess this is what you mentioned  \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374#437548",
      "votes": 2,
      "replies": [
        {
          "id": 445131,
          "postDate": "2018-12-25T18:08:24.870Z",
          "content": "<p>Yes, that's it! Thanks for pointing it out.</p>",
          "rawMarkdown": "Yes, that's it! Thanks for pointing it out."
        }
      ]
    },
    {
      "id": 450408,
      "postDate": "2019-01-04T21:54:37.637Z",
      "content": "<p>The external data can be found in this discussion</p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a></p>",
      "rawMarkdown": "The external data can be found in this discussion\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984"
    },
    {
      "id": 445908,
      "postDate": "2018-12-27T07:38:01.753Z",
      "content": "<p>thanks for sharing...great help for people starting late</p>",
      "rawMarkdown": "thanks for sharing...great help for people starting late"
    },
    {
      "id": 445224,
      "postDate": "2018-12-26T02:17:36.753Z",
      "content": "<p>Thanks for this - very nice. \nQuestion: do you use all of the external data or do you throw away the more common classes?</p>",
      "rawMarkdown": "Thanks for this - very nice. \nQuestion: do you use all of the external data or do you throw away the more common classes?",
      "replies": [
        {
          "id": 445292,
          "postDate": "2018-12-26T06:28:21.353Z",
          "content": "<p>I used all of the external data.</p>",
          "rawMarkdown": "I used all of the external data."
        }
      ]
    },
    {
      "id": 445108,
      "postDate": "2018-12-25T16:36:13.983Z",
      "content": "<p>Thank you for your generous share at this stage of the competition.</p>\n\n<p>My understanding is that the train and test dataset do not come from same distribution. This can explain the confusion about \"Optimal threshold searching does not work for me\", \"Using sigmoid threshold 0.5 cannot produce decent results\" and the \"Weird observation\".</p>\n\n<p>BTW, are you using RGB or RGBY?</p>",
      "rawMarkdown": "Thank you for your generous share at this stage of the competition.\n\nMy understanding is that the train and test dataset do not come from same distribution. This can explain the confusion about \"Optimal threshold searching does not work for me\", \"Using sigmoid threshold 0.5 cannot produce decent results\" and the \"Weird observation\".\n\nBTW, are you using RGB or RGBY?\n",
      "replies": [
        {
          "id": 445129,
          "postDate": "2018-12-25T18:06:35.513Z",
          "content": "<p>I use RGBY. I initialized the extra depth conv layer with random initialization. Some mentioned you can copy the existing weights but I haven't tried it.</p>",
          "rawMarkdown": "I use RGBY. I initialized the extra depth conv layer with random initialization. Some mentioned you can copy the existing weights but I haven't tried it."
        },
        {
          "id": 445138,
          "postDate": "2018-12-25T18:17:17.783Z",
          "content": "<p>Plus, we had some discussion within the team and we are also considering that, is it possible there are some noise labels, at least, in the extenal data?</p>",
          "rawMarkdown": "Plus, we had some discussion within the team and we are also considering that, is it possible there are some noise labels, at least, in the extenal data?"
        },
        {
          "id": 445142,
          "postDate": "2018-12-25T18:32:14.710Z",
          "content": "<p>Do you see better results with yellow included? The problem is that it is not clear how to recover the yellow in the external data, isn't it, according to the discussion in the external data thread?</p>",
          "rawMarkdown": "Do you see better results with yellow included? The problem is that it is not clear how to recover the yellow in the external data, isn't it, according to the discussion in the external data thread?"
        },
        {
          "id": 445144,
          "postDate": "2018-12-25T18:36:50.977Z",
          "content": "<p>Some noise in the data and labels is fine for deep neural nets. The only exception I can think of is leaks, wrong labels in train data will result in wrong predictions in the corresponding test data if data are same but labels are not. I don't think this is issue because there is no leak in private lb</p>",
          "rawMarkdown": "Some noise in the data and labels is fine for deep neural nets. The only exception I can think of is leaks, wrong labels in train data will result in wrong predictions in the corresponding test data if data are same but labels are not. I don't think this is issue because there is no leak in private lb"
        },
        {
          "id": 445175,
          "postDate": "2018-12-25T20:36:55.567Z",
          "content": "<p>I just used the green channel from the yellow images in the external data. Use red channel gave slightly worse result. But I am not sure whether there are some randomness..</p>",
          "rawMarkdown": "I just used the green channel from the yellow images in the external data. Use red channel gave slightly worse result. But I am not sure whether there are some randomness..\n"
        },
        {
          "id": 445341,
          "postDate": "2018-12-26T08:41:57.337Z",
          "content": "<p>I also get worse result when use  only red channel in yellow images, do you know reasons? \ni maybe try used the green channel like you do</p>",
          "rawMarkdown": "I also get worse result when use  only red channel in yellow images, do you know reasons? \ni maybe try used the green channel like you do"
        }
      ]
    },
    {
      "id": 444976,
      "postDate": "2018-12-25T08:54:20.657Z",
      "content": "<p>hi! could do you tell me if your external data is processed specially?because i didn't got the better result with  external data</p>",
      "rawMarkdown": "hi! could do you tell me if your external data is processed specially?because i didn't got the better result with  external data",
      "replies": [
        {
          "id": 445210,
          "postDate": "2018-12-26T00:37:37.253Z",
          "content": "<p>I followed the comments from the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#latest-445200\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#latest-445200</a></p>",
          "rawMarkdown": "I followed the comments from the https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#latest-445200"
        },
        {
          "id": 445415,
          "postDate": "2018-12-26T12:07:28.860Z",
          "content": "<p>@ManyFoldCV ，Hi，thanks for sharing. \nSo, you are downloading RGB images and then convert each image to gray scale using <code>Image.open('img_HPAv18_some_name_blue.jpg').convert('L')</code> ?</p>",
          "rawMarkdown": "@ManyFoldCV ，Hi，thanks for sharing. \nSo, you are downloading RGB images and then convert each image to gray scale using `Image.open('img_HPAv18_some_name_blue.jpg').convert('L')` ?"
        },
        {
          "id": 445820,
          "postDate": "2018-12-27T05:04:28.767Z",
          "content": "<p>I extract only one channel from each image. In the link I pasted above you can find the answer.</p>",
          "rawMarkdown": "I extract only one channel from each image. In the link I pasted above you can find the answer."
        }
      ]
    },
    {
      "id": 450925,
      "postDate": "2019-01-06T04:49:04.383Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 446393,
      "postDate": "2018-12-28T01:50:44.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 445937,
      "postDate": "2018-12-27T08:38:19.333Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 444988,
      "postDate": "2018-12-25T09:18:26.887Z",
      "content": "<p>Thanks for sharing ! </p>",
      "rawMarkdown": "Thanks for sharing ! ",
      "votes": 3
    },
    {
      "id": 445883,
      "postDate": "2018-12-27T06:48:56.900Z",
      "content": "<p>Thank you for sharing :) </p>",
      "rawMarkdown": "Thank you for sharing :) ",
      "votes": 1
    },
    {
      "id": 445790,
      "postDate": "2018-12-27T04:13:56.650Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !\n\n",
      "votes": 1
    },
    {
      "id": 445323,
      "postDate": "2018-12-26T08:07:38.320Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 452862,
      "postDate": "2019-01-09T09:28:57.637Z",
      "content": "<p>Thanks a lot for sharing!</p>",
      "rawMarkdown": "Thanks a lot for sharing!"
    },
    {
      "id": 450558,
      "postDate": "2019-01-05T08:59:30.937Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing."
    },
    {
      "id": 448609,
      "postDate": "2019-01-01T16:29:03.973Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 448521,
      "postDate": "2019-01-01T12:08:57.283Z",
      "content": "<p>Thank you for your help</p>",
      "rawMarkdown": "Thank you for your help"
    },
    {
      "id": 448442,
      "postDate": "2019-01-01T07:43:43.757Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 446588,
      "postDate": "2018-12-28T10:15:53.627Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !"
    },
    {
      "id": 445956,
      "postDate": "2018-12-27T09:15:34.847Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing"
    },
    {
      "id": 445911,
      "postDate": "2018-12-27T07:42:02.347Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 445227,
      "postDate": "2018-12-26T02:27:40.607Z",
      "content": "<p>Thanks, for sharing.</p>",
      "rawMarkdown": "Thanks, for sharing."
    }
  ],
  "comments": [
    {
      "id": 444910,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-12-25T06:24:05.597000",
      "content": "<p>@ManyFoldCV Thanks for the hints!</p>\n\n<p>What public/validation score do you get (or you estimate) if we only use kaggle train data without leak, without external for a single model without TTA, without ensemble?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 445134,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-25T18:11:50.310000",
          "content": "<p>I've learned so much from you Heng :D</p>\n\n<p>Using only the official data I got a public score 0.472. At that time my local f1 formula had a problem so I don't have a local stat.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 446001,
      "author_name": "Ildoo Kim",
      "author_url": "",
      "post_date": "2018-12-27T10:34:03.447000",
      "content": "<p>How do you ensemble models? Caculating average of outputs? Voting?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 446010,
          "author_name": "ChienYiChi",
          "author_url": "",
          "post_date": "2018-12-27T10:44:43.953000",
          "content": "<p>for me，simply averaging the output worse than single fold model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446013,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2018-12-27T10:47:14.070000",
          "content": "<p>There are two options:\n1. Your k-fold split is bad (you probably need more subtle stratification or something like that).\n2. You single fold model is by chance overfitted to the public LB, but average of folds is still more robust, so it will be better on private LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 446033,
          "author_name": "ChienYiChi",
          "author_url": "",
          "post_date": "2018-12-27T11:27:45.773000",
          "content": "<p>Thanks for your reply\nAbout the k fold split，I use the Multilabel Stratification  python package by @Trent ,so may be that no the problem.\nhope my model can survive in private LB. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446166,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-27T16:29:07.460000",
          "content": "<p>Average or weighted average the proba</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 445922,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2018-12-27T08:12:37.350000",
      "content": "<p>I think this is the one you are looking for:</p>\n\n<pre>train_df_orig=train_df.copy()    \nlows = [15,15,15,8,9,10,8,9,10,8,9,10,17,20,24,26,15,27,15,20,24,17,8,15,27,27,27]\nfor i in lows:\n    target = str(i)\n    indicies = train_df_orig.loc[train_df_orig['Target'] == target].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.startswith(target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.endswith(\" \"+target)].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.contains(\" \"+target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n</pre>",
      "votes": 4,
      "replies": [
        {
          "id": 446169,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-27T16:29:53.103000",
          "content": "<p>Thanks <a href=\"/ldm314\">@ldm314</a>, that's the one.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446460,
          "author_name": "Phil Butcher",
          "author_url": "",
          "post_date": "2018-12-28T05:22:54.550000",
          "content": "<p>Thanks for the snippet. Out of curiosity how did you choose the frequency to oversample the classes in your 'lows' list?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 446234,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2018-12-27T18:46:17.470000",
      "content": "<p>So what is the 'correct' way of processing the external data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 446302,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-27T20:52:19.863000",
          "content": "<p>I processed it in the way we discussed on previous threads - basically a straight copy of the corresponding channels, adding R+G for yellow. When I added the external data to the challenge data (a LOT of data) my results were not optimal. So I tried adding the rare classes first and kept adding until the results started to degrade. That helped. </p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 446317,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-27T21:32:04.460000",
          "content": "<p>Hi <a href=\"/maw501\">@maw501</a>, pete is giving the correct way :DDD</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446322,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-27T21:54:45.140000",
          "content": "<p>Thanks Pete and ManyFoldCV.</p>\n\n<p>So basically download using something like the following (note the line marked <em>*</em> which just extracts a single channel):</p>\n\n<pre><code>def download(id_):\n    try:\n        image_dir, _, image_id = id_.partition('_')\n        arr_out = np.zeros((512,512,3))\n        for i, c in enumerate(['red', 'green', 'blue']):\n            url = f'http://v18.proteinatlas.org/images/{image_dir}/{image_id}_{c}.jpg'\n            r = requests.get(url)\n            image = Image.open(BytesIO(r.content)).resize((512, 512), PIL.Image.LANCZOS)\n            arr_out[:,:,i] = (np.array(image)[:,:,i]).astype('uint8') # ***\n        img_pil = Image.fromarray(arr_out.astype('uint8')) \n        img_pil.save(f'{hpa_dir}/{id_}.png')\n    except:\n        print(f'{id_} broke...')\n</code></pre>\n\n<p>Are you adding hpa to your validation sets as well?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446481,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-28T06:43:00.210000",
          "content": "<p>I did. I worried about possible duplicates but did not act on it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446557,
          "author_name": "Shiv Gowda",
          "author_url": "",
          "post_date": "2018-12-28T09:25:48.530000",
          "content": "<p>Here is what I use to download and store the images for protein of interest(green) and the other 3 markers(red/blue/yellow).  I am not combining them into one image like you do, rather I am saving them as separate images just like in the train and the test set.</p>\n\n<pre><code>def downloadImages(id):\n    colors = ['red','green','blue','yellow']\n    img_path = v18_url + id.replace('_', '/', 1)\n    for i, color in enumerate(colors):\n        response = requests.get(img_path + \"_\" + color + \".jpg\", allow_redirects=True)\n        im = Image.open(BytesIO(response.content)).resize((512,512))\n        arr = np.asarray(im)\n        if i == 3: # Yellow\n            im = Image.fromarray(((arr[:,:,0] + arr[:,:,1])/2).astype(np.uint8))\n       else :\n           im = Image.fromarray(arr[:,:,i])\n        im.save(DATA_DIR + id + \"_\" + color + \".png\")\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446788,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-28T17:38:10.737000",
          "content": "<p><a href=\"/maw501\">@maw501</a> I suggest to save the RGB space JPEG image to the disk first then process it if you have enough space. It will be about ~60G, 280k images. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446813,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-28T18:21:09.643000",
          "content": "<p>Hi <a href=\"/manyfoldcv\">@manyfoldcv</a> - I've already got the hpa...just finding training pretty unstable with it atm. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450287,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2019-01-04T16:27:17.060000",
          "content": "<p>which classes do you keep? From hpa, I've added Nuclear membrane and every class that has less samples which gave 0.55 for 5-fold model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 450295,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2019-01-04T16:38:58.980000",
          "content": "<p>Keeping all at the minute but using a sampling strategy to even out the distributions a little.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451038,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-01-06T10:31:06.223000",
          "content": "<p>Yellow=R+G or Yellow=(R+G)/2. I've tried using later for hpa only. Is this correct?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451455,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2019-01-07T06:03:58.950000",
          "content": "<p><a href=\"/maw501\">@maw501</a> thanks not sure if weighted sampling is worth ceasing my iterative stratified kfold runs at this point. <a href=\"/valanm\">@valanm</a> i am using the yellow channel image provided in the hpa, and none that don't have a y-channel image. It seems pete is doing R+G=Y and and Shiv is doing (R+G)/2=Y. Dividing by 2 seems to be more correct, as an attempt to stay relatively close in scale to the other channels. However, pete's score may beg to differ, assuming he's not dividing by 2</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 445269,
      "author_name": "roush",
      "author_url": "",
      "post_date": "2018-12-26T04:57:36.627000",
      "content": "<p>I just wanted to say thanks. It's my first time trying CNN's and I only just started a couple days ago. With under 2 weeks to go, I'm still hoping to do well by learning from people like you. If only training didn't take forever haha.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 445293,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-26T06:30:33.457000",
          "content": "<p>Thanks. To get a better score you need to do some more experiments. The current sharing from me (and my team) are just helping to save some time on doing duplicated jobs. (Try loss1, try loss2, try loss3... These might not be very helpful for learning and practice purpose...)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446414,
          "author_name": "Hy li",
          "author_url": "",
          "post_date": "2018-12-28T03:01:58.827000",
          "content": "<p>Thnx for sharing!\n  Did u train only using the HPA data (containing 70k+ samples). My  LB score went down after training using the HPA data... I wonder whether there are some additional process before training using the HPA data ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 444999,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2018-12-25T10:18:41.787000",
      "content": "<p>Thanks for sharing, ManyFoldCV!<br>\nI have one question.<br>\n\"External data is not a must for 0.58+ (based on my teammate's result)\"<br>\nis this score model ensemble or single model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 445132,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-25T18:09:31.640000",
          "content": "<p>It should be a single model with 1024 input... It's from my teammate Pavel I think.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 445147,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2018-12-25T18:49:34.143000",
          "content": "<p>Single model with only train data reached 0.58+....it is so amazing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 445888,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-27T07:05:59.227000",
          "content": "<p>It is an ensembled result, but without the external data. Sorry about the wrong information.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 445891,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2018-12-27T07:14:55.737000",
          "content": "<p>Don't worry, no problem.<br>\nThanks ManyFoldCV!<br>\nYour information is very helpful for me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 445057,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2018-12-25T14:06:08.290000",
      "content": "<blockquote>\n  <ol>\n  <li>Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere…)  </li>\n  </ol>\n</blockquote>\n\n<p>I guess this is what you mentioned <br>\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374#437548\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374#437548</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 445131,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-25T18:08:24.870000",
          "content": "<p>Yes, that's it! Thanks for pointing it out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 450408,
      "author_name": "joe1Dataset",
      "author_url": "",
      "post_date": "2019-01-04T21:54:37.637000",
      "content": "<p>The external data can be found in this discussion</p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445908,
      "author_name": "Prachi",
      "author_url": "",
      "post_date": "2018-12-27T07:38:01.753000",
      "content": "<p>thanks for sharing...great help for people starting late</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445224,
      "author_name": "pete",
      "author_url": "",
      "post_date": "2018-12-26T02:17:36.753000",
      "content": "<p>Thanks for this - very nice. \nQuestion: do you use all of the external data or do you throw away the more common classes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 445292,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-26T06:28:21.353000",
          "content": "<p>I used all of the external data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 445108,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2018-12-25T16:36:13.983000",
      "content": "<p>Thank you for your generous share at this stage of the competition.</p>\n\n<p>My understanding is that the train and test dataset do not come from same distribution. This can explain the confusion about \"Optimal threshold searching does not work for me\", \"Using sigmoid threshold 0.5 cannot produce decent results\" and the \"Weird observation\".</p>\n\n<p>BTW, are you using RGB or RGBY?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 445129,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-25T18:06:35.513000",
          "content": "<p>I use RGBY. I initialized the extra depth conv layer with random initialization. Some mentioned you can copy the existing weights but I haven't tried it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445138,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-25T18:17:17.783000",
          "content": "<p>Plus, we had some discussion within the team and we are also considering that, is it possible there are some noise labels, at least, in the extenal data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445142,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2018-12-25T18:32:14.710000",
          "content": "<p>Do you see better results with yellow included? The problem is that it is not clear how to recover the yellow in the external data, isn't it, according to the discussion in the external data thread?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445144,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2018-12-25T18:36:50.977000",
          "content": "<p>Some noise in the data and labels is fine for deep neural nets. The only exception I can think of is leaks, wrong labels in train data will result in wrong predictions in the corresponding test data if data are same but labels are not. I don't think this is issue because there is no leak in private lb</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445175,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-25T20:36:55.567000",
          "content": "<p>I just used the green channel from the yellow images in the external data. Use red channel gave slightly worse result. But I am not sure whether there are some randomness..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445341,
          "author_name": "zhengjie",
          "author_url": "",
          "post_date": "2018-12-26T08:41:57.337000",
          "content": "<p>I also get worse result when use  only red channel in yellow images, do you know reasons? \ni maybe try used the green channel like you do</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 444976,
      "author_name": "ML sutdy ha ha",
      "author_url": "",
      "post_date": "2018-12-25T08:54:20.657000",
      "content": "<p>hi! could do you tell me if your external data is processed specially?because i didn't got the better result with  external data</p>",
      "votes": 0,
      "replies": [
        {
          "id": 445210,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-26T00:37:37.253000",
          "content": "<p>I followed the comments from the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#latest-445200\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#latest-445200</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445415,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-12-26T12:07:28.860000",
          "content": "<p>@ManyFoldCV ，Hi，thanks for sharing. \nSo, you are downloading RGB images and then convert each image to gray scale using <code>Image.open('img_HPAv18_some_name_blue.jpg').convert('L')</code> ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445820,
          "author_name": "Kulbear",
          "author_url": "",
          "post_date": "2018-12-27T05:04:28.767000",
          "content": "<p>I extract only one channel from each image. In the link I pasted above you can find the answer.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 450925,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-06T04:49:04.383000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446393,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T01:50:44.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445937,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-27T08:38:19.333000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 444988,
      "author_name": "Spytensor",
      "author_url": "",
      "post_date": "2018-12-25T09:18:26.887000",
      "content": "<p>Thanks for sharing ! </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 445883,
      "author_name": "Chandramowli J",
      "author_url": "",
      "post_date": "2018-12-27T06:48:56.900000",
      "content": "<p>Thank you for sharing :) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 445790,
      "author_name": "Aaveg Barole",
      "author_url": "",
      "post_date": "2018-12-27T04:13:56.650000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 445323,
      "author_name": "zhangboshen",
      "author_url": "",
      "post_date": "2018-12-26T08:07:38.320000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 452862,
      "author_name": "irina",
      "author_url": "",
      "post_date": "2019-01-09T09:28:57.637000",
      "content": "<p>Thanks a lot for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450558,
      "author_name": "Abe Tatsuhiro",
      "author_url": "",
      "post_date": "2019-01-05T08:59:30.937000",
      "content": "<p>Thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448609,
      "author_name": "Norman Chen",
      "author_url": "",
      "post_date": "2019-01-01T16:29:03.973000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448521,
      "author_name": "hiroshi2727",
      "author_url": "",
      "post_date": "2019-01-01T12:08:57.283000",
      "content": "<p>Thank you for your help</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448442,
      "author_name": "Just for fun",
      "author_url": "",
      "post_date": "2019-01-01T07:43:43.757000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446588,
      "author_name": "chchchuang ",
      "author_url": "",
      "post_date": "2018-12-28T10:15:53.627000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445956,
      "author_name": "Femi",
      "author_url": "",
      "post_date": "2018-12-27T09:15:34.847000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445911,
      "author_name": "weizhipeng",
      "author_url": "",
      "post_date": "2018-12-27T07:42:02.347000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445227,
      "author_name": "xftts",
      "author_url": "",
      "post_date": "2018-12-26T02:27:40.607000",
      "content": "<p>Thanks, for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "444881": "1. External data is not a must for 0.58+ (based on my teammate's result), but it does help (make your life easier).\n2. Image size matters, use larger image size if you can.\n3. Personally found Adam works well in our case... You can play with it...\n4. Augmentation (rotation, flip and shear seem to be enough, maybe colorjitter?)\n5. Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere...)\n6. Reduce your learning rate, with a proper schedule, to escape from local minima...\n7. I personally always got some worse results when I tried to combine simple BCE with some other fancy loss (Focal, Soft F1 loss...)...\n8. Optimal threshold searching does not work for me, up to now.\n9. Ensemble your many folds (and different models if possible).\n10. I tried to ensemble my 256 5-fold and got 0.556, then I try to ensemble my 512 5-fold and got 586. I added the 256 result to the 512 result I can reach 0.59X (forgot) ...\n11. The above score can be reached with BN-Inception...\n12. I use https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812 as my starter template as I joined this competition pretty late, like, about 2 weeks ago.\n\nBonus: a weird observation --&gt; using varied batch size, when the model converge to some BCE (say, 0.7), the exact same settings and model, the one with larger batch size will give better local f1 validation score, though the LB scores are not very different, potentially there are some noise labels (in the external?) ?\n\nOther info:\nI use one 1080Ti + 1080, and some GCP credits to use P100.\n\nMy question:\n1. Using sigmoid threshold 0.5 cannot produce decent results, I am using some smaller value but I'd like to understand why this happens, any help is appreciated.\n\n\nUpdate 12-25:\nFind the \"correct\" way of processing the external data is critical!",
    "444910": "@ManyFoldCV Thanks for the hints!\n\nWhat public/validation score do you get (or you estimate) if we only use kaggle train data without leak, without external for a single model without TTA, without ensemble?",
    "446001": "How do you ensemble models? Caculating average of outputs? Voting?",
    "445922": "I think this is the one you are looking for:\n<pre>train_df_orig=train_df.copy()    \nlows = [15,15,15,8,9,10,8,9,10,8,9,10,17,20,24,26,15,27,15,20,24,17,8,15,27,27,27]\nfor i in lows:\n    target = str(i)\n    indicies = train_df_orig.loc[train_df_orig['Target'] == target].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.startswith(target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.endswith(\" \"+target)].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.contains(\" \"+target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n</pre>",
    "446234": "So what is the 'correct' way of processing the external data?",
    "445269": "I just wanted to say thanks. It's my first time trying CNN's and I only just started a couple days ago. With under 2 weeks to go, I'm still hoping to do well by learning from people like you. If only training didn't take forever haha.",
    "444999": "Thanks for sharing, ManyFoldCV!<br>\nI have one question.<br>\n\"External data is not a must for 0.58+ (based on my teammate's result)\"<br>\nis this score model ensemble or single model?",
    "445057": "&gt; 5. Oversampling (I use the code snippet from @brian , but I cannot find the link, it is in the dicussion forum, somewhere…)  \n\nI guess this is what you mentioned  \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74374#437548",
    "450408": "The external data can be found in this discussion\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984",
    "445908": "thanks for sharing...great help for people starting late",
    "445224": "Thanks for this - very nice. \nQuestion: do you use all of the external data or do you throw away the more common classes?",
    "445108": "Thank you for your generous share at this stage of the competition.\n\nMy understanding is that the train and test dataset do not come from same distribution. This can explain the confusion about \"Optimal threshold searching does not work for me\", \"Using sigmoid threshold 0.5 cannot produce decent results\" and the \"Weird observation\".\n\nBTW, are you using RGB or RGBY?\n",
    "444976": "hi! could do you tell me if your external data is processed specially?because i didn't got the better result with  external data",
    "450925": "",
    "446393": "",
    "445937": "",
    "444988": "Thanks for sharing ! ",
    "445883": "Thank you for sharing :) ",
    "445790": "Thanks for sharing !\n\n",
    "445323": "Thanks for sharing!",
    "452862": "Thanks a lot for sharing!",
    "450558": "Thank you for sharing.",
    "448609": "Thanks",
    "448521": "Thank you for your help",
    "448442": "Thanks for sharing.",
    "446588": "Thanks for sharing !",
    "445956": "Thank you for sharing",
    "445911": "Thanks for sharing",
    "445227": "Thanks, for sharing."
  }
}