{
  "id": 76820,
  "title": "Something about reaching 0.6+ with single model",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/76820",
  "author_name": "pascal1129",
  "post_date": "2019-01-07T05:48:17.661000",
  "votes": 37,
  "comment_count": 35,
  "views": 0,
  "content": "<p>0、My situation: single GTX1080ti (sometimes 2), PyTorch 1.0;</p>\n\n<p>1、My best single model now can get 0.59+, and i think reaching 0.60 is possible;</p>\n\n<p>2、I use SGD, and i think lr schedule is important. I trained each model for about 30-35 epochs. My base_lr is 0.20, maybe is too big; (In addition, my batchsize is 32*8 with accumulating gradenit tech);</p>\n\n<p>3、I am sure that searching threshold  for each class is quite uesful;</p>\n\n<p>4、rgb is better than rgby, based on my experimental results ;</p>\n\n<p>5、I wasetd too much time tring advanced networks, sucha as SeNet, RseNeXt. Now I only use Res34 with 512 size, i think res18 and res34 is quite enough;</p>\n\n<p>6、I use upsample + downsample + log-dammped weights;</p>\n\n<p>7、TTA is useful;</p>\n\n<p>8、I think external data is not essential to getting 0.60;</p>\n\n<p>9、I reproduced 2 papers, which bring me 2 percent increase;</p>\n\n<p>10、My model has used too much regularization tricks to prevent overfitting, which restricted my public leadboad score, so I gave up using dropout, resize and crop. It may be a little late for me to realize the situation.</p>\n\n<p>11、Although dropout can be deleted, BN1d is essential to the fc layer in my network. I have tried to delete BN1d in the fc layer, and then my loss boomed, so don't delete your bn layer and have a try in other ways;</p>\n\n<p>I spent quite lots of time to find parameters that work for my network, which is painful. The deadline is comming, and i have other things to do, i may have no eonuth time to finish the competition, so i share my thougts, which i hope can bring you improvements.</p>",
  "messages": [
    {
      "id": 451445,
      "postDate": "2019-01-07T05:48:17.660Z",
      "content": "<p>0、My situation: single GTX1080ti (sometimes 2), PyTorch 1.0;</p>\n\n<p>1、My best single model now can get 0.59+, and i think reaching 0.60 is possible;</p>\n\n<p>2、I use SGD, and i think lr schedule is important. I trained each model for about 30-35 epochs. My base_lr is 0.20, maybe is too big; (In addition, my batchsize is 32*8 with accumulating gradenit tech);</p>\n\n<p>3、I am sure that searching threshold  for each class is quite uesful;</p>\n\n<p>4、rgb is better than rgby, based on my experimental results ;</p>\n\n<p>5、I wasetd too much time tring advanced networks, sucha as SeNet, RseNeXt. Now I only use Res34 with 512 size, i think res18 and res34 is quite enough;</p>\n\n<p>6、I use upsample + downsample + log-dammped weights;</p>\n\n<p>7、TTA is useful;</p>\n\n<p>8、I think external data is not essential to getting 0.60;</p>\n\n<p>9、I reproduced 2 papers, which bring me 2 percent increase;</p>\n\n<p>10、My model has used too much regularization tricks to prevent overfitting, which restricted my public leadboad score, so I gave up using dropout, resize and crop. It may be a little late for me to realize the situation.</p>\n\n<p>11、Although dropout can be deleted, BN1d is essential to the fc layer in my network. I have tried to delete BN1d in the fc layer, and then my loss boomed, so don't delete your bn layer and have a try in other ways;</p>\n\n<p>I spent quite lots of time to find parameters that work for my network, which is painful. The deadline is comming, and i have other things to do, i may have no eonuth time to finish the competition, so i share my thougts, which i hope can bring you improvements.</p>",
      "rawMarkdown": "0、My situation: single GTX1080ti (sometimes 2), PyTorch 1.0;\n\n1、My best single model now can get 0.59+, and i think reaching 0.60 is possible;\n\n2、I use SGD, and i think lr schedule is important. I trained each model for about 30-35 epochs. My base_lr is 0.20, maybe is too big; (In addition, my batchsize is 32*8 with accumulating gradenit tech);\n\n3、I am sure that searching threshold  for each class is quite uesful;\n\n4、rgb is better than rgby, based on my experimental results ;\n\n5、I wasetd too much time tring advanced networks, sucha as SeNet, RseNeXt. Now I only use Res34 with 512 size, i think res18 and res34 is quite enough;\n\n6、I use upsample + downsample + log-dammped weights;\n\n7、TTA is useful;\n\n8、I think external data is not essential to getting 0.60;\n\n9、I reproduced 2 papers, which bring me 2 percent increase;\n\n10、My model has used too much regularization tricks to prevent overfitting, which restricted my public leadboad score, so I gave up using dropout, resize and crop. It may be a little late for me to realize the situation.\n\n11、Although dropout can be deleted, BN1d is essential to the fc layer in my network. I have tried to delete BN1d in the fc layer, and then my loss boomed, so don't delete your bn layer and have a try in other ways;\n\nI spent quite lots of time to find parameters that work for my network, which is painful. The deadline is comming, and i have other things to do, i may have no eonuth time to finish the competition, so i share my thougts, which i hope can bring you improvements.\n",
      "votes": 37
    },
    {
      "id": 453183,
      "postDate": "2019-01-09T20:22:39.543Z",
      "content": "<p>thx for sharing, I wonder how your local F1 and loss (val and train) look? If it's too much to ask, no worries.</p>",
      "rawMarkdown": "thx for sharing, I wonder how your local F1 and loss (val and train) look? If it's too much to ask, no worries.",
      "votes": 1
    },
    {
      "id": 452110,
      "postDate": "2019-01-08T07:48:30.547Z",
      "content": "<p>hi，can u add my qq 1124963678， i need some help.</p>",
      "rawMarkdown": "hi，can u add my qq 1124963678， i need some help.",
      "votes": -6
    },
    {
      "id": 452762,
      "postDate": "2019-01-09T06:03:30.267Z",
      "content": "<p><a href=\"/pascal1129\">@pascal1129</a>,Thx for sharing,and could you give some explanation about item 6 \"upsample + downsample + log-dammped weights\",I don't understand very well.</p>",
      "rawMarkdown": "@pascal1129,Thx for sharing,and could you give some explanation about item 6 \"upsample + downsample + log-dammped weights\",I don't understand very well.",
      "replies": [
        {
          "id": 452785,
          "postDate": "2019-01-09T06:44:24.823Z",
          "content": "<p>add the samples of rare class, reduce the samples of class which is too much, give loss function a weight in order to balance differenet classes</p>",
          "rawMarkdown": "add the samples of rare class, reduce the samples of class which is too much, give loss function a weight in order to balance differenet classes"
        }
      ]
    },
    {
      "id": 451686,
      "postDate": "2019-01-07T13:44:28.827Z",
      "content": "<p>Thanks for sharing !\nCan you list the two papers you mentioned ? Probably won't have time to try but I am curious. </p>",
      "rawMarkdown": "Thanks for sharing !\nCan you list the two papers you mentioned ? Probably won't have time to try but I am curious. ",
      "replies": [
        {
          "id": 451694,
          "postDate": "2019-01-07T14:04:09.937Z",
          "content": "<p>After this competition, i will share my solution.</p>",
          "rawMarkdown": "After this competition, i will share my solution.",
          "votes": 4
        }
      ]
    },
    {
      "id": 451673,
      "postDate": "2019-01-07T13:22:00.887Z",
      "content": "<p>Hey <a href=\"/pascal1129\">@pascal1129</a> </p>\n\n<p>How do you handle BatchNorm with gradient accumulation? I thought they were incompatible and you would have to use GroupNorm or get rid of BatchNorm to accumulate gradients. </p>",
      "rawMarkdown": "Hey @pascal1129 \n\nHow do you handle BatchNorm with gradient accumulation? I thought they were incompatible and you would have to use GroupNorm or get rid of BatchNorm to accumulate gradients. ",
      "replies": [
        {
          "id": 451678,
          "postDate": "2019-01-07T13:28:42.367Z",
          "content": "<p>Can you proveide more infomation? I haven't thought about the problem before. I just use BN in the model structure, and used accumulating gradeint in the train stage. I don't know what is needed to be paid attention to.</p>",
          "rawMarkdown": "Can you proveide more infomation? I haven't thought about the problem before. I just use BN in the model structure, and used accumulating gradeint in the train stage. I don't know what is needed to be paid attention to."
        },
        {
          "id": 451855,
          "postDate": "2019-01-07T19:20:23.417Z",
          "content": "<p>As far as I know, batch norm statistics get updated on each forward pass, so no problem if you don't do .backward() every time.</p>",
          "rawMarkdown": "As far as I know, batch norm statistics get updated on each forward pass, so no problem if you don't do .backward() every time.",
          "votes": 1
        },
        {
          "id": 452367,
          "postDate": "2019-01-08T16:16:56.543Z",
          "content": "<p>Thanks.</p>",
          "rawMarkdown": "Thanks."
        },
        {
          "id": 452816,
          "postDate": "2019-01-09T07:38:03.310Z",
          "content": "<p>Took me a minute to get back to this, but thanks for the responses. I retried using gradient accumulation and it is working as expected. Thanks! I had another issue going on at the first time I tried to accumulate gradients.</p>",
          "rawMarkdown": "Took me a minute to get back to this, but thanks for the responses. I retried using gradient accumulation and it is working as expected. Thanks! I had another issue going on at the first time I tried to accumulate gradients."
        },
        {
          "id": 452914,
          "postDate": "2019-01-09T11:05:51.503Z",
          "content": "<p>would it be ok like this?\n`</p>\n\n<pre><code>loss += loss / accumulation_steps\n\nif (i+1) % accumulation_steps == 0:  \n\n     loss.backward()\n\n     optimizer.step()             \n\n     model.zero_grad()`\n</code></pre>",
          "rawMarkdown": "would it be ok like this?\n`\n\n    loss += loss / accumulation_steps\n\n    if (i+1) % accumulation_steps == 0:  \n\n         loss.backward()\n\n         optimizer.step()             \n\n         model.zero_grad()`",
          "votes": 1
        },
        {
          "id": 453025,
          "postDate": "2019-01-09T14:31:37.580Z",
          "content": "<p>```\nfor i,(images,target) in enumerate(train_loader):</p>\n\n<pre><code># 1. input output\nimages = images.cuda(non_blocking=True)\ntarget = torch.from_numpy(np.array(target)).float().cuda(non_blocking=True)\noutputs = model(images)\nloss = criterion(outputs,target)\n\n# 2.1 loss regularization\nloss = loss/accumulation_steps   \n\n# 2.2 back propagation\nloss.backward()\n\n# 3. update parameters of net\nif(i%accumulation_steps)==0:\n    optimizer.step()        # update parameters of net\n    optimizer.zero_grad()   # reset gradient\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "```\nfor i,(images,target) in enumerate(train_loader):\n\n    # 1. input output\n    images = images.cuda(non_blocking=True)\n    target = torch.from_numpy(np.array(target)).float().cuda(non_blocking=True)\n    outputs = model(images)\n    loss = criterion(outputs,target)\n\n    # 2.1 loss regularization\n    loss = loss/accumulation_steps   \n\n    # 2.2 back propagation\n    loss.backward()\n\n    # 3. update parameters of net\n    if(i%accumulation_steps)==0:\n        optimizer.step()        # update parameters of net\n        optimizer.zero_grad()   # reset gradient\n```",
          "votes": 3
        },
        {
          "id": 454670,
          "postDate": "2019-01-12T00:33:16.767Z",
          "content": "<p>Thx for sharing the code. Is there any reference paper of this, I’m new to DL :)</p>",
          "rawMarkdown": "Thx for sharing the code. Is there any reference paper of this, I’m new to DL :)"
        }
      ]
    },
    {
      "id": 451638,
      "postDate": "2019-01-07T12:09:32.780Z",
      "content": "<p>Thx  for sharing！  May I ask what‘s your TTA strategy in detail？ Cause  TTA  lower my LB score = = </p>",
      "rawMarkdown": "Thx  for sharing！  May I ask what‘s your TTA strategy in detail？ Cause  TTA  lower my LB score = = ",
      "replies": [
        {
          "id": 451645,
          "postDate": "2019-01-07T12:31:33.783Z",
          "content": "<p>Indeed, several days aog, TTA lower my score too. After improving my model, TTA works.</p>\n\n<p>I only use HorizontalFlip and VerticalFlip in TTA.</p>\n\n<p>For more details about TTA transform, you can see my github, Focus on the input[6]:\n<a href=\"https://github.com/pascal1129/CV_Notes/blob/master/codes/torchvision.transforms.ipynb\">https://github.com/pascal1129/CV_Notes/blob/master/codes/torchvision.transforms.ipynb</a></p>",
          "rawMarkdown": "Indeed, several days aog, TTA lower my score too. After improving my model, TTA works.\n\nI only use HorizontalFlip and VerticalFlip in TTA.\n\nFor more details about TTA transform, you can see my github, Focus on the input[6]:\nhttps://github.com/pascal1129/CV_Notes/blob/master/codes/torchvision.transforms.ipynb"
        }
      ]
    },
    {
      "id": 451524,
      "postDate": "2019-01-07T08:24:05.340Z",
      "content": "<p>Can U  share  the score  without extranal data？</p>",
      "rawMarkdown": "Can U  share  the score  without extranal data？",
      "replies": [
        {
          "id": 451529,
          "postDate": "2019-01-07T08:29:47.537Z",
          "content": "<p>I did not record the relevant results.</p>",
          "rawMarkdown": "I did not record the relevant results."
        }
      ]
    },
    {
      "id": 451521,
      "postDate": "2019-01-07T08:21:14.667Z",
      "content": "<p>Thx for  sharing, Could you share your lr schedule strategy? </p>",
      "rawMarkdown": "Thx for  sharing, Could you share your lr schedule strategy? ",
      "replies": [
        {
          "id": 451528,
          "postDate": "2019-01-07T08:27:50.957Z",
          "content": "<p>I used compose of <a href=\"https://pytorch.org/docs/stable/optim.html#how-to-adjust-learning-rate\">https://pytorch.org/docs/stable/optim.html#how-to-adjust-learning-rate</a></p>",
          "rawMarkdown": "I used compose of https://pytorch.org/docs/stable/optim.html#how-to-adjust-learning-rate"
        }
      ]
    },
    {
      "id": 451496,
      "postDate": "2019-01-07T07:54:49.980Z",
      "content": "<p>Thanks for sharing! I wanted to ask one question:\n How to search threshold for each class?</p>",
      "rawMarkdown": "Thanks for sharing! I wanted to ask one question:\n How to search threshold for each class?",
      "replies": [
        {
          "id": 451512,
          "postDate": "2019-01-07T08:11:20.653Z",
          "content": "<p>For example, try different thresholds on one class form 0.01 to 0.99 by step 0.01 , and use the threshold which has the best f1-score.</p>\n\n<p><img src=\"https://s2.ax1x.com/2019/01/07/FbIauT.png\" alt=\"FbIauT.png\"></p>",
          "rawMarkdown": "For example, try different thresholds on one class form 0.01 to 0.99 by step 0.01 , and use the threshold which has the best f1-score.\n\n\n\n![FbIauT.png](https://s2.ax1x.com/2019/01/07/FbIauT.png)",
          "votes": 1
        },
        {
          "id": 451530,
          "postDate": "2019-01-07T08:29:49.253Z",
          "content": "<p>Trying from 0.01 to 0.99 by step 0.01 have a combination of 100^28. This is an intractable.</p>",
          "rawMarkdown": "Trying from 0.01 to 0.99 by step 0.01 have a combination of 100^28. This is an intractable."
        },
        {
          "id": 451534,
          "postDate": "2019-01-07T08:34:44.717Z",
          "content": "<p>If you used suitable method, I think is 100 x 28, not 100^28.</p>",
          "rawMarkdown": "If you used suitable method, I think is 100 x 28, not 100^28.",
          "votes": 1
        },
        {
          "id": 451556,
          "postDate": "2019-01-07T09:05:01.740Z",
          "content": "<p>Thanks for your help!</p>",
          "rawMarkdown": "Thanks for your help!"
        },
        {
          "id": 451599,
          "postDate": "2019-01-07T10:30:58.383Z",
          "content": "<p>search individually for each class and take the best threshold is fine imo</p>",
          "rawMarkdown": "search individually for each class and take the best threshold is fine imo"
        },
        {
          "id": 451610,
          "postDate": "2019-01-07T10:54:19.293Z",
          "content": "<p>I just tried  threshold searching, but it lowered my lb... I don't know why</p>\n\n<p>I just share the code here</p>\n\n<p>```\nimport numpy as np\nfrom sklearn.metrics import f1_score\nnum_classes=28\ndef optimise_f1_thresholds(y, p, verbose=True, resolution=100):\n    def mf(x):\n        p1 = np.zeros_like(p)\n        for i in range(num_classes):\n            p1[:, i] = (p[:, i] &gt; x[i]).astype(np.int)\n        score = f1_score(y, p1, average='samples')\n        return score</p>\n\n<pre><code>x = [0.2]*num_classes\nfor i in range(num_classes):\n    best_i1 = 0\n    best_score = 0\n    for i1 in range(resolution):\n        i1 /= resolution\n        x[i] = i1\n        score = mf(x)\n        if score &gt; best_score:\n            best_i1 = i1\n            best_score = score\n    x[i] = best_i1\n    if verbose:\n        print(i, best_i1, best_score)\n\nreturn x\n</code></pre>\n\n<p>optimise_f1_thresholds(target, predicted)\n```</p>",
          "rawMarkdown": "I just tried  threshold searching, but it lowered my lb... I don't know why\n\nI just share the code here\n\n```\nimport numpy as np\nfrom sklearn.metrics import f1_score\nnum_classes=28\ndef optimise_f1_thresholds(y, p, verbose=True, resolution=100):\n    def mf(x):\n        p1 = np.zeros_like(p)\n        for i in range(num_classes):\n            p1[:, i] = (p[:, i] &gt; x[i]).astype(np.int)\n        score = f1_score(y, p1, average='samples')\n        return score\n\n    x = [0.2]*num_classes\n    for i in range(num_classes):\n        best_i1 = 0\n        best_score = 0\n        for i1 in range(resolution):\n            i1 /= resolution\n            x[i] = i1\n            score = mf(x)\n            if score &gt; best_score:\n                best_i1 = i1\n                best_score = score\n        x[i] = best_i1\n        if verbose:\n            print(i, best_i1, best_score)\n\n    return x\noptimise_f1_thresholds(target, predicted)\n```",
          "votes": 2
        },
        {
          "id": 451784,
          "postDate": "2019-01-07T16:35:07.100Z",
          "content": "<p>I've been using this code, supplied by Brian:\n<code>\nfrom sklearn.metrics import f1_score\nthresholds = np.linspace(0, 1, 1000)\nscore = 0.0\ntest_threshold=0.5*np.ones(28)\nbest_threshold=np.zeros(28)\nbest_val = np.zeros(28)\nfor i in range(28):\n    for threshold in thresholds:\n        test_threshold[i] = threshold\n        max_val = np.max(preds_y)\n        val_predict = (preds_y &gt; test_threshold)\n        score = f1_score(valid_y &gt; 0.5, val_predict, average='macro')\n        if score &gt; best_val[i]:\n            best_threshold[i] = threshold\n            best_val[i] = score\n    print(\"Threshold[%d] %0.6f, F1: %0.6f\" % (i,best_threshold[i],best_val[i]))\n    test_threshold[i] = best_threshold[i]\nprint(\"Best threshold: \")\nprint(best_threshold)\nprint(\"Best f1:\")\nprint(best_val)\n</code>\nThe f1 it displays is always an overestimate for me, but the thresholds it gives usually improve the score compared to a single threshold, though not always.</p>",
          "rawMarkdown": "I've been using this code, supplied by Brian:\n```\nfrom sklearn.metrics import f1_score\nthresholds = np.linspace(0, 1, 1000)\nscore = 0.0\ntest_threshold=0.5*np.ones(28)\nbest_threshold=np.zeros(28)\nbest_val = np.zeros(28)\nfor i in range(28):\n    for threshold in thresholds:\n        test_threshold[i] = threshold\n        max_val = np.max(preds_y)\n        val_predict = (preds_y &gt; test_threshold)\n        score = f1_score(valid_y &gt; 0.5, val_predict, average='macro')\n        if score &gt; best_val[i]:\n            best_threshold[i] = threshold\n            best_val[i] = score\n    print(\"Threshold[%d] %0.6f, F1: %0.6f\" % (i,best_threshold[i],best_val[i]))\n    test_threshold[i] = best_threshold[i]\nprint(\"Best threshold: \")\nprint(best_threshold)\nprint(\"Best f1:\")\nprint(best_val)\n```\nThe f1 it displays is always an overestimate for me, but the thresholds it gives usually improve the score compared to a single threshold, though not always.",
          "votes": 1
        },
        {
          "id": 451978,
          "postDate": "2019-01-08T01:54:57.350Z",
          "content": "<p>thanks very much <a href=\"/interneuron\">@interneuron</a> @Peterzhang, I just want to ask one question:\nwhich data set should we use to search the threshold, all training set or just validation set?</p>",
          "rawMarkdown": "thanks very much @interneuron @Peterzhang, I just want to ask one question:\nwhich data set should we use to search the threshold, all training set or just validation set?",
          "votes": 1
        },
        {
          "id": 452015,
          "postDate": "2019-01-08T03:16:43.583Z",
          "content": "<p>Using the training set usually gives a slightly higher score than valid, but using the external data my training set is ~80k so it takes me an hour or so to get those thresholds. Valid takes about 10 min.</p>",
          "rawMarkdown": "Using the training set usually gives a slightly higher score than valid, but using the external data my training set is ~80k so it takes me an hour or so to get those thresholds. Valid takes about 10 min."
        }
      ]
    },
    {
      "id": 451958,
      "postDate": "2019-01-08T00:56:30.573Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 452021,
          "postDate": "2019-01-08T03:27:50.347Z",
          "content": "<p>only rgb</p>",
          "rawMarkdown": "only rgb"
        },
        {
          "id": 452886,
          "postDate": "2019-01-09T10:07:45.467Z",
          "content": "<p>I don't understand.\nDo you just skip yellow channel or you convert rgby to rgb in some way?\nSeems strange to skip whole channel. Do you have a hypothesis why that helps ?</p>",
          "rawMarkdown": "I don't understand.\nDo you just skip yellow channel or you convert rgby to rgb in some way?\nSeems strange to skip whole channel. Do you have a hypothesis why that helps ?"
        },
        {
          "id": 452888,
          "postDate": "2019-01-09T10:10:55.363Z",
          "content": "<p>drop yellow channel</p>",
          "rawMarkdown": "drop yellow channel"
        }
      ]
    },
    {
      "id": 453431,
      "postDate": "2019-01-10T07:28:19.970Z",
      "content": "<p>Thanks for help!</p>",
      "rawMarkdown": "Thanks for help!"
    }
  ],
  "comments": [
    {
      "id": 453183,
      "author_name": "Miroslav Valan",
      "author_url": "",
      "post_date": "2019-01-09T20:22:39.543000",
      "content": "<p>thx for sharing, I wonder how your local F1 and loss (val and train) look? If it's too much to ask, no worries.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 452110,
      "author_name": "Zhen Cao",
      "author_url": "",
      "post_date": "2019-01-08T07:48:30.547000",
      "content": "<p>hi，can u add my qq 1124963678， i need some help.</p>",
      "votes": -6,
      "replies": []
    },
    {
      "id": 452762,
      "author_name": "Newt Scamander",
      "author_url": "",
      "post_date": "2019-01-09T06:03:30.267000",
      "content": "<p><a href=\"/pascal1129\">@pascal1129</a>,Thx for sharing,and could you give some explanation about item 6 \"upsample + downsample + log-dammped weights\",I don't understand very well.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 452785,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-09T06:44:24.823000",
          "content": "<p>add the samples of rare class, reduce the samples of class which is too much, give loss function a weight in order to balance differenet classes</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 451686,
      "author_name": "tkuanlun350",
      "author_url": "",
      "post_date": "2019-01-07T13:44:28.827000",
      "content": "<p>Thanks for sharing !\nCan you list the two papers you mentioned ? Probably won't have time to try but I am curious. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 451694,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-07T14:04:09.937000",
          "content": "<p>After this competition, i will share my solution.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 451673,
      "author_name": "David Wagner",
      "author_url": "",
      "post_date": "2019-01-07T13:22:00.887000",
      "content": "<p>Hey <a href=\"/pascal1129\">@pascal1129</a> </p>\n\n<p>How do you handle BatchNorm with gradient accumulation? I thought they were incompatible and you would have to use GroupNorm or get rid of BatchNorm to accumulate gradients. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 451678,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-07T13:28:42.367000",
          "content": "<p>Can you proveide more infomation? I haven't thought about the problem before. I just use BN in the model structure, and used accumulating gradeint in the train stage. I don't know what is needed to be paid attention to.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451855,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-01-07T19:20:23.417000",
          "content": "<p>As far as I know, batch norm statistics get updated on each forward pass, so no problem if you don't do .backward() every time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 452367,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-08T16:16:56.543000",
          "content": "<p>Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452816,
          "author_name": "David Wagner",
          "author_url": "",
          "post_date": "2019-01-09T07:38:03.310000",
          "content": "<p>Took me a minute to get back to this, but thanks for the responses. I retried using gradient accumulation and it is working as expected. Thanks! I had another issue going on at the first time I tried to accumulate gradients.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452914,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-01-09T11:05:51.503000",
          "content": "<p>would it be ok like this?\n`</p>\n\n<pre><code>loss += loss / accumulation_steps\n\nif (i+1) % accumulation_steps == 0:  \n\n     loss.backward()\n\n     optimizer.step()             \n\n     model.zero_grad()`\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 453025,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-09T14:31:37.580000",
          "content": "<p>```\nfor i,(images,target) in enumerate(train_loader):</p>\n\n<pre><code># 1. input output\nimages = images.cuda(non_blocking=True)\ntarget = torch.from_numpy(np.array(target)).float().cuda(non_blocking=True)\noutputs = model(images)\nloss = criterion(outputs,target)\n\n# 2.1 loss regularization\nloss = loss/accumulation_steps   \n\n# 2.2 back propagation\nloss.backward()\n\n# 3. update parameters of net\nif(i%accumulation_steps)==0:\n    optimizer.step()        # update parameters of net\n    optimizer.zero_grad()   # reset gradient\n</code></pre>\n\n<p>```</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 454670,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-01-12T00:33:16.767000",
          "content": "<p>Thx for sharing the code. Is there any reference paper of this, I’m new to DL :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 451638,
      "author_name": "Cris Lee",
      "author_url": "",
      "post_date": "2019-01-07T12:09:32.780000",
      "content": "<p>Thx  for sharing！  May I ask what‘s your TTA strategy in detail？ Cause  TTA  lower my LB score = = </p>",
      "votes": 0,
      "replies": [
        {
          "id": 451645,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-07T12:31:33.783000",
          "content": "<p>Indeed, several days aog, TTA lower my score too. After improving my model, TTA works.</p>\n\n<p>I only use HorizontalFlip and VerticalFlip in TTA.</p>\n\n<p>For more details about TTA transform, you can see my github, Focus on the input[6]:\n<a href=\"https://github.com/pascal1129/CV_Notes/blob/master/codes/torchvision.transforms.ipynb\">https://github.com/pascal1129/CV_Notes/blob/master/codes/torchvision.transforms.ipynb</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 451524,
      "author_name": "excllent123",
      "author_url": "",
      "post_date": "2019-01-07T08:24:05.340000",
      "content": "<p>Can U  share  the score  without extranal data？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 451529,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-07T08:29:47.537000",
          "content": "<p>I did not record the relevant results.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 451521,
      "author_name": "Peterzhang",
      "author_url": "",
      "post_date": "2019-01-07T08:21:14.667000",
      "content": "<p>Thx for  sharing, Could you share your lr schedule strategy? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 451528,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-07T08:27:50.957000",
          "content": "<p>I used compose of <a href=\"https://pytorch.org/docs/stable/optim.html#how-to-adjust-learning-rate\">https://pytorch.org/docs/stable/optim.html#how-to-adjust-learning-rate</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 451496,
      "author_name": "Femi",
      "author_url": "",
      "post_date": "2019-01-07T07:54:49.980000",
      "content": "<p>Thanks for sharing! I wanted to ask one question:\n How to search threshold for each class?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 451512,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-07T08:11:20.653000",
          "content": "<p>For example, try different thresholds on one class form 0.01 to 0.99 by step 0.01 , and use the threshold which has the best f1-score.</p>\n\n<p><img src=\"https://s2.ax1x.com/2019/01/07/FbIauT.png\" alt=\"FbIauT.png\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451530,
          "author_name": "Ildoo Kim",
          "author_url": "",
          "post_date": "2019-01-07T08:29:49.253000",
          "content": "<p>Trying from 0.01 to 0.99 by step 0.01 have a combination of 100^28. This is an intractable.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451534,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-07T08:34:44.717000",
          "content": "<p>If you used suitable method, I think is 100 x 28, not 100^28.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451556,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-07T09:05:01.740000",
          "content": "<p>Thanks for your help!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451599,
          "author_name": "xuan",
          "author_url": "",
          "post_date": "2019-01-07T10:30:58.383000",
          "content": "<p>search individually for each class and take the best threshold is fine imo</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 451610,
          "author_name": "Peterzhang",
          "author_url": "",
          "post_date": "2019-01-07T10:54:19.293000",
          "content": "<p>I just tried  threshold searching, but it lowered my lb... I don't know why</p>\n\n<p>I just share the code here</p>\n\n<p>```\nimport numpy as np\nfrom sklearn.metrics import f1_score\nnum_classes=28\ndef optimise_f1_thresholds(y, p, verbose=True, resolution=100):\n    def mf(x):\n        p1 = np.zeros_like(p)\n        for i in range(num_classes):\n            p1[:, i] = (p[:, i] &gt; x[i]).astype(np.int)\n        score = f1_score(y, p1, average='samples')\n        return score</p>\n\n<pre><code>x = [0.2]*num_classes\nfor i in range(num_classes):\n    best_i1 = 0\n    best_score = 0\n    for i1 in range(resolution):\n        i1 /= resolution\n        x[i] = i1\n        score = mf(x)\n        if score &gt; best_score:\n            best_i1 = i1\n            best_score = score\n    x[i] = best_i1\n    if verbose:\n        print(i, best_i1, best_score)\n\nreturn x\n</code></pre>\n\n<p>optimise_f1_thresholds(target, predicted)\n```</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 451784,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-01-07T16:35:07.100000",
          "content": "<p>I've been using this code, supplied by Brian:\n<code>\nfrom sklearn.metrics import f1_score\nthresholds = np.linspace(0, 1, 1000)\nscore = 0.0\ntest_threshold=0.5*np.ones(28)\nbest_threshold=np.zeros(28)\nbest_val = np.zeros(28)\nfor i in range(28):\n    for threshold in thresholds:\n        test_threshold[i] = threshold\n        max_val = np.max(preds_y)\n        val_predict = (preds_y &gt; test_threshold)\n        score = f1_score(valid_y &gt; 0.5, val_predict, average='macro')\n        if score &gt; best_val[i]:\n            best_threshold[i] = threshold\n            best_val[i] = score\n    print(\"Threshold[%d] %0.6f, F1: %0.6f\" % (i,best_threshold[i],best_val[i]))\n    test_threshold[i] = best_threshold[i]\nprint(\"Best threshold: \")\nprint(best_threshold)\nprint(\"Best f1:\")\nprint(best_val)\n</code>\nThe f1 it displays is always an overestimate for me, but the thresholds it gives usually improve the score compared to a single threshold, though not always.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 451978,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-08T01:54:57.350000",
          "content": "<p>thanks very much <a href=\"/interneuron\">@interneuron</a> @Peterzhang, I just want to ask one question:\nwhich data set should we use to search the threshold, all training set or just validation set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 452015,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-01-08T03:16:43.583000",
          "content": "<p>Using the training set usually gives a slightly higher score than valid, but using the external data my training set is ~80k so it takes me an hour or so to get those thresholds. Valid takes about 10 min.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 451958,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-08T00:56:30.573000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 452021,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-08T03:27:50.347000",
          "content": "<p>only rgb</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452886,
          "author_name": "Alfonsas Juršėnas",
          "author_url": "",
          "post_date": "2019-01-09T10:07:45.467000",
          "content": "<p>I don't understand.\nDo you just skip yellow channel or you convert rgby to rgb in some way?\nSeems strange to skip whole channel. Do you have a hypothesis why that helps ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 452888,
          "author_name": "pascal1129",
          "author_url": "",
          "post_date": "2019-01-09T10:10:55.363000",
          "content": "<p>drop yellow channel</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 453431,
      "author_name": "Blue Yun",
      "author_url": "",
      "post_date": "2019-01-10T07:28:19.970000",
      "content": "<p>Thanks for help!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "451445": "0、My situation: single GTX1080ti (sometimes 2), PyTorch 1.0;\n\n1、My best single model now can get 0.59+, and i think reaching 0.60 is possible;\n\n2、I use SGD, and i think lr schedule is important. I trained each model for about 30-35 epochs. My base_lr is 0.20, maybe is too big; (In addition, my batchsize is 32*8 with accumulating gradenit tech);\n\n3、I am sure that searching threshold  for each class is quite uesful;\n\n4、rgb is better than rgby, based on my experimental results ;\n\n5、I wasetd too much time tring advanced networks, sucha as SeNet, RseNeXt. Now I only use Res34 with 512 size, i think res18 and res34 is quite enough;\n\n6、I use upsample + downsample + log-dammped weights;\n\n7、TTA is useful;\n\n8、I think external data is not essential to getting 0.60;\n\n9、I reproduced 2 papers, which bring me 2 percent increase;\n\n10、My model has used too much regularization tricks to prevent overfitting, which restricted my public leadboad score, so I gave up using dropout, resize and crop. It may be a little late for me to realize the situation.\n\n11、Although dropout can be deleted, BN1d is essential to the fc layer in my network. I have tried to delete BN1d in the fc layer, and then my loss boomed, so don't delete your bn layer and have a try in other ways;\n\nI spent quite lots of time to find parameters that work for my network, which is painful. The deadline is comming, and i have other things to do, i may have no eonuth time to finish the competition, so i share my thougts, which i hope can bring you improvements.\n",
    "453183": "thx for sharing, I wonder how your local F1 and loss (val and train) look? If it's too much to ask, no worries.",
    "452110": "hi，can u add my qq 1124963678， i need some help.",
    "452762": "@pascal1129,Thx for sharing,and could you give some explanation about item 6 \"upsample + downsample + log-dammped weights\",I don't understand very well.",
    "451686": "Thanks for sharing !\nCan you list the two papers you mentioned ? Probably won't have time to try but I am curious. ",
    "451673": "Hey @pascal1129 \n\nHow do you handle BatchNorm with gradient accumulation? I thought they were incompatible and you would have to use GroupNorm or get rid of BatchNorm to accumulate gradients. ",
    "451638": "Thx  for sharing！  May I ask what‘s your TTA strategy in detail？ Cause  TTA  lower my LB score = = ",
    "451524": "Can U  share  the score  without extranal data？",
    "451521": "Thx for  sharing, Could you share your lr schedule strategy? ",
    "451496": "Thanks for sharing! I wanted to ask one question:\n How to search threshold for each class?",
    "451958": "",
    "453431": "Thanks for help!"
  }
}