{
  "id": 101528,
  "title": "ResNet34 not converging",
  "url": "/competitions/recursion-cellular-image-classification/discussion/101528",
  "author_name": "",
  "post_date": "2019-07-26T14:21:00.524163300Z",
  "votes": 2,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi guys, as suggested by many top scorers, we should train a network and do individual predictions based on cell-type. I followed the advice, trained a ResNet34 on both sites without any validation and managed to achieve an training accuracy of 96%.</p>\n\n<p>I then saved the model and loaded the model to predict cell-type <code>HEPG2</code>. I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case, it's \n<code>\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n</code></p>\n\n<p>Below is the loss and accuracy I have been achieving. I notice it doesn't converge and fluctuates a lot even after 50 epochs. Here is a small snapshot. </p>\n\n<p><code>\nEpoch: 1    Train loss: 7.768    Valid loss: 11.482 Valid accuracy: 0.018\nEpoch: 2    Train loss: 5.849    Valid loss: 9.355  Valid accuracy: 0.021\nEpoch: 3    Train loss: 5.797    Valid loss: 8.638  Valid accuracy: 0.019\nEpoch: 4    Train loss: 5.882    Valid loss: 6.722  Valid accuracy: 0.019\nEpoch: 5    Train loss: 5.990    Valid loss: 7.152  Valid accuracy: 0.014\nEpoch: 6    Train loss: 6.076    Valid loss: 6.602  Valid accuracy: 0.019\nEpoch: 7    Train loss: 6.145    Valid loss: 6.531  Valid accuracy: 0.010\nEpoch: 8    Train loss: 6.178    Valid loss: 10.853 Valid accuracy: 0.001\nEpoch: 9    Train loss: 6.202    Valid loss: 7.589  Valid accuracy: 0.011\nEpoch: 10   Train loss: 6.204    Valid loss: 7.592  Valid accuracy: 0.007\nEpoch: 11   Train loss: 6.194    Valid loss: 6.135  Valid accuracy: 0.014\nEpoch: 12   Train loss: 6.192    Valid loss: 6.804  Valid accuracy: 0.007\nEpoch: 13   Train loss: 6.181    Valid loss: 6.213  Valid accuracy: 0.014\n</code></p>\n\n<p>Any advice what might be wrong? Thank you.</p>",
  "messages": [
    {
      "id": "584823",
      "postDate": "07/26/2019 14:21:00",
      "content": "<p>Hi guys, as suggested by many top scorers, we should train a network and do individual predictions based on cell-type. I followed the advice, trained a ResNet34 on both sites without any validation and managed to achieve an training accuracy of 96%.</p>\n\n<p>I then saved the model and loaded the model to predict cell-type <code>HEPG2</code>. I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case, it's \n<code>\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n</code></p>\n\n<p>Below is the loss and accuracy I have been achieving. I notice it doesn't converge and fluctuates a lot even after 50 epochs. Here is a small snapshot. </p>\n\n<p><code>\nEpoch: 1    Train loss: 7.768    Valid loss: 11.482 Valid accuracy: 0.018\nEpoch: 2    Train loss: 5.849    Valid loss: 9.355  Valid accuracy: 0.021\nEpoch: 3    Train loss: 5.797    Valid loss: 8.638  Valid accuracy: 0.019\nEpoch: 4    Train loss: 5.882    Valid loss: 6.722  Valid accuracy: 0.019\nEpoch: 5    Train loss: 5.990    Valid loss: 7.152  Valid accuracy: 0.014\nEpoch: 6    Train loss: 6.076    Valid loss: 6.602  Valid accuracy: 0.019\nEpoch: 7    Train loss: 6.145    Valid loss: 6.531  Valid accuracy: 0.010\nEpoch: 8    Train loss: 6.178    Valid loss: 10.853 Valid accuracy: 0.001\nEpoch: 9    Train loss: 6.202    Valid loss: 7.589  Valid accuracy: 0.011\nEpoch: 10   Train loss: 6.204    Valid loss: 7.592  Valid accuracy: 0.007\nEpoch: 11   Train loss: 6.194    Valid loss: 6.135  Valid accuracy: 0.014\nEpoch: 12   Train loss: 6.192    Valid loss: 6.804  Valid accuracy: 0.007\nEpoch: 13   Train loss: 6.181    Valid loss: 6.213  Valid accuracy: 0.014\n</code></p>\n\n<p>Any advice what might be wrong? Thank you.</p>",
      "rawMarkdown": "Hi guys, as suggested by many top scorers, we should train a network and do individual predictions based on cell-type. I followed the advice, trained a ResNet34 on both sites without any validation and managed to achieve an training accuracy of 96%.\n\nI then saved the model and loaded the model to predict cell-type `HEPG2`. I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case, it's \n```\ntrain len: 5536 \t experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214 \t experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429 \t experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n```\n\nBelow is the loss and accuracy I have been achieving. I notice it doesn't converge and fluctuates a lot even after 50 epochs. Here is a small snapshot. \n\n```\nEpoch: 1\tTrain loss: 7.768\t Valid loss: 11.482\tValid accuracy: 0.018\nEpoch: 2\tTrain loss: 5.849\t Valid loss: 9.355\tValid accuracy: 0.021\nEpoch: 3\tTrain loss: 5.797\t Valid loss: 8.638\tValid accuracy: 0.019\nEpoch: 4\tTrain loss: 5.882\t Valid loss: 6.722\tValid accuracy: 0.019\nEpoch: 5\tTrain loss: 5.990\t Valid loss: 7.152\tValid accuracy: 0.014\nEpoch: 6\tTrain loss: 6.076\t Valid loss: 6.602\tValid accuracy: 0.019\nEpoch: 7\tTrain loss: 6.145\t Valid loss: 6.531\tValid accuracy: 0.010\nEpoch: 8\tTrain loss: 6.178\t Valid loss: 10.853\tValid accuracy: 0.001\nEpoch: 9\tTrain loss: 6.202\t Valid loss: 7.589\tValid accuracy: 0.011\nEpoch: 10\tTrain loss: 6.204\t Valid loss: 7.592\tValid accuracy: 0.007\nEpoch: 11\tTrain loss: 6.194\t Valid loss: 6.135\tValid accuracy: 0.014\nEpoch: 12\tTrain loss: 6.192\t Valid loss: 6.804\tValid accuracy: 0.007\nEpoch: 13\tTrain loss: 6.181\t Valid loss: 6.213\tValid accuracy: 0.014\n```\n\nAny advice what might be wrong? Thank you.",
      "votes": null
    },
    {
      "id": "585525",
      "postDate": "07/27/2019 15:29:12",
      "content": "<p>Are you using Adam optimizer? </p>",
      "rawMarkdown": "Are you using Adam optimizer?",
      "votes": null
    },
    {
      "id": "585532",
      "postDate": "07/27/2019 15:46:06",
      "content": "<p>hi <a href=\"/joven1997\">@joven1997</a>, yes i am. and also Exponential lr scheduler, what would you recommend?</p>",
      "rawMarkdown": "hi @joven1997, yes i am. and also Exponential lr scheduler, what would you recommend?",
      "votes": null
    },
    {
      "id": "590541",
      "postDate": "08/02/2019 10:06:41",
      "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> Sorry for late reply, have you solved your error?</p>",
      "rawMarkdown": "wjshenggggg Sorry for late reply, have you solved your error?",
      "votes": null
    },
    {
      "id": "590728",
      "postDate": "08/02/2019 14:15:55",
      "content": "<p>hi <a href=\"/joven1997\">@joven1997</a> no worries. i have not, actually. </p>",
      "rawMarkdown": "hi @joven1997 no worries. i have not, actually.",
      "votes": null
    },
    {
      "id": "590790",
      "postDate": "08/02/2019 15:41:04",
      "content": "<p>Did you normalize your image to mean 0 std 1 before you fed them into your network?</p>",
      "rawMarkdown": "Did you normalize your image to mean 0 std 1 before you fed them into your network?",
      "votes": null
    },
    {
      "id": "590859",
      "postDate": "08/02/2019 18:07:00",
      "content": "<p>hi <a href=\"/lintseju\">@lintseju</a>! no i did not. May i know how to normalize the images for my pretraine resnet model when there are 6 channels? thanks! </p>",
      "rawMarkdown": "hi @lintseju! no i did not. May i know how to normalize the images for my pretraine resnet model when there are 6 channels? thanks!",
      "votes": null
    },
    {
      "id": "590860",
      "postDate": "08/02/2019 18:08:03",
      "content": "<p>i found a recent discussion thread by Lorenzo Fabbri, he took the stats from <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/main.py\">https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/main.py</a>. is that what you are doing too? Thanks!</p>",
      "rawMarkdown": "i found a recent discussion thread by Lorenzo Fabbri, he took the stats from https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/main.py. is that what you are doing too? Thanks!",
      "votes": null
    },
    {
      "id": "591312",
      "postDate": "08/03/2019 13:33:38",
      "content": "<p>I normalized by channel. You can try, it should help.</p>",
      "rawMarkdown": "I normalized by channel. You can try, it should help.",
      "votes": null
    },
    {
      "id": "591371",
      "postDate": "08/03/2019 15:17:55",
      "content": "<p>thanks BlackSwan, it is interesting how the tensors are already between -1 and 1 after I call <code>img = torch.cat([self._load_img_as_tensor(img_path) for img_path in paths])</code>. I wonder would further normalization be necessary then.</p>",
      "rawMarkdown": "thanks BlackSwan, it is interesting how the tensors are already between -1 and 1 after I call `img = torch.cat([self._load_img_as_tensor(img_path) for img_path in paths])`. I wonder would further normalization be necessary then.",
      "votes": null
    },
    {
      "id": "591372",
      "postDate": "08/03/2019 15:23:08",
      "content": "<p>in fact, my <code>torch.max</code> for a random image changed from <code>0.89</code> to <code>13.1</code> after normalizing...</p>",
      "rawMarkdown": "in fact, my `torch.max` for a random image changed from `0.89` to `13.1` after normalizing...",
      "votes": null
    },
    {
      "id": "591374",
      "postDate": "08/03/2019 15:33:41",
      "content": "<p>here is to reproduce the result:\n```\nGLOBAL_PIXEL_STATS = (np.array([6.74696984, 14.74640167, 10.51260864,\n                                10.45369445,  5.49959796, 9.81545561]),\n                       np.array([7.95876312, 12.17305868, 5.86172946,\n                                 7.83451711, 4.701167, 5.43130431]))</p>\n\n<p>path = DATA_DIR</p>\n\n<p>def get_mean_std():\n    SIZE = 512\n    mean_, std_ = GLOBAL_PIXEL_STATS</p>\n\n<pre><code>print(mean_, std_)\n\nMEAN = []\n\nfor i in range(len(mean_)):\n    t = torch.full((SIZE, SIZE), mean_[i])\n    MEAN += t,\n\nMEAN = torch.stack(MEAN)\n\nSTD = []\nfor i in range(len(std_)):\n    t = torch.full((SIZE, SIZE), std_[i])\n    STD += t,\n\nSTD = torch.stack(STD)\n\nreturn MEAN, STD \n</code></pre>\n\n<p>MEAN, STD = get_mean_std()</p>\n\n<p>def pil2tensor(image,dtype):\n    a = np.asarray(image)\n    if a.ndim==2 : a = np.expand_dims(a,2)\n    a = np.transpose(a, (2, 0, 1))\n    return torch.from_numpy(a.astype(dtype, copy=False))</p>\n\n<p>def image_path(dataset, experiment, plate,\n               address, site, channel, base_path=path):</p>\n\n<pre><code>return os.path.join(base_path, dataset, experiment, \"Plate{}\".format(plate),\n                    \"{}_s{}_w{}.png\".format(address, site, channel))\n</code></pre>\n\n<p>def open_6_channel(dataset, experiment, plate, address, site, base_path=path):\n    img = torch.cat([pil2tensor(Image.open(image_path(dataset, experiment, plate, address, site, i)), np.float32) for i \n            in range(1,7)])\n    return img.sub_(MEAN).div_(STD)</p>\n\n<p>img = open_6_channel('train', 'RPE-01', '1', 'B03', '1', DATA_DIR)</p>\n\n<p>torch.max(img)\n<code>``\ncredits to Darragh for this script for</code>open_6_channel`</p>",
      "rawMarkdown": "here is to reproduce the result:\n```\nGLOBAL_PIXEL_STATS = (np.array([6.74696984, 14.74640167, 10.51260864,\n                                10.45369445,  5.49959796, 9.81545561]),\n                       np.array([7.95876312, 12.17305868, 5.86172946,\n                                 7.83451711, 4.701167, 5.43130431]))\n\npath = DATA_DIR\n\ndef get_mean_std():\n    SIZE = 512\n    mean_, std_ = GLOBAL_PIXEL_STATS\n    \n    print(mean_, std_)\n    \n    MEAN = []\n\n    for i in range(len(mean_)):\n        t = torch.full((SIZE, SIZE), mean_[i])\n        MEAN += t,\n\n    MEAN = torch.stack(MEAN)\n\n    STD = []\n    for i in range(len(std_)):\n        t = torch.full((SIZE, SIZE), std_[i])\n        STD += t,\n\n    STD = torch.stack(STD)\n\n    return MEAN, STD \n    \nMEAN, STD = get_mean_std()\n\ndef pil2tensor(image,dtype):\n    a = np.asarray(image)\n    if a.ndim==2 : a = np.expand_dims(a,2)\n    a = np.transpose(a, (2, 0, 1))\n    return torch.from_numpy(a.astype(dtype, copy=False))\n\ndef image_path(dataset, experiment, plate,\n               address, site, channel, base_path=path):\n\n    return os.path.join(base_path, dataset, experiment, \"Plate{}\".format(plate),\n                        \"{}_s{}_w{}.png\".format(address, site, channel))\n\ndef open_6_channel(dataset, experiment, plate, address, site, base_path=path):\n    img = torch.cat([pil2tensor(Image.open(image_path(dataset, experiment, plate, address, site, i)), np.float32) for i \n            in range(1,7)])\n    return img.sub_(MEAN).div_(STD)\n\nimg = open_6_channel('train', 'RPE-01', '1', 'B03', '1', DATA_DIR)\n\ntorch.max(img)\n```\ncredits to Darragh for this script for `open_6_channel`",
      "votes": null
    },
    {
      "id": "592246",
      "postDate": "08/05/2019 04:06:46",
      "content": "<p>hi <a href=\"/lintseju\">@lintseju</a>, I wonder can you please take a look at how i do my normalization. Please let me know if it is wrong. To reiterate, <code>img.sub_(MEAN).div_(STD)</code> returns a value greater than 1 compared to not doing any normalization. thank you! :) </p>",
      "rawMarkdown": "hi @lintseju, I wonder can you please take a look at how i do my normalization. Please let me know if it is wrong. To reiterate, `img.sub_(MEAN).div_(STD)` returns a value greater than 1 compared to not doing any normalization. thank you! :)",
      "votes": null
    },
    {
      "id": "592868",
      "postDate": "08/05/2019 23:23:28",
      "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> it seems you have a nice performance now, so the improvement doesn't relate to this topic? ;)</p>",
      "rawMarkdown": "@wjshenggggg it seems you have a nice performance now, so the improvement doesn't relate to this topic? ;)",
      "votes": null
    },
    {
      "id": "592912",
      "postDate": "08/06/2019 01:07:37",
      "content": "<p>thanks for the kind words, <a href=\"/ratthachat\">@ratthachat</a>! I think i was lucky but still learning how to do things the right way! :) </p>",
      "rawMarkdown": "thanks for the kind words, @ratthachat! I think i was lucky but still learning how to do things the right way! :)",
      "votes": null
    },
    {
      "id": "596155",
      "postDate": "08/10/2019 08:20:38",
      "content": "<p>Hi <a href=\"/wjshenggggg\">@wjshenggggg</a>! May I ask you whether you're still using ResNet__? Before moving to Metric Learning I'd like to get a decent score (~0.6) with <em>simple</em> classification. Unfortunately, my attempts so far have all failed...</p>\n\n<p>According to <a href=\"/yuval6967\">@yuval6967</a> <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/100397\">here</a>, it's possible to get these kind of scores without relying on metric learning.</p>",
      "rawMarkdown": "Hi @wjshenggggg! May I ask you whether you're still using ResNet__? Before moving to Metric Learning I'd like to get a decent score (~0.6) with *simple* classification. Unfortunately, my attempts so far have all failed...\n\nAccording to @yuval6967 [here](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/100397), it's possible to get these kind of scores without relying on metric learning.",
      "votes": null
    },
    {
      "id": "596302",
      "postDate": "08/10/2019 12:58:50",
      "content": "<p>hi Lorenzo. Im using ResNet34 so far with <code>ArcFaceLoss</code>. Unfortunately, I am unable to reproduce my result, which is pretty concerning to me. I ran the same kernel that produced my 0.4 score but the validation score became 0.2. I think i should try the simple classification again.</p>",
      "rawMarkdown": "hi Lorenzo. Im using ResNet34 so far with `ArcFaceLoss`. Unfortunately, I am unable to reproduce my result, which is pretty concerning to me. I ran the same kernel that produced my 0.4 score but the validation score became 0.2. I think i should try the simple classification again.",
      "votes": null
    },
    {
      "id": "596306",
      "postDate": "08/10/2019 13:03:15",
      "content": "<p>That's weird... Were you able to get 0.4 with simple classification, though? Again, thanks.</p>",
      "rawMarkdown": "That's weird... Were you able to get 0.4 with simple classification, though? Again, thanks.",
      "votes": null
    },
    {
      "id": "596393",
      "postDate": "08/10/2019 15:28:08",
      "content": "<p>simple classification does not work for me. even less than 0.2! it's far worse now seeing i want to train with multiple GPUs and getting a correct LR is very tricky! I have been spending more than 2 days trying to get the optimal LR using multiple GPUs. Do you know yuval is using any trick that involves controls? </p>",
      "rawMarkdown": "simple classification does not work for me. even less than 0.2! it's far worse now seeing i want to train with multiple GPUs and getting a correct LR is very tricky! I have been spending more than 2 days trying to get the optimal LR using multiple GPUs. Do you know yuval is using any trick that involves controls?",
      "votes": null
    },
    {
      "id": "596419",
      "postDate": "08/10/2019 16:07:11",
      "content": "<p>Ah, so you always used Metric Learning... I stick to 1 GPU since ResNet18 is quite fast for training (not to mention I thought the deadline for the credits was in August and not in July!).</p>\n\n<p>I do not know, actually. He just posted one kernel so far. I don't think he ever mentioned anything specific, just that you can get 0.6 with DenseNet121, if I remember correctly.</p>",
      "rawMarkdown": "Ah, so you always used Metric Learning... I stick to 1 GPU since ResNet18 is quite fast for training (not to mention I thought the deadline for the credits was in August and not in July!).\n\nI do not know, actually. He just posted one kernel so far. I don't think he ever mentioned anything specific, just that you can get 0.6 with DenseNet121, if I remember correctly.",
      "votes": null
    },
    {
      "id": "596497",
      "postDate": "08/10/2019 18:13:37",
      "content": "<p>ya and i tried using pure densenet121 without any metric learning and i couldnt achieve what he did - that was why i switched. i wonder what am i doing wrongly, learning rate or what...</p>",
      "rawMarkdown": "ya and i tried using pure densenet121 without any metric learning and i couldnt achieve what he did - that was why i switched. i wonder what am i doing wrongly, learning rate or what...",
      "votes": null
    },
    {
      "id": "596499",
      "postDate": "08/10/2019 18:18:01",
      "content": "<p>I've being asking myself the same for a month now :) For sure, having many GPUs allows for more testing. But maybe there is something really simple they they all do (LB &gt; 0.5). Dunno</p>",
      "rawMarkdown": "I've being asking myself the same for a month now :) For sure, having many GPUs allows for more testing. But maybe there is something really simple they they all do (LB &gt; 0.5). Dunno",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 585525,
      "author_name": "joven1997",
      "author_url": "",
      "post_date": "07/27/2019 15:29:12",
      "content": "<p>Are you using Adam optimizer? </p>",
      "votes": null,
      "replies": [
        {
          "id": 585532,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/27/2019 15:46:06",
          "content": "<p>hi <a href=\"/joven1997\">@joven1997</a>, yes i am. and also Exponential lr scheduler, what would you recommend?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 590541,
          "author_name": "joven1997",
          "author_url": "",
          "post_date": "08/02/2019 10:06:41",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> Sorry for late reply, have you solved your error?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 590728,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/02/2019 14:15:55",
          "content": "<p>hi <a href=\"/joven1997\">@joven1997</a> no worries. i have not, actually. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596155,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/10/2019 08:20:38",
          "content": "<p>Hi <a href=\"/wjshenggggg\">@wjshenggggg</a>! May I ask you whether you're still using ResNet__? Before moving to Metric Learning I'd like to get a decent score (~0.6) with <em>simple</em> classification. Unfortunately, my attempts so far have all failed...</p>\n\n<p>According to <a href=\"/yuval6967\">@yuval6967</a> <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/100397\">here</a>, it's possible to get these kind of scores without relying on metric learning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596302,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/10/2019 12:58:50",
          "content": "<p>hi Lorenzo. Im using ResNet34 so far with <code>ArcFaceLoss</code>. Unfortunately, I am unable to reproduce my result, which is pretty concerning to me. I ran the same kernel that produced my 0.4 score but the validation score became 0.2. I think i should try the simple classification again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596306,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/10/2019 13:03:15",
          "content": "<p>That's weird... Were you able to get 0.4 with simple classification, though? Again, thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596393,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/10/2019 15:28:08",
          "content": "<p>simple classification does not work for me. even less than 0.2! it's far worse now seeing i want to train with multiple GPUs and getting a correct LR is very tricky! I have been spending more than 2 days trying to get the optimal LR using multiple GPUs. Do you know yuval is using any trick that involves controls? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596419,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/10/2019 16:07:11",
          "content": "<p>Ah, so you always used Metric Learning... I stick to 1 GPU since ResNet18 is quite fast for training (not to mention I thought the deadline for the credits was in August and not in July!).</p>\n\n<p>I do not know, actually. He just posted one kernel so far. I don't think he ever mentioned anything specific, just that you can get 0.6 with DenseNet121, if I remember correctly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596497,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/10/2019 18:13:37",
          "content": "<p>ya and i tried using pure densenet121 without any metric learning and i couldnt achieve what he did - that was why i switched. i wonder what am i doing wrongly, learning rate or what...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596499,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/10/2019 18:18:01",
          "content": "<p>I've being asking myself the same for a month now :) For sure, having many GPUs allows for more testing. But maybe there is something really simple they they all do (LB &gt; 0.5). Dunno</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 590790,
      "author_name": "lintseju",
      "author_url": "",
      "post_date": "08/02/2019 15:41:04",
      "content": "<p>Did you normalize your image to mean 0 std 1 before you fed them into your network?</p>",
      "votes": null,
      "replies": [
        {
          "id": 590859,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/02/2019 18:07:00",
          "content": "<p>hi <a href=\"/lintseju\">@lintseju</a>! no i did not. May i know how to normalize the images for my pretraine resnet model when there are 6 channels? thanks! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 590860,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/02/2019 18:08:03",
          "content": "<p>i found a recent discussion thread by Lorenzo Fabbri, he took the stats from <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/main.py\">https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/main.py</a>. is that what you are doing too? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591312,
          "author_name": "lintseju",
          "author_url": "",
          "post_date": "08/03/2019 13:33:38",
          "content": "<p>I normalized by channel. You can try, it should help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591371,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/03/2019 15:17:55",
          "content": "<p>thanks BlackSwan, it is interesting how the tensors are already between -1 and 1 after I call <code>img = torch.cat([self._load_img_as_tensor(img_path) for img_path in paths])</code>. I wonder would further normalization be necessary then.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591372,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/03/2019 15:23:08",
          "content": "<p>in fact, my <code>torch.max</code> for a random image changed from <code>0.89</code> to <code>13.1</code> after normalizing...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591374,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/03/2019 15:33:41",
          "content": "<p>here is to reproduce the result:\n```\nGLOBAL_PIXEL_STATS = (np.array([6.74696984, 14.74640167, 10.51260864,\n                                10.45369445,  5.49959796, 9.81545561]),\n                       np.array([7.95876312, 12.17305868, 5.86172946,\n                                 7.83451711, 4.701167, 5.43130431]))</p>\n\n<p>path = DATA_DIR</p>\n\n<p>def get_mean_std():\n    SIZE = 512\n    mean_, std_ = GLOBAL_PIXEL_STATS</p>\n\n<pre><code>print(mean_, std_)\n\nMEAN = []\n\nfor i in range(len(mean_)):\n    t = torch.full((SIZE, SIZE), mean_[i])\n    MEAN += t,\n\nMEAN = torch.stack(MEAN)\n\nSTD = []\nfor i in range(len(std_)):\n    t = torch.full((SIZE, SIZE), std_[i])\n    STD += t,\n\nSTD = torch.stack(STD)\n\nreturn MEAN, STD \n</code></pre>\n\n<p>MEAN, STD = get_mean_std()</p>\n\n<p>def pil2tensor(image,dtype):\n    a = np.asarray(image)\n    if a.ndim==2 : a = np.expand_dims(a,2)\n    a = np.transpose(a, (2, 0, 1))\n    return torch.from_numpy(a.astype(dtype, copy=False))</p>\n\n<p>def image_path(dataset, experiment, plate,\n               address, site, channel, base_path=path):</p>\n\n<pre><code>return os.path.join(base_path, dataset, experiment, \"Plate{}\".format(plate),\n                    \"{}_s{}_w{}.png\".format(address, site, channel))\n</code></pre>\n\n<p>def open_6_channel(dataset, experiment, plate, address, site, base_path=path):\n    img = torch.cat([pil2tensor(Image.open(image_path(dataset, experiment, plate, address, site, i)), np.float32) for i \n            in range(1,7)])\n    return img.sub_(MEAN).div_(STD)</p>\n\n<p>img = open_6_channel('train', 'RPE-01', '1', 'B03', '1', DATA_DIR)</p>\n\n<p>torch.max(img)\n<code>``\ncredits to Darragh for this script for</code>open_6_channel`</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592246,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/05/2019 04:06:46",
          "content": "<p>hi <a href=\"/lintseju\">@lintseju</a>, I wonder can you please take a look at how i do my normalization. Please let me know if it is wrong. To reiterate, <code>img.sub_(MEAN).div_(STD)</code> returns a value greater than 1 compared to not doing any normalization. thank you! :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592868,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "08/05/2019 23:23:28",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> it seems you have a nice performance now, so the improvement doesn't relate to this topic? ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592912,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/06/2019 01:07:37",
          "content": "<p>thanks for the kind words, <a href=\"/ratthachat\">@ratthachat</a>! I think i was lucky but still learning how to do things the right way! :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "584823": "Hi guys, as suggested by many top scorers, we should train a network and do individual predictions based on cell-type. I followed the advice, trained a ResNet34 on both sites without any validation and managed to achieve an training accuracy of 96%.\n\nI then saved the model and loaded the model to predict cell-type `HEPG2`. I am doing my validation based on the bottom N experiments by cell-type where N = the number of experiments in test by cell-typ. For example, in this case, it's \n```\ntrain len: 5536 \t experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214 \t experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429 \t experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n```\n\nBelow is the loss and accuracy I have been achieving. I notice it doesn't converge and fluctuates a lot even after 50 epochs. Here is a small snapshot. \n\n```\nEpoch: 1\tTrain loss: 7.768\t Valid loss: 11.482\tValid accuracy: 0.018\nEpoch: 2\tTrain loss: 5.849\t Valid loss: 9.355\tValid accuracy: 0.021\nEpoch: 3\tTrain loss: 5.797\t Valid loss: 8.638\tValid accuracy: 0.019\nEpoch: 4\tTrain loss: 5.882\t Valid loss: 6.722\tValid accuracy: 0.019\nEpoch: 5\tTrain loss: 5.990\t Valid loss: 7.152\tValid accuracy: 0.014\nEpoch: 6\tTrain loss: 6.076\t Valid loss: 6.602\tValid accuracy: 0.019\nEpoch: 7\tTrain loss: 6.145\t Valid loss: 6.531\tValid accuracy: 0.010\nEpoch: 8\tTrain loss: 6.178\t Valid loss: 10.853\tValid accuracy: 0.001\nEpoch: 9\tTrain loss: 6.202\t Valid loss: 7.589\tValid accuracy: 0.011\nEpoch: 10\tTrain loss: 6.204\t Valid loss: 7.592\tValid accuracy: 0.007\nEpoch: 11\tTrain loss: 6.194\t Valid loss: 6.135\tValid accuracy: 0.014\nEpoch: 12\tTrain loss: 6.192\t Valid loss: 6.804\tValid accuracy: 0.007\nEpoch: 13\tTrain loss: 6.181\t Valid loss: 6.213\tValid accuracy: 0.014\n```\n\nAny advice what might be wrong? Thank you.",
    "585525": "Are you using Adam optimizer?",
    "585532": "hi @joven1997, yes i am. and also Exponential lr scheduler, what would you recommend?",
    "590541": "wjshenggggg Sorry for late reply, have you solved your error?",
    "590728": "hi @joven1997 no worries. i have not, actually.",
    "590790": "Did you normalize your image to mean 0 std 1 before you fed them into your network?",
    "590859": "hi @lintseju! no i did not. May i know how to normalize the images for my pretraine resnet model when there are 6 channels? thanks!",
    "590860": "i found a recent discussion thread by Lorenzo Fabbri, he took the stats from https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/main.py. is that what you are doing too? Thanks!",
    "591312": "I normalized by channel. You can try, it should help.",
    "591371": "thanks BlackSwan, it is interesting how the tensors are already between -1 and 1 after I call `img = torch.cat([self._load_img_as_tensor(img_path) for img_path in paths])`. I wonder would further normalization be necessary then.",
    "591372": "in fact, my `torch.max` for a random image changed from `0.89` to `13.1` after normalizing...",
    "591374": "here is to reproduce the result:\n```\nGLOBAL_PIXEL_STATS = (np.array([6.74696984, 14.74640167, 10.51260864,\n                                10.45369445,  5.49959796, 9.81545561]),\n                       np.array([7.95876312, 12.17305868, 5.86172946,\n                                 7.83451711, 4.701167, 5.43130431]))\n\npath = DATA_DIR\n\ndef get_mean_std():\n    SIZE = 512\n    mean_, std_ = GLOBAL_PIXEL_STATS\n    \n    print(mean_, std_)\n    \n    MEAN = []\n\n    for i in range(len(mean_)):\n        t = torch.full((SIZE, SIZE), mean_[i])\n        MEAN += t,\n\n    MEAN = torch.stack(MEAN)\n\n    STD = []\n    for i in range(len(std_)):\n        t = torch.full((SIZE, SIZE), std_[i])\n        STD += t,\n\n    STD = torch.stack(STD)\n\n    return MEAN, STD \n    \nMEAN, STD = get_mean_std()\n\ndef pil2tensor(image,dtype):\n    a = np.asarray(image)\n    if a.ndim==2 : a = np.expand_dims(a,2)\n    a = np.transpose(a, (2, 0, 1))\n    return torch.from_numpy(a.astype(dtype, copy=False))\n\ndef image_path(dataset, experiment, plate,\n               address, site, channel, base_path=path):\n\n    return os.path.join(base_path, dataset, experiment, \"Plate{}\".format(plate),\n                        \"{}_s{}_w{}.png\".format(address, site, channel))\n\ndef open_6_channel(dataset, experiment, plate, address, site, base_path=path):\n    img = torch.cat([pil2tensor(Image.open(image_path(dataset, experiment, plate, address, site, i)), np.float32) for i \n            in range(1,7)])\n    return img.sub_(MEAN).div_(STD)\n\nimg = open_6_channel('train', 'RPE-01', '1', 'B03', '1', DATA_DIR)\n\ntorch.max(img)\n```\ncredits to Darragh for this script for `open_6_channel`",
    "592246": "hi @lintseju, I wonder can you please take a look at how i do my normalization. Please let me know if it is wrong. To reiterate, `img.sub_(MEAN).div_(STD)` returns a value greater than 1 compared to not doing any normalization. thank you! :)",
    "592868": "@wjshenggggg it seems you have a nice performance now, so the improvement doesn't relate to this topic? ;)",
    "592912": "thanks for the kind words, @ratthachat! I think i was lucky but still learning how to do things the right way! :)",
    "596155": "Hi @wjshenggggg! May I ask you whether you're still using ResNet__? Before moving to Metric Learning I'd like to get a decent score (~0.6) with *simple* classification. Unfortunately, my attempts so far have all failed...\n\nAccording to @yuval6967 [here](https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/100397), it's possible to get these kind of scores without relying on metric learning.",
    "596302": "hi Lorenzo. Im using ResNet34 so far with `ArcFaceLoss`. Unfortunately, I am unable to reproduce my result, which is pretty concerning to me. I ran the same kernel that produced my 0.4 score but the validation score became 0.2. I think i should try the simple classification again.",
    "596306": "That's weird... Were you able to get 0.4 with simple classification, though? Again, thanks.",
    "596393": "simple classification does not work for me. even less than 0.2! it's far worse now seeing i want to train with multiple GPUs and getting a correct LR is very tricky! I have been spending more than 2 days trying to get the optimal LR using multiple GPUs. Do you know yuval is using any trick that involves controls?",
    "596419": "Ah, so you always used Metric Learning... I stick to 1 GPU since ResNet18 is quite fast for training (not to mention I thought the deadline for the credits was in August and not in July!).\n\nI do not know, actually. He just posted one kernel so far. I don't think he ever mentioned anything specific, just that you can get 0.6 with DenseNet121, if I remember correctly.",
    "596497": "ya and i tried using pure densenet121 without any metric learning and i couldnt achieve what he did - that was why i switched. i wonder what am i doing wrongly, learning rate or what...",
    "596499": "I've being asking myself the same for a month now :) For sure, having many GPUs allows for more testing. But maybe there is something really simple they they all do (LB &gt; 0.5). Dunno"
  },
  "source": "meta"
}