{
  "id": 133219,
  "title": "Some thoughts",
  "url": "/competitions/bengaliai-cv19/discussion/133219",
  "author_name": "CPMP",
  "post_date": "2020-03-01T12:46:45.011000",
  "votes": 56,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Disclaimer: I am not at all a computer vision expert, never got a gold in an image competition.  Therefore take the following with a large grain of salt.  </p>\n\n<p>Let me first say that this forum is a goldmine for whomever wants to start.  I'm very impressed by how much Qishen Ha has shared in particular.  He could have kept his secret sauce entirely hidden.  Needless to say I'm reading and rereading everything he and other top performers are sharing.  </p>\n\n<p>The following is a little giveback to the community.</p>\n\n<p>The first comment I have is on the use of feather to save image data.  Why not use numpy arrays?  In my test using numpy is 3x faster.  Here is how I saved train data:</p>\n\n<pre><code>k = 0\ntrain_idx = []\ntrain_values = np.zeros((n_image, 137 * 236), dtype = \"uint8\")\nfor i in range(n_file):\n    print(\"reading train file\",i)\n    directory = \"../input/bengaliai-cv19/train_image_data_\"+str(i)+\".parquet\"\n    train_f = pd.read_parquet(directory, engine = \"pyarrow\")\n    train_f.set_index('image_id', inplace=True)\n    train_idx.append(train_f.index.copy())\n    train_values[i * n_file_image : (i + 1) * n_file_image, :] = (255 - train_f.values)\n    del train_f\n    gc.collect()\n\ntrain_values = train_values.reshape((-1, 137, 236))\n\nnp.save('../input/bengaliai-cv19/train.npy', train_values)\n</code></pre>\n\n<p>The second comment is that many public kernels call sklearn metric with the wrong argument order.  For some reason, Pytorch and Scikit-learn have different ordering.  In sklearn ground truth is the first argument...</p>\n\n<p>The third comment is about the use of accuracy as metric while training.  I see it in some public kernels and discussions.  Why not use the competition metric directly?</p>\n\n<p>Here is the code I use.  First, just wrap sklearn metric.</p>\n\n<pre><code>def get_recall(y_true, y_pred):\n    pred_labels = np.argmax(y_pred, axis=1)\n    res = recall_score(y_true, pred_labels, average='macro')\n    return res\n</code></pre>\n\n<p>Then use it this way in your training or validation loop, if <code>all_preds</code> is a numpy array containing the predictions:</p>\n\n<pre><code>    all_preds = np.split(all_preds,\n                         np.cumsum([n_grapheme, n_vowel, n_consonant]), \n                         axis=1\n                        )\n    recall_grapheme = get_recall(y[:, 0], all_preds[0], )\n    recall_vowel = get_recall(y[:, 1], all_preds[1], )\n    recall_consonant = get_recall(y[:, 2], all_preds[2], )\n    recall = np.average([recall_grapheme, recall_vowel, recall_consonant], \n                        weights=[2, 1, 1])\n</code></pre>\n\n<p>Monitoring this makes much more sense than monitoring accuracy or logloss.</p>\n\n<p>Fourth comment is about dealing with one channel input to 3 channels input pretrained models.  I saw the addition of a convolution layer before the model, or hacking the first level of the model.  There is is a simpler way that Chris Deotte gave for Keras: just concatenate copies of the input.  Here is a Pytorch code, just start the forward method of your model with:</p>\n\n<pre><code>def forward(self, x):\n    h = torch.cat([x, x, x], dim=1)\n    h = self.base_model.features(h)\n</code></pre>\n\n<p>Sure, this is probably way too late for most, but maybe it will help some of you.</p>",
  "messages": [
    {
      "id": 760541,
      "postDate": "2020-03-01T12:46:45.010Z",
      "content": "<p>Disclaimer: I am not at all a computer vision expert, never got a gold in an image competition.  Therefore take the following with a large grain of salt.  </p>\n\n<p>Let me first say that this forum is a goldmine for whomever wants to start.  I'm very impressed by how much Qishen Ha has shared in particular.  He could have kept his secret sauce entirely hidden.  Needless to say I'm reading and rereading everything he and other top performers are sharing.  </p>\n\n<p>The following is a little giveback to the community.</p>\n\n<p>The first comment I have is on the use of feather to save image data.  Why not use numpy arrays?  In my test using numpy is 3x faster.  Here is how I saved train data:</p>\n\n<pre><code>k = 0\ntrain_idx = []\ntrain_values = np.zeros((n_image, 137 * 236), dtype = \"uint8\")\nfor i in range(n_file):\n    print(\"reading train file\",i)\n    directory = \"../input/bengaliai-cv19/train_image_data_\"+str(i)+\".parquet\"\n    train_f = pd.read_parquet(directory, engine = \"pyarrow\")\n    train_f.set_index('image_id', inplace=True)\n    train_idx.append(train_f.index.copy())\n    train_values[i * n_file_image : (i + 1) * n_file_image, :] = (255 - train_f.values)\n    del train_f\n    gc.collect()\n\ntrain_values = train_values.reshape((-1, 137, 236))\n\nnp.save('../input/bengaliai-cv19/train.npy', train_values)\n</code></pre>\n\n<p>The second comment is that many public kernels call sklearn metric with the wrong argument order.  For some reason, Pytorch and Scikit-learn have different ordering.  In sklearn ground truth is the first argument...</p>\n\n<p>The third comment is about the use of accuracy as metric while training.  I see it in some public kernels and discussions.  Why not use the competition metric directly?</p>\n\n<p>Here is the code I use.  First, just wrap sklearn metric.</p>\n\n<pre><code>def get_recall(y_true, y_pred):\n    pred_labels = np.argmax(y_pred, axis=1)\n    res = recall_score(y_true, pred_labels, average='macro')\n    return res\n</code></pre>\n\n<p>Then use it this way in your training or validation loop, if <code>all_preds</code> is a numpy array containing the predictions:</p>\n\n<pre><code>    all_preds = np.split(all_preds,\n                         np.cumsum([n_grapheme, n_vowel, n_consonant]), \n                         axis=1\n                        )\n    recall_grapheme = get_recall(y[:, 0], all_preds[0], )\n    recall_vowel = get_recall(y[:, 1], all_preds[1], )\n    recall_consonant = get_recall(y[:, 2], all_preds[2], )\n    recall = np.average([recall_grapheme, recall_vowel, recall_consonant], \n                        weights=[2, 1, 1])\n</code></pre>\n\n<p>Monitoring this makes much more sense than monitoring accuracy or logloss.</p>\n\n<p>Fourth comment is about dealing with one channel input to 3 channels input pretrained models.  I saw the addition of a convolution layer before the model, or hacking the first level of the model.  There is is a simpler way that Chris Deotte gave for Keras: just concatenate copies of the input.  Here is a Pytorch code, just start the forward method of your model with:</p>\n\n<pre><code>def forward(self, x):\n    h = torch.cat([x, x, x], dim=1)\n    h = self.base_model.features(h)\n</code></pre>\n\n<p>Sure, this is probably way too late for most, but maybe it will help some of you.</p>",
      "rawMarkdown": "Disclaimer: I am not at all a computer vision expert, never got a gold in an image competition.  Therefore take the following with a large grain of salt.  \n\nLet me first say that this forum is a goldmine for whomever wants to start.  I'm very impressed by how much Qishen Ha has shared in particular.  He could have kept his secret sauce entirely hidden.  Needless to say I'm reading and rereading everything he and other top performers are sharing.  \n\nThe following is a little giveback to the community.\n\nThe first comment I have is on the use of feather to save image data.  Why not use numpy arrays?  In my test using numpy is 3x faster.  Here is how I saved train data:\n\n    k = 0\n    train_idx = []\n    train_values = np.zeros((n_image, 137 * 236), dtype = \"uint8\")\n    for i in range(n_file):\n        print(\"reading train file\",i)\n        directory = \"../input/bengaliai-cv19/train_image_data_\"+str(i)+\".parquet\"\n        train_f = pd.read_parquet(directory, engine = \"pyarrow\")\n        train_f.set_index('image_id', inplace=True)\n        train_idx.append(train_f.index.copy())\n        train_values[i * n_file_image : (i + 1) * n_file_image, :] = (255 - train_f.values)\n        del train_f\n        gc.collect()\n\n    train_values = train_values.reshape((-1, 137, 236))\n\n    np.save('../input/bengaliai-cv19/train.npy', train_values)\n\nThe second comment is that many public kernels call sklearn metric with the wrong argument order.  For some reason, Pytorch and Scikit-learn have different ordering.  In sklearn ground truth is the first argument...\n\nThe third comment is about the use of accuracy as metric while training.  I see it in some public kernels and discussions.  Why not use the competition metric directly?\n\nHere is the code I use.  First, just wrap sklearn metric.\n\n    def get_recall(y_true, y_pred):\n        pred_labels = np.argmax(y_pred, axis=1)\n        res = recall_score(y_true, pred_labels, average='macro')\n        return res\n\nThen use it this way in your training or validation loop, if `all_preds` is a numpy array containing the predictions:\n\n        all_preds = np.split(all_preds,\n                             np.cumsum([n_grapheme, n_vowel, n_consonant]), \n                             axis=1\n                            )\n        recall_grapheme = get_recall(y[:, 0], all_preds[0], )\n        recall_vowel = get_recall(y[:, 1], all_preds[1], )\n        recall_consonant = get_recall(y[:, 2], all_preds[2], )\n        recall = np.average([recall_grapheme, recall_vowel, recall_consonant], \n                            weights=[2, 1, 1])\n\nMonitoring this makes much more sense than monitoring accuracy or logloss.\n\nFourth comment is about dealing with one channel input to 3 channels input pretrained models.  I saw the addition of a convolution layer before the model, or hacking the first level of the model.  There is is a simpler way that Chris Deotte gave for Keras: just concatenate copies of the input.  Here is a Pytorch code, just start the forward method of your model with:\n\n    def forward(self, x):\n        h = torch.cat([x, x, x], dim=1)\n        h = self.base_model.features(h)\n\nSure, this is probably way too late for most, but maybe it will help some of you.",
      "votes": 55
    },
    {
      "id": 766048,
      "postDate": "2020-03-07T15:35:54.987Z",
      "content": "<p>To convert 1 to 3 channels using Keras directly in the GPU, just use something like this in the model definition:</p>\n\n<pre><code>input1ch = Input(shape=(224, 224, 1))\ninput3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels\n</code></pre>",
      "rawMarkdown": "To convert 1 to 3 channels using Keras directly in the GPU, just use something like this in the model definition:\n\n    input1ch = Input(shape=(224, 224, 1))\n    input3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels",
      "votes": 5,
      "replies": [
        {
          "id": 766060,
          "postDate": "2020-03-07T16:04:09.117Z",
          "content": "<p>That way is used in many pytorch models too.  But it adds learnable parameters.  I don't know if it is good or no.</p>",
          "rawMarkdown": "That way is used in many pytorch models too.  But it adds learnable parameters.  I don't know if it is good or no.",
          "votes": 1
        },
        {
          "id": 766080,
          "postDate": "2020-03-07T16:28:19.753Z",
          "content": "<p>well, this is a trick i sometimes use:</p>\n\n<p>```\n1.  add a 1 to 3 conversion conv2d layer (or many layer)</p>\n\n<ol>\n<li><p>freeze all imagenet pretrained weights. Now only the conversion conv2d and the last fc logit layer are trainable.</p></li>\n<li><p>start to train. this force the conversion conv2d to produce activation signal to match the magnitude that is suitable for the pretrained weights.</p></li>\n<li><p>when there can be not more improvement, unfreeze everything and do the training as usual.</p></li>\n</ol>\n\n<p>```</p>",
          "rawMarkdown": "well, this is a trick i sometimes use:\n\n```\n1.  add a 1 to 3 conversion conv2d layer (or many layer)\n\n2. freeze all imagenet pretrained weights. Now only the conversion conv2d and the last fc logit layer are trainable.\n\n3. start to train. this force the conversion conv2d to produce activation signal to match the magnitude that is suitable for the pretrained weights.\n\n4. when there can be not more improvement, unfreeze everything and do the training as usual.\n\n```",
          "votes": 4
        },
        {
          "id": 766117,
          "postDate": "2020-03-07T17:18:17.640Z",
          "content": "<p>It adds only 6 learnable parameters:</p>\n\n<hr>\n\n<p>Layer (type)                 Output Shape              Param #   </p>\n\n<hr>\n\n<p>input_1 (InputLayer)         (None, 128, 128, 1)       0         </p>\n\n<hr>\n\n<p>conv2d_1 (Conv2D)            (None, 128, 128, 3)       6         </p>\n\n<hr>\n\n<p>efficientnet-b1 (Model)      (None, 4, 4, 1280)        6575232   </p>\n\n<hr>",
          "rawMarkdown": "It adds only 6 learnable parameters:\n\n_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n_________________________________________________________________\ninput_1 (InputLayer)         (None, 128, 128, 1)       0         \n_________________________________________________________________\nconv2d_1 (Conv2D)            (None, 128, 128, 3)       6         \n_________________________________________________________________\nefficientnet-b1 (Model)      (None, 4, 4, 1280)        6575232   \n_________________________________________________________________",
          "votes": 2
        },
        {
          "id": 766292,
          "postDate": "2020-03-08T00:11:06.897Z",
          "content": "<p>Sure, but are these required?  I thought of what Heng proposes, but did not implement it.</p>",
          "rawMarkdown": "Sure, but are these required?  I thought of what Heng proposes, but did not implement it."
        }
      ]
    },
    {
      "id": 761417,
      "postDate": "2020-03-02T13:48:22.023Z",
      "content": "<p>😲 oops. I really didn't notice the different ordering between Pytorch and Scikit-learn With wrong implementation, I have cv/lb gap around 0.011. After fixing this, I have a smaller gap around 0.006. Thanks for pointing this out.</p>",
      "rawMarkdown": "😲 oops. I really didn't notice the different ordering between Pytorch and Scikit-learn With wrong implementation, I have cv/lb gap around 0.011. After fixing this, I have a smaller gap around 0.006. Thanks for pointing this out.",
      "votes": 5,
      "replies": [
        {
          "id": 761553,
          "postDate": "2020-03-02T16:48:51.130Z",
          "content": "<p>Same. Lol, thats embarrassing) </p>",
          "rawMarkdown": "Same. Lol, thats embarrassing) "
        },
        {
          "id": 761570,
          "postDate": "2020-03-02T17:21:44.690Z",
          "rawMarkdown": ""
        },
        {
          "id": 762647,
          "postDate": "2020-03-03T17:20:54.123Z",
          "content": "<p>you mean 0.006?</p>",
          "rawMarkdown": "you mean 0.006?",
          "votes": 1
        },
        {
          "id": 767743,
          "postDate": "2020-03-10T03:42:21.887Z",
          "content": "<p>Yes 0.006. That was a typo.</p>",
          "rawMarkdown": "Yes 0.006. That was a typo."
        }
      ]
    },
    {
      "id": 760565,
      "postDate": "2020-03-01T13:29:07.627Z",
      "content": "<p>I saved using .npy</p>\n\n<p><a href=\"https://www.kaggle.com/mks2192/bengali-ai-train\">https://www.kaggle.com/mks2192/bengali-ai-train</a></p>",
      "rawMarkdown": "I saved using .npy\n\nhttps://www.kaggle.com/mks2192/bengali-ai-train",
      "votes": 3,
      "replies": [
        {
          "id": 761195,
          "postDate": "2020-03-02T08:35:13.463Z",
          "content": "<p>Good, I was surprised no one shared it before.  I looked at notebooks and forum, but not datasets.</p>",
          "rawMarkdown": "Good, I was surprised no one shared it before.  I looked at notebooks and forum, but not datasets.",
          "votes": 1
        },
        {
          "id": 761344,
          "postDate": "2020-03-02T12:06:52.147Z",
          "content": "<p>Here is the kernal. I have removed all rows and colums which are having vales more than 230. which reduces the images size to a extent</p>\n\n<p><a href=\"https://www.kaggle.com/mks2192/data-preparation\">https://www.kaggle.com/mks2192/data-preparation</a></p>",
          "rawMarkdown": "Here is the kernal. I have removed all rows and colums which are having vales more than 230. which reduces the images size to a extent\n\nhttps://www.kaggle.com/mks2192/data-preparation"
        }
      ]
    },
    {
      "id": 760610,
      "postDate": "2020-03-01T14:21:50Z",
      "content": "<p>using repeat save memory</p>\n\n<p>```\ndef forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)</p>\n\n<p>```</p>",
      "rawMarkdown": "using repeat save memory\n\n```\ndef forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)\n\n```",
      "votes": 4,
      "replies": [
        {
          "id": 760620,
          "postDate": "2020-03-01T14:32:33.653Z",
          "content": "<p>Thanks, I learned something today.</p>",
          "rawMarkdown": "Thanks, I learned something today."
        },
        {
          "id": 760638,
          "postDate": "2020-03-01T15:00:03.443Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> You made me research the topic a bit.  Actually, <code>repeat</code> duplicates x 3 times as well.  What saves memory is <code>expand</code> but you have to specify the right number of elements for all target dimensions.</p>",
          "rawMarkdown": "@hengck23 You made me research the topic a bit.  Actually, `repeat` duplicates x 3 times as well.  What saves memory is `expand` but you have to specify the right number of elements for all target dimensions.",
          "votes": 1
        },
        {
          "id": 760665,
          "postDate": "2020-03-01T15:38:47.353Z",
          "content": "<p>oops. i could have made a mistake.  you should use \"expand\" to save memory</p>\n\n<p><a href=\"https://stackoverflow.com/questions/44593141/stacking-copies-of-an-array-a-torch-tensor-efficiently\">https://stackoverflow.com/questions/44593141/stacking-copies-of-an-array-a-torch-tensor-efficiently</a></p>",
          "rawMarkdown": "oops. i could have made a mistake.  you should use \"expand\" to save memory\n\nhttps://stackoverflow.com/questions/44593141/stacking-copies-of-an-array-a-torch-tensor-efficiently",
          "votes": 3
        },
        {
          "id": 768168,
          "postDate": "2020-03-10T14:07:53.360Z",
          "content": "<p>You don't have to specify all the dimensions, instead you could pass -1 for those that need not to be changed. E.g.  <code>F.conv2d(input_pad, tmp_kernel.expand(c, -1, -1, -1), groups=c, padding=0, stride=1)</code></p>\n\n<p><a href=\"https://github.com/kornia/kornia/blob/master/kornia/filters/filter.py#L86\">https://github.com/kornia/kornia/blob/master/kornia/filters/filter.py#L86</a> </p>",
          "rawMarkdown": "You don't have to specify all the dimensions, instead you could pass -1 for those that need not to be changed. E.g.  `F.conv2d(input_pad, tmp_kernel.expand(c, -1, -1, -1), groups=c, padding=0, stride=1)`\n\nhttps://github.com/kornia/kornia/blob/master/kornia/filters/filter.py#L86 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 762750,
      "postDate": "2020-03-03T19:26:24.310Z",
      "content": "<p>Thanks for pointing out that recall problem! Gonna try it now!</p>",
      "rawMarkdown": "Thanks for pointing out that recall problem! Gonna try it now!",
      "votes": 1
    },
    {
      "id": 760605,
      "postDate": "2020-03-01T14:17:24.400Z",
      "content": "<p>Nice one, thanks. I always use npy images when it's possible. Also agree with 3 channels greyscale input (always marginal improvement <a href=\"https://stackoverflow.com/questions/51995977/how-can-i-use-a-pre-trained-neural-network-with-grayscale-images/54777347#54777347\">compared</a> to alternatives). \nGood luck with your CV gold!</p>",
      "rawMarkdown": "Nice one, thanks. I always use npy images when it's possible. Also agree with 3 channels greyscale input (always marginal improvement [compared](https://stackoverflow.com/questions/51995977/how-can-i-use-a-pre-trained-neural-network-with-grayscale-images/54777347#54777347) to alternatives). \nGood luck with your CV gold!",
      "votes": 1,
      "replies": [
        {
          "id": 761190,
          "postDate": "2020-03-02T08:33:54.820Z",
          "content": "<p>Thanks, I will need more than luck to get a gold.  I'm here for learning, I don't expect to shine by any mean.</p>",
          "rawMarkdown": "Thanks, I will need more than luck to get a gold.  I'm here for learning, I don't expect to shine by any mean."
        }
      ]
    },
    {
      "id": 768087,
      "postDate": "2020-03-10T12:57:11.320Z",
      "content": "<p>Yes, it will be helpful. Thanks for  sharing </p>",
      "rawMarkdown": "Yes, it will be helpful. Thanks for  sharing "
    },
    {
      "id": 767654,
      "postDate": "2020-03-10T00:44:21.693Z",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> - Thanks so much for this post. It has brought forth discussion from a great group of expert ML folks, especially those who specialize in deep learning, NN models. This is a relatively new area for me and this competition has surfaced some interesting approaches to modeling.</p>\n\n<p>Your observation about using the competition metric for training is very insightful. </p>\n\n<p>I also appreciate the contributors of this thread sharing their collective wisdom! People like <a href=\"/titericz\">@titericz</a> , <a href=\"/hengck23\">@hengck23</a>  etc. The masters are at work and sharing what they know!</p>",
      "rawMarkdown": "@cpmpml - Thanks so much for this post. It has brought forth discussion from a great group of expert ML folks, especially those who specialize in deep learning, NN models. This is a relatively new area for me and this competition has surfaced some interesting approaches to modeling.\n\nYour observation about using the competition metric for training is very insightful. \n\nI also appreciate the contributors of this thread sharing their collective wisdom! People like @titericz , @hengck23  etc. The masters are at work and sharing what they know!\n"
    },
    {
      "id": 767012,
      "postDate": "2020-03-09T03:13:49.707Z",
      "content": "<p>Good job!</p>",
      "rawMarkdown": "Good job!"
    },
    {
      "id": 766136,
      "postDate": "2020-03-07T17:51:12.123Z",
      "content": "<p>This is what I used as a starter. Just take average weights of the first layer and combine them into 1.\n<code>\nmodel = models.resnet18(pretrained=True)\ncn1 = nn.Parameter(torch.mean(model.conv1.weight, dim=1, keepdim=True))\nmodel.conv1 = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3,\n                               bias=False)\nmodel.conv1.weight = cn1\nmodel.fc = nn.Linear(512,186)\n</code></p>",
      "rawMarkdown": "This is what I used as a starter. Just take average weights of the first layer and combine them into 1.\n```\nmodel = models.resnet18(pretrained=True)\ncn1 = nn.Parameter(torch.mean(model.conv1.weight, dim=1, keepdim=True))\nmodel.conv1 = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3,\n                               bias=False)\nmodel.conv1.weight = cn1\nmodel.fc = nn.Linear(512,186)\n```"
    },
    {
      "id": 767033,
      "postDate": "2020-03-09T04:10:38.717Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 760573,
      "postDate": "2020-03-01T13:34:35.843Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 767407,
      "postDate": "2020-03-09T15:36:20.320Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 766048,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-03-07T15:35:54.987000",
      "content": "<p>To convert 1 to 3 channels using Keras directly in the GPU, just use something like this in the model definition:</p>\n\n<pre><code>input1ch = Input(shape=(224, 224, 1))\ninput3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels\n</code></pre>",
      "votes": 5,
      "replies": [
        {
          "id": 766060,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-07T16:04:09.117000",
          "content": "<p>That way is used in many pytorch models too.  But it adds learnable parameters.  I don't know if it is good or no.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766080,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-07T16:28:19.753000",
          "content": "<p>well, this is a trick i sometimes use:</p>\n\n<p>```\n1.  add a 1 to 3 conversion conv2d layer (or many layer)</p>\n\n<ol>\n<li><p>freeze all imagenet pretrained weights. Now only the conversion conv2d and the last fc logit layer are trainable.</p></li>\n<li><p>start to train. this force the conversion conv2d to produce activation signal to match the magnitude that is suitable for the pretrained weights.</p></li>\n<li><p>when there can be not more improvement, unfreeze everything and do the training as usual.</p></li>\n</ol>\n\n<p>```</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 766117,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2020-03-07T17:18:17.640000",
          "content": "<p>It adds only 6 learnable parameters:</p>\n\n<hr>\n\n<p>Layer (type)                 Output Shape              Param #   </p>\n\n<hr>\n\n<p>input_1 (InputLayer)         (None, 128, 128, 1)       0         </p>\n\n<hr>\n\n<p>conv2d_1 (Conv2D)            (None, 128, 128, 3)       6         </p>\n\n<hr>\n\n<p>efficientnet-b1 (Model)      (None, 4, 4, 1280)        6575232   </p>\n\n<hr>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 766292,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-08T00:11:06.897000",
          "content": "<p>Sure, but are these required?  I thought of what Heng proposes, but did not implement it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761417,
      "author_name": "syoya",
      "author_url": "",
      "post_date": "2020-03-02T13:48:22.023000",
      "content": "<p>😲 oops. I really didn't notice the different ordering between Pytorch and Scikit-learn With wrong implementation, I have cv/lb gap around 0.011. After fixing this, I have a smaller gap around 0.006. Thanks for pointing this out.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 761553,
          "author_name": "Kupchanski",
          "author_url": "",
          "post_date": "2020-03-02T16:48:51.130000",
          "content": "<p>Same. Lol, thats embarrassing) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 761570,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2020-03-02T17:21:44.690000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762647,
          "author_name": "sajwankit",
          "author_url": "",
          "post_date": "2020-03-03T17:20:54.123000",
          "content": "<p>you mean 0.006?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767743,
          "author_name": "syoya",
          "author_url": "",
          "post_date": "2020-03-10T03:42:21.887000",
          "content": "<p>Yes 0.006. That was a typo.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 760565,
      "author_name": "Manoj",
      "author_url": "",
      "post_date": "2020-03-01T13:29:07.627000",
      "content": "<p>I saved using .npy</p>\n\n<p><a href=\"https://www.kaggle.com/mks2192/bengali-ai-train\">https://www.kaggle.com/mks2192/bengali-ai-train</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 761195,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-02T08:35:13.463000",
          "content": "<p>Good, I was surprised no one shared it before.  I looked at notebooks and forum, but not datasets.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 761344,
          "author_name": "Manoj",
          "author_url": "",
          "post_date": "2020-03-02T12:06:52.147000",
          "content": "<p>Here is the kernal. I have removed all rows and colums which are having vales more than 230. which reduces the images size to a extent</p>\n\n<p><a href=\"https://www.kaggle.com/mks2192/data-preparation\">https://www.kaggle.com/mks2192/data-preparation</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 760610,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-01T14:21:50",
      "content": "<p>using repeat save memory</p>\n\n<p>```\ndef forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)</p>\n\n<p>```</p>",
      "votes": 4,
      "replies": [
        {
          "id": 760620,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-01T14:32:33.653000",
          "content": "<p>Thanks, I learned something today.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 760638,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-01T15:00:03.443000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> You made me research the topic a bit.  Actually, <code>repeat</code> duplicates x 3 times as well.  What saves memory is <code>expand</code> but you have to specify the right number of elements for all target dimensions.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 760665,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-01T15:38:47.353000",
          "content": "<p>oops. i could have made a mistake.  you should use \"expand\" to save memory</p>\n\n<p><a href=\"https://stackoverflow.com/questions/44593141/stacking-copies-of-an-array-a-torch-tensor-efficiently\">https://stackoverflow.com/questions/44593141/stacking-copies-of-an-array-a-torch-tensor-efficiently</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 768168,
          "author_name": "old-ufo",
          "author_url": "",
          "post_date": "2020-03-10T14:07:53.360000",
          "content": "<p>You don't have to specify all the dimensions, instead you could pass -1 for those that need not to be changed. E.g.  <code>F.conv2d(input_pad, tmp_kernel.expand(c, -1, -1, -1), groups=c, padding=0, stride=1)</code></p>\n\n<p><a href=\"https://github.com/kornia/kornia/blob/master/kornia/filters/filter.py#L86\">https://github.com/kornia/kornia/blob/master/kornia/filters/filter.py#L86</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 762750,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-03-03T19:26:24.310000",
      "content": "<p>Thanks for pointing out that recall problem! Gonna try it now!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 760605,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2020-03-01T14:17:24.400000",
      "content": "<p>Nice one, thanks. I always use npy images when it's possible. Also agree with 3 channels greyscale input (always marginal improvement <a href=\"https://stackoverflow.com/questions/51995977/how-can-i-use-a-pre-trained-neural-network-with-grayscale-images/54777347#54777347\">compared</a> to alternatives). \nGood luck with your CV gold!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 761190,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-02T08:33:54.820000",
          "content": "<p>Thanks, I will need more than luck to get a gold.  I'm here for learning, I don't expect to shine by any mean.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 768087,
      "author_name": "Kranti Kumar",
      "author_url": "",
      "post_date": "2020-03-10T12:57:11.320000",
      "content": "<p>Yes, it will be helpful. Thanks for  sharing </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767654,
      "author_name": "Bill Holst",
      "author_url": "",
      "post_date": "2020-03-10T00:44:21.693000",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> - Thanks so much for this post. It has brought forth discussion from a great group of expert ML folks, especially those who specialize in deep learning, NN models. This is a relatively new area for me and this competition has surfaced some interesting approaches to modeling.</p>\n\n<p>Your observation about using the competition metric for training is very insightful. </p>\n\n<p>I also appreciate the contributors of this thread sharing their collective wisdom! People like <a href=\"/titericz\">@titericz</a> , <a href=\"/hengck23\">@hengck23</a>  etc. The masters are at work and sharing what they know!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767012,
      "author_name": "Alex Ross",
      "author_url": "",
      "post_date": "2020-03-09T03:13:49.707000",
      "content": "<p>Good job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 766136,
      "author_name": "garyong",
      "author_url": "",
      "post_date": "2020-03-07T17:51:12.123000",
      "content": "<p>This is what I used as a starter. Just take average weights of the first layer and combine them into 1.\n<code>\nmodel = models.resnet18(pretrained=True)\ncn1 = nn.Parameter(torch.mean(model.conv1.weight, dim=1, keepdim=True))\nmodel.conv1 = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3,\n                               bias=False)\nmodel.conv1.weight = cn1\nmodel.fc = nn.Linear(512,186)\n</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767033,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-09T04:10:38.717000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 760573,
      "author_name": "forthelast",
      "author_url": "",
      "post_date": "2020-03-01T13:34:35.843000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 767407,
      "author_name": "Alake Edah",
      "author_url": "",
      "post_date": "2020-03-09T15:36:20.320000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "760541": "Disclaimer: I am not at all a computer vision expert, never got a gold in an image competition.  Therefore take the following with a large grain of salt.  \n\nLet me first say that this forum is a goldmine for whomever wants to start.  I'm very impressed by how much Qishen Ha has shared in particular.  He could have kept his secret sauce entirely hidden.  Needless to say I'm reading and rereading everything he and other top performers are sharing.  \n\nThe following is a little giveback to the community.\n\nThe first comment I have is on the use of feather to save image data.  Why not use numpy arrays?  In my test using numpy is 3x faster.  Here is how I saved train data:\n\n    k = 0\n    train_idx = []\n    train_values = np.zeros((n_image, 137 * 236), dtype = \"uint8\")\n    for i in range(n_file):\n        print(\"reading train file\",i)\n        directory = \"../input/bengaliai-cv19/train_image_data_\"+str(i)+\".parquet\"\n        train_f = pd.read_parquet(directory, engine = \"pyarrow\")\n        train_f.set_index('image_id', inplace=True)\n        train_idx.append(train_f.index.copy())\n        train_values[i * n_file_image : (i + 1) * n_file_image, :] = (255 - train_f.values)\n        del train_f\n        gc.collect()\n\n    train_values = train_values.reshape((-1, 137, 236))\n\n    np.save('../input/bengaliai-cv19/train.npy', train_values)\n\nThe second comment is that many public kernels call sklearn metric with the wrong argument order.  For some reason, Pytorch and Scikit-learn have different ordering.  In sklearn ground truth is the first argument...\n\nThe third comment is about the use of accuracy as metric while training.  I see it in some public kernels and discussions.  Why not use the competition metric directly?\n\nHere is the code I use.  First, just wrap sklearn metric.\n\n    def get_recall(y_true, y_pred):\n        pred_labels = np.argmax(y_pred, axis=1)\n        res = recall_score(y_true, pred_labels, average='macro')\n        return res\n\nThen use it this way in your training or validation loop, if `all_preds` is a numpy array containing the predictions:\n\n        all_preds = np.split(all_preds,\n                             np.cumsum([n_grapheme, n_vowel, n_consonant]), \n                             axis=1\n                            )\n        recall_grapheme = get_recall(y[:, 0], all_preds[0], )\n        recall_vowel = get_recall(y[:, 1], all_preds[1], )\n        recall_consonant = get_recall(y[:, 2], all_preds[2], )\n        recall = np.average([recall_grapheme, recall_vowel, recall_consonant], \n                            weights=[2, 1, 1])\n\nMonitoring this makes much more sense than monitoring accuracy or logloss.\n\nFourth comment is about dealing with one channel input to 3 channels input pretrained models.  I saw the addition of a convolution layer before the model, or hacking the first level of the model.  There is is a simpler way that Chris Deotte gave for Keras: just concatenate copies of the input.  Here is a Pytorch code, just start the forward method of your model with:\n\n    def forward(self, x):\n        h = torch.cat([x, x, x], dim=1)\n        h = self.base_model.features(h)\n\nSure, this is probably way too late for most, but maybe it will help some of you.",
    "766048": "To convert 1 to 3 channels using Keras directly in the GPU, just use something like this in the model definition:\n\n    input1ch = Input(shape=(224, 224, 1))\n    input3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels",
    "761417": "😲 oops. I really didn't notice the different ordering between Pytorch and Scikit-learn With wrong implementation, I have cv/lb gap around 0.011. After fixing this, I have a smaller gap around 0.006. Thanks for pointing this out.",
    "760565": "I saved using .npy\n\nhttps://www.kaggle.com/mks2192/bengali-ai-train",
    "760610": "using repeat save memory\n\n```\ndef forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)\n\n```",
    "762750": "Thanks for pointing out that recall problem! Gonna try it now!",
    "760605": "Nice one, thanks. I always use npy images when it's possible. Also agree with 3 channels greyscale input (always marginal improvement [compared](https://stackoverflow.com/questions/51995977/how-can-i-use-a-pre-trained-neural-network-with-grayscale-images/54777347#54777347) to alternatives). \nGood luck with your CV gold!",
    "768087": "Yes, it will be helpful. Thanks for  sharing ",
    "767654": "@cpmpml - Thanks so much for this post. It has brought forth discussion from a great group of expert ML folks, especially those who specialize in deep learning, NN models. This is a relatively new area for me and this competition has surfaced some interesting approaches to modeling.\n\nYour observation about using the competition metric for training is very insightful. \n\nI also appreciate the contributors of this thread sharing their collective wisdom! People like @titericz , @hengck23  etc. The masters are at work and sharing what they know!\n",
    "767012": "Good job!",
    "766136": "This is what I used as a starter. Just take average weights of the first layer and combine them into 1.\n```\nmodel = models.resnet18(pretrained=True)\ncn1 = nn.Parameter(torch.mean(model.conv1.weight, dim=1, keepdim=True))\nmodel.conv1 = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3,\n                               bias=False)\nmodel.conv1.weight = cn1\nmodel.fc = nn.Linear(512,186)\n```",
    "767033": "",
    "760573": "Thanks for sharing!",
    "767407": "Thanks for sharing!"
  }
}