{
  "id": 35456,
  "title": "My approach using coordinates as regression target ",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/35456",
  "author_name": "",
  "post_date": "2017-06-28T18:33:45.732650700Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Firstly, thank you very much @threeplusone for the sea lion coordinates. I learned a lot of thing from your code.</p>\n\n<p>I'm not so high in the leader board but I guess my approach may be worth mentioning.\nI tried to feed the given information as precise as possible. So I used each sea lions' coordinates as regression target.</p>\n\n<p>I scaled down the input to 1/2 size, and applied UNET. But at the end of network, I replaced sigmoid with two separate NN path:\n       - 1x1 Conv2D with depth 1024 to predict whether sea lion is nearby in this pixel\n       - 1x1 Conv2D with depth 1024 to predict two numbers - relative offset of the nearest sea from this pixel.\nLoss function for offset is like sum(true_hit * (RMSE between true coordinate and predicted coordiante))</p>\n\n<p>This works quite well for isolated sea lions but didn't work well when multiple sea lions are nearby. The attached images shows predicted location of sea lions and brighter red is more confident prediction.</p>\n\n<p>This approach is quite slow. After seeing other solutions, I realized that I used too high depth for conv layers - higher depth yielded a little bit better solution, so I kept increasing depth for UNET. But now it looks like that tuning the simpler model with scale/augmentation was better use of time.</p>\n\n<p>Anyway this approach got me find 70% of sea lions on validation set, and I noticed that the orientations of predicted point clusters are usually aligned with sea lion orientation(used skimage.measure.label and regionprop). So I rotated the sea lion patch and cut 1:2 area only - this reduced the amount of data for second pipeline that classifies sea lion type and also increased the classification accuracy by 10%.</p>\n\n<p>I got RMSE 16 from validation set, but public leader board score was just 25. I couldn't get it any higher even after a month of tuning.</p>\n\n<p>And after seeing mrgroom's solution, I decided to try similar approach - split image to 512x512 patch and predict the counts. This method is much simpler but yielded LB score around 25 also :(</p>\n\n<p>I wasn't thinking of merging two results until two days before deadline at all. \nBut after merging two results, I got public LB score around 21.13.</p>",
  "messages": [
    {
      "id": "197078",
      "postDate": "06/28/2017 18:33:45",
      "content": "<p>Firstly, thank you very much @threeplusone for the sea lion coordinates. I learned a lot of thing from your code.</p>\n\n<p>I'm not so high in the leader board but I guess my approach may be worth mentioning.\nI tried to feed the given information as precise as possible. So I used each sea lions' coordinates as regression target.</p>\n\n<p>I scaled down the input to 1/2 size, and applied UNET. But at the end of network, I replaced sigmoid with two separate NN path:\n       - 1x1 Conv2D with depth 1024 to predict whether sea lion is nearby in this pixel\n       - 1x1 Conv2D with depth 1024 to predict two numbers - relative offset of the nearest sea from this pixel.\nLoss function for offset is like sum(true_hit * (RMSE between true coordinate and predicted coordiante))</p>\n\n<p>This works quite well for isolated sea lions but didn't work well when multiple sea lions are nearby. The attached images shows predicted location of sea lions and brighter red is more confident prediction.</p>\n\n<p>This approach is quite slow. After seeing other solutions, I realized that I used too high depth for conv layers - higher depth yielded a little bit better solution, so I kept increasing depth for UNET. But now it looks like that tuning the simpler model with scale/augmentation was better use of time.</p>\n\n<p>Anyway this approach got me find 70% of sea lions on validation set, and I noticed that the orientations of predicted point clusters are usually aligned with sea lion orientation(used skimage.measure.label and regionprop). So I rotated the sea lion patch and cut 1:2 area only - this reduced the amount of data for second pipeline that classifies sea lion type and also increased the classification accuracy by 10%.</p>\n\n<p>I got RMSE 16 from validation set, but public leader board score was just 25. I couldn't get it any higher even after a month of tuning.</p>\n\n<p>And after seeing mrgroom's solution, I decided to try similar approach - split image to 512x512 patch and predict the counts. This method is much simpler but yielded LB score around 25 also :(</p>\n\n<p>I wasn't thinking of merging two results until two days before deadline at all. \nBut after merging two results, I got public LB score around 21.13.</p>",
      "rawMarkdown": "Firstly, thank you very much @threeplusone for the sea lion coordinates. I learned a lot of thing from your code.\n\nI'm not so high in the leader board but I guess my approach may be worth mentioning.\nI tried to feed the given information as precise as possible. So I used each sea lions' coordinates as regression target.\n\nI scaled down the input to 1/2 size, and applied UNET. But at the end of network, I replaced sigmoid with two separate NN path:\n       - 1x1 Conv2D with depth 1024 to predict whether sea lion is nearby in this pixel\n       - 1x1 Conv2D with depth 1024 to predict two numbers - relative offset of the nearest sea from this pixel.\nLoss function for offset is like sum(true_hit * (RMSE between true coordinate and predicted coordiante))\n\nThis works quite well for isolated sea lions but didn't work well when multiple sea lions are nearby. The attached images shows predicted location of sea lions and brighter red is more confident prediction.\n\nThis approach is quite slow. After seeing other solutions, I realized that I used too high depth for conv layers - higher depth yielded a little bit better solution, so I kept increasing depth for UNET. But now it looks like that tuning the simpler model with scale/augmentation was better use of time.\n\nAnyway this approach got me find 70% of sea lions on validation set, and I noticed that the orientations of predicted point clusters are usually aligned with sea lion orientation(used skimage.measure.label and regionprop). So I rotated the sea lion patch and cut 1:2 area only - this reduced the amount of data for second pipeline that classifies sea lion type and also increased the classification accuracy by 10%.\n\nI got RMSE 16 from validation set, but public leader board score was just 25. I couldn't get it any higher even after a month of tuning.\n\nAnd after seeing mrgroom's solution, I decided to try similar approach - split image to 512x512 patch and predict the counts. This method is much simpler but yielded LB score around 25 also :(\n\nI wasn't thinking of merging two results until two days before deadline at all. \nBut after merging two results, I got public LB score around 21.13.",
      "votes": null
    },
    {
      "id": "197375",
      "postDate": "06/29/2017 13:31:17",
      "content": "<p>@JandJ: Congrats.\nWould you mind sharing the code of your \"two path NN\" ?</p>",
      "rawMarkdown": "JandJ: Congrats.\nWould you mind sharing the code of your \"two path NN\" ?",
      "votes": null
    },
    {
      "id": "197473",
      "postDate": "06/29/2017 16:57:14",
      "content": "<p>This was the first time that I modified existing model significantly, so it may be in wrong direction in implementation.\nBut using coordinates itself seems to be good idea (it wasn't good for clustered sea lions, but may be good for other applications), so I want to get some feedback by sharing.</p>\n\n<p>BTW, calculating loss using multiple target coordinates didn't work out really. NN just predicts center of multiple sea lions. <br>\nAnd smaller depth u-net (even no u-net) yielded similar(less accurate) results. It's trade off between computing time and small bit of accuracy. </p>\n\n<pre>lambda_cls = 1\nlambda_offset = .001\n\n# y_true : (:, 256,256, 1), y_pred(:,256,256, 1)\n# y_true is sorted to have hit first, so just compare to the first entry\ndef cls_loss(y_true, y_pred):\n    #print(K.int_shape(y_true), K.int_shape(y_pred))\n    return lambda_cls * K.mean(K.binary_crossentropy(y_pred[:, :, :], y_true))\n\n# y_true: (batch_size, 256, 256, 3*k), y_pred(:,256,256,3)\n# y_pred[:,:,:,0] is hit/miss around target area. y_pred[...,1:3] is row/col offset from center of each blk to nearest sea lion\n# y_true has coordinates of k number of targets.\n# Goal is not to predict all true coord but not to penalize when NN detects non-nearest points\ndef offset_loss(y_true, y_pred):\n    dists = []\n    max_dist = 999999\n\n    for k in range(hp.p1_cnt_per_blk):\n        r_diff = K.pow(y_true[:, :, :, k*3+1] - y_pred[:, :, :, 1], 2)\n        c_diff = K.pow(y_true[:, :, :, k*3+2] - y_pred[:, :, :, 2], 2)\n        dist = (r_diff+c_diff)/2 * y_true[:, :, :, k*3]\n        dist += max_dist * (1-y_true[:, :, :, k*3]) # Handle empty entry\n        dists.append(dist)\n    os_stacked = K.stack(dists, axis=3)\n    os_loss = K.min(os_stacked, axis=-1)\n    return lambda_offset * K.sum(y_true[:, :, :, 0] * os_loss)/(0.01+K.sum(y_true[:, :, :, 0]))\n\n    def get_model():\n        input = Input(shape=hp.p1_vgg_in_dim)\n        vgg = VGG16(include_top=False, weights='imagenet', input_tensor=input, input_shape=hp.p1_vgg_in_dim)\n        x = vgg.layers[hp.p1_vgg_layers_to_keep].output\n        for layer in vgg.layers[:hp.p1_vgg_layers_to_keep]: \n            layer.trainable=False\n        # After VGG 9 layers, output is (128,128,256)\n        conv1 = x\n        pool1 = MaxPooling2D(pool_size=(2,2))(conv1)\n\n        conv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-1')(pool1))\n        conv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-2')(conv2))\n        pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)  #=(64,64,3)\n\n        conv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-1')(pool2))\n        conv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-2')(conv3))\n        pool3 = MaxPooling2D(pool_size=(2, 2))(conv3)  # =(32,32,3)\n\n        conv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-1')(pool3)\n        conv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-2')(conv4)\n\n        up1 = concatenate([UpSampling2D(size=(2, 2))(conv4), conv3], axis=3, name='conc1')\n        conv5 = BatchNormalization()(Conv2D(512, (3,3), activation='relu', padding='same', name='conv5-1')(up1))\n        conv5 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv5-2')(conv5))\n\n        up2 = concatenate([UpSampling2D(size=(2, 2))(conv5), conv2], axis=3, name='conc2')\n        conv6 = BatchNormalization()(Conv2D(256, (3,3), activation='relu', padding='same', name='conv6-1')(up2))\n        conv6 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv6-2')(conv6))\n\n        cls_conv_f = BatchNormalization()(\n            Conv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n        cls_conv = Conv2D(1, (1, 1), activation='sigmoid', name='cls')(cls_conv_f)\n\n        off_conv_f = BatchNormalization()(\n            Conv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n        off_conv_f = Conv2D(2, (1, 1), activation='linear', kernel_initializer='zero', padding='same')\\\n                    (off_conv_f)\n        off_conv = Concatenate(name='offset')([cls_conv, off_conv_f])\n\n        model = Model(inputs=[input], outputs=[cls_conv, off_conv])\n        model.compile(optimizer=Adam(lr=hp.p1_lr), loss={'cls':cls_loss, 'offset':offset_loss})\n\n        return model\n\n</pre>",
      "rawMarkdown": "This was the first time that I modified existing model significantly, so it may be in wrong direction in implementation.\nBut using coordinates itself seems to be good idea (it wasn't good for clustered sea lions, but may be good for other applications), so I want to get some feedback by sharing.\n\nBTW, calculating loss using multiple target coordinates didn't work out really. NN just predicts center of multiple sea lions.  \nAnd smaller depth u-net (even no u-net) yielded similar(less accurate) results. It's trade off between computing time and small bit of accuracy. \n\n<pre>lambda_cls = 1\nlambda_offset = .001\n\n# y_true : (:, 256,256, 1), y_pred(:,256,256, 1)\n# y_true is sorted to have hit first, so just compare to the first entry\ndef cls_loss(y_true, y_pred):\n\t#print(K.int_shape(y_true), K.int_shape(y_pred))\n\treturn lambda_cls * K.mean(K.binary_crossentropy(y_pred[:, :, :], y_true))\n\n# y_true: (batch_size, 256, 256, 3*k), y_pred(:,256,256,3)\n# y_pred[:,:,:,0] is hit/miss around target area. y_pred[...,1:3] is row/col offset from center of each blk to nearest sea lion\n# y_true has coordinates of k number of targets.\n# Goal is not to predict all true coord but not to penalize when NN detects non-nearest points\ndef offset_loss(y_true, y_pred):\n\tdists = []\n\tmax_dist = 999999\n\n\tfor k in range(hp.p1_cnt_per_blk):\n\t\tr_diff = K.pow(y_true[:, :, :, k*3+1] - y_pred[:, :, :, 1], 2)\n\t\tc_diff = K.pow(y_true[:, :, :, k*3+2] - y_pred[:, :, :, 2], 2)\n\t\tdist = (r_diff+c_diff)/2 * y_true[:, :, :, k*3]\n\t\tdist += max_dist * (1-y_true[:, :, :, k*3]) # Handle empty entry\n\t\tdists.append(dist)\n\tos_stacked = K.stack(dists, axis=3)\n\tos_loss = K.min(os_stacked, axis=-1)\n\treturn lambda_offset * K.sum(y_true[:, :, :, 0] * os_loss)/(0.01+K.sum(y_true[:, :, :, 0]))\n\n\tdef get_model():\n\t\tinput = Input(shape=hp.p1_vgg_in_dim)\n\t\tvgg = VGG16(include_top=False, weights='imagenet', input_tensor=input, input_shape=hp.p1_vgg_in_dim)\n\t\tx = vgg.layers[hp.p1_vgg_layers_to_keep].output\n\t\tfor layer in vgg.layers[:hp.p1_vgg_layers_to_keep]: \n\t\t\tlayer.trainable=False\n\t\t# After VGG 9 layers, output is (128,128,256)\n\t\tconv1 = x\n\t\tpool1 = MaxPooling2D(pool_size=(2,2))(conv1)\n\n\t\tconv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-1')(pool1))\n\t\tconv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-2')(conv2))\n\t\tpool2 = MaxPooling2D(pool_size=(2, 2))(conv2)  #=(64,64,3)\n\n\t\tconv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-1')(pool2))\n\t\tconv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-2')(conv3))\n\t\tpool3 = MaxPooling2D(pool_size=(2, 2))(conv3)  # =(32,32,3)\n\n\t\tconv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-1')(pool3)\n\t\tconv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-2')(conv4)\n\n\t\tup1 = concatenate([UpSampling2D(size=(2, 2))(conv4), conv3], axis=3, name='conc1')\n\t\tconv5 = BatchNormalization()(Conv2D(512, (3,3), activation='relu', padding='same', name='conv5-1')(up1))\n\t\tconv5 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv5-2')(conv5))\n\n\t\tup2 = concatenate([UpSampling2D(size=(2, 2))(conv5), conv2], axis=3, name='conc2')\n\t\tconv6 = BatchNormalization()(Conv2D(256, (3,3), activation='relu', padding='same', name='conv6-1')(up2))\n\t\tconv6 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv6-2')(conv6))\n\n\t\tcls_conv_f = BatchNormalization()(\n\t\t\tConv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n\t\tcls_conv = Conv2D(1, (1, 1), activation='sigmoid', name='cls')(cls_conv_f)\n\n\t\toff_conv_f = BatchNormalization()(\n\t\t\tConv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n\t\toff_conv_f = Conv2D(2, (1, 1), activation='linear', kernel_initializer='zero', padding='same')\\\n\t\t\t\t\t(off_conv_f)\n\t\toff_conv = Concatenate(name='offset')([cls_conv, off_conv_f])\n\n\t\tmodel = Model(inputs=[input], outputs=[cls_conv, off_conv])\n\t\tmodel.compile(optimizer=Adam(lr=hp.p1_lr), loss={'cls':cls_loss, 'offset':offset_loss})\n\n\t\treturn model\n\n</pre>",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 197375,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "06/29/2017 13:31:17",
      "content": "<p>@JandJ: Congrats.\nWould you mind sharing the code of your \"two path NN\" ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 197473,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "06/29/2017 16:57:14",
          "content": "<p>This was the first time that I modified existing model significantly, so it may be in wrong direction in implementation.\nBut using coordinates itself seems to be good idea (it wasn't good for clustered sea lions, but may be good for other applications), so I want to get some feedback by sharing.</p>\n\n<p>BTW, calculating loss using multiple target coordinates didn't work out really. NN just predicts center of multiple sea lions. <br>\nAnd smaller depth u-net (even no u-net) yielded similar(less accurate) results. It's trade off between computing time and small bit of accuracy. </p>\n\n<pre>lambda_cls = 1\nlambda_offset = .001\n\n# y_true : (:, 256,256, 1), y_pred(:,256,256, 1)\n# y_true is sorted to have hit first, so just compare to the first entry\ndef cls_loss(y_true, y_pred):\n    #print(K.int_shape(y_true), K.int_shape(y_pred))\n    return lambda_cls * K.mean(K.binary_crossentropy(y_pred[:, :, :], y_true))\n\n# y_true: (batch_size, 256, 256, 3*k), y_pred(:,256,256,3)\n# y_pred[:,:,:,0] is hit/miss around target area. y_pred[...,1:3] is row/col offset from center of each blk to nearest sea lion\n# y_true has coordinates of k number of targets.\n# Goal is not to predict all true coord but not to penalize when NN detects non-nearest points\ndef offset_loss(y_true, y_pred):\n    dists = []\n    max_dist = 999999\n\n    for k in range(hp.p1_cnt_per_blk):\n        r_diff = K.pow(y_true[:, :, :, k*3+1] - y_pred[:, :, :, 1], 2)\n        c_diff = K.pow(y_true[:, :, :, k*3+2] - y_pred[:, :, :, 2], 2)\n        dist = (r_diff+c_diff)/2 * y_true[:, :, :, k*3]\n        dist += max_dist * (1-y_true[:, :, :, k*3]) # Handle empty entry\n        dists.append(dist)\n    os_stacked = K.stack(dists, axis=3)\n    os_loss = K.min(os_stacked, axis=-1)\n    return lambda_offset * K.sum(y_true[:, :, :, 0] * os_loss)/(0.01+K.sum(y_true[:, :, :, 0]))\n\n    def get_model():\n        input = Input(shape=hp.p1_vgg_in_dim)\n        vgg = VGG16(include_top=False, weights='imagenet', input_tensor=input, input_shape=hp.p1_vgg_in_dim)\n        x = vgg.layers[hp.p1_vgg_layers_to_keep].output\n        for layer in vgg.layers[:hp.p1_vgg_layers_to_keep]: \n            layer.trainable=False\n        # After VGG 9 layers, output is (128,128,256)\n        conv1 = x\n        pool1 = MaxPooling2D(pool_size=(2,2))(conv1)\n\n        conv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-1')(pool1))\n        conv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-2')(conv2))\n        pool2 = MaxPooling2D(pool_size=(2, 2))(conv2)  #=(64,64,3)\n\n        conv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-1')(pool2))\n        conv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-2')(conv3))\n        pool3 = MaxPooling2D(pool_size=(2, 2))(conv3)  # =(32,32,3)\n\n        conv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-1')(pool3)\n        conv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-2')(conv4)\n\n        up1 = concatenate([UpSampling2D(size=(2, 2))(conv4), conv3], axis=3, name='conc1')\n        conv5 = BatchNormalization()(Conv2D(512, (3,3), activation='relu', padding='same', name='conv5-1')(up1))\n        conv5 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv5-2')(conv5))\n\n        up2 = concatenate([UpSampling2D(size=(2, 2))(conv5), conv2], axis=3, name='conc2')\n        conv6 = BatchNormalization()(Conv2D(256, (3,3), activation='relu', padding='same', name='conv6-1')(up2))\n        conv6 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv6-2')(conv6))\n\n        cls_conv_f = BatchNormalization()(\n            Conv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n        cls_conv = Conv2D(1, (1, 1), activation='sigmoid', name='cls')(cls_conv_f)\n\n        off_conv_f = BatchNormalization()(\n            Conv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n        off_conv_f = Conv2D(2, (1, 1), activation='linear', kernel_initializer='zero', padding='same')\\\n                    (off_conv_f)\n        off_conv = Concatenate(name='offset')([cls_conv, off_conv_f])\n\n        model = Model(inputs=[input], outputs=[cls_conv, off_conv])\n        model.compile(optimizer=Adam(lr=hp.p1_lr), loss={'cls':cls_loss, 'offset':offset_loss})\n\n        return model\n\n</pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "197078": "Firstly, thank you very much @threeplusone for the sea lion coordinates. I learned a lot of thing from your code.\n\nI'm not so high in the leader board but I guess my approach may be worth mentioning.\nI tried to feed the given information as precise as possible. So I used each sea lions' coordinates as regression target.\n\nI scaled down the input to 1/2 size, and applied UNET. But at the end of network, I replaced sigmoid with two separate NN path:\n       - 1x1 Conv2D with depth 1024 to predict whether sea lion is nearby in this pixel\n       - 1x1 Conv2D with depth 1024 to predict two numbers - relative offset of the nearest sea from this pixel.\nLoss function for offset is like sum(true_hit * (RMSE between true coordinate and predicted coordiante))\n\nThis works quite well for isolated sea lions but didn't work well when multiple sea lions are nearby. The attached images shows predicted location of sea lions and brighter red is more confident prediction.\n\nThis approach is quite slow. After seeing other solutions, I realized that I used too high depth for conv layers - higher depth yielded a little bit better solution, so I kept increasing depth for UNET. But now it looks like that tuning the simpler model with scale/augmentation was better use of time.\n\nAnyway this approach got me find 70% of sea lions on validation set, and I noticed that the orientations of predicted point clusters are usually aligned with sea lion orientation(used skimage.measure.label and regionprop). So I rotated the sea lion patch and cut 1:2 area only - this reduced the amount of data for second pipeline that classifies sea lion type and also increased the classification accuracy by 10%.\n\nI got RMSE 16 from validation set, but public leader board score was just 25. I couldn't get it any higher even after a month of tuning.\n\nAnd after seeing mrgroom's solution, I decided to try similar approach - split image to 512x512 patch and predict the counts. This method is much simpler but yielded LB score around 25 also :(\n\nI wasn't thinking of merging two results until two days before deadline at all. \nBut after merging two results, I got public LB score around 21.13.",
    "197375": "JandJ: Congrats.\nWould you mind sharing the code of your \"two path NN\" ?",
    "197473": "This was the first time that I modified existing model significantly, so it may be in wrong direction in implementation.\nBut using coordinates itself seems to be good idea (it wasn't good for clustered sea lions, but may be good for other applications), so I want to get some feedback by sharing.\n\nBTW, calculating loss using multiple target coordinates didn't work out really. NN just predicts center of multiple sea lions.  \nAnd smaller depth u-net (even no u-net) yielded similar(less accurate) results. It's trade off between computing time and small bit of accuracy. \n\n<pre>lambda_cls = 1\nlambda_offset = .001\n\n# y_true : (:, 256,256, 1), y_pred(:,256,256, 1)\n# y_true is sorted to have hit first, so just compare to the first entry\ndef cls_loss(y_true, y_pred):\n\t#print(K.int_shape(y_true), K.int_shape(y_pred))\n\treturn lambda_cls * K.mean(K.binary_crossentropy(y_pred[:, :, :], y_true))\n\n# y_true: (batch_size, 256, 256, 3*k), y_pred(:,256,256,3)\n# y_pred[:,:,:,0] is hit/miss around target area. y_pred[...,1:3] is row/col offset from center of each blk to nearest sea lion\n# y_true has coordinates of k number of targets.\n# Goal is not to predict all true coord but not to penalize when NN detects non-nearest points\ndef offset_loss(y_true, y_pred):\n\tdists = []\n\tmax_dist = 999999\n\n\tfor k in range(hp.p1_cnt_per_blk):\n\t\tr_diff = K.pow(y_true[:, :, :, k*3+1] - y_pred[:, :, :, 1], 2)\n\t\tc_diff = K.pow(y_true[:, :, :, k*3+2] - y_pred[:, :, :, 2], 2)\n\t\tdist = (r_diff+c_diff)/2 * y_true[:, :, :, k*3]\n\t\tdist += max_dist * (1-y_true[:, :, :, k*3]) # Handle empty entry\n\t\tdists.append(dist)\n\tos_stacked = K.stack(dists, axis=3)\n\tos_loss = K.min(os_stacked, axis=-1)\n\treturn lambda_offset * K.sum(y_true[:, :, :, 0] * os_loss)/(0.01+K.sum(y_true[:, :, :, 0]))\n\n\tdef get_model():\n\t\tinput = Input(shape=hp.p1_vgg_in_dim)\n\t\tvgg = VGG16(include_top=False, weights='imagenet', input_tensor=input, input_shape=hp.p1_vgg_in_dim)\n\t\tx = vgg.layers[hp.p1_vgg_layers_to_keep].output\n\t\tfor layer in vgg.layers[:hp.p1_vgg_layers_to_keep]: \n\t\t\tlayer.trainable=False\n\t\t# After VGG 9 layers, output is (128,128,256)\n\t\tconv1 = x\n\t\tpool1 = MaxPooling2D(pool_size=(2,2))(conv1)\n\n\t\tconv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-1')(pool1))\n\t\tconv2 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv2-2')(conv2))\n\t\tpool2 = MaxPooling2D(pool_size=(2, 2))(conv2)  #=(64,64,3)\n\n\t\tconv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-1')(pool2))\n\t\tconv3 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv3-2')(conv3))\n\t\tpool3 = MaxPooling2D(pool_size=(2, 2))(conv3)  # =(32,32,3)\n\n\t\tconv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-1')(pool3)\n\t\tconv4 = Conv2D(512, (3, 3), activation='relu', padding='same', name='conv4-2')(conv4)\n\n\t\tup1 = concatenate([UpSampling2D(size=(2, 2))(conv4), conv3], axis=3, name='conc1')\n\t\tconv5 = BatchNormalization()(Conv2D(512, (3,3), activation='relu', padding='same', name='conv5-1')(up1))\n\t\tconv5 = BatchNormalization()(Conv2D(512, (3, 3), activation='relu', padding='same', name='conv5-2')(conv5))\n\n\t\tup2 = concatenate([UpSampling2D(size=(2, 2))(conv5), conv2], axis=3, name='conc2')\n\t\tconv6 = BatchNormalization()(Conv2D(256, (3,3), activation='relu', padding='same', name='conv6-1')(up2))\n\t\tconv6 = BatchNormalization()(Conv2D(256, (3, 3), activation='relu', padding='same', name='conv6-2')(conv6))\n\n\t\tcls_conv_f = BatchNormalization()(\n\t\t\tConv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n\t\tcls_conv = Conv2D(1, (1, 1), activation='sigmoid', name='cls')(cls_conv_f)\n\n\t\toff_conv_f = BatchNormalization()(\n\t\t\tConv2D(1024, (1, 1), activation='relu', padding='same')(conv6))\n\t\toff_conv_f = Conv2D(2, (1, 1), activation='linear', kernel_initializer='zero', padding='same')\\\n\t\t\t\t\t(off_conv_f)\n\t\toff_conv = Concatenate(name='offset')([cls_conv, off_conv_f])\n\n\t\tmodel = Model(inputs=[input], outputs=[cls_conv, off_conv])\n\t\tmodel.compile(optimizer=Adam(lr=hp.p1_lr), loss={'cls':cls_loss, 'offset':offset_loss})\n\n\t\treturn model\n\n</pre>"
  },
  "source": "meta"
}