{
  "id": 29789,
  "title": "Basic UNet approach, good for small objects",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/writeups/alno-kostia-deepsystems-basic-unet-approach-good-f",
  "author_name": "",
  "post_date": "2017-03-08T20:36:24.833087500Z",
  "votes": 29,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Here is my approach which I used before teaming up with @alno and deepsystems.io. It gives ~0.476 public LB score, or ~0.495 if water from <a href=\"https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb\">this kernel</a> by Vladimir Osin is used. This score is without voting of several models, which can improve the score quite a bit for many classes. Here are the public LB scores by class:</p>\n\n<pre><code>0.07679 Buildings\n0.02179 Structures\n0.07970 Road\n0.02548 Track\n0.04688 Trees\n0.07776 Crops\n0.07540 Waterway\n0.03897 Standing Water\n0.02582 Vehicle Large\n0.00713 Vehicle Small\n0.47572 Total\n</code></pre>\n\n<p>So it's a far cry from 0.55698 on the public LB which we managed to achieve as a team - kudos to my teammates Alexey Noskov, Renat Bashirov and Ruslan Baikulov.</p>\n\n<h2>General approach</h2>\n\n<p>UNet network with batch-normalization added, training with Adam optimizer with\na loss that is a sum of 0.1 cross-entropy and 0.9 dice loss.\nInput for UNet was a 116 by 116 pixel patch, output was 64 by 64 pixels,\nso there were 16 additional pixels on each side that just provided context for\nthe prediction.\nBatch size was 128, learning rate was set to 0.0001\n(but loss was multiplied by the batch size).\nLearning rate was divided by 5 on the 25-th epoch\nand then again by 5 on the 50-th epoch,\nmost models were trained for 70-100 epochs.\nPatches that formed a batch were selected completely randomly across all images.\nDuring one epoch, network saw patches that covered about one half\nof the whole training set area. Best results for individual classes\nwere achieved when training on related classes, for example buildings\nand structures, roads and tracks, two kinds of vehicles.</p>\n\n<p>Augmentations included small rotations for some classes\n(±10-25 degrees for houses, structures and both vehicle classes),\nfull rotations and vertical/horizontal flips\nfor other classes. Small amount of dropout (0.1) was used in some cases.\nAlignment between channels was fixed with the help of\n<code>cv2.findTransformECC</code>, and lower-resolution layers were upscaled to\nmatch RGB size. In most cases, 12 channels were used (RGB, P, M),\nwhile in some cases just RGB and P or all 20 channels made results\nslightly better.</p>\n\n<h2>Validation</h2>\n\n<p>Validation was very hard, especially for both water and both vehicle\nclasses. In most cases, validation was performed on 5 images\n(6140_3_1, 6110_1_2, 6160_2_1, 6170_0_4, 6100_2_2), while other 20 were used\nfor training. Re-training the model with the same parameters on all 25 images\nimproved LB score.</p>\n\n<h2>Some details</h2>\n\n<ul>\n<li>This setup provides good results for small-scale classes\n(houses, structures, small vehicles), reasonable\nresults for most other classes and overfits quite badly on waterway.</li>\n<li>Man-made structures performed significantly better if training polygons\nwere made bigger by 0.5 pixel before producing training masks.</li>\n<li>For some classes (e.g. vehicles), it helped a bit to make the first\ndownscaling in UNet 4x instead of default 2x,\nand also made training 1.5x faster.</li>\n<li>Averaging of predictions (of one model)\nwith small shifts (1/3 of the 64 pixel step) were used\nfor some classes.</li>\n<li>Predictions on the edges of the input image (closer than 16 pixels to the\nborder) were bad for some classes and were left empty in this case.</li>\n<li>All models were implemented in pytorch, training for 70 epochs took about\n5 hours, submission generation took about 30 minutes without averaging,\nor about 5 hours with averaging.</li>\n</ul>\n\n<h2>Other things tried</h2>\n\n<p>A lot of things that either did not bring noticeable improvements,\nor made things worse:</p>\n\n<ul>\n<li>Losses: jaccard instead of dice, trying to predict distance to the border\nof the objects.</li>\n<li>Color augmentations.</li>\n<li>Oversampling of rare classes.</li>\n<li>Passing lower-resolution channels directly to lower-resolution layers in UNet.</li>\n<li>Varying UNet filter sizes, activations, number of layers and upscale/downscale\nsteps, using deconvolutions instead of upsampling.</li>\n<li>Learning rate decay.</li>\n<li>Models: VGG-like modules for UNet, SegNet, DenseNet</li>\n</ul>",
  "messages": [
    {
      "id": "166227",
      "postDate": "03/08/2017 20:36:24",
      "content": "<p>Here is my approach which I used before teaming up with @alno and deepsystems.io. It gives ~0.476 public LB score, or ~0.495 if water from <a href=\"https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb\">this kernel</a> by Vladimir Osin is used. This score is without voting of several models, which can improve the score quite a bit for many classes. Here are the public LB scores by class:</p>\n\n<pre><code>0.07679 Buildings\n0.02179 Structures\n0.07970 Road\n0.02548 Track\n0.04688 Trees\n0.07776 Crops\n0.07540 Waterway\n0.03897 Standing Water\n0.02582 Vehicle Large\n0.00713 Vehicle Small\n0.47572 Total\n</code></pre>\n\n<p>So it's a far cry from 0.55698 on the public LB which we managed to achieve as a team - kudos to my teammates Alexey Noskov, Renat Bashirov and Ruslan Baikulov.</p>\n\n<h2>General approach</h2>\n\n<p>UNet network with batch-normalization added, training with Adam optimizer with\na loss that is a sum of 0.1 cross-entropy and 0.9 dice loss.\nInput for UNet was a 116 by 116 pixel patch, output was 64 by 64 pixels,\nso there were 16 additional pixels on each side that just provided context for\nthe prediction.\nBatch size was 128, learning rate was set to 0.0001\n(but loss was multiplied by the batch size).\nLearning rate was divided by 5 on the 25-th epoch\nand then again by 5 on the 50-th epoch,\nmost models were trained for 70-100 epochs.\nPatches that formed a batch were selected completely randomly across all images.\nDuring one epoch, network saw patches that covered about one half\nof the whole training set area. Best results for individual classes\nwere achieved when training on related classes, for example buildings\nand structures, roads and tracks, two kinds of vehicles.</p>\n\n<p>Augmentations included small rotations for some classes\n(±10-25 degrees for houses, structures and both vehicle classes),\nfull rotations and vertical/horizontal flips\nfor other classes. Small amount of dropout (0.1) was used in some cases.\nAlignment between channels was fixed with the help of\n<code>cv2.findTransformECC</code>, and lower-resolution layers were upscaled to\nmatch RGB size. In most cases, 12 channels were used (RGB, P, M),\nwhile in some cases just RGB and P or all 20 channels made results\nslightly better.</p>\n\n<h2>Validation</h2>\n\n<p>Validation was very hard, especially for both water and both vehicle\nclasses. In most cases, validation was performed on 5 images\n(6140_3_1, 6110_1_2, 6160_2_1, 6170_0_4, 6100_2_2), while other 20 were used\nfor training. Re-training the model with the same parameters on all 25 images\nimproved LB score.</p>\n\n<h2>Some details</h2>\n\n<ul>\n<li>This setup provides good results for small-scale classes\n(houses, structures, small vehicles), reasonable\nresults for most other classes and overfits quite badly on waterway.</li>\n<li>Man-made structures performed significantly better if training polygons\nwere made bigger by 0.5 pixel before producing training masks.</li>\n<li>For some classes (e.g. vehicles), it helped a bit to make the first\ndownscaling in UNet 4x instead of default 2x,\nand also made training 1.5x faster.</li>\n<li>Averaging of predictions (of one model)\nwith small shifts (1/3 of the 64 pixel step) were used\nfor some classes.</li>\n<li>Predictions on the edges of the input image (closer than 16 pixels to the\nborder) were bad for some classes and were left empty in this case.</li>\n<li>All models were implemented in pytorch, training for 70 epochs took about\n5 hours, submission generation took about 30 minutes without averaging,\nor about 5 hours with averaging.</li>\n</ul>\n\n<h2>Other things tried</h2>\n\n<p>A lot of things that either did not bring noticeable improvements,\nor made things worse:</p>\n\n<ul>\n<li>Losses: jaccard instead of dice, trying to predict distance to the border\nof the objects.</li>\n<li>Color augmentations.</li>\n<li>Oversampling of rare classes.</li>\n<li>Passing lower-resolution channels directly to lower-resolution layers in UNet.</li>\n<li>Varying UNet filter sizes, activations, number of layers and upscale/downscale\nsteps, using deconvolutions instead of upsampling.</li>\n<li>Learning rate decay.</li>\n<li>Models: VGG-like modules for UNet, SegNet, DenseNet</li>\n</ul>",
      "rawMarkdown": "Here is my approach which I used before teaming up with @alno and deepsystems.io. It gives ~0.476 public LB score, or ~0.495 if water from [this kernel](https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb) by Vladimir Osin is used. This score is without voting of several models, which can improve the score quite a bit for many classes. Here are the public LB scores by class:\n \n    0.07679 Buildings\n    0.02179 Structures\n    0.07970 Road\n    0.02548 Track\n    0.04688 Trees\n    0.07776 Crops\n    0.07540 Waterway\n    0.03897 Standing Water\n    0.02582 Vehicle Large\n    0.00713 Vehicle Small\n    0.47572 Total\n\nSo it's a far cry from 0.55698 on the public LB which we managed to achieve as a team - kudos to my teammates Alexey Noskov, Renat Bashirov and Ruslan Baikulov.\n\nGeneral approach\n----------------\n\nUNet network with batch-normalization added, training with Adam optimizer with\na loss that is a sum of 0.1 cross-entropy and 0.9 dice loss.\nInput for UNet was a 116 by 116 pixel patch, output was 64 by 64 pixels,\nso there were 16 additional pixels on each side that just provided context for\nthe prediction.\nBatch size was 128, learning rate was set to 0.0001\n(but loss was multiplied by the batch size).\nLearning rate was divided by 5 on the 25-th epoch\nand then again by 5 on the 50-th epoch,\nmost models were trained for 70-100 epochs.\nPatches that formed a batch were selected completely randomly across all images.\nDuring one epoch, network saw patches that covered about one half\nof the whole training set area. Best results for individual classes\nwere achieved when training on related classes, for example buildings\nand structures, roads and tracks, two kinds of vehicles.\n\nAugmentations included small rotations for some classes\n(±10-25 degrees for houses, structures and both vehicle classes),\nfull rotations and vertical/horizontal flips\nfor other classes. Small amount of dropout (0.1) was used in some cases.\nAlignment between channels was fixed with the help of\n``cv2.findTransformECC``, and lower-resolution layers were upscaled to\nmatch RGB size. In most cases, 12 channels were used (RGB, P, M),\nwhile in some cases just RGB and P or all 20 channels made results\nslightly better.\n\n\nValidation\n----------\n\nValidation was very hard, especially for both water and both vehicle\nclasses. In most cases, validation was performed on 5 images\n(6140_3_1, 6110_1_2, 6160_2_1, 6170_0_4, 6100_2_2), while other 20 were used\nfor training. Re-training the model with the same parameters on all 25 images\nimproved LB score.\n\n\nSome details\n------------\n\n* This setup provides good results for small-scale classes\n  (houses, structures, small vehicles), reasonable\n  results for most other classes and overfits quite badly on waterway.\n* Man-made structures performed significantly better if training polygons\n  were made bigger by 0.5 pixel before producing training masks.\n* For some classes (e.g. vehicles), it helped a bit to make the first\n  downscaling in UNet 4x instead of default 2x,\n  and also made training 1.5x faster.\n* Averaging of predictions (of one model)\n  with small shifts (1/3 of the 64 pixel step) were used\n  for some classes.\n* Predictions on the edges of the input image (closer than 16 pixels to the\n  border) were bad for some classes and were left empty in this case.\n* All models were implemented in pytorch, training for 70 epochs took about\n  5 hours, submission generation took about 30 minutes without averaging,\n  or about 5 hours with averaging.\n\n\nOther things tried\n------------------\n\nA lot of things that either did not bring noticeable improvements,\nor made things worse:\n\n* Losses: jaccard instead of dice, trying to predict distance to the border\n  of the objects.\n* Color augmentations.\n* Oversampling of rare classes.\n* Passing lower-resolution channels directly to lower-resolution layers in UNet.\n* Varying UNet filter sizes, activations, number of layers and upscale/downscale\n  steps, using deconvolutions instead of upsampling.\n* Learning rate decay.\n* Models: VGG-like modules for UNet, SegNet, DenseNet",
      "votes": null
    },
    {
      "id": "166384",
      "postDate": "03/09/2017 12:19:10",
      "content": "<p>Can you clarify how do you implement the ignored border of the patches in the UNet? Maybe you just ignore the border when predicting  or is also the border of the mask ignored somehow during training, in that case some example how to do it would be nice.</p>",
      "rawMarkdown": "Can you clarify how do you implement the ignored border of the patches in the UNet? Maybe you just ignore the border when predicting  or is also the border of the mask ignored somehow during training, in that case some example how to do it would be nice.",
      "votes": null
    },
    {
      "id": "166385",
      "postDate": "03/09/2017 12:24:31",
      "content": "<p>I ignored border both during training and during predicting, throwing it away in the output layer. Here is an example (it's much simpler than UNet but the idea is exactly the same), <code>x[:, :, b:-b, b:-b]</code> does that:</p>\n\n<pre>class MiniNet(BaseNet):\n    def __init__(self, hps):\n        super().__init__(hps)\n        self.conv1 = nn.Conv2d(hps.n_channels, 4, 1)\n        self.conv2 = nn.Conv2d(4, 8, 3, padding=1)\n        self.conv3 = nn.Conv2d(8, hps.n_classes, 3, padding=1)\n\n    def forward(self, x):\n        x = F.relu(self.conv1(x))\n        x = F.relu(self.conv2(x))\n        x = self.conv3(x)\n        b = self.hps.patch_border\n        return F.sigmoid(x[:, :, b:-b, b:-b])\n</pre>",
      "rawMarkdown": "I ignored border both during training and during predicting, throwing it away in the output layer. Here is an example (it's much simpler than UNet but the idea is exactly the same), ``x[:, :, b:-b, b:-b]`` does that:\n\n<pre>\nclass MiniNet(BaseNet):\n    def __init__(self, hps):\n        super().__init__(hps)\n        self.conv1 = nn.Conv2d(hps.n_channels, 4, 1)\n        self.conv2 = nn.Conv2d(4, 8, 3, padding=1)\n        self.conv3 = nn.Conv2d(8, hps.n_classes, 3, padding=1)\n\n    def forward(self, x):\n        x = F.relu(self.conv1(x))\n        x = F.relu(self.conv2(x))\n        x = self.conv3(x)\n        b = self.hps.patch_border\n        return F.sigmoid(x[:, :, b:-b, b:-b])\n</pre>",
      "votes": null
    },
    {
      "id": "166405",
      "postDate": "03/09/2017 14:48:19",
      "content": "<p>I used similar approach, for me, the last layers in the Unet represented in Keras looked like:</p>\n\n<pre><code>conv9 = Convolution2D(32, 3, 3, border_mode='same', init='he_uniform')(conv9)\ncrop9 = Cropping2D(cropping=((16, 16), (16, 16)))(conv9)\nconv9 = BatchNormalization(mode=0, axis=1)(crop9)\nconv9 = keras.layers.advanced_activations.ELU()(conv9)\nconv10 = Convolution2D(num_mask_channels, 1, 1, activation='sigmoid')(conv9)\n\nmodel = Model(input=inputs, output=conv10)\n</code></pre>\n\n<p>where 16 is anumber of pixels that I cut on each side, thus it was used both in train and in test time.</p>",
      "rawMarkdown": "I used similar approach, for me, the last layers in the Unet represented in Keras looked like:\n\n    conv9 = Convolution2D(32, 3, 3, border_mode='same', init='he_uniform')(conv9)\n    crop9 = Cropping2D(cropping=((16, 16), (16, 16)))(conv9)\n    conv9 = BatchNormalization(mode=0, axis=1)(crop9)\n    conv9 = keras.layers.advanced_activations.ELU()(conv9)\n    conv10 = Convolution2D(num_mask_channels, 1, 1, activation='sigmoid')(conv9)\n\n    model = Model(input=inputs, output=conv10)\n\n\nwhere 16 is anumber of pixels that I cut on each side, thus it was used both in train and in test time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 166384,
      "author_name": "cubbus",
      "author_url": "",
      "post_date": "03/09/2017 12:19:10",
      "content": "<p>Can you clarify how do you implement the ignored border of the patches in the UNet? Maybe you just ignore the border when predicting  or is also the border of the mask ignored somehow during training, in that case some example how to do it would be nice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 166385,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "03/09/2017 12:24:31",
          "content": "<p>I ignored border both during training and during predicting, throwing it away in the output layer. Here is an example (it's much simpler than UNet but the idea is exactly the same), <code>x[:, :, b:-b, b:-b]</code> does that:</p>\n\n<pre>class MiniNet(BaseNet):\n    def __init__(self, hps):\n        super().__init__(hps)\n        self.conv1 = nn.Conv2d(hps.n_channels, 4, 1)\n        self.conv2 = nn.Conv2d(4, 8, 3, padding=1)\n        self.conv3 = nn.Conv2d(8, hps.n_classes, 3, padding=1)\n\n    def forward(self, x):\n        x = F.relu(self.conv1(x))\n        x = F.relu(self.conv2(x))\n        x = self.conv3(x)\n        b = self.hps.patch_border\n        return F.sigmoid(x[:, :, b:-b, b:-b])\n</pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 166405,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "03/09/2017 14:48:19",
          "content": "<p>I used similar approach, for me, the last layers in the Unet represented in Keras looked like:</p>\n\n<pre><code>conv9 = Convolution2D(32, 3, 3, border_mode='same', init='he_uniform')(conv9)\ncrop9 = Cropping2D(cropping=((16, 16), (16, 16)))(conv9)\nconv9 = BatchNormalization(mode=0, axis=1)(crop9)\nconv9 = keras.layers.advanced_activations.ELU()(conv9)\nconv10 = Convolution2D(num_mask_channels, 1, 1, activation='sigmoid')(conv9)\n\nmodel = Model(input=inputs, output=conv10)\n</code></pre>\n\n<p>where 16 is anumber of pixels that I cut on each side, thus it was used both in train and in test time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "166227": "Here is my approach which I used before teaming up with @alno and deepsystems.io. It gives ~0.476 public LB score, or ~0.495 if water from [this kernel](https://www.kaggle.com/resolut/dstl-satellite-imagery-feature-detection/waterway-0-095-lb) by Vladimir Osin is used. This score is without voting of several models, which can improve the score quite a bit for many classes. Here are the public LB scores by class:\n \n    0.07679 Buildings\n    0.02179 Structures\n    0.07970 Road\n    0.02548 Track\n    0.04688 Trees\n    0.07776 Crops\n    0.07540 Waterway\n    0.03897 Standing Water\n    0.02582 Vehicle Large\n    0.00713 Vehicle Small\n    0.47572 Total\n\nSo it's a far cry from 0.55698 on the public LB which we managed to achieve as a team - kudos to my teammates Alexey Noskov, Renat Bashirov and Ruslan Baikulov.\n\nGeneral approach\n----------------\n\nUNet network with batch-normalization added, training with Adam optimizer with\na loss that is a sum of 0.1 cross-entropy and 0.9 dice loss.\nInput for UNet was a 116 by 116 pixel patch, output was 64 by 64 pixels,\nso there were 16 additional pixels on each side that just provided context for\nthe prediction.\nBatch size was 128, learning rate was set to 0.0001\n(but loss was multiplied by the batch size).\nLearning rate was divided by 5 on the 25-th epoch\nand then again by 5 on the 50-th epoch,\nmost models were trained for 70-100 epochs.\nPatches that formed a batch were selected completely randomly across all images.\nDuring one epoch, network saw patches that covered about one half\nof the whole training set area. Best results for individual classes\nwere achieved when training on related classes, for example buildings\nand structures, roads and tracks, two kinds of vehicles.\n\nAugmentations included small rotations for some classes\n(±10-25 degrees for houses, structures and both vehicle classes),\nfull rotations and vertical/horizontal flips\nfor other classes. Small amount of dropout (0.1) was used in some cases.\nAlignment between channels was fixed with the help of\n``cv2.findTransformECC``, and lower-resolution layers were upscaled to\nmatch RGB size. In most cases, 12 channels were used (RGB, P, M),\nwhile in some cases just RGB and P or all 20 channels made results\nslightly better.\n\n\nValidation\n----------\n\nValidation was very hard, especially for both water and both vehicle\nclasses. In most cases, validation was performed on 5 images\n(6140_3_1, 6110_1_2, 6160_2_1, 6170_0_4, 6100_2_2), while other 20 were used\nfor training. Re-training the model with the same parameters on all 25 images\nimproved LB score.\n\n\nSome details\n------------\n\n* This setup provides good results for small-scale classes\n  (houses, structures, small vehicles), reasonable\n  results for most other classes and overfits quite badly on waterway.\n* Man-made structures performed significantly better if training polygons\n  were made bigger by 0.5 pixel before producing training masks.\n* For some classes (e.g. vehicles), it helped a bit to make the first\n  downscaling in UNet 4x instead of default 2x,\n  and also made training 1.5x faster.\n* Averaging of predictions (of one model)\n  with small shifts (1/3 of the 64 pixel step) were used\n  for some classes.\n* Predictions on the edges of the input image (closer than 16 pixels to the\n  border) were bad for some classes and were left empty in this case.\n* All models were implemented in pytorch, training for 70 epochs took about\n  5 hours, submission generation took about 30 minutes without averaging,\n  or about 5 hours with averaging.\n\n\nOther things tried\n------------------\n\nA lot of things that either did not bring noticeable improvements,\nor made things worse:\n\n* Losses: jaccard instead of dice, trying to predict distance to the border\n  of the objects.\n* Color augmentations.\n* Oversampling of rare classes.\n* Passing lower-resolution channels directly to lower-resolution layers in UNet.\n* Varying UNet filter sizes, activations, number of layers and upscale/downscale\n  steps, using deconvolutions instead of upsampling.\n* Learning rate decay.\n* Models: VGG-like modules for UNet, SegNet, DenseNet",
    "166384": "Can you clarify how do you implement the ignored border of the patches in the UNet? Maybe you just ignore the border when predicting  or is also the border of the mask ignored somehow during training, in that case some example how to do it would be nice.",
    "166385": "I ignored border both during training and during predicting, throwing it away in the output layer. Here is an example (it's much simpler than UNet but the idea is exactly the same), ``x[:, :, b:-b, b:-b]`` does that:\n\n<pre>\nclass MiniNet(BaseNet):\n    def __init__(self, hps):\n        super().__init__(hps)\n        self.conv1 = nn.Conv2d(hps.n_channels, 4, 1)\n        self.conv2 = nn.Conv2d(4, 8, 3, padding=1)\n        self.conv3 = nn.Conv2d(8, hps.n_classes, 3, padding=1)\n\n    def forward(self, x):\n        x = F.relu(self.conv1(x))\n        x = F.relu(self.conv2(x))\n        x = self.conv3(x)\n        b = self.hps.patch_border\n        return F.sigmoid(x[:, :, b:-b, b:-b])\n</pre>",
    "166405": "I used similar approach, for me, the last layers in the Unet represented in Keras looked like:\n\n    conv9 = Convolution2D(32, 3, 3, border_mode='same', init='he_uniform')(conv9)\n    crop9 = Cropping2D(cropping=((16, 16), (16, 16)))(conv9)\n    conv9 = BatchNormalization(mode=0, axis=1)(crop9)\n    conv9 = keras.layers.advanced_activations.ELU()(conv9)\n    conv10 = Convolution2D(num_mask_channels, 1, 1, activation='sigmoid')(conv9)\n\n    model = Model(input=inputs, output=conv10)\n\n\nwhere 16 is anumber of pixels that I cut on each side, thus it was used both in train and in test time."
  },
  "source": "meta"
}