{
  "id": 75984,
  "title": "[LB 0.337] Siamese network prototype using fastai",
  "url": "/competitions/humpback-whale-identification/discussion/75984",
  "author_name": "Radek Osmulski",
  "post_date": "2018-12-28T08:32:31.695000",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I added a <a href=\"https://github.com/radekosmulski/whale/blob/master/siamese_network_prototype.ipynb\">siamese network prototype</a> to my <a href=\"https://github.com/radekosmulski/whale\">starter pack repository</a>. I only train it for around 20 minutes and have not experimented with it much.</p>\n\n<p>Here are a couple of ideas that might be useful for improving the result:</p>\n\n<ul>\n<li>training for longer</li>\n<li>use a bigger cnn</li>\n<li>use a more complex head, potentially one created using the fastai <code>create_head</code> function</li>\n<li>due to the way I create the validation set, the task ends up being a zero shot learning for 1285 whales, and the model generalizes poorly to whales it has not seen - could creating a validation set in a different way be helpful here?</li>\n<li>is there another way one could add the new_whale class that could work better?</li>\n<li>sampling pairs in a different manner (for instance, to present the model with progressively harder pairs to train on)</li>\n<li>incorporating whales of the new_whale class into training</li>\n<li>experiment with different ways of processing features in <code>process_features</code>, is global max pooling the best approach to take here?</li>\n<li>taking inspiration from <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">the great write up of the winning solution to the playground challenge</a> and porting selected ideas</li>\n</ul>\n\n<p>I share this as is for two reasons. First of all, I don't think I will have a lot of time to work on this in near future. Secondly, it is fun to discover things on your own. I realize the code is messy but it shouldn't take that long if someone would want to go through each line and understand what it does. I feel there is quite a lot of potential in this approach so playing with this further might be quite rewarding.</p>\n\n<p>As a side note, holding other things constant I trained the model for a little bit longer with resnet50 as the base cnn and my score on the LB increased to 0.461. That is with not doing anything to how the new whale class is predicted nor to the makeup of the train set.</p>",
  "messages": [
    {
      "id": 446522,
      "postDate": "2018-12-28T08:32:31.697Z",
      "content": "<p>I added a <a href=\"https://github.com/radekosmulski/whale/blob/master/siamese_network_prototype.ipynb\">siamese network prototype</a> to my <a href=\"https://github.com/radekosmulski/whale\">starter pack repository</a>. I only train it for around 20 minutes and have not experimented with it much.</p>\n\n<p>Here are a couple of ideas that might be useful for improving the result:</p>\n\n<ul>\n<li>training for longer</li>\n<li>use a bigger cnn</li>\n<li>use a more complex head, potentially one created using the fastai <code>create_head</code> function</li>\n<li>due to the way I create the validation set, the task ends up being a zero shot learning for 1285 whales, and the model generalizes poorly to whales it has not seen - could creating a validation set in a different way be helpful here?</li>\n<li>is there another way one could add the new_whale class that could work better?</li>\n<li>sampling pairs in a different manner (for instance, to present the model with progressively harder pairs to train on)</li>\n<li>incorporating whales of the new_whale class into training</li>\n<li>experiment with different ways of processing features in <code>process_features</code>, is global max pooling the best approach to take here?</li>\n<li>taking inspiration from <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">the great write up of the winning solution to the playground challenge</a> and porting selected ideas</li>\n</ul>\n\n<p>I share this as is for two reasons. First of all, I don't think I will have a lot of time to work on this in near future. Secondly, it is fun to discover things on your own. I realize the code is messy but it shouldn't take that long if someone would want to go through each line and understand what it does. I feel there is quite a lot of potential in this approach so playing with this further might be quite rewarding.</p>\n\n<p>As a side note, holding other things constant I trained the model for a little bit longer with resnet50 as the base cnn and my score on the LB increased to 0.461. That is with not doing anything to how the new whale class is predicted nor to the makeup of the train set.</p>",
      "rawMarkdown": "I added a [siamese network prototype](https://github.com/radekosmulski/whale/blob/master/siamese_network_prototype.ipynb) to my [starter pack repository](https://github.com/radekosmulski/whale). I only train it for around 20 minutes and have not experimented with it much.\n\nHere are a couple of ideas that might be useful for improving the result:\n\n- training for longer\n- use a bigger cnn\n- use a more complex head, potentially one created using the fastai `create_head` function\n- due to the way I create the validation set, the task ends up being a zero shot learning for 1285 whales, and the model generalizes poorly to whales it has not seen - could creating a validation set in a different way be helpful here?\n- is there another way one could add the new_whale class that could work better?\n- sampling pairs in a different manner (for instance, to present the model with progressively harder pairs to train on)\n- incorporating whales of the new_whale class into training\n- experiment with different ways of processing features in `process_features`, is global max pooling the best approach to take here?\n- taking inspiration from [the great write up of the winning solution to the playground challenge](https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563) and porting selected ideas\n\nI share this as is for two reasons. First of all, I don't think I will have a lot of time to work on this in near future. Secondly, it is fun to discover things on your own. I realize the code is messy but it shouldn't take that long if someone would want to go through each line and understand what it does. I feel there is quite a lot of potential in this approach so playing with this further might be quite rewarding.\n\nAs a side note, holding other things constant I trained the model for a little bit longer with resnet50 as the base cnn and my score on the LB increased to 0.461. That is with not doing anything to how the new whale class is predicted nor to the makeup of the train set.",
      "votes": 16
    },
    {
      "id": 447483,
      "postDate": "2018-12-29T23:03:38.690Z",
      "content": "<p>Awesome, was looking for this!</p>",
      "rawMarkdown": "Awesome, was looking for this!",
      "votes": 1
    },
    {
      "id": 446665,
      "postDate": "2018-12-28T13:02:51.127Z",
      "content": "<p>Thank you very much radek! I really appreciate your work.</p>",
      "rawMarkdown": "Thank you very much radek! I really appreciate your work.",
      "votes": 1
    },
    {
      "id": 446558,
      "postDate": "2018-12-28T09:26:13.817Z",
      "content": "<p>Thanks for sharing! I was thinking about the same thing.</p>",
      "rawMarkdown": "Thanks for sharing! I was thinking about the same thing.",
      "votes": 1
    },
    {
      "id": 461361,
      "postDate": "2019-01-25T20:53:02.483Z",
      "content": "<p>Thanks Radek for sharing the amazing work.\nThe best I could get with your implementation is 0.49 in LB. That was by OS the training set into 15 samples, so that no single images will be excluded from training. I thought that maybe relying on transformation will eventually make it overall better  than discarding a large portion of single whales from training. And it helped to increase the ranking by ~0.1. I tried also to increase the SZ into 448 and use only 10% of the validation set (~257 images) which helped a bit more. The threshold of new whale became more reasonable too (~0.933) and NN found only ~30-40% of the test set as New whales which made me thinking, perhaps there is no bug in your code after all.</p>\n\n<p>Now I think that to get a substantial improvement with your code the sampling pairs method be modified in a different manner (presenting the model with progressively harder pairs to train on). This is the secret sauce that Martin mentioned in his notebook. Quoting him:</p>\n\n<p><em>Pairs of images of different whales are selected to be difficult for the network to distinguish at a given stage of the training. This is inspired from adversarial training: find pairs of images that are from different whales, but that are still very similar from the model perspective.\nImplementing this strategy while training a Siamese Neural Network is what makes the largest contribution to the model accuracy. Other details contribute somewhat to the accuracy, but have a much smaller impact.</em></p>\n\n<p>I found Andrew Ng too (emphasizing this)[ <a href=\"https://www.coursera.org/lecture/convolutional-neural-networks/triplet-loss-HuUtN\">https://www.coursera.org/lecture/convolutional-neural-networks/triplet-loss-HuUtN</a> ] in his DL course (9:00 to 11:00). </p>\n\n<p>This is the transcript of this two minute video segment:</p>\n\n<p>Now, how do you actually choose these triplets to form your training set? One of the problems if you choose A, P, and N randomly from your training set subject to A and P being from the same person, and A and N being different persons, one of the problems is that if you choose them so that they're at random, then this constraint is very easy to satisfy. Because given two randomly chosen pictures of people, chances are A and N are much different than A and P. I hope you still recognize this notation, this d(A, P) was what we had written on the last year's slides as this encoding. So this is just equal to this squared known distance between the encodings that we have on the previous slide. But if A and N are two randomly chosen different persons, then there is a very high chance that this will be much bigger more than the margin alpha that that term on the left. And so, the neural network won't learn much from it. So to construct a training set, what you want to do is to choose triplets A, P, and N that are hard to train on. So in particular, what you want is for all triplets that this constraint be satisfied. So, a triplet that is hard will be if you choose values for A, P, and N so that maybe d(A, P) is actually quite close to d(A,N). So in that case, the learning algorithm has to try extra hard to take this thing on the right and try to push it up or take this thing on the left and try to push it down so that there is at least a margin of alpha between the left side and the right side. And the effect of choosing these triplets is that it increases the computational efficiency of your learning algorithm. If you choose your triplets randomly, then too many triplets would be really easy, and so, gradient descent won't do anything because your neural network will just get them right, pretty much all the time. And it's only by using hard triplets that the gradient descent procedure has to do some work to try to push these quantities further away from those quantities. And if you're interested, the details are presented in this paper by Florian Schroff, Dmitry Kalinichenko, and James Philbin, where they have a system called FaceNet,which is where a lot of the ideas I'm presenting in this video come from. </p>",
      "rawMarkdown": "Thanks Radek for sharing the amazing work.\nThe best I could get with your implementation is 0.49 in LB. That was by OS the training set into 15 samples, so that no single images will be excluded from training. I thought that maybe relying on transformation will eventually make it overall better  than discarding a large portion of single whales from training. And it helped to increase the ranking by ~0.1. I tried also to increase the SZ into 448 and use only 10% of the validation set (~257 images) which helped a bit more. The threshold of new whale became more reasonable too (~0.933) and NN found only ~30-40% of the test set as New whales which made me thinking, perhaps there is no bug in your code after all.\n\nNow I think that to get a substantial improvement with your code the sampling pairs method be modified in a different manner (presenting the model with progressively harder pairs to train on). This is the secret sauce that Martin mentioned in his notebook. Quoting him:\n\n*Pairs of images of different whales are selected to be difficult for the network to distinguish at a given stage of the training. This is inspired from adversarial training: find pairs of images that are from different whales, but that are still very similar from the model perspective.\nImplementing this strategy while training a Siamese Neural Network is what makes the largest contribution to the model accuracy. Other details contribute somewhat to the accuracy, but have a much smaller impact.*\n\nI found Andrew Ng too (emphasizing this)[ https://www.coursera.org/lecture/convolutional-neural-networks/triplet-loss-HuUtN ] in his DL course (9:00 to 11:00). \n\n\nThis is the transcript of this two minute video segment:\n\nNow, how do you actually choose these triplets to form your training set? One of the problems if you choose A, P, and N randomly from your training set subject to A and P being from the same person, and A and N being different persons, one of the problems is that if you choose them so that they're at random, then this constraint is very easy to satisfy. Because given two randomly chosen pictures of people, chances are A and N are much different than A and P. I hope you still recognize this notation, this d(A, P) was what we had written on the last year's slides as this encoding. So this is just equal to this squared known distance between the encodings that we have on the previous slide. But if A and N are two randomly chosen different persons, then there is a very high chance that this will be much bigger more than the margin alpha that that term on the left. And so, the neural network won't learn much from it. So to construct a training set, what you want to do is to choose triplets A, P, and N that are hard to train on. So in particular, what you want is for all triplets that this constraint be satisfied. So, a triplet that is hard will be if you choose values for A, P, and N so that maybe d(A, P) is actually quite close to d(A,N). So in that case, the learning algorithm has to try extra hard to take this thing on the right and try to push it up or take this thing on the left and try to push it down so that there is at least a margin of alpha between the left side and the right side. And the effect of choosing these triplets is that it increases the computational efficiency of your learning algorithm. If you choose your triplets randomly, then too many triplets would be really easy, and so, gradient descent won't do anything because your neural network will just get them right, pretty much all the time. And it's only by using hard triplets that the gradient descent procedure has to do some work to try to push these quantities further away from those quantities. And if you're interested, the details are presented in this paper by Florian Schroff, Dmitry Kalinichenko, and James Philbin, where they have a system called FaceNet,which is where a lot of the ideas I'm presenting in this video come from. ",
      "votes": 2
    },
    {
      "id": 447680,
      "postDate": "2018-12-30T10:16:37.490Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 447714,
          "postDate": "2018-12-30T11:43:46.220Z",
          "content": "<p>Hi Abhishek,</p>\n\n<p>Thank you very much for raising this. Yes, I think you are right. I pushed a new version where this should now be corrected.</p>\n\n<p>As is, the results change only very slightly... that is because the model is very poor in identifying similar images, even if comparing against self! (it will usually find many other images it claims are more similar to any given image than that image itself). </p>",
          "rawMarkdown": "Hi Abhishek,\n\nThank you very much for raising this. Yes, I think you are right. I pushed a new version where this should now be corrected.\n\nAs is, the results change only very slightly... that is because the model is very poor in identifying similar images, even if comparing against self! (it will usually find many other images it claims are more similar to any given image than that image itself). "
        },
        {
          "id": 447738,
          "postDate": "2018-12-30T12:24:10.110Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 446707,
      "postDate": "2018-12-28T14:49:43.740Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 447483,
      "author_name": "Edwin Villanueva",
      "author_url": "",
      "post_date": "2018-12-29T23:03:38.690000",
      "content": "<p>Awesome, was looking for this!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 446665,
      "author_name": "Tommy Jiang",
      "author_url": "",
      "post_date": "2018-12-28T13:02:51.127000",
      "content": "<p>Thank you very much radek! I really appreciate your work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 446558,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2018-12-28T09:26:13.817000",
      "content": "<p>Thanks for sharing! I was thinking about the same thing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 461361,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-01-25T20:53:02.483000",
      "content": "<p>Thanks Radek for sharing the amazing work.\nThe best I could get with your implementation is 0.49 in LB. That was by OS the training set into 15 samples, so that no single images will be excluded from training. I thought that maybe relying on transformation will eventually make it overall better  than discarding a large portion of single whales from training. And it helped to increase the ranking by ~0.1. I tried also to increase the SZ into 448 and use only 10% of the validation set (~257 images) which helped a bit more. The threshold of new whale became more reasonable too (~0.933) and NN found only ~30-40% of the test set as New whales which made me thinking, perhaps there is no bug in your code after all.</p>\n\n<p>Now I think that to get a substantial improvement with your code the sampling pairs method be modified in a different manner (presenting the model with progressively harder pairs to train on). This is the secret sauce that Martin mentioned in his notebook. Quoting him:</p>\n\n<p><em>Pairs of images of different whales are selected to be difficult for the network to distinguish at a given stage of the training. This is inspired from adversarial training: find pairs of images that are from different whales, but that are still very similar from the model perspective.\nImplementing this strategy while training a Siamese Neural Network is what makes the largest contribution to the model accuracy. Other details contribute somewhat to the accuracy, but have a much smaller impact.</em></p>\n\n<p>I found Andrew Ng too (emphasizing this)[ <a href=\"https://www.coursera.org/lecture/convolutional-neural-networks/triplet-loss-HuUtN\">https://www.coursera.org/lecture/convolutional-neural-networks/triplet-loss-HuUtN</a> ] in his DL course (9:00 to 11:00). </p>\n\n<p>This is the transcript of this two minute video segment:</p>\n\n<p>Now, how do you actually choose these triplets to form your training set? One of the problems if you choose A, P, and N randomly from your training set subject to A and P being from the same person, and A and N being different persons, one of the problems is that if you choose them so that they're at random, then this constraint is very easy to satisfy. Because given two randomly chosen pictures of people, chances are A and N are much different than A and P. I hope you still recognize this notation, this d(A, P) was what we had written on the last year's slides as this encoding. So this is just equal to this squared known distance between the encodings that we have on the previous slide. But if A and N are two randomly chosen different persons, then there is a very high chance that this will be much bigger more than the margin alpha that that term on the left. And so, the neural network won't learn much from it. So to construct a training set, what you want to do is to choose triplets A, P, and N that are hard to train on. So in particular, what you want is for all triplets that this constraint be satisfied. So, a triplet that is hard will be if you choose values for A, P, and N so that maybe d(A, P) is actually quite close to d(A,N). So in that case, the learning algorithm has to try extra hard to take this thing on the right and try to push it up or take this thing on the left and try to push it down so that there is at least a margin of alpha between the left side and the right side. And the effect of choosing these triplets is that it increases the computational efficiency of your learning algorithm. If you choose your triplets randomly, then too many triplets would be really easy, and so, gradient descent won't do anything because your neural network will just get them right, pretty much all the time. And it's only by using hard triplets that the gradient descent procedure has to do some work to try to push these quantities further away from those quantities. And if you're interested, the details are presented in this paper by Florian Schroff, Dmitry Kalinichenko, and James Philbin, where they have a system called FaceNet,which is where a lot of the ideas I'm presenting in this video come from. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 447680,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-30T10:16:37.490000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 447714,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2018-12-30T11:43:46.220000",
          "content": "<p>Hi Abhishek,</p>\n\n<p>Thank you very much for raising this. Yes, I think you are right. I pushed a new version where this should now be corrected.</p>\n\n<p>As is, the results change only very slightly... that is because the model is very poor in identifying similar images, even if comparing against self! (it will usually find many other images it claims are more similar to any given image than that image itself). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 447738,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-30T12:24:10.110000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 446707,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T14:49:43.740000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "446522": "I added a [siamese network prototype](https://github.com/radekosmulski/whale/blob/master/siamese_network_prototype.ipynb) to my [starter pack repository](https://github.com/radekosmulski/whale). I only train it for around 20 minutes and have not experimented with it much.\n\nHere are a couple of ideas that might be useful for improving the result:\n\n- training for longer\n- use a bigger cnn\n- use a more complex head, potentially one created using the fastai `create_head` function\n- due to the way I create the validation set, the task ends up being a zero shot learning for 1285 whales, and the model generalizes poorly to whales it has not seen - could creating a validation set in a different way be helpful here?\n- is there another way one could add the new_whale class that could work better?\n- sampling pairs in a different manner (for instance, to present the model with progressively harder pairs to train on)\n- incorporating whales of the new_whale class into training\n- experiment with different ways of processing features in `process_features`, is global max pooling the best approach to take here?\n- taking inspiration from [the great write up of the winning solution to the playground challenge](https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563) and porting selected ideas\n\nI share this as is for two reasons. First of all, I don't think I will have a lot of time to work on this in near future. Secondly, it is fun to discover things on your own. I realize the code is messy but it shouldn't take that long if someone would want to go through each line and understand what it does. I feel there is quite a lot of potential in this approach so playing with this further might be quite rewarding.\n\nAs a side note, holding other things constant I trained the model for a little bit longer with resnet50 as the base cnn and my score on the LB increased to 0.461. That is with not doing anything to how the new whale class is predicted nor to the makeup of the train set.",
    "447483": "Awesome, was looking for this!",
    "446665": "Thank you very much radek! I really appreciate your work.",
    "446558": "Thanks for sharing! I was thinking about the same thing.",
    "461361": "Thanks Radek for sharing the amazing work.\nThe best I could get with your implementation is 0.49 in LB. That was by OS the training set into 15 samples, so that no single images will be excluded from training. I thought that maybe relying on transformation will eventually make it overall better  than discarding a large portion of single whales from training. And it helped to increase the ranking by ~0.1. I tried also to increase the SZ into 448 and use only 10% of the validation set (~257 images) which helped a bit more. The threshold of new whale became more reasonable too (~0.933) and NN found only ~30-40% of the test set as New whales which made me thinking, perhaps there is no bug in your code after all.\n\nNow I think that to get a substantial improvement with your code the sampling pairs method be modified in a different manner (presenting the model with progressively harder pairs to train on). This is the secret sauce that Martin mentioned in his notebook. Quoting him:\n\n*Pairs of images of different whales are selected to be difficult for the network to distinguish at a given stage of the training. This is inspired from adversarial training: find pairs of images that are from different whales, but that are still very similar from the model perspective.\nImplementing this strategy while training a Siamese Neural Network is what makes the largest contribution to the model accuracy. Other details contribute somewhat to the accuracy, but have a much smaller impact.*\n\nI found Andrew Ng too (emphasizing this)[ https://www.coursera.org/lecture/convolutional-neural-networks/triplet-loss-HuUtN ] in his DL course (9:00 to 11:00). \n\n\nThis is the transcript of this two minute video segment:\n\nNow, how do you actually choose these triplets to form your training set? One of the problems if you choose A, P, and N randomly from your training set subject to A and P being from the same person, and A and N being different persons, one of the problems is that if you choose them so that they're at random, then this constraint is very easy to satisfy. Because given two randomly chosen pictures of people, chances are A and N are much different than A and P. I hope you still recognize this notation, this d(A, P) was what we had written on the last year's slides as this encoding. So this is just equal to this squared known distance between the encodings that we have on the previous slide. But if A and N are two randomly chosen different persons, then there is a very high chance that this will be much bigger more than the margin alpha that that term on the left. And so, the neural network won't learn much from it. So to construct a training set, what you want to do is to choose triplets A, P, and N that are hard to train on. So in particular, what you want is for all triplets that this constraint be satisfied. So, a triplet that is hard will be if you choose values for A, P, and N so that maybe d(A, P) is actually quite close to d(A,N). So in that case, the learning algorithm has to try extra hard to take this thing on the right and try to push it up or take this thing on the left and try to push it down so that there is at least a margin of alpha between the left side and the right side. And the effect of choosing these triplets is that it increases the computational efficiency of your learning algorithm. If you choose your triplets randomly, then too many triplets would be really easy, and so, gradient descent won't do anything because your neural network will just get them right, pretty much all the time. And it's only by using hard triplets that the gradient descent procedure has to do some work to try to push these quantities further away from those quantities. And if you're interested, the details are presented in this paper by Florian Schroff, Dmitry Kalinichenko, and James Philbin, where they have a system called FaceNet,which is where a lot of the ideas I'm presenting in this video come from. ",
    "447680": "",
    "446707": ""
  }
}