{
  "id": 81085,
  "title": "Prototypical Networks as a Fine Grained Classifier",
  "url": "/competitions/humpback-whale-identification/discussion/81085",
  "author_name": "daisukelab",
  "post_date": "2019-02-19T03:08:08.627000",
  "votes": 54,
  "comment_count": 95,
  "views": 0,
  "content": "<p>Though competition is approaching to finish, so ... it's kind of late but let me share what I have been working on this couple of months.\nI'd like to introduce Prototypical Networks (ProtoNets), a metrics learning basically proposed for few-shot problems.\nIt can learn distances between classes as same as Siamese networks but between many classes.</p>\n\n<ul>\n<li>Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.\n<a href=\"https://arxiv.org/pdf/1703.05175.pdf\">https://arxiv.org/pdf/1703.05175.pdf</a></li>\n</ul>\n\n<p>My intuition was it could be better discriminator to learn differences between many class samples than to learn just between two classes, and it could be universal discriminator/classifier applicable to variety of problems. While conventional classifier is limited to tell differences between trained classes, metrics learning model is free to handle unknown classes, and measure the degree of deviation from normal = trained classes.\nThere are many few-shot learning models proposed recently, which you can find in this good article \"Advances in few-shot learning: a guided tour\" by Oscar Knagg:\n<a href=\"https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\">https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77</a></p>\n\n<p>Then I picked Prototypical Networks because solution is quite simple and practical. But basically it doesn't use pre-trained model in the paper though tested problem was image few-shot classification, and though we know that pre-trained model is essential for better image interpretation/ representation.\nMain contribution of my work is applying ImageNet models to ProtoNets.</p>\n\n<ul>\n<li><p>Repository:\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf\">https://github.com/daisukelab/protonet-fine-grained-clf</a></p></li>\n<li><p>Notebook is ready to reproduce public LB 0.748 with 100 epochs (takes <strike>an hour or so</strike> 4-5 hours):\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/Example_Humpback_Whale_Identification.ipynb\">https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/Example_Humpback_Whale_Identification.ipynb</a></p></li>\n</ul>\n\n<p>The best public LB score among my local models is 0.878, and ensemble of them is my score 0.907 as of now. Yes, nothing other than ProtoNets is used in my solution. I even don't tweak training/test samples so much; usual augmentation, normalization, TTA... these pushed my score totally.</p>\n\n<p>The reason why I'm so much interested in ProtoNets is, it is easy to be understood by non-ML engineers in real world projects.\nI think it is almost proved to work fine without big effort in this competition, then you can also try. :)</p>\n\n<hr>\n\n<p>Update Feb-21: Fixed training time \"an hour or so\" --&gt; \"4-5 hours\", excuse me to write old info...</p>\n\n<p>Update Feb-24: Bug fix - unneeded log() caused calculating wrong distances.</p>\n\n<p>Update Feb-26: Added option for data sampling &amp; Bug fix.\n- Fixed softmax was too hard, this was introduced by Feb-24 change. Fix for user of softmax result.\n- Added <code>get_training_datalists(sampling_type)</code> is added: <code>'exhaustive'</code> will use all training samples, <code>'more_than_two'</code> is original implementation.</p>\n\n<p>Update Mar-7: Added best single model that marked LB score: Private/Public = 0.88599/0.87523</p>",
  "messages": [
    {
      "id": 474174,
      "postDate": "2019-02-19T03:08:08.627Z",
      "content": "<p>Though competition is approaching to finish, so ... it's kind of late but let me share what I have been working on this couple of months.\nI'd like to introduce Prototypical Networks (ProtoNets), a metrics learning basically proposed for few-shot problems.\nIt can learn distances between classes as same as Siamese networks but between many classes.</p>\n\n<ul>\n<li>Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.\n<a href=\"https://arxiv.org/pdf/1703.05175.pdf\">https://arxiv.org/pdf/1703.05175.pdf</a></li>\n</ul>\n\n<p>My intuition was it could be better discriminator to learn differences between many class samples than to learn just between two classes, and it could be universal discriminator/classifier applicable to variety of problems. While conventional classifier is limited to tell differences between trained classes, metrics learning model is free to handle unknown classes, and measure the degree of deviation from normal = trained classes.\nThere are many few-shot learning models proposed recently, which you can find in this good article \"Advances in few-shot learning: a guided tour\" by Oscar Knagg:\n<a href=\"https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\">https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77</a></p>\n\n<p>Then I picked Prototypical Networks because solution is quite simple and practical. But basically it doesn't use pre-trained model in the paper though tested problem was image few-shot classification, and though we know that pre-trained model is essential for better image interpretation/ representation.\nMain contribution of my work is applying ImageNet models to ProtoNets.</p>\n\n<ul>\n<li><p>Repository:\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf\">https://github.com/daisukelab/protonet-fine-grained-clf</a></p></li>\n<li><p>Notebook is ready to reproduce public LB 0.748 with 100 epochs (takes <strike>an hour or so</strike> 4-5 hours):\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/Example_Humpback_Whale_Identification.ipynb\">https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/Example_Humpback_Whale_Identification.ipynb</a></p></li>\n</ul>\n\n<p>The best public LB score among my local models is 0.878, and ensemble of them is my score 0.907 as of now. Yes, nothing other than ProtoNets is used in my solution. I even don't tweak training/test samples so much; usual augmentation, normalization, TTA... these pushed my score totally.</p>\n\n<p>The reason why I'm so much interested in ProtoNets is, it is easy to be understood by non-ML engineers in real world projects.\nI think it is almost proved to work fine without big effort in this competition, then you can also try. :)</p>\n\n<hr>\n\n<p>Update Feb-21: Fixed training time \"an hour or so\" --&gt; \"4-5 hours\", excuse me to write old info...</p>\n\n<p>Update Feb-24: Bug fix - unneeded log() caused calculating wrong distances.</p>\n\n<p>Update Feb-26: Added option for data sampling &amp; Bug fix.\n- Fixed softmax was too hard, this was introduced by Feb-24 change. Fix for user of softmax result.\n- Added <code>get_training_datalists(sampling_type)</code> is added: <code>'exhaustive'</code> will use all training samples, <code>'more_than_two'</code> is original implementation.</p>\n\n<p>Update Mar-7: Added best single model that marked LB score: Private/Public = 0.88599/0.87523</p>",
      "rawMarkdown": "Though competition is approaching to finish, so ... it's kind of late but let me share what I have been working on this couple of months.\nI'd like to introduce Prototypical Networks (ProtoNets), a metrics learning basically proposed for few-shot problems.\nIt can learn distances between classes as same as Siamese networks but between many classes.\n\n- Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.\nhttps://arxiv.org/pdf/1703.05175.pdf\n\nMy intuition was it could be better discriminator to learn differences between many class samples than to learn just between two classes, and it could be universal discriminator/classifier applicable to variety of problems. While conventional classifier is limited to tell differences between trained classes, metrics learning model is free to handle unknown classes, and measure the degree of deviation from normal = trained classes.\nThere are many few-shot learning models proposed recently, which you can find in this good article \"Advances in few-shot learning: a guided tour\" by Oscar Knagg:\nhttps://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\n\nThen I picked Prototypical Networks because solution is quite simple and practical. But basically it doesn't use pre-trained model in the paper though tested problem was image few-shot classification, and though we know that pre-trained model is essential for better image interpretation/ representation.\nMain contribution of my work is applying ImageNet models to ProtoNets.\n\n- Repository:\nhttps://github.com/daisukelab/protonet-fine-grained-clf\n\n- Notebook is ready to reproduce public LB 0.748 with 100 epochs (takes <strike>an hour or so</strike> 4-5 hours):\nhttps://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/Example_Humpback_Whale_Identification.ipynb\n\nThe best public LB score among my local models is 0.878, and ensemble of them is my score 0.907 as of now. Yes, nothing other than ProtoNets is used in my solution. I even don't tweak training/test samples so much; usual augmentation, normalization, TTA... these pushed my score totally.\n\nThe reason why I'm so much interested in ProtoNets is, it is easy to be understood by non-ML engineers in real world projects.\nI think it is almost proved to work fine without big effort in this competition, then you can also try. :)\n\n----------\nUpdate Feb-21: Fixed training time \"an hour or so\" --&gt; \"4-5 hours\", excuse me to write old info...\n\nUpdate Feb-24: Bug fix - unneeded log() caused calculating wrong distances.\n\nUpdate Feb-26: Added option for data sampling &amp; Bug fix.\n- Fixed softmax was too hard, this was introduced by Feb-24 change. Fix for user of softmax result.\n- Added `get_training_datalists(sampling_type)` is added: `'exhaustive'` will use all training samples, `'more_than_two'` is original implementation.\n\nUpdate Mar-7: Added best single model that marked LB score: Private/Public = 0.88599/0.87523",
      "votes": 54
    },
    {
      "id": 476951,
      "postDate": "2019-02-23T14:08:15.113Z",
      "content": "<p>@Haider Alwasiti</p>\n\n<p>the formula is so much like one-class SVM. And the results too!</p>\n\n<p>visualization of learning  process of prototype network:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/476951/11385/animated.gif\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "@Haider Alwasiti\n \nthe formula is so much like one-class SVM. And the results too!\n\nvisualization of learning  process of prototype network:\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/476951/11385/animated.gif",
      "votes": 13,
      "replies": [
        {
          "id": 477117,
          "postDate": "2019-02-23T23:29:16.820Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>\nThanks for the interesting visualization.</p>\n\n<p>btw, when you've suggested to use <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79524#466297\">sample-to-cluster instead of sample-to-sample  in metric learning</a>:</p>\n\n<p>&gt; if metric learning is used, is it sample-to-sample or\n&gt; sample-to-cluster? (each train id may have more than one images) if\n&gt; the image are not pose-aligned, sample-to-sample is difficult.</p>\n\n<p>This was nothing but a Prototypical Network, especially with our dataset with small number of  images/class :)</p>\n\n<p>I haven't take your suggestion seriously, since I thought that would be almost the same like classification... </p>\n\n<p>Now I see, that for 5-20 images/class, one can use ProtoNets with more efficiency than classification. </p>\n\n<p><strong>Actually, it is</strong>  like classification, but for such small sample size it is better... </p>",
          "rawMarkdown": "@hengck23\nThanks for the interesting visualization.\n\nbtw, when you've suggested to use [sample-to-cluster instead of sample-to-sample  in metric learning][1]:\n\n&gt; if metric learning is used, is it sample-to-sample or\n&gt; sample-to-cluster? (each train id may have more than one images) if\n&gt; the image are not pose-aligned, sample-to-sample is difficult.\n\nThis was nothing but a Prototypical Network, especially with our dataset with small number of  images/class :)\n\nI haven't take your suggestion seriously, since I thought that would be almost the same like classification... \n\nNow I see, that for 5-20 images/class, one can use ProtoNets with more efficiency than classification. \n\n**Actually, it is**  like classification, but for such small sample size it is better... \n\n  [1]: https://www.kaggle.com/c/humpback-whale-identification/discussion/79524#466297",
          "votes": 2
        }
      ]
    },
    {
      "id": 478010,
      "postDate": "2019-02-25T16:00:10.523Z",
      "content": "<p>i make several changes, i was able to get about \n  - LB 0.80 for resnet18 using 224 input.  (threshold at 30% new-whale)\n  - LB 0.85 for resnet18 using 384 input.  (threshold at 30% new-whale)\n  - LB 0.88 for resnet18 using 640 input.  (threshold at 30% new-whale)\n  - LB  0.91 for resnet18 using 800input.  (threshold at 30% new-whale)  </p>\n\n<p>I am not sure which modification actually affect results:</p>\n\n<p>(i ran daisukela baseline code and get LB0.759 for resnet18 224 input crop out of 256)</p>\n\n<ul>\n<li><p>strong augmentation, see code</p></li>\n<li><p>add linear layer after average pooling of resnet18</p></li>\n<li><p>change the order of prototype update:</p>\n\n<ol><li>compute embedding</li>\n<li>compute distance from current batch samples to all 5004 previous prototype. I keep a buffer of all 5004 class prototype.</li>\n<li>softmax distance and compute loss, update network parameters</li>\n<li>update prototype  </li></ol></li>\n<li><p>use all train id samples (even if it is single sample class). new-whale is excluded for this version reported here.</p></li>\n<li><p>I note that my convergence is very much slower than that of daisukela</p></li>\n</ul>",
      "rawMarkdown": "i make several changes, i was able to get about \n  - LB 0.80 for resnet18 using 224 input.  (threshold at 30% new-whale)\n  - LB 0.85 for resnet18 using 384 input.  (threshold at 30% new-whale)\n  - LB 0.88 for resnet18 using 640 input.  (threshold at 30% new-whale)\n  - LB  0.91 for resnet18 using 800input.  (threshold at 30% new-whale)  \n\nI am not sure which modification actually affect results:\n\n(i ran daisukela baseline code and get LB0.759 for resnet18 224 input crop out of 256)\n\n\n- strong augmentation, see code\n\n- add linear layer after average pooling of resnet18\n\n- change the order of prototype update:\n \n    1.  compute embedding\n    2. compute distance from current batch samples to all 5004 previous prototype. I keep a buffer of all 5004 class prototype.\n    3. softmax distance and compute loss, update network parameters\n    4. update prototype  \n\n\n- use all train id samples (even if it is single sample class). new-whale is excluded for this version reported here.\n\n- I note that my convergence is very much slower than that of daisukela\n\n",
      "votes": 9,
      "replies": [
        {
          "id": 478036,
          "postDate": "2019-02-25T16:30:52.263Z",
          "content": "<p>One thing that I noticed is that if you change the Avg Pooling to Max Pooling the convergence is muuuch slower. </p>\n\n<p>Furthermore, I also tried adding a fully connected layer after the pooling but it did not change much if I recall correctly. \nEven though for other models like CosFace and triplet network (batch hard) the fully connected improved the results considerably.</p>",
          "rawMarkdown": "One thing that I noticed is that if you change the Avg Pooling to Max Pooling the convergence is muuuch slower. \n\nFurthermore, I also tried adding a fully connected layer after the pooling but it did not change much if I recall correctly. \nEven though for other models like CosFace and triplet network (batch hard) the fully connected improved the results considerably."
        },
        {
          "id": 478039,
          "postDate": "2019-02-25T16:33:14.587Z",
          "content": "<p>I reckon the best pooling layer is the <a href=\"https://arxiv.org/pdf/1711.02512.pdf\">Generalized Mean Pooling</a> or something like it.</p>",
          "rawMarkdown": "I reckon the best pooling layer is the [Generalized Mean Pooling][1] or something like it.\n\n[1]: https://arxiv.org/pdf/1711.02512.pdf",
          "votes": 1
        },
        {
          "id": 478051,
          "postDate": "2019-02-25T16:51:04.680Z",
          "content": "<p>this is another option:</p>\n\n<p>\"Alpha pooling for fine-grained recognition\"\n<a href=\"https://github.com/cvjena/alpha_pooling\">https://github.com/cvjena/alpha_pooling</a></p>",
          "rawMarkdown": "this is another option:\n\n \"Alpha pooling for fine-grained recognition\"\nhttps://github.com/cvjena/alpha_pooling\n",
          "votes": 3
        },
        {
          "id": 478070,
          "postDate": "2019-02-25T17:27:40.623Z",
          "content": "<p>I didn't know about this one, I'll take a look.</p>\n\n<p>Another useful trick is to remove the last ReLu before pooling. ReLu admits only positive values, therefore it restrains the \"usable\" section of the embedding space. Take a look:\n<img src=\"https://i.imgur.com/SBxMP91.png\" alt=\"ReLu\"></p>\n\n<p>Image from <a href=\"https://arxiv.org/abs/1704.08063\">SphereFace paper</a></p>",
          "rawMarkdown": "I didn't know about this one, I'll take a look.\n\nAnother useful trick is to remove the last ReLu before pooling. ReLu admits only positive values, therefore it restrains the \"usable\" section of the embedding space. Take a look:\n![ReLu](https://i.imgur.com/SBxMP91.png)\n\nImage from [SphereFace paper](https://arxiv.org/abs/1704.08063)",
          "votes": 5
        },
        {
          "id": 478128,
          "postDate": "2019-02-25T19:21:10.737Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>\nresnet50 ; sz 224 ; LB 0.829\nresnet18  ;  sz 384  ; LB 0.826\nwith the same code of this kernel</p>",
          "rawMarkdown": "@hengck23\nresnet50 ; sz 224 ; LB 0.829\nresnet18  ;  sz 384  ; LB 0.826\nwith the same code of this kernel",
          "votes": 2
        },
        {
          "id": 478225,
          "postDate": "2019-02-25T22:45:23.150Z",
          "content": "<p>Hi <a href=\"/hwasiti\">@hwasiti</a>,  do you refer 'this kernel' to  <a href=\"/daisukelab\">@daisukelab</a> kernel?</p>",
          "rawMarkdown": "Hi @hwasiti,  do you refer 'this kernel' to  @daisukelab kernel?"
        }
      ]
    },
    {
      "id": 479195,
      "postDate": "2019-02-27T02:30:33.683Z",
      "content": "<p>beyond prototype, e.g. after training a prototype network, </p>\n\n<ul>\n<li><p>use it to create prototype  from train/test samples</p></li>\n<li><p>given a test image, compute distances from all the prototypes</p></li>\n<li><p>you can build another network to make prediction for 1004 classification based on the distances or distances+other__features as input</p></li>\n</ul>\n\n<p>in fact instead of thresholding, you can make a new-whale predictor as:\ndistances  --&gt; CNN --&gt; new-whale or not</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/479195/11435/prototype.png\" alt=\"enter image description here\"></p>\n\n<p><a href=\"https://arxiv.org/pdf/1806.03018.pdf\">https://arxiv.org/pdf/1806.03018.pdf</a></p>",
      "rawMarkdown": "beyond prototype, e.g. after training a prototype network, \n\n- use it to create prototype  from train/test samples\n\n- given a test image, compute distances from all the prototypes\n\n- you can build another network to make prediction for 1004 classification based on the distances or distances+other__features as input\n\nin fact instead of thresholding, you can make a new-whale predictor as:\ndistances  --&gt; CNN --&gt; new-whale or not\n\n\n  ![enter image description here][1]\n\nhttps://arxiv.org/pdf/1806.03018.pdf\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/479195/11435/prototype.png",
      "votes": 7
    },
    {
      "id": 477699,
      "postDate": "2019-02-25T05:36:15.540Z",
      "content": "<p>i run your code. it seems that you have completely ignored new-whales in training.</p>\n\n<p>actually the new-whale can also be treated as single class single train train (1-shot train sample)</p>\n\n<p>some of the new whale have distinct whale features which may help to improve your embedding space.</p>",
      "rawMarkdown": "i run your code. it seems that you have completely ignored new-whales in training.\n\nactually the new-whale can also be treated as single class single train train (1-shot train sample)\n\nsome of the new whale have distinct whale features which may help to improve your embedding space.",
      "votes": 5,
      "replies": [
        {
          "id": 477781,
          "postDate": "2019-02-25T09:00:47.390Z",
          "content": "<p>Hi, training data sampling is following Martin’s, using classes with 2 samples or more only.\nThis is basically because ProtoNets training require two samples for a class minimum. I tried to include 1 sample classes also, but it was not successful. Maybe augmentation was not enough or similar reasons.\nBut it should be possible, I will add option for this tonight.</p>",
          "rawMarkdown": "Hi, training data sampling is following Martin’s, using classes with 2 samples or more only.\nThis is basically because ProtoNets training require two samples for a class minimum. I tried to include 1 sample classes also, but it was not successful. Maybe augmentation was not enough or similar reasons.\nBut it should be possible, I will add option for this tonight."
        },
        {
          "id": 477791,
          "postDate": "2019-02-25T09:12:53.193Z",
          "content": "<p>You are right. But the code need modyfication. In oryginal idea we need 'support' and 'query' for each class (so we need at least 2 examples per class).   Novel have theoretically unique ID so we cannot use them as new class. </p>\n\n<p>I think we can treat them as a 'support' which would confuse <code>query</code> but would not be classified against any class</p>\n\n<p>In general, so many ideas, so less time :P</p>",
          "rawMarkdown": "You are right. But the code need modyfication. In oryginal idea we need 'support' and 'query' for each class (so we need at least 2 examples per class).   Novel have theoretically unique ID so we cannot use them as new class. \n\nI think we can treat them as a 'support' which would confuse `query` but would not be classified against any class\n\nIn general, so many ideas, so less time :P",
          "votes": 4
        },
        {
          "id": 477805,
          "postDate": "2019-02-25T09:32:45.617Z",
          "content": "<p>Hey, I have a plan. ;)\nDataset class is just looking at pandas dataframe, it can fairly easy to extend to accommodate single sample classes, as if it has multiple samples.\nRegarding new whales, assigning fake id is also easy... I hope so, going home...</p>",
          "rawMarkdown": "Hey, I have a plan. ;)\nDataset class is just looking at pandas dataframe, it can fairly easy to extend to accommodate single sample classes, as if it has multiple samples.\nRegarding new whales, assigning fake id is also easy... I hope so, going home...",
          "votes": 1
        },
        {
          "id": 477821,
          "postDate": "2019-02-25T10:10:47.827Z",
          "content": "<p>Here it is. Will commit maybe tomorrow.\n<code>train.py</code> to be extended to have single images doubled.\nAnd new_whale is replaced with fake id. Now 37,098 training images...\nI need stronger augmentation.</p>\n\n<pre><code>df = pd.read_csv(DATA_PATH+'/train.csv')\n\n# Assign fake Id to new_whale\nn_new_whale = len(df[df.Id == 'new_whale'])\ndf.at[df.Id == 'new_whale', 'Id'] = [f'new{i:05d}' for i in range(n_new_whale)]\n\nids = df.Id.values\nclasses = sorted(list(set(ids)))\nimages = df.Image.values\nall_cls2imgs = {cls:images[ids == cls] for cls in classes}\n\n# Duplicate all the single image classes\nsingle_images = [image for image, _id in zip(images, ids) if len(all_cls2imgs[_id]) == 1]\nsingle_labels = [_id   for image, _id in zip(images, ids) if len(all_cls2imgs[_id]) == 1]\n\ntrn_images = list(images) + single_images\ntrn_labels = list(ids) + single_labels\n</code></pre>",
          "rawMarkdown": "Here it is. Will commit maybe tomorrow.\n`train.py` to be extended to have single images doubled.\nAnd new_whale is replaced with fake id. Now 37,098 training images...\nI need stronger augmentation.\n\n    df = pd.read_csv(DATA_PATH+'/train.csv')\n    \n    # Assign fake Id to new_whale\n    n_new_whale = len(df[df.Id == 'new_whale'])\n    df.at[df.Id == 'new_whale', 'Id'] = [f'new{i:05d}' for i in range(n_new_whale)]\n    \n    ids = df.Id.values\n    classes = sorted(list(set(ids)))\n    images = df.Image.values\n    all_cls2imgs = {cls:images[ids == cls] for cls in classes}\n    \n    # Duplicate all the single image classes\n    single_images = [image for image, _id in zip(images, ids) if len(all_cls2imgs[_id]) == 1]\n    single_labels = [_id   for image, _id in zip(images, ids) if len(all_cls2imgs[_id]) == 1]\n    \n    trn_images = list(images) + single_images\n    trn_labels = list(ids) + single_labels",
          "votes": 5
        },
        {
          "id": 477823,
          "postDate": "2019-02-25T10:18:22.947Z",
          "content": "<p>This paper here <a href=\"https://arxiv.org/pdf/1803.00676.pdf\">https://arxiv.org/pdf/1803.00676.pdf</a> uses new entries in training as a type of semi-supervised ProtoNet. I read it sometime ago but if I'm not mistaken,  it uses soft-KNN during episode training to \"create\" labels for the unlabeled images. With this I guess we can use both new_whales/Ids with 1 image and test images.</p>",
          "rawMarkdown": "This paper here https://arxiv.org/pdf/1803.00676.pdf uses new entries in training as a type of semi-supervised ProtoNet. I read it sometime ago but if I'm not mistaken,  it uses soft-KNN during episode training to \"create\" labels for the unlabeled images. With this I guess we can use both new_whales/Ids with 1 image and test images.",
          "votes": 2
        },
        {
          "id": 477825,
          "postDate": "2019-02-25T10:22:12.763Z",
          "content": "<p>Currently my models are pretty poor. Maybe I am doing something wrong ; I do not have much time to look at unfortunately.  </p>\n\n<p>But this is how I generate batches:\nsupport ( n classes )       : x1  x2  x3  ...  xn\nquery ( same n classes) : y1  y2  y3  ... yn\nnew whale ( random ) :    w1  w2 w3 ... n</p>\n\n<p>for the n queries, you compute the distance to the support and the distance to the new whale. Then the loss is a softmax like function where the logits are the negative distances. So new whales are just treated as out of distribution samples. </p>\n\n<p>Up to now I use all classes, even the one with a single image. Whatever I am using ( direct classification or metric learning) I always have more or less the same score ! so I may have a problem somewhere else. I tried several distances:\n- euclidean,\n- cosine (normalized features dot product)\n- direct features dot product, aka local classification,\n- margin based things, </p>\n\n<p>Also using proxies instead of a support set works. The idea is to save the prototypes directly in a matrix (number of classes x features size ) and directly compute the distances based on these proxies, and of course learn these proxies. For computing the loss you extract on the fly the proxies that are present in the batch. At the end the net starts to look at a normal classification network. </p>",
          "rawMarkdown": "Currently my models are pretty poor. Maybe I am doing something wrong ; I do not have much time to look at unfortunately.  \n\nBut this is how I generate batches:\nsupport ( n classes )       : x1  x2  x3  ...  xn\nquery ( same n classes) : y1  y2  y3  ... yn\nnew whale ( random ) :    w1  w2 w3 ... n\n\nfor the n queries, you compute the distance to the support and the distance to the new whale. Then the loss is a softmax like function where the logits are the negative distances. So new whales are just treated as out of distribution samples. \n\nUp to now I use all classes, even the one with a single image. Whatever I am using ( direct classification or metric learning) I always have more or less the same score ! so I may have a problem somewhere else. I tried several distances:\n- euclidean,\n- cosine (normalized features dot product)\n- direct features dot product, aka local classification,\n- margin based things, \n\nAlso using proxies instead of a support set works. The idea is to save the prototypes directly in a matrix (number of classes x features size ) and directly compute the distances based on these proxies, and of course learn these proxies. For computing the loss you extract on the fly the proxies that are present in the batch. At the end the net starts to look at a normal classification network. \n\n\n"
        },
        {
          "id": 477849,
          "postDate": "2019-02-25T11:09:28.867Z",
          "content": "<p>Wow! <a href=\"/daisukelab\">@daisukelab</a>, that was fast. Thanks.</p>",
          "rawMarkdown": "Wow! @daisukelab, that was fast. Thanks."
        },
        {
          "id": 477864,
          "postDate": "2019-02-25T11:49:06.713Z",
          "content": "<p>HI <a href=\"/jeandebleau\">@jeandebleau</a>, actually I have tried everything, using all classes including <code>new_whale</code>, or using except new_whale, ... and come to conclusion once that the best is non <code>new_whale</code> &amp; class &gt;=2 samples.\nSo it's no time but starting with simplest problem setting would be better if you suffer from getting better score...\nP.S. When I was using all classes, I made mistake to assign all new whales as one class new_whale... Thanks to Heng it seems to be learning fine this time...</p>\n\n<p>Hi <a href=\"/sheriytm\">@sheriytm</a>, you can do that by setting this.  None by default to train from Imagenet pre-trained weight.</p>\n\n<pre><code>args.init_weight = 'your_model.pth'\n</code></pre>",
          "rawMarkdown": "HI @jeandebleau, actually I have tried everything, using all classes including `new_whale`, or using except new_whale, ... and come to conclusion once that the best is non `new_whale` &amp; class &gt;=2 samples.\nSo it's no time but starting with simplest problem setting would be better if you suffer from getting better score...\nP.S. When I was using all classes, I made mistake to assign all new whales as one class new_whale... Thanks to Heng it seems to be learning fine this time...\n\nHi @sheriytm, you can do that by setting this.  None by default to train from Imagenet pre-trained weight.\n\n    args.init_weight = 'your_model.pth'",
          "votes": 1
        },
        {
          "id": 477867,
          "postDate": "2019-02-25T11:55:07.663Z",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a>. I deleted my earlier question because I saw where I needed to make the change.</p>",
          "rawMarkdown": "Thanks @daisukelab. I deleted my earlier question because I saw where I needed to make the change."
        },
        {
          "id": 477887,
          "postDate": "2019-02-25T12:28:55.793Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> , thanks ! I will give a last try by removing classes with samples &lt; 2 and new whales. The networks that I implemented work very good on other tasks. So I guess I made a mistake somewhere in between... I do not care about my score ; I already learned a lot !</p>",
          "rawMarkdown": "@daisukelab , thanks ! I will give a last try by removing classes with samples &lt; 2 and new whales. The networks that I implemented work very good on other tasks. So I guess I made a mistake somewhere in between... I do not care about my score ; I already learned a lot !"
        }
      ]
    },
    {
      "id": 476921,
      "postDate": "2019-02-23T13:22:33.057Z",
      "content": "<p>The idea of this network is so interesting.. </p>\n\n<p>Reading this <a href=\"https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\">blog</a> , I remembered that me and @Iafoss came to the conclusion of modifying the approach of one shot into few shot without realizing that it is a thing.. </p>\n\n<p>Quote from the blog post:\n</p>\n\n<p>&gt; The meaning of this is that the prediction of the model, y^, is the\n&gt; weighted sum of the labels, y_i, of the support set, where the weights\n&gt; are a pairwise similarity function, a(x^, x_i), between the query\n&gt; example, x^, and a support set samples, x_i. The labels y_i in this\n&gt; equation are one-hot encoded label vectors. Notice that if we choose\n&gt; a(x^, x_i) to be 1/k for the closest k samples to the query sample and\n&gt; 0 otherwise we recover the k-nearest-neighbours algorithm\n&gt; \n&gt; The key thing to note is that Matching Networks are end-to-end\n&gt; differentiable provided the attention function a(x^, x_i) is\n&gt; differentiable.</p>\n\n<p>I tried to inspect the nearest <a href=\"https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit/comments#468512\">neighbors of my one shot prediction</a> of test set , I thought that if we consider a simple knn the result will be better (green colored will enhance score, red is not).</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/476921/11397/Screenshot%20from%202019-02-09%2011-10-55.png\" alt=\"enter image description here\"></p>\n\n<p>Iafoss even suggested to make it as weighted knn. Later, It did not work for us,  maybe the weight function wasn't perfect. But, this was  nothing but few shot learning.. The Matching Network of the few shot learning goes one step further, by training the model with this approach too and not just in prediction..</p>",
      "rawMarkdown": "The idea of this network is so interesting.. \n\nReading this [blog][1] , I remembered that me and @Iafoss came to the conclusion of modifying the approach of one shot into few shot without realizing that it is a thing.. \n\nQuote from the blog post:\n![enter image description here][2]\n\n&gt; The meaning of this is that the prediction of the model, y^, is the\n&gt; weighted sum of the labels, y_i, of the support set, where the weights\n&gt; are a pairwise similarity function, a(x^, x_i), between the query\n&gt; example, x^, and a support set samples, x_i. The labels y_i in this\n&gt; equation are one-hot encoded label vectors. Notice that if we choose\n&gt; a(x^, x_i) to be 1/k for the closest k samples to the query sample and\n&gt; 0 otherwise we recover the k-nearest-neighbours algorithm\n&gt; \n&gt; The key thing to note is that Matching Networks are end-to-end\n&gt; differentiable provided the attention function a(x^, x_i) is\n&gt; differentiable.\n\nI tried to inspect the nearest [neighbors of my one shot prediction][3] of test set , I thought that if we consider a simple knn the result will be better (green colored will enhance score, red is not).\n\n![enter image description here][4]\n\nIafoss even suggested to make it as weighted knn. Later, It did not work for us,  maybe the weight function wasn't perfect. But, this was  nothing but few shot learning.. The Matching Network of the few shot learning goes one step further, by training the model with this approach too and not just in prediction..\n\n\n  [1]: https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\n  [2]: https://cdn-images-1.medium.com/max/1600/1*Quo_tUQ2kE4v0c-y7n3RCA.png\n  [3]: https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit/comments#468512\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/476921/11397/Screenshot%20from%202019-02-09%2011-10-55.png",
      "votes": 5
    },
    {
      "id": 480572,
      "postDate": "2019-02-28T11:16:24.970Z",
      "content": "<p>new paper today:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1902.10441.pdf\">https://arxiv.org/pdf/1902.10441.pdf</a>\n\"Fix Your Features: Stationary and Maximally Discriminative Embeddings using\nRegular Polytope (Fixed Classifier) Networks\"</p>\n\n<p>close to one of my idea below:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/480572/11457/poly.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "new paper today:\n\nhttps://arxiv.org/pdf/1902.10441.pdf\n\"Fix Your Features: Stationary and Maximally Discriminative Embeddings using\nRegular Polytope (Fixed Classifier) Networks\"\n\n  close to one of my idea below:\n\n\n![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/480572/11457/poly.png",
      "votes": 3,
      "replies": [
        {
          "id": 480621,
          "postDate": "2019-02-28T12:40:11.720Z",
          "content": "<p>I have trained smth like this - for me it overfitted much more than standard approach</p>",
          "rawMarkdown": "I have trained smth like this - for me it overfitted much more than standard approach",
          "votes": 1
        },
        {
          "id": 480678,
          "postDate": "2019-02-28T14:13:03.387Z",
          "content": "<p>I read the paper and it is quite nice. But I'm at a loss on something. If I were to use a Orthoplex the number of dimensions would K / 2, right? So that means I need to encode my labels as a \"two-hot encoding\" where we have K / 2 dimensions and two labels per dimension like ([0, -1], [0, 1], [-1, 0], [1, 0])?</p>",
          "rawMarkdown": "I read the paper and it is quite nice. But I'm at a loss on something. If I were to use a Orthoplex the number of dimensions would K / 2, right? So that means I need to encode my labels as a \"two-hot encoding\" where we have K / 2 dimensions and two labels per dimension like ([0, -1], [0, 1], [-1, 0], [1, 0])?"
        }
      ]
    },
    {
      "id": 476741,
      "postDate": "2019-02-22T16:24:37.013Z",
      "content": "<p>papers worth looking at:</p>\n\n<p>[1] baseline line:  Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.</p>\n\n<p>miniImageNet :  49.42 /  68.20  (1-shot/5-shot)</p>\n\n<ul>\n<li>Prototpyical Networks  + deformation augmentation (mixup, ghost, stitched, montage, and partly erased images, etc.)\n\"IMAGE DEFORMATION META-NETWORKS FOR ONESHOT  LEARNING\"\n<a href=\"https://openreview.net/pdf?id=Sylw7nCqFQ\">https://openreview.net/pdf?id=Sylw7nCqFQ</a></li>\n</ul>\n\n<p>miniImageNet :  57.71 / 74.34 (1-shot/5-shot)</p>\n\n<p>(another paper with almost same idea and same results: <a href=\"http://yugangjiang.info/publication/19AAAI-oneshot.pdf\">http://yugangjiang.info/publication/19AAAI-oneshot.pdf</a>)</p>\n\n<ul>\n<li>Prototpyical Networks  + semisupervised</li>\n</ul>\n\n<p>\"Semi-Supervised Few-Shot Learning with Prototypical Networks\"\n<a href=\"http://metalearning.ml/2017/papers/metalearn17_boney.pdf\">http://metalearning.ml/2017/papers/metalearn17_boney.pdf</a></p>\n\n<p>miniImageNet : 54.05 / 70.92 (1-shot/5-shot)</p>\n\n<hr>\n\n<p>after reading several papers, i think we can do this:</p>\n\n<p>for ids (classes) with many train samples, employ classification</p>\n\n<p>for ids with  5 train samples, employ, 5-shot</p>\n\n<p>for ids with single train samples, employ, 1-shot, metric learning etc ...</p>",
      "rawMarkdown": "papers worth looking at:\n\n[1] baseline line:  Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.\n\nminiImageNet :  49.42 /  68.20  (1-shot/5-shot)\n\n\n- Prototpyical Networks  + deformation augmentation (mixup, ghost, stitched, montage, and partly erased images, etc.)\n\"IMAGE DEFORMATION META-NETWORKS FOR ONESHOT  LEARNING\"\nhttps://openreview.net/pdf?id=Sylw7nCqFQ\n\nminiImageNet :  57.71 / 74.34 (1-shot/5-shot)\n\n(another paper with almost same idea and same results: http://yugangjiang.info/publication/19AAAI-oneshot.pdf)\n\n\n- Prototpyical Networks  + semisupervised\n\n\"Semi-Supervised Few-Shot Learning with Prototypical Networks\"\nhttp://metalearning.ml/2017/papers/metalearn17_boney.pdf\n\nminiImageNet : 54.05 / 70.92 (1-shot/5-shot)\n\n\n----\n\nafter reading several papers, i think we can do this:\n\nfor ids (classes) with many train samples, employ classification\n\nfor ids with  5 train samples, employ, 5-shot\n\n\nfor ids with single train samples, employ, 1-shot, metric learning etc ...",
      "votes": 3,
      "replies": [
        {
          "id": 476814,
          "postDate": "2019-02-23T08:16:20.133Z",
          "content": "<p>Thanks for sharing paper info and great idea, I was thinking about better ensemble algorithm between classifier &amp; ProtoNets results. I'd like to try weighting between classes based on number of class samples...</p>",
          "rawMarkdown": "Thanks for sharing paper info and great idea, I was thinking about better ensemble algorithm between classifier &amp; ProtoNets results. I'd like to try weighting between classes based on number of class samples..."
        },
        {
          "id": 476866,
          "postDate": "2019-02-23T10:56:24.167Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a></p>\n\n<p>My experiment shows that:</p>\n\n<ol>\n<li><p>if you can identify some strong match (train-test match) and give it a pesudo label and add it to training, results improve</p></li>\n<li><p>classifier, metric learning, prototype network are detecting different matches. it seems that they can be teacher-student pair and improve each other results. In particular:</p>\n\n<ul><li><p>classifier is good for cases with &gt;5 samples and of different view</p></li>\n<li><p>metric learning is good for matching similar  pose. e.g. we learn how to match train-poseA/train-poseB pair in training. If the test is of poseA and the ground-truth match is in poseB, metric learning gives good results</p></li>\n<li><p>protype network seems to be in-between the two cases.</p></li></ul></li>\n</ol>\n\n<p>Now i am experimenting with large margin prototype network ,etc</p>",
          "rawMarkdown": "@daisukelab\n \nMy experiment shows that:\n\n1. if you can identify some strong match (train-test match) and give it a pesudo label and add it to training, results improve\n\n2. classifier, metric learning, prototype network are detecting different matches. it seems that they can be teacher-student pair and improve each other results. In particular:\n\n- classifier is good for cases with &gt;5 samples and of different view\n\n- metric learning is good for matching similar  pose. e.g. we learn how to match train-poseA/train-poseB pair in training. If the test is of poseA and the ground-truth match is in poseB, metric learning gives good results\n\n- protype network seems to be in-between the two cases.\n\nNow i am experimenting with large margin prototype network ,etc",
          "votes": 3
        }
      ]
    },
    {
      "id": 474575,
      "postDate": "2019-02-19T15:18:11.760Z",
      "content": "<p>Greate work, thanks for sharing! I was investigating other One-Short-Learning technique and my single model can get 0.924 LB, so better. But I'm using 448 image-size. So maybe both our technique have similar accuracy.</p>\n\n<p>I'm not sure but look like you have not been using any detection there?</p>",
      "rawMarkdown": "Greate work, thanks for sharing! I was investigating other One-Short-Learning technique and my single model can get 0.924 LB, so better. But I'm using 448 image-size. So maybe both our technique have similar accuracy.\n\nI'm not sure but look like you have not been using any detection there?",
      "votes": 3,
      "replies": [
        {
          "id": 474582,
          "postDate": "2019-02-19T15:38:49.847Z",
          "content": "<p>Hi, congratulations for your single model, sounds great! Regarding image size, unfortunately I couldn't make my models trained well with bigger size. I'm looking forward to see your solution after competition if possible.\nAnd regarding cropping image by using bounding box, I forgot to mention but I'm using that for testing phase. Training with original image is better for my models... In the github code, I tried to keep it simple.</p>",
          "rawMarkdown": "Hi, congratulations for your single model, sounds great! Regarding image size, unfortunately I couldn't make my models trained well with bigger size. I'm looking forward to see your solution after competition if possible.\nAnd regarding cropping image by using bounding box, I forgot to mention but I'm using that for testing phase. Training with original image is better for my models... In the github code, I tried to keep it simple."
        },
        {
          "id": 474589,
          "postDate": "2019-02-19T15:45:27.757Z",
          "content": "<p>if you have the keypoint point or box detected, use one shot for part-to-part matching. each part is crop of full resolution image.</p>\n\n<p>An example is to divide whale into 4 overlapping subcells.</p>\n\n<p>in this way, you can do full resolution matching</p>",
          "rawMarkdown": "if you have the keypoint point or box detected, use one shot for part-to-part matching. each part is crop of full resolution image.\n\nAn example is to divide whale into 4 overlapping subcells.\n\nin this way, you can do full resolution matching",
          "votes": 3
        },
        {
          "id": 474961,
          "postDate": "2019-02-20T03:48:25.560Z",
          "content": "<p>Thank you for great suggestion, it sounds promising. I'd like to try :)\n<a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/78453#latest-471775\">https://www.kaggle.com/c/humpback-whale-identification/discussion/78453#latest-471775</a>\nThis cropped &amp; masked image would be the best I guess...</p>",
          "rawMarkdown": "Thank you for great suggestion, it sounds promising. I'd like to try :)\nhttps://www.kaggle.com/c/humpback-whale-identification/discussion/78453#latest-471775\nThis cropped &amp; masked image would be the best I guess..."
        },
        {
          "id": 475188,
          "postDate": "2019-02-20T12:09:46.633Z",
          "content": "<p>Hi Bartek, what kind of model architecture and image size of your single model that can achieve a score of 0.924?Is it keras, pytorch or tensorflow-based?</p>",
          "rawMarkdown": "Hi Bartek, what kind of model architecture and image size of your single model that can achieve a score of 0.924?Is it keras, pytorch or tensorflow-based?\n",
          "votes": 12
        },
        {
          "id": 475795,
          "postDate": "2019-02-21T07:58:56.213Z",
          "content": "<p>It's PyTorch, Se-resnext101. Image size is 448. Maybe this information would help anybody :P</p>",
          "rawMarkdown": "It's PyTorch, Se-resnext101. Image size is 448. Maybe this information would help anybody :P",
          "votes": 4
        },
        {
          "id": 477457,
          "postDate": "2019-02-24T16:02:51.843Z",
          "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, thank you for your post! If I may ask, did you get 0.878LB with a single model/simple prediction or did you use TTA? What about you <a href=\"/melgor\">@melgor</a>? Thank you guys</p>",
          "rawMarkdown": "Hi @daisukelab, thank you for your post! If I may ask, did you get 0.878LB with a single model/simple prediction or did you use TTA? What about you @melgor? Thank you guys"
        },
        {
          "id": 477567,
          "postDate": "2019-02-24T21:32:22.583Z",
          "content": "<p>Hi, I used TTA for 0.878LB, simple single model prediction was 0.869LB.</p>",
          "rawMarkdown": "Hi, I used TTA for 0.878LB, simple single model prediction was 0.869LB.",
          "votes": 2
        }
      ]
    },
    {
      "id": 474333,
      "postDate": "2019-02-19T08:32:31.593Z",
      "content": "<p>Many thanks, i am going through the prototypical nets paper and your implementation seems great. Quick question, do you see improved performance using larger model architectures (DenseNet101, etc). </p>\n\n<p>I also found the youtube video on the same topic: <a href=\"https://www.youtube.com/watch?v=wcKL05DomBU\">https://www.youtube.com/watch?v=wcKL05DomBU</a></p>",
      "rawMarkdown": "Many thanks, i am going through the prototypical nets paper and your implementation seems great. Quick question, do you see improved performance using larger model architectures (DenseNet101, etc). \n\nI also found the youtube video on the same topic: https://www.youtube.com/watch?v=wcKL05DomBU",
      "votes": 3,
      "replies": [
        {
          "id": 474389,
          "postDate": "2019-02-19T09:35:46.627Z",
          "content": "<p>Thanks for youtube video, and regarding larger model, that's one thing I couldn't make it successful.\nSo far ResNets are the best, and it's not always larger is better as far as tried a lot.\nThe same thing is applicable to image size, I have to mention here that the 384 is not optimal, training stops improving.\nOne of my plan is porting to fast.ai, but mabe I cannot make it in time...</p>",
          "rawMarkdown": "Thanks for youtube video, and regarding larger model, that's one thing I couldn't make it successful.\nSo far ResNets are the best, and it's not always larger is better as far as tried a lot.\nThe same thing is applicable to image size, I have to mention here that the 384 is not optimal, training stops improving.\nOne of my plan is porting to fast.ai, but mabe I cannot make it in time...",
          "votes": 1
        }
      ]
    },
    {
      "id": 474200,
      "postDate": "2019-02-19T04:20:19.803Z",
      "content": "<p>thanks! i will definitely take a look at this.</p>\n\n<p>it is also worth searching for papers that reference this paper.</p>\n\n<p>i found this improved version: <a href=\"https://arxiv.org/pdf/1803.00676.pdf\">https://arxiv.org/pdf/1803.00676.pdf</a></p>",
      "rawMarkdown": "thanks! i will definitely take a look at this.\n\nit is also worth searching for papers that reference this paper.\n\ni found this improved version: https://arxiv.org/pdf/1803.00676.pdf",
      "votes": 3,
      "replies": [
        {
          "id": 474375,
          "postDate": "2019-02-19T09:23:31.237Z",
          "content": "<p>Hi, firstly let me say thanks to your many interesting posts, I'm learning from you a lot!\nAnd thanks again for sharing the link of updated study, it's interesting to apply semi-supervised learning with ProtoNets. I'll try if time permits... It's only 10 days left. BTW, your single model is great, it's now 0.908...</p>",
          "rawMarkdown": "Hi, firstly let me say thanks to your many interesting posts, I'm learning from you a lot!\nAnd thanks again for sharing the link of updated study, it's interesting to apply semi-supervised learning with ProtoNets. I'll try if time permits... It's only 10 days left. BTW, your single model is great, it's now 0.908..."
        }
      ]
    },
    {
      "id": 477205,
      "postDate": "2019-02-24T05:53:26.117Z",
      "content": "<p>Everyone, I updated repository due to rather serious bug.\nThis will cause nan for some of test samples when you calculate softmax of distances.</p>\n\n<ul>\n<li><p>Commit info.\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/commit/1b8685c1c995d41b10744df56a733311a53fa497\">https://github.com/daisukelab/protonet-fine-grained-clf/commit/1b8685c1c995d41b10744df56a733311a53fa497</a></p></li>\n<li><p>You can check if your result contains nan or not as follows, this will output list of nan-result indexes:</p>\n\n<p>np.where(np.isnan(test_preds))[0]</p></li>\n</ul>\n\n<p>I apologize if you spend time for issue related to this...</p>",
      "rawMarkdown": "Everyone, I updated repository due to rather serious bug.\nThis will cause nan for some of test samples when you calculate softmax of distances.\n\n- Commit info.\nhttps://github.com/daisukelab/protonet-fine-grained-clf/commit/1b8685c1c995d41b10744df56a733311a53fa497\n\n- You can check if your result contains nan or not as follows, this will output list of nan-result indexes:\n\n    np.where(np.isnan(test_preds))[0]\n\nI apologize if you spend time for issue related to this...",
      "votes": 4,
      "replies": [
        {
          "id": 478269,
          "postDate": "2019-02-26T00:48:58.167Z",
          "content": "<p>Last minute update again:</p>\n\n<p>Update Feb-26: Added option for data sampling &amp; Bug fix.</p>\n\n<ul>\n<li>Fixed softmax was too hard, this was introduced by Feb-24 change. Fix for user of softmax result.</li>\n<li>Added <code>get_training_datalists(sampling_type)</code> is added: 'exhaustive' will use all training samples, 'morethantwo' is original implementation.</li>\n</ul>",
          "rawMarkdown": "Last minute update again:\n\nUpdate Feb-26: Added option for data sampling &amp; Bug fix.\n\n- Fixed softmax was too hard, this was introduced by Feb-24 change. Fix for user of softmax result.\n- Added `get_training_datalists(sampling_type)` is added: 'exhaustive' will use all training samples, 'morethantwo' is original implementation.",
          "votes": 2
        }
      ]
    },
    {
      "id": 477268,
      "postDate": "2019-02-24T09:10:48.710Z",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a>, thanks for sharing. I ran your example code on colab for 40 epochs and scored 0.470 on the LB. Running it again with slight modifications to see if I can improve the score.</p>",
      "rawMarkdown": "@daisukelab, thanks for sharing. I ran your example code on colab for 40 epochs and scored 0.470 on the LB. Running it again with slight modifications to see if I can improve the score.",
      "votes": 1,
      "replies": [
        {
          "id": 477291,
          "postDate": "2019-02-24T09:52:13.630Z",
          "content": "<p>Hi, Just FYI I can find log like this when getting LB score over 0.7.</p>\n\n<pre><code>... categorical_accuracy=0.994, val_1-shot_10-way_acc=1\n</code></pre>",
          "rawMarkdown": "Hi, Just FYI I can find log like this when getting LB score over 0.7.\n\n    ... categorical_accuracy=0.994, val_1-shot_10-way_acc=1",
          "votes": 1
        },
        {
          "id": 477321,
          "postDate": "2019-02-24T10:35:10.493Z",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a>. Here is the last entry I can see, it is still running.</p>\n\n<p>&gt;Epoch 32: ... loss=0.747, categorical_accuracy=0.79, val_1-shot_10-way_acc=0.999]</p>",
          "rawMarkdown": "Thanks @daisukelab. Here is the last entry I can see, it is still running.\n\n&gt;Epoch 32: ... loss=0.747, categorical_accuracy=0.79, val_1-shot_10-way_acc=0.999]",
          "votes": 1
        },
        {
          "id": 477329,
          "postDate": "2019-02-24T10:51:57.787Z",
          "content": "<p>I think <code>categorical_accuracy</code> is not matured enough, it means your model:\n- Randomly picks 50 class, then try to classify in between 50 class and 1-0.79=0.21 will fail.\n- If your model is applied to full 5004 classes, model will fail more...</p>\n\n<p>So I'm checking model to be almost perfect in 50-way 1-shot problem. Then it has good metrics even with much more classes.</p>",
          "rawMarkdown": "I think `categorical_accuracy` is not matured enough, it means your model:\n- Randomly picks 50 class, then try to classify in between 50 class and 1-0.79=0.21 will fail.\n- If your model is applied to full 5004 classes, model will fail more...\n\nSo I'm checking model to be almost perfect in 50-way 1-shot problem. Then it has good metrics even with much more classes.",
          "votes": 2
        },
        {
          "id": 477337,
          "postDate": "2019-02-24T11:00:38.037Z",
          "content": "<p>Yeah. I am hoping it would improve within the remaining 18 epochs :-) I did set the num of class to 100.</p>\n\n<p>For the previous run that scored 0.470, with class=50, the final numbers are:-</p>\n\n<blockquote>\n  <p>loss=0.542, categorical_accuracy=0.841, val_1-shot_10-way_acc=0.997</p>\n</blockquote>",
          "rawMarkdown": "Yeah. I am hoping it would improve within the remaining 18 epochs :-) I did set the num of class to 100.\n\nFor the previous run that scored 0.470, with class=50, the final numbers are:-\n\n&gt;loss=0.542, categorical_accuracy=0.841, val_1-shot_10-way_acc=0.997"
        }
      ]
    },
    {
      "id": 475650,
      "postDate": "2019-02-21T03:21:38.617Z",
      "content": "<p>Hey daisukelab, thanks again for such a nice implementation of prototype networks. I’m quite interested in the properties of the prototypes so it’s great to have working network to play with after the competition. </p>\n\n<p>Quick question, you mention that your notebook takes an hour or so to reproduce your submission, that’s for inference and not training right?</p>",
      "rawMarkdown": "Hey daisukelab, thanks again for such a nice implementation of prototype networks. I’m quite interested in the properties of the prototypes so it’s great to have working network to play with after the competition. \n\nQuick question, you mention that your notebook takes an hour or so to reproduce your submission, that’s for inference and not training right?",
      "votes": 1,
      "replies": [
        {
          "id": 475766,
          "postDate": "2019-02-21T07:00:06.323Z",
          "content": "<p>Thank you for your question, I noticed that I wrote wrong information about training time, it actually takes 4-5 hours. I didn't notice that it have changed when cleaning code...\nMy apologies to everybody who expected short training time...</p>",
          "rawMarkdown": "Thank you for your question, I noticed that I wrote wrong information about training time, it actually takes 4-5 hours. I didn't notice that it have changed when cleaning code...\nMy apologies to everybody who expected short training time...",
          "votes": 1
        },
        {
          "id": 475777,
          "postDate": "2019-02-21T07:27:08.830Z",
          "content": "<p>No worries at all, this morning I finished training a densenet121 after about 18 hours with size 128 and scored 0.56, which is quite good imo. I’ll report back when ResNet 34 at 256 is done tomorrow. </p>",
          "rawMarkdown": "No worries at all, this morning I finished training a densenet121 after about 18 hours with size 128 and scored 0.56, which is quite good imo. I’ll report back when ResNet 34 at 256 is done tomorrow. ",
          "votes": 4
        },
        {
          "id": 475786,
          "postDate": "2019-02-21T07:44:49.660Z",
          "content": "<p>That's why I stopped exploring bigger model :) I hope you will see good scores.</p>",
          "rawMarkdown": "That's why I stopped exploring bigger model :) I hope you will see good scores.",
          "votes": 1
        },
        {
          "id": 476984,
          "postDate": "2019-02-23T16:01:12.897Z",
          "content": "<p>Resnet34 took about 26 hours to train at 256 with default settings from your repo and got a 0.812. The training time was only so long because I am dumb and forgot that when I use my windows machine I set workers to 0 to avoid a forking pickle error. I ran resnet34 again on my Linux machine and trained it again with 8 workers for the same result in ~7 hours. </p>",
          "rawMarkdown": "Resnet34 took about 26 hours to train at 256 with default settings from your repo and got a 0.812. The training time was only so long because I am dumb and forgot that when I use my windows machine I set workers to 0 to avoid a forking pickle error. I ran resnet34 again on my Linux machine and trained it again with 8 workers for the same result in ~7 hours. ",
          "votes": 3
        },
        {
          "id": 477108,
          "postDate": "2019-02-23T22:48:04.050Z",
          "content": "<p>Thanks for sharing your result, one more thing I could share is k-way (the number of class) to increase from sample code's 50. Making it larger would make the score better a little more as far as I tried...</p>",
          "rawMarkdown": "Thanks for sharing your result, one more thing I could share is k-way (the number of class) to increase from sample code's 50. Making it larger would make the score better a little more as far as I tried...",
          "votes": 1
        }
      ]
    },
    {
      "id": 474212,
      "postDate": "2019-02-19T04:41:38.533Z",
      "content": "<p>Thanks! Very interesting!\n\" ensemble of them is my score 0.907 as of now\" For ensemble of them, do you mean you have several models to predict, and choose the final prediction by the possibility from mix of the prediction? </p>",
      "rawMarkdown": "Thanks! Very interesting!\n\" ensemble of them is my score 0.907 as of now\" For ensemble of them, do you mean you have several models to predict, and choose the final prediction by the possibility from mix of the prediction? ",
      "votes": 1,
      "replies": [
        {
          "id": 474377,
          "postDate": "2019-02-19T09:25:55.847Z",
          "content": "<p>Yes just applying ensemble of many predictions with different parameters. I'm taking mean of predicted distances instead of softmax probabilities.</p>",
          "rawMarkdown": "Yes just applying ensemble of many predictions with different parameters. I'm taking mean of predicted distances instead of softmax probabilities.",
          "votes": 2
        }
      ]
    },
    {
      "id": 474182,
      "postDate": "2019-02-19T03:29:39.387Z",
      "content": "<p>Excellent! Thanks for showing this. I've been tinkering with protonets and a few other fgc methods as well, but nothing like what you have done.</p>",
      "rawMarkdown": "Excellent! Thanks for showing this. I've been tinkering with protonets and a few other fgc methods as well, but nothing like what you have done.",
      "votes": 1
    },
    {
      "id": 485548,
      "postDate": "2019-03-07T14:38:04.747Z",
      "content": "<p>Hi <a href=\"/sheriytm\">@sheriytm</a> and <a href=\"/hwasiti\">@hwasiti</a>, I guess you might want to know that I've uploaded my best single model source code for your reference.</p>\n\n<p><a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/k_Submission.ipynb\">https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/k_Submission.ipynb</a></p>",
      "rawMarkdown": "Hi @sheriytm and @hwasiti, I guess you might want to know that I've uploaded my best single model source code for your reference.\n\nhttps://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/k_Submission.ipynb",
      "votes": 2,
      "replies": [
        {
          "id": 485567,
          "postDate": "2019-03-07T14:57:38.937Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>, thanks so much for sharing.</p>",
          "rawMarkdown": "@daisukelab, thanks so much for sharing."
        }
      ]
    },
    {
      "id": 477999,
      "postDate": "2019-02-25T15:45:27.823Z",
      "content": "<p>i find  that it is possible to do this for toy data:</p>\n\n<p>fix the prototype vector (no need to update).  e.g i use 11110000000 for class1, 0000111100000000 for class2, etc. 0000000 is for background class (new whale). this can ensure that the prototype center are \"farthest apart\" (inter-class variance). rubbish background data are projected to the origin and class data are projected outwards the origin, in orthogonal subspace</p>\n\n<p>But i haven't prove that  it will work for real image in this challenge.</p>\n\n<p>this means that if your feature space is \"big enough\" random prototype works</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/477999/11409/proto.png\" alt=\"enter image description here\"></p>\n\n<p>black is background class classification space</p>",
      "rawMarkdown": "i find  that it is possible to do this for toy data:\n\nfix the prototype vector (no need to update).  e.g i use 11110000000 for class1, 0000111100000000 for class2, etc. 0000000 is for background class (new whale). this can ensure that the prototype center are \"farthest apart\" (inter-class variance). rubbish background data are projected to the origin and class data are projected outwards the origin, in orthogonal subspace\n\nBut i haven't prove that  it will work for real image in this challenge.\n\nthis means that if your feature space is \"big enough\" random prototype works\n\n\n\n ![enter image description here][1]\n\n\nblack is background class classification space\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/477999/11409/proto.png",
      "votes": 2
    },
    {
      "id": 476117,
      "postDate": "2019-02-21T16:20:00.533Z",
      "content": "<p>i now read the \"Prototpyical Networks for Few-shot Learning\" in detail.</p>\n\n<p>my feeling is that this is similar to center loss (if you are using prototype = center of support).</p>\n\n<p>I feel that the following could be related:</p>\n\n<p>\"Rethinking Feature Distribution for Loss Functions in Image Classification\"\n- Weitao Wan, Yuanyi Zhong, Tianpeng Li, Jiansheng Chen</p>\n\n<p>basically it assume Gaussian distribution around  each center.  Gaussian distribution is enforced by KL divergence loss.\nThe center are mde far apart for margin loss</p>",
      "rawMarkdown": "i now read the \"Prototpyical Networks for Few-shot Learning\" in detail.\n\nmy feeling is that this is similar to center loss (if you are using prototype = center of support).\n\nI feel that the following could be related:\n\n\"Rethinking Feature Distribution for Loss Functions in Image Classification\"\n- Weitao Wan, Yuanyi Zhong, Tianpeng Li, Jiansheng Chen\n\n\nbasically it assume Gaussian distribution around  each center.  Gaussian distribution is enforced by KL divergence loss.\nThe center are mde far apart for margin loss",
      "votes": 2,
      "replies": [
        {
          "id": 476220,
          "postDate": "2019-02-21T19:47:45.963Z",
          "content": "<p>There is also a \"gaussian\" version of the prototypical network :\n<a href=\"https://arxiv.org/abs/1708.02735\">https://arxiv.org/abs/1708.02735</a>\nwhere you learn both means and covariance matrices. However for a large number of classes, it may take way too much memory: 512 features , 5004 classes = 512x512x5004 float is a lot !</p>\n\n<p>There is also the proxies approach which seem to work well:\n<a href=\"http://openaccess.thecvf.com/content_ICCV_2017/papers/Movshovitz-Attias_No_Fuss_Distance_ICCV_2017_paper.pdf\">http://openaccess.thecvf.com/content_ICCV_2017/papers/Movshovitz-Attias_No_Fuss_Distance_ICCV_2017_paper.pdf</a></p>\n\n<p>The only thing that bother me conceptually and that I do not understand well is: what happens when a class is formed of multiple clusters ? basically a multimodal distribution ?</p>\n\n<p>They talk a little about this problem here (Figure 13) \"soft k nearest neighbor loss\", and especially compare with triplet loss: \n<a href=\"https://arxiv.org/pdf/1902.01889.pdf\">https://arxiv.org/pdf/1902.01889.pdf</a></p>",
          "rawMarkdown": "There is also a \"gaussian\" version of the prototypical network :\nhttps://arxiv.org/abs/1708.02735\nwhere you learn both means and covariance matrices. However for a large number of classes, it may take way too much memory: 512 features , 5004 classes = 512x512x5004 float is a lot !\n\nThere is also the proxies approach which seem to work well:\nhttp://openaccess.thecvf.com/content_ICCV_2017/papers/Movshovitz-Attias_No_Fuss_Distance_ICCV_2017_paper.pdf\n\nThe only thing that bother me conceptually and that I do not understand well is: what happens when a class is formed of multiple clusters ? basically a multimodal distribution ?\n\nThey talk a little about this problem here (Figure 13) \"soft k nearest neighbor loss\", and especially compare with triplet loss: \nhttps://arxiv.org/pdf/1902.01889.pdf\n\n",
          "votes": 5
        },
        {
          "id": 476825,
          "postDate": "2019-02-23T09:12:37.013Z",
          "content": "<p>Thank you <a href=\"/hengck23\">@hengck23</a> and <a href=\"/jeandebleau\">@jeandebleau</a>, very interesting papers.\nI have also checked center loss paper, really interesting to encourage classifiers to gain margin, and paper also shows that it can be solution to adversarial sample issue which I was worried about.\nAnd it sounds like center loss and its variants are also simple way to bring to real world applications.\nThanks again! very informative.</p>",
          "rawMarkdown": "Thank you @hengck23 and @jeandebleau, very interesting papers.\nI have also checked center loss paper, really interesting to encourage classifiers to gain margin, and paper also shows that it can be solution to adversarial sample issue which I was worried about.\nAnd it sounds like center loss and its variants are also simple way to bring to real world applications.\nThanks again! very informative."
        }
      ]
    },
    {
      "id": 474990,
      "postDate": "2019-02-20T05:07:50.513Z",
      "content": "<p>Thanks for sharing! I notice that the training procedure of  Prototypical Networks you are presenting in the github is training on one-shot task, which means that there are one sample per class on a batch during training. It seems that from this point of view, the prototypical network is equivalent to Matching Network, since in the training stage, prototype vector is just the sample embedding itself. The difference is on test time, when \"proto_net.make_prototypes(trn_dl)\" is called. I wonder if i am wrong cause i am working on matching network. thanks :)</p>",
      "rawMarkdown": "Thanks for sharing! I notice that the training procedure of  Prototypical Networks you are presenting in the github is training on one-shot task, which means that there are one sample per class on a batch during training. It seems that from this point of view, the prototypical network is equivalent to Matching Network, since in the training stage, prototype vector is just the sample embedding itself. The difference is on test time, when \"proto_net.make_prototypes(trn_dl)\" is called. I wonder if i am wrong cause i am working on matching network. thanks :)",
      "votes": 2,
      "replies": [
        {
          "id": 475186,
          "postDate": "2019-02-20T12:09:01.310Z",
          "content": "<p>Hey dear Matching Networks user, nice to see you! :)\nYes I'm using ProtoNets with one-shot in training time, equivalent to Matching Networks as you said, and as described in the ProtoNets paper.\nAnd as exactly you mention, I'm feeding many shot, so to say as-much-as-samples-there-shot. :)\nThis is actually against the original paper says - train/test n would be better to be equal.\nI have to say that I didn't test n=1 in test time, due to intuition it would be better taking mean of many-shots than 1-shot. If training n=5 or so, I will try the same n=5 in test time, but 1-shot result is reported too bad.\nBut I would be trying some more different attempts, n=1 could be best with selected one samples from classes... Thanks for your intriguing comment.</p>",
          "rawMarkdown": "Hey dear Matching Networks user, nice to see you! :)\nYes I'm using ProtoNets with one-shot in training time, equivalent to Matching Networks as you said, and as described in the ProtoNets paper.\nAnd as exactly you mention, I'm feeding many shot, so to say as-much-as-samples-there-shot. :)\nThis is actually against the original paper says - train/test n would be better to be equal.\nI have to say that I didn't test n=1 in test time, due to intuition it would be better taking mean of many-shots than 1-shot. If training n=5 or so, I will try the same n=5 in test time, but 1-shot result is reported too bad.\nBut I would be trying some more different attempts, n=1 could be best with selected one samples from classes... Thanks for your intriguing comment.",
          "votes": 1
        }
      ]
    },
    {
      "id": 483254,
      "postDate": "2019-03-04T11:29:07.150Z",
      "content": "<p>I noticed the other other 2 types of few shots learning (Matching Networks and the Model-agnostic Meta-Learning) had been implemented in the code too (i guess by the original author). Did you try them?</p>",
      "rawMarkdown": "I noticed the other other 2 types of few shots learning (Matching Networks and the Model-agnostic Meta-Learning) had been implemented in the code too (i guess by the original author). Did you try them?",
      "replies": [
        {
          "id": 483277,
          "postDate": "2019-03-04T12:15:06.540Z",
          "content": "<p>Hi, at first, I have to emphasize that most of the implemented code are credit to the author Oscar Knagg.</p>\n\n<p>And he wrote great article for the implementation for reproducing 3 papers:\n- <a href=\"https://towardsdatascience.com/advances-in-few-shot-learning-reproducing-results-in-pytorch-aba70dee541d\">Advances in few-shot learning: reproducing results in PyTorch</a>, Towards Data Science</p>\n\n<p>And I haven't tried much for MAML and Matching Networks yet...\nYou can find original repository below, and experimentation code is there for these networks.\n<a href=\"https://github.com/oscarknagg/few-shot\">https://github.com/oscarknagg/few-shot</a></p>",
          "rawMarkdown": "Hi, at first, I have to emphasize that most of the implemented code are credit to the author Oscar Knagg.\n\nAnd he wrote great article for the implementation for reproducing 3 papers:\n- [Advances in few-shot learning: reproducing results in PyTorch](https://towardsdatascience.com/advances-in-few-shot-learning-reproducing-results-in-pytorch-aba70dee541d), Towards Data Science\n\nAnd I haven't tried much for MAML and Matching Networks yet...\nYou can find original repository below, and experimentation code is there for these networks.\nhttps://github.com/oscarknagg/few-shot"
        }
      ]
    },
    {
      "id": 480335,
      "postDate": "2019-02-28T04:13:35.370Z",
      "content": "<p>Will it be better to use the bounding box?</p>",
      "rawMarkdown": "Will it be better to use the bounding box?",
      "replies": [
        {
          "id": 480339,
          "postDate": "2019-02-28T04:21:30.617Z",
          "content": "<p>If you are using the code on github, it doesn't work fine when training with cropped images. Test would be better with cropped images.\nI suspect that augmentation is not enough, because it works even in training if augmentation code is replaced with Martin's solution.</p>",
          "rawMarkdown": "If you are using the code on github, it doesn't work fine when training with cropped images. Test would be better with cropped images.\nI suspect that augmentation is not enough, because it works even in training if augmentation code is replaced with Martin's solution."
        }
      ]
    },
    {
      "id": 479914,
      "postDate": "2019-02-27T15:09:55.190Z",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a>\nDid you use FP32 or FP16 in your code?</p>",
      "rawMarkdown": "@daisukelab\nDid you use FP32 or FP16 in your code?",
      "replies": [
        {
          "id": 479961,
          "postDate": "2019-02-27T16:13:08.467Z",
          "content": "<p>Default, should be FP32</p>",
          "rawMarkdown": "Default, should be FP32",
          "votes": 1
        },
        {
          "id": 480000,
          "postDate": "2019-02-27T17:00:14.227Z",
          "content": "<p>It's too late to ask now, but what do you mean by default? Is there any argument to change to assign the model as fp16? </p>",
          "rawMarkdown": "It's too late to ask now, but what do you mean by default? Is there any argument to change to assign the model as fp16? "
        },
        {
          "id": 480007,
          "postDate": "2019-02-27T17:07:26.187Z",
          "content": "<p>I think nothing fed to program for changing to fp16...</p>",
          "rawMarkdown": "I think nothing fed to program for changing to fp16..."
        },
        {
          "id": 480035,
          "postDate": "2019-02-27T17:49:21.800Z",
          "content": "<p>I'm  sorry, what do you mean ?</p>",
          "rawMarkdown": "I'm  sorry, what do you mean ?"
        },
        {
          "id": 480076,
          "postDate": "2019-02-27T18:51:53.057Z",
          "content": "<p>I've been trying to get the protonet to work with fastai, I doubt I will get it before the contest ends, but that could be a nice way to use fp16.</p>",
          "rawMarkdown": "I've been trying to get the protonet to work with fastai, I doubt I will get it before the contest ends, but that could be a nice way to use fp16.",
          "votes": 2
        },
        {
          "id": 480187,
          "postDate": "2019-02-27T22:10:29.437Z",
          "content": "<p>This commit log could be helpful for porting to FP16, first thing I did after forking original implementation was changing base precision from double to float. Nothing is done other than this regarding floating point precision.</p>\n\n<p><a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/commit/b9e749f97fcd2b6b7cd0c02e3f0a422a067877f0\">https://github.com/daisukelab/protonet-fine-grained-clf/commit/b9e749f97fcd2b6b7cd0c02e3f0a422a067877f0</a></p>",
          "rawMarkdown": "This commit log could be helpful for porting to FP16, first thing I did after forking original implementation was changing base precision from double to float. Nothing is done other than this regarding floating point precision.\n\nhttps://github.com/daisukelab/protonet-fine-grained-clf/commit/b9e749f97fcd2b6b7cd0c02e3f0a422a067877f0",
          "votes": 2
        },
        {
          "id": 480193,
          "postDate": "2019-02-27T22:32:58.670Z",
          "content": "<p>Thanks dl</p>",
          "rawMarkdown": "Thanks dl"
        },
        {
          "id": 480249,
          "postDate": "2019-02-28T01:24:35.277Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>\nThanks\nI think there is no need to change the input from float to FP16. The model will do that automatically if it is converted into FP16</p>\n\n<p>Doing so, we can get almost double the batch size for free which could be very helpful (x1.8 to be exact) </p>\n\n<p>From <a href=\"/iafoss\">@iafoss</a> pytorch + fastai  <a href=\"https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit?scriptVersionId=10640082\">kernel</a>:\n<code>m = model.module if isinstance(model,FP16) else model</code></p>\n\n<p>I think this is the way to convert into half precision. Wrap the model inside module.. </p>\n\n<p><a href=\"/interneuron\">@interneuron</a>\nThat would be very interesting..\nYou are right... There are a lot of nice stuff we can get, once it is ported to fastai..\nIf you will be able to share the code after the comp ends, please give a hint in the fastai forum with a post :)</p>",
          "rawMarkdown": "@daisukelab\nThanks\nI think there is no need to change the input from float to FP16. The model will do that automatically if it is converted into FP16\n\nDoing so, we can get almost double the batch size for free which could be very helpful (x1.8 to be exact) \n\nFrom @iafoss pytorch + fastai  [kernel][1]:\n` m = model.module if isinstance(model,FP16) else model`\n\nI think this is the way to convert into half precision. Wrap the model inside module.. \n\n@interneuron\nThat would be very interesting..\nYou are right... There are a lot of nice stuff we can get, once it is ported to fastai..\nIf you will be able to share the code after the comp ends, please give a hint in the fastai forum with a post :)\n\n\n  [1]: https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit?scriptVersionId=10640082",
          "votes": 1
        },
        {
          "id": 480265,
          "postDate": "2019-02-28T01:53:20.110Z",
          "content": "<p>If I get it working I will certainly post the code, even if not I will anyway as a curiosity. My knowledge of fastai is far from complete. </p>",
          "rawMarkdown": "If I get it working I will certainly post the code, even if not I will anyway as a curiosity. My knowledge of fastai is far from complete. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 479277,
      "postDate": "2019-02-27T03:31:57.607Z",
      "content": "<p>Thanks for sharing… :).\nIn the function prepare_submission you set new_whale_thresh=-1.85.How to confirm the new_whale_thresh value？</p>",
      "rawMarkdown": "Thanks for sharing… :).\nIn the function prepare_submission you set new_whale_thresh=-1.85.How to confirm the new_whale_thresh value？",
      "replies": [
        {
          "id": 479288,
          "postDate": "2019-02-27T03:43:43.703Z",
          "content": "<p>That's as I wrote in the notebook. </p>\n\n<pre><code>Adjust threshold so that it contains new_whale for 0.30 ~ 0.37 which depends on how much the test set would have new_whale you think.\n</code></pre>\n\n<p>This <code>0.3373115577889447</code> is the new_whale rate, you can check the code. :)</p>\n\n<pre><code>app_whale_n1_k50_q1_epoch100 0.3373115577889447 3146\n</code></pre>",
          "rawMarkdown": "That's as I wrote in the notebook. \n\n    Adjust threshold so that it contains new_whale for 0.30 ~ 0.37 which depends on how much the test set would have new_whale you think.\n\nThis `0.3373115577889447` is the new_whale rate, you can check the code. :)\n\n    app_whale_n1_k50_q1_epoch100 0.3373115577889447 3146",
          "votes": 1
        },
        {
          "id": 479393,
          "postDate": "2019-02-27T05:31:26.917Z",
          "content": "<p>Sorry, I didn’t see it carefully, thank you.</p>",
          "rawMarkdown": "Sorry, I didn’t see it carefully, thank you."
        },
        {
          "id": 479917,
          "postDate": "2019-02-27T15:12:33.253Z",
          "content": "<p>In your example of new whale rate\n3146 are the new whales predicted in the test set\nwhich is 0.3373 of the test set\nright?</p>\n\n<p>But the test set = 7960\n7960 * 0.3373 = 2685\nand not 3146</p>\n\n<p>Did I miss something?</p>",
          "rawMarkdown": "In your example of new whale rate\n3146 are the new whales predicted in the test set\nwhich is 0.3373 of the test set\nright?\n\nBut the test set = 7960\n7960 * 0.3373 = 2685\nand not 3146\n\nDid I miss something?"
        },
        {
          "id": 479971,
          "postDate": "2019-02-27T16:23:49.597Z",
          "content": "<p>Excuse me but 3rd number is not related to new whale, number of whales found in the submission. It shows how much whales your model found in test set as the top prediction result.</p>\n\n<pre><code>len(set(pd.read_csv(f'subs/{submission_filename}.csv.gz').Id.str.split().apply(lambda x: x[0]).values))\n</code></pre>",
          "rawMarkdown": "Excuse me but 3rd number is not related to new whale, number of whales found in the submission. It shows how much whales your model found in test set as the top prediction result.\n\n    len(set(pd.read_csv(f'subs/{submission_filename}.csv.gz').Id.str.split().apply(lambda x: x[0]).values))",
          "votes": 1
        }
      ]
    },
    {
      "id": 478819,
      "postDate": "2019-02-26T17:09:37.730Z",
      "content": "<p>Just wondering if anyone tried combined softmax loss and prototype loss, i.e multi task classification and one shot leaening?</p>",
      "rawMarkdown": "Just wondering if anyone tried combined softmax loss and prototype loss, i.e multi task classification and one shot leaening?",
      "replies": [
        {
          "id": 479005,
          "postDate": "2019-02-26T22:54:16Z",
          "content": "<p>By that you mean adding a fully-connected layer on top of the final embedding and computing two losses: (1) softmax loss on the last layer (5004 neurons) and (2) prototype loss on the distance of the embedding (pre-ultimate layer)?</p>\n\n<p>So if you are creating episodes like 100 images per batch 1-shot, 50-way it would be training softmax on 100 images (2 images per class, 50 classes) and prototype loss on 50 images.</p>",
          "rawMarkdown": "By that you mean adding a fully-connected layer on top of the final embedding and computing two losses: (1) softmax loss on the last layer (5004 neurons) and (2) prototype loss on the distance of the embedding (pre-ultimate layer)?\n\nSo if you are creating episodes like 100 images per batch 1-shot, 50-way it would be training softmax on 100 images (2 images per class, 50 classes) and prototype loss on 50 images."
        }
      ]
    },
    {
      "id": 475082,
      "postDate": "2019-02-20T08:26:26.427Z",
      "content": "<p>Greate paper, I'm sure I'm gonna use this for my thesis!</p>",
      "rawMarkdown": "Greate paper, I'm sure I'm gonna use this for my thesis!"
    },
    {
      "id": 474915,
      "postDate": "2019-02-20T01:50:34.970Z",
      "content": "<p>Thank you <a href=\"/daisukelab\">@daisukelab</a>, I found ProtoNet too, but I'm new guy to this area, can't adapt to this competition, because the episodic sampling and how to choice <em>K</em>-Way <em>N</em>-Shot, this is a very good tutorial for me ! </p>",
      "rawMarkdown": "Thank you @daisukelab, I found ProtoNet too, but I'm new guy to this area, can't adapt to this competition, because the episodic sampling and how to choice *K*-Way *N*-Shot, this is a very good tutorial for me ! "
    },
    {
      "id": 474768,
      "postDate": "2019-02-19T19:22:41.160Z",
      "content": "<p>Very helpful, Thanks for sharing <a href=\"/daisukelab\">@daisukelab</a></p>",
      "rawMarkdown": "Very helpful, Thanks for sharing @daisukelab"
    },
    {
      "id": 474560,
      "postDate": "2019-02-19T14:49:24.387Z",
      "content": "<p>Thanks for sharing。\nIt seems that  your example miss some function？</p>",
      "rawMarkdown": "Thanks for sharing。\nIt seems that  your example miss some function？",
      "replies": [
        {
          "id": 474572,
          "postDate": "2019-02-19T15:08:56.693Z",
          "content": "<p>Hi, you will need to install some modules, please find in readme.\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf\">https://github.com/daisukelab/protonet-fine-grained-clf</a>\nOr something could be missing in readme, but everything should be found from pip anyway.\nExcuse me it is not useful to list in requirement.txt.</p>",
          "rawMarkdown": "Hi, you will need to install some modules, please find in readme.\nhttps://github.com/daisukelab/protonet-fine-grained-clf\nOr something could be missing in readme, but everything should be found from pip anyway.\nExcuse me it is not useful to list in requirement.txt.",
          "votes": 1
        },
        {
          "id": 474912,
          "postDate": "2019-02-20T01:40:59.743Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 474188,
      "postDate": "2019-02-19T03:50:22.137Z",
      "content": "<p>Sugoi!. \nThank for sharing. :D. </p>",
      "rawMarkdown": "Sugoi!. \nThank for sharing. :D. ",
      "votes": 1
    },
    {
      "id": 477753,
      "postDate": "2019-02-25T07:27:03.337Z",
      "content": "<p>Thanks for sharing... :)</p>",
      "rawMarkdown": "Thanks for sharing... :)"
    },
    {
      "id": 476851,
      "postDate": "2019-02-23T10:14:34.710Z",
      "content": "<p>Thanks!!</p>",
      "rawMarkdown": "Thanks!!"
    },
    {
      "id": 476012,
      "postDate": "2019-02-21T13:53:38.923Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 476951,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-23T14:08:15.113000",
      "content": "<p>@Haider Alwasiti</p>\n\n<p>the formula is so much like one-class SVM. And the results too!</p>\n\n<p>visualization of learning  process of prototype network:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/476951/11385/animated.gif\" alt=\"enter image description here\"></p>",
      "votes": 13,
      "replies": [
        {
          "id": 477117,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-23T23:29:16.820000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>\nThanks for the interesting visualization.</p>\n\n<p>btw, when you've suggested to use <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79524#466297\">sample-to-cluster instead of sample-to-sample  in metric learning</a>:</p>\n\n<p>&gt; if metric learning is used, is it sample-to-sample or\n&gt; sample-to-cluster? (each train id may have more than one images) if\n&gt; the image are not pose-aligned, sample-to-sample is difficult.</p>\n\n<p>This was nothing but a Prototypical Network, especially with our dataset with small number of  images/class :)</p>\n\n<p>I haven't take your suggestion seriously, since I thought that would be almost the same like classification... </p>\n\n<p>Now I see, that for 5-20 images/class, one can use ProtoNets with more efficiency than classification. </p>\n\n<p><strong>Actually, it is</strong>  like classification, but for such small sample size it is better... </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 478010,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-25T16:00:10.523000",
      "content": "<p>i make several changes, i was able to get about \n  - LB 0.80 for resnet18 using 224 input.  (threshold at 30% new-whale)\n  - LB 0.85 for resnet18 using 384 input.  (threshold at 30% new-whale)\n  - LB 0.88 for resnet18 using 640 input.  (threshold at 30% new-whale)\n  - LB  0.91 for resnet18 using 800input.  (threshold at 30% new-whale)  </p>\n\n<p>I am not sure which modification actually affect results:</p>\n\n<p>(i ran daisukela baseline code and get LB0.759 for resnet18 224 input crop out of 256)</p>\n\n<ul>\n<li><p>strong augmentation, see code</p></li>\n<li><p>add linear layer after average pooling of resnet18</p></li>\n<li><p>change the order of prototype update:</p>\n\n<ol><li>compute embedding</li>\n<li>compute distance from current batch samples to all 5004 previous prototype. I keep a buffer of all 5004 class prototype.</li>\n<li>softmax distance and compute loss, update network parameters</li>\n<li>update prototype  </li></ol></li>\n<li><p>use all train id samples (even if it is single sample class). new-whale is excluded for this version reported here.</p></li>\n<li><p>I note that my convergence is very much slower than that of daisukela</p></li>\n</ul>",
      "votes": 9,
      "replies": [
        {
          "id": 478036,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-25T16:30:52.263000",
          "content": "<p>One thing that I noticed is that if you change the Avg Pooling to Max Pooling the convergence is muuuch slower. </p>\n\n<p>Furthermore, I also tried adding a fully connected layer after the pooling but it did not change much if I recall correctly. \nEven though for other models like CosFace and triplet network (batch hard) the fully connected improved the results considerably.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478039,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-25T16:33:14.587000",
          "content": "<p>I reckon the best pooling layer is the <a href=\"https://arxiv.org/pdf/1711.02512.pdf\">Generalized Mean Pooling</a> or something like it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 478051,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-02-25T16:51:04.680000",
          "content": "<p>this is another option:</p>\n\n<p>\"Alpha pooling for fine-grained recognition\"\n<a href=\"https://github.com/cvjena/alpha_pooling\">https://github.com/cvjena/alpha_pooling</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 478070,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-25T17:27:40.623000",
          "content": "<p>I didn't know about this one, I'll take a look.</p>\n\n<p>Another useful trick is to remove the last ReLu before pooling. ReLu admits only positive values, therefore it restrains the \"usable\" section of the embedding space. Take a look:\n<img src=\"https://i.imgur.com/SBxMP91.png\" alt=\"ReLu\"></p>\n\n<p>Image from <a href=\"https://arxiv.org/abs/1704.08063\">SphereFace paper</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 478128,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-25T19:21:10.737000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>\nresnet50 ; sz 224 ; LB 0.829\nresnet18  ;  sz 384  ; LB 0.826\nwith the same code of this kernel</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 478225,
          "author_name": "Steinhafen",
          "author_url": "",
          "post_date": "2019-02-25T22:45:23.150000",
          "content": "<p>Hi <a href=\"/hwasiti\">@hwasiti</a>,  do you refer 'this kernel' to  <a href=\"/daisukelab\">@daisukelab</a> kernel?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 479195,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-27T02:30:33.683000",
      "content": "<p>beyond prototype, e.g. after training a prototype network, </p>\n\n<ul>\n<li><p>use it to create prototype  from train/test samples</p></li>\n<li><p>given a test image, compute distances from all the prototypes</p></li>\n<li><p>you can build another network to make prediction for 1004 classification based on the distances or distances+other__features as input</p></li>\n</ul>\n\n<p>in fact instead of thresholding, you can make a new-whale predictor as:\ndistances  --&gt; CNN --&gt; new-whale or not</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/479195/11435/prototype.png\" alt=\"enter image description here\"></p>\n\n<p><a href=\"https://arxiv.org/pdf/1806.03018.pdf\">https://arxiv.org/pdf/1806.03018.pdf</a></p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 477699,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-25T05:36:15.540000",
      "content": "<p>i run your code. it seems that you have completely ignored new-whales in training.</p>\n\n<p>actually the new-whale can also be treated as single class single train train (1-shot train sample)</p>\n\n<p>some of the new whale have distinct whale features which may help to improve your embedding space.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 477781,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-25T09:00:47.390000",
          "content": "<p>Hi, training data sampling is following Martin’s, using classes with 2 samples or more only.\nThis is basically because ProtoNets training require two samples for a class minimum. I tried to include 1 sample classes also, but it was not successful. Maybe augmentation was not enough or similar reasons.\nBut it should be possible, I will add option for this tonight.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 477791,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-02-25T09:12:53.193000",
          "content": "<p>You are right. But the code need modyfication. In oryginal idea we need 'support' and 'query' for each class (so we need at least 2 examples per class).   Novel have theoretically unique ID so we cannot use them as new class. </p>\n\n<p>I think we can treat them as a 'support' which would confuse <code>query</code> but would not be classified against any class</p>\n\n<p>In general, so many ideas, so less time :P</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 477805,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-25T09:32:45.617000",
          "content": "<p>Hey, I have a plan. ;)\nDataset class is just looking at pandas dataframe, it can fairly easy to extend to accommodate single sample classes, as if it has multiple samples.\nRegarding new whales, assigning fake id is also easy... I hope so, going home...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 477821,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-25T10:10:47.827000",
          "content": "<p>Here it is. Will commit maybe tomorrow.\n<code>train.py</code> to be extended to have single images doubled.\nAnd new_whale is replaced with fake id. Now 37,098 training images...\nI need stronger augmentation.</p>\n\n<pre><code>df = pd.read_csv(DATA_PATH+'/train.csv')\n\n# Assign fake Id to new_whale\nn_new_whale = len(df[df.Id == 'new_whale'])\ndf.at[df.Id == 'new_whale', 'Id'] = [f'new{i:05d}' for i in range(n_new_whale)]\n\nids = df.Id.values\nclasses = sorted(list(set(ids)))\nimages = df.Image.values\nall_cls2imgs = {cls:images[ids == cls] for cls in classes}\n\n# Duplicate all the single image classes\nsingle_images = [image for image, _id in zip(images, ids) if len(all_cls2imgs[_id]) == 1]\nsingle_labels = [_id   for image, _id in zip(images, ids) if len(all_cls2imgs[_id]) == 1]\n\ntrn_images = list(images) + single_images\ntrn_labels = list(ids) + single_labels\n</code></pre>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 477823,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-25T10:18:22.947000",
          "content": "<p>This paper here <a href=\"https://arxiv.org/pdf/1803.00676.pdf\">https://arxiv.org/pdf/1803.00676.pdf</a> uses new entries in training as a type of semi-supervised ProtoNet. I read it sometime ago but if I'm not mistaken,  it uses soft-KNN during episode training to \"create\" labels for the unlabeled images. With this I guess we can use both new_whales/Ids with 1 image and test images.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 477825,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "2019-02-25T10:22:12.763000",
          "content": "<p>Currently my models are pretty poor. Maybe I am doing something wrong ; I do not have much time to look at unfortunately.  </p>\n\n<p>But this is how I generate batches:\nsupport ( n classes )       : x1  x2  x3  ...  xn\nquery ( same n classes) : y1  y2  y3  ... yn\nnew whale ( random ) :    w1  w2 w3 ... n</p>\n\n<p>for the n queries, you compute the distance to the support and the distance to the new whale. Then the loss is a softmax like function where the logits are the negative distances. So new whales are just treated as out of distribution samples. </p>\n\n<p>Up to now I use all classes, even the one with a single image. Whatever I am using ( direct classification or metric learning) I always have more or less the same score ! so I may have a problem somewhere else. I tried several distances:\n- euclidean,\n- cosine (normalized features dot product)\n- direct features dot product, aka local classification,\n- margin based things, </p>\n\n<p>Also using proxies instead of a support set works. The idea is to save the prototypes directly in a matrix (number of classes x features size ) and directly compute the distances based on these proxies, and of course learn these proxies. For computing the loss you extract on the fly the proxies that are present in the batch. At the end the net starts to look at a normal classification network. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 477849,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-25T11:09:28.867000",
          "content": "<p>Wow! <a href=\"/daisukelab\">@daisukelab</a>, that was fast. Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 477864,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-25T11:49:06.713000",
          "content": "<p>HI <a href=\"/jeandebleau\">@jeandebleau</a>, actually I have tried everything, using all classes including <code>new_whale</code>, or using except new_whale, ... and come to conclusion once that the best is non <code>new_whale</code> &amp; class &gt;=2 samples.\nSo it's no time but starting with simplest problem setting would be better if you suffer from getting better score...\nP.S. When I was using all classes, I made mistake to assign all new whales as one class new_whale... Thanks to Heng it seems to be learning fine this time...</p>\n\n<p>Hi <a href=\"/sheriytm\">@sheriytm</a>, you can do that by setting this.  None by default to train from Imagenet pre-trained weight.</p>\n\n<pre><code>args.init_weight = 'your_model.pth'\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 477867,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-25T11:55:07.663000",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a>. I deleted my earlier question because I saw where I needed to make the change.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 477887,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "2019-02-25T12:28:55.793000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> , thanks ! I will give a last try by removing classes with samples &lt; 2 and new whales. The networks that I implemented work very good on other tasks. So I guess I made a mistake somewhere in between... I do not care about my score ; I already learned a lot !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 476921,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-02-23T13:22:33.057000",
      "content": "<p>The idea of this network is so interesting.. </p>\n\n<p>Reading this <a href=\"https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\">blog</a> , I remembered that me and @Iafoss came to the conclusion of modifying the approach of one shot into few shot without realizing that it is a thing.. </p>\n\n<p>Quote from the blog post:\n</p>\n\n<p>&gt; The meaning of this is that the prediction of the model, y^, is the\n&gt; weighted sum of the labels, y_i, of the support set, where the weights\n&gt; are a pairwise similarity function, a(x^, x_i), between the query\n&gt; example, x^, and a support set samples, x_i. The labels y_i in this\n&gt; equation are one-hot encoded label vectors. Notice that if we choose\n&gt; a(x^, x_i) to be 1/k for the closest k samples to the query sample and\n&gt; 0 otherwise we recover the k-nearest-neighbours algorithm\n&gt; \n&gt; The key thing to note is that Matching Networks are end-to-end\n&gt; differentiable provided the attention function a(x^, x_i) is\n&gt; differentiable.</p>\n\n<p>I tried to inspect the nearest <a href=\"https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit/comments#468512\">neighbors of my one shot prediction</a> of test set , I thought that if we consider a simple knn the result will be better (green colored will enhance score, red is not).</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/476921/11397/Screenshot%20from%202019-02-09%2011-10-55.png\" alt=\"enter image description here\"></p>\n\n<p>Iafoss even suggested to make it as weighted knn. Later, It did not work for us,  maybe the weight function wasn't perfect. But, this was  nothing but few shot learning.. The Matching Network of the few shot learning goes one step further, by training the model with this approach too and not just in prediction..</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 480572,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-28T11:16:24.970000",
      "content": "<p>new paper today:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1902.10441.pdf\">https://arxiv.org/pdf/1902.10441.pdf</a>\n\"Fix Your Features: Stationary and Maximally Discriminative Embeddings using\nRegular Polytope (Fixed Classifier) Networks\"</p>\n\n<p>close to one of my idea below:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/480572/11457/poly.png\" alt=\"enter image description here\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 480621,
          "author_name": "old-ufo",
          "author_url": "",
          "post_date": "2019-02-28T12:40:11.720000",
          "content": "<p>I have trained smth like this - for me it overfitted much more than standard approach</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 480678,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-28T14:13:03.387000",
          "content": "<p>I read the paper and it is quite nice. But I'm at a loss on something. If I were to use a Orthoplex the number of dimensions would K / 2, right? So that means I need to encode my labels as a \"two-hot encoding\" where we have K / 2 dimensions and two labels per dimension like ([0, -1], [0, 1], [-1, 0], [1, 0])?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 476741,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-22T16:24:37.013000",
      "content": "<p>papers worth looking at:</p>\n\n<p>[1] baseline line:  Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.</p>\n\n<p>miniImageNet :  49.42 /  68.20  (1-shot/5-shot)</p>\n\n<ul>\n<li>Prototpyical Networks  + deformation augmentation (mixup, ghost, stitched, montage, and partly erased images, etc.)\n\"IMAGE DEFORMATION META-NETWORKS FOR ONESHOT  LEARNING\"\n<a href=\"https://openreview.net/pdf?id=Sylw7nCqFQ\">https://openreview.net/pdf?id=Sylw7nCqFQ</a></li>\n</ul>\n\n<p>miniImageNet :  57.71 / 74.34 (1-shot/5-shot)</p>\n\n<p>(another paper with almost same idea and same results: <a href=\"http://yugangjiang.info/publication/19AAAI-oneshot.pdf\">http://yugangjiang.info/publication/19AAAI-oneshot.pdf</a>)</p>\n\n<ul>\n<li>Prototpyical Networks  + semisupervised</li>\n</ul>\n\n<p>\"Semi-Supervised Few-Shot Learning with Prototypical Networks\"\n<a href=\"http://metalearning.ml/2017/papers/metalearn17_boney.pdf\">http://metalearning.ml/2017/papers/metalearn17_boney.pdf</a></p>\n\n<p>miniImageNet : 54.05 / 70.92 (1-shot/5-shot)</p>\n\n<hr>\n\n<p>after reading several papers, i think we can do this:</p>\n\n<p>for ids (classes) with many train samples, employ classification</p>\n\n<p>for ids with  5 train samples, employ, 5-shot</p>\n\n<p>for ids with single train samples, employ, 1-shot, metric learning etc ...</p>",
      "votes": 3,
      "replies": [
        {
          "id": 476814,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-23T08:16:20.133000",
          "content": "<p>Thanks for sharing paper info and great idea, I was thinking about better ensemble algorithm between classifier &amp; ProtoNets results. I'd like to try weighting between classes based on number of class samples...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 476866,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-02-23T10:56:24.167000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a></p>\n\n<p>My experiment shows that:</p>\n\n<ol>\n<li><p>if you can identify some strong match (train-test match) and give it a pesudo label and add it to training, results improve</p></li>\n<li><p>classifier, metric learning, prototype network are detecting different matches. it seems that they can be teacher-student pair and improve each other results. In particular:</p>\n\n<ul><li><p>classifier is good for cases with &gt;5 samples and of different view</p></li>\n<li><p>metric learning is good for matching similar  pose. e.g. we learn how to match train-poseA/train-poseB pair in training. If the test is of poseA and the ground-truth match is in poseB, metric learning gives good results</p></li>\n<li><p>protype network seems to be in-between the two cases.</p></li></ul></li>\n</ol>\n\n<p>Now i am experimenting with large margin prototype network ,etc</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 474575,
      "author_name": "Bartek",
      "author_url": "",
      "post_date": "2019-02-19T15:18:11.760000",
      "content": "<p>Greate work, thanks for sharing! I was investigating other One-Short-Learning technique and my single model can get 0.924 LB, so better. But I'm using 448 image-size. So maybe both our technique have similar accuracy.</p>\n\n<p>I'm not sure but look like you have not been using any detection there?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 474582,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-19T15:38:49.847000",
          "content": "<p>Hi, congratulations for your single model, sounds great! Regarding image size, unfortunately I couldn't make my models trained well with bigger size. I'm looking forward to see your solution after competition if possible.\nAnd regarding cropping image by using bounding box, I forgot to mention but I'm using that for testing phase. Training with original image is better for my models... In the github code, I tried to keep it simple.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 474589,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-02-19T15:45:27.757000",
          "content": "<p>if you have the keypoint point or box detected, use one shot for part-to-part matching. each part is crop of full resolution image.</p>\n\n<p>An example is to divide whale into 4 overlapping subcells.</p>\n\n<p>in this way, you can do full resolution matching</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 474961,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-20T03:48:25.560000",
          "content": "<p>Thank you for great suggestion, it sounds promising. I'd like to try :)\n<a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/78453#latest-471775\">https://www.kaggle.com/c/humpback-whale-identification/discussion/78453#latest-471775</a>\nThis cropped &amp; masked image would be the best I guess...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 475188,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-20T12:09:46.633000",
          "content": "<p>Hi Bartek, what kind of model architecture and image size of your single model that can achieve a score of 0.924?Is it keras, pytorch or tensorflow-based?</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 475795,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-02-21T07:58:56.213000",
          "content": "<p>It's PyTorch, Se-resnext101. Image size is 448. Maybe this information would help anybody :P</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 477457,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-24T16:02:51.843000",
          "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, thank you for your post! If I may ask, did you get 0.878LB with a single model/simple prediction or did you use TTA? What about you <a href=\"/melgor\">@melgor</a>? Thank you guys</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 477567,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-24T21:32:22.583000",
          "content": "<p>Hi, I used TTA for 0.878LB, simple single model prediction was 0.869LB.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 474333,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2019-02-19T08:32:31.593000",
      "content": "<p>Many thanks, i am going through the prototypical nets paper and your implementation seems great. Quick question, do you see improved performance using larger model architectures (DenseNet101, etc). </p>\n\n<p>I also found the youtube video on the same topic: <a href=\"https://www.youtube.com/watch?v=wcKL05DomBU\">https://www.youtube.com/watch?v=wcKL05DomBU</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 474389,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-19T09:35:46.627000",
          "content": "<p>Thanks for youtube video, and regarding larger model, that's one thing I couldn't make it successful.\nSo far ResNets are the best, and it's not always larger is better as far as tried a lot.\nThe same thing is applicable to image size, I have to mention here that the 384 is not optimal, training stops improving.\nOne of my plan is porting to fast.ai, but mabe I cannot make it in time...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 474200,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-19T04:20:19.803000",
      "content": "<p>thanks! i will definitely take a look at this.</p>\n\n<p>it is also worth searching for papers that reference this paper.</p>\n\n<p>i found this improved version: <a href=\"https://arxiv.org/pdf/1803.00676.pdf\">https://arxiv.org/pdf/1803.00676.pdf</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 474375,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-19T09:23:31.237000",
          "content": "<p>Hi, firstly let me say thanks to your many interesting posts, I'm learning from you a lot!\nAnd thanks again for sharing the link of updated study, it's interesting to apply semi-supervised learning with ProtoNets. I'll try if time permits... It's only 10 days left. BTW, your single model is great, it's now 0.908...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 477205,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "2019-02-24T05:53:26.117000",
      "content": "<p>Everyone, I updated repository due to rather serious bug.\nThis will cause nan for some of test samples when you calculate softmax of distances.</p>\n\n<ul>\n<li><p>Commit info.\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/commit/1b8685c1c995d41b10744df56a733311a53fa497\">https://github.com/daisukelab/protonet-fine-grained-clf/commit/1b8685c1c995d41b10744df56a733311a53fa497</a></p></li>\n<li><p>You can check if your result contains nan or not as follows, this will output list of nan-result indexes:</p>\n\n<p>np.where(np.isnan(test_preds))[0]</p></li>\n</ul>\n\n<p>I apologize if you spend time for issue related to this...</p>",
      "votes": 4,
      "replies": [
        {
          "id": 478269,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-26T00:48:58.167000",
          "content": "<p>Last minute update again:</p>\n\n<p>Update Feb-26: Added option for data sampling &amp; Bug fix.</p>\n\n<ul>\n<li>Fixed softmax was too hard, this was introduced by Feb-24 change. Fix for user of softmax result.</li>\n<li>Added <code>get_training_datalists(sampling_type)</code> is added: 'exhaustive' will use all training samples, 'morethantwo' is original implementation.</li>\n</ul>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 477268,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-02-24T09:10:48.710000",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a>, thanks for sharing. I ran your example code on colab for 40 epochs and scored 0.470 on the LB. Running it again with slight modifications to see if I can improve the score.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 477291,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-24T09:52:13.630000",
          "content": "<p>Hi, Just FYI I can find log like this when getting LB score over 0.7.</p>\n\n<pre><code>... categorical_accuracy=0.994, val_1-shot_10-way_acc=1\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 477321,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-24T10:35:10.493000",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a>. Here is the last entry I can see, it is still running.</p>\n\n<p>&gt;Epoch 32: ... loss=0.747, categorical_accuracy=0.79, val_1-shot_10-way_acc=0.999]</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 477329,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-24T10:51:57.787000",
          "content": "<p>I think <code>categorical_accuracy</code> is not matured enough, it means your model:\n- Randomly picks 50 class, then try to classify in between 50 class and 1-0.79=0.21 will fail.\n- If your model is applied to full 5004 classes, model will fail more...</p>\n\n<p>So I'm checking model to be almost perfect in 50-way 1-shot problem. Then it has good metrics even with much more classes.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 477337,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-24T11:00:38.037000",
          "content": "<p>Yeah. I am hoping it would improve within the remaining 18 epochs :-) I did set the num of class to 100.</p>\n\n<p>For the previous run that scored 0.470, with class=50, the final numbers are:-</p>\n\n<blockquote>\n  <p>loss=0.542, categorical_accuracy=0.841, val_1-shot_10-way_acc=0.997</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 475650,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-02-21T03:21:38.617000",
      "content": "<p>Hey daisukelab, thanks again for such a nice implementation of prototype networks. I’m quite interested in the properties of the prototypes so it’s great to have working network to play with after the competition. </p>\n\n<p>Quick question, you mention that your notebook takes an hour or so to reproduce your submission, that’s for inference and not training right?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 475766,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-21T07:00:06.323000",
          "content": "<p>Thank you for your question, I noticed that I wrote wrong information about training time, it actually takes 4-5 hours. I didn't notice that it have changed when cleaning code...\nMy apologies to everybody who expected short training time...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 475777,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-02-21T07:27:08.830000",
          "content": "<p>No worries at all, this morning I finished training a densenet121 after about 18 hours with size 128 and scored 0.56, which is quite good imo. I’ll report back when ResNet 34 at 256 is done tomorrow. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 475786,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-21T07:44:49.660000",
          "content": "<p>That's why I stopped exploring bigger model :) I hope you will see good scores.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 476984,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-02-23T16:01:12.897000",
          "content": "<p>Resnet34 took about 26 hours to train at 256 with default settings from your repo and got a 0.812. The training time was only so long because I am dumb and forgot that when I use my windows machine I set workers to 0 to avoid a forking pickle error. I ran resnet34 again on my Linux machine and trained it again with 8 workers for the same result in ~7 hours. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 477108,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-23T22:48:04.050000",
          "content": "<p>Thanks for sharing your result, one more thing I could share is k-way (the number of class) to increase from sample code's 50. Making it larger would make the score better a little more as far as I tried...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 474212,
      "author_name": "chen danxia",
      "author_url": "",
      "post_date": "2019-02-19T04:41:38.533000",
      "content": "<p>Thanks! Very interesting!\n\" ensemble of them is my score 0.907 as of now\" For ensemble of them, do you mean you have several models to predict, and choose the final prediction by the possibility from mix of the prediction? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 474377,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-19T09:25:55.847000",
          "content": "<p>Yes just applying ensemble of many predictions with different parameters. I'm taking mean of predicted distances instead of softmax probabilities.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 474182,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-02-19T03:29:39.387000",
      "content": "<p>Excellent! Thanks for showing this. I've been tinkering with protonets and a few other fgc methods as well, but nothing like what you have done.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 485548,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "2019-03-07T14:38:04.747000",
      "content": "<p>Hi <a href=\"/sheriytm\">@sheriytm</a> and <a href=\"/hwasiti\">@hwasiti</a>, I guess you might want to know that I've uploaded my best single model source code for your reference.</p>\n\n<p><a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/k_Submission.ipynb\">https://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/k_Submission.ipynb</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 485567,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-03-07T14:57:38.937000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>, thanks so much for sharing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 477999,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-25T15:45:27.823000",
      "content": "<p>i find  that it is possible to do this for toy data:</p>\n\n<p>fix the prototype vector (no need to update).  e.g i use 11110000000 for class1, 0000111100000000 for class2, etc. 0000000 is for background class (new whale). this can ensure that the prototype center are \"farthest apart\" (inter-class variance). rubbish background data are projected to the origin and class data are projected outwards the origin, in orthogonal subspace</p>\n\n<p>But i haven't prove that  it will work for real image in this challenge.</p>\n\n<p>this means that if your feature space is \"big enough\" random prototype works</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/477999/11409/proto.png\" alt=\"enter image description here\"></p>\n\n<p>black is background class classification space</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 476117,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-21T16:20:00.533000",
      "content": "<p>i now read the \"Prototpyical Networks for Few-shot Learning\" in detail.</p>\n\n<p>my feeling is that this is similar to center loss (if you are using prototype = center of support).</p>\n\n<p>I feel that the following could be related:</p>\n\n<p>\"Rethinking Feature Distribution for Loss Functions in Image Classification\"\n- Weitao Wan, Yuanyi Zhong, Tianpeng Li, Jiansheng Chen</p>\n\n<p>basically it assume Gaussian distribution around  each center.  Gaussian distribution is enforced by KL divergence loss.\nThe center are mde far apart for margin loss</p>",
      "votes": 2,
      "replies": [
        {
          "id": 476220,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "2019-02-21T19:47:45.963000",
          "content": "<p>There is also a \"gaussian\" version of the prototypical network :\n<a href=\"https://arxiv.org/abs/1708.02735\">https://arxiv.org/abs/1708.02735</a>\nwhere you learn both means and covariance matrices. However for a large number of classes, it may take way too much memory: 512 features , 5004 classes = 512x512x5004 float is a lot !</p>\n\n<p>There is also the proxies approach which seem to work well:\n<a href=\"http://openaccess.thecvf.com/content_ICCV_2017/papers/Movshovitz-Attias_No_Fuss_Distance_ICCV_2017_paper.pdf\">http://openaccess.thecvf.com/content_ICCV_2017/papers/Movshovitz-Attias_No_Fuss_Distance_ICCV_2017_paper.pdf</a></p>\n\n<p>The only thing that bother me conceptually and that I do not understand well is: what happens when a class is formed of multiple clusters ? basically a multimodal distribution ?</p>\n\n<p>They talk a little about this problem here (Figure 13) \"soft k nearest neighbor loss\", and especially compare with triplet loss: \n<a href=\"https://arxiv.org/pdf/1902.01889.pdf\">https://arxiv.org/pdf/1902.01889.pdf</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 476825,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-23T09:12:37.013000",
          "content": "<p>Thank you <a href=\"/hengck23\">@hengck23</a> and <a href=\"/jeandebleau\">@jeandebleau</a>, very interesting papers.\nI have also checked center loss paper, really interesting to encourage classifiers to gain margin, and paper also shows that it can be solution to adversarial sample issue which I was worried about.\nAnd it sounds like center loss and its variants are also simple way to bring to real world applications.\nThanks again! very informative.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 474990,
      "author_name": "gengshi",
      "author_url": "",
      "post_date": "2019-02-20T05:07:50.513000",
      "content": "<p>Thanks for sharing! I notice that the training procedure of  Prototypical Networks you are presenting in the github is training on one-shot task, which means that there are one sample per class on a batch during training. It seems that from this point of view, the prototypical network is equivalent to Matching Network, since in the training stage, prototype vector is just the sample embedding itself. The difference is on test time, when \"proto_net.make_prototypes(trn_dl)\" is called. I wonder if i am wrong cause i am working on matching network. thanks :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 475186,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-20T12:09:01.310000",
          "content": "<p>Hey dear Matching Networks user, nice to see you! :)\nYes I'm using ProtoNets with one-shot in training time, equivalent to Matching Networks as you said, and as described in the ProtoNets paper.\nAnd as exactly you mention, I'm feeding many shot, so to say as-much-as-samples-there-shot. :)\nThis is actually against the original paper says - train/test n would be better to be equal.\nI have to say that I didn't test n=1 in test time, due to intuition it would be better taking mean of many-shots than 1-shot. If training n=5 or so, I will try the same n=5 in test time, but 1-shot result is reported too bad.\nBut I would be trying some more different attempts, n=1 could be best with selected one samples from classes... Thanks for your intriguing comment.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 483254,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-03-04T11:29:07.150000",
      "content": "<p>I noticed the other other 2 types of few shots learning (Matching Networks and the Model-agnostic Meta-Learning) had been implemented in the code too (i guess by the original author). Did you try them?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 483277,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-03-04T12:15:06.540000",
          "content": "<p>Hi, at first, I have to emphasize that most of the implemented code are credit to the author Oscar Knagg.</p>\n\n<p>And he wrote great article for the implementation for reproducing 3 papers:\n- <a href=\"https://towardsdatascience.com/advances-in-few-shot-learning-reproducing-results-in-pytorch-aba70dee541d\">Advances in few-shot learning: reproducing results in PyTorch</a>, Towards Data Science</p>\n\n<p>And I haven't tried much for MAML and Matching Networks yet...\nYou can find original repository below, and experimentation code is there for these networks.\n<a href=\"https://github.com/oscarknagg/few-shot\">https://github.com/oscarknagg/few-shot</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 480335,
      "author_name": "zhousheng",
      "author_url": "",
      "post_date": "2019-02-28T04:13:35.370000",
      "content": "<p>Will it be better to use the bounding box?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 480339,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-28T04:21:30.617000",
          "content": "<p>If you are using the code on github, it doesn't work fine when training with cropped images. Test would be better with cropped images.\nI suspect that augmentation is not enough, because it works even in training if augmentation code is replaced with Martin's solution.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 479914,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-02-27T15:09:55.190000",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a>\nDid you use FP32 or FP16 in your code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 479961,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-27T16:13:08.467000",
          "content": "<p>Default, should be FP32</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 480000,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-27T17:00:14.227000",
          "content": "<p>It's too late to ask now, but what do you mean by default? Is there any argument to change to assign the model as fp16? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 480007,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-27T17:07:26.187000",
          "content": "<p>I think nothing fed to program for changing to fp16...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 480035,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-27T17:49:21.800000",
          "content": "<p>I'm  sorry, what do you mean ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 480076,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-02-27T18:51:53.057000",
          "content": "<p>I've been trying to get the protonet to work with fastai, I doubt I will get it before the contest ends, but that could be a nice way to use fp16.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 480187,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-27T22:10:29.437000",
          "content": "<p>This commit log could be helpful for porting to FP16, first thing I did after forking original implementation was changing base precision from double to float. Nothing is done other than this regarding floating point precision.</p>\n\n<p><a href=\"https://github.com/daisukelab/protonet-fine-grained-clf/commit/b9e749f97fcd2b6b7cd0c02e3f0a422a067877f0\">https://github.com/daisukelab/protonet-fine-grained-clf/commit/b9e749f97fcd2b6b7cd0c02e3f0a422a067877f0</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 480193,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-02-27T22:32:58.670000",
          "content": "<p>Thanks dl</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 480249,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-28T01:24:35.277000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>\nThanks\nI think there is no need to change the input from float to FP16. The model will do that automatically if it is converted into FP16</p>\n\n<p>Doing so, we can get almost double the batch size for free which could be very helpful (x1.8 to be exact) </p>\n\n<p>From <a href=\"/iafoss\">@iafoss</a> pytorch + fastai  <a href=\"https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit?scriptVersionId=10640082\">kernel</a>:\n<code>m = model.module if isinstance(model,FP16) else model</code></p>\n\n<p>I think this is the way to convert into half precision. Wrap the model inside module.. </p>\n\n<p><a href=\"/interneuron\">@interneuron</a>\nThat would be very interesting..\nYou are right... There are a lot of nice stuff we can get, once it is ported to fastai..\nIf you will be able to share the code after the comp ends, please give a hint in the fastai forum with a post :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 480265,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-02-28T01:53:20.110000",
          "content": "<p>If I get it working I will certainly post the code, even if not I will anyway as a curiosity. My knowledge of fastai is far from complete. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 479277,
      "author_name": "zhousheng",
      "author_url": "",
      "post_date": "2019-02-27T03:31:57.607000",
      "content": "<p>Thanks for sharing… :).\nIn the function prepare_submission you set new_whale_thresh=-1.85.How to confirm the new_whale_thresh value？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 479288,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-27T03:43:43.703000",
          "content": "<p>That's as I wrote in the notebook. </p>\n\n<pre><code>Adjust threshold so that it contains new_whale for 0.30 ~ 0.37 which depends on how much the test set would have new_whale you think.\n</code></pre>\n\n<p>This <code>0.3373115577889447</code> is the new_whale rate, you can check the code. :)</p>\n\n<pre><code>app_whale_n1_k50_q1_epoch100 0.3373115577889447 3146\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 479393,
          "author_name": "zhousheng",
          "author_url": "",
          "post_date": "2019-02-27T05:31:26.917000",
          "content": "<p>Sorry, I didn’t see it carefully, thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 479917,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-27T15:12:33.253000",
          "content": "<p>In your example of new whale rate\n3146 are the new whales predicted in the test set\nwhich is 0.3373 of the test set\nright?</p>\n\n<p>But the test set = 7960\n7960 * 0.3373 = 2685\nand not 3146</p>\n\n<p>Did I miss something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 479971,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-27T16:23:49.597000",
          "content": "<p>Excuse me but 3rd number is not related to new whale, number of whales found in the submission. It shows how much whales your model found in test set as the top prediction result.</p>\n\n<pre><code>len(set(pd.read_csv(f'subs/{submission_filename}.csv.gz').Id.str.split().apply(lambda x: x[0]).values))\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 478819,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-26T17:09:37.730000",
      "content": "<p>Just wondering if anyone tried combined softmax loss and prototype loss, i.e multi task classification and one shot leaening?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 479005,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-26T22:54:16",
          "content": "<p>By that you mean adding a fully-connected layer on top of the final embedding and computing two losses: (1) softmax loss on the last layer (5004 neurons) and (2) prototype loss on the distance of the embedding (pre-ultimate layer)?</p>\n\n<p>So if you are creating episodes like 100 images per batch 1-shot, 50-way it would be training softmax on 100 images (2 images per class, 50 classes) and prototype loss on 50 images.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 475082,
      "author_name": "Vahid Mostofi",
      "author_url": "",
      "post_date": "2019-02-20T08:26:26.427000",
      "content": "<p>Greate paper, I'm sure I'm gonna use this for my thesis!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474915,
      "author_name": "Tony Zhao",
      "author_url": "",
      "post_date": "2019-02-20T01:50:34.970000",
      "content": "<p>Thank you <a href=\"/daisukelab\">@daisukelab</a>, I found ProtoNet too, but I'm new guy to this area, can't adapt to this competition, because the episodic sampling and how to choice <em>K</em>-Way <em>N</em>-Shot, this is a very good tutorial for me ! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474768,
      "author_name": "Karan Jakhar",
      "author_url": "",
      "post_date": "2019-02-19T19:22:41.160000",
      "content": "<p>Very helpful, Thanks for sharing <a href=\"/daisukelab\">@daisukelab</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474560,
      "author_name": "excllent123",
      "author_url": "",
      "post_date": "2019-02-19T14:49:24.387000",
      "content": "<p>Thanks for sharing。\nIt seems that  your example miss some function？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 474572,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-02-19T15:08:56.693000",
          "content": "<p>Hi, you will need to install some modules, please find in readme.\n<a href=\"https://github.com/daisukelab/protonet-fine-grained-clf\">https://github.com/daisukelab/protonet-fine-grained-clf</a>\nOr something could be missing in readme, but everything should be found from pip anyway.\nExcuse me it is not useful to list in requirement.txt.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 474912,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-20T01:40:59.743000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 474188,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2019-02-19T03:50:22.137000",
      "content": "<p>Sugoi!. \nThank for sharing. :D. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 477753,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-25T07:27:03.337000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476851,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T10:14:34.710000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476012,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T13:53:38.923000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "474174": "Though competition is approaching to finish, so ... it's kind of late but let me share what I have been working on this couple of months.\nI'd like to introduce Prototypical Networks (ProtoNets), a metrics learning basically proposed for few-shot problems.\nIt can learn distances between classes as same as Siamese networks but between many classes.\n\n- Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.\nhttps://arxiv.org/pdf/1703.05175.pdf\n\nMy intuition was it could be better discriminator to learn differences between many class samples than to learn just between two classes, and it could be universal discriminator/classifier applicable to variety of problems. While conventional classifier is limited to tell differences between trained classes, metrics learning model is free to handle unknown classes, and measure the degree of deviation from normal = trained classes.\nThere are many few-shot learning models proposed recently, which you can find in this good article \"Advances in few-shot learning: a guided tour\" by Oscar Knagg:\nhttps://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\n\nThen I picked Prototypical Networks because solution is quite simple and practical. But basically it doesn't use pre-trained model in the paper though tested problem was image few-shot classification, and though we know that pre-trained model is essential for better image interpretation/ representation.\nMain contribution of my work is applying ImageNet models to ProtoNets.\n\n- Repository:\nhttps://github.com/daisukelab/protonet-fine-grained-clf\n\n- Notebook is ready to reproduce public LB 0.748 with 100 epochs (takes <strike>an hour or so</strike> 4-5 hours):\nhttps://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/Example_Humpback_Whale_Identification.ipynb\n\nThe best public LB score among my local models is 0.878, and ensemble of them is my score 0.907 as of now. Yes, nothing other than ProtoNets is used in my solution. I even don't tweak training/test samples so much; usual augmentation, normalization, TTA... these pushed my score totally.\n\nThe reason why I'm so much interested in ProtoNets is, it is easy to be understood by non-ML engineers in real world projects.\nI think it is almost proved to work fine without big effort in this competition, then you can also try. :)\n\n----------\nUpdate Feb-21: Fixed training time \"an hour or so\" --&gt; \"4-5 hours\", excuse me to write old info...\n\nUpdate Feb-24: Bug fix - unneeded log() caused calculating wrong distances.\n\nUpdate Feb-26: Added option for data sampling &amp; Bug fix.\n- Fixed softmax was too hard, this was introduced by Feb-24 change. Fix for user of softmax result.\n- Added `get_training_datalists(sampling_type)` is added: `'exhaustive'` will use all training samples, `'more_than_two'` is original implementation.\n\nUpdate Mar-7: Added best single model that marked LB score: Private/Public = 0.88599/0.87523",
    "476951": "@Haider Alwasiti\n \nthe formula is so much like one-class SVM. And the results too!\n\nvisualization of learning  process of prototype network:\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/476951/11385/animated.gif",
    "478010": "i make several changes, i was able to get about \n  - LB 0.80 for resnet18 using 224 input.  (threshold at 30% new-whale)\n  - LB 0.85 for resnet18 using 384 input.  (threshold at 30% new-whale)\n  - LB 0.88 for resnet18 using 640 input.  (threshold at 30% new-whale)\n  - LB  0.91 for resnet18 using 800input.  (threshold at 30% new-whale)  \n\nI am not sure which modification actually affect results:\n\n(i ran daisukela baseline code and get LB0.759 for resnet18 224 input crop out of 256)\n\n\n- strong augmentation, see code\n\n- add linear layer after average pooling of resnet18\n\n- change the order of prototype update:\n \n    1.  compute embedding\n    2. compute distance from current batch samples to all 5004 previous prototype. I keep a buffer of all 5004 class prototype.\n    3. softmax distance and compute loss, update network parameters\n    4. update prototype  \n\n\n- use all train id samples (even if it is single sample class). new-whale is excluded for this version reported here.\n\n- I note that my convergence is very much slower than that of daisukela\n\n",
    "479195": "beyond prototype, e.g. after training a prototype network, \n\n- use it to create prototype  from train/test samples\n\n- given a test image, compute distances from all the prototypes\n\n- you can build another network to make prediction for 1004 classification based on the distances or distances+other__features as input\n\nin fact instead of thresholding, you can make a new-whale predictor as:\ndistances  --&gt; CNN --&gt; new-whale or not\n\n\n  ![enter image description here][1]\n\nhttps://arxiv.org/pdf/1806.03018.pdf\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/479195/11435/prototype.png",
    "477699": "i run your code. it seems that you have completely ignored new-whales in training.\n\nactually the new-whale can also be treated as single class single train train (1-shot train sample)\n\nsome of the new whale have distinct whale features which may help to improve your embedding space.",
    "476921": "The idea of this network is so interesting.. \n\nReading this [blog][1] , I remembered that me and @Iafoss came to the conclusion of modifying the approach of one shot into few shot without realizing that it is a thing.. \n\nQuote from the blog post:\n![enter image description here][2]\n\n&gt; The meaning of this is that the prediction of the model, y^, is the\n&gt; weighted sum of the labels, y_i, of the support set, where the weights\n&gt; are a pairwise similarity function, a(x^, x_i), between the query\n&gt; example, x^, and a support set samples, x_i. The labels y_i in this\n&gt; equation are one-hot encoded label vectors. Notice that if we choose\n&gt; a(x^, x_i) to be 1/k for the closest k samples to the query sample and\n&gt; 0 otherwise we recover the k-nearest-neighbours algorithm\n&gt; \n&gt; The key thing to note is that Matching Networks are end-to-end\n&gt; differentiable provided the attention function a(x^, x_i) is\n&gt; differentiable.\n\nI tried to inspect the nearest [neighbors of my one shot prediction][3] of test set , I thought that if we consider a simple knn the result will be better (green colored will enhance score, red is not).\n\n![enter image description here][4]\n\nIafoss even suggested to make it as weighted knn. Later, It did not work for us,  maybe the weight function wasn't perfect. But, this was  nothing but few shot learning.. The Matching Network of the few shot learning goes one step further, by training the model with this approach too and not just in prediction..\n\n\n  [1]: https://towardsdatascience.com/advances-in-few-shot-learning-a-guided-tour-36bc10a68b77\n  [2]: https://cdn-images-1.medium.com/max/1600/1*Quo_tUQ2kE4v0c-y7n3RCA.png\n  [3]: https://www.kaggle.com/iafoss/similarity-densenet121-0-805lb-kernel-time-limit/comments#468512\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/476921/11397/Screenshot%20from%202019-02-09%2011-10-55.png",
    "480572": "new paper today:\n\nhttps://arxiv.org/pdf/1902.10441.pdf\n\"Fix Your Features: Stationary and Maximally Discriminative Embeddings using\nRegular Polytope (Fixed Classifier) Networks\"\n\n  close to one of my idea below:\n\n\n![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/480572/11457/poly.png",
    "476741": "papers worth looking at:\n\n[1] baseline line:  Original paper: Prototpyical Networks for Few-shot Learning, Snell et al.\n\nminiImageNet :  49.42 /  68.20  (1-shot/5-shot)\n\n\n- Prototpyical Networks  + deformation augmentation (mixup, ghost, stitched, montage, and partly erased images, etc.)\n\"IMAGE DEFORMATION META-NETWORKS FOR ONESHOT  LEARNING\"\nhttps://openreview.net/pdf?id=Sylw7nCqFQ\n\nminiImageNet :  57.71 / 74.34 (1-shot/5-shot)\n\n(another paper with almost same idea and same results: http://yugangjiang.info/publication/19AAAI-oneshot.pdf)\n\n\n- Prototpyical Networks  + semisupervised\n\n\"Semi-Supervised Few-Shot Learning with Prototypical Networks\"\nhttp://metalearning.ml/2017/papers/metalearn17_boney.pdf\n\nminiImageNet : 54.05 / 70.92 (1-shot/5-shot)\n\n\n----\n\nafter reading several papers, i think we can do this:\n\nfor ids (classes) with many train samples, employ classification\n\nfor ids with  5 train samples, employ, 5-shot\n\n\nfor ids with single train samples, employ, 1-shot, metric learning etc ...",
    "474575": "Greate work, thanks for sharing! I was investigating other One-Short-Learning technique and my single model can get 0.924 LB, so better. But I'm using 448 image-size. So maybe both our technique have similar accuracy.\n\nI'm not sure but look like you have not been using any detection there?",
    "474333": "Many thanks, i am going through the prototypical nets paper and your implementation seems great. Quick question, do you see improved performance using larger model architectures (DenseNet101, etc). \n\nI also found the youtube video on the same topic: https://www.youtube.com/watch?v=wcKL05DomBU",
    "474200": "thanks! i will definitely take a look at this.\n\nit is also worth searching for papers that reference this paper.\n\ni found this improved version: https://arxiv.org/pdf/1803.00676.pdf",
    "477205": "Everyone, I updated repository due to rather serious bug.\nThis will cause nan for some of test samples when you calculate softmax of distances.\n\n- Commit info.\nhttps://github.com/daisukelab/protonet-fine-grained-clf/commit/1b8685c1c995d41b10744df56a733311a53fa497\n\n- You can check if your result contains nan or not as follows, this will output list of nan-result indexes:\n\n    np.where(np.isnan(test_preds))[0]\n\nI apologize if you spend time for issue related to this...",
    "477268": "@daisukelab, thanks for sharing. I ran your example code on colab for 40 epochs and scored 0.470 on the LB. Running it again with slight modifications to see if I can improve the score.",
    "475650": "Hey daisukelab, thanks again for such a nice implementation of prototype networks. I’m quite interested in the properties of the prototypes so it’s great to have working network to play with after the competition. \n\nQuick question, you mention that your notebook takes an hour or so to reproduce your submission, that’s for inference and not training right?",
    "474212": "Thanks! Very interesting!\n\" ensemble of them is my score 0.907 as of now\" For ensemble of them, do you mean you have several models to predict, and choose the final prediction by the possibility from mix of the prediction? ",
    "474182": "Excellent! Thanks for showing this. I've been tinkering with protonets and a few other fgc methods as well, but nothing like what you have done.",
    "485548": "Hi @sheriytm and @hwasiti, I guess you might want to know that I've uploaded my best single model source code for your reference.\n\nhttps://github.com/daisukelab/protonet-fine-grained-clf/blob/master/app/whale/k_Submission.ipynb",
    "477999": "i find  that it is possible to do this for toy data:\n\nfix the prototype vector (no need to update).  e.g i use 11110000000 for class1, 0000111100000000 for class2, etc. 0000000 is for background class (new whale). this can ensure that the prototype center are \"farthest apart\" (inter-class variance). rubbish background data are projected to the origin and class data are projected outwards the origin, in orthogonal subspace\n\nBut i haven't prove that  it will work for real image in this challenge.\n\nthis means that if your feature space is \"big enough\" random prototype works\n\n\n\n ![enter image description here][1]\n\n\nblack is background class classification space\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/477999/11409/proto.png",
    "476117": "i now read the \"Prototpyical Networks for Few-shot Learning\" in detail.\n\nmy feeling is that this is similar to center loss (if you are using prototype = center of support).\n\nI feel that the following could be related:\n\n\"Rethinking Feature Distribution for Loss Functions in Image Classification\"\n- Weitao Wan, Yuanyi Zhong, Tianpeng Li, Jiansheng Chen\n\n\nbasically it assume Gaussian distribution around  each center.  Gaussian distribution is enforced by KL divergence loss.\nThe center are mde far apart for margin loss",
    "474990": "Thanks for sharing! I notice that the training procedure of  Prototypical Networks you are presenting in the github is training on one-shot task, which means that there are one sample per class on a batch during training. It seems that from this point of view, the prototypical network is equivalent to Matching Network, since in the training stage, prototype vector is just the sample embedding itself. The difference is on test time, when \"proto_net.make_prototypes(trn_dl)\" is called. I wonder if i am wrong cause i am working on matching network. thanks :)",
    "483254": "I noticed the other other 2 types of few shots learning (Matching Networks and the Model-agnostic Meta-Learning) had been implemented in the code too (i guess by the original author). Did you try them?",
    "480335": "Will it be better to use the bounding box?",
    "479914": "@daisukelab\nDid you use FP32 or FP16 in your code?",
    "479277": "Thanks for sharing… :).\nIn the function prepare_submission you set new_whale_thresh=-1.85.How to confirm the new_whale_thresh value？",
    "478819": "Just wondering if anyone tried combined softmax loss and prototype loss, i.e multi task classification and one shot leaening?",
    "475082": "Greate paper, I'm sure I'm gonna use this for my thesis!",
    "474915": "Thank you @daisukelab, I found ProtoNet too, but I'm new guy to this area, can't adapt to this competition, because the episodic sampling and how to choice *K*-Way *N*-Shot, this is a very good tutorial for me ! ",
    "474768": "Very helpful, Thanks for sharing @daisukelab",
    "474560": "Thanks for sharing。\nIt seems that  your example miss some function？",
    "474188": "Sugoi!. \nThank for sharing. :D. ",
    "477753": "Thanks for sharing... :)",
    "476851": "Thanks!!",
    "476012": "Thanks for sharing!"
  }
}