{
  "id": 305746,
  "title": "How to know if model trained with Triplet Loss can generate \"good\" embeddings?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/305746",
  "author_name": "",
  "post_date": "2022-02-06T18:33:06.029449400Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm pretty new to Deep Learning, so I don't have much experience.<br>\nI read some materials across the web and found that Triplet Loss (more precisely Semi Hard Triplet Loss) can achieve good results for recognition tasks. I decided to use <code>margin = 1.0</code>.</p>\n<p>After running the model with 100 epochs, the loss goes from 1.0 to 0.2 for both training and validation loss. Therefore I used the model to predict embeddings and used those embeddings as features to run a classification model like KNN or SVM. The problem is that I found very low MAP@5 (~0.004).</p>\n<p>Seems that the embeddings are not being useful at all to run a classification model, even after big decrease in train/val losses and a possible \"convergence\".</p>\n<p>I found in the internet that Triplet Loss networks are bit tricky to train. <br>\nAny tips on this?<br>\nHow would I know if the embeddings are \"useful\" to discriminate the classes?</p>",
  "messages": [
    {
      "id": "1678716",
      "postDate": "02/06/2022 18:33:06",
      "content": "<p>I'm pretty new to Deep Learning, so I don't have much experience.<br>\nI read some materials across the web and found that Triplet Loss (more precisely Semi Hard Triplet Loss) can achieve good results for recognition tasks. I decided to use <code>margin = 1.0</code>.</p>\n<p>After running the model with 100 epochs, the loss goes from 1.0 to 0.2 for both training and validation loss. Therefore I used the model to predict embeddings and used those embeddings as features to run a classification model like KNN or SVM. The problem is that I found very low MAP@5 (~0.004).</p>\n<p>Seems that the embeddings are not being useful at all to run a classification model, even after big decrease in train/val losses and a possible \"convergence\".</p>\n<p>I found in the internet that Triplet Loss networks are bit tricky to train. <br>\nAny tips on this?<br>\nHow would I know if the embeddings are \"useful\" to discriminate the classes?</p>",
      "rawMarkdown": "I'm pretty new to Deep Learning, so I don't have much experience.\nI read some materials across the web and found that Triplet Loss (more precisely Semi Hard Triplet Loss) can achieve good results for recognition tasks. I decided to use `margin = 1.0`.\n\nAfter running the model with 100 epochs, the loss goes from 1.0 to 0.2 for both training and validation loss. Therefore I used the model to predict embeddings and used those embeddings as features to run a classification model like KNN or SVM. The problem is that I found very low MAP@5 (~0.004).\n\nSeems that the embeddings are not being useful at all to run a classification model, even after big decrease in train/val losses and a possible \"convergence\".\n\nI found in the internet that Triplet Loss networks are bit tricky to train. \nAny tips on this?\nHow would I know if the embeddings are \"useful\" to discriminate the classes?",
      "votes": null
    },
    {
      "id": "1678820",
      "postDate": "02/06/2022 20:00:52",
      "content": "<p>This is very big qestion.</p>\n<blockquote>\n  <p>How would I know if the embeddings are \"useful\" to discriminate the classes</p>\n</blockquote>\n<p>It should increase your metric/KPI :)</p>\n<blockquote>\n  <p>Any tips on this?</p>\n</blockquote>\n<p>Loss value in triplet loss is not really useful, if you look at one-sample loss - very easy get 0 for not hard samples. The main problem of triplet loss - hard samples mining, there are many different technic and they are close one to magic/euristic.</p>\n<p>Also, KNN and SVM may be is not good choice for comparing embeddings, I'd recommend to see at dot product between embeddings or something like this</p>",
      "rawMarkdown": "This is very big qestion.\n\n> How would I know if the embeddings are \"useful\" to discriminate the classes\n\nIt should increase your metric/KPI :)\n\n> Any tips on this?\n\nLoss value in triplet loss is not really useful, if you look at one-sample loss - very easy get 0 for not hard samples. The main problem of triplet loss - hard samples mining, there are many different technic and they are close one to magic/euristic.\n\nAlso, KNN and SVM may be is not good choice for comparing embeddings, I'd recommend to see at dot product between embeddings or something like this",
      "votes": null
    },
    {
      "id": "1678830",
      "postDate": "02/06/2022 20:09:03",
      "content": "<p>Thanks for your help. Calculating dot product seems a good idea. </p>",
      "rawMarkdown": "Thanks for your help. Calculating dot product seems a good idea.",
      "votes": null
    },
    {
      "id": "1684312",
      "postDate": "02/10/2022 12:30:19",
      "content": "<p>you can visualize embeddings by using TSNE by compressing your embeddings into a two dimensional space.<br>\n<code>from sklearn.manifold import TSNE</code><br>\n<code>embeddings = TSNE(n_components=2,init = 'pca').fit_transform(embeddings)</code><br>\n<code>sns.scatterplot(x=embeddings[:,0],y=embeddings[:,1],hue= labels  ,legend='full')</code></p>",
      "rawMarkdown": "you can visualize embeddings by using TSNE by compressing your embeddings into a two dimensional space.\n`from sklearn.manifold import TSNE`\n`embeddings = TSNE(n_components=2,init = 'pca').fit_transform(embeddings)`\n`sns.scatterplot(x=embeddings[:,0],y=embeddings[:,1],hue= labels  ,legend='full')`",
      "votes": null
    },
    {
      "id": "1684420",
      "postDate": "02/10/2022 13:37:03",
      "content": "<p>Thanks for the tip! That makes sense</p>",
      "rawMarkdown": "Thanks for the tip! That makes sense",
      "votes": null
    },
    {
      "id": "1748261",
      "postDate": "04/07/2022 12:31:44",
      "content": "<p>Training triplet loss sometimes will crash, which means input different images while output same embedding. You can use some other sampling strategies, e.g. distance weighted sampling…</p>",
      "rawMarkdown": "Training triplet loss sometimes will crash, which means input different images while output same embedding. You can use some other sampling strategies, e.g. distance weighted sampling...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1678820,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "02/06/2022 20:00:52",
      "content": "<p>This is very big qestion.</p>\n<blockquote>\n  <p>How would I know if the embeddings are \"useful\" to discriminate the classes</p>\n</blockquote>\n<p>It should increase your metric/KPI :)</p>\n<blockquote>\n  <p>Any tips on this?</p>\n</blockquote>\n<p>Loss value in triplet loss is not really useful, if you look at one-sample loss - very easy get 0 for not hard samples. The main problem of triplet loss - hard samples mining, there are many different technic and they are close one to magic/euristic.</p>\n<p>Also, KNN and SVM may be is not good choice for comparing embeddings, I'd recommend to see at dot product between embeddings or something like this</p>",
      "votes": null,
      "replies": [
        {
          "id": 1678830,
          "author_name": "igorkf",
          "author_url": "",
          "post_date": "02/06/2022 20:09:03",
          "content": "<p>Thanks for your help. Calculating dot product seems a good idea. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1684312,
      "author_name": "mrabetyacine",
      "author_url": "",
      "post_date": "02/10/2022 12:30:19",
      "content": "<p>you can visualize embeddings by using TSNE by compressing your embeddings into a two dimensional space.<br>\n<code>from sklearn.manifold import TSNE</code><br>\n<code>embeddings = TSNE(n_components=2,init = 'pca').fit_transform(embeddings)</code><br>\n<code>sns.scatterplot(x=embeddings[:,0],y=embeddings[:,1],hue= labels  ,legend='full')</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1684420,
          "author_name": "igorkf",
          "author_url": "",
          "post_date": "02/10/2022 13:37:03",
          "content": "<p>Thanks for the tip! That makes sense</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1748261,
      "author_name": "jisongxie",
      "author_url": "",
      "post_date": "04/07/2022 12:31:44",
      "content": "<p>Training triplet loss sometimes will crash, which means input different images while output same embedding. You can use some other sampling strategies, e.g. distance weighted sampling…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1678716": "I'm pretty new to Deep Learning, so I don't have much experience.\nI read some materials across the web and found that Triplet Loss (more precisely Semi Hard Triplet Loss) can achieve good results for recognition tasks. I decided to use `margin = 1.0`.\n\nAfter running the model with 100 epochs, the loss goes from 1.0 to 0.2 for both training and validation loss. Therefore I used the model to predict embeddings and used those embeddings as features to run a classification model like KNN or SVM. The problem is that I found very low MAP@5 (~0.004).\n\nSeems that the embeddings are not being useful at all to run a classification model, even after big decrease in train/val losses and a possible \"convergence\".\n\nI found in the internet that Triplet Loss networks are bit tricky to train. \nAny tips on this?\nHow would I know if the embeddings are \"useful\" to discriminate the classes?",
    "1678820": "This is very big qestion.\n\n> How would I know if the embeddings are \"useful\" to discriminate the classes\n\nIt should increase your metric/KPI :)\n\n> Any tips on this?\n\nLoss value in triplet loss is not really useful, if you look at one-sample loss - very easy get 0 for not hard samples. The main problem of triplet loss - hard samples mining, there are many different technic and they are close one to magic/euristic.\n\nAlso, KNN and SVM may be is not good choice for comparing embeddings, I'd recommend to see at dot product between embeddings or something like this",
    "1678830": "Thanks for your help. Calculating dot product seems a good idea.",
    "1684312": "you can visualize embeddings by using TSNE by compressing your embeddings into a two dimensional space.\n`from sklearn.manifold import TSNE`\n`embeddings = TSNE(n_components=2,init = 'pca').fit_transform(embeddings)`\n`sns.scatterplot(x=embeddings[:,0],y=embeddings[:,1],hue= labels  ,legend='full')`",
    "1684420": "Thanks for the tip! That makes sense",
    "1748261": "Training triplet loss sometimes will crash, which means input different images while output same embedding. You can use some other sampling strategies, e.g. distance weighted sampling..."
  },
  "source": "meta"
}