{
  "id": 170214,
  "title": "Alternative to Softmax?",
  "url": "/competitions/landmark-retrieval-2020/discussion/170214",
  "author_name": "",
  "post_date": "2020-07-26T21:59:01.257735600Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So we have like 80k classes. </p>\n\n<p>A softmax activation from a 2048 embedding to only 10k classes is 20 million params in the network. This seems silly. </p>\n\n<p>Are there any alternatives to this? I'd rather not try triplet loss. </p>",
  "messages": [
    {
      "id": "946836",
      "postDate": "07/26/2020 21:59:01",
      "content": "<p>So we have like 80k classes. </p>\n\n<p>A softmax activation from a 2048 embedding to only 10k classes is 20 million params in the network. This seems silly. </p>\n\n<p>Are there any alternatives to this? I'd rather not try triplet loss. </p>",
      "rawMarkdown": "So we have like 80k classes. \n\nA softmax activation from a 2048 embedding to only 10k classes is 20 million params in the network. This seems silly. \n\nAre there any alternatives to this? I'd rather not try triplet loss.",
      "votes": null
    },
    {
      "id": "946955",
      "postDate": "07/27/2020 01:16:34",
      "content": "<p>I tried triplet loss but it didn't converge and is slow.</p>",
      "rawMarkdown": "I tried triplet loss but it didn't converge and is slow.",
      "votes": null
    },
    {
      "id": "948132",
      "postDate": "07/27/2020 17:25:36",
      "content": "<p>Yeah I read that triplet loss is not very effective. Guess we're stuck with ArcFace softmax</p>",
      "rawMarkdown": "Yeah I read that triplet loss is not very effective. Guess we're stuck with ArcFace softmax",
      "votes": null
    },
    {
      "id": "948148",
      "postDate": "07/27/2020 17:33:00",
      "content": "<p>This can be a good try: <a href=\"https://www.kaggle.com/akensert/glr-triplet-semi-hard-loss-with-distributed-tf\">https://www.kaggle.com/akensert/glr-triplet-semi-hard-loss-with-distributed-tf</a></p>",
      "rawMarkdown": "This can be a good try: https://www.kaggle.com/akensert/glr-triplet-semi-hard-loss-with-distributed-tf",
      "votes": null
    },
    {
      "id": "951533",
      "postDate": "07/30/2020 07:38:54",
      "content": "<p>From what I've read there are only 3 ways of learning deep global features:\n- image-level classification with Softmax (usually with ArcFace projection before)\n- n-tuplet learning (n=2: siamese, n=3: triplet, n&gt;3 contrastive learning, check out <a href=\"https://arxiv.org/abs/1711.02512\">https://arxiv.org/abs/1711.02512</a>), which implies hard sample mining and is difficult to optimize properly\n- listwise loss (<a href=\"https://arxiv.org/abs/1906.07589\">https://arxiv.org/abs/1906.07589</a>), code is available here (<a href=\"https://github.com/almazan/deep-image-retrieval\">https://github.com/almazan/deep-image-retrieval</a>), intuitively the best but results are somewhat disappointing</p>",
      "rawMarkdown": "From what I've read there are only 3 ways of learning deep global features:\n- image-level classification with Softmax (usually with ArcFace projection before)\n- n-tuplet learning (n=2: siamese, n=3: triplet, n&gt;3 contrastive learning, check out https://arxiv.org/abs/1711.02512), which implies hard sample mining and is difficult to optimize properly\n- listwise loss (https://arxiv.org/abs/1906.07589), code is available here (https://github.com/almazan/deep-image-retrieval), intuitively the best but results are somewhat disappointing",
      "votes": null
    },
    {
      "id": "953579",
      "postDate": "07/31/2020 22:17:41",
      "content": "<p>hierarchical softmax?</p>",
      "rawMarkdown": "hierarchical softmax?",
      "votes": null
    },
    {
      "id": "953581",
      "postDate": "07/31/2020 22:21:52",
      "content": "<p>or, you can encode the output in few categories, say 100. and then based on which one of the 100 categories relais, you can get more deep in that categories.\nsay, first level categories -&gt; bridges, constructions, ...\nsecond level categories, -&gt; if bridge, what name actually is</p>",
      "rawMarkdown": "or, you can encode the output in few categories, say 100. and then based on which one of the 100 categories relais, you can get more deep in that categories.\nsay, first level categories -&gt; bridges, constructions, ...\nsecond level categories, -&gt; if bridge, what name actually is",
      "votes": null
    },
    {
      "id": "954523",
      "postDate": "08/01/2020 19:55:19",
      "content": "<p>that's a good idea, the challenge is getting the labelled data</p>",
      "rawMarkdown": "that's a good idea, the challenge is getting the labelled data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 946955,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "07/27/2020 01:16:34",
      "content": "<p>I tried triplet loss but it didn't converge and is slow.</p>",
      "votes": null,
      "replies": [
        {
          "id": 948132,
          "author_name": "hooong",
          "author_url": "",
          "post_date": "07/27/2020 17:25:36",
          "content": "<p>Yeah I read that triplet loss is not very effective. Guess we're stuck with ArcFace softmax</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948148,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "07/27/2020 17:33:00",
          "content": "<p>This can be a good try: <a href=\"https://www.kaggle.com/akensert/glr-triplet-semi-hard-loss-with-distributed-tf\">https://www.kaggle.com/akensert/glr-triplet-semi-hard-loss-with-distributed-tf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 951533,
      "author_name": "dgominski",
      "author_url": "",
      "post_date": "07/30/2020 07:38:54",
      "content": "<p>From what I've read there are only 3 ways of learning deep global features:\n- image-level classification with Softmax (usually with ArcFace projection before)\n- n-tuplet learning (n=2: siamese, n=3: triplet, n&gt;3 contrastive learning, check out <a href=\"https://arxiv.org/abs/1711.02512\">https://arxiv.org/abs/1711.02512</a>), which implies hard sample mining and is difficult to optimize properly\n- listwise loss (<a href=\"https://arxiv.org/abs/1906.07589\">https://arxiv.org/abs/1906.07589</a>), code is available here (<a href=\"https://github.com/almazan/deep-image-retrieval\">https://github.com/almazan/deep-image-retrieval</a>), intuitively the best but results are somewhat disappointing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953579,
      "author_name": "alfredomaussa",
      "author_url": "",
      "post_date": "07/31/2020 22:17:41",
      "content": "<p>hierarchical softmax?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953581,
      "author_name": "alfredomaussa",
      "author_url": "",
      "post_date": "07/31/2020 22:21:52",
      "content": "<p>or, you can encode the output in few categories, say 100. and then based on which one of the 100 categories relais, you can get more deep in that categories.\nsay, first level categories -&gt; bridges, constructions, ...\nsecond level categories, -&gt; if bridge, what name actually is</p>",
      "votes": null,
      "replies": [
        {
          "id": 954523,
          "author_name": "hooong",
          "author_url": "",
          "post_date": "08/01/2020 19:55:19",
          "content": "<p>that's a good idea, the challenge is getting the labelled data</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "946836": "So we have like 80k classes. \n\nA softmax activation from a 2048 embedding to only 10k classes is 20 million params in the network. This seems silly. \n\nAre there any alternatives to this? I'd rather not try triplet loss.",
    "946955": "I tried triplet loss but it didn't converge and is slow.",
    "948132": "Yeah I read that triplet loss is not very effective. Guess we're stuck with ArcFace softmax",
    "948148": "This can be a good try: https://www.kaggle.com/akensert/glr-triplet-semi-hard-loss-with-distributed-tf",
    "951533": "From what I've read there are only 3 ways of learning deep global features:\n- image-level classification with Softmax (usually with ArcFace projection before)\n- n-tuplet learning (n=2: siamese, n=3: triplet, n&gt;3 contrastive learning, check out https://arxiv.org/abs/1711.02512), which implies hard sample mining and is difficult to optimize properly\n- listwise loss (https://arxiv.org/abs/1906.07589), code is available here (https://github.com/almazan/deep-image-retrieval), intuitively the best but results are somewhat disappointing",
    "953579": "hierarchical softmax?",
    "953581": "or, you can encode the output in few categories, say 100. and then based on which one of the 100 categories relais, you can get more deep in that categories.\nsay, first level categories -&gt; bridges, constructions, ...\nsecond level categories, -&gt; if bridge, what name actually is",
    "954523": "that's a good idea, the challenge is getting the labelled data"
  },
  "source": "meta"
}