{
  "id": 74143,
  "title": "Heterogeneity within \"new_whale\"",
  "url": "/competitions/humpback-whale-identification/discussion/74143",
  "author_name": "",
  "post_date": "2018-12-09T05:22:21.659469500Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>If I were to approach the problem with a re-identification model, how should I treat the images identified as new whale? Are they also different from each other?</p>",
  "messages": [
    {
      "id": "435936",
      "postDate": "12/09/2018 05:22:21",
      "content": "<p>If I were to approach the problem with a re-identification model, how should I treat the images identified as new whale? Are they also different from each other?</p>",
      "rawMarkdown": "If I were to approach the problem with a re-identification model, how should I treat the images identified as new whale? Are they also different from each other?",
      "votes": null
    },
    {
      "id": "435946",
      "postDate": "12/09/2018 05:46:16",
      "content": "<p>I looked at these images, it seems to me there are many whales within this category.</p>",
      "rawMarkdown": "I looked at these images, it seems to me there are many whales within this category.",
      "votes": null
    },
    {
      "id": "436033",
      "postDate": "12/09/2018 11:40:46",
      "content": "<p>Hi!</p>\n\n<p>Two possibilities I've been thinking about are: (i) Consider each \"new whale\" sample as a unique identity, (ii) Consider the samples tagged as \"new whale\" a definition of the maximum distance between a particular identity and the space of other identities.</p>\n\n<p>Although both approaches could be similar in concept, they entail differences in terms of modeling.</p>\n\n<p>Hope this was useful!</p>",
      "rawMarkdown": "Hi!\n\nTwo possibilities I've been thinking about are: (i) Consider each \"new whale\" sample as a unique identity, (ii) Consider the samples tagged as \"new whale\" a definition of the maximum distance between a particular identity and the space of other identities.\n\nAlthough both approaches could be similar in concept, they entail differences in terms of modeling.\n\nHope this was useful!",
      "votes": null
    },
    {
      "id": "436145",
      "postDate": "12/09/2018 17:20:04",
      "content": "<p>And what constitutes a distance between two whales? </p>",
      "rawMarkdown": "And what constitutes a distance between two whales?",
      "votes": null
    },
    {
      "id": "436169",
      "postDate": "12/09/2018 19:07:48",
      "content": "<p>It could be the cosine distance between two normalized embeddings. For an initial reference on the concept, you could use this video (seen on <a href=\"https://www.kaggle.com/ashishpatel26/triplet-loss-network-for-humpback-whale-prediction\">https://www.kaggle.com/ashishpatel26/triplet-loss-network-for-humpback-whale-prediction</a> ):</p>\n\n<p><a href=\"https://youtu.be/LN3RdUFPYyI\">https://youtu.be/LN3RdUFPYyI</a></p>",
      "rawMarkdown": "It could be the cosine distance between two normalized embeddings. For an initial reference on the concept, you could use this video (seen on https://www.kaggle.com/ashishpatel26/triplet-loss-network-for-humpback-whale-prediction ):\n\nhttps://youtu.be/LN3RdUFPYyI",
      "votes": null
    },
    {
      "id": "436589",
      "postDate": "12/10/2018 15:36:01",
      "content": "<p>Good point!\nA conservative approach is to use only (\"new_whale\", non \"new_whale\") pairs as negative samples and do not use (\"new_whale\", \"new_whale\") pairs as both negative and positive pairs.\nAfter training embedding, we can manually check some (\"new_whale\", \"new_whale\") pairs whose distances are relatively small to see whether they also different from each other or not.</p>",
      "rawMarkdown": "Good point!\nA conservative approach is to use only (\"new_whale\", non \"new_whale\") pairs as negative samples and do not use (\"new_whale\", \"new_whale\") pairs as both negative and positive pairs.\nAfter training embedding, we can manually check some (\"new_whale\", \"new_whale\") pairs whose distances are relatively small to see whether they also different from each other or not.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 435946,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "12/09/2018 05:46:16",
      "content": "<p>I looked at these images, it seems to me there are many whales within this category.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 436033,
      "author_name": "chelus",
      "author_url": "",
      "post_date": "12/09/2018 11:40:46",
      "content": "<p>Hi!</p>\n\n<p>Two possibilities I've been thinking about are: (i) Consider each \"new whale\" sample as a unique identity, (ii) Consider the samples tagged as \"new whale\" a definition of the maximum distance between a particular identity and the space of other identities.</p>\n\n<p>Although both approaches could be similar in concept, they entail differences in terms of modeling.</p>\n\n<p>Hope this was useful!</p>",
      "votes": null,
      "replies": [
        {
          "id": 436145,
          "author_name": "badtyprr",
          "author_url": "",
          "post_date": "12/09/2018 17:20:04",
          "content": "<p>And what constitutes a distance between two whales? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436169,
          "author_name": "chelus",
          "author_url": "",
          "post_date": "12/09/2018 19:07:48",
          "content": "<p>It could be the cosine distance between two normalized embeddings. For an initial reference on the concept, you could use this video (seen on <a href=\"https://www.kaggle.com/ashishpatel26/triplet-loss-network-for-humpback-whale-prediction\">https://www.kaggle.com/ashishpatel26/triplet-loss-network-for-humpback-whale-prediction</a> ):</p>\n\n<p><a href=\"https://youtu.be/LN3RdUFPYyI\">https://youtu.be/LN3RdUFPYyI</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436589,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "12/10/2018 15:36:01",
      "content": "<p>Good point!\nA conservative approach is to use only (\"new_whale\", non \"new_whale\") pairs as negative samples and do not use (\"new_whale\", \"new_whale\") pairs as both negative and positive pairs.\nAfter training embedding, we can manually check some (\"new_whale\", \"new_whale\") pairs whose distances are relatively small to see whether they also different from each other or not.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "435936": "If I were to approach the problem with a re-identification model, how should I treat the images identified as new whale? Are they also different from each other?",
    "435946": "I looked at these images, it seems to me there are many whales within this category.",
    "436033": "Hi!\n\nTwo possibilities I've been thinking about are: (i) Consider each \"new whale\" sample as a unique identity, (ii) Consider the samples tagged as \"new whale\" a definition of the maximum distance between a particular identity and the space of other identities.\n\nAlthough both approaches could be similar in concept, they entail differences in terms of modeling.\n\nHope this was useful!",
    "436145": "And what constitutes a distance between two whales?",
    "436169": "It could be the cosine distance between two normalized embeddings. For an initial reference on the concept, you could use this video (seen on https://www.kaggle.com/ashishpatel26/triplet-loss-network-for-humpback-whale-prediction ):\n\nhttps://youtu.be/LN3RdUFPYyI",
    "436589": "Good point!\nA conservative approach is to use only (\"new_whale\", non \"new_whale\") pairs as negative samples and do not use (\"new_whale\", \"new_whale\") pairs as both negative and positive pairs.\nAfter training embedding, we can manually check some (\"new_whale\", \"new_whale\") pairs whose distances are relatively small to see whether they also different from each other or not."
  },
  "source": "meta"
}