{
  "id": 318307,
  "title": "I did terribly on this competition.",
  "url": "/competitions/happy-whale-and-dolphin/discussion/318307",
  "author_name": "",
  "post_date": "2022-04-11T17:44:26.505647700Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The Happywhale Competition this year had images from 28 different species, with over 15K individuals. Some photos are distant images of the backs of humpback whales and some are closeups of spinner dolphin dorsal fins. </p>\n<p>For a baseline model, I just threw <em>all</em> the training data into the pot, trained it overnight, and tried that. Well, I wasn’t surprised that I had terrible results: generating the transform that accurately locates photos in one of 15K compartments is slow work! </p>\n<p>I was confused that, even in the early days of the contest, the leaderboard was filling up with people getting 60-70% accuracy. My first baseline had about 12% accuracy! </p>\n<h2>Simple! Let’s try complexity!</h2>\n<p>I quickly (maybe hastily) concluded that the way forward was <em>many models</em>. Instead of throwing all the data into one model and trying to generate 15K tight clusters, I’d use multiple layers: </p>\n<ul>\n<li>I trained a model to find, in the photo, a Region Of Interest aka tried to crop the image tightly around the whale</li>\n<li>I trained a model to classify the camera viewpoint (I had 6 categories: anterior, posterior, port full, starboard full, port dorsal fin, starboard dorsal fin)</li>\n<li>I trained a model to identify the species or species-type in the photo (by species-type I mean I had categories such as “blackfish” that included several species of small toothed whales)  </li>\n</ul>\n<p>How complex was my multimodel approach? </p>\n<p><img src=\"https://www.knowing.net/images/happywhale_multimodels.png\" alt=\"multimodel architecture\"></p>\n<p><em>sigh</em></p>\n<p>Realize that this architecture requires <code>number of viewpoints * number of species</code> separate embedding models! </p>\n<p>Which might have been tractable had I developed the ability to pretty-rapidly generate good embeddings for a particular viewpoint and species! It would have been too much for me to do locally, but I was planning on building an Azure ML Pipeline, spending some money to spin up a half dozen training machines, and basking in glory. All I had to do was develop some kick-ass embedding code.</p>\n<p>And that. Just. Didn’t. Work. </p>\n<h2>Am I getting close? Who knows?</h2>\n<p>The big problem with embeddings and the idea that “every individual has an area in high-dimensional space” is that when you generate a point in that high-dimensional space, it’s hard to say if it’s near or far from where you ultimately want it. </p>\n<p>For one thing, “distance” can mean two common things when working with embeddings: </p>\n<ul>\n<li>Euclidean distance, which is the distance in space between the two points</li>\n<li>Cosine distance, which is the difference in the angles to the points, as measured from the origin</li>\n</ul>\n<p>Since this is a story of failure, I’m not going to try to justify why I was biased towards Cosine distance. But in the end, I tried almost everything with <em>both</em>. (And got crappy results.)</p>\n<p>No matter which distance measure you choose, at the moment you generate an embedding for your target photo, what do you compare it to? <em>Eventually</em> you want it to be “near other photos of the same individual” but in the <em>moment</em>, what do you know? </p>\n<h2>Overwhelmed by triplets</h2>\n<p>What you can do is: instead of just the target photo, you have a <em>triplet</em> of photos: </p>\n<ul>\n<li>the target photo, </li>\n<li>another photo of the same individual (a positive match), and </li>\n<li>a photo you know is <em>not</em> the same individual (a negative match)</li>\n</ul>\n<p>Generate embeddings for each in the triplet. Now, early in the training process these will presumably spread out all over your n-dimensional space. <em>But</em> you know that you’d <em>rather</em> have the positive match close and the negative match far. So you can change the weights to further that goal. </p>\n<p>And that, dear friends, is called <em>triplet loss</em> and it’s been my world for the past month. </p>\n<p>Again, since this is a story of failure, I shouldn’t lecture on “tips and tricks,” but I’ll just point out one obvious challenge with triplet loss: <em>most</em> negative matches are going to be pretty obvious. Similarly, since cetacean photography often has images taken 1/10th of a second after the previous and some individuals have obvious features (mutilated fin edges, often), <em>some</em> positive matches are also obvious. What you <em>want</em> to train on are “the positive match that looks the least like this photo” and “the negative match that looks the most like this photo.” Trying to find <em>those</em> photos is called <em>triplet mining</em> and, yeesh, lemme’ tell ya’. Or, rather, what can I tell you? It didn’t work for me. </p>\n<p>tl;dr: I spent a month trying to apply an embeddings-based approach to cetacean reidentification and never had good results. I tried many standard things -- Siamese networks, Euclidean distance, Cosine distance, alternatives to ArcFace loss, soft and hard triplet mining, etc. In all cases, I would get good results with simplified datasets (MNIST-Fashion, a hand-picked set of individual cetaceans) but when I applied it to the Happywhale training data, I had very poor results. </p>",
  "messages": [
    {
      "id": "1752413",
      "postDate": "04/11/2022 17:44:26",
      "content": "<p>The Happywhale Competition this year had images from 28 different species, with over 15K individuals. Some photos are distant images of the backs of humpback whales and some are closeups of spinner dolphin dorsal fins. </p>\n<p>For a baseline model, I just threw <em>all</em> the training data into the pot, trained it overnight, and tried that. Well, I wasn’t surprised that I had terrible results: generating the transform that accurately locates photos in one of 15K compartments is slow work! </p>\n<p>I was confused that, even in the early days of the contest, the leaderboard was filling up with people getting 60-70% accuracy. My first baseline had about 12% accuracy! </p>\n<h2>Simple! Let’s try complexity!</h2>\n<p>I quickly (maybe hastily) concluded that the way forward was <em>many models</em>. Instead of throwing all the data into one model and trying to generate 15K tight clusters, I’d use multiple layers: </p>\n<ul>\n<li>I trained a model to find, in the photo, a Region Of Interest aka tried to crop the image tightly around the whale</li>\n<li>I trained a model to classify the camera viewpoint (I had 6 categories: anterior, posterior, port full, starboard full, port dorsal fin, starboard dorsal fin)</li>\n<li>I trained a model to identify the species or species-type in the photo (by species-type I mean I had categories such as “blackfish” that included several species of small toothed whales)  </li>\n</ul>\n<p>How complex was my multimodel approach? </p>\n<p><img src=\"https://www.knowing.net/images/happywhale_multimodels.png\" alt=\"multimodel architecture\"></p>\n<p><em>sigh</em></p>\n<p>Realize that this architecture requires <code>number of viewpoints * number of species</code> separate embedding models! </p>\n<p>Which might have been tractable had I developed the ability to pretty-rapidly generate good embeddings for a particular viewpoint and species! It would have been too much for me to do locally, but I was planning on building an Azure ML Pipeline, spending some money to spin up a half dozen training machines, and basking in glory. All I had to do was develop some kick-ass embedding code.</p>\n<p>And that. Just. Didn’t. Work. </p>\n<h2>Am I getting close? Who knows?</h2>\n<p>The big problem with embeddings and the idea that “every individual has an area in high-dimensional space” is that when you generate a point in that high-dimensional space, it’s hard to say if it’s near or far from where you ultimately want it. </p>\n<p>For one thing, “distance” can mean two common things when working with embeddings: </p>\n<ul>\n<li>Euclidean distance, which is the distance in space between the two points</li>\n<li>Cosine distance, which is the difference in the angles to the points, as measured from the origin</li>\n</ul>\n<p>Since this is a story of failure, I’m not going to try to justify why I was biased towards Cosine distance. But in the end, I tried almost everything with <em>both</em>. (And got crappy results.)</p>\n<p>No matter which distance measure you choose, at the moment you generate an embedding for your target photo, what do you compare it to? <em>Eventually</em> you want it to be “near other photos of the same individual” but in the <em>moment</em>, what do you know? </p>\n<h2>Overwhelmed by triplets</h2>\n<p>What you can do is: instead of just the target photo, you have a <em>triplet</em> of photos: </p>\n<ul>\n<li>the target photo, </li>\n<li>another photo of the same individual (a positive match), and </li>\n<li>a photo you know is <em>not</em> the same individual (a negative match)</li>\n</ul>\n<p>Generate embeddings for each in the triplet. Now, early in the training process these will presumably spread out all over your n-dimensional space. <em>But</em> you know that you’d <em>rather</em> have the positive match close and the negative match far. So you can change the weights to further that goal. </p>\n<p>And that, dear friends, is called <em>triplet loss</em> and it’s been my world for the past month. </p>\n<p>Again, since this is a story of failure, I shouldn’t lecture on “tips and tricks,” but I’ll just point out one obvious challenge with triplet loss: <em>most</em> negative matches are going to be pretty obvious. Similarly, since cetacean photography often has images taken 1/10th of a second after the previous and some individuals have obvious features (mutilated fin edges, often), <em>some</em> positive matches are also obvious. What you <em>want</em> to train on are “the positive match that looks the least like this photo” and “the negative match that looks the most like this photo.” Trying to find <em>those</em> photos is called <em>triplet mining</em> and, yeesh, lemme’ tell ya’. Or, rather, what can I tell you? It didn’t work for me. </p>\n<p>tl;dr: I spent a month trying to apply an embeddings-based approach to cetacean reidentification and never had good results. I tried many standard things -- Siamese networks, Euclidean distance, Cosine distance, alternatives to ArcFace loss, soft and hard triplet mining, etc. In all cases, I would get good results with simplified datasets (MNIST-Fashion, a hand-picked set of individual cetaceans) but when I applied it to the Happywhale training data, I had very poor results. </p>",
      "rawMarkdown": "The Happywhale Competition this year had images from 28 different species, with over 15K individuals. Some photos are distant images of the backs of humpback whales and some are closeups of spinner dolphin dorsal fins. \n\nFor a baseline model, I just threw _all_ the training data into the pot, trained it overnight, and tried that. Well, I wasn’t surprised that I had terrible results: generating the transform that accurately locates photos in one of 15K compartments is slow work! \n\nI was confused that, even in the early days of the contest, the leaderboard was filling up with people getting 60-70% accuracy. My first baseline had about 12% accuracy! \n\n## Simple! Let’s try complexity!\n\nI quickly (maybe hastily) concluded that the way forward was _many models_. Instead of throwing all the data into one model and trying to generate 15K tight clusters, I’d use multiple layers: \n\n*  I trained a model to find, in the photo, a Region Of Interest aka tried to crop the image tightly around the whale\n* I trained a model to classify the camera viewpoint (I had 6 categories: anterior, posterior, port full, starboard full, port dorsal fin, starboard dorsal fin)\n* I trained a model to identify the species or species-type in the photo (by species-type I mean I had categories such as “blackfish” that included several species of small toothed whales)  \n\nHow complex was my multimodel approach? \n\n<img src=\"https://www.knowing.net/images/happywhale_multimodels.png\" alt=\"multimodel architecture\" width=\"800\"/>\n\n\n*sigh*\n\nRealize that this architecture requires `number of viewpoints * number of species` separate embedding models! \n\nWhich might have been tractable had I developed the ability to pretty-rapidly generate good embeddings for a particular viewpoint and species! It would have been too much for me to do locally, but I was planning on building an Azure ML Pipeline, spending some money to spin up a half dozen training machines, and basking in glory. All I had to do was develop some kick-ass embedding code.\n\nAnd that. Just. Didn’t. Work. \n\n## Am I getting close? Who knows?\n\nThe big problem with embeddings and the idea that “every individual has an area in high-dimensional space” is that when you generate a point in that high-dimensional space, it’s hard to say if it’s near or far from where you ultimately want it. \n\nFor one thing, “distance” can mean two common things when working with embeddings: \n\n* Euclidean distance, which is the distance in space between the two points\n* Cosine distance, which is the difference in the angles to the points, as measured from the origin\n\nSince this is a story of failure, I’m not going to try to justify why I was biased towards Cosine distance. But in the end, I tried almost everything with _both_. (And got crappy results.)\n\nNo matter which distance measure you choose, at the moment you generate an embedding for your target photo, what do you compare it to? _Eventually_ you want it to be “near other photos of the same individual” but in the _moment_, what do you know? \n\n## Overwhelmed by triplets\n\nWhat you can do is: instead of just the target photo, you have a _triplet_ of photos: \n\n* the target photo, \n* another photo of the same individual (a positive match), and \n* a photo you know is _not_ the same individual (a negative match)\n\nGenerate embeddings for each in the triplet. Now, early in the training process these will presumably spread out all over your n-dimensional space. *But* you know that you’d _rather_ have the positive match close and the negative match far. So you can change the weights to further that goal. \n\nAnd that, dear friends, is called _triplet loss_ and it’s been my world for the past month. \n\nAgain, since this is a story of failure, I shouldn’t lecture on “tips and tricks,” but I’ll just point out one obvious challenge with triplet loss: _most_ negative matches are going to be pretty obvious. Similarly, since cetacean photography often has images taken 1/10th of a second after the previous and some individuals have obvious features (mutilated fin edges, often), _some_ positive matches are also obvious. What you _want_ to train on are “the positive match that looks the least like this photo” and “the negative match that looks the most like this photo.” Trying to find _those_ photos is called _triplet mining_ and, yeesh, lemme’ tell ya’. Or, rather, what can I tell you? It didn’t work for me. \n\n \ntl;dr: I spent a month trying to apply an embeddings-based approach to cetacean reidentification and never had good results. I tried many standard things -- Siamese networks, Euclidean distance, Cosine distance, alternatives to ArcFace loss, soft and hard triplet mining, etc. In all cases, I would get good results with simplified datasets (MNIST-Fashion, a hand-picked set of individual cetaceans) but when I applied it to the Happywhale training data, I had very poor results.",
      "votes": null
    },
    {
      "id": "1752506",
      "postDate": "04/11/2022 20:16:31",
      "content": "<p>Sounds like you tried a lot of things…and learned may'be a lot in the proces?</p>\n<p>So if that is the case .. than did you really do that terrible??</p>",
      "rawMarkdown": "Sounds like you tried a lot of things...and learned may'be a lot in the proces?\n\nSo if that is the case .. than did you really do that terrible??",
      "votes": null
    },
    {
      "id": "1752523",
      "postDate": "04/11/2022 21:22:48",
      "content": "<p>I hope that once the leaders do their postmortems I'll be able to revisit things and get a handle on what I did wrong. I certainly feel like I tackled the problem from a lot of different angles, so it will be interesting to see if I was just missing some crucial ingredient or if I was too impatient and should have just trained slower and longer. </p>\n<p>One thing I'd be very interested in would be not even the final solutions but the manner in which people were able to get 60-70% baselines so early in the process. </p>",
      "rawMarkdown": "I hope that once the leaders do their postmortems I'll be able to revisit things and get a handle on what I did wrong. I certainly feel like I tackled the problem from a lot of different angles, so it will be interesting to see if I was just missing some crucial ingredient or if I was too impatient and should have just trained slower and longer. \n\nOne thing I'd be very interested in would be not even the final solutions but the manner in which people were able to get 60-70% baselines so early in the process.",
      "votes": null
    },
    {
      "id": "1752534",
      "postDate": "04/11/2022 21:51:01",
      "content": "<p>Meh, fit-predict is much more easier :)</p>\n<p>Honestly, you did a lot of work and got the most important thing - experience. But if you have bad results for a lot of methods - may be problem at step before - data? Here we have kernels with quite good score on ArcFace, I'd recommend repeat it and go back step-by-step to your solution.</p>\n<p>Simple is better than complex (C) The Zen of Python. Start to do complex stuff after you have good simple.</p>",
      "rawMarkdown": "Meh, fit-predict is much more easier :)\n\nHonestly, you did a lot of work and got the most important thing - experience. But if you have bad results for a lot of methods - may be problem at step before - data? Here we have kernels with quite good score on ArcFace, I'd recommend repeat it and go back step-by-step to your solution.\n\nSimple is better than complex (C) The Zen of Python. Start to do complex stuff after you have good simple.",
      "votes": null
    },
    {
      "id": "1752660",
      "postDate": "04/12/2022 03:01:59",
      "content": "<p>Follow the giant, then you go furher.</p>",
      "rawMarkdown": "Follow the giant, then you go furher.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1752506,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/11/2022 20:16:31",
      "content": "<p>Sounds like you tried a lot of things…and learned may'be a lot in the proces?</p>\n<p>So if that is the case .. than did you really do that terrible??</p>",
      "votes": null,
      "replies": [
        {
          "id": 1752523,
          "author_name": "lobrien",
          "author_url": "",
          "post_date": "04/11/2022 21:22:48",
          "content": "<p>I hope that once the leaders do their postmortems I'll be able to revisit things and get a handle on what I did wrong. I certainly feel like I tackled the problem from a lot of different angles, so it will be interesting to see if I was just missing some crucial ingredient or if I was too impatient and should have just trained slower and longer. </p>\n<p>One thing I'd be very interested in would be not even the final solutions but the manner in which people were able to get 60-70% baselines so early in the process. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1752534,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "04/11/2022 21:51:01",
      "content": "<p>Meh, fit-predict is much more easier :)</p>\n<p>Honestly, you did a lot of work and got the most important thing - experience. But if you have bad results for a lot of methods - may be problem at step before - data? Here we have kernels with quite good score on ArcFace, I'd recommend repeat it and go back step-by-step to your solution.</p>\n<p>Simple is better than complex (C) The Zen of Python. Start to do complex stuff after you have good simple.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1752660,
      "author_name": "shawndong98",
      "author_url": "",
      "post_date": "04/12/2022 03:01:59",
      "content": "<p>Follow the giant, then you go furher.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1752413": "The Happywhale Competition this year had images from 28 different species, with over 15K individuals. Some photos are distant images of the backs of humpback whales and some are closeups of spinner dolphin dorsal fins. \n\nFor a baseline model, I just threw _all_ the training data into the pot, trained it overnight, and tried that. Well, I wasn’t surprised that I had terrible results: generating the transform that accurately locates photos in one of 15K compartments is slow work! \n\nI was confused that, even in the early days of the contest, the leaderboard was filling up with people getting 60-70% accuracy. My first baseline had about 12% accuracy! \n\n## Simple! Let’s try complexity!\n\nI quickly (maybe hastily) concluded that the way forward was _many models_. Instead of throwing all the data into one model and trying to generate 15K tight clusters, I’d use multiple layers: \n\n*  I trained a model to find, in the photo, a Region Of Interest aka tried to crop the image tightly around the whale\n* I trained a model to classify the camera viewpoint (I had 6 categories: anterior, posterior, port full, starboard full, port dorsal fin, starboard dorsal fin)\n* I trained a model to identify the species or species-type in the photo (by species-type I mean I had categories such as “blackfish” that included several species of small toothed whales)  \n\nHow complex was my multimodel approach? \n\n<img src=\"https://www.knowing.net/images/happywhale_multimodels.png\" alt=\"multimodel architecture\" width=\"800\"/>\n\n\n*sigh*\n\nRealize that this architecture requires `number of viewpoints * number of species` separate embedding models! \n\nWhich might have been tractable had I developed the ability to pretty-rapidly generate good embeddings for a particular viewpoint and species! It would have been too much for me to do locally, but I was planning on building an Azure ML Pipeline, spending some money to spin up a half dozen training machines, and basking in glory. All I had to do was develop some kick-ass embedding code.\n\nAnd that. Just. Didn’t. Work. \n\n## Am I getting close? Who knows?\n\nThe big problem with embeddings and the idea that “every individual has an area in high-dimensional space” is that when you generate a point in that high-dimensional space, it’s hard to say if it’s near or far from where you ultimately want it. \n\nFor one thing, “distance” can mean two common things when working with embeddings: \n\n* Euclidean distance, which is the distance in space between the two points\n* Cosine distance, which is the difference in the angles to the points, as measured from the origin\n\nSince this is a story of failure, I’m not going to try to justify why I was biased towards Cosine distance. But in the end, I tried almost everything with _both_. (And got crappy results.)\n\nNo matter which distance measure you choose, at the moment you generate an embedding for your target photo, what do you compare it to? _Eventually_ you want it to be “near other photos of the same individual” but in the _moment_, what do you know? \n\n## Overwhelmed by triplets\n\nWhat you can do is: instead of just the target photo, you have a _triplet_ of photos: \n\n* the target photo, \n* another photo of the same individual (a positive match), and \n* a photo you know is _not_ the same individual (a negative match)\n\nGenerate embeddings for each in the triplet. Now, early in the training process these will presumably spread out all over your n-dimensional space. *But* you know that you’d _rather_ have the positive match close and the negative match far. So you can change the weights to further that goal. \n\nAnd that, dear friends, is called _triplet loss_ and it’s been my world for the past month. \n\nAgain, since this is a story of failure, I shouldn’t lecture on “tips and tricks,” but I’ll just point out one obvious challenge with triplet loss: _most_ negative matches are going to be pretty obvious. Similarly, since cetacean photography often has images taken 1/10th of a second after the previous and some individuals have obvious features (mutilated fin edges, often), _some_ positive matches are also obvious. What you _want_ to train on are “the positive match that looks the least like this photo” and “the negative match that looks the most like this photo.” Trying to find _those_ photos is called _triplet mining_ and, yeesh, lemme’ tell ya’. Or, rather, what can I tell you? It didn’t work for me. \n\n \ntl;dr: I spent a month trying to apply an embeddings-based approach to cetacean reidentification and never had good results. I tried many standard things -- Siamese networks, Euclidean distance, Cosine distance, alternatives to ArcFace loss, soft and hard triplet mining, etc. In all cases, I would get good results with simplified datasets (MNIST-Fashion, a hand-picked set of individual cetaceans) but when I applied it to the Happywhale training data, I had very poor results.",
    "1752506": "Sounds like you tried a lot of things...and learned may'be a lot in the proces?\n\nSo if that is the case .. than did you really do that terrible??",
    "1752523": "I hope that once the leaders do their postmortems I'll be able to revisit things and get a handle on what I did wrong. I certainly feel like I tackled the problem from a lot of different angles, so it will be interesting to see if I was just missing some crucial ingredient or if I was too impatient and should have just trained slower and longer. \n\nOne thing I'd be very interested in would be not even the final solutions but the manner in which people were able to get 60-70% baselines so early in the process.",
    "1752534": "Meh, fit-predict is much more easier :)\n\nHonestly, you did a lot of work and got the most important thing - experience. But if you have bad results for a lot of methods - may be problem at step before - data? Here we have kernels with quite good score on ArcFace, I'd recommend repeat it and go back step-by-step to your solution.\n\nSimple is better than complex (C) The Zen of Python. Start to do complex stuff after you have good simple.",
    "1752660": "Follow the giant, then you go furher."
  },
  "source": "meta"
}