{
  "id": 320477,
  "title": "High scorer willing to share model `summary`?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/320477",
  "author_name": "",
  "post_date": "2022-04-21T20:00:25.267365300Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Would anyone in the medal tier be willing to share a <code>torchinfo</code> / TF <code>summary</code> of their final model? A few weeks ago I admitted to doing terribly in this competition and after reading the high-scorer post-mortems I tried to switch to a higher-resolution, more capable backbone (<code>efficientnet-b5</code>), which is the main fundamental thing that jumped out. Initial 30-epoch training with that doesn’t seem to do any better than what I’d seen and I’m wondering if I’m just missing something fundamental about the head or maybe the inference logic. (By the end, I was generating embeddings for every image in the training set, putting individual embeddings into one FAISS index (nearest photo) and centroids into another (nearest centerpoint). For me, neither does well.)</p>\n<p>I’m not looking for near-final LB/CV score, I’ll be very happy if I can just get &gt;.6 Happywhale score. At least then I’ll have a sense of what the heck I was doing wrong.</p>",
  "messages": [
    {
      "id": "1763780",
      "postDate": "04/21/2022 20:00:25",
      "content": "<p>Would anyone in the medal tier be willing to share a <code>torchinfo</code> / TF <code>summary</code> of their final model? A few weeks ago I admitted to doing terribly in this competition and after reading the high-scorer post-mortems I tried to switch to a higher-resolution, more capable backbone (<code>efficientnet-b5</code>), which is the main fundamental thing that jumped out. Initial 30-epoch training with that doesn’t seem to do any better than what I’d seen and I’m wondering if I’m just missing something fundamental about the head or maybe the inference logic. (By the end, I was generating embeddings for every image in the training set, putting individual embeddings into one FAISS index (nearest photo) and centroids into another (nearest centerpoint). For me, neither does well.)</p>\n<p>I’m not looking for near-final LB/CV score, I’ll be very happy if I can just get &gt;.6 Happywhale score. At least then I’ll have a sense of what the heck I was doing wrong.</p>",
      "rawMarkdown": "Would anyone in the medal tier be willing to share a `torchinfo` / TF `summary` of their final model? A few weeks ago I admitted to doing terribly in this competition and after reading the high-scorer post-mortems I tried to switch to a higher-resolution, more capable backbone (`efficientnet-b5`), which is the main fundamental thing that jumped out. Initial 30-epoch training with that doesn’t seem to do any better than what I’d seen and I’m wondering if I’m just missing something fundamental about the head or maybe the inference logic. (By the end, I was generating embeddings for every image in the training set, putting individual embeddings into one FAISS index (nearest photo) and centroids into another (nearest centerpoint). For me, neither does well.)\n\nI’m not looking for near-final LB/CV score, I’ll be very happy if I can just get >.6 Happywhale score. At least then I’ll have a sense of what the heck I was doing wrong.",
      "votes": null
    },
    {
      "id": "1763791",
      "postDate": "04/21/2022 20:17:43",
      "content": "<p>Hi. An easy way to boost your CV LB with your existing pipeline is to just train with all the 4 public datasets concatenated. For example </p>\n<pre><code>train_dataset = get_training_dataset(train_files1 + train_files2 \n    + train_files3 + train_files4)\nmodel.fit(train_dataset, epochs=20)\n</code></pre>\n<p>Then afterward, infer each dataset separately and average the embeddings</p>\n<pre><code>test_files1 = tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-test*.tfrec')\ntest_dataset1 = get_test_dataset(test_files1)\ntest_embeddings1 = embed_model.predict(test_dataset1)\n\ntest_files2 = tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-test*.tfrec')\ntest_dataset2 = get_test_dataset(test_files2)\ntest_embeddings2 = embed_model.predict(test_dataset2)\n\ntest_files3 = tf.io.gfile.glob(GCS_PATH3 + '/happywhale-2022-test*.tfrec')\ntest_dataset3 = get_test_dataset(test_files3)\ntest_embeddings3 = embed_model.predict(test_dataset3)\n\ntest_files4 = tf.io.gfile.glob(GCS_PATH4 + '/happywhale-2022-test*.tfrec')\ntest_dataset4 = get_test_dataset(test_files4)\ntest_embeddings4 = embed_model.predict(test_dataset4)\n</code></pre>\n<p>Then</p>\n<pre><code>test_embeddings = (test_embeddings1 + test_embeddings2 \n        + test_embeddings3 + test_embeddings4)/4\n</code></pre>\n<p>And create the train embeddings the same way. More info in my team's solution write up <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Hi. An easy way to boost your CV LB with your existing pipeline is to just train with all the 4 public datasets concatenated. For example \n\n    train_dataset = get_training_dataset(train_files1 + train_files2 \n        + train_files3 + train_files4)\n    model.fit(train_dataset, epochs=20)\n\nThen afterward, infer each dataset separately and average the embeddings\n\n    test_files1 = tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-test*.tfrec')\n    test_dataset1 = get_test_dataset(test_files1)\n    test_embeddings1 = embed_model.predict(test_dataset1)\n\n    test_files2 = tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-test*.tfrec')\n    test_dataset2 = get_test_dataset(test_files2)\n    test_embeddings2 = embed_model.predict(test_dataset2)\n\n    test_files3 = tf.io.gfile.glob(GCS_PATH3 + '/happywhale-2022-test*.tfrec')\n    test_dataset3 = get_test_dataset(test_files3)\n    test_embeddings3 = embed_model.predict(test_dataset3)\n\n    test_files4 = tf.io.gfile.glob(GCS_PATH4 + '/happywhale-2022-test*.tfrec')\n    test_dataset4 = get_test_dataset(test_files4)\n    test_embeddings4 = embed_model.predict(test_dataset4)\n\nThen\n\n    test_embeddings = (test_embeddings1 + test_embeddings2 \n            + test_embeddings3 + test_embeddings4)/4\n\nAnd create the train embeddings the same way. More info in my team's solution write up [here][1]\n\n[1]: https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298",
      "votes": null
    },
    {
      "id": "1763964",
      "postDate": "04/22/2022 02:35:29",
      "content": "<p>It sounds like the fundamental thing you missed is using a good cropping algorithm. Using the <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184\" target=\"_blank\">\"fullbody\" annotations by Jan Bre</a> significantly improved the performance of our models. That change alone takes the public solution from ~0.55 to 0.7.</p>",
      "rawMarkdown": "It sounds like the fundamental thing you missed is using a good cropping algorithm. Using the [\"fullbody\" annotations by Jan Bre](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184) significantly improved the performance of our models. That change alone takes the public solution from ~0.55 to 0.7.",
      "votes": null
    },
    {
      "id": "1763966",
      "postDate": "04/22/2022 02:37:30",
      "content": "<p>Exactly. Datasets made from cropping was powerful. And if you used multiple datasets made from different types of crops it boosted even more.</p>",
      "rawMarkdown": "Exactly. Datasets made from cropping was powerful. And if you used multiple datasets made from different types of crops it boosted even more.",
      "votes": null
    },
    {
      "id": "1764346",
      "postDate": "04/22/2022 12:12:06",
      "content": "<p>I am thinking about this line <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>test_embeddings = (test_embeddings1 + test_embeddings2 <br>\n          + test_embeddings3 + test_embeddings4)/4</p>\n</blockquote>\n<p>How to understand this? Each embedding (vector in a multidimentional space) represents different world (from different dataset) and then we averaging them … I am trying to understand logic behind this operation. 😄😄 As I understand we just trying to find … vector position which is an average of all feature representations … so something between … them …</p>",
      "rawMarkdown": "I am thinking about this line @cdeotte \n\n> test_embeddings = (test_embeddings1 + test_embeddings2 \n        + test_embeddings3 + test_embeddings4)/4\n\nHow to understand this? Each embedding (vector in a multidimentional space) represents different world (from different dataset) and then we averaging them ... I am trying to understand logic behind this operation. 😄😄 As I understand we just trying to find ... vector position which is an average of all feature representations ... so something between ... them ...",
      "votes": null
    },
    {
      "id": "1764435",
      "postDate": "04/22/2022 13:24:40",
      "content": "<p>Yes. The first thing to consider is that each embedding is created from the same model. (The difference is just inferring different dataset, i.e. a different view of the fish). For EffNetB7, each embeddings has 2560 columns where each column indicates the presence of a certain feature in the image. </p>\n<p>So for example, column 1 (of each embedding) may indicate the present of a triangle shaped scar. When we average the 4 embeddings, we ask \"do any of the different crops (i.e. views of the fish) have a triangle shaped scar?\" </p>\n<p>This is very similar to global mean pooling 2D which searches an image for a certain feature.</p>",
      "rawMarkdown": "Yes. The first thing to consider is that each embedding is created from the same model. (The difference is just inferring different dataset, i.e. a different view of the fish). For EffNetB7, each embeddings has 2560 columns where each column indicates the presence of a certain feature in the image. \n\nSo for example, column 1 (of each embedding) may indicate the present of a triangle shaped scar. When we average the 4 embeddings, we ask \"do any of the different crops (i.e. views of the fish) have a triangle shaped scar?\" \n\nThis is very similar to global mean pooling 2D which searches an image for a certain feature.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1763791,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/21/2022 20:17:43",
      "content": "<p>Hi. An easy way to boost your CV LB with your existing pipeline is to just train with all the 4 public datasets concatenated. For example </p>\n<pre><code>train_dataset = get_training_dataset(train_files1 + train_files2 \n    + train_files3 + train_files4)\nmodel.fit(train_dataset, epochs=20)\n</code></pre>\n<p>Then afterward, infer each dataset separately and average the embeddings</p>\n<pre><code>test_files1 = tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-test*.tfrec')\ntest_dataset1 = get_test_dataset(test_files1)\ntest_embeddings1 = embed_model.predict(test_dataset1)\n\ntest_files2 = tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-test*.tfrec')\ntest_dataset2 = get_test_dataset(test_files2)\ntest_embeddings2 = embed_model.predict(test_dataset2)\n\ntest_files3 = tf.io.gfile.glob(GCS_PATH3 + '/happywhale-2022-test*.tfrec')\ntest_dataset3 = get_test_dataset(test_files3)\ntest_embeddings3 = embed_model.predict(test_dataset3)\n\ntest_files4 = tf.io.gfile.glob(GCS_PATH4 + '/happywhale-2022-test*.tfrec')\ntest_dataset4 = get_test_dataset(test_files4)\ntest_embeddings4 = embed_model.predict(test_dataset4)\n</code></pre>\n<p>Then</p>\n<pre><code>test_embeddings = (test_embeddings1 + test_embeddings2 \n        + test_embeddings3 + test_embeddings4)/4\n</code></pre>\n<p>And create the train embeddings the same way. More info in my team's solution write up <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1764346,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "04/22/2022 12:12:06",
          "content": "<p>I am thinking about this line <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>test_embeddings = (test_embeddings1 + test_embeddings2 <br>\n          + test_embeddings3 + test_embeddings4)/4</p>\n</blockquote>\n<p>How to understand this? Each embedding (vector in a multidimentional space) represents different world (from different dataset) and then we averaging them … I am trying to understand logic behind this operation. 😄😄 As I understand we just trying to find … vector position which is an average of all feature representations … so something between … them …</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764435,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/22/2022 13:24:40",
          "content": "<p>Yes. The first thing to consider is that each embedding is created from the same model. (The difference is just inferring different dataset, i.e. a different view of the fish). For EffNetB7, each embeddings has 2560 columns where each column indicates the presence of a certain feature in the image. </p>\n<p>So for example, column 1 (of each embedding) may indicate the present of a triangle shaped scar. When we average the 4 embeddings, we ask \"do any of the different crops (i.e. views of the fish) have a triangle shaped scar?\" </p>\n<p>This is very similar to global mean pooling 2D which searches an image for a certain feature.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1763964,
      "author_name": "lextoumbourou",
      "author_url": "",
      "post_date": "04/22/2022 02:35:29",
      "content": "<p>It sounds like the fundamental thing you missed is using a good cropping algorithm. Using the <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184\" target=\"_blank\">\"fullbody\" annotations by Jan Bre</a> significantly improved the performance of our models. That change alone takes the public solution from ~0.55 to 0.7.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1763966,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/22/2022 02:37:30",
          "content": "<p>Exactly. Datasets made from cropping was powerful. And if you used multiple datasets made from different types of crops it boosted even more.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1763780": "Would anyone in the medal tier be willing to share a `torchinfo` / TF `summary` of their final model? A few weeks ago I admitted to doing terribly in this competition and after reading the high-scorer post-mortems I tried to switch to a higher-resolution, more capable backbone (`efficientnet-b5`), which is the main fundamental thing that jumped out. Initial 30-epoch training with that doesn’t seem to do any better than what I’d seen and I’m wondering if I’m just missing something fundamental about the head or maybe the inference logic. (By the end, I was generating embeddings for every image in the training set, putting individual embeddings into one FAISS index (nearest photo) and centroids into another (nearest centerpoint). For me, neither does well.)\n\nI’m not looking for near-final LB/CV score, I’ll be very happy if I can just get >.6 Happywhale score. At least then I’ll have a sense of what the heck I was doing wrong.",
    "1763791": "Hi. An easy way to boost your CV LB with your existing pipeline is to just train with all the 4 public datasets concatenated. For example \n\n    train_dataset = get_training_dataset(train_files1 + train_files2 \n        + train_files3 + train_files4)\n    model.fit(train_dataset, epochs=20)\n\nThen afterward, infer each dataset separately and average the embeddings\n\n    test_files1 = tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-test*.tfrec')\n    test_dataset1 = get_test_dataset(test_files1)\n    test_embeddings1 = embed_model.predict(test_dataset1)\n\n    test_files2 = tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-test*.tfrec')\n    test_dataset2 = get_test_dataset(test_files2)\n    test_embeddings2 = embed_model.predict(test_dataset2)\n\n    test_files3 = tf.io.gfile.glob(GCS_PATH3 + '/happywhale-2022-test*.tfrec')\n    test_dataset3 = get_test_dataset(test_files3)\n    test_embeddings3 = embed_model.predict(test_dataset3)\n\n    test_files4 = tf.io.gfile.glob(GCS_PATH4 + '/happywhale-2022-test*.tfrec')\n    test_dataset4 = get_test_dataset(test_files4)\n    test_embeddings4 = embed_model.predict(test_dataset4)\n\nThen\n\n    test_embeddings = (test_embeddings1 + test_embeddings2 \n            + test_embeddings3 + test_embeddings4)/4\n\nAnd create the train embeddings the same way. More info in my team's solution write up [here][1]\n\n[1]: https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298",
    "1763964": "It sounds like the fundamental thing you missed is using a good cropping algorithm. Using the [\"fullbody\" annotations by Jan Bre](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184) significantly improved the performance of our models. That change alone takes the public solution from ~0.55 to 0.7.",
    "1763966": "Exactly. Datasets made from cropping was powerful. And if you used multiple datasets made from different types of crops it boosted even more.",
    "1764346": "I am thinking about this line @cdeotte \n\n> test_embeddings = (test_embeddings1 + test_embeddings2 \n        + test_embeddings3 + test_embeddings4)/4\n\nHow to understand this? Each embedding (vector in a multidimentional space) represents different world (from different dataset) and then we averaging them ... I am trying to understand logic behind this operation. 😄😄 As I understand we just trying to find ... vector position which is an average of all feature representations ... so something between ... them ...",
    "1764435": "Yes. The first thing to consider is that each embedding is created from the same model. (The difference is just inferring different dataset, i.e. a different view of the fish). For EffNetB7, each embeddings has 2560 columns where each column indicates the presence of a certain feature in the image. \n\nSo for example, column 1 (of each embedding) may indicate the present of a triangle shaped scar. When we average the 4 embeddings, we ask \"do any of the different crops (i.e. views of the fish) have a triangle shaped scar?\" \n\nThis is very similar to global mean pooling 2D which searches an image for a certain feature."
  },
  "source": "meta"
}