{
  "id": 396103,
  "title": "Finetuning Google bird-vocalization-classifier model",
  "url": "/competitions/birdclef-2023/discussion/396103",
  "author_name": "",
  "post_date": "2023-03-20T10:44:09.954654800Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have a question about how to finetune Google bird-vocalization-classifier model. As mentioned <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1/discussion/393788#2176598\" target=\"_blank\">here</a>, the model itself can't be tuned due to being ported from JAX. Therefore, it's suggested to take output embeddings from the model, and train another (smaller) model on it. Output embeddings have shape of <code>(1, 1280)</code>. Am I correct, that the another (smaller) model should be 1-dimensional, for example, composed from Conv1d layers?</p>",
  "messages": [
    {
      "id": "2189236",
      "postDate": "03/20/2023 10:44:09",
      "content": "<p>I have a question about how to finetune Google bird-vocalization-classifier model. As mentioned <a href=\"https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1/discussion/393788#2176598\" target=\"_blank\">here</a>, the model itself can't be tuned due to being ported from JAX. Therefore, it's suggested to take output embeddings from the model, and train another (smaller) model on it. Output embeddings have shape of <code>(1, 1280)</code>. Am I correct, that the another (smaller) model should be 1-dimensional, for example, composed from Conv1d layers?</p>",
      "rawMarkdown": "I have a question about how to finetune Google bird-vocalization-classifier model. As mentioned [here](https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1/discussion/393788#2176598), the model itself can't be tuned due to being ported from JAX. Therefore, it's suggested to take output embeddings from the model, and train another (smaller) model on it. Output embeddings have shape of `(1, 1280)`. Am I correct, that the another (smaller) model should be 1-dimensional, for example, composed from Conv1d layers?",
      "votes": null
    },
    {
      "id": "2189545",
      "postDate": "03/20/2023 15:38:12",
      "content": "<p>Hi, Araik;<br>\nThe signature takes a 5s audio segment and produces one flat 1280-dimensional embedding (without a time dimension). You can concatenate multiple embeddings along the batch dimension if you'd like to use a convolutional model.</p>",
      "rawMarkdown": "Hi, Araik;\nThe signature takes a 5s audio segment and produces one flat 1280-dimensional embedding (without a time dimension). You can concatenate multiple embeddings along the batch dimension if you'd like to use a convolutional model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2189545,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "03/20/2023 15:38:12",
      "content": "<p>Hi, Araik;<br>\nThe signature takes a 5s audio segment and produces one flat 1280-dimensional embedding (without a time dimension). You can concatenate multiple embeddings along the batch dimension if you'd like to use a convolutional model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2189236": "I have a question about how to finetune Google bird-vocalization-classifier model. As mentioned [here](https://www.kaggle.com/models/google/bird-vocalization-classifier/frameworks/TensorFlow2/variations/bird-vocalization-classifier/versions/1/discussion/393788#2176598), the model itself can't be tuned due to being ported from JAX. Therefore, it's suggested to take output embeddings from the model, and train another (smaller) model on it. Output embeddings have shape of `(1, 1280)`. Am I correct, that the another (smaller) model should be 1-dimensional, for example, composed from Conv1d layers?",
    "2189545": "Hi, Araik;\nThe signature takes a 5s audio segment and produces one flat 1280-dimensional embedding (without a time dimension). You can concatenate multiple embeddings along the batch dimension if you'd like to use a convolutional model."
  },
  "source": "meta"
}