{
  "id": 502472,
  "title": "Tips For Multilabel Knowledge Distillation",
  "url": "/competitions/birdclef-2024/discussion/502472",
  "author_name": "",
  "post_date": "2024-05-13T15:51:18.603137Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Knowledge distillation seems like a promising approach for training a large model and distilling it to a model that's small enough to do inference in the allotted time. I found this <a href=\"https://arxiv.org/pdf/2308.06453\" target=\"_blank\">paper</a> which has a method for the multilabel case. However, I'm having difficulty actually producing a quality student model. </p>\n<p>Has anyone found anything that works well for multilabel knowledge distillation?</p>",
  "messages": [
    {
      "id": "2811090",
      "postDate": "05/13/2024 15:51:18",
      "content": "<p>Knowledge distillation seems like a promising approach for training a large model and distilling it to a model that's small enough to do inference in the allotted time. I found this <a href=\"https://arxiv.org/pdf/2308.06453\" target=\"_blank\">paper</a> which has a method for the multilabel case. However, I'm having difficulty actually producing a quality student model. </p>\n<p>Has anyone found anything that works well for multilabel knowledge distillation?</p>",
      "rawMarkdown": "Knowledge distillation seems like a promising approach for training a large model and distilling it to a model that's small enough to do inference in the allotted time. I found this [paper](https://arxiv.org/pdf/2308.06453) which has a method for the multilabel case. However, I'm having difficulty actually producing a quality student model. \n\nHas anyone found anything that works well for multilabel knowledge distillation?",
      "votes": null
    },
    {
      "id": "2811169",
      "postDate": "05/13/2024 16:09:55",
      "content": "<p>The 4th place approach from last year uses softmax distillation and only the primary labels so I guess I could try that.<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412753</a></p>",
      "rawMarkdown": "The 4th place approach from last year uses softmax distillation and only the primary labels so I guess I could try that.\nhttps://www.kaggle.com/competitions/birdclef-2023/discussion/412753",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2811169,
      "author_name": "willrice",
      "author_url": "",
      "post_date": "05/13/2024 16:09:55",
      "content": "<p>The 4th place approach from last year uses softmax distillation and only the primary labels so I guess I could try that.<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412753\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412753</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2811090": "Knowledge distillation seems like a promising approach for training a large model and distilling it to a model that's small enough to do inference in the allotted time. I found this [paper](https://arxiv.org/pdf/2308.06453) which has a method for the multilabel case. However, I'm having difficulty actually producing a quality student model. \n\nHas anyone found anything that works well for multilabel knowledge distillation?",
    "2811169": "The 4th place approach from last year uses softmax distillation and only the primary labels so I guess I could try that.\nhttps://www.kaggle.com/competitions/birdclef-2023/discussion/412753"
  },
  "source": "meta"
}