{
  "id": 486805,
  "title": "Knowledge Distillation",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/486805",
  "author_name": "",
  "post_date": "2024-03-26T11:51:08.019998800Z",
  "votes": 8,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello Everyone,</p>\n<p>Knowledge distillation seems to be giving good improvments, (i.e single model from 0.37 to 0.34 LB) <br>\nI have shared <a href=\"https://www.kaggle.com/datasets/nartaa/knowledge-distillation/data\" target=\"_blank\">this distilled knowledge dataset</a> that is based on 0.3 LB ensemble predictions.</p>\n<p>To find out how to use it, you can check out the code in the <a href=\"https://www.kaggle.com/code/nartaa/features-head-starter\" target=\"_blank\">Features+Head Starter</a></p>",
  "messages": [
    {
      "id": "2717129",
      "postDate": "03/26/2024 11:51:08",
      "content": "<p>Hello Everyone,</p>\n<p>Knowledge distillation seems to be giving good improvments, (i.e single model from 0.37 to 0.34 LB) <br>\nI have shared <a href=\"https://www.kaggle.com/datasets/nartaa/knowledge-distillation/data\" target=\"_blank\">this distilled knowledge dataset</a> that is based on 0.3 LB ensemble predictions.</p>\n<p>To find out how to use it, you can check out the code in the <a href=\"https://www.kaggle.com/code/nartaa/features-head-starter\" target=\"_blank\">Features+Head Starter</a></p>",
      "rawMarkdown": "Hello Everyone,\n\nKnowledge distillation seems to be giving good improvments, (i.e single model from 0.37 to 0.34 LB) \nI have shared [this distilled knowledge dataset](https://www.kaggle.com/datasets/nartaa/knowledge-distillation/data) that is based on 0.3 LB ensemble predictions.\n\nTo find out how to use it, you can check out the code in the [Features+Head Starter] (https://www.kaggle.com/code/nartaa/features-head-starter)",
      "votes": null
    },
    {
      "id": "2717377",
      "postDate": "03/26/2024 14:40:16",
      "content": "<p>Thanks for providing the dataset. Are these OOF predictions from the ensembled models or just train in sample? </p>",
      "rawMarkdown": "Thanks for providing the dataset. Are these OOF predictions from the ensembled models or just train in sample?",
      "votes": null
    },
    {
      "id": "2717402",
      "postDate": "03/26/2024 14:50:18",
      "content": "<p>You are welcome.<br>\nNot OOF, every fold predicts on all training samples.</p>",
      "rawMarkdown": "You are welcome.\nNot OOF, every fold predicts on all training samples.",
      "votes": null
    },
    {
      "id": "2717527",
      "postDate": "03/26/2024 16:10:14",
      "content": "<p>Does it have any effect on your integration?</p>",
      "rawMarkdown": "Does it have any effect on your integration?",
      "votes": null
    },
    {
      "id": "2717560",
      "postDate": "03/26/2024 16:37:45",
      "content": "<p>I ran out of GPU time this week, so I couldn't say, but it should</p>",
      "rawMarkdown": "I ran out of GPU time this week, so I couldn't say, but it should",
      "votes": null
    },
    {
      "id": "2718176",
      "postDate": "03/27/2024 02:45:35",
      "content": "<p>What do you mean by KD here in this context? Are you feeding a model the ensemble's predictions on train.csv as a ground truth? In essence, the smaller/simplier model will just learn to emulate the ensemble - don't really see the benefit of it when the simpler model is then placed as part of that ensemble</p>",
      "rawMarkdown": "What do you mean by KD here in this context? Are you feeding a model the ensemble's predictions on train.csv as a ground truth? In essence, the smaller/simplier model will just learn to emulate the ensemble - don't really see the benefit of it when the simpler model is then placed as part of that ensemble",
      "votes": null
    },
    {
      "id": "2718237",
      "postDate": "03/27/2024 03:22:18",
      "content": "<p>Thanks for sharing!<br>\nDoes the cv also improve?</p>",
      "rawMarkdown": "Thanks for sharing!\nDoes the cv also improve?",
      "votes": null
    },
    {
      "id": "2718694",
      "postDate": "03/27/2024 07:49:40",
      "content": "<p>You are welcome!<br>\nYes, CV and LB both improve nicely.</p>",
      "rawMarkdown": "You are welcome!\nYes, CV and LB both improve nicely.",
      "votes": null
    },
    {
      "id": "2718725",
      "postDate": "03/27/2024 08:01:11",
      "content": "<blockquote>\n  <p>What do you mean by KD here in this context? </p>\n</blockquote>\n<p>Probably in this context it would mean training on pseudo labeling based on best model's predictions, the term knowledge distillation includes true labels in the loss function, but both methods distill knowledge from pseudo labels.</p>\n<blockquote>\n  <p>Are you feeding a model the ensemble's predictions on train.csv as a ground truth?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>don't really see the benefit of it when the simpler model is then placed as part of that ensemble</p>\n</blockquote>\n<p>I have similliar intuition, but knowledge distillation is a phenomena that works against intuition, just like ensemble. I've looked at previous competitions, they utilize this method. I will conclude next week if it would actually give better LB on ensemble.</p>",
      "rawMarkdown": "> What do you mean by KD here in this context? \n\nProbably in this context it would mean training on pseudo labeling based on best model's predictions, the term knowledge distillation includes true labels in the loss function, but both methods distill knowledge from pseudo labels.\n\n> Are you feeding a model the ensemble's predictions on train.csv as a ground truth?\n\nYes.\n\n>  don't really see the benefit of it when the simpler model is then placed as part of that ensemble\n\nI have similliar intuition, but knowledge distillation is a phenomena that works against intuition, just like ensemble. I've looked at previous competitions, they utilize this method. I will conclude next week if it would actually give better LB on ensemble.",
      "votes": null
    },
    {
      "id": "2719734",
      "postDate": "03/27/2024 22:16:07",
      "content": "<p>Thank you for the reply, eager to see what happens (also out of GPU allocation this week) - If it improves I'd be keen to understand why!</p>",
      "rawMarkdown": "Thank you for the reply, eager to see what happens (also out of GPU allocation this week) - If it improves I'd be keen to understand why!",
      "votes": null
    },
    {
      "id": "2722397",
      "postDate": "03/29/2024 15:13:48",
      "content": "<p>I retrained the KER with KD and got a single model 0.33 LB<br>\n<a href=\"https://www.kaggle.com/code/jimmyisme1/features-head-ker-only-lb-0-33\" target=\"_blank\">https://www.kaggle.com/code/jimmyisme1/features-head-ker-only-lb-0-33</a></p>",
      "rawMarkdown": "I retrained the KER with KD and got a single model 0.33 LB\nhttps://www.kaggle.com/code/jimmyisme1/features-head-ker-only-lb-0-33",
      "votes": null
    },
    {
      "id": "2722691",
      "postDate": "03/29/2024 18:06:33",
      "content": "<p>Nicely done!</p>",
      "rawMarkdown": "Nicely done!",
      "votes": null
    },
    {
      "id": "2723259",
      "postDate": "03/30/2024 04:55:11",
      "content": "<p>Would it be better if I trained more models with KD and then fused them together？</p>",
      "rawMarkdown": "Would it be better if I trained more models with KD and then fused them together？",
      "votes": null
    },
    {
      "id": "2723390",
      "postDate": "03/30/2024 07:02:45",
      "content": "<p>I think so, that would be my next move.</p>",
      "rawMarkdown": "I think so, that would be my next move.",
      "votes": null
    },
    {
      "id": "2725208",
      "postDate": "03/31/2024 12:13:38",
      "content": "<p>how about your cv?</p>",
      "rawMarkdown": "how about your cv?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2717377,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "03/26/2024 14:40:16",
      "content": "<p>Thanks for providing the dataset. Are these OOF predictions from the ensembled models or just train in sample? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2717402,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "03/26/2024 14:50:18",
          "content": "<p>You are welcome.<br>\nNot OOF, every fold predicts on all training samples.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2717527,
      "author_name": "gentlezdh",
      "author_url": "",
      "post_date": "03/26/2024 16:10:14",
      "content": "<p>Does it have any effect on your integration?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2717560,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "03/26/2024 16:37:45",
          "content": "<p>I ran out of GPU time this week, so I couldn't say, but it should</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2718176,
      "author_name": "exjustice",
      "author_url": "",
      "post_date": "03/27/2024 02:45:35",
      "content": "<p>What do you mean by KD here in this context? Are you feeding a model the ensemble's predictions on train.csv as a ground truth? In essence, the smaller/simplier model will just learn to emulate the ensemble - don't really see the benefit of it when the simpler model is then placed as part of that ensemble</p>",
      "votes": null,
      "replies": [
        {
          "id": 2718725,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "03/27/2024 08:01:11",
          "content": "<blockquote>\n  <p>What do you mean by KD here in this context? </p>\n</blockquote>\n<p>Probably in this context it would mean training on pseudo labeling based on best model's predictions, the term knowledge distillation includes true labels in the loss function, but both methods distill knowledge from pseudo labels.</p>\n<blockquote>\n  <p>Are you feeding a model the ensemble's predictions on train.csv as a ground truth?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>don't really see the benefit of it when the simpler model is then placed as part of that ensemble</p>\n</blockquote>\n<p>I have similliar intuition, but knowledge distillation is a phenomena that works against intuition, just like ensemble. I've looked at previous competitions, they utilize this method. I will conclude next week if it would actually give better LB on ensemble.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2719734,
              "author_name": "exjustice",
              "author_url": "",
              "post_date": "03/27/2024 22:16:07",
              "content": "<p>Thank you for the reply, eager to see what happens (also out of GPU allocation this week) - If it improves I'd be keen to understand why!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2718237,
      "author_name": "johnlemon3",
      "author_url": "",
      "post_date": "03/27/2024 03:22:18",
      "content": "<p>Thanks for sharing!<br>\nDoes the cv also improve?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2718694,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "03/27/2024 07:49:40",
          "content": "<p>You are welcome!<br>\nYes, CV and LB both improve nicely.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2722397,
      "author_name": "jimmyisme1",
      "author_url": "",
      "post_date": "03/29/2024 15:13:48",
      "content": "<p>I retrained the KER with KD and got a single model 0.33 LB<br>\n<a href=\"https://www.kaggle.com/code/jimmyisme1/features-head-ker-only-lb-0-33\" target=\"_blank\">https://www.kaggle.com/code/jimmyisme1/features-head-ker-only-lb-0-33</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2722691,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "03/29/2024 18:06:33",
          "content": "<p>Nicely done!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2723259,
          "author_name": "majiaqi111",
          "author_url": "",
          "post_date": "03/30/2024 04:55:11",
          "content": "<p>Would it be better if I trained more models with KD and then fused them together？</p>",
          "votes": null,
          "replies": [
            {
              "id": 2723390,
              "author_name": "nartaa",
              "author_url": "",
              "post_date": "03/30/2024 07:02:45",
              "content": "<p>I think so, that would be my next move.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2725208,
          "author_name": "gentlezdh",
          "author_url": "",
          "post_date": "03/31/2024 12:13:38",
          "content": "<p>how about your cv?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2717129": "Hello Everyone,\n\nKnowledge distillation seems to be giving good improvments, (i.e single model from 0.37 to 0.34 LB) \nI have shared [this distilled knowledge dataset](https://www.kaggle.com/datasets/nartaa/knowledge-distillation/data) that is based on 0.3 LB ensemble predictions.\n\nTo find out how to use it, you can check out the code in the [Features+Head Starter] (https://www.kaggle.com/code/nartaa/features-head-starter)",
    "2717377": "Thanks for providing the dataset. Are these OOF predictions from the ensembled models or just train in sample?",
    "2717402": "You are welcome.\nNot OOF, every fold predicts on all training samples.",
    "2717527": "Does it have any effect on your integration?",
    "2717560": "I ran out of GPU time this week, so I couldn't say, but it should",
    "2718176": "What do you mean by KD here in this context? Are you feeding a model the ensemble's predictions on train.csv as a ground truth? In essence, the smaller/simplier model will just learn to emulate the ensemble - don't really see the benefit of it when the simpler model is then placed as part of that ensemble",
    "2718237": "Thanks for sharing!\nDoes the cv also improve?",
    "2718694": "You are welcome!\nYes, CV and LB both improve nicely.",
    "2718725": "> What do you mean by KD here in this context? \n\nProbably in this context it would mean training on pseudo labeling based on best model's predictions, the term knowledge distillation includes true labels in the loss function, but both methods distill knowledge from pseudo labels.\n\n> Are you feeding a model the ensemble's predictions on train.csv as a ground truth?\n\nYes.\n\n>  don't really see the benefit of it when the simpler model is then placed as part of that ensemble\n\nI have similliar intuition, but knowledge distillation is a phenomena that works against intuition, just like ensemble. I've looked at previous competitions, they utilize this method. I will conclude next week if it would actually give better LB on ensemble.",
    "2719734": "Thank you for the reply, eager to see what happens (also out of GPU allocation this week) - If it improves I'd be keen to understand why!",
    "2722397": "I retrained the KER with KD and got a single model 0.33 LB\nhttps://www.kaggle.com/code/jimmyisme1/features-head-ker-only-lb-0-33",
    "2722691": "Nicely done!",
    "2723259": "Would it be better if I trained more models with KD and then fused them together？",
    "2723390": "I think so, that would be my next move.",
    "2725208": "how about your cv?"
  },
  "source": "meta"
}