{
  "id": 125046,
  "title": "Teacher-Student",
  "url": "/competitions/bengaliai-cv19/discussion/125046",
  "author_name": "Mighty Rains",
  "post_date": "2020-01-08T09:15:57.014000",
  "votes": 5,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Has anybody tried to train multiple ensembles of models into a single teacher and got a good performance?\nI mean not only in this competition, but also in others/outside kaggle.</p>",
  "messages": [
    {
      "id": 713416,
      "postDate": "2020-01-08T09:15:57.013Z",
      "content": "<p>Has anybody tried to train multiple ensembles of models into a single teacher and got a good performance?\nI mean not only in this competition, but also in others/outside kaggle.</p>",
      "rawMarkdown": "Has anybody tried to train multiple ensembles of models into a single teacher and got a good performance?\nI mean not only in this competition, but also in others/outside kaggle.",
      "votes": 5
    },
    {
      "id": 713808,
      "postDate": "2020-01-08T17:11:28.003Z",
      "content": "<p>Paper with current SOTA on imagenet (updated yesterday Jan 7th 2020) uses it <a href=\"https://arxiv.org/abs/1911.04252\">https://arxiv.org/abs/1911.04252</a></p>",
      "rawMarkdown": "Paper with current SOTA on imagenet (updated yesterday Jan 7th 2020) uses it https://arxiv.org/abs/1911.04252",
      "votes": 3
    },
    {
      "id": 716923,
      "postDate": "2020-01-12T13:15:31.737Z",
      "content": "<p>I tried that out, although on a very superficial level. I trained a resnet18 on an ensemble of 5fold resnet34 scoring 0.9656 in LB and the student got 0.9527 LB. A densenet121 is being tutored at the moment by the same ensemble of teachers. I'll update the post when it does.</p>",
      "rawMarkdown": "I tried that out, although on a very superficial level. I trained a resnet18 on an ensemble of 5fold resnet34 scoring 0.9656 in LB and the student got 0.9527 LB. A densenet121 is being tutored at the moment by the same ensemble of teachers. I'll update the post when it does.",
      "votes": 1
    },
    {
      "id": 716899,
      "postDate": "2020-01-12T12:37:54.667Z",
      "content": "<p>I'm trying to distil from B3 to B0 right now. If things work well, I might go further and run distillation from an ensemble of teachers ✊ </p>",
      "rawMarkdown": "I'm trying to distil from B3 to B0 right now. If things work well, I might go further and run distillation from an ensemble of teachers ✊ ",
      "votes": 1,
      "replies": [
        {
          "id": 716925,
          "postDate": "2020-01-12T13:18:17.597Z",
          "content": "<p>What loss are you using? I did a blend of KL divergence loss between teacher-student and a CCE between student-true labels.</p>",
          "rawMarkdown": "What loss are you using? I did a blend of KL divergence loss between teacher-student and a CCE between student-true labels."
        },
        {
          "id": 717059,
          "postDate": "2020-01-12T17:05:10.183Z",
          "content": "<p>Yes, same as you, adding a kl loss between teacher and student outputs for every mini-batch iteration along with the usual CE. \nMy 1st experiment show a slight bump (0.03) on the student model on local val.</p>",
          "rawMarkdown": "Yes, same as you, adding a kl loss between teacher and student outputs for every mini-batch iteration along with the usual CE. \nMy 1st experiment show a slight bump (0.03) on the student model on local val.",
          "votes": 1
        },
        {
          "id": 717071,
          "postDate": "2020-01-12T17:28:54.810Z",
          "content": "<p>sounds like EfficientNet worked for you</p>",
          "rawMarkdown": "sounds like EfficientNet worked for you",
          "votes": 1
        },
        {
          "id": 717096,
          "postDate": "2020-01-12T18:15:23.550Z",
          "content": "<p>I'm a big fan of ResNe(X)ts but the strict resources constraints in this comp. forced me to make EffNets work 😄 </p>",
          "rawMarkdown": "I'm a big fan of ResNe(X)ts but the strict resources constraints in this comp. forced me to make EffNets work 😄 ",
          "votes": 1
        },
        {
          "id": 717098,
          "postDate": "2020-01-12T18:19:29.913Z",
          "content": "<p>glad to know that you made EfficientNets more efficient -&gt; Efficient(er)Net  😄 </p>",
          "rawMarkdown": "glad to know that you made EfficientNets more efficient -&gt; Efficient(er)Net  😄 ",
          "votes": 3
        }
      ]
    },
    {
      "id": 717775,
      "postDate": "2020-01-13T15:52:38.197Z",
      "content": "<p>The densenet121 tutored by the resnet34 ensemble scored 0.9574 LB.</p>",
      "rawMarkdown": "The densenet121 tutored by the resnet34 ensemble scored 0.9574 LB."
    },
    {
      "id": 716919,
      "postDate": "2020-01-12T13:08:49.247Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 713808,
      "author_name": "Miroslav Valan",
      "author_url": "",
      "post_date": "2020-01-08T17:11:28.003000",
      "content": "<p>Paper with current SOTA on imagenet (updated yesterday Jan 7th 2020) uses it <a href=\"https://arxiv.org/abs/1911.04252\">https://arxiv.org/abs/1911.04252</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 716923,
      "author_name": "Mighty Rains",
      "author_url": "",
      "post_date": "2020-01-12T13:15:31.737000",
      "content": "<p>I tried that out, although on a very superficial level. I trained a resnet18 on an ensemble of 5fold resnet34 scoring 0.9656 in LB and the student got 0.9527 LB. A densenet121 is being tutored at the moment by the same ensemble of teachers. I'll update the post when it does.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 716899,
      "author_name": "NguyenThanhNhan",
      "author_url": "",
      "post_date": "2020-01-12T12:37:54.667000",
      "content": "<p>I'm trying to distil from B3 to B0 right now. If things work well, I might go further and run distillation from an ensemble of teachers ✊ </p>",
      "votes": 1,
      "replies": [
        {
          "id": 716925,
          "author_name": "Mighty Rains",
          "author_url": "",
          "post_date": "2020-01-12T13:18:17.597000",
          "content": "<p>What loss are you using? I did a blend of KL divergence loss between teacher-student and a CCE between student-true labels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717059,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2020-01-12T17:05:10.183000",
          "content": "<p>Yes, same as you, adding a kl loss between teacher and student outputs for every mini-batch iteration along with the usual CE. \nMy 1st experiment show a slight bump (0.03) on the student model on local val.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717071,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-12T17:28:54.810000",
          "content": "<p>sounds like EfficientNet worked for you</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717096,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2020-01-12T18:15:23.550000",
          "content": "<p>I'm a big fan of ResNe(X)ts but the strict resources constraints in this comp. forced me to make EffNets work 😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717098,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-12T18:19:29.913000",
          "content": "<p>glad to know that you made EfficientNets more efficient -&gt; Efficient(er)Net  😄 </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 717775,
      "author_name": "Mighty Rains",
      "author_url": "",
      "post_date": "2020-01-13T15:52:38.197000",
      "content": "<p>The densenet121 tutored by the resnet34 ensemble scored 0.9574 LB.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 716919,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-12T13:08:49.247000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "713416": "Has anybody tried to train multiple ensembles of models into a single teacher and got a good performance?\nI mean not only in this competition, but also in others/outside kaggle.",
    "713808": "Paper with current SOTA on imagenet (updated yesterday Jan 7th 2020) uses it https://arxiv.org/abs/1911.04252",
    "716923": "I tried that out, although on a very superficial level. I trained a resnet18 on an ensemble of 5fold resnet34 scoring 0.9656 in LB and the student got 0.9527 LB. A densenet121 is being tutored at the moment by the same ensemble of teachers. I'll update the post when it does.",
    "716899": "I'm trying to distil from B3 to B0 right now. If things work well, I might go further and run distillation from an ensemble of teachers ✊ ",
    "717775": "The densenet121 tutored by the resnet34 ensemble scored 0.9574 LB.",
    "716919": ""
  }
}