{
  "id": 90412,
  "title": "How to pre-train the model",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/90412",
  "author_name": "",
  "post_date": "2019-04-23T16:16:22.880989700Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I don't have any experience with pretraining of such models. There are some suggestions on the forum that it's a good strategy to use noisy data as pretraining, so my question is how to do that? Are there any rules of thumbs, or good practices?</p>\n\n<p>For now, I'm wondering how long should such pre-training last and isn't there any risk that we will forget what we learned during pretraining, or that we won't be able to generalize to a different, less noisy domain?</p>",
  "messages": [
    {
      "id": "521916",
      "postDate": "04/23/2019 16:16:22",
      "content": "<p>I don't have any experience with pretraining of such models. There are some suggestions on the forum that it's a good strategy to use noisy data as pretraining, so my question is how to do that? Are there any rules of thumbs, or good practices?</p>\n\n<p>For now, I'm wondering how long should such pre-training last and isn't there any risk that we will forget what we learned during pretraining, or that we won't be able to generalize to a different, less noisy domain?</p>",
      "rawMarkdown": "I don't have any experience with pretraining of such models. There are some suggestions on the forum that it's a good strategy to use noisy data as pretraining, so my question is how to do that? Are there any rules of thumbs, or good practices?\n\nFor now, I'm wondering how long should such pre-training last and isn't there any risk that we will forget what we learned during pretraining, or that we won't be able to generalize to a different, less noisy domain?",
      "votes": null
    },
    {
      "id": "522072",
      "postDate": "04/23/2019 20:39:56",
      "content": "<p>There are a bunch of tricks on how you can actually overcome something that people call \"catastrophic forgetting\" - a phenomenon when the model \"forgets\" the past knowledge. I would recommend going through the <a href=\"https://arxiv.org/abs/1801.06146\">ULMFiT</a> paper that provides a lot of practical methods on how to finetune a model and make it use the past knowledge rather than simply overwrite the weights.</p>",
      "rawMarkdown": "There are a bunch of tricks on how you can actually overcome something that people call \"catastrophic forgetting\" - a phenomenon when the model \"forgets\" the past knowledge. I would recommend going through the [ULMFiT](https://arxiv.org/abs/1801.06146) paper that provides a lot of practical methods on how to finetune a model and make it use the past knowledge rather than simply overwrite the weights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 522072,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "04/23/2019 20:39:56",
      "content": "<p>There are a bunch of tricks on how you can actually overcome something that people call \"catastrophic forgetting\" - a phenomenon when the model \"forgets\" the past knowledge. I would recommend going through the <a href=\"https://arxiv.org/abs/1801.06146\">ULMFiT</a> paper that provides a lot of practical methods on how to finetune a model and make it use the past knowledge rather than simply overwrite the weights.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "521916": "I don't have any experience with pretraining of such models. There are some suggestions on the forum that it's a good strategy to use noisy data as pretraining, so my question is how to do that? Are there any rules of thumbs, or good practices?\n\nFor now, I'm wondering how long should such pre-training last and isn't there any risk that we will forget what we learned during pretraining, or that we won't be able to generalize to a different, less noisy domain?",
    "522072": "There are a bunch of tricks on how you can actually overcome something that people call \"catastrophic forgetting\" - a phenomenon when the model \"forgets\" the past knowledge. I would recommend going through the [ULMFiT](https://arxiv.org/abs/1801.06146) paper that provides a lot of practical methods on how to finetune a model and make it use the past knowledge rather than simply overwrite the weights."
  },
  "source": "meta"
}