{
  "id": 240586,
  "title": "Tuning on Longer Lengths",
  "url": "/competitions/bms-molecular-translation/discussion/240586",
  "author_name": "Andrew Shao",
  "post_date": "2021-05-20T13:01:15.802000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am currently trying the strategy of training initially on shorter(140) tokens and then fine-tuning on the longer ones, to speed up training. However, when I switch from 140 -&gt; Full length, I find that the performance actually worsens(Converges to a lower point than the previous checkpoint). Has anyone else experienced this. If so, how did you combat it?</p>",
  "messages": [
    {
      "id": 1316321,
      "postDate": "2021-05-20T13:01:15.803Z",
      "content": "<p>I am currently trying the strategy of training initially on shorter(140) tokens and then fine-tuning on the longer ones, to speed up training. However, when I switch from 140 -&gt; Full length, I find that the performance actually worsens(Converges to a lower point than the previous checkpoint). Has anyone else experienced this. If so, how did you combat it?</p>",
      "rawMarkdown": "I am currently trying the strategy of training initially on shorter(140) tokens and then fine-tuning on the longer ones, to speed up training. However, when I switch from 140 -> Full length, I find that the performance actually worsens(Converges to a lower point than the previous checkpoint). Has anyone else experienced this. If so, how did you combat it?",
      "votes": 3
    },
    {
      "id": 1316457,
      "postDate": "2021-05-20T15:14:07.520Z",
      "content": "<p>Do you use a fixed image size, or patches from the original image?<br>\nIf the former, then when you change to longer inchis, then you have to resize your images a lot more than for the shorter ones. Your model has never experienced this scale of molecules, so your encoder output most probably is gibberish. That's just my theory.</p>",
      "rawMarkdown": "Do you use a fixed image size, or patches from the original image?\nIf the former, then when you change to longer inchis, then you have to resize your images a lot more than for the shorter ones. Your model has never experienced this scale of molecules, so your encoder output most probably is gibberish. That's just my theory.",
      "votes": 2
    },
    {
      "id": 1317829,
      "postDate": "2021-05-21T18:12:09.770Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1316457,
      "author_name": "nofreewill42",
      "author_url": "",
      "post_date": "2021-05-20T15:14:07.520000",
      "content": "<p>Do you use a fixed image size, or patches from the original image?<br>\nIf the former, then when you change to longer inchis, then you have to resize your images a lot more than for the shorter ones. Your model has never experienced this scale of molecules, so your encoder output most probably is gibberish. That's just my theory.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1317829,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-21T18:12:09.770000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1316321": "I am currently trying the strategy of training initially on shorter(140) tokens and then fine-tuning on the longer ones, to speed up training. However, when I switch from 140 -> Full length, I find that the performance actually worsens(Converges to a lower point than the previous checkpoint). Has anyone else experienced this. If so, how did you combat it?",
    "1316457": "Do you use a fixed image size, or patches from the original image?\nIf the former, then when you change to longer inchis, then you have to resize your images a lot more than for the shorter ones. Your model has never experienced this scale of molecules, so your encoder output most probably is gibberish. That's just my theory.",
    "1317829": ""
  }
}