{
  "id": 232401,
  "title": "Possible Way To Speed Up Training - Idea ",
  "url": "/competitions/bms-molecular-translation/discussion/232401",
  "author_name": "Darien Schettler",
  "post_date": "2021-04-13T15:17:11.607000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi there.</p>\n<p>I had an idea for speeding up training. Please let me know if this makes sense as I will be attempting to implement it in the near future.</p>\n<hr>\n<p><strong>First we create 3 datasets to work from</strong></p>\n<ol>\n<li>All examples where the inchi string is less than a reasonable length (i.e. &lt;100 tokens).</li>\n<li>All examples where the inchi string is greater than a reasonable length (i.e. &gt;100 tokens)</li>\n<li>All examples</li>\n</ol>\n<blockquote>\n  <p><em>EDIT: Depending on how we do this, and the size of your instance(s), this might allow certain datasets to be cached in memory… although I doubt it. I'm still only utilizing Kaggle and caching to memory won't be possible regardless of how the dataset is sharded.</em></p>\n</blockquote>\n<hr>\n<p><strong>Then we follow this training protocol:</strong></p>\n<ol>\n<li><p>Initially we will train on dataset 1 which has a smaller max length and larger proportion of the examples (probably). <em>The benefit is 3-fold</em>:</p>\n<ul>\n<li>Decoding is limited to a relatively small sequence length compared to the max (100 v. ~300). This will result in our decoder predicting much less than if we trained directly on all sequences where for every example we would be predicting ~300 tokens. </li>\n<li>Training on smaller sequences will let us converge faster. We know that when an LSTM is converging, it will naturally optimize for shorter examples first before moving onto longer examples. We are essentially just embracing this by removing the longer sequences.</li>\n<li>Training on a smaller number of examples means we will converge faster. This is a trivial benefit as we are not actually done at this point…<br><br></li></ul></li>\n<li><p>Fine-tune on longer INCHI dataset.</p>\n<ul>\n<li>Our longer inchi dataset will most likely be smaller than our small inchi dataset in terms of the number of examples so we should see improvement relatively quickly.</li>\n<li>The decoding step will now take much longer… but at least we will be actually using the loss values accumulated past token 100.<br><br></li></ul></li>\n<li><p>Fine-tune on all examples</p>\n<ul>\n<li>Now that our model has been fine-tuned on the longer sequences we may have lost some of the intuition w.r.t. smaller sequences and overfit to the more complicated molecules.</li>\n<li>By finetuning we should be able to converge even further.</li></ul></li>\n</ol>\n<hr>\n<p><strong>Does this make sense? What do you think?</strong></p>",
  "messages": [
    {
      "id": 1272567,
      "postDate": "2021-04-13T15:17:11.607Z",
      "content": "<p>Hi there.</p>\n<p>I had an idea for speeding up training. Please let me know if this makes sense as I will be attempting to implement it in the near future.</p>\n<hr>\n<p><strong>First we create 3 datasets to work from</strong></p>\n<ol>\n<li>All examples where the inchi string is less than a reasonable length (i.e. &lt;100 tokens).</li>\n<li>All examples where the inchi string is greater than a reasonable length (i.e. &gt;100 tokens)</li>\n<li>All examples</li>\n</ol>\n<blockquote>\n  <p><em>EDIT: Depending on how we do this, and the size of your instance(s), this might allow certain datasets to be cached in memory… although I doubt it. I'm still only utilizing Kaggle and caching to memory won't be possible regardless of how the dataset is sharded.</em></p>\n</blockquote>\n<hr>\n<p><strong>Then we follow this training protocol:</strong></p>\n<ol>\n<li><p>Initially we will train on dataset 1 which has a smaller max length and larger proportion of the examples (probably). <em>The benefit is 3-fold</em>:</p>\n<ul>\n<li>Decoding is limited to a relatively small sequence length compared to the max (100 v. ~300). This will result in our decoder predicting much less than if we trained directly on all sequences where for every example we would be predicting ~300 tokens. </li>\n<li>Training on smaller sequences will let us converge faster. We know that when an LSTM is converging, it will naturally optimize for shorter examples first before moving onto longer examples. We are essentially just embracing this by removing the longer sequences.</li>\n<li>Training on a smaller number of examples means we will converge faster. This is a trivial benefit as we are not actually done at this point…<br><br></li></ul></li>\n<li><p>Fine-tune on longer INCHI dataset.</p>\n<ul>\n<li>Our longer inchi dataset will most likely be smaller than our small inchi dataset in terms of the number of examples so we should see improvement relatively quickly.</li>\n<li>The decoding step will now take much longer… but at least we will be actually using the loss values accumulated past token 100.<br><br></li></ul></li>\n<li><p>Fine-tune on all examples</p>\n<ul>\n<li>Now that our model has been fine-tuned on the longer sequences we may have lost some of the intuition w.r.t. smaller sequences and overfit to the more complicated molecules.</li>\n<li>By finetuning we should be able to converge even further.</li></ul></li>\n</ol>\n<hr>\n<p><strong>Does this make sense? What do you think?</strong></p>",
      "rawMarkdown": "Hi there.\n\nI had an idea for speeding up training. Please let me know if this makes sense as I will be attempting to implement it in the near future.\n\n---\n\n**First we create 3 datasets to work from**\n\n1. All examples where the inchi string is less than a reasonable length (i.e. <100 tokens).\n2. All examples where the inchi string is greater than a reasonable length (i.e. >100 tokens)\n3. All examples\n\n> *EDIT: Depending on how we do this, and the size of your instance(s), this might allow certain datasets to be cached in memory... although I doubt it. I'm still only utilizing Kaggle and caching to memory won't be possible regardless of how the dataset is sharded.*\n\n---\n\n**Then we follow this training protocol:**\n\n1. Initially we will train on dataset 1 which has a smaller max length and larger proportion of the examples (probably). *The benefit is 3-fold*:\n  * Decoding is limited to a relatively small sequence length compared to the max (100 v. ~300). This will result in our decoder predicting much less than if we trained directly on all sequences where for every example we would be predicting ~300 tokens. \n  * Training on smaller sequences will let us converge faster. We know that when an LSTM is converging, it will naturally optimize for shorter examples first before moving onto longer examples. We are essentially just embracing this by removing the longer sequences.\n  * Training on a smaller number of examples means we will converge faster. This is a trivial benefit as we are not actually done at this point...<br><br>\n\n2. Fine-tune on longer INCHI dataset.\n  * Our longer inchi dataset will most likely be smaller than our small inchi dataset in terms of the number of examples so we should see improvement relatively quickly.\n  * The decoding step will now take much longer... but at least we will be actually using the loss values accumulated past token 100.<br><br>\n\n3. Fine-tune on all examples\n  * Now that our model has been fine-tuned on the longer sequences we may have lost some of the intuition w.r.t. smaller sequences and overfit to the more complicated molecules.\n  * By finetuning we should be able to converge even further.\n\n---\n\n**Does this make sense? What do you think?**\n  ",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1272567": "Hi there.\n\nI had an idea for speeding up training. Please let me know if this makes sense as I will be attempting to implement it in the near future.\n\n---\n\n**First we create 3 datasets to work from**\n\n1. All examples where the inchi string is less than a reasonable length (i.e. <100 tokens).\n2. All examples where the inchi string is greater than a reasonable length (i.e. >100 tokens)\n3. All examples\n\n> *EDIT: Depending on how we do this, and the size of your instance(s), this might allow certain datasets to be cached in memory... although I doubt it. I'm still only utilizing Kaggle and caching to memory won't be possible regardless of how the dataset is sharded.*\n\n---\n\n**Then we follow this training protocol:**\n\n1. Initially we will train on dataset 1 which has a smaller max length and larger proportion of the examples (probably). *The benefit is 3-fold*:\n  * Decoding is limited to a relatively small sequence length compared to the max (100 v. ~300). This will result in our decoder predicting much less than if we trained directly on all sequences where for every example we would be predicting ~300 tokens. \n  * Training on smaller sequences will let us converge faster. We know that when an LSTM is converging, it will naturally optimize for shorter examples first before moving onto longer examples. We are essentially just embracing this by removing the longer sequences.\n  * Training on a smaller number of examples means we will converge faster. This is a trivial benefit as we are not actually done at this point...<br><br>\n\n2. Fine-tune on longer INCHI dataset.\n  * Our longer inchi dataset will most likely be smaller than our small inchi dataset in terms of the number of examples so we should see improvement relatively quickly.\n  * The decoding step will now take much longer... but at least we will be actually using the loss values accumulated past token 100.<br><br>\n\n3. Fine-tune on all examples\n  * Now that our model has been fine-tuned on the longer sequences we may have lost some of the intuition w.r.t. smaller sequences and overfit to the more complicated molecules.\n  * By finetuning we should be able to converge even further.\n\n---\n\n**Does this make sense? What do you think?**\n  "
  }
}