{
  "id": 230170,
  "title": "transformer(decoder)+resnet baseline 55+",
  "url": "/competitions/bms-molecular-translation/discussion/230170",
  "author_name": "zzz0534",
  "post_date": "2021-04-02T11:05:14.090000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>*** some important data ***<br>\nencoder : resnet34<br>\npretrained : true<br>\ndecoder : transformer.decoder<br>\nlayers : 1<br>\nfine_tuning : true<br>\nepoches : 2<br>\ntrain samples : 900 k<br>\nvalid samples : 100 k<br>\nlr : 1e-4<br>\nschedule: cosine<br>\nembedding dimension : 512<br>\nevery epoch train time : 2.5h<br>\ninference time(~1600 k) : 12.5h<br>\nCV = LB:55+<br>\ndata split type from: <a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-preprocess-2</a><br>\n*** learning source ***<br>\nlearn a lot from below kaggers sharing and very thanks for their sharing:<br>\n<a href=\"https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention\" target=\"_blank\">https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention</a><br>\n<a href=\"https://www.kaggle.com/towardsentropy/pytorch-cnn-to-gru-strong-baseline-55-6\" target=\"_blank\">https://www.kaggle.com/towardsentropy/pytorch-cnn-to-gru-strong-baseline-55-6</a></p>\n<p>*** some experiments or puzzles  ***</p>\n<ol>\n<li><p>Speed<br>\nTraining 1 layer decoder, 900k(train) + 100k(valid+ieference) about 2.5h(in kaggle gpu)<br>\nPredicting the whole data will cost about 12.5h (in colab gpu)<br>\nIf you want to test more layer, prepare the time will cost alot. I have tried 3 layer model, and the inference process is so slow which cost 30h.</p></li>\n<li><p>Beam search<br>\nIn fact, I find the teacher forcing data can get 30+ good grades but the inference data only 50. So I want to use beam search to get a close grade. However, for transformer ,I couldn't find a efficient and simple way. Learn a lot from <a href=\"https://github.com/budzianowski/PyTorch-Beam-Search-Decoding\" target=\"_blank\">https://github.com/budzianowski/PyTorch-Beam-Search-Decoding</a>. It is more friendly for LSTM like. So I test a easy version but the time cost is very huge. I give up this idea.</p></li>\n<li><p>More epoches<br>\nMany sharing good results are from many epoches like 5, 10 and even 18! But I found my way don't have a good performance . After training 5 epoches, the CV is just smaller about 3 points.</p></li>\n<li><p>Whole data<br>\nI train the whole data for 3 epoches, and get the same result (It is amazing!)</p></li>\n</ol>\n<p>The train process is in here : <a href=\"https://www.kaggle.com/zzz0534/photo-trans2\" target=\"_blank\">https://www.kaggle.com/zzz0534/photo-trans2</a><br>\nIf you are interested in transformer and want to try it, welcome to use this to do your reference.<br>\nWelcome to share your different thoughts! I hope I can find some small technique to break the bottle.</p>",
  "messages": [
    {
      "id": 1260703,
      "postDate": "2021-04-02T11:05:14.090Z",
      "content": "<p>*** some important data ***<br>\nencoder : resnet34<br>\npretrained : true<br>\ndecoder : transformer.decoder<br>\nlayers : 1<br>\nfine_tuning : true<br>\nepoches : 2<br>\ntrain samples : 900 k<br>\nvalid samples : 100 k<br>\nlr : 1e-4<br>\nschedule: cosine<br>\nembedding dimension : 512<br>\nevery epoch train time : 2.5h<br>\ninference time(~1600 k) : 12.5h<br>\nCV = LB:55+<br>\ndata split type from: <a href=\"https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/inchi-preprocess-2</a><br>\n*** learning source ***<br>\nlearn a lot from below kaggers sharing and very thanks for their sharing:<br>\n<a href=\"https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention\" target=\"_blank\">https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention</a><br>\n<a href=\"https://www.kaggle.com/towardsentropy/pytorch-cnn-to-gru-strong-baseline-55-6\" target=\"_blank\">https://www.kaggle.com/towardsentropy/pytorch-cnn-to-gru-strong-baseline-55-6</a></p>\n<p>*** some experiments or puzzles  ***</p>\n<ol>\n<li><p>Speed<br>\nTraining 1 layer decoder, 900k(train) + 100k(valid+ieference) about 2.5h(in kaggle gpu)<br>\nPredicting the whole data will cost about 12.5h (in colab gpu)<br>\nIf you want to test more layer, prepare the time will cost alot. I have tried 3 layer model, and the inference process is so slow which cost 30h.</p></li>\n<li><p>Beam search<br>\nIn fact, I find the teacher forcing data can get 30+ good grades but the inference data only 50. So I want to use beam search to get a close grade. However, for transformer ,I couldn't find a efficient and simple way. Learn a lot from <a href=\"https://github.com/budzianowski/PyTorch-Beam-Search-Decoding\" target=\"_blank\">https://github.com/budzianowski/PyTorch-Beam-Search-Decoding</a>. It is more friendly for LSTM like. So I test a easy version but the time cost is very huge. I give up this idea.</p></li>\n<li><p>More epoches<br>\nMany sharing good results are from many epoches like 5, 10 and even 18! But I found my way don't have a good performance . After training 5 epoches, the CV is just smaller about 3 points.</p></li>\n<li><p>Whole data<br>\nI train the whole data for 3 epoches, and get the same result (It is amazing!)</p></li>\n</ol>\n<p>The train process is in here : <a href=\"https://www.kaggle.com/zzz0534/photo-trans2\" target=\"_blank\">https://www.kaggle.com/zzz0534/photo-trans2</a><br>\nIf you are interested in transformer and want to try it, welcome to use this to do your reference.<br>\nWelcome to share your different thoughts! I hope I can find some small technique to break the bottle.</p>",
      "rawMarkdown": "*** some important data ***\nencoder : resnet34\npretrained : true\ndecoder : transformer.decoder\nlayers : 1\nfine_tuning : true\nepoches : 2\ntrain samples : 900 k\nvalid samples : 100 k\nlr : 1e-4\nschedule: cosine\nembedding dimension : 512\nevery epoch train time : 2.5h\ninference time(~1600 k) : 12.5h\nCV = LB:55+\ndata split type from: https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\n*** learning source ***\nlearn a lot from below kaggers sharing and very thanks for their sharing:\nhttps://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention\nhttps://www.kaggle.com/towardsentropy/pytorch-cnn-to-gru-strong-baseline-55-6\n\n*** some experiments or puzzles  ***\n1. Speed\nTraining 1 layer decoder, 900k(train) + 100k(valid+ieference) about 2.5h(in kaggle gpu)\nPredicting the whole data will cost about 12.5h (in colab gpu)\nIf you want to test more layer, prepare the time will cost alot. I have tried 3 layer model, and the inference process is so slow which cost 30h.\n\n2. Beam search\nIn fact, I find the teacher forcing data can get 30+ good grades but the inference data only 50. So I want to use beam search to get a close grade. However, for transformer ,I couldn't find a efficient and simple way. Learn a lot from https://github.com/budzianowski/PyTorch-Beam-Search-Decoding. It is more friendly for LSTM like. So I test a easy version but the time cost is very huge. I give up this idea.\n\n3. More epoches\nMany sharing good results are from many epoches like 5, 10 and even 18! But I found my way don't have a good performance . After training 5 epoches, the CV is just smaller about 3 points.\n\n4. Whole data\nI train the whole data for 3 epoches, and get the same result (It is amazing!)\n\nThe train process is in here : https://www.kaggle.com/zzz0534/photo-trans2\nIf you are interested in transformer and want to try it, welcome to use this to do your reference.\nWelcome to share your different thoughts! I hope I can find some small technique to break the bottle.",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1260703": "*** some important data ***\nencoder : resnet34\npretrained : true\ndecoder : transformer.decoder\nlayers : 1\nfine_tuning : true\nepoches : 2\ntrain samples : 900 k\nvalid samples : 100 k\nlr : 1e-4\nschedule: cosine\nembedding dimension : 512\nevery epoch train time : 2.5h\ninference time(~1600 k) : 12.5h\nCV = LB:55+\ndata split type from: https://www.kaggle.com/yasufuminakama/inchi-preprocess-2\n*** learning source ***\nlearn a lot from below kaggers sharing and very thanks for their sharing:\nhttps://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention\nhttps://www.kaggle.com/towardsentropy/pytorch-cnn-to-gru-strong-baseline-55-6\n\n*** some experiments or puzzles  ***\n1. Speed\nTraining 1 layer decoder, 900k(train) + 100k(valid+ieference) about 2.5h(in kaggle gpu)\nPredicting the whole data will cost about 12.5h (in colab gpu)\nIf you want to test more layer, prepare the time will cost alot. I have tried 3 layer model, and the inference process is so slow which cost 30h.\n\n2. Beam search\nIn fact, I find the teacher forcing data can get 30+ good grades but the inference data only 50. So I want to use beam search to get a close grade. However, for transformer ,I couldn't find a efficient and simple way. Learn a lot from https://github.com/budzianowski/PyTorch-Beam-Search-Decoding. It is more friendly for LSTM like. So I test a easy version but the time cost is very huge. I give up this idea.\n\n3. More epoches\nMany sharing good results are from many epoches like 5, 10 and even 18! But I found my way don't have a good performance . After training 5 epoches, the CV is just smaller about 3 points.\n\n4. Whole data\nI train the whole data for 3 epoches, and get the same result (It is amazing!)\n\nThe train process is in here : https://www.kaggle.com/zzz0534/photo-trans2\nIf you are interested in transformer and want to try it, welcome to use this to do your reference.\nWelcome to share your different thoughts! I hope I can find some small technique to break the bottle."
  }
}