{
  "id": 245704,
  "title": "6th Place Solution - General Overview",
  "url": "/competitions/bms-molecular-translation/discussion/245704",
  "author_name": "Bojan Tunguz",
  "post_date": "2021-06-12T00:16:08.261000",
  "votes": 30,
  "comment_count": 0,
  "views": 0,
  "content": "<p>MODEL:</p>\n<p>Our basic models are similar to the general modeling that has been used in this competition: a combination of high-end vision models with either a recurrent net or a transformer for the chemical formula output. Early on we have experimented with several combination of the two, but for our final effort we’ve used the following three main combinations:</p>\n<p>Efficientnetv1(b5,b7) + transformer <br>\nEfficientnetv2(m,l) + transformer <br>\nB7 + lstm</p>\n<p>When blending, we found that blending lstm with transformers-based models did particularly well, oftentimes giving us an overall boost up to 0.2 on CV/LB.</p>\n<p>In the early stages of the competition, we blended probabilities, but subsequently this approach became unwieldy and we never managed to get it fully working for the entire solution. </p>\n<p>The kind of “blending” that <em>did</em> work was blending based on the correct form of the formula. It was noticed early on that the correctness of the prediction is proportional to the number of the formulas that are encoded correctly, </p>\n<p>SOLUTION:</p>\n<p>Training schedule:</p>\n<p>Initially we had used a simple cosine annealing for our training schedule, but then we switched to a modified semi-manual plateau training schedule with restarts. This is something that Bojan had used before in a few previous competitions (Bengali and Alaska), and found it very effective. The automated version of this training involves just the “ordinary” plateau schedule, with retraining once the training stops. For the retraining we would re-load the best weights, and start the training process again with the biggest learning rate. This was a very time-consuming training regiment, as one epoch on DGX-1/DGX-Station would take anywhere between 2 and 12 hours, depending on the model architecture and the resolution. </p>\n<p>Local validation scheme:</p>\n<p>We had set up our training with a 5-fold CV in mind, but due to the extremely long training time, we never managed to train more than one fold. So most of our models were based on the same 80-20 split of the training data. During the last week we retrained our models on the 96% of the data, with the 4% left out for validation, This worked well initially, but due to some quirks with the sampling of the 4% that we had used, the validation score started becoming unreasonably low (0.23) at one point, and from then on it was not possible to use local validation any more.<br>\nWe had noticed that our local oof validation and the LB score were very close (within less than 0.01 of each other), so we decided to mostly rely on our LB score for validation. This was not too risky in this competition, as it was a) an image competition, with b) an enormous dataset, and c) synthetic data. As there was very little shakeup in the end, our assumptions proved to be justified.</p>\n<p>1.Train a rotate detector. And fix the orientation in the test set. It makes CV-LB correlated very well.</p>\n<ol>\n<li><p>Dataprocess<br>\n2.1 We cut the edge to get more effective pixels, The width of the edge varies between 26 and 41 pixels in both train and test set, so we decided to uniformly cut 25 pixels in training and inference:<br>\nedge = 25<br>\nimage_raw = image_raw[edge:-edge, edge:-edge]<br>\n2.2 We noticed that most of the validation errors were for the formulas that were longer than 150 characters, so in our training we decided to oversample those data points. We had used different oversampling ratex, for instance for len(Inchi)&gt;150, len(Inchi)&gt;100, len(Inchi) &gt; 200, len(Inchi) &gt; 250, etc.<br>\n2.3 For most of our models we had kept the same h/w ratio, but later on in the competition we also included models that were trained on the square images.<br>\n3.4 Test set is more noisy than the train set, and it’s salt and pepper noise.<br>\ndef addSaltNoise(self,image,SNR=0.99):<br>\n  if SNR==1:<br>\n      return image<br>\n  h,w = image.shape<br>\n  noiseSize = int(h*w * (1 - SNR))<br>\n  for k in range(0, noiseSize):<br>\n      xi = int(random.uniform(0, image.shape[1]))<br>\n      xj = int(random.uniform(0, image.shape[0]))<br>\n      image[xj, xi] = 0<br>\n  return image</p></li>\n<li><p>Use different resolutions (256x256, 256x384,384x512,512x512,640x640).</p></li>\n</ol>\n<p>4.Check InCHi valid status. And try to get as many valid InCHi as possible. </p>\n<p>The final blend was created from 20+ predictions results. We ordered the predictions by local CV. For one image, find the first valid result, if there is no valid prediction in these 20+ results, then set the result with the lowest CV prediction.</p>\n<ol>\n<li><p>Use a different number of transformer layers. We’ve tried three, four, five and six layers. Our single best model used three layers, but it is conceivable that the V2 640x640 model with four layers could have done even better, but we ran out of time to fully tune it. </p></li>\n<li><p>Weights checkpoint blending. For our best single model (B7, 3 x transformer layers) we managed to blend about 6 different weight checkpoints. That helped our score by about 0.04. We also tried blending three different checkpoints for the B7 lstm model, but that only had a marginal impact on our model accuracy. </p></li>\n</ol>\n<p>WHAT DIDN’T WORK:</p>\n<p>Harder augmentation<br>\nVFlip, HFlip, Rotate<br>\nRandomLines augmentation<br>\nPseudolabeling<br>\nTraining more image rotation detectors<br>\nPreprocess images to close some lines and remove salt noise:<br>\n   img = cv2.imread(fn,0)<br>\n   kernel = np.ones((2,2),np.uint8)<br>\n   erosion = cv2.fastNlMeansDenoising(255-img, h=25)<br>\n   erosion = cv2.dilate(erosion,kernel,iterations = 1)<br>\n   erosion = 255-cv2.erode(erosion,kernel,iterations = 1)</p>\n<p>Multi-model probability blend had minimum boost in score.<br>\nFinetune a model specialized in small (384 pixels in any size).<br>\nFinetune a model specialized in small InChI strings (=150).</p>\n<p>OTHER THINGS WE WISH WE COULD HAVE DONE:</p>\n<p>Train all 5 folds<br>\nTrain models with different “mild” augmentations and blend those<br>\nTTA</p>\n<p>HARDWARE:</p>\n<p>We had been fortunate to have access to very generous hardware resources provided to us by Nvidia. Most of the models were trained on one DGX Station and up to three DGX-1 servers. Despite what our team’s name may suggest, we DID NOT have access to 1024 A100s. :) Halfway through the competition we got access to one DGX A100 station, which has 4 A100 GPUs, and close to the end of the competition we had access to another one of those machines. Full disclosure: Bojan is working with the Nvidia DGX A100 sales team as an influencer. These are awesome machines, and I’ll post another brief post about them soon.</p>",
  "messages": [
    {
      "id": 1345924,
      "postDate": "2021-06-12T00:16:08.260Z",
      "content": "<p>MODEL:</p>\n<p>Our basic models are similar to the general modeling that has been used in this competition: a combination of high-end vision models with either a recurrent net or a transformer for the chemical formula output. Early on we have experimented with several combination of the two, but for our final effort we’ve used the following three main combinations:</p>\n<p>Efficientnetv1(b5,b7) + transformer <br>\nEfficientnetv2(m,l) + transformer <br>\nB7 + lstm</p>\n<p>When blending, we found that blending lstm with transformers-based models did particularly well, oftentimes giving us an overall boost up to 0.2 on CV/LB.</p>\n<p>In the early stages of the competition, we blended probabilities, but subsequently this approach became unwieldy and we never managed to get it fully working for the entire solution. </p>\n<p>The kind of “blending” that <em>did</em> work was blending based on the correct form of the formula. It was noticed early on that the correctness of the prediction is proportional to the number of the formulas that are encoded correctly, </p>\n<p>SOLUTION:</p>\n<p>Training schedule:</p>\n<p>Initially we had used a simple cosine annealing for our training schedule, but then we switched to a modified semi-manual plateau training schedule with restarts. This is something that Bojan had used before in a few previous competitions (Bengali and Alaska), and found it very effective. The automated version of this training involves just the “ordinary” plateau schedule, with retraining once the training stops. For the retraining we would re-load the best weights, and start the training process again with the biggest learning rate. This was a very time-consuming training regiment, as one epoch on DGX-1/DGX-Station would take anywhere between 2 and 12 hours, depending on the model architecture and the resolution. </p>\n<p>Local validation scheme:</p>\n<p>We had set up our training with a 5-fold CV in mind, but due to the extremely long training time, we never managed to train more than one fold. So most of our models were based on the same 80-20 split of the training data. During the last week we retrained our models on the 96% of the data, with the 4% left out for validation, This worked well initially, but due to some quirks with the sampling of the 4% that we had used, the validation score started becoming unreasonably low (0.23) at one point, and from then on it was not possible to use local validation any more.<br>\nWe had noticed that our local oof validation and the LB score were very close (within less than 0.01 of each other), so we decided to mostly rely on our LB score for validation. This was not too risky in this competition, as it was a) an image competition, with b) an enormous dataset, and c) synthetic data. As there was very little shakeup in the end, our assumptions proved to be justified.</p>\n<p>1.Train a rotate detector. And fix the orientation in the test set. It makes CV-LB correlated very well.</p>\n<ol>\n<li><p>Dataprocess<br>\n2.1 We cut the edge to get more effective pixels, The width of the edge varies between 26 and 41 pixels in both train and test set, so we decided to uniformly cut 25 pixels in training and inference:<br>\nedge = 25<br>\nimage_raw = image_raw[edge:-edge, edge:-edge]<br>\n2.2 We noticed that most of the validation errors were for the formulas that were longer than 150 characters, so in our training we decided to oversample those data points. We had used different oversampling ratex, for instance for len(Inchi)&gt;150, len(Inchi)&gt;100, len(Inchi) &gt; 200, len(Inchi) &gt; 250, etc.<br>\n2.3 For most of our models we had kept the same h/w ratio, but later on in the competition we also included models that were trained on the square images.<br>\n3.4 Test set is more noisy than the train set, and it’s salt and pepper noise.<br>\ndef addSaltNoise(self,image,SNR=0.99):<br>\n  if SNR==1:<br>\n      return image<br>\n  h,w = image.shape<br>\n  noiseSize = int(h*w * (1 - SNR))<br>\n  for k in range(0, noiseSize):<br>\n      xi = int(random.uniform(0, image.shape[1]))<br>\n      xj = int(random.uniform(0, image.shape[0]))<br>\n      image[xj, xi] = 0<br>\n  return image</p></li>\n<li><p>Use different resolutions (256x256, 256x384,384x512,512x512,640x640).</p></li>\n</ol>\n<p>4.Check InCHi valid status. And try to get as many valid InCHi as possible. </p>\n<p>The final blend was created from 20+ predictions results. We ordered the predictions by local CV. For one image, find the first valid result, if there is no valid prediction in these 20+ results, then set the result with the lowest CV prediction.</p>\n<ol>\n<li><p>Use a different number of transformer layers. We’ve tried three, four, five and six layers. Our single best model used three layers, but it is conceivable that the V2 640x640 model with four layers could have done even better, but we ran out of time to fully tune it. </p></li>\n<li><p>Weights checkpoint blending. For our best single model (B7, 3 x transformer layers) we managed to blend about 6 different weight checkpoints. That helped our score by about 0.04. We also tried blending three different checkpoints for the B7 lstm model, but that only had a marginal impact on our model accuracy. </p></li>\n</ol>\n<p>WHAT DIDN’T WORK:</p>\n<p>Harder augmentation<br>\nVFlip, HFlip, Rotate<br>\nRandomLines augmentation<br>\nPseudolabeling<br>\nTraining more image rotation detectors<br>\nPreprocess images to close some lines and remove salt noise:<br>\n   img = cv2.imread(fn,0)<br>\n   kernel = np.ones((2,2),np.uint8)<br>\n   erosion = cv2.fastNlMeansDenoising(255-img, h=25)<br>\n   erosion = cv2.dilate(erosion,kernel,iterations = 1)<br>\n   erosion = 255-cv2.erode(erosion,kernel,iterations = 1)</p>\n<p>Multi-model probability blend had minimum boost in score.<br>\nFinetune a model specialized in small (384 pixels in any size).<br>\nFinetune a model specialized in small InChI strings (=150).</p>\n<p>OTHER THINGS WE WISH WE COULD HAVE DONE:</p>\n<p>Train all 5 folds<br>\nTrain models with different “mild” augmentations and blend those<br>\nTTA</p>\n<p>HARDWARE:</p>\n<p>We had been fortunate to have access to very generous hardware resources provided to us by Nvidia. Most of the models were trained on one DGX Station and up to three DGX-1 servers. Despite what our team’s name may suggest, we DID NOT have access to 1024 A100s. :) Halfway through the competition we got access to one DGX A100 station, which has 4 A100 GPUs, and close to the end of the competition we had access to another one of those machines. Full disclosure: Bojan is working with the Nvidia DGX A100 sales team as an influencer. These are awesome machines, and I’ll post another brief post about them soon.</p>",
      "rawMarkdown": "MODEL:\n\nOur basic models are similar to the general modeling that has been used in this competition: a combination of high-end vision models with either a recurrent net or a transformer for the chemical formula output. Early on we have experimented with several combination of the two, but for our final effort we’ve used the following three main combinations:\n\nEfficientnetv1(b5,b7) + transformer \nEfficientnetv2(m,l) + transformer \nB7 + lstm\n\nWhen blending, we found that blending lstm with transformers-based models did particularly well, oftentimes giving us an overall boost up to 0.2 on CV/LB.\n\nIn the early stages of the competition, we blended probabilities, but subsequently this approach became unwieldy and we never managed to get it fully working for the entire solution. \n\nThe kind of “blending” that *did* work was blending based on the correct form of the formula. It was noticed early on that the correctness of the prediction is proportional to the number of the formulas that are encoded correctly, \n \nSOLUTION:\n\nTraining schedule:\n\nInitially we had used a simple cosine annealing for our training schedule, but then we switched to a modified semi-manual plateau training schedule with restarts. This is something that Bojan had used before in a few previous competitions (Bengali and Alaska), and found it very effective. The automated version of this training involves just the “ordinary” plateau schedule, with retraining once the training stops. For the retraining we would re-load the best weights, and start the training process again with the biggest learning rate. This was a very time-consuming training regiment, as one epoch on DGX-1/DGX-Station would take anywhere between 2 and 12 hours, depending on the model architecture and the resolution. \n\nLocal validation scheme:\n\nWe had set up our training with a 5-fold CV in mind, but due to the extremely long training time, we never managed to train more than one fold. So most of our models were based on the same 80-20 split of the training data. During the last week we retrained our models on the 96% of the data, with the 4% left out for validation, This worked well initially, but due to some quirks with the sampling of the 4% that we had used, the validation score started becoming unreasonably low (0.23) at one point, and from then on it was not possible to use local validation any more.\nWe had noticed that our local oof validation and the LB score were very close (within less than 0.01 of each other), so we decided to mostly rely on our LB score for validation. This was not too risky in this competition, as it was a) an image competition, with b) an enormous dataset, and c) synthetic data. As there was very little shakeup in the end, our assumptions proved to be justified.\n \n1.Train a rotate detector. And fix the orientation in the test set. It makes CV-LB correlated very well.\n\n2. Dataprocess\n2.1 We cut the edge to get more effective pixels, The width of the edge varies between 26 and 41 pixels in both train and test set, so we decided to uniformly cut 25 pixels in training and inference:\nedge = 25\nimage_raw = image_raw[edge:-edge, edge:-edge]\n2.2 We noticed that most of the validation errors were for the formulas that were longer than 150 characters, so in our training we decided to oversample those data points. We had used different oversampling ratex, for instance for len(Inchi)>150, len(Inchi)>100, len(Inchi) > 200, len(Inchi) > 250, etc.\n2.3 For most of our models we had kept the same h/w ratio, but later on in the competition we also included models that were trained on the square images.\n3.4 Test set is more noisy than the train set, and it’s salt and pepper noise.\ndef addSaltNoise(self,image,SNR=0.99):\n      if SNR==1:\n          return image\n      h,w = image.shape\n      noiseSize = int(h*w * (1 - SNR))\n      for k in range(0, noiseSize):\n          xi = int(random.uniform(0, image.shape[1]))\n          xj = int(random.uniform(0, image.shape[0]))\n          image[xj, xi] = 0\n      return image\n\n3. Use different resolutions (256x256, 256x384,384x512,512x512,640x640).\n\n4.Check InCHi valid status. And try to get as many valid InCHi as possible. \n\nThe final blend was created from 20+ predictions results. We ordered the predictions by local CV. For one image, find the first valid result, if there is no valid prediction in these 20+ results, then set the result with the lowest CV prediction.\n\n5. Use a different number of transformer layers. We’ve tried three, four, five and six layers. Our single best model used three layers, but it is conceivable that the V2 640x640 model with four layers could have done even better, but we ran out of time to fully tune it. \n\n6. Weights checkpoint blending. For our best single model (B7, 3 x transformer layers) we managed to blend about 6 different weight checkpoints. That helped our score by about 0.04. We also tried blending three different checkpoints for the B7 lstm model, but that only had a marginal impact on our model accuracy. \n \nWHAT DIDN’T WORK:\n\nHarder augmentation\nVFlip, HFlip, Rotate\nRandomLines augmentation\nPseudolabeling\nTraining more image rotation detectors\nPreprocess images to close some lines and remove salt noise:\n   img = cv2.imread(fn,0)\n   kernel = np.ones((2,2),np.uint8)\n   erosion = cv2.fastNlMeansDenoising(255-img, h=25)\n   erosion = cv2.dilate(erosion,kernel,iterations = 1)\n   erosion = 255-cv2.erode(erosion,kernel,iterations = 1)\n\nMulti-model probability blend had minimum boost in score.\nFinetune a model specialized in small (<384 pixels in any size) and other in big images (>384 pixels in any size).\nFinetune a model specialized in small InChI strings (<150) and other in big strings (>=150).\n \n \nOTHER THINGS WE WISH WE COULD HAVE DONE:\n\nTrain all 5 folds\nTrain models with different “mild” augmentations and blend those\nTTA\n \nHARDWARE:\n\nWe had been fortunate to have access to very generous hardware resources provided to us by Nvidia. Most of the models were trained on one DGX Station and up to three DGX-1 servers. Despite what our team’s name may suggest, we DID NOT have access to 1024 A100s. :) Halfway through the competition we got access to one DGX A100 station, which has 4 A100 GPUs, and close to the end of the competition we had access to another one of those machines. Full disclosure: Bojan is working with the Nvidia DGX A100 sales team as an influencer. These are awesome machines, and I’ll post another brief post about them soon.\n",
      "votes": 30
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1345924": "MODEL:\n\nOur basic models are similar to the general modeling that has been used in this competition: a combination of high-end vision models with either a recurrent net or a transformer for the chemical formula output. Early on we have experimented with several combination of the two, but for our final effort we’ve used the following three main combinations:\n\nEfficientnetv1(b5,b7) + transformer \nEfficientnetv2(m,l) + transformer \nB7 + lstm\n\nWhen blending, we found that blending lstm with transformers-based models did particularly well, oftentimes giving us an overall boost up to 0.2 on CV/LB.\n\nIn the early stages of the competition, we blended probabilities, but subsequently this approach became unwieldy and we never managed to get it fully working for the entire solution. \n\nThe kind of “blending” that *did* work was blending based on the correct form of the formula. It was noticed early on that the correctness of the prediction is proportional to the number of the formulas that are encoded correctly, \n \nSOLUTION:\n\nTraining schedule:\n\nInitially we had used a simple cosine annealing for our training schedule, but then we switched to a modified semi-manual plateau training schedule with restarts. This is something that Bojan had used before in a few previous competitions (Bengali and Alaska), and found it very effective. The automated version of this training involves just the “ordinary” plateau schedule, with retraining once the training stops. For the retraining we would re-load the best weights, and start the training process again with the biggest learning rate. This was a very time-consuming training regiment, as one epoch on DGX-1/DGX-Station would take anywhere between 2 and 12 hours, depending on the model architecture and the resolution. \n\nLocal validation scheme:\n\nWe had set up our training with a 5-fold CV in mind, but due to the extremely long training time, we never managed to train more than one fold. So most of our models were based on the same 80-20 split of the training data. During the last week we retrained our models on the 96% of the data, with the 4% left out for validation, This worked well initially, but due to some quirks with the sampling of the 4% that we had used, the validation score started becoming unreasonably low (0.23) at one point, and from then on it was not possible to use local validation any more.\nWe had noticed that our local oof validation and the LB score were very close (within less than 0.01 of each other), so we decided to mostly rely on our LB score for validation. This was not too risky in this competition, as it was a) an image competition, with b) an enormous dataset, and c) synthetic data. As there was very little shakeup in the end, our assumptions proved to be justified.\n \n1.Train a rotate detector. And fix the orientation in the test set. It makes CV-LB correlated very well.\n\n2. Dataprocess\n2.1 We cut the edge to get more effective pixels, The width of the edge varies between 26 and 41 pixels in both train and test set, so we decided to uniformly cut 25 pixels in training and inference:\nedge = 25\nimage_raw = image_raw[edge:-edge, edge:-edge]\n2.2 We noticed that most of the validation errors were for the formulas that were longer than 150 characters, so in our training we decided to oversample those data points. We had used different oversampling ratex, for instance for len(Inchi)>150, len(Inchi)>100, len(Inchi) > 200, len(Inchi) > 250, etc.\n2.3 For most of our models we had kept the same h/w ratio, but later on in the competition we also included models that were trained on the square images.\n3.4 Test set is more noisy than the train set, and it’s salt and pepper noise.\ndef addSaltNoise(self,image,SNR=0.99):\n      if SNR==1:\n          return image\n      h,w = image.shape\n      noiseSize = int(h*w * (1 - SNR))\n      for k in range(0, noiseSize):\n          xi = int(random.uniform(0, image.shape[1]))\n          xj = int(random.uniform(0, image.shape[0]))\n          image[xj, xi] = 0\n      return image\n\n3. Use different resolutions (256x256, 256x384,384x512,512x512,640x640).\n\n4.Check InCHi valid status. And try to get as many valid InCHi as possible. \n\nThe final blend was created from 20+ predictions results. We ordered the predictions by local CV. For one image, find the first valid result, if there is no valid prediction in these 20+ results, then set the result with the lowest CV prediction.\n\n5. Use a different number of transformer layers. We’ve tried three, four, five and six layers. Our single best model used three layers, but it is conceivable that the V2 640x640 model with four layers could have done even better, but we ran out of time to fully tune it. \n\n6. Weights checkpoint blending. For our best single model (B7, 3 x transformer layers) we managed to blend about 6 different weight checkpoints. That helped our score by about 0.04. We also tried blending three different checkpoints for the B7 lstm model, but that only had a marginal impact on our model accuracy. \n \nWHAT DIDN’T WORK:\n\nHarder augmentation\nVFlip, HFlip, Rotate\nRandomLines augmentation\nPseudolabeling\nTraining more image rotation detectors\nPreprocess images to close some lines and remove salt noise:\n   img = cv2.imread(fn,0)\n   kernel = np.ones((2,2),np.uint8)\n   erosion = cv2.fastNlMeansDenoising(255-img, h=25)\n   erosion = cv2.dilate(erosion,kernel,iterations = 1)\n   erosion = 255-cv2.erode(erosion,kernel,iterations = 1)\n\nMulti-model probability blend had minimum boost in score.\nFinetune a model specialized in small (<384 pixels in any size) and other in big images (>384 pixels in any size).\nFinetune a model specialized in small InChI strings (<150) and other in big strings (>=150).\n \n \nOTHER THINGS WE WISH WE COULD HAVE DONE:\n\nTrain all 5 folds\nTrain models with different “mild” augmentations and blend those\nTTA\n \nHARDWARE:\n\nWe had been fortunate to have access to very generous hardware resources provided to us by Nvidia. Most of the models were trained on one DGX Station and up to three DGX-1 servers. Despite what our team’s name may suggest, we DID NOT have access to 1024 A100s. :) Halfway through the competition we got access to one DGX A100 station, which has 4 A100 GPUs, and close to the end of the competition we had access to another one of those machines. Full disclosure: Bojan is working with the Nvidia DGX A100 sales team as an influencer. These are awesome machines, and I’ll post another brief post about them soon.\n"
  }
}