{
  "id": 243767,
  "title": "Interesting Experiments & Learnings",
  "url": "/competitions/bms-molecular-translation/discussion/243767",
  "author_name": "Gabriel Lindenmaier",
  "post_date": "2021-06-04T00:23:57.206000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello There,</p>\n<p>I hope everyone is (halfway) satisfied with there final ranking.</p>\n<p>The Conclusion section below functions as TLDR. <br>\nAll sections are mostly independent, so read what seems interesting.</p>\n<p>This is my summary of the interesting parts of my paper reproductions and experiments. I focused on those, instead of a high score, so I decided to not submit anything but experiment until the end - I will come back to that.</p>\n<p>I post this text to hopefully give some of you useful ideas. Every section adds something new I haven't read in the discussions (Didn't really check notebooks out).<br>\nI have implemented everything myself as long as there was no library for that purpose (<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a>, <a href=\"https://hydra.cc/docs/intro/\" target=\"_blank\">hydra</a>, <a href=\"https://github.com/catalyst-team/catalyst\" target=\"_blank\">catalyst</a>, …); except for the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223423\" target=\"_blank\">path template</a> by <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>.</p>\n<h1>Tokenization &amp; Cross-Validation Split</h1>\n<p>I used this regex:<br>\n<code>r'([A-Z][a-z]?|(?=\\d{3})\\d|\\d{1,2}|/[a-z]|\\W)'; token_list = re.findall(..)</code></p>\n<p>Note, that it splits e.g. 107 into 1 &amp; 07 or 167 into 1 &amp; 67. That is for generalisation to all numbers between 100-199 as there weren't InchIs for all of those in the training set. It also rises the number of occurrences of the numerical tokens.<br>\nYou could of course use a extra token for the 10² and 10¹ positions to not get e.g. 07. So 107 -&gt; ħ ð 7. But that seemed like over-engineering. </p>\n<p><strong>5-Fold CV-split:</strong><br>\nI used the InchI token lengths split into buckets and the number of the atoms and InChI sections in the InchI string for the CV split. I treated it as multi-label classification for the CV splitting. As I didn't find a library which worked with this amount of 'labels' and data-rows I used a simplified own splitter. I sorted the feature label rows lexically (like in Excel when you sort after column A, then B…). Then I assigned the first row to fold 1, the second to fold 2, etc. the sixth to fold 1 again. That assured that similar features co-occurred at least somewhat through the folds.</p>\n<h1>General Architecture</h1>\n<p>I used a CNN-Encoder and Transformer-Decoder architecture.<br>\n<strong>Encoder</strong>: For the CNN I used <a href=\"https://arxiv.org/abs/2006.14090v4\" target=\"_blank\">GENet</a> Normal ('gernet_m' in timm library). It is supposed to be a very fast CNN on GPU with a decent ImageNet performance of <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv\" target=\"_blank\">80.7% top-1 accuracy</a>. This is true as far as I can tell. GENet Normal is 4x faster as EfficientNet V2S (version optimized for Tensor Cores, <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv\" target=\"_blank\">83.8% top-1</a>) on my Titan RTX with 17M vs 19.8M parameters (no classification heads &amp; feature boom/upscale layer). </p>\n<p><strong>Decoder</strong>: I used <a href=\"https://arxiv.org/abs/2009.04534\" target=\"_blank\">Pay-Attention-when-Required</a> (PAR) Transformer. It uses more  feedforward blocks after every attention block. And it has more of the former at the end for further feature refinement. It uses <a href=\"https://arxiv.org/abs/1901.02860\" target=\"_blank\">Transformer-XL</a> attention and works as well as it with up to ~50% speed-up. <br>\n'A' = Attention block; 'F' = FeedForward Block</p>\n<ul>\n<li>Normal Transformer: N * 'AF' </li>\n<li>PAR-Transformer for enwik8: 5 * 'AFF' + 9 * 'F'</li>\n<li>Own PAR-Transformer enwik8 baseline: 4 * 'AFFF' + 5 * 'F'</li>\n<li>Not sure what the best version of encoder-decoder architecture is. It works indeed, contrary to my hypothesis a few days ago. For this kind of data something along this might be a good starting point: 4 * 'AFF' + 'FF'. Two 'F' instead of three seem to work better (for me). </li>\n</ul>\n<p><strong>Didn't work:</strong></p>\n<ul>\n<li>Using a 'AFF' PAR Transformer after the CNN for global attention did never work out. Maybe only with shallow CNN. I used adaptive embeddings described below (like in the decoder). I tried a lot of variations to get it to work.</li>\n<li>CBAM attention in the GENet (applied through timm library by modifiying build config for gernet_m. It is also slows training down too much. And don't even get started with Triplet-Attention. It is even far slower.</li>\n<li>Better performance with GENet Large then Normal - but back then I still used not-pretrained weights &amp; the 'AFF' encoder block.</li>\n<li>Using higher channel/block count in GENet for earlier building blocks (2nd basic block). I thought might be useful as the images are all fine lines. The results weren't showing a significant difference</li>\n</ul>\n<p><strong>Unused:</strong></p>\n<ul>\n<li>SiLU activation instead of ReLU. It works better (1 percentage point in top-1 token accuracy). But the pretrained model works (of course) even better. Don't know why they didn't use it in the original. Additional 11% speed overhead. But maybe on ImageNet not that useful for this CNN?!</li>\n</ul>\n<p><strong>Verdict:</strong></p>\n<ul>\n<li><p>Best estimated performance possible with everything I used (see also other sections): 1.5 LD.<br>\n(Although I just realized a few hours ago that I forgot to change mean &amp; std for mormalization back to those of ImageNet with pretrained model. Maybe that improves performance.)</p>\n<p>GENent normal as used + adaptive positional embeddings (see below): 18.17M Parameters.<br>\nDecoder: varies 43M-47M. Speed: 3:10 hours with 40k validation set only &amp; 224x284 images on Titan RTX with <br>\nmixed-precision training. BUT this is mostly the layer depth. Swapping 'F' with 'A' doesn't slow <br>\ndecoder down [sic!]. (Decoder to deep). Also Transformer-XL  attention is ~1/3 slower than <br>\nvanilla one in my implementation. So use better encoder for better LD.</p></li>\n</ul>\n<h1>Adaptive Positional Embeddings</h1>\n<p>From <a href=\"https://arxiv.org/pdf/1910.04396.pdf\" target=\"_blank\">On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention</a>, a paper suggested by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (Thank you!)<br>\nIt uses the global average pooled vector of the CNN to compute two scale factors (feedforward-net + sigmoid) for the horizontal &amp; vertical positional embeddings. For the later it uses the standard transformer sinusoidal positional embeddings. The positional embeddings are for the CNN output image features when you attend to them via attention. Improves results a little bit - I guess it helps for images with originally different side ratios. Probably far better for areas with really distorted objects like in the paper.<br>\n<strong>Didn't work:</strong> Adaptive positional embeddings for the decoder (text sequence) in all kind of variations.</p>\n<h1>COOL Training (Model-Head Ensembling)</h1>\n<p>From <a href=\"https://arxiv.org/pdf/1609.02226.pdf\" target=\"_blank\">Fitted Learning: Models with Awareness of their Limits</a><br>\nIt uses several output-heads (degree of overcompleteness (DOO)) for one class. Not directly an ensemble as the outputs are in 'competition' with each other during training. And during evaluation you multiply the Softmax results of the heads for one class together.<br>\nI was quite suspicious while reading the paper as they only trained it for MNIST &amp; CIFAR10. They don't have much data, and are easy anyway. BUT far more importantly they have only 10 classes!<br>\nBut it was very easy to implement: For n class-head expand &amp; use target-label 1/n. Train as given. For evaluation reshape tensor, multiply class heads together.<br>\nNeither 16, nor 4 heads/DOO did work. This includes using Weight-Dropout (WD) for the output heads like in the <a href=\"https://www.paperswithcode.com/method/awd-lstm\" target=\"_blank\">AWD-LSTM</a>, which made performance a bit better. 2 percentage-points worse for 4 heads &amp; 1 for 4 heads + WD.</p>\n<h1>Augmentations</h1>\n<p>I used for all augmentations the probability 0f 15%, except the noise, which was 33%. The order in which they are presented is equal to their order in the pipeline. My goal was to mimic the worse images with the augmentations and force generalisation</p>\n<ul>\n<li>Inserting a white rectangle at a random position between 5-20% of the image height/width</li>\n<li>Dropping between 5-20% of the pixels in the images - making them white</li>\n<li>Using torchvisions <em>RandomPerpective</em> with distortion_scale=0.3 to transform perspective slightly</li>\n<li>Adding random noise which was based on random black pixels of 0.01-0.2% of the image which where then rescaled to mimic the look of original noise after scaling (down &amp; up). Regarding the percentages I estimated the noise in the test-set.</li>\n<li>Normalization of course<br>\n==&gt; First 3 augmentations improved validation score (can't determine anymore how much).</li>\n</ul>\n<p><strong>What didn't work:</strong></p>\n<ul>\n<li>Scaling the images down &amp; up to  mimic images of extreme sizes. Although I used <em>Nearest</em> as interpolation which made the quality even worse. Might work with better interpolation</li>\n<li><a href=\"https://imgaug.readthedocs.io/en/latest/source/overview/geometric.html?highlight=jigsaw#jigsaw\" target=\"_blank\">JigSaw</a>, which splits the image into several equally-sized rectangles and shifts them around.  This is supposed to enforce better global feature learning through its adversarial interruption of local features. Even a very 'friendly' JigSaw split of 4 rectangles didn't work out.</li>\n<li>Higher augmentation probabilities (values &amp; usage)</li>\n</ul>\n<h1>Miscellaneous</h1>\n<ul>\n<li><a href=\"https://www.paperswithcode.com/method/label-smoothing\" target=\"_blank\">Label smoothing</a> did not work, neither epsilon = 0.1 nor 0.01 (see also <em>COOL Training</em>)</li>\n<li>Adding the original image size ration to the encoder output did not work</li>\n<li>I used the PIL library for image resizing. Note that among many other libraries torchvision has a faulty implementation for resizing which lets horizontal/vertical lines vanish - even with bicubic interpolation. But you can use PIL images directly with torchvision and it works fine.</li>\n<li>dSiLU activation (derivative formula of SiLU as activation) in the last feedforward blocks, as suggested by this <a href=\"https://arxiv.org/abs/1702.03118\" target=\"_blank\">Reinforcement-Learning paper</a>. Might be the fickle RL, different task or non-reproducibility.</li>\n<li>Couldn't reproduce Transformer-N from <a href=\"https://arxiv.org/pdf/2104.03474.pdf\" target=\"_blank\">Revisiting Simple Neural Probabilistic Language Models</a>. But had to guess exact layer architecture.</li>\n</ul>\n<h1>Conclusion</h1>\n<ul>\n<li>GENet is a very fast image classification CNN which gives decent performance, but not enough for really high rankings. Would use more powerful encoder and shallower decoder (speed) next time - or something entirely different.</li>\n<li>Transformer-Decoder can benefit from more then one feedforward-block between attention blocks</li>\n<li>Adaptive positional embedding for CNN encoder features seems likely to be reproducible and gives small performance boost for molecule images</li>\n<li>Augmentations that mimic worse looking images work well. Distorting image perspective a bit also works.</li>\n<li>Did not work: COOL-Training, dSiLU activation in last layers, JigSaw augmentation, attention within CNN, label smoothing, shallow transformer encoder after deep CNN encoder for global attention</li>\n</ul>\n<h1>My View</h1>\n<p>I had quite a lot of fun doing all this, although I have not much to show for it (except my code should I publish it). Given that I was knew to computer vision apart from my very first project years ago, I had the opportunity to learn quite some things.<br>\nI also enjoyed most of the discussions.</p>\n<p>(Why did I not submit anything: Wanted to save myself the stress behind reaching a decent score within the last days - I also had neither a rotation classifier, already resized test images or the code for test-set inference, only validation)</p>\n<p>At least in the next competition I will go for a high score and see how motivating it is. Just doing endless training to get the highest possible performance didn't motivate me in this competition. I will have to see if the normal way of <em>kaggling</em> is something for me.</p>\n<p>Note: If you use anything novel for you from my learnings here then please link this post.</p>\n<p>Thank You :)</p>",
  "messages": [
    {
      "id": 1334993,
      "postDate": "2021-06-04T00:23:57.207Z",
      "content": "<p>Hello There,</p>\n<p>I hope everyone is (halfway) satisfied with there final ranking.</p>\n<p>The Conclusion section below functions as TLDR. <br>\nAll sections are mostly independent, so read what seems interesting.</p>\n<p>This is my summary of the interesting parts of my paper reproductions and experiments. I focused on those, instead of a high score, so I decided to not submit anything but experiment until the end - I will come back to that.</p>\n<p>I post this text to hopefully give some of you useful ideas. Every section adds something new I haven't read in the discussions (Didn't really check notebooks out).<br>\nI have implemented everything myself as long as there was no library for that purpose (<a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a>, <a href=\"https://hydra.cc/docs/intro/\" target=\"_blank\">hydra</a>, <a href=\"https://github.com/catalyst-team/catalyst\" target=\"_blank\">catalyst</a>, …); except for the <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/223423\" target=\"_blank\">path template</a> by <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>.</p>\n<h1>Tokenization &amp; Cross-Validation Split</h1>\n<p>I used this regex:<br>\n<code>r'([A-Z][a-z]?|(?=\\d{3})\\d|\\d{1,2}|/[a-z]|\\W)'; token_list = re.findall(..)</code></p>\n<p>Note, that it splits e.g. 107 into 1 &amp; 07 or 167 into 1 &amp; 67. That is for generalisation to all numbers between 100-199 as there weren't InchIs for all of those in the training set. It also rises the number of occurrences of the numerical tokens.<br>\nYou could of course use a extra token for the 10² and 10¹ positions to not get e.g. 07. So 107 -&gt; ħ ð 7. But that seemed like over-engineering. </p>\n<p><strong>5-Fold CV-split:</strong><br>\nI used the InchI token lengths split into buckets and the number of the atoms and InChI sections in the InchI string for the CV split. I treated it as multi-label classification for the CV splitting. As I didn't find a library which worked with this amount of 'labels' and data-rows I used a simplified own splitter. I sorted the feature label rows lexically (like in Excel when you sort after column A, then B…). Then I assigned the first row to fold 1, the second to fold 2, etc. the sixth to fold 1 again. That assured that similar features co-occurred at least somewhat through the folds.</p>\n<h1>General Architecture</h1>\n<p>I used a CNN-Encoder and Transformer-Decoder architecture.<br>\n<strong>Encoder</strong>: For the CNN I used <a href=\"https://arxiv.org/abs/2006.14090v4\" target=\"_blank\">GENet</a> Normal ('gernet_m' in timm library). It is supposed to be a very fast CNN on GPU with a decent ImageNet performance of <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv\" target=\"_blank\">80.7% top-1 accuracy</a>. This is true as far as I can tell. GENet Normal is 4x faster as EfficientNet V2S (version optimized for Tensor Cores, <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv\" target=\"_blank\">83.8% top-1</a>) on my Titan RTX with 17M vs 19.8M parameters (no classification heads &amp; feature boom/upscale layer). </p>\n<p><strong>Decoder</strong>: I used <a href=\"https://arxiv.org/abs/2009.04534\" target=\"_blank\">Pay-Attention-when-Required</a> (PAR) Transformer. It uses more  feedforward blocks after every attention block. And it has more of the former at the end for further feature refinement. It uses <a href=\"https://arxiv.org/abs/1901.02860\" target=\"_blank\">Transformer-XL</a> attention and works as well as it with up to ~50% speed-up. <br>\n'A' = Attention block; 'F' = FeedForward Block</p>\n<ul>\n<li>Normal Transformer: N * 'AF' </li>\n<li>PAR-Transformer for enwik8: 5 * 'AFF' + 9 * 'F'</li>\n<li>Own PAR-Transformer enwik8 baseline: 4 * 'AFFF' + 5 * 'F'</li>\n<li>Not sure what the best version of encoder-decoder architecture is. It works indeed, contrary to my hypothesis a few days ago. For this kind of data something along this might be a good starting point: 4 * 'AFF' + 'FF'. Two 'F' instead of three seem to work better (for me). </li>\n</ul>\n<p><strong>Didn't work:</strong></p>\n<ul>\n<li>Using a 'AFF' PAR Transformer after the CNN for global attention did never work out. Maybe only with shallow CNN. I used adaptive embeddings described below (like in the decoder). I tried a lot of variations to get it to work.</li>\n<li>CBAM attention in the GENet (applied through timm library by modifiying build config for gernet_m. It is also slows training down too much. And don't even get started with Triplet-Attention. It is even far slower.</li>\n<li>Better performance with GENet Large then Normal - but back then I still used not-pretrained weights &amp; the 'AFF' encoder block.</li>\n<li>Using higher channel/block count in GENet for earlier building blocks (2nd basic block). I thought might be useful as the images are all fine lines. The results weren't showing a significant difference</li>\n</ul>\n<p><strong>Unused:</strong></p>\n<ul>\n<li>SiLU activation instead of ReLU. It works better (1 percentage point in top-1 token accuracy). But the pretrained model works (of course) even better. Don't know why they didn't use it in the original. Additional 11% speed overhead. But maybe on ImageNet not that useful for this CNN?!</li>\n</ul>\n<p><strong>Verdict:</strong></p>\n<ul>\n<li><p>Best estimated performance possible with everything I used (see also other sections): 1.5 LD.<br>\n(Although I just realized a few hours ago that I forgot to change mean &amp; std for mormalization back to those of ImageNet with pretrained model. Maybe that improves performance.)</p>\n<p>GENent normal as used + adaptive positional embeddings (see below): 18.17M Parameters.<br>\nDecoder: varies 43M-47M. Speed: 3:10 hours with 40k validation set only &amp; 224x284 images on Titan RTX with <br>\nmixed-precision training. BUT this is mostly the layer depth. Swapping 'F' with 'A' doesn't slow <br>\ndecoder down [sic!]. (Decoder to deep). Also Transformer-XL  attention is ~1/3 slower than <br>\nvanilla one in my implementation. So use better encoder for better LD.</p></li>\n</ul>\n<h1>Adaptive Positional Embeddings</h1>\n<p>From <a href=\"https://arxiv.org/pdf/1910.04396.pdf\" target=\"_blank\">On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention</a>, a paper suggested by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (Thank you!)<br>\nIt uses the global average pooled vector of the CNN to compute two scale factors (feedforward-net + sigmoid) for the horizontal &amp; vertical positional embeddings. For the later it uses the standard transformer sinusoidal positional embeddings. The positional embeddings are for the CNN output image features when you attend to them via attention. Improves results a little bit - I guess it helps for images with originally different side ratios. Probably far better for areas with really distorted objects like in the paper.<br>\n<strong>Didn't work:</strong> Adaptive positional embeddings for the decoder (text sequence) in all kind of variations.</p>\n<h1>COOL Training (Model-Head Ensembling)</h1>\n<p>From <a href=\"https://arxiv.org/pdf/1609.02226.pdf\" target=\"_blank\">Fitted Learning: Models with Awareness of their Limits</a><br>\nIt uses several output-heads (degree of overcompleteness (DOO)) for one class. Not directly an ensemble as the outputs are in 'competition' with each other during training. And during evaluation you multiply the Softmax results of the heads for one class together.<br>\nI was quite suspicious while reading the paper as they only trained it for MNIST &amp; CIFAR10. They don't have much data, and are easy anyway. BUT far more importantly they have only 10 classes!<br>\nBut it was very easy to implement: For n class-head expand &amp; use target-label 1/n. Train as given. For evaluation reshape tensor, multiply class heads together.<br>\nNeither 16, nor 4 heads/DOO did work. This includes using Weight-Dropout (WD) for the output heads like in the <a href=\"https://www.paperswithcode.com/method/awd-lstm\" target=\"_blank\">AWD-LSTM</a>, which made performance a bit better. 2 percentage-points worse for 4 heads &amp; 1 for 4 heads + WD.</p>\n<h1>Augmentations</h1>\n<p>I used for all augmentations the probability 0f 15%, except the noise, which was 33%. The order in which they are presented is equal to their order in the pipeline. My goal was to mimic the worse images with the augmentations and force generalisation</p>\n<ul>\n<li>Inserting a white rectangle at a random position between 5-20% of the image height/width</li>\n<li>Dropping between 5-20% of the pixels in the images - making them white</li>\n<li>Using torchvisions <em>RandomPerpective</em> with distortion_scale=0.3 to transform perspective slightly</li>\n<li>Adding random noise which was based on random black pixels of 0.01-0.2% of the image which where then rescaled to mimic the look of original noise after scaling (down &amp; up). Regarding the percentages I estimated the noise in the test-set.</li>\n<li>Normalization of course<br>\n==&gt; First 3 augmentations improved validation score (can't determine anymore how much).</li>\n</ul>\n<p><strong>What didn't work:</strong></p>\n<ul>\n<li>Scaling the images down &amp; up to  mimic images of extreme sizes. Although I used <em>Nearest</em> as interpolation which made the quality even worse. Might work with better interpolation</li>\n<li><a href=\"https://imgaug.readthedocs.io/en/latest/source/overview/geometric.html?highlight=jigsaw#jigsaw\" target=\"_blank\">JigSaw</a>, which splits the image into several equally-sized rectangles and shifts them around.  This is supposed to enforce better global feature learning through its adversarial interruption of local features. Even a very 'friendly' JigSaw split of 4 rectangles didn't work out.</li>\n<li>Higher augmentation probabilities (values &amp; usage)</li>\n</ul>\n<h1>Miscellaneous</h1>\n<ul>\n<li><a href=\"https://www.paperswithcode.com/method/label-smoothing\" target=\"_blank\">Label smoothing</a> did not work, neither epsilon = 0.1 nor 0.01 (see also <em>COOL Training</em>)</li>\n<li>Adding the original image size ration to the encoder output did not work</li>\n<li>I used the PIL library for image resizing. Note that among many other libraries torchvision has a faulty implementation for resizing which lets horizontal/vertical lines vanish - even with bicubic interpolation. But you can use PIL images directly with torchvision and it works fine.</li>\n<li>dSiLU activation (derivative formula of SiLU as activation) in the last feedforward blocks, as suggested by this <a href=\"https://arxiv.org/abs/1702.03118\" target=\"_blank\">Reinforcement-Learning paper</a>. Might be the fickle RL, different task or non-reproducibility.</li>\n<li>Couldn't reproduce Transformer-N from <a href=\"https://arxiv.org/pdf/2104.03474.pdf\" target=\"_blank\">Revisiting Simple Neural Probabilistic Language Models</a>. But had to guess exact layer architecture.</li>\n</ul>\n<h1>Conclusion</h1>\n<ul>\n<li>GENet is a very fast image classification CNN which gives decent performance, but not enough for really high rankings. Would use more powerful encoder and shallower decoder (speed) next time - or something entirely different.</li>\n<li>Transformer-Decoder can benefit from more then one feedforward-block between attention blocks</li>\n<li>Adaptive positional embedding for CNN encoder features seems likely to be reproducible and gives small performance boost for molecule images</li>\n<li>Augmentations that mimic worse looking images work well. Distorting image perspective a bit also works.</li>\n<li>Did not work: COOL-Training, dSiLU activation in last layers, JigSaw augmentation, attention within CNN, label smoothing, shallow transformer encoder after deep CNN encoder for global attention</li>\n</ul>\n<h1>My View</h1>\n<p>I had quite a lot of fun doing all this, although I have not much to show for it (except my code should I publish it). Given that I was knew to computer vision apart from my very first project years ago, I had the opportunity to learn quite some things.<br>\nI also enjoyed most of the discussions.</p>\n<p>(Why did I not submit anything: Wanted to save myself the stress behind reaching a decent score within the last days - I also had neither a rotation classifier, already resized test images or the code for test-set inference, only validation)</p>\n<p>At least in the next competition I will go for a high score and see how motivating it is. Just doing endless training to get the highest possible performance didn't motivate me in this competition. I will have to see if the normal way of <em>kaggling</em> is something for me.</p>\n<p>Note: If you use anything novel for you from my learnings here then please link this post.</p>\n<p>Thank You :)</p>",
      "rawMarkdown": "Hello There,\n\nI hope everyone is (halfway) satisfied with there final ranking.\n\nThe Conclusion section below functions as TLDR. \nAll sections are mostly independent, so read what seems interesting.\n\nThis is my summary of the interesting parts of my paper reproductions and experiments. I focused on those, instead of a high score, so I decided to not submit anything but experiment until the end - I will come back to that.\n\nI post this text to hopefully give some of you useful ideas. Every section adds something new I haven't read in the discussions (Didn't really check notebooks out).\nI have implemented everything myself as long as there was no library for that purpose ([timm](https://github.com/rwightman/pytorch-image-models), [hydra](https://hydra.cc/docs/intro/), [catalyst](https://github.com/catalyst-team/catalyst), ...); except for the [path template](https://www.kaggle.com/c/bms-molecular-translation/discussion/223423) by @ihelon.\n\n# Tokenization & Cross-Validation Split\nI used this regex:\n`r'([A-Z][a-z]?|(?=\\d{3})\\d|\\d{1,2}|/[a-z]|\\W)'; token_list = re.findall(..)`\n\nNote, that it splits e.g. 107 into 1 & 07 or 167 into 1 & 67. That is for generalisation to all numbers between 100-199 as there weren't InchIs for all of those in the training set. It also rises the number of occurrences of the numerical tokens.\nYou could of course use a extra token for the 10² and 10¹ positions to not get e.g. 07. So 107 -> ħ ð 7. But that seemed like over-engineering. \n\n**5-Fold CV-split:**\nI used the InchI token lengths split into buckets and the number of the atoms and InChI sections in the InchI string for the CV split. I treated it as multi-label classification for the CV splitting. As I didn't find a library which worked with this amount of 'labels' and data-rows I used a simplified own splitter. I sorted the feature label rows lexically (like in Excel when you sort after column A, then B...). Then I assigned the first row to fold 1, the second to fold 2, etc. the sixth to fold 1 again. That assured that similar features co-occurred at least somewhat through the folds.\n\n# General Architecture\nI used a CNN-Encoder and Transformer-Decoder architecture.\n**Encoder**: For the CNN I used [GENet](https://arxiv.org/abs/2006.14090v4) Normal ('gernet_m' in timm library). It is supposed to be a very fast CNN on GPU with a decent ImageNet performance of [80.7% top-1 accuracy](https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv). This is true as far as I can tell. GENet Normal is 4x faster as EfficientNet V2S (version optimized for Tensor Cores, [83.8% top-1](https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv)) on my Titan RTX with 17M vs 19.8M parameters (no classification heads & feature boom/upscale layer). \n\n**Decoder**: I used [Pay-Attention-when-Required](https://arxiv.org/abs/2009.04534) (PAR) Transformer. It uses more  feedforward blocks after every attention block. And it has more of the former at the end for further feature refinement. It uses [Transformer-XL](https://arxiv.org/abs/1901.02860) attention and works as well as it with up to ~50% speed-up. \n'A' = Attention block; 'F' = FeedForward Block\n- Normal Transformer: N * 'AF' \n- PAR-Transformer for enwik8: 5 * 'AFF' + 9 * 'F'\n- Own PAR-Transformer enwik8 baseline: 4 * 'AFFF' + 5 * 'F'\n- Not sure what the best version of encoder-decoder architecture is. It works indeed, contrary to my hypothesis a few days ago. For this kind of data something along this might be a good starting point: 4 * 'AFF' + 'FF'. Two 'F' instead of three seem to work better (for me). \n\n**Didn't work:**\n- Using a 'AFF' PAR Transformer after the CNN for global attention did never work out. Maybe only with shallow CNN. I used adaptive embeddings described below (like in the decoder). I tried a lot of variations to get it to work.\n- CBAM attention in the GENet (applied through timm library by modifiying build config for gernet_m. It is also slows training down too much. And don't even get started with Triplet-Attention. It is even far slower.\n- Better performance with GENet Large then Normal - but back then I still used not-pretrained weights & the 'AFF' encoder block.\n- Using higher channel/block count in GENet for earlier building blocks (2nd basic block). I thought might be useful as the images are all fine lines. The results weren't showing a significant difference\n\n**Unused:**\n- SiLU activation instead of ReLU. It works better (1 percentage point in top-1 token accuracy). But the pretrained model works (of course) even better. Don't know why they didn't use it in the original. Additional 11% speed overhead. But maybe on ImageNet not that useful for this CNN?!\n\n**Verdict:**\n- Best estimated performance possible with everything I used (see also other sections): 1.5 LD.\n(Although I just realized a few hours ago that I forgot to change mean & std for mormalization back to those of ImageNet with pretrained model. Maybe that improves performance.)\n\n GENent normal as used + adaptive positional embeddings (see below): 18.17M Parameters.\nDecoder: varies 43M-47M. Speed: 3:10 hours with 40k validation set only & 224x284 images on Titan RTX with \nmixed-precision training. BUT this is mostly the layer depth. Swapping 'F' with 'A' doesn't slow \ndecoder down [sic!]. (Decoder to deep). Also Transformer-XL  attention is ~1/3 slower than \nvanilla one in my implementation. So use better encoder for better LD.\n\n# Adaptive Positional Embeddings\nFrom [On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention](https://arxiv.org/pdf/1910.04396.pdf), a paper suggested by @hengck23 (Thank you!)\nIt uses the global average pooled vector of the CNN to compute two scale factors (feedforward-net + sigmoid) for the horizontal & vertical positional embeddings. For the later it uses the standard transformer sinusoidal positional embeddings. The positional embeddings are for the CNN output image features when you attend to them via attention. Improves results a little bit - I guess it helps for images with originally different side ratios. Probably far better for areas with really distorted objects like in the paper.\n**Didn't work:** Adaptive positional embeddings for the decoder (text sequence) in all kind of variations.\n\n# COOL Training (Model-Head Ensembling)\nFrom [Fitted Learning: Models with Awareness of their Limits](https://arxiv.org/pdf/1609.02226.pdf)\nIt uses several output-heads (degree of overcompleteness (DOO)) for one class. Not directly an ensemble as the outputs are in 'competition' with each other during training. And during evaluation you multiply the Softmax results of the heads for one class together.\nI was quite suspicious while reading the paper as they only trained it for MNIST & CIFAR10. They don't have much data, and are easy anyway. BUT far more importantly they have only 10 classes!\nBut it was very easy to implement: For n class-head expand & use target-label 1/n. Train as given. For evaluation reshape tensor, multiply class heads together.\nNeither 16, nor 4 heads/DOO did work. This includes using Weight-Dropout (WD) for the output heads like in the [AWD-LSTM](https://www.paperswithcode.com/method/awd-lstm), which made performance a bit better. 2 percentage-points worse for 4 heads & 1 for 4 heads + WD.\n\n# Augmentations\nI used for all augmentations the probability 0f 15%, except the noise, which was 33%. The order in which they are presented is equal to their order in the pipeline. My goal was to mimic the worse images with the augmentations and force generalisation\n- Inserting a white rectangle at a random position between 5-20% of the image height/width\n- Dropping between 5-20% of the pixels in the images - making them white\n- Using torchvisions *RandomPerpective* with distortion_scale=0.3 to transform perspective slightly\n- Adding random noise which was based on random black pixels of 0.01-0.2% of the image which where then rescaled to mimic the look of original noise after scaling (down & up). Regarding the percentages I estimated the noise in the test-set.\n- Normalization of course\n==> First 3 augmentations improved validation score (can't determine anymore how much).\n\n**What didn't work:**\n- Scaling the images down & up to  mimic images of extreme sizes. Although I used *Nearest* as interpolation which made the quality even worse. Might work with better interpolation\n- [JigSaw](https://imgaug.readthedocs.io/en/latest/source/overview/geometric.html?highlight=jigsaw#jigsaw), which splits the image into several equally-sized rectangles and shifts them around.  This is supposed to enforce better global feature learning through its adversarial interruption of local features. Even a very 'friendly' JigSaw split of 4 rectangles didn't work out.\n- Higher augmentation probabilities (values & usage)\n\n# Miscellaneous \n- [Label smoothing](https://www.paperswithcode.com/method/label-smoothing) did not work, neither epsilon = 0.1 nor 0.01 (see also *COOL Training*)\n- Adding the original image size ration to the encoder output did not work\n- I used the PIL library for image resizing. Note that among many other libraries torchvision has a faulty implementation for resizing which lets horizontal/vertical lines vanish - even with bicubic interpolation. But you can use PIL images directly with torchvision and it works fine.\n- dSiLU activation (derivative formula of SiLU as activation) in the last feedforward blocks, as suggested by this [Reinforcement-Learning paper](https://arxiv.org/abs/1702.03118). Might be the fickle RL, different task or non-reproducibility.\n- Couldn't reproduce Transformer-N from [Revisiting Simple Neural Probabilistic Language Models](https://arxiv.org/pdf/2104.03474.pdf). But had to guess exact layer architecture.\n\n# Conclusion\n- GENet is a very fast image classification CNN which gives decent performance, but not enough for really high rankings. Would use more powerful encoder and shallower decoder (speed) next time - or something entirely different.\n- Transformer-Decoder can benefit from more then one feedforward-block between attention blocks\n- Adaptive positional embedding for CNN encoder features seems likely to be reproducible and gives small performance boost for molecule images\n- Augmentations that mimic worse looking images work well. Distorting image perspective a bit also works.\n- Did not work: COOL-Training, dSiLU activation in last layers, JigSaw augmentation, attention within CNN, label smoothing, shallow transformer encoder after deep CNN encoder for global attention\n\n# My View\nI had quite a lot of fun doing all this, although I have not much to show for it (except my code should I publish it). Given that I was knew to computer vision apart from my very first project years ago, I had the opportunity to learn quite some things.\nI also enjoyed most of the discussions.\n\n(Why did I not submit anything: Wanted to save myself the stress behind reaching a decent score within the last days - I also had neither a rotation classifier, already resized test images or the code for test-set inference, only validation)\n\nAt least in the next competition I will go for a high score and see how motivating it is. Just doing endless training to get the highest possible performance didn't motivate me in this competition. I will have to see if the normal way of *kaggling* is something for me.\n\nNote: If you use anything novel for you from my learnings here then please link this post.\n\nThank You :)",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1334993": "Hello There,\n\nI hope everyone is (halfway) satisfied with there final ranking.\n\nThe Conclusion section below functions as TLDR. \nAll sections are mostly independent, so read what seems interesting.\n\nThis is my summary of the interesting parts of my paper reproductions and experiments. I focused on those, instead of a high score, so I decided to not submit anything but experiment until the end - I will come back to that.\n\nI post this text to hopefully give some of you useful ideas. Every section adds something new I haven't read in the discussions (Didn't really check notebooks out).\nI have implemented everything myself as long as there was no library for that purpose ([timm](https://github.com/rwightman/pytorch-image-models), [hydra](https://hydra.cc/docs/intro/), [catalyst](https://github.com/catalyst-team/catalyst), ...); except for the [path template](https://www.kaggle.com/c/bms-molecular-translation/discussion/223423) by @ihelon.\n\n# Tokenization & Cross-Validation Split\nI used this regex:\n`r'([A-Z][a-z]?|(?=\\d{3})\\d|\\d{1,2}|/[a-z]|\\W)'; token_list = re.findall(..)`\n\nNote, that it splits e.g. 107 into 1 & 07 or 167 into 1 & 67. That is for generalisation to all numbers between 100-199 as there weren't InchIs for all of those in the training set. It also rises the number of occurrences of the numerical tokens.\nYou could of course use a extra token for the 10² and 10¹ positions to not get e.g. 07. So 107 -> ħ ð 7. But that seemed like over-engineering. \n\n**5-Fold CV-split:**\nI used the InchI token lengths split into buckets and the number of the atoms and InChI sections in the InchI string for the CV split. I treated it as multi-label classification for the CV splitting. As I didn't find a library which worked with this amount of 'labels' and data-rows I used a simplified own splitter. I sorted the feature label rows lexically (like in Excel when you sort after column A, then B...). Then I assigned the first row to fold 1, the second to fold 2, etc. the sixth to fold 1 again. That assured that similar features co-occurred at least somewhat through the folds.\n\n# General Architecture\nI used a CNN-Encoder and Transformer-Decoder architecture.\n**Encoder**: For the CNN I used [GENet](https://arxiv.org/abs/2006.14090v4) Normal ('gernet_m' in timm library). It is supposed to be a very fast CNN on GPU with a decent ImageNet performance of [80.7% top-1 accuracy](https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv). This is true as far as I can tell. GENet Normal is 4x faster as EfficientNet V2S (version optimized for Tensor Cores, [83.8% top-1](https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv)) on my Titan RTX with 17M vs 19.8M parameters (no classification heads & feature boom/upscale layer). \n\n**Decoder**: I used [Pay-Attention-when-Required](https://arxiv.org/abs/2009.04534) (PAR) Transformer. It uses more  feedforward blocks after every attention block. And it has more of the former at the end for further feature refinement. It uses [Transformer-XL](https://arxiv.org/abs/1901.02860) attention and works as well as it with up to ~50% speed-up. \n'A' = Attention block; 'F' = FeedForward Block\n- Normal Transformer: N * 'AF' \n- PAR-Transformer for enwik8: 5 * 'AFF' + 9 * 'F'\n- Own PAR-Transformer enwik8 baseline: 4 * 'AFFF' + 5 * 'F'\n- Not sure what the best version of encoder-decoder architecture is. It works indeed, contrary to my hypothesis a few days ago. For this kind of data something along this might be a good starting point: 4 * 'AFF' + 'FF'. Two 'F' instead of three seem to work better (for me). \n\n**Didn't work:**\n- Using a 'AFF' PAR Transformer after the CNN for global attention did never work out. Maybe only with shallow CNN. I used adaptive embeddings described below (like in the decoder). I tried a lot of variations to get it to work.\n- CBAM attention in the GENet (applied through timm library by modifiying build config for gernet_m. It is also slows training down too much. And don't even get started with Triplet-Attention. It is even far slower.\n- Better performance with GENet Large then Normal - but back then I still used not-pretrained weights & the 'AFF' encoder block.\n- Using higher channel/block count in GENet for earlier building blocks (2nd basic block). I thought might be useful as the images are all fine lines. The results weren't showing a significant difference\n\n**Unused:**\n- SiLU activation instead of ReLU. It works better (1 percentage point in top-1 token accuracy). But the pretrained model works (of course) even better. Don't know why they didn't use it in the original. Additional 11% speed overhead. But maybe on ImageNet not that useful for this CNN?!\n\n**Verdict:**\n- Best estimated performance possible with everything I used (see also other sections): 1.5 LD.\n(Although I just realized a few hours ago that I forgot to change mean & std for mormalization back to those of ImageNet with pretrained model. Maybe that improves performance.)\n\n GENent normal as used + adaptive positional embeddings (see below): 18.17M Parameters.\nDecoder: varies 43M-47M. Speed: 3:10 hours with 40k validation set only & 224x284 images on Titan RTX with \nmixed-precision training. BUT this is mostly the layer depth. Swapping 'F' with 'A' doesn't slow \ndecoder down [sic!]. (Decoder to deep). Also Transformer-XL  attention is ~1/3 slower than \nvanilla one in my implementation. So use better encoder for better LD.\n\n# Adaptive Positional Embeddings\nFrom [On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention](https://arxiv.org/pdf/1910.04396.pdf), a paper suggested by @hengck23 (Thank you!)\nIt uses the global average pooled vector of the CNN to compute two scale factors (feedforward-net + sigmoid) for the horizontal & vertical positional embeddings. For the later it uses the standard transformer sinusoidal positional embeddings. The positional embeddings are for the CNN output image features when you attend to them via attention. Improves results a little bit - I guess it helps for images with originally different side ratios. Probably far better for areas with really distorted objects like in the paper.\n**Didn't work:** Adaptive positional embeddings for the decoder (text sequence) in all kind of variations.\n\n# COOL Training (Model-Head Ensembling)\nFrom [Fitted Learning: Models with Awareness of their Limits](https://arxiv.org/pdf/1609.02226.pdf)\nIt uses several output-heads (degree of overcompleteness (DOO)) for one class. Not directly an ensemble as the outputs are in 'competition' with each other during training. And during evaluation you multiply the Softmax results of the heads for one class together.\nI was quite suspicious while reading the paper as they only trained it for MNIST & CIFAR10. They don't have much data, and are easy anyway. BUT far more importantly they have only 10 classes!\nBut it was very easy to implement: For n class-head expand & use target-label 1/n. Train as given. For evaluation reshape tensor, multiply class heads together.\nNeither 16, nor 4 heads/DOO did work. This includes using Weight-Dropout (WD) for the output heads like in the [AWD-LSTM](https://www.paperswithcode.com/method/awd-lstm), which made performance a bit better. 2 percentage-points worse for 4 heads & 1 for 4 heads + WD.\n\n# Augmentations\nI used for all augmentations the probability 0f 15%, except the noise, which was 33%. The order in which they are presented is equal to their order in the pipeline. My goal was to mimic the worse images with the augmentations and force generalisation\n- Inserting a white rectangle at a random position between 5-20% of the image height/width\n- Dropping between 5-20% of the pixels in the images - making them white\n- Using torchvisions *RandomPerpective* with distortion_scale=0.3 to transform perspective slightly\n- Adding random noise which was based on random black pixels of 0.01-0.2% of the image which where then rescaled to mimic the look of original noise after scaling (down & up). Regarding the percentages I estimated the noise in the test-set.\n- Normalization of course\n==> First 3 augmentations improved validation score (can't determine anymore how much).\n\n**What didn't work:**\n- Scaling the images down & up to  mimic images of extreme sizes. Although I used *Nearest* as interpolation which made the quality even worse. Might work with better interpolation\n- [JigSaw](https://imgaug.readthedocs.io/en/latest/source/overview/geometric.html?highlight=jigsaw#jigsaw), which splits the image into several equally-sized rectangles and shifts them around.  This is supposed to enforce better global feature learning through its adversarial interruption of local features. Even a very 'friendly' JigSaw split of 4 rectangles didn't work out.\n- Higher augmentation probabilities (values & usage)\n\n# Miscellaneous \n- [Label smoothing](https://www.paperswithcode.com/method/label-smoothing) did not work, neither epsilon = 0.1 nor 0.01 (see also *COOL Training*)\n- Adding the original image size ration to the encoder output did not work\n- I used the PIL library for image resizing. Note that among many other libraries torchvision has a faulty implementation for resizing which lets horizontal/vertical lines vanish - even with bicubic interpolation. But you can use PIL images directly with torchvision and it works fine.\n- dSiLU activation (derivative formula of SiLU as activation) in the last feedforward blocks, as suggested by this [Reinforcement-Learning paper](https://arxiv.org/abs/1702.03118). Might be the fickle RL, different task or non-reproducibility.\n- Couldn't reproduce Transformer-N from [Revisiting Simple Neural Probabilistic Language Models](https://arxiv.org/pdf/2104.03474.pdf). But had to guess exact layer architecture.\n\n# Conclusion\n- GENet is a very fast image classification CNN which gives decent performance, but not enough for really high rankings. Would use more powerful encoder and shallower decoder (speed) next time - or something entirely different.\n- Transformer-Decoder can benefit from more then one feedforward-block between attention blocks\n- Adaptive positional embedding for CNN encoder features seems likely to be reproducible and gives small performance boost for molecule images\n- Augmentations that mimic worse looking images work well. Distorting image perspective a bit also works.\n- Did not work: COOL-Training, dSiLU activation in last layers, JigSaw augmentation, attention within CNN, label smoothing, shallow transformer encoder after deep CNN encoder for global attention\n\n# My View\nI had quite a lot of fun doing all this, although I have not much to show for it (except my code should I publish it). Given that I was knew to computer vision apart from my very first project years ago, I had the opportunity to learn quite some things.\nI also enjoyed most of the discussions.\n\n(Why did I not submit anything: Wanted to save myself the stress behind reaching a decent score within the last days - I also had neither a rotation classifier, already resized test images or the code for test-set inference, only validation)\n\nAt least in the next competition I will go for a high score and see how motivating it is. Just doing endless training to get the highest possible performance didn't motivate me in this competition. I will have to see if the normal way of *kaggling* is something for me.\n\nNote: If you use anything novel for you from my learnings here then please link this post.\n\nThank You :)"
  }
}