{
  "id": 460257,
  "title": "[93th place solution] BPP as embedding layer",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460257",
  "author_name": "Ángel Jacinto Sánchez Ruiz",
  "post_date": "2023-12-08T12:33:04.870000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The code:<br>\n<a href=\"https://www.kaggle.com/code/sacuscreed/bpp-emb-train-0-17894\" target=\"_blank\">https://www.kaggle.com/code/sacuscreed/bpp-emb-train-0-17894</a></p>\n<p>Inspiration:<br>\nRNAdegformer: accurate prediction of mRNA degradation at nucleotide resolution with deep learning<br>\nShujun He, Baizhen Gao, Rushant Sabnis and Qing Sun<br>\nCorresponding author. Qing Sun, Department of Chemical Engineering, Texas A&amp;M University, 100 Spence St., 77843 TX, USA. Tel.: 979-845-3401;<br>\nE-mail: <a href=\"mailto:sunqing@tamu.edu\">sunqing@tamu.edu</a></p>\n<p>I came to this solution impressed by the power of pseudoknots and looking for a way to take profit from bpp files like they did. I had previously tried a more direct approach to the one described in that work using a custom SelfAttentionHead, but I failed trying to update scores properly.<br>\nSo I thought about whether the relationships between nucleotides contained in the bpp files (along with a little help) could replicate the performance of using pseudoknots without using them explicitly. Such pseudoknots are mainly obtained from such relations after all.</p>\n<p>Procedure:<br>\nFor each sequence scores will be computed as the average of five channels:<br>\n-The bpp, each nucleotide should take more into account other nucleotides as more probable his paring is.<br>\n-The Distance matrix and his powders, each nucleotide should take less into account other nucleotides as more distant they are.<br>\n-The diagonal matrix, all the nucleotides should also remember at some grade what they are.<br>\nThe final scores will result from a convolution across these channels and will represent how much to consider each other nucleotides for the computation of the final  nucleotides representations. The weight of each channel will be learned from the model.<br>\nThese scores will multiply the initial representations from nn.embedding to obtain the final representations that will be positional encoded and feed the transformer.<br>\nThis embedding will take into account all the sequence into each nucleotide and so should have in principle a good generalization.<br>\nJust for curiosity, I've subtracted the optimized weights for that convolution on each of the 4 fold models with these results:</p>\n<h1>Conv weights for the score channels</h1>\n<p>import numpy as np<br>\nweights = np.array([[0.5241,-0.0928,-0.0836,-0.0736,1.0798],<br>\n                    [0.5324,-0.0924,-0.0832,-0.0732,1.0797],<br>\n                    [0.5317,-0.0891,-0.0799,-0.0699,1.0901],<br>\n                    [0.5283,-0.0943,-0.0851,-0.0751,1.0932]])<br>\nweights.mean(axis=0)<br>\narray([ 0.529125, -0.09215 , -0.08295 , -0.07295 ,  1.0857  ])<br>\nweights.std(axis=0)<br>\narray([0.00328966, 0.00189803, 0.00189803, 0.00189803, 0.00605021])</p>\n<p>So as expected the weights for bpp, and diagonal channels are positive and the weights for distance channels are negative. Being the importance of each distance channel lower than the previous one and significantly less important than the bpp channel.</p>",
  "messages": [
    {
      "id": 2553671,
      "postDate": "2023-12-08T12:33:04.870Z",
      "content": "<p>The code:<br>\n<a href=\"https://www.kaggle.com/code/sacuscreed/bpp-emb-train-0-17894\" target=\"_blank\">https://www.kaggle.com/code/sacuscreed/bpp-emb-train-0-17894</a></p>\n<p>Inspiration:<br>\nRNAdegformer: accurate prediction of mRNA degradation at nucleotide resolution with deep learning<br>\nShujun He, Baizhen Gao, Rushant Sabnis and Qing Sun<br>\nCorresponding author. Qing Sun, Department of Chemical Engineering, Texas A&amp;M University, 100 Spence St., 77843 TX, USA. Tel.: 979-845-3401;<br>\nE-mail: <a href=\"mailto:sunqing@tamu.edu\">sunqing@tamu.edu</a></p>\n<p>I came to this solution impressed by the power of pseudoknots and looking for a way to take profit from bpp files like they did. I had previously tried a more direct approach to the one described in that work using a custom SelfAttentionHead, but I failed trying to update scores properly.<br>\nSo I thought about whether the relationships between nucleotides contained in the bpp files (along with a little help) could replicate the performance of using pseudoknots without using them explicitly. Such pseudoknots are mainly obtained from such relations after all.</p>\n<p>Procedure:<br>\nFor each sequence scores will be computed as the average of five channels:<br>\n-The bpp, each nucleotide should take more into account other nucleotides as more probable his paring is.<br>\n-The Distance matrix and his powders, each nucleotide should take less into account other nucleotides as more distant they are.<br>\n-The diagonal matrix, all the nucleotides should also remember at some grade what they are.<br>\nThe final scores will result from a convolution across these channels and will represent how much to consider each other nucleotides for the computation of the final  nucleotides representations. The weight of each channel will be learned from the model.<br>\nThese scores will multiply the initial representations from nn.embedding to obtain the final representations that will be positional encoded and feed the transformer.<br>\nThis embedding will take into account all the sequence into each nucleotide and so should have in principle a good generalization.<br>\nJust for curiosity, I've subtracted the optimized weights for that convolution on each of the 4 fold models with these results:</p>\n<h1>Conv weights for the score channels</h1>\n<p>import numpy as np<br>\nweights = np.array([[0.5241,-0.0928,-0.0836,-0.0736,1.0798],<br>\n                    [0.5324,-0.0924,-0.0832,-0.0732,1.0797],<br>\n                    [0.5317,-0.0891,-0.0799,-0.0699,1.0901],<br>\n                    [0.5283,-0.0943,-0.0851,-0.0751,1.0932]])<br>\nweights.mean(axis=0)<br>\narray([ 0.529125, -0.09215 , -0.08295 , -0.07295 ,  1.0857  ])<br>\nweights.std(axis=0)<br>\narray([0.00328966, 0.00189803, 0.00189803, 0.00189803, 0.00605021])</p>\n<p>So as expected the weights for bpp, and diagonal channels are positive and the weights for distance channels are negative. Being the importance of each distance channel lower than the previous one and significantly less important than the bpp channel.</p>",
      "rawMarkdown": "The code:\nhttps://www.kaggle.com/code/sacuscreed/bpp-emb-train-0-17894\n\nInspiration:\nRNAdegformer: accurate prediction of mRNA degradation at nucleotide resolution with deep learning\nShujun He, Baizhen Gao, Rushant Sabnis and Qing Sun\nCorresponding author. Qing Sun, Department of Chemical Engineering, Texas A&M University, 100 Spence St., 77843 TX, USA. Tel.: 979-845-3401;\nE-mail: sunqing@tamu.edu\n\nI came to this solution impressed by the power of pseudoknots and looking for a way to take profit from bpp files like they did. I had previously tried a more direct approach to the one described in that work using a custom SelfAttentionHead, but I failed trying to update scores properly.\nSo I thought about whether the relationships between nucleotides contained in the bpp files (along with a little help) could replicate the performance of using pseudoknots without using them explicitly. Such pseudoknots are mainly obtained from such relations after all.\n\nProcedure:\nFor each sequence scores will be computed as the average of five channels:\n-The bpp, each nucleotide should take more into account other nucleotides as more probable his paring is.\n-The Distance matrix and his powders, each nucleotide should take less into account other nucleotides as more distant they are.\n-The diagonal matrix, all the nucleotides should also remember at some grade what they are.\nThe final scores will result from a convolution across these channels and will represent how much to consider each other nucleotides for the computation of the final  nucleotides representations. The weight of each channel will be learned from the model.\nThese scores will multiply the initial representations from nn.embedding to obtain the final representations that will be positional encoded and feed the transformer.\nThis embedding will take into account all the sequence into each nucleotide and so should have in principle a good generalization.\nJust for curiosity, I've subtracted the optimized weights for that convolution on each of the 4 fold models with these results:\n\n# Conv weights for the score channels\nimport numpy as np\nweights = np.array([[0.5241,-0.0928,-0.0836,-0.0736,1.0798],\n                    [0.5324,-0.0924,-0.0832,-0.0732,1.0797],\n                    [0.5317,-0.0891,-0.0799,-0.0699,1.0901],\n                    [0.5283,-0.0943,-0.0851,-0.0751,1.0932]])\nweights.mean(axis=0)\narray([ 0.529125, -0.09215 , -0.08295 , -0.07295 ,  1.0857  ])\nweights.std(axis=0)\narray([0.00328966, 0.00189803, 0.00189803, 0.00189803, 0.00605021])\n\nSo as expected the weights for bpp, and diagonal channels are positive and the weights for distance channels are negative. Being the importance of each distance channel lower than the previous one and significantly less important than the bpp channel.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2553671": "The code:\nhttps://www.kaggle.com/code/sacuscreed/bpp-emb-train-0-17894\n\nInspiration:\nRNAdegformer: accurate prediction of mRNA degradation at nucleotide resolution with deep learning\nShujun He, Baizhen Gao, Rushant Sabnis and Qing Sun\nCorresponding author. Qing Sun, Department of Chemical Engineering, Texas A&M University, 100 Spence St., 77843 TX, USA. Tel.: 979-845-3401;\nE-mail: sunqing@tamu.edu\n\nI came to this solution impressed by the power of pseudoknots and looking for a way to take profit from bpp files like they did. I had previously tried a more direct approach to the one described in that work using a custom SelfAttentionHead, but I failed trying to update scores properly.\nSo I thought about whether the relationships between nucleotides contained in the bpp files (along with a little help) could replicate the performance of using pseudoknots without using them explicitly. Such pseudoknots are mainly obtained from such relations after all.\n\nProcedure:\nFor each sequence scores will be computed as the average of five channels:\n-The bpp, each nucleotide should take more into account other nucleotides as more probable his paring is.\n-The Distance matrix and his powders, each nucleotide should take less into account other nucleotides as more distant they are.\n-The diagonal matrix, all the nucleotides should also remember at some grade what they are.\nThe final scores will result from a convolution across these channels and will represent how much to consider each other nucleotides for the computation of the final  nucleotides representations. The weight of each channel will be learned from the model.\nThese scores will multiply the initial representations from nn.embedding to obtain the final representations that will be positional encoded and feed the transformer.\nThis embedding will take into account all the sequence into each nucleotide and so should have in principle a good generalization.\nJust for curiosity, I've subtracted the optimized weights for that convolution on each of the 4 fold models with these results:\n\n# Conv weights for the score channels\nimport numpy as np\nweights = np.array([[0.5241,-0.0928,-0.0836,-0.0736,1.0798],\n                    [0.5324,-0.0924,-0.0832,-0.0732,1.0797],\n                    [0.5317,-0.0891,-0.0799,-0.0699,1.0901],\n                    [0.5283,-0.0943,-0.0851,-0.0751,1.0932]])\nweights.mean(axis=0)\narray([ 0.529125, -0.09215 , -0.08295 , -0.07295 ,  1.0857  ])\nweights.std(axis=0)\narray([0.00328966, 0.00189803, 0.00189803, 0.00189803, 0.00605021])\n\nSo as expected the weights for bpp, and diagonal channels are positive and the weights for distance channels are negative. Being the importance of each distance channel lower than the previous one and significantly less important than the bpp channel."
  }
}