{
  "id": 460383,
  "title": "[18th place solution] Transformer + bpp conv + relative position bias",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460383",
  "author_name": "Bohan Yoon",
  "post_date": "2023-12-09T01:44:21.137000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the hosts for organizing this competition. This competition was a good experience for me.</p>\n<p><strong>Features</strong></p>\n<p>I only used RNA sequence and bpp matrix provided by host. I tried other features, but it didn't have much effect.</p>\n<p><strong>Model</strong></p>\n<blockquote>\n  <p>x1 = rna_sequence<br>\n          x2 = bpp_matrix<br>\n          x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=True)(x1)<br>\n          x2 = tf.expand_dims(x2, axis=1)<br>\n          x2 = tf.keras.layers.Conv2D(16, 16, activation='gelu', padding='same', data_format='channels_last')(x2)<br>\n          pos = RelativePositionBias(16, 32, 128)(x, x)<br>\n          for i in range(12):<br>\n              x = rel_transformer_block(hidden_dim, hidden_dim*4, head=16, drop_rate=0.2, dtype=dtype)([x, x2, pos])<br>\n          x = tf.keras.layers.Dense(2)(x)</p>\n</blockquote>\n<p><strong>Training</strong></p>\n<ul>\n<li>Epoch : 80</li>\n<li>Batch size : 128</li>\n<li>learning rate : 1e-3, with Cosine Decay</li>\n<li>Optimizer : Ranger</li>\n<li>loss : mae loss</li>\n<li>only single model</li>\n</ul>\n<p><strong>Mystery</strong></p>\n<p>When I added the conv2d layer to my model, the learning speed slowed down noticeably. Does anyone know why?</p>",
  "messages": [
    {
      "id": 2554269,
      "postDate": "2023-12-09T01:44:21.137Z",
      "content": "<p>Thanks to Kaggle and the hosts for organizing this competition. This competition was a good experience for me.</p>\n<p><strong>Features</strong></p>\n<p>I only used RNA sequence and bpp matrix provided by host. I tried other features, but it didn't have much effect.</p>\n<p><strong>Model</strong></p>\n<blockquote>\n  <p>x1 = rna_sequence<br>\n          x2 = bpp_matrix<br>\n          x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=True)(x1)<br>\n          x2 = tf.expand_dims(x2, axis=1)<br>\n          x2 = tf.keras.layers.Conv2D(16, 16, activation='gelu', padding='same', data_format='channels_last')(x2)<br>\n          pos = RelativePositionBias(16, 32, 128)(x, x)<br>\n          for i in range(12):<br>\n              x = rel_transformer_block(hidden_dim, hidden_dim*4, head=16, drop_rate=0.2, dtype=dtype)([x, x2, pos])<br>\n          x = tf.keras.layers.Dense(2)(x)</p>\n</blockquote>\n<p><strong>Training</strong></p>\n<ul>\n<li>Epoch : 80</li>\n<li>Batch size : 128</li>\n<li>learning rate : 1e-3, with Cosine Decay</li>\n<li>Optimizer : Ranger</li>\n<li>loss : mae loss</li>\n<li>only single model</li>\n</ul>\n<p><strong>Mystery</strong></p>\n<p>When I added the conv2d layer to my model, the learning speed slowed down noticeably. Does anyone know why?</p>",
      "rawMarkdown": "Thanks to Kaggle and the hosts for organizing this competition. This competition was a good experience for me.\n\n**Features**\n\nI only used RNA sequence and bpp matrix provided by host. I tried other features, but it didn't have much effect.\n\n**Model**\n\n> \n        x1 = rna_sequence\n        x2 = bpp_matrix\n        x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=True)(x1)\n        x2 = tf.expand_dims(x2, axis=1)\n        x2 = tf.keras.layers.Conv2D(16, 16, activation='gelu', padding='same', data_format='channels_last')(x2)\n        pos = RelativePositionBias(16, 32, 128)(x, x)\n        for i in range(12):\n            x = rel_transformer_block(hidden_dim, hidden_dim*4, head=16, drop_rate=0.2, dtype=dtype)([x, x2, pos])\n        x = tf.keras.layers.Dense(2)(x)\n\n**Training**\n\n- Epoch : 80\n- Batch size : 128\n- learning rate : 1e-3, with Cosine Decay\n- Optimizer : Ranger\n- loss : mae loss\n- only single model\n\n**Mystery**\n\nWhen I added the conv2d layer to my model, the learning speed slowed down noticeably. Does anyone know why?",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2554269": "Thanks to Kaggle and the hosts for organizing this competition. This competition was a good experience for me.\n\n**Features**\n\nI only used RNA sequence and bpp matrix provided by host. I tried other features, but it didn't have much effect.\n\n**Model**\n\n> \n        x1 = rna_sequence\n        x2 = bpp_matrix\n        x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=True)(x1)\n        x2 = tf.expand_dims(x2, axis=1)\n        x2 = tf.keras.layers.Conv2D(16, 16, activation='gelu', padding='same', data_format='channels_last')(x2)\n        pos = RelativePositionBias(16, 32, 128)(x, x)\n        for i in range(12):\n            x = rel_transformer_block(hidden_dim, hidden_dim*4, head=16, drop_rate=0.2, dtype=dtype)([x, x2, pos])\n        x = tf.keras.layers.Dense(2)(x)\n\n**Training**\n\n- Epoch : 80\n- Batch size : 128\n- learning rate : 1e-3, with Cosine Decay\n- Optimizer : Ranger\n- loss : mae loss\n- only single model\n\n**Mystery**\n\nWhen I added the conv2d layer to my model, the learning speed slowed down noticeably. Does anyone know why?"
  }
}