{
  "id": 454575,
  "title": "Gibbs modified attention",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/454575",
  "author_name": "Louis Sanna",
  "post_date": "2023-11-10T20:07:42.058000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello all, </p>\n<p>Under the advice of the chatGPT, I modified the attention of a base transformer to focus on stable pair as measured by their gibbs energy. It did not improve the model, so I'm sharing it should some of you have some feedback on it.</p>\n<p>(I have no idea wether is make sense from chemical/biological point of view).</p>\n<p><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2779691/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2779691/</a><br>\n<img src=\"https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&amp;p=PMC3&amp;id=2779691_2278tbl1.jpg\" alt=\"https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&amp;p=PMC3&amp;id=2779691_2278tbl1.jpg\"></p>\n<p>Here is the code, I took <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> great notebook as starter.</p>\n<p>To compute energy:</p>\n<pre><code> tensorflow  tf\n\n\ngibbs_energy_tensor = tf.constant([\n    [, , , , ],    \n    [, , , , ],    \n    [, , , , ],   \n    [, , , , ],   \n    [, , , , ], \n], dtype=tf.float32)\n\nmax_gibbs_energy = tf.reduce_max(gibbs_energy_tensor)\n\n ():\n     energies / (max_gibbs_energy * )\n\n ():\n    \n    sequence = tf.pad(sequence, [[, ], [, ]], constant_values=)\n\n    \n    sequence_int = tf.cast(sequence, tf.int32)\n\n    \n    left_pairs_int = sequence_int[:, :-]   \n    middle_pairs_int = sequence_int[:, :-]  \n    right_pairs_int = sequence_int[:, :]   \n\n    \n    left_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([left_pairs_int, middle_pairs_int], axis=))\n    right_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([middle_pairs_int, right_pairs_int], axis=))\n\n    \n    summed_energies = left_energy_pairs + right_energy_pairs\n\n     summed_energies\n\n ():\n    raw_energies = map_sequence_to_pairwise_gibbs_energy(batched_sequence)\n\n    normalized_energies = max_normalize_tensor(raw_energies)\n\n     normalized_energies\n\n\n\n\nbatched_sequence = tf.constant([[, , ], [, , ]], dtype=tf.float32)\nnormalized_energies = map_batch_to_pairwise_gibbs_energy_normalized(batched_sequence)\n(, normalized_energies)\n</code></pre>\n<p>Modified transformer</p>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n        ().__init__()\n        self.att = tf.keras.layers.MultiHeadAttention(num_heads=num_heads, key_dim=dim//num_heads)\n        self.ffn = tf.keras.Sequential(\n            [\n                tf.keras.layers.Dense(feed_forward_dim, activation=),\n                tf.keras.layers.Dense(dim),\n            ]\n        )\n        self.layernorm1 = tf.keras.layers.LayerNormalization(epsilon=)\n        self.layernorm2 = tf.keras.layers.LayerNormalization(epsilon=)\n        self.dropout1 = tf.keras.layers.Dropout(rate)\n        self.dropout2 = tf.keras.layers.Dropout(rate)\n        self.alpha = self.add_weight(name=,\n                                     shape=(),\n                                     trainable=alpha_trainable,\n                                     initializer=tf.keras.initializers.Constant(value=init_alpha),\n                                     constraint=tf.keras.constraints.MinMaxNorm(min_value=, max_value=, axis=))\n        self.supports_masking = \n        self.num_heads = num_heads\n\n     ():\n\n        att_mask = tf.expand_dims(mask, axis=-)\n        att_mask = tf.repeat(att_mask, repeats=tf.shape(att_mask)[], axis=-)\n\n        copy_gibbs_scores = tf.identity(gibbs_scores)\n        gibbs_tensor = tf.expand_dims(copy_gibbs_scores, -)\n\n        \n        \n        \n        \n        \n        \n\n        gibbs_tensor = tf.broadcast_to(copy_gibbs_scores, tf.shape(attn_output))\n\n        left = tf.multiply(tf.multiply(self.alpha, gibbs_tensor), attn_output)\n        one_minus_alpha = tf.subtract(, self.alpha)\n        right = tf.multiply(one_minus_alpha, attn_output)\n        attn_output = tf.add(left, right)\n\n        attn_output = self.dropout1(attn_output, training=training)\n        out1 = self.layernorm1(inputs + attn_output)\n        ffn_output = self.ffn(out1)\n        ffn_output = self.dropout2(ffn_output, training=training)\n         self.layernorm2(out1 + ffn_output)\n</code></pre>\n<p>How to use:</p>\n<pre><code> ():\n     strategy.scope():\n        inp = tf.keras.Input([max_len])\n\n        gibbs_inp = tf.identity(inp)\n        gibbs_scores_layer = PairwiseGibbsEnergyLayer(trainable=)\n        gibbs_scores = gibbs_scores_layer(gibbs_inp)\n        gibbs_scores = tf.keras.layers.Reshape((max_len, ))(gibbs_scores)\n        \n\n        x = inp\n\n        x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=)(x)\n        x = positional_encoding_layer(num_vocab=num_vocab, maxlen=, hidden_dim=hidden_dim)(x)\n\n         _  ():\n          x = thermo_transformer_block(hidden_dim, , hidden_dim*)(x, gibbs_scores=gibbs_scores)\n\n        x = tf.keras.layers.Dropout()(x)\n        x = tf.keras.layers.Dense()(x)\n\n        model = tf.keras.Model(inp, x)\n        loss = loss_fn\n        optimizer = tf.keras.optimizers.AdamW(learning_rate=)\n        model.(loss=loss, optimizer=optimizer, steps_per_execution = )\n         model\n</code></pre>\n<p>Callback to check the value of alpha, it lets the model free to use the original attention or not. The model picks 1 for the first layer, meaning gibbs seems useful, but the validation score does not improve in later epochs.</p>\n<pre><code> (tf.keras.callbacks.Callback):\n     ():\n          ().__init__()\n     ():\n        alphas = []\n         layer  self.model.layers:\n             (layer, thermo_transformer_block):\n                alpha_value = layer.alpha.numpy()\n                alphas.append(alpha_value)\n                ()\n</code></pre>\n<p>By the way I would happy to join a team, I'm a beginner in machine learning but quite experienced in software engineering.</p>",
  "messages": [
    {
      "id": 2520402,
      "postDate": "2023-11-10T20:07:42.057Z",
      "content": "<p>Hello all, </p>\n<p>Under the advice of the chatGPT, I modified the attention of a base transformer to focus on stable pair as measured by their gibbs energy. It did not improve the model, so I'm sharing it should some of you have some feedback on it.</p>\n<p>(I have no idea wether is make sense from chemical/biological point of view).</p>\n<p><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2779691/\" target=\"_blank\">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2779691/</a><br>\n<img src=\"https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&amp;p=PMC3&amp;id=2779691_2278tbl1.jpg\" alt=\"https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&amp;p=PMC3&amp;id=2779691_2278tbl1.jpg\"></p>\n<p>Here is the code, I took <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> great notebook as starter.</p>\n<p>To compute energy:</p>\n<pre><code> tensorflow  tf\n\n\ngibbs_energy_tensor = tf.constant([\n    [, , , , ],    \n    [, , , , ],    \n    [, , , , ],   \n    [, , , , ],   \n    [, , , , ], \n], dtype=tf.float32)\n\nmax_gibbs_energy = tf.reduce_max(gibbs_energy_tensor)\n\n ():\n     energies / (max_gibbs_energy * )\n\n ():\n    \n    sequence = tf.pad(sequence, [[, ], [, ]], constant_values=)\n\n    \n    sequence_int = tf.cast(sequence, tf.int32)\n\n    \n    left_pairs_int = sequence_int[:, :-]   \n    middle_pairs_int = sequence_int[:, :-]  \n    right_pairs_int = sequence_int[:, :]   \n\n    \n    left_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([left_pairs_int, middle_pairs_int], axis=))\n    right_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([middle_pairs_int, right_pairs_int], axis=))\n\n    \n    summed_energies = left_energy_pairs + right_energy_pairs\n\n     summed_energies\n\n ():\n    raw_energies = map_sequence_to_pairwise_gibbs_energy(batched_sequence)\n\n    normalized_energies = max_normalize_tensor(raw_energies)\n\n     normalized_energies\n\n\n\n\nbatched_sequence = tf.constant([[, , ], [, , ]], dtype=tf.float32)\nnormalized_energies = map_batch_to_pairwise_gibbs_energy_normalized(batched_sequence)\n(, normalized_energies)\n</code></pre>\n<p>Modified transformer</p>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n        ().__init__()\n        self.att = tf.keras.layers.MultiHeadAttention(num_heads=num_heads, key_dim=dim//num_heads)\n        self.ffn = tf.keras.Sequential(\n            [\n                tf.keras.layers.Dense(feed_forward_dim, activation=),\n                tf.keras.layers.Dense(dim),\n            ]\n        )\n        self.layernorm1 = tf.keras.layers.LayerNormalization(epsilon=)\n        self.layernorm2 = tf.keras.layers.LayerNormalization(epsilon=)\n        self.dropout1 = tf.keras.layers.Dropout(rate)\n        self.dropout2 = tf.keras.layers.Dropout(rate)\n        self.alpha = self.add_weight(name=,\n                                     shape=(),\n                                     trainable=alpha_trainable,\n                                     initializer=tf.keras.initializers.Constant(value=init_alpha),\n                                     constraint=tf.keras.constraints.MinMaxNorm(min_value=, max_value=, axis=))\n        self.supports_masking = \n        self.num_heads = num_heads\n\n     ():\n\n        att_mask = tf.expand_dims(mask, axis=-)\n        att_mask = tf.repeat(att_mask, repeats=tf.shape(att_mask)[], axis=-)\n\n        copy_gibbs_scores = tf.identity(gibbs_scores)\n        gibbs_tensor = tf.expand_dims(copy_gibbs_scores, -)\n\n        \n        \n        \n        \n        \n        \n\n        gibbs_tensor = tf.broadcast_to(copy_gibbs_scores, tf.shape(attn_output))\n\n        left = tf.multiply(tf.multiply(self.alpha, gibbs_tensor), attn_output)\n        one_minus_alpha = tf.subtract(, self.alpha)\n        right = tf.multiply(one_minus_alpha, attn_output)\n        attn_output = tf.add(left, right)\n\n        attn_output = self.dropout1(attn_output, training=training)\n        out1 = self.layernorm1(inputs + attn_output)\n        ffn_output = self.ffn(out1)\n        ffn_output = self.dropout2(ffn_output, training=training)\n         self.layernorm2(out1 + ffn_output)\n</code></pre>\n<p>How to use:</p>\n<pre><code> ():\n     strategy.scope():\n        inp = tf.keras.Input([max_len])\n\n        gibbs_inp = tf.identity(inp)\n        gibbs_scores_layer = PairwiseGibbsEnergyLayer(trainable=)\n        gibbs_scores = gibbs_scores_layer(gibbs_inp)\n        gibbs_scores = tf.keras.layers.Reshape((max_len, ))(gibbs_scores)\n        \n\n        x = inp\n\n        x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=)(x)\n        x = positional_encoding_layer(num_vocab=num_vocab, maxlen=, hidden_dim=hidden_dim)(x)\n\n         _  ():\n          x = thermo_transformer_block(hidden_dim, , hidden_dim*)(x, gibbs_scores=gibbs_scores)\n\n        x = tf.keras.layers.Dropout()(x)\n        x = tf.keras.layers.Dense()(x)\n\n        model = tf.keras.Model(inp, x)\n        loss = loss_fn\n        optimizer = tf.keras.optimizers.AdamW(learning_rate=)\n        model.(loss=loss, optimizer=optimizer, steps_per_execution = )\n         model\n</code></pre>\n<p>Callback to check the value of alpha, it lets the model free to use the original attention or not. The model picks 1 for the first layer, meaning gibbs seems useful, but the validation score does not improve in later epochs.</p>\n<pre><code> (tf.keras.callbacks.Callback):\n     ():\n          ().__init__()\n     ():\n        alphas = []\n         layer  self.model.layers:\n             (layer, thermo_transformer_block):\n                alpha_value = layer.alpha.numpy()\n                alphas.append(alpha_value)\n                ()\n</code></pre>\n<p>By the way I would happy to join a team, I'm a beginner in machine learning but quite experienced in software engineering.</p>",
      "rawMarkdown": "Hello all, \n\nUnder the advice of the chatGPT, I modified the attention of a base transformer to focus on stable pair as measured by their gibbs energy. It did not improve the model, so I'm sharing it should some of you have some feedback on it.\n\n(I have no idea wether is make sense from chemical/biological point of view).\n\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC2779691/\n![https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&p=PMC3&id=2779691_2278tbl1.jpg](https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&p=PMC3&id=2779691_2278tbl1.jpg)\n\nHere is the code, I took @shlomoron great notebook as starter.\n\nTo compute energy:\n```python\nimport tensorflow as tf\n\n# Gibbs energy tensor\ngibbs_energy_tensor = tf.constant([\n    [0.0, 0.0, 0.0, 0.0, 0.0],    # Padding row\n    [0.0, 0.0, 0.0, 0.0, 4.42],    # Row for 'A' (A, C, G, U)\n    [0.0, 0.0, 0.0, 5.53, 0.37],   # Row for 'C' (A, C, G, U)\n    [0.0, 0.0, 5.53, 0.0, 4.45],   # Row for 'G' (A, C, G, U)\n    [0.0, 4.42, 0.37, 4.45, 5.82], # Row for 'U' (A, C, G, U)\n], dtype=tf.float32)\n\nmax_gibbs_energy = tf.reduce_max(gibbs_energy_tensor)\n\ndef max_normalize_tensor(energies):\n    return energies / (max_gibbs_energy * 2)\n\ndef map_sequence_to_pairwise_gibbs_energy(sequence):\n    # Ensure the sequence is padded with 0s to avoid out of bounds issues\n    sequence = tf.pad(sequence, [[0, 0], [1, 1]], constant_values=0)\n\n    # Convert the sequence to integer type for indexing\n    sequence_int = tf.cast(sequence, tf.int32)\n\n    # Create pairs of nucleotides as indices for each sequence in the batch\n    left_pairs_int = sequence_int[:, :-2]   # Left elements of the pairs\n    middle_pairs_int = sequence_int[:, 1:-1]  # Middle elements, will be summed\n    right_pairs_int = sequence_int[:, 2:]   # Right elements of the pairs\n\n    # Look up the Gibbs energy for each pair\n    left_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([left_pairs_int, middle_pairs_int], axis=2))\n    right_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([middle_pairs_int, right_pairs_int], axis=2))\n\n    # Sum the energies for the middle nucleotides\n    summed_energies = left_energy_pairs + right_energy_pairs\n\n    return summed_energies\n\ndef map_batch_to_pairwise_gibbs_energy_normalized(batched_sequence):\n    raw_energies = map_sequence_to_pairwise_gibbs_energy(batched_sequence)\n\n    normalized_energies = max_normalize_tensor(raw_energies)\n\n    return normalized_energies\n\n# Example usage:\n# batched_sequence must be a 2D tensor with shape [batch_size, sequence_length]\n# For example:\nbatched_sequence = tf.constant([[4, 4, 4], [4, 1, 2]], dtype=tf.float32)\nnormalized_energies = map_batch_to_pairwise_gibbs_energy_normalized(batched_sequence)\nprint(\"normalized_energies\", normalized_energies)\n```\n\nModified transformer\n\n```python\nclass thermo_transformer_block(tf.keras.layers.Layer):\n    def __init__(self, dim, num_heads, feed_forward_dim, rate=0.1, init_alpha=0.5, alpha_trainable=True):\n        super().__init__()\n        self.att = tf.keras.layers.MultiHeadAttention(num_heads=num_heads, key_dim=dim//num_heads)\n        self.ffn = tf.keras.Sequential(\n            [\n                tf.keras.layers.Dense(feed_forward_dim, activation=\"relu\"),\n                tf.keras.layers.Dense(dim),\n            ]\n        )\n        self.layernorm1 = tf.keras.layers.LayerNormalization(epsilon=1e-6)\n        self.layernorm2 = tf.keras.layers.LayerNormalization(epsilon=1e-6)\n        self.dropout1 = tf.keras.layers.Dropout(rate)\n        self.dropout2 = tf.keras.layers.Dropout(rate)\n        self.alpha = self.add_weight(name='alpha',\n                                     shape=(),\n                                     trainable=alpha_trainable,\n                                     initializer=tf.keras.initializers.Constant(value=init_alpha),\n                                     constraint=tf.keras.constraints.MinMaxNorm(min_value=0, max_value=1, axis=None))\n        self.supports_masking = True\n        self.num_heads = num_heads\n\n    def call(self, inputs, training, mask, gibbs_scores=None):\n\n        att_mask = tf.expand_dims(mask, axis=-1)\n        att_mask = tf.repeat(att_mask, repeats=tf.shape(att_mask)[1], axis=-1)\n\n        copy_gibbs_scores = tf.identity(gibbs_scores)\n        gibbs_tensor = tf.expand_dims(copy_gibbs_scores, -1)\n\n        #Other implem where we modify attention mask\n        #gibbs_binary_mask = tf.cast(tf.math.not_equal(copy_gibbs_scores, 0), dtype=tf.float32)\n        #gibbs_binary_mask = tf.repeat(gibbs_binary_mask, repeats=tf.shape(att_mask)[1], axis=-1)\n        #att_mask = tf.cast(att_mask, tf.float32)\n        #combined_mask = tf.math.multiply(att_mask, gibbs_binary_mask)\n        #attn_output = self.att(inputs, inputs, attention_mask=combined_mask)\n\n        gibbs_tensor = tf.broadcast_to(copy_gibbs_scores, tf.shape(attn_output))\n\n        left = tf.multiply(tf.multiply(self.alpha, gibbs_tensor), attn_output)\n        one_minus_alpha = tf.subtract(1.0, self.alpha)\n        right = tf.multiply(one_minus_alpha, attn_output)\n        attn_output = tf.add(left, right)\n\n        attn_output = self.dropout1(attn_output, training=training)\n        out1 = self.layernorm1(inputs + attn_output)\n        ffn_output = self.ffn(out1)\n        ffn_output = self.dropout2(ffn_output, training=training)\n        return self.layernorm2(out1 + ffn_output)\n```\n\nHow to use:\n\n```python\ndef get_model(hidden_dim = 384, max_len = 206):\n    with strategy.scope():\n        inp = tf.keras.Input([max_len])\n\n        gibbs_inp = tf.identity(inp)\n        gibbs_scores_layer = PairwiseGibbsEnergyLayer(trainable=False)\n        gibbs_scores = gibbs_scores_layer(gibbs_inp)\n        gibbs_scores = tf.keras.layers.Reshape((max_len, 1))(gibbs_scores)\n        # print(\"gibbs_scores\", gibbs_scores.shape)\n\n        x = inp\n\n        x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=True)(x)\n        x = positional_encoding_layer(num_vocab=num_vocab, maxlen=500, hidden_dim=hidden_dim)(x)\n\n        for _ in range(12):\n          x = thermo_transformer_block(hidden_dim, 6, hidden_dim*4)(x, gibbs_scores=gibbs_scores)\n\n        x = tf.keras.layers.Dropout(0.5)(x)\n        x = tf.keras.layers.Dense(2)(x)\n\n        model = tf.keras.Model(inp, x)\n        loss = loss_fn\n        optimizer = tf.keras.optimizers.AdamW(learning_rate=0.0005)\n        model.compile(loss=loss, optimizer=optimizer, steps_per_execution = 100)\n        return model\n\n```\n\nCallback to check the value of alpha, it lets the model free to use the original attention or not. The model picks 1 for the first layer, meaning gibbs seems useful, but the validation score does not improve in later epochs.\n\n```python\nclass PrintAlphaCallback(tf.keras.callbacks.Callback):\n    def __init__(self):\n          super().__init__()\n    def on_epoch_end(self, epoch, logs=None):\n        alphas = []\n        for layer in self.model.layers:\n            if isinstance(layer, thermo_transformer_block):\n                alpha_value = layer.alpha.numpy()\n                alphas.append(alpha_value)\n                print(f'Alpha for layer {layer.name}: {alpha_value}')\n```\n\nBy the way I would happy to join a team, I'm a beginner in machine learning but quite experienced in software engineering.",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2520402": "Hello all, \n\nUnder the advice of the chatGPT, I modified the attention of a base transformer to focus on stable pair as measured by their gibbs energy. It did not improve the model, so I'm sharing it should some of you have some feedback on it.\n\n(I have no idea wether is make sense from chemical/biological point of view).\n\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC2779691/\n![https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&p=PMC3&id=2779691_2278tbl1.jpg](https://www.ncbi.nlm.nih.gov/core/lw/2.0/html/tileshop_pmc/tileshop_pmc_inline.html?title=Click%20on%20image%20to%20zoom&p=PMC3&id=2779691_2278tbl1.jpg)\n\nHere is the code, I took @shlomoron great notebook as starter.\n\nTo compute energy:\n```python\nimport tensorflow as tf\n\n# Gibbs energy tensor\ngibbs_energy_tensor = tf.constant([\n    [0.0, 0.0, 0.0, 0.0, 0.0],    # Padding row\n    [0.0, 0.0, 0.0, 0.0, 4.42],    # Row for 'A' (A, C, G, U)\n    [0.0, 0.0, 0.0, 5.53, 0.37],   # Row for 'C' (A, C, G, U)\n    [0.0, 0.0, 5.53, 0.0, 4.45],   # Row for 'G' (A, C, G, U)\n    [0.0, 4.42, 0.37, 4.45, 5.82], # Row for 'U' (A, C, G, U)\n], dtype=tf.float32)\n\nmax_gibbs_energy = tf.reduce_max(gibbs_energy_tensor)\n\ndef max_normalize_tensor(energies):\n    return energies / (max_gibbs_energy * 2)\n\ndef map_sequence_to_pairwise_gibbs_energy(sequence):\n    # Ensure the sequence is padded with 0s to avoid out of bounds issues\n    sequence = tf.pad(sequence, [[0, 0], [1, 1]], constant_values=0)\n\n    # Convert the sequence to integer type for indexing\n    sequence_int = tf.cast(sequence, tf.int32)\n\n    # Create pairs of nucleotides as indices for each sequence in the batch\n    left_pairs_int = sequence_int[:, :-2]   # Left elements of the pairs\n    middle_pairs_int = sequence_int[:, 1:-1]  # Middle elements, will be summed\n    right_pairs_int = sequence_int[:, 2:]   # Right elements of the pairs\n\n    # Look up the Gibbs energy for each pair\n    left_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([left_pairs_int, middle_pairs_int], axis=2))\n    right_energy_pairs = tf.gather_nd(gibbs_energy_tensor, tf.stack([middle_pairs_int, right_pairs_int], axis=2))\n\n    # Sum the energies for the middle nucleotides\n    summed_energies = left_energy_pairs + right_energy_pairs\n\n    return summed_energies\n\ndef map_batch_to_pairwise_gibbs_energy_normalized(batched_sequence):\n    raw_energies = map_sequence_to_pairwise_gibbs_energy(batched_sequence)\n\n    normalized_energies = max_normalize_tensor(raw_energies)\n\n    return normalized_energies\n\n# Example usage:\n# batched_sequence must be a 2D tensor with shape [batch_size, sequence_length]\n# For example:\nbatched_sequence = tf.constant([[4, 4, 4], [4, 1, 2]], dtype=tf.float32)\nnormalized_energies = map_batch_to_pairwise_gibbs_energy_normalized(batched_sequence)\nprint(\"normalized_energies\", normalized_energies)\n```\n\nModified transformer\n\n```python\nclass thermo_transformer_block(tf.keras.layers.Layer):\n    def __init__(self, dim, num_heads, feed_forward_dim, rate=0.1, init_alpha=0.5, alpha_trainable=True):\n        super().__init__()\n        self.att = tf.keras.layers.MultiHeadAttention(num_heads=num_heads, key_dim=dim//num_heads)\n        self.ffn = tf.keras.Sequential(\n            [\n                tf.keras.layers.Dense(feed_forward_dim, activation=\"relu\"),\n                tf.keras.layers.Dense(dim),\n            ]\n        )\n        self.layernorm1 = tf.keras.layers.LayerNormalization(epsilon=1e-6)\n        self.layernorm2 = tf.keras.layers.LayerNormalization(epsilon=1e-6)\n        self.dropout1 = tf.keras.layers.Dropout(rate)\n        self.dropout2 = tf.keras.layers.Dropout(rate)\n        self.alpha = self.add_weight(name='alpha',\n                                     shape=(),\n                                     trainable=alpha_trainable,\n                                     initializer=tf.keras.initializers.Constant(value=init_alpha),\n                                     constraint=tf.keras.constraints.MinMaxNorm(min_value=0, max_value=1, axis=None))\n        self.supports_masking = True\n        self.num_heads = num_heads\n\n    def call(self, inputs, training, mask, gibbs_scores=None):\n\n        att_mask = tf.expand_dims(mask, axis=-1)\n        att_mask = tf.repeat(att_mask, repeats=tf.shape(att_mask)[1], axis=-1)\n\n        copy_gibbs_scores = tf.identity(gibbs_scores)\n        gibbs_tensor = tf.expand_dims(copy_gibbs_scores, -1)\n\n        #Other implem where we modify attention mask\n        #gibbs_binary_mask = tf.cast(tf.math.not_equal(copy_gibbs_scores, 0), dtype=tf.float32)\n        #gibbs_binary_mask = tf.repeat(gibbs_binary_mask, repeats=tf.shape(att_mask)[1], axis=-1)\n        #att_mask = tf.cast(att_mask, tf.float32)\n        #combined_mask = tf.math.multiply(att_mask, gibbs_binary_mask)\n        #attn_output = self.att(inputs, inputs, attention_mask=combined_mask)\n\n        gibbs_tensor = tf.broadcast_to(copy_gibbs_scores, tf.shape(attn_output))\n\n        left = tf.multiply(tf.multiply(self.alpha, gibbs_tensor), attn_output)\n        one_minus_alpha = tf.subtract(1.0, self.alpha)\n        right = tf.multiply(one_minus_alpha, attn_output)\n        attn_output = tf.add(left, right)\n\n        attn_output = self.dropout1(attn_output, training=training)\n        out1 = self.layernorm1(inputs + attn_output)\n        ffn_output = self.ffn(out1)\n        ffn_output = self.dropout2(ffn_output, training=training)\n        return self.layernorm2(out1 + ffn_output)\n```\n\nHow to use:\n\n```python\ndef get_model(hidden_dim = 384, max_len = 206):\n    with strategy.scope():\n        inp = tf.keras.Input([max_len])\n\n        gibbs_inp = tf.identity(inp)\n        gibbs_scores_layer = PairwiseGibbsEnergyLayer(trainable=False)\n        gibbs_scores = gibbs_scores_layer(gibbs_inp)\n        gibbs_scores = tf.keras.layers.Reshape((max_len, 1))(gibbs_scores)\n        # print(\"gibbs_scores\", gibbs_scores.shape)\n\n        x = inp\n\n        x = tf.keras.layers.Embedding(num_vocab, hidden_dim, mask_zero=True)(x)\n        x = positional_encoding_layer(num_vocab=num_vocab, maxlen=500, hidden_dim=hidden_dim)(x)\n\n        for _ in range(12):\n          x = thermo_transformer_block(hidden_dim, 6, hidden_dim*4)(x, gibbs_scores=gibbs_scores)\n\n        x = tf.keras.layers.Dropout(0.5)(x)\n        x = tf.keras.layers.Dense(2)(x)\n\n        model = tf.keras.Model(inp, x)\n        loss = loss_fn\n        optimizer = tf.keras.optimizers.AdamW(learning_rate=0.0005)\n        model.compile(loss=loss, optimizer=optimizer, steps_per_execution = 100)\n        return model\n\n```\n\nCallback to check the value of alpha, it lets the model free to use the original attention or not. The model picks 1 for the first layer, meaning gibbs seems useful, but the validation score does not improve in later epochs.\n\n```python\nclass PrintAlphaCallback(tf.keras.callbacks.Callback):\n    def __init__(self):\n          super().__init__()\n    def on_epoch_end(self, epoch, logs=None):\n        alphas = []\n        for layer in self.model.layers:\n            if isinstance(layer, thermo_transformer_block):\n                alpha_value = layer.alpha.numpy()\n                alphas.append(alpha_value)\n                print(f'Alpha for layer {layer.name}: {alpha_value}')\n```\n\nBy the way I would happy to join a team, I'm a beginner in machine learning but quite experienced in software engineering."
  }
}