{
  "id": 567587,
  "title": "Insights of a training pipeline for RNA 3D Folding prediction",
  "url": "/competitions/stanford-rna-3d-folding/discussion/567587",
  "author_name": "",
  "post_date": "2025-03-11T02:30:04.231883800Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am using this <a href=\"https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook\" target=\"_blank\">notebook</a> as reference</p>\n<p>Today, I was exploring his training pipeline, which has two components, the first one is to create a new evaluation metric based on the Huber Loss function, this function is particularly important because:</p>\n<ol>\n<li><p>Handles the padding we set at the beginning to make sure the sequence has the same length, we only calculate loss on real sequence and not padding</p></li>\n<li><p>We use the smooth L1 Loss that combines being less sensitive to outliers and better convergence to near minima</p></li>\n<li><p>Normalization: we get the average loss per nucleotide position, making possible the comparison</p></li>\n</ol>\n<pre><code> ():\n    \n    \n    mask = torch.zeros_like(target, dtype=torch.)\n     i, length  (seq_lengths):\n        mask[i, :length, :] = \n\n    \n    loss = nn.SmoothL1Loss(reduction=)(output, target)\n\n    \n    masked_loss = loss * mask.()\n\n    \n     masked_loss.() / mask.()  mask.() &gt;   \n</code></pre>\n<h1>Training pipeline:</h1>\n<ol>\n<li><p>Training Loop Structure: This follows the standard PyTorch training pattern: zero gradients → forward pass → calculate loss → backward pass → update parameters.</p></li>\n<li><p>Gradient Clipping: This helps stabilize training by preventing large gradient updates, which is especially important for sequence models like this RNA folding model.</p></li>\n<li><p>Device Handling: The function correctly moves tensors to the specified device, allowing training on either CPU or GPU.</p></li>\n<li><p>Loss Tracking: The function tracks and returns the average loss, which is useful for monitoring training progress.</p></li>\n</ol>\n<pre><code>def train_epoch(model, dataloader, optimizer, device):\n    \n    Performs one training epoch  the model using the provided dataloader.\n\n    This   through all batches in the dataloader, running the \n     training cycle (forward pass, loss calculation, backpropagation, \n     parameter updates)  each batch. It handles device placement, gradient \n    clipping,  loss accumulation.\n\n    Parameters:\n    -----------\n    model : torch..Module\n        The neural network model  train (typically RNAFoldingModel)\n\n    dataloader : torch.utils.data.DataLoader\n        DataLoader providing batches of training data with features, targets,\n        ids,  sequence lengths\n\n    optimizer : torch.optim.Optimizer\n        The optimization algorithm used   model parameters\n\n    device : torch.device\n        The device (CPU  GPU)  perform computations \n\n    Returns:\n    --------\n    float\n        Average loss value  the epoch,  float()   batches were processed\n\n    Notes:\n    ------\n    - Sets the model  training  with model.train()\n    - Skips batches where targets are None\n    - Applies gradient clipping with max_norm=  prevent exploding gradients\n    - Accumulates loss   calculate the epoch average\n    \n    model.train()\n    epoch_loss = \n    batches = \n     features, targets, ids, seq_lengths in dataloader:\n         targets  None:\n            \n        optimizer.zero_grad()\n        features = features.(device)\n        targets = targets.(device)\n        outputs = model(features, seq_lengths)\n        loss = smooth_l1_loss(outputs, targets, seq_lengths)\n        loss.backward()\n        torch..utils.clip_grad_norm_(model.parameters(), max_norm=)\n        optimizer.step()\n        epoch_loss += loss.item()\n        batches += \n     epoch_loss / batches  batches &gt;   float()\n</code></pre>",
  "messages": [
    {
      "id": "3146553",
      "postDate": "03/11/2025 02:30:04",
      "content": "<p>I am using this <a href=\"https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook\" target=\"_blank\">notebook</a> as reference</p>\n<p>Today, I was exploring his training pipeline, which has two components, the first one is to create a new evaluation metric based on the Huber Loss function, this function is particularly important because:</p>\n<ol>\n<li><p>Handles the padding we set at the beginning to make sure the sequence has the same length, we only calculate loss on real sequence and not padding</p></li>\n<li><p>We use the smooth L1 Loss that combines being less sensitive to outliers and better convergence to near minima</p></li>\n<li><p>Normalization: we get the average loss per nucleotide position, making possible the comparison</p></li>\n</ol>\n<pre><code> ():\n    \n    \n    mask = torch.zeros_like(target, dtype=torch.)\n     i, length  (seq_lengths):\n        mask[i, :length, :] = \n\n    \n    loss = nn.SmoothL1Loss(reduction=)(output, target)\n\n    \n    masked_loss = loss * mask.()\n\n    \n     masked_loss.() / mask.()  mask.() &gt;   \n</code></pre>\n<h1>Training pipeline:</h1>\n<ol>\n<li><p>Training Loop Structure: This follows the standard PyTorch training pattern: zero gradients → forward pass → calculate loss → backward pass → update parameters.</p></li>\n<li><p>Gradient Clipping: This helps stabilize training by preventing large gradient updates, which is especially important for sequence models like this RNA folding model.</p></li>\n<li><p>Device Handling: The function correctly moves tensors to the specified device, allowing training on either CPU or GPU.</p></li>\n<li><p>Loss Tracking: The function tracks and returns the average loss, which is useful for monitoring training progress.</p></li>\n</ol>\n<pre><code>def train_epoch(model, dataloader, optimizer, device):\n    \n    Performs one training epoch  the model using the provided dataloader.\n\n    This   through all batches in the dataloader, running the \n     training cycle (forward pass, loss calculation, backpropagation, \n     parameter updates)  each batch. It handles device placement, gradient \n    clipping,  loss accumulation.\n\n    Parameters:\n    -----------\n    model : torch..Module\n        The neural network model  train (typically RNAFoldingModel)\n\n    dataloader : torch.utils.data.DataLoader\n        DataLoader providing batches of training data with features, targets,\n        ids,  sequence lengths\n\n    optimizer : torch.optim.Optimizer\n        The optimization algorithm used   model parameters\n\n    device : torch.device\n        The device (CPU  GPU)  perform computations \n\n    Returns:\n    --------\n    float\n        Average loss value  the epoch,  float()   batches were processed\n\n    Notes:\n    ------\n    - Sets the model  training  with model.train()\n    - Skips batches where targets are None\n    - Applies gradient clipping with max_norm=  prevent exploding gradients\n    - Accumulates loss   calculate the epoch average\n    \n    model.train()\n    epoch_loss = \n    batches = \n     features, targets, ids, seq_lengths in dataloader:\n         targets  None:\n            \n        optimizer.zero_grad()\n        features = features.(device)\n        targets = targets.(device)\n        outputs = model(features, seq_lengths)\n        loss = smooth_l1_loss(outputs, targets, seq_lengths)\n        loss.backward()\n        torch..utils.clip_grad_norm_(model.parameters(), max_norm=)\n        optimizer.step()\n        epoch_loss += loss.item()\n        batches += \n     epoch_loss / batches  batches &gt;   float()\n</code></pre>",
      "rawMarkdown": "I am using this [notebook](https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook) as reference\n\nToday, I was exploring his training pipeline, which has two components, the first one is to create a new evaluation metric based on the Huber Loss function, this function is particularly important because:\n\n1. Handles the padding we set at the beginning to make sure the sequence has the same length, we only calculate loss on real sequence and not padding\n\n2. We use the smooth L1 Loss that combines being less sensitive to outliers and better convergence to near minima\n\n3. Normalization: we get the average loss per nucleotide position, making possible the comparison\n\n```\ndef smooth_l1_loss(output, target, seq_lengths):\n    \"\"\"\n    Calculate masked Smooth L1 Loss for variable-length sequences.\n    \n    This function computes the Smooth L1 Loss (Huber Loss) between predicted and\n    target 3D coordinates, while ignoring padded positions in the sequences.\n    It ensures that only valid nucleotide positions contribute to the loss.\n    \n    Parameters:\n    -----------\n    output : torch.Tensor\n        Model predictions of shape (batch_size, max_seq_length, 3)\n        containing predicted 3D coordinates (x, y, z)\n    \n    target : torch.Tensor\n        Ground truth values of shape (batch_size, max_seq_length, 3)\n        containing actual 3D coordinates (x, y, z)\n    \n    seq_lengths : list\n        List of actual sequence lengths for each sample in the batch,\n        used to identify which positions are real vs. padding\n    \n    Returns:\n    --------\n    torch.Tensor\n        A scalar tensor containing the average Smooth L1 Loss computed\n        only over valid (non-padded) positions\n    \n    Notes:\n    ------\n    - Smooth L1 Loss combines L1 and L2 losses, being less sensitive to outliers\n      than MSE while providing better convergence properties than MAE\n    - The loss is normalized by the number of actual elements to ensure\n      fair comparison across batches with different sequence lengths\n    - Returns 0 if there are no valid positions (safety check)\n    \"\"\"\n    # Create a boolean mask for valid positions (non-padding)\n    mask = torch.zeros_like(target, dtype=torch.bool)\n    for i, length in enumerate(seq_lengths):\n        mask[i, :length, :] = True\n    \n    # Calculate Smooth L1 Loss without reduction\n    loss = nn.SmoothL1Loss(reduction='none')(output, target)\n    \n    # Apply mask to zero out loss from padding positions\n    masked_loss = loss * mask.float()\n    \n    # Return average loss over valid positions only\n    return masked_loss.sum() / mask.sum() if mask.sum() > 0 else 0\n```\n\n# Training pipeline:\n\n1. Training Loop Structure: This follows the standard PyTorch training pattern: zero gradients → forward pass → calculate loss → backward pass → update parameters.\n\n2. Gradient Clipping: This helps stabilize training by preventing large gradient updates, which is especially important for sequence models like this RNA folding model.\n\n3. Device Handling: The function correctly moves tensors to the specified device, allowing training on either CPU or GPU.\n\n4. Loss Tracking: The function tracks and returns the average loss, which is useful for monitoring training progress.\n\n```\ndef train_epoch(model, dataloader, optimizer, device):\n    \"\"\"\n    Performs one training epoch on the model using the provided dataloader.\n    \n    This function iterates through all batches in the dataloader, running the \n    complete training cycle (forward pass, loss calculation, backpropagation, \n    and parameter updates) for each batch. It handles device placement, gradient \n    clipping, and loss accumulation.\n    \n    Parameters:\n    -----------\n    model : torch.nn.Module\n        The neural network model to train (typically RNAFoldingModel)\n    \n    dataloader : torch.utils.data.DataLoader\n        DataLoader providing batches of training data with features, targets,\n        ids, and sequence lengths\n    \n    optimizer : torch.optim.Optimizer\n        The optimization algorithm used to update model parameters\n    \n    device : torch.device\n        The device (CPU or GPU) to perform computations on\n    \n    Returns:\n    --------\n    float\n        Average loss value for the epoch, or float('inf') if no batches were processed\n    \n    Notes:\n    ------\n    - Sets the model to training mode with model.train()\n    - Skips batches where targets are None\n    - Applies gradient clipping with max_norm=5.0 to prevent exploding gradients\n    - Accumulates loss values to calculate the epoch average\n    \"\"\"\n    model.train()\n    epoch_loss = 0\n    batches = 0\n    for features, targets, ids, seq_lengths in dataloader:\n        if targets is None:\n            continue\n        optimizer.zero_grad()\n        features = features.to(device)\n        targets = targets.to(device)\n        outputs = model(features, seq_lengths)\n        loss = smooth_l1_loss(outputs, targets, seq_lengths)\n        loss.backward()\n        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=5.0)\n        optimizer.step()\n        epoch_loss += loss.item()\n        batches += 1\n    return epoch_loss / batches if batches > 0 else float('inf')\n```",
      "votes": null
    },
    {
      "id": "3146709",
      "postDate": "03/11/2025 07:06:19",
      "content": "<p>this is wrong.</p>\n<ul>\n<li>we do not need absolute location. during evaluation, the rpediction is align to the ground truth before comparing the location</li>\n<li>the train data are not normalised to common xyz axis</li>\n</ul>",
      "rawMarkdown": "this is wrong.\n- we do not need absolute location. during evaluation, the rpediction is align to the ground truth before comparing the location\n- the train data are not normalised to common xyz axis",
      "votes": null
    },
    {
      "id": "3147152",
      "postDate": "03/11/2025 18:04:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, Thank you so much for your comment. Can you show me where I made the mistake, and do you know anything that can help me understand and fix the error?</p>\n<p>I appreciate your comment, I am trying to learn as much as I can.</p>",
      "rawMarkdown": "Hi @hengck23, Thank you so much for your comment. Can you show me where I made the mistake, and do you know anything that can help me understand and fix the error?\n\nI appreciate your comment, I am trying to learn as much as I can.",
      "votes": null
    },
    {
      "id": "3147165",
      "postDate": "03/11/2025 18:19:49",
      "content": "<p>see host example: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune</a></p>\n<p>in self-study for ML/AI in model building, the trick is NOT to do your own things but repeat and understand top implementations. </p>\n<p>yes! all enginerring work start from copying !!! </p>\n<p>but it is not blind copy …. we want to discover the devils in detail.</p>",
      "rawMarkdown": "see host example: https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune\n\nin self-study for ML/AI in model building, the trick is NOT to do your own things but repeat and understand top implementations. \n\nyes! all enginerring work start from copying !!! \n\nbut it is not blind copy .... we want to discover the devils in detail.",
      "votes": null
    },
    {
      "id": "3147184",
      "postDate": "03/11/2025 18:38:23",
      "content": "<p>Thank you! I will take your advice!! I appreciate your comment. </p>",
      "rawMarkdown": "Thank you! I will take your advice!! I appreciate your comment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3146709,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/11/2025 07:06:19",
      "content": "<p>this is wrong.</p>\n<ul>\n<li>we do not need absolute location. during evaluation, the rpediction is align to the ground truth before comparing the location</li>\n<li>the train data are not normalised to common xyz axis</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3147152,
          "author_name": "pastorsoto",
          "author_url": "",
          "post_date": "03/11/2025 18:04:25",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, Thank you so much for your comment. Can you show me where I made the mistake, and do you know anything that can help me understand and fix the error?</p>\n<p>I appreciate your comment, I am trying to learn as much as I can.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3147165,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/11/2025 18:19:49",
              "content": "<p>see host example: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune</a></p>\n<p>in self-study for ML/AI in model building, the trick is NOT to do your own things but repeat and understand top implementations. </p>\n<p>yes! all enginerring work start from copying !!! </p>\n<p>but it is not blind copy …. we want to discover the devils in detail.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3147184,
                  "author_name": "pastorsoto",
                  "author_url": "",
                  "post_date": "03/11/2025 18:38:23",
                  "content": "<p>Thank you! I will take your advice!! I appreciate your comment. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3146553": "I am using this [notebook](https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook) as reference\n\nToday, I was exploring his training pipeline, which has two components, the first one is to create a new evaluation metric based on the Huber Loss function, this function is particularly important because:\n\n1. Handles the padding we set at the beginning to make sure the sequence has the same length, we only calculate loss on real sequence and not padding\n\n2. We use the smooth L1 Loss that combines being less sensitive to outliers and better convergence to near minima\n\n3. Normalization: we get the average loss per nucleotide position, making possible the comparison\n\n```\ndef smooth_l1_loss(output, target, seq_lengths):\n    \"\"\"\n    Calculate masked Smooth L1 Loss for variable-length sequences.\n    \n    This function computes the Smooth L1 Loss (Huber Loss) between predicted and\n    target 3D coordinates, while ignoring padded positions in the sequences.\n    It ensures that only valid nucleotide positions contribute to the loss.\n    \n    Parameters:\n    -----------\n    output : torch.Tensor\n        Model predictions of shape (batch_size, max_seq_length, 3)\n        containing predicted 3D coordinates (x, y, z)\n    \n    target : torch.Tensor\n        Ground truth values of shape (batch_size, max_seq_length, 3)\n        containing actual 3D coordinates (x, y, z)\n    \n    seq_lengths : list\n        List of actual sequence lengths for each sample in the batch,\n        used to identify which positions are real vs. padding\n    \n    Returns:\n    --------\n    torch.Tensor\n        A scalar tensor containing the average Smooth L1 Loss computed\n        only over valid (non-padded) positions\n    \n    Notes:\n    ------\n    - Smooth L1 Loss combines L1 and L2 losses, being less sensitive to outliers\n      than MSE while providing better convergence properties than MAE\n    - The loss is normalized by the number of actual elements to ensure\n      fair comparison across batches with different sequence lengths\n    - Returns 0 if there are no valid positions (safety check)\n    \"\"\"\n    # Create a boolean mask for valid positions (non-padding)\n    mask = torch.zeros_like(target, dtype=torch.bool)\n    for i, length in enumerate(seq_lengths):\n        mask[i, :length, :] = True\n    \n    # Calculate Smooth L1 Loss without reduction\n    loss = nn.SmoothL1Loss(reduction='none')(output, target)\n    \n    # Apply mask to zero out loss from padding positions\n    masked_loss = loss * mask.float()\n    \n    # Return average loss over valid positions only\n    return masked_loss.sum() / mask.sum() if mask.sum() > 0 else 0\n```\n\n# Training pipeline:\n\n1. Training Loop Structure: This follows the standard PyTorch training pattern: zero gradients → forward pass → calculate loss → backward pass → update parameters.\n\n2. Gradient Clipping: This helps stabilize training by preventing large gradient updates, which is especially important for sequence models like this RNA folding model.\n\n3. Device Handling: The function correctly moves tensors to the specified device, allowing training on either CPU or GPU.\n\n4. Loss Tracking: The function tracks and returns the average loss, which is useful for monitoring training progress.\n\n```\ndef train_epoch(model, dataloader, optimizer, device):\n    \"\"\"\n    Performs one training epoch on the model using the provided dataloader.\n    \n    This function iterates through all batches in the dataloader, running the \n    complete training cycle (forward pass, loss calculation, backpropagation, \n    and parameter updates) for each batch. It handles device placement, gradient \n    clipping, and loss accumulation.\n    \n    Parameters:\n    -----------\n    model : torch.nn.Module\n        The neural network model to train (typically RNAFoldingModel)\n    \n    dataloader : torch.utils.data.DataLoader\n        DataLoader providing batches of training data with features, targets,\n        ids, and sequence lengths\n    \n    optimizer : torch.optim.Optimizer\n        The optimization algorithm used to update model parameters\n    \n    device : torch.device\n        The device (CPU or GPU) to perform computations on\n    \n    Returns:\n    --------\n    float\n        Average loss value for the epoch, or float('inf') if no batches were processed\n    \n    Notes:\n    ------\n    - Sets the model to training mode with model.train()\n    - Skips batches where targets are None\n    - Applies gradient clipping with max_norm=5.0 to prevent exploding gradients\n    - Accumulates loss values to calculate the epoch average\n    \"\"\"\n    model.train()\n    epoch_loss = 0\n    batches = 0\n    for features, targets, ids, seq_lengths in dataloader:\n        if targets is None:\n            continue\n        optimizer.zero_grad()\n        features = features.to(device)\n        targets = targets.to(device)\n        outputs = model(features, seq_lengths)\n        loss = smooth_l1_loss(outputs, targets, seq_lengths)\n        loss.backward()\n        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=5.0)\n        optimizer.step()\n        epoch_loss += loss.item()\n        batches += 1\n    return epoch_loss / batches if batches > 0 else float('inf')\n```",
    "3146709": "this is wrong.\n- we do not need absolute location. during evaluation, the rpediction is align to the ground truth before comparing the location\n- the train data are not normalised to common xyz axis",
    "3147152": "Hi @hengck23, Thank you so much for your comment. Can you show me where I made the mistake, and do you know anything that can help me understand and fix the error?\n\nI appreciate your comment, I am trying to learn as much as I can.",
    "3147165": "see host example: https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune\n\nin self-study for ML/AI in model building, the trick is NOT to do your own things but repeat and understand top implementations. \n\nyes! all enginerring work start from copying !!! \n\nbut it is not blind copy .... we want to discover the devils in detail.",
    "3147184": "Thank you! I will take your advice!! I appreciate your comment."
  },
  "source": "meta"
}