{
  "id": 567025,
  "title": "3rd day: DataLoaders and padding",
  "url": "/competitions/stanford-rna-3d-folding/discussion/567025",
  "author_name": "Pastor Soto",
  "post_date": "2025-03-08T01:23:37.875000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Third day of kaggle competition, I am following this <a href=\"https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook\" target=\"_blank\">notebook</a> that shows how to make a submission and prepare the training loop.</p>\n<p>Today I was able to put together the data loaders in which there are important concepts to evaluate:</p>\n<ol>\n<li><p>Padding: </p>\n<ul>\n<li>Ensures all sequences in a batch have the same length by adding zeros.</li></ul></li>\n<li><p>Stacking:</p>\n<ul>\n<li>Combines padded sequences into a single tensor for efficient processing.</li></ul></li>\n<li><p>Handling Variable-Length Sequences:</p>\n<ul>\n<li>Allows the model to process sequences of different lengths in the same batch.</li></ul></li>\n</ol>\n<pre><code> ():\n     ():\n        .data = data\n     ():\n         (.data)\n     ():\n        item = .data[idx]\n        features = torch.tensor(item[], dtype=torch.float32)\n         item[]   :\n            target = torch.tensor(item[][], dtype=torch.float32)\n             features, target, item[]\n        :\n             features, , item[]\n</code></pre>\n<p>Overall, the code wasn't difficult to understand but it is important to have the concepts clear to make a good understanding of how this work.</p>",
  "messages": [
    {
      "id": 3144128,
      "postDate": "2025-03-08T01:23:37.877Z",
      "content": "<p>Third day of kaggle competition, I am following this <a href=\"https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook\" target=\"_blank\">notebook</a> that shows how to make a submission and prepare the training loop.</p>\n<p>Today I was able to put together the data loaders in which there are important concepts to evaluate:</p>\n<ol>\n<li><p>Padding: </p>\n<ul>\n<li>Ensures all sequences in a batch have the same length by adding zeros.</li></ul></li>\n<li><p>Stacking:</p>\n<ul>\n<li>Combines padded sequences into a single tensor for efficient processing.</li></ul></li>\n<li><p>Handling Variable-Length Sequences:</p>\n<ul>\n<li>Allows the model to process sequences of different lengths in the same batch.</li></ul></li>\n</ol>\n<pre><code> ():\n     ():\n        .data = data\n     ():\n         (.data)\n     ():\n        item = .data[idx]\n        features = torch.tensor(item[], dtype=torch.float32)\n         item[]   :\n            target = torch.tensor(item[][], dtype=torch.float32)\n             features, target, item[]\n        :\n             features, , item[]\n</code></pre>\n<p>Overall, the code wasn't difficult to understand but it is important to have the concepts clear to make a good understanding of how this work.</p>",
      "rawMarkdown": "Third day of kaggle competition, I am following this [notebook](https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook) that shows how to make a submission and prepare the training loop.\n\nToday I was able to put together the data loaders in which there are important concepts to evaluate:\n\n1. Padding: \n    - Ensures all sequences in a batch have the same length by adding zeros.\n\n2. Stacking:\n     - Combines padded sequences into a single tensor for efficient processing.\n\n3. Handling Variable-Length Sequences:\n      - Allows the model to process sequences of different lengths in the same batch.\n\n```\nclass RNADataset(Dataset):\n    def __init__(self, data):\n        self.data = data\n    def __len__(self):\n        return len(self.data)\n    def __getitem__(self, idx):\n        item = self.data[idx]\n        features = torch.tensor(item['features'], dtype=torch.float32)\n        if item['structures'] is not None:\n            target = torch.tensor(item['structures'][0], dtype=torch.float32)\n            return features, target, item['id']\n        else:\n            return features, None, item['id']\n```\n\nOverall, the code wasn't difficult to understand but it is important to have the concepts clear to make a good understanding of how this work.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3144128": "Third day of kaggle competition, I am following this [notebook](https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition/notebook) that shows how to make a submission and prepare the training loop.\n\nToday I was able to put together the data loaders in which there are important concepts to evaluate:\n\n1. Padding: \n    - Ensures all sequences in a batch have the same length by adding zeros.\n\n2. Stacking:\n     - Combines padded sequences into a single tensor for efficient processing.\n\n3. Handling Variable-Length Sequences:\n      - Allows the model to process sequences of different lengths in the same batch.\n\n```\nclass RNADataset(Dataset):\n    def __init__(self, data):\n        self.data = data\n    def __len__(self):\n        return len(self.data)\n    def __getitem__(self, idx):\n        item = self.data[idx]\n        features = torch.tensor(item['features'], dtype=torch.float32)\n        if item['structures'] is not None:\n            target = torch.tensor(item['structures'][0], dtype=torch.float32)\n            return features, target, item['id']\n        else:\n            return features, None, item['id']\n```\n\nOverall, the code wasn't difficult to understand but it is important to have the concepts clear to make a good understanding of how this work."
  }
}