{
  "id": 575358,
  "title": "not able to map the sequence to numeric data it take to much resouce ",
  "url": "/competitions/stanford-rna-3d-folding/discussion/575358",
  "author_name": "Shivanshu Prajapati",
  "post_date": "2025-04-28T02:07:11.908000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I cannot map the sequence to numeric data because it requires too many resources.</p>\n<p>here is the method I m using  \"<br>\ndef create_seq(train_seq):<br>\n    for base in [\"A\",\"G\",\"C\",\"U\"]:<br>\n        train_seq[f\"{base.lower()}_count\"]=train_seq[\"sequence\"].str.count(base)</p>\n<pre><code>nucleotide_map = {: , : , \n                  : , : ,\n                  : }  \ndef one_hott_encode(seq):\n     [nucleotide_map.get(nt,)  nt in seq]\ntrain_se=train_se.apply(one_hott_encode)\n train_seq\n</code></pre>\n<p>\" now converting the sequence hot it take many computing power \"<br>\nimport numpy as np<br>\nimport torch<br>\nfrom torch.nn.utils.rnn import pad_sequence</p>\n<h1>Convert sequence_hot column to list of Tensors</h1>\n<p>x_train_seq = [torch.tensor(seq, dtype=torch.float32) for seq in x_train['sequence_hot']]<br>\nx_test_seq  = [torch.tensor(seq, dtype=torch.float32) for seq in x_test['sequence_hot']]</p>\n<h1>Pad sequences so they are all the same length</h1>\n<p>x_train_seq_padded = pad_sequence(x_train_seq, batch_first=True)<br>\nx_test_seq_padded  = pad_sequence(x_test_seq, batch_first=True)</p>\n<h1>Convert nucleotide counts to tensor</h1>\n<p>x_train_counts = torch.tensor(x_train[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)<br>\nx_test_counts  = torch.tensor(x_test[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)</p>\n<h1>Concatenate padded sequence and count features</h1>\n<p>x_train_tensor = torch.cat((x_train_seq_padded, x_train_counts), dim=1)<br>\nx_test_tensor  = torch.cat((x_test_seq_padded,  x_test_counts), dim=1)<br>\n\"</p>",
  "messages": [
    {
      "id": 3188650,
      "postDate": "2025-04-28T02:07:11.910Z",
      "content": "<p>I cannot map the sequence to numeric data because it requires too many resources.</p>\n<p>here is the method I m using  \"<br>\ndef create_seq(train_seq):<br>\n    for base in [\"A\",\"G\",\"C\",\"U\"]:<br>\n        train_seq[f\"{base.lower()}_count\"]=train_seq[\"sequence\"].str.count(base)</p>\n<pre><code>nucleotide_map = {: , : , \n                  : , : ,\n                  : }  \ndef one_hott_encode(seq):\n     [nucleotide_map.get(nt,)  nt in seq]\ntrain_se=train_se.apply(one_hott_encode)\n train_seq\n</code></pre>\n<p>\" now converting the sequence hot it take many computing power \"<br>\nimport numpy as np<br>\nimport torch<br>\nfrom torch.nn.utils.rnn import pad_sequence</p>\n<h1>Convert sequence_hot column to list of Tensors</h1>\n<p>x_train_seq = [torch.tensor(seq, dtype=torch.float32) for seq in x_train['sequence_hot']]<br>\nx_test_seq  = [torch.tensor(seq, dtype=torch.float32) for seq in x_test['sequence_hot']]</p>\n<h1>Pad sequences so they are all the same length</h1>\n<p>x_train_seq_padded = pad_sequence(x_train_seq, batch_first=True)<br>\nx_test_seq_padded  = pad_sequence(x_test_seq, batch_first=True)</p>\n<h1>Convert nucleotide counts to tensor</h1>\n<p>x_train_counts = torch.tensor(x_train[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)<br>\nx_test_counts  = torch.tensor(x_test[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)</p>\n<h1>Concatenate padded sequence and count features</h1>\n<p>x_train_tensor = torch.cat((x_train_seq_padded, x_train_counts), dim=1)<br>\nx_test_tensor  = torch.cat((x_test_seq_padded,  x_test_counts), dim=1)<br>\n\"</p>",
      "rawMarkdown": "I cannot map the sequence to numeric data because it requires too many resources.\n\nhere is the method I m using  \"\ndef create_seq(train_seq):\n    for base in [\"A\",\"G\",\"C\",\"U\"]:\n        train_seq[f\"{base.lower()}_count\"]=train_seq[\"sequence\"].str.count(base)\n        \n    nucleotide_map = {'A': 1, 'G': 4, \n                      'C': 2, 'U': 3,\n                      'T': 3}  # Treat T as U\n    def one_hott_encode(seq):\n        return [nucleotide_map.get(nt,0) for nt in seq]\n    train_seq[\"sequence_hot\"]=train_seq[\"sequence\"].apply(one_hott_encode)\n    return train_seq\n\n\" now converting the sequence hot it take many computing power \"\nimport numpy as np\nimport torch\nfrom torch.nn.utils.rnn import pad_sequence\n\n# Convert sequence_hot column to list of Tensors\nx_train_seq = [torch.tensor(seq, dtype=torch.float32) for seq in x_train['sequence_hot']]\nx_test_seq  = [torch.tensor(seq, dtype=torch.float32) for seq in x_test['sequence_hot']]\n\n# Pad sequences so they are all the same length\nx_train_seq_padded = pad_sequence(x_train_seq, batch_first=True)\nx_test_seq_padded  = pad_sequence(x_test_seq, batch_first=True)\n\n# Convert nucleotide counts to tensor\nx_train_counts = torch.tensor(x_train[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)\nx_test_counts  = torch.tensor(x_test[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)\n\n# Concatenate padded sequence and count features\nx_train_tensor = torch.cat((x_train_seq_padded, x_train_counts), dim=1)\nx_test_tensor  = torch.cat((x_test_seq_padded,  x_test_counts), dim=1)\n\"",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3188650": "I cannot map the sequence to numeric data because it requires too many resources.\n\nhere is the method I m using  \"\ndef create_seq(train_seq):\n    for base in [\"A\",\"G\",\"C\",\"U\"]:\n        train_seq[f\"{base.lower()}_count\"]=train_seq[\"sequence\"].str.count(base)\n        \n    nucleotide_map = {'A': 1, 'G': 4, \n                      'C': 2, 'U': 3,\n                      'T': 3}  # Treat T as U\n    def one_hott_encode(seq):\n        return [nucleotide_map.get(nt,0) for nt in seq]\n    train_seq[\"sequence_hot\"]=train_seq[\"sequence\"].apply(one_hott_encode)\n    return train_seq\n\n\" now converting the sequence hot it take many computing power \"\nimport numpy as np\nimport torch\nfrom torch.nn.utils.rnn import pad_sequence\n\n# Convert sequence_hot column to list of Tensors\nx_train_seq = [torch.tensor(seq, dtype=torch.float32) for seq in x_train['sequence_hot']]\nx_test_seq  = [torch.tensor(seq, dtype=torch.float32) for seq in x_test['sequence_hot']]\n\n# Pad sequences so they are all the same length\nx_train_seq_padded = pad_sequence(x_train_seq, batch_first=True)\nx_test_seq_padded  = pad_sequence(x_test_seq, batch_first=True)\n\n# Convert nucleotide counts to tensor\nx_train_counts = torch.tensor(x_train[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)\nx_test_counts  = torch.tensor(x_test[['a_count','g_count','c_count','u_count']].values, dtype=torch.float32)\n\n# Concatenate padded sequence and count features\nx_train_tensor = torch.cat((x_train_seq_padded, x_train_counts), dim=1)\nx_test_tensor  = torch.cat((x_test_seq_padded,  x_test_counts), dim=1)\n\""
  }
}