{
  "id": 448959,
  "title": "duplicated samples in training dataset ",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/448959",
  "author_name": "yuanzhe zhou",
  "post_date": "2023-10-22T10:10:32.489000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>It seems that there are samples with same sequences but difference SNR, how did you hanlde them?</p>",
  "messages": [
    {
      "id": 2492180,
      "postDate": "2023-10-22T10:10:32.490Z",
      "content": "<p>It seems that there are samples with same sequences but difference SNR, how did you hanlde them?</p>",
      "rawMarkdown": "It seems that there are samples with same sequences but difference SNR, how did you hanlde them?",
      "votes": 4
    },
    {
      "id": 2493726,
      "postDate": "2023-10-23T14:36:28.413Z",
      "content": "<p>I just remove them, but I still have the same question with you: why this situation exist?</p>",
      "rawMarkdown": "I just remove them, but I still have the same question with you: why this situation exist?",
      "replies": [
        {
          "id": 2496432,
          "postDate": "2023-10-24T03:54:07.513Z",
          "content": "<p>The sample with highest snr is better in my opinion.</p>",
          "rawMarkdown": "The sample with highest snr is better in my opinion.",
          "replies": [
            {
              "id": 2497944,
              "postDate": "2023-10-25T03:03:35.023Z",
              "content": "<p>BTW, can I join your team?😀</p>",
              "rawMarkdown": "BTW, can I join your team?😀"
            }
          ]
        },
        {
          "id": 2497507,
          "postDate": "2023-10-24T17:09:09.843Z",
          "content": "<p>The replicate data are from independent experiments and should be quite similar to each other. You could average them, pick the replicate with highest signal to noise, or leave them as different training examples to help reduce overfitting. </p>\n<p>Since there aren’t too many of these replicates it may not matter too much.</p>\n<p>But  I you try different ways to handle replicates feel free to share results!</p>",
          "rawMarkdown": "The replicate data are from independent experiments and should be quite similar to each other. You could average them, pick the replicate with highest signal to noise, or leave them as different training examples to help reduce overfitting. \n\nSince there aren’t too many of these replicates it may not matter too much.\n\nBut  I you try different ways to handle replicates feel free to share results!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2501062,
      "postDate": "2023-10-27T08:14:50.513Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2501063,
          "postDate": "2023-10-27T08:15:17.447Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2492837,
      "postDate": "2023-10-22T23:57:32.693Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2493726,
      "author_name": "imaginist",
      "author_url": "",
      "post_date": "2023-10-23T14:36:28.413000",
      "content": "<p>I just remove them, but I still have the same question with you: why this situation exist?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2496432,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2023-10-24T03:54:07.513000",
          "content": "<p>The sample with highest snr is better in my opinion.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2497944,
              "author_name": "imaginist",
              "author_url": "",
              "post_date": "2023-10-25T03:03:35.023000",
              "content": "<p>BTW, can I join your team?😀</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2497507,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2023-10-24T17:09:09.843000",
          "content": "<p>The replicate data are from independent experiments and should be quite similar to each other. You could average them, pick the replicate with highest signal to noise, or leave them as different training examples to help reduce overfitting. </p>\n<p>Since there aren’t too many of these replicates it may not matter too much.</p>\n<p>But  I you try different ways to handle replicates feel free to share results!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2501062,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-27T08:14:50.513000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2501063,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-10-27T08:15:17.447000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2492837,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-22T23:57:32.693000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2492180": "It seems that there are samples with same sequences but difference SNR, how did you hanlde them?",
    "2493726": "I just remove them, but I still have the same question with you: why this situation exist?",
    "2501062": "",
    "2492837": ""
  }
}