{
  "id": 440698,
  "title": "Simple Clustering of the RNA sequences as a minimal start?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/440698",
  "author_name": "",
  "post_date": "2023-09-15T20:39:41.018499500Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have no idea on how to solve the whole problem, but just looking at the RNA string sequences, I'm wondering about using an algorithm called K-medoids to at least get some minimal, useful data-mining by clustering the RNA-sequence strings. Basically, you could use K-medoids (a discrete version of K-means that works in arbitrary metric spaces), and then use the \"edit-distance\" metric: the edit-distance dist(A, B) = # of minimum single-character changes to transform string A into string B, and vice versa. Of course this doesn't solve the prediction problem, but I think this could be a relatively easy way to get some minimum data mining out of the RNA sequences.</p>\n<p>I believe the \"pyclustering library\" (<a href=\"https://pyclustering.github.io/\" target=\"_blank\">https://pyclustering.github.io/</a>) has a section on the k-medoids algorithm, and the edit-distance metric could be programmed as a function in Python to use in the clustering algorithm.</p>",
  "messages": [
    {
      "id": "2440948",
      "postDate": "09/15/2023 20:39:41",
      "content": "<p>I have no idea on how to solve the whole problem, but just looking at the RNA string sequences, I'm wondering about using an algorithm called K-medoids to at least get some minimal, useful data-mining by clustering the RNA-sequence strings. Basically, you could use K-medoids (a discrete version of K-means that works in arbitrary metric spaces), and then use the \"edit-distance\" metric: the edit-distance dist(A, B) = # of minimum single-character changes to transform string A into string B, and vice versa. Of course this doesn't solve the prediction problem, but I think this could be a relatively easy way to get some minimum data mining out of the RNA sequences.</p>\n<p>I believe the \"pyclustering library\" (<a href=\"https://pyclustering.github.io/\" target=\"_blank\">https://pyclustering.github.io/</a>) has a section on the k-medoids algorithm, and the edit-distance metric could be programmed as a function in Python to use in the clustering algorithm.</p>",
      "rawMarkdown": "I have no idea on how to solve the whole problem, but just looking at the RNA string sequences, I'm wondering about using an algorithm called K-medoids to at least get some minimal, useful data-mining by clustering the RNA-sequence strings. Basically, you could use K-medoids (a discrete version of K-means that works in arbitrary metric spaces), and then use the \"edit-distance\" metric: the edit-distance dist(A, B) = # of minimum single-character changes to transform string A into string B, and vice versa. Of course this doesn't solve the prediction problem, but I think this could be a relatively easy way to get some minimum data mining out of the RNA sequences.\n\nI believe the \"pyclustering library\" (https://pyclustering.github.io/) has a section on the k-medoids algorithm, and the edit-distance metric could be programmed as a function in Python to use in the clustering algorithm.",
      "votes": null
    },
    {
      "id": "2442692",
      "postDate": "09/17/2023 07:12:22",
      "content": "<p>Keep working on it….</p>",
      "rawMarkdown": "Keep working on it....",
      "votes": null
    },
    {
      "id": "2444740",
      "postDate": "09/18/2023 12:59:18",
      "content": "<p>Look at: <a href=\"https://en.wikipedia.org/wiki/Sequence_clustering\" target=\"_blank\">https://en.wikipedia.org/wiki/Sequence_clustering</a><br>\nCH-hit:  <a href=\"https://www.kaggle.com/code/alexandervc/cd-hit-sequence-clustering\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/cd-hit-sequence-clustering</a><br>\nDiamond + graph clustering: <a href=\"https://www.kaggle.com/code/alexandervc/cafa5-23-groups-and-folds-diamond-igraph\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/cafa5-23-groups-and-folds-diamond-igraph</a></p>",
      "rawMarkdown": "Look at: https://en.wikipedia.org/wiki/Sequence_clustering\nCH-hit:  https://www.kaggle.com/code/alexandervc/cd-hit-sequence-clustering\nDiamond + graph clustering: https://www.kaggle.com/code/alexandervc/cafa5-23-groups-and-folds-diamond-igraph",
      "votes": null
    },
    {
      "id": "2445448",
      "postDate": "09/18/2023 20:45:33",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/Alexander\" target=\"_blank\">@Alexander</a>! I'll check these resources out.</p>",
      "rawMarkdown": "Thanks @Alexander! I'll check these resources out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2442692,
      "author_name": "josealexanderrubio",
      "author_url": "",
      "post_date": "09/17/2023 07:12:22",
      "content": "<p>Keep working on it….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2444740,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/18/2023 12:59:18",
      "content": "<p>Look at: <a href=\"https://en.wikipedia.org/wiki/Sequence_clustering\" target=\"_blank\">https://en.wikipedia.org/wiki/Sequence_clustering</a><br>\nCH-hit:  <a href=\"https://www.kaggle.com/code/alexandervc/cd-hit-sequence-clustering\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/cd-hit-sequence-clustering</a><br>\nDiamond + graph clustering: <a href=\"https://www.kaggle.com/code/alexandervc/cafa5-23-groups-and-folds-diamond-igraph\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/cafa5-23-groups-and-folds-diamond-igraph</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2445448,
          "author_name": "raveenaj",
          "author_url": "",
          "post_date": "09/18/2023 20:45:33",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/Alexander\" target=\"_blank\">@Alexander</a>! I'll check these resources out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2440948": "I have no idea on how to solve the whole problem, but just looking at the RNA string sequences, I'm wondering about using an algorithm called K-medoids to at least get some minimal, useful data-mining by clustering the RNA-sequence strings. Basically, you could use K-medoids (a discrete version of K-means that works in arbitrary metric spaces), and then use the \"edit-distance\" metric: the edit-distance dist(A, B) = # of minimum single-character changes to transform string A into string B, and vice versa. Of course this doesn't solve the prediction problem, but I think this could be a relatively easy way to get some minimum data mining out of the RNA sequences.\n\nI believe the \"pyclustering library\" (https://pyclustering.github.io/) has a section on the k-medoids algorithm, and the edit-distance metric could be programmed as a function in Python to use in the clustering algorithm.",
    "2442692": "Keep working on it....",
    "2444740": "Look at: https://en.wikipedia.org/wiki/Sequence_clustering\nCH-hit:  https://www.kaggle.com/code/alexandervc/cd-hit-sequence-clustering\nDiamond + graph clustering: https://www.kaggle.com/code/alexandervc/cafa5-23-groups-and-folds-diamond-igraph",
    "2445448": "Thanks @Alexander! I'll check these resources out."
  },
  "source": "meta"
}