{
  "id": 582377,
  "title": "Description of my solution for the competition / Descripción de mi solución para la competición",
  "url": "/competitions/stanford-rna-3d-folding/discussion/582377",
  "author_name": "Rem1210",
  "post_date": "2025-05-30T18:10:32.160000",
  "votes": 23,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I’m sharing the steps I followed to develop my solution (score: 0.484, 6th place at the end of the submission phase). This is an initial summary that I plan to expand in future updates.</p>\n<hr>\n<p><strong>Considerations:</strong></p>\n<ul>\n<li>No external models were used.  </li>\n<li>No additional data was used beyond what was provided for the training phase.  </li>\n<li>The solution is based solely on the input RNA sequences (e.g., AAGGCCUU…), with no extra information.</li>\n</ul>\n<p><strong>Note:</strong></p>\n<p>Given the limited time and available resources (CPU i5-6500, 24 GB RAM, no GPU), I’m aware that each step can be significantly improved.</p>\n<hr>\n<p><strong>Process overview:</strong></p>\n<p>Briefly, the process consisted of the following steps:</p>\n<ol>\n<li>Data acquisition, unification, and selection.  </li>\n<li>TM-score calculation between each sequence and all others, grouped by sequence length to reduce computational cost.  </li>\n<li>Generation of a distance matrix, clustering into n groups (classes), and selection of the most representative sequences in each group.  </li>\n<li>Training of a model (Keras) to classify sequences into 5 groups based on the input sequence.  </li>\n<li>Model inference. For each input sequence, the model predicts the most likely groups (top 5), and returns the 3D sequence of the most similar entry within each predicted group (based on sequence length and best TM-score within that group).</li>\n</ol>\n<hr>\n<p><strong>Additional note:</strong></p>\n<p>Step 2 is computationally expensive, especially for long sequences. I considered optimizing this step by incorporating PDB data or using data augmentation techniques, but discarded those options due to hardware limitations. These ideas could be implemented as future extensions.</p>\n<hr>\n<p>It’s been a pleasure to take part in this competition — I’ve learned a lot and really enjoyed the process. Many thanks to the organizers and the Kaggle community for making this challenge possible. Best of luck to the other participants!</p>\n<hr>\n<h2>Versión en español</h2>\n<h3>1. Descripción de mi solución para la competición</h3>\n<p>Comparto los pasos que seguí para desarrollar mi solución (score: 0.484, 6.º puesto al cierre de la fase de envíos). Este es un resumen inicial que trataré de ir ampliando en futuras actualizaciones.</p>\n<p><strong>Consideraciones:</strong></p>\n<ul>\n<li>No se utilizó ningún modelo externo.  </li>\n<li>No se emplearon datos adicionales más allá de los proporcionados para la fase de entrenamiento.  </li>\n<li>La solución se basa únicamente en las secuencias de ARN de entrada (ej. AAGGCCUU…), sin información extra.</li>\n</ul>\n<p><strong>Nota:</strong></p>\n<p>Dado el tiempo limitado y los recursos disponibles (CPU i5-6500, 24 GB de RAM, sin GPU), soy consciente de que cada uno de los pasos es claramente mejorable.</p>\n<hr>\n<p><strong>Proceso seguido:</strong></p>\n<p>De forma resumida, el proceso consistió en los siguientes pasos:</p>\n<ol>\n<li>Obtención, unificación y selección de datos.  </li>\n<li>Cálculo del TM-score entre cada secuencia y todas las demás, agrupando por tamaño de secuencia para tratar de reducir el coste computacional.  </li>\n<li>Generación de la matriz de distancias, <em>clustering</em> (división en <em>n</em> grupos o clases) y selección de las secuencias más representativas de cada grupo.  </li>\n<li>Entrenamiento de un modelo (Keras) para clasificar las secuencias en 5 grupos a partir de la secuencia de entrada.  </li>\n<li>Ejecución del modelo. Para cada nueva secuencia, se predice su grupo (los 5 mejores) y se devuelve la secuencia 3D de la entrada más similar dentro de ese grupo (según el tamaño de secuencia y el mejor TM-score dentro de ese rango), para cada uno de los 5 mejores grupos predichos.</li>\n</ol>\n<hr>\n<p><strong>Nota adicional:</strong></p>\n<p>El paso 2 es computacionalmente muy costoso, especialmente para secuencias largas. Se exploró la posibilidad de optimizar este paso incorporando datos de la base PDB o mediante <em>data augmentation</em>, pero se descartó debido a las limitaciones de hardware. Estas mejoras podrían implementarse como una extensión futura.</p>\n<hr>\n<p>Ha sido un placer participar en esta competición, he aprendido mucho y he disfrutado del proceso. Agradezco a los organizadores y a la comunidad de Kaggle por hacer posible este reto. ¡Mucha suerte al resto de participantes!</p>",
  "messages": [
    {
      "id": 3213996,
      "postDate": "2025-05-30T18:10:32.160Z",
      "content": "<p>I’m sharing the steps I followed to develop my solution (score: 0.484, 6th place at the end of the submission phase). This is an initial summary that I plan to expand in future updates.</p>\n<hr>\n<p><strong>Considerations:</strong></p>\n<ul>\n<li>No external models were used.  </li>\n<li>No additional data was used beyond what was provided for the training phase.  </li>\n<li>The solution is based solely on the input RNA sequences (e.g., AAGGCCUU…), with no extra information.</li>\n</ul>\n<p><strong>Note:</strong></p>\n<p>Given the limited time and available resources (CPU i5-6500, 24 GB RAM, no GPU), I’m aware that each step can be significantly improved.</p>\n<hr>\n<p><strong>Process overview:</strong></p>\n<p>Briefly, the process consisted of the following steps:</p>\n<ol>\n<li>Data acquisition, unification, and selection.  </li>\n<li>TM-score calculation between each sequence and all others, grouped by sequence length to reduce computational cost.  </li>\n<li>Generation of a distance matrix, clustering into n groups (classes), and selection of the most representative sequences in each group.  </li>\n<li>Training of a model (Keras) to classify sequences into 5 groups based on the input sequence.  </li>\n<li>Model inference. For each input sequence, the model predicts the most likely groups (top 5), and returns the 3D sequence of the most similar entry within each predicted group (based on sequence length and best TM-score within that group).</li>\n</ol>\n<hr>\n<p><strong>Additional note:</strong></p>\n<p>Step 2 is computationally expensive, especially for long sequences. I considered optimizing this step by incorporating PDB data or using data augmentation techniques, but discarded those options due to hardware limitations. These ideas could be implemented as future extensions.</p>\n<hr>\n<p>It’s been a pleasure to take part in this competition — I’ve learned a lot and really enjoyed the process. Many thanks to the organizers and the Kaggle community for making this challenge possible. Best of luck to the other participants!</p>\n<hr>\n<h2>Versión en español</h2>\n<h3>1. Descripción de mi solución para la competición</h3>\n<p>Comparto los pasos que seguí para desarrollar mi solución (score: 0.484, 6.º puesto al cierre de la fase de envíos). Este es un resumen inicial que trataré de ir ampliando en futuras actualizaciones.</p>\n<p><strong>Consideraciones:</strong></p>\n<ul>\n<li>No se utilizó ningún modelo externo.  </li>\n<li>No se emplearon datos adicionales más allá de los proporcionados para la fase de entrenamiento.  </li>\n<li>La solución se basa únicamente en las secuencias de ARN de entrada (ej. AAGGCCUU…), sin información extra.</li>\n</ul>\n<p><strong>Nota:</strong></p>\n<p>Dado el tiempo limitado y los recursos disponibles (CPU i5-6500, 24 GB de RAM, sin GPU), soy consciente de que cada uno de los pasos es claramente mejorable.</p>\n<hr>\n<p><strong>Proceso seguido:</strong></p>\n<p>De forma resumida, el proceso consistió en los siguientes pasos:</p>\n<ol>\n<li>Obtención, unificación y selección de datos.  </li>\n<li>Cálculo del TM-score entre cada secuencia y todas las demás, agrupando por tamaño de secuencia para tratar de reducir el coste computacional.  </li>\n<li>Generación de la matriz de distancias, <em>clustering</em> (división en <em>n</em> grupos o clases) y selección de las secuencias más representativas de cada grupo.  </li>\n<li>Entrenamiento de un modelo (Keras) para clasificar las secuencias en 5 grupos a partir de la secuencia de entrada.  </li>\n<li>Ejecución del modelo. Para cada nueva secuencia, se predice su grupo (los 5 mejores) y se devuelve la secuencia 3D de la entrada más similar dentro de ese grupo (según el tamaño de secuencia y el mejor TM-score dentro de ese rango), para cada uno de los 5 mejores grupos predichos.</li>\n</ol>\n<hr>\n<p><strong>Nota adicional:</strong></p>\n<p>El paso 2 es computacionalmente muy costoso, especialmente para secuencias largas. Se exploró la posibilidad de optimizar este paso incorporando datos de la base PDB o mediante <em>data augmentation</em>, pero se descartó debido a las limitaciones de hardware. Estas mejoras podrían implementarse como una extensión futura.</p>\n<hr>\n<p>Ha sido un placer participar en esta competición, he aprendido mucho y he disfrutado del proceso. Agradezco a los organizadores y a la comunidad de Kaggle por hacer posible este reto. ¡Mucha suerte al resto de participantes!</p>",
      "rawMarkdown": "I’m sharing the steps I followed to develop my solution (score: 0.484, 6th place at the end of the submission phase). This is an initial summary that I plan to expand in future updates.\n\n---\n\n**Considerations:**\n\n- No external models were used.  \n- No additional data was used beyond what was provided for the training phase.  \n- The solution is based solely on the input RNA sequences (e.g., AAGGCCUU...), with no extra information.\n\n**Note:**\n\nGiven the limited time and available resources (CPU i5-6500, 24 GB RAM, no GPU), I’m aware that each step can be significantly improved.\n\n---\n\n**Process overview:**\n\nBriefly, the process consisted of the following steps:\n\n1. Data acquisition, unification, and selection.  \n2. TM-score calculation between each sequence and all others, grouped by sequence length to reduce computational cost.  \n3. Generation of a distance matrix, clustering into n groups (classes), and selection of the most representative sequences in each group.  \n4. Training of a model (Keras) to classify sequences into 5 groups based on the input sequence.  \n5. Model inference. For each input sequence, the model predicts the most likely groups (top 5), and returns the 3D sequence of the most similar entry within each predicted group (based on sequence length and best TM-score within that group).\n\n---\n\n**Additional note:**\n\nStep 2 is computationally expensive, especially for long sequences. I considered optimizing this step by incorporating PDB data or using data augmentation techniques, but discarded those options due to hardware limitations. These ideas could be implemented as future extensions.\n\n---\n\nIt’s been a pleasure to take part in this competition — I’ve learned a lot and really enjoyed the process. Many thanks to the organizers and the Kaggle community for making this challenge possible. Best of luck to the other participants!\n\n\n\n---\n\n## Versión en español\n\n### 1. Descripción de mi solución para la competición\n\nComparto los pasos que seguí para desarrollar mi solución (score: 0.484, 6.º puesto al cierre de la fase de envíos). Este es un resumen inicial que trataré de ir ampliando en futuras actualizaciones.\n\n**Consideraciones:**\n\n- No se utilizó ningún modelo externo.  \n- No se emplearon datos adicionales más allá de los proporcionados para la fase de entrenamiento.  \n- La solución se basa únicamente en las secuencias de ARN de entrada (ej. AAGGCCUU...), sin información extra.\n\n**Nota:**\n\nDado el tiempo limitado y los recursos disponibles (CPU i5-6500, 24 GB de RAM, sin GPU), soy consciente de que cada uno de los pasos es claramente mejorable.\n\n---\n\n**Proceso seguido:**\n\nDe forma resumida, el proceso consistió en los siguientes pasos:\n\n1. Obtención, unificación y selección de datos.  \n2. Cálculo del TM-score entre cada secuencia y todas las demás, agrupando por tamaño de secuencia para tratar de reducir el coste computacional.  \n3. Generación de la matriz de distancias, *clustering* (división en *n* grupos o clases) y selección de las secuencias más representativas de cada grupo.  \n4. Entrenamiento de un modelo (Keras) para clasificar las secuencias en 5 grupos a partir de la secuencia de entrada.  \n5. Ejecución del modelo. Para cada nueva secuencia, se predice su grupo (los 5 mejores) y se devuelve la secuencia 3D de la entrada más similar dentro de ese grupo (según el tamaño de secuencia y el mejor TM-score dentro de ese rango), para cada uno de los 5 mejores grupos predichos.\n\n---\n\n**Nota adicional:**\n\nEl paso 2 es computacionalmente muy costoso, especialmente para secuencias largas. Se exploró la posibilidad de optimizar este paso incorporando datos de la base PDB o mediante *data augmentation*, pero se descartó debido a las limitaciones de hardware. Estas mejoras podrían implementarse como una extensión futura.\n\n---\n\nHa sido un placer participar en esta competición, he aprendido mucho y he disfrutado del proceso. Agradezco a los organizadores y a la comunidad de Kaggle por hacer posible este reto. ¡Mucha suerte al resto de participantes!\n",
      "votes": 23
    },
    {
      "id": 3277689,
      "postDate": "2025-08-28T13:58:23.463Z",
      "content": "<p>Amazing! you actually didn't use fine-tuning😂</p>",
      "rawMarkdown": "Amazing! you actually didn't use fine-tuning😂",
      "votes": 1
    },
    {
      "id": 3214721,
      "postDate": "2025-06-01T02:05:39.913Z",
      "content": "<p>I can hardly believe that unsupervised methods can achieve such great results. This approach is simply amazing, and I've learned a lot from it. I always assumed that almost all top-ranked participants used fine-tuning or distillation of SOTAs like protenix or drfold.</p>\n<p>However, I'm not sure if unsupervised methods can still perform this well if the private test dataset, previous training dataset, or CASP16 sequences differ too much. How about your CV? like kaggle training datasets? </p>\n<p>I'm really looking forward to it in september.</p>",
      "rawMarkdown": "I can hardly believe that unsupervised methods can achieve such great results. This approach is simply amazing, and I've learned a lot from it. I always assumed that almost all top-ranked participants used fine-tuning or distillation of SOTAs like protenix or drfold.\n\nHowever, I'm not sure if unsupervised methods can still perform this well if the private test dataset, previous training dataset, or CASP16 sequences differ too much. How about your CV? like kaggle training datasets? \n\nI'm really looking forward to it in september.",
      "votes": 1,
      "replies": [
        {
          "id": 3214957,
          "postDate": "2025-06-01T10:37:31.773Z",
          "content": "<p>Only the training data was used (train_sequences, train_sequencesv2, validation_sequences, and labels), followed by a 20% train-test split to train the model.</p>\n<p>We’ll see how it goes in September. I’d also love to hear about other participants’ approaches — hopefully some of you will share them here!</p>",
          "rawMarkdown": "Only the training data was used (train_sequences, train_sequencesv2, validation_sequences, and labels), followed by a 20% train-test split to train the model.\n\nWe’ll see how it goes in September. I’d also love to hear about other participants’ approaches — hopefully some of you will share them here!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3214485,
      "postDate": "2025-05-31T16:22:59.957Z",
      "content": "<p>If I understood correctly is you approach fully unsupervised? You selected the most similar candidate from the pool of training sequences based on the characteristics you mentioned? What if there is no sequence that has the same length as the input sequence among the set of training sequences? Thanks for sharing :)</p>",
      "rawMarkdown": "If I understood correctly is you approach fully unsupervised? You selected the most similar candidate from the pool of training sequences based on the characteristics you mentioned? What if there is no sequence that has the same length as the input sequence among the set of training sequences? Thanks for sharing :)",
      "votes": 1,
      "replies": [
        {
          "id": 3214496,
          "postDate": "2025-05-31T16:51:41.563Z",
          "content": "<p>Once the distance matrix is generated and the clusters are formed using k-medoids:</p>\n<pre><code> sklearn_extra.cluster  Kmedoids\n\nX=tm_matrix\n\nkmedoids = KMedoids(n_clusters=n_clusters, metric=, init=  ,random_state=,method= ) \nkmedoids.fit(X) \n</code></pre>\n<p>The network is simply a classification network (softmax activation). From there, the five best clusters are selected.</p>\n<p>If there is no sequence with the exact length, it is simply padded with zeros or pruned (I understand that this is clearly something that could be improved, but the truth is that I tried several methods and this gave me better results…)</p>\n<pre><code> ():   \n    m = (points)\n     m == n:\n         points  \n\n     m &gt; n: \n        new_points= points[:n]\n    :  \n        new_points = np.concatenate([points, np.zeros((n-m,))])\n     new_points\n</code></pre>",
          "rawMarkdown": "Once the distance matrix is generated and the clusters are formed using k-medoids:\n\n```python\nfrom sklearn_extra.cluster import Kmedoids\n\nX=tm_matrix\n\nkmedoids = KMedoids(n_clusters=n_clusters, metric='precomputed', init= 'k-medoids++' ,random_state=42,method= 'pam') \nkmedoids.fit(X) \n```\nThe network is simply a classification network (softmax activation). From there, the five best clusters are selected.\n\nIf there is no sequence with the exact length, it is simply padded with zeros or pruned (I understand that this is clearly something that could be improved, but the truth is that I tried several methods and this gave me better results…)\n\n```python\ndef transform_3d_structure(points, n):   \n    m = len(points)\n    if m == n:\n        return points  \n        \n    if m > n: \n        new_points= points[:n]\n    else:  \n        new_points = np.concatenate([points, np.zeros((n-m,3))])\n    return new_points\n```\n\n",
          "votes": 1,
          "replies": [
            {
              "id": 3214541,
              "postDate": "2025-05-31T18:24:01.227Z",
              "content": "<p>I am surprised your method worked so well. This might indicate that the training samples come from a pretty similar distribution to the test set.</p>",
              "rawMarkdown": "I am surprised your method worked so well. This might indicate that the training samples come from a pretty similar distribution to the test set."
            },
            {
              "id": 3214569,
              "postDate": "2025-05-31T19:46:16.830Z",
              "content": "<p>Yes, although it could also indicate that, in general, samples in nature are always distributed forming a series or set of similar patterns. I take it for granted that this must be the case, and it is the hypothesis I have been working with. I guess we will check it shortly when it is validated with new data. Greetings.</p>",
              "rawMarkdown": "Yes, although it could also indicate that, in general, samples in nature are always distributed forming a series or set of similar patterns. I take it for granted that this must be the case, and it is the hypothesis I have been working with. I guess we will check it shortly when it is validated with new data. Greetings.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 3214040,
      "postDate": "2025-05-30T20:34:36.787Z",
      "content": "<p>I think this is an excellent result achieved with limited resources. It offered a fresh perspective and was truly insightful. Excellent!</p>",
      "rawMarkdown": "I think this is an excellent result achieved with limited resources. It offered a fresh perspective and was truly insightful. Excellent!\n",
      "votes": 1
    },
    {
      "id": 3272261,
      "postDate": "2025-08-20T16:00:51.350Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3277689,
      "author_name": "DaoHe Liu",
      "author_url": "",
      "post_date": "2025-08-28T13:58:23.463000",
      "content": "<p>Amazing! you actually didn't use fine-tuning😂</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3214721,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2025-06-01T02:05:39.913000",
      "content": "<p>I can hardly believe that unsupervised methods can achieve such great results. This approach is simply amazing, and I've learned a lot from it. I always assumed that almost all top-ranked participants used fine-tuning or distillation of SOTAs like protenix or drfold.</p>\n<p>However, I'm not sure if unsupervised methods can still perform this well if the private test dataset, previous training dataset, or CASP16 sequences differ too much. How about your CV? like kaggle training datasets? </p>\n<p>I'm really looking forward to it in september.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3214957,
          "author_name": "Rem1210",
          "author_url": "",
          "post_date": "2025-06-01T10:37:31.773000",
          "content": "<p>Only the training data was used (train_sequences, train_sequencesv2, validation_sequences, and labels), followed by a 20% train-test split to train the model.</p>\n<p>We’ll see how it goes in September. I’d also love to hear about other participants’ approaches — hopefully some of you will share them here!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3214485,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2025-05-31T16:22:59.957000",
      "content": "<p>If I understood correctly is you approach fully unsupervised? You selected the most similar candidate from the pool of training sequences based on the characteristics you mentioned? What if there is no sequence that has the same length as the input sequence among the set of training sequences? Thanks for sharing :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3214496,
          "author_name": "Rem1210",
          "author_url": "",
          "post_date": "2025-05-31T16:51:41.563000",
          "content": "<p>Once the distance matrix is generated and the clusters are formed using k-medoids:</p>\n<pre><code> sklearn_extra.cluster  Kmedoids\n\nX=tm_matrix\n\nkmedoids = KMedoids(n_clusters=n_clusters, metric=, init=  ,random_state=,method= ) \nkmedoids.fit(X) \n</code></pre>\n<p>The network is simply a classification network (softmax activation). From there, the five best clusters are selected.</p>\n<p>If there is no sequence with the exact length, it is simply padded with zeros or pruned (I understand that this is clearly something that could be improved, but the truth is that I tried several methods and this gave me better results…)</p>\n<pre><code> ():   \n    m = (points)\n     m == n:\n         points  \n\n     m &gt; n: \n        new_points= points[:n]\n    :  \n        new_points = np.concatenate([points, np.zeros((n-m,))])\n     new_points\n</code></pre>",
          "votes": 1,
          "replies": [
            {
              "id": 3214541,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2025-05-31T18:24:01.227000",
              "content": "<p>I am surprised your method worked so well. This might indicate that the training samples come from a pretty similar distribution to the test set.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3214569,
              "author_name": "Rem1210",
              "author_url": "",
              "post_date": "2025-05-31T19:46:16.830000",
              "content": "<p>Yes, although it could also indicate that, in general, samples in nature are always distributed forming a series or set of similar patterns. I take it for granted that this must be the case, and it is the hypothesis I have been working with. I guess we will check it shortly when it is validated with new data. Greetings.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3214040,
      "author_name": "Shun Kuraishi",
      "author_url": "",
      "post_date": "2025-05-30T20:34:36.787000",
      "content": "<p>I think this is an excellent result achieved with limited resources. It offered a fresh perspective and was truly insightful. Excellent!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3272261,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-08-20T16:00:51.350000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3213996": "I’m sharing the steps I followed to develop my solution (score: 0.484, 6th place at the end of the submission phase). This is an initial summary that I plan to expand in future updates.\n\n---\n\n**Considerations:**\n\n- No external models were used.  \n- No additional data was used beyond what was provided for the training phase.  \n- The solution is based solely on the input RNA sequences (e.g., AAGGCCUU...), with no extra information.\n\n**Note:**\n\nGiven the limited time and available resources (CPU i5-6500, 24 GB RAM, no GPU), I’m aware that each step can be significantly improved.\n\n---\n\n**Process overview:**\n\nBriefly, the process consisted of the following steps:\n\n1. Data acquisition, unification, and selection.  \n2. TM-score calculation between each sequence and all others, grouped by sequence length to reduce computational cost.  \n3. Generation of a distance matrix, clustering into n groups (classes), and selection of the most representative sequences in each group.  \n4. Training of a model (Keras) to classify sequences into 5 groups based on the input sequence.  \n5. Model inference. For each input sequence, the model predicts the most likely groups (top 5), and returns the 3D sequence of the most similar entry within each predicted group (based on sequence length and best TM-score within that group).\n\n---\n\n**Additional note:**\n\nStep 2 is computationally expensive, especially for long sequences. I considered optimizing this step by incorporating PDB data or using data augmentation techniques, but discarded those options due to hardware limitations. These ideas could be implemented as future extensions.\n\n---\n\nIt’s been a pleasure to take part in this competition — I’ve learned a lot and really enjoyed the process. Many thanks to the organizers and the Kaggle community for making this challenge possible. Best of luck to the other participants!\n\n\n\n---\n\n## Versión en español\n\n### 1. Descripción de mi solución para la competición\n\nComparto los pasos que seguí para desarrollar mi solución (score: 0.484, 6.º puesto al cierre de la fase de envíos). Este es un resumen inicial que trataré de ir ampliando en futuras actualizaciones.\n\n**Consideraciones:**\n\n- No se utilizó ningún modelo externo.  \n- No se emplearon datos adicionales más allá de los proporcionados para la fase de entrenamiento.  \n- La solución se basa únicamente en las secuencias de ARN de entrada (ej. AAGGCCUU...), sin información extra.\n\n**Nota:**\n\nDado el tiempo limitado y los recursos disponibles (CPU i5-6500, 24 GB de RAM, sin GPU), soy consciente de que cada uno de los pasos es claramente mejorable.\n\n---\n\n**Proceso seguido:**\n\nDe forma resumida, el proceso consistió en los siguientes pasos:\n\n1. Obtención, unificación y selección de datos.  \n2. Cálculo del TM-score entre cada secuencia y todas las demás, agrupando por tamaño de secuencia para tratar de reducir el coste computacional.  \n3. Generación de la matriz de distancias, *clustering* (división en *n* grupos o clases) y selección de las secuencias más representativas de cada grupo.  \n4. Entrenamiento de un modelo (Keras) para clasificar las secuencias en 5 grupos a partir de la secuencia de entrada.  \n5. Ejecución del modelo. Para cada nueva secuencia, se predice su grupo (los 5 mejores) y se devuelve la secuencia 3D de la entrada más similar dentro de ese grupo (según el tamaño de secuencia y el mejor TM-score dentro de ese rango), para cada uno de los 5 mejores grupos predichos.\n\n---\n\n**Nota adicional:**\n\nEl paso 2 es computacionalmente muy costoso, especialmente para secuencias largas. Se exploró la posibilidad de optimizar este paso incorporando datos de la base PDB o mediante *data augmentation*, pero se descartó debido a las limitaciones de hardware. Estas mejoras podrían implementarse como una extensión futura.\n\n---\n\nHa sido un placer participar en esta competición, he aprendido mucho y he disfrutado del proceso. Agradezco a los organizadores y a la comunidad de Kaggle por hacer posible este reto. ¡Mucha suerte al resto de participantes!\n",
    "3277689": "Amazing! you actually didn't use fine-tuning😂",
    "3214721": "I can hardly believe that unsupervised methods can achieve such great results. This approach is simply amazing, and I've learned a lot from it. I always assumed that almost all top-ranked participants used fine-tuning or distillation of SOTAs like protenix or drfold.\n\nHowever, I'm not sure if unsupervised methods can still perform this well if the private test dataset, previous training dataset, or CASP16 sequences differ too much. How about your CV? like kaggle training datasets? \n\nI'm really looking forward to it in september.",
    "3214485": "If I understood correctly is you approach fully unsupervised? You selected the most similar candidate from the pool of training sequences based on the characteristics you mentioned? What if there is no sequence that has the same length as the input sequence among the set of training sequences? Thanks for sharing :)",
    "3214040": "I think this is an excellent result achieved with limited resources. It offered a fresh perspective and was truly insightful. Excellent!\n",
    "3272261": ""
  }
}