{
  "id": 663297,
  "title": "Lesson Learned and a Brief Solution",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/663297",
  "author_name": "",
  "post_date": "2025-12-17T08:19:25.013725200Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Huge thanks to the organizers for the competition! It was a great excuse to ask around my girlfriend with lots of questions about TCRs (this is her field, and we competed together). I really enjoyed the challenge. I think I would’ve enjoyed it even more if the competition had lasted a few months longer. There just wasn’t enough time to try many ideas that sounded fun, like MIL or DL approaches for full repertoires, so I ended up going back to more classical ML stuff.</p>\n<p>I also learned a painful lesson. I messed up the final submission assembly. I couldn’t properly merge everything into the final submission. Polars is awesome most of the time, but once you need indexing, it can get tricky and it really likes to lose sorting after each filtering, etc. In the end, my last submission had unsorted TCRs, and I realized I just wouldn’t have time to fix the submission pipeline before the deadline. Good reminder for the future: don’t leave critical things for the very last moment.</p>\n<p>I finished in a painful spot on the leaderboard, 11th place. If I had merged everything in time, I’d probably be a few places higher. The late submission is visible on the screenshot. Still, congrats to the team that took 10th place!</p>\n<p>Very briefly about the solution. I mostly avoided k-mers, except in datasets where motifs were explicitly simulated. That was a conscious choice, as I wanted to better understand TCRs, clonal structure, and the differences between antiviral and autoimmune responses. Most of my approach was about clustering TCRs, computing statistics, selecting enriched clusters /I tried tons of metics, look at the screenshot!/, and then training models on top of those features. It's interesting that there was one highly significant cluster in the first dataset, but I never tried to cluster it with dataset 8 (dataset 1 mixed in proinsulin TSR, right?)For clustering, I used embeddings from a randomly chosen model and also some Hamming-distance based approaches.</p>\n<p>I really hope the organizers will share per-dataset and per-task score breakdowns. It would be super interesting to dig into the results and compare approaches in more detail.</p>",
  "messages": [
    {
      "id": "3377955",
      "postDate": "12/17/2025 08:19:25",
      "content": "<p>Huge thanks to the organizers for the competition! It was a great excuse to ask around my girlfriend with lots of questions about TCRs (this is her field, and we competed together). I really enjoyed the challenge. I think I would’ve enjoyed it even more if the competition had lasted a few months longer. There just wasn’t enough time to try many ideas that sounded fun, like MIL or DL approaches for full repertoires, so I ended up going back to more classical ML stuff.</p>\n<p>I also learned a painful lesson. I messed up the final submission assembly. I couldn’t properly merge everything into the final submission. Polars is awesome most of the time, but once you need indexing, it can get tricky and it really likes to lose sorting after each filtering, etc. In the end, my last submission had unsorted TCRs, and I realized I just wouldn’t have time to fix the submission pipeline before the deadline. Good reminder for the future: don’t leave critical things for the very last moment.</p>\n<p>I finished in a painful spot on the leaderboard, 11th place. If I had merged everything in time, I’d probably be a few places higher. The late submission is visible on the screenshot. Still, congrats to the team that took 10th place!</p>\n<p>Very briefly about the solution. I mostly avoided k-mers, except in datasets where motifs were explicitly simulated. That was a conscious choice, as I wanted to better understand TCRs, clonal structure, and the differences between antiviral and autoimmune responses. Most of my approach was about clustering TCRs, computing statistics, selecting enriched clusters /I tried tons of metics, look at the screenshot!/, and then training models on top of those features. It's interesting that there was one highly significant cluster in the first dataset, but I never tried to cluster it with dataset 8 (dataset 1 mixed in proinsulin TSR, right?)For clustering, I used embeddings from a randomly chosen model and also some Hamming-distance based approaches.</p>\n<p>I really hope the organizers will share per-dataset and per-task score breakdowns. It would be super interesting to dig into the results and compare approaches in more detail.</p>",
      "rawMarkdown": "Huge thanks to the organizers for the competition! It was a great excuse to ask around my girlfriend with lots of questions about TCRs (this is her field, and we competed together). I really enjoyed the challenge. I think I would’ve enjoyed it even more if the competition had lasted a few months longer. There just wasn’t enough time to try many ideas that sounded fun, like MIL or DL approaches for full repertoires, so I ended up going back to more classical ML stuff.\n\nI also learned a painful lesson. I messed up the final submission assembly. I couldn’t properly merge everything into the final submission. Polars is awesome most of the time, but once you need indexing, it can get tricky and it really likes to lose sorting after each filtering, etc. In the end, my last submission had unsorted TCRs, and I realized I just wouldn’t have time to fix the submission pipeline before the deadline. Good reminder for the future: don’t leave critical things for the very last moment.\n\nI finished in a painful spot on the leaderboard, 11th place. If I had merged everything in time, I’d probably be a few places higher. The late submission is visible on the screenshot. Still, congrats to the team that took 10th place!\n\nVery briefly about the solution. I mostly avoided k-mers, except in datasets where motifs were explicitly simulated. That was a conscious choice, as I wanted to better understand TCRs, clonal structure, and the differences between antiviral and autoimmune responses. Most of my approach was about clustering TCRs, computing statistics, selecting enriched clusters /I tried tons of metics, look at the screenshot!/, and then training models on top of those features. It's interesting that there was one highly significant cluster in the first dataset, but I never tried to cluster it with dataset 8 (dataset 1 mixed in proinsulin TSR, right?)For clustering, I used embeddings from a randomly chosen model and also some Hamming-distance based approaches.\n\nI really hope the organizers will share per-dataset and per-task score breakdowns. It would be super interesting to dig into the results and compare approaches in more detail.",
      "votes": null
    },
    {
      "id": "3378012",
      "postDate": "12/17/2025 11:13:21",
      "content": "<p>Can you provide information on this package from where you get the Cooccurence metrics?</p>",
      "rawMarkdown": "Can you provide information on this package from where you get the Cooccurence metrics?",
      "votes": null
    },
    {
      "id": "3378055",
      "postDate": "12/17/2025 12:52:28",
      "content": "<p>This is just my competition repository, I asked copilot to generate a readme for  implemented score functions, because there were too many of them.</p>",
      "rawMarkdown": "This is just my competition repository, I asked copilot to generate a readme for  implemented score functions, because there were too many of them.",
      "votes": null
    },
    {
      "id": "3378093",
      "postDate": "12/17/2025 14:29:19",
      "content": "<p>My model performed worse after adding k-mers. I suspect the k-mer approach may have led to overfitting, since each dataset has only ~400 repertoires relative to the high-dimensional feature space introduced by k-mer representations.\nI really appreciate that you chose to avoid k-mers and instead focused on capturing clonal structure through clustering TCRs!</p>",
      "rawMarkdown": "My model performed worse after adding k-mers. I suspect the k-mer approach may have led to overfitting, since each dataset has only ~400 repertoires relative to the high-dimensional feature space introduced by k-mer representations.\nI really appreciate that you chose to avoid k-mers and instead focused on capturing clonal structure through clustering TCRs!",
      "votes": null
    },
    {
      "id": "3378417",
      "postDate": "12/18/2025 05:35:07",
      "content": "<p>Ooh I hope you end up open-sourcing this, I'd love to read through the code :D </p>",
      "rawMarkdown": "Ooh I hope you end up open-sourcing this, I'd love to read through the code :D",
      "votes": null
    },
    {
      "id": "3378445",
      "postDate": "12/18/2025 08:09:08",
      "content": "<p>I’d like to, but right now it’s a mess :с Tight deadlines turned my initial pipeline idea into a bunch of ugly notebooks. I’ll clean it up first. </p>",
      "rawMarkdown": "I’d like to, but right now it’s a mess :с Tight deadlines turned my initial pipeline idea into a bunch of ugly notebooks. I’ll clean it up first.",
      "votes": null
    },
    {
      "id": "3378448",
      "postDate": "12/18/2025 08:17:56",
      "content": "<p>Yes, overfitting is a real issue here. Doing CV on clusters without leakage or score inflation is even trickier than with k-mers. In an ideal world you’d build clusters on the training folds and then somehow project the test fold to them, which is quite painful both computationally and logically. I couldn't handle it with embeddings clustering, so I considered the embeddings clusters as external information and did it on the entire dataset at once (including the test). However, with Hamming clustering, it was possible to do it separately on folds.</p>",
      "rawMarkdown": "Yes, overfitting is a real issue here. Doing CV on clusters without leakage or score inflation is even trickier than with k-mers. In an ideal world you’d build clusters on the training folds and then somehow project the test fold to them, which is quite painful both computationally and logically. I couldn't handle it with embeddings clustering, so I considered the embeddings clusters as external information and did it on the entire dataset at once (including the test). However, with Hamming clustering, it was possible to do it separately on folds.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3378012,
      "author_name": "raghvendrakul",
      "author_url": "",
      "post_date": "12/17/2025 11:13:21",
      "content": "<p>Can you provide information on this package from where you get the Cooccurence metrics?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3378055,
          "author_name": "gottalottarock",
          "author_url": "",
          "post_date": "12/17/2025 12:52:28",
          "content": "<p>This is just my competition repository, I asked copilot to generate a readme for  implemented score functions, because there were too many of them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3378093,
      "author_name": "kenchanhodgkin",
      "author_url": "",
      "post_date": "12/17/2025 14:29:19",
      "content": "<p>My model performed worse after adding k-mers. I suspect the k-mer approach may have led to overfitting, since each dataset has only ~400 repertoires relative to the high-dimensional feature space introduced by k-mer representations.\nI really appreciate that you chose to avoid k-mers and instead focused on capturing clonal structure through clustering TCRs!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3378448,
          "author_name": "gottalottarock",
          "author_url": "",
          "post_date": "12/18/2025 08:17:56",
          "content": "<p>Yes, overfitting is a real issue here. Doing CV on clusters without leakage or score inflation is even trickier than with k-mers. In an ideal world you’d build clusters on the training folds and then somehow project the test fold to them, which is quite painful both computationally and logically. I couldn't handle it with embeddings clustering, so I considered the embeddings clusters as external information and did it on the entire dataset at once (including the test). However, with Hamming clustering, it was possible to do it separately on folds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3378417,
      "author_name": "sajayr",
      "author_url": "",
      "post_date": "12/18/2025 05:35:07",
      "content": "<p>Ooh I hope you end up open-sourcing this, I'd love to read through the code :D </p>",
      "votes": null,
      "replies": [
        {
          "id": 3378445,
          "author_name": "gottalottarock",
          "author_url": "",
          "post_date": "12/18/2025 08:09:08",
          "content": "<p>I’d like to, but right now it’s a mess :с Tight deadlines turned my initial pipeline idea into a bunch of ugly notebooks. I’ll clean it up first. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3377955": "Huge thanks to the organizers for the competition! It was a great excuse to ask around my girlfriend with lots of questions about TCRs (this is her field, and we competed together). I really enjoyed the challenge. I think I would’ve enjoyed it even more if the competition had lasted a few months longer. There just wasn’t enough time to try many ideas that sounded fun, like MIL or DL approaches for full repertoires, so I ended up going back to more classical ML stuff.\n\nI also learned a painful lesson. I messed up the final submission assembly. I couldn’t properly merge everything into the final submission. Polars is awesome most of the time, but once you need indexing, it can get tricky and it really likes to lose sorting after each filtering, etc. In the end, my last submission had unsorted TCRs, and I realized I just wouldn’t have time to fix the submission pipeline before the deadline. Good reminder for the future: don’t leave critical things for the very last moment.\n\nI finished in a painful spot on the leaderboard, 11th place. If I had merged everything in time, I’d probably be a few places higher. The late submission is visible on the screenshot. Still, congrats to the team that took 10th place!\n\nVery briefly about the solution. I mostly avoided k-mers, except in datasets where motifs were explicitly simulated. That was a conscious choice, as I wanted to better understand TCRs, clonal structure, and the differences between antiviral and autoimmune responses. Most of my approach was about clustering TCRs, computing statistics, selecting enriched clusters /I tried tons of metics, look at the screenshot!/, and then training models on top of those features. It's interesting that there was one highly significant cluster in the first dataset, but I never tried to cluster it with dataset 8 (dataset 1 mixed in proinsulin TSR, right?)For clustering, I used embeddings from a randomly chosen model and also some Hamming-distance based approaches.\n\nI really hope the organizers will share per-dataset and per-task score breakdowns. It would be super interesting to dig into the results and compare approaches in more detail.",
    "3378012": "Can you provide information on this package from where you get the Cooccurence metrics?",
    "3378055": "This is just my competition repository, I asked copilot to generate a readme for  implemented score functions, because there were too many of them.",
    "3378093": "My model performed worse after adding k-mers. I suspect the k-mer approach may have led to overfitting, since each dataset has only ~400 repertoires relative to the high-dimensional feature space introduced by k-mer representations.\nI really appreciate that you chose to avoid k-mers and instead focused on capturing clonal structure through clustering TCRs!",
    "3378417": "Ooh I hope you end up open-sourcing this, I'd love to read through the code :D",
    "3378445": "I’d like to, but right now it’s a mess :с Tight deadlines turned my initial pipeline idea into a bunch of ugly notebooks. I’ll clean it up first.",
    "3378448": "Yes, overfitting is a real issue here. Doing CV on clusters without leakage or score inflation is even trickier than with k-mers. In an ideal world you’d build clusters on the training folds and then somehow project the test fold to them, which is quite painful both computationally and logically. I couldn't handle it with embeddings clustering, so I considered the embeddings clusters as external information and did it on the entire dataset at once (including the test). However, with Hamming clustering, it was possible to do it separately on folds."
  },
  "source": "meta"
}