{
  "id": 647280,
  "title": "Task 2 importance meaning",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/647280",
  "author_name": "",
  "post_date": "2025-12-01T00:31:10.676186600Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>For task 2 where we need to sort out the top sequence linked to the disease, is it the sequence that is linked to the positivity of a disease (E.g: If this sequence exists in the repertoire, this person is sick), or the sequence that is linked to the health of the person (E.g: if this sequence exists in the repertoire, the person is healthy), or the one that is decisive for both ?</p>\n<p>For example, when using logistic regression, positive coefficients means that the sequence is linked highly to the disease positivity. But negative coefficients means that the sequence is linked highly to the healthyness. However, if we only look at the absolute value of the coefficients, a sequence is decisive for the model to decide whether a person is sick or healhy but is not linked to the disease or health of the person. For example, on lightgbm or xgboost's feature importance, a sequence that has a high score on feature importance means that the machine relies on that feature a lot but the value itself does not tells us whether it's linked to the positivity or the negativity of the label.</p>\n<p>My question is, for task 2, which one is need to be predicted ?\nA. The order of sequence linked to the positivity of the disease (On logistic regression, the coefficient with highest positive values)\nB. The order of sequence linked to the negativity of the disease (On logistic regression, the coefficient with highest negative values)\nC. The order of sequence that is used by the ML to predict (On logistic regression, the coefficient with highest absolute values, or in Lightgbm or xgboost, the feature importance)</p>\n<p>The description for task 2 are a bit ambiguous and the baseline itself uses the absolute value and the signed value. Clarifications about task 2 would be very much helpful. Thank you</p>",
  "messages": [
    {
      "id": "3356468",
      "postDate": "12/01/2025 00:31:10",
      "content": "<p>For task 2 where we need to sort out the top sequence linked to the disease, is it the sequence that is linked to the positivity of a disease (E.g: If this sequence exists in the repertoire, this person is sick), or the sequence that is linked to the health of the person (E.g: if this sequence exists in the repertoire, the person is healthy), or the one that is decisive for both ?</p>\n<p>For example, when using logistic regression, positive coefficients means that the sequence is linked highly to the disease positivity. But negative coefficients means that the sequence is linked highly to the healthyness. However, if we only look at the absolute value of the coefficients, a sequence is decisive for the model to decide whether a person is sick or healhy but is not linked to the disease or health of the person. For example, on lightgbm or xgboost's feature importance, a sequence that has a high score on feature importance means that the machine relies on that feature a lot but the value itself does not tells us whether it's linked to the positivity or the negativity of the label.</p>\n<p>My question is, for task 2, which one is need to be predicted ?\nA. The order of sequence linked to the positivity of the disease (On logistic regression, the coefficient with highest positive values)\nB. The order of sequence linked to the negativity of the disease (On logistic regression, the coefficient with highest negative values)\nC. The order of sequence that is used by the ML to predict (On logistic regression, the coefficient with highest absolute values, or in Lightgbm or xgboost, the feature importance)</p>\n<p>The description for task 2 are a bit ambiguous and the baseline itself uses the absolute value and the signed value. Clarifications about task 2 would be very much helpful. Thank you</p>",
      "rawMarkdown": "For task 2 where we need to sort out the top sequence linked to the disease, is it the sequence that is linked to the positivity of a disease (E.g: If this sequence exists in the repertoire, this person is sick), or the sequence that is linked to the health of the person (E.g: if this sequence exists in the repertoire, the person is healthy), or the one that is decisive for both ?\n\nFor example, when using logistic regression, positive coefficients means that the sequence is linked highly to the disease positivity. But negative coefficients means that the sequence is linked highly to the healthyness. However, if we only look at the absolute value of the coefficients, a sequence is decisive for the model to decide whether a person is sick or healhy but is not linked to the disease or health of the person. For example, on lightgbm or xgboost's feature importance, a sequence that has a high score on feature importance means that the machine relies on that feature a lot but the value itself does not tells us whether it's linked to the positivity or the negativity of the label.\n\nMy question is, for task 2, which one is need to be predicted ?\nA. The order of sequence linked to the positivity of the disease (On logistic regression, the coefficient with highest positive values)\nB. The order of sequence linked to the negativity of the disease (On logistic regression, the coefficient with highest negative values)\nC. The order of sequence that is used by the ML to predict (On logistic regression, the coefficient with highest absolute values, or in Lightgbm or xgboost, the feature importance)\n\nThe description for task 2 are a bit ambiguous and the baseline itself uses the absolute value and the signed value. Clarifications about task 2 would be very much helpful. Thank you",
      "votes": null
    },
    {
      "id": "3357852",
      "postDate": "12/01/2025 12:07:19",
      "content": "<p>Good question, and great that this can also clarify to other participants if it was ambiguous to some of them. </p>\n<p>For Task-2, the ranked sequences that are associated with the positive label (\"positivity of disease\" as you say) are of interest. The <a href=\"https://www.kaggle.com/code/ckanduri/example-baseline-predictor-using-code-template\" target=\"_blank\">example baseline provided</a> mainly uses the learnt coefficients (from a model trained to predict label being positive) to score each sequence (<code>scores.append(np.dot(counts, coefficients))</code> in <code>score_all_sequences method</code>). You are probably referring to the  <code>get_feature_importance</code> method, which we ended up not using anywhere 🤔. </p>",
      "rawMarkdown": "Good question, and great that this can also clarify to other participants if it was ambiguous to some of them. \n\nFor Task-2, the ranked sequences that are associated with the positive label (\"positivity of disease\" as you say) are of interest. The [example baseline provided](https://www.kaggle.com/code/ckanduri/example-baseline-predictor-using-code-template) mainly uses the learnt coefficients (from a model trained to predict label being positive) to score each sequence (`scores.append(np.dot(counts, coefficients))` in `score_all_sequences method`). You are probably referring to the  `get_feature_importance` method, which we ended up not using anywhere 🤔.",
      "votes": null
    },
    {
      "id": "3360378",
      "postDate": "12/02/2025 05:45:09",
      "content": "<p>Thank you for the clarification, I must have misread the notebook, sorry for that. There are also some biology related questions I want to ask that is also related to the baseline notebook. In the baseline, it only uses score from junction_aa. I thought that even the same junction_aa could have different properties when paired with different v_call and j_call. However, since the baseline only used junction_aa, does this imply that my hypothesis is wrong ?</p>",
      "rawMarkdown": "Thank you for the clarification, I must have misread the notebook, sorry for that. There are also some biology related questions I want to ask that is also related to the baseline notebook. In the baseline, it only uses score from junction_aa. I thought that even the same junction_aa could have different properties when paired with different v_call and j_call. However, since the baseline only used junction_aa, does this imply that my hypothesis is wrong ?",
      "votes": null
    },
    {
      "id": "3360444",
      "postDate": "12/02/2025 07:17:24",
      "content": "<p>Good observation and we agree with your thinking. Since the baseline is intended to serve only as a simple example and inspiration, we did not consider the combination of v_call and j_call there. We used another published method as baseline (that you can see on leaderboard and from the <a href=\"https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\" target=\"_blank\">pre-registered protocol</a> linked from competition overview page)  that considers sequences in combination with v_call and j_call. </p>",
      "rawMarkdown": "Good observation and we agree with your thinking. Since the baseline is intended to serve only as a simple example and inspiration, we did not consider the combination of v_call and j_call there. We used another published method as baseline (that you can see on leaderboard and from the [pre-registered protocol](https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf) linked from competition overview page)  that considers sequences in combination with v_call and j_call.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3357852,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "12/01/2025 12:07:19",
      "content": "<p>Good question, and great that this can also clarify to other participants if it was ambiguous to some of them. </p>\n<p>For Task-2, the ranked sequences that are associated with the positive label (\"positivity of disease\" as you say) are of interest. The <a href=\"https://www.kaggle.com/code/ckanduri/example-baseline-predictor-using-code-template\" target=\"_blank\">example baseline provided</a> mainly uses the learnt coefficients (from a model trained to predict label being positive) to score each sequence (<code>scores.append(np.dot(counts, coefficients))</code> in <code>score_all_sequences method</code>). You are probably referring to the  <code>get_feature_importance</code> method, which we ended up not using anywhere 🤔. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3360378,
          "author_name": "darrenamadeusmartin",
          "author_url": "",
          "post_date": "12/02/2025 05:45:09",
          "content": "<p>Thank you for the clarification, I must have misread the notebook, sorry for that. There are also some biology related questions I want to ask that is also related to the baseline notebook. In the baseline, it only uses score from junction_aa. I thought that even the same junction_aa could have different properties when paired with different v_call and j_call. However, since the baseline only used junction_aa, does this imply that my hypothesis is wrong ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3360444,
              "author_name": "ckanduri",
              "author_url": "",
              "post_date": "12/02/2025 07:17:24",
              "content": "<p>Good observation and we agree with your thinking. Since the baseline is intended to serve only as a simple example and inspiration, we did not consider the combination of v_call and j_call there. We used another published method as baseline (that you can see on leaderboard and from the <a href=\"https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\" target=\"_blank\">pre-registered protocol</a> linked from competition overview page)  that considers sequences in combination with v_call and j_call. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3356468": "For task 2 where we need to sort out the top sequence linked to the disease, is it the sequence that is linked to the positivity of a disease (E.g: If this sequence exists in the repertoire, this person is sick), or the sequence that is linked to the health of the person (E.g: if this sequence exists in the repertoire, the person is healthy), or the one that is decisive for both ?\n\nFor example, when using logistic regression, positive coefficients means that the sequence is linked highly to the disease positivity. But negative coefficients means that the sequence is linked highly to the healthyness. However, if we only look at the absolute value of the coefficients, a sequence is decisive for the model to decide whether a person is sick or healhy but is not linked to the disease or health of the person. For example, on lightgbm or xgboost's feature importance, a sequence that has a high score on feature importance means that the machine relies on that feature a lot but the value itself does not tells us whether it's linked to the positivity or the negativity of the label.\n\nMy question is, for task 2, which one is need to be predicted ?\nA. The order of sequence linked to the positivity of the disease (On logistic regression, the coefficient with highest positive values)\nB. The order of sequence linked to the negativity of the disease (On logistic regression, the coefficient with highest negative values)\nC. The order of sequence that is used by the ML to predict (On logistic regression, the coefficient with highest absolute values, or in Lightgbm or xgboost, the feature importance)\n\nThe description for task 2 are a bit ambiguous and the baseline itself uses the absolute value and the signed value. Clarifications about task 2 would be very much helpful. Thank you",
    "3357852": "Good question, and great that this can also clarify to other participants if it was ambiguous to some of them. \n\nFor Task-2, the ranked sequences that are associated with the positive label (\"positivity of disease\" as you say) are of interest. The [example baseline provided](https://www.kaggle.com/code/ckanduri/example-baseline-predictor-using-code-template) mainly uses the learnt coefficients (from a model trained to predict label being positive) to score each sequence (`scores.append(np.dot(counts, coefficients))` in `score_all_sequences method`). You are probably referring to the  `get_feature_importance` method, which we ended up not using anywhere 🤔.",
    "3360378": "Thank you for the clarification, I must have misread the notebook, sorry for that. There are also some biology related questions I want to ask that is also related to the baseline notebook. In the baseline, it only uses score from junction_aa. I thought that even the same junction_aa could have different properties when paired with different v_call and j_call. However, since the baseline only used junction_aa, does this imply that my hypothesis is wrong ?",
    "3360444": "Good observation and we agree with your thinking. Since the baseline is intended to serve only as a simple example and inspiration, we did not consider the combination of v_call and j_call there. We used another published method as baseline (that you can see on leaderboard and from the [pre-registered protocol](https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf) linked from competition overview page)  that considers sequences in combination with v_call and j_call."
  },
  "source": "meta"
}