{
  "id": 516092,
  "title": "[Beware overfitting]Focus on more generalization capabilities",
  "url": "/competitions/leash-BELKA/discussion/516092",
  "author_name": "",
  "post_date": "2024-07-01T09:33:59.188452100Z",
  "votes": 26,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We found the following:</p>\n<ol>\n<li>Sequence models work better than graph models in no-share</li>\n<li>3D graph models are not often better than 2D graph models</li>\n<li>Equivariant GNN seem to have the best generalization</li>\n<li>And when using augment with scaffold, it seems better than other augment…</li>\n</ol>\n<p>Our current best lb (0.480) is produced by Pretrained-Bert + 2D-GIN + PaiNN; but the best offline score (lb-0.447) is produced by 5 completely different models; the correlation between them is only 0.68; weighting them will get a lower lb (lb-0.473) score.</p>\n<p>We believe this is a dangerous sign of overfitting….<br>\nCome on, it's time to sharing cv &amp; lb</p>\n<p>It's our sub-modeling result, all use 5folds-skf, shuffle=True, seed=0.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>cv</th>\n<th>lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mamba</td>\n<td>0.651</td>\n<td>0.412</td>\n</tr>\n<tr>\n<td>gin2d</td>\n<td>0.649</td>\n<td>0.434</td>\n</tr>\n<tr>\n<td>PaiNN</td>\n<td>0.597</td>\n<td>0.443</td>\n</tr>\n<tr>\n<td>PreBert1</td>\n<td>0.612</td>\n<td>0.372</td>\n</tr>\n<tr>\n<td>PreBert2</td>\n<td>0.643</td>\n<td>0.418</td>\n</tr>\n<tr>\n<td>SGMP</td>\n<td>0.661</td>\n<td>0.396</td>\n</tr>\n<tr>\n<td>GraphFormer</td>\n<td>0.613</td>\n<td>0.422</td>\n</tr>\n<tr>\n<td>EquiformerV2</td>\n<td>0.583</td>\n<td>0.353</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2898693",
      "postDate": "07/01/2024 09:33:59",
      "content": "<p>We found the following:</p>\n<ol>\n<li>Sequence models work better than graph models in no-share</li>\n<li>3D graph models are not often better than 2D graph models</li>\n<li>Equivariant GNN seem to have the best generalization</li>\n<li>And when using augment with scaffold, it seems better than other augment…</li>\n</ol>\n<p>Our current best lb (0.480) is produced by Pretrained-Bert + 2D-GIN + PaiNN; but the best offline score (lb-0.447) is produced by 5 completely different models; the correlation between them is only 0.68; weighting them will get a lower lb (lb-0.473) score.</p>\n<p>We believe this is a dangerous sign of overfitting….<br>\nCome on, it's time to sharing cv &amp; lb</p>\n<p>It's our sub-modeling result, all use 5folds-skf, shuffle=True, seed=0.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>cv</th>\n<th>lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mamba</td>\n<td>0.651</td>\n<td>0.412</td>\n</tr>\n<tr>\n<td>gin2d</td>\n<td>0.649</td>\n<td>0.434</td>\n</tr>\n<tr>\n<td>PaiNN</td>\n<td>0.597</td>\n<td>0.443</td>\n</tr>\n<tr>\n<td>PreBert1</td>\n<td>0.612</td>\n<td>0.372</td>\n</tr>\n<tr>\n<td>PreBert2</td>\n<td>0.643</td>\n<td>0.418</td>\n</tr>\n<tr>\n<td>SGMP</td>\n<td>0.661</td>\n<td>0.396</td>\n</tr>\n<tr>\n<td>GraphFormer</td>\n<td>0.613</td>\n<td>0.422</td>\n</tr>\n<tr>\n<td>EquiformerV2</td>\n<td>0.583</td>\n<td>0.353</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "We found the following:\n1. Sequence models work better than graph models in no-share\n2. 3D graph models are not often better than 2D graph models\n3. Equivariant GNN seem to have the best generalization\n4. And when using augment with scaffold, it seems better than other augment...\n\nOur current best lb (0.480) is produced by Pretrained-Bert + 2D-GIN + PaiNN; but the best offline score (lb-0.447) is produced by 5 completely different models; the correlation between them is only 0.68; weighting them will get a lower lb (lb-0.473) score.\n\nWe believe this is a dangerous sign of overfitting....\nCome on, it's time to sharing cv & lb\n\nIt's our sub-modeling result, all use 5folds-skf, shuffle=True, seed=0.\n\n| model | cv | lb |\n|-----|-----|-----|\n| mamba | 0.651 | 0.412 |\n| gin2d | 0.649 | 0.434 |\n| PaiNN| 0.597 | 0.443 |\n| PreBert1 | 0.612 | 0.372 |\n| PreBert2| 0.643 | 0.418 |\n| SGMP| 0.661 | 0.396 |\n| GraphFormer | 0.613 | 0.422 |\n| EquiformerV2| 0.583 | 0.353 |",
      "votes": null
    },
    {
      "id": "2898836",
      "postDate": "07/01/2024 11:07:18",
      "content": "<p>I developed my own way (algorithm ) to do this and it gives 75-98% accuracy with different proteins.  But it only shows 0.03 on leaderboard. What's wrong that I am doing?</p>",
      "rawMarkdown": "I developed my own way (algorithm ) to do this and it gives 75-98% accuracy with different proteins.  But it only shows 0.03 on leaderboard. What's wrong that I am doing?",
      "votes": null
    },
    {
      "id": "2898887",
      "postDate": "07/01/2024 11:35:34",
      "content": "<p>Very Informative and Useful. Thanks.</p>",
      "rawMarkdown": "Very Informative and Useful. Thanks.",
      "votes": null
    },
    {
      "id": "2898911",
      "postDate": "07/01/2024 11:58:14",
      "content": "<p>maybe you can use mAP rather than accuracy</p>",
      "rawMarkdown": "maybe you can use mAP rather than accuracy",
      "votes": null
    },
    {
      "id": "2898956",
      "postDate": "07/01/2024 12:25:25",
      "content": "<p>Are you using all the data or a subset? Many millions of negatives can look unnecessary but probably they're not.</p>",
      "rawMarkdown": "Are you using all the data or a subset? Many millions of negatives can look unnecessary but probably they're not.",
      "votes": null
    },
    {
      "id": "2911242",
      "postDate": "07/08/2024 07:01:12",
      "content": "<p>I only have 1 model.  I should have some other models to average over. </p>",
      "rawMarkdown": "I only have 1 model.  I should have some other models to average over.",
      "votes": null
    },
    {
      "id": "2913612",
      "postDate": "07/09/2024 15:17:13",
      "content": "<p>Hey, I'm really curious, who got CV (well, holdout) non-share scores for every model?</p>\n<p>Likewise, who predicted share vs non-share with different models or ensemble weights?</p>\n<p>If you did, do you think it helped? (Random variation is still probably the biggest single factor in final place)</p>\n<p>I never got around to the first as I didn't spend any time last couple months. I did do the second, though. </p>",
      "rawMarkdown": "Hey, I'm really curious, who got CV (well, holdout) non-share scores for every model?\n\nLikewise, who predicted share vs non-share with different models or ensemble weights?\n\nIf you did, do you think it helped? (Random variation is still probably the biggest single factor in final place)\n\nI never got around to the first as I didn't spend any time last couple months. I did do the second, though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2898836,
      "author_name": "laboratorycomputer",
      "author_url": "",
      "post_date": "07/01/2024 11:07:18",
      "content": "<p>I developed my own way (algorithm ) to do this and it gives 75-98% accuracy with different proteins.  But it only shows 0.03 on leaderboard. What's wrong that I am doing?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2898911,
          "author_name": "wenxuanxx",
          "author_url": "",
          "post_date": "07/01/2024 11:58:14",
          "content": "<p>maybe you can use mAP rather than accuracy</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2898956,
          "author_name": "sacuscreed",
          "author_url": "",
          "post_date": "07/01/2024 12:25:25",
          "content": "<p>Are you using all the data or a subset? Many millions of negatives can look unnecessary but probably they're not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2898887,
      "author_name": "drravikumarc",
      "author_url": "",
      "post_date": "07/01/2024 11:35:34",
      "content": "<p>Very Informative and Useful. Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2911242,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "07/08/2024 07:01:12",
      "content": "<p>I only have 1 model.  I should have some other models to average over. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2913612,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/09/2024 15:17:13",
      "content": "<p>Hey, I'm really curious, who got CV (well, holdout) non-share scores for every model?</p>\n<p>Likewise, who predicted share vs non-share with different models or ensemble weights?</p>\n<p>If you did, do you think it helped? (Random variation is still probably the biggest single factor in final place)</p>\n<p>I never got around to the first as I didn't spend any time last couple months. I did do the second, though. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2898693": "We found the following:\n1. Sequence models work better than graph models in no-share\n2. 3D graph models are not often better than 2D graph models\n3. Equivariant GNN seem to have the best generalization\n4. And when using augment with scaffold, it seems better than other augment...\n\nOur current best lb (0.480) is produced by Pretrained-Bert + 2D-GIN + PaiNN; but the best offline score (lb-0.447) is produced by 5 completely different models; the correlation between them is only 0.68; weighting them will get a lower lb (lb-0.473) score.\n\nWe believe this is a dangerous sign of overfitting....\nCome on, it's time to sharing cv & lb\n\nIt's our sub-modeling result, all use 5folds-skf, shuffle=True, seed=0.\n\n| model | cv | lb |\n|-----|-----|-----|\n| mamba | 0.651 | 0.412 |\n| gin2d | 0.649 | 0.434 |\n| PaiNN| 0.597 | 0.443 |\n| PreBert1 | 0.612 | 0.372 |\n| PreBert2| 0.643 | 0.418 |\n| SGMP| 0.661 | 0.396 |\n| GraphFormer | 0.613 | 0.422 |\n| EquiformerV2| 0.583 | 0.353 |",
    "2898836": "I developed my own way (algorithm ) to do this and it gives 75-98% accuracy with different proteins.  But it only shows 0.03 on leaderboard. What's wrong that I am doing?",
    "2898887": "Very Informative and Useful. Thanks.",
    "2898911": "maybe you can use mAP rather than accuracy",
    "2898956": "Are you using all the data or a subset? Many millions of negatives can look unnecessary but probably they're not.",
    "2911242": "I only have 1 model.  I should have some other models to average over.",
    "2913612": "Hey, I'm really curious, who got CV (well, holdout) non-share scores for every model?\n\nLikewise, who predicted share vs non-share with different models or ensemble weights?\n\nIf you did, do you think it helped? (Random variation is still probably the biggest single factor in final place)\n\nI never got around to the first as I didn't spend any time last couple months. I did do the second, though."
  },
  "source": "meta"
}