{
  "id": 413729,
  "title": "Improve lb score by using different channels than the ones used for training",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/413729",
  "author_name": "",
  "post_date": "2023-05-29T22:38:33.540326400Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have a feeling that the channels used in the validation fragment are not aligned to the ones used during training.  If I train on channels 29-31 and validate on 28-30 or 30-32 on a held out fragment, the dice score either goes down or stays roughly the same.  </p>\n<p>However, I have a model that uses 29-31 channels that when submitted to the competition produces a score of .37.  If I use 28-30, it improves the score to .51.  27-29 improves the score to .55.  26-28 improves the score to .56.  And, 25-27 improves the score to .57. </p>\n<p>A .37 to .57 gap is pretty large and I don't see these results when training on the public training set using two fragments and validating on the third.</p>\n<p>I also wonder if using different channels only works for the public test set and doesn't work for the private test set.</p>\n<p>Let me know in the comments if you have experienced something similar with your models.  Maybe this is unique to mine.  I also urge the competition hosts to investigate if there is anything wrong with the fourth fragment used to judge this competition. </p>",
  "messages": [
    {
      "id": "2280144",
      "postDate": "05/29/2023 22:38:33",
      "content": "<p>I have a feeling that the channels used in the validation fragment are not aligned to the ones used during training.  If I train on channels 29-31 and validate on 28-30 or 30-32 on a held out fragment, the dice score either goes down or stays roughly the same.  </p>\n<p>However, I have a model that uses 29-31 channels that when submitted to the competition produces a score of .37.  If I use 28-30, it improves the score to .51.  27-29 improves the score to .55.  26-28 improves the score to .56.  And, 25-27 improves the score to .57. </p>\n<p>A .37 to .57 gap is pretty large and I don't see these results when training on the public training set using two fragments and validating on the third.</p>\n<p>I also wonder if using different channels only works for the public test set and doesn't work for the private test set.</p>\n<p>Let me know in the comments if you have experienced something similar with your models.  Maybe this is unique to mine.  I also urge the competition hosts to investigate if there is anything wrong with the fourth fragment used to judge this competition. </p>",
      "rawMarkdown": "I have a feeling that the channels used in the validation fragment are not aligned to the ones used during training.  If I train on channels 29-31 and validate on 28-30 or 30-32 on a held out fragment, the dice score either goes down or stays roughly the same.  \n\nHowever, I have a model that uses 29-31 channels that when submitted to the competition produces a score of .37.  If I use 28-30, it improves the score to .51.  27-29 improves the score to .55.  26-28 improves the score to .56.  And, 25-27 improves the score to .57. \n\nA .37 to .57 gap is pretty large and I don't see these results when training on the public training set using two fragments and validating on the third.\n\nI also wonder if using different channels only works for the public test set and doesn't work for the private test set.\n\nLet me know in the comments if you have experienced something similar with your models.  Maybe this is unique to mine.  I also urge the competition hosts to investigate if there is anything wrong with the fourth fragment used to judge this competition.",
      "votes": null
    },
    {
      "id": "2280168",
      "postDate": "05/29/2023 23:23:13",
      "content": "<blockquote>\n  <p>I also wonder if using different channels only works for the public test set and doesn't work for the private test set.</p>\n</blockquote>\n<p>Private test set uses wider area of the same fragments as public leaderboard. So, I'm pretty sure in general channel selection has same effect on public and private score. There is of course always a little bit of unknown in public/private ink ratio etc.</p>",
      "rawMarkdown": "> I also wonder if using different channels only works for the public test set and doesn't work for the private test set.\n\nPrivate test set uses wider area of the same fragments as public leaderboard. So, I'm pretty sure in general channel selection has same effect on public and private score. There is of course always a little bit of unknown in public/private ink ratio etc.",
      "votes": null
    },
    {
      "id": "2280747",
      "postDate": "05/30/2023 10:28:01",
      "content": "<p>Not sure if you have looked at this post -<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348</a></p>\n<p>There is a lot in it, but a couple of things I took from it - where the pattern of ink distribution is similar for slices used the model seems to perform better.  And the best slices are not the same for all fragments in training so probably this is true also for private test.  There is also - \"The distance d between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified.\"</p>\n<p>So perhaps some analysis of the private test fragment slices to determine which could be the best to use for inference either based on a theoretical ink distribution or similarity of distribution to what was used in training fold or both. Not sure if anyone has compared results for models trained on only one of the train fragments to see if any are better than the others on private test.</p>\n<p>For public test - The sample slices available to download in the test folders are simply copied from training fragment one.  So would seem likely to perform well on models trained on fragment one. </p>",
      "rawMarkdown": "Not sure if you have looked at this post -\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348\n\nThere is a lot in it, but a couple of things I took from it - where the pattern of ink distribution is similar for slices used the model seems to perform better.  And the best slices are not the same for all fragments in training so probably this is true also for private test.  There is also - \"The distance d between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified.\"\n\nSo perhaps some analysis of the private test fragment slices to determine which could be the best to use for inference either based on a theoretical ink distribution or similarity of distribution to what was used in training fold or both. Not sure if anyone has compared results for models trained on only one of the train fragments to see if any are better than the others on private test.\n\nFor public test - The sample slices available to download in the test folders are simply copied from training fragment one.  So would seem likely to perform well on models trained on fragment one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2280168,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "05/29/2023 23:23:13",
      "content": "<blockquote>\n  <p>I also wonder if using different channels only works for the public test set and doesn't work for the private test set.</p>\n</blockquote>\n<p>Private test set uses wider area of the same fragments as public leaderboard. So, I'm pretty sure in general channel selection has same effect on public and private score. There is of course always a little bit of unknown in public/private ink ratio etc.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2280747,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "05/30/2023 10:28:01",
      "content": "<p>Not sure if you have looked at this post -<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348</a></p>\n<p>There is a lot in it, but a couple of things I took from it - where the pattern of ink distribution is similar for slices used the model seems to perform better.  And the best slices are not the same for all fragments in training so probably this is true also for private test.  There is also - \"The distance d between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified.\"</p>\n<p>So perhaps some analysis of the private test fragment slices to determine which could be the best to use for inference either based on a theoretical ink distribution or similarity of distribution to what was used in training fold or both. Not sure if anyone has compared results for models trained on only one of the train fragments to see if any are better than the others on private test.</p>\n<p>For public test - The sample slices available to download in the test folders are simply copied from training fragment one.  So would seem likely to perform well on models trained on fragment one. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2280144": "I have a feeling that the channels used in the validation fragment are not aligned to the ones used during training.  If I train on channels 29-31 and validate on 28-30 or 30-32 on a held out fragment, the dice score either goes down or stays roughly the same.  \n\nHowever, I have a model that uses 29-31 channels that when submitted to the competition produces a score of .37.  If I use 28-30, it improves the score to .51.  27-29 improves the score to .55.  26-28 improves the score to .56.  And, 25-27 improves the score to .57. \n\nA .37 to .57 gap is pretty large and I don't see these results when training on the public training set using two fragments and validating on the third.\n\nI also wonder if using different channels only works for the public test set and doesn't work for the private test set.\n\nLet me know in the comments if you have experienced something similar with your models.  Maybe this is unique to mine.  I also urge the competition hosts to investigate if there is anything wrong with the fourth fragment used to judge this competition.",
    "2280168": "> I also wonder if using different channels only works for the public test set and doesn't work for the private test set.\n\nPrivate test set uses wider area of the same fragments as public leaderboard. So, I'm pretty sure in general channel selection has same effect on public and private score. There is of course always a little bit of unknown in public/private ink ratio etc.",
    "2280747": "Not sure if you have looked at this post -\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348\n\nThere is a lot in it, but a couple of things I took from it - where the pattern of ink distribution is similar for slices used the model seems to perform better.  And the best slices are not the same for all fragments in training so probably this is true also for private test.  There is also - \"The distance d between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified.\"\n\nSo perhaps some analysis of the private test fragment slices to determine which could be the best to use for inference either based on a theoretical ink distribution or similarity of distribution to what was used in training fold or both. Not sure if anyone has compared results for models trained on only one of the train fragments to see if any are better than the others on private test.\n\nFor public test - The sample slices available to download in the test folders are simply copied from training fragment one.  So would seem likely to perform well on models trained on fragment one."
  },
  "source": "meta"
}