{
  "id": 409770,
  "title": "My experimental results, which channels you need?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/409770",
  "author_name": "yoyobar",
  "post_date": "2023-05-12T14:06:35.802000",
  "votes": 34,
  "comment_count": 36,
  "views": 0,
  "content": "<h1>I was testing which channels is my model prefers ,these are my experimental procedures.</h1>\n<h2>1. Train a model in several consequent channels.</h2>\n<p>-&gt;Using 12 channel between 12.tif and 38.tif to training and got mean CV:0.561</p>\n<h2>2. use val dataset to get the CVs in different channels set.</h2>\n<table>\n<thead>\n<tr>\n<th>range</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>12:24</td>\n<td>0.261201</td>\n</tr>\n<tr>\n<td>14:26</td>\n<td>0.413033</td>\n</tr>\n<tr>\n<td>16:28</td>\n<td>0.515143</td>\n</tr>\n<tr>\n<td>18:30</td>\n<td>0.564901</td>\n</tr>\n<tr>\n<td>20:32</td>\n<td>0.596330</td>\n</tr>\n<tr>\n<td>24:36</td>\n<td><strong>0.608991</strong></td>\n</tr>\n<tr>\n<td>26:38</td>\n<td>0.594269</td>\n</tr>\n</tbody>\n</table>\n<h2>3. Draw a line graph to show which channel is better.</h2>\n<p>`</p>\n<pre><code>import numpy as np\nimport matplotlib.pyplot as plt\n\nstart=12\nload=12\nnum=26\n\ndata=np.zeros(num+start)\npossibility=np.zeros(num+start)\nindex=np.arange(start+num)\nscore=[[12,0.261201],[14,0.413033],[16,0.515143],[18,0.564901],[20,0.596330],\n[22,0.615607],[24,0.608991],[26,0.594269]]\n\nfor (k,v) in score:\n    data[k:k+load]+=v\n    possibility[k:k+load]+=1\n\nplt.plot(index[start-2:],data[start-2:],label=\"channel_prefer_count\")\nplt.plot(index[start-2:],possibility[start-2:],label=\"the possibility of channel appear in training\")\nplt.plot(index[start-2:],(data[start-2:]/possibility[start-2:])*6-1,label=\"prefer:(channel/pos)*6-1\")#get mean\n\nplt.legend()\nplt.xlabel(\"channel_index\")\nplt.show()\n</code></pre>\n<p>`<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F68fcfd9d0f307261a31d4dda96a7791f%2FFigure_1.png?generation=1683898996998474&amp;alt=media\" alt=\"\"></p>\n<p>This is vary interesting ,my model prefers deeper features. Therefore, I trained another model which using 12 channels from 16.tif to 42.tif and got mean CV:0.626.</p>\n<table>\n<thead>\n<tr>\n<th>range</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>16:28</td>\n<td>0.490768</td>\n</tr>\n<tr>\n<td>18:30</td>\n<td>0.631685</td>\n</tr>\n<tr>\n<td>20:32</td>\n<td>0.616809</td>\n</tr>\n<tr>\n<td>22:34</td>\n<td><strong>0.652383</strong></td>\n</tr>\n<tr>\n<td>24:36</td>\n<td>0.648402</td>\n</tr>\n<tr>\n<td>26:38</td>\n<td>0.628398</td>\n</tr>\n<tr>\n<td>28:40</td>\n<td>0.557015</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2Fa752b5fc32c29a152ebb47ae3e4bff3e%2FFigure_1.png?generation=1683899398335747&amp;alt=media\" alt=\"\"></p>\n<p>It seems good :D</p>\n<h2>Therefore, I think most of the features are between 18 and 38 channels. :)</h2>",
  "messages": [
    {
      "id": 2256471,
      "postDate": "2023-05-12T14:06:35.803Z",
      "content": "<h1>I was testing which channels is my model prefers ,these are my experimental procedures.</h1>\n<h2>1. Train a model in several consequent channels.</h2>\n<p>-&gt;Using 12 channel between 12.tif and 38.tif to training and got mean CV:0.561</p>\n<h2>2. use val dataset to get the CVs in different channels set.</h2>\n<table>\n<thead>\n<tr>\n<th>range</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>12:24</td>\n<td>0.261201</td>\n</tr>\n<tr>\n<td>14:26</td>\n<td>0.413033</td>\n</tr>\n<tr>\n<td>16:28</td>\n<td>0.515143</td>\n</tr>\n<tr>\n<td>18:30</td>\n<td>0.564901</td>\n</tr>\n<tr>\n<td>20:32</td>\n<td>0.596330</td>\n</tr>\n<tr>\n<td>24:36</td>\n<td><strong>0.608991</strong></td>\n</tr>\n<tr>\n<td>26:38</td>\n<td>0.594269</td>\n</tr>\n</tbody>\n</table>\n<h2>3. Draw a line graph to show which channel is better.</h2>\n<p>`</p>\n<pre><code>import numpy as np\nimport matplotlib.pyplot as plt\n\nstart=12\nload=12\nnum=26\n\ndata=np.zeros(num+start)\npossibility=np.zeros(num+start)\nindex=np.arange(start+num)\nscore=[[12,0.261201],[14,0.413033],[16,0.515143],[18,0.564901],[20,0.596330],\n[22,0.615607],[24,0.608991],[26,0.594269]]\n\nfor (k,v) in score:\n    data[k:k+load]+=v\n    possibility[k:k+load]+=1\n\nplt.plot(index[start-2:],data[start-2:],label=\"channel_prefer_count\")\nplt.plot(index[start-2:],possibility[start-2:],label=\"the possibility of channel appear in training\")\nplt.plot(index[start-2:],(data[start-2:]/possibility[start-2:])*6-1,label=\"prefer:(channel/pos)*6-1\")#get mean\n\nplt.legend()\nplt.xlabel(\"channel_index\")\nplt.show()\n</code></pre>\n<p>`<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F68fcfd9d0f307261a31d4dda96a7791f%2FFigure_1.png?generation=1683898996998474&amp;alt=media\" alt=\"\"></p>\n<p>This is vary interesting ,my model prefers deeper features. Therefore, I trained another model which using 12 channels from 16.tif to 42.tif and got mean CV:0.626.</p>\n<table>\n<thead>\n<tr>\n<th>range</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>16:28</td>\n<td>0.490768</td>\n</tr>\n<tr>\n<td>18:30</td>\n<td>0.631685</td>\n</tr>\n<tr>\n<td>20:32</td>\n<td>0.616809</td>\n</tr>\n<tr>\n<td>22:34</td>\n<td><strong>0.652383</strong></td>\n</tr>\n<tr>\n<td>24:36</td>\n<td>0.648402</td>\n</tr>\n<tr>\n<td>26:38</td>\n<td>0.628398</td>\n</tr>\n<tr>\n<td>28:40</td>\n<td>0.557015</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2Fa752b5fc32c29a152ebb47ae3e4bff3e%2FFigure_1.png?generation=1683899398335747&amp;alt=media\" alt=\"\"></p>\n<p>It seems good :D</p>\n<h2>Therefore, I think most of the features are between 18 and 38 channels. :)</h2>",
      "rawMarkdown": "# I was testing which channels is my model prefers ,these are my experimental procedures.\n\n## 1. Train a model in several consequent channels.\n->Using 12 channel between 12.tif and 38.tif to training and got mean CV:0.561\n## 2. use val dataset to get the CVs in different channels set.\n| range        | CV           |   \n| ------------- |:------:| \n| 12:24  | 0.261201 |  \n| 14:26 | 0.413033 |  \n| 16:28 | 0.515143  |  \n| 18:30 | 0.564901 |  \n| 20:32 | 0.596330   |  \n|24:36 |**0.608991** |\n|26:38 | 0.594269|\n## 3. Draw a line graph to show which channel is better.\n`\n\n    import numpy as np\n    import matplotlib.pyplot as plt\n\n    start=12\n    load=12\n    num=26\n\n    data=np.zeros(num+start)\n    possibility=np.zeros(num+start)\n    index=np.arange(start+num)\n    score=[[12,0.261201],[14,0.413033],[16,0.515143],[18,0.564901],[20,0.596330],\n    [22,0.615607],[24,0.608991],[26,0.594269]]\n\n    for (k,v) in score:\n        data[k:k+load]+=v\n        possibility[k:k+load]+=1\n\n    plt.plot(index[start-2:],data[start-2:],label=\"channel_prefer_count\")\n    plt.plot(index[start-2:],possibility[start-2:],label=\"the possibility of channel appear in training\")\n    plt.plot(index[start-2:],(data[start-2:]/possibility[start-2:])*6-1,label=\"prefer:(channel/pos)*6-1\")#get mean\n\n    plt.legend()\n    plt.xlabel(\"channel_index\")\n    plt.show()\n`![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F68fcfd9d0f307261a31d4dda96a7791f%2FFigure_1.png?generation=1683898996998474&alt=media)\n\nThis is vary interesting ,my model prefers deeper features. Therefore, I trained another model which using 12 channels from 16.tif to 42.tif and got mean CV:0.626.\n\n| range| CV|   \n| ------------- |:------:| \n|16:28|0.490768|\n|18:30|0.631685|\n|20:32|0.616809|\n|22:34|**0.652383**|\n|24:36|0.648402|\n|26:38|0.628398|\n|28:40|0.557015|\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2Fa752b5fc32c29a152ebb47ae3e4bff3e%2FFigure_1.png?generation=1683899398335747&alt=media)\n\nIt seems good :D\n## Therefore, I think most of the features are between 18 and 38 channels. :)",
      "votes": 35
    },
    {
      "id": 2303738,
      "postDate": "2023-06-15T12:50:04.997Z",
      "content": "<p>Congratulations on winning the gold medal. Can you share your solution? I borrowed from your notebook, but my score did not improve afterwards, so I want to learn something from your solution. Thank you!</p>",
      "rawMarkdown": "Congratulations on winning the gold medal. Can you share your solution? I borrowed from your notebook, but my score did not improve afterwards, so I want to learn something from your solution. Thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 2303809,
          "postDate": "2023-06-15T13:37:49.213Z",
          "content": "<p>Of course! This is our <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417383\" target=\"_blank\">solution</a></p>",
          "rawMarkdown": "Of course! This is our [solution](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417383)",
          "replies": [
            {
              "id": 2303822,
              "postDate": "2023-06-15T13:48:13.113Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2256759,
      "postDate": "2023-05-12T17:34:42.423Z",
      "content": "<p>In this case does cv mean per fragment or is your validation scheme different?</p>",
      "rawMarkdown": "In this case does cv mean per fragment or is your validation scheme different?",
      "votes": 1,
      "replies": [
        {
          "id": 2256881,
          "postDate": "2023-05-12T19:53:51.230Z",
          "content": "<p>I trained my model in the same validation scheme. I used 1,2 fragments for training and 3 fragment for validation.</p>",
          "rawMarkdown": "I trained my model in the same validation scheme. I used 1,2 fragments for training and 3 fragment for validation.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2256562,
      "postDate": "2023-05-12T15:25:46.357Z",
      "content": "<p>is there a chance that the best channels are location dependent?<br>\ni.e. at some x,y location it is 22:34, while at other location it is different at 24:36</p>",
      "rawMarkdown": "is there a chance that the best channels are location dependent?\ni.e. at some x,y location it is 22:34, while at other location it is different at 24:36",
      "votes": 1,
      "replies": [
        {
          "id": 2256578,
          "postDate": "2023-05-12T15:39:24.853Z",
          "content": "<p>I think it is possible,but it is hard to prove.Of course we can add TTS here to solve this problem.</p>",
          "rawMarkdown": "I think it is possible,but it is hard to prove.Of course we can add TTS here to solve this problem.",
          "replies": [
            {
              "id": 2257498,
              "postDate": "2023-05-13T12:01:15.283Z",
              "content": "<p>check this out <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348#2235071\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348#2235071</a> if you are looking for a proof</p>",
              "rawMarkdown": "check this out https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348#2235071 if you are looking for a proof",
              "votes": 2
            }
          ]
        },
        {
          "id": 2256684,
          "postDate": "2023-05-12T16:48:54.123Z",
          "content": "<p>do you mean for one fragment some specific area have better features on  different slice range ?<br>\nif yes,  <a href=\"https://www.kaggle.com/code/henrikbrk/fragment-flattening\" target=\"_blank\">for sure</a></p>",
          "rawMarkdown": "do you mean for one fragment some specific area have better features on  different slice range ?\nif yes,  [for sure](https://www.kaggle.com/code/henrikbrk/fragment-flattening)",
          "votes": 1,
          "replies": [
            {
              "id": 2257726,
              "postDate": "2023-05-13T16:01:02.083Z",
              "content": "<p>my experiments shows that attention pooling (at each location) over z slice gives the best results.<br>\nusing a Unet-resnet34 with encoder pooling at each scale (i.e each location can have different z pool for different scale), you can easily get CV (fragment_id) &gt;0.61</p>",
              "rawMarkdown": "my experiments shows that attention pooling (at each location) over z slice gives the best results.\nusing a Unet-resnet34 with encoder pooling at each scale (i.e each location can have different z pool for different scale), you can easily get CV (fragment_id) >0.61",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2256505,
      "postDate": "2023-05-12T14:29:34.123Z",
      "content": "<p>thanks for sharing ; does your cv correlates with lb ?</p>",
      "rawMarkdown": "thanks for sharing ; does your cv correlates with lb ?",
      "votes": 1,
      "replies": [
        {
          "id": 2256511,
          "postDate": "2023-05-12T14:35:07.847Z",
          "content": "<p>Sure,first one: CV:0.608 LB:0.64, second one I am submitting now. </p>",
          "rawMarkdown": "Sure,first one: CV:0.608 LB:0.64, second one I am submitting now. "
        },
        {
          "id": 2257189,
          "postDate": "2023-05-13T05:41:13.740Z",
          "content": "<p>Second one: CV:0.652 LB:0.61. It got worse…</p>",
          "rawMarkdown": "Second one: CV:0.652 LB:0.61. It got worse...",
          "votes": -1,
          "replies": [
            {
              "id": 2258108,
              "postDate": "2023-05-13T21:59:47.517Z",
              "content": "<p>why don't you probe the best channels on the public test set?<br>\n(but results for private test will be different unless that are from the same fragment)</p>\n<p>this enable you to find the upper limit of improvement if there is any.<br>\n(you can use this limit to 'evaluate' you automatic/adaptive channel selection later)</p>",
              "rawMarkdown": "why don't you probe the best channels on the public test set?\n(but results for private test will be different unless that are from the same fragment)\n\nthis enable you to find the upper limit of improvement if there is any.\n(you can use this limit to 'evaluate' you automatic/adaptive channel selection later)",
              "votes": 1
            },
            {
              "id": 2258197,
              "postDate": "2023-05-14T01:28:20.357Z",
              "content": "<p>This is a good idea:D  ,but I don't have enough submit times for testing.(When I train a model ,I have to find the best threshold because it is very unstable….)</p>",
              "rawMarkdown": "This is a good idea:D  ,but I don't have enough submit times for testing.(When I train a model ,I have to find the best threshold because it is very unstable....)",
              "votes": -1
            }
          ]
        }
      ]
    },
    {
      "id": 2265037,
      "postDate": "2023-05-19T00:19:47.660Z",
      "content": "<p>be very careful to find best parameters (e.g. z crop region or threshold) using public LB score.</p>\n<p>it is not known if the public test data is random ~10% pixel or ~10% region (i.e. continous crop) or others (e.g. if the test fragement has 30 characters, maybe 3 are used for public).<br>\nBut if you want, it is not difficult to probe.</p>\n<p>As shown in fragment train 1,2,3 some region has more FP then others, i.e. the error are not uniformly distributed.</p>\n<p>but there are good news:</p>\n<ul>\n<li>public and private test comes from same fragment surface volume(i.e. same material)</li>\n<li>the ratio of ink and non-ink pixels probably quite the same in private and test (would be a good idea if threshold is based on num of detected pixels or some detected trained image characteristics)</li>\n</ul>",
      "rawMarkdown": "be very careful to find best parameters (e.g. z crop region or threshold) using public LB score.\n\nit is not known if the public test data is random ~10% pixel or ~10% region (i.e. continous crop) or others (e.g. if the test fragement has 30 characters, maybe 3 are used for public).\nBut if you want, it is not difficult to probe.\n\nAs shown in fragment train 1,2,3 some region has more FP then others, i.e. the error are not uniformly distributed.\n\nbut there are good news:\n- public and private test comes from same fragment surface volume(i.e. same material)\n- the ratio of ink and non-ink pixels probably quite the same in private and test (would be a good idea if threshold is based on num of detected pixels or some detected trained image characteristics)\n",
      "votes": 2,
      "replies": [
        {
          "id": 2265094,
          "postDate": "2023-05-19T01:22:13.830Z",
          "content": "<p>as stated <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/402699\" target=\"_blank\">here</a>, if i udnerstand correctly contiguous pixel = region<br>\nthe more i advance in this competion the less i think the public leaderboard is relevant especially with only 10% test data…</p>",
          "rawMarkdown": "as stated [here](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/402699), if i udnerstand correctly contiguous pixel = region\nthe more i advance in this competion the less i think the public leaderboard is relevant especially with only 10% test data...",
          "votes": 1,
          "replies": [
            {
              "id": 2265103,
              "postDate": "2023-05-19T01:29:46.177Z",
              "content": "<p>It might be relevant if the user only has a small number of submissions, but I think 100+ submissions on the 10% region is going to lead to a big shakeup between private and public lb.</p>",
              "rawMarkdown": "It might be relevant if the user only has a small number of submissions, but I think 100+ submissions on the 10% region is going to lead to a big shakeup between private and public lb."
            },
            {
              "id": 2265129,
              "postDate": "2023-05-19T02:26:37.113Z",
              "content": "<p>no. no shakeup. because private and public data is from same fragment</p>",
              "rawMarkdown": "no. no shakeup. because private and public data is from same fragment\n"
            },
            {
              "id": 2276627,
              "postDate": "2023-05-27T04:47:35.803Z",
              "content": "<p>I do think there will be a shakeup, based on my experiments, it is possible for a certain region(10% of the whole data) to have a f1 score of 0.85, in other words, the score on different regions are not distributed evenly.</p>",
              "rawMarkdown": "I do think there will be a shakeup, based on my experiments, it is possible for a certain region(10% of the whole data) to have a f1 score of 0.85, in other words, the score on different regions are not distributed evenly."
            },
            {
              "id": 2276789,
              "postDate": "2023-05-27T07:56:09.023Z",
              "content": "<p>there are two test data: fragmnent4a and fragment4b.<br>\nit is easy to determine with is public and private</p>\n<p>you can actually estimate the private score by:</p>\n<ul>\n<li>make predict on the private fragment</li>\n<li>use online learning to distill the knowledge/results from private to public etc</li>\n</ul>\n<p>i think the top kagglers can avoid shakeup easily using this kind of methods or others.<br>\nthe main reason is that i think they already know what is the 4th fragment from the dataset paper, etc.</p>\n<hr>\n<p>a simple this will be if the threshold used has the same proportional of detected pixels in public and private. this can determine if the thresold is is overfitting the public or not</p>\n<p>there will not be shake up for the team ranking but the fbeta score of the private fragment may be lower</p>",
              "rawMarkdown": "there are two test data: fragmnent4a and fragment4b.\nit is easy to determine with is public and private\n\nyou can actually estimate the private score by:\n- make predict on the private fragment\n- use online learning to distill the knowledge/results from private to public etc\n\ni think the top kagglers can avoid shakeup easily using this kind of methods or others.\nthe main reason is that i think they already know what is the 4th fragment from the dataset paper, etc.\n\n---\n\na simple this will be if the threshold used has the same proportional of detected pixels in public and private. this can determine if the thresold is is overfitting the public or not\n\nthere will not be shake up for the team ranking but the fbeta score of the private fragment may be lower\n"
            }
          ]
        },
        {
          "id": 2265157,
          "postDate": "2023-05-19T03:04:11.780Z",
          "content": "<blockquote>\n  <p>public and private test comes from same fragment surface volume(i.e. same material)</p>\n</blockquote>\n<p>Where is this evidence described?</p>",
          "rawMarkdown": ">public and private test comes from same fragment surface volume(i.e. same material)\n\nWhere is this evidence described?",
          "replies": [
            {
              "id": 2265164,
              "postDate": "2023-05-19T03:17:35.750Z",
              "content": "<p>i deduce from the dateset paper and other reading (from other websites, etc) about the competition…<br>\nbut you can prove it by probing with a \"same patch classifier\" or some image characteristics (histogram, textures, etc)</p>\n<hr>\n<p>on a side note, from the paper, test data 4th fragement has 25 characters</p>",
              "rawMarkdown": "i deduce from the dateset paper and other reading (from other websites, etc) about the competition...\nbut you can prove it by probing with a \"same patch classifier\" or some image characteristics (histogram, textures, etc)\n\n---\n\non a side note, from the paper, test data 4th fragement has 25 characters",
              "votes": 2
            },
            {
              "id": 2265165,
              "postDate": "2023-05-19T03:19:26.430Z",
              "content": "<p>\"The dataset contains 3d x-ray scans of four such fragments at 4µm resolution,\"</p>\n<p>further, it is stated in kaggle website:<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection</a></p>\n<p>FOUR</p>\n<hr>\n<p>i feed all information to chatgpt and ask with the prompt \"answer as you are the organizer of the kaggle competitor, are the public and private set from the same fragment?\"</p>",
              "rawMarkdown": "\"The dataset contains 3d x-ray scans of four such fragments at 4µm resolution,\"\n\nfurther, it is stated in kaggle website:\nhttps://www.kaggle.com/competitions/vesuvius-challenge-ink-detection\n\nFOUR\n\n----\n\ni feed all information to chatgpt and ask with the prompt \"answer as you are the organizer of the kaggle competitor, are the public and private set from the same fragment?\""
            },
            {
              "id": 2265188,
              "postDate": "2023-05-19T04:01:38.867Z",
              "content": "<p>Thanks for the reply.<br>\nI suspect that one of the train sets and public are duplicates… (fold0: val_ink_id=1, fold1: val_ink_id=2_a, fold2: val_ink_id=3, fold3: val_ink_id=2_b, fold4: val_ink_id=2_c)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Faffe8f056be992699d583d058ad69264%2F2023-05-19%2012.52.33.png?generation=1684468721336559&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Thanks for the reply.\nI suspect that one of the train sets and public are duplicates... (fold0: val_ink_id=1, fold1: val_ink_id=2_a, fold2: val_ink_id=3, fold3: val_ink_id=2_b, fold4: val_ink_id=2_c)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Faffe8f056be992699d583d058ad69264%2F2023-05-19%2012.52.33.png?generation=1684468721336559&alt=media)"
            },
            {
              "id": 2265189,
              "postDate": "2023-05-19T04:05:33.320Z",
              "content": "<p>do you mean oen fragment (better say a  region rotated?) is in both test and train?<br>\nif yes, i am pretty sure that a part could be, i overfitted each fold and used the trained weight from each  and i had very similar result like the one you posted, but i thought that was stupid…</p>",
              "rawMarkdown": "do you mean oen fragment (better say a  region rotated?) is in both test and train?\nif yes, i am pretty sure that a part could be, i overfitted each fold and used the trained weight from each  and i had very similar result like the one you posted, but i thought that was stupid..."
            },
            {
              "id": 2265524,
              "postDate": "2023-05-19T10:04:29.333Z",
              "content": "<p>i don't think they are ducplicate. But results are hidden test are correlated to some of the train fragment.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/txc3Fxh/Selection-999-2060.png\" alt=\"Selection-999-2060\"></a><br>\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/4WcFhgf/Selection-999-2059.png\" alt=\"Selection-999-2059\"></a></p>\n<p>The public test lb score only have 2 to 3 characters like this. and i think lb 0.65 results looks like these.<br>\nthey are very sensitive to threshold.</p>\n<p>It is likely that results for the other whole fragment (private test set) can be very different if you get the threshold wrong</p>",
              "rawMarkdown": "i don't think they are ducplicate. But results are hidden test are correlated to some of the train fragment.\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/txc3Fxh/Selection-999-2060.png\" alt=\"Selection-999-2060\" border=\"0\"></a>\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/4WcFhgf/Selection-999-2059.png\" alt=\"Selection-999-2059\" border=\"0\"></a>\n\nThe public test lb score only have 2 to 3 characters like this. and i think lb 0.65 results looks like these.\nthey are very sensitive to threshold.\n\nIt is likely that results for the other whole fragment (private test set) can be very different if you get the threshold wrong",
              "votes": 1
            },
            {
              "id": 2265551,
              "postDate": "2023-05-19T10:32:31.447Z",
              "content": "<p>I hope that is not the case.<br>\nIf not, Trust CV is important because the public is so small.</p>",
              "rawMarkdown": "I hope that is not the case.\nIf not, Trust CV is important because the public is so small."
            },
            {
              "id": 2265553,
              "postDate": "2023-05-19T10:33:12.250Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F9ea018cb2ed4482f05f5bf4f76144f1a%2F2023-05-19%2018-18-13.png?generation=1684491526749471&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F6e2082ac5964e4e1f1a52104e3ba0e14%2F2023-05-19%2018-18-48.png?generation=1684491540955731&amp;alt=media\" alt=\"\"><br>\nThis is my lb 0.67 results looks like.(it trained on 1,3 and half of 2 fragment)<br>\nI think a and b testing fragment basically is the 1 training fragment because their masks are the same.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F0ddb426eb11af75f040b930f1318a4b3%2F2023-05-19%2018-23-31.png?generation=1684492022716611&amp;alt=media\" alt=\"\">  </p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F9ea018cb2ed4482f05f5bf4f76144f1a%2F2023-05-19%2018-18-13.png?generation=1684491526749471&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F6e2082ac5964e4e1f1a52104e3ba0e14%2F2023-05-19%2018-18-48.png?generation=1684491540955731&alt=media)\nThis is my lb 0.67 results looks like.(it trained on 1,3 and half of 2 fragment)\nI think a and b testing fragment basically is the 1 training fragment because their masks are the same.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F0ddb426eb11af75f040b930f1318a4b3%2F2023-05-19%2018-23-31.png?generation=1684492022716611&alt=media)  "
            },
            {
              "id": 2265563,
              "postDate": "2023-05-19T10:40:16.170Z",
              "content": "<p>The test shown may be different from the public test.</p>",
              "rawMarkdown": "The test shown may be different from the public test."
            },
            {
              "id": 2265570,
              "postDate": "2023-05-19T10:49:10.123Z",
              "content": "<p>i got this results</p>\n<pre><code>TTA    threshold   lb\nyes    0.33    0.62\nyes    0.40    0.65\nyes    0.60    0   \nno    0.40    0 ????\n\n#i hope thaere is no bug in my code\n</code></pre>\n<pre><code>validation fragment_id=1            \nCFG.stride=56            \nCFG.window=224    \n\nbce=0.24270            \np_sum  th   prec   recall   fpr   dice   score            \n----------------------------------------------            \n0.33, 0.10, 0.374, 0.683, 0.131,  0.483,  0.411            \n0.15, 0.20, 0.547, 0.445, 0.042,  0.491,  0.523            \n0.09, 0.30, 0.647, 0.327, 0.021,  0.434,  0.541            \n0.06, 0.40, 0.730, 0.243, 0.010,  0.364,  0.521            \n0.04, 0.50, 0.792, 0.178, 0.005,  0.291,  0.468            \n0.03, 0.60, 0.854, 0.128, 0.003,  0.223,  0.401            \n0.02, 0.70, 0.906, 0.081, 0.001,  0.148,  0.297            \n0.01, 0.80, 0.947, 0.035, 0.000,  0.068,  0.153            \n0.00, 0.90, 0.992, 0.002, 0.000,  0.004,  0.009            \n</code></pre>",
              "rawMarkdown": "i got this results\n\n```\nTTA\tthreshold\tlb\nyes\t0.33 \t0.62\nyes\t0.40 \t0.65\nyes\t0.60 \t0\t\nno\t0.40 \t0 ????\n\n#i hope thaere is no bug in my code\n\n```\n\n```\nvalidation fragment_id=1\t\t\t\nCFG.stride=56\t\t\t\nCFG.window=224\t\n\t\t\nbce=0.24270\t\t\t\np_sum  th   prec   recall   fpr   dice   score\t\t\t\n----------------------------------------------\t\t\t\n0.33, 0.10, 0.374, 0.683, 0.131,  0.483,  0.411\t\t\t\n0.15, 0.20, 0.547, 0.445, 0.042,  0.491,  0.523\t\t\t\n0.09, 0.30, 0.647, 0.327, 0.021,  0.434,  0.541\t\t\t\n0.06, 0.40, 0.730, 0.243, 0.010,  0.364,  0.521\t\t\t\n0.04, 0.50, 0.792, 0.178, 0.005,  0.291,  0.468\t\t\t\n0.03, 0.60, 0.854, 0.128, 0.003,  0.223,  0.401\t\t\t\n0.02, 0.70, 0.906, 0.081, 0.001,  0.148,  0.297\t\t\t\n0.01, 0.80, 0.947, 0.035, 0.000,  0.068,  0.153\t\t\t\n0.00, 0.90, 0.992, 0.002, 0.000,  0.004,  0.009\t\t\t\n\n```"
            },
            {
              "id": 2265573,
              "postDate": "2023-05-19T10:54:14.090Z",
              "content": "<p>lowering the threshold below 0.5 wil affect the private LB I feel</p>",
              "rawMarkdown": "lowering the threshold below 0.5 wil affect the private LB I feel"
            }
          ]
        }
      ]
    },
    {
      "id": 2260223,
      "postDate": "2023-05-15T14:38:26.963Z",
      "content": "<p>Would like to know what you mean by selecting partial channels? Does it mean that only some of the channels are selected for training? Or what? I don't really understand it, sorry I'm a beginner, maybe the question is rather elementary.</p>",
      "rawMarkdown": "Would like to know what you mean by selecting partial channels? Does it mean that only some of the channels are selected for training? Or what? I don't really understand it, sorry I'm a beginner, maybe the question is rather elementary.",
      "replies": [
        {
          "id": 2263555,
          "postDate": "2023-05-17T17:19:05.410Z",
          "content": "<p>There are 65 slices per fragment. so we select some of the middle slices for our training because the deeper slices would not contain ink and shallower slices are close to air so they may be damaged</p>",
          "rawMarkdown": "There are 65 slices per fragment. so we select some of the middle slices for our training because the deeper slices would not contain ink and shallower slices are close to air so they may be damaged",
          "votes": 2
        }
      ]
    },
    {
      "id": 2256690,
      "postDate": "2023-05-12T16:53:37.800Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2257208,
      "postDate": "2023-05-13T06:21:50.583Z",
      "content": "<p>thanks for sharing with this post</p>",
      "rawMarkdown": "thanks for sharing with this post",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2303738,
      "author_name": "WangXuC",
      "author_url": "",
      "post_date": "2023-06-15T12:50:04.997000",
      "content": "<p>Congratulations on winning the gold medal. Can you share your solution? I borrowed from your notebook, but my score did not improve afterwards, so I want to learn something from your solution. Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2303809,
          "author_name": "yoyobar",
          "author_url": "",
          "post_date": "2023-06-15T13:37:49.213000",
          "content": "<p>Of course! This is our <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417383\" target=\"_blank\">solution</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2303822,
              "author_name": "WangXuC",
              "author_url": "",
              "post_date": "2023-06-15T13:48:13.113000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2256759,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2023-05-12T17:34:42.423000",
      "content": "<p>In this case does cv mean per fragment or is your validation scheme different?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2256881,
          "author_name": "yoyobar",
          "author_url": "",
          "post_date": "2023-05-12T19:53:51.230000",
          "content": "<p>I trained my model in the same validation scheme. I used 1,2 fragments for training and 3 fragment for validation.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2256562,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-12T15:25:46.357000",
      "content": "<p>is there a chance that the best channels are location dependent?<br>\ni.e. at some x,y location it is 22:34, while at other location it is different at 24:36</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2256578,
          "author_name": "yoyobar",
          "author_url": "",
          "post_date": "2023-05-12T15:39:24.853000",
          "content": "<p>I think it is possible,but it is hard to prove.Of course we can add TTS here to solve this problem.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2257498,
              "author_name": "Pavel Hanchar",
              "author_url": "",
              "post_date": "2023-05-13T12:01:15.283000",
              "content": "<p>check this out <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348#2235071\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/403348#2235071</a> if you are looking for a proof</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2256684,
          "author_name": "Séraphin Lampion",
          "author_url": "",
          "post_date": "2023-05-12T16:48:54.123000",
          "content": "<p>do you mean for one fragment some specific area have better features on  different slice range ?<br>\nif yes,  <a href=\"https://www.kaggle.com/code/henrikbrk/fragment-flattening\" target=\"_blank\">for sure</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2257726,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-13T16:01:02.083000",
              "content": "<p>my experiments shows that attention pooling (at each location) over z slice gives the best results.<br>\nusing a Unet-resnet34 with encoder pooling at each scale (i.e each location can have different z pool for different scale), you can easily get CV (fragment_id) &gt;0.61</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2256505,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2023-05-12T14:29:34.123000",
      "content": "<p>thanks for sharing ; does your cv correlates with lb ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2256511,
          "author_name": "yoyobar",
          "author_url": "",
          "post_date": "2023-05-12T14:35:07.847000",
          "content": "<p>Sure,first one: CV:0.608 LB:0.64, second one I am submitting now. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2257189,
          "author_name": "yoyobar",
          "author_url": "",
          "post_date": "2023-05-13T05:41:13.740000",
          "content": "<p>Second one: CV:0.652 LB:0.61. It got worse…</p>",
          "votes": -1,
          "replies": [
            {
              "id": 2258108,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-13T21:59:47.517000",
              "content": "<p>why don't you probe the best channels on the public test set?<br>\n(but results for private test will be different unless that are from the same fragment)</p>\n<p>this enable you to find the upper limit of improvement if there is any.<br>\n(you can use this limit to 'evaluate' you automatic/adaptive channel selection later)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2258197,
              "author_name": "yoyobar",
              "author_url": "",
              "post_date": "2023-05-14T01:28:20.357000",
              "content": "<p>This is a good idea:D  ,but I don't have enough submit times for testing.(When I train a model ,I have to find the best threshold because it is very unstable….)</p>",
              "votes": -1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2265037,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-19T00:19:47.660000",
      "content": "<p>be very careful to find best parameters (e.g. z crop region or threshold) using public LB score.</p>\n<p>it is not known if the public test data is random ~10% pixel or ~10% region (i.e. continous crop) or others (e.g. if the test fragement has 30 characters, maybe 3 are used for public).<br>\nBut if you want, it is not difficult to probe.</p>\n<p>As shown in fragment train 1,2,3 some region has more FP then others, i.e. the error are not uniformly distributed.</p>\n<p>but there are good news:</p>\n<ul>\n<li>public and private test comes from same fragment surface volume(i.e. same material)</li>\n<li>the ratio of ink and non-ink pixels probably quite the same in private and test (would be a good idea if threshold is based on num of detected pixels or some detected trained image characteristics)</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2265094,
          "author_name": "Séraphin Lampion",
          "author_url": "",
          "post_date": "2023-05-19T01:22:13.830000",
          "content": "<p>as stated <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/402699\" target=\"_blank\">here</a>, if i udnerstand correctly contiguous pixel = region<br>\nthe more i advance in this competion the less i think the public leaderboard is relevant especially with only 10% test data…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2265103,
              "author_name": "Kyle Peters",
              "author_url": "",
              "post_date": "2023-05-19T01:29:46.177000",
              "content": "<p>It might be relevant if the user only has a small number of submissions, but I think 100+ submissions on the 10% region is going to lead to a big shakeup between private and public lb.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265129,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-19T02:26:37.113000",
              "content": "<p>no. no shakeup. because private and public data is from same fragment</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2276627,
              "author_name": "Feng Qilong",
              "author_url": "",
              "post_date": "2023-05-27T04:47:35.803000",
              "content": "<p>I do think there will be a shakeup, based on my experiments, it is possible for a certain region(10% of the whole data) to have a f1 score of 0.85, in other words, the score on different regions are not distributed evenly.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2276789,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-27T07:56:09.023000",
              "content": "<p>there are two test data: fragmnent4a and fragment4b.<br>\nit is easy to determine with is public and private</p>\n<p>you can actually estimate the private score by:</p>\n<ul>\n<li>make predict on the private fragment</li>\n<li>use online learning to distill the knowledge/results from private to public etc</li>\n</ul>\n<p>i think the top kagglers can avoid shakeup easily using this kind of methods or others.<br>\nthe main reason is that i think they already know what is the 4th fragment from the dataset paper, etc.</p>\n<hr>\n<p>a simple this will be if the threshold used has the same proportional of detected pixels in public and private. this can determine if the thresold is is overfitting the public or not</p>\n<p>there will not be shake up for the team ranking but the fbeta score of the private fragment may be lower</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2265157,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2023-05-19T03:04:11.780000",
          "content": "<blockquote>\n  <p>public and private test comes from same fragment surface volume(i.e. same material)</p>\n</blockquote>\n<p>Where is this evidence described?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2265164,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-19T03:17:35.750000",
              "content": "<p>i deduce from the dateset paper and other reading (from other websites, etc) about the competition…<br>\nbut you can prove it by probing with a \"same patch classifier\" or some image characteristics (histogram, textures, etc)</p>\n<hr>\n<p>on a side note, from the paper, test data 4th fragement has 25 characters</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2265165,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-19T03:19:26.430000",
              "content": "<p>\"The dataset contains 3d x-ray scans of four such fragments at 4µm resolution,\"</p>\n<p>further, it is stated in kaggle website:<br>\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection</a></p>\n<p>FOUR</p>\n<hr>\n<p>i feed all information to chatgpt and ask with the prompt \"answer as you are the organizer of the kaggle competitor, are the public and private set from the same fragment?\"</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265188,
              "author_name": "tattaka",
              "author_url": "",
              "post_date": "2023-05-19T04:01:38.867000",
              "content": "<p>Thanks for the reply.<br>\nI suspect that one of the train sets and public are duplicates… (fold0: val_ink_id=1, fold1: val_ink_id=2_a, fold2: val_ink_id=3, fold3: val_ink_id=2_b, fold4: val_ink_id=2_c)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Faffe8f056be992699d583d058ad69264%2F2023-05-19%2012.52.33.png?generation=1684468721336559&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265189,
              "author_name": "Séraphin Lampion",
              "author_url": "",
              "post_date": "2023-05-19T04:05:33.320000",
              "content": "<p>do you mean oen fragment (better say a  region rotated?) is in both test and train?<br>\nif yes, i am pretty sure that a part could be, i overfitted each fold and used the trained weight from each  and i had very similar result like the one you posted, but i thought that was stupid…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265524,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-19T10:04:29.333000",
              "content": "<p>i don't think they are ducplicate. But results are hidden test are correlated to some of the train fragment.</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/txc3Fxh/Selection-999-2060.png\" alt=\"Selection-999-2060\"></a><br>\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/4WcFhgf/Selection-999-2059.png\" alt=\"Selection-999-2059\"></a></p>\n<p>The public test lb score only have 2 to 3 characters like this. and i think lb 0.65 results looks like these.<br>\nthey are very sensitive to threshold.</p>\n<p>It is likely that results for the other whole fragment (private test set) can be very different if you get the threshold wrong</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2265551,
              "author_name": "tattaka",
              "author_url": "",
              "post_date": "2023-05-19T10:32:31.447000",
              "content": "<p>I hope that is not the case.<br>\nIf not, Trust CV is important because the public is so small.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265553,
              "author_name": "yoyobar",
              "author_url": "",
              "post_date": "2023-05-19T10:33:12.250000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F9ea018cb2ed4482f05f5bf4f76144f1a%2F2023-05-19%2018-18-13.png?generation=1684491526749471&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F6e2082ac5964e4e1f1a52104e3ba0e14%2F2023-05-19%2018-18-48.png?generation=1684491540955731&amp;alt=media\" alt=\"\"><br>\nThis is my lb 0.67 results looks like.(it trained on 1,3 and half of 2 fragment)<br>\nI think a and b testing fragment basically is the 1 training fragment because their masks are the same.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F0ddb426eb11af75f040b930f1318a4b3%2F2023-05-19%2018-23-31.png?generation=1684492022716611&amp;alt=media\" alt=\"\">  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265563,
              "author_name": "tattaka",
              "author_url": "",
              "post_date": "2023-05-19T10:40:16.170000",
              "content": "<p>The test shown may be different from the public test.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265570,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-19T10:49:10.123000",
              "content": "<p>i got this results</p>\n<pre><code>TTA    threshold   lb\nyes    0.33    0.62\nyes    0.40    0.65\nyes    0.60    0   \nno    0.40    0 ????\n\n#i hope thaere is no bug in my code\n</code></pre>\n<pre><code>validation fragment_id=1            \nCFG.stride=56            \nCFG.window=224    \n\nbce=0.24270            \np_sum  th   prec   recall   fpr   dice   score            \n----------------------------------------------            \n0.33, 0.10, 0.374, 0.683, 0.131,  0.483,  0.411            \n0.15, 0.20, 0.547, 0.445, 0.042,  0.491,  0.523            \n0.09, 0.30, 0.647, 0.327, 0.021,  0.434,  0.541            \n0.06, 0.40, 0.730, 0.243, 0.010,  0.364,  0.521            \n0.04, 0.50, 0.792, 0.178, 0.005,  0.291,  0.468            \n0.03, 0.60, 0.854, 0.128, 0.003,  0.223,  0.401            \n0.02, 0.70, 0.906, 0.081, 0.001,  0.148,  0.297            \n0.01, 0.80, 0.947, 0.035, 0.000,  0.068,  0.153            \n0.00, 0.90, 0.992, 0.002, 0.000,  0.004,  0.009            \n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265573,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2023-05-19T10:54:14.090000",
              "content": "<p>lowering the threshold below 0.5 wil affect the private LB I feel</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2260223,
      "author_name": "Caicaicheng",
      "author_url": "",
      "post_date": "2023-05-15T14:38:26.963000",
      "content": "<p>Would like to know what you mean by selecting partial channels? Does it mean that only some of the channels are selected for training? Or what? I don't really understand it, sorry I'm a beginner, maybe the question is rather elementary.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2263555,
          "author_name": "Piyush Jain",
          "author_url": "",
          "post_date": "2023-05-17T17:19:05.410000",
          "content": "<p>There are 65 slices per fragment. so we select some of the middle slices for our training because the deeper slices would not contain ink and shallower slices are close to air so they may be damaged</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2256690,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-12T16:53:37.800000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2257208,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-13T06:21:50.583000",
      "content": "<p>thanks for sharing with this post</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2256471": "# I was testing which channels is my model prefers ,these are my experimental procedures.\n\n## 1. Train a model in several consequent channels.\n->Using 12 channel between 12.tif and 38.tif to training and got mean CV:0.561\n## 2. use val dataset to get the CVs in different channels set.\n| range        | CV           |   \n| ------------- |:------:| \n| 12:24  | 0.261201 |  \n| 14:26 | 0.413033 |  \n| 16:28 | 0.515143  |  \n| 18:30 | 0.564901 |  \n| 20:32 | 0.596330   |  \n|24:36 |**0.608991** |\n|26:38 | 0.594269|\n## 3. Draw a line graph to show which channel is better.\n`\n\n    import numpy as np\n    import matplotlib.pyplot as plt\n\n    start=12\n    load=12\n    num=26\n\n    data=np.zeros(num+start)\n    possibility=np.zeros(num+start)\n    index=np.arange(start+num)\n    score=[[12,0.261201],[14,0.413033],[16,0.515143],[18,0.564901],[20,0.596330],\n    [22,0.615607],[24,0.608991],[26,0.594269]]\n\n    for (k,v) in score:\n        data[k:k+load]+=v\n        possibility[k:k+load]+=1\n\n    plt.plot(index[start-2:],data[start-2:],label=\"channel_prefer_count\")\n    plt.plot(index[start-2:],possibility[start-2:],label=\"the possibility of channel appear in training\")\n    plt.plot(index[start-2:],(data[start-2:]/possibility[start-2:])*6-1,label=\"prefer:(channel/pos)*6-1\")#get mean\n\n    plt.legend()\n    plt.xlabel(\"channel_index\")\n    plt.show()\n`![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2F68fcfd9d0f307261a31d4dda96a7791f%2FFigure_1.png?generation=1683898996998474&alt=media)\n\nThis is vary interesting ,my model prefers deeper features. Therefore, I trained another model which using 12 channels from 16.tif to 42.tif and got mean CV:0.626.\n\n| range| CV|   \n| ------------- |:------:| \n|16:28|0.490768|\n|18:30|0.631685|\n|20:32|0.616809|\n|22:34|**0.652383**|\n|24:36|0.648402|\n|26:38|0.628398|\n|28:40|0.557015|\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8293010%2Fa752b5fc32c29a152ebb47ae3e4bff3e%2FFigure_1.png?generation=1683899398335747&alt=media)\n\nIt seems good :D\n## Therefore, I think most of the features are between 18 and 38 channels. :)",
    "2303738": "Congratulations on winning the gold medal. Can you share your solution? I borrowed from your notebook, but my score did not improve afterwards, so I want to learn something from your solution. Thank you!",
    "2256759": "In this case does cv mean per fragment or is your validation scheme different?",
    "2256562": "is there a chance that the best channels are location dependent?\ni.e. at some x,y location it is 22:34, while at other location it is different at 24:36",
    "2256505": "thanks for sharing ; does your cv correlates with lb ?",
    "2265037": "be very careful to find best parameters (e.g. z crop region or threshold) using public LB score.\n\nit is not known if the public test data is random ~10% pixel or ~10% region (i.e. continous crop) or others (e.g. if the test fragement has 30 characters, maybe 3 are used for public).\nBut if you want, it is not difficult to probe.\n\nAs shown in fragment train 1,2,3 some region has more FP then others, i.e. the error are not uniformly distributed.\n\nbut there are good news:\n- public and private test comes from same fragment surface volume(i.e. same material)\n- the ratio of ink and non-ink pixels probably quite the same in private and test (would be a good idea if threshold is based on num of detected pixels or some detected trained image characteristics)\n",
    "2260223": "Would like to know what you mean by selecting partial channels? Does it mean that only some of the channels are selected for training? Or what? I don't really understand it, sorry I'm a beginner, maybe the question is rather elementary.",
    "2256690": "",
    "2257208": "thanks for sharing with this post"
  }
}