{
  "id": 407283,
  "title": "Some Insights",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/407283",
  "author_name": "Yassine Alouini",
  "post_date": "2023-05-05T21:33:04.972000",
  "votes": 18,
  "comment_count": 21,
  "views": 0,
  "content": "<p>A new challenge, new insights. </p>\n<p>The scroll challenge consists in three main prizes:</p>\n<ol>\n<li>700k $ Grand Prize: the main prize, consists in understanding unopened scrolls, ends by end of 2023.</li>\n<li>100k $ Ink Detection: the Kaggle challenge, consists in ink detection from detached fragments, ends by 14th of June 2023.</li>\n<li>35k $ Segmentation Tooling: consists in developing segmentation tools, ends by 15th of May 2023.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F59c4caff37181cb9bc3ecfc1158275f8%2FScreenshot_from_2023-05-04_21-47-58.png?generation=1683322531308350&amp;alt=media\" alt=\"Scroll Challenge Prizes\"></p>\n<p>In what follows, notice that we will mainly focus on the <strong>Ink Detection Progress Prize (</strong>ends in June 14th). You can find more details about it from the official website’s <a href=\"https://scrollprize.org/ink_detection\" target=\"_blank\">section</a>.</p>\n<p><a href=\"https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm\" target=\"_blank\">https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm</a></p>\n<p>Let’s start by exploring the data!</p>\n<h2>Data</h2>\n<p>In this Kaggle competition, images have been extracted from <strong>3D X-ray scans</strong> of <strong>detached fragments</strong> of ancient papyrus rolls.</p>\n<p>This is in contrast with the general competition in which you need to work with 3D X-ray scans of unopened scrolls.</p>\n<p>Here is the data workflow: </p>\n<p>scrolls ⇒ detached fragments ⇒ 3D X-ray of these scans.</p>\n<p>Indeed, these scrolls are among thousands that have been carbonized during the eruption of the <a href=\"https://en.wikipedia.org/wiki/Mount_Etna\" target=\"_blank\">Etna</a> almost 2000 years ago. </p>\n<p>More details about the technical process can be found in the official website <a href=\"https://scrollprize.org/\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F6d137a22e5fd6eecd486e87d70872237%2Ftechnical_overview.png?generation=1683322247082263&amp;alt=media\" alt=\"Technical Overview\"></p>\n<h2>Technical Overview</h2>\n<p>The data comes in two folders:  <strong>train</strong> and <strong>test</strong>. It has the following hierarchy:</p>\n<pre><code>[train/test]/[fragment_id]\n</code></pre>\n<p>The content of the folders is organized in the following manner:</p>\n<h3>Training</h3>\n<p>For training, we have:</p>\n<ul>\n<li>3 folders with the fragment ids: <strong>1</strong>, <strong>2</strong>, and <strong>3</strong>.</li>\n<li>Each fragment contains <strong>65</strong> TIFF files.</li>\n<li>Each TIFF file has the following dimensions:</li>\n<li>These <strong>65</strong> TIFF files constitute a 3D volume.</li>\n</ul>\n<h3>Testing</h3>\n<ul>\n<li>2 folders this time, <strong>a</strong> and <strong>b</strong>.</li>\n<li>Again, <strong>65</strong> TIFF files.</li>\n</ul>\n<h2>Task</h2>\n<p>This competition's objective is to predict ink masks given a 3D volume (via TIFF files).</p>\n<p>It is thus a <strong>binary</strong> (there is ink or not) <strong>semantic segmentation (</strong>one mask to predict) <strong>task</strong>.</p>\n<h2>Evaluation</h2>\n<p>To evaluate this model, the <a href=\"https://machinelearningmastery.com/fbeta-measure-for-machine-learning/\" target=\"_blank\">\\( F_{\\beta}\\) score</a> is used. How does that relate to other segmentation metrics? Let’s find out.</p>\n<p>A quick reminder, for segmentation, a popular evaluation metric is the <a href=\"https://en.wikipedia.org/wiki/Jaccard_index\" target=\"_blank\">IoU</a> metric (short for <strong>intersection over union</strong>). We compute the pixels that are shared between the true mask \\( X \\) and the predicted one \\( Y \\). The formula from there is: </p>\n<p>$$IoU := \\frac{|X \\cap Y|}{|X| + |Y|}$$</p>\n<p>This formula translates to </p>\n<p>$$\\frac{TP}{TP+FP+FN}$$</p>\n<p>where \\( TP \\)  is true positives, i.e. ink pixels correctly predicted, \\( FP \\) is false positives, i.e. ink pixels are predicted but aren’t there, \\( FN \\) ink pixels aren’t predicted but there are.</p>\n<p>This is easy to see from the <strong>IoU</strong> formula since the intersection means that both prediction and reality agree and the union means we add up predictions (\\( TP \\) and \\( FP \\)) and reality (\\( TP \\) and \\( FN \\)) but only once \\( TP \\) . Notice that \\( TN \\)  is what is outside \\( X \\) and \\( Y \\).</p>\n<p>From there, we can make the connection with the \\( F_{1} \\) score. It is the harmonic mean of <strong>precision</strong> and <strong>recall</strong>:</p>\n<p>$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</p>\n<p>$$ R:= \\frac{TP}{TP + FN} $$</p>\n<p>$$ P:= \\frac{TP}{TP + FP} $$</p>\n<p>\\( F_{1} \\) then simplifies to: </p>\n<p>$$ F_{1} = 2\\frac{(TP)^{2}}{(TP)^{2} + FP*TP + (TP)^{2} + FN * TP} = 2\\frac{TP}{TP + FP + TP + FN} = \\frac{2TP}{2TP + FP + FN}$$</p>\n<p>Finally, if we set \\( a = TP \\) and \\( b = FP + FN \\), then \\( IoU=\\frac{a}{a+b} \\)  and \\( F_{1}=\\frac{2a}{2a+b} \\). Thus \\( F_{1}=\\frac{2IoU}{IoU + 1} \\).</p>\n<p>We can change this computation a bit to penalize more <strong>precision</strong> than <strong>recall</strong> or vice-versa by introducing a \\( \\beta \\) parameter. </p>\n<p>This weighting is controlled via the  \\( \\beta \\) parameter: \\( \\beta \\lt 1 \\) lends more weight to <strong>precision</strong>, while \\( \\beta \\gt 1 \\) lends more weight to <strong>recall</strong>. That’s how the  \\( \\beta \\) score works. </p>\n<h2>Models</h2>\n<p>Few models to try during this competition:</p>\n<ul>\n<li><strong><a href=\"https://paperswithcode.com/paper/u-net-convolutional-networks-for-biomedical\" target=\"_blank\">U-Net</a></strong>: this is a very popular model for segmentation tasks and especially for medical segmentation tasks. Indeed, a lot of medical segmentation tasks also take as input a 3D volume. If you use the <a href=\"https://smp.readthedocs.io/en/latest/\" target=\"_blank\">PyTorch segmentation models</a> library, you can try various encoders. One example: <a href=\"https://paperswithcode.com/model/seresnext?variant=seresnext50-32x4d\" target=\"_blank\">SEResNeXt</a>.</li>\n<li><strong>SAM</strong>: a new segmentation model that could be interesting to fine-tune.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F0eb2e1e62c21b97d00b19bb3fbbfbe3c%2Fsam_vesiuvius.png?generation=1683323143316102&amp;alt=media\" alt=\"sam_vesiuvius.png\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F7e7059b38ea98b633d2e2647d4b34ebe%2FScreenshot%20from%202023-04-24%2020-50-31.png?generation=1683529028426903&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F756aefb4fe81010e409d9d8746a2021d%2FScreenshot%20from%202023-04-24%2021-02-22.png?generation=1683529047376100&amp;alt=media\" alt=\"\"></p>\n<h2>Keywords &amp; Concepts</h2>\n<p>Here are some keywords and concepts to help you see through this competition:</p>\n<h3>3D X-Ray</h3>\n<p>A 3D <a href=\"https://en.wikipedia.org/wiki/X-ray\" target=\"_blank\">X-Ray</a> is an imaging technique that consists in recovering hidden 3D structure for a given object. The resulting 3D image is often saved.</p>\n<h3>Run-length Encoding</h3>\n<p>This is a popular encoding technique used in segmentation tasks. It consists in only</p>\n<p>encoding the start and end pixels of a connected segment. This makes the predictions</p>\n<p>much smaller, i.e. save few pixels instead of saving each one in the image.</p>\n<p>It is shortened as RLE.</p>\n<p>Notice that there are at methods to do the encoding: F and C methods. These two methods differ in the way how pixels are processed: from left to right first (C) or from top to bottom first (F). </p>\n<h3>Sorensen-Dice, IoU, \\( F_{1} \\)</h3>\n<p>These are two metrics that define more or less the same thing and are used in segmentation.</p>\n<h3>Fβ Score</h3>\n<p>This is a <a href=\"https://en.wikipedia.org/wiki/F-score#F%CE%B2_score\" target=\"_blank\">generalization</a> of the \\( F_{1} \\) score where the weights given to the precision and recall can be different (they are equal for \\( F_{1} \\))</p>\n<h3>TIFF</h3>\n<p><a href=\"https://en.wikipedia.org/wiki/TIFF\" target=\"_blank\">TIFF</a> (Tag Image File Format) is a file format for storing raster data. It is a popular format for storing scanned documents and in particular medical records (X-rays for example). To read a TIFF file, you can use the following code snippet:</p>\n<h3>Semantic Segmentation</h3>\n<p>This is a specific dense prediction task, i.e. a task that consists in predicting labels for all the pixels of the image (in contrast with say classification where we only predict a few classes for the whole image).</p>\n<p>Semantic in this context means that we don’t distinguish between various instances of the same object. So for example if we have two scrolls, we will predict them as belonging to the same mask. For instance segmentation, we will predict two different instances.</p>\n<h3>Herculaneum Papyri</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fb3920b9fcf667a9a1b77f0bdf42459ad%2FUntitled.png?generation=1683322095180986&amp;alt=media\" alt=\"Untitled\"></p>\n<p><a href=\"https://en.wikipedia.org/wiki/Herculaneum_papyri\" target=\"_blank\">https://en.wikipedia.org/wiki/Herculaneum_papyri</a> </p>\n<h2>Useful Previous Competitions</h2>\n<p>Here are some previous Kaggle competitions that can be useful: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview</a></li>\n</ul>\n<h2>Resources</h2>\n<p>Some additional resources if you want to dig further:</p>\n<ul>\n<li>Interesting video: <a href=\"https://www.youtube.com/watch?v=PpNq2cFotyY&amp;ab_channel=VisCenter\" target=\"_blank\">https://www.youtube.com/watch?v=PpNq2cFotyY&amp;ab_channel=VisCenter</a></li>\n<li>Official website: <a href=\"https://scrollprize.org/\" target=\"_blank\">https://scrollprize.org/</a></li>\n<li>For more details about <strong>RLE</strong>, someone made a great topic <a href=\"https://www.notion.so/Vesuvius-challenge-ink-detection-95d3f080bd8741ae84f33fa03e79e0b8\" target=\"_blank\">here</a>.</li>\n<li>For more details about <strong>semantic segmentation</strong>, check this great <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/396178\" target=\"_blank\">discussion thread</a>.</li>\n<li>Run-length encoding and decoding: <a href=\"https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script\" target=\"_blank\">https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script</a></li>\n<li>Same as above but more efficient: <a href=\"https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook\" target=\"_blank\">https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook</a></li>\n<li>How to work with TIFF files (warning, my own work): <a href=\"https://www.kaggle.com/code/yassinealouini/working-with-tiff-files\" target=\"_blank\">https://www.kaggle.com/code/yassinealouini/working-with-tiff-files</a> and discussion thread <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681</a></li>\n<li>Various segmentation metrics detailed (warning, my own work again): <a href=\"https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics\" target=\"_blank\">https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook\" target=\"_blank\">https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook</a></li>\n<li><a href=\"https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world\" target=\"_blank\">https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world</a></li>\n</ul>",
  "messages": [
    {
      "id": 2247305,
      "postDate": "2023-05-05T21:33:04.973Z",
      "content": "<p>A new challenge, new insights. </p>\n<p>The scroll challenge consists in three main prizes:</p>\n<ol>\n<li>700k $ Grand Prize: the main prize, consists in understanding unopened scrolls, ends by end of 2023.</li>\n<li>100k $ Ink Detection: the Kaggle challenge, consists in ink detection from detached fragments, ends by 14th of June 2023.</li>\n<li>35k $ Segmentation Tooling: consists in developing segmentation tools, ends by 15th of May 2023.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F59c4caff37181cb9bc3ecfc1158275f8%2FScreenshot_from_2023-05-04_21-47-58.png?generation=1683322531308350&amp;alt=media\" alt=\"Scroll Challenge Prizes\"></p>\n<p>In what follows, notice that we will mainly focus on the <strong>Ink Detection Progress Prize (</strong>ends in June 14th). You can find more details about it from the official website’s <a href=\"https://scrollprize.org/ink_detection\" target=\"_blank\">section</a>.</p>\n<p><a href=\"https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm\" target=\"_blank\">https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm</a></p>\n<p>Let’s start by exploring the data!</p>\n<h2>Data</h2>\n<p>In this Kaggle competition, images have been extracted from <strong>3D X-ray scans</strong> of <strong>detached fragments</strong> of ancient papyrus rolls.</p>\n<p>This is in contrast with the general competition in which you need to work with 3D X-ray scans of unopened scrolls.</p>\n<p>Here is the data workflow: </p>\n<p>scrolls ⇒ detached fragments ⇒ 3D X-ray of these scans.</p>\n<p>Indeed, these scrolls are among thousands that have been carbonized during the eruption of the <a href=\"https://en.wikipedia.org/wiki/Mount_Etna\" target=\"_blank\">Etna</a> almost 2000 years ago. </p>\n<p>More details about the technical process can be found in the official website <a href=\"https://scrollprize.org/\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F6d137a22e5fd6eecd486e87d70872237%2Ftechnical_overview.png?generation=1683322247082263&amp;alt=media\" alt=\"Technical Overview\"></p>\n<h2>Technical Overview</h2>\n<p>The data comes in two folders:  <strong>train</strong> and <strong>test</strong>. It has the following hierarchy:</p>\n<pre><code>[train/test]/[fragment_id]\n</code></pre>\n<p>The content of the folders is organized in the following manner:</p>\n<h3>Training</h3>\n<p>For training, we have:</p>\n<ul>\n<li>3 folders with the fragment ids: <strong>1</strong>, <strong>2</strong>, and <strong>3</strong>.</li>\n<li>Each fragment contains <strong>65</strong> TIFF files.</li>\n<li>Each TIFF file has the following dimensions:</li>\n<li>These <strong>65</strong> TIFF files constitute a 3D volume.</li>\n</ul>\n<h3>Testing</h3>\n<ul>\n<li>2 folders this time, <strong>a</strong> and <strong>b</strong>.</li>\n<li>Again, <strong>65</strong> TIFF files.</li>\n</ul>\n<h2>Task</h2>\n<p>This competition's objective is to predict ink masks given a 3D volume (via TIFF files).</p>\n<p>It is thus a <strong>binary</strong> (there is ink or not) <strong>semantic segmentation (</strong>one mask to predict) <strong>task</strong>.</p>\n<h2>Evaluation</h2>\n<p>To evaluate this model, the <a href=\"https://machinelearningmastery.com/fbeta-measure-for-machine-learning/\" target=\"_blank\">\\( F_{\\beta}\\) score</a> is used. How does that relate to other segmentation metrics? Let’s find out.</p>\n<p>A quick reminder, for segmentation, a popular evaluation metric is the <a href=\"https://en.wikipedia.org/wiki/Jaccard_index\" target=\"_blank\">IoU</a> metric (short for <strong>intersection over union</strong>). We compute the pixels that are shared between the true mask \\( X \\) and the predicted one \\( Y \\). The formula from there is: </p>\n<p>$$IoU := \\frac{|X \\cap Y|}{|X| + |Y|}$$</p>\n<p>This formula translates to </p>\n<p>$$\\frac{TP}{TP+FP+FN}$$</p>\n<p>where \\( TP \\)  is true positives, i.e. ink pixels correctly predicted, \\( FP \\) is false positives, i.e. ink pixels are predicted but aren’t there, \\( FN \\) ink pixels aren’t predicted but there are.</p>\n<p>This is easy to see from the <strong>IoU</strong> formula since the intersection means that both prediction and reality agree and the union means we add up predictions (\\( TP \\) and \\( FP \\)) and reality (\\( TP \\) and \\( FN \\)) but only once \\( TP \\) . Notice that \\( TN \\)  is what is outside \\( X \\) and \\( Y \\).</p>\n<p>From there, we can make the connection with the \\( F_{1} \\) score. It is the harmonic mean of <strong>precision</strong> and <strong>recall</strong>:</p>\n<p>$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</p>\n<p>$$ R:= \\frac{TP}{TP + FN} $$</p>\n<p>$$ P:= \\frac{TP}{TP + FP} $$</p>\n<p>\\( F_{1} \\) then simplifies to: </p>\n<p>$$ F_{1} = 2\\frac{(TP)^{2}}{(TP)^{2} + FP*TP + (TP)^{2} + FN * TP} = 2\\frac{TP}{TP + FP + TP + FN} = \\frac{2TP}{2TP + FP + FN}$$</p>\n<p>Finally, if we set \\( a = TP \\) and \\( b = FP + FN \\), then \\( IoU=\\frac{a}{a+b} \\)  and \\( F_{1}=\\frac{2a}{2a+b} \\). Thus \\( F_{1}=\\frac{2IoU}{IoU + 1} \\).</p>\n<p>We can change this computation a bit to penalize more <strong>precision</strong> than <strong>recall</strong> or vice-versa by introducing a \\( \\beta \\) parameter. </p>\n<p>This weighting is controlled via the  \\( \\beta \\) parameter: \\( \\beta \\lt 1 \\) lends more weight to <strong>precision</strong>, while \\( \\beta \\gt 1 \\) lends more weight to <strong>recall</strong>. That’s how the  \\( \\beta \\) score works. </p>\n<h2>Models</h2>\n<p>Few models to try during this competition:</p>\n<ul>\n<li><strong><a href=\"https://paperswithcode.com/paper/u-net-convolutional-networks-for-biomedical\" target=\"_blank\">U-Net</a></strong>: this is a very popular model for segmentation tasks and especially for medical segmentation tasks. Indeed, a lot of medical segmentation tasks also take as input a 3D volume. If you use the <a href=\"https://smp.readthedocs.io/en/latest/\" target=\"_blank\">PyTorch segmentation models</a> library, you can try various encoders. One example: <a href=\"https://paperswithcode.com/model/seresnext?variant=seresnext50-32x4d\" target=\"_blank\">SEResNeXt</a>.</li>\n<li><strong>SAM</strong>: a new segmentation model that could be interesting to fine-tune.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F0eb2e1e62c21b97d00b19bb3fbbfbe3c%2Fsam_vesiuvius.png?generation=1683323143316102&amp;alt=media\" alt=\"sam_vesiuvius.png\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F7e7059b38ea98b633d2e2647d4b34ebe%2FScreenshot%20from%202023-04-24%2020-50-31.png?generation=1683529028426903&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F756aefb4fe81010e409d9d8746a2021d%2FScreenshot%20from%202023-04-24%2021-02-22.png?generation=1683529047376100&amp;alt=media\" alt=\"\"></p>\n<h2>Keywords &amp; Concepts</h2>\n<p>Here are some keywords and concepts to help you see through this competition:</p>\n<h3>3D X-Ray</h3>\n<p>A 3D <a href=\"https://en.wikipedia.org/wiki/X-ray\" target=\"_blank\">X-Ray</a> is an imaging technique that consists in recovering hidden 3D structure for a given object. The resulting 3D image is often saved.</p>\n<h3>Run-length Encoding</h3>\n<p>This is a popular encoding technique used in segmentation tasks. It consists in only</p>\n<p>encoding the start and end pixels of a connected segment. This makes the predictions</p>\n<p>much smaller, i.e. save few pixels instead of saving each one in the image.</p>\n<p>It is shortened as RLE.</p>\n<p>Notice that there are at methods to do the encoding: F and C methods. These two methods differ in the way how pixels are processed: from left to right first (C) or from top to bottom first (F). </p>\n<h3>Sorensen-Dice, IoU, \\( F_{1} \\)</h3>\n<p>These are two metrics that define more or less the same thing and are used in segmentation.</p>\n<h3>Fβ Score</h3>\n<p>This is a <a href=\"https://en.wikipedia.org/wiki/F-score#F%CE%B2_score\" target=\"_blank\">generalization</a> of the \\( F_{1} \\) score where the weights given to the precision and recall can be different (they are equal for \\( F_{1} \\))</p>\n<h3>TIFF</h3>\n<p><a href=\"https://en.wikipedia.org/wiki/TIFF\" target=\"_blank\">TIFF</a> (Tag Image File Format) is a file format for storing raster data. It is a popular format for storing scanned documents and in particular medical records (X-rays for example). To read a TIFF file, you can use the following code snippet:</p>\n<h3>Semantic Segmentation</h3>\n<p>This is a specific dense prediction task, i.e. a task that consists in predicting labels for all the pixels of the image (in contrast with say classification where we only predict a few classes for the whole image).</p>\n<p>Semantic in this context means that we don’t distinguish between various instances of the same object. So for example if we have two scrolls, we will predict them as belonging to the same mask. For instance segmentation, we will predict two different instances.</p>\n<h3>Herculaneum Papyri</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fb3920b9fcf667a9a1b77f0bdf42459ad%2FUntitled.png?generation=1683322095180986&amp;alt=media\" alt=\"Untitled\"></p>\n<p><a href=\"https://en.wikipedia.org/wiki/Herculaneum_papyri\" target=\"_blank\">https://en.wikipedia.org/wiki/Herculaneum_papyri</a> </p>\n<h2>Useful Previous Competitions</h2>\n<p>Here are some previous Kaggle competitions that can be useful: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview</a></li>\n</ul>\n<h2>Resources</h2>\n<p>Some additional resources if you want to dig further:</p>\n<ul>\n<li>Interesting video: <a href=\"https://www.youtube.com/watch?v=PpNq2cFotyY&amp;ab_channel=VisCenter\" target=\"_blank\">https://www.youtube.com/watch?v=PpNq2cFotyY&amp;ab_channel=VisCenter</a></li>\n<li>Official website: <a href=\"https://scrollprize.org/\" target=\"_blank\">https://scrollprize.org/</a></li>\n<li>For more details about <strong>RLE</strong>, someone made a great topic <a href=\"https://www.notion.so/Vesuvius-challenge-ink-detection-95d3f080bd8741ae84f33fa03e79e0b8\" target=\"_blank\">here</a>.</li>\n<li>For more details about <strong>semantic segmentation</strong>, check this great <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/396178\" target=\"_blank\">discussion thread</a>.</li>\n<li>Run-length encoding and decoding: <a href=\"https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script\" target=\"_blank\">https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script</a></li>\n<li>Same as above but more efficient: <a href=\"https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook\" target=\"_blank\">https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook</a></li>\n<li>How to work with TIFF files (warning, my own work): <a href=\"https://www.kaggle.com/code/yassinealouini/working-with-tiff-files\" target=\"_blank\">https://www.kaggle.com/code/yassinealouini/working-with-tiff-files</a> and discussion thread <a href=\"https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681</a></li>\n<li>Various segmentation metrics detailed (warning, my own work again): <a href=\"https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics\" target=\"_blank\">https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics</a>.</li>\n<li><a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook\" target=\"_blank\">https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook</a></li>\n<li><a href=\"https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world\" target=\"_blank\">https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world</a></li>\n</ul>",
      "rawMarkdown": "A new challenge, new insights. \n\nThe scroll challenge consists in three main prizes:\n\n1. 700k $ Grand Prize: the main prize, consists in understanding unopened scrolls, ends by end of 2023.\n2. 100k $ Ink Detection: the Kaggle challenge, consists in ink detection from detached fragments, ends by 14th of June 2023.\n3. 35k $ Segmentation Tooling: consists in developing segmentation tools, ends by 15th of May 2023.\n\n![Scroll Challenge Prizes](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F59c4caff37181cb9bc3ecfc1158275f8%2FScreenshot_from_2023-05-04_21-47-58.png?generation=1683322531308350&alt=media)\n\nIn what follows, notice that we will mainly focus on the **Ink Detection Progress Prize (**ends in June 14th). You can find more details about it from the official website’s [section](https://scrollprize.org/ink_detection).\n\n[https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm](https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm)\n\nLet’s start by exploring the data!\n\n## Data\n\nIn this Kaggle competition, images have been extracted from **3D X-ray scans** of **detached fragments** of ancient papyrus rolls.\n\nThis is in contrast with the general competition in which you need to work with 3D X-ray scans of unopened scrolls.\n\nHere is the data workflow: \n\nscrolls ⇒ detached fragments ⇒ 3D X-ray of these scans.\n\nIndeed, these scrolls are among thousands that have been carbonized during the eruption of the [Etna](https://en.wikipedia.org/wiki/Mount_Etna) almost 2000 years ago. \n\nMore details about the technical process can be found in the official website [here](https://scrollprize.org/).\n\n![Technical Overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F6d137a22e5fd6eecd486e87d70872237%2Ftechnical_overview.png?generation=1683322247082263&alt=media)\n\n## Technical Overview\n\nThe data comes in two folders:  **train** and **test**. It has the following hierarchy:\n\n```bash\n[train/test]/[fragment_id]\n```\n\nThe content of the folders is organized in the following manner:\n\n### Training\n\nFor training, we have:\n\n- 3 folders with the fragment ids: **1**, **2**, and **3**.\n- Each fragment contains **65** TIFF files.\n- Each TIFF file has the following dimensions:\n- These **65** TIFF files constitute a 3D volume.\n\n### Testing\n\n- 2 folders this time, **a** and **b**.\n- Again, **65** TIFF files.\n\n## Task\n\nThis competition's objective is to predict ink masks given a 3D volume (via TIFF files).\n\nIt is thus a **binary** (there is ink or not) **semantic segmentation (**one mask to predict) **task**.\n\n## Evaluation\n\nTo evaluate this model, the [\\\\( F_{\\beta}\\\\) score](https://machinelearningmastery.com/fbeta-measure-for-machine-learning/) is used. How does that relate to other segmentation metrics? Let’s find out.\n\nA quick reminder, for segmentation, a popular evaluation metric is the [IoU](https://en.wikipedia.org/wiki/Jaccard_index) metric (short for **intersection over union**). We compute the pixels that are shared between the true mask \\\\( X \\\\) and the predicted one \\\\( Y \\\\). The formula from there is: \n\n$$IoU := \\frac{|X \\cap Y|}{|X| + |Y|}$$\n\n\nThis formula translates to \n\n$$\\frac{TP}{TP+FP+FN}$$\n\nwhere \\\\( TP \\\\)  is true positives, i.e. ink pixels correctly predicted, \\\\( FP \\\\) is false positives, i.e. ink pixels are predicted but aren’t there, \\\\( FN \\\\) ink pixels aren’t predicted but there are.\n\nThis is easy to see from the **IoU** formula since the intersection means that both prediction and reality agree and the union means we add up predictions (\\\\( TP \\\\) and \\\\( FP \\\\)) and reality (\\\\( TP \\\\) and \\\\( FN \\\\)) but only once \\\\( TP \\\\) . Notice that \\\\( TN \\\\)  is what is outside \\\\( X \\\\) and \\\\( Y \\\\).\n\nFrom there, we can make the connection with the \\\\( F_{1} \\\\) score. It is the harmonic mean of **precision** and **recall**:\n\n\n$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$\n\n\n$$ R:= \\frac{TP}{TP + FN} $$\n\n$$ P:= \\frac{TP}{TP + FP} $$\n\n\n\\\\( F_{1} \\\\) then simplifies to: \n\n$$ F_{1} = 2\\frac{(TP)^{2}}{(TP)^{2} + FP*TP + (TP)^{2} + FN * TP} = 2\\frac{TP}{TP + FP + TP + FN} = \\frac{2TP}{2TP + FP + FN}$$\n\nFinally, if we set \\\\( a = TP \\\\) and \\\\( b = FP + FN \\\\), then \\\\( IoU=\\frac{a}{a+b} \\\\)  and \\\\( F_{1}=\\frac{2a}{2a+b} \\\\). Thus \\\\( F_{1}=\\frac{2IoU}{IoU + 1} \\\\).\n\nWe can change this computation a bit to penalize more **precision** than **recall** or vice-versa by introducing a \\\\( \\beta \\\\) parameter. \n\nThis weighting is controlled via the  \\\\( \\beta \\\\) parameter: \\\\( \\beta \\lt 1 \\\\) lends more weight to **precision**, while \\\\( \\beta \\gt 1 \\\\) lends more weight to **recall**. That’s how the  \\\\( \\beta \\\\) score works. \n\n## Models\n\nFew models to try during this competition:\n\n- **[U-Net](https://paperswithcode.com/paper/u-net-convolutional-networks-for-biomedical)**: this is a very popular model for segmentation tasks and especially for medical segmentation tasks. Indeed, a lot of medical segmentation tasks also take as input a 3D volume. If you use the [PyTorch segmentation models](https://smp.readthedocs.io/en/latest/) library, you can try various encoders. One example: [SEResNeXt](https://paperswithcode.com/model/seresnext?variant=seresnext50-32x4d).\n- **SAM**: a new segmentation model that could be interesting to fine-tune.\n\n![sam_vesiuvius.png](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F0eb2e1e62c21b97d00b19bb3fbbfbe3c%2Fsam_vesiuvius.png?generation=1683323143316102&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F7e7059b38ea98b633d2e2647d4b34ebe%2FScreenshot%20from%202023-04-24%2020-50-31.png?generation=1683529028426903&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F756aefb4fe81010e409d9d8746a2021d%2FScreenshot%20from%202023-04-24%2021-02-22.png?generation=1683529047376100&alt=media)\n\n## Keywords & Concepts\n\nHere are some keywords and concepts to help you see through this competition:\n\n### 3D X-Ray\n\nA 3D [X-Ray](https://en.wikipedia.org/wiki/X-ray) is an imaging technique that consists in recovering hidden 3D structure for a given object. The resulting 3D image is often saved.\n\n### Run-length Encoding\n\nThis is a popular encoding technique used in segmentation tasks. It consists in only\n\nencoding the start and end pixels of a connected segment. This makes the predictions\n\nmuch smaller, i.e. save few pixels instead of saving each one in the image.\n\nIt is shortened as RLE.\n\nNotice that there are at methods to do the encoding: F and C methods. These two methods differ in the way how pixels are processed: from left to right first (C) or from top to bottom first (F). \n\n### Sorensen-Dice, IoU, \\\\( F_{1} \\\\)\n\nThese are two metrics that define more or less the same thing and are used in segmentation.\n\n### Fβ Score\n\nThis is a [generalization](https://en.wikipedia.org/wiki/F-score#F%CE%B2_score) of the \\\\( F_{1} \\\\) score where the weights given to the precision and recall can be different (they are equal for \\\\( F_{1} \\\\))\n\n### TIFF\n\n[TIFF](https://en.wikipedia.org/wiki/TIFF) (Tag Image File Format) is a file format for storing raster data. It is a popular format for storing scanned documents and in particular medical records (X-rays for example). To read a TIFF file, you can use the following code snippet:\n\n \n\n### Semantic Segmentation\n\nThis is a specific dense prediction task, i.e. a task that consists in predicting labels for all the pixels of the image (in contrast with say classification where we only predict a few classes for the whole image).\n\nSemantic in this context means that we don’t distinguish between various instances of the same object. So for example if we have two scrolls, we will predict them as belonging to the same mask. For instance segmentation, we will predict two different instances.\n\n### Herculaneum Papyri\n\n![Untitled](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fb3920b9fcf667a9a1b77f0bdf42459ad%2FUntitled.png?generation=1683322095180986&alt=media)\n\n[https://en.wikipedia.org/wiki/Herculaneum_papyri](https://en.wikipedia.org/wiki/Herculaneum_papyri) \n\n## Useful Previous Competitions\n\nHere are some previous Kaggle competitions that can be useful: \n\n- [https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview)\n\n## Resources\n\nSome additional resources if you want to dig further:\n\n- Interesting video: [https://www.youtube.com/watch?v=PpNq2cFotyY&ab_channel=VisCenter](https://www.youtube.com/watch?v=PpNq2cFotyY&ab_channel=VisCenter)\n- Official website: [https://scrollprize.org/](https://scrollprize.org/)\n- For more details about **RLE**, someone made a great topic [here](https://www.notion.so/Vesuvius-challenge-ink-detection-95d3f080bd8741ae84f33fa03e79e0b8).\n- For more details about **semantic segmentation**, check this great [discussion thread](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/396178).\n- Run-length encoding and decoding: [https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script](https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script)\n- Same as above but more efficient: [https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook](https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook)\n- How to work with TIFF files (warning, my own work): [https://www.kaggle.com/code/yassinealouini/working-with-tiff-files](https://www.kaggle.com/code/yassinealouini/working-with-tiff-files) and discussion thread [https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681](https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681)\n- Various segmentation metrics detailed (warning, my own work again): [https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics](https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics).\n- [https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook](https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook)\n- [https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world](https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world)",
      "votes": 18
    },
    {
      "id": 2254640,
      "postDate": "2023-05-11T06:20:05.207Z",
      "content": "<p>omg so interesting posts, ill reed it and tried to give a feedback</p>",
      "rawMarkdown": "omg so interesting posts, ill reed it and tried to give a feedback",
      "votes": 3,
      "replies": [
        {
          "id": 2254960,
          "postDate": "2023-05-11T11:23:30.290Z",
          "content": "<p>That's the spirit! Let me know if you find new additional interesting things to add. 👌</p>",
          "rawMarkdown": "That's the spirit! Let me know if you find new additional interesting things to add. 👌",
          "votes": 1
        }
      ]
    },
    {
      "id": 2263463,
      "postDate": "2023-05-17T16:03:54.797Z",
      "content": "<p>interesting, great explanation <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> </p>",
      "rawMarkdown": "interesting, great explanation @yassinealouini ",
      "votes": 1,
      "replies": [
        {
          "id": 2263620,
          "postDate": "2023-05-17T18:32:54.930Z",
          "content": "<p>Thanks, glad it helps!</p>",
          "rawMarkdown": "Thanks, glad it helps!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2249900,
      "postDate": "2023-05-08T07:16:54.760Z",
      "content": "<p>What is the way to do inline LaTeX? $F_{1}$ doesn't seem to work… 🤔</p>",
      "rawMarkdown": "What is the way to do inline LaTeX? $F_{1}$ doesn't seem to work... 🤔",
      "votes": 1,
      "replies": [
        {
          "id": 2250521,
          "postDate": "2023-05-08T15:45:51.650Z",
          "content": "<p>It's <code>\\\\(  F_{1} \\\\)</code></p>\n<p>This is inline math with \\(  F_{1} \\). See <a href=\"https://www.kaggle.com/general/581\" target=\"_blank\">this post</a> for a full rundown.</p>",
          "rawMarkdown": "It's `\\\\(  F_{1} \\\\)`\n\nThis is inline math with \\\\(  F_{1} \\\\). See [this post](https://www.kaggle.com/general/581) for a full rundown.",
          "votes": 2,
          "replies": [
            {
              "id": 2251262,
              "postDate": "2023-05-09T08:06:11.937Z",
              "content": "<p>Thanks very much for the link! 👍</p>",
              "rawMarkdown": "Thanks very much for the link! 👍",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2249898,
      "postDate": "2023-05-08T07:11:37.227Z",
      "content": "<p>Also, it seems that this equation breaks the LaTeX block part =&gt;</p>\n<p>$$  F_{1}:= \\frac{1}{\\frac{1}{2}<em>(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{P</em>R}{R+P} $$</p>\n<p>notice that it works without the * in frac:</p>\n<p>$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</p>",
      "rawMarkdown": "Also, it seems that this equation breaks the LaTeX block part =>\n\n$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{P*R}{R+P} $$\n\n\nnotice that it works without the * in frac:\n\n$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$",
      "votes": 1,
      "replies": [
        {
          "id": 2257298,
          "postDate": "2023-05-13T08:20:00.620Z",
          "content": "<p>Have anyone found a solution to this mystery? Is that a bug in the implementation or am I doing something wrong here? <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> maybe you have the answer? Thanks!</p>",
          "rawMarkdown": "Have anyone found a solution to this mystery? Is that a bug in the implementation or am I doing something wrong here? @wcukierski maybe you have the answer? Thanks!",
          "votes": 1,
          "replies": [
            {
              "id": 2257735,
              "postDate": "2023-05-13T16:14:36.473Z",
              "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> sorry, can you be more specific about what you are trying to do?</p>",
              "rawMarkdown": "@yassinealouini sorry, can you be more specific about what you are trying to do?",
              "votes": 1
            },
            {
              "id": 2258550,
              "postDate": "2023-05-14T10:08:30.083Z",
              "content": "<p>Trying to get the equation above render correctly. If I add a * between the R and P symbols it breaks. I have no idea though so thought it might bit an implementation bug. </p>",
              "rawMarkdown": "Trying to get the equation above render correctly. If I add a * between the R and P symbols it breaks. I have no idea though so thought it might bit an implementation bug. ",
              "votes": 1
            },
            {
              "id": 2264634,
              "postDate": "2023-05-18T15:36:57.660Z",
              "content": "<p>do you mean smth like that? </p>\n<p>$$ F_{1}:= \\frac{1}{\\frac{1}{2}\\times(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</p>\n<p><code>$$ F_{1}:= \\frac{1}{\\frac{1}{2}\\times(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</code></p>",
              "rawMarkdown": "do you mean smth like that? \n\n$$ F_{1}:= \\frac{1}{\\frac{1}{2}\\times(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$\n\n`$$ F_{1}:= \\frac{1}{\\frac{1}{2}\\times(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$`"
            },
            {
              "id": 2270541,
              "postDate": "2023-05-23T08:03:06.703Z",
              "content": "<p>Almost! The problem is when adding the <code>*</code> symbol between <code>P</code> and <code>R</code>.</p>",
              "rawMarkdown": "Almost! The problem is when adding the `*` symbol between `P` and `R`."
            }
          ]
        }
      ]
    },
    {
      "id": 2247335,
      "postDate": "2023-05-05T22:08:40.450Z",
      "content": "<p>Notice that the formatting is a bit chaotic for now, I am fixing it.</p>",
      "rawMarkdown": "Notice that the formatting is a bit chaotic for now, I am fixing it.",
      "votes": 1
    },
    {
      "id": 2263639,
      "postDate": "2023-05-17T18:51:50.343Z",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> this is well explained and interesting at the same time 🙌🏼🙌🏼</p>",
      "rawMarkdown": "@yassinealouini this is well explained and interesting at the same time 🙌🏼🙌🏼",
      "votes": 2,
      "replies": [
        {
          "id": 2263685,
          "postDate": "2023-05-17T19:54:15.530Z",
          "content": "<p>Awesome, hope you can use some of these insights.</p>",
          "rawMarkdown": "Awesome, hope you can use some of these insights.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2263633,
      "postDate": "2023-05-17T18:48:04.010Z",
      "content": "<p>Thanks for sharing this great explanation <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a></p>",
      "rawMarkdown": "Thanks for sharing this great explanation @yassinealouini",
      "votes": 2,
      "replies": [
        {
          "id": 2263682,
          "postDate": "2023-05-17T19:53:45.673Z",
          "content": "<p>Great! Hope you find it useful.</p>",
          "rawMarkdown": "Great! Hope you find it useful.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2251264,
      "postDate": "2023-05-09T08:07:37.280Z",
      "content": "<p>Update using <a href=\"https://www.kaggle.com/general/581\" target=\"_blank\">https://www.kaggle.com/general/581</a>. Thanks again <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a>!</p>",
      "rawMarkdown": "Update using https://www.kaggle.com/general/581. Thanks again @wcukierski!",
      "votes": 2
    },
    {
      "id": 2270539,
      "postDate": "2023-05-23T08:00:59.317Z",
      "content": "<p>[<strong>UPDATE</strong>] There is this cool web-based tool to render a volumetric scroll: <a href=\"https://github.com/tomhsiao1260/volume-viewer\" target=\"_blank\">https://github.com/tomhsiao1260/volume-viewer</a><br>\nby <a href=\"https://github.com/tomhsiao1260\" target=\"_blank\">https://github.com/tomhsiao1260</a>.</p>\n<p>I will try to play a bit with this and share some insights.</p>",
      "rawMarkdown": "[**UPDATE**] There is this cool web-based tool to render a volumetric scroll: https://github.com/tomhsiao1260/volume-viewer\nby https://github.com/tomhsiao1260.\n\nI will try to play a bit with this and share some insights."
    },
    {
      "id": 2249877,
      "postDate": "2023-05-08T06:58:47.253Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2254640,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-11T06:20:05.207000",
      "content": "<p>omg so interesting posts, ill reed it and tried to give a feedback</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2254960,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2023-05-11T11:23:30.290000",
          "content": "<p>That's the spirit! Let me know if you find new additional interesting things to add. 👌</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2263463,
      "author_name": "Tamanna Akter Swarna",
      "author_url": "",
      "post_date": "2023-05-17T16:03:54.797000",
      "content": "<p>interesting, great explanation <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2263620,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2023-05-17T18:32:54.930000",
          "content": "<p>Thanks, glad it helps!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2249900,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-05-08T07:16:54.760000",
      "content": "<p>What is the way to do inline LaTeX? $F_{1}$ doesn't seem to work… 🤔</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2250521,
          "author_name": "Will Cukierski",
          "author_url": "",
          "post_date": "2023-05-08T15:45:51.650000",
          "content": "<p>It's <code>\\\\(  F_{1} \\\\)</code></p>\n<p>This is inline math with \\(  F_{1} \\). See <a href=\"https://www.kaggle.com/general/581\" target=\"_blank\">this post</a> for a full rundown.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2251262,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2023-05-09T08:06:11.937000",
              "content": "<p>Thanks very much for the link! 👍</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2249898,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-05-08T07:11:37.227000",
      "content": "<p>Also, it seems that this equation breaks the LaTeX block part =&gt;</p>\n<p>$$  F_{1}:= \\frac{1}{\\frac{1}{2}<em>(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{P</em>R}{R+P} $$</p>\n<p>notice that it works without the * in frac:</p>\n<p>$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2257298,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2023-05-13T08:20:00.620000",
          "content": "<p>Have anyone found a solution to this mystery? Is that a bug in the implementation or am I doing something wrong here? <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> maybe you have the answer? Thanks!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2257735,
              "author_name": "Will Cukierski",
              "author_url": "",
              "post_date": "2023-05-13T16:14:36.473000",
              "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> sorry, can you be more specific about what you are trying to do?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2258550,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2023-05-14T10:08:30.083000",
              "content": "<p>Trying to get the equation above render correctly. If I add a * between the R and P symbols it breaks. I have no idea though so thought it might bit an implementation bug. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2264634,
              "author_name": "Ioannis M",
              "author_url": "",
              "post_date": "2023-05-18T15:36:57.660000",
              "content": "<p>do you mean smth like that? </p>\n<p>$$ F_{1}:= \\frac{1}{\\frac{1}{2}\\times(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</p>\n<p><code>$$ F_{1}:= \\frac{1}{\\frac{1}{2}\\times(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$</code></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2270541,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2023-05-23T08:03:06.703000",
              "content": "<p>Almost! The problem is when adding the <code>*</code> symbol between <code>P</code> and <code>R</code>.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2247335,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-05-05T22:08:40.450000",
      "content": "<p>Notice that the formatting is a bit chaotic for now, I am fixing it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2263639,
      "author_name": "Muhammad Bilal Hussain",
      "author_url": "",
      "post_date": "2023-05-17T18:51:50.343000",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> this is well explained and interesting at the same time 🙌🏼🙌🏼</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2263685,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2023-05-17T19:54:15.530000",
          "content": "<p>Awesome, hope you can use some of these insights.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2263633,
      "author_name": "Ms. Nancy Al Aswad",
      "author_url": "",
      "post_date": "2023-05-17T18:48:04.010000",
      "content": "<p>Thanks for sharing this great explanation <a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2263682,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2023-05-17T19:53:45.673000",
          "content": "<p>Great! Hope you find it useful.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2251264,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-05-09T08:07:37.280000",
      "content": "<p>Update using <a href=\"https://www.kaggle.com/general/581\" target=\"_blank\">https://www.kaggle.com/general/581</a>. Thanks again <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a>!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2270539,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-05-23T08:00:59.317000",
      "content": "<p>[<strong>UPDATE</strong>] There is this cool web-based tool to render a volumetric scroll: <a href=\"https://github.com/tomhsiao1260/volume-viewer\" target=\"_blank\">https://github.com/tomhsiao1260/volume-viewer</a><br>\nby <a href=\"https://github.com/tomhsiao1260\" target=\"_blank\">https://github.com/tomhsiao1260</a>.</p>\n<p>I will try to play a bit with this and share some insights.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2249877,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-08T06:58:47.253000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2247305": "A new challenge, new insights. \n\nThe scroll challenge consists in three main prizes:\n\n1. 700k $ Grand Prize: the main prize, consists in understanding unopened scrolls, ends by end of 2023.\n2. 100k $ Ink Detection: the Kaggle challenge, consists in ink detection from detached fragments, ends by 14th of June 2023.\n3. 35k $ Segmentation Tooling: consists in developing segmentation tools, ends by 15th of May 2023.\n\n![Scroll Challenge Prizes](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F59c4caff37181cb9bc3ecfc1158275f8%2FScreenshot_from_2023-05-04_21-47-58.png?generation=1683322531308350&alt=media)\n\nIn what follows, notice that we will mainly focus on the **Ink Detection Progress Prize (**ends in June 14th). You can find more details about it from the official website’s [section](https://scrollprize.org/ink_detection).\n\n[https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm](https://scrollprize.org/img/tutorials/ink-detection-anim2-dark.webm)\n\nLet’s start by exploring the data!\n\n## Data\n\nIn this Kaggle competition, images have been extracted from **3D X-ray scans** of **detached fragments** of ancient papyrus rolls.\n\nThis is in contrast with the general competition in which you need to work with 3D X-ray scans of unopened scrolls.\n\nHere is the data workflow: \n\nscrolls ⇒ detached fragments ⇒ 3D X-ray of these scans.\n\nIndeed, these scrolls are among thousands that have been carbonized during the eruption of the [Etna](https://en.wikipedia.org/wiki/Mount_Etna) almost 2000 years ago. \n\nMore details about the technical process can be found in the official website [here](https://scrollprize.org/).\n\n![Technical Overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F6d137a22e5fd6eecd486e87d70872237%2Ftechnical_overview.png?generation=1683322247082263&alt=media)\n\n## Technical Overview\n\nThe data comes in two folders:  **train** and **test**. It has the following hierarchy:\n\n```bash\n[train/test]/[fragment_id]\n```\n\nThe content of the folders is organized in the following manner:\n\n### Training\n\nFor training, we have:\n\n- 3 folders with the fragment ids: **1**, **2**, and **3**.\n- Each fragment contains **65** TIFF files.\n- Each TIFF file has the following dimensions:\n- These **65** TIFF files constitute a 3D volume.\n\n### Testing\n\n- 2 folders this time, **a** and **b**.\n- Again, **65** TIFF files.\n\n## Task\n\nThis competition's objective is to predict ink masks given a 3D volume (via TIFF files).\n\nIt is thus a **binary** (there is ink or not) **semantic segmentation (**one mask to predict) **task**.\n\n## Evaluation\n\nTo evaluate this model, the [\\\\( F_{\\beta}\\\\) score](https://machinelearningmastery.com/fbeta-measure-for-machine-learning/) is used. How does that relate to other segmentation metrics? Let’s find out.\n\nA quick reminder, for segmentation, a popular evaluation metric is the [IoU](https://en.wikipedia.org/wiki/Jaccard_index) metric (short for **intersection over union**). We compute the pixels that are shared between the true mask \\\\( X \\\\) and the predicted one \\\\( Y \\\\). The formula from there is: \n\n$$IoU := \\frac{|X \\cap Y|}{|X| + |Y|}$$\n\n\nThis formula translates to \n\n$$\\frac{TP}{TP+FP+FN}$$\n\nwhere \\\\( TP \\\\)  is true positives, i.e. ink pixels correctly predicted, \\\\( FP \\\\) is false positives, i.e. ink pixels are predicted but aren’t there, \\\\( FN \\\\) ink pixels aren’t predicted but there are.\n\nThis is easy to see from the **IoU** formula since the intersection means that both prediction and reality agree and the union means we add up predictions (\\\\( TP \\\\) and \\\\( FP \\\\)) and reality (\\\\( TP \\\\) and \\\\( FN \\\\)) but only once \\\\( TP \\\\) . Notice that \\\\( TN \\\\)  is what is outside \\\\( X \\\\) and \\\\( Y \\\\).\n\nFrom there, we can make the connection with the \\\\( F_{1} \\\\) score. It is the harmonic mean of **precision** and **recall**:\n\n\n$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$\n\n\n$$ R:= \\frac{TP}{TP + FN} $$\n\n$$ P:= \\frac{TP}{TP + FP} $$\n\n\n\\\\( F_{1} \\\\) then simplifies to: \n\n$$ F_{1} = 2\\frac{(TP)^{2}}{(TP)^{2} + FP*TP + (TP)^{2} + FN * TP} = 2\\frac{TP}{TP + FP + TP + FN} = \\frac{2TP}{2TP + FP + FN}$$\n\nFinally, if we set \\\\( a = TP \\\\) and \\\\( b = FP + FN \\\\), then \\\\( IoU=\\frac{a}{a+b} \\\\)  and \\\\( F_{1}=\\frac{2a}{2a+b} \\\\). Thus \\\\( F_{1}=\\frac{2IoU}{IoU + 1} \\\\).\n\nWe can change this computation a bit to penalize more **precision** than **recall** or vice-versa by introducing a \\\\( \\beta \\\\) parameter. \n\nThis weighting is controlled via the  \\\\( \\beta \\\\) parameter: \\\\( \\beta \\lt 1 \\\\) lends more weight to **precision**, while \\\\( \\beta \\gt 1 \\\\) lends more weight to **recall**. That’s how the  \\\\( \\beta \\\\) score works. \n\n## Models\n\nFew models to try during this competition:\n\n- **[U-Net](https://paperswithcode.com/paper/u-net-convolutional-networks-for-biomedical)**: this is a very popular model for segmentation tasks and especially for medical segmentation tasks. Indeed, a lot of medical segmentation tasks also take as input a 3D volume. If you use the [PyTorch segmentation models](https://smp.readthedocs.io/en/latest/) library, you can try various encoders. One example: [SEResNeXt](https://paperswithcode.com/model/seresnext?variant=seresnext50-32x4d).\n- **SAM**: a new segmentation model that could be interesting to fine-tune.\n\n![sam_vesiuvius.png](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F0eb2e1e62c21b97d00b19bb3fbbfbe3c%2Fsam_vesiuvius.png?generation=1683323143316102&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F7e7059b38ea98b633d2e2647d4b34ebe%2FScreenshot%20from%202023-04-24%2020-50-31.png?generation=1683529028426903&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F756aefb4fe81010e409d9d8746a2021d%2FScreenshot%20from%202023-04-24%2021-02-22.png?generation=1683529047376100&alt=media)\n\n## Keywords & Concepts\n\nHere are some keywords and concepts to help you see through this competition:\n\n### 3D X-Ray\n\nA 3D [X-Ray](https://en.wikipedia.org/wiki/X-ray) is an imaging technique that consists in recovering hidden 3D structure for a given object. The resulting 3D image is often saved.\n\n### Run-length Encoding\n\nThis is a popular encoding technique used in segmentation tasks. It consists in only\n\nencoding the start and end pixels of a connected segment. This makes the predictions\n\nmuch smaller, i.e. save few pixels instead of saving each one in the image.\n\nIt is shortened as RLE.\n\nNotice that there are at methods to do the encoding: F and C methods. These two methods differ in the way how pixels are processed: from left to right first (C) or from top to bottom first (F). \n\n### Sorensen-Dice, IoU, \\\\( F_{1} \\\\)\n\nThese are two metrics that define more or less the same thing and are used in segmentation.\n\n### Fβ Score\n\nThis is a [generalization](https://en.wikipedia.org/wiki/F-score#F%CE%B2_score) of the \\\\( F_{1} \\\\) score where the weights given to the precision and recall can be different (they are equal for \\\\( F_{1} \\\\))\n\n### TIFF\n\n[TIFF](https://en.wikipedia.org/wiki/TIFF) (Tag Image File Format) is a file format for storing raster data. It is a popular format for storing scanned documents and in particular medical records (X-rays for example). To read a TIFF file, you can use the following code snippet:\n\n \n\n### Semantic Segmentation\n\nThis is a specific dense prediction task, i.e. a task that consists in predicting labels for all the pixels of the image (in contrast with say classification where we only predict a few classes for the whole image).\n\nSemantic in this context means that we don’t distinguish between various instances of the same object. So for example if we have two scrolls, we will predict them as belonging to the same mask. For instance segmentation, we will predict two different instances.\n\n### Herculaneum Papyri\n\n![Untitled](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fb3920b9fcf667a9a1b77f0bdf42459ad%2FUntitled.png?generation=1683322095180986&alt=media)\n\n[https://en.wikipedia.org/wiki/Herculaneum_papyri](https://en.wikipedia.org/wiki/Herculaneum_papyri) \n\n## Useful Previous Competitions\n\nHere are some previous Kaggle competitions that can be useful: \n\n- [https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview)\n\n## Resources\n\nSome additional resources if you want to dig further:\n\n- Interesting video: [https://www.youtube.com/watch?v=PpNq2cFotyY&ab_channel=VisCenter](https://www.youtube.com/watch?v=PpNq2cFotyY&ab_channel=VisCenter)\n- Official website: [https://scrollprize.org/](https://scrollprize.org/)\n- For more details about **RLE**, someone made a great topic [here](https://www.notion.so/Vesuvius-challenge-ink-detection-95d3f080bd8741ae84f33fa03e79e0b8).\n- For more details about **semantic segmentation**, check this great [discussion thread](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/396178).\n- Run-length encoding and decoding: [https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script](https://www.kaggle.com/code/paulorzp/run-length-encode-and-decode/script)\n- Same as above but more efficient: [https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook](https://www.kaggle.com/code/xhlulu/efficient-mask2rle/notebook)\n- How to work with TIFF files (warning, my own work): [https://www.kaggle.com/code/yassinealouini/working-with-tiff-files](https://www.kaggle.com/code/yassinealouini/working-with-tiff-files) and discussion thread [https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681](https://www.kaggle.com/competitions/hubmap-organ-segmentation/discussion/332681)\n- Various segmentation metrics detailed (warning, my own work again): [https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics](https://www.kaggle.com/code/yassinealouini/all-the-segmentation-metrics).\n- [https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook](https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial/notebook)\n- [https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world](https://precise.zeiss.com/en/welcome-to-the-3d-scanning-world)",
    "2254640": "omg so interesting posts, ill reed it and tried to give a feedback",
    "2263463": "interesting, great explanation @yassinealouini ",
    "2249900": "What is the way to do inline LaTeX? $F_{1}$ doesn't seem to work... 🤔",
    "2249898": "Also, it seems that this equation breaks the LaTeX block part =>\n\n$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{P*R}{R+P} $$\n\n\nnotice that it works without the * in frac:\n\n$$  F_{1}:= \\frac{1}{\\frac{1}{2}*(\\frac{1}{P}+\\frac{1}{R})} = 2\\frac{PR}{R+P} $$",
    "2247335": "Notice that the formatting is a bit chaotic for now, I am fixing it.",
    "2263639": "@yassinealouini this is well explained and interesting at the same time 🙌🏼🙌🏼",
    "2263633": "Thanks for sharing this great explanation @yassinealouini",
    "2251264": "Update using https://www.kaggle.com/general/581. Thanks again @wcukierski!",
    "2270539": "[**UPDATE**] There is this cool web-based tool to render a volumetric scroll: https://github.com/tomhsiao1260/volume-viewer\nby https://github.com/tomhsiao1260.\n\nI will try to play a bit with this and share some insights.",
    "2249877": ""
  }
}