{
  "id": 399825,
  "title": "How large are the original X-ray images before tomography?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/399825",
  "author_name": "Ludwig Maes",
  "post_date": "2023-04-05T17:51:06.692000",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>EDIT: I edited the question because it renumbers questions to start again from 1 for each series of questions, so I am prepending alphabetic characters…</p>\n<p>The question is about the actual Vesuvius half-scroll:</p>\n<p>A: 1-5</p>\n<ol>\n<li>How many X-rays were taken, at what step angle? I am assuming 360 as usual.</li>\n<li>What was the pixel resolution of the original X-ray images?</li>\n<li>Whats the digital bit depth on the original X-ray images?</li>\n<li>How large would this dataset be in GB?</li>\n<li>At what X-ray energy was this acquired?</li>\n</ol>\n<p>The \"air\" part (in the video the fragment stack actually looks like</p>\n<ul>\n<li>glass/plastic</li>\n<li>fragment and then</li>\n<li>paper</li>\n</ul>\n<p>B: 1</p>\n<ol>\n<li>Was the text/ink side of the fragment facing the glass panel or the paper?</li>\n</ol>\n<p>The \"air\" part seems to have structure or ringing artifacts</p>\n<p>C: 1-4</p>\n<ol>\n<li><p>Are these the intrinsic result of X-ray imaging?</p></li>\n<li><p>Or are these the result of scattering or diffraction processes?</p></li>\n<li><p>Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?</p></li>\n<li><p>Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?</p></li>\n</ol>\n<p>Consider a hypothetical halfspace of a more radiotranslucent material succeeded by a more radio-opaque material. Depending on the shift of a virtual voxel grid normal to the interface (shifting the voxel grid into or out of the phase boundary), voxels that remain completely within the same material will report similar radiodensity values, but the voxel on the boundary will interpolate. No material with the interpolated radiodensity value is present of course… Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.</p>\n<p>D: 1</p>\n<ol>\n<li>Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?</li>\n</ol>",
  "messages": [
    {
      "id": 2211157,
      "postDate": "2023-04-05T20:46:41.440Z",
      "content": "<blockquote>\n  <p>How many X-rays were taken, at what step angle? I am assuming 360 as usual.</p>\n</blockquote>\n<p>[0, 360] (end inclusive) with a step angle of 0.1 for 3601 projections per offset position. Number of offset positions depends on the sample. For example, Scroll 1 had 2 horizontal and 21 vertical offset positions. Some of these vertical offsets are redundant (this will be explained in <a href=\"https://github.com/educelab/EduceLab-Scrolls\" target=\"_blank\">the paper</a>), but the raw dataset is 3601 * 2 * 21 = 151242 projection images for one scroll.</p>\n<blockquote>\n  <p>What was the pixel resolution of the original X-ray images?</p>\n</blockquote>\n<p>7.91um for the scrolls, 3.24um for the fragments. Due to beam divergence these could be larger by up to 4%, but given where the samples were scanned in the beam, they're likely very close to the nominal pixel size. Note that these were scanned in parallel beam, so the reconstructed pixel size is approximately equal to the detector pixel size.</p>\n<blockquote>\n  <p>Whats the digital bit depth on the original X-ray images?</p>\n</blockquote>\n<p>16-bits of effective depth, stored in 32-bit floats by the acquisition system.</p>\n<blockquote>\n  <p>How large would this dataset be in GB?</p>\n</blockquote>\n<p>Depends on the dataset. The Scroll 1 raw data mentioned above is about 1.9TBs. This says nothing of the reconstructed data.</p>\n<blockquote>\n  <p>At what X-ray energy was this acquired?</p>\n</blockquote>\n<p>54keV for everything. Additionally 88keV for the fragments. We are working to get these released as well and you'll find that this is referenced in the paper.</p>\n<blockquote>\n  <p>The \"air\" part (in the video the fragment stack actually looks like</p>\n  <p>glass/plastic<br>\n    fragment and then<br>\n    paper</p>\n</blockquote>\n<p>Here's a marked up slice from Fragment 1:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14195746%2F943d184b6e2535659dd997057d0612ef%2Ffrag1-example-markup.png?generation=1680724353192735&amp;alt=media\" alt=\"\"></p>\n<p>If you're looking at the slice stack here on Kaggle, note that these are <a href=\"https://scrollprize.org/tutorial3#surface-volumes\" target=\"_blank\">surface volumes</a> and not the original reconstructions. The tissue paper and acrylic frame and not represented in the Kaggle data.</p>\n<blockquote>\n  <p>Was the text/ink side of the fragment facing the glass panel or the paper?</p>\n</blockquote>\n<p>No. The ink side of the fragment is faced away from the tissue paper and acrylic frame. In the marked up image above, it would be on the \"top\" edge of the fragment, where the green line is pointing.</p>\n<blockquote>\n  <p>The \"air\" part seems to have structure or ringing artifacts</p>\n</blockquote>\n<p>These are <a href=\"http://www.edboas.com/science/CT/0012.pdf\" target=\"_blank\">fairly common artifacts</a> in CT scan reconstructions, particularly those reconstructed using filtered backprojection.</p>\n<blockquote>\n  <p>Are these the intrinsic result of X-ray imaging?<br>\n  Or are these the result of scattering or diffraction processes?</p>\n</blockquote>\n<p>These datasets are derived from x-ray attenuation data, i.e. \"normal\" x-ray images.</p>\n<blockquote>\n  <p>Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?</p>\n</blockquote>\n<p>That is incorrect. Segmentation, meshing, and flattening were used to align the photograph-derived ink labels onto the 3D surface of the fragment AND to generate the optimized surface volumes provided on Kaggle. The original reconstructed volumes are available <a href=\"https://scrollprize.org/data\" target=\"_blank\">here</a> if you would like to attempt ink detection without those steps. I would personally be very interested in how you frame the problem without some form of virtual unwrapping.</p>\n<blockquote>\n  <p>Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?</p>\n</blockquote>\n<p>Though virtual unwrapping was used, it is worth clarifying that all fragment datasets used in the Kaggle competition were prepared in the same way. That includes the private dataset that we use for validating results.</p>\n<blockquote>\n  <p>…Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.</p>\n  <p>Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?</p>\n</blockquote>\n<p>That is correct in principle. Multi-energy acquisition could theoretically improve contrast between the ink and papyrus. In practice, this is much harder than it sounds. A few things to consider for the scans themselves are:</p>\n<ul>\n<li>The edge you want to capture is at a low energy (subtle carbon differences), but the scrolls are dense enough that you get very little transmission at these energies with x-ray tubes. Practically, you end up having to scan at higher energies (50kV+)</li>\n<li>Even at these energies, scan times can be long for moderate micro-CT resolutions (12+ hours for a 23um scan of a scroll)</li>\n<li>While that's not the worst scan length as a matter of logistics, you have to consider things like thermal effects on the sample during extended scans. Some materials can expand and contract on the order of microns multiple times over a 12 hour period. This will cause reconstruction problems.</li>\n<li>To really take advantage of multi-energy scans, the reconstructions themselves need to be aligned so you can do voxel-to-voxel comparisons across scans. Any non-linear scaling effects in the reconstruction (e.g. thermal effects) make this harder.</li>\n<li>Some of these things can be solved by capturing at a beam line, which is a brilliant light source, but beam lines have their own limitations. For example, the Diamond beam line we worked on has a lower energy limit of 54keV. Other beam lines can capture at lower energies, but they tend to also have smaller fields-of-view and/or pixel sizes. It's really tricky to find the beam line with the right properties <em>for this task</em>.</li>\n</ul>\n<p>You also need to consider logistics:</p>\n<ul>\n<li>Access to the materials must be negotiated. </li>\n<li>Transport of materials to a scanner or a scanner to the materials is non-trivial. </li>\n<li>One does not simply walk into a beam line. It requires either (a) writing an acceptable scientific proposal, or (b) a lot of money. </li>\n<li>Once you have beam time scheduled, you must additionally make sure your material access window lines up with the beam line window.</li>\n<li>Every sample needs supports (cases, frames, etc.) that are custom-designed and do not risk damage to the physical object. The first 90% of the work is preparing for the scan. The last 90% of the work is scanning and processing the data.</li>\n</ul>\n<p>Anyway, we've thought a lot about multi-energy scanning, and we try to make it happen when we can, but it's just hard to get good data given all of the constraints.</p>",
      "rawMarkdown": "> How many X-rays were taken, at what step angle? I am assuming 360 as usual.\n\n[0, 360] (end inclusive) with a step angle of 0.1 for 3601 projections per offset position. Number of offset positions depends on the sample. For example, Scroll 1 had 2 horizontal and 21 vertical offset positions. Some of these vertical offsets are redundant (this will be explained in [the paper](https://github.com/educelab/EduceLab-Scrolls)), but the raw dataset is 3601 * 2 * 21 = 151242 projection images for one scroll.\n\n> What was the pixel resolution of the original X-ray images?\n\n7.91um for the scrolls, 3.24um for the fragments. Due to beam divergence these could be larger by up to 4%, but given where the samples were scanned in the beam, they're likely very close to the nominal pixel size. Note that these were scanned in parallel beam, so the reconstructed pixel size is approximately equal to the detector pixel size.\n\n> Whats the digital bit depth on the original X-ray images?\n\n16-bits of effective depth, stored in 32-bit floats by the acquisition system.\n\n> How large would this dataset be in GB?\n\nDepends on the dataset. The Scroll 1 raw data mentioned above is about 1.9TBs. This says nothing of the reconstructed data.\n\n> At what X-ray energy was this acquired?\n\n54keV for everything. Additionally 88keV for the fragments. We are working to get these released as well and you'll find that this is referenced in the paper.\n\n> The \"air\" part (in the video the fragment stack actually looks like\n>\n>   glass/plastic\n>   fragment and then\n>   paper\n\nHere's a marked up slice from Fragment 1:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14195746%2F943d184b6e2535659dd997057d0612ef%2Ffrag1-example-markup.png?generation=1680724353192735&alt=media)\n\nIf you're looking at the slice stack here on Kaggle, note that these are [surface volumes](https://scrollprize.org/tutorial3#surface-volumes) and not the original reconstructions. The tissue paper and acrylic frame and not represented in the Kaggle data.\n\n> Was the text/ink side of the fragment facing the glass panel or the paper?\n\nNo. The ink side of the fragment is faced away from the tissue paper and acrylic frame. In the marked up image above, it would be on the \"top\" edge of the fragment, where the green line is pointing.\n\n> The \"air\" part seems to have structure or ringing artifacts\n\nThese are [fairly common artifacts](http://www.edboas.com/science/CT/0012.pdf) in CT scan reconstructions, particularly those reconstructed using filtered backprojection.\n\n> Are these the intrinsic result of X-ray imaging?\n> Or are these the result of scattering or diffraction processes?\n\nThese datasets are derived from x-ray attenuation data, i.e. \"normal\" x-ray images.\n\n> Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?\n\nThat is incorrect. Segmentation, meshing, and flattening were used to align the photograph-derived ink labels onto the 3D surface of the fragment AND to generate the optimized surface volumes provided on Kaggle. The original reconstructed volumes are available [here](https://scrollprize.org/data) if you would like to attempt ink detection without those steps. I would personally be very interested in how you frame the problem without some form of virtual unwrapping.\n\n> Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?\n\nThough virtual unwrapping was used, it is worth clarifying that all fragment datasets used in the Kaggle competition were prepared in the same way. That includes the private dataset that we use for validating results.\n\n> ...Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.\n>\n> Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?\n\nThat is correct in principle. Multi-energy acquisition could theoretically improve contrast between the ink and papyrus. In practice, this is much harder than it sounds. A few things to consider for the scans themselves are:\n* The edge you want to capture is at a low energy (subtle carbon differences), but the scrolls are dense enough that you get very little transmission at these energies with x-ray tubes. Practically, you end up having to scan at higher energies (50kV+)\n* Even at these energies, scan times can be long for moderate micro-CT resolutions (12+ hours for a 23um scan of a scroll)\n* While that's not the worst scan length as a matter of logistics, you have to consider things like thermal effects on the sample during extended scans. Some materials can expand and contract on the order of microns multiple times over a 12 hour period. This will cause reconstruction problems.\n* To really take advantage of multi-energy scans, the reconstructions themselves need to be aligned so you can do voxel-to-voxel comparisons across scans. Any non-linear scaling effects in the reconstruction (e.g. thermal effects) make this harder.\n* Some of these things can be solved by capturing at a beam line, which is a brilliant light source, but beam lines have their own limitations. For example, the Diamond beam line we worked on has a lower energy limit of 54keV. Other beam lines can capture at lower energies, but they tend to also have smaller fields-of-view and/or pixel sizes. It's really tricky to find the beam line with the right properties _for this task_.\n\nYou also need to consider logistics:\n* Access to the materials must be negotiated. \n* Transport of materials to a scanner or a scanner to the materials is non-trivial. \n* One does not simply walk into a beam line. It requires either (a) writing an acceptable scientific proposal, or (b) a lot of money. \n* Once you have beam time scheduled, you must additionally make sure your material access window lines up with the beam line window.\n* Every sample needs supports (cases, frames, etc.) that are custom-designed and do not risk damage to the physical object. The first 90% of the work is preparing for the scan. The last 90% of the work is scanning and processing the data.\n\nAnyway, we've thought a lot about multi-energy scanning, and we try to make it happen when we can, but it's just hard to get good data given all of the constraints.",
      "votes": 12,
      "replies": [
        {
          "id": 2213437,
          "postDate": "2023-04-07T15:17:04.723Z",
          "content": "<p>That was a formidable reply!</p>\n<p>I comprehend this type of research involves a lot difficulties and challenges: from bureaucracy (justifiable given these are unique fragile artifacts) to technological.</p>\n<p>My first series of questions (A 1-5) was mostly motivated by an idea for compressing the original dataset of X-ray images (both lossy and lossless).</p>\n<p>I forgot to ask about the dimensions in pixels of those X-ray images, i.e. for one of the 151242 projection images.</p>\n<p>Are the 2 horizontal and 21 vertical offset positions different image sensors during the same run, or one and the same image sensor during different runs but with the rotating stage translated to different heights?</p>\n<p>I assume the 3601'th angular position is to compare with the 1st, for calibration and for estimating the thermal distortions? (in theory it would seem that for each angle say 123.4 degrees there is a second angle 303.4 degrees which in a perfect world ought to encode the same transmission information.</p>\n<p>Is it known if the backprojection algorithm A) effectively \"simply\" tracks the sinusoidal movement of a given voxel, or B) optimizes for self-consistency so that it also considers the densities along the full ray (necessitating voxel densities from the previous fitting iteration)? I assume it tries to fit a scalar density by voxel, without any angular dependence of voxel opacities.</p>\n<p>EDIT: forgot to ask, is the X-ray beam switched on and off for each position with the sample held still (causing acceleration and deceleration forces at each step), or is the beam constantly on with the sample rotating continuously?</p>\n<p>Does the sensor utilize a global shutter, and in the case of continous illumination and rotation, what duty cycle (or percentage of the time) is the pixel integrating photons? Or is it a rolling shutter, in that case I am also interested in the duty cycle?</p>\n<p>I believe the dataset can be efficiently compressed such that the datastructure permits consulting voxels, but also permits consulting pixels from the original X-ray images.</p>\n<p>Thanks again for your thorough answer,</p>",
          "rawMarkdown": "That was a formidable reply!\n\nI comprehend this type of research involves a lot difficulties and challenges: from bureaucracy (justifiable given these are unique fragile artifacts) to technological.\n\nMy first series of questions (A 1-5) was mostly motivated by an idea for compressing the original dataset of X-ray images (both lossy and lossless).\n\nI forgot to ask about the dimensions in pixels of those X-ray images, i.e. for one of the 151242 projection images.\n\nAre the 2 horizontal and 21 vertical offset positions different image sensors during the same run, or one and the same image sensor during different runs but with the rotating stage translated to different heights?\n\nI assume the 3601'th angular position is to compare with the 1st, for calibration and for estimating the thermal distortions? (in theory it would seem that for each angle say 123.4 degrees there is a second angle 303.4 degrees which in a perfect world ought to encode the same transmission information.\n\nIs it known if the backprojection algorithm A) effectively \"simply\" tracks the sinusoidal movement of a given voxel, or B) optimizes for self-consistency so that it also considers the densities along the full ray (necessitating voxel densities from the previous fitting iteration)? I assume it tries to fit a scalar density by voxel, without any angular dependence of voxel opacities.\n\nEDIT: forgot to ask, is the X-ray beam switched on and off for each position with the sample held still (causing acceleration and deceleration forces at each step), or is the beam constantly on with the sample rotating continuously?\n\nDoes the sensor utilize a global shutter, and in the case of continous illumination and rotation, what duty cycle (or percentage of the time) is the pixel integrating photons? Or is it a rolling shutter, in that case I am also interested in the duty cycle?\n\nI believe the dataset can be efficiently compressed such that the datastructure permits consulting voxels, but also permits consulting pixels from the original X-ray images.\n\nThanks again for your thorough answer,\n",
          "replies": [
            {
              "id": 2213508,
              "postDate": "2023-04-07T16:11:54.477Z",
              "content": "<blockquote>\n  <p>I forgot to ask about the dimensions in pixels of those X-ray images, i.e. for one of the 151242 projection images.</p>\n</blockquote>\n<p>Ahh, right! That's important. 2560x2160 for all scans. The Diamond <a href=\"https://www.diamond.ac.uk/Instruments/Imaging-and-Microscopy/I12/Detectors-at-I12.html#Imaging%20cameras\" target=\"_blank\">high-res camera</a> uses swappable optical modules with the same x-ray camera back. We used Module 2 for the scrolls and Module 3 for the fragments.</p>\n<blockquote>\n  <p>Are the 2 horizontal and 21 vertical offset positions different image sensors during the same run, or on and the same image sensor during different runs but with the rotating stage translated to different heights?</p>\n</blockquote>\n<p>Same sensor, multiple translated positions of the sample w.r.t. the camera.</p>\n<blockquote>\n  <p>I assume the 3601'th angular position is to compare with the 1st, for calibration and for estimating the thermal distortions?</p>\n</blockquote>\n<p>I'm honestly not sure how they use it for reconstruction correction at Diamond. They use an in-house reconstruction system called <a href=\"https://github.com/DiamondLightSource/Savu\" target=\"_blank\">Savu</a> which has more plugins than I can count. I'm also not sure how important thermal effects are in a beamline. The x-ray generation method isn't going to produce as much heat (at least in the sample chamber) and the sample chamber is a room that's bigger than my living room at home (i.e. it has lots of space for absorbing heat fluctuations). I'm sure there's some, particularly at 1-3um resolutions, but not sure how much. A post-scan is definitely a common technique in commercial scanners, though, and I would bet they at least have it in their toolkit even if they don't always need it.</p>\n<blockquote>\n  <p>(in theory it would seem that for each angle say 123.4 degrees there is a second angle 303.4 degrees which in a perfect world ought to encode the same transmission information.</p>\n</blockquote>\n<p>Yes, and when you have a parallel beam, you can exploit this to reduce the number of offset positions you need to capture. If you translate off center-of-rotation +X, then the projections you get for theta &gt; 180 degrees are the same as what you would get if you translated off center-of-rotation -X. We called this \"annular scanning\" during the session because each horizontal offset position adds an annulus to the total scanned volume, but I'm not sure if there's a better name.</p>\n<blockquote>\n  <p>I it known if the backprojection algorithm A) effectively \"simply\" tracks the sinusoidal movement of a given voxel, or B) optimizes for self-consistency so that it also considers the densities along the full ray (necessitating voxel densities from the previous fitting iteration)? I assume it tries to fit a scalar density by voxel, without any angular dependence of voxel opacities.</p>\n</blockquote>\n<p>Backprojection is an analytically-derived reconstruction approach built on the <a href=\"https://en.wikipedia.org/wiki/Projection-slice_theorem\" target=\"_blank\">projection-slice/central slice theorem</a>. It relies on a bunch of stuff about the equivalence between object projections and the Fourier-transformed object, but the takeaway is that there's a formula that lets you compute a voxel's density directly (single step) from all projected samples of that voxel. </p>\n<p>What you're describing sounds more like an <a href=\"https://en.wikipedia.org/wiki/Tomographic_reconstruction#Iterative_Reconstruction_Algorithm\" target=\"_blank\">iterative, forward projection algorithm</a> where you iteratively reconstruct the volume by minimizing the error between the original projections and forward projections of the reconstruction <a href=\"https://en.wikipedia.org/wiki/Tomographic_reconstruction#Gallery\" target=\"_blank\">[example]</a>. If I'm understanding you correctly, both approaches necessarily take into effect the angular dependence of voxel opacities. I found Chapter 3 of <a href=\"https://doi.org/10.1117/3.2197756\" target=\"_blank\">Hsieh's <em>Computed Tomography</em> book</a> particularly good for understanding the math and intuition around the various reconstruction approaches.</p>",
              "rawMarkdown": "> I forgot to ask about the dimensions in pixels of those X-ray images, i.e. for one of the 151242 projection images.\n\nAhh, right! That's important. 2560x2160 for all scans. The Diamond [high-res camera](https://www.diamond.ac.uk/Instruments/Imaging-and-Microscopy/I12/Detectors-at-I12.html#Imaging%20cameras) uses swappable optical modules with the same x-ray camera back. We used Module 2 for the scrolls and Module 3 for the fragments.\n\n> Are the 2 horizontal and 21 vertical offset positions different image sensors during the same run, or on and the same image sensor during different runs but with the rotating stage translated to different heights?\n\nSame sensor, multiple translated positions of the sample w.r.t. the camera.\n\n> I assume the 3601'th angular position is to compare with the 1st, for calibration and for estimating the thermal distortions?\n\nI'm honestly not sure how they use it for reconstruction correction at Diamond. They use an in-house reconstruction system called [Savu](https://github.com/DiamondLightSource/Savu) which has more plugins than I can count. I'm also not sure how important thermal effects are in a beamline. The x-ray generation method isn't going to produce as much heat (at least in the sample chamber) and the sample chamber is a room that's bigger than my living room at home (i.e. it has lots of space for absorbing heat fluctuations). I'm sure there's some, particularly at 1-3um resolutions, but not sure how much. A post-scan is definitely a common technique in commercial scanners, though, and I would bet they at least have it in their toolkit even if they don't always need it.\n\n> (in theory it would seem that for each angle say 123.4 degrees there is a second angle 303.4 degrees which in a perfect world ought to encode the same transmission information.\n\nYes, and when you have a parallel beam, you can exploit this to reduce the number of offset positions you need to capture. If you translate off center-of-rotation +X, then the projections you get for theta > 180 degrees are the same as what you would get if you translated off center-of-rotation -X. We called this \"annular scanning\" during the session because each horizontal offset position adds an annulus to the total scanned volume, but I'm not sure if there's a better name.\n\n> I it known if the backprojection algorithm A) effectively \"simply\" tracks the sinusoidal movement of a given voxel, or B) optimizes for self-consistency so that it also considers the densities along the full ray (necessitating voxel densities from the previous fitting iteration)? I assume it tries to fit a scalar density by voxel, without any angular dependence of voxel opacities.\n\nBackprojection is an analytically-derived reconstruction approach built on the [projection-slice/central slice theorem](https://en.wikipedia.org/wiki/Projection-slice_theorem). It relies on a bunch of stuff about the equivalence between object projections and the Fourier-transformed object, but the takeaway is that there's a formula that lets you compute a voxel's density directly (single step) from all projected samples of that voxel. \n\nWhat you're describing sounds more like an [iterative, forward projection algorithm](https://en.wikipedia.org/wiki/Tomographic_reconstruction#Iterative_Reconstruction_Algorithm) where you iteratively reconstruct the volume by minimizing the error between the original projections and forward projections of the reconstruction [[example]](https://en.wikipedia.org/wiki/Tomographic_reconstruction#Gallery). If I'm understanding you correctly, both approaches necessarily take into effect the angular dependence of voxel opacities. I found Chapter 3 of [Hsieh's _Computed Tomography_ book](https://doi.org/10.1117/3.2197756) particularly good for understanding the math and intuition around the various reconstruction approaches."
            },
            {
              "id": 2213705,
              "postDate": "2023-04-07T19:33:25.290Z",
              "content": "<p>So the original X-ray dataset for a single fragment should be 2560x2160x3601x2x21 = 836.3 GPixels, correct? in the original format of 32-bit floats that would be about 3.042 TB, or about 1.5TB converted to 16-bit integers. Does this sound right?</p>",
              "rawMarkdown": "So the original X-ray dataset for a single fragment should be 2560x2160x3601x2x21 = 836.3 GPixels, correct? in the original format of 32-bit floats that would be about 3.042 TB, or about 1.5TB converted to 16-bit integers. Does this sound right?"
            },
            {
              "id": 2213730,
              "postDate": "2023-04-07T20:00:23.783Z",
              "content": "<p>On small change: It is actually stored as U16 and not 32F. I was misinterpreting the fundamental data type before. That doesn't really change much about the data (still 16-bits of dynamic range), but it does change the conversation about data sizes.</p>\n<p>Otherwise, yeah, that's about right for the Scroll 1. They're all natively stored in Diamond's Nexus file format, which as I understand it is a formalized metadata layer around HDF5. The 1.9TBs I quoted before was calculated by getting the size of a single HDF5 file (one set of 3601 projections) and multiplying by 42. The actual sizes of each chunk vary somewhat (some of the hdf's also contain flatfields and dark fields). Total size for Scroll 1 projections is somewhere between 1.5-1.9TBs.</p>\n<p>And just to be clear, these are the raw x-ray projections, not the reconstructed slices.</p>",
              "rawMarkdown": "On small change: It is actually stored as U16 and not 32F. I was misinterpreting the fundamental data type before. That doesn't really change much about the data (still 16-bits of dynamic range), but it does change the conversation about data sizes.\n\nOtherwise, yeah, that's about right for the Scroll 1. They're all natively stored in Diamond's Nexus file format, which as I understand it is a formalized metadata layer around HDF5. The 1.9TBs I quoted before was calculated by getting the size of a single HDF5 file (one set of 3601 projections) and multiplying by 42. The actual sizes of each chunk vary somewhat (some of the hdf's also contain flatfields and dark fields). Total size for Scroll 1 projections is somewhere between 1.5-1.9TBs.\n\nAnd just to be clear, these are the raw x-ray projections, not the reconstructed slices.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2210923,
      "postDate": "2023-04-05T17:51:06.693Z",
      "content": "<p>EDIT: I edited the question because it renumbers questions to start again from 1 for each series of questions, so I am prepending alphabetic characters…</p>\n<p>The question is about the actual Vesuvius half-scroll:</p>\n<p>A: 1-5</p>\n<ol>\n<li>How many X-rays were taken, at what step angle? I am assuming 360 as usual.</li>\n<li>What was the pixel resolution of the original X-ray images?</li>\n<li>Whats the digital bit depth on the original X-ray images?</li>\n<li>How large would this dataset be in GB?</li>\n<li>At what X-ray energy was this acquired?</li>\n</ol>\n<p>The \"air\" part (in the video the fragment stack actually looks like</p>\n<ul>\n<li>glass/plastic</li>\n<li>fragment and then</li>\n<li>paper</li>\n</ul>\n<p>B: 1</p>\n<ol>\n<li>Was the text/ink side of the fragment facing the glass panel or the paper?</li>\n</ol>\n<p>The \"air\" part seems to have structure or ringing artifacts</p>\n<p>C: 1-4</p>\n<ol>\n<li><p>Are these the intrinsic result of X-ray imaging?</p></li>\n<li><p>Or are these the result of scattering or diffraction processes?</p></li>\n<li><p>Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?</p></li>\n<li><p>Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?</p></li>\n</ol>\n<p>Consider a hypothetical halfspace of a more radiotranslucent material succeeded by a more radio-opaque material. Depending on the shift of a virtual voxel grid normal to the interface (shifting the voxel grid into or out of the phase boundary), voxels that remain completely within the same material will report similar radiodensity values, but the voxel on the boundary will interpolate. No material with the interpolated radiodensity value is present of course… Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.</p>\n<p>D: 1</p>\n<ol>\n<li>Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?</li>\n</ol>",
      "rawMarkdown": "EDIT: I edited the question because it renumbers questions to start again from 1 for each series of questions, so I am prepending alphabetic characters...\n\nThe question is about the actual Vesuvius half-scroll:\n\nA: 1-5\n1. How many X-rays were taken, at what step angle? I am assuming 360 as usual.\n2. What was the pixel resolution of the original X-ray images?\n3. Whats the digital bit depth on the original X-ray images?\n4. How large would this dataset be in GB?\n5. At what X-ray energy was this acquired?\n\nThe \"air\" part (in the video the fragment stack actually looks like\n- glass/plastic\n- fragment and then\n- paper\n\nB: 1\n6. Was the text/ink side of the fragment facing the glass panel or the paper?\n\nThe \"air\" part seems to have structure or ringing artifacts\n\nC: 1-4\n7. Are these the intrinsic result of X-ray imaging?\n8. Or are these the result of scattering or diffraction processes?\n\n9. Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?\n10. Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?\n\nConsider a hypothetical halfspace of a more radiotranslucent material succeeded by a more radio-opaque material. Depending on the shift of a virtual voxel grid normal to the interface (shifting the voxel grid into or out of the phase boundary), voxels that remain completely within the same material will report similar radiodensity values, but the voxel on the boundary will interpolate. No material with the interpolated radiodensity value is present of course... Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.\n\nD: 1\n11. Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?\n",
      "votes": 6
    },
    {
      "id": 2223183,
      "postDate": "2023-04-16T00:17:02.813Z",
      "content": "<p><a href=\"https://www.kaggle.com/csparker\" target=\"_blank\">@csparker</a>,<br>\nI think it is a right thread to post my questions, instead of openning a new thread.<br>\nAre scrolls made from papyrus?</p>",
      "rawMarkdown": "@csparker,\nI think it is a right thread to post my questions, instead of openning a new thread.\nAre scrolls made from papyrus?",
      "replies": [
        {
          "id": 2223190,
          "postDate": "2023-04-16T00:35:48.147Z",
          "content": "<p>Yes, they are made from papyrus.</p>",
          "rawMarkdown": "Yes, they are made from papyrus.",
          "votes": 1,
          "replies": [
            {
              "id": 2223197,
              "postDate": "2023-04-16T01:08:08.567Z",
              "content": "<p><a href=\"https://www.kaggle.com/csparker\" target=\"_blank\">@csparker</a>, thank you for the response.<br>\nI want to understand something.<br>\nAFAIK, the papyrus is made from two layers. The process of making papyrus involved layering the strips of pith in two directions, perpendicular to each other. The first layer of strips was laid parallel to each other, with the strips overlapping slightly to form a continuous sheet. The second layer of strips was then laid perpendicular to the first layer, and the two layers were pressed together to create a single sheet of papyrus.<br>\nThe upper layer of papyrus is the layer that would have been written on. This layer is typically composed of strips that run parallel to the length of the scroll, and perpendicular to the strips in the lower layer.<br>\nIs that correct?<br>\nIf yes, I have another question. Looking into the fragment 1, I see more perpendicular strips than a vertical. The vertical strips always intersects with horizontal ones. If I understand how papyrus is done and the resolution of the image, I would anticipate that the accurate unwrapping of the scroll should contain only one direction of strips in the most of slices.</p>",
              "rawMarkdown": "@csparker, thank you for the response.\nI want to understand something.\nAFAIK, the papyrus is made from two layers. The process of making papyrus involved layering the strips of pith in two directions, perpendicular to each other. The first layer of strips was laid parallel to each other, with the strips overlapping slightly to form a continuous sheet. The second layer of strips was then laid perpendicular to the first layer, and the two layers were pressed together to create a single sheet of papyrus.\nThe upper layer of papyrus is the layer that would have been written on. This layer is typically composed of strips that run parallel to the length of the scroll, and perpendicular to the strips in the lower layer.\nIs that correct?\nIf yes, I have another question. Looking into the fragment 1, I see more perpendicular strips than a vertical. The vertical strips always intersects with horizontal ones. If I understand how papyrus is done and the resolution of the image, I would anticipate that the accurate unwrapping of the scroll should contain only one direction of strips in the most of slices."
            }
          ]
        }
      ]
    },
    {
      "id": 2210946,
      "postDate": "2023-04-05T18:15:27.133Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2211157,
      "author_name": "Seth P.",
      "author_url": "",
      "post_date": "2023-04-05T20:46:41.440000",
      "content": "<blockquote>\n  <p>How many X-rays were taken, at what step angle? I am assuming 360 as usual.</p>\n</blockquote>\n<p>[0, 360] (end inclusive) with a step angle of 0.1 for 3601 projections per offset position. Number of offset positions depends on the sample. For example, Scroll 1 had 2 horizontal and 21 vertical offset positions. Some of these vertical offsets are redundant (this will be explained in <a href=\"https://github.com/educelab/EduceLab-Scrolls\" target=\"_blank\">the paper</a>), but the raw dataset is 3601 * 2 * 21 = 151242 projection images for one scroll.</p>\n<blockquote>\n  <p>What was the pixel resolution of the original X-ray images?</p>\n</blockquote>\n<p>7.91um for the scrolls, 3.24um for the fragments. Due to beam divergence these could be larger by up to 4%, but given where the samples were scanned in the beam, they're likely very close to the nominal pixel size. Note that these were scanned in parallel beam, so the reconstructed pixel size is approximately equal to the detector pixel size.</p>\n<blockquote>\n  <p>Whats the digital bit depth on the original X-ray images?</p>\n</blockquote>\n<p>16-bits of effective depth, stored in 32-bit floats by the acquisition system.</p>\n<blockquote>\n  <p>How large would this dataset be in GB?</p>\n</blockquote>\n<p>Depends on the dataset. The Scroll 1 raw data mentioned above is about 1.9TBs. This says nothing of the reconstructed data.</p>\n<blockquote>\n  <p>At what X-ray energy was this acquired?</p>\n</blockquote>\n<p>54keV for everything. Additionally 88keV for the fragments. We are working to get these released as well and you'll find that this is referenced in the paper.</p>\n<blockquote>\n  <p>The \"air\" part (in the video the fragment stack actually looks like</p>\n  <p>glass/plastic<br>\n    fragment and then<br>\n    paper</p>\n</blockquote>\n<p>Here's a marked up slice from Fragment 1:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14195746%2F943d184b6e2535659dd997057d0612ef%2Ffrag1-example-markup.png?generation=1680724353192735&amp;alt=media\" alt=\"\"></p>\n<p>If you're looking at the slice stack here on Kaggle, note that these are <a href=\"https://scrollprize.org/tutorial3#surface-volumes\" target=\"_blank\">surface volumes</a> and not the original reconstructions. The tissue paper and acrylic frame and not represented in the Kaggle data.</p>\n<blockquote>\n  <p>Was the text/ink side of the fragment facing the glass panel or the paper?</p>\n</blockquote>\n<p>No. The ink side of the fragment is faced away from the tissue paper and acrylic frame. In the marked up image above, it would be on the \"top\" edge of the fragment, where the green line is pointing.</p>\n<blockquote>\n  <p>The \"air\" part seems to have structure or ringing artifacts</p>\n</blockquote>\n<p>These are <a href=\"http://www.edboas.com/science/CT/0012.pdf\" target=\"_blank\">fairly common artifacts</a> in CT scan reconstructions, particularly those reconstructed using filtered backprojection.</p>\n<blockquote>\n  <p>Are these the intrinsic result of X-ray imaging?<br>\n  Or are these the result of scattering or diffraction processes?</p>\n</blockquote>\n<p>These datasets are derived from x-ray attenuation data, i.e. \"normal\" x-ray images.</p>\n<blockquote>\n  <p>Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?</p>\n</blockquote>\n<p>That is incorrect. Segmentation, meshing, and flattening were used to align the photograph-derived ink labels onto the 3D surface of the fragment AND to generate the optimized surface volumes provided on Kaggle. The original reconstructed volumes are available <a href=\"https://scrollprize.org/data\" target=\"_blank\">here</a> if you would like to attempt ink detection without those steps. I would personally be very interested in how you frame the problem without some form of virtual unwrapping.</p>\n<blockquote>\n  <p>Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?</p>\n</blockquote>\n<p>Though virtual unwrapping was used, it is worth clarifying that all fragment datasets used in the Kaggle competition were prepared in the same way. That includes the private dataset that we use for validating results.</p>\n<blockquote>\n  <p>…Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.</p>\n  <p>Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?</p>\n</blockquote>\n<p>That is correct in principle. Multi-energy acquisition could theoretically improve contrast between the ink and papyrus. In practice, this is much harder than it sounds. A few things to consider for the scans themselves are:</p>\n<ul>\n<li>The edge you want to capture is at a low energy (subtle carbon differences), but the scrolls are dense enough that you get very little transmission at these energies with x-ray tubes. Practically, you end up having to scan at higher energies (50kV+)</li>\n<li>Even at these energies, scan times can be long for moderate micro-CT resolutions (12+ hours for a 23um scan of a scroll)</li>\n<li>While that's not the worst scan length as a matter of logistics, you have to consider things like thermal effects on the sample during extended scans. Some materials can expand and contract on the order of microns multiple times over a 12 hour period. This will cause reconstruction problems.</li>\n<li>To really take advantage of multi-energy scans, the reconstructions themselves need to be aligned so you can do voxel-to-voxel comparisons across scans. Any non-linear scaling effects in the reconstruction (e.g. thermal effects) make this harder.</li>\n<li>Some of these things can be solved by capturing at a beam line, which is a brilliant light source, but beam lines have their own limitations. For example, the Diamond beam line we worked on has a lower energy limit of 54keV. Other beam lines can capture at lower energies, but they tend to also have smaller fields-of-view and/or pixel sizes. It's really tricky to find the beam line with the right properties <em>for this task</em>.</li>\n</ul>\n<p>You also need to consider logistics:</p>\n<ul>\n<li>Access to the materials must be negotiated. </li>\n<li>Transport of materials to a scanner or a scanner to the materials is non-trivial. </li>\n<li>One does not simply walk into a beam line. It requires either (a) writing an acceptable scientific proposal, or (b) a lot of money. </li>\n<li>Once you have beam time scheduled, you must additionally make sure your material access window lines up with the beam line window.</li>\n<li>Every sample needs supports (cases, frames, etc.) that are custom-designed and do not risk damage to the physical object. The first 90% of the work is preparing for the scan. The last 90% of the work is scanning and processing the data.</li>\n</ul>\n<p>Anyway, we've thought a lot about multi-energy scanning, and we try to make it happen when we can, but it's just hard to get good data given all of the constraints.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2213437,
          "author_name": "Ludwig Maes",
          "author_url": "",
          "post_date": "2023-04-07T15:17:04.723000",
          "content": "<p>That was a formidable reply!</p>\n<p>I comprehend this type of research involves a lot difficulties and challenges: from bureaucracy (justifiable given these are unique fragile artifacts) to technological.</p>\n<p>My first series of questions (A 1-5) was mostly motivated by an idea for compressing the original dataset of X-ray images (both lossy and lossless).</p>\n<p>I forgot to ask about the dimensions in pixels of those X-ray images, i.e. for one of the 151242 projection images.</p>\n<p>Are the 2 horizontal and 21 vertical offset positions different image sensors during the same run, or one and the same image sensor during different runs but with the rotating stage translated to different heights?</p>\n<p>I assume the 3601'th angular position is to compare with the 1st, for calibration and for estimating the thermal distortions? (in theory it would seem that for each angle say 123.4 degrees there is a second angle 303.4 degrees which in a perfect world ought to encode the same transmission information.</p>\n<p>Is it known if the backprojection algorithm A) effectively \"simply\" tracks the sinusoidal movement of a given voxel, or B) optimizes for self-consistency so that it also considers the densities along the full ray (necessitating voxel densities from the previous fitting iteration)? I assume it tries to fit a scalar density by voxel, without any angular dependence of voxel opacities.</p>\n<p>EDIT: forgot to ask, is the X-ray beam switched on and off for each position with the sample held still (causing acceleration and deceleration forces at each step), or is the beam constantly on with the sample rotating continuously?</p>\n<p>Does the sensor utilize a global shutter, and in the case of continous illumination and rotation, what duty cycle (or percentage of the time) is the pixel integrating photons? Or is it a rolling shutter, in that case I am also interested in the duty cycle?</p>\n<p>I believe the dataset can be efficiently compressed such that the datastructure permits consulting voxels, but also permits consulting pixels from the original X-ray images.</p>\n<p>Thanks again for your thorough answer,</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2213508,
              "author_name": "Seth P.",
              "author_url": "",
              "post_date": "2023-04-07T16:11:54.477000",
              "content": "<blockquote>\n  <p>I forgot to ask about the dimensions in pixels of those X-ray images, i.e. for one of the 151242 projection images.</p>\n</blockquote>\n<p>Ahh, right! That's important. 2560x2160 for all scans. The Diamond <a href=\"https://www.diamond.ac.uk/Instruments/Imaging-and-Microscopy/I12/Detectors-at-I12.html#Imaging%20cameras\" target=\"_blank\">high-res camera</a> uses swappable optical modules with the same x-ray camera back. We used Module 2 for the scrolls and Module 3 for the fragments.</p>\n<blockquote>\n  <p>Are the 2 horizontal and 21 vertical offset positions different image sensors during the same run, or on and the same image sensor during different runs but with the rotating stage translated to different heights?</p>\n</blockquote>\n<p>Same sensor, multiple translated positions of the sample w.r.t. the camera.</p>\n<blockquote>\n  <p>I assume the 3601'th angular position is to compare with the 1st, for calibration and for estimating the thermal distortions?</p>\n</blockquote>\n<p>I'm honestly not sure how they use it for reconstruction correction at Diamond. They use an in-house reconstruction system called <a href=\"https://github.com/DiamondLightSource/Savu\" target=\"_blank\">Savu</a> which has more plugins than I can count. I'm also not sure how important thermal effects are in a beamline. The x-ray generation method isn't going to produce as much heat (at least in the sample chamber) and the sample chamber is a room that's bigger than my living room at home (i.e. it has lots of space for absorbing heat fluctuations). I'm sure there's some, particularly at 1-3um resolutions, but not sure how much. A post-scan is definitely a common technique in commercial scanners, though, and I would bet they at least have it in their toolkit even if they don't always need it.</p>\n<blockquote>\n  <p>(in theory it would seem that for each angle say 123.4 degrees there is a second angle 303.4 degrees which in a perfect world ought to encode the same transmission information.</p>\n</blockquote>\n<p>Yes, and when you have a parallel beam, you can exploit this to reduce the number of offset positions you need to capture. If you translate off center-of-rotation +X, then the projections you get for theta &gt; 180 degrees are the same as what you would get if you translated off center-of-rotation -X. We called this \"annular scanning\" during the session because each horizontal offset position adds an annulus to the total scanned volume, but I'm not sure if there's a better name.</p>\n<blockquote>\n  <p>I it known if the backprojection algorithm A) effectively \"simply\" tracks the sinusoidal movement of a given voxel, or B) optimizes for self-consistency so that it also considers the densities along the full ray (necessitating voxel densities from the previous fitting iteration)? I assume it tries to fit a scalar density by voxel, without any angular dependence of voxel opacities.</p>\n</blockquote>\n<p>Backprojection is an analytically-derived reconstruction approach built on the <a href=\"https://en.wikipedia.org/wiki/Projection-slice_theorem\" target=\"_blank\">projection-slice/central slice theorem</a>. It relies on a bunch of stuff about the equivalence between object projections and the Fourier-transformed object, but the takeaway is that there's a formula that lets you compute a voxel's density directly (single step) from all projected samples of that voxel. </p>\n<p>What you're describing sounds more like an <a href=\"https://en.wikipedia.org/wiki/Tomographic_reconstruction#Iterative_Reconstruction_Algorithm\" target=\"_blank\">iterative, forward projection algorithm</a> where you iteratively reconstruct the volume by minimizing the error between the original projections and forward projections of the reconstruction <a href=\"https://en.wikipedia.org/wiki/Tomographic_reconstruction#Gallery\" target=\"_blank\">[example]</a>. If I'm understanding you correctly, both approaches necessarily take into effect the angular dependence of voxel opacities. I found Chapter 3 of <a href=\"https://doi.org/10.1117/3.2197756\" target=\"_blank\">Hsieh's <em>Computed Tomography</em> book</a> particularly good for understanding the math and intuition around the various reconstruction approaches.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2213705,
              "author_name": "Ludwig Maes",
              "author_url": "",
              "post_date": "2023-04-07T19:33:25.290000",
              "content": "<p>So the original X-ray dataset for a single fragment should be 2560x2160x3601x2x21 = 836.3 GPixels, correct? in the original format of 32-bit floats that would be about 3.042 TB, or about 1.5TB converted to 16-bit integers. Does this sound right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2213730,
              "author_name": "Seth P.",
              "author_url": "",
              "post_date": "2023-04-07T20:00:23.783000",
              "content": "<p>On small change: It is actually stored as U16 and not 32F. I was misinterpreting the fundamental data type before. That doesn't really change much about the data (still 16-bits of dynamic range), but it does change the conversation about data sizes.</p>\n<p>Otherwise, yeah, that's about right for the Scroll 1. They're all natively stored in Diamond's Nexus file format, which as I understand it is a formalized metadata layer around HDF5. The 1.9TBs I quoted before was calculated by getting the size of a single HDF5 file (one set of 3601 projections) and multiplying by 42. The actual sizes of each chunk vary somewhat (some of the hdf's also contain flatfields and dark fields). Total size for Scroll 1 projections is somewhere between 1.5-1.9TBs.</p>\n<p>And just to be clear, these are the raw x-ray projections, not the reconstructed slices.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2223183,
      "author_name": "Yuri Kreinin",
      "author_url": "",
      "post_date": "2023-04-16T00:17:02.813000",
      "content": "<p><a href=\"https://www.kaggle.com/csparker\" target=\"_blank\">@csparker</a>,<br>\nI think it is a right thread to post my questions, instead of openning a new thread.<br>\nAre scrolls made from papyrus?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2223190,
          "author_name": "Seth P.",
          "author_url": "",
          "post_date": "2023-04-16T00:35:48.147000",
          "content": "<p>Yes, they are made from papyrus.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2223197,
              "author_name": "Yuri Kreinin",
              "author_url": "",
              "post_date": "2023-04-16T01:08:08.567000",
              "content": "<p><a href=\"https://www.kaggle.com/csparker\" target=\"_blank\">@csparker</a>, thank you for the response.<br>\nI want to understand something.<br>\nAFAIK, the papyrus is made from two layers. The process of making papyrus involved layering the strips of pith in two directions, perpendicular to each other. The first layer of strips was laid parallel to each other, with the strips overlapping slightly to form a continuous sheet. The second layer of strips was then laid perpendicular to the first layer, and the two layers were pressed together to create a single sheet of papyrus.<br>\nThe upper layer of papyrus is the layer that would have been written on. This layer is typically composed of strips that run parallel to the length of the scroll, and perpendicular to the strips in the lower layer.<br>\nIs that correct?<br>\nIf yes, I have another question. Looking into the fragment 1, I see more perpendicular strips than a vertical. The vertical strips always intersects with horizontal ones. If I understand how papyrus is done and the resolution of the image, I would anticipate that the accurate unwrapping of the scroll should contain only one direction of strips in the most of slices.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2210946,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-05T18:15:27.133000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2211157": "> How many X-rays were taken, at what step angle? I am assuming 360 as usual.\n\n[0, 360] (end inclusive) with a step angle of 0.1 for 3601 projections per offset position. Number of offset positions depends on the sample. For example, Scroll 1 had 2 horizontal and 21 vertical offset positions. Some of these vertical offsets are redundant (this will be explained in [the paper](https://github.com/educelab/EduceLab-Scrolls)), but the raw dataset is 3601 * 2 * 21 = 151242 projection images for one scroll.\n\n> What was the pixel resolution of the original X-ray images?\n\n7.91um for the scrolls, 3.24um for the fragments. Due to beam divergence these could be larger by up to 4%, but given where the samples were scanned in the beam, they're likely very close to the nominal pixel size. Note that these were scanned in parallel beam, so the reconstructed pixel size is approximately equal to the detector pixel size.\n\n> Whats the digital bit depth on the original X-ray images?\n\n16-bits of effective depth, stored in 32-bit floats by the acquisition system.\n\n> How large would this dataset be in GB?\n\nDepends on the dataset. The Scroll 1 raw data mentioned above is about 1.9TBs. This says nothing of the reconstructed data.\n\n> At what X-ray energy was this acquired?\n\n54keV for everything. Additionally 88keV for the fragments. We are working to get these released as well and you'll find that this is referenced in the paper.\n\n> The \"air\" part (in the video the fragment stack actually looks like\n>\n>   glass/plastic\n>   fragment and then\n>   paper\n\nHere's a marked up slice from Fragment 1:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14195746%2F943d184b6e2535659dd997057d0612ef%2Ffrag1-example-markup.png?generation=1680724353192735&alt=media)\n\nIf you're looking at the slice stack here on Kaggle, note that these are [surface volumes](https://scrollprize.org/tutorial3#surface-volumes) and not the original reconstructions. The tissue paper and acrylic frame and not represented in the Kaggle data.\n\n> Was the text/ink side of the fragment facing the glass panel or the paper?\n\nNo. The ink side of the fragment is faced away from the tissue paper and acrylic frame. In the marked up image above, it would be on the \"top\" edge of the fragment, where the green line is pointing.\n\n> The \"air\" part seems to have structure or ringing artifacts\n\nThese are [fairly common artifacts](http://www.edboas.com/science/CT/0012.pdf) in CT scan reconstructions, particularly those reconstructed using filtered backprojection.\n\n> Are these the intrinsic result of X-ray imaging?\n> Or are these the result of scattering or diffraction processes?\n\nThese datasets are derived from x-ray attenuation data, i.e. \"normal\" x-ray images.\n\n> Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?\n\nThat is incorrect. Segmentation, meshing, and flattening were used to align the photograph-derived ink labels onto the 3D surface of the fragment AND to generate the optimized surface volumes provided on Kaggle. The original reconstructed volumes are available [here](https://scrollprize.org/data) if you would like to attempt ink detection without those steps. I would personally be very interested in how you frame the problem without some form of virtual unwrapping.\n\n> Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?\n\nThough virtual unwrapping was used, it is worth clarifying that all fragment datasets used in the Kaggle competition were prepared in the same way. That includes the private dataset that we use for validating results.\n\n> ...Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.\n>\n> Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?\n\nThat is correct in principle. Multi-energy acquisition could theoretically improve contrast between the ink and papyrus. In practice, this is much harder than it sounds. A few things to consider for the scans themselves are:\n* The edge you want to capture is at a low energy (subtle carbon differences), but the scrolls are dense enough that you get very little transmission at these energies with x-ray tubes. Practically, you end up having to scan at higher energies (50kV+)\n* Even at these energies, scan times can be long for moderate micro-CT resolutions (12+ hours for a 23um scan of a scroll)\n* While that's not the worst scan length as a matter of logistics, you have to consider things like thermal effects on the sample during extended scans. Some materials can expand and contract on the order of microns multiple times over a 12 hour period. This will cause reconstruction problems.\n* To really take advantage of multi-energy scans, the reconstructions themselves need to be aligned so you can do voxel-to-voxel comparisons across scans. Any non-linear scaling effects in the reconstruction (e.g. thermal effects) make this harder.\n* Some of these things can be solved by capturing at a beam line, which is a brilliant light source, but beam lines have their own limitations. For example, the Diamond beam line we worked on has a lower energy limit of 54keV. Other beam lines can capture at lower energies, but they tend to also have smaller fields-of-view and/or pixel sizes. It's really tricky to find the beam line with the right properties _for this task_.\n\nYou also need to consider logistics:\n* Access to the materials must be negotiated. \n* Transport of materials to a scanner or a scanner to the materials is non-trivial. \n* One does not simply walk into a beam line. It requires either (a) writing an acceptable scientific proposal, or (b) a lot of money. \n* Once you have beam time scheduled, you must additionally make sure your material access window lines up with the beam line window.\n* Every sample needs supports (cases, frames, etc.) that are custom-designed and do not risk damage to the physical object. The first 90% of the work is preparing for the scan. The last 90% of the work is scanning and processing the data.\n\nAnyway, we've thought a lot about multi-energy scanning, and we try to make it happen when we can, but it's just hard to get good data given all of the constraints.",
    "2210923": "EDIT: I edited the question because it renumbers questions to start again from 1 for each series of questions, so I am prepending alphabetic characters...\n\nThe question is about the actual Vesuvius half-scroll:\n\nA: 1-5\n1. How many X-rays were taken, at what step angle? I am assuming 360 as usual.\n2. What was the pixel resolution of the original X-ray images?\n3. Whats the digital bit depth on the original X-ray images?\n4. How large would this dataset be in GB?\n5. At what X-ray energy was this acquired?\n\nThe \"air\" part (in the video the fragment stack actually looks like\n- glass/plastic\n- fragment and then\n- paper\n\nB: 1\n6. Was the text/ink side of the fragment facing the glass panel or the paper?\n\nThe \"air\" part seems to have structure or ringing artifacts\n\nC: 1-4\n7. Are these the intrinsic result of X-ray imaging?\n8. Or are these the result of scattering or diffraction processes?\n\n9. Is my assumption correct or incorrect that for the ink-detection challenge no virtual meshing and virtual flattening was applied?\n10. Assuming no virtual meshing and flattening (with unknown Jacobians) was used to produce the training TIFF's, may we assume the test dataset followed the same process without meshing and flattening?\n\nConsider a hypothetical halfspace of a more radiotranslucent material succeeded by a more radio-opaque material. Depending on the shift of a virtual voxel grid normal to the interface (shifting the voxel grid into or out of the phase boundary), voxels that remain completely within the same material will report similar radiodensity values, but the voxel on the boundary will interpolate. No material with the interpolated radiodensity value is present of course... Given this issue the most obvious thing to do would be to add X-ray energy \"color\" channels, so that material compositions would correspond to regions in a higher dimensional space such that interpolated values (on a path between such regions) are significantly less likely to cross so disambiguation becomes much more tractable.\n\nD: 1\n11. Apart from dataset size, is there a reason the scrolls aren't scanned at different X-ray energies?\n",
    "2223183": "@csparker,\nI think it is a right thread to post my questions, instead of openning a new thread.\nAre scrolls made from papyrus?",
    "2210946": ""
  }
}