{
  "id": 421339,
  "title": "U-NET Research Paper Deep Analysis | HuBMAP",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/421339",
  "author_name": "AyushS9020",
  "post_date": "2023-07-05T01:57:19.858000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>1 | Pre-Reqeuistics 🎓</h1>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>🔴</td>\n<td><code>High Level Recommended</code></td>\n</tr>\n<tr>\n<td>🟡</td>\n<td><code>Medium Level Recommended</code></td>\n</tr>\n<tr>\n<td>🟢</td>\n<td><code>Low Level Recommended</code></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>Concept</th>\n<th></th>\n<th>Usage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Convolution</td>\n<td>🔴</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Convolution Transpose</td>\n<td>🔴</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Concatenation</td>\n<td>🔴</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Adam</td>\n<td>🔴</td>\n<td><code>Optimizer</code></td>\n</tr>\n<tr>\n<td>Padding</td>\n<td>🔴</td>\n<td><code>Runtime Processing</code></td>\n</tr>\n<tr>\n<td>Strides</td>\n<td>🔴</td>\n<td><code>Runtime Processing</code></td>\n</tr>\n<tr>\n<td>ReLU</td>\n<td>🟡</td>\n<td><code>Activation Function</code></td>\n</tr>\n<tr>\n<td>Sigmoid</td>\n<td>🟡</td>\n<td><code>Activation Function</code></td>\n</tr>\n<tr>\n<td>Binary Cross Entropy</td>\n<td>🟡</td>\n<td><code>Loss</code></td>\n</tr>\n<tr>\n<td>Dropout</td>\n<td>🟢</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Pooling</td>\n<td>🟢</td>\n<td><code>Network Layer</code></td>\n</tr>\n</tbody>\n</table>\n<h1>2 | Starting 🌱</h1>\n<p>We know that <code>CNNs</code> are showing great results in the field of <code>Vision</code> related tasks</p>\n<ul>\n<li><strong>Image Classification</strong> - CNNs have been used to <code>achieve accuracies</code> of over 99% on <code>image classification</code> tasks.</li>\n<li><strong>Object Detection</strong> - CNNs have been used to <code>achieve accuracies</code> of over 90% on <code>object detection</code> tasks</li>\n<li><strong>Semantic Segmentation</strong> - CNNs have been used to <code>achieve accuracies</code> of over 80% on <code>semantic segmentation</code> tasks</li>\n</ul>\n<p>While <code>CNNs</code> have already <code>existed for a long time</code>, their <code>success was limited</code> due to the</p>\n<ul>\n<li>size of the available training sets</li>\n<li>size of the considered networks.</li>\n</ul>\n<p>The <code>normal classification</code> task was actually <code>pretty easy</code> for the network to learn. But as the field of <code>Image Segmentation</code>, increased, it became <code>difficult</code> for the model to <code>learn both the class label and position/localization</code> .<code>Localization</code> is a <code>challenging task</code>, because it requires the model to <code>learn both the appearance</code> of the <code>objects/features</code> of interest, and their <code>spatial location</code> in the image.</p>\n<p>The major use of <code>Image Semantic Segmentation</code> is in <code>Bio-Medical'. In</code>Bio-Medical<code>image processing, the</code>desired output<code>is not just a</code>classification of the image<code>, but also the</code>location of the objects` or features of interest in the image</p>\n<p>For example, in <code>Bio-Medical</code> image processing, it is often <code>important</code> to be able to <code>localize tumors</code> or <code>other abnormalities</code> in <code>images of tissues</code> or <code>organs</code>. This is because the <code>location of these abnormalities</code> can be used to <code>diagnose diseases</code> or to <code>plan treatments</code>.</p>\n<h1>3 | Inspiration 💡</h1>\n<p><strong><a href=\"https://papers.nips.cc/paper_files/paper/2012/file/459a4ddcb586f24efd9395aa7662bc7c-Paper.pdf\" target=\"_blank\">Deep Neural Networks Segment Neuronal Membranes in Electron Microscopy Images</a></strong> used the <code>method of masking to predict the correct labels</code> of the image</p>\n<p><code>Ciresan</code> model had many advantages such as</p>\n<ul>\n<li>Localization - The network was able to identify the location of an object in an image.</li>\n<li>Larger Training Data - the <code>number of patches</code> that was extracted from a <code>training image</code> was <code>much larger than the number of training images themselves</code>.</li>\n</ul>\n<p>This was because a <code>patch</code> is a <code>small</code>, <code>rectangular section</code> of an image. For example, a patch of size 64x64 can be extracted from any image. If we have a training set of 100 images, we can extract a total of 100x64x64 = 32,768,000 <code>patches</code>.</p>\n<p>But the technique used by <code>Ciresan</code> had some <code>drawbacks</code></p>\n<ul>\n<li>Long Training Time - The model was <code>quite slow</code> because the <code>network must be run separately for each patch</code>, which means that the model <code>must be run multiple times for each image</code>. This can be very slow, especially for large images.</li>\n<li>Redundancy -  When an <code>image</code> is <code>divided into patches</code>, there is a lot of <code>overlap between the patches</code>. This is because the <code>patches are typically small</code>, and they are <code>often arranged in a regular grid</code>.</li>\n<li>Localization-Context Trade-Off - The problem of <code>balancing the accuracy of localization</code> with the <code>amount of context</code> .</li>\n<li>* <code>Larger patches</code> allow for more <code>accurate localization</code>, but they <code>also reduce the amount of context</code> that is available to the network. This is because <code>max-pooling</code> layers are typically used to <code>down-sample the image</code>.</li>\n<li>* <code>Smaller patches</code>, allow for <code>more context</code>, but they also <code>reduce the accuracy of localization</code>. This is because <code>smaller patches contain less information</code> about the location of the object.</li>\n</ul>\n<h1>4 | Advantages ✅</h1>\n<ul>\n<li><p>Less Training Data - The model is <code>modified</code> such that it works with <code>very few training images</code> and yields <code>more precise segmentations</code>.</p></li>\n<li><p>Touching Objects - <code>weighted loss</code> is applied, where the <code>separating background labels between touching cells</code> obtain a <code>large weight</code> in the loss function.</p></li>\n</ul>\n<h1>5 | Parameters and Formulations 🧬</h1>\n<h2>5.1 | Weights Intialization</h2>\n<p>The weights are intialized such that each feature map in the network has approximately unit variance. This was achieved by intializing the weights by<br>\n a <code>Gaussian Distribution(1 , sqrt(2/N))</code></p>\n<h2>5.1 | Optimizer</h2>\n<p>The <code>Adam optimizer</code> is used to train the $UNet$ model, and the <code>momentum parameter</code> is set to 0.99.</p>\n<p>The <code>Adam optimizer</code> is a <code>Stochastic</code> <code>Gradient</code> <code>Descent</code> (SGD) <code>optimizer</code> that was introduced in 2014. The <code>momentum parameter</code> controls <code>how much the optimizer remembers its past gradients</code>. Though, a momentum value of 0.99 is a <code>relatively high value</code>. This means that the Adam <code>optimizer</code> will be more likely to <code>follow its past gradients</code>, which can <code>help to improve the stability of training</code> and <code>prevent the model from overfitting</code>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2F55fdd37178b63612afb111403208297d%2FOptimizer.png?generation=1688521271597503&amp;alt=media\"></p>\n<h2>5.2 | Prediction</h2>\n<p>The <code>energy function</code> is computed by a <code>pixel-wise soft-max</code> over the final feature map</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd5ba65b1f7d31619da7823a977567643%2FPredictioon.png?generation=1688521293363603&amp;alt=media\"></p>\n<p>![]()</p>\n<h3>Where</h3>\n<ul>\n<li>a_k(x) =  Feature Channel k</li>\n<li>X 𝛜 Ω = Pixel Position</li>\n<li>K = Number Of Labels</li>\n</ul>\n<h3>Assumptions</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Faa1e84d19521a53ab934ba4562eb0946%2FAssumpetions.png?generation=1688521344912976&amp;alt=media\"></p>\n<h2>5.3 | Loss</h2>\n<p>The cross entropy then penalizes at each position the deviation of p^`_x (x) from 1</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd7ce2dae61705f5d1c91310a3e7844bd%2FLos.png?generation=1688521366011103&amp;alt=media\"></p>\n<h1>6 | Experiments 🥼</h1>\n<p>The network is tested on 2 different <code>segmentation tasks</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Error</th>\n<th>Accuracy</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Neural Structures in Electron Microscopic Recordings</td>\n<td>0.0003529</td>\n<td>0.9203</td>\n</tr>\n<tr>\n<td>Light Microscopic Images</td>\n<td></td>\n<td>92%</td>\n</tr>\n</tbody>\n</table>\n<h1>8 | Acknowledgements 🤝</h1>\n<p>This study was supported by the Excellence Initiative of the German Federal and State governments (EXC 294) and by the BMBF (Fkz 0316185B).</p>\n<h1>9 | Refrences 🔗</h1>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Authors</th>\n<th>Place</th>\n<th>Year</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Deep neural net-works segment neuronal membranes in electron microscopy images</td>\n<td><code>Ciresan</code> <code>D.C.</code> <code>Gambardella</code> <code>L.M.</code> <code>Giusti A.</code> <code>Schmidhuber J.</code></td>\n<td>NIPS. pp. 2852 2860</td>\n<td>2012</td>\n</tr>\n<tr>\n<td>Discriminative un-supervised feature learning with convolutional neural networks</td>\n<td><code>Dosovitskiy</code> <code>A.</code> <code>Springenberg</code> <code>J.T.</code> <code>Riedmiller</code> <code>M.</code> <code>Brox</code> <code>T.</code></td>\n<td>NIPS</td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Rich feature hierarchies for ac-curate object detection and semantic segmentation.</td>\n<td><code>Girshick</code> <code>R.</code> <code>Donahue</code> <code>J.</code> <code>Darrell</code> <code>T.</code> <code>Malik</code> <code>J.</code></td>\n<td>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Hyper-columns for object segmentation and  ne-grained localization</td>\n<td><code>Hariharan</code> <code>B.</code> <code>Arbelez</code> <code>P.</code> <code>Girshick</code> <code>R.</code> <code>Malik</code> <code>J.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Delving deep into rectifiers: Surpassing human-level performance on image-net classification</td>\n<td><code>He</code>, <code>K.</code>, <code>Zhang</code>, <code>X.</code>, <code>Ren</code>, <code>S.</code>, <code>Sun</code>, <code>J.</code></td>\n<td></td>\n<td>2015</td>\n</tr>\n<tr>\n<td>Convolutional architecture for fast feature embedding</td>\n<td><code>Jia</code> <code>Y.</code> <code>Shelhamer</code> <code>E.</code> <code>Donahue</code> <code>J.</code> <code>Karayev</code> <code>S.</code> <code>Long</code> <code>J.</code> <code>Girshick</code> <code>R.</code> <code>Guadar-rama</code> <code>S.</code> <code>Darrell</code> <code>T.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Image-net classification with deep convolutional neural networks</td>\n<td><code>Krizhevsky</code> <code>A.</code> <code>Sutskever</code> <code>I.</code> <code>Hinton</code> <code>G.E.</code></td>\n<td>NIPS. pp. 1106 1114</td>\n<td>2012</td>\n</tr>\n<tr>\n<td>Backpropagation applied to handwritten zip code recognition. Neural Computation</td>\n<td><code>LeCun</code> <code>Y.</code> <code>Boser</code> <code>B.</code> <code>Denker</code> <code>J.S.</code> <code>Henderson</code> <code>D.</code> <code>Howard</code> <code>R.E.</code> <code>Hubbard</code> <code>W.</code> <code>Jackel</code> <code>L.D.</code></td>\n<td>1(4) 541 551</td>\n<td>1989</td>\n</tr>\n<tr>\n<td>Fully convolutional networks for semantic segmentation</td>\n<td><code>Long</code> <code>J.</code> <code>Shelhamer</code> <code>E.</code> <code>Darrell</code> <code>T.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td>A benchmark for comparison of cell tracking algorithms</td>\n<td><code>Maska</code> <code>M.</code> <code>(...)</code> <code>de Solorzano</code> <code>C.O.</code></td>\n<td>Bioinformatics 30, 1609 1617</td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Very deep convolutional networks for large-scale image recognition</td>\n<td><code>Simonyan</code> <code>K.</code> <code>Zisserman</code> <code>A.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td><a href=\"http://www.codesolorzano.com/celltrackingchallenge/Cell_Tracking_Challenge/Welcome.html\" target=\"_blank\">Web page of the cell tracking challenge</a></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><a href=\"http://brainiac2.mit.edu/isbi_challenge/\" target=\"_blank\">Web page of the em segmentation challenge</a></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2330418,
      "postDate": "2023-07-05T01:57:19.860Z",
      "content": "<h1>1 | Pre-Reqeuistics 🎓</h1>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>🔴</td>\n<td><code>High Level Recommended</code></td>\n</tr>\n<tr>\n<td>🟡</td>\n<td><code>Medium Level Recommended</code></td>\n</tr>\n<tr>\n<td>🟢</td>\n<td><code>Low Level Recommended</code></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>Concept</th>\n<th></th>\n<th>Usage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Convolution</td>\n<td>🔴</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Convolution Transpose</td>\n<td>🔴</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Concatenation</td>\n<td>🔴</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Adam</td>\n<td>🔴</td>\n<td><code>Optimizer</code></td>\n</tr>\n<tr>\n<td>Padding</td>\n<td>🔴</td>\n<td><code>Runtime Processing</code></td>\n</tr>\n<tr>\n<td>Strides</td>\n<td>🔴</td>\n<td><code>Runtime Processing</code></td>\n</tr>\n<tr>\n<td>ReLU</td>\n<td>🟡</td>\n<td><code>Activation Function</code></td>\n</tr>\n<tr>\n<td>Sigmoid</td>\n<td>🟡</td>\n<td><code>Activation Function</code></td>\n</tr>\n<tr>\n<td>Binary Cross Entropy</td>\n<td>🟡</td>\n<td><code>Loss</code></td>\n</tr>\n<tr>\n<td>Dropout</td>\n<td>🟢</td>\n<td><code>Network Layer</code></td>\n</tr>\n<tr>\n<td>Pooling</td>\n<td>🟢</td>\n<td><code>Network Layer</code></td>\n</tr>\n</tbody>\n</table>\n<h1>2 | Starting 🌱</h1>\n<p>We know that <code>CNNs</code> are showing great results in the field of <code>Vision</code> related tasks</p>\n<ul>\n<li><strong>Image Classification</strong> - CNNs have been used to <code>achieve accuracies</code> of over 99% on <code>image classification</code> tasks.</li>\n<li><strong>Object Detection</strong> - CNNs have been used to <code>achieve accuracies</code> of over 90% on <code>object detection</code> tasks</li>\n<li><strong>Semantic Segmentation</strong> - CNNs have been used to <code>achieve accuracies</code> of over 80% on <code>semantic segmentation</code> tasks</li>\n</ul>\n<p>While <code>CNNs</code> have already <code>existed for a long time</code>, their <code>success was limited</code> due to the</p>\n<ul>\n<li>size of the available training sets</li>\n<li>size of the considered networks.</li>\n</ul>\n<p>The <code>normal classification</code> task was actually <code>pretty easy</code> for the network to learn. But as the field of <code>Image Segmentation</code>, increased, it became <code>difficult</code> for the model to <code>learn both the class label and position/localization</code> .<code>Localization</code> is a <code>challenging task</code>, because it requires the model to <code>learn both the appearance</code> of the <code>objects/features</code> of interest, and their <code>spatial location</code> in the image.</p>\n<p>The major use of <code>Image Semantic Segmentation</code> is in <code>Bio-Medical'. In</code>Bio-Medical<code>image processing, the</code>desired output<code>is not just a</code>classification of the image<code>, but also the</code>location of the objects` or features of interest in the image</p>\n<p>For example, in <code>Bio-Medical</code> image processing, it is often <code>important</code> to be able to <code>localize tumors</code> or <code>other abnormalities</code> in <code>images of tissues</code> or <code>organs</code>. This is because the <code>location of these abnormalities</code> can be used to <code>diagnose diseases</code> or to <code>plan treatments</code>.</p>\n<h1>3 | Inspiration 💡</h1>\n<p><strong><a href=\"https://papers.nips.cc/paper_files/paper/2012/file/459a4ddcb586f24efd9395aa7662bc7c-Paper.pdf\" target=\"_blank\">Deep Neural Networks Segment Neuronal Membranes in Electron Microscopy Images</a></strong> used the <code>method of masking to predict the correct labels</code> of the image</p>\n<p><code>Ciresan</code> model had many advantages such as</p>\n<ul>\n<li>Localization - The network was able to identify the location of an object in an image.</li>\n<li>Larger Training Data - the <code>number of patches</code> that was extracted from a <code>training image</code> was <code>much larger than the number of training images themselves</code>.</li>\n</ul>\n<p>This was because a <code>patch</code> is a <code>small</code>, <code>rectangular section</code> of an image. For example, a patch of size 64x64 can be extracted from any image. If we have a training set of 100 images, we can extract a total of 100x64x64 = 32,768,000 <code>patches</code>.</p>\n<p>But the technique used by <code>Ciresan</code> had some <code>drawbacks</code></p>\n<ul>\n<li>Long Training Time - The model was <code>quite slow</code> because the <code>network must be run separately for each patch</code>, which means that the model <code>must be run multiple times for each image</code>. This can be very slow, especially for large images.</li>\n<li>Redundancy -  When an <code>image</code> is <code>divided into patches</code>, there is a lot of <code>overlap between the patches</code>. This is because the <code>patches are typically small</code>, and they are <code>often arranged in a regular grid</code>.</li>\n<li>Localization-Context Trade-Off - The problem of <code>balancing the accuracy of localization</code> with the <code>amount of context</code> .</li>\n<li>* <code>Larger patches</code> allow for more <code>accurate localization</code>, but they <code>also reduce the amount of context</code> that is available to the network. This is because <code>max-pooling</code> layers are typically used to <code>down-sample the image</code>.</li>\n<li>* <code>Smaller patches</code>, allow for <code>more context</code>, but they also <code>reduce the accuracy of localization</code>. This is because <code>smaller patches contain less information</code> about the location of the object.</li>\n</ul>\n<h1>4 | Advantages ✅</h1>\n<ul>\n<li><p>Less Training Data - The model is <code>modified</code> such that it works with <code>very few training images</code> and yields <code>more precise segmentations</code>.</p></li>\n<li><p>Touching Objects - <code>weighted loss</code> is applied, where the <code>separating background labels between touching cells</code> obtain a <code>large weight</code> in the loss function.</p></li>\n</ul>\n<h1>5 | Parameters and Formulations 🧬</h1>\n<h2>5.1 | Weights Intialization</h2>\n<p>The weights are intialized such that each feature map in the network has approximately unit variance. This was achieved by intializing the weights by<br>\n a <code>Gaussian Distribution(1 , sqrt(2/N))</code></p>\n<h2>5.1 | Optimizer</h2>\n<p>The <code>Adam optimizer</code> is used to train the $UNet$ model, and the <code>momentum parameter</code> is set to 0.99.</p>\n<p>The <code>Adam optimizer</code> is a <code>Stochastic</code> <code>Gradient</code> <code>Descent</code> (SGD) <code>optimizer</code> that was introduced in 2014. The <code>momentum parameter</code> controls <code>how much the optimizer remembers its past gradients</code>. Though, a momentum value of 0.99 is a <code>relatively high value</code>. This means that the Adam <code>optimizer</code> will be more likely to <code>follow its past gradients</code>, which can <code>help to improve the stability of training</code> and <code>prevent the model from overfitting</code>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2F55fdd37178b63612afb111403208297d%2FOptimizer.png?generation=1688521271597503&amp;alt=media\"></p>\n<h2>5.2 | Prediction</h2>\n<p>The <code>energy function</code> is computed by a <code>pixel-wise soft-max</code> over the final feature map</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd5ba65b1f7d31619da7823a977567643%2FPredictioon.png?generation=1688521293363603&amp;alt=media\"></p>\n<p>![]()</p>\n<h3>Where</h3>\n<ul>\n<li>a_k(x) =  Feature Channel k</li>\n<li>X 𝛜 Ω = Pixel Position</li>\n<li>K = Number Of Labels</li>\n</ul>\n<h3>Assumptions</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Faa1e84d19521a53ab934ba4562eb0946%2FAssumpetions.png?generation=1688521344912976&amp;alt=media\"></p>\n<h2>5.3 | Loss</h2>\n<p>The cross entropy then penalizes at each position the deviation of p^`_x (x) from 1</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd7ce2dae61705f5d1c91310a3e7844bd%2FLos.png?generation=1688521366011103&amp;alt=media\"></p>\n<h1>6 | Experiments 🥼</h1>\n<p>The network is tested on 2 different <code>segmentation tasks</code></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Error</th>\n<th>Accuracy</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Neural Structures in Electron Microscopic Recordings</td>\n<td>0.0003529</td>\n<td>0.9203</td>\n</tr>\n<tr>\n<td>Light Microscopic Images</td>\n<td></td>\n<td>92%</td>\n</tr>\n</tbody>\n</table>\n<h1>8 | Acknowledgements 🤝</h1>\n<p>This study was supported by the Excellence Initiative of the German Federal and State governments (EXC 294) and by the BMBF (Fkz 0316185B).</p>\n<h1>9 | Refrences 🔗</h1>\n<table>\n<thead>\n<tr>\n<th>Paper</th>\n<th>Authors</th>\n<th>Place</th>\n<th>Year</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Deep neural net-works segment neuronal membranes in electron microscopy images</td>\n<td><code>Ciresan</code> <code>D.C.</code> <code>Gambardella</code> <code>L.M.</code> <code>Giusti A.</code> <code>Schmidhuber J.</code></td>\n<td>NIPS. pp. 2852 2860</td>\n<td>2012</td>\n</tr>\n<tr>\n<td>Discriminative un-supervised feature learning with convolutional neural networks</td>\n<td><code>Dosovitskiy</code> <code>A.</code> <code>Springenberg</code> <code>J.T.</code> <code>Riedmiller</code> <code>M.</code> <code>Brox</code> <code>T.</code></td>\n<td>NIPS</td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Rich feature hierarchies for ac-curate object detection and semantic segmentation.</td>\n<td><code>Girshick</code> <code>R.</code> <code>Donahue</code> <code>J.</code> <code>Darrell</code> <code>T.</code> <code>Malik</code> <code>J.</code></td>\n<td>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Hyper-columns for object segmentation and  ne-grained localization</td>\n<td><code>Hariharan</code> <code>B.</code> <code>Arbelez</code> <code>P.</code> <code>Girshick</code> <code>R.</code> <code>Malik</code> <code>J.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Delving deep into rectifiers: Surpassing human-level performance on image-net classification</td>\n<td><code>He</code>, <code>K.</code>, <code>Zhang</code>, <code>X.</code>, <code>Ren</code>, <code>S.</code>, <code>Sun</code>, <code>J.</code></td>\n<td></td>\n<td>2015</td>\n</tr>\n<tr>\n<td>Convolutional architecture for fast feature embedding</td>\n<td><code>Jia</code> <code>Y.</code> <code>Shelhamer</code> <code>E.</code> <code>Donahue</code> <code>J.</code> <code>Karayev</code> <code>S.</code> <code>Long</code> <code>J.</code> <code>Girshick</code> <code>R.</code> <code>Guadar-rama</code> <code>S.</code> <code>Darrell</code> <code>T.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Image-net classification with deep convolutional neural networks</td>\n<td><code>Krizhevsky</code> <code>A.</code> <code>Sutskever</code> <code>I.</code> <code>Hinton</code> <code>G.E.</code></td>\n<td>NIPS. pp. 1106 1114</td>\n<td>2012</td>\n</tr>\n<tr>\n<td>Backpropagation applied to handwritten zip code recognition. Neural Computation</td>\n<td><code>LeCun</code> <code>Y.</code> <code>Boser</code> <code>B.</code> <code>Denker</code> <code>J.S.</code> <code>Henderson</code> <code>D.</code> <code>Howard</code> <code>R.E.</code> <code>Hubbard</code> <code>W.</code> <code>Jackel</code> <code>L.D.</code></td>\n<td>1(4) 541 551</td>\n<td>1989</td>\n</tr>\n<tr>\n<td>Fully convolutional networks for semantic segmentation</td>\n<td><code>Long</code> <code>J.</code> <code>Shelhamer</code> <code>E.</code> <code>Darrell</code> <code>T.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td>A benchmark for comparison of cell tracking algorithms</td>\n<td><code>Maska</code> <code>M.</code> <code>(...)</code> <code>de Solorzano</code> <code>C.O.</code></td>\n<td>Bioinformatics 30, 1609 1617</td>\n<td>2014</td>\n</tr>\n<tr>\n<td>Very deep convolutional networks for large-scale image recognition</td>\n<td><code>Simonyan</code> <code>K.</code> <code>Zisserman</code> <code>A.</code></td>\n<td></td>\n<td>2014</td>\n</tr>\n<tr>\n<td><a href=\"http://www.codesolorzano.com/celltrackingchallenge/Cell_Tracking_Challenge/Welcome.html\" target=\"_blank\">Web page of the cell tracking challenge</a></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><a href=\"http://brainiac2.mit.edu/isbi_challenge/\" target=\"_blank\">Web page of the em segmentation challenge</a></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# 1 | Pre-Reqeuistics 🎓\n\n|||\n|---|---\n|🔴|`High Level Recommended`\n|🟡|`Medium Level Recommended`\n|🟢|`Low Level Recommended`\n\n|Concept||Usage|\n|---|---|---\n|Convolution|🔴|`Network Layer`\n|Convolution Transpose|🔴|`Network Layer`\n|Concatenation|🔴|`Network Layer`\n|Adam|🔴|`Optimizer`\n|Padding|🔴|`Runtime Processing`\n|Strides|🔴|`Runtime Processing`\n|ReLU|🟡|`Activation Function`\n|Sigmoid|🟡|`Activation Function`\n|Binary Cross Entropy|🟡|`Loss`\n|Dropout|🟢|`Network Layer`\n|Pooling|🟢|`Network Layer`\n\n# 2 | Starting 🌱\n\nWe know that `CNNs` are showing great results in the field of `Vision` related tasks\n\n* **Image Classification** - CNNs have been used to `achieve accuracies` of over 99% on `image classification` tasks.\n* **Object Detection** - CNNs have been used to `achieve accuracies` of over 90% on `object detection` tasks\n* **Semantic Segmentation** - CNNs have been used to `achieve accuracies` of over 80% on `semantic segmentation` tasks\n\nWhile `CNNs` have already `existed for a long time`, their `success was limited` due to the\n* size of the available training sets\n* size of the considered networks.\n\nThe `normal classification` task was actually `pretty easy` for the network to learn. But as the field of `Image Segmentation`, increased, it became `difficult` for the model to `learn both the class label and position/localization` .`Localization` is a `challenging task`, because it requires the model to `learn both the appearance` of the `objects/features` of interest, and their `spatial location` in the image.\n\nThe major use of `Image Semantic Segmentation` is in `Bio-Medical'. In `Bio-Medical` image processing, the `desired output` is not just a `classification of the image`, but also the `location of the objects` or features of interest in the image\n\nFor example, in `Bio-Medical` image processing, it is often `important` to be able to `localize tumors` or `other abnormalities` in `images of tissues` or `organs`. This is because the `location of these abnormalities` can be used to `diagnose diseases` or to `plan treatments`.\n\n# 3 | Inspiration 💡\n\n**[Deep Neural Networks Segment Neuronal Membranes in Electron Microscopy Images](https://papers.nips.cc/paper_files/paper/2012/file/459a4ddcb586f24efd9395aa7662bc7c-Paper.pdf)** used the `method of masking to predict the correct labels` of the image\n\n`Ciresan` model had many advantages such as\n* Localization - The network was able to identify the location of an object in an image.\n* Larger Training Data - the `number of patches` that was extracted from a `training image` was `much larger than the number of training images themselves`.\n\nThis was because a `patch` is a `small`, `rectangular section` of an image. For example, a patch of size 64x64 can be extracted from any image. If we have a training set of 100 images, we can extract a total of 100x64x64 = 32,768,000 `patches`.\n\nBut the technique used by `Ciresan` had some `drawbacks`\n* Long Training Time - The model was `quite slow` because the `network must be run separately for each patch`, which means that the model `must be run multiple times for each image`. This can be very slow, especially for large images.\n* Redundancy -  When an `image` is `divided into patches`, there is a lot of `overlap between the patches`. This is because the `patches are typically small`, and they are `often arranged in a regular grid`.\n* Localization-Context Trade-Off - The problem of `balancing the accuracy of localization` with the `amount of context` .\n* * `Larger patches` allow for more `accurate localization`, but they `also reduce the amount of context` that is available to the network. This is because `max-pooling` layers are typically used to `down-sample the image`.\n* * `Smaller patches`, allow for `more context`, but they also `reduce the accuracy of localization`. This is because `smaller patches contain less information` about the location of the object.\n\n# 4 | Advantages ✅ \n\n* Less Training Data - The model is `modified` such that it works with `very few training images` and yields `more precise segmentations`.\n\n* Touching Objects - `weighted loss` is applied, where the `separating background labels between touching cells` obtain a `large weight` in the loss function.\n\n# 5 | Parameters and Formulations 🧬 \n\n## 5.1 | Weights Intialization\n\nThe weights are intialized such that each feature map in the network has approximately unit variance. This was achieved by intializing the weights by\n a `Gaussian Distribution(1 , sqrt(2/N))`\n\n## 5.1 | Optimizer\n\nThe `Adam optimizer` is used to train the $UNet$ model, and the `momentum parameter` is set to 0.99.\n\nThe `Adam optimizer` is a `Stochastic` `Gradient` `Descent` (SGD) `optimizer` that was introduced in 2014. The `momentum parameter` controls `how much the optimizer remembers its past gradients`. Though, a momentum value of 0.99 is a `relatively high value`. This means that the Adam `optimizer` will be more likely to `follow its past gradients`, which can `help to improve the stability of training` and `prevent the model from overfitting`.\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2F55fdd37178b63612afb111403208297d%2FOptimizer.png?generation=1688521271597503&alt=media\">\n\n## 5.2 | Prediction\n\nThe `energy function` is computed by a `pixel-wise soft-max` over the final feature map\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd5ba65b1f7d31619da7823a977567643%2FPredictioon.png?generation=1688521293363603&alt=media\">\n\n![]()\n### Where\n\n* a_k(x) =  Feature Channel k\n* X 𝛜 Ω = Pixel Position\n* K = Number Of Labels\n\n### Assumptions\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Faa1e84d19521a53ab934ba4562eb0946%2FAssumpetions.png?generation=1688521344912976&alt=media\">\n\n## 5.3 | Loss\n\nThe cross entropy then penalizes at each position the deviation of p^`_x (x) from 1\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd7ce2dae61705f5d1c91310a3e7844bd%2FLos.png?generation=1688521366011103&alt=media\">\n\n# 6 | Experiments 🥼\n\nThe network is tested on 2 different `segmentation tasks`\n\n||Error|Accuracy\n|---|---|---\n|Neural Structures in Electron Microscopic Recordings|0.0003529|0.9203\n|Light Microscopic Images||92%\n\n# 8 | Acknowledgements 🤝 \n\nThis study was supported by the Excellence Initiative of the German Federal and State governments (EXC 294) and by the BMBF (Fkz 0316185B).\n\n# 9 | Refrences 🔗\n\n|Paper|Authors|Place|Year|\n|---|---|---|\n|Deep neural net-works segment neuronal membranes in electron microscopy images|`Ciresan` `D.C.` `Gambardella` `L.M.` `Giusti A.` `Schmidhuber J.`|NIPS. pp. 2852 2860|2012\n|Discriminative un-supervised feature learning with convolutional neural networks|`Dosovitskiy` `A.` `Springenberg` `J.T.` `Riedmiller` `M.` `Brox` `T.`|NIPS|2014\n|Rich feature hierarchies for ac-curate object detection and semantic segmentation.|`Girshick` `R.` `Donahue` `J.` `Darrell` `T.` `Malik` `J.`| Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)|2014\n|Hyper-columns for object segmentation and  ne-grained localization |`Hariharan` `B.` `Arbelez` `P.` `Girshick` `R.` `Malik` `J.`||2014\n|Delving deep into rectifiers: Surpassing human-level performance on image-net classification|`He`, `K.`, `Zhang`, `X.`, `Ren`, `S.`, `Sun`, `J.`||2015\n|Convolutional architecture for fast feature embedding |`Jia` `Y.` `Shelhamer` `E.` `Donahue` `J.` `Karayev` `S.` `Long` `J.` `Girshick` `R.` `Guadar-rama` `S.` `Darrell` `T.`||2014\n|Image-net classification with deep convolutional neural networks|`Krizhevsky` `A.` `Sutskever` `I.` `Hinton` `G.E.`|NIPS. pp. 1106 1114|2012\n|Backpropagation applied to handwritten zip code recognition. Neural Computation|`LeCun` `Y.` `Boser` `B.` `Denker` `J.S.` `Henderson` `D.` `Howard` `R.E.` `Hubbard` `W.` `Jackel` `L.D.`|1(4) 541 551|1989\n|Fully convolutional networks for semantic segmentation |`Long` `J.` `Shelhamer` `E.` `Darrell` `T.`||2014\n|A benchmark for comparison of cell tracking algorithms |`Maska` `M.` `(...)` `de Solorzano` `C.O.`|Bioinformatics 30, 1609 1617|2014|Image segmentation with cascaded hierarchical models and logistic disjunctive normal networks.|`Seyedhosseini` `M.` `Sajjadi` `M.` `Tasdizen` `T.`|Computer Vision (ICCV), 2013 IEEE International Conference on. pp.|2013\n|Very deep convolutional networks for large-scale image recognition|`Simonyan` `K.` `Zisserman` `A.`||2014\n|[Web page of the cell tracking challenge](http://www.codesolorzano.com/celltrackingchallenge/Cell_Tracking_Challenge/Welcome.html)\n|[Web page of the em segmentation challenge](http://brainiac2.mit.edu/isbi_challenge/)",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2330418": "# 1 | Pre-Reqeuistics 🎓\n\n|||\n|---|---\n|🔴|`High Level Recommended`\n|🟡|`Medium Level Recommended`\n|🟢|`Low Level Recommended`\n\n|Concept||Usage|\n|---|---|---\n|Convolution|🔴|`Network Layer`\n|Convolution Transpose|🔴|`Network Layer`\n|Concatenation|🔴|`Network Layer`\n|Adam|🔴|`Optimizer`\n|Padding|🔴|`Runtime Processing`\n|Strides|🔴|`Runtime Processing`\n|ReLU|🟡|`Activation Function`\n|Sigmoid|🟡|`Activation Function`\n|Binary Cross Entropy|🟡|`Loss`\n|Dropout|🟢|`Network Layer`\n|Pooling|🟢|`Network Layer`\n\n# 2 | Starting 🌱\n\nWe know that `CNNs` are showing great results in the field of `Vision` related tasks\n\n* **Image Classification** - CNNs have been used to `achieve accuracies` of over 99% on `image classification` tasks.\n* **Object Detection** - CNNs have been used to `achieve accuracies` of over 90% on `object detection` tasks\n* **Semantic Segmentation** - CNNs have been used to `achieve accuracies` of over 80% on `semantic segmentation` tasks\n\nWhile `CNNs` have already `existed for a long time`, their `success was limited` due to the\n* size of the available training sets\n* size of the considered networks.\n\nThe `normal classification` task was actually `pretty easy` for the network to learn. But as the field of `Image Segmentation`, increased, it became `difficult` for the model to `learn both the class label and position/localization` .`Localization` is a `challenging task`, because it requires the model to `learn both the appearance` of the `objects/features` of interest, and their `spatial location` in the image.\n\nThe major use of `Image Semantic Segmentation` is in `Bio-Medical'. In `Bio-Medical` image processing, the `desired output` is not just a `classification of the image`, but also the `location of the objects` or features of interest in the image\n\nFor example, in `Bio-Medical` image processing, it is often `important` to be able to `localize tumors` or `other abnormalities` in `images of tissues` or `organs`. This is because the `location of these abnormalities` can be used to `diagnose diseases` or to `plan treatments`.\n\n# 3 | Inspiration 💡\n\n**[Deep Neural Networks Segment Neuronal Membranes in Electron Microscopy Images](https://papers.nips.cc/paper_files/paper/2012/file/459a4ddcb586f24efd9395aa7662bc7c-Paper.pdf)** used the `method of masking to predict the correct labels` of the image\n\n`Ciresan` model had many advantages such as\n* Localization - The network was able to identify the location of an object in an image.\n* Larger Training Data - the `number of patches` that was extracted from a `training image` was `much larger than the number of training images themselves`.\n\nThis was because a `patch` is a `small`, `rectangular section` of an image. For example, a patch of size 64x64 can be extracted from any image. If we have a training set of 100 images, we can extract a total of 100x64x64 = 32,768,000 `patches`.\n\nBut the technique used by `Ciresan` had some `drawbacks`\n* Long Training Time - The model was `quite slow` because the `network must be run separately for each patch`, which means that the model `must be run multiple times for each image`. This can be very slow, especially for large images.\n* Redundancy -  When an `image` is `divided into patches`, there is a lot of `overlap between the patches`. This is because the `patches are typically small`, and they are `often arranged in a regular grid`.\n* Localization-Context Trade-Off - The problem of `balancing the accuracy of localization` with the `amount of context` .\n* * `Larger patches` allow for more `accurate localization`, but they `also reduce the amount of context` that is available to the network. This is because `max-pooling` layers are typically used to `down-sample the image`.\n* * `Smaller patches`, allow for `more context`, but they also `reduce the accuracy of localization`. This is because `smaller patches contain less information` about the location of the object.\n\n# 4 | Advantages ✅ \n\n* Less Training Data - The model is `modified` such that it works with `very few training images` and yields `more precise segmentations`.\n\n* Touching Objects - `weighted loss` is applied, where the `separating background labels between touching cells` obtain a `large weight` in the loss function.\n\n# 5 | Parameters and Formulations 🧬 \n\n## 5.1 | Weights Intialization\n\nThe weights are intialized such that each feature map in the network has approximately unit variance. This was achieved by intializing the weights by\n a `Gaussian Distribution(1 , sqrt(2/N))`\n\n## 5.1 | Optimizer\n\nThe `Adam optimizer` is used to train the $UNet$ model, and the `momentum parameter` is set to 0.99.\n\nThe `Adam optimizer` is a `Stochastic` `Gradient` `Descent` (SGD) `optimizer` that was introduced in 2014. The `momentum parameter` controls `how much the optimizer remembers its past gradients`. Though, a momentum value of 0.99 is a `relatively high value`. This means that the Adam `optimizer` will be more likely to `follow its past gradients`, which can `help to improve the stability of training` and `prevent the model from overfitting`.\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2F55fdd37178b63612afb111403208297d%2FOptimizer.png?generation=1688521271597503&alt=media\">\n\n## 5.2 | Prediction\n\nThe `energy function` is computed by a `pixel-wise soft-max` over the final feature map\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd5ba65b1f7d31619da7823a977567643%2FPredictioon.png?generation=1688521293363603&alt=media\">\n\n![]()\n### Where\n\n* a_k(x) =  Feature Channel k\n* X 𝛜 Ω = Pixel Position\n* K = Number Of Labels\n\n### Assumptions\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Faa1e84d19521a53ab934ba4562eb0946%2FAssumpetions.png?generation=1688521344912976&alt=media\">\n\n## 5.3 | Loss\n\nThe cross entropy then penalizes at each position the deviation of p^`_x (x) from 1\n\n<img src = \"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9711370%2Fd7ce2dae61705f5d1c91310a3e7844bd%2FLos.png?generation=1688521366011103&alt=media\">\n\n# 6 | Experiments 🥼\n\nThe network is tested on 2 different `segmentation tasks`\n\n||Error|Accuracy\n|---|---|---\n|Neural Structures in Electron Microscopic Recordings|0.0003529|0.9203\n|Light Microscopic Images||92%\n\n# 8 | Acknowledgements 🤝 \n\nThis study was supported by the Excellence Initiative of the German Federal and State governments (EXC 294) and by the BMBF (Fkz 0316185B).\n\n# 9 | Refrences 🔗\n\n|Paper|Authors|Place|Year|\n|---|---|---|\n|Deep neural net-works segment neuronal membranes in electron microscopy images|`Ciresan` `D.C.` `Gambardella` `L.M.` `Giusti A.` `Schmidhuber J.`|NIPS. pp. 2852 2860|2012\n|Discriminative un-supervised feature learning with convolutional neural networks|`Dosovitskiy` `A.` `Springenberg` `J.T.` `Riedmiller` `M.` `Brox` `T.`|NIPS|2014\n|Rich feature hierarchies for ac-curate object detection and semantic segmentation.|`Girshick` `R.` `Donahue` `J.` `Darrell` `T.` `Malik` `J.`| Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)|2014\n|Hyper-columns for object segmentation and  ne-grained localization |`Hariharan` `B.` `Arbelez` `P.` `Girshick` `R.` `Malik` `J.`||2014\n|Delving deep into rectifiers: Surpassing human-level performance on image-net classification|`He`, `K.`, `Zhang`, `X.`, `Ren`, `S.`, `Sun`, `J.`||2015\n|Convolutional architecture for fast feature embedding |`Jia` `Y.` `Shelhamer` `E.` `Donahue` `J.` `Karayev` `S.` `Long` `J.` `Girshick` `R.` `Guadar-rama` `S.` `Darrell` `T.`||2014\n|Image-net classification with deep convolutional neural networks|`Krizhevsky` `A.` `Sutskever` `I.` `Hinton` `G.E.`|NIPS. pp. 1106 1114|2012\n|Backpropagation applied to handwritten zip code recognition. Neural Computation|`LeCun` `Y.` `Boser` `B.` `Denker` `J.S.` `Henderson` `D.` `Howard` `R.E.` `Hubbard` `W.` `Jackel` `L.D.`|1(4) 541 551|1989\n|Fully convolutional networks for semantic segmentation |`Long` `J.` `Shelhamer` `E.` `Darrell` `T.`||2014\n|A benchmark for comparison of cell tracking algorithms |`Maska` `M.` `(...)` `de Solorzano` `C.O.`|Bioinformatics 30, 1609 1617|2014|Image segmentation with cascaded hierarchical models and logistic disjunctive normal networks.|`Seyedhosseini` `M.` `Sajjadi` `M.` `Tasdizen` `T.`|Computer Vision (ICCV), 2013 IEEE International Conference on. pp.|2013\n|Very deep convolutional networks for large-scale image recognition|`Simonyan` `K.` `Zisserman` `A.`||2014\n|[Web page of the cell tracking challenge](http://www.codesolorzano.com/celltrackingchallenge/Cell_Tracking_Challenge/Welcome.html)\n|[Web page of the em segmentation challenge](http://brainiac2.mit.edu/isbi_challenge/)"
  }
}