{
  "id": 541067,
  "title": "75th Place Solution",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/writeups/brightskies-75th-place-solution",
  "author_name": "",
  "post_date": "2024-10-17T11:28:41.480883400Z",
  "votes": 15,
  "comment_count": 8,
  "views": 0,
  "content": "<p>It was a challenging and exciting journey, and the high level of competition truly pushed everyone to excel. First, I want to thank my teammates, <a href=\"https://www.kaggle.com/hossamelkady99999\" target=\"_blank\">@hossamelkady99999</a>, <a href=\"https://www.kaggle.com/aymanelghotni\" target=\"_blank\">@aymanelghotni</a>, <a href=\"https://www.kaggle.com/youssefhany10\" target=\"_blank\">@youssefhany10</a>, <a href=\"https://www.kaggle.com/amgedelmasry001\" target=\"_blank\">@amgedelmasry001</a>  who showed great effort throughout the competition. I also want to extend my thanks to all contributers for their invaluable contributions, support, and inspiration throughout the competition. Additionally, a heartfelt thank you to the organizers for putting together such a well-structured and engaging competition. Your efforts made this a great learning experience for all involved. </p>\n<h2>Overview of the Approach</h2>\n<h3>Data Augmentations</h3>\n<p>Thanks to <a href=\"https://www.kaggle.com/itsuki9180\" target=\"_blank\">@itsuki9180</a> for providing baseline to start from and submission notebook</p>\n<ul>\n<li>Brightness and Contrast Adjustment: Randomly adjusts the brightness and contrast within a limit of ±0.2 </li>\n<li>Blurring and Noise: Applies one of the following transformations: </li>\n<li>Motion blur, median blur, or Gaussian blur with a limit of 5. </li>\n<li>Adds Gaussian noise with a variance limit between 5 and 30. </li>\n<li>Geometric Distortions: Applies one of the following transformations: </li>\n<li>Optical distortion with a limit of 1.0. </li>\n<li>Grid distortion with 5 steps and a distortion limit of 1.0. </li>\n<li>Elastic transformation with an alpha of 3. </li>\n<li>Shift, Scale, and Rotate: Randomly shifts, scales, and rotates the image with a shift and scale limit of 10%, and a rotation limit of 15 degrees. </li>\n<li>Coarse Dropout: Randomly removes parts of the image by dropping out up to 3 areas, with a maximum size of 12x12 pixels and a minimum size of 8x8 pixels. </li>\n</ul>\n<h3>Hyperparameters and Training Strategy</h3>\n<ul>\n<li>Optimizer: AdamW </li>\n<li>Learning Rate: 0.0002 </li>\n<li>Weight Decay: 0.01 </li>\n<li>Batch Size: 4 </li>\n<li>Cross-Validations: 5 Folds</li>\n<li>Learning Rate Scheduler: Cosine Annealing with Warm Restarts </li>\n<li>Mixup Regularization: Applied during training </li>\n</ul>\n<h3>Data Preprocessing</h3>\n<p>We started by converting all DICOM images to PNG format, which allowed for more standardized handling across models. For the baseline, we selected <strong>10 evenly spaced slices</strong> from each modality and stacked them across the channel axis, creating a tensor of shape <strong>(bs, 30, 512, 512)</strong>, where bs represents batch size, and 30 comes from the 10 slices per modality (Sagittal T1, Sagittal T2, and Axial). </p>\n<p>For feature extraction, we used <strong>EfficientNet-B5 / EfficientNet-V2</strong> as the backbone, followed by a fully connected (FC) layer to classify into 75 output classes.</p>\n<h3>Multi-Instance Learning (MIL) plus (2.5D)</h3>\n<p>To improve performance, we introduced <strong>Multi-Instance Learning (MIL)</strong>. Instead of concatenating the slices, we preserved them per modality. Thus, for each modality, we extracted 10 slices, leading to an input shape of <strong>(bs, 3, 10, 512, 512)</strong>. The three channels correspond to the Sagittal T1, Sagittal T2, and Axial modalities. </p>\n<p>We experimented with various sequence models, including <strong>Bi-LSTMs</strong>, <strong>GRUs</strong>, and <strong>attention mechanisms</strong>. Among these, the attention mechanism achieved the best results, pushing the public leaderboard score to <strong>0.55%</strong>. </p>\n<h3>YOLO for Preprocessing and Cropping</h3>\n<p>Next, we trained <strong>YOLO v8</strong> to predict specific levels across the modalities (Sagittal T1, Sagittal T2, and Axial) using the provided coordinates. After predicting the coordinates, we cropped the images to <strong>100x100</strong> pixels and stored them in the following structure:</p>\n<pre><code>L_L/ \nL_L/  \n...\nL_S/ \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F682a0a18e2322d86b835784c61163cc9%2Fyolo_l.jpg?generation=1729089212492676&amp;alt=media\" alt=\"\"><br>\nWe then calculated the mean number of slices per modality per level, which averaged around <strong>7 slices</strong>. With this information, we reshaped the data into <strong>(bs, 5, 3, 7, 100, 100)</strong>, where 5 represents the levels, 3 represents the modalities, and 7 represents the slices.</p>\n<h3>Enhancing the Attention Mechanism</h3>\n<p>Initially, we attempted using <strong>1 attention block after the encoder</strong> with an input shape of <strong>(bs*5, 3, n_features)</strong>, but it did not improve the baseline score. We hypothesized that the model might not be receiving enough data per sequence. </p>\n<p>To address this, we augmented the input by adding adjacent slices to each selected slice. For example, if we selected slices at intervals of [0, 3, 6, 9, 12, 15, 18] from a folder with 20 slices, we also included adjacent slices like [2, 4] for slice 3. For edge cases, we padded with zeros. This resulted in an input shape of <strong>(bs, 5, 3, 7, 3, 100, 100)</strong>, where the extra dimension represents the adjacent slices (input channels). </p>\n<p>After the encoder, the shape became <strong>(bs*5, 3x7, n_features)</strong>. By increasing the sequence length by a factor of 7, we saw an improvement in the public leaderboard score to <strong>0.43%</strong>.</p>\n<h3>Separate Encoders for Modalities</h3>\n<p>We also tried separating the encoders for Axial and Sagittal (T1 and T2) modalities. This resulted in two encoders handling the three modalities. Although this approach showed minor improvements in the overall score, it did not have a significant impact on cross-validation (CV) performance. However, this step was necessary for further approaches. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc60bb0557270ad9f11cae5af97c313d5%2Fmodel.png?generation=1729089435504890&amp;alt=media\" alt=\"\"></p>\n<h3>Data Sorting</h3>\n<p>We discovered some issues with data sorting thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for showing how to sort it using IPP and we discovered also incorrect metadata, which included duplicate modalities.We merged the duplicates and corrected the sorting. This refinement resulted in a slight improvement, increasing the score to <strong>0.42%</strong>.</p>\n<h2>Additional Approaches</h2>\n<h3>Regression Models</h3>\n<p>We also trained a <strong>regression</strong> model using the coordinates improved thanks to <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for improving the coordinates data, but this model did not perform better than YOLO. The results were comparable to those achieved by YOLO v8. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fdc481eae1eeba7cb7e2066ec3e547909%2Fpred_coord.png?generation=1729090680548927&amp;alt=media\" alt=\"\"></p>\n<h3>Segmentation of Vertebra</h3>\n<p>We attempted to segment the vertebrae and stack it over the channel, but this approach led to worse results than our best submission. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F1f1d8ec89f54f98e2e34389bf473a352%2F7%201.png?generation=1729090606914059&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F00d991945776fb148e6fb611a4231ec2%2F7_seg%201.png?generation=1729090627547931&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ff55f572fdb47602f4bf147ebf6d0713a%2F1%20(1).png?generation=1729090647580937&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F66d92f87e6873a0cfe45a185fed5c8b0%2F1_seg.png?generation=1729090660449277&amp;alt=media\" alt=\"\"></p>\n<h3>Disease Prediction with YOLO</h3>\n<p>Another approach involved training YOLO to predict disease coordinates. We used the coordinates to extract a <strong>50x50</strong> crop around the predicted disease location. Additionally, we projected the disease's <strong>x, y coordinates</strong> into the other two modalities (T1, T2) thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for showing how projection could be done, and applied the same cropping strategy.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F2abcc472f73773f30ff6a3eae6506155%2FYolo_disease%201.jpg?generation=1729090703285103&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F95c2adfaf454f4383a11be509831a7bd%2Fpr_l1.png?generation=1729090716063007&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F6abdea05ad39c2a4ae043aa0562961d7%2Fpr_L4.png?generation=1729090728003675&amp;alt=media\" alt=\"\"></p>\n<h3>Cascaded Models</h3>\n<p>We experimented with a cascaded model approach: </p>\n<ul>\n<li><p>The first model predicted <strong>Normal vs. Abnormal</strong> conditions, trained with Binary Cross-Entropy (BCE) loss. </p></li>\n<li><p>The second model predicted Abnormalities <strong>Moderate vs. Severe</strong> conditions. </p></li>\n</ul>\n<p>Each model outputs a value between 0 and 1. We then concatenated the outputs to get the final prediction. We tried two approaches: </p>\n<ul>\n<li><p>Using probability-based evaluation, which achieved a local score of <strong>38%</strong>. </p></li>\n<li><p>Using a deep fully connected (FC) layer to directly predict the final probabilities, yielding similar results. </p></li>\n<li><p>Unfortunately, neither of these approaches performed well on the public or private leaderboard. </p></li>\n</ul>\n<h2>ORCHA: Our Experiment Tracking and Automation Tool</h2>\n<p>Throughout this competition, we used an internal tool that’s currently under development, which was pivotal in streamlining our workflow. This tool helped us manage experiments, automate repetitive tasks, and utilize resources efficiently, allowing us to focus more on refining model performance.</p>\n<p>Key aspects of this tool that enhanced our process include:</p>\n<ul>\n<li>Experiment Tracking: It allowed us to log and organize experiment data systematically, making it easier to reference previous runs and assess the impact of hyperparameter changes or architecture tweaks.</li>\n<li>Automation: By automating routine tasks like data preprocessing and model training, we could focus on optimizing our models without getting bogged down by manual tasks.</li>\n<li>Execution Efficiency: The tool managed resources autonomously, handling tasks such as training and evaluation in the background, ensuring optimal use of computational resources without constant user input.</li>\n<li>Customization: It was highly adaptable, which made it easier to tailor the tool to fit the specific needs of different competitions and use cases.</li>\n</ul>\n<p>Although still in development, ORCHA has been instrumental in improving our efficiency and enabling us to iterate faster. We recommend that others consider developing similar tools tailored to their specific workflows, as even early-stage solutions can significantly enhance productivity and simplify complex tasks.</p>\n<p>For more information, feel free to contact us at ayman.elghotni@brightskiesinc.com.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ffa5827eb280d44ce370206b6f932ec4a%2Ftr1.png?generation=1729090862548327&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F584343bc84ff675d884de4207e00949a%2Ftr2.png?generation=1729090872614220&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc9b56b52f6ecf9fcd9362dd314009271%2Ft_p.png?generation=1729088297606266&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3020300",
      "postDate": "10/17/2024 11:28:41",
      "content": "<p>It was a challenging and exciting journey, and the high level of competition truly pushed everyone to excel. First, I want to thank my teammates, <a href=\"https://www.kaggle.com/hossamelkady99999\" target=\"_blank\">@hossamelkady99999</a>, <a href=\"https://www.kaggle.com/aymanelghotni\" target=\"_blank\">@aymanelghotni</a>, <a href=\"https://www.kaggle.com/youssefhany10\" target=\"_blank\">@youssefhany10</a>, <a href=\"https://www.kaggle.com/amgedelmasry001\" target=\"_blank\">@amgedelmasry001</a>  who showed great effort throughout the competition. I also want to extend my thanks to all contributers for their invaluable contributions, support, and inspiration throughout the competition. Additionally, a heartfelt thank you to the organizers for putting together such a well-structured and engaging competition. Your efforts made this a great learning experience for all involved. </p>\n<h2>Overview of the Approach</h2>\n<h3>Data Augmentations</h3>\n<p>Thanks to <a href=\"https://www.kaggle.com/itsuki9180\" target=\"_blank\">@itsuki9180</a> for providing baseline to start from and submission notebook</p>\n<ul>\n<li>Brightness and Contrast Adjustment: Randomly adjusts the brightness and contrast within a limit of ±0.2 </li>\n<li>Blurring and Noise: Applies one of the following transformations: </li>\n<li>Motion blur, median blur, or Gaussian blur with a limit of 5. </li>\n<li>Adds Gaussian noise with a variance limit between 5 and 30. </li>\n<li>Geometric Distortions: Applies one of the following transformations: </li>\n<li>Optical distortion with a limit of 1.0. </li>\n<li>Grid distortion with 5 steps and a distortion limit of 1.0. </li>\n<li>Elastic transformation with an alpha of 3. </li>\n<li>Shift, Scale, and Rotate: Randomly shifts, scales, and rotates the image with a shift and scale limit of 10%, and a rotation limit of 15 degrees. </li>\n<li>Coarse Dropout: Randomly removes parts of the image by dropping out up to 3 areas, with a maximum size of 12x12 pixels and a minimum size of 8x8 pixels. </li>\n</ul>\n<h3>Hyperparameters and Training Strategy</h3>\n<ul>\n<li>Optimizer: AdamW </li>\n<li>Learning Rate: 0.0002 </li>\n<li>Weight Decay: 0.01 </li>\n<li>Batch Size: 4 </li>\n<li>Cross-Validations: 5 Folds</li>\n<li>Learning Rate Scheduler: Cosine Annealing with Warm Restarts </li>\n<li>Mixup Regularization: Applied during training </li>\n</ul>\n<h3>Data Preprocessing</h3>\n<p>We started by converting all DICOM images to PNG format, which allowed for more standardized handling across models. For the baseline, we selected <strong>10 evenly spaced slices</strong> from each modality and stacked them across the channel axis, creating a tensor of shape <strong>(bs, 30, 512, 512)</strong>, where bs represents batch size, and 30 comes from the 10 slices per modality (Sagittal T1, Sagittal T2, and Axial). </p>\n<p>For feature extraction, we used <strong>EfficientNet-B5 / EfficientNet-V2</strong> as the backbone, followed by a fully connected (FC) layer to classify into 75 output classes.</p>\n<h3>Multi-Instance Learning (MIL) plus (2.5D)</h3>\n<p>To improve performance, we introduced <strong>Multi-Instance Learning (MIL)</strong>. Instead of concatenating the slices, we preserved them per modality. Thus, for each modality, we extracted 10 slices, leading to an input shape of <strong>(bs, 3, 10, 512, 512)</strong>. The three channels correspond to the Sagittal T1, Sagittal T2, and Axial modalities. </p>\n<p>We experimented with various sequence models, including <strong>Bi-LSTMs</strong>, <strong>GRUs</strong>, and <strong>attention mechanisms</strong>. Among these, the attention mechanism achieved the best results, pushing the public leaderboard score to <strong>0.55%</strong>. </p>\n<h3>YOLO for Preprocessing and Cropping</h3>\n<p>Next, we trained <strong>YOLO v8</strong> to predict specific levels across the modalities (Sagittal T1, Sagittal T2, and Axial) using the provided coordinates. After predicting the coordinates, we cropped the images to <strong>100x100</strong> pixels and stored them in the following structure:</p>\n<pre><code>L_L/ \nL_L/  \n...\nL_S/ \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F682a0a18e2322d86b835784c61163cc9%2Fyolo_l.jpg?generation=1729089212492676&amp;alt=media\" alt=\"\"><br>\nWe then calculated the mean number of slices per modality per level, which averaged around <strong>7 slices</strong>. With this information, we reshaped the data into <strong>(bs, 5, 3, 7, 100, 100)</strong>, where 5 represents the levels, 3 represents the modalities, and 7 represents the slices.</p>\n<h3>Enhancing the Attention Mechanism</h3>\n<p>Initially, we attempted using <strong>1 attention block after the encoder</strong> with an input shape of <strong>(bs*5, 3, n_features)</strong>, but it did not improve the baseline score. We hypothesized that the model might not be receiving enough data per sequence. </p>\n<p>To address this, we augmented the input by adding adjacent slices to each selected slice. For example, if we selected slices at intervals of [0, 3, 6, 9, 12, 15, 18] from a folder with 20 slices, we also included adjacent slices like [2, 4] for slice 3. For edge cases, we padded with zeros. This resulted in an input shape of <strong>(bs, 5, 3, 7, 3, 100, 100)</strong>, where the extra dimension represents the adjacent slices (input channels). </p>\n<p>After the encoder, the shape became <strong>(bs*5, 3x7, n_features)</strong>. By increasing the sequence length by a factor of 7, we saw an improvement in the public leaderboard score to <strong>0.43%</strong>.</p>\n<h3>Separate Encoders for Modalities</h3>\n<p>We also tried separating the encoders for Axial and Sagittal (T1 and T2) modalities. This resulted in two encoders handling the three modalities. Although this approach showed minor improvements in the overall score, it did not have a significant impact on cross-validation (CV) performance. However, this step was necessary for further approaches. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc60bb0557270ad9f11cae5af97c313d5%2Fmodel.png?generation=1729089435504890&amp;alt=media\" alt=\"\"></p>\n<h3>Data Sorting</h3>\n<p>We discovered some issues with data sorting thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for showing how to sort it using IPP and we discovered also incorrect metadata, which included duplicate modalities.We merged the duplicates and corrected the sorting. This refinement resulted in a slight improvement, increasing the score to <strong>0.42%</strong>.</p>\n<h2>Additional Approaches</h2>\n<h3>Regression Models</h3>\n<p>We also trained a <strong>regression</strong> model using the coordinates improved thanks to <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for improving the coordinates data, but this model did not perform better than YOLO. The results were comparable to those achieved by YOLO v8. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fdc481eae1eeba7cb7e2066ec3e547909%2Fpred_coord.png?generation=1729090680548927&amp;alt=media\" alt=\"\"></p>\n<h3>Segmentation of Vertebra</h3>\n<p>We attempted to segment the vertebrae and stack it over the channel, but this approach led to worse results than our best submission. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F1f1d8ec89f54f98e2e34389bf473a352%2F7%201.png?generation=1729090606914059&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F00d991945776fb148e6fb611a4231ec2%2F7_seg%201.png?generation=1729090627547931&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ff55f572fdb47602f4bf147ebf6d0713a%2F1%20(1).png?generation=1729090647580937&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F66d92f87e6873a0cfe45a185fed5c8b0%2F1_seg.png?generation=1729090660449277&amp;alt=media\" alt=\"\"></p>\n<h3>Disease Prediction with YOLO</h3>\n<p>Another approach involved training YOLO to predict disease coordinates. We used the coordinates to extract a <strong>50x50</strong> crop around the predicted disease location. Additionally, we projected the disease's <strong>x, y coordinates</strong> into the other two modalities (T1, T2) thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for showing how projection could be done, and applied the same cropping strategy.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F2abcc472f73773f30ff6a3eae6506155%2FYolo_disease%201.jpg?generation=1729090703285103&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F95c2adfaf454f4383a11be509831a7bd%2Fpr_l1.png?generation=1729090716063007&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F6abdea05ad39c2a4ae043aa0562961d7%2Fpr_L4.png?generation=1729090728003675&amp;alt=media\" alt=\"\"></p>\n<h3>Cascaded Models</h3>\n<p>We experimented with a cascaded model approach: </p>\n<ul>\n<li><p>The first model predicted <strong>Normal vs. Abnormal</strong> conditions, trained with Binary Cross-Entropy (BCE) loss. </p></li>\n<li><p>The second model predicted Abnormalities <strong>Moderate vs. Severe</strong> conditions. </p></li>\n</ul>\n<p>Each model outputs a value between 0 and 1. We then concatenated the outputs to get the final prediction. We tried two approaches: </p>\n<ul>\n<li><p>Using probability-based evaluation, which achieved a local score of <strong>38%</strong>. </p></li>\n<li><p>Using a deep fully connected (FC) layer to directly predict the final probabilities, yielding similar results. </p></li>\n<li><p>Unfortunately, neither of these approaches performed well on the public or private leaderboard. </p></li>\n</ul>\n<h2>ORCHA: Our Experiment Tracking and Automation Tool</h2>\n<p>Throughout this competition, we used an internal tool that’s currently under development, which was pivotal in streamlining our workflow. This tool helped us manage experiments, automate repetitive tasks, and utilize resources efficiently, allowing us to focus more on refining model performance.</p>\n<p>Key aspects of this tool that enhanced our process include:</p>\n<ul>\n<li>Experiment Tracking: It allowed us to log and organize experiment data systematically, making it easier to reference previous runs and assess the impact of hyperparameter changes or architecture tweaks.</li>\n<li>Automation: By automating routine tasks like data preprocessing and model training, we could focus on optimizing our models without getting bogged down by manual tasks.</li>\n<li>Execution Efficiency: The tool managed resources autonomously, handling tasks such as training and evaluation in the background, ensuring optimal use of computational resources without constant user input.</li>\n<li>Customization: It was highly adaptable, which made it easier to tailor the tool to fit the specific needs of different competitions and use cases.</li>\n</ul>\n<p>Although still in development, ORCHA has been instrumental in improving our efficiency and enabling us to iterate faster. We recommend that others consider developing similar tools tailored to their specific workflows, as even early-stage solutions can significantly enhance productivity and simplify complex tasks.</p>\n<p>For more information, feel free to contact us at ayman.elghotni@brightskiesinc.com.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ffa5827eb280d44ce370206b6f932ec4a%2Ftr1.png?generation=1729090862548327&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F584343bc84ff675d884de4207e00949a%2Ftr2.png?generation=1729090872614220&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc9b56b52f6ecf9fcd9362dd314009271%2Ft_p.png?generation=1729088297606266&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "It was a challenging and exciting journey, and the high level of competition truly pushed everyone to excel. First, I want to thank my teammates, @hossamelkady99999, @aymanelghotni, @youssefhany10, @amgedelmasry001  who showed great effort throughout the competition. I also want to extend my thanks to all contributers for their invaluable contributions, support, and inspiration throughout the competition. Additionally, a heartfelt thank you to the organizers for putting together such a well-structured and engaging competition. Your efforts made this a great learning experience for all involved. \n\n\n## Overview of the Approach \n### Data Augmentations \nThanks to @itsuki9180 for providing baseline to start from and submission notebook\n- Brightness and Contrast Adjustment: Randomly adjusts the brightness and contrast within a limit of ±0.2 \n- Blurring and Noise: Applies one of the following transformations: \n- Motion blur, median blur, or Gaussian blur with a limit of 5. \n- Adds Gaussian noise with a variance limit between 5 and 30. \n- Geometric Distortions: Applies one of the following transformations: \n- Optical distortion with a limit of 1.0. \n- Grid distortion with 5 steps and a distortion limit of 1.0. \n- Elastic transformation with an alpha of 3. \n- Shift, Scale, and Rotate: Randomly shifts, scales, and rotates the image with a shift and scale limit of 10%, and a rotation limit of 15 degrees. \n- Coarse Dropout: Randomly removes parts of the image by dropping out up to 3 areas, with a maximum size of 12x12 pixels and a minimum size of 8x8 pixels. \n\n### Hyperparameters and Training Strategy \n- Optimizer: AdamW \n- Learning Rate: 0.0002 \n- Weight Decay: 0.01 \n- Batch Size: 4 \n- Cross-Validations: 5 Folds\n- Learning Rate Scheduler: Cosine Annealing with Warm Restarts \n- Mixup Regularization: Applied during training \n\n### Data Preprocessing\nWe started by converting all DICOM images to PNG format, which allowed for more standardized handling across models. For the baseline, we selected **10 evenly spaced slices** from each modality and stacked them across the channel axis, creating a tensor of shape **(bs, 30, 512, 512)**, where bs represents batch size, and 30 comes from the 10 slices per modality (Sagittal T1, Sagittal T2, and Axial). \n\nFor feature extraction, we used **EfficientNet-B5 / EfficientNet-V2** as the backbone, followed by a fully connected (FC) layer to classify into 75 output classes.\n\n### Multi-Instance Learning (MIL) plus (2.5D) \n\nTo improve performance, we introduced **Multi-Instance Learning (MIL)**. Instead of concatenating the slices, we preserved them per modality. Thus, for each modality, we extracted 10 slices, leading to an input shape of **(bs, 3, 10, 512, 512)**. The three channels correspond to the Sagittal T1, Sagittal T2, and Axial modalities. \n\nWe experimented with various sequence models, including **Bi-LSTMs**, **GRUs**, and **attention mechanisms**. Among these, the attention mechanism achieved the best results, pushing the public leaderboard score to **0.55%**. \n\n### YOLO for Preprocessing and Cropping \n\nNext, we trained **YOLO v8** to predict specific levels across the modalities (Sagittal T1, Sagittal T2, and Axial) using the provided coordinates. After predicting the coordinates, we cropped the images to **100x100** pixels and stored them in the following structure:\n\n```\nL1_L2/ (T1, T2, Axial)\nL2_L3/ (T1, T2, Axial) \n...\nL5_S1/ (T1, T2, Axial)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F682a0a18e2322d86b835784c61163cc9%2Fyolo_l.jpg?generation=1729089212492676&alt=media)\nWe then calculated the mean number of slices per modality per level, which averaged around **7 slices**. With this information, we reshaped the data into **(bs, 5, 3, 7, 100, 100)**, where 5 represents the levels, 3 represents the modalities, and 7 represents the slices.\n\n### Enhancing the Attention Mechanism\nInitially, we attempted using **1 attention block after the encoder** with an input shape of **(bs*5, 3, n_features)**, but it did not improve the baseline score. We hypothesized that the model might not be receiving enough data per sequence. \n\nTo address this, we augmented the input by adding adjacent slices to each selected slice. For example, if we selected slices at intervals of [0, 3, 6, 9, 12, 15, 18] from a folder with 20 slices, we also included adjacent slices like [2, 4] for slice 3. For edge cases, we padded with zeros. This resulted in an input shape of **(bs, 5, 3, 7, 3, 100, 100)**, where the extra dimension represents the adjacent slices (input channels). \n\nAfter the encoder, the shape became **(bs*5, 3x7, n_features)**. By increasing the sequence length by a factor of 7, we saw an improvement in the public leaderboard score to **0.43%**.\n\n### Separate Encoders for Modalities \n\nWe also tried separating the encoders for Axial and Sagittal (T1 and T2) modalities. This resulted in two encoders handling the three modalities. Although this approach showed minor improvements in the overall score, it did not have a significant impact on cross-validation (CV) performance. However, this step was necessary for further approaches. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc60bb0557270ad9f11cae5af97c313d5%2Fmodel.png?generation=1729089435504890&alt=media)\n### Data Sorting \n\nWe discovered some issues with data sorting thanks to @hengck23 for showing how to sort it using IPP and we discovered also incorrect metadata, which included duplicate modalities.We merged the duplicates and corrected the sorting. This refinement resulted in a slight improvement, increasing the score to **0.42%**.\n\n## Additional Approaches \n\n### Regression Models \n\nWe also trained a **regression** model using the coordinates improved thanks to @brendanartley for improving the coordinates data, but this model did not perform better than YOLO. The results were comparable to those achieved by YOLO v8. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fdc481eae1eeba7cb7e2066ec3e547909%2Fpred_coord.png?generation=1729090680548927&alt=media)\n### Segmentation of Vertebra \n\nWe attempted to segment the vertebrae and stack it over the channel, but this approach led to worse results than our best submission. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F1f1d8ec89f54f98e2e34389bf473a352%2F7%201.png?generation=1729090606914059&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F00d991945776fb148e6fb611a4231ec2%2F7_seg%201.png?generation=1729090627547931&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ff55f572fdb47602f4bf147ebf6d0713a%2F1%20(1).png?generation=1729090647580937&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F66d92f87e6873a0cfe45a185fed5c8b0%2F1_seg.png?generation=1729090660449277&alt=media)\n### Disease Prediction with YOLO \n\nAnother approach involved training YOLO to predict disease coordinates. We used the coordinates to extract a **50x50** crop around the predicted disease location. Additionally, we projected the disease's **x, y coordinates** into the other two modalities (T1, T2) thanks to @hengck23 for showing how projection could be done, and applied the same cropping strategy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F2abcc472f73773f30ff6a3eae6506155%2FYolo_disease%201.jpg?generation=1729090703285103&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F95c2adfaf454f4383a11be509831a7bd%2Fpr_l1.png?generation=1729090716063007&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F6abdea05ad39c2a4ae043aa0562961d7%2Fpr_L4.png?generation=1729090728003675&alt=media)\n### Cascaded Models \n\nWe experimented with a cascaded model approach: \n\n- The first model predicted **Normal vs. Abnormal** conditions, trained with Binary Cross-Entropy (BCE) loss. \n\n- The second model predicted Abnormalities **Moderate vs. Severe** conditions. \n\nEach model outputs a value between 0 and 1. We then concatenated the outputs to get the final prediction. We tried two approaches: \n\n- Using probability-based evaluation, which achieved a local score of **38%**. \n\n- Using a deep fully connected (FC) layer to directly predict the final probabilities, yielding similar results. \n\n- Unfortunately, neither of these approaches performed well on the public or private leaderboard. \n\n## ORCHA: Our Experiment Tracking and Automation Tool \n\nThroughout this competition, we used an internal tool that’s currently under development, which was pivotal in streamlining our workflow. This tool helped us manage experiments, automate repetitive tasks, and utilize resources efficiently, allowing us to focus more on refining model performance.\n\nKey aspects of this tool that enhanced our process include:\n- Experiment Tracking: It allowed us to log and organize experiment data systematically, making it easier to reference previous runs and assess the impact of hyperparameter changes or architecture tweaks.\n- Automation: By automating routine tasks like data preprocessing and model training, we could focus on optimizing our models without getting bogged down by manual tasks.\n- Execution Efficiency: The tool managed resources autonomously, handling tasks such as training and evaluation in the background, ensuring optimal use of computational resources without constant user input.\n- Customization: It was highly adaptable, which made it easier to tailor the tool to fit the specific needs of different competitions and use cases.\n\nAlthough still in development, ORCHA has been instrumental in improving our efficiency and enabling us to iterate faster. We recommend that others consider developing similar tools tailored to their specific workflows, as even early-stage solutions can significantly enhance productivity and simplify complex tasks.\n\nFor more information, feel free to contact us at ayman.elghotni@brightskiesinc.com.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ffa5827eb280d44ce370206b6f932ec4a%2Ftr1.png?generation=1729090862548327&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F584343bc84ff675d884de4207e00949a%2Ftr2.png?generation=1729090872614220&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc9b56b52f6ecf9fcd9362dd314009271%2Ft_p.png?generation=1729088297606266&alt=media)",
      "votes": null
    },
    {
      "id": "3020304",
      "postDate": "10/17/2024 11:33:26",
      "content": "<p>It was a hell of a learning experience, thank you to all the organizers and to the helpful community of participants of this competition!</p>",
      "rawMarkdown": "It was a hell of a learning experience, thank you to all the organizers and to the helpful community of participants of this competition!",
      "votes": null
    },
    {
      "id": "3020345",
      "postDate": "10/17/2024 12:24:28",
      "content": "<p>This competition provided an amazing learning experience that I will cherish. I want to extend my heartfelt thanks to all the organizers for their hard work and dedication, as well as to the supportive community of participants who made this event so enriching. A special shoutout to my teammates for their collaboration and commitment throughout the process; your support made all the difference!</p>",
      "rawMarkdown": "This competition provided an amazing learning experience that I will cherish. I want to extend my heartfelt thanks to all the organizers for their hard work and dedication, as well as to the supportive community of participants who made this event so enriching. A special shoutout to my teammates for their collaboration and commitment throughout the process; your support made all the difference!",
      "votes": null
    },
    {
      "id": "3020726",
      "postDate": "10/17/2024 20:19:01",
      "content": "<p>That’s a great solution and description. Good job guys 👏</p>",
      "rawMarkdown": "That’s a great solution and description. Good job guys 👏",
      "votes": null
    },
    {
      "id": "3021127",
      "postDate": "10/18/2024 07:58:54",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": null
    },
    {
      "id": "3022277",
      "postDate": "10/19/2024 12:37:37",
      "content": "<p>Thank you so much </p>",
      "rawMarkdown": "Thank you so much",
      "votes": null
    },
    {
      "id": "3023142",
      "postDate": "10/20/2024 09:15:22",
      "content": "<p>This competition provided an amazing learning experience. I want to thank all the organizers and the supportive community of participants. A special shoutout to my teammates for their collaboration and commitment throughout the process.</p>",
      "rawMarkdown": "This competition provided an amazing learning experience. I want to thank all the organizers and the supportive community of participants. A special shoutout to my teammates for their collaboration and commitment throughout the process.",
      "votes": null
    },
    {
      "id": "3023161",
      "postDate": "10/20/2024 09:26:22",
      "content": "<p>I've gained tremendously throughout the competition, and I extend my deepest gratitude to the community for their engaging discussions. A special thank you goes to my teammates for their hard work and support. Additionally, I must acknowledge our invaluable tool, ORCHA, which played a crucial role in our experiments. Thank you all for your support and contributions.</p>",
      "rawMarkdown": "I've gained tremendously throughout the competition, and I extend my deepest gratitude to the community for their engaging discussions. A special thank you goes to my teammates for their hard work and support. Additionally, I must acknowledge our invaluable tool, ORCHA, which played a crucial role in our experiments. Thank you all for your support and contributions.",
      "votes": null
    },
    {
      "id": "3023217",
      "postDate": "10/20/2024 10:29:34",
      "content": "<p>Great work! Congratulations!</p>",
      "rawMarkdown": "Great work! Congratulations!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3020304,
      "author_name": "aymanelghotni",
      "author_url": "",
      "post_date": "10/17/2024 11:33:26",
      "content": "<p>It was a hell of a learning experience, thank you to all the organizers and to the helpful community of participants of this competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3020345,
      "author_name": "hossamelkady99999",
      "author_url": "",
      "post_date": "10/17/2024 12:24:28",
      "content": "<p>This competition provided an amazing learning experience that I will cherish. I want to extend my heartfelt thanks to all the organizers for their hard work and dedication, as well as to the supportive community of participants who made this event so enriching. A special shoutout to my teammates for their collaboration and commitment throughout the process; your support made all the difference!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3020726,
      "author_name": "mahmoudelkarargy",
      "author_url": "",
      "post_date": "10/17/2024 20:19:01",
      "content": "<p>That’s a great solution and description. Good job guys 👏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3021127,
      "author_name": "vaibhaavtiwari",
      "author_url": "",
      "post_date": "10/18/2024 07:58:54",
      "content": "<p>Great Work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3022277,
          "author_name": "mohamednabil33",
          "author_url": "",
          "post_date": "10/19/2024 12:37:37",
          "content": "<p>Thank you so much </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3023142,
      "author_name": "youssefhany10",
      "author_url": "",
      "post_date": "10/20/2024 09:15:22",
      "content": "<p>This competition provided an amazing learning experience. I want to thank all the organizers and the supportive community of participants. A special shoutout to my teammates for their collaboration and commitment throughout the process.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3023161,
      "author_name": "amgedelmasry001",
      "author_url": "",
      "post_date": "10/20/2024 09:26:22",
      "content": "<p>I've gained tremendously throughout the competition, and I extend my deepest gratitude to the community for their engaging discussions. A special thank you goes to my teammates for their hard work and support. Additionally, I must acknowledge our invaluable tool, ORCHA, which played a crucial role in our experiments. Thank you all for your support and contributions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3023217,
      "author_name": "asmahachaichi",
      "author_url": "",
      "post_date": "10/20/2024 10:29:34",
      "content": "<p>Great work! Congratulations!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3020300": "It was a challenging and exciting journey, and the high level of competition truly pushed everyone to excel. First, I want to thank my teammates, @hossamelkady99999, @aymanelghotni, @youssefhany10, @amgedelmasry001  who showed great effort throughout the competition. I also want to extend my thanks to all contributers for their invaluable contributions, support, and inspiration throughout the competition. Additionally, a heartfelt thank you to the organizers for putting together such a well-structured and engaging competition. Your efforts made this a great learning experience for all involved. \n\n\n## Overview of the Approach \n### Data Augmentations \nThanks to @itsuki9180 for providing baseline to start from and submission notebook\n- Brightness and Contrast Adjustment: Randomly adjusts the brightness and contrast within a limit of ±0.2 \n- Blurring and Noise: Applies one of the following transformations: \n- Motion blur, median blur, or Gaussian blur with a limit of 5. \n- Adds Gaussian noise with a variance limit between 5 and 30. \n- Geometric Distortions: Applies one of the following transformations: \n- Optical distortion with a limit of 1.0. \n- Grid distortion with 5 steps and a distortion limit of 1.0. \n- Elastic transformation with an alpha of 3. \n- Shift, Scale, and Rotate: Randomly shifts, scales, and rotates the image with a shift and scale limit of 10%, and a rotation limit of 15 degrees. \n- Coarse Dropout: Randomly removes parts of the image by dropping out up to 3 areas, with a maximum size of 12x12 pixels and a minimum size of 8x8 pixels. \n\n### Hyperparameters and Training Strategy \n- Optimizer: AdamW \n- Learning Rate: 0.0002 \n- Weight Decay: 0.01 \n- Batch Size: 4 \n- Cross-Validations: 5 Folds\n- Learning Rate Scheduler: Cosine Annealing with Warm Restarts \n- Mixup Regularization: Applied during training \n\n### Data Preprocessing\nWe started by converting all DICOM images to PNG format, which allowed for more standardized handling across models. For the baseline, we selected **10 evenly spaced slices** from each modality and stacked them across the channel axis, creating a tensor of shape **(bs, 30, 512, 512)**, where bs represents batch size, and 30 comes from the 10 slices per modality (Sagittal T1, Sagittal T2, and Axial). \n\nFor feature extraction, we used **EfficientNet-B5 / EfficientNet-V2** as the backbone, followed by a fully connected (FC) layer to classify into 75 output classes.\n\n### Multi-Instance Learning (MIL) plus (2.5D) \n\nTo improve performance, we introduced **Multi-Instance Learning (MIL)**. Instead of concatenating the slices, we preserved them per modality. Thus, for each modality, we extracted 10 slices, leading to an input shape of **(bs, 3, 10, 512, 512)**. The three channels correspond to the Sagittal T1, Sagittal T2, and Axial modalities. \n\nWe experimented with various sequence models, including **Bi-LSTMs**, **GRUs**, and **attention mechanisms**. Among these, the attention mechanism achieved the best results, pushing the public leaderboard score to **0.55%**. \n\n### YOLO for Preprocessing and Cropping \n\nNext, we trained **YOLO v8** to predict specific levels across the modalities (Sagittal T1, Sagittal T2, and Axial) using the provided coordinates. After predicting the coordinates, we cropped the images to **100x100** pixels and stored them in the following structure:\n\n```\nL1_L2/ (T1, T2, Axial)\nL2_L3/ (T1, T2, Axial) \n...\nL5_S1/ (T1, T2, Axial)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F682a0a18e2322d86b835784c61163cc9%2Fyolo_l.jpg?generation=1729089212492676&alt=media)\nWe then calculated the mean number of slices per modality per level, which averaged around **7 slices**. With this information, we reshaped the data into **(bs, 5, 3, 7, 100, 100)**, where 5 represents the levels, 3 represents the modalities, and 7 represents the slices.\n\n### Enhancing the Attention Mechanism\nInitially, we attempted using **1 attention block after the encoder** with an input shape of **(bs*5, 3, n_features)**, but it did not improve the baseline score. We hypothesized that the model might not be receiving enough data per sequence. \n\nTo address this, we augmented the input by adding adjacent slices to each selected slice. For example, if we selected slices at intervals of [0, 3, 6, 9, 12, 15, 18] from a folder with 20 slices, we also included adjacent slices like [2, 4] for slice 3. For edge cases, we padded with zeros. This resulted in an input shape of **(bs, 5, 3, 7, 3, 100, 100)**, where the extra dimension represents the adjacent slices (input channels). \n\nAfter the encoder, the shape became **(bs*5, 3x7, n_features)**. By increasing the sequence length by a factor of 7, we saw an improvement in the public leaderboard score to **0.43%**.\n\n### Separate Encoders for Modalities \n\nWe also tried separating the encoders for Axial and Sagittal (T1 and T2) modalities. This resulted in two encoders handling the three modalities. Although this approach showed minor improvements in the overall score, it did not have a significant impact on cross-validation (CV) performance. However, this step was necessary for further approaches. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc60bb0557270ad9f11cae5af97c313d5%2Fmodel.png?generation=1729089435504890&alt=media)\n### Data Sorting \n\nWe discovered some issues with data sorting thanks to @hengck23 for showing how to sort it using IPP and we discovered also incorrect metadata, which included duplicate modalities.We merged the duplicates and corrected the sorting. This refinement resulted in a slight improvement, increasing the score to **0.42%**.\n\n## Additional Approaches \n\n### Regression Models \n\nWe also trained a **regression** model using the coordinates improved thanks to @brendanartley for improving the coordinates data, but this model did not perform better than YOLO. The results were comparable to those achieved by YOLO v8. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fdc481eae1eeba7cb7e2066ec3e547909%2Fpred_coord.png?generation=1729090680548927&alt=media)\n### Segmentation of Vertebra \n\nWe attempted to segment the vertebrae and stack it over the channel, but this approach led to worse results than our best submission. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F1f1d8ec89f54f98e2e34389bf473a352%2F7%201.png?generation=1729090606914059&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F00d991945776fb148e6fb611a4231ec2%2F7_seg%201.png?generation=1729090627547931&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ff55f572fdb47602f4bf147ebf6d0713a%2F1%20(1).png?generation=1729090647580937&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F66d92f87e6873a0cfe45a185fed5c8b0%2F1_seg.png?generation=1729090660449277&alt=media)\n### Disease Prediction with YOLO \n\nAnother approach involved training YOLO to predict disease coordinates. We used the coordinates to extract a **50x50** crop around the predicted disease location. Additionally, we projected the disease's **x, y coordinates** into the other two modalities (T1, T2) thanks to @hengck23 for showing how projection could be done, and applied the same cropping strategy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F2abcc472f73773f30ff6a3eae6506155%2FYolo_disease%201.jpg?generation=1729090703285103&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F95c2adfaf454f4383a11be509831a7bd%2Fpr_l1.png?generation=1729090716063007&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F6abdea05ad39c2a4ae043aa0562961d7%2Fpr_L4.png?generation=1729090728003675&alt=media)\n### Cascaded Models \n\nWe experimented with a cascaded model approach: \n\n- The first model predicted **Normal vs. Abnormal** conditions, trained with Binary Cross-Entropy (BCE) loss. \n\n- The second model predicted Abnormalities **Moderate vs. Severe** conditions. \n\nEach model outputs a value between 0 and 1. We then concatenated the outputs to get the final prediction. We tried two approaches: \n\n- Using probability-based evaluation, which achieved a local score of **38%**. \n\n- Using a deep fully connected (FC) layer to directly predict the final probabilities, yielding similar results. \n\n- Unfortunately, neither of these approaches performed well on the public or private leaderboard. \n\n## ORCHA: Our Experiment Tracking and Automation Tool \n\nThroughout this competition, we used an internal tool that’s currently under development, which was pivotal in streamlining our workflow. This tool helped us manage experiments, automate repetitive tasks, and utilize resources efficiently, allowing us to focus more on refining model performance.\n\nKey aspects of this tool that enhanced our process include:\n- Experiment Tracking: It allowed us to log and organize experiment data systematically, making it easier to reference previous runs and assess the impact of hyperparameter changes or architecture tweaks.\n- Automation: By automating routine tasks like data preprocessing and model training, we could focus on optimizing our models without getting bogged down by manual tasks.\n- Execution Efficiency: The tool managed resources autonomously, handling tasks such as training and evaluation in the background, ensuring optimal use of computational resources without constant user input.\n- Customization: It was highly adaptable, which made it easier to tailor the tool to fit the specific needs of different competitions and use cases.\n\nAlthough still in development, ORCHA has been instrumental in improving our efficiency and enabling us to iterate faster. We recommend that others consider developing similar tools tailored to their specific workflows, as even early-stage solutions can significantly enhance productivity and simplify complex tasks.\n\nFor more information, feel free to contact us at ayman.elghotni@brightskiesinc.com.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Ffa5827eb280d44ce370206b6f932ec4a%2Ftr1.png?generation=1729090862548327&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2F584343bc84ff675d884de4207e00949a%2Ftr2.png?generation=1729090872614220&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14071460%2Fc9b56b52f6ecf9fcd9362dd314009271%2Ft_p.png?generation=1729088297606266&alt=media)",
    "3020304": "It was a hell of a learning experience, thank you to all the organizers and to the helpful community of participants of this competition!",
    "3020345": "This competition provided an amazing learning experience that I will cherish. I want to extend my heartfelt thanks to all the organizers for their hard work and dedication, as well as to the supportive community of participants who made this event so enriching. A special shoutout to my teammates for their collaboration and commitment throughout the process; your support made all the difference!",
    "3020726": "That’s a great solution and description. Good job guys 👏",
    "3021127": "Great Work!",
    "3022277": "Thank you so much",
    "3023142": "This competition provided an amazing learning experience. I want to thank all the organizers and the supportive community of participants. A special shoutout to my teammates for their collaboration and commitment throughout the process.",
    "3023161": "I've gained tremendously throughout the competition, and I extend my deepest gratitude to the community for their engaging discussions. A special thank you goes to my teammates for their hard work and support. Additionally, I must acknowledge our invaluable tool, ORCHA, which played a crucial role in our experiments. Thank you all for your support and contributions.",
    "3023217": "Great work! Congratulations!"
  },
  "source": "meta"
}