{
  "id": 280572,
  "title": "8th place solution",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/writeups/arturhugo-8th-place-solution",
  "author_name": "",
  "post_date": "2021-10-21T19:35:01.800285500Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Well, that was unexpected. First of all I'd like to thank Kaggle and the host to make this competition. I'd also like to thank everyone who discussed leaderboard shakeup on the forum, because it really helped me understand what just happened. This was my first competition ever and I was participating as part of an assignment from a ML class I'm taking, so I got really confused with the results.</p>\n<p>Nonetheless, I'm here to document and share my solution. But before that, I need to mention that the code is messy and all in just one notebook (<a href=\"https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook\" target=\"_blank\">here</a>), I had one week to make the submission. Also, the idea of this assignment was to exercise what we had learned in class so far, so I just went with what we had just seen (we were seeing feature extraction and bag-of-words approach at the time).</p>\n<h2>1. Extracting features</h2>\n<p>The first step was to extract the features of all images of the training set. I used the ORB feature detector from OpenCV because it would take to long to run the solution if I used SIFT. I should have preprocessed all the data just once and serialized to use as input afterwards, but I was still figuring kaggle out at the time, so I didn't think of it.</p>\n<p>So in this step, I iterate over the training data going one sample at a time. For each sample, I iterate over every image (FLAIR, T1w, T1Gd, T2), extract its features and stack them making a feature matrix of <em>n_features</em> by <em>descriptor_size</em> dimensions. At the end, I have a list of <em>train_size</em> matrices with varied number of rows.</p>\n<h2>2. Creating the visual vocabulary</h2>\n<p>Once I have all the features for all the samples, I need to establish what I will consider a word for my bag-of-words. So in this step I stack all the feature matrices of the training set making a <em>n_features0</em> + <em>n_features1</em> + … + <em>n_featuresN</em> by <em>descriptor_size</em> matrix (<em>all_features</em>).</p>\n<p>I run a clustering algorithm over <em>all_features</em> and the clusters of the fitted model is my vocabulary. I used the MiniBatchKMeans from scikit-learn because I was running in memory issues with the amount of descriptors used to fit the model.</p>\n<h2>3. Word frequency histograms</h2>\n<p>With the vocabulary in hand, I just needed to run through each sample's feature matrix to convert it to a histogram of word frequency. For each feature in a matrix I used the clustering model to predict it's cluster and incremented the correspond position in the histogram by 1/<em>n_features</em>. This process is later repeated for the test set in order to get the predictions.</p>\n<h2>4. The classifier and the predictions</h2>\n<p>Here I just used an SVC from scikit-learn with all the parameters set to default. I trained the model in the training set histograms and did the predictions on the test set histograms. I set probability to true when making the predictions.</p>\n<h2>Final considerations</h2>\n<p>I could see with this competition that the problem posed is really complex and, more likely than not, my rather simple approach of bag-of-words is not suited for this task. One thing I found interesting when looking at the probabilities of the predictions is that no test case had a split more pronounced then 0.6/0.4, which I think means the model is not confident in what it is predicting.</p>\n<p>I know that I was really lucky with the private leaderboard score, since I did a couple of late submissions to see if the performance was consistent and they got 0.58 and 0.59 in the private leaderboard, instead of the previous 0.6. But it seems that no model was able to perform well in the task, so I think the results serve the purpose of showing that there is a long way to go in order to solve the task at hand.</p>\n<h2>Notebook</h2>\n<p><a href=\"https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook\" target=\"_blank\">https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook</a></p>",
  "messages": [
    {
      "id": "1552991",
      "postDate": "10/21/2021 19:35:01",
      "content": "<p>Well, that was unexpected. First of all I'd like to thank Kaggle and the host to make this competition. I'd also like to thank everyone who discussed leaderboard shakeup on the forum, because it really helped me understand what just happened. This was my first competition ever and I was participating as part of an assignment from a ML class I'm taking, so I got really confused with the results.</p>\n<p>Nonetheless, I'm here to document and share my solution. But before that, I need to mention that the code is messy and all in just one notebook (<a href=\"https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook\" target=\"_blank\">here</a>), I had one week to make the submission. Also, the idea of this assignment was to exercise what we had learned in class so far, so I just went with what we had just seen (we were seeing feature extraction and bag-of-words approach at the time).</p>\n<h2>1. Extracting features</h2>\n<p>The first step was to extract the features of all images of the training set. I used the ORB feature detector from OpenCV because it would take to long to run the solution if I used SIFT. I should have preprocessed all the data just once and serialized to use as input afterwards, but I was still figuring kaggle out at the time, so I didn't think of it.</p>\n<p>So in this step, I iterate over the training data going one sample at a time. For each sample, I iterate over every image (FLAIR, T1w, T1Gd, T2), extract its features and stack them making a feature matrix of <em>n_features</em> by <em>descriptor_size</em> dimensions. At the end, I have a list of <em>train_size</em> matrices with varied number of rows.</p>\n<h2>2. Creating the visual vocabulary</h2>\n<p>Once I have all the features for all the samples, I need to establish what I will consider a word for my bag-of-words. So in this step I stack all the feature matrices of the training set making a <em>n_features0</em> + <em>n_features1</em> + … + <em>n_featuresN</em> by <em>descriptor_size</em> matrix (<em>all_features</em>).</p>\n<p>I run a clustering algorithm over <em>all_features</em> and the clusters of the fitted model is my vocabulary. I used the MiniBatchKMeans from scikit-learn because I was running in memory issues with the amount of descriptors used to fit the model.</p>\n<h2>3. Word frequency histograms</h2>\n<p>With the vocabulary in hand, I just needed to run through each sample's feature matrix to convert it to a histogram of word frequency. For each feature in a matrix I used the clustering model to predict it's cluster and incremented the correspond position in the histogram by 1/<em>n_features</em>. This process is later repeated for the test set in order to get the predictions.</p>\n<h2>4. The classifier and the predictions</h2>\n<p>Here I just used an SVC from scikit-learn with all the parameters set to default. I trained the model in the training set histograms and did the predictions on the test set histograms. I set probability to true when making the predictions.</p>\n<h2>Final considerations</h2>\n<p>I could see with this competition that the problem posed is really complex and, more likely than not, my rather simple approach of bag-of-words is not suited for this task. One thing I found interesting when looking at the probabilities of the predictions is that no test case had a split more pronounced then 0.6/0.4, which I think means the model is not confident in what it is predicting.</p>\n<p>I know that I was really lucky with the private leaderboard score, since I did a couple of late submissions to see if the performance was consistent and they got 0.58 and 0.59 in the private leaderboard, instead of the previous 0.6. But it seems that no model was able to perform well in the task, so I think the results serve the purpose of showing that there is a long way to go in order to solve the task at hand.</p>\n<h2>Notebook</h2>\n<p><a href=\"https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook\" target=\"_blank\">https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook</a></p>",
      "rawMarkdown": "Well, that was unexpected. First of all I'd like to thank Kaggle and the host to make this competition. I'd also like to thank everyone who discussed leaderboard shakeup on the forum, because it really helped me understand what just happened. This was my first competition ever and I was participating as part of an assignment from a ML class I'm taking, so I got really confused with the results.\n\nNonetheless, I'm here to document and share my solution. But before that, I need to mention that the code is messy and all in just one notebook ([here](https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook)), I had one week to make the submission. Also, the idea of this assignment was to exercise what we had learned in class so far, so I just went with what we had just seen (we were seeing feature extraction and bag-of-words approach at the time).\n\n## 1. Extracting features\n\nThe first step was to extract the features of all images of the training set. I used the ORB feature detector from OpenCV because it would take to long to run the solution if I used SIFT. I should have preprocessed all the data just once and serialized to use as input afterwards, but I was still figuring kaggle out at the time, so I didn't think of it.\n\nSo in this step, I iterate over the training data going one sample at a time. For each sample, I iterate over every image (FLAIR, T1w, T1Gd, T2), extract its features and stack them making a feature matrix of _n_features_ by _descriptor_size_ dimensions. At the end, I have a list of _train_size_ matrices with varied number of rows.\n\n## 2. Creating the visual vocabulary\n\nOnce I have all the features for all the samples, I need to establish what I will consider a word for my bag-of-words. So in this step I stack all the feature matrices of the training set making a _n_features0_ + _n_features1_ + ... + _n_featuresN_ by _descriptor_size_ matrix (_all_features_).\n\nI run a clustering algorithm over _all_features_ and the clusters of the fitted model is my vocabulary. I used the MiniBatchKMeans from scikit-learn because I was running in memory issues with the amount of descriptors used to fit the model.\n\n## 3. Word frequency histograms\n\nWith the vocabulary in hand, I just needed to run through each sample's feature matrix to convert it to a histogram of word frequency. For each feature in a matrix I used the clustering model to predict it's cluster and incremented the correspond position in the histogram by 1/_n_features_. This process is later repeated for the test set in order to get the predictions.\n\n## 4. The classifier and the predictions\n\nHere I just used an SVC from scikit-learn with all the parameters set to default. I trained the model in the training set histograms and did the predictions on the test set histograms. I set probability to true when making the predictions.\n\n## Final considerations\n\nI could see with this competition that the problem posed is really complex and, more likely than not, my rather simple approach of bag-of-words is not suited for this task. One thing I found interesting when looking at the probabilities of the predictions is that no test case had a split more pronounced then 0.6/0.4, which I think means the model is not confident in what it is predicting.\n\nI know that I was really lucky with the private leaderboard score, since I did a couple of late submissions to see if the performance was consistent and they got 0.58 and 0.59 in the private leaderboard, instead of the previous 0.6. But it seems that no model was able to perform well in the task, so I think the results serve the purpose of showing that there is a long way to go in order to solve the task at hand.\n\n## Notebook\n\nhttps://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook",
      "votes": null
    },
    {
      "id": "2934113",
      "postDate": "07/24/2024 07:42:20",
      "content": "<p>Very interesting approach where most of works are focusing on deep learning techniques. Good work!</p>",
      "rawMarkdown": "Very interesting approach where most of works are focusing on deep learning techniques. Good work!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2934113,
      "author_name": "wheresmadog",
      "author_url": "",
      "post_date": "07/24/2024 07:42:20",
      "content": "<p>Very interesting approach where most of works are focusing on deep learning techniques. Good work!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1552991": "Well, that was unexpected. First of all I'd like to thank Kaggle and the host to make this competition. I'd also like to thank everyone who discussed leaderboard shakeup on the forum, because it really helped me understand what just happened. This was my first competition ever and I was participating as part of an assignment from a ML class I'm taking, so I got really confused with the results.\n\nNonetheless, I'm here to document and share my solution. But before that, I need to mention that the code is messy and all in just one notebook ([here](https://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook)), I had one week to make the submission. Also, the idea of this assignment was to exercise what we had learned in class so far, so I just went with what we had just seen (we were seeing feature extraction and bag-of-words approach at the time).\n\n## 1. Extracting features\n\nThe first step was to extract the features of all images of the training set. I used the ORB feature detector from OpenCV because it would take to long to run the solution if I used SIFT. I should have preprocessed all the data just once and serialized to use as input afterwards, but I was still figuring kaggle out at the time, so I didn't think of it.\n\nSo in this step, I iterate over the training data going one sample at a time. For each sample, I iterate over every image (FLAIR, T1w, T1Gd, T2), extract its features and stack them making a feature matrix of _n_features_ by _descriptor_size_ dimensions. At the end, I have a list of _train_size_ matrices with varied number of rows.\n\n## 2. Creating the visual vocabulary\n\nOnce I have all the features for all the samples, I need to establish what I will consider a word for my bag-of-words. So in this step I stack all the feature matrices of the training set making a _n_features0_ + _n_features1_ + ... + _n_featuresN_ by _descriptor_size_ matrix (_all_features_).\n\nI run a clustering algorithm over _all_features_ and the clusters of the fitted model is my vocabulary. I used the MiniBatchKMeans from scikit-learn because I was running in memory issues with the amount of descriptors used to fit the model.\n\n## 3. Word frequency histograms\n\nWith the vocabulary in hand, I just needed to run through each sample's feature matrix to convert it to a histogram of word frequency. For each feature in a matrix I used the clustering model to predict it's cluster and incremented the correspond position in the histogram by 1/_n_features_. This process is later repeated for the test set in order to get the predictions.\n\n## 4. The classifier and the predictions\n\nHere I just used an SVC from scikit-learn with all the parameters set to default. I trained the model in the training set histograms and did the predictions on the test set histograms. I set probability to true when making the predictions.\n\n## Final considerations\n\nI could see with this competition that the problem posed is really complex and, more likely than not, my rather simple approach of bag-of-words is not suited for this task. One thing I found interesting when looking at the probabilities of the predictions is that no test case had a split more pronounced then 0.6/0.4, which I think means the model is not confident in what it is predicting.\n\nI know that I was really lucky with the private leaderboard score, since I did a couple of late submissions to see if the performance was consistent and they got 0.58 and 0.59 in the private leaderboard, instead of the previous 0.6. But it seems that no model was able to perform well in the task, so I think the results serve the purpose of showing that there is a long way to go in order to solve the task at hand.\n\n## Notebook\n\nhttps://www.kaggle.com/arturhcpereira/rsna-miccai-submission/notebook",
    "2934113": "Very interesting approach where most of works are focusing on deep learning techniques. Good work!"
  },
  "source": "meta"
}