{
  "id": 581212,
  "title": "#2: Pseudolabelling and Resnet-50",
  "url": "/competitions/forams-classification-2025/writeups/ambrosm-2-pseudolabelling-and-resnet-50",
  "author_name": "",
  "post_date": "2025-05-29T05:46:36.783614Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I thank the competition hosts for organizing this challenging competition and learning opportunity. The diagram summarizes my approach:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F843acf737a1a330a3e654c4c42eba6da%2Fpipeline2.png?generation=1748496646746480&amp;alt=media\" alt=\"pipeline\"></p>\n<h1>Feature engineering and data augmentation</h1>\n<p>During the first days of the competition, I tried my luck with linear models and feature engineering based on the distribution of the mass in the 3d images. The notebook <a href=\"https://www.kaggle.com/code/ambrosm/forams-linear-discriminant-analysis\" target=\"_blank\">Forams Linear Discriminant Analysis</a> shows the approach. After publishing this notebook, I added a few more features with the intention of engineering features which do not depend on the spatial orientation of the sample. For instance, the eigenvalues of the inertia tensor are invariant under rotation, and they help distinguish objects with a round appearance (e.g., classes 10 and 11) from flat objects (e.g., classes 5 and 6). Logistic regression with these features got a public leaderboard score of 0.66296.</p>\n<p>An EDA shows that while most class 10 samples are spheres with an empty interior, a few samples of this class are filled with dirt. The dirt in the interior evidently should not matter for classification, but produces noise in features such as mass and inertia tensor. This observation suggests that classification should be based on views of the surface rather than the mass in the interior.</p>\n<p>I then computed six 127x127 surface views for every 128x128x128 cube. Flipping and rotating the surface views augments the training data (the image below shows different views of the first training sample). Resnet-50 then embeds every surface view into a 2048-dimensional space. These 2048 Resnet activations are fed into a neural network with a single hidden layer and 14 softmax outputs. This model reaches a public score of 0.82340.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F62fc4f60c87d9829dd1982b89dddaaa9%2Fviews.png?generation=1748496891410102&amp;alt=media\" alt=\"surface views\"></p>\n<h1>Semi-supervised learning</h1>\n<p>I first tried to use scikit-learn's LabelPropagation to implement semi-supervised learning. Its results weren't good. Pseudolabelling had more success: I compared the predictions of my first four models. For 41 % of the test samples, all four models agreed (i.e., they predicted the same class). I added these 41 % of the test samples to the labelled data and trained a new supervised neural network, which got a public score of 0.86621. This classifier never predicts class 14.</p>\n<h1>Outlier detection</h1>\n<p>Part of the competition task was finding samples which belong to none of the 14 given species. Experiments with IsolationForest had little success, and so I decided to assign samples to class 14 simply if the classifier showed low confidence. </p>\n<h1>Future work</h1>\n<p>My neural network models work only with six axis-parallel views onto the 3d object (top, bottom, left, right, front, back). It will be interesting to see how the classification improves if the data are augmented with views from other angles.</p>\n<p>Source code is <a href=\"https://www.kaggle.com/code/ambrosm/forams-pseudolabeling-and-resnet-50?scriptVersionId=241567496\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": "3211968",
      "postDate": "05/29/2025 05:46:36",
      "content": "<p>I thank the competition hosts for organizing this challenging competition and learning opportunity. The diagram summarizes my approach:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F843acf737a1a330a3e654c4c42eba6da%2Fpipeline2.png?generation=1748496646746480&amp;alt=media\" alt=\"pipeline\"></p>\n<h1>Feature engineering and data augmentation</h1>\n<p>During the first days of the competition, I tried my luck with linear models and feature engineering based on the distribution of the mass in the 3d images. The notebook <a href=\"https://www.kaggle.com/code/ambrosm/forams-linear-discriminant-analysis\" target=\"_blank\">Forams Linear Discriminant Analysis</a> shows the approach. After publishing this notebook, I added a few more features with the intention of engineering features which do not depend on the spatial orientation of the sample. For instance, the eigenvalues of the inertia tensor are invariant under rotation, and they help distinguish objects with a round appearance (e.g., classes 10 and 11) from flat objects (e.g., classes 5 and 6). Logistic regression with these features got a public leaderboard score of 0.66296.</p>\n<p>An EDA shows that while most class 10 samples are spheres with an empty interior, a few samples of this class are filled with dirt. The dirt in the interior evidently should not matter for classification, but produces noise in features such as mass and inertia tensor. This observation suggests that classification should be based on views of the surface rather than the mass in the interior.</p>\n<p>I then computed six 127x127 surface views for every 128x128x128 cube. Flipping and rotating the surface views augments the training data (the image below shows different views of the first training sample). Resnet-50 then embeds every surface view into a 2048-dimensional space. These 2048 Resnet activations are fed into a neural network with a single hidden layer and 14 softmax outputs. This model reaches a public score of 0.82340.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F62fc4f60c87d9829dd1982b89dddaaa9%2Fviews.png?generation=1748496891410102&amp;alt=media\" alt=\"surface views\"></p>\n<h1>Semi-supervised learning</h1>\n<p>I first tried to use scikit-learn's LabelPropagation to implement semi-supervised learning. Its results weren't good. Pseudolabelling had more success: I compared the predictions of my first four models. For 41 % of the test samples, all four models agreed (i.e., they predicted the same class). I added these 41 % of the test samples to the labelled data and trained a new supervised neural network, which got a public score of 0.86621. This classifier never predicts class 14.</p>\n<h1>Outlier detection</h1>\n<p>Part of the competition task was finding samples which belong to none of the 14 given species. Experiments with IsolationForest had little success, and so I decided to assign samples to class 14 simply if the classifier showed low confidence. </p>\n<h1>Future work</h1>\n<p>My neural network models work only with six axis-parallel views onto the 3d object (top, bottom, left, right, front, back). It will be interesting to see how the classification improves if the data are augmented with views from other angles.</p>\n<p>Source code is <a href=\"https://www.kaggle.com/code/ambrosm/forams-pseudolabeling-and-resnet-50?scriptVersionId=241567496\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I thank the competition hosts for organizing this challenging competition and learning opportunity. The diagram summarizes my approach:\n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F843acf737a1a330a3e654c4c42eba6da%2Fpipeline2.png?generation=1748496646746480&alt=media)\n\n# Feature engineering and data augmentation\n\nDuring the first days of the competition, I tried my luck with linear models and feature engineering based on the distribution of the mass in the 3d images. The notebook [Forams Linear Discriminant Analysis](https://www.kaggle.com/code/ambrosm/forams-linear-discriminant-analysis) shows the approach. After publishing this notebook, I added a few more features with the intention of engineering features which do not depend on the spatial orientation of the sample. For instance, the eigenvalues of the inertia tensor are invariant under rotation, and they help distinguish objects with a round appearance (e.g., classes 10 and 11) from flat objects (e.g., classes 5 and 6). Logistic regression with these features got a public leaderboard score of 0.66296.\n\nAn EDA shows that while most class 10 samples are spheres with an empty interior, a few samples of this class are filled with dirt. The dirt in the interior evidently should not matter for classification, but produces noise in features such as mass and inertia tensor. This observation suggests that classification should be based on views of the surface rather than the mass in the interior.\n\nI then computed six 127x127 surface views for every 128x128x128 cube. Flipping and rotating the surface views augments the training data (the image below shows different views of the first training sample). Resnet-50 then embeds every surface view into a 2048-dimensional space. These 2048 Resnet activations are fed into a neural network with a single hidden layer and 14 softmax outputs. This model reaches a public score of 0.82340.\n\n![surface views](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F62fc4f60c87d9829dd1982b89dddaaa9%2Fviews.png?generation=1748496891410102&alt=media)\n\n# Semi-supervised learning\n\nI first tried to use scikit-learn's LabelPropagation to implement semi-supervised learning. Its results weren't good. Pseudolabelling had more success: I compared the predictions of my first four models. For 41 % of the test samples, all four models agreed (i.e., they predicted the same class). I added these 41 % of the test samples to the labelled data and trained a new supervised neural network, which got a public score of 0.86621. This classifier never predicts class 14.\n\n# Outlier detection\n\nPart of the competition task was finding samples which belong to none of the 14 given species. Experiments with IsolationForest had little success, and so I decided to assign samples to class 14 simply if the classifier showed low confidence. \n\n# Future work\n\nMy neural network models work only with six axis-parallel views onto the 3d object (top, bottom, left, right, front, back). It will be interesting to see how the classification improves if the data are augmented with views from other angles.\n\nSource code is [here](https://www.kaggle.com/code/ambrosm/forams-pseudolabeling-and-resnet-50?scriptVersionId=241567496).",
      "votes": null
    },
    {
      "id": "3212040",
      "postDate": "05/29/2025 08:24:45",
      "content": "<p>Thanks for the description. It is greatly appreciated.</p>",
      "rawMarkdown": "Thanks for the description. It is greatly appreciated.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3212040,
      "author_name": "qimcenter",
      "author_url": "",
      "post_date": "05/29/2025 08:24:45",
      "content": "<p>Thanks for the description. It is greatly appreciated.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3211968": "I thank the competition hosts for organizing this challenging competition and learning opportunity. The diagram summarizes my approach:\n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F843acf737a1a330a3e654c4c42eba6da%2Fpipeline2.png?generation=1748496646746480&alt=media)\n\n# Feature engineering and data augmentation\n\nDuring the first days of the competition, I tried my luck with linear models and feature engineering based on the distribution of the mass in the 3d images. The notebook [Forams Linear Discriminant Analysis](https://www.kaggle.com/code/ambrosm/forams-linear-discriminant-analysis) shows the approach. After publishing this notebook, I added a few more features with the intention of engineering features which do not depend on the spatial orientation of the sample. For instance, the eigenvalues of the inertia tensor are invariant under rotation, and they help distinguish objects with a round appearance (e.g., classes 10 and 11) from flat objects (e.g., classes 5 and 6). Logistic regression with these features got a public leaderboard score of 0.66296.\n\nAn EDA shows that while most class 10 samples are spheres with an empty interior, a few samples of this class are filled with dirt. The dirt in the interior evidently should not matter for classification, but produces noise in features such as mass and inertia tensor. This observation suggests that classification should be based on views of the surface rather than the mass in the interior.\n\nI then computed six 127x127 surface views for every 128x128x128 cube. Flipping and rotating the surface views augments the training data (the image below shows different views of the first training sample). Resnet-50 then embeds every surface view into a 2048-dimensional space. These 2048 Resnet activations are fed into a neural network with a single hidden layer and 14 softmax outputs. This model reaches a public score of 0.82340.\n\n![surface views](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F62fc4f60c87d9829dd1982b89dddaaa9%2Fviews.png?generation=1748496891410102&alt=media)\n\n# Semi-supervised learning\n\nI first tried to use scikit-learn's LabelPropagation to implement semi-supervised learning. Its results weren't good. Pseudolabelling had more success: I compared the predictions of my first four models. For 41 % of the test samples, all four models agreed (i.e., they predicted the same class). I added these 41 % of the test samples to the labelled data and trained a new supervised neural network, which got a public score of 0.86621. This classifier never predicts class 14.\n\n# Outlier detection\n\nPart of the competition task was finding samples which belong to none of the 14 given species. Experiments with IsolationForest had little success, and so I decided to assign samples to class 14 simply if the classifier showed low confidence. \n\n# Future work\n\nMy neural network models work only with six axis-parallel views onto the 3d object (top, bottom, left, right, front, back). It will be interesting to see how the classification improves if the data are augmented with views from other angles.\n\nSource code is [here](https://www.kaggle.com/code/ambrosm/forams-pseudolabeling-and-resnet-50?scriptVersionId=241567496).",
    "3212040": "Thanks for the description. It is greatly appreciated."
  },
  "source": "meta"
}