{
  "id": 425729,
  "title": "3rd place solution",
  "url": "/competitions/planttraits2023/discussion/425729",
  "author_name": "",
  "post_date": "2023-07-20T07:20:59.649926600Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, I would like to express deep gratitude to the competition organisers and the Kaggle team.The competition covered a wide range of knowledge, including data processing, deep learning, and combining image and weather data, making it one of the most challenging and interesting contests in my own Kaggle history.Without my wonderful teammates this ranking would not have been possible and it was indeed a pleasure for me to work with them.</p>\n<h1>Overview</h1>\n<h2>The key points of my solution are:</h2>\n<ol>\n<li>Reasonable data processing</li>\n<li>Powerful visual model and a simple way to fuse data</li>\n</ol>\n<h2>Data processing</h2>\n<p>We found large numerical differences in climate data.This leads to two problems. The first is that our model is difficult to converge over a large range of data and does not find the local optimal solution quickly. Secondly, the huge scale difference between features can lead to certain features dominating the training and the model is not able to learn the relevant features in the data well.</p>\n<p>Therefore, we preprocess the climate data and label data during the training process, i.e., log the data first and then normalize them second. This allows us to narrow down the climate data and label data to a reasonable range. For the image data we only use imagenet's normalization parameters for normalization. It is worth noting that we only used the normalized parameters in the training set to restore the predicted values.</p>\n<h2>Overall structure of the model</h2>\n<ul>\n<li><p>Image encoder<br>\nConvnext xlarge(imagenet pre-trained)</p></li>\n<li><p>Climate encoder<br>\nSimple one fully-connected layer 、LayerNorm、and ReLU</p></li>\n<li><p>Prediction head<br>\nSimple three fully-connected layer 、LayerNorm，and ReLU<br>\nWe directly concatenate the output of Image enconder and climate encoder as input to prediction head.</p></li>\n<li><p>Image data augmentation<br>\nHorizontalFlip(p=0.5)<br>\nVerticalFlip(p=0.5)</p></li>\n<li><p>Other detailed training settings<br>\nInput image size : 512×512<br>\n20 epochs<br>\nLoss：SmoothL1 loss<br>\nOptimizer：AdamW</p></li>\n</ul>\n<h2>Some other findings</h2>\n<p>The commonly used image TTA strategy and Ensemble prediction approach did not work in this competition.</p>",
  "messages": [
    {
      "id": "2351495",
      "postDate": "07/20/2023 07:20:59",
      "content": "<p>First of all, I would like to express deep gratitude to the competition organisers and the Kaggle team.The competition covered a wide range of knowledge, including data processing, deep learning, and combining image and weather data, making it one of the most challenging and interesting contests in my own Kaggle history.Without my wonderful teammates this ranking would not have been possible and it was indeed a pleasure for me to work with them.</p>\n<h1>Overview</h1>\n<h2>The key points of my solution are:</h2>\n<ol>\n<li>Reasonable data processing</li>\n<li>Powerful visual model and a simple way to fuse data</li>\n</ol>\n<h2>Data processing</h2>\n<p>We found large numerical differences in climate data.This leads to two problems. The first is that our model is difficult to converge over a large range of data and does not find the local optimal solution quickly. Secondly, the huge scale difference between features can lead to certain features dominating the training and the model is not able to learn the relevant features in the data well.</p>\n<p>Therefore, we preprocess the climate data and label data during the training process, i.e., log the data first and then normalize them second. This allows us to narrow down the climate data and label data to a reasonable range. For the image data we only use imagenet's normalization parameters for normalization. It is worth noting that we only used the normalized parameters in the training set to restore the predicted values.</p>\n<h2>Overall structure of the model</h2>\n<ul>\n<li><p>Image encoder<br>\nConvnext xlarge(imagenet pre-trained)</p></li>\n<li><p>Climate encoder<br>\nSimple one fully-connected layer 、LayerNorm、and ReLU</p></li>\n<li><p>Prediction head<br>\nSimple three fully-connected layer 、LayerNorm，and ReLU<br>\nWe directly concatenate the output of Image enconder and climate encoder as input to prediction head.</p></li>\n<li><p>Image data augmentation<br>\nHorizontalFlip(p=0.5)<br>\nVerticalFlip(p=0.5)</p></li>\n<li><p>Other detailed training settings<br>\nInput image size : 512×512<br>\n20 epochs<br>\nLoss：SmoothL1 loss<br>\nOptimizer：AdamW</p></li>\n</ul>\n<h2>Some other findings</h2>\n<p>The commonly used image TTA strategy and Ensemble prediction approach did not work in this competition.</p>",
      "rawMarkdown": "First of all, I would like to express deep gratitude to the competition organisers and the Kaggle team.The competition covered a wide range of knowledge, including data processing, deep learning, and combining image and weather data, making it one of the most challenging and interesting contests in my own Kaggle history.Without my wonderful teammates this ranking would not have been possible and it was indeed a pleasure for me to work with them.\n\n# Overview\n## The key points of my solution are:\n1. Reasonable data processing\n2. Powerful visual model and a simple way to fuse data\n\n## Data processing\nWe found large numerical differences in climate data.This leads to two problems. The first is that our model is difficult to converge over a large range of data and does not find the local optimal solution quickly. Secondly, the huge scale difference between features can lead to certain features dominating the training and the model is not able to learn the relevant features in the data well.\n\nTherefore, we preprocess the climate data and label data during the training process, i.e., log the data first and then normalize them second. This allows us to narrow down the climate data and label data to a reasonable range. For the image data we only use imagenet's normalization parameters for normalization. It is worth noting that we only used the normalized parameters in the training set to restore the predicted values.\n\n## Overall structure of the model\n- Image encoder\n   Convnext xlarge(imagenet pre-trained)\n- Climate encoder\n   Simple one fully-connected layer 、LayerNorm、and ReLU\n- Prediction head\n   Simple three fully-connected layer 、LayerNorm，and ReLU\n   We directly concatenate the output of Image enconder and climate encoder as input to prediction head.\n\n- Image data augmentation\n  HorizontalFlip(p=0.5)\n  VerticalFlip(p=0.5)\n- Other detailed training settings\n  Input image size : 512×512\n  20 epochs\n  Loss：SmoothL1 loss\n  Optimizer：AdamW\n\n## Some other findings\nThe commonly used image TTA strategy and Ensemble prediction approach did not work in this competition.",
      "votes": null
    },
    {
      "id": "2351515",
      "postDate": "07/20/2023 07:32:30",
      "content": "<p>good job. I want to know that  Input image size : 512×512 means the cropped image size or the original input size？</p>",
      "rawMarkdown": "good job. I want to know that  Input image size : 512×512 means the cropped image size or the original input size？",
      "votes": null
    },
    {
      "id": "2351518",
      "postDate": "07/20/2023 07:35:43",
      "content": "<p>The original size of the image is 512x512.</p>",
      "rawMarkdown": "The original size of the image is 512x512.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2351515,
      "author_name": "ninglee65536",
      "author_url": "",
      "post_date": "07/20/2023 07:32:30",
      "content": "<p>good job. I want to know that  Input image size : 512×512 means the cropped image size or the original input size？</p>",
      "votes": null,
      "replies": [
        {
          "id": 2351518,
          "author_name": "hubulai",
          "author_url": "",
          "post_date": "07/20/2023 07:35:43",
          "content": "<p>The original size of the image is 512x512.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2351495": "First of all, I would like to express deep gratitude to the competition organisers and the Kaggle team.The competition covered a wide range of knowledge, including data processing, deep learning, and combining image and weather data, making it one of the most challenging and interesting contests in my own Kaggle history.Without my wonderful teammates this ranking would not have been possible and it was indeed a pleasure for me to work with them.\n\n# Overview\n## The key points of my solution are:\n1. Reasonable data processing\n2. Powerful visual model and a simple way to fuse data\n\n## Data processing\nWe found large numerical differences in climate data.This leads to two problems. The first is that our model is difficult to converge over a large range of data and does not find the local optimal solution quickly. Secondly, the huge scale difference between features can lead to certain features dominating the training and the model is not able to learn the relevant features in the data well.\n\nTherefore, we preprocess the climate data and label data during the training process, i.e., log the data first and then normalize them second. This allows us to narrow down the climate data and label data to a reasonable range. For the image data we only use imagenet's normalization parameters for normalization. It is worth noting that we only used the normalized parameters in the training set to restore the predicted values.\n\n## Overall structure of the model\n- Image encoder\n   Convnext xlarge(imagenet pre-trained)\n- Climate encoder\n   Simple one fully-connected layer 、LayerNorm、and ReLU\n- Prediction head\n   Simple three fully-connected layer 、LayerNorm，and ReLU\n   We directly concatenate the output of Image enconder and climate encoder as input to prediction head.\n\n- Image data augmentation\n  HorizontalFlip(p=0.5)\n  VerticalFlip(p=0.5)\n- Other detailed training settings\n  Input image size : 512×512\n  20 epochs\n  Loss：SmoothL1 loss\n  Optimizer：AdamW\n\n## Some other findings\nThe commonly used image TTA strategy and Ensemble prediction approach did not work in this competition.",
    "2351515": "good job. I want to know that  Input image size : 512×512 means the cropped image size or the original input size？",
    "2351518": "The original size of the image is 512x512."
  },
  "source": "meta"
}