{
  "id": 643289,
  "title": "top10 submission",
  "url": "/competitions/the-3lc-cotton-weed-detection-challenge/writeups/top10-submission",
  "author_name": "",
  "post_date": "2025-11-28T19:18:01.697Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>👋 Hi everyone, I would like to provide the pipeline of the ninjas MIREA x MISIS team.</p>\n<p>Initially, we wrote to the organizers about what is possible and what is not. And we got this answer:</p>\n<p>This is a DATA QUALITY competition.</p>\n<p>Following is what is Allowed:</p>\n<ol>\n<li>Modify training data: add missing labels, adjust bad labels, delete wrong ones</li>\n<li>Change all hyperparameters and training strategies</li>\n<li>Confidence threshold optimization</li>\n</ol>\n<p>The following is Not Allowed:</p>\n<ol>\n<li>The trained model must run inference out of the box on test data with no modifications</li>\n<li>Pseudo-labeling test data - only train/val data can be used in the training process</li>\n</ol>\n<p>Core principle: Fix the data quality on train/val sets, tune your training, but inference must be standard single-pass YOLOv8n prediction.Hope this clarifies!</p>\n<p>In fact, these are very serious restrictions, because it was forbidden:</p>\n<ol>\n<li>Take a different model</li>\n<li>Take another dataset (only change the markup)</li>\n<li>Use SAHI, TTA, sliding windows, etc</li>\n</ol>\n<p>Our final result uses only data from train/val, we did not use other datasets, did not change the model, etc. The main goal that we promoted on private is that there are errors in the data and they have poor markup (perhaps because of this, we did not manage to move up on private).</p>\n<p>The main advantage of our solution was to find the best parameters and use the SGD optimizer (because it provides better convergence)</p>\n<p>Initially, we got 0.85 on the public and then there was the idea of fine-tuning the 0.85 model</p>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/weed-detection-model\" target=\"_blank\">Learn model Yolov8n 0.85</a></h2>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/predicts-weeds\" target=\"_blank\">Model 0.85 predict</a></h2>\n<p>Next, we fine-tuned our models, used OPTUNA and learn and validate on our images(created with inaturalist.org). (since it is not forbidden to validate and search for better parameters on another dataset)</p>\n<h2><a href=\"https://www.kaggle.com/datasets/antonoof/internetdataweed\" target=\"_blank\">Internet Dataset</a></h2>\n<p>Next, we take best params and tuned yolo8n(0.85 score public)</p>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data\" target=\"_blank\">Fine-Tunung Learn Optuna</a></h2>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts?scriptVersionId=281421591\" target=\"_blank\">Fine-Tunung Predict Optuna</a></h2>\n<p>The task was on 3LC, honestly we tried to use it, it's really a very powerful and powerful method that I will use in the future, but it seemed to us that 3LC improves our local score and worsens public and private score (maybe there was a mistake). We tried using other datasets, etc., and as a result, we agreed on the solution above.</p>\n<p>🎉 I would like to express my deep gratitude to the organizers. This is a really good and cool experience. I really hope for an honest assessment and kick from people who have used other datasets, cheats, models or ways to improve the result, because we tried to work honestly.</p>",
  "messages": [
    {
      "id": "3351947",
      "postDate": "11/28/2025 18:48:57",
      "content": "<p>👋 Hi everyone, I would like to provide the pipeline of the ninjas MIREA x MISIS team.</p>\n<p>Initially, we wrote to the organizers about what is possible and what is not. And we got this answer:</p>\n<p>This is a DATA QUALITY competition.</p>\n<p>Following is what is Allowed:</p>\n<ol>\n<li>Modify training data: add missing labels, adjust bad labels, delete wrong ones</li>\n<li>Change all hyperparameters and training strategies</li>\n<li>Confidence threshold optimization</li>\n</ol>\n<p>The following is Not Allowed:</p>\n<ol>\n<li>The trained model must run inference out of the box on test data with no modifications</li>\n<li>Pseudo-labeling test data - only train/val data can be used in the training process</li>\n</ol>\n<p>Core principle: Fix the data quality on train/val sets, tune your training, but inference must be standard single-pass YOLOv8n prediction.Hope this clarifies!</p>\n<p>In fact, these are very serious restrictions, because it was forbidden:</p>\n<ol>\n<li>Take a different model</li>\n<li>Take another dataset (only change the markup)</li>\n<li>Use SAHI, TTA, sliding windows, etc</li>\n</ol>\n<p>Our final result uses only data from train/val, we did not use other datasets, did not change the model, etc. The main goal that we promoted on private is that there are errors in the data and they have poor markup (perhaps because of this, we did not manage to move up on private).</p>\n<p>The main advantage of our solution was to find the best parameters and use the SGD optimizer (because it provides better convergence)</p>\n<p>Initially, we got 0.85 on the public and then there was the idea of fine-tuning the 0.85 model</p>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/weed-detection-model\" target=\"_blank\">Learn model Yolov8n 0.85</a></h2>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/predicts-weeds\" target=\"_blank\">Model 0.85 predict</a></h2>\n<p>Next, we fine-tuned our models, used OPTUNA and learn and validate on our images(created with inaturalist.org). (since it is not forbidden to validate and search for better parameters on another dataset)</p>\n<h2><a href=\"https://www.kaggle.com/datasets/antonoof/internetdataweed\" target=\"_blank\">Internet Dataset</a></h2>\n<p>Next, we take best params and tuned yolo8n(0.85 score public)</p>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data\" target=\"_blank\">Fine-Tunung Learn Optuna</a></h2>\n<h2>✔️ <a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts?scriptVersionId=281421591\" target=\"_blank\">Fine-Tunung Predict Optuna</a></h2>\n<p>The task was on 3LC, honestly we tried to use it, it's really a very powerful and powerful method that I will use in the future, but it seemed to us that 3LC improves our local score and worsens public and private score (maybe there was a mistake). We tried using other datasets, etc., and as a result, we agreed on the solution above.</p>\n<p>🎉 I would like to express my deep gratitude to the organizers. This is a really good and cool experience. I really hope for an honest assessment and kick from people who have used other datasets, cheats, models or ways to improve the result, because we tried to work honestly.</p>",
      "rawMarkdown": "👋 Hi everyone, I would like to provide the pipeline of the ninjas MIREA x MISIS team.\n\nInitially, we wrote to the organizers about what is possible and what is not. And we got this answer:\n\nThis is a DATA QUALITY competition.\n\nFollowing is what is Allowed:\n1. Modify training data: add missing labels, adjust bad labels, delete wrong ones\n2. Change all hyperparameters and training strategies\n3. Confidence threshold optimization\n\nThe following is Not Allowed:\n1. The trained model must run inference out of the box on test data with no modifications\n2. Pseudo-labeling test data - only train/val data can be used in the training process\n\nCore principle: Fix the data quality on train/val sets, tune your training, but inference must be standard single-pass YOLOv8n prediction.Hope this clarifies!\n\nIn fact, these are very serious restrictions, because it was forbidden:\n1. Take a different model\n2. Take another dataset (only change the markup)\n3. Use SAHI, TTA, sliding windows, etc\n\nOur final result uses only data from train/val, we did not use other datasets, did not change the model, etc. The main goal that we promoted on private is that there are errors in the data and they have poor markup (perhaps because of this, we did not manage to move up on private).\n\nThe main advantage of our solution was to find the best parameters and use the SGD optimizer (because it provides better convergence)\n\nInitially, we got 0.85 on the public and then there was the idea of fine-tuning the 0.85 model\n\n## ✔️ [Learn model Yolov8n 0.85](https://www.kaggle.com/code/antonoof/weed-detection-model)\n\n## ✔️ [Model 0.85 predict](https://www.kaggle.com/code/antonoof/predicts-weeds)\n\nNext, we fine-tuned our models, used OPTUNA and learn and validate on our images(created with inaturalist.org). (since it is not forbidden to validate and search for better parameters on another dataset)\n\n## [Internet Dataset](https://www.kaggle.com/datasets/antonoof/internetdataweed)\n\nNext, we take best params and tuned yolo8n(0.85 score public)\n\n## ✔️ [Fine-Tunung Learn Optuna](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data)\n\n## ✔️ [Fine-Tunung Predict Optuna](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts?scriptVersionId=281421591)\n\nThe task was on 3LC, honestly we tried to use it, it's really a very powerful and powerful method that I will use in the future, but it seemed to us that 3LC improves our local score and worsens public and private score (maybe there was a mistake). We tried using other datasets, etc., and as a result, we agreed on the solution above.\n\n🎉 I would like to express my deep gratitude to the organizers. This is a really good and cool experience. I really hope for an honest assessment and kick from people who have used other datasets, cheats, models or ways to improve the result, because we tried to work honestly.",
      "votes": null
    },
    {
      "id": "3351973",
      "postDate": "11/28/2025 19:13:06",
      "content": "<p>I'm sorry, I do not know why the links to the final model are not being corrected, I mistakenly attached another optuna prediction with a bad score. I have attached the final training and model prediction:</p>\n<h2><a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data\" target=\"_blank\">Optuna params learn</a></h2>\n<h2><a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts\" target=\"_blank\">model predict</a></h2>",
      "rawMarkdown": "I'm sorry, I do not know why the links to the final model are not being corrected, I mistakenly attached another optuna prediction with a bad score. I have attached the final training and model prediction:\n\n## [Optuna params learn](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data)\n\n## [model predict](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts)",
      "votes": null
    },
    {
      "id": "3352011",
      "postDate": "11/28/2025 19:58:34",
      "content": "<h1>Bonus:  <a href=\"https://www.kaggle.com/code/antonoof/top-1-2-private-clear\" target=\"_blank\">top1 score submisison</a></h1>",
      "rawMarkdown": "# Bonus:  [top1 score submisison](https://www.kaggle.com/code/antonoof/top-1-2-private-clear)",
      "votes": null
    },
    {
      "id": "3352034",
      "postDate": "11/28/2025 20:27:07",
      "content": "<p>Did you by any chance correct any incorrect labels from the train/val dataset and again there were images with multiple annotation like 12 annotation just for a single class especially palmer_amarath</p>",
      "rawMarkdown": "Did you by any chance correct any incorrect labels from the train/val dataset and again there were images with multiple annotation like 12 annotation just for a single class especially palmer_amarath",
      "votes": null
    },
    {
      "id": "3352036",
      "postDate": "11/28/2025 20:31:58",
      "content": "<p>Hi, we've reviewed the entire train/val, we've seen some minor bugs, but it didn't seem critical to us. We tried to fix it using 3LC, we used our(local pseudo) three lines of code, but it seemed to us that this was a bad idea, because the labels were getting better, the score was growing locally and it was getting worse in public, so we assumed that the test data also had bad markup and did not touch train/val as a result.</p>",
      "rawMarkdown": "Hi, we've reviewed the entire train/val, we've seen some minor bugs, but it didn't seem critical to us. We tried to fix it using 3LC, we used our(local pseudo) three lines of code, but it seemed to us that this was a bad idea, because the labels were getting better, the score was growing locally and it was getting worse in public, so we assumed that the test data also had bad markup and did not touch train/val as a result.",
      "votes": null
    },
    {
      "id": "3352037",
      "postDate": "11/28/2025 20:34:58",
      "content": "<p>so what you majorly focused on was fine-tuning the model  parameters to yield the better results </p>",
      "rawMarkdown": "so what you majorly focused on was fine-tuning the model  parameters to yield the better results",
      "votes": null
    },
    {
      "id": "3352047",
      "postDate": "11/28/2025 20:47:58",
      "content": "<p>data understanding, fine tuning improvement, optuna, SGD and parameters. Conf -&gt; 0 =&gt; better score</p>",
      "rawMarkdown": "data understanding, fine tuning improvement, optuna, SGD and parameters. Conf -> 0 => better score",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3351973,
      "author_name": "antonoof",
      "author_url": "",
      "post_date": "11/28/2025 19:13:06",
      "content": "<p>I'm sorry, I do not know why the links to the final model are not being corrected, I mistakenly attached another optuna prediction with a bad score. I have attached the final training and model prediction:</p>\n<h2><a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data\" target=\"_blank\">Optuna params learn</a></h2>\n<h2><a href=\"https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts\" target=\"_blank\">model predict</a></h2>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3352011,
      "author_name": "antonoof",
      "author_url": "",
      "post_date": "11/28/2025 19:58:34",
      "content": "<h1>Bonus:  <a href=\"https://www.kaggle.com/code/antonoof/top-1-2-private-clear\" target=\"_blank\">top1 score submisison</a></h1>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3352034,
      "author_name": "bingos",
      "author_url": "",
      "post_date": "11/28/2025 20:27:07",
      "content": "<p>Did you by any chance correct any incorrect labels from the train/val dataset and again there were images with multiple annotation like 12 annotation just for a single class especially palmer_amarath</p>",
      "votes": null,
      "replies": [
        {
          "id": 3352036,
          "author_name": "antonoof",
          "author_url": "",
          "post_date": "11/28/2025 20:31:58",
          "content": "<p>Hi, we've reviewed the entire train/val, we've seen some minor bugs, but it didn't seem critical to us. We tried to fix it using 3LC, we used our(local pseudo) three lines of code, but it seemed to us that this was a bad idea, because the labels were getting better, the score was growing locally and it was getting worse in public, so we assumed that the test data also had bad markup and did not touch train/val as a result.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3352037,
              "author_name": "bingos",
              "author_url": "",
              "post_date": "11/28/2025 20:34:58",
              "content": "<p>so what you majorly focused on was fine-tuning the model  parameters to yield the better results </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3352047,
                  "author_name": "antonoof",
                  "author_url": "",
                  "post_date": "11/28/2025 20:47:58",
                  "content": "<p>data understanding, fine tuning improvement, optuna, SGD and parameters. Conf -&gt; 0 =&gt; better score</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3351947": "👋 Hi everyone, I would like to provide the pipeline of the ninjas MIREA x MISIS team.\n\nInitially, we wrote to the organizers about what is possible and what is not. And we got this answer:\n\nThis is a DATA QUALITY competition.\n\nFollowing is what is Allowed:\n1. Modify training data: add missing labels, adjust bad labels, delete wrong ones\n2. Change all hyperparameters and training strategies\n3. Confidence threshold optimization\n\nThe following is Not Allowed:\n1. The trained model must run inference out of the box on test data with no modifications\n2. Pseudo-labeling test data - only train/val data can be used in the training process\n\nCore principle: Fix the data quality on train/val sets, tune your training, but inference must be standard single-pass YOLOv8n prediction.Hope this clarifies!\n\nIn fact, these are very serious restrictions, because it was forbidden:\n1. Take a different model\n2. Take another dataset (only change the markup)\n3. Use SAHI, TTA, sliding windows, etc\n\nOur final result uses only data from train/val, we did not use other datasets, did not change the model, etc. The main goal that we promoted on private is that there are errors in the data and they have poor markup (perhaps because of this, we did not manage to move up on private).\n\nThe main advantage of our solution was to find the best parameters and use the SGD optimizer (because it provides better convergence)\n\nInitially, we got 0.85 on the public and then there was the idea of fine-tuning the 0.85 model\n\n## ✔️ [Learn model Yolov8n 0.85](https://www.kaggle.com/code/antonoof/weed-detection-model)\n\n## ✔️ [Model 0.85 predict](https://www.kaggle.com/code/antonoof/predicts-weeds)\n\nNext, we fine-tuned our models, used OPTUNA and learn and validate on our images(created with inaturalist.org). (since it is not forbidden to validate and search for better parameters on another dataset)\n\n## [Internet Dataset](https://www.kaggle.com/datasets/antonoof/internetdataweed)\n\nNext, we take best params and tuned yolo8n(0.85 score public)\n\n## ✔️ [Fine-Tunung Learn Optuna](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data)\n\n## ✔️ [Fine-Tunung Predict Optuna](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts?scriptVersionId=281421591)\n\nThe task was on 3LC, honestly we tried to use it, it's really a very powerful and powerful method that I will use in the future, but it seemed to us that 3LC improves our local score and worsens public and private score (maybe there was a mistake). We tried using other datasets, etc., and as a result, we agreed on the solution above.\n\n🎉 I would like to express my deep gratitude to the organizers. This is a really good and cool experience. I really hope for an honest assessment and kick from people who have used other datasets, cheats, models or ways to improve the result, because we tried to work honestly.",
    "3351973": "I'm sorry, I do not know why the links to the final model are not being corrected, I mistakenly attached another optuna prediction with a bad score. I have attached the final training and model prediction:\n\n## [Optuna params learn](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-data)\n\n## [model predict](https://www.kaggle.com/code/antonoof/fine-tuning-optuna-predicts)",
    "3352011": "# Bonus:  [top1 score submisison](https://www.kaggle.com/code/antonoof/top-1-2-private-clear)",
    "3352034": "Did you by any chance correct any incorrect labels from the train/val dataset and again there were images with multiple annotation like 12 annotation just for a single class especially palmer_amarath",
    "3352036": "Hi, we've reviewed the entire train/val, we've seen some minor bugs, but it didn't seem critical to us. We tried to fix it using 3LC, we used our(local pseudo) three lines of code, but it seemed to us that this was a bad idea, because the labels were getting better, the score was growing locally and it was getting worse in public, so we assumed that the test data also had bad markup and did not touch train/val as a result.",
    "3352037": "so what you majorly focused on was fine-tuning the model  parameters to yield the better results",
    "3352047": "data understanding, fine tuning improvement, optuna, SGD and parameters. Conf -> 0 => better score"
  },
  "source": "meta"
}