{
  "id": 549715,
  "title": "How to learn the most from this competition ?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/549715",
  "author_name": "",
  "post_date": "2024-12-03T15:32:59.559712400Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>So this is my first ever attempt at a \"featured competition\", I have joined this competition with the goal of learn the most, but as of now I just have learned a bit about the data and object detection, I have not previously worked with CNN but I have a good understanding of how they work. I was wondering what are the approaches you guys take for these competition in general, I’m trying to figure out the best approach for tackling competitions like this, as I have only particpated in tabular competitions, where I would:<br>\n    1. Understand the data,<br>\n    2. Perform feature engineering,<br>\n    3. Build a pipeline, and<br>\n    4. Fit a model.</p>\n<p>My questions are :</p>\n<ol>\n<li><p>What does the typical process look like for featured competitions (like those involving computer vision, object detection, etc.)?</p>\n<ul>\n<li>How do you guys approach the problem from start to finish?</li></ul></li>\n<li><p>How do you prepare for competitions?</p>\n<ul>\n<li>Do you read research papers to understand the domain or task?</li>\n<li>Do you analyze and learn from other people’s notebooks or code?</li></ul></li>\n<li><p>How do you plan for a competition if you’re not already proficient in the required techniques?</p>\n<ul>\n<li>For example, if you were new to object detection, how would you go about learning it?</li>\n<li>What steps would you take to become skilled enough to achieve a competitive position on the leaderboard?</li></ul></li>\n</ol>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": "3062442",
      "postDate": "12/03/2024 15:32:59",
      "content": "<p>So this is my first ever attempt at a \"featured competition\", I have joined this competition with the goal of learn the most, but as of now I just have learned a bit about the data and object detection, I have not previously worked with CNN but I have a good understanding of how they work. I was wondering what are the approaches you guys take for these competition in general, I’m trying to figure out the best approach for tackling competitions like this, as I have only particpated in tabular competitions, where I would:<br>\n    1. Understand the data,<br>\n    2. Perform feature engineering,<br>\n    3. Build a pipeline, and<br>\n    4. Fit a model.</p>\n<p>My questions are :</p>\n<ol>\n<li><p>What does the typical process look like for featured competitions (like those involving computer vision, object detection, etc.)?</p>\n<ul>\n<li>How do you guys approach the problem from start to finish?</li></ul></li>\n<li><p>How do you prepare for competitions?</p>\n<ul>\n<li>Do you read research papers to understand the domain or task?</li>\n<li>Do you analyze and learn from other people’s notebooks or code?</li></ul></li>\n<li><p>How do you plan for a competition if you’re not already proficient in the required techniques?</p>\n<ul>\n<li>For example, if you were new to object detection, how would you go about learning it?</li>\n<li>What steps would you take to become skilled enough to achieve a competitive position on the leaderboard?</li></ul></li>\n</ol>\n<p>Thank you.</p>",
      "rawMarkdown": "So this is my first ever attempt at a \"featured competition\", I have joined this competition with the goal of learn the most, but as of now I just have learned a bit about the data and object detection, I have not previously worked with CNN but I have a good understanding of how they work. I was wondering what are the approaches you guys take for these competition in general, I’m trying to figure out the best approach for tackling competitions like this, as I have only particpated in tabular competitions, where I would:\n    1. Understand the data,\n    2. Perform feature engineering,\n    3. Build a pipeline, and\n    4. Fit a model.\n\nMy questions are :\n1. What does the typical process look like for featured competitions (like those involving computer vision, object detection, etc.)?\n  - How do you guys approach the problem from start to finish?\n\n2. How do you prepare for competitions?\n  - Do you read research papers to understand the domain or task?\n  - Do you analyze and learn from other people’s notebooks or code?\n\n3. How do you plan for a competition if you’re not already proficient in the required techniques?\n  - For example, if you were new to object detection, how would you go about learning it?\n  - What steps would you take to become skilled enough to achieve a competitive position on the leaderboard?\n\nThank you.",
      "votes": null
    },
    {
      "id": "3063016",
      "postDate": "12/04/2024 05:10:11",
      "content": "<p>Start here to understand the data (I may be a little biased on this one):<br>\n<a href=\"https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization\" target=\"_blank\">https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization</a></p>\n<p>Go through this notebook line by line and understand what it does.  Submit it, and then figure out ways to make it better:<br>\n<a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">https://www.kaggle.com/code/fnands/baseline-unet-train-submit</a></p>\n<p>Visit this site:<br>\n<a href=\"https://cryoetdataportal.czscience.com/competition\" target=\"_blank\">https://cryoetdataportal.czscience.com/competition</a><br>\nAnd this:<br>\n<a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks</a><br>\nAnd this:<br>\n<a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a></p>\n<p>Read the discussion posts.  Pay special attention to any posts by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.  For instance:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350</a></p>\n<p>Hope this helps!</p>",
      "rawMarkdown": "Start here to understand the data (I may be a little biased on this one):\n[https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization](https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization)\n\nGo through this notebook line by line and understand what it does.  Submit it, and then figure out ways to make it better:\nhttps://www.kaggle.com/code/fnands/baseline-unet-train-submit\n\nVisit this site:\n[https://cryoetdataportal.czscience.com/competition](https://cryoetdataportal.czscience.com/competition)\nAnd this:\n[https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks)\nAnd this:\n[https://cryoetdataportal.czscience.com/datasets/10441](https://cryoetdataportal.czscience.com/datasets/10441)\n\nRead the discussion posts.  Pay special attention to any posts by @hengck23.  For instance:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350)\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "3063074",
      "postDate": "12/04/2024 06:13:53",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> for this. </p>",
      "rawMarkdown": "Thank you so much @davidlist for this.",
      "votes": null
    },
    {
      "id": "3063577",
      "postDate": "12/04/2024 16:16:06",
      "content": "<p>Hi David, could you please share some knowledge on how to get data labels from the '10441' dataset</p>",
      "rawMarkdown": "Hi David, could you please share some knowledge on how to get data labels from the '10441' dataset",
      "votes": null
    },
    {
      "id": "3063738",
      "postDate": "12/04/2024 20:06:45",
      "content": "<p>I'm only just downloading it now for the first time.  First impression is the directory structure is a bit different so understanding how to use it may be its own little project.</p>",
      "rawMarkdown": "I'm only just downloading it now for the first time.  First impression is the directory structure is a bit different so understanding how to use it may be its own little project.",
      "votes": null
    },
    {
      "id": "3064711",
      "postDate": "12/05/2024 23:28:55",
      "content": "<p>I think this notebook goes a long way towards an answer:<br>\n<a href=\"https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook\" target=\"_blank\">https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook</a></p>",
      "rawMarkdown": "I think this notebook goes a long way towards an answer:\n[https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook](https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook)",
      "votes": null
    },
    {
      "id": "3067164",
      "postDate": "12/09/2024 01:57:58",
      "content": "<p>Or actually, just use this notebook.  <a href=\"https://www.kaggle.com/code/davidlist/simulated-data-and-labels\" target=\"_blank\">https://www.kaggle.com/code/davidlist/simulated-data-and-labels</a>  :)</p>",
      "rawMarkdown": "Or actually, just use this notebook.  [https://www.kaggle.com/code/davidlist/simulated-data-and-labels](https://www.kaggle.com/code/davidlist/simulated-data-and-labels)  :)",
      "votes": null
    },
    {
      "id": "3067310",
      "postDate": "12/09/2024 06:30:06",
      "content": "<p>thanks a lot. </p>",
      "rawMarkdown": "thanks a lot.",
      "votes": null
    },
    {
      "id": "3087229",
      "postDate": "01/03/2025 09:00:58",
      "content": "<p>Hi! Thank you for your advice. I've gone through the baseline U-Net model. Could you please share some suggestions or directions on how to further improve it?</p>",
      "rawMarkdown": "Hi! Thank you for your advice. I've gone through the baseline U-Net model. Could you please share some suggestions or directions on how to further improve it?",
      "votes": null
    },
    {
      "id": "3087837",
      "postDate": "01/04/2025 00:06:10",
      "content": "<p>There are quite a few bugs in that notebook that affect performance, but probably the most important exercise relates to understanding the concept behind a precision-recall curve and how to use that to optimize your f4 score.  Doing that should get you well into the 0.500's.</p>",
      "rawMarkdown": "There are quite a few bugs in that notebook that affect performance, but probably the most important exercise relates to understanding the concept behind a precision-recall curve and how to use that to optimize your f4 score.  Doing that should get you well into the 0.500's.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3063016,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "12/04/2024 05:10:11",
      "content": "<p>Start here to understand the data (I may be a little biased on this one):<br>\n<a href=\"https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization\" target=\"_blank\">https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization</a></p>\n<p>Go through this notebook line by line and understand what it does.  Submit it, and then figure out ways to make it better:<br>\n<a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">https://www.kaggle.com/code/fnands/baseline-unet-train-submit</a></p>\n<p>Visit this site:<br>\n<a href=\"https://cryoetdataportal.czscience.com/competition\" target=\"_blank\">https://cryoetdataportal.czscience.com/competition</a><br>\nAnd this:<br>\n<a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks</a><br>\nAnd this:<br>\n<a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a></p>\n<p>Read the discussion posts.  Pay special attention to any posts by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.  For instance:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350</a></p>\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3063074,
          "author_name": "aquibahmad7",
          "author_url": "",
          "post_date": "12/04/2024 06:13:53",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> for this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3063577,
          "author_name": "stronggaylocksman",
          "author_url": "",
          "post_date": "12/04/2024 16:16:06",
          "content": "<p>Hi David, could you please share some knowledge on how to get data labels from the '10441' dataset</p>",
          "votes": null,
          "replies": [
            {
              "id": 3063738,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "12/04/2024 20:06:45",
              "content": "<p>I'm only just downloading it now for the first time.  First impression is the directory structure is a bit different so understanding how to use it may be its own little project.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3064711,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "12/05/2024 23:28:55",
              "content": "<p>I think this notebook goes a long way towards an answer:<br>\n<a href=\"https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook\" target=\"_blank\">https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook</a></p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3067164,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "12/09/2024 01:57:58",
              "content": "<p>Or actually, just use this notebook.  <a href=\"https://www.kaggle.com/code/davidlist/simulated-data-and-labels\" target=\"_blank\">https://www.kaggle.com/code/davidlist/simulated-data-and-labels</a>  :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3067310,
                  "author_name": "stronggaylocksman",
                  "author_url": "",
                  "post_date": "12/09/2024 06:30:06",
                  "content": "<p>thanks a lot. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3087229,
          "author_name": "biancochiu",
          "author_url": "",
          "post_date": "01/03/2025 09:00:58",
          "content": "<p>Hi! Thank you for your advice. I've gone through the baseline U-Net model. Could you please share some suggestions or directions on how to further improve it?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3087837,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "01/04/2025 00:06:10",
              "content": "<p>There are quite a few bugs in that notebook that affect performance, but probably the most important exercise relates to understanding the concept behind a precision-recall curve and how to use that to optimize your f4 score.  Doing that should get you well into the 0.500's.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3062442": "So this is my first ever attempt at a \"featured competition\", I have joined this competition with the goal of learn the most, but as of now I just have learned a bit about the data and object detection, I have not previously worked with CNN but I have a good understanding of how they work. I was wondering what are the approaches you guys take for these competition in general, I’m trying to figure out the best approach for tackling competitions like this, as I have only particpated in tabular competitions, where I would:\n    1. Understand the data,\n    2. Perform feature engineering,\n    3. Build a pipeline, and\n    4. Fit a model.\n\nMy questions are :\n1. What does the typical process look like for featured competitions (like those involving computer vision, object detection, etc.)?\n  - How do you guys approach the problem from start to finish?\n\n2. How do you prepare for competitions?\n  - Do you read research papers to understand the domain or task?\n  - Do you analyze and learn from other people’s notebooks or code?\n\n3. How do you plan for a competition if you’re not already proficient in the required techniques?\n  - For example, if you were new to object detection, how would you go about learning it?\n  - What steps would you take to become skilled enough to achieve a competitive position on the leaderboard?\n\nThank you.",
    "3063016": "Start here to understand the data (I may be a little biased on this one):\n[https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization](https://www.kaggle.com/code/davidlist/experiment-ts-6-4-visualization)\n\nGo through this notebook line by line and understand what it does.  Submit it, and then figure out ways to make it better:\nhttps://www.kaggle.com/code/fnands/baseline-unet-train-submit\n\nVisit this site:\n[https://cryoetdataportal.czscience.com/competition](https://cryoetdataportal.czscience.com/competition)\nAnd this:\n[https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks)\nAnd this:\n[https://cryoetdataportal.czscience.com/datasets/10441](https://cryoetdataportal.czscience.com/datasets/10441)\n\nRead the discussion posts.  Pay special attention to any posts by @hengck23.  For instance:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547350)\n\nHope this helps!",
    "3063074": "Thank you so much @davidlist for this.",
    "3063577": "Hi David, could you please share some knowledge on how to get data labels from the '10441' dataset",
    "3063738": "I'm only just downloading it now for the first time.  First impression is the directory structure is a bit different so understanding how to use it may be its own little project.",
    "3064711": "I think this notebook goes a long way towards an answer:\n[https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook](https://www.kaggle.com/code/davidlist/ts-0-simulated-data-visualization/notebook)",
    "3067164": "Or actually, just use this notebook.  [https://www.kaggle.com/code/davidlist/simulated-data-and-labels](https://www.kaggle.com/code/davidlist/simulated-data-and-labels)  :)",
    "3067310": "thanks a lot.",
    "3087229": "Hi! Thank you for your advice. I've gone through the baseline U-Net model. Could you please share some suggestions or directions on how to further improve it?",
    "3087837": "There are quite a few bugs in that notebook that affect performance, but probably the most important exercise relates to understanding the concept behind a precision-recall curve and how to use that to optimize your f4 score.  Doing that should get you well into the 0.500's."
  },
  "source": "meta"
}