{
  "id": 560269,
  "title": "Seeking Assistance",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/560269",
  "author_name": "Alaa Sweed",
  "post_date": "2025-01-30T10:25:26.100000",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi, i just want to ask my fellow Kagglers how do they know what to do, where and when to do it regarding everything in this competition.</p>\n<p>I find it very interesting, fun , and wonderful to learn and participate in. <br>\nCan you tell me how can i learn about this type of analysis to be able to know what to do when faced with Data and objectives like that! I read so many genius amazing Notebooks and Discussions and studied it a lot but i still find the approaches and code hard to understand and when i open a notebook to start by myself i literally freeze because i don't know the process and approaches to such a problem!</p>\n<p>Thank you in advance for your advices 🙏♥️</p>\n<p>Note: idc about submission date and LB, i just want to learn about it even if the comp ended!</p>",
  "messages": [
    {
      "id": 3111744,
      "postDate": "2025-01-31T14:41:04.960Z",
      "content": "<p>Hey there!</p>\n<p>First off, huge props for jumping into Kaggle—it's awesome that you’re excited to learn and grow, and that mindset will definitely help you get where you want to be. It’s perfectly normal to feel overwhelmed when you're starting out, especially when you see all the advanced notebooks and discussions. But trust me, every Kaggle pro was once in your shoes!</p>\n<p>The most important thing to keep in mind is <strong>to start simple</strong>. It’s easy to get lost in the complexity of others’ work, but you don’t have to do everything at once. A solid foundation will help you build your confidence and skills over time.</p>\n<p>Here’s how I’d recommend approaching it:</p>\n<ol>\n<li><p><strong>Start with the basics</strong>: Focus on understanding basic concepts like data cleaning, exploration, feature engineering, and model building. Try to replicate simple notebooks step by step. Don’t worry if you don’t get it right away; every little thing you do will help you understand the bigger picture.</p></li>\n<li><p><strong>Follow a simple process</strong>: Many top Kagglers follow a basic workflow:</p>\n<ul>\n<li><strong>Exploration</strong>: Load your data and explore it (basic statistics, distributions, correlations).</li>\n<li><strong>Preprocessing</strong>: Handle missing data, outliers, and any data wrangling.</li>\n<li><strong>Modeling</strong>: Start with basic models like Resnet18 before moving to more advanced ones.</li>\n<li><strong>Evaluation</strong>: Check your model’s performance using the right metrics.</li></ul></li>\n<li><p><strong>Learn one concept at a time</strong>: For example, maybe this week focus on data exploration and visualization. Next week, dive into feature engineering or basic machine learning models. Trying to learn everything at once can be a recipe for frustration.</p></li>\n<li><p><strong>Don’t stress about the code</strong>: Instead, try to focus on <strong>why</strong> certain steps are being taken. Once you grasp the “why” behind it, understanding the code will follow more naturally. Try running the code in a notebook and modify it to see how things change.</p></li>\n<li><p><br>\nThis recommendation from LLM is actually wrong. I think working on a competition with LB feedback and a lively community is 100x better than working on a small project</p></li>\n<li><p><strong>Ask for help</strong>: Kaggle is full of people who are more than willing to help. Don’t hesitate to ask for clarification when something isn’t clear. You’ll often find that others are happy to walk you through their thought process.</p></li>\n</ol>\n<p>Finally, remember that you don’t have to rush. You’re in this to learn, and the competition part will always be secondary to gaining a solid understanding of how to approach data problems. Stay curious and enjoy the process—don’t be afraid to make mistakes, because that’s where the real learning happens!</p>\n<p>Good luck, and keep going! You'll get there! ✨</p>\n<p>(written by LLM, slightly edited)</p>",
      "rawMarkdown": "Hey there!\n\nFirst off, huge props for jumping into Kaggle—it's awesome that you’re excited to learn and grow, and that mindset will definitely help you get where you want to be. It’s perfectly normal to feel overwhelmed when you're starting out, especially when you see all the advanced notebooks and discussions. But trust me, every Kaggle pro was once in your shoes!\n\nThe most important thing to keep in mind is **to start simple**. It’s easy to get lost in the complexity of others’ work, but you don’t have to do everything at once. A solid foundation will help you build your confidence and skills over time.\n\nHere’s how I’d recommend approaching it:\n\n1. **Start with the basics**: Focus on understanding basic concepts like data cleaning, exploration, feature engineering, and model building. Try to replicate simple notebooks step by step. Don’t worry if you don’t get it right away; every little thing you do will help you understand the bigger picture.\n\n2. **Follow a simple process**: Many top Kagglers follow a basic workflow:\n   - **Exploration**: Load your data and explore it (basic statistics, distributions, correlations).\n   - **Preprocessing**: Handle missing data, outliers, and any data wrangling.\n   - **Modeling**: Start with basic models like Resnet18 before moving to more advanced ones.\n   - **Evaluation**: Check your model’s performance using the right metrics.\n\n3. **Learn one concept at a time**: For example, maybe this week focus on data exploration and visualization. Next week, dive into feature engineering or basic machine learning models. Trying to learn everything at once can be a recipe for frustration.\n\n4. **Don’t stress about the code**: Instead, try to focus on **why** certain steps are being taken. Once you grasp the “why” behind it, understanding the code will follow more naturally. Try running the code in a notebook and modify it to see how things change.\n\n5. ~~**Start small with your own projects**: After you get a feel for the process, try working on small datasets that interest you. You don’t need to be in a competition to practice! Gradually, you’ll feel more comfortable tackling larger, more complex problems.~~\nThis recommendation from LLM is actually wrong. I think working on a competition with LB feedback and a lively community is 100x better than working on a small project\n\n6. **Ask for help**: Kaggle is full of people who are more than willing to help. Don’t hesitate to ask for clarification when something isn’t clear. You’ll often find that others are happy to walk you through their thought process.\n\nFinally, remember that you don’t have to rush. You’re in this to learn, and the competition part will always be secondary to gaining a solid understanding of how to approach data problems. Stay curious and enjoy the process—don’t be afraid to make mistakes, because that’s where the real learning happens!\n\nGood luck, and keep going! You'll get there! ✨\n\n(written by LLM, slightly edited)",
      "votes": 9,
      "replies": [
        {
          "id": 3111878,
          "postDate": "2025-01-31T17:51:16.293Z",
          "content": "<p>Thank you for taking the time to give this advice, i certainly sometimes feel overwhelmed and want to learn everything all at once, but i won't rush it and i will try to figure everything out step by step and with practice.<br>\nMuch appreciation! 🙏</p>",
          "rawMarkdown": "Thank you for taking the time to give this advice, i certainly sometimes feel overwhelmed and want to learn everything all at once, but i won't rush it and i will try to figure everything out step by step and with practice.\nMuch appreciation! 🙏"
        }
      ]
    },
    {
      "id": 3112179,
      "postDate": "2025-02-01T03:10:23.507Z",
      "content": "<p>(1) find something that can run:</p>\n<ul>\n<li>read and process data</li>\n<li>train a model</li>\n<li>submit notebook with published results</li>\n</ul>\n<p>you can find it in public notebook, etc ….<br>\nrun it and make sure you can get same published results.</p>\n<hr>\n<p>(2) ask people or ask chatgpt anything about (1) </p>",
      "rawMarkdown": "(1) find something that can run:\n- read and process data\n- train a model\n- submit notebook with published results\n\nyou can find it in public notebook, etc ....\nrun it and make sure you can get same published results.\n\n---\n\n(2) ask people or ask chatgpt anything about (1) \n",
      "votes": 5,
      "replies": [
        {
          "id": 3112641,
          "postDate": "2025-02-01T14:30:49.513Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3110723,
      "postDate": "2025-01-30T10:25:26.100Z",
      "content": "<p>Hi, i just want to ask my fellow Kagglers how do they know what to do, where and when to do it regarding everything in this competition.</p>\n<p>I find it very interesting, fun , and wonderful to learn and participate in. <br>\nCan you tell me how can i learn about this type of analysis to be able to know what to do when faced with Data and objectives like that! I read so many genius amazing Notebooks and Discussions and studied it a lot but i still find the approaches and code hard to understand and when i open a notebook to start by myself i literally freeze because i don't know the process and approaches to such a problem!</p>\n<p>Thank you in advance for your advices 🙏♥️</p>\n<p>Note: idc about submission date and LB, i just want to learn about it even if the comp ended!</p>",
      "rawMarkdown": "Hi, i just want to ask my fellow Kagglers how do they know what to do, where and when to do it regarding everything in this competition.\n\nI find it very interesting, fun , and wonderful to learn and participate in. \nCan you tell me how can i learn about this type of analysis to be able to know what to do when faced with Data and objectives like that! I read so many genius amazing Notebooks and Discussions and studied it a lot but i still find the approaches and code hard to understand and when i open a notebook to start by myself i literally freeze because i don't know the process and approaches to such a problem!\n\nThank you in advance for your advices 🙏♥️\n\nNote: idc about submission date and LB, i just want to learn about it even if the comp ended!",
      "votes": 5
    },
    {
      "id": 3111453,
      "postDate": "2025-01-31T07:14:42.203Z",
      "content": "<p>This post has my recommendation for how to get started:  <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715</a></p>\n<p>FYI, the origin of at least some of the notebooks will become more clear if you check out the CZII site.</p>",
      "rawMarkdown": "This post has my recommendation for how to get started:  [https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715)\n\nFYI, the origin of at least some of the notebooks will become more clear if you check out the CZII site.",
      "votes": 3,
      "replies": [
        {
          "id": 3111625,
          "postDate": "2025-01-31T11:38:14.167Z",
          "content": "<p>I will follow your comment step by step, i read all your notebooks regarding this competition! also i just took a look at the CZII site and it has many interesting Datasets and things to learn, I'll try my best to make an analysis on my own then start to edit a baseline notebook to make it better before i submit if i am gonna submit! <br>\nI also am reading the GitHub repo of the competition to understand everything brick by brick…<br>\nThank you for your time and advice truly ! 🙏</p>",
          "rawMarkdown": "I will follow your comment step by step, i read all your notebooks regarding this competition! also i just took a look at the CZII site and it has many interesting Datasets and things to learn, I'll try my best to make an analysis on my own then start to edit a baseline notebook to make it better before i submit if i am gonna submit! \nI also am reading the GitHub repo of the competition to understand everything brick by brick...\nThank you for your time and advice truly ! 🙏",
          "votes": 3
        }
      ]
    },
    {
      "id": 3110867,
      "postDate": "2025-01-30T15:05:08.050Z",
      "content": "<p>I am new myself.  Just started learning about PyTorch and Tensorflow. I am looking to work with you on a team.  Either on your team or my team.  The only thing I was able to do in this competition is using copick, create dataloaders, and successfully upload a pre-trained model in Kaggle notebook.</p>",
      "rawMarkdown": "I am new myself.  Just started learning about PyTorch and Tensorflow. I am looking to work with you on a team.  Either on your team or my team.  The only thing I was able to do in this competition is using copick, create dataloaders, and successfully upload a pre-trained model in Kaggle notebook.",
      "votes": 2,
      "replies": [
        {
          "id": 3111069,
          "postDate": "2025-01-30T19:57:33.230Z",
          "content": "<p>Sure we can collaborate on anything even if this competition ended before we could, we can try something else and share knowledge !</p>",
          "rawMarkdown": "Sure we can collaborate on anything even if this competition ended before we could, we can try something else and share knowledge !",
          "votes": 1
        }
      ]
    },
    {
      "id": 3110855,
      "postDate": "2025-01-30T14:46:15.590Z",
      "content": "<p>I was in the same place as you when I first started Kaggle and I agree, it is VERY overwhelming looking at other people's code and wondering \"how did they possibly think to do to that?\".  My best piece of advice is to just start working with the dataset.  Use what you currently know to create your preprocessing pipeline, choose a model architecture, build the training loop and validation pipeline, and so on and just start from there.  Once you have an end-to-end solution that works, start asking yourself \"What could I do different in my next experiment that may improve the results\" and give that a try.  </p>\n<p>It all just takes practice.  It will feel very overwhelming at first looking at other people's solutions but I recommend ignoring the current competition code, taking a starter notebook that builds the segmentation dataset using copick, and then build a baseline solution.  Nothing fancy, just a very simple solution using something like a 2D architecture (maybe resnet50 to start) and see how it does.  Then slowly improve it from there.  The only way to get past the freeze point when you open a notebook is to just start building something even it is just a visualization.</p>\n<p>A lot of the code you are lookin at now was built on months of work and research since the competition started, so it makes sense it seems like a lot, but just know when any competition starts the initial code is very basic just to get some kind of baseline score to try and improve on.  At the start of the competition a score of 0.635 was ruling the leaderboard and now its pushing close to 0.8.</p>",
      "rawMarkdown": "I was in the same place as you when I first started Kaggle and I agree, it is VERY overwhelming looking at other people's code and wondering \"how did they possibly think to do to that?\".  My best piece of advice is to just start working with the dataset.  Use what you currently know to create your preprocessing pipeline, choose a model architecture, build the training loop and validation pipeline, and so on and just start from there.  Once you have an end-to-end solution that works, start asking yourself \"What could I do different in my next experiment that may improve the results\" and give that a try.  \n\nIt all just takes practice.  It will feel very overwhelming at first looking at other people's solutions but I recommend ignoring the current competition code, taking a starter notebook that builds the segmentation dataset using copick, and then build a baseline solution.  Nothing fancy, just a very simple solution using something like a 2D architecture (maybe resnet50 to start) and see how it does.  Then slowly improve it from there.  The only way to get past the freeze point when you open a notebook is to just start building something even it is just a visualization.\n\nA lot of the code you are lookin at now was built on months of work and research since the competition started, so it makes sense it seems like a lot, but just know when any competition starts the initial code is very basic just to get some kind of baseline score to try and improve on.  At the start of the competition a score of 0.635 was ruling the leaderboard and now its pushing close to 0.8.",
      "votes": 2,
      "replies": [
        {
          "id": 3111066,
          "postDate": "2025-01-30T19:50:44.327Z",
          "content": "<p>Thank you very much for your advice and encouragement you're literally the best!!! , can you provide me with any resources and notebooks to start in a beginner way in this competition! i am taking a look at copick github documentation but do i need something else ? <a href=\"https://www.kaggle.com/connorjd\" target=\"_blank\">@connorjd</a> </p>",
          "rawMarkdown": "Thank you very much for your advice and encouragement you're literally the best!!! , can you provide me with any resources and notebooks to start in a beginner way in this competition! i am taking a look at copick github documentation but do i need something else ? @connorjd ",
          "replies": [
            {
              "id": 3111144,
              "postDate": "2025-01-30T22:01:48.400Z",
              "content": "<p>Hello, I have started with <a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">this</a>, from <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a>.</p>",
              "rawMarkdown": "Hello, I have started with [this](https://www.kaggle.com/code/fnands/baseline-unet-train-submit), from @fnands.",
              "votes": 1
            },
            {
              "id": 3111628,
              "postDate": "2025-01-31T11:40:21.137Z",
              "content": "<p>I started with this and many similar but i posted this thread because i didn't understand almost anything they did, but now i encourage you to take a look at this <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715</a></p>",
              "rawMarkdown": "I started with this and many similar but i posted this thread because i didn't understand almost anything they did, but now i encourage you to take a look at this https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3111744,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2025-01-31T14:41:04.960000",
      "content": "<p>Hey there!</p>\n<p>First off, huge props for jumping into Kaggle—it's awesome that you’re excited to learn and grow, and that mindset will definitely help you get where you want to be. It’s perfectly normal to feel overwhelmed when you're starting out, especially when you see all the advanced notebooks and discussions. But trust me, every Kaggle pro was once in your shoes!</p>\n<p>The most important thing to keep in mind is <strong>to start simple</strong>. It’s easy to get lost in the complexity of others’ work, but you don’t have to do everything at once. A solid foundation will help you build your confidence and skills over time.</p>\n<p>Here’s how I’d recommend approaching it:</p>\n<ol>\n<li><p><strong>Start with the basics</strong>: Focus on understanding basic concepts like data cleaning, exploration, feature engineering, and model building. Try to replicate simple notebooks step by step. Don’t worry if you don’t get it right away; every little thing you do will help you understand the bigger picture.</p></li>\n<li><p><strong>Follow a simple process</strong>: Many top Kagglers follow a basic workflow:</p>\n<ul>\n<li><strong>Exploration</strong>: Load your data and explore it (basic statistics, distributions, correlations).</li>\n<li><strong>Preprocessing</strong>: Handle missing data, outliers, and any data wrangling.</li>\n<li><strong>Modeling</strong>: Start with basic models like Resnet18 before moving to more advanced ones.</li>\n<li><strong>Evaluation</strong>: Check your model’s performance using the right metrics.</li></ul></li>\n<li><p><strong>Learn one concept at a time</strong>: For example, maybe this week focus on data exploration and visualization. Next week, dive into feature engineering or basic machine learning models. Trying to learn everything at once can be a recipe for frustration.</p></li>\n<li><p><strong>Don’t stress about the code</strong>: Instead, try to focus on <strong>why</strong> certain steps are being taken. Once you grasp the “why” behind it, understanding the code will follow more naturally. Try running the code in a notebook and modify it to see how things change.</p></li>\n<li><p><br>\nThis recommendation from LLM is actually wrong. I think working on a competition with LB feedback and a lively community is 100x better than working on a small project</p></li>\n<li><p><strong>Ask for help</strong>: Kaggle is full of people who are more than willing to help. Don’t hesitate to ask for clarification when something isn’t clear. You’ll often find that others are happy to walk you through their thought process.</p></li>\n</ol>\n<p>Finally, remember that you don’t have to rush. You’re in this to learn, and the competition part will always be secondary to gaining a solid understanding of how to approach data problems. Stay curious and enjoy the process—don’t be afraid to make mistakes, because that’s where the real learning happens!</p>\n<p>Good luck, and keep going! You'll get there! ✨</p>\n<p>(written by LLM, slightly edited)</p>",
      "votes": 9,
      "replies": [
        {
          "id": 3111878,
          "author_name": "Alaa Sweed",
          "author_url": "",
          "post_date": "2025-01-31T17:51:16.293000",
          "content": "<p>Thank you for taking the time to give this advice, i certainly sometimes feel overwhelmed and want to learn everything all at once, but i won't rush it and i will try to figure everything out step by step and with practice.<br>\nMuch appreciation! 🙏</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3112179,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-02-01T03:10:23.507000",
      "content": "<p>(1) find something that can run:</p>\n<ul>\n<li>read and process data</li>\n<li>train a model</li>\n<li>submit notebook with published results</li>\n</ul>\n<p>you can find it in public notebook, etc ….<br>\nrun it and make sure you can get same published results.</p>\n<hr>\n<p>(2) ask people or ask chatgpt anything about (1) </p>",
      "votes": 5,
      "replies": [
        {
          "id": 3112641,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-02-01T14:30:49.513000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3111453,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2025-01-31T07:14:42.203000",
      "content": "<p>This post has my recommendation for how to get started:  <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715</a></p>\n<p>FYI, the origin of at least some of the notebooks will become more clear if you check out the CZII site.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3111625,
          "author_name": "Alaa Sweed",
          "author_url": "",
          "post_date": "2025-01-31T11:38:14.167000",
          "content": "<p>I will follow your comment step by step, i read all your notebooks regarding this competition! also i just took a look at the CZII site and it has many interesting Datasets and things to learn, I'll try my best to make an analysis on my own then start to edit a baseline notebook to make it better before i submit if i am gonna submit! <br>\nI also am reading the GitHub repo of the competition to understand everything brick by brick…<br>\nThank you for your time and advice truly ! 🙏</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3110867,
      "author_name": "JohnKennedyMLOps",
      "author_url": "",
      "post_date": "2025-01-30T15:05:08.050000",
      "content": "<p>I am new myself.  Just started learning about PyTorch and Tensorflow. I am looking to work with you on a team.  Either on your team or my team.  The only thing I was able to do in this competition is using copick, create dataloaders, and successfully upload a pre-trained model in Kaggle notebook.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3111069,
          "author_name": "Alaa Sweed",
          "author_url": "",
          "post_date": "2025-01-30T19:57:33.230000",
          "content": "<p>Sure we can collaborate on anything even if this competition ended before we could, we can try something else and share knowledge !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3110855,
      "author_name": "Connor",
      "author_url": "",
      "post_date": "2025-01-30T14:46:15.590000",
      "content": "<p>I was in the same place as you when I first started Kaggle and I agree, it is VERY overwhelming looking at other people's code and wondering \"how did they possibly think to do to that?\".  My best piece of advice is to just start working with the dataset.  Use what you currently know to create your preprocessing pipeline, choose a model architecture, build the training loop and validation pipeline, and so on and just start from there.  Once you have an end-to-end solution that works, start asking yourself \"What could I do different in my next experiment that may improve the results\" and give that a try.  </p>\n<p>It all just takes practice.  It will feel very overwhelming at first looking at other people's solutions but I recommend ignoring the current competition code, taking a starter notebook that builds the segmentation dataset using copick, and then build a baseline solution.  Nothing fancy, just a very simple solution using something like a 2D architecture (maybe resnet50 to start) and see how it does.  Then slowly improve it from there.  The only way to get past the freeze point when you open a notebook is to just start building something even it is just a visualization.</p>\n<p>A lot of the code you are lookin at now was built on months of work and research since the competition started, so it makes sense it seems like a lot, but just know when any competition starts the initial code is very basic just to get some kind of baseline score to try and improve on.  At the start of the competition a score of 0.635 was ruling the leaderboard and now its pushing close to 0.8.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3111066,
          "author_name": "Alaa Sweed",
          "author_url": "",
          "post_date": "2025-01-30T19:50:44.327000",
          "content": "<p>Thank you very much for your advice and encouragement you're literally the best!!! , can you provide me with any resources and notebooks to start in a beginner way in this competition! i am taking a look at copick github documentation but do i need something else ? <a href=\"https://www.kaggle.com/connorjd\" target=\"_blank\">@connorjd</a> </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3111144,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2025-01-30T22:01:48.400000",
              "content": "<p>Hello, I have started with <a href=\"https://www.kaggle.com/code/fnands/baseline-unet-train-submit\" target=\"_blank\">this</a>, from <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a>.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3111628,
              "author_name": "Alaa Sweed",
              "author_url": "",
              "post_date": "2025-01-31T11:40:21.137000",
              "content": "<p>I started with this and many similar but i posted this thread because i didn't understand almost anything they did, but now i encourage you to take a look at this <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3111744": "Hey there!\n\nFirst off, huge props for jumping into Kaggle—it's awesome that you’re excited to learn and grow, and that mindset will definitely help you get where you want to be. It’s perfectly normal to feel overwhelmed when you're starting out, especially when you see all the advanced notebooks and discussions. But trust me, every Kaggle pro was once in your shoes!\n\nThe most important thing to keep in mind is **to start simple**. It’s easy to get lost in the complexity of others’ work, but you don’t have to do everything at once. A solid foundation will help you build your confidence and skills over time.\n\nHere’s how I’d recommend approaching it:\n\n1. **Start with the basics**: Focus on understanding basic concepts like data cleaning, exploration, feature engineering, and model building. Try to replicate simple notebooks step by step. Don’t worry if you don’t get it right away; every little thing you do will help you understand the bigger picture.\n\n2. **Follow a simple process**: Many top Kagglers follow a basic workflow:\n   - **Exploration**: Load your data and explore it (basic statistics, distributions, correlations).\n   - **Preprocessing**: Handle missing data, outliers, and any data wrangling.\n   - **Modeling**: Start with basic models like Resnet18 before moving to more advanced ones.\n   - **Evaluation**: Check your model’s performance using the right metrics.\n\n3. **Learn one concept at a time**: For example, maybe this week focus on data exploration and visualization. Next week, dive into feature engineering or basic machine learning models. Trying to learn everything at once can be a recipe for frustration.\n\n4. **Don’t stress about the code**: Instead, try to focus on **why** certain steps are being taken. Once you grasp the “why” behind it, understanding the code will follow more naturally. Try running the code in a notebook and modify it to see how things change.\n\n5. ~~**Start small with your own projects**: After you get a feel for the process, try working on small datasets that interest you. You don’t need to be in a competition to practice! Gradually, you’ll feel more comfortable tackling larger, more complex problems.~~\nThis recommendation from LLM is actually wrong. I think working on a competition with LB feedback and a lively community is 100x better than working on a small project\n\n6. **Ask for help**: Kaggle is full of people who are more than willing to help. Don’t hesitate to ask for clarification when something isn’t clear. You’ll often find that others are happy to walk you through their thought process.\n\nFinally, remember that you don’t have to rush. You’re in this to learn, and the competition part will always be secondary to gaining a solid understanding of how to approach data problems. Stay curious and enjoy the process—don’t be afraid to make mistakes, because that’s where the real learning happens!\n\nGood luck, and keep going! You'll get there! ✨\n\n(written by LLM, slightly edited)",
    "3112179": "(1) find something that can run:\n- read and process data\n- train a model\n- submit notebook with published results\n\nyou can find it in public notebook, etc ....\nrun it and make sure you can get same published results.\n\n---\n\n(2) ask people or ask chatgpt anything about (1) \n",
    "3110723": "Hi, i just want to ask my fellow Kagglers how do they know what to do, where and when to do it regarding everything in this competition.\n\nI find it very interesting, fun , and wonderful to learn and participate in. \nCan you tell me how can i learn about this type of analysis to be able to know what to do when faced with Data and objectives like that! I read so many genius amazing Notebooks and Discussions and studied it a lot but i still find the approaches and code hard to understand and when i open a notebook to start by myself i literally freeze because i don't know the process and approaches to such a problem!\n\nThank you in advance for your advices 🙏♥️\n\nNote: idc about submission date and LB, i just want to learn about it even if the comp ended!",
    "3111453": "This post has my recommendation for how to get started:  [https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715)\n\nFYI, the origin of at least some of the notebooks will become more clear if you check out the CZII site.",
    "3110867": "I am new myself.  Just started learning about PyTorch and Tensorflow. I am looking to work with you on a team.  Either on your team or my team.  The only thing I was able to do in this competition is using copick, create dataloaders, and successfully upload a pre-trained model in Kaggle notebook.",
    "3110855": "I was in the same place as you when I first started Kaggle and I agree, it is VERY overwhelming looking at other people's code and wondering \"how did they possibly think to do to that?\".  My best piece of advice is to just start working with the dataset.  Use what you currently know to create your preprocessing pipeline, choose a model architecture, build the training loop and validation pipeline, and so on and just start from there.  Once you have an end-to-end solution that works, start asking yourself \"What could I do different in my next experiment that may improve the results\" and give that a try.  \n\nIt all just takes practice.  It will feel very overwhelming at first looking at other people's solutions but I recommend ignoring the current competition code, taking a starter notebook that builds the segmentation dataset using copick, and then build a baseline solution.  Nothing fancy, just a very simple solution using something like a 2D architecture (maybe resnet50 to start) and see how it does.  Then slowly improve it from there.  The only way to get past the freeze point when you open a notebook is to just start building something even it is just a visualization.\n\nA lot of the code you are lookin at now was built on months of work and research since the competition started, so it makes sense it seems like a lot, but just know when any competition starts the initial code is very basic just to get some kind of baseline score to try and improve on.  At the start of the competition a score of 0.635 was ruling the leaderboard and now its pushing close to 0.8."
  }
}