{
  "id": 520125,
  "title": "A Rookie Seeking StreetSmart Tips for EDA, Cross-Validation",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/520125",
  "author_name": "",
  "post_date": "2024-07-14T15:29:10.100803Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I’m diving into the RSNA 2024 Lumbar Spine Degenerative Classification challenge as a way to cut my teeth on machine learning.<br>\nEven very confused, I am loving it, and that a huuuuge upside for my soul ❤️</p>\n<p>The complexity is making my head spin. I need some wisdom from the veterans here.</p>\n<p>Here’s what’s perplexing me:</p>\n<ul>\n<li><strong>Cross-validation</strong>: I know it’s key for assessing model performance. But where does it slot into the workflow? Before or after exploratory data analysis (EDA)?</li>\n<li><strong>EDA for medical imaging</strong>: I can handle basic EDA, but DICOM files and 3D MRI scans are throwing me off. How should I approach these beasts?</li>\n<li><strong>StratifiedKFold</strong>: This competition involves multi-label classification. What’s the right way to implement StratifiedKFold for this scenario?</li>\n<li><strong>Feature engineering</strong>: With image data, what features should I engineer? And when should I do this—before or after cross-validation?</li>\n<li><strong>GroupKFold</strong>: I’ve seen mentions of GroupKFold. Given we have multiple images per patient, is this method relevant here?</li>\n<li><strong>Workflow order</strong>: I need a clear sequence for this competition. How do I fit EDA, preprocessing, cross-validation, and model building together optimally?</li>\n</ul>\n<p>I’m pumped for this challenge and the learning it promises. Just want to ensure I’m on solid ground. Any pointers, resources, or advice would be gold.</p>\n<p>Thank you all, and best of luck in the competition</p>",
  "messages": [
    {
      "id": "2921703",
      "postDate": "07/14/2024 15:29:10",
      "content": "<p>I’m diving into the RSNA 2024 Lumbar Spine Degenerative Classification challenge as a way to cut my teeth on machine learning.<br>\nEven very confused, I am loving it, and that a huuuuge upside for my soul ❤️</p>\n<p>The complexity is making my head spin. I need some wisdom from the veterans here.</p>\n<p>Here’s what’s perplexing me:</p>\n<ul>\n<li><strong>Cross-validation</strong>: I know it’s key for assessing model performance. But where does it slot into the workflow? Before or after exploratory data analysis (EDA)?</li>\n<li><strong>EDA for medical imaging</strong>: I can handle basic EDA, but DICOM files and 3D MRI scans are throwing me off. How should I approach these beasts?</li>\n<li><strong>StratifiedKFold</strong>: This competition involves multi-label classification. What’s the right way to implement StratifiedKFold for this scenario?</li>\n<li><strong>Feature engineering</strong>: With image data, what features should I engineer? And when should I do this—before or after cross-validation?</li>\n<li><strong>GroupKFold</strong>: I’ve seen mentions of GroupKFold. Given we have multiple images per patient, is this method relevant here?</li>\n<li><strong>Workflow order</strong>: I need a clear sequence for this competition. How do I fit EDA, preprocessing, cross-validation, and model building together optimally?</li>\n</ul>\n<p>I’m pumped for this challenge and the learning it promises. Just want to ensure I’m on solid ground. Any pointers, resources, or advice would be gold.</p>\n<p>Thank you all, and best of luck in the competition</p>",
      "rawMarkdown": "I’m diving into the RSNA 2024 Lumbar Spine Degenerative Classification challenge as a way to cut my teeth on machine learning.\nEven very confused, I am loving it, and that a huuuuge upside for my soul ❤️\n\nThe complexity is making my head spin. I need some wisdom from the veterans here.\n\nHere’s what’s perplexing me:\n\n- **Cross-validation**: I know it’s key for assessing model performance. But where does it slot into the workflow? Before or after exploratory data analysis (EDA)?\n- **EDA for medical imaging**: I can handle basic EDA, but DICOM files and 3D MRI scans are throwing me off. How should I approach these beasts?\n- **StratifiedKFold**: This competition involves multi-label classification. What’s the right way to implement StratifiedKFold for this scenario?\n- **Feature engineering**: With image data, what features should I engineer? And when should I do this—before or after cross-validation?\n- **GroupKFold**: I’ve seen mentions of GroupKFold. Given we have multiple images per patient, is this method relevant here?\n- **Workflow order**: I need a clear sequence for this competition. How do I fit EDA, preprocessing, cross-validation, and model building together optimally?\n\nI’m pumped for this challenge and the learning it promises. Just want to ensure I’m on solid ground. Any pointers, resources, or advice would be gold.\n\nThank you all, and best of luck in the competition",
      "votes": null
    },
    {
      "id": "2921802",
      "postDate": "07/14/2024 16:56:26",
      "content": "<p>Btw, i am listening to don’t let me down  🤣🤣<br>\nCrashin', hit a wall<br>\nRight now, I need a miracle Hurry up now, I need a miracle<br>\nStranded, reachin' out I call your name, but you're not around<br>\nI say your name, but you're not around<br>\nI need ya, I need ya, I need you right now 🤣</p>",
      "rawMarkdown": "Btw, i am listening to don’t let me down  🤣🤣\nCrashin', hit a wall\nRight now, I need a miracle Hurry up now, I need a miracle\nStranded, reachin' out I call your name, but you're not around\nI say your name, but you're not around\nI need ya, I need ya, I need you right now 🤣",
      "votes": null
    },
    {
      "id": "2921934",
      "postDate": "07/14/2024 18:43:37",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/tinkerturbo\" target=\"_blank\">@tinkerturbo</a>, I recommend studying 5 notebooks that look interesting and identifying patterns. This will help build a strategy for you. </p>",
      "rawMarkdown": "Hello @tinkerturbo, I recommend studying 5 notebooks that look interesting and identifying patterns. This will help build a strategy for you.",
      "votes": null
    },
    {
      "id": "2922005",
      "postDate": "07/14/2024 19:43:06",
      "content": "<p>Thx Carlos, l'll do that.</p>\n<p>I am currently watching </p>\n<p><a href=\"https://www.youtube.com/watch?v=kI4yxAL2wtM&amp;ab_channel=AbhishekThakur\" target=\"_blank\">https://www.youtube.com/watch?v=kI4yxAL2wtM&amp;ab_channel=AbhishekThakur</a></p>\n<p>Talks # 14: Martin Henze; Knowledge is Power: Understanding your Data through EDA and Visualisations </p>",
      "rawMarkdown": "Thx Carlos, l'll do that.\n\nI am currently watching \n\nhttps://www.youtube.com/watch?v=kI4yxAL2wtM&ab_channel=AbhishekThakur\n\nTalks # 14: Martin Henze; Knowledge is Power: Understanding your Data through EDA and Visualisations",
      "votes": null
    },
    {
      "id": "2922126",
      "postDate": "07/14/2024 22:33:05",
      "content": "<p>If you are just starting don't think about all this. Focus on one main goal: build a working pipeline and submit to leaderboard asap. EDA, cross validation, etc don't matter till you have a pipeline that takes the input and produces the output prediction. It can be a very simple pipeline but it needs to be the first priority. Once you make a submission make another post and ask community for more help on next steps.</p>",
      "rawMarkdown": "If you are just starting don't think about all this. Focus on one main goal: build a working pipeline and submit to leaderboard asap. EDA, cross validation, etc don't matter till you have a pipeline that takes the input and produces the output prediction. It can be a very simple pipeline but it needs to be the first priority. Once you make a submission make another post and ask community for more help on next steps.",
      "votes": null
    },
    {
      "id": "2922146",
      "postDate": "07/14/2024 23:22:02",
      "content": "<p>This sounds practical, love it. <br>\nI'll get that pipeline up no matter the quality and get on the leaderboard. <br>\nEDA and cross-validation can wait. Over time, I'll learn and improve. Let's get it done!<br>\nThx sigint</p>",
      "rawMarkdown": "This sounds practical, love it. \nI'll get that pipeline up no matter the quality and get on the leaderboard. \nEDA and cross-validation can wait. Over time, I'll learn and improve. Let's get it done!\nThx sigint",
      "votes": null
    },
    {
      "id": "2922147",
      "postDate": "07/14/2024 23:25:54",
      "content": "<p><a href=\"https://www.kaggle.com/snehalverma10\" target=\"_blank\">@snehalverma10</a> <br>\nI just noticed that I can tag in Kaggle 😆<br>\nThx again </p>",
      "rawMarkdown": "snehalverma10 \nI just noticed that I can tag in Kaggle 😆\nThx again",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2921802,
      "author_name": "tinkerturbo",
      "author_url": "",
      "post_date": "07/14/2024 16:56:26",
      "content": "<p>Btw, i am listening to don’t let me down  🤣🤣<br>\nCrashin', hit a wall<br>\nRight now, I need a miracle Hurry up now, I need a miracle<br>\nStranded, reachin' out I call your name, but you're not around<br>\nI say your name, but you're not around<br>\nI need ya, I need ya, I need you right now 🤣</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2921934,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "07/14/2024 18:43:37",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/tinkerturbo\" target=\"_blank\">@tinkerturbo</a>, I recommend studying 5 notebooks that look interesting and identifying patterns. This will help build a strategy for you. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2922005,
          "author_name": "tinkerturbo",
          "author_url": "",
          "post_date": "07/14/2024 19:43:06",
          "content": "<p>Thx Carlos, l'll do that.</p>\n<p>I am currently watching </p>\n<p><a href=\"https://www.youtube.com/watch?v=kI4yxAL2wtM&amp;ab_channel=AbhishekThakur\" target=\"_blank\">https://www.youtube.com/watch?v=kI4yxAL2wtM&amp;ab_channel=AbhishekThakur</a></p>\n<p>Talks # 14: Martin Henze; Knowledge is Power: Understanding your Data through EDA and Visualisations </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2922126,
      "author_name": "snehalverma10",
      "author_url": "",
      "post_date": "07/14/2024 22:33:05",
      "content": "<p>If you are just starting don't think about all this. Focus on one main goal: build a working pipeline and submit to leaderboard asap. EDA, cross validation, etc don't matter till you have a pipeline that takes the input and produces the output prediction. It can be a very simple pipeline but it needs to be the first priority. Once you make a submission make another post and ask community for more help on next steps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2922146,
          "author_name": "tinkerturbo",
          "author_url": "",
          "post_date": "07/14/2024 23:22:02",
          "content": "<p>This sounds practical, love it. <br>\nI'll get that pipeline up no matter the quality and get on the leaderboard. <br>\nEDA and cross-validation can wait. Over time, I'll learn and improve. Let's get it done!<br>\nThx sigint</p>",
          "votes": null,
          "replies": [
            {
              "id": 2922147,
              "author_name": "tinkerturbo",
              "author_url": "",
              "post_date": "07/14/2024 23:25:54",
              "content": "<p><a href=\"https://www.kaggle.com/snehalverma10\" target=\"_blank\">@snehalverma10</a> <br>\nI just noticed that I can tag in Kaggle 😆<br>\nThx again </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2921703": "I’m diving into the RSNA 2024 Lumbar Spine Degenerative Classification challenge as a way to cut my teeth on machine learning.\nEven very confused, I am loving it, and that a huuuuge upside for my soul ❤️\n\nThe complexity is making my head spin. I need some wisdom from the veterans here.\n\nHere’s what’s perplexing me:\n\n- **Cross-validation**: I know it’s key for assessing model performance. But where does it slot into the workflow? Before or after exploratory data analysis (EDA)?\n- **EDA for medical imaging**: I can handle basic EDA, but DICOM files and 3D MRI scans are throwing me off. How should I approach these beasts?\n- **StratifiedKFold**: This competition involves multi-label classification. What’s the right way to implement StratifiedKFold for this scenario?\n- **Feature engineering**: With image data, what features should I engineer? And when should I do this—before or after cross-validation?\n- **GroupKFold**: I’ve seen mentions of GroupKFold. Given we have multiple images per patient, is this method relevant here?\n- **Workflow order**: I need a clear sequence for this competition. How do I fit EDA, preprocessing, cross-validation, and model building together optimally?\n\nI’m pumped for this challenge and the learning it promises. Just want to ensure I’m on solid ground. Any pointers, resources, or advice would be gold.\n\nThank you all, and best of luck in the competition",
    "2921802": "Btw, i am listening to don’t let me down  🤣🤣\nCrashin', hit a wall\nRight now, I need a miracle Hurry up now, I need a miracle\nStranded, reachin' out I call your name, but you're not around\nI say your name, but you're not around\nI need ya, I need ya, I need you right now 🤣",
    "2921934": "Hello @tinkerturbo, I recommend studying 5 notebooks that look interesting and identifying patterns. This will help build a strategy for you.",
    "2922005": "Thx Carlos, l'll do that.\n\nI am currently watching \n\nhttps://www.youtube.com/watch?v=kI4yxAL2wtM&ab_channel=AbhishekThakur\n\nTalks # 14: Martin Henze; Knowledge is Power: Understanding your Data through EDA and Visualisations",
    "2922126": "If you are just starting don't think about all this. Focus on one main goal: build a working pipeline and submit to leaderboard asap. EDA, cross validation, etc don't matter till you have a pipeline that takes the input and produces the output prediction. It can be a very simple pipeline but it needs to be the first priority. Once you make a submission make another post and ask community for more help on next steps.",
    "2922146": "This sounds practical, love it. \nI'll get that pipeline up no matter the quality and get on the leaderboard. \nEDA and cross-validation can wait. Over time, I'll learn and improve. Let's get it done!\nThx sigint",
    "2922147": "snehalverma10 \nI just noticed that I can tag in Kaggle 😆\nThx again"
  },
  "source": "meta"
}