{
  "id": 104407,
  "title": "Are you new to machine learning? Ask questions here!",
  "url": "/competitions/understanding_cloud_organization/discussion/104407",
  "author_name": "inversion",
  "post_date": "2019-08-16T16:53:36.760000",
  "votes": 23,
  "comment_count": 47,
  "views": 0,
  "content": "<p>Are you new to machine learning or image segmentation problems? Feel free to ask them here. No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>",
  "messages": [
    {
      "id": 600863,
      "postDate": "2019-08-16T16:53:36.760Z",
      "content": "<p>Are you new to machine learning or image segmentation problems? Feel free to ask them here. No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!</p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>",
      "rawMarkdown": "Are you new to machine learning or image segmentation problems? Feel free to ask them here. No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!",
      "votes": 23
    },
    {
      "id": 617951,
      "postDate": "2019-09-04T16:49:12.930Z",
      "content": "<p>Hello I am confused as how to use to the encoded pixels. The usual way I would use images with a CNN, would be to take the images in the training folder, augment the data, shape it so its constant and then plug it in the CNN model for training. But here we have images and we have encoded pixels in the train csv file with the label, so I am confused as to how do you use these encoded pixels?</p>",
      "rawMarkdown": "Hello I am confused as how to use to the encoded pixels. The usual way I would use images with a CNN, would be to take the images in the training folder, augment the data, shape it so its constant and then plug it in the CNN model for training. But here we have images and we have encoded pixels in the train csv file with the label, so I am confused as to how do you use these encoded pixels?",
      "votes": 3,
      "replies": [
        {
          "id": 628271,
          "postDate": "2019-09-17T04:36:44.113Z",
          "content": "<p>You can see here, in this notebook: <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools</a>.\nCheck function rle_decode.</p>",
          "rawMarkdown": "You can see here, in this notebook: https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools.\nCheck function rle_decode.",
          "votes": 1
        }
      ]
    },
    {
      "id": 607273,
      "postDate": "2019-08-24T23:49:21.570Z",
      "content": "<p>Question №1:\nNo matter which model i use, validation loss and dice always jump like on pictures below. It's OK? What could be a reason of those jumps?\nI tried:\n- Another model;\n- Lower learning rate;\n- Learning rate with decay;\n- Different input image sizes;\n- Augmentation.</p>\n\n<p>Those techniques help me to improve my score but jumps are exist anyway.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2F91a63edff491533297c27c2654b4b16c%2Fbad_dice.png?generation=1566689355399191&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2Fc6147640717215c30d936430e7d8bcee%2Fbad_loss.png?generation=1566689394406313&amp;alt=media\" alt=\"\"></p>\n\n<p>Question №2:\nDice on validation set and test set are completely different. I have 0.331 dice on training set, 0.339 on validation set, 0.455 on leaderboard in my best submission. What could be the reason for this difference? I understand that test set may have different distribution but i don't know how model that bad on training data can be good on test data.</p>",
      "rawMarkdown": "Question №1:\nNo matter which model i use, validation loss and dice always jump like on pictures below. It's OK? What could be a reason of those jumps?\nI tried:\n- Another model;\n- Lower learning rate;\n- Learning rate with decay;\n- Different input image sizes;\n- Augmentation.\n\nThose techniques help me to improve my score but jumps are exist anyway.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2F91a63edff491533297c27c2654b4b16c%2Fbad_dice.png?generation=1566689355399191&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2Fc6147640717215c30d936430e7d8bcee%2Fbad_loss.png?generation=1566689394406313&amp;alt=media)\n\nQuestion №2:\nDice on validation set and test set are completely different. I have 0.331 dice on training set, 0.339 on validation set, 0.455 on leaderboard in my best submission. What could be the reason for this difference? I understand that test set may have different distribution but i don't know how model that bad on training data can be good on test data.\n",
      "votes": 4
    },
    {
      "id": 649637,
      "postDate": "2019-10-15T15:42:42.883Z",
      "content": "<p>I don't understand why there are some 'nan' in colume 'EncodedPixels' of training data.</p>",
      "rawMarkdown": "I don't understand why there are some 'nan' in colume 'EncodedPixels' of training data.",
      "votes": 1
    },
    {
      "id": 649634,
      "postDate": "2019-10-15T15:40:24.067Z",
      "content": "<p>I do not understand about \"EncodedPixels\". \nwhy didn't use (xmin, ymin, xmax, ymax) to decide the rectangles' position? I think this is better and use less room.</p>",
      "rawMarkdown": "I do not understand about \"EncodedPixels\". \nwhy didn't use (xmin, ymin, xmax, ymax) to decide the rectangles' position? I think this is better and use less room.",
      "votes": 1,
      "replies": [
        {
          "id": 652016,
          "postDate": "2019-10-18T08:19:55.660Z",
          "content": "<p><a href=\"/yixinchen1\">@yixinchen1</a> Yes that is possible too, But when your bounding box (ground truth of any class) has a different shape other then rectangle or square (xmin, ymin, xmax, ymax) can't work over there. So runlength encoding comes to rescue.</p>",
          "rawMarkdown": "@yixinchen1 Yes that is possible too, But when your bounding box (ground truth of any class) has a different shape other then rectangle or square (xmin, ymin, xmax, ymax) can't work over there. So runlength encoding comes to rescue.",
          "votes": 4
        },
        {
          "id": 652059,
          "postDate": "2019-10-18T09:20:31.923Z",
          "content": "<p>OK. I guess I know it. Thanks!</p>",
          "rawMarkdown": "OK. I guess I know it. Thanks!",
          "votes": 2
        }
      ]
    },
    {
      "id": 644601,
      "postDate": "2019-10-09T03:53:59Z",
      "content": "<p>Hi,\nhow we can use a optimized size of the minimal cloud size?\nif you can break a \"flower\" cloud in small pieces, there could be other small types encapsulated?\nregards,\nAdrian</p>",
      "rawMarkdown": "Hi,\nhow we can use a optimized size of the minimal cloud size?\nif you can break a \"flower\" cloud in small pieces, there could be other small types encapsulated?\nregards,\nAdrian",
      "votes": 1
    },
    {
      "id": 644435,
      "postDate": "2019-10-08T20:32:50.233Z",
      "content": "<p>Are there any ratings of what library to use?\nSo far it is understandable that it is a good way to use some gradient approach like XGBoost, LightGBM etc.\nLast trends it are in Fast AI + Keras Conv2d and CatBoost\nWhat is the best to be used in this case? Are there any \"new\" gradient library upcoming or specific for this topic upcoming on the market?</p>\n\n<p>My personal opinion: Big companies who release the \"public\" versions are using a better and improved ones for their business and using Kaggle competitions just to make assumption for improving their own.</p>\n\n<p>regards,\nAdrian</p>",
      "rawMarkdown": "Are there any ratings of what library to use?\nSo far it is understandable that it is a good way to use some gradient approach like XGBoost, LightGBM etc.\nLast trends it are in Fast AI + Keras Conv2d and CatBoost\nWhat is the best to be used in this case? Are there any \"new\" gradient library upcoming or specific for this topic upcoming on the market?\n\nMy personal opinion: Big companies who release the \"public\" versions are using a better and improved ones for their business and using Kaggle competitions just to make assumption for improving their own.\n\nregards,\nAdrian\n",
      "votes": 1
    },
    {
      "id": 621557,
      "postDate": "2019-09-08T16:50:05.547Z",
      "content": "<p>Hello, </p>\n\n<p>I've been looking into using torchvision Deep Lab v3 for semantic segmentation. Very much a novice with ML. The concept was to have it convert the rle data into mask images, (re)train the pre-trained model by comparing the train jpg's with their image masks. Was thinking one model for each cloud class to keep things relatively simple. </p>\n\n<p>My questions: \n    1. Is this a viable tool / concept for the competition? \n    2. Is there a more appropriate tool / concept I should pursue?\n    3. How does Deep Lab v3 consume both train image and mask in the first place??? [Using one cloud type mask dataset per model] I've not found any example scripts so am now wondering if this is the wrong tool. </p>\n\n<p>I'd be happy just to get something working and able to make a submission. Perhaps to place above the bottom 10 scores for once.</p>\n\n<p>Thank You, </p>\n\n<p>Steven</p>",
      "rawMarkdown": "Hello, \n\nI've been looking into using torchvision Deep Lab v3 for semantic segmentation. Very much a novice with ML. The concept was to have it convert the rle data into mask images, (re)train the pre-trained model by comparing the train jpg's with their image masks. Was thinking one model for each cloud class to keep things relatively simple. \n\nMy questions: \n    1. Is this a viable tool / concept for the competition? \n    2. Is there a more appropriate tool / concept I should pursue?\n    3. How does Deep Lab v3 consume both train image and mask in the first place??? [Using one cloud type mask dataset per model] I've not found any example scripts so am now wondering if this is the wrong tool. \n\nI'd be happy just to get something working and able to make a submission. Perhaps to place above the bottom 10 scores for once.\n\nThank You, \n\nSteven",
      "votes": 1,
      "replies": [
        {
          "id": 634035,
          "postDate": "2019-09-25T18:12:12.590Z",
          "content": "<ol>\n<li>You can use Deeplab in tensorflow, keras, or pytorch. There are more git repository about this.</li>\n</ol>",
          "rawMarkdown": "1. You can use Deeplab in tensorflow, keras, or pytorch. There are more git repository about this.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 616277,
      "postDate": "2019-09-02T23:59:01.560Z",
      "content": "<p>I read this sentence from Data Description: \" IMPORTANT: Your prediction masks should be scaled down to 350 x 525 px.\", but I don't understand what it means. Can someone give an idea for this?</p>",
      "rawMarkdown": "I read this sentence from Data Description: \" IMPORTANT: Your prediction masks should be scaled down to 350 x 525 px.\", but I don't understand what it means. Can someone give an idea for this?\n",
      "votes": 1,
      "replies": [
        {
          "id": 625939,
          "postDate": "2019-09-13T15:49:24.783Z",
          "content": "<p>Did you get an answer for this Mengyu? I think I know what it means but I am uncertain too.</p>",
          "rawMarkdown": "Did you get an answer for this Mengyu? I think I know what it means but I am uncertain too."
        },
        {
          "id": 634032,
          "postDate": "2019-09-25T18:04:19.877Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 615043,
      "postDate": "2019-09-01T11:50:59.897Z",
      "content": "<p>Can someone suggest  a good YouTube channel to learn Deep Learning  !!</p>",
      "rawMarkdown": "Can someone suggest  a good YouTube channel to learn Deep Learning  !!",
      "votes": 1,
      "replies": [
        {
          "id": 622111,
          "postDate": "2019-09-09T09:35:07.900Z",
          "content": "<p><a href=\"https://www.youtube.com/watch?v=XfoYk_Z5AkI\">https://www.youtube.com/watch?v=XfoYk_Z5AkI</a></p>\n\n<p>This is a great course from Jeremy Howard</p>",
          "rawMarkdown": "https://www.youtube.com/watch?v=XfoYk_Z5AkI\n\nThis is a great course from Jeremy Howard",
          "votes": 2
        }
      ]
    },
    {
      "id": 603988,
      "postDate": "2019-08-20T23:25:52.143Z",
      "content": "<p>I recently started learning Data Science by my own , my first step was Machine Learning ( YouTube channels , coursers kaggle and so on  ....) .My question is what is the right path to became a data scientist  I mean what is next ? I checked internet I still confused .</p>\n\n<p>PS : using R </p>",
      "rawMarkdown": "I recently started learning Data Science by my own , my first step was Machine Learning ( YouTube channels , coursers kaggle and so on  ....) .My question is what is the right path to became a data scientist  I mean what is next ? I checked internet I still confused .\n\nPS : using R \n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 604386,
          "postDate": "2019-08-21T10:33:40.100Z",
          "content": "<p>Real world tasks are really different from online courses but they are necessary so don't stop taking those. Keep finding new material to learn and don't try to understand 100% of it.</p>\n\n<p>The only way I can recommend is to do practical stuff, like build image classifier, predict something using tabular data, build a simple CNN from scratch. Kaggle can be overwhelming if you're new, so don't try to create submission when you start. Just give yourself fefew days to understand the data, use R/tableau/python to create a visual-story of data.</p>\n\n<p>Also, read other competition's kernels. If kaggle is a university then kernels are like books. You just have to know which one to pick.</p>\n\n<p>PS: Maybe try to move to Python, if you have experience in other similar languages like Java or C# or javascript</p>",
          "rawMarkdown": "Real world tasks are really different from online courses but they are necessary so don't stop taking those. Keep finding new material to learn and don't try to understand 100% of it.\n\nThe only way I can recommend is to do practical stuff, like build image classifier, predict something using tabular data, build a simple CNN from scratch. Kaggle can be overwhelming if you're new, so don't try to create submission when you start. Just give yourself fefew days to understand the data, use R/tableau/python to create a visual-story of data.\n\nAlso, read other competition's kernels. If kaggle is a university then kernels are like books. You just have to know which one to pick.\n\nPS: Maybe try to move to Python, if you have experience in other similar languages like Java or C# or javascript",
          "votes": 5
        },
        {
          "id": 604696,
          "postDate": "2019-08-21T16:56:27.970Z",
          "content": "<p><a href=\"/mukul1904\">@mukul1904</a> thanks  for your advises </p>",
          "rawMarkdown": "@mukul1904 thanks  for your advises ",
          "votes": 1
        },
        {
          "id": 625907,
          "postDate": "2019-09-13T15:17:37.127Z",
          "content": "<p>Thanks for sharing</p>",
          "rawMarkdown": "Thanks for sharing"
        }
      ]
    },
    {
      "id": 645344,
      "postDate": "2019-10-10T02:50:12.870Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> Thank you for this topic！ I have a question, that is 'The images in the test_images folder is the 25% of the test data or the 100% of the test data?'</p>",
      "rawMarkdown": "@inversion Thank you for this topic！ I have a question, that is 'The images in the test_images folder is the 25% of the test data or the 100% of the test data?'",
      "votes": 2,
      "replies": [
        {
          "id": 645793,
          "postDate": "2019-10-10T13:36:10.033Z",
          "content": "<p>It is 100% of the test data. But we only show the score on 25% until the end of the competition, then the final leaderboard rank is based on the other 75%</p>",
          "rawMarkdown": "It is 100% of the test data. But we only show the score on 25% until the end of the competition, then the final leaderboard rank is based on the other 75%",
          "votes": 1
        },
        {
          "id": 649632,
          "postDate": "2019-10-15T15:37:25.543Z",
          "content": "<p>Does that mean that the final leaderboard will change a little and we can get the other 75% by the end of test and submit it again?</p>",
          "rawMarkdown": "Does that mean that the final leaderboard will change a little and we can get the other 75% by the end of test and submit it again?",
          "votes": 1
        },
        {
          "id": 649685,
          "postDate": "2019-10-15T16:54:52.580Z",
          "content": "<p>No, not quite. Final score will be calculated automatically based on private leaderboard which is 75% of test data. Sometimes results change a little bit comparing to public leaderboard, in other cases change is huge (this is called shakeup here on kaggle).</p>\n\n<p>Any way, you won't need to submit again, you'll see you private score after competition ends.</p>",
          "rawMarkdown": "No, not quite. Final score will be calculated automatically based on private leaderboard which is 75% of test data. Sometimes results change a little bit comparing to public leaderboard, in other cases change is huge (this is called shakeup here on kaggle).\n\nAny way, you won't need to submit again, you'll see you private score after competition ends.",
          "votes": 1
        }
      ]
    },
    {
      "id": 633750,
      "postDate": "2019-09-25T11:23:58.933Z",
      "content": "<p>I'm a late arrival to this comp and a ML 'noob'.</p>\n\n<p>Why does the training data have <code>EncodedPixels</code> ? Why not have a bunch of images that basically define the categories for the training data?  In other words, if a satellite image has more than one category in it, why not discard the image (entirely) or crop it to the desired region and classify it as flowers, for example?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3027464%2F6fea485299ee5981a4d7779d1f866fdd%2Fflowers.PNG?generation=1569410626958676&amp;alt=media\" alt=\"\"></p>\n\n<p>If I understand right, the <code>EncodedPixels</code> are the region(s) that human volunteers have mapped out already. But these mapped regions are in a format that is unintelligible to me :(</p>\n\n<p>TIA to anyone who can set me straight after they have finished rolling eyes.🙄</p>",
      "rawMarkdown": "I'm a late arrival to this comp and a ML 'noob'.\n\nWhy does the training data have `EncodedPixels` ? Why not have a bunch of images that basically define the categories for the training data?  In other words, if a satellite image has more than one category in it, why not discard the image (entirely) or crop it to the desired region and classify it as flowers, for example?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3027464%2F6fea485299ee5981a4d7779d1f866fdd%2Fflowers.PNG?generation=1569410626958676&amp;alt=media)\n\nIf I understand right, the `EncodedPixels` are the region(s) that human volunteers have mapped out already. But these mapped regions are in a format that is unintelligible to me :(\n\nTIA to anyone who can set me straight after they have finished rolling eyes.🙄",
      "votes": 2,
      "replies": [
        {
          "id": 643138,
          "postDate": "2019-10-07T06:52:43.920Z",
          "content": "<p>What you're trying to say is that all image segmentation tasks ~ image classification after cropping. Which is impractical and slow at the time of deployment. You may want to look up multi-class image segmentation tasks to better understand why we are given an <code>EncodedPixels</code> column. </p>\n\n<blockquote>\n  <p>if a satellite image has more than one category in it, why not discard the image (entirely)</p>\n</blockquote>\n\n<p>Basic EDA will let you know most of the images have at least 2 types of cloud so by your suggestion a lot of valuable data will simply be lost. Also, the problem with cropping is how to know what to crop at the time of deployment? You only have an image and you want to detect types of clouds in it, no one's going to give you cropped clouds. </p>\n\n<p>I think by now, you might've figured all this out. Good luck. </p>",
          "rawMarkdown": "What you're trying to say is that all image segmentation tasks ~ image classification after cropping. Which is impractical and slow at the time of deployment. You may want to look up multi-class image segmentation tasks to better understand why we are given an `EncodedPixels` column. \n\n&gt; if a satellite image has more than one category in it, why not discard the image (entirely)\n\nBasic EDA will let you know most of the images have at least 2 types of cloud so by your suggestion a lot of valuable data will simply be lost. Also, the problem with cropping is how to know what to crop at the time of deployment? You only have an image and you want to detect types of clouds in it, no one's going to give you cropped clouds. \n\nI think by now, you might've figured all this out. Good luck. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 616791,
      "postDate": "2019-09-03T13:31:19.280Z",
      "content": "<p>How do I get validation prediction results during training using Keras? Do I need to customise a Callback?</p>",
      "rawMarkdown": "How do I get validation prediction results during training using Keras? Do I need to customise a Callback?",
      "votes": 2,
      "replies": [
        {
          "id": 627760,
          "postDate": "2019-09-16T10:36:40.927Z",
          "content": "<p>Hi Yirun Zhang!</p>\n\n<p>You can define a custom class that inherits from the <code>keras.callbacks.Callback</code> class. You just need to override two methods <code>__init__()</code> and <code>on_epoch_end</code>. A simple example would be as follows:</p>\n\n<p>`class TestCallback(Callback):\n    def <strong>init</strong>(self, test_data):\n        self.test_data = test_data</p>\n\n<pre><code>def on_epoch_end(self, epoch, logs={}):\n    x, y = self.test_data\n    loss, acc = self.model.evaluate(x, y, verbose=0)\n    print('\\nTesting loss: {}, acc: {}\\n'.format(loss, acc))`\n</code></pre>\n\n<p>After which you can pass it on to the <code>.fit()</code> method when training your model as follows:</p>\n\n<p><code>model.fit(X_train, Y_train, validation_data=(X_val, Y_val), \n          callbacks=[TestCallback((X_test, Y_test))])</code></p>\n\n<p>Regards,\nSri Yogesh.</p>",
          "rawMarkdown": "Hi Yirun Zhang!\n\nYou can define a custom class that inherits from the `keras.callbacks.Callback` class. You just need to override two methods `__init__()` and `on_epoch_end`. A simple example would be as follows:\n\n`class TestCallback(Callback):\n    def __init__(self, test_data):\n        self.test_data = test_data\n\n    def on_epoch_end(self, epoch, logs={}):\n        x, y = self.test_data\n        loss, acc = self.model.evaluate(x, y, verbose=0)\n        print('\\nTesting loss: {}, acc: {}\\n'.format(loss, acc))`\n\nAfter which you can pass it on to the `.fit()` method when training your model as follows:\n\n`model.fit(X_train, Y_train, validation_data=(X_val, Y_val), \n          callbacks=[TestCallback((X_test, Y_test))])`\n\nRegards,\nSri Yogesh.\n\n"
        }
      ]
    },
    {
      "id": 605814,
      "postDate": "2019-08-22T19:20:12.323Z",
      "content": "<p>Here's a very detailed (and much appreciated) article on how you can start your ML journey using Kaggle:\n<a href=\"https://towardsdatascience.com/use-kaggle-to-start-and-guide-your-ml-data-science-journey-f09154baba35\">https://towardsdatascience.com/use-kaggle-to-start-and-guide-your-ml-data-science-journey-f09154baba35</a></p>",
      "rawMarkdown": "Here's a very detailed (and much appreciated) article on how you can start your ML journey using Kaggle:\nhttps://towardsdatascience.com/use-kaggle-to-start-and-guide-your-ml-data-science-journey-f09154baba35",
      "votes": 2
    },
    {
      "id": 604716,
      "postDate": "2019-08-21T17:22:38.190Z",
      "content": "<p>Can someone suggest a learning path for ML and AI ?</p>",
      "rawMarkdown": "Can someone suggest a learning path for ML and AI ?",
      "votes": 2,
      "replies": [
        {
          "id": 604877,
          "postDate": "2019-08-21T20:50:17.083Z",
          "content": "<p>that's what I am looking for, this is my question purpose </p>",
          "rawMarkdown": "that's what I am looking for, this is my question purpose ",
          "votes": 1
        },
        {
          "id": 626471,
          "postDate": "2019-09-14T11:39:09.920Z",
          "content": "<p>Depends on what you want to learn. Target things that you are passionate in solving and do courses, competitions and read editorials based on them, every day. You'll get better eventually. For the long term, always learn the math hidden and learn to code out things from scratch (numpy) instead of relying on libraries and so called 'top down approaches'.</p>",
          "rawMarkdown": "Depends on what you want to learn. Target things that you are passionate in solving and do courses, competitions and read editorials based on them, every day. You'll get better eventually. For the long term, always learn the math hidden and learn to code out things from scratch (numpy) instead of relying on libraries and so called 'top down approaches'.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1779361,
      "postDate": "2022-05-06T10:36:16.600Z",
      "content": "<p>ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0, with gufunc signature (n?,k),(k,m?)-&gt;(n?,m?) (size 34 is different from 55)</p>",
      "rawMarkdown": "ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0, with gufunc signature (n?,k),(k,m?)->(n?,m?) (size 34 is different from 55)\n"
    },
    {
      "id": 1776862,
      "postDate": "2022-05-04T08:50:59.970Z",
      "content": "<p>Bonjour à tous!<br>\nJe suis nouveau dans la science des données, j'apprend depuis 4 mois sur coursera et je travaille pour un cabinet d'audit et conseil. <br>\nJ'ai un souci, je suis à la recherche des jeux de données sur la détection des cas de fraude de TVA car nous voulons améliorer le modèle. J'ai cherché dans plusieurs plateformes mais en vain y compris même kaggle.<br>\nAlors n'ayant pas encore des repères assez développés, je sollicite votre aide afin d'en obtenir.<br>\nJe vous remercie en avance de vos réponses.<br>\nCordialement.</p>",
      "rawMarkdown": "Bonjour à tous!\nJe suis nouveau dans la science des données, j'apprend depuis 4 mois sur coursera et je travaille pour un cabinet d'audit et conseil. \nJ'ai un souci, je suis à la recherche des jeux de données sur la détection des cas de fraude de TVA car nous voulons améliorer le modèle. J'ai cherché dans plusieurs plateformes mais en vain y compris même kaggle.\nAlors n'ayant pas encore des repères assez développés, je sollicite votre aide afin d'en obtenir.\nJe vous remercie en avance de vos réponses.\nCordialement."
    },
    {
      "id": 671031,
      "postDate": "2019-11-12T07:16:58.227Z",
      "content": "<p>How do I generate masked images from pixel values stored in csv file?</p>",
      "rawMarkdown": "How do I generate masked images from pixel values stored in csv file?\n"
    },
    {
      "id": 670223,
      "postDate": "2019-11-11T07:35:04.303Z",
      "content": "<p>how to approach this problem??</p>",
      "rawMarkdown": "how to approach this problem??"
    },
    {
      "id": 668069,
      "postDate": "2019-11-08T00:04:03.733Z",
      "content": "<p>Do you guys know of any good resources(Kernels, discussions, blogs, etc..) for ensembling segmentation models?</p>",
      "rawMarkdown": "Do you guys know of any good resources(Kernels, discussions, blogs, etc..) for ensembling segmentation models?"
    },
    {
      "id": 660541,
      "postDate": "2019-10-29T09:42:56.133Z",
      "content": "<p>In submission file, do we have to predict for all cloud types for every image in test folder?</p>",
      "rawMarkdown": "In submission file, do we have to predict for all cloud types for every image in test folder?"
    },
    {
      "id": 653726,
      "postDate": "2019-10-20T21:56:00.500Z",
      "content": "<p>I am new to machine competition and started to learn EDA but i am confused is there anything like EDA in image classification .So please help me how to start image classification competition </p>",
      "rawMarkdown": "I am new to machine competition and started to learn EDA but i am confused is there anything like EDA in image classification .So please help me how to start image classification competition "
    },
    {
      "id": 650658,
      "postDate": "2019-10-16T15:20:29.743Z",
      "content": "<p>Hi,\nAfter reading this forum, I studied the Unet architecture to get prepared for the competition, and I wonder if I can ask you a technical explanation. </p>\n\n<ul>\n<li>I understand that in Unet the prediction is based on a pixel by pixel. The networs sees the true value of each pixel during training from the label mask provided</li>\n<li>a cross-entropy loss function is calculated using the prediction (output of the sigmoid function) and the true value of the pixel (given by the label mask)\nMy <strong>question</strong>:  in case of an image 100 x 100 pixels I can imagine there will be around 10000 final activations. What error will be backpropagated?  I guess 10000 individual errors for each image.  In other words a pixel py pixel classification is like predicting 10000 classes?  (or the task is easier than predicting 10000 classes since the true classes are only 1 and 0)</li>\n<li>(I understand there is an extra element added to the loss function IoU)</li>\n</ul>\n\n<p>any help, links or suggestion would be really appreciated\nRegards</p>",
      "rawMarkdown": "Hi,\nAfter reading this forum, I studied the Unet architecture to get prepared for the competition, and I wonder if I can ask you a technical explanation. \n\n- I understand that in Unet the prediction is based on a pixel by pixel. The networs sees the true value of each pixel during training from the label mask provided\n- a cross-entropy loss function is calculated using the prediction (output of the sigmoid function) and the true value of the pixel (given by the label mask)\nMy **question**:  in case of an image 100 x 100 pixels I can imagine there will be around 10000 final activations. What error will be backpropagated?  I guess 10000 individual errors for each image.  In other words a pixel py pixel classification is like predicting 10000 classes?  (or the task is easier than predicting 10000 classes since the true classes are only 1 and 0)\n- (I understand there is an extra element added to the loss function IoU)\n\nany help, links or suggestion would be really appreciated\nRegards"
    },
    {
      "id": 614200,
      "postDate": "2019-08-31T07:45:40.807Z",
      "content": "<p>Can someone suggest a starting point for building a good ML algorithm on matching jobs with service provider.</p>",
      "rawMarkdown": "Can someone suggest a starting point for building a good ML algorithm on matching jobs with service provider.",
      "replies": [
        {
          "id": 634037,
          "postDate": "2019-09-25T18:12:57.003Z",
          "content": "<p>You should learn about segmentation problems</p>",
          "rawMarkdown": "You should learn about segmentation problems",
          "votes": 1
        },
        {
          "id": 645281,
          "postDate": "2019-10-10T00:47:01.117Z",
          "content": "<p>If you look for a similar problem in kaggle and refer to notebooks written by others, it might help.\nsee this <a href=\"https://cloud.google.com/solutions/talent-solution/\">https://cloud.google.com/solutions/talent-solution/</a></p>",
          "rawMarkdown": "If you look for a similar problem in kaggle and refer to notebooks written by others, it might help.\nsee this https://cloud.google.com/solutions/talent-solution/"
        }
      ]
    },
    {
      "id": 624115,
      "postDate": "2019-09-11T17:25:27.177Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true,
      "replies": [
        {
          "id": 626014,
          "postDate": "2019-09-13T17:31:35.067Z",
          "content": "<p>Your question is really fine in this thread. When you are classifying fruit or animal or digit in MNIST dataset you are doing classification of whole image. You may do some preprocessing first to cut some parts but this is not really important at first, this is to enhance your results later (more advanced).</p>\n\n<p>This competition is different - you need to find your objects (they are parts of the whole image and you don't know where they are) and classify them also. So key words here are segmentation, U-Net (algorithm to do that). And also I believe you can find helpful to do this first manually to feel how you can find clouds and classify them yourself. Try it on zooniverse:</p>\n\n<p><a href=\"https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel\">https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel</a></p>",
          "rawMarkdown": "Your question is really fine in this thread. When you are classifying fruit or animal or digit in MNIST dataset you are doing classification of whole image. You may do some preprocessing first to cut some parts but this is not really important at first, this is to enhance your results later (more advanced).\n\nThis competition is different - you need to find your objects (they are parts of the whole image and you don't know where they are) and classify them also. So key words here are segmentation, U-Net (algorithm to do that). And also I believe you can find helpful to do this first manually to feel how you can find clouds and classify them yourself. Try it on zooniverse:\n\nhttps://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel",
          "votes": 3
        },
        {
          "id": 632913,
          "postDate": "2019-09-24T08:02:18.980Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 615045,
      "postDate": "2019-09-01T11:56:59.620Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 617951,
      "author_name": "Adi Mithani",
      "author_url": "",
      "post_date": "2019-09-04T16:49:12.930000",
      "content": "<p>Hello I am confused as how to use to the encoded pixels. The usual way I would use images with a CNN, would be to take the images in the training folder, augment the data, shape it so its constant and then plug it in the CNN model for training. But here we have images and we have encoded pixels in the train csv file with the label, so I am confused as to how do you use these encoded pixels?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 628271,
          "author_name": "L3viEvil",
          "author_url": "",
          "post_date": "2019-09-17T04:36:44.113000",
          "content": "<p>You can see here, in this notebook: <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools</a>.\nCheck function rle_decode.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 607273,
      "author_name": "Dmitry Larionov",
      "author_url": "",
      "post_date": "2019-08-24T23:49:21.570000",
      "content": "<p>Question №1:\nNo matter which model i use, validation loss and dice always jump like on pictures below. It's OK? What could be a reason of those jumps?\nI tried:\n- Another model;\n- Lower learning rate;\n- Learning rate with decay;\n- Different input image sizes;\n- Augmentation.</p>\n\n<p>Those techniques help me to improve my score but jumps are exist anyway.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2F91a63edff491533297c27c2654b4b16c%2Fbad_dice.png?generation=1566689355399191&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2Fc6147640717215c30d936430e7d8bcee%2Fbad_loss.png?generation=1566689394406313&amp;alt=media\" alt=\"\"></p>\n\n<p>Question №2:\nDice on validation set and test set are completely different. I have 0.331 dice on training set, 0.339 on validation set, 0.455 on leaderboard in my best submission. What could be the reason for this difference? I understand that test set may have different distribution but i don't know how model that bad on training data can be good on test data.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 649637,
      "author_name": "Yixinchen",
      "author_url": "",
      "post_date": "2019-10-15T15:42:42.883000",
      "content": "<p>I don't understand why there are some 'nan' in colume 'EncodedPixels' of training data.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 649634,
      "author_name": "Yixinchen",
      "author_url": "",
      "post_date": "2019-10-15T15:40:24.067000",
      "content": "<p>I do not understand about \"EncodedPixels\". \nwhy didn't use (xmin, ymin, xmax, ymax) to decide the rectangles' position? I think this is better and use less room.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 652016,
          "author_name": "Yogesh Mahawar",
          "author_url": "",
          "post_date": "2019-10-18T08:19:55.660000",
          "content": "<p><a href=\"/yixinchen1\">@yixinchen1</a> Yes that is possible too, But when your bounding box (ground truth of any class) has a different shape other then rectangle or square (xmin, ymin, xmax, ymax) can't work over there. So runlength encoding comes to rescue.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 652059,
          "author_name": "Yixinchen",
          "author_url": "",
          "post_date": "2019-10-18T09:20:31.923000",
          "content": "<p>OK. I guess I know it. Thanks!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 644601,
      "author_name": "Adrian Zinovei",
      "author_url": "",
      "post_date": "2019-10-09T03:53:59",
      "content": "<p>Hi,\nhow we can use a optimized size of the minimal cloud size?\nif you can break a \"flower\" cloud in small pieces, there could be other small types encapsulated?\nregards,\nAdrian</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 644435,
      "author_name": "Adrian Zinovei",
      "author_url": "",
      "post_date": "2019-10-08T20:32:50.233000",
      "content": "<p>Are there any ratings of what library to use?\nSo far it is understandable that it is a good way to use some gradient approach like XGBoost, LightGBM etc.\nLast trends it are in Fast AI + Keras Conv2d and CatBoost\nWhat is the best to be used in this case? Are there any \"new\" gradient library upcoming or specific for this topic upcoming on the market?</p>\n\n<p>My personal opinion: Big companies who release the \"public\" versions are using a better and improved ones for their business and using Kaggle competitions just to make assumption for improving their own.</p>\n\n<p>regards,\nAdrian</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621557,
      "author_name": "I am Kiera",
      "author_url": "",
      "post_date": "2019-09-08T16:50:05.547000",
      "content": "<p>Hello, </p>\n\n<p>I've been looking into using torchvision Deep Lab v3 for semantic segmentation. Very much a novice with ML. The concept was to have it convert the rle data into mask images, (re)train the pre-trained model by comparing the train jpg's with their image masks. Was thinking one model for each cloud class to keep things relatively simple. </p>\n\n<p>My questions: \n    1. Is this a viable tool / concept for the competition? \n    2. Is there a more appropriate tool / concept I should pursue?\n    3. How does Deep Lab v3 consume both train image and mask in the first place??? [Using one cloud type mask dataset per model] I've not found any example scripts so am now wondering if this is the wrong tool. </p>\n\n<p>I'd be happy just to get something working and able to make a submission. Perhaps to place above the bottom 10 scores for once.</p>\n\n<p>Thank You, </p>\n\n<p>Steven</p>",
      "votes": 1,
      "replies": [
        {
          "id": 634035,
          "author_name": "L3viEvil",
          "author_url": "",
          "post_date": "2019-09-25T18:12:12.590000",
          "content": "<ol>\n<li>You can use Deeplab in tensorflow, keras, or pytorch. There are more git repository about this.</li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 616277,
      "author_name": "Mengyu",
      "author_url": "",
      "post_date": "2019-09-02T23:59:01.560000",
      "content": "<p>I read this sentence from Data Description: \" IMPORTANT: Your prediction masks should be scaled down to 350 x 525 px.\", but I don't understand what it means. Can someone give an idea for this?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 625939,
          "author_name": "Ashleigh Trinh",
          "author_url": "",
          "post_date": "2019-09-13T15:49:24.783000",
          "content": "<p>Did you get an answer for this Mengyu? I think I know what it means but I am uncertain too.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 634032,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-25T18:04:19.877000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 615043,
      "author_name": "Ibrahim Chaoudi ",
      "author_url": "",
      "post_date": "2019-09-01T11:50:59.897000",
      "content": "<p>Can someone suggest  a good YouTube channel to learn Deep Learning  !!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 622111,
          "author_name": "Anna Novikova",
          "author_url": "",
          "post_date": "2019-09-09T09:35:07.900000",
          "content": "<p><a href=\"https://www.youtube.com/watch?v=XfoYk_Z5AkI\">https://www.youtube.com/watch?v=XfoYk_Z5AkI</a></p>\n\n<p>This is a great course from Jeremy Howard</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 603988,
      "author_name": "Ibrahim Chaoudi ",
      "author_url": "",
      "post_date": "2019-08-20T23:25:52.143000",
      "content": "<p>I recently started learning Data Science by my own , my first step was Machine Learning ( YouTube channels , coursers kaggle and so on  ....) .My question is what is the right path to became a data scientist  I mean what is next ? I checked internet I still confused .</p>\n\n<p>PS : using R </p>",
      "votes": 1,
      "replies": [
        {
          "id": 604386,
          "author_name": "Mukul",
          "author_url": "",
          "post_date": "2019-08-21T10:33:40.100000",
          "content": "<p>Real world tasks are really different from online courses but they are necessary so don't stop taking those. Keep finding new material to learn and don't try to understand 100% of it.</p>\n\n<p>The only way I can recommend is to do practical stuff, like build image classifier, predict something using tabular data, build a simple CNN from scratch. Kaggle can be overwhelming if you're new, so don't try to create submission when you start. Just give yourself fefew days to understand the data, use R/tableau/python to create a visual-story of data.</p>\n\n<p>Also, read other competition's kernels. If kaggle is a university then kernels are like books. You just have to know which one to pick.</p>\n\n<p>PS: Maybe try to move to Python, if you have experience in other similar languages like Java or C# or javascript</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 604696,
          "author_name": "Ibrahim Chaoudi ",
          "author_url": "",
          "post_date": "2019-08-21T16:56:27.970000",
          "content": "<p><a href=\"/mukul1904\">@mukul1904</a> thanks  for your advises </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 625907,
          "author_name": "Azeez Idris",
          "author_url": "",
          "post_date": "2019-09-13T15:17:37.127000",
          "content": "<p>Thanks for sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 645344,
      "author_name": "lifengnan",
      "author_url": "",
      "post_date": "2019-10-10T02:50:12.870000",
      "content": "<p><a href=\"/inversion\">@inversion</a> Thank you for this topic！ I have a question, that is 'The images in the test_images folder is the 25% of the test data or the 100% of the test data?'</p>",
      "votes": 2,
      "replies": [
        {
          "id": 645793,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2019-10-10T13:36:10.033000",
          "content": "<p>It is 100% of the test data. But we only show the score on 25% until the end of the competition, then the final leaderboard rank is based on the other 75%</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 649632,
          "author_name": "Yixinchen",
          "author_url": "",
          "post_date": "2019-10-15T15:37:25.543000",
          "content": "<p>Does that mean that the final leaderboard will change a little and we can get the other 75% by the end of test and submit it again?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 649685,
          "author_name": "Anna Novikova",
          "author_url": "",
          "post_date": "2019-10-15T16:54:52.580000",
          "content": "<p>No, not quite. Final score will be calculated automatically based on private leaderboard which is 75% of test data. Sometimes results change a little bit comparing to public leaderboard, in other cases change is huge (this is called shakeup here on kaggle).</p>\n\n<p>Any way, you won't need to submit again, you'll see you private score after competition ends.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 633750,
      "author_name": "MattV",
      "author_url": "",
      "post_date": "2019-09-25T11:23:58.933000",
      "content": "<p>I'm a late arrival to this comp and a ML 'noob'.</p>\n\n<p>Why does the training data have <code>EncodedPixels</code> ? Why not have a bunch of images that basically define the categories for the training data?  In other words, if a satellite image has more than one category in it, why not discard the image (entirely) or crop it to the desired region and classify it as flowers, for example?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3027464%2F6fea485299ee5981a4d7779d1f866fdd%2Fflowers.PNG?generation=1569410626958676&amp;alt=media\" alt=\"\"></p>\n\n<p>If I understand right, the <code>EncodedPixels</code> are the region(s) that human volunteers have mapped out already. But these mapped regions are in a format that is unintelligible to me :(</p>\n\n<p>TIA to anyone who can set me straight after they have finished rolling eyes.🙄</p>",
      "votes": 2,
      "replies": [
        {
          "id": 643138,
          "author_name": "timetraveller",
          "author_url": "",
          "post_date": "2019-10-07T06:52:43.920000",
          "content": "<p>What you're trying to say is that all image segmentation tasks ~ image classification after cropping. Which is impractical and slow at the time of deployment. You may want to look up multi-class image segmentation tasks to better understand why we are given an <code>EncodedPixels</code> column. </p>\n\n<blockquote>\n  <p>if a satellite image has more than one category in it, why not discard the image (entirely)</p>\n</blockquote>\n\n<p>Basic EDA will let you know most of the images have at least 2 types of cloud so by your suggestion a lot of valuable data will simply be lost. Also, the problem with cropping is how to know what to crop at the time of deployment? You only have an image and you want to detect types of clouds in it, no one's going to give you cropped clouds. </p>\n\n<p>I think by now, you might've figured all this out. Good luck. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 616791,
      "author_name": "Yirun Zhang",
      "author_url": "",
      "post_date": "2019-09-03T13:31:19.280000",
      "content": "<p>How do I get validation prediction results during training using Keras? Do I need to customise a Callback?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 627760,
          "author_name": "Sri Yogesh",
          "author_url": "",
          "post_date": "2019-09-16T10:36:40.927000",
          "content": "<p>Hi Yirun Zhang!</p>\n\n<p>You can define a custom class that inherits from the <code>keras.callbacks.Callback</code> class. You just need to override two methods <code>__init__()</code> and <code>on_epoch_end</code>. A simple example would be as follows:</p>\n\n<p>`class TestCallback(Callback):\n    def <strong>init</strong>(self, test_data):\n        self.test_data = test_data</p>\n\n<pre><code>def on_epoch_end(self, epoch, logs={}):\n    x, y = self.test_data\n    loss, acc = self.model.evaluate(x, y, verbose=0)\n    print('\\nTesting loss: {}, acc: {}\\n'.format(loss, acc))`\n</code></pre>\n\n<p>After which you can pass it on to the <code>.fit()</code> method when training your model as follows:</p>\n\n<p><code>model.fit(X_train, Y_train, validation_data=(X_val, Y_val), \n          callbacks=[TestCallback((X_test, Y_test))])</code></p>\n\n<p>Regards,\nSri Yogesh.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 605814,
      "author_name": "Nityesh Agarwal",
      "author_url": "",
      "post_date": "2019-08-22T19:20:12.323000",
      "content": "<p>Here's a very detailed (and much appreciated) article on how you can start your ML journey using Kaggle:\n<a href=\"https://towardsdatascience.com/use-kaggle-to-start-and-guide-your-ml-data-science-journey-f09154baba35\">https://towardsdatascience.com/use-kaggle-to-start-and-guide-your-ml-data-science-journey-f09154baba35</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 604716,
      "author_name": "AbhishAnk Tiwari - Aspiring Data Scientist",
      "author_url": "",
      "post_date": "2019-08-21T17:22:38.190000",
      "content": "<p>Can someone suggest a learning path for ML and AI ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 604877,
          "author_name": "Ibrahim Chaoudi ",
          "author_url": "",
          "post_date": "2019-08-21T20:50:17.083000",
          "content": "<p>that's what I am looking for, this is my question purpose </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 626471,
          "author_name": "Sayantan Das",
          "author_url": "",
          "post_date": "2019-09-14T11:39:09.920000",
          "content": "<p>Depends on what you want to learn. Target things that you are passionate in solving and do courses, competitions and read editorials based on them, every day. You'll get better eventually. For the long term, always learn the math hidden and learn to code out things from scratch (numpy) instead of relying on libraries and so called 'top down approaches'.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1779361,
      "author_name": "ADITHI RAO",
      "author_url": "",
      "post_date": "2022-05-06T10:36:16.600000",
      "content": "<p>ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0, with gufunc signature (n?,k),(k,m?)-&gt;(n?,m?) (size 34 is different from 55)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1776862,
      "author_name": "BINIAKOUNOU ",
      "author_url": "",
      "post_date": "2022-05-04T08:50:59.970000",
      "content": "<p>Bonjour à tous!<br>\nJe suis nouveau dans la science des données, j'apprend depuis 4 mois sur coursera et je travaille pour un cabinet d'audit et conseil. <br>\nJ'ai un souci, je suis à la recherche des jeux de données sur la détection des cas de fraude de TVA car nous voulons améliorer le modèle. J'ai cherché dans plusieurs plateformes mais en vain y compris même kaggle.<br>\nAlors n'ayant pas encore des repères assez développés, je sollicite votre aide afin d'en obtenir.<br>\nJe vous remercie en avance de vos réponses.<br>\nCordialement.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 671031,
      "author_name": "Altair029",
      "author_url": "",
      "post_date": "2019-11-12T07:16:58.227000",
      "content": "<p>How do I generate masked images from pixel values stored in csv file?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 670223,
      "author_name": "Taksh Kamlesh",
      "author_url": "",
      "post_date": "2019-11-11T07:35:04.303000",
      "content": "<p>how to approach this problem??</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 668069,
      "author_name": "Nischal sanil",
      "author_url": "",
      "post_date": "2019-11-08T00:04:03.733000",
      "content": "<p>Do you guys know of any good resources(Kernels, discussions, blogs, etc..) for ensembling segmentation models?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 660541,
      "author_name": "Taksh Kamlesh",
      "author_url": "",
      "post_date": "2019-10-29T09:42:56.133000",
      "content": "<p>In submission file, do we have to predict for all cloud types for every image in test folder?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 653726,
      "author_name": "Hare Krishna Singh",
      "author_url": "",
      "post_date": "2019-10-20T21:56:00.500000",
      "content": "<p>I am new to machine competition and started to learn EDA but i am confused is there anything like EDA in image classification .So please help me how to start image classification competition </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 650658,
      "author_name": "Guido Salimbeni",
      "author_url": "",
      "post_date": "2019-10-16T15:20:29.743000",
      "content": "<p>Hi,\nAfter reading this forum, I studied the Unet architecture to get prepared for the competition, and I wonder if I can ask you a technical explanation. </p>\n\n<ul>\n<li>I understand that in Unet the prediction is based on a pixel by pixel. The networs sees the true value of each pixel during training from the label mask provided</li>\n<li>a cross-entropy loss function is calculated using the prediction (output of the sigmoid function) and the true value of the pixel (given by the label mask)\nMy <strong>question</strong>:  in case of an image 100 x 100 pixels I can imagine there will be around 10000 final activations. What error will be backpropagated?  I guess 10000 individual errors for each image.  In other words a pixel py pixel classification is like predicting 10000 classes?  (or the task is easier than predicting 10000 classes since the true classes are only 1 and 0)</li>\n<li>(I understand there is an extra element added to the loss function IoU)</li>\n</ul>\n\n<p>any help, links or suggestion would be really appreciated\nRegards</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 614200,
      "author_name": "Sandeep Kumar  ",
      "author_url": "",
      "post_date": "2019-08-31T07:45:40.807000",
      "content": "<p>Can someone suggest a starting point for building a good ML algorithm on matching jobs with service provider.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 634037,
          "author_name": "L3viEvil",
          "author_url": "",
          "post_date": "2019-09-25T18:12:57.003000",
          "content": "<p>You should learn about segmentation problems</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 645281,
          "author_name": "BaekDong CHA",
          "author_url": "",
          "post_date": "2019-10-10T00:47:01.117000",
          "content": "<p>If you look for a similar problem in kaggle and refer to notebooks written by others, it might help.\nsee this <a href=\"https://cloud.google.com/solutions/talent-solution/\">https://cloud.google.com/solutions/talent-solution/</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 624115,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-11T17:25:27.177000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 626014,
          "author_name": "Anna Novikova",
          "author_url": "",
          "post_date": "2019-09-13T17:31:35.067000",
          "content": "<p>Your question is really fine in this thread. When you are classifying fruit or animal or digit in MNIST dataset you are doing classification of whole image. You may do some preprocessing first to cut some parts but this is not really important at first, this is to enhance your results later (more advanced).</p>\n\n<p>This competition is different - you need to find your objects (they are parts of the whole image and you don't know where they are) and classify them also. So key words here are segmentation, U-Net (algorithm to do that). And also I believe you can find helpful to do this first manually to feel how you can find clouds and classify them yourself. Try it on zooniverse:</p>\n\n<p><a href=\"https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel\">https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 632913,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-24T08:02:18.980000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 615045,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-01T11:56:59.620000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "600863": "Are you new to machine learning or image segmentation problems? Feel free to ask them here. No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with!\n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!",
    "617951": "Hello I am confused as how to use to the encoded pixels. The usual way I would use images with a CNN, would be to take the images in the training folder, augment the data, shape it so its constant and then plug it in the CNN model for training. But here we have images and we have encoded pixels in the train csv file with the label, so I am confused as to how do you use these encoded pixels?",
    "607273": "Question №1:\nNo matter which model i use, validation loss and dice always jump like on pictures below. It's OK? What could be a reason of those jumps?\nI tried:\n- Another model;\n- Lower learning rate;\n- Learning rate with decay;\n- Different input image sizes;\n- Augmentation.\n\nThose techniques help me to improve my score but jumps are exist anyway.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2F91a63edff491533297c27c2654b4b16c%2Fbad_dice.png?generation=1566689355399191&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3046684%2Fc6147640717215c30d936430e7d8bcee%2Fbad_loss.png?generation=1566689394406313&amp;alt=media)\n\nQuestion №2:\nDice on validation set and test set are completely different. I have 0.331 dice on training set, 0.339 on validation set, 0.455 on leaderboard in my best submission. What could be the reason for this difference? I understand that test set may have different distribution but i don't know how model that bad on training data can be good on test data.\n",
    "649637": "I don't understand why there are some 'nan' in colume 'EncodedPixels' of training data.",
    "649634": "I do not understand about \"EncodedPixels\". \nwhy didn't use (xmin, ymin, xmax, ymax) to decide the rectangles' position? I think this is better and use less room.",
    "644601": "Hi,\nhow we can use a optimized size of the minimal cloud size?\nif you can break a \"flower\" cloud in small pieces, there could be other small types encapsulated?\nregards,\nAdrian",
    "644435": "Are there any ratings of what library to use?\nSo far it is understandable that it is a good way to use some gradient approach like XGBoost, LightGBM etc.\nLast trends it are in Fast AI + Keras Conv2d and CatBoost\nWhat is the best to be used in this case? Are there any \"new\" gradient library upcoming or specific for this topic upcoming on the market?\n\nMy personal opinion: Big companies who release the \"public\" versions are using a better and improved ones for their business and using Kaggle competitions just to make assumption for improving their own.\n\nregards,\nAdrian\n",
    "621557": "Hello, \n\nI've been looking into using torchvision Deep Lab v3 for semantic segmentation. Very much a novice with ML. The concept was to have it convert the rle data into mask images, (re)train the pre-trained model by comparing the train jpg's with their image masks. Was thinking one model for each cloud class to keep things relatively simple. \n\nMy questions: \n    1. Is this a viable tool / concept for the competition? \n    2. Is there a more appropriate tool / concept I should pursue?\n    3. How does Deep Lab v3 consume both train image and mask in the first place??? [Using one cloud type mask dataset per model] I've not found any example scripts so am now wondering if this is the wrong tool. \n\nI'd be happy just to get something working and able to make a submission. Perhaps to place above the bottom 10 scores for once.\n\nThank You, \n\nSteven",
    "616277": "I read this sentence from Data Description: \" IMPORTANT: Your prediction masks should be scaled down to 350 x 525 px.\", but I don't understand what it means. Can someone give an idea for this?\n",
    "615043": "Can someone suggest  a good YouTube channel to learn Deep Learning  !!",
    "603988": "I recently started learning Data Science by my own , my first step was Machine Learning ( YouTube channels , coursers kaggle and so on  ....) .My question is what is the right path to became a data scientist  I mean what is next ? I checked internet I still confused .\n\nPS : using R \n\n\n",
    "645344": "@inversion Thank you for this topic！ I have a question, that is 'The images in the test_images folder is the 25% of the test data or the 100% of the test data?'",
    "633750": "I'm a late arrival to this comp and a ML 'noob'.\n\nWhy does the training data have `EncodedPixels` ? Why not have a bunch of images that basically define the categories for the training data?  In other words, if a satellite image has more than one category in it, why not discard the image (entirely) or crop it to the desired region and classify it as flowers, for example?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3027464%2F6fea485299ee5981a4d7779d1f866fdd%2Fflowers.PNG?generation=1569410626958676&amp;alt=media)\n\nIf I understand right, the `EncodedPixels` are the region(s) that human volunteers have mapped out already. But these mapped regions are in a format that is unintelligible to me :(\n\nTIA to anyone who can set me straight after they have finished rolling eyes.🙄",
    "616791": "How do I get validation prediction results during training using Keras? Do I need to customise a Callback?",
    "605814": "Here's a very detailed (and much appreciated) article on how you can start your ML journey using Kaggle:\nhttps://towardsdatascience.com/use-kaggle-to-start-and-guide-your-ml-data-science-journey-f09154baba35",
    "604716": "Can someone suggest a learning path for ML and AI ?",
    "1779361": "ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0, with gufunc signature (n?,k),(k,m?)->(n?,m?) (size 34 is different from 55)\n",
    "1776862": "Bonjour à tous!\nJe suis nouveau dans la science des données, j'apprend depuis 4 mois sur coursera et je travaille pour un cabinet d'audit et conseil. \nJ'ai un souci, je suis à la recherche des jeux de données sur la détection des cas de fraude de TVA car nous voulons améliorer le modèle. J'ai cherché dans plusieurs plateformes mais en vain y compris même kaggle.\nAlors n'ayant pas encore des repères assez développés, je sollicite votre aide afin d'en obtenir.\nJe vous remercie en avance de vos réponses.\nCordialement.",
    "671031": "How do I generate masked images from pixel values stored in csv file?\n",
    "670223": "how to approach this problem??",
    "668069": "Do you guys know of any good resources(Kernels, discussions, blogs, etc..) for ensembling segmentation models?",
    "660541": "In submission file, do we have to predict for all cloud types for every image in test folder?",
    "653726": "I am new to machine competition and started to learn EDA but i am confused is there anything like EDA in image classification .So please help me how to start image classification competition ",
    "650658": "Hi,\nAfter reading this forum, I studied the Unet architecture to get prepared for the competition, and I wonder if I can ask you a technical explanation. \n\n- I understand that in Unet the prediction is based on a pixel by pixel. The networs sees the true value of each pixel during training from the label mask provided\n- a cross-entropy loss function is calculated using the prediction (output of the sigmoid function) and the true value of the pixel (given by the label mask)\nMy **question**:  in case of an image 100 x 100 pixels I can imagine there will be around 10000 final activations. What error will be backpropagated?  I guess 10000 individual errors for each image.  In other words a pixel py pixel classification is like predicting 10000 classes?  (or the task is easier than predicting 10000 classes since the true classes are only 1 and 0)\n- (I understand there is an extra element added to the loss function IoU)\n\nany help, links or suggestion would be really appreciated\nRegards",
    "614200": "Can someone suggest a starting point for building a good ML algorithm on matching jobs with service provider.",
    "624115": "",
    "615045": ""
  }
}