{
  "id": 55299,
  "title": "Welcome!",
  "url": "/competitions/avito-demand-prediction/discussion/55299",
  "author_name": "Sohier Dane",
  "post_date": "2018-04-24T22:08:39.492000",
  "votes": 28,
  "comment_count": 57,
  "views": 0,
  "content": "<p>Welcome to the Avito Demand Prediction Challenge, where you're tasked with estimating the success of online ads. \nYou'll need to integrate information from tabular, text, and image data. \nThe images are stored in zip archives that are too large to fit in memory of a kernel if loaded all at once, so you may find this <a href=\"https://www.kaggle.com/sohier/getting-started-loading-the-images\">starter kernel</a> about loading single files from the archives useful.\n(Note, due to the size of the files, it takes a few minutes for data to load up in Kernels.)</p>\n\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": 318961,
      "postDate": "2018-04-24T22:08:39.493Z",
      "content": "<p>Welcome to the Avito Demand Prediction Challenge, where you're tasked with estimating the success of online ads. \nYou'll need to integrate information from tabular, text, and image data. \nThe images are stored in zip archives that are too large to fit in memory of a kernel if loaded all at once, so you may find this <a href=\"https://www.kaggle.com/sohier/getting-started-loading-the-images\">starter kernel</a> about loading single files from the archives useful.\n(Note, due to the size of the files, it takes a few minutes for data to load up in Kernels.)</p>\n\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Welcome to the Avito Demand Prediction Challenge, where you're tasked with estimating the success of online ads. \nYou'll need to integrate information from tabular, text, and image data. \nThe images are stored in zip archives that are too large to fit in memory of a kernel if loaded all at once, so you may find this [starter kernel](https://www.kaggle.com/sohier/getting-started-loading-the-images) about loading single files from the archives useful.\n(Note, due to the size of the files, it takes a few minutes for data to load up in Kernels.)\n\nHappy Kaggling!",
      "votes": 28
    },
    {
      "id": 319000,
      "postDate": "2018-04-25T02:07:04.930Z",
      "content": "<p>Is there any way to get a description of how the dependent variable is created?  I'm interested in learning in more detail what we are actually predicting.</p>",
      "rawMarkdown": "Is there any way to get a description of how the dependent variable is created?  I'm interested in learning in more detail what we are actually predicting.",
      "votes": 22,
      "replies": [
        {
          "id": 322409,
          "postDate": "2018-05-02T22:08:25.557Z",
          "content": "<p>Me too, please let me know if you find anything!</p>",
          "rawMarkdown": "Me too, please let me know if you find anything!"
        }
      ]
    },
    {
      "id": 319063,
      "postDate": "2018-04-25T06:04:54.450Z",
      "content": "<p>Seems like a cool competition! Using RMSE for probabilities however is a weird choice I haven't seen before.</p>",
      "rawMarkdown": "Seems like a cool competition! Using RMSE for probabilities however is a weird choice I haven't seen before.",
      "votes": 19,
      "replies": [
        {
          "id": 319087,
          "postDate": "2018-04-25T07:37:52.743Z",
          "content": "<p>Agreed! That's the first thing I noticed when looking at the competition. It would be interesting to know what is the rationale behind that choice because right now it just seems bizarre to me.</p>",
          "rawMarkdown": "Agreed! That's the first thing I noticed when looking at the competition. It would be interesting to know what is the rationale behind that choice because right now it just seems bizarre to me.",
          "votes": 2
        },
        {
          "id": 320088,
          "postDate": "2018-04-27T12:50:09.233Z",
          "content": "<p>As for the RMSE. Maybe this condition served to choose the evaluation of the competition.\n'Since the errors are squared before they are averaged, the RMSE gives a relatively high weight to large errors. This means the RMSE should be more useful when large errors are particularly undesirable.'\nAlthough maybe I'm wrong.</p>",
          "rawMarkdown": "As for the RMSE. Maybe this condition served to choose the evaluation of the competition.\n'Since the errors are squared before they are averaged, the RMSE gives a relatively high weight to large errors. This means the RMSE should be more useful when large errors are particularly undesirable.'\nAlthough maybe I'm wrong.\n"
        },
        {
          "id": 320118,
          "postDate": "2018-04-27T14:48:03.773Z",
          "content": "<p>@Evgeniy The same is true for logloss, a very common metric for classification competitions</p>",
          "rawMarkdown": "@Evgeniy The same is true for logloss, a very common metric for classification competitions"
        },
        {
          "id": 320400,
          "postDate": "2018-04-28T15:22:01.310Z",
          "content": "<p>Logloss won't work if the ground truth isn't 0/1 </p>",
          "rawMarkdown": "Logloss won't work if the ground truth isn't 0/1 ",
          "votes": 1
        },
        {
          "id": 321078,
          "postDate": "2018-04-30T14:34:04.573Z",
          "content": "<p>What about logloss doesn't work in this situation if the ground truth isn't 0/1? You can naively apply the same formula as the 0-1 ground truth case case, as long as the response is in [0,1]. </p>\n\n<p>For example if there is lots of data with few unique covariate combinations, it's very common to fit a binomial GLM (whose loss function is logloss) on the aggregated data with ground truth equal to percentage of observed successes in a cell and weight equal to number of observations in the cell. It's equivalent to \"unrolling\" the data to have 0-1 response and the same weight per row.</p>",
          "rawMarkdown": "What about logloss doesn't work in this situation if the ground truth isn't 0/1? You can naively apply the same formula as the 0-1 ground truth case case, as long as the response is in [0,1]. \n\nFor example if there is lots of data with few unique covariate combinations, it's very common to fit a binomial GLM (whose loss function is logloss) on the aggregated data with ground truth equal to percentage of observed successes in a cell and weight equal to number of observations in the cell. It's equivalent to \"unrolling\" the data to have 0-1 response and the same weight per row.",
          "votes": -1
        },
        {
          "id": 321214,
          "postDate": "2018-04-30T20:59:54.093Z",
          "content": "<p>Binary log loss is a special case of <a href=\"https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence\">Kullback-Leibrer divergence</a> where p(x) is assumed to be either 0 or 1. In this case, the entropy of the p(x) is exactly zero and the loss is just a negative log of the correct class probability. Transition to a non-binary case is fairly straightforward - we just need to treat p(x) as a true *deal_probability*.</p>",
          "rawMarkdown": "Binary log loss is a special case of [Kullback-Leibrer divergence][1] where p(x) is assumed to be either 0 or 1. In this case, the entropy of the p(x) is exactly zero and the loss is just a negative log of the correct class probability. Transition to a non-binary case is fairly straightforward - we just need to treat p(x) as a true *deal_probability*.\n\n  [1]: https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence",
          "votes": 6
        },
        {
          "id": 321277,
          "postDate": "2018-04-30T23:03:54.260Z",
          "content": "<p>@Dmitriy Danevskiy thanks for the suggestion. Unfortunately we currently don't support that in our metrics library, we will keep it in mind for the future. </p>",
          "rawMarkdown": "@Dmitriy Danevskiy thanks for the suggestion. Unfortunately we currently don't support that in our metrics library, we will keep it in mind for the future. "
        }
      ]
    },
    {
      "id": 319903,
      "postDate": "2018-04-27T04:40:31.753Z",
      "content": "<p>Just curious where does the deal_probability on train data set come from, why is not it a 0/1 label?</p>",
      "rawMarkdown": "Just curious where does the deal_probability on train data set come from, why is not it a 0/1 label?",
      "votes": 11,
      "replies": [
        {
          "id": 321199,
          "postDate": "2018-04-30T20:03:16.120Z",
          "content": "<p>May be the best way to look at deal_probability is to read it as ad_effectiveness. Closer it is to 1, more effective it is. </p>",
          "rawMarkdown": "May be the best way to look at deal_probability is to read it as ad_effectiveness. Closer it is to 1, more effective it is. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 319610,
      "postDate": "2018-04-26T12:40:34.523Z",
      "content": "<p>Hi, can you put some light on what method/strategy was used to estimate  these  deal probability.</p>",
      "rawMarkdown": "Hi, can you put some light on what method/strategy was used to estimate  these  deal probability.",
      "votes": 7,
      "replies": [
        {
          "id": 319811,
          "postDate": "2018-04-26T21:49:36.407Z",
          "content": "<p>Would definitely appreciate that as well.</p>",
          "rawMarkdown": "Would definitely appreciate that as well.",
          "votes": 3
        }
      ]
    },
    {
      "id": 323826,
      "postDate": "2018-05-06T10:41:38.570Z",
      "content": "<p>What are the columns params_1, params_2, params_3 signify and also what does image_top_1 represent?</p>",
      "rawMarkdown": "What are the columns params_1, params_2, params_3 signify and also what does image_top_1 represent?",
      "votes": 3,
      "replies": [
        {
          "id": 323854,
          "postDate": "2018-05-06T12:18:25.827Z",
          "content": "<p>Would definitely appreciate some help here, too, with how to interpret those columns.</p>",
          "rawMarkdown": "Would definitely appreciate some help here, too, with how to interpret those columns."
        },
        {
          "id": 328535,
          "postDate": "2018-05-14T15:08:57.350Z",
          "content": "<p>Looking at the data, params_1,(2)(3) are just different levels of details about the product. For example in one of the rows: parent_category_name - personal things, category_name - clothes, shoes, accessories, param_1 - women clothes, param_2 - jeans, param_3 - 26 (size). In another example, they are selling a car and in params say that it is used, the model of the car, etc. \nNot sure about top_image_1 though. Might be some code assigned to each image at Avito.  </p>",
          "rawMarkdown": "Looking at the data, params_1,(2)(3) are just different levels of details about the product. For example in one of the rows: parent_category_name - personal things, category_name - clothes, shoes, accessories, param_1 - women clothes, param_2 - jeans, param_3 - 26 (size). In another example, they are selling a car and in params say that it is used, the model of the car, etc. \nNot sure about top_image_1 though. Might be some code assigned to each image at Avito. \t",
          "votes": 3
        },
        {
          "id": 333883,
          "postDate": "2018-05-26T00:55:33.847Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 333920,
          "postDate": "2018-05-26T03:40:37.923Z",
          "content": "<p>@zanhanock you can see possible explanation here\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57388\">https://www.kaggle.com/c/avito-demand-prediction/discussion/57388</a></p>",
          "rawMarkdown": "@zanhanock you can see possible explanation here\nhttps://www.kaggle.com/c/avito-demand-prediction/discussion/57388"
        },
        {
          "id": 333922,
          "postDate": "2018-05-26T03:42:58.357Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 319576,
      "postDate": "2018-04-26T10:39:57.517Z",
      "content": "<p>Is there any possibility of data sharing through Torrent, I have downloaded data through kaggle-API but the zip file is corrupted. [my <code>test_jpg.zip</code> filr is corrupted.]\nIs anyone facing similar issue ?? \nSomeone please share sha-hash to check.\nhow to start download from middle of a file using kaggle-API, or verify it.</p>\n\n<p>need help</p>",
      "rawMarkdown": "Is there any possibility of data sharing through Torrent, I have downloaded data through kaggle-API but the zip file is corrupted. [my `test_jpg.zip` filr is corrupted.]\nIs anyone facing similar issue ?? \nSomeone please share sha-hash to check.\nhow to start download from middle of a file using kaggle-API, or verify it.\n\nneed help",
      "votes": 1,
      "replies": [
        {
          "id": 319821,
          "postDate": "2018-04-26T22:35:01.770Z",
          "content": "<p>For those having trouble downloading files, here are SHA-256 checksum of files:\n<code>\n$ sha256sum *\n218950a51c6e87a7f3f4cfd7f4cccc459185d9ee5791598f35d5947b7368ec35  periods_test.csv.zip\ncda15d926826094364fca8decd90ad9cb786effdfa418dc3c258a7dc3c9af98a  periods_train.csv.zip\nf87c970b58e57dfc7b34fcd92ab0ecee44f69ccd76d35af105e0bb687495d6b2  sample_submission.csv\n63342b3f6828bd73092e345aacddc0810b0e43b4f1c558c0fb679315ce717ad4  test_active.csv.zip\nbe5123ad3a767c6d3a0467d170a70630ce3fbdeb5dbf4972f9e8960de02f3c23  test.csv.zip\n372cb16f10580e3b54eac7336079973bcea234b7ebffcaf323cae70aad0c3397  test_jpg.zip\n8ea9ac44a96e08af1016a618c11399db762f2ad53bc08ffc4bf781e15427d9d8  train_active.csv.zip\nbeeeeb4bf51df3a97ebfa5318f02d3a658efb1f0ac6e2ac312b9565b8496fcf6  train.csv.zip\naa0a8cc02c8e59020b943fae7aeedaaecd4f25d2460e150cac882ebf52f43519  train_jpg.zip\n</code></p>\n\n<p>I hope it helps.</p>",
          "rawMarkdown": "For those having trouble downloading files, here are SHA-256 checksum of files:\n```\n$ sha256sum *\n218950a51c6e87a7f3f4cfd7f4cccc459185d9ee5791598f35d5947b7368ec35  periods_test.csv.zip\ncda15d926826094364fca8decd90ad9cb786effdfa418dc3c258a7dc3c9af98a  periods_train.csv.zip\nf87c970b58e57dfc7b34fcd92ab0ecee44f69ccd76d35af105e0bb687495d6b2  sample_submission.csv\n63342b3f6828bd73092e345aacddc0810b0e43b4f1c558c0fb679315ce717ad4  test_active.csv.zip\nbe5123ad3a767c6d3a0467d170a70630ce3fbdeb5dbf4972f9e8960de02f3c23  test.csv.zip\n372cb16f10580e3b54eac7336079973bcea234b7ebffcaf323cae70aad0c3397  test_jpg.zip\n8ea9ac44a96e08af1016a618c11399db762f2ad53bc08ffc4bf781e15427d9d8  train_active.csv.zip\nbeeeeb4bf51df3a97ebfa5318f02d3a658efb1f0ac6e2ac312b9565b8496fcf6  train.csv.zip\naa0a8cc02c8e59020b943fae7aeedaaecd4f25d2460e150cac882ebf52f43519  train_jpg.zip\n```\n\nI hope it helps.",
          "votes": 1
        }
      ]
    },
    {
      "id": 319404,
      "postDate": "2018-04-26T01:27:31.767Z",
      "content": "<p>Seems like a great task to work on!</p>",
      "rawMarkdown": "Seems like a great task to work on!",
      "votes": 1
    },
    {
      "id": 319497,
      "postDate": "2018-04-26T07:23:23.710Z",
      "content": "<p>The rules mention open source software. What about commercial software? I am thinking tools like AWS Rekognition, or other tool to interpret the image files.</p>",
      "rawMarkdown": "The rules mention open source software. What about commercial software? I am thinking tools like AWS Rekognition, or other tool to interpret the image files.",
      "votes": 2
    },
    {
      "id": 350907,
      "postDate": "2018-06-30T18:08:31.130Z",
      "content": "<p>Dear Sohier Dane, (to competition organizer)</p>\n\n<p>I am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.</p>\n\n<p>Main content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.</p>\n\n<p>I was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives are possible content candidates.</p>\n\n<p>If using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.</p>\n\n<p>Please let me know, Best regards, Kweonwoo</p>",
      "rawMarkdown": "Dear Sohier Dane, (to competition organizer)\n\nI am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.\n\nMain content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.\n\nI was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives are possible content candidates.\n\nIf using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.\n\nPlease let me know, Best regards, Kweonwoo",
      "replies": [
        {
          "id": 351247,
          "postDate": "2018-07-01T17:47:09.143Z",
          "content": "<p>Kweonwoo,</p>\n\n<p>That sounds like a great project! All public Kaggle kernels are released under the Apache 2.0 open source license. As for the data, this is covered under sections 7.A and 7.B of the competition rules.</p>\n\n<p>Best,</p>\n\n<p>Sohier</p>",
          "rawMarkdown": "Kweonwoo,\n\nThat sounds like a great project! All public Kaggle kernels are released under the Apache 2.0 open source license. As for the data, this is covered under sections 7.A and 7.B of the competition rules.\n\nBest,\n\nSohier",
          "votes": 1
        }
      ]
    },
    {
      "id": 350510,
      "postDate": "2018-06-29T21:42:01.563Z",
      "content": "<p>Hi @Sohier, I see the medals are released just before, I ranked 95/1917 but got bronze, and the leaderboard shows I'm still in silver. I just get confused, could you explain this? Thank you!</p>",
      "rawMarkdown": "Hi @Sohier, I see the medals are released just before, I ranked 95/1917 but got bronze, and the leaderboard shows I'm still in silver. I just get confused, could you explain this? Thank you!",
      "replies": [
        {
          "id": 350516,
          "postDate": "2018-06-29T21:57:23.097Z",
          "content": "<p>Sorry about the confusion, <a href=\"/bangda\">@bangda</a>. It can take some time for the final medals to propagate through our systems once the leaderboard has been finalized. Your profile should properly show silver by early next week; please message me again if that doesn't happen. This is definitely an area we could improve upon so thank you for your patience.</p>",
          "rawMarkdown": "Sorry about the confusion, @bangda. It can take some time for the final medals to propagate through our systems once the leaderboard has been finalized. Your profile should properly show silver by early next week; please message me again if that doesn't happen. This is definitely an area we could improve upon so thank you for your patience.",
          "votes": 1
        },
        {
          "id": 350567,
          "postDate": "2018-06-30T01:01:44.780Z",
          "content": "<p>@Sohier， aha, similar issue here, my rank 191/1917 seems got no medal, but the leaderboard shows the bronze, the last one of bronze :-)</p>",
          "rawMarkdown": "@Sohier， aha, similar issue here, my rank 191/1917 seems got no medal, but the leaderboard shows the bronze, the last one of bronze :-)"
        },
        {
          "id": 353375,
          "postDate": "2018-07-06T15:19:48.740Z",
          "content": "<p>Thank you for your patience! I'm going to touch base with our engineering team, the correct medals should have come through by now. </p>",
          "rawMarkdown": "Thank you for your patience! I'm going to touch base with our engineering team, the correct medals should have come through by now. ",
          "votes": 1
        },
        {
          "id": 353846,
          "postDate": "2018-07-08T02:26:32.827Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 354943,
          "postDate": "2018-07-10T14:29:55.100Z",
          "content": "<p>Hi Sohier, the medal is still not changed...Could you double check with engineering team? Thank you!</p>",
          "rawMarkdown": "Hi Sohier, the medal is still not changed...Could you double check with engineering team? Thank you!"
        },
        {
          "id": 354950,
          "postDate": "2018-07-10T14:46:07.527Z",
          "content": "<p>I filed a bug, but don't have an ETA on the solution yet.</p>",
          "rawMarkdown": "I filed a bug, but don't have an ETA on the solution yet.",
          "votes": 1
        },
        {
          "id": 354951,
          "postDate": "2018-07-10T14:49:21.373Z",
          "content": "<p>Cool, great to know, hope it could be fixed soon;)</p>",
          "rawMarkdown": "Cool, great to know, hope it could be fixed soon;)"
        },
        {
          "id": 361046,
          "postDate": "2018-07-23T18:38:48.460Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 346307,
      "postDate": "2018-06-21T13:02:52.043Z",
      "content": "<p>Hi Sohier,\nis it fine if I blend the output of my model with the output of another person model (made public in his kernel) and submit the result?\nThanks in advance!!</p>",
      "rawMarkdown": "Hi Sohier,\nis it fine if I blend the output of my model with the output of another person model (made public in his kernel) and submit the result?\nThanks in advance!!",
      "replies": [
        {
          "id": 346374,
          "postDate": "2018-06-21T15:29:54.473Z",
          "content": "<p>Yes, working off of public kernels is fine.</p>",
          "rawMarkdown": "Yes, working off of public kernels is fine."
        },
        {
          "id": 346375,
          "postDate": "2018-06-21T15:31:05.550Z",
          "content": "<p>Thanks!!</p>",
          "rawMarkdown": "Thanks!!"
        }
      ]
    },
    {
      "id": 340325,
      "postDate": "2018-06-08T22:48:29.673Z",
      "content": "<p>Let's play!</p>",
      "rawMarkdown": "Let's play!"
    },
    {
      "id": 326658,
      "postDate": "2018-05-10T05:34:54.250Z",
      "content": "<p>Interesting challenge! :)</p>",
      "rawMarkdown": "Interesting challenge! :)",
      "replies": [
        {
          "id": 327025,
          "postDate": "2018-05-10T17:39:10.683Z",
          "content": "<p>I see</p>",
          "rawMarkdown": "I see"
        }
      ]
    },
    {
      "id": 319760,
      "postDate": "2018-04-26T19:30:43.613Z",
      "content": "<p>Hi, my excel couldn't recognize Russian. Can anyone help me with that?</p>",
      "rawMarkdown": "Hi, my excel couldn't recognize Russian. Can anyone help me with that?",
      "replies": [
        {
          "id": 319813,
          "postDate": "2018-04-26T21:51:33.623Z",
          "content": "<p>For the title and description, you can try something like TextBlob to convert it to English. For categorical features, you can encode them.</p>",
          "rawMarkdown": "For the title and description, you can try something like TextBlob to convert it to English. For categorical features, you can encode them.",
          "votes": 3
        }
      ]
    },
    {
      "id": 319472,
      "postDate": "2018-04-26T06:13:53.553Z",
      "content": "<p>Thx</p>",
      "rawMarkdown": "Thx"
    },
    {
      "id": 319454,
      "postDate": "2018-04-26T04:13:31.787Z",
      "content": "<p>TeSt</p>",
      "rawMarkdown": "TeSt"
    },
    {
      "id": 319391,
      "postDate": "2018-04-26T00:37:43.150Z",
      "content": "<p>+1</p>",
      "rawMarkdown": "+1"
    },
    {
      "id": 319364,
      "postDate": "2018-04-25T22:03:14.543Z",
      "content": "<p>Any one who is Interested to form a team of 2 ?</p>",
      "rawMarkdown": "Any one who is Interested to form a team of 2 ?",
      "replies": [
        {
          "id": 319829,
          "postDate": "2018-04-26T23:07:44.163Z",
          "content": "<p>I am interested</p>",
          "rawMarkdown": "I am interested"
        },
        {
          "id": 323997,
          "postDate": "2018-05-06T21:54:01.643Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 345393,
      "postDate": "2018-06-19T20:00:58.480Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 319261,
      "postDate": "2018-04-25T16:10:08.133Z",
      "rawMarkdown": "",
      "votes": 9,
      "isDeleted": true,
      "replies": [
        {
          "id": 319812,
          "postDate": "2018-04-26T21:50:18.383Z",
          "content": "<p>Seconded.</p>",
          "rawMarkdown": "Seconded."
        }
      ]
    },
    {
      "id": 350904,
      "postDate": "2018-06-30T18:03:04.300Z",
      "content": "<p>Thank you for this competition!</p>",
      "rawMarkdown": "Thank you for this competition!"
    },
    {
      "id": 340315,
      "postDate": "2018-06-08T21:57:58.410Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 320486,
      "postDate": "2018-04-28T20:07:55.877Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 319673,
      "postDate": "2018-04-26T15:16:45.730Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 319351,
      "postDate": "2018-04-25T20:59:00.070Z",
      "content": "<p>Amazing challenge thank you </p>",
      "rawMarkdown": "Amazing challenge thank you "
    }
  ],
  "comments": [
    {
      "id": 319000,
      "author_name": "RyanCaldwell",
      "author_url": "",
      "post_date": "2018-04-25T02:07:04.930000",
      "content": "<p>Is there any way to get a description of how the dependent variable is created?  I'm interested in learning in more detail what we are actually predicting.</p>",
      "votes": 22,
      "replies": [
        {
          "id": 322409,
          "author_name": "Cristian Ionescu",
          "author_url": "",
          "post_date": "2018-05-02T22:08:25.557000",
          "content": "<p>Me too, please let me know if you find anything!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319063,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "2018-04-25T06:04:54.450000",
      "content": "<p>Seems like a cool competition! Using RMSE for probabilities however is a weird choice I haven't seen before.</p>",
      "votes": 19,
      "replies": [
        {
          "id": 319087,
          "author_name": "jairsan",
          "author_url": "",
          "post_date": "2018-04-25T07:37:52.743000",
          "content": "<p>Agreed! That's the first thing I noticed when looking at the competition. It would be interesting to know what is the rationale behind that choice because right now it just seems bizarre to me.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 320088,
          "author_name": "Evgeniy Malishev",
          "author_url": "",
          "post_date": "2018-04-27T12:50:09.233000",
          "content": "<p>As for the RMSE. Maybe this condition served to choose the evaluation of the competition.\n'Since the errors are squared before they are averaged, the RMSE gives a relatively high weight to large errors. This means the RMSE should be more useful when large errors are particularly undesirable.'\nAlthough maybe I'm wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320118,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2018-04-27T14:48:03.773000",
          "content": "<p>@Evgeniy The same is true for logloss, a very common metric for classification competitions</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 320400,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2018-04-28T15:22:01.310000",
          "content": "<p>Logloss won't work if the ground truth isn't 0/1 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 321078,
          "author_name": "Jonny Lomond",
          "author_url": "",
          "post_date": "2018-04-30T14:34:04.573000",
          "content": "<p>What about logloss doesn't work in this situation if the ground truth isn't 0/1? You can naively apply the same formula as the 0-1 ground truth case case, as long as the response is in [0,1]. </p>\n\n<p>For example if there is lots of data with few unique covariate combinations, it's very common to fit a binomial GLM (whose loss function is logloss) on the aggregated data with ground truth equal to percentage of observed successes in a cell and weight equal to number of observations in the cell. It's equivalent to \"unrolling\" the data to have 0-1 response and the same weight per row.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 321214,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2018-04-30T20:59:54.093000",
          "content": "<p>Binary log loss is a special case of <a href=\"https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence\">Kullback-Leibrer divergence</a> where p(x) is assumed to be either 0 or 1. In this case, the entropy of the p(x) is exactly zero and the loss is just a negative log of the correct class probability. Transition to a non-binary case is fairly straightforward - we just need to treat p(x) as a true *deal_probability*.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 321277,
          "author_name": "Wendy Kan",
          "author_url": "",
          "post_date": "2018-04-30T23:03:54.260000",
          "content": "<p>@Dmitriy Danevskiy thanks for the suggestion. Unfortunately we currently don't support that in our metrics library, we will keep it in mind for the future. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319903,
      "author_name": "Validation",
      "author_url": "",
      "post_date": "2018-04-27T04:40:31.753000",
      "content": "<p>Just curious where does the deal_probability on train data set come from, why is not it a 0/1 label?</p>",
      "votes": 11,
      "replies": [
        {
          "id": 321199,
          "author_name": "Parva Thakkar",
          "author_url": "",
          "post_date": "2018-04-30T20:03:16.120000",
          "content": "<p>May be the best way to look at deal_probability is to read it as ad_effectiveness. Closer it is to 1, more effective it is. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 319610,
      "author_name": "dd MlStart",
      "author_url": "",
      "post_date": "2018-04-26T12:40:34.523000",
      "content": "<p>Hi, can you put some light on what method/strategy was used to estimate  these  deal probability.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 319811,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-04-26T21:49:36.407000",
          "content": "<p>Would definitely appreciate that as well.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 323826,
      "author_name": "Jatin Balodhi",
      "author_url": "",
      "post_date": "2018-05-06T10:41:38.570000",
      "content": "<p>What are the columns params_1, params_2, params_3 signify and also what does image_top_1 represent?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 323854,
          "author_name": "Cristian Ionescu",
          "author_url": "",
          "post_date": "2018-05-06T12:18:25.827000",
          "content": "<p>Would definitely appreciate some help here, too, with how to interpret those columns.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 328535,
          "author_name": "Diana Kolusheva",
          "author_url": "",
          "post_date": "2018-05-14T15:08:57.350000",
          "content": "<p>Looking at the data, params_1,(2)(3) are just different levels of details about the product. For example in one of the rows: parent_category_name - personal things, category_name - clothes, shoes, accessories, param_1 - women clothes, param_2 - jeans, param_3 - 26 (size). In another example, they are selling a car and in params say that it is used, the model of the car, etc. \nNot sure about top_image_1 though. Might be some code assigned to each image at Avito.  </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 333883,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-26T00:55:33.847000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333920,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-05-26T03:40:37.923000",
          "content": "<p>@zanhanock you can see possible explanation here\n<a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57388\">https://www.kaggle.com/c/avito-demand-prediction/discussion/57388</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 333922,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-26T03:42:58.357000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319576,
      "author_name": "Mohammad Nishat Hussain",
      "author_url": "",
      "post_date": "2018-04-26T10:39:57.517000",
      "content": "<p>Is there any possibility of data sharing through Torrent, I have downloaded data through kaggle-API but the zip file is corrupted. [my <code>test_jpg.zip</code> filr is corrupted.]\nIs anyone facing similar issue ?? \nSomeone please share sha-hash to check.\nhow to start download from middle of a file using kaggle-API, or verify it.</p>\n\n<p>need help</p>",
      "votes": 1,
      "replies": [
        {
          "id": 319821,
          "author_name": "Eric Bouteillon",
          "author_url": "",
          "post_date": "2018-04-26T22:35:01.770000",
          "content": "<p>For those having trouble downloading files, here are SHA-256 checksum of files:\n<code>\n$ sha256sum *\n218950a51c6e87a7f3f4cfd7f4cccc459185d9ee5791598f35d5947b7368ec35  periods_test.csv.zip\ncda15d926826094364fca8decd90ad9cb786effdfa418dc3c258a7dc3c9af98a  periods_train.csv.zip\nf87c970b58e57dfc7b34fcd92ab0ecee44f69ccd76d35af105e0bb687495d6b2  sample_submission.csv\n63342b3f6828bd73092e345aacddc0810b0e43b4f1c558c0fb679315ce717ad4  test_active.csv.zip\nbe5123ad3a767c6d3a0467d170a70630ce3fbdeb5dbf4972f9e8960de02f3c23  test.csv.zip\n372cb16f10580e3b54eac7336079973bcea234b7ebffcaf323cae70aad0c3397  test_jpg.zip\n8ea9ac44a96e08af1016a618c11399db762f2ad53bc08ffc4bf781e15427d9d8  train_active.csv.zip\nbeeeeb4bf51df3a97ebfa5318f02d3a658efb1f0ac6e2ac312b9565b8496fcf6  train.csv.zip\naa0a8cc02c8e59020b943fae7aeedaaecd4f25d2460e150cac882ebf52f43519  train_jpg.zip\n</code></p>\n\n<p>I hope it helps.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 319404,
      "author_name": "Siddhant Bansal",
      "author_url": "",
      "post_date": "2018-04-26T01:27:31.767000",
      "content": "<p>Seems like a great task to work on!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 319497,
      "author_name": "VictorDZurkowski",
      "author_url": "",
      "post_date": "2018-04-26T07:23:23.710000",
      "content": "<p>The rules mention open source software. What about commercial software? I am thinking tools like AWS Rekognition, or other tool to interpret the image files.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 350907,
      "author_name": "kweonwooj",
      "author_url": "",
      "post_date": "2018-06-30T18:08:31.130000",
      "content": "<p>Dear Sohier Dane, (to competition organizer)</p>\n\n<p>I am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.</p>\n\n<p>Main content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.</p>\n\n<p>I was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives are possible content candidates.</p>\n\n<p>If using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.</p>\n\n<p>Please let me know, Best regards, Kweonwoo</p>",
      "votes": 0,
      "replies": [
        {
          "id": 351247,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-07-01T17:47:09.143000",
          "content": "<p>Kweonwoo,</p>\n\n<p>That sounds like a great project! All public Kaggle kernels are released under the Apache 2.0 open source license. As for the data, this is covered under sections 7.A and 7.B of the competition rules.</p>\n\n<p>Best,</p>\n\n<p>Sohier</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 350510,
      "author_name": "bangda",
      "author_url": "",
      "post_date": "2018-06-29T21:42:01.563000",
      "content": "<p>Hi @Sohier, I see the medals are released just before, I ranked 95/1917 but got bronze, and the leaderboard shows I'm still in silver. I just get confused, could you explain this? Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 350516,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-06-29T21:57:23.097000",
          "content": "<p>Sorry about the confusion, <a href=\"/bangda\">@bangda</a>. It can take some time for the final medals to propagate through our systems once the leaderboard has been finalized. Your profile should properly show silver by early next week; please message me again if that doesn't happen. This is definitely an area we could improve upon so thank you for your patience.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 350567,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-06-30T01:01:44.780000",
          "content": "<p>@Sohier， aha, similar issue here, my rank 191/1917 seems got no medal, but the leaderboard shows the bronze, the last one of bronze :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 353375,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-07-06T15:19:48.740000",
          "content": "<p>Thank you for your patience! I'm going to touch base with our engineering team, the correct medals should have come through by now. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 353846,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-08T02:26:32.827000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354943,
          "author_name": "bangda",
          "author_url": "",
          "post_date": "2018-07-10T14:29:55.100000",
          "content": "<p>Hi Sohier, the medal is still not changed...Could you double check with engineering team? Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354950,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-07-10T14:46:07.527000",
          "content": "<p>I filed a bug, but don't have an ETA on the solution yet.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 354951,
          "author_name": "bangda",
          "author_url": "",
          "post_date": "2018-07-10T14:49:21.373000",
          "content": "<p>Cool, great to know, hope it could be fixed soon;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 361046,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-23T18:38:48.460000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 346307,
      "author_name": "Andrea Rapuzzi",
      "author_url": "",
      "post_date": "2018-06-21T13:02:52.043000",
      "content": "<p>Hi Sohier,\nis it fine if I blend the output of my model with the output of another person model (made public in his kernel) and submit the result?\nThanks in advance!!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 346374,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-06-21T15:29:54.473000",
          "content": "<p>Yes, working off of public kernels is fine.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346375,
          "author_name": "Andrea Rapuzzi",
          "author_url": "",
          "post_date": "2018-06-21T15:31:05.550000",
          "content": "<p>Thanks!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 340325,
      "author_name": "Everton Leite",
      "author_url": "",
      "post_date": "2018-06-08T22:48:29.673000",
      "content": "<p>Let's play!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326658,
      "author_name": "Ankit Kulshreshtha",
      "author_url": "",
      "post_date": "2018-05-10T05:34:54.250000",
      "content": "<p>Interesting challenge! :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 327025,
          "author_name": "Vikas Kumar",
          "author_url": "",
          "post_date": "2018-05-10T17:39:10.683000",
          "content": "<p>I see</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319760,
      "author_name": "Di Ye",
      "author_url": "",
      "post_date": "2018-04-26T19:30:43.613000",
      "content": "<p>Hi, my excel couldn't recognize Russian. Can anyone help me with that?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 319813,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-04-26T21:51:33.623000",
          "content": "<p>For the title and description, you can try something like TextBlob to convert it to English. For categorical features, you can encode them.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 319472,
      "author_name": "Windforce",
      "author_url": "",
      "post_date": "2018-04-26T06:13:53.553000",
      "content": "<p>Thx</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319454,
      "author_name": "Rafael Batista",
      "author_url": "",
      "post_date": "2018-04-26T04:13:31.787000",
      "content": "<p>TeSt</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319391,
      "author_name": "adamrocker",
      "author_url": "",
      "post_date": "2018-04-26T00:37:43.150000",
      "content": "<p>+1</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319364,
      "author_name": "vvr",
      "author_url": "",
      "post_date": "2018-04-25T22:03:14.543000",
      "content": "<p>Any one who is Interested to form a team of 2 ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 319829,
          "author_name": "Jagadish",
          "author_url": "",
          "post_date": "2018-04-26T23:07:44.163000",
          "content": "<p>I am interested</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323997,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-06T21:54:01.643000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 345393,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-19T20:00:58.480000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319261,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-25T16:10:08.133000",
      "content": "",
      "votes": 9,
      "replies": [
        {
          "id": 319812,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-04-26T21:50:18.383000",
          "content": "<p>Seconded.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 350904,
      "author_name": "YiwenTang",
      "author_url": "",
      "post_date": "2018-06-30T18:03:04.300000",
      "content": "<p>Thank you for this competition!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 340315,
      "author_name": "Cheparukhin",
      "author_url": "",
      "post_date": "2018-06-08T21:57:58.410000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 320486,
      "author_name": "yuxiang1515",
      "author_url": "",
      "post_date": "2018-04-28T20:07:55.877000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319673,
      "author_name": "Sergey Enikeev",
      "author_url": "",
      "post_date": "2018-04-26T15:16:45.730000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319351,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-25T20:59:00.070000",
      "content": "<p>Amazing challenge thank you </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "318961": "Welcome to the Avito Demand Prediction Challenge, where you're tasked with estimating the success of online ads. \nYou'll need to integrate information from tabular, text, and image data. \nThe images are stored in zip archives that are too large to fit in memory of a kernel if loaded all at once, so you may find this [starter kernel](https://www.kaggle.com/sohier/getting-started-loading-the-images) about loading single files from the archives useful.\n(Note, due to the size of the files, it takes a few minutes for data to load up in Kernels.)\n\nHappy Kaggling!",
    "319000": "Is there any way to get a description of how the dependent variable is created?  I'm interested in learning in more detail what we are actually predicting.",
    "319063": "Seems like a cool competition! Using RMSE for probabilities however is a weird choice I haven't seen before.",
    "319903": "Just curious where does the deal_probability on train data set come from, why is not it a 0/1 label?",
    "319610": "Hi, can you put some light on what method/strategy was used to estimate  these  deal probability.",
    "323826": "What are the columns params_1, params_2, params_3 signify and also what does image_top_1 represent?",
    "319576": "Is there any possibility of data sharing through Torrent, I have downloaded data through kaggle-API but the zip file is corrupted. [my `test_jpg.zip` filr is corrupted.]\nIs anyone facing similar issue ?? \nSomeone please share sha-hash to check.\nhow to start download from middle of a file using kaggle-API, or verify it.\n\nneed help",
    "319404": "Seems like a great task to work on!",
    "319497": "The rules mention open source software. What about commercial software? I am thinking tools like AWS Rekognition, or other tool to interpret the image files.",
    "350907": "Dear Sohier Dane, (to competition organizer)\n\nI am planning to write a beginner's guide to kaggle for ML community in South Korea, and include this competition as a main competition case study.\n\nMain content of the publication will be EDA of this competition, and winner's solutions shared on github with appropriate licenses by competition winners.\n\nI was wondering if I can include EDA part of this competition in it. Visualizations of raw data or its derivatives are possible content candidates.\n\nIf using the data as-is is a sensitive matter and not allowed, I plan to include executable codes written by me with no image displays so that raw data or its derivatives are not included in the publication, but readers can follow the code and reproduce those on their computer.\n\nPlease let me know, Best regards, Kweonwoo",
    "350510": "Hi @Sohier, I see the medals are released just before, I ranked 95/1917 but got bronze, and the leaderboard shows I'm still in silver. I just get confused, could you explain this? Thank you!",
    "346307": "Hi Sohier,\nis it fine if I blend the output of my model with the output of another person model (made public in his kernel) and submit the result?\nThanks in advance!!",
    "340325": "Let's play!",
    "326658": "Interesting challenge! :)",
    "319760": "Hi, my excel couldn't recognize Russian. Can anyone help me with that?",
    "319472": "Thx",
    "319454": "TeSt",
    "319391": "+1",
    "319364": "Any one who is Interested to form a team of 2 ?",
    "345393": "",
    "319261": "",
    "350904": "Thank you for this competition!",
    "340315": "Thanks",
    "320486": "Thanks",
    "319673": "Thanks",
    "319351": "Amazing challenge thank you "
  }
}