{
  "id": 320026,
  "title": "DOLG Army: 7th Place Solution Summary",
  "url": "/competitions/happy-whale-and-dolphin/writeups/dolg-army-dolg-army-7th-place-solution-summary",
  "author_name": "",
  "post_date": "2022-04-21T03:22:58.023Z",
  "votes": 32,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all, </p>\n<p>Congratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition.It was as always very tough last week with huge LB changing everybody , fortunately for us , our ideas worked and we ended on right side of the private LB.</p>\n<p>Congratulations to <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> for becoming Kaggle Competitions GM. Thanks a lot to my teammates <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/navjotbansal\" target=\"_blank\">@navjotbansal</a>. Such a great team effort 🤜.</p>\n<p>Our solution consists of the following major components:</p>\n<ol>\n<li>Data Recipe</li>\n<li>Modelling</li>\n<li>Progressive Psuedo Labelling</li>\n<li>PostProcessing</li>\n<li>Ensemble</li>\n</ol>\n<p>All our models are trained on TPUs using more or less the same tensorflow pipeline that KS shared in the beginning of the competition.</p>\n<h2>Data Recipe</h2>\n<p>There were different datasets available publically, which we thought could help us in diversity. Inspired from <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> 's solution in Chaii, we decided to make our own data recipe. We used the following combination of datasets to train our models:</p>\n<ul>\n<li>Fullbody Annotations</li>\n<li>Fullbody Annotations + Backfins concatenated horizontally  </li>\n<li>Fullbody Annotations + Original Images concatenated horizontally</li>\n</ul>\n<p>Fullbody Annotations<br>\n<a href=\"https://ibb.co/LPtzkGQ\"><img src=\"https://i.ibb.co/KGqj01N/fullbody.png\" alt=\"fullbody\"></a><br><a target=\"_blank\" href=\"https://imgbb.com/\"></a><br></p>\n<p>Fullbody Annotations + Backfins concatenated horizontally <br>\n<a href=\"https://ibb.co/kH7CW6G\"><img src=\"https://i.ibb.co/r7jr1fQ/fullbody-backfin.png\" alt=\"fullbody-backfin\"></a><br><a target=\"_blank\" href=\"https://imgbb.com/\"></a><br></p>\n<p>Fullbody Annotations + Original Images concatenated horizontally<br>\n<a href=\"https://ibb.co/mz25mhN\"><img src=\"https://i.ibb.co/dKV4Skm/fullbody-original.png\" alt=\"fullbody-original\"></a><br><a target=\"_blank\" href=\"https://imgbb.com/\"></a><br></p>\n<p>We used simple augmentations:</p>\n<ul>\n<li>horizontal flip</li>\n<li>random pixel based augmentation (brightness, contrast, HSV)</li>\n<li>cutout</li>\n</ul>\n<h2>Modelling</h2>\n<p>We used a combination of DOLG (with EFFNet backbone) and normal EFFNets with CurricularFace loss. We used <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 's implementation of DOLG model, <a href=\"https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099\" target=\"_blank\">implemented in PyTorch</a> which we ported to TensorFlow because it worked better than the public implementation in this competition. We also used multiple heads in all the models: one for species classification and another for individual classification. Species classification head was trained with normal softmax loss while the individual classification head was trained with CurricularFace loss.</p>\n<p>Models in our final submission:</p>\n<ul>\n<li>DOLG B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins</li>\n<li>DOLG B6/B7, image_sizes: (896, 896x2), dataset: Fullbody Annotations + Original Images</li>\n<li>DOLG B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations</li>\n<li>EFFNet B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins</li>\n<li>EFFNet B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations</li>\n</ul>\n<p>All the models were trained on psuedo labelled data from our best ensemble. During inference we also use hflip as TTA.</p>\n<h2>PostProcessing</h2>\n<p>We used a second stage model for getting better confidence scores which is less susceptible to threshold changes. The idea was to use predictions/embeddings/features to get a better confidence score than just using nearest distances. After candidate generation (top100 neighbors based on knn distances) we engineered the following features for our 2nd stage model (features were generated from the models stated above):</p>\n<ul>\n<li>species probabilites for each image_id</li>\n<li>top3 nearest distances for each (image_id, unique individual_id) pair present in the candidates</li>\n<li>distance of each image_id from centroid of each unique individual_id present in the candidates</li>\n<li>rank of each unique individual_id present in the candidates</li>\n<li>sum of top3 neighbor distances </li>\n<li>OOF predictions</li>\n</ul>\n<p>We trained a 5 folds XGB and LightGBM models on the above features and used a ensemble of their predicitons as final confidence scores.</p>\n<h2>Ensemble</h2>\n<p>We used simple weighted voting approach using the confidence scores obtained from above where the weights were optimized using 5 fold OOFs.</p>\n<h2>Acknowledgements</h2>\n<ul>\n<li>Thanks a lot to <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> for sharing the datasets (The real hero 🔥)</li>\n<li>Special mention to Google TPU Research Program (<a href=\"https://sites.research.google/trc/\" target=\"_blank\">https://sites.research.google/trc/</a>) which helped us in training bigger models in the last few days of the competition by supporting us with V3 TPUs on GCP.</li>\n</ul>",
  "messages": [
    {
      "id": "1761241",
      "postDate": "04/19/2022 18:21:43",
      "content": "<p>Hi all, </p>\n<p>Congratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition.It was as always very tough last week with huge LB changing everybody , fortunately for us , our ideas worked and we ended on right side of the private LB.</p>\n<p>Congratulations to <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> for becoming Kaggle Competitions GM. Thanks a lot to my teammates <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> <a href=\"https://www.kaggle.com/navjotbansal\" target=\"_blank\">@navjotbansal</a>. Such a great team effort 🤜.</p>\n<p>Our solution consists of the following major components:</p>\n<ol>\n<li>Data Recipe</li>\n<li>Modelling</li>\n<li>Progressive Psuedo Labelling</li>\n<li>PostProcessing</li>\n<li>Ensemble</li>\n</ol>\n<p>All our models are trained on TPUs using more or less the same tensorflow pipeline that KS shared in the beginning of the competition.</p>\n<h2>Data Recipe</h2>\n<p>There were different datasets available publically, which we thought could help us in diversity. Inspired from <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> 's solution in Chaii, we decided to make our own data recipe. We used the following combination of datasets to train our models:</p>\n<ul>\n<li>Fullbody Annotations</li>\n<li>Fullbody Annotations + Backfins concatenated horizontally  </li>\n<li>Fullbody Annotations + Original Images concatenated horizontally</li>\n</ul>\n<p>Fullbody Annotations<br>\n<a href=\"https://ibb.co/LPtzkGQ\"><img src=\"https://i.ibb.co/KGqj01N/fullbody.png\" alt=\"fullbody\"></a><br><a target=\"_blank\" href=\"https://imgbb.com/\"></a><br></p>\n<p>Fullbody Annotations + Backfins concatenated horizontally <br>\n<a href=\"https://ibb.co/kH7CW6G\"><img src=\"https://i.ibb.co/r7jr1fQ/fullbody-backfin.png\" alt=\"fullbody-backfin\"></a><br><a target=\"_blank\" href=\"https://imgbb.com/\"></a><br></p>\n<p>Fullbody Annotations + Original Images concatenated horizontally<br>\n<a href=\"https://ibb.co/mz25mhN\"><img src=\"https://i.ibb.co/dKV4Skm/fullbody-original.png\" alt=\"fullbody-original\"></a><br><a target=\"_blank\" href=\"https://imgbb.com/\"></a><br></p>\n<p>We used simple augmentations:</p>\n<ul>\n<li>horizontal flip</li>\n<li>random pixel based augmentation (brightness, contrast, HSV)</li>\n<li>cutout</li>\n</ul>\n<h2>Modelling</h2>\n<p>We used a combination of DOLG (with EFFNet backbone) and normal EFFNets with CurricularFace loss. We used <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 's implementation of DOLG model, <a href=\"https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099\" target=\"_blank\">implemented in PyTorch</a> which we ported to TensorFlow because it worked better than the public implementation in this competition. We also used multiple heads in all the models: one for species classification and another for individual classification. Species classification head was trained with normal softmax loss while the individual classification head was trained with CurricularFace loss.</p>\n<p>Models in our final submission:</p>\n<ul>\n<li>DOLG B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins</li>\n<li>DOLG B6/B7, image_sizes: (896, 896x2), dataset: Fullbody Annotations + Original Images</li>\n<li>DOLG B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations</li>\n<li>EFFNet B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins</li>\n<li>EFFNet B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations</li>\n</ul>\n<p>All the models were trained on psuedo labelled data from our best ensemble. During inference we also use hflip as TTA.</p>\n<h2>PostProcessing</h2>\n<p>We used a second stage model for getting better confidence scores which is less susceptible to threshold changes. The idea was to use predictions/embeddings/features to get a better confidence score than just using nearest distances. After candidate generation (top100 neighbors based on knn distances) we engineered the following features for our 2nd stage model (features were generated from the models stated above):</p>\n<ul>\n<li>species probabilites for each image_id</li>\n<li>top3 nearest distances for each (image_id, unique individual_id) pair present in the candidates</li>\n<li>distance of each image_id from centroid of each unique individual_id present in the candidates</li>\n<li>rank of each unique individual_id present in the candidates</li>\n<li>sum of top3 neighbor distances </li>\n<li>OOF predictions</li>\n</ul>\n<p>We trained a 5 folds XGB and LightGBM models on the above features and used a ensemble of their predicitons as final confidence scores.</p>\n<h2>Ensemble</h2>\n<p>We used simple weighted voting approach using the confidence scores obtained from above where the weights were optimized using 5 fold OOFs.</p>\n<h2>Acknowledgements</h2>\n<ul>\n<li>Thanks a lot to <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> for sharing the datasets (The real hero 🔥)</li>\n<li>Special mention to Google TPU Research Program (<a href=\"https://sites.research.google/trc/\" target=\"_blank\">https://sites.research.google/trc/</a>) which helped us in training bigger models in the last few days of the competition by supporting us with V3 TPUs on GCP.</li>\n</ul>",
      "rawMarkdown": "Hi all, \n\nCongratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition.It was as always very tough last week with huge LB changing everybody , fortunately for us , our ideas worked and we ended on right side of the private LB.\n\nCongratulations to @ks2019 for becoming Kaggle Competitions GM. Thanks a lot to my teammates @nischaydnk @tanulsingh077 @navjotbansal. Such a great team effort 🤜.\n\nOur solution consists of the following major components:\n1. Data Recipe\n2. Modelling\n3. Progressive Psuedo Labelling\n4. PostProcessing\n5. Ensemble\n\nAll our models are trained on TPUs using more or less the same tensorflow pipeline that KS shared in the beginning of the competition.\n\n## Data Recipe\nThere were different datasets available publically, which we thought could help us in diversity. Inspired from @thedrcat 's solution in Chaii, we decided to make our own data recipe. We used the following combination of datasets to train our models:\n- Fullbody Annotations\n- Fullbody Annotations + Backfins concatenated horizontally  \n- Fullbody Annotations + Original Images concatenated horizontally\n\nFullbody Annotations\n<a href=\"https://ibb.co/LPtzkGQ\"><img src=\"https://i.ibb.co/KGqj01N/fullbody.png\" alt=\"fullbody\" border=\"0\"></a><br /><a target='_blank' href='https://imgbb.com/'></a><br/>\n\nFullbody Annotations + Backfins concatenated horizontally \n<a href=\"https://ibb.co/kH7CW6G\"><img src=\"https://i.ibb.co/r7jr1fQ/fullbody-backfin.png\" alt=\"fullbody-backfin\" border=\"0\"></a><br /><a target='_blank' href='https://imgbb.com/'></a><br/>\n\nFullbody Annotations + Original Images concatenated horizontally\n<a href=\"https://ibb.co/mz25mhN\"><img src=\"https://i.ibb.co/dKV4Skm/fullbody-original.png\" alt=\"fullbody-original\" border=\"0\"></a><br /><a target='_blank' href='https://imgbb.com/'></a><br/>\n\nWe used simple augmentations:\n- horizontal flip\n- random pixel based augmentation (brightness, contrast, HSV)\n- cutout\n\n## Modelling\nWe used a combination of DOLG (with EFFNet backbone) and normal EFFNets with CurricularFace loss. We used @christofhenkel 's implementation of DOLG model, [implemented in PyTorch](https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099) which we ported to TensorFlow because it worked better than the public implementation in this competition. We also used multiple heads in all the models: one for species classification and another for individual classification. Species classification head was trained with normal softmax loss while the individual classification head was trained with CurricularFace loss.\n\nModels in our final submission:\n- DOLG B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins\n- DOLG B6/B7, image_sizes: (896, 896x2), dataset: Fullbody Annotations + Original Images\n- DOLG B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations\n- EFFNet B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins\n- EFFNet B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations\n\nAll the models were trained on psuedo labelled data from our best ensemble. During inference we also use hflip as TTA.\n\n## PostProcessing\nWe used a second stage model for getting better confidence scores which is less susceptible to threshold changes. The idea was to use predictions/embeddings/features to get a better confidence score than just using nearest distances. After candidate generation (top100 neighbors based on knn distances) we engineered the following features for our 2nd stage model (features were generated from the models stated above):\n- species probabilites for each image_id\n- top3 nearest distances for each (image_id, unique individual_id) pair present in the candidates\n- distance of each image_id from centroid of each unique individual_id present in the candidates\n- rank of each unique individual_id present in the candidates\n- sum of top3 neighbor distances \n- OOF predictions\n\nWe trained a 5 folds XGB and LightGBM models on the above features and used a ensemble of their predicitons as final confidence scores.\n\n## Ensemble\nWe used simple weighted voting approach using the confidence scores obtained from above where the weights were optimized using 5 fold OOFs.\n\n## Acknowledgements\n- Thanks a lot to @jpbremer for sharing the datasets (The real hero 🔥)\n- Special mention to Google TPU Research Program (https://sites.research.google/trc/) which helped us in training bigger models in the last few days of the competition by supporting us with V3 TPUs on GCP.",
      "votes": null
    },
    {
      "id": "1761383",
      "postDate": "04/19/2022 21:10:12",
      "content": "<p>Thanks for your sharing. Can you share your source code? I'm really interested in stacking part</p>",
      "rawMarkdown": "Thanks for your sharing. Can you share your source code? I'm really interested in stacking part",
      "votes": null
    },
    {
      "id": "1761429",
      "postDate": "04/19/2022 22:48:02",
      "content": "<p>Good job, I am also interested in PostProcessing part.</p>",
      "rawMarkdown": "Good job, I am also interested in PostProcessing part.",
      "votes": null
    },
    {
      "id": "1763110",
      "postDate": "04/21/2022 09:25:12",
      "content": "<p>Hi, sorry for late reply. Taking some time off from kaggle after intense competition. We'll make the complete source code public in few days.  </p>",
      "rawMarkdown": "Hi, sorry for late reply. Taking some time off from kaggle after intense competition. We'll make the complete source code public in few days.",
      "votes": null
    },
    {
      "id": "1763118",
      "postDate": "04/21/2022 09:29:39",
      "content": "<p>Here is the code for feature engineering part. We trained the LGBM and XGBoost model on these features. Hyperparameters tuned with optuna to optimise competition metric.  This gave us ~ 0.01-0.015 boost in public &amp; private lb.</p>\n<pre><code>KNN = 100\nnn_cols = [f'nn-{x}' for x in range(KNN)]\ndist_cols = [f'dist-{x}' for x in range(KNN)]\nreciprocals = 1/(np.arange(1,6))\n\ndef generate_preds(row,threshold=0.5):\n\n    nns = list(zip(row[dist_cols],row[nn_cols].values))\n    nns.append((threshold,-1))\n    preds = []\n    for dist,pred in sorted(nns):\n        if pred not in preds:\n            preds.append(pred)\n        if len(preds)==5:\n            break\n    if len(preds)&lt;5:\n        preds = preds+[-1,-1,-1,-1,-1]\n        preds = preds[:5]\n    return np.array(preds)\n\n\ndef map_per_sample(row):\n    return ((row.target==row['preds'])*reciprocals).sum()\n\ndef get_sample_neighbours(data,MODE):\n\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    allowed_targets = set(train_targets)\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(train_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = train_targets[idxs]\n    ids = train_ids[idxs]\n\n    val_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    val_predictions['nn_ids'] = train_ids[idxs].reshape(-1)\n    val_predictions['nn_distances'] = distances.reshape(-1)\n    val_predictions['nn_rank'] = val_predictions.index//KNN\n    val_predictions['image'] = val_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    if MODE=='TRAIN':\n        val_predictions['target'] = val_predictions['nn_rank'].apply(lambda x: val_targets[x])\n    else:\n        val_predictions['target'] = -1\n    val_predictions['nn_rank'] = val_predictions.index%KNN\n    return val_predictions\n\ndef get_centroid_neighbours(data):\n\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    train_embeddings_df = pd.DataFrame(train_embeddings)\n    train_embeddings_df['target'] = train_targets\n    centroid_embeddings = train_embeddings_df.groupby('target').mean()\n    centroid_targets = centroid_embeddings.index.values\n    centroid_embeddings = centroid_embeddings.values\n\n    from sklearn.neighbors import NearestNeighbors\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(centroid_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = centroid_targets[idxs]\n\n    centroid_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    centroid_predictions['nn_distances'] = distances.reshape(-1)\n    centroid_predictions['nn_rank'] = centroid_predictions.index//KNN\n    centroid_predictions['image'] = centroid_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    centroid_predictions['nn_rank'] = centroid_predictions.index%KNN\n    centroid_predictions = centroid_predictions.set_index(['image','nn_target']).nn_distances.to_dict()\n\n    return centroid_predictions\n\ndef get_target_counts(data):\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    target_counts = pd.Series(train_targets).value_counts().to_dict()\n    return target_counts\n\ndef engineer_features(train_pattern,test_pattern,fold=0,MODE='TEST'):\n\n    print(\"Loading Data\")\n    val_species_ids,val_species,test_species_ids,test_species = load_species(fold=fold)\n    val_species = {x:y for x,y in zip(val_species_ids,val_species)}\n    test_species = {x:y for x,y in zip(test_species_ids,test_species)}\n\n    if MODE=='TEST':\n        data = load_test_embeds(train_pattern,test_pattern,fold=fold)\n        species = test_species\n    else:\n        data = load_embeds(train_pattern,test_pattern,fold=fold)\n        species = val_species\n\n    spec_df = pd.DataFrame(species).T.reset_index()\n    spec_df.columns = ['image']+[f'spec_{i}' for i in range(0,26)]\n\n    print(\"Getting Sample Neighbours\")\n    val_predictions = get_sample_neighbours(data,MODE)\n\n    print(\"Getting Centroid Neighbours\")\n    centroid_predictions = get_centroid_neighbours(data)\n\n    print(\"Getting Target Counts\")\n    target_counts = get_target_counts(data)\n\n    print(len(species),len(val_predictions))\n\n    train_df = val_predictions.groupby(['image','nn_target']).head(3)\n    train_df['nn_rank_target'] = train_df.groupby(['image','nn_target']).nn_distances.rank(method='first')\n    train_df = pd.pivot_table(train_df, values='nn_distances', index=['image','target', 'nn_target'],\n                        columns=['nn_rank_target'], aggfunc=np.sum).reset_index().fillna(1)\n    train_df = train_df.rename(columns={1:'nn1',2:'nn2',3:'nn3','nn_target':'individual'}).reset_index(drop=True)\n    train_df['centroid'] = train_df.apply(lambda row:centroid_predictions[(row.image,row.individual)]\n                                              if (row.image,row.individual) in centroid_predictions else 1\n                                          ,axis=1)\n    train_df['target_counts'] = train_df.apply(lambda row:target_counts[row.individual],axis=1)\n    train_df['individual_species'] = train_df.individual.map(species_individual_mappings)\n    train_df['species_conf'] = train_df.apply(lambda row:species[row.image][row.individual_species],axis=1)\n#     train_df[[f'species_{i}' for i in range(26)]] = train_df.apply(lambda row: species[row.image],axis=1)\n    return train_df,spec_df\n</code></pre>",
      "rawMarkdown": "Here is the code for feature engineering part. We trained the LGBM and XGBoost model on these features. Hyperparameters tuned with optuna to optimise competition metric.  This gave us ~ 0.01-0.015 boost in public & private lb.\n\n```\nKNN = 100\nnn_cols = [f'nn-{x}' for x in range(KNN)]\ndist_cols = [f'dist-{x}' for x in range(KNN)]\nreciprocals = 1/(np.arange(1,6))\n\ndef generate_preds(row,threshold=0.5):\n    \n    nns = list(zip(row[dist_cols],row[nn_cols].values))\n    nns.append((threshold,-1))\n    preds = []\n    for dist,pred in sorted(nns):\n        if pred not in preds:\n            preds.append(pred)\n        if len(preds)==5:\n            break\n    if len(preds)<5:\n        preds = preds+[-1,-1,-1,-1,-1]\n        preds = preds[:5]\n    return np.array(preds)\n    \n    \ndef map_per_sample(row):\n    return ((row.target==row['preds'])*reciprocals).sum()\n\ndef get_sample_neighbours(data,MODE):\n    \n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    allowed_targets = set(train_targets)\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(train_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = train_targets[idxs]\n    ids = train_ids[idxs]\n\n    val_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    val_predictions['nn_ids'] = train_ids[idxs].reshape(-1)\n    val_predictions['nn_distances'] = distances.reshape(-1)\n    val_predictions['nn_rank'] = val_predictions.index//KNN\n    val_predictions['image'] = val_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    if MODE=='TRAIN':\n        val_predictions['target'] = val_predictions['nn_rank'].apply(lambda x: val_targets[x])\n    else:\n        val_predictions['target'] = -1\n    val_predictions['nn_rank'] = val_predictions.index%KNN\n    return val_predictions\n\ndef get_centroid_neighbours(data):\n    \n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    train_embeddings_df = pd.DataFrame(train_embeddings)\n    train_embeddings_df['target'] = train_targets\n    centroid_embeddings = train_embeddings_df.groupby('target').mean()\n    centroid_targets = centroid_embeddings.index.values\n    centroid_embeddings = centroid_embeddings.values\n\n    from sklearn.neighbors import NearestNeighbors\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(centroid_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = centroid_targets[idxs]\n\n    centroid_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    centroid_predictions['nn_distances'] = distances.reshape(-1)\n    centroid_predictions['nn_rank'] = centroid_predictions.index//KNN\n    centroid_predictions['image'] = centroid_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    centroid_predictions['nn_rank'] = centroid_predictions.index%KNN\n    centroid_predictions = centroid_predictions.set_index(['image','nn_target']).nn_distances.to_dict()\n    \n    return centroid_predictions\n\ndef get_target_counts(data):\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    target_counts = pd.Series(train_targets).value_counts().to_dict()\n    return target_counts\n\ndef engineer_features(train_pattern,test_pattern,fold=0,MODE='TEST'):\n    \n    print(\"Loading Data\")\n    val_species_ids,val_species,test_species_ids,test_species = load_species(fold=fold)\n    val_species = {x:y for x,y in zip(val_species_ids,val_species)}\n    test_species = {x:y for x,y in zip(test_species_ids,test_species)}\n    \n    if MODE=='TEST':\n        data = load_test_embeds(train_pattern,test_pattern,fold=fold)\n        species = test_species\n    else:\n        data = load_embeds(train_pattern,test_pattern,fold=fold)\n        species = val_species\n        \n    spec_df = pd.DataFrame(species).T.reset_index()\n    spec_df.columns = ['image']+[f'spec_{i}' for i in range(0,26)]\n    \n    print(\"Getting Sample Neighbours\")\n    val_predictions = get_sample_neighbours(data,MODE)\n    \n    print(\"Getting Centroid Neighbours\")\n    centroid_predictions = get_centroid_neighbours(data)\n    \n    print(\"Getting Target Counts\")\n    target_counts = get_target_counts(data)\n    \n    print(len(species),len(val_predictions))\n    \n    train_df = val_predictions.groupby(['image','nn_target']).head(3)\n    train_df['nn_rank_target'] = train_df.groupby(['image','nn_target']).nn_distances.rank(method='first')\n    train_df = pd.pivot_table(train_df, values='nn_distances', index=['image','target', 'nn_target'],\n                        columns=['nn_rank_target'], aggfunc=np.sum).reset_index().fillna(1)\n    train_df = train_df.rename(columns={1:'nn1',2:'nn2',3:'nn3','nn_target':'individual'}).reset_index(drop=True)\n    train_df['centroid'] = train_df.apply(lambda row:centroid_predictions[(row.image,row.individual)]\n                                              if (row.image,row.individual) in centroid_predictions else 1\n                                          ,axis=1)\n    train_df['target_counts'] = train_df.apply(lambda row:target_counts[row.individual],axis=1)\n    train_df['individual_species'] = train_df.individual.map(species_individual_mappings)\n    train_df['species_conf'] = train_df.apply(lambda row:species[row.image][row.individual_species],axis=1)\n#     train_df[[f'species_{i}' for i in range(26)]] = train_df.apply(lambda row: species[row.image],axis=1)\n    return train_df,spec_df\n\n```",
      "votes": null
    },
    {
      "id": "1765975",
      "postDate": "04/24/2022 04:24:37",
      "content": "<p>Great job on the competition. I had heard that not very much computing resources went into your training pipeline. Is this true? How many iterations did your team go through to get such good results?</p>",
      "rawMarkdown": "Great job on the competition. I had heard that not very much computing resources went into your training pipeline. Is this true? How many iterations did your team go through to get such good results?",
      "votes": null
    },
    {
      "id": "1766206",
      "postDate": "04/24/2022 09:47:04",
      "content": "<p>Thanks for your reply！</p>",
      "rawMarkdown": "Thanks for your reply！",
      "votes": null
    },
    {
      "id": "1767557",
      "postDate": "04/25/2022 13:21:23",
      "content": "<p>No, It's not that. There is some miscommunication. We didn't have any personal hardware (GPUs). We used TPUs to train our models, as some other top teams did. But, TPUs are good compute resources, available on Kaggle and Colab.  Almost all of my kaggle finishes are with the models trained on TPUs/GPUs provided by Kaggle and Colab, except few in which I teamed up with someone who has good hardware. This is what Tanul wanted to convey in his post. <br>\nWe have been competing actively in this competition for last 1.5 months. So, our results has been results of many iterations. At first, our postprocessing was fixed (the same which I released in first week of start of the competition), and we tuned the architecture/training strategy for both DOLG &amp; Non-DOLG models. Then, a lot of time went in understanding which preprocessing works better and designing data recipe, experimenting with psuedolabelling. Last week went in mostly designing &amp; optimising postprocessing.</p>",
      "rawMarkdown": "No, It's not that. There is some miscommunication. We didn't have any personal hardware (GPUs). We used TPUs to train our models, as some other top teams did. But, TPUs are good compute resources, available on Kaggle and Colab.  Almost all of my kaggle finishes are with the models trained on TPUs/GPUs provided by Kaggle and Colab, except few in which I teamed up with someone who has good hardware. This is what Tanul wanted to convey in his post. \nWe have been competing actively in this competition for last 1.5 months. So, our results has been results of many iterations. At first, our postprocessing was fixed (the same which I released in first week of start of the competition), and we tuned the architecture/training strategy for both DOLG & Non-DOLG models. Then, a lot of time went in understanding which preprocessing works better and designing data recipe, experimenting with psuedolabelling. Last week went in mostly designing & optimising postprocessing.",
      "votes": null
    },
    {
      "id": "1767615",
      "postDate": "04/25/2022 14:02:34",
      "content": "<p><a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a>  thanks for your sharing. Can you share your code for training LGB/XGB. I'm not familiar with stacking and want to learn.</p>",
      "rawMarkdown": "ks2019  thanks for your sharing. Can you share your code for training LGB/XGB. I'm not familiar with stacking and want to learn.",
      "votes": null
    },
    {
      "id": "1852856",
      "postDate": "07/12/2022 12:39:06",
      "content": "<p>I wanted to ask about the your 2nd stage model, I bit confused what was the target used for it<br>\nLike suppose you have 1 image and its 100 possible matches (founded via knn), then you generate features for that image using the matches, now here what target do you use, and is it the knn distance or the cosine similarity?</p>",
      "rawMarkdown": "I wanted to ask about the your 2nd stage model, I bit confused what was the target used for it\nLike suppose you have 1 image and its 100 possible matches (founded via knn), then you generate features for that image using the matches, now here what target do you use, and is it the knn distance or the cosine similarity?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1761383,
      "author_name": "cuongnn218",
      "author_url": "",
      "post_date": "04/19/2022 21:10:12",
      "content": "<p>Thanks for your sharing. Can you share your source code? I'm really interested in stacking part</p>",
      "votes": null,
      "replies": [
        {
          "id": 1761429,
          "author_name": "librauee",
          "author_url": "",
          "post_date": "04/19/2022 22:48:02",
          "content": "<p>Good job, I am also interested in PostProcessing part.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1763110,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "04/21/2022 09:25:12",
          "content": "<p>Hi, sorry for late reply. Taking some time off from kaggle after intense competition. We'll make the complete source code public in few days.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1763118,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "04/21/2022 09:29:39",
          "content": "<p>Here is the code for feature engineering part. We trained the LGBM and XGBoost model on these features. Hyperparameters tuned with optuna to optimise competition metric.  This gave us ~ 0.01-0.015 boost in public &amp; private lb.</p>\n<pre><code>KNN = 100\nnn_cols = [f'nn-{x}' for x in range(KNN)]\ndist_cols = [f'dist-{x}' for x in range(KNN)]\nreciprocals = 1/(np.arange(1,6))\n\ndef generate_preds(row,threshold=0.5):\n\n    nns = list(zip(row[dist_cols],row[nn_cols].values))\n    nns.append((threshold,-1))\n    preds = []\n    for dist,pred in sorted(nns):\n        if pred not in preds:\n            preds.append(pred)\n        if len(preds)==5:\n            break\n    if len(preds)&lt;5:\n        preds = preds+[-1,-1,-1,-1,-1]\n        preds = preds[:5]\n    return np.array(preds)\n\n\ndef map_per_sample(row):\n    return ((row.target==row['preds'])*reciprocals).sum()\n\ndef get_sample_neighbours(data,MODE):\n\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    allowed_targets = set(train_targets)\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(train_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = train_targets[idxs]\n    ids = train_ids[idxs]\n\n    val_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    val_predictions['nn_ids'] = train_ids[idxs].reshape(-1)\n    val_predictions['nn_distances'] = distances.reshape(-1)\n    val_predictions['nn_rank'] = val_predictions.index//KNN\n    val_predictions['image'] = val_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    if MODE=='TRAIN':\n        val_predictions['target'] = val_predictions['nn_rank'].apply(lambda x: val_targets[x])\n    else:\n        val_predictions['target'] = -1\n    val_predictions['nn_rank'] = val_predictions.index%KNN\n    return val_predictions\n\ndef get_centroid_neighbours(data):\n\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    train_embeddings_df = pd.DataFrame(train_embeddings)\n    train_embeddings_df['target'] = train_targets\n    centroid_embeddings = train_embeddings_df.groupby('target').mean()\n    centroid_targets = centroid_embeddings.index.values\n    centroid_embeddings = centroid_embeddings.values\n\n    from sklearn.neighbors import NearestNeighbors\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(centroid_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = centroid_targets[idxs]\n\n    centroid_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    centroid_predictions['nn_distances'] = distances.reshape(-1)\n    centroid_predictions['nn_rank'] = centroid_predictions.index//KNN\n    centroid_predictions['image'] = centroid_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    centroid_predictions['nn_rank'] = centroid_predictions.index%KNN\n    centroid_predictions = centroid_predictions.set_index(['image','nn_target']).nn_distances.to_dict()\n\n    return centroid_predictions\n\ndef get_target_counts(data):\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    target_counts = pd.Series(train_targets).value_counts().to_dict()\n    return target_counts\n\ndef engineer_features(train_pattern,test_pattern,fold=0,MODE='TEST'):\n\n    print(\"Loading Data\")\n    val_species_ids,val_species,test_species_ids,test_species = load_species(fold=fold)\n    val_species = {x:y for x,y in zip(val_species_ids,val_species)}\n    test_species = {x:y for x,y in zip(test_species_ids,test_species)}\n\n    if MODE=='TEST':\n        data = load_test_embeds(train_pattern,test_pattern,fold=fold)\n        species = test_species\n    else:\n        data = load_embeds(train_pattern,test_pattern,fold=fold)\n        species = val_species\n\n    spec_df = pd.DataFrame(species).T.reset_index()\n    spec_df.columns = ['image']+[f'spec_{i}' for i in range(0,26)]\n\n    print(\"Getting Sample Neighbours\")\n    val_predictions = get_sample_neighbours(data,MODE)\n\n    print(\"Getting Centroid Neighbours\")\n    centroid_predictions = get_centroid_neighbours(data)\n\n    print(\"Getting Target Counts\")\n    target_counts = get_target_counts(data)\n\n    print(len(species),len(val_predictions))\n\n    train_df = val_predictions.groupby(['image','nn_target']).head(3)\n    train_df['nn_rank_target'] = train_df.groupby(['image','nn_target']).nn_distances.rank(method='first')\n    train_df = pd.pivot_table(train_df, values='nn_distances', index=['image','target', 'nn_target'],\n                        columns=['nn_rank_target'], aggfunc=np.sum).reset_index().fillna(1)\n    train_df = train_df.rename(columns={1:'nn1',2:'nn2',3:'nn3','nn_target':'individual'}).reset_index(drop=True)\n    train_df['centroid'] = train_df.apply(lambda row:centroid_predictions[(row.image,row.individual)]\n                                              if (row.image,row.individual) in centroid_predictions else 1\n                                          ,axis=1)\n    train_df['target_counts'] = train_df.apply(lambda row:target_counts[row.individual],axis=1)\n    train_df['individual_species'] = train_df.individual.map(species_individual_mappings)\n    train_df['species_conf'] = train_df.apply(lambda row:species[row.image][row.individual_species],axis=1)\n#     train_df[[f'species_{i}' for i in range(26)]] = train_df.apply(lambda row: species[row.image],axis=1)\n    return train_df,spec_df\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1766206,
          "author_name": "librauee",
          "author_url": "",
          "post_date": "04/24/2022 09:47:04",
          "content": "<p>Thanks for your reply！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1767615,
          "author_name": "cuongnn218",
          "author_url": "",
          "post_date": "04/25/2022 14:02:34",
          "content": "<p><a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a>  thanks for your sharing. Can you share your code for training LGB/XGB. I'm not familiar with stacking and want to learn.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1765975,
      "author_name": "anandparthiban",
      "author_url": "",
      "post_date": "04/24/2022 04:24:37",
      "content": "<p>Great job on the competition. I had heard that not very much computing resources went into your training pipeline. Is this true? How many iterations did your team go through to get such good results?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1767557,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "04/25/2022 13:21:23",
          "content": "<p>No, It's not that. There is some miscommunication. We didn't have any personal hardware (GPUs). We used TPUs to train our models, as some other top teams did. But, TPUs are good compute resources, available on Kaggle and Colab.  Almost all of my kaggle finishes are with the models trained on TPUs/GPUs provided by Kaggle and Colab, except few in which I teamed up with someone who has good hardware. This is what Tanul wanted to convey in his post. <br>\nWe have been competing actively in this competition for last 1.5 months. So, our results has been results of many iterations. At first, our postprocessing was fixed (the same which I released in first week of start of the competition), and we tuned the architecture/training strategy for both DOLG &amp; Non-DOLG models. Then, a lot of time went in understanding which preprocessing works better and designing data recipe, experimenting with psuedolabelling. Last week went in mostly designing &amp; optimising postprocessing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1852856,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "07/12/2022 12:39:06",
      "content": "<p>I wanted to ask about the your 2nd stage model, I bit confused what was the target used for it<br>\nLike suppose you have 1 image and its 100 possible matches (founded via knn), then you generate features for that image using the matches, now here what target do you use, and is it the knn distance or the cosine similarity?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1761241": "Hi all, \n\nCongratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition.It was as always very tough last week with huge LB changing everybody , fortunately for us , our ideas worked and we ended on right side of the private LB.\n\nCongratulations to @ks2019 for becoming Kaggle Competitions GM. Thanks a lot to my teammates @nischaydnk @tanulsingh077 @navjotbansal. Such a great team effort 🤜.\n\nOur solution consists of the following major components:\n1. Data Recipe\n2. Modelling\n3. Progressive Psuedo Labelling\n4. PostProcessing\n5. Ensemble\n\nAll our models are trained on TPUs using more or less the same tensorflow pipeline that KS shared in the beginning of the competition.\n\n## Data Recipe\nThere were different datasets available publically, which we thought could help us in diversity. Inspired from @thedrcat 's solution in Chaii, we decided to make our own data recipe. We used the following combination of datasets to train our models:\n- Fullbody Annotations\n- Fullbody Annotations + Backfins concatenated horizontally  \n- Fullbody Annotations + Original Images concatenated horizontally\n\nFullbody Annotations\n<a href=\"https://ibb.co/LPtzkGQ\"><img src=\"https://i.ibb.co/KGqj01N/fullbody.png\" alt=\"fullbody\" border=\"0\"></a><br /><a target='_blank' href='https://imgbb.com/'></a><br/>\n\nFullbody Annotations + Backfins concatenated horizontally \n<a href=\"https://ibb.co/kH7CW6G\"><img src=\"https://i.ibb.co/r7jr1fQ/fullbody-backfin.png\" alt=\"fullbody-backfin\" border=\"0\"></a><br /><a target='_blank' href='https://imgbb.com/'></a><br/>\n\nFullbody Annotations + Original Images concatenated horizontally\n<a href=\"https://ibb.co/mz25mhN\"><img src=\"https://i.ibb.co/dKV4Skm/fullbody-original.png\" alt=\"fullbody-original\" border=\"0\"></a><br /><a target='_blank' href='https://imgbb.com/'></a><br/>\n\nWe used simple augmentations:\n- horizontal flip\n- random pixel based augmentation (brightness, contrast, HSV)\n- cutout\n\n## Modelling\nWe used a combination of DOLG (with EFFNet backbone) and normal EFFNets with CurricularFace loss. We used @christofhenkel 's implementation of DOLG model, [implemented in PyTorch](https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099) which we ported to TensorFlow because it worked better than the public implementation in this competition. We also used multiple heads in all the models: one for species classification and another for individual classification. Species classification head was trained with normal softmax loss while the individual classification head was trained with CurricularFace loss.\n\nModels in our final submission:\n- DOLG B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins\n- DOLG B6/B7, image_sizes: (896, 896x2), dataset: Fullbody Annotations + Original Images\n- DOLG B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations\n- EFFNet B5/B6/B7, image_sizes: (786, 786x2), (896, 896x2), dataset: Fullbody Annotations + Backfins\n- EFFNet B5/B6/B7, image_sizes: (1024, 1024), dataset: Fullbody Annotations\n\nAll the models were trained on psuedo labelled data from our best ensemble. During inference we also use hflip as TTA.\n\n## PostProcessing\nWe used a second stage model for getting better confidence scores which is less susceptible to threshold changes. The idea was to use predictions/embeddings/features to get a better confidence score than just using nearest distances. After candidate generation (top100 neighbors based on knn distances) we engineered the following features for our 2nd stage model (features were generated from the models stated above):\n- species probabilites for each image_id\n- top3 nearest distances for each (image_id, unique individual_id) pair present in the candidates\n- distance of each image_id from centroid of each unique individual_id present in the candidates\n- rank of each unique individual_id present in the candidates\n- sum of top3 neighbor distances \n- OOF predictions\n\nWe trained a 5 folds XGB and LightGBM models on the above features and used a ensemble of their predicitons as final confidence scores.\n\n## Ensemble\nWe used simple weighted voting approach using the confidence scores obtained from above where the weights were optimized using 5 fold OOFs.\n\n## Acknowledgements\n- Thanks a lot to @jpbremer for sharing the datasets (The real hero 🔥)\n- Special mention to Google TPU Research Program (https://sites.research.google/trc/) which helped us in training bigger models in the last few days of the competition by supporting us with V3 TPUs on GCP.",
    "1761383": "Thanks for your sharing. Can you share your source code? I'm really interested in stacking part",
    "1761429": "Good job, I am also interested in PostProcessing part.",
    "1763110": "Hi, sorry for late reply. Taking some time off from kaggle after intense competition. We'll make the complete source code public in few days.",
    "1763118": "Here is the code for feature engineering part. We trained the LGBM and XGBoost model on these features. Hyperparameters tuned with optuna to optimise competition metric.  This gave us ~ 0.01-0.015 boost in public & private lb.\n\n```\nKNN = 100\nnn_cols = [f'nn-{x}' for x in range(KNN)]\ndist_cols = [f'dist-{x}' for x in range(KNN)]\nreciprocals = 1/(np.arange(1,6))\n\ndef generate_preds(row,threshold=0.5):\n    \n    nns = list(zip(row[dist_cols],row[nn_cols].values))\n    nns.append((threshold,-1))\n    preds = []\n    for dist,pred in sorted(nns):\n        if pred not in preds:\n            preds.append(pred)\n        if len(preds)==5:\n            break\n    if len(preds)<5:\n        preds = preds+[-1,-1,-1,-1,-1]\n        preds = preds[:5]\n    return np.array(preds)\n    \n    \ndef map_per_sample(row):\n    return ((row.target==row['preds'])*reciprocals).sum()\n\ndef get_sample_neighbours(data,MODE):\n    \n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    allowed_targets = set(train_targets)\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(train_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = train_targets[idxs]\n    ids = train_ids[idxs]\n\n    val_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    val_predictions['nn_ids'] = train_ids[idxs].reshape(-1)\n    val_predictions['nn_distances'] = distances.reshape(-1)\n    val_predictions['nn_rank'] = val_predictions.index//KNN\n    val_predictions['image'] = val_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    if MODE=='TRAIN':\n        val_predictions['target'] = val_predictions['nn_rank'].apply(lambda x: val_targets[x])\n    else:\n        val_predictions['target'] = -1\n    val_predictions['nn_rank'] = val_predictions.index%KNN\n    return val_predictions\n\ndef get_centroid_neighbours(data):\n    \n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    train_embeddings_df = pd.DataFrame(train_embeddings)\n    train_embeddings_df['target'] = train_targets\n    centroid_embeddings = train_embeddings_df.groupby('target').mean()\n    centroid_targets = centroid_embeddings.index.values\n    centroid_embeddings = centroid_embeddings.values\n\n    from sklearn.neighbors import NearestNeighbors\n    neigh = NearestNeighbors(n_neighbors = KNN,metric='cosine')\n    neigh.fit(centroid_embeddings)\n    distances,idxs = neigh.kneighbors(val_embeddings, KNN, return_distance=True)\n    targets = centroid_targets[idxs]\n\n    centroid_predictions = pd.DataFrame(targets.reshape(-1,1),columns=['nn_target'])\n    centroid_predictions['nn_distances'] = distances.reshape(-1)\n    centroid_predictions['nn_rank'] = centroid_predictions.index//KNN\n    centroid_predictions['image'] = centroid_predictions['nn_rank'].apply(lambda x: val_ids[x])\n    centroid_predictions['nn_rank'] = centroid_predictions.index%KNN\n    centroid_predictions = centroid_predictions.set_index(['image','nn_target']).nn_distances.to_dict()\n    \n    return centroid_predictions\n\ndef get_target_counts(data):\n    train_embeddings,train_targets,train_ids,val_embeddings,val_targets,val_ids = data\n    target_counts = pd.Series(train_targets).value_counts().to_dict()\n    return target_counts\n\ndef engineer_features(train_pattern,test_pattern,fold=0,MODE='TEST'):\n    \n    print(\"Loading Data\")\n    val_species_ids,val_species,test_species_ids,test_species = load_species(fold=fold)\n    val_species = {x:y for x,y in zip(val_species_ids,val_species)}\n    test_species = {x:y for x,y in zip(test_species_ids,test_species)}\n    \n    if MODE=='TEST':\n        data = load_test_embeds(train_pattern,test_pattern,fold=fold)\n        species = test_species\n    else:\n        data = load_embeds(train_pattern,test_pattern,fold=fold)\n        species = val_species\n        \n    spec_df = pd.DataFrame(species).T.reset_index()\n    spec_df.columns = ['image']+[f'spec_{i}' for i in range(0,26)]\n    \n    print(\"Getting Sample Neighbours\")\n    val_predictions = get_sample_neighbours(data,MODE)\n    \n    print(\"Getting Centroid Neighbours\")\n    centroid_predictions = get_centroid_neighbours(data)\n    \n    print(\"Getting Target Counts\")\n    target_counts = get_target_counts(data)\n    \n    print(len(species),len(val_predictions))\n    \n    train_df = val_predictions.groupby(['image','nn_target']).head(3)\n    train_df['nn_rank_target'] = train_df.groupby(['image','nn_target']).nn_distances.rank(method='first')\n    train_df = pd.pivot_table(train_df, values='nn_distances', index=['image','target', 'nn_target'],\n                        columns=['nn_rank_target'], aggfunc=np.sum).reset_index().fillna(1)\n    train_df = train_df.rename(columns={1:'nn1',2:'nn2',3:'nn3','nn_target':'individual'}).reset_index(drop=True)\n    train_df['centroid'] = train_df.apply(lambda row:centroid_predictions[(row.image,row.individual)]\n                                              if (row.image,row.individual) in centroid_predictions else 1\n                                          ,axis=1)\n    train_df['target_counts'] = train_df.apply(lambda row:target_counts[row.individual],axis=1)\n    train_df['individual_species'] = train_df.individual.map(species_individual_mappings)\n    train_df['species_conf'] = train_df.apply(lambda row:species[row.image][row.individual_species],axis=1)\n#     train_df[[f'species_{i}' for i in range(26)]] = train_df.apply(lambda row: species[row.image],axis=1)\n    return train_df,spec_df\n\n```",
    "1765975": "Great job on the competition. I had heard that not very much computing resources went into your training pipeline. Is this true? How many iterations did your team go through to get such good results?",
    "1766206": "Thanks for your reply！",
    "1767557": "No, It's not that. There is some miscommunication. We didn't have any personal hardware (GPUs). We used TPUs to train our models, as some other top teams did. But, TPUs are good compute resources, available on Kaggle and Colab.  Almost all of my kaggle finishes are with the models trained on TPUs/GPUs provided by Kaggle and Colab, except few in which I teamed up with someone who has good hardware. This is what Tanul wanted to convey in his post. \nWe have been competing actively in this competition for last 1.5 months. So, our results has been results of many iterations. At first, our postprocessing was fixed (the same which I released in first week of start of the competition), and we tuned the architecture/training strategy for both DOLG & Non-DOLG models. Then, a lot of time went in understanding which preprocessing works better and designing data recipe, experimenting with psuedolabelling. Last week went in mostly designing & optimising postprocessing.",
    "1767615": "ks2019  thanks for your sharing. Can you share your code for training LGB/XGB. I'm not familiar with stacking and want to learn.",
    "1852856": "I wanted to ask about the your 2nd stage model, I bit confused what was the target used for it\nLike suppose you have 1 image and its 100 possible matches (founded via knn), then you generate features for that image using the matches, now here what target do you use, and is it the knn distance or the cosine similarity?"
  },
  "source": "meta"
}