Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology (LMRL Workshop, NeurIPS 2021)

Last update: Dec 24, 2022

Overview

Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology

Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology, LMRL Workshop, NeurIPS 2021. [Workshop] [arXiv]
Richard. J. Chen, Rahul G. Krishnan

@article{chen2022self,
  title={Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology},
  author={Chen, Richard J and Krishnan, Rahul G},
  journal={Learning Meaningful Representations of Life, NeurIPS 2021},
  year={2021}
}

Summary / Main Findings:

In head-to-head comparison of SimCLR versus DINO, DINO learns more effective pretrained representations for histopathology - likely due to 1) not needing negative samples (histopathology has lots of potential class imbalance), 2) capturing better inductive biases about the part-whole hierarchies of how cells are spatially organized in tissue.
ImageNet features do lag behind SSL methods (in terms of data-efficiency), but are better than you think on patch/slide-level tasks. Transfer learning with ImageNet features (from a truncated ResNet-50 after 3rd residual block) gives very decent performance using the CLAM package.
SSL may help mitigate domain shift from site-specific H&E stainining protocols. With vanilla data augmentations, global structure of morphological subtypes (within each class) are more well-preserved than ImageNet features via 2D UMAP scatter plots.
Self-supervised ViTs are able to localize cell location quite well w/o any supervision. Our results show that ViTs are able to localize visual concepts in histopathology in introspecting the attention heads.

Updates

Stay tuned for more updates :).

TBA: Pretrained SimCLR and DINO models on TCGA-Lung (Larger working paper, in submission).
TBA: Pretrained SimCLR and DINO models on TCGA-PanCancer (Larger working paper, in submission).
TBA: PEP8-compliance (cleaning and organizing code).
03/04/2022: Reproducible and largely-working codebase that I'm satisfied with and have heavily tested.

Pre-Reqs

We use Git LFS to version-control large files in this repository (e.g. - images, embeddings, checkpoints). After installing, to pull these large files, please run:

git lfs pull

Pretrained Models

SIMCLR and DINO models were trained for 100 epochs using their vanilla training recipes in their respective papers. These models were developed on 2,055,742 patches (256 x 256 resolution at 20X magnification) extracted from diagnostic slides in the TCGA-BRCA dataset, and evaluated via K-NN on patch-level datasets in histopathology.

Note: Results should be taken-in w.r.t. to the size of dataset and duraration of training epochs. Ideally, longer training with larger batch sizes would demonstrate larger gains in SSL performance.

Arch	SSL Method	Dataset	Epochs	Dim	K-NN	Download
ResNet-50	Transfer	ImageNet	N/A	1024	0.935	N/A
ResNet-50	SimCLR	TCGA-BRCA	100	2048	0.938	Backbone
ViT-S/16	DINO	TCGA-BRCA	100	384	0.941	Backbone

Data Download + Data Preprocessing

CRC-100K: Train and test data can be downloaded as is via this Zenodo link.
BreastPathQ: Train and test data can be downloaded from the official Grand Challenge link.
TCGA-BRCA: To download diagnostic WSIs (formatted as .svs files) and associated clinical metadata, please refer to the NIH Genomic Data Commons Data Portal and the cBioPortal. WSIs for each cancer type can be downloaded using the GDC Data Transfer Tool.

For CRC-100K and BreastPathQ, pre-extracted embeddings are already available and processed in ./embeddings_patch_library. See patch_extraction_utils.py on how these patch datasets were processed.

Additional Datasets + Custom Implementation: This codebase is flexible for feature extraction on a variety of different patch datasets. To extend this work, simply modify patch_extraction_utils.py with a custom Dataset Loader for your dataset. As an example, we include BCSS (results not yet updated in this work).

BCSS (v1): You can download the BCSS dataset from the official Grand Challenge link. For this dataset, we manually developed the train and test dataset splits and labels using majority-voting. Reproducibility for the raw BCSS dataset may be not exact, but we include the pre-extracted embeddings of this dataset in ./embeddings_patch_library (denoted as version 1).

Evaluation: K-NN Patch-Level Classification on CRC-100K + BreastPathQ

Run the notebook patch_extraction.ipynb, followed by patch_evaluation.ipynb. The evaluation notebook should run "out-of-the-box" with Git LFS.

Evaluation: Slide-Level Classification on TCGA-BRCA (IDC versus ILC)

Install the CLAM Package, followed by using the 10-fold cross-validation splits made available in ./slide_evaluation/10foldcv_subtype/tcga_brca. Tensorboard train + validation logs can visualized via:

tensorboard --logdir ./slide_evaluation/results/

Visualization: Creating UMAPs

Install umap-learn (can be tricky to install if you have incompatible dependencies), followed by using the following code snippet in patch_extraction_utils.py, and is used in patch_extraction.ipynb to create Figure 4.

Visualization: Attention Maps

Attention visualizations (reproducing Figure 3) can be performed via walking through the following notebook at attention_visualization_256.ipynb.

Issues

Please open new threads or report issues directly (for urgent blockers) to [email protected].
Immediate response to minor issues may not be available.

Acknowledgements, License & Usage

Part of this work was performed while at Microsoft Research. We thank the BioML group at Microsoft Research New England for their insightful feedback.
This work is still under submission in a formal proceeding. Still, if you found our work useful in your research, please consider citing our paper at:

@article{chen2022self,
  title={Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology},
  author={Chen, Richard J and Krishnan, Rahul G},
  journal={Learning Meaningful Representations of Life, NeurIPS 2021},
  year={2021}
}

Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology (LMRL Workshop, NeurIPS 2021)

Related tags

Overview

Self-Supervised Vision Transformers Learn Visual Concepts in Histopathology

Updates

Pre-Reqs

Pretrained Models

Data Download + Data Preprocessing

Evaluation: K-NN Patch-Level Classification on CRC-100K + BreastPathQ

Evaluation: Slide-Level Classification on TCGA-BRCA (IDC versus ILC)

Visualization: Creating UMAPs

Visualization: Attention Maps

Issues

Acknowledgements, License & Usage

Owner

Richard Chen

Christmas face app for Decathlon xmas coding party!

Pgn2tex - Scripts to convert pgn files to latex document. Useful to build books or pdf from pgn studies

A simple implementation of Kalman filter in Multi Object Tracking

Experiments with differentiable stacks and queues in PyTorch

HistoSeg : Quick attention with multi-loss function for multi-structure segmentation in digital histology images

Official PyTorch implementation of paper: Standardized Max Logits: A Simple yet Effective Approach for Identifying Unexpected Road Obstacles in Urban-Scene Segmentation (ICCV 2021 Oral Presentation)

D²Conv3D: Dynamic Dilated Convolutions for Object Segmentation in Videos

official Pytorch implementation of ICCV 2021 paper FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting.

Source code for EquiDock: Independent SE(3)-Equivariant Models for End-to-End Rigid Protein Docking (ICLR 2022)

Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework

StableSims is an open-source project aimed at simulating MakerDAO's Dai stablecoin system

Tensorflow implementation of the paper "HumanGPS: Geodesic PreServing Feature for Dense Human Correspondences", CVPR 2021.

Code for the paper "Learning-Augmented Algorithms for Online Steiner Tree"

Notes, programming assignments and quizzes from all courses within the Coursera Deep Learning specialization offered by deeplearning.ai

An open-source Deep Learning Engine for Healthcare that aims to treat & prevent major diseases

Curvlearn, a Tensorflow based non-Euclidean deep learning framework.

A Real-World Benchmark for Reinforcement Learning based Recommender System

Sub-tomogram-Detection - Deep learning based model for Cyro ET Sub-tomogram-Detection

Source code for "Taming Visually Guided Sound Generation" (Oral at the BMVC 2021)

This is the reference implementation for "Coresets via Bilevel Optimization for Continual Learning and Streaming"